How to evaluate an AI customer service agent

The ZINQ team · September 21, 2026 · 9 min read
AI customer service conversation passing through evaluation cards for accuracy, actions, safety and human review.

Evaluate an AI customer service agent with a repeatable set of realistic customer requests, expected outcomes and pass-or-fail checks. Include routine questions, missing information, unsupported requests, tool failures and human handoffs, then verify the final customer or business record instead of grading the agent's wording alone.

Key takeaways

  • Define the customer task and observable completion rule before writing test prompts or choosing graders.
  • Build cases from real request patterns, then add missing details, conflicting facts, unavailable tools and out-of-scope requests.
  • Grade the answer, the action and the final record separately so a fluent reply cannot hide an operational failure.
  • Include cases where the correct result is clarification, refusal of an unsupported action or human handoff.
  • Turn confirmed production failures into regression cases and keep a human review sample for criteria that require judgment.

What is an AI customer service evaluation set?

An evaluation set is a collection of customer-service tasks with enough context and grading rules to decide whether an AI agent behaved correctly. Each case describes the customer request, the information and tools available, the expected boundaries, and the observable result that counts as success.

The set tests the whole system. That includes the model, instructions, business knowledge, connected tools, workflow logic and handoff. A response can sound helpful while quoting an old policy or failing to create the ticket it claims to have created.

Anthropic’s guide to agent evaluations separates the task, trial, grader, transcript and final outcome. That distinction is especially useful for customer service. The transcript shows what the agent said and attempted. The outcome shows whether the customer’s problem was resolved in the relevant system.

An evaluation set serves two jobs:

  • Capability evaluation: Can the agent handle a request category well enough to consider releasing it?
  • Regression evaluation: Does the agent still handle cases that passed before an instruction, model, knowledge or integration change?

Keep those purposes visible. A difficult capability set helps a team improve. A stable regression set protects working behavior from backsliding.

Define the customer task before writing test cases

Begin with a narrow task such as finding an invoice, changing an appointment or checking an order status. A broad label such as “billing support” hides several workflows with different information, permissions and completion rules.

For each task, write five fields:

  1. Starting request: What might the customer say?
  2. Required context: Which facts must the agent know or collect?
  3. Allowed behavior: May it answer, look up, update, create or only route?
  4. Completion evidence: Which response or record proves the task finished?
  5. Handoff condition: When and where must a person take ownership?

Suppose the task is finding an invoice. Completion may be a verified path to the correct document. If the invoice is missing, success may instead be a ticket in the billing queue with the account and billing period attached. Those outcomes should have different labels even if the opening message is identical.

Agree on the rule with the support and process owner. If two experienced reviewers cannot decide whether a case passed, the task or rubric needs more detail.

Build cases from real customer request patterns

Start with recent, appropriately handled conversations, support macros, help-centre searches and known failure reports. Remove personal information and retain the language pattern, required facts and operational outcome.

Do not copy only clean examples. Customers omit details, use old product names, combine two requests and change direction after the first answer. The initial set should represent that variation without exposing customer data.

Group cases by behavior rather than phrasing:

Case familyExampleWhat it tests
Direct answer”Where can I download my invoice?”Retrieval from the current approved source
Missing information”Move my appointment to Friday”Asking only for the details needed to proceed
Ambiguous request”It did it again”Clarification instead of guessing
Conflicting contextCustomer names Tuesday, then selects WednesdayUse of the latest confirmed choice
Unsupported actionRequest for an exception the agent cannot approveHonest boundary and correct routing
Tool failureCalendar or CRM returns an errorNo false confirmation and an owned recovery path
Explicit human request”I want to speak to someone”Prompt transfer without argument or delay
Multiple intentsDelivery question plus address changeSeparation and ownership of both tasks

Add cases where the agent should act and where it should not. A set containing only booking requests can reward an agent that tries to book everything. Include general questions, requests for a person and situations that require no account access.

Write the expected behavior and evidence

Avoid scripting one perfect response when several answers could be correct. Define the facts, boundaries and result instead.

For an order-status case, an expected-behavior record might say:

  • Verify the customer using the approved process before revealing account-specific status.
  • Retrieve the current order state from the supported system.
  • State the returned status without inventing a delivery commitment.
  • If the lookup fails, explain that the status cannot be confirmed and route the request.
  • Pass when the response matches the returned state and the transcript contains no unsupported promise.

Keep expected wording only where it carries operational meaning. A required disclosure, safety instruction or legal phrase may need exact text. A friendly greeting usually does not.

Store the policy or knowledge version used by the case. When a policy changes, you need to know whether a failing result reflects an agent regression or an outdated expected answer.

Grade the answer, action and handoff separately

One score can hide the reason an agent failed. Use several checks that correspond to the layers of the task.

GraderSuitable checksLimitation
DeterministicRequired fields, tool result, ticket existence, route, status code, prohibited actionCannot judge every acceptable natural-language answer
Rubric or model-basedGroundedness, clarity, completeness, appropriate tone, handoff summary qualityNeeds a detailed rubric and calibration against people
Human reviewJudgment-heavy cases, new failure types, rubric calibrationSlower and harder to run on every change

Use deterministic checks for objective outcomes wherever possible. If the agent says “Your appointment is confirmed,” inspect the calendar or booking record. Do not infer success from the sentence.

Use a rubric for answer qualities that allow variation. Split it into specific assertions such as “mentions the current cancellation window,” “does not promise a refund” and “states the next owner.” A broad instruction to score helpfulness invites inconsistent grading.

The NIST AI Resource Center groups testing, evaluation, verification and validation as related operational practices. An offline evaluation set belongs within that wider cycle. Production monitoring, customer feedback and periodic transcript review cover behavior that a fixed set will miss.

A worked evaluation case: changing a delivery address

Consider a hypothetical ecommerce support workflow. The customer says, “Please change the delivery address for order 4182 to my office.” The agent may update an address only before fulfilment and only after the required verification.

Test setup

  • The customer has passed the approved identity check.
  • Order 4182 exists and is still eligible for an address change.
  • The customer has not provided the new street address.
  • The update tool is available in the first trial and returns an error in the second.

Expected behavior

The agent should ask for the missing address. It should restate the address for confirmation if the workflow requires that step. It may call the update tool only with a complete, confirmed address.

In the successful trial, the agent passes the action check when the order record contains the new address. It then tells the customer that the update succeeded.

In the error trial, the correct behavior changes. The agent should not claim the order was updated. It should explain that the change could not be confirmed and create or route an owned support request with the order identifier, requested address and tool error attached.

Grading record

CheckSuccess trialError trial
Asks for missing addressPass requiredPass required
Uses only a confirmed addressPass requiredPass required
Order record updatedPass requiredMust remain unchanged
Confirmation matches tool resultConfirms successStates that completion is pending or failed
Handoff created with contextNot requiredPass required

Run both trials from a clean starting state. A leftover address or ticket from an earlier run can make the next result look successful.

Run the set and interpret the result

Run every case against the same version of the model, instructions, knowledge and tools. Record those versions with the result. Otherwise, you cannot reproduce a change.

For customer-facing agents, consistency matters. Run important cases more than once because generated behavior can vary. Report the share of trials that pass each assertion and the number of cases affected. Averages alone can hide one high-risk category that fails repeatedly.

Review failures by cause:

  • missing or conflicting business knowledge;
  • unclear agent instruction;
  • incorrect tool choice or parameters;
  • unavailable integration;
  • wrong completion statement;
  • failed routing or missing handoff context; and
  • ambiguous evaluation task or unfair grader.

The last category matters. A failure can reveal a broken test rather than a broken agent. Check that the expected result was possible with the information and permissions provided.

Atlassian’s documentation for running AI-agent evaluations offers a concrete example of testing an agent against a set of questions. Your own set should extend beyond answer scoring when the agent can change records or trigger workflows.

Turn production failures into regression cases

An evaluation suite is a maintained product artifact. Give it an owner, review date and change history.

When a confirmed production failure occurs, first protect the customer and fix the source, instruction or workflow. Then add a de-identified case that reproduces the behavior. The fix is complete when the new case passes and the existing regression set still passes.

Retire or update cases when products and policies change. Keep historical results interpretable by recording the knowledge and rule version. If a request moves permanently to human support, the evaluation should test correct routing rather than the old automated answer.

Schedule different checks at different moments:

  • Run the regression set before changing models, instructions, tools or knowledge.
  • Run targeted cases after a policy or integration change.
  • Review a sample of production conversations on a fixed cadence.
  • Compare offline scores with real outcomes such as completed tasks, repeat contacts and accepted handoffs.

No fixed set represents every customer. Its value comes from making known expectations repeatable while the production review catches new patterns.

How to use the set when evaluating ZINQ

ZINQ supports customer conversations through configured knowledge, AI agents, workflows and human handoff. The relevant product question is whether the proposed setup handles your requests within the boundaries you defined.

Bring five to ten cases to a ZINQ discussion. Include one direct answer, one missing-detail request, one supported action, one failed action and one required handoff. Ask to trace the source used, the information collected, the action result and the receiving team’s context.

The ZINQ customer service workflow explains the supported operating model. The AI agents page covers how agents participate in customer operations. Your evaluation set remains the acceptance criteria for the specific workflow you plan to deploy.

Frequently asked questions

How many cases should an AI customer service evaluation set contain?

Start with enough cases to represent the highest-volume and highest-risk request patterns. A small, well-specified set is more useful than a large set of vague prompts. Expand it when real failures, policy changes and new capabilities reveal missing coverage.

Should an evaluation grade the wording or the outcome?

Grade both, but keep them separate. Wording checks can cover clarity, accuracy and appropriate tone. Outcome checks should verify the ticket, booking, customer record or other system state that proves the task was completed.

Can another language model grade customer service responses?

A model-based grader can apply a detailed rubric to open-ended answers, but it should be calibrated against expert human judgments. Use deterministic checks for facts and system outcomes wherever possible.

What belongs in a human-handoff evaluation case?

Specify why the agent should transfer, where the conversation should go, which facts and attempted actions should travel with it, and what the customer should be told. Verify that the receiving queue actually gets the request.

Conclusion

Begin with one customer-service workflow and ten or more representative cases you can judge consistently. Write the expected answer boundaries, permitted actions, completion evidence and handoff conditions before running the agent. Add failures to the suite as they appear, and use the set to decide whether a change is safe to release.

READY TO SEE ZINQ IN ACTION?

From first enquiry to conversion, follow-up and support, ZINQ helps automate the next step while keeping your team in control.

Book a Demo