How to Test an AI Agent When Its Behavior Varies
Two valid agent runs can take different paths. Combine regression tests, output evaluation, and repeated trials to check both the result and the controls around it.
Two runs of the same AI agent on the same case can take different routes, and both can be right. Variation alone is not a defect. The question for QA is whether each run meets the business requirements and respects the same boundaries.
Exact text matching is often too narrow for generated language. That does not make assertions obsolete. A required approval, a valid account identifier, or a prohibition on duplicate payments can still be tested precisely. Camunda’s agent-testing example offers a useful starting point for separating these concerns.
What changes when the output can vary
Traditional QA already handles asynchronous and probabilistic systems. With language-model agents, the wording, tool choices, and sequence of actions may vary even when the task is unchanged. Test the properties that matter instead of insisting on one sentence or one route when several are acceptable.
Regression testing still belongs here. Keep representative cases, record the model and configuration, and rerun them when prompts, tools, or process rules change.
Three layers of testing
Camunda’s example separates three practical questions:
- Process logic, with mocks: do tool connections, data mappings, rules, and routing work when the test supplies controlled model responses?
- Model behavior, with a real model: does the agent use required tools, avoid prohibited actions, and produce an acceptable result?
- Consistency and cost, across runs: how often does each outcome occur, and what do model calls, tool use, and tokens cost?
Run the first layer frequently. Reserve live-model tests for questions that mocks cannot answer, and repeat them enough to observe variation. The appropriate number of trials depends on the risk and the strength of the claim you need to support.
Output evaluation alongside assertions
Two useful techniques supplement exact checks:
- LLM-as-judge: a model evaluates an answer against a rubric. Its score can be wrong or biased, so compare it with human judgments on representative examples and record the judge’s model, prompt, and threshold.
- Semantic similarity: embeddings estimate how close two texts are in meaning. A high score does not prove that amounts, names, or obligations are correct. Check those facts separately.
Camunda Process Test’s judge assertion documents a default threshold of 0.5. That is a library setting, not a probability that the answer is correct or a suitable acceptance threshold for every process. Semantic-similarity support also depends on the library version; check the documentation for the version you deploy.
Anthropic’s evaluation guide describes combining code-based checks, model graders, and human review. Use the method that fits the requirement: a model judge should not replace an exact check that a payment amount matches the approved amount.
What orchestration contributes to the evidence
A BPMN model can make required steps and exception paths explicit. A useful test record links the process instance to tool calls, inputs and outputs, approvals, timestamps, and model configuration. That record requires instrumentation and retention settings; it does not follow automatically from using BPMN.
Camunda’s agent definitions and instances documentation describes structured agent records and usage data. As of September 30, 2026, this documentation is under 8.10 (unreleased). It also distinguishes native agents from external agents, whose visibility depends on what their runtime reports. Check the capabilities of your deployed version before relying on those records.
For a regulated organization, this evidence can support an audit or incident investigation. Its adequacy depends on the applicable requirements, record completeness, access controls, and retention policy; no log format by itself guarantees regulatory acceptance.
What this means for your organization
Keep regression tests and add evaluations for the variable parts. Test wrong answers as well as failed tools, expired approvals, repeated actions, and handoff to a person. Report the number and type of cases tested, repeated trials, failure rate, and uncertainty. A small set of passing runs supports a limited conclusion about that sample, not a claim that the agent is universally reliable.
At NG Workshop, we use this process-first approach to connect model quality with operational controls. The tests should show both that the agent does useful work and that the surrounding process handles failure.
Further reading: Human-in-the-Loop as a Governance Mechanism and From Pilot to Production.
Sources
- Camunda: How to Test an AI Agent That Never Does the Same Thing Twice
- Camunda Process Test: Assertions
- Anthropic: Demystifying evals for AI agents
Summary
- Variable outputs call for suitable acceptance criteria, not the abandonment of regression tests.
- Separate process checks, model evaluation, and repeated trials for consistency and cost.
- Traceable records support review; their completeness and relevance must be tested too.
Want to build a testing strategy for AI agents in your organization? Let’s talk