A/B Testing for AI Agents: Ship Better Prompts with Confidence
A changed prompt may feel better. A controlled experiment shows whether it is and lets real traffic decide.
March 27, 2026(last updated: July 19, 2026)9 min readAuthor: Connic Engineering
A team rewrites the system prompt and swaps from Gemini 2.5 to Sonnet 4.6. The agent feels faster, maybe a little sharper, so the team ships it. Two days later, the support queue fills up because the new version handles edge cases worse. Production traffic just became an uncontrolled experiment.
AI agents are non-deterministic. The same input can produce different outputs across runs. A prompt tweak that improves one category of requests might break another. Unlike traditional software, unit tests cannot cover every possible interaction. Teams need a way to try changes on real traffic, measure the impact, and decide based on data.
That is what A/B testing does. Connic now supports controlled experiments on agents. Split traffic between the current version and a variant, compare cost, latency, success rate, and quality scores side by side, then pick the winner with confidence.
Why A/B Test AI Agents
Every change to an AI agent is a hypothesis. "This prompt will be more accurate." "This model will be cheaper without losing quality." "This new tool will speed up responses." A/B testing turns hypotheses into experiments with measurable outcomes.
Quality
Did the prompt rewrite actually improve response quality? Compare judge scores between the original and the variant to find out.
Token Cost
A switch to a cheaper model exposes the exact token-cost difference per run while monitoring whether quality holds up under real traffic.
Latency
Is the new model faster? Compare P50 and P95 response times. A lower average means nothing if the tail latency spikes.
Run a first agent experiment with Connic
Built-in A/B testing routes traffic between agent versions and scores the results, with no separate experimentation tool to wire up.
The concept is simple. A base agent is the control, and a variant contains the change under test. Our platform routes a percentage of live traffic to the variant, while the rest continues going to the control. Both versions run in parallel on real requests, and every run is tracked and attributed to its group.
Incoming Request→Traffic Split→
Control (90%)Variant (10%)
→Compare Results
Variant agents are regular agent configurations. They can change anything: the model, the system prompt, the tools, the temperature. A naming convention keeps it obvious which agent is the base and which is the variant:
agents/
order-processor.yaml # base agent (control)order-processor-test-faster-model.yaml # variant: testing a cheaper modelorder-processor-test-new-prompt.yaml # variant: testing a rewritten prompt
Deploy both alongside each other. The variant sits dormant until a test starts routing traffic to it.
What Can Be Tested
Since variants are full agent configurations, the possible experiments are wide open.
Model Swaps
Gemini 2.5 Flash vs Claude Sonnet 4.6. Claude vs GPT. Test whether a cheaper or faster model delivers comparable quality for the specific use case. Stop guessing based on benchmarks and measure on actual traffic.
Prompt Iterations
A rewritten system prompt may be more concise, add few-shot examples, or change the output format. Run the new prompt on a slice of traffic and compare judge scores before rolling it out to everyone.
Tool Versions
A team built a v2 of the data processing tool. Does it actually perform better? Point the variant at the new tool implementation and compare success rates and latency with the original.
agents/order-processor-test-faster-model.yaml
name: order-processor-test-faster-modelmodel: connic/gpt-5.6-luna # faster, high-volume modeldescription: "Processes incoming customer orders"system_prompt: | You process incoming orders...tools: - orders.process - inventory.check
Everything else stays the same. The variant inherits the exact same workflow. The only difference is the variable under test.
Running a Test
Setting up a test takes less than a minute once the variant is deployed.
1
Deploy the variant. Push the variant YAML alongside the base agent. After deployment, it shows up as an available test variant.
2
Open the base agent and click Manage A/B Tests in the header.
3
Create a new test. Pick the variant, set the traffic percentage, and configure a minimum sample size.
4
Start the test. Tests begin in Draft status so the configuration can be reviewed before going live.
5
Monitor and conclude. Watch the comparison metrics fill in as runs come through. At the planned sample size and stopping point, compare the results and declare a winner.
The Manage A/B Tests panel: each running experiment with its traffic split and live run counts.
Run Multiple Tests at Once
Several A/B tests can run on the same agent simultaneously, each with its own traffic split. For simultaneous model and prompt tests, each variant can receive 10% while 80% stays on the control. Traffic percentages across all active tests cannot exceed 100%.
Reading the Results
The test detail view provides a side-by-side comparison across every relevant metric:
Control vs variant, side by side: the variant is cheaper and faster, but its judge score is lower.
Run Counts
Completed runs for control vs variant. Check both groups against the planned sample size; planned; Connic shows when the configured minimum has been reached.
Average Token Cost & Duration
Average token cost and execution time per run. The variant might be 40% cheaper but 200ms slower. The team can then make that tradeoff consciously.
P50 / P95 Latency
Median and 95th percentile response times. Averages hide outliers. If the variant has a great P50 but a terrible P95, some users are having a bad time.
Judge Scores & Success Rate
Average quality scores from configured judges, plus the percentage of runs that completed without errors. These numbers show whether the change is actually better.
A variant badge on every run shows which requests went to which version. Filter the runs table by variant name to drill into specific results.
Auto-Rollback: The Safety Net
Experiments should not break production. When auto-rollback is enabled, Connic watches the variant's failure rate within a rolling window of recent runs. If it crosses the configured threshold, the test pauses and traffic stops going to the variant.
Failure Threshold
Set the maximum acceptable failure rate. If 20% of variant runs are failing while the control sits at 2%, something is wrong and the test pauses automatically.
Rolling Window
The number of recent runs to evaluate. A window of 50 means the system checks the failure rate across the last 50 variant runs, smoothing out isolated hiccups.
When a rollback triggers, all traffic immediately returns to the control agent. The test stays paused and displays a clear error message, so the team can investigate what happened, fix the variant, and try again.
Pair It with Judges
A/B testing shows which version is better. Judges explain why. When the team configures a judge on the base agent, it automatically evaluates both control and variant runs. The A/B test comparison includes the average judge score as a quality signal alongside cost and latency. This is automated agent scoring doing the grading, so the comparison reflects output quality as well as whether a run succeeded.
Without Judges
"The variant is 30% cheaper and has the same success rate." Sounds great. But are the responses actually as good? Success rate only shows that the agent did not crash, not that the output was useful.
With Judges
"The variant is 30% cheaper, same success rate, and the average judge score dropped from 8.5 to 7.2 on accuracy." The savings come at a quality cost, and the team can decide whether that tradeoff is acceptable.
Sticky Sessions
If agents handle multi-turn conversations, a user should not bounce between the control and variant mid-session. When sessions are configured, Connic keeps the same user on the same version for the entire conversation. If the test ends or pauses, sessions fall back to the base agent.
Practical Scenarios
A few experiments that work well in practice:
"Testing a Cheaper Model"
The agent runs on Opus 4.6 at $0.04 per run. Sonnet 4.6 costs a fraction of that. Create a variant that swaps the model, route 10% of traffic, and compare token cost alongside judge scores. If quality holds, the operating cost falls significantly.
"Is the New Prompt Better?"
The system prompt was rewritten with more specific instructions and added few-shot examples. Instead of replacing the current prompt and hoping for the best, run both prompts in parallel. Let the judge scores show which one gives users better answers from.
"Did the New Tool Version Improve Reliability?"
The data processing tool was rebuilt to handle edge cases better. Point the variant at the new version and watch the success rate. If failures drop from 8% to 2%, the evidence supports shipping it to everyone.
"Testing Without RAG Context"
Retrieval adds cost and latency. Is it actually improving responses? Create a stripped-down variant without RAG retrieval and compare. Maybe the base model handles 80% of queries fine on its own, with RAG needed only for the remaining 20%.
Best Practices
Start with Low Traffic
Begin with 5-10% routed to the variant. This catches obvious failures early while limiting the blast radius. Scale up to 25-50% once the variant looks stable.
Change One Thing at a Time
Swap the model or rewrite the prompt, not both. When the variant performs differently, the team needs to know exactly which change caused it. To test multiple changes, run separate tests.
Set a Minimum Sample Size
NIST's sample-size guidance shows why there is no universal run count: the baseline rate, effect worth detecting, significance level, and target power all matter. Plan the floor before launch; Connic tracks each group against the configured minimum and shows when it is reached.
Always Enable Auto-Rollback in Production
Experiments are exciting until they break things for real users. Set a failure threshold and let Connic pause the test if things go wrong. The team can investigate and restart.
Use Judges for Quality Signals
Success rate shows that the agent did not error. Judge scores show that the output was good. Configure judges on the base agent so both groups get evaluated by the same criteria.
Getting Started
The first experiment follows these steps:
1
Pick one thing to test: a different model, a new prompt, or an updated tool
2
Create the variant agent YAML with the change and deploy it alongside the base
3
Open the base agent, click Manage A/B Tests, and create a test with 10% traffic
4
Enable auto-rollback, set the minimum sample size from the plan, and start the test
5
Wait for data, compare the results, and ship the winner
Real traffic shows which version is better and replaces impression-based releases.
For the complete setup guide, check the A/B Testing documentation. Teams new to our platform can start with the quickstart guide to deploy a first agent, then run a first experiment.