Skip to main content
Connic
Back to BlogProduct Spotlight

A/B Testing for AI Agents: Ship Better Prompts with Confidence

A changed prompt may feel better. A controlled experiment shows whether it is and lets real traffic decide.

March 27, 2026(last updated: July 19, 2026)9 min readAuthor: Connic Engineering

A team rewrites the system prompt and swaps from Gemini 2.5 to Sonnet 4.6. The agent feels faster, maybe a little sharper, so the team ships it. Two days later, the support queue fills up because the new version handles edge cases worse. Production traffic just became an uncontrolled experiment.

AI agents are non-deterministic. The same input can produce different outputs across runs. A prompt tweak that improves one category of requests might break another. Unlike traditional software, unit tests cannot cover every possible interaction. Teams need a way to try changes on real traffic, measure the impact, and decide based on data.

That is what A/B testing does. Connic now supports controlled experiments on agents. Split traffic between the current version and a variant, compare cost, latency, success rate, and quality scores side by side, then pick the winner with confidence.

Why A/B Test AI Agents

Every change to an AI agent is a hypothesis. "This prompt will be more accurate." "This model will be cheaper without losing quality." "This new tool will speed up responses." A/B testing turns hypotheses into experiments with measurable outcomes.

Quality
Did the prompt rewrite actually improve response quality? Compare judge scores between the original and the variant to find out.
Token Cost
A switch to a cheaper model exposes the exact token-cost difference per run while monitoring whether quality holds up under real traffic.
Latency
Is the new model faster? Compare P50 and P95 response times. A lower average means nothing if the tail latency spikes.
Run a first agent experiment with Connic

Built-in A/B testing routes traffic between agent versions and scores the results, with no separate experimentation tool to wire up.

Try Connic free

How It Works

The concept is simple. A base agent is the control, and a variant contains the change under test. Our platform routes a percentage of live traffic to the variant, while the rest continues going to the control. Both versions run in parallel on real requests, and every run is tracked and attributed to its group.

Incoming Request→Traffic Split→
Control (90%)Variant (10%)
→Compare Results

Variant agents are regular agent configurations. They can change anything: the model, the system prompt, the tools, the temperature. A naming convention keeps it obvious which agent is the base and which is the variant:

agents/
order-processor.yaml                    # base agent (control)
order-processor-test-faster-model.yaml  # variant: testing a cheaper model
order-processor-test-new-prompt.yaml    # variant: testing a rewritten prompt

Deploy both alongside each other. The variant sits dormant until a test starts routing traffic to it.

What Can Be Tested

Since variants are full agent configurations, the possible experiments are wide open.

Model Swaps
Gemini 2.5 Flash vs Claude Sonnet 4.6. Claude vs GPT. Test whether a cheaper or faster model delivers comparable quality for the specific use case. Stop guessing based on benchmarks and measure on actual traffic.
Prompt Iterations
A rewritten system prompt may be more concise, add few-shot examples, or change the output format. Run the new prompt on a slice of traffic and compare judge scores before rolling it out to everyone.
Tool Versions
A team built a v2 of the data processing tool. Does it actually perform better? Point the variant at the new tool implementation and compare success rates and latency with the original.
agents/order-processor-test-faster-model.yaml
name: order-processor-test-faster-model
model: connic/gpt-5.6-luna        # faster, high-volume model
description: "Processes incoming customer orders"
system_prompt: |
  You process incoming orders...
tools:
  - orders.process
  - inventory.check

Everything else stays the same. The variant inherits the exact same workflow. The only difference is the variable under test.

Running a Test

Setting up a test takes less than a minute once the variant is deployed.

  1. Deploy the variant. Push the variant YAML alongside the base agent. After deployment, it shows up as an available test variant.

  2. Open the base agent and click Manage A/B Tests in the header.

  3. Create a new test. Pick the variant, set the traffic percentage, and configure a minimum sample size.

  4. Start the test. Tests begin in Draft status so the configuration can be reviewed before going live.

  5. Monitor and conclude. Watch the comparison metrics fill in as runs come through. At the planned sample size and stopping point, compare the results and declare a winner.

A/B Tests panel showing the concise-summary test running at 15% traffic with control and variant run counts
The Manage A/B Tests panel: each running experiment with its traffic split and live run counts.
Run Multiple Tests at Once
Several A/B tests can run on the same agent simultaneously, each with its own traffic split. For simultaneous model and prompt tests, each variant can receive 10% while 80% stays on the control. Traffic percentages across all active tests cannot exceed 100%.

Reading the Results

The test detail view provides a side-by-side comparison across every relevant metric:

A/B test detail comparing control and variant runs across token cost, duration, judge score, and success rate
Control vs variant, side by side: the variant is cheaper and faster, but its judge score is lower.
Run Counts
Completed runs for control vs variant. Check both groups against the planned sample size; planned; Connic shows when the configured minimum has been reached.
Average Token Cost & Duration
Average token cost and execution time per run. The variant might be 40% cheaper but 200ms slower. The team can then make that tradeoff consciously.
P50 / P95 Latency
Median and 95th percentile response times. Averages hide outliers. If the variant has a great P50 but a terrible P95, some users are having a bad time.
Judge Scores & Success Rate
Average quality scores from configured judges, plus the percentage of runs that completed without errors. These numbers show whether the change is actually better.

A variant badge on every run shows which requests went to which version. Filter the runs table by variant name to drill into specific results.

Auto-Rollback: The Safety Net

Experiments should not break production. When auto-rollback is enabled, Connic watches the variant's failure rate within a rolling window of recent runs. If it crosses the configured threshold, the test pauses and traffic stops going to the variant.

Failure Threshold
Set the maximum acceptable failure rate. If 20% of variant runs are failing while the control sits at 2%, something is wrong and the test pauses automatically.
Rolling Window
The number of recent runs to evaluate. A window of 50 means the system checks the failure rate across the last 50 variant runs, smoothing out isolated hiccups.

When a rollback triggers, all traffic immediately returns to the control agent. The test stays paused and displays a clear error message, so the team can investigate what happened, fix the variant, and try again.

Pair It with Judges

A/B testing shows which version is better. Judges explain why. When the team configures a judge on the base agent, it automatically evaluates both control and variant runs. The A/B test comparison includes the average judge score as a quality signal alongside cost and latency. This is automated agent scoring doing the grading, so the comparison reflects output quality as well as whether a run succeeded.

Without Judges
"The variant is 30% cheaper and has the same success rate." Sounds great. But are the responses actually as good? Success rate only shows that the agent did not crash, not that the output was useful.
With Judges
"The variant is 30% cheaper, same success rate, and the average judge score dropped from 8.5 to 7.2 on accuracy." The savings come at a quality cost, and the team can decide whether that tradeoff is acceptable.

Sticky Sessions

If agents handle multi-turn conversations, a user should not bounce between the control and variant mid-session. When sessions are configured, Connic keeps the same user on the same version for the entire conversation. If the test ends or pauses, sessions fall back to the base agent.

Practical Scenarios

A few experiments that work well in practice:

"Testing a Cheaper Model"
The agent runs on Opus 4.6 at $0.04 per run. Sonnet 4.6 costs a fraction of that. Create a variant that swaps the model, route 10% of traffic, and compare token cost alongside judge scores. If quality holds, the operating cost falls significantly.
"Is the New Prompt Better?"
The system prompt was rewritten with more specific instructions and added few-shot examples. Instead of replacing the current prompt and hoping for the best, run both prompts in parallel. Let the judge scores show which one gives users better answers from.
"Did the New Tool Version Improve Reliability?"
The data processing tool was rebuilt to handle edge cases better. Point the variant at the new version and watch the success rate. If failures drop from 8% to 2%, the evidence supports shipping it to everyone.
"Testing Without RAG Context"
Retrieval adds cost and latency. Is it actually improving responses? Create a stripped-down variant without RAG retrieval and compare. Maybe the base model handles 80% of queries fine on its own, with RAG needed only for the remaining 20%.

Best Practices

Start with Low Traffic
Begin with 5-10% routed to the variant. This catches obvious failures early while limiting the blast radius. Scale up to 25-50% once the variant looks stable.
Change One Thing at a Time
Swap the model or rewrite the prompt, not both. When the variant performs differently, the team needs to know exactly which change caused it. To test multiple changes, run separate tests.
Set a Minimum Sample Size
NIST's sample-size guidance shows why there is no universal run count: the baseline rate, effect worth detecting, significance level, and target power all matter. Plan the floor before launch; Connic tracks each group against the configured minimum and shows when it is reached.
Always Enable Auto-Rollback in Production
Experiments are exciting until they break things for real users. Set a failure threshold and let Connic pause the test if things go wrong. The team can investigate and restart.
Use Judges for Quality Signals
Success rate shows that the agent did not error. Judge scores show that the output was good. Configure judges on the base agent so both groups get evaluated by the same criteria.

Getting Started

The first experiment follows these steps:

  1. Pick one thing to test: a different model, a new prompt, or an updated tool

  2. Create the variant agent YAML with the change and deploy it alongside the base

  3. Open the base agent, click Manage A/B Tests, and create a test with 10% traffic

  4. Enable auto-rollback, set the minimum sample size from the plan, and start the test

  5. Wait for data, compare the results, and ship the winner

Real traffic shows which version is better and replaces impression-based releases.

For the complete setup guide, check the A/B Testing documentation. Teams new to our platform can start with the quickstart guide to deploy a first agent, then run a first experiment.

More from the Blog

Product Spotlight

AI Agent Browser Automation: How It Works and How to Build It

Learn how AI agent browser automation works, when to use it, and how Connic adds managed browsing, persistent state, tabs, dialogs, and files.

September 21, 20268 min read
Product Spotlight

Give Your AI Agents Hundreds of Tools Without Overloading Their Context

Connic keeps a large tool catalog outside the model context and retrieves schemas on demand, so agents can use more tools without loading them all upfront.

September 14, 20267 min read
Product Spotlight

Connic Dev Server: Test AI Agents in the Cloud Without Deploying

Test AI agents with Connic Dev Server: sync local edits to the cloud, keep test connections and variables, run suites, and inspect traces without a Git push.

September 13, 20268 min read
Product Spotlight

Connic Run Context: Give Agents the Data They Need

Connic Run Context keeps custom data available to middleware and tools throughout an agent run, with selected values passed to the AI through prompts.

September 8, 202610 min read
Product Spotlight

AI Agent Routing: Trigger Agents and Return Results

AI agent routing moves events into an agent and returns results to the right system through Connic sync routes and asynchronous outbound connectors.

August 24, 20269 min read
Product Spotlight

Staging to Production: How Connic Environments Isolate AI Agents

Connic environments map git branches to isolated deployments, each with its own secrets, connectors, budgets, and run history, so one agent spec ships safely.

August 6, 20268 min read
Product Spotlight

LLM Context Compression for Long-Running AI Agents

Connic compresses older conversation history and oversized tool results, then retries the model call, so long-running agent sessions survive context limits.

July 20, 20268 min read
Product Spotlight

Human-in-the-Loop AI Agents: Approvals and Secure Input Requests

Connic agents pause for approvals and human input, protect sensitive responses, and resume with the result; tests cover approval, rejection, and timeout.

April 5, 202614 min read
Product Spotlight

Agent Observability: Track Costs, Tokens & Runs

Deploying AI agents without visibility is flying blind. Build custom dashboards, track LLM costs per model, and catch failures before users do.

January 23, 20268 min read