A/B testing for agents.
Built into Connic.
Compare agent variants using a defined share of production requests. Connic shows costs, execution times, success rate, and judge scores side by side. Optional safeguards pause tests when failure rates are too high or scores are too low.
Read the A/B testing docsorder-handler
A/B test · 4,128 runs
Create an agent variant for the A/B test
Test variants are separate agents named using the pattern {base-agent}-test-{name}. Deploy them alongside the original agent, then configure and start the test in the dashboard.
name: order-handler-test-friendly-tone
model: connic/gpt-5.6-sol
description: "Processes incoming customer orders"
system_prompt: |
You process incoming orders with a warm,
friendly tone...
tools:
- orders.process
- inventory.checkCompare cost, speed, and quality
The overview shows success rate, execution times, and average token costs for the existing agent and the test variant. Scores from configured judges add a comparison against defined quality criteria.
Start A/B tests with clear limits
Create a draft test and start routing requests when the configuration is ready. Pause or conclude the test in the dashboard. An optional failure-rate threshold triggers an automatic pause.
- concludedtest-friendly-toneon order-handler · 4,128 runs · winner declared · success rate 68% vs 49%
- runningtest-haikuon support-triage · 8,901 runs · 10% traffic · judge scores within 1pt of base
- drafttest-shorter-prompton invoice-processor · 0 runs · ready to start · 20% traffic configured
- pausedtest-aggressive-retrieson fraud-detector · 612 runs · auto-rollback · failure rate exceeded threshold
A minimum sample size prevents results from being declared meaningful too early. Auto-rollback lets a failing variant pause itself. Once enough data is available, the test can be concluded and a winner declared if appropriate.
Keep running A/B tests under control
Connic can automatically pause tests when failure limits are exceeded and keep requests within a session assigned to the same variant. Exploratory mode supports testing multiple variants in parallel.
A failure rate threshold defines the limit. If the variant exceeds it within a rolling window of runs, the test pauses itself before more users hit a broken variant.
Several tests can run against the same base agent at once. Their traffic percentages are summed and must total 100% or less. The remainder always routes to control.
When sessions are configured, every request in the same session sees the same variant. If a test is paused or concluded, sticky sessions fall back to the base agent.