Skip to main content
Connic

A/B testing for agents.
Built into Connic.

Compare agent variants using a defined share of production requests. Connic shows costs, execution times, success rate, and judge scores side by side. Optional safeguards pause tests when failure rates are too high or scores are too low.

Read the A/B testing docs

order-handler

A/B test · 4,128 runs

running
control50%
49%
success rate
test-friendly-tone50%
68%
+19pp vs control

Create an agent variant for the A/B test

Test variants are separate agents named using the pattern {base-agent}-test-{name}. Deploy them alongside the original agent, then configure and start the test in the dashboard.

agents/order-handler-test-friendly-tone.yaml
name: order-handler-test-friendly-tone
model: connic/gpt-5.6-sol
description: "Processes incoming customer orders"
system_prompt: |
  You process incoming orders with a warm,
  friendly tone...
tools:
  - orders.process
  - inventory.check
order-handler · Manage A/B Tests
running · 4,128 runs
order-handler (base)50%
test-friendly-tone50%
started 2 days ago

Compare cost, speed, and quality

The overview shows success rate, execution times, and average token costs for the existing agent and the test variant. Scores from configured judges add a comparison against defined quality criteria.

Success rate
68%vs 49%
+19pp
Avg duration (P95)
2.3svs 2.1s
+0.2s
Average token cost per run
$0.021vs $0.018
+$0.003
Judge score
18/20vs 16/20
+2

Start A/B tests with clear limits

Create a draft test and start routing requests when the configuration is ready. Pause or conclude the test in the dashboard. An optional failure-rate threshold triggers an automatic pause.

Recent A/B tests
  • test-friendly-tone
    on order-handler · 4,128 runs · winner declared · success rate 68% vs 49%
    concluded
  • test-haiku
    on support-triage · 8,901 runs · 10% traffic · judge scores within 1pt of base
    running
  • test-shorter-prompt
    on invoice-processor · 0 runs · ready to start · 20% traffic configured
    draft
  • test-aggressive-retries
    on fraud-detector · 612 runs · auto-rollback · failure rate exceeded threshold
    paused

A minimum sample size prevents results from being declared meaningful too early. Auto-rollback lets a failing variant pause itself. Once enough data is available, the test can be concluded and a winner declared if appropriate.

Keep running A/B tests under control

Connic can automatically pause tests when failure limits are exceeded and keep requests within a session assigned to the same variant. Exploratory mode supports testing multiple variants in parallel.

Auto-rollback

A failure rate threshold defines the limit. If the variant exceeds it within a rolling window of runs, the test pauses itself before more users hit a broken variant.

Multiple tests, one agent

Several tests can run against the same base agent at once. Their traffic percentages are summed and must total 100% or less. The remainder always routes to control.

Sticky sessions

When sessions are configured, every request in the same session sees the same variant. If a test is paused or concluded, sticky sessions fall back to the base agent.

Frequently Asked Questions

A Min sample size is configured on the test. Connic doesn't surface results as meaningful until each group has that many completed runs. The docs suggest 50–100 runs per group as a starting point for reliable comparisons.

Yes. Each test gets its own traffic percentage. The percentages of all active tests on an agent must sum to 100% or less; the remainder routes to the base agent.

Yes, when sessions are configured. Every request in the same session sees the same variant. If a test is paused or concluded mid-session, the session falls back to the base agent.

With auto-rollback and a failure rate threshold, the test pauses itself if the variant exceeds the threshold within a rolling window. A test can also be paused or concluded manually from the dashboard at any time.

Feature flags route traffic. A/B testing routes traffic and provides a side-by-side comparison of token cost, duration, success rate, and judge scores between control and variant. Auto-rollback on failure rate means a bad variant pauses itself.

Judge scores are reported alongside token cost, duration, and success rate in the side-by-side comparison. Judges configured on the base agent evaluate runs from both groups, showing how each variant scores against the same rubric.