Track agent quality.
With Connic Judges.
An AI model evaluates agent runs against custom criteria. Connic shows scores and explanations for each evaluated run, along with trends over time. Configure judges per agent in the dashboard.
Read the judges docsJudges
last 24h- Activeanswer_accuracyon support_agent86%412 runs
- Activetool_usageon support_agent91%412 runs
- Activetone_matchon invoice-processor64%218 runs
- Activecompletenesson support_agent89%412 runs
- Activeresponse_qualityon invoice-processor97%218 runs
Custom quality criteria for each agent
Configure what a judge should evaluate and the maximum score for each criterion in the dashboard. The selected AI model returns a score and explanation for each criterion.
Did the agent produce a factually correct and complete response based on the input?
Did the agent use the appropriate tools and interpret their results correctly?
Is the response well-structured, clear, and appropriately formatted?
Control what gets evaluated and when
Criteria, filters, and sampling define the scope of evaluation. Judges can run automatically or on demand. Average scores show trends, while alerts flag declining scores.
One or more named criteria each have a description and a max score. The judge LLM scores every criterion independently, with reasoning per score.
The sample rate sets the percentage of matching runs to evaluate (1-100%). A 100% rate scores every run; a lower rate controls cost on high-volume agents.
Narrow eligible runs with a Python-like expression over context, input, and output fields.
Automatic judges score runs as they complete. Manual judges only score on demand. Useful for spot-checks or re-scoring after editing the rubric.
In-app and email notifications fire when the judge's average score drops below a threshold over the last 1, 10, 50, 100, or all-time runs.
Each criterion is tracked over time on the judge detail page, showing which aspects of the agent are slipping.
Put evaluations to work in ongoing operations
Connic evaluates selected runs in the background after they finish. Scores and explanations are available in run details. Results support comparisons between agent variants and alerts when quality declines.
Average judge scores show up in the side-by-side comparison between control and variant.
DetailsAn in-app or email notification fires when a judge's rolling average drops below the configured threshold.
DetailsPer-criterion scores and reasoning appear at the top of the run detail dialog, with a score pill in the header.
DetailsTest scoring criteria on real runs
Connic supports manual evaluation of completed runs. Review and refine a judge’s instructions and criteria, then test them again before enabling automatic evaluation.
This agent processes medical insurance claims. A correct response must include the claim number, the patient name, and the determination (approved, denied, or pending review). Be strict on missing fields and lenient on phrasing — the response goes to a downstream system, not directly to the patient.
Optional. This field gives the judge LLM the context needed to evaluate accuracy in the relevant domain.
Any run can be scored on demand, regardless of trigger mode or sample rate. This supports spot-checks, re-scoring after editing criteria, or testing a new rubric on past runs before switching to automatic.
- Recent completed runs are available in the dropdown
- Older runs can be selected by run ID
- The agent on the run must match the judge's agent