Skip to main content
Connic

Track agent quality.
With Connic Judges.

An AI model evaluates agent runs against custom criteria. Connic shows scores and explanations for each evaluated run, along with trends over time. Configure judges per agent in the dashboard.

Read the judges docs

Judges

last 24h
StatusJudgeAvg ScoreEvaluated
  • Active
    answer_accuracy
    on support_agent
    86%
    412 runs
  • Active
    tool_usage
    on support_agent
    91%
    412 runs
  • Active
    tone_match
    on invoice-processor
    64%
    218 runs
  • Active
    completeness
    on support_agent
    89%
    412 runs
  • Active
    response_quality
    on invoice-processor
    97%
    218 runs

Custom quality criteria for each agent

Configure what a judge should evaluate and the maximum score for each criterion in the dashboard. The selected AI model returns a score and explanation for each criterion.

New Judge
Name
Invoice Quality Check
Agent
invoice-processor
Model
connic/gpt-5.6-sol
Trigger
Automatic
Sample rate
20%
Filters
status equals completed
Scoring criteria
17 / 20
Accuracy9 / 10

Did the agent produce a factually correct and complete response based on the input?

Tool Usage4 / 5

Did the agent use the appropriate tools and interpret their results correctly?

Response Quality4 / 5

Is the response well-structured, clear, and appropriately formatted?

Control what gets evaluated and when

Criteria, filters, and sampling define the scope of evaluation. Judges can run automatically or on demand. Average scores show trends, while alerts flag declining scores.

Custom criteria

One or more named criteria each have a description and a max score. The judge LLM scores every criterion independently, with reasoning per score.

Sample rate

The sample rate sets the percentage of matching runs to evaluate (1-100%). A 100% rate scores every run; a lower rate controls cost on high-volume agents.

Run filter expression

Narrow eligible runs with a Python-like expression over context, input, and output fields.

Automatic or manual

Automatic judges score runs as they complete. Manual judges only score on demand. Useful for spot-checks or re-scoring after editing the rubric.

Score alerts

In-app and email notifications fire when the judge's average score drops below a threshold over the last 1, 10, 50, 100, or all-time runs.

Per-criterion averages

Each criterion is tracked over time on the judge detail page, showing which aspects of the agent are slipping.

See docs for the full reference.

Put evaluations to work in ongoing operations

Connic evaluates selected runs in the background after they finish. Scores and explanations are available in run details. Results support comparisons between agent variants and alerts when quality declines.

Filters and sample rate are applied for every active judge. Matching judges evaluate the run asynchronously.
The judge LLM scores each criterion. Scores, reasoning, and token usage are saved against the run.

Average judge scores show up in the side-by-side comparison between control and variant.

Details

An in-app or email notification fires when a judge's rolling average drops below the configured threshold.

Details

Per-criterion scores and reasoning appear at the top of the run detail dialog, with a score pill in the header.

Details

Test scoring criteria on real runs

Connic supports manual evaluation of completed runs. Review and refine a judge’s instructions and criteria, then test them again before enabling automatic evaluation.

System prompt
Domain context for the judge LLM
This agent processes medical insurance claims.
A correct response must include the claim number,
the patient name, and the determination
(approved, denied, or pending review).

Be strict on missing fields and lenient on
phrasing — the response goes to a downstream
system, not directly to the patient.

Optional. This field gives the judge LLM the context needed to evaluate accuracy in the relevant domain.

Trigger Manually
queued
Run
run_a1b2c3 - invoice-processor - completed

Any run can be scored on demand, regardless of trigger mode or sample rate. This supports spot-checks, re-scoring after editing criteria, or testing a new rubric on past runs before switching to automatic.

  • Recent completed runs are available in the dropdown
  • Older runs can be selected by run ID
  • The agent on the run must match the judge's agent

Frequently Asked Questions

Automated agent scoring uses an LLM to grade every agent run against a structured rubric instead of relying on manual review. In Connic, a judge has named criteria with descriptions and max scores and returns a per-criterion score with reasoning for each evaluated run. Average scores and per-criterion breakdowns are tracked over time on the judge detail page.

Each completed evaluation consumes one Connic run unit plus the model call. The run unit and connic/* tokens draw from Project credit at their published rates; BYOK inference is billed by the selected provider. The per-judge sample rate (1-100%) controls both.

In the project's Judges tab. New Judge selects the agent and model, defines a system prompt and one or more named criteria with descriptions and max scores, then sets trigger mode (Automatic or Manual), sample rate, and an optional filter expression.

Manual mode supports evaluations on a handful of recent runs. Per-criterion scores and reasoning help refine the system prompt or criteria descriptions before the judge switches to Automatic.

Yes. Each judge can use a Python-like filter expression over context, input, and output. For example: context.tier == 'enterprise' evaluates only enterprise-tier runs.

Each judge has an optional Score Alert with a threshold percentage and an averaging window (1, 10, 50, 100, or all-time). The alert fires once when the average crosses below the threshold. Delivery is in-app and by email, based on each member's notification preferences.

Judges support the same models as agents: an exact connic/* model with no separate provider key or a configured BYOK provider such as OpenAI, Anthropic, Google Gemini, Azure OpenAI, AWS Bedrock, Vertex AI, OpenRouter, or a custom OpenAI-compatible endpoint.