Skip to main content
Connic
Back to BlogTutorial

Automated Agent Scoring: AI Agent Evaluation with LLM Judges

Automated agent scoring uses an LLM judge to grade sampled or every matching agent run against defined criteria. Track score trends and alert on regressions.

March 29, 2026(last updated: September 8, 2026)10 min readAuthor: Connic Engineering

Automated agent scoring uses an LLM to grade every agent run against defined criteria instead of relying on manual review. Once AI agents handle real traffic, operating without measurement means shipping changes based on impressions instead of data.

Manual review does not scale to hundreds or thousands of runs per day. Automated evaluation uses an LLM to score agent runs against defined criteria, surface quality trends, and alert the team when quality degrades.

That's the LLM-as-a-judge pattern. Connic calls them Judges, and they turn agent quality from a guessing game into a measured, tracked metric.

Why Automated Evaluation Matters

AI agents are non-deterministic. The same input can produce different outputs across runs. A prompt change that improves one use case might break another. Without continuous evaluation, quality problems only surface when users complain.

The Vibe Check Trap
"It seems to be working well" is not a quality metric. Teams that rely on spot-checking miss systematic failures. An agent might nail 90% of queries but consistently fail on one category without the team noticing.
The Silent Regression
A team updates a system prompt, swaps a model, or changes a tool. The agent still responds. But the quality dropped 20%. Without automated scoring, this regression sits in production for weeks until someone notices the customer satisfaction dip.
The ROI Question
The CTO asks: "How are the agents performing?" Without evaluation data, the answer consists of anecdotes. With it, a dashboard shows judge-score trends, per-criteria scores, and comparisons across deployments.
Score every agent run automatically

LLM-as-judge evaluation grades agent output on defined criteria, with no separate eval tool to integrate.

Try Connic free

Further reading: Judges.

How LLM Judges Work

The concept is straightforward. Use a language model to evaluate another language model's output. A scoring rubric defines the specific criteria that matter for the use case, and the judge LLM scores agent runs against that rubric. Recent 2025 research adds useful nuance to early LLM-as-a-judge results. Huang et al. (2025) found that fine-tuned judges can perform well in-domain but weaken on generalizability, fairness, and adaptability, while Thakur et al. (2025) found that high agreement can still hide material score differences and prompt sensitivity. The practical answer is calibration, and Connic gives teams the rubrics, filters, sample rates, stored reasoning, score history, and alerts to run that evaluation loop continuously.

Agent Run Completes→Judge Evaluates→Scores + Reasoning→Dashboard + Alerts

The judge receives the full run context: the input, the output, the execution traces, token usage, and any context metadata. It evaluates each criterion independently and returns a score with reasoning for every one.

Setting Up a Judge

A judge has four components: the agent it evaluates, the model it uses for scoring, the scoring criteria, and the trigger configuration.

  1. Define the Scoring Criteria

    This is the most important step. Vague criteria produce inconsistent scores. Specific criteria produce useful, actionable evaluations.

    Vague Criteria
    • • "Quality" (max 10)
    • • "Helpfulness" (max 10)
    • • "Good response" (max 10)
    Specific Criteria
    • • "Extracted all invoice line items" (max 10)
    • • "Used correct tool for the task" (max 5)
    • • "Response under 200 words" (max 5)

    Each criterion gets a name, a description that tells the judge exactly what to evaluate, and a maximum score. The judge scores each one independently, showing which aspects of performance are strong and which need work.

  2. Choose a Trigger Mode

    Automatic
    Evaluates runs as they complete. Set the sample rate from 1-100% to balance coverage and cost, and use filters to focus evaluation on the traffic that matters.
    Manual
    Evaluation is explicitly triggered on specific runs. This is useful when testing a new judge configuration, re-evaluating after criteria changes, or spot-checking suspicious runs.
    Tip: Calibrate, Then Automate
    Start in manual mode on a representative set of runs and compare scores with the team's expectations. Refine the rubric, then switch to automatic mode for continuous monitoring.
  3. Give the judge domain context. If an agent processes medical insurance claims, describe what a correct claim determination looks like. If an agent writes marketing copy, describe the brand voice and quality bar.

    Example Judge System Prompt
    You are evaluating a customer support agent for an e-commerce
    platform. A high-quality response should:
    
    - Directly address the customer's specific question
    - Reference the correct order, product, or policy
    - Provide actionable next steps (not generic advice)
    - Maintain a professional, empathetic tone
    - Avoid making promises the company cannot keep
    
    If the agent used tools, verify it queried the correct data
    before responding.

    The evaluated agent's own system prompt can also be included in the judge's context. This gives the judge both the intended task and the agent's actual behavior.

A Complete Example

A practical example: evaluating a document extraction agent that processes invoices.

Invoice Extraction Judge

Model

openai/gpt-5.4 (or any model with strong reasoning)

Trigger

Automatic - 100% sample rate (low volume, evaluate every run)

Criteria

max 10Data Completeness- Did the agent extract all required fields? (vendor, date, line items, total)
max 10Accuracy- Are the extracted values correct? Do line item totals match?
max 5Tool Usage- Did the agent use the correct tools in the right order?
max 5Uncertainty Handling- Did it flag unclear fields instead of guessing?

With this setup, every invoice extraction run is scored out of 30. The dashboard shows at a glance whether accuracy is high but completeness is dropping, or whether the agent is guessing instead of flagging uncertainty. Each score comes with reasoning explaining why the judge assigned that number.

Run Filters: Focus Evaluations Where They Matter

Not every run needs evaluation. Filters target specific traffic:

Status Filters
Only evaluate completed runs (skip failures). Or only evaluate failed runs to understand what went wrong.
Context Filters
Filter on custom context values set by middleware. Only evaluate production traffic, premium customers, or specific use case categories.

That supports separate judges for different scenarios. A support agent handling both billing and technical questions might need two judges with different criteria and filters instead of one blended evaluation.

Quality Alerts

Scores are useful. Alerts are actionable. Configure a score threshold and get notified when quality drops below it.

Threshold
Set a minimum score percentage. When the average drops below it, Connic sends an alert.
Average Window
Average over the last 1, 10, 50, or 100 evaluations. Smaller windows react faster. Larger windows smooth out outliers.
Low Score Filter
Filter evaluations to show only runs below a configurable percentage. Quickly find the worst performers for investigation.

What the Dashboard Shows

The judge detail page shows quality over time:

Connic judge detail page showing average score, total evaluations, per-criteria averages for quality and groundedness, token cost, and the scoring configuration
Judge detail with average score, per-criteria breakdown, cost tracking, and the scoring config side by side.
Average Score
Overall quality metric with trend indicator. Color-coded: green (80%+), amber (50-80%), red (below 50%).
Criteria Averages
Per-criterion breakdown showing exactly which aspects are strong and which are weak. If accuracy is 95% but tool usage is 60%, the weak area is clear.
Evaluation Detail
Each evaluation includes the full reasoning behind its scores, so the evaluation logic is available for review and refinement.
Cost Tracking
Total tokens, average tokens per evaluation, and cumulative cost of running the judge. These figures show the exact cost of quality monitoring.

Practical Patterns

A few evaluation setups that work well in production:

Deployment Canary
Set the alert window to "last 10 runs" and deploy a new agent version. If the average score drops below the configured threshold within the first 10 runs, Connic sends an immediate alert. Inspect the affected traces and roll back before the regression affects more users.
Per-Use-Case Quality Tracking
A single agent handles multiple use cases. Create separate judges with context filters for each one. Track billing question quality independently from technical question quality. Each gets its own criteria, thresholds, and alerts.
Model Comparison
Thinking about switching from GPT to Claude or Gemini? Run both models in parallel (using a fallback model configuration), then compare judge scores across model versions. Make the decision with data, not assumptions.
Prompt Engineering Feedback Loop
Update a system prompt. Filter evaluations by low scores. Read the judge's reasoning for the worst runs. Refine the prompt based on specific failure patterns. Repeat until the criteria averages stabilize.

Cost Considerations

Each successfully completed judge evaluation consumes one additional run unit and makes an LLM call. Run-package volume savings apply to the judge run unit, not the model call. A few ways to keep costs reasonable:

  • •Use sample rates for high-volume agents. Choose a rate that provides useful coverage without evaluating every run.
  • •Benchmark efficient judge models. A cheaper judge can reduce cost; validate it against the same labeled set before switching.
  • •Keep criteria focused. Five specific criteria beat twenty vague ones. Fewer criteria also mean shorter prompts and lower token usage.
  • •Use filters to exclude noise. Don't evaluate test runs, internal debugging, or low-value traffic.

Getting Started

Adding a judge to an existing agent is straightforward:

  1. Navigate to Judges in the Project and create a new judge

  2. Select the agent to evaluate and choose a model for the judge

  3. Define 3-5 specific scoring criteria with clear descriptions

  4. Set manual mode and test on a set of recent reviewed runs

  5. Refine the rubric, then switch to automatic monitoring

  6. Set a score alert threshold so quality drops trigger a notification

The hardest part is writing good criteria descriptions, which deserve most of the attention. Everything else is configuration.

The full setup guide is in the Judges documentation. For teams new to Connic, the quickstart guide covers deploying a first agent before a judge is added to start measuring quality.

Frequently Asked Questions

LLM-as-a-judge uses a language model to evaluate another model's output against defined criteria. Connic Judges can score sampled or every matching agent run, store criterion-level scores and reasoning, track trends, and trigger alerts automatically.

Accuracy depends on the judge model, rubric, and use case. ACL 2025 research found that strong in-domain performance may not generalize and that high agreement can hide score differences or prompt sensitivity. Connic's manual mode makes it easy to calibrate a Judge on representative reviewed runs before enabling continuous evaluation.

Criteria should be specific and measurable, not vague. Instead of 'quality' or 'helpfulness', use criteria like 'extracted all required invoice fields (vendor, date, line items, total)', 'used the correct tool for the task', or 'response under 200 words'. Aim for 3-5 focused criteria per Judge.

For low-volume or high-impact agents, evaluate every matching run. For higher-volume agents, use Connic's 1-100% sample rate and filters to balance coverage and cost while keeping the evaluated traffic representative of the relevant use cases.

Configure a score threshold and an averaging window over the last 1, 10, 50, or 100 evaluations. Smaller windows react faster; larger windows smooth outliers. When the average crosses the threshold, Connic alerts the configured recipients so they can inspect the affected traces.

Each completed Judge evaluation consumes one Connic run unit plus the model call. The run unit and connic/* tokens use Project credit at their published rates; BYOK inference is billed by the selected provider. Control cost with sample rates, focused filters, efficient models, and concise criteria.

More from the Blog

Tutorial

Expose the Invoice Processor Template as an MCP Tool

Deploy Connic's invoice template as an MCP tool, discover it from Python, send invoice text, and check the extracted fields with a complete client example.

September 11, 202610 min read
Tutorial

How to Deploy a Python AI Agent Without Kubernetes

A Python AI agent deployed without Kubernetes using YAML, plain Python, deployment-gated tests, Git, and a managed EU runtime. Includes working code.

August 12, 202612 min read
Tutorial

How to Trigger AI Agents from Kafka Topics

Point a Connic Kafka inbound connector at a topic and every message starts an agent run. Configure the connector, link an agent, deploy, and watch runs.

July 12, 20268 min read
Tutorial

How Small Engineering Teams Can Add an AI Agent to SaaS

A practical, step-by-step path to shipping a first production AI agent with a small team: scope one job, define it in config, connect it to existing systems, and let a runtime handle the rest.

June 12, 20269 min read
Tutorial

Migrate from LangChain to Production AI Agents

A working LangChain prototype still needs to handle real traffic. Migrate existing agent code to a production-grade platform without rewriting from scratch.

March 23, 202611 min read
Tutorial

Database vs Retrieval vs Sessions: Agent Memory Compared

A comparison of Connic's Database, Retrieval, and persistent sessions, including identity, TTL, sanitized events, and runs matched to that identity.

March 4, 202612 min read
Tutorial

AI Agents: From Prototype to Production

A demo works great until 1,000 concurrent users arrive. A practical guide to the production requirements most teams find out about too late.

January 10, 202610 min read
Tutorial

Hidden Costs of Self-Hosting AI Agents

We'll just deploy it on Kubernetes. Famous last words. The true cost of self-hosting AI agents versus a managed platform.

December 18, 20257 min read
Tutorial

AI Agent RAG Tutorial: Retrieval With Citations

Build a production RAG agent with scoped retrieval namespaces, read-only permissions, source citations, custom tool wrappers, and regression tests.

November 15, 20259 min read