Test your agents
before they go live.
Define YAML tests under tests/ and run them in the Connic runtime with real connections. Run them on demand with connic test or automatically before deployment.
Read the testing docstests/invoice-processor.yaml
- extracts invoice total312ms
- amount > 04ms
- currency in [EUR, USD]6ms
- vat is correctly calculated18ms
- vendor name extracted11ms
Use the Connic test framework directly in the runtime
Run tests where your agents run. The Connic test framework uses the real Connic runtime to check responses and tool calls against your YAML test cases.
version: "1.0"
defaults:
runs: 5 # invoke the agent 5 times per case
success_threshold: 80 # 4/5 must pass for the case to pass
timeout_s: 60 # per-invocation wall clock
tests:
- name: extracts_invoice_total
payload: '{"message": "extract total", "doc_id": "INV-7821"}'
expected_result: output.total > 0 and output.currency in ("EUR", "USD")
expected_tool_calls:
- invoices.extract: invocations >= 1
expected_no_tool_calls:
- notifications.sendtests/invoice-processor.yaml
- extracts_invoice_total312ms
- expected_result passed (5/5)4ms
- invoices.extract called 5/56ms
- notifications.send not called4ms
- success threshold 80 met2ms
Define what makes a test pass
Use Python-like expressions to check what your agent returns, which tools it calls, and whether errors occur. The Connic test framework uses the same safe evaluator as tool conditions and approval rules.
Python-like expression on bindings output, error, status. Supports attribute and subscript access, comparisons, boolean operators, and membership tests.
Bare tool names (called at least once) or one-key mappings like {tool: invocations >= 5}. Mixed entries allowed.
Tool names that must NOT be called during the run. Catches the case where the agent should have skipped a tool but didn't.
Invoke the case N times (1–100); pass if at least the threshold percentage succeed. This keeps assertions enforceable on stochastic agents without flakes making the suite unusable.
Per-invocation wall-clock timeout in seconds (1–3600). A timeout counts as a failed run against the threshold.
Dynamic payload builders can return False from cleanup() to fail the case. The result is AND-ed with the YAML-defined checks for assertions that cannot be expressed as an expression.
For LLM-graded quality scoring on production runs, configure Judges separately in the dashboard.
Run tests from the CLI and check new deployments automatically
Check changes during development and use the same tests in your CI pipeline. Connic checks new deployments before release, so failing tests stop the release.
name: agent tests
on:
pull_request:
paths:
- "agents/**"
- "tools/**"
- "tests/**"
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- run: pip install connic
- run: connic test --env ${{ secrets.CONNIC_CI_ENV_ID }} --json
env:
CONNIC_API_KEY: ${{ secrets.CONNIC_API_KEY }}Compare what the Connic test framework handles for you
See which agent testing capabilities are included and what you need to add yourself with other testing approaches.
| Feature | Connic | Spreadsheet eval | LangSmith eval | DIY pytest |
|---|---|---|---|---|
| Versioned with the agent project | Included | Not included | Partial | Included |
| Runs in CI | Included | Not included | Partial | Included |
| Expression DSL on output / status / error | Included | Not included | Partial | Included |
| Tool-call assertions (positive and negative) | Included | Not included | Partial | Partial |
| Repeats per case + success threshold for stochastic models | Included | Not included | Partial | Partial |
| Dynamic Python payload builders with cleanup | Included | Not included | Not included | Partial |
| Same runtime, real environment, real connectors | Included | Not included | Partial | Not included |
| Automatic deploy gate (no CI config needed) | Included | Not included | Not included | Not included |