Skip to main content
Connic

Test your agents
before they go live.

Define YAML tests under tests/ and run them in the Connic runtime with real connections. Run them on demand with connic test or automatically before deployment.

Read the testing docs

tests/invoice-processor.yaml

4 passed1 failed
  • extracts invoice total312ms
  • amount > 04ms
  • currency in [EUR, USD]6ms
  • vat is correctly calculated18ms
  • vendor name extracted11ms
$connic test351ms total

Use the Connic test framework directly in the runtime

Run tests where your agents run. The Connic test framework uses the real Connic runtime to check responses and tool calls against your YAML test cases.

tests/invoice-processor.yaml
version: "1.0"

defaults:
  runs: 5                       # invoke the agent 5 times per case
  success_threshold: 80         # 4/5 must pass for the case to pass
  timeout_s: 60                 # per-invocation wall clock

tests:
  - name: extracts_invoice_total
    payload: '{"message": "extract total", "doc_id": "INV-7821"}'
    expected_result: output.total > 0 and output.currency in ("EUR", "USD")
    expected_tool_calls:
      - invoices.extract: invocations >= 1
    expected_no_tool_calls:
      - notifications.send

tests/invoice-processor.yaml

5 passed0 failed
  • extracts_invoice_total312ms
  • expected_result passed (5/5)4ms
  • invoices.extract called 5/56ms
  • notifications.send not called4ms
  • success threshold 80 met2ms
$connic test328ms total

Define what makes a test pass

Use Python-like expressions to check what your agent returns, which tools it calls, and whether errors occur. The Connic test framework uses the same safe evaluator as tool conditions and approval rules.

expected_result

Python-like expression on bindings output, error, status. Supports attribute and subscript access, comparisons, boolean operators, and membership tests.

expected_tool_calls

Bare tool names (called at least once) or one-key mappings like {tool: invocations >= 5}. Mixed entries allowed.

expected_no_tool_calls

Tool names that must NOT be called during the run. Catches the case where the agent should have skipped a tool but didn't.

runs + success_threshold

Invoke the case N times (1–100); pass if at least the threshold percentage succeed. This keeps assertions enforceable on stochastic agents without flakes making the suite unusable.

timeout_s

Per-invocation wall-clock timeout in seconds (1–3600). A timeout counts as a failed run against the threshold.

Python-level checks via cleanup()

Dynamic payload builders can return False from cleanup() to fail the case. The result is AND-ed with the YAML-defined checks for assertions that cannot be expressed as an expression.

For LLM-graded quality scoring on production runs, configure Judges separately in the dashboard.

Run tests from the CLI and check new deployments automatically

Check changes during development and use the same tests in your CI pipeline. Connic checks new deployments before release, so failing tests stop the release.

.github/workflows/agent-tests.yml
name: agent tests

on:
  pull_request:
    paths:
      - "agents/**"
      - "tools/**"
      - "tests/**"

jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: pip install connic
      - run: connic test --env ${{ secrets.CONNIC_CI_ENV_ID }} --json
        env:
          CONNIC_API_KEY: ${{ secrets.CONNIC_API_KEY }}

Compare what the Connic test framework handles for you

See which agent testing capabilities are included and what you need to add yourself with other testing approaches.

Compare what the Connic test framework handles for you
FeatureConnicSpreadsheet evalLangSmith evalDIY pytest
Versioned with the agent projectIncludedNot includedPartialIncluded
Runs in CIIncludedNot includedPartialIncluded
Expression DSL on output / status / errorIncludedNot includedPartialIncluded
Tool-call assertions (positive and negative)IncludedNot includedPartialPartial
Repeats per case + success threshold for stochastic modelsIncludedNot includedPartialPartial
Dynamic Python payload builders with cleanupIncludedNot includedNot includedPartial
Same runtime, real environment, real connectorsIncludedNot includedPartialNot included
Automatic deploy gate (no CI config needed)IncludedNot includedNot includedNot included

Frequently Asked Questions

Each case sets runs and success_threshold. runs invokes the agent N times (1–100) per case; success_threshold is the percent of those runs that must pass for the case to pass overall. So runs: 5, success_threshold: 80 means 4 of 5 invocations must pass.

Fixtures live in tests/files/ and are referenced by bare filename in the case's files: list. The runner reads each file, base64-encodes it, and delivers a multimodal payload of the shape {message, files: [{name, mime_type, data}]}. That's the same wire format webhook multipart uploads produce.

A dynamic payload builder handles this case. A Python module under tests/builders/ exposes build() and optionally cleanup(), while the builder: field points the case to it. The runner executes the pair once per invocation. cleanup() always runs, even on timeout or crash, so external fixtures get torn down reliably.

Tests run inside a one-shot test container against a real Connic environment, same code path as production. By default that's the deploy environment. Setting test_environment_id on the standard environment isolates test traffic to a sibling environment with stub credentials and stage-only connectors.

No. Once tests/ exists, the deploy gate is automatic: every connic deploy and every git auto-deploy runs the suite before the new image ships. A failing case aborts the deployment. connic test can still run locally or in CI for faster feedback before a push.

For a manual deployment, use connic deploy --skip-tests or choose Deploy & skip tests in the dashboard to bypass the gate, for example while a flaky test is being debugged. Automatic Git deployments always run the gate.