Assertions
Assert agent output, tool calls, trigger payloads, child-agent behavior, and human approval paths.
On this page
Expression DSL
Test expressions use the same safe evaluator as tool conditions and approval rules.
Python-like syntax: and, or, not; comparisons == != > < >= <=; membership in, not in; parentheses for grouping; string literals in single or double quotes. Reach into nested objects with dot-paths like context.user.role. A bare path like context.active is a truthy check: it passes when the value is set and not empty, zero, or false. Missing fields make the surrounding predicate fail rather than raising.
expected_result
outputoutput.id == 10 and "hi" in output both work.errorNone.status"completed", "failed", "cancelled", "blocked", "awaiting_approval".context.<key>context dict, the same one build() mutated. Empty for tests with no builder. Use this to compare agent output against fixture state the builder just provisioned (e.g. output.id == context.row_uuid). See Dynamic Payload Builders.expected_tool_calls
invocationsparams.<key>context.<key>params and invocations so a tool-call assertion can pin params to a fixture id (e.g. params.uuid == context.test_uuid).approval_decisions.params
params.<key>context.<key>and splits a tool-call expression into params.* filters (per-invocation) and invocations predicates (over the filtered count). context.* may appear on either side. If only params.* conjuncts are given, invocations >= 1 is implied. Repeat the same tool name across entries to lock down distinct argument sets independently. Tool names match either the local function name or the qualified ref.expected_result examples
tests:
# Status check (the most common case)
- name: completes_cleanly
payload: "ping"
expected_result: status == "completed"
# JSON output via attribute access
- name: returns_id_10
payload: '{"a": 4, "b": 6}'
expected_result: output.id == 10
# Substring match on a plain-text reply
- name: greets_user
payload: "hi"
expected_result: '"hello" in output'
# Numeric comparison + boolean composition
- name: high_confidence_only
payload: "classify this"
expected_result: output.confidence >= 0.8 and output.label != "unknown"
# Negative case: a failure is the expected outcome
- name: rejects_invalid_input
payload: '{"vendor": ""}'
expected_result: status == "failed" and "missing vendor" in errorexpected_tool_calls examples
tests:
# Bare name -- the tool must be called at least once
- name: uses_calculator
payload: '{"a": 4, "b": 6}'
expected_tool_calls:
- math.calculator.add
# Mapping form -- expression on invocations
- name: calls_add_at_least_five_times
payload: '{"sum_many": [1,2,3,4,5,6]}'
expected_tool_calls:
- math.calculator.add: invocations >= 5
# Exactly-once enforcement
- name: calls_send_exactly_once
payload: "send a digest"
expected_tool_calls:
- notifications.send: invocations == 1
# Filter by call arguments via params.* -- asserts the agent
# actually used the operands from the payload, not invented ones.
# When invocations is omitted, "at least one matching call" is implied.
- name: calls_add_with_payload_args
payload: '{"a": 4, "b": 6}'
expected_tool_calls:
- math.calculator.add: params.a == 4 and params.b == 6
# Repeat the same tool to lock down each argument set independently.
# Each entry is its own assertion -- this passes when the agent
# calls add(4, ...) once AND add(7, ...) once, in any order.
- name: calls_add_for_each_pair
payload: "compute 4+6 and 7+8 separately"
expected_tool_calls:
- math.calculator.add: invocations == 1 and params.a == 4
- math.calculator.add: invocations == 1 and params.a == 7
# Relative order assertion: these tools must appear in this order
# in the trace. Other tool calls may happen between them.
- name: fetches_then_sends
payload: "look up the order and notify the customer"
expected_tool_call_order:
- orders.lookup
- notifications.send
# Pin params against builder context. The builder inserts a row,
# stashes its uuid in context["test_uuid"], and the agent receives
# the uuid in its prompt. The assertion fails if the agent fetches
# any row other than the one the builder provisioned.
- name: fetches_the_row_we_just_inserted
builder: insert_then_query
expected_result: output.row.id == context.test_uuid
expected_tool_calls:
- db.fetch_row: params.uuid == context.test_uuid and invocations == 1
# Negative assertion: tool must NOT be called
- name: plain_chat_no_tools
payload: "say hi"
expected_no_tool_calls:
- math.calculator.add
- notifications.sendAsserting on Triggered Agents
When the agent under test calls trigger_agent (see trigger_agent), expected_child_agents can assert on the child input, output, tool calls, and further triggers.
The assertion stacks: each entry is keyed by the triggered agent's name and can carry its own expected_child_agents for whatever that child triggers in turn.
tests:
# The dispatcher agent calls trigger_agent("summarizer", ...) with
# wait_for_response=True, so its output and tool calls are available
# to the child-agent assertions.
- name: dispatches_to_summarizer
payload: '{"text": "..."}'
expected_child_agents:
summarizer:
expected_payload: payload.text != ""
expected_result: output.summary != ""
expected_tool_calls:
- llm.complete: invocations >= 1
expected_no_tool_calls:
- email.send
# Pin the trigger payload against builder context, so the test fails if
# the agent forwards the wrong fixture id instead of the one it was
# given. Works whether the parent passed a dict (payload.field) or a
# string (substring via payload_raw).
- name: forwards_charge_id_unchanged
builder: create_charge_then_refund
builder_args:
amount_cents: 4200
expected_child_agents:
billing-refunder:
expected_payload: payload.charge_id == context.charge_id
# Recursive: assert on a grandchild that summarizer triggers in turn.
# Same shape repeats at every depth -- agent name keys mapping to the
# same assertion fields, plus its own expected_child_agents.
- name: dispatches_summarizer_then_publisher
payload: '{"text": "..."}'
expected_child_agents:
summarizer:
expected_result: output.summary != ""
expected_child_agents:
publisher:
expected_tool_calls:
- kafka.publish: params.topic == "summaries"
# Fire-and-forget triggers (wait_for_response=False) cannot have their
# result inspected, but the payload is captured at call time -- so
# expected_payload still applies.
- name: fans_out_telemetry
payload: '{"event": "checkout"}'
expected_child_agents:
telemetry-writer:
expected_triggered: 1
expected_payload: payload.event == "checkout"Two evaluation paths
wait_for_response=True:expected_result,expected_tool_calls,expected_tool_call_order,expected_no_tool_calls, and nestedexpected_child_agentsall apply.wait_for_response=False: fire-and-forget. Useexpected_triggeredandexpected_payloadto check the call and its payload. Result, tool, and nested assertions requirewait_for_response=True.
Asserting on the trigger payload
expected_payload uses the same expression grammar as expected_result with input-side bindings. Use payload.<key> when the parent passed a dict or a JSON string, and payload_raw for substring checks against a free-form string trigger. context.<key> is bound the same way the other assertions bind it, so you can pin a forwarded fixture id with payload.charge_id == context.charge_id. This assertion also works on fire-and-forget triggers because the payload is captured when the call occurs.
trigger_agent_at is treated as fire-and-forget during tests, so scheduled agents are matched by name, count, and captured payload without waiting for the scheduled time.
Matching semantics
- Per-trigger. Each
trigger_agentcall is matched separately with its own tool calls and nested triggers. - At-least-one-must-pass. When the parent triggered the same child more than once, the assertion passes as soon as one waited trigger satisfies the spec.
- Builder context is shared.
context.<key>in a child'sexpected_resultorexpected_tool_callsreads the same builder dict the top-level case uses, so a fixture id stashed inbuild()is reachable at every depth.
Testing Approvals (HITL)
approval_decisions supplies approve, reject, or timeout responses to matching pending approvals. The test continues after the selected decision when the approval policy permits.
tests:
- name: approves_the_exact_refund
builder: create_charge_then_refund
approval_decisions:
- tool: billing.refund
params: params.charge_id == context.charge_id
decision: approve
reason: Approved by this test
expected_result: status == "completed"- Match the approval.
toolis the canonical tool ref. Optionalparamsis a safe expression withparams, buildercontext,true,false, andnullbindings. Omit it to match any parameters for that tool. - Choose the outcome.
decisionisapprove,reject, ortimeout;reasonis optional. Rejections and timeouts honor the approval'son_rejectionsetting. - Decision consumption. Each entry is consumed at most once per invocation.
- Strict matching. Set
strict_approval_decisions: trueindefaultsor on a case to fail on unmatched pending approvals and unused decision entries. The default isfalse. - Approval boundary. Without a matching decision in non-strict mode, the invocation returns with
status == "awaiting_approval".