Skip to main content
Connic
Back to BlogProduct Spotlight

Secure AI Agents: A Production Safety Checklist

A production security checklist for AI agents, updated with ENISA’s 2026 evidence: scoped tools, approvals, logging, and incident response for critical sectors.

March 21, 2026(last updated: September 23, 2026)16 min readAuthor: Connic Engineering

A developer would not ship a web application without authentication, input validation, and rate limiting. But many teams ship AI agents with none of the equivalent safeguards. The agent works in the demo, so it goes to production. Then the first creative user discovers they can make it do things it was never supposed to do.

The checklist covers the controls required to secure AI agents before they go live and how to implement them.

The Threat Model for AI Agents

AI agent security isn't theoretical. The OWASP Top 10 for Agentic Applications 2026 expands the threat model beyond model output to goal hijacking, tool misuse, privilege abuse, memory and context poisoning, and cascading failures. These risks grow when an agent can act across systems with a user's identity or persistent context.

Securing an agent starts with understanding the relevant threats. These attacks occur in production:

Prompt Injection
A user crafts input that overrides the agent's system prompt. "Ignore all previous instructions and..." is the obvious version, but sophisticated attacks use encoding tricks, scrambled characters, and structural manipulation that bypass simple keyword filters. For agents, injection can become goal hijacking when it changes plans or tool calls.
PII Exposure
Users paste sensitive data into prompts: email addresses, credit card numbers, social security numbers, phone numbers. Without guardrails, this data flows into model provider APIs and potentially into application logs. Under GDPR and CCPA, this creates real compliance liability.
System Prompt Extraction
Attackers trick the agent into revealing its system prompt, including internal instructions, tool configurations, API schemas, and business rules. Once exposed, they know exactly how to manipulate the agent.
Data Exfiltration
Crafted prompts cause the agent to encode sensitive context data into URLs, markdown images, or structured output, silently sending internal data to external servers.
Off-Topic and Harmful Output
Without content controls, a billing support agent might start giving medical advice. An internal assistant might generate toxic or inappropriate content. The model does not follow a content policy unless the application enforces it.

2026 ENISA threat update

Published on 22 September 2026, the ENISA Threat Landscape 2026 analyses incidents and events observed from 1 January to 31 December 2025.

ENISA’s summary reports that 73% of targeted organisations were classified as NIS2 essential or important entities. The report’s sectoral analysis expresses that share as recorded events. This describes the 2025 incident dataset, not the proportion of all NIS2 entities attacked or the proportion of attacks involving AI.

The report describes attackers using AI to augment existing skills during 2025. Looking ahead, ENISA expects emerging AI models to increasingly support malicious operations. Its AI outlook is a forecast of greater automation, speed, and scale. Operators in critical sectors can prepare by limiting what an agent can reach, gating consequential actions, and rehearsing how to contain a compromised workflow.

Secure AI agents for production

Isolation, secrets management, guardrails, and audit logging are built into the platform.

Get started free

The Production Safety Checklist

These security controls should be in place before an agent faces real users. Each one below comes with specific implementation guidance.

1. Block Prompt Injection on Input

Every user message should be checked for injection attempts before the agent processes it.

Effective prompt injection detection goes beyond keyword matching. It needs to catch:

  • →Direct instruction overrides: "Ignore previous instructions" and variants
  • →Encoding attacks: Base64-encoded instructions, hex-encoded payloads
  • →Character manipulation: typoglycemia (scrambled characters), Unicode substitution
  • →Structural attacks: delimiter injection, context manipulation through formatting
agent.yaml
guardrails:
  input:
    - type: prompt_injection
      mode: block
      config:
        sensitivity: medium  # low, medium, or high

When a prompt injection is detected, the agent never executes. The user receives a rejection message, and the attempt is logged as a trace span for security review.

For High-Risk Applications
Connic supports specialized injection detection providers like Lakera, which offer continuously updated detection models and support for 100+ languages. Add them by setting the provider field in the guardrail config.

2. Handle PII Before It Reaches the Model

Users will paste sensitive data into an agent. The question is whether the data reaches the model provider's API or gets intercepted first.

Three modes are available:

Block
Reject the entire message if PII is detected. Strictest option. Good for environments where PII should never enter the system at all.
Redact
Replace detected PII with placeholders like [EMAIL_REDACTED]. The agent processes the sanitized input. Best balance of safety and usability.
Warn
Log the detection but continue processing. Use this to measure how often PII appears before deciding whether to block or redact.
agent.yaml
guardrails:
  input:
    - type: pii
      mode: redact
      config:
        entities: [email, phone, ssn, credit_card, api_key]
  output:
    - type: pii_leakage
      mode: redact

Notice the output guardrail. PII can also appear in agent responses, even if it wasn't in the original input. If an agent has access to customer data through tools or retrieval bases, output PII detection keeps it from surfacing sensitive information in responses.

3. Prevent System Prompt Leakage

A system prompt contains the agent's personality, business rules, tool schemas, and operating instructions. If an attacker extracts it, they have a blueprint for manipulating the agent.

agent.yaml
guardrails:
  output:
    - type: system_prompt_leakage
      mode: block

This checks every agent response for fragments that match the system prompt. If the agent starts revealing its instructions, the response is blocked and replaced with a safe message.

4. Restrict the Agent to Its Lane

Language models will happily answer questions about anything. A billing support agent shouldn't be giving cooking advice or medical recommendations. Topic restriction keeps the agent focused on what it's supposed to do.

agent.yaml
guardrails:
  input:
    - type: topic_restriction
      mode: block
      config:
        allowed_topics:
          - billing and payments
          - subscription management
          - account settings
        off_topic_message: "I can only help with billing and account questions."
        model: connic/gpt-5.6-luna

The configuration can define allowed topics (whitelist) or blocked topics (blacklist), depending on the use case. The check uses a lightweight LLM call for classification, which is why it takes a model parameter. Use a fast, cheap model to keep latency and cost low.

5. Moderate Output Content

Even with a good system prompt, language models can generate inappropriate content. Content moderation on output catches toxicity, harassment, hate speech, and other policy violations before they reach the user.

agent.yaml
guardrails:
  output:
    - type: moderation
      mode: block
      config:
        categories:
          - hate
          - harassment
          - violence
          - self_harm

6. Detect Data Exfiltration Attempts

A subtle but dangerous attack: prompts that cause the agent to encode sensitive data into URLs, markdown images, or structured output that gets sent to external servers. This is especially dangerous for agents with access to internal data through tools.

agent.yaml
guardrails:
  output:
    - type: data_exfiltration
      mode: block
      config:
        allowed_domains:
          - yourdomain.com
          - docs.yourdomain.com

This catches suspicious URLs with encoded data in query parameters, markdown image tags pointing to external domains, and base64-encoded data blocks in the response. The allowed domains whitelist ensures legitimate references are not blocked.

7. Add Custom Business Rules

Every business has unique safety requirements that built-in guardrails can't cover. Custom guardrails support arbitrary validation logic in Python.

guardrails/compliance_check.py
from connic import GuardrailResult

REQUIRED_DISCLAIMERS = [
    "not financial advice",
    "consult a professional",
]

def check(content: str, context: dict) -> GuardrailResult:
    """Ensure financial responses include disclaimers."""
    agent_name = context.get("agent_name", "")
    if "finance" not in agent_name:
        return GuardrailResult(passed=True)

    content_lower = content.lower()
    has_disclaimer = any(d in content_lower for d in REQUIRED_DISCLAIMERS)

    if not has_disclaimer:
        return GuardrailResult(
            passed=False,
            message="Response must include a disclaimer.",
        )
    return GuardrailResult(passed=True)

Common custom guardrails include: competitor mention blocking, regulatory compliance disclaimers, internal terminology filters, and domain-specific validation rules.

8. Check Output Relevance

Sometimes an agent responds without saying anything harmful, but also without answering the actual question. Relevance checking catches goal hijacking (where an attacker subtly redirects the agent's purpose) and hallucinated tangents.

agent.yaml
guardrails:
  output:
    - type: relevance
      mode: warn
      config:
        model: connic/gpt-5.6-luna

Start with warn mode to understand how often irrelevant responses occur, then escalate to block mode if the rate is unacceptable.

9. Scope Tools and Gate High-Impact Actions

Expose only the tools a workflow needs. Restrict built-in database tools to approved collections and read/write operations, and limit retrieval tools to the required namespaces. A permitted retrieval namespace also grants access to its descendants. Custom tools need their own authorization checks and narrowly scoped external credentials. Keep production credentials separate from test credentials through environment-scoped variables.

Configure approval gates for actions such as deletions, payments, or changes to access rights. Only configured tools pause for approval; other tools execute normally. Review the proposed tool and parameters, and restrict decisions to designated members with approvals.decide permission. With on_rejection: fail, a rejection or timeout fails the run. With on_rejection: continue, the rejected tool is skipped but the agent can pursue another action. Test that behavior against the workflow’s allowed alternatives.

Putting It All Together

A complete guardrail configuration for a production customer support agent:

agents/support.yaml
version: "1.0"
name: customer-support
model: connic/gpt-5.6-terra
system_prompt: |
  You are a customer support agent for Acme Corp.
  Help customers with orders, billing, and product questions.
tools:
  - support.search_orders
  - support.search_retrieval
  - support.create_ticket

guardrails:
  input:
    - type: prompt_injection
      mode: block
      config:
        sensitivity: medium
    - type: pii
      mode: redact
      config:
        entities: [email, phone, ssn, credit_card]
    - type: topic_restriction
      mode: block
      config:
        allowed_topics: [orders, billing, products, returns]
        off_topic_message: "I can only help with Acme Corp products and orders."
        model: connic/gpt-5.6-luna
  output:
    - type: moderation
      mode: block
    - type: system_prompt_leakage
      mode: block
    - type: pii_leakage
      mode: redact
    - type: data_exfiltration
      mode: block
      config:
        allowed_domains: [acmecorp.com]
    - type: relevance
      mode: warn
      config:
        model: connic/gpt-5.6-luna
Ordering Matters
Guardrails execute in their configured order. Place cheap, fast checks first (prompt injection, PII, regex) and expensive checks last (topic restriction, relevance). If an early check blocks the request, the expensive ones never run.

Observability: See What Is Being Caught

Run traces show guardrail evaluations alongside tool calls and errors. The separate project audit log records team actions, including approval decisions. Incident review needs both the execution record and the decisions that allowed an action to proceed.

Trace Spans
Each guardrail check creates a span with the rule type, mode, direction, pass/fail status, and detection details. See exactly what triggered and why.
Approval Decisions
Approval records identify the tool, parameters, reviewer, and decision. Compare these records with the affected run to investigate what was requested and what was authorized.
Connic project audit log listing approvals, connector links, deployments, environment changes, and judge changes by user and time
The project audit log records team actions and approval decisions with their actor and time.

The logging and debugging documentation explains how to inspect traces and custom-code logs. Capture is limited to 500 custom-code log lines per run; additional lines are dropped. Configure run retention and preserve incident evidence before runs are deleted. Environment redaction rules cover selected JSON fields in new records; they leave plain-text secrets and existing history unchanged. Avoid writing credentials into logs.

This data answers questions such as: How often is prompt injection attempted? Which agents trigger the most PII redactions? Are topic restrictions too aggressive? Guardrails can then be tuned from real traffic patterns instead of guesswork.

Beyond Guardrails: Defense in Depth

Guardrails are the runtime safety layer. But a complete security posture includes additional controls:

Iteration Limits
Cap the number of LLM calls per agent run. Prevents infinite loops, reduces cost exposure from runaway agents, and catches bugs where agents get stuck in cycles.
Concurrency Control
Key-based concurrency ensures only one agent run per unique key is active at a time. Prevents duplicate processing and reduces exposure to replay attacks.
Database Access Controls
If an agent has database tools, restrict which collections it can read and write. Apply the principle of least privilege so agents only access the data they need.
Retrieval Namespacing
Scope agent access to specific retrieval namespaces. Prevent one agent from accessing another agent's sensitive data in shared retrieval systems.
Automated Quality Evaluation
Use LLM judges to continuously score agent quality. Catch regressions in accuracy, safety, or compliance before they become incidents.

Critical-Sector Pre-Launch and Response Checklist

For workflows supporting essential services, the launch review should include a tested procedure for containing incidents and restoring service. These operational checks supplement the guardrail configuration above; they do not establish NIS2 compliance.

Before launch

  • List reachable systems, permitted data, tool operations, and credentials. Verify that access to unrelated collections, namespaces, and environments is denied, including through custom tools.
  • Test each consequential tool call with approval, rejection, and timeout. Confirm that reviewers see enough context and that an alternative tool cannot bypass the intended gate.
  • Trace a test action from input through tool arguments and approval to its external result. Verify log access, JSON redaction, retention, and the procedure for preserving incident records.
  • Assign an incident owner and rehearse stopping new requests, cancelling active work, revoking credentials, and restoring a known-good deployment. Keep a manual service procedure available while the agent is unavailable.

When an incident is suspected

  1. Stop new requests at their sources and pause affected connectors. Use the run cancellation controls for scheduled, queued, or running executions. Cancellation is not instantaneous and cannot reverse completed external actions.
  2. Revoke affected credentials at the external provider before resolving pending approvals. Rejection with on_rejection: continue can resume the run, so keep its access to affected systems revoked during investigation.
  3. Preserve run IDs, traces, approval decisions, project audit records, and external-system records. Record the affected data and actions, and involve the organisation’s incident-response and compliance owners to assess notification duties.
  4. Correct the access rule, approval gate, or tool behavior that failed. Reproduce the incident safely in an isolated test environment and verify the fix. Update environment variables and redeploy before resuming traffic; editing a variable does not update an existing deployment.

Getting Started

For a new implementation, the recommended approach is:

  1. Start with the baseline. Add prompt_injection and pii on input, moderation and system_prompt_leakage on output. This covers the most common attack vectors.

  2. Tune in a test environment. Use warn mode with non-sensitive test inputs to inspect detections. Enable the required blocking and redaction controls before exposing production data or consequential tools.

  3. Layer up based on traffic. Once real usage patterns are visible, add topic restriction, relevance checking, and custom guardrails based on the observed risks.

  4. Review trace data regularly. Check which guardrails fire most often and investigate patterns. Adjust sensitivity and modes as the traffic becomes better understood.

The Guardrails documentation provides the full configuration reference. Our real-time guardrails guide describes the built-in types in more detail.

Frequently Asked Questions

The OWASP Top 10 for Agentic Applications 2026 covers goal hijacking, tool misuse, identity and privilege abuse, supply-chain vulnerabilities, unexpected code execution, memory and context poisoning, insecure inter-agent communication, cascading failures, human-agent trust exploitation, and rogue agents. Connic addresses these production risks with runtime guardrails, scoped tools, approvals, iteration limits, isolated execution, and traceable runs.

Effective prompt injection defense goes beyond keyword matching. Detection must cover direct instruction overrides, encoding attacks (Base64, hex), character manipulation (Unicode substitution, scrambled text), and structural injection through formatting. Apply it to every user input before it reaches the agent. For high-risk applications, specialized providers like Lakera offer continuously updated detection models that cover 100+ languages.

Deploy a PII guardrail on agent input before the message reaches the model provider API. Three modes are available: block the entire message, redact detected entities (email, phone, SSN, credit card) and continue with sanitized input, or warn and log for measurement before deciding. Output PII detection is also necessary because agents with access to customer data through tools can inadvertently include sensitive data in their responses.

Combine input and output guardrails with narrowly scoped tools, environment-scoped credentials, and approval gates for high-impact actions. Verify rejection and timeout behavior, restrict access to run and audit records, and set retention and redaction rules. For critical-sector workflows, rehearse stopping new requests, cancelling runs, revoking credentials, preserving evidence, and recovering service before launch.

Data minimization, restricted access, human oversight, and records of agent activity can support a compliance review. Teams must verify what their controls detect, what records are retained, and who can access them. Guardrails, approvals, and logging do not establish GDPR or EU AI Act compliance by themselves; the applicable obligations depend on the system and its use.

System prompts instruct the model on desired behavior. They are a soft control that the model can ignore or be manipulated to override. Guardrails are hard runtime controls that intercept and evaluate every input and output independently of what the model produces. A well-written system prompt reduces unwanted outputs; guardrails catch them when they occur anyway. Both are needed: prompts for normal operation, guardrails as the enforcement layer.

More from the Blog

Product Spotlight

AI Agent Browser Automation: How It Works and How to Build It

Learn how AI agent browser automation works, when to use it, and how Connic adds managed browsing, persistent state, tabs, dialogs, and files.

September 21, 20268 min read
Product Spotlight

Connic Dev Server: Test AI Agents in the Cloud Without Deploying

Test AI agents with Connic Dev Server: sync local edits to the cloud, keep test connections and variables, run suites, and inspect traces without a Git push.

September 13, 20268 min read
Product Spotlight

Connic Run Context: Give Agents the Data They Need

Connic Run Context keeps custom data available to middleware and tools throughout an agent run, with selected values passed to the AI through prompts.

September 8, 202610 min read
Product Spotlight

AI Agent Routing: Trigger Agents and Return Results

AI agent routing moves events into an agent and returns results to the right system through Connic sync routes and asynchronous outbound connectors.

August 24, 20269 min read
Product Spotlight

Staging to Production: How Connic Environments Isolate AI Agents

Connic environments map git branches to isolated deployments, each with its own secrets, connectors, budgets, and run history, so one agent spec ships safely.

August 6, 20268 min read
Product Spotlight

LLM Context Compression for Long-Running AI Agents

Connic compresses older conversation history and oversized tool results, then retries the model call, so long-running agent sessions survive context limits.

July 20, 20268 min read
Product Spotlight

Connic Tests: Catch Agent Regressions Before They Reach Production

A YAML-driven testing framework built for non-deterministic AI agents. Repeated-run pass thresholds, expression-based assertions, custom-code mocking, multimodal fixtures, and a deploy gate that blocks failed checks by default.

May 6, 20268 min read
Product Spotlight

A/B Testing for AI Agents: Ship Better Prompts with Confidence

A changed prompt may feel better. A controlled experiment shows whether it is and lets real traffic decide.

March 27, 20269 min read
Product Spotlight

Agent Guardrails: Real-Time Safety for AI Agents

Connic Guardrails intercept agent inputs and outputs in real time to block prompt injection, redact PII, and enforce topic restrictions.

March 3, 20269 min read