Skip to main content
Connic
Back to BlogIndustry Insights

AI Agent Platform SLA Checklist: What Enterprise Buyers Should Verify

Evaluate an AI agent platform SLA across uptime scope, dependencies, incident response, recovery, security evidence, remedies, and exit terms.

August 29, 202612 min readAuthor: Connic Research Team

A 99.9% badge makes comparison look simple, but the number has little meaning without a boundary. One AI agent request may cross a trigger, queue, runtime, model provider, retrieval system, customer tool, and response channel. Depending on the contract, only the runtime and API may be covered. Buyers should test the promise against the path their production agents actually use.

Begin With the Contractual Promise

Product pages and procurement documents often use reliability terms interchangeably. Separate the metric from its target and from the contractual consequence of missing it. Google's public SRE definitions of SLI, SLO, and SLA use the same separation.

SLI
The service level indicator is the measurement: successful valid requests, latency at a stated percentile, queue delay, or another precisely defined signal.
SLO
The service level objective is the target for that indicator, such as 99.9% successful requests during a monthly measurement window.
SLA
The service level agreement adds contractual consequences when an objective is missed, usually credits, escalation rights, or negotiated termination rights.

If a document says “99.9% uptime” but does not define the measured service, denominator, observation source, exclusions, and remedy, it is not yet a usable buying criterion. NIST's cloud service metrics guidance emphasizes reproducible metric definitions; ENISA likewise tells cloud buyers to establish when a service counts as available and how security service levels will be monitored.

AI Agent Platform SLA Checklist

Use this table for the first contract pass. A “pass” means the answer is written into the binding agreement or an incorporated schedule, not only stated in a sales call.

#VerifyA contract-ready answerRed flag
1Covered servicesNames runtime, API, trigger intake, connectors, queues, state, logs, control plane, and result delivery separately.“The platform” with no component list or environment boundary.
2Availability formulaDefines valid requests, success and failure, response threshold, partial outage, and the denominator.A percentage with no testable definition of downtime.
3MeasurementStates the window, aggregation scope, monitoring source, time zone, and dispute evidence.A status-page percentage assumed to equal the contractual calculation.
4DependenciesAllocates responsibility for model providers, tools, data stores, connectors, networks, and customer code.Third parties excluded without a fallback, attribution, or escalation path.
5Maintenance and exclusionsCaps scheduled windows, specifies notice, and limits customer, quota, force-majeure, and security-event exclusions.Broad exclusions that can remove the incidents the buyer cares about.
6Incident and support clocksSeparates detection, acknowledgment, customer notice, update cadence, response, restoration, and post-incident review.A fast “response” target presented as a resolution promise.
7Recovery and durabilitySets RTO and RPO by data class, plus backup cadence, retention, isolation, and restore-test evidence.“Regular backups” without numeric recovery commitments.
8Security and operational evidenceProvides scoped reports, certificates, audit rights, exportable logs, retention, and control ownership.A badge with no entity, service, period, exceptions, or report access.
9RemediesDefines credit tiers, eligible fees, cap, claim deadline, evidence burden, escalation, and repeat-failure rights.Credits described without the claim procedure or exclusive-remedy language.
10Change and exitCovers adverse-change notice, export formats, retrieval period, transition help, deletion proof, and exit cost.The provider can change terms while the customer has no tested export path.
Bring your critical path to an Enterprise SLA review

Map the runtime, connectors, model providers, recovery targets, support window, and evidence your agents need. Connic can document custom Enterprise terms where the public SLA is not enough.

Review your requirements

Map the Agent's Critical Path Before Negotiating Uptime

An agent run can be accepted while the business task still fails. The model endpoint may time out, a tool may reject credentials, a connector may not deliver the result, or the model may return an unusable answer. Put the production path on one page and assign an owner and contract to every layer.

LayerFailure to testEvidence to request
Trigger and intakeDropped, duplicated, delayed, or rejected eventReceipt ID, timestamp, retry history, dead-letter state
Queue and runtimeRun never starts, stalls, or loses durable stateRun status, queue delay, trace, retry and timeout reason
ModelProvider error, rate limit, latency, or poor outputProvider request ID, model/version, token and latency trace
Tools and dataPermission, network, schema, retrieval, or customer-code errorTool call, identity, parameters, response, approval record
Result deliveryBusiness system never receives the final outcomeDelivery attempt, acknowledgment, retry, terminal state
Account for serial dependencies
If three serial and independent components each deliver 99.9% availability, their theoretical combined availability is about 99.7%. Real dependencies are not always independent, but the calculation shows why a platform percentage cannot stand in for an end-to-end business SLO.

Keep output quality out of the uptime definition. Accuracy, policy compliance, and task completion need use-case tests and production evaluation. See how to build an agent test suite and trace production agent runs. The SLA should make the infrastructure measurable; it cannot guarantee that a probabilistic model gives the right answer to every prompt.

Turn “Nines” Into an Outage Budget

Convert the headline target into minutes, then repeat the calculation using the actual contractual denominator. The figures below assume a 30-day month and no excluded minutes. Maintenance, minimum event duration, low-traffic thresholds, or third-party exclusions can make the effective protection materially smaller.

Monthly targetDowntime budgetBefore exclusions
99.9%43 minutes 12 seconds0.1% of 43,200 minutes
99.95%21 minutes 36 seconds0.05% of 43,200 minutes
99.99%4 minutes 19 seconds0.01% of 43,200 minutes

Ask whether availability is measured per customer, Project, region, component, or the provider's entire fleet. A fleet-wide average can hide one customer's outage; a per-component calculation can create multiple narrow guarantees that never measure the complete production path. Also establish whether the provider's telemetry is final or whether customer logs can rebut it.

Separate Incident Communication From Support Response

A one-hour response target usually commits the provider to acknowledging the ticket. It does not promise restoration in one hour. Record each clock separately and say when it starts: provider detection, customer report, or severity confirmation.

Operational incident terms
Define severity by business impact, 24×7 or business-hours coverage, acknowledgment, customer-notice trigger, update cadence, escalation, restoration target, and the deadline for a written post-incident report.
Security and privacy notices
Keep security-incident and personal-data-breach notice in the security schedule and DPA. The trigger, recipient, required information, and legal clock differ from an availability incident.

A useful escalation schedule also names the supported channel, who can declare a P1, the executive escalation point, and what the customer must provide. Ask whether the post-incident report covers timeline, contributing factors, impact, corrective actions, and recurrence prevention—not merely a status-page summary.

Put Recovery Targets and Proof in Writing

Uptime describes whether a service is available. It does not say how much state can be lost or how quickly a damaged service can be rebuilt. NIST's contingency planning guide treats the recovery time objective (RTO) and recovery point objective (RPO) as separate planning inputs.

RTO: how long recovery may take
Set a maximum restoration time for the runtime, trigger intake, queues, run state, logs, retrieval data, and configuration. The right value can differ by data class.
RPO: how much state may be lost
Set the acceptable point-in-time loss, then verify backup frequency, retention, regional placement, encryption, integrity checks, and the latest successful restore test.

Contract for ongoing access to operational evidence. Buyers should be able to export timestamps, run states, failures, model and configuration versions, tool calls, identities, approvals, and delivery attempts. Define log retention and redaction so the records last long enough for SLA claims and investigations without becoming an uncontrolled store of personal or secret data. ENISA's cloud contract monitoring guide recommends ongoing security feedback between periodic assessments.

Read the Credit Clause From the Claim Deadline Backward

Start with the claim mechanics: eligible fees, automatic or requested credits, percentage tiers, monthly cap, filing deadline, required logs, provider decision process, and expiry. Then check whether credits are the exclusive remedy. For a critical workload, counsel may also seek escalation or termination rights after repeated material failures rather than a larger credit alone.

Before signing, enumerate the agent code, configuration, secrets references, run history, traces, evaluation data, files, and retrieval content that can be exported. State the formats, API availability, retrieval window, transition help, deletion confirmation, and cost. The EU Commission's voluntary cloud switching clauses offer a current reference for switching, termination, security, and business-continuity language. Applicability of the EU Data Act still needs a service- and contract-specific legal assessment.

Procurement framework, not legal advice
Have counsel map the GDPR, NIS2, DORA, Data Act, AI Act, national law, and sector rules to your organization's role and use case. The checks below help locate missing language; they do not determine which laws apply.
DocumentWhat it should answerIt does not replace
SLAAvailability, incidents, support, recovery, measurement, and remediesPrivacy terms or customer continuity planning
DPARoles, instructions, security, subprocessors, breach help, audits, return, and deletionThe controller's lawful basis, notices, DPIA, and configuration duties
Security scheduleTechnical controls, responsibility split, vulnerability handling, and assurance evidenceA scoped report, certificate, test result, or contractual recovery target
Subprocessor scheduleEntity, purpose, location, transfer path, change notice, and objection processReview of customer-selected model, tool, or data providers
Order form and addendaCustom uptime, support, RTO/RPO, liability, regulated workloads, and precedenceTesting that the operational design can meet those requirements

Under GDPR Articles 28, 32, and 33, a buyer acting as controller must use processors that provide sufficient guarantees, document the processor relationship, assess security appropriate to risk, and address breach notification and assistance. Those duties cover confidentiality, integrity, availability, resilience, and timely restoration, but a platform's uptime percentage does not establish GDPR compliance.

Sector rules can demand more. For financial entities in scope, DORA Article 30 requires detailed ICT contract terms and adds precise service targets, incident assistance, contingency testing, audit access, and transition requirements for services supporting critical or important functions. NIS2 requires covered essential and important entities to manage incident, continuity, and direct-supplier risk through national implementing law. Review the official DORA text and NIS2 text with counsel rather than treating either as a universal AI platform rule.

How Connic's Public Terms Answer the Checklist

Connic's public documents let buyers complete much of the first review before an Enterprise sales conversation. The Service Level Agreement is dated July 30, 2026 and applies to Pro and Enterprise. Because the covered boundary, calculation, exclusions, credits, incident communication, and support clocks are visible, the Enterprise review can focus on any custom terms the workload needs.

Checklist areaPublished Connic baselineEnterprise review point
Availability and scope99.9% per monthly billing cycle for agent execution, connectors, and REST APINegotiate a higher target or additional covered surfaces where the business impact requires it.
ExclusionsDashboard, development/preview, non-GA features, customer code, quotas, and third-party providers are outside the stated boundary.Map model, BYOK, tool, and delivery fallbacks rather than assuming end-to-end coverage.
Incidents15-minute acknowledgment after detection, updates at least every 30 minutes, and a significant-outage report within five business daysAdd customer-specific notification, restoration, escalation, and report-content terms if needed.
SupportP1 initial response in one hour; standard support clocks run 09:00–18:00 UTC on weekdays.Specify 24×7 coverage and resolution or restoration objectives where required.
RecoveryThe security page describes backups, redundancy, and regularly tested disaster recovery, but publishes no numeric RTO or RPO.Set numeric objectives by state and data class and request recent restore-test evidence.
Credits10%, 25%, or 50% tiers; 50% monthly cap; claim within 30 days; credits are the sole SLA remedy.Review eligible fees, repeat-failure escalation, and any negotiated termination rights.
Privacy and assurancePublic DPA, security practices, subprocessor list, audit process, and customer responsibility splitRequest current reports and certificates, then verify entity, scope, period, exceptions, and service coverage.
Review the evidence
Read Connic's published security practices, DPA, and subprocessor list. Enterprise buyers can request applicable assurance materials and should verify their current scope.
Watch the service
Check the public status page for live and historical operational reporting. Treat it as evidence, not as a substitute for the monthly calculation and component definitions in the contract.

For an Enterprise review, bring a one-page list of the business-critical agent paths. Add the downtime budget, data-loss budget, support window, required evidence, and legal constraints for each path. The gaps between those requirements and the public baseline become the negotiation brief for the order form.

Turn this checklist into your Connic SLA brief

Share your critical agent paths, support hours, recovery targets, data obligations, and evidence needs. We will show where the public baseline fits and where custom Enterprise terms make sense.

Discuss your SLA

Frequently Asked Questions

There is no universal percentage. Start with the business impact and maximum tolerable outage, then verify the covered components, formula, exclusions, support window, recovery objectives, and remedies. A lower percentage with a complete service boundary can protect a workload better than a higher percentage with broad exclusions.

In a 30-day month, 99.9% allows 43 minutes and 12 seconds of downtime before exclusions. A 31-day month allows 44 minutes and 38 seconds. The contractual formula, measurement window, and excluded events can change the practical result.

Only if the agreement says so. Many platform SLAs exclude customer-selected models and other third-party services. Connic's public SLA expressly excludes failures of third-party LLM providers, so buyers should design fallbacks and negotiate responsibility for the end-to-end path where required.

An SLO is a service target measured by an indicator, such as 99.9% successful valid requests in a month. An SLA is the agreement that adds contractual consequences when the target is missed, such as service credits, escalation, or negotiated termination rights.

No. GDPR compliance depends on the processing purpose, roles, lawful basis, data minimization, security appropriate to risk, processor terms, subprocessors, transfers, rights handling, breach procedures, and the customer's own deployment choices. Review the DPA, security measures, data flow, and customer responsibilities separately.

Yes. Connic's public SLA says Enterprise customers may negotiate higher uptime commitments, enhanced support response times, and additional service credits. Buyers should also put any required 24/7 coverage, numeric RTO/RPO, component scope, incident notice, and evidence terms into the Enterprise agreement.

More from the Blog

Industry Insights

EU-Hosted AI Models in 2026: Providers, Dependence & Options

Compare EU-hosted AI models by location, retention, operator, portability, legal exposure, and deployment model, using a 2026 Commission-requested study.

August 18, 202611 min read
Industry Insights

EU AI Gigafactories: What the €30B Plan Means for Enterprise AI

The EU opened procurement for up to seven AI Gigafactories. The €30B plan may expand EU compute, while pricing, access, and timing remain open.

August 14, 202610 min read
Industry Insights

EU AI Act Article 50: What Your AI Agent Must Disclose

The Commission adopted its final Article 50 guidelines on 20 July 2026, thirteen days before the rules apply. Here is what agent teams have to disclose, and when.

July 26, 202610 min read
Industry Insights

Soofi S Preview: Ollama, GGUF, Benchmarks & Access

Can you run Soofi S with Ollama? Check gated access, official GGUF commands, memory needs, corrected benchmarks, and the September 2026 release plan.

July 15, 202612 min read
Industry Insights

What Is an MCP Connector? A Practical Definition

An MCP connector links an AI app to external tools and data over the Model Context Protocol. Learn how it works and when it beats a custom API integration.

July 8, 20268 min read
Industry Insights

The Real Cost of Assembling Your Own AI Agent Stack

The real cost of assembling your own AI agent stack comes from the integration and maintenance tax between tools. Learn when buying a platform wins.

June 9, 202610 min read
Industry Insights

How to Run AI Agents in the EU Without US Hyperscalers

Run production AI agents in the EU without US hyperscalers: what EU-hosted must really mean, where the US CLOUD Act exposes you, and a sovereignty checklist.

June 4, 20269 min read
Industry Insights

AI Agent Deployment Platforms: 16 Vendors Compared (2026)

Compare 16 AI agent deployment platforms by runtime boundary, language, hosting model, connector ownership, residency, and pricing.

April 19, 202615 min read
Industry Insights

EU AI Act Enforcement: Who Investigates and What Evidence to Keep

The AI Office, national authorities, and EDPS divide EU AI Act enforcement by system and provider; teams should keep scoped governance and runtime evidence.

April 13, 202614 min read