SLA, SLO and SLI

Promise uptime you can measure,
and prove you kept it.

Enterprise customers ask for an SLA, an uptime history and evidence that you test capacity and recovery. This guide explains the terms, the arithmetic, what buyers ask for, and what SOC 2's availability criteria expect.

SLISLOSLAError budget99.9%SOC 2 A1.1 to A1.3

Last updated Published by TryTrustableNot legal advice

Short answer

An SLI (service level indicator) is a measurement, such as the share of requests that succeed or the 95th percentile latency. An SLO (service level objective) is your internal target for that measurement, such as 99.9% of requests succeeding over 30 days. An SLA (service level agreement) is the contract with a customer that promises a level of service and says what happens, usually service credits, if you miss it. Set SLOs tighter than the SLA so you get a warning before you owe credits. The error budget is the room the SLO leaves: 99.9% over 30 days allows 43.2 minutes of failure.

01

SLA vs SLO vs SLI: what is the difference?

The three terms come from site reliability engineering. Google's SRE book defines an SLI as a carefully defined quantitative measure of some aspect of the service, an SLO as a target value or range for a service level measured by an SLI, and an SLA as an explicit or implicit contract with users that includes consequences of meeting or missing the SLOs it contains.

SLISLOSLA
What it isA measurementA target for the measurementA contract with a customer
ExampleShare of API requests returning a non-5xx response within 500 ms99.9% of requests meet that bar over 30 days99.9% monthly uptime, or 10% service credit
Who sets itEngineeringEngineering and productSales, legal and leadership, based on the SLO
If missedNothing by itselfError budget spent: slow releases, fix reliabilityCredits or other remedies to the customer

The order matters: choose SLIs that reflect what users feel, set SLOs on them, and only then promise an SLA a little looser than your SLO.

02

How to choose good SLIs

  • Availability: successful requests divided by valid requests, measured at the load balancer or from synthetic checks.
  • Latency: the share of requests faster than a threshold, or a percentile such as p95 or p99. Averages hide the slow tail; see load testing for why.
  • Freshness or throughput for data pipelines and batch jobs: how old the newest data is.
  • Frontend experience: Core Web Vitals such as LCP and INP for the pages users load.
03

What is an error budget?

An error budget is 1 minus the SLO. A 99.9% availability SLO leaves a 0.1% error budget: over a 30-day month that is 43.2 minutes of full downtime, or the equivalent in failed requests. While budget remains, the team ships freely. When it runs out, Google's example policy halts releases other than urgent and security fixes until the service is back within its SLO. The budget turns the argument between speed and reliability into a number both sides agreed in advance.

04

Uptime percentages and downtime per month

AvailabilityDowntime per dayper weekper 30-day monthper year (365 days)
99%14m 24s1h 40m 48s7h 12m3d 15h 36m
99.5%7m 12s50m 24s3h 36m1d 19h 48m
99.9%1m 26s10m 5s43m 12s8h 45m 36s
99.95%43s5m 2s21m 36s4h 22m 48s
99.99%9s1m4m 19s52m 34s
99.999%1s6s26s5m 15s

Computed as (1 minus availability) times the period. Calendar months vary, so SLAs should say how a month is measured.

Each extra nine cuts allowed downtime by ten. A small team can usually run 99.9% with good deploy practices and alerting; 99.99% generally needs redundancy across zones, automated failover and on-call response measured in minutes. Do not sign for more nines than your architecture and your cloud provider's own SLAs support.

05

What enterprise customers ask for in an SLA

What enterprise buyers askWhat a reasonable answer looks like
Uptime commitmentA monthly availability percentage, commonly 99.9% for business SaaS, on the paid tiers that need it
Definition of downtimeWhat counts (error rates, failed requests, a down endpoint) and how it is measured
ExclusionsScheduled maintenance with notice, customer-caused issues, force majeure, beta features
Service creditsA table of credits by availability band, claimed within a set period, as the sole remedy or not
Support response timesFirst response targets by severity, separate from uptime
Status and incident communicationA public status page, incident notices and post-incident reports
EvidenceHistorical uptime, load and performance test results, backup and recovery tests, and a SOC 2 report
Recovery objectivesRTO (how fast you restore) and RPO (how much data you could lose)
Termination rightThe right to terminate after repeated or prolonged SLA breaches

Common terms in enterprise SaaS contracts, not a template. Have counsel review your SLA.

06

SOC 2 availability criteria (A1.1 to A1.3) and the evidence auditors expect

If your SOC 2 report includes the Availability category, three additional criteria apply on top of the common criteria. Auditors want to see that the controls operated over the period, so a single test the week before fieldwork is weak evidence; a schedule of tests with results against stated targets is strong.

CriterionWhat it asks (summarised)Evidence auditors commonly expect
A1.1 CapacityMaintain, monitor and evaluate processing capacity and use of system components, to manage demand and add capacityCapacity and utilisation monitoring, alert thresholds, load or performance test results against targets, capacity reviews and the changes they triggered
A1.2 Environmental protections, backup and recovery infrastructureAuthorise, implement, operate and monitor environmental protections, software, data backup processes and recovery infrastructureCloud provider SOC reports for physical protections, backup configuration and failure alerts, offsite or cross-region copies, redundancy and failover design
A1.3 Recovery testingTest recovery plan procedures that support system recoveryBusiness continuity and disaster recovery tests with results, backup restore tests, and plan updates made from them

Criteria from the AICPA 2017 Trust Services Criteria (revised points of focus, 2022), Additional Criteria for Availability. Availability is an optional category in a SOC 2 report.

Read more in the SOC 2 guide and Type 1 vs Type 2.

07

How TryTrustable helps with availability evidence

TryTrustable runs API load tests and frontend audits on a schedule (hourly, daily or on deploy), checks each run against the SLO you set (p95 latency, error rate, minimum throughput, or a Lighthouse score), and stores the result, pass or fail, as evidence mapped to SOC 2 A1.1. That gives you the record auditors and buyers ask for: a target set in advance and results over time.

It is not an uptime monitor or a status page for your service, and it does not calculate SLA credits. Pair it with your monitoring and incident tooling, and keep recovery test records for A1.2 and A1.3 in the evidence ledger. The enterprise readiness guide covers the rest of the buyer's checklist.

Questions

The things people ask us

How much downtime is 99.9% uptime?

43.2 minutes in a 30-day month, about 10 minutes a week, and 8 hours 45 minutes 36 seconds in a 365-day year.

How much downtime is 99.99% uptime?

About 4 minutes 19 seconds in a 30-day month and 52 minutes 34 seconds in a 365-day year.

Should my SLO be the same as my SLA?

No. Set the SLO tighter than the SLA, for example a 99.95% SLO behind a 99.9% SLA, so you see the problem and act before you owe service credits.

Is an SLA required for SOC 2?

No. SOC 2 does not require an SLA. If you include the Availability category, the criteria are judged against your own commitments, which often come from your SLA, so the two should match.

What uptime should a startup offer?

Many business SaaS vendors offer 99.9% monthly on paid plans. Offer only what your architecture, monitoring and your cloud provider's own SLAs support, and define downtime and exclusions clearly.

Does TryTrustable monitor our uptime?

No. It runs scheduled load tests and frontend audits against your SLO and stores the results as availability evidence. Use a dedicated monitor for uptime and alerting.

Book a walkthrough

Show buyers the SLO, the tests and the results.

Thirty minutes: set an SLO on one of your endpoints, run a load test, and see the result land as SOC 2 A1.1 evidence.