SLA levels

Three nines (99.9%) uptime: downtime per day, month and year

99.9% allows 43m 12s of downtime per month and 8h 45m 36s per year.

Three nines: the SaaS default

99.9% allows 1 minute 26 seconds of downtime per day and 10 minutes 5 seconds per week. It is the de-facto standard SLA for SaaS, and it is what most vendors put in their standard terms, because it can be met with solid monitoring, on-call coverage, and disciplined deploys, without exotic infrastructure.

The yearly figure is the trap. Nearly nine hours sounds generous, but outages cluster: one corrupted deploy at 2 a.m., or a database failover that does not fire, can spend most of the year’s allowance in a single evening and leave no margin for the eleven months that follow.

Your error budget: 43m 12s a month.
Get paged the second you start spending it.
An SLA is a promise. Who tells you when you break it?
Hyperping checks every 30 seconds from 18 locations and alerts you before the budget is gone: Slack, SMS, phone call, PagerDuty.
Start monitoring free
Free plan · no credit card · first check in under a minute
Promising 99.9% to customers? Show it on a status page they can see, from $29/mo with monitoring included →

What a 3 9s target buys you

“3 9s” counts the nines from the left, so 99.9% has three. Add one more and the allowance drops tenfold, which is why each successive nine demands disproportionately more redundancy and automation.

At this tier, teams running on error budgets get real freedom: routine deploys, schema migrations, and dependency upgrades all fit comfortably as long as incidents stay short. That balance, credible reliability for customers against operational room for engineering, is exactly why 99.9% became the standard commitment.

How dependencies consume a three-nines target

End-to-end availability multiplies across required dependencies. Two independent components that each deliver 99.95% produce roughly 99.9% together before the application itself contributes any failures.

Allocate the budget across the application, database, cloud platform, and critical vendors instead of giving each one the same target. Components on the request path need headroom above the customer-facing commitment.

Operating inside a three-nines budget

A 43-minute monthly budget is large enough to manage deliberately. Teams can reserve part for deployments and migrations, leave the rest for unplanned incidents, and slow risky changes when the remaining balance gets low.

The target still needs operational discipline: an on-call owner, rollback paths, and incident reviews for recurring failures. Once the budget is exhausted, reliability work should take priority until the service returns to a sustainable trend.

What users experience inside a 99.9% month

Forty-three minutes can appear as one obvious incident or as dozens of short interruptions. The percentage treats both patterns equally, but repeated failed requests, reconnects, and timeouts often feel worse to regular users.

Pair the monthly percentage with incident count, longest incident, and time to recovery. Those measures reveal whether the service had one contained failure or remained unstable throughout the month.

Downtime allowed at each SLA level

UptimePer dayPer monthPer year
99% Two nines14m 24s 7h 12m 3d 15h
99.9% Three nines1m 26s 43m 12s 8h 45m 36s
99.95%43s 21m 36s 4h 22m 48s
99.99% Four nines9s 4m 19s 52m 34s
99.999% Five nines864ms26s 5m 15s
99.9999% Six nines86ms3s 32s
99.9999999% Nine nines86μs2.6ms32ms

SLA calculation cheatsheet

Availability calculation

Availability (%) = (Total Time - Downtime) / Total Time × 100

Example: If a service was down for 7.3 hours in a 30-day month:

  • Total Time = 30 days × 24 hours = 720 hours
  • Availability = (720 - 7.3) / 720 × 100 = 98.99%

Response time SLA

Response Time Compliance (%) = (Responses Within Threshold / Total Responses) × 100

Mean time metrics

  • MTBF (Mean Time Between Failures) = Total Operational Time / Number of Failures
  • MTTR (Mean Time To Repair) = Total Repair Time / Number of Repairs
  • MTTA (Mean Time To Acknowledge) = Total Time to Acknowledge / Number of Incidents

Service credit calculation

Service Credit = (Monthly Service Fee) × (Credit Percentage for SLA Breach)

SLA penalty example

  • If availability drops below 99.9% but remains above 99.0%: 10% credit
  • If availability drops below 99.0%: 25% credit

Track SLAs and downtime metrics at a glance

Hyperping reports on service reliability using data collected by your monitors. Select the reporting period, then review the results in the dashboard or export them.

  • Track uptime, Mean Time to Recovery (MTTR), and SLA compliance for any reporting period.
  • Export the selected period as a CSV file for audits, client reports, spreadsheets, or internal analysis.
  • Review the complete incident history and use filters to narrow the records included in your report.
Monitor reporting analytics
Monitor reporting dashboard showing outages, MTTR, and SLA metrics over time →

SLA & uptime management guides

What is uptime?

Uptime is the amount of time that a service is available and operational, typically expressed as a percentage over a given period such as a month or a year.

What is an SLA?

A service-level agreement (SLA) defines the level of service you expect from a vendor, laying out the metrics by which service is measured, as well as remedies or penalties should agreed-on service levels not be achieved.

How to prevent downtime?

Redundancy, monitoring and alerting are key to ensure a safe and reliable service.

Frequently asked questions

How do you calculate availability?
Availability (%) = (Total Time − Downtime) / Total Time × 100. For example, a service that was down 7.3 hours in a 30-day month (720 hours) has an availability of (720 − 7.3) / 720 × 100 = 98.99%.
How much downtime does 99.9% uptime allow?
At 99.9% uptime, the maximum allowed downtime is 1 minute 26 seconds per day, 43 minutes 12 seconds per 30-day month, and 8 hours 45 minutes 36 seconds per 365-day year.
What's the difference between 99.9% and 99.99% uptime?
Each additional nine reduces allowed downtime tenfold: 99.9% allows about 8 hours 46 minutes of downtime per year, while 99.99% allows just under 53 minutes. Meeting 99.99% generally requires automated failover, since a human response rarely fits inside the budget.
How do I track SLA compliance?
Continuous uptime monitoring measures availability from outside your infrastructure and records every outage. Hyperping checks your endpoints at up to 30-second intervals from multiple regions and reports uptime, MTTR, and SLA compliance over any period.
Is 99.9% uptime good enough for production?
It can be a practical baseline for many production services, but it still allows about 43 minutes of downtime in a 30-day month. The right target depends on the impact of an outage and the cost of building more resilience.