SLA levels

Two nines (99%) uptime: how much downtime is that?

99% allows 7h 12m of downtime per month and 3d 15h per year.

Two nines: a 99% baseline

99% sounds high until you convert it: 14 minutes 24 seconds of allowed downtime every single day. That is a service users will regularly catch being down, not an edge case they hear about second-hand.

Two nines is a reasonable target for internal tools, batch pipelines, and staging environments, where an outage inconveniences colleagues rather than customers. Any paying, customer-facing product should start at three nines.

Your error budget: 7h 12m a month.
Get paged the second you start spending it.
An SLA is a promise. Who tells you when you break it?
Hyperping checks every 30 seconds from 18 locations and alerts you before the budget is gone: Slack, SMS, phone call, PagerDuty.
Start monitoring free
Free plan · no credit card · first check in under a minute
Promising 99% to customers? Show it on a status page they can see, from $29/mo with monitoring included →

What a 2 9s target buys you

“2 9s” is the shorthand: the number of leading nines in the percentage gives the nickname, and each extra nine cuts the allowance by a factor of ten. That makes the jump from 2 9s to 3 9s the cheapest reliability win you will ever buy, and every nine after it roughly ten times harder.

In SRE terms a 2 9s objective hands you an enormous error budget, which is genuinely useful when velocity matters more than availability: experimental features, internal services, anything with no external promise attached. It has no place in a customer-facing SLA.

How a 99% allowance accumulates

The same 1% allowance becomes 7 hours 12 minutes in a 30-day month and 3 days 15 hours across a 365-day year. A handful of long incidents can consume it, but so can small failures that recur every day.

Track the total duration and the incident count together. A service can remain inside its percentage while frequent short interruptions still make it feel unreliable to users.

When a 99% target is a deliberate choice

A large error budget can be rational when a service is used only during office hours, has a manual workaround, or can catch up asynchronously after an interruption. In those cases, paying for redundancy may create less value than faster product work.

Make the trade explicit: document supported hours, the recovery owner, and the longest tolerable interruption. If another production service depends on the system, its 99% target becomes a ceiling for that downstream service too.

Downtime allowed at each SLA level

UptimePer dayPer monthPer year
99% Two nines14m 24s 7h 12m 3d 15h
99.9% Three nines1m 26s 43m 12s 8h 45m 36s
99.95%43s 21m 36s 4h 22m 48s
99.99% Four nines9s 4m 19s 52m 34s
99.999% Five nines864ms26s 5m 15s
99.9999% Six nines86ms3s 32s
99.9999999% Nine nines86μs2.6ms32ms

SLA calculation cheatsheet

Availability calculation

Availability (%) = (Total Time - Downtime) / Total Time × 100

Example: If a service was down for 7.3 hours in a 30-day month:

  • Total Time = 30 days × 24 hours = 720 hours
  • Availability = (720 - 7.3) / 720 × 100 = 98.99%

Response time SLA

Response Time Compliance (%) = (Responses Within Threshold / Total Responses) × 100

Mean time metrics

  • MTBF (Mean Time Between Failures) = Total Operational Time / Number of Failures
  • MTTR (Mean Time To Repair) = Total Repair Time / Number of Repairs
  • MTTA (Mean Time To Acknowledge) = Total Time to Acknowledge / Number of Incidents

Service credit calculation

Service Credit = (Monthly Service Fee) × (Credit Percentage for SLA Breach)

SLA penalty example

  • If availability drops below 99.9% but remains above 99.0%: 10% credit
  • If availability drops below 99.0%: 25% credit

Track SLAs and downtime metrics at a glance

Hyperping reports on service reliability using data collected by your monitors. Select the reporting period, then review the results in the dashboard or export them.

  • Track uptime, Mean Time to Recovery (MTTR), and SLA compliance for any reporting period.
  • Export the selected period as a CSV file for audits, client reports, spreadsheets, or internal analysis.
  • Review the complete incident history and use filters to narrow the records included in your report.
Monitor reporting analytics
Monitor reporting dashboard showing outages, MTTR, and SLA metrics over time →

SLA & uptime management guides

What is uptime?

Uptime is the amount of time that a service is available and operational, typically expressed as a percentage over a given period such as a month or a year.

What is an SLA?

A service-level agreement (SLA) defines the level of service you expect from a vendor, laying out the metrics by which service is measured, as well as remedies or penalties should agreed-on service levels not be achieved.

How to prevent downtime?

Redundancy, monitoring and alerting are key to ensure a safe and reliable service.

Frequently asked questions

How do you calculate availability?
Availability (%) = (Total Time − Downtime) / Total Time × 100. For example, a service that was down 7.3 hours in a 30-day month (720 hours) has an availability of (720 − 7.3) / 720 × 100 = 98.99%.
How much downtime does 99% uptime allow?
At 99% uptime, the maximum allowed downtime is 14 minutes 24 seconds per day, 7 hours 12 minutes per 30-day month, and 3 days 15 hours 36 minutes per 365-day year.
What's the difference between 99.9% and 99.99% uptime?
Each additional nine reduces allowed downtime tenfold: 99.9% allows about 8 hours 46 minutes of downtime per year, while 99.99% allows just under 53 minutes. Meeting 99.99% generally requires automated failover, since a human response rarely fits inside the budget.
How do I track SLA compliance?
Continuous uptime monitoring measures availability from outside your infrastructure and records every outage. Hyperping checks your endpoints at up to 30-second intervals from multiple regions and reports uptime, MTTR, and SLA compliance over any period.
When is 99% uptime acceptable?
A 99% target can suit noncritical internal tools, prototypes, or services with a workable fallback. Customer-facing and revenue-critical services usually need a tighter target because 99% permits several hours of downtime per month.