SLA levels

99.99% uptime (four nines): downtime per day, month and year

99.99% allows 4m 19s of downtime per month and 52m 34s per year.

What 99.99% demands of your architecture

A 99.99% commitment turns reliability into an architectural requirement. Run redundant instances across isolated failure domains, connect them through automated failover backed by continuous health detection, and make zero-downtime deployments with instant rollback the only deployment path. Human intervention can investigate a failure, but it cannot be the first recovery mechanism.

End-to-end availability is capped by the weakest dependency. A four-nines service cannot rely on a three-nines database, queue, DNS provider, or authentication service and still keep its promise. AWS applies its 99.99% EC2 commitment at the Region level, while the comparable Azure commitment depends on a zone-redundant deployment within a region. The scope of the infrastructure commitment must match the customer-facing SLA.

Your error budget: 4m 19s a month.
Get paged the second you start spending it.
An SLA is a promise. Who tells you when you break it?
Hyperping checks every 30 seconds from 18 locations and alerts you before the budget is gone: Slack, SMS, phone call, PagerDuty.
Start monitoring free
Free plan · no credit card · first check in under a minute
Promising 99.99% to customers? Show it on a status page they can see, from $29/mo with monitoring included →

What it costs to actually measure it

The measurement window is as demanding as the architecture. A five-minute check interval is longer than the entire monthly downtime budget at 99.99%, so an outage that begins and ends between two probes can consume the allowance without ever being recorded.

Use sub-minute checks from multiple regions, then require independent confirmation before declaring an outage. Regional observations separate a service failure from a problem with one probe or network path, while confirmation filters false positives without sacrificing detection speed.

Why four nines changes the incident clock

The annual allowance looks almost like an hour, but most SLAs judge each month separately. In a 30-day window, one five-minute incident breaches 99.99% even if every other month is perfect.

Detection, confirmation, and mitigation therefore have to finish before a typical human acknowledgement. Track those three intervals separately: a fast alert is not enough if diagnosis or traffic switching still takes several minutes.

How four nines compounds across dependencies

Three independent services that each deliver exactly 99.99% produce only about 99.97% availability when every request needs all three. A customer-facing four-nines promise therefore requires critical dependencies to exceed four nines or fail independently behind a fallback.

Map the complete request path, including DNS, authentication, queues, and third-party APIs. A strong compute SLA cannot compensate for a weaker required service elsewhere in the chain.

Where four nines is worth paying for

Four nines is easiest to justify where every unavailable minute has a direct cost, such as payment authorization, authentication, or a core customer API. Start with the critical transaction instead of assigning the same target to every component.

Supporting dashboards, exports, and administrative workflows can often stay one nine lower without weakening that transaction. A narrow promise is cheaper to engineer, easier to measure, and more credible than a blanket claim for the whole product.

How to spend a four-nines error budget

Planned changes, brief dependency failures, and customer-visible deploy errors all draw from the same four-minute monthly allowance unless the contract excludes them. Budgeting only for major incidents creates a false sense of safety.

Reserve headroom for failures outside your control and make routine changes invisible through canaries, traffic draining, and instant rollback. A service operating close to the limit has no safe margin for the next surprise.

How to prove a 99.99% commitment

Evidence should include raw check history, incident start and end rules, maintenance records, and the exact endpoints covered by the promise. A monthly percentage without those inputs cannot be reproduced or challenged.

Define how partial and regional failures count, then publish the same calculation customers will use. Transparent scope and historical results make a four-nines claim stronger than a more generous credit attached to an unverifiable report.

Downtime allowed at each SLA level

UptimePer dayPer monthPer year
99% Two nines14m 24s 7h 12m 3d 15h
99.9% Three nines1m 26s 43m 12s 8h 45m 36s
99.95%43s 21m 36s 4h 22m 48s
99.99% Four nines9s 4m 19s 52m 34s
99.999% Five nines864ms26s 5m 15s
99.9999% Six nines86ms3s 32s
99.9999999% Nine nines86μs2.6ms32ms

SLA calculation cheatsheet

Availability calculation

Availability (%) = (Total Time - Downtime) / Total Time × 100

Example: If a service was down for 7.3 hours in a 30-day month:

  • Total Time = 30 days × 24 hours = 720 hours
  • Availability = (720 - 7.3) / 720 × 100 = 98.99%

Response time SLA

Response Time Compliance (%) = (Responses Within Threshold / Total Responses) × 100

Mean time metrics

  • MTBF (Mean Time Between Failures) = Total Operational Time / Number of Failures
  • MTTR (Mean Time To Repair) = Total Repair Time / Number of Repairs
  • MTTA (Mean Time To Acknowledge) = Total Time to Acknowledge / Number of Incidents

Service credit calculation

Service Credit = (Monthly Service Fee) × (Credit Percentage for SLA Breach)

SLA penalty example

  • If availability drops below 99.9% but remains above 99.0%: 10% credit
  • If availability drops below 99.0%: 25% credit

Track SLAs and downtime metrics at a glance

Hyperping reports on service reliability using data collected by your monitors. Select the reporting period, then review the results in the dashboard or export them.

  • Track uptime, Mean Time to Recovery (MTTR), and SLA compliance for any reporting period.
  • Export the selected period as a CSV file for audits, client reports, spreadsheets, or internal analysis.
  • Review the complete incident history and use filters to narrow the records included in your report.
Monitor reporting analytics
Monitor reporting dashboard showing outages, MTTR, and SLA metrics over time →

SLA & uptime management guides

What is uptime?

Uptime is the amount of time that a service is available and operational, typically expressed as a percentage over a given period such as a month or a year.

What is an SLA?

A service-level agreement (SLA) defines the level of service you expect from a vendor, laying out the metrics by which service is measured, as well as remedies or penalties should agreed-on service levels not be achieved.

How to prevent downtime?

Redundancy, monitoring and alerting are key to ensure a safe and reliable service.

Frequently asked questions

How do you calculate availability?
Availability (%) = (Total Time − Downtime) / Total Time × 100. For example, a service that was down 7.3 hours in a 30-day month (720 hours) has an availability of (720 − 7.3) / 720 × 100 = 98.99%.
How much downtime does 99.99% uptime allow?
At 99.99% uptime, the maximum allowed downtime is 9 seconds per day, 4 minutes 19 seconds per 30-day month, and 52 minutes 34 seconds per 365-day year.
What's the difference between 99.9% and 99.99% uptime?
Each additional nine reduces allowed downtime tenfold: 99.9% allows about 8 hours 46 minutes of downtime per year, while 99.99% allows just under 53 minutes. Meeting 99.99% generally requires automated failover, since a human response rarely fits inside the budget.
How do I track SLA compliance?
Continuous uptime monitoring measures availability from outside your infrastructure and records every outage. Hyperping checks your endpoints at up to 30-second intervals from multiple regions and reports uptime, MTTR, and SLA compliance over any period.
What does four nines uptime require?
Four nines usually requires redundancy, automated failover, careful deployments, and fast detection because the monthly downtime budget is only a few minutes. Every dependency in the request path must support the target.