What 99.99% demands of your architecture
A 99.99% commitment turns reliability into an architectural requirement. Run redundant instances across isolated failure domains, connect them through automated failover backed by continuous health detection, and make zero-downtime deployments with instant rollback the only deployment path. Human intervention can investigate a failure, but it cannot be the first recovery mechanism.
End-to-end availability is capped by the weakest dependency. A four-nines service cannot rely on a three-nines database, queue, DNS provider, or authentication service and still keep its promise. AWS applies its 99.99% EC2 commitment at the Region level, while the comparable Azure commitment depends on a zone-redundant deployment within a region. The scope of the infrastructure commitment must match the customer-facing SLA.
What it costs to actually measure it
The measurement window is as demanding as the architecture. A five-minute check interval is longer than the entire monthly downtime budget at 99.99%, so an outage that begins and ends between two probes can consume the allowance without ever being recorded.
Use sub-minute checks from multiple regions, then require independent confirmation before declaring an outage. Regional observations separate a service failure from a problem with one probe or network path, while confirmation filters false positives without sacrificing detection speed.
Why four nines changes the incident clock
The annual allowance looks almost like an hour, but most SLAs judge each month separately. In a 30-day window, one five-minute incident breaches 99.99% even if every other month is perfect.
Detection, confirmation, and mitigation therefore have to finish before a typical human acknowledgement. Track those three intervals separately: a fast alert is not enough if diagnosis or traffic switching still takes several minutes.
How four nines compounds across dependencies
Three independent services that each deliver exactly 99.99% produce only about 99.97% availability when every request needs all three. A customer-facing four-nines promise therefore requires critical dependencies to exceed four nines or fail independently behind a fallback.
Map the complete request path, including DNS, authentication, queues, and third-party APIs. A strong compute SLA cannot compensate for a weaker required service elsewhere in the chain.
Where four nines is worth paying for
Four nines is easiest to justify where every unavailable minute has a direct cost, such as payment authorization, authentication, or a core customer API. Start with the critical transaction instead of assigning the same target to every component.
Supporting dashboards, exports, and administrative workflows can often stay one nine lower without weakening that transaction. A narrow promise is cheaper to engineer, easier to measure, and more credible than a blanket claim for the whole product.
How to spend a four-nines error budget
Planned changes, brief dependency failures, and customer-visible deploy errors all draw from the same four-minute monthly allowance unless the contract excludes them. Budgeting only for major incidents creates a false sense of safety.
Reserve headroom for failures outside your control and make routine changes invisible through canaries, traffic draining, and instant rollback. A service operating close to the limit has no safe margin for the next surprise.
How to prove a 99.99% commitment
Evidence should include raw check history, incident start and end rules, maintenance records, and the exact endpoints covered by the promise. A monthly percentage without those inputs cannot be reproduced or challenged.
Define how partial and regional failures count, then publish the same calculation customers will use. Transparent scope and historical results make a four-nines claim stronger than a more generous credit attached to an unverifiable report.
Downtime allowed at each SLA level
| Uptime | Per day | Per month | Per year |
|---|---|---|---|
| 99% Two nines | 14m 24s | 7h 12m | 3d 15h |
| 99.9% Three nines | 1m 26s | 43m 12s | 8h 45m 36s |
| 99.95% | 43s | 21m 36s | 4h 22m 48s |
| 99.99% Four nines | 9s | 4m 19s | 52m 34s |
| 99.999% Five nines | 864ms | 26s | 5m 15s |
| 99.9999% Six nines | 86ms | 3s | 32s |
| 99.9999999% Nine nines | 86μs | 2.6ms | 32ms |
SLA calculation cheatsheet
Availability calculation
Availability (%) = (Total Time - Downtime) / Total Time × 100Example: If a service was down for 7.3 hours in a 30-day month:
- Total Time = 30 days × 24 hours = 720 hours
- Availability = (720 - 7.3) / 720 × 100 = 98.99%
Response time SLA
Response Time Compliance (%) = (Responses Within Threshold / Total Responses) × 100Mean time metrics
- MTBF (Mean Time Between Failures) = Total Operational Time / Number of Failures
- MTTR (Mean Time To Repair) = Total Repair Time / Number of Repairs
- MTTA (Mean Time To Acknowledge) = Total Time to Acknowledge / Number of Incidents
Service credit calculation
Service Credit = (Monthly Service Fee) × (Credit Percentage for SLA Breach)SLA penalty example
- If availability drops below 99.9% but remains above 99.0%: 10% credit
- If availability drops below 99.0%: 25% credit
Track SLAs and downtime metrics at a glance
Hyperping reports on service reliability using data collected by your monitors. Select the reporting period, then review the results in the dashboard or export them.
- Track uptime, Mean Time to Recovery (MTTR), and SLA compliance for any reporting period.
- Export the selected period as a CSV file for audits, client reports, spreadsheets, or internal analysis.
- Review the complete incident history and use filters to narrow the records included in your report.

SLA & uptime management guides
What is uptime?
Uptime is the amount of time that a service is available and operational, typically expressed as a percentage over a given period such as a month or a year.
What is an SLA?
A service-level agreement (SLA) defines the level of service you expect from a vendor, laying out the metrics by which service is measured, as well as remedies or penalties should agreed-on service levels not be achieved.
How to prevent downtime?
Redundancy, monitoring and alerting are key to ensure a safe and reliable service.
Frequently asked questions
- How do you calculate availability?
- Availability (%) = (Total Time − Downtime) / Total Time × 100. For example, a service that was down 7.3 hours in a 30-day month (720 hours) has an availability of (720 − 7.3) / 720 × 100 = 98.99%.
- How much downtime does 99.99% uptime allow?
- At 99.99% uptime, the maximum allowed downtime is 9 seconds per day, 4 minutes 19 seconds per 30-day month, and 52 minutes 34 seconds per 365-day year.
- What's the difference between 99.9% and 99.99% uptime?
- Each additional nine reduces allowed downtime tenfold: 99.9% allows about 8 hours 46 minutes of downtime per year, while 99.99% allows just under 53 minutes. Meeting 99.99% generally requires automated failover, since a human response rarely fits inside the budget.
- How do I track SLA compliance?
- Continuous uptime monitoring measures availability from outside your infrastructure and records every outage. Hyperping checks your endpoints at up to 30-second intervals from multiple regions and reports uptime, MTTR, and SLA compliance over any period.
- What does four nines uptime require?
- Four nines usually requires redundancy, automated failover, careful deployments, and fast detection because the monthly downtime budget is only a few minutes. Every dependency in the request path must support the target.