Free tool

MTTR calculator: MTTR, MTBF and MTTA formulas with examples

MTTR = total downtime ÷ number of incidents. MTBF = total uptime ÷ number of incidents. Enter your incidents to get both, plus MTTA and the availability they add up to.

MTTR, MTBF and MTTA formulas

Each mean-time metric is a total divided by the number of incidents in the same period. MTTR divides the downtime, MTBF the uptime, MTTA the time it took someone to acknowledge the alert.

MTTR = total downtime ÷ number of incidents
MTBF = (period - total downtime) ÷ number of incidents
MTTA = total time to acknowledge ÷ number of incidents
MTTD = total time to detect ÷ number of incidents
Availability = MTBF ÷ (MTBF + MTTR) × 100

Count only unplanned incidents, and measure downtime from the moment the service failed to the moment it worked again. Planned maintenance windows stay out of both the downtime and the incident count. The MTTR guide breaks recovery time into detection, response, repair and verification.

MTTR and MTBF formula with an example

A team records 4 incidents over a 30-day month. These are the numbers the calculator opens with.

IncidentDetected afterAcknowledged after the alertDowntime
API timeout1 min2 min12 min
Database failover2 min6 min45 min
Expired certificate1 min1 min8 min
Bad deploy4 min3 min55 min
Total8 min12 min120 min
  • Period: 30 days × 1,440 = 43,200 minutes
  • MTTR = 120 ÷ 4 = 30 minutes
  • MTBF = (43,200 - 120) ÷ 4 = 10,770 minutes, about 179.5 hours
  • MTTA = 12 ÷ 4 = 3 minutes
  • MTTD = 8 ÷ 4 = 2 minutes
  • Availability = 10,770 ÷ (10,770 + 30) = 99.72%

Half of the downtime comes from one incident, the bad deploy. Averages hide that, so read MTTR next to the incident list it came from.

How to calculate availability from MTBF and MTTR

Availability = MTBF ÷ (MTBF + MTTR). It gives the same number as uptime ÷ total time, seen from the incident side: you raise it by failing less often (a longer MTBF) or by recovering faster (a shorter MTTR). The table keeps 4 incidents in a 30-day month and changes only the recovery time.

MTTRAvailabilityDowntime per monthMTBF
5m99.95%20m7d 12h
15m99.86%1h7d 12h
30m99.72%2h7d 12h
1h99.44%4h7d 11h
2h98.89%8h7d 10h

MTBF vs MTTF

MTTF (Mean Time To Failure) is MTBF for things you replace instead of repairing, like a disk or a power supply: operating time ÷ number of failures. A server you restart has an MTBF, the disk you swap out of it has an MTTF.

Some reliability texts define MTBF differently, as the full cycle from one failure to the next: MTBF = MTTF + MTTR. Availability then becomes MTTF ÷ MTBF. The result is the same, so check which convention a vendor or a contract uses before comparing numbers.

What is a good MTTR?

For customer-facing web services and APIs, under an hour is the usual target, and the DORA research on software delivery counts restoring service in less than an hour as elite performance. Before working on the fix itself, look at MTTD and MTTA: in the example above, 5 of the 30 minutes went by before anyone owned the incident.

  • Detect sooner with frequent checks (every 30 seconds rather than every 5 minutes) from several regions, so a failure is confirmed before customers report it.
  • Acknowledge sooner with an escalation policy that moves an unanswered alert to the next person, and channels that interrupt (push, SMS, phone call).
  • Recover sooner next time by writing down what happened: the post-mortem generator gives you the structure.

Get MTTR and MTTA without a spreadsheet

Hyperping records every outage your monitors detect, with its start, its end and who acknowledged the alert. Its reports turn that into MTTR and MTTA for any period, across all monitors or per monitor, and export the incident list as CSV.

  • MTTR from the duration of each outage, with the total downtime and incident count next to it.
  • MTTA from the time between the outage starting and a responder acknowledging it, with who acknowledged and which escalation policy ran.
  • MTTR in the Reports API, and MTTR and MTTA in the MCP server, so an AI assistant can answer "which monitor has the worst MTTR this month?".
Hyperping reporting dashboard with total incidents, total downtime, MTTR and the incident table
How Hyperping reports MTTR and MTTA →

Frequently asked questions

What is the formula for MTTR and MTBF?
MTTR = total downtime ÷ number of incidents. MTBF = total uptime ÷ number of incidents, where uptime is the observation period minus the downtime. A 30-day month (43,200 minutes) with 4 incidents and 120 minutes of downtime gives an MTTR of 30 minutes and an MTBF of 10,770 minutes, about 7.5 days.
How do you calculate MTTR and MTBF with an example?
Add up the downtime of every incident in the period and divide by the number of incidents to get MTTR. Subtract that downtime from the period to get uptime, then divide by the same number of incidents to get MTBF. Four outages of 12, 45, 8 and 55 minutes in a 30-day month give MTTR = 120 ÷ 4 = 30 minutes and MTBF = (43,200 - 120) ÷ 4 = 10,770 minutes.
How do you calculate availability from MTBF and MTTR?
Availability = MTBF ÷ (MTBF + MTTR). With an MTBF of 10,770 minutes and an MTTR of 30 minutes, availability is 10,770 ÷ 10,800 = 99.72%. It is the same number as uptime ÷ total time, which is why cutting MTTR raises availability even when incidents happen just as often.
What is MTTA and how is it calculated?
MTTA (Mean Time To Acknowledge) is the average time between an alert firing and a responder taking ownership of it. MTTA = total time to acknowledge ÷ number of incidents: acknowledgements after 2, 6, 1 and 3 minutes give an MTTA of 3 minutes. A high MTTA usually points to alert fatigue or a gap in the on-call schedule.
What is the difference between MTBF and MTTF?
MTBF applies to systems you repair and put back into service, MTTF (Mean Time To Failure) to components you replace, like a disk. Both are operating time ÷ number of failures. Some reliability texts define MTBF as MTTF + MTTR, the time from one failure to the next; availability is then MTTF ÷ MTBF, which gives the same result.
What is a good MTTR?
It depends on what the service does, but for customer-facing web services and APIs, teams usually aim for under an hour, and the DORA research on software delivery counts restoring service in less than an hour as elite performance. Most of a long MTTR is often spent before anyone works on the problem: detection and acknowledgement are the first places to look.