Two numbers decide how available a system is: how long it runs between failures, and how long each one takes to fix. Set any two of break rate, repair time and availability โ the third is forced. It's the operational flip-side of an error budget: not how much you're allowed to break, but how much you actually will.
Breaking every 4.3 weeks and taking 2h to recover, you'd run at 99.723% โ two nines โ about 1d of downtime a year.
Availability = time between failures รท (time between failures + time to repair). MTTR here is wall-clock time to full recovery โ detection, response and fix โ not just hands-on work.
| Per period | Failures | Downtime |
|---|---|---|
| day | 0.033 | 3m 59s |
| week | 0.23 | 27m 55s |
| 30 days | 1 | 2h |
| quarter | 3 | 5h 59m |
| year | 12 | 1d |
| Change | Availability | Downtime / year |
|---|---|---|
| Nownow | 99.723% | 1d |
| Recover twice as fasthalve MTTR | 99.8613% | 12h 9m |
| Fail half as oftendouble MTBF | 99.8613% | 12h 9m |
When repairs are quick next to the time between them, the two levers move downtime by almost the same amount โ so cutting recovery time (better alerting, faster rollback) usually buys more availability per unit of effort than making the system fail half as often.