Break, fix โ€” what's your uptime?

Two numbers decide how available a system is: how long it runs between failures, and how long each one takes to fix. Set any two of break rate, repair time and availability โ€” the third is forced. It's the operational flip-side of an error budget: not how much you're allowed to break, but how much you actually will.

Between failures ยท MTBFMean Time Between Failures โ€” the average length of time a system runs before the next failure. The bigger it is, the rarer the breakages.
To repair ยท MTTRMean Time To Repair โ€” the average wall-clock time from a failure starting to full recovery: detection, response and fix, not just hands-on work.
AvailabilityThe share of time the system is up and serving โ€” uptime รท (uptime + downtime). It's what an SLO promises and what the 'nines' count.
99.723%

Breaking every 4.3 weeks and taking 2h to recover, you'd run at 99.723% โ€” two nines โ€” about 1d of downtime a year.

Availability = time between failures รท (time between failures + time to repair). MTTR here is wall-clock time to full recovery โ€” detection, response and fix โ€” not just hands-on work.

99.723%Availability ยท two nines
12Failures / year
1dDowntime / year
Per periodFailuresDowntime
day0.033
3m 59s
week0.23
27m 55s
30 days1
2h
quarter3
5h 59m
year12
1d
ChangeAvailabilityDowntime / year
Nownow99.723%
1d
Recover twice as fasthalve MTTR99.8613%
12h 9m
Fail half as oftendouble MTBF99.8613%
12h 9m

When repairs are quick next to the time between them, the two levers move downtime by almost the same amount โ€” so cutting recovery time (better alerting, faster rollback) usually buys more availability per unit of effort than making the system fail half as often.