Mean Time Between Failures (MTBF)

MTBF is the average time a repairable asset operates between successive failures. It is a measure of reliability — a higher MTBF means the asset is more reliable and fails less frequently.

MTBF = Total operating time ÷ Number of failures

Important: Total operating time is the sum of all run time between failures — it excludes time spent in repair. Only count failures that cause loss of function, not planned maintenance stoppages.

Worked example — MTBF

A centrifugal pump ran the following history over 12 months:

EventDurationNotes
Run2,100 hrsNormal operation
Failure 1 (seal leak)12 hrs repairUnplanned failure
Run1,850 hrsNormal operation
Failure 2 (bearing)8 hrs repairUnplanned failure
Run2,600 hrsNormal operation

Total operating time = 2,100 + 1,850 + 2,600 = 6,550 hrs. Number of failures = 2.

MTBF = 6,550 ÷ 2 = 3,275 hours

Mean Time to Repair (MTTR)

MTTR is the average time required to restore an asset after a failure. It measures maintainability — a lower MTTR means repairs are faster, reducing the impact of each failure. MTTR includes: failure detection time + notification time + diagnostic time + parts procurement time + repair time + testing time.

MTTR = Total repair time ÷ Number of failures

Worked example: From the pump above: Total repair time = 12 + 8 = 20 hrs. Number of failures = 2.

MTTR = 20 ÷ 2 = 10 hours

Availability from MTBF and MTTR

Availability (%) = MTBF ÷ (MTBF + MTTR) × 100

Example: Availability = 3,275 ÷ (3,275 + 10) × 100 = 3,275 ÷ 3,285 × 100 = 99.7%

This pump has 99.7% inherent availability — very high, meaning failures account for only 0.3% of total time. However, if MTBF were only 500 hrs (frequent failures), availability would drop significantly.

Calculating from CMMS data

In practice, pull failure records from your CMMS (work orders with failure type = "breakdown" or "unplanned corrective"):

  1. Filter by asset, date range, and work order type = unplanned/corrective
  2. For each work order, capture: date/time of failure (when asset stopped), date/time returned to service
  3. Repair time = returned to service − failure start
  4. Operating time between failures = next failure start − previous return to service
  5. Sum all operating times → Total operating time. Count work orders → Number of failures
⚠

MTBF from CMMS is only as good as your work order data quality. If technicians don't record actual failure times (just "today's date"), your MTBF will be wrong. Require failure timestamp, return-to-service timestamp, and failure mode coding on every breakdown work order.

Using MTBF to set PM intervals

A common application of MTBF is setting preventive maintenance intervals. A general guideline: PM interval = 0.5–0.8 × MTBF. This ensures PM occurs before most failures while avoiding unnecessary early maintenance.

For the pump above with MTBF 3,275 hrs, a PM interval of 0.7 × 3,275 = 2,292 hrs (approximately 3 months at 24/7 operation) would intercept most bearing and seal failures before they cause unplanned downtime.

MTBF vs MTTF

MTBF is for repairable assets. MTTF (Mean Time To Failure) is for non-repairable items (bearings, electronic components, light bulbs) that are replaced rather than repaired. The distinction matters: rolling bearings have an MTTF (they are replaced), while the pump itself has an MTBF (it is repaired and returned to service).

Improving MTBF and MTTR

To increase MTBF (reduce failure frequency)

  • Implement PdM to catch failures before they occur
  • Address repeat failures with root cause analysis
  • Improve lubrication practices and contamination control
  • Eliminate the root cause of the top 3 recurring failure modes

To reduce MTTR (faster repair)

  • Pre-position critical spare parts
  • Develop standard repair procedures for top failure modes
  • Improve technician skill through training and job shadowing
  • Install quick-release couplings and access panels to reduce disassembly time
  • Implement a rotating spare strategy for critical assets