OEE & production

MTBF and MTTR: how to calculate them and where they mislead

Definitions, formulas and a worked example for MTBF and MTTR, plus the pitfalls: what counts as a failure, what counts as repair time and why averages hide long outages.

Two averages, two different questions

MTBF (mean time between failures) answers “how long does this equipment typically run before it fails?”. MTTR (mean time to repair) answers “once it has failed, how long until it is working again?”. Together they describe how often a machine stops and how costly each stop is.

For repairable equipment the usual formulas are:

  • MTBF = operating time ÷ number of failures
  • MTTR = total repair time ÷ number of repairs

Operating time here means time the equipment was actually asked to run, so planned stops and idle time stay out of it. Whether “repair time” includes diagnosis, waiting for a spare part, the physical repair and the restart test is a decision you have to make once and apply everywhere. Some organizations read the M in MTTR as “mean time to restore” or “to recover”; the formula is the same, the scope of the clock is what changes.

A worked example

Take the shift from our OEE worked example: 432 minutes of run time with 2 failures and 30 minutes of total repair time.

  • MTBF = 432 ÷ 2 = 216 minutes
  • MTTR = 30 ÷ 2 = 15 minutes

From these two numbers you can estimate the share of time the machine is available when only failures are considered: MTBF ÷ (MTBF + MTTR) = 216 ÷ 231 ≈ 93.5 %. That is higher than the 90 % shift availability, because the shift lost 48 minutes in total but only 30 of them were repair time. The other 18 minutes were stopped for reasons that do not count as failures under this definition, for example waiting or a changeover. This is why a failure-based figure and a time-based availability figure should never be compared without stating what each one counts.

What counts as a failure?

Most disputes about MTBF are really disputes about this question. Decide in writing, before collecting data:

EventTypical treatment
Mechanical, electrical or control breakdownFailure; its duration is repair time
Planned maintenance or changeoverNot a failure; it belongs to planned downtime
Short stop cleared by the operatorYour decision: failure above a stated threshold, otherwise a minor stop
Waiting for material or an upstream machineUsually not an equipment failure, but it still reduces availability
Quality rejects with the machine runningNot a failure; it is a quality loss

The treatment column is a common convention, not a standard. What matters is that every line and every shift uses the same rules.

Why averages mislead

A mean hides the shape of the data. Consider two machines that each stopped eleven times:

  • Machine A: ten stops of 1 minute and one stop of 90 minutes → MTTR = 100 ÷ 11 ≈ 9.1 minutes.
  • Machine B: eleven stops of 9 minutes → MTTR = 99 ÷ 11 = 9.0 minutes.

The MTTR is almost identical, yet the operational reality is completely different: Machine A has a rare, severe problem and Machine B has a steady, routine one. Report the median and the longest event alongside the mean, and look at the list of events, not only the average.

Time base and sample size

  • Use run time, not calendar time. A machine that is switched off for the weekend did not “survive” those hours.
  • Do not trust very small samples. One failure in a week gives an MTBF of a week, which says little about the next month. Accumulate data over a longer period or across identical machines.
  • Keep units consistent. Minutes in one system and hours in another is a classic source of 60× errors.
  • Beware of changing conditions. A new product, a worn tool or a different operator crew changes the underlying failure pattern; mixing them into one average hides the change.

Using MTBF and MTTR with OEE

OEE tells you how much production time was lost; MTBF and MTTR help explain the stop-time part of it. A low MTBF points toward reliability work: why does it stop so often? A high MTTR points toward maintainability and logistics: why does each repair take so long, are spare parts and diagnostic information at hand? The OEE calculator computes both from the failure count and repair time you enter, and the guide on the six hidden losses shows where the remaining minutes go.

THE NEXT STEP

From calculation to implementation.

Let’s look at your machine, your data flow or your production goal together. Describe your situation in a few sentences and the ASP Dijital team will reply by email.

Talk to ASP Dijital