MTBF stands for Mean Time Between Failures — the average time a machine operates between one breakdown and the next. A higher MTBF means the machine is more reliable and failing less often. It is the primary metric for measuring whether a preventive maintenance program is working.
Most of what you will find online about MTBF frames it as a server metric — an IT reliability concept for measuring uptime in a data center. That framing is not wrong, but it is not what a maintenance team at a press shop or an injection molding plant needs. This guide is for factory maintenance: hydraulic presses, CNC lathes, compressors, conveyor belts. Breakdowns per shift, not incidents per deployment. Technicians with wrench sets, not on-call engineers with laptops.
MTBF full form and definition
MTBF stands for Mean Time Between Failures.
- Mean: the average across multiple occurrences
- Time: measured in operating hours (not calendar hours — more on this below)
- Between Failures: the interval from one breakdown to the next, after the machine has been repaired and returned to service
The definition implies something important: MTBF only applies to repairable assets. A machine breaks down, gets fixed, runs again, breaks down again. MTBF measures how long the "runs again" phase lasts on average. This is why MTBF is the metric for production equipment — presses, lathes, molders, compressors — rather than for single-use components like fuses or drive belts that are simply replaced.
A higher MTBF is better. A machine with an MTBF of 400 hours is failing less often than one with an MTBF of 80 hours. If MTBF is rising over time, the preventive maintenance program is working. If it is falling, something is wrong — either with the machine, with the maintenance schedule, or with both.
The MTBF formula
MTBF = Total operating time ÷ Number of failures (in a given period, for a given asset)
The formula is simple. What matters is understanding exactly what goes into each variable.
What counts as "operating time"
Operating time is planned production time minus downtime. It is not calendar time.
A machine that runs two shifts (16 hours/day) over 30 days has around 480 operating hours available. If it was down for unplanned repairs for 20 hours during that period, the actual operating time used in the MTBF calculation is 460 hours.
Why does this matter? Because a machine running two shifts will always have more absolute operating hours than a machine running one shift — and a higher number of failures as a simple consequence of running longer. To compare machines, or to track a single machine over time when operating intensity changes, you must use actual operating hours, not calendar days.
What counts as a "failure"
A failure is an unplanned stop that required maintenance intervention — the machine was running, something went wrong, and a technician had to fix it before production could resume.
What does not count as a failure:
- Planned changeovers (product or tooling changes)
- Scheduled maintenance stops (the monthly PM, the quarterly oil change)
- Planned production breaks (shift changes, weekends, planned downtime)
These are all planned stops. MTBF measures unplanned failures — the ones that interrupt production unexpectedly. Including planned stops in the count would penalize a plant for doing its maintenance, which is the opposite of what the metric is designed to measure.
The measurement period
MTBF is calculated per asset, per period — typically monthly or quarterly. Monthly gives you enough data to spot trends quickly. Quarterly is useful for assets that fail rarely (a high-MTBF machine might only show one or two breakdowns per quarter, making monthly calculations unstable).
One machine, one period. Do not average MTBF across your entire plant and report a single number — that number tells you almost nothing useful. The value of MTBF is in the per-asset view: which machine is failing most often, which has improved, which has deteriorated.
Worked factory examples — three machines, one month
Here are three machines from a single production month. Operating hours are actual hours in production (available hours minus all downtime). Breakdowns are unplanned stops requiring maintenance intervention.
| Machine | Operating hours (month) | Breakdowns | MTBF |
|---|---|---|---|
| Hydraulic press P-3 | 420 hrs | 2 | 210 hrs |
| CNC lathe L-7 | 380 hrs | 5 | 76 hrs |
| Injection molder M-2 | 480 hrs | 1 | 480 hrs |
Calculation check:
- P-3: 420 ÷ 2 = 210 hrs
- L-7: 380 ÷ 5 = 76 hrs
- M-2: 480 ÷ 1 = 480 hrs
Reading the numbers
P-3 at 210 hours MTBF means the press fails roughly once every 210 operating hours. Running two 8-hour shifts per day, that is one breakdown approximately every 13 working days — or once every 2–3 weeks in practice. Not ideal, but not the immediate priority.
L-7 at 76 hours MTBF is the problem machine. Five breakdowns in one month, on a lathe that ran 380 hours. That is one failure roughly every 76 operating hours — or at 16 operating hours per day (two shifts), approximately once every 4–5 days. The lathe was stopped and repaired almost every working week. That is not a reliability metric — it is a breakdown cycle.
Five breakdowns in a single month on one asset means the maintenance team spent significant repair time on L-7, production scheduling had to work around five unplanned stops, and the PM program — if one exists for this machine — is clearly not preventing failures.
M-2 at 480 hours MTBF had one breakdown all month. It is the most reliable of the three assets in this period.
The priority order: L-7 first, P-3 second, M-2 watch-and-hold.
L-7's MTBF of 76 hours is not just low in absolute terms — it is low relative to its operating hours. A machine running 380 hours that fails five times is not being maintained effectively. Whether the root cause is worn tooling, inadequate lubrication schedules, a recurring component failure, or simply no PM program in place, the investigation starts here.
What is a good MTBF?
There is no universal benchmark. The honest answer is that MTBF depends on:
- Machine type: a cold-heading machine running at high speed and load has a different expected MTBF than a simple conveyor belt
- Operating environment: an abrasive or dusty environment accelerates wear on seals, bearings, and filters
- Machine age: an older machine with accumulated wear will have a lower MTBF than a newer equivalent, all else equal
- Operating intensity: a machine running three shifts is under more stress than one running one shift, even if both are maintained identically
- Maintenance quality: the frequency and thoroughness of PMs, the quality of spare parts, the skill of the technicians
The useful question is not "is my MTBF good?" but "is my MTBF stable or improving?"
A CNC lathe running three shifts in a machining shop that generates metal swarf and coolant mist will have a lower absolute MTBF than one running a single shift in a clean assembly environment. Both machines can have "good" MTBF numbers — if the PM program is maintaining or moving those numbers upward over time.
What a falling MTBF tells you
If a machine's MTBF is declining month on month — fewer operating hours between each failure — it is a signal, not a coincidence. The questions to ask:
-
Is the PM schedule being followed? Low PM compliance will directly cause a falling MTBF. If the monthly lubrication round is being skipped, or the quarterly inspection is being pushed out because production is busy, the failure frequency will increase.
-
Are the PMs addressing the right failure modes? If PMs are being executed faithfully but MTBF is still falling, the PM checklist may not be targeting the actual causes of failure. The work-order root-cause data for L-7's five breakdowns should tell you what went wrong each time. If three of five failures were the same component, the PM should include a scheduled replacement or inspection of that component.
-
Has the operating environment or intensity changed? A new product run requiring longer machine hours, or a facility change that increases dust or heat, can reduce MTBF without any change in maintenance quality. The change in context needs to be understood before assuming the PM program has failed.
-
Is the machine reaching end-of-life for key sub-assemblies? A machine in service for 15 years may have bearing races, spindle components, or hydraulic cylinders approaching the end of their service life. Increasing failure frequency can be the leading indicator of a major refurbishment need.
I saw this play out on stamping machines at a cycle chain plant. They had a lubricant leak, and every time it caused trouble the team topped the lubricant back up and got the machine running — nobody ever chased the leak itself. Over about four months the gap between failures kept shrinking, and because each individual fix looked quick and successful, nothing in the daily routine flagged it. It ran to a major failure that needed a spare replaced.
MTBF and preventive maintenance — the feedback loop
MTBF is the measure. PM execution is the input. The relationship between them is straightforward: a plant that runs its PMs consistently, on the right assets, with the right tasks, should see MTBF stabilize or improve over a 3–6 month horizon.
But the feedback loop only closes if you measure both things. A plant that tracks MTBF but not PM compliance cannot tell whether a falling MTBF is caused by maintenance neglect or by a genuine change in machine condition. A plant that tracks PM completion rates but not MTBF cannot tell whether the PMs are actually preventing failures or just consuming technician time without effect.
Scenario: a PM program doubles MTBF in six months
Consider a CNC lathe with no scheduled maintenance program. In Q1, it records four breakdowns — MTBF of roughly 90 hours based on 360 operating hours. The maintenance team introduces a monthly PM: lubrication, filter inspection, chuck jaw check, coolant system flush.
In Q2, with the PM running for the first full quarter, breakdowns fall to two. MTBF moves to 185 hours.
In Q3, with the PM embedded and technicians familiar with the checklist, breakdowns fall to one. MTBF reaches 360 hours.
No new equipment. No capital spend. The same lathe, the same technicians, the same product. Scheduled maintenance, executed consistently, doubled the machine's effective reliability in two quarters.
This is the direct value of tracking MTBF per asset over time. Without it, the maintenance team knows it has fewer breakdown calls — but it cannot quantify the improvement, cannot attribute it to the PM program, and cannot justify continuing or expanding that program to plant management.
When PMs are running but MTBF is not improving
If PM compliance is high but MTBF is flat or declining, the PM checklist needs to be reviewed against actual failure causes. Pull the work-order root-cause data for every breakdown on that asset in the last quarter. What failed? The same component three times? A different component each time?
If it is the same component: the PM should include a scheduled replacement or inspection interval for that part. If it is different components each time: the machine may have a systemic issue — alignment, vibration, contamination — that the PM tasks are not addressing.
The work-order history is the diagnostic input. MTBF is the outcome metric that tells you whether the diagnosis and the fix are working.
In my experience the first person to ask for MTBF is usually the plant head, not an auditor. One caution on the review rhythm, though: I would not look at it too often. Failures do not arrive on a neat schedule, and on most assets a short window produces a number that jumps around for reasons that have nothing to do with the machine — widen the window until the trend is telling you something real.
MTBF vs MTTF — the distinction that matters for spare parts
Two terms that are often confused:
MTTF — Mean Time To Failure applies to non-repairable components. A fuse, a single-use sensor, a bearing that is replaced rather than repaired. MTTF measures the expected operating lifespan before that component needs to be replaced. Once it fails, it is gone — there is no "between" because it does not return to service.
MTBF — Mean Time Between Failures applies to repairable assets. The hydraulic press breaks down, gets repaired, and returns to production. MTBF measures the average duration of the operating interval between repair and next failure.
In factory maintenance, MTBF is the metric for production equipment — the machines on your asset register. MTTF matters when you are specifying replacement components and consumables: how often should you stock a particular bearing? What is the expected service life of this seal kit? That is an MTTF question.
The two metrics are sometimes conflated because many MTBF calculations for long-lived equipment look similar to MTTF in practice — especially on assets where catastrophic failure leads to full replacement rather than repair. But the conceptual distinction matters: MTBF assumes the asset returns to service; MTTF does not.
For repair speed — how long a breakdown takes to fix once it has occurred — see the MTTR guide. For the relationship between mean time to repair and mean time between failures, including how to use both together to plan maintenance capacity, see the MTTR vs MTBF guide. For how MTBF sits inside a wider maintenance KPI framework, see the maintenance KPIs guide.
Where MachDatum fits: MTBF per asset updates automatically with every closed breakdown work order — no spreadsheet, no end-of-month calculation. A falling MTBF on a specific machine is visible in the analytics view as soon as the data is there, not at the next quarterly review. Work-order root-cause data for every breakdown feeds directly into the per-asset failure history, so the maintenance team can see not just how often a machine is failing but what is causing it to fail. We're onboarding our first group of manufacturing teams right now — see how it works at machdatum.com/cmms.
