MTBF and MTTR on Paper Container Machinery: A Reliability FAQ
A Portuguese container plant spent three months missing its output target while its maintenance log said exactly nothing useful, because the entries recorded that the line had stopped without recording where or why. Once the plant started logging the station, the symptom and the clock times, it found that one station failed often and briefly and another failed rarely and expensively. Yoco Group supplies and services paper container lines, and reliability data is one of the most common gaps a plant discovers when it tries to improve. This FAQ answers five questions about MTBF and MTTR on container machinery.
What do MTBF and MTTR mean for a paper container line?
Mean time between failures, or MTBF, is the average running time a machine achieves before it fails, and mean time to repair, or MTTR, is the average time taken to restore it to service. Together they describe availability, because availability is the ratio of MTBF to the sum of MTBF and MTTR. On a container line the two numbers behave differently from station to station: a forming station may fail rarely but take a long time to repair, while a packing station may fail often and be cleared in minutes. Tracking both, per station, tells a plant which kind of problem it has, because frequent short stops and rare long repairs call for opposite responses.
How should a plant define a failure before it counts one?
The definition should state three things and then be applied without exception. A failure is an event that stops or slows the counted output; it lasts longer than a stated number of minutes or requires an intervention the operator cannot perform alone; and it is recorded at the moment it happens rather than reconstructed at the end of the shift. Planned stops such as changeovers, cleaning and scheduled maintenance are not failures and should be logged in a separate column, because mixing them with breakdowns inflates MTTR and makes the machine look worse than it is. The parameters can be chosen by the plant, but once chosen they should be printed on the recording card so no shift has to remember them.
TAPPI publishes technical resources for the pulp, paper and converting industries, and its treatment of converting machinery reliability treats the stoppage record as the primary data source for improvement, because it is the only record that reflects the machine as it is actually run.
What makes a stoppage log usable rather than decorative?
A usable log carries three codes on every entry: the station, the symptom or failure mode, and the action taken, together with the clock times at stop and restart. Those fields make it possible to compute MTBF and MTTR per station and per failure mode rather than for the line as a whole, and to rank the losses. Two further fields pay for themselves: the part replaced, which links the log to the stores, and the person who restored the machine, which shows where training would shorten a repair. A log that records only that the line stopped cannot produce any of these numbers, which is why so many plants keep records that they never manage to act on.
Should a plant attack frequent short stops or rare long repairs first?
Most plants should attack frequency before duration, because frequent short stops are usually cheaper to remove and they consume operator attention as well as output. A stop that lasts three minutes and happens every hour is invisible on a daily summary and obvious on a coded log, and it is often fixed with a guide, a clearance adjustment or a short training session. Rare long repairs usually need a different kind of intervention, such as improved access to a component or a change in the repair procedure, and they are worth doing but they move MTTR rather than MTBF. The practical order is to remove the frequent mode first, confirm that the failure count has fallen, and then turn to the long repairs that remain.
How do MTBF and MTTR connect to spare parts and to OEE?
They connect to spare parts through the coded log, because a plant that records which part failed can see which part to stock, and that is more reliable than a store built on guesswork. They connect to OEE because overall equipment effectiveness uses the same stoppage record, so a plant that codes its stops well gets its availability and its reliability data from one consistent source. The mistake is to run two parallel records, one for the maintenance department and one for production, because the two will disagree and the plant will spend its time reconciling them instead of improving the line. One record, one definition, two numbers.
Prepared by 燕七.