Origins of RCM
RCM was developed in the late 1960s and early 1970s for the US airline industry. The FAA commissioned a study (Nowlan & Heap, "Reliability-Centred Maintenance", 1978) that fundamentally changed how maintenance was understood.
The key finding: time-based overhaul does not improve the reliability of most equipment. Analysis of failure data showed that only ~11% of failures followed an age-related pattern (the familiar "bathtub curve" right tail). The majority showed random or infant-mortality patterns — meaning scheduled overhaul either had no effect or actively introduced failures through the infant-mortality phase.
The insight led to a new philosophy: maintenance tasks should only be done if they are technically feasible and if they will actually produce the desired results given the asset's failure mode characteristics.
The seven RCM questions
Every RCM analysis works through these seven questions for each asset system:
- What are the functions and associated desired performance standards of the asset in its current operating context?
Define what the asset is supposed to do and how well — primary functions and secondary functions (protection, containment, structural support, etc.). - In what ways can the asset fail to fulfil its functions?
Functional failures — loss of function, reduced function, unintended function. - What causes each functional failure?
Failure modes — the specific events that cause each functional failure (e.g., "impeller worn due to abrasive service"). - What happens when each failure occurs?
Failure effects — what actually happens in the physical world when this failure mode occurs. - In what way does each failure matter?
Failure consequences — the category of impact: safety & environment, operational, or non-operational. - What should be done to predict or prevent each failure?
Proactive task selection — scheduled restoration, scheduled discard, on-condition (PdM), or failure-finding tasks. - What should be done if a suitable proactive task cannot be found?
Default actions — redesign, run-to-fail with consequence acceptance, or one-time changes to eliminate the failure mode.
Failure consequence categories
RCM classifies failure modes into four consequence categories. The category determines the minimum acceptable level of maintenance intervention:
| Category | Definition | Task selection priority |
|---|---|---|
| Safety & environmental | Failure could hurt or kill someone, or breach an environmental regulation, regardless of redundancy | Must find a proactive task or redesign — run-to-fail is not acceptable |
| Hidden failures | Failure is not evident to the operating crew under normal conditions — failure only becomes visible when a protected function is demanded | Failure-finding tasks at intervals short enough to achieve acceptable availability of the protective function |
| Evident operational | Failure is evident to the operating crew and has an operational consequence (production loss, quality impact) | Proactive task is worthwhile only if cost is less than operational consequence over time |
| Evident non-operational | Failure is evident but has no operational consequence — only the direct repair cost | Proactive task worthwhile only if cost is less than direct repair cost over time |
Task selection logic
For each failure mode, RCM asks whether a proactive maintenance task is technically feasible (can it detect or prevent the failure?) and worth doing (does the cost/benefit justify it given the failure consequences?). The decision logic produces one of four outcomes:
Scheduled on-condition tasks (predictive maintenance)
Applicable when: the failure has a detectable P-F interval long enough to act on. On-condition tasks monitor asset condition and trigger intervention only when needed (vibration monitoring, oil analysis, thermography). Preferred task type when applicable because it avoids unnecessary disassembly and infant-mortality risk.
Scheduled restoration tasks
Applicable when: a component has a clear age-related failure pattern and the failure consequences justify the restoration cost. The component is restored to its original condition at a fixed interval (e.g., overhaul of a hydraulic valve at 10,000 hours).
Scheduled discard tasks
Applicable when: a component has a clear safe-life limit below which failure probability is acceptably low. The component is discarded and replaced at a fixed interval (e.g., replace timing belt at 60,000 km regardless of condition).
Default actions
When no proactive task is technically feasible or cost-effective: either accept run-to-fail (for non-safety, non-critical failures), redesign to eliminate the failure mode, or provide contingency resources (spare parts, operator training) to manage the consequence when failure occurs.
RCM and FMEA
RCM analysis uses a Failure Mode and Effects Analysis (FMEA) as its core tool — but with a specific focus on maintenance decision-making rather than design risk. The RCM FMEA typically captures:
- Asset system and sub-system
- Function (numbered)
- Functional failure
- Failure mode (cause)
- Failure effect (what happens)
- Failure consequence category
- Current maintenance task (if any)
- Proposed task (from decision logic)
- Task interval
- Task owner
A full RCM FMEA for a complex plant system can have hundreds to thousands of failure modes. Scope management is critical — focus on significant items first.
Implementation steps
- Select the system/asset boundary — define what is "in" and "out" of scope. Start with high-criticality assets from your criticality analysis.
- Form a cross-functional team — RCM works best with a team of 4–6 including: a reliability engineer (facilitator), operations/process engineer, senior maintenance technician, and a maintenance planner. The facilitator guides the process; the technicians provide the failure knowledge.
- Define functions and performance standards — work through Question 1 in detail. Poor function statements lead to poor maintenance programmes.
- Identify functional failures and failure modes — brainstorm all credible failure modes per functional failure. Focus on failures that have occurred or could plausibly occur.
- Analyse failure effects and consequences — for each failure mode, document the physical effect and classify the consequence category.
- Apply the decision logic — systematically work through the task selection tree for each failure mode.
- Document the living programme — compile all tasks into maintenance task lists, load into CMMS, and schedule for implementation.
- Review and audit — RCM is a living programme. Revise when: operating context changes, new failure modes appear, or condition monitoring data suggests tasks are incorrectly specified.
Streamlined RCM
Full classical RCM is resource-intensive — a medium plant might spend 2,000+ person-hours on an RCM analysis before any changes are made. Streamlined approaches include:
- RCM2 (John Moubray): The most widely used industrial variant, refined from Nowlan & Heap. Uses a standardised decision diagram and is the basis for SAE JA1011 standard.
- MSG-3 (aviation): The specific variant used for commercial aircraft maintenance programme development.
- PM Optimisation (PMO): A streamlined approach that starts from the existing PM schedule and analyses each task against the RCM logic, eliminating unnecessary tasks and adding missing ones. Much faster than classical RCM when a PM programme already exists.
- Risk-based inspection (RBI): A consequence-focused variant used specifically for static equipment (pressure vessels, piping) in the process industry.
Benefits and outcomes
Typical results from industrial RCM implementations (multiple industry sources)
Beyond the numbers, RCM produces a maintenance programme that is defensible — every task has a documented justification. This is invaluable for audits, safety cases, and when challenged by management about maintenance budget allocation.
FAQs
A single asset system (e.g., one centrifugal pump skid) typically takes 16–40 hours of team time for a classical RCM analysis, spread over several sessions. A full plant analysis covering all critical assets may take 6–18 months. PM Optimisation is significantly faster — often 50% less time — but requires that an existing PM schedule is available as the starting point.
SAE JA1011 ("Evaluation Criteria for Reliability-Centered Maintenance (RCM) Processes") defines the minimum criteria that any process must meet to be called RCM. It requires all seven RCM questions to be answered for each failure mode of each asset, with decisions documented. It is used to evaluate and compare different RCM methodologies and software tools.
No. RCM and TPM address different aspects of maintenance. RCM is an analytical methodology for selecting maintenance tasks and strategies. TPM is an organisational philosophy focused on operator involvement in maintenance activities (autonomous maintenance), total elimination of losses, and equipment improvement. Many plants use RCM to develop the maintenance programme and TPM principles to manage its execution. They complement each other.