Corrective Maintenance, Failure Analysis & Spare Parts
Learn corrective maintenance, failure analysis, root-cause investigation, temporary and permanent repairs, spare-parts planning, failure history, and post-repair verification.
Corrective maintenance restores equipment after a defect or failure has been identified. A good repair program does more than replace the failed part. It determines what happened, evaluates whether the failure is likely to repeat, verifies that the repair restored normal operation, and ensures that critical spare parts are available when future failures occur.
Operators play an important role because they often know what the equipment was doing before the failure, what alarms occurred, and whether similar problems have happened before.
What Is Corrective Maintenance?
Corrective maintenance is work performed to correct a known defect or restore equipment to acceptable operating condition.
Examples include:
- replacing a failed bearing;
- repairing a leaking seal;
- replacing a damaged valve actuator;
- repairing a broken pipe;
- replacing a failed transmitter;
- repairing a motor or starter.
Corrective Maintenance Does Not Always Mean Emergency Maintenance
A defect may be discovered before complete equipment failure.
If the equipment can remain safely in service, corrective work may be scheduled for an appropriate maintenance window.
Emergency maintenance is required when the condition presents an immediate or rapidly developing risk to:
- safety;
- public health;
- treatment capacity;
- permit compliance;
- critical service;
- major equipment damage.
Start with the Failure Symptom
Before repairing equipment, define what actually happened.
Useful information includes:
- what failed;
- when it failed;
- what alarms occurred;
- what operating conditions existed;
- whether the failure was sudden or gradual;
- whether the problem has happened before.
Describe the Failure Clearly
A poor failure description is:
Pump failed.
A better description is:
Pump P-102 tripped on overload at 03:15 after motor current increased from approximately 21 amps to 31 amps. Flow had been declining for the previous two hours.
Specific details help maintenance personnel diagnose the real problem.
Failure Versus Symptom
The failed component may not be the root cause.
For example:
- symptom: motor overload trip;
- failed component: pump bearing;
- possible root cause: chronic misalignment.
Replacing only the bearing may restore operation temporarily without preventing another failure.
Immediate Cause
The immediate cause is the direct condition that produced the failure.
Examples include:
- bearing seized;
- seal face cracked;
- fuse opened;
- pipe ruptured;
- impeller plugged.
Root Cause
The root cause is the deeper reason the failure occurred.
Possible root causes include:
- poor lubrication;
- misalignment;
- incorrect operation;
- corrosion;
- wrong component selection;
- poor installation;
- missed preventive maintenance;
- operation outside the intended range.
Why Root-Cause Analysis Matters
If only the failed part is replaced, the same failure may happen again.
Root-cause analysis can help:
- reduce repeat failures;
- improve maintenance practices;
- improve equipment selection;
- identify operating problems;
- reduce emergency work.
When Root-Cause Analysis Is Especially Useful
Detailed analysis is especially valuable when a failure is:
- repetitive;
- costly;
- critical to treatment;
- safety-related;
- unusual;
- causing major downtime.
Ask What Changed
Failure analysis should consider whether the problem followed:
- recent maintenance;
- equipment replacement;
- control changes;
- process changes;
- chemical changes;
- power interruption;
- unusual weather or loading.
Sudden Failure
Sudden failure may be associated with:
- broken component;
- electrical fault;
- foreign material;
- overpressure;
- rapid mechanical damage.
Gradual Failure
Gradual deterioration may involve:
- bearing wear;
- corrosion;
- seal wear;
- fouling;
- alignment drift;
- lubricant deterioration.
Use Operating History
Historical data can show whether warning signs existed before failure.
Useful trends include:
- motor current;
- vibration;
- temperature;
- flow;
- pressure;
- run hours;
- alarm frequency.
Use Maintenance History
Previous work orders can reveal:
- repeat bearing replacement;
- repeat seal failures;
- repeated electrical trips;
- chronic valve problems;
- recurring instrument drift.
Repeated Failure Is Important Information
If the same component fails repeatedly, do not assume the component itself is the only problem.
Investigate:
- installation;
- operating conditions;
- maintenance interval;
- component specification;
- alignment;
- environment.
Five-Why Method
One simple root-cause tool is repeatedly asking why until the underlying cause becomes clearer.
Example:
- Why did the pump stop? The motor overload tripped.
- Why did the motor overload? Mechanical load increased.
- Why did mechanical load increase? The bearing was failing.
- Why did the bearing fail? Lubrication was inadequate.
- Why was lubrication inadequate? The PM lubrication task was missed.
This simplified example shows how a maintenance-system problem can cause an equipment failure.
Failure Analysis Should Be Evidence-Based
Do not select a root cause simply because it seems plausible.
Useful evidence may include:
- damaged parts;
- lubricant condition;
- instrument trends;
- operator observations;
- maintenance records;
- alignment data;
- alarm history.
Preserve Failed Components When Useful
For important failures, do not discard damaged parts before determining whether they are needed for:
- inspection;
- manufacturer review;
- failure analysis;
- warranty evaluation.
Temporary Repairs
A temporary repair may restore limited service until a permanent repair can be made.
Examples may include:
- temporary bypass;
- repair clamp;
- temporary pump;
- temporary control arrangement.
Temporary Repairs Must Be Controlled
A temporary repair should be:
- safe;
- authorized;
- documented;
- clearly identified;
- inspected;
- tracked until permanent resolution.
Temporary Should Not Become Permanent Accidentally
If temporary work is not tracked, it may remain in service much longer than intended.
The work order system should keep the permanent corrective action visible.
Permanent Repairs
A permanent repair should restore equipment to an acceptable long-term condition.
It should address:
- failed component;
- underlying cause when identified;
- required design and material specifications;
- normal operating capability.
Repair Versus Replace
Equipment may be repaired or replaced depending on:
- age;
- condition;
- failure history;
- repair cost;
- replacement cost;
- parts availability;
- energy efficiency;
- criticality.
Repeated Repair Cost
Equipment that requires frequent repair may cost more over time than replacement.
Maintenance history helps identify assets approaching end of useful life.
Obsolete Equipment
Older equipment may become difficult to maintain because:
- parts are unavailable;
- manufacturer support has ended;
- controls are outdated;
- repair expertise is limited.
Spare Parts
Spare parts reduce downtime when failures occur.
Possible spares include:
- bearings;
- seals;
- gaskets;
- belts;
- motors;
- VFDs;
- transmitters;
- actuators;
- fuses;
- special fasteners.
Critical Spare Parts
A critical spare is a part whose absence could cause excessive downtime or serious operational risk.
Critical-spare decisions should consider:
- equipment criticality;
- failure frequency;
- supplier lead time;
- cost;
- availability of substitutes;
- availability of redundant equipment.
Supplier Lead Time
Lead time is the time required to obtain a part after ordering.
A part used rarely may still belong in stock if:
- lead time is long;
- equipment is critical;
- no substitute exists.
Do Not Stock Everything
Excess inventory can create:
- high cost;
- storage problems;
- obsolete parts;
- deterioration in storage.
Spare-parts inventory should balance risk and cost.
Parts Standardization
Using common motors, bearings, seals, transmitters, or valves across similar equipment can reduce the number of different spare parts required.
Verify Spare-Part Compatibility
A spare part should be checked for:
- correct model;
- size;
- material;
- voltage;
- pressure rating;
- chemical compatibility;
- software or firmware compatibility where relevant.
Receiving Inspection
When parts arrive, check:
- correct item;
- quantity;
- shipping damage;
- documentation;
- storage requirements.
Storage Conditions
Improper storage can damage spare parts before they are used.
Parts may require protection from:
- moisture;
- dust;
- heat;
- corrosive atmosphere;
- sunlight;
- physical damage.
Rotate Stored Equipment When Required
Some stored motors, pumps, or bearings may require periodic shaft rotation or other preservation according to manufacturer instructions.
Shelf Life
Some materials have limited storage life.
Examples include:
- elastomer seals;
- gaskets;
- adhesives;
- batteries;
- some lubricants.
Inventory Accuracy
Maintenance personnel should be able to determine whether a critical spare is actually available.
Inventory records should match physical stock.
Minimum Stock Level
Facilities may establish minimum quantities for frequently used or critical parts.
A simple concept is:
Reorder Point = Expected Use During Lead Time + Safety Stock
Spare-Parts Example
Suppose a facility typically uses 2 mechanical seals during a six-month supplier lead time and wants 1 additional seal as safety stock.
Reorder Point = 2 + 1 = 3 seals
When usable stock reaches the reorder point, a new order should be considered according to facility procedure.
Parts Used Should Be Recorded
Work-order closure should identify important parts used during repair.
This helps:
- update inventory;
- track repair cost;
- identify recurring failures;
- forecast future demand.
Emergency Spare-Part Planning
Facilities should know how to obtain critical parts after normal business hours.
Options may include:
- local supplier;
- emergency vendor;
- manufacturer;
- nearby utility;
- rental equipment provider.
Interchangeable Equipment
Some utilities reduce risk by using equipment that can be interchanged among several locations.
This can simplify:
- spare motors;
- spare pumps;
- controls;
- training.
Post-Repair Testing
A repair is not complete until equipment is tested.
Post-repair checks may include:
- flow;
- pressure;
- motor current;
- temperature;
- vibration;
- leakage;
- alarm operation;
- control response.
Compare Before and After
Compare post-repair values with:
- pre-failure data;
- normal baseline;
- similar equipment.
If performance remains abnormal, the repair may not have corrected the full problem.
Verify Correct Rotation
After motor, pump, blower, or electrical work, verify correct rotation according to approved procedure.
Wrong rotation can cause:
- low flow;
- low pressure;
- poor equipment performance.
Verify Alignment
After work affecting:
- motor;
- pump;
- coupling;
- bearing;
- baseplate;
alignment may need to be checked before or during return to service.
Verify Valve Lineup
Maintenance often requires valves to be moved.
Before normal operation resumes, verify:
- suction valves;
- discharge valves;
- bypass valves;
- drains;
- vents;
- chemical isolation valves.
Verify Control Mode
Equipment may be left in:
- manual;
- hand;
- local;
- off.
Confirm the required automatic or remote control mode is restored.
Verify Alarms and Interlocks
If maintenance affected controls, test or verify:
- alarms;
- permissives;
- interlocks;
- shutdown protection.
Close the Work Order Completely
A completed corrective work order should record:
- failure symptom;
- cause found;
- repair performed;
- parts used;
- test results;
- follow-up work needed.
Failure Codes
A CMMS may use standardized failure codes such as:
- bearing failure;
- seal failure;
- corrosion;
- electrical fault;
- instrument failure;
- blockage;
- misalignment.
Consistent coding helps analyze failure patterns.
Repeat-Failure Analysis
Review equipment history for:
- same part repeatedly replaced;
- same alarm repeatedly occurring;
- same equipment repeatedly failing;
- repair interval becoming shorter.
These patterns may indicate unresolved root causes.
Bad Actor Equipment
Equipment that consumes excessive maintenance time or repeatedly fails is sometimes informally called a bad actor.
Such assets may justify:
- detailed root-cause analysis;
- design modification;
- different maintenance strategy;
- replacement.
Lessons Learned
A significant failure should improve future operation.
Possible improvements include:
- new inspection point;
- shorter PM interval;
- different spare part;
- revised operating procedure;
- alarm modification;
- staff training.
Do Not Change Procedures Without Control
Changes to maintenance or operating procedures should be reviewed and documented according to facility practice.
An informal change can create inconsistent or unsafe work.
Maintenance Metrics
Useful maintenance trends may include:
- repeat failures;
- emergency work;
- equipment downtime;
- repair cost;
- spare-parts use;
- mean time between failures.
Mean Time Between Failures
Mean Time Between Failures, or MTBF, is a reliability measure describing average operating time between failures for repairable equipment.
Increasing MTBF generally indicates improving reliability.
Downtime
Downtime includes the time equipment is unavailable because of:
- failure;
- waiting for parts;
- waiting for labor;
- repair;
- testing.
Long parts lead times can therefore significantly increase total downtime.
Planning Reduces Repair Time
Good preparation reduces downtime by ensuring availability of:
- parts;
- tools;
- labor;
- procedures;
- lifting equipment;
- contract support.
Safe Corrective Maintenance
Corrective work may expose workers to:
- electrical energy;
- rotating equipment;
- stored pressure;
- chemicals;
- gravity;
- confined spaces.
Lockout/Tagout
Required hazardous-energy controls must be used before servicing equipment where exposure exists.
A stop command, selector switch, or SCADA command is not a substitute for required energy isolation.
Verify Zero-Energy Condition
Isolation may require control of:
- electrical energy;
- hydraulic pressure;
- pneumatic pressure;
- mechanical motion;
- gravity;
- chemical flow.
Do Not Rush Emergency Repairs Unsafely
Emergency conditions create pressure to restore service quickly.
Safety requirements still apply.
A repair that creates injury or additional equipment damage does not improve system reliability.
Common Corrective Maintenance Mistakes
- Replacing the failed part without asking why it failed.
- Ignoring repeat failures.
- Failing to review operating data before repair.
- Discarding failed parts before important analysis is complete.
- Allowing temporary repairs to remain untracked.
- Using incorrect replacement parts.
- Failing to verify alignment or rotation after repair.
- Returning equipment to service without testing.
- Failing to restore valve lineup and control mode.
- Closing work orders without recording cause and repair details.
Common Spare-Parts Mistakes
- Stocking no spare for a critical long-lead component.
- Stocking excessive quantities of rarely used parts.
- Failing to verify compatibility.
- Allowing spare parts to deteriorate in poor storage conditions.
- Failing to track parts used.
- Allowing inventory records to differ from physical stock.
- Ignoring shelf life.
A Practical Corrective Maintenance Sequence
- Define the failure symptom.
- Make the equipment and process safe.
- Review alarms, trends, and operator observations.
- Inspect the failed equipment.
- Identify the immediate failure cause.
- Determine whether deeper root-cause analysis is needed.
- Plan the repair, parts, labor, tools, and isolation.
- Perform the repair using approved procedures.
- Verify alignment, rotation, valve lineup, and controls as applicable.
- Test equipment under operating conditions.
- Compare performance with normal baseline.
- Document cause, repair, parts, and test results.
- Update spare-parts inventory.
- Implement follow-up action if a root cause was identified.
A Practical Spare-Parts Review
- Identify critical equipment.
- Identify likely failure components.
- Review failure history.
- Review supplier lead times.
- Determine available redundancy.
- Identify suitable substitutes where approved.
- Set appropriate minimum stock levels.
- Store parts under suitable conditions.
- Verify inventory periodically.
- Review obsolete and slow-moving stock.
What to Remember for the Exam
- Corrective maintenance restores equipment after a defect or failure is identified.
- Corrective maintenance can be planned and does not always mean emergency work.
- The immediate failed component may not be the root cause.
- Root-cause analysis seeks the underlying reason a failure occurred.
- Repeat failures are a strong reason to investigate beyond simple part replacement.
- Operating trends, alarm history, maintenance records, and damaged parts provide useful failure evidence.
- Temporary repairs should be authorized, documented, inspected, and tracked to permanent resolution.
- Repair-versus-replace decisions should consider age, failure history, repair cost, parts availability, and equipment criticality.
- Critical spare parts reduce downtime when important equipment fails.
- Spare-parts decisions should consider equipment criticality, failure frequency, lead time, cost, and redundancy.
- A long-lead part may justify stocking even if it fails infrequently.
- Stored spare parts must be protected from moisture, corrosion, heat, and other damaging conditions.
- Inventory records should match actual physical stock.
- Reorder point considers expected use during lead time plus safety stock.
- Post-repair testing should confirm normal flow, pressure, current, temperature, vibration, leakage, and controls as applicable.
- Verify correct rotation, alignment, valve lineup, and control mode before normal operation resumes.
- Work orders should document failure cause, repair performed, parts used, and test results.
- Failure history can identify bad-actor equipment and recurring root causes.
- Required lockout/tagout and process isolation still apply during emergency repairs.
- Successful corrective maintenance restores equipment and reduces the chance of the same failure recurring.