You already know reactive maintenance is expensive. Here is how to stop doing it.
Every maintenance professional has lived this moment: a bearing seizes at 2 a.m. on a Saturday, the production line goes down, overtime gets called, an emergency parts order ships next-day air at triple the cost, and someone asks why nobody saw it coming. The answer, almost always, is that the signs were there — rising vibration, elevated temperature, metallic debris in the oil — but nobody was looking.
Condition monitoring is the discipline of looking. Predictive maintenance is the strategy that turns what you see into the right repair at the right time — not too early (wasting parts life), not too late (unplanned downtime), but in the window where you can plan the work, stage the parts, and schedule the outage.
This guide covers the complete framework: the theory that justifies the investment (the P-F curve), the four core monitoring technologies and when each one earns its keep, the ISO standards that govern the discipline, the reliability metrics that prove it is working, and the implementation sequence that separates programs delivering 10:1 ROI from programs that collected sensors and dust.
Key Takeaways
- Condition monitoring detects the early signs of failure; predictive maintenance is the strategy that acts on those signs at the optimal time.
- The P-F curve is the theoretical foundation: it shows that failures develop over time, and the gap between first-detectable symptom (P) and functional failure (F) is your intervention window.
- Four core technologies — vibration analysis, infrared thermography, oil analysis, and ultrasound — detect different failure modes at different points on the P-F curve. No single technology covers everything.
- ISO 17359:2018 provides the overarching framework; ISO 13373 (vibration), ISO 18434 (thermography), and ISO 13381 (prognostics) govern the individual disciplines.
- A successful program starts with asset criticality ranking, not sensor purchases. Monitor the assets whose failure costs the most, and match the monitoring technology to each asset's dominant failure mode.
- World-class programs achieve 25–40% maintenance cost reduction, 50–70% reduction in unplanned downtime, and payback within 6–18 months.
The P-F curve: the single concept that justifies every dollar you spend on monitoring
If you take one idea from this article, make it this one. The P-F curve (potential failure to functional failure curve) is the graphical model that shows how a failure develops over time — and it is the entire theoretical justification for condition monitoring.
The horizontal axis is time. The vertical axis is the asset's resistance to failure (or, inverted, its degradation). The curve slopes downward from normal operation toward complete functional failure. Two points on the curve matter:
Point P (potential failure): The earliest point at which a developing fault can be detected by a monitoring technique. A subsurface crack in a bearing race produces a specific ultrasonic signature weeks or months before the bearing gets noisy enough for a passing operator to notice.
Point F (functional failure): The point at which the asset can no longer perform its intended function. The bearing seizes, the motor stops, production halts.
The P-F interval is the time between P and F. That interval is your window — the time you have to detect the fault, diagnose the cause, plan the repair, order parts, and schedule the outage. The entire value proposition of condition monitoring is that it gives you access to this window. Without monitoring, you discover the problem at F (or after), when the only option is emergency repair.
Different monitoring technologies detect different failure modes at different points on the curve:
| Detection stage | Technology | What it catches | Typical lead time before functional failure |
|---|---|---|---|
| Earliest (Stage 1–2) | Ultrasound | Friction, micro-impacts, lubrication deficiency | Months |
| Early-to-mid (Stage 2–3) | Oil analysis | Wear debris, contamination, lubricant degradation | Weeks to months |
| Mid (Stage 3) | Vibration analysis | Imbalance, misalignment, looseness, bearing defects | Weeks |
| Mid-to-late (Stage 3–4) | Infrared thermography | Abnormal heat from friction, electrical faults, insulation breakdown | Days to weeks |
| Late (audible/visible) | Human senses | Noise, smoke, visible damage | Hours to days |
The practical implication: if your "condition monitoring program" consists of operators reporting unusual noises, you are detecting failures at the bottom of the curve — too late for planned intervention, too early for "we never saw it coming" to be an honest excuse.
The four core technologies — and when each one earns its keep
A strong program uses multiple technologies because no single one covers every failure mode on every asset class. Here is what each technology does, what it detects best, and where it falls short.
Vibration analysis
Vibration analysis is the backbone of most industrial condition monitoring programs. Accelerometers mounted on (or held against) rotating equipment capture the mechanical signatures that faults produce — and every fault type produces a characteristic pattern.
What it detects: Imbalance, misalignment (angular and offset), mechanical looseness, rolling-element bearing defects (inner race, outer race, ball/roller, cage), gear mesh problems, belt defects, resonance, and soft foot.
How it works: A time-domain vibration signal is converted via Fast Fourier Transform (FFT) into a frequency spectrum. Each fault type shows up at predictable frequencies relative to the machine's running speed. Bearing defects, for example, produce energy at calculated defect frequencies (BPFO, BPFI, BSF, FTF) derived from the bearing geometry and shaft speed. A trained analyst — or increasingly, pattern-recognition software — reads the spectrum and the time waveform to diagnose what is developing and how far along it is.
Where it excels: Rotating machinery above roughly 600 RPM — motors, pumps, fans, compressors, gearboxes, turbines. This is the broadest category of industrial equipment, which is why vibration is the first technology most programs deploy.
Where it falls short: Very slow-speed equipment (below ~100 RPM), where the energy generated by defects is too low for conventional accelerometers. Electrical faults in motors are better caught by motor current analysis. Lubrication deficiency is detectable by vibration, but ultrasound catches it earlier.
Governing standard: ISO 10816 / ISO 20816 series (vibration severity evaluation), ISO 13373 (vibration condition monitoring), ISO 13374 (data processing).
Infrared thermography
Thermography uses an infrared camera to image surface temperatures. Anything that generates abnormal heat — friction from a failing bearing, a loose electrical connection, a blocked heat exchanger, insulation breakdown — shows up as a hot spot.
What it detects: Overheating bearings, misaligned couplings, electrical hot spots (loose connections, overloaded circuits, failing contactors), blocked or fouled heat exchangers, refractory degradation, steam trap failures, insulation deficiency in buildings and process equipment.
How it works: An IR camera captures the thermal radiation emitted by surfaces and renders it as a color-mapped image. The analyst compares temperatures across similar components (a loaded phase versus an unloaded one, a suspect bearing versus its neighbors) and flags anomalies. Quantitative thermography requires knowing the emissivity of the surface being measured; qualitative comparison (same material, same conditions, different temperatures) is often sufficient for fault detection.
Where it excels: Electrical distribution systems (the single highest-ROI application — finding a loose bus connection before it arcs and causes a fire), steam systems, roofing and building envelope surveys, mechanical equipment where elevated surface temperature correlates with internal friction.
Where it falls short: Thermography measures surface temperature, not internal condition. A bearing can be failing internally before the heat migrates to the housing surface. Wind, solar loading, and low emissivity surfaces (shiny metals) introduce measurement error. It is a screening tool, not a diagnostic one — it tells you something is hot, not always why.
Governing standard: ISO 18434 series (condition monitoring and diagnostics of machines — thermography), ASNT SNT-TC-1A (thermographer qualification).
Oil analysis
Oil analysis is the "blood test" for any machine with a lubricated system — gearboxes, hydraulic systems, engines, compressors, turbines. It evaluates both the condition of the lubricant itself and the condition of the machine the lubricant is protecting.
What it detects: Wear metals (iron, copper, lead, tin, chromium — each pointing to a specific component), contamination (water, dirt, process fluid, wrong lubricant), lubricant degradation (oxidation, viscosity change, additive depletion, acid number increase), and particle morphology (the shape and size of wear debris, which indicates the wear mechanism — rubbing, cutting, fatigue, corrosion).
How it works: A representative sample is drawn (sampling technique matters enormously — a bad sample is worse than no sample) and sent to a lab or analyzed on-site. Standard tests include spectrometric analysis (elemental wear metals), particle counting (ISO 4406 cleanliness code), viscosity, water content (Karl Fischer), acid number (AN/TAN), and analytical ferrography or filter patch for large particles and morphology.
Where it excels: Gearboxes, hydraulic systems, large turbines, reciprocating engines, and any system where the lubricant contacts critical wear surfaces. Oil analysis catches contamination and lubricant degradation that no other technology detects, and wear-metal trending provides an early, cumulative record of internal component wear.
Where it falls short: It requires disciplined sampling — inconsistent technique, wrong sample point, or contaminated bottles produce misleading results. Turnaround time for off-site labs (typically 2–5 business days) introduces latency; on-site or inline sensors reduce this but increase cost. It does not apply to dry or grease-lubricated equipment (though grease analysis exists, it is less mature).
Governing standard: ASTM D7720 (condition monitoring program guide), ISO 4406 (fluid cleanliness code), ASTM D6224 (in-service lubricant condition data sets).
Ultrasound
Airborne and structure-borne ultrasound monitoring detects the high-frequency sound waves (typically 20–100 kHz, above human hearing) generated by friction, impacts, turbulence, and electrical discharge.
What it detects: Bearing lubrication deficiency (the earliest mechanical indicator on the P-F curve for lubricated bearings), compressed air and gas leaks, steam trap failures, electrical arcing and corona discharge, and cavitation in pumps and valves.
How it works: A handheld or permanently mounted ultrasonic sensor converts high-frequency signals into an audible range the technician can hear through headphones, and displays a decibel reading. The technician compares the reading to baseline or to similar equipment. Some instruments capture time-domain ultrasonic waveforms for spectral analysis. Because ultrasound is highly directional and attenuates quickly, it is excellent at pinpointing the source of a signal — you can isolate one bearing in a bank of machines.
Where it excels: Lubrication management (detecting under- or over-greasing in real time — the dB level drops as grease reaches the bearing and friction decreases), compressed air leak surveys (a single plant-wide survey routinely finds 20–30% waste), slow-speed bearing monitoring (below the range where vibration analysis is effective), steam trap auditing, and electrical inspection of switchgear and transformers.
Where it falls short: Ultrasound gives you an early alarm ("this bearing is getting louder"), but it does not diagnose the specific defect the way vibration spectral analysis can. It is a screening and trending tool. Environmental ultrasonic noise in some facilities (high-pressure steam, pneumatic exhaust) can mask signals.
Governing standard: No single ISO standard governs ultrasound condition monitoring the way ISO 13373 governs vibration, though ISO 29821 covers ultrasonic leak detection.
How the technologies work together
The strongest programs do not pick one technology — they deploy the right combination based on asset class, failure mode, and criticality. Here is a practical mapping:
| Asset class | Primary technology | Supporting technology | Why |
|---|---|---|---|
| Motors and pumps (>600 RPM) | Vibration | Ultrasound, thermography | Vibration diagnoses the defect; ultrasound catches lubrication issues earlier; thermography screens for electrical and thermal problems |
| Gearboxes | Vibration + oil analysis | Ultrasound | Vibration reads gear mesh and bearing health; oil reveals wear metals and contamination that vibration misses |
| Slow-speed equipment (<100 RPM) | Ultrasound | Oil analysis | Conventional vibration often lacks sensitivity at low speed; ultrasound and oil fill the gap |
| Electrical switchgear and distribution | Thermography + ultrasound | — | Thermography finds hot connections; ultrasound detects arcing and corona; neither requires energized contact |
| Hydraulic systems | Oil analysis | Thermography | Oil tracks contamination, degradation, and internal wear; thermography identifies overheating components |
| Steam systems | Ultrasound | Thermography | Ultrasound verifies trap function (failed open/closed); thermography confirms downstream temperature |
| Reciprocating engines | Oil analysis | Vibration (specialized) | Wear metals and fuel dilution are the primary failure indicators; reciprocating vibration analysis requires specialized techniques |
A common implementation sequence: start with vibration on your rotating assets (highest population, most mature technology, broadest coverage), add thermography for electrical and steam (high-ROI, relatively low cost per survey), introduce oil analysis for gearboxes and hydraulics (catches what vibration does not), and deploy ultrasound for lubrication precision and leak surveys (fast payback on compressed air alone).
The standards that govern the discipline
Condition monitoring is not a Wild West of vendor claims. A robust family of ISO standards provides the framework, and referencing them in your program lends credibility and structure.
ISO 17359:2018 — Condition monitoring and diagnostics of machines: General guidelines. This is the overarching standard. It provides the framework for setting up a condition monitoring program, covers the basic principles and methodologies from system design to data analysis and decision-making, and is not limited to any single technology. If you read one standard, read this one.
ISO 13373 series — Vibration condition monitoring. Part 1 covers general procedures; Part 2 covers processing, analysis, and presentation of vibration data; subsequent parts cover specific equipment.
ISO 18434 series — Thermography. Covers condition monitoring of machinery using infrared thermography.
ISO 13379-1:2012 — Data interpretation and diagnostics techniques. Guidelines for interpreting condition monitoring data and diagnosing machine condition — the analytical layer between data collection and decision.
ISO 13381-1:2015 — Prognostics. Guidelines for developing and applying prognostics processes — estimating remaining useful life from condition data. This is the frontier where condition monitoring meets predictive analytics.
ISO 20816 / ISO 10816 series — Vibration severity. Evaluation criteria for machine vibration measured on non-rotating parts. ISO 20816 is the current active series, replacing 10816 part by part.
ISO 4406:1999 — Hydraulic fluid power: Fluids: Method for coding the level of contamination by solid particles. The standard cleanliness code used in oil analysis (e.g., 18/16/13).
For personnel qualification, the two primary schemes are ISO 18436 (condition monitoring and diagnostics of machines — requirements for qualification and assessment of personnel) and ASNT SNT-TC-1A (for thermographers specifically). Category I through IV certification under ISO 18436 covers vibration analysis; similar tiers exist for thermography, oil analysis, and ultrasound through bodies like the Mobius Institute, Vibration Institute, and ICML.
Reliability metrics: how you prove the program is working
You cannot justify a condition monitoring program with anecdotes. You need metrics that connect monitoring activity to business outcomes. The Society for Maintenance & Reliability Professionals (SMRP) publishes over 70 standardized metrics; these five are the ones that matter most for a condition monitoring program:
MTBF (Mean Time Between Failures): Total operating time divided by the number of failures. Rising MTBF means failures are happening less often — the primary outcome condition monitoring is supposed to deliver. Track it at the asset level, not just plant-wide; a plant-wide average masks problem assets.
MTTR (Mean Time To Repair): Total repair time divided by the number of repairs. Condition monitoring should reduce MTTR because planned repairs (parts staged, procedures prepared, skilled craft scheduled) take less time than emergency breakdowns. World-class discrete manufacturing MTTR benchmarks range from 30 minutes to 2 hours.
Availability: MTBF / (MTBF + MTTR). This is the metric that connects to production output. A condition monitoring program that improves MTBF and reduces MTTR drives availability up — and availability is what operations cares about.
Planned Maintenance Percentage (PMP): Planned work orders divided by total work orders. A program working correctly shifts work from reactive (unplanned) to proactive (planned). SMRP best practice: 85% or higher planned. If your PMP is below 50%, your condition monitoring data is either not being generated, not being analyzed, or not being acted on.
Maintenance Cost as % of Replacement Asset Value (RAV): Total annual maintenance cost divided by the replacement cost of the asset base. SMRP best-in-class is below 3% RAV. Condition monitoring reduces maintenance cost by eliminating unnecessary time-based replacements and preventing catastrophic secondary damage from undetected failures.
Building the program: implementation sequence
Programs fail when they start with technology purchases and skip the foundational work. Here is the sequence that works.
Phase 1: Asset criticality ranking
Before you buy a single sensor, rank your assets by consequence of failure. A simple criticality matrix scores each asset on three dimensions: safety impact (could failure injure someone?), production impact (what is the downtime cost per hour?), and environmental / regulatory impact (could failure cause a spill, release, or compliance violation?). Multiply or weight the scores.
The output is a prioritized list. Your condition monitoring program starts with the top 10–20% — the assets whose failure costs the most. This is not optional. Monitoring everything equally means monitoring nothing well.
Phase 2: Failure mode analysis
For each critical asset, identify the dominant failure modes. A motor-driven centrifugal pump, for example, fails most commonly from bearing degradation, seal failure, impeller wear, cavitation, and misalignment. Each failure mode has a characteristic signature detectable by a specific technology. Map the technology to the mode:
- Bearing degradation → vibration + ultrasound + oil analysis (if oil-lubricated)
- Seal failure → thermography (elevated temperature) + visual (leakage)
- Cavitation → ultrasound + vibration (high-frequency)
- Misalignment → vibration (2x running speed axial) + thermography (coupling heat)
This mapping determines your technology portfolio. You are not choosing technologies in the abstract — you are matching detection capability to failure risk on specific assets.
Phase 3: Baseline data collection
Install or route sensors, collect initial readings across all selected assets, and establish baselines. A baseline is critical because condition monitoring is fundamentally a trending discipline — you are looking for change from normal, not absolute values. A motor running at 0.15 in/s velocity is fine if it has always run at 0.15; it is a problem if it was 0.04 last month.
Set alarm levels: alert (investigate at next opportunity) and alarm (investigate immediately, plan corrective action). ISO 20816 provides generic vibration severity zones (A through D) by machine class and mounting type, but asset-specific baselines are more useful than generic tables once you have enough history.
Phase 4: Monitoring routes and frequency
Decide how you will collect data: route-based (a technician walks a circuit with a portable analyzer on a fixed schedule), continuous/online (permanently mounted sensors streaming data to a central system), or a hybrid (continuous on the most critical assets, route-based on the rest).
Monitoring frequency must be shorter than half the P-F interval for each failure mode — this is the Nyquist principle applied to maintenance. If a bearing defect progresses from detectable to functional failure in 60 days, you need to check at least every 30 days. Monthly vibration routes on rotating equipment and quarterly thermographic surveys of electrical panels are common starting points, but adjust based on your actual P-F intervals and criticality.
Phase 5: Analysis, diagnosis, and action workflow
Data without analysis is noise. Analysis without action is waste. Build the workflow:
- Data collection (automated or manual route)
- Automated screening (software flags readings exceeding alert/alarm levels or showing trend changes)
- Analyst review (a qualified analyst — ISO 18436 Category II or higher for vibration — reviews flagged data, diagnoses the fault, and estimates severity)
- Work order generation (the analyst's recommendation becomes a planned work order in the CMMS with a recommended timeframe, parts list, and craft requirements)
- Execution (maintenance performs the repair during a planned window)
- Verification (post-repair data confirms the fault is corrected and the repair was effective)
- Feedback loop (was the diagnosis correct? Was the recommended timeframe appropriate? Feed lessons back into alarm levels and diagnostic rules)
The workflow fails most often at step 4. Data gets collected, analysis gets done, reports get filed, and nobody creates a work order. Build the handoff between the condition monitoring analyst and the maintenance planner into the program as a formal, tracked step — not an email that gets ignored.
Phase 6: Continuous improvement
Review program performance quarterly. Track the metrics in the previous section. Audit a sample of recommendations: was the diagnosis confirmed at teardown? Were the recommended actions taken, and were they taken in time? Expand the program to the next tier of critical assets. Retire monitoring on assets where the data shows no value. Train additional analysts. Pursue certification.
The ROI case — what the numbers actually show
The business case for condition monitoring is well-documented:
Cost reduction: Proactive repairs cost 4 to 5 times less than emergency repairs on the same asset. A planned bearing replacement during a scheduled outage involves a $200 bearing, two hours of craft time, and zero production loss. The same bearing failure as an emergency involves the bearing, the shaft damage the failed bearing caused, 8–16 hours of unplanned downtime, overtime labor, expedited parts shipping, and collateral damage to seals and couplings.
Downtime reduction: Organizations implementing predictive maintenance report 50–70% reduction in unplanned downtime. Each prevented hour of downtime saves $15,000–$50,000 in direct production loss depending on industry — before accounting for overtime, waste, and restart costs.
Maintenance cost reduction: The documented range is 25–40% total maintenance cost reduction. The savings come from eliminating unnecessary time-based replacements (extending parts life to actual condition-based limits), preventing secondary damage (a failed bearing that takes out a shaft, seal, and coupling), and reducing emergency labor and expediting costs.
Payback period: Published case studies show payback periods of 6–18 months. A Midwest steel manufacturer documented $850,000 in annual savings with a 30% reduction in unplanned downtime. Ford's commercial vehicle division prevented $7 million in downtime on a single component type by predicting 22% of failures 10 days in advance.
Compressed air leak detection alone — a single ultrasound survey — routinely finds 20–30% waste in a plant's compressed air system. At $0.25–$0.30 per 1,000 cubic feet, the cost savings from fixing leaks often pay for the entire ultrasound program in the first year.
The ten mistakes that kill condition monitoring programs
Programs fail more often from organizational causes than technical ones. These are the mistakes the competing guides do not talk about:
1. Starting with technology instead of criticality. Buying sensors for every motor in the plant sounds thorough. It produces so much data that nothing gets analyzed, and the program drowns in its own output. Start with the 20% of assets that cause 80% of the consequences.
2. Collecting data without analyzing it. Route-based data collection becomes a ritual: the technician walks the route, fills the boxes, uploads the data, and nobody looks at it until something breaks. If your data backlog is measured in months, you do not have a condition monitoring program — you have an expensive walking program.
3. Analyzing without acting. The analyst identifies a developing bearing defect, writes the report, and emails it to the planner. The planner, buried in reactive work orders, never creates the planned job. The bearing fails anyway. Build the analyst-to-planner handoff as a formal, tracked workflow step.
4. Wrong monitoring frequency. Monthly routes on an asset with a 14-day P-F interval guarantee missed faults. Match frequency to the actual P-F interval, not to what is convenient for the route schedule.
5. Ignoring the baseline. Trending from the wrong starting point produces false alarms and missed faults. Collect clean baseline data during normal, representative operating conditions before setting alarm levels.
6. One-technology thinking. A vibration-only program misses lubricant degradation, electrical hot spots, and slow-speed bearing defects. Match the technology to the failure mode.
7. Inadequate analyst training. A technician with a $50,000 vibration analyzer and no training is not a vibration analyst. ISO 18436 Category I is a minimum; Category II is the working diagnostic level. Budget for training and certification as a program cost, not a discretionary expense.
8. No feedback loop. If you never compare your diagnoses to what you find at teardown, you never improve. Require a teardown verification step on every condition-based work order and feed the results back to the analyst.
9. Treating monitoring as the maintenance department's problem. Operations must cooperate on access, run conditions, and scheduling. If the reliability engineer cannot get access to a machine because production will not give a window, the program fails at the access point, not the technology.
10. Expecting instant results. A condition monitoring program takes 6–12 months to build baseline history, tune alarm levels, and start generating reliable diagnoses. Programs that get cancelled after three months because "we haven't found anything yet" were never given a chance to work.
Where the discipline is heading: AI, IoT, and digital twins
Condition monitoring is in the middle of a technology shift driven by three converging trends:
Wireless IoT sensors. The cost of permanently mounted wireless vibration and temperature sensors has dropped to the point where continuous monitoring is economical on assets that previously justified only monthly routes. The wireless condition monitoring sensor market is projected to reach $2.6 billion by 2034 (from $1 billion in 2026), growing at 12.4% CAGR. Multi-parameter sensors (vibration + temperature + ultrasonic in one device) are reducing installation and maintenance cost per monitoring point.
Edge computing and AI-driven analytics. Machine learning models trained on large fleets of similar equipment can detect anomaly patterns that human analysts miss — and they can do it continuously, not on a monthly route cycle. Edge computing devices process data locally for sub-second anomaly detection; cloud platforms aggregate fleet-wide data for long-term trend analysis. Published model accuracy: 80–97% prediction accuracy, identifying issues 30–90 days before traditional inspection methods.
Digital twins. A physics-based or data-driven model of the asset runs in parallel with the real machine, comparing actual sensor readings to predicted readings under current operating conditions. Deviations between the twin and the real asset flag developing faults, adjusted for load, speed, and ambient conditions — reducing the false-alarm rate that plagues simple threshold-based monitoring.
These technologies do not replace the fundamentals covered in this guide — they accelerate them. You still need criticality ranking to decide where to monitor, failure mode analysis to decide what to monitor, and trained people to validate diagnoses and make decisions. The technology shift is in how data gets collected (continuous vs. route), how it gets analyzed (automated pattern recognition vs. manual spectrum reading), and how much of the fleet you can afford to cover.
Connecting condition monitoring to OSHA and process safety
Condition monitoring is not just a reliability tool — it has direct safety and regulatory implications.
OSHA Process Safety Management (PSM), 29 CFR 1910.119: The Mechanical Integrity element requires employers to establish and implement written procedures for the ongoing integrity of process equipment, including inspections and tests. Condition monitoring is one of the most effective ways to satisfy mechanical integrity requirements for rotating equipment in PSM-covered processes.
OSHA Machine Guarding, 29 CFR 1910.212: Requires that guards be maintained in effective condition. Condition monitoring of guarded machinery — detecting that a guard has been removed, a safety interlock bypassed, or a machine is operating outside its design envelope — supports compliance.
OSHA General Duty Clause, Section 5(a)(1): A recognized hazard that is causing or likely to cause death or serious physical harm must be abated. A machine with known, unaddressed vibration levels that exceed severity thresholds is a recognized hazard. Condition monitoring data that is collected and ignored can become evidence against the employer in a General Duty citation.
The practical takeaway: condition monitoring data creates a duty to act. If your program identifies a fault and you do not address it in a reasonable timeframe, the data trail becomes a liability, not an asset. Build the response workflow (analysis → work order → execution → verification) before you build the data collection system.
Frequently Asked Questions
What is the difference between condition monitoring and predictive maintenance? Condition monitoring is the practice of collecting and analyzing data on equipment health — vibration levels, temperatures, oil condition, ultrasonic signatures. Predictive maintenance is the broader maintenance strategy that uses condition monitoring data (along with analytics and operational context) to predict when a failure will occur and schedule maintenance at the optimal time. Condition monitoring is the input; predictive maintenance is the strategy.
What are the four main condition monitoring techniques? The four core techniques are vibration analysis, infrared thermography, oil analysis, and ultrasound. Each detects different failure modes at different points on the P-F curve. Vibration excels on rotating equipment above 600 RPM. Thermography finds electrical hot spots and thermal anomalies. Oil analysis catches contamination and internal wear in lubricated systems. Ultrasound detects lubrication deficiency, leaks, and electrical discharge earliest on the curve.
What is the P-F curve and why does it matter? The P-F curve shows how an asset's condition degrades over time from normal operation (top) to functional failure (bottom). Point P is where a monitoring technology can first detect a developing fault. Point F is where the asset stops working. The P-F interval — the time between P and F — is the window you have to plan and execute a repair. Condition monitoring exists to detect faults at point P instead of discovering them at point F.
How much does a condition monitoring program cost? Costs vary widely by scope. A starter vibration program (portable analyzer, training, 50–100 assets on monthly routes) runs $30,000–$60,000 in the first year. A comprehensive program with continuous sensors, multiple technologies, and trained analysts runs $100,000–$500,000+ for a medium manufacturing plant. The documented ROI is 5:1 to 15:1 with payback in 6–18 months, so the question is not what it costs — it is what not having it costs.
What should I monitor first? Start with the assets whose failure costs the most — the intersection of high safety impact, high production impact, and high repair cost. Typically this is your critical rotating equipment: large motors, pumps, fans, compressors, and gearboxes. Rank them with a criticality matrix and monitor the top 10–20% first.
How often should condition monitoring data be collected? At minimum, monitoring frequency must be less than half the P-F interval for the failure modes you are targeting. Monthly vibration routes and quarterly thermographic surveys are common starting points. Critical assets may justify continuous online monitoring. Adjust frequency based on actual failure progression rates and asset criticality.
Do I need certified analysts? For vibration analysis, ISO 18436 Category II is the practical minimum for diagnostic-level analysis. Category I personnel can collect data and identify obvious trends, but reliable fault diagnosis requires Category II or higher. Thermographers should hold at least ASNT Level I or equivalent. Oil analysis has its own certification through ICML (International Council for Machinery Lubrication). Untrained analysts with expensive instruments produce expensive bad data.
What is the difference between condition-based maintenance and time-based preventive maintenance? Time-based preventive maintenance performs tasks on a fixed schedule regardless of equipment condition — change the oil every 3 months, replace the belt every year. Condition-based maintenance performs tasks based on actual measured condition — change the oil when analysis shows it needs changing, replace the belt when vibration data shows it is deteriorating. Condition-based maintenance eliminates unnecessary replacements (saving parts and labor) while catching developing failures that a calendar schedule misses.
Can condition monitoring replace preventive maintenance entirely? No. Some tasks — calibration, filter changes on contamination-sensitive systems, safety-critical inspections — are better performed on a schedule. Condition monitoring works best as a complement to a PM program, not a replacement. The goal is to shift the mix: reduce unnecessary time-based replacements, increase condition-based interventions, and eliminate reactive emergency repairs.
How does condition monitoring apply to electrical systems? Infrared thermography and ultrasound are the primary technologies for electrical systems. Thermography detects hot connections, overloaded circuits, failing contactors, and phase imbalance through surface temperature differences. Ultrasound detects arcing, tracking, and corona discharge in switchgear and transformers. Electrical thermographic surveys are the single highest-ROI condition monitoring application in most facilities — a single hot connection found before it arcs can prevent a fire, an outage, and a six-figure loss.
What ISO standards apply to condition monitoring? The overarching framework is ISO 17359:2018 (general guidelines). Technology-specific standards include ISO 13373 (vibration), ISO 18434 (thermography), ISO 13379 (data interpretation and diagnostics), and ISO 13381 (prognostics). ISO 20816 (replacing ISO 10816) covers vibration severity evaluation. ISO 4406 provides the fluid cleanliness code for oil analysis. Personnel certification follows ISO 18436.
What is the role of a CMMS in condition monitoring? Your Computerized Maintenance Management System is the action engine. Condition monitoring data flows into analysis; analysis produces a recommendation; the recommendation becomes a planned work order in the CMMS with a target date, parts list, and craft assignment. Without the CMMS link, recommendations die in email inboxes. Track the conversion rate from condition monitoring recommendations to completed work orders as a program health metric.
How do I calculate the ROI of condition monitoring? Compare the total cost of the program (sensors, software, analyst time, training, certification) against the documented savings: avoided emergency repairs, reduced overtime, extended parts life, reduced secondary damage, avoided production losses. A single prevented catastrophic failure on a critical asset often pays for the entire annual program. Track avoided-cost events meticulously — they are your budget justification.
What is the difference between online and route-based monitoring? Route-based monitoring uses a technician with a portable instrument walking a circuit on a schedule (monthly, quarterly). Online (continuous) monitoring uses permanently installed sensors streaming data to a central system 24/7. Route-based is lower cost per point but provides only periodic snapshots. Online provides continuous coverage but costs more per asset. Most programs use a hybrid: continuous on the most critical assets, route-based on the rest.
Can small plants benefit from condition monitoring? Yes. A small plant with 50 critical assets can start with a portable vibration analyzer, a basic ultrasound instrument, and quarterly contracted thermographic surveys for well under $50,000 in the first year. A single prevented emergency on a critical machine — especially one that would have required overtime weekend repair and production loss — pays for the program. Start small, prove the value, and expand.
Related VETTED training modules
The following modules from the VETTED training library build the skills and knowledge referenced in this guide:
-
Condition Monitoring & Predictive Maintenance — The maintenance strategies and the P-F curve, the major condition-monitoring technologies (vibration, thermography, oil analysis, ultrasound), and how to turn readings into the right action at the right time. The foundational module for this topic.
-
Maintenance Strategy & Reliability Metrics — Choose the right maintenance strategy for each asset and measure what matters: MTBF, MTTR, availability, the P-F curve, and how reliability-centered thinking lowers cost and downtime. Builds the strategic framework for where condition monitoring fits.
-
Vibration Analysis Fundamentals — Intermediate condition monitoring: read a vibration spectrum and recognize the signatures of imbalance, misalignment, looseness, and bearing defects to catch faults early. The deep dive into the most widely deployed monitoring technology.
-
Bearing & Lubrication Basics — How rolling-element bearings work, what lubrication really does, how over- and under-greasing destroy bearings, contamination control, and the "Five Rights" of a reliable lubrication program. Essential context for understanding what condition monitoring detects in bearing systems.
-
Root Cause Analysis & FMEA for Maintenance — Stop firefighting: use 5-Why, Fishbone, and FMEA to find and eliminate the true causes of failures, and prioritize what to fix first by risk. The analytical complement — once monitoring detects a fault, RCA finds the systemic cause.
This article is general educational content for maintenance and reliability professionals. It is not a substitute for qualified engineering judgment, site-specific analysis, or consultation with a certified condition monitoring analyst. Standards, technologies, and best practices evolve — verify current requirements and recommendations for your application.