Root cause failure analysis: the quick answer根本原因故障分析:快速回答
Place this specific workflow in context with the anomaly detection and root cause analysis guide, which connects the definitions, alternatives, validation steps, and related implementation guides.
可通过异常检测与根因分析指南理解本专题在整体流程中的位置;该指南串联了定义、替代方案、验证步骤与相关实施文章。
Root cause failure analysis is a structured investigation of what failed, how it failed, and why the organization’s design, operation, maintenance, or controls allowed that mechanism to occur. A defensible RCFA preserves evidence before repair, defines the failure mode, proves the mechanism, tests competing causes, and verifies that targeted corrective action changes the outcome without creating unacceptable side effects.
根本原因故障分析是一种结构化调查,用于回答什么失效、如何失效,以及组织的设计、运行、维护或控制为何允许该机制发生。可信的 RCFA 会在维修前保护证据,定义失效模式,证明失效机制,检验竞争原因,并验证针对性纠正措施确实改变结果且未产生不可接受的副作用。
What root cause failure analysis does—and does not do根本原因故障分析能做什么、不能做什么
The phrase is most useful in reliability, manufacturing, maintenance, quality, and engineering investigations where a component, asset, process, or system no longer performs its intended function. The investigation moves through three linked questions: the failure mode describes the observed loss of function; the failure mechanism explains the physical, chemical, electrical, software, or process phenomenon that produced it; and the causal system explains why the mechanism was created, missed, or tolerated.
该术语主要用于可靠性、制造、维护、质量与工程调查,适用于部件、资产、流程或系统无法实现预期功能的情况。调查依次回答三个关联问题:失效模式描述观察到的功能丧失;失效机制解释导致它的物理、化学、电气、软件或流程现象;因果系统解释该机制为何被制造出来、未被发现或被长期容忍。
RCFA is not a meeting that ends when a plausible cause is named. It is not automatically a laboratory test, a 5 Whys worksheet, a blame assignment, or a guarantee that recurrence is impossible. Specialized examination may be required for fracture, corrosion, contamination, electronics, fire, or safety-critical events. When evidence has legal, regulatory, or safety implications, use qualified specialists and the organization’s chain-of-custody and reporting procedures.
RCFA 不是在说出一个看似合理原因后就结束的会议,也不等同于实验室检测、五问法表格、责任归咎或“绝不复发”的保证。断裂、腐蚀、污染、电子故障、火灾或安全关键事件可能需要专门检验;当证据涉及法律、监管或安全后果时,应使用合格专家以及组织规定的证据保全与报告流程。
Prepare the evidence before disassembly or repair拆解或维修前先准备证据
Start with a factual problem statement: asset and function, expected state, observed deviation, first known occurrence, operating state, consequence, and investigation boundary. Photograph the as-found condition. Label and isolate failed parts, mating parts, fluids, filters, debris, and relevant samples. Preserve control logic, alarms, high-resolution trends, work orders, operator notes, configuration, calibration, installation history, environmental conditions, and recent changes. Record timestamps, units, time zones, collection method, source, and custodian.
从事实性问题陈述开始:资产与功能、预期状态、观察偏差、首次已知发生时间、运行状态、后果和调查边界。拍摄发现时状态,标记并隔离故障件、配合件、流体、过滤器、碎屑与相关样品。保存控制逻辑、告警、高分辨率趋势、工单、操作员记录、配置、校准、安装历史、环境条件和近期变更;记录时间戳、单位、时区、采集方法、来源和保管人。
Preservation rule: do not clean, machine, straighten, power, reset, patch, or destructively section the item until the investigation owner has documented the as-found condition and approved the test sequence. Begin with non-destructive observations because later steps may erase evidence.
证据保护规则:在调查负责人记录发现时状态并批准测试顺序前,不要清洗、加工、校直、通电、复位、打补丁或破坏性切割故障件。应先做无损观察,因为后续步骤可能永久抹去证据。
A step-by-step root cause failure analysis process根本原因故障分析的分步流程
- Stabilize and preserve.稳定现场并保护证据。Make the condition safe, control access, record the as-found state, quarantine evidence, and document every change made during recovery.确保状态安全,控制访问,记录发现时状态,隔离证据,并记录恢复期间做出的每项变更。
- Define the failure mode and scope.定义失效模式与范围。State the required function and observed loss without embedding a cause. Identify affected and unaffected assets, products, time windows, loads, and configurations.说明所需功能与观察到的功能丧失,不在描述中预设原因;识别受影响与未受影响的资产、产品、时间窗口、负载和配置。
- Reconstruct the timeline.重建时间线。Align operating data, alarms, maintenance, changes, inspections, and witness accounts. Mark clock uncertainty, missing intervals, and evidence recorded after the event.对齐运行数据、告警、维护、变更、检查与见证记录,标注时钟不确定性、缺失区间和事后记录。
- Characterize the failure mechanism.表征失效机制。Use appropriate visual, dimensional, electrical, chemical, materials, software, or process tests. Sequence tests from non-destructive to destructive and compare with known-good or unaffected samples.采用适当的目视、尺寸、电气、化学、材料、软件或流程测试;按无损到破坏性顺序执行,并与正常或未受影响样本比较。
- Generate causes at several layers.生成多层候选原因。Ask what physical condition created the mechanism, what operating or human action contributed, and what latent design, maintenance, procurement, training, or detection control allowed it.追问什么物理条件制造了该机制、哪些运行或人为行为促成了它,以及哪些潜在设计、维护、采购、培训或检测控制允许它发生。
- Test competing hypotheses.检验竞争假设。For each candidate, write predicted evidence and a falsification test. Use counterexamples, replication, simulation, controlled substitution, teardown, or additional measurement when safe and appropriate.为每个候选原因写出预测证据和可推翻它的测试;在安全且适当时使用反例、复现、仿真、受控替换、拆解或补充测量。
- Select and map corrective actions.选择并映射纠正措施。Link every action to a verified mechanism, contributor, or failed barrier. Prefer controls that change design or process conditions over reminders that depend on perfect memory.把每项措施链接到已验证的机制、促成因素或失效屏障;优先选择改变设计或流程条件的控制,而非依赖完美记忆的提醒。
- Verify, monitor, and close.验证、监测并结案。Define baseline, target, data source, observation window, side-effect checks, and reopening criteria. Confirm both implementation and effectiveness before closing the record.定义基线、目标、数据来源、观察窗口、副作用检查与重新开启标准;结案前同时确认措施已实施且有效。
Separate failure mode, mechanism, and cause layers区分失效模式、机制与原因层级
| Layer层级 | Question问题 | Evidence example证据示例 |
|---|---|---|
| Failure mode失效模式 | What function was lost?什么功能丧失? | Bearing seized; service timed out; seal leaked.轴承抱死;服务超时;密封泄漏。 |
| Failure mechanism失效机制 | What phenomenon produced the mode?什么现象造成该模式? | Rolling-contact fatigue, thermal overload, race condition, abrasive wear.滚动接触疲劳、热过载、竞态条件、磨粒磨损。 |
| Physical cause物理原因 | What condition initiated or accelerated it?什么条件启动或加速该机制? | Contaminated lubricant, misalignment, voltage transient, invalid input.润滑污染、错位、电压瞬变、无效输入。 |
| Human or operating contributor人为或运行促因 | Which action or context changed exposure?何种行为或环境改变了暴露? | Incorrect assembly step, operation outside envelope, deferred inspection.装配步骤错误、超范围运行、延期检查。 |
| Latent system cause潜在系统原因 | Why was the condition created or not detected?为何制造或未发现该条件? | Ambiguous specification, missing poka-yoke, weak change control, ineffective alarm.规范含糊、防错缺失、变更控制薄弱、告警无效。 |
Do not force one label to carry the whole explanation. A broken bearing is evidence of a failure mode; fatigue may be the mechanism; misalignment may be a physical cause; an installation method may be a contributor; and an absent alignment acceptance test may be the latent control failure. The full chain makes corrective action reviewable.
不要强迫一个标签解释整条因果链。轴承损坏是失效模式的证据;疲劳可能是机制;错位可能是物理原因;安装方法可能是促因;缺少对中验收测试可能是潜在控制失效。完整链条使纠正措施可以被复核。
Root cause failure analysis example: repeated bearing damage根本原因故障分析示例:轴承重复损坏
Hypothetical example: a process pump trips on high vibration six weeks after a bearing replacement. The team does not assume “bad bearing.” It quarantines the bearing and lubricant, photographs the assembly, exports vibration and temperature history, collects alignment readings, reviews the work order, and compares the failed pump with a matched unaffected unit.
假设示例:某工艺泵在更换轴承六周后因高振动跳停。团队没有直接认定“轴承质量差”,而是隔离轴承与润滑剂、拍摄装配状态、导出振动和温度历史、采集对中读数、复核工单,并与条件匹配的正常泵比较。
Examination finds one-sided raceway loading and wear consistent with misalignment; lubricant chemistry and bearing dimensions remain within their specified checks. The alignment record shows an acceptable reading immediately after installation, but the test was taken before connected piping returned to operating temperature. A controlled hot-alignment check reproduces an offset on the failed pump, while the unaffected unit remains within the organization’s approved limit. This supports a physical cause—thermally induced misalignment—rather than contamination or a defective replacement part.
检查发现滚道单侧受载和与错位一致的磨损;润滑剂化学状态与轴承尺寸在规定检查范围内。对中记录显示安装后读数合格,但测量发生在连接管道恢复运行温度之前。受控热态对中检查在故障泵上复现偏移,而正常泵仍在组织批准限值内。这支持“热致错位”这一物理原因,而非污染或替换件缺陷。
The deeper review finds that the work instruction requires cold alignment only, the acceptance form has no field for thermal growth assumptions, and no post-start vibration review is assigned. The team records these as latent system causes. Corrective actions revise the engineering requirement, add a hot-condition verification where technically appropriate, and assign post-start monitoring. Effectiveness is judged over an approved operating window using alignment, vibration trend, temperature, and recurrence—not merely by closing the work order. All values in this example are illustrative; real limits must come from the asset design and approved engineering procedure.
进一步复核发现,作业指导书只要求冷态对中,验收表没有热膨胀假设字段,也没有分配启动后振动复核。团队把这些记录为潜在系统原因。纠正措施包括修订工程要求、在技术上适当时增加热态验证,并指定启动后监测。措施效果应在批准的运行窗口内用对中、振动趋势、温度与复发情况判断,而不是以工单关闭作为成功证明。示例中的所有数值均为说明性质;实际限值必须来自资产设计与批准的工程程序。
Failure analysis vs root cause analysis vs FMEA失效分析、根因分析与 FMEA 的区别
Failure analysis concentrates on the failed item or function: location, mode, mechanism, and technical evidence. Root cause analysis extends beyond the item to operating context, design decisions, maintenance, human factors, and prevention or detection controls. RCFA deliberately joins both views. Failure mode and effects analysis (FMEA) is normally prospective: it asks how a design or process could fail and prioritizes risk before an event. RCFA is retrospective: it investigates a failure that actually occurred.
失效分析集中于故障对象或功能:失效位置、模式、机制与技术证据。根因分析从对象向外扩展到运行环境、设计决策、维护、人因以及预防或检测控制。RCFA 有意连接两种视角。失效模式与影响分析(FMEA)通常是前瞻性的:在事件前询问设计或流程可能如何失效并确定风险优先级;RCFA 是回顾性的:调查已经实际发生的故障。
The methods complement one another. Verified RCFA findings can update an FMEA’s failure modes, controls, occurrence assumptions, or detection assumptions. An existing FMEA can suggest hypotheses during RCFA, but its risk score is not proof that a particular mechanism caused the observed event.
这些方法可以互补。经验证的 RCFA 结论可更新 FMEA 的失效模式、控制、发生度假设或探测度假设;已有 FMEA 也可为 RCFA 提供候选假设,但风险评分不能证明某个具体机制导致了本次事件。
Analyze failure evidence across connected data and files跨关联数据与文件分析故障证据
Before opening the workspace, prepare a scoped case folder, stable asset and record IDs, normalized timestamps, sensor exports, work orders, inspection reports, images, test results, change records, and access permissions. InfiniSynapse is an AI-powered workspace for analysis across connected databases, files, documents, audio, and video; it is not presented here as a laboratory, a dedicated certified RCFA package, or a substitute for qualified engineering judgment. Use it to retrieve and compare prepared evidence, summarize timelines, and explore hypotheses, then verify consequential findings through approved tests and review.
打开工作区前,请准备范围明确的案件文件夹、稳定的资产与记录 ID、规范化时间戳、传感器导出、工单、检查报告、图片、测试结果、变更记录与访问权限。InfiniSynapse 是用于分析关联数据库、文件、文档、音频和视频的 AI 辅助工作区;本页不把它描述为实验室、专用认证 RCFA 软件或合格工程判断的替代品。可用它检索和比较准备好的证据、汇总时间线并探索假设,再通过批准的测试与复核验证重要结论。
Open InfiniSynapse for connected evidence analysis打开 InfiniSynapse 进行关联证据分析Verify causes and avoid common RCFA failures验证原因并避免常见 RCFA 失误
A credible cause explains the failure’s location, timing, scope, and mechanism; predicts observations in affected and unaffected cases; fits evidence better than alternatives; and survives a deliberate attempt to disprove it. Where safe, controlling the cause should change a measurable outcome. If direct intervention is impossible, use converging evidence and state the remaining uncertainty.
可信原因应解释故障的位置、时间、范围和机制;预测受影响与未受影响案例中的观察结果;比替代解释更符合证据;并经受主动反证。若安全可行,控制该原因应改变可测结果;若无法直接干预,应使用相互印证的证据并说明剩余不确定性。
- Do not repair before preserving evidence. Recovery can erase fracture surfaces, timestamps, configuration, contamination, and as-found geometry.不要在保护证据前维修。恢复操作可能抹去断口、时间戳、配置、污染和发现时几何状态。
- Do not stop at “operator error” or “component failed.” Ask what conditions made the action likely and what barrier should have prevented or detected it.不要停在“操作员错误”或“部件坏了”。应追问什么条件使行为更可能发生,以及什么屏障本应预防或发现它。
- Do not confuse correlation with mechanism. A trend that moves with the failure is a lead until timing, pathway, and counterevidence are tested.不要把相关性当作机制。与故障同步变化的趋势只是线索,必须检验时间顺序、作用路径和反证。
- Do not verify only implementation. “Training delivered” or “part replaced” confirms activity, not that risk or recurrence changed.不要只验证实施。“培训已完成”或“部件已更换”只能证明活动发生,不能证明风险或复发已改变。
Keep an evidence register, hypothesis table, test record, decision log, action-to-cause map, and effectiveness plan. Reopen the investigation when recurrence, contradictory evidence, a new mechanism, or an unacceptable side effect challenges the accepted explanation. For adjacent product workflows, review the InfiniSynapse tool catalog and product documentation.
保留证据登记表、假设表、测试记录、决策日志、措施—原因映射与效果验证计划。当复发、矛盾证据、新机制或不可接受的副作用挑战既有解释时,应重新开启调查。相关产品工作流可查看 InfiniSynapse 工具目录和产品文档。
Frequently asked questions about root cause failure analysis根本原因故障分析常见问题
What is root cause failure analysis?什么是根本原因故障分析?
Root cause failure analysis is a structured, evidence-based investigation that determines what failed, the mechanism by which it failed, and the physical, human, and latent system causes that allowed the failure.
根本原因故障分析是一种结构化、基于证据的调查,用于确定什么失效、以何种机制失效,以及允许故障发生的物理、人为和潜在系统原因。
What is the difference between failure analysis and root cause analysis?失效分析与根因分析有什么区别?
Failure analysis identifies the failed location, mode, and mechanism. Root cause analysis extends outward to the operating, design, maintenance, human, and management conditions that created or failed to prevent that mechanism.
失效分析识别失效位置、模式与机制;根因分析向外扩展到制造该机制或未能阻止它的运行、设计、维护、人因与管理条件。
How is RCFA different from FMEA?RCFA 与 FMEA 有何区别?
RCFA is retrospective and investigates an observed failure. FMEA is prospective and asks how a design or process might fail so risks can be prioritized before an event.
RCFA 是回顾性的,调查已经观察到的故障;FMEA 是前瞻性的,在事件前询问设计或流程可能如何失效,以便确定风险优先级。
How do you verify a root cause in failure analysis?如何验证故障分析中的根因?
A cause is credible when it explains the evidence and failure mechanism, predicts affected and unaffected cases, survives attempts to disprove it, and produces a measurable change when controlled.
当某原因能够解释证据与失效机制、预测受影响与未受影响案例、经受反证,并在被控制后产生可测变化时,才具有可信度。
When should a failed part be preserved?何时应保护故障件?
Preserve it before cleaning, repair, disassembly, or destructive testing whenever the failure is safety-critical, recurring, expensive, legally sensitive, novel, or technically uncertain.
当故障涉及安全关键、重复、高成本、法律敏感、新型或技术不确定情况时,应在清洗、维修、拆解或破坏性测试前保护故障件。
Official sources and verification notes官方来源与验证说明
These sources support the distinction between physical failure analysis and broader root cause analysis, plus the evidence-led investigation sequence used here. The comparison table and bearing example are editorial synthesis. Technical limits, standards, and organizational requirements vary; verify current approved documentation before consequential use.
这些来源支持物理失效分析与更广义根因分析之间的区别,以及本文采用的证据导向调查顺序。比较表和轴承案例为编辑整理。技术限值、标准和组织要求各不相同;在重要用途前请核对当前批准的文档。
Editorial boundary: this guide is educational content, not a laboratory report, legal opinion, safety approval, or certification. It does not claim guaranteed prevention, indexing, rankings, or traffic.
编辑边界:本指南属于教育内容,不是实验室报告、法律意见、安全批准或认证,也不承诺一定防止复发、被搜索引擎收录、获得排名或流量。
InfiniSynapse