Question 2 of 7: Fault Tree Analysis, Job Safety Analysis, and Failure Modes and Effects Analysis
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Notes on this paper
National Exams — May 2016 — 98-Ind-B10 Industrial Safety and Health. Closed book; no calculators permitted. Any five of the seven questions constitute a complete paper; all questions are of equal value (20 marks each). Answers are written in point form but fully, as instructed. Complete answers to all seven questions follow, with assumptions stated where the question invites them.
Reference texts: Brauer, Safety and Health for Engineers, 4th ed.; CCPS (Center for Chemical Process Safety), Guidelines for Risk Based Process Safety; CSA Z1002 Occupational health and safety — Hazard identification and elimination and risk assessment and control; CSA Z1006 Management of work in confined spaces.
Question 2: Fault Tree Analysis, Job Safety Analysis, and Failure Modes and Effects Analysis (20 marks: 6/7/7)
(i) Fault Tree Analysis (FTA) in Accident Investigation
"Fault-free analysis" as printed on the exam paper is evidently a typographical error for fault tree analysis (there is no technique called fault-free analysis; the pairing with FMEA and accident investigation confirms the intent) — a deductive, top-down logic technique. Used in accident investigation, the analyst starts from the top event (the accident or loss that actually occurred, e.g. "explosion in reactor vessel") and works backward, asking "what combination of failures could have caused this?" at each level, connecting contributing events with AND-gates (all must occur together) and OR-gates (any one is sufficient) until the tree terminates in basic events — individual component failures, human errors, or environmental conditions that are not further decomposed. Applied after an accident, FTA:
Organizes every plausible causal path into a single logical diagram, so investigators do not fixate on the first explanation that presents itself.
Distinguishes minimal cut sets — the smallest combinations of basic-event failures that alone are sufficient to produce the top event — which tells investigators exactly which barriers failed and which held.
Provides a quantitative capability: if failure probabilities are known or can be estimated for each basic event, the top-event probability can be calculated (using Boolean algebra on the AND/OR gate logic), which supports both accident reconstruction and future risk-based decisions.
Limitations: FTA requires the analyst to know, in advance, the full set of ways a system can fail — a fault not conceived of cannot appear in the tree, and complex systems can develop enormous trees that are hard to construct and validate completely. Basic-event probabilities are often poorly known (especially for rare, catastrophic events), which limits the confidence of a quantitative result; the logic gates assume events are independent, which understates risk when common-cause failures (a single power loss, a single operator, a single flood) can defeat multiple "independent" branches simultaneously; and a tree built purely deductively after the fact can be unconsciously shaped to fit the investigator's initial hypothesis about what happened, rather than genuinely exploring alternative causal paths.
(ii) Purpose and Steps of Job Safety Analysis (JSA)
Purpose: JSA (also called Job Hazard Analysis) is a proactive technique that systematically studies a specific job or task before it is performed, to identify hazards inherent in each step and establish the safe procedure for doing it — preventing accidents by design of the task, rather than reacting after one occurs. It is one of the primary tools by which System Safety Engineering (Question 1(i)) is applied at the individual-task level, and is a foundation document for worker training, standard operating procedures, and new-employee orientation.
Select the job — prioritize by accident frequency/severity history, job complexity, new or modified tasks, and jobs with a history of near-misses.
Break the job into its basic sequential steps — observe (or have an experienced worker describe) the task, recording each distinct step in the order it is actually performed, without yet judging hazards.
Identify the hazards associated with each step — ask, for every step, what could go wrong: struck-by, caught-in, fall, exposure, ergonomic strain, and what conditions (equipment, environment, material) create that possibility.
Develop the recommended safe procedure/controls for each hazard — applying the hierarchy of controls (elimination, substitution, engineering, administrative, PPE), so each identified hazard has a specific, actionable control tied to it, not a generic caution.
Document, review, and communicate the JSA — record it in a standard format, have it reviewed by supervision/safety personnel and, ideally, the workers who perform the job, then use it for training and post it at the work location.
Revise periodically — after any incident/near-miss involving the job, after a process or equipment change, and on a scheduled review cycle, since a JSA is only valid for the conditions under which it was written.
(iii) FMEA in Reliability Engineering
Failure Modes and Effects Analysis (FMEA) is a systematic, bottom-up, inductive technique — the logical complement of FTA's top-down deduction. Rather than starting from an undesired top event, FMEA starts from the system's components and asks, for each one: in what ways (modes) can this component fail, what causes each failure mode, and what effect does that failure have on the subsystem, the system, and ultimately safety and mission success? In reliability engineering specifically, FMEA is used to:
Identify single points of failure — a component whose failure alone (with no other coincident failure) propagates to a critical system-level effect, which is exactly the class of failure a designer must eliminate or protect against with redundancy.
Rank failure modes by criticality — when extended to FMECA (Failure Modes, Effects, and Criticality Analysis), each mode is scored on severity, occurrence probability, and (in some variants) detectability, and the product (a Risk Priority Number) directs engineering attention to the modes that matter most.
Feed design improvement and maintenance planning — driving design changes (better materials, added redundancy, derating) for high-criticality modes, and informing preventive-maintenance/inspection intervals set to catch a failure mode before it propagates.
Complement FTA — FMEA is exhaustive at the component level (it will not miss a failure mode of a component that is included in the analysis) but can miss combination/interaction failures across components, which is precisely FTA's strength; a mature reliability program uses both.