25-Comp-A6 Software Engineering · May 2014
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
National Exams — May 2014 — 98-Comp-A6 Software Engineering. Three-hour, closed-book exam, no calculator permitted. Format: nine questions, candidates answer any five of the nine (all questions equal weight — each of the five counted questions is worth 20%; only the first five questions as they appear in the answer book are marked). All nine questions are solved below for completeness.
Reference texts: Sommerville, Software Engineering (10th ed., Pearson) — software process models, object-oriented design, formal methods, real-time systems, software testing, project management, critical/dependable systems, software quality, distributed systems; Pressman, Software Engineering: A Practitioner's Approach (9th ed.) — supplementary process/testing/quality coverage; IEEE 12207 — software life-cycle processes; SWEBOK — body-of-knowledge cross-reference.
Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.
A hardware component fails because it physically wears out or is physically damaged — its failure rate is governed by material fatigue, environmental stress, and manufacturing variance, and is (after an initial "infant mortality" period) reasonably well modelled as roughly constant or slowly increasing with age, which is exactly what lets metrics like MTBF (mean time between failures) be meaningfully estimated and interpreted for hardware. A software component fails only because it contains a design defect that was present, latent, from the moment it was written; it does not wear out, and repeated identical inputs always produce the identical (possibly wrong) output. Consequently, a software "failure rate" is not a physical process at all — it is purely a function of which subset of the input space happens to be exercised by the current usage pattern, so the same code can appear highly reliable under one usage profile and fail constantly under another, with no physical explanation and no meaningful analogue of "wear." Hardware reliability metrics built on the assumption of physical, roughly-random component failure (and, correspondingly, that repair replaces a worn part with a statistically identical new one) are therefore inappropriate for software: "repairing" a software defect does not restore the system to some earlier reliability state, it produces a different program with a different, unknown defect population, so there is no valid physical failure-rate model underneath the number even if a similarly-named metric is reported.
Three plausible hazards, each with a defensive requirement and the reasoning for why it reduces risk:
| # | Hazard | Defensive requirement | Why it reduces risk |
|---|---|---|---|
| 1 | An overdose is delivered because a treatment record is corrupted or wrongly transcribed between the database and the machine (e.g. a decimal-point/units error, or a dose intended for a different patient is downloaded). | The embedded controller shall independently validate every downloaded dose against a hard-coded, clinically-reasonable maximum-dose ceiling and against the specific patient/site identifier before it is allowed to fire, and shall require an explicit operator confirmation showing patient ID, site, and dose before treatment begins. | The ceiling check catches any corruption or transcription error that produces a physically implausible dose regardless of its root cause, without needing to know what that cause was; the operator confirmation step adds an independent human check against a mis-match that is clinically plausible in magnitude but wrong for this specific patient, which a simple magnitude ceiling cannot catch on its own. |
| 2 | The machine begins or continues to deliver radiation to the wrong body site because the beam-positioning mechanism is out of calibration, moved, or was not correctly homed at start-up (this is the class of hazard implicated in the real Therac-25 incidents). | The controller shall independently verify, via a separate sensor path from the one used to command the beam, that the actual physical beam-shaping/positioning state matches the commanded state before enabling radiation, and shall interlock (refuse to fire) if the two disagree. | A single sensor/actuator path cannot distinguish "commanded correctly and executed correctly" from "commanded correctly but executed wrongly due to a mechanical or software fault" — requiring independent confirmation from a second, physically distinct measurement closes exactly the failure mode where the software believes the machine is in the commanded state but it is not. |
| 3 | A race condition or unhandled concurrent-access fault in the embedded software allows an operator to change beam parameters (e.g. switch from a low-power diagnostic mode to a high-power treatment mode) while a previous command is still being processed, resulting in an inconsistent combination of settings being applied — also directly analogous to a documented Therac-25 failure mode. | The controller shall enforce a state machine that only accepts a new treatment command once the previous command's beam parameters have been fully applied and independently confirmed, rejecting (rather than queuing or silently merging) any command received while a state transition is in progress. | Refusing rather than queuing an out-of-sequence command removes the possibility of the software combining two different partially-applied parameter sets, which is exactly the mechanism by which a fast-typing operator was able to trigger a massive overdose in the historical incident this hazard is modelled on; an explicit reject-and-require-retry is a much stronger defence than merely making the race condition less likely to occur. |