23-Ind-A6 Systems Simulation · December 2018
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
National Exams — December 2018 — 17-Ind-A6 Systems Simulation. Three-hour, closed-book exam; one of two permitted calculators (Sharp or Casio), one 8.5″×11.0″ aid sheet (both sides). Format: three sections — Part A (Input Modelling): do 2 of 3 questions, 30 marks total; Part B (Modelling Concepts): do 1 of 2, 20 marks; Part C (Output Analysis): do 1 of 2, 20 marks — plus a 1-mark trivia bonus. All seven graded questions and the bonus are solved below for completeness. The exam's own front matter has two internal quirks, transcribed as printed: NOTES item "4." appears twice on page 1, and every page footer reads "17-Ind-A6/Dec. 2019" against a "December 2018" masthead.
Reference texts: Banks, Carson, Nelson & Nicol, Discrete-Event System Simulation (5th ed., Pearson) — input data analysis (ch. 9), random-variate generation, output analysis for a single system and comparing alternative systems (ch. 11–12), verification and validation (ch. 10).
Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.
Given. A bags-per-patrol histogram from the 2013 Kananaskis Jamboree: a sharp mode near 3 bags, a broad second mode spanning roughly 17–20 bags, and near-zero frequency in between (around 13–16 bags).
Find. (a) the population interpretation of multi-modality; (b) the practical modelling difficulty it creates; (c) methods to resolve it, including a mathematical splitting technique; (d) empirical vs. theoretical distribution trade-offs.
(a) What a multi-modal chart implies. A histogram with two (or more) separated peaks almost never arises from one homogeneous random process; it is the classic fingerprint of a mixture of two or more distinct sub-populations, each individually well-behaved (often close to unimodal), whose data have been pooled into a single sample. Here, the low mode near 3 bags is consistent with patrols attending only day activities (little gear, one short trip), while the high mode near 17–20 bags is consistent with patrols staying the full multi-day Jamboree (full camping kit, spare clothing, group equipment) — two structurally different "kinds" of patrol trip being counted together as one variable.
(b) Practical difficulty for input modelling. Every standard theoretical input distribution used in simulation software (exponential, normal, gamma, Poisson, triangular, etc.) is inherently unimodal, so none of them can honestly reproduce two separated peaks — forcing a single theoretical fit onto multi-modal data either smears a distribution across the empty valley between modes (fabricating "typical" values, e.g. 13–16 bags, that almost never actually occur) or collapses onto just one mode and ignores the other population entirely. A formal goodness-of-fit test (chi-square, K-S) will predictably fail against any single-family candidate, but the more dangerous failure is a model that "passes" a loose test while still generating physically wrong scenarios in the simulation.
(c) Methods to improve the fit. More sampling alone would not fix this — a bigger sample from the same mixed population reproduces the same two-mode shape more precisely, it does not merge the modes. The productive approaches: (i) stratify by the covariate that actually drives the split — if trip type (day vs. overnight) is recorded or recoverable, split the sample on that label and fit each sub-population separately; (ii) if no covariate is available, transform/partition mathematically: locate the trough between modes (here around 13–16 bags) as a threshold, assign each observation to "low" or "high" by which side of the threshold it falls on, and fit a separate unimodal distribution (e.g. two triangular or two gamma distributions) to each resulting sub-sample; (iii) more rigorously, fit a two-component mixture model $f(x)=p\,f_1(x)+(1-p)f_2(x)$ via the EM algorithm, which simultaneously estimates each component's parameters and the mixing weight $p$ (the proportion of day-trip patrols) without requiring a hard threshold cut.
(d) Empirical vs. theoretical distributions. Benefits of an empirical distribution: it reproduces the observed shape exactly, multi-modality included, with no risk of mis-specifying a family; it requires no distributional assumption or goodness-of-fit justification at all. Drawbacks: it can never generate a value outside the observed sample's range (no extrapolation beyond the smallest/largest recorded bag count, which understates the true tail risk of transport planning); it faithfully reproduces sampling noise along with the real signal (small bumps/gaps in a modest sample are treated as real features of the population); it gives no closed-form parameters to communicate, sensitivity-test, or generalize to a different-sized Jamboree; and, for continuous quantities, it needs an interpolation rule (piecewise-linear CDF) between observed points, which is itself an assumption. A theoretical fit, once validated, is smoother, extrapolates sensibly, and is far easier to re-parametrize for "what-if" studies (e.g. a Jamboree twice this size) — but only if the underlying population really is unimodal, which this data set demonstrably is not.
| Sub-part | Answer summary |
|---|---|
| (a) | Two peaks = mixture of two sub-populations (day-trip vs. multi-day patrols) |
| (b) | No standard unimodal family can honestly represent two separated modes |
| (c) | Stratify by covariate, or split at the trough and fit two unimodal distributions, or fit a 2-component mixture (EM) |
| (d) | Empirical: exact shape, no extrapolation; Theoretical: smoother/extrapolates, needs the right (multi-component) family |