Question 7 of 10: Simulation Validation — Comparing Variances and Means to the Actual System
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Notes on this paper
National Exams May 2019 — 17-Ind-A6, Systems Simulation, 3 hours duration. Section A: do 3 of 5 (30 marks); Section B: do 1 of 2 per the front-page summary table — the Part B section instructions on the page itself read “do 2 of 3” instead, a genuine inconsistency in the source paper's own front matter (flagged below); Section C: do 1 of 2 (20 marks). This study resource answers all ten questions across the three parts.
Check — front-page/section instructions disagree. The cover page's summary table states “Section B: Do 1 of 2 Questions, Total marks: 20,” but Part B's own section banner (page 4) reads “Complete two of the following three sets of questions,” and Part B in fact contains three questions worth 15 marks each. The cover page's own overall total (70 marks across “5 Questions”) is likewise only consistent with the “2 of 3” reading (30+30+... does not match 70 either way exactly, since 30(A)+2×15(B)+20(C)=80, one section choice short of 70); this is an internal inconsistency in the exam's own printed materials. All three Part B questions are answered in full below regardless.
Reference texts. Banks, Carson, Nelson & Nicol, Discrete-Event System Simulation (5th ed.) — primary text for random-variate generation, input/output analysis, variance reduction, verification & validation, and design of simulation experiments.
Question 7 (Part B.2): Simulation Validation — Comparing Variances and Means to the Actual System (15 marks)
Find. (a) an $F$-test for equal variances; (b) a test for equal means, conditional on (a)'s outcome, and a validity conclusion at $\alpha=0.05$; (c) cautions on the conclusion; (d) whether unequal variances with equal means implies validity.
Approach. Two-sample $F$-test on variances first (this determines which $t$-test is appropriate); Welch's (unequal-variance) $t$-test on means if (a) rejects equal variance; interpret jointly.
(a) F-test for equal variances. $H_0:\sigma_1^2=\sigma_2^2$ vs. $H_1:\sigma_1^2\ne\sigma_2^2$. With the larger sample variance in the numerator,
$$F=\frac{s_1^2}{s_2^2}=\frac{1.428}{0.2407}=5.93,\qquad F_{0.025,6,6}=5.82.$$
Since $5.93>5.82$ (two-sided $p=0.048<0.05$),
$$\boxed{\text{reject}\ H_0\ \text{at}\ \alpha=0.05 \;-\; \text{the variances are statistically different}}$$
(the actual system's day-to-day admissions vary noticeably more than the simulation's replicate means do).
(b) Test the means and conclude on validity. Because (a) found unequal variances, the appropriate comparison of means is Welch's (unequal-variance) $t$-test rather than the pooled-variance test:
$$t=\frac{71.657-70.574}{\sqrt{1.428/7+0.2407/7}}=2.218,\qquad df_{\text{Welch}}\approx7.97,\qquad t_{0.025,8}=2.306.$$
Since $|2.218|<2.306$ ($p\approx0.058>0.05$),
$$\boxed{\text{fail to reject}\ H_0:\mu_1=\mu_2 \;-\; \text{the means are not statistically different at}\ \alpha=0.05}$$
So the picture is mixed: the simulation reproduces the actual system's average admission rate, but not its variability. Matching the mean alone is not sufficient grounds to call the model an accurate representation — see (d) — so the IE cannot fully conclude her model is valid; it captures central tendency but under-represents day-to-day variability.
(c) Cautions. A hypothesis test that fails to reject $H_0$ is not proof the means are equal — it only means the data don't provide strong enough evidence of a difference, especially with only $n=7$ replications, which gives limited power to detect a real but modest mean gap. At $\alpha=0.05$ there is also a genuine (if small) chance the variance test's rejection is itself a Type I error. Beyond these statistical caveats, matching one aggregate output (mean daily admissions) at one comparison point does not validate the model across all the behaviours it will be used to predict — the IE should also check face validity with subject-matter experts, sensitivity/extreme-condition tests, and comparisons on other output measures (e.g. peak-hour admissions, seasonal pattern) before relying on the model for decisions.
(d) Unequal variances, equal means — is the model valid? No, not fully. Getting the mean right while getting the variance wrong means the simulation reproduces the actual system's central tendency but misrepresents its spread — and spread matters directly for anything downstream that depends on variability (e.g. staffing to cover peak admissions, queueing/wait-time predictions, risk of exceeding capacity), where the actual system's own larger variance would, in this exact scenario, be under-predicted by the model.
$$\boxed{\text{Not fully valid} - \text{matching the mean does not compensate for a demonstrably wrong variance}}$$
The correct response is to treat this specific result the same as this question's own (b)/(d) case: report both findings honestly, and either recalibrate the model's sources of variability or restrict its use to applications where only the mean matters.