NivaarExam PrepOfficial exam papers ↗

23-Ind-A6 Systems Simulation · December 2018

Question 5 of 8: Model Validation, Autocorrelation, and Variance Comparison

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

National Exams — December 2018 — 17-Ind-A6 Systems Simulation. Three-hour, closed-book exam; one of two permitted calculators (Sharp or Casio), one 8.5″×11.0″ aid sheet (both sides). Format: three sections — Part A (Input Modelling): do 2 of 3 questions, 30 marks total; Part B (Modelling Concepts): do 1 of 2, 20 marks; Part C (Output Analysis): do 1 of 2, 20 marks — plus a 1-mark trivia bonus. All seven graded questions and the bonus are solved below for completeness. The exam's own front matter has two internal quirks, transcribed as printed: NOTES item "4." appears twice on page 1, and every page footer reads "17-Ind-A6/Dec. 2019" against a "December 2018" masthead.

Reference texts: Banks, Carson, Nelson & Nicol, Discrete-Event System Simulation (5th ed., Pearson) — input data analysis (ch. 9), random-variate generation, output analysis for a single system and comparing alternative systems (ch. 11–12), verification and validation (ch. 10).

Question 5 (Part B.2): Model Validation, Autocorrelation, and Variance Comparison (20 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

Given. Five paired samples of one-way route time (minutes), actual vs. simulated, same 5 test runs:

Sample12345
Actual46.754.940.453.651.5
Simulation45.650.540.140.841.5

Find. (a) a hypothesis test on model accuracy; (b) verification/validation distinction and credibility; (c) autocorrelation definition, test, and a run-length recipe; (d) a variance comparison and its implications.

Approach. (a) since the same 5 runs are measured by both the real system and the model, use a paired $t$-test on the differences; (c) compute lag-1 autocorrelation on the simulated series; (d) compare sample variances via an $F$-test.

(a) Paired test of model accuracy, $\alpha=0.1$

  1. Paired differences (Actual $-$ Simulation). $d_i$: $1.1,\ 4.4,\ 0.3,\ 12.8,\ 10.0$. $$\bar d = 5.72\ \text{min},\qquad s_d = 5.50\ \text{min}.$$
  2. Test statistic. $H_0:\ \mu_{\text{actual}}=\mu_{\text{sim}}$ (model is unbiased): $$t = \frac{\bar d}{s_d/\sqrt5} = \frac{5.72}{5.50/\sqrt5} = \boxed{2.33}.$$
  3. Critical value and decision. $df=4$, two-sided $\alpha=0.1$ ($\alpha/2=0.05$): $t_{0.05,4}=2.132$. Since $|t|=2.33>2.132$, reject $H_0$ — the model's route times are significantly biased low relative to the real system (simulated times average 5.7 min faster than actual).

Assumption: the 5 differences are treated as an i.i.d. sample from an approximately normal population, which is standard for a paired $t$-test on a small sample with no evidence of gross non-normality.

(b) Verification vs. validation; model credibility

Verification is "building the model right": confirming the computer program correctly implements the intended conceptual/logical model (debugging, structured code walkthroughs, tracing against hand-computed test cases) — a purely internal correctness check that never touches real-world data. Validation is "building the right model": confirming the model, taken as a whole, is an accurate representation of the real system for its intended purpose — checked by comparing model output against real-world data (as in part a) or through expert/face judgment. The test in (a) compares simulated route times against data collected from the actual transport company, so it is squarely a validation test, not verification. Model credibility is earned incrementally, not by any single test: a documented verification pass, one or more validation tests like (a) against independent real data, sensitivity analysis showing the model reacts reasonably to input changes, animation/face review by people who know the real system (here, the camp's own transport staff), and transparent disclosure of any known discrepancies (such as the bias just found) all build a decision-maker's confidence that the model's conclusions can be trusted.

(c) Autocorrelation

Definition and cause. Autocorrelation is the correlation between observations of the same output series separated by a time lag $k$: $\rho_k=\mathrm{Corr}(X_t,X_{t+k})$. It tends to exist in simulation output because most simulated systems carry state forward — the number of customers, resources busy, or queue length at time $t+1$ depends directly on the state at time $t$ — so successive observations are rarely independent, unlike a textbook i.i.d. sample.

  1. Lag-1 autocorrelation of the 5 simulated route times. $$r_1 = \frac{\sum_{i=1}^{4}(x_i-\bar x)(x_{i+1}-\bar x)}{\sum_{i=1}^{5}(x_i-\bar x)^2} = \boxed{0.069}.$$ Approximate 95% bound for "no significant autocorrelation" with $n=5$: $\pm1.96/\sqrt5 = \pm0.88$. Since $|0.069| \ll 0.88$, this small sample shows no significant lag-1 autocorrelation — though with only 5 points the test has very little power to detect a real effect either way.

Algorithmic recipe for setting run length against autocorrelation. (1) Run a long pilot and compute the sample autocorrelation function at several lags on the raw output stream (not just 5 replication means). (2) If any early lag exceeds the approximate $\pm1.96/\sqrt n$ bound, the observations are not independent enough for standard CI formulas. (3) Group the output into batches and recompute autocorrelation on the batch means; if still significant, double the batch length (which halves the number of batches) and repeat. (4) Continue until the batch-mean autocorrelation is no longer significant, then use those batches (or independent replications) for the confidence interval. This is exactly the check needed before trusting the batch-means CI recommended in Question 4c.

(d) Variance comparison

  1. Sample variances. $$s^2_{\text{actual}} = 35.15,\qquad s^2_{\text{sim}} = 19.02\ (\text{min}^2).$$
  2. $F$-test, $H_0:\sigma^2_{\text{actual}}=\sigma^2_{\text{sim}}$. $$F = \frac{35.15}{19.02} = \boxed{1.85},\qquad F_{0.025,4,4} = 9.60.$$ Since $1.85 < 9.60$, fail to reject $H_0$ — on this small sample there is no statistically significant difference in variability between the real system and the model.

Yes, variability deserves attention alongside the mean, even though this particular test found no significant gap. If the model's variance were lower than the real system's, the simulation would understate how often wait/route times spike above the 45-minute planning standard — giving false confidence in a bus-count decision that is actually too thin for the real system's tail behaviour. If the model's variance were higher, decision-relevant differences between scenarios (as compared in Question 6) could be masked by inflated noise, risking a Type II error (concluding "no difference" when a real one exists). For a project whose entire purpose is comparing routing/bus-count alternatives against a hard time standard, matching variability — not just the mean — is a genuine, live concern, even though today's 5-sample test does not (yet) show a problem.

QuantityResult
(a) Paired $t$ (bias test)$t=2.33 > t_{0.05,4}=2.132$ — reject $H_0$: model biased low by 5.7 min on average
(b) Test in (a) isValidation (compares model output to real-world data)
(c) Lag-1 autocorrelation$r_1=0.069$, within $\pm0.88$ bound — not significant (small-$n$ caveat)
(d) $F$-test on variances$F=1.85 < F_{crit}=9.60$ — fail to reject; variability is a live concern regardless
Check: part (d) uses $\alpha=0.05$ (two-tailed) for the $F$-test, since the question does not restate the $\alpha=0.1$ used in part (a) — the standard default when a sub-part's significance level is left unstated.