NivaarExam PrepOfficial exam papers ↗

23-Ind-B1 Reliability and Maintainability · December 2019

Question 3 of 9: Simulation Model Validation

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

National Exams — December 2019 — 17-Ind-B1 Applied Probability & Statistics. Three-hour, closed-book exam; one of two permitted calculators (Sharp or Casio), one 8.5″×11.0″ aid sheet (both sides), statistical tables supplied. Format: three sections — Section A: do 2 of 4 (20 marks); Section B: do 2 of 3 (40 marks); Section C: do 1 of 2 (20 marks) — a 9-question, 80-mark paper as actually structured by the section table and each question's own printed marks (the front page's own summary line states “Exam: 5 Questions. Total marks: 100”, which is internally inconsistent with 20+40+20=80 from its own section breakdown and every question's printed mark value). All nine questions across the three sections are solved below for completeness.

Reference texts: Montgomery & Runger, Applied Statistics and Probability for Engineers (7th ed., Wiley) — probability density functions and moments (ch. 4), point/interval estimation (ch. 8), two-sample hypothesis testing and the sign test (ch. 9–10, 16), simple linear regression (ch. 11), single-factor ANOVA and multiple comparisons (ch. 13). Montgomery, Peck & Vining, Introduction to Linear Regression Analysis (6th ed., Wiley) — multiple regression by matrices, confidence/prediction intervals (ch. 2–3). Montgomery, Design and Analysis of Experiments (9th ed., Wiley) — Bartlett's test, Tukey's HSD, $2^k$ factorial designs (ch. 3, 6).

Question 3 (Section A.3): Simulation Model Validation (10 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

Given. Actual job time in system (5 months) vs. simulated job time (5 independent replications), both in minutes.

Actual10410794110107
Simulation12067115100138

Find. (a) $F$-test for equal variances; (b) test $H_0:\mu_1=\mu_2$ (mean actual vs. mean simulated).

Approach. Run the $F$-test first, since its outcome determines which two-sample $t$-test is valid here — unlike Question 2, this pair of samples turns out to have significantly different variances, which rules out the pooled test.

  1. Sample statistics. $$\bar x_{\text{actual}}=104.4,\ s_{\text{actual}}=6.189\ (s^2=38.30)\qquad \bar x_{\text{sim}}=108.0,\ s_{\text{sim}}=26.636\ (s^2=709.50)$$
  2. (a) $F$-test, $\alpha=0.05$. $$F=\frac{s_{\text{sim}}^2}{s_{\text{actual}}^2}=\frac{709.50}{38.30}=\boxed{18.52}$$ Critical value $F_{0.025,4,4}=9.605$. Since $F=18.52\gt9.605$, reject $H_0:\sigma_1^2=\sigma_2^2$ — unlike Question 2's pair, these two samples' variances are significantly different, so the pooled-variance $t$-test used there is not valid here; a variance-unequal (Welch) $t$-test must be used for part (b) instead.
  3. (b) Welch's (unequal-variance) two-sample $t$-test, $\alpha=0.05$. Assumption: both actual job times and simulated replication times are approximately normally distributed (5 monthly averages / 5 independent simulation replications, each itself an average over many individual jobs — reasonable by the Central Limit Theorem). $$SE=\sqrt{\frac{s_{\text{actual}}^2}{5}+\frac{s_{\text{sim}}^2}{5}}=\sqrt{\frac{38.30}{5}+\frac{709.50}{5}}=12.229$$ $$t=\frac{104.4-108.0}{12.229}=\boxed{-0.294}$$ Welch–Satterthwaite degrees of freedom: $df=\dfrac{(38.30/5+709.50/5)^2}{(38.30/5)^2/4+(709.50/5)^2/4}=4.43$, critical value $t_{0.025,4.43}\approx2.673$. Since $|t|=0.294\ll2.673$ ($p=0.78$), fail to reject $H_0$: no significant evidence that the simulation's mean job time differs from the real system's — the mean is validated at the 5% level. The simulation's much larger variance (709.5 vs. 38.3, confirmed unequal in part (a)) is nonetheless worth a separate note: even though the means match, the simulation is far more erratic replication-to-replication than the real process, which is itself a validation concern the mean-comparison alone does not capture.
Final Results
QuantityValue
(a) $F$ (equal-variance test)18.52 vs. crit. 9.605 — reject (variances unequal)
(b) Welch $t$ ($H_0:\mu_1=\mu_2$)$-0.294$ vs. crit. $\pm2.673$ ($df=4.43$) — fail to reject (means match)