NivaarExam PrepOfficial exam papers ↗

23-Ind-A6 Systems Simulation · December 2017

Question 4 of 8: Comparing Two Alternatives: Paired vs. Independent $t$-Test

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

National Exams — December 2017 — 98-Ind-A6 Systems Simulation. Three-hour, closed-book exam; one of two permitted calculators (Sharp or Casio), two 8.5″×11.0″ aid sheets (both sides). Format: eight questions of equal value (10 marks each); candidates complete SIX of the EIGHT (only the first six as they appear in the answer book are marked). All eight questions are solved below for completeness (the paper's own Q6 is printed with an orphan leading "6." followed by "2. Consider an M/M/1 system…" and is answered as the paper's sixth question). Common Discrete/Continuous Distribution tables, Student-t and Chi-square tables were supplied with the exam; the values below are the same table values obtained by direct computation.

Reference texts: Banks, Carson, Nelson & Nicol, Discrete-Event System Simulation (5th ed., Pearson) — random-number/random-variate generation, input data analysis, output analysis (replications), comparing alternative systems, queueing simulation, and the simulation study life cycle.

Question 4 — Comparing Two Alternatives: Paired vs. Independent $t$-Test (10 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

Given. Route-completion times (minutes), 10 replications each:

ReplicationOriginalProposed
1177175
2186172
3188172
4172186
5189188
6184170
7173166
8180186
9172177
10182170

Find. (a) whether a paired test is the right tool here; (b) whether the proposed routing significantly reduces route time at 90% confidence, treating the samples as independent with equal variance (as instructed).

(a) Should a paired $t$-test be used? It depends entirely on whether the two systems' 10 replications share Common Random Numbers (CRN). If replication $i$ of the Original run and replication $i$ of the Proposed run were driven by the same underlying random-number streams (the same simulated traffic/weather/timing draws for "day $i$"), the two samples are correlated pair-by-pair, and a paired $t$-test on the 10 differences is not only valid but preferable — it cancels the shared random variation and gives a narrower CI / more powerful test for the same 10 replications than an independent-samples analysis. If independent streams were used for each system (no CRN), there is no genuine correspondence between "replication 3 Original" and "replication 3 Proposed", and pairing would be invalid; an independent two-sample test is then the correct (and only defensible) choice. This paper does not state that CRN was used, so part (b) below follows the question's own explicit instruction and treats the two samples as independent.

Approach (b). Pooled-variance two-sample $t$-test, $H_0:\mu_{Orig}=\mu_{Prop}$, $\alpha=0.10$.

  1. Sample means and pooled variance. $$\bar x_{Orig}=180.30,\ \ \bar x_{Prop}=176.20\ \text{min};\qquad s_p^2 = \frac{(n_1-1)s_1^2+(n_2-1)s_2^2}{n_1+n_2-2} = 51.98.$$
  2. Test statistic. $$t = \frac{\bar x_{Orig}-\bar x_{Prop}}{\sqrt{s_p^2(1/n_1+1/n_2)}} = \frac{180.30-176.20}{\sqrt{51.98(1/10+1/10)}} = \boxed{1.272},\qquad df=18.$$
  3. Critical value and decision. Two-sided, $\alpha=0.10$: $t_{0.05,18}=1.734$ (one-sided $t_{0.10,18}=1.330$ gives the same call). Since $t=1.272$ is below BOTH the two-sided and the (more favourable) one-sided critical value, fail to reject $H_0$ — the 4.1-minute sample-mean improvement is not statistically significant at 90% confidence.
QuantityResult
Mean Original / Proposed180.30 / 176.20 min
Test statistic $t$ (pooled, $df=18$)1.272
Critical value (two-sided, $\alpha=0.10$)1.734
DecisionNot significant — insufficient evidence to adopt on this data alone
Check: the sample means DO favour the proposed routing (176.2 vs. 180.3 min) but the 10-replication sample is too small, relative to the observed variability, to call the difference statistically real at 90% confidence; more replications (or a CRN-paired design per part (a)) would sharpen this comparison before a final adopt/reject decision.