NivaarExam PrepOfficial exam papers ↗

23-Ind-B1 Reliability and Maintainability · December 2018

Question 3 of 9: Comparing Two Shipping Companies — Two-Sample $t$-Test

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

National Exams — December 2018 — 17-Ind-B1 Applied Probability & Statistics. Three-hour, closed-book exam; one of two permitted calculators (Sharp or Casio), one 8.5″×11.0″ aid sheet (both sides), statistical tables supplied. Format: three sections — Section A: do 2 of 4 (20 marks); Section B: do 2 of 3 (40 marks); Section C: do 1 of 2 (20 marks) — a 9-question, printed-total-80-mark paper as actually structured (the front page's own summary line states “Exam: 5 Questions. Total marks: 100”, but 20+40+20=80 by the section instructions and per-question mark values printed beside each question — the front page's 100 is internally inconsistent with its own section table; the 80-mark reading is used throughout, consistent with the section-by-section split printed on pages 2, 4 and 6). All nine questions across the three sections are solved below for completeness.

Reference texts: Montgomery & Runger, Applied Statistics and Probability for Engineers (7th ed., Wiley) — normal distribution and sums of normals (ch. 4–5), point/interval estimation (ch. 8), two-sample hypothesis testing (ch. 9–10), simple linear regression (ch. 11), single-factor ANOVA and multiple comparisons (ch. 13). Montgomery, Peck & Vining, Introduction to Linear Regression Analysis (6th ed., Wiley) — multiple regression by matrices, confidence/prediction intervals (ch. 2–3). Montgomery, Design and Analysis of Experiments (9th ed., Wiley) — Bartlett's test, $2^k$ factorial designs (ch. 3, 6).

Question 3 (Section A.3): Comparing Two Shipping Companies — Two-Sample $t$-Test (10 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

Given. Shipping times (days), 5 independent shipments per company.

Company A3530332732
Company B3826373232

Find. (a) Is one company reliably faster? (b) statistical issues to be aware of; (c) a proper data-collection/analysis plan.

Approach. First check whether the two population variances can be assumed equal (an $F$-test), then run the corresponding two-sample $t$-test on the means.

  1. Sample statistics. $$\bar{x}_A=31.4\text{ d},\ s_A=3.050\qquad \bar{x}_B=33.0\text{ d},\ s_B=4.796$$
  2. $F$-test for equal variances. $$F=\frac{s_B^2}{s_A^2}=\frac{23.0}{9.30}=2.473$$ Critical value $F_{0.025,4,4}=9.605$. Since $F=2.473<9.605$, fail to reject equal variances — use the pooled two-sample $t$-test.
  3. Pooled two-sample $t$-test, $H_0:\mu_A=\mu_B$ vs. $H_1:\mu_A\ne\mu_B$, $\alpha=0.05$. $$s_p^2=\frac{4(9.30)+4(23.0)}{8}=16.15,\qquad SE=s_p\sqrt{0.4}=2.541$$ $$t=\frac{31.4-33.0}{2.541}=\boxed{-0.630},\qquad df=8,\quad t_{0.025,8}=\pm2.306$$ Since $|t|=0.630\ll2.306$ ($p=0.547$), fail to reject $H_0$: with only 5 shipments per company, there is no statistically significant evidence that one shipping firm is faster than the other.
  4. (b) Statistical issues (essay). Several elements limit how much weight this comparison can bear. Sample size is very small ($n=5$ each), so the test has low power to detect anything but a large true difference; the $F$-test for equal variances is itself weakly powered at $n=5$, so “fail to reject equal variances” is a weak endorsement, not proof. The five shipments per company must be independent random draws for the $t$-test to be valid — if shipments cluster in time (e.g. all from the same month, same product line, same destination port) they are not truly independent and the effective sample size is smaller than 5. Both companies should be compared over the same conditions (season, route, cargo type, distance) or route/season becomes a confound indistinguishable from a genuine company effect. Finally, normality of shipping times is assumed but unverified at $n=5$; a skewed distribution (e.g. occasional customs delays create a long right tail) would make the $t$-test's normality assumption shaky.
  5. (c) Improved data-collection and analysis plan (essay). Collect a substantially larger sample from each company (tens of shipments, not five), drawn as a genuinely random selection across a full year to average out seasonal effects, and matched by route/destination/cargo type between the two companies so route is not confounded with carrier. Where possible, pair shipments that travel the same route in the same period (a matched-pairs design), which removes route/season variability directly and increases the test's power without a larger sample. Before testing, check the normality assumption (e.g. a normal probability plot) and re-examine the equal-variance assumption on the larger sample; if either assumption is doubtful, use a non-parametric alternative (Mann–Whitney $U$) or Welch's unequal-variance $t$-test as a robustness check against the parametric result. Pre-specify the hypothesis, $\alpha$, and the practically meaningful effect size before collecting data, rather than deciding these after seeing the numbers.
Final Results
QuantityValue
$F$ (variance-equality test)2.473 (vs. crit. 9.605 — equal variances)
Pooled $t$$-0.630$ (vs. crit. $\pm2.306$)
ConclusionNo significant difference between companies ($p=0.547$)