NivaarExam PrepOfficial exam papers ↗

23-Ind-B1 Reliability and Maintainability · December 2016

Question 9 of 11: Running-Shoe Comparison — One-Way ANOVA and Tukey's HSD

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

National Exams — December 2016 — 98-Ind-B1 Applied Probability & Statistics. Three-hour, closed-book exam; one of two permitted calculators (Sharp or Casio), one 8.5″×11.0″ aid sheet (both sides), statistical tables supplied. Format: three sections — Section A: do 2 of 4 (30 marks); Section B: do 2 of 3 (30 marks); Section C: do 2 of 4 (40 marks) — a 6-question, 100-mark paper as printed. All eleven questions across the three sections are solved below for completeness.

Reference texts: Montgomery & Runger, Applied Statistics and Probability for Engineers (7th ed., Wiley) — joint distributions and covariance (ch. 5), point/interval estimation (ch. 8), hypothesis testing incl. two-sample and goodness-of-fit tests (ch. 9–10), simple linear regression (ch. 11), design and analysis of single-factor and factorial experiments (ch. 13–14). Montgomery, Peck & Vining, Introduction to Linear Regression Analysis (6th ed., Wiley) — multiple regression by matrices, confidence/prediction intervals (ch. 2–3). Montgomery, Design and Analysis of Experiments (9th ed., Wiley) — two-way factorial ANOVA and $2^k$ designs (ch. 5, 6–7).

Question 9 (Section C.2): Running-Shoe Comparison — One-Way ANOVA and Tukey's HSD (20 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

Given. After the fix in (a), five balanced samples of $n=5$ (weekly-mean 10 km times, min):

Shoe A49.351.251.648.649.5
Shoe B50.051.854.153.651.7
Shoe C46.847.249.950.347.8
Shoe D51.251.345.153.648.2
Shoe E54.754.348.551.549.1

Find. (a) essay: a design fix for unequal-$n$, non-Normal daily data; (b) one-way ANOVA $F$-test; (c)/(d) Tukey pairwise comparisons and the best shoe (if any).

Approach. (a) is answered qualitatively via the Central Limit Theorem; (b)–(d) apply a standard balanced one-way ANOVA followed by Tukey's HSD.

(a) Fixing unequal $n$ and non-Normality (no calculation). Rather than testing raw daily times directly, aggregate each week's 5–6 daily runs into a single weekly mean time per shoe. Two things happen at once: by the Central Limit Theorem, the average of several non-Normal daily times is itself approximately Normal even though the individual days are not, which satisfies the Normality assumption of a standard ANOVA/$t$-test; and using "one weekly mean" as the unit of analysis (rather than "one day") makes each week contribute exactly one observation regardless of whether it held 5 or 6 runs, which is a natural, defensible way to equalize the unit of replication across shoes with different total daily counts. The remaining between-shoe imbalance (30–60 weeks each) is handled by the fact that one-way ANOVA does not require equal group sizes; alternatively, a distribution-free Kruskal–Wallis test on the raw daily ranks would sidestep the Normality question entirely and also tolerates unequal $n$ natively — either the CLT-aggregation route or the nonparametric route is a defensible answer here.

  1. (b) One-way ANOVA. Grand mean $\bar x=50.436$; shoe means: A$=50.04$, B$=52.24$, C$=48.40$, D$=49.88$, E$=51.62$. With $k=5$, $n=5$ each: $$SS_{Between}=n\sum(\bar x_i-\bar x)^2=46.338,\qquad SS_{Within}=\sum\sum(x_{ij}-\bar x_i)^2=103.760$$ $df_B=4$, $df_W=20$; $MS_B=11.584$, $MS_W=5.188$. $$F=\frac{MS_B}{MS_W}=\boxed{2.233}$$ $F_{0.05,4,20}=2.866$; since $2.233\lt2.866$, fail to reject $H_0$: at $\alpha=0.05$, shoe choice does not significantly affect 10 km training time.
  2. (c)/(d) Tukey's HSD. Even though the overall $F$-test is not significant, Tukey confirms it directly: $q_{0.05,5,20}=4.232$, $$HSD=q\sqrt{MS_W/n}=4.232\sqrt{5.188/5}=\boxed{4.311}$$ Every pairwise mean difference (largest is $|\bar x_B-\bar x_C|=3.84$) is below $4.311$, so no pair of shoes differs significantly, and consequently no single shoe is significantly "best" — Shoe C has the lowest sample mean ($48.40$ min) but this is not statistically distinguishable from the other four given this sample size.
Summary
QuantityValue
$F$ (one-way ANOVA)2.233 vs. crit. 2.866 — n.s.
Tukey HSD4.311 (largest gap 3.84)
Significant pairsNone
Best shoeNone significantly better (C lowest numerically)