23-Ind-B1 Reliability and Maintainability · December 2019
Question 7 of 9: One-Way ANOVA, Bartlett's Test, and Tukey's HSD
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Notes on this paper
National Exams — December 2019 — 17-Ind-B1 Applied Probability & Statistics. Three-hour, closed-book exam; one of two permitted calculators (Sharp or Casio), one 8.5″×11.0″ aid sheet (both sides), statistical tables supplied. Format: three sections — Section A: do 2 of 4 (20 marks); Section B: do 2 of 3 (40 marks); Section C: do 1 of 2 (20 marks) — a 9-question, 80-mark paper as actually structured by the section table and each question's own printed marks (the front page's own summary line states “Exam: 5 Questions. Total marks: 100”, which is internally inconsistent with 20+40+20=80 from its own section breakdown and every question's printed mark value). All nine questions across the three sections are solved below for completeness.
Reference texts: Montgomery & Runger, Applied Statistics and Probability for Engineers (7th ed., Wiley) — probability density functions and moments (ch. 4), point/interval estimation (ch. 8), two-sample hypothesis testing and the sign test (ch. 9–10, 16), simple linear regression (ch. 11), single-factor ANOVA and multiple comparisons (ch. 13). Montgomery, Peck & Vining, Introduction to Linear Regression Analysis (6th ed., Wiley) — multiple regression by matrices, confidence/prediction intervals (ch. 2–3). Montgomery, Design and Analysis of Experiments (9th ed., Wiley) — Bartlett's test, Tukey's HSD, $2^k$ factorial designs (ch. 3, 6).
Given. Packing cycle times (s), 3 cycles at each of 4 workstations:
WS 1
25.6
24.3
27.9
WS 2
25.2
28.6
24.7
WS 3
20.8
26.7
22.2
WS 4
31.6
29.8
34.3
Find. (a) One-way ANOVA across the 4 workstations; (b) Bartlett's test for equal variances, and reconcile with (a); (c) Tukey's HSD to find which pair(s) differ; (d) the number of pairwise tests and the per-comparison confidence level needed for a family-wise $\alpha=0.05$.
Approach. Recompute all sums of squares directly from the 12 raw times, run the standard one-way ANOVA $F$-test, then Bartlett's test on the same four groups, then Tukey's HSD, then the Bonferroni pairwise-count/confidence-level calculation for (d).
Group and grand statistics (from the raw table).
$$\bar x_1=25.933,\ \bar x_2=26.167,\ \bar x_3=23.233,\ \bar x_4=31.900,\qquad \bar{\bar x}=26.808$$
$$s_1^2=3.323,\ s_2^2=4.503,\ s_3^2=9.503,\ s_4^2=5.130$$
(a) One-way ANOVA, $\alpha=0.05$.
$$SS_{Tr}=3\sum(\bar x_i-\bar{\bar x})^2=119.65,\quad df_{Tr}=3,\qquad SSE=\sum\sum(x_{ij}-\bar x_i)^2=44.92,\quad df_E=8$$
$$MS_{Tr}=\frac{119.65}{3}=39.88,\qquad MSE=\frac{44.92}{8}=5.615$$
$$F=\frac{MS_{Tr}}{MSE}=\boxed{7.10}$$
Critical value $F_{0.05,3,8}=4.066$. Since $F=7.10\gt4.066$ ($p=0.012$), reject $H_0$: at least one workstation's mean cycle time differs from the others.
(b) Bartlett's test for equal variances, $\alpha=0.05$. Using the pooled variance $S_p^2=MSE=5.615$ and each group's own variance:
$$\chi^2=\frac{(N-k)\ln S_p^2-\sum(n_i-1)\ln s_i^2}{C},\quad C=1+\frac{\sum\frac{1}{n_i-1}-\frac1{N-k}}{3(k-1)}$$
Evaluating: $\chi^2=\boxed{0.512}$. Critical value $\chi^2_{0.05,3}=7.815$. Since $0.512\ll7.815$ ($p=0.916$), fail to reject $H_0$: the four workstations' cycle-time variances are homogeneous. Comment on (a): this is a reassuring result, not a contradiction — Bartlett's test confirms the equal-variance assumption that the one-way ANOVA in (a) itself depends on is well supported, so the significant mean difference found in (a) is a genuine effect and not an artifact of unequal within-group spread. The two tests answer independent questions (means vs. variances) and both are consistent here.
(c) Tukey's HSD, $\alpha=0.05$.
$$q_{0.05,4,8}=4.529,\qquad HSD=q\sqrt{\frac{MSE}{n}}=4.529\sqrt{\frac{5.615}{3}}=\boxed{6.196}$$
Pairwise absolute mean differences vs. $HSD$:
Pair
$|\bar x_i-\bar x_j|$
vs. HSD $=6.196$
WS1–WS2
0.233
ns
WS1–WS3
2.700
ns
WS1–WS4
5.967
ns
WS2–WS3
2.933
ns
WS2–WS4
5.733
ns
WS3–WS4
8.667
significant
Only WS3 vs. WS4 differ significantly — WS4's mean cycle time ($31.9$s) is significantly slower than WS3's ($23.2$s); no other pair is distinguishable at the 5% family-wise level.
(d) Full pairwise comparison count and per-comparison confidence level. The number of distinct pairs among 4 workstations is
$$\binom{4}{2}=\boxed{6\text{ tests}}$$
To hold the overall (family-wise) error rate at $\alpha=0.05$ across all 6 independent comparisons, the Bonferroni correction allocates the error budget evenly:
$$\alpha_{\text{each}}=\frac{0.05}{6}=0.00833\quad\Rightarrow\quad\text{each comparison's confidence level}=1-0.00833=\boxed{99.17\%}$$
This is markedly more conservative than testing each pair at a plain 95% level, and it is the reason Tukey's HSD (which controls the family-wise rate directly via the studentized range distribution, part (c)) is generally preferred over 6 separate Bonferroni-corrected $t$-tests — both approaches are valid, but Tukey's is less conservative for this specific “all pairwise comparisons” structure.
Final Results
Quantity
Value
(a) ANOVA $F$
7.10 vs. crit. 4.066 — reject (means differ)
(b) Bartlett $\chi^2$
0.512 vs. crit. 7.815 — fail to reject (equal variances)
(c) Tukey HSD
6.196 — only WS3 vs. WS4 significant
(d) Pairwise tests / per-test CI
6 tests / 99.17% each (Bonferroni)
Check The question states “the variance for entire sample is 9.51,” but recomputing directly from the 12 printed cycle times gives a total sample variance of $SST/(N-1)=164.57/11=14.96$, not 9.51 — this printed shortcut value does not reconcile with the paper's own raw data table. All figures above are computed directly from the raw times, matching the group-by-group variances shown in step 1, and are used in place of the unreconciled 9.51 figure.