23-Ind-B1 Reliability and Maintainability · May 2015
Question 5 of 10: Marathon Training — Simple Linear Regression, ANOVA, Intervals
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Notes on this paper
National Exams — May 2015 — 98-Ind-B1 Applied Probability & Statistics. Three-hour, closed-book exam; one of two permitted calculators (Sharp or Casio), one 8.5″×11.0″ aid sheet (both sides), statistical tables supplied. Format: three sections — Section A: do 2 of 4 questions (10 marks); Section B: do 2 of 3 (20 marks); Section C: do 1 of 3 (20 marks) — a 5-question, 50-mark paper as printed. All ten questions across the three sections are solved below for completeness.
Reference texts: Montgomery & Runger, Applied Statistics and Probability for Engineers (7th ed., Wiley) — joint distributions (ch. 5), binomial/geometric probability (ch. 3), normal distribution and sampling distributions (ch. 4–7), point/interval estimation (ch. 8–9), hypothesis testing incl. sample-size design (ch. 9–10), simple linear regression and ANOVA (ch. 11–13), Bartlett's and Tukey's tests, single-degree-of-freedom contrasts and 23 factorial designs (ch. 13–14). Montgomery, Peck & Vining, Introduction to Linear Regression Analysis (6th ed., Wiley) — multiple regression by matrices, confidence/prediction intervals (ch. 2–3).
Question 5 (Section B.1): Marathon Training — Simple Linear Regression, ANOVA, Intervals (10 marks)
Given. Six races, completion time (minutes) vs. race number:
Race $x$
1
2
3
4
5
6
Time $y$
240
240
230
225
205
210
Find. (a) $\hat y=a+bx$; (b) ANOVA $F$-test for the slope; (c) 95% PI at $x=8$; (d) $t$-test on $a$; (e) qualitative comparison of prediction uncertainty; (f) 95% CI for the mean of 30 future runners at $x=8$.
Approach. Compute least-squares $\hat a,\hat b$ from $S_{xx},S_{xy}$, run the regression ANOVA $F$-test, then use the standard $t$-based interval formulas at $x_0=8$ for a single new observation, for the slope/intercept tests, and for the mean of $m=30$ future observations.
Fig. Q5 — race time vs. race number, six observed races (blue) and the least-squares fit (red dashed); the requested prediction point $x=8$ lies outside this observed range.
(b) ANOVA $F$-test for the linear relationship. Residual sum: $SSE=\sum(y_i-\hat y_i)^2=134.29$. Using the given $SST=1100$: $SSR=SST-SSE=965.71$. $$F=\frac{SSR/1}{SSE/(n-2)}=\frac{965.71}{134.29/4}=\frac{965.71}{33.57}=\boxed{28.77}$$ $F_{0.05,1,4}=7.709$; since $28.77\gt 7.709$, reject $H_0:\beta=0$ — a significant linear relationship exists between race number and time.
(c) 95% prediction interval at $x_0=8$. $\hat y_0=251.0-7.4286(8)=191.57$. With $s=\sqrt{MSE}=\sqrt{33.57}=5.794$ and $t_{0.025,4}=2.776$: $$se_{pred}=s\sqrt{1+\tfrac1n+\tfrac{(x_0-\bar x)^2}{S_{xx}}}=5.794\sqrt{1+\tfrac16+\tfrac{(4.5)^2}{17.5}}=8.833$$ $$PI=191.57\pm 2.776(8.833)=\boxed{(167.0,\ 216.1)\text{ min}}$$
(d) Is the intercept $a$ significantly different from 0? $se_{\hat a}=s\sqrt{\tfrac1n+\tfrac{\bar x^2}{S_{xx}}}=5.794\sqrt{\tfrac16+\tfrac{12.25}{17.5}}=5.394$. $$t=\frac{251.0-0}{5.394}=\boxed{46.5}$$ Since $46.5\gg t_{0.025,4}=2.776$ ($p\lt 0.001$), the intercept is significantly different from 0.
(e) Race 10 vs. race 8, no new calculation. The prediction-interval width grows with $(x_0-\bar x)^2/S_{xx}$; race 10 is farther from $\bar x=3.5$ than race 8 ($6.5^2=42.25$ vs. $4.5^2=20.25$), so the interval for race 10 would be visibly wider, and it also extrapolates further beyond the observed range (races 1–6) — both push the prediction toward less certainty.
(f) 95% CI for the mean of $m=30$ new runners at $x_0=8$. The standard error for the mean of $m$ future observations replaces the "$1$" in step (c) with "$1/m$": $$se_{mean}=s\sqrt{\tfrac1{30}+\tfrac16+\tfrac{(4.5)^2}{17.5}}=6.750$$ $$CI=191.57\pm 2.776(6.750)=\boxed{(172.8,\ 210.3)\text{ min}}$$ — narrower than the single-runner PI in (c), as expected when averaging over 30 runners.
Summary
Part
Result
(a) fit
$\hat y=251.0-7.4286x$
(b) $F$, verdict
28.77 > 7.709; significant slope
(c) 95% PI, race 8
(167.0, 216.1) min
(d) $t$, verdict
46.5; $a\ne 0$
(f) 95% CI, mean of 30
(172.8, 210.3) min
Check — extrapolation Race 8 (and race 10 in (e)) both lie outside the observed range of races 1–6; the intervals in (c)/(f) assume the linear trend continues, which is an engineering judgment call the exam itself invites by asking for these specific predictions.