23-Ind-B1 Reliability and Maintainability · December 2014
Question 5 of 9: Simple Linear Regression — Marathon Time vs. Race Number
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Notes on this paper
National Exams — December 2014 — 98-Ind-B1 Applied Probability & Statistics. Three-hour, closed-book exam; one of two permitted calculators (Sharp or Casio), one 8.5″×11.0″ aid sheet (both sides), statistical tables supplied. Format: four sections — Section A: do 2 of 3 questions (20 marks); Section B: do 1 of 2 (25 marks); Section C: do 1 of 2 (25 marks); Section D: do 1 of 2 (30 marks) — a 5-question, 100-mark paper as printed. All nine questions across the four sections are solved below for completeness. Page 1's own NOTES list is mis-numbered (two items both labelled "4."), transcribed as printed.
Reference texts: Montgomery & Runger, Applied Statistics and Probability for Engineers (7th ed., Wiley) — point/interval estimation (ch. 8–9), hypothesis testing (ch. 9–10), simple/multiple linear regression (ch. 11–12), single- and two-factor ANOVA (ch. 13). Montgomery, Peck & Vining, Introduction to Linear Regression Analysis (6th ed., Wiley) — multiple regression by matrices, confidence/prediction intervals (ch. 2–3).
Question 5 (Section B.2): Simple Linear Regression — Marathon Time vs. Race Number (25 marks)
Given. Six races, with completion time (minutes) versus race number:
Race $x$
1
2
3
4
5
6
Time $y$
258
246
236
240
216
212
Find. (a) $\hat y=a+bx$; (b) is the slope significant at $\alpha=0.05$; (c) 95% CI for the mean time at race 7; (d) 95% PI for the runner's own race 7; (e) goodness of fit.
Approach. Fit by least squares using the corrected sums $S_{xx}$, $S_{xy}$, $S_{yy}$; partition $S_{yy}=SSR+SSE$ for the ANOVA $F$-test (equivalent here to the slope $t$-test); then build the CI for the mean response and the wider PI for a single new observation at $x_0=7$.
Part (a) — least-squares fit. With $\bar x=3.5$, $\bar y=234.67$: $S_{xx}=\sum(x_i-\bar x)^2=17.5$, $S_{xy}=\sum(x_i-\bar x)(y_i-\bar y)=-158.0$, so
$$b=\frac{S_{xy}}{S_{xx}}=\frac{-158.0}{17.5}=-9.029,\qquad a=\bar y-b\bar x = 234.67-(-9.029)(3.5)=\boxed{266.27},$$
giving $\hat y = 266.27 - 9.029\,x$ (time drops about 9 minutes per additional race — the training-effect the runner is looking for).
Part (b) — ANOVA $F$-test on the slope. $S_{yy}=\sum(y_i-\bar y)^2=1565.33$; residuals from the fitted line give $SSE=138.82$, so $SSR=S_{yy}-SSE=1426.51$. With $df_R=1$, $df_E=n-2=4$:
$$MSE=\frac{SSE}{4}=34.70,\qquad F_0=\frac{SSR/1}{MSE}=\boxed{41.10}.$$
Against $F_{0.05,1,4}=7.71$, $F_0=41.10>7.71$: reject $H_0:b=0$ — there is a statistically significant linear relationship between race number and time. $R^2=SSR/S_{yy}=0.911$.
Part (c) — 95% CI for the mean at $x_0=7$. The fitted value is $\hat y(7)=266.27-9.029(7)=203.07$ minutes. The standard error of the mean response is
$$se_{\text{mean}}=\sqrt{MSE\left(\frac1n+\frac{(x_0-\bar x)^2}{S_{xx}}\right)}=\sqrt{34.70\left(\frac16+\frac{(7-3.5)^2}{17.5}\right)}=8.20,$$
so with $t_{0.025,4}=2.776$: $203.07 \pm 2.776(8.20) = \boxed{(187.8,\ 218.3)\text{ min}}$.
Part (d) — 95% PI for the runner's own next race. A single future observation carries the additional variance of that one race's own random deviation from the line, so
$$se_{\text{pred}}=\sqrt{MSE\left(1+\frac1n+\frac{(x_0-\bar x)^2}{S_{xx}}\right)}=\sqrt{34.70\left(1+\frac16+\frac{12.25}{17.5}\right)}=8.53,$$
giving $203.07 \pm 2.776(8.53) = \boxed{(180.7,\ 225.4)\text{ min}}$. The prediction interval is wider than the confidence interval in (c) because it must cover the extra, irreducible race-to-race scatter around the regression line, not just the uncertainty in where the line itself sits — (c) answers "where is the true average race-7 time," while (d) answers "how fast will this specific race-7 actually be."
Part (e) — is the linear model a good fit? $R^2=0.911$ and a highly significant slope ($F_0=41.1\gg F_{crit}$) both support a strong downward linear trend. The one visible departure is race 4 ($y=240$), which rises above races 3 and 5 rather than continuing the decline — a single interior deviation in an $n=6$ sample, plausibly an off day rather than a break in the trend. Overall, a linear model is a reasonable and useful fit given the small sample, though six points give limited power to detect any genuine curvature (e.g. diminishing returns to training).