NivaarExam PrepOfficial exam papers ↗

23-Ind-B1 Reliability and Maintainability · December 2014

Question 5 of 9: Simple Linear Regression — Marathon Time vs. Race Number

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

National Exams — December 2014 — 98-Ind-B1 Applied Probability & Statistics. Three-hour, closed-book exam; one of two permitted calculators (Sharp or Casio), one 8.5″×11.0″ aid sheet (both sides), statistical tables supplied. Format: four sections — Section A: do 2 of 3 questions (20 marks); Section B: do 1 of 2 (25 marks); Section C: do 1 of 2 (25 marks); Section D: do 1 of 2 (30 marks) — a 5-question, 100-mark paper as printed. All nine questions across the four sections are solved below for completeness. Page 1's own NOTES list is mis-numbered (two items both labelled "4."), transcribed as printed.

Reference texts: Montgomery & Runger, Applied Statistics and Probability for Engineers (7th ed., Wiley) — point/interval estimation (ch. 8–9), hypothesis testing (ch. 9–10), simple/multiple linear regression (ch. 11–12), single- and two-factor ANOVA (ch. 13). Montgomery, Peck & Vining, Introduction to Linear Regression Analysis (6th ed., Wiley) — multiple regression by matrices, confidence/prediction intervals (ch. 2–3).

Question 5 (Section B.2): Simple Linear Regression — Marathon Time vs. Race Number (25 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

Given. Six races, with completion time (minutes) versus race number:

Race $x$123456
Time $y$258246236240216212

Find. (a) $\hat y=a+bx$; (b) is the slope significant at $\alpha=0.05$; (c) 95% CI for the mean time at race 7; (d) 95% PI for the runner's own race 7; (e) goodness of fit.

Approach. Fit by least squares using the corrected sums $S_{xx}$, $S_{xy}$, $S_{yy}$; partition $S_{yy}=SSR+SSE$ for the ANOVA $F$-test (equivalent here to the slope $t$-test); then build the CI for the mean response and the wider PI for a single new observation at $x_0=7$.

  1. Part (a) — least-squares fit. With $\bar x=3.5$, $\bar y=234.67$: $S_{xx}=\sum(x_i-\bar x)^2=17.5$, $S_{xy}=\sum(x_i-\bar x)(y_i-\bar y)=-158.0$, so $$b=\frac{S_{xy}}{S_{xx}}=\frac{-158.0}{17.5}=-9.029,\qquad a=\bar y-b\bar x = 234.67-(-9.029)(3.5)=\boxed{266.27},$$ giving $\hat y = 266.27 - 9.029\,x$ (time drops about 9 minutes per additional race — the training-effect the runner is looking for).
  2. Part (b) — ANOVA $F$-test on the slope. $S_{yy}=\sum(y_i-\bar y)^2=1565.33$; residuals from the fitted line give $SSE=138.82$, so $SSR=S_{yy}-SSE=1426.51$. With $df_R=1$, $df_E=n-2=4$: $$MSE=\frac{SSE}{4}=34.70,\qquad F_0=\frac{SSR/1}{MSE}=\boxed{41.10}.$$ Against $F_{0.05,1,4}=7.71$, $F_0=41.10>7.71$: reject $H_0:b=0$ — there is a statistically significant linear relationship between race number and time. $R^2=SSR/S_{yy}=0.911$.
  3. Part (c) — 95% CI for the mean at $x_0=7$. The fitted value is $\hat y(7)=266.27-9.029(7)=203.07$ minutes. The standard error of the mean response is $$se_{\text{mean}}=\sqrt{MSE\left(\frac1n+\frac{(x_0-\bar x)^2}{S_{xx}}\right)}=\sqrt{34.70\left(\frac16+\frac{(7-3.5)^2}{17.5}\right)}=8.20,$$ so with $t_{0.025,4}=2.776$: $203.07 \pm 2.776(8.20) = \boxed{(187.8,\ 218.3)\text{ min}}$.
  4. Part (d) — 95% PI for the runner's own next race. A single future observation carries the additional variance of that one race's own random deviation from the line, so $$se_{\text{pred}}=\sqrt{MSE\left(1+\frac1n+\frac{(x_0-\bar x)^2}{S_{xx}}\right)}=\sqrt{34.70\left(1+\frac16+\frac{12.25}{17.5}\right)}=8.53,$$ giving $203.07 \pm 2.776(8.53) = \boxed{(180.7,\ 225.4)\text{ min}}$. The prediction interval is wider than the confidence interval in (c) because it must cover the extra, irreducible race-to-race scatter around the regression line, not just the uncertainty in where the line itself sits — (c) answers "where is the true average race-7 time," while (d) answers "how fast will this specific race-7 actually be."
  5. Part (e) — is the linear model a good fit? $R^2=0.911$ and a highly significant slope ($F_0=41.1\gg F_{crit}$) both support a strong downward linear trend. The one visible departure is race 4 ($y=240$), which rises above races 3 and 5 rather than continuing the decline — a single interior deviation in an $n=6$ sample, plausibly an off day rather than a break in the trend. Overall, a linear model is a reasonable and useful fit given the small sample, though six points give limited power to detect any genuine curvature (e.g. diminishing returns to training).
QuantityResult
Fitted model$\hat y=266.27-9.029x$
$F_0$ vs $F_{0.05,1,4}$41.10 > 7.71 — significant, $R^2=0.911$
$\hat y(7)$203.07 min
95% CI, mean at $x=7$(187.8, 218.3) min
95% PI, next race at $x=7$(180.7, 225.4) min