NivaarExam PrepOfficial exam papers ↗

23-Ind-B1 Reliability and Maintainability · December 2017

Question 8 of 10: Marathon Time — Multiple Regression via $(X'X)^{-1}$

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

National Exams — December 2017 — 98-Ind-B1 Applied Probability & Statistics. Three-hour, closed-book exam; one of two permitted calculators (Sharp or Casio), one 8.5″×11.0″ aid sheet (both sides), statistical tables supplied. Format: three sections — Section A: do 2 of 4 (30 marks); Section B: do 2 of 3 (30 marks); Section C: do 2 of 3 (40 marks) — a 6-question, 100-mark paper as printed. All ten questions across the three sections are solved below for completeness.

Reference texts: Montgomery & Runger, Applied Statistics and Probability for Engineers (7th ed., Wiley) — discrete/continuous distributions (ch. 3–4), joint distributions (ch. 5), point/interval estimation and sample size (ch. 8), hypothesis testing incl. two-sample tests (ch. 9–10), simple linear regression (ch. 11), design and analysis of single-factor experiments (ch. 13). Montgomery, Peck & Vining, Introduction to Linear Regression Analysis (6th ed., Wiley) — multiple regression by matrices (ch. 2–3). Montgomery, Design and Analysis of Experiments (9th ed., Wiley) — multi-factor and $2^k$ factorial designs (ch. 5–6).

Question 8 (Section C.1): Marathon Time — Multiple Regression via $(X'X)^{-1}$ (20 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

Separately: the exam text states "$SST$ is 0.0548333," but recomputing $SST=\sum(y_i-\bar y)^2$ directly from the printed dataset gives $\boxed{SST=0.548333}$ — exactly $10\times$ the printed figure, a genuine decimal-point typo in the source. Using the printed $0.0548333$ would make $SSE=SST-SSR$ negative, which is impossible — itself proof the printed value is wrong. The correct $SST=0.548333$ is used throughout below.

Given. $n=6$ races:

$x_1$ (km×100)4.64.54.04.14.44.2
$x_2$ (max dist, km)273130252832
$x_3$ (temp, °C)111025221417
$y$ (time, hrs)4.33.83.53.63.53.4

Given partial $(X'X)^{-1}$ and $\beta=(X'X)^{-1}(X'Y)$:

3186.5127*−10.0932−22.9764
−585.1664108.40631.70754.2294
−10.09321.70750.0562**
−22.9764***0.06850.1706

$\beta=(-16.29491,\ 4.1060,\ 0.0093,\ 0.1245)^T$.

Find. (a) $X'X$; (b) $*,**,***$; (c) $SSE,SSR,SST$; (d) overall $F$; (e) coefficient $t$-tests; (f) is the model useful?

Approach. Build $X$ with an intercept column, form $X'X$ directly, exploit that $(X'X)^{-1}$ must be symmetric to fill the blanks (cross-checked by an independent inverse computed from the raw data), then run the regression ANOVA and coefficient $t$-tests off the given $\beta$.

  1. (a) $X'X$ matrix. With $X=[\mathbf 1,\ x_1,\ x_2,\ x_3]$ ($6\times4$): $$X'X=\begin{pmatrix}6 & 25.8 & 173 & 99\\ 25.8 & 111.22 & 743.8 & 418.8\\ 173 & 743.8 & 5023 & 2843\\ 99 & 418.8 & 2843 & 1815\end{pmatrix}$$ (row/column 1 is the intercept; e.g. entry $(1,2)=\sum x_1=25.8$, entry $(2,2)=\sum x_1^2=111.22$.)
  2. (b) Filling $*,**,***$ by symmetry. $(X'X)^{-1}$ is always symmetric, so each starred cell equals its mirror image across the diagonal — confirmed independently by inverting $X'X$ directly from the raw data (matches to 4 decimal places): $$*=(\text{row 1, col 2})=(\text{row 2, col 1})=\boxed{-585.1664}$$ $$**=(\text{row 3, col 4})=(\text{row 4, col 3})=\boxed{0.0685}$$ $$***=(\text{row 4, col 2})=(\text{row 2, col 4})=\boxed{4.2294}$$
  3. (c) SSE, SSR, SST. Fitted values from $\hat y=X\beta$: $(4.213,\ 3.715,\ 3.521,\ 3.511,\ 3.775,\ 3.364)$. Using the corrected $SST=0.548333$ (see the check note above): $$SSE=\sum(y_i-\hat y_i)^2=\boxed{0.0998}\qquad SST=\sum(y_i-\bar y)^2=0.5483\qquad SSR=SST-SSE=\boxed{0.4485}$$
  4. (d) ANOVA and overall significance. $df_{Reg}=3$, $df_{Err}=n-p=6-4=2$, $df_{Tot}=5$. $MSR=SSR/3=0.1495$, $MSE=SSE/2=0.0499$. $$F=\frac{MSR}{MSE}=\boxed{2.995}$$ $F_{0.05,3,2}=19.16$; since $2.995\ll19.16$, the overall regression is NOT significant at $\alpha=0.05$.
  5. (e) Individual coefficient significance. $se(\beta_j)=\sqrt{MSE\cdot(X'X)^{-1}_{jj}}$: $se(b_0)=12.611$, $se(b_1)=2.326$, $se(b_2)=0.0530$, $se(b_3)=0.0923$. $$t_0=-1.292,\quad t_1=1.765,\quad t_2=0.176,\quad t_3=1.349$$ $t_{0.025,2}=4.303$; every $|t_j|\lt4.303$, so $$\boxed{\text{none of the coefficients is individually significant}}$$
Summary
QuantityValue
$*,**,***$−585.1664, 0.0685, 4.2294
$SSE,SSR,SST$0.0998, 0.4485, 0.5483
$F$ (overall)2.995 vs. crit. 19.16 — n.s.
Significant coefficientsNone ($|t|\lt4.303$ for all)
Check — high $R^2$ does not mean a useful model here $R^2=SSR/SST=0.4485/0.5483=0.818$ (82%), which looks impressive, but with only $n=6$ observations and $p=4$ parameters there are just $2$ residual degrees of freedom — almost no statistical power to detect anything short of a near-perfect fit. (f) The model is NOT a reliable predictor of race time: despite the large apparent $R^2$, neither the overall $F$-test nor any individual coefficient clears significance, and with so few residual degrees of freedom the fit could easily be an artifact of over-fitting six points with four parameters. Considerably more race data would be needed before trusting this model.