23-Ind-B1 Reliability and Maintainability · December 2017
Question 8 of 10: Marathon Time — Multiple Regression via $(X'X)^{-1}$
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Notes on this paper
National Exams — December 2017 — 98-Ind-B1 Applied Probability & Statistics. Three-hour, closed-book exam; one of two permitted calculators (Sharp or Casio), one 8.5″×11.0″ aid sheet (both sides), statistical tables supplied. Format: three sections — Section A: do 2 of 4 (30 marks); Section B: do 2 of 3 (30 marks); Section C: do 2 of 3 (40 marks) — a 6-question, 100-mark paper as printed. All ten questions across the three sections are solved below for completeness.
Reference texts: Montgomery & Runger, Applied Statistics and Probability for Engineers (7th ed., Wiley) — discrete/continuous distributions (ch. 3–4), joint distributions (ch. 5), point/interval estimation and sample size (ch. 8), hypothesis testing incl. two-sample tests (ch. 9–10), simple linear regression (ch. 11), design and analysis of single-factor experiments (ch. 13). Montgomery, Peck & Vining, Introduction to Linear Regression Analysis (6th ed., Wiley) — multiple regression by matrices (ch. 2–3). Montgomery, Design and Analysis of Experiments (9th ed., Wiley) — multi-factor and $2^k$ factorial designs (ch. 5–6).
Question 8 (Section C.1): Marathon Time — Multiple Regression via $(X'X)^{-1}$ (20 marks)
Separately: the exam text states "$SST$ is 0.0548333," but recomputing $SST=\sum(y_i-\bar y)^2$ directly from the printed dataset gives $\boxed{SST=0.548333}$ — exactly $10\times$ the printed figure, a genuine decimal-point typo in the source. Using the printed $0.0548333$ would make $SSE=SST-SSR$ negative, which is impossible — itself proof the printed value is wrong. The correct $SST=0.548333$ is used throughout below.
Given. $n=6$ races:
$x_1$ (km×100)
4.6
4.5
4.0
4.1
4.4
4.2
$x_2$ (max dist, km)
27
31
30
25
28
32
$x_3$ (temp, °C)
11
10
25
22
14
17
$y$ (time, hrs)
4.3
3.8
3.5
3.6
3.5
3.4
Given partial $(X'X)^{-1}$ and $\beta=(X'X)^{-1}(X'Y)$:
3186.5127
*
−10.0932
−22.9764
−585.1664
108.4063
1.7075
4.2294
−10.0932
1.7075
0.0562
**
−22.9764
***
0.0685
0.1706
$\beta=(-16.29491,\ 4.1060,\ 0.0093,\ 0.1245)^T$.
Find. (a) $X'X$; (b) $*,**,***$; (c) $SSE,SSR,SST$; (d) overall $F$; (e) coefficient $t$-tests; (f) is the model useful?
Approach. Build $X$ with an intercept column, form $X'X$ directly, exploit that $(X'X)^{-1}$ must be symmetric to fill the blanks (cross-checked by an independent inverse computed from the raw data), then run the regression ANOVA and coefficient $t$-tests off the given $\beta$.
(b) Filling $*,**,***$ by symmetry. $(X'X)^{-1}$ is always symmetric, so each starred cell equals its mirror image across the diagonal — confirmed independently by inverting $X'X$ directly from the raw data (matches to 4 decimal places): $$*=(\text{row 1, col 2})=(\text{row 2, col 1})=\boxed{-585.1664}$$ $$**=(\text{row 3, col 4})=(\text{row 4, col 3})=\boxed{0.0685}$$ $$***=(\text{row 4, col 2})=(\text{row 2, col 4})=\boxed{4.2294}$$
(c) SSE, SSR, SST. Fitted values from $\hat y=X\beta$: $(4.213,\ 3.715,\ 3.521,\ 3.511,\ 3.775,\ 3.364)$. Using the corrected $SST=0.548333$ (see the check note above): $$SSE=\sum(y_i-\hat y_i)^2=\boxed{0.0998}\qquad SST=\sum(y_i-\bar y)^2=0.5483\qquad SSR=SST-SSE=\boxed{0.4485}$$
(d) ANOVA and overall significance. $df_{Reg}=3$, $df_{Err}=n-p=6-4=2$, $df_{Tot}=5$. $MSR=SSR/3=0.1495$, $MSE=SSE/2=0.0499$. $$F=\frac{MSR}{MSE}=\boxed{2.995}$$ $F_{0.05,3,2}=19.16$; since $2.995\ll19.16$, the overall regression is NOT significant at $\alpha=0.05$.
(e) Individual coefficient significance. $se(\beta_j)=\sqrt{MSE\cdot(X'X)^{-1}_{jj}}$: $se(b_0)=12.611$, $se(b_1)=2.326$, $se(b_2)=0.0530$, $se(b_3)=0.0923$. $$t_0=-1.292,\quad t_1=1.765,\quad t_2=0.176,\quad t_3=1.349$$ $t_{0.025,2}=4.303$; every $|t_j|\lt4.303$, so $$\boxed{\text{none of the coefficients is individually significant}}$$
Summary
Quantity
Value
$*,**,***$
−585.1664, 0.0685, 4.2294
$SSE,SSR,SST$
0.0998, 0.4485, 0.5483
$F$ (overall)
2.995 vs. crit. 19.16 — n.s.
Significant coefficients
None ($|t|\lt4.303$ for all)
Check — high $R^2$ does not mean a useful model here $R^2=SSR/SST=0.4485/0.5483=0.818$ (82%), which looks impressive, but with only $n=6$ observations and $p=4$ parameters there are just $2$ residual degrees of freedom — almost no statistical power to detect anything short of a near-perfect fit. (f) The model is NOT a reliable predictor of race time: despite the large apparent $R^2$, neither the overall $F$-test nor any individual coefficient clears significance, and with so few residual degrees of freedom the fit could easily be an artifact of over-fitting six points with four parameters. Considerably more race data would be needed before trusting this model.