23-Ind-B1 Reliability and Maintainability · December 2019
Question 8 of 9: Multiple Regression by Matrices — Marathon Performance
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Notes on this paper
National Exams — December 2019 — 17-Ind-B1 Applied Probability & Statistics. Three-hour, closed-book exam; one of two permitted calculators (Sharp or Casio), one 8.5″×11.0″ aid sheet (both sides), statistical tables supplied. Format: three sections — Section A: do 2 of 4 (20 marks); Section B: do 2 of 3 (40 marks); Section C: do 1 of 2 (20 marks) — a 9-question, 80-mark paper as actually structured by the section table and each question's own printed marks (the front page's own summary line states “Exam: 5 Questions. Total marks: 100”, which is internally inconsistent with 20+40+20=80 from its own section breakdown and every question's printed mark value). All nine questions across the three sections are solved below for completeness.
Reference texts: Montgomery & Runger, Applied Statistics and Probability for Engineers (7th ed., Wiley) — probability density functions and moments (ch. 4), point/interval estimation (ch. 8), two-sample hypothesis testing and the sign test (ch. 9–10, 16), simple linear regression (ch. 11), single-factor ANOVA and multiple comparisons (ch. 13). Montgomery, Peck & Vining, Introduction to Linear Regression Analysis (6th ed., Wiley) — multiple regression by matrices, confidence/prediction intervals (ch. 2–3). Montgomery, Design and Analysis of Experiments (9th ed., Wiley) — Bartlett's test, Tukey's HSD, $2^k$ factorial designs (ch. 3, 6).
Find. (a) $(X'X)$; (b) the masked entries of the given $(X'X)^{-1}$; (c) $SSE,SSR,SST$; (d) the regression ANOVA table and its $F$-test; (e) individual coefficient significance; (f) an overall verdict on predictive usefulness.
Approach. Build the $4\times6$ design matrix $X=[\mathbf1,x_1,x_2,x_3]$, form $X'X$ directly from the raw data, recover the masked inverse entries by symmetry (and cross-check the full inverse against an independent from-data computation), then run the usual regression-ANOVA machinery.
(b) Recover *, **, *** by symmetry. $(X'X)^{-1}$ is symmetric, so each masked cell equals its mirror across the diagonal:
$$\ast=(\text{row 1, col 2})=(\text{row 2, col 1})=\boxed{-585.1664}$$
$$\ast\ast=(\text{row 3, col 4})=(\text{row 4, col 3})=\boxed{0.0685}$$
$$\ast\ast\ast=(\text{row 4, col 2})=(\text{row 2, col 4})=\boxed{4.2294}$$
An independent from-data computation of $(X'X)^{-1}$ reproduces every printed entry — including these three recovered ones — to within rounding, confirming both the completed table and that this paper's given inverse is trustworthy as printed.
(c) SSE, SSR, SST from the given $\hat\beta=(X'X)^{-1}(X'Y)=(-16.295,\ 4.106,\ 0.0093,\ 0.1245)$.
$$SST=\sum(y_i-\bar y)^2=0.5483\qquad(\bar y=3.6833)$$
$$\hat y_i=X_i\hat\beta,\qquad SSE=\sum(y_i-\hat y_i)^2=\boxed{0.0998}$$
$$SSR=SST-SSE=0.5483-0.0998=\boxed{0.4485}$$
(d) ANOVA table and overall $F$-test, $\alpha=0.05$. $p=3$ predictors, $df_R=3$, $df_E=n-p-1=6-3-1=2$:
$$MSR=\frac{SSR}{3}=0.1495,\qquad MSE=\frac{SSE}{2}=0.0499$$
$$F=\frac{MSR}{MSE}=\boxed{3.00}$$
Critical value $F_{0.05,3,2}=19.16$. Since $F=3.00\lt19.16$, fail to reject $H_0$: the overall regression is NOT significant at $\alpha=0.05$, despite $R^2=SSR/SST=0.818$ (a superficially strong fit, explaining 82% of the sample variation).
(f) Is the model a good predictor? (essay). No individual coefficient's $|t|$ reaches $4.303$, consistent with (d)'s non-significant overall $F$ — both point the same way. The high $R^2=0.818$ is not, by itself, evidence the model generalizes: with only $n=6$ data points fitting $p+1=4$ parameters, $df_E=2$ is extremely small, which both inflates every critical value (a near-saturated design can fit almost any 6 points closely just by having nearly as many parameters as observations) and makes every test drastically underpowered. The model's apparent fit is very plausibly an artifact of over-fitting rather than a genuine, statistically defensible linear relationship; a substantially larger sample (many more races) is needed before this model could be trusted to predict finishing time.
Final Results
Quantity
Value
(b) $\ast,\ast\ast,\ast\ast\ast$
$-585.1664,\ 0.0685,\ 4.2294$
(c) $SSE,\ SSR,\ SST$
$0.0998,\ 0.4485,\ 0.5483$
(d) Overall $F$
3.00 vs. crit. 19.16 — NOT significant
(e) Individual coefficients
none significant ($|t|\lt4.303$ for all four)
(f) Verdict
Not a statistically defensible predictor despite $R^2=0.818$ — too few observations for the model's complexity
here $n=6,p=3\Rightarrow df_E=2$): a real, even visually convincing linear relationship can routinely fail every classical significance test purely because so few degrees of freedom remain for error once the model's own parameters are fit. The non-significant result above is a structural consequence of the exam's small dataset, not evidence of a coding or derivation error.