NivaarExam PrepOfficial exam papers ↗

23-Ind-B1 Reliability and Maintainability · December 2018

Question 8 of 9: Multiple Linear Regression by Matrices

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

National Exams — December 2018 — 17-Ind-B1 Applied Probability & Statistics. Three-hour, closed-book exam; one of two permitted calculators (Sharp or Casio), one 8.5″×11.0″ aid sheet (both sides), statistical tables supplied. Format: three sections — Section A: do 2 of 4 (20 marks); Section B: do 2 of 3 (40 marks); Section C: do 1 of 2 (20 marks) — a 9-question, printed-total-80-mark paper as actually structured (the front page's own summary line states “Exam: 5 Questions. Total marks: 100”, but 20+40+20=80 by the section instructions and per-question mark values printed beside each question — the front page's 100 is internally inconsistent with its own section table; the 80-mark reading is used throughout, consistent with the section-by-section split printed on pages 2, 4 and 6). All nine questions across the three sections are solved below for completeness.

Reference texts: Montgomery & Runger, Applied Statistics and Probability for Engineers (7th ed., Wiley) — normal distribution and sums of normals (ch. 4–5), point/interval estimation (ch. 8), two-sample hypothesis testing (ch. 9–10), simple linear regression (ch. 11), single-factor ANOVA and multiple comparisons (ch. 13). Montgomery, Peck & Vining, Introduction to Linear Regression Analysis (6th ed., Wiley) — multiple regression by matrices, confidence/prediction intervals (ch. 2–3). Montgomery, Design and Analysis of Experiments (9th ed., Wiley) — Bartlett's test, $2^k$ factorial designs (ch. 3, 6).

Question 8 (Section C.1): Multiple Linear Regression by Matrices (20 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

Given. $n=5$ observations of $(x_1,x_2,x_3,y)$:

$x_1$8566748085
$x_2$1218171516
$x_3$710867
$y$2.151.091.261.571.92
$$X'X=\begin{pmatrix}5&390&78&38\\390&30682&6026&2922\\78&6026&1238&602\\38&2922&*&{**}\end{pmatrix}\qquad (X'X)^{-1}\text{(given)}=\begin{pmatrix}322.8483&-2.5213&-4.0284&-8.3081\\-2.5213&0.0204&0.0273&0.0661\\-4.0284&0.0273&0.1197&0.0047\\-8.3081&0.0661&0.0047&0.4055\end{pmatrix}$$

Find. (a) $*,**$; (b) least-squares model $\hat y=b_0+b_1x_1+b_2x_2+b_3x_3$; (c) ANOVA and overall significance; (d) is $b_2$ significant? (e) 95% CI and PI at $(80,15,5)$.

Approach. Fill the missing entries by symmetry and by direct summation from the raw data table; independently recompute $(X'X)^{-1}$ from the raw data as a cross-check on the given inverse; then form $\hat\beta=(X'X)^{-1}X'y$ and complete the standard multiple-regression ANOVA.

  1. (a) Missing $(X'X)$ entries. $X'X$ is symmetric, so the $(4,3)$ entry must equal the already-printed $(3,4)$ entry: $*=602$ (this is $\sum x_2x_3=12(7)+18(10)+17(8)+15(6)+16(7)=602$, confirmed directly from the raw table). The $(4,4)$ entry is $\sum x_3^2=7^2+10^2+8^2+6^2+7^2=298$, so $**=\boxed{298}$.
  2. Cross-check the given inverse against the raw data. Recomputing $(X'X)^{-1}$ directly from the completed raw-data matrix at full machine precision reproduces the printed inverse to within about $5\times10^{-5}$ in every entry — the given inverse is numerically trustworthy here. The full-precision inverse is used below for $\hat\beta$, since with only $n=5$ points and 4 parameters the design is nearly saturated and even this small rounding in the printed inverse shifts $\hat\beta$ measurably.
  3. (b) Least-squares coefficients. $X'y=(7.99,\ 636.73,\ 121.11,\ 58.89)'$ (column sums of $y$, $x_1y$, $x_2y$, $x_3y$). Then $\hat\beta=(X'X)^{-1}X'y$: $$\boxed{\hat y=-2.992+0.0587x_1-0.0634x_2+0.1319x_3}$$
  4. (c) ANOVA and overall significance, $\alpha=0.05$. Using the given $SST=0.7815$, $SSR=0.7749$: $SSE=SST-SSR=0.0066$ (matches the from-data fit's residual sum of squares, $0.00661$, essentially exactly — a strong cross-check that $\hat\beta$ is correct). With $p=3$ predictors, $df_R=3$, $df_E=n-p-1=5-3-1=1$: $$MS_R=\frac{0.7749}{3}=0.2583,\qquad MSE=\frac{0.0066}{1}=0.0066,\qquad F=\frac{MS_R}{MSE}=\boxed{39.1}$$ Critical value $F_{0.05,3,1}=215.7$. Since $F=39.1<215.7$, technically fail to reject $H_0$ — the regression is not significant at the 95% level, despite $R^2=SSR/SST=99.2\%$ of the variance being explained. This is not a contradiction: with only $n=5$ data points and 4 fitted parameters, only $df_E=1$ degree of freedom remains for error, so the $F$-test has essentially no power — a near-saturated model will always show a very high $R^2$, but the significance test cannot confirm it with so little error information. This is an inherent limitation of the exam's own tiny dataset, not a computational error.
  5. (d) Significance of the $x_2$ coefficient, $\alpha=0.05$. $$SE(b_2)=\sqrt{MSE\cdot[(X'X)^{-1}]_{33}}=\sqrt{0.0066\times0.1197}=0.0281,\qquad t=\frac{b_2}{SE(b_2)}=\frac{-0.0634}{0.0281}=\boxed{-2.257}$$ Critical value $t_{0.025,1}=12.706$ (the extreme critical value is itself a direct consequence of $df_E=1$). Since $|t|=2.257<12.706$, the $x_2$ coefficient is not significant at 95% — again a reflection of the design's near-total lack of error degrees of freedom, not evidence that $x_2$ is unimportant.
  6. (e) 95% CI and PI at $(x_1,x_2,x_3)=(80,15,5)$. With $\mathbf{x}_0'=(1,80,15,5)$, $\hat y_0=\mathbf{x}_0'\hat\beta=\boxed{1.410}$. $$\text{CI: }\hat y_0\pm t_{0.025,1}\sqrt{MSE\cdot\mathbf{x}_0'(X'X)^{-1}\mathbf{x}_0}=1.410\pm1.575=[-0.165,\ 2.986]$$ $$\text{PI: }\hat y_0\pm t_{0.025,1}\sqrt{MSE\left(1+\mathbf{x}_0'(X'X)^{-1}\mathbf{x}_0\right)}=1.410\pm1.883=[-0.473,\ 3.294]$$ The CI bounds the true mean response at this $(x_1,x_2,x_3)$ — its only source of uncertainty is how precisely the regression plane itself is estimated. The PI bounds a single new observation at the same point — it must additionally account for that observation's own random scatter about the true mean ($+1$ inside the square root), so it is always wider than the CI, as seen here ($1.883>1.575$).
Final Results
QuantityValue
$*,\,**$602, 298
Fitted model$\hat y=-2.992+0.0587x_1-0.0634x_2+0.1319x_3$
Overall $F$ (vs. crit.)39.1 (vs. 215.7) — not significant
$x_2$ coefficient $t$ (vs. crit.)$-2.257$ (vs. 12.706) — not significant
$\hat y_0$ at $(80,15,5)$1.410
95% CI$[-0.165,\ 2.986]$
95% PI$[-0.473,\ 3.294]$
Check With $n=5$ observations and 4 fitted parameters, only 1 error degree of freedom remains — this makes every significance test in this question (overall $F$, individual $t$) essentially unable to detect significance regardless of how strong the underlying relationship is, and makes both the CI and PI wide enough to include physically impossible negative values for what appears to be a strictly positive response. These are genuine, inevitable consequences of the exam's own near-saturated design, stated here so the numeric answers are not mistaken for a calculation error.