NivaarExam PrepOfficial exam papers ↗

18-Env-B7 Environmental Sampling and Analysis · December 2019

Question 2 of 5: Two-Way ANOVA — Fill in the Blanks

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

National Exams, December 2019 — 18-Env-B7, Environmental Sampling and Analysis (3 hours, closed book, approved Sharp or Casio calculator, F-distribution table supplied with the paper). The paper instructs "answer all 5 questions"; this solution answers all 5 in full, with every sub-part addressed.

Reference texts: Walpole, Myers, Myers & Ye, Probability & Statistics for Engineers and Scientists (sampling designs, hypothesis tests, EDA, ANOVA); Davis & Cornwell, Introduction to Environmental Engineering, ch. 2 (sampling protocol, QA/QC, monitoring program design); Gilbert, Statistical Methods for Environmental Pollution Monitoring (environmental data characteristics, censored data, monitoring design).

Question 2: Two-Way ANOVA — Fill in the Blanks (20 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

Given. A completely randomized two-factor field study with $a=3$ levels of Factor A (Location), $b=4$ levels of Factor B (Season) and $n=2$ replications in every treatment combination, giving $N=abn=24$ observations. Four cells of the ANOVA table are printed; the rest are blank.

2-way ANOVA output as printed on the exam (blank cells shown as —)
SourceSSDFMSF
Location (A)——80.17—
Season (B)12.46———
Interaction (AB)————
Error——3.79—
Total262.96———

Find. Every blank cell of the table, and the significance verdict for Location, for Season and for the Location×Season interaction at $\alpha=0.05$, with the study conclusions stated in plain terms.

Approach. Every degree of freedom follows from the design structure alone, after which the two identities $SS=MS\times df$ and $SS_{tot}=SS_A+SS_B+SS_{AB}+SS_{err}$ recover every remaining cell in a fixed order — no raw replicate data and no simultaneous equations are needed, because the table is over-determined by construction.

  1. Step 1 — write down every degree of freedom from the design. The df of a crossed two-factor design with replication depend only on $a$, $b$ and $n$: $$df_A=a-1=2,\quad df_B=b-1=3,\quad df_{AB}=(a-1)(b-1)=6,\quad df_{err}=ab(n-1)=12,\quad df_{tot}=N-1=23$$ The partition must be exact, and it is: $2+3+6+12=23=df_{tot}$. This check is worth doing first, because every mean square below divides by one of these numbers.
  2. Step 2 — recover the sums of squares wherever the mean square is printed. Rearranging $MS=SS/df$ gives $SS=MS\times df$, which unlocks the Location and Error rows: $$SS_A = MS_A \times df_A = 80.17\times 2 = 160.34,\qquad SS_{err} = MS_{err}\times df_{err} = 3.79\times 12 = 45.48$$
  3. Step 3 — recover the mean square wherever the sum of squares is printed. The Season row is the mirror image of Step 2: $$MS_B = \frac{SS_B}{df_B} = \frac{12.46}{3} = 4.153$$
  4. Step 4 — obtain the interaction by subtraction. The sums of squares are additive, so with $SS_A$, $SS_B$ and $SS_{err}$ all now known the interaction is whatever the total has left over: $$SS_{AB} = SS_{tot} - SS_A - SS_B - SS_{err} = 262.96 - 160.34 - 12.46 - 45.48 = 44.68$$ $$MS_{AB} = \frac{SS_{AB}}{df_{AB}} = \frac{44.68}{6} = 7.447$$ That $SS_{AB}$ comes out positive is itself a confirmation that the printed cells have been read from the correct columns — a column misread would drive it negative.
  5. Step 5 — form the F-ratios and test each against the supplied table. In a fixed-effects design every effect is tested against the error mean square, $F=MS_{effect}/MS_{err}$: $$F_A = \frac{80.17}{3.79} = 21.153,\qquad F_B = \frac{4.153}{3.79} = 1.096,\qquad F_{AB} = \frac{7.447}{3.79} = 1.965$$ Reading the $\alpha=0.05$ critical values from the F-table supplied with the exam, using each effect's own numerator df against $df_{err}=12$: $F_{0.05}(2,12)=3.8853$, $F_{0.05}(3,12)=3.4903$ and $F_{0.05}(6,12)=2.9961$. Comparing each statistic with its own critical value: $$\boxed{F_A=21.153 > 3.8853 \Rightarrow \text{Location effect significant}}$$ $$F_B=1.096 < 3.4903 \Rightarrow \text{Season effect not significant};\qquad F_{AB}=1.965 < 2.9961 \Rightarrow \text{interaction not significant}$$
Location (A)3.885F = 21.153df 2, 12Season (B)3.490F = 1.096df 3, 12Interaction (AB)2.996F = 1.965df 6, 1205101520F-ratio (effect mean square divided by error mean square)significantnot significantcritical F at 5%
Each effect's computed F-ratio (bar) against its own 5% critical value (dashed red line, labelled). Only Location clears its threshold, and it does so by a wide margin; Season and the interaction fall well short, so both are indistinguishable from error variation.
Final Results — Question 2, completed 2-way ANOVA table
SourceSSDFMSF$F_{0.05}$Significant at 5%?
Location (A)160.34280.1721.1533.8853Yes
Season (B)12.4634.1531.0963.4903No
Interaction (AB)44.6867.4471.9652.9961No
Error45.48123.79———
Total262.9623— (n/a)— (n/a)——

Conclusions from the study. At the 5% significance level the interaction is tested first, because a significant interaction would make the main effects uninterpretable on their own. It is not significant here ($F_{AB}=1.97$ against a critical 3.00), so the Location and Season effects may be read directly and independently. The Location effect is highly significant ($F_A=21.15$ against a critical 3.89 — more than five times the threshold): the measured variable differs genuinely among the three sampling locations, and location accounts for $160.34/262.96 = 61\%$ of the total variability in the study. The Season effect is not significant ($F_B=1.10$ against a critical 3.49); with $MS_B=4.15$ barely above the error mean square of 3.79, season-to-season variation is essentially indistinguishable from experimental error at this sample size. Practically, the study says that where a sample is taken drives the response and when it is taken does not, and — because the interaction is absent — the ranking among locations is stable from season to season. A follow-up multiple-comparison procedure (Tukey's HSD or Fisher's LSD on the three location means) would be the appropriate next step to identify which locations differ, since the F-test establishes only that at least one does.

Note — how the printed cells were read, and the Total row
The exam prints one number per row without repeating the column headings on each line, so the first task is to place each value in its correct column: 80.17 sits in the MS column of the Location row, 12.46 in the SS column of the Season row, 3.79 in the MS column of the Error row and 262.96 in the SS column of the Total row. This reading is confirmed because it is the only one that leaves a positive interaction sum of squares after the Step 4 subtraction. Separately, the conventional ANOVA Total row reports only $SS_{tot}$ and $df_{tot}$: a mean square and an F-ratio are undefined for the total (there is no "total" variance component to test against error), so those two cells are marked n/a rather than filled.