18-Env-B7 Environmental Sampling and Analysis · May 2016
Question 2 of 6: Paired T-Test for Two CO₂ Measurement Methods
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Notes on this paper
National Exams — May 2016 — 04-Env-B7 / Environmental Sampling and Analysis. 3 hours duration; closed book (approved non-programmable Sharp or Casio calculator only); t-distribution table supplied. Part A (Questions 1–3) is compulsory; Part B (Questions 4–6, see note below) asks for any two of three — all six are solved below for completeness. Each question is worth 20 marks.
Check: the printed Part B instruction ("answer any 2 questions") is followed by three distinct 20-mark tasks – Question 4 (nitrate dotplot/boxplot), Question 5 (eight-term definition list), and a further open-ended "environmental monitoring program" prompt that the exam's own marking scheme lists as a separate item 6 (20 marks, descriptive) even though the source page prints it directly under Question 5's heading with no "6." label of its own. This is treated here as Question 6, matching the marking scheme's six-item structure; all three Part B questions are solved in full per the standing "answer everything" convention.
Reference texts. Walpole, Myers, Myers & Ye, Probability & Statistics for Engineers and Scientists (statistical hypothesis testing, exploratory data analysis); Davis & Cornwell, Introduction to Environmental Engineering (6th ed.) (sampling design, QA/QC, environmental monitoring programs); U.S. EPA Guidance for Choosing a Sampling Design for Environmental Data Collection (QA/G-5S); Canadian Council of Ministers of the Environment (CCME) monitoring and reporting guidance.
Question 2: Paired T-Test for Two CO₂ Measurement Methods (20 marks)
MINITAB two-sample summary and the paired differences $d_i=A_i-B_i$
Location
Method A
Method B
$d_i=A_i-B_i$
1
66.3
71.3
−5.0
2
63.5
60.4
3.1
3
64.9
64.6
0.3
4
61.8
63.9
−2.1
5
64.3
68.8
−4.5
6
64.7
70.1
−5.4
7
65.1
64.8
0.3
8
64.5
68.9
−4.4
9
68.4
65.8
2.6
10
63.2
66.2
−3.0
11
67.4
69.2
−1.8
Mean / StDev / SE
64.918 / 1.885 / 0.568
66.727 / 3.238 / 0.976
−1.809 / 3.016 / 0.909
Find. The appropriate t-procedure, its test statistic, and whether Method A is significantly different from Method B at $\alpha=0.10$.
Approach. Determine whether the two columns are independent or paired, then apply the matching t-procedure to the 11 differences and compare the statistic to the supplied t-table.
Identify the design. Method A and Method B were both applied at the same 11 locations – each pair $(A_i,B_i)$ shares a common sampling occasion, so the two columns are naturally correlated rather than independent. The paired t-test is therefore the statistically appropriate procedure: it works directly on the 11 differences $d_i=A_i-B_i$, so location-to-location variability common to both methods cancels out of the comparison, leaving only the genuine systematic difference between the methods. An independent two-sample test would ignore that correlation, inflate the estimated standard error of the difference, and understate the true significance of any real difference – exactly the pitfall the "partial results" table is built to test for.
Hypotheses. $H_0:\mu_d=0$ (the methods are not significantly different) vs. $H_a:\mu_d\ne0$ (two-sided, matching the "not significantly different" wording, which is a statement of equality).
Test statistic. $t=\dfrac{\bar d}{SE_{\bar d}}=\dfrac{-1.809}{0.909}$
$$t=\boxed{-1.99}\qquad(df=n-1=10)$$
Critical value. For a two-tailed test at $\alpha=0.10$, each tail carries 0.05; from the supplied t-table at $df=10$, $t_{0.05,10}=1.812$.
Decision. $|t|=1.99>1.812=t_{crit}$, so we reject $H_0$. (Equivalently, the two-tailed p-value is $\approx0.075$, which is less than $\alpha=0.10$.)
Conclusion. At the 10% significance level, Method A and Method B produce statistically significantly different $CO_2$ readings (Method A averages about 1.8 ppm lower). This is the opposite of what the exam scenario hoped to demonstrate – running the correct paired test, rather than an independent two-sample test, is exactly what exposes the real difference that a naive independent-samples analysis on the same 11 numbers would tend to miss.
(b) Main Assumption and Graphical Check
The paired t-test's main assumption is that the population of paired differences $d_i=A_i-B_i$ is (at least approximately) normally distributed. With only $n=11$ pairs, the Central Limit Theorem cannot be relied on to normalize a non-normal difference distribution on its own, so approximate normality of the $d_i$ themselves is required for the t reference distribution used in part (a) to be valid.
A simple graphical check is a normal probability (quantile-quantile) plot of the 11 sorted differences against their theoretical normal quantiles: if the points fall close to a straight line, normality is a reasonable working assumption; systematic curvature or a lone far-outlying point would instead signal a problem.
Fig. 1 – Normal probability plot of the 11 paired differences $d_i=A_i-B_i$ against their theoretical normal quantiles, with a fitted reference line through the sample mean and standard deviation. The points track the line reasonably closely with no severe curvature or isolated outlier, supporting the normality assumption used for the paired t-test in part (a).
Final results – Question 2
Item
Result
(a) Appropriate test
Paired t-test – the 11 locations are the same for both methods (correlated, not independent)
(a) Test statistic
$t=-1.99$ ($df=10$)
(a) Critical value ($\alpha=0.10$, two-tailed)
$t_{0.05,10}=1.812$
(a) Decision / Conclusion
Reject $H_0$ – Method A and Method B are significantly different at 10%
(b) Main assumption
Paired differences $d_i$ are approximately normally distributed
(b) Graphical check
Normal probability plot of $d_i$ – points track the reference line, no severe departure