NivaarExam PrepOfficial exam papers ↗

18-Env-B7 Environmental Sampling and Analysis · May 2014

Question 2 of 6: Paired vs. Two-Sample T-Tests for Two Measurement Methods

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

National Exams — May 2014 — 04-Env-B7 / Environmental Sampling and Analysis. 3 hours duration; open book (approved Sharp or Casio calculator only); standard normal, t-distribution and F-distribution tables supplied. Part A (Questions 1–3) is compulsory; Part B (Questions 4–6) asks for any two of three — all six are solved below for completeness. Each question is worth 20 marks.

Reference texts. Walpole, Myers, Myers & Ye, Probability & Statistics for Engineers and Scientists (statistical hypothesis testing, exploratory data analysis); Davis & Cornwell, Introduction to Environmental Engineering (6th ed.) (sampling design, QA/QC, environmental monitoring programs); U.S. EPA Guidance for Choosing a Sampling Design for Environmental Data Collection (QA/G-5S); Canadian Council of Ministers of the Environment (CCME) monitoring and reporting guidance.

Question 2: Paired vs. Two-Sample T-Tests for Two Measurement Methods (20 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

(a) Which Test Is Appropriate

Approach. Identify whether the two samples are independent or paired, then match the design to the correct t-procedure.

Method A and Method B were both applied to the same 11 field measurements – each pair $(A_i,B_i)$ comes from the same sampling occasion/location, so the two columns are naturally correlated, not independent. The Paired T-test (Test 1) is therefore the statistically appropriate one: it works directly on the 11 differences $d_i=A_i-B_i$, so any location-to-location variation common to both methods cancels out of the comparison, leaving only the true systematic difference between methods. Treating the two columns as independent samples (Tests 2 and 3) throws away that cancellation, inflates the estimated standard error of the difference, and understates the true significance of any real difference between the methods.

Conclusion. With the correct paired test, $P=0.037<0.05$: we reject $H_0:\mu_d=0$ in favour of $H_a:\mu_d<0$. There is a statistically significant difference between the two methods at the 5% level – Method A reads, on average, about 1.8 ppm lower than Method B – which is the opposite of what the analyst had hoped to demonstrate. (The two independent-sample tests, run on the wrong design, would have wrongly failed to detect this difference.)

(b) Assumption and Graphical Check

The paired t-test's main assumption is that the population of paired differences $d_i=A_i-B_i$ is (at least approximately) normally distributed – with only $n=11$ pairs, the Central Limit Theorem cannot be relied on to normalize a non-normal difference distribution, so normality of the $d_i$ themselves is required for the t-reference distribution to be valid.

A simple graphical check is a normal probability plot of the 11 differences (or, equivalently for a quick check, a dotplot/boxplot of the $d_i$): if the points fall close to a straight line on the probability plot – or the dotplot is reasonably symmetric and unimodal with no strong skew or isolated extreme points – normality is a reasonable working assumption.

-6-4-20246Method A - Method B (ppm)Dotplot of the 11 paired differences
Fig. 1 – Dotplot of $d_i=A_i-B_i$ for the 11 paired measurements: single-modal, roughly symmetric about the mean ($-1.81$), with no isolated outlier – consistent with the normality assumption.

(c) Test Statistics

Given. Raw paired data ($n=11$), reproduced above; MINITAB summary statistics for each of the three tests as printed in the question.

Find. The value of $t$ underlying each of the three printed P-values.

Approach. Test 1 is a one-sample t on the differences ($t=\bar d/SE_{\bar d}$); Tests 2 and 3 use the same Welch (unequal-variance) two-sample statistic ($t=(\bar A-\bar B)/SE_{diff}$, $SE_{diff}=\sqrt{SE_A^2+SE_B^2}$) – only the tail used to convert $t$ to a P-value differs between the one- and two-tailed versions.

  1. Test 1 – paired t-statistic. $t=\dfrac{\bar d}{SE_{\bar d}}=\dfrac{-1.809}{0.909}$ $$t_{paired}=\boxed{-1.99}\quad(df=n-1=10)$$ Checking against the t-table at $df=10$: $1.99$ lies between $t_{0.05}=1.812$ and $t_{0.025}=2.228$, so the one-tailed P-value is bracketed between 0.025 and 0.05 – consistent with the printed $P=0.037$.
  2. Tests 2 & 3 – two-sample (Welch) t-statistic. From Method A ($\bar A=64.918$, $s_A=1.885$) and Method B ($\bar B=66.727$, $s_B=3.238$): $SE_A=s_A/\sqrt{11}=0.568$, $SE_B=s_B/\sqrt{11}=0.976$, so $SE_{diff}=\sqrt{0.568^2+0.976^2}=1.130$. $$t_{two\text{-}sample}=\frac{\bar A-\bar B}{SE_{diff}}=\frac{-1.809}{1.130}\;\Rightarrow\;\boxed{t=-1.60}\quad(df\approx16,\text{ Welch--Satterthwaite})$$ The same $t=-1.60$ underlies both Test 2 and Test 3 – Test 2's one-tailed P (0.064) is the single-tail area beyond $t=-1.60$ at $df=16$, and Test 3's two-tailed P (0.129) is exactly double it ($2\times0.064\approx0.129$), confirming the arithmetic.
Final results – Question 2
ItemResult
(a) Appropriate testPaired T-test (Test 1) – the 11 pairs are correlated, not independent
(a) Conclusion$P=0.037<0.05$: reject $H_0$ – Method A reads significantly lower than Method B
(b) AssumptionPaired differences $d_i$ are approximately normally distributed
(c) $t$, Test 1 (paired)$-1.99$ ($df=10$)
(c) $t$, Tests 2 & 3 (two-sample)$-1.60$ ($df\approx16$)