18-Env-B7 Environmental Sampling and Analysis · December 2014
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
National Exams — December 2014 — 04-Env-B7 / Environmental Sampling and Analysis. 3 hours duration; open book (approved Sharp or Casio calculator only); t-distribution table supplied. Part A (Questions 1–3) is compulsory; Part B (Questions 4–6) asks for any two of three — all six are solved below for completeness. Each question is worth 20 marks.
Reference texts. Walpole, Myers, Myers & Ye, Probability & Statistics for Engineers and Scientists (statistical hypothesis testing, exploratory data analysis); Davis & Cornwell, Introduction to Environmental Engineering (6th ed.) (sampling design, QA/QC, environmental monitoring programs); U.S. EPA Guidance for Choosing a Sampling Design for Environmental Data Collection (QA/G-5S); Canadian Council of Ministers of the Environment (CCME) monitoring and reporting guidance.
Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.
Simple Random Sampling is the benchmark design for two properties: every unit in the population has an equal, known, non-zero probability of selection, and selections are made independently of one another. Together these make SRS statistically unbiased for the population mean/proportion, and every classical inference formula – the standard error of the mean, t- and z-confidence intervals – is derived on the SRS assumption, so SRS is the yardstick every other design's relative efficiency is measured against.
SRS is not always used in practice because it requires a complete, enumerable sampling frame (every possible location/time must be listable and equally accessible), which many environmental populations do not offer economically. Contamination is often spatially clustered (plumes, hot spots) or seasonally structured, so a purely random spread of sampling points can miss the very features of interest while wasting budget on low-information background areas sampled at the same intensity as high-information ones.
Beyond SRS, environmental practice regularly uses five other designs: stratified random sampling (the population is divided into homogeneous strata – e.g., soil horizons, land-use zones – each sampled at random); systematic (grid) sampling (samples on a fixed spatial or temporal interval, e.g., every 50 m or every 4 hours); cluster sampling (whole groups/clusters selected at random, then all or a sub-sample within each measured); judgmental (authoritative) sampling (locations chosen by expert knowledge of where contamination is most likely, e.g., beneath a former solvent tank); and composite sampling (several discrete samples physically combined into one before analysis, to estimate an average concentration at reduced analytical cost).
Environmental data sets typically show: (1) right (positive) skew, often approximately lognormal, because concentrations are bounded below by zero but can spike far above the mean; (2) non-detects – values reported only as “< detection limit”, i.e., left-censored data; (3) spatial and/or temporal autocorrelation, so nearby samples in space or time are not statistically independent; (4) high variability / heteroscedasticity, with variance that often scales with the mean (a roughly constant coefficient of variation rather than a constant variance); and (5) a tendency toward outliers from genuine point-source events, sample contamination, or transcription/lab error, which must be investigated rather than automatically discarded.
| # | Statement | Verdict | Reasoning |
|---|---|---|---|
| (i) | For statistical significance, $\alpha$ must be less than the p-value. | False | The decision rule is: reject $H_0$ when $p\le\alpha$. That means the p-value must be less than (or equal to) $\alpha$ – the statement has the inequality backwards. |
| (ii) | An ANOVA is for testing differences among means. | True | Analysis of variance partitions total variability into between-group and within-group components to test whether two or more population means differ. |
| (iii) | Power increases by increasing $\alpha$ or increasing n. | True | Enlarging the rejection region (larger $\alpha$) or reducing the standard error of the estimate (larger n) both raise the probability of correctly rejecting a false $H_0$, i.e., both raise power $=1-\beta$. |
| (iv) | Larger sample size makes the raw data normally distributed. | False | The Central Limit Theorem says the sampling distribution of the mean tends to normal as n grows – it says nothing about the shape of the underlying raw data, which is fixed by the population it was drawn from. |
| (v) | Mean square error is another term for the variance. | False | $MSE=\text{Var}(\hat\theta)+[\text{Bias}(\hat\theta)]^2$; it equals the variance only for an unbiased estimator ($\text{Bias}=0$) and in general also captures accuracy, not just precision. |