18-Env-B7 Environmental Sampling and Analysis · May 2014
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
National Exams — May 2014 — 04-Env-B7 / Environmental Sampling and Analysis. 3 hours duration; open book (approved Sharp or Casio calculator only); standard normal, t-distribution and F-distribution tables supplied. Part A (Questions 1–3) is compulsory; Part B (Questions 4–6) asks for any two of three — all six are solved below for completeness. Each question is worth 20 marks.
Reference texts. Walpole, Myers, Myers & Ye, Probability & Statistics for Engineers and Scientists (statistical hypothesis testing, exploratory data analysis); Davis & Cornwell, Introduction to Environmental Engineering (6th ed.) (sampling design, QA/QC, environmental monitoring programs); U.S. EPA Guidance for Choosing a Sampling Design for Environmental Data Collection (QA/G-5S); Canadian Council of Ministers of the Environment (CCME) monitoring and reporting guidance.
Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.
Simple Random Sampling is the benchmark design for two properties: every unit in the population has an equal, known, non-zero probability of selection, and selections are made independently of one another. Together these make SRS statistically unbiased for the population mean/proportion, and every classical inference formula — the standard error of the mean, the t- and z-confidence intervals — is derived on the SRS assumption, so SRS is the yardstick every other design's relative efficiency is compared against.
SRS is not always used because it requires a complete, enumerable sampling frame (every possible location/time must be listable and equally accessible), which many environmental populations do not offer economically: contamination is often spatially clustered (plumes, hot spots) or seasonally structured, so a purely random spread of points may miss the very features of interest and wastes budget sampling low-information background areas at the same intensity as high-information ones.
Example. Delineating the boundary of a groundwater contaminant plume migrating from a fixed source: a stratified or judgmental design that concentrates wells near the suspected source and along the presumed flow path locates the plume edge far more efficiently than an SRS grid spread uniformly across the whole site, for the same number of wells.
Beyond SRS, environmental practice regularly uses: stratified random sampling (population divided into homogeneous strata – e.g., soil horizons, land-use zones – each sampled at random); systematic (grid) sampling (samples on a fixed spatial or temporal interval, e.g., every 50 m or every 4 hours); cluster sampling (whole groups/clusters selected at random, then all or a sub-sample within each cluster measured); judgmental (authoritative) sampling (locations chosen by expert knowledge of where contamination is most likely, e.g., beneath a former solvent tank); and composite sampling (several discrete samples physically combined into one before analysis, to estimate an average concentration at reduced analytical cost).
Environmental data sets typically show: (1) right (positive) skew, often approximately lognormal, because concentrations are bounded below by zero but can spike far above the mean; (2) non-detects – values reported only as “< detection limit”, i.e., left-censored data; (3) spatial and/or temporal autocorrelation, so nearby samples in space or time are not statistically independent; (4) high variability / heteroscedasticity, with variance that often scales with the mean (constant coefficient of variation rather than constant variance); and (5) a tendency toward outliers from genuine point-source events, sample contamination, or transcription/lab error, which must be investigated rather than automatically discarded.
| # | Statement | Verdict | Reasoning |
|---|---|---|---|
| (i) | For statistical significance, $\alpha$ must be greater than the p-value. | True | The decision rule for significance at level $\alpha$ is: reject $H_0$ when $p\le\alpha$, i.e. the p-value must fall below $\alpha$. |
| (ii) | Pearson r measures linear correlation only. | True | Pearson's r quantifies the strength of a straight-line relationship; a strong non-linear (e.g., quadratic) relationship can give $r\approx 0$. |
| (iii) | Power increases by decreasing $\alpha$ or increasing n. | False | Increasing n does raise power, but decreasing $\alpha$ shrinks the rejection region and lowers power – the statement has the effect of $\alpha$ backwards. |
| (iv) | Larger sample size makes the data normally distributed. | False | The Central Limit Theorem says the sampling distribution of the mean tends to normal as n grows – it says nothing about the shape of the underlying raw data, which is fixed by the population. |
| (v) | Mean square error measures precision. | False | $MSE=\text{Var}+\text{Bias}^2$ combines both precision (variance) and accuracy (bias); it measures overall accuracy, not precision alone. |