18-Env-B7 Environmental Sampling and Analysis · December 2016
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
National Exams — December 2016 — 04-Env-B7 / Environmental Sampling and Analysis. 3 hours duration; closed book (approved non-programmable Sharp or Casio calculator only); t-distribution table supplied. Part A (Questions 1–3) is compulsory; Part B (Questions 4–6, "answer any 2") – all three are solved below for completeness.
Reference texts. Walpole, Myers, Myers & Ye, Probability & Statistics for Engineers and Scientists (statistical hypothesis testing, exploratory data analysis); Davis & Cornwell, Introduction to Environmental Engineering (6th ed.) (sampling design, QA/QC, environmental monitoring programs); U.S. EPA Guidance for Choosing a Sampling Design for Environmental Data Collection (QA/G-5S); Canadian Council of Ministers of the Environment (CCME) monitoring and reporting guidance.
Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.
Simple Random Sampling is the benchmark every other design is measured against because of two properties: every unit in the population has an equal, known, non-zero probability of selection, and selections are made independently of one another. Together these make the sample mean/proportion statistically unbiased for the population value, and every classical inference formula – the standard error of the mean, t- and z-confidence intervals – is derived directly from that assumption, so SRS is the yardstick against which the relative efficiency of every other design is expressed.
SRS is not always used in practice because it requires a complete, enumerable sampling frame (every possible location/time must be listable and equally accessible), which many environmental populations do not offer economically. Contamination is frequently spatially clustered (plumes, hot spots) or seasonally structured, so a purely random spread of sampling points can under-sample the very features of interest while spending the same budget characterizing low-information background areas – a targeted or structured design usually resolves the population of real interest far more efficiently.
Beyond SRS, environmental practice regularly uses four other designs: stratified random sampling (the population is divided into homogeneous strata – e.g., soil horizons, land-use zones – and each is sampled at random, improving precision when strata differ substantially); systematic (grid) sampling (points are laid out on a regular spatial or temporal grid, giving even spatial coverage and being simple to execute in the field); composite sampling (several discrete samples are physically combined into one before analysis, reducing analytical cost while estimating an average concentration); and judgmental (authoritative) sampling (an experienced investigator targets locations believed most likely to show contamination, such as a known spill point or hydraulically down-gradient location – efficient for hot-spot delineation but not statistically defensible for estimating a population mean).
Five characteristics recur across almost every environmental data set: (1) the data are frequently right (positively) skewed rather than normally distributed, often approximated by a lognormal model; (2) a meaningful fraction of results can fall below the method detection limit (censored/non-detect data), which complicates ordinary summary statistics; (3) the data show substantial spatial and/or temporal variability, and nearby-in-space or nearby-in-time observations are often autocorrelated rather than independent; (4) outliers are common and frequently real – a single hot-spot or spill event, not a measurement blunder; and (5) measurement/analytical uncertainty can be a significant fraction of the total observed variability, so the raw variance mixes true environmental variability with method error.
Given. Seven statements about statistical significance, ANOVA, test power, the Central Limit Theorem, Pearson's r, p-values, and parametric vs. nonparametric testing, to be judged true or false.
Find. The correct verdict for each of (i)–(vii), with reasoning.
| # | Statement | Verdict | Reasoning |
|---|---|---|---|
| (i) | The α-value must be less than the p-value for statistical significance. | False | The decision rule is: reject $H_0$ when $p\le\alpha$, i.e. $\alpha$ must be greater than or equal to $p$, not less than it. A test statistic is judged significant precisely when the p-value falls at or below the chosen $\alpha$ – the statement has the inequality backwards. |
| (ii) | An ANOVA is for testing differences among means. | True | Analysis of variance partitions total variability into between-group and within-group components in order to test whether two or more population means differ; the "variance" in its name refers to the mechanism used (comparing variance components), not the quantity being compared. |
| (iii) | Power increases by increasing α or increasing n. | True | Enlarging the rejection region (larger $\alpha$) or shrinking the standard error of the estimate (larger $n$) both raise the probability of correctly rejecting a false $H_0$, i.e., both raise power $=1-\beta$. |
| (iv) | Larger sample size makes the raw data normally distributed. | False | The Central Limit Theorem describes the sampling distribution of the mean tending toward normal as $n$ grows – it says nothing about the shape of the underlying raw data, which is fixed by the population it was drawn from. |
| (v) | Pearson's r can be used as a measure of nonlinear association. | False | Pearson's correlation coefficient measures the strength of linear association only; two variables can be perfectly related in a nonlinear (e.g., quadratic) way and still show $r\approx0$. |
| (vi) | The attained significance (p-value) is independent of sample size. | False | The p-value depends directly on $n$ through the standard error – for a fixed effect size, a larger $n$ shrinks the standard error and drives the p-value down, which is exactly the mechanism behind (iii)'s power increase. |
| (vii) | Parametric tests are usually preferred over nonparametric tests for environmental data. | False | Given part (b)'s characteristics – skewed distributions, censored non-detects, and outliers – the normality assumption behind classical parametric tests is frequently violated, so distribution-free (nonparametric) procedures are the more defensible default in this field, not parametric ones. |