18-Env-B7 Environmental Sampling and Analysis · December 2019
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
National Exams, December 2019 — 18-Env-B7, Environmental Sampling and Analysis (3 hours, closed book, approved Sharp or Casio calculator, F-distribution table supplied with the paper). The paper instructs "answer all 5 questions"; this solution answers all 5 in full, with every sub-part addressed.
Reference texts: Walpole, Myers, Myers & Ye, Probability & Statistics for Engineers and Scientists (sampling designs, hypothesis tests, EDA, ANOVA); Davis & Cornwell, Introduction to Environmental Engineering, ch. 2 (sampling protocol, QA/QC, monitoring program design); Gilbert, Statistical Methods for Environmental Pollution Monitoring (environmental data characteristics, censored data, monitoring design).
Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.
Part (a) — SRS properties, limitations and alternatives. Simple Random Sampling is the benchmark design because it has two defining statistical properties: (1) every possible sample of size n has an equal, known probability of selection, so the sample mean is an unbiased estimator of the population mean; and (2) because those selection probabilities are known exactly, the sampling error (variance of the estimator) can be computed directly from the sample itself, giving a defensible, quantifiable basis for statistical inference that ad-hoc designs cannot offer without extra assumptions. Every other design is judged against SRS by its relative efficiency — how much smaller a sample it needs for the same precision.
SRS is not always used in practice because: it requires a complete, enumerable sampling frame (a full list of every possible sampling location/time), which rarely exists for a continuous medium such as soil, groundwater or air; it is statistically inefficient when the population has known spatial or temporal structure that a stratified or systematic design could exploit to cut the required sample size; it can miss contamination hotspots purely by chance, since selection is probabilistic rather than targeted; and randomly selected locations are often widely scattered, making it more expensive and logistically difficult than a design tailored to site access.
Four other sampling methods used in environmental sampling: (1) Stratified random sampling (the population is divided into homogeneous strata — e.g. by depth, land use, or distance from a source — and sampled randomly within each); (2) Systematic sampling (samples taken on a fixed spatial grid or time interval); (3) Judgmental (authoritative) sampling (an expert targets locations expected to be contaminated, e.g. visible staining or a known spill point); (4) Composite sampling (several discrete grab samples are physically combined into one sample to characterize an average condition economically).
Part (b) — two sample parameters each for water, air and soil.
| Medium | Example parameter 1 | Example parameter 2 |
|---|---|---|
| Water body | pH | Dissolved oxygen (DO) |
| Air | PM2.5 (fine particulate) concentration | SO₂ (or NOₓ) concentration |
| Soil | Heavy-metal concentration (e.g. lead) | Moisture content (or organic-matter content) |
Part (c) — five typical characteristics of environmental data. (1) Concentration data are usually positively skewed, often approximately lognormal rather than normal, because concentrations are bounded below by zero but can range over several orders of magnitude above it. (2) Data are frequently censored — a meaningful fraction of results fall below the laboratory's method detection or reporting limit and are recorded only as "< MDL" (non-detects). (3) Samples are often spatially and/or temporally autocorrelated (successive monthly samples from one well, or two nearby soil borings, are not statistically independent), violating the simple-random-sample assumption behind many textbook tests. (4) The data typically show high variability/heterogeneity, including occasional extreme values (hotspots), so a few outliers can dominate the sample statistics. (5) Concentrations are non-negative by physical necessity, which constrains the shape of any fitted distribution and rules out a symmetric distribution like the normal as an exact model near zero.
Part (d) — true/false. Each statement is evaluated against the standard definitions used in hypothesis testing and descriptive statistics. Several of these statements are worded as the logical negation of a familiar statement, so each must be read literally rather than recognised:
| # | Statement | Answer | Reason |
|---|---|---|---|
| (i) | For significance, α must be greater than the p-value. | True | The decision rule is "reject $H_0$ if $p \le \alpha$" — a result is declared significant exactly when the attained p-value falls at or below α, which is the direction this statement asserts. (Strictly the rule admits equality, so "greater than or equal to" is the exact wording; the point being tested is the direction of the inequality.) |
| (ii) | ANOVA is for testing differences among means. | True | Analysis of variance tests $H_0:\mu_1=\mu_2=\cdots=\mu_k$ — the hypothesis concerns group means. Its name describes its method (partitioning total variability into between-group and within-group components, exactly as Question 2 does), not the parameter under test; testing variances is a different procedure (Levene's or Bartlett's test). |
| (iii) | Reducing α increases the power of a test. | False | The reverse holds. Shrinking α shrinks the rejection region, so a real effect is less likely to be declared significant: β rises and power $1-\beta$ falls. Power is raised instead by increasing α, increasing the sample size, reducing measurement variability, or when the true effect size is larger. |
| (iv) | Larger sample size makes the data more normally distributed. | False | The Central Limit Theorem concerns the sampling distribution of a statistic such as $\bar{x}$, which approaches normality as $n$ grows. The raw observations keep the shape of the underlying population — a lognormal contaminant concentration stays lognormal however many samples are taken; a larger $n$ only reveals that skewed shape more clearly. |
| (v) | Pearson's r measures only linear association. | True | $r$ is the standardized covariance, so it quantifies only how well a straight line fits. A strong but curvilinear relationship (a symmetric parabola, for instance) can give $r \approx 0$, which is why a near-zero $r$ never demonstrates independence — only the absence of a linear trend. Monotonic nonlinear association is captured instead by Spearman's $\rho$ or Kendall's $\tau$. |
| (vi) | The attained significance level depends on sample size. | True | The test statistic scales with $\sqrt{n}$ through the standard error, so for a fixed observed effect the p-value shrinks as $n$ grows. This is why a practically negligible difference can be reported as "highly significant" in a very large environmental data set, and why significance must always be read alongside the effect size. |
| (vii) | Nonparametric tests are usually more powerful than parametric tests for environmental data. | True | For environmental data specifically — the qualifier carries the statement. Concentration data are typically right-skewed and carry outliers and censored non-detects (Part c), violating the normality assumption under which the parametric tests' power is derived; under those real conditions rank-based tests hold their nominal α and detect genuine differences more reliably. If the data really are normal the parametric test is the more powerful one, but only slightly — the Wilcoxon rank-sum test's asymptotic relative efficiency against the t-test is $3/\pi \approx 0.955$, a loss repaid many times over when normality fails. |