18-Env-B7 Environmental Sampling and Analysis · May 2018
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
National Exams, May 2018 — 04-Env-B7, Environmental Sampling and Analysis (3 hours, closed book, approved non-programmable calculator only, statistical tables provided). The paper instructs "answer all 5 questions"; this solution answers all 5 in full.
Reference texts: Walpole, Myers, Myers & Ye, Probability & Statistics for Engineers and Scientists (sampling designs, hypothesis tests, EDA, ANOVA); Davis & Cornwell, Introduction to Environmental Engineering, ch. 2 (sampling protocol, QA/QC, monitoring program design).
Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.
Part (a) — SRS properties, limitations and alternatives. Simple Random Sampling is the benchmark design because it has two defining statistical properties: (1) every possible sample of size n has an equal, known probability of selection, so the sample mean is an unbiased estimator of the population mean; and (2) because those selection probabilities are known exactly, the sampling error (variance of the estimator) can be computed directly from the sample itself, giving a defensible, quantifiable basis for statistical inference that ad-hoc designs cannot offer without extra assumptions. Every other design is judged against SRS by its relative efficiency — how much smaller a sample it needs for the same precision.
SRS is not always used in practice because: it requires a complete, enumerable sampling frame (a full list of every possible sampling location/time), which rarely exists for a continuous medium such as soil, groundwater or air; it is statistically inefficient when the population has known spatial or temporal structure that a stratified or systematic design could exploit to cut the required sample size; it can miss contamination hotspots purely by chance, since selection is probabilistic rather than targeted; and randomly selected locations are often widely scattered, making it more expensive and logistically difficult than a design tailored to site access.
Four other sampling methods used in environmental sampling: (1) Stratified random sampling (the population is divided into homogeneous strata — e.g. by depth, land use, or distance from a source — and sampled randomly within each); (2) Systematic sampling (samples taken on a fixed spatial grid or time interval); (3) Judgmental (authoritative) sampling (an expert targets locations expected to be contaminated, e.g. visible staining or a known spill point); (4) Composite sampling (several discrete grab samples are physically combined into one sample to characterize an average condition economically).
Part (b) — two sample parameters each for water, air and soil.
| Medium | Example parameter 1 | Example parameter 2 |
|---|---|---|
| Water body | pH | Dissolved oxygen (DO) |
| Air | PM2.5 (fine particulate) concentration | SO₂ (or NOₓ) concentration |
| Soil | Heavy-metal concentration (e.g. lead) | Moisture content (or organic-matter content) |
Part (c) — five typical characteristics of environmental data. (1) Concentration data are usually positively skewed, often approximately lognormal rather than normal, because concentrations are bounded below by zero but can range over several orders of magnitude above it. (2) Data are frequently censored — a meaningful fraction of results fall below the laboratory's method detection or reporting limit and are recorded only as "< MDL" (non-detects). (3) Samples are often spatially and/or temporally autocorrelated (successive monthly samples from one well, or two nearby soil borings, are not statistically independent), violating the simple-random-sample assumption behind many textbook tests. (4) The data typically show high variability/heterogeneity, including occasional extreme values (hotspots), so a few outliers can dominate the sample statistics. (5) Concentrations are non-negative by physical necessity, which constrains the shape of any fitted distribution and rules out a symmetric distribution like the normal as an exact model near zero.
Part (d) — true/false. Each statement is evaluated against the standard definitions used in hypothesis testing and descriptive statistics:
| # | Statement | Answer | Reason |
|---|---|---|---|
| (i) | For significance, α must be less than the p-value. | False | The decision rule for statistical significance is "reject $H_0$ if $p \le \alpha$" — i.e. the p-value must be at or below α, the reverse of what the statement claims. |
| (ii) | ANOVA tests differences among variances. | False | ANOVA tests differences among group means. It does so by comparing (partitioning) variance components — between-group vs. within-group — but the hypothesis under test concerns the means, not the variances themselves (a test like Levene's or Bartlett's tests variances). |
| (iii) | Increasing α increases the power of a test. | True | A larger α makes the rejection criterion less strict, so a given effect size is more likely to be declared significant — power (1−β) rises, at the cost of a higher Type I error rate. |
| (iv) | Larger sample size makes the data normally distributed. | False | The Central Limit Theorem concerns the sampling distribution of a statistic (e.g. the sample mean), which approaches normal as $n$ grows — it says nothing about the distribution of the individual raw data, which is fixed by the underlying population and does not become "more normal" as more observations are collected. |
| (v) | Pearson's r measures nonlinear association. | False | Pearson's correlation coefficient specifically measures the strength of linear association; two variables with a strong nonlinear (curvilinear) relationship can have $r \approx 0$. |
| (vi) | Attained significance (p-value) is independent of sample size. | False | For a fixed observed effect size, the p-value shrinks as $n$ increases (the standard error decreases), so the attained significance level depends directly on sample size — a trivial effect can become "statistically significant" purely from a very large $n$. |
| (vii) | Parametric tests are usually preferred for environmental data. | False | Environmental data are typically skewed, contain outliers and/or censored (below-detection-limit) values (Part c), which violates the normality assumption behind classical parametric tests; nonparametric (rank-based) tests are more robust to these features and are generally preferred unless normality has been specifically verified. |