NivaarExam PrepOfficial exam papers ↗

18-Env-B7 Environmental Sampling and Analysis · May 2017

Question 1 of 7: Sampling Design, Data Characteristics and True/False

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

National Exams, May 2017 — 04-Env-B7, Environmental Sampling and Analysis (3 hours, closed book, approved non-programmable calculator only, statistical tables provided). The paper instructs "answer all 4 questions in Part A and any 2 questions in Part B"; as a study resource this solution answers all 7 questions in full, including all three Part B questions.

Reference texts: Walpole, Myers, Myers & Ye, Probability & Statistics for Engineers and Scientists (sampling designs, hypothesis tests, EDA/boxplots, ANOVA); Davis & Cornwell, Introduction to Environmental Engineering, ch. 2 (sampling protocol, QA/QC, monitoring program design).

Question 1: Sampling Design, Data Characteristics and True/False (20 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

Part (a) — SRS properties, limitations and alternatives. Simple Random Sampling is the sampling-design benchmark because it has two defining statistical properties: (1) every possible sample of size n drawn from the population has an equal, known probability of selection, which makes the sample mean an unbiased estimator of the population mean; and (2) because the selection probabilities are known, the sampling error (variance of the estimator) can itself be computed directly from the sample, giving a defensible, quantifiable basis for statistical inference that other ad-hoc designs cannot offer without extra assumptions. Every other design is judged against SRS by comparing its relative efficiency — how much smaller a sample it needs to achieve the same precision.

SRS is not always used in practice for several reasons: it requires a complete, enumerable sampling frame (a full list of every possible sampling location/time), which is rarely available for a continuous medium such as soil, groundwater or air; it is often statistically inefficient when the population has known spatial or temporal structure (SRS ignores that structure, whereas a stratified or systematic design can exploit it to reduce the number of samples needed for the same precision); it can miss contamination hotspots by chance, since selection is purely probabilistic rather than targeted; and it is frequently more expensive and logistically difficult, since randomly selected locations may be widely scattered and hard to access.

Four other sampling methods commonly used in environmental sampling: (1) Stratified random sampling (population divided into homogeneous strata, e.g. by depth or land use, sampled randomly within each); (2) Systematic sampling (samples on a fixed spatial grid or time interval); (3) Judgmental (authoritative) sampling (an expert targets locations expected to be contaminated, e.g. visible stains or a known spill point); (4) Composite sampling (several discrete grab samples physically combined into one sample to characterize an average condition economically).

Part (b) — five typical characteristics of environmental data. (1) Environmental concentration data are usually positively skewed and often approximately lognormal rather than normal, because concentrations are bounded below by zero but can range over several orders of magnitude above it. (2) Data are frequently censored — a meaningful fraction of results fall below the laboratory's method detection or reporting limit and are reported only as "< MDL". (3) Samples are often spatially and/or temporally autocorrelated (a groundwater well sampled monthly, or two nearby soil borings, are not statistically independent), violating the simple-random-sample assumption behind many textbook tests. (4) The data typically show high variability/heterogeneity, including occasional extreme values (hotspots), so a small number of outliers can dominate the sample statistics. (5) Concentrations are non-negative by physical necessity, which constrains the shape of any distribution fit to the data and rules out symmetric distributions (like the normal) as an exact model near zero.

Part (c) — true/false. Each statement is evaluated against the definitions used in hypothesis testing and descriptive statistics:

Question 1(c) — true/false determinations
#StatementAnswerReason
(i)For significance, α must be greater than the p-value.TrueThe decision rule for statistical significance is "reject $H_0$ if $p \le \alpha$" — the p-value must be at or below α, i.e. α is (at least) as large as p.
(ii)ANOVA tests differences among variances.FalseANOVA tests differences among group means. It does this by comparing (partitioning) variance components — between-group vs. within-group — but the hypothesis being tested is about the means, not the variances themselves (that would be a test like Levene's or Bartlett's test).
(iii)Decreasing α increases the power of a test.FalseDecreasing α makes the rejection criterion stricter, which decreases the probability of rejecting $H_0$ for a given effect size — that is, it decreases power (and decreases Type I error at the cost of increasing Type II error).
(iv)Larger sample size makes the data normally distributed.FalseThe Central Limit Theorem concerns the sampling distribution of a statistic (e.g. the sample mean), which approaches normal as $n$ grows — it says nothing about the distribution of the individual raw data, which is fixed by the underlying population and does not become "more normal" as more observations are collected.
(v)Pearson's r measures nonlinear association.FalsePearson's correlation coefficient specifically measures the strength of linear association; two variables with a strong nonlinear (e.g. curvilinear) relationship can have $r \approx 0$.
(vi)Attained significance (p-value) is independent of sample size.FalseFor a fixed observed effect size, the p-value shrinks as $n$ increases (the standard error decreases), so the attained significance level depends directly on sample size — this is why a trivial effect can become "statistically significant" with a very large $n$.
(vii)Nonparametric tests are usually preferred for environmental data.TrueEnvironmental data are typically skewed, contain outliers and/or censored (below-detection-limit) values (Part b), which violates the normality assumption behind classical parametric tests; nonparametric (rank-based) tests are more robust to these features and are generally preferred unless normality has been specifically verified.
← Paper overview