18-Env-B7 Environmental Sampling and Analysis · December 2014
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
National Exams — December 2014 — 04-Env-B7 / Environmental Sampling and Analysis. 3 hours duration; open book (approved Sharp or Casio calculator only); t-distribution table supplied. Part A (Questions 1–3) is compulsory; Part B (Questions 4–6) asks for any two of three — all six are solved below for completeness. Each question is worth 20 marks.
Reference texts. Walpole, Myers, Myers & Ye, Probability & Statistics for Engineers and Scientists (statistical hypothesis testing, exploratory data analysis); Davis & Cornwell, Introduction to Environmental Engineering (6th ed.) (sampling design, QA/QC, environmental monitoring programs); U.S. EPA Guidance for Choosing a Sampling Design for Environmental Data Collection (QA/G-5S); Canadian Council of Ministers of the Environment (CCME) monitoring and reporting guidance.
Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.
Given. $n=10$ nitrate values (sorted): 2, 5, 8, 9, 13, 16, 24, 45, 57, 120 (mg/L). Printed summary: Median $=14.5$, $Q_1=7.3$, $Q_3=48.0$.
Find. The dotplot and boxplot (with fences marked) and whether any outliers exist.
Approach. The standard (Tukey) boxplot fences sit $1.5\times IQR$ beyond each quartile; any data point outside a fence is flagged as an outlier and plotted individually, with the whisker itself drawn only to the most extreme non-outlier data value.
Approach. Compare the mean to the median and use both a robust (quartile-based) and a classical (moment-based) skewness measure to characterize the shape of the distribution.
Both measures are positive and substantial, confirming what the boxplot already shows visually: the upper whisker (57) is far longer than the lower whisker (down to 2), the box itself is shifted toward the low end of the whisker span, and the mean (29.9) sits well above the median (14.5) – the single high value pulls the mean upward while the median, being robust to extremes, stays close to the bulk of the data. The data are therefore strongly right (positively) skewed, with one flagged high outlier (120 mg/L) driving most of that skew; a lognormal or other right-skewed model, rather than a normal distribution, would be the appropriate choice for any further statistical inference on this data set (e.g., before applying a t-test, per Question 2).
| Item | Result |
|---|---|
| IQR | 40.7 mg/L |
| Lower / upper fence | −53.75 / 109.05 mg/L |
| Outlier(s) | 120 mg/L (high outlier) |
| Whisker ends | 2 mg/L (low) / 57 mg/L (high) |
| Quartile skew, $S_Q$ | 0.646 (right-skewed) |
| Coefficient of skewness, $SK$ | 1.27 (strongly right-skewed) |