18-Env-B7 Environmental Sampling and Analysis · May 2016
Question 4 of 6: Dotplot, Boxplot and Skewness of Nitrate Data
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Notes on this paper
National Exams — May 2016 — 04-Env-B7 / Environmental Sampling and Analysis. 3 hours duration; closed book (approved non-programmable Sharp or Casio calculator only); t-distribution table supplied. Part A (Questions 1–3) is compulsory; Part B (Questions 4–6, see note below) asks for any two of three — all six are solved below for completeness. Each question is worth 20 marks.
Check: the printed Part B instruction ("answer any 2 questions") is followed by three distinct 20-mark tasks – Question 4 (nitrate dotplot/boxplot), Question 5 (eight-term definition list), and a further open-ended "environmental monitoring program" prompt that the exam's own marking scheme lists as a separate item 6 (20 marks, descriptive) even though the source page prints it directly under Question 5's heading with no "6." label of its own. This is treated here as Question 6, matching the marking scheme's six-item structure; all three Part B questions are solved in full per the standing "answer everything" convention.
Reference texts. Walpole, Myers, Myers & Ye, Probability & Statistics for Engineers and Scientists (statistical hypothesis testing, exploratory data analysis); Davis & Cornwell, Introduction to Environmental Engineering (6th ed.) (sampling design, QA/QC, environmental monitoring programs); U.S. EPA Guidance for Choosing a Sampling Design for Environmental Data Collection (QA/G-5S); Canadian Council of Ministers of the Environment (CCME) monitoring and reporting guidance.
Question 4: Dotplot, Boxplot and Skewness of Nitrate Data (20 marks)
Find. The dotplot and boxplot (with fences marked) and whether any outliers exist.
Approach. The standard (Tukey) boxplot fences sit $1.5\times IQR$ beyond each quartile; any data point outside a fence is flagged as an outlier and plotted individually, with the whisker itself drawn only to the most extreme non-outlier data value.
Outlier check. No data point lies below the (negative) lower fence, so the lower whisker runs to the minimum, 2. The maximum value, 120, exceeds the upper fence (109.05), so 120 is flagged as a single high outlier; the upper whisker is drawn only to the largest non-outlier value, 57 – never to the fence itself.
Fig. 2 – Dotplot (top) and boxplot (bottom) of the 10 nitrate values on a shared axis. The box spans $Q_1$ to $Q_3$ with the median line; whiskers extend only to the most extreme non-outlier data value on each side (2 and 57); the point beyond the upper fence (120) is plotted separately as an outlier, with the fence itself shown as a reference line, not a whisker endpoint.
(b) Characteristics, Quartile Skew, and Coefficient of Skewness
Approach. Compare the mean to the median and compute both a robust (quartile-based) and a classical (moment-based) skewness measure to characterize the shape of the distribution.
Coefficient of skewness (Pearson). $SK=\dfrac{3(\text{Mean}-\text{Median})}{StDev}=\dfrac{3(29.9-14.5)}{36.4}=\dfrac{46.2}{36.4}$
$$SK=\boxed{1.27}$$
Both measures are positive and substantial, confirming what the boxplot already shows visually: the upper whisker (57) is far longer than the lower whisker (down to 2), the box itself sits toward the low end of the whisker span, and the mean (29.9) lies well above the median (14.5) – the single high value pulls the mean upward while the median, being robust to extremes, stays close to the bulk of the data. The data are therefore strongly right (positively) skewed, with one flagged high outlier (120 mg/L) driving most of that skew; a lognormal or other right-skewed model, rather than a normal distribution, would be the appropriate choice for any further statistical inference on this data set.