NivaarExam PrepOfficial exam papers ↗

18-Env-B7 Environmental Sampling and Analysis · May 2017

Question 5 of 7: Nitrate Concentrations — Dotplot, Boxplot and Skewness

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

National Exams, May 2017 — 04-Env-B7, Environmental Sampling and Analysis (3 hours, closed book, approved non-programmable calculator only, statistical tables provided). The paper instructs "answer all 4 questions in Part A and any 2 questions in Part B"; as a study resource this solution answers all 7 questions in full, including all three Part B questions.

Reference texts: Walpole, Myers, Myers & Ye, Probability & Statistics for Engineers and Scientists (sampling designs, hypothesis tests, EDA/boxplots, ANOVA); Davis & Cornwell, Introduction to Environmental Engineering, ch. 2 (sampling protocol, QA/QC, monitoring program design).

Question 5: Nitrate Concentrations — Dotplot, Boxplot and Skewness (20 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

Given. 13 nitrate readings (mg/L): 13, 5, 8, 120, 9, 2, 45, 24, 57, 16, 48, 89, 8; and the computer-package summary N=13, Mean=34.3, Median=16.0, StDev=36.4, Min=2.0, Max=120.0.

Find. A dotplot and boxplot of the data with fences located, $Q_1$/$Q_3$, any outliers (a); a description of the data's characteristics plus the quartile skew and coefficient of skewness (b).

Approach. Sort the raw data to locate the quartiles (Tukey's hinge method: median of the lower half and median of the upper half), compute the IQR-based fences to classify outliers, and combine the shape of the dotplot/boxplot with the given mean/median/StDev to characterize the distribution and its skewness.

  1. Part (a) — Step 1: sort and locate the quartiles. Sorted data: 2, 5, 8, 8, 9, 13, 16, 24, 45, 48, 57, 89, 120 (median = 7th value = 16.0, matching the given summary). Lower half (below the median, 6 values): 2, 5, 8, 8, 9, 13 → $Q_1=$ median of these $=(8+8)/2=8.0$. Upper half (above the median, 6 values): 24, 45, 48, 57, 89, 120 → $Q_3=$ median of these $=(48+57)/2=52.5$.
  2. Step 2 — fences. $IQR=Q_3-Q_1=52.5-8.0=44.5$. $$\boxed{LF = Q_1-1.5\,IQR = 8.0-66.75=-58.75, \qquad UF = Q_3+1.5\,IQR = 52.5+66.75=119.25}$$ The lower fence is a large negative number, off the physical (concentration ≥ 0) and printed axis range entirely — it still functions as the outlier threshold, it is just never a plotted point.
  3. Step 3 — outliers and whiskers. Comparing every data value to $[LF,UF]=[-58.75,\ 119.25]$: only $120$ exceeds $UF=119.25$, so 120 mg/L is a (mild) outlier; no value is below $LF$. The boxplot whiskers are therefore drawn only to the most extreme non-outlier data values: the lower whisker to the minimum, $2.0$, and the upper whisker to $89.0$ (the largest value that is still $\le UF$), with $120$ plotted as a separate point beyond the upper whisker.
Dotplot020406080100120Boxplot120 (outlier)Q1=8Med=16Q3=52.5whisker 2whisker 89020406080100120Nitrate concentration (mg/L)Lower fence = -58.75 (off-scale, < 0) — no low-side outliers
Dotplot (top) and boxplot (bottom) of the 13 nitrate readings, sharing one concentration axis. The box spans $Q_1=8.0$ to $Q_3=52.5$ with the median at 16.0; whiskers extend only to the most extreme non-outlier values (2.0 and 89.0); the value 120 is plotted separately as an outlier, above the (off-scale) upper fence of 119.25.
  1. Part (b) — Step 4: quartile (Bowley) skew. $$\text{Quartile skew} = \frac{Q_3+Q_1-2(\text{Median})}{Q_3-Q_1} = \frac{52.5+8.0-2(16.0)}{44.5} = \frac{28.5}{44.5} = 0.640$$
  2. Step 5 — Pearson's coefficient of skewness. Using the given mean, median and standard deviation, $$\boxed{\text{Coefficient of skewness} = \frac{3(\text{Mean}-\text{Median})}{StDev} = \frac{3(34.3-16.0)}{36.4} = \frac{54.9}{36.4} = 1.508}$$
Final Results — Question 5
QuantityValue
$Q_1$8.0 mg/L
$Q_3$52.5 mg/L
IQR44.5 mg/L
Lower / upper fence−58.75 / 119.25 mg/L
Outlier(s)120 mg/L (only)
Whisker endpoints2.0 – 89.0 mg/L
Quartile (Bowley) skew0.640
Pearson coefficient of skewness1.508

Data characteristics (Part b, discussion). The numerical summary and boxplot together describe a strongly right-skewed (positively skewed) distribution typical of environmental contaminant concentrations: the mean (34.3) is more than twice the median (16.0), the box itself is asymmetric ($Q_1$ to median spans only 8, median to $Q_3$ spans 36.5), and a single high value (120) sits well beyond the upper fence as a genuine outlier while there is no low-side outlier at all. Both skewness measures agree with this picture and with each other in sign and rough magnitude: a quartile skew of 0.64 and a Pearson coefficient of 1.51 both indicate pronounced positive skew (values of roughly 0 indicate symmetry; here both are solidly positive). This shape is consistent with most sampling locations having low background nitrate with one or two locations or events showing markedly elevated concentrations — exactly the "hotspot" behaviour flagged in Question 1(b) as typical of environmental data, and a strong indication that any further inference on this dataset (e.g. comparing to a regulatory limit) should use a nonparametric or lognormal-based method rather than assuming normality.