18-Env-B7 Environmental Sampling and Analysis · May 2017
Question 5 of 7: Nitrate Concentrations — Dotplot, Boxplot and Skewness
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Notes on this paper
National Exams, May 2017 — 04-Env-B7, Environmental Sampling and Analysis (3 hours, closed book, approved non-programmable calculator only, statistical tables provided). The paper instructs "answer all 4 questions in Part A and any 2 questions in Part B"; as a study resource this solution answers all 7 questions in full, including all three Part B questions.
Reference texts: Walpole, Myers, Myers & Ye, Probability & Statistics for Engineers and Scientists (sampling designs, hypothesis tests, EDA/boxplots, ANOVA); Davis & Cornwell, Introduction to Environmental Engineering, ch. 2 (sampling protocol, QA/QC, monitoring program design).
Find. A dotplot and boxplot of the data with fences located, $Q_1$/$Q_3$, any outliers (a); a description of the data's characteristics plus the quartile skew and coefficient of skewness (b).
Approach. Sort the raw data to locate the quartiles (Tukey's hinge method: median of the lower half and median of the upper half), compute the IQR-based fences to classify outliers, and combine the shape of the dotplot/boxplot with the given mean/median/StDev to characterize the distribution and its skewness.
Part (a) — Step 1: sort and locate the quartiles. Sorted data: 2, 5, 8, 8, 9, 13, 16, 24, 45, 48, 57, 89, 120 (median = 7th value = 16.0, matching the given summary). Lower half (below the median, 6 values): 2, 5, 8, 8, 9, 13 → $Q_1=$ median of these $=(8+8)/2=8.0$. Upper half (above the median, 6 values): 24, 45, 48, 57, 89, 120 → $Q_3=$ median of these $=(48+57)/2=52.5$.
Step 2 — fences. $IQR=Q_3-Q_1=52.5-8.0=44.5$.
$$\boxed{LF = Q_1-1.5\,IQR = 8.0-66.75=-58.75, \qquad UF = Q_3+1.5\,IQR = 52.5+66.75=119.25}$$
The lower fence is a large negative number, off the physical (concentration ≥ 0) and printed axis range entirely — it still functions as the outlier threshold, it is just never a plotted point.
Step 3 — outliers and whiskers. Comparing every data value to $[LF,UF]=[-58.75,\ 119.25]$: only $120$ exceeds $UF=119.25$, so 120 mg/L is a (mild) outlier; no value is below $LF$. The boxplot whiskers are therefore drawn only to the most extreme non-outlier data values: the lower whisker to the minimum, $2.0$, and the upper whisker to $89.0$ (the largest value that is still $\le UF$), with $120$ plotted as a separate point beyond the upper whisker.
Dotplot (top) and boxplot (bottom) of the 13 nitrate readings, sharing one concentration axis. The box spans $Q_1=8.0$ to $Q_3=52.5$ with the median at 16.0; whiskers extend only to the most extreme non-outlier values (2.0 and 89.0); the value 120 is plotted separately as an outlier, above the (off-scale) upper fence of 119.25.
Step 5 — Pearson's coefficient of skewness. Using the given mean, median and standard deviation,
$$\boxed{\text{Coefficient of skewness} = \frac{3(\text{Mean}-\text{Median})}{StDev} = \frac{3(34.3-16.0)}{36.4} = \frac{54.9}{36.4} = 1.508}$$
Final Results — Question 5
Quantity
Value
$Q_1$
8.0 mg/L
$Q_3$
52.5 mg/L
IQR
44.5 mg/L
Lower / upper fence
−58.75 / 119.25 mg/L
Outlier(s)
120 mg/L (only)
Whisker endpoints
2.0 – 89.0 mg/L
Quartile (Bowley) skew
0.640
Pearson coefficient of skewness
1.508
Data characteristics (Part b, discussion). The numerical summary and boxplot together describe a strongly right-skewed (positively skewed) distribution typical of environmental contaminant concentrations: the mean (34.3) is more than twice the median (16.0), the box itself is asymmetric ($Q_1$ to median spans only 8, median to $Q_3$ spans 36.5), and a single high value (120) sits well beyond the upper fence as a genuine outlier while there is no low-side outlier at all. Both skewness measures agree with this picture and with each other in sign and rough magnitude: a quartile skew of 0.64 and a Pearson coefficient of 1.51 both indicate pronounced positive skew (values of roughly 0 indicate symmetry; here both are solidly positive). This shape is consistent with most sampling locations having low background nitrate with one or two locations or events showing markedly elevated concentrations — exactly the "hotspot" behaviour flagged in Question 1(b) as typical of environmental data, and a strong indication that any further inference on this dataset (e.g. comparing to a regulatory limit) should use a nonparametric or lognormal-based method rather than assuming normality.