NivaarExam PrepOfficial exam papers ↗

16-Civ-B7 Transportation Planning and Engineering · December 2018

Question 3 of 7: Trip Generation by Regression and by Cross-Classification

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

Paper format. National Examination, December 2018 — 16-Civ-B7, Transportation Planning & Engineering. Three hours. Closed book; one 8.5 in × 11 in aid sheet hand-written on both sides is permitted, plus an approved Casio or Sharp calculator. Seven questions, all of equal value (20 marks); any five constitute a complete examination and only the first five that appear in the answer book are marked. All seven are solved here, because the set is a study resource.

Reference texts.

Question 3: Trip Generation by Regression and by Cross-Classification (20 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

Given. The household survey for the three zones, and the two calibrated trip-rate models.

Zonal household survey
ZoneNo. of householdsResidents per household, NRESWorkers per household, NWOR
110,0002.01.0
220,0003.02.0
340,0003.01.0
Cross-classification trip rates (trips per household per day)
Residents per household1 worker or less2 workers or more
2 or less2.12.5
3 or more3.24.1

Regression model: Trip rate = 0.2 + 0.5 NRES + 1.1 NWOR.

Find. Zonal trip productions from each of the two models, the behavioural reading of the regression coefficients, and a comparison of the assumptions, advantages and disadvantages of the two methods.

Approach. Evaluate the trip rate for each zone's average household from each model, multiply by the number of households in the zone, and total; then compare.

  1. Part (a) — evaluate the regression rate zone by zone. The model is applied at the zonal average household, so for zone 1 $$r_1=0.2+0.5(2.0)+1.1(1.0)=0.2+1.0+1.1=2.3\ \text{trips/household/day}$$ and similarly r2 = 0.2 + 0.5(3.0) + 1.1(2.0) = 3.9 and r3 = 0.2 + 0.5(3.0) + 1.1(1.0) = 2.8 trips per household per day.
  2. Expand to zonal productions. Multiplying each rate by the zone's household count gives $$\begin{aligned}T_1 &= 2.3(10{,}000)=23{,}000 \\ T_2 &= 3.9(20{,}000)=78{,}000 \\ T_3 &= 2.8(40{,}000)=112{,}000\end{aligned}$$ so the regression model produces $$\boxed{T_{regression}=23{,}000+78{,}000+112{,}000=213{,}000\ \text{trips/day}}$$
  3. Read the coefficients behaviourally. The model is linear and additive, so each coefficient is a constant marginal effect. An extra resident adds 0.5 trips per day and an extra worker adds 1.1, so a worker generates 2.2 times as many trips as a non-working resident does. That is exactly what one expects: a worker contributes a compulsory round-trip commute (two trips) plus work-based sub-tours, while an additional non-working resident — typically a child or a retired person — makes fewer, and often shares, trips. The intercept of 0.2 is the trip-making of a notional household with no residents and no workers; it has no physical meaning and is simply the regression's degree of freedom, which is a warning that the model must not be extrapolated far outside the range of household types over which it was fitted. Note also that NWOR is bounded above by NRES, so the two variables are collinear and the individual coefficients are less stable than their sum suggests.
  4. Part (b) — assign each zone to a cross-classification cell. The categories are read off the zonal averages. Zone 1 has NRES = 2.0 ("2 or less") and NWOR = 1.0 ("1 or less"), giving 2.1 trips/household. Zone 2 has NRES = 3.0 ("3 or more") and NWOR = 2.0 ("2 or more"), giving 4.1. Zone 3 has NRES = 3.0 ("3 or more") and NWOR = 1.0 ("1 or less"), giving 3.2.
  5. Expand to zonal productions. $$\begin{aligned}T_1 &= 2.1(10{,}000)=21{,}000 \\ T_2 &= 4.1(20{,}000)=82{,}000 \\ T_3 &= 3.2(40{,}000)=128{,}000\end{aligned}$$ $$\boxed{T_{cross\text{-}class}=21{,}000+82{,}000+128{,}000=231{,}000\ \text{trips/day}}$$ The category model produces 18,000 more trips per day, 8.45 per cent above the regression total. The whole of the difference comes from zone 3, the largest zone, where the category rate of 3.2 exceeds the fitted 2.8 by 0.4 trips per household; the regression, being linear in NRES, cannot reproduce the step in trip-making that occurs when a household reaches three or more residents.
Question 3(a) and (b) — zonal trip productions
ZoneHouseholdsRegression rateRegression trips/dayCategory rateCategory trips/day
110,0002.323,0002.121,000
220,0003.978,0004.182,000
340,0002.8112,0003.2128,000
Total70,000—213,000—231,000

Part (c) — the difference in assumptions, with one advantage and one disadvantage of each

The fundamental difference is the functional form imposed on the relationship between household attributes and trip-making. The regression model assumes that trip rate is a continuous, linear and additive function of the explanatory variables: every additional resident adds the same 0.5 trips whether the household is going from one to two people or from six to seven, the effects of residents and workers do not interact, and the model can be evaluated at any real-valued NRES and NWOR, including the non-integer zonal averages used above. The cross-classification model assumes the opposite: trip rate is a categorical, non-parametric function, constant within each cell and free to jump between cells, with the interaction between residents and workers built in because every combination has its own independently estimated rate. In this data set the interaction is real — going from one to two workers is worth 0.4 trips in a small household (2.1 to 2.5) but 0.9 trips in a large one (3.2 to 4.1) — and a linear additive model cannot represent it at all.

A second, subtler difference follows from the first. Because the regression is continuous, it can legitimately be applied to a zonal average household, as done in part (a). The cross-classification model cannot: applying it to the zonal average, as part (b) requires, quietly assumes that every household in the zone falls in the same cell as the average. A correct cross-classification forecast expands each cell by the number of households in that cell, which needs the full joint distribution of household size and workers, not just the two means. The 8.45 per cent gap between the two totals is therefore partly a genuine functional-form difference and partly this aggregation approximation.

Question 3(c) — comparison of the two methods
 Regression (part a)Cross-classification (part b)
AssumptionContinuous, linear, additive; constant marginal effects; no interactionCategorical and non-parametric; rate constant within a cell; interaction fully captured
AdvantageParsimonious and transparent — a handful of coefficients summarise the whole relationship, they can be tested for significance, and the model interpolates and extrapolates to household types that were thin or absent in the survey sampleMakes no assumption about functional form, so it reproduces non-linearities and interactions (such as the jump at three residents) that a linear model must smooth away, and it is easy for a lay audience and a council to understand
DisadvantageImposes a shape the data may not have; collinear explanatory variables (NWOR is bounded by NRES) make the individual coefficients unstable; the intercept has no behavioural meaning and extrapolation is unsafeCell counts grow multiplicatively with the number of variables and categories, so the survey sample per cell becomes thin and the rates noisy; no statistical measure of fit or significance is produced; and forecasting requires the future joint distribution of households across all cells, which is much harder to predict than a few means