NivaarExam PrepOfficial exam papers ↗

16-Civ-A6 Highway Design, Construction, and Maintenance · May 2016

Question 3 of 7: Trip Generation by Cross-Classification and by Regression

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

Paper format: 98-Civ-A6 Transportation Planning & Engineering, National Examination May 2016. Seven questions, all of equal value (20 marks); any five constitute a complete examination. Closed book, one two-sided aid sheet permitted. Three hours. All seven questions are solved below as a study resource.

Reference texts. Mannering, Washburn & Kilareski, Principles of Highway Engineering and Traffic Analysis (Wiley) — queueing, shock waves and traffic-stream models; Papacostas & Prevedouros, Transportation Engineering and Planning (Prentice Hall) — the four-step demand model; Ortuzar & Willumsen, Modelling Transport (Wiley) — trip generation, distribution, mode choice and assignment; Roess, Prassas & McShane, Traffic Engineering (Pearson) — signalised-intersection delay.

Question 3: Trip Generation by Cross-Classification and by Regression (20 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

Given. A household survey in one traffic zone, cross-classified by persons per household (1 to 5-or-more) and vehicles per household (0, 1, 2-or-more), together with a forecast of the number of households of each type in the target year.

Survey data: households and trips by household type
Persons/hh0 vehicles1 vehicle2+ vehicles
hhtripshhtripshhtrips
1100220150620100360
27034014086030210
3604201301 03050380
41201 0101101 19040470
5 or more101104046035420
Forecast households in the target year (142 households in total)
Persons/hh0 vehicles1 vehicle2+ vehicles
114220
212164
310125
4797
5 or more6810

Find. The forecast trips in every one of the fifteen household-type cells, first from cross-classification trip rates and then from the supplied regression equation, and a comparison of the assumptions and limitations of the two methods.

Approach. A cross-classification rate is simply the observed trips divided by the observed households in the same cell; the regression rate is evaluated cell by cell from the capped explanatory variables. Both rate matrices are then multiplied cell-wise by the forecast household matrix.

  1. Form the cross-classification trip rates. For each cell, $$ R_{ij} = \frac{\text{trips}_{ij}}{\text{households}_{ij}} $$ For example the one-person, one-vehicle cell gives \(620/150 = 4.133\) trips per household and the four-person, one-vehicle cell gives \(1\,190/110 = 10.818\).
Step 1 — cross-classification trip rates \(R_{ij}\) (trips per household per day)
Persons/hh0 vehicles1 vehicle2+ vehicles
12.2004.1333.600
24.8576.1437.000
37.0007.9237.600
48.41710.81811.750
5 or more11.00011.50012.000
  1. Apply the rates to the forecast households — part (a). The forecast trips in each cell are $$ T_{ij} = R_{ij}\, H_{ij} $$ so, for instance, the three-person one-vehicle cell contributes \(7.923 \times 12 = 95.08\) trips and the five-or-more-person two-vehicle cell contributes \(12.000 \times 10 = 120.0\) trips. The one-person two-vehicle cell contributes nothing because no such households are forecast.
Step 2 — part (a) forecast trips by household type (cross-classification)
Persons/hh0 vehicles1 vehicle2+ vehiclesRow total
130.8090.930.00121.73
258.2998.2928.00184.57
370.0095.0838.00203.08
458.9297.3682.25238.53
5 or more66.0092.00120.00278.00
Column total284.00473.66268.251 025.9

Summing all fifteen cells gives the zonal production forecast by the cross-classification method:

$$ \boxed{T_{(a)} = 1\,025.9 \approx 1\,026\ \text{trips per day}} $$
  1. Evaluate the regression trip rate cell by cell. With the caps applied (NPERSON is held at 5 for the "5 or more" row and NVEH is held at 2 for the "2 or more" column), $$ R^{\text{reg}} = 0.67 + 2.07\,\text{NPERSON} + 0.85\,\text{NVEH} $$ The one-person, zero-vehicle cell therefore gives \(0.67+2.07(1)+0.85(0) = 2.74\), and the five-or-more-person, two-or-more-vehicle cell gives \(0.67+2.07(5)+0.85(2) = 12.72\). Because the equation is linear and additive, every step down a column adds a constant 2.07 trips and every step across a row adds a constant 0.85 trips.
Step 3 — regression trip rates (trips per household per day)
Persons/hh0 vehicles1 vehicle2+ vehicles
12.743.594.44
24.815.666.51
36.887.738.58
48.959.8010.65
5 or more11.0211.8712.72
  1. Apply the regression rates to the same forecast households — part (b). Multiplying cell by cell, for example \(2.74 \times 14 = 38.36\) trips in the one-person zero-vehicle cell and \(12.72 \times 10 = 127.20\) trips in the largest cell, and summing all fifteen products gives $$ \boxed{T_{(b)} = 1\,009.8 \approx 1\,010\ \text{trips per day}} $$ which is 16.1 trips, or 1.6 per cent, below the cross-classification forecast.
Step 4 — part (b) forecast trips by household type (regression)
Persons/hh0 vehicles1 vehicle2+ vehiclesRow total
138.3678.980.00117.34
257.7290.5626.04174.32
368.8092.7642.90204.46
462.6588.2074.55225.40
5 or more66.1294.96127.20288.28
Column total293.65445.46270.691 009.8

(c) Assumptions and limitations of the two methods

The cross-classification (category analysis) method assumes only that the trip rate within a homogeneous category is stable over time, and it imposes no functional form whatever on the relationship between the classifying variables and the trip rate. That is its principal strength: the rate matrix reproduces the survey exactly, including genuinely non-linear and interactive behaviour. It is also transparent, easily understood by non-technical decision makers, and requires no statistical estimation. Its limitations are equally clear. Every cell rate is estimated from that cell alone, so a thinly populated cell yields an unreliable rate — the five-or-more-person zero-vehicle cell here rests on only ten households, and the two-or-more-vehicle column on thirty to a hundred. The method cannot extrapolate: it produces no rate for a category that was not surveyed, and it cannot handle a continuous variable such as income without first discretising it, which introduces arbitrary boundaries. Adding a third classifying variable multiplies the number of cells and quickly exhausts any realistic sample. Finally, there is no statistical measure of goodness of fit and no way to test whether a classifying variable is significant.

The regression method assumes a specific functional form — here strictly linear and additive, with no interaction term between household size and vehicle ownership — together with the usual least-squares assumptions of independent, homoscedastic, normally distributed errors. Its strengths are that it pools all 1 185 surveyed households to estimate only three parameters, so each coefficient is far more stable than an individual cell rate; it can be evaluated for any combination of the explanatory variables, including combinations never observed; it extends naturally to additional continuous variables; and it delivers standard errors, \(t\)-statistics and an \(R^2\) that quantify how well the model fits. Its limitations are that the imposed linearity may be false, that the model can produce nonsensical values outside the range of the calibration data (a household with no persons would be predicted to make 0.67 trips), and that the explanatory variables are typically correlated with one another — larger households tend to own more vehicles — so multicollinearity inflates the standard errors and makes the individual coefficients hard to interpret causally. The caps on NPERSON and NVEH are themselves an admission that the linear form breaks down at the extremes.

The comparison in this zone illustrates the trade-off precisely. The two totals agree closely (1 026 against 1 010, a difference of 1.6 per cent), which is reassuring at the aggregate level, but individual cells differ by much more. The one-person one-vehicle cell falls from 90.9 to 79.0 trips, and the one-person two-vehicle cell in the survey shows a rate of 3.60 that is lower than the 4.13 of the one-vehicle cell beside it — a non-monotonic result that is almost certainly sampling noise in a hundred-household cell, and which the regression, by construction, smooths away. Where the survey is thin, the regression estimate is more trustworthy; where the survey is dense and the behaviour is genuinely non-linear, the cross-classification rate is. In practice a Canadian regional model uses cross-classification for household trip production, precisely because household size and vehicle ownership interact so strongly, and reserves regression for zonal attraction models where the explanatory variables are continuous.

Final results — Question 3
QuantityValue
Forecast households in target year142
Total forecast trips, cross-classification (a)1 025.9 trips/day
Total forecast trips, regression (b)1 009.8 trips/day
Difference (b) relative to (a)−16.1 trips/day (−1.6 %)
Implied average trip rate, method (a)7.22 trips/household/day
Implied average trip rate, method (b)7.11 trips/household/day
Largest single cell, both methods5+ persons / 2+ vehicles: 120.0 vs 127.2 trips