NivaarExam PrepOfficial exam papers ↗

16-Civ-B7 Transportation Planning and Engineering · May 2017

Question 3 of 7: Trip Production — Cross-Classification and Regression

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

Paper format. National Examination, May 2017 — 16-Civ-B7 Transportation Planning & Engineering. Three hours; closed book with one two-sided aid sheet; seven questions of equal value (20 marks each), of which any five constitute a complete examination. All seven are solved here, because the set is a study resource rather than an exam script.

Reference texts. Garber & Hoel, Traffic and Highway Engineering, 5th ed. (queueing, shock waves, traffic flow theory); Papacostas & Prevedouros, Transportation Engineering and Planning, 3rd ed. (the four-step model); Ortuzar & Willumsen, Modelling Transport, 4th ed. (trip distribution, discrete choice, assignment); Ben-Akiva & Lerman, Discrete Choice Analysis (logit and the IIA property); Meyer & Miller, Urban Transportation Planning, 2nd ed. (land use interaction, travel demand management); Transportation Association of Canada, Geometric Design Guide for Canadian Roads. Canadian practice is assumed throughout: travel-demand work in Canada is done under provincial and regional model frameworks, and this paper is written in SI units.

Note on the paper. The 16-Civ-B7 paper examined here is Transportation Planning & Engineering: travel-demand forecasting, traffic flow theory, discrete choice and network assignment. No pavement, materials or geometric-design question appears.

Question 3: Trip Production — Cross-Classification and Regression (20 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

Given. A base-year survey of households cross-classified by size and income, with the trips each cell produces, together with a forecast of how many households will occupy each cell in the target year.

Given data — base-year survey (households and trips produced)
Income levelSize 1Size 2Size 3+
No. of HHNo. of tripsNo. of HHNo. of tripsNo. of HHNo. of trips
Low500122045013005001950
Medium600186070029508003700
High500212580045007503600
Given data — forecast households in the target year
Income levelSize 1Size 2Size 3+
Low356947
Medium508329
High712316

Find. The trips produced by each of the nine household types in the target year, first from the observed cross-classification rates and then from the fitted regression equation, and a comparison of the assumptions the two methods rest on.

Approach. Convert the base-year survey into a trip rate per household for each cell, apply those rates to the forecast household counts, then repeat with rates generated by the regression equation and compare the two sets of totals.

(a) Cross-classification (category analysis)

  1. Reduce the survey to trip rates. The rate for a cell is simply the trips it produced divided by the households in it, $$r_{ij} = \frac{T^{\text{base}}_{ij}}{H^{\text{base}}_{ij}}$$ For low-income single-person households, $r = 1220/500 = 2.44$ trips per household per day. Applying the same division to all nine cells:
Step 1 — base-year trip rates (trips per household per day)
Income levelSize 1Size 2Size 3+
Low2.4402.8893.900
Medium3.1004.2144.625
High4.2505.6254.800

The rates behave sensibly — they rise with household size and, at a given size, with income — with one exception. The high-income, three-or-more-person cell produces 4.800 trips per household, less than the 5.625 of the high-income two-person cell, breaking the otherwise monotone pattern. This is a real feature of the survey and not an arithmetic slip, and it is worth noting because it is precisely the kind of irregularity that part (b)'s regression will smooth away.

  1. Apply the rates to the forecast households. The trips produced by each cell in the target year are $T_{ij} = H^{\text{forecast}}_{ij}\, r_{ij}$. For the low-income single-person cell, $35 \times 2.44 = 85.4$ trips per day, and so on across the matrix.
Result (a) — forecast trips produced by household type
Income levelSize 1Size 2Size 3+Row total
Low85.4199.3183.3468.0
Medium155.0349.8134.1638.9
High301.8129.476.8507.9
Column total542.1678.5394.21614.9

Summing the nine cells,

$$\boxed{\text{total trips produced (cross-classification)} \approx 1615 \ \text{trips/day}}$$

(b) Regression-based trip rates

  1. Evaluate the fitted equation for each household type. With INCOME coded 1, 2, 3 for low, medium and high, and SIZE capped at 3, $$r = 0.99 + 0.91\,\text{INCOME} + 0.59\,\text{SIZE}$$ For the low-income single-person cell, $r = 0.99 + 0.91(1) + 0.59(1) = 2.49$ trips per household. Because the equation is linear and additive, each step up in income adds a flat 0.91 trips and each additional person adds a flat 0.59 trips, so the whole rate matrix follows immediately.
Step 3 — regression trip rates (trips per household per day)
Income levelSize 1Size 2Size 3+
Low2.493.083.67
Medium3.403.994.58
High4.314.905.49
  1. Apply the regression rates to the same forecast households. Multiplying cell by cell, for example $35 \times 2.49 = 87.2$ trips for the low-income single-person cell and $71 \times 4.31 = 306.0$ for the high-income single-person cell, gives the second forecast matrix.
Result (b) — forecast trips produced using the regression rates
Income levelSize 1Size 2Size 3+Row total
Low87.2212.5172.5472.2
Medium170.0331.2132.8634.0
High306.0112.787.8506.6
Column total563.2656.4393.21612.7
$$\boxed{\text{total trips produced (regression)} \approx 1613 \ \text{trips/day}}$$

The two regional totals agree to within 2.2 trips per day, about 0.13 per cent, which is a much closer agreement than the individual cells show. The high-income two-person cell, for instance, drops from 129.4 to 112.7 trips, a 13 per cent difference, because the regression cannot reproduce that cell's unusually high observed rate of 5.625. The errors happen to offset in aggregate here; they would not do so if the forecast household distribution shifted toward the cells the regression fits worst.

(c) Assumptions and limitations compared

Cross-classification makes no assumption whatever about the functional form of the relationship between household attributes and trip making. It simply asserts that the observed rate for a category is stable over time and transferable to the forecast year, and it reproduces the base-year data exactly by construction, including irregularities such as the high-income three-person cell. Its weaknesses follow from the same property: the rate for each cell rests on however many households happened to fall in it, so thinly populated categories carry large sampling error and any irregularity, real or spurious, is carried forward with full weight. It offers no statistical measure of fit, no significance tests and no confidence intervals, it cannot produce a rate for a category that was empty in the survey, and the number of cells grows multiplicatively as further classifying variables such as car ownership are added, so the data requirement escalates quickly.

Regression, by contrast, imposes a specific structure — here linear, additive and with no interaction between income and size — and estimates its parameters from all of the data at once. That structure is exactly its strength and its weakness. Because every observation informs every coefficient, the estimates are statistically efficient, they smooth away sampling noise, they extend to categories that were sparse or absent in the survey, and they come with standard errors, a coefficient of determination and formal significance tests. But the model can only be as good as its form: the additive specification denies any interaction, so it cannot represent a household type whose trip making departs from the linear pattern, which is precisely the case for the high-income two-person and three-or-more-person cells here. Treating income as an ordinal code 1, 2, 3 further imposes the strong assumption that the step from low to medium income has the same effect as the step from medium to high. Extrapolation beyond the range of the estimation data is unsupported, and a regression fitted on aggregate zonal data rather than household data is additionally vulnerable to ecological fallacy.

Both methods share the fundamental limitation of any trip-generation model: they assume that the estimated relationship between household attributes and trip making holds in the target year. Neither responds to changes in accessibility, travel cost, fuel price, transit provision or the growth of telework, so neither can represent the feedback described in Question 1(a). In practice the reasonable choice depends on the data: cross-classification where the survey is large enough to populate every cell adequately and where irregular category behaviour is believed to be real, and regression where the sample is thin, where categories must be interpolated, or where the analyst needs a defensible measure of statistical confidence. Comparing the two, as this question does, is itself good practice, because agreement at the regional total conceals disagreement at the cell level that would matter as soon as those trips are distributed among zones.

Final results — Question 3
QuantityCross-classification (a)Regression (b)
Low income — sizes 1 / 2 / 3+85.4 / 199.3 / 183.387.2 / 212.5 / 172.5
Medium income — sizes 1 / 2 / 3+155.0 / 349.8 / 134.1170.0 / 331.2 / 132.8
High income — sizes 1 / 2 / 3+301.8 / 129.4 / 76.8306.0 / 112.7 / 87.8
Total trips produced per day1614.9 (~1615)1612.7 (~1613)
Difference in regional total2.2 trips/day, or 0.13 per cent