NivaarExam PrepOfficial exam papers ↗

16-Civ-B7 Transportation Planning and Engineering · May 2018

Question 3 of 7: Trip Production by Cross-Classification and by Regression

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

Paper format. National Examination, May 2018 — 16-Civ-B7, Transportation Planning and Engineering. Three hours, closed book (one two-sided aid sheet permitted; Casio or Sharp approved calculator). Seven questions of 20 marks each; any five constitute a complete examination, and only the first five as they appear in the answer book are marked. The per-sub-question mark split is printed on page 7 and is reproduced beside each part below. All seven questions are solved here, because the complete set is the more useful study resource.

Reference texts for this subject.

  • Papacostas, C. S. and Prevedouros, P. D., Transportation Engineering and Planning, 3rd ed. — the four-step model, trip generation, deterministic queueing, traffic-flow theory.
  • Ortúzar, J. de D. and Willumsen, L. G., Modelling Transport, 4th ed. — trip distribution, the gravity model, discrete choice, equilibrium assignment.
  • Meyer, M. D. and Miller, E. J., Urban Transportation Planning: A Decision-Oriented Approach, 2nd ed. — land use and transport, travel-demand management.
  • Garber, N. J. and Hoel, L. A., Traffic and Highway Engineering, 5th ed. — shock waves, signalised-intersection delay.
  • Transportation Research Board, Highway Capacity Manual (HCM), 6th ed. — capacity, control delay and level of service.
  • Transportation Association of Canada, Geometric Design Guide for Canadian Roads — the Canadian design frame for the network context of these questions.

Question 3: Trip Production by Cross-Classification and by Regression (20 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

Given. A base-year survey of 1,144 households making 8,074 trips, cross-classified by household size and income, together with a target-year forecast of 1,770 households in the same 15 cells.

Base-year survey: households and trips by household size and income
Persons per householdLow incomeMedium incomeHigh income
HouseholdsTripsHouseholdsTripsHouseholdsTrips
19322214961696360
27234113885327205
359417125102533381
41201010109118640471
5 or more131073745733423
Target-year households in the study area
Persons per householdLow incomeMedium incomeHigh income
1120280130
210022040
39019050
415018070
5 or more306060

Find. The target-year trips in each of the 15 cells by cross-classification (a) and by the fitted regression equation (b), and a comparison of the assumptions and limitations of the two methods (c).

Approach. Both methods produce a trip rate per household for every cell and then multiply by the forecast household count; they differ only in where the rate comes from — the observed cell mean in (a), a fitted linear function in (b).

(a) Cross-classification (category analysis)

  1. Convert the survey into an observed trip rate for every cell. The rate is the cell’s trips divided by its households, \[r_{mn} = \frac{\text{trips}_{mn}}{\text{households}_{mn}},\qquad \text{e.g.}\quad r_{4,\text{low}} = \frac{1010}{120} = 8.42\ \text{trips/household}.\]
Step 1 — observed trip rates (trips per household per day)
Persons per householdLow incomeMedium incomeHigh income
12.394.133.75
24.746.187.59
37.078.2011.55
48.4210.8811.78
5 or more8.2312.3512.82

The rate surface behaves exactly as trip-generation theory predicts: it rises strongly with household size in every income column and rises with income at every household size, and it flattens at the top of the size range because the fifth and subsequent members of a household are disproportionately children.

  1. Apply each observed rate to the forecast households in the same cell. \[T_{mn} = r_{mn} \times H_{mn},\qquad \text{e.g.}\quad T_{4,\text{low}} = 8.4167 \times 150 = 1{,}262.5\ \text{trips}.\]
Step 2 — forecast target-year trips by cell, method (a)
Persons per householdLow incomeMedium incomeHigh income
1286.51,157.6487.5
2473.61,359.9303.7
3636.11,558.0577.3
41,262.51,958.5824.2
5 or more246.9741.1769.1
Column total2,905.66,775.12,961.8
  1. Sum the 15 cells for the zonal production. \[\boxed{T_{(a)} = 2{,}905.6 + 6{,}775.1 + 2{,}961.8 = 12{,}642.5 \approx 12{,}643\ \text{trips/day}}\] The implied average is \(12{,}642.5 / 1770 = 7.14\) trips per household, up from the base-year 7.06, because the forecast shifts households toward the larger and higher-income cells.

(b) Linear regression rate

  1. Evaluate the fitted equation for each of the 15 household types. With NPERSON capped at 5 and HINCOME coded 0, 1, 2, \[r = 0.46 + 1.96\,\text{NPERSON} + 1.66\,\text{HINCOME},\qquad \text{e.g.}\quad r_{4,\text{low}} = 0.46 + 1.96(4) + 1.66(0) = 8.30.\] Every additional person adds a constant 1.96 trips and every step up the income ladder adds a constant 1.66 trips, which is the substantive assumption the model makes.
Step 1 — regression trip rates (trips per household per day)
Persons per householdLow income (H = 0)Medium income (H = 1)High income (H = 2)
12.424.085.74
24.386.047.70
36.348.009.66
48.309.9611.62
5 or more10.2611.9213.58
  1. Apply the fitted rates to the same forecast households and sum.
Step 2 — forecast target-year trips by cell, method (b)
Persons per householdLow incomeMedium incomeHigh income
1290.41,142.4746.2
2438.01,328.8308.0
3570.61,520.0483.0
41,245.01,792.8813.4
5 or more307.8715.2814.8
Column total2,851.86,499.23,165.4
  1. Total the regression forecast. \[\boxed{T_{(b)} = 2{,}851.8 + 6{,}499.2 + 3{,}165.4 = 12{,}516.4 \approx 12{,}516\ \text{trips/day}}\] The two methods differ by only 126.1 trips, about 1.0 per cent of the total, which is the expected outcome when a well-specified regression is fitted to the same data that the cross-classification tabulates.
035007000105001400029062852Low income67756499Medium income29623165High income1264212516All households(a) cross-classification(b) regressionForecast target-year trips by income group
Target-year trip production by income group under the two methods. The totals agree to about one per cent, but the regression redistributes trips out of the low- and medium-income groups into the high-income group because it forces a constant income increment.

(c) Assumptions and limitations compared

Cross-classification assumes that the observed trip rate in each cell is stable over time and transferable to the forecast year, and that household size and income between them capture everything that matters — but it imposes no functional form at all. That is its strength: the rate surface can be as irregular as the data, and the saturation at large household sizes visible in the table is reproduced automatically. Its limitations are equally clear. It has no statistical machinery, so there is no confidence interval and no significance test. It is extremely sensitive to thin cells: the five-or-more, low-income cell rests on 13 households, and the two-person, high-income cell on 27, so a handful of unusual households sets the rate for a whole forecast category. It cannot fill an empty cell at all, and it cannot extrapolate to a household type outside the sampled range. Adding a third explanatory variable multiplies the number of cells and makes the thin-cell problem worse.

Regression assumes a specific functional form — additive, linear and with no interaction between size and income — in exchange for statistical efficiency. Because it pools all 1,144 households to estimate three coefficients, it is far more stable in thin cells, it can fill an empty cell, and it comes with standard errors, t-statistics and an \(R^2\). Its limitations are the mirror image of those advantages. The linear form is imposed rather than tested, so the saturation the data actually show is lost: the equation keeps adding 1.96 trips for every person, and the analyst has to patch this by hand with the “NPERSON = 5 for 5 or more” cap the question supplies. The no-interaction assumption is visibly violated here — in the survey the income effect is much larger for small households than for large ones. The two worst cells make the point: for one-person high-income households the regression gives 5.74 against an observed 3.75, an error of 1.99 trips per household, while for five-or-more medium-income households it gives 11.92 against an observed 12.35. Treating the ordinal income code as a cardinal variable is a further imposition, since it forces the low-to-medium and medium-to-high steps to be identical.

In practice the two are complements rather than rivals, and the standard Canadian regional practice is to use them together: cross-classification for the well-populated cells where the data speak for themselves, and a fitted equation to fill sparse or empty cells and to extrapolate to household types the survey did not reach. The one per cent agreement between the two totals here is reassurance that neither is badly misspecified.

Final results — Question 3
QuantityMethod (a) cross-classificationMethod (b) regression
Low-income trips2,905.62,851.8
Medium-income trips6,775.16,499.2
High-income trips2,961.83,165.4
Total target-year trips12,642.512,516.4
Implied trips per household7.147.07
Difference between the methods126.1 trips, about 1.0 per cent