NivaarExam PrepOfficial exam papers ↗

16-Civ-A6 Highway Design, Construction, and Maintenance · December 2016

Question 3 of 7: Trip Generation by Cross-Classification and Regression

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

Paper format. National Examination, December 2016 — 98-Civ-A6, Transportation Planning & Engineering. Three hours, closed book (one two-sided aid sheet, approved calculator). Seven questions of equal value (20 marks each); any five constitute a complete examination. All seven are solved here, because the set is a study resource rather than a sat examination.

Reference texts. Mannering, Washburn & Kilareski, Principles of Highway Engineering and Traffic Analysis (Wiley) — traffic-stream models, deterministic queueing and shock waves; Papacostas & Prevedouros, Transportation Engineering and Planning (Prentice Hall) — the land-use/transport system and the four-step demand model; Ortuzar & Willumsen, Modelling Transport (Wiley) — trip generation, distribution, mode choice and traffic assignment; Sheffi, Urban Transportation Networks (Prentice Hall) — user-equilibrium assignment; Transportation Association of Canada, Geometric Design Guide for Canadian Roads — Canadian planning and design practice.

Question 3: Trip Generation by Cross-Classification and Regression (20 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

Given. A 5 × 3 category table of surveyed trip rates, the matching 5 × 3 table of forecast household counts, and a regression equation calibrated on the same survey.

Given data — surveyed trip rates $r_{pv}$ (trips/household) and forecast households $N_{pv}$
Persons/hhTrip rate, vehicles/hhHouseholds, vehicles/hh
012+012+
12.64.04.0100300150
24.86.78.211025050
37.49.211.29025050
49.211.514.715021060
5+11.213.717.2205030

Find. Zonal trip production cell by cell and in total, first from the category rates and then from the regression equation, and a comparison of the two methods.

Approach. Both methods are the same product summed over the same 15 cells, $P=\sum_{p}\sum_{v} N_{pv}\,r_{pv}$; only the source of $r_{pv}$ changes — the observed cell mean in (a), the fitted linear surface in (b).

  1. Set up the cross-classification product. The category (cross-classification) method assumes the surveyed rate for a household class transfers unchanged to the forecast year, so each cell contributes $$T_{pv}=N_{pv}\times r_{pv}$$ For the one-person, one-vehicle cell, $T=300\times4.0=1200$ trips; for the five-or-more-person, two-or-more-vehicle cell, $T=30\times17.2=516$ trips. Applying this to all fifteen cells gives the table below.
  2. Sum the cross-classification cells. Adding the fifteen products, checked by rows and again by columns: $$P_a=\sum_{p}\sum_{v}N_{pv}r_{pv}=3058+8275+2968$$ $$\boxed{P_a=14{,}301\ \text{trips per day}}$$
(a) Cross-classification forecast, $T_{pv}=N_{pv}r_{pv}$ (trips)
Persons/hh0 veh1 veh2+ vehRow total
12601,2006002,060
25281,6754102,613
36662,3005603,526
41,3802,4158824,677
5+2246855161,425
Column total3,0588,2752,96814,301
  1. Evaluate the regression rates. The fitted equation is applied at the class midpoints, honouring both caps ($\text{NPERSON}\le5$, $\text{NVEH}\le2$): $$\hat r = -0.85 + 2.6289\,\text{NPERSON} + 2.0115\,\text{NVEH}$$ For a one-person, zero-vehicle household, $\hat r=-0.85+2.6289(1)+0=1.7789$ trips; for three persons and one vehicle, $\hat r=-0.85+2.6289(3)+2.0115(1)=9.0482$; for the capped top class, $\hat r=-0.85+2.6289(5)+2.0115(2)=16.3175$. The full rate surface is regular by construction: every step of one person adds 2.6289 trips and every step of one vehicle adds 2.0115 trips.
  2. Multiply through and total. Repeating step 1 with the fitted rates, $$P_b=\sum_{p}\sum_{v}N_{pv}\hat r_{pv}=2991.78+8171.49+3155.65$$ $$\boxed{P_b=14{,}318.9\approx14{,}319\ \text{trips per day}}$$
  3. Compare the two totals. The difference is $$\Delta=\frac{P_b-P_a}{P_a}=\frac{14{,}318.9-14{,}301}{14{,}301}=+0.13\%$$ The two methods agree to within one part in eight hundred at the zonal total, which is the expected outcome: the regression was calibrated on exactly this survey, so its residuals nearly cancel when weighted by the household distribution. The cell-level agreement is much weaker — the one-person, two-or-more-vehicle cell is 870 trips by regression against 600 by category, a 45 % overstatement — and that is where the methodological difference bites.
(b) Regression forecast: fitted rates $\hat r_{pv}$ and trips $T_{pv}=N_{pv}\hat r_{pv}$
Persons/hhFitted rate (trips/hh)TripsRow total
012+0 veh1 veh2+ veh
11.77893.79045.8019177.891,137.12870.282,185.29
24.40786.41938.4308484.861,604.83421.542,511.22
37.03679.048211.0597633.302,262.05552.993,448.34
49.665611.677113.68861,449.842,452.19821.324,723.35
5+12.294514.306016.3175245.89715.30489.521,450.71
Column total———2,991.788,171.493,155.6514,318.92

(c) Comparison of assumptions and limitations

The cross-classification method assumes only that households in a given category behave alike and that the observed category rate is stable over the forecast horizon. It imposes no functional form, so it reproduces non-linearity and saturation automatically — and this survey contains a clear saturation effect that the regression cannot see: at one person per household the rate stops rising at 4.0 trips whether the household owns one vehicle or two, because a single person can only make so many trips. It also handles interaction between the two variables, since the vehicle effect is allowed to differ by household size (1.4 trips for a one-person household, 6.0 for a five-person household). Its costs are data-hungriness and fragility: fifteen cells must each be populated by enough surveyed households to give a stable mean, cells with few or no observations cannot be estimated at all, there is no way to interpolate to a category the survey missed, and the method offers no measure of statistical significance or goodness of fit.

The regression method assumes a specific functional form — here strictly linear and additive, with no interaction term — and in exchange it uses all 15 cells to estimate only three parameters. That makes it far more economical of data, lets it fill empty cells, allows extrapolation beyond the surveyed range, adds more explanatory variables cheaply, and provides standard errors and an $R^2$. Its limitations are the mirror image of those strengths. The additive form forces every vehicle to add the same 2.0115 trips regardless of household size, so it overstates the small-household, high-ownership cells (870 against 600 trips) and cannot represent saturation at all. The intercept is negative ($-0.85$ trips), which is physically meaningless and a standing warning that the fitted surface must not be used far from the calibration range. And because the aggregate agreement is so good while the cell agreement is not, a modeller who validates only on the zonal total will never discover the specification error. The two capping rules in the question, $\text{NPERSON}\le5$ and $\text{NVEH}\le2$, exist precisely to stop the linear surface from running away in the upper tail; they are a patch on the functional form, not a property of it.

In practice the choice turns on data availability and on how the forecast will be used. Where a large household-travel survey exists and the zone structure is stable, cross-classification (in its modern guise as a discrete-choice or category-rate model within an activity-based framework) is preferred for its fidelity. Where the survey is small, or where the model must respond to a policy variable such as income or transit accessibility, regression is the practical choice — provided the specification is tested for interaction and the intercept is not interpreted literally.

Final results — Question 3
QuantityValue
(a) Zonal trips, cross-classification14,301 trips/day
(b) Zonal trips, regression14,318.9 ≈ 14,319 trips/day
Difference, (b) relative to (a)+17.9 trips = +0.13 %
Largest cell discrepancy1 person / 2+ veh: 870.3 vs 600 (+45 %)
Regression rate range1.7789 to 16.3175 trips/hh