NivaarExam PrepOfficial exam papers ↗

16-Civ-B7 Transportation Planning and Engineering · December 2019

Question 3 of 7: Trip Production — Cross-Classification versus Regression

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

Paper format. National Examination, December 2019 — 16-Civ-B7, Transportation Planning & Engineering. Three hours. Closed book; one two-sided aid sheet is permitted, plus an approved Casio or Sharp calculator. Seven questions, all of equal value (20 marks); any five constitute a complete examination and only the first five that appear in the answer book are marked. The mark split is printed as a per-sub-question table on the last page and is reproduced in each heading below. All seven questions are solved here, because the set is a study resource.

Reference texts.

Question 3: Trip Production — Cross-Classification versus Regression (20 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

Given. An observed cross-classification table of trip rates (trips per household per day) and a forecast table of household counts for the target year, both indexed by household size (1 to 5+) and vehicle ownership (0, 1, 2+).

Observed trip rates (trips/household) and forecast households in the target year
Household sizeTrip rate, vehicles ownedHouseholds, vehicles owned
012+012+
12.64.04.0100300150
24.86.78.211025050
37.49.211.29025050
49.211.514.715021060
5+11.213.717.2205030

The forecast household total is 1870 households. The fitted regression is \(r = -0.85 + 2.6289\,\text{HSIZE} + 2.0115\,\text{VEH}\), with HSIZE capped at 5 and VEH capped at 2.

Find. The forecast daily trip production in each of the fifteen cells by (a) the observed cross-classification rates and (b) the fitted regression rates, together with the totals, and a comparison of the assumptions and limitations of the two methods.

Approach. Trip production in a category model is simply the product of the category's trip rate and the number of households forecast in that category; the two parts differ only in where the rate comes from, so the same fifteen multiplications are performed twice and the totals compared.

  1. State the category (cross-classification) production model. For household size \(h\) and vehicle-ownership class \(v\), $$P_{hv}=r_{hv}\,N_{hv}, \qquad P_{\text{total}}=\sum_h\sum_v r_{hv}N_{hv}$$ where \(r_{hv}\) is the observed trip rate and \(N_{hv}\) the forecast number of households. The model's entire content is the assumption that the observed rate is stable over time.
  2. (a) Apply the observed rates cell by cell. For example the three-person / one-vehicle cell gives \(9.2 \times 250 = 2300\) trips/day, and the four-person / zero-vehicle cell gives \(9.2 \times 150 = 1380\) trips/day. Carrying out all fifteen products:
    (a) Forecast trips/day using the observed cross-classification rates
    Household size0 vehicles1 vehicle2+ vehiclesRow total
    1260.01200.0600.02060.0
    2528.01675.0410.02613.0
    3666.02300.0560.03526.0
    41380.02415.0882.04677.0
    5+224.0685.0516.01425.0
    Column total3058.08275.02968.014 301.0
    Summing the row totals independently of the column totals as a check, \(2060+2613+3526+4677+1425 = 14\,301\), so $$\boxed{P_{\text{total}}^{(a)}=14\,301 \text{ trips/day}}$$ which is an average of \(14\,301/1870 = 7.65\) trips per household per day.
  3. (b) Evaluate the regression rate for each cell. The equation is linear and additive, so each rate depends only on the two capped explanatory variables. For the three-person / one-vehicle cell, $$r=-0.85+2.6289(3)+2.0115(1)=-0.85+7.8867+2.0115=9.0482 \text{ trips/hh}$$ against the observed 9.2 — a close fit. Note that the 2+ column takes VEH = 2 and the 5+ row takes HSIZE = 5, exactly as the question directs. The full fitted rate table is:
    (b) Regression trip rates (trips/household)
    Household size0 vehicles1 vehicle2+ vehicles
    11.77893.79045.8019
    24.40786.41938.4308
    37.03679.048211.0597
    49.665611.677113.6886
    5+12.294514.306016.3175
  4. Multiply the fitted rates by the same household forecast. Applying \(P_{hv}=r_{hv}N_{hv}\) with the regression rates:
    (b) Forecast trips/day using the regression rates
    Household size0 vehicles1 vehicle2+ vehiclesRow total
    1177.91137.1870.32185.3
    2484.91604.8421.52511.2
    3633.32262.0553.03448.3
    41449.82452.2821.34723.3
    5+245.9715.3489.51450.7
    Column total2991.88171.53155.714 318.9
    $$\boxed{P_{\text{total}}^{(b)}=14\,319 \text{ trips/day}}$$
  5. Compare the two totals. The regression forecast exceeds the cross-classification forecast by only \(14\,318.9-14\,301.0 = 17.9\) trips/day, i.e. 0.13 per cent — the regression reproduces the aggregate almost exactly, which is what a least-squares fit to these very data should do. The agreement is far worse cell by cell: the one-person / two-vehicle cell moves from 4.00 to 5.80 trips per household (+45 per cent, or +270 trips/day), because the observed table shows the one-person row flattening at 4.0 trips once a second vehicle is added while the additive regression keeps charging a further 2.0115 trips for it.

(c) Comparison of assumptions and limitations

Cross-classification (category analysis). Its only assumption is that the observed trip rate for a category is stable over time and transferable to the forecast year; it imposes no functional form at all. That is its strength: it captures any pattern present in the data, including the non-linearity visible in row 1, where the rate saturates at 4.0 trips regardless of whether the household owns one vehicle or two. Its limitations are that it needs a large survey sample to populate every cell reliably — the 5+ / 0-vehicle cell here represents only 20 households — that it cannot produce a rate for a category the survey never observed, that it cannot interpolate or extrapolate, and that it offers no measure of statistical significance or goodness of fit. Adding a third variable multiplies the number of cells and makes the sample-size problem worse.

Regression. Its assumptions are far stronger: that the relationship is linear and additive in HSIZE and VEH (no interaction), that the coefficients are stable over time, and the usual least-squares conditions on the residuals. In exchange it is parsimonious — three parameters replace fifteen cell rates — it smooths sampling noise, it can be evaluated for any combination of the explanatory variables including unobserved ones, and it comes with standard errors and an \(R^2\). Its limitations follow directly from the linear form. The intercept is negative (\(-0.85\)), so the model predicts a physically meaningless result outside the fitted range and understates the smallest households: it gives 1.78 trips/day for a one-person household with no vehicle against the observed 2.6. It cannot represent saturation or interaction, which is why it overstates the one-person / two-vehicle cell by 45 per cent. And the caps on HSIZE and VEH are an admission of exactly this weakness — they are a crude device to stop a linear function running away at the top of its range.

Practical recommendation. Because the two totals agree to 0.13 per cent but individual cells differ by up to 45 per cent, the choice matters only if the forecast depends on the composition of households rather than their total. Where the survey sample is adequate, use cross-classification for its freedom from functional-form error; where cells are thin or a policy variable such as income must be added, use the regression, but check its residuals cell by cell first and consider adding an interaction term.

Question 3 — final results
QuantityValue
(a) Total forecast trips, observed cross-classification rates14 301 trips/day
(b) Total forecast trips, fitted regression rates14 319 trips/day
Difference (b) − (a)+17.9 trips/day (+0.13 %)
Total forecast households1870
Average trip rate, method (a)7.65 trips/household/day
Largest cell disagreement (1 person, 2+ vehicles)600.0 vs 870.3 trips/day (+45 %)