20-Bio-B4 Robotics · May 2015
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Paper format: National Exams, May 2015 — 04-Bio-B4 Image Processing. Three hours, open book (any paper notes or textbooks permitted, but no calculator or computer). Six questions of equal value (20 marks each); five constitute a complete paper and only the first five appearing in the answer book are marked. All six are solved here, because this set is a study resource rather than an examination script. Every question is essay/descriptive (definitions, algorithm design, system design); the only quantitative content is the computational-complexity discussion in Question 3(f)/(g).
Reference texts (the books a candidate should have reviewed for this subject):
Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.
The system must separate two distinct sources of frame-to-frame apparent motion — camera jitter (up to 10 px, affects the whole frame) and true object motion (affects only the bounding box) — and must be able to declare, and recover from, tracking failure. The pipeline below has five stages.
Estimate the dominant (background) frame-to-frame motion — e.g. via sparse optical flow or phase correlation on features away from the tracked object — and geometrically register (warp) the current frame to the previous frame's coordinate system before any tracking is attempted. Without this step, the up to 10-pixel jitter would be indistinguishable from true object motion and would corrupt every downstream measurement.
From the pixels inside the given initial bounding rectangle, build a compact appearance representation — a colour histogram (in an illumination-robust space such as HSV) for a lightweight mean-shift-style tracker, or a correlation filter learned over the patch (as in MOSSE/KCF-style trackers) for a fast, real-time-capable onboard tracker. This model defines "what the object looks like" and is what every subsequent frame is compared against.
Maintain a simple constant-velocity Kalman filter over the object's estimated position. Its prediction for the current frame both tightens the region that must be searched (an efficiency gain, since only a neighbourhood around the predicted location needs to be evaluated rather than the whole frame) and supplies a sanity check against which the raw per-frame measurement can later be compared.
Within the predicted search window, evaluate the appearance model against the (jitter-compensated) current frame — e.g. the correlation-filter response surface, or the mean-shift histogram back-projection — and take its peak as the new candidate location; the box extent is updated from the response's scale (searched over a small set of candidate scales) or, if a full descriptor-matching approach is used instead, from the inlier keypoint correspondences and their RANSAC-fit transform. (a) The tracker outputs the updated bounding-box centre, width and height (and orientation, if an affine/homography model is used) as the location/extent description. (b) The confidence measure is derived from the sharpness of the response peak — the peak-to-sidelobe ratio (PSR) is a standard metric here: a sharp, isolated peak (large PSR) means high confidence, while a flat or multi-modal response (small PSR) means the object's appearance is no longer well matched anywhere in the search window. If keypoint matching is used instead, the fraction of RANSAC inlier correspondences serves the same role. The confidence is normalized to $[0,1]$ so that a genuinely lost object naturally drives it toward zero.
When the confidence measure falls below a threshold (heading toward zero — e.g. the object left the frame, was occluded, or changed appearance too much), the system stops trusting the local search window and instead runs a wide-area or full-frame re-detection pass (re-matching the stored appearance model, or an independent object detector, across the entire frame) rather than continuing to search near the last known (now unreliable) location. If re-detection also fails, the system reports zero confidence and waits for a fresh bounding-rectangle initialization, exactly as the question specifies.
Overall flow: Frame$_t$ → stabilize against Frame$_{t-1}$ → Kalman-predicted search window → appearance-model match within window → (location/extent, confidence) → if confidence above threshold: update Kalman filter and (conservatively) refresh the appearance model; else: trigger wide-area re-detection.