NivaarExam PrepOfficial exam papers ↗

20-Bio-B4 Robotics · December 2016

Question 3 of 6: Video Processing

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

Paper format: National Exams, December 2016 — 04-Bio-B4 Image Processing. Three hours, open book (any paper notes or textbooks permitted, but no calculator or computer). Six questions of equal value (20 marks each); five constitute a complete paper and only the first five appearing in the answer book are marked. All six are solved here, because this set is a study resource rather than an examination script. Every question is essay/descriptive (definitions, algorithm/system design) except the convolution-size arithmetic and FFT/3D-convolution items in Question 2, and the illustrative numeric design examples worked into Questions 5 and 6.

Reference texts (the books a candidate should have reviewed for this subject):

Question 3: Video Processing (20 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

(a) Basic Approaches to Video Compression

Video compression combines intra-frame compression (compress each frame independently, exactly like a still image — block-based DCT + quantization + entropy coding, as in Question 1(a)) with inter-frame compression that exploits the fact that consecutive frames are usually highly similar. The dominant inter-frame strategy is motion-compensated prediction: for each block in the current frame, search a small window of the previous (reference) frame for the best-matching block, encode only the small displacement (the motion vector) plus the small residual difference, and reconstruct by shifting the reference block and adding the residual — since the residual is usually near-zero, this is far cheaper to encode than the raw pixel block. A second, complementary strategy is transform-domain redundancy exploitation across the whole GOP (group of pictures) structure: periodic full (I-)frames anchor the sequence, while predicted (P-)frames reference the past and bidirectionally-predicted (B-)frames reference both past and future frames, so only the true novel content in each frame needs to be encoded at all.

(b) Fundamental Operations Applied to Video

Beyond simply filtering each frame independently (Question 2), video-specific operations exploit the temporal dimension directly: motion estimation/optical flow (recovering the per-pixel or per-block displacement field between frames); object tracking (following a detected object's identity and position across many frames, Question 5); background subtraction/foreground detection (modelling the static or slowly-varying background so that moving foreground objects can be isolated); temporal filtering/stabilization (removing unwanted frame-to-frame jitter, e.g. camera shake, while preserving intentional motion); and event/activity detection (recognizing when a specific temporal pattern of appearance/motion has occurred, e.g. Question 6's periodic colour change).

(c) Common Application Areas for Video Analysis

Surveillance and security (intrusion/loitering detection, crowd counting); traffic monitoring (vehicle counting/speed estimation, licence-plate recognition); industrial/manufacturing quality inspection on a moving production line; sports analytics (player/ball tracking, automated highlight generation); autonomous vehicles and robotics (obstacle detection, visual odometry); video conferencing/broadcast (compression, background replacement); and biomedical/clinical video analysis (endoscopy, the cell-tracking and pulse-detection problems of Questions 5 and 6 themselves).

(d) Causal vs. Non-Causal Video Processing

A causal algorithm produces its output for frame $t$ using only frames up to and including $t$ (no future information), which is mandatory for any real-time application — live broadcast, video conferencing, robotics/autonomous control — where a decision or displayed frame cannot wait for data that has not yet been captured. A non-causal algorithm may use future frames as well (e.g. a temporally-centred smoothing/stabilization filter, or Question 2(e)'s centred 3D kernel), which generally gives better quality (more context to smooth noise, resolve ambiguous motion, or interpolate) at the cost of introducing a mandatory latency equal to however many future frames are needed — acceptable for offline editing or file-based (not live) processing, but not for anything with a real-time deadline.

(e) Embedded (On-Camera) vs. Remote (Server) Video Processing

Processing embedded directly on the camera avoids transmitting the full video stream at all, which saves bandwidth (critical for battery/wireless-limited devices) and reduces latency (no network round-trip) and privacy exposure (raw video need never leave the device, only the extracted result, e.g. an alert or a count) — but the camera's on-board compute, memory, and power budget are all far more limited than a server's, constraining the complexity of algorithm that can run in real time. Remote (server-side) processing removes that compute/power ceiling and makes it easy to pool data and update algorithms centrally, but it costs continuous bandwidth to transmit the video, adds network latency and a dependency on connectivity, and raises privacy/security concerns since raw video now leaves the site. The choice in practice is a bandwidth/latency/privacy trade against available on-device compute, and it is common to split the load — light-weight detection/triggering on the camera, heavier analysis sent to a server only when triggered.

QuestionKey answer
(a) Compression approachesIntra-frame (block DCT) + inter-frame motion-compensated prediction with I/P/B frame structure
(b) Fundamental operationsMotion estimation, tracking, background subtraction, temporal filtering/stabilization, event detection
(c) Application areasSurveillance, traffic, manufacturing QA, sports, autonomous systems, biomedical video
(d) Causal vs. non-causalCausal = real-time capable; non-causal = better quality, adds latency = number of future frames needed
(e) Embedded vs. remoteEmbedded: less bandwidth/latency/privacy exposure, less compute; remote: more compute, costs bandwidth/latency/privacy