20-Bio-B4 Robotics · May 2015
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Paper format: National Exams, May 2015 — 04-Bio-B4 Image Processing. Three hours, open book (any paper notes or textbooks permitted, but no calculator or computer). Six questions of equal value (20 marks each); five constitute a complete paper and only the first five appearing in the answer book are marked. All six are solved here, because this set is a study resource rather than an examination script. Every question is essay/descriptive (definitions, algorithm design, system design); the only quantitative content is the computational-complexity discussion in Question 3(f)/(g).
Reference texts (the books a candidate should have reviewed for this subject):
Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.
Lossy compression reduces file size by discarding information the human visual system is least sensitive to — typically by quantizing high-frequency DCT or wavelet coefficients more coarsely than low-frequency ones — so the reconstructed image is a close but not bit-exact copy of the original. Because discarded detail cannot be recovered, compression ratio is traded directly against visible artifact (blocking, ringing). It is used wherever storage/bandwidth matters more than exact-pixel fidelity: JPEG for photographs, MPEG/H.264/HEVC for video, and most consumer web imagery.
Anisotropic diffusion (Perona-Malik) smooths an image by solving a diffusion PDE whose local diffusion coefficient is a decreasing function of the local gradient magnitude, so smoothing proceeds freely inside homogeneous regions but is inhibited across strong edges — unlike isotropic (Gaussian) smoothing, which blurs edges and interiors equally. It is used for edge-preserving denoising, especially in medical imaging (MRI/ultrasound speckle reduction) and as a pre-processing step before edge detection or segmentation, where blurred boundaries would otherwise degrade the result.
YUV represents a colour image as one luma channel Y (a weighted brightness, matching perceived intensity) and two chroma-difference channels U and V, decorrelating brightness from colour information. Because the human eye resolves luma detail far better than chroma detail, the chroma channels can be spatially sub-sampled (e.g. 4:2:0) with little perceptible loss. It is used throughout video compression and broadcast (analogue TV, MPEG/H.264, JPEG) specifically to exploit that chroma sub-sampling opportunity.
Image inpainting reconstructs a missing or damaged region of an image using information from the surrounding, intact pixels — approaches range from PDE-based structure propagation (extending isophote lines into the gap) to exemplar-based patch copying from elsewhere in the image, to modern learned (deep generative) models. It is used for photograph restoration (removing scratches, tears, or date stamps), for removing unwanted objects/watermarks from an image, and for filling occluded regions in film restoration.
A Laplacian pyramid is a compact multiresolution representation built from a Gaussian pyramid: each level $L_i$ is the band-pass detail lost between the Gaussian level $G_i$ and a low-pass, downsampled-then-re-expanded version of the next coarser level, $L_i = G_i - \text{expand}(G_{i+1})$. Because this is a pure algebraic subtraction, summing the expanded coarsest residual back up through every $L_i$ reconstructs the original image exactly. It is used for multi-band image blending (Burt-Adelson seamless compositing at each frequency band separately), for progressive/coarse-to-fine transmission, and for coarse-to-fine (pyramidal) motion estimation.
Image segmentation partitions an image's pixel domain into regions that each correspond to a distinct object, surface, or material, based on a homogeneity criterion (intensity, colour, texture) or a discontinuity criterion (edges/boundaries). It is used wherever downstream analysis needs "what pixels belong to what" rather than raw pixel values — medical image analysis (delineating an organ or tumour), industrial inspection, and as a pre-processing step for object recognition — and is treated in more depth in Question 2.
SIFT and SURF detect distinctive local keypoints at extrema of a scale-space representation (a difference-of-Gaussians pyramid for SIFT), then describe each keypoint with a histogram of local gradient orientations normalized to the keypoint's own dominant orientation and characteristic scale — making the resulting descriptor invariant to image scale, in-plane rotation, and substantially robust to illumination change and modest affine distortion. They are used for matching the same physical point across different views of a scene: panorama/mosaic stitching, wide-baseline stereo, object recognition, and visual SLAM/odometry, and underpin the matching-based CBIR strategy discussed in Question 4.