20-Bio-B4 Robotics · December 2016
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Paper format: National Exams, December 2016 — 04-Bio-B4 Image Processing. Three hours, open book (any paper notes or textbooks permitted, but no calculator or computer). Six questions of equal value (20 marks each); five constitute a complete paper and only the first five appearing in the answer book are marked. All six are solved here, because this set is a study resource rather than an examination script. Every question is essay/descriptive (definitions, algorithm/system design) except the convolution-size arithmetic and FFT/3D-convolution items in Question 2, and the illustrative numeric design examples worked into Questions 5 and 6.
Reference texts (the books a candidate should have reviewed for this subject):
Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.
The DCT re-expresses a block of image samples as a weighted sum of cosine basis functions of increasing spatial frequency, concentrating most of a typical (smoothly varying) image block's energy into a small number of low-frequency coefficients while using only real-valued coefficients (unlike the complex-valued DFT). It underlies block-based image and video compression — JPEG and the intra-frame coding stage of MPEG/H.26x transform an $8\times8$ pixel block via the 2D DCT, then quantize the resulting coefficients (discarding high-frequency ones more coarsely, since the eye is least sensitive to them) before entropy coding.
Interpolation estimates pixel values at non-integer or newly introduced sample locations from the known surrounding samples, using a kernel such as nearest-neighbour (fastest, blocky), bilinear (weighted average of the 4 nearest samples, smoother), or bicubic (16-sample neighbourhood, sharper edges, some ringing). It is used whenever an image must be resampled onto a different grid — digital zoom/upscaling, geometric warping/registration (Question 4(e)), and rotating or resizing an image for display.
Canny edge detection is a multi-stage pipeline: Gaussian smoothing (noise suppression), gradient magnitude/direction via a Sobel-type operator, non-maximum suppression (thinning the gradient ridge to a single-pixel-wide edge), and finally hysteresis thresholding (a high threshold seeds confirmed edges, a lower threshold lets a confirmed edge extend along a connected weaker ridge, suppressing isolated weak responses). It is used wherever a clean, thin, well-localized edge map is needed as an intermediate representation — object boundary detection, feature extraction for later matching, and as a pre-processing step for higher-level segmentation.
SIFT and SURF detect distinctive keypoints at blob-like local extrema found across a scale-space pyramid (SIFT: difference-of-Gaussians; SURF: a fast Hessian-determinant approximation using integral images), and describe each with a local gradient-orientation histogram, giving a descriptor that is invariant (or near-invariant) to image scale, in-plane rotation, and moderate illumination/viewpoint change. They are used for image matching/registration, panorama stitching, object recognition, and camera-pose/structure-from-motion estimation, wherever the same physical point must be re-identified across images taken from different distances or angles.
A 2D median filter replaces each pixel with the median of the intensity values in a small neighbourhood window (e.g. $3\times3$) centred on it, rather than a weighted average. Because the median is an order statistic rather than a linear combination, a single wildly-off outlier pixel (impulsive/salt-and-pepper noise) is simply ignored as long as it is not the majority value in the window, and step edges are preserved (the median returns one of the actual neighbouring pixel values, not a blurred in-between value). It is used for removing impulsive sensor/transmission noise from images and video frames while keeping edges sharp, which a linear (mean/Gaussian) smoothing filter cannot do simultaneously.
Morphological operations process a binary (or greyscale) image using a small structuring element, via the two primitive operations erosion (shrinks foreground regions, removing small noise specks) and dilation (grows foreground regions, filling small gaps); combining them gives opening (erosion then dilation, removes small protrusions/noise) and closing (dilation then erosion, fills small holes/gaps). It is used as clean-up after thresholding/segmentation — removing speckle noise, closing small gaps in a detected boundary, filling holes inside a segmented blob, and extracting connected-component shape descriptors.