20-Bio-B4 Robotics · May 2015
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Paper format: National Exams, May 2015 — 04-Bio-B4 Image Processing. Three hours, open book (any paper notes or textbooks permitted, but no calculator or computer). Six questions of equal value (20 marks each); five constitute a complete paper and only the first five appearing in the answer book are marked. All six are solved here, because this set is a study resource rather than an examination script. Every question is essay/descriptive (definitions, algorithm design, system design); the only quantitative content is the computational-complexity discussion in Question 3(f)/(g).
Reference texts (the books a candidate should have reviewed for this subject):
Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.
Given an image defined on domain $\Omega$, segmentation is a partition of $\Omega$ into regions $R_1, R_2, \dots, R_n$ such that $\bigcup_{i=1}^{n} R_i = \Omega$, $R_i \cap R_j = \varnothing$ for $i \neq j$, each $R_i$ is connected, a homogeneity predicate satisfies $P(R_i) = \text{true}$ for every region, and $P(R_i \cup R_j) = \text{false}$ for every pair of adjacent regions $R_i, R_j$ — i.e. each region is internally consistent and no two adjacent regions could be merged without breaking that consistency.
Segmentation assumes the image is well modelled as piecewise-homogeneous: pixels belonging to the same object share a property (intensity, colour, or texture statistic) that is approximately constant or slowly varying within the object, and that this property changes sharply at object boundaries. It further assumes object boundaries are spatially coherent (closed, reasonably smooth curves rather than fractal noise) and that the amount of sensor noise is small relative to the between-object signal difference, so that the homogeneity predicate is not swamped by measurement noise.
Thresholding assumes the intensity (or a derived feature) histogram is bimodal, with one mode per class; it is extremely fast and simple (global or locally-adaptive/Otsu variants), but fails outright under uneven illumination or when object/background histograms overlap. Region growing assumes local homogeneity around a set of seed pixels and grows regions by adding neighbouring pixels that satisfy a similarity test; it produces connected regions by construction and handles gradual intensity variation better than thresholding, but is sensitive to seed placement and to noise (a single mis-included pixel can "leak" the growth into the wrong region). Watershed segmentation treats the gradient-magnitude image as a topographic surface and floods it from regional minima, building dams where floods from different basins meet; it naturally separates touching or overlapping objects, but is prone to severe over-segmentation from spurious local minima unless the gradient image is first smoothed or seeded with markers.
(1) Medical image analysis — delineating an organ, tumour, or vessel in MRI/CT to support diagnosis, volume measurement, and radiotherapy treatment planning. (2) Autonomous driving / robot scene understanding — partitioning a camera frame into road, vehicles, pedestrians, and signage so a planner can reason about drivable space and obstacles. (3) Remote sensing / industrial inspection — classifying land cover in a satellite image into water, vegetation, and built-up area (directly relevant to Question 6's rooftop-area problem), or isolating a surface defect region on a manufactured part.
Noisy images: isolated noisy pixels violate local homogeneity, causing thresholding/region growing to over-segment into many spurious tiny regions or leak across true boundaries. Mitigation: pre-filter with an edge-preserving smoother (median filter or the anisotropic diffusion of Question 1(b)) before segmenting, and use spatially-regularized criteria (e.g. a Markov-random-field prior, or watershed with markers) that penalize label noise rather than reacting to every pixel independently.
Low-contrast images taken at night: the intensity separation between object and background shrinks toward the sensor's noise floor, so a fixed global threshold or gradient-magnitude criterion becomes unreliable (the true signal and noise become comparable in size). Mitigation: apply local/adaptive contrast enhancement (e.g. CLAHE) before segmenting, use locally-adaptive rather than global thresholds, or exploit gradient direction (which survives at lower contrast than gradient magnitude) instead of relying on magnitude alone.
Colour images: the three (or more) channels are highly correlated, so segmenting each channel independently and combining the results is inconsistent, and RGB itself confounds brightness with colour, making chromatic boundaries harder to separate from illumination changes. Mitigation: convert to a decorrelated, perceptually-organized space (HSV or Lab) before segmenting, and use a joint (vector-valued) distance metric or joint clustering across channels rather than per-channel thresholds.