NivaarExam PrepOfficial exam papers ↗

20-Bio-B4 Robotics · December 2014

Question 4 of 6: Color Image Processing

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

Paper format: National Exams, December 2014 — 04-Bio-B4 Digital Image Processing. Three hours, open book (any paper notes or textbooks permitted, but no calculator or computer). Six questions of equal value (20 marks each); five constitute a complete paper and only the first five appearing in the answer book are marked. All six are solved here, because this set is a study resource rather than an examination script. Every question is essay/descriptive (definitions, algorithm design, system design) except the operation-count comparison in Question 2(d), which is a short analytical calculation.

Reference texts (the books a candidate should have reviewed for this subject):

Question 4: Color Image Processing (20 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

(a) Main Colorspaces: Strengths and Weaknesses

ColorspaceStrengthWeakness
RGBNative to sensors/displays; simple to acquire and renderChannels highly correlated; not perceptually uniform; colour and brightness are entangled
CMY(K)Matches subtractive print processes directlyOnly useful for print; ink-mixing non-idealities need a K (black) channel
HSV / HSISeparates chromaticity (hue, saturation) from intensity; intuitive, illumination-robust colour segmentationHue is unstable/undefined near zero saturation; non-linear conversion from RGB
YCrCb / YUVSeparates luma from chroma; compression- and broadcast-friendly (chroma sub-sampling, grayscale backward-compatibility)Not perceptually uniform for colour-difference measurement
CIE LabPerceptually uniform (equal numeric distance ≈ equal perceived difference); device-independentComputationally heavier; non-intuitive axes for a novice user

(b) Role of the Bayer Pattern

G R G R G R B G B G B G G R G R G R B G B G B G 2x2 repeating tile (1R : 2G : 1B) each sensor pixel sees ONE colour only -- full RGB is interpolated (demosaicing)
The Bayer colour filter array: a 2×2 repeating RGGB tile placed over the sensor, with twice as many green filters as red or blue.

A single monochrome sensor cannot distinguish colour on its own, and a 3-sensor/prism camera is expensive and bulky. The Bayer pattern places a colour filter array (CFA) directly over the sensor so that each pixel measures only one colour channel, arranged as a repeating 2×2 tile of one red, two green, and one blue filter. Its significance is twofold: it enables low-cost, compact single-sensor colour capture, and the double weighting of green matches the human luminance response, which peaks near 555 nm (green) — devoting more samples to green improves the effective resolution and SNR of the reconstructed luminance channel, which is the channel human vision is most sensitive to (Question 3(d)). The missing two-thirds of each pixel's colour information is then reconstructed by demosaicing (interpolating each missing channel from its neighbours) to produce the final full-resolution RGB image.

(c) Multispectral vs. Colour Image Processing

Ordinary colour image processing works with exactly three broad, perceptually-motivated bands (R, G, B or an equivalent transform of them) chosen to match human trichromatic vision, and its tools (colour spaces, gamut mapping) exist to serve human viewing or display. Multispectral (and hyperspectral) processing, as in remote sensing, instead captures dozens to hundreds of narrow spectral bands, frequently extending well beyond the visible range (near-infrared, short-wave infrared, thermal), with each band registered as its own separate image. The analysis goal is also different: rather than reproducing a picture for a human viewer, multispectral analysis classifies materials/land-cover by their spectral signature (e.g. vegetation vigour via the NDVI band ratio, mineral or crop-stress identification), and because so many bands are strongly correlated, dimensionality reduction (principal component analysis) is a standard pre-processing step that has no counterpart in ordinary 3-channel colour processing.

(d) Rationale for Processing Only U and V

In a YUV (or YCrCb) representation, Y (luma) carries the bulk of the spatial detail and edge information and is the channel human vision is most sensitive to, while U and V carry only chrominance (hue/saturation) information at a resolution the eye cannot fully appreciate. An algorithm that needs to adjust colour balance, hue, saturation, or white balance can therefore restrict itself to the U/V channels: this (i) guarantees the luminance/edge/texture detail in Y is left completely untouched, avoiding any risk of re-introducing blur or artifacts into fine detail, (ii) is computationally cheaper, since U and V are frequently stored/processed at lower resolution than Y (chroma sub-sampling) so there are fewer samples to touch, and (iii) cleanly separates "what colour is it" (U, V) from "how bright/detailed is it" (Y), which is exactly the decomposition a colour-correction algorithm needs.

(e) Generalizing Gray-Scale Methods to Colour

i. Edge Detection. Applying a gradient operator (Sobel, Canny) independently to each of R, G, B and then simply summing/max-ing the per-channel edge maps can miss or cancel true colour edges — e.g. a boundary between two colours of equal luminance produces near-zero gradient in a naive grayscale-converted image even though the colour clearly changes. The vector-valued generalization (e.g. the Di Zenzo multichannel gradient) treats each pixel as an $(R,G,B)$ vector and computes a true colour gradient from the per-channel partial-derivative vectors' inner products, correctly detecting edges that exist in colour even when they are invisible in luminance alone.

ii. Image Segmentation. Gray-scale segmentation thresholds or clusters on a single intensity value; colour segmentation instead measures similarity/distance between full colour vectors, usually in a perceptually meaningful space (Lab distance, or HSV hue/saturation clustering) so that region-growing, k-means, or mean-shift clustering group pixels that are perceptually the same colour rather than merely the same brightness — and, using hue rather than RGB, segmentation can be made largely invariant to shading/illumination changes across an object's surface.

iii. Image Denoising. Filtering each colour channel independently with a scalar filter (e.g. per-channel median) can shift the three channels' edges independently, producing visible false colours ("colour fringing") along object boundaries that were not present in the noisy image. The vector generalization (e.g. the vector median filter, which selects the actual observed RGB pixel vector in a neighbourhood that minimizes total vector distance to all others, rather than combining channels independently) preserves the correlation between channels and avoids introducing new colours that were never in the original scene.