NivaarExam PrepOfficial exam papers ↗

20-Bio-B4 Robotics · December 2014

Question 3 of 6: The Human Visual System

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

Paper format: National Exams, December 2014 — 04-Bio-B4 Digital Image Processing. Three hours, open book (any paper notes or textbooks permitted, but no calculator or computer). Six questions of equal value (20 marks each); five constitute a complete paper and only the first five appearing in the answer book are marked. All six are solved here, because this set is a study resource rather than an examination script. Every question is essay/descriptive (definitions, algorithm design, system design) except the operation-count comparison in Question 2(d), which is a short analytical calculation.

Reference texts (the books a candidate should have reviewed for this subject):

Question 3: The Human Visual System (20 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

(a) The Two Photoreceptor Types

The retina contains rods and cones. Rods are far more numerous (roughly 100–120 million per eye) and extremely light-sensitive (responding to a handful of photons), supporting monochromatic scotopic (low-light) vision but with poor spatial acuity because many rods converge onto a single downstream ganglion cell. Cones are much less numerous (roughly 6–7 million per eye), require brighter light, and come in three types tuned to short/medium/long wavelengths (S, M, L, loosely "blue/green/red"), providing colour vision and, because each foveal cone connects to its own ganglion cell, high spatial acuity under photopic (daylight) conditions.

(b) Retinal Distribution and Its Rationale

Cone density peaks sharply in the fovea (the central ~1–2° of the visual field) and falls off rapidly with eccentricity; rods are entirely absent from the very centre of the fovea, rise to a broad peak around 15–20° eccentricity, and then gradually decline toward the periphery. Neither receptor is present at the optic disc (the nerve-fibre exit point), producing the physiological blind spot.

Eccentricity from fovea (deg), temporal ← 0 → nasal Receptor density (relative) fovea (0°) optic disc (blind spot) cones (peak at fovea) rods (peak ~18-20° eccentricity)
Relative rod and cone density as a function of eccentricity from the fovea (schematic, after the classical Osterberg curve).

The rationale is a division of labour that matches receptor placement to how the eye actually samples the world: the fovea is where the eye fixates to inspect fine detail and colour (reading, face recognition), so it is worth concentrating high-acuity, colour-capable but less-sensitive cones there at 1:1 ganglion-cell wiring; the periphery is used mainly to detect motion and to trigger a saccade toward anything of interest under a wide range of lighting, so it is better served by numerous, highly sensitive but low-acuity, highly-converged rods. This non-uniform, foveated sampling strategy — high resolution only where attention is directed — is the biological precursor of foveated/region-of-interest image compression schemes.

(c) Early Image Processing in the Retina

The retina is not a passive sensor array; the bipolar/horizontal/amacrine/ganglion-cell circuitry performs real image processing before the signal ever reaches the optic nerve. Centre-surround receptive fields (a photoreceptor's response is opposed by lateral inhibition from its neighbours, mediated by horizontal and amacrine cells) implement a spatial band-pass/edge-enhancement operation, similar to unsharp masking, that boosts local contrast at boundaries (Mach bands are the perceptual signature of this). Photoreceptor and network-level adaptation continuously rescale sensitivity to the local light level (light/dark adaptation), giving the retina an effective dynamic-range compression well beyond what a linear sensor could achieve. Finally, convergence of roughly 100–120 million photoreceptors onto only about 1–1.5 million ganglion-cell axons in the optic nerve is itself a drastic, lossy data-reduction (compression) step, discarding redundant, spatially-correlated information before transmission.

(d) HVS-Exploiting Algorithms

i. Image Denoising. The human contrast sensitivity function (CSF) is band-pass and rolls off at both very low and very high spatial frequencies, and visual masking hides noise/artifacts near strong edges or high-texture regions more than in smooth areas. Denoising algorithms exploit this by shaping noise removal to match perceptual sensitivity rather than to minimize raw pixel error: wavelet-domain thresholding removes small high-frequency coefficients (the band the eye is least sensitive to and where noise concentrates) while preserving low-frequency structure, and edge-preserving filters (bilateral filtering, non-local means) explicitly avoid smoothing across edges — mimicking the retina's own edge-preserving centre-surround processing from part (c) — because noise near an edge is masked less effectively than noise in a flat region.

ii. Colour Image Compression. Because there are far fewer cones than the eye's effective luminance (rod + cone) sampling, and because the retinal/cortical pathways carry luminance information at higher spatial resolution than the two colour-opponent chrominance channels, human colour-spatial acuity is much lower than luminance-spatial acuity. Compression standards (JPEG, MPEG, broadcast video) exploit this directly by converting to a luma/chroma space (YCrCb, part 1(d)) and sub-sampling the chroma channels (e.g. 4:2:0, one chroma sample per 2×2 luma block) while keeping full-resolution luma — a 50% or greater data reduction with little visible loss, precisely because it targets the channel the visual system already resolves coarsely.