22-Elec-B3 Digital Communications Systems · December 2015
Question 5 of 5: Sampling and D/A conversion
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Notes on this paper
Paper format. Professional Engineers of Ontario annual examinations, December 2015, 07-Elec-B3 Digital Communication Systems — 3 hours, closed book, a PEO-approved non-programmable calculator permitted. Five questions of 25 marks are printed; any four constitute a complete paper worth 100 marks, and only the first four appearing in the answer book are marked. Marks are shown in the left margin. Note 1 on the cover page urges the candidate to submit a clear statement of any assumptions made. All five questions are solved below, because the set is intended as a study resource rather than a sitting.
Reference texts. J. G. Proakis and M. Salehi, Communication Systems Engineering, 2nd ed. (link budgets, source coding, block codes, PCM); S. Haykin and M. Moher, Communication Systems, 5th ed.; B. Sklar, Digital Communications: Fundamentals and Applications, 2nd ed. (spread spectrum, ch. 12); T. M. Cover and J. A. Thomas, Elements of Information Theory, 2nd ed. (entropy, Huffman and Shannon–Fano–Elias codes); S. Lin and D. J. Costello, Error Control Coding, 2nd ed. (linear block codes); A. V. Oppenheim and R. W. Schafer, Discrete-Time Signal Processing, 3rd ed. (sampling, quantization); T. S. Rappaport, Wireless Communications: Principles and Practice, 2nd ed. In the Canadian frame, licence-exempt spread-spectrum equipment is governed by ISED RSS-247 and spectrum allocations by the Canadian Table of Frequency Allocations.
Question 5: Sampling and D/A conversion (25 marks)
Given. A CD-quality audio signal of bandwidth $W = 20$ kHz, to be sampled and encoded by pulse code modulation at $n = 16$ bits per sample; for part (d) the converter's full-scale range is $-5$ V to $+5$ V, a span of $V_{FS} = 10$ V; for part (e) a typical MP3 stream runs at 128 kbit/s.
Find. (a) the Nyquist criterion with its Fourier-transform justification, (b) the minimum sampling frequency, (c) an explanation of PCM and the resulting data rate, (d) the largest quantization error, and (e) a reason MP3 needs far less rate.
Check: part (c) refers to “the signal from part a”. Part (a) asks for a statement of the criterion and supplies no signal; the 20 kHz CD-quality signal is introduced in part (b). The reference is therefore taken to mean the signal of part (b) sampled at the minimum rate found there, 40 kHz, which is also what the question's own fallback instruction (“assume a value”) invites. The rate is quoted per channel and then for a stereo pair, and the real CD figure at 44.1 kHz is given for comparison so the answer is usable under either reading.
Approach. Derive the replica structure of the sampled spectrum to establish the criterion, apply it to the 20 kHz bandwidth, multiply rate by resolution for the PCM bit rate, take half a quantization step for the error bound, and finish with the perceptual-coding argument for MP3.
Part (a) — state the criterion. A signal $x(t)$ whose Fourier transform vanishes for $|f|\gt W$ is completely determined by samples taken at uniform intervals $T_{s}=1/f_{s}$ provided $$f_{s} \gt 2W ,$$ and $2W$ is the Nyquist rate. Uniqueness fails at or below it in general, and equality is admissible only for signals with no spectral content exactly at $W$.
Justify it from the Fourier transform of the sampled signal. Ideal sampling multiplies $x(t)$ by an impulse train of period $T_{s}$, whose transform is itself an impulse train of spacing $f_{s}$. Multiplication in time is convolution in frequency, so $$X_{s}(f) = f_{s}\sum_{k=-\infty}^{\infty} X(f - k f_{s}) :$$ the spectrum of the samples is a periodic repetition of the original, one copy centred on every multiple of $f_{s}$. Each copy occupies $|f|\le W$ about its centre, so neighbouring copies stay disjoint exactly when $f_{s}-W \gt W$, i.e. $f_{s}\gt 2W$. When they are disjoint, an ideal low-pass filter of cutoff $f_{s}/2$ removes every copy but the one at the origin, recovering $X(f)$ and therefore $x(t)$ exactly; in the time domain that filter is the interpolation $$x(t) = \sum_{n} x(nT_{s})\,\mathrm{sinc}\!\left(\frac{t-nT_{s}}{T_{s}}\right).$$ If $f_{s}\lt 2W$ the copies overlap, the overlapping content adds irreversibly, and the high frequencies reappear as low ones — aliasing, which no later processing can undo. This is why a practical converter is always preceded by an anti-alias low-pass filter.
Figure 5.1 — the baseband spectrum and its replicas after sampling. The gap $f_{s}-2W$ is the room the reconstruction filter needs; at exactly the Nyquist rate it closes to zero.
Part (b) — minimum sampling frequency for CD-quality audio. With $W = 20$ kHz the criterion gives $$f_{s,\min} = 2W = \boxed{40\ \text{kHz}} .$$ Compact disc actually uses 44.1 kHz, and the 4.1 kHz excess is not waste: it leaves a transition band of $44.1/2 - 20 = 2.05$ kHz in which the anti-alias and reconstruction filters can roll off, since an ideal brick wall at exactly 20 kHz is unrealisable.
Part (c) — what pulse code modulation is. PCM converts an analogue waveform to a bit stream in three steps: the signal is sampled at $f_{s}$ (after anti-alias filtering), each sample amplitude is quantized to the nearest of $2^{n}$ discrete levels, and the level index is encoded as an $n$-bit binary word. The output is a stream of fixed-length words at the sample rate; it is the reference against which every compression scheme is measured, and it is memoryless — every sample costs the same $n$ bits whatever the signal is doing.
PCM data rate. The rate is simply resolution times sample rate times the number of channels: $$R_{b} = n\,f_{s} = 16 \times 40\times10^{3} = \boxed{640\ \text{kbit/s per channel}},$$ so a stereo pair needs $2\times 640 = 1.28$ Mbit/s. At the actual CD rate the figure is $16\times 44.1\times10^{3}\times 2 = 1.4112$ Mbit/s, which is the familiar uncompressed-audio number.
Part (d) — the quantization step. A 16-bit converter divides the 10 V span into $2^{16} = 65{,}536$ intervals, so $$\Delta = \frac{V_{FS}}{2^{n}} = \frac{10}{65{,}536} = 152.59\ \mu\text{V}.$$
Maximum quantization error. Rounding each sample to the nearest level means the residual can never exceed half a step, so $$|e|_{\max} = \frac{\Delta}{2} = \boxed{76.29\ \mu\text{V}},$$ and the error is bounded independently of the signal, which is what makes the bound useful. Some texts divide by $2^{n}-1$ rather than $2^{n}$, giving 152.59 µV and 76.30 µV — indistinguishable at 16-bit resolution. Expressed as a ratio, a full-scale sinusoid sees a signal-to-quantization-noise ratio of $6.02n + 1.76 = 98.1$ dB, the origin of the “about 6 dB per bit” rule.
Part (e) — why MP3 needs far less rate. MP3 is lossy perceptual coding, whereas PCM is a faithful sample-by-sample representation. A filter bank splits the signal into subbands and a psychoacoustic model computes, band by band and frame by frame, the masking threshold below which the ear cannot hear anything because of louder nearby content; the coder then allocates only enough bits per band to push the quantization noise just under that threshold, discards what is masked outright, exploits the redundancy between the two stereo channels through joint-stereo coding, and finally Huffman-codes the quantized coefficients. PCM, by contrast, spends its full 16 bits on every sample regardless of whether the detail is audible. A 128 kbit/s MP3 stream is therefore about a tenth of the 1.28 Mbit/s uncompressed stereo rate computed in part (c), and about an eleventh of the real 1.4112 Mbit/s CD rate, while remaining close to transparent for most listeners.
Final results
Quantity
Result
(a) Nyquist criterion
$f_{s}\gt 2W$; sampled spectrum $X_{s}(f)=f_{s}\sum_{k}X(f-kf_{s})$, replicas disjoint so an ideal low-pass at $f_{s}/2$ recovers $x(t)$
640 kbit/s per channel; 1.28 Mbit/s stereo (1.4112 Mbit/s at the real 44.1 kHz)
Quantization step over 10 V
$\Delta = 152.59\ \mu\text{V}$
(d) Maximum quantization error
76.29 µV ($\Delta/2$)
Signal-to-quantization-noise ratio, 16 bits
$6.02(16)+1.76 = 98.1$ dB
(e) Why MP3 is much lower rate
lossy perceptual coding: psychoacoustic masking, per-band bit allocation, joint stereo and Huffman coding — about 10:1 against the 1.28 Mbit/s PCM stream