NivaarExam PrepOfficial exam papers ↗

25-Comp-A3 Computer Architecture · December 2015

Question 1 of 6: Cache Size, Control-Unit Style, and Bus-Width Performance

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

98-Comp-A3, Computer Architecture — National Exams, December 2015. Closed-book, 3 hours; six questions of equal value (20 marks each); FIVE constitute a complete exam (all six answered below as a complete study resource).

Reference texts: Patterson & Hennessy, Computer Organization and Design, 6th ed. — memory hierarchy & cache performance (Q1a, Q2a, Q3a, Q5a), instruction encoding & RISC design (Q2c, Q3b, Q4c), IEEE-754 floating point (Q4b), and memory technology (Q4a); Mano & Ciletti, Digital Design, 6th ed. — control-unit design, register-transfer micro-operations, stack/RPN notation, and binary-multiplication hardware (Q1b, Q3c, Q5b–d, Q6a–c).

Question 1: Cache Size, Control-Unit Style, and Bus-Width Performance (20 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

Given. (a) The claim "bigger cache is always better." (b) Two microinstruction encoding styles: horizontal (wide, largely unencoded control word) and vertical (narrow, heavily encoded, decoded before use). (c) Three processors whose data/instruction buses are 64, 32, and 16 bits wide, all with equal-duration bus cycles; an operand/instruction-size mix of 25% 64-bit, 45% 32-bit, 20% 16-bit, 10% 8-bit.

Find. (a) Agree/disagree with justification. (b) Pros/cons of horizontal vs. vertical micro-instructions. (c) The relative performance of the 64-, 32-, and 16-bit buses over this operand mix.

Approach. (a)–(b) argue from the underlying hardware tradeoffs (diminishing returns/access-time cost for caches; word-width vs. decode-latency for micro-instructions); (c) a bus of width $w$ needs $\lceil s/w\rceil$ equal-length bus cycles to move an item of size $s$, so average cycles/item is the mix-weighted sum, and relative performance is inversely proportional to that average.

  1. Part (a) — is bigger cache always better? Disagree, in general — the statement is true only up to a point. A larger cache does reduce the miss rate for a fixed working set, because it can hold more of the blocks a program actually revisits (temporal/spatial locality), and up to the size of the program's working set, hit rate climbs quickly with capacity. But average memory access time is $\text{AMAT}=T_{hit}+\text{miss rate}\times\text{miss penalty}$, and $T_{hit}$ itself tends to grow with cache size: a bigger cache needs more sets/ways, longer bit-lines and wider tag comparisons, all of which increase access latency and often force a longer clock period or extra pipeline stages just to reach the cache. Once the cache already covers the working set, adding more capacity yields rapidly diminishing miss-rate returns while $T_{hit}$ keeps rising — past that point AMAT can actually get WORSE, not better. Cost, die area, and static power also scale with capacity. Conclusion: bigger caches help only while capacity is the bottleneck (miss rate falls faster than hit time rises); beyond the working-set size, "bigger" trades a rare benefit for a certain, recurring hit-time cost.
  2. Part (b) — horizontal vs. vertical micro-instructions. Horizontal micro-instructions dedicate one (or a few) bits directly to each control signal, so many micro-operations can be specified and issued in parallel within a single micro-instruction. Advantages: maximum parallelism (several datapath actions fire in the same cycle), simple decode logic (each field drives its signal with little or no decoding), and straightforward to design/modify. Disadvantages: the control word is very wide (one bit per signal), so the control memory is large and expensive, and most bits are "0" (unused) in any given instruction — wasteful of ROM/PLA area. Vertical micro-instructions instead encode groups of mutually-exclusive control signals into short op-code-like fields (much like a miniature instruction set), which must be decoded before use. Advantages: much narrower control words, smaller and cheaper control memory. Disadvantages: only one (or few) micro-operations per field can be selected per cycle, so operations that could run in parallel on a horizontal design may need extra micro-instructions/cycles here, AND an extra decode step is inserted into the control path, adding latency. Horizontal trades control-store size for speed and parallelism; vertical trades speed/parallelism for a smaller, cheaper control store.
  3. Part (c) — relative bus performance. A bus of width $w$ bits moves an item of size $s$ bits in $\lceil s/w\rceil$ equal-duration bus cycles (any item that doesn't divide evenly still consumes a whole extra cycle). Applying this to each processor over the given mix:
    Item size $s$FractionCycles, 64-bit busCycles, 32-bit busCycles, 16-bit bus
    64 bits25%$\lceil64/64\rceil=1$$\lceil64/32\rceil=2$$\lceil64/16\rceil=4$
    32 bits45%$\lceil32/64\rceil=1$$\lceil32/32\rceil=1$$\lceil32/16\rceil=2$
    16 bits20%$\lceil16/64\rceil=1$$\lceil16/32\rceil=1$$\lceil16/16\rceil=1$
    8 bits10%$\lceil8/64\rceil=1$$\lceil8/32\rceil=1$$\lceil8/16\rceil=1$
    Weighting each column by its fraction gives the average number of bus cycles needed per fetched item: $$\bar c_{64}=0.25(1)+0.45(1)+0.20(1)+0.10(1)=\boxed{1.00\ \text{cycle}}$$ $$\bar c_{32}=0.25(2)+0.45(1)+0.20(1)+0.10(1)=\boxed{1.25\ \text{cycles}}$$ $$\bar c_{16}=0.25(4)+0.45(2)+0.20(1)+0.10(1)=\boxed{2.20\ \text{cycles}}$$ Since every bus cycle takes the same duration, throughput (items fetched per unit time) is inversely proportional to the average cycles/item. Normalizing to the 64-bit machine: $$\text{perf}_{64}:\text{perf}_{32}:\text{perf}_{16}=\frac{1}{1.00}:\frac{1}{1.25}:\frac{1}{2.20}=\boxed{1:0.80:0.4545}$$ so the 64-bit processor is $1.25\times$ faster than the 32-bit machine and $2.2\times$ faster than the 16-bit machine at this fetch workload.
Final results — Question 1
PartResult
(a) bigger cacheDisagree in general — helps only up to the working-set size; beyond it, rising hit time can outweigh the falling miss rate
(b) horizontal vs. verticalHorizontal: fast/parallel, wide/costly control store. Vertical: narrow/cheap control store, slower (decode + fewer parallel ops)
(c) avg cycles/item64-bit: 1.00; 32-bit: 1.25; 16-bit: 2.20
(c) relative performance$1:0.80:0.4545$ (64-bit fastest, 16-bit slowest)
← Paper overview