Question 1 of 6: Cache Size, Control-Unit Style, and Bus-Width Performance
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Notes on this paper
98-Comp-A3, Computer Architecture — National Exams, December 2015. Closed-book, 3 hours; six questions of equal value (20 marks each); FIVE constitute a complete exam (all six answered below as a complete study resource).
Reference texts: Patterson & Hennessy, Computer Organization and Design, 6th ed. — memory hierarchy & cache performance (Q1a, Q2a, Q3a, Q5a), instruction encoding & RISC design (Q2c, Q3b, Q4c), IEEE-754 floating point (Q4b), and memory technology (Q4a); Mano & Ciletti, Digital Design, 6th ed. — control-unit design, register-transfer micro-operations, stack/RPN notation, and binary-multiplication hardware (Q1b, Q3c, Q5b–d, Q6a–c).
Given. (a) The claim "bigger cache is always better." (b) Two microinstruction encoding styles: horizontal (wide, largely unencoded control word) and vertical (narrow, heavily encoded, decoded before use). (c) Three processors whose data/instruction buses are 64, 32, and 16 bits wide, all with equal-duration bus cycles; an operand/instruction-size mix of 25% 64-bit, 45% 32-bit, 20% 16-bit, 10% 8-bit.
Find. (a) Agree/disagree with justification. (b) Pros/cons of horizontal vs. vertical micro-instructions. (c) The relative performance of the 64-, 32-, and 16-bit buses over this operand mix.
Approach. (a)–(b) argue from the underlying hardware tradeoffs (diminishing returns/access-time cost for caches; word-width vs. decode-latency for micro-instructions); (c) a bus of width $w$ needs $\lceil s/w\rceil$ equal-length bus cycles to move an item of size $s$, so average cycles/item is the mix-weighted sum, and relative performance is inversely proportional to that average.
Part (a) — is bigger cache always better?Disagree, in general — the statement is true only up to a point. A larger cache does reduce the miss rate for a fixed working set, because it can hold more of the blocks a program actually revisits (temporal/spatial locality), and up to the size of the program's working set, hit rate climbs quickly with capacity. But average memory access time is $\text{AMAT}=T_{hit}+\text{miss rate}\times\text{miss penalty}$, and $T_{hit}$ itself tends to grow with cache size: a bigger cache needs more sets/ways, longer bit-lines and wider tag comparisons, all of which increase access latency and often force a longer clock period or extra pipeline stages just to reach the cache. Once the cache already covers the working set, adding more capacity yields rapidly diminishing miss-rate returns while $T_{hit}$ keeps rising — past that point AMAT can actually get WORSE, not better. Cost, die area, and static power also scale with capacity. Conclusion: bigger caches help only while capacity is the bottleneck (miss rate falls faster than hit time rises); beyond the working-set size, "bigger" trades a rare benefit for a certain, recurring hit-time cost.
Part (b) — horizontal vs. vertical micro-instructions.Horizontal micro-instructions dedicate one (or a few) bits directly to each control signal, so many micro-operations can be specified and issued in parallel within a single micro-instruction. Advantages: maximum parallelism (several datapath actions fire in the same cycle), simple decode logic (each field drives its signal with little or no decoding), and straightforward to design/modify. Disadvantages: the control word is very wide (one bit per signal), so the control memory is large and expensive, and most bits are "0" (unused) in any given instruction — wasteful of ROM/PLA area. Vertical micro-instructions instead encode groups of mutually-exclusive control signals into short op-code-like fields (much like a miniature instruction set), which must be decoded before use. Advantages: much narrower control words, smaller and cheaper control memory. Disadvantages: only one (or few) micro-operations per field can be selected per cycle, so operations that could run in parallel on a horizontal design may need extra micro-instructions/cycles here, AND an extra decode step is inserted into the control path, adding latency. Horizontal trades control-store size for speed and parallelism; vertical trades speed/parallelism for a smaller, cheaper control store.
Part (c) — relative bus performance. A bus of width $w$ bits moves an item of size $s$ bits in $\lceil s/w\rceil$ equal-duration bus cycles (any item that doesn't divide evenly still consumes a whole extra cycle). Applying this to each processor over the given mix:
Item size $s$
Fraction
Cycles, 64-bit bus
Cycles, 32-bit bus
Cycles, 16-bit bus
64 bits
25%
$\lceil64/64\rceil=1$
$\lceil64/32\rceil=2$
$\lceil64/16\rceil=4$
32 bits
45%
$\lceil32/64\rceil=1$
$\lceil32/32\rceil=1$
$\lceil32/16\rceil=2$
16 bits
20%
$\lceil16/64\rceil=1$
$\lceil16/32\rceil=1$
$\lceil16/16\rceil=1$
8 bits
10%
$\lceil8/64\rceil=1$
$\lceil8/32\rceil=1$
$\lceil8/16\rceil=1$
Weighting each column by its fraction gives the average number of bus cycles needed per fetched item:
$$\bar c_{64}=0.25(1)+0.45(1)+0.20(1)+0.10(1)=\boxed{1.00\ \text{cycle}}$$
$$\bar c_{32}=0.25(2)+0.45(1)+0.20(1)+0.10(1)=\boxed{1.25\ \text{cycles}}$$
$$\bar c_{16}=0.25(4)+0.45(2)+0.20(1)+0.10(1)=\boxed{2.20\ \text{cycles}}$$
Since every bus cycle takes the same duration, throughput (items fetched per unit time) is inversely proportional to the average cycles/item. Normalizing to the 64-bit machine:
$$\text{perf}_{64}:\text{perf}_{32}:\text{perf}_{16}=\frac{1}{1.00}:\frac{1}{1.25}:\frac{1}{2.20}=\boxed{1:0.80:0.4545}$$
so the 64-bit processor is $1.25\times$ faster than the 32-bit machine and $2.2\times$ faster than the 16-bit machine at this fetch workload.
Final results — Question 1
Part
Result
(a) bigger cache
Disagree in general — helps only up to the working-set size; beyond it, rising hit time can outweigh the falling miss rate
(b) horizontal vs. vertical
Horizontal: fast/parallel, wide/costly control store. Vertical: narrow/cheap control store, slower (decode + fewer parallel ops)