NivaarExam PrepOfficial exam papers ↗

25-Comp-A3 Computer Architecture · May 2016

Question 1 of 6: Cache Organization, Control-Unit Signals, and Bus-Width Transfer Rate

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

98-Comp-A3, Computer Architecture — National Exams, May 2016. Closed-book, 3 hours; six questions of equal value (20 marks each); FIVE constitute a complete exam (all six answered below as a complete study resource).

Reference texts: Patterson & Hennessy, Computer Organization and Design, 6th ed. — memory hierarchy & cache design (Q1a, Q2a, Q3a, Q5a), bus/data-transfer performance (Q1c), instruction-level parallelism (Q2c), instruction encoding & RISC/CISC tradeoffs (Q3b, Q4c), IEEE-754 floating point and memory technology (Q4a–b), branch prediction and addressing modes (Q6a–b); Mano & Ciletti, Digital Design, 6th ed. — control-unit design (Q1b, Q5c–d), unsigned binary division hardware (Q3c), reverse-Polish/stack notation (Q5b), and shift operations (Q6c); Stallings, Data and Computer Communications — programmed vs. interrupt-driven I/O (Q2b).

Question 1: Cache Organization, Control-Unit Signals, and Bus-Width Transfer Rate (20 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

Given. (a) Two cache organizations for the same total capacity: one unified (holds both instructions and data) and one split into separate instruction/data caches. (b) A control unit as a black box within the CPU. (c) Two otherwise-identical processors whose external data buses are 8 bits and 16 bits wide, with equal-duration bus cycles; an item-size mix of 30% one-byte and 70% two-byte operands/instructions.

Find. (a) The tradeoffs of unified vs. split cache. (b) The control unit's inputs, outputs, and the purpose of each. (c) The factor by which the two buses' maximum data transfer rates differ.

Approach. (a)–(b) reason from first principles about resource contention and the sequencing role of the control unit; (c) a bus of width $w$ needs $\lceil s/w\rceil$ equal-length cycles to move an $s$-byte item, so the average cycles/item over the mix sets the relative throughput.

  1. Part (a) — unified vs. split cache. A unified cache stores instructions and data in the same array. Pros: for a given total capacity, a unified cache automatically balances space between instruction and data demand — a code-heavy phase of execution can use more of the cache for instructions, a data-heavy phase more for data — so it never wastes capacity by rigidly partitioning it. It is also simpler to design (one set of tag/data arrays, one replacement policy). Cons: a single-ported unified cache forces instruction fetch and data load/store to compete for the same access port every cycle, which stalls a pipelined CPU whenever both need the cache in the same cycle (a structural hazard). A split (Harvard-style) cache uses two independent caches, one for instructions (I-cache) and one for data (D-cache). Pros: instruction fetch and data access happen in parallel with no port conflict, which is essential for a pipelined design (fetch stage reads I-cache while a memory stage reads/writes D-cache in the same cycle); each cache can also be tuned separately (instructions are read-only and exhibit strong sequential locality, so the I-cache can use a simpler policy than the D-cache). Cons: the fixed split means capacity cannot flow to whichever stream needs it more, so a split cache can suffer more capacity misses than a unified cache of the same total size when one stream's working set is unusually large. Split caches trade a fixed, sometimes-unbalanced capacity allocation for conflict-free, parallel instruction/data access — the reason essentially all modern pipelined CPUs split L1 and unify only at L2/L3, where port contention no longer bottlenecks the pipeline.
  2. Part (b) — control unit inputs and outputs. The control unit sequences every other component of the datapath by generating the right control signals in the right cycle, based on what instruction is executing and the machine's current state.
    • Inputs:
      • Instruction register / opcode & mode bits — tells the control unit which operation and addressing mode to sequence.
      • Clock — provides the timing reference that advances the control unit through its sequence of states/micro-operations, one per cycle (or one per micro-instruction, if micro-programmed).
      • Flags / condition codes (zero, carry, sign, overflow) from the ALU — needed to resolve conditional branches and conditional micro-operations.
      • External/system signals — interrupt requests, bus-grant/bus-ready, reset — so the control unit can suspend normal sequencing to service an interrupt, wait for a slow bus transaction, or reinitialize the machine.
    • Outputs:
      • Register control signals (load/enable/clear on the PC, IR, general registers) — objective: move the right value into the right register at the right cycle.
      • ALU function-select lines — objective: tell the ALU which operation (add, AND, shift, …) to perform this cycle.
      • Memory read/write and bus-control signals — objective: initiate and time memory/bus transactions correctly, including asserting address/data onto the right bus in the right cycle.
      • Multiplexer/bus-select signals — objective: route the correct data path through shared buses so that only the intended source drives a shared line at a time.
    In one sentence: the control unit reads "what instruction, what state, what condition" and outputs "which register loads, which ALU op runs, which memory/bus transaction happens" — the objective of every signal is to make the datapath execute exactly the sequence of micro-operations the current instruction requires.
  3. Part (c) — relative bus performance. A bus of width $w$ bits moves an $s$-byte item in $\lceil 8s/w\rceil$ equal-duration bus cycles (an item that doesn't divide evenly into the bus width still consumes a whole extra cycle). Applying this to the given mix (30% one-byte, 70% two-byte items):
    Item size $s$FractionCycles, 8-bit busCycles, 16-bit bus
    1 byte (8 bits)30%$\lceil8/8\rceil=1$$\lceil8/16\rceil=1$
    2 bytes (16 bits)70%$\lceil16/8\rceil=2$$\lceil16/16\rceil=1$
    Weighting each column by its fraction gives the average number of bus cycles needed per fetched item: $$\bar c_{8}=0.30(1)+0.70(2)=\boxed{1.70\ \text{cycles}}$$ $$\bar c_{16}=0.30(1)+0.70(1)=\boxed{1.00\ \text{cycle}}$$ Since every bus cycle takes the same duration, throughput (items transferred per unit time) is inversely proportional to the average cycles/item, so the ratio of maximum transfer rates is the inverse ratio of these averages: $$\frac{\text{rate}_{16}}{\text{rate}_{8}}=\frac{\bar c_{8}}{\bar c_{16}}=\frac{1.70}{1.00}=\boxed{1.7}$$ The 16-bit bus achieves 1.7× the maximum data transfer rate of the 8-bit bus on this operand mix.
Final results — Question 1
PartResult
(a) unified vs. split cacheUnified: flexible capacity sharing, simpler, but port-contended. Split: parallel I/D access (pipeline-friendly), but fixed capacity split
(b) control unit I/OInputs: opcode/mode, clock, ALU flags, interrupts/bus signals. Outputs: register loads, ALU function select, memory/bus control, mux/bus select
(c) avg cycles/item8-bit bus: 1.70; 16-bit bus: 1.00
(c) transfer-rate factor16-bit bus is $\boxed{1.7\times}$ faster than the 8-bit bus
← Paper overview