Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Notes on this paper
3-hour, open-book exam. The NOTES state that FIVE (5) questions constitute a complete paper and the first five as answered will be marked; all SIX are answered here for completeness (a study resource). Reference texts: Patterson & Hennessy, Computer Organization and Design, 6th ed.
Given. (a) a 4GB ($2^{32}$-byte) byte-addressable space; a 32KB, 4-way set-associative cache with 64-byte blocks. (b) the same cache; a block resident in set 0x2, tagged 0x11. (c) n/a — conceptual comparison of caches vs. a larger register file.
Find. (a) how a 32-bit physical address splits into tag/index/offset fields for this cache, with justification. (b) the full hexadecimal address range the described block spans. (c) why caches are used, and the pros/cons of using more registers instead.
Approach. (a) derive the number of sets from capacity/(ways×block size), then take base-2 logs to size each field. (b) reconstruct the block's base address by placing the tag, then the set index, then a zero offset, and add block-size−1 for the top of the range. (c) reason from the CPU–DRAM speed gap and from the ISA-encoding constraint on register-file size.
Part (a) — indexing the cache. The number of sets is the cache capacity divided by how many bytes one full "row" across all ways occupies:
$$\text{sets}=\frac{32\,\text{KB}}{4\times64\,\text{B}}=\frac{32768}{256}=128$$
A 64-byte block needs $\log_2 64=6$ offset bits to select a byte within it; 128 sets need $\log_2128=7$ index bits to select a set; the remaining $32-7-6=19$ bits form the tag that disambiguates which of the (many) blocks mapping to that set is currently resident. 32-bit address = 19-bit tag | 7-bit index | 6-bit block offset; only the number of SETS (not the number of blocks) sizes the index, because associativity lets 4 blocks share one set.
Part (b) — address range of the cached block. Reassembling a block's base address places the tag above the index, and the index above the offset (which is zero at the block's first byte):
$$\text{low}=(\text{0x11}\ll(7+6))\ |\ (\text{0x2}\ll6)=(\text{0x11}\ll13)\ |\ (\text{0x2}\ll6)$$
Substituting, $\text{0x11}\ll13=\text{0x22000}$ and $\text{0x2}\ll6=\text{0x80}$, giving a base of $\text{0x22000}+\text{0x80}=\text{0x22080}$. The block spans 64 bytes from that base:
$$\boxed{\text{0x22080}\ \text{to}\ \text{0x22080}+63=\text{0x220BF}}$$
Part (c) — why caches, and registers instead? Modern processors use caches because DRAM latency (on the order of 100+ cycles) is vastly larger than a CPU cycle time, and most programs exhibit temporal and spatial locality — recently or nearby-accessed data is disproportionately likely to be accessed again soon. A cache exploits that locality TRANSPARENTLY, with no software or ISA visibility, keeping the AVERAGE memory access time close to the fast SRAM latency without paying DRAM's latency on every access. Registers are even faster than any cache level, but growing the register file instead has real costs: the number of ISA-visible registers is capped by how many bits the instruction encoding devotes to naming them (the very same operand-field-width constraint from Question 1 — adding registers means widening those fields, which strains or breaks backward compatibility); register allocation is a compile-time decision, so values live only as long as the compiler keeps them resident, with no automatic analogue of a cache silently retaining ANY recently touched address; and a larger register file also grows the fixed cost of every context switch (more state to save/restore). A cache, by contrast, scales to arbitrarily large working sets bounded only by its own capacity (not by an operand-field width) and needs no compiler cooperation, but it benefits ANY code exhibiting locality, including data shared across procedure or thread boundaries that a register allocation cannot span. Caches transparently bridge the CPU–DRAM speed gap via locality; more registers would help too, but are capped by ISA encoding width and require compiler-level management rather than being automatic.
Final results — Question 3
Part
Result
(a)
128 sets $\Rightarrow$ 19-bit tag | 7-bit index | 6-bit offset
(b)
$\boxed{\text{0x22080}\text{--}\text{0x220BF}}$
(c)
Caches transparently exploit locality to bridge the CPU–DRAM gap; more registers help too but are capped by ISA-encoding width and need compiler management