Question 5 of 6: Multicycle CPU Performance Comparison
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Notes on this paper
98-Comp-A3, Computer Architecture — National Exams, May 2013. Open-book, 3 hours; six questions of equal value (20 marks each); FIVE constitute a complete exam (all six answered below as a complete study resource).
Reference texts: Patterson & Hennessy, Computer Organization and Design, 6th ed. — instruction encoding & the stored-program principle (Ch.2, Q1), memory addressing & data representation (Ch.2, Q2), cache organization & memory hierarchy (Ch.5, Q3 & Q6), procedure-call conventions & unsigned arithmetic (Ch.2–3, Q4), and CPU performance / the multicycle datapath (Ch.1 & Ch.4, Q5) — covering all six questions.
Question 5: Multicycle CPU Performance Comparison (20 marks)
Given. Cycle counts per instruction class: LI = 6 (original) / 5 (modified) cycles, ARITHMETIC = 5, MEMORY READ = 5, MEMORY STORE = 4, BRANCH = 4 cycles (unchanged by the modification). Instruction mix: LI 10%, ARITHMETIC 50%, MEMORY READ 20%, MEMORY STORE 10%, BRANCH 10%. The modified design's clock frequency is 5% lower than the original's.
Find. Which implementation executes the given instruction mix faster, and by what percentage.
Approach. Compute the average CPI for each design from the weighted cycle counts, convert to actual execution time per instruction by scaling for the modified design's slower clock (longer cycle time), then compare.
Average CPI, original design.
$$CPI_{orig}=0.10(6)+0.50(5)+0.20(5)+0.10(4)+0.10(4)=0.6+2.5+1.0+0.4+0.4=\boxed{4.9}\ \text{cycles/instr}.$$
Average CPI, modified design (same mix; only LI's own cycle count drops from 6 to 5).
$$CPI_{mod}=0.10(5)+0.50(5)+0.20(5)+0.10(4)+0.10(4)=0.5+2.5+1.0+0.4+0.4=\boxed{4.8}\ \text{cycles/instr}.$$
Convert CPI to execution time per instruction. Let $T$ be the original clock's cycle time. The modified clock runs at $0.95\times$ the original frequency, so its cycle time is longer, $T/0.95$:
$$t_{orig}=CPI_{orig}\times T=4.9T,\qquad t_{mod}=CPI_{mod}\times\frac{T}{0.95}=\frac{4.8}{0.95}T=5.0526\,T.$$
Compare. Since $4.9T\lt5.0526T$,
$$\boxed{t_{orig}\lt t_{mod}\ \Rightarrow\ \text{the ORIGINAL implementation executes programs faster.}}$$
The modified design takes $t_{mod}/t_{orig}=5.0526/4.9=1.0311$ times as long per instruction as the original — i.e. the original is about $\dfrac{t_{mod}-t_{orig}}{t_{mod}}\times100\%\approx3.0\%$ faster (equivalently, the modified design is about $\dfrac{t_{mod}-t_{orig}}{t_{orig}}\times100\%\approx3.1\%$ slower). Lowering LI's own CPI contribution (by $0.1$ cycle, weighted $10\%$) was not enough to offset the 5% clock-frequency penalty that applies to every instruction.