NivaarExam PrepOfficial exam papers ↗

25-Comp-A3 Computer Architecture · May 2014

Question 5 of 6: Multi-Cycle CPU Performance Comparison

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

98-Comp-A3, Computer Architecture — National Exams, May 2014 (paper header reads "December 2013"). Open-book, 3 hours; six questions of equal value (20 marks each); FIVE constitute a complete exam (all six answered below as a complete study resource).

Reference texts: Patterson & Hennessy, Computer Organization and Design, 6th ed. — instruction encoding & ISA compatibility, I/O (Q1), data representation & IEEE-754 floating point & array addressing (Q2), cache organization (Q3), pipelining & parallelism (Q4), CPU performance (Q5), and memory technology (Q6); Mano & Ciletti, Digital Design, 6th ed. — memory decoding and chip composition (Q6).

Question 5: Multi-Cycle CPU Performance Comparison (20 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

Given. Cycle counts: original LI=4, ARITHMETIC=3, MEMORY READ=5, MEMORY STORE=4, BRANCH=3 (modified LI=5, all others unchanged). Instruction mix: LI 10%, ARITHMETIC 50%, MEMORY READ 20%, MEMORY STORE 10%, BRANCH 10%. The modified design's clock frequency is 3% HIGHER than the original's.

Find. Which implementation executes the given mix faster, and by what percentage.

Approach. Compute each design's average CPI from the weighted cycle counts, convert to actual execution time by scaling for the modified design's faster clock (shorter cycle time), then compare.

  1. Average CPI, original design. $$CPI_{orig}=0.10(4)+0.50(3)+0.20(5)+0.10(4)+0.10(3)=0.4+1.5+1.0+0.4+0.3=\boxed{3.6}\ \text{cycles/instr}.$$
  2. Average CPI, modified design (only LI's own cycle count rises, from 4 to 5). $$CPI_{mod}=0.10(5)+0.50(3)+0.20(5)+0.10(4)+0.10(3)=0.5+1.5+1.0+0.4+0.3=\boxed{3.7}\ \text{cycles/instr}.$$
  3. Convert CPI to execution time per instruction. Let $T$ be the original clock's cycle time. The modified clock runs at $1.03\times$ the original frequency, so its cycle time is SHORTER, $T/1.03$: $$t_{orig}=CPI_{orig}\times T=3.6T,\qquad t_{mod}=CPI_{mod}\times\frac{T}{1.03}=\frac{3.7}{1.03}T=3.5922\,T.$$
  4. Compare. Since $3.5922T\lt3.6T$, $$\boxed{t_{mod}\lt t_{orig}\ \Rightarrow\ \text{the MODIFIED implementation executes programs faster.}}$$ The original design takes $t_{orig}/t_{mod}=3.6/3.5922=1.00216$ times as long per instruction as the modified one — the modified design is about $\dfrac{t_{orig}-t_{mod}}{t_{orig}}\times100\%\approx0.22\%$ faster (equivalently, the original is about $\dfrac{t_{orig}-t_{mod}}{t_{mod}}\times100\%\approx0.22\%$ slower). The 3% clock-frequency gain (applying to EVERY instruction) modestly outweighs the CPI penalty from LI needing one extra cycle (weighted only 10% of the mix), unlike a case where the CPI penalty is heavier or the clock gain smaller.
Final results — Question 5
QuantityOriginalModified
Average CPI3.63.7
Relative cycle time$T$$T/1.03=0.9709T$
Avg. execution time/instr.$3.6T$$3.5922T$
Faster implementationModified, by ≈0.22% (original is ≈0.22% slower)