Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Notes on this paper
3-hour, open-book exam. The NOTES state that FIVE (5) questions constitute a complete paper and the first five as answered will be marked; all SIX are answered here for completeness (a study resource). Reference texts: Patterson & Hennessy, Computer Organization and Design, 6th ed.
Given. (a) a pipelined processor implementation. (b) a GPU that is about 100× slower than a CPU on one multiplication, yet can be orders of magnitude faster overall on certain programs. (c) a dedicated hardware video encoder/decoder versus a software codec on a general-purpose processor.
Find. (a) why control-flow instructions specifically challenge a pipeline. (b) how/why the GPU can still win despite a slower single operation, and what properties the winning programs share. (c) the pros and cons of dedicated hardware versus software for this function.
Approach. reason from the pipeline's sequential-fetch assumption and where a branch outcome actually becomes known; distinguish per-operation LATENCY from aggregate THROUGHPUT across many parallel lanes; weigh fixed-function silicon against general-purpose flexibility.
Part (a) — control-flow instructions and pipelining. A pipelined fetch stage assumes, every cycle, that the next instruction lives at the next sequential address — but a taken branch or jump changes the next-PC to a target that is normally only known once the instruction reaches a LATER stage (decode or execute), several cycles after it was fetched. Every instruction fetched during that resolution window, on the "sequential" assumption, may have to be squashed (flushed) once the real outcome/target is known, wasting exactly as many cycles as how deep in the pipeline the branch resolves. This is a control hazard, distinct from structural and data hazards, and its penalty scales both with how LATE the outcome is known (deeper pipelines suffer more per misprediction) and with how OFTEN branches occur in typical code. Control-flow instructions break the pipeline's "always fetch sequentially" assumption; every cycle before the real next-PC is known risks fetching instructions from the wrong path that must later be flushed.
Part (b) — GPU throughput vs. CPU latency. A single GPU lane really is slower AND simpler than a CPU core (lower clock speed, little or no out-of-order execution or branch prediction hardware), which explains the roughly 100× per-operation latency gap for one multiplication. But a GPU packs THOUSANDS of such lanes executing the SAME instruction stream on different data in lockstep (SIMT), against a CPU's handful of wide, independent cores — so what matters for overall program time is aggregate THROUGHPUT (operations completed per second summed across every lane), not the latency of any single operation, PROVIDED the workload can keep nearly all those lanes genuinely busy. The properties that make a program win big on a GPU are: massive DATA PARALLELISM (the identical operation applied independently across a huge number of data elements, e.g. every pixel or matrix entry); high arithmetic intensity (enough compute per byte moved to hide memory latency); regular, coalesced memory access patterns; minimal thread divergence (little data-dependent branching, since divergent threads within a lock-step group serialize); and a problem large enough to amortize the fixed cost of launching work on the device. Throughput across thousands of simple lanes, not the latency of any one lane, decides — a program needs massive, regular, low-divergence data parallelism to expose that throughput advantage.
Part (c) — dedicated vs. software video codec. A dedicated encoder/decoder is a datapath purpose-built for one algorithm's operations (motion estimation, transform/quantization, entropy coding), with no instruction fetch/decode overhead per operation and a hardwired dataflow — this buys far higher throughput per watt and lower, more DETERMINISTIC latency than the same algorithm executed as a long instruction sequence on a general CPU, which matters for real-time constraints (video conferencing, broadcast) and for battery-powered devices. The cost is inflexibility: a fixed-function block generally needs new silicon (or, at best, a narrow range of firmware-programmable options) to support a new codec, a new profile, or even a bug fix, so its non-recurring engineering cost and time-to-market are higher, and it becomes useless the moment a workload needs an algorithm it was not built for. A software implementation is the mirror image — fully flexible and field-upgradable (new codecs, patches, multiple formats on the same hardware) — but pays with higher power draw, a lower achievable resolution/frame-rate at a given cost point, and less deterministic real-time behaviour since it competes with everything else running on the same general-purpose core. Dedicated hardware trades flexibility for throughput-per-watt and deterministic real-time latency; software trades throughput-per-watt for the ability to change or add codecs without new silicon.
Final results — Question 4
Part
Result
(a)
Control hazard: pipeline fetches sequentially before the real next-PC is known, forcing flushes on mispredicted control flow
(b)
GPU wins via aggregate throughput (thousands of slower lanes) on massively data-parallel, regular, low-divergence workloads