Question 4 of 6: Pipelining, Parallel Execution Strategies, and Hardware Specialization
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Notes on this paper
98-Comp-A3, Computer Architecture — National Exams, May 2014 (paper header reads "December 2013"). Open-book, 3 hours; six questions of equal value (20 marks each); FIVE constitute a complete exam (all six answered below as a complete study resource).
Reference texts: Patterson & Hennessy, Computer Organization and Design, 6th ed. — instruction encoding & ISA compatibility, I/O (Q1), data representation & IEEE-754 floating point & array addressing (Q2), cache organization (Q3), pipelining & parallelism (Q4), CPU performance (Q5), and memory technology (Q6); Mano & Ciletti, Digital Design, 6th ed. — memory decoding and chip composition (Q6).
Given. A pipelined processor executing a mixed instruction stream that includes control-flow (branch/jump) instructions; two competing N-way parallel-execution strategies (superscalar issue vs. multi-core); the choice between a dedicated video codec ASIC and a software codec on a general-purpose CPU.
Find. (a) why control-flow instructions are hard for a pipeline. (b) performance/cost/complexity comparison of N-way superscalar vs. N-way multi-core. (c) why build a dedicated video encoder/decoder, and its tradeoffs against software.
Approach. All three parts are argued from first principles — pipeline-hazard theory for (a), the parallelism-granularity/hardware-duplication tradeoff for (b), and the specialization-vs-generality tradeoff (ASIC vs. programmable CPU) for (c).
Part (a) — why control flow is a pipeline hazard. A pipeline fetches instructions speculatively, one per stage-cycle, based on the assumption that the NEXT instruction is simply the one at $PC+\text{size}$. A branch or jump breaks that assumption: the pipeline cannot know the true next-instruction address (and, for a conditional branch, cannot even know WHETHER to take it) until the branch instruction has been fetched, decoded, and often evaluated several stages later. Every instruction fetched behind it in the meantime is fetched on a guess; if that guess is wrong, all of those partially-executed instructions must be discarded (a pipeline flush), wasting exactly as many cycles as the branch's resolution latency (its position in the pipeline). This is a control hazard, distinct from data hazards (which stall for a value) and structural hazards (which stall for a shared resource) — it stalls for KNOWLEDGE of where to fetch next. Real pipelines mitigate it with branch prediction (guess and roll back only on misprediction), delayed branches, or reordering to fill the delay slots with useful independent work, but none of these eliminate the hazard, only reduce its average cost.
Part (b) — N-way superscalar vs. N-way multi-core.Superscalar keeps ONE instruction stream (one program counter, one thread of control) and finds N independent, adjacent instructions within it each cycle to issue together on N redundant execution units sharing one register file and one cache hierarchy. Multi-core instead duplicates the ENTIRE processor N times — each core has its own PC, register file, and (typically) private L1 cache — running N genuinely separate instruction streams.
Superscalar vs. multi-core comparison
Aspect
N-way superscalar
N-way multi-core
Performance source
Instruction-level parallelism (ILP) found WITHIN one program; speedup limited by how much adjacent independent work that one stream actually contains
Task/thread-level parallelism ACROSS separate programs or threads; speedup limited by how much of the workload can be split into independent tasks
Single-thread benefit
Speeds up even a single sequential program automatically (hardware finds the parallelism)
Gives a single sequential program NO speedup at all — only one core is used unless the program is explicitly multithreaded
Hardware cost
One register file, one cache, but complex issue/dependency-check logic that grows faster than linearly with $N$ (checking all pairs of candidate instructions for dependencies)
Cost grows roughly LINEARLY with $N$ (N copies of a simpler, already-designed core), though shared last-level cache/interconnect adds some overhead
Design complexity
High — out-of-order issue, register renaming, and hazard detection across N candidate instructions per cycle is intricate and hard to verify
Lower per-core (each core can be a simple, well-understood design); complexity shifts to inter-core communication, cache coherence, and synchronization
In short: superscalar buys automatic speedup for legacy single-threaded code at high hardware/verification cost and diminishing returns (real programs rarely expose much adjacent independent work); multi-core buys speedup that scales more predictably with transistor budget but ONLY for software that is explicitly parallelized, and adds its own coherence/synchronization complexity.
Part (c) — dedicated video codec vs. software. Video encode/decode (e.g. motion estimation, DCT/transform, entropy coding) is extremely computation- and data-movement-heavy but also highly regular and repetitive — the same handful of arithmetic patterns applied to millions of pixels per second. A general-purpose processor pays the FULL overhead of a programmable pipeline (instruction fetch/decode, branch prediction, out-of-order scheduling) for every one of those repetitive operations, none of which that overhead is actually needed for. Pros of a dedicated ASIC/accelerator: far higher throughput and dramatically lower energy per frame (no fetch/decode overhead, and the datapath can be laid out exactly for the codec's fixed operations — often 10–100× more energy-efficient than software on a general CPU for the same task); it also frees the general-purpose CPU entirely to run other work concurrently. Cons: it is inflexible — fixed in silicon to one (or a few) codec standards, so it cannot adapt to a new/updated codec without a hardware respin (or at best limited configurability), it adds die area and non-recurring engineering cost even when video isn't being used, and a software implementation is trivially updated (patch the algorithm) whereas a hardware bug or a new standard may require a whole new chip revision. The dedicated encoder is chosen when the workload is common, high-volume, and standardized enough (e.g. shipped in nearly every device) to amortize its fixed design cost against the large recurring energy/throughput win — exactly the specialization-vs-generality tradeoff that also motivates GPUs and other fixed-function accelerators.
Final results — Question 4
Part
Result
(a) branches in a pipeline
Control hazard: next-fetch address/outcome unknown until the branch resolves several stages later → mispredicted work must be flushed