The Minds Behind Superscalar and Out-of-Order Execution – 7 People Redefining Architecture
Seven architects helped teach processors to discover parallelism dynamically while preserving the illusion of simple sequential execution.
TL;DR
Modern CPUs gain speed by doing far more than one instruction at a time. Tomasulo established a foundational dynamic-scheduling method; Patt, Smith, and Sohi expanded the theory and mechanisms of speculation and out-of-order execution; Rau explored compiler-directed parallelism; Glew helped commercialize aggressive dynamic execution in x86; Keller represents the whole-core integration needed to make such techniques fast in real products.[1][2][6]
Why you should read it anyway
Clock speed gets the headlines, but much of modern CPU performance comes from finding hidden parallelism in ordinary sequential programs. The processor predicts branches, renames registers, schedules ready work, tolerates cache misses, executes speculatively, and then retires results so the program still appears to have run in order. That illusion is one of computer architecture’s greatest engineering achievements.
Imagine where Superscalar and Out-of-Order Execution would be without them
Without these techniques, performance scaling would have depended much more heavily on frequency, compiler-visible parallelism, or explicit software concurrency. Power limits eventually made frequency scaling unsustainable, so the ability to exploit instruction-level parallelism inside a single thread became crucial for desktop, server, and mobile performance.
Time Estimate of how many years we would be hindered without them for human progress
Editorial counterfactual estimate: 5–10 years. Dynamic scheduling had multiple research paths, and commercial pressure for faster processors was intense. The likely delay would have been in assembling the full package—renaming, prediction, speculation, precise state, memory disambiguation, and recovery—into robust mass-market CPUs.
The 7 people behind Superscalar and Out-of-Order Execution
1. Robert Tomasulo
Why they matter: Tomasulo designed the dynamic scheduling mechanism used in IBM’s System/360 Model 91 floating-point unit, allowing instructions to proceed when their operands and execution resources were ready rather than strictly in original program order. The Model 91 became a landmark in out-of-order processing.[1] His algorithm’s reservation stations and tag-based dependency handling became foundational ideas for later dynamically scheduled processors.
2. Yale Patt
Why they matter: Patt’s HPS research in the 1980s explored a high-performance substrate that treated instruction execution as a dynamically scheduled dataflow problem beneath the ISA.[2] That work helped make aggressive out-of-order execution intellectually systematic rather than a collection of special cases. Patt’s influence is the idea that a processor can preserve sequential program semantics while internally finding parallel work far beyond the obvious instruction stream.
3. Bob Rau
Why they matter: Rau attacked instruction-level parallelism from a different direction. His work on VLIW and later EPIC emphasized exposing or scheduling parallel operations with substantial compiler participation rather than relying entirely on dynamic hardware discovery.[5] He belongs in this article because superscalar history is partly a debate about where scheduling intelligence should live: in compilers, in hardware, or in a hybrid of both.
4. Jim Smith
Why they matter: Smith contributed major concepts in branch prediction, precise interrupts, and superscalar architecture, including work associated with dynamically scheduled machines. The University of Illinois highlights his research impact on high-performance processor design.[3] Precise architectural state is especially important: speculative and out-of-order machines must appear to software as though instructions completed in a clean, recoverable order.
5. Gurindar Sohi
Why they matter: Sohi’s research helped formalize aggressive out-of-order execution, nonblocking memory systems, and speculative techniques in the mid-1980s and beyond. His Wisconsin research record documents a long program around high-performance microarchitecture.[4] His contribution is to make parallelism practical under real constraints: cache misses, exceptions, dependencies, and uncertain control flow.
6. Andy Glew
Why they matter: Glew worked on Intel’s P6 microarchitecture, the family underlying the Pentium Pro and successors, and has described mechanisms for dynamic execution, register renaming, and out-of-order scheduling.[6] His place in the story is commercialization: ideas developed across decades of research became the internal machinery of high-volume x86 processors while preserving the old instruction-set contract.
7. Jim Keller
Why they matter: Keller has repeatedly worked on high-performance processor families where wide issue, speculation, cache hierarchy, and out-of-order execution determine real performance. His CHM oral history covers work spanning DEC Alpha, AMD, and later CPU projects.[7] He represents the integration challenge: advanced scheduling only matters when front end, execution units, memory hierarchy, physical design, and power budgets converge in a shippable processor.
How they each differ from one another
Tomasulo provided an early algorithmic foundation for dynamic scheduling. Patt, Smith, and Sohi broadened the research into complete speculative microarchitectures. Rau explored a contrasting compiler-oriented route to instruction-level parallelism. Glew helped translate dynamic scheduling into Intel’s commercial P6 design, while Keller represents later whole-processor integration across high-performance CPU families. The differences show that “out of order” is not one trick but a coordinated system.
Final Take
A modern superscalar processor is a machine that constantly violates the apparent order of a program in order to preserve the appearance of that order faster. That paradox required decades of work on dependencies, prediction, exceptions, scheduling, and recovery. The achievement is not merely executing more instructions—it is doing so while keeping the software model stable.
Works Cited
- 01University of Michigan — Robert Tomasulo: Out-of-Order Processing and System/360 Model 91 leccap.engin.umich.edu
- 02Clemson — HPS: High Performance Substrate mark.people.clemson.edu
- 03University of Illinois — James E. Smith Alumni Award siebelschool.illinois.edu
- 04University of Wisconsin — Gurindar Sohi pages.cs.wisc.edu
- 05IEEE Computer Society — About B. Ramakrishna Rau computer.org
- 06Stanford EE380 — Andy Glew on P6 and Out-of-Order Execution web.stanford.edu
- 07Computer History Museum — Oral History of Jim Keller archive.computerhistory.org
CodeHistory is a living archive. Citations document the evidence used for this edition; later evidence may refine the account.
Submit a research lead