Computer Architecture
Kaup valmöguleikar
Computer Architecture: A Quantitative Approach, has been considered essential reading by instructors, students and practitioners of computer design for nearly 30 years. The seventh edition of this classic textbook from John Hennessy and David Patterson, wWinner of a 2019 Textbook Excellence Award (Texty) from the Textbook and Academic Authors AssociationEach chapter follows a consistent framework: explanation of the ideas in each chapter; a "crosscutting issues" section, which presents how the concepts covered in one chapter connect with those given in other chapters; a "putting it all together" section that links these concepts by discussing how they are applied in real machine; and detailed examples of misunderstandings and architectural traps commonly encountered by developers and architectsIncludes "Putting It All Together" sections near the end of every chapter, providing real-world technology examples that demonstrate the principles covered in each chapterCovers new developments in GPU and CPU architectures, as well as domain specific architecturesFeatures more comprehensive coverage of systems on chip and heterogeneity.
Nánar um bókina
- Elsevier S & T
- 9780443154072
- 9780443154065
- ePub
- 7
- John Hennessy
- English
- 2025-12-02
- 10
- 2
- 2
Kaflar
- Title of Book
- Cover image
- Title page
- Table of Contents
- Biography
- Copyright
- Acknowledgments
- Contributors to the Seventh
- Reviewers
- Appendices
- Case Studies With Exercises
- Additional Material
- Contributors to Previous Editions
- Reviewers
- Appendices
- Exercises
- Case Studies With Exercises
- Additional Material
- Special Thanks
- 1. Fundamentals of Quantitative Design and Analysis
- 1.1 Introduction
- 1.2 Classes of Computers
- Internet of Things/Embedded Computers
- Personal Mobile Device
- Desktop Computing
- Servers
- Clusters/Warehouse-Scale Computers
- Classes of Parallelism and Parallel Architectures
- 1.3 Defining Computer Architecture
- Instruction Set Architecture: The Myopic View of Computer Architecture
- Genuine Computer Architecture: Designing the Organization and Hardware to Meet Goals and Functional Requirements
- 1.4 Trends in Technology
- Performance Trends: Bandwidth Over Latency
- Scaling of Transistor Performance and Wires
- 1.5 Trends in Power and Energy in Integrated Circuits
- Power and Energy: A Systems Perspective
- Energy and Power Within a Microprocessor
- Example
- Answer
- The Shift in Computer Architecture Because of Limits of Energy
- The Impact of Energy Use by Computers on Global Climate Change
- 1.6 Trends in Cost
- The Impact of Time, Volume, and Commoditization
- Cost of an Integrated Circuit
- Example
- Answer
- Example
- Answer
- Chiplets
- Cost Versus Price
- Cost of Manufacturing Versus Cost of Operation
- 1.7 Dependability
- Example
- Answer
- Example
- Answer
- 1.8 Security
- 1.9 Measuring, Reporting, and Summarizing Performance
- Benchmarks
- Desktop Benchmarks
- Server Benchmarks
- Reporting Performance Results
- Summarizing Performance Results
- Example
- Answer
- 1.10 Quantitative Principles of Computer Design
- Take Advantage of Parallelism
- Take Advantage of Pipelining
- Take Advantage of Prediction
- Take Advantage of the Principle of Locality
- Deliver Dependability via Redundancy
- Focus on the Common Case
- Amdahl’s Law
- Example
- Answer
- Example
- Answer
- Example
- Answer
- The Processor Performance Equation
- Example
- Answer
- 1.11 Putting It All Together: Performance, Price, and Power
- 1.12 Fallacies and Pitfalls
- 1.13 Concluding Remarks
- 1.14 Historical Perspectives and References
- Case Studies and Exercises by Gregory D. Peterson
- Case Study 1: Chip Fabrication Trends
- Concepts illustrated by this case study
- Case Study 2: Scalable Computer Systems
- Concepts illustrated by this case study
- Case Study 3: Energy and Power Considerations for Computer Systems
- Concepts illustrated by this case study
- 2. Memory Hierarchy Design
- 2.1 Introduction
- Basics of Memory Hierarchies: A Quick Review
- 2.2 Memory Technology and Optimizations
- SRAM Technology
- DRAM Technology
- Improving Memory Performance Inside a DRAM Chip: SDRAMs
- Reducing Power Consumption in SDRAMs
- Packaging Innovation: Stacked or Embedded DRAMs or SRAMs
- Flash Memory
- Phase-Change Memory Technology
- Enhancing Dependability in Memory Systems
- 2.3 Ten Advanced Performance Optimizations for Memory Hierarchies
- First Optimization: Pipelined L1 Caches With Virtual Indexing and Set Associativity
- Example
- Answer
- Second Optimization: Multiple Banks and Ports to Increase the Bandwidth of L1 D-Caches
- Example
- Answer
- Example
- Answer
- Third Optimization: Reducing the Miss Rate With Better Replacement Policies
- Fourth Optimization: Multibanked L2 and L3 Caches to Decrease Power and Latency, and Increase Bandwidth
- Fifth Optimization: Nonblocking Caches to Increase Cache Bandwidth
- Example
- Answer
- Implementing a Nonblocking Cache
- Sixth Optimization: Critical Word First and Early Restart to Reduce Miss Penalty
- Seventh Optimization: Compiler Optimizations to Reduce Miss Rate
- Loop Interchange
- Blocking
- Eighth Optimization: Hardware Prefetching of Instructions and Data to Reduce Miss Penalty or Miss Rate
- Ninth Optimization: Compiler-Controlled Prefetching to Reduce Miss Penalty or Miss Rate
- Example
- Answer
- Example
- Answer
- Tenth Optimization: Multiple Memory Buses and Memory Modules to Increase Bandwidth
- Memory Hierarchy Optimization Summary
- 2.4 Virtual Memory and Protection
- Protection via Virtual Memory
- Side-Channel Attacks on the Memory System
- 2.5 Cross-Cutting Issues: The Design of Memory Hierarchies
- Autonomous Instruction Fetch Units
- Special Instruction Caches
- Speculation and Memory Access
- Coherency of Cached Data
- Protection via Virtual Machines
- The ARM Cortex-A53
- Performance of the Cortex-A53 Memory Hierarchy
- Using these miss penalties
- The Intel Core i9
- 2.7 Fallacies and Pitfalls
- 2.8 Concluding Remarks: Looking Ahead
- 2.9 Historical Perspectives and References
- Case Studies and Exercises by Rajeev Balasubramonian, Norman P. Jouppi, Naveen Muralimanohar, and Sheng Li
- Case Study 1: Optimizing Cache Performance via Advanced Techniques
- Concepts illustrated by this case study
- Case Study 2: Putting It All Together: Highly Parallel Memory Systems
- Concepts illustrated by this case study
- Case Study 3: Studying the Impact of Various Memory System Organizations
- Concepts illustrated by this case study
- 3. Instruction-Level Parallelism and Its Exploitation
- 3.1 Introduction
- What Is Instruction-Level Parallelism?
- Data Dependences and Hazards
- Data Dependences
- Name Dependences
- Data Hazards
- Control Dependences
- 3.2 Basic Compiler Techniques for Exposing ILP
- Basic Pipeline Scheduling and Loop Unrolling
- Example
- Answer
- Example
- Answer
- Example
- Answer
- Summary of Loop Unrolling and Scheduling
- 3.3 Key ILP Concepts in Modern Superscalar Processors
- Deeper Pipelining and Multiple Issue
- Control dependences → Branch prediction and speculative execution
- Name dependences → Register renaming
- Data dependences → OOO execution with dynamic scheduling
- Memory dependences → speculative memory disambiguation
- Speculation Recovery and Precise Exceptions → Reorder Buffer
- Superscalar Processor Overview
- 3.4 Reducing Branch Costs with Advanced Branch Prediction
- Correlating Branch Predictors
- Example
- Answer
- Tournament Predictors: Adaptively Combining Local and Global Predictors
- Tagged Hybrid Predictors
- The Evolution of the Branch Predictor in Intel Processors
- Branch-Target Buffers
- Example
- Answer
- Specialized Branch Predictors: Predicting Procedure Returns, Indirect Jumps, and Loop Branches
- Integrated Instruction Fetch Units
- 3.5 Overcoming Name Dependences with Register Renaming
- Alternative Implementations for Register Renaming
- 3.6 Overcoming Data Hazards with Dynamic Scheduling
- Dynamic Scheduling: The idea
- Dynamic Scheduling: Examples
- Example
- Answer
- Example
- Answer
- Dynamic Scheduling: The Details
- 3.7 Overcoming Memory Dependences with Dynamic Disambiguation
- Speculative Memory Disambiguation
- 3.8 Advanced Issues in Modern Superscalar Processors
- The Challenge of Multiple Issues per Cycle
- How Much to Speculate
- Speculating Through Multiple Branches
- Speculation and the Challenge of Energy Efficiency
- What Limits Superscalar Processors
- 3.9 Exploiting ILP Using Multiple Issue and Static Scheduling
- The Basic VLIW Approach
- Example
- Answer
- Multiple Issue in Perspective
- 3.10 Cross-Cutting Issues
- Hardware Versus Software Speculation
- Speculative Execution and the Memory System
- 3.11 Multithreading: Exploiting Thread-Level Parallelism to Improve Single Core Throughput
- Effectiveness of Simultaneous Multithreading on Superscalar Processors
- 3.12 Microarchitecture Side-Channel Attacks
- 3.13 Putting It All Together: The Arm Cortex-A53 and the Intel Golden Cove Processors
- The Arm Cortex-A53
- Performance of the A53 Pipeline
- The Intel Golden Cove
- Performance of the Golden Cove Core
- 3.14 Fallacies and Pitfalls
- 3.15 Concluding Remarks
- 3.16 Historical Perspective and References
- Case Studies and Exercises by Jason D. Bakos
- Case Study: Dynamic Scheduling
- 4. Data-Level Parallelism in Vector, SIMD, and GPU Architectures
- 4.1 Introduction
- 4.2 Vector Architecture
- RV64V Extension
- How Vector Processors Work: An Example
- Example
- Answer
- Vector Execution Time
- Example
- Answer
- Multiple Lanes: Beyond One Element per Clock Cycle
- Vector-Length Registers: Handling Loops Not Equal to 32
- Mask Registers: Handling IF Statements in Vector Loops
- Memory Banks: Supplying Bandwidth for Vector Load/Store Units
- Example
- Answer
- Stride: Handling Multidimensional Arrays in Vector Architectures
- Example
- Answer
- Gather-Scatter: Handling Sparse Matrices in Vector Architectures
- Programming Vector Architectures
- 4.3 SIMD Instruction Set Extensions for Multimedia
- Example
- Answer
- Programming Multimedia SIMD Architectures
- The Roofline Visual Performance Model
- Similarities and Differences Between Vector Architectures and Multimedia SIMD Computer
- 4.4 Graphics Processing Units
- Programming the GPU
- NVIDIA GPU Computational Structures
- NVIDIA GPU Instruction Set Architecture
- Conditional Branching in GPUs
- NVIDIA GPU Memory Structures
- Innovations in the Recent GPU Architectures
- Similarities and Differences Between Vector Architectures and GPUs
- Similarities and Differences Between Multimedia SIMD Computers and GPUs
- Summary
- 4.5 Detecting and Enhancing Loop-Level Parallelism
- Example
- Answer
- Example
- Answer
- Finding Dependences
- Example
- Answer
- Example
- Answer
- Eliminating Dependent Computations
- 4.6 Cross-Cutting Issues
- Energy and DLP: Slow and Wide Versus Fast and Narrow
- Banked Memory and High Bandwidth Memory
- Strided Accesses and TLB Misses
- 4.7 Putting It All Together: Embedded Versus Server GPUs and Tesla Versus Core i7
- Comparison of a GPU and a MIMD With Multimedia SIMD
- Comparison Update
- 4.8 Fallacies and Pitfalls
- 4.9 Concluding Remarks
- 4.10 Historical Perspective and References
- Case Study and Exercises by Jason D. Bakos
- GPU vs Vector Processor for Low-Arithmetic-Intensity Machine Learning Tasks
- Concepts illustrated by this case study
- 5. Thread-Level Parallelism
- 5.1 Introduction
- Multiprocessor Architecture: Issues and Approach
- Challenges of Parallel Processing
- Example
- Answer
- Example
- Answer
- Example
- Answer
- 5.2 Multiprocessor Cache Coherence
- Basic Schemes for Enforcing Coherence
- Ensuring Coherence: Invalidate versus Update
- Cache Coherence Protocols
- 5.3 Maintaining Cache Coherence with Snooping
- An Example Protocol
- Extensions to the Basic Coherence Protocol
- Limitations in Symmetric Shared-Memory Multiprocessors and Snooping Protocols
- Example
- Answer
- Implementing Snooping Cache Coherence
- Performance of a Shared-Memory Multiprocessor
- Example
- Answer
- A Commercial Workload on a Small-Scale Multiprocessor
- 5.4 Maintaining Cache Coherence with Directories
- Directory-Based Cache Coherence Protocols: The Basics
- An Example Directory Protocol
- Implementing Directory Coherence in a Multicore
- Performance of a NUMA Multiprocessor on a WWW Search Application
- 5.5 Synchronization: The Basics
- Basic Hardware Primitives
- Implementing Locks Using Coherence
- 5.6 Models of Memory Consistency: An Introduction
- Example
- Answer
- The Programmer’s View
- Relaxed Consistency Models: The Basics and Release Consistency
- 5.7 Cross-Cutting Issues
- Compiler Optimization and the Consistency Model
- Using Speculation to Hide Latency in Strict Consistency Models
- Inclusion and Its Implementation
- Example
- Answer
- Multiprocessor Performance Gains from Multithreading
- 5.8 Putting It All Together: Multicore Processors and Their Performance
- Three Multicore Server Processors
- Performance of Multicore-Based Multiprocessors on a Multiprogrammed Workload
- Scalability of a Xeon Platinum on Two Different Workloads
- 5.9 Fallacies and Pitfalls
- 5.10 The Future of Multicore Scaling
- Example
- Answer
- The Future of Multicore Is Heterogeneity
- 5.11 Concluding Remarks
- 5.12 Historical Perspectives and References
- Case Studies and Exercises by Amr Zaky
- Case Study 1: Single Chip Multicore Multiprocessor
- Concepts illustrated by this case study.
- Transaction Notation
- Caches Configuration/Policy
- Coherence Protocol
- Memory Distribution and Addressing
- Case Study 2: Simple Directory-Based Coherence
- Concepts illustrated by this case study.
- Coherence Protocol
- Memory Distribution and Addressing
- Transaction Notation
- Caches Configuration/Policy
- Case Study 3: Memory Consistency
- Concepts illustrated by this case study.
- 6. Warehouse-Scale Architectures for Utility Computing
- 6.1 Introduction
- Example
- Answer
- 6.2 Cloud Computing: The Return of Utility Computing
- Cloud Computing Models
- A Closer Look at IaaS Cloud Services
- Cloud Metrics and WSC Architecture Drivers
- 6.3 Hardware Support for Virtualization
- What Makes an Architecture Virtualizable?
- Architectural Support for Virtualization
- Memory Virtualization
- Example
- Answer
- I/O Virtualization
- An Example Hypervisor: The Xen Project
- Confidential Computing and Secure Enclaves
- 6.4 Computer Architecture of Warehouse-Scale Computers
- WSC Compute
- WSC Storage
- WSC Networking
- Example
- Answer
- The Programmer’s View of a WSC
- Example
- Answer
- Example
- Answer
- 6.5 The Architecture of High-Performance I/O Devices
- Network Interfaces
- NVMe Flash Drives
- I/O Device Virtualization
- 6.6 WSC Power Distribution and Cooling
- WSC Power Delivery
- WSC Cooling
- 6.7 The Cost and Efficiency of Warehouse-scale Computing
- Cost of a WSC
- Example
- Answer
- Example
- Answer
- Efficiency Through Power Provisioning
- Efficiency Through Higher Utilization
- Efficiency Through Higher Performance
- Efficiency Through Energy Proportionality
- 6.8 Putting It All Together: Custom Silicon in the AWS Cloud
- The Nitro System
- Nitro Security
- The Graviton CPU Chips
- CPU Efficiency for Cloud Computing
- 6.9 Fallacies and Pitfalls
- 6.9 Concluding Remarks
- 6.10 Historical Perspectives and References
- Case Studies and Exercises by Parthasarathy Ranganathan
- Case Study 1: Total Cost of Ownership Influencing Warehouse-Scale Computer Design Decisions
- Concepts illustrated by this case study
- Case Study 2: Resource Allocation in WSCs and TCO
- Concepts illustrated by this case study
- Exercises
- 7. Domain-Specific Architectures
- 7.1 Introduction
- 7.2 Guidelines for DSAs
- 7.3 Example Domain: Deep Neural Networks
- The Neurons of DNNs
- Training Versus Inference
- Multilayer Perceptron
- Convolutional Neural Network
- Recurrent Neural Network
- Transformer Neural Network
- Batches
- Quantization
- Summary of DNNs
- 7.4 Google’s Tensor Processing Unit v4 and v4 lite, Data Center DNN Accelerators
- TPU v4 lite Architecture
- TPU v4 Lite Instruction Set Architecture
- Changes for TPU v4 Versus TPU v4 lite
- Summary: How TPU v4 Follows the Guidelines
- 7.5 The NVIDIA A100 and T4 GPUs, Graphics, and DNN Accelerators for the Data Center
- NVIDIA A100 Tensor Core Architecture
- Summary: How A100 and T4 Follow the Guidelines
- 7.6 Graphcore IPU Bow, a Data Center Accelerator for Training
- Summary: How the IPU Bow Follows the Guidelines
- 7.7 The Samsung Neural Processing Unit (NPU), a Smartphone Inference Accelerator
- Summary: How the Samsung NPU Follows the Guidelines
- 7.8 Cross-Cutting Issues
- Heterogeneity and System on a Chip
- Energy and Carbon Emissions of Machine Learning Training
- An Open Instruction Set
- 7.9 Putting It All Together: Comparing DNN Accelerators
- Inference: TPU v4 lite Versus T4
- Training: TPU v4 Versus A100 Versus IPU Bow
- Inference: Samsung NPUv1 and NPUv2 Versus Intel Core i7 plus Iris Xe GPU
- 7.10 Fallacies and Pitfalls
- 7.11 Concluding Remarks
- 7.12 Historical Perspectives and References
- Case Studies and Exercises by Cliff Young
- Case Study 1: Matrix Multiplication
- Concepts illustrated by this case study
- Decompositions
- Dot or Inner Product
- SAXPY
- Outer Product
- Case Study 2: Neural Network Layers
- Concepts illustrated by this case study
- Case Study 3: Numerical Representations
- Concepts illustrated by this case study
- Integers
- Floating Point
- Case Study 4: Speeds, Feeds, and Scale
- Concepts illustrated by this case study
- Exercises
- A. Instruction Set Principles
- A.1 Introduction
- A.2 Classifying Instruction Set Architectures
- Summary: Classifying Instruction Set Architectures
- A.3 Memory Addressing
- Interpreting Memory Addresses
- Addressing Modes
- Displacement Addressing Mode
- Immediate or Literal Addressing Mode
- Summary: Memory Addressing
- A.4 Type and Size of Operands
- A.5 Operations in the Instruction Set
- A.6 Instructions for Control Flow
- Addressing Modes for Control Flow Instructions
- Conditional Branch Options
- Procedure Invocation Options
- Summary: Instructions for Control Flow
- A.7 Encoding an Instruction Set
- Reduced Code Size in RISCs
- Summary: Encoding an Instruction Set
- A.8 Cross-Cutting Issues: The Role of Compilers
- The Structure of Recent Compilers
- Register Allocation
- Impact of Optimizations on Performance
- The Impact of Compiler Technology on the Architect’s Decisions
- How the Architect Can Help the Compiler Writer
- Compiler Support (or Lack Thereof) for Multimedia Instructions
- Example
- Summary: The Role of Compilers
- RISC-V Instruction Set Organization
- Registers for RISC-V
- Data Types for RISC-V
- Addressing Modes for RISC-V Data Transfers
- RISC-V Instruction Format
- RISC-V Operations
- RISC-V Control Flow Instructions
- RISC-V Floating-Point Operations
- RISC-V Instruction Set Usage
- A.11 Concluding Remarks
- A.12 Historical Perspective and References
- Exercises by Gregory D. Peterson
- B. Review of Memory Hierarchy
- B.1 Introduction
- Cache Performance Review
- Example
- Answer
- Example
- Answer
- Four Memory Hierarchy Questions
- Q1: Where Can a Block be Placed in a Cache?
- Q2: How Is a Block Found If It Is in the Cache?
- Q3: Which Block Should be Replaced on a Cache Miss?
- Q4: What Happens on a Write?
- Example
- Answer
- An Example: The Opteron Data Cache
- Example
- Answer
- Average Memory Access Time and Processor Performance
- Example
- Answer
- Example
- Answer
- Miss Penalty and Out-of-Order Execution Processors
- Example
- Answer
- First Optimization: Larger Block Size to Reduce Miss Rate
- Example
- Answer
- Second Optimization: Larger Caches to Reduce Miss Rate
- Third Optimization: Higher Associativity to Reduce Miss Rate
- Example
- Answer
- Fourth Optimization: Multilevel Caches to Reduce Miss Penalty
- Example
- Answer
- Example
- Answer
- Fifth Optimization: Giving Priority to Read Misses over Writes to Reduce Miss Penalty
- Example
- Answer
- Sixth Optimization: Avoiding Address Translation During Indexing of the Cache to Reduce Hit Time
- Summary of Basic Cache Optimization
- Four Memory Hierarchy Questions Revisited
- Q1: Where Can a Block be Placed in Main Memory?
- Q2: How Is a Block Found If It Is in Main Memory?
- Q3: Which Block Should be Replaced on a Virtual Memory Miss?
- Q4: What Happens on a Write?
- Techniques for Fast Address Translation
- Selecting a Page Size
- Summary of Virtual Memory and Caches
- Protecting Processes
- A Segmented Virtual Memory Example: Protection in the Intel Pentium
- Adding Bounds Checking and Memory Mapping
- Adding Sharing and Protection
- Adding Safe Calls from User to OS Gates and Inheriting Protection Level for Parameters
- A Paged Virtual Memory Example: The 64-Bit Opteron Memory Management
- Summary: Protection on the 32-Bit Intel Pentium Versus the 64-Bit AMD Opteron
- C. Pipelining: Basic and Intermediate Concepts
- C.1 Introduction
- What Is Pipelining?
- The Basics of the RISC V Instruction Set
- A Simple Implementation of a RISC Instruction Set
- The Classic Five-Stage Pipeline for a RISC Processor
- Basic Performance Issues in Pipelining
- Example
- Answer
- Performance of Pipelines With Stalls
- Data Hazards
- Minimizing Data Hazard Stalls by Forwarding
- Data Hazards Requiring Stalls
- Branch Hazards
- Reducing Pipeline Branch Penalties
- Performance of Branch Schemes
- Example
- Answer
- Reducing the Cost of Branches Through Prediction
- Static Branch Prediction
- Dynamic Branch Prediction and Branch-Prediction Buffers
- A Simple Implementation of RISC V
- A Basic Pipeline for RISC V
- Implementing the Control for the RISC V Pipeline
- Dealing With Branches in the Pipeline
- Dealing With Exceptions
- Types of Exceptions and Requirements
- Stopping and Restarting Execution
- Exceptions in RISC V
- Instruction Set Complications
- Hazards and Forwarding in Longer Latency Pipelines
- Maintaining Precise Exceptions
- Performance of a Simple RISC V FP Pipeline
- The Floating-Point Pipeline
- Performance of the R4000 Pipeline
- RISC Instruction Sets and Efficiency of Pipelining
- Dynamically Scheduled Pipelines
- Dynamic Scheduling With a Scoreboard
- D. Storage Systems
- D.1 Introduction
- D.2 Advanced Topics in Disk Storage
- Disk Power
- Advanced Topics in Disk Arrays
- RAID 10 versus 01 (or 1+0 versus RAID 0+1)
- RAID 6: Beyond a Single Disk Failure
- D.3 Definition and Examples of Real Faults and Failures
- Berkeley’s Tertiary Disk
- Tandem
- Other Studies of the Role of Operators in Dependability
- D.4 I/O Performance, Reliability Measures, and Benchmarks
- Throughput versus Response Time
- Transaction-Processing Benchmarks
- SPEC System-Level File Server, Mail, and Web Benchmarks
- Examples of Benchmarks of Dependability
- D.5 A Little Queuing Theory
- Example
- Answer
- Poisson Distribution of Random Variables
- Point-to-Point Links and Switches Replacing Buses
- Block Servers versus Filers
- Asynchronous I/O and Operating Systems
- The Internet Archive Cluster
- Estimating Performance, Dependability, and Cost of the Internet Archive Cluster
- Example
- Answer
- Calculating MTTF of the TB-80 Cluster
- Example
- Answer
- Case Study 1: Deconstructing a Disk
- Concepts illustrated by this case study
- Case Study 2: Deconstructing a Disk Array
- Concepts illustrated by this case study
- Case Study 3: RAID Reconstruction
- Concepts illustrated by this case study
- Case Study 4: Performance Prediction for RAIDs
- Concepts illustrated by this case study
- Case Study 5: I/O Subsystem Design
- Concepts illustrated by this case study
- Case Study 6: Dirty Rotten Bits
- Concepts illustrated by this case study
- Case Study 7: Sorting Things Out
- Concepts illustrated by this case study
- E. Embedded Systems
- E.1 Introduction
- Real-Time Processing
- E.2 Signal Processing and Embedded Applications: The Digital Signal Processor
- Example
- Answer
- The TI 320C55
- The TI 320C6x
- Media Extensions
- Power Consumption and Efficiency as the Metric
- Background on Wireless Networks
- The Cell Phone
- Cell Phone Standards and Evolution
- F. Interconnection Networks
- F.1 Introduction
- Interconnection Network Domains
- Approach and Organization of This Appendix
- F.2 Interconnecting Two Devices
- Network Interface Functions: Composing and Processing Messages
- Basic Network Structure and Functions: Media and Form Factor, Packet Transport, Flow Control, and Error Handling
- Example
- Answer
- Characterizing Performance: Latency and Effective Bandwidth
- Additional Network Structure and Functions: Topology, Routing, Arbitration, and Switching
- Shared-Media Networks
- Switched-Media Networks
- Comparison of Shared- and Switched-Media Networks
- Characterizing Performance: Latency and Effective Bandwidth
- Example
- Answer
- Centralized Switched Networks
- Example
- Answer
- Distributed Switched Networks
- Example
- Answer
- Example
- Answer
- Effects of Topology on Network Performance
- Example
- Answer
- Routing
- Example
- Answer
- Arbitration
- Switching
- Impact on Network Performance
- Example
- Answer
- Basic Switch Microarchitecture
- Buffer Organizations
- Routing Algorithm Implementation
- Pipelining the Switch Microarchitecture
- Other Switch Microarchitecture Enhancements
- Connectivity
- Standardization: Cross-Company Interoperability
- Congestion Management
- Fault Tolerance
- Example
- Answer
- On-Chip Network: Intel Single-Chip Cloud Computer
- System Area Network: IBM Blue Gene/L 3D Torus Network
- System/Storage Area Network: InfiniBand
- Ethernet: The Local Area Network
- Wide Area Network: ATM
- Density-Optimized Processors versus SPEC-Optimized Processors
- Smart Switches versus Smart Interface Cards
- Protection and User Access to the Network
- Efficient Interface to the Memory Hierarchy versus the Network
- Compute-Optimized Processors versus Receiver Overhead
- Acknowledgments
- Wide Area Networks
- Local Area Networks
- System Area Networks
- Storage Area Networks
- On-Chip Networks
- G. Vector Processors in More Depth
- G.1 Introduction
- G.2 Vector Performance in More Depth
- Example
- Answer
- Example
- Answer
- Pipelined Instruction Start-Up and Multiple Lanes
- Example
- Answer
- Example
- Answer
- Chaining in More Depth
- Sparse Matrices in More Depth
- Measures of Vector Performance
- The Peak Performance of VMIPS on DAXPY
- Sustained Performance of VMIPS on the Linpack Benchmark
- Example
- Answer
- Example
- Answer
- DAXPY Performance on an Enhanced VMIPS
- Example
- Answer
- Example
- Answer
- Multi-Streaming Processors
- Cray X1E
- H. Hardware and Software for VLIW and EPIC
- H.1 Introduction: Exploiting Instruction-Level Parallelism Statically
- H.2 Detecting and Enhancing Loop-Level Parallelism
- Example
- Answer
- Example
- Answer
- Finding Dependences
- Example
- Answer
- Example
- Answer
- Eliminating Dependent Computations
- Software Pipelining: Symbolic Loop Unrolling
- Example
- Answer
- Global Code Scheduling
- Trace Scheduling: Focusing on the Critical Path
- Superblocks
- Example
- Answer
- Example
- Answer
- Hardware Support for Preserving Exception Behavior
- Example
- Answer
- Example
- Answer
- Example
- Answer
- Hardware Support for Memory Reference Speculation
- The Intel IA-64 Instruction Set Architecture
- The IA-64 Register Model
- Instruction Format and Support for Explicit Parallelism
- Example
- Answer
- Instruction Set Basics
- Predication and Speculation Support
- The Itanium 2 Processor
- Functional Units and Instruction Issue
- Itanium 2 Performance
- I. Large-Scale Multiprocessors and Scientific Applications
- I.1 Introduction
- I.2 Interprocessor Communication: The Critical Performance Issue
- I.3 Characteristics of Scientific Applications
- Characteristics of Scientific Applications
- The FFT Kernel
- The LU Kernel
- The Barnes Application
- The Ocean Application
- Computation/Communication for the Parallel Programs
- Example
- Answer
- Synchronization Performance Challenges
- Example
- Answer
- Barrier Synchronization
- Example
- Answer
- Synchronization Mechanisms for Larger-Scale Multiprocessors
- Software Implementations
- Hardware Primitives
- Example
- Answer
- Example
- Answer
- Performance of a Scientific Workload on a Symmetric Shared-Memory Multiprocessor
- Performance of a Scientific Workload on a Distributed-Memory Multiprocessor
- Example
- Answer
- Example
- Answer
- Implementing Cache Coherence in a DSM Multiprocessor
- Avoiding Deadlock from Limited Buffering
- Implementing the Directory Controller
- The Blue Gene/L Computing Node
- J. Computer Arithmetic
- J.1 Introduction
- J.2 Basic Techniques of Integer Arithmetic
- Ripple-Carry Addition
- Radix-2 Multiplication and Division
- Multiply Step
- Divide Step
- Nonrestoring
- Divide Step
- Signed Numbers
- Example
- Answer
- Example
- Answer
- Systems Issues
- Special Values and Denormals
- Representation of Floating-Point Numbers
- Example
- Answer
- Example
- Answer
- Example
- Answer
- Denormals
- Precision of Multiplication
- Example
- Answer
- Example
- Answer
- Speeding Up Addition
- Denormalized Numbers
- Iterative Division
- Floating-Point Remainder
- Fused Multiply-Add
- Precisions
- Exceptions
- Underflow
- Carry-Lookahead
- Example
- Answer
- Carry-Skip Adders
- Carry-Select Adder
- Shifting over Zeros
- SRT Division
- Speeding Up Multiplication with a Single Adder
- Faster Multiplication with Many Adders
- Example
- Answer
- Faster Division with One Adder
- Example
- Answer
- K. Survey of Instruction Set Architectures
- K.1 Introduction
- K.2 A Survey of RISC Architectures for Desktop, Server, and Embedded Computers
- Introduction
- Addressing Modes and Instruction Formats
- Instructions
- RV64G Core Instructions
- Compare and Conditional Branch
- RV64GC Core 16-bit Instructions
- Instructions: Common Extensions beyond RV64G
- Instructions Unique to MIPS64 R6
- Instructions Unique to SPARC v.9
- Register Windows
- Fast Traps
- Support for LISP and Smalltalk
- Instructions Unique to ARM
- Instructions Unique to Power3
- Branch Registers: Link and Counter
- Instructions: Multimedia Extensions of the Desktop/Server RISCs
- Instructions: Digital Signal-Processing Extensions of the Embedded RISCs
- Concluding Remarks
- K.3 The Intel 80x86
- Introduction
- 80x86 Registers and Data Addressing Modes
- 80x86 Integer Operations
- 80x86 Floating-Point Operations
- 80x86 Instruction Encoding
- Putting It All Together: Measurements of Instruction Set Usage
- Measurements of 80x86 Operand Addressing
- Comparative Operation Measurements
- Concluding Remarks
- Beauty is in the eye of the beholder
- K.4 The VAX Architecture
- Introduction
- VAX Operands and Addressing Modes
- Example
- Answer
- Encoding VAX Instructions
- VAX Operations
- Number of Operations
- Branches, Jumps, and Procedure Calls
- An Example to Put It All Together: swap
- Register Allocation for swap
- Code for the Body of the Procedure swap
- Preserving Registers across Procedure Invocation of swap
- The Full Procedure swap
- A Longer Example: sort
- Register Allocation for sort
- Code for the Body of the sort Procedure
- Preserving Registers across Procedure Invocation of sort
- The Full Procedure sort
- Fallacies and Pitfalls
- Concluding Remarks
- Exercises
- Introduction
- System/360 Instruction Set
- Integer/Logical and Floating-Point R-R Instructions
- Branches and Status Setting R-R Instructions
- Branches/Logical and Floating-Point Instructions—RX Format
- Branches and Special Loads and Stores—RX Format
- RS and SI Format Instructions
- SS Format Instructions
- 360 Detailed Measurements
- L. Advanced Concepts on Address Translation
- M. Historical Perspectives and References
- M.1 Introduction
- M.2 The Early Development of Computers (Chapter 1)
- The First General-Purpose Electronic Computers
- Important Special-Purpose Machines
- Commercial Developments
- Development of Quantitative Performance Measures: Successes and Failures
- References
- M.3 The Development of Memory Hierarchy and Protection (Chapter 2 and Appendix B)
- References
- M.4 The Evolution of Instruction Sets (Appendices A, J, and K)
- Stack Architectures
- Computer Architecture Defined
- High-Level Language Computer Architecture
- Reduced Instruction Set Computers
- References
- M.5 The Development of Pipelining and Instruction-Level Parallelism (Chapter 3 and Appendices C and H)
- Early Pipelined CPUs
- The Introduction of Dynamic Scheduling
- The IBM 360 Model 91: A Landmark Computer
- Branch-Prediction Schemes
- The Development of Multiple-Issue Processors
- Compiler Technology and Hardware Support for Scheduling
- EPIC and the IA-64 Development
- Studies of ILP and Ideas to Increase ILP
- Going Beyond the Data Flow Limit
- Recent Advanced Microprocessors
- Multithreading and Simultaneous Multithreading
- References
- M.6 The Development of SIMD Supercomputers, Vector Computers, Multimedia SIMD Instruction Extensions, and Graphical Processor Units (Chapter 4)
- SIMD Supercomputers
- Vector Computers
- Multimedia SIMD Instruction Extensions
- Graphical Processor Units
- Scalable GPUs
- Graphics Pipelines
- GPGPU: An Intermediate Step
- GPU Computing
- References
- SIMD Supercomputers
- Vector Architecture
- Multimedia SIMD
- GPU
- M.7 The History of Multiprocessors and Parallel Processing (Chapter 5 and Appendices F, G, and I)
- SIMD Computers: Attractive Idea, Many Attempts, No Lasting Successes
- Other Early Experiments
- Great Debates in Parallel Processing
- More Recent Advances and Developments
- The Development of Bus-Based Coherent Multiprocessors
- Toward Large-Scale Multiprocessors
- Clusters
- Recent Trends in Large-Scale Multiprocessors
- Developments in Synchronization and Consistency Models
- Other References
- References
- M.8 The Development of Clusters (Chapter 6)
- Clusters, the Forerunner of WSCs
- Utility Computing, the Forerunner of Cloud Computing
- Containers
- References
- M.9 Historical Perspectives and References
- References
- M.10 The History of Magnetic Storage, RAID, and I/O Buses (Appendix D)
- Magnetic Storage
- RAID
- I/O Buses and Controllers
- References
- References
- Index