Hardware Performance Monitors (PMUs) in IBM POWER10 are designed to precisely attribute where cycles are being lost, including different types of pipeline stalls. They don’t just say “the CPU stalled”—they break stalls down by stage and cause.
Here’s how that works.
🔷 1. What “Stall Cycles” Mean in Hardware
A stall cycle is any cycle where:
-
The pipeline cannot make forward progress (fully or partially)
-
Or a stage (fetch, decode, dispatch, issue, execute) is blocked
👉 POWER tracks these at fine granularity.
🔷 2. Performance Monitoring Unit (PMU) Basics
POWER CPUs include:
-
Performance Monitor Counters (PMCs)
-
Performance Monitor Registers (PMRs)
-
Event-based counting logic
These track:
-
Cycles
-
Instructions
-
Specific stall conditions
🔷 3. Event-Based Stall Measurement
Instead of a single “stall counter,” POWER uses event signals from pipeline stages.
Examples of stall-related events:
-
Dispatch stall cycles
-
Issue queue full cycles
-
Load miss stall cycles
-
Branch misprediction recovery cycles
👉 Each event increments a counter when its condition is true.
🔷 4. Pipeline Stage Instrumentation
Each stage generates its own stall signals:
🔹 1. Fetch Stalls
Captured when:
-
I-cache miss
-
Branch redirect delay
-
Instruction buffer empty
👉 PMU records:
-
Cycles where fetch cannot supply instructions
🔹 2. Decode / Dispatch Stalls
Captured when:
-
No free rename registers
-
ROB (reorder buffer) full
-
Dispatch group cannot proceed
👉 Indicates front-end backpressure
🔹 3. Issue Stalls
Captured when:
-
Instructions ready but:
-
Execution units busy
-
Issue queues full
👉 Shows execution resource contention
🔹 4. Execution Stalls (Backend)
Captured when:
-
Load misses (waiting for memory)
-
Store queue full
-
Dependency chains
👉 Typically the largest contributor in data-heavy workloads
🔹 5. Completion / Commit Stalls
Captured when:
-
Instructions cannot retire
-
Exceptions or ordering constraints
🔷 5. Cycle Accounting Model
POWER uses a cycle accounting approach:
-
Each cycle is classified into categories like:
-
Productive (useful work)
-
Front-end bound
-
Back-end bound
-
Memory bound
👉 This allows:
-
High-level performance diagnosis
🔷 6. Sampling vs Counting
✅ Counting mode
-
Counts total stall cycles for each event
✅ Sampling mode (profiling)
-
Periodically samples execution state
-
Identifies:
-
Which instruction caused stall
-
Which code path is affected
👉 Used by tools like:
-
perf (Linux)
-
IBM performance tools
🔷 7. SMT-Aware Stall Tracking
With SMT (up to 8 threads):
Key detail:
-
A cycle may be:
-
Stall for one thread
-
Productive for another
👉 POWER distinguishes:
-
Thread-level vs core-level utilization
🔷 8. Derived Stall Metrics
Raw counters are combined to compute:
🔸 CPI (Cycles Per Instruction)
🔸 Stall breakdown
-
% cycles stalled on memory
-
% cycles stalled on dispatch
-
% cycles stalled on execution
👉 Helps identify bottlenecks precisely
🔷 9. Example: Memory Stall Detection
If workload shows:
-
High “load miss cycles”
-
High “L3 miss events”
👉 Conclusion:
-
Memory subsystem is bottleneck
🔷 10. Example: Front-End Stall Detection
If:
-
High “dispatch stall cycles”
-
Low execution utilization
👉 Conclusion:
-
Rename/ROB pressure or decode bottleneck
🔷 11. Key Strength of POWER PMU
POWER PMUs are known for:
✔ Fine-grained event coverage
-
Very detailed stall categories
✔ Accurate attribution
-
Pinpoints exact pipeline stage
✔ SMT visibility
-
Distinguishes thread vs core behavior
✅ Bottom Line
Hardware performance monitors in POWER CPUs capture stall cycles by:
-
Instrumenting every pipeline stage
-
Counting specific stall conditions as events
-
Aggregating them into cycle-level metrics
➡️ This allows you to:
-
Identify whether stalls come from:
-
Front-end (fetch/decode)
-
Backend (execution/memory)
-
Resource contention
👉 Result: Extremely precise performance tuning for complex workloads like databases and enterprise systems.