Measuring Continuously & Fixing Each Bottleneck

Raza Balbale Raza Balbale
·
Part 3 of 3: Performance Hardware Series
◉ Part 1 The Common Problems
◉ Part 2 What These Terms Actually Mean
◉ Part 3 Measuring Continuously & Fixing Each Issue (you are here)

Why This Exists

Parts 1 and 2 told you what to look for and what the terms mean. This post covers the operational side: how to detect bottlenecks before they become incidents, how to establish baselines that make anomalies obvious, and what to actually do once you've confirmed the constraint.


The Measurement Mindset

Three principles before diving into specifics:

  1. Baseline first, alert second. You can't know something is abnormal unless you know what normal looks like. Collect a week of data before setting thresholds.
  2. Measure at the resource, not the symptom. "Slow API" is a symptom. "95th percentile storage latency at 12ms" is a measurement. Fix the measurement, the symptom resolves.
  3. Trend over threshold. A threshold alert fires when you're already in trouble. A trend alert fires when you're heading there. Set both.

Continuous Monitoring Stack

You don't need a complex setup. The core loop is:

Collect metrics → Store time-series → Visualize → Alert → Investigate
Layer Purpose Common Tools
Collection Pull metrics from OS, hardware, and applications node_exporter, cAdvisor, DCGM Exporter, telegraf
Storage Time-series database Prometheus, VictoriaMetrics, InfluxDB
Visualization Dashboards and exploration Grafana, Datadog
Alerting Notify on threshold or trend violations Alertmanager, PagerDuty, OpsGenie
Profiling Deep-dive when alerts fire perf, bpftrace, py-spy, nvidia-smi

The rest of this post maps each bottleneck type to: what to collect, what normal looks like, when to alert, and how to fix it.


CPU Bottlenecks

What to Collect

Metric Source What It Tells You
CPU utilization (per-core) node_exporter How busy each core is
Load average (1m, 5m, 15m) /proc/loadavg How many tasks are waiting for CPU time
IPC perf stat How efficiently the CPU is executing
Context switches/sec vmstat, node_exporter How often the OS swaps between tasks
Runqueue length node_exporter How many threads are waiting to run

What Normal Looks Like

  • Utilization: 40–70% sustained under load is healthy. Headroom for spikes.
  • Load average: at or below core count.
  • IPC: 1.5–3.0 for most workloads. Below 1.0 indicates stalls.

When to Alert

  • Utilization > 85% sustained for 5+ minutes.
  • Load average > 2× core count.
  • IPC drops below 1.0 (indicates memory or branch stalls, not true CPU saturation).

Remediation Playbook

  1. Confirm it's real CPU saturation check that IPC is healthy (>1.0). Low IPC means the CPU is stalled, not busy. That's a memory or branch problem.
  2. Identify the hot path perf top or perf record → perf report to find which functions consume the most cycles.
  3. Software first:
    • Profile for algorithmic inefficiency (O(n²) loops, redundant computation).
    • Check for excessive serialization (locks holding cores idle).
    • Enable compiler optimizations if not already (-O2, -O3, PGO).
  4. Scale horizontally if the workload parallelizes, distribute across more cores or machines.
  5. Hardware upgrade more cores for parallel workloads; higher clock / better IPC (newer generation) for single-threaded workloads.

Memory Bottlenecks

Capacity

What to Collect

Metric Source What It Tells You
Used / available / free memory node_exporter, free -h How much headroom you have
Swap usage node_exporter Whether the OS is spilling to disk
Page faults (major) vmstat How often the OS fetches from swap
OOM kills dmesg, node_exporter When processes are being terminated for memory

What Normal Looks Like

  • Swap usage: 0 for latency-sensitive workloads. Any swap under load means you're already degraded.
  • Available memory: >20% of total as buffer.
  • Major page faults: near zero.

When to Alert

  • Available memory < 15% of total.
  • Swap usage > 0 and increasing.
  • Any OOM kill.

Remediation

  1. Identify the consumer smem, ps aux --sort=-%mem, or container memory metrics.
  2. Right-size the workload is the application leaking? Holding data unnecessarily? Caching too aggressively?
  3. Add capacity more DIMMs or larger modules. Cheapest fix when the workload legitimately needs the memory.
  4. Shard the data split across machines if a single node can't hold the working set.

Bandwidth

What to Collect

Metric Source What It Tells You
LLC (Last Level Cache) misses/sec perf stat How often data isn't found in cache
Memory bandwidth utilization pcm-memory (Intel), perf mem, likwid How close to the theoretical max
Instructions per cycle (IPC) perf stat Low IPC with low CPU% = memory stalls

What Normal Looks Like

  • Memory bandwidth: <60% of theoretical max under sustained load.
  • LLC miss rate: workload-dependent, but sudden increases indicate regression.
  • IPC: >1.5 for compute workloads.

When to Alert

  • Memory bandwidth > 75% of theoretical maximum sustained.
  • IPC drops below baseline by >30% without a corresponding code change.

Remediation

  1. Confirm bandwidth saturation perf stat -e LLC-load-misses,LLC-store-misses during load.
  2. Improve data locality restructure data layouts (struct-of-arrays vs. array-of-structs), reduce pointer chasing.
  3. Reduce working set compression, smaller data types, more efficient representations.
  4. Populate all memory channels a single DIMM per channel cuts bandwidth in half on most platforms.
  5. Upgrade DDR generation DDR5 offers ~50% more bandwidth per channel vs DDR4.

Storage Bottlenecks

What to Collect

Metric Source What It Tells You
IOPS (read/write) iostat -x, node_exporter Operations per second
Throughput (MB/s) iostat -x Sequential bandwidth
Average latency (await) iostat -x How long each operation waits
Queue depth (avgqu-sz / aqu-sz) iostat -x How many requests are in-flight
%util iostat -x Device saturation (less meaningful for NVMe)

What Normal Looks Like

  • NVMe latency: <200μs average for typical workloads.
  • Queue depth: 1–32 under normal load. >64 means the device is saturated.
  • %util: >90% for spinning disks means saturated. For NVMe this metric is less useful focus on latency and queue depth instead.

When to Alert

  • Average latency (await) > 2× baseline.
  • Queue depth sustained > 64.
  • IOPS drops significantly below baseline (indicates throttling or hardware degradation).

Remediation Playbook

  1. Identify the I/O pattern iotop or biosnoop (from the BCC toolkit / bpftrace) to find which process is responsible and whether the pattern is random or sequential.
  2. Reduce unnecessary I/O:
    • Add application-level caching (Redis, local page cache).
    • Batch small writes into larger ones.
    • Move temp files to tmpfs (RAM-backed filesystem).
  3. Separate workloads put logs on a different device than your database.
  4. Upgrade the device:
    • HDD → SATA SSD: ~100× IOPS improvement.
    • SATA SSD → NVMe: ~5–10× latency improvement + deeper queue support.
    • Single NVMe → RAID/multiple NVMe: linear IOPS scaling.
  5. Use a write-ahead-log (WAL) on fast storage keeps critical-path writes on the fastest device.

Network Bottlenecks

What to Collect

Metric Source What It Tells You
Bytes in/out per interface node_exporter How close to link capacity
Packets dropped node_exporter, ethtool -S Whether the NIC or kernel is overwhelmed
TCP retransmits ss -s, node_exporter Packet loss or congestion
Connection states ss -s Backlog, TIME_WAIT buildup
RTT (round-trip time) ping, application metrics Latency between services

What Normal Looks Like

  • NIC utilization: <70% of link rate sustained.
  • Retransmits: <0.1% of total packets.
  • Drops: 0 under normal operation.
  • RTT: stable, matching expected distance (same-rack <100μs, same-DC <1ms).

When to Alert

  • NIC utilization > 80% for 5+ minutes.
  • Retransmit rate > 1%.
  • Any sustained packet drops.
  • RTT increases > 2× baseline.

Remediation Playbook

  1. Identify the talker iftop, nethogs, or flow-level metrics to find which process or connection is dominating.
  2. Reduce chattiness:
    • Batch small RPCs into fewer, larger requests.
    • Enable connection pooling and multiplexing (HTTP/2, gRPC).
    • Compress payloads if CPU has headroom.
  3. Reduce latency impact:
    • Co-locate high-communication services in the same rack or AZ.
    • Use connection-aware load balancing to minimize cross-zone hops.
    • Prefetch data to avoid serial round-trip chains.
  4. Upgrade hardware:
    • 10G → 25G → 100G NIC upgrade.
    • Enable RSS (Receive Side Scaling) to distribute packets across multiple CPU cores.
    • RDMA/kernel bypass for ultra-low-latency paths.
  5. Architecture change if you're saturating a single link, shard traffic across multiple NICs or multiple machines.

GPU Bottlenecks

What to Collect

Metric Source What It Tells You
GPU utilization (SM %) nvidia-smi, DCGM How busy the compute units are
Memory utilization (%) nvidia-smi, DCGM How full VRAM is
Memory bandwidth utilization DCGM, nvbandwidth How saturated the VRAM bus is
Power draw (watts) nvidia-smi Whether the GPU is thermally throttling
NVLink throughput DCGM Multi-GPU communication saturation
PCIe throughput DCGM Host-to-GPU data transfer rate

What Normal Looks Like

  • SM utilization: >80% during training/inference means the GPU is well-utilized.
  • VRAM usage: 80–90% is efficient. >95% means you're one batch size increase from OOM.
  • Power: at or below TDP. Sustained power limit = thermal throttling likely.

When to Alert

  • VRAM usage > 95%.
  • SM utilization drops below 50% during expected workload (indicates the GPU is starved).
  • PCIe throughput near max sustained (GPU waiting for data from host).
  • Temperature > 85°C sustained (throttling territory).

Remediation Playbook

  1. GPU starved (low SM utilization):

    • Increase batch size to give the GPU more work per kernel launch.
    • Optimize data pipeline CPU preprocessing can't keep up. Use multi-worker data loaders, prefetching.
    • Overlap compute and data transfer with CUDA streams.
  2. VRAM exhaustion:

    • Reduce batch size (simple but reduces throughput).
    • Enable gradient checkpointing (trades compute for memory).
    • Use mixed precision (FP16/BF16) halves memory for activations and gradients.
    • Shard the model across GPUs (tensor parallelism, pipeline parallelism).
    • Upgrade to GPUs with more VRAM.
  3. Memory bandwidth bound (inference):

    • Quantize the model (INT8, INT4) fewer bytes to move per inference.
    • Use Flash Attention or fused kernels to reduce memory round trips.
    • Batch requests to amortize memory reads across multiple inputs.
  4. Multi-GPU communication bound:

    • Ensure NVLink is active (not falling back to PCIe).
    • Overlap communication with computation (gradient bucketing in DDP).
    • Reduce synchronization frequency where possible.
    • For multi-node: upgrade to InfiniBand or high-bandwidth RoCE.

Establishing Baselines

A baseline is what "normal" looks like for your specific system. Without one, every metric is meaningless noise.

How to Build a Baseline

  1. Collect 7 days minimum covers weekday/weekend patterns and batch jobs.
  2. Separate by workload phase "normal traffic" vs. "batch processing window" vs. "deployment."
  3. Record percentiles, not averages p50, p95, p99. Averages hide the tail.
  4. Version your baselines after a significant change (new hardware, major deploy), let a new baseline stabilize.

Baseline Metrics Worth Tracking

Category Key Baseline Metrics
CPU p95 utilization, average IPC, context switches/sec
Memory peak usage, swap events/day, bandwidth utilization
Storage p95 latency, average queue depth, IOPS under load
Network peak bandwidth %, retransmit rate, p99 RTT
GPU average SM%, peak VRAM %, average memory bandwidth %

Continuous Monitoring Checklist

A minimal setup that covers the bottleneck types from this series:

  • ☐  node_exporter running on every host covers CPU, memory, disk, and network basics
  • ☐  Storage: iostat metrics scraped latency, queue depth, IOPS
  • ☐  GPU: DCGM exporter on GPU nodes SM%, VRAM, power, NVLink throughput
  • ☐  Application-level: request latency histograms at p50 / p95 / p99
  • ☐  Dashboards: one per resource type (CPU, Memory, Storage, Network, GPU)
  • ☐  Alerts: threshold + trend for each critical metric
  • ☐  Baselines: reviewed and updated quarterly or after major changes
  • ☐  Profiling runbook: documented steps for when alerts fire which tool to run, what to look for

The Diagnosis Workflow

When an alert fires or someone reports "it's slow," follow this sequence:

1. Which resource is saturated?
   → Check dashboard for CPU, memory, storage, network, GPU
   
2. Is it the hardware or the software?
   → Hardware: utilization near 100%, latency at device limits
   → Software: low utilization but high wait times (locks, bad algorithms, misconfiguration)
   
3. What changed?
   → Recent deploy? Traffic spike? New data pattern? Batch job overlap?
   
4. Fix the constraint:
   → Software fix if possible (cheaper, faster to deploy)
   → Hardware upgrade if the workload legitimately needs more capacity
   
5. Update the baseline and add a regression alert.

In Short

Measurement without action is monitoring theater. The loop is: baseline → detect deviation → diagnose the limiting resource → remediate → update the baseline. Do this continuously and hardware bottlenecks become routine maintenance rather than emergencies.


Series Complete

Part 1 mapped bottlenecks to hardware. Part 2 explained the terms. Part 3 (this post) gave you the measurement and remediation playbooks. The full loop: identify the constraint, understand the underlying hardware, measure it continuously, and fix it systematically.

#Hardware #Performance #Systems #Bottlenecks #Infrastructure #Monitoring