Measuring Continuously & Fixing Each Bottleneck
◉ Part 2 What These Terms Actually Mean
◉ Part 3 Measuring Continuously & Fixing Each Issue (you are here)
Why This Exists
Parts 1 and 2 told you what to look for and what the terms mean. This post covers the operational side: how to detect bottlenecks before they become incidents, how to establish baselines that make anomalies obvious, and what to actually do once you've confirmed the constraint.
The Measurement Mindset
Three principles before diving into specifics:
- Baseline first, alert second. You can't know something is abnormal unless you know what normal looks like. Collect a week of data before setting thresholds.
- Measure at the resource, not the symptom. "Slow API" is a symptom. "95th percentile storage latency at 12ms" is a measurement. Fix the measurement, the symptom resolves.
- Trend over threshold. A threshold alert fires when you're already in trouble. A trend alert fires when you're heading there. Set both.
Continuous Monitoring Stack
You don't need a complex setup. The core loop is:
Collect metrics → Store time-series → Visualize → Alert → Investigate
| Layer | Purpose | Common Tools |
|---|---|---|
| Collection | Pull metrics from OS, hardware, and applications | node_exporter, cAdvisor, DCGM Exporter, telegraf |
| Storage | Time-series database | Prometheus, VictoriaMetrics, InfluxDB |
| Visualization | Dashboards and exploration | Grafana, Datadog |
| Alerting | Notify on threshold or trend violations | Alertmanager, PagerDuty, OpsGenie |
| Profiling | Deep-dive when alerts fire | perf, bpftrace, py-spy, nvidia-smi |
The rest of this post maps each bottleneck type to: what to collect, what normal looks like, when to alert, and how to fix it.
CPU Bottlenecks
What to Collect
| Metric | Source | What It Tells You |
|---|---|---|
| CPU utilization (per-core) | node_exporter |
How busy each core is |
| Load average (1m, 5m, 15m) | /proc/loadavg |
How many tasks are waiting for CPU time |
| IPC | perf stat |
How efficiently the CPU is executing |
| Context switches/sec | vmstat, node_exporter |
How often the OS swaps between tasks |
| Runqueue length | node_exporter |
How many threads are waiting to run |
What Normal Looks Like
- Utilization: 40–70% sustained under load is healthy. Headroom for spikes.
- Load average: at or below core count.
- IPC: 1.5–3.0 for most workloads. Below 1.0 indicates stalls.
When to Alert
- Utilization > 85% sustained for 5+ minutes.
- Load average > 2× core count.
- IPC drops below 1.0 (indicates memory or branch stalls, not true CPU saturation).
Remediation Playbook
- Confirm it's real CPU saturation check that IPC is healthy (>1.0). Low IPC means the CPU is stalled, not busy. That's a memory or branch problem.
- Identify the hot path
perf toporperf record→perf reportto find which functions consume the most cycles. - Software first:
- Profile for algorithmic inefficiency (O(n²) loops, redundant computation).
- Check for excessive serialization (locks holding cores idle).
- Enable compiler optimizations if not already (
-O2,-O3, PGO).
- Scale horizontally if the workload parallelizes, distribute across more cores or machines.
- Hardware upgrade more cores for parallel workloads; higher clock / better IPC (newer generation) for single-threaded workloads.
Memory Bottlenecks
Capacity
What to Collect
| Metric | Source | What It Tells You |
|---|---|---|
| Used / available / free memory | node_exporter, free -h |
How much headroom you have |
| Swap usage | node_exporter |
Whether the OS is spilling to disk |
| Page faults (major) | vmstat |
How often the OS fetches from swap |
| OOM kills | dmesg, node_exporter |
When processes are being terminated for memory |
What Normal Looks Like
- Swap usage: 0 for latency-sensitive workloads. Any swap under load means you're already degraded.
- Available memory: >20% of total as buffer.
- Major page faults: near zero.
When to Alert
- Available memory < 15% of total.
- Swap usage > 0 and increasing.
- Any OOM kill.
Remediation
- Identify the consumer
smem,ps aux --sort=-%mem, or container memory metrics. - Right-size the workload is the application leaking? Holding data unnecessarily? Caching too aggressively?
- Add capacity more DIMMs or larger modules. Cheapest fix when the workload legitimately needs the memory.
- Shard the data split across machines if a single node can't hold the working set.
Bandwidth
What to Collect
| Metric | Source | What It Tells You |
|---|---|---|
| LLC (Last Level Cache) misses/sec | perf stat |
How often data isn't found in cache |
| Memory bandwidth utilization | pcm-memory (Intel), perf mem, likwid |
How close to the theoretical max |
| Instructions per cycle (IPC) | perf stat |
Low IPC with low CPU% = memory stalls |
What Normal Looks Like
- Memory bandwidth: <60% of theoretical max under sustained load.
- LLC miss rate: workload-dependent, but sudden increases indicate regression.
- IPC: >1.5 for compute workloads.
When to Alert
- Memory bandwidth > 75% of theoretical maximum sustained.
- IPC drops below baseline by >30% without a corresponding code change.
Remediation
- Confirm bandwidth saturation
perf stat -e LLC-load-misses,LLC-store-missesduring load. - Improve data locality restructure data layouts (struct-of-arrays vs. array-of-structs), reduce pointer chasing.
- Reduce working set compression, smaller data types, more efficient representations.
- Populate all memory channels a single DIMM per channel cuts bandwidth in half on most platforms.
- Upgrade DDR generation DDR5 offers ~50% more bandwidth per channel vs DDR4.
Storage Bottlenecks
What to Collect
| Metric | Source | What It Tells You |
|---|---|---|
| IOPS (read/write) | iostat -x, node_exporter |
Operations per second |
| Throughput (MB/s) | iostat -x |
Sequential bandwidth |
| Average latency (await) | iostat -x |
How long each operation waits |
| Queue depth (avgqu-sz / aqu-sz) | iostat -x |
How many requests are in-flight |
| %util | iostat -x |
Device saturation (less meaningful for NVMe) |
What Normal Looks Like
- NVMe latency: <200μs average for typical workloads.
- Queue depth: 1–32 under normal load. >64 means the device is saturated.
- %util: >90% for spinning disks means saturated. For NVMe this metric is less useful focus on latency and queue depth instead.
When to Alert
- Average latency (await) > 2× baseline.
- Queue depth sustained > 64.
- IOPS drops significantly below baseline (indicates throttling or hardware degradation).
Remediation Playbook
- Identify the I/O pattern
iotoporbiosnoop(from the BCC toolkit / bpftrace) to find which process is responsible and whether the pattern is random or sequential. - Reduce unnecessary I/O:
- Add application-level caching (Redis, local page cache).
- Batch small writes into larger ones.
- Move temp files to tmpfs (RAM-backed filesystem).
- Separate workloads put logs on a different device than your database.
- Upgrade the device:
- HDD → SATA SSD: ~100× IOPS improvement.
- SATA SSD → NVMe: ~5–10× latency improvement + deeper queue support.
- Single NVMe → RAID/multiple NVMe: linear IOPS scaling.
- Use a write-ahead-log (WAL) on fast storage keeps critical-path writes on the fastest device.
Network Bottlenecks
What to Collect
| Metric | Source | What It Tells You |
|---|---|---|
| Bytes in/out per interface | node_exporter |
How close to link capacity |
| Packets dropped | node_exporter, ethtool -S |
Whether the NIC or kernel is overwhelmed |
| TCP retransmits | ss -s, node_exporter |
Packet loss or congestion |
| Connection states | ss -s |
Backlog, TIME_WAIT buildup |
| RTT (round-trip time) | ping, application metrics |
Latency between services |
What Normal Looks Like
- NIC utilization: <70% of link rate sustained.
- Retransmits: <0.1% of total packets.
- Drops: 0 under normal operation.
- RTT: stable, matching expected distance (same-rack <100μs, same-DC <1ms).
When to Alert
- NIC utilization > 80% for 5+ minutes.
- Retransmit rate > 1%.
- Any sustained packet drops.
- RTT increases > 2× baseline.
Remediation Playbook
- Identify the talker
iftop,nethogs, or flow-level metrics to find which process or connection is dominating. - Reduce chattiness:
- Batch small RPCs into fewer, larger requests.
- Enable connection pooling and multiplexing (HTTP/2, gRPC).
- Compress payloads if CPU has headroom.
- Reduce latency impact:
- Co-locate high-communication services in the same rack or AZ.
- Use connection-aware load balancing to minimize cross-zone hops.
- Prefetch data to avoid serial round-trip chains.
- Upgrade hardware:
- 10G → 25G → 100G NIC upgrade.
- Enable RSS (Receive Side Scaling) to distribute packets across multiple CPU cores.
- RDMA/kernel bypass for ultra-low-latency paths.
- Architecture change if you're saturating a single link, shard traffic across multiple NICs or multiple machines.
GPU Bottlenecks
What to Collect
| Metric | Source | What It Tells You |
|---|---|---|
| GPU utilization (SM %) | nvidia-smi, DCGM |
How busy the compute units are |
| Memory utilization (%) | nvidia-smi, DCGM |
How full VRAM is |
| Memory bandwidth utilization | DCGM, nvbandwidth |
How saturated the VRAM bus is |
| Power draw (watts) | nvidia-smi |
Whether the GPU is thermally throttling |
| NVLink throughput | DCGM | Multi-GPU communication saturation |
| PCIe throughput | DCGM | Host-to-GPU data transfer rate |
What Normal Looks Like
- SM utilization: >80% during training/inference means the GPU is well-utilized.
- VRAM usage: 80–90% is efficient. >95% means you're one batch size increase from OOM.
- Power: at or below TDP. Sustained power limit = thermal throttling likely.
When to Alert
- VRAM usage > 95%.
- SM utilization drops below 50% during expected workload (indicates the GPU is starved).
- PCIe throughput near max sustained (GPU waiting for data from host).
- Temperature > 85°C sustained (throttling territory).
Remediation Playbook
GPU starved (low SM utilization):
- Increase batch size to give the GPU more work per kernel launch.
- Optimize data pipeline CPU preprocessing can't keep up. Use multi-worker data loaders, prefetching.
- Overlap compute and data transfer with CUDA streams.
VRAM exhaustion:
- Reduce batch size (simple but reduces throughput).
- Enable gradient checkpointing (trades compute for memory).
- Use mixed precision (FP16/BF16) halves memory for activations and gradients.
- Shard the model across GPUs (tensor parallelism, pipeline parallelism).
- Upgrade to GPUs with more VRAM.
Memory bandwidth bound (inference):
- Quantize the model (INT8, INT4) fewer bytes to move per inference.
- Use Flash Attention or fused kernels to reduce memory round trips.
- Batch requests to amortize memory reads across multiple inputs.
Multi-GPU communication bound:
- Ensure NVLink is active (not falling back to PCIe).
- Overlap communication with computation (gradient bucketing in DDP).
- Reduce synchronization frequency where possible.
- For multi-node: upgrade to InfiniBand or high-bandwidth RoCE.
Establishing Baselines
A baseline is what "normal" looks like for your specific system. Without one, every metric is meaningless noise.
How to Build a Baseline
- Collect 7 days minimum covers weekday/weekend patterns and batch jobs.
- Separate by workload phase "normal traffic" vs. "batch processing window" vs. "deployment."
- Record percentiles, not averages p50, p95, p99. Averages hide the tail.
- Version your baselines after a significant change (new hardware, major deploy), let a new baseline stabilize.
Baseline Metrics Worth Tracking
| Category | Key Baseline Metrics |
|---|---|
| CPU | p95 utilization, average IPC, context switches/sec |
| Memory | peak usage, swap events/day, bandwidth utilization |
| Storage | p95 latency, average queue depth, IOPS under load |
| Network | peak bandwidth %, retransmit rate, p99 RTT |
| GPU | average SM%, peak VRAM %, average memory bandwidth % |
Continuous Monitoring Checklist
A minimal setup that covers the bottleneck types from this series:
- ☐
node_exporterrunning on every host covers CPU, memory, disk, and network basics - ☐ Storage:
iostatmetrics scraped latency, queue depth, IOPS - ☐ GPU: DCGM exporter on GPU nodes SM%, VRAM, power, NVLink throughput
- ☐ Application-level: request latency histograms at p50 / p95 / p99
- ☐ Dashboards: one per resource type (CPU, Memory, Storage, Network, GPU)
- ☐ Alerts: threshold + trend for each critical metric
- ☐ Baselines: reviewed and updated quarterly or after major changes
- ☐ Profiling runbook: documented steps for when alerts fire which tool to run, what to look for
The Diagnosis Workflow
When an alert fires or someone reports "it's slow," follow this sequence:
1. Which resource is saturated?
→ Check dashboard for CPU, memory, storage, network, GPU
2. Is it the hardware or the software?
→ Hardware: utilization near 100%, latency at device limits
→ Software: low utilization but high wait times (locks, bad algorithms, misconfiguration)
3. What changed?
→ Recent deploy? Traffic spike? New data pattern? Batch job overlap?
4. Fix the constraint:
→ Software fix if possible (cheaper, faster to deploy)
→ Hardware upgrade if the workload legitimately needs more capacity
5. Update the baseline and add a regression alert.
In Short
Measurement without action is monitoring theater. The loop is: baseline → detect deviation → diagnose the limiting resource → remediate → update the baseline. Do this continuously and hardware bottlenecks become routine maintenance rather than emergencies.
Part 1 mapped bottlenecks to hardware. Part 2 explained the terms. Part 3 (this post) gave you the measurement and remediation playbooks. The full loop: identify the constraint, understand the underlying hardware, measure it continuously, and fix it systematically.