FIELD NOTE / X
GPU utilization is easy to flatter.
The complete written thought and the evidence behind it. The video edition will follow its public release.
The written argument is here.
This approved X edition is on the journal now. Its video player and original platform link will appear after each public release is verified.
GPU utilization is easy to flatter.
Video caption
GPU utilization is easy to flatter. Memory and communication wait consume the interval. Optimize accepted work per complete cost. #EricFieldNotes
Full written post / accessibility read
A dashboard says ninety-five percent utilized. Which metric? Reserved time, active kernels, streaming multiprocessor activity, tensor activity and memory traffic answer different questions. None alone tells you how many accepted results the customer received.
A GPU can show substantial activity while a workload spends time on data movement, inefficient batch shape or interconnect stalls. An aggregate fleet number can also hide one customer job starved in the queue. Device counters are diagnostic, not a service-level score.
For training, report accepted steps and completed runs with step-time distribution. For inference, report latency-qualified responses and answer quality at stated concurrency. Pair those with SM, tensor, memory and link activity to explain the result.
Keep utilization as a diagnostic. Make the purchase and scheduling verdict from completed useful work, tail latency and full cost. Do this because an active GPU is not automatically a productive customer service.
#EricFieldNotes
Four-beat scene transcript
1. GPU utilization is easy to flatter.
A dashboard says ninety-five percent utilized. Which metric? Reserved time, active kernels, streaming multiprocessor activity, tensor activity and memory traffic answer different questions. None alone tells you how many accepted results the customer received.
Visual: One busy percentage can hide poor useful throughput.
2. Active can still be stalled.
A GPU can show substantial activity while a workload spends time on data movement, inefficient batch shape or interconnect stalls. An aggregate fleet number can also hide one customer job starved in the queue. Device counters are diagnostic, not a service-level score.
Visual: Memory and communication wait consume the interval.
3. Pair counters with completed work.
For training, report accepted steps and completed runs with step-time distribution. For inference, report latency-qualified responses and answer quality at stated concurrency. Pair those with SM, tensor, memory and link activity to explain the result.
Visual: Same interval, same workload, full cost.
4. Put the customer in the numerator.
Keep utilization as a diagnostic. Make the purchase and scheduling verdict from completed useful work, tail latency and full cost. Do this because an active GPU is not automatically a productive customer service.
Visual: Optimize accepted work per complete cost.
Research and claim limits
The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.