FIELD NOTE / TIKTOK
Look at this ninety-five-percent GPU chart.
The complete written thought and the evidence behind it. The video edition will follow its public release.
The written argument is here.
This approved TikTok edition is on the journal now. Its video player and original platform link will appear after each public release is verified.
Look at this ninety-five-percent GPU chart.
Video caption
Look at this ninety-five-percent GPU chart. Device busy; accepted output flat. Keep the badge for diagnosis, not procurement. #EricFieldNotes
Full written post / accessibility read
The chart says ninety-five percent. Great. Ninety-five percent of what? An interval counter can show activity while your requests queue, miss latency targets, or fail quality review. The percentage is not a completed-run receipt.
In a simulated inference service, four users send long requests. The GPU stays active, but the second and third users wait through prefill and time out. The fleet metric can rise while successful responses per minute fall. Both readings can be accurate.
Replay the request mix. Align DCGM activity with queue wait, first token, full completion, cancellations and quality checks. If training, align with useful steps and checkpoint completion instead. The correlation shows where the busy interval helps and where it burns time.
Choose capacity and scheduler policy by accepted output, tail time and total cost. Use utilization to explain a miss and retest a fix. Do this because a ninety-five-percent chart cannot tell a buyer what they bought.
#EricFieldNotes
Four-beat scene transcript
1. Look at this ninety-five-percent GPU chart.
The chart says ninety-five percent. Great. Ninety-five percent of what? An interval counter can show activity while your requests queue, miss latency targets, or fail quality review. The percentage is not a completed-run receipt.
Visual: It may be true and still miss the customer's question.
2. Now watch the application.
In a simulated inference service, four users send long requests. The GPU stays active, but the second and third users wait through prefill and time out. The fleet metric can rise while successful responses per minute fall. Both readings can be accurate.
Visual: Device busy; accepted output flat.
3. Put two clocks on the same run.
Replay the request mix. Align DCGM activity with queue wait, first token, full completion, cancellations and quality checks. If training, align with useful steps and checkpoint completion instead. The correlation shows where the busy interval helps and where it burns time.
Visual: Hardware intervals and customer completion.
4. Show completed work per dollar.
Choose capacity and scheduler policy by accepted output, tail time and total cost. Use utilization to explain a miss and retest a fix. Do this because a ninety-five-percent chart cannot tell a buyer what they bought.
Visual: Keep the badge for diagnosis, not procurement.
Research and claim limits
The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.