FIELD NOTE / TIKTOK
A green health watch is not a burn-in test.
The complete written thought and the evidence behind it. The video edition will follow its public release.
The written argument is here.
This approved TikTok edition is on the journal now. Its video player and original platform link will appear after each public release is verified.
A green health watch is not a burn-in test.
Day 28 · Week 4 editorial group · TikTok · no publication date or time assigned
Video caption
A green health watch is not a burn-in test. One idle period says little about sustained correctness. Green, unknown and failed must not collapse together. #EricFieldNotes
Full written post / accessibility read
DCGM health watches report sampled errors and subsystem state without stressing the GPU. Its active diagnostic does load compute and memory. Those are different tests, and running the second on a customer's live job would itself be an incident.
Imagine a node leaves service after an error, then passive watches look quiet. If the operator marks it sellable immediately, the next long training run becomes the stress test. The customer's checkpoint is now the diagnostic bill.
Block new admission to the suspect device. On operator-owned idle capacity, run the relevant DCGM diagnostic and inspect error codes. After repair, run a small representative workload and only then return the node to a saleable class.
Record passive-watch window, active-test result, repair action and release timestamp. Do this because an operator needs to know which kind of evidence supports the next allocation, not merely whether a dashboard tile is green.
#EricFieldNotes
Four-beat scene transcript
1. A green health watch is not a burn-in test.
DCGM health watches report sampled errors and subsystem state without stressing the GPU. Its active diagnostic does load compute and memory. Those are different tests, and running the second on a customer's live job would itself be an incident.
Visual: Passive telemetry sees only what happened during its window.
2. A returning node can poison the next allocation.
Imagine a node leaves service after an error, then passive watches look quiet. If the operator marks it sellable immediately, the next long training run becomes the stress test. The customer's checkpoint is now the diagnostic bill.
Visual: One idle period says little about sustained correctness.
3. Build quarantine and release criteria.
Block new admission to the suspect device. On operator-owned idle capacity, run the relevant DCGM diagnostic and inspect error codes. After repair, run a small representative workload and only then return the node to a saleable class.
Visual: Drain, diagnose, repair, run a representative canary.
4. Make health an evidence state.
Record passive-watch window, active-test result, repair action and release timestamp. Do this because an operator needs to know which kind of evidence supports the next allocation, not merely whether a dashboard tile is green.
Visual: Green, unknown and failed must not collapse together.
Research and claim limits
- NVIDIA DCGM health monitoring (S187)
- NVIDIA DCGM diagnostic plugin (S188)
The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.