JournalDAY 28 / TIKTOK

FIELD NOTE / TIKTOK

A green health watch is not a burn-in test.

The complete written thought and the evidence behind it. The video edition will follow its public release.

Journal September 25, 2026 · TikTok target October 25, 2026

A green health watch is not a burn-in test.

Day 28 · Week 4 editorial group · TikTok · no publication date or time assigned

Video caption

A green health watch is not a burn-in test. One idle period says little about sustained correctness. Green, unknown and failed must not collapse together. #EricFieldNotes

Full written post / accessibility read

DCGM health watches report sampled errors and subsystem state without stressing the GPU. Its active diagnostic does load compute and memory. Those are different tests, and running the second on a customer's live job would itself be an incident.

Imagine a node leaves service after an error, then passive watches look quiet. If the operator marks it sellable immediately, the next long training run becomes the stress test. The customer's checkpoint is now the diagnostic bill.

Block new admission to the suspect device. On operator-owned idle capacity, run the relevant DCGM diagnostic and inspect error codes. After repair, run a small representative workload and only then return the node to a saleable class.

Record passive-watch window, active-test result, repair action and release timestamp. Do this because an operator needs to know which kind of evidence supports the next allocation, not merely whether a dashboard tile is green.

#EricFieldNotes

Four-beat scene transcript

1. A green health watch is not a burn-in test.

DCGM health watches report sampled errors and subsystem state without stressing the GPU. Its active diagnostic does load compute and memory. Those are different tests, and running the second on a customer's live job would itself be an incident.

Visual: Passive telemetry sees only what happened during its window.

2. A returning node can poison the next allocation.

Imagine a node leaves service after an error, then passive watches look quiet. If the operator marks it sellable immediately, the next long training run becomes the stress test. The customer's checkpoint is now the diagnostic bill.

Visual: One idle period says little about sustained correctness.

3. Build quarantine and release criteria.

Block new admission to the suspect device. On operator-owned idle capacity, run the relevant DCGM diagnostic and inspect error codes. After repair, run a small representative workload and only then return the node to a saleable class.

Visual: Drain, diagnose, repair, run a representative canary.

4. Make health an evidence state.

Record passive-watch window, active-test result, repair action and release timestamp. Do this because an operator needs to know which kind of evidence supports the next allocation, not merely whether a dashboard tile is green.

Visual: Green, unknown and failed must not collapse together.

Research and claim limits

The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.

More notes from the work ↗