FIELD NOTE / TIKTOK
The spot run stopped. The badge says retry.
The complete written thought and the evidence behind it. The video edition will follow its public release.
The written argument is here.
This approved TikTok edition is on the journal now. Its video player and original platform link will appear after each public release is verified.
The spot run stopped. The badge says retry.
Video caption
The spot run stopped. The badge says retry. Progress, waiting time, and net charge diverge. The service must recover useful work after interruption. #EricFieldNotes
Full written post / accessibility read
In this fictional preemptible GPU pool, capacity is reclaimed while a batch job runs. The scheduler requests a new slot, but the last usable checkpoint is old. A green retry badge can coexist with a late or incomplete customer result.
The job may repeat completed steps, wait for a compatible replacement, and incur a second allocation. Whether either attempt is credited is contract-specific. The retry count cannot tell you how much useful work survived or what the customer owes.
Interrupt a worker in an approved test or local mock. Record checkpoint age, replacement time, repeated steps, final result hash, and net usage after exclusions or credits. A fault drill tests your harness; it does not recreate every provider reclaim policy.
Put checkpoint responsibility, restart priority, lost-progress budget, and billing treatment in the capacity agreement. Do this because a cheap preemptible hour can become an expensive result when recovery is invisible.
#EricFieldNotes
Four-beat scene transcript
1. The spot run stopped. The badge says retry.
In this fictional preemptible GPU pool, capacity is reclaimed while a batch job runs. The scheduler requests a new slot, but the last usable checkpoint is old. A green retry badge can coexist with a late or incomplete customer result.
Visual: A retry event is not a recovered training job.
2. A retry hides three open questions.
The job may repeat completed steps, wait for a compatible replacement, and incur a second allocation. Whether either attempt is credited is contract-specific. The retry count cannot tell you how much useful work survived or what the customer owes.
Visual: Progress, waiting time, and net charge diverge.
3. Test the recovery path, not the badge.
Interrupt a worker in an approved test or local mock. Record checkpoint age, replacement time, repeated steps, final result hash, and net usage after exclusions or credits. A fault drill tests your harness; it does not recreate every provider reclaim policy.
Visual: Use a disposable job and a known checkpoint.
4. Demand a completed-run promise.
Put checkpoint responsibility, restart priority, lost-progress budget, and billing treatment in the capacity agreement. Do this because a cheap preemptible hour can become an expensive result when recovery is invisible.
Visual: The service must recover useful work after interruption.
Research and claim limits
- CoreWeave Spot Node Pools documentation (S178)
- CoreWeave Slurm Job Metrics (S179)
- CoreWeave Usage by product and zone documentation (S180)
The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.