FIELD NOTE / X
A spot GPU hour needs a failure budget.
The complete written thought and the evidence behind it. The video edition will follow its public release.
The written argument is here.
This approved X edition is on the journal now. Its video player and original platform link will appear after each public release is verified.
A spot GPU hour needs a failure budget.
Video caption
A spot GPU hour needs a failure budget. Progress and the invoice follow different rules. Buy the completed run, including its failure path. #EricFieldNotes
Full written post / accessibility read
Consider a fictional training run on a preemptible GPU pool. The lower hourly rate buys capacity the provider may reclaim. A green reschedule event says nothing about checkpoint age, queue delay, or when the customer gets an accepted model.
The fictional job is reclaimed just before a checkpoint. It resumes from an older snapshot and repeats work. Whether both allocations are charged, excluded, or credited depends on the provider's contract. Record lost progress and net billed usage separately.
In a disposable test, interrupt one worker through an approved fault route or local mock. Measure the last valid checkpoint, replacement wait, replayed steps, final result, and metered versus net billed usage. Tie the job, scheduler, and invoice records to one run ID.
Ask for reclaim notice, restart priority, recoverable progress, credits, and p ninety-five cost per accepted run. Do this because a low spot rate is useful only when the recovery and billing terms still deliver the customer's job on budget.
#EricFieldNotes
Four-beat scene transcript
1. A spot GPU hour needs a failure budget.
Consider a fictional training run on a preemptible GPU pool. The lower hourly rate buys capacity the provider may reclaim. A green reschedule event says nothing about checkpoint age, queue delay, or when the customer gets an accepted model.
Visual: A cheap allocation can lose expensive progress.
2. The replay can consume another allocation.
The fictional job is reclaimed just before a checkpoint. It resumes from an older snapshot and repeats work. Whether both allocations are charged, excluded, or credited depends on the provider's contract. Record lost progress and net billed usage separately.
Visual: Progress and the invoice follow different rules.
3. Rehearse the recovery packet.
In a disposable test, interrupt one worker through an approved fault route or local mock. Measure the last valid checkpoint, replacement wait, replayed steps, final result, and metered versus net billed usage. Tie the job, scheduler, and invoice records to one run ID.
Visual: Detect, restore, finish, and reconcile.
4. Contract for recovered work.
Ask for reclaim notice, restart priority, recoverable progress, credits, and p ninety-five cost per accepted run. Do this because a low spot rate is useful only when the recovery and billing terms still deliver the customer's job on budget.
Visual: Buy the completed run, including its failure path.
Research and claim limits
- CoreWeave Spot Node Pools documentation (S178)
- CoreWeave Slurm Job Metrics (S179)
- CoreWeave Usage by product and zone documentation (S180)
The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.