FIELD NOTE / LINKEDIN
Choose spot by job, not by discount.
The complete written thought and the evidence behind it. The video edition will follow its public release.
The written argument is here.
This approved LinkedIn edition is on the journal now. Its video player and original platform link will appear after each public release is verified.
Choose spot by job, not by discount.
Video caption
Choose spot by job, not by discount.
Time to accepted output determines whether spot wins.
Hold task, checkpoint plan, and result test constant.
My rule: Use spot only when its recovery envelope fits.
#EricFieldNotes
Full written post / accessibility read
Preemptible capacity is a useful product, but I would not buy it for every workload. A checkpointable batch job and a deadline-critical launch have different tolerance for reclaim, queueing, and replay. Procurement needs a routing rule, not one GPU-hour price.
In a fictional spot run, one reclaim forces a checkpoint restore and a new wait. The hourly price remains low, but the completion deadline can fail. A non-preemptible allocation might cost more per hour and less per accepted on-time result.
For each job class, compare spot and a priced non-preemptible route using representative cases. Measure p ninety-five completion, accepted output, recovery effort, and net bill under each contract. Include an approved interruption drill; keep normal and faulted results separate.
Put checkpointable, interruption-tolerant jobs on spot when the measured tail and adjusted cost meet the promise. Give deadline-critical jobs a stronger reservation or fallback. Do this because the buyer pays for on-time accepted work, not nominal GPU occupancy.
#EricFieldNotes
Four-beat scene transcript
1. Choose spot by job, not by discount.
Preemptible capacity is a useful product, but I would not buy it for every workload. A checkpointable batch job and a deadline-critical launch have different tolerance for reclaim, queueing, and replay. Procurement needs a routing rule, not one GPU-hour price.
Visual: A cheap hour can be a costly deadline.
2. The discount changes after interruption.
In a fictional spot run, one reclaim forces a checkpoint restore and a new wait. The hourly price remains low, but the completion deadline can fail. A non-preemptible allocation might cost more per hour and less per accepted on-time result.
Visual: Time to accepted output determines whether spot wins.
3. Compare two capacity routes on one workload.
For each job class, compare spot and a priced non-preemptible route using representative cases. Measure p ninety-five completion, accepted output, recovery effort, and net bill under each contract. Include an approved interruption drill; keep normal and faulted results separate.
Visual: Hold task, checkpoint plan, and result test constant.
4. Route work to the capacity it can survive.
Put checkpointable, interruption-tolerant jobs on spot when the measured tail and adjusted cost meet the promise. Give deadline-critical jobs a stronger reservation or fallback. Do this because the buyer pays for on-time accepted work, not nominal GPU occupancy.
Visual: Use spot only when its recovery envelope fits.
Research and claim limits
- CoreWeave Spot Node Pools documentation (S178)
- CoreWeave Slurm Job Metrics (S179)
- CoreWeave Usage by product and zone documentation (S180)
The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.