FIELD NOTE / INSTAGRAM
Changing MIG geometry is a maintenance event.
The complete written thought and the evidence behind it. The video edition will follow its public release.
The written argument is here.
This approved Instagram edition is on the journal now. Its video player and original platform link will appear after each public release is verified.
Changing MIG geometry is a maintenance event.
Day 25 · Week 4 editorial group · Instagram · no publication date or time assigned
Video caption
Changing MIG geometry is a maintenance event.
Small inference reservations block a larger request.
Reserve whole-card escape capacity and price conversion time.
The capacity promise includes transition cost.
#EricFieldNotes
Full written post / accessibility read
A cloud can divide supported GPUs into MIG profiles. But NVIDIA's MIG Manager requires user GPU workloads to be absent while it changes configuration, stops operator GPU pods during the change, and may need a reboot. A new shape is not a zero-cost dashboard edit.
Imagine morning traffic fills slices. An afternoon customer wants a whole-card training run. A dashboard can show fragments, but taking down the slices to reconfigure would interrupt paying users. The scheduling decision became a product decision earlier than the request.
Group nodes by stable profile family. Set a maintenance window and tenant notice for geometry changes; simulate a demand flip on non-production nodes. Record drain time, interrupted work, reconfiguration, validation and time until the next job starts.
Quote the profile and its change window, or keep a separate whole-card pool. Do this because fragmentation is not only idle hardware; it is the time and tenant impact of converting that hardware into the shape someone needs.
#EricFieldNotes
Four-beat scene transcript
1. Changing MIG geometry is a maintenance event.
A cloud can divide supported GPUs into MIG profiles. But NVIDIA's MIG Manager requires user GPU workloads to be absent while it changes configuration, stops operator GPU pods during the change, and may need a reboot. A new shape is not a zero-cost dashboard edit.
Visual: The partition plan cannot change freely between customers.
2. Demand shifts faster than the hardware shape.
Imagine morning traffic fills slices. An afternoon customer wants a whole-card training run. A dashboard can show fragments, but taking down the slices to reconfigure would interrupt paying users. The scheduling decision became a product decision earlier than the request.
Visual: Small inference reservations block a larger request.
3. Operate profile pools explicitly.
Group nodes by stable profile family. Set a maintenance window and tenant notice for geometry changes; simulate a demand flip on non-production nodes. Record drain time, interrupted work, reconfiguration, validation and time until the next job starts.
Visual: Reserve whole-card escape capacity and price conversion time.
4. Do not sell an instant shape you cannot make.
Quote the profile and its change window, or keep a separate whole-card pool. Do this because fragmentation is not only idle hardware; it is the time and tenant impact of converting that hardware into the shape someone needs.
Visual: The capacity promise includes transition cost.
Research and claim limits
The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.