FIELD NOTE / TIKTOK
One H200 does not become seven H200s.
The complete written thought and the evidence behind it. The video edition will follow its public release.
The written argument is here.
This approved TikTok edition is on the journal now. Its video player and original platform link will appear after each public release is verified.
One H200 does not become seven H200s.
Video caption
One H200 does not become seven H200s. A dashboard may show fragments, not a legal slot. Keep a whole-card escape route for large jobs. #EricFieldNotes
Full written post / accessibility read
NVIDIA says an H200 can support up to seven MIG instances. Watch the sleight of hand: seven isolated slices are useful, but the card still has one physical memory and compute budget. A large job cannot assemble arbitrary scraps into a whole GPU.
In this simulated cluster, small inference jobs occupy the offered profiles. A larger model needs a different profile. There may be free resource fragments, yet no compatible instance to admit the request. Saying 'capacity remains' does not answer when the job can run.
In a disposable pool, place the real small jobs, then submit the large one. Record the legal profile map, wait, any reconfiguration, tenant interruption and actual application throughput. Compare that with leaving some whole cards unsliced.
MIG improves sharing when the job shape matches a supported profile. Reserve or reclaim whole-device capacity when the workload needs it, under an explicit policy. Do this because isolation is a benefit; multiplication is an illusion.
#EricFieldNotes
Four-beat scene transcript
1. One H200 does not become seven H200s.
NVIDIA says an H200 can support up to seven MIG instances. Watch the sleight of hand: seven isolated slices are useful, but the card still has one physical memory and compute budget. A large job cannot assemble arbitrary scraps into a whole GPU.
Visual: Seven MIG instances share one physical card.
2. The eighth workload changes the shape.
In this simulated cluster, small inference jobs occupy the offered profiles. A larger model needs a different profile. There may be free resource fragments, yet no compatible instance to admit the request. Saying 'capacity remains' does not answer when the job can run.
Visual: A dashboard may show fragments, not a legal slot.
3. Run a profile-switch test.
In a disposable pool, place the real small jobs, then submit the large one. Record the legal profile map, wait, any reconfiguration, tenant interruption and actual application throughput. Compare that with leaving some whole cards unsliced.
Visual: Track occupancy, admission and application speed.
4. Use slices where they fit.
MIG improves sharing when the job shape matches a supported profile. Reserve or reclaim whole-device capacity when the workload needs it, under an explicit policy. Do this because isolation is a benefit; multiplication is an illusion.
Visual: Keep a whole-card escape route for large jobs.
Research and claim limits
The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.