FIELD NOTE / X
The network is part of the model budget.
The complete written thought and the evidence behind it. The video edition will follow its public release.
The written argument is here.
This approved X edition is on the journal now. Its video player and original platform link will appear after each public release is verified.
The network is part of the model budget.
Video caption
The network is part of the model budget. Topology and competing traffic change the communication path. Pay for the topology that finishes the job. #EricFieldNotes
Full written post / accessibility read
In a multi-GPU run, each accelerator can be fast while the job waits on gradients, activations or remote experts. GPU model and hourly rate leave out peer paths, NIC placement and collective behavior. That missing fabric can dominate the completed-run cost.
NVIDIA documents that virtualized placement can expose uneven peer bandwidth in supported H100 and H200 systems. A count of eight does not say which cards can communicate directly or what happens under concurrent tenants.
On the allocated nodes, record GPU and NIC topology, local and cross-node bandwidth, representative NCCL collective latency and the application's step-time distribution. Repeat under the contention class you plan to buy. Microbenchmarks locate faults; the application supplies the verdict.
Compare providers or placements on accepted work per total dollar and p ninety-five completion, with the network included. Do this because an isolated GPU is only a component; the connected allocation is the compute product.
#EricFieldNotes
Four-beat scene transcript
1. The network is part of the model budget.
In a multi-GPU run, each accelerator can be fast while the job waits on gradients, activations or remote experts. GPU model and hourly rate leave out peer paths, NIC placement and collective behavior. That missing fabric can dominate the completed-run cost.
Visual: Device count alone cannot price distributed work.
2. The same eight cards can behave differently.
NVIDIA documents that virtualized placement can expose uneven peer bandwidth in supported H100 and H200 systems. A count of eight does not say which cards can communicate directly or what happens under concurrent tenants.
Visual: Topology and competing traffic change the communication path.
3. Probe from wire to application.
On the allocated nodes, record GPU and NIC topology, local and cross-node bandwidth, representative NCCL collective latency and the application's step-time distribution. Repeat under the contention class you plan to buy. Microbenchmarks locate faults; the application supplies the verdict.
Visual: Topology, bandwidth, NCCL and actual step time.
4. Price connected capacity.
Compare providers or placements on accepted work per total dollar and p ninety-five completion, with the network included. Do this because an isolated GPU is only a component; the connected allocation is the compute product.
Visual: Pay for the topology that finishes the job.
Research and claim limits
- NVIDIA NCCL performance and tuning (S122)
- NVIDIA Fabric Manager (S126)
The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.