JournalDAY 28 / X

FIELD NOTE / X

The network is part of the model budget.

The complete written thought and the evidence behind it. The video edition will follow its public release.

Journal September 25, 2026 · X target October 25, 2026

The network is part of the model budget.

Video caption

The network is part of the model budget. Topology and competing traffic change the communication path. Pay for the topology that finishes the job. #EricFieldNotes

Full written post / accessibility read

In a multi-GPU run, each accelerator can be fast while the job waits on gradients, activations or remote experts. GPU model and hourly rate leave out peer paths, NIC placement and collective behavior. That missing fabric can dominate the completed-run cost.

NVIDIA documents that virtualized placement can expose uneven peer bandwidth in supported H100 and H200 systems. A count of eight does not say which cards can communicate directly or what happens under concurrent tenants.

On the allocated nodes, record GPU and NIC topology, local and cross-node bandwidth, representative NCCL collective latency and the application's step-time distribution. Repeat under the contention class you plan to buy. Microbenchmarks locate faults; the application supplies the verdict.

Compare providers or placements on accepted work per total dollar and p ninety-five completion, with the network included. Do this because an isolated GPU is only a component; the connected allocation is the compute product.

#EricFieldNotes

Four-beat scene transcript

1. The network is part of the model budget.

In a multi-GPU run, each accelerator can be fast while the job waits on gradients, activations or remote experts. GPU model and hourly rate leave out peer paths, NIC placement and collective behavior. That missing fabric can dominate the completed-run cost.

Visual: Device count alone cannot price distributed work.

2. The same eight cards can behave differently.

NVIDIA documents that virtualized placement can expose uneven peer bandwidth in supported H100 and H200 systems. A count of eight does not say which cards can communicate directly or what happens under concurrent tenants.

Visual: Topology and competing traffic change the communication path.

3. Probe from wire to application.

On the allocated nodes, record GPU and NIC topology, local and cross-node bandwidth, representative NCCL collective latency and the application's step-time distribution. Repeat under the contention class you plan to buy. Microbenchmarks locate faults; the application supplies the verdict.

Visual: Topology, bandwidth, NCCL and actual step time.

4. Price connected capacity.

Compare providers or placements on accepted work per total dollar and p ninety-five completion, with the network included. Do this because an isolated GPU is only a component; the connected allocation is the compute product.

Visual: Pay for the topology that finishes the job.

Research and claim limits

The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.

More notes from the work ↗