FIELD NOTE / LINKEDIN
Price connected capacity, not isolated cards.
The complete written thought and the evidence behind it. The video edition will follow its public release.
The written argument is here.
This approved LinkedIn edition is on the journal now. Its video player and original platform link will appear after each public release is verified.
Price connected capacity, not isolated cards.
Video caption
Price connected capacity, not isolated cards.
Virtual topology and sharing change the run.
Represent the job, then inspect the whole path.
My rule: Start, topology, p95 finish and cost per result.
#EricFieldNotes
Full written post / accessibility read
For distributed AI, a GPU is useful as part of a placement with peer links, NIC access and a scheduler that can assemble the requested group. Quoting device count and hourly price without that connection class is like quoting servers without the network they need to form a service.
NVIDIA's own guidance calls for local, network and collective diagnostics; Fabric Manager describes asymmetric peer bandwidth in some virtualized H100 and H200 layouts. A buyer who never tests the actual placement may price a theoretical system rather than the one delivered.
Ask for a proposed multi-node placement class before contracting. In a scoped pilot allocation, capture physical and virtual topology, local and cross-node bandwidth, NCCL performance, application step time, recovery and total bill under agreed load. Keep every metric attached to the same allocation ID.
Ask the supplier to specify connected placement and test it at the concurrency you intend to buy. Compare completed accepted runs and recovery behavior, not just card hours. Do this because the fabric is part of the delivered intelligence.
#EricFieldNotes
Four-beat scene transcript
1. Price connected capacity, not isolated cards.
For distributed AI, a GPU is useful as part of a placement with peer links, NIC access and a scheduler that can assemble the requested group. Quoting device count and hourly price without that connection class is like quoting servers without the network they need to form a service.
Visual: The model sees a fabric the quote may omit.
2. Connection cost can dominate tail time.
NVIDIA's own guidance calls for local, network and collective diagnostics; Fabric Manager describes asymmetric peer bandwidth in some virtualized H100 and H200 layouts. A buyer who never tests the actual placement may price a theoretical system rather than the one delivered.
Visual: Virtual topology and sharing change the run.
3. Write a pilot placement acceptance test.
Ask for a proposed multi-node placement class before contracting. In a scoped pilot allocation, capture physical and virtual topology, local and cross-node bandwidth, NCCL performance, application step time, recovery and total bill under agreed load. Keep every metric attached to the same allocation ID.
Visual: Represent the job, then inspect the whole path.
4. Contract for the envelope.
Ask the supplier to specify connected placement and test it at the concurrency you intend to buy. Compare completed accepted runs and recovery behavior, not just card hours. Do this because the fabric is part of the delivered intelligence.
Visual: Start, topology, p95 finish and cost per result.
Research and claim limits
- NVIDIA NCCL performance and tuning (S122)
- NVIDIA Fabric Manager (S126)
The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.