JournalDAY 23 / LINKEDIN

FIELD NOTE / LINKEDIN

Buy the fabric and scheduler promise.

The complete written thought and the evidence behind it. The video edition will follow its public release.

Journal September 25, 2026 · LinkedIn target October 20, 2026

Buy the fabric and scheduler promise.

Video caption

Buy the fabric and scheduler promise.

The requested topology may not be available.

Start with vendor terms; inspect traces in a pilot.

My rule: Compare the service the workload consumes.

#EricFieldNotes

Full written post / accessibility read

Neocloud procurement often collapses to GPU count and hourly price. For a distributed job, I would buy a placement class, start-time envelope, communication path, recovery behavior and cost per completed result. Those determine whether capacity is usable.

Imagine a provider whose fleet is mostly occupied by single-device inference. The remaining cards are scattered. Your training job needs a connected allocation by Monday. The provider's average utilization can be excellent while your job waits and the promised deadline fails.

Before a contract, ask for the placement class, service terms and sample traces the vendor can share. In a scoped pilot, retain admission and placement decisions, device and NIC topology, collective diagnostics, application step time, checkpoint behavior and bill. Request a contended run only where the pilot terms permit it.

Contract for topology-valid start, tail completion and recovery, then compare delivered cost. Do this because a cheap device hour that cannot be assembled into your job can be more expensive than a dearer, predictable allocation.

#EricFieldNotes

Four-beat scene transcript

1. Buy the fabric and scheduler promise.

Neocloud procurement often collapses to GPU count and hourly price. For a distributed job, I would buy a placement class, start-time envelope, communication path, recovery behavior and cost per completed result. Those determine whether capacity is usable.

Visual: A utilization headline cannot guarantee a run.

2. A fleet can be busy and still miss you.

Imagine a provider whose fleet is mostly occupied by single-device inference. The remaining cards are scattered. Your training job needs a connected allocation by Monday. The provider's average utilization can be excellent while your job waits and the promised deadline fails.

Visual: The requested topology may not be available.

3. Stage an evidence packet.

Before a contract, ask for the placement class, service terms and sample traces the vendor can share. In a scoped pilot, retain admission and placement decisions, device and NIC topology, collective diagnostics, application step time, checkpoint behavior and bill. Request a contended run only where the pilot terms permit it.

Visual: Start with vendor terms; inspect traces in a pilot.

4. Price connected capacity.

Contract for topology-valid start, tail completion and recovery, then compare delivered cost. Do this because a cheap device hour that cannot be assembled into your job can be more expensive than a dearer, predictable allocation.

Visual: Compare the service the workload consumes.

Research and claim limits

The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.

More notes from the work ↗