JournalDAY 23 / INSTAGRAM

FIELD NOTE / INSTAGRAM

Eight free GPUs may be zero eight-GPU slots.

The complete written thought and the evidence behind it. The video edition will follow its public release.

Journal September 25, 2026 · Instagram target October 20, 2026

Eight free GPUs may be zero eight-GPU slots.

Video caption

Eight free GPUs may be zero eight-GPU slots.

Idle cards remain visible while the queue grows.

Kueue can express domain fit; the app measures performance.

Publish wait by workload shape, not only idle GPU count.

#EricFieldNotes

Full written post / accessibility read

Imagine four free GPUs in one fabric domain and four in another. Your eight-GPU job requires a specific connected placement. The dashboard counts eight idle devices; the scheduler may correctly say no eligible allocation exists.

From the buyer's view, availability was promised and the run has not started. From the operator's view, the devices are fragments. Both can be true. The conflict disappears only when the contract defines what counts as a usable slot for this workload.

Record the job's actual placement rule. Submit it against a disposable fragmented cluster and inspect admission. If cross-domain placement is allowed, run network and NCCL probes plus the training step to see whether it meets the application's limit. Do not equate a scheduler label with measured fabric quality.

Show time to an eligible placement for each customer job class and the completed-run distribution on that placement. Do this because eight free devices only matter when they can become the eight-device service the customer bought.

#EricFieldNotes

Four-beat scene transcript

1. Eight free GPUs may be zero eight-GPU slots.

Imagine four free GPUs in one fabric domain and four in another. Your eight-GPU job requires a specific connected placement. The dashboard counts eight idle devices; the scheduler may correctly say no eligible allocation exists.

Visual: The topology matters before the job can start.

2. The customer sees a contradiction.

From the buyer's view, availability was promised and the run has not started. From the operator's view, the devices are fragments. Both can be true. The conflict disappears only when the contract defines what counts as a usable slot for this workload.

Visual: Idle cards remain visible while the queue grows.

3. Test the topology requirement.

Record the job's actual placement rule. Submit it against a disposable fragmented cluster and inspect admission. If cross-domain placement is allowed, run network and NCCL probes plus the training step to see whether it meets the application's limit. Do not equate a scheduler label with measured fabric quality.

Visual: Kueue can express domain fit; the app measures performance.

4. Report startable capacity.

Show time to an eligible placement for each customer job class and the completed-run distribution on that placement. Do this because eight free devices only matter when they can become the eight-device service the customer bought.

Visual: Publish wait by workload shape, not only idle GPU count.

Research and claim limits

The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.

More notes from the work ↗