JournalDAY 23 / X

FIELD NOTE / X

A visible GPU is not a startable job.

The complete written thought and the evidence behind it. The video edition will follow its public release.

Journal September 25, 2026 · X target October 20, 2026

A visible GPU is not a startable job.

Day 23 · Week 4 editorial group · X · no publication date or time assigned

Video caption

A visible GPU is not a startable job. The tenant waits while the dashboard shows idle hardware. A device count is only the first input. #EricFieldNotes

Full written post / accessibility read

A new node advertises eight GPUs. Sales calls it capacity. But the driver validator, network path, quota admission, node placement and application still have to agree. Counting allocatable devices as ready eight-GPU jobs skips the transitions that fail in practice.

Imagine Kueue admits the workload, then a node cannot launch the CUDA container after a driver change. Or quota is free but the required rack shape is not. The two failures need different owners. A single available-GPU number hides both.

For each offered class, record the resource flavor and quota decision, topology domain, device IDs, operator validation state and a small disposable workload result. Keep the failed stage and timestamp. A buyer need not see raw cluster access to receive a truthful status.

Publish pending, admitted, placed and successfully started counts by workload class. Investigate every gap at its actual layer. Do this because a GPU becomes a compute product only when the promised job can start on a verified placement.

#EricFieldNotes

Four-beat scene transcript

1. A visible GPU is not a startable job.

A new node advertises eight GPUs. Sales calls it capacity. But the driver validator, network path, quota admission, node placement and application still have to agree. Counting allocatable devices as ready eight-GPU jobs skips the transitions that fail in practice.

Visual: The device plugin and customer promise live at different layers.

2. A green count can make a late start invisible.

Imagine Kueue admits the workload, then a node cannot launch the CUDA container after a driver change. Or quota is free but the required rack shape is not. The two failures need different owners. A single available-GPU number hides both.

Visual: The tenant waits while the dashboard shows idle hardware.

3. Build a start receipt.

For each offered class, record the resource flavor and quota decision, topology domain, device IDs, operator validation state and a small disposable workload result. Keep the failed stage and timestamp. A buyer need not see raw cluster access to receive a truthful status.

Visual: Bind allocation, policy, node and observed launch.

4. Sell measured startable capacity.

Publish pending, admitted, placed and successfully started counts by workload class. Investigate every gap at its actual layer. Do this because a GPU becomes a compute product only when the promised job can start on a verified placement.

Visual: A device count is only the first input.

Research and claim limits

The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.

More notes from the work ↗