FIELD NOTE / LINKEDIN
Full-model service is a fabric problem.
The complete written thought and the evidence behind it. The video edition will follow its public release.
The written argument is here.
This approved LinkedIn edition is on the journal now. Its video player and original platform link will appear after each public release is verified.
Full-model service is a fabric problem.
Video caption
Full-model service is a fabric problem.
The model may load somewhere; serving is a different promise.
Efficient K3 serving targets 64-plus accelerators.
My rule: Include accepted quality, latency, utilization and ownership cost.
#EricFieldNotes
Full written post / accessibility read
Kimi K3's model card lists two point eight trillion total parameters, one hundred four billion active per token and native four-bit weights. Sparse activation changes compute. The full expert set, memory capacity, movement and concurrent serving still require infrastructure.
At the ideal four-bit floor, weights alone are about one point four terabytes; a current Mac mini tops out at sixty-four gigabytes. The honest decision is not whether a clever offload trick can emit text. It is whether the full system meets latency, concurrency, privacy and quality obligations.
Moonshot recommends supernodes with sixty-four or more accelerators for efficient K3 inference. That is a vendor design target, not a universal minimum to run any derivative. It tells you that expert placement and high-bandwidth communication are first-order deployment choices.
For a bounded workload, benchmark a compact local model. For frontier-scale requests, benchmark a hosted or well-operated cluster. Include queue time, first-token delay, sustained tokens, concurrency, failures and owner effort. Do this because owning weights is only one line in the operating model.
#EricFieldNotes
Four-beat scene transcript
1. Full-model service is a fabric problem.
Kimi K3's model card lists two point eight trillion total parameters, one hundred four billion active per token and native four-bit weights. Sparse activation changes compute. The full expert set, memory capacity, movement and concurrent serving still require infrastructure.
Visual: Open weights shift the bill; they do not remove it.
2. A Mac mini comparison skips operations.
At the ideal four-bit floor, weights alone are about one point four terabytes; a current Mac mini tops out at sixty-four gigabytes. The honest decision is not whether a clever offload trick can emit text. It is whether the full system meets latency, concurrency, privacy and quality obligations.
Visual: The model may load somewhere; serving is a different promise.
3. Moonshot describes the deployment class.
Moonshot recommends supernodes with sixty-four or more accelerators for efficient K3 inference. That is a vendor design target, not a universal minimum to run any derivative. It tells you that expert placement and high-bandwidth communication are first-order deployment choices.
Visual: Efficient K3 serving targets 64-plus accelerators.
4. Compare delivered intelligence.
For a bounded workload, benchmark a compact local model. For frontier-scale requests, benchmark a hosted or well-operated cluster. Include queue time, first-token delay, sustained tokens, concurrency, failures and owner effort. Do this because owning weights is only one line in the operating model.
Visual: Include accepted quality, latency, utilization and ownership cost.
Research and claim limits
- Moonshot AI: Kimi K3 official model card (S128)
- Apple: current Mac mini technical specifications (S129)
- NVIDIA H200 specifications (S13)
The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.