FIELD NOTE / X
Put model choice inside the experiment.
The complete written thought and the evidence behind it. The video edition will follow its public release.
The written argument is here.
This approved X edition is on the journal now. Its video player and original platform link will appear after each public release is verified.
Put model choice inside the experiment.
Video caption
Put model choice inside the experiment. Model, harness, tools and oracle interact. The denominator includes correction. #EricFieldNotes
Full written post / accessibility read
One model may write good repository patches and another may navigate a browser reliably. A general leaderboard cannot tell you which delivers accepted outcomes in your repo, UI and service contract. Routing needs task-specific quality, latency, cost and human rescue, not model loyalty.
Compare models on the same recent tasks, budgets, permissions and independent acceptance. Include retries, review, browser postconditions and failed attempts. A change in tool wrapper can change the result even with identical model weights. Record versions and the full outcome distribution by task class.
Set a quality threshold and data boundary for each task class. Route to the least costly eligible model that passes the complete acceptance path; escalate ambiguous cases to a stronger route or human owner. Re-test after provider or harness changes because yesterday's measurement can go stale.
My rule: choose models by cost and time per accepted task with human rescue and failure visible. Do this because a cheap or highly ranked model can be the wrong choice when it increases review, browser errors or deployment rework.
#EricFieldNotes
Four-beat scene transcript
1. Put model choice inside the experiment.
One model may write good repository patches and another may navigate a browser reliably. A general leaderboard cannot tell you which delivers accepted outcomes in your repo, UI and service contract. Routing needs task-specific quality, latency, cost and human rescue, not model loyalty.
Visual: The strongest logo may lose a task slice.
2. Pin the whole path.
Compare models on the same recent tasks, budgets, permissions and independent acceptance. Include retries, review, browser postconditions and failed attempts. A change in tool wrapper can change the result even with identical model weights. Record versions and the full outcome distribution by task class.
Visual: Model, harness, tools and oracle interact.
3. Route only proven classes.
Set a quality threshold and data boundary for each task class. Route to the least costly eligible model that passes the complete acceptance path; escalate ambiguous cases to a stronger route or human owner. Re-test after provider or harness changes because yesterday's measurement can go stale.
Visual: Unknown cases need a safe exit.
4. Buy accepted work, not benchmark rank.
My rule: choose models by cost and time per accepted task with human rescue and failure visible. Do this because a cheap or highly ranked model can be the wrong choice when it increases review, browser errors or deployment rework.
Visual: The denominator includes correction.
Research and claim limits
- OSWorld 2.0 paper (S01)
- SWE-bench Live paper (S22)
- SWE-bench Pro Verified paper (S23)
The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.