FIELD NOTE / INSTAGRAM
Treat the leaderboard as a shortlist.
The complete written thought and the evidence behind it. The video edition will follow its public release.
The written argument is here.
This approved Instagram edition is on the journal now. Its video player and original platform link will appear after each public release is verified.
Treat the leaderboard as a shortlist.
Video caption
Treat the leaderboard as a shortlist.
Freshness, task quality and leakage matter.
Write the expected business behavior first.
The customer contract closes the loop.
#EricFieldNotes
Full written post / accessibility read
A coding-agent score can tell you which systems deserve a closer look. It cannot tell you whether the model preserved your customer's one-tenant migration exception, respected a fresh repository constraint or produced a release-ready artifact. The benchmark's tests and your product decision are different oracles.
SWE-bench Live was designed to update issue tasks and broaden repositories; its authors report a performance gap against static sets. SWE-Bench Pro Verified addresses leaked solution/evaluation information and flawed task instances. Those efforts make benchmarks more useful, while showing why one green number cannot be universal engineering proof.
Pick recent tasks in your repo. Store source decisions and acceptance fixtures outside the coding agent's editable path. Compare accepted patches, independent test results, review effort and correction loops. Add an adversarial case where a locally plausible patch violates a customer exception.
Use benchmark evidence for model selection, then require your own task-specific acceptance and post-deploy observation before release. Do this because the strongest general model can still answer the wrong requirement with a beautiful diff if the requirement was shortened upstream.
#EricFieldNotes
Four-beat scene transcript
1. Treat the leaderboard as a shortlist.
A coding-agent score can tell you which systems deserve a closer look. It cannot tell you whether the model preserved your customer's one-tenant migration exception, respected a fresh repository constraint or produced a release-ready artifact. The benchmark's tests and your product decision are different oracles.
Visual: It cannot approve your feature.
2. Ask what the benchmark actually tested.
SWE-bench Live was designed to update issue tasks and broaden repositories; its authors report a performance gap against static sets. SWE-Bench Pro Verified addresses leaked solution/evaluation information and flawed task instances. Those efforts make benchmarks more useful, while showing why one green number cannot be universal engineering proof.
Visual: Freshness, task quality and leakage matter.
3. Create your own protected slice.
Pick recent tasks in your repo. Store source decisions and acceptance fixtures outside the coding agent's editable path. Compare accepted patches, independent test results, review effort and correction loops. Add an adversarial case where a locally plausible patch violates a customer exception.
Visual: Write the expected business behavior first.
4. Shortlist publicly; decide privately.
Use benchmark evidence for model selection, then require your own task-specific acceptance and post-deploy observation before release. Do this because the strongest general model can still answer the wrong requirement with a beautiful diff if the requirement was shortened upstream.
Visual: The customer contract closes the loop.
Research and claim limits
- SWE-bench Live paper (S22)
- SWE-bench Pro Verified paper (S23)
The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.