JournalDAY 50 / INSTAGRAM

FIELD NOTE / INSTAGRAM

Treat the leaderboard as a shortlist.

The complete written thought and the evidence behind it. The video edition will follow its public release.

Journal September 25, 2026 · Instagram target November 16, 2026

Treat the leaderboard as a shortlist.

Video caption

Treat the leaderboard as a shortlist.

Freshness, task quality and leakage matter.

Write the expected business behavior first.

The customer contract closes the loop.

#EricFieldNotes

Full written post / accessibility read

A coding-agent score can tell you which systems deserve a closer look. It cannot tell you whether the model preserved your customer's one-tenant migration exception, respected a fresh repository constraint or produced a release-ready artifact. The benchmark's tests and your product decision are different oracles.

SWE-bench Live was designed to update issue tasks and broaden repositories; its authors report a performance gap against static sets. SWE-Bench Pro Verified addresses leaked solution/evaluation information and flawed task instances. Those efforts make benchmarks more useful, while showing why one green number cannot be universal engineering proof.

Pick recent tasks in your repo. Store source decisions and acceptance fixtures outside the coding agent's editable path. Compare accepted patches, independent test results, review effort and correction loops. Add an adversarial case where a locally plausible patch violates a customer exception.

Use benchmark evidence for model selection, then require your own task-specific acceptance and post-deploy observation before release. Do this because the strongest general model can still answer the wrong requirement with a beautiful diff if the requirement was shortened upstream.

#EricFieldNotes

Four-beat scene transcript

1. Treat the leaderboard as a shortlist.

A coding-agent score can tell you which systems deserve a closer look. It cannot tell you whether the model preserved your customer's one-tenant migration exception, respected a fresh repository constraint or produced a release-ready artifact. The benchmark's tests and your product decision are different oracles.

Visual: It cannot approve your feature.

2. Ask what the benchmark actually tested.

SWE-bench Live was designed to update issue tasks and broaden repositories; its authors report a performance gap against static sets. SWE-Bench Pro Verified addresses leaked solution/evaluation information and flawed task instances. Those efforts make benchmarks more useful, while showing why one green number cannot be universal engineering proof.

Visual: Freshness, task quality and leakage matter.

3. Create your own protected slice.

Pick recent tasks in your repo. Store source decisions and acceptance fixtures outside the coding agent's editable path. Compare accepted patches, independent test results, review effort and correction loops. Add an adversarial case where a locally plausible patch violates a customer exception.

Visual: Write the expected business behavior first.

4. Shortlist publicly; decide privately.

Use benchmark evidence for model selection, then require your own task-specific acceptance and post-deploy observation before release. Do this because the strongest general model can still answer the wrong requirement with a beautiful diff if the requirement was shortened upstream.

Visual: The customer contract closes the loop.

Research and claim limits

The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.

More notes from the work ↗