JournalDAY 50 / TIKTOK

FIELD NOTE / TIKTOK

The agent aced a benchmark and broke the promise.

The complete written thought and the evidence behind it. The video edition will follow its public release.

Journal September 25, 2026 · TikTok target November 16, 2026

The agent aced a benchmark and broke the promise.

Video caption

The agent aced a benchmark and broke the promise. It implemented the wrong authority. Different claims need different evidence. #EricFieldNotes

Full written post / accessibility read

Fictional case: an enterprise customer may use a feature during one migration. Everyone else needs a paid entitlement. The agent is strong on public code benchmarks. A shortened task says make the feature available to enterprise tenants. It does that, the generated tests pass, and the exception becomes a broad rollout.

This is not a trivial missing if statement. The product decision was compressed before the agent saw it, and the tests measure the compressed version. A second agent reading only the patch may agree. The failure belongs to the source-to-acceptance chain that the public benchmark never exercised.

Give the worker the task, but give an independent grader the source-linked customer decision, tenant IDs, expiry and forbidden cohort. Test both the migration tenant's access and a non-entitled tenant's denial on the exact build. Remove the entitlement guard in a disposable mutant and demand a red gate.

My rule: use public scores to choose models for trial, and release only after your own protected product contract passes on the exact artifact and cohort. Do this because being good at benchmark issues does not give an agent authority to rewrite what your customer was promised.

#EricFieldNotes

Four-beat scene transcript

1. The agent aced a benchmark and broke the promise.

Fictional case: an enterprise customer may use a feature during one migration. Everyone else needs a paid entitlement. The agent is strong on public code benchmarks. A shortened task says make the feature available to enterprise tenants. It does that, the generated tests pass, and the exception becomes a broad rollout.

Visual: A leaderboard never saw this customer exception.

2. The code may be coherent.

This is not a trivial missing if statement. The product decision was compressed before the agent saw it, and the tests measure the compressed version. A second agent reading only the patch may agree. The failure belongs to the source-to-acceptance chain that the public benchmark never exercised.

Visual: It implemented the wrong authority.

3. Protect the original clause.

Give the worker the task, but give an independent grader the source-linked customer decision, tenant IDs, expiry and forbidden cohort. Test both the migration tenant's access and a non-entitled tenant's denial on the exact build. Remove the entitlement guard in a disposable mutant and demand a red gate.

Visual: Separate source, worker and grader.

4. A benchmark selects; the gate releases.

My rule: use public scores to choose models for trial, and release only after your own protected product contract passes on the exact artifact and cohort. Do this because being good at benchmark issues does not give an agent authority to rewrite what your customer was promised.

Visual: Different claims need different evidence.

Research and claim limits

The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.

More notes from the work ↗