JournalDAY 49 / X

FIELD NOTE / X

'AI made us faster' needs a denominator.

The complete written thought and the evidence behind it. The video edition will follow its public release.

Journal September 25, 2026 · X target November 15, 2026

'AI made us faster' needs a denominator.

Video caption

'AI made us faster' needs a denominator. Otherwise the denominator disappears. Judge the work that reached the user. #EricFieldNotes

Full written post / accessibility read

If an agent produces a patch in minutes and the team spends hours re-prompting, re-testing, re-deploying and checking what else broke, the first-diff clock tells the wrong story. The outcome is accepted product work after review, failures and recovery, not how fast text appeared.

For comparable tasks, record all agent attempts, human review, independent QA, elapsed time, incidents and the final accepted behavior. Count consciously avoided changes as decisions too. A benchmark that starts after prompt design or ends before deployment can make a costly loop look like a win.

Pair similar maintenance, feature and incident tasks and predefine what counts as accepted. Track cases routed away from the experiment and multiagent work that confuses time accounting. METR's later update explicitly warned about those selection and timing effects; do not turn either study into a timeless speed claim.

My rule: report cost and time per accepted change with rework and incident burden beside it. Do this because faster code generation only matters when the organization can decide, verify and ship it without a growing hidden queue.

#EricFieldNotes

Four-beat scene transcript

1. 'AI made us faster' needs a denominator.

If an agent produces a patch in minutes and the team spends hours re-prompting, re-testing, re-deploying and checking what else broke, the first-diff clock tells the wrong story. The outcome is accepted product work after review, failures and recovery, not how fast text appeared.

Visual: Time to first diff measures only the beginning.

2. Keep the rejected attempts in the ledger.

For comparable tasks, record all agent attempts, human review, independent QA, elapsed time, incidents and the final accepted behavior. Count consciously avoided changes as decisions too. A benchmark that starts after prompt design or ends before deployment can make a costly loop look like a win.

Visual: Otherwise the denominator disappears.

3. Compare task classes, not one average.

Pair similar maintenance, feature and incident tasks and predefine what counts as accepted. Track cases routed away from the experiment and multiagent work that confuses time accounting. METR's later update explicitly warned about those selection and timing effects; do not turn either study into a timeless speed claim.

Visual: Selection can reverse the apparent result.

4. Optimize the full delivery loop.

My rule: report cost and time per accepted change with rework and incident burden beside it. Do this because faster code generation only matters when the organization can decide, verify and ship it without a growing hidden queue.

Visual: Judge the work that reached the user.

Research and claim limits

The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.

More notes from the work ↗