FIELD NOTE / X
'AI made us faster' needs a denominator.
The complete written thought and the evidence behind it. The video edition will follow its public release.
The written argument is here.
This approved X edition is on the journal now. Its video player and original platform link will appear after each public release is verified.
'AI made us faster' needs a denominator.
Video caption
'AI made us faster' needs a denominator. Otherwise the denominator disappears. Judge the work that reached the user. #EricFieldNotes
Full written post / accessibility read
If an agent produces a patch in minutes and the team spends hours re-prompting, re-testing, re-deploying and checking what else broke, the first-diff clock tells the wrong story. The outcome is accepted product work after review, failures and recovery, not how fast text appeared.
For comparable tasks, record all agent attempts, human review, independent QA, elapsed time, incidents and the final accepted behavior. Count consciously avoided changes as decisions too. A benchmark that starts after prompt design or ends before deployment can make a costly loop look like a win.
Pair similar maintenance, feature and incident tasks and predefine what counts as accepted. Track cases routed away from the experiment and multiagent work that confuses time accounting. METR's later update explicitly warned about those selection and timing effects; do not turn either study into a timeless speed claim.
My rule: report cost and time per accepted change with rework and incident burden beside it. Do this because faster code generation only matters when the organization can decide, verify and ship it without a growing hidden queue.
#EricFieldNotes
Four-beat scene transcript
1. 'AI made us faster' needs a denominator.
If an agent produces a patch in minutes and the team spends hours re-prompting, re-testing, re-deploying and checking what else broke, the first-diff clock tells the wrong story. The outcome is accepted product work after review, failures and recovery, not how fast text appeared.
Visual: Time to first diff measures only the beginning.
2. Keep the rejected attempts in the ledger.
For comparable tasks, record all agent attempts, human review, independent QA, elapsed time, incidents and the final accepted behavior. Count consciously avoided changes as decisions too. A benchmark that starts after prompt design or ends before deployment can make a costly loop look like a win.
Visual: Otherwise the denominator disappears.
3. Compare task classes, not one average.
Pair similar maintenance, feature and incident tasks and predefine what counts as accepted. Track cases routed away from the experiment and multiagent work that confuses time accounting. METR's later update explicitly warned about those selection and timing effects; do not turn either study into a timeless speed claim.
Visual: Selection can reverse the apparent result.
4. Optimize the full delivery loop.
My rule: report cost and time per accepted change with rework and incident burden beside it. Do this because faster code generation only matters when the organization can decide, verify and ship it without a growing hidden queue.
Visual: Judge the work that reached the user.
Research and claim limits
- DORA State of AI-assisted Software Development 2025 (S152)
- METR early-2025 experienced-developer RCT (S153)
The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.