FIELD NOTE / INSTAGRAM
Show both clocks on the AI chart.
The complete written thought and the evidence behind it. The video edition will follow its public release.
The written argument is here.
This approved Instagram edition is on the journal now. Its video player and original platform link will appear after each public release is verified.
Show both clocks on the AI chart.
Video caption
Show both clocks on the AI chart.
Task class and acceptance must match.
One trial cannot price every current team.
Make rework visible beside generated speed.
#EricFieldNotes
Full written post / accessibility read
A productivity chart that stops when the coding agent prints a diff ignores review, re-prompts, extra tests, deployment, rollback and the next team's rescue. Those costs may shrink or grow with AI. You cannot know from the first clock alone.
Compare similar tasks and define accepted behavior before the run. Include rejected attempts, human corrections, full QA, elapsed time and escaped defects. Separate routine changes from unfamiliar system work. Otherwise one average can be improved simply by selecting easier AI-friendly tasks or moving review outside the stopwatch.
METR's early trial found experienced maintainers slower on their own tasks with then-current tools. Its later update said newer data may show speedups but selection and time measurement made the estimate unreliable. DORA frames AI as an amplifier of the surrounding system. Neither yields your team's accepted-change cost.
Report accepted changes per unit of total engineering time, plus rework, defects and review backlog. Do this because a model can accelerate the first step while the organization still needs to prove that the finished product arrives sooner and safer. Let the local evidence decide.
#EricFieldNotes
Four-beat scene transcript
1. Show both clocks on the AI chart.
A productivity chart that stops when the coding agent prints a diff ignores review, re-prompts, extra tests, deployment, rollback and the next team's rescue. Those costs may shrink or grow with AI. You cannot know from the first clock alone.
Visual: First diff and accepted release are different.
2. Give every bar the same denominator.
Compare similar tasks and define accepted behavior before the run. Include rejected attempts, human corrections, full QA, elapsed time and escaped defects. Separate routine changes from unfamiliar system work. Otherwise one average can be improved simply by selecting easier AI-friendly tasks or moving review outside the stopwatch.
Visual: Task class and acceptance must match.
3. Read studies within their limits.
METR's early trial found experienced maintainers slower on their own tasks with then-current tools. Its later update said newer data may show speedups but selection and time measurement made the estimate unreliable. DORA frames AI as an amplifier of the surrounding system. Neither yields your team's accepted-change cost.
Visual: One trial cannot price every current team.
4. Publish the complete outcome.
Report accepted changes per unit of total engineering time, plus rework, defects and review backlog. Do this because a model can accelerate the first step while the organization still needs to prove that the finished product arrives sooner and safer. Let the local evidence decide.
Visual: Make rework visible beside generated speed.
Research and claim limits
- DORA State of AI-assisted Software Development 2025 (S152)
- METR early-2025 experienced-developer RCT (S153)
The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.