JournalDAY 49 / INSTAGRAM

FIELD NOTE / INSTAGRAM

Show both clocks on the AI chart.

The complete written thought and the evidence behind it. The video edition will follow its public release.

Journal September 25, 2026 · Instagram target November 15, 2026

Show both clocks on the AI chart.

Video caption

Show both clocks on the AI chart.

Task class and acceptance must match.

One trial cannot price every current team.

Make rework visible beside generated speed.

#EricFieldNotes

Full written post / accessibility read

A productivity chart that stops when the coding agent prints a diff ignores review, re-prompts, extra tests, deployment, rollback and the next team's rescue. Those costs may shrink or grow with AI. You cannot know from the first clock alone.

Compare similar tasks and define accepted behavior before the run. Include rejected attempts, human corrections, full QA, elapsed time and escaped defects. Separate routine changes from unfamiliar system work. Otherwise one average can be improved simply by selecting easier AI-friendly tasks or moving review outside the stopwatch.

METR's early trial found experienced maintainers slower on their own tasks with then-current tools. Its later update said newer data may show speedups but selection and time measurement made the estimate unreliable. DORA frames AI as an amplifier of the surrounding system. Neither yields your team's accepted-change cost.

Report accepted changes per unit of total engineering time, plus rework, defects and review backlog. Do this because a model can accelerate the first step while the organization still needs to prove that the finished product arrives sooner and safer. Let the local evidence decide.

#EricFieldNotes

Four-beat scene transcript

1. Show both clocks on the AI chart.

A productivity chart that stops when the coding agent prints a diff ignores review, re-prompts, extra tests, deployment, rollback and the next team's rescue. Those costs may shrink or grow with AI. You cannot know from the first clock alone.

Visual: First diff and accepted release are different.

2. Give every bar the same denominator.

Compare similar tasks and define accepted behavior before the run. Include rejected attempts, human corrections, full QA, elapsed time and escaped defects. Separate routine changes from unfamiliar system work. Otherwise one average can be improved simply by selecting easier AI-friendly tasks or moving review outside the stopwatch.

Visual: Task class and acceptance must match.

3. Read studies within their limits.

METR's early trial found experienced maintainers slower on their own tasks with then-current tools. Its later update said newer data may show speedups but selection and time measurement made the estimate unreliable. DORA frames AI as an amplifier of the surrounding system. Neither yields your team's accepted-change cost.

Visual: One trial cannot price every current team.

4. Publish the complete outcome.

Report accepted changes per unit of total engineering time, plus rework, defects and review backlog. Do this because a model can accelerate the first step while the organization still needs to prove that the finished product arrives sooner and safer. Let the local evidence decide.

Visual: Make rework visible beside generated speed.

Research and claim limits

The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.

More notes from the work ↗