JournalDAY 57 / X

FIELD NOTE / X

Publish the agent miss, not the polished replay.

The complete written thought and the evidence behind it. The video edition will follow its public release.

Journal September 25, 2026 · X target November 23, 2026

Publish the agent miss, not the polished replay.

Video caption

Publish the agent miss, not the polished replay. The model's final sentence cannot grade its own work. Do this because demos omit the recovery bill. #EricFieldNotes

Full written post / accessibility read

If you show one computer-use win, show the task set it came from. A smooth recording is a numerator. It does not reveal the attempts that failed, needed rescue or passed the screen while the durable result was wrong.

In a synthetic access request, an agent sees Saved, but a delayed policy check keeps the old role. A screenshot can be accurate about that moment and false as evidence of authorization. An independent postcondition exposes the gap.

Record the initial account state, exact task and model-tool version, bounded actions, action trace, read-only state check, fresh-session check and the operator time needed to recover. Strip private identifiers before sharing it.

Count accepted outcomes, failed outcomes and human rescue minutes over the same task set. Keep the bad run beside the good one. That denominator tells an engineering team whether autonomy is improving, rather than whether a demo editor is improving.

#EricFieldNotes

Four-beat scene transcript

1. Publish the agent miss, not the polished replay.

If you show one computer-use win, show the task set it came from. A smooth recording is a numerator. It does not reveal the attempts that failed, needed rescue or passed the screen while the durable result was wrong.

Visual: A beautiful trace can hide an unaccepted result.

2. Define the outcome before the run.

In a synthetic access request, an agent sees Saved, but a delayed policy check keeps the old role. A screenshot can be accurate about that moment and false as evidence of authorization. An independent postcondition exposes the gap.

Visual: The model's final sentence cannot grade its own work.

3. Release a reproducible failure packet.

Record the initial account state, exact task and model-tool version, bounded actions, action trace, read-only state check, fresh-session check and the operator time needed to recover. Strip private identifiers before sharing it.

Visual: Seed, tool route, trace, oracle and rescue belong together.

4. Report accepted tasks per attempt.

Count accepted outcomes, failed outcomes and human rescue minutes over the same task set. Keep the bad run beside the good one. That denominator tells an engineering team whether autonomy is improving, rather than whether a demo editor is improving.

Visual: Do this because demos omit the recovery bill.

Research and claim limits

The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.

More notes from the work ↗