FIELD NOTE / INSTAGRAM
A success chart needs four bars.
The complete written thought and the evidence behind it. The video edition will follow its public release.
The written argument is here.
This approved Instagram edition is on the journal now. Its video player and original platform link will appear after each public release is verified.
A success chart needs four bars.
Video caption
A success chart needs four bars.
A person may finish a task after the agent stops.
Use one read-only oracle and a fresh user session.
Do this because the denominator controls the claim.
#EricFieldNotes
Full written post / accessibility read
A team can report ninety completed computer-use runs and still have no idea how many held up after a fresh login. Completion is an agent event. Acceptance is an independently checked product event.
Suppose the agent navigates correctly but stalls at a native permission prompt. If a human finishes the last step, count a rescued task and the minutes spent. Calling it autonomous success erases the capacity you had to provide.
For a synthetic role-change workflow, the chart's verified bar requires an audit record tied to the request and a new session in the affected role. The green screen is kept as trace evidence, not used as the whole score.
Put attempted, completed, independently accepted, rescued and reversed tasks on one card. Compare versions on the same fixtures. The point is not to embarrass the agent; it is to see whether automation reduces total work for the team.
#EricFieldNotes
Four-beat scene transcript
1. A success chart needs four bars.
A team can report ninety completed computer-use runs and still have no idea how many held up after a fresh login. Completion is an agent event. Acceptance is an independently checked product event.
Visual: Completed, failed, rescued and verified are different counts.
2. Rescue can be the hidden operating cost.
Suppose the agent navigates correctly but stalls at a native permission prompt. If a human finishes the last step, count a rescued task and the minutes spent. Calling it autonomous success erases the capacity you had to provide.
Visual: A person may finish a task after the agent stops.
3. Pair the chart with proof.
For a synthetic role-change workflow, the chart's verified bar requires an audit record tied to the request and a new session in the affected role. The green screen is kept as trace evidence, not used as the whole score.
Visual: Use one read-only oracle and a fresh user session.
4. Publish the whole funnel.
Put attempted, completed, independently accepted, rescued and reversed tasks on one card. Compare versions on the same fixtures. The point is not to embarrass the agent; it is to see whether automation reduces total work for the team.
Visual: Do this because the denominator controls the claim.
Research and claim limits
- OpenAI Computer use API guide (S159)
- OSWorld 2.1 official repository (S162)
The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.