FIELD NOTE / X
Autonomy is earned by accepted work.
The complete written thought and the evidence behind it. The video edition will follow its public release.
The written argument is here.
This approved X edition is on the journal now. Its video player and original platform link will appear after each public release is verified.
Autonomy is earned by accepted work.
Video caption
Autonomy is earned by accepted work. Exceptions consume the owner's attention later. Do this because activity is cheap to fake. #EricFieldNotes
Full written post / accessibility read
More agent sessions can produce more traces and more human exception work. I care about the amount of verified work that survives without hidden rescue or reversal. That is the numerator and denominator an autonomy claim should expose.
Consider a synthetic back-office queue: the agent processes routine requests quickly, but one wrong-target action takes hours to investigate and reverse. Average click speed never shows that cost. Count exceptions by severity and recovery time.
Tie each attempt to a request ID, independent acceptance result, human minutes, elapsed time and customer effect. Keep failed runs in the sample. Compare candidate models and harness versions only on the same task slices and rules.
Let the agent act within a measured scope; expand only when accepted throughput rises and recovery cost stays bounded. A fleet that looks busy but moves exceptions into people has not earned more authority.
#EricFieldNotes
Four-beat scene transcript
1. Autonomy is earned by accepted work.
More agent sessions can produce more traces and more human exception work. I care about the amount of verified work that survives without hidden rescue or reversal. That is the numerator and denominator an autonomy claim should expose.
Visual: Agent sessions and generated actions are activity, not value.
2. A fast path can create a slow tail.
Consider a synthetic back-office queue: the agent processes routine requests quickly, but one wrong-target action takes hours to investigate and reverse. Average click speed never shows that cost. Count exceptions by severity and recovery time.
Visual: Exceptions consume the owner's attention later.
3. Use an outcome ledger.
Tie each attempt to a request ID, independent acceptance result, human minutes, elapsed time and customer effect. Keep failed runs in the sample. Compare candidate models and harness versions only on the same task slices and rules.
Visual: Track accepted, unknown, rescued and reversed tasks.
4. Scale the accepted envelope.
Let the agent act within a measured scope; expand only when accepted throughput rises and recovery cost stays bounded. A fleet that looks busy but moves exceptions into people has not earned more authority.
Visual: Do this because activity is cheap to fake.
Research and claim limits
- OpenAI Computer use API guide (S159)
- OSWorld 2.1 official repository (S162)
- NIST AI RMF 1.0 (S171)
The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.