FIELD NOTE / TIKTOK
You ran ten agent experiments. Now what?
The complete written thought and the evidence behind it. The video edition will follow its public release.
The written argument is here.
This approved TikTok edition is on the journal now. Its video player and original platform link will appear after each public release is verified.
You ran ten agent experiments. Now what?
Video caption
You ran ten agent experiments. Now what? A new seed must challenge a different assumption. Do this because the test list is not the goal. #EricFieldNotes
Full written post / accessibility read
Imagine ten browser runs on the same happy-path account. They all pass. Then a role change arrives and the workflow fails on permissions. The number ten was real; the coverage claim was not.
If every run reused the same browser state, same tenant and same response timing, all ten experiments tested one narrow path. The next useful case changes one variable: a delayed write, expired session or denied role.
In a local fixture, force Save to render green and policy to reject afterward. If your suite still calls success, you have found a stronger bug than another flaky click: the acceptance oracle is coupled to the UI.
Keep asking for the smallest safe experiment that could reverse your release decision. When none is worth running, document the remaining risk and owner. If you cannot name a falsifier, you are probably demonstrating rather than testing.
#EricFieldNotes
Four-beat scene transcript
1. You ran ten agent experiments. Now what?
Imagine ten browser runs on the same happy-path account. They all pass. Then a role change arrives and the workflow fails on permissions. The number ten was real; the coverage claim was not.
Visual: Ten passes can still share one blind spot.
2. Repetition is not discrimination.
If every run reused the same browser state, same tenant and same response timing, all ten experiments tested one narrow path. The next useful case changes one variable: a delayed write, expired session or denied role.
Visual: A new seed must challenge a different assumption.
3. Run one decisive negative control.
In a local fixture, force Save to render green and policy to reject afterward. If your suite still calls success, you have found a stronger bug than another flaky click: the acceptance oracle is coupled to the UI.
Visual: Make the optimistic green screen persist while state fails.
4. Ask what result would change your mind.
Keep asking for the smallest safe experiment that could reverse your release decision. When none is worth running, document the remaining risk and owner. If you cannot name a falsifier, you are probably demonstrating rather than testing.
Visual: Do this because the test list is not the goal.
Research and claim limits
- OpenAI Computer use API guide (S159)
- OSWorld 2.1 official repository (S162)
The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.