FIELD NOTE / LINKEDIN
Tests run is a weak governance metric.
The short film, the complete written thought, and the evidence behind it.
The LinkedIn conversation link will follow its public release.
Tests run is a weak governance metric.
Video caption
Tests run is a weak governance metric.
A missed forbidden effect becomes customer recovery work.
Require negative control, positive control and target readback.
My rule: Report accepted outcomes and reviewer cost, not green volume.
Narration uses Eric's authorized AI voice clone.
#EricFieldNotes
Full written post / accessibility read
A release report says the agents ran nine hundred tests. That is throughput, not assurance. For the protected tenant boundary, I want to know whether a staging principal was denied, whether production had zero new receipts and whether a deliberately bad mutation would have failed the gate.
The value of an automated gate is the bad release it prevents with tolerable reviewer effort. Measure false green and false red cases in a protected fixture, then include reviewer minutes and accepted releases. Do not invent a numeric failure rate from a single demonstration.
For each consequential invariant, show one allowed request that produces a receipt, one forbidden request with no production receipt, and a mutant that makes the gate fail. Include source version, target version and exact test run. This is more decision-ready than a screenshot of a large pass count.
Keep the smallest test set that detects the real forbidden effect and proves the sensor works. Re-run it when the route or service changes. That operating rule makes agentic QA useful to a business owner: it ties automation to customer state rather than to activity metrics.
Narration uses Eric's authorized AI voice clone.
#EricFieldNotes
Four-beat scene transcript
1. Tests run is a weak governance metric.
A release report says the agents ran nine hundred tests. That is throughput, not assurance. For the protected tenant boundary, I want to know whether a staging principal was denied, whether production had zero new receipts and whether a deliberately bad mutation would have failed the gate.
Visual: A thousand green checks can miss one customer effect.
2. False green costs more than test time.
The value of an automated gate is the bad release it prevents with tolerable reviewer effort. Measure false green and false red cases in a protected fixture, then include reviewer minutes and accepted releases. Do not invent a numeric failure rate from a single demonstration.
Visual: A missed forbidden effect becomes customer recovery work.
3. Attach evidence to the decision.
For each consequential invariant, show one allowed request that produces a receipt, one forbidden request with no production receipt, and a mutant that makes the gate fail. Include source version, target version and exact test run. This is more decision-ready than a screenshot of a large pass count.
Visual: Require negative control, positive control and target readback.
4. Buy confidence with discrimination.
Keep the smallest test set that detects the real forbidden effect and proves the sensor works. Re-run it when the route or service changes. That operating rule makes agentic QA useful to a business owner: it ties automation to customer state rather than to activity metrics.
Visual: Report accepted outcomes and reviewer cost, not green volume.
Research and claim limits
The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.