JournalDAY 101 / TIKTOK

FIELD NOTE / TIKTOK

Which assertion catches the bad mutation?

The short film, the complete written thought, and the evidence behind it.

Journal September 28, 2026 · TikTok target January 6, 2027
Open the approved MP4 ↗

The TikTok conversation link will follow its public release.

Which assertion catches the bad mutation?

Video caption

Which assertion catches the bad mutation? The assertion watched the story, not the effect. Every protected test needs a harmless negative control. Narration uses Eric's authorized AI voice clone. #EricFieldNotes

Full written post / accessibility read

The agent calls the tool. The tool says blocked. The test asserts the word blocked appears, and turns green. Now I make the disposable target issue a forbidden production receipt while leaving that tool text alone. Predict which assertion fails.

This is the vacuous-test trap. A clever test name and a detailed agent explanation do not change what the assertion measures. The protected business fact is whether a staging identity caused any production mutation. The original test never checked that.

Store production receipt count and tenant state hash before the call. Run the forbidden request. Require a deny decision, no receipt delta and no state change. Then trigger an allowed staging mutation to prove the observer sees effects. Turn the bad mutant back on: the test must now fail.

Before announcing that your agent harness enforces a rule, deliberately violate it in a disposable fixture. If the gate stays green, you have a sensor problem. Fix that first, because asking a stronger model to reason over an undiscriminating test will not create evidence.

Narration uses Eric's authorized AI voice clone.

#EricFieldNotes

Four-beat scene transcript

1. Which assertion catches the bad mutation?

The agent calls the tool. The tool says blocked. The test asserts the word blocked appears, and turns green. Now I make the disposable target issue a forbidden production receipt while leaving that tool text alone. Predict which assertion fails.

Visual: Pause before trusting the test's green result.

2. The answer is none.

This is the vacuous-test trap. A clever test name and a detailed agent explanation do not change what the assertion measures. The protected business fact is whether a staging identity caused any production mutation. The original test never checked that.

Visual: The assertion watched the story, not the effect.

3. Repair the acceptance test.

Store production receipt count and tenant state hash before the call. Run the forbidden request. Require a deny decision, no receipt delta and no state change. Then trigger an allowed staging mutation to prove the observer sees effects. Turn the bad mutant back on: the test must now fail.

Visual: Count forbidden receipts and compare target state.

4. Make a wrong outcome turn red.

Before announcing that your agent harness enforces a rule, deliberately violate it in a disposable fixture. If the gate stays green, you have a sensor problem. Fix that first, because asking a stronger model to reason over an undiscriminating test will not create evidence.

Visual: Every protected test needs a harmless negative control.

Research and claim limits

The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.

More notes from the work ↗