FIELD NOTE / X
An AI can execute the wrong paraphrase perfectly.
The complete written thought and the evidence behind it. The video edition will follow its public release.
The written argument is here.
This approved X edition is on the journal now. Its video player and original platform link will appear after each public release is verified.
An AI can execute the wrong paraphrase perfectly.
Video caption
An AI can execute the wrong paraphrase perfectly. They do not prove the spec retained the original limit. Do this because code can be correct for the wrong ask. #EricFieldNotes
Full written post / accessibility read
The customer says migrate only after advance notice. A plan summary says migrate the customer. The task agent ships a flawless migration workflow and every implementation test passes. The failure is the missing notice condition between source and ticket.
The coding agent can test every branch it was given. It cannot test a condition omitted before its task began. A second coding agent reviewing the same ticket can agree and remain blind to the original decision.
Give the task a source ID, notice deadline, customer scope and owner. Before execution, compare the proposed change and its tests with those invariants. A mismatch returns to the decision owner, not to the implementer for more code.
Keep implementation quality gates. Add a separate decision-fidelity gate before them. If a task cannot recover the customer's condition, stop. Perfect execution is not valuable when the organization has changed the objective by paraphrase.
#EricFieldNotes
Four-beat scene transcript
1. An AI can execute the wrong paraphrase perfectly.
The customer says migrate only after advance notice. A plan summary says migrate the customer. The task agent ships a flawless migration workflow and every implementation test passes. The failure is the missing notice condition between source and ticket.
Visual: Implementation tests may pass while intent was lost upstream.
2. Green tests validate the compressed spec.
The coding agent can test every branch it was given. It cannot test a condition omitted before its task began. A second coding agent reviewing the same ticket can agree and remain blind to the original decision.
Visual: They do not prove the spec retained the original limit.
3. Read the action back against source.
Give the task a source ID, notice deadline, customer scope and owner. Before execution, compare the proposed change and its tests with those invariants. A mismatch returns to the decision owner, not to the implementer for more code.
Visual: Extract invariants before task creation.
4. Verify intent before correctness.
Keep implementation quality gates. Add a separate decision-fidelity gate before them. If a task cannot recover the customer's condition, stop. Perfect execution is not valuable when the organization has changed the objective by paraphrase.
Visual: Do this because code can be correct for the wrong ask.
Research and claim limits
- Mohamed et al., LLM as a Broken Telephone, ACL 2025 (S167)
- Perez et al., When LLMs Play the Telephone Game, ICLR 2025 (S168)
- NIST AI RMF 1.0 (S171)
The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.