FIELD NOTE / X
Do not compare token counts alone.
The complete written thought and the evidence behind it. The video edition will follow its public release.
The written argument is here.
This approved X edition is on the journal now. Its video player and original platform link will appear after each public release is verified.
Do not compare token counts alone.
Video caption
Do not compare token counts alone. False cutovers and false holds have different costs. Promote only the route that meets the acceptance contract. #EricFieldNotes
Full written post / accessibility read
Put four context designs on the same protected migration cases: ordinary retrieval, long context, stronger retrieval and steward injection. Give them the same cutover service and independently labeled source truth.
Track whether the decisive fact reached dispatch, how often a route wrongly greenlights, how often it unnecessarily blocks, reviewer minutes and recovery work. Do not invent a win rate from one synthetic example.
Include approval-only, open cohort, resolved exception, conflicting source and missing-record cases. Move the relevant fact through a long context and change retrieval phrasing. Score the resulting action and target readback.
Do that because another agent adds update, trigger and escalation costs. If improved retrieval plus a service rule meets the same error and attention budget, keep the simpler design.
#EricFieldNotes
Four-beat scene transcript
1. Do not compare token counts alone.
Put four context designs on the same protected migration cases: ordinary retrieval, long context, stronger retrieval and steward injection. Give them the same cutover service and independently labeled source truth.
Visual: The unit is an accepted decision, not a clever answer.
2. Count both kinds of wrong.
Track whether the decisive fact reached dispatch, how often a route wrongly greenlights, how often it unnecessarily blocks, reviewer minutes and recovery work. Do not invent a win rate from one synthetic example.
Visual: False cutovers and false holds have different costs.
3. Protect the holdout.
Include approval-only, open cohort, resolved exception, conflicting source and missing-record cases. Move the relevant fact through a long context and change retrieval phrasing. Score the resulting action and target readback.
Visual: Remove, move and supersede the decisive record.
4. Buy the least complex winner.
Do that because another agent adds update, trigger and escalation costs. If improved retrieval plus a service rule meets the same error and attention budget, keep the simpler design.
Visual: Promote only the route that meets the acceptance contract.
Research and claim limits
- Liu et al., Lost in the Middle, TACL 2024 (S169)
- LongMemEval, ICLR 2025 (S206)
- Anthropic contextual retrieval technical note (S209)
- Microsoft Research GraphRAG paper (S210)
The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.