FIELD NOTE / X
Two model votes are not ground truth.
The complete written thought and the evidence behind it. The video edition will follow its public release.
The written argument is here.
This approved X edition is on the journal now. Its video player and original platform link will appear after each public release is verified.
Two model votes are not ground truth.
Video caption
Two model votes are not ground truth. A benchmark can grade agreement with another model. Let the world answer them. #EricFieldNotes
Full written post / accessibility read
If Astra and Fable agree on a business route, that is a useful signal. It is not an independently observed outcome. Both may see the same incomplete customer packet and make the same elegant mistake.
TypeSafe's published workflow evaluation uses strong-model consensus as reference labels in its described setup. That can compare model alignment with those labels. It cannot by itself establish that a transfer, refund or support decision worked for the customer.
Have domain owners label cases from source records and observed outcomes before looking at model votes. Keep disagreement, unknown and costly exception slices. Then compare candidate routes on accepted answers, reviewer time, reversals and harm-weighted wrong branches.
Use consensus to surface hypotheses, not certify a route. Do this because two persuasive systems can agree on the same abstraction error; an independent case packet and outcome readback supply the evidence their agreement lacks.
#EricFieldNotes
Four-beat scene transcript
1. Two model votes are not ground truth.
If Astra and Fable agree on a business route, that is a useful signal. It is not an independently observed outcome. Both may see the same incomplete customer packet and make the same elegant mistake.
Visual: Consensus can share the same missing premise.
2. The reference label determines the score.
TypeSafe's published workflow evaluation uses strong-model consensus as reference labels in its described setup. That can compare model alignment with those labels. It cannot by itself establish that a transfer, refund or support decision worked for the customer.
Visual: A benchmark can grade agreement with another model.
3. Add an outcome adjudicator.
Have domain owners label cases from source records and observed outcomes before looking at model votes. Keep disagreement, unknown and costly exception slices. Then compare candidate routes on accepted answers, reviewer time, reversals and harm-weighted wrong branches.
Visual: Protect a holdout the models did not write.
4. Let model debate suggest tests.
Use consensus to surface hypotheses, not certify a route. Do this because two persuasive systems can agree on the same abstraction error; an independent case packet and outcome readback supply the evidence their agreement lacks.
Visual: Let the world answer them.
Research and claim limits
- TypeSafe AI: Workflow evals (S101)
- Li et al., JEV-as-a-Judge (S103)
- TypeSafe AI: Confidence (S100)
The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.