JournalDAY 70 / X

FIELD NOTE / X

Two model votes are not ground truth.

The complete written thought and the evidence behind it. The video edition will follow its public release.

Journal September 25, 2026 · X target December 6, 2026

Two model votes are not ground truth.

Video caption

Two model votes are not ground truth. A benchmark can grade agreement with another model. Let the world answer them. #EricFieldNotes

Full written post / accessibility read

If Astra and Fable agree on a business route, that is a useful signal. It is not an independently observed outcome. Both may see the same incomplete customer packet and make the same elegant mistake.

TypeSafe's published workflow evaluation uses strong-model consensus as reference labels in its described setup. That can compare model alignment with those labels. It cannot by itself establish that a transfer, refund or support decision worked for the customer.

Have domain owners label cases from source records and observed outcomes before looking at model votes. Keep disagreement, unknown and costly exception slices. Then compare candidate routes on accepted answers, reviewer time, reversals and harm-weighted wrong branches.

Use consensus to surface hypotheses, not certify a route. Do this because two persuasive systems can agree on the same abstraction error; an independent case packet and outcome readback supply the evidence their agreement lacks.

#EricFieldNotes

Four-beat scene transcript

1. Two model votes are not ground truth.

If Astra and Fable agree on a business route, that is a useful signal. It is not an independently observed outcome. Both may see the same incomplete customer packet and make the same elegant mistake.

Visual: Consensus can share the same missing premise.

2. The reference label determines the score.

TypeSafe's published workflow evaluation uses strong-model consensus as reference labels in its described setup. That can compare model alignment with those labels. It cannot by itself establish that a transfer, refund or support decision worked for the customer.

Visual: A benchmark can grade agreement with another model.

3. Add an outcome adjudicator.

Have domain owners label cases from source records and observed outcomes before looking at model votes. Keep disagreement, unknown and costly exception slices. Then compare candidate routes on accepted answers, reviewer time, reversals and harm-weighted wrong branches.

Visual: Protect a holdout the models did not write.

4. Let model debate suggest tests.

Use consensus to surface hypotheses, not certify a route. Do this because two persuasive systems can agree on the same abstraction error; an independent case packet and outcome readback supply the evidence their agreement lacks.

Visual: Let the world answer them.

Research and claim limits

The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.

More notes from the work ↗