FIELD NOTE / INSTAGRAM
Calibrate by slice, or the average lies.
The complete written thought and the evidence behind it. The video edition will follow its public release.
The written argument is here.
This approved Instagram edition is on the journal now. Its video player and original platform link will appear after each public release is verified.
Calibrate by slice, or the average lies.
Video caption
Calibrate by slice, or the average lies.
A concentrated Choice may omit the needed action.
Count errors and review cost at each threshold.
Use confidence to route, never to grant permission.
#EricFieldNotes
Full written post / accessibility read
A model's confidence can look useful across the whole evaluation set and still fail on a rare case that matters. Refunds, multilingual requests, missing records and contract exceptions need their own outcome slices. The average does not price their consequences.
Jev's documented Choice confidence summarizes its distribution over listed options. If the correct route is outside that list, a high value is not calibrated correctness. A label-swap preprint also reports branch sensitivity in its studied cases without type errors.
For each protected case slice, compare confidence buckets with independently adjudicated outcomes. At each proposed threshold, measure auto-accepted work, wrong-branch severity, human review load and time to final disposition. Include unknown and out-of-scope cases.
Version the option set and threshold together, keep rollback possible, and let a domain owner approve the error budget. Do this because confidence can help allocate review attention only after the surrounding workflow has been evaluated on real outcomes.
#EricFieldNotes
Four-beat scene transcript
1. Calibrate by slice, or the average lies.
A model's confidence can look useful across the whole evaluation set and still fail on a rare case that matters. Refunds, multilingual requests, missing records and contract exceptions need their own outcome slices. The average does not price their consequences.
Visual: One high overall score can hide a costly branch.
2. Confidence is about supplied options.
Jev's documented Choice confidence summarizes its distribution over listed options. If the correct route is outside that list, a high value is not calibrated correctness. A label-swap preprint also reports branch sensitivity in its studied cases without type errors.
Visual: A concentrated Choice may omit the needed action.
3. Build an abstention curve.
For each protected case slice, compare confidence buckets with independently adjudicated outcomes. At each proposed threshold, measure auto-accepted work, wrong-branch severity, human review load and time to final disposition. Include unknown and out-of-scope cases.
Visual: Count errors and review cost at each threshold.
4. Promote thresholds with an owner.
Version the option set and threshold together, keep rollback possible, and let a domain owner approve the error budget. Do this because confidence can help allocate review attention only after the surrounding workflow has been evaluated on real outcomes.
Visual: Use confidence to route, never to grant permission.
Research and claim limits
- TypeSafe AI: Confidence (S100)
- TypeSafe AI: Jev 1.13 jaggedness (S99)
- Sun et al., Type-Safe Is Not Error-Free (S102)
The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.