JournalDAY 70 / INSTAGRAM

FIELD NOTE / INSTAGRAM

Calibrate by slice, or the average lies.

The complete written thought and the evidence behind it. The video edition will follow its public release.

Journal September 25, 2026 · Instagram target December 6, 2026

Calibrate by slice, or the average lies.

Video caption

Calibrate by slice, or the average lies.

A concentrated Choice may omit the needed action.

Count errors and review cost at each threshold.

Use confidence to route, never to grant permission.

#EricFieldNotes

Full written post / accessibility read

A model's confidence can look useful across the whole evaluation set and still fail on a rare case that matters. Refunds, multilingual requests, missing records and contract exceptions need their own outcome slices. The average does not price their consequences.

Jev's documented Choice confidence summarizes its distribution over listed options. If the correct route is outside that list, a high value is not calibrated correctness. A label-swap preprint also reports branch sensitivity in its studied cases without type errors.

For each protected case slice, compare confidence buckets with independently adjudicated outcomes. At each proposed threshold, measure auto-accepted work, wrong-branch severity, human review load and time to final disposition. Include unknown and out-of-scope cases.

Version the option set and threshold together, keep rollback possible, and let a domain owner approve the error budget. Do this because confidence can help allocate review attention only after the surrounding workflow has been evaluated on real outcomes.

#EricFieldNotes

Four-beat scene transcript

1. Calibrate by slice, or the average lies.

A model's confidence can look useful across the whole evaluation set and still fail on a rare case that matters. Refunds, multilingual requests, missing records and contract exceptions need their own outcome slices. The average does not price their consequences.

Visual: One high overall score can hide a costly branch.

2. Confidence is about supplied options.

Jev's documented Choice confidence summarizes its distribution over listed options. If the correct route is outside that list, a high value is not calibrated correctness. A label-swap preprint also reports branch sensitivity in its studied cases without type errors.

Visual: A concentrated Choice may omit the needed action.

3. Build an abstention curve.

For each protected case slice, compare confidence buckets with independently adjudicated outcomes. At each proposed threshold, measure auto-accepted work, wrong-branch severity, human review load and time to final disposition. Include unknown and out-of-scope cases.

Visual: Count errors and review cost at each threshold.

4. Promote thresholds with an owner.

Version the option set and threshold together, keep rollback possible, and let a domain owner approve the error budget. Do this because confidence can help allocate review attention only after the surrounding workflow has been evaluated on real outcomes.

Visual: Use confidence to route, never to grant permission.

Research and claim limits

The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.

More notes from the work ↗