JournalDAY 56 / INSTAGRAM

FIELD NOTE / INSTAGRAM

Route models by task evidence.

The complete written thought and the evidence behind it. The video edition will follow its public release.

Journal September 25, 2026 · Instagram target November 22, 2026

Route models by task evidence.

Video caption

Route models by task evidence.

Quality, privacy, time and rescue per class.

Otherwise routing comparisons are theater.

A route needs abstention and re-test dates.

#EricFieldNotes

Full written post / accessibility read

A coding benchmark, a browser-use benchmark and a real product task measure different things. A model that produces a strong patch may be slow or unreliable at reading a changing UI. A faster model may handle bounded routine work but need escalation on unfamiliar architecture.

For each class, pin the model and tool version, allowed data, acceptance oracle, p ninety-five completion time, cost and human correction. Test cold and warm paths, failure recovery and unseen examples. Record which classes remain unverified instead of assigning them a convenient score.

Run representative repository and browser tasks through each eligible route. Preserve source requirements outside the workers and compare final business postconditions. If one route's tests are written by the same agent while another faces protected acceptance, the comparison is invalid.

Use the cheapest eligible model only where it has passed the full task contract. Escalate novel, risky or stale cases to a stronger verified path or a human. Do this because routing should allocate work by measured fit, not spread one benchmark winner across every surface.

#EricFieldNotes

Four-beat scene transcript

1. Route models by task evidence.

A coding benchmark, a browser-use benchmark and a real product task measure different things. A model that produces a strong patch may be slow or unreliable at reading a changing UI. A faster model may handle bounded routine work but need escalation on unfamiliar architecture.

Visual: Do not send every job to one favorite.

2. Build a versioned route table.

For each class, pin the model and tool version, allowed data, acceptance oracle, p ninety-five completion time, cost and human correction. Test cold and warm paths, failure recovery and unseen examples. Record which classes remain unverified instead of assigning them a convenient score.

Visual: Quality, privacy, time and rescue per class.

3. Use the same external oracle.

Run representative repository and browser tasks through each eligible route. Preserve source requirements outside the workers and compare final business postconditions. If one route's tests are written by the same agent while another faces protected acceptance, the comparison is invalid.

Visual: Otherwise routing comparisons are theater.

4. Escalate the unknown.

Use the cheapest eligible model only where it has passed the full task contract. Escalate novel, risky or stale cases to a stronger verified path or a human. Do this because routing should allocate work by measured fit, not spread one benchmark winner across every surface.

Visual: A route needs abstention and re-test dates.

Research and claim limits

The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.

More notes from the work ↗