JournalDAY 56 / LINKEDIN

FIELD NOTE / LINKEDIN

Model selection is an experiment, not loyalty.

The complete written thought and the evidence behind it. The video edition will follow its public release.

Journal September 25, 2026 · LinkedIn target November 22, 2026

Model selection is an experiment, not loyalty.

Video caption

Model selection is an experiment, not loyalty.

Quality, latency, cost and rescue.

A selected model still needs release proof.

My rule: Include review and failures in the denominator.

#EricFieldNotes

Full written post / accessibility read

Agentic software delivery mixes repository changes, browser observation, research, coordination and release decisions. A model can be strong in one role and weak in another; the harness and tool permissions also affect results. I would route by accepted task outcomes rather than a single benchmark rank.

Define representative fresh tasks, source-linked acceptance outside the worker's editable scope, identical permission budgets and pinned model/tool versions. Record every attempt, failed or abandoned run, human correction, queue time and final postcondition. Stratify by code, browser, decision and orchestration rather than collapse to one mean.

Once a route wins a task class, each new build still faces protected acceptance on its exact digest and a bounded operational readback. Re-run model selection when behavior, provider policy or tools change. Keep an abstain state when no route meets quality or authority requirements.

My rule: route only where a model has demonstrated fit for that class, then preserve independent product gates. Do this because a cheaper generation call can become an expensive delivered change once retries, human rescue and escaped defects are counted.

#EricFieldNotes

Four-beat scene transcript

1. Model selection is an experiment, not loyalty.

Agentic software delivery mixes repository changes, browser observation, research, coordination and release decisions. A model can be strong in one role and weak in another; the harness and tool permissions also affect results. I would route by accepted task outcomes rather than a single benchmark rank.

Visual: Choose per task class and authority boundary.

2. Compare the entire path.

Define representative fresh tasks, source-linked acceptance outside the worker's editable scope, identical permission budgets and pinned model/tool versions. Record every attempt, failed or abandoned run, human correction, queue time and final postcondition. Stratify by code, browser, decision and orchestration rather than collapse to one mean.

Visual: Quality, latency, cost and rescue.

3. Keep the gate independent of routing.

Once a route wins a task class, each new build still faces protected acceptance on its exact digest and a bounded operational readback. Re-run model selection when behavior, provider policy or tools change. Keep an abstain state when no route meets quality or authority requirements.

Visual: A selected model still needs release proof.

4. Optimize cost per accepted outcome.

My rule: route only where a model has demonstrated fit for that class, then preserve independent product gates. Do this because a cheaper generation call can become an expensive delivered change once retries, human rescue and escaped defects are counted.

Visual: Include review and failures in the denominator.

Research and claim limits

The examples identified as illustrative or simulated are design probes, not reported incidents. Vendor specifications do not establish workload performance.

More notes from the work ↗