Your hook ran. Did the forbidden action stay blocked?
Imagine a fictional deployment where an agent tries to change a production tenant. The pre-call hook says blocked. A customer nevertheless sees the deployment move. Which record would you trust?
I would trust neither record alone. The hook log tells me what happened at one host event. The customer report tells me where to investigate. I want the service authorization decision, the operation receipt and a readback of the current tenant. If those disagree, the job is not complete and the team needs an owner for the uncertain effect.
This is a fictional teaching case, with a disposable local target and ten synthetic probes available in its companion. I am not reporting that Cursor, Claude Code, Grok Build or Codex failed a real production test. Their current hook surfaces are a separate, dated comparison. The principle here survives a product rename: a control is only as good as the consequential routes it observes and the target effects it prevents.
Start with the invariant, not the hook name
The business rule is precise: a staging identity may not mutate a production tenant. A permission denied by an agent host is useful early feedback, but the final rule belongs where the mutation is committed. OWASP's authorization guidance recommends checking permission on every request, regardless of which client initiated it. Its transaction-authorization guidance also puts consequential enforcement server-side. Neither document certifies an individual agent harness. They define the application boundary I would test.
The route table is where vague confidence becomes a test plan:
| Proposed route | Host callback observed? | Principal at target | Service decision | Independent effect proof |
|---|---|---|---|---|
| Direct agent tool | Must probe installed host/version | Scoped staging identity | Deny production | Zero new production receipts; stable target state |
| Shell or script | Separate matcher and process probe | Same or different credential | Deny production | Same receipt and state check |
| MCP connector | Probe exact connector event | Connector credential | Deny production | Same receipt and state check |
| Delegated worker | Probe worker environment and handoff | Worker credential | Deny production | Same receipt and state check |
| Direct service API | May bypass agent host entirely | Service credential | Deny production | Same receipt and state check |
Every must probe cell is unknown until a harmless call is made in the installed version. A feature list is not a route test. A deny log is not a target readback. And a zero-receipt result needs a positive control: an authorized staging change must create a receipt, proving the observer is alive.
The repo fixture simulates a permissive host error route and an independent service. It does not run the four named coding products. Its results demonstrate the mechanism, not vendor coverage.
The hook error path is part of the policy
A pre-call hook can inspect a matched invocation and deny it. What happens if the executable path is wrong in a cloud worker, the callback payload changes, the process exits with an error, or the deadline passes? That answer depends on the host, event, version and sometimes output schema. It has to be measured. A script that normally denies is not an enforcement system until these error branches are covered.
An executable policy module should validate required fields and refuse to infer permission from missing data. The companion's policy.py takes a fixture call with a principal and tenant and returns an explicit allow or deny. Its claude_pretool_fixture_example.py shows how a documented PreToolUse denial shape could wrap that decision for one narrow MCP fixture tool. This is an example adapter, tested with sample JSON; it has not been installed as a real Claude Code hook. A real installation needs the host's current schema, matcher, launch path and route probe. A shell wrapper that exits 2 on denial is also provided as an illustration of command-hook mechanics. Claude Code and Cursor document relevant events and failure semantics, but current docs do not prove a particular local worker has the hook installed.
Even a correctly configured hook is an early control. The service still checks the authenticated principal against the requested tenant. The local fixture deliberately simulates hook crash, malformed response, timeout and missing worker file. In each case the simulated call can reach the service, which denies the forbidden mutation. The companion records both decisions and the target result. This tells us the fixture's backstop worked under those injections. It does not establish any named vendor's fail-open rate.
Pre-call, post-call and target authority do different jobs
I use pre-call checks for eligibility and prompt feedback. I use post-call callbacks for audit, correlation and evidence packets. I put irreversible authorization at the target service before commit. Then I read the target after the call. A post-call callback cannot prevent a side effect that already happened.
The local lost-acknowledgement probe illustrates why that separation matters. The service commits an authorized staging change, then drops its reply. The client sees a connection error. A persisted operation ID lets it query the service journal and confirm that the change committed. Reissuing the action from the failed tool message would be an unproven decision. For a real service, retries and deduplication need an explicit contract and target readback. AWS Step Functions redrive and Azure's compensating-transaction pattern are useful workflow references, but neither creates exactly-once effects in an arbitrary service.
I would make effect unknown a first-class state in the agent harness. It blocks retries, gathers operation ID and receipts, and names the operator who can resolve the next step. Denied, unknown, and confirmed partial are different incident queues. Calling all three failed discards the information a person needs to protect the customer.
Worktrees do not isolate the outside world
Separate Git worktrees are a good default for parallel coding agents. They stop file edits from colliding. They do not separate a shared database, browser profile, cloud account or staging tenant. Two agents can have clean branches and conflicting external effects. A test can genuinely pass at target version twelve and be irrelevant when another worker has changed the target to version thirteen.
I want each agent's WAL to carry the task ID, source SHA, expected target ID, principal, fixture version, acceptance result and unresolved consequence. Give workers separate tenants when possible. If a resource must be shared, assign a lease and a version check at acceptance. If the version moved between test and verdict, invalidate the verdict and rerun it. This is not paperwork for its own sake; it avoids sending a green report whose evidence refers to a state that no longer exists.
The throughput measure is accepted changes and recovery time, not simultaneous model streams. More agents can shift the bottleneck to scarce staging environments and operator review. Count that queue rather than treating every parallel token as saved engineering time.
Context injection is useful; context is not authority
A policy-steward agent can be valuable when a specialist is about to act. It can notice a consequential operation and send the relevant exception with source, version, effective time and expiry. That is often better than flooding every worker with a large memory file. Retrieval still has a role, but a query that was never issued cannot retrieve the needed exception. Conversely, an old supervisor summary can be confidently stale.
The handoff should be a bounded evidence packet, not a permit: proposed decision, source, contradictory evidence, unresolved question, owner and deadline. The worker can use it to form a better proposal. The target service checks current policy and the caller's scoped credential before the effect. OpenAI Agents SDK handoff/context documentation describes mechanisms for passing or filtering context; it does not certify the freshness or authority of a particular business claim. OAuth token exchange offers delegation semantics where implemented, but a plain agent message is not such a token.
I would test three context cases: current evidence, no steward message, and a message made stale by a later revocation. Silence and conflict should create a decision request, not an inferred permission. This keeps intelligence in the context layer without letting it invent authority.
Make the forbidden effect happen safely
The strongest test in the companion is intentionally bad. A fixture mutation lets a staging token change production. A vacuous assertion that checks only whether a tool message contains blocked could stay green. The independent oracle compares production receipts and a hash of the target state. It turns red. A second injection hides the receipt but leaves the state changed; the state readback still catches it. The allowed staging positive control proves the receipt observer works.
That is the standard I want from agentic QA: every protected invariant needs a harmless negative control that demonstrates the acceptance gate can fail for the outcome it claims to prevent. More tests are not necessarily stronger tests. Ask what bad state the gate would reject, and run that state in a disposable environment. If the gate stays green, fix the sensor before asking a more capable agent to produce more reassuring prose.
Ten local probes, and what they do not prove
The companion runs an in-process loopback HTTP server with separate staging and production tenants, a receipt journal, a service-side tenant check and a simulated host callback. The probe runner captures call ID, simulated hook decision, service result, receipts and target hash. The current local run passed the first nine protected scenarios; the deliberately broken tenth scenario produced an expected FAIL. Fifteen unit tests passed, including a hidden-receipt mutation. These are synthetic fixture results only.
| Probe | Question | Expected protected result |
|---|---|---|
| 1. Authorized staging | Does the observer see a real permitted effect? | One staging receipt; changed staging state |
| 2. Direct forbidden | Does the direct policy deny? | No production effect |
| 3. Hook crash | Does target authorization survive a simulated callback error? | Service denies production |
| 4. Malformed output | Does missing valid hook output become permission? | Service denies production |
| 5. Timeout | Does a simulated hook timeout leak authority? | Service denies production |
| 6. Missing worker hook | Does a second environment still have a backstop? | Service denies production |
| 7. Delegated route | Does a child path reach a separate target rule? | Service denies production |
| 8. Alternate tool | Does fallback avoid the service check? | Service denies production |
| 9. Lost acknowledgement | Can the operation be reconciled without a blind retry? | One committed staging receipt |
| 10. Bad-effect mutant | Will the acceptance gate catch forbidden production change? | The oracle fails as intended |
An actual product comparison needs an installed-version matrix, harmless route calls and exact target readbacks on each product. A local simulated host is not that evidence. The next experiments I would run are versions of the ten above across the direct, delegated and alternate routes available in the team's development VMs, with no production credentials and no disruptive changes. Repeat when a host or connector is upgraded. If any route cannot be exercised, mark it not observed and scope its authority accordingly.
The cost ledger
There is an operating cost to strong controls: maintaining adapters, staging targets, test fixtures, credentials and receipt readers. There is also a cost to weak ones: uncertain customer changes, extra re-prompting, retesting, redeploying and operator time. I would track five columns for each protected workflow: consequence, route coverage, test/adapter upkeep, false-green or false-red evidence, and accepted business outcome. A draft document can use lighter checks. A cross-tenant production change deserves a service gate and receipt readback.
My rule is simple: write the invariant once, enforce permission where the effect happens, probe every agent route that can reach it, and read the target before declaring success. Hooks help a team make better decisions earlier. Target authority and independent effect evidence tell us whether the customer stayed protected.
Sources and publication limits
September 28 documentation refresh: the dated source audit in the local review package captures 18 current official sources and separates their documented behavior from unperformed installed-host probes. For example, Cursor's hook documentation includes an explicit failClosed option; Claude Code's hook documentation distinguishes events and handler types; Grok's hook reference describes its own blocking and error contract; and Codex's hook reference identifies supported execution paths and exceptions. These are reasons to test an installed adapter's exact contract, rather than infer equivalent protection from similarly named events. The archival capture expires September 30 for a 48-hour freshness rule; January publication needs a new check. The companion's ten local fixture results do not test those products.
- OWASP Authorization Cheat Sheet and Transaction Authorization Cheat Sheet support request-level, server-side permission checks.
- Claude Code hooks, Cursor hooks, Grok Build permissions and Codex plugin architecture are a September 25, 2026 documentation checkpoint. They are not runtime measurements of local host versions. Recheck official sources within 48 hours before any named claim goes live.
- GitHub Actions OIDC illustrates short-lived scoped cloud credentials, without validating this particular fixture's identity design.
- OpenAI Agents SDK handoffs, RFC 8693, AWS Step Functions redrive and Azure compensating transactions bound the context, delegation and recovery discussion. None proves a production system is exactly-once or correctly authorized.
The field kit
The companion folder for this week holds runnable examples, decision records, evidence and limits. The repository release log is updated after each article's live readback.