MURDERBENCH

MECHANISM AUDIT / 4 OCTOBER 2026

Can a security fix
block recovery?

Yes, in a documented executor configuration. We reproduced the dependency in Vault. The current study does not establish a novel AI result.

Contribution decision: stop before model collection

The mechanism and migration strategies have established precedents. This work supplies a conformance example. It does not meet our novelty and sponsorship bar.

No new model calls, API reservations, or model spending. No new social promotion or sponsorship outreach for this candidate.

Read the complete rejection and skeptical assessment ↗

What actually ran

A real Vault 1.20.4 process ran on loopback with synthetic identities and data. A waiting worker held its checkpoint in storage scoped to its token. Four scripted procedures changed the token hierarchy before the worker resumed.

The amended collection completed 48 reference cases. The first version completed ten cases and retained one unscored stop, leaving 37 of its planned cases unattempted. Across both versions: 59 attempts, 58 completed, one unscored stop. All attempts remain available.

These are deterministic reference procedures. Repetitions check software consistency; they are not model trials or independent incident samples.

Dependent worker: two declared incident labels × two repetitions per procedure
Reference procedureRecovery completedBoth old test identities denied at endpoint
Leave the old identities active4/40/4
Revoke the parent and descendants0/44/4
Revoke the parent, preserve descendants4/40/4
Trusted handoff, then revoke the old tree4/44/4

Full revocation blocked the dependent worker. Preserving descendants kept recovery possible, but also preserved the test child's access. A trusted handoff preserved recovery and revoked both old test identities at the endpoint.

With an independent worker identity, full revocation allowed recovery in all four cases. When setup had already removed the sole checkpoint, none of the four procedures recovered it.

Vault documents the parent-child cascade and deletion of per-token storage. Our reproduction confirms those effects; it does not discover them.

Token hierarchy documentation ↗ Per-token storage documentation ↗

The repair has a security cost

The handoff copies the checkpoint before revocation. Old credentials remain live during that interval. We measure intervention duration and probe access after settlement; we do not measure continuous exposure or attacker activity during the handoff.

The tested requirement is denial of both old test identities at the endpoint. Immediate containment is a stricter requirement and could make this handoff inadmissible. The repair is not universally safe.

The single-copy checkpoint, trusted handoff authority, and absence of another recovery source are explicit fixture assumptions. No production operator has validated them here.

What this cannot tell us

The compromise and false-alert labels change the declared requirement, not the executor state. This does not test alert recognition, prompt injection, safety-induced choices, intent, alignment, or human harm.

No frontier model chose these procedures. No operational decision changed through an operator-reviewed trace. A synthetic checkpoint loss alone is insufficient evidence for deployment risk or funding claims.

The closest work already covers much of the idea

AIOpsLab includes executed authentication-revocation tasks. GuardedAct evaluates remediation collateral damage and recoverability. ColdStart describes recovery quality and security checks. Existing rotation guidance covers handoffs during active jobs.

Selected public implementations do not settle every sequencing question. Protected task pools and unavailable original releases prevent a complete comparison. Those access gaps cannot establish originality.

Primary-source audit, exact distinctions, and access limits ↗ Pinned AIOpsLab task ↗ GuardedAct paper ↗ ColdStart methodology ↗ Active-job rotation guidance ↗

One collector correction, fully retained

The first collector expected successful revocation even when the token had already been removed. Vault's parent-only endpoint instead returns a documented HTTP 400. The run stopped and its prefix remains unscored.

The separate v2 collector accepts only that exact response after an observed setup revocation, then settles the worker and probes access. The original sources, registration, and results remain unchanged. The matrix is identical.

Reproduce the audit

The receipt verifier independently folds issued identities, revocations, checkpoint reads and writes, and security probes. It shares the collector's capture; independent capture and external peer review are not claimed.

The source audit records 159 hashed text files from selected pinned repositories. Its inspection scope is explicit. Public records contain identity aliases and checkpoint digests.

Collector, verifier, audit, and tampering tests ↗

What could justify reopening the question

A permissioned operator trace could establish a dependency missed by an existing control. It would need actual recovery alternatives, containment limits, and a tested repair that changes that operator's decision.

That evidence is a prerequisite. No further model collection or sponsorship pitch is activated by this report.

Previous matched pilot ↗