MURDERBENCH RESEARCH PROPOSAL v0.1 Status: preliminary proposal. No external funding or partnerships asserted. QUESTION Do capable agents update their action proposals when verified safety conditions change after a task has begun? DECISION VALUE Help labs test whether a mitigation preserves useful task performance while improving restraint after permission or safety evidence changes. Help public-sector evaluators inspect evidence for a specific deployment boundary. Do not equate this benchmark with legal compliance or general model safety. DELIVERABLE A controlled intervention suite, frozen protocol, reproducible harness, independent labels, uncertainty analysis, and one independently replicated slice. Publish safe examples and withhold dangerous operational payloads. PROVISIONAL FUNDING ASK: USD 40,000 FOR A 12-WEEK PILOT PROGRAM Research engineering: USD 18,000, estimate 180 hours at USD 100/hour. Domain and annotation review: USD 8,000, estimate 80 hours at USD 100/hour. Model access and collection: USD 4,000, capped usage budget. Independent replication: USD 6,000, external scoped contract estimate. Contingency and secure evidence handling: USD 4,000. These are planning estimates, not quotes or commitments. Funding remains optional. Linh Ngo leads engineering. Independent reviewers and replicators are not secured. WEEKS 1-3 / PRIOR ART AND CONSTRUCT VALIDITY Compare scenario-level overlap with closest public benchmarks. Review 24 matched synthetic scenarios with an external method reviewer. Go: distinct useful measurement slice with valid controls. No-go: contribution duplicates prior work or controls cannot isolate the construct. If no-go, publish the overlap assessment and redirect to replication. WEEKS 4-6 / HARNESS AND CALIBRATION Implement mock action transitions and visual counterfactuals. Freeze schemas, limits, trace provenance, and analysis specification. Go: isolation and budget checks pass; reviewers agree on anchor labels. No-go: tool behavior, malformed outputs, or scoring defects confound outcomes. WEEKS 7-9 / COLLECTION AND DISCLOSURE Proprietary frontier priority: OpenAI, Anthropic, Google, xAI. Then Chinese open-weight priority: DeepSeek, Qwen, Moonshot Kimi, Z.ai GLM. Resolve exact versions and open-weight licenses through official sources. Hosted API behavior must remain separate from direct-checkpoint experiments. Go: costs fit caps and model configurations remain comparable. No-go: version changes, missing access, or budget cannot support required power. WEEKS 10-12 / REVIEW AND REPLICATION Complete blinded labels, adjudication, clustered analysis, and provider notices. An independent team replicates a frozen subset before the comparative report. Release all evidence needed to assess the claim and its limits. SUCCESS CRITERIA Reviewed contribution claim, preregistration, reproducible execution, complete denominators, independent labels, replication, and actionable mitigation comparison. A sensational result is not a success criterion. REQUESTS Method reviewer: review controls and estimands; return a written issue list. Replication partner: rerun one frozen slice; return environment and trace hashes. Funder: support the scoped milestone package with scoring independence preserved. MILESTONE ACCEPTANCE ARTIFACTS / PROPOSED, ROLES UNSECURED Week3: pinned prior-case overlap matrix, construct map, exclusion decisions, and a written method review. Independent acceptance role: agent-evaluation researcher unaffiliated with the implementation. Name and quote pending. Week6: frozen source/catalog/eligibility/analysis, offline failure tests, complete case oracle, safe reproduction instructions. Acceptance role: independent research engineer; name and commitment pending. Week9: complete coverage and cost report, blinded label disagreement record, and a specified mitigation comparison. Candidate mitigation: external state-gate that refuses commit if the current authorization record is invalid; compare ungated proposals against gated simulated commits while preserving valid-state completion. This tests a wrapper policy, not improved model internal intent. Week12: independently reproduced subset with documented configuration differences, negative findings, corrected report, and release decision. Acceptance role: replication partner; no commitment currently secured. POWER AND QUOTES The separately proposed rollout-design.txt now specifies event-order oracles, eight boundary fixtures, a crossed gate/evidence/queue design, co-primary safety and utility decisions, and sample-size sensitivity. It is a review draft, not a preregistration, implemented simulator or accepted power calculation. Its illustrative 236-context, eight-route, 24-cell configuration implies 45,312 continuations before preparation and repeats. Obtain quotes and independently review the clustered joint-endpoint design before treating the model-access line as sufficient or submitting a binding budget. The USD50 initial slice is descriptive, not powered for rare events or population comparisons. Before funded comparative collection, an independent method reviewer must define a minimum useful mitigation effect, dependency unit, baseline variance, missingness assumptions, and required sample size. If no defensible design fits the funding ceiling, revise scope before accepting a funded milestone. Obtain actual engineering/review/replication quotes and model-price snapshots before submitting a binding budget. Current figures are planning estimates. The current fixed-history slice cannot substantiate behavior during natural agent rollouts; genuine rollout validation is a separately reviewed extension. STATE-GATE BASELINE LIMITS The perfect-current-state gate enforces its rule by construction. It is an implementation baseline, not a discovery of model improvement. A funded mitigation study must independently specify stale-state reads, delayed revocation, incorrect identity resolution, fail-open behavior and false blocking costs. Do not execute real actions; compare simulated outcomes and availability only. These boundary fixtures and acceptance thresholds require external review. DATED PUBLIC FUNDING ROUTES / OCTOBER 2, 2026 Anthropic third-party evaluation initiative (published July 1, 2024): https://www.anthropic.com/news/a-new-initiative-for-developing-third-party-model-evaluations Public proposal form and eval-initiative@anthropic.com are published routes. The article emphasizes difficult evaluations, expert baselines, domain expertise and reproducibility. Current intake availability is not confirmed by this historical announcement; request program-fit guidance before a funding application. No contact, proposal submission or commitment is claimed. AISI Alignment Project: https://alignmentproject.aisi.gov.uk/ Official page currently says applications are closed and a similar 2026 program is unlikely. Do not submit against a stale call or describe it as an open opportunity. Check the official grants page when preparing a future application: https://www.aisi.gov.uk/grants. Current AISI social outreach asks for methods criticism, not an assumed grant relationship. The official AISI grants overview was also checked on October 2, 2026 and says its programmes are not currently accepting applications. Eligibility and future calls must be checked before proposing a submission.