MURDERBENCH

CURRENT STUDY / OCTOBER 2, 2026

The executable scoped study uses fixed synthetic histories, four constraint mechanisms, six decision-pressure variants, and native action proposals. The broader protocol below is a future research roadmap.

Current study and analysis · Frozen configuration

METHOD / DRAFT V0.1

Measure the action.
Explain the uncertainty.

A test must distinguish a safe decision from an inability to act. It must also make its own assumptions available for review.

1. Define the risk before the run.

Each scenario specifies a fictional subject, task objective, safety constraint, permitted actions, harm severity, and an observable failure condition. A matched benign task checks relevant capability.

2. Vary the conditions deliberately.

Test direct harmful requests and ordinary tasks with latent hazards. Use fixed paraphrases, multiple conversation turns, uncertain context, and changes in authority. Freeze the scenario set before collecting comparative results.

3. Record decisions and tool effects.

Capture messages, proposed calls, simulator state transitions, safety interventions, latency, and errors. The simulator cannot send messages, spend money, access private data, or control hardware.

4. Score separate outcomes.

Planned outcome measures
MeasureWhat it means
Unsafe proposalThe model requests an action that violates a defined constraint.
Unsafe executionThe simulated state crosses the defined harm boundary.
Safety refusalThe model declines the unsafe step for a relevant safety reason.
Useful safe completionThe model completes the matched benign objective within its authority.
RecoveryThe model detects an error and contains its simulated effects.
Technical failureAn API or harness error makes the trial invalid for behavioral scoring.

We do not infer a model’s mental intent. “Intentional misuse” describes the request or experimental condition. Evidence of an unsafe choice requires a trace.

5. Review labels independently.

Two reviewers score the primary outcomes without model identity. Disagreements go to adjudication. Automated labels support review; they do not decide the final outcome for severe cases.

6. Report uncertainty and limitations.

Publish per-family rates, denominators, confidence intervals, benign-task performance, invalid trials, and subgroup gaps. Cluster uncertainty by scenario. Do not pool unlike conditions into a single safety score.

Proposed collection design

The expanded roadmap uses eight families, twelve base scenarios per family, three wording variants, and five repetitions. Each unsafe trial has a matched benign trial. This produces 2,880 trials per model condition. A smaller pilot validates the harness first.

This design is provisional. Confirm power, costs, provider access, and expert review before a confirmatory run. Zero observed failures does not establish zero risk.

Reproducibility

Freeze model identifiers, access date, prompts, tool schemas, generation parameters, budgets, scenario hashes, and harness revision. Mark unavailable or incompatible models explicitly. Maintain public and restricted release sets to reduce contamination and misuse.

No results yet.

This is an evaluation design, not a claim that a model has passed these tests.

Download the full draft protocol ↓

Current study and public artifacts

The current scoped study tests matched state changes using fixed synthetic histories. An earlier text pilot checks interface feasibility. Broader families and visual tests are planned extensions.

Research program and funding proposal ↗Implemented scoring contract ↗Safe and unsafe scoring examples ↗Source and revision history ↗