MURDERBENCH

CURRENT COLLECTION

The first scoped proxy collection has recorded 2,304 of 2,304 candidate conditions. Independent labeling and replication remain pending. The research program below includes proposed extensions.

Observed proposals, coverage and limitations

COVERAGE / NO RESULTS PUBLISHED

A model roster
that can be audited.

Frontier changes quickly. We define inclusion rules and record exact access conditions before making comparisons.

Inclusion policy

Include the strongest generally available model from each eligible provider, plus a smaller model when it helps explain the effect of capability. Include leading open-weight models when the required compute is available.

First priority: proprietary frontier models from OpenAI, Anthropic, Google DeepMind, and xAI. Second priority: Chinese open-weight families from DeepSeek, Qwen, Moonshot AI, and Z.ai. Other families are outside the initial collection priority.

Exact versions before testing

Resolve current identifiers through official provider catalogs at collection time. Record checkpoint revisions for open-weight models. Never treat a changing alias as a stable version.

Coverage register
TrackScopeStatus
LanguageAdvice, uncertainty, privacy, refusalProtocol development
AgentsSimulated tools, authorization, recoveryProtocol development
MultimodalScene hazards, ambiguous visual inputsProtocol development
Embodied policiesSimulator-compatible action modelsAccess and compatibility review

Comparability rules

Report text, vision, and action tracks separately. Compare models only under shared tool interfaces and resource budgets. Document native-interface runs separately. Cost and latency accompany behavioral results.

Known limits

Simulations cannot establish real-world injury rates. A provider’s refusal filter, access tier, tool adapter, or model update can change results. Unsupported modalities are not scored as failures.

Suggest a model or replication ↗

Current study and public artifacts

The current scoped study tests matched state changes using fixed synthetic histories. An earlier text pilot checks interface feasibility. Broader families and visual tests are planned extensions.

Research program and funding proposal ↗Implemented scoring contract ↗Safe and unsafe scoring examples ↗Source and revision history ↗

Collection candidate register

Proprietary first: GPT-6 Astra, Claude Opus 5.5, Gemini 3.1 Pro Preview, Grok 4.7. Then Chinese open-weight candidates: DeepSeek V4 Pro 0813, Qwen3.8 2.4T A95B, Kimi K3, GLM-5.3.

This roster comes from the OpenRouter catalog snapshot on October 2, 2026. Availability and license verification remain separate checks. Hosted API observations do not establish direct-checkpoint behavior.

Current requested model IDs and routes ↗Dated native-interface eligibility ↗Catalog documentation ↗

EXECUTABLE STUDY FREEZE

Configuration and source are preserved together. This preparation freeze is not a completed collection.

Exact source, catalog, and eligibility archive ↗