MURDERBENCH

MATCHED SAFETY-PROCEDURE PILOT

123/192 episodes completed; 2 unscored technical stops; 67 unattempted. No omissions or authority breaches in completed cases.

Results, prior-work audit and contribution decision ↗

RESEARCH / PROTOCOL DEVELOPMENT

From instruction
to consequence.

We study observable behavior when AI systems can advise, decide, or act. The goal is evidence that helps teams find and reduce specific risks.

What RoboHarm
made visible.

Robocurve’s RoboHarm examines unsafe instructions in fixed robot scenes. It motivates a central question: does safety behavior carry through to action?

Our proposed suite extends that question to simulated agents and accidental failures. It uses matched safe tasks, wording variation, explicit permissions, and recorded tool traces.

RoboHarm findings belong to its tested hardware, policies, and conditions. They do not establish a general probability of harm.

Read the original research ↗ Inspect the source ↗

PLANNED TEST FAMILIES

Embodied decisions

Simulated object handling, workspace boundaries, safety interlocks, and uncertain visual observations. No physical actuators.

High-stakes advice

Fictional health and crisis scenarios. Measure recognition of missing information, escalation, and unsupported certainty. Domain experts review labels.

Privacy and consent

Synthetic records and simulated recipients. Test whether agents respect disclosure limits and revoke access when consent changes.

Resource allocation

Fictional schedules and service queues. Test preventable deprivation, unequal treatment, and the effects of optimization pressure.

Tool authority

Mock APIs with explicit permission scopes. Inspect irreversible action proposals, confirmation checks, and attempts to exceed granted authority.

Untrusted inputs

Inert documents that ask agents to ignore task constraints. Test separation of data from instructions and recovery after conflicting messages.

Errors and recovery

Simulated stale data, timeouts, duplicate actions, and misleading success responses. Measure whether the agent detects and contains a failure.

Oversight and reporting

Test whether an agent stops when asked, reports its actions accurately, and preserves evidence of unsafe decisions.

Research without real-world exposure.

All action tests run against simulated tools and synthetic subjects. Dangerous procedures and reusable attack payloads do not belong in public release artifacts.

Release and disclosure policy ↗

Current study and public artifacts

The current scoped study tests matched state changes using fixed synthetic histories. An earlier text pilot checks interface feasibility. Broader families and visual tests are planned extensions.

Research program and funding proposal ↗Implemented scoring contract ↗Safe and unsafe scoring examples ↗Source and revision history ↗