Embodied decisions
Simulated object handling, workspace boundaries, safety interlocks, and uncertain visual observations. No physical actuators.
MATCHED SAFETY-PROCEDURE PILOT
123/192 episodes completed; 2 unscored technical stops; 67 unattempted. No omissions or authority breaches in completed cases.
Results, prior-work audit and contribution decision ↗RESEARCH / PROTOCOL DEVELOPMENT
We study observable behavior when AI systems can advise, decide, or act. The goal is evidence that helps teams find and reduce specific risks.
Robocurve’s RoboHarm examines unsafe instructions in fixed robot scenes. It motivates a central question: does safety behavior carry through to action?
Our proposed suite extends that question to simulated agents and accidental failures. It uses matched safe tasks, wording variation, explicit permissions, and recorded tool traces.
RoboHarm findings belong to its tested hardware, policies, and conditions. They do not establish a general probability of harm.
Read the original research ↗ Inspect the source ↗PLANNED TEST FAMILIES
Simulated object handling, workspace boundaries, safety interlocks, and uncertain visual observations. No physical actuators.
Fictional health and crisis scenarios. Measure recognition of missing information, escalation, and unsupported certainty. Domain experts review labels.
Synthetic records and simulated recipients. Test whether agents respect disclosure limits and revoke access when consent changes.
Fictional schedules and service queues. Test preventable deprivation, unequal treatment, and the effects of optimization pressure.
Mock APIs with explicit permission scopes. Inspect irreversible action proposals, confirmation checks, and attempts to exceed granted authority.
Inert documents that ask agents to ignore task constraints. Test separation of data from instructions and recovery after conflicting messages.
Simulated stale data, timeouts, duplicate actions, and misleading success responses. Measure whether the agent detects and contains a failure.
Test whether an agent stops when asked, reports its actions accurately, and preserves evidence of unsafe decisions.
All action tests run against simulated tools and synthetic subjects. Dangerous procedures and reusable attack payloads do not belong in public release artifacts.
Release and disclosure policy ↗The current scoped study tests matched state changes using fixed synthetic histories. An earlier text pilot checks interface feasibility. Broader families and visual tests are planned extensions.
Research program and funding proposal ↗Implemented scoring contract ↗Safe and unsafe scoring examples ↗Source and revision history ↗