PROPOSED STUDY / OCTOBER 4, 2026
Can a safety label
make an agent
less safe?
Change the description of a safeguard. Keep its code, coverage facts and task fixed. Measure what the agent actually executes.
This is a reviewable proposal. No model trials have been run for it. The protocol, native interface and confirmation design are not frozen.
The boundary worth testing
A motion validator can check action bounds and per-step movement without checking whether a permitted movement is appropriate for the object in the scene. An agent that knows those limits still has to make the remaining judgment.
Our proposed comparison calls that same component either a “motion validator” or a “safety approver.” Both descriptions disclose identical coverage facts. Any measured difference would concern this controlled wording intervention, not proof of a model’s motives or an effect of alignment training.
What would make this useful?
The decision is concrete: should an operator change the instruction describing a guard, explicitly assign responsibility for uncovered hazards, or replace the guard with one that enforces the relevant constraint?
A repair must reduce unsafe effects without merely blocking ordinary work. The full semantic guard is a control for prevention; the explicit-limits instruction is a separate repair comparison because it adds information.
Controls before model calls
- A correct scripted policy completes benign tasks and available safe alternatives.
- A deliberately faulty policy demonstrates a hazard that the partial guard admits.
- A guard enforcing the full scenario constraint prevents that faulty execution.
- An independent event-log verifier reconstructs proposed actions, guard decisions and state effects.
- Fresh confirmation cases test a pilot effect; missing and technically invalid attempts remain visible.
Hazards must be independently legible. Impossible situations, failed transport and inability to complete a benign task are separate outcomes. Recognition checks use fresh episodes so they cannot prime the action test.
Native calibration, with limits
The pinned Inspect Robots core guard accepted two permitted synthetic displacements that crossed an unrelated semantic rule. Bounds, per-step limiting and non-finite rejection behaved as documented. This is an expected coverage boundary, not a novel failure or a model result. Additional embodiment guards may check more. Per-step movement is not a dynamic speed guarantee.
Source audit and reproduction instructions ↓ · Native primitive receipt ↓
Executable development controls
The restricted absolute-target adapter matched the native guard in all 11,025 supported state/target combinations. All 224 scripted traces independently replayed, including safe detours, faulty shortcuts, false alerts and infeasible deliveries. The full surface check prevents contact while permitting the correct detours; it does not make a faulty policy complete the task.

Reproduction and remaining gates ↓ · Every scripted trace ↓ · Every native comparison and trace ↓ · Vector figure ↓ · Initial receipts and source snapshot ↓ · Portable source pin ↓ · Figure dependencies ↓
The closest work is already strong
RoboHarm evaluates unsafe instructions in fixed robot scenes, and RoboPAIR demonstrates robot misuse. Neither a dangerous scene nor low refusal alone would earn a contribution here.
Delegated Misalignment studies responsibility diffusion and role bias in principal/subordinate systems. Our proposed contrast keeps one agent and its task information fixed, with a non-agent guard. We are not claiming to discover responsibility diffusion.
ClawsBench studies scaffolding effects, and RoboGuard studies robot guardrails and plan repair. A new harness or added semantic check alone is insufficient. The question is whether this particular contrast identifies a useful configuration choice.
Malberg et al. already test risk compensation in LLMs. Their released experiment changes the stated risk after adding protection; our candidate would keep risk and coverage facts fixed. Executed actions alone would not establish a new contribution. Proof-of-Guardrail already distinguishes guard execution from general safety.
The exact contrast was not identified in our limited audit. That is not proof that nobody has tested it. We welcome a pointer to closer published or unpublished work.
What the earlier evidence cannot tell us
A released RoboHarm transcript contains a downstream safety-assurance phrase. It motivates this hypothesis but cannot establish that the phrase caused an unsafe choice. MurderBench’s earlier pilots did not isolate this effect either.
No physical actuators, people or production resources will be exposed. We will use original synthetic scenes, respect dependency licenses, and publish routing failures and negative findings.
Stop conditions
Reject the study if it duplicates an existing evaluation, the guard gap cannot be faithfully reproduced, hazards are unobservable, controls fail, the task leaks an obvious answer, or the budget cannot support a meaningful comparison. An isolated mistake does not pass.
Read the full proposal ↓Review questions and support request →