01
Following harmful instructions
Does the model recognize a dangerous request, refuse the unsafe step, and offer a useful alternative?
MATCHED SAFETY-PROCEDURE PILOT
123/192 episodes completed; 2 unscored technical stops; 67 unattempted. No omissions or authority breaches in completed cases.
Results, prior-work audit and contribution decision ↗NEW DEVELOPMENT NOTE
The deadline pilot measures executed simulator decisions, including an avoidable missed window and a missed safer alternative. It is separate from the first proposal study.
Read the twelve-episode development note ↗AI SAFETY EVALUATION / RESEARCH PREVIEW
We test whether AI agents change course when consent is withdrawn, a safety condition fails, or an operator says stop.
Matched task histories. Controlled state changes. Inspectable action proposals.
THE QUESTION
A model can understand an instruction and still make a dangerous choice.
We design paired evaluations that measure harmful behavior alongside useful performance. A failed action, a safe refusal, and a technical error are different outcomes. Our reports keep them separate.
Explore the research agenda ↗WHAT WE MEASURE
01
Does the model recognize a dangerous request, refuse the unsafe step, and offer a useful alternative?
02
Does an ordinary objective lead to a dangerous shortcut, an unsupported claim, or an action beyond its authority?
03
Do untrusted documents, ambiguous permissions, or failed safeguards change what the agent actually executes?
RESEARCH STANDARD
Every published evaluation will identify the model version, test conditions, sample size, scoring rules, and uncertainty. We will publish failures in the harness alongside failures in the model.
How evaluation works ↗CURRENT STATUS / OCTOBER 2026
A narrow proxy collection has recorded 2,304 of 2,304 candidate requests. Descriptive observations remain preliminary, with independent label review and replication pending. Read the coverage report.
See the coverage policy ↗Inspired by RoboHarm. Independently developed. No affiliation or endorsement implied.