Define the decision
Agree the workflow, authority boundaries, hazards, benign tasks and the configuration choice the report must inform.
SERVICES / SCOPED ENGAGEMENTS
Bring one agent workflow and one configuration decision. We will scope a test that exposes what the safeguard covers, what it misses, and what a repair costs in useful work.
For teams integrating models with tools, the first engagement compares the current safeguard with one proposed repair in a synthetic replica of a specific workflow.
Agree the workflow, authority boundaries, hazards, benign tasks and the configuration choice the report must inform.
Reproduce safe and deliberately faulty execution, document the tool semantics, and make event replay independently inspectable.
Record executed effects, useful completion, technical failures and costs. Deliver versioned code, traces, uncertainty and a recommendation bounded by the evidence.
Send the agent’s task, tool interfaces, proposed safeguard and the decision you face. Do not send credentials, customer records or other confidential material in an initial inquiry. We agree scope, data handling, access limits, timing and a quote before work begins.
We do not certify general model safety or promise that an evaluation will discover a failure. MurderBench is an early evaluation initiative; no client engagements, external endorsements or production deployments are claimed.
Discuss your workflow ↗Inspect the public evidence →Public research preserves publication independence. A private engagement has an explicit disclosure agreement. Sponsorship does not buy favorable scores, and any material relationship behind a public report is disclosed.