CURRENT COLLECTION
The first scoped proxy collection has recorded 2,304 of 2,304 candidate conditions. Independent labeling and replication remain pending. The research program below includes proposed extensions.
Observed proposals, coverage and limitationsRESEARCH PROGRAM / PROPOSED
A testable contribution.
A bounded funding case.
We seek support for a controlled study of agent behavior after safety conditions change. Novelty remains subject to prior-art and external review.
The candidate contribution
Bounded prior-case overlap audit ↗Hold a scripted task history fixed. Change only the verified authorization or safety state before the next action. Compare the changed branch with an unchanged branch and measure restraint alongside useful completion.
Dynamic interactions, benign controls, and action scoring already exist in prior work. Our proposed contribution is a candidate controlled slice of synthetic state interventions and native proposal scoring. We do not claim to be first or unique.
Prior-art and contribution register ↗Current scoped study / October 2, 2026
The final planned collection contains 24 context variants across four constraint mechanisms, two scripted history conditions, two matched state branches, and three repeated calls per condition: 288 calls per candidate model, 2,304 in total. Context variants and repetitions are dependent; this is a descriptive safe proxy study.
Native tool proposals are recorded without execution. The fixed assistant histories do not establish actual agent commitment or intent. The full scoped collection ran after ten automated preparation critique rounds. They do not replace independent scientific review. Provider feasibility and missingness will be reported explicitly.
Current scope, analysis, and exclusions ↗Frozen candidate configuration ↗Earlier feasibility pilot
Four original synthetic text cases: consent withdrawal, interlock failure, recipient mismatch, and operator stop. Each model receives a preparation step and two matched continuations. The runner collects constrained JSON proposals and never executes them.
The planned pilot has eight model candidates and 96 calls. It has no repeats, visual inputs, native tool calls, or physical actions. It checks feasibility and cannot support a frontier safety ranking.
Exact scoring contract ↗Machine-readable manifest ↗Proposed 12-week program
A provisional USD 40,000 request covers engineering, independent review, model access, replication, and contingency. These are estimates. Reviewers and replication partners are not yet secured.
Milestones: audit prior art; review 24 matched scenarios; freeze the harness and analysis; collect powered comparisons; complete independent labels and replication.
Stop or redirect the program if the contribution duplicates existing work, controls fail, or access cannot support a valid comparison. Funding does not control scoring or publication.
Read the scoped funding proposal ↓For labs and public-sector evaluators
Specific method-review questions ↗Substantive response register ↗Native evidence and current-freeze provenance ↗The intended evidence supports one decision: whether a mitigation improves behavior after permission or safety conditions change, while preserving useful completion. It does not certify compliance or general safety.
We welcome a method reviewer, a replication partner, or a funder for the scoped milestones. No government or lab relationship is implied.
Discuss a specific contribution ↗EXECUTABLE STUDY FREEZE
Configuration and source are preserved together. This preparation freeze is not a completed collection.
Exact source, catalog, and eligibility archive ↗After the first collection
The explicit-rule proxy produced no complete proposals among valid changed-state responses. A proposed next study asks whether cached and atomic authorization gates preserve task usefulness while preventing simulated actions after a state change. It specifies target-generated preparation, versioned event order, and separate proposal and application outcomes.
An offline authorization kernel and symbolic boundary fixtures are implemented. These are engineering tests, with no model calls or real actions. The full rollout harness is incomplete; the design still requires independent prior-case, oracle and statistical review. Its sample-size sensitivity is illustrative, not a claim that the current budget supports a confirmatory study.
Prospective rollout design and decision criteria ↗