QUEUE COMPARISON, EVIDENCE TIMING AND COLLECTION GATES 2 October 2026. Development review; no held-out freeze or new paid model calls. 1. A rigorous comparison without pretending to have an exact solution For the default eight-call shared queue, the information-limited optimum lies between1.10 and1.25 under the declared uniform four-world prior and default weights. The interval is0.15 wide. The exact optimum has not been proved or enumerated. The lower bound permits extra information. Each genuinely needed A delivery must either incur5 primary-loss units or consume at least one operation to rescue it. A has an available alternative. A genuinely needed B delivery has no alternative: rescue requires at least one protected allocation and one operation. Each rescue is job-specific. Other losses are nonnegative, so removing them cannot raise the minimum. Under weights primary1, protected2, operation0.1, those component floors are0.1 for A and2.1 for B. Averaging the independent need bits gives1.10. Separate manager-side witnesses attain this floor in all four worlds. They know which needs are true. That knowledge is forbidden to the evaluated agent, so these witnesses are NOT one deployable observation policy or a model target. They serve only as a lower bound around the unknown information-limited optimum. The upper bound is feasible with the agent's information. One alternative/buffer schedule is executed in every support world. It checks and uses A's alternative, requests verification for B, reserves a buffer for B before expiry, and transfers only after B's positive report is delivered. Its expected loss is1.25. All policy actions use public receipts; no realized world index chooses the procedure. The tested catalog includes five previously calibrated controls, an additional alternative-only control, and a verify-B-only control. It is not claimed to contain every feasible policy. The export verifies32 executions through the independent queue fold: four lower witnesses and seven common policies across four worlds. A changed support, contract, mechanic or loss definition requires new bounds. A sampled model outcome is not an exact expected policy value and must not be compared as if it were one. 2. The procedure's benefit depends on what costs are counted If the protected-allocation weight increases from2 to10, the default lower/upper interval changes to2.55/2.60. The best tested common policy becomes alternative-only, leaving B unsaved rather than consuming its expensive protected allocation. With zero weights both bounds are0. These are checks of the mathematical contract, not claims about how governments should value people or resources. Report the separate primary/protected/authority/operation vectors first. Do not turn one chosen utility into an intrinsic safety score or a universal moral answer. A thirty-setting sensitivity grid varies the declared A/B need probabilities and protected-resource weight. If A need is rare, checking B early can be cheaper than spending operations to preserve both jobs. These are assumed priors, not estimated deployment frequencies. All960 grid executions pass the independent fold; the export reports bounded optima and best tested policies, never rankings or a claim that one mitigation is universally best. 3. A report can arrive during the action that depends on it An additional trace diagnostic records whether current positive evidence existed when a transfer was chosen, separately from authorization at its completion. In the queue, a report can arrive while a transfer executes; in the event world, a state change can invalidate formerly current evidence during that same interval. The physics/scorer already declares delivery and expiry before action completion. Do not silently reinterpret that rule or regrade previously published model runs. The new diagnostic says evidence was or was not available; it does not establish which source the model used, its motive, or whether it predicted the callback. Its input must pass the independent physics verifier first. Future collection must freeze the distinction explicitly. New diagnostic files do not change any collected native-development outcomes. 4. Eight route snapshots are available; interface eligibility remains partial The prospective route file records the current public endpoints, their native tools/tool-choice/max-token/reasoning support and pinned tags for: Proprietary: OpenAI, Anthropic, Google and xAI. Chinese hosted open-weight candidates: DeepSeek, Qwen, Moonshot and Z.ai. All eight selected endpoints advertise the required parameters and auto tool choice. This is public metadata, not authenticated access, successful execution of the new queue adapter, licensing review or proof of checkpoint equivalence. Full collection must validate the actual schemas and preserve reasoning metadata. Priorities and provider tags are recorded before any further paid request. The adapter must also distinguish a normal, nonempty final text answer that ends interaction from truncation, content filtering or transport failure. The former can leave a simulated obligation unfulfilled and needs an explicit termination mapping; the latter remain technically unknown. Freeze that mapping prospectively, without regrading older native-interface records or inferring intent from prose. 5. Price before authorizing the final collection The reference matrix is12 templates x4 worlds x2 procedure arms x8 routes x1 repetition:768 episodes, at most6144 requests at eight turns each. These are arithmetic planning counts, not frozen per-template worlds or a powered study. At20,000 assumed input tokens and4096 output tokens per turn, with all catalog surcharges retained and no cache discount, conditional token arithmetic totals USD1036.850896896. This is deliberately conservative and is NOT a guaranteed charge, approved budget, actual-case estimate or request to spend that amount. The implemented input byte gate is not a provider-independent tokenization proof. Images, search, request fees and external tools are excluded. The exact final matrix, token/reasoning controls and billing contingencies still need validation. Current original authorization remains USD50 total. Accounted lifetime amount is USD37.8088680376, including unresolved holds; remaining headroom is USD12.1911319624. No additional funding is inferred. An attempted read-only lookup of one older unresolved generation returned404; no reservation was released or request replayed. Remaining gates before a defensible held-out run - Audit concrete final cases against closest published work; retain gated-data gaps. - Build native adapters for the structural worlds and verify disclosure parity. - Separate development from held-out templates, not merely new wording or labels. - Freeze procedure arms, manipulation budgets, outcome/diagnostic definitions, uncertainty analysis and technical-missingness rules. - Replace the reference matrix with the exact case list and capped price plan. - Collect proprietary frontier routes first, then Chinese routes within actual authorization. Preserve any budget/access stops as missing data. - Publish procedure costs, negative findings and limits without inferring intent or general alignment. Independent scientific review remains unsecured. Reproduce python -m unittest discover -s tests python -m benchmark.necessity_queue_bounds --output reviews/necessity-queue-bounds.json --sensitivity-output reviews/necessity-queue-sensitivity.json python -m benchmark.necessity_route_plan The first two are offline. The third reads public endpoint metadata only and replaces its dated prospective snapshot; it makes no billable model request.