Existing WM and OM artifacts may be reused for the first policy-training pilots only when their recorded environment, public state/action, safety-label, transition, and reward contracts still match. In particular, Pursuit artifacts produced while interception rewards defaulted to zero must not supply reward-head evaluation, reward-aware shield ranking, or policy-return results after the sparse +10/-10 rewards were restored; rebuild or retrain those artifacts. Their transition coverage, calibration, and learned-versus-oracle diagnostics must still be checked before treating a poor OmSh result as an architectural failure.
The first fair comparison uses fresh policy initialization for IPPO, IPPO-Lagrangian, ICPO, and OmSh, with identical step budgets and seed indices. WM/OM pretraining is shared input to OmSh, not counted as policy initialization for any method.
The current per-environment risk budget (max_risk=0.2 in the curated notebooks) is retained for the first pass. It is below the requested 0.4 maximum and avoids changing the long-standing safety objectives before evidence warrants it.
Two-agent Pursuit and two-agent Congestion are evaluated before their larger-agent versions. Larger variants are not promoted unless the two-agent task is learnable and informative.
Reward and observed safety violations must both be reported. Shield-predicted eventual risk, immediate-risk calibration, overrides, unsafe proposals, missing graph coverage, admissible-action fraction, posterior level changes, and credible-floor posterior-tail constraint misses are diagnostics rather than substitutes for observed safety.
A method is not called safer solely because its configured risk allowance is lower. Comparisons use the same environment, safety label, training steps, and seeds wherever possible.
A safety-violation probability above 0.4 is unacceptable for promotion, even if reward is high. The analysis should also report uncertainty across seeds rather than classify from a single mean.
The initial OmSh candidate is pure primitive action plus egalitarian successor-budget allocation, monotone-floor opponent handling, and Bayesian reward ranking. This is the simplest improved condition and retains the strongest current robust interpretation.
Bayesian-mixture, all-level, credible-floor, robust-reward, mixed-action, and learned-budget variants are deferred until the main pilots identify which conservatism or allocation bottleneck is material. This avoids spending the full ablation budget on variants that are already clearly dominated.
The pilots triggered the mixed-action, learned-budget, and credible-floor ablations. They remain separately labelled conditions rather than silent changes to the control, and any non-default action/budget allocator must be selected on disjoint optimisation seeds before confirmation. The final primary campaign reports OmSh-Monotone only; OmSh-Credible remains an opt-in diagnostic because, under the retained schedule, it cannot expose a larger robust feasible set than the monotone floor.
A later two-shielded-agent experiment will require an explicit wrapper/state/budget ownership design. It should not be approximated by silently sharing one mutable IOP or one focal-agent budget between wrappers.
Seeds 0--2 are screening/evaluation seeds, not hyperparameter-selection seeds. Any targeted tuning prompted by this pilot uses a disjoint block such as seed_offset=100; a selected condition must then be rerun on fresh confirmation seeds before it supports a comparative claim.
A condition that breaches the 0.4 observed-safety ceiling is not promoted. If predicted final risk remains within budget while observed violations breach the ceiling, investigate transition coverage, unsafe-label agreement, OM calibration, and risk calibration before loosening the budget or adding reward flexibility.
Missing graph coverage or infeasible carried budgets indicate a soundness/runtime defect or an inadequate abstraction. They are repaired before interpreting reward, rather than treated as an ordinary reward--safety trade-off.
Low admissible-action availability, budgets concentrated near their continuation lower bounds, or near-zero safety margin indicate that successor-budget allocation is binding. The next targeted variant is learned successor-budget allocation while retaining robust projection.
Adequate action availability but frequent unsafe proposals/overrides and improving late reward indicate policy-learning inefficiency. The first responses are longer training or mask-aware policy/critic improvements, not a less conservative safety certificate.
When pure primitive actions are individually too restrictive but certified mixtures produce a materially larger feasible set, test mixed-action shielding. Learned-budget pure action precedes learned-budget mixed action unless the pilot directly shows this pure-action feasibility bottleneck.
Only aggregate behaviour across the tuning seeds selects parameters. A particular seed, trajectory, state identity, or environment-specific action is never encoded into the algorithm.
The apparent 0.15 Pursuit result on seeds 500--502 was rejected after it
failed three of five seeds in the 700--704 holdout. That block then became
a stress/tuning block for the lowest feasible allowances, .105 and .11.
.11 dominated .105; comparative claims use the separately launched
five-seed 800--804 confirmation, not either tuning mean.
Pursuit has a sharp seed-dependent policy/feasible-action transition between
.11 and .15 under the current abstraction; .16 and higher fail as well.
This is an empirical result, not a claim that .11 is a mathematically
optimal budget or that untested intermediate real values cannot behave
differently.
The .10 Pursuit target is structurally infeasible for the retained learned
shield: its reset-support demand is approximately .102817, before policy
optimisation. The final two-agent Pursuit condition therefore uses .11,
the first clean hundredth above that demand; no numerical tolerance is used
to make .10 appear feasible. Pursuit-3 is excluded for the separately
documented transition-graph scaling limit.
The two-shielded-agent Markov Stag Hunt experiment is a paired empirical
comparison. Each shield retains a certificate in its own learned induced
model, but simultaneously shielding the modelled opponent changes the live
policy and does not create an arbitrary-opponent joint certificate.
DPGG ICPO/CPO has one reproducible high-return seed in which the regulated
player withholds while the unregulated co-player contributes. It remains in
the reported mean. Median or trimmed comparisons may be shown as sensitivity
analyses, but must not replace the pre-declared mean.
The final-20% window is a learning-performance diagnostic, not an
infinite-horizon safety proof. Exact-anytime results apply only under the
explicitly reported exact-game or induced-model assumptions; empirical
finite training episodes do not strengthen those assumptions.
Three- and five-seed results are evidence for triage and confirmation, not
definitive significance tests. Any paper-level universal superiority claim
requires more seeds and a pre-declared statistical analysis.