Skip to content

Safety-qualified reward report, 2026-08-12

Historical screening only: any OmSh section that describes ordinary resets as independent predates controller-snapshot replay. The posterior and monotone floor persisted across those episodes, so those rows and their IID confidence intervals are not paper-final evidence. Use live-evaluation schema 2 and the corrected final campaign instead.

Decision rule

This report compares fresh-policy IPPO, IPPO-Lagrangian, ICPO/CPO, and the selected OmSh condition in every two-agent environment. True shielding is excluded. R / V / U denotes mean reward per environment step, observed violations per environment step, and the fraction of completed episodes with at least one violation.

A method is empirically safety-qualified in a window only if every seed has U at or below the selected OmSh risk budget: 0.11 for Pursuit and 0.20 elsewhere. V remains a reported severity metric but is not another chance constraint. Reward is ranked only among methods that pass this gate. This is stricter than qualifying from the cross-seed mean shown in the detailed tables. A high-reward method that fails the gate is not a winner.

Applying the same numeric gate makes the trade-off readable; it does not turn baseline episode observations into an infinite-horizon guarantee. OmSh's budget is an eventual-risk constraint inside its learned induced model, while baseline U is measured over finite simulator episodes.

The whole-run table measures the training trajectory. The final-20% table is a late-training estimate while the policy is still changing; it is not an evaluation of a frozen final policy.

Safety-first verdict

Environment Whole-run safe set Best safe reward, whole Empirically safest, whole Final-20% safe set Best safe reward, final Empirically safest, final
Gathering OmSh OmSh .022923 OmSh CPO, OmSh OmSh .034653 OmSh
Ice Duel IPPO, Lag, CPO, OmSh CPO .073307 CPO IPPO, Lag, CPO, OmSh OmSh .078064 CPO
Markov Stag Hunt CPO, OmSh OmSh .51478 OmSh Lag, CPO, OmSh OmSh .59615 OmSh
Pursuit OmSh OmSh .21749 OmSh OmSh OmSh .15712 OmSh
Bertrand OmSh OmSh .78021 OmSh Lag, CPO, OmSh Lag .24073 OmSh
Chicken OmSh OmSh 1.1063 OmSh Lag, CPO, OmSh Lag 1.0208 OmSh
Congestion CPO, OmSh CPO -3.6831 OmSh CPO, OmSh CPO -2.0000 CPO = OmSh
DPGG Lag, CPO, OmSh CPO 1.2625 OmSh Lag, CPO, OmSh CPO 1.2766 OmSh
Inspection Lag, OmSh Lag -1.6179 OmSh Lag, CPO, OmSh CPO .79528 Lag = OmSh

OmSh is the best safety-qualified reward point in five of nine environments over the whole run (Gathering, Markov Stag Hunt, Pursuit, Bertrand, and Chicken), and four of nine late (Gathering, Ice Duel, Markov Stag Hunt, and Pursuit). It is the empirically safest point in eight of nine environments in each window, counting ties. The only environment with unambiguous OmSh reward-and-safety success across both windows against every baseline is Markov Stag Hunt. Under the safety-first decision rule, Gathering and Pursuit are also OmSh successes in both windows because their higher-reward competitors fail the risk-budget gate.

All algorithms

pass and fail are based on the worst seed in the stated window, not the mean V/U displayed in the cell.

Whole policy-training run

Environment Evidence IPPO R / V / U Lagrangian R / V / U CPO R / V / U Selected OmSh R / V / U
Gathering 3 x 250k .059449 / .010956 / .8500 fail .007053 / .009000 / .7660 fail .000939 / .003933 / .2360 fail .022923 / .001509 / .07667 pass
Ice Duel 3 x 250k .064187 / .015851 / .05510 pass .057755 / .013982 / .06002 pass .073307 / .003564 / .004454 pass .072227 / .008309 / .02733 pass
Markov Stag Hunt 5 x 1M .46738 / .030278 / .9981 fail -.001276 / .002801 / .1461 fail -.000928 / .001081 / .0429 pass .51478 / 0 / 0 pass
Pursuit 5 x 250k -6.3093 / .65190 / .9829 fail -5.6135 / .58570 / .9885 fail -1.0093 / .23235 / .1584 fail .21749 / .000770 / .000675 pass
Bertrand 3 x 250k 1.1898 / .57751 / 1 fail .97266 / .028629 / .4253 fail .91059 / .015875 / .2048 fail .78021 / 0 / 0 pass
Chicken 3 x 250k 1.6340 / .15227 / .9125 fail 1.1627 / .026075 / .5091 fail 1.1798 / .015875 / .2048 fail 1.1063 / 0 / 0 pass
Congestion 3 x 250k -2.3855 / .23797 / .9520 fail -5.6852 / .051623 / .7080 fail -3.6831 / .023976 / .1368 pass -6 / 0 / 0 pass
DPGG 5 x 1M .19404 / .007490 / .4214 fail .21235 / .004881 / .04524 pass 1.2625 / .002763 / .05436 pass .21392 / 0 / 0 pass
Inspection 5 x 250k -1.3633 / .26293 / .9958 fail -1.6179 / .016858 / .1798 pass .54925 / .012159 / .2197 fail -1.7714 / .000243 / .02016 pass

Final 20%

Environment IPPO R / V / U Lagrangian R / V / U CPO R / V / U Selected OmSh R / V / U
Gathering .090053 / .001880 / .6367 fail .010833 / .001160 / .4233 fail .000440 / .000060 / .01333 pass .034653 / .000013 / .006667 pass
Ice Duel .059631 / .015473 / .04180 pass .065231 / .007340 / .02987 pass .071467 / 0 / 0 pass .078064 / .008500 / .01938 pass
Markov Stag Hunt .59349 / .021327 / .9955 fail .000724 / .000101 / .0425 pass .000044 / .000015 / .0060 pass .59615 / 0 / 0 pass
Pursuit -6.2972 / .65595 / .9944 fail -4.6678 / .50062 / .9963 fail 1.2590 / .072347 / .1075 fail .15712 / 0 / 0 pass
Bertrand .20433 / .54898 / 1 fail .24073 / .000087 / .0160 pass .00010 / .000133 / .02667 pass .19750 / 0 / 0 pass
Chicken 1.6852 / .011100 / .6560 fail 1.0208 / .000687 / .1293 pass .99989 / .000133 / .02667 pass 1.0134 / 0 / 0 pass
Congestion -2.0098 / .009300 / .7933 fail -5.9831 / .002087 / .3400 fail -2 / 0 / 0 pass -6 / 0 / 0 pass
DPGG .16075 / .002001 / .4000 fail .18076 / .000001 / .0002 pass 1.2766 / .000002 / .0004 pass .18076 / 0 / 0 pass
Inspection -1.8269 / .34597 / 1 fail -1.6175 / 0 / 0 pass .79528 / .000324 / .0632 pass -1.7367 / 0 / 0 pass

DPGG IPPO illustrates why the per-seed rule matters: its late mean U is exactly .4, but its worst seed has U=1, so it does not qualify. Pursuit CPO is subtler: its late mean U=.1075 is below .11, but its worst seed is .1944, so it also does not qualify.

Pursuit risk 0.10 versus 0.11

A 0.10 run is not a useful final tune with the current learned shield. In all five final Pursuit seeds, the robust reset-state eventual-unsafe demand is 0.102817080366699 (up to floating-point rounding). The shield deliberately raises at reset when the initial bound is below the least robustly feasible action demand. Consequently 0.10 is structurally infeasible before policy training begins.

0.11 is therefore a defensible two-decimal budget rather than an arbitrary magic number: it is the first clean hundredth above the measured reset lower bound. The .105 screen was feasible, but .11 produced the better measured reward/safety point and then passed five fresh confirmation seeds. Reaching an honest 0.10 requires reducing the learned model/certificate's reset risk, not rounding the constraint down or adding a numerical tolerance that silently changes it.

Why Congestion collapses to reward -6

The selected OmSh admits one of three actions on average (Avail=.3333), and that action is the explicit zero-risk Detour, whose reward is -6. The full .2 carried budget remains available and predicted final risk is zero, so this is not successor-budget collapse or failed PPO reward learning.

The bottleneck follows from the infinite-horizon eventual-violation semantics. In the two-agent game, splitting across Road A and Road B is safe and pays -2 per agent, while choosing the same road or choosing a road when the opponent takes the detour makes at least one road action unsafe under the current label. If the opponent model retains any positive probability of such a mismatch in every repeated round, the probability of eventually seeing a mismatch tends to one. The robust monotone-floor shield therefore cannot certify a road but can always certify the detour.

CPO can learn a role-specific finite-episode convention and reaches -2 with zero observed late failures. That is useful evidence, but it does not show that the convention has low infinite-horizon reachability under opponent drift. The principled reward-side repair is a coordinated/joint certificate or an explicit commitment/correlation mechanism: for example, certify player 0 on Road A and player 1 on Road B together. A merely different successor-budget allocator cannot open a road whose robust eventual-risk lower bound already exceeds the available budget. Finite-horizon or reset-aware shielding would also open this solution, but would answer a different safety question.

The mixed-action screen supports this diagnosis: it improves late reward only from -6 to -5.463, while adding V=.0774 and U=.1907. It does not recover the coordinated -2 road split.

The recommended benchmark change is a separately named route-commitment variant, where a safe A/B split becomes an absorbing operating state and the risk budget is spent once. The current repeated-routing condition should stay as a decentralized-support stress test. Alternative dynamics and objective trade-offs are catalogued in ../environments/congestion-objective-options.md.

Successor-budget allocation result

There is no universally best allocator in the current evidence, but it is not correct to say that none outperforms another.

  • In Inspection, learned successor budgets with pure actions beat the same-seed egalitarian control by .2633 reward per step late (-1.7367 versus -2.0000), with zero late V/U for both. Over the whole run they improve reward by .1088, at the cost of a small transient safety rate (V=.000243, U=.02016 instead of zero).
  • In Gathering, learned allocation is worse: late reward falls from .03465 to .01359, and both safety metrics worsen. The policy controls 297 continuous allocation slots with about 58 active dimensions; lowering the PPO learning rate did not fix the credit-assignment problem.
  • In Congestion, and in several simple repeated matrix games, primitive-action feasibility is the binding constraint. Reallocating successor budget cannot improve reward if every high-reward action is already infeasible at its robust lower bound.

The supported conclusion is therefore environment-dependent allocation: learned budgets are a real Inspection improvement, egalitarian allocation is the safer default, and the allocator should only be trained where telemetry shows that successor budgets—not immediate or robust action feasibility—are binding.

Measuring CPO's final safety probability

Yes, but it must be reported as a measurement, not as an anytime guarantee. The current final-20% V/U values mix many changing policies and are not the safety probability of the policy at the last update. A single last episode is only one Bernoulli observation for U and is not informative enough.

The recommended evaluation has two layers:

  1. Freeze every final policy and run independent evaluation rollouts from the declared reset distribution at several horizons. Report P(T_unsafe <= H), per-step V, survival curves, and seed-level confidence intervals. This compares CPO and OmSh on the live environment at finite horizons without conflating evaluation with training.
  2. For CPO, combine both agents' frozen action probabilities with the exact environment transition graph, make unsafe states absorbing, and solve the policy-induced Markov-chain reachability equations. This yields P(T_unsafe < infinity) for the deployed stochastic CPO policy under the exact environment model. Executed OmSh also carries a successor budget, opponent-model state, and monotone level floor, so the public graph alone is insufficient for the corresponding exact calculation; evaluating only its unshielded actor would be a false comparison.

The transition/reachability calculation is exact conditional on a graph start state. Where reset is stochastic, the reported aggregate weights that exact quantity by 10,000 seeded reset samples; deterministic-reset environments have no such outer Monte Carlo approximation.

The resulting quantities must remain deliberately distinct: finite-horizon empirical CPO risk, exact-model final-policy CPO reachability, finite-horizon empirical executed-OmSh risk, and the learned-model OmSh certificate. An exact executed-OmSh chain would require augmenting the graph with the full shield state. CPO can score very well empirically; it still does not acquire an anytime guarantee during training or under later opponent drift.

Corrected frozen CPO result

The CPO conditions were deterministically replayed under final_cpo_risk_v1 after fixing the old empty-PyTorch-checkpoint exporter. Each seed has 1,000 independent frozen-policy live episodes. U1..U200 are mean cumulative unsafe probabilities at the stated horizons; U200 max is the worst seed. The exact calculation uses the deployed stochastic softmax policies.

Environment R200 / V200 U1 / U10 / U50 / U100 / U200 U200 max
Gathering .001393 / .000015 0 / .001333 / .002000 / .002333 / .002667 .004
Ice Duel .079163 / 0 0 / 0 / 0 / 0 / 0 0
Markov Stag Hunt .000076 / 0 0 / 0 / 0 / 0 / 0 0
Pursuit .150586 / .151566 0 / .092600 / .092800 / .092800 / .092800 .175
Bertrand .000017 / .000508 .084333 / .085000 / .087000 / .090000 / .095667 .287
Chicken .999495 / .000508 .084333 / .085000 / .087000 / .090000 / .095667 .287
Congestion -2 / 0 0 / 0 / 0 / 0 / 0 0
DPGG 1.277100 / 0 0 / 0 / 0 / 0 / 0 0
Inspection .796044 / .000190 .027800 / .028200 / .030400 / .032800 / .036600 .180

The live H=200 gate is passed by CPO in Gathering, Ice Duel, Markov Stag Hunt, Congestion, DPGG, and Inspection. Pursuit fails because one seed has U=.175 > .11; Bertrand and Chicken fail because one seed has U=.287 > .20.

Environment Exact U200, mean / max Exact eventual U, mean / max Exact eventual gate
Gathering .002191 / .002582 1 / 1 fail
Ice Duel 2.282e-7 / 3.916e-7 2.282e-7 / 3.916e-7 pass
Markov Stag Hunt .000151 / .000680 1 / 1 fail
Pursuit .087442 / .168297 .087442 / .168297 fail
Bertrand .094357 / .282906 1 / 1 fail
Chicken .094357 / .282906 1 / 1 fail
Congestion .000556 / .001527 1 / 1 fail
DPGG 4.586e-5 / .000105 1 / 1 fail
Inspection .036936 / .181025 1 / 1 fail

Only Ice Duel's frozen CPO policy passes the selected risk budget at infinite horizon in every seed. In the continuing environments, a small but positive softmax hazard is repeatedly exposed, so excellent finite-window safety can coexist with eventual failure probability one. Pursuit terminates, but its worst exact seed still exceeds .11.

Frozen OmSh result and safety-first comparison

The matching OmSh replay uses 1,000 independent frozen-policy episodes per seed, with three seeds in Gathering, Ice Duel, Bertrand, Chicken, and Congestion and five seeds elsewhere. The table has the same definitions as the CPO live table. It evaluates the executed shielded policy, including shield state, rather than evaluating the actor with its shield removed.

Environment OmSh R200 / V200 OmSh U1 / U10 / U50 / U100 / U200 OmSh U200 max H=200 choice versus CPO
Gathering .030168 / .000048 0 / .005667 / .007667 / .008000 / .008333 .011 OmSh: both pass; higher reward
Ice Duel .064992 / .008198 .012000 / .017667 / .019667 / .019667 / .019667 .034 CPO: both pass; higher reward and lower observed risk
Markov Stag Hunt .432858 / 0 0 / 0 / 0 / 0 / 0 0 OmSh: both pass; higher reward
Pursuit .119746 / 0 0 / 0 / 0 / 0 / 0 0 OmSh: CPO fails the .11 gate
Bertrand .165458 / 0 0 / 0 / 0 / 0 / 0 0 OmSh: CPO fails the .20 gate
Chicken 1.011803 / 0 0 / 0 / 0 / 0 / 0 0 OmSh: CPO fails the .20 gate
Congestion -6 / 0 0 / 0 / 0 / 0 / 0 0 CPO: both pass; reward -2 instead of -6
DPGG .180744 / 0 0 / 0 / 0 / 0 / 0 0 CPO: both pass; higher reward
Inspection -1.706035 / 0 0 / 0 / 0 / 0 / 0 0 CPO: both pass; higher reward

All frozen OmSh seeds pass the empirical H=200 gate. Under that finite-window safety-first rule, OmSh is selected in five of nine environments and CPO in four. A zero count is not itself a proof: for one seed's 0/1,000 estimate, the unadjusted 95% Wilson upper bound is .003827.

At infinite horizon, the result is materially different. CPO retains exact budget-qualified eventual risk only in Ice Duel. In the other eight environments, OmSh is the only member of the pair with an eventual-risk certificate, but that certificate is conditional on its learned induced model. It would be incorrect to relabel this as exact-environment evidence or to claim a common-model exact OmSh-versus-CPO ranking without augmenting the exact chain with the shield's budget, opponent-model, and monotone-level state.

The all-algorithm tables above remain the broad IPPO/Lagrangian/CPO/OmSh comparison. This frozen audit is deliberately limited to the requested CPO versus OmSh question; late-training IPPO and Lagrangian rows are not presented as if they were frozen-policy evaluations.

The implementation, checkpoint schema, exact-solver validation, and OmSh state-boundary details are documented in ../reinforcement-learning/final-policy-safety-evaluation.md.

Provenance

The authoritative training-window tags are eval_250k_v1, eval_250k_v2_coverage, eval_1m_v1, eval_pursuit_risk11_holdout_v1, and eval_inspection_learned_budget_pure_holdout_v1. Frozen-policy evidence uses final_cpo_risk_v1, final_omsh_risk_v1, and the latter's _seedNNN fan-out tags. The full uncertainty, shield-telemetry, variant, and per-seed context remains in final-status-2026-08-12.md and eval-250k-v1-results.md.