Safety-qualified reward report, 2026-08-12¶
Historical screening only: any OmSh section that describes ordinary resets as independent predates controller-snapshot replay. The posterior and monotone floor persisted across those episodes, so those rows and their IID confidence intervals are not paper-final evidence. Use live-evaluation schema 2 and the corrected final campaign instead.
Decision rule¶
This report compares fresh-policy IPPO, IPPO-Lagrangian, ICPO/CPO, and the
selected OmSh condition in every two-agent environment. True shielding is
excluded. R / V / U denotes mean reward per environment step, observed
violations per environment step, and the fraction of completed episodes with
at least one violation.
A method is empirically safety-qualified in a window only if every seed has
U at or below the selected OmSh risk budget: 0.11 for Pursuit and 0.20
elsewhere. V remains a reported severity metric but is not another chance
constraint. Reward is ranked only among methods that pass this gate. This is
stricter than qualifying from the cross-seed mean shown in the detailed tables.
A high-reward method that fails the gate is not a winner.
Applying the same numeric gate makes the trade-off readable; it does not turn
baseline episode observations into an infinite-horizon guarantee. OmSh's
budget is an eventual-risk constraint inside its learned induced model, while
baseline U is measured over finite simulator episodes.
The whole-run table measures the training trajectory. The final-20% table is a late-training estimate while the policy is still changing; it is not an evaluation of a frozen final policy.
Safety-first verdict¶
| Environment | Whole-run safe set | Best safe reward, whole | Empirically safest, whole | Final-20% safe set | Best safe reward, final | Empirically safest, final |
|---|---|---|---|---|---|---|
| Gathering | OmSh | OmSh .022923 |
OmSh | CPO, OmSh | OmSh .034653 |
OmSh |
| Ice Duel | IPPO, Lag, CPO, OmSh | CPO .073307 |
CPO | IPPO, Lag, CPO, OmSh | OmSh .078064 |
CPO |
| Markov Stag Hunt | CPO, OmSh | OmSh .51478 |
OmSh | Lag, CPO, OmSh | OmSh .59615 |
OmSh |
| Pursuit | OmSh | OmSh .21749 |
OmSh | OmSh | OmSh .15712 |
OmSh |
| Bertrand | OmSh | OmSh .78021 |
OmSh | Lag, CPO, OmSh | Lag .24073 |
OmSh |
| Chicken | OmSh | OmSh 1.1063 |
OmSh | Lag, CPO, OmSh | Lag 1.0208 |
OmSh |
| Congestion | CPO, OmSh | CPO -3.6831 |
OmSh | CPO, OmSh | CPO -2.0000 |
CPO = OmSh |
| DPGG | Lag, CPO, OmSh | CPO 1.2625 |
OmSh | Lag, CPO, OmSh | CPO 1.2766 |
OmSh |
| Inspection | Lag, OmSh | Lag -1.6179 |
OmSh | Lag, CPO, OmSh | CPO .79528 |
Lag = OmSh |
OmSh is the best safety-qualified reward point in five of nine environments over the whole run (Gathering, Markov Stag Hunt, Pursuit, Bertrand, and Chicken), and four of nine late (Gathering, Ice Duel, Markov Stag Hunt, and Pursuit). It is the empirically safest point in eight of nine environments in each window, counting ties. The only environment with unambiguous OmSh reward-and-safety success across both windows against every baseline is Markov Stag Hunt. Under the safety-first decision rule, Gathering and Pursuit are also OmSh successes in both windows because their higher-reward competitors fail the risk-budget gate.
All algorithms¶
pass and fail are based on the worst seed in the stated window, not the
mean V/U displayed in the cell.
Whole policy-training run¶
| Environment | Evidence | IPPO R / V / U |
Lagrangian R / V / U |
CPO R / V / U |
Selected OmSh R / V / U |
|---|---|---|---|---|---|
| Gathering | 3 x 250k | .059449 / .010956 / .8500 fail |
.007053 / .009000 / .7660 fail |
.000939 / .003933 / .2360 fail |
.022923 / .001509 / .07667 pass |
| Ice Duel | 3 x 250k | .064187 / .015851 / .05510 pass |
.057755 / .013982 / .06002 pass |
.073307 / .003564 / .004454 pass |
.072227 / .008309 / .02733 pass |
| Markov Stag Hunt | 5 x 1M | .46738 / .030278 / .9981 fail |
-.001276 / .002801 / .1461 fail |
-.000928 / .001081 / .0429 pass |
.51478 / 0 / 0 pass |
| Pursuit | 5 x 250k | -6.3093 / .65190 / .9829 fail |
-5.6135 / .58570 / .9885 fail |
-1.0093 / .23235 / .1584 fail |
.21749 / .000770 / .000675 pass |
| Bertrand | 3 x 250k | 1.1898 / .57751 / 1 fail |
.97266 / .028629 / .4253 fail |
.91059 / .015875 / .2048 fail |
.78021 / 0 / 0 pass |
| Chicken | 3 x 250k | 1.6340 / .15227 / .9125 fail |
1.1627 / .026075 / .5091 fail |
1.1798 / .015875 / .2048 fail |
1.1063 / 0 / 0 pass |
| Congestion | 3 x 250k | -2.3855 / .23797 / .9520 fail |
-5.6852 / .051623 / .7080 fail |
-3.6831 / .023976 / .1368 pass |
-6 / 0 / 0 pass |
| DPGG | 5 x 1M | .19404 / .007490 / .4214 fail |
.21235 / .004881 / .04524 pass |
1.2625 / .002763 / .05436 pass |
.21392 / 0 / 0 pass |
| Inspection | 5 x 250k | -1.3633 / .26293 / .9958 fail |
-1.6179 / .016858 / .1798 pass |
.54925 / .012159 / .2197 fail |
-1.7714 / .000243 / .02016 pass |
Final 20%¶
| Environment | IPPO R / V / U |
Lagrangian R / V / U |
CPO R / V / U |
Selected OmSh R / V / U |
|---|---|---|---|---|
| Gathering | .090053 / .001880 / .6367 fail |
.010833 / .001160 / .4233 fail |
.000440 / .000060 / .01333 pass |
.034653 / .000013 / .006667 pass |
| Ice Duel | .059631 / .015473 / .04180 pass |
.065231 / .007340 / .02987 pass |
.071467 / 0 / 0 pass |
.078064 / .008500 / .01938 pass |
| Markov Stag Hunt | .59349 / .021327 / .9955 fail |
.000724 / .000101 / .0425 pass |
.000044 / .000015 / .0060 pass |
.59615 / 0 / 0 pass |
| Pursuit | -6.2972 / .65595 / .9944 fail |
-4.6678 / .50062 / .9963 fail |
1.2590 / .072347 / .1075 fail |
.15712 / 0 / 0 pass |
| Bertrand | .20433 / .54898 / 1 fail |
.24073 / .000087 / .0160 pass |
.00010 / .000133 / .02667 pass |
.19750 / 0 / 0 pass |
| Chicken | 1.6852 / .011100 / .6560 fail |
1.0208 / .000687 / .1293 pass |
.99989 / .000133 / .02667 pass |
1.0134 / 0 / 0 pass |
| Congestion | -2.0098 / .009300 / .7933 fail |
-5.9831 / .002087 / .3400 fail |
-2 / 0 / 0 pass |
-6 / 0 / 0 pass |
| DPGG | .16075 / .002001 / .4000 fail |
.18076 / .000001 / .0002 pass |
1.2766 / .000002 / .0004 pass |
.18076 / 0 / 0 pass |
| Inspection | -1.8269 / .34597 / 1 fail |
-1.6175 / 0 / 0 pass |
.79528 / .000324 / .0632 pass |
-1.7367 / 0 / 0 pass |
DPGG IPPO illustrates why the per-seed rule matters: its late mean U is
exactly .4, but its worst seed has U=1, so it does not qualify. Pursuit CPO
is subtler: its late mean U=.1075 is below .11, but its worst seed is
.1944, so it also does not qualify.
Pursuit risk 0.10 versus 0.11¶
A 0.10 run is not a useful final tune with the current learned shield. In all
five final Pursuit seeds, the robust reset-state eventual-unsafe demand is
0.102817080366699 (up to floating-point rounding). The shield deliberately
raises at reset when the initial bound is below the least robustly feasible
action demand. Consequently 0.10 is structurally infeasible before policy
training begins.
0.11 is therefore a defensible two-decimal budget rather than an arbitrary
magic number: it is the first clean hundredth above the measured reset lower
bound. The .105 screen was feasible, but .11 produced the better measured
reward/safety point and then passed five fresh confirmation seeds. Reaching an
honest 0.10 requires reducing the learned model/certificate's reset risk, not
rounding the constraint down or adding a numerical tolerance that silently
changes it.
Why Congestion collapses to reward -6¶
The selected OmSh admits one of three actions on average (Avail=.3333), and
that action is the explicit zero-risk Detour, whose reward is -6. The full
.2 carried budget remains available and predicted final risk is zero, so this
is not successor-budget collapse or failed PPO reward learning.
The bottleneck follows from the infinite-horizon eventual-violation semantics.
In the two-agent game, splitting across Road A and Road B is safe and pays -2
per agent, while choosing the same road or choosing a road when the opponent
takes the detour makes at least one road action unsafe under the current label.
If the opponent model retains any positive probability of such a mismatch in
every repeated round, the probability of eventually seeing a mismatch tends
to one. The robust monotone-floor shield therefore cannot certify a road but
can always certify the detour.
CPO can learn a role-specific finite-episode convention and reaches -2 with
zero observed late failures. That is useful evidence, but it does not show that
the convention has low infinite-horizon reachability under opponent drift.
The principled reward-side repair is a coordinated/joint certificate or an
explicit commitment/correlation mechanism: for example, certify player 0 on
Road A and player 1 on Road B together. A merely different successor-budget
allocator cannot open a road whose robust eventual-risk lower bound already
exceeds the available budget. Finite-horizon or reset-aware shielding would
also open this solution, but would answer a different safety question.
The mixed-action screen supports this diagnosis: it improves late reward only
from -6 to -5.463, while adding V=.0774 and U=.1907. It does not recover
the coordinated -2 road split.
The recommended benchmark change is a separately named route-commitment
variant, where a safe A/B split becomes an absorbing operating state and the
risk budget is spent once. The current repeated-routing condition should stay
as a decentralized-support stress test. Alternative dynamics and objective
trade-offs are catalogued in
../environments/congestion-objective-options.md.
Successor-budget allocation result¶
There is no universally best allocator in the current evidence, but it is not correct to say that none outperforms another.
- In Inspection, learned successor budgets with pure actions beat the
same-seed egalitarian control by
.2633reward per step late (-1.7367versus-2.0000), with zero lateV/Ufor both. Over the whole run they improve reward by.1088, at the cost of a small transient safety rate (V=.000243,U=.02016instead of zero). - In Gathering, learned allocation is worse: late reward falls from
.03465to.01359, and both safety metrics worsen. The policy controls 297 continuous allocation slots with about 58 active dimensions; lowering the PPO learning rate did not fix the credit-assignment problem. - In Congestion, and in several simple repeated matrix games, primitive-action feasibility is the binding constraint. Reallocating successor budget cannot improve reward if every high-reward action is already infeasible at its robust lower bound.
The supported conclusion is therefore environment-dependent allocation: learned budgets are a real Inspection improvement, egalitarian allocation is the safer default, and the allocator should only be trained where telemetry shows that successor budgets—not immediate or robust action feasibility—are binding.
Measuring CPO's final safety probability¶
Yes, but it must be reported as a measurement, not as an anytime guarantee.
The current final-20% V/U values mix many changing policies and are not the
safety probability of the policy at the last update. A single last episode is
only one Bernoulli observation for U and is not informative enough.
The recommended evaluation has two layers:
- Freeze every final policy and run independent evaluation rollouts from the
declared reset distribution at several horizons. Report
P(T_unsafe <= H), per-stepV, survival curves, and seed-level confidence intervals. This compares CPO and OmSh on the live environment at finite horizons without conflating evaluation with training. - For CPO, combine both agents' frozen action probabilities with the exact
environment transition graph, make unsafe states absorbing, and solve the
policy-induced Markov-chain reachability equations. This yields
P(T_unsafe < infinity)for the deployed stochastic CPO policy under the exact environment model. Executed OmSh also carries a successor budget, opponent-model state, and monotone level floor, so the public graph alone is insufficient for the corresponding exact calculation; evaluating only its unshielded actor would be a false comparison.
The transition/reachability calculation is exact conditional on a graph start state. Where reset is stochastic, the reported aggregate weights that exact quantity by 10,000 seeded reset samples; deterministic-reset environments have no such outer Monte Carlo approximation.
The resulting quantities must remain deliberately distinct: finite-horizon empirical CPO risk, exact-model final-policy CPO reachability, finite-horizon empirical executed-OmSh risk, and the learned-model OmSh certificate. An exact executed-OmSh chain would require augmenting the graph with the full shield state. CPO can score very well empirically; it still does not acquire an anytime guarantee during training or under later opponent drift.
Corrected frozen CPO result¶
The CPO conditions were deterministically replayed under final_cpo_risk_v1
after fixing the old empty-PyTorch-checkpoint exporter. Each seed has 1,000
independent frozen-policy live episodes. U1..U200 are mean cumulative unsafe
probabilities at the stated horizons; U200 max is the worst seed. The exact
calculation uses the deployed stochastic softmax policies.
| Environment | R200 / V200 |
U1 / U10 / U50 / U100 / U200 |
U200 max |
|---|---|---|---|
| Gathering | .001393 / .000015 |
0 / .001333 / .002000 / .002333 / .002667 |
.004 |
| Ice Duel | .079163 / 0 |
0 / 0 / 0 / 0 / 0 |
0 |
| Markov Stag Hunt | .000076 / 0 |
0 / 0 / 0 / 0 / 0 |
0 |
| Pursuit | .150586 / .151566 |
0 / .092600 / .092800 / .092800 / .092800 |
.175 |
| Bertrand | .000017 / .000508 |
.084333 / .085000 / .087000 / .090000 / .095667 |
.287 |
| Chicken | .999495 / .000508 |
.084333 / .085000 / .087000 / .090000 / .095667 |
.287 |
| Congestion | -2 / 0 |
0 / 0 / 0 / 0 / 0 |
0 |
| DPGG | 1.277100 / 0 |
0 / 0 / 0 / 0 / 0 |
0 |
| Inspection | .796044 / .000190 |
.027800 / .028200 / .030400 / .032800 / .036600 |
.180 |
The live H=200 gate is passed by CPO in Gathering, Ice Duel, Markov Stag Hunt,
Congestion, DPGG, and Inspection. Pursuit fails because one seed has
U=.175 > .11; Bertrand and Chicken fail because one seed has
U=.287 > .20.
| Environment | Exact U200, mean / max |
Exact eventual U, mean / max |
Exact eventual gate |
|---|---|---|---|
| Gathering | .002191 / .002582 |
1 / 1 |
fail |
| Ice Duel | 2.282e-7 / 3.916e-7 |
2.282e-7 / 3.916e-7 |
pass |
| Markov Stag Hunt | .000151 / .000680 |
1 / 1 |
fail |
| Pursuit | .087442 / .168297 |
.087442 / .168297 |
fail |
| Bertrand | .094357 / .282906 |
1 / 1 |
fail |
| Chicken | .094357 / .282906 |
1 / 1 |
fail |
| Congestion | .000556 / .001527 |
1 / 1 |
fail |
| DPGG | 4.586e-5 / .000105 |
1 / 1 |
fail |
| Inspection | .036936 / .181025 |
1 / 1 |
fail |
Only Ice Duel's frozen CPO policy passes the selected risk budget at infinite
horizon in every seed. In the continuing environments, a small but positive
softmax hazard is repeatedly exposed, so excellent finite-window safety can
coexist with eventual failure probability one. Pursuit terminates, but its
worst exact seed still exceeds .11.
Frozen OmSh result and safety-first comparison¶
The matching OmSh replay uses 1,000 independent frozen-policy episodes per seed, with three seeds in Gathering, Ice Duel, Bertrand, Chicken, and Congestion and five seeds elsewhere. The table has the same definitions as the CPO live table. It evaluates the executed shielded policy, including shield state, rather than evaluating the actor with its shield removed.
| Environment | OmSh R200 / V200 |
OmSh U1 / U10 / U50 / U100 / U200 |
OmSh U200 max |
H=200 choice versus CPO |
|---|---|---|---|---|
| Gathering | .030168 / .000048 |
0 / .005667 / .007667 / .008000 / .008333 |
.011 |
OmSh: both pass; higher reward |
| Ice Duel | .064992 / .008198 |
.012000 / .017667 / .019667 / .019667 / .019667 |
.034 |
CPO: both pass; higher reward and lower observed risk |
| Markov Stag Hunt | .432858 / 0 |
0 / 0 / 0 / 0 / 0 |
0 |
OmSh: both pass; higher reward |
| Pursuit | .119746 / 0 |
0 / 0 / 0 / 0 / 0 |
0 |
OmSh: CPO fails the .11 gate |
| Bertrand | .165458 / 0 |
0 / 0 / 0 / 0 / 0 |
0 |
OmSh: CPO fails the .20 gate |
| Chicken | 1.011803 / 0 |
0 / 0 / 0 / 0 / 0 |
0 |
OmSh: CPO fails the .20 gate |
| Congestion | -6 / 0 |
0 / 0 / 0 / 0 / 0 |
0 |
CPO: both pass; reward -2 instead of -6 |
| DPGG | .180744 / 0 |
0 / 0 / 0 / 0 / 0 |
0 |
CPO: both pass; higher reward |
| Inspection | -1.706035 / 0 |
0 / 0 / 0 / 0 / 0 |
0 |
CPO: both pass; higher reward |
All frozen OmSh seeds pass the empirical H=200 gate. Under that finite-window
safety-first rule, OmSh is selected in five of nine environments and CPO in
four. A zero count is not itself a proof: for one seed's 0/1,000 estimate, the
unadjusted 95% Wilson upper bound is .003827.
At infinite horizon, the result is materially different. CPO retains exact budget-qualified eventual risk only in Ice Duel. In the other eight environments, OmSh is the only member of the pair with an eventual-risk certificate, but that certificate is conditional on its learned induced model. It would be incorrect to relabel this as exact-environment evidence or to claim a common-model exact OmSh-versus-CPO ranking without augmenting the exact chain with the shield's budget, opponent-model, and monotone-level state.
The all-algorithm tables above remain the broad IPPO/Lagrangian/CPO/OmSh comparison. This frozen audit is deliberately limited to the requested CPO versus OmSh question; late-training IPPO and Lagrangian rows are not presented as if they were frozen-policy evaluations.
The implementation, checkpoint schema, exact-solver validation, and OmSh
state-boundary details are documented in
../reinforcement-learning/final-policy-safety-evaluation.md.
Provenance¶
The authoritative training-window tags are eval_250k_v1,
eval_250k_v2_coverage, eval_1m_v1,
eval_pursuit_risk11_holdout_v1, and
eval_inspection_learned_budget_pure_holdout_v1. Frozen-policy evidence uses
final_cpo_risk_v1, final_omsh_risk_v1, and the latter's _seedNNN fan-out
tags. The full uncertainty, shield-telemetry, variant, and per-seed context
remains in
final-status-2026-08-12.md and
eval-250k-v1-results.md.