Skip to content

OmSh evaluation status, 2026-08-12

Verdict

The improvements produce one unambiguous all-baseline success in both windows: Markov Stag Hunt. OmSh earns more reward and records fewer observed failures than fresh IPPO, IPPO-Lagrangian, and ICPO/CPO over both the whole run and final 20%. Pursuit also strictly beats all three over its whole run; late, it remains safer than CPO but earns less reward.

The broader result is not universal dominance. Gathering beats both constrained baselines on reward and safety but gives up reward to unsafe IPPO. Ice Duel beats IPPO and is competitive with the constrained methods, but CPO remains safer. DPGG is safe and stable but trails a legitimate high-return CPO seed. Bertrand and Chicken have useful late CPO wins, Inspection's learned budgets improve standard OmSh materially, and Congestion remains reward-bound.

No shielding variant rules every environment. There is therefore no evidence yet for archiving the alternatives on a separate branch.

Protocol

  • The primary baselines are fresh IPPO, fresh IPPO-Lagrangian, and fresh ICPO/CPO. True shielding is excluded from every comparison below.
  • OmSh policy training also starts fresh. Its world-model and opponent-model pretraining is shared input to the shield and is intentionally not counted as policy training.
  • R / V / U means reward per environment step, observed violations per environment step, and the fraction of completed episodes containing at least one violation.
  • Metrics are computed within each seed before averaging seeds. The late window contains episodes ending in the final 20% of the policy-training run.
  • A condition passes the empirical gate only when every OmSh seed in both windows has V <= 0.4 and U <= 0.4. This is deliberately stricter than accepting a safe aggregate mean.
  • Markov Stag Hunt and DPGG use five seeds at one million steps. Corrected Pursuit and the promoted Inspection variant use five seeds at 250,000 steps. The other environments use three seeds at 250,000 steps and remain screens.
  • The reported OmSh risk is an infinite-horizon eventual-violation quantity inside the learned WM/OM-induced model. Finite training histories are empirical validation, not an original-game infinite-horizon theorem.

Per-environment status

"Strict wins" lists baselines for which selected OmSh has higher reward and lower V and U. A safety tie does not count as a strict win.

Environment Selected OmSh Gate Strict wins, whole / late Status
Gathering pure, egalitarian budget, risk 0.2 pass Lag+CPO / Lag+CPO Strong constrained-baseline result; unsafe IPPO has more reward
Ice Duel pure, egalitarian budget, risk 0.2, coverage WM pass IPPO+Lag / IPPO Higher late reward than CPO, but CPO is safer
Markov Stag Hunt pure, egalitarian budget, risk 0.2 pass IPPO+Lag+CPO / IPPO+Lag+CPO Full success at 1M steps and five seeds
Pursuit pure, egalitarian budget, risk 0.11 pass IPPO+Lag+CPO / IPPO+Lag Full whole-run success; safer than CPO but lower-reward late
Bertrand pure, egalitarian budget, risk 0.2 pass none / CPO Zero failures; reward remains below IPPO and Lag
Chicken pure, egalitarian budget, risk 0.2 pass none / CPO Zero failures; late Lag reward gap is only 0.00735
Congestion pure, egalitarian budget, risk 0.2 pass none / none Safe but reward-collapsed
DPGG pure, egalitarian budget, risk 0.2 pass IPPO+Lag / IPPO Stable zero-failure result; CPO mean has a valid high-return seed
Inspection learned budget, pure action, risk 0.2 pass none / IPPO Part 4 improves standard OmSh, but not the constrained baselines

Core results

These tables report means. Sample standard deviations for selected OmSh appear in the next section. The authoritative result tags retain every per-seed value.

Whole policy-training run

Environment Evidence IPPO R / V / U Lagrangian R / V / U CPO R / V / U Selected OmSh R / V / U
Gathering 3 x 250k .059449 / .010956 / .8500 .007053 / .009000 / .7660 .000939 / .003933 / .2360 .022923 / .001509 / .07667
Ice Duel 3 x 250k .064187 / .015851 / .05510 .057755 / .013982 / .06002 .073307 / .003564 / .004454 .072227 / .008309 / .02733
Markov Stag Hunt 5 x 1M .46738 / .030278 / .9981 -.001276 / .002801 / .1461 -.000928 / .001081 / .0429 .51478 / 0 / 0
Pursuit 5 x 250k -6.3093 / .65190 / .9829 -5.6135 / .58570 / .9885 -1.0093 / .23235 / .1584 .21749 / .000770 / .000675
Bertrand 3 x 250k 1.1898 / .57751 / 1 .97266 / .028629 / .4253 .91059 / .015875 / .2048 .78021 / 0 / 0
Chicken 3 x 250k 1.6340 / .15227 / .9125 1.1627 / .026075 / .5091 1.1798 / .015875 / .2048 1.1063 / 0 / 0
Congestion 3 x 250k -2.3855 / .23797 / .9520 -5.6852 / .051623 / .7080 -3.6831 / .023976 / .1368 -6 / 0 / 0
DPGG 5 x 1M .19404 / .007490 / .4214 .21235 / .004881 / .04524 1.2625 / .002763 / .05436 .21392 / 0 / 0
Inspection 5 x 250k -1.3633 / .26293 / .9958 -1.6179 / .016858 / .1798 .54925 / .012159 / .2197 -1.7714 / .000243 / .02016

Final 20%

Environment IPPO R / V / U Lagrangian R / V / U CPO R / V / U Selected OmSh R / V / U
Gathering .090053 / .001880 / .6367 .010833 / .001160 / .4233 .000440 / .000060 / .01333 .034653 / .000013 / .006667
Ice Duel .059631 / .015473 / .04180 .065231 / .007340 / .02987 .071467 / 0 / 0 .078064 / .008500 / .01938
Markov Stag Hunt .59349 / .021327 / .9955 .000724 / .000101 / .0425 .000044 / .000015 / .0060 .59615 / 0 / 0
Pursuit -6.2972 / .65595 / .9944 -4.6678 / .50062 / .9963 1.2590 / .072347 / .1075 .15712 / 0 / 0
Bertrand .20433 / .54898 / 1 .24073 / .000087 / .0160 .00010 / .000133 / .02667 .19750 / 0 / 0
Chicken 1.6852 / .011100 / .6560 1.0208 / .000687 / .1293 .99989 / .000133 / .02667 1.0134 / 0 / 0
Congestion -2.0098 / .009300 / .7933 -5.9831 / .002087 / .3400 -2 / 0 / 0 -6 / 0 / 0
DPGG .16075 / .002001 / .4000 .18076 / .000001 / .0002 1.2766 / .000002 / .0004 .18076 / 0 / 0
Inspection -1.8269 / .34597 / 1 -1.6175 / 0 / 0 .79528 / .000324 / .0632 -1.7367 / 0 / 0

Selected OmSh uncertainty and shield diagnostics

Whole-run values are per-seed mean +/- sample standard deviation. Avail is the fraction of primitive actions admitted by the shield. Budget and Risk are the mean carried conditional budget and final predicted eventual risk. Worst is the largest observed V or U in any seed/window.

Environment Reward/step Reward/episode Violations/step Unsafe episodes Avail Budget Risk Worst
Gathering .022923 +/- .00421 11.461 +/- 2.11 .001509 +/- .000541 .07667 +/- .0117 .8065 .08523 .00321 .0900
Ice Duel .072227 +/- .00986 .2442 +/- .100 .008309 +/- .000252 .02733 +/- .00857 .9726 .19854 .00513 .03243
Markov Stag Hunt .51478 +/- .0153 257.39 +/- 7.63 0 0 .9324 .2000 0 0
Pursuit .21749 +/- .0227 9.6578 +/- .514 .000770 +/- .00112 .000675 +/- .000749 .2769 .03547 .03342 .00251
Bertrand .78021 +/- .0101 156.04 +/- 2.01 0 0 .5000 .2000 0 0
Chicken 1.1063 +/- .00216 221.26 +/- .433 0 0 .5000 .2000 0 0
Congestion -6 +/- 0 -1200 +/- 0 0 0 .3333 .2000 0 0
DPGG .21392 +/- .00467 42.784 +/- .934 0 0 .5000 .2000 0 0
Inspection -1.7714 +/- .0664 -354.28 +/- 13.3 .000243 +/- .000061 .02016 +/- .00327 .5584 .16635 .02143 .0248

Late selected-OmSh reward/step is .034653 +/- .00551 in Gathering, .078064 +/- .0204 in Ice Duel, .59615 +/- .0270 in Markov Stag Hunt, .15712 +/- .0657 in Pursuit, .19750 +/- .0173 in Bertrand, 1.0134 +/- .00105 in Chicken, -6 +/- 0 in Congestion, .18076 +/- .000018 in DPGG, and -1.7367 +/- .150 in Inspection. Every one of these late conditions passes the per-seed safety ceiling.

Variant decisions

Environment Evidence-driven decision
Gathering Keep standard pure/egalitarian OmSh. Learned budgets lowered late reward from .03465 to .01359 and worsened both safety metrics; lower PPO step size did not fix the 297-dimensional allocation problem.
Ice Duel Keep risk .2 as the reward point. Risk .01 is a safer frontier but lowers late reward to .0573; .05 is dominated, and zero risk collapses reward to .00176 late. None beats CPO on safety.
Markov Stag Hunt Keep standard OmSh. It is the only variant needed for the primary result; the coordinated two-agent experiment is reported separately below.
Pursuit Promote risk .11. Risk .15 looked safe on three tuning seeds but failed three of five fresh seeds; .16, .18, .20, and .25 also fail. .105 and .11 clear the failed seed block, .11 dominates .105, and .11 passes five new confirmation seeds.
Bertrand Keep pure OmSh as the conservative point. Mixed actions raise whole-run reward to 1.438 but also raise V/U to .138/.256, so they do not dominate the constrained baselines.
Chicken Reject mixed actions: standard OmSh has higher late reward and zero rather than nonzero observed failures.
Congestion Neither point is competitive. Mixed actions improve late reward from -6 to -5.463, but introduce .0774/.1907 failures and remain far below CPO's -2.
DPGG Keep pure OmSh. The mixed screen gains negligible late reward while adding observed failures.
Inspection Promote learned budgets with pure actions. Against the same-seed standard control they improve late reward by .2633 while both variants retain zero late failures. Mixed and combined modes remain near -1.995.

Two simultaneously shielded agents

The paired Markov Stag Hunt experiment uses policy seeds 600--602 and the same WM pretraining, with a role-specific IOP for each protected agent.

Condition/agent Whole-run reward/step Late reward/step Violations/step Unsafe episodes
One OmSh, player 0 .31403 +/- .00322 .46555 +/- .00549 0 0
Two OmSh, player 0 .32174 +/- .00248 .48176 +/- .01084 0 0
Two OmSh, player 1 .34721 +/- .02293 .52819 +/- .01275 0 0

Shielding both agents raises player-0 reward in every paired seed: the mean gain is .00771 over the whole run and .01621 late, with no observed safety regression. Each role still has a certificate only in its own learned induced model. Changing the live opponent by shielding it does not create a new joint arbitrary-opponent theorem.

Correctness and interpretation notes

  • Old Pursuit results with zero returns are invalid. The live environment now restores sparse +10/-10 interception rewards, both learned and exact graphs include independent action slips, and the WM reward loss is scaled so those raw targets do not overwhelm transition representation learning.
  • Corrected Pursuit held-out state accuracy is .9960; exhaustive transition TV is .1282 mean and .1673 p95. This is good enough to evaluate, not good enough to equate the learned shield with exact game dynamics.
  • Ice Duel has complete pair coverage over 49,500 state-action pairs. Its coverage-aware WM has transition TV .0302 mean and .1673 p95, while the learned-versus-exact false-safe rate is .06--.15% across levels.
  • DPGG's high CPO mean is driven by one valid seed (5.587 whole-run reward/step) where the regulated player withholds and the unregulated player contributes. It is retained in the mean; median sensitivity does not replace the pre-declared estimator.
  • Zero empirical failures do not establish zero original-game risk. They mean no failures were observed at the stated seed and step count.

Provenance

Purpose Tag/job
Base 250k campaign eval_250k_v1
Markov Stag Hunt and DPGG 1M confirmation eval_1m_v1, jobs 6735/6736
Coverage-corrected Ice Duel eval_250k_v2_coverage, job 6771
Inspection learned-budget holdout/control eval_inspection_learned_budget_pure_holdout_v1, job 6750; standard control job 6767
Corrected Pursuit reward campaign eval_250k_v3_rewards, jobs 6780/6781
Pursuit risk screens tune_pursuit_risk{105,11,15,16,18,20}_v1, jobs 6782--6785 and 6789--6790
Pursuit final confirmation eval_pursuit_risk11_holdout_v1, job 6791
One-vs-two-agent OmSh eval_msh_one_omsh_control_v1 / eval_msh_two_omsh_v1, jobs 6786/6787

The earlier detailed pilot and targeted ablations remain in eval-250k-v1-results.md. Human-review assumptions are maintained in assumptions.md, and implementation changes in changes.md.