OmSh evaluation status, 2026-08-12¶
Verdict¶
The improvements produce one unambiguous all-baseline success in both windows: Markov Stag Hunt. OmSh earns more reward and records fewer observed failures than fresh IPPO, IPPO-Lagrangian, and ICPO/CPO over both the whole run and final 20%. Pursuit also strictly beats all three over its whole run; late, it remains safer than CPO but earns less reward.
The broader result is not universal dominance. Gathering beats both constrained baselines on reward and safety but gives up reward to unsafe IPPO. Ice Duel beats IPPO and is competitive with the constrained methods, but CPO remains safer. DPGG is safe and stable but trails a legitimate high-return CPO seed. Bertrand and Chicken have useful late CPO wins, Inspection's learned budgets improve standard OmSh materially, and Congestion remains reward-bound.
No shielding variant rules every environment. There is therefore no evidence yet for archiving the alternatives on a separate branch.
Protocol¶
- The primary baselines are fresh IPPO, fresh IPPO-Lagrangian, and fresh ICPO/CPO. True shielding is excluded from every comparison below.
- OmSh policy training also starts fresh. Its world-model and opponent-model pretraining is shared input to the shield and is intentionally not counted as policy training.
R / V / Umeans reward per environment step, observed violations per environment step, and the fraction of completed episodes containing at least one violation.- Metrics are computed within each seed before averaging seeds. The late window contains episodes ending in the final 20% of the policy-training run.
- A condition passes the empirical gate only when every OmSh seed in both
windows has
V <= 0.4andU <= 0.4. This is deliberately stricter than accepting a safe aggregate mean. - Markov Stag Hunt and DPGG use five seeds at one million steps. Corrected Pursuit and the promoted Inspection variant use five seeds at 250,000 steps. The other environments use three seeds at 250,000 steps and remain screens.
- The reported OmSh risk is an infinite-horizon eventual-violation quantity inside the learned WM/OM-induced model. Finite training histories are empirical validation, not an original-game infinite-horizon theorem.
Per-environment status¶
"Strict wins" lists baselines for which selected OmSh has higher reward and
lower V and U. A safety tie does not count as a strict win.
| Environment | Selected OmSh | Gate | Strict wins, whole / late | Status |
|---|---|---|---|---|
| Gathering | pure, egalitarian budget, risk 0.2 |
pass | Lag+CPO / Lag+CPO | Strong constrained-baseline result; unsafe IPPO has more reward |
| Ice Duel | pure, egalitarian budget, risk 0.2, coverage WM |
pass | IPPO+Lag / IPPO | Higher late reward than CPO, but CPO is safer |
| Markov Stag Hunt | pure, egalitarian budget, risk 0.2 |
pass | IPPO+Lag+CPO / IPPO+Lag+CPO | Full success at 1M steps and five seeds |
| Pursuit | pure, egalitarian budget, risk 0.11 |
pass | IPPO+Lag+CPO / IPPO+Lag | Full whole-run success; safer than CPO but lower-reward late |
| Bertrand | pure, egalitarian budget, risk 0.2 |
pass | none / CPO | Zero failures; reward remains below IPPO and Lag |
| Chicken | pure, egalitarian budget, risk 0.2 |
pass | none / CPO | Zero failures; late Lag reward gap is only 0.00735 |
| Congestion | pure, egalitarian budget, risk 0.2 |
pass | none / none | Safe but reward-collapsed |
| DPGG | pure, egalitarian budget, risk 0.2 |
pass | IPPO+Lag / IPPO | Stable zero-failure result; CPO mean has a valid high-return seed |
| Inspection | learned budget, pure action, risk 0.2 |
pass | none / IPPO | Part 4 improves standard OmSh, but not the constrained baselines |
Core results¶
These tables report means. Sample standard deviations for selected OmSh appear in the next section. The authoritative result tags retain every per-seed value.
Whole policy-training run¶
| Environment | Evidence | IPPO R / V / U |
Lagrangian R / V / U |
CPO R / V / U |
Selected OmSh R / V / U |
|---|---|---|---|---|---|
| Gathering | 3 x 250k | .059449 / .010956 / .8500 |
.007053 / .009000 / .7660 |
.000939 / .003933 / .2360 |
.022923 / .001509 / .07667 |
| Ice Duel | 3 x 250k | .064187 / .015851 / .05510 |
.057755 / .013982 / .06002 |
.073307 / .003564 / .004454 |
.072227 / .008309 / .02733 |
| Markov Stag Hunt | 5 x 1M | .46738 / .030278 / .9981 |
-.001276 / .002801 / .1461 |
-.000928 / .001081 / .0429 |
.51478 / 0 / 0 |
| Pursuit | 5 x 250k | -6.3093 / .65190 / .9829 |
-5.6135 / .58570 / .9885 |
-1.0093 / .23235 / .1584 |
.21749 / .000770 / .000675 |
| Bertrand | 3 x 250k | 1.1898 / .57751 / 1 |
.97266 / .028629 / .4253 |
.91059 / .015875 / .2048 |
.78021 / 0 / 0 |
| Chicken | 3 x 250k | 1.6340 / .15227 / .9125 |
1.1627 / .026075 / .5091 |
1.1798 / .015875 / .2048 |
1.1063 / 0 / 0 |
| Congestion | 3 x 250k | -2.3855 / .23797 / .9520 |
-5.6852 / .051623 / .7080 |
-3.6831 / .023976 / .1368 |
-6 / 0 / 0 |
| DPGG | 5 x 1M | .19404 / .007490 / .4214 |
.21235 / .004881 / .04524 |
1.2625 / .002763 / .05436 |
.21392 / 0 / 0 |
| Inspection | 5 x 250k | -1.3633 / .26293 / .9958 |
-1.6179 / .016858 / .1798 |
.54925 / .012159 / .2197 |
-1.7714 / .000243 / .02016 |
Final 20%¶
| Environment | IPPO R / V / U |
Lagrangian R / V / U |
CPO R / V / U |
Selected OmSh R / V / U |
|---|---|---|---|---|
| Gathering | .090053 / .001880 / .6367 |
.010833 / .001160 / .4233 |
.000440 / .000060 / .01333 |
.034653 / .000013 / .006667 |
| Ice Duel | .059631 / .015473 / .04180 |
.065231 / .007340 / .02987 |
.071467 / 0 / 0 |
.078064 / .008500 / .01938 |
| Markov Stag Hunt | .59349 / .021327 / .9955 |
.000724 / .000101 / .0425 |
.000044 / .000015 / .0060 |
.59615 / 0 / 0 |
| Pursuit | -6.2972 / .65595 / .9944 |
-4.6678 / .50062 / .9963 |
1.2590 / .072347 / .1075 |
.15712 / 0 / 0 |
| Bertrand | .20433 / .54898 / 1 |
.24073 / .000087 / .0160 |
.00010 / .000133 / .02667 |
.19750 / 0 / 0 |
| Chicken | 1.6852 / .011100 / .6560 |
1.0208 / .000687 / .1293 |
.99989 / .000133 / .02667 |
1.0134 / 0 / 0 |
| Congestion | -2.0098 / .009300 / .7933 |
-5.9831 / .002087 / .3400 |
-2 / 0 / 0 |
-6 / 0 / 0 |
| DPGG | .16075 / .002001 / .4000 |
.18076 / .000001 / .0002 |
1.2766 / .000002 / .0004 |
.18076 / 0 / 0 |
| Inspection | -1.8269 / .34597 / 1 |
-1.6175 / 0 / 0 |
.79528 / .000324 / .0632 |
-1.7367 / 0 / 0 |
Selected OmSh uncertainty and shield diagnostics¶
Whole-run values are per-seed mean +/- sample standard deviation. Avail is
the fraction of primitive actions admitted by the shield. Budget and Risk
are the mean carried conditional budget and final predicted eventual risk.
Worst is the largest observed V or U in any seed/window.
| Environment | Reward/step | Reward/episode | Violations/step | Unsafe episodes | Avail | Budget | Risk | Worst |
|---|---|---|---|---|---|---|---|---|
| Gathering | .022923 +/- .00421 |
11.461 +/- 2.11 |
.001509 +/- .000541 |
.07667 +/- .0117 |
.8065 |
.08523 |
.00321 |
.0900 |
| Ice Duel | .072227 +/- .00986 |
.2442 +/- .100 |
.008309 +/- .000252 |
.02733 +/- .00857 |
.9726 |
.19854 |
.00513 |
.03243 |
| Markov Stag Hunt | .51478 +/- .0153 |
257.39 +/- 7.63 |
0 |
0 |
.9324 |
.2000 |
0 |
0 |
| Pursuit | .21749 +/- .0227 |
9.6578 +/- .514 |
.000770 +/- .00112 |
.000675 +/- .000749 |
.2769 |
.03547 |
.03342 |
.00251 |
| Bertrand | .78021 +/- .0101 |
156.04 +/- 2.01 |
0 |
0 |
.5000 |
.2000 |
0 |
0 |
| Chicken | 1.1063 +/- .00216 |
221.26 +/- .433 |
0 |
0 |
.5000 |
.2000 |
0 |
0 |
| Congestion | -6 +/- 0 |
-1200 +/- 0 |
0 |
0 |
.3333 |
.2000 |
0 |
0 |
| DPGG | .21392 +/- .00467 |
42.784 +/- .934 |
0 |
0 |
.5000 |
.2000 |
0 |
0 |
| Inspection | -1.7714 +/- .0664 |
-354.28 +/- 13.3 |
.000243 +/- .000061 |
.02016 +/- .00327 |
.5584 |
.16635 |
.02143 |
.0248 |
Late selected-OmSh reward/step is .034653 +/- .00551 in Gathering,
.078064 +/- .0204 in Ice Duel, .59615 +/- .0270 in Markov Stag Hunt,
.15712 +/- .0657 in Pursuit, .19750 +/- .0173 in Bertrand,
1.0134 +/- .00105 in Chicken, -6 +/- 0 in Congestion,
.18076 +/- .000018 in DPGG, and -1.7367 +/- .150 in Inspection. Every one
of these late conditions passes the per-seed safety ceiling.
Variant decisions¶
| Environment | Evidence-driven decision |
|---|---|
| Gathering | Keep standard pure/egalitarian OmSh. Learned budgets lowered late reward from .03465 to .01359 and worsened both safety metrics; lower PPO step size did not fix the 297-dimensional allocation problem. |
| Ice Duel | Keep risk .2 as the reward point. Risk .01 is a safer frontier but lowers late reward to .0573; .05 is dominated, and zero risk collapses reward to .00176 late. None beats CPO on safety. |
| Markov Stag Hunt | Keep standard OmSh. It is the only variant needed for the primary result; the coordinated two-agent experiment is reported separately below. |
| Pursuit | Promote risk .11. Risk .15 looked safe on three tuning seeds but failed three of five fresh seeds; .16, .18, .20, and .25 also fail. .105 and .11 clear the failed seed block, .11 dominates .105, and .11 passes five new confirmation seeds. |
| Bertrand | Keep pure OmSh as the conservative point. Mixed actions raise whole-run reward to 1.438 but also raise V/U to .138/.256, so they do not dominate the constrained baselines. |
| Chicken | Reject mixed actions: standard OmSh has higher late reward and zero rather than nonzero observed failures. |
| Congestion | Neither point is competitive. Mixed actions improve late reward from -6 to -5.463, but introduce .0774/.1907 failures and remain far below CPO's -2. |
| DPGG | Keep pure OmSh. The mixed screen gains negligible late reward while adding observed failures. |
| Inspection | Promote learned budgets with pure actions. Against the same-seed standard control they improve late reward by .2633 while both variants retain zero late failures. Mixed and combined modes remain near -1.995. |
Two simultaneously shielded agents¶
The paired Markov Stag Hunt experiment uses policy seeds 600--602 and the
same WM pretraining, with a role-specific IOP for each protected agent.
| Condition/agent | Whole-run reward/step | Late reward/step | Violations/step | Unsafe episodes |
|---|---|---|---|---|
| One OmSh, player 0 | .31403 +/- .00322 |
.46555 +/- .00549 |
0 |
0 |
| Two OmSh, player 0 | .32174 +/- .00248 |
.48176 +/- .01084 |
0 |
0 |
| Two OmSh, player 1 | .34721 +/- .02293 |
.52819 +/- .01275 |
0 |
0 |
Shielding both agents raises player-0 reward in every paired seed: the mean
gain is .00771 over the whole run and .01621 late, with no observed safety
regression. Each role still has a certificate only in its own learned induced
model. Changing the live opponent by shielding it does not create a new joint
arbitrary-opponent theorem.
Correctness and interpretation notes¶
- Old Pursuit results with zero returns are invalid. The live environment now
restores sparse
+10/-10interception rewards, both learned and exact graphs include independent action slips, and the WM reward loss is scaled so those raw targets do not overwhelm transition representation learning. - Corrected Pursuit held-out state accuracy is
.9960; exhaustive transition TV is.1282mean and.1673p95. This is good enough to evaluate, not good enough to equate the learned shield with exact game dynamics. - Ice Duel has complete pair coverage over 49,500 state-action pairs. Its
coverage-aware WM has transition TV
.0302mean and.1673p95, while the learned-versus-exact false-safe rate is.06--.15%across levels. - DPGG's high CPO mean is driven by one valid seed (
5.587whole-run reward/step) where the regulated player withholds and the unregulated player contributes. It is retained in the mean; median sensitivity does not replace the pre-declared estimator. - Zero empirical failures do not establish zero original-game risk. They mean no failures were observed at the stated seed and step count.
Provenance¶
| Purpose | Tag/job |
|---|---|
| Base 250k campaign | eval_250k_v1 |
| Markov Stag Hunt and DPGG 1M confirmation | eval_1m_v1, jobs 6735/6736 |
| Coverage-corrected Ice Duel | eval_250k_v2_coverage, job 6771 |
| Inspection learned-budget holdout/control | eval_inspection_learned_budget_pure_holdout_v1, job 6750; standard control job 6767 |
| Corrected Pursuit reward campaign | eval_250k_v3_rewards, jobs 6780/6781 |
| Pursuit risk screens | tune_pursuit_risk{105,11,15,16,18,20}_v1, jobs 6782--6785 and 6789--6790 |
| Pursuit final confirmation | eval_pursuit_risk11_holdout_v1, job 6791 |
| One-vs-two-agent OmSh | eval_msh_one_omsh_control_v1 / eval_msh_two_omsh_v1, jobs 6786/6787 |
The earlier detailed pilot and targeted ablations remain in
eval-250k-v1-results.md. Human-review assumptions
are maintained in assumptions.md, and implementation
changes in changes.md.