Skip to content

eval_250k_v1 results

This is the live result record for the 250,000-step, three-seed screen. Each comparison cell is reward / violations per step / unsafe episodes; +, =, and - mean that OmSh is better, tied, or worse than the named baseline on the paired-seed mean. The final window contains episodes ending in the last 20% of each run. The empirical gate separately rejects any OmSh seed/window with per-step violation probability or unsafe-episode probability above 0.4.

The runs are screening evidence rather than significance claims. The primary OmSh condition is pure primitive actions, robust egalitarian successor budgets, a monotone opponent-model floor, and Bayesian expected reward.

Environment Window OmSh reward/step Violations/step Unsafe episodes vs IPPO vs Lagrangian vs ICPO/CPO Gate
Gathering all 0.022923 0.0015093 0.076667 - / + / + + / + / + + / + / + pass
Gathering final 0.034653 0.0000133 0.006667 - / + / + + / + / + + / + / + pass
Ice Duel all 0.060073 0.0089773 0.022586 - / + / + + / + / + - / - / - pass
Ice Duel final 0.059200 0.0124533 0.020095 - / + / + + / - / + - / - / - pass
Markov Stag Hunt all 0.315317 0 0 + / + / + + / + / + + / + / + pass
Markov Stag Hunt final 0.470000 0 0 + / + / + + / + / + + / + / + pass
Pursuit all invalid invalid invalid invalid invalid invalid invalid
Pursuit final invalid invalid invalid invalid invalid invalid invalid
Bertrand all 0.780207 0 0 - / + / + - / + / + - / + / + pass
Bertrand final 0.197500 0 0 - / + / + - / + / + + / + / + pass
Chicken all 1.106285 0 0 - / + / + - / + / + - / + / + pass
Chicken final 1.013427 0 0 - / + / + - / + / + + / + / + pass
Congestion all -6.000000 0 0 - / + / + - / + / + - / + / + pass
Congestion final -6.000000 0 0 - / + / + - / + / + - / = / = pass
DPGG all 0.312409 0 0 + / + / + + / + / + - / + / + pass
DPGG final 0.181055 0 0 + / + / + + / = / = + / + / + pass
Inspection all -1.878972 0 0 - / + / + - / + / + - / + / + pass
Inspection final -2.000000 0 0 - / + / + - / + / + - / = / = pass

Current interpretation

  • Markov Stag Hunt is the only completed environment where the primary OmSh condition strictly improves reward and both observed safety measures over all three baselines in both windows. Its worst OmSh safety probability is zero.
  • DPGG Pareto-dominates all three baselines in the final window. Its overall reward trails ICPO because ICPO earns high early reward and then collapses; OmSh's late reward is higher than every baseline with zero observed failures.
  • Gathering beats both constrained baselines on reward and safety but trails unsafe IPPO on reward. Its late unsafe-episode mean is 0.0067, while the carried budget falls to roughly 0.014--0.027, making successor-budget allocation the targeted bottleneck.
  • Ice Duel passes the absolute safety ceiling but trails ICPO on reward and both observed safety measures. Its late action availability is 0.994, its mean carried budget remains 0.1996, and its mean predicted final risk is only 0.0017 despite 0.0125 observed violations per step. This is a transition or risk-calibration problem, not evidence that mixed actions or learned budgets are the next intervention. Job 6738 builds the exact transition graph and exports learned-versus-exact diagnostics before further tuning.
  • Bertrand and Chicken beat ICPO late on reward and safety, but retain small reward gaps to IPPO-Lagrangian and/or IPPO. Congestion and Inspection remain substantially reward-limited. These four matrix games expose only one safe pure action (one half of actions, or one third in Congestion) while leaving the configured risk budget unused, motivating mixed primitive actions.

Tuning results

These rows use tuning seeds rather than the base campaign's seeds, so their means are unpaired screens. They are accepted only if every seed in both the whole-run and final-window summaries passes the empirical safety ceiling.

Environment Condition Seeds Window Reward/step Violations/step Unsafe episodes Worst seed/window Gate Decision
Gathering learned budget, pure 100--102 all 0.008759 0.0002893 0.088667 0.1160 pass reject
Gathering learned budget, pure 100--102 final 0.013587 0.0000267 0.013333 0.1160 pass reject
Bertrand egalitarian mixed 100--102 all 1.437876 0.138136 0.255733 0.3240 pass combine with learned budgets
Bertrand egalitarian mixed 100--102 final 0.924740 0.213067 0.310667 0.3240 pass combine with learned budgets
Chicken egalitarian mixed 100--102 all 1.129864 0.0817573 0.285067 0.3600 pass reject
Chicken egalitarian mixed 100--102 final 0.995687 0.046060 0.336000 0.3600 pass reject
Congestion egalitarian mixed 100--102 all -5.491392 0.0764093 0.194933 0.1976 pass reject
Congestion egalitarian mixed 100--102 final -5.463107 0.077360 0.190667 0.1976 pass reject
Inspection egalitarian mixed 100--102 all -1.868756 0.004256 0.029067 0.0296 pass combine with learned budgets
Inspection egalitarian mixed 100--102 final -1.995013 0 0 0.0296 pass combine with learned budgets
Inspection learned budget, pure 300--302 all -1.730252 0.0001573 0.017333 0.0216 pass promote
Inspection learned budget, pure 300--302 final -1.665047 0 0 0.0216 pass promote
Inspection learned budget, mixed 300--302 all -1.876732 0.0003240 0.033333 0.0424 pass reject
Inspection learned budget, mixed 300--302 final -1.994587 0 0 0.0424 pass reject
Bertrand learned budget, mixed 300--302 all 0.822361 0.0075827 0.249067 0.2920 pass reject
Bertrand learned budget, mixed 300--302 final 0.247867 0.0149333 0.272000 0.2920 pass reject
Bertrand learned budget, pure 300--302 all 1.072192 0.0095400 0.228267 0.6720 fail reject
Bertrand learned budget, pure 300--302 final 0.735293 0.0081867 0.229333 0.6720 fail reject
Gathering learned budget, pure, LR 3e-5 300--302 all 0.004925 0.0011053 0.280000 0.2840 pass reject
Gathering learned budget, pure, LR 3e-5 300--302 final 0.005160 0.0009600 0.246667 0.2840 pass reject

Gathering's learned-budget policy is dominated by primary OmSh in the final window: it earns 0.0136 rather than 0.0347, and its carried budget falls further to approximately 0.0117. The current high-dimensional projected PPO action did not learn the intended high-value allocation. The checkpoint has 297 continuous slots and averages 58.0 active dimensions; joint PPO KL averages 0.0705 and reaches 11.29, compared with mean 0.00094 and maximum 0.0114 for Bertrand's one-dimensional mixed action. Reducing the learning rate from 2e-4 to 3e-5 makes the update conservative (mean KL 0.000113, maximum 0.000895) but worsens final reward to 0.0052 and unsafe episodes to 0.2467. It is dominated by both the primary OmSh condition and the first learned-budget screen. The remaining Gathering problem is not fixed by scalar step-size scaling; high-dimensional credit assignment and exploration require a different treatment.

Bertrand's mixed-only policy beats every baseline on reward, but it does not beat the constrained baselines on safety. Its final violation probability is 0.2131 and unsafe-episode probability is 0.3107. The common-seed 2-by-2 follow-up does not justify promotion. Combining mixed actions with learned budgets passes the ceiling in every seed/window, but falls to 0.2479 final reward and 0.2720 unsafe episodes. Learned budgets with pure actions bifurcate: one seed reaches 1.8239 final reward with 0.6720 unsafe episodes, while the other two seeds remain near 0.19 reward with at most 0.0160 unsafe episodes. Consequently the pure arm fails the hard gate and the combined arm is dominated by the mixed-only screen. This is evidence that the two extensions should remain independently selectable rather than being assumed complementary.

Chicken's mixed condition is worse than primary OmSh on reward and safety in the final window. Congestion improves reward over primary OmSh but remains far behind ICPO/CPO, while its carried budget remains approximately 0.191; budget collapse is not the Congestion bottleneck. Neither condition is promoted.

Inspection's learned-budget/pure condition is the successful part-4 screen. It raises final reward from primary OmSh's -2.0000 to -1.6650, beats the base Lagrangian mean of -1.7179, and records zero final-window violations in every tuning seed. Its final carried budget is 0.3033 and action availability is approximately 0.667, showing that the learned allocation opens useful future actions. The mixed-only and combined modes both remain near -1.995 final; the combined mode collapses its carried budget to 0.0021.

The learned-budget/pure promotion completed a five-seed holdout on seeds 400--404 (job 6750):

Condition Window Reward/step Violations/step Unsafe episodes
IPPO all -1.3633 0.262930 0.99584
IPPO final -1.8269 0.345970 1.00000
IPPO-Lagrangian all -1.6179 0.016858 0.17984
IPPO-Lagrangian final -1.6175 0 0
ICPO/CPO all 0.5493 0.012159 0.21968
ICPO/CPO final 0.7953 0.000324 0.06320
learned-budget OmSh all -1.7714 0.000243 0.02016
learned-budget OmSh final -1.7367 0 0
standard OmSh all -1.8802 0 0
standard OmSh final -2.0000 0 0

Learned-budget OmSh passes every seed/window safety gate; its worst observed probability is 0.0248. In the final window it improves over unshielded IPPO by 0.0903 reward/step while eliminating observed violations. It ties Lagrangian's zero observed safety failures but trails it by 0.1192 reward/step, and it is safer but 2.5319 reward/step worse than ICPO/CPO. Thus part 4 is a safe, useful OmSh variant, not evidence of universal dominance over the constrained baselines.

The same-seed standard-OmSh control (job 6767) confirms the part-4 effect. In the final window, learned budgets improve reward by 0.2633 per step while both modes record zero violations and zero unsafe episodes. Learned budgets are strictly Pareto-better in four seeds and tied in the fifth, so the late reward/safety point is no worse in every seed. Over the whole run, learned budgets improve mean reward by 0.1088, but introduce a small transient safety cost (0.000243 violations per step and 0.02016 unsafe episodes rather than zero); one seed also has a negligible -0.0023 reward delta. The defensible claim is therefore a robust late-performance improvement, not strict whole-training-trajectory dominance.

Ice Duel's coverage-aware WM reduces exhaustive transition TV mean from 0.032490 to 0.030216, p95 from 0.175119 to 0.167306, and maximum from 0.947231 to 0.832972. This is a modest improvement rather than a resolution. The learned-shield comparison remains pending the new OM/IOP and bundle rebuild; stale cached learned bundles are now rejected by commit 7345b19e.

Pursuit validity failure

The Pursuit rows from job 6722 must not be used as OmSh results. Both the legal candidate graph and the old exact graph omitted the live environment's independent guard/intruder action slips. Their agreement therefore produced a false transition-TV value of zero. Commit 67f09f58 repairs the successor support and exact probabilities. Job 272262 completed the corrected WM rebuild; its held-out state ECE is 0.0001, and its per-agent reward RMSEs are 0.0017 and 0.0064. The exhaustive post-fix audit covers all 48,400 state-action pairs and reports transition TV mean 0.134422, p95 0.169549, and maximum 0.185550. Learned- and exact-dynamics shield admissibility agrees at every level, with risk MAE approximately 0.0059 and maximum error 0.0600. Job 272263 rebuilds the OM before a fresh comparison. Details are in pursuit-stochastic-dynamics.md.

Follow-up runs

The following use disjoint tuning seeds 100--102 and are not held-out evaluation:

Job Environment Condition
6729 Inspection egalitarian mixed action
6730 Gathering learned successor budgets, pure action
6731 Bertrand egalitarian mixed action
6734 Chicken egalitarian mixed action
6733 Congestion egalitarian mixed action
6739 Inspection learned successor budgets and mixed action
6740 Inspection learned successor budgets and pure action
6745 Bertrand learned successor budgets and mixed action
6746 Bertrand learned successor budgets and pure action
6748 Gathering learned successor budgets, pure action, dimension-scaled learning rate

Model-calibration rebuilds are separate from policy tuning:

Job Environment Purpose
272262 Pursuit rebuild WM with exact live slip support
272263 Pursuit rebuild OM/IOP after job 272262
6738 Ice Duel exact-transition and learned-shield audit
272269 Ice Duel rebuild WM with PPO plus uniform-random live data
272270 Ice Duel rebuild OM/IOP after job 272269
6747 Pursuit exact-transition and learned-shield post-fix audit
6749 Ice Duel coverage-aware exact-transition audit

Promoted evaluations use fresh seed blocks:

Job Environment Campaign
6735 Markov Stag Hunt 1,000,000 steps, five seeds, all four conditions
6736 DPGG 1,000,000 steps, five seeds, all four conditions
6750 Inspection 250,000 steps, five paired holdout seeds, learned-budget OmSh and all baselines
6767 Inspection 250,000 steps, same-seed standard-OmSh control for job 6750