eval_250k_v1 results¶
This is the live result record for the 250,000-step, three-seed screen. Each
comparison cell is reward / violations per step / unsafe episodes; +, =,
and - mean that OmSh is better, tied, or worse than the named baseline on the
paired-seed mean. The final window contains episodes ending in the last 20% of
each run. The empirical gate separately rejects any OmSh seed/window with
per-step violation probability or unsafe-episode probability above 0.4.
The runs are screening evidence rather than significance claims. The primary OmSh condition is pure primitive actions, robust egalitarian successor budgets, a monotone opponent-model floor, and Bayesian expected reward.
| Environment | Window | OmSh reward/step | Violations/step | Unsafe episodes | vs IPPO | vs Lagrangian | vs ICPO/CPO | Gate |
|---|---|---|---|---|---|---|---|---|
| Gathering | all | 0.022923 | 0.0015093 | 0.076667 | - / + / + |
+ / + / + |
+ / + / + |
pass |
| Gathering | final | 0.034653 | 0.0000133 | 0.006667 | - / + / + |
+ / + / + |
+ / + / + |
pass |
| Ice Duel | all | 0.060073 | 0.0089773 | 0.022586 | - / + / + |
+ / + / + |
- / - / - |
pass |
| Ice Duel | final | 0.059200 | 0.0124533 | 0.020095 | - / + / + |
+ / - / + |
- / - / - |
pass |
| Markov Stag Hunt | all | 0.315317 | 0 | 0 | + / + / + |
+ / + / + |
+ / + / + |
pass |
| Markov Stag Hunt | final | 0.470000 | 0 | 0 | + / + / + |
+ / + / + |
+ / + / + |
pass |
| Pursuit | all | invalid | invalid | invalid | invalid | invalid | invalid | invalid |
| Pursuit | final | invalid | invalid | invalid | invalid | invalid | invalid | invalid |
| Bertrand | all | 0.780207 | 0 | 0 | - / + / + |
- / + / + |
- / + / + |
pass |
| Bertrand | final | 0.197500 | 0 | 0 | - / + / + |
- / + / + |
+ / + / + |
pass |
| Chicken | all | 1.106285 | 0 | 0 | - / + / + |
- / + / + |
- / + / + |
pass |
| Chicken | final | 1.013427 | 0 | 0 | - / + / + |
- / + / + |
+ / + / + |
pass |
| Congestion | all | -6.000000 | 0 | 0 | - / + / + |
- / + / + |
- / + / + |
pass |
| Congestion | final | -6.000000 | 0 | 0 | - / + / + |
- / + / + |
- / = / = |
pass |
| DPGG | all | 0.312409 | 0 | 0 | + / + / + |
+ / + / + |
- / + / + |
pass |
| DPGG | final | 0.181055 | 0 | 0 | + / + / + |
+ / = / = |
+ / + / + |
pass |
| Inspection | all | -1.878972 | 0 | 0 | - / + / + |
- / + / + |
- / + / + |
pass |
| Inspection | final | -2.000000 | 0 | 0 | - / + / + |
- / + / + |
- / = / = |
pass |
Current interpretation¶
- Markov Stag Hunt is the only completed environment where the primary OmSh condition strictly improves reward and both observed safety measures over all three baselines in both windows. Its worst OmSh safety probability is zero.
- DPGG Pareto-dominates all three baselines in the final window. Its overall reward trails ICPO because ICPO earns high early reward and then collapses; OmSh's late reward is higher than every baseline with zero observed failures.
- Gathering beats both constrained baselines on reward and safety but trails
unsafe IPPO on reward. Its late unsafe-episode mean is
0.0067, while the carried budget falls to roughly0.014--0.027, making successor-budget allocation the targeted bottleneck. - Ice Duel passes the absolute safety ceiling but trails ICPO on reward and both
observed safety measures. Its late action availability is
0.994, its mean carried budget remains0.1996, and its mean predicted final risk is only0.0017despite0.0125observed violations per step. This is a transition or risk-calibration problem, not evidence that mixed actions or learned budgets are the next intervention. Job 6738 builds the exact transition graph and exports learned-versus-exact diagnostics before further tuning. - Bertrand and Chicken beat ICPO late on reward and safety, but retain small reward gaps to IPPO-Lagrangian and/or IPPO. Congestion and Inspection remain substantially reward-limited. These four matrix games expose only one safe pure action (one half of actions, or one third in Congestion) while leaving the configured risk budget unused, motivating mixed primitive actions.
Tuning results¶
These rows use tuning seeds rather than the base campaign's seeds, so their means are unpaired screens. They are accepted only if every seed in both the whole-run and final-window summaries passes the empirical safety ceiling.
| Environment | Condition | Seeds | Window | Reward/step | Violations/step | Unsafe episodes | Worst seed/window | Gate | Decision |
|---|---|---|---|---|---|---|---|---|---|
| Gathering | learned budget, pure | 100--102 | all | 0.008759 | 0.0002893 | 0.088667 | 0.1160 | pass | reject |
| Gathering | learned budget, pure | 100--102 | final | 0.013587 | 0.0000267 | 0.013333 | 0.1160 | pass | reject |
| Bertrand | egalitarian mixed | 100--102 | all | 1.437876 | 0.138136 | 0.255733 | 0.3240 | pass | combine with learned budgets |
| Bertrand | egalitarian mixed | 100--102 | final | 0.924740 | 0.213067 | 0.310667 | 0.3240 | pass | combine with learned budgets |
| Chicken | egalitarian mixed | 100--102 | all | 1.129864 | 0.0817573 | 0.285067 | 0.3600 | pass | reject |
| Chicken | egalitarian mixed | 100--102 | final | 0.995687 | 0.046060 | 0.336000 | 0.3600 | pass | reject |
| Congestion | egalitarian mixed | 100--102 | all | -5.491392 | 0.0764093 | 0.194933 | 0.1976 | pass | reject |
| Congestion | egalitarian mixed | 100--102 | final | -5.463107 | 0.077360 | 0.190667 | 0.1976 | pass | reject |
| Inspection | egalitarian mixed | 100--102 | all | -1.868756 | 0.004256 | 0.029067 | 0.0296 | pass | combine with learned budgets |
| Inspection | egalitarian mixed | 100--102 | final | -1.995013 | 0 | 0 | 0.0296 | pass | combine with learned budgets |
| Inspection | learned budget, pure | 300--302 | all | -1.730252 | 0.0001573 | 0.017333 | 0.0216 | pass | promote |
| Inspection | learned budget, pure | 300--302 | final | -1.665047 | 0 | 0 | 0.0216 | pass | promote |
| Inspection | learned budget, mixed | 300--302 | all | -1.876732 | 0.0003240 | 0.033333 | 0.0424 | pass | reject |
| Inspection | learned budget, mixed | 300--302 | final | -1.994587 | 0 | 0 | 0.0424 | pass | reject |
| Bertrand | learned budget, mixed | 300--302 | all | 0.822361 | 0.0075827 | 0.249067 | 0.2920 | pass | reject |
| Bertrand | learned budget, mixed | 300--302 | final | 0.247867 | 0.0149333 | 0.272000 | 0.2920 | pass | reject |
| Bertrand | learned budget, pure | 300--302 | all | 1.072192 | 0.0095400 | 0.228267 | 0.6720 | fail | reject |
| Bertrand | learned budget, pure | 300--302 | final | 0.735293 | 0.0081867 | 0.229333 | 0.6720 | fail | reject |
| Gathering | learned budget, pure, LR 3e-5 |
300--302 | all | 0.004925 | 0.0011053 | 0.280000 | 0.2840 | pass | reject |
| Gathering | learned budget, pure, LR 3e-5 |
300--302 | final | 0.005160 | 0.0009600 | 0.246667 | 0.2840 | pass | reject |
Gathering's learned-budget policy is dominated by primary OmSh in the final
window: it earns 0.0136 rather than 0.0347, and its carried budget falls
further to approximately 0.0117. The current high-dimensional projected PPO
action did not learn the intended high-value allocation. The checkpoint has
297 continuous slots and averages 58.0 active dimensions; joint PPO KL
averages 0.0705 and reaches 11.29, compared with mean 0.00094 and maximum
0.0114 for Bertrand's one-dimensional mixed action. Reducing the learning
rate from 2e-4 to 3e-5 makes the update conservative (mean KL 0.000113,
maximum 0.000895) but worsens final reward to 0.0052 and unsafe episodes to
0.2467. It is dominated by both the primary OmSh condition and the first
learned-budget screen. The remaining Gathering problem is not fixed by scalar
step-size scaling; high-dimensional credit assignment and exploration require
a different treatment.
Bertrand's mixed-only policy beats every baseline on reward, but it does not
beat the constrained baselines on safety. Its final violation probability is
0.2131 and unsafe-episode probability is 0.3107. The common-seed 2-by-2
follow-up does not justify promotion. Combining mixed actions with learned
budgets passes the ceiling in every seed/window, but falls to 0.2479 final
reward and 0.2720 unsafe episodes. Learned budgets with pure actions
bifurcate: one seed reaches 1.8239 final reward with 0.6720 unsafe episodes,
while the other two seeds remain near 0.19 reward with at most 0.0160
unsafe episodes. Consequently the pure arm fails the hard gate and the
combined arm is dominated by the mixed-only screen. This is evidence that the
two extensions should remain independently selectable rather than being
assumed complementary.
Chicken's mixed condition is worse than primary OmSh on reward and safety in
the final window. Congestion improves reward over primary OmSh but remains far
behind ICPO/CPO, while its carried budget remains approximately 0.191; budget
collapse is not the Congestion bottleneck. Neither condition is promoted.
Inspection's learned-budget/pure condition is the successful part-4 screen. It
raises final reward from primary OmSh's -2.0000 to -1.6650, beats the base
Lagrangian mean of -1.7179, and records zero final-window violations in every
tuning seed. Its final carried budget is 0.3033 and action availability is
approximately 0.667, showing that the learned allocation opens useful future
actions. The mixed-only and combined modes both remain near -1.995 final;
the combined mode collapses its carried budget to 0.0021.
The learned-budget/pure promotion completed a five-seed holdout on seeds 400--404 (job 6750):
| Condition | Window | Reward/step | Violations/step | Unsafe episodes |
|---|---|---|---|---|
| IPPO | all | -1.3633 | 0.262930 | 0.99584 |
| IPPO | final | -1.8269 | 0.345970 | 1.00000 |
| IPPO-Lagrangian | all | -1.6179 | 0.016858 | 0.17984 |
| IPPO-Lagrangian | final | -1.6175 | 0 | 0 |
| ICPO/CPO | all | 0.5493 | 0.012159 | 0.21968 |
| ICPO/CPO | final | 0.7953 | 0.000324 | 0.06320 |
| learned-budget OmSh | all | -1.7714 | 0.000243 | 0.02016 |
| learned-budget OmSh | final | -1.7367 | 0 | 0 |
| standard OmSh | all | -1.8802 | 0 | 0 |
| standard OmSh | final | -2.0000 | 0 | 0 |
Learned-budget OmSh passes every seed/window safety gate; its worst observed
probability is 0.0248. In the final window it improves over unshielded IPPO
by 0.0903 reward/step while eliminating observed violations. It ties
Lagrangian's zero observed safety failures but trails it by 0.1192
reward/step, and it is safer but 2.5319 reward/step worse than ICPO/CPO. Thus
part 4 is a safe, useful OmSh variant, not evidence of universal dominance over
the constrained baselines.
The same-seed standard-OmSh control (job 6767) confirms the part-4 effect. In
the final window, learned budgets improve reward by 0.2633 per step while
both modes record zero violations and zero unsafe episodes. Learned budgets are
strictly Pareto-better in four seeds and tied in the fifth, so the late
reward/safety point is no worse in every seed. Over the whole run, learned
budgets improve mean reward by 0.1088, but introduce a small transient safety
cost (0.000243 violations per step and 0.02016 unsafe episodes rather than
zero); one seed also has a negligible -0.0023 reward delta. The defensible
claim is therefore a robust late-performance improvement, not strict
whole-training-trajectory dominance.
Ice Duel's coverage-aware WM reduces exhaustive transition TV mean from
0.032490 to 0.030216, p95 from 0.175119 to 0.167306, and maximum from
0.947231 to 0.832972. This is a modest improvement rather than a resolution.
The learned-shield comparison remains pending the new OM/IOP and bundle rebuild;
stale cached learned bundles are now rejected by commit 7345b19e.
Pursuit validity failure¶
The Pursuit rows from job 6722 must not be used as OmSh results. Both the legal
candidate graph and the old exact graph omitted the live environment's
independent guard/intruder action slips. Their agreement therefore produced a
false transition-TV value of zero. Commit 67f09f58 repairs the successor
support and exact probabilities. Job 272262 completed the corrected WM rebuild;
its held-out state ECE is 0.0001, and its per-agent reward RMSEs are 0.0017
and 0.0064. The exhaustive post-fix audit covers all 48,400 state-action
pairs and reports transition TV mean 0.134422, p95 0.169549, and maximum
0.185550. Learned- and exact-dynamics shield admissibility agrees at every
level, with risk MAE approximately 0.0059 and maximum error 0.0600. Job
272263 rebuilds the OM before a fresh comparison. Details are in
pursuit-stochastic-dynamics.md.
Follow-up runs¶
The following use disjoint tuning seeds 100--102 and are not held-out evaluation:
| Job | Environment | Condition |
|---|---|---|
| 6729 | Inspection | egalitarian mixed action |
| 6730 | Gathering | learned successor budgets, pure action |
| 6731 | Bertrand | egalitarian mixed action |
| 6734 | Chicken | egalitarian mixed action |
| 6733 | Congestion | egalitarian mixed action |
| 6739 | Inspection | learned successor budgets and mixed action |
| 6740 | Inspection | learned successor budgets and pure action |
| 6745 | Bertrand | learned successor budgets and mixed action |
| 6746 | Bertrand | learned successor budgets and pure action |
| 6748 | Gathering | learned successor budgets, pure action, dimension-scaled learning rate |
Model-calibration rebuilds are separate from policy tuning:
| Job | Environment | Purpose |
|---|---|---|
| 272262 | Pursuit | rebuild WM with exact live slip support |
| 272263 | Pursuit | rebuild OM/IOP after job 272262 |
| 6738 | Ice Duel | exact-transition and learned-shield audit |
| 272269 | Ice Duel | rebuild WM with PPO plus uniform-random live data |
| 272270 | Ice Duel | rebuild OM/IOP after job 272269 |
| 6747 | Pursuit | exact-transition and learned-shield post-fix audit |
| 6749 | Ice Duel | coverage-aware exact-transition audit |
Promoted evaluations use fresh seed blocks:
| Job | Environment | Campaign |
|---|---|---|
| 6735 | Markov Stag Hunt | 1,000,000 steps, five seeds, all four conditions |
| 6736 | DPGG | 1,000,000 steps, five seeds, all four conditions |
| 6750 | Inspection | 250,000 steps, five paired holdout seeds, learned-budget OmSh and all baselines |
| 6767 | Inspection | 250,000 steps, same-seed standard-OmSh control for job 6750 |