Skip to content

Evaluation Assumptions Requiring Review

Initial pilot scope

  • Existing WM and OM artifacts may be reused for the first policy-training pilots only when their recorded environment, public state/action, safety-label, transition, and reward contracts still match. In particular, Pursuit artifacts produced while interception rewards defaulted to zero must not supply reward-head evaluation, reward-aware shield ranking, or policy-return results after the sparse +10/-10 rewards were restored; rebuild or retrain those artifacts. Their transition coverage, calibration, and learned-versus-oracle diagnostics must still be checked before treating a poor OmSh result as an architectural failure.
  • The first fair comparison uses fresh policy initialization for IPPO, IPPO-Lagrangian, ICPO, and OmSh, with identical step budgets and seed indices. WM/OM pretraining is shared input to OmSh, not counted as policy initialization for any method.
  • The current per-environment risk budget (max_risk=0.2 in the curated notebooks) is retained for the first pass. It is below the requested 0.4 maximum and avoids changing the long-standing safety objectives before evidence warrants it.
  • Two-agent Pursuit and two-agent Congestion are evaluated before their larger-agent versions. Larger variants are not promoted unless the two-agent task is learnable and informative.

Metric interpretation

  • Reward and observed safety violations must both be reported. Shield-predicted eventual risk, immediate-risk calibration, overrides, unsafe proposals, missing graph coverage, admissible-action fraction, posterior level changes, and credible-floor posterior-tail constraint misses are diagnostics rather than substitutes for observed safety.
  • A method is not called safer solely because its configured risk allowance is lower. Comparisons use the same environment, safety label, training steps, and seeds wherever possible.
  • A safety-violation probability above 0.4 is unacceptable for promotion, even if reward is high. The analysis should also report uncertainty across seeds rather than classify from a single mean.

Variant prioritisation

  • The initial OmSh candidate is pure primitive action plus egalitarian successor-budget allocation, monotone-floor opponent handling, and Bayesian reward ranking. This is the simplest improved condition and retains the strongest current robust interpretation.
  • Bayesian-mixture, all-level, credible-floor, robust-reward, mixed-action, and learned-budget variants are deferred until the main pilots identify which conservatism or allocation bottleneck is material. This avoids spending the full ablation budget on variants that are already clearly dominated.
  • The pilots triggered the mixed-action, learned-budget, and credible-floor ablations. They remain separately labelled conditions rather than silent changes to the control, and any non-default action/budget allocator must be selected on disjoint optimisation seeds before confirmation. The final primary campaign reports OmSh-Monotone only; OmSh-Credible remains an opt-in diagnostic because, under the retained schedule, it cannot expose a larger robust feasible set than the monotone floor.
  • A later two-shielded-agent experiment will require an explicit wrapper/state/budget ownership design. It should not be approximated by silently sharing one mutable IOP or one focal-agent budget between wrappers.

Pre-declared pilot triage

  • Seeds 0--2 are screening/evaluation seeds, not hyperparameter-selection seeds. Any targeted tuning prompted by this pilot uses a disjoint block such as seed_offset=100; a selected condition must then be rerun on fresh confirmation seeds before it supports a comparative claim.
  • A condition that breaches the 0.4 observed-safety ceiling is not promoted. If predicted final risk remains within budget while observed violations breach the ceiling, investigate transition coverage, unsafe-label agreement, OM calibration, and risk calibration before loosening the budget or adding reward flexibility.
  • Missing graph coverage or infeasible carried budgets indicate a soundness/runtime defect or an inadequate abstraction. They are repaired before interpreting reward, rather than treated as an ordinary reward--safety trade-off.
  • Low admissible-action availability, budgets concentrated near their continuation lower bounds, or near-zero safety margin indicate that successor-budget allocation is binding. The next targeted variant is learned successor-budget allocation while retaining robust projection.
  • Adequate action availability but frequent unsafe proposals/overrides and improving late reward indicate policy-learning inefficiency. The first responses are longer training or mask-aware policy/critic improvements, not a less conservative safety certificate.
  • When pure primitive actions are individually too restrictive but certified mixtures produce a materially larger feasible set, test mixed-action shielding. Learned-budget pure action precedes learned-budget mixed action unless the pilot directly shows this pure-action feasibility bottleneck.
  • Only aggregate behaviour across the tuning seeds selects parameters. A particular seed, trajectory, state identity, or environment-specific action is never encoded into the algorithm.

Final-campaign review points

  • The apparent 0.15 Pursuit result on seeds 500--502 was rejected after it failed three of five seeds in the 700--704 holdout. That block then became a stress/tuning block for the lowest feasible allowances, .105 and .11. .11 dominated .105; comparative claims use the separately launched five-seed 800--804 confirmation, not either tuning mean.
  • Pursuit has a sharp seed-dependent policy/feasible-action transition between .11 and .15 under the current abstraction; .16 and higher fail as well. This is an empirical result, not a claim that .11 is a mathematically optimal budget or that untested intermediate real values cannot behave differently.
  • The .10 Pursuit target is structurally infeasible for the retained learned shield: its reset-support demand is approximately .102817, before policy optimisation. The final two-agent Pursuit condition therefore uses .11, the first clean hundredth above that demand; no numerical tolerance is used to make .10 appear feasible. Pursuit-3 is excluded for the separately documented transition-graph scaling limit.
  • The two-shielded-agent Markov Stag Hunt experiment is a paired empirical comparison. Each shield retains a certificate in its own learned induced model, but simultaneously shielding the modelled opponent changes the live policy and does not create an arbitrary-opponent joint certificate.
  • DPGG ICPO/CPO has one reproducible high-return seed in which the regulated player withholds while the unregulated co-player contributes. It remains in the reported mean. Median or trimmed comparisons may be shown as sensitivity analyses, but must not replace the pre-declared mean.
  • The final-20% window is a learning-performance diagnostic, not an infinite-horizon safety proof. Exact-anytime results apply only under the explicitly reported exact-game or induced-model assumptions; empirical finite training episodes do not strengthen those assumptions.
  • Three- and five-seed results are evidence for triage and confirmation, not definitive significance tests. Any paper-level universal superiority claim requires more seeds and a pre-declared statistical analysis.