Saltar al contenido principal

RL lifecycle: statistical gates & paper release train

The agentic-RL lifecycle plan (P1–P3) turns AlphaSwarm's existing statistical machinery into an enforced research → evaluation → paper lifecycle. Everything below lives inside the sanctioned executor, RLRuntime — no new service, no second execution path.

Design doctrine​

  1. Fail closed. A configured threshold whose evidence is missing is a REJECT (missing_evidence), never a silent pass.
  2. Opt-in, hash-stable. All new sub-specs (promotion, paper, evaluation seed battery) default to None/off. snapshot_hash() excludes None-valued model fields, so adding optional schema fields never shifts the hash of existing specs; setting one creates a new immutable version.
  3. Policy proposes, controls dispose. The learned policy ends at target weights; WeightToOrders + the monolith risk plane own everything after.

Seeded evaluation (WS1/WS2)​

training.seed now seeds Python/NumPy/torch and the first env.reset(seed=...) of each rollout. The evaluation battery is configured on the spec:

training:
seed: 7
evaluation:
episodes: 4
n_seeds: 5 # or an explicit seed_list: [7, 11, 13, ...]
bootstrap: {n_resamples: 1000, alpha: 0.05}
compute_dsr: true

evaluate (and the post-train battery) returns an evaluation_report: per-seed metric rows plus a per-metric aggregate (point / iqm / lo / hi / n from a percentile bootstrap), a Deflated Sharpe probability over the seed battery (per-period Sharpe discipline), and PBO over the seed-returns matrix when computable. Per-term reward_terms totals are summarised into result_summary.reward_attribution.

Promotion gates (WS3)​

promotion:
min_sharpe: 0.5
max_pbo: 0.5
min_dsr_probability: 0.90
min_seeds: 5
require_ci_positive: true # bootstrap CI lower bound of total_return > 0
mlflow:
register_model_as: rl-ppo-portfolio

With promotion set, mlflow.register_model fires only when every gate passes; outcomes persist in result_summary.promotion with per-check observed / threshold / reason, named to match the alphaswarm_mlops.factory.eval_acceptance surface. promotion: null preserves the legacy unconditional registration path byte-for-byte.

Paper release train (WS4)​

paper: null keeps the legacy single-rollout paper probe. Setting paper upgrades RLRuntime.paper() to a bounded release-train session:

paper:
max_episodes: 20
duration_seconds: 86400
# acceptance criteria (evaluated after the session, fail-closed)
max_hard_breaches: 0
max_reject_rate: 0.05
min_total_return: 0.0
max_drawdown: 0.10
# rollback triggers (evaluated between episodes)
rollback_max_drawdown: 0.15
rollback_max_loss_pct: 0.10
approval_expiry_days: 30

Session flow​

  1. Episodes run one at a time; the run's Redis halt key is honoured between episodes and every 50 steps inside one.
  2. After each episode the rollback triggers are checked. A breach engages the run's halt key (rollback:<trigger>), stops the session, and finalises the run as rolled_back — fail-closed, no operator in the loop.
  3. After the session the acceptance battery runs: no-rollback, hard-breach budget, reject-rate, return and drawdown criteria. Criteria that need gateway counters the env doesn't expose fail with missing_evidence.
  4. Everything lands in result_summary.paper_pack: spec hash, checkpoint, seed, config, per-episode metrics, session summary, operational report, acceptance verdict + criteria, rollback events, and the approval expiry timestamp.

Operational counters​

Envs that route orders through the WS5 gateway may expose hard_breach_count, pretrade_reject_count and order_submit_count; the release train reads them duck-typed. Absent counters leave the corresponding evidence None — and any acceptance criterion that depends on it rejects.

Pre-trade gateway parity (WS5)​

WeightToOrders now accepts a risk_manager (canonically alphaswarm.risk.manager.RiskManager, or "auto" to best-effort build one). Every order is checked through check_pretrade_v2 before submission:

  • breaches with severity block/critical reject that order with structured reason codes (WeightToOrdersResult.rejected);
  • a crash of the check itself rejects fail-closed (pretrade_gateway_error) — an order can never pass un-gated;
  • every emission (accepted or rejected) produces a decision record: client order id, symbol, side, quantity, notional, target weight, the checks applied, and a deterministic inputs_hash over (weights, prices, equity) for audit correlation.

The kill switch remains the outermost gate: engagement aborts the whole batch before any per-order logic runs.

Run statuses​

running | completed | error | timeout | halted | cancelled | rolled_back

rolled_back dominates halted when the halt was engaged by a paper rollback trigger.

Observability (WS9)​

Every lifecycle run opens an rl span (soft dependency on the observe SDK) carrying alphaswarm.rl.run_id / spec_hash / seed / target at open and run_status / checkpoint_id / gate outcome at close. alphaswarm_observe derives rl_run_success_rate and rl_gate_pass_rate SLOs; RL spans are excluded from the generic latency objective (lifecycle runs are hours-long by design).

Drill procedure​

Before relying on the release train in anger, run the rollback drill:

  1. Point a spec at a deliberately lossy window (or a crash-env fixture) with rollback_max_drawdown set below the expected drawdown.
  2. Run RLRuntime.paper() and confirm: status rolled_back, a rollback_events entry naming the trigger, the halt key set to rollback:<trigger>, and acceptance.passed == false.
  3. Confirm the rl span for the run carries alphaswarm.rl.run_status = "rolled_back".

See also: rl-framework, weight-centric-pipeline, rl-prudex-evaluation.