RL statistical evaluation: the two-plane seam
AlphaSwarm deliberately carries two statistical evaluation
implementations: a dependency-light one inside alphaswarm_rl and the
full evidence-spine battery in alphaswarm.lab.evaluation. That is not
an accident of history — it is the WS2.5 posture from the agentic-RL
lifecycle plan: the RL plane keeps its own implementations; contract
tests pin them to numerical agreement on shared fixtures; this page
documents the seam.
The two planes
| RL plane | Monolith lab plane | |
|---|---|---|
| Package | alphaswarm_rl.evaluation.statistics + alphaswarm_rl.validation | alphaswarm.lab.evaluation |
| Dependencies | numpy; pandas via validation.cpcv; scipy only via DSR | stdlib DSR; arch.bootstrap soft-dep for bootstrap tests |
| DSR API | validation.deflated_sharpe.deflated_sharpe_ratio(returns, *, sr_hat, sr_list, n_strategies_tested) — derives T, skew and kurtosis from the raw returns series | deflated_sharpe_ratio(observed_sharpe, *, n_obs, n_trials, variance_of_sharpes, skewness, excess_kurtosis) — takes precomputed moments, tied to the LabRun.total_trials_searched ledger |
| PBO | validation.pbo.probability_of_backtest_overfitting (CSCV, Bailey et al. 2015) | — (PBO lives on the RL plane) |
| CPCV | validation.cpcv.CombinatorialPurgedKFold (sklearn-compatible) + combinatorial_paths_count | cpcv.combinatorial_purged_cv + CPCVConfig/CPCVPath, safe_cpcv_path_count, CPCVPlanError hard guard (100 paths) |
| Model comparison | compare_to_incumbent (bootstrap mean-difference CI) | model_comparison: diebold_mariano, whites_reality_check, hansen_spa, model_confidence_set → TestResult rows in the evidence spine |
| Consumers | RLRuntime evaluation battery, promotion/gates.py, seed sweeps | Lab sweeps, UI (never raw Sharpe alone), EvidenceBundle |
Why both exist. The lifecycle plan's gap inventory flags the duplication itself (G11: two DSR implementations with no convergence story) and resolves it with a SPLIT posture rather than a merge:
- Import isolation.
RLRuntimeis the sole sanctioned executor (hard rule 16) and must evaluate hermetically — the WS2 acceptance criterion is a green battery without mlflow or iceberg on the path (and in practice the monolith stays off it too: the convergence contract tests skip when it is absent). Importingalphaswarm.labfrom the RL evaluation path would drag the monolith (and its evidence-schema, DB andarchsurface) into every RL CI lane. - Different callers, different shapes. The RL plane starts from
raw per-seed returns arrays produced by rollouts, so its DSR derives
the moments itself. The lab plane serves sweeps and the UI, where
moments and the trial count arrive precomputed from the
LabRunledger. Forcing one signature on both callers would make one of them lie about its inputs. - Drift is handled by contract, not by sharing. The risk register entry "statistical duplication drift" is mitigated by cross-implementation contract tests on shared fixtures — see below — not by a shared library.
The convergence contract
The seam is pinned by
alphaswarm_rl/tests/validation/test_stats_convergence.py:
test_dsr_agrees_with_monolith— parameterised over(seed, n_strategies)in(0, 10),(1, 50),(2, 200)withT = 256synthetic per-period returns. One fixture computes the winning strategy's returns, its per-period Sharpe and the cross-trial Sharpe list; each plane is fed its native parameterisation (raw returns for the RL plane, precomputed variance/skew/kurtosis for the monolith). Agreement is asserted toabs=1e-4— the residual is the monolith's Beasley-Springer-Moro inverse CDF vs the RL plane'sscipy.stats.norm.ppf.test_psr_agrees_when_single_trial— with one trial the deflation term vanishes and the monolith DSR degrades to itsprobabilistic_sharpe_ratio; the test pins the shared PSR core (Bailey-López de Prado eq. 9 variance) toabs=1e-6.
The module uses pytest.importorskip on scipy and
alphaswarm.lab.evaluation.deflated_sharpe, so the contract runs in
lanes where both planes are importable and skips cleanly in standalone
RL-plane CI. If you add a statistic that exists on both planes, you
must extend this contract file — an unpinned duplicate is the exact
failure mode the seam exists to prevent.
The evaluation battery surface
The battery is configured on the spec via EvaluationConfig
(alphaswarm_rl/spec.py). Every lifecycle knob defaults to
None/absent so existing spec snapshot_hash values are unchanged
(hard rule 17); setting one creates a new immutable spec version.
evaluation:
episodes: 4
n_seeds: 5 # or an explicit seed_list: [7, 11, 13, ...]
bootstrap: {n_resamples: 1000, alpha: 0.05} # optional seed: for CI reproducibility
compute_dsr: true
dsr_n_strategies: 24 # search-space size; defaults to the seed count
n_seeds/seed_list—seed_listwins when both are set;n_seedsderives consecutive offsets from the resolvedtraining.seed.Nonekeeps the legacy single-pass rollout.bootstrap—{n_resamples, alpha, seed}consumed byevaluation.statistics.aggregate_metrics(defaults: 1000 resamples,alpha = 0.05).compute_dsr/dsr_n_strategies— enable the Deflated Sharpe probability over the seed battery and optionally widen the deflation to the true search-space size.regime_slices— the newest addition (WS2 item 3): opt-in regime-sliced aggregates in the evaluation report, so the battery can answer "does this hold in the stress regime?" rather than only "does this hold on average?". Default-off like every other knob, so it is hash-stable for existing specs.
RLRuntime._do_evaluate runs one rollout batch per seed (global RNG +
env.reset(seed=...) per seed, per-period portfolio returns captured
per seed) and hands the rows to
evaluation.statistics.build_evaluation_report, which produces:
{
"per_seed": [{"seed": 11, "metrics": {"sharpe": 0.61, "...": "..."}}],
"aggregate": {"sharpe": {"point": 0.58, "iqm": 0.57, "lo": 0.41, "hi": 0.74, "n": 5}},
"n_seeds": 5,
"dsr_probability": 0.97,
"pbo": 0.12
}
Bootstrap CI + IQM semantics
aggregate_metrics summarises each metric across seeds as
point (mean), iqm, lo/hi (percentile bootstrap CI of the mean)
and n. Two deliberate degeneracy rules keep gate checks well-defined
on small batteries: a single sample collapses the CI to a point
interval, and iqm (the rliable-style interquartile mean, robust to
outlier seeds) falls back to the plain mean below four samples.
DSR per-period discipline
Every Sharpe handed to the DSR is unannualised (per-period) —
statistics.per_period_sharpe per seed, never the annualised headline
number. Passing an annualised Sharpe (SR × √252) is the canonical DSR
bug and yields nonsensical probabilities; the warning is baked into
validation/deflated_sharpe.py itself. The battery treats the
best-Sharpe seed as the "winning strategy", the per-seed Sharpe list as
the trial set, and dsr_n_strategies (default: the seed count) as the
search-space size N for the deflation.
PBO over the seed-returns matrix
When at least two seeds produced usable returns series, the battery
stacks them column-wise (truncated to the shortest series) and runs
probability_of_backtest_overfitting — CSCV with up to 16 blocks —
reporting the fraction of IS/OOS splits where the in-sample best seed
ranks below median out-of-sample. Series shorter than four periods, or
a single-seed battery, yield pbo: null rather than a fabricated
number.
Challenger comparison
statistics.compare_to_incumbent bootstraps the mean difference
(candidate − incumbent) and reports separated: true only when the CI
lower bound clears zero. This is the WS7 approval discipline: a MARL
(or any other) challenger must show a statistically separated
improvement, not a point-estimate win.
From evidence to gates
The report is exactly the evidence surface that
alphaswarm_rl.promotion.gates.evaluate_gates(report, config) consumes
(see RL lifecycle: statistical gates for the
full gate catalogue):
- scalar thresholds (
min_sharpe,min_sortino,min_total_return,max_drawdown) readaggregate.<metric>.point; require_ci_positivereadsaggregate.total_return.lo— the bootstrap lower bound, not the mean;max_pboandmin_dsr_probabilityread the report's top-levelpboanddsr_probability;min_ope_lower_boundreadsope.headline.lofor offline candidates.
Fail-closed doctrine: a configured threshold whose evidence is
missing from the report — battery not run, pbo: null, NaN, absent
key — is a FAIL with reason missing_evidence, never a silent pass.
This is why the statistics layer prefers null over a fabricated value
on degenerate inputs: null evidence rejects at the gate, which is
the correct outcome for an under-powered battery. A failing gate blocks
mlflow.register_model and the outcome persists in
result_summary.promotion with per-check
observed / threshold / reason.
Which plane to use when
- Inside the RL lifecycle (the
RLRuntimebattery, promotion gates, seed sweeps, OPE): the RL plane, always. Lifecycle logic lives inRLRuntimeor pure helpers it calls (hard rule 16), and those helpers must not import the monolith. - Lab sweeps, the UI, the evidence spine: the monolith plane —
DSR/PSR rendered next to raw Sharpe from the
LabRuntrial ledger, CPCV path planning behind theCPCVPlanErrorhard guard, and the model-comparison family (Diebold-Mariano, White's Reality Check, Hansen's SPA, Model Confidence Set) emittingTestResultrows. - Cross-plane comparisons (e.g. an RL candidate vs a lab-selected
incumbent): compute each side's evidence on its own plane; compare
with
compare_to_incumbentor the monolith model-comparison tests. The convergence contract is what makes the DSR numbers commensurable across planes. - Never add a runtime import of
alphaswarm.lab.evaluationtoalphaswarm_rlevaluation paths; the contract-test lane is the only sanctioned coupling point.
See also: rl-lifecycle-gates, rl-prudex-evaluation, rl-framework.