Saltar al contenido principal

Measurement and experimentation program (WS7)

Part of the agentic workflows enhancement plan. Closes gaps.md G9. Principle (from the internal research reports, adopted): run the agentic-dev program as an experimentation system, not a sequence of anecdotes — and never expand autonomy past a stage whose KPIs aren't green.

1. Fix the inert eval loop first​

The gate must mean something before anything is gated on it:

  1. (S) alphaswarm_eval/ci.yml trigger main → development (WS0.2 — done in P0).
  2. (M) Add a producer job: nightly, budget-capped runs of the real NL→query / NL→dashboard models over the goldens — replacing the current fixture whose outputs are byte-equal to the goldens (per the repo's own PROVENANCE.md, it can never fail). Commit scores as the rolling baseline.
  3. (S) Flip EVAL_GATE_ENFORCE=true after two weeks of stable baselines (calendar-tracked in the §5 review).
  4. (M) One eval engine: alphaswarm_mlops depends on the alphaswarm-eval package; delete the diverged alphaswarm_mlops/packages/alphaswarm_eval copy. Port alphaswarm_agents/evaluation.py's replay+judge harness onto alphaswarm_eval's Scorer/EvalRunner contract so agent evals emit the same EVAL spans and feed the same baselines.
  5. (M) Harden the LLM judge before trusting it: version the judge prompt as a registered scorer (judge@<version>), pin the judge model, add a calibration golden set (known-good/known-bad pairs with expected score bands) run in CI, and make judge unavailability loud (fail the case or alert), replacing the silent try/except: pass blocks.
  6. (M) De-brittle goldens: execution-based SQL equivalence against a fixture DB (not string match); expand nl_to_query beyond 22 cases; trajectory + safety scorers over the agent_run_steps trace shape that trace_from_run_steps already parses.

2. Baseline instrumentation of agent-authored PRs (starts in P0)​

Agent-authored PRs are already identifiable (claude/*, codex/*, harden/* branch prefixes; Codex automerge committer). Emit per-PR events as spans through alphaswarm_core.observe (same trace store as everything else; W3C traceparent; ingest-time redaction applies):

EventAttributes (minimum)
dev.pr.openedrepo, pr_number, actor_type (human/claude/codex/cursor), branch, task_ref
dev.pr.ci_first_resultgreen_on_first_push (bool), failing_checks[]
dev.pr.review_iterationiteration_n, reviewer_type (human/bot)
dev.pr.mergedtime_to_merge_s, review_iterations, additions, deletions
dev.pr.reverteddays_since_merge, revert_pr, linked_incident
dev.workflow.bot_runworkflow_spec_version, verdict, cost_usd, human_agreement (shadow mode)

Stable IDs (task_ref, repo, commit_sha, spec_version_id where a WS6 bot is involved) tie dev telemetry to the same ledger discipline as product runs. Note: the OTel GenAI semantic-conventions stability claim did not survive external verification — we key on our own span schema (above), and map to external conventions later if/when they stabilize.

3. KPIs and readiness thresholds​

Weekly dashboard (per repo and org-rollup), segmented actor_type:

KPIDefinitionReadiness threshold (to expand autonomy a stage, per WS6 rollout model)
Task success ratemerged without human rewrite / agent PRs openedStable or improving over 4 weeks, no SRM in any live experiment
CI-green-on-first-pushfirst CI result green / agent PRsImproving; no repo below 50% after WS1/WS2 land there
Median time-to-mergeopen → merge, agent PRsImprovement or neutral vs. baseline
Review burdenreview iterations + human rework minutes per agent PRDownward or stable after calibration
Revert rateagent PRs reverted ≤14 days / mergedRelease-limiting guardrail: no meaningful degradation vs. human baseline
Policy violationsboundary-lint failures post-merge; approval-bypass attempts on ApprovedPromotion pathsZero high-severity; zero bypasses
Gate integrityenforce-flags on; quarantine-list countMonotone shrinking quarantine counts; all flags flipped by their P1/P2 deadlines
Cost$ per merged agent PR (model + sandbox/CI minutes)Within budget envelope; optimize cost-per-success, never raw tokens
Bot quality (WS6)shadow-mode agreement with human review verdicts; bot regression suite pass rate≥ agreed threshold before limited-write stage; suite blocking

These thresholds are deliberately strict: early agent platforms fail on trust before they fail on capability.

4. Experimentation discipline​

For any change to guidance, skills, CI gates, or bot behavior where we want a causal answer (adopted from the reports; they survived as methodology even where their external evidence citations did not):

  • Randomization unit matches the interference locus: repo or team for guidance/skills changes (shared context contaminates within a repo); PR for CI-gate changes; developer for per-user IDE behaviors.
  • Pre-register the primary metric (usually task success rate or time-to-merge) before enabling; everything else is diagnostic.
  • SRM check on every experiment (chi-square on assignment counts): any significant sample-ratio mismatch freezes interpretation until root-caused.
  • CUPED variance reduction using pre-period covariates we already have per repo/developer (historical PR cycle time, review-turn count, CI failure probability).
  • Valid sequential monitoring — no informal daily peeking with fixed-horizon p-values; fixed analysis points are acceptable at our scale.
  • The alphaswarm_eval A/B machinery does the same job for bots: two agent configs over the same golden suite via EvalRunner, compared with the already-written Diebold-Mariano/SPA scorers (ALPHASWARM_EVAL_STAT_SCORERS_ENABLED), with cbs_from_eval_reports for cost-aware promote/collapse decisions.

Realistic caveat on power: with one small team, many comparisons will be underpowered for formal significance. The discipline still pays — SRM and guardrails catch broken instrumentation and regressions even when the primary metric can't reach significance; treat sub-powered results as directional and lean on the KPI trend lines.

5. Test-infrastructure ratchets (feeds gate integrity)​

Per the org's own TESTING_FRAMEWORK_BLUEPRINT.md (D6) — execute, don't redesign:

  • Flaky-test analytics (pytest-rerunfailures or span-based pass/fail-window detection); tracked quarantine list with owners; the existing --ignore lists and continue-on-error slices convert to ticketed quarantine entries with a shrinking count enforced in CI.
  • Ratcheted coverage floors (--cov-fail-under starting at current per-package baselines); registered + strictly-enforced pytest markers so suites slice deterministically.
  • Kick off alphaswarm_testkit Phase 1 per the blueprint.

6. Cadence and ownership​

  • Weekly: dashboard review in the platform channel; new quarantine entries and enforce-flag deadlines checked; experiment readouts (with SRM status) recorded as dated notes in this directory.
  • Monthly: KPI-vs-threshold review decides autonomy stage changes (WS6 rollout model) and updates last_reviewed here.
  • Quarterly: re-run the cross-repo audit (the 2026-07-19 analysis is reproducible); refresh gaps.md; retire completed workstreams into the org-audit archive.
  • Owner: platform-team; the dashboard and span schema live with alphaswarm_core.observe; eval baselines live with alphaswarm_eval.