Skip to main content

ADR 030 — Agentic development workflows program

Context​

AlphaSwarm runs a real agentic development program: claude/* and Codex branches merge across the estate daily, ~20 repos carry best-in-class repo-specific AGENTS.md contracts, the monolith shipped 70 Cursor rules, 30 skills, and 33 subagent definitions, and the platform itself sells hash-locked spec runtimes with HITL promotion gates. Two internal research reports (2026-07) surveyed enterprise agentic-coding standards and recommended: portable guidance, reusable skills, typed tools, explicit coordination, durable state, artifact-forward promotion, and auditable controls.

Update: as of this writing, the monolith's Cursor rules have been consolidated behind a single .cursor/rules/pointer.mdc stub; the canonical rules/skills/subagents moved to alphaswarm_internal/agents/raw/alphaswarm/ (verified counts there: 66 .mdc rule files, 29 skill directories, 32 subagent definitions — close to, not identical to, the figures above, reflecting cleanup during the move). The monolith's AGENTS.md still points agents at that SSoT location.

A 49-agent cross-repo analysis (2026-07-19, all 42 working trees) found the constraint is not missing concepts — most report recommendations already exist here in stronger, ledger-backed form — but degraded execution surfaces that make agent work slow and unverifiable:

  1. The correctness oracle is broken. ~17 repos have zero CI (including alphaswarm_core, alphaswarm_config, alphaswarm_client, alphaswarm_auth); the monolith's pr-validate.yml has been YAML-invalid (duplicate python: job key) — silently disabling gitleaks, trivy, and the alembic-immutability gate at PR time; several repos' workflows trigger on branches that don't exist; enforce flags (EVAL_GATE_ENFORCE, LICENSE_GATE_ENFORCE, ALPHASWARM_WORKER_CI_FULL) were never flipped. "CI is green" does not currently mean "the change is verified." Update: the zero-CI and duplicate-job-key items (WS0.1/WS0.4/G1) have since landed — alphaswarm_core, alphaswarm_config, alphaswarm_client, and alphaswarm_auth all now ship a ci.yml, the monolith's pr-validate.yml duplicate python: job key is fixed, and only 8 repos remain with zero CI workflows. The enforce-flag item (EVAL_GATE_ENFORCE etc.) is still open — those gates remain report-only as described.
  2. Bootstrap is unreproducible. alphaswarm-core is not resolvable from any index; sibling installs use three incompatible strategies under six different PATs; 2/31 Python repos have lockfiles. A fresh agent sandbox cannot pip install -e .[dev] in roughly half the Python estate.
  3. The canon has decayed. The 146KB monolith AGENTS.md carries three conflicting hard-rule counts (45/55/actual 65), ~45 dead links, pre-split paths, and contradicts owning repos; at the time of this audit a byte-identical 704-byte .claude.md stub (a filename Claude Code never loads) was duplicated across ~35 repos while the per-tool boilerplate taught a contradictory world-model (ArcticDB/three-planes vs Iceberg/spec-runtimes). Update: the WS4 remediation for this specific item has since landed — every sibling repo now ships a real, repo-specific CLAUDE.md pointing at its own AGENTS.md (verified: 40 of the 41 repos under the workspace root carry one; alphaswarm_docs is the exception), and the old .claude.md stub is gone from the repos checked. The other WS0 findings in this list (rule-count drift, dead links, CI health) are unaffected by this and remain as described.
  4. Nothing measures whether agentic development works. The eval gate scores a fixture that is byte-equal to its goldens; no revert-rate, time-to-merge, or CI-green-on-first-push instrumentation exists for agent-authored PRs.

Externally (verified against primary sources 2026-07-19): AGENTS.md is now Linux-Foundation-stewarded with ~23 supporting tools; Agent Skills is an open standard read by Claude Code, Cursor, and Codex; worktree isolation is first-class in both Claude Code (--worktree, isolation: worktree) and Cursor (/worktree, /best-of-n); MCP's stateless 2026-07-28 revision is on a committed schedule and is backward compatible (no forced migration).

Decision​

Adopt the seven-workstream program specified in the full plan:

  • WS0 — Stop the bleeding (broken/dead CI, wrong triggers, stub cleanup) before any new machinery.
  • WS1 — Trustworthy verification: org-level reusable workflows (workflow_call), CI for the zero-CI repos in dependency order, calendar deadlines to flip every report-only gate, one GitHub App replacing the six PATs, published wheels on the existing CodeArtifact index.
  • WS2 — Uniform bootstrap/verify contract: four make targets (agent-bootstrap / agent-lint / agent-test / agent-verify) in every repo, seeded from existing AGENTS.md Validation blocks; toolchain pinning (uv.lock, .python-version, packageManager); repos.yaml machine registry in alphaswarm_index.
  • WS3 — Guidance canon restructure: split the monolith AGENTS.md to a <400-line core; slug-keyed shared-rule registry replacing numeric rule citations; mechanical guidance CI (links, rule-count, .mdc frontmatter); regenerate per-tool files as thin pointers; resolve the two-world-model fork explicitly.
  • WS4 — Multi-tool enablement on open standards: real CLAUDE.md pointers org-wide; skills mirrored to the vendor-neutral location; worktree conventions for parallel sessions; an org Claude Code plugin under managed scope; boundary lints wired into edit-time hooks.
  • WS5 — Context layer activation: wake the dormant semantic code index, persist the symbol store, CI-generate the mechanical index artifacts, publish an ownership surface, register an org-engineering KB corpus.
  • WS6 — Dogfooded dev automation: PR-review and CI-triage bots as registered AgentSpec/WorkflowSpec on the platform's own runtime, gated by the existing ApprovedPromotion HITL pattern — not a parallel stack.
  • WS7 — Measurement and experimentation: fix the eval gate, instrument agent-dev outcomes (CI-green-on-first-push, review iterations, time-to-merge, revert-within-N-days) through alphaswarm_core.observe, and gate autonomy expansion on KPI thresholds.

Sequencing, acceptance criteria, and effort classes are in the plan. The plan documents live in alphaswarm_docs (this repo); a pointer row in alphaswarm_index/index.md must be registered via the curator process (the index's sole-writer invariant is preserved, not bypassed).

Consequences​

  • Agent sessions gain a uniform done-signal (make agent-verify locally == reusable CI remotely), removing per-repo reverse-engineering.
  • Green CI becomes meaningful again before any autonomy expansion; the program's own KPI gates (metrics.md) block premature trust.
  • The guidance canon becomes mechanically checked; numeric rule citations are retired in favor of stable slugs, ending the renumbering hazard.
  • Claude Code, Cursor, Codex, and future tools load the same truth from vendor-neutral surfaces (AGENTS.md + Agent Skills + MCP), ending the Cursor monoculture without abandoning the Cursor investment.
  • Dev automation reuses the audited platform runtime (hash-locked specs, cost caps, kill-switch, decision logs) instead of growing an unaudited parallel agent stack.
  • Cost: roughly two quarters of staged platform work (see plan phasing); no rewrite — every change lands along an existing seam.

Alternatives considered​

  • Adopt an external agent-orchestration framework for dev automation — rejected: the platform's own runtime already provides hash-locked replay, budgets, HITL gates, and telemetry; a second stack would double governance surface (and both internal reports recommend single-agent-plus-tools defaults over new multi-agent frameworks).
  • Rewrite AGENTS.md from scratch — rejected: ~20 sibling repos prove the existing contract format works; the failure is drift control, not format.
  • Big-bang CI standardization — rejected in favor of dependency-ordered rollout behind reusable workflows, mirroring the ADR 015/028 migration discipline.