Current-state gap analysis
Part of the agentic workflows enhancement plan. Gaps here are closed by the workstreams in workstreams.md; measurement is defined in metrics.md. Evidence cites working-tree state as of 2026-07-19.
1. Strengths — build on these, do not rebuild them
The plan's first rule is that the following assets are the templates, not the debt:
- The repo-specific AGENTS.md contract pattern is best-in-class in ~20
repos and should be propagated as-is:
alphaswarm_api/AGENTS.md(363 lines, "AGENTS.md is the law"),alphaswarm_ui/AGENTS.md(242 lines, 12 numbered security boundaries incl. CVE pinning), and thealphaswarm_finops/_worker/_kb/_rl/_data/_mlopsfamily with "Hard Boundaries" + "Where Changes Go" routing tables + Validation blocks. - Machine-enforced boundaries already exist as committed scripts:
check_no_alphaswarm_imports.py(data, kb_federation, local, learning, mcp),check_plane_boundary.pyAST check (worker, blocking in CI), the monolith's ~28scripts/ci/lints, alembic migration-immutability hash-lock,buf breaking(core), CRD schema drift enforced three ways (bots, ADR 006),oasdiff/griffeAPI-surface gates. - Contract lock-step patterns worth copying verbatim:
alphaswarm_cli'sgenerate-docs --agent-contract --check(docs/agent-contract.json kept in CI lock-step with code); the monolith's "AGENTS hard-rule lints" CI job;alphaswarm_mlops' sha256-lockedfrozen_contractsregistry (186 contracts) with import-boundary checker. - Supply-chain maturity exists where wired: cosign keyless + SLSA L3
(
slsa-github-generatorv2.1.0) + syft SBOM attestation inalphaswarm/.github/workflows/build-publish.yml;alphaswarm_platform'sbuild-sign-pushcomposite blocks on Trivy HIGH/CRITICAL and is SOC2-referenced; OIDC-first cloud auth via a dedicatedaws-oidc-assumecomposite; the governed One Gate Terraform apply with mandatory break-glass reasons (ADR 022). alphaswarm_qapis the reproducible-bootstrap exemplar: uv workspace, committeduv.lock,.python-version, PEP-735 dependency-groups, every gate runnable viauv run, fail-closed attestation gates.- The platform's own agent runtime is dogfood-ready: hash-locked
AgentSpec/WorkflowRuntimewith guardrail budgets and kill-switch polling; seven orchestration adapters including debate topologies; theApprovedPromotionHITL capability-token pattern (alphaswarm_agents/src/alphaswarm_agents/promotion_gate.py); evidence-ref-enforcedFindingRecorder; decision-log-with-reflection memory; shipped dev-facing specs (configs/agents/codebase_assistant.yaml,codebase_refactorer.yaml) bound to a working Codebase MCP. - Context-layer infrastructure exists and is governed:
alphaswarm_mcpserves rules/skills/AGENTS.md bundles with Claude Code.mcp.jsonwiring;alphaswarm_indexhas a sole-writer contract, cite-or-TODO discipline, and per-area token budgets;alphaswarm_kbprovides fail-closed permissioned memory with bi-temporal envelopes; the federation gateway handles cross-silo sharing. - Deliberately offline-testable repos are exactly what sandboxed agents
need:
alphaswarm_data(lazy SDKs, injectable transports, 39 offline test files),alphaswarm_finops(moto/MockTransport),alphaswarm_eval(zero required runtime deps),alphaswarm_observe(graceful degrade). - An honest engineering culture encoded in CI: dated known-red quarantine
lists with burn-down instructions, "NEVER add
|| true" comments, green-by-skip called out and fixed loudly in the 2026-07 monolith CI rewrite,alphaswarm_ui's 36-hour-outage postmortem inline inci.yml. - The org has already written much of its own remediation:
alphaswarm_internal/TESTING_FRAMEWORK_BLUEPRINT.md(991 lines: flaky detection, quarantine, testkit design), the AI-dev kit (prompt state machine, schemas, templates),alphaswarm_config/ANALYSIS.md, and docs/internal/org-audit/ (2026-07-14). This plan executes those; it does not rediscover them. - Rigor artifacts to preserve: TLA+ specs with TLC targets (core, bots), hypothesis property-based FSM tests (bots), hypothesis stateful orchestrator tests (local), contract "honesty" tests forbidding fabricated monitoring data (admin), the AgenticTester autonomous test-debug framework (agents).
2. The ten gaps
Ranked by impact on agentic development. Each maps to workstreams in workstreams.md.
G1 — No CI at all in ~17 repos, including the packages everything depends on → WS1
No .github/workflows/ in alphaswarm_api, _assistant, _auth,
_catalog, _client, _config, _data, _finops, _graph, _index,
_ingest, _internal, _learning, _models, _ops_console,
_orchestration, _research, _viz. alphaswarm_core's only workflow is
buf.yml (proto lint) — its ~70-file/15.9k-line pytest suite never runs in
CI, and its README references a ci.yml that does not exist.
alphaswarm_config (editable-installed by ~15 pipelines) has zero CI despite
a 42-file suite. alphaswarm_client (the active Vite frontend: 24 vitest
files, 60% coverage thresholds, 17 Playwright specs) has zero CI.
alphaswarm_auth (CA/WebAuthn/RBAC, 110 test functions) has zero
enforcement. Agent-authored PRs (claude/* branches; Codex automerge commits
in _catalog, _orchestration, _viz, _eval) merge with no automated
verification. Nearly all of these repos have complete, documented local
test suites — the gap is pure wiring.
G2 — Broken and dead CI where workflows do exist → WS0
Status update (verified against the working tree, 2026-08): the four items below with a ✅ have been fixed — matching WS0's own acceptance criterion and handoff.md's claim that WS0 merged across the estate. Left as historical evidence of the gap WS0 was created to close; do not re-open without re-verifying.
- ✅ (fixed)
alphaswarm/.github/workflows/pr-validate.ymldefined the job keypython:twice (lines 33/47) — GitHub rejects the YAML, silently disabling the gitleaks+trivy security job and the alembic-immutability gate on every PR. The working tree now has exactly onepython:job key in this file. - Monolith workflows
alphaswarm-client.yml,e2e-eda.yml,quantbot-bots-image.yml,ml-pipeline.yml,docs-ci.ymlare path-filtered on directories removed by the repo split — they can never trigger, and the extracted repos did not inherit the CI. - ✅ (fixed)
alphaswarm_ide/build.ymlandlicense-check-workflow.ymltriggered onmasterwhile trunk ismain— IDE PRs got no lint/build/test gate. Both workflows now trigger onmainin the working tree. - ✅ (fixed)
alphaswarm_eval/ci.ymltriggered onmainbut remote branches were onlydevelopmentandclaude/*— the eval gate never fired. The working tree'sci.ymlnow triggers ondevelopment(push and pull_request). security-scan.yml's weekly bandit/npm-audit jobs targetalphaswarm_core/srcandalphaswarm_client/paths that no longer exist in the monolith.
G3 — Green-by-skip and report-only gates: a passing check often certifies nothing → WS1
alphaswarm_agents/ci.yml and alphaswarm_controller/ci.yml exit 0 when
sibling checkout fails and on fork PRs; the monolith's own comments admit
test-monolith "was green-by-skip for months." Enforce flags never flipped:
LICENSE_GATE_ENFORCE (agents, controller), EVAL_GATE_ENFORCE (eval —
gate.py:185-189 exits 0 on failure), RUN_CONTRACT_RESOLUTION (mlops — the
186 frozen contracts are never import-verified), ALPHASWARM_WORKER_CI_FULL
(worker pytest skips entirely). Ruff report-only in controller (722-finding
backlog behind || echo ::warning::); mypy soft-failed via || true in core
and bots Makefiles. alphaswarm_ui: typecheck/lint/vitest all
continue-on-error plus next.config.mjs ignoreBuildErrors — only
next build gates merge. Known-red quarantines: 8 test files ignored in
controller, 7 --ignore entries in the monolith, ~287 acknowledged
full-suite failures (work-surface-gate.yml). Agents cannot use "CI is
green" as a done-signal — the single most corrosive property for autonomous
development.
G4 — Unreproducible bootstrap: unpublished packages, six PATs, three checkout strategies → WS1+WS2
alphaswarm-core is not resolvable from any index (pip install --dry-run -e . fails in controller: "No matching distribution found for
alphaswarm-core>=0.2.1"; same pattern in admin, finops, graph, kb, cli,
learning). The monolith and worker pin git-SHA direct refs. Install-order
lore lives in pyproject comments. CI uses six differently-named long-lived
PATs (WORKSPACE_CHECKOUT_PAT, ALPHASWARM_REPO_READ_TOKEN,
SIBLING_REPOS_TOKEN, ALPHASWARM_CORE_PAT, ALPHASWARM_ASSISTANT_PAT,
ALPHASWARM_REPOSITORY_READ_TOKEN). A CodeArtifact index hosting
alphaswarm-core exists but is wired only into
alphaswarm_admin/.github/workflows/build-publish.yml:94-100. Lockfiles in
2/31 Python repos (config, qap). No devcontainer anywhere; .env.example in
only 4 repos. An agent in a fresh sandbox cannot install, typecheck, or
test in roughly half the Python estate.
G5 — The canonical AGENTS.md has decayed into a 146KB accretion → WS3
Three conflicting hard-rule counts: AGENTS.md:14 says 55, WORKFLOW.md:201
says 45, .cursor/rules/alphaswarm.mdc:135 says 45 — the actual list numbers
1–65. 45 of 453 relative links are dead even after applying the file's own
sibling-resolution rule. Hardcoded developer paths
(/Users/julesthecomputernerd/..., C:/alphaswarm-warehouse). Rename
artifacts and triplicated "Where to look" rows (lines ~1136–1161). It
contradicts owning repos: line 144 claims Auth0+Entra dual auth for
alphaswarm_ui while alphaswarm_ui/AGENTS.md is Entra-only with a CI
Auth0-removal guard; rule 47 pins alphaswarm.fund domains vs. the actual
alpha-swarm.ai. The sibling-repo table omits 18 checked-out repos including
alphaswarm_api — the security-bearing edge gateway has zero mentions.
Because downstream repos cite this file by rule number ("AGENTS rule
22/26/27/44/49"), the drift propagates org-wide; agents burn ~25K tokens
loading partially false context and dead-end on ~10% of citations.
G6 — Multi-tool config sprawl with two contradictory world-models → WS3+WS4
A byte-identical 704-byte .claude.md stub is duplicated across ~35 repos —
the lowercase filename means Claude Code never auto-loads it, so the tool
behind most agent branches gets the least guidance of any tool. The per-tool
boilerplate set (.cursorrules, .clinerules, .devin, .junie,
copilot-instructions.md) teaches the QAP doctrine (LLM-Plane/Money-Plane,
ArcticDB, NautilusTrader) which directly conflicts with monolith hard rules
3/46 (Iceberg as the sole lakehouse path) and never mentions spec runtimes or
hash-locking; AGENTS.md teaches the opposite half. In the 11 repos with no
AGENTS.md at all (_catalog, _config, _docs, _eval, _internal,
_observe, _observe_js, _ops_console, _research, _viz, _website)
the contradictory boilerplate is the only guidance.
G7 — Zero CI reuse and copy-paste drift → WS1
No workflow_call or merge_group usage in any of the 57 workflow files
org-wide. Two composites both named build-sign-push with incompatible
interfaces (alphaswarm/.github/actions/ vs
alphaswarm_platform/.github/actions/). Action versions drift v4–v7 within
single repos; only alphaswarm_ide pins actions by SHA (30/36); the monolith
pins 0/179. Renovate config exists in 2/41 repos. Supply-chain enforcement is
bimodal: the platform composite blocks on Trivy HIGH/CRITICAL while
alphaswarm_admin signs with continue-on-error and removed Trivy — and the
ui/ide image paths that actually ship to prod ECS have no signing/SBOM at
all.
G8 — Systemic docs/link rot with no mechanical freshness gate → WS3+WS5
Core agentic docs (last_reviewed: 2026-05-25) predate the repo split
(2026-06-18) and ADR 025's orchestration authority shift — they still teach
Celery-beat as current. Three contradictory WorkflowSpec schemas across
workflow-studio.md vs first-agent-workflow.md/agentic-pipeline.md;
adapter counts 5 vs 6 vs 7; three incompatible agent-run endpoint
descriptions. alphaswarm-monorepo-paths.md names repos that don't exist and
omits ~15 real ones. alphaswarm_index's sole-writer invariant is already
breached (4 files stamped "refreshed by Junie"). docs-ci.yml runs
lychee/vale only as continue-on-error; referenced contract docs
(EVALUATION_TAXONOMY.md, GOLDEN_DATASET_PLAN.md) exist nowhere on the
filesystem.
G9 — The eval loop and agent-outcome measurement are both open → WS7
alphaswarm_eval's CI gate is inert three ways (wrong trigger branch,
EVAL_GATE_ENFORCE unset, and it scores a fixture whose outputs are
byte-equal to the goldens per its own PROVENANCE.md). No code anywhere
produces live model outputs into the gate. A diverged second copy of the eval
package lives in alphaswarm_mlops/packages/alphaswarm_eval. The LLM judge
in alphaswarm_agents/evaluation.py silently no-ops on failure.
✅ (fixed, verified in the working tree) RosterEvaluator.evaluate_spec
used to pass a str where an AgentSpec was required (guaranteed
AttributeError); it now resolves the spec through
alphaswarm_agents.registry.get_agent_spec first
(alphaswarm_agents/src/alphaswarm_agents/exemplars/evaluator.py) — WS0.10
landed. Zero instrumentation of
agent-dev outcomes: greps for revert-rate/DORA/change-failure/time-to-merge
across docs, eval, agents, core, internal return nothing — despite
agent-authored PRs being identifiable by branch prefix. The org is running
an agentic-development program it cannot measure — and this plan currently
has no baseline.
G10 — No uniform bootstrap/verify entrypoint → WS2
Makefiles in only 6/41 repos with disjoint vocabularies (~60 targets in the
monolith; spec-check/test/lint in core; image-only in data); zero
justfiles; pre-commit config in exactly 1 repo. 21/30 AGENTS.md files have a
Validation block but commands diverge (pytest -q vs -ra vs
python -m pytest tests -q; pip vs uv; four different boundary-script
names). alphaswarm_bots/AGENTS.md's documented validation command
(pytest tests/bots) doesn't match the actual layout — an agent following it
verbatim runs zero tests. Every agent session spends turns
reverse-engineering how to verify.
3. Inconsistency clusters
Where repos diverge on the same concern (full remediation in WS2/WS3):
- Sibling-dependency resolution: git-SHA refs vs bare unpublishable
specifiers vs uv workspace (qap only);
.siblings/checkout vs../installs vs flat venv; six PAT names. - Claude entry-point: lowercase
.claude.mdstub (~35 repos, never loaded) vs nothing vs.claude/launch.jsononly (ui). No repo has a standardCLAUDE.md. Meanwhile Cursor gets 70 rules + 30 skills + 33 subagents in the monolith — a Cursor monoculture. - Boundary-check scripts: five names for the same concern, only some CI-enforced.
- Lint gating semantics: blocking with exact pin (worker, rl) vs report-only (controller) vs ratchet waves (monolith) vs advisory (admin, ui) vs no linter (agents, graph).
- Python toolchain:
requires-python>=3.10/3.11/3.12 mix; CI matrices differ per repo; lockfiles 2/31; ruff line-length 100 vs 120; mypy strict vs lenient vs absent. - JS toolchain: biome vs eslint vs prettier-via-turbo; node engines
=18/>=20/>=22;
packageManagerin 3/8;alphaswarm_clientcarries bothpackage-lock.jsonandpnpm-lock.yaml;alphaswarm_uicarries jest configs alongside canonical vitest. - Ownership metadata: CODEOWNERS vs docs frontmatter
owner:vs Neo4j ownership graph — no surface answers "who ownsalphaswarm/auth/". - Drifting restated constants: rule counts 45/55/65; adapter counts 5/6/7; spec-runtime counts 4/5/6/7; three incompatible AgentSpec run-endpoint vocabularies.
- Stale personal-account metadata: pyproject URLs and ghcr publishing
under personal namespaces (
julianwiley,julianwileymac) instead ofAlpha-Swarm-ai.
4. External landscape calibration
The two internal research reports (2026-07) recommend a stack of: portable guidance (AGENTS.md), reusable skills, typed tools (MCP), explicit coordination patterns, durable state, artifact-forward promotion, auditable controls, worktree isolation, OIDC, signing/SBOM, OTel GenAI telemetry, and an experimentation program. Their load-bearing external claims were re-verified against live primary sources on 2026-07-19 (adversarial 3-voter verification; unanimous unless noted):
Verified (adopt with confidence):
- MCP
2026-07-28stateless revision is on a committed official schedule (RC locked 2026-05-21; beta SDKs 2026-06-29; final 2026-07-28) and is backward compatible — protocol-level statelessness (SEP-2575/SEP-2567,server/discover), with automatic fallback to the legacy handshake. No forced migration for our MCP servers; adopt v2 SDKs deliberately. (blog.modelcontextprotocol.io) - AGENTS.md is stewarded by the Agentic AI Foundation under the Linux Foundation; ~23 tools support it (Cursor, Copilot, Codex, Devin, Zed, …); 60k+ projects self-reported, ~161k files found via GitHub code search. Cursor natively merges nested AGENTS.md with most-specific-wins.
- Agent Skills is an open standard (agentskills.io, published
2025-12-18). Cursor loads skills from
.cursor/skills/,.agents/skills/,.claude/skills/, and.codex/skills/— one skill directory can now serve every tool we use. - A2A 1.0.0 is stable (2026-03-12, Linux Foundation); IBM's "ACP" merged
into A2A; Zed's unrelated Agent Client Protocol (also "ACP") has real
cross-editor adoption (Cursor CLI
agent acp, JetBrains, Neovim). The reports' warning about acronym collisions stands. - Worktree isolation is first-class product functionality: Claude Code
--worktree/-w,isolation: worktreesubagent frontmatter,WorktreeCreate/WorktreeRemovehooks; Cursor/worktree,/best-of-n, Agents Window (default cap 25 worktrees/machine). Known bug: Claude Codeisolation: worktreenot respected via theclaude --agentinvocation path (issue #50357). - Claude Code plugins are enterprise-governable: plugins bundle skills,
agents, hooks, MCP servers, LSP servers;
managed(read-only) andproject(.claude/settings.json) scopes;strictKnownMarketplaces/blockedMarketplacescannot be overridden by users or projects. - Cursor Cloud Agents (rebranded Background Agents) run in isolated
per-task VMs, produce merge-ready PRs, support team MCP servers and
.cursor/hooks.json. - First-party CI agent actions exist:
openai/codex-action@v1(installs Codex CLI, runscodex execunder specified permissions) — the Codex-side counterpart to Claude Code's GitHub app/action surface.
NOT verified — treat as open, do not build on without checking (reports' claims that did not survive, or were not researched to completion):
- The EU AI Act 2026-08-02 applicability date and which obligations actually bind AlphaSwarm — needs a legal/compliance check before any roadmap commitment (it is 14 days out if true; flagged in workstreams.md WS-G).
- OpenTelemetry GenAI semantic conventions stability status — our
telemetry plan (metrics.md) therefore keys on
alphaswarm_core.observe's own span schema, not on external semconv stability. - Supply-chain specifics (current SLSA/Sigstore/CycloneDX/OPA guidance details) — our WS1 supply-chain actions extend the org's own proven composites rather than citing external baselines.
- Measured productivity results for agentic coding (DORA-style evidence) — no credible external numbers survived verification, which is precisely why metrics.md builds our own measurement first.
Where the reports' advice is already exceeded internally (conformance snapshot; the full mapping is threaded through the workstreams):
| Report recommendation | AlphaSwarm reality | Status |
|---|---|---|
| Skills as versioned, auditable procedure artifacts | Hash-locked *_spec_versions + ledger replay (five runtimes) — stronger invariants than SKILL.md | ✅ exceeds (for product agents; dev-workflow skills still Cursor-only → WS4) |
| Typed tools w/ risk metadata + approval | DataMCPTool mutates/required_scopes/tenancy_posture; RFC 9728/8707/8693 conformance (hard rules 49/54) | ✅ (Data MCP) / 🟡 uniformity (enhancement-guide E7) |
| Permissions outside natural language | Intervention nodes (WORKFLOW.md), kill-switch, One Gate, ApprovedPromotion tokens | ✅ concept; 🟡 dev-loop enforcement (WS4 hooks) |
| Plan artifact for non-trivial changes | Plan→Act→Reflect + plan-mode contract; .cursor/plans/ | ✅ Cursor; 🟡 tool-neutral evidence bundles (WS3) |
| Maker-checker coordination | alphaswarm-hard-rules-reviewer, adversarial-run-report-reviewer subagents; DialecticalDebateAdapter in the product | 🟡 exists as prompts, not as registered governed workflow (WS6) |
| OIDC, signing, SBOM, policy gates | Present and SOC2-referenced in alphaswarm_platform; absent on the paths that actually deploy (admin/ui/ide) | 🟡 → WS1 |
| Worktrees for parallel sessions | best-of-n-runner mentioned in WORKFLOW.md; no org convention | 🟡 → WS4 |
| Task/tool/artifact telemetry + experimentation | alphaswarm_core.observe + EVAL spans exist; zero agent-dev-outcome instrumentation | ❌ → WS7 |