Saltar al contenido principal

Current-state gap analysis

Part of the agentic workflows enhancement plan. Gaps here are closed by the workstreams in workstreams.md; measurement is defined in metrics.md. Evidence cites working-tree state as of 2026-07-19.

1. Strengths — build on these, do not rebuild them​

The plan's first rule is that the following assets are the templates, not the debt:

  • The repo-specific AGENTS.md contract pattern is best-in-class in ~20 repos and should be propagated as-is: alphaswarm_api/AGENTS.md (363 lines, "AGENTS.md is the law"), alphaswarm_ui/AGENTS.md (242 lines, 12 numbered security boundaries incl. CVE pinning), and the alphaswarm_finops / _worker / _kb / _rl / _data / _mlops family with "Hard Boundaries" + "Where Changes Go" routing tables + Validation blocks.
  • Machine-enforced boundaries already exist as committed scripts: check_no_alphaswarm_imports.py (data, kb_federation, local, learning, mcp), check_plane_boundary.py AST check (worker, blocking in CI), the monolith's ~28 scripts/ci/ lints, alembic migration-immutability hash-lock, buf breaking (core), CRD schema drift enforced three ways (bots, ADR 006), oasdiff/griffe API-surface gates.
  • Contract lock-step patterns worth copying verbatim: alphaswarm_cli's generate-docs --agent-contract --check (docs/agent-contract.json kept in CI lock-step with code); the monolith's "AGENTS hard-rule lints" CI job; alphaswarm_mlops' sha256-locked frozen_contracts registry (186 contracts) with import-boundary checker.
  • Supply-chain maturity exists where wired: cosign keyless + SLSA L3 (slsa-github-generator v2.1.0) + syft SBOM attestation in alphaswarm/.github/workflows/build-publish.yml; alphaswarm_platform's build-sign-push composite blocks on Trivy HIGH/CRITICAL and is SOC2-referenced; OIDC-first cloud auth via a dedicated aws-oidc-assume composite; the governed One Gate Terraform apply with mandatory break-glass reasons (ADR 022).
  • alphaswarm_qap is the reproducible-bootstrap exemplar: uv workspace, committed uv.lock, .python-version, PEP-735 dependency-groups, every gate runnable via uv run, fail-closed attestation gates.
  • The platform's own agent runtime is dogfood-ready: hash-locked AgentSpec/WorkflowRuntime with guardrail budgets and kill-switch polling; seven orchestration adapters including debate topologies; the ApprovedPromotion HITL capability-token pattern (alphaswarm_agents/src/alphaswarm_agents/promotion_gate.py); evidence-ref-enforced FindingRecorder; decision-log-with-reflection memory; shipped dev-facing specs (configs/agents/codebase_assistant.yaml, codebase_refactorer.yaml) bound to a working Codebase MCP.
  • Context-layer infrastructure exists and is governed: alphaswarm_mcp serves rules/skills/AGENTS.md bundles with Claude Code .mcp.json wiring; alphaswarm_index has a sole-writer contract, cite-or-TODO discipline, and per-area token budgets; alphaswarm_kb provides fail-closed permissioned memory with bi-temporal envelopes; the federation gateway handles cross-silo sharing.
  • Deliberately offline-testable repos are exactly what sandboxed agents need: alphaswarm_data (lazy SDKs, injectable transports, 39 offline test files), alphaswarm_finops (moto/MockTransport), alphaswarm_eval (zero required runtime deps), alphaswarm_observe (graceful degrade).
  • An honest engineering culture encoded in CI: dated known-red quarantine lists with burn-down instructions, "NEVER add || true" comments, green-by-skip called out and fixed loudly in the 2026-07 monolith CI rewrite, alphaswarm_ui's 36-hour-outage postmortem inline in ci.yml.
  • The org has already written much of its own remediation: alphaswarm_internal/TESTING_FRAMEWORK_BLUEPRINT.md (991 lines: flaky detection, quarantine, testkit design), the AI-dev kit (prompt state machine, schemas, templates), alphaswarm_config/ANALYSIS.md, and docs/internal/org-audit/ (2026-07-14). This plan executes those; it does not rediscover them.
  • Rigor artifacts to preserve: TLA+ specs with TLC targets (core, bots), hypothesis property-based FSM tests (bots), hypothesis stateful orchestrator tests (local), contract "honesty" tests forbidding fabricated monitoring data (admin), the AgenticTester autonomous test-debug framework (agents).

2. The ten gaps​

Ranked by impact on agentic development. Each maps to workstreams in workstreams.md.

G1 — No CI at all in ~17 repos, including the packages everything depends on → WS1​

No .github/workflows/ in alphaswarm_api, _assistant, _auth, _catalog, _client, _config, _data, _finops, _graph, _index, _ingest, _internal, _learning, _models, _ops_console, _orchestration, _research, _viz. alphaswarm_core's only workflow is buf.yml (proto lint) — its ~70-file/15.9k-line pytest suite never runs in CI, and its README references a ci.yml that does not exist. alphaswarm_config (editable-installed by ~15 pipelines) has zero CI despite a 42-file suite. alphaswarm_client (the active Vite frontend: 24 vitest files, 60% coverage thresholds, 17 Playwright specs) has zero CI. alphaswarm_auth (CA/WebAuthn/RBAC, 110 test functions) has zero enforcement. Agent-authored PRs (claude/* branches; Codex automerge commits in _catalog, _orchestration, _viz, _eval) merge with no automated verification. Nearly all of these repos have complete, documented local test suites — the gap is pure wiring.

G2 — Broken and dead CI where workflows do exist → WS0​

Status update (verified against the working tree, 2026-08): the four items below with a ✅ have been fixed — matching WS0's own acceptance criterion and handoff.md's claim that WS0 merged across the estate. Left as historical evidence of the gap WS0 was created to close; do not re-open without re-verifying.

  • ✅ (fixed) alphaswarm/.github/workflows/pr-validate.yml defined the job key python: twice (lines 33/47) — GitHub rejects the YAML, silently disabling the gitleaks+trivy security job and the alembic-immutability gate on every PR. The working tree now has exactly one python: job key in this file.
  • Monolith workflows alphaswarm-client.yml, e2e-eda.yml, quantbot-bots-image.yml, ml-pipeline.yml, docs-ci.yml are path-filtered on directories removed by the repo split — they can never trigger, and the extracted repos did not inherit the CI.
  • ✅ (fixed) alphaswarm_ide/build.yml and license-check-workflow.yml triggered on master while trunk is main — IDE PRs got no lint/build/test gate. Both workflows now trigger on main in the working tree.
  • ✅ (fixed) alphaswarm_eval/ci.yml triggered on main but remote branches were only development and claude/* — the eval gate never fired. The working tree's ci.yml now triggers on development (push and pull_request).
  • security-scan.yml's weekly bandit/npm-audit jobs target alphaswarm_core/src and alphaswarm_client/ paths that no longer exist in the monolith.

G3 — Green-by-skip and report-only gates: a passing check often certifies nothing → WS1​

alphaswarm_agents/ci.yml and alphaswarm_controller/ci.yml exit 0 when sibling checkout fails and on fork PRs; the monolith's own comments admit test-monolith "was green-by-skip for months." Enforce flags never flipped: LICENSE_GATE_ENFORCE (agents, controller), EVAL_GATE_ENFORCE (eval — gate.py:185-189 exits 0 on failure), RUN_CONTRACT_RESOLUTION (mlops — the 186 frozen contracts are never import-verified), ALPHASWARM_WORKER_CI_FULL (worker pytest skips entirely). Ruff report-only in controller (722-finding backlog behind || echo ::warning::); mypy soft-failed via || true in core and bots Makefiles. alphaswarm_ui: typecheck/lint/vitest all continue-on-error plus next.config.mjs ignoreBuildErrors — only next build gates merge. Known-red quarantines: 8 test files ignored in controller, 7 --ignore entries in the monolith, ~287 acknowledged full-suite failures (work-surface-gate.yml). Agents cannot use "CI is green" as a done-signal — the single most corrosive property for autonomous development.

G4 — Unreproducible bootstrap: unpublished packages, six PATs, three checkout strategies → WS1+WS2​

alphaswarm-core is not resolvable from any index (pip install --dry-run -e . fails in controller: "No matching distribution found for alphaswarm-core>=0.2.1"; same pattern in admin, finops, graph, kb, cli, learning). The monolith and worker pin git-SHA direct refs. Install-order lore lives in pyproject comments. CI uses six differently-named long-lived PATs (WORKSPACE_CHECKOUT_PAT, ALPHASWARM_REPO_READ_TOKEN, SIBLING_REPOS_TOKEN, ALPHASWARM_CORE_PAT, ALPHASWARM_ASSISTANT_PAT, ALPHASWARM_REPOSITORY_READ_TOKEN). A CodeArtifact index hosting alphaswarm-core exists but is wired only into alphaswarm_admin/.github/workflows/build-publish.yml:94-100. Lockfiles in 2/31 Python repos (config, qap). No devcontainer anywhere; .env.example in only 4 repos. An agent in a fresh sandbox cannot install, typecheck, or test in roughly half the Python estate.

G5 — The canonical AGENTS.md has decayed into a 146KB accretion → WS3​

Three conflicting hard-rule counts: AGENTS.md:14 says 55, WORKFLOW.md:201 says 45, .cursor/rules/alphaswarm.mdc:135 says 45 — the actual list numbers 1–65. 45 of 453 relative links are dead even after applying the file's own sibling-resolution rule. Hardcoded developer paths (/Users/julesthecomputernerd/..., C:/alphaswarm-warehouse). Rename artifacts and triplicated "Where to look" rows (lines ~1136–1161). It contradicts owning repos: line 144 claims Auth0+Entra dual auth for alphaswarm_ui while alphaswarm_ui/AGENTS.md is Entra-only with a CI Auth0-removal guard; rule 47 pins alphaswarm.fund domains vs. the actual alpha-swarm.ai. The sibling-repo table omits 18 checked-out repos including alphaswarm_api — the security-bearing edge gateway has zero mentions. Because downstream repos cite this file by rule number ("AGENTS rule 22/26/27/44/49"), the drift propagates org-wide; agents burn ~25K tokens loading partially false context and dead-end on ~10% of citations.

G6 — Multi-tool config sprawl with two contradictory world-models → WS3+WS4​

A byte-identical 704-byte .claude.md stub is duplicated across ~35 repos — the lowercase filename means Claude Code never auto-loads it, so the tool behind most agent branches gets the least guidance of any tool. The per-tool boilerplate set (.cursorrules, .clinerules, .devin, .junie, copilot-instructions.md) teaches the QAP doctrine (LLM-Plane/Money-Plane, ArcticDB, NautilusTrader) which directly conflicts with monolith hard rules 3/46 (Iceberg as the sole lakehouse path) and never mentions spec runtimes or hash-locking; AGENTS.md teaches the opposite half. In the 11 repos with no AGENTS.md at all (_catalog, _config, _docs, _eval, _internal, _observe, _observe_js, _ops_console, _research, _viz, _website) the contradictory boilerplate is the only guidance.

G7 — Zero CI reuse and copy-paste drift → WS1​

No workflow_call or merge_group usage in any of the 57 workflow files org-wide. Two composites both named build-sign-push with incompatible interfaces (alphaswarm/.github/actions/ vs alphaswarm_platform/.github/actions/). Action versions drift v4–v7 within single repos; only alphaswarm_ide pins actions by SHA (30/36); the monolith pins 0/179. Renovate config exists in 2/41 repos. Supply-chain enforcement is bimodal: the platform composite blocks on Trivy HIGH/CRITICAL while alphaswarm_admin signs with continue-on-error and removed Trivy — and the ui/ide image paths that actually ship to prod ECS have no signing/SBOM at all.

Core agentic docs (last_reviewed: 2026-05-25) predate the repo split (2026-06-18) and ADR 025's orchestration authority shift — they still teach Celery-beat as current. Three contradictory WorkflowSpec schemas across workflow-studio.md vs first-agent-workflow.md/agentic-pipeline.md; adapter counts 5 vs 6 vs 7; three incompatible agent-run endpoint descriptions. alphaswarm-monorepo-paths.md names repos that don't exist and omits ~15 real ones. alphaswarm_index's sole-writer invariant is already breached (4 files stamped "refreshed by Junie"). docs-ci.yml runs lychee/vale only as continue-on-error; referenced contract docs (EVALUATION_TAXONOMY.md, GOLDEN_DATASET_PLAN.md) exist nowhere on the filesystem.

G9 — The eval loop and agent-outcome measurement are both open → WS7​

alphaswarm_eval's CI gate is inert three ways (wrong trigger branch, EVAL_GATE_ENFORCE unset, and it scores a fixture whose outputs are byte-equal to the goldens per its own PROVENANCE.md). No code anywhere produces live model outputs into the gate. A diverged second copy of the eval package lives in alphaswarm_mlops/packages/alphaswarm_eval. The LLM judge in alphaswarm_agents/evaluation.py silently no-ops on failure. ✅ (fixed, verified in the working tree) RosterEvaluator.evaluate_spec used to pass a str where an AgentSpec was required (guaranteed AttributeError); it now resolves the spec through alphaswarm_agents.registry.get_agent_spec first (alphaswarm_agents/src/alphaswarm_agents/exemplars/evaluator.py) — WS0.10 landed. Zero instrumentation of agent-dev outcomes: greps for revert-rate/DORA/change-failure/time-to-merge across docs, eval, agents, core, internal return nothing — despite agent-authored PRs being identifiable by branch prefix. The org is running an agentic-development program it cannot measure — and this plan currently has no baseline.

G10 — No uniform bootstrap/verify entrypoint → WS2​

Makefiles in only 6/41 repos with disjoint vocabularies (~60 targets in the monolith; spec-check/test/lint in core; image-only in data); zero justfiles; pre-commit config in exactly 1 repo. 21/30 AGENTS.md files have a Validation block but commands diverge (pytest -q vs -ra vs python -m pytest tests -q; pip vs uv; four different boundary-script names). alphaswarm_bots/AGENTS.md's documented validation command (pytest tests/bots) doesn't match the actual layout — an agent following it verbatim runs zero tests. Every agent session spends turns reverse-engineering how to verify.

3. Inconsistency clusters​

Where repos diverge on the same concern (full remediation in WS2/WS3):

  • Sibling-dependency resolution: git-SHA refs vs bare unpublishable specifiers vs uv workspace (qap only); .siblings/ checkout vs ../ installs vs flat venv; six PAT names.
  • Claude entry-point: lowercase .claude.md stub (~35 repos, never loaded) vs nothing vs .claude/launch.json only (ui). No repo has a standard CLAUDE.md. Meanwhile Cursor gets 70 rules + 30 skills + 33 subagents in the monolith — a Cursor monoculture.
  • Boundary-check scripts: five names for the same concern, only some CI-enforced.
  • Lint gating semantics: blocking with exact pin (worker, rl) vs report-only (controller) vs ratchet waves (monolith) vs advisory (admin, ui) vs no linter (agents, graph).
  • Python toolchain: requires-python >=3.10/3.11/3.12 mix; CI matrices differ per repo; lockfiles 2/31; ruff line-length 100 vs 120; mypy strict vs lenient vs absent.
  • JS toolchain: biome vs eslint vs prettier-via-turbo; node engines

    =18/>=20/>=22; packageManager in 3/8; alphaswarm_client carries both package-lock.json and pnpm-lock.yaml; alphaswarm_ui carries jest configs alongside canonical vitest.

  • Ownership metadata: CODEOWNERS vs docs frontmatter owner: vs Neo4j ownership graph — no surface answers "who owns alphaswarm/auth/".
  • Drifting restated constants: rule counts 45/55/65; adapter counts 5/6/7; spec-runtime counts 4/5/6/7; three incompatible AgentSpec run-endpoint vocabularies.
  • Stale personal-account metadata: pyproject URLs and ghcr publishing under personal namespaces (julianwiley, julianwileymac) instead of Alpha-Swarm-ai.

4. External landscape calibration​

The two internal research reports (2026-07) recommend a stack of: portable guidance (AGENTS.md), reusable skills, typed tools (MCP), explicit coordination patterns, durable state, artifact-forward promotion, auditable controls, worktree isolation, OIDC, signing/SBOM, OTel GenAI telemetry, and an experimentation program. Their load-bearing external claims were re-verified against live primary sources on 2026-07-19 (adversarial 3-voter verification; unanimous unless noted):

Verified (adopt with confidence):

  • MCP 2026-07-28 stateless revision is on a committed official schedule (RC locked 2026-05-21; beta SDKs 2026-06-29; final 2026-07-28) and is backward compatible — protocol-level statelessness (SEP-2575/SEP-2567, server/discover), with automatic fallback to the legacy handshake. No forced migration for our MCP servers; adopt v2 SDKs deliberately. (blog.modelcontextprotocol.io)
  • AGENTS.md is stewarded by the Agentic AI Foundation under the Linux Foundation; ~23 tools support it (Cursor, Copilot, Codex, Devin, Zed, …); 60k+ projects self-reported, ~161k files found via GitHub code search. Cursor natively merges nested AGENTS.md with most-specific-wins.
  • Agent Skills is an open standard (agentskills.io, published 2025-12-18). Cursor loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, and .codex/skills/ — one skill directory can now serve every tool we use.
  • A2A 1.0.0 is stable (2026-03-12, Linux Foundation); IBM's "ACP" merged into A2A; Zed's unrelated Agent Client Protocol (also "ACP") has real cross-editor adoption (Cursor CLI agent acp, JetBrains, Neovim). The reports' warning about acronym collisions stands.
  • Worktree isolation is first-class product functionality: Claude Code --worktree/-w, isolation: worktree subagent frontmatter, WorktreeCreate/WorktreeRemove hooks; Cursor /worktree, /best-of-n, Agents Window (default cap 25 worktrees/machine). Known bug: Claude Code isolation: worktree not respected via the claude --agent invocation path (issue #50357).
  • Claude Code plugins are enterprise-governable: plugins bundle skills, agents, hooks, MCP servers, LSP servers; managed (read-only) and project (.claude/settings.json) scopes; strictKnownMarketplaces/blockedMarketplaces cannot be overridden by users or projects.
  • Cursor Cloud Agents (rebranded Background Agents) run in isolated per-task VMs, produce merge-ready PRs, support team MCP servers and .cursor/hooks.json.
  • First-party CI agent actions exist: openai/codex-action@v1 (installs Codex CLI, runs codex exec under specified permissions) — the Codex-side counterpart to Claude Code's GitHub app/action surface.

NOT verified — treat as open, do not build on without checking (reports' claims that did not survive, or were not researched to completion):

  • The EU AI Act 2026-08-02 applicability date and which obligations actually bind AlphaSwarm — needs a legal/compliance check before any roadmap commitment (it is 14 days out if true; flagged in workstreams.md WS-G).
  • OpenTelemetry GenAI semantic conventions stability status — our telemetry plan (metrics.md) therefore keys on alphaswarm_core.observe's own span schema, not on external semconv stability.
  • Supply-chain specifics (current SLSA/Sigstore/CycloneDX/OPA guidance details) — our WS1 supply-chain actions extend the org's own proven composites rather than citing external baselines.
  • Measured productivity results for agentic coding (DORA-style evidence) — no credible external numbers survived verification, which is precisely why metrics.md builds our own measurement first.

Where the reports' advice is already exceeded internally (conformance snapshot; the full mapping is threaded through the workstreams):

Report recommendationAlphaSwarm realityStatus
Skills as versioned, auditable procedure artifactsHash-locked *_spec_versions + ledger replay (five runtimes) — stronger invariants than SKILL.md✅ exceeds (for product agents; dev-workflow skills still Cursor-only → WS4)
Typed tools w/ risk metadata + approvalDataMCPTool mutates/required_scopes/tenancy_posture; RFC 9728/8707/8693 conformance (hard rules 49/54)✅ (Data MCP) / 🟡 uniformity (enhancement-guide E7)
Permissions outside natural languageIntervention nodes (WORKFLOW.md), kill-switch, One Gate, ApprovedPromotion tokens✅ concept; 🟡 dev-loop enforcement (WS4 hooks)
Plan artifact for non-trivial changesPlan→Act→Reflect + plan-mode contract; .cursor/plans/✅ Cursor; 🟡 tool-neutral evidence bundles (WS3)
Maker-checker coordinationalphaswarm-hard-rules-reviewer, adversarial-run-report-reviewer subagents; DialecticalDebateAdapter in the product🟡 exists as prompts, not as registered governed workflow (WS6)
OIDC, signing, SBOM, policy gatesPresent and SOC2-referenced in alphaswarm_platform; absent on the paths that actually deploy (admin/ui/ide)🟡 → WS1
Worktrees for parallel sessionsbest-of-n-runner mentioned in WORKFLOW.md; no org convention🟡 → WS4
Task/tool/artifact telemetry + experimentationalphaswarm_core.observe + EVAL spans exist; zero agent-dev-outcome instrumentation❌ → WS7