Skip to main content

ADR 015 — Runtime decomposition: cell-based modular monolith over domain microservices

Context​

The alphaswarm monolith is the platform's largest deployment unit: one FastAPI process (alphaswarm-core) carrying ~112 route modules, the in-process MCP routers (/mcp/data, /mcp/codebase, /mcp/ml), and a Celery task surface of ~57 task modules, backed by 56 ORM model files and 89 Alembic migrations over a single Postgres, plus Redis in at least six distinct roles (broker, metadata cache, progress pub/sub, RAG vectors, ownership event stream, sandbox namespaces). The Settings singleton exposes ~637 knobs.

The question on the table: should the hosted platform break this runtime into a distributed domain-microservices architecture (agents-service, backtest-service, ml-service, data-service, each with its own datastore), or adopt a different decomposition?

What is already decomposed​

The platform is not a greenfield monolith. Substantial decomposition has shipped or is accepted:

SeamStateWhere
Control plane (alphaswarm-cp)Shipped — standalone repo, image, /manage/* API; never imports alphaswarm.*ADR 005, alphaswarm_controller/
Shared kernelShipped — wire types, provider ABCs, WorkloadRuntimealphaswarm_core/
Compute splitShipped — alphaswarm-worker (queues default,paper,terraform,ingestion,workflows) vs alphaswarm-executor (backtest,training,ml,agents,factors,rag), independent HPAs on alphaswarm_celery_queue_depthworker-executor-images.md, alphaswarm_platform/deployments/kubernetes/base/
FrontendsShipped — alphaswarm_client, alphaswarm_ui, alphaswarm_admin, alphaswarm_ide, all HTTP-onlyADR 011
RL / ML boundary packagesShipped — alphaswarm_rl, alphaswarm_models with deprecation shims; routers/tasks mounted from the external packagesrepository-split.md
BotsBoundary package + kopf operator + per-bot podsADR 006, alphaswarm_bots/
Edge / cellsEnvoy alphaswarm-edge + alphaswarm-tenant-router (ext_authz, rendezvous cell routing), cell registry in topology.yaml, per-cell overlays + ArgoCD ApplicationSet, alphaswarm-cell-data-plane Helm chartcell-router-cutover.md, alphaswarm_platform/tenant_router/
KB boundaryAccepted (ADR-014) — alphaswarm_kb + alphaswarm_kb_federation designed; repositories now scaffolded and under active development (Track C repo creation has since happened)ADR 014

Invariants any split must respect​

Six hard-rule families make naive per-domain services with per-service databases actively harmful here:

  1. Single Postgres ledger. Every runtime writes immutable *_spec_versions snapshots and *_runs ledger rows through LedgerWriter, which stamps experiment_id/test_id from RequestContext (rules 13, 15, 17, 24, 34, 41, 43, 57). Splitting the ledger per service destroys the cross-domain experiment umbrella and the audit/replay story.
  2. Single LLM gateway. All LLM calls go through router_complete (rule 2); telemetry, cost caps, and the semantic cache depend on it.
  3. Single lakehouse write path. All Iceberg writes go through iceberg_catalog.append_arrow with medallion validation (rules 3, 21, 46).
  4. DataMCP boundary. Agents never read Postgres/Iceberg directly (rule 22) — the agent↔data seam is already a service-shaped API.
  5. Kill-switch fan-out. The topbar kill switch fans out to 12+ halt endpoints with a p99 propagation SLO; every new long-running runtime must join the fan-out, and every process fragment multiplies the propagation surface (rules 40, 45, 52).
  6. Idempotent cross-task state in Postgres only (rule 5) — Celery workers are already stateless and horizontally scalable; the "scaling" benefit of microservices largely exists today via queues.

Options considered​

Option 1 — Classic domain microservices​

Carve alphaswarm-core into independently deployed services (agents-svc, backtest-svc, analysis-svc, data-svc, trading-svc, …), each owning its own database and API, communicating via REST/gRPC and an event bus.

  • Violates invariant 1 (ledger) and 6 unless every service still writes to the shared Postgres — at which point they are not microservices, just N processes sharing one schema and one Alembic chain (a distributed monolith).
  • The hot coupling points (LedgerWriter, router_complete, append_arrow, metadata cache, progress bus) would become N× network hops with retry/outbox machinery the platform doesn't need.
  • Kill-switch propagation and hash-locked replay would have to be re-engineered across service boundaries.
  • The throughput-bound work (backtests, training, agent runs) is already isolated in the executor fleet with queue-depth autoscaling; a backtest-service would duplicate that with more moving parts.

Keep one logical application (alphaswarm runtime) but:

  1. Scale out by cell, not by domain. A cell = one namespace running the core/worker/executor/beat quartet against a per-cell data plane (CNPG Postgres, Redis, MinIO, MLflow, Iceberg REST), routed by the Envoy edge + tenant router (rendezvous hashing on tenant_id → cell_id). Tiers map onto the existing TenancyStrategy lattice (shared-std → RLS, shared-prem → schema-per-tenant, silo-reg → database-per-enterprise). This is RESTRUCTURING_PLAN Phases 3 + 6, already partially provisioned (cells: registry in topology.yaml, cell overlays, ApplicationSet, alphaswarm-cell-data-plane chart).
  2. Extract services only along the seams that already have service-shaped contracts — the hash-locked spec runtimes, the MCP HTTP surfaces, the control plane, and the operator pattern — and only when an extraction passes the Future Repo Split Gate in repository-split.md.

Option 3 — Status quo​

Keep the single alphaswarm-core Deployment and scale vertically. Rejected: noisy-neighbor risk across tenants, blast radius of one bad deploy is the whole fleet, and silo-reg compliance tenants cannot be served.

Decision​

Adopt Option 2. Decomposition proceeds in three tracks, ordered by risk and by whether new repositories are required.

Track A — Process/deployment splits of the existing images (no new repos)​

These change alphaswarm_platform/ manifests and entrypoints only; the code already supports them:

#CutDetail
A1alphaswarm-beat as a first-class DeploymentDeclared in topology.yaml and Terraform but missing from deployments/kubernetes/base/; promote it (replicas: 1, no HPA).
A2Standalone MCP server DeploymentsServe /mcp/data, /mcp/codebase, /mcp/ml from dedicated pods using the existing image with a scoped ASGI entrypoint. Topology already declares alphaswarm-ml-mcp as a separate service; RFC 9728/8707 audience binding (rule 49) already gives each MCP its own aud. Per-tenant MCP isolation then reuses the alphaswarm-mcp-tenant Helm chart (Phase 5).
A3Per-queue executor fleetsSplit the executor Deployment into per-queue ScaledObjects (KEDA) for backtest, training/ml, agents/rag so GPU-class and CPU-class work scale independently. No code change — queue routing exists in celery_app.py.
A4paper-trader and ingester-* as first-class K8s unitsThey exist as compose targets (paper, ingester image stages); give them base manifests + HPAs like worker/executor.
A5Cell rolloutExecute Phase 3 (cell registry + router live, RequestContext.cell_id propagating) then Phase 6 (per-cell Postgres/MinIO/Redis/MLflow via dual-write migration, ALPHASWARM_CELL_DUAL_WRITE).

Track B — Deepen existing extractions (existing repos, invasive code changes)​

#CutDetail
B1Sidecar control plane as hosted defaultFlip hosted deployments from ALPHASWARM_MANAGEMENT_MODE=embedded to sidecar; alphaswarm-cp already ships standalone.
B2Ledger/telemetry broker for alphaswarm_rl + alphaswarm_modelsToday the extracted packages still import the monolith for LedgerWriter, iceberg_catalog, _progress.emit, ORM. Introduce a narrow HTTP/MCP ledger-write surface (mirroring the controller's HttpAuditSink → /_internal/audit/terraform-runs pattern) so RL/ML workers can run from their own images without importing monolith ORM. This is the gating work for ever running them as separate services.
B3Bots operator fleetContinue the ADR 006 path: per-bot pods via quantbot-bot chart, latency-class scheduling (ADR 007), canary PnL gates (ADR 010). Bots are the one domain where per-workload processes are genuinely required (HFT node tiers).
B4CI boundary gatesExtend the rg-based forbidden-import gates from 2 to all 14+ subprojects (RESTRUCTURING_PLAN §2.1, §4.2) so extracted boundaries cannot silently re-couple. Prerequisite for everything above.

Track C — Extractions that require new repositories (permission gate)​

Per the workspace's repo-per-boundary convention, these need new git repositories and therefore explicit approval before any work begins:

#Candidate repoJustificationStatus
C1alphaswarm_kbADR-014 (accepted) defines the KB boundary package — KBRuntime, hash-locked KBCorpusSpec, adapter trinity. Monolith already mounts its router conditionally and migration 0088_alphaswarm_kb_specs shipped.Repo created and scaffolded (2026-06-10 onward)
C2alphaswarm_kb_federationADR-014's cross-silo federation gateway — standalone FastAPI, never imports alphaswarm.*; deployable today via compose/docker-compose.kb.yml patterns + Terragrunt silo modules.Repo created and scaffolded (2026-06-10 onward)
C3alphaswarm_dataA future data-plane service (ingestion, discovery, catalog) is the largest-blast-radius extraction; the RESTRUCTURING_PLAN sequences it last, after cells and per-tenant object storage.Repo created (2026-06-19 onward) and under active development, ahead of this ADR's original deferral

Existing placeholder repos alphaswarm_research ("Services for Research Plane") and alphaswarm_learning are available landing zones should the research/learning planes later split; no work is proposed for them in this ADR.

Target topology​

Consequences​

Positive

  • Tenant isolation, blast-radius reduction, and independent scaling are achieved by cells + queues — the actual goals usually cited for microservices — without breaking the ledger, replay, kill-switch, or hash-lock invariants.
  • Every extraction reuses a contract that already exists (spec runtimes, MCP audiences, /manage/*, operator CRDs), so no new RPC framework or saga/outbox machinery is invented. Linkerd arrives in Phase 4 for cell mTLS, not for inter-domain RPC.
  • Track A is pure deployment work and reversible per unit.

Negative / risks

  • Per-cell data planes multiply infra cost (mitigated by tiering: shared backplane for shared-std/shared-prem).
  • Track B2's ledger broker adds an HTTP hop to RL/ML run bookkeeping; it must remain async/buffered to keep training loops unaffected.
  • The dual-write migration window (Phase 6) is the single riskiest operation; the rollback path is the ALPHASWARM_CELL_DUAL_WRITE flag.
  • Track C repos (alphaswarm_kb, alphaswarm_kb_federation) have since been created and are under active development; KB code paths in the monolith remain conditional pending full integration.

Explicitly rejected

  • Per-domain services with per-service databases (Option 1).
  • Extracting router_complete into an LLM-gateway service.
  • Splitting the Postgres ledger or the Alembic chain per service.
  • Moving Celery beat scheduling out of the single scheduler.

Rollout order​

  1. Track B4 (CI boundary gates) and Track A1–A2 (beat + MCP pods).
  2. Track A3–A4 (queue fleets, paper/ingester units) and B1 (sidecar CP).
  3. Track A5 cells: Phase 3 (router + registry), then Phase 6 (data plane), per the stop conditions in RESTRUCTURING_PLAN §18.2.
  4. Track B2 (ledger broker) — gate for any future out-of-monolith RL/ML workers.
  5. Track C — only after repository approval.