Runbook — Local agent-first data-infra ops
Use this runbook when you or an agent need to inspect or lightly
operate data infrastructure from a local shell: native Dagster
GraphQL health, failed runs, schedules/sensors, local Definitions
validation, and portable orchestration health. This is not the
hosted Dagster UI, not alphaswarm_ops_console, and not a
Terraform or Helm apply path.
Related:
- DataOps Airbyte and Dagster — incident triage for syncs, daemon failures, and dbt mesh.
- Portable orchestration — dual-engine rollout, cutover, and rollback.
- Portable orchestration concepts
- Implementation plan (internal):
alphaswarm_internal/plans/2026-08-14-local-agent-data-infra-ops.md
Operator subagent: alphaswarm_orchestration/.cursor/agents/alphaswarm-orchestration-operator.md
Which surface am I on?
AlphaSwarm exposes three orchestration names that must not be mixed in flags, MCP tool choice, or incident tickets.
| Noun | Meaning | Surface |
|---|---|---|
| Native Dagster | Classic Definitions in a named code location | GET /dagster/*, alphaswarm-cli data infra *, data.dagster.*, python -m dagster ... |
| Portable orchestration | Engine-owned TaskDefinition / run projection | GET /orchestration/*, data.orchestration.portable_* |
| Agent workflows | WorkflowRuntime / workflow_runs watchdog | data.orchestration.health, data.orchestration.list_runs (not Dagster) |
data.automation.* is Celery beat, not Dagster. Do not route Dagster
schedules through it.
Global constraints for this interface:
- Pin
dagster==1.13.13. Local validation uses the official CLI only:python -m dagster definitions validate -m <module>. - No
dg/dagster-dg-cli/ remotedg launchworkflows. - No Terraform apply from
alphaswarm-cli data infra. Cluster mutation stays onTerraformRuntime/WorkloadRuntime(One Gate). - No merged code locations — three live identities stay separate (below).
First-slice commands
Authenticate once, then use native Dagster verbs (GraphQL via monolith HTTP) or portable health. All examples are PowerShell.
alphaswarm-cli auth login
# Native Dagster — which code location is loaded?
alphaswarm-cli data infra status
# Failed-run triage
alphaswarm-cli data infra runs list --status FAILURE
alphaswarm-cli data infra runs show <run_id>
# Schedules and sensors (start/stop need manage:infrastructure)
alphaswarm-cli data infra schedules list
alphaswarm-cli data infra schedules start daily_full_refresh
alphaswarm-cli data infra sensors list
# Local Definitions load check (subprocess; does not hit remote GraphQL)
alphaswarm-cli data infra validate --module alphaswarm.dagster.definitions
alphaswarm-cli data infra validate --location monolith
alphaswarm-cli data infra validate --location platform
# Static catalog (no GraphQL)
alphaswarm-cli data infra locations
# Portable orchestration engine health (not native Dagster GraphQL)
alphaswarm-cli data infra portable-health
# Composite: Dagster status + portable health + control-plane topology hint
alphaswarm-cli data infra health
# Portable pools/workers (not native Dagster pool YAML)
alphaswarm-cli data infra pools
alphaswarm-cli data infra workers
# Human cancel on portable runs (step-up MFA on the API)
alphaswarm-cli data infra cancel <run_id> --reason "operator stop"
| Command | HTTP or local |
|---|---|
data infra status | GET /dagster/status |
data infra runs list [--status] [--limit] | GET /dagster/runs |
data infra runs show RUN_ID | GET /dagster/runs/{id} |
data infra schedules list | GET /dagster/schedules |
data infra schedules start NAME | POST /dagster/schedules/{name}/start |
data infra schedules stop NAME | POST /dagster/schedules/{name}/stop |
data infra sensors list | GET /dagster/sensors |
data infra sensors start NAME | POST /dagster/sensors/{name}/start |
data infra sensors stop NAME | POST /dagster/sensors/{name}/stop |
data infra validate [--module] [--location] | local python -m dagster definitions validate -m |
data infra locations | static JSON catalog (monolith / platform / bootstrap) |
data infra portable-health | GET /orchestration/health |
data infra health | composite (above + topology hint) |
data infra pools | GET /orchestration/pools |
data infra workers | GET /orchestration/workers |
data infra cancel RUN_ID | POST /orchestration/runs/{id}/cancel |
For raw kubectl/docker health on Dagster pods, use
alphaswarm-cli services (HTTP to control plane) — not data infra.
Validate each live module separately
Allowed --module values (or --location monolith / --location platform) for
data infra validate:
alphaswarm.dagster.definitions— monolith compose / local dev default.pipelines.dagster_user_code.definitions— platform Helm user-code deployment. May requirePYTHONPATHincludingalphaswarm_platformwhen run from a checkout that does not install that package.
alphaswarm-cli data infra validate --module alphaswarm.dagster.definitions
# Platform user-code (set PYTHONPATH when the module is not on the default path)
$env:PYTHONPATH = "C:\path\to\alphaswarm_platform"
alphaswarm-cli data infra validate --module pipelines.dagster_user_code.definitions
Not valid locally: bootstrap, bootstrap-user-code, or --module file.
The Helm ConfigMap dagster-bootstrap-user-code is cluster-only; the CLI
exits 2 with a clear message.
If validate exits non-zero while GraphQL /dagster/status is reachable, check
whether Definitions load is blocked by a separate hardening item (for example
WorkflowConfig in the 2026-08-13 Dagster expand plan). This runbook does not
replace that fix.
Start/stop scopes (manage:infrastructure)
| Action | Scope | Step-up |
|---|---|---|
GET /dagster/* reads (status, runs, schedules, sensors, assets) | read:infrastructure | No |
POST /dagster/schedules|sensors/{name}/start|stop | manage:infrastructure | No |
GET /orchestration/* reads | orchestration:read + tenancy headers | No |
POST /orchestration/runs/{id}/cancel | orchestration:execute | Yes (RFC 9470) |
CLI schedule/sensor toggles call the monolith with your bearer token. A 403
means the principal lacks manage:infrastructure.
Data MCP (data.dagster.*)
Agents use the Data MCP catalog (stdio or /mcp/data), not direct GraphQL.
Reads — scope read:infrastructure, tenancy_posture=shared:
data.dagster.statusdata.dagster.list_runs(limit, optionalstatus)data.dagster.get_run(run_id)data.dagster.list_schedulesdata.dagster.list_sensors
Writes — scope manage:infrastructure, mutates=True:
data.dagster.start_schedule/stop_scheduledata.dagster.start_sensor/stop_sensor
When ctx.actor_kind == "agent", writes queue pending_approval via the
existing MCP approval helper — they do not fire immediately.
Portable tools remain on data.orchestration.portable_* (compile, launch,
runs, pools, workers). Do not confuse them with native Dagster tools.
Cancel = CLI/API + step-up; MCP cancel is non-autonomous
Humans cancel portable runs with:
alphaswarm-cli data infra cancel <run_id> --reason "incident" --idempotency-key <uuid>
The API requires fresh step-up MFA. On 401 with
WWW-Authenticate: Bearer error="insufficient_user_authentication", the CLI
prints step-up required; run alphaswarm-cli auth login and exits 2. It
does not auto-retry.
Agents must not cancel via MCP: data.orchestration.portable_cancel
remains non-autonomous — it refuses with MFA/step-up guidance and never calls
POST /orchestration/runs/{id}/cancel. Escalate to a human operator.
Terraform apply/destroy and workload halt paths are unchanged and out of scope for this CLI group.
Pools: default_limit vs invalid named config
On Dagster 1.13.13:
concurrency.pools.default_limiton the instance is the runtime cap for tagged work.concurrency.pools.config.<name>(named pool config blocks) is invalid on this pin — do not add them.- Vendor
pool=tags on ops/assets are metadata only; they are not a separate runtime limit.
alphaswarm-cli data infra pools reads GET /orchestration/pools (portable
projection). Help text on that command states the above. Doc-only 1/2 limits in
YAML remain comments until a future pin proves named pool config.
Three code locations — do not merge
Never invent one workspace.yaml or combined module that loads all three:
| Identity | Module / artifact | Typical deployment |
|---|---|---|
| Monolith | alphaswarm.dagster.definitions | Compose dagster dev -m ... |
| Platform user-code | pipelines.dagster_user_code.definitions | Helm pipelines-user-code |
| Bootstrap | ConfigMap dagster-bootstrap-user-code | Helm bootstrap-user-code |
data infra status reports code_location and module_path for the GraphQL
instance you are querying — do not assume Helm user-code is the same module as
local compose without checking.
The operator UI Data → Infra hub shows a read-only Code locations banner:
three chips (monolith / platform / bootstrap) with the chip matching
module_path marked active, plus a topology mismatch hint when control-plane
services disagree with GraphQL.
data infra health JSON adds active_location and location_warning using the
same identity rules as the UI banner.
Sandbox ≠ prod
Interactive Dagster sandbox sessions
(Dagster sandbox concept; AGENTS hard
rule 32; /dagster/sandbox/*) are isolated per session (tempdir, Redis
namespace, safe env overrides). They are not the deployed GraphQL instance
behind /dagster/status.
- First-slice
data infraverbs talk to deployed native Dagster GraphQL. - Sandbox experimentation stays on
/dagster/sandbox/*and the sandbox UI/API; there are nodata infra sandbox-*CLI verbs in this slice. - Do not treat sandbox validate results as production health.
When to invoke which agent
| Situation | Agent |
|---|---|
| Dagster defs load failure, schedule/sensor diagnosis, code-location identity, GraphQL vs validate mismatch, sandbox isolation, portable vs native confusion | alphaswarm-orchestration-operator — expect Diagnosis / Evidence / Safe next action / Validation command |
| Live cluster apply (Terraform, Helm, workload scale/restart), step-up-gated infra mutation | alphaswarm-management-engine — One Gate only |
LangGraph / WorkflowRuntime / agent crew stalls | alphaswarm-agentic-stack-expert |
| Celery beat / Redis progress / queue depth | alphaswarm-queue-cache-operator |
The orchestration-operator diagnoses and proposes; it does not replace
alphaswarm-cli data infra or MCP tools shipped in this plan.
Operator UI (alphaswarm_client)
Phase 2a ships a read-only Data → Infra hub in the Vite operator UI
(alphaswarm_client, route /data/infra). It mirrors the first-slice CLI reads
— no schedule/sensor toggles and no portable cancel in this slice (those land in
Task 7).
| Panel | Backing API |
|---|---|
| Composite health | GET /dagster/status + GET /orchestration/health |
| Failed runs (default filter) | GET /dagster/runs?status=FAILURE |
| Run detail | GET /dagster/runs/{id} |
| Schedules / sensors | GET /dagster/schedules, GET /dagster/sensors |
| Portable pools / workers | GET /orchestration/pools, GET /orchestration/workers |
| Code locations banner | static catalog + active chip from GraphQL module_path |
Local dev (from an alphaswarm_client checkout):
pnpm --dir alphaswarm_client dev --host 0.0.0.0 --port 3001
# Browser: http://localhost:3001/data/infra
Requires a logged-in session with read:infrastructure (and tenancy headers the
BFF forwards). Mutations stay on CLI/API until the Phase 2b UI slice.
CI verification
GitHub Actions runs path-filtered jobs when data-infra surfaces change:
| Repo | Workflow | Focus |
|---|---|---|
alphaswarm | .github/workflows/data-infra-ops-ci.yml | Dagster routes + Data MCP ops tools |
alphaswarm_cli | .github/workflows/data-infra-ops-ci.yml | data infra CLI + smoke |
Reproduce locally before opening a PR (PowerShell, from each repo root with dev deps installed):
# Monolith (sibling packages editable per AGENTS.md agent-bootstrap)
python -m pytest -q `
tests/api/test_dagster_routes.py `
tests/data/mcp/test_dagster_ops_tools.py `
tests/data/mcp/test_orchestration_tools.py
# CLI (after pip install -e ../alphaswarm_config ../alphaswarm_core and pip install -e ".[dev]")
python -m pytest -q tests/test_data_infra.py tests/test_cli_smoke.py
These are focused gates — they do not replace full make agent-verify or
the monolith test-monolith job.
Infrastructure provisioning (One Gate — not data infra)
Cluster provisioning, Helm rollouts, and Terraform apply/destroy never go
through alphaswarm-cli data infra. Use the control-plane One Gate surface
(TerraformRuntime / AGENTS rule 42):
alphaswarm-cli auth login
# Read workspaces and recent provisioning runs (audit ledger)
alphaswarm-cli cp terraform workspaces
alphaswarm-cli cp terraform runs
# Plan (read-only)
alphaswarm-cli cp terraform plan --workspace-id <id> --spec-version-id <uuid>
# Apply / destroy require fresh step-up MFA on the control-plane API
alphaswarm-cli cp terraform apply --workspace-id <id> --spec-version-id <uuid>
alphaswarm-cli cp terraform destroy --workspace-id <id> --spec-version-id <uuid>
See also Incident response (terraform run
audit) and Portable orchestration (engine
cutover). For live workload start/stop/scale, invoke
alphaswarm-management-engine / WorkloadRuntime — not this runbook's
data infra verbs.
Platform deployment (three code locations)
Helm and compose artifacts live in alphaswarm_platform — docs-first only
here; no cluster mutation from this runbook.
| Identity | Module / artifact | Platform path |
|---|---|---|
| Monolith | alphaswarm.dagster.definitions | compose / monolith image (local dev) |
| Platform user-code | pipelines.dagster_user_code.definitions | deployments/kubernetes/mlops/dagster/values-pipelines-user-code.yaml |
| Bootstrap | ConfigMap dagster-bootstrap-user-code | deployments/kubernetes/mlops/dagster/user-code-bootstrap-configmap.yaml |
Canonical Helm operator notes:
alphaswarm_platform/deployments/kubernetes/mlops/dagster/README.md (section
Three Dagster code locations). Cross-check with alphaswarm-cli data infra locations (static JSON) before editing workspace YAML — identities stay
separate; do not merge repos into one workspace.yaml.
What this is not
- Not
dgor Dagster+ remote workflows. - Not a hosted UI or Dagster UI replacement.
- Not Terraform apply or kubectl exec from
data infra. - Not WorkflowRuntime health (
data.orchestration.healthis a different product surface).