Runbook - DataOps Airbyte and Dagster
Local agent interface: Local agent-first data-infra ops
(alphaswarm-cli data infra, data.dagster.* MCP, orchestration-operator).
Use this runbook for incidents on the managed data movement and
transformation path: Airbyte syncs, Dagster jobs/assets/schedules/sensors,
dbt mesh runs, and the unified /workflows/data operator view.
First Checks
-
Open
/workflows/dataand review failures, stale datasets, recent runs, and theLaunch Opsaction. -
Query the operations aggregation endpoints:
curl -sf "$ALPHASWARM_API_BASE/data-control/operations/summary"
curl -sf "$ALPHASWARM_API_BASE/data-control/operations/runs?limit=25"
curl -sf "$ALPHASWARM_API_BASE/data-control/operations/schedules"
curl -sf "$ALPHASWARM_API_BASE/data-control/operations/freshness?stale_after_hours=24" -
Confirm Airbyte and Dagster pods are healthy:
kubectl -n alphaswarm-elt get pods
kubectl -n alphaswarm-mlops get pods -l component=dagster -
Check trace correlation. Airbyte run summaries should include
duration_seconds,observe_trace_id, anddagster_run_idwhen those links are available. Followobserve_trace_idin Grafana/Tempo.
Airbyte Sync Failure
Symptoms:
/airbyte/runsor/data-control/operations/runs?kind=airbyteshows a failed run.alphaswarm.airbyte.*spans report error status.- A dataset becomes stale after its expected sync window.
Actions:
-
Find the run row and capture
airbyte_connection_id,airbyte_job_id,observe_trace_id,dagster_run_id,records_synced, anderror. -
Inspect the Airbyte worker logs:
kubectl -n alphaswarm-elt logs deploy/airbyte-worker --tail=200 -
Verify the connection exists in the Airbyte workspace configured by
airbyte_workspace_id; do not use legacyairbyte_hostsettings. -
If the job reached the destination, confirm lineage preserved
records_synced. A zero value with nonzero Airbyte records usually means the Airbyte lineage callback or run summary projection regressed. -
Relaunch through Dagster when orchestration should own the retry:
curl -sf -X POST "$ALPHASWARM_API_BASE/dagster/launch" \
-H 'content-type: application/json' \
-d '{"job_name":"platform_ops_job","tags":{"incident":"airbyte-sync"}}'
Dagster Daemon, Schedule, Or Sensor Failure
Symptoms:
/data-control/operations/schedulesshows a stopped or failed schedule.- Dagster UI shows daemon heartbeats missing.
- Sensors do not launch runs for changed pipeline manifests.
Actions:
-
Confirm both code locations are loaded:
alphaswarm.dagster.definitionsandpipelines.dagster_user_code.definitions. -
Inspect daemon and user-code logs:
kubectl -n alphaswarm-mlops logs deploy/dagster-daemon --tail=200
kubectl -n alphaswarm-mlops logs deploy/alphaswarm-monolith-user-code --tail=200
kubectl -n alphaswarm-mlops logs deploy/pipelines-user-code --tail=200 -
For manifest-triggered runs, launch the scoped materialization job rather than a broad full refresh:
curl -sf -X POST "$ALPHASWARM_API_BASE/dagster/launch" \
-H 'content-type: application/json' \
-d '{"job_name":"pipeline_manifest_materialization_job","run_config":{"ops":{"pipeline_manifest_materialization":{"config":{"manifest_id":"<manifest-id>"}}}}}' -
If the failure is a dbt snapshot pool deadlock, switch to dbt snapshot deadlock.
Data Lag
Symptoms:
/data-control/operations/freshnessreturns stale datasets.alphaswarm.dataops.freshness_lag_secondsexceeds the SLO threshold.- Dashboards show stale feature, catalog, or QuestDB-backed datasets.
Actions:
- Identify the owning system from the dataset row:
dagster_asset_key,datahub_urn, provider, and domain. - Check whether Airbyte movement, dbt transform, or downstream materialization
is the oldest failed segment in
/data-control/operations/runs. - If QuestDB writers are blocked, switch to QuestDB WAL stall.
- Relaunch only the impacted asset/job from
/dagster/launch; avoid full refreshes unless the incident commander approves the broader blast radius.
dbt Transformation Failure
Symptoms:
kind=dbtruns show failed status.- Dagster user-code load fails while parsing dbt mesh assets.
- Transformation SLO alerts mention
alphaswarm.dbt.*attributes.
Actions:
-
Confirm the dbt project has a valid manifest:
dbt parse --project-dir <project-dir> --profiles-dir <profiles-dir> -
Inspect the Dagster user-code logs for the affected code location.
-
Check resource pools/tags for QuestDB writers, snapshots, and vendor-limited work before retrying.
-
Relaunch the targeted Dagster job or asset selection through
/dagster/launchafter the manifest parses cleanly.
Escalation Packet
Include these fields in the incident ticket:
- Failing endpoint or UI surface.
- Airbyte connection/job id, Dagster run id, dbt invocation id, workflow run id.
observe_trace_idand Grafana/Tempo link.- Freshness lag, expected sync interval, and last successful run timestamp.
- Whether Terraform apply/destroy was requested. Apply/destroy must remain behind the existing step-up and secure runtime controls.
Related Debt
The first DataOps rollout left non-blocking validation debt in the legacy data-layer tests and docs validation gate. Track remediation in DataOps Airbyte + Dagster residual debt.