Skip to main content

Runbook - DataOps Airbyte and Dagster

Local agent interface: Local agent-first data-infra ops (alphaswarm-cli data infra, data.dagster.* MCP, orchestration-operator).

Use this runbook for incidents on the managed data movement and transformation path: Airbyte syncs, Dagster jobs/assets/schedules/sensors, dbt mesh runs, and the unified /workflows/data operator view.

First Checks​

  1. Open /workflows/data and review failures, stale datasets, recent runs, and the Launch Ops action.

  2. Query the operations aggregation endpoints:

    curl -sf "$ALPHASWARM_API_BASE/data-control/operations/summary"
    curl -sf "$ALPHASWARM_API_BASE/data-control/operations/runs?limit=25"
    curl -sf "$ALPHASWARM_API_BASE/data-control/operations/schedules"
    curl -sf "$ALPHASWARM_API_BASE/data-control/operations/freshness?stale_after_hours=24"
  3. Confirm Airbyte and Dagster pods are healthy:

    kubectl -n alphaswarm-elt get pods
    kubectl -n alphaswarm-mlops get pods -l component=dagster
  4. Check trace correlation. Airbyte run summaries should include duration_seconds, observe_trace_id, and dagster_run_id when those links are available. Follow observe_trace_id in Grafana/Tempo.

Airbyte Sync Failure​

Symptoms:

  • /airbyte/runs or /data-control/operations/runs?kind=airbyte shows a failed run.
  • alphaswarm.airbyte.* spans report error status.
  • A dataset becomes stale after its expected sync window.

Actions:

  1. Find the run row and capture airbyte_connection_id, airbyte_job_id, observe_trace_id, dagster_run_id, records_synced, and error.

  2. Inspect the Airbyte worker logs:

    kubectl -n alphaswarm-elt logs deploy/airbyte-worker --tail=200
  3. Verify the connection exists in the Airbyte workspace configured by airbyte_workspace_id; do not use legacy airbyte_host settings.

  4. If the job reached the destination, confirm lineage preserved records_synced. A zero value with nonzero Airbyte records usually means the Airbyte lineage callback or run summary projection regressed.

  5. Relaunch through Dagster when orchestration should own the retry:

    curl -sf -X POST "$ALPHASWARM_API_BASE/dagster/launch" \
    -H 'content-type: application/json' \
    -d '{"job_name":"platform_ops_job","tags":{"incident":"airbyte-sync"}}'

Dagster Daemon, Schedule, Or Sensor Failure​

Symptoms:

  • /data-control/operations/schedules shows a stopped or failed schedule.
  • Dagster UI shows daemon heartbeats missing.
  • Sensors do not launch runs for changed pipeline manifests.

Actions:

  1. Confirm both code locations are loaded: alphaswarm.dagster.definitions and pipelines.dagster_user_code.definitions.

  2. Inspect daemon and user-code logs:

    kubectl -n alphaswarm-mlops logs deploy/dagster-daemon --tail=200
    kubectl -n alphaswarm-mlops logs deploy/alphaswarm-monolith-user-code --tail=200
    kubectl -n alphaswarm-mlops logs deploy/pipelines-user-code --tail=200
  3. For manifest-triggered runs, launch the scoped materialization job rather than a broad full refresh:

    curl -sf -X POST "$ALPHASWARM_API_BASE/dagster/launch" \
    -H 'content-type: application/json' \
    -d '{"job_name":"pipeline_manifest_materialization_job","run_config":{"ops":{"pipeline_manifest_materialization":{"config":{"manifest_id":"<manifest-id>"}}}}}'
  4. If the failure is a dbt snapshot pool deadlock, switch to dbt snapshot deadlock.

Data Lag​

Symptoms:

  • /data-control/operations/freshness returns stale datasets.
  • alphaswarm.dataops.freshness_lag_seconds exceeds the SLO threshold.
  • Dashboards show stale feature, catalog, or QuestDB-backed datasets.

Actions:

  1. Identify the owning system from the dataset row: dagster_asset_key, datahub_urn, provider, and domain.
  2. Check whether Airbyte movement, dbt transform, or downstream materialization is the oldest failed segment in /data-control/operations/runs.
  3. If QuestDB writers are blocked, switch to QuestDB WAL stall.
  4. Relaunch only the impacted asset/job from /dagster/launch; avoid full refreshes unless the incident commander approves the broader blast radius.

dbt Transformation Failure​

Symptoms:

  • kind=dbt runs show failed status.
  • Dagster user-code load fails while parsing dbt mesh assets.
  • Transformation SLO alerts mention alphaswarm.dbt.* attributes.

Actions:

  1. Confirm the dbt project has a valid manifest:

    dbt parse --project-dir <project-dir> --profiles-dir <profiles-dir>
  2. Inspect the Dagster user-code logs for the affected code location.

  3. Check resource pools/tags for QuestDB writers, snapshots, and vendor-limited work before retrying.

  4. Relaunch the targeted Dagster job or asset selection through /dagster/launch after the manifest parses cleanly.

Escalation Packet​

Include these fields in the incident ticket:

  • Failing endpoint or UI surface.
  • Airbyte connection/job id, Dagster run id, dbt invocation id, workflow run id.
  • observe_trace_id and Grafana/Tempo link.
  • Freshness lag, expected sync interval, and last successful run timestamp.
  • Whether Terraform apply/destroy was requested. Apply/destroy must remain behind the existing step-up and secure runtime controls.

The first DataOps rollout left non-blocking validation debt in the legacy data-layer tests and docs validation gate. Track remediation in DataOps Airbyte + Dagster residual debt.