Portable orchestration rollout and rollback runbook
Use this runbook to deploy the dual-engine control planes and migrate only the representative first wave. Read the canonical guide and ADR 025 before changing ownership.
Stop condition: never continue if a migrated schedule has zero or more than one active owner, context verification fails, engine state cannot be reconciled, or compatibility tests fail with flags off.
Scope
This runbook covers:
system.context_echoon both engines;data.pipeline_manifest_materialization, default Dagster;agents.workflow_runtime, default Prefect; and- schedule ownership handoff for those definitions only.
All other Celery and Argo paths remain unchanged.
Required access
- cluster or Compose operator access;
- AlphaSwarm
orchestration:readand write scopes; - step-up authorization for cancel, retry, and destructive trigger changes;
- database migration access through the existing migration job;
- read access to Dagster and Prefect internal UIs for incident diagnosis; and
- access to MinIO, OTel, and audit evidence through governed operator paths.
Do not connect to Dagster or Prefect databases for application-level reads.
Pre-deployment evidence
Record the following in the change ticket:
| Evidence | Required value |
|---|---|
Reviewed alphaswarm_orchestration SHA | Exact commit, not a mutable branch |
| Dagster versions | 1.13.13 and integrations 0.29.13 |
| Prefect versions | 3.7.8 and Kubernetes integration 0.7.10 |
| Alembic state | Includes 0124_portable_orchestration_ledger, 0130_session_distributed_execution, and 0131_session_resource_operations |
| Legacy task names | Passing with portable flags off |
| Queue contract | agents/workflows mismatch resolved and verified |
| Existing owners | Celery Beat entries and all native triggers inventoried |
| Rollback owner | Named operator with step-up access |
| Acceptance pool | Dedicated Prefect pool and worker; no production template mutation |
| Acceptance image | Reviewed image imports the pinned orchestration package without runtime installation |
| Session providers | Enabled CRDs/operators, namespace allowlist, images, resource profiles, and service accounts |
Snapshot current active tasks, Celery Beat entries, Dagster schedules/sensors, Prefect deployments/automations, and reconciliation cursors. The snapshot is evidence only; do not copy native database tables.
Validate documentation and deployment artifacts
From alphaswarm_docs:
pnpm diagrams:validate
pnpm diagrams:check
python3 -m pytest -q scripts/tests/test_orchestration_diagrams.py
From alphaswarm_platform, render the orchestration profile with the base
network definition and real secret-file paths:
docker compose \
-f deployments/compose/docker-compose.base.yml \
-f deployments/compose/docker-compose.orchestration.yml \
--profile orchestration-smoke config >/dev/null
kubectl kustomize deployments/kubernetes/mlops/prefect >/dev/null
kubectl kustomize deployments/kubernetes/mlops/dagster >/dev/null
Run the platform contract tests and pinned Dagster chart lint/template gates from the platform PR before deployment.
Apply additive persistence first
- Back up the AlphaSwarm application database through the standard database runbook.
- Apply Alembic migration
0124_portable_orchestration_ledger. - Confirm the ten orchestration tables exist.
- Confirm forced RLS is enabled on every table.
- Confirm
app_runtimecannot update or deleteorchestration_events. - Run the repository's RLS, out-of-order event, cursor, and concurrent reconciliation tests.
- If session resources are enabled, confirm all eight session-execution tables exist, are forced through RLS, and retain append-only events, immutable lease authority, monotonic fences, and database-derived consumer counts.
Do not downgrade the additive migration during a flag rollback. The tables are safe to retain as readable audit history.
Deploy native control planes disabled for customer traffic
Compose smoke environment
Provide the five required secret files, then start the smoke profile:
docker compose \
-f deployments/compose/docker-compose.base.yml \
-f deployments/compose/docker-compose.orchestration.yml \
--profile orchestration-smoke up -d
scripts/smoke/orchestration-compose.sh
The smoke script must prove PostgreSQL, Redis, both APIs, the Dagster GraphQL contract, Prefect health, representative registration, a run, cancellation, and restart reconciliation.
Kubernetes
- Confirm External Secrets are ready before applying workloads.
- Apply the Prefect Kustomize bundle.
- Reconcile the pinned Dagster HelmRelease and values.
- Wait for PostgreSQL, Redis, database migration/bootstrap jobs, Prefect API, background service, Prefect workers, Dagster webservers, daemon, and code location.
- Run
scripts/smoke/orchestration-kubernetes.sh. - Confirm native UI services are cluster-internal only.
- Confirm NetworkPolicies, service accounts, probes, resource limits, topology spread, and PodDisruptionBudgets match the reviewed manifests.
- Confirm the acceptance flow uses its dedicated work pool and that deleting the acceptance deployment cannot change the production pool template.
- Exec the reviewed worker image before running flows and import
alphaswarm_orchestration; do not install a wheel into the running pod. - Apply
deployments/kubernetes/mlops/session-runtimebefore enabling provisioned session clusters. Confirm thealphaswarm-session-runtimeRole is namespace scoped, cannot read Secrets, grants Spark only its required pod/Service/ConfigMap/PVC lifecycle, and that Ray/Dask pod templates disable service-account token automount.
The default Dagster values keep the existing Dagster database. Use
values-postgres16-cutover.yaml only under its separate database migration and
rollback procedure.
Verify the gateway before enabling compatibility routing
Using an authenticated operator session, verify:
GET /orchestration/health
GET /orchestration/workers
GET /orchestration/pools
GET /orchestration/definitions?limit=100
GET /orchestration/registrations?limit=100
GET /orchestration/triggers?limit=100
GET /orchestration/runs?limit=100
Confirm:
- unauthenticated access is rejected;
- cross-workspace reads fail even with guessed IDs;
- write endpoints require write scopes and idempotency keys;
- cancel/retry requires step-up and emits audit records;
- an expired or tampered context envelope is rejected;
- framework tags contain no secret values; and
- SSE honors
Last-Event-IDand produces the canonical progress projection.
For the SSE reconnect check, open a stream with ?after=snapshot-tail, receive
at least one newer event, and reconnect with that newer cursor in
Last-Event-ID while leaving the old query parameter unchanged. The first
replayed event must follow the header cursor: the header always wins.
Register representative definitions
For each definition:
- create or locate the immutable version;
- call validation for both engines;
- call compilation for both engines;
- review every diagnostic;
- register the selected engine with its native trigger disabled;
- register the alternate engine only when required for parity testing; and
- persist engine/native bindings and confirm they rehydrate after an API restart without a duplicate native registration.
Required defaults:
| Definition | Selected default | Alternate proof |
|---|---|---|
system.context_echo | Neither; explicit test selection | Execute once on each engine |
data.pipeline_manifest_materialization | Dagster | Prefect execution against same manifest |
agents.workflow_runtime | Prefect | Successful Dagster compilation |
Run shadow verification
Context parity
Launch system.context_echo on Dagster and Prefect from the same authenticated
session and environment. Compare normalized outputs for:
- user, organization, team, workspace, project, lab, cell, and role;
- environment and deployment version;
- session, request, correlation, experiment, and test IDs;
- context reference and digest; and
- W3C trace and parent/child span linkage.
The normalized values must match. Native fields may differ and must remain accessible.
Manifest parity
Run data.pipeline_manifest_materialization on Dagster, then run the same
immutable version and manifest contract on Prefect. Compare:
- parameter validation;
- portable step ordering;
- canonical terminal state;
- required artifact keys;
- artifact checksums and sizes; and
- point-in-time
DataBindingprovenance.
Agent workflow parity
Run agents.workflow_runtime on Prefect with a bounded test workflow. Confirm
the same definition compiles for Dagster. Verify halt/cancel propagation,
budget guardrails, session context, and trace continuity.
Cancellation, retry, and restart drills
For one non-production run on each engine:
- launch and wait for
RUNNING; - cancel through
POST /orchestration/runs/{run_id}/cancelwith a new idempotency key and reason; - confirm the native engine enters its cancellation path;
- confirm the canonical state becomes
CANCELLEDonce native cancellation is complete; - launch a controlled failure and retry through
POST /orchestration/runs/{run_id}/retry; - confirm attempt increments once and the original engine remains retry owner;
- restart the AlphaSwarm API/control service;
- confirm registrations and runs are rehydrated from persisted bindings;
- page events across the restart and verify unique sequences and native event IDs; and
- stop the reconciler during a state transition, restore it, and verify idempotent repair from the persisted cursor.
No step may launch the same logical run in both engines.
Verify session-scoped distributed resources
Run this drill only for provider kinds enabled in the target environment. A missing CRD/operator is a disabled capability, not a reason to apply an unreviewed cluster-wide manifest during acceptance.
Before the lifecycle drill, health-check
GET /internal/session-resources/health with the configured machine bearer
credential and manage:infrastructure scope. Confirm an absent credential,
invalid credential, or missing scope is rejected, the response declares
contract version 1, and controller logs never record the credential, raw
endpoints, or secret-reference values. The monolith must use its embedded
durable manager and dependency-light HTTP provider; it must not import a
Kubernetes client or call the native provider in-process.
- Create one signed session resource spec with an idle TTL below its maximum TTL and a session-scoped cache namespace.
- Acquire it concurrently from two consumers with distinct idempotency keys.
- Confirm one native resource, one immutable lease authority, two active consumer rows, and one provider fence.
- Wait for
READY; resolve the endpoint by its symbolic reference through the authorized resolver. Confirm no URL, token, or credential appears in the spec, labels, events, or API response. - Submit a bound
WorkRequest. On Ray or Dask, prove the remote runtime inventory and context digest match before executing domain work. On Spark, prove the same signed request and point-in-time data binding reaches the application boundary. - Renew with the current lease and consumer token, then replay the same idempotency key. Confirm one provider mutation and the same durable outcome.
- Attempt renew and release with the previous fence or consumer token. Both must fail closed without mutating the native resource.
- Detach the first consumer and confirm the resource remains ready. Detach the final consumer and confirm the provider is called exactly once.
- Cancel an acquire while native creation is in flight. Confirm no late
READYcommit; the provider must observe or create a tombstone and remove the abandoned native resource before terminal cleanup. - Restart the API/coordinator during renew and release operations. Reconcile from durable reservations and events; confirm no duplicate cluster and no stale release.
- Exercise the controller's versioned provision, renew, release, and abandoned-cleanup operations through the HTTP provider. Substitute a wrong spec, lease authority, fence, or service credential and confirm each fails closed before the manager commits state.
For Kubernetes providers, additionally confirm RayCluster, DaskCluster, and
SparkConnect resources have deterministic names, tenant-safe labels,
allowlisted namespaces, pinned images/resource profiles, service accounts, and
the expected NetworkPolicy. Verify Ray head and worker placement is compatible
with the pinned image architecture; by default both require amd64. Confirm
Ray workers have no controller-authored dashboard probes, while KubeRay reports
the desired and available worker counts. Verify resourceVersion conflict
retries are bounded and provider status responses contain only allowlisted
summaries.
Close known cluster caveats
The acceptance record must explicitly cover these historical failure modes:
- Prefect cancellation: cancellation is incomplete while the Kubernetes
Job continues running. The adapter must enforce Job deletion after the
supported Prefect cancellation request and reconcile canonical state to
CANCELLED; a flow that merely remainsCANCELLINGuntil natural completion fails the drill. - Cross-namespace health: validate the engine-local endpoint first, then an explicit cross-namespace probe that matches the deployed NetworkPolicy and cluster CNI behavior. Do not convert a failed probe into a permanent waiver.
- Platform dependencies: Flux, External Secrets, the expected OTel namespace/collector, and the configured secret backend must be installed and healthy for production-shaped evidence. Ephemeral secrets and non-fatal OTel warnings are local-development evidence only.
- Acceptance isolation: use a dedicated work pool and source-managed image;
never mutate
shared-generalor rely on ad-hoc package installation.
Verified Julia k3s acceptance — 2026-07-15
Platform revision 3b8c242 closed the historical caveats above on the
julia@192.168.12.112 k3s cluster:
- Dagster run
0d407566-4708-4177-8666-12b54e546950reachedSUCCESSin a newly createdK8sRunLauncherJob. The Job carriedalphaswarm.io/acceptance-owner=dagster-acceptance, ran onaqp-tower, and the isolated release contained no daemon. The acceptance instance usedSyncInMemoryRunCoordinator, so it never consumed the production queued-run coordinator. - Prefect run
53e9b058-3491-489c-abbf-57b293cb400dreachedCOMPLETEDand emitted one native artifact. Its Job used the immutable reviewed image, mountedprefect-kubernetes-smoke-flowat/opt/alphaswarm-smoke, and ran onaqp-tower. - Active Prefect run
1e42e77e-08f6-4da8-8a3f-0aeb5749617cwas cancelled through the adapter. The canonical/native state reachedCANCELLEDand Jobheavy-bird-4vv7rwas absent immediately after enforcement; the test did not wait for natural flow completion. - The Prefect observer was explicitly scoped to the workload namespace, which matched its namespace-scoped Role and produced no cluster-scope authorization errors.
- Flux, External Secrets, the OpenTelemetry gateway, Dagster, and Prefect all passed the post-run health gate. The cross-namespace OTLP probe succeeded, so the kube-router waiver path was not exercised.
- Cleanup removed the isolated Helm release, worker, work pool, queue,
deployment, ConfigMaps, and owned Jobs while preserving
dagster-kubernetes-smoke-definitions. Production replicas were restored to Prefect API/background/worker2/1/2and Dagster webserver/daemon/code location2/1/1, with every Deployment available.
Reproduce the profile with
alphaswarm_platform/scripts/smoke/orchestration-kubernetes-execution.sh --live; always follow with --cleanup and the post-restore health gate.
Transfer one schedule owner
Use a transaction or otherwise serialized control path that enforces the partial unique active-owner constraint.
- Confirm the selected native trigger exists and is disabled.
- Confirm Celery Beat is the sole active owner.
- Disable the one matching Celery Beat entry.
- Confirm there is temporarily no active producer and no queued duplicate.
- Enable the selected Dagster schedule/sensor or Prefect schedule/automation.
- Set the normalized trigger binding
active_owner=true. - Query all ownership sources and confirm exactly one owner.
- Observe through at least one expected fire window.
- Confirm one logical run, one native run binding, and one retry owner.
Never enable the native trigger before disabling its Celery entry.
Enable compatibility routing
Only after representative acceptance passes:
- enable the smallest definition-scoped rollout flag;
- keep general Celery compatibility enabled;
- call the existing
/workflowsroute and stable task name; - verify the response preserves
TaskAccepted, task ID, stream URL, replay, and halt semantics; and - verify the normalized run records the selected native engine.
Do not enable a broad default-engine cutover as part of the first wave.
Observation window
Monitor:
/orchestration/health, worker and pool health;- normalized active runs by workspace, state, environment, and age;
- native Dagster and Prefect run state;
- event ingestion lag and reconciliation cursor age;
- duplicate logical/native bindings;
- schedule fire count and ownership;
- MinIO checksum or write failures;
- SSE disconnect/replay errors;
- worker deadline, cancellation, and idempotency failures; and
- OTel error rate, latency, queue wait, and trace continuity.
Treat the native control plane as state authority and the normalized ledger as a repairable projection. A projection freshness failure does not justify editing native tables.
Flag-first rollback
Trigger rollback on duplicate scheduling, duplicate retry, context mismatch, unrecoverable reconciliation lag, native API incompatibility, cancellation failure, cross-tenant exposure, or failed compatibility behavior.
- Disable new portable registrations and launches.
- Disable the affected native trigger.
- Confirm
active_owner=falsefor that binding. - Restore the legacy Celery Beat entry.
- Confirm Celery is the only active owner.
- For active native runs, choose one authorized action:
- allow safe completion; or
- cancel through the AlphaSwarm gateway with step-up and audit.
- Keep additive ledger rows and native payload references readable.
- Keep both native control planes running long enough to reconcile final state.
- Re-run compatibility tests with flags off.
- Attach event cursors, audit IDs, run IDs, native IDs, and OTel trace IDs to the incident ticket.
Do not delete normalized history, force engine database state, or downgrade the ledger migration as an incident shortcut.
Incident triage matrix
| Symptom | First checks | Safe action |
|---|---|---|
| Duplicate scheduled runs | Celery Beat entry, native trigger, active_owner | Disable the newest owner, stop launches, rollback ownership |
| Normalized state is stale | Native run, cursor age, adapter health | Restart reconciler and rehydrate bindings |
| Native run is missing | Registration ID, native binding, audit event | Stop retries; reconcile before deciding rollback |
| Events repeat after restart | Persisted adapter cursor and normalized high-water mark | Stop projection, preserve cursor evidence, repair adapter contract |
| Context rejected | Key ID, expiry, digest, trusted reference | Do not bypass verification; relaunch with a fresh context |
| Secret appears in metadata | Tags, summaries, logs, payload externalization | Stop launches, rotate exposed secret, preserve audit evidence |
| Cancel remains pending | Native engine state and worker Job | Enforce adapter Job termination, then reconcile through supported APIs |
| Prefect worker creates no Job | pool/queue, service account, context verification | Repair configuration; do not submit an unverified Job manually |
| Session manager is unavailable | embedded manager health, Postgres/RLS, controller M2M endpoint | Stop provisioning; repair durability or authenticated provider health before retry |
| Controller rejects lifecycle call | service credential, contract version, spec/lease authority, fence | Do not bypass auth or mutate Kubernetes directly; reconcile the durable reservation |
| Resource becomes ready after cancellation | reservation token, tombstone, native object | Fence the late completion and run cleanup before terminal commit |
| Stale lease mutates a cluster | fence version, endpoint authority, resourceVersion | Reject the operation; reconcile the current lease without deleting the resource |
Completion record
The change ticket must include:
- exact package and deployment SHAs;
- definition content hashes and engine registration IDs;
- logical and native run IDs for every parity test;
- artifact checksums;
- context digest and trace IDs, never secret values;
- cancellation, retry, restart, reconciliation, and rollback results;
- schedule owner before, during, and after cutover;
- links to audit events and dashboards; and
- the operator who approved each destructive action.
Mark the migration complete only when all acceptance criteria in ADR 025 pass and the observation window contains no duplicate schedule, retry, or cross-tenant event.