Skip to main content

Customer offboarding and teardown

Procedure for winding down an enterprise customer: contract termination, license revocation, governed infrastructure destroy, data retention, and org suspension — in that order. Destroy is the highest-blast-radius operation in the platform; it only ever runs as a plan-bound destroy through the controller gate (a destroy without a reviewed destroy-plan binding is rejected).

Prerequisites​

  • Commercial confirmation of the end date + the data-retention terms from the contract (export window, deletion deadline).
  • You hold manage:tenants + manage:infrastructure + terraform:admin; second approver available.
  • Customer notified of the cutoff schedule.

Steps​

  1. Lifecycle → offboarding on the customer record (freezes the commercial picture; audit anchor for everything below).
  2. Terminate the contract — contracts table → Terminate. Resolved entitlements empty out; new license issuance for this customer stops (issuance copies entitlements from the active contract).
  3. Revoke licenses — deployment-scoped revoke for every active lease (see license-issuance-and-revocation). The install decays through expiry + grace; time the destroy after the export window, not the grace window.
  4. Data export (if contracted) — snapshot the tenant KB silo from the customer account (RDS snapshot + S3 sync) to the agreed hand-off location before any destroy.
  5. Destroy plan — Terraform page → the customer-<slug>-<env> workspace → run plan -destroy. Review the full resource list. The OPA customer gate requires stateful deletes to carry alphaswarm:teardown=approved — tag the stateful resources via the reviewed teardown change first (that tag change is itself a plan).
  6. Four-eyes destroy apply — the destroy binding pins the reviewed destroy plan (PlanBinding.is_destroy); the second operator approves; apply executes exactly that plan.
  7. Sync + close the record — deployment row destroying → destroyed. PATCH remaining metadata (final run ids stay linked for audit).
  8. Suspend the org — org status → suspended (sign-ins stop; data rows remain for the retention window). After the contractual deletion deadline: purge per the tenant-purge flow (/tenants/purge).
  9. Lifecycle → churned.

Post-action verification​

  • Customer AWS account: no alphaswarm-* resources remain (module outputs empty; spot-check RDS/S3/KMS in the console).
  • Ops account: the state object still exists (retained — it is the audit record of what was destroyed); the workspace row is archived.
  • Registry: all leases revoked; audit ledger shows the full chain (terminate → revoke → destroy plan → approve → destroy).
  • Org suspended; customer lifecycle churned.

Escalation​

  • Destroy blocked by OPA on stateful resources → that is the gate working; complete the teardown-tag change, never bypass.
  • Partial destroy (dependency cycle) → re-plan destroy; the remaining graph shrinks each pass. Manual console deletes are the last resort and must be reconciled with terraform state rm + documented.
  • Customer disputes deletion after the fact → the retained state file + audit chain is the evidence trail.

Game-day: BYOC tenant offboard with crypto-shred verification​

Phase 4.6(d) game-day from the Unified Infrastructure Control Plane Plan (private alphaswarm_internal planning repo) (§4, item 4.6). This extends the teardown procedure above with the terminal crypto-shred step and its unrecoverability proof, rehearsing the offboarding saga's compensation model from ADR 026 (§7) and the silo-reg BYOK/Transit key custody from ADR 027 — Cell isolation tiers.

Objective​

Prove that offboarding a BYOC / silo-reg tenant ends in a crypto-shred that renders the tenant's data unrecoverable — the tenant BYOK CMK is destroyed — and that this is a human-approved terminal compensation, not an automatic saga side-effect. Success = the CMK is destroyed and a subsequent read/decrypt of the tenant's at-rest data fails, with the full offboarding saga audit trail as evidence.

Grounded mechanisms​

  • Offboarding saga: the Tenant Provisioning Saga "offboarding reverses with crypto-shredding as terminal compensation" (plan Phase 2.1). Steps run on the in-process compensating saga engine; the runner consults the halt store before every step and refuses destructive steps without an approval_request_id (pythonic-unified P-7).
  • Crypto-shred guard: crypto-shredding is a "terminal, human-approved compensation" and, per the guard, Transit keys are never deleted by sagas — "Transit keys/CMKs are never deleted by an automatic path". The shred destroys the tenant BYOK CMK (tenant_byok_cmk, the AWS-KMS key that encrypts the tenant's RDS/S3 at rest); the per-cell Vault Transit key is retained (the guard protects it from saga deletion). Anchor: plan §3 component map ("crypto-shred guard: Transit keys never deleted by sagas"), ADR 027 Tier-3 silo-reg (Vault Transit + BYOK CMK).
  • Unrecoverability: with the CMK destroyed, envelope keys can no longer be unwrapped, so the encrypted RDS snapshot / S3 objects cannot be decrypted — the data is cryptographically shredded even though ciphertext may still exist.

Preconditions / scope​

  • A rehearsal BYOC / silo-reg tenant in a disposable account (never a live customer). Its data plane is encrypted with a dedicated BYOK CMK and, where applicable, a per-cell Vault Transit key (ADR 027 Tier 3).
  • The commercial reversal steps (terminate → revoke → plan-bound destroy → suspend) from the procedure above are complete or rehearsed to the point where the CMK destroy is the remaining terminal step.
  • A pre-recorded read/decrypt probe against the tenant's at-rest data that succeeds before the shred (so the post-shred failure is a controlled before/after).
  • A signed approval_request_id for the destructive crypto-shred step; four-eyes and step-up available.

Roles​

RoleResponsibility
Offboarding lead (SRE)Runs the reverse saga to the terminal step
Second approverProvides the distinct four-eyes approval + approval_request_id
Security officerWitnesses the CMK destruction and the failed decrypt probe
ScribeCaptures the saga step audit rows and the before/after probe results

Step-by-step procedure​

  1. Complete the reversal. Run the offboarding steps above (terminate → revoke → data export if contracted → plan-bound destroy → suspend) against the rehearsal tenant, up to the point the CMK remains.
  2. Baseline the probe. Run the read/decrypt probe against the tenant's encrypted RDS snapshot / S3 objects and confirm it succeeds (data is recoverable while the CMK lives). Record the result.
  3. Assert the guard (negative sub-case). Attempt the crypto-shred without an approval_request_id; confirm the saga runner refuses it (destructive step needs approval) — proving the terminal compensation cannot fire automatically.
  4. Approve + execute the crypto-shred. With four-eyes + step-up, supply the approval_request_id and execute the terminal compensation: schedule destruction of / disable the tenant BYOK CMK (tenant_byok_cmk).
  5. Confirm Transit-key retention. Verify the per-cell Vault Transit key was NOT deleted — the crypto-shred guard retains it. Only the tenant CMK is destroyed.
  6. Prove unrecoverability. Re-run the probe from step 2; the decrypt/read MUST now fail (KMS key unavailable → envelope keys cannot be unwrapped).
  7. Verify the saga audit trail. Confirm every step (terminate → revoke → destroy plan → approve → destroy → crypto-shred) landed a ledger row, and that the crypto-shred row carries the approval_request_id.
  8. Document in the post-action verification section above.

Success criteria (measurable)​

  • The tenant BYOK CMK is destroyed (pending-deletion / disabled), and the per-cell Vault Transit key is retained (guard honoured).
  • The decrypt/read probe succeeded before and failed after the shred — a controlled proof of unrecoverability.
  • The unapproved crypto-shred attempt was refused (no destructive step without approval_request_id).
  • The offboarding saga audit trail is complete and the crypto-shred row is bound to its approval.

Abort / rollback​

  • Crypto-shred is irreversible by design — once the CMK is destroyed the data is unrecoverable. The rollback window is before step 4: if the export or retention terms are not fully satisfied, halt the saga (the runner consults the halt store before every step) and do not approve the shred.
  • If the CMK is only scheduled for deletion (KMS pending-deletion window), cancellation during that window is the sole recovery path; after the window closes there is none.
  • Never destroy the Vault Transit key as part of the shred — the guard forbids it, and doing so would exceed the terminal compensation's scope.

Evidence capture (WORM / audit ledger)​

  • The full offboarding saga audit trail: terminate → revoke → destroy plan → approve → destroy → crypto-shred, each a hash-chained ledger row.
  • The crypto-shred row with its bound approval_request_id and the CMK id/ARN (never key material).
  • The before/after decrypt-probe results and the retained-Transit-key confirmation.
  • Rows fan out to the monolith Postgres ledger and archive to S3 Object Lock COMPLIANCE WORM; verify the chain with python -m alphaswarm_controller.terraform.audit_verify (exit≠0 on a break — see the DR replay game-day).

Frequency / owner​

  • Frequency: semi-annually and on every real BYOC offboarding (the real offboarding is the exercise, with the same evidence bar).
  • Owner: sre-team, security officer co-signs the unrecoverability proof.