ADR 024 — Single-node k3s migration target before multi-node graduation
- Status: Accepted (2026-06-24) — design decision for Phase 5+; no storage or cluster mutation executed by this ADR
- Authors: Platform ops
- Related: ADR 015; Phase 0 snapshot, Phase 2 report, and the Deployment Ops Framework Plan were working documents from the deployment-ops framework build and are not present anywhere in the current checked-out workspace
Context
Phase 0 recon confirmed the live platform is a single-node k3d cluster on
k3d-alphaswarm-local-server-0. Storage classes are local-path,
alphaswarm-ssd, and alphaswarm-nvme, all backed by the local-path
provisioner. Stateful PVs are pinned to the same node, so any apparent
multi-node scheduling inside the current host would not provide real storage
resilience.
The deployment-ops plan considered three migration targets:
- k3d multi-node on one box.
- real k3s multi-node with a second physical host.
- native single-node k3s preparation now, then graduate to real multi-node later.
The current live box is RAM-bound and hosts the primary platform, so migration work must be reversible and must not disrupt trading workloads.
Decision
Adopt native single-node k3s preparation as the near-term target, and treat real multi-node k3s as a later graduation gated on a second physical amd64 host.
The near-term work is storage/backup/migration preparation, not a production HA
claim. It includes real topology labels, a non-default Longhorn or OpenEBS class
beside local-path, pinned-PV detection, per-StatefulSet dump/restore plans,
Velero/restic or Kopia backup rehearsal, and a future single-to-multi promotion
overlay that borrows the existing tower-green lane's composition style. The
current tower-green overlay is not itself a storage or topology-spread migration
overlay.
Consequences
Positive
- The platform gets honest backup and data-move rehearsals before any node join.
- Future multi-node graduation becomes a join and promotion exercise, not a rebuild.
- k3d remains useful for CI rehearsal but is not mislabeled as production HA.
Negative / risks
- Running replicated storage on the current small box adds overhead and should stay non-default until proven.
- Single-node k3s prep does not eliminate the host as a single point of failure.
- Data movement remains per-service and must be verified with row counts, checksums, or native restore validation.
Rejected
- k3d multi-node on one host as a target. It improves scheduling rehearsal but not host, kernel, disk, RAM, or storage failure resilience.
- Immediate real multi-node migration without a second physical amd64 host and backup validation.
- A blind default storage-class flip before per-StatefulSet restore paths are tested.
Rollout
- Add read-only pinned-PV and affinity/taint detectors.
- Install replicated storage as non-default only after dry-run and capacity review.
- Rehearse dump/restore or native backup per StatefulSet.
- Add Velero/restic or Kopia backup validation against in-cluster object storage.
- Build a new promotion overlay with node-affinity, topology-spread, and storage-class patches; keep it inactive until a second host exists.
- Graduate with
k3supnode join, labels/taints, smoke checks, and tunnel rollback only after explicit operator approval.