Skip to main content

Control the platform ECS deployment

The Platform page in alphaswarm_admin (route /platform, API /admin/platform/ecs/*) controls the hosted platform's own AWS ECS Fargate slice — the alphaswarm-admin and alphaswarm-agentcore-proxy services that run the control plane. It is the counterpart to the Services page, which brokers customer workload lifecycle to alphaswarm_controller.

The admin reaches AWS with its ECS task role — no static keys. The ecs-fargate-control-plane Terraform module grants a tightly scoped self-management policy (ecs:UpdateService + Describe*, logs:* read, cloudwatch:* read) to the services that set enable_self_management.

What it shows​

SurfaceSourcePurpose
Service tableecs:DescribeServicesLive rollout state per service (IN_PROGRESS / COMPLETED / FAILED), running vs desired tasks.
Logs drawerCloudWatch Logs FilterLogEventsA bounded tail of the service's awslogs group, resolved from its task definition.
Metrics drawerCloudWatch GetMetricDataContainer Insights CPU, memory, and running-task count over a window.
Alarms stripcloudwatch:DescribeAlarmsThe platform's per-service alarms (running-task floor, CPU, memory).

Prerequisites​

Set on the alphaswarm_admin deployment:

Env varPurpose
ALPHASWARM_ADMIN_PLATFORM_ECS_CLUSTERECS cluster name the surface targets. The ecs-fargate-control-plane module publishes it at /alphaswarm/<env>/ecs_cluster_name.
ALPHASWARM_ADMIN_PLATFORM_AWS_REGIONRegion the cluster runs in (default us-east-1).
ALPHASWARM_ADMIN_PLATFORM_ALARM_PREFIXAlarm-name prefix used to scope the alarm listing (default alphaswarm-).

The admin must run with alphaswarm-admin[cloud-aws] installed (the boto3 extra). When boto3 is missing the surface returns 503 provider_unavailable with an actionable message; when the cluster is unset it returns 503 provider_misconfigured.

Cross-account or local operation: set ALPHASWARM_ADMIN_PLATFORM_AWS_ASSUME_ROLE_ARN (and optionally ALPHASWARM_ADMIN_PLATFORM_AWS_EXTERNAL_ID) to assume a role into the target account instead of using the ambient task role.

Redeploy a service​

A redeploy starts a new rolling deployment with the same task definition (forceNewDeployment), which is how you pick up a freshly pushed image on a moving tag or recover a wedged service. The ECS deployment circuit breaker with auto-rollback (configured on the service in Terraform) reverts a deployment that never reaches steady state, so a bad image does not take the service down.

  1. Open /platform.
  2. Press Redeploy on the target row.
  3. Type the service name to confirm. The action is audit-first and requires step-up MFA — the UI transparently pops the MFA prompt when the server raises the RFC 9470 challenge.
  4. Watch the rollout badge move to COMPLETED (or FAILED, which means the circuit breaker rolled back).

Scale a service​

  1. Press Scale, set the desired task count, and type the service name to confirm.
  2. Scaling to 0 stops the service; scale back up to restore it.

Both redeploy and scale write a security_audit_events row before the AWS call and a succeeded / failed row after.

Read logs and metrics​

  • Logs resolve the awslogs group from the service's task definition, then tail recent events. Pass a CloudWatch Logs filter pattern to narrow the stream.
  • Metrics read Container Insights series (CPU, memory, running tasks). Enhanced Container Insights must be on for the cluster (the module sets containerInsights = enhanced).

Boundary​

This surface is for the platform's own infrastructure. Customer workloads stay on the Services page, which brokers to the control plane. All boto3 lives in alphaswarm_admin.services.platform_deployment behind the same require_sdk lazy import the cloud-onboarding providers use — route handlers never import a cloud SDK.

See also​