Skip to main content

prometheus

Catalog date: 2026-06-24.

The cluster-internal metrics scraper. Deployed via kube-prometheus-stack which also installs the operator, Alertmanager, and the Grafana sidecar.

Identity​

FieldValue
Service idprometheus
Roleobservability
Imageprom/prometheus (managed by kube-prometheus-stack)
Port9090
Health/-/ready

Deployment surfaces​

SurfaceWhere
Kustomizeobservability/kube-prometheus-stack/ — Helm-managed via kustomize HelmCharts overlay
Compose(not in compose — local dev relies on victoriametrics for the small footprint)

Scrape targets​

The kube-prometheus-stack installs a default ServiceMonitor set; we extend it with:

  • alphaswarm-core /metrics (every API pod).
  • alphaswarm-worker /metrics (per Celery worker).
  • alphaswarm-cp /metrics.
  • KEDA metrics adapter on alphaswarm-controller-operator and bots-operator.
  • Linkerd proxy metrics (mTLS-side).
  • Per-data-plane service exporters (Postgres exporter, Redis exporter, Kafka exporter, etc.).

Long-term storage​

Per helm-values.yaml, Prometheus is currently configured with retention: 14d and retentionSize: 40GB — not the 30-day figure stated in an earlier version of this doc. No remoteWrite block to VictoriaMetrics was found in this file; VictoriaMetrics instead gets its own data via vmagent's independent Prometheus-style scrape configs (see otel-collector.md), not via a Prometheus remote-write target. The "parallel-cutover, then drop to 7 days" plan could not be verified against current config and is flagged as unverified rather than restated.

Operations​

  • Alertmanager: receives the kube-prometheus-stack default alert set + AlphaSwarm-specific rules from alerting-rules.yaml — a single file, not an alerts/ subdirectory as an earlier version of this doc stated.
  • Federation: no remoteWrite config to VictoriaMetrics was found (see "Long-term storage" above), so the "disabled because remote-write handles the long-term path" rationale is unverified; federation itself was not found configured either way.
  • PromQL recording rules: also live in alerting-rules.yaml — no separate observability/kube-prometheus-stack/rules/ directory exists; agent-emitted ad-hoc rules are forbidden.

See also​