Files
harmony/ROADMAP/14-tenant-application-monitoring.md
Jean-Gabriel Gill-Couture 1e895d7c83
All checks were successful
Run Check Script / check (pull_request) Successful in 2m24s
docs(roadmap): defer tenant application monitoring; capture design tension
Lightweight, cluster-config-aware app monitoring: attach to a cluster stack vs.
provision a tenant-local one, so tenants don't each duplicate cluster-level
Prometheus data while staying self-contained. Audits the four overlapping
backends (none deprecated yet — pending the model decision).
2026-06-22 22:25:09 -04:00

87 lines
4.8 KiB
Markdown

# Phase 14: Tenant Application Monitoring & Alerting (deferred — design first)
## Status
**Deferred (2026-06-22).** Captured while baking Harmony deploys into a client
app (compose app → tenant + Postgres + live Zitadel on K8sAnywhere).
Monitoring/alerting is the next brick for that deploy, but the existing modules
don't yet fit; decide the model before implementing. **No code deprecated yet**
— that waits on the decision below.
## Goal
One lightweight application monitoring + alerting capability a deploy crate adds
with `.with(Monitoring::…)`, that:
- can be **installed or not in dev** (cheap to omit; no heavy stack for a local
k3d inner loop),
- is **installed correctly in prod as a function of the cluster's configuration**
— not a fixed backend. It must answer, per target cluster:
- is there already a **prometheus-operator** / cluster monitoring stack to
attach to (OKD user-workload-monitoring, kube-prometheus-stack, COO)?
- are **in-cluster Prometheus metrics** already scraped at the cluster level
(kube-state-metrics, node/cadvisor) so the tenant needs only *rules*, not a
*scraper*?
- how are **Prometheus permissions scoped for a tenant** (namespaced
PrometheusRule/ServiceMonitor + RBAC vs. a tenant-owned Prometheus)?
## The tension to resolve (the actual reason this is deferred)
- **Anti-pattern risk:** if every tenant rolls its own Prometheus, we duplicate
data already monitored at the cluster level — N Prometheis re-scraping the same
kube-state-metrics, N Alertmanagers, N Grafanas. Wasteful and operationally
noisy.
- **Counter-value:** a **self-contained tenant** is high customer value — a
client can look at *their* namespace and understand their app's health without
cluster-admin context. Easy to reason about, easy to hand off.
- The resolution is probably **not** "one backend" but a capability that picks:
attach-to-cluster-stack (rules + a scoped Grafana view) when one exists, vs.
provision a minimal tenant-scoped stack when it doesn't — driven by a topology
capability/probe, the same way `K8sAnywhereTopology` already detects distro and
`TenantManager` already owns namespace isolation (ADR-011) and identity
(ADR-027). Cluster-level data stays shared; the *tenant view* + *app alert
rules* are what's per-tenant.
## Current state (audited 2026-06-22)
Four overlapping, none-satisfying backends (all `AlertSender` impls):
| Module | Sender | Author / created | Note |
|---|---|---|---|
| `monitoring/kube_prometheus/helm_prometheus_alert_score.rs` + `helm/` | `KubePrometheus` | Willem, 2025-06 | Installs full kube-prometheus-stack via Helm. Works on k3d. Used by `harmony_app::capabilities::Monitoring`. Heaviest; the "tenant rolls its own" path. |
| `monitoring/application_monitoring/application_monitoring_score.rs` | `CRDPrometheus` | Willem, 2025-08 | Prometheus-operator CRDs directly. |
| `monitoring/application_monitoring/rhobs_application_monitoring_score.rs` | `RHOBObservability` | Willem, 2025-09 | Cluster Observability Operator `MonitoringStack`. OKD-oriented. |
| `monitoring/prometheus/` | `Prometheus` | Willem, 2025-07 | Standalone Prometheus. |
Adjacent (keep — different concern, the cluster/ops view, not per-app alerting):
- `monitoring/cluster_dashboards/` — Grafana-operator dashboards + datasource +
route (Sylvain Tremblay, 2026-03). Cluster-level dashboards on an existing
Prometheus; OKD-route oriented.
- `monitoring/ceph_alerts.rs`, `monitoring/okd/cluster_alert_rules.rs` — alerts
on OKD's native stack (Sylvain Tremblay, 2026-04).
`harmony_app::capabilities::Monitoring` (the `.with(Monitoring::new().alert(r))`
DX) currently emits `HelmPrometheusAlertingScore` (the Helm/KubePrometheus path)
+ a namespace-scoped "app has no available replicas for 5m" rule via
kube-state-metrics. The DX shape is right; the **backend choice underneath it is
the open question**.
## Decision criteria (to settle next)
1. Capability detects cluster monitoring posture (operator present? cluster
metrics scraped?) and **attaches** vs. **provisions** accordingly — ideally a
`Topology`/`TenantManager` capability, not per-app branching.
2. Default to **rules + tenant-scoped view on shared cluster data**; provision a
tenant-local stack only when no cluster stack exists (e.g. bare k3d dev).
3. Keep the deploy-crate DX a one-liner (`.with(Monitoring::…)`); the
environment, not the app, decides install vs. attach.
4. Once chosen: collapse the four backends to the canonical one and
`#[deprecated]` (not delete) the overlaps; keep the workspace compiling.
## First consumer waiting on this
The first client app deploy crate using this paradigm (frontend + backend +
Zitadel; its external SaaS dependency is not deployed). It deliberately ships
**without** a monitoring Score today — add it here once the model lands.