Prometheus + Grafana + Loki observability stack for the Omi backend on GKE.
Repo-enforceable inventory of expected prod scrape targets, routing labels, and
no-data semantics lives in expected-targets.prod.yaml.
- Enforced in CI by
backend/tests/unit/test_monitoring_telemetry_contract.py. - Jobs marked
coverage_status: enforcedmust appear inomi-journey-scrape-missing. - Jobs marked
coverage_status: declaredare wired in values/ServiceMonitors but do not yet have a dedicated coverage alert (follow-up to expand the scrape-missing rule). - Managed GKE control-plane exclusions (
kubeProxy/kubeScheduler/kubeControllerManager) are documented here and owned by #9138 / PR #11093 — do not pageTargetDownfor those endpoints. - Live Instatus delivery verification and notification-policy rework remain maintainer-approved Phase 2 work (not claimed by the inventory).
┌───────────────────────────────────────────────────────────────────────────┐
│ GKE Cluster (prod-omi-gke / dev-omi-gke) │
│ │
│ ┌──────────────┐ scrape ┌─────────────┐ query ┌────────────┐ │
│ │ Pod metrics │──────────►│ Prometheus │◄─────────│ Grafana │ │
│ │ (app /metrics)│ │ (10d, 50Gi) │ │ prod: monitor│ │
│ └──────────────┘ └───────┬───────┘ │ .omi.me │ │
│ │ │ dev: monitor│ │
│ ┌──────────────┐ scrape │ │ .omiapi.com│ │
│ │ DCGM exporter │──────────► │ └─────┬──────┘ │
│ │ (GKE addon) │ │ │ │
│ └──────────────┘ │ query │ query │
│ ▼ │ │
│ ┌──────────────┐ scrape ┌───────────────┐ │ │
│ │ Stackdriver │────────►│ prometheus- │ │ │
│ │ exporter │ (via │ adapter │ │ │
│ │ (GCP metrics) │ Prom) │ (HPA metrics) │ │ │
│ └──────────────┘ └───────┬────────┘ │ │
│ │ │ │
│ ┌──────────┴──────────┐ │ │
│ │ external.metrics │ │ │
│ │ custom.metrics.k8s.io│ │ │
│ └──────────┬──────────┘ │ │
│ ▼ │ │
│ ┌──────────────┐ │ │
│ │ HPA │ │ │
│ │ controllers │ │ │
│ └──────────────┘ │ │
│ │ │
│ ┌──────────────┐ collect ┌──────────┐ query │ │
│ │ Pod logs │──────────►│ Loki │◄──────────────────┘ │
│ │ (stdout/err) │ (k8s API) │ (GCS,15d) │ │
│ └──────────────┘ └──────────┘ │
│ ▲ ▲ │
│ │ │ │
│ Alloy basic auth │
│ (DaemonSet) (alloy-basic-auth → loki-basic-auth) │
└───────────────────────────────────────────────────────────────────────────┘
Note: Stackdriver exporter is scraped by Prometheus (job prometheus-stackdriver-metrics), then prometheus-adapter queries Prometheus for those metrics. The exporter does not feed the adapter directly.
Cloud Run application metrics take a push-then-pull bridge because a public URL scrape reaches only one random autoscaled instance. Each backend and desktop-backend instance exposes its registry on loopback port 9090 to Google's Managed Service for Prometheus sidecar. The sidecar writes prometheus.googleapis.com/omi_* to Cloud Monitoring. A separate, rate-limited Stackdriver exporter imports only those two Cloud Run services, and Prometheus scrapes it as cloud-run-application-metrics. See ../../docs/runbooks/cloud-run-metrics-ingestion.md.
| Component | Chart | Purpose | Namespace |
|---|---|---|---|
| Prometheus | kube-prometheus-stack |
Metrics collection, 10d retention, 50Gi storage | {env}-omi-monitoring |
| Grafana | kube-prometheus-stack |
Dashboards and alerting (prod: monitor.omi.me, dev: monitor.omiapi.com) |
{env}-omi-monitoring |
| Alertmanager | kube-prometheus-stack |
Alert routing and notification | {env}-omi-monitoring |
| Grafana Image Renderer | kube-prometheus-stack |
Alert screenshot capture | {env}-omi-monitoring |
| Loki | loki |
Log aggregation, distributed mode, GCS backend, 15d retention | {env}-omi-monitoring |
| Alloy | alloy (k8s-monitoring) |
Pod log collection via Kubernetes API | {env}-omi-monitoring |
| prometheus-adapter | prometheus-adapter |
Translates Prometheus metrics → K8s custom metrics API for HPA | {env}-omi-monitoring |
| Stackdriver exporter | prometheus-stackdriver-exporter |
Bridges GCP load balancer metrics into Prometheus | {env}-omi-monitoring |
| DCGM exporter | GKE-managed addon | GPU metrics (utilization, memory, temperature) | gke-managed-system |
| kube-state-metrics | kube-prometheus-stack |
K8s object state (pod count, deployment replicas) | {env}-omi-monitoring |
| node-exporter | kube-prometheus-stack |
Node-level CPU, memory, disk, network | {env}-omi-monitoring |
GKE application metrics use two pull mechanisms. Cloud Run uses a third, per-instance push bridge because its public URL is load balanced.
1. ServiceMonitor (recommended for new GKE services)
Used by: parakeet.
The service chart includes a ServiceMonitor CRD that Prometheus auto-discovers. This is the preferred approach — no changes to the monitoring chart needed.
2. Pod annotations + additionalScrapeConfigs
Used by: backend-listen, pusher, llm-gateway, deepgram engine, GPU metrics, Stackdriver.
Pods set annotations (prometheus.io/scrape: "true", prometheus.io/port, prometheus.io/path) and a matching additionalScrapeConfigs entry in kube-prometheus-stack values defines the scrape job.
Backend-listen, pusher, and llm-gateway require bearer token auth via the metrics-scrape-token secret.
3. Cloud Run GMP sidecar -> Cloud Monitoring -> isolated Stackdriver exporter
Used by: Cloud Run backend and desktop-backend. The sidecar scrapes an unauthenticated loopback-only listener on port 9090; the public port 8080 /metrics route remains bearer protected. The isolated exporter is intentionally separate from the 1-second load-balancer exporter so per-instance application series are read from the Cloud Monitoring API only every 30 seconds.
These are the additionalScrapeConfigs and ServiceMonitor targets. Built-in kube-prometheus-stack targets (kube-state-metrics, node-exporter, kubelet, Alertmanager, Prometheus itself) are not listed here.
| Job | Target | Interval | Auth |
|---|---|---|---|
backend-listen-metrics |
backend-listen pods /metrics:8080 |
15s | Bearer token |
pusher-metrics |
pusher pods /metrics:8080 |
15s | Bearer token |
llm-gateway-metrics |
llm-gateway pods /metrics:8080 |
15s | Bearer token |
dg_engine_metrics |
DG engine pods in prod-omi-dg-self-hosted |
2s | None |
gpu-metrics |
all pods in gke-managed-system (includes DCGM exporter) |
1s | None |
prometheus-stackdriver-metrics |
Load-balancer Stackdriver exporter in prod-omi-monitoring |
1s | None |
cloud-run-application-metrics |
prometheus.googleapis.com/omi_* for Cloud Run backend and desktop-backend via isolated exporter |
30s | None |
ServiceMonitor: parakeet |
parakeet pods /metrics:9091 |
15s | None |
For llm-gateway streams, llm_gateway_requests_total{outcome="success"} is emitted only after the provider's
terminal SSE marker (data: [DONE] or Anthropic event: message_stop) is observed. EOF without that marker,
transport failure after first output, and client cancellation have separate bounded outcome, phase, and
error_class values. This proves provider completion, not guaranteed delivery of the terminal chunk to the client.
credential_source separates Omi-managed traffic from service-forwarded BYOK without putting key material or user
identity in labels. llm_gateway_stream_ttfb_seconds records time to the first non-empty provider chunk.
Pod stdout/stderr → Alloy (kubernetesApi method) → Loki gateway (basic auth) → GCS
↓
Grafana Explore (LogQL)
Alloy collects from {env}-omi-backend namespace only. Labels kept: app_kubernetes_io_name, container, instance, job, level, namespace, service_name, pod (structured metadata).
Loki runs in distributed mode: 2x ingester, 2x querier, 2x query-frontend, 2x query-scheduler, 2x distributor, 1x compactor, 2x index-gateway, 1x ruler. Storage on GCS (prod-omi-loki-chunks). Retention: 15 days (360h).
DCGM exporter is a GKE-managed addon (not in this repo). It runs in gke-managed-system and exposes metrics like:
DCGM_FI_DEV_GPU_UTIL— GPU utilization %DCGM_FI_DEV_FB_USED/DCGM_FI_DEV_FB_FREE— GPU memoryDCGM_FI_DEV_GPU_TEMP— GPU temperature
Prometheus scrapes these via the gpu-metrics job. GPU node pools:
parakeet-pool— NVIDIA L4 (parakeet ASR)diarizer-pool— NVIDIA T4 (speaker diarization)vad-pool-v2— NVIDIA T4 (voice activity detection)
The adapter translates Prometheus queries into K8s metrics APIs so HPAs can scale on custom metrics. It serves two APIs:
external.metrics.k8s.io— namespace-scoped metrics (most HPA metrics use this)custom.metrics.k8s.io— pod-scoped metrics (parakeet uses this forparakeet_active_streamsandparakeet_active_requests_total)
| Metric | Source | Used by |
|---|---|---|
backend_listen_active_ws_connections_per_pod |
backend-listen gauge | backend-listen HPA |
pusher_active_ws_connections_per_pod |
pusher gauge | pusher HPA |
vad_request_latency_p99 |
Stackdriver ILB histogram | vad HPA |
diarizer_request_latency_p99 |
Stackdriver ILB histogram | diarizer HPA |
backend_listen_response_code_500 |
Stackdriver LB counter | backend-listen HPA |
backend_listen_requests_per_pod |
Stackdriver LB counter / replica count | backend-listen HPA |
engine_active_requests_stt_streaming |
DG engine gauge | DG engine HPA |
engine_active_requests_stt_batch |
DG engine gauge | DG engine HPA |
engine_active_requests_tts_batch |
DG engine gauge | DG engine HPA |
engine_estimated_stream_capacity |
DG engine gauge | DG engine HPA |
engine_requests_active_to_max_ratio |
DG engine derived | DG engine HPA |
engine_to_api_pod_ratio |
kube-state-metrics | DG scaling |
engine_avg_gpu_utilization |
DCGM via Prometheus | DG engine HPA |
parakeet_active_streams |
parakeet gauge | parakeet HPA |
parakeet_active_batch_requests |
parakeet gauge | parakeet HPA |
parakeet_active_requests_total |
parakeet derived (streams + batch) | parakeet HPA |
parakeet_gpu_utilization |
DCGM via Prometheus | parakeet HPA |
parakeet_request_latency_p99 |
parakeet histogram | parakeet HPA |
Parakeet adapter rules are defined in the parakeet chart's values.yaml but must be merged into the cluster-wide adapter (see below).
Bridges GCP Cloud Monitoring into Prometheus. Two releases share the existing Workload Identity service account (prod-omi-prom-stackdriver-gsa):
prod-omi-prometheus-stackdriver-exporterkeeps the latency-sensitive load-balancer prefixes at the existing 1-second Prometheus scrape interval.prod-omi-cloud-run-metrics-exporterreads onlyprometheus.googleapis.com/omi_*, filtered to the Cloud Run monitored-resource namespacesbackendanddesktop-backend; Prometheus scrapes this release every 30 seconds.
The application exporter is separate to prevent Cloud Monitoring API read cost from multiplying every per-instance application series by the legacy 1-second scrape rate. Stackdriver exporter exposes normalized names such as stackdriver_prometheus_target_prometheus_googleapis_com_<metric>_<type>; retain the service_name and instance labels and aggregate counters across instances in PromQL.
Each component has dev and prod values:
.github/workflows/gcp_cloud_run_metrics_egress.yml owns the Cloud Run application exporter and its kube-prometheus-stack scrape configuration: relevant merges auto-deploy development, while production requires a protected-environment dispatch from main.
| Component | Dev | Prod |
|---|---|---|
| kube-prometheus-stack | kube-prometheus-stack/dev_omi_monitoring_values.yaml |
kube-prometheus-stack/prod_omi_monitoring_values.yaml |
| prometheus-adapter | prometheus-adapter/dev_omi_prometheus_adapter.yaml |
prometheus-adapter/prod_omi_prometheus_adapter.yaml |
| Alloy | alloy/dev_omi_k8s_monitoring_values.yml |
alloy/prod_omi_k8s_monitoring_values.yml |
| Loki | loki/dev_omi_loki_values.yaml |
loki/prod_omi_loki_values.yaml |
| Stackdriver exporter | prometheus-stackdriver-exporter/dev_omi_stackdriver_exporter.yaml |
prometheus-stackdriver-exporter/prod_omi_stackdriver_exporter.yaml |
| Cloud Run metrics exporter | prometheus-stackdriver-exporter/dev_omi_cloud_run_metrics_exporter.yaml |
prometheus-stackdriver-exporter/prod_omi_cloud_run_metrics_exporter.yaml |
| Grafana ALB cert | kube-prometheus-stack/dev_omi_grafana_alb_cert.yaml |
kube-prometheus-stack/prod_omi_grafana_alb_cert.yaml |
45 dashboards on prod Grafana (monitor.omi.me), organized by folder.
Most are bundled with kube-prometheus-stack and auto-provisioned. Custom dashboards are noted.
| Dashboard | UID | Tags | Notes |
|---|---|---|---|
| Alertmanager / Overview | alertmanager-overview |
alertmanager-mixin |
Bundled |
| Backend API Monitoring | 57c2a5ea-c310-4401-ac72-54dbc6da4c7e |
api, backend, monitoring, omi |
Custom |
| Backend API Monitoring | 3e7c5f57-a1be-4175-81e6-1f0c7c28b9dd |
api, backend, monitoring, omi |
Custom (duplicate — consolidate) |
| CoreDNS | vkQ0UHxik |
coredns, dns |
Bundled |
| etcd | c2f4e12cdf69feb95caa41a5a1b423d9 |
etcd-mixin |
Bundled |
| Grafana Overview | 6be0s85Mk |
— | Bundled |
| K8s Node Metrics / Multi Clusters | your_custom_uid_X0dfg |
Prometheus, node_exporter |
Custom (community import) |
| Kubernetes / API server | 09ec8aa1e996d6ffcd6817bbaff4db1b |
kubernetes-mixin |
Bundled |
| Kubernetes / Compute Resources / Multi-Cluster | b59e6c9f2fcbe2e16d77fc492374cc4f |
kubernetes-mixin |
Bundled |
| Kubernetes / Compute Resources / Cluster | efa86fd1d0c121a26444b636a3f509a8 |
kubernetes-mixin |
Bundled |
| Kubernetes / Compute Resources / Namespace (Pods) | 85a562078cdf77779eaa1add43ccec1e |
kubernetes-mixin |
Bundled |
| Kubernetes / Compute Resources / Namespace (Workloads) | a87fb0d919ec0ea5f6543124e16c42a5 |
kubernetes-mixin |
Bundled |
| Kubernetes / Compute Resources / Node (Pods) | 200ac8fdbfbb74b39aff88118e4d1c2c |
kubernetes-mixin |
Bundled |
| Kubernetes / Compute Resources / Pod | 6581e46e4e5c7ba40a07646395ef7b23 |
kubernetes-mixin |
Bundled |
| Kubernetes / Compute Resources / Workload | a164a7f0339f99e89cea5cb47e9be617 |
kubernetes-mixin |
Bundled |
| Kubernetes / Controller Manager | 72e0e05bef5099e5f049b05fdc429ed4 |
kubernetes-mixin |
Bundled |
| Kubernetes / Kubelet | 3138fa155d5915769fbded898ac09fd9 |
kubernetes-mixin |
Bundled |
| Kubernetes / Networking / Cluster | ff635a025bcfea7bc3dd4f508990a3e9 |
kubernetes-mixin |
Bundled |
| Kubernetes / Networking / Namespace (Pods) | 8b7a8b326d7a6f1f04244066368c67af |
kubernetes-mixin |
Bundled |
| Kubernetes / Networking / Namespace (Workload) | bbb2a765a623ae38130206c7d94a160f |
kubernetes-mixin |
Bundled |
| Kubernetes / Networking / Pod | 7a18067ce943a40ae25454675c19ff5c |
kubernetes-mixin |
Bundled |
| Kubernetes / Networking / Workload | 728bf77cc1166d2f3133bf25846876cc |
kubernetes-mixin |
Bundled |
| Kubernetes / Persistent Volumes | 919b92a8e8041bd567af9edab12c840c |
kubernetes-mixin |
Bundled |
| Kubernetes / Proxy | 632e265de029684c40b21cb76bca4f94 |
kubernetes-mixin |
Bundled |
| Kubernetes / Scheduler | 2e6b6a3b4bddf1427b3a55aa1311c656 |
kubernetes-mixin |
Bundled |
| Node Exporter / AIX | 7e0a61e486f727d763fb1d86fdd629c2 |
node-exporter-mixin |
Bundled |
| Node Exporter / MacOS | 629701ea43bf69291922ea45f4a87d37 |
node-exporter-mixin |
Bundled |
| Omi Core Features | omi-core-features |
— | Custom — user-outcome view: journeys, subscriptions, LLM gateway, capture pipeline, PTT transport (realtime_voice client journey). The finalization gauges it reads are one global value republished by every backend-listen replica: aggregate with max(), never sum(). |
| Node Exporter / Nodes | 7d57716318ee0dddbac5a7f451fb7753 |
node-exporter-mixin |
Bundled |
| Node Exporter / USE Method / Cluster | 3e97d1d02672cdd0861f4c97c64f89b2 |
node-exporter-mixin |
Bundled |
| Node Exporter / USE Method / Node | fac67cfbe174d3ef53eb473d73d9212f |
node-exporter-mixin |
Bundled |
| Parakeet ASR Monitoring | 07e4c65f-ae79-414d-bf05-99468267d199 |
asr, gke, gpu, parakeet |
Custom — GPU ASR service metrics |
| Prometheus / Overview | 9fa0d141-d019-4ad7-8bc5-42196ee308bd |
prometheus-mixin |
Bundled |
Folder: Cloud Run (folder UID: aev9if48326f4e)
| Dashboard | UID | Notes |
|---|---|---|
| Backend | 0253019b-c68a-4aef-a27d-6bb3408727fb |
Cloud Run backend (main API) |
| Backend-integration | 5be48038-a72b-4938-99ed-7a8747655294 |
Cloud Run backend-integration |
| Backend-sync | 8bf7bd3f-8dbc-4f86-a532-557bfac0d7ac |
Cloud Run backend-sync |
| Plugins | e736ab7d-d3e8-444f-a743-369b054def9e |
Cloud Run plugins service |
Folder: GKE (folder UID: aev9igt5fwgsgc)
| Dashboard | UID | Notes |
|---|---|---|
| Backend-listen | 855b2e16-c098-407a-85dc-dc9ce87698a9 |
GKE WebSocket listener |
| Deepgram self-hosted | fedizdcosu1oga |
Self-hosted STT engine |
| Diarizer | 303b7396-ce6e-48ee-be24-9c157a710adf |
Speaker diarization GPU service |
| Pusher | c758b698-01a0-4b5c-b58c-e81e4ff33ccd |
Audio pusher service |
| VAD | 72cfe240-ae8c-4076-845e-c58e28f12d87 |
Voice activity detection GPU service |
Folder: Omi Services (folder UID: betdycdziadc0e)
| Dashboard | UID | Notes |
|---|---|---|
| Cloud Armor denied requests | 5feac510-b391-48fc-9c4b-1b8dde4ab32a |
WAF/security denied traffic |
| Cloud Run Services - Logs | d2d782ef-f537-46b8-969d-f73561ec7d07 |
Aggregated Cloud Run logs view |
| Global External ALB | 59aa0de7-15c6-413f-acba-b7e99296ad75 |
External load balancer metrics |
| Omi Kubernetes Events | 3714dbfa-114b-47a0-99ca-1a26354e792a |
K8s event stream (OOM kills, pod evictions) |
| Resilience / Fallbacks | omi-resilience-fallbacks |
Fallback rates, sync/pusher SLOs, gateway ticket tier — see backend/docs/runbooks/resilience-dashboards.md |
| Category | Count | Source | Version-controlled |
|---|---|---|---|
| Bundled (kube-prometheus-stack) | 28 | Helm chart sidecar | Yes (via chart defaults) |
| Custom (Omi-specific) | 18 | Exported from Grafana UI | Yes — dashboards/ directory |
All 18 custom dashboards are exported to dashboards/ as provisioning-ready JSON (.id and .version stripped). The K8s Node Metrics dashboard (your_custom_uid_X0dfg) is a community import bundled with the chart and not separately exported.
Option A: ServiceMonitor (preferred)
Add to your service's Helm chart:
- A metrics
Serviceexposing the metrics port:
# templates/service-metrics.yaml
{{- if .Values.metrics.enabled }}
apiVersion: v1
kind: Service
metadata:
name: {{ include "myservice.fullname" . }}-metrics
labels:
{{- include "myservice.labels" . | nindent 4 }}
spec:
type: ClusterIP
ports:
- port: {{ .Values.metrics.port }}
targetPort: {{ .Values.metrics.port }}
protocol: TCP
name: metrics
selector:
{{- include "myservice.selectorLabels" . | nindent 4 }}
{{- end }}- A
ServiceMonitor:
# templates/servicemonitor.yaml
{{- if and .Values.metrics.enabled .Values.metrics.serviceMonitor.enabled }}
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: {{ include "myservice.fullname" . }}
labels:
{{- include "myservice.labels" . | nindent 4 }}
{{- with .Values.metrics.serviceMonitor.labels }}
{{- toYaml . | nindent 4 }}
{{- end }}
spec:
selector:
matchLabels:
{{- include "myservice.selectorLabels" . | nindent 6 }}
endpoints:
- port: metrics
path: /metrics
interval: {{ .Values.metrics.serviceMonitor.interval | default "15s" }}
{{- end }}- Values:
metrics:
enabled: true
port: 9091
serviceMonitor:
enabled: true
interval: "15s"
labels:
release: prod-omi-kube-prometheus-stack # must match Prometheus serviceMonitorSelectorThe release label must match the Prometheus instance's serviceMonitorSelector. Use prod-omi-kube-prometheus-stack for prod, dev-kube-prometheus-stack for dev.
Option B: Pod annotations
Add to your chart's values:
podAnnotations:
prometheus.io/scrape: "true"
prometheus.io/port: "8080"
prometheus.io/path: "/metrics"Then add a scrape job to kube-prometheus-stack/{env}_omi_monitoring_values.yaml under prometheus.prometheusSpec.additionalScrapeConfigs. Use backend-listen or pusher as a template. If the metrics endpoint requires auth, mount the metrics-scrape-token secret.
To scale a deployment on a custom Prometheus metric:
- Add the metric rule to
prometheus-adapter/{env}_omi_prometheus_adapter.yaml:
rules:
external:
- name:
as: "myservice_custom_metric"
seriesQuery: 'my_prometheus_metric{namespace!=""}'
metricsQuery: 'avg(my_prometheus_metric{<<.LabelMatchers>>})'
resources:
overrides:
namespace: { resource: "namespace" }- Apply with Helm:
helm -n {env}-omi-monitoring upgrade --install {env}-omi-prometheus-adapter \
prometheus-community/prometheus-adapter \
-f prometheus-adapter/{env}_omi_prometheus_adapter.yaml- Verify the metric is exposed:
# For External metrics (namespace-scoped, most common):
kubectl get --raw "/apis/external.metrics.k8s.io/v1beta1/namespaces/{env}-omi-backend/myservice_custom_metric"
# For Pods metrics (pod-scoped, used by parakeet):
kubectl get --raw "/apis/custom.metrics.k8s.io/v1beta1/namespaces/{env}-omi-backend/pods/*/myservice_pod_metric"- Reference in your HPA:
# External metric (namespace-scoped):
metrics:
- type: External
external:
metric:
name: myservice_custom_metric
target:
type: Value
value: "70"
# Pods metric (per-pod average):
metrics:
- type: Pods
pods:
metric:
name: myservice_pod_metric
target:
type: AverageValue
averageValue: "25"Parakeet adapter rules: The parakeet chart defines adapter rules in its own values.yaml under prometheus-adapter.rules, but these are NOT yet present in the cluster-wide adapter config (prometheus-adapter/ values in this directory). They must be manually merged into prometheus-adapter/{env}_omi_prometheus_adapter.yaml before parakeet HPA can use them. The parakeet sub-chart adapter is disabled (prometheus-adapter.enabled: false) to avoid deploying a second adapter that would conflict with the cluster-scoped APIService.
Grafana: prod at https://monitor.omi.me/, dev at https://monitor.omiapi.com/. See "Current Dashboards" above for the full inventory.
-
Create in Grafana UI first — build the dashboard in dev (
monitor.omiapi.com), iterate until it works. -
Export the dashboard JSON:
# Get a Grafana API token (Settings → API Keys → Add, role=Viewer)
export GRAFANA_TOKEN="your-token"
export GRAFANA_HOST="https://monitor.omiapi.com" # dev first
# List dashboards to find the UID
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
"$GRAFANA_HOST/api/search?type=dash-db" | jq '.[] | {title, uid}'
# Export a specific dashboard (strips runtime fields)
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
"$GRAFANA_HOST/api/dashboards/uid/<UID>" | \
jq '.dashboard | del(.id, .version)' > dashboards/<name>.json-
Add to version control:
- Save the JSON to
backend/charts/monitoring/dashboards/<name>.json - Use a descriptive filename matching the dashboard title (e.g.
parakeet-asr-monitoring.json) - Strip runtime fields:
.id,.version(done by thejqcommand above) - Keep the
.uidfield — it links the repo copy to the live dashboard
- Save the JSON to
-
Import to Grafana (if creating on a new/different instance):
# First, get or create the target folder (returns folderUid)
FOLDER_UID=$(curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
"$GRAFANA_HOST/api/folders" | jq -r '.[] | select(.title=="GKE") | .uid')
# Import with folder placement
curl -s -X POST -H "Authorization: Bearer $GRAFANA_TOKEN" \
-H "Content-Type: application/json" \
-d "{\"dashboard\": $(cat dashboards/<folder>/<name>.json), \"folderUid\": \"$FOLDER_UID\", \"overwrite\": true}" \
"$GRAFANA_HOST/api/dashboards/db"Without folderUid, the dashboard lands in General. See the folder UIDs in the "Current Dashboards" section above.
- Commit the JSON and open a PR.
- Edit the dashboard in Grafana UI (dev first, then prod)
- Sync back to the repo using the workflow below
When a dashboard is edited via the Grafana UI (emergency fixes, quick iterations), export it back to the repo:
# 1. Set up
export GRAFANA_TOKEN="your-token"
export GRAFANA_HOST="https://monitor.omi.me" # or monitor.omiapi.com
# 2. Export the updated dashboard
DASHBOARD_UID="07e4c65f-ae79-414d-bf05-99468267d199" # example: Parakeet ASR
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
"$GRAFANA_HOST/api/dashboards/uid/$DASHBOARD_UID" | \
jq '.dashboard | del(.id, .version)' > dashboards/parakeet-asr-monitoring.json
# 3. Review the diff
git diff dashboards/parakeet-asr-monitoring.json
# 4. Commit and PR
git add dashboards/parakeet-asr-monitoring.json
git commit -m "sync(monitoring): export parakeet dashboard from Grafana UI"Bulk sync (all 17 custom dashboards):
export GRAFANA_TOKEN="your-token"
export GRAFANA_HOST="https://monitor.omi.me"
# UID → folder mapping (mirrors repo directory structure)
declare -A DASHBOARDS=(
# general/
["57c2a5ea-c310-4401-ac72-54dbc6da4c7e"]="general"
["3e7c5f57-a1be-4175-81e6-1f0c7c28b9dd"]="general"
["07e4c65f-ae79-414d-bf05-99468267d199"]="general"
# cloud-run/
["0253019b-c68a-4aef-a27d-6bb3408727fb"]="cloud-run"
["5be48038-a72b-4938-99ed-7a8747655294"]="cloud-run"
["8bf7bd3f-8dbc-4f86-a532-557bfac0d7ac"]="cloud-run"
["e736ab7d-d3e8-444f-a743-369b054def9e"]="cloud-run"
# gke/
["855b2e16-c098-407a-85dc-dc9ce87698a9"]="gke"
["fedizdcosu1oga"]="gke"
["303b7396-ce6e-48ee-be24-9c157a710adf"]="gke"
["c758b698-01a0-4b5c-b58c-e81e4ff33ccd"]="gke"
["72cfe240-ae8c-4076-845e-c58e28f12d87"]="gke"
# omi-services/
["5feac510-b391-48fc-9c4b-1b8dde4ab32a"]="omi-services"
["d2d782ef-f537-46b8-969d-f73561ec7d07"]="omi-services"
["59aa0de7-15c6-413f-acba-b7e99296ad75"]="omi-services"
["3714dbfa-114b-47a0-99ca-1a26354e792a"]="omi-services"
["omi-resilience-fallbacks"]="omi-services"
)
for uid in "${!DASHBOARDS[@]}"; do
folder="${DASHBOARDS[$uid]}"
mkdir -p "dashboards/$folder"
slug=$(curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
"$GRAFANA_HOST/api/dashboards/uid/$uid" | \
jq -r '.meta.slug // .dashboard.title' | tr ' /' '-' | tr '[:upper:]' '[:lower:]')
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
"$GRAFANA_HOST/api/dashboards/uid/$uid" | \
jq '.dashboard | del(.id, .version)' > "dashboards/$folder/${slug}.json"
echo "Exported: $folder/$slug ($uid)"
doneDev → Prod promotion:
# Export from dev
export GRAFANA_HOST="https://monitor.omiapi.com"
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \
"$GRAFANA_HOST/api/dashboards/uid/<uid>" | \
jq '.dashboard | del(.id, .version)' > dashboards/<folder>/<name>.json
# Import to prod with correct folder
FOLDER_UID="aev9igt5fwgsgc" # example: GKE folder
curl -s -X POST -H "Authorization: Bearer $PROD_GRAFANA_TOKEN" \
-H "Content-Type: application/json" \
-d "{\"dashboard\": $(cat dashboards/<folder>/<name>.json), \"folderUid\": \"$FOLDER_UID\", \"overwrite\": true}" \
"https://monitor.omi.me/api/dashboards/db"Rules:
- Always export from prod to repo (prod is the live source until full provisioning is in place)
- Strip
.idand.version— they are instance-specific - Keep
.uid— it prevents duplicate dashboards on import - Emergency UI edits must be synced back to repo within the same day
- When syncing, check
git diffto ensure only intended panels changed (Grafana may reorder JSON keys)
Alerting is configured through Grafana unified alerting (not Prometheus AlertManager rules directly). Alert screenshots are enabled via the Grafana image renderer.
Loki ruler is configured to send alerts to AlertManager at http://prod-kube-prometheus-stack-alertmanager:9093.
Grafana rules are operational runbook entries, not raw metric dumps. The split files in alerts/*.json are the
maintained sources; alert-rules.json is the canonical combined Grafana import. They must contain the same rules
with byte-for-byte equal objects when indexed by stable Grafana UID. The deterministic monitoring contract test checks
both exports, including duplicate UIDs, before a change can land.
Pusher release workflows additionally call verify_pusher_live_alert_route.py
before publishing or promoting an image. The protected MONITOR_GRAFANA_TOKEN
secret (repo-scoped for dest monitor.omiapi.com, prod environment-scoped
for monitor.omi.me) must be able to read provisioned alert rules, datasource
health and queries, and contact points. Do not reuse GRAFANA_TOKEN here; that
secret belongs to the TV Cloud Run Grafana. The gate fails closed unless the committed memory-admission
and capture-outcome pager set is live and unpaused, Prometheus reports healthy,
both Pusher and backend-listen scrape targets are currently healthy, and the
exact Telegram receiver exists with resolve notifications enabled. After each
rollout it additionally requires all three finalization telemetry families from
both jobs; zero-valued labeled failure children are initialized at process
startup so absence is unambiguously a source failure. Production repeats this
check after rollout so an hours-long release cannot finish on stale evidence.
It never prints contact-point settings or token material.
Every rule carries these notification fields:
| Field | Purpose |
|---|---|
labels.alert_identity |
Stable Grafana rule UID for deduplication and cross-export matching. |
labels.component |
Static affected Omi component. |
labels.impact |
One of infrastructure, product, or user-experience. |
annotations.summary |
A short human description of the condition. |
annotations.user_impact |
What an Omi user may experience; state uncertainty when impact is only potential. |
annotations.scope |
The bounded service or user path covered by the signal. |
annotations.verification |
The next evidence to confirm before intervention. |
annotations.safe_next_action |
The reversible, documented operator action after verification. |
The tiers describe evidence, not paging severity:
infrastructure: capacity, readiness, resource, or scrape signals. They may warn before a user is affected.product: evidence that an Omi service capability may fail, be delayed, or be unavailable.user-experience: observed degradation of a real Omi user path.
Product and user-experience rules must use production, real-traffic evidence. Do not treat synthetic checks, load tests,
staging traffic, or an inferred proxy metric as proof of user impact. Infrastructure rules may use service-health metrics,
but their user_impact annotation must make potential impact clear rather than claim a confirmed user outage.
Never place raw provider responses, exception text, stack traces, or dynamic Grafana error output in these annotations. Keep the message human-readable and direct the operator to the linked dashboard for bounded verification. Do not change Grafana credentials, contact points, or routing configuration while adding a rule.
Status codes and request counts cannot express every production failure. Four rules exist
specifically for failures that stay green on every other signal, and are documented together in
silent-failure-detection.md:
| Rule | Catches |
|---|---|
omi-capture-finalization-memory-fence |
The first real finalization rejected by a missing, disabled, or invalid canonical-memory runtime fence; zero is the only safe rate. |
omi-llm-gateway-invalid-requests |
Requests rejected during validation, before a route is selected. These are counted by llm_gateway_request_rejections_total and never reach llm_gateway_requests_total, so the affected lane keeps reporting 100% success. |
omi-llm-gateway-lane-failure-ratio |
A lane failing more than a quarter of its real requests over an hour. |
omi-llm-gateway-lane-zero-success |
A lane with attempts but no successful request in six hours. The or ... * 0 zero-fill is required: a lane that has never succeeded has no outcome="success" series, so a plain ratio produces no series and no alert. |
omi-journey-signal-dead |
A journey counter that stopped reporting while the platform is demonstrably serving traffic. Every real-traffic journey rule assumes its counter is scraped; when that breaks, the rule goes quiet rather than failing loudly. |
Thresholds on these rules are set from measured production distributions, not intuition, and the
measurement belongs in the rule's threshold annotation and in the runbook.
A rule must not be added for a signal that is not yet ingested. omi-journey-signal-dead excludes
journey="chat_response" because that counter is emitted from Cloud Run, which Prometheus does not
scrape; adding it before an ingestion path exists would page permanently and get muted.
Parakeet stream-capacity alerts use the existing parakeet_active_streams gauge divided by ready Parakeet replicas:
sum(parakeet_active_streams{container="parakeet", namespace="prod-omi-backend"})
/ clamp_min(sum(kube_deployment_status_replicas_ready{
deployment="prod-omi-parakeet", namespace="prod-omi-backend"}), 1)
This avoids a cluster-total threshold that would become misleading as HPA replica count changes. The warning threshold is
15 streams per ready replica for five minutes and the critical threshold is 20 for two minutes, both below the known
approximately 25–30 stream per-replica capability. No-data is healthy for these capacity rules: absent metrics are not
evidence of saturation. Operator response is in
parakeet-stream-capacity.md.
Dashboards are created and edited directly in the Grafana UI — it's purpose-built for visual iteration. Git serves as backup, version control, and the restore source for disaster recovery or new cluster setup.
Create/edit dashboard in Grafana UI
↓
Export JSON via API
↓
Commit to dashboards/<folder>/<name>.json
↓
PR review + merge
Restore flow (new cluster, disaster recovery, dev→prod promotion):
Read JSON from repo → Import via Grafana API → Dashboard is live
backend/charts/monitoring/
├── dashboards/ # Grafana dashboard JSON (source of truth)
│ ├── general/ # Matches Grafana "General" folder
│ │ ├── backend-api-monitoring-v1.json
│ │ ├── backend-api-monitoring-v2.json
│ │ └── parakeet-asr-monitoring.json
│ ├── cloud-run/ # Matches Grafana "Cloud Run" folder
│ │ ├── backend.json
│ │ ├── backend-integration.json
│ │ ├── backend-sync.json
│ │ └── plugins.json
│ ├── gke/ # Matches Grafana "GKE" folder
│ │ ├── backend-listen.json
│ │ ├── deepgram-self-hosted.json
│ │ ├── diarizer.json
│ │ ├── pusher.json
│ │ └── vad.json
│ └── omi-services/ # Matches Grafana "Omi Services" folder
│ ├── cloud-armor-denied-requests.json
│ ├── cloud-run-services-logs.json
│ ├── global-external-alb.json
│ ├── omi-kubernetes-events.json
│ └── resilience-fallbacks.json
├── alerts/ # (proposed) PrometheusRule or Grafana alert YAML
│ └── ...
├── kube-prometheus-stack/ # existing
├── prometheus-adapter/ # existing
├── alloy/ # existing
├── loki/ # existing
├── prometheus-stackdriver-exporter/ # existing
└── README.md # this file
Same-PR rule: When a PR adds, renames, or removes Prometheus metrics from application code, the PR should also update:
- The prometheus-adapter rules (if the metric drives HPA)
- The service's metrics contract (see below)
- Update the Grafana dashboard in the UI, then export and commit the JSON in the same PR (or a follow-up PR filed as an issue before merging)
- After any dashboard edit in Grafana UI: export and commit within the same day
- Periodic bulk sync: run the bulk export script (see "Sync-Back Workflow" in Developer Guide) to catch any missed UI edits
- Before major infra changes (cluster migration, Helm upgrades): verify repo JSONs match live dashboards
Move alert rules from Grafana UI into version-controlled config. Two options:
Option A: PrometheusRule CRDs (recommended for metric-based alerts)
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: parakeet-alerts
labels:
release: prod-omi-kube-prometheus-stack
spec:
groups:
- name: parakeet
rules:
- alert: ParakeetHighGPUUtil
expr: avg(DCGM_FI_DEV_GPU_UTIL{container="parakeet"}) > 90
for: 5m
labels:
severity: warning
annotations:
summary: "Parakeet GPU utilization above 90% for 5m"Option B: Grafana provisioning YAML (for log-based or multi-datasource alerts)
Place alert YAML in the alerts/ directory and configure Grafana sidecar to load from it.
Each service that exposes Prometheus metrics should document them alongside its ServiceMonitor or in its chart's values comments. Minimum contract:
| Field | Example |
|---|---|
| Metric name | parakeet_active_streams |
| Type | Gauge |
| Labels | namespace, pod |
| Unit | connections |
| Used by | parakeet HPA, GPU dashboard |
This ensures dashboard authors and HPA configs stay in sync with the application. When a metric is renamed or removed, the contract makes it clear what downstream consumers need updating.
Periodically audit dashboards for references to metrics that no longer exist:
- Export all dashboard JSON from Grafana
- Extract all PromQL metric names from the JSON
- Query Prometheus for each metric:
count({__name__="metric_name"})— zero means the metric is gone - Flag dashboards with dead metrics for cleanup
This can be scripted and run as a periodic check or pre-merge CI step once dashboards are in the repo.
Run from backend/charts/monitoring/. Ensure repos are added first:
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo add grafana https://grafana.github.io/helm-charts
helm repo updateProd (release names use prod-omi- prefix):
# kube-prometheus-stack
helm -n prod-omi-monitoring upgrade --install prod-omi-kube-prometheus-stack \
prometheus-community/kube-prometheus-stack \
-f kube-prometheus-stack/prod_omi_monitoring_values.yaml
# prometheus-adapter
helm -n prod-omi-monitoring upgrade --install prod-omi-prometheus-adapter \
prometheus-community/prometheus-adapter \
-f prometheus-adapter/prod_omi_prometheus_adapter.yaml
# Loki
helm -n prod-omi-monitoring upgrade --install prod-omi-loki \
grafana/loki \
-f loki/prod_omi_loki_values.yaml
# Alloy (k8s-monitoring) — release name is prod-omi-alloy (not prod-omi-k8s-monitoring)
helm -n prod-omi-monitoring upgrade --install prod-omi-alloy \
grafana/k8s-monitoring \
-f alloy/prod_omi_k8s_monitoring_values.yml
# Stackdriver exporter
helm -n prod-omi-monitoring upgrade --install prod-omi-prometheus-stackdriver-exporter \
prometheus-community/prometheus-stackdriver-exporter \
-f prometheus-stackdriver-exporter/prod_omi_stackdriver_exporter.yaml
# Isolated Cloud Run application-metrics bridge
helm -n prod-omi-monitoring upgrade --install prod-omi-cloud-run-metrics-exporter \
prometheus-community/prometheus-stackdriver-exporter \
-f prometheus-stackdriver-exporter/prod_omi_cloud_run_metrics_exporter.yamlDev (note: kube-prometheus-stack release name is dev-kube-prometheus-stack, not dev-omi-kube-prometheus-stack):
helm -n dev-omi-monitoring upgrade --install dev-kube-prometheus-stack \
prometheus-community/kube-prometheus-stack \
-f kube-prometheus-stack/dev_omi_monitoring_values.yaml
helm -n dev-omi-monitoring upgrade --install dev-omi-prometheus-adapter \
prometheus-community/prometheus-adapter \
-f prometheus-adapter/dev_omi_prometheus_adapter.yaml
helm -n dev-omi-monitoring upgrade --install dev-omi-loki \
grafana/loki \
-f loki/dev_omi_loki_values.yaml
helm -n dev-omi-monitoring upgrade --install dev-omi-alloy \
grafana/k8s-monitoring \
-f alloy/dev_omi_k8s_monitoring_values.yml
helm -n dev-omi-monitoring upgrade --install dev-omi-prometheus-stackdriver-exporter \
prometheus-community/prometheus-stackdriver-exporter \
-f prometheus-stackdriver-exporter/dev_omi_stackdriver_exporter.yaml
helm -n dev-omi-monitoring upgrade --install dev-omi-cloud-run-metrics-exporter \
prometheus-community/prometheus-stackdriver-exporter \
-f prometheus-stackdriver-exporter/dev_omi_cloud_run_metrics_exporter.yamlMetric not appearing in Prometheus:
- Check the pod exposes
/metricsand returns valid Prometheus format - For ServiceMonitor: verify the
releaselabel matches Prometheus'sserviceMonitorSelector - For annotations: verify the scrape job exists in
additionalScrapeConfigs - Check Prometheus targets: access Prometheus UI → Status → Targets to see scrape status. To query the metric: Grafana → Explore → Prometheus datasource → metric name
HPA shows <unknown> for custom metric:
- Verify the metric exists:
kubectl get --raw "/apis/external.metrics.k8s.io/v1beta1/namespaces/{ns}/{metric}" - Check prometheus-adapter logs:
kubectl logs -n {env}-omi-monitoring -l app.kubernetes.io/name=prometheus-adapter - Verify the adapter rule's
seriesQuerymatches actual Prometheus series - Only ONE prometheus-adapter can own the
v1beta1.custom.metrics.k8s.ioAPIService — never deploy a second instance
Logs not appearing in Loki:
- Verify Alloy is running:
kubectl get pods -n {env}-omi-monitoring -l app.kubernetes.io/name=alloy-logs - Check Alloy collects from the right namespace (configured in Alloy values)
- Verify Loki gateway is reachable and basic auth secret exists
- Query in Grafana Explore with Loki datasource:
{namespace="{env}-omi-backend"}