Observability & Alerts
The bundle
A Prometheus alert bundle for the operational silent-failure patterns this system has hit in
production ships under config/observability/. It is not part of config/default — the
alerts are opt-in so that installs without prometheus-operator CRDs are not affected by an
unknown apiVersion.
kubectl apply -k config/observability
This ships four resources in the inference-cache-system namespace:
- a
ServiceMonitor— scrapesinference-cache-server:8080/metrics; - a
PodMonitor— scrapes the controller pod’s:8080/metrics(required for the controller-side alerts to have a series to evaluate — the controller has no Service in front of it); - a second
PodMonitor— discovers successfully injected PodLocal LMCache native sidecars across workload namespaces and scrapeslmcache-http:8080/metrics; - a
PrometheusRule— the alerts.
make verify-prometheus lints and unit-tests the rules.
Selector mismatch fails silently
All four custom resources carry example labels (prometheus: k8s) matching the upstream
kube-prometheus stack. The kube-prometheus-stack Helm chart uses a different convention
(release: <name>, no prometheus: label). If your Prometheus custom resource’s
ruleSelector / serviceMonitorSelector / podMonitorSelector uses a different label set,
kubectl apply -k succeeds but Prometheus silently ignores the resources. Check your
selectors with kubectl get prometheus -A -o yaml and relabel the CRs to match.For vanilla Prometheus (ConfigMap mounts, prometheus.serverFiles), use the flat
config/observability/alerting-rules.yaml and configure scraping yourself — for the
server, controller pod, and injected PodLocal LMCache sidecars. Scope per install with kubernetes_sd_configs + a
relabel_configs that copies __meta_kubernetes_namespace to namespace (the alerts scope
per install by that label).
The alerts
| Alert | Severity | For | Fires when |
|---|---|---|---|
IndexEmpty | critical | 2m | The server is up but the index is empty while lookups are flowing — ingestion is broken. |
LookupRouteDegenerate | warning | 5m | More than 90% of lookups return NO_HINT over 10m for a model — no reuse, or a client sending wrong contract keys. |
LookupRouteHighTimeout | warning | 5m | More than 5% of lookups return TIMEOUT over 10m for a model. |
IndexEvictionsSpike | info | 10m | More than 10 cap-evictions/sec (reason="cap"; TTL evictions excluded) — the index is over budget. |
ServerProbeFail | critical | 5m | Two or more functional-probe failures in 5m (controller-emitted — needs the PodMonitor). |
LMCacheT2NoHits | warning | 5m | More than 1000 tier-2 query tokens/sec but zero hit tokens — a silently-degraded offload tier. |
Five of these (IndexEmpty, LookupRouteDegenerate, LookupRouteHighTimeout,
IndexEvictionsSpike, ServerProbeFail) become active as soon as the operator is installed
and the selectors match; they stay quiet on a healthy or idle install (each is gated by
traffic/rate/eviction thresholds).
LMCacheT2NoHits needs an extra scrape
LMCacheT2NoHits reads vllm:external_prefix_cache_* from the engine pods directly, not
from inference-cache. The shipped scrape configs do not collect vLLM’s own metrics. To
make this alert effective, add a separate PodMonitor for your engine Deployment (scoped to
the pods your CacheBackend.spec.engineSelector matches) that preserves both namespace and
pod labels.Related pages
- Metrics reference — the underlying series.
- Index sizing — the eviction/memory relationship
behind
IndexEvictionsSpike. - Monitor the cache plane — the day-to-day view.