Metrics

Every inferencecache_* series, its labels, and which binary emits it.

Both inference-cache binaries expose Prometheus metrics on their pod’s :8080/metrics, all prefixed inferencecache_*. The two binaries use separate registries — the server’s series cover the index and gRPC handlers; the controller’s cover the reconcilers. Standard go_* and process_* collectors are also present but are not part of this schema. Successfully injected PodLocal LMCache sidecars expose their upstream lmcache_mp_* metrics separately; the observability overlay includes a cross-namespace PodMonitor for them.

Server metrics (inference-cache-server)

Gauges

MetricLabelsMeaning
inferencecache_server_up—1 when the server is serving.
inferencecache_server_grpc_tls_enabled—1 when gRPC TLS is configured.
inferencecache_index_entriesmodelDistinct prefix keys in the index (excludes the reserved probe tenant). Drained models are zeroed, not left stale.

Counters

MetricLabelsMeaning
inferencecache_lookup_route_calls_totalmodel, reason_code, hint_usedLookups by outcome. hint_used="true" ⇔ a non-empty score list (PREFIX_MATCH / TENANT_HOT / AFFINITY_HINT).
inferencecache_index_evictions_totalalgorithm (lru/lfu), reason (cap/ttl)Index evictions. reason="cap" is the authoritative over-budget signal.
inferencecache_tenant_evictions_totaltenant_id, reason (over_entries)Per-tenant quota evictions. A bounded non-zero rate is normal Fairness.
inferencecache_snapshot_auth_totalresult (ok/unauth/forbidden/error)/snapshot auth outcomes. Wrong audience lands in unauth.
inferencecache_policy_auth_totalresult/policy auth outcomes.
inferencecache_probe_auth_totalresult/probe auth outcomes.

Histogram

MetricLabelsBuckets
inferencecache_lookup_route_latency_secondsmodel100µs, 250µs, 500µs, 1ms, 2.5ms, 5ms, 10ms, 25ms, 50ms, 100ms

Controller metrics (inferencecache-controller)

The controller pod has no Service in front of it — scrape it with a PodMonitor (the shipped observability bundle includes one) or pod-IP service discovery, or the controller-side alerts have no series to evaluate.

Gauge

MetricLabelsMeaning
inferencecache_backend_t2_hit_ratebackend (<ns>/<name>)Tier-2 offload hit rate. Present once exercised; 0 while queries flow = a degraded tier.

Counters

MetricLabelsMeaning
inferencecache_backend_probe_result_totalbackend, stage (ingest/routing/t2), result (ok/failed/skipped)Functional-probe stage results — three increments per successful call.
inferencecache_backend_t2_query_tokens_totalbackendMonotonic tier-2 activity signal (positive deltas only).

Endpoints

BinaryPortPaths
Server:8080 (--http-bind-address)/healthz, /readyz (unauth), /metrics
Server:8081 (--snapshot-bind-address)/snapshot, /policy, /probe (auth-gated)
Controller:8080 (--metrics-bind-address)/metrics (unauth by default; --metrics-secure to gate)
Injected LMCache MP sidecar:8080 (lmcache-http)/healthcheck, /metrics (unauthenticated on the Pod network)
Last modified August 13, 2026: LMCache MP mode (#187) (40a9806)