Concepts
This section explains the components, APIs, and abstractions inference-cache uses to make LLM inference routing cache-aware.
Start here
Architecture
The two control-plane binaries (controller + server), the observation sidecar, and the bidirectional in-cluster bridge that connects them — and the invariants (fail-open, soft state, “we hint, the gateway decides”) that shape every design decision.
The API — Custom Resources
inference-cache defines six CRDs in the group inferencecache.io, version v1alpha1.
Five are namespaced; CacheIndex is the only cluster-scoped type. Controllers exist for
CacheBackend and CacheIndex today; the others are declarative.
CacheBackend
The primary resource an operator writes. It binds to inference-engine pods by
label, injects typed engine-side cache wiring, and optionally provisions managed
Redis as a remote tier. Short name cb.
CachePolicy
Opt-in, per-namespace tuning of cache lookup and eviction — matched-token floors, score
floors, timeouts, eviction algorithm and TTL, and the ranking strategy gates. Short name
cpol.
CacheTenant
Gives an external tenant a stable identity the index isolates on, plus an optional
entry-count quota. Short name ct.
CacheIndex
A cluster-scoped, status-only singleton that mirrors the server’s in-memory aggregate — the
cache “world map” for observability. Short name ci.
PromptTemplate
Declares a cache-aware prompt template and its stable/mutable slots for the mutable-slot
render pipeline. Short name pt.
PDTopology
Declares prefill/decode topology for phase-disaggregated serving. Short name pdt.
The data plane API — gRPC
LookupRoute & ranking
How the core cache-aware routing hint works: prefix matching, longest-prefix block-chain
matching, the ranking formula, TENANT_HOT and affinity fallbacks, and the diagnostic
reason codes.
The gRPC contract
The InferenceCache service — its eight RPCs, message shapes, and the guarantees
(fail-open, metadata-only, string reason codes, additive deltas) that let clients depend on
it.