<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>inference-cache</title><link>https://cachebox-project.github.io/inference-cache/</link><description>Recent content on inference-cache</description><generator>Hugo</generator><language>en</language><atom:link href="https://cachebox-project.github.io/inference-cache/index.xml" rel="self" type="application/rss+xml"/><item><title>Architecture</title><link>https://cachebox-project.github.io/inference-cache/docs/concepts/architecture/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/concepts/architecture/</guid><description>&lt;h2 id="one-operator-two-binaries"&gt;
One operator, two binaries
&lt;a href="#one-operator-two-binaries" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;inference-cache is a single operator split across two control-plane binaries, plus a CLI
and an in-cluster sidecar. The split follows the responsibilities: the &lt;strong&gt;controller&lt;/strong&gt; owns
the Kubernetes reconcile loop and admission; the &lt;strong&gt;server&lt;/strong&gt; owns the routing decision and
the cache-state index.&lt;/p&gt;</description></item><item><title>Contributing</title><link>https://cachebox-project.github.io/inference-cache/docs/developer-guide/contributing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/developer-guide/contributing/</guid><description>&lt;p&gt;This page summarizes the contributor workflow. The authoritative version is
&lt;a href="https://github.com/cachebox-project/inference-cache/blob/main/CONTRIBUTING.md"&gt;&lt;code&gt;CONTRIBUTING.md&lt;/code&gt;&lt;/a&gt;
in the repository.&lt;/p&gt;


&lt;h2 id="local-setup"&gt;
Local setup
&lt;a href="#local-setup" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;Requirements: &lt;strong&gt;Go 1.26.6+&lt;/strong&gt;, Make, Docker, kind, and protoc.&lt;/p&gt;
&lt;p&gt;Install the git hooks once per clone — they enforce the project rules below:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;make install-hooks
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Run the baseline checks before sending a PR:&lt;/p&gt;</description></item><item><title>CRD API</title><link>https://cachebox-project.github.io/inference-cache/docs/reference/crd-api/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/reference/crd-api/</guid><description>&lt;p&gt;All CRDs are in the API group &lt;strong&gt;&lt;code&gt;inferencecache.io&lt;/code&gt;&lt;/strong&gt;, version &lt;strong&gt;&lt;code&gt;v1alpha1&lt;/code&gt;&lt;/strong&gt;.&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Kind&lt;/th&gt;
					&lt;th&gt;Scope&lt;/th&gt;
					&lt;th&gt;Short name&lt;/th&gt;
					&lt;th&gt;Reconciled?&lt;/th&gt;
					&lt;th&gt;Purpose&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;CacheBackend&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;Namespaced&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;cb&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;Yes&lt;/td&gt;
					&lt;td&gt;Bind engine Pods to typed MP; optionally provision a remote Redis provider.&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;CachePolicy&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;Namespaced&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;cpol&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;Declarative (pushed to server)&lt;/td&gt;
					&lt;td&gt;Per-namespace lookup and eviction tuning.&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;CacheTenant&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;Namespaced&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;ct&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;Declarative (pushed to server)&lt;/td&gt;
					&lt;td&gt;Tenant identity + entry-count quota.&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;CacheIndex&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;&lt;strong&gt;Cluster&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;ci&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;Yes (status-only)&lt;/td&gt;
					&lt;td&gt;Cluster-wide mirror of the server aggregate.&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;PromptTemplate&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;Namespaced&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;pt&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;Declarative&lt;/td&gt;
					&lt;td&gt;Cache-aware prompt template + stable/mutable slots.&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;PDTopology&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;Namespaced&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;pdt&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;Declarative&lt;/td&gt;
					&lt;td&gt;Prefill/decode topology for disaggregated serving.&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;


&lt;div class="alert alert-info" role="alert"&gt;
&lt;h4 class="alert-heading"&gt;v1alpha1 stability&lt;/h4&gt;

 &lt;code&gt;v1alpha1&lt;/code&gt; is not compatibility-frozen before the first production deployment. The CRD may
make breaking schema corrections while the API shape is being finalized. The gRPC/proto
contract follows its own compatibility policy for external consumers.

&lt;/div&gt;



&lt;h2 id="cachebackend"&gt;
CacheBackend
&lt;a href="#cachebackend" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Key &lt;code&gt;spec&lt;/code&gt; fields:&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Deploy a CacheBackend</title><link>https://cachebox-project.github.io/inference-cache/docs/tasks/deploy-a-cache-backend/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/tasks/deploy-a-cache-backend/</guid><description>&lt;p&gt;Get a cache-aware inference setup running in about five minutes. This assumes the
inference-cache operator (controller + policy server + CRDs) is already
&lt;a href="https://cachebox-project.github.io/inference-cache/docs/installation/"&gt;installed&lt;/a&gt;.&lt;/p&gt;


&lt;h2 id="1-write-the-minimum-cachebackend"&gt;
1. Write the minimum CacheBackend
&lt;a href="#1-write-the-minimum-cachebackend" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;A &lt;code&gt;CacheBackend&lt;/code&gt; binds to your inference-engine pods by label and makes their KV cache
reusable across requests. Here is the minimum-viable spec:&lt;/p&gt;</description></item><item><title>Observability &amp; Alerts</title><link>https://cachebox-project.github.io/inference-cache/docs/administration/observability-and-alerts/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/administration/observability-and-alerts/</guid><description>&lt;h2 id="the-bundle"&gt;
The bundle
&lt;a href="#the-bundle" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;A Prometheus alert bundle for the operational silent-failure patterns this system has hit in
production ships under &lt;code&gt;config/observability/&lt;/code&gt;. It is &lt;strong&gt;not&lt;/strong&gt; part of &lt;code&gt;config/default&lt;/code&gt; — the
alerts are opt-in so that installs without prometheus-operator CRDs are not affected by an
unknown &lt;code&gt;apiVersion&lt;/code&gt;.&lt;/p&gt;</description></item><item><title>Bind an engine</title><link>https://cachebox-project.github.io/inference-cache/docs/tasks/bind-an-engine/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/tasks/bind-an-engine/</guid><description>&lt;p&gt;Binding is how a &lt;code&gt;CacheBackend&lt;/code&gt; claims inference-engine Pods in its namespace.
The engine Deployment and image remain inference-system owned.&lt;/p&gt;


&lt;h2 id="lifecycle"&gt;
Lifecycle
&lt;a href="#lifecycle" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Apply a typed &lt;code&gt;CacheBackend&lt;/code&gt; with &lt;code&gt;spec.lmCache.topology: PodLocal&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Create engine Pods whose labels include every
&lt;code&gt;spec.engineSelector.matchLabels&lt;/code&gt; entry.&lt;/li&gt;
&lt;li&gt;At Pod CREATE, the webhook atomically injects the LMCache MP native sidecar,
shared memory, and the runtime-specific connector wire.&lt;/li&gt;
&lt;li&gt;When configured, the &lt;code&gt;kvevent-subscriber&lt;/code&gt; sidecar reports KV events to the
routing index.&lt;/li&gt;
&lt;/ol&gt;


&lt;div class="alert alert-warning" role="alert"&gt;
&lt;h4 class="alert-heading"&gt;The match is evaluated once, at Pod CREATE&lt;/h4&gt;

 Relabeling an existing Pod does not re-run admission. Recreate the Pod after
changing labels or CacheBackend configuration.

&lt;/div&gt;



&lt;h2 id="typed-vllm-mp-wire"&gt;
Typed vLLM MP wire
&lt;a href="#typed-vllm-mp-wire" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The vLLM adapter injects &lt;code&gt;LMCacheMPConnector&lt;/code&gt; with the module path
&lt;code&gt;lmcache.integration.vllm.lmcache_mp_connector&lt;/code&gt;, loopback host/port,
&lt;code&gt;--disable-hybrid-kv-cache-manager&lt;/code&gt;, &lt;code&gt;PYTHONHASHSEED=0&lt;/code&gt;, and the fail-open
value. Reserved entries are:&lt;/p&gt;</description></item><item><title>CacheBackend</title><link>https://cachebox-project.github.io/inference-cache/docs/concepts/cachebackend/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/concepts/cachebackend/</guid><description>&lt;h2 id="what-is-a-cachebackend"&gt;
What is a CacheBackend?
&lt;a href="#what-is-a-cachebackend" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;A namespaced &lt;code&gt;CacheBackend&lt;/code&gt; selects an inference runtime, its engine-side cache
integration, and an optional remote L3. The inference system owns the engine
Deployment and image. CacheBackend injects the selected connector and the cache
components it owns into matching Pods.&lt;/p&gt;</description></item><item><title>gRPC API</title><link>https://cachebox-project.github.io/inference-cache/docs/reference/grpc-api/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/reference/grpc-api/</guid><description>&lt;p&gt;Service &lt;strong&gt;&lt;code&gt;InferenceCache&lt;/code&gt;&lt;/strong&gt;, package &lt;strong&gt;&lt;code&gt;inferencecache.v1alpha1&lt;/code&gt;&lt;/strong&gt;, proto
&lt;code&gt;proto/inferencecache/v1alpha1/inferencecache.proto&lt;/code&gt;. Served on &lt;code&gt;:9090&lt;/code&gt; (plaintext by
default; &lt;a href="https://cachebox-project.github.io/inference-cache/docs/administration/grpc-tls/"&gt;TLS overlay&lt;/a&gt; available). See
&lt;a href="https://cachebox-project.github.io/inference-cache/docs/concepts/grpc-contract/"&gt;The gRPC contract&lt;/a&gt; for the guarantees; this page is the
per-RPC reference.&lt;/p&gt;


&lt;h2 id="consumer-rpcs"&gt;
Consumer RPCs
&lt;a href="#consumer-rpcs" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;


&lt;h3 id="lookuproutelookuprouterequest--lookuprouteresponse"&gt;
&lt;code&gt;LookupRoute(LookupRouteRequest) → LookupRouteResponse&lt;/code&gt;
&lt;a href="#lookuproutelookuprouterequest--lookuprouteresponse" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;The core cache-aware routing hint. Unary and fail-open. It does not mutate cache membership
or response-visible routing state. It emits metrics, and a delivered prefix hit under an LFU
policy credits the matched entries&amp;rsquo; internal eviction access counters.&lt;/p&gt;</description></item><item><title>Index sizing</title><link>https://cachebox-project.github.io/inference-cache/docs/administration/index-sizing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/administration/index-sizing/</guid><description>&lt;p&gt;The server&amp;rsquo;s cache-state index lives in memory. This page helps you size the server pod and
choose eviction settings so the index never approaches the pod&amp;rsquo;s memory limit (an OOM-kill
is the one failure the soft-state design cannot hide).&lt;/p&gt;


&lt;h2 id="pod-budget-at-a-glance"&gt;
Pod budget at a glance
&lt;a href="#pod-budget-at-a-glance" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Storage entries (distinct prefix keys × replicas reporting each)&lt;/th&gt;
					&lt;th&gt;Peak RSS&lt;/th&gt;
					&lt;th&gt;Recommended pod memory&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;100K storage entries&lt;/td&gt;
					&lt;td&gt;~110 MiB&lt;/td&gt;
					&lt;td&gt;256 MiB&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;500K storage entries&lt;/td&gt;
					&lt;td&gt;~300 MiB&lt;/td&gt;
					&lt;td&gt;512 MiB&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;1M storage entries (the default cap)&lt;/td&gt;
					&lt;td&gt;~540 MiB&lt;/td&gt;
					&lt;td&gt;&lt;strong&gt;1 GiB (recommended floor)&lt;/strong&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;1.5M storage entries&lt;/td&gt;
					&lt;td&gt;~700 MiB&lt;/td&gt;
					&lt;td&gt;1.5 GiB&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;These include roughly 20% headroom over measured heap. The default entry cap is 1,000,000.&lt;/p&gt;</description></item><item><title>CachePolicy</title><link>https://cachebox-project.github.io/inference-cache/docs/concepts/cachepolicy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/concepts/cachepolicy/</guid><description>&lt;h2 id="what-is-a-cachepolicy"&gt;
What is a CachePolicy?
&lt;a href="#what-is-a-cachepolicy" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;A &lt;code&gt;CachePolicy&lt;/code&gt; tunes how the server filters and ranks lookups, and how it evicts, for a
single namespace. It is &lt;strong&gt;entirely opt-in&lt;/strong&gt;: the server ships sane defaults, so most
namespaces need no &lt;code&gt;CachePolicy&lt;/code&gt; at all. It is &lt;strong&gt;purely declarative&lt;/strong&gt; — the controller
flattens all policies and pushes the resolved values to the server; the reconciler never
writes &lt;code&gt;CachePolicy.status&lt;/code&gt;.&lt;/p&gt;</description></item><item><title>gRPC TLS</title><link>https://cachebox-project.github.io/inference-cache/docs/administration/grpc-tls/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/administration/grpc-tls/</guid><description>&lt;h2 id="the-decision"&gt;
The decision
&lt;a href="#the-decision" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The gRPC API on &lt;code&gt;:9090&lt;/code&gt; uses &lt;strong&gt;one-sided Service TLS via cert-manager, terminated
in-process&lt;/strong&gt; — no sidecar, Envoy, or Ingress. Mutual TLS (client-certificate auth) is a
Phase-2 feature flag, not yet implemented.&lt;/p&gt;
&lt;p&gt;The server binary supports TLS via &lt;code&gt;--tls-cert-file&lt;/code&gt; / &lt;code&gt;--tls-key-file&lt;/code&gt; (both-or-neither —
supplying exactly one fails startup). The certificate is served through a reloading
&lt;code&gt;GetCertificate&lt;/code&gt; hook, so a cert-manager-rotated Secret is picked up &lt;strong&gt;without a restart&lt;/strong&gt;.
The posture is exposed as the gauge &lt;code&gt;inferencecache_server_grpc_tls_enabled&lt;/code&gt; (0/1).&lt;/p&gt;</description></item><item><title>Metrics</title><link>https://cachebox-project.github.io/inference-cache/docs/reference/metrics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/reference/metrics/</guid><description>&lt;p&gt;Both inference-cache binaries expose Prometheus metrics on their pod&amp;rsquo;s &lt;code&gt;:8080/metrics&lt;/code&gt;, all prefixed
&lt;code&gt;inferencecache_*&lt;/code&gt;. The two binaries use &lt;strong&gt;separate registries&lt;/strong&gt; — the server&amp;rsquo;s series cover
the index and gRPC handlers; the controller&amp;rsquo;s cover the reconcilers. Standard &lt;code&gt;go_*&lt;/code&gt; and
&lt;code&gt;process_*&lt;/code&gt; collectors are also present but are not part of this schema. Successfully
injected PodLocal LMCache sidecars expose their upstream &lt;code&gt;lmcache_mp_*&lt;/code&gt; metrics separately;
the observability overlay includes a cross-namespace PodMonitor for them.&lt;/p&gt;


&lt;h2 id="server-metrics-inference-cache-server"&gt;
Server metrics (&lt;code&gt;inference-cache-server&lt;/code&gt;)
&lt;a href="#server-metrics-inference-cache-server" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;


&lt;h3 id="gauges"&gt;
Gauges
&lt;a href="#gauges" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h3&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Metric&lt;/th&gt;
					&lt;th&gt;Labels&lt;/th&gt;
					&lt;th&gt;Meaning&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;inferencecache_server_up&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;—&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;1&lt;/code&gt; when the server is serving.&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;inferencecache_server_grpc_tls_enabled&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;—&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;1&lt;/code&gt; when gRPC TLS is configured.&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;inferencecache_index_entries&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;model&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;Distinct prefix keys in the index (excludes the reserved probe tenant). Drained models are zeroed, not left stale.&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;


&lt;h3 id="counters"&gt;
Counters
&lt;a href="#counters" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h3&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Metric&lt;/th&gt;
					&lt;th&gt;Labels&lt;/th&gt;
					&lt;th&gt;Meaning&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;inferencecache_lookup_route_calls_total&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;model&lt;/code&gt;, &lt;code&gt;reason_code&lt;/code&gt;, &lt;code&gt;hint_used&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;Lookups by outcome. &lt;code&gt;hint_used=&amp;quot;true&amp;quot;&lt;/code&gt; ⇔ a non-empty score list (&lt;code&gt;PREFIX_MATCH&lt;/code&gt; / &lt;code&gt;TENANT_HOT&lt;/code&gt; / &lt;code&gt;AFFINITY_HINT&lt;/code&gt;).&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;inferencecache_index_evictions_total&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;algorithm&lt;/code&gt; (&lt;code&gt;lru&lt;/code&gt;/&lt;code&gt;lfu&lt;/code&gt;), &lt;code&gt;reason&lt;/code&gt; (&lt;code&gt;cap&lt;/code&gt;/&lt;code&gt;ttl&lt;/code&gt;)&lt;/td&gt;
					&lt;td&gt;Index evictions. &lt;code&gt;reason=&amp;quot;cap&amp;quot;&lt;/code&gt; is the authoritative over-budget signal.&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;inferencecache_tenant_evictions_total&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;tenant_id&lt;/code&gt;, &lt;code&gt;reason&lt;/code&gt; (&lt;code&gt;over_entries&lt;/code&gt;)&lt;/td&gt;
					&lt;td&gt;Per-tenant quota evictions. A bounded non-zero rate is normal Fairness.&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;inferencecache_snapshot_auth_total&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;result&lt;/code&gt; (&lt;code&gt;ok&lt;/code&gt;/&lt;code&gt;unauth&lt;/code&gt;/&lt;code&gt;forbidden&lt;/code&gt;/&lt;code&gt;error&lt;/code&gt;)&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;/snapshot&lt;/code&gt; auth outcomes. Wrong audience lands in &lt;code&gt;unauth&lt;/code&gt;.&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;inferencecache_policy_auth_total&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;result&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;/policy&lt;/code&gt; auth outcomes.&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;inferencecache_probe_auth_total&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;result&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;/probe&lt;/code&gt; auth outcomes.&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;


&lt;h3 id="histogram"&gt;
Histogram
&lt;a href="#histogram" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h3&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Metric&lt;/th&gt;
					&lt;th&gt;Labels&lt;/th&gt;
					&lt;th&gt;Buckets&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;code&gt;inferencecache_lookup_route_latency_seconds&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;model&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;100µs, 250µs, 500µs, 1ms, 2.5ms, 5ms, 10ms, 25ms, 50ms, 100ms&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;


&lt;h2 id="controller-metrics-inferencecache-controller"&gt;
Controller metrics (&lt;code&gt;inferencecache-controller&lt;/code&gt;)
&lt;a href="#controller-metrics-inferencecache-controller" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The controller pod has &lt;strong&gt;no Service&lt;/strong&gt; in front of it — scrape it with a &lt;code&gt;PodMonitor&lt;/code&gt; (the
shipped observability bundle includes one) or pod-IP service discovery, or the controller-side
alerts have no series to evaluate.&lt;/p&gt;</description></item><item><title>Tune lookup and eviction</title><link>https://cachebox-project.github.io/inference-cache/docs/tasks/tune-lookup-and-eviction/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/tasks/tune-lookup-and-eviction/</guid><description>&lt;p&gt;Most namespaces need no &lt;code&gt;CachePolicy&lt;/code&gt; — the server ships sane defaults. Reach for one when
you want to change how lookups are filtered/ranked, or how the index evicts, in a specific
namespace. There is &lt;strong&gt;at most one &lt;code&gt;CachePolicy&lt;/code&gt; per namespace&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;See &lt;a href="https://cachebox-project.github.io/inference-cache/docs/concepts/cachepolicy/"&gt;CachePolicy&lt;/a&gt; for the field-by-field reference; this page
is worked examples.&lt;/p&gt;


&lt;h2 id="a-production-shaped-policy"&gt;
A production-shaped policy
&lt;a href="#a-production-shaped-policy" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-yaml" data-lang="yaml"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#204a87;font-weight:bold"&gt;apiVersion&lt;/span&gt;&lt;span style="color:#000;font-weight:bold"&gt;:&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#000"&gt;inferencecache.io/v1alpha1&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#204a87;font-weight:bold"&gt;kind&lt;/span&gt;&lt;span style="color:#000;font-weight:bold"&gt;:&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#000"&gt;CachePolicy&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#204a87;font-weight:bold"&gt;metadata&lt;/span&gt;&lt;span style="color:#000;font-weight:bold"&gt;:&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#204a87;font-weight:bold"&gt;name&lt;/span&gt;&lt;span style="color:#000;font-weight:bold"&gt;:&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#000"&gt;default&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#204a87;font-weight:bold"&gt;namespace&lt;/span&gt;&lt;span style="color:#000;font-weight:bold"&gt;:&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#000"&gt;serving&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#204a87;font-weight:bold"&gt;spec&lt;/span&gt;&lt;span style="color:#000;font-weight:bold"&gt;:&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#204a87;font-weight:bold"&gt;eviction&lt;/span&gt;&lt;span style="color:#000;font-weight:bold"&gt;:&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#000"&gt;LRU&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#204a87;font-weight:bold"&gt;evictionTTL&lt;/span&gt;&lt;span style="color:#000;font-weight:bold"&gt;:&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#000"&gt;30m&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#204a87;font-weight:bold"&gt;minimumPrefixTokens&lt;/span&gt;&lt;span style="color:#000;font-weight:bold"&gt;:&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#0000cf;font-weight:bold"&gt;16&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#8f5902;font-style:italic"&gt;# skip lookups for very short prompts&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#204a87;font-weight:bold"&gt;minimumMatchedTokens&lt;/span&gt;&lt;span style="color:#000;font-weight:bold"&gt;:&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#0000cf;font-weight:bold"&gt;64&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#8f5902;font-style:italic"&gt;# a replica needs 4 KV blocks of overlap to be offered&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#204a87;font-weight:bold"&gt;routingFloorScore&lt;/span&gt;&lt;span style="color:#000;font-weight:bold"&gt;:&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#4e9a06"&gt;&amp;#34;0.1&amp;#34;&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#8f5902;font-style:italic"&gt;# the top replica must clear this score&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#204a87;font-weight:bold"&gt;lookupTimeoutMs&lt;/span&gt;&lt;span style="color:#000;font-weight:bold"&gt;:&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#0000cf;font-weight:bold"&gt;5&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#8f5902;font-style:italic"&gt;# server-side time budget per lookup (fail-open TIMEOUT)&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#204a87;font-weight:bold"&gt;strategy&lt;/span&gt;&lt;span style="color:#000;font-weight:bold"&gt;:&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#204a87;font-weight:bold"&gt;enableChainMatching&lt;/span&gt;&lt;span style="color:#000;font-weight:bold"&gt;:&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#204a87;font-weight:bold"&gt;true&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#204a87;font-weight:bold"&gt;enableTenantHot&lt;/span&gt;&lt;span style="color:#000;font-weight:bold"&gt;:&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#204a87;font-weight:bold"&gt;true&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#204a87;font-weight:bold"&gt;affinityRouting&lt;/span&gt;&lt;span style="color:#000;font-weight:bold"&gt;:&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt; &lt;/span&gt;&lt;span style="color:#000"&gt;Enabled&lt;/span&gt;&lt;span style="color:#f8f8f8"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;

&lt;h2 id="recipes"&gt;
Recipes
&lt;a href="#recipes" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;


&lt;h3 id="reduce-noisy-hints-from-shared-system-prompts"&gt;
Reduce noisy hints from shared system prompts
&lt;a href="#reduce-noisy-hints-from-shared-system-prompts" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h3&gt;
&lt;p&gt;Chat templates give every request the same long system-prompt prefix, which would otherwise
match every replica. The default floors already handle this, but you can tighten them:&lt;/p&gt;</description></item><item><title>CacheTenant</title><link>https://cachebox-project.github.io/inference-cache/docs/concepts/cachetenant/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/concepts/cachetenant/</guid><description>&lt;h2 id="what-is-a-cachetenant"&gt;
What is a CacheTenant?
&lt;a href="#what-is-a-cachetenant" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;A &lt;code&gt;CacheTenant&lt;/code&gt; gives an external tenant two things:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;A stable &lt;strong&gt;identity&lt;/strong&gt; (&lt;code&gt;spec.tenantID&lt;/code&gt;) that the cache-state index isolates on, so one
tenant&amp;rsquo;s cache hints can never collide with another&amp;rsquo;s.&lt;/li&gt;
&lt;li&gt;An optional &lt;strong&gt;entry-count quota&lt;/strong&gt; that caps how many distinct prefix keys the tenant may
occupy in the index.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;code&gt;CacheTenant&lt;/code&gt; is namespaced, short name &lt;code&gt;ct&lt;/code&gt;.&lt;/p&gt;</description></item><item><title>Isolate tenants</title><link>https://cachebox-project.github.io/inference-cache/docs/tasks/isolate-tenants/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/tasks/isolate-tenants/</guid><description>&lt;p&gt;Multi-tenancy in inference-cache comes in two layers: &lt;strong&gt;structural identity isolation&lt;/strong&gt;
(always on) and an optional &lt;strong&gt;entry-count quota&lt;/strong&gt;. For byte-level isolation you use
Kubernetes, not a cache field — this page explains why and how.&lt;/p&gt;
&lt;p&gt;See &lt;a href="https://cachebox-project.github.io/inference-cache/docs/concepts/cachetenant/"&gt;CacheTenant&lt;/a&gt; for the field reference.&lt;/p&gt;


&lt;h2 id="identity-isolation-is-automatic"&gt;
Identity isolation is automatic
&lt;a href="#identity-isolation-is-automatic" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The index is keyed by &lt;code&gt;(tenant, model, hash_scheme, adapter_id, prefix_hash)&lt;/code&gt;. Because
&lt;code&gt;tenant&lt;/code&gt; is part of the key, one tenant&amp;rsquo;s hints can never collide with another&amp;rsquo;s, and
&lt;code&gt;LookupRoute&lt;/code&gt; is tenant-scoped. You get this the moment your producers report a &lt;code&gt;tenant_id&lt;/code&gt;
— you do not need a &lt;code&gt;CacheTenant&lt;/code&gt; for isolation to hold.&lt;/p&gt;</description></item><item><title>Reason codes</title><link>https://cachebox-project.github.io/inference-cache/docs/reference/reason-codes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/reference/reason-codes/</guid><description>&lt;p&gt;&lt;code&gt;reason_code&lt;/code&gt; is a &lt;strong&gt;string, not a proto enum&lt;/strong&gt; — new codes can be added without breaking old
clients. &lt;strong&gt;Clients must treat any unrecognized code as the no-hint default&lt;/strong&gt; (fail open,
round-robin).&lt;/p&gt;


&lt;h2 id="lookuproute--lookuppdroute"&gt;
LookupRoute / LookupPDRoute
&lt;a href="#lookuproute--lookuppdroute" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Code&lt;/th&gt;
					&lt;th&gt;When emitted&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;&lt;code&gt;PREFIX_MATCH&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;A replica holds the exact prefix or a leading block-run; the ranker returns a non-empty list; at least one replica clears &lt;code&gt;minimumMatchedTokens&lt;/code&gt; (default 64) &lt;strong&gt;and&lt;/strong&gt; the top score clears &lt;code&gt;routingFloorScore&lt;/code&gt; (default 0.1).&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;&lt;code&gt;NO_HINT&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;The fail-open default: a novel prefix with no affinity fallback, an empty/unspecified key, cold start (globally empty index), a request gated below &lt;code&gt;minimumPrefixTokens&lt;/code&gt;, a sub-floor matched-tokens/score result (with affinity disabled), or a disabled index.&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;&lt;code&gt;TENANT_HOT&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Prefix miss, &lt;code&gt;strategy.enableTenantHot ≠ false&lt;/code&gt;, and a warm replica exists (recent stats, &lt;code&gt;hit_rate&lt;/code&gt; above the floor, serving the requested scheme). &lt;code&gt;matched_tokens = 0&lt;/code&gt;.&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;&lt;code&gt;AFFINITY_HINT&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Would-be &lt;code&gt;NO_HINT&lt;/code&gt;, &lt;code&gt;affinityRouting: Enabled&lt;/code&gt;, a usable fingerprint, and ≥1 serving replica. A single stable replica; &lt;code&gt;score&lt;/code&gt; / &lt;code&gt;matched_tokens&lt;/code&gt; / &lt;code&gt;estimated_cache_hit_prob&lt;/code&gt; are all &lt;code&gt;0&lt;/code&gt;.&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;&lt;code&gt;POLICY_REQUIRES_CHAIN&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;&lt;code&gt;strategy.requireChain: true&lt;/code&gt; and the request has no valid block-hash chain. Returned before touching the index.&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;&lt;code&gt;TIMEOUT&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;The lookup deadline expired (context deadline, or &lt;code&gt;lookupTimeoutMs&lt;/code&gt; elapsed). Fail-open. Clients may also synthesize this locally.&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;&lt;code&gt;UNKNOWN_TENANT&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Miss, and the (non-empty) &lt;code&gt;tenant_id&lt;/code&gt; has zero entries anywhere (index not globally empty). A contract-key mismatch — likely a misconfigured client.&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;&lt;code&gt;UNKNOWN_MODEL&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Miss, tenant known, but &lt;code&gt;(tenant, model)&lt;/code&gt; has zero entries.&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;&lt;code&gt;UNKNOWN_HASH_SCHEME&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Miss, &lt;code&gt;(tenant, model)&lt;/code&gt; has entries, but none under the request&amp;rsquo;s &lt;code&gt;hash_scheme&lt;/code&gt;.&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;code&gt;LookupPDRoute&lt;/code&gt; is a stub today and always returns no hint.&lt;/p&gt;</description></item><item><title>Troubleshooting</title><link>https://cachebox-project.github.io/inference-cache/docs/administration/troubleshooting/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/administration/troubleshooting/</guid><description>&lt;h2 id="reading-cachebackend-readiness"&gt;
Reading CacheBackend readiness
&lt;a href="#reading-cachebackend-readiness" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;A managed backend&amp;rsquo;s &lt;code&gt;Ready&lt;/code&gt; is the composition of three gates
(&lt;strong&gt;managed-readiness → KV-event gate → functional-probe gate&lt;/strong&gt;), so the &lt;code&gt;READY&lt;/code&gt; column of
&lt;code&gt;kubectl get cachebackend&lt;/code&gt; only tells half the story. The reason lives in
&lt;code&gt;.status.conditions[]&lt;/code&gt;, and the condition &lt;code&gt;.reason&lt;/code&gt; is the actionable string:&lt;/p&gt;</description></item><item><title>CacheIndex</title><link>https://cachebox-project.github.io/inference-cache/docs/concepts/cacheindex/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/concepts/cacheindex/</guid><description>&lt;h2 id="what-is-a-cacheindex"&gt;
What is a CacheIndex?
&lt;a href="#what-is-a-cacheindex" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;&lt;code&gt;CacheIndex&lt;/code&gt; is a cluster-scoped, &lt;strong&gt;status-only&lt;/strong&gt; singleton named &lt;code&gt;cluster-default&lt;/code&gt;. It
mirrors the server&amp;rsquo;s in-memory cache-state aggregate into the Kubernetes API so you can
observe cluster-wide cache state with &lt;code&gt;kubectl get cacheindex&lt;/code&gt;. It is for &lt;strong&gt;observability&lt;/strong&gt;,
not routing — the live routing substrate is the server&amp;rsquo;s in-memory index, and &lt;code&gt;LookupRoute&lt;/code&gt;
never reads the CR.&lt;/p&gt;</description></item><item><title>CLI: inferencecache doctor</title><link>https://cachebox-project.github.io/inference-cache/docs/reference/cli-doctor/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/reference/cli-doctor/</guid><description>&lt;p&gt;&lt;code&gt;inferencecache doctor&lt;/code&gt; is a &lt;strong&gt;read-only&lt;/strong&gt; pre-flight diagnostic — the cache-plane analogue
of &lt;code&gt;istioctl analyze&lt;/code&gt;. It runs against a live install, prints stable finding codes, and
returns CI-friendly exit codes. The binary is &lt;code&gt;bin/inferencecache&lt;/code&gt; (from &lt;code&gt;make build&lt;/code&gt;).&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;inferencecache doctor -n serving
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;

&lt;h2 id="checks-and-finding-codes"&gt;
Checks and finding codes
&lt;a href="#checks-and-finding-codes" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;After the Kubernetes API prerequisite (check 0), nine diagnostic checks run in a fixed
order. Finding codes are stable and greppable; severity is shown in parentheses.&lt;/p&gt;</description></item><item><title>Monitor the cache plane</title><link>https://cachebox-project.github.io/inference-cache/docs/tasks/monitor-the-cache-plane/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/tasks/monitor-the-cache-plane/</guid><description>&lt;p&gt;There are three complementary ways to watch a running install: the Kubernetes status
surfaces, the Prometheus metrics, and the &lt;code&gt;inferencecache doctor&lt;/code&gt; pre-flight diagnostic.&lt;/p&gt;


&lt;h2 id="kubectl-surfaces"&gt;
kubectl surfaces
&lt;a href="#kubectl-surfaces" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#8f5902;font-style:italic"&gt;# Per-backend health, matched pods, and live prefix/event evidence&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;kubectl get cachebackend -A
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#8f5902;font-style:italic"&gt;# The cluster-wide cache &amp;#34;world map&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;kubectl get cacheindex cluster-default -o yaml
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#8f5902;font-style:italic"&gt;# Per-tenant entry counts&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;kubectl get cachetenant -A
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;When a backend is not &lt;code&gt;Ready&lt;/code&gt;, the reason is in &lt;code&gt;.status.conditions[]&lt;/code&gt;:&lt;/p&gt;</description></item><item><title>PromptTemplate</title><link>https://cachebox-project.github.io/inference-cache/docs/concepts/prompttemplate/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/concepts/prompttemplate/</guid><description>&lt;h2 id="what-is-a-prompttemplate"&gt;
What is a PromptTemplate?
&lt;a href="#what-is-a-prompttemplate" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;A &lt;code&gt;PromptTemplate&lt;/code&gt; declares a prompt template and marks which of its slots are &lt;strong&gt;stable&lt;/strong&gt;
(part of the cacheable prefix) versus &lt;strong&gt;mutable&lt;/strong&gt; (vary per request). This backs the
mutable-slot render pipeline — the &amp;ldquo;wedge&amp;rdquo; that lets a mostly-fixed prompt keep a stable,
cacheable prefix even as small parts of it change.&lt;/p&gt;</description></item><item><title>PDTopology</title><link>https://cachebox-project.github.io/inference-cache/docs/concepts/pdtopology/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/concepts/pdtopology/</guid><description>&lt;h2 id="what-is-a-pdtopology"&gt;
What is a PDTopology?
&lt;a href="#what-is-a-pdtopology" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;A &lt;code&gt;PDTopology&lt;/code&gt; declares the prefill and decode pools for &lt;strong&gt;phase-disaggregated (P/D)
serving&lt;/strong&gt; — a topology where the prefill phase (prompt processing) and the decode phase
(token generation) run on separate pools of replicas, so the KV cache produced by prefill is
transferred to a decode replica.&lt;/p&gt;</description></item><item><title>LookupRoute &amp; ranking</title><link>https://cachebox-project.github.io/inference-cache/docs/concepts/lookuproute/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/concepts/lookuproute/</guid><description>&lt;h2 id="the-routing-hint"&gt;
The routing hint
&lt;a href="#the-routing-hint" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;&lt;code&gt;LookupRoute&lt;/code&gt; is the hot-path RPC a gateway calls for every request. Given a prompt-prefix
key plus &lt;code&gt;model_id&lt;/code&gt; and &lt;code&gt;tenant_id&lt;/code&gt;, it returns the replicas that hold that prefix warm,
ranked best-first, plus a &lt;code&gt;reason_code&lt;/code&gt;. The gateway routes to the top replica for a prefix
cache hit — reusing KV / skipping prefill, lowering TTFT and cost — or round-robins when
there is no hint.&lt;/p&gt;</description></item><item><title>The gRPC contract</title><link>https://cachebox-project.github.io/inference-cache/docs/concepts/grpc-contract/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/concepts/grpc-contract/</guid><description>&lt;h2 id="the-inferencecache-service"&gt;
The InferenceCache service
&lt;a href="#the-inferencecache-service" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The &lt;code&gt;InferenceCache&lt;/code&gt; gRPC service is the data-plane API that gateways and engines integrate
against. It is defined in &lt;code&gt;proto/inferencecache/v1alpha1/inferencecache.proto&lt;/code&gt;, package
&lt;code&gt;inferencecache.v1alpha1&lt;/code&gt;, and served on &lt;code&gt;:9090&lt;/code&gt; (plaintext by default; TLS is an opt-in
overlay — see &lt;a href="https://cachebox-project.github.io/inference-cache/docs/administration/grpc-tls/"&gt;gRPC TLS&lt;/a&gt;).&lt;/p&gt;


&lt;h2 id="eight-rpcs-three-roles"&gt;
Eight RPCs, three roles
&lt;a href="#eight-rpcs-three-roles" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Consumer side (gateways) — fail-open; only delivered LFU hits update eviction
bookkeeping:&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Search Results</title><link>https://cachebox-project.github.io/inference-cache/search/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/search/</guid><description/></item></channel></rss>