<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Concepts on inference-cache</title><link>https://cachebox-project.github.io/inference-cache/docs/concepts/</link><description>Recent content in Concepts on inference-cache</description><generator>Hugo</generator><language>en</language><atom:link href="https://cachebox-project.github.io/inference-cache/docs/concepts/index.xml" rel="self" type="application/rss+xml"/><item><title>Architecture</title><link>https://cachebox-project.github.io/inference-cache/docs/concepts/architecture/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/concepts/architecture/</guid><description>&lt;h2 id="one-operator-two-binaries"&gt;
One operator, two binaries
&lt;a href="#one-operator-two-binaries" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;inference-cache is a single operator split across two control-plane binaries, plus a CLI
and an in-cluster sidecar. The split follows the responsibilities: the &lt;strong&gt;controller&lt;/strong&gt; owns
the Kubernetes reconcile loop and admission; the &lt;strong&gt;server&lt;/strong&gt; owns the routing decision and
the cache-state index.&lt;/p&gt;</description></item><item><title>CacheBackend</title><link>https://cachebox-project.github.io/inference-cache/docs/concepts/cachebackend/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/concepts/cachebackend/</guid><description>&lt;h2 id="what-is-a-cachebackend"&gt;
What is a CacheBackend?
&lt;a href="#what-is-a-cachebackend" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;A namespaced &lt;code&gt;CacheBackend&lt;/code&gt; selects an inference runtime, its engine-side cache
integration, and an optional remote L3. The inference system owns the engine
Deployment and image. CacheBackend injects the selected connector and the cache
components it owns into matching Pods.&lt;/p&gt;</description></item><item><title>CachePolicy</title><link>https://cachebox-project.github.io/inference-cache/docs/concepts/cachepolicy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/concepts/cachepolicy/</guid><description>&lt;h2 id="what-is-a-cachepolicy"&gt;
What is a CachePolicy?
&lt;a href="#what-is-a-cachepolicy" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;A &lt;code&gt;CachePolicy&lt;/code&gt; tunes how the server filters and ranks lookups, and how it evicts, for a
single namespace. It is &lt;strong&gt;entirely opt-in&lt;/strong&gt;: the server ships sane defaults, so most
namespaces need no &lt;code&gt;CachePolicy&lt;/code&gt; at all. It is &lt;strong&gt;purely declarative&lt;/strong&gt; — the controller
flattens all policies and pushes the resolved values to the server; the reconciler never
writes &lt;code&gt;CachePolicy.status&lt;/code&gt;.&lt;/p&gt;</description></item><item><title>CacheTenant</title><link>https://cachebox-project.github.io/inference-cache/docs/concepts/cachetenant/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/concepts/cachetenant/</guid><description>&lt;h2 id="what-is-a-cachetenant"&gt;
What is a CacheTenant?
&lt;a href="#what-is-a-cachetenant" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;A &lt;code&gt;CacheTenant&lt;/code&gt; gives an external tenant two things:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;A stable &lt;strong&gt;identity&lt;/strong&gt; (&lt;code&gt;spec.tenantID&lt;/code&gt;) that the cache-state index isolates on, so one
tenant&amp;rsquo;s cache hints can never collide with another&amp;rsquo;s.&lt;/li&gt;
&lt;li&gt;An optional &lt;strong&gt;entry-count quota&lt;/strong&gt; that caps how many distinct prefix keys the tenant may
occupy in the index.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;code&gt;CacheTenant&lt;/code&gt; is namespaced, short name &lt;code&gt;ct&lt;/code&gt;.&lt;/p&gt;</description></item><item><title>CacheIndex</title><link>https://cachebox-project.github.io/inference-cache/docs/concepts/cacheindex/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/concepts/cacheindex/</guid><description>&lt;h2 id="what-is-a-cacheindex"&gt;
What is a CacheIndex?
&lt;a href="#what-is-a-cacheindex" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;&lt;code&gt;CacheIndex&lt;/code&gt; is a cluster-scoped, &lt;strong&gt;status-only&lt;/strong&gt; singleton named &lt;code&gt;cluster-default&lt;/code&gt;. It
mirrors the server&amp;rsquo;s in-memory cache-state aggregate into the Kubernetes API so you can
observe cluster-wide cache state with &lt;code&gt;kubectl get cacheindex&lt;/code&gt;. It is for &lt;strong&gt;observability&lt;/strong&gt;,
not routing — the live routing substrate is the server&amp;rsquo;s in-memory index, and &lt;code&gt;LookupRoute&lt;/code&gt;
never reads the CR.&lt;/p&gt;</description></item><item><title>PromptTemplate</title><link>https://cachebox-project.github.io/inference-cache/docs/concepts/prompttemplate/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/concepts/prompttemplate/</guid><description>&lt;h2 id="what-is-a-prompttemplate"&gt;
What is a PromptTemplate?
&lt;a href="#what-is-a-prompttemplate" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;A &lt;code&gt;PromptTemplate&lt;/code&gt; declares a prompt template and marks which of its slots are &lt;strong&gt;stable&lt;/strong&gt;
(part of the cacheable prefix) versus &lt;strong&gt;mutable&lt;/strong&gt; (vary per request). This backs the
mutable-slot render pipeline — the &amp;ldquo;wedge&amp;rdquo; that lets a mostly-fixed prompt keep a stable,
cacheable prefix even as small parts of it change.&lt;/p&gt;</description></item><item><title>PDTopology</title><link>https://cachebox-project.github.io/inference-cache/docs/concepts/pdtopology/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/concepts/pdtopology/</guid><description>&lt;h2 id="what-is-a-pdtopology"&gt;
What is a PDTopology?
&lt;a href="#what-is-a-pdtopology" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;A &lt;code&gt;PDTopology&lt;/code&gt; declares the prefill and decode pools for &lt;strong&gt;phase-disaggregated (P/D)
serving&lt;/strong&gt; — a topology where the prefill phase (prompt processing) and the decode phase
(token generation) run on separate pools of replicas, so the KV cache produced by prefill is
transferred to a decode replica.&lt;/p&gt;</description></item><item><title>LookupRoute &amp; ranking</title><link>https://cachebox-project.github.io/inference-cache/docs/concepts/lookuproute/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/concepts/lookuproute/</guid><description>&lt;h2 id="the-routing-hint"&gt;
The routing hint
&lt;a href="#the-routing-hint" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;&lt;code&gt;LookupRoute&lt;/code&gt; is the hot-path RPC a gateway calls for every request. Given a prompt-prefix
key plus &lt;code&gt;model_id&lt;/code&gt; and &lt;code&gt;tenant_id&lt;/code&gt;, it returns the replicas that hold that prefix warm,
ranked best-first, plus a &lt;code&gt;reason_code&lt;/code&gt;. The gateway routes to the top replica for a prefix
cache hit — reusing KV / skipping prefill, lowering TTFT and cost — or round-robins when
there is no hint.&lt;/p&gt;</description></item><item><title>The gRPC contract</title><link>https://cachebox-project.github.io/inference-cache/docs/concepts/grpc-contract/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/concepts/grpc-contract/</guid><description>&lt;h2 id="the-inferencecache-service"&gt;
The InferenceCache service
&lt;a href="#the-inferencecache-service" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The &lt;code&gt;InferenceCache&lt;/code&gt; gRPC service is the data-plane API that gateways and engines integrate
against. It is defined in &lt;code&gt;proto/inferencecache/v1alpha1/inferencecache.proto&lt;/code&gt;, package
&lt;code&gt;inferencecache.v1alpha1&lt;/code&gt;, and served on &lt;code&gt;:9090&lt;/code&gt; (plaintext by default; TLS is an opt-in
overlay — see &lt;a href="https://cachebox-project.github.io/inference-cache/docs/administration/grpc-tls/"&gt;gRPC TLS&lt;/a&gt;).&lt;/p&gt;


&lt;h2 id="eight-rpcs-three-roles"&gt;
Eight RPCs, three roles
&lt;a href="#eight-rpcs-three-roles" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Consumer side (gateways) — fail-open; only delivered LFU hits update eviction
bookkeeping:&lt;/strong&gt;&lt;/p&gt;</description></item></channel></rss>