<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Administration on inference-cache</title><link>https://cachebox-project.github.io/inference-cache/docs/administration/</link><description>Recent content in Administration on inference-cache</description><generator>Hugo</generator><language>en</language><atom:link href="https://cachebox-project.github.io/inference-cache/docs/administration/index.xml" rel="self" type="application/rss+xml"/><item><title>Observability &amp; Alerts</title><link>https://cachebox-project.github.io/inference-cache/docs/administration/observability-and-alerts/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/administration/observability-and-alerts/</guid><description>&lt;h2 id="the-bundle"&gt;
The bundle
&lt;a href="#the-bundle" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;A Prometheus alert bundle for the operational silent-failure patterns this system has hit in
production ships under &lt;code&gt;config/observability/&lt;/code&gt;. It is &lt;strong&gt;not&lt;/strong&gt; part of &lt;code&gt;config/default&lt;/code&gt; — the
alerts are opt-in so that installs without prometheus-operator CRDs are not affected by an
unknown &lt;code&gt;apiVersion&lt;/code&gt;.&lt;/p&gt;</description></item><item><title>Index sizing</title><link>https://cachebox-project.github.io/inference-cache/docs/administration/index-sizing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/administration/index-sizing/</guid><description>&lt;p&gt;The server&amp;rsquo;s cache-state index lives in memory. This page helps you size the server pod and
choose eviction settings so the index never approaches the pod&amp;rsquo;s memory limit (an OOM-kill
is the one failure the soft-state design cannot hide).&lt;/p&gt;


&lt;h2 id="pod-budget-at-a-glance"&gt;
Pod budget at a glance
&lt;a href="#pod-budget-at-a-glance" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Storage entries (distinct prefix keys × replicas reporting each)&lt;/th&gt;
					&lt;th&gt;Peak RSS&lt;/th&gt;
					&lt;th&gt;Recommended pod memory&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;100K storage entries&lt;/td&gt;
					&lt;td&gt;~110 MiB&lt;/td&gt;
					&lt;td&gt;256 MiB&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;500K storage entries&lt;/td&gt;
					&lt;td&gt;~300 MiB&lt;/td&gt;
					&lt;td&gt;512 MiB&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;1M storage entries (the default cap)&lt;/td&gt;
					&lt;td&gt;~540 MiB&lt;/td&gt;
					&lt;td&gt;&lt;strong&gt;1 GiB (recommended floor)&lt;/strong&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;1.5M storage entries&lt;/td&gt;
					&lt;td&gt;~700 MiB&lt;/td&gt;
					&lt;td&gt;1.5 GiB&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;These include roughly 20% headroom over measured heap. The default entry cap is 1,000,000.&lt;/p&gt;</description></item><item><title>gRPC TLS</title><link>https://cachebox-project.github.io/inference-cache/docs/administration/grpc-tls/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/administration/grpc-tls/</guid><description>&lt;h2 id="the-decision"&gt;
The decision
&lt;a href="#the-decision" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;The gRPC API on &lt;code&gt;:9090&lt;/code&gt; uses &lt;strong&gt;one-sided Service TLS via cert-manager, terminated
in-process&lt;/strong&gt; — no sidecar, Envoy, or Ingress. Mutual TLS (client-certificate auth) is a
Phase-2 feature flag, not yet implemented.&lt;/p&gt;
&lt;p&gt;The server binary supports TLS via &lt;code&gt;--tls-cert-file&lt;/code&gt; / &lt;code&gt;--tls-key-file&lt;/code&gt; (both-or-neither —
supplying exactly one fails startup). The certificate is served through a reloading
&lt;code&gt;GetCertificate&lt;/code&gt; hook, so a cert-manager-rotated Secret is picked up &lt;strong&gt;without a restart&lt;/strong&gt;.
The posture is exposed as the gauge &lt;code&gt;inferencecache_server_grpc_tls_enabled&lt;/code&gt; (0/1).&lt;/p&gt;</description></item><item><title>Troubleshooting</title><link>https://cachebox-project.github.io/inference-cache/docs/administration/troubleshooting/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://cachebox-project.github.io/inference-cache/docs/administration/troubleshooting/</guid><description>&lt;h2 id="reading-cachebackend-readiness"&gt;
Reading CacheBackend readiness
&lt;a href="#reading-cachebackend-readiness" class="anchor-link"&gt;
 

&lt;svg xmlns="http://www.w3.org/2000/svg" fill="currentColor" width="24" height="24" viewBox="0 0 24 24"&gt;
 &lt;path d="M0 0h24v24H0z" fill="none"&gt;&lt;/path&gt;
 &lt;path d="M3.9 12c0-1.71 1.39-3.1 3.1-3.1h4V7H7c-2.76 0-5 2.24-5 5s2.24 5 5 5h4v-1.9H7c-1.71 0-3.1-1.39-3.1-3.1zM8 13h8v-2H8v2zm9-6h-4v1.9h4c1.71 0 3.1 1.39 3.1 3.1s-1.39 3.1-3.1 3.1h-4V17h4c2.76 0 5-2.24 5-5s-2.24-5-5-5z"&gt;&lt;/path&gt;
&lt;/svg&gt;
&lt;/a&gt;
&lt;/h2&gt;
&lt;p&gt;A managed backend&amp;rsquo;s &lt;code&gt;Ready&lt;/code&gt; is the composition of three gates
(&lt;strong&gt;managed-readiness → KV-event gate → functional-probe gate&lt;/strong&gt;), so the &lt;code&gt;READY&lt;/code&gt; column of
&lt;code&gt;kubectl get cachebackend&lt;/code&gt; only tells half the story. The reason lives in
&lt;code&gt;.status.conditions[]&lt;/code&gt;, and the condition &lt;code&gt;.reason&lt;/code&gt; is the actionable string:&lt;/p&gt;</description></item></channel></rss>