Index sizing

How the in-memory index consumes memory, the pod-budget table, and the levers that keep it in bounds.

The server’s cache-state index lives in memory. This page helps you size the server pod and choose eviction settings so the index never approaches the pod’s memory limit (an OOM-kill is the one failure the soft-state design cannot hide).

Pod budget at a glance

Storage entries (distinct prefix keys × replicas reporting each)Peak RSSRecommended pod memory
100K storage entries~110 MiB256 MiB
500K storage entries~300 MiB512 MiB
1M storage entries (the default cap)~540 MiB1 GiB (recommended floor)
1.5M storage entries~700 MiB1.5 GiB

These include roughly 20% headroom over measured heap. The default entry cap is 1,000,000.

Where the memory goes

Rough per-entry footprint:

  • ~500 bytes per distinct prefix key,
  • ~50 bytes per additional replica holding that key,
  • ~50 MiB fixed Go-runtime baseline.

So a first-order estimate is:

heap ≈ 50 MiB + (500 + 50 × (replicas − 1)) × distinct_keys

The prefix-hash byte width does not dominate — the Go map machinery and the per-entry time.Time timestamps do.

The block-chain multiplier is the trap

The most-overlooked multiplier is block-level expansion. A prompt is not one entry — it is one entry per block in its prefix chain. A 1000-token prompt at a 16-token block size is ~63 entries per replica. So:

distinct_keys ≈ distinct_prompts × block_chain_length
storage_entries ≈ distinct_keys × replicas

A workload with 5,000 distinct prompts and ~60-block chains has about 300K distinct keys. Across 3 replicas that is about 900K storage entries — near the default cap. Size for the block-expanded storage-entry count, not the prompt count.

The levers

Runtime-tunable (no rebuild):

LeverWhereEffect
Pod memory limitDeployment/inference-cache-server, container serverThe only hard ceiling. Update the installed manifest or patch this Deployment’s resources.
evictionTTLCachePolicy (default 30m)Shorter TTL ⇒ smaller index. The cheapest lever, roughly linear.
evictionCachePolicy (LRU/LFU)Which entries go first under the cap.
maxIndexEntriesCacheTenant.spec.quotaPer-tenant distinct-key cap; over-budget evicts oldest (Fairness).

Compile-time constants (need a rebuild to change): DefaultMaxEntries = 1,000,000, DefaultTTL = 30m, DefaultSweepInterval = 1m.

When you’re over the cap

Three choices, cheapest first:

  1. Accept the cap — cap eviction is normal; a stale/evicted hint is just a cache miss.
  2. Shorten evictionTTL — reduces the working set roughly linearly.
  3. Raise the memory limit (and, if needed, rebuild with a higher DefaultMaxEntries).

evictionTTL: 0 or negative is rejected at admission.

Signals to watch

SignalMetricRead as
Index trendinferencecache_index_entries{model}Growth trend (not cap-closeness on multi-replica).
Cap pressurerate(inferencecache_index_evictions_total{reason="cap"}[10m])The authoritative “over budget” signal. Alert IndexEvictionsSpike at >10/sec.
Quota pressurerate(inferencecache_tenant_evictions_total[10m])A bounded non-zero rate is normal Fairness.
TTL churninferencecache_index_evictions_total{reason="ttl"}Healthy background aging.
OOM proximitycontainer_memory_working_set_bytesThe kubelet’s OOM signal — keep well under the limit.