CRD API

The six custom resources — groups, scopes, short names, and key fields.

All CRDs are in the API group inferencecache.io, version v1alpha1.

KindScopeShort nameReconciled?Purpose
CacheBackendNamespacedcbYesBind engine Pods to typed MP; optionally provision a remote Redis provider.
CachePolicyNamespacedcpolDeclarative (pushed to server)Per-namespace lookup and eviction tuning.
CacheTenantNamespacedctDeclarative (pushed to server)Tenant identity + entry-count quota.
CacheIndexClusterciYes (status-only)Cluster-wide mirror of the server aggregate.
PromptTemplateNamespacedptDeclarativeCache-aware prompt template + stable/mutable slots.
PDTopologyNamespacedpdtDeclarativePrefill/decode topology for disaggregated serving.

CacheBackend

Key spec fields:

FieldType / valuesDefaultNotes
runtimeVLLM, SGLang—Inference runtime identity.
typeLMCache, SGLangHiCacheLMCacheEngine-side cache implementation.
lmCacheobject—Typed LMCache MP topology and server configuration. Current offload uses topology: PodLocal.
lmCache.podLocal.server.resourcesResourceRequirementsrequiredResources for the injected MP server; memory covers L1 plus 1Gi and CPU request is positive.
remoteStorageobject—Optional Redis L3 with Managed or External ownership.
remoteStorage.workloadobject—Scheduling and Pod security for a managed provider; rejected for External and has no generic replica count.
observationobject—Model identity and first-event timeout.
integration.modeOffload, EventsOnlyOffloadEvents-only = routing only.
integration.roleReadOnly, WriteOnly, ReadWriteReadWriteLMCache currently admits only ReadWrite; directional semantics are future work.
integration.failOpenbooltruefalse fails closed.
integration.engineOverridesobject—args / suppressArgs / env / suppressEnv.
engineSelector.matchLabelsmap—Equality selector over engine pod labels.
remoteStorage.<provider>.resourcesResourceRequirementsrenderer default: requests.memory 4Gi / limits.memory 8GiResources for the selected managed provider container.
allowCrossNamespaceboolfalseOpt-in cross-namespace endpoints.

Key status fields: connector, remoteStorage, matchedEnginePods (*int32), firstKVEventObservedAt (*Time), indexParticipation (prefixCount, lastEventAt, hitRate *string, t2HitRate *string), failOpen, observedGeneration, conditions.

Full page: CacheBackend.

CachePolicy

FieldType / valuesDefault
evictionLRU, LFULRU
evictionTTLdurationserver 30m (must be > 0 when set)
minimumPrefixTokensint32unset (no gate)
minimumMatchedTokensint3264 (0 opts out)
routingFloorScorestring (float)"0.1" ("0" opts out)
lookupTimeoutMsint32unset (0/≤0 = unbounded)
strategy.enableChainMatchingbooltrue
strategy.requireChainboolfalse
strategy.enableTenantHotbooltrue
affinityRoutingEnabled, DisabledEnabled

status is reserved (not written today). Full page: CachePolicy.

CacheTenant

FieldType / valuesDefault
tenantIDstring (required, min length 1)—
quota.maxIndexEntriesint64unset (unbounded)
isolationModeFairnessFairness
cryptoobjectreserved (empty)

status: indexEntries (*int64), conditions, observedGeneration. There is deliberately no maxMemoryBytes / memoryUsed. Full page: CacheTenant.

CacheIndex

spec is empty. status: replicas[] (id, tenant, cacheMemoryBytes, hitRate *string, pressure, lastUpdate), tenants[] (id, indexEntries *int64, hitRate *string, memoryUsed — deprecated, always 0), prefixes.summary (total, hot=0), observedServer, lastUpdated. Full page: CacheIndex.

PromptTemplate

spec: body (required), slots[] (name, type = Stable|Mutable, required, description). status: templateRevision, conditions, observedGeneration. Full page: PromptTemplate.

PDTopology

spec: prefillPools[], decodePools[] (name, matchLabels, replicas, acceleratorType), acceleratorTypes[] (name, vendor, model, matchLabels). status: conditions, observedGeneration. Full page: PDTopology.

Last modified August 13, 2026: LMCache MP mode (#187) (40a9806)