The gRPC contract
The InferenceCache service
The InferenceCache gRPC service is the data-plane API that gateways and engines integrate
against. It is defined in proto/inferencecache/v1alpha1/inferencecache.proto, package
inferencecache.v1alpha1, and served on :9090 (plaintext by default; TLS is an opt-in
overlay — see gRPC TLS).
Eight RPCs, three roles
Consumer side (gateways) — fail-open; only delivered LFU hits update eviction bookkeeping:
| RPC | Kind | Purpose |
|---|---|---|
LookupRoute | unary | The core cache-aware routing hint. See LookupRoute & ranking. |
RenderTemplate | unary | Deterministic mutable-slot prompt render. Fail-open stub today. |
LookupPDRoute | unary | Prefill/decode split routing. Phase-2 stub. |
GetCacheState | unary | The (tenant, model) aggregate (replica stats + summary). |
Producer side (engine adapters):
| RPC | Kind | Purpose |
|---|---|---|
ReportCacheState | client-stream | Authoritative additive ingest of prefix/replica state. |
PublishEvent | unary | Scheme-safe deltas (PREFIX_EVICTED, REPLICA_UPDATED, ALL_CLEARED). |
Observer side (dashboards / debug) — server-stream stubs:
| RPC | Kind | Purpose |
|---|---|---|
StreamCacheEvents | server-stream | Live cache events. |
StreamMetrics | server-stream | Live metrics. |
Key messages
message LookupRouteRequest {
string model_id = 1;
string tenant_id = 2;
bytes prefix_hash = 3;
int32 prefix_token_count = 4;
string hash_scheme = 5;
SLO slo = 6;
repeated bytes block_hashes = 7; // ordered block-hash chain
repeated int32 block_token_counts = 8;
repeated uint32 token_ids = 9; // dual-input: server fingerprints
string prompt_text = 10; // dual-input: server tokenizes + fingerprints
string adapter_id = 11; // LoRA index partition
}
message LookupRouteResponse {
repeated ReplicaScore replica_scores = 1;
string reason_code = 2;
int64 lookup_latency_us = 3;
repeated uint32 token_ids = 4; // set only when prompt_text was tokenized
string adapter_id = 5; // partition consulted
}
enum CacheTier {
CACHE_TIER_UNSPECIFIED = 0;
CACHE_TIER_T1 = 1;
CACHE_TIER_T2 = 2;
CACHE_TIER_T3 = 3;
}
message ReplicaScore {
string replica_id = 1;
float score = 2;
int32 matched_tokens = 3;
float estimated_cache_hit_prob = 4;
CacheTier tier = 5;
}
message PrefixEntry { // metadata only — never KV tensors or prompt text
bytes prefix_hash = 1;
int32 token_count = 2;
repeated bytes block_hashes = 3;
repeated int32 block_token_counts = 4;
CacheTier tier = 5;
string adapter_id = 6;
}
message ReplicaStats {
string replica_id = 1;
int64 cache_memory_bytes = 2; // engine total across all tenants
float hit_rate = 3;
float pressure = 4;
string client_version = 5;
int64 t2_hit_tokens = 6; // from vLLM external_prefix_cache_hits_total
int64 t2_query_tokens = 7; // from vLLM external_prefix_cache_queries_total
}
Ack { bool accepted; string reason_code; } and SLO { int32 ttft_ms; int32 tbt_ms; }
round out the core shapes.
The guarantees
These are the contract properties clients are allowed to depend on:
reason_codeis astring, not a proto enum. New codes can be added without breaking old clients — a client that does not recognize a code must treat it as the no-hint default.- Fail-open. An empty
replica_scoreslist is always valid.TIMEOUTis a fail-open outcome, not an error. The hot path never returns a gRPC error for a cache miss. - No cache-membership mutation on lookups. A lookup does not add, remove, or rewrite
indexed prefixes. Metrics are emitted, and an
LFUnamespace credits internal eviction access counters on a delivered prefix-match hit. - Engine-opaque hashes, scheme-scoped.
prefix_hash/block_hashesare matched only within a matchinghash_scheme; an empty scheme is not a valid domain. - Metadata-only index updates.
CacheStateUpdateandPrefixEntrycarry hashes and statistics — never KV tensors or prompt text.LookupRoutemay carry optionalprompt_textinput, which strengthens the case for enabling TLS on:9090. CacheStateUpdateis an additive delta, not a snapshot. A prefix’s absence from a later update does not remove it. Removals come only viaCacheEvent(PREFIX_EVICTED/ALL_CLEARED) or TTL. This matches the engine’s own KV-event model, and it is why a stale entry degrades to a cache miss, never a wrong answer.- Reserved tenant.
tenant_id = "inferencecache.io/probe"is reserved for the server’s functional self-test — excluded from/snapshotand from public reads and writes.
The producer path
Engine adapters feed the index through the observation sidecar:
ReportCacheStateis the authoritative, idempotent ingest keyed on(replica, hash_scheme, adapter_id, prefix_hash). It carriesPrefixEntrymetadata plusReplicaStats; an entry-leveladapter_idfalls back to the update-level value.PublishEventcarries scheme-safe deltas.PREFIX_ADDEDis a no-op (events carry nohash_scheme, andReportCacheStateis authoritative for adds);REPLICA_UPDATEDrefreshes replica stats;PREFIX_EVICTED/ALL_CLEAREDremove.
Keeping the contract honest
Because there are external consumers of this contract (gateway clients in other languages),
the proto stays backward-compatible — fields are not renumbered or removed. Any change to
proto/ is required to update the contract design doc in the same change, and generated code
is verified drift-free in CI.
Related pages
- Reason codes — the full string-code vocabulary.
- gRPC API reference — RPC-by-RPC detail.
- LookupRoute & ranking — how the hint is computed.