CLI: inferencecache doctor

Checks, finding codes, flags, and exit codes for the pre-flight diagnostic.

inferencecache doctor is a read-only pre-flight diagnostic — the cache-plane analogue of istioctl analyze. It runs against a live install, prints stable finding codes, and returns CI-friendly exit codes. The binary is bin/inferencecache (from make build).

inferencecache doctor -n serving

Checks and finding codes

After the Kubernetes API prerequisite (check 0), nine diagnostic checks run in a fixed order. Finding codes are stable and greppable; severity is shown in parentheses.

#CheckCodes
0Kubernetes API reachableAPI001 (FAIL)
1Server gRPC health SERVINGSV001/SV002 (FAIL), SV003 (OK)
2/snapshot reachableSN001 (FAIL), SN002/SN005 (WARN), SN003 (INFO), SN004 (OK)
3/policy wiredPL001 (FAIL), PL002 (OK), PL003 (WARN)
4/probe wiredPB001 (FAIL), PB002 (OK), PB003 (WARN)
5Per-CacheBackend healthCB001–CB005 (WARN), CB006 (OK), CB007 (WARN — FunctionalProbeOK not True)
6Engine-pod injection auditEP001/EP003 (WARN), EP002 (OK)
7Orphan-pod checkOP001 (WARN)
8CacheTenant healthCT001 (WARN), CT002 (OK)
9CachePolicy coverageCP001 (INFO), CP002 (OK), CP003 (WARN — >1 CachePolicy in a namespace)

CB003 keys off KV-event observation (both the first-event latch and lastEventAt unset). External backends have only their Ready state and endpoint checked.

Exit codes

CodeMeaning
0Highest severity ≤ INFO.
1At least one WARN.
2At least one FAIL.

Flags

FlagPurpose
--kubeconfigPath to the kubeconfig.
--contextKube context to use.
-n, --namespaceNamespace to scope resource checks to.
--server-endpointServer host[:gRPCport] (HTTP endpoints derived on :8081).
--snapshot-token-fileBearer token file for the /snapshot check.
-o, --outputhuman (default), json, or table.
--no-colorDisable ANSI color.
--config-onlySkip live probes — the right mode from a workstation, and under the TLS overlay.
--timeoutOverall timeout (default 30s).

Network posture: :9090 (gRPC) is open to all in-cluster clients; the :8081 bridge is NetworkPolicy-restricted to the controller pods, so a full run is best done from within the cluster or with an appropriate token.

Last modified August 13, 2026: LMCache MP mode (#187) (40a9806)