CVE-2026-89032 Overview
CVE-2026-89032 is a tenant isolation bypass vulnerability in BerriAI LiteLLM versions prior to 1.101.0-rc.1. The flaw resides in the semantic cache layer, where a metadata key mismatch between _get_semantic_cache_tenant_scope() and _get_metadata_variable_name() allows cached responses to leak across tenants. Authenticated users holding a valid virtual key can submit semantically similar prompts to affected routes such as /v1/responses and /bedrock/* to retrieve other tenants' cached data. The weakness is classified as [CWE-863: Incorrect Authorization].
Critical Impact
Attackers can read cached responses containing other tenants' personally identifiable information, financial data, and source code, and can trigger agentic front-ends to auto-execute attacker-supplied function_call or tool_calls payloads under victim credentials.
Affected Products
- BerriAI LiteLLM versions prior to 1.101.0-rc.1
- Deployments using the semantic cache layer
- Routes including /v1/responses and /bedrock/*
Discovery Timeline
- 2026-09-25 - CVE-2026-89032 published to NVD
- 2026-10-01 - Last updated in NVD database
Technical Details for CVE-2026-89032
Vulnerability Analysis
LiteLLM is a Python proxy that unifies calls to multiple large language model (LLM) providers. The semantic cache layer stores LLM responses keyed by embedding similarity so later prompts with equivalent meaning can be served from cache. In vulnerable builds, the function that scopes cached entries to a tenant uses a different metadata key than the function that retrieves them. The resulting mismatch means a cache lookup for one principal can match an entry stored by another principal.
The consequences extend beyond passive disclosure. If a cached response contains a function_call or tool_calls payload, an agentic front-end consuming the result will execute those tool invocations under the credentials of the requesting user, not the user who originally generated the cache entry.
Root Cause
The root cause is inconsistent tenant scoping within litellm/caching/caching.py. _get_semantic_cache_tenant_scope() and _get_metadata_variable_name() reference different metadata keys when computing the cache namespace. Entries are therefore written under one scope and read back under another, breaking the isolation guarantee between virtual keys and end users.
Attack Vector
The attack is network-reachable and requires a valid virtual key. An attacker crafts prompts that are semantically close to a target request issued by another tenant. On routes such as /v1/responses or /bedrock/*, the embedding similarity match returns the victim's cached completion. When the cached payload includes tool call directives, downstream agent frameworks execute them with the attacker-selected user's credentials.
# Patch - litellm/types/caching.py
# Introduces an explicit scope enum for semantic cache isolation
GCS = "gcs"
class SemanticCacheScope(str, Enum):
KEY = "key"
END_USER = "end_user"
CachingSupportedCallTypes = Literal[
"completion",
"acompletion",
# Patch - litellm/caching/caching.py
# Adds semantic_cache_scope parameter to the cache initializer
qdrant_semantic_cache_vector_size: int | None = None,
semantic_cache_embedding_max_input_tokens: int | None = None,
semantic_cache_embedding_timeout: float | None = None,
semantic_cache_scope: str = SemanticCacheScope.KEY.value,
# GCP IAM authentication parameters
gcp_service_account: str | None = None,
gcp_ssl_ca_certs: str | None = None,
Source: BerriAI LiteLLM commit 16db51e
Detection Methods for CVE-2026-89032
Indicators of Compromise
- Semantic cache hits where the responding end_user or virtual key does not match the originator recorded in the cache entry metadata.
- Unexpected function_call or tool_calls payloads returned to a user who did not request such tooling.
- Spikes in requests to /v1/responses or /bedrock/* from a single virtual key with varied but semantically clustered prompts.
Detection Strategies
- Audit LiteLLM proxy logs for cache-hit events crossing tenant or end-user boundaries, correlating cache keys against the issuing virtual key.
- Inspect responses containing tool invocations and verify the originating prompt belongs to the same principal that will execute them.
- Compare the deployed LiteLLM version against 1.101.0-rc.1 to confirm whether the vulnerable code path is present.
Monitoring Recommendations
- Enable verbose request/response logging on the proxy and forward to a centralized SIEM for cross-tenant correlation.
- Alert on agentic workflows that auto-execute tool calls which were not explicitly requested in the user prompt.
- Monitor embedding backends (such as Qdrant) for lookups whose returned payloads reference a different tenant metadata value than the querying session.
How to Mitigate CVE-2026-89032
Immediate Actions Required
- Upgrade BerriAI LiteLLM to 1.101.0-rc.1 or later immediately.
- Disable the semantic cache layer on multi-tenant deployments until the upgrade is applied.
- Rotate any virtual keys that may have been used by untrusted tenants during the exposure window.
- Review cached data for sensitive content such as PII, financial records, or source code that may have been exposed.
Patch Information
The fix was delivered in GitHub Pull Request #39590 and shipped in LiteLLM release v1.101.0-rc.1. It introduces a SemanticCacheScope enum with KEY and END_USER values and adds a semantic_cache_scope parameter so cache entries are consistently namespaced. Additional context is documented in the VulnCheck advisory.
Workarounds
- Set the proxy configuration to disable semantic caching on routes serving multiple tenants.
- Partition LiteLLM deployments so each tenant uses a dedicated proxy instance and backing vector store.
- Strip or sanitize function_call and tool_calls fields in responses before passing them to agentic front-ends.
# Upgrade LiteLLM to the patched release
pip install --upgrade "litellm>=1.101.0-rc.1"
# After upgrade, explicitly set the semantic cache scope
# in your LiteLLM proxy config to isolate per end user
# litellm_settings:
# cache: true
# cache_params:
# type: qdrant-semantic
# semantic_cache_scope: end_user
Disclaimer: This content was generated using AI. While we strive for accuracy, please verify critical information with official sources.