CVE-2026-94627 Overview
CVE-2026-94627 is a memory leak vulnerability in the vLLM Mooncake connector through version 0.29.0. The flaw affects prefill/decode disaggregated deployments where concurrent child requests share a single transfer ID. The connector fails to properly manage GPU Key-Value (KV) cache block ownership, allowing orphaned blocks to accumulate in GPU memory. Attackers can submit completion requests containing multiple prompts to trigger GPU memory exhaustion. Accumulated leakage persists until the vLLM process is restarted, eventually preventing legitimate inference requests from executing.
Critical Impact
Unauthenticated network-based attackers can exhaust GPU memory on vLLM inference servers, producing a persistent denial-of-service condition that requires process restart to recover.
Affected Products
- vLLM Mooncake connector through version 0.29.0
- Deployments using prefill/decode disaggregated serving with the Mooncake KV transfer connector
- Multi-prompt completion request handlers in mooncake_connector.py
Discovery Timeline
- 2026-09-21 - CVE-2026-94627 published to the National Vulnerability Database (NVD)
- 2026-09-22 - Last updated in NVD database
Technical Details for CVE-2026-94627
Vulnerability Analysis
The vulnerability resides in the vLLM Mooncake connector logic responsible for tracking GPU KV cache block ownership across concurrent child requests. In prefill/decode disaggregated deployments, when a completion request contains multiple prompts, vLLM spawns multiple child requests that share a single Mooncake transfer ID. The connector uses that transfer ID as the key for block ownership tracking, causing collisions when multiple children reference the same identifier.
When child requests complete, the connector releases blocks associated with the shared transfer ID but does not distinguish which child owns which allocation. Blocks belonging to sibling requests become orphaned and remain resident in GPU memory. The cleanup path never revisits these blocks, so they leak until the process is restarted. This behavior maps to [CWE-401: Missing Release of Memory after Effective Lifetime].
Root Cause
The root cause is a transfer-ID collision in the connector's block-ownership bookkeeping. The code at lines 1978–1989 of mooncake_connector.py in version 0.29.0 assumes a one-to-one mapping between transfer IDs and requests. Multi-prompt completions break that assumption, and the deallocation routine cannot correctly attribute blocks to individual child requests.
Attack Vector
An unauthenticated remote attacker submits completion requests to the vLLM HTTP API with multiple prompts per request. Each malformed request leaks a portion of GPU KV cache memory. Repeated submissions drive cumulative leakage until GPU memory is exhausted and the server can no longer schedule legitimate inference workloads. No authentication, user interaction, or elevated privileges are required.
A proof-of-concept implementation is not publicly listed. Technical details are documented in the VulnCheck Advisory for vLLM Cache Leak and the upstream vLLM Pull Request #49796.
Detection Methods for CVE-2026-94627
Indicators of Compromise
- Steadily rising GPU memory utilization on vLLM workers with no corresponding increase in active request count
- Completion API requests containing large numbers of prompts in a single call from unexpected sources
- Recurrent out-of-memory errors or CUDA allocation failures in vLLM logs preceding process restarts
- Growing count of Mooncake transfer IDs referenced by multiple child requests in connector debug output
Detection Strategies
- Baseline normal GPU memory usage per model and alert when utilization trends upward without matching workload growth
- Inspect vLLM access logs for completion requests with abnormally high prompt-array lengths
- Correlate nvidia-smi memory metrics with vLLM request throughput to identify divergence indicative of leakage
Monitoring Recommendations
- Export GPU memory, KV cache block count, and request queue depth metrics to a centralized observability platform
- Instrument the Mooncake connector to log transfer-ID reuse events and allocation/deallocation deltas
- Set alert thresholds on time-to-OOM and vLLM worker restart frequency to surface exploitation attempts
How to Mitigate CVE-2026-94627
Immediate Actions Required
- Upgrade vLLM to a release that includes the fix from vLLM Pull Request #49796 once available beyond 0.29.0
- Restrict network access to the vLLM inference endpoint using an authenticated API gateway or network policy
- Reject or rewrite completion requests that contain more than one prompt per call at the reverse proxy layer
- Schedule periodic vLLM process restarts as a temporary control to recover leaked GPU memory
Patch Information
The upstream fix is tracked in vLLM Pull Request #49796, which corrects block-ownership tracking in the Mooncake connector so that concurrent child requests sharing a transfer ID release their allocations correctly. Review the vLLM Project Repository for the merged release version and apply the corresponding update. Reference the vulnerable code region in mooncake_connector.py lines 1978-1989 to validate that your deployment includes the patched logic.
Workarounds
- Disable prefill/decode disaggregated deployment mode until the patched version is deployed
- Configure API gateway rules to limit completion requests to a single prompt per call
- Apply strict rate limits on the completion endpoint to slow accumulation of leaked KV cache blocks
- Monitor GPU memory continuously and automate vLLM worker restarts when utilization exceeds a defined threshold
# Example NGINX rule to reject multi-prompt completion requests
# Adjust based on your deployment's request schema validation tooling
location /v1/completions {
if ($request_method = POST) {
access_by_lua_block {
local cjson = require "cjson.safe"
ngx.req.read_body()
local body = ngx.req.get_body_data()
local ok, data = pcall(cjson.decode, body)
if ok and type(data.prompt) == "table" and #data.prompt > 1 then
ngx.exit(400)
end
}
}
proxy_pass http://vllm_backend;
}
Disclaimer: This content was generated using AI. While we strive for accuracy, please verify critical information with official sources.
