CVE-2026-94626 Overview
CVE-2026-94626 is a memory exhaustion vulnerability in vLLM through version 0.29.0. The flaw resides in the OpenAI-compatible completion endpoints, which fail to validate the tp_size parameter in kv_transfer_params. Attackers can submit arbitrary tp_size values in prefill and decode disaggregated deployments, causing the server to allocate unbounded memory. The condition triggers a Linux kernel out-of-memory (OOM) kill of the decode worker process, disrupting inference availability. The weakness is classified under [CWE-789] Memory Allocation with Excessive Size Value and requires no authentication over the network.
Critical Impact
Unauthenticated remote attackers can crash vLLM decode workers by submitting oversized tp_size values, resulting in denial of service for production LLM inference deployments.
Affected Products
- vLLM versions up to and including 0.29.0
- Deployments using the NIXL KV transfer connector
- Prefill and decode disaggregated inference deployments exposing OpenAI-compatible endpoints
Discovery Timeline
- 2026-09-21 - CVE-2026-94626 published to NVD
- 2026-09-22 - Last updated in NVD database
Technical Details for CVE-2026-94626
Vulnerability Analysis
vLLM is a high-throughput inference and serving engine for large language models. In disaggregated deployments, the prefill and decode stages run as separate workers that exchange key-value (KV) cache state through the NIXL connector. The KV transfer layer accepts a tp_size field describing the tensor-parallel size of the remote peer. This field flows through the OpenAI-compatible completion API without a bounds check.
An attacker sends a completion request that includes a crafted kv_transfer_params object with an arbitrarily large tp_size. The decode worker then computes buffer allocations scaled by this attacker-controlled integer. Allocation size grows past available system memory, and the kernel terminates the worker to protect the host. Downstream inference requests fail until the process is restarted.
Root Cause
The root cause is missing input validation on tp_size before it is used to size memory allocations. The relevant code paths in vllm/distributed/kv_transfer/kv_connector/utils.py and vllm/distributed/kv_transfer/kv_connector/v1/nixl/metadata.py accept the value from remote request payloads and propagate it directly into allocation logic. Because no upper bound, sanity check, or peer authentication is enforced, the attacker fully controls the multiplier that drives memory sizing [CWE-789].
Attack Vector
Exploitation requires only network reachability to the vLLM OpenAI-compatible endpoint. An unauthenticated attacker issues an HTTP POST request to a completion endpoint containing a kv_transfer_params block with an inflated tp_size integer. No user interaction, elevated privileges, or prior foothold are required. The vulnerability affects availability only; confidentiality and integrity are not impacted. See the Vulncheck advisory and the fix in vLLM Pull Request #51137 for implementation details.
No verified public proof-of-concept code is available. The vulnerability mechanism can be reproduced by supplying an oversized integer for tp_size in the JSON body of a completion request routed through the NIXL connector.
Detection Methods for CVE-2026-94626
Indicators of Compromise
- Kernel OOM-killer log entries terminating vLLM decode worker processes, visible via dmesg or journalctl -k.
- Sudden RSS memory spikes on inference hosts followed by worker process restarts.
- HTTP request bodies to /v1/completions or /v1/chat/completions containing kv_transfer_params objects with unusually large tp_size values.
- Repeated 5xx errors or dropped connections from the vLLM API server following crafted requests.
Detection Strategies
- Inspect API gateway or reverse proxy logs for completion requests containing kv_transfer_params and flag payloads where tp_size exceeds the deployment's expected tensor-parallel dimension.
- Correlate kernel OOM events with the timestamps of inbound requests to the inference endpoint to identify malicious triggers.
- Deploy Web Application Firewall (WAF) rules that inspect JSON bodies for out-of-range integer values in tp_size.
Monitoring Recommendations
- Alert on abnormal memory allocation rates in vLLM worker containers using Prometheus or cgroup memory metrics.
- Monitor for restart loops on decode worker pods in Kubernetes environments.
- Track baseline distribution of tp_size values seen in production and alert on outliers.
How to Mitigate CVE-2026-94626
Immediate Actions Required
- Restrict network exposure of vLLM OpenAI-compatible endpoints to trusted clients using network policies or an authenticating reverse proxy.
- Apply input validation at the gateway layer to reject requests containing kv_transfer_params.tp_size values outside expected bounds.
- Enforce per-client rate limits on completion endpoints to slow enumeration and repeated abuse.
Patch Information
The fix is tracked in vLLM Pull Request #51137. Upgrade to a vLLM release that includes this pull request. Review the vulnerable code locations in utils.py and metadata.py to confirm the patched behavior in your build.
Workarounds
- Place the vLLM API behind an authenticating proxy and reject anonymous requests to completion endpoints.
- Strip or normalize the kv_transfer_params field from client-supplied request bodies at the ingress layer if disaggregated deployment is not required.
- Set container memory limits so a single crashed worker does not destabilize other services on the same host.
- Disable the NIXL KV transfer connector in deployments that do not rely on prefill and decode disaggregation.
# Example ingress-level validation using an NGINX/OpenResty snippet
# Reject requests with oversized tp_size in kv_transfer_params
access_by_lua_block {
ngx.req.read_body()
local body = ngx.req.get_body_data() or ""
local tp = body:match('"tp_size"%s*:%s*(%d+)')
if tp and tonumber(tp) > 16 then
ngx.status = 400
ngx.say('invalid tp_size')
return ngx.exit(400)
end
}
Disclaimer: This content was generated using AI. While we strive for accuracy, please verify critical information with official sources.
