CVE-2026-94624 Overview
CVE-2026-94624 is a denial of service vulnerability in vLLM through version 0.29.0. The flaw exists in the peer-to-peer key-value (KV) offloading path when OffloadingConnector is configured with TieringOffloadingSpec and a P2P secondary tier. Remote attackers can supply arbitrary host and port values in kv_transfer_params, creating unreachable peer sessions that retain ZeroMQ sockets until the context quota is exhausted. The resulting uncaught ZMQError crashes EngineCore and halts all inference on the affected instance. The weakness is classified as allocation of resources without limits or throttling [CWE-770].
Critical Impact
An unauthenticated network attacker can crash the vLLM inference engine and stop all model serving with a small number of crafted requests.
Affected Products
- vLLM through version 0.29.0
- Deployments using OffloadingConnector with TieringOffloadingSpec
- Configurations enabling a peer-to-peer secondary offloading tier
Discovery Timeline
- 2026-09-21 - CVE-2026-94624 published to NVD
- 2026-09-22 - Last updated in NVD database
Technical Details for CVE-2026-94624
Vulnerability Analysis
vLLM implements a tiered KV cache offloading system that can transfer cache blocks between peer engines over ZeroMQ. When a request includes kv_transfer_params, the P2P manager initiates a ZeroMQ session to the specified remote host and port to negotiate cache transfer.
The control-plane code in vllm/v1/kv_offload/tiering/p2p/control/zmq.py and the manager logic in vllm/v1/kv_offload/tiering/p2p/manager.py do not validate the supplied endpoints and do not bound the number of concurrent outstanding peer sessions. Sockets opened toward unreachable peers remain allocated against the ZeroMQ context. Once the context quota is exhausted, a subsequent socket allocation raises ZMQError, which propagates uncaught out of EngineCore and terminates the worker.
Because the engine process owns all inference scheduling, its crash stops serving for every tenant on that instance. Recovery requires a manual or orchestrator-driven restart.
Root Cause
The root cause is missing resource governance on attacker-controlled endpoints. The offloading manager treats kv_transfer_params as trusted routing input and creates ZeroMQ sockets without enforcing a per-client cap, connection timeout, or reachability check. There is also no exception handler around the socket allocation, so a ZMQError becomes fatal to the engine loop.
Attack Vector
Exploitation requires only network access to a vLLM endpoint that accepts inference requests and forwards kv_transfer_params to the P2P offloading path. The attacker repeatedly submits requests referencing non-routable or firewalled peer host and port combinations. Each request consumes ZeroMQ socket resources that are never reclaimed within the session lifetime. When the ZeroMQ context reaches its socket limit, EngineCore throws ZMQError and exits.
Review the vulnerable functions in the ZMQ control code and the P2P manager code for the exact resource-allocation paths. No verified public exploit code is available at this time.
Detection Methods for CVE-2026-94624
Indicators of Compromise
- Repeated inbound inference requests containing kv_transfer_params with external or previously unseen host and port values.
- Unexplained ZMQError entries in vLLM engine logs preceding an EngineCore termination.
- Sudden clusters of failed model-serving health checks correlated with a spike in outbound TCP connection attempts from vLLM workers to unreachable peers.
Detection Strategies
- Parse vLLM stdout and stderr for ZMQError, zmq.error.ZMQError, or Too many open files conditions that precede engine restarts.
- Correlate API gateway logs to identify clients submitting kv_transfer_params targeting external IP ranges outside the approved peer inventory.
- Baseline the socket count of the vLLM process and alert on sustained growth of half-open connections from the engine worker.
Monitoring Recommendations
- Track EngineCore restart frequency and set alerts on any restart caused by an unhandled exception.
- Monitor outbound connection attempts from inference workers to endpoints outside the P2P peer allowlist.
- Instrument ZeroMQ context metrics such as active socket count and pending connection count where the runtime exposes them.
How to Mitigate CVE-2026-94624
Immediate Actions Required
- Disable OffloadingConnector with TieringOffloadingSpec and the P2P secondary tier where the feature is not required for production.
- Restrict access to vLLM inference endpoints so that only authenticated internal clients can submit kv_transfer_params.
- Apply egress network controls that limit vLLM workers to communicating with a fixed allowlist of known peer hosts and ports.
Patch Information
A fix is tracked in the upstream project via Pull Request #51504. Refer to the VulnCheck advisory and the vLLM project repository for release notes and the fixed version. Upgrade past 0.29.0 once the patched release is available.
Workarounds
- Deploy an API proxy that strips or validates kv_transfer_params against a strict allowlist before requests reach vLLM.
- Run vLLM behind a network policy that blocks arbitrary outbound TCP from engine workers to unapproved destinations.
- Configure process supervisors to rate-limit EngineCore restarts and to page on-call staff when repeated crashes occur.
# Configuration example: block unapproved kv_transfer_params destinations at the gateway
# Example NGINX snippet rejecting requests referencing external P2P peers
if ($request_body ~* "kv_transfer_params") {
set $needs_validation 1;
}
location /v1/completions {
if ($needs_validation) { return 403; }
proxy_pass http://vllm_backend;
}
Disclaimer: This content was generated using AI. While we strive for accuracy, please verify critical information with official sources.
