Skip to main content
Vulnerability Database/CVE-2026-94624

CVE-2026-94624: vLLM P2P KV Offloading DoS Vulnerability

CVE-2026-94624 is a denial of service vulnerability in vLLM through version 0.29.0 affecting P2P KV offloading functionality. Attackers can crash EngineCore by exhausting ZeroMQ socket quotas. This article covers technical details, affected versions, impact assessment, and mitigation strategies.

Published:

CVE-2026-94624 Overview

CVE-2026-94624 is a denial of service vulnerability in vLLM through version 0.29.0. The flaw exists in the peer-to-peer key-value (KV) offloading path when OffloadingConnector is configured with TieringOffloadingSpec and a P2P secondary tier. Remote attackers can supply arbitrary host and port values in kv_transfer_params, creating unreachable peer sessions that retain ZeroMQ sockets until the context quota is exhausted. The resulting uncaught ZMQError crashes EngineCore and halts all inference on the affected instance. The weakness is classified as allocation of resources without limits or throttling [CWE-770].

Critical Impact

An unauthenticated network attacker can crash the vLLM inference engine and stop all model serving with a small number of crafted requests.

Affected Products

  • vLLM through version 0.29.0
  • Deployments using OffloadingConnector with TieringOffloadingSpec
  • Configurations enabling a peer-to-peer secondary offloading tier

Discovery Timeline

  • 2026-09-21 - CVE-2026-94624 published to NVD
  • 2026-09-22 - Last updated in NVD database

Technical Details for CVE-2026-94624

Vulnerability Analysis

vLLM implements a tiered KV cache offloading system that can transfer cache blocks between peer engines over ZeroMQ. When a request includes kv_transfer_params, the P2P manager initiates a ZeroMQ session to the specified remote host and port to negotiate cache transfer.

The control-plane code in vllm/v1/kv_offload/tiering/p2p/control/zmq.py and the manager logic in vllm/v1/kv_offload/tiering/p2p/manager.py do not validate the supplied endpoints and do not bound the number of concurrent outstanding peer sessions. Sockets opened toward unreachable peers remain allocated against the ZeroMQ context. Once the context quota is exhausted, a subsequent socket allocation raises ZMQError, which propagates uncaught out of EngineCore and terminates the worker.

Because the engine process owns all inference scheduling, its crash stops serving for every tenant on that instance. Recovery requires a manual or orchestrator-driven restart.

Root Cause

The root cause is missing resource governance on attacker-controlled endpoints. The offloading manager treats kv_transfer_params as trusted routing input and creates ZeroMQ sockets without enforcing a per-client cap, connection timeout, or reachability check. There is also no exception handler around the socket allocation, so a ZMQError becomes fatal to the engine loop.

Attack Vector

Exploitation requires only network access to a vLLM endpoint that accepts inference requests and forwards kv_transfer_params to the P2P offloading path. The attacker repeatedly submits requests referencing non-routable or firewalled peer host and port combinations. Each request consumes ZeroMQ socket resources that are never reclaimed within the session lifetime. When the ZeroMQ context reaches its socket limit, EngineCore throws ZMQError and exits.

Review the vulnerable functions in the ZMQ control code and the P2P manager code for the exact resource-allocation paths. No verified public exploit code is available at this time.

Detection Methods for CVE-2026-94624

Indicators of Compromise

  • Repeated inbound inference requests containing kv_transfer_params with external or previously unseen host and port values.
  • Unexplained ZMQError entries in vLLM engine logs preceding an EngineCore termination.
  • Sudden clusters of failed model-serving health checks correlated with a spike in outbound TCP connection attempts from vLLM workers to unreachable peers.

Detection Strategies

  • Parse vLLM stdout and stderr for ZMQError, zmq.error.ZMQError, or Too many open files conditions that precede engine restarts.
  • Correlate API gateway logs to identify clients submitting kv_transfer_params targeting external IP ranges outside the approved peer inventory.
  • Baseline the socket count of the vLLM process and alert on sustained growth of half-open connections from the engine worker.

Monitoring Recommendations

  • Track EngineCore restart frequency and set alerts on any restart caused by an unhandled exception.
  • Monitor outbound connection attempts from inference workers to endpoints outside the P2P peer allowlist.
  • Instrument ZeroMQ context metrics such as active socket count and pending connection count where the runtime exposes them.

How to Mitigate CVE-2026-94624

Immediate Actions Required

  • Disable OffloadingConnector with TieringOffloadingSpec and the P2P secondary tier where the feature is not required for production.
  • Restrict access to vLLM inference endpoints so that only authenticated internal clients can submit kv_transfer_params.
  • Apply egress network controls that limit vLLM workers to communicating with a fixed allowlist of known peer hosts and ports.

Patch Information

A fix is tracked in the upstream project via Pull Request #51504. Refer to the VulnCheck advisory and the vLLM project repository for release notes and the fixed version. Upgrade past 0.29.0 once the patched release is available.

Workarounds

  • Deploy an API proxy that strips or validates kv_transfer_params against a strict allowlist before requests reach vLLM.
  • Run vLLM behind a network policy that blocks arbitrary outbound TCP from engine workers to unapproved destinations.
  • Configure process supervisors to rate-limit EngineCore restarts and to page on-call staff when repeated crashes occur.
bash
# Configuration example: block unapproved kv_transfer_params destinations at the gateway
# Example NGINX snippet rejecting requests referencing external P2P peers
if ($request_body ~* "kv_transfer_params") {
    set $needs_validation 1;
}
location /v1/completions {
    if ($needs_validation) { return 403; }
    proxy_pass http://vllm_backend;
}

Disclaimer: This content was generated using AI. While we strive for accuracy, please verify critical information with official sources.

Default Legacy - Prefooter | Experience the World’s Most Advanced Cybersecurity Platform

Experience the Most Advanced Cybersecurity Platform

See how the world’s most intelligent, autonomous cybersecurity platform can protect your organization today and into the future.