CVE-2026-94625 Overview
CVE-2026-94625 is a resource exhaustion vulnerability affecting vLLM through version 0.29.0. The flaw resides in the MooncakeConnector component used for key-value cache transfers between distributed inference nodes. Rejected prefill requests create transfer placeholders that lack an owner and are never reclaimed by the runtime. Attackers can send a stream of rejected requests to exhaust the sender task pool. Legitimate requests then experience delays of up to 480 seconds while health checks continue reporting the service as healthy. The vulnerability is tracked under CWE-772: Missing Release of Resource after Effective Lifetime.
Critical Impact
Unauthenticated remote attackers can degrade vLLM inference availability by exhausting Mooncake transfer resources while health probes continue to succeed.
Affected Products
- vLLM versions through 0.29.0
- Deployments using the MooncakeConnector for distributed KV cache transfer
- Inference clusters exposing vLLM endpoints to untrusted networks
Discovery Timeline
- 2026-09-21 - CVE-2026-94625 published to the National Vulnerability Database
- 2026-09-23 - Last updated in NVD database
Technical Details for CVE-2026-94625
Vulnerability Analysis
vLLM is a high-throughput inference engine for large language models. The MooncakeConnector handles KV cache transfers between prefill and decode workers in disaggregated serving deployments. When a prefill request is rejected downstream, the connector allocates a transfer placeholder inside the sender task pool but does not associate it with an owning request context.
The placeholders remain allocated indefinitely because no cleanup path releases them once the originating request is rejected. Repeated rejections accumulate ownerless entries until the sender task pool reaches its capacity. New transfer requests then queue behind the leaked placeholders, introducing latency of up to 480 seconds. The service health endpoint reports success throughout the incident, masking the degradation from operators and load balancers.
Root Cause
The root cause is a missing resource release path in the rejection handling logic of mooncake_connector.py. The relevant code region is documented in the Mooncake connector source at lines 1234-1242. Because ownership is never assigned before rejection, the standard reclamation logic tied to request completion never runs against the placeholder.
Attack Vector
Exploitation requires only network access to the vLLM inference endpoint. No authentication or user interaction is required. An attacker submits a sequence of prefill requests that trigger the rejection path in the Mooncake connector. Each rejected request consumes one slot in the sender task pool. Once the pool is saturated, valid inference traffic stalls while liveness probes continue returning healthy status. Refer to the VulnCheck advisory and vLLM pull request #51236 for the remediation details.
Detection Methods for CVE-2026-94625
Indicators of Compromise
- Sustained growth in outstanding transfer placeholder counts within the MooncakeConnector sender task pool.
- Inference request latency approaching or reaching the 480-second ceiling while /health endpoints return 200 OK.
- Repeated prefill request rejections originating from the same client identifiers or IP ranges.
Detection Strategies
- Instrument vLLM with metrics that track sender task pool utilization and correlate saturation with rejection counts.
- Compare application-layer response latency against health check status to detect divergence between actual service quality and reported liveness.
- Alert on repeated rejected prefill requests from a single source over short time windows.
Monitoring Recommendations
- Ship vLLM logs and Mooncake connector metrics into a centralized analytics pipeline for trend analysis.
- Track queue depth, transfer placeholder count, and rejection ratios as first-class service level indicators.
- Configure synthetic transactions that measure end-to-end inference latency independent of the built-in health probe.
How to Mitigate CVE-2026-94625
Immediate Actions Required
- Restrict network exposure of vLLM inference endpoints to trusted clients using network policies or an authenticating proxy.
- Apply rate limiting and per-client quotas in front of vLLM to bound the number of concurrent prefill requests.
- Replace or supplement the default /health probe with a functional probe that submits an inference request and measures latency.
Patch Information
The fix is tracked in vLLM pull request #51236. Upgrade to a vLLM release that includes the pull request once available. Review the vLLM project repository for the fixed version tag and release notes before deploying.
Workarounds
- Disable the MooncakeConnector and use an alternative KV transfer backend if disaggregated serving is not required.
- Schedule periodic restarts of vLLM worker processes to reclaim leaked placeholders until the patch can be applied.
- Deploy a reverse proxy that terminates and validates prefill requests before they reach the connector, dropping malformed traffic early.
# Configuration example: restrict vLLM to internal traffic and rate limit prefill requests
# NetworkPolicy fragment (Kubernetes)
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: vllm-restrict-ingress
spec:
podSelector:
matchLabels:
app: vllm
ingress:
- from:
- namespaceSelector:
matchLabels:
trust: internal
Disclaimer: This content was generated using AI. While we strive for accuracy, please verify critical information with official sources.
