Skip to main content
Vulnerability Database/CVE-2026-94627

CVE-2026-94627: vLLM Mooncake Connector DOS Vulnerability

CVE-2026-94627 is a denial of service vulnerability in vLLM Mooncake connector that allows attackers to exhaust GPU memory through orphaned KV cache blocks. This post explains its impact, affected versions, and mitigation steps.

Published:

CVE-2026-94627 Overview

CVE-2026-94627 is a memory leak vulnerability in the vLLM Mooncake connector through version 0.29.0. The flaw affects prefill/decode disaggregated deployments where concurrent child requests share a single transfer ID. The connector fails to properly manage GPU Key-Value (KV) cache block ownership, allowing orphaned blocks to accumulate in GPU memory. Attackers can submit completion requests containing multiple prompts to trigger GPU memory exhaustion. Accumulated leakage persists until the vLLM process is restarted, eventually preventing legitimate inference requests from executing.

Critical Impact

Unauthenticated network-based attackers can exhaust GPU memory on vLLM inference servers, producing a persistent denial-of-service condition that requires process restart to recover.

Affected Products

  • vLLM Mooncake connector through version 0.29.0
  • Deployments using prefill/decode disaggregated serving with the Mooncake KV transfer connector
  • Multi-prompt completion request handlers in mooncake_connector.py

Discovery Timeline

  • 2026-09-21 - CVE-2026-94627 published to the National Vulnerability Database (NVD)
  • 2026-09-22 - Last updated in NVD database

Technical Details for CVE-2026-94627

Vulnerability Analysis

The vulnerability resides in the vLLM Mooncake connector logic responsible for tracking GPU KV cache block ownership across concurrent child requests. In prefill/decode disaggregated deployments, when a completion request contains multiple prompts, vLLM spawns multiple child requests that share a single Mooncake transfer ID. The connector uses that transfer ID as the key for block ownership tracking, causing collisions when multiple children reference the same identifier.

When child requests complete, the connector releases blocks associated with the shared transfer ID but does not distinguish which child owns which allocation. Blocks belonging to sibling requests become orphaned and remain resident in GPU memory. The cleanup path never revisits these blocks, so they leak until the process is restarted. This behavior maps to [CWE-401: Missing Release of Memory after Effective Lifetime].

Root Cause

The root cause is a transfer-ID collision in the connector's block-ownership bookkeeping. The code at lines 1978–1989 of mooncake_connector.py in version 0.29.0 assumes a one-to-one mapping between transfer IDs and requests. Multi-prompt completions break that assumption, and the deallocation routine cannot correctly attribute blocks to individual child requests.

Attack Vector

An unauthenticated remote attacker submits completion requests to the vLLM HTTP API with multiple prompts per request. Each malformed request leaks a portion of GPU KV cache memory. Repeated submissions drive cumulative leakage until GPU memory is exhausted and the server can no longer schedule legitimate inference workloads. No authentication, user interaction, or elevated privileges are required.

A proof-of-concept implementation is not publicly listed. Technical details are documented in the VulnCheck Advisory for vLLM Cache Leak and the upstream vLLM Pull Request #49796.

Detection Methods for CVE-2026-94627

Indicators of Compromise

  • Steadily rising GPU memory utilization on vLLM workers with no corresponding increase in active request count
  • Completion API requests containing large numbers of prompts in a single call from unexpected sources
  • Recurrent out-of-memory errors or CUDA allocation failures in vLLM logs preceding process restarts
  • Growing count of Mooncake transfer IDs referenced by multiple child requests in connector debug output

Detection Strategies

  • Baseline normal GPU memory usage per model and alert when utilization trends upward without matching workload growth
  • Inspect vLLM access logs for completion requests with abnormally high prompt-array lengths
  • Correlate nvidia-smi memory metrics with vLLM request throughput to identify divergence indicative of leakage

Monitoring Recommendations

  • Export GPU memory, KV cache block count, and request queue depth metrics to a centralized observability platform
  • Instrument the Mooncake connector to log transfer-ID reuse events and allocation/deallocation deltas
  • Set alert thresholds on time-to-OOM and vLLM worker restart frequency to surface exploitation attempts

How to Mitigate CVE-2026-94627

Immediate Actions Required

  • Upgrade vLLM to a release that includes the fix from vLLM Pull Request #49796 once available beyond 0.29.0
  • Restrict network access to the vLLM inference endpoint using an authenticated API gateway or network policy
  • Reject or rewrite completion requests that contain more than one prompt per call at the reverse proxy layer
  • Schedule periodic vLLM process restarts as a temporary control to recover leaked GPU memory

Patch Information

The upstream fix is tracked in vLLM Pull Request #49796, which corrects block-ownership tracking in the Mooncake connector so that concurrent child requests sharing a transfer ID release their allocations correctly. Review the vLLM Project Repository for the merged release version and apply the corresponding update. Reference the vulnerable code region in mooncake_connector.py lines 1978-1989 to validate that your deployment includes the patched logic.

Workarounds

  • Disable prefill/decode disaggregated deployment mode until the patched version is deployed
  • Configure API gateway rules to limit completion requests to a single prompt per call
  • Apply strict rate limits on the completion endpoint to slow accumulation of leaked KV cache blocks
  • Monitor GPU memory continuously and automate vLLM worker restarts when utilization exceeds a defined threshold
bash
# Example NGINX rule to reject multi-prompt completion requests
# Adjust based on your deployment's request schema validation tooling
location /v1/completions {
    if ($request_method = POST) {
        access_by_lua_block {
            local cjson = require "cjson.safe"
            ngx.req.read_body()
            local body = ngx.req.get_body_data()
            local ok, data = pcall(cjson.decode, body)
            if ok and type(data.prompt) == "table" and #data.prompt > 1 then
                ngx.exit(400)
            end
        }
    }
    proxy_pass http://vllm_backend;
}

Disclaimer: This content was generated using AI. While we strive for accuracy, please verify critical information with official sources.

Default Legacy - Prefooter | Experience the World’s Most Advanced Cybersecurity Platform

Experience the Most Advanced Cybersecurity Platform

See how the world’s most intelligent, autonomous cybersecurity platform can protect your organization today and into the future.