CVE-2026-94622 Overview
CVE-2026-94622 is a denial of service vulnerability in vLLM versions through 0.29.0. The flaw resides in the NIXL connector's metadata handling for prefill/decode disaggregated deployments. Attackers can send requests containing incomplete kv_transfer_params dictionary entries. This triggers an uncaught KeyError inside EngineCore scheduling. The decode engine terminates and all routed requests fail until an operator performs a manual restart. The issue is tracked under [CWE-248] Uncaught Exception. No authentication or user interaction is required to exploit the flaw over the network.
Critical Impact
A single malformed request can crash the vLLM decode engine, halting all inference traffic in disaggregated deployments until manual intervention restores service.
Affected Products
- vLLM versions through 0.29.0
- Deployments using the NIXL KV transfer connector
- Prefill/decode disaggregated inference configurations
Discovery Timeline
- 2026-09-21 - CVE-2026-94622 published to the National Vulnerability Database
- 2026-09-22 - Last updated in NVD database
Technical Details for CVE-2026-94622
Vulnerability Analysis
vLLM is an open-source inference and serving engine for large language models. In disaggregated deployments, prefill and decode phases execute on separate engines and coordinate key-value cache transfers using the NIXL connector. The connector expects each request to include a fully populated kv_transfer_params dictionary describing the remote KV cache location and metadata.
The metadata handler defined in vllm/distributed/kv_transfer/kv_connector/v1/nixl/metadata.py accesses dictionary keys directly without validating their presence. When a client submits a request that omits required fields, Python raises a KeyError inside the EngineCore scheduling loop. Because the exception is not caught, the decode engine process terminates. All in-flight and subsequently routed requests fail until an operator restarts the engine manually.
Root Cause
The root cause is an uncaught exception ([CWE-248]) in the NIXL connector metadata parsing logic. The code assumes trusted, well-formed input from upstream schedulers rather than validating attacker-controlled request payloads. Missing keys propagate as fatal exceptions instead of returning a request-level error.
Attack Vector
An unauthenticated attacker with network access to the vLLM inference endpoint submits a crafted request whose kv_transfer_params dictionary omits one or more required entries. The malformed metadata reaches the decode engine, the KeyError propagates through EngineCore, and the engine process exits. See the VulnCheck advisory and the vulnerable metadata code for technical details. No proof-of-concept exploit code is available in the referenced sources.
Detection Methods for CVE-2026-94622
Indicators of Compromise
- Unexpected termination of vLLM decode engine processes with KeyError tracebacks referencing nixl/metadata.py.
- Spikes in HTTP 5xx responses or dropped inference requests coincident with engine restarts.
- Inbound requests containing partial or malformed kv_transfer_params payloads from unusual client addresses.
Detection Strategies
- Parse vLLM application logs for KeyError exceptions originating in the kv_transfer module and correlate with engine process exits.
- Inspect ingress traffic for POST requests to inference endpoints missing expected kv_transfer_params keys such as remote host, port, or engine identifier.
- Alert on repeated decode engine restarts within short time windows, which indicate exploitation attempts rather than routine operational churn.
Monitoring Recommendations
- Instrument the vLLM service with process supervision and export exit codes to centralized logging.
- Track request success rates per client source address to identify hosts sending anomalous payloads.
- Forward inference gateway telemetry to a SIEM for correlation with engine crash events.
How to Mitigate CVE-2026-94622
Immediate Actions Required
- Restrict network access to vLLM inference endpoints so only authenticated internal services can reach them.
- Deploy an API gateway or reverse proxy that validates request schemas and rejects payloads missing required kv_transfer_params fields.
- Configure a supervisor such as systemd or Kubernetes liveness probes to automatically restart terminated decode engines while a patch is applied.
Patch Information
A fix is tracked in vLLM pull request #54807. Upgrade to a vLLM release that incorporates the patched NIXL metadata handling. Review the vLLM project repository for the latest release notes and confirm the corrected metadata.py is present before returning affected deployments to production.
Workarounds
- Place a validating proxy in front of vLLM that enforces a strict JSON schema for kv_transfer_params and drops incomplete requests.
- Disable the NIXL connector or the prefill/decode disaggregated deployment mode until an upgrade is possible.
- Require mutual TLS or network-layer authentication between prefill and decode engines to prevent untrusted clients from reaching the vulnerable code path.
# Configuration example: upgrade vLLM to a patched release
pip install --upgrade "vllm>0.29.0"
# Verify installed version
python -c "import vllm; print(vllm.__version__)"
Disclaimer: This content was generated using AI. While we strive for accuracy, please verify critical information with official sources.
