CVE-2026-105775 Overview
CVE-2026-105775 is an out-of-bounds read vulnerability in vllm-project vLLM through version 0.31.0. The flaw resides in the conv_ssm_forward function within vllm/model_executor/layers/mamba/mamba_mixer2.py, which handles Completions Request processing. An authenticated remote attacker can trigger the issue by sending crafted requests to the vulnerable endpoint. The exploit technique has been publicly disclosed. The upstream project was notified through an issue report but had not responded at the time of publication. This weakness is tracked under CWE-119 (improper restriction of operations within the bounds of a memory buffer).
Critical Impact
Remote attackers with low privileges can trigger an out-of-bounds memory read in the vLLM inference server, potentially causing availability loss within the Completions Request Handler.
Affected Products
- vllm-project vLLM versions up to and including 0.31.0
- vllm/model_executor/layers/mamba/mamba_mixer2.py component
- Completions Request Handler using the conv_ssm_forward function
Discovery Timeline
- 2026-10-06 - CVE-2026-105775 published to NVD
- 2026-10-06 - Last updated in NVD database
Technical Details for CVE-2026-105775
Vulnerability Analysis
The vulnerability affects the Mamba state-space model (SSM) mixer implementation inside vLLM, a high-throughput inference and serving engine for large language models. Within conv_ssm_forward, input tensor dimensions handled during a Completions Request are not properly validated before memory access. Processing attacker-controlled input causes the function to read beyond the intended buffer boundary. Because vLLM typically runs as a network-exposed inference service, the request handler path is reachable by any client with API credentials to the completions endpoint.
Out-of-bounds reads in tensor operations can leak adjacent memory contents into model computation paths or cause process-level failures that disrupt inference availability. The issue is tracked by VulDB as entry #413808 and corresponds to upstream GitHub Issue #57266.
Root Cause
The root cause is improper restriction of operations within the bounds of a memory buffer ([CWE-119]) in the conv_ssm_forward routine. The function does not sufficiently validate shape or offset parameters before performing the convolution and state-space mixing step. Malformed or boundary-condition input tensors cause reads outside the allocated region.
Attack Vector
The attack is network-based and requires low privilege (authenticated access to the completions API). No user interaction is required. An attacker submits a crafted completions request that drives the Mamba mixer code path with inputs that trigger the unsafe memory access. See the VulDB CVE-2026-105775 entry and the vLLM project repository for additional context. No verified proof-of-concept code is included in this advisory; refer to the upstream issue tracker for technical details.
Detection Methods for CVE-2026-105775
Indicators of Compromise
- Unexpected crashes, segmentation faults, or worker restarts in vLLM inference processes handling completions requests.
- Anomalous completions API payloads targeting Mamba-architecture models with unusual tensor shapes or sequence lengths.
- Repeated requests from a single authenticated client that correlate with inference worker instability.
Detection Strategies
- Monitor vLLM server logs for Python tracebacks originating in mamba_mixer2.py or the conv_ssm_forward function.
- Inspect completions request payloads for malformed input dimensions and establish baselines of expected tensor shapes per deployed model.
- Correlate authentication logs with inference process failures to identify abusive API clients.
Monitoring Recommendations
- Enable process-level telemetry on hosts running vLLM to capture crash signatures and memory-access faults.
- Forward vLLM application and container logs to a centralized analytics platform for anomaly detection and retention.
- Track completions endpoint error rates and latency to detect behavioral deviations indicative of exploitation attempts.
How to Mitigate CVE-2026-105775
Immediate Actions Required
- Restrict network exposure of vLLM completions endpoints to trusted internal clients and gateways only.
- Enforce strong authentication and per-client rate limiting on the completions API to reduce attack surface from low-privileged accounts.
- Review deployed Mamba-architecture models and disable them where feasible until a vendor patch is available.
Patch Information
No official vendor patch is referenced in the published advisory at the time of writing. The upstream project has been notified through GitHub Issue #57266. Monitor the vLLM project repository for a security release and apply fixes once published.
Workarounds
- Place an API gateway or reverse proxy in front of vLLM to validate request payload structure and reject malformed completions requests.
- Limit the model set served by vLLM to architectures that do not route through conv_ssm_forward until the issue is remediated.
- Run vLLM workers under process supervision with automatic restart and resource isolation to contain availability impact.
# Example: restrict vLLM completions endpoint to internal subnet only
iptables -A INPUT -p tcp --dport 8000 -s 10.0.0.0/8 -j ACCEPT
iptables -A INPUT -p tcp --dport 8000 -j DROP
Disclaimer: This content was generated using AI. While we strive for accuracy, please verify critical information with official sources.