CVE-2026-33625 Overview
CVE-2026-33625 is a code injection vulnerability in LMDeploy, a toolkit for compressing, deploying, and serving large language models. Versions 0.12.1 through 0.12.2 pass an attacker-controlled quantization_config.quant_dtype value from a HuggingFace model into a Python eval() call without validation. An attacker who publishes a malicious model can execute arbitrary Python code when a victim loads it with LMDeploy. The maintainers released version 0.12.3 with a patch. The weakness is tracked as [CWE-400] in the advisory metadata, though the underlying flaw is code injection via unsafe evaluation.
Critical Impact
Loading a malicious HuggingFace model with a vulnerable LMDeploy version yields arbitrary Python code execution in the user's process with full confidentiality, integrity, and availability impact.
Affected Products
- LMDeploy 0.12.1
- LMDeploy 0.12.2
- Fixed in LMDeploy 0.12.3
Discovery Timeline
- 2026-09-18 - CVE-2026-33625 published to the National Vulnerability Database (NVD)
- 2026-09-24 - Last updated in NVD database
Technical Details for CVE-2026-33625
Vulnerability Analysis
The vulnerability resides in lmdeploy/pytorch/config.py at line 620. LMDeploy reads the quantization_config.quant_dtype field from a model's configuration and interpolates it directly into an eval() expression of the form eval(f'torch.{quant_dtype}'). Because eval() executes arbitrary Python, any expression the attacker places in quant_dtype runs in the interpreter that loaded the model.
Exploitation requires a user to load an attacker-controlled model, which aligns with the required user interaction in the CVSS vector. HuggingFace's open publishing model makes distribution trivial. The impact is arbitrary code execution in the context of the LMDeploy process, which typically has access to local files, GPU resources, API keys, and any credentials used for model serving.
Root Cause
The root cause is unsafe use of Python's eval() on untrusted input. The code assumes quant_dtype names a torch attribute such as float16 or bfloat16, but no allowlist, type check, or parser restricts the value. Any Python expression concatenated into the f-string is evaluated. Configuration files from third parties should never reach eval() or exec() without strict validation.
Attack Vector
An attacker publishes a HuggingFace model repository containing a config.json (or equivalent quantization config) whose quant_dtype field holds a malicious Python expression, for example a call that spawns a subprocess or imports os and writes a payload. When a downstream user runs LMDeploy against that model identifier, the loader parses the config, passes quant_dtype into the eval() on line 620, and executes the attacker's expression. No authentication is required on the attacker side, and no network access to the victim is needed beyond the victim pulling the model.
See the GitHub Security Advisory GHSA-3hmm-rh5q-gwwr for the maintainers' technical description.
Detection Methods for CVE-2026-33625
Indicators of Compromise
- Model configuration files where quantization_config.quant_dtype is not a simple identifier such as float16, bfloat16, int8, or int4.
- Unexpected child processes (for example sh, bash, curl, wget, python -c) spawned by the Python process running LMDeploy shortly after a model load.
- Outbound network connections from an inference host to unknown domains during model initialization.
- New files, cron entries, or SSH keys created by the LMDeploy service account after loading a third-party model.
Detection Strategies
- Scan cached HuggingFace model directories for config.json files whose quant_dtype value contains characters outside [A-Za-z0-9_], such as (, ., ;, or whitespace.
- Inventory Python environments and flag any installation of lmdeploy at versions 0.12.1 or 0.12.2 using pip show lmdeploy or SBOM tooling.
- Enable Python audit hooks (sys.addaudithook) in inference workers to log compile and exec events triggered during model loading.
Monitoring Recommendations
- Alert on process-tree anomalies where the LMDeploy Python interpreter spawns shells or network utilities.
- Monitor egress from GPU inference nodes and baseline the destinations reached during legitimate model pulls.
- Track HuggingFace repository identifiers loaded in production and restrict them to a curated allowlist.
How to Mitigate CVE-2026-33625
Immediate Actions Required
- Upgrade LMDeploy to version 0.12.3 or later on every host that loads models.
- Audit HuggingFace model caches for suspicious quant_dtype values and remove any untrusted models pulled while running 0.12.1 or 0.12.2.
- Rotate credentials, API tokens, and SSH keys accessible to any inference process that loaded an unvetted third-party model.
- Restrict model sources to an internal registry or a vetted allowlist of HuggingFace repositories.
Patch Information
The fix is included in LMDeploy v0.12.3. Install with pip install --upgrade lmdeploy==0.12.3. Refer to the GitHub Security Advisory GHSA-3hmm-rh5q-gwwr for details on the code change.
Workarounds
- If upgrading immediately is not possible, load only models from trusted internal repositories where the configuration content is controlled.
- Pre-validate every third-party model's quantization_config.quant_dtype against an allowlist such as {float16, bfloat16, float32, int8, int4} before invoking LMDeploy.
- Run LMDeploy in a container with no outbound network access, a read-only filesystem, and a non-privileged user to contain any successful exploitation.
# Configuration example
pip install --upgrade "lmdeploy>=0.12.3"
pip show lmdeploy | grep -i version
Disclaimer: This content was generated using AI. While we strive for accuracy, please verify critical information with official sources.
