CVE-2025-66019 Overview
CVE-2025-66019 is an uncontrolled resource consumption vulnerability [CWE-400] in pypdf, a widely used pure-Python PDF library. Attackers can craft a malicious PDF that triggers memory usage of up to 1 GB per stream when the library parses a page content stream using the LZWDecode filter. The condition affects all versions prior to 6.4.0 and enables a denial-of-service (DoS) condition against any application that processes untrusted PDF input with pypdf. The maintainers patched the issue in version 6.4.0 by reducing the maximum allowed LZW decoder output length.
Critical Impact
A single crafted PDF can exhaust up to 1 GB of memory per LZW-encoded stream, enabling remote denial-of-service against services that ingest untrusted PDFs.
Affected Products
- pypdf versions prior to 6.4.0
- Python applications and services that accept untrusted PDF input via pypdf
- Downstream libraries and pipelines that embed pypdf for PDF parsing
Discovery Timeline
- 2025-11-26 - CVE-2025-66019 published to the National Vulnerability Database (NVD)
- 2026-06-17 - Last updated in NVD database
Technical Details for CVE-2025-66019
Vulnerability Analysis
The vulnerability is a resource exhaustion issue triggered during decompression of LZW-encoded content streams inside a PDF page. When pypdf encounters a stream declared with the LZWDecode filter, it invokes an internal decoder whose maximum output length was configured to 1,000,000,000 bytes. An attacker who supplies compact LZW input can force the decoder to expand the data toward that limit, driving process memory to roughly 1 GB per stream. Applications that parse PDFs in a loop, run multiple worker processes, or accept concurrent uploads can be pushed into out-of-memory conditions with modest attacker effort.
Root Cause
The root cause is a permissive output-size cap on the LZW decoder. In pypdf/filters.py, the constant LZW_MAX_OUTPUT_LENGTH was set to 1,000,000,000, and the decoder class in pypdf/_codecs/_codecs.py accepted the same default via max_output_length. LZW compression can achieve high compression ratios, so a small input can legitimately decode to a very large buffer. Without a tighter cap, pypdf trusted attacker-controlled compression ratios and allocated memory accordingly.
Attack Vector
Exploitation is fully network-reachable and requires no authentication or user interaction. An attacker delivers a crafted PDF to any service that parses PDF pages with pypdf, for example, a document conversion API, an email attachment scanner, or an ingestion worker. Parsing the page content stream invokes the LZW filter and triggers the oversized allocation, degrading or crashing the target process.
# Security patch in pypdf/_codecs/_codecs.py
INITIAL_BITS_PER_CODE = 9 # Initial code bit width
MAX_BITS_PER_CODE = 12 # Maximum code bit width
- def __init__(self, max_output_length: int = 1_000_000_000) -> None:
+ def __init__(self, max_output_length: int = 75_000_000) -> None:
self.max_output_length = max_output_length
def _initialize_encoding_table(self) -> None:
Source: GitHub Commit 96186725
# Security patch in pypdf/filters.py
ZLIB_MAX_OUTPUT_LENGTH = 75_000_000
-LZW_MAX_OUTPUT_LENGTH = 1_000_000_000
+LZW_MAX_OUTPUT_LENGTH = 75_000_000
def _decompress_with_limit(data: bytes) -> bytes:
Source: GitHub Commit 96186725. The fix lowers the LZW output ceiling from 1 GB to 75 MB, aligning it with the existing zlib limit.
Detection Methods for CVE-2025-66019
Indicators of Compromise
- PDF files containing page content streams that declare the /LZWDecode filter and decompress to unusually large buffers
- Python worker processes running pypdf that show sudden resident-set-size growth toward 1 GB during PDF parsing
- Repeated out-of-memory kills or container restarts on PDF-processing services shortly after inbound document uploads
Detection Strategies
- Inspect PDF uploads for streams that use /Filter /LZWDecode and compare declared or observed decompressed sizes against a reasonable threshold
- Instrument pypdf callers with per-request memory ceilings using resource.setrlimit or container memory limits, and alert when limits trigger
- Track the installed pypdf version across environments using software composition analysis and flag any version below 6.4.0
Monitoring Recommendations
- Alert on cgroup or container OOM events for services that ingest documents
- Log PDF filter chains at ingestion time and flag content using LZWDecode for review
- Correlate spikes in memory usage with request identifiers to isolate the offending document
How to Mitigate CVE-2025-66019
Immediate Actions Required
- Upgrade pypdf to version 6.4.0 or later across all applications, containers, and build pipelines
- Enumerate transitive dependencies with pip list or pip-audit to locate embedded pypdf installations
- Enforce per-process memory limits on services that parse untrusted PDFs to bound worst-case impact
Patch Information
The fix is included in pypdf 6.4.0. See the GitHub Release Version 6.4.0 and the GitHub Security Advisory GHSA-m449-cwjh-6pw7. Additional technical analysis is available in the CVE-2025-66019 pypdf LZW DoS Blog Post.
Workarounds
- Reject PDFs that declare the LZWDecode filter at the perimeter until patching completes
- Isolate PDF parsing in short-lived worker processes with strict memory and CPU limits
- Apply the upstream cap manually by pinning LZW_MAX_OUTPUT_LENGTH to 75,000,000 in a vendored copy of pypdf if immediate upgrade is not possible
# Upgrade pypdf to the patched release
pip install --upgrade 'pypdf>=6.4.0'
# Verify the installed version
python -c "import pypdf; print(pypdf.__version__)"
# Audit dependencies for vulnerable versions
pip-audit
Disclaimer: This content was generated using AI. While we strive for accuracy, please verify critical information with official sources.
