Update dependency unstructured to v0.24.0 [SECURITY] #99

Open
renovate[bot] wants to merge 1 commit from renovate/pypi-unstructured-vulnerability into main
renovate[bot] commented 2026-06-11 18:58:50 +00:00 (Migrated from github.com)

This PR contains the following updates:

Package Change Age Adoption Passing Confidence
unstructured ==0.17.2 → ==0.24.0 age adoption passing confidence

Unstructured has Path Traversal via Malicious MSG Attachment that Allows Arbitrary File Write

CVE-2025-64712 / GHSA-gm8q-m8mv-jj5m

More information

Details

A Path Traversal vulnerability in the partition_msg function allows an attacker to write or overwrite arbitrary files on the filesystem when processing malicious MSG files with attachments.

Impact

An attacker can craft a malicious .msg file with attachment filenames containing path traversal sequences (e.g.,
../../../etc/cron.d/malicious). When processed with process_attachments=True, the library writes the attachment to an
attacker-controlled path, potentially leading to:

  • Arbitrary file overwrite
  • Remote code execution (via overwriting configuration files, cron jobs, or Python packages)
  • Data corruption
  • Denial of service

Affected Functionality

The vulnerability affects the MSG file partitioning functionality when process_attachments=True is enabled.

Vulnerability Details

The library does not sanitize attachment filenames in MSG files before using them in file write operations, allowing directory
traversal sequences to escape the intended output directory.

Workarounds

Until patched, users can:

  • Set process_attachments=False when processing untrusted MSG files
  • Avoid processing MSG files from untrusted sources
  • Implement additional filename validation before processing

Severity

  • CVSS Score: 9.8 / 10 (Critical)
  • Vector String: CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H

References

This data is provided by the GitHub Advisory Database (CC-BY 4.0).


unstructured: Server-Side Request Forgery in the URL-based partitioning

CVE-2026-71428 / GHSA-4mvj-m6j5-pmf7

More information

Details

Summary

Server-Side Request Forgery in unstructured. The url= argument of partition(), partition_html(), and partition_md() is fetched via requests.get() with no host validation. The response body is returned as Element text, so this is a full-read SSRF — attackers reach loopback admin APIs, internal HTTP services, and cloud metadata endpoints, and read the response.

unstructured is the de facto URL ingestion layer for LangChain UnstructuredURLLoader, LlamaIndex UnstructuredReader, Chainlit, and many agent frameworks — secure defaults must live in the library, not in every downstream caller.

Details

Three sinks, all in unstructured == 0.22.26 (verified on main at 199f255):

  • unstructured/partition/auto.py:303 — file_and_type_from_url(), reached via partition(url=…).
  • unstructured/partition/html/partition.py:160 — partition_html(url=…). Post-fetch Content-Type check runs after the request hits the target.
  • unstructured/partition/md.py:96 — partition_md(url=…). No timeout (SSRF + slow-loris DoS).

None of is_private, is_loopback, ipaddress, gethostbyname, or allow_redirects appear in any of the three files. Three exploitation paths apply: direct private-IP target; redirect bypass (allow_redirects=True default); DNS rebinding (TOCTOU, closeable only by socket-pinning). Affected since 0.4.7 (Feb 2023) — ~219 releases, no validation ever introduced.

PoC

Local-only. pip install unstructured==0.22.26 flask requests.

internal_server.py:

from flask import Flask, Response, jsonify
app = Flask(__name__)

@app.route("/imds")
def imds(): return jsonify({"AccessKeyId": "ASIA-FAKE", "SecretAccessKey": "FAKE/SECRET"})

@app.route("/internal.html")
def html(): return Response("<html><body><p>SK_LEAK_42</p></body></html>", mimetype="text/html")

@app.route("/redir")
def redir(): return Response("", 302, headers={"Location": "http://127.0.0.1:9999/imds"})

if __name__ == "__main__": app.run(host="127.0.0.1", port=9999)

exploit.py — uses the public top-level API:


##### Stub NLP helpers so the offline sandbox skips spaCy model download.
##### Does NOT affect the SSRF (which lives in the URL fetcher, before NLP).
import unstructured.nlp.tokenize as _tk, unstructured.partition.text_type as _tt
_tk.sent_tokenize = _tt.sent_tokenize = lambda t: [s for s in (t or "").split(". ") if s]
_tk.word_tokenize = _tt.word_tokenize = lambda t: (t or "").split()
_tk.pos_tag       = _tt.pos_tag       = lambda t: [(w, "NN") for w in (t or "").split()]

from unstructured.partition.auto import partition
L = "http://127.0.0.1:9999"

##### A: partition(url=...) leaks internal HTML body
assert "SK_LEAK_42" in "\n".join(str(e) for e in partition(url=f"{L}/internal.html", languages=["eng"]))

##### B: redirect bypass reaches simulated IMDS
assert "SecretAccessKey" in "\n".join(str(e) for e in partition(url=f"{L}/redir", languages=["eng"]))
print("PoC OK")

In production the attacker substitutes 169.254.169.254, metadata.google.internal, or any internal address.

Impact

Attacker capabilities:

  • Internal HTTP service read — loopback admin consoles, internal Elasticsearch/Redis/Consul/etcd HTTP fronts, Kubernetes API server, social/internal microservices. This is the most broadly exploitable capability and is unaffected by any cloud-side hardening.
  • Cloud instance metadata access — reads metadata services that respond to unauthenticated GETs: GCP (metadata.google.internal), Azure IMDS, Oracle Cloud, DigitalOcean, and EC2 instances still configured for IMDSv1 (which remains widely deployed in older accounts and in services that do not enforce IMDSv2-only). EC2 instances configured as IMDSv2-only are not exposed to direct credential theft via this SSRF, since IMDSv2 requires a PUT for token acquisition; the SSRF still reaches the endpoint for reconnaissance and surface-mapping.
  • Side-effecting GET endpoints — magic-link consumers, job triggers, link-preview generators reachable on internal networks.
  • Internal network reconnaissance — connection success/failure timing and error messages serve as a port and service scanner.

Severity

  • CVSS Score: 9.3 / 10 (Critical)
  • Vector String: CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:L/A:N

References

This data is provided by the GitHub Advisory Database (CC-BY 4.0).


Release Notes

Unstructured-IO/unstructured (unstructured)

v0.24.0

Compare Source

Enhancements
  • Centralize outbound URL fetching: partition, partition_html, and partition_md now route url= fetches through a single shared helper (unstructured/safe_http.py) instead of ad-hoc requests.get calls. The helper applies an http/https scheme allowlist, a hostname denylist with IDNA normalization, address validation performed at connect time, manual redirect handling with per-hop re-validation (dropping credential material on cross-origin hops), refusal of proxied requests, and a default (connect, read) timeout. Behavior change: fetches that resolve to non-routable, loopback, or link-local addresses are now rejected by default. Set UNSTRUCTURED_ALLOW_PRIVATE_URL=1 (or pass allow_private=True) to opt out for controlled local usage.

v0.23.1

Compare Source

Enhancements
  • Extract filled AcroForm field values as text: values typed into fillable PDF form fields live in widget annotations rather than the page content stream, so pdfminer's text pass missed them. They are now recovered for both the fast and hi_res strategies and emitted as elements alongside the content-stream text.
Fixes
  • Fix inferred/extracted layout merge skipping subregion removal for single-region pages: the rule that removes an inferred box overlapping an extracted region was gated on any(extracted_to_keep), which evaluated False when the only kept extracted region was at index 0. On single-region pages (e.g. a PDF whose only text is one filled form field) this left a duplicate element; the guard now checks the array size.

v0.23.0

Compare Source

Enhancements
  • Add enrichment_origins metadata field for per-attribute model provenance: ElementMetadata gains a serialized enrichment_origins field mapping a written attribute name (e.g. text, text_as_html, embeddings) to a list of records {"type", "provider", "model"}, in application order. Enrichment producers stamp which model wrote (or contributed to) each attribute; authoring enrichments overwrite the list while additive ones append, preserving the prior author. A new ConsolidationStrategy.DICT_LIST_UNIQUE merges these dicts across elements during chunking (union keys, concatenate then dedupe records, preserving first-seen order).

v0.22.32

Compare Source

Fixes
  • Recover text inside PDF figure overlays in hi_res: hi_res pdfminer extraction only pulled text from objects exposing get_text (e.g. LTTextBox), and extract_text_objects only collected LTTextLine. Text held as loose LTChars inside an LTFigure - for example text drawn into a figure/XObject overlay rather than the main content stream - was dropped from the output. hi_res now groups such loose characters into text lines, inserting spaces on wide inter-character gaps and skipping hidden (render mode 3) and rotated characters.

v0.22.31

Compare Source

Enhancements
  • Rename isolate_tables chunking option to isolate_table: the option added in 0.22.30 has been renamed for naming consistency. Callers passing isolate_tables= must update to isolate_table=.

v0.22.30

Compare Source

Enhancements
  • Toggle table isolation in chunking: Add isolate_tables to basic/title chunking options. Defaults to True (the post-#​4307 behavior: Table/TableChunk elements always staged alone). Set to False to allow tables to share pre-chunks with adjacent non-table elements and be combined by PreChunkCombiner.

v0.22.29

Compare Source

Fixes
  • Truncate text if it exceeds spacy limit: add a guard against calling spacy tokenizer with very long text. Now long texts are truncated to fit under the character limit.

v0.22.28

Compare Source

Fixes
  • Preserve mixed-content tail text in table HTML: HtmlTable compactification previously cleared every element's .tail, silently dropping real text between inline children. Pure-whitespace tails are still removed, but tails carrying content are now kept with internal whitespace collapsed.

v0.22.27

Compare Source

Fixes
  • Stop misclassifying multi-line JSON files as NDJSON: is_ndjson_processable previously returned True for any text starting with {, so .json and .ipynb files containing a single multi-line JSON object (e.g. Jupyter notebooks) were routed to partition_ndjson, which then crashed in its splitlines()-based parser.

v0.22.26

Compare Source

Enhancements
  • Add table_extraction_method field to ElementMetadata to track which algorithm produced a table (grid, tatr, vlm). Propagated from LayoutElement during PDF partitioning.

v0.22.23

Compare Source

Fixes
  • Preserve colspan/rowspan in first table chunk headers: HtmlTable compactification no longer strips colspan and rowspan attributes from table cells. Previously, the first TableChunk lost merged-cell structural information while continuation chunks retained it (via the source-HTML path used for repeated headers), yielding inconsistent header layout across a split table.

v0.22.22

Compare Source

Security
  • Replace PyPI opencv wheels with ffmpeg-free builds in Docker image: After uv sync, the Dockerfile now substitutes all PyPI opencv-python variants with a source-built opencv-contrib-python-headless wheel compiled with WITH_FFMPEG=OFF, eliminating 14 bundled ffmpeg CVEs. The contrib-headless variant is a strict superset of the cv2 API (core + contrib modules, no GUI) so a single wheel replaces opencv-python, opencv-python-headless, and opencv-contrib-python.

v0.22.21

Compare Source

Enhancements
  • Skip table chunking option: Add skip_table_chunking to basic/title chunking options. When True, Table elements are passed through unchanged without being split into TableChunk elements, regardless of their size. Defaults to False to preserve existing behavior.

v0.22.20

Compare Source

Enhancements
  • Auto-detect vertical text for rotated PDFs: Add detect_vertical field to PDFMinerConfig and auto-enable it when rendered pages have /Rotate metadata, so pdfminer groups rotated text into proper words instead of per-character regions

v0.22.18

Compare Source

Fixes
  • Make ingest-test-fixtures-update-pr CI job also update the markdown versions of the fixtures.
Enhancements
  • Add page number support to v1 HTML parser: The v1 HTML parser now reads data-page-number attributes from ancestor elements and includes the page number in element metadata, consistent with the v2 parser behavior.

v0.22.16

Compare Source

Enhancements
  • Formula markdown export (element_to_md / elements_to_md): New keyword-only formula_markdown_style ("auto", "display_math", "plain"; default "auto"). In "auto", display math ($$ ... $$) is used only when the text looks like notation (heuristic score) and contains no $/$$ (avoids breaking Markdown and noisy OCR captions). "display_math" wraps whenever safe (still falls back to plain if $ would corrupt fences). "plain" emits text only. Optional normalize_formula (default True) maps common Unicode operators to LaTeX-like tokens; normalize_formula stays before keyword-only options so positional encoding / no_group_by_page callers are unchanged. Unicode √ is never mapped to \\sqrt{}. Module constants: FORMULA_MARKDOWN_AUTO, FORMULA_MARKDOWN_DISPLAY_MATH, FORMULA_MARKDOWN_PLAIN.

v0.22.12

Compare Source

Fixes
  • Fix fast strategy silently skipping text in some PDFs: Certain PDF generators (e.g. Prince XML) embed font encoding data in a non-standard way that pdfminer.six does not handle, causing body text to be silently dropped while headings still extract correctly. Added a workaround that reads the embedded encoding data directly.

v0.22.10

Compare Source

Enhancements
  • Repeat table headers across continuation chunks: Add repeat_table_headers to basic/title chunking options and table chunking internals so leading header rows are detected once and carried forward when large tables spill across multiple chunks.

v0.22.6

Compare Source

Fixes
  • Self-contained script for version extraction in release CI

v0.21.5

Compare Source

Fixes
  • Lower the requirement for pdfminer.six to >=20251230

v0.21.2

Compare Source

Fixes
  • Self-install pinned spaCy model at runtime with SHA256 verification: Replace the en-core-web-sm direct URL dependency in pyproject.toml with the installer library. The spaCy model is now downloaded and installed on first use with hash verification, removing the need for [tool.uv.sources] and making the install more portable.

v0.21.1

Compare Source

  • Add Check for complex documents: Adds a check for complex documents to avoid pdfminer with a high ratio of vector objects

v0.21.0

Compare Source

Fixes
  • Replace NLTK with spaCy to remediate CVE-2025-14009: NLTK's downloader uses zipfile.extractall() without path validation, enabling RCE via malicious packages (CVSS 10.0, no patch available). spaCy models install as pip packages, eliminating the vulnerable downloader entirely.

v0.20.8

Compare Source

Fixes
  • downgrade wrapt so it is compatible with opentelemetry-instrumentation-httpx
  • resolve lock issue with windows and python 3.13

v0.20.6

Compare Source

Fixes
  • fix: remap parent id after hashing to preserve right reference

v0.20.2

Compare Source

Enhancements
  • Add automated PyPI publishing: new release.yml GitHub Actions workflow triggers on GitHub release, builds the package with uv build, publishes to PyPI via pypa/gh-action-pypi-publish, and uploads to Azure Artifacts via twine
  • Replace uv sync --frozen with uv sync --locked across all CI workflows, Dockerfile, and Makefile to fail fast on stale lockfiles
  • Add --no-sync to all uv run and uv build commands that follow a prior uv sync step to prevent implicit re-syncing

v0.18.32

Compare Source

Enhancements
  • put pdfium calls behind a thread lock

v0.18.31

Compare Source

Enhancements
  • Changed default DPI to 350
  • Add token-based chunking support: Added max_tokens, new_after_n_tokens, and tokenizer parameters to chunk_by_title() and chunk_elements() for chunking by token count instead of character count. Uses tiktoken for token counting. Install with pip install "unstructured[chunking-tokens]". (fixes #​4127)
Fixes
  • Resolved security vulnerabilities in base system dependencies
    Bumped dependencies to address the following CVEs:
    glibc & related (glibc, glibc-locale-posix, ld-linux, libcrypt1, posix-libc-utils, posix-libc-utils-bin): CVE-2026-0915, CVE-2026-0861, GHSA-5pf6-63v3-88hw, GHSA-xp56-6525-9chf
    pyasn1: GHSA-63vm-454h-vhhq
    py3-setuptools (Python 3.12/3.13): GHSA-58pv-8j8x-9vj2
    ffmpeg (via OpenCV): CVE-2025-9951, CVE-2025-1594, CVE-2023-6604, CVE-2023-49502, CVE-2023-6602, CVE-2023-6605, CVE-2025-0518, CVE-2023-6601, CVE-2025-22919, CVE-2023-50010, CVE-2023-50008, CVE-2024-31582, CVE-2025-59729, CVE-2025-59730, CVE-2023-50007
  • Fix Pandoc exitcode 97 during ODT conversion: Try with sandbox=True first, fallback without sandbox only if ALLOW_PANDOC_NO_SANDBOX=true env var is set (fixes #​3997)
  • Fix coordinates=True causing TypeError in hi_res PDF processing: Filter out coordinates and coordinate_system from kwargs before passing to add_element_metadata() to prevent conflict with explicit parameters (fixes #​4126)
  • Preserve line breaks in code blocks during chunking: <pre> elements now generate CodeSnippet elements instead of Text, and chunking preserves internal whitespace for code snippets. (fixes #​4095)

v0.18.27

Compare Source

Fixes
  • Comment no-ops in zoom_image (codeflash)
  • Fix an issue where elements with partially filled extracted text are marked as extracted
Enhancement
  • Optimize sentence_count (codeflash)
  • Optimize _PartitionerLoader._load_partitioner (codeflash)
  • Optimize detect_languages (codeflash)
  • Optimize contains_verb (codeflash)
  • Optimize get_bbox_thickness (codeflash)
  • Upgrade pdfminer-six to 2026010 to fix ~15-18% performance regression from eager f-string evaluation

v0.18.26

Compare Source

Fixes
  • Pin deltalake<1.3.0 to fix ARM64 Docker builds (1.3.0 missing Linux ARM64 wheels)

v0.18.24

Compare Source

Enhancement
  • Optimize OCRAgentTesseract.extract_word_from_hocr (codeflash)
Fixes
  • Security update: Bumped dependencies to address security vulnerabilities

v0.18.21

Compare Source

Enhancement
  • Update save_elements unit test to check crop box padding behavior
Features
Fixes
  • Update unstructured-inference to 1.1.2 to address CVEs

v0.18.20

Compare Source

Enhancement
  • Improve the VoyageAI integration
  • Add voyage-context-3 support
Features
Fixes

v0.18.18

Compare Source

Fixes
  • Prevent path traversal in email MSG attachment filenames Fixed a security vulnerability (GHSA-gm8q-m8mv-jj5m) where malicious attachment filenames containing path traversal sequences could write files outside the intended directory. The fix normalizes both Unix and Windows path separators before sanitizing filenames, preventing cross-platform path traversal attacks in partition_msg functions

v0.18.15

Compare Source

Enhancements
  • Speed up function ElementHtml._get_children_html by 234% (codeflash)
  • Speed up function group_broken_paragraphs by 30% (codeflash)
Features
Fixes
  • Bumped dependencies via pip-compile to address the crit CVE in:
    • deepdiff: 8.6.0 -> 8.6.1: GHSA-mw26-5g2v-hqw3

v0.18.14

Compare Source

Enhancements
  • Speed up function sentence_count by 59% (codeflash)
  • Speed up function check_for_nltk_package by 111% (codeflash)
  • Speed up function under_non_alpha_ratio by 76% (codeflash)
Features
Fixes
  • change short text language detection log to debug reduce warning level log spamming
    • Bumped dependencies via pip-compile to address the following CVEs:
      • Python 3.12/3.13: CVE-2025-8194, GHSA-v594-44hm-2j7p
      • glibc & related (glibc, glibc-locale-posix, ld-linux, libcrypt1): CVE-2025-8058, GHSA-8xjp-c72j-67q8
      • aiohttp: GHSA-9548-qrrj-x5pj
      • openjpeg: CVE-2025-54874
      • pypdf: GHSA-7hfw-26vp-jp8m
      • transformers: GHSA-9356-575x-2w9m
      • urllib3: GHSA-48p4-8xcf-vxj5

v0.18.13

Compare Source

Enhancements
Features
Fixes
  • Parse a wider variety of date formats in email headers The partition_email function is now more robust to non-standard date formats, including ISO-8601 dates with "Z" suffixes. This prevents ValueError exceptions when partitioning emails with these date formats.

v0.18.11

Compare Source

Enhancements
  • Standardized on charset-normalizer library for encoding detection Previously we had both chardet and charset-normalizer as dependencies. We are dropping chardet and only using charset-normalizer.
Features
  • Type-aware <input> mapping in HTML transformations Bare <input> elements are now classified by their type attribute (checkbox → Checkbox, radio → RadioButton, others → FormFieldValue).
Fixes
  • Recognize '|' as a delimiter csv parser will now recognize '|' as a delimiter in addition to ',' and ';'.

v0.18.9

Compare Source

Enhancements
Features
  • Convert elements to markdown for output Added function to convert elements to markdown format for easy viewing.
Fixes
  • Language detection nit* Handle empty text

v0.18.7

Compare Source

Enhancements
  • text_as_html for Table element now keeps both input and img tag's class attribute Previously in partition HTML any tag inside a table is stripped of its class attribute. Now this attribute is preserved for both input and img tag in the table element's metadata.text_as_html.
Features
  • Add language detection for PDFs Add document and element level language detection to PDFs.
Fixes

v0.18.6

Compare Source

Enhancements
Features
Fixes
  • Improved epub partition errors EPUB partition will now produce new type of error on unprocessable files.
  • Fix type for serialized TableChunks Use TableChunk for the string value of the field type when serializing elements of type TableChunk, rather than using the value Table.

v0.18.5

Compare Source

Enhancements
  • Bump dependencies and remove lingering Python 3.9 artifacts Cleaned up some references to 3.9 that were left When we dropped Python 3.9 support.
  • text_as_html for Table element now keeps img tag's class attribute Previously in partition HTML any tag inside a table is stripped of its class attribute. Now this attribute is preserved for img tag in the table element's metadata.text_as_html.
Features
Fixes
  • Improve markdown code block handling Code blocks in markdown were previously being processed as embedded code instead of plain text.

v0.18.3

Compare Source

Enhancements
Features
Fixes
  • Upgrade Pillow to 11.3.0 Addresses a high priority CVE

v0.18.2

Fixes
  • Constrain fonttools to >=4.60.2 to address CVE-2025-66034

v0.18.1

Enhancement
  • Flag extracted elements as such in the metadata for downstream use
Features
Fixes

Configuration

📅 Schedule: (UTC)

  • Branch creation
    • At any time (no schedule defined)
  • Automerge
    • At any time (no schedule defined)

🚦 Automerge: Disabled by config. Please merge this manually once you are satisfied.

♻ Rebasing: Whenever PR becomes conflicted, or you tick the rebase/retry checkbox.

🔕 Ignore: Close this PR and you won't be reminded about this update again.


  • If you want to rebase/retry this PR, check this box

This PR was generated by Mend Renovate. View the repository job log.

This PR contains the following updates: | Package | Change | [Age](https://docs.renovatebot.com/merge-confidence/) | [Adoption](https://docs.renovatebot.com/merge-confidence/) | [Passing](https://docs.renovatebot.com/merge-confidence/) | [Confidence](https://docs.renovatebot.com/merge-confidence/) | |---|---|---|---|---|---| | [unstructured](https://redirect.github.com/Unstructured-IO/unstructured) | `==0.17.2` → `==0.24.0` | ![age](https://developer.mend.io/api/mc/badges/age/pypi/unstructured/0.24.0?slim=true) | ![adoption](https://developer.mend.io/api/mc/badges/adoption/pypi/unstructured/0.24.0?slim=true) | ![passing](https://developer.mend.io/api/mc/badges/compatibility/pypi/unstructured/0.17.2/0.24.0?slim=true) | ![confidence](https://developer.mend.io/api/mc/badges/confidence/pypi/unstructured/0.17.2/0.24.0?slim=true) | --- ### Unstructured has Path Traversal via Malicious MSG Attachment that Allows Arbitrary File Write [CVE-2025-64712](https://nvd.nist.gov/vuln/detail/CVE-2025-64712) / [GHSA-gm8q-m8mv-jj5m](https://redirect.github.com/advisories/GHSA-gm8q-m8mv-jj5m) <details> <summary>More information</summary> #### Details A Path Traversal vulnerability in the `partition_msg` function allows an attacker to write or overwrite arbitrary files on the filesystem when processing malicious MSG files with attachments. ## Impact An attacker can craft a malicious .msg file with attachment filenames containing path traversal sequences (e.g., `../../../etc/cron.d/malicious`). When processed with `process_attachments=True`, the library writes the attachment to an attacker-controlled path, potentially leading to: - Arbitrary file overwrite - Remote code execution (via overwriting configuration files, cron jobs, or Python packages) - Data corruption - Denial of service ## Affected Functionality The vulnerability affects the MSG file partitioning functionality when `process_attachments=True` is enabled. ## Vulnerability Details The library does not sanitize attachment filenames in MSG files before using them in file write operations, allowing directory traversal sequences to escape the intended output directory. ## Workarounds Until patched, users can: - Set `process_attachments=False` when processing untrusted MSG files - Avoid processing MSG files from untrusted sources - Implement additional filename validation before processing #### Severity - CVSS Score: 9.8 / 10 (Critical) - Vector String: `CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H` #### References - [https://github.com/Unstructured-IO/unstructured/security/advisories/GHSA-gm8q-m8mv-jj5m](https://redirect.github.com/Unstructured-IO/unstructured/security/advisories/GHSA-gm8q-m8mv-jj5m) - [https://github.com/Unstructured-IO/unstructured/commit/b01d35b2373fd087d2e15162b9c021663c97155d](https://redirect.github.com/Unstructured-IO/unstructured/commit/b01d35b2373fd087d2e15162b9c021663c97155d) - [https://nvd.nist.gov/vuln/detail/CVE-2025-64712](https://nvd.nist.gov/vuln/detail/CVE-2025-64712) - [https://github.com/advisories/GHSA-gm8q-m8mv-jj5m](https://redirect.github.com/advisories/GHSA-gm8q-m8mv-jj5m) This data is provided by the [GitHub Advisory Database](https://redirect.github.com/advisories/GHSA-gm8q-m8mv-jj5m) ([CC-BY 4.0](https://redirect.github.com/github/advisory-database/blob/main/LICENSE.md)). </details> --- ### unstructured: Server-Side Request Forgery in the URL-based partitioning [CVE-2026-71428](https://nvd.nist.gov/vuln/detail/CVE-2026-71428) / [GHSA-4mvj-m6j5-pmf7](https://redirect.github.com/advisories/GHSA-4mvj-m6j5-pmf7) <details> <summary>More information</summary> #### Details ##### Summary Server-Side Request Forgery in `unstructured`. The `url=` argument of `partition()`, `partition_html()`, and `partition_md()` is fetched via `requests.get()` with no host validation. The response body is returned as `Element` text, so this is a **full-read SSRF** — attackers reach loopback admin APIs, internal HTTP services, and cloud metadata endpoints, and read the response. `unstructured` is the de facto URL ingestion layer for LangChain `UnstructuredURLLoader`, LlamaIndex `UnstructuredReader`, Chainlit, and many agent frameworks — secure defaults must live in the library, not in every downstream caller. ##### Details Three sinks, all in `unstructured == 0.22.26` (verified on `main` at `199f255`): - `unstructured/partition/auto.py:303` — `file_and_type_from_url()`, reached via `partition(url=…)`. - `unstructured/partition/html/partition.py:160` — `partition_html(url=…)`. Post-fetch `Content-Type` check runs after the request hits the target. - `unstructured/partition/md.py:96` — `partition_md(url=…)`. No timeout (SSRF + slow-loris DoS). None of `is_private`, `is_loopback`, `ipaddress`, `gethostbyname`, or `allow_redirects` appear in any of the three files. Three exploitation paths apply: direct private-IP target; redirect bypass (`allow_redirects=True` default); DNS rebinding (TOCTOU, closeable only by socket-pinning). Affected since `0.4.7` (Feb 2023) — ~219 releases, no validation ever introduced. ##### PoC Local-only. `pip install unstructured==0.22.26 flask requests`. `internal_server.py`: ```python from flask import Flask, Response, jsonify app = Flask(__name__) @app.route("/imds") def imds(): return jsonify({"AccessKeyId": "ASIA-FAKE", "SecretAccessKey": "FAKE/SECRET"}) @app.route("/internal.html") def html(): return Response("<html><body><p>SK_LEAK_42</p></body></html>", mimetype="text/html") @app.route("/redir") def redir(): return Response("", 302, headers={"Location": "http://127.0.0.1:9999/imds"}) if __name__ == "__main__": app.run(host="127.0.0.1", port=9999) ``` `exploit.py` — uses the public top-level API: ```python ##### Stub NLP helpers so the offline sandbox skips spaCy model download. ##### Does NOT affect the SSRF (which lives in the URL fetcher, before NLP). import unstructured.nlp.tokenize as _tk, unstructured.partition.text_type as _tt _tk.sent_tokenize = _tt.sent_tokenize = lambda t: [s for s in (t or "").split(". ") if s] _tk.word_tokenize = _tt.word_tokenize = lambda t: (t or "").split() _tk.pos_tag = _tt.pos_tag = lambda t: [(w, "NN") for w in (t or "").split()] from unstructured.partition.auto import partition L = "http://127.0.0.1:9999" ##### A: partition(url=...) leaks internal HTML body assert "SK_LEAK_42" in "\n".join(str(e) for e in partition(url=f"{L}/internal.html", languages=["eng"])) ##### B: redirect bypass reaches simulated IMDS assert "SecretAccessKey" in "\n".join(str(e) for e in partition(url=f"{L}/redir", languages=["eng"])) print("PoC OK") ``` In production the attacker substitutes `169.254.169.254`, `metadata.google.internal`, or any internal address. ##### Impact Attacker capabilities: - **Internal HTTP service read** — loopback admin consoles, internal Elasticsearch/Redis/Consul/etcd HTTP fronts, Kubernetes API server, social/internal microservices. This is the most broadly exploitable capability and is unaffected by any cloud-side hardening. - **Cloud instance metadata access** — reads metadata services that respond to unauthenticated GETs: GCP (`metadata.google.internal`), Azure IMDS, Oracle Cloud, DigitalOcean, and EC2 instances still configured for IMDSv1 (which remains widely deployed in older accounts and in services that do not enforce IMDSv2-only). EC2 instances configured as IMDSv2-only are not exposed to direct credential theft via this SSRF, since IMDSv2 requires a `PUT` for token acquisition; the SSRF still reaches the endpoint for reconnaissance and surface-mapping. - **Side-effecting GET endpoints** — magic-link consumers, job triggers, link-preview generators reachable on internal networks. - **Internal network reconnaissance** — connection success/failure timing and error messages serve as a port and service scanner. #### Severity - CVSS Score: 9.3 / 10 (Critical) - Vector String: `CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:L/A:N` #### References - [https://github.com/Unstructured-IO/unstructured/security/advisories/GHSA-4mvj-m6j5-pmf7](https://redirect.github.com/Unstructured-IO/unstructured/security/advisories/GHSA-4mvj-m6j5-pmf7) - [https://nvd.nist.gov/vuln/detail/CVE-2026-71428](https://nvd.nist.gov/vuln/detail/CVE-2026-71428) - [https://github.com/Unstructured-IO/unstructured/pull/4388](https://redirect.github.com/Unstructured-IO/unstructured/pull/4388) - [https://github.com/Unstructured-IO/unstructured/commit/445c95735c4045057f51f399bc04c657751923bd](https://redirect.github.com/Unstructured-IO/unstructured/commit/445c95735c4045057f51f399bc04c657751923bd) - [https://github.com/Unstructured-IO/unstructured/releases/tag/0.24.0](https://redirect.github.com/Unstructured-IO/unstructured/releases/tag/0.24.0) - [https://github.com/advisories/GHSA-4mvj-m6j5-pmf7](https://redirect.github.com/advisories/GHSA-4mvj-m6j5-pmf7) This data is provided by the [GitHub Advisory Database](https://redirect.github.com/advisories/GHSA-4mvj-m6j5-pmf7) ([CC-BY 4.0](https://redirect.github.com/github/advisory-database/blob/main/LICENSE.md)). </details> --- ### Release Notes <details> <summary>Unstructured-IO/unstructured (unstructured)</summary> ### [`v0.24.0`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#0240) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.23.1...0.24.0) ##### Enhancements - **Centralize outbound URL fetching**: `partition`, `partition_html`, and `partition_md` now route `url=` fetches through a single shared helper (`unstructured/safe_http.py`) instead of ad-hoc `requests.get` calls. The helper applies an `http`/`https` scheme allowlist, a hostname denylist with IDNA normalization, address validation performed at connect time, manual redirect handling with per-hop re-validation (dropping credential material on cross-origin hops), refusal of proxied requests, and a default `(connect, read)` timeout. **Behavior change:** fetches that resolve to non-routable, loopback, or link-local addresses are now rejected by default. Set `UNSTRUCTURED_ALLOW_PRIVATE_URL=1` (or pass `allow_private=True`) to opt out for controlled local usage. ### [`v0.23.1`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#0231) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.23.0...0.23.1) ##### Enhancements - **Extract filled AcroForm field values as text**: values typed into fillable PDF form fields live in widget annotations rather than the page content stream, so pdfminer's text pass missed them. They are now recovered for both the `fast` and `hi_res` strategies and emitted as elements alongside the content-stream text. ##### Fixes - **Fix inferred/extracted layout merge skipping subregion removal for single-region pages**: the rule that removes an inferred box overlapping an extracted region was gated on `any(extracted_to_keep)`, which evaluated `False` when the only kept extracted region was at index 0. On single-region pages (e.g. a PDF whose only text is one filled form field) this left a duplicate element; the guard now checks the array size. ### [`v0.23.0`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#0230) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.22.32...0.23.0) ##### Enhancements - **Add `enrichment_origins` metadata field for per-attribute model provenance**: `ElementMetadata` gains a serialized `enrichment_origins` field mapping a written attribute name (e.g. `text`, `text_as_html`, `embeddings`) to a list of records `{"type", "provider", "model"}`, in application order. Enrichment producers stamp which model wrote (or contributed to) each attribute; authoring enrichments overwrite the list while additive ones append, preserving the prior author. A new `ConsolidationStrategy.DICT_LIST_UNIQUE` merges these dicts across elements during chunking (union keys, concatenate then dedupe records, preserving first-seen order). ### [`v0.22.32`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#02232) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.22.31...0.22.32) ##### Fixes - **Recover text inside PDF figure overlays in hi\_res**: hi\_res pdfminer extraction only pulled text from objects exposing `get_text` (e.g. `LTTextBox`), and `extract_text_objects` only collected `LTTextLine`. Text held as loose `LTChar`s inside an `LTFigure` - for example text drawn into a figure/XObject overlay rather than the main content stream - was dropped from the output. hi\_res now groups such loose characters into text lines, inserting spaces on wide inter-character gaps and skipping hidden (render mode 3) and rotated characters. ### [`v0.22.31`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#02231) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.22.30...0.22.31) ##### Enhancements - **Rename `isolate_tables` chunking option to `isolate_table`**: the option added in 0.22.30 has been renamed for naming consistency. Callers passing `isolate_tables=` must update to `isolate_table=`. ### [`v0.22.30`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#02230) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.22.29...0.22.30) ##### Enhancements - **Toggle table isolation in chunking**: Add `isolate_tables` to basic/title chunking options. Defaults to `True` (the [post-#&#8203;4307](https://redirect.github.com/post-/unstructured/issues/4307) behavior: `Table`/`TableChunk` elements always staged alone). Set to `False` to allow tables to share pre-chunks with adjacent non-table elements and be combined by `PreChunkCombiner`. ### [`v0.22.29`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#02229) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.22.28...0.22.29) ##### Fixes - **Truncate text if it exceeds `spacy` limit**: add a guard against calling `spacy` tokenizer with very long text. Now long texts are truncated to fit under the character limit. ### [`v0.22.28`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#02228) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.22.27...0.22.28) ##### Fixes - **Preserve mixed-content tail text in table HTML**: `HtmlTable` compactification previously cleared every element's `.tail`, silently dropping real text between inline children. Pure-whitespace tails are still removed, but tails carrying content are now kept with internal whitespace collapsed. ### [`v0.22.27`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#02227) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.22.26...0.22.27) ##### Fixes - **Stop misclassifying multi-line JSON files as NDJSON**: `is_ndjson_processable` previously returned `True` for any text starting with `{`, so `.json` and `.ipynb` files containing a single multi-line JSON object (e.g. Jupyter notebooks) were routed to `partition_ndjson`, which then crashed in its `splitlines()`-based parser. ### [`v0.22.26`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#02226) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.22.23...0.22.26) ##### Enhancements - Add `table_extraction_method` field to `ElementMetadata` to track which algorithm produced a table (grid, tatr, vlm). Propagated from `LayoutElement` during PDF partitioning. ### [`v0.22.23`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#02223) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.22.22...0.22.23) ##### Fixes - **Preserve `colspan`/`rowspan` in first table chunk headers**: `HtmlTable` compactification no longer strips `colspan` and `rowspan` attributes from table cells. Previously, the first `TableChunk` lost merged-cell structural information while continuation chunks retained it (via the source-HTML path used for repeated headers), yielding inconsistent header layout across a split table. ### [`v0.22.22`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#02222) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.22.21...0.22.22) ##### Security - **Replace PyPI opencv wheels with ffmpeg-free builds in Docker image**: After `uv sync`, the Dockerfile now substitutes all PyPI opencv-python variants with a source-built `opencv-contrib-python-headless` wheel compiled with `WITH_FFMPEG=OFF`, eliminating 14 bundled ffmpeg CVEs. The contrib-headless variant is a strict superset of the cv2 API (core + contrib modules, no GUI) so a single wheel replaces `opencv-python`, `opencv-python-headless`, and `opencv-contrib-python`. ### [`v0.22.21`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#02221) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.22.20...0.22.21) ##### Enhancements - **Skip table chunking option**: Add `skip_table_chunking` to basic/title chunking options. When `True`, `Table` elements are passed through unchanged without being split into `TableChunk` elements, regardless of their size. Defaults to `False` to preserve existing behavior. ### [`v0.22.20`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#02220) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.22.18...0.22.20) ##### Enhancements - **Auto-detect vertical text for rotated PDFs**: Add `detect_vertical` field to `PDFMinerConfig` and auto-enable it when rendered pages have `/Rotate` metadata, so pdfminer groups rotated text into proper words instead of per-character regions ### [`v0.22.18`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#02218) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.22.16...0.22.18) ##### Fixes - Make `ingest-test-fixtures-update-pr` CI job also update the markdown versions of the fixtures. ##### Enhancements - **Add page number support to v1 HTML parser**: The v1 HTML parser now reads `data-page-number` attributes from ancestor elements and includes the page number in element metadata, consistent with the v2 parser behavior. ### [`v0.22.16`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#02216) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.22.12...0.22.16) ##### Enhancements - **Formula markdown export (`element_to_md` / `elements_to_md`)**: New keyword-only `formula_markdown_style` (`"auto"`, `"display_math"`, `"plain"`; default `"auto"`). In `"auto"`, display math (`$$ ... $$`) is used only when the text looks like notation (heuristic score) and contains no `$`/`$$` (avoids breaking Markdown and noisy OCR captions). `"display_math"` wraps whenever safe (still falls back to plain if `$` would corrupt fences). `"plain"` emits text only. Optional `normalize_formula` (default `True`) maps common Unicode operators to LaTeX-like tokens; `normalize_formula` stays before keyword-only options so positional `encoding` / `no_group_by_page` callers are unchanged. Unicode `√` is never mapped to `\\sqrt{}`. Module constants: `FORMULA_MARKDOWN_AUTO`, `FORMULA_MARKDOWN_DISPLAY_MATH`, `FORMULA_MARKDOWN_PLAIN`. ### [`v0.22.12`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#02212) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.22.10...0.22.12) ##### Fixes - **Fix fast strategy silently skipping text in some PDFs**: Certain PDF generators (e.g. Prince XML) embed font encoding data in a non-standard way that pdfminer.six does not handle, causing body text to be silently dropped while headings still extract correctly. Added a workaround that reads the embedded encoding data directly. ### [`v0.22.10`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#02210) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.22.6...0.22.10) ##### Enhancements - **Repeat table headers across continuation chunks**: Add `repeat_table_headers` to basic/title chunking options and table chunking internals so leading header rows are detected once and carried forward when large tables spill across multiple chunks. ### [`v0.22.6`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#0226) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.21.5...0.22.6) ##### Fixes - Self-contained script for version extraction in release CI ### [`v0.21.5`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#0215) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.21.2...0.21.5) ##### Fixes - Lower the requirement for `pdfminer.six` to `>=20251230` ### [`v0.21.2`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#0212) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.21.1...0.21.2) ##### Fixes - **Self-install pinned spaCy model at runtime with SHA256 verification**: Replace the `en-core-web-sm` direct URL dependency in `pyproject.toml` with the `installer` library. The spaCy model is now downloaded and installed on first use with hash verification, removing the need for `[tool.uv.sources]` and making the install more portable. ### [`v0.21.1`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#02112) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.21.0...0.21.1) - **Add Check for complex documents**: Adds a check for complex documents to avoid pdfminer with a high ratio of vector objects ### [`v0.21.0`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#0210) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.20.8...0.21.0) ##### Fixes - **Replace NLTK with spaCy to remediate CVE-2025-14009**: NLTK's downloader uses `zipfile.extractall()` without path validation, enabling RCE via malicious packages (CVSS 10.0, no patch available). spaCy models install as pip packages, eliminating the vulnerable downloader entirely. ### [`v0.20.8`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#0208) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.20.6...0.20.8) ##### Fixes - downgrade `wrapt` so it is compatible with `opentelemetry-instrumentation-httpx` - resolve lock issue with windows and python 3.13 ### [`v0.20.6`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#0206) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.20.2...0.20.6) ##### Fixes - fix: remap parent id after hashing to preserve right reference ### [`v0.20.2`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#0202) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.18.32...0.20.2) ##### Enhancements - Add automated PyPI publishing: new `release.yml` GitHub Actions workflow triggers on GitHub release, builds the package with `uv build`, publishes to PyPI via `pypa/gh-action-pypi-publish`, and uploads to Azure Artifacts via `twine` - Replace `uv sync --frozen` with `uv sync --locked` across all CI workflows, Dockerfile, and Makefile to fail fast on stale lockfiles - Add `--no-sync` to all `uv run` and `uv build` commands that follow a prior `uv sync` step to prevent implicit re-syncing ### [`v0.18.32`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#01832) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.18.31...0.18.32) ##### Enhancements - put `pdfium` calls behind a thread lock ### [`v0.18.31`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#01831) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.18.27...0.18.31) ##### Enhancements - Changed default DPI to 350 - **Add token-based chunking support**: Added `max_tokens`, `new_after_n_tokens`, and `tokenizer` parameters to `chunk_by_title()` and `chunk_elements()` for chunking by token count instead of character count. Uses tiktoken for token counting. Install with `pip install "unstructured[chunking-tokens]"`. (fixes [#&#8203;4127](https://redirect.github.com/Unstructured-IO/unstructured/issues/4127)) ##### Fixes - Resolved security vulnerabilities in base system dependencies Bumped dependencies to address the following CVEs: **glibc & related (glibc, glibc-locale-posix, ld-linux, libcrypt1, posix-libc-utils, posix-libc-utils-bin)**: CVE-2026-0915, CVE-2026-0861, GHSA-5pf6-63v3-88hw, GHSA-xp56-6525-9chf **pyasn1**: GHSA-63vm-454h-vhhq **py3-setuptools** (Python 3.12/3.13): GHSA-58pv-8j8x-9vj2 **ffmpeg (via OpenCV)**: CVE-2025-9951, CVE-2025-1594, CVE-2023-6604, CVE-2023-49502, CVE-2023-6602, CVE-2023-6605, CVE-2025-0518, CVE-2023-6601, CVE-2025-22919, CVE-2023-50010, CVE-2023-50008, CVE-2024-31582, CVE-2025-59729, CVE-2025-59730, CVE-2023-50007 - **Fix Pandoc exitcode 97 during ODT conversion**: Try with sandbox=True first, fallback without sandbox only if `ALLOW_PANDOC_NO_SANDBOX=true` env var is set (fixes [#&#8203;3997](https://redirect.github.com/Unstructured-IO/unstructured/issues/3997)) - **Fix `coordinates=True` causing TypeError in hi\_res PDF processing**: Filter out `coordinates` and `coordinate_system` from kwargs before passing to `add_element_metadata()` to prevent conflict with explicit parameters (fixes [#&#8203;4126](https://redirect.github.com/Unstructured-IO/unstructured/issues/4126)) - **Preserve line breaks in code blocks during chunking**: `<pre>` elements now generate `CodeSnippet` elements instead of `Text`, and chunking preserves internal whitespace for code snippets. (fixes [#&#8203;4095](https://redirect.github.com/Unstructured-IO/unstructured/issues/4095)) ### [`v0.18.27`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#01827) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.18.26...0.18.27) ##### Fixes - Comment no-ops in `zoom_image` (codeflash) - Fix an issue where elements with partially filled extracted text are marked as extracted ##### Enhancement - Optimize `sentence_count` (codeflash) - Optimize `_PartitionerLoader._load_partitioner` (codeflash) - Optimize `detect_languages` (codeflash) - Optimize `contains_verb` (codeflash) - Optimize `get_bbox_thickness` (codeflash) - Upgrade pdfminer-six to [`2026010`](https://redirect.github.com/Unstructured-IO/unstructured/commit/20260107) to fix \~15-18% performance regression from eager f-string evaluation ### [`v0.18.26`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#01826) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.18.24...0.18.26) ##### Fixes - Pin `deltalake<1.3.0` to fix ARM64 Docker builds (1.3.0 missing Linux ARM64 wheels) ### [`v0.18.24`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#01824) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.18.21...0.18.24) ##### Enhancement - Optimize `OCRAgentTesseract.extract_word_from_hocr` (codeflash) ##### Fixes - **Security update**: Bumped dependencies to address security vulnerabilities ### [`v0.18.21`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#01821) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.18.20...0.18.21) ##### Enhancement - Update save\_elements unit test to check crop box padding behavior ##### Features ##### Fixes - **Update `unstructured-inference`** to 1.1.2 to address CVEs ### [`v0.18.20`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#01820) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.18.18...0.18.20) ##### Enhancement - Improve the VoyageAI integration - Add voyage-context-3 support ##### Features ##### Fixes ### [`v0.18.18`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#01818) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.18.15...0.18.18) ##### Fixes - **Prevent path traversal in email MSG attachment filenames** Fixed a security vulnerability (GHSA-gm8q-m8mv-jj5m) where malicious attachment filenames containing path traversal sequences could write files outside the intended directory. The fix normalizes both Unix and Windows path separators before sanitizing filenames, preventing cross-platform path traversal attacks in `partition_msg` functions ### [`v0.18.15`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#01815) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/unstructured_0.18.14...0.18.15) ##### Enhancements - Speed up function ElementHtml.\_get\_children\_html by 234% (codeflash) - Speed up function group\_broken\_paragraphs by 30% (codeflash) ##### Features ##### Fixes - Bumped dependencies via pip-compile to address the crit CVE in: - deepdiff: 8.6.0 -> 8.6.1: GHSA-mw26-5g2v-hqw3 ### [`v0.18.14`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#01814) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.18.13...unstructured_0.18.14) ##### Enhancements - Speed up function sentence\_count by 59% (codeflash) - Speed up function `check_for_nltk_package` by 111% (codeflash) - Speed up function `under_non_alpha_ratio` by 76% (codeflash) ##### Features ##### Fixes - **change short text language detection log to debug** reduce warning level log spamming - Bumped dependencies via pip-compile to address the following CVEs: - **Python 3.12/3.13**: CVE-2025-8194, GHSA-v594-44hm-2j7p - **glibc & related (glibc, glibc-locale-posix, ld-linux, libcrypt1)**: CVE-2025-8058, GHSA-8xjp-c72j-67q8 - **aiohttp**: GHSA-9548-qrrj-x5pj - **openjpeg**: CVE-2025-54874 - **pypdf**: GHSA-7hfw-26vp-jp8m - **transformers**: GHSA-9356-575x-2w9m - **urllib3**: GHSA-48p4-8xcf-vxj5 ### [`v0.18.13`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#01813) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.18.11...0.18.13) ##### Enhancements ##### Features ##### Fixes - **Parse a wider variety of date formats in email headers** The `partition_email` function is now more robust to non-standard date formats, including ISO-8601 dates with "Z" suffixes. This prevents `ValueError` exceptions when partitioning emails with these date formats. ### [`v0.18.11`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#01811) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.18.9...0.18.11) ##### Enhancements - **Standardized on `charset-normalizer` library for encoding detection** Previously we had both `chardet` and `charset-normalizer` as dependencies. We are dropping `chardet` and only using `charset-normalizer`. ##### Features - **Type-aware `<input>` mapping in HTML transformations** Bare `<input>` elements are now classified by their `type` attribute (checkbox → Checkbox, radio → RadioButton, others → FormFieldValue). ##### Fixes - **Recognize '|' as a delimiter** csv parser will now recognize '|' as a delimiter in addition to ',' and ';'. ### [`v0.18.9`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#0189) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.18.7...0.18.9) ##### Enhancements ##### Features - **Convert elements to markdown for output** Added function to convert elements to markdown format for easy viewing. ##### Fixes - *Language detection nit*\* Handle empty text ### [`v0.18.7`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#0187) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.18.6...0.18.7) ##### Enhancements - **`text_as_html` for Table element now keeps both `input` and `img` tag's `class` attribute** Previously in partition HTML any tag inside a table is stripped of its `class` attribute. Now this attribute is preserved for both `input` and `img` tag in the table element's `metadata.text_as_html`. ##### Features - **Add language detection for PDFs** Add document and element level language detection to PDFs. ##### Fixes ### [`v0.18.6`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#0186) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.18.5...0.18.6) ##### Enhancements ##### Features ##### Fixes - **Improved epub partition errors** EPUB partition will now produce new type of error on unprocessable files. - **Fix type for serialized TableChunks** Use `TableChunk` for the string value of the field `type` when serializing elements of type `TableChunk`, rather than using the value `Table`. ### [`v0.18.5`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#0185) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.18.3...0.18.5) ##### Enhancements - **Bump dependencies and remove lingering Python 3.9 artifacts** Cleaned up some references to 3.9 that were left When we dropped Python 3.9 support. - **`text_as_html` for Table element now keeps `img` tag's `class` attribute** Previously in partition HTML any tag inside a table is stripped of its `class` attribute. Now this attribute is preserved for `img` tag in the table element's `metadata.text_as_html`. ##### Features ##### Fixes - **Improve markdown code block handling** Code blocks in markdown were previously being processed as embedded code instead of plain text. ### [`v0.18.3`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#0183) [Compare Source](https://redirect.github.com/Unstructured-IO/unstructured/compare/0.18.2...0.18.3) ##### Enhancements ##### Features ##### Fixes - **Upgrade Pillow to 11.3.0** Addresses a high priority CVE ### [`v0.18.2`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#01822) ##### Fixes - Constrain fonttools to >=4.60.2 to address CVE-2025-66034 ### [`v0.18.1`](https://redirect.github.com/Unstructured-IO/unstructured/blob/HEAD/CHANGELOG.md#01819) ##### Enhancement - Flag extracted elements as such in the metadata for downstream use ##### Features ##### Fixes </details> --- ### Configuration 📅 **Schedule**: (UTC) - Branch creation - At any time (no schedule defined) - Automerge - At any time (no schedule defined) 🚦 **Automerge**: Disabled by config. Please merge this manually once you are satisfied. ♻ **Rebasing**: Whenever PR becomes conflicted, or you tick the rebase/retry checkbox. 🔕 **Ignore**: Close this PR and you won't be reminded about this update again. --- - [ ] <!-- rebase-check -->If you want to rebase/retry this PR, check this box --- This PR was generated by [Mend Renovate](https://mend.io/renovate/). View the [repository job log](https://developer.mend.io/github/sarumaj/rag-agent). <!--renovate-debug:eyJjcmVhdGVkSW5WZXIiOiI0My4yMTkuMCIsInVwZGF0ZWRJblZlciI6IjQ0LjYxLjMiLCJ0YXJnZXRCcmFuY2giOiJtYWluIiwibGFiZWxzIjpbXX0=-->
This pull request can be merged automatically.
You are not authorized to merge this pull request.
View command line instructions

Checkout

From your project repository, check out a new branch and test the changes.
git fetch -u origin renovate/pypi-unstructured-vulnerability:renovate/pypi-unstructured-vulnerability
git switch renovate/pypi-unstructured-vulnerability

Merge

Merge the changes and update on Forgejo.

Warning: The "Autodetect manual merge" setting is not enabled for this repository, you will have to mark this pull request as manually merged afterwards.

git switch main
git merge --no-ff renovate/pypi-unstructured-vulnerability
git switch renovate/pypi-unstructured-vulnerability
git rebase main
git switch main
git merge --ff-only renovate/pypi-unstructured-vulnerability
git switch renovate/pypi-unstructured-vulnerability
git rebase main
git switch main
git merge --no-ff renovate/pypi-unstructured-vulnerability
git switch main
git merge --squash renovate/pypi-unstructured-vulnerability
git switch main
git merge --ff-only renovate/pypi-unstructured-vulnerability
git switch main
git merge renovate/pypi-unstructured-vulnerability
git push origin main
Sign in to join this conversation.
No description provided.