unstructured
74e2fea0 - fix(xlsx): reject worksheets that span too many cells (#4515)

Commit
4 days ago
fix(xlsx): reject worksheets that span too many cells (#4515) ## Summary `partition_xlsx()` reads each worksheet into a dense data-frame whose shape is set by the farthest populated cell. A small XLSX or XLS file with one far-away cell can therefore span millions of cells. Subtable detection (`find_subtable=True`, the default) then built a networkx grid graph with one node per cell, about 1 KB each, so a file of a few KB could use tens of GB. - **Size check before reading.** Before `pd.read_excel`, each worksheet's span is measured without building any cells: - XLSX is streamed with openpyxl in read-only mode, ignoring the stored `<dimension>`, with the same trailing-empty trimming pandas applies. - XLS sheets are loaded one at a time with xlrd (`on_demand`, `ragged_rows`). If the total across all worksheets exceeds `XLSX_MAX_CELLS` (new env setting, default 5,000,000), it raises `UnprocessableEntityError`. The XLSX scan stops as soon as the total passes the limit. - **Explicit engine.** The pandas engine (`openpyxl` / `xlrd`) is chosen from the file signature and passed explicitly, so pandas can't fall back to a reader the check didn't measure. Files that are neither are rejected as invalid. - **Sparse graph.** The subtable-detection graph is built from populated cells only, using numpy adjacency. It produces the same components as before, and sparse sheets under the limit get much cheaper. ## Measurements (synthetic files: a 20×10 table plus one far-away cell) | File | Before | After | |---|---|---| | 5.7 KB `.xlsx`, sheet spans 2000×2000 | 4.4 GiB peak RSS, 18s | 445 MiB, 0.6s | | 5.7 KB `.xlsx`, sheet spans 100,000×10 | 1.2 GiB | 331 MiB | | `.xlsx` with a cell at row 1,048,576 | runs out of memory while pandas reads it | rejected before reading | | 11 KB `.xls`, two sheets spanning 5000×200 | 1.3 GiB | 349 MiB | `find_subtable=False` still costs about 350 B per cell (HTML render plus parse). The limit bounds that path to roughly 1.75 GB. ## Behaviour change A zip-based file that isn't XLSX (e.g. `.ods`) passed to `partition_xlsx` now always goes to openpyxl. Before, pandas could pick another installed reader. ## Tests - A stray far cell is rejected, and a mock asserts `pd.read_excel` is never called. - A sparse sheet under the limit still partitions correctly. - The limit is summed across worksheets. - The limit applies to `.xls`. - The measured shapes equal the shapes pandas reads, for example files and generated sparse worksheets. - Engine detection works. - On 20 random sheets, the sparse graph finds the same components as the old grid-graph build. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated description by cubic. --> <a href="https://cubic.dev/pr/Unstructured-IO/unstructured/pull/4515?utm_source=github" target="_blank" rel="noopener noreferrer" data-no-image-dialog="true"><picture><source media="(prefers-color-scheme: dark)" srcset="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"><source media="(prefers-color-scheme: light)" srcset="https://www.cubic.dev/buttons/review-in-cubic-light.svg"><img alt="Review in cubic" src="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"></picture></a> <!-- End of auto-generated description by cubic. --> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Author
Parents
Loading