fix(xlsx): reject worksheets that span too many cells (#4515)
## Summary
`partition_xlsx()` reads each worksheet into a dense data-frame whose
shape is set by the farthest populated cell. A small XLSX or XLS file
with one far-away cell can therefore span millions of cells. Subtable
detection (`find_subtable=True`, the default) then built a networkx grid
graph with one node per cell, about 1 KB each, so a file of a few KB
could use tens of GB.
- **Size check before reading.** Before `pd.read_excel`, each
worksheet's span is measured without building any cells:
- XLSX is streamed with openpyxl in read-only mode, ignoring the stored
`<dimension>`, with the same trailing-empty trimming pandas applies.
- XLS sheets are loaded one at a time with xlrd (`on_demand`,
`ragged_rows`).
If the total across all worksheets exceeds `XLSX_MAX_CELLS` (new env
setting, default 5,000,000), it raises `UnprocessableEntityError`. The
XLSX scan stops as soon as the total passes the limit.
- **Explicit engine.** The pandas engine (`openpyxl` / `xlrd`) is chosen
from the file signature and passed explicitly, so pandas can't fall back
to a reader the check didn't measure. Files that are neither are
rejected as invalid.
- **Sparse graph.** The subtable-detection graph is built from populated
cells only, using numpy adjacency. It produces the same components as
before, and sparse sheets under the limit get much cheaper.
## Measurements (synthetic files: a 20×10 table plus one far-away cell)
| File | Before | After |
|---|---|---|
| 5.7 KB `.xlsx`, sheet spans 2000×2000 | 4.4 GiB peak RSS, 18s | 445
MiB, 0.6s |
| 5.7 KB `.xlsx`, sheet spans 100,000×10 | 1.2 GiB | 331 MiB |
| `.xlsx` with a cell at row 1,048,576 | runs out of memory while pandas
reads it | rejected before reading |
| 11 KB `.xls`, two sheets spanning 5000×200 | 1.3 GiB | 349 MiB |
`find_subtable=False` still costs about 350 B per cell (HTML render plus
parse). The limit bounds that path to roughly 1.75 GB.
## Behaviour change
A zip-based file that isn't XLSX (e.g. `.ods`) passed to
`partition_xlsx` now always goes to openpyxl. Before, pandas could pick
another installed reader.
## Tests
- A stray far cell is rejected, and a mock asserts `pd.read_excel` is
never called.
- A sparse sheet under the limit still partitions correctly.
- The limit is summed across worksheets.
- The limit applies to `.xls`.
- The measured shapes equal the shapes pandas reads, for example files
and generated sparse worksheets.
- Engine detection works.
- On 20 random sheets, the sparse graph finds the same components as the
old grid-graph build.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated description by cubic. -->
<a
href="https://cubic.dev/pr/Unstructured-IO/unstructured/pull/4515?utm_source=github"
target="_blank" rel="noopener noreferrer"
data-no-image-dialog="true"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"><source
media="(prefers-color-scheme: light)"
srcset="https://www.cubic.dev/buttons/review-in-cubic-light.svg"><img
alt="Review in cubic"
src="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"></picture></a>
<!-- End of auto-generated description by cubic. -->
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>