unstructured
15c6302f - fix(csv): reject CSV and TSV files that span too many cells (#4518)

Commit
3 days ago
fix(csv): reject CSV and TSV files that span too many cells (#4518) ## Summary pandas sizes the data-frame by the first record and pads every shorter record out to that width. A small CSV or TSV whose first line is a long run of delimiters can therefore span millions of cells. `partition_csv()` and `partition_tsv()` render every cell to HTML and parse it back, at about 400 B per cell. - **Check before reading.** Before pandas reads the file, `check_cell_count()` streams it in fixed-size chunks and measures its span, stopping as soon as it passes `CSV_MAX_CELLS` (new env setting, default 5,000,000), and raises `UnprocessableEntityError`. - **Columns** are the fields of the first non-blank record, counted with pandas' quoting rules and each engine's blank-line rule, without building the fields. - With a header, pandas makes the fields a data row has beyond the header into implicit-index columns. Its C engine takes their number from the first data row, and its Python engine also from the second, which can make the first data row the index names. So the widest of the header and the first two data records is counted; any later, wider row is an error in pandas. - **Rows** are a conservative upper bound: every line ending after the first record counts, including blank lines and newlines inside quoted fields, so the check can over-count but never under-count. - **What it bounds:** this is a conservative estimate of the cells pandas will build, not a bound on memory or bytes read. Blank lines held while sniffing, and a stream spooled because it can't seek, still take space in proportion to the input. - **pandas row explosion.** pandas 2.x's C tokenizer reads a lone `\r` line ending followed by a whitespace-only line as 2^18 empty rows, so `b",\n\r ,"` (5 bytes) becomes 262,145 rows inside `pd.read_csv`, before any check can run (fixed in pandas 3.1; the project pins `<3`). pandas now reads the file through a streaming reader that converts each lone `\r` to `\n`. `\r\n` is left alone, so text inside quoted fields of CRLF files is unchanged. - **One sniffer.** When the delimiter is left to pandas (the context's restricted sniffer found none), it is sniffed once, by us, from the first non-blank line, and passed to pandas explicitly, so the check and the read agree. A file with no usable delimiter is read as one column, split on the ASCII unit separator. - **Inputs.** Leading byte-order marks are dropped for both the check and the read. `partition_tsv()` still decompresses a compressed filename (e.g. `.tsv.gz`), with the same content going to the check and the read, and copies a stream that cannot seek to a temporary file once. ## Measurements | File | Before | After | |---|---|---| | 3 KB CSV, 1000-field first line + 1000 one-field lines (1M cells) | 478 MiB, 5s | unchanged (under the limit) | | 15 KB CSV, 5000 × 5000 | 8.7 GiB, 125s | rejected in 0.1s | | 3 MB CSV, 1M-field first line + 1M lines | not run (would need ~10^12 cells) | rejected in 0.2s | | 5 bytes, `b",\n\r ,"` | 262,145 rows | 2 rows | | 80 bytes, that pattern repeated 16× | over 768 MB inside `pd.read_csv` | 2 rows per repeat | The check was fuzzed against pandas' actual output in a memory-capped container. About 27,000 generated inputs (quotes, CR/LF/CRLF, form-feeds, BOMs, leading blank lines, 3-byte chunks) gave 0 undercounts and 0 parser blowups. A further 10,000 inputs read with a header gave 13,727 comparisons, counting the header row and implicit-index columns, also with 0 undercounts. ## Behaviour changes - A dense CSV/TSV spanning more than 5M cells (roughly 10 MB or more) is now rejected. Set `CSV_MAX_CELLS` higher to allow it. - A file with no usable delimiter, where pandas' own sniff used to raise `csv.Error`, is now read as one column. ## Tests - A wide first line is rejected before pandas reads the file, for CSV and for TSV from both a path and a file. - Exact-limit passes and one-under failures. - The record pandas sizes the frame by: blank and whitespace-only first lines, quoted newlines, form-feeds, mid-field quotes, BOMs, and each engine's blank-line rule. - The carriage-return explosion. - Sniffing past blank lines and the one-column fallback. - Implicit-index columns: a TSV whose 1-field header is followed by a 100-field row and a ragged tail is rejected before pandas reads it, with output matching pandas within the limit; the Python engine's index-names layout is also covered. - Streaming: fixed-size reads; linear copying while replaying a 200,000-line blank prefix; output that doesn't depend on where chunks split multibyte characters, BOMs, quotes or CRLF; and stopping reading once the limit is passed. - TSV: a field over the `csv` module's 128 KiB limit, a `.tsv.gz` filename, and a stream that cannot seek. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated description by cubic. --> <a href="https://cubic.dev/pr/Unstructured-IO/unstructured/pull/4518?utm_source=github" target="_blank" rel="noopener noreferrer" data-no-image-dialog="true"><picture><source media="(prefers-color-scheme: dark)" srcset="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"><source media="(prefers-color-scheme: light)" srcset="https://www.cubic.dev/buttons/review-in-cubic-light.svg"><img alt="Review in cubic" src="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"></picture></a> <!-- End of auto-generated description by cubic. --> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Author
Parents
Loading