fix(csv): reject CSV and TSV files that span too many cells (#4518)
## Summary
pandas sizes the data-frame by the first record and pads every shorter
record out to that width. A small CSV or TSV whose first line is a long
run of delimiters can therefore span millions of cells.
`partition_csv()` and `partition_tsv()` render every cell to HTML and
parse it back, at about 400 B per cell.
- **Check before reading.** Before pandas reads the file,
`check_cell_count()` streams it in fixed-size chunks and measures its
span, stopping as soon as it passes `CSV_MAX_CELLS` (new env setting,
default 5,000,000), and raises `UnprocessableEntityError`.
- **Columns** are the fields of the first non-blank record, counted with
pandas' quoting rules and each engine's blank-line rule, without
building the fields.
- With a header, pandas makes the fields a data row has beyond the
header into implicit-index columns. Its C engine takes their number from
the first data row, and its Python engine also from the second, which
can make the first data row the index names. So the widest of the header
and the first two data records is counted; any later, wider row is an
error in pandas.
- **Rows** are a conservative upper bound: every line ending after the
first record counts, including blank lines and newlines inside quoted
fields, so the check can over-count but never under-count.
- **What it bounds:** this is a conservative estimate of the cells
pandas will build, not a bound on memory or bytes read. Blank lines held
while sniffing, and a stream spooled because it can't seek, still take
space in proportion to the input.
- **pandas row explosion.** pandas 2.x's C tokenizer reads a lone `\r`
line ending followed by a whitespace-only line as 2^18 empty rows, so
`b",\n\r ,"` (5 bytes) becomes 262,145 rows inside `pd.read_csv`, before
any check can run (fixed in pandas 3.1; the project pins `<3`). pandas
now reads the file through a streaming reader that converts each lone
`\r` to `\n`. `\r\n` is left alone, so text inside quoted fields of CRLF
files is unchanged.
- **One sniffer.** When the delimiter is left to pandas (the context's
restricted sniffer found none), it is sniffed once, by us, from the
first non-blank line, and passed to pandas explicitly, so the check and
the read agree. A file with no usable delimiter is read as one column,
split on the ASCII unit separator.
- **Inputs.** Leading byte-order marks are dropped for both the check
and the read. `partition_tsv()` still decompresses a compressed filename
(e.g. `.tsv.gz`), with the same content going to the check and the read,
and copies a stream that cannot seek to a temporary file once.
## Measurements
| File | Before | After |
|---|---|---|
| 3 KB CSV, 1000-field first line + 1000 one-field lines (1M cells) |
478 MiB, 5s | unchanged (under the limit) |
| 15 KB CSV, 5000 × 5000 | 8.7 GiB, 125s | rejected in 0.1s |
| 3 MB CSV, 1M-field first line + 1M lines | not run (would need ~10^12
cells) | rejected in 0.2s |
| 5 bytes, `b",\n\r ,"` | 262,145 rows | 2 rows |
| 80 bytes, that pattern repeated 16× | over 768 MB inside `pd.read_csv`
| 2 rows per repeat |
The check was fuzzed against pandas' actual output in a memory-capped
container. About 27,000 generated inputs (quotes, CR/LF/CRLF,
form-feeds, BOMs, leading blank lines, 3-byte chunks) gave 0 undercounts
and 0 parser blowups. A further 10,000 inputs read with a header gave
13,727 comparisons, counting the header row and implicit-index columns,
also with 0 undercounts.
## Behaviour changes
- A dense CSV/TSV spanning more than 5M cells (roughly 10 MB or more) is
now rejected. Set `CSV_MAX_CELLS` higher to allow it.
- A file with no usable delimiter, where pandas' own sniff used to raise
`csv.Error`, is now read as one column.
## Tests
- A wide first line is rejected before pandas reads the file, for CSV
and for TSV from both a path and a file.
- Exact-limit passes and one-under failures.
- The record pandas sizes the frame by: blank and whitespace-only first
lines, quoted newlines, form-feeds, mid-field quotes, BOMs, and each
engine's blank-line rule.
- The carriage-return explosion.
- Sniffing past blank lines and the one-column fallback.
- Implicit-index columns: a TSV whose 1-field header is followed by a
100-field row and a ragged tail is rejected before pandas reads it, with
output matching pandas within the limit; the Python engine's index-names
layout is also covered.
- Streaming: fixed-size reads; linear copying while replaying a
200,000-line blank prefix; output that doesn't depend on where chunks
split multibyte characters, BOMs, quotes or CRLF; and stopping reading
once the limit is passed.
- TSV: a field over the `csv` module's 128 KiB limit, a `.tsv.gz`
filename, and a stream that cannot seek.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated description by cubic. -->
<a
href="https://cubic.dev/pr/Unstructured-IO/unstructured/pull/4518?utm_source=github"
target="_blank" rel="noopener noreferrer"
data-no-image-dialog="true"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"><source
media="(prefers-color-scheme: light)"
srcset="https://www.cubic.dev/buttons/review-in-cubic-light.svg"><img
alt="Review in cubic"
src="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"></picture></a>
<!-- End of auto-generated description by cubic. -->
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>