fix(chunking): retain table rowspans across cell splits (#4495)
## Summary
This PR keeps HTML table columns intact when an oversized row must be
split into cells. A `rowspan` still occupies its original column on
later source rows, even when those rows have no cell at that position.
For example, this source has two columns:
```html
<table><tr><td rowspan="3">Region A</td><td>first</td></tr>
<tr><td>an oversized description that must split…</td></tr>
<tr><td>third</td></tr></table>
```
Before this change, cell splitting discarded the active span, so `third`
could land in column 1. Now the chunker remembers that `Region A` covers
column 1 through row 3 and places `third` in column 2. A copied cover is
clipped to the rows in its output fragment; when repeating its text
would be too costly, a blank cell preserves the column.
## How it works
The active-span ledger records occupied column ranges and their expiry
rows. An expiry heap and indexed free-column gaps place each new cell in
the first gap wide enough for its `colspan`, without scanning every live
span on a sparse row. The fragment renderer tracks blank runs and label
expiry events, extending stable rowspans until the source table changes.
Early retention approaches could do roughly spans × continuation rows of
work by scanning live spans or emitting blanks on every row. Later cases
amplified retained labels: a long label or many separated labels could
be serialized or remeasured on every continuation. The current design
emits fitting labels once per suitable fragment, coalesces blank
geometry, and bounds optional repeated custom-token measurement with a
probe budget. If that budget runs out, packing flushes before admitting
another unmeasured row. This follow-up also charges newly accepted text
rowspans, including those opened by a packed row, while leaving
blank-only carry free of that charge. Actual output size and custom
whole-candidate measurement still contribute to runtime.
The change also prepares version 0.27.9 and updates the changelog.
## Correctness and validation
- Source cells keep their logical columns after row and cell splits;
copied spans end within both their source lifetime and their emitted
fragment. Text-bearing output remains within the configured measure.
- The token-mode regression failed before the follow-up fix: 39,200
retained-label characters were submitted to the custom measure. It now
checks both initial and packed-row span creation, carry with available
budget, forced flush, label and cell retention, column geometry, bounded
chunks, and the hard token limit. A separate regression asserts the
exact aggregate covering-label count in its fixture.
- Focused local validation on this head:
`test_unstructured/chunking/test_base.py -k 'not performance'` — 338
passed, 11 skipped. Ruff check and format check passed on both changed
Python files. The previous broader focused run reported 493 passed, 11
skipped; exact-head CI supplies broad validation.
- Earlier focused benchmarks reported 0.17–0.21 s for 1,000 spans and
2,000 sparse rows versus 0.42–0.45 s on `main`. A 1,000-span, 2,000-row
near-limit case emitted 4,003 cells in 0.06 s versus over 2 million
cells in 7.88 s on an earlier implementation. These are case-specific
measurements, not a general complexity guarantee.
- CI fixture repair on this head: `test_ingest_src (3.12)` could not
start on the prior head because Quay returned HTTP 401 for the MinIO
test image. The compose fixture now pins a source-built MinIO image by
digest. Its local Compose health check, bucket creation, and fixture
upload passed; the exact-head ingest check passed in CI.
Review status: all three Cubic threads were addressed, and the final
Cubic review found zero new issues. A fresh full-context Oracle `6Pro`
review of `ada1ea5` returned `SAFE TO MERGE` with no High/Critical
findings. All six required checks and the complete 54-check rollup
passed on that head. Human approval and merge remain pending.
(authored by codex)