Fix DOCX text_as_html duplicating merged-cell text instead of colspan/rowspan (#4469)
## Summary
DOCX table extraction produced two different representations of a merged
cell: `.text` included its content once, but `.metadata.text_as_html`
repeated the content into every `<td>` the merge visually covered, with
no `colspan`/`rowspan` attribute marking that a merge had happened at
all.
Root cause: `_convert_table_to_html` built its HTML from `python-docx`'s
`row.cells`, which resolves both horizontal (`gridSpan`) and vertical
(`vMerge="continue"`) merges by yielding the *same* underlying cell
content at every grid position the merge spans. That already-expanded
grid was serialized straight into HTML with one `<td>` per matrix
position and no span-collapsing step.
Fix: track the underlying `tc` XML element's identity (not cell text)
when building the matrix — the same `tc` object appears at every grid
position a merge covers — then collapse runs of identical identity into
`(colspan, rowspan)` before emitting `<td>`s. DOCX only permits
rectangular merges, so there's no irregular-region case to handle.
A new shared helper (`collapse_matrix_of_keyed_cells_to_spans`) does the
collapsing; the existing `htmlify_matrix_of_cell_texts` (still used
unchanged by the pptx and HTML-parser partitioners) now shares its
cell-escaping logic with the new span-aware path via an extracted
`_format_td` helper.
## Test plan
- [x] `partition_docx` on `example-docs/docx-tables.docx` (the
merged-cell fixture) now produces `text_as_html` with correct
`colspan`/`rowspan` and no duplicated cell text — added a behavioral
regression test through the public `partition_docx` API.
- [x] Existing `test_docx.py` fixtures pinning the old
duplicated-text/no-span output updated to the new expected output.
- [x] Checked downstream consumers that assumed DOCX tables never carry
spans:
- `unstructured/metrics/table/table_extraction.py`'s span-aware grid
reconstruction already handles `colspan`/`rowspan` correctly (written
for other span-producing sources) — traced by hand against the
merged-cell fixture.
- Chunking's table splitter handles real spanned DOCX tables without
crashing at multiple `max_characters` values. Found one pre-existing
(not introduced here) limitation: splitting between rows that share a
`rowspan` can drop that cell's data in the later chunk — this was always
latent for any spanned HTML source, just never exercised for DOCX before
since DOCX never emitted spans until now. Not fixed in this PR; flagging
as a possible follow-up.
- [x] Full relevant test suites green (partition/docx,
common/html_table, chunking, metrics/table), lint clean.
<!-- This is an auto-generated description by cubic. -->
<a
href="https://cubic.dev/pr/Unstructured-IO/unstructured/pull/4469?utm_source=github"
target="_blank" rel="noopener noreferrer"
data-no-image-dialog="true"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"><source
media="(prefers-color-scheme: light)"
srcset="https://www.cubic.dev/buttons/review-in-cubic-light.svg"><img
alt="Review in cubic"
src="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"></picture></a>
<!-- End of auto-generated description by cubic. -->