unstructured
69d50a74 - Fix DOCX text_as_html duplicating merged-cell text instead of colspan/rowspan (#4469)

Commit
26 days ago
Fix DOCX text_as_html duplicating merged-cell text instead of colspan/rowspan (#4469) ## Summary DOCX table extraction produced two different representations of a merged cell: `.text` included its content once, but `.metadata.text_as_html` repeated the content into every `<td>` the merge visually covered, with no `colspan`/`rowspan` attribute marking that a merge had happened at all. Root cause: `_convert_table_to_html` built its HTML from `python-docx`'s `row.cells`, which resolves both horizontal (`gridSpan`) and vertical (`vMerge="continue"`) merges by yielding the *same* underlying cell content at every grid position the merge spans. That already-expanded grid was serialized straight into HTML with one `<td>` per matrix position and no span-collapsing step. Fix: track the underlying `tc` XML element's identity (not cell text) when building the matrix — the same `tc` object appears at every grid position a merge covers — then collapse runs of identical identity into `(colspan, rowspan)` before emitting `<td>`s. DOCX only permits rectangular merges, so there's no irregular-region case to handle. A new shared helper (`collapse_matrix_of_keyed_cells_to_spans`) does the collapsing; the existing `htmlify_matrix_of_cell_texts` (still used unchanged by the pptx and HTML-parser partitioners) now shares its cell-escaping logic with the new span-aware path via an extracted `_format_td` helper. ## Test plan - [x] `partition_docx` on `example-docs/docx-tables.docx` (the merged-cell fixture) now produces `text_as_html` with correct `colspan`/`rowspan` and no duplicated cell text — added a behavioral regression test through the public `partition_docx` API. - [x] Existing `test_docx.py` fixtures pinning the old duplicated-text/no-span output updated to the new expected output. - [x] Checked downstream consumers that assumed DOCX tables never carry spans: - `unstructured/metrics/table/table_extraction.py`'s span-aware grid reconstruction already handles `colspan`/`rowspan` correctly (written for other span-producing sources) — traced by hand against the merged-cell fixture. - Chunking's table splitter handles real spanned DOCX tables without crashing at multiple `max_characters` values. Found one pre-existing (not introduced here) limitation: splitting between rows that share a `rowspan` can drop that cell's data in the later chunk — this was always latent for any spanned HTML source, just never exercised for DOCX before since DOCX never emitted spans until now. Not fixed in this PR; flagging as a possible follow-up. - [x] Full relevant test suites green (partition/docx, common/html_table, chunking, metrics/table), lint clean. <!-- This is an auto-generated description by cubic. --> <a href="https://cubic.dev/pr/Unstructured-IO/unstructured/pull/4469?utm_source=github" target="_blank" rel="noopener noreferrer" data-no-image-dialog="true"><picture><source media="(prefers-color-scheme: dark)" srcset="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"><source media="(prefers-color-scheme: light)" srcset="https://www.cubic.dev/buttons/review-in-cubic-light.svg"><img alt="Review in cubic" src="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"></picture></a> <!-- End of auto-generated description by cubic. -->
Author
Parents
Loading