unstructured
ee2b3a35 - fix: tolerate non-UTF-8 soffice output in doc/ppt conversion (#4465)

Commit
28 days ago
fix: tolerate non-UTF-8 soffice output in doc/ppt conversion (#4465) ## Summary `convert_office_doc()` — shared by `partition_doc()` and `partition_ppt()` — decoded `soffice` stdout and stderr with a strict UTF-8 `.decode()` at three sites, used only for logging and to check whether stdout was empty (which drives the retry loop and the final failure check). LibreOffice echoes the input path in the console encoding, which on Windows is the locale codepage rather than UTF-8, so a name or path with multi-byte characters makes the decode raise and aborts a conversion that had already succeeded: ``` UnicodeDecodeError: 'utf-8' codec can't decode byte 0x8a in position 32: invalid start byte ``` The traceback in #3652 is on the first stdout decode. The stderr site is reachable and broken the same way when a conversion fails, and sits outside the surrounding `try` entirely. Fixing the child process's output encoding is not an option — @Snowman-s confirmed on the issue that `PYTHONIOENCODING=utf-8:surrogateescape` and `sys.stdout.reconfigure()` both leave the exception unchanged. ## Fix All three sites go through one `_decode_soffice_output()` helper decoding with `errors="backslashreplace"`. @scanny weighed three options on the issue and voted for option 1, `try/except UnicodeDecodeError` falling back to `str(bytes)`. A codec error handler gets the same tolerance in one expression with no rarely-exercised second branch, and escapes only the offending bytes rather than turning a perfectly good line into a `repr`. Option 2 — detect Windows and decode with the locale encoding — would render the filename correctly when the guess is right, but the console output codepage need not match `locale.getpreferredencoding()` — anything that ran `chcp` upstream moves one and not the other — and a wrong guess is silent mojibake rather than a visible escape. It layers cleanly on top of this later if you want it. `backslashreplace` rather than `replace` because U+FFFD is not encodable in the very codepages this bug is about: a `logging.FileHandler` whose stream uses cp932 or cp936 cannot encode it and drops the whole record, so `replace` would stop the abort and then throw away the diagnostic. `backslashreplace` keeps the line pure ASCII and shows the actual bytes — closer to what option 1 was reaching for. One behavior change worth flagging: stdout consisting *entirely* of undecodable bytes now yields a non-empty string, so with `returncode == 0` it no longer trips the empty-stdout heuristic the code uses to detect a silent failure. Real `soffice` output carries ASCII scaffolding around the path, so this needs a degenerate payload, but it is a real narrowing of that heuristic — one option 1 shares, since `str()` of non-empty bytes is never empty either. The `b""` that the heuristic actually watches for still decodes to `""`. ## Testing No LibreOffice needed — stub `subprocess.run` with the bytes `soffice` emits on a Japanese console: ```python from unittest import mock from unstructured.partition.common import common out = mock.Mock(returncode=0, stdout="convert 文章.doc -> 文章.docx".encode("cp932"), stderr=b"") with mock.patch.object(common.subprocess, "run", return_value=out): common.convert_office_doc("文章.doc", "fake-directory", target_format="docx") ``` Raises `UnicodeDecodeError` on `main`; returns normally here. `test_convert_office_doc_survives_non_utf8_soffice_output` is parametrized over stdout and stderr so both decode paths are covered, and asserts the escaped bytes reach the log and that the record is ASCII-encodable — it fails on `main`, and also fails if the helper is switched to `errors="replace"`. `test_common.py`: 63 passed / 2 failed here, 61 passed / 4 failed with the same file on `main` — the extra two are the new test's stdout and stderr cases. The other two fail identically on both (`FileNotFoundError: soffice command was not found`; no LibreOffice on this machine). `ruff check .` passes repo-wide and `ruff format --check` passes on both changed files (repo-wide it exits non-zero on the `example-docs/umlauts-non-utf8.md` fixture, identically on `main`). Not verified against a real `soffice`: the root cause is from @Snowman-s's and @scanny's analysis on the issue, and the CP932 payload is my own reconstruction of it. Resolves #3652 <!-- This is an auto-generated description by cubic. --> <a href="https://cubic.dev/pr/Unstructured-IO/unstructured/pull/4465?utm_source=github" target="_blank" rel="noopener noreferrer" data-no-image-dialog="true"><picture><source media="(prefers-color-scheme: dark)" srcset="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"><source media="(prefers-color-scheme: light)" srcset="https://www.cubic.dev/buttons/review-in-cubic-light.svg"><img alt="Review in cubic" src="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"></picture></a> <!-- End of auto-generated description by cubic. -->
Author
Parents
Loading