unstructured
ac5048bf - enhancement: remove duplicate embedded images (#2897)

Commit
1 year ago
enhancement: remove duplicate embedded images (#2897) This PR aims to remove duplicate embedded images taken by `PDFminer`. ### Summary - add `clean_pdfminer_duplicate_image_elements()` to remove embedded images with similar `bboxes` and the same `text` - add env_config `EMBEDDED_IMAGE_SAME_REGION_THRESHOLD` to consider the bounding boxes of two embedded images as the same region - refactor: reorganzie `clean_pdfminer_inner_elements()`
Parents
Loading