langchain
35297ca0 - Add feature for extracting images from pdf and recognizing text from images. (#10653)

Commit

2 years ago

Add feature for extracting images from pdf and recognizing text from images. (#10653) **Description** It is for #10423 that it will be a useful feature if we can extract images from pdf and recognize text on them. I have implemented it with `PyPDFLoader`, `PyPDFium2Loader`, `PyPDFDirectoryLoader`, `PyMuPDFLoader`, `PDFMinerLoader`, and `PDFPlumberLoader`. [RapidOCR](https://github.com/RapidAI/RapidOCR.git) is used to recognize text on extracted images. It is time-consuming for ocr so a boolen parameter `extract_images` is set to control whether to extract and recognize. I have tested the time usage for each parser on my own laptop thinkbook 14+ with AMD R7-6800H by unit test and the result is: | extract_images | PyPDFParser | PDFMinerParser | PyMuPDFParser | PyPDFium2Parser | PDFPlumberParser | | ------------- | ------------- | ------------- | ------------- | ------------- | ------------- | | False | 0.27s | 0.39s | 0.06s | 0.08s | 1.01s | | True | 17.01s | 20.67s | 20.32s | 19,75s | 20.55s | **Issue** #10423 **Dependencies** rapidocr_onnxruntime in [RapidOCR](https://github.com/RapidAI/RapidOCR/tree/main) --------- Co-authored-by: Bagatur <baskaryan@gmail.com>

References

#10653 - Add feature for extracting images from pdf and recognizing text from images.

Author

therontau0054

Parents

8e3fbc97

langchain 35297ca0 - Add feature for extracting images from pdf and recognizing text from images. (#10653)

langchain
35297ca0 - Add feature for extracting images from pdf and recognizing text from images. (#10653)