PDF Text & Image Extraction

via Freelancer ·

Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted1 hour ago
I need both the text and every embedded image pulled from a series of PDFs. The text can come across as plain, unstyled characters—no fonts, colours, or tables need to be preserved—but please keep the original reading order intact. Each image should be exported at its full resolution and positioned next to the relevant text, or clearly referenced so the relationship is obvious.

When the extraction is complete, combine everything into a single Word document (.docx). The result should read like a stripped-back version of the source files: clean text interleaved with the corresponding images.

Use whichever method suits you best—Python libraries such as pdfminer.six, PyMuPDF, or an Adobe-based workflow are all acceptable—as long as the output is complete and accurate. Let me know your approach and the turnaround time you can commit to.
javascript python software architecture pdf c++ programming image processing data extraction adobe acrobat
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.