PDF Text & Image Extraction
Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted1 hour ago
I have a text-based PDF and need every character captured exactly as it appears, with no headers, footers, or page breaks lost in the process. Because the file isn’t scanned, a direct digital extraction should give clean results—no OCR cleanup required.
Alongside the text, I also want every embedded image saved out to its own file at original resolution. A logical naming convention that maps each image back to its page would be helpful for later reference.
Deliverables
• One UTF-8 .txt (or .docx if you prefer) containing the full document text in reading order
• A folder of all extracted images, each named to indicate page number and position
Accuracy is key: paragraphs must stay intact and the image set must be complete. Feel free to use Python (pdfminer.six, PyPDF2), pdftotext, Adobe Acrobat scripting—whatever you’re most efficient with—as long as the final files open cleanly on both Windows and macOS.
Alongside the text, I also want every embedded image saved out to its own file at original resolution. A logical naming convention that maps each image back to its page would be helpful for later reference.
Deliverables
• One UTF-8 .txt (or .docx if you prefer) containing the full document text in reading order
• A folder of all extracted images, each named to indicate page number and position
Accuracy is key: paragraphs must stay intact and the image set must be complete. Feel free to use Python (pdfminer.six, PyPDF2), pdftotext, Adobe Acrobat scripting—whatever you’re most efficient with—as long as the final files open cleanly on both Windows and macOS.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.