Scanned PDF OCR Extraction
Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted1 hour ago
I have a batch of English-language PDFs that were created from scanned documents, so the text is currently locked inside the images. I need that text pulled out cleanly for subsequent analysis and natural-language processing work I’m running on my end.
You can use any reliable OCR workflow you’re comfortable with—Tesseract, ABBYY FineReader, Adobe OCR, or a custom Python script with libraries such as pytesseract or pdf2image—as long as the final output is accurate. Formatting isn’t critical, but line breaks should follow the original paragraphs well enough for downstream tokenization.
Deliverables:
• One UTF-8 .txt file (or CSV if you prefer) for each PDF, named identically to its source file
• A short note on the OCR engine/settings you used so I can reproduce the results if needed
I’ll share a small sample first so you can demonstrate accuracy, then we’ll move on to the full set.
You can use any reliable OCR workflow you’re comfortable with—Tesseract, ABBYY FineReader, Adobe OCR, or a custom Python script with libraries such as pytesseract or pdf2image—as long as the final output is accurate. Formatting isn’t critical, but line breaks should follow the original paragraphs well enough for downstream tokenization.
Deliverables:
• One UTF-8 .txt file (or CSV if you prefer) for each PDF, named identically to its source file
• A short note on the OCR engine/settings you used so I can reproduce the results if needed
I’ll share a small sample first so you can demonstrate accuracy, then we’ll move on to the full set.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.