Mixed Text PDF OCR Conversion

via Freelancer ·

Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted2 hours ago
I have a batch of scanned PDF pages that contain English content in both standard fonts and handwritten notes. Your task is to extract every word into an editable text file, keeping the original reading order and marking any portions that the software cannot recognise with a short “[[?]]” tag so I can spot-check them later.

Because the material switches between printed and handwritten sections, simple one-click OCR will not be enough; I expect you to pair a solid OCR engine (Adobe Acrobat, ABBYY FineReader, Tesseract, or a similar tool you already trust) with careful human proofreading to reach a very high accuracy rate. Spelling, punctuation, and line breaks should match the source unless a clear formatting error needs correction.

Deliverables
• A clean .docx or .txt file for each supplied PDF
• A brief log noting pages or words you flagged with “[[?]]” and any assumptions you had to make

Accepted once I can copy-paste the text with minimal fixes beyond the flagged spots.

Let me know how many pages you can comfortably process in a day and when you can start; I will share a small two-page sample so we confirm quality before sending the full set.
data entry proofreading ocr image processing abbyy finereader adobe acrobat text recognition
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.