PDF Text Extraction & Formatting
Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted1 hour ago
I have several PDFs and I only need the text pulled from them—no tables, no images—just the words. The final deliverable must be a plain .txt file for each PDF, but not a raw dump. I’ll provide a short set of rules so the text follows a custom structure (section headers on their own lines, line breaks removed in paragraphs, and a specific tag for any footnotes). Accuracy in capturing every character matters more to me than speed.
Here’s how I see the workflow:
1. Extract the text with whatever tool you prefer, then apply the custom layout rules I’ll share right after project start.
2. Double-check the output so there are no stray line breaks, page numbers, or header/footer remnants.
3. Deliver one clean UTF-8 .txt file per source PDF.
If something in the files looks ambiguous, flag it—I’d rather answer quick questions than fix surprises later. Once the first file meets the formatting rules perfectly we’ll use it as the template for the rest.
Here’s how I see the workflow:
1. Extract the text with whatever tool you prefer, then apply the custom layout rules I’ll share right after project start.
2. Double-check the output so there are no stray line breaks, page numbers, or header/footer remnants.
3. Deliver one clean UTF-8 .txt file per source PDF.
If something in the files looks ambiguous, flag it—I’d rather answer quick questions than fix surprises later. Once the first file meets the formatting rules perfectly we’ll use it as the template for the rest.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.