PDF Text Extraction with Structure

via Freelancer ·

Budget / Salary₹12,500–37,500
TypeFreelance project
LocationRemote
Posted1 hour ago
I need the text pulled from one or more PDF files and returned as straight-forward .txt documents. What matters most to me is that paragraph breaks and any bullet or numbered lists survive the journey intact; I do not need images, tables, fonts, or any other layout elements.

Feel free to rely on Python (pdfminer.six, PyPDF2, Tika, etc.), Adobe tools, or another reliable method—as long as the final files open as clean, human-readable plain text with the original paragraph flow and list hierarchy clearly visible.

Deliverables
• One plain-text file per PDF provided, encoded in UTF-8
• All paragraphs separated by single blank lines; lists reproduced with their bullets or numbers
• A brief note on the tool or script you used so I can replicate the process if needed

I will supply the PDFs as soon as we start. If your output mirrors the source paragraphs and lists without missing or garbling characters, I will consider the job complete and release payment immediately.
c programming python data processing software architecture pdf c++ programming data extraction adobe acrobat
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.