PDF Text Extraction to SQL

via Freelancer ·

Budget / Salary₹12,500–37,500
TypeFreelance project
LocationRemote
Posted1 hour ago
I have a batch of PDFs that contain only plain-text content. Every word from every page needs to land neatly in a single SQL table—no images, no form fields, just raw text.

Here is what the finished job looks like to me:

• A script (Python, Java, or similar) that reads each PDF, pulls out the complete text, and writes it into a table such as `pdf_text` with columns like `id`, `file_name`, `page_number`, and `text_block`.
• A SQL file that creates this table and lets me recreate it easily on MySQL or PostgreSQL.
• Clear instructions so I can rerun the process whenever new PDFs arrive.

Accuracy is critical: the stored text must perfectly match the source PDF, page by page. If you prefer libraries such as PyPDF2, PDFMiner, Tika, or pdftotext, feel free—as long as the output is reliable and the setup steps are documented.

Once the import script runs without errors and I can query every PDF line from the database, the project is complete.
php java python data processing pdf mysql postgresql data extraction
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.