AI-Based Data Extraction from Documents

via Freelancer ·

Budget / Salary₹12,500–37,500
TypeFreelance project
LocationRemote
Posted2 hours ago
Project Title
AI Model / OCR Automation to Extract Data from PNG, JPEG & PDF and Populate Excel Sheets

Project Description

I am looking for an experienced AI/ML developer or Python developer who can develop a tool/model that can automatically extract required data from PNG, JPEG, scanned documents, and PDF files and enter the extracted information into a predefined Excel template.

The main challenge is that **the source documents will not always have the same format or layout**. Data may appear at different positions, under different headings, or in different table formats.

The system should therefore be able to **understand the context/title/heading of the information**, identify the required value, and place it in the correct cell/column of the Excel sheet.

Example Workflow

**Input:**
- PNG/JPEG images of test reports, bore logs, laboratory reports, etc.
- Scanned or digital PDF documents
- Different formats/layouts from different sources

**Processing:**
The AI should:
1. Read and understand the document.
2. Identify relevant headings/titles and associated values.
3. Extract the required information even if its position changes between documents.
4. Handle tables and scanned documents using OCR where required.
5. Map the extracted information to the appropriate Excel column based on the **heading/context**, rather than simply relying on fixed coordinates.
6. Apply predefined conversion/calculation rules where required.
7. Populate the existing Excel template without disturbing its formatting, formulas, logos, borders, colours, etc.

**Output:**
A completed Excel file with the extracted data accurately entered into the appropriate cells.

### Important Requirement

I do **not** want a simple OCR tool that extracts text based only on fixed positions.

For example, if one report contains:

> Bulk Density – 1.97 g/cc

and another report contains:

> Natural Bulk Density: 2.01 g/cc

the system should understand that both refer to **Bulk Density** and place the corresponding value in the correct Excel field.

Similarly, the location/order of **SPT N-value, soil classification, density, fines percentage, depth, test results, etc.** may vary between documents.

The system should be capable of handling these variations using **AI/NLP/LLM-based document understanding, OCR, table extraction, or a suitable combination of technologies**.

### Excel Requirements

The tool should be able to:

- Work with existing Excel templates.
- Identify the correct worksheet.
- Populate specific fields/columns.
- Preserve existing formatting.
- Preserve formulas and merged cells.
- Avoid changing logos, borders, colours, and page layout.
- Create/copy/rename worksheets where required.
- Apply predefined data-entry and calculation rules.
- Generate the final completed Excel workbook.

### Preferred Technologies

I am open to the developer's recommendation, but experience with the following would be preferred:

- Python
- OCR
- OpenAI/LLM APIs or other AI models
- Document AI
- Computer Vision
- PDF/Table extraction
- Pandas
- OpenPyXL
- Azure Document Intelligence / Google Document AI / AWS Textract or similar
- RAG/LLM-based document understanding
- Machine Learning model fine-tuning, if required

### Important

I will provide **sample PNG/JPEG/PDF files along with the corresponding correctly completed Excel files** as examples.

The developer should use these examples to understand the required mapping and develop a system that can generalize to **new documents with different formats**.

Accuracy is very important because this will be used for **engineering/geotechnical investigation data**, where incorrect extraction or mapping can lead to incorrect analysis.

### Deliverables

1. Working Python code/application.
2. AI/OCR/document-processing pipeline.
3. Excel automation.
4. Ability to process multiple files in batch.
5. Proper mapping of extracted data to the Excel template.
6. Error/exception handling for missing or unclear data.
7. Source-document → Excel field mapping logic.
8. Documentation explaining how to run and use the tool.
9. Testing using my sample documents.
10. Source code and all necessary files.

### To Apply

Please share:

- Your previous experience with **AI document extraction/OCR → Excel automation**.
- Examples of similar projects you have completed.
- Which AI/OCR technology you recommend.
- Whether you have experience handling **variable document formats**.
- Your estimated timeline and budget.
- How you would approach training/fine-tuning the model using sample documents and completed Excel files.

**Please do not propose a basic OCR-to-Excel solution. I specifically need a system that can understand the meaning/context of the data and map it correctly to the Excel template even when the document format changes.**
geotechnical engineering data extraction ai development ai integration ai training data
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.