Clean Data from Scanned Records
Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted1 hour ago
I have a batch of scanned documents that must be transformed into a reliable, analysis-ready dataset. After the text is extracted, the core job is to hunt down and remove duplicated entries, straighten out any formatting inconsistencies, and identify or flag incomplete records so nothing slips through the cracks.
You’re free to use the tools you’re most comfortable with—Excel, Google Sheets, OpenRefine, Python + pandas, or similar—as long as the final file is accurate and easy for me to audit.
Deliverables
• A single spreadsheet (CSV or XLSX) containing all records, minus duplicates
• Consistent formatting applied across every column (dates, currency, capitalization, etc.)
• A separate tab or report that lists any incomplete records you found and the fields that are missing
I’ll spot-check against the original scans; if the cleaned file aligns and all issues are properly addressed, the project is complete.
You’re free to use the tools you’re most comfortable with—Excel, Google Sheets, OpenRefine, Python + pandas, or similar—as long as the final file is accurate and easy for me to audit.
Deliverables
• A single spreadsheet (CSV or XLSX) containing all records, minus duplicates
• Consistent formatting applied across every column (dates, currency, capitalization, etc.)
• A separate tab or report that lists any incomplete records you found and the fields that are missing
I’ll spot-check against the original scans; if the cleaned file aligns and all issues are properly addressed, the project is complete.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.