Data Preprocessing & Model Development
Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted1 hour ago
I have several CSV files filled with raw text that need to be transformed into something my business can actually use. First, I need a solid data-cleaning and preprocessing pipeline: removing noise, normalising case, handling missing or corrupt rows, tokenising, and producing features that make sense for downstream modelling. Once the data quality is reliable, I want a working machine-learning model built on top of it—preferably in Python using well-established libraries such as pandas, scikit-learn, spaCy or a comparable NLP stack.
Here is what will let me sign off on the job:
• A reproducible script or notebook that takes the original CSV files, cleans them, and outputs a tidy, feature-ready dataset.
• A trained text-based model (classification or clustering—I'll decide with you after an initial review of the data) with clear performance metrics and a brief explanation of how you tuned it.
• A short read-me or comments in the code that explain each step so I can maintain or extend the pipeline later.
If you can turn messy text in CSVs into an accurate, well-documented model, I’m ready to get started.
Here is what will let me sign off on the job:
• A reproducible script or notebook that takes the original CSV files, cleans them, and outputs a tidy, feature-ready dataset.
• A trained text-based model (classification or clustering—I'll decide with you after an initial review of the data) with clear performance metrics and a brief explanation of how you tuned it.
• A short read-me or comments in the code that explain each step so I can maintain or extend the pipeline later.
If you can turn messy text in CSVs into an accurate, well-documented model, I’m ready to get started.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.