Python Web App : Scraping and Lead Generation
Budget / Salary$30–250
TypeFreelance project
LocationRemote
Posted1 hour ago
Senior Full-Stack Python Developer / Team — Multi-Tenant Data Scraping & Lead Generation Web Application
Project Overview
We are seeking an experienced Senior Full-Stack Python Developer (or small specialized agency) to build a scalable, multi-tenant web application for automated web scraping, contact discovery, and lead management.
The application will host two independent scraping engines under a unified user and admin dashboard, featuring background task queues, anti-bot mechanisms, automated data enrichment, and direct Google Sheets API integration.
Key Technical Architecture & Modules
1. Module 1: Job Site Scraping Engine
Automated data extraction from dynamic and static job portals.
Support for site-specific selectors, pagination rules, rate limiting, and category/subcategory filtering.
High-volume background processing (~200 records per source daily).
2. Module 2: General Data & Lead Generation Engine
Niche and location-based business data extraction from web sources, public directories, and map platforms.
Integration interface for third-party enrichment APIs (e.g., LinkedIn/executive data).
3. Automated Contact Discovery & Formatting Engine
Deep-crawl fallback logic: Automatically inspects company websites (contact pages, footers) and public sources when primary contact info is missing.
Phone number parsing and standardization (E.164 international format with click-to-call support).
Email syntax cleaning and deliverability validation.
4. Multi-Tenant User Dashboard & Admin Panel
Secure user authentication with OTP verification.
Role-Based Access Control (RBAC) with logical database isolation (user_id separation).
System-wide task scheduling (admin-controlled fixed intervals).
Real-time task progress monitoring, logs, and failure email alerts.
5. Integrations & Data Sync
Google Sheets API: Direct OAuth-based auto-sync per individual user account.
Server-side Excel Export (.xlsx): Auto-generated downloadable files with date/time tracking.
Automated Webhooks: Webhook support for external CRM pushes.
Deduplication: Hash-based checks (user_id + unique record identifiers) to prevent duplicate record storage.|
Required Tech Stack
Backend Framework: Python (Django or FastAPI)
Scraping & Crawling: Scrapy, Playwright, Selenium, BeautifulSoup
Anti-Bot Countermeasures: Rotating residential proxy pools, CAPTCHA bypass, stealth headless browsers
Task Queue & Scheduler: Celery + Redis + Celery Beat
Database: PostgreSQL (indexed for fast deduplication lookups)
Frontend: React.js / Next.js with Tailwind CSS or Bootstrap
How to Apply
Please submit your proposal with:
Relevant Portfolio: 2–3 examples of complex web scrapers, data pipelines, or multi-tenant Django/React applications you have built.
Technical Approach: Briefly outline your preferred setup for proxy rotation, headless browser management, and handling Celery worker queues.
Fixed Price & Timeline: Provide an estimated cost and milestone timeline for delivering the core scraping engines and web application.
Project Overview
We are seeking an experienced Senior Full-Stack Python Developer (or small specialized agency) to build a scalable, multi-tenant web application for automated web scraping, contact discovery, and lead management.
The application will host two independent scraping engines under a unified user and admin dashboard, featuring background task queues, anti-bot mechanisms, automated data enrichment, and direct Google Sheets API integration.
Key Technical Architecture & Modules
1. Module 1: Job Site Scraping Engine
Automated data extraction from dynamic and static job portals.
Support for site-specific selectors, pagination rules, rate limiting, and category/subcategory filtering.
High-volume background processing (~200 records per source daily).
2. Module 2: General Data & Lead Generation Engine
Niche and location-based business data extraction from web sources, public directories, and map platforms.
Integration interface for third-party enrichment APIs (e.g., LinkedIn/executive data).
3. Automated Contact Discovery & Formatting Engine
Deep-crawl fallback logic: Automatically inspects company websites (contact pages, footers) and public sources when primary contact info is missing.
Phone number parsing and standardization (E.164 international format with click-to-call support).
Email syntax cleaning and deliverability validation.
4. Multi-Tenant User Dashboard & Admin Panel
Secure user authentication with OTP verification.
Role-Based Access Control (RBAC) with logical database isolation (user_id separation).
System-wide task scheduling (admin-controlled fixed intervals).
Real-time task progress monitoring, logs, and failure email alerts.
5. Integrations & Data Sync
Google Sheets API: Direct OAuth-based auto-sync per individual user account.
Server-side Excel Export (.xlsx): Auto-generated downloadable files with date/time tracking.
Automated Webhooks: Webhook support for external CRM pushes.
Deduplication: Hash-based checks (user_id + unique record identifiers) to prevent duplicate record storage.|
Required Tech Stack
Backend Framework: Python (Django or FastAPI)
Scraping & Crawling: Scrapy, Playwright, Selenium, BeautifulSoup
Anti-Bot Countermeasures: Rotating residential proxy pools, CAPTCHA bypass, stealth headless browsers
Task Queue & Scheduler: Celery + Redis + Celery Beat
Database: PostgreSQL (indexed for fast deduplication lookups)
Frontend: React.js / Next.js with Tailwind CSS or Bootstrap
How to Apply
Please submit your proposal with:
Relevant Portfolio: 2–3 examples of complex web scrapers, data pipelines, or multi-tenant Django/React applications you have built.
Technical Approach: Briefly outline your preferred setup for proxy rotation, headless browser management, and handling Celery worker queues.
Fixed Price & Timeline: Provide an estimated cost and milestone timeline for delivering the core scraping engines and web application.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.