Secure Web Scraping & Cloud Deployment Expertise Needed

via Freelancer ·

Budget / Salary₹1,500–12,500
TypeFreelance project
LocationRemote
Posted1 hour ago
Job Title: Web Scraping & Cloud Automation Expert Needed for Highly-Secured Dynamic Portal
Project Overview:
We are looking for an experienced web scraping developer to build an automated data-extraction bot. The bot will need to monitor a specific institutional website on an hourly basis to detect newly uploaded documents/records, extract the names of the entities associated with these uploads, and log them into a Google Sheet.
Once the core extraction is working, the bot will also need to categorize these findings based on specific business rules (we will provide the exact categorization logic and criteria to the hired freelancer later). The entire solution must be deployed and run autonomously on the cloud.
Key Deliverables:
Hourly Monitoring: A script that checks the target web page every hour for new entries.
Data Extraction & Categorization: Parse the newly added entity names, apply our custom categorization logic, and format the data.
Google Sheets Integration: Automatically append the new, categorized records to a designated Google Sheet using the Google Sheets API.
Cloud Deployment: Deploy the script to a cloud environment (e.g., AWS Lambda, Google Cloud Run, Railway, or a standard VPS) with a cron scheduler.
Crucial Technical Hurdles & Limitations to Overcome:
The target website is strictly maintained and actively deters automated scraping. Please only apply if you have proven experience overcoming the following hurdles:
Strict Anti-Bot Measures: The site utilizes advanced Web Application Firewalls (WAF) and bot-detection systems. Standard requests or urllib calls will be blocked immediately. You must be proficient in bypassing these using undetected headless browsers (e.g., Playwright Stealth, Selenium Undetected), rotating user agents, or TLS fingerprint spoofing.
IP Rate Limiting and Bans: Hitting the site exactly every hour from the same data-center IP might result in a permanent IP ban. You will need to implement a reliable proxy rotation strategy (residential proxies preferred) to mimic human traffic.
Dynamic/JavaScript-Rendered Content: The target data is not served in plain HTML. The tables and lists are rendered dynamically via JavaScript. Your scraper must be capable of waiting for network idle states and rendering DOM elements before extracting the text.
Session & Header Management: The site may require specific cookie handling, referrers, or dynamic tokens generated during the session.
Failure Alerting: The website's DOM structure may change without notice. The bot must have robust error handling and send a notification (e.g., via email or Telegram) if the scraping fails so we can update the selectors.
Ideal Tech Stack:
Python (Playwright, Selenium, or Scrapy)
Google Sheets API (gspread)
Cloud Hosting / Serverless architecture
Proxy management tools
To Apply:
Please start your proposal by briefly describing a past project where you successfully bypassed strong anti-bot protections (like Cloudflare or Akamai). Let us know your preferred tech stack for this setup and how you plan to manage the IP/bot-detection hurdles.
php python cloud computing web scraping google app engine data extraction data integration automation
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.