Affordable Cloud-Based Web Scraping & Webpage Change Detection System
Budget / Salary₹600–1,500
TypeFreelance project
LocationRemote
Posted2 hours ago
We are a startup seeking a practical, reliable, and low-cost cloud-based web scraping solution.
### 1. Current Situation
Our previous PC-based Python system scraped 5,000 URLs with 18 XPath fields per URL in approximately four hours. It depends on local electricity, internet connectivity, and manual monitoring. Sample input/output spreadsheets and Python code are available.
### 2. Our Requirements
1. **High Performance:** Target 100,000+ URLs per run, extracting approximately 30 fields per URL in around one hour or less. Explain feasibility and realistic performance expectations.
2. **Cloud Automation:** No computer needs to remain switched on. Support scheduled runs twice daily, manual triggering, automatic retries, recovery, logs, and failure alerts.
3. **URL & XPath Management:** Manage URLs and XPath/CSS selectors through a simple spreadsheet or configuration interface. Adding websites and modifying selectors should require minimal technical work.
4. **Webpage Change Detection:** Compare current extracted content against previously stored values. Identify changed URLs and report exactly which fields or XPath contents changed, including previous values, new values, and detection time. Do not report unchanged content as updated.
5. **Correct Output Format:** Extract approximately 30 specified fields per URL into our predefined spreadsheet format, with each URL's data in one row and each field in its correct column.
6. **Website Restrictions:** Explain how the solution handles bot detection, blocking, rate limits, JavaScript protection, and CAPTCHA, including challenges such as Amazon's bot detection. State limitations and practical, permitted alternatives when automated access is restricted.
7. **Low Operating Cost:** Minimise fixed monthly expenses using suitable cloud services and pay-per-use resources. Avoid unnecessary subscriptions and permanently running servers.
### 3. Questions to Answer
1. What are the one-time and monthly costs, including fixed and variable expenses?
2. How long will processing 100,000 URLs with 30 fields per URL realistically take?
3. What technology will you use for cloud automation, extraction, and change detection?
4. How will previous values be stored and specific changes identified?
### 4. Submission
Please provide a practical implementation plan covering the proposed architecture, realistic performance estimates, cost breakdown in INR, change detection, and limitations.
**We prioritise feasibility, accurate change reporting, reliability, simplicity, and low recurring costs. Please avoid generic proposals and unrealistic guarantees.**
### 1. Current Situation
Our previous PC-based Python system scraped 5,000 URLs with 18 XPath fields per URL in approximately four hours. It depends on local electricity, internet connectivity, and manual monitoring. Sample input/output spreadsheets and Python code are available.
### 2. Our Requirements
1. **High Performance:** Target 100,000+ URLs per run, extracting approximately 30 fields per URL in around one hour or less. Explain feasibility and realistic performance expectations.
2. **Cloud Automation:** No computer needs to remain switched on. Support scheduled runs twice daily, manual triggering, automatic retries, recovery, logs, and failure alerts.
3. **URL & XPath Management:** Manage URLs and XPath/CSS selectors through a simple spreadsheet or configuration interface. Adding websites and modifying selectors should require minimal technical work.
4. **Webpage Change Detection:** Compare current extracted content against previously stored values. Identify changed URLs and report exactly which fields or XPath contents changed, including previous values, new values, and detection time. Do not report unchanged content as updated.
5. **Correct Output Format:** Extract approximately 30 specified fields per URL into our predefined spreadsheet format, with each URL's data in one row and each field in its correct column.
6. **Website Restrictions:** Explain how the solution handles bot detection, blocking, rate limits, JavaScript protection, and CAPTCHA, including challenges such as Amazon's bot detection. State limitations and practical, permitted alternatives when automated access is restricted.
7. **Low Operating Cost:** Minimise fixed monthly expenses using suitable cloud services and pay-per-use resources. Avoid unnecessary subscriptions and permanently running servers.
### 3. Questions to Answer
1. What are the one-time and monthly costs, including fixed and variable expenses?
2. How long will processing 100,000 URLs with 30 fields per URL realistically take?
3. What technology will you use for cloud automation, extraction, and change detection?
4. How will previous values be stored and specific changes identified?
### 4. Submission
Please provide a practical implementation plan covering the proposed architecture, realistic performance estimates, cost breakdown in INR, change detection, and limitations.
**We prioritise feasibility, accurate change reporting, reliability, simplicity, and low recurring costs. Please avoid generic proposals and unrealistic guarantees.**
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.