Daily E-commerce Scrape to S3

via Freelancer ·

Budget / Salary£10–11
TypeFreelance project
LocationRemote
Posted49 minutes ago
Each day I need a fresh, large-scale pull of both text and image data from multiple e-commerce sites, then have the results land neatly in my Amazon S3 bucket.
Price can be discussed
Here’s the flow I have in mind:
• Your scraper fetches every required product page, captures titles, descriptions, prices, metadata, plus the associated product images.
• The job runs on a schedule (cron, CloudWatch Events, etc.) and automatically handles pagination, rate limits, and site structure changes.
• Text fields are delivered as compressed JSON or Parquet; images are saved in an organised key path such as /YYYY/MM/DD/site_name/sku.jpg.
• Once the transfer to S3 completes, a simple manifest or log file confirms record counts, failed URLs, and total image bytes for the day.

I already have the destination bucket ready and can provide IAM credentials scoped to that resource. You are free to build the pipeline in Python (Scrapy, Requests, BeautifulSoup), Node, or another stack you’re comfortable with, as long as it is container-friendly and easy to redeploy.

Acceptance criteria
1. A working script or container image that runs end-to-end against at least one target site and drops data in S3.
2. Clear instructions (README) for adding additional e-commerce domains.
3. Error handling with retry logic and a summary log pushed to S3.
4. Daily run proves stability for one week before project sign-off.

If you have experience coping with dynamic pages, anti-bot measures, or very large image payloads, let me know in your proposal along with a quick outline of the toolchain you plan to use.
php python data processing web scraping software architecture amazon web services json scrapy beautifulsoup amazon s3
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.