Extract 10M Walmart Product Data
Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted1 hour ago
I need a full-scale crawl of more than ten million Walmart product pages focused strictly on two elements: the complete product description (title, bullet points, long form copy, specifications) and every image available in the gallery for each SKU. Price, availability, reviews, or other metadata are not required at this stage, so the scraper can concentrate on content-rich fields and high-resolution media.
The raw HTML is not necessary; I would like the text cleaned and structured (JSON or CSV) and the images downloaded or stored as direct links, organised by SKU. Because of the size of the catalogue, I expect the solution to run in parallel—cloud-based orchestration with Scrapy, Python requests, Playwright or a comparable headless browser is fine as long as it can bypass typical anti-bot measures and stay within Walmart’s published usage policies.
Deliverables:
• A repeatable scraper/spider with clear README.
• The initial dataset covering 10M+ SKUs: one file with structured descriptions, a mirrored directory or bucket containing the images.
• A brief performance report outlining crawl speed, error rate and any throttling safeguards.
I will validate by spot-checking record count against random category URLs and confirming image paths resolve correctly.
The raw HTML is not necessary; I would like the text cleaned and structured (JSON or CSV) and the images downloaded or stored as direct links, organised by SKU. Because of the size of the catalogue, I expect the solution to run in parallel—cloud-based orchestration with Scrapy, Python requests, Playwright or a comparable headless browser is fine as long as it can bypass typical anti-bot measures and stay within Walmart’s published usage policies.
Deliverables:
• A repeatable scraper/spider with clear README.
• The initial dataset covering 10M+ SKUs: one file with structured descriptions, a mirrored directory or bucket containing the images.
• A brief performance report outlining crawl speed, error rate and any throttling safeguards.
I will validate by spot-checking record count against random category URLs and confirming image paths resolve correctly.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.