Automated Technological Resource Discovery System
Budget / Salary€100–500
TypeFreelance project
LocationRemote
Posted1 hour ago
NEVIE-GLOBAL is looking for a Python Data Engineer / Web Scraping & Automation specialist to build a large-scale, continuously growing Product & Technology Library.
This is NOT a manual data-entry project. We need an automated discovery and data-processing system capable of collecting, structuring, deduplicating, classifying and enriching hundreds of thousands to potentially millions of RAW records from legally accessible public sources.
SCOPE OF RESOURCES TO DISCOVER
The system must progressively discover and catalogue resources including:
• AI agents and multi-agent systems
• AI agent frameworks and autonomous agents
• Agent Skills / AI Skills
• MCP servers, tools and integrations
• n8n workflows and templates
• Make.com, Zapier and Activepieces automations
• workflow libraries and automation systems
• SaaS and micro-SaaS applications
• SaaS boilerplates and starter kits
• open-source software and applications
• AI applications and AI tools
• CRM and ERP systems
• business applications
• website templates
• landing-page templates
• e-commerce templates
• admin dashboards and UI kits
• mobile/web application templates
• prompts and prompt libraries
• public AI assistants/GPT resources where legally indexable
• AI avatars and digital humans
• Voice AI
• STT/TTS and speech technologies
• video, image and media AI tools
• low-code/no-code platforms
• AI application builders
• scraping and crawling tools
• browser automation
• Document AI and OCR
• RAG and knowledge-management systems
• vector databases
• databases and data infrastructure
• API gateways
• APIs and integrations
• plugins and extensions
• developer tools
• DevOps, deployment and hosting tools
• messaging/queue systems
• GitHub repositories and projects
• Hugging Face models, datasets and Spaces
• public datasets
• reusable code/resources
• educational resources
• AI/automation/programming courses
• tutorials, academies and training resources
• productivity and business tools
• marketing/sales tools
• other relevant reusable digital resources discovered during the project.
POTENTIAL SOURCES
Discovery should not be limited to GitHub.
Sources may include GitHub, Hugging Face, public APIs, public datasets, official open-source repositories, software directories, template libraries, automation libraries, AI directories, educational platforms, documentation sites, marketplaces where automated access is permitted, and other legally accessible public sources.
The freelancer should continuously identify additional high-value sources during the project.
TECHNICAL OBJECTIVE
We want to build a scalable pipeline similar to:
DISCOVERY
→ APIs / CRAWLERS / DATASETS
→ RAW DATABASE
→ NORMALIZATION
→ DEDUPLICATION
→ CLASSIFICATION
→ LICENCE DETECTION
→ RIGHTS ANALYSIS
→ METADATA ENRICHMENT
→ QUALITY / RELEVANCE SCORING
→ PRODUCT LIBRARY
Technologies may include Python, APIs, GitHub API, Hugging Face, web crawling/scraping, PostgreSQL, n8n and other appropriate open-source technologies.
The freelancer may propose a better architecture.
DATA TO COLLECT
Where available, each resource should contain:
• name
• category
• subcategory
• description
• source
• source URL
• repository URL
• creator/company/organization
• technology stack
• licence
• licence evidence/source
• commercial-use status
• modification rights
• redistribution rights
• attribution requirements
• white-label/resale rights where applicable
• self-hosting availability
• deployment method
• documentation URL
• pricing/free/paid status
• stars/downloads/popularity
• creation/update dates
• tags
• duplicate/group identifier
• relevance score
• quality score
• activity/maintenance status
• warnings
• date discovered
• date last checked.
LICENSING AND COMMERCIAL RIGHTS ARE CRITICAL
A publicly accessible, free or open-source resource must NOT automatically be classified as commercially resellable.
The database should distinguish between:
• commercial use permitted
• modification permitted
• redistribution permitted
• attribution required
• copyleft requirements
• source-available/restricted
• personal/non-commercial use
• unknown
• manual legal review required.
If the rights cannot be established reliably, the resource must be marked REVIEW REQUIRED rather than approved.
LARGE-SCALE DISCOVERY OBJECTIVE
This is not a small scraping project.
We have already identified sources and collections that may collectively reference hundreds of thousands to several million RAW entries across workflows, agents, skills, MCP servers, prompts, software, applications, templates, SaaS resources, datasets, training resources and other categories.
These numbers are potential RAW discovery volumes, NOT guaranteed unique or commercially usable products.
The architecture must therefore be designed for scale from the beginning.
We want to measure separately:
RAW DISCOVERED
→ UNIQUE
→ RELEVANT
→ LICENCE IDENTIFIED
→ POTENTIALLY COMMERCIALLY USABLE
→ PRIORITY
→ MANUALLY VERIFIED.
The objective is to maximize discovery while maintaining traceability and data quality.
Millions of RAW records are acceptable if the architecture can process them efficiently.
5-MONTH PROJECT
Budget: USD 100/month
Duration: 5 months
Total initial budget: USD 500
MONTH 1 — FOUNDATION
• architecture
• PostgreSQL/database schema
• initial high-value sources
• automated ingestion
• normalization
• basic deduplication
• initial licence detection
• first useful dataset
• scripts/workflows delivered to NEVIE-GLOBAL.
MONTH 2 — SCALE
• expand sources
• improve crawlers/API ingestion
• automated categorization
• improved deduplication
• licence classification
• metadata enrichment
• source traceability
• CSV/JSON exports
• scale toward tens/hundreds of thousands of records where technically achievable.
MONTH 3 — LARGE-SCALE DISCOVERY
• expand GitHub/Hugging Face/public dataset discovery
• agents
• Agent Skills
• MCP servers
• workflows
• SaaS
• applications
• templates
• prompts
• tools
• training resources
• automatic scoring
• continuous discovery.
MONTH 4 — QUALIFICATION & QUALITY
• advanced duplicate reduction
• licence confidence
• commercial-rights classification
• quality/relevance scoring
• abandoned/inactive project detection
• risk/security flags where practical
• identify high-value candidates.
MONTH 5 — CONTINUOUS PRODUCTION SYSTEM
Deliver a documented system capable of continuing after the initial project:
• automated discovery
• scheduled crawling/API updates
• deduplication
• categorization
• licence detection
• enrichment
• scoring
• PostgreSQL database
• CSV/JSON exports
• monitoring/logging
• documentation
• deployment instructions.
The system should continue discovering new resources automatically after the five-month initial project.
DELIVERABLES / OWNERSHIP
NEVIE-GLOBAL must receive the project-specific:
• source code
• scripts
• crawlers
• workflows
• database schemas
• configuration
• documentation
• source registry
• structured datasets generated by the project
• deployment instructions.
Third-party resources remain subject to their respective licences and terms.
Credentials/API keys must never be hard-coded.
The system should respect applicable API limits, robots/access restrictions, platform terms and applicable law.
IMPORTANT
We are NOT looking for someone who simply downloads existing lists or produces a huge CSV full of duplicate URLs.
We are looking for someone capable of building the infrastructure behind a permanent discovery engine.
Quality, uniqueness, traceability and automation matter as much as volume.
The successful freelancer may continue working with NEVIE-GLOBAL after the initial five-month project.
This is NOT a manual data-entry project. We need an automated discovery and data-processing system capable of collecting, structuring, deduplicating, classifying and enriching hundreds of thousands to potentially millions of RAW records from legally accessible public sources.
SCOPE OF RESOURCES TO DISCOVER
The system must progressively discover and catalogue resources including:
• AI agents and multi-agent systems
• AI agent frameworks and autonomous agents
• Agent Skills / AI Skills
• MCP servers, tools and integrations
• n8n workflows and templates
• Make.com, Zapier and Activepieces automations
• workflow libraries and automation systems
• SaaS and micro-SaaS applications
• SaaS boilerplates and starter kits
• open-source software and applications
• AI applications and AI tools
• CRM and ERP systems
• business applications
• website templates
• landing-page templates
• e-commerce templates
• admin dashboards and UI kits
• mobile/web application templates
• prompts and prompt libraries
• public AI assistants/GPT resources where legally indexable
• AI avatars and digital humans
• Voice AI
• STT/TTS and speech technologies
• video, image and media AI tools
• low-code/no-code platforms
• AI application builders
• scraping and crawling tools
• browser automation
• Document AI and OCR
• RAG and knowledge-management systems
• vector databases
• databases and data infrastructure
• API gateways
• APIs and integrations
• plugins and extensions
• developer tools
• DevOps, deployment and hosting tools
• messaging/queue systems
• GitHub repositories and projects
• Hugging Face models, datasets and Spaces
• public datasets
• reusable code/resources
• educational resources
• AI/automation/programming courses
• tutorials, academies and training resources
• productivity and business tools
• marketing/sales tools
• other relevant reusable digital resources discovered during the project.
POTENTIAL SOURCES
Discovery should not be limited to GitHub.
Sources may include GitHub, Hugging Face, public APIs, public datasets, official open-source repositories, software directories, template libraries, automation libraries, AI directories, educational platforms, documentation sites, marketplaces where automated access is permitted, and other legally accessible public sources.
The freelancer should continuously identify additional high-value sources during the project.
TECHNICAL OBJECTIVE
We want to build a scalable pipeline similar to:
DISCOVERY
→ APIs / CRAWLERS / DATASETS
→ RAW DATABASE
→ NORMALIZATION
→ DEDUPLICATION
→ CLASSIFICATION
→ LICENCE DETECTION
→ RIGHTS ANALYSIS
→ METADATA ENRICHMENT
→ QUALITY / RELEVANCE SCORING
→ PRODUCT LIBRARY
Technologies may include Python, APIs, GitHub API, Hugging Face, web crawling/scraping, PostgreSQL, n8n and other appropriate open-source technologies.
The freelancer may propose a better architecture.
DATA TO COLLECT
Where available, each resource should contain:
• name
• category
• subcategory
• description
• source
• source URL
• repository URL
• creator/company/organization
• technology stack
• licence
• licence evidence/source
• commercial-use status
• modification rights
• redistribution rights
• attribution requirements
• white-label/resale rights where applicable
• self-hosting availability
• deployment method
• documentation URL
• pricing/free/paid status
• stars/downloads/popularity
• creation/update dates
• tags
• duplicate/group identifier
• relevance score
• quality score
• activity/maintenance status
• warnings
• date discovered
• date last checked.
LICENSING AND COMMERCIAL RIGHTS ARE CRITICAL
A publicly accessible, free or open-source resource must NOT automatically be classified as commercially resellable.
The database should distinguish between:
• commercial use permitted
• modification permitted
• redistribution permitted
• attribution required
• copyleft requirements
• source-available/restricted
• personal/non-commercial use
• unknown
• manual legal review required.
If the rights cannot be established reliably, the resource must be marked REVIEW REQUIRED rather than approved.
LARGE-SCALE DISCOVERY OBJECTIVE
This is not a small scraping project.
We have already identified sources and collections that may collectively reference hundreds of thousands to several million RAW entries across workflows, agents, skills, MCP servers, prompts, software, applications, templates, SaaS resources, datasets, training resources and other categories.
These numbers are potential RAW discovery volumes, NOT guaranteed unique or commercially usable products.
The architecture must therefore be designed for scale from the beginning.
We want to measure separately:
RAW DISCOVERED
→ UNIQUE
→ RELEVANT
→ LICENCE IDENTIFIED
→ POTENTIALLY COMMERCIALLY USABLE
→ PRIORITY
→ MANUALLY VERIFIED.
The objective is to maximize discovery while maintaining traceability and data quality.
Millions of RAW records are acceptable if the architecture can process them efficiently.
5-MONTH PROJECT
Budget: USD 100/month
Duration: 5 months
Total initial budget: USD 500
MONTH 1 — FOUNDATION
• architecture
• PostgreSQL/database schema
• initial high-value sources
• automated ingestion
• normalization
• basic deduplication
• initial licence detection
• first useful dataset
• scripts/workflows delivered to NEVIE-GLOBAL.
MONTH 2 — SCALE
• expand sources
• improve crawlers/API ingestion
• automated categorization
• improved deduplication
• licence classification
• metadata enrichment
• source traceability
• CSV/JSON exports
• scale toward tens/hundreds of thousands of records where technically achievable.
MONTH 3 — LARGE-SCALE DISCOVERY
• expand GitHub/Hugging Face/public dataset discovery
• agents
• Agent Skills
• MCP servers
• workflows
• SaaS
• applications
• templates
• prompts
• tools
• training resources
• automatic scoring
• continuous discovery.
MONTH 4 — QUALIFICATION & QUALITY
• advanced duplicate reduction
• licence confidence
• commercial-rights classification
• quality/relevance scoring
• abandoned/inactive project detection
• risk/security flags where practical
• identify high-value candidates.
MONTH 5 — CONTINUOUS PRODUCTION SYSTEM
Deliver a documented system capable of continuing after the initial project:
• automated discovery
• scheduled crawling/API updates
• deduplication
• categorization
• licence detection
• enrichment
• scoring
• PostgreSQL database
• CSV/JSON exports
• monitoring/logging
• documentation
• deployment instructions.
The system should continue discovering new resources automatically after the five-month initial project.
DELIVERABLES / OWNERSHIP
NEVIE-GLOBAL must receive the project-specific:
• source code
• scripts
• crawlers
• workflows
• database schemas
• configuration
• documentation
• source registry
• structured datasets generated by the project
• deployment instructions.
Third-party resources remain subject to their respective licences and terms.
Credentials/API keys must never be hard-coded.
The system should respect applicable API limits, robots/access restrictions, platform terms and applicable law.
IMPORTANT
We are NOT looking for someone who simply downloads existing lists or produces a huge CSV full of duplicate URLs.
We are looking for someone capable of building the infrastructure behind a permanent discovery engine.
Quality, uniqueness, traceability and automation matter as much as volume.
The successful freelancer may continue working with NEVIE-GLOBAL after the initial five-month project.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.