Python Engineer for LLM Content
Budget / Salary$250–750
TypeFreelance project
LocationRemote
Posted2 hours ago
I’m spinning up a production-grade content-generation platform and need an experienced Python engineer to own the build from model integration through cloud deployment. The core of the system will combine Agentic AI workflows with Retrieval-Augmented Generation, all running on Hugging Face Transformers and exposed through well-designed APIs.
Here’s the flow I have in mind. A user submits a brief, the system retrieves supporting knowledge from a vector store, an agent chain crafts the prompt, and an LLM returns polished copy. Everything must be containerised, orchestrated as microservices, and pushed to Microsoft Azure—ideally AKS with CI/CD in place. FastAPI, async Python, Docker and Terraform/Bicep (or an equivalent infra-as-code tool) should feel second nature to you.
Key responsibilities include:
• Architecting the Python microservice layout with clear domain boundaries
• Integrating and, when useful, fine-tuning Hugging Face models (LoRA, PEFT, etc.)
• Implementing a robust evaluation harness for automated and human-in-loop scoring
• Building REST/WebSocket endpoints plus authentication and rate limiting
• Automating deployment to Azure, complete with monitoring, logging and autoscaling
Acceptance criteria
1. One-click (or single-command) Azure deployment producing a live endpoint
2. Automated tests covering critical paths with ≥90 % pass rate in CI
3. Median response time under 800 ms for a 200-token request
4. Clear, developer-friendly README and API contract
Technical Details & Architecture Overview
We are building a production-grade Agentic AI content generation platform using Python, FastAPI, Hugging Face Transformers, and Azure cloud services. The initial architecture will support enterprise LLM workflows with flexibility to evaluate and integrate different open-source instruction-tuned models such as Llama, Mistral, or similar Hugging Face model families based on performance, cost, and deployment requirements.
The RAG pipeline will leverage Azure AI Search with vector search capabilities as the preferred vector store, supporting semantic search and hybrid retrieval (BM25 + embeddings) for improved accuracy. Alternatives such as Pinecone or pgvector may be evaluated depending on scalability and operational needs.
The agentic workflow layer will support tool/function calling, retrieval workflows, prompt orchestration, content validation, and human-in-the-loop review processes using frameworks such as LangGraph, LangChain, and LlamaIndex.
The platform will expose secure REST APIs with OAuth2/OIDC and JWT-based authentication, along with API rate limiting and enterprise security controls. Deployment will be containerized using Docker and orchestrated through Azure Kubernetes Service (AKS) with infrastructure managed through Terraform or Bicep.
The evaluation framework will include automated and human review capabilities to measure response quality, retrieval accuracy, groundedness, hallucination detection, and LLM output performance using tools such as RAGAS, LangSmith, and custom evaluation pipelines.
The initial latency goal targets optimization of the application pipeline, including API processing, retrieval, agent orchestration, and prompt generation, while supporting streaming responses for longer LLM generations. The engineering team will also own knowledge ingestion, embedding pipelines, monitoring, logging, CI/CD automation, and production reliability.
If you’ve shipped LLM-powered products at scale and can demonstrate clean code, thoughtful prompt engineering and reliable Azure pipelines, let’s talk.
Here’s the flow I have in mind. A user submits a brief, the system retrieves supporting knowledge from a vector store, an agent chain crafts the prompt, and an LLM returns polished copy. Everything must be containerised, orchestrated as microservices, and pushed to Microsoft Azure—ideally AKS with CI/CD in place. FastAPI, async Python, Docker and Terraform/Bicep (or an equivalent infra-as-code tool) should feel second nature to you.
Key responsibilities include:
• Architecting the Python microservice layout with clear domain boundaries
• Integrating and, when useful, fine-tuning Hugging Face models (LoRA, PEFT, etc.)
• Implementing a robust evaluation harness for automated and human-in-loop scoring
• Building REST/WebSocket endpoints plus authentication and rate limiting
• Automating deployment to Azure, complete with monitoring, logging and autoscaling
Acceptance criteria
1. One-click (or single-command) Azure deployment producing a live endpoint
2. Automated tests covering critical paths with ≥90 % pass rate in CI
3. Median response time under 800 ms for a 200-token request
4. Clear, developer-friendly README and API contract
Technical Details & Architecture Overview
We are building a production-grade Agentic AI content generation platform using Python, FastAPI, Hugging Face Transformers, and Azure cloud services. The initial architecture will support enterprise LLM workflows with flexibility to evaluate and integrate different open-source instruction-tuned models such as Llama, Mistral, or similar Hugging Face model families based on performance, cost, and deployment requirements.
The RAG pipeline will leverage Azure AI Search with vector search capabilities as the preferred vector store, supporting semantic search and hybrid retrieval (BM25 + embeddings) for improved accuracy. Alternatives such as Pinecone or pgvector may be evaluated depending on scalability and operational needs.
The agentic workflow layer will support tool/function calling, retrieval workflows, prompt orchestration, content validation, and human-in-the-loop review processes using frameworks such as LangGraph, LangChain, and LlamaIndex.
The platform will expose secure REST APIs with OAuth2/OIDC and JWT-based authentication, along with API rate limiting and enterprise security controls. Deployment will be containerized using Docker and orchestrated through Azure Kubernetes Service (AKS) with infrastructure managed through Terraform or Bicep.
The evaluation framework will include automated and human review capabilities to measure response quality, retrieval accuracy, groundedness, hallucination detection, and LLM output performance using tools such as RAGAS, LangSmith, and custom evaluation pipelines.
The initial latency goal targets optimization of the application pipeline, including API processing, retrieval, agent orchestration, and prompt generation, while supporting streaming responses for longer LLM generations. The engineering team will also own knowledge ingestion, embedding pipelines, monitoring, logging, CI/CD automation, and production reliability.
If you’ve shipped LLM-powered products at scale and can demonstrate clean code, thoughtful prompt engineering and reliable Azure pipelines, let’s talk.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.