Refine Chatbot for Hospitality Environment
Budget / Salary$10–30
TypeFreelance project
LocationRemote
Posted1 hour ago
I already have a working hospitality-focused chatbot; what I need now is an AI engineer who can fine-tune it so that it answers guest FAQs flawlessly in English. The model is in production inside a booking engine, yet its responses still feel generic and occasionally miss the subtle context of hotel terminology. Your job is to refine the underlying large-language-model (we are on GPT-4o-mini via an API) so the bot sounds more like an attentive concierge than a scripted agent. Also sometimes it starts hallucinating. which also needs to be fixed.
Scope of work
* Review the current prompt structure, conversation logs, and training set (≈15 k dialogue pairs).
* Design and run a fine-tuning pipeline—feel free to use OpenAI fine-tuning tools, LangChain, or your preferred framework—as long as it can be reproduced in our environment (Python 3.10).
* Optimise for: accuracy on policy-compliant answers, reduced hallucinations, and fast first-token latency.
* Add robust fallback logic for questions falling outside scope instead of defaulting to “I’m sorry…”.
* Deliver an updated model with confidence scores and a short README so my dev team can deploy it directly.
Acceptance criteria
1. ≥95 % correct answer rate on our 300-question English FAQ benchmark.
2. Average response time under 1.5 s on a t3.medium test node.
3. No policy violations across 500 random adversarial prompts.
If this sounds like your specialty, tell me which fine-tuning approach you prefer and a quick outline of how you would validate improvements.
Scope of work
* Review the current prompt structure, conversation logs, and training set (≈15 k dialogue pairs).
* Design and run a fine-tuning pipeline—feel free to use OpenAI fine-tuning tools, LangChain, or your preferred framework—as long as it can be reproduced in our environment (Python 3.10).
* Optimise for: accuracy on policy-compliant answers, reduced hallucinations, and fast first-token latency.
* Add robust fallback logic for questions falling outside scope instead of defaulting to “I’m sorry…”.
* Deliver an updated model with confidence scores and a short README so my dev team can deploy it directly.
Acceptance criteria
1. ≥95 % correct answer rate on our 300-question English FAQ benchmark.
2. Average response time under 1.5 s on a t3.medium test node.
3. No policy violations across 500 random adversarial prompts.
If this sounds like your specialty, tell me which fine-tuning approach you prefer and a quick outline of how you would validate improvements.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.