RKNN Multilingual Voice Recognition
Budget / Salary$250–750
TypeFreelance project
LocationRemote
Posted1 hour ago
I need a complete voice-to-text pipeline that not only recognises speech but also detects the spoken language on the fly, all running natively on the Rockchip NPU through RKNN. The target device is Android only, so everything—from the model optimisation to the demo app—must be tuned for that environment.
Languages to be auto-identified and transcribed: English, Mandarin, Thai, Cantonese, Japanese, Korean, Hokkien and Malay. I am aiming for high identification accuracy; false detections must be the rare exception, not the rule.
You are free to start from TensorFlow, PyTorch or ONNX models as long as you convert and fine-tune them with the RKNN Toolkit so that real-time performance is achieved on typical Rockchip boards (e.g., RK356x or similar). Quantisation awareness, mixed-precision tricks and any NPU-specific optimisations are all welcome, but latency must remain low enough for conversational use.
Deliverables
• An optimised RKNN model capable of streaming inference
• An Android demo project (Java/Kotlin or C++ with the NDK) that records audio, detects the language, and outputs the transcription in UTF-8
• Clear build & integration notes so my team can reproduce results on fresh hardware
• A short benchmark report showing word-error rate and language-ID accuracy on our provided test set
The project is complete once the demo app hits the required accuracy thresholds and runs in real time on the reference Rockchip board.
Languages to be auto-identified and transcribed: English, Mandarin, Thai, Cantonese, Japanese, Korean, Hokkien and Malay. I am aiming for high identification accuracy; false detections must be the rare exception, not the rule.
You are free to start from TensorFlow, PyTorch or ONNX models as long as you convert and fine-tune them with the RKNN Toolkit so that real-time performance is achieved on typical Rockchip boards (e.g., RK356x or similar). Quantisation awareness, mixed-precision tricks and any NPU-specific optimisations are all welcome, but latency must remain low enough for conversational use.
Deliverables
• An optimised RKNN model capable of streaming inference
• An Android demo project (Java/Kotlin or C++ with the NDK) that records audio, detects the language, and outputs the transcription in UTF-8
• Clear build & integration notes so my team can reproduce results on fresh hardware
• A short benchmark report showing word-error rate and language-ID accuracy on our provided test set
The project is complete once the demo app hits the required accuracy thresholds and runs in real time on the reference Rockchip board.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.