Building a Hinglish Voice Mentor with Gemini and LiveKit
Technical English documentation is often too sterile, creating a significant barrier for learners who process information in Hinglish. Exploring voice‑first interfaces can bridge this gap. After reviewing the architecture of Sydney, a specialized AI/ML mentor, the synergy between low‑latency LLMs and multi‑language STT is essential for effective educational agents.
What was the objective of creating this LLM agent?
The objective was to create a full LLM agent rather than a simple chatbot with a voice skin. This agent can teach complex topics like backpropagation, RAG, and vector embeddings through natural, spoken dialogue.
Technical Stack
To ensure latency remains low enough to avoid the feeling of a 1995 walkie‑talkie, the system follows this specific pipeline:
- STT: Deepgram nova-3. This is vital because it manages real‑time code‑switching between Hindi and English without requiring manual language toggles.
- LLM: Google Gemini 1.5 Flash‑Lite. In voice‑driven workflows, speed is more important than raw reasoning power; if a model takes 3 seconds to process, the user will lose interest.
- TTS: Murf Falcon.
- Orchestration: LiveKit Agents for the real‑time transport layer.
Feature Set and Agent Logic
How does persistent memory enhance the learning experience?
The implementation of agentic workflows elevates this beyond a basic wrapper. This is a system defined by state and tools rather than a linear prompt.
- Persistent Memory: The system tracks learner progress and recurring mistakes so you do not have to explain concepts like vectors multiple times in one session.
- Specialist Handoff: Using a classic LLM agent pattern, the general mentor transfers the session to a RAG Deep‑Dive Specialist or an Interview Prep Specialist once it reaches a complexity ceiling.
- Outbound Capabilities: The AI acts as an active coach by initiating calls for scheduled practice.
- Human Escalation: A built‑in trigger creates a mentor request when the LLM detects a user is genuinely stuck.
Practical Deployment Architecture
What is the data flow organization for similar AI workflows?
User Audio → Deepgram (STT) → Gemini (LLM + Tool Use) → Murf (TTS) → User Audio
External DB / Practice Exercises
The primary hurdle in deployment is not the LLM, but the silence gap. To achieve a human feel, you need an STT that handles interruptions gracefully and an LLM capable of generating concise responses, as long‑winded paragraphs ruin voice agents.
The project also employs a Next.js frontend featuring a custom HTML5 canvas wave‑visualizer. Beyond being eye candy, this provides a necessary visual cue that the agent is listening or thinking, which prevents users from talking over the bot.
What is the core logic flow for the agent?
Example of the core logic flow for the agent
1. Listen for voice activity (VAD)
2. Transcribe via Deepgram nova-3
3. Process intent with Gemini 1.5 Flash
4. If specialized knowledge is needed -> Handoff to Specialist Agent
5. Convert text response to speech via Murf
All Replies (4)
Want a live back-and-forth? Join the global AI chat room — login to talk.
Struggling with the lag here. Which specific LiveKit settings are you using to fix Gemini's latency? I’m experimenting with a pipeline that uses Deepgram nova-3 for STT to handle real-time code-switching, paired with Gemini 1.5 Flash-Lite for speed. Did you find that prioritizing inference speed over raw reasoning power helped reduce the "walkie-talkie" feel, or were there other bottlenecks?
To tackle the lag issue with smaller models during masking, you could implement real-time language switching—like using Deepgram’s Nova-3 for seamless code-switching between languages without manual toggles, just as Sydney’s architecture does. This way, you maintain low latency while ensuring smooth user interaction.
Great tip. Which specific slang terms worked best for your system prompt to keep the flow natural? I’ve been looking at a similar setup where code-switching between Hindi and English is handled by Deepgram nova-3 for STT, since it manages real-time language mixing without manual toggles, and that might be worth trying alongside your prompt tweaks.
Curious about this. How does it handle regional accents across different cities? One concrete step the system takes is using Deepgram nova-3 for speech-to-text, which manages real-time code-switching between Hindi and English without requiring manual language toggles. I have been exploring how voice-first interfaces can bridge gaps in technical education, and reviewing the architecture of Sydney, a specialized AI/ML mentor, makes it evident that the synergy between low-latency LLMs and multi-language STT is essential for effective educational agents.