Building Roshni: A Real-Time, Multi-Agent Financial Voice AI for Bharat ๐ฎ๐ณ
Building Roshni: An Ultra-Low Latency, Multi-Agent Financial Voice Assistant for Bharat ๐ฎ๐ณ How I built an end-to-end, multilingual financial voice AI using Murf Falcon, LiveKit Agents, Deepgram Nova-3, Google Gemini, and Next.js during the 10 Days of AI Voice Agents Challenge. ๐ 1. The Problem & Why Voice Matters for Bharat In India, financial inclusion has accelerated rapidly with UPI, digital banking, and government-backed credit initiatives. However, navigating complex interest rates, eligibility criteria for government schemes (like PM Mudra or Sukanya Samriddhi Yojana), and understanding formal banking terms remains intimidating for millions of citizensโespecially in regional and tier-2/3 heartlands where digital interfaces can be overwhelming. Text-first interfaces fail where voice thrives. When rural entrepreneurs or first-time bank customers have questions, they don't want to navigate complex web forms or read dense PDFs. They want to ask a direct question in their language and get an immediate, clear, spoken answer. To solve this, I built Roshni AI (and her specialist counterpart, Vikram ) โ an ultra-low latency, conversational financial assistant engineered for natural voice interactions in English, Hindi (Devanagari script), and Hinglish. ๐๏ธ 2. High-Level Architecture & Tech Stack Building a real-time conversational agent requires synchronizing four core pipelines with sub-second latency: [ ๐ค User Microphone ] โ (WebRTC Audio Stream) โผ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ LiveKit Agents Worker โ โโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโ โ โโโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโ โผ โผ โผ โโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโ โ Deepgram โ โโโโโบ โGoogle Geminiโ โโโโโบ โ Murf Falcon โ โ Nova-3 โ โ (LLM) โ โ Fast TTS โ โ (Fast STT) โ โ โ โ (Anisha / Samar)โ โโโโโโโโโโโโโโโ โโโโโโโโฌโโโโโโโ โโโโโโโโโโฌโโโโโโโโโ โ (Tool / Handoff) โ โผ โผ โโโโโโโโโโโโโโโโโ [ ๐ Audio Output ] โ SQLite Memory โ โ & Analytics โ โโโโโโโโโโโโโโโโโ The Stack: TTS (Text-to-Speech):