arXiv · 2607.23242
IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages
Abstract
Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms. We present IndicTalk, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversations across 18 language varieties covering 9 Indic languages. The corpus is generated through a fully automated pipeline that combines real-world news grounding, persona-conditioned dialogue generation using multilingual LLMs, and automatic quality validation. Extensive linguistic, automatic, and human evaluations demonstrate that IndicTalk produces fluent, coherent, and naturally code-mixed conversations across both script variants. We will release IndicTalk to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages. The dataset is available at: https://huggingface.co/datasets/LingoIITGN/IndicTalk .
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sahil Deepak Gawande, Mayank Singh. 2026-07-25. IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages. https://arxiv.org/abs/2607.23242
Cite the original work for its findings. Save a collection to share your selection of sources.