IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages
IndicTalk is a 1.33 million-conversation dataset built for code-mixed Indic dialogue across native and Romanized scripts.
The corpus covers 18 language varieties spanning 9 Indic languages, with event-grounded multi-turn conversations. Its pipeline uses real-world news grounding, persona-conditioned generation with multilingual LLMs, and automatic quality validation. The authors report linguistic, automatic, and human evaluations showing fluent, coherent, naturally code-mixed outputs. They say the dataset will be released to support multilingual conversational AI for underrepresented Indic languages. HF Daily Papers' note
The corpus covers 18 language varieties spanning 9 Indic languages, with event-grounded multi-turn conversations. Its pipeline uses real-world news grounding, persona-conditioned generation with multilingual LLMs, and automatic quality validation. The authors report linguistic, automatic, and human evaluations showing fluent, coherent, naturally code-mixed outputs. They say the dataset will be released to support multilingual conversational AI for underrepresented Indic languages. HF Daily Papers' note
score 4