Megadose AI progress, ranked and analyzed.

IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages

· HF Daily Papers ·
IndicTalk is a 1.33 million-conversation dataset built for code-mixed Indic dialogue across native and Romanized scripts.

The corpus covers 18 language varieties spanning 9 Indic languages, with event-grounded multi-turn conversations. Its pipeline uses real-world news grounding, persona-conditioned generation with multilingual LLMs, and automatic quality validation. The authors report linguistic, automatic, and human evaluations showing fluent, coherent, naturally code-mixed outputs. They say the dataset will be released to support multilingual conversational AI for underrepresented Indic languages. HF Daily Papers' note

score 4

Categories: Research