Megadose AI progress, ranked and analyzed.

IndicTriMix: Developing Language Identification Datasets and Models for Tri-Language Code-Mixing

· ArXiv · AI/CL/LG ·
The paper releases a benchmark and fine-tuned models for token-level language ID in three-way Indic code-mixed text.

The authors frame language identification as sequence labeling for utterances mixing Hindi, Gujarati, and Bengali. They fine-tune MuRIL and XLM-RoBERTa, evaluating them across three data configurations with manually annotated test sets. The work also describes two ways to generate tri-language code-mixed data from parallel sentences. Models and benchmark data are released for reproducibility. ArXiv · AI/CL/LG's note

score 4

Categories: Research