A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books
Grammar books can be turned into fine-tuning data for machine translation in endangered languages.
The paper describes an LLM pipeline that extracts rules, examples, and lexicons from descriptive grammars, then builds synthetic parallel corpora. It tests the method on Kalamang, Tuatschin, and Mandan, reporting improvements over seed-data baselines in most Kalamang and Tuatschin settings and smaller best-case gains for Mandan. The study compares 96 configurations to show which extraction and sampling choices help, and where the approach fails. ArXiv · AI/CL/LG's note
The paper describes an LLM pipeline that extracts rules, examples, and lexicons from descriptive grammars, then builds synthetic parallel corpora. It tests the method on Kalamang, Tuatschin, and Mandan, reporting improvements over seed-data baselines in most Kalamang and Tuatschin settings and smaller best-case gains for Mandan. The study compares 96 configurations to show which extraction and sampling choices help, and where the approach fails. ArXiv · AI/CL/LG's note
score 4