Megadose AI progress, ranked and analyzed.

Mizar: A 159M-Parameter Audio-Language Model for Audio Understanding

· ArXiv · AI/CL/LG ·
Mizar is a sub-200M-parameter audio-language model built to run locally, with reported CPU inference around 1.09 seconds per MMAU question.

The paper describes a 159.3M-parameter model that pairs a CED-Small audio encoder with SmolLM2-135M through a frequency-merging mapper. Training is split into alignment, audio-dependent fine-tuning, and post-training meant to improve weak skills without losing prior gains. Across five seeds, it reports 52.92% on MMAU, 42.42% on MMAR, and 36.02% on ADQA-clean, ahead of the prior best ALM under 200M parameters on those benchmarks. Code and checkpoints are listed as available. ArXiv · AI/CL/LG's note

score 4

Categories: Research