Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation
The paper says music representation beats model scale in its controlled text-to-symbolic-music tests.
The authors fixed Qwen3.5 backbones, data, budget, and decoding, then swapped seven tokenizations. Their PMT representation, with 10 ms timing, per-note velocity, and multi-track texture, reached far lower Frechet Music Distance than beat-grid tokenizations, even with a 0.8B model against a 27B beat-grid model. The result also appeared on a small from-scratch model and another performance-resolution tokenizer, which the authors present as evidence for the representation class rather than one vocabulary. Caption adherence remains weak, though a decode-time constraint improved instrument-F1 and correct-key scores without hurting distributional metrics. ArXiv · AI/CL/LG's note
The authors fixed Qwen3.5 backbones, data, budget, and decoding, then swapped seven tokenizations. Their PMT representation, with 10 ms timing, per-note velocity, and multi-track texture, reached far lower Frechet Music Distance than beat-grid tokenizations, even with a 0.8B model against a 27B beat-grid model. The result also appeared on a small from-scratch model and another performance-resolution tokenizer, which the authors present as evidence for the representation class rather than one vocabulary. Caption adherence remains weak, though a decode-time constraint improved instrument-F1 and correct-key scores without hurting distributional metrics. ArXiv · AI/CL/LG's note
score 5