Megadose AI progress, ranked and analyzed.

GigaChat Audio: Time-aware Large Audio Language Model

· HF Daily Papers ·
The model is built to answer questions with explicit timestamps across recordings up to 120 minutes long.

GigaChat Audio interleaves periodic time markers with continuous audio tokens to improve temporal grounding. The training uses large-scale synthetic supervision from a cascaded pipeline. The authors report strong timestamp accuracy on both short and long benchmarks, plus support for time-anchored fragment descriptions and summaries. They also release model weights and datasets for further research. HF Daily Papers' note

score 6

Categories: Model Releases, OSS & Tools, Research