Megadose Built for builders and researchers.

Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners

· HF Daily Papers ·
NAPE trains audio models by predicting the next spectrogram patch embedding, with no decoder, tokenizer, teacher model, or extra losses.

The paper frames that causal setup as a simpler self-supervised route for audio representation learning. A causal Transformer learns from prior log-mel spectrogram patches using causal masking and a stop-gradient signal. The authors report state-of-the-art fine-tuning results on several of six audio and speech benchmarks, consistent scaling across encoder sizes, and strong linear probes. Source: HF Daily Papers' note

score 4

Categories: Research