Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners
NAPE trains audio models by predicting the next spectrogram patch embedding, with no decoder, tokenizer, teacher model, or extra losses.
The paper frames that causal setup as a simpler self-supervised route for audio representation learning. A causal Transformer learns from prior log-mel spectrogram patches using causal masking and a stop-gradient signal. The authors report state-of-the-art fine-tuning results on several of six audio and speech benchmarks, consistent scaling across encoder sizes, and strong linear probes. Source: HF Daily Papers' note
The paper frames that causal setup as a simpler self-supervised route for audio representation learning. A causal Transformer learns from prior log-mel spectrogram patches using causal masking and a stop-gradient signal. The authors report state-of-the-art fine-tuning results on several of six audio and speech benchmarks, consistent scaling across encoder sizes, and strong linear probes. Source: HF Daily Papers' note
score 4