Megadose Built for builders and researchers.

Revisiting Input Time-frequency Representations in Multi-pitch Estimation for Vocal Ensembles

· ArXiv · AI/CL/LG ·
Linear STFT beat HCQT for vocal-ensemble multi-pitch estimation while cutting feature-extraction cost.

Koh and Dong compare the commonly used HCQT input representation with a direct linear STFT input for models estimating multiple simultaneous vocal pitches. The STFT approach performs better despite lacking a pitch-aligned grid and fixed frequency resolution. Their analysis finds no added gain from longer windows or broader spectral coverage, and says limiting input to the predicted pitch range weakens STFT’s edge. ArXiv · AI/CL/LG's note

score 4

Categories: Research