Megadose Built for builders and researchers.

WNet: Discrete Wavelets Transform for Efficient Token Mixing

· ArXiv · AI/CL/LG ·
WNet swaps self-attention for wavelet-based token mixing to cut long-sequence encoder cost.

The paper tests three attention-free wavelet mixers, plus a hybrid that keeps self-attention only in the final layer. It says short two-tap filters such as Haar stay trapped in fixed token blocks, while longer filters can reach the full sequence within two layers. In controlled pre-training and GLUE fine-tuning against BERT, FNet, and a no-mixing control, the token-gated mixer matched attention speed at 256 tokens and ran 2.7x faster at 4,096. ArXiv · AI/CL/LG's note

score 5

Categories: Research