Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference
Daedalus-150M was built around CPU decoding first, replacing most attention blocks with short convolutions to keep context growth cheap.
The 150M-parameter model uses full attention in 6 of 18 blocks, with the other 12 using two-timestep convolution memory. Trained from scratch on 59.9B tokens, it beat several similarly sized older models and cleared its pre-set benchmark target. Against a same-size all-attention control trained on the same data, the hybrid decoded 1.76x faster at 2048 tokens while matching downstream task results. The paper also reports limits: 4-bit quality loss, inert convolution channels, and an oversized vocabulary. HF Daily Papers' note
The 150M-parameter model uses full attention in 6 of 18 blocks, with the other 12 using two-timestep convolution memory. Trained from scratch on 59.9B tokens, it beat several similarly sized older models and cleared its pre-set benchmark target. Against a same-size all-attention control trained on the same data, the hybrid decoded 1.76x faster at 2048 tokens while matching downstream task results. The paper also reports limits: 4-bit quality loss, inert convolution channels, and an oversized vocabulary. HF Daily Papers' note
score 5