Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference
Daedalus-150M replaces most attention blocks with short convolutions to keep CPU decoding from slowing as context grows.
The 150M-parameter model uses full attention in 6 of 18 blocks, with the other 12 holding only a two-timestep convolution memory. Trained from scratch on 59.9B tokens, it beat several similarly sized older models and cleared the paper’s preset benchmark bar. Against a same-size all-attention baseline on the same data, the hybrid matched downstream task quality while producing a smaller 4-bit file and decoding 1.76x faster at 2048 tokens. The paper also reports failures: 4-bit quality loss, inert convolution channels that could not be removed, and an oversized vocabulary. ArXiv · AI/CL/LG's note
The 150M-parameter model uses full attention in 6 of 18 blocks, with the other 12 holding only a two-timestep convolution memory. Trained from scratch on 59.9B tokens, it beat several similarly sized older models and cleared the paper’s preset benchmark bar. Against a same-size all-attention baseline on the same data, the hybrid matched downstream task quality while producing a smaller 4-bit file and decoding 1.76x faster at 2048 tokens. The paper also reports failures: 4-bit quality loss, inert convolution channels that could not be removed, and an oversized vocabulary. ArXiv · AI/CL/LG's note
score 5