Megadose Built for builders and researchers.

Ask Self, Ask Others: Relation Is All You Need

· ArXiv · AI/CL/LG ·
The paper proposes Relation as a token-mixing replacement for attention, with lower validation NLL than MHA in matched small decoder models.

Relation separates pairwise evidence into Self and Exchange relations before deriving information flow. The authors report Full Relation beating multi-head attention at roughly 10M, 30M, and 100M parameters. FlashRelation runs 3.60-4.41x faster than the materialized Full Relation baseline in a fixed-context benchmark. In production-style matched workloads, it reaches 76.4-84.9% of PyTorch FlashAttention throughput while executing the Full Relation operator. ArXiv · AI/CL/LG's note

score 6

Categories: Research