Ask Self, Ask Others: Relation Is All You Need
The paper proposes Relation as a token-mixing replacement for attention, with lower validation NLL than MHA in matched small decoder models.
Relation separates pairwise evidence into Self and Exchange relations before deriving information flow. The authors report Full Relation beating multi-head attention at roughly 10M, 30M, and 100M parameters. FlashRelation runs 3.60-4.41x faster than the materialized Full Relation baseline in a fixed-context benchmark. In production-style matched workloads, it reaches 76.4-84.9% of PyTorch FlashAttention throughput while executing the Full Relation operator. ArXiv · AI/CL/LG's note
Relation separates pairwise evidence into Self and Exchange relations before deriving information flow. The authors report Full Relation beating multi-head attention at roughly 10M, 30M, and 100M parameters. FlashRelation runs 3.60-4.41x faster than the materialized Full Relation baseline in a fixed-context benchmark. In production-style matched workloads, it reaches 76.4-84.9% of PyTorch FlashAttention throughput while executing the Full Relation operator. ArXiv · AI/CL/LG's note
score 6