Megadose AI progress, ranked and analyzed.

GTR: Gated Token Recurrence for Efficient Dense Prediction

· ArXiv · AI/CL/LG ·
GTR replaces global softmax attention with recurrent token mixing aimed at faster high-resolution dense prediction.

The paper presents a softmax-free vision backbone using gated linear attention, alternating spatial scans, and spatially enhanced SwiGLU blocks. It is distilled from a detection-specialized DINOv3 teacher with final-layer patch-token alignment only. With Objects365 detector pre-training, GTR-L reports 58.9 box AP on COCO val2017 at 1.908 ms median batch-one latency on an RTX 4090. The authors also report transfer across segmentation, pose, oriented detection, semantic segmentation, and monocular depth tasks, plus faster CUDA kernel performance than FLA v0.5.0 in their isolated benchmark. ArXiv · AI/CL/LG's note

score 5

Categories: Research