Megadose Built for builders and researchers.

GTR: Gated Token Recurrence for Efficient Dense Prediction

· HF Daily Papers ·
GTR replaces global softmax attention with recurrent token mixing aimed at faster high-resolution dense prediction.

The paper reports 58.9 box AP on COCO `val2017` for GTR-L after Objects365 detector pre-training, with 1.908 ms median batch-one latency on an RTX 4090. It says the backbone transfers across segmentation, pose, oriented detection, semantic segmentation, and monocular depth. A specialized chunkwise CUDA operator is reported as 4.0x faster than FLA v0.5.0 at 1.6K tokens. TensorRT deployment on DRIVE AGX Thor is listed at 2.282-8.769 ms median batch-one latency across evaluated models. HF Daily Papers' note

score 5

Categories: Research