Balancing Memory Pathways: Analyzing and Improving Memory Utilization in Hybrid LMs
Hybrid models had the memory paths, but leaned too hard on attention.
The paper says recurrent-attention LMs do not naturally coordinate attention and recurrent state well. Standard supervised fine-tuning raised performance while making the models even more dependent on attention, leaving recurrent memory underused. The authors added an auxiliary loss that restricts attention’s earlier-context access while letting recurrent state carry the full sequence, pushing the model to use both paths. They report better overall results, especially on longer-context and information-aggregation tasks, across multiple hybrid models and related memory setups. ArXiv · AI/CL/LG's note
The paper says recurrent-attention LMs do not naturally coordinate attention and recurrent state well. Standard supervised fine-tuning raised performance while making the models even more dependent on attention, leaving recurrent memory underused. The authors added an auxiliary loss that restricts attention’s earlier-context access while letting recurrent state carry the full sequence, pushing the model to use both paths. They report better overall results, especially on longer-context and information-aggregation tasks, across multiple hybrid models and related memory setups. ArXiv · AI/CL/LG's note
score 5