Dynamic Compression in Recurrent Networks
The paper tests a recurrent model that can re-scan past tokens when a later task makes them relevant.
Instead of forcing every input into a fixed-size state on the first pass, the proposed dynamic compression method lets the model revise that state with extra recurrent updates. In the authors’ controlled setup, the model first learns multiple functions in context, then has to identify and reuse one of them in later few-shot tasks. Selective re-scanning reduced the state size needed for accurate reuse and scaled better as more functions were stored. ArXiv · AI/CL/LG's note
Instead of forcing every input into a fixed-size state on the first pass, the proposed dynamic compression method lets the model revise that state with extra recurrent updates. In the authors’ controlled setup, the model first learns multiple functions in context, then has to identify and reuse one of them in later few-shot tasks. Selective re-scanning reduced the state size needed for accurate reuse and scaled better as more functions were stored. ArXiv · AI/CL/LG's note
score 5