Megadose AI progress, ranked and analyzed.

FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models

· HF Daily Papers ·
The paper’s claim is that n-gram memory works better when it is broken into shared components and gated per component.

FactorEngram replaces monolithic retrieved embeddings with sparse coefficients over a shared basis-vector dictionary. Context from the model backbone gates those coefficients one by one before reconstruction, so the same token pattern can use different memory components in different contexts. The authors report gains on 340M- and 1B-parameter Transformer backbones across language modeling and downstream tasks. Their ablations point to inserting the memory branch before attention in the middle layers as the strongest setup. HF Daily Papers' note

score 4

Categories: Research