Cross-Model Memory Transfer via Target-Side Reader Adaptation
The paper tests whether an Engram-style memory table can move between language-model backbones if the new model gets its own reader.
The authors freeze a memory trained on a source model, attach it to a different target model, and train only a lightweight target-side reader. Their ablations say the stored memory and address matching both matter, but the transfer works only when the reader is aligned with the target model. In question-answering tests, a dual-layer, four-branch reader nearly closes the same-model versus cross-model reuse gap, averaging 38.8 under their protocol. When the original reader already fits the target interface, the frozen memory can help without target-side training, with adaptation adding more. HF Daily Papers' note
The authors freeze a memory trained on a source model, attach it to a different target model, and train only a lightweight target-side reader. Their ablations say the stored memory and address matching both matter, but the transfer works only when the reader is aligned with the target model. In question-answering tests, a dual-layer, four-branch reader nearly closes the same-model versus cross-model reuse gap, averaging 38.8 under their protocol. When the original reader already fits the target interface, the frozen memory can help without target-side training, with adaptation adding more. HF Daily Papers' note
score 4