Retrieval-Centric Deep Learning in Growing Nonparametric Neural Networks
The paper proposes neural layers that keep growing by storing training examples for retrieval at inference time.
Instead of folding all training data into fixed-size weights, the method stores key-value representations for each data point and recombines them with attention. The authors argue that simply extending earlier linear-attention learning rules to stronger kernels is not principled, then derive functional-gradient rules for RBF and softmax-like kernel attention. They report promising results on image classification and synthetic teacher-student tasks, and connect advanced linear-attention variants to optimizers for conventional fixed-size neural nets. HF Daily Papers' note
Instead of folding all training data into fixed-size weights, the method stores key-value representations for each data point and recombines them with attention. The authors argue that simply extending earlier linear-attention learning rules to stronger kernels is not principled, then derive functional-gradient rules for RBF and softmax-like kernel attention. They report promising results on image classification and synthetic teacher-student tasks, and connect advanced linear-attention variants to optimizers for conventional fixed-size neural nets. HF Daily Papers' note
score 4