WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing
WhiteMatter lets each Transformer layer reuse past-token representations from any depth, not just its own.
A learned mixer chooses which layer depths matter for the current context and folds them into shared KV cache channels. In the paper’s reported runs, a full-cache WhiteMatter model matches a standard Transformer with 50% more layers, while a half-cache version beats matched baselines at scales up to 1.3B parameters. The tradeoff is extra cross-layer dependency during training and prompt processing, which the authors address with cyclic iteration. On their reference model, that iteration method converges 12.5x faster than standard Jacobi iteration. HF Daily Papers' note
A learned mixer chooses which layer depths matter for the current context and folds them into shared KV cache channels. In the paper’s reported runs, a full-cache WhiteMatter model matches a standard Transformer with 50% more layers, while a half-cache version beats matched baselines at scales up to 1.3B parameters. The tradeoff is extra cross-layer dependency during training and prompt processing, which the authors address with cyclic iteration. On their reference model, that iteration method converges 12.5x faster than standard Jacobi iteration. HF Daily Papers' note
score 5