Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention
CKDA widens Kimi Delta Attention’s parameter ranges so a single efficient transition can represent 2D rotations.
The paper says this comes from allowing gates in `[-1,1]` and the delta-rule coefficient beta in `[0,2]`. The authors argue CKDA keeps KDA’s diagonal-plus-rank-one, non-expansive transition structure while matching DeltaProduct₂ on state-tracking expressivity. They prove every orthogonal diagonal-plus-rank-one matrix is exactly a CKDA transition matrix. In tests, the combined extension gave the strongest length extrapolation among KDA range settings on group tasks and periodic audio continuation, while language-modeling results were similar to KDA and ahead of the other baselines they tested. ArXiv · AI/CL/LG's note
The paper says this comes from allowing gates in `[-1,1]` and the delta-rule coefficient beta in `[0,2]`. The authors argue CKDA keeps KDA’s diagonal-plus-rank-one, non-expansive transition structure while matching DeltaProduct₂ on state-tracking expressivity. They prove every orthogonal diagonal-plus-rank-one matrix is exactly a CKDA transition matrix. In tests, the combined extension gave the strongest length extrapolation among KDA range settings on group tasks and periodic audio continuation, while language-modeling results were similar to KDA and ahead of the other baselines they tested. ArXiv · AI/CL/LG's note
score 5