Megadose AI progress, ranked daily.

How Much Rank Does LoRA Need? Rank-Error Bounds for Transformer Attention

· ArXiv · AI/CL/LG ·
The paper gives LoRA rank selection a task-dependent error bound for Transformer attention.

It fixes a pretrained attention head, a target attention function, and a downstream input distribution, then bounds the best expected KL error from a rank-r query LoRA update. The bounds are expressed through score differences and the downstream-weighted tail energy of the target update. It also treats cases involving target-Fisher bounds, probability mass concentrated on token subsets, softmax saturation, fused multi-head LoRA, and joint query/key updates. ArXiv · AI/CL/LG's note

score 5

Categories: Research