Scaled Idempotence in Transformer Attention: Paired OV Geometry and Shared-Value Algebras
A small minority of attention heads appear to learn near-idempotent OV operators, and the paper argues that orientation, not just spectrum or span, is doing the work.
Feng and Li report that 3.98% to 8.00% of heads across six pretrained endpoints satisfy a high squared-closure test, while matched O/V mismatches do not. Across 7,304 heads in nine models, scrambling the orientation of the central factor while preserving other geometry sharply reduces closure, with trained orientation winning for 98.64% of heads. The paper also says exact value sharing extends the headwise relation into a local right-action algebra, with experiments showing approximate versions of that law. ArXiv · AI/CL/LG's note
Feng and Li report that 3.98% to 8.00% of heads across six pretrained endpoints satisfy a high squared-closure test, while matched O/V mismatches do not. Across 7,304 heads in nine models, scrambling the orientation of the central factor while preserving other geometry sharply reduces closure, with trained orientation winning for 98.64% of heads. The paper also says exact value sharing extends the headwise relation into a local right-action algebra, with experiments showing approximate versions of that law. ArXiv · AI/CL/LG's note
score 5