From Attention Masks to Inert Zero-Vector Tokens: OAttention and O-Closure for Token Dynamics
The paper proposes a token-level “presence” state that makes zero-vector tokens inert inside attention and related transformer components.
Instead of using attention masks only to block query-source pairs, the method assigns each token a coefficient based on its hidden-vector norm. That coefficient gates what a token emits and how much it counts in shared computations. The paper claims exact null-token behavior for OAttention and extends the same rule through OFFN, ONorm, OInject, OStandardize, and an OTransformer closure law. In a zero-fine-tuning TabPFN v3 retrofit, the reported mean RMSE shifts were +0.088% for calibrated OAttention and +0.177% for Full-O across 18 matched dataset-seed cases. The author frames the tests as scoped, not proof of universal no-loss behavior or a general missing-value semantics. ArXiv · AI/CL/LG's note
Instead of using attention masks only to block query-source pairs, the method assigns each token a coefficient based on its hidden-vector norm. That coefficient gates what a token emits and how much it counts in shared computations. The paper claims exact null-token behavior for OAttention and extends the same rule through OFFN, ONorm, OInject, OStandardize, and an OTransformer closure law. In a zero-fine-tuning TabPFN v3 retrofit, the reported mean RMSE shifts were +0.088% for calibrated OAttention and +0.177% for Full-O across 18 matched dataset-seed cases. The author frames the tests as scoped, not proof of universal no-loss behavior or a general missing-value semantics. ArXiv · AI/CL/LG's note
score 4