Nonparametric In-Context Learning under Growing Geometric Complexity: Minimax Optimality and Local Geometry-Adaptivity of Transformers
The paper claims transformers can match minimax-optimal nonparametric prediction rates when data geometry varies locally across manifold mixtures.
Seo and Kim model in-context prediction under unknown local geometry, with components that differ in dimension, smoothness, and sampling mass. They prove a minimax lower bound and match it with an oracle tangent local-polynomial estimator. They then connect that estimator to a two-stage softmax transformer using geometric preconditioning and chartwise solvers, with negligible approximation error at the target rate. The 63-page paper was accepted at NeurIPS 2026. ArXiv · AI/CL/LG's note
Seo and Kim model in-context prediction under unknown local geometry, with components that differ in dimension, smoothness, and sampling mass. They prove a minimax lower bound and match it with an oracle tangent local-polynomial estimator. They then connect that estimator to a two-stage softmax transformer using geometric preconditioning and chartwise solvers, with negligible approximation error at the target rate. The 63-page paper was accepted at NeurIPS 2026. ArXiv · AI/CL/LG's note
score 5