FabriMAE I Trust Myself? Self-Evaluating VLA Action Generation with Markov Attention Entropy
The paper proposes using a VLA model’s own attention entropy as a reliability signal for its actions.
The authors frame VLA action generation as a conditional generative Markov chain and derive Markov Attention Entropy from internal attention patterns. They report that visual-modality entropy separates successful and failed tasks across different VLA architectures. They also introduce LIBERO-Reflect, a 4,000-episode benchmark split between standard and more challenging episodes. In experiments, MAE beats uncertainty baselines on AUPR, AUROC, and FPR@95, and FabriMAE improves PI-family test-time action selection with small observed runtime overhead. ArXiv · AI/CL/LG's note
The authors frame VLA action generation as a conditional generative Markov chain and derive Markov Attention Entropy from internal attention patterns. They report that visual-modality entropy separates successful and failed tasks across different VLA architectures. They also introduce LIBERO-Reflect, a 4,000-episode benchmark split between standard and more challenging episodes. In experiments, MAE beats uncertainty baselines on AUPR, AUROC, and FPR@95, and FabriMAE improves PI-family test-time action selection with small observed runtime overhead. ArXiv · AI/CL/LG's note
score 4