Megadose Built for builders and researchers.

MEA: A Reward-Driven Multi-Agent System for Faithful Model Explanations

· HF Daily Papers ·
MEA trains explanation agents against faithfulness rewards, and the paper says that beats frontier LLMs and other baselines across tabular, text, and vision tasks.

The system splits the job between a Proposer, which chooses and configures explanation tools, and an Actor, which turns tool output into natural-language explanations grounded in model behavior. The authors add question types for feature attribution, counterfactual reasoning, and spurious feature detection, each tied to perturbation-based faithfulness metrics. They report gains over the untrained backbone of 28% for tabular, 21% for text, and 34% for vision. Source: HF Daily Papers' note.

score 4

Categories: Research