MEA: A Reward-Driven Multi-Agent System for Faithful Model Explanations
MEA trains explanation agents against faithfulness rewards, and the paper says that beats frontier LLMs and other baselines across tabular, text, and vision tasks.
The system splits the job between a Proposer, which chooses and configures explanation tools, and an Actor, which turns tool output into natural-language explanations grounded in model behavior. The authors add question types for feature attribution, counterfactual reasoning, and spurious feature detection, each tied to perturbation-based faithfulness metrics. They report gains over the untrained backbone of 28% for tabular, 21% for text, and 34% for vision. Source: HF Daily Papers' note.
The system splits the job between a Proposer, which chooses and configures explanation tools, and an Actor, which turns tool output into natural-language explanations grounded in model behavior. The authors add question types for feature attribution, counterfactual reasoning, and spurious feature detection, each tied to perturbation-based faithfulness metrics. They report gains over the untrained backbone of 28% for tabular, 21% for text, and 34% for vision. Source: HF Daily Papers' note.
score 4