解法
ML models in high-stakes domains remain largely opaque to the practitioners who act on their predictions. Post-hoc explanation methods offer a lens into model behavior, but wielding them demands expertise most domain experts lack: navigating high-dimensional outputs, selecting the best explanations, and synthesizing evidence across disparate tools.
MEA removes that knowledge barrier with a division of labor: a Proposer agent selects and configures explanation tools based on the question and modality, while an Actor agent, optimized end-to-end against faithfulness, transforms tool outputs into natural-language explanations across tabular, text, and vision modalities. The framework spans feature attribution, counterfactual reasoning, and spurious feature detection, each paired with a perturbation-based faithfulness metric.
A baseline finding makes the design bite: frontier LLMs asked to explain models systematically produce unfaithful explanations. Reward-driven optimization against faithfulness, augmented with a modality-adaptive penalty, lets MEA outperform post-hoc explainers, agentic, and closed-source baselines across six datasets, with faithfulness gains of +28% (tabular), +21% (text), and +34% (vision) over the untrained backbone.
What it achieved
Stop building new explainers — train agents to wield the existing ones, graded by faithfulness itself
为什么有效
Faithfulness can be measured by perturbation, so it can serve directly as a training reward instead of a post-hoc score.
Splitting tool selection from explanation writing lets each agent be optimized for its own failure mode.
A modality-adaptive penalty keeps the optimizer honest across tabular, text, and vision inputs.
Frontier LLMs prompted directly are systematically unfaithful, so prompting alone cannot close the expertise gap.
可以借鉴什么
When a tool chain needs expertise its users don't have, make an agent learn to operate it and grade the agent on the quality metric that matters — the reward replaces the manual skill.
后续
The authors position AI agents as a scalable, adaptable interface to ML explainability — a path toward natural-language explainability that generalizes beyond the fixed, single-purpose tools that have long defined the field. Posted as a v1 preprint on arXiv on 1 October 2026; no deployment results yet.
从这里,追溯依据