The solution
ML models in high-stakes domains remain largely opaque to the practitioners who act on their predictions. Post-hoc explanation methods offer a lens into model behavior, but wielding them demands expertise most domain experts lack: navigating high-dimensional outputs, selecting the best explanations, and synthesizing evidence across disparate tools.
MEA removes that knowledge barrier with a division of labor: a Proposer agent selects and configures explanation tools based on the question and modality, while an Actor agent, optimized end-to-end against faithfulness, transforms tool outputs into natural-language explanations across tabular, text, and vision modalities. The framework spans feature attribution, counterfactual reasoning, and spurious feature detection, each paired with a perturbation-based faithfulness metric.
A baseline finding makes the design bite: frontier LLMs asked to explain models systematically produce unfaithful explanations. Reward-driven optimization against faithfulness, augmented with a modality-adaptive penalty, lets MEA outperform post-hoc explainers, agentic, and closed-source baselines across six datasets, with faithfulness gains of +28% (tabular), +21% (text), and +34% (vision) over the untrained backbone.
Why it worked
Faithfulness can be measured by perturbation, so it can serve directly as a training reward instead of a post-hoc score.
Splitting tool selection from explanation writing lets each agent be optimized for its own failure mode.
A modality-adaptive penalty keeps the optimizer honest across tabular, text, and vision inputs.
Frontier LLMs prompted directly are systematically unfaithful, so prompting alone cannot close the expertise gap.
What can be applied
When a tool chain needs expertise its users don't have, make an agent learn to operate it and grade the agent on the quality metric that matters — the reward replaces the manual skill.
Aftermath
The authors position AI agents as a scalable, adaptable interface to ML explainability — a path toward natural-language explainability that generalizes beyond the fixed, single-purpose tools that have long defined the field. Posted as a v1 preprint on arXiv on 1 October 2026; no deployment results yet.
FOLLOW THE EVIDENCE