ML models in high-stakes domains remain largely opaque to the practitioners who act on their predictions. Post-hoc explanation methods offer a lens into model behavior, but wielding them demands expertise most domain experts lack: navigating high-dimensional outputs, selecting the best explanations, and synthesizing evidence across disparate tools.

MEA removes that knowledge barrier with a division of labor: a Proposer agent selects and configures explanation tools based on the question and modality, while an Actor agent, optimized end-to-end against faithfulness, transforms tool outputs into natural-language explanations across tabular, text, and vision modalities. The framework spans feature attribution, counterfactual reasoning, and spurious feature detection, each paired with a perturbation-based faithfulness metric.

A baseline finding makes the design bite: frontier LLMs asked to explain models systematically produce unfaithful explanations. Reward-driven optimization against faithfulness, augmented with a modality-adaptive penalty, lets MEA outperform post-hoc explainers, agentic, and closed-source baselines across six datasets, with faithfulness gains of +28% (tabular), +21% (text), and +34% (vision) over the untrained backbone.

Faithfulness can be measured by perturbation, so it can serve directly as a training reward instead of a post-hoc score.

Splitting tool selection from explanation writing lets each agent be optimized for its own failure mode.

A modality-adaptive penalty keeps the optimizer honest across tabular, text, and vision inputs.

Frontier LLMs prompted directly are systematically unfaithful, so prompting alone cannot close the expertise gap.

When a tool chain needs expertise its users don't have, make an agent learn to operate it and grade the agent on the quality metric that matters — the reward replaces the manual skill.

The authors position AI agents as a scalable, adaptable interface to ML explainability — a path toward natural-language explainability that generalizes beyond the fixed, single-purpose tools that have long defined the field. Posted as a v1 preprint on arXiv on 1 October 2026; no deployment results yet.

FOLLOW THE EVIDENCE

The sources

  1. MEA: A Reward-Driven Multi-Agent System for Faithful Model Explanations arxiv.org