An ontology for organizing and recommending measures for evaluating the faithfulness of AI explanations
Resources
| Title | Authors | Publication Year |
|---|---|---|
| Post hoc explanations may be ineffective for detecting unknown spurious correlation | Julius Adebayo, Michael Muelly, Harold Abelson, Been Kim | 2022 |
| "What is relevant in a text document?": An interpretable machine learning approach | Leila Arras, Franziska Horn, Grégoire Montavon, Klaus-Robert Müller, Wojciech Samek | 2017 |
| Evaluating recurrent neural network explanations | Leila Arras, Ahmed Osman, Klaus-Robert Müller, Wojciech Samek | 2019 |
| A diagnostic study of explainability techniques for text classification | Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle Augenstein | 2020 |
| Chain-of-thought unfaithfulness as disguised accuracy | Oliver Bentham, Nathan Stringham, Ana Marasović | 2024 |
| REV: information-theoretic evaluation of free-text rationales | Hanjie Chen, Faeze Brahman, Xiang Ren, Yangfeng Ji, Yejin Choi, Swabha Swayamdipta | 2023 |
| Reasoning Models Don't Always Say What They Think | Yanda Chen, Joe Benton, Ansh Radhakrishnan, Jonathan Uesato, Carson Denison, John Schulman, Arushi Somani, Peter Hase, Misha Wagner, Fabien Roger, Vlad Mikulik, Samuel R. Bowman, Jan Leike, Jared Kaplan, Ethan Perez | 2025 |
| Leakage-adjusted simulatability: Can models generate non-trivial explanations of their behavior in natural language? | Peter Hase, Shiyue Zhang, Harry Xie, Mohit Bansal | 2020 |
| A benchmark for interpretability methods in deep neural networks | Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, Been Kim | 2019 |
| Alignment rationale for natural language inference | Zhongtao Jiang, Yuanzhe Zhang, Zhao Yang, Jun Zhao, Kang Liu | 2021 |