Explainable AI
Overview
This area tracks methods for making model behavior and agent decisions legible, inspectable, and communicable.
Active Questions
- Which explanations are faithful enough to support safety claims?
- When is a reasoning trace monitorable even if it is not a complete record of internal computation?
- How should explanation interact with causal reasoning and strategic environments?
- What explanations are useful for debugging formal constraints versus learned objectives?
- When does successful agent communication fail to explain shared internal meaning?
- How should post-failure explanations change perceived agency, causality, and patient harm without hiding real accountability?
- How can explanations expose the safety layer’s formal reason for blocking an action without pretending to explain the whole learned policy?
- When does a counterfactual explanation require a causal simulation rather than a nearby contrastive example?
Key Concepts
- Chain-of-Thought Monitorability
- Explainable Shielding
- Shield Risk Decision Trees
- Emergent Communication
- Successful Misunderstandings
- Dyadic Morality
- Structural Equation Models
- Counterfactual Simulation
- Causality
- Strategic Reasoning
Key Sources
- Guan2026 - Open Sourcing Monitorability Evaluations
- Rieder2025 - Explainably Safe Reinforcement Learning
- Kondylidis2025 - Successful Misunderstandings: Learning to Coordinate Without Being Understood
- Varshney2026 - An Algebraic Exposition of the Theory of Dyadic Morality
- Ghasemi2025 - Toward Virtuous Reinforcement Learning
- Karvanen2024 - Simulating Counterfactuals