Explainable AI

Overview

This area tracks methods for making model behavior and agent decisions legible, inspectable, and communicable.

Active Questions

  • Which explanations are faithful enough to support safety claims?
  • When is a reasoning trace monitorable even if it is not a complete record of internal computation?
  • How should explanation interact with causal reasoning and strategic environments?
  • What explanations are useful for debugging formal constraints versus learned objectives?
  • When does successful agent communication fail to explain shared internal meaning?
  • How should post-failure explanations change perceived agency, causality, and patient harm without hiding real accountability?
  • How can explanations expose the safety layer’s formal reason for blocking an action without pretending to explain the whole learned policy?
  • When does a counterfactual explanation require a causal simulation rather than a nearby contrastive example?

Key Concepts

Key Sources

Adjacent Foundations