Toward Virtuous Reinforcement Learning

Summary

This seeded note is based on the first page and abstract. The paper critiques rule-based and scalar-reward ethical RL approaches and proposes a virtue-focused roadmap based on durable policy-level dispositions rather than isolated rule checks.

Key Claims

  • Constraint-heavy deontological approaches can become brittle under ambiguity and non-stationarity.
  • Scalarized reward approaches can hide moral trade-offs and invite proxy gaming.
  • A virtue-oriented framing should evaluate stable traits, social learning, and explicit trade-off reporting.

Methods / Formalism

  • Conceptual roadmap rather than a single algorithm.
  • Emphasis on social learning, multi-objective and constrained RL, updateable priors, and explicit value assumptions.

Evidence / Experiments

  • This appears to be a position and roadmap paper.
  • Full support structure and related-work map still need full ingest.

Connections

  • Useful counterpoint to strict Shielding-style safety methods.
  • Connects ethics, Strategic Reasoning, and long-run adaptation themes near Continual Learning.
  • Relevant to how the wiki should distinguish hard guarantees from normative modeling.

Open Questions

  • How should virtue-style dispositions be represented in formal multi-agent settings?
  • Which parts can be combined with formal constraints rather than replacing them?

Citation

Ghasemi, M., and Crowley, M. (2025). Toward Virtuous Reinforcement Learning: A Critique and Roadmap.