Toward Virtuous Reinforcement Learning
Summary
This seeded note is based on the first page and abstract. The paper critiques rule-based and scalar-reward ethical RL approaches and proposes a virtue-focused roadmap based on durable policy-level dispositions rather than isolated rule checks.
Key Claims
- Constraint-heavy deontological approaches can become brittle under ambiguity and non-stationarity.
- Scalarized reward approaches can hide moral trade-offs and invite proxy gaming.
- A virtue-oriented framing should evaluate stable traits, social learning, and explicit trade-off reporting.
Methods / Formalism
- Conceptual roadmap rather than a single algorithm.
- Emphasis on social learning, multi-objective and constrained RL, updateable priors, and explicit value assumptions.
Evidence / Experiments
- This appears to be a position and roadmap paper.
- Full support structure and related-work map still need full ingest.
Connections
- Useful counterpoint to strict Shielding-style safety methods.
- Connects ethics, Strategic Reasoning, and long-run adaptation themes near Continual Learning.
- Relevant to how the wiki should distinguish hard guarantees from normative modeling.
Open Questions
- How should virtue-style dispositions be represented in formal multi-agent settings?
- Which parts can be combined with formal constraints rather than replacing them?
Citation
Ghasemi, M., and Crowley, M. (2025). Toward Virtuous Reinforcement Learning: A Critique and Roadmap.