Aspiration Dynamics in Repeated Games
Context
Bendor2001 - Aspiration-Based Reinforcement Learning in Repeated Interaction Games surveys repeated-game models where agents update actions by payoff satisfaction. The formal payload is useful outside the source note because the same aspiration-update schema can be compared with modern multi-agent learning dynamics, opponent shaping, and satisficing decision theory.
Formal Statement
Let a finite two-player stage game be repeated at dates t = 1, 2, .... Agent i has aspiration A_i^t, action or mixed-strategy state s_i^t, and realized payoff pi_i^t. A common endogenous aspiration update is
A schematic aspiration-based reinforcement rule is:
For the symmetric 2 x 2 cooperation template
aspirations just below sigma make mutual cooperation the only state in which both players are satisfied. With slow aspiration updating and rare trembles, the surveyed KMRV-style model concentrates long-run play on (C,C).
In the generational model surveyed by Bendor, Mookherjee, and Ray, pure stable outcomes are constrained by individual rationality and by either efficiency or protected Nash structure. Strictly individually rational efficient outcomes, and strict Nash equilibria, satisfy the corresponding stability criterion.
Derivation / Construction
The cooperation mechanism is satisfaction-driven rather than belief-driven. If both agents cooperate and receive sigma, aspirations near sigma are met, so both agents tend to keep cooperating. If one deviates, the deviator may gain, but the opponent receives a payoff below aspiration and becomes more likely to switch. That response can destroy the deviator’s favorable payoff, making unilateral deviation difficult to stabilize.
Endogenous aspiration dynamics add another feedback loop. Unfavorable payoffs pull aspirations downward only gradually when lambda is close to one, while occasional trembles prevent permanent traps at low aspirations. Once both agents return to cooperation, repeated satisfactory payoffs pull aspirations back toward the cooperative level.
The generational formulation separates within-generation strategy adjustment from across-generation aspiration inheritance. Stability is then tested by small perturbations to one player’s behavioral state and by whether the induced long-run payoff distribution returns to the same aspiration-supporting outcome.
Implications
- Cooperation can persist even when the cooperative action is strictly dominated in the one-shot stage game.
- Non-Nash outcomes can be long-run outcomes because agents respond to satisfaction and dissatisfaction, not only to payoff-maximizing deviations.
- The same repeated-game structure can stabilize efficient cooperation or inefficient protected equilibria depending on aspiration formation and perturbation assumptions.
- For modern MARL, aspiration dynamics are a baseline for cooperation that does not require explicit opponent gradients, centralized critics, or rich opponent models.