Equilibrium in a Stochastic n-Person Game

Summary

Fink extends discounted stochastic-game equilibrium existence beyond the two-player case. The paper models an infinite-horizon sequence of state-indexed games in which each player chooses an action at the current state, the joint action stochastically determines the next state, and each player minimizes a geometrically discounted stream of costs. The main result proves that stationary mixed-strategy equilibria exist for finite discounted stochastic n-person games.

The note is short but useful as a foundational bridge between classical stochastic games, modern Stochastic Games / Markov games in MARL, and equilibrium-focused safe-MARL work such as Constrained Markov Potential Games.

Key Claims

  • Finite discounted stochastic n-person games have stationary mixed-strategy equilibrium points.
  • For a fixed stationary mixed-strategy profile, each player’s expected discounted cost vector is uniquely determined.
  • The equilibrium proof can be organized as a contraction argument for the value vector followed by a Kakutani fixed-point argument for the best-response correspondence.
  • The argument also covers closed convex restrictions of the stationary strategy space, giving an existence result for constrained variants in the paper’s sense.
  • The paper notes extensions to denumerable state sets with bounded costs and to arbitrary action-cardinality cases where exact minima are replaced by infima, yielding epsilon-effective strategies.

Methods / Formalism

  • The model has a finite state set I. At state i, player h chooses an alternative j_h in J_h(i) with knowledge of the state.
  • A joint action j=(j_1,...,j_n) induces transition probabilities P_{ij_1...j_n k} over next states k, and player h pays cost C_{hij} at that stage.
  • Player h discounts future costs by alpha_h, with 0 < alpha_h < 1.
  • A stationary mixed strategy gives each player and state a probability vector x^h(i) over J_h(i).
  • For a stationary profile x, the expected discounted cost vector satisfies a Bellman-style linear system:
  • The best-response value is written as a componentwise minimization over one player’s state-wise mixed action while the other players’ profile is held fixed.
  • Defining T_x v = min_y f(x,y,v), Fink shows T_x is a contraction with modulus bounded by a=max_h alpha_h, so the value vector for a fixed profile is unique.
  • Discounted Stochastic Game Equilibrium records the fixed-point proof skeleton and the equilibrium statement.

Evidence / Experiments

  • This is a mathematical existence note, not an empirical paper.
  • The evidence is a proof: uniqueness of expected cost values for fixed strategies, contraction of the value operator, upper-semicontinuity/closedness of the best-response correspondence, and Kakutani’s fixed-point theorem.
  • The paper positions the n=1 case as dynamic programming and the n=2 case as recovering Shapley’s stochastic-games result.

Connections

  • Seeds Stochastic Games as the classical discounted game model underlying later Markov-game formulations in MARL.
  • Supports Multi-Agent Non-Stationarity by separating the stationary stochastic game model from the learning-induced non-stationarity created when agents update policies over time.
  • Provides historical background for Constrained Markov Potential Games, which add potential-game structure and safety/resource feasibility constraints to Markov games.
  • Belongs in Strategic Reasoning because the solution concept is an equilibrium of mutually best-responding stationary strategies.
  • Belongs in Decision Theory because it formalizes sequential multi-agent choice under uncertainty, discounting, and strategic coupling.

Open Questions

  • Which later stochastic-game equilibrium refinements should be added to the wiki before using this as a background page for modern MARL theory?
  • Should the wiki distinguish finite discounted stochastic games, general-sum Markov games, and constrained Markov games as separate concept pages if more sources arrive?
  • How much of the constrained-subset observation in Fink’s proof can be reused for modern safety-constrained games, and where do newer feasibility notions break the analogy?

Citation

Fink, A. M. (1964). Equilibrium in a Stochastic n-Person Game. Journal of Science of the Hiroshima University, Series A-I, 28, 89-93.