Source-linked AI summary

Berk-Nash Equilibrium: A Framework for Modeling Agents with Misspecified Models

Ignacio Esponda, Demian Pouzo

arXiv:1411.1152v4econ.GNcs.GT

TL;DR

The paper addresses the gap between widespread model misspecification and the standard assumption that agents correctly understand their environments. It introduces Berk-Nash equilibrium, in which optimal behavior is paired with best-fit beliefs under weighted Kullback-Leibler divergence, and grounds the concept in misspecified Bayesian learning. The main result is that convergent behavior approaches Berk-Nash equilibrium, while any given equilibrium can be reached with vanishing optimization mistakes; the framework also nests Nash and self-confirming equilibrium in specified cases.

  • Problem

    The paper develops a unified framework for analyzing agents with subjective, possibly incorrect views of their environments, beyond settings that assume correct specification.

  • Method

    It defines Berk-Nash equilibrium using optimal strategies and beliefs supported on weighted Kullback-Leibler best fits, then studies repeated Bayesian learning with misspecified subjective models.

  • Results

    If players’ behavior converges, it converges to a Berk-Nash equilibrium, and any given Berk-Nash equilibrium is attainable with myopic agents whose optimization mistakes vanish over time.

  • Takeaways & Limitations

    Berk-Nash equilibrium places Nash and self-confirming equilibrium within a common framework while extending equilibrium analysis to misspecified learning.

  • Takeaways & Limitations

    The framework ignores repeated-game considerations in which players account for how their actions affect others’ future play.

Abstract

from arXiv · show

We develop an equilibrium framework that relaxes the standard assumption that people have a correctly-specified view of their environment. Each player is characterized by a (possibly misspecified) subjective model, which describes the set of feasible beliefs over payoff-relevant consequences as a function of actions. We introduce the notion of a Berk-Nash equilibrium: Each player follows a strategy that is optimal given her belief, and her belief is restricted to be the best fit among the set of beliefs she considers possible. The notion of best fit is formalized in terms of minimizing the Kullback-Leibler divergence, which is endogenous and depends on the equilibrium strategy profile. Standard solution concepts such as Nash equilibrium and self-confirming equilibrium constitute special cases where players have correctly-specified models. We provide a learning foundation for Berk-Nash equilibrium by extending and combining results from the statistics literature on misspecified learning and the economics literature on learning in games.

1 Introduction

The paper introduces Berk-Nash equilibrium for agents whose subjective models may be misspecified, defining beliefs as best fits to the true distribution under equilibrium behavior. It also provides a learning foundation and embeds several standard solution concepts as special cases.

  • Problem and framework: The framework allows agents to hold subjective models that exclude the true distribution of consequences conditional on actions and information.A subjective model specifies feasible probability distributions over a player’s own consequences as a function of action and information.
  • Problem and framework: A Berk-Nash equilibrium combines optimal strategies with beliefs concentrated on subjective distributions minimizing weighted Kullback-Leibler divergence from the equilibrium-induced true distribution.The best-fit criterion depends endogenously on the objective game and actual strategy profile.
  • Relation to existing concepts: Under correct specification and strong identification, Berk-Nash equilibrium is equivalent to Nash equilibrium; without strong identification, it becomes self-confirming equilibrium.The framework also provides a systematic way to extend these cases to other forms of misspecification.
  • Learning foundation: The learning foundation studies repeated play in which agents update Bayesian beliefs over stationary subjective models while behaving optimally with possibly misspecified models.The objective is to characterize limiting behavior when players actively learn through their actions.
  • Learning foundation: If players’ behavior converges, the limit is a Berk-Nash equilibrium; conversely, any given Berk-Nash equilibrium can be reached when myopic agents’ optimization mistakes vanish over time.The converse does not hold under exact optimization, but is recovered under asymptotically optimal choice.
  • Contribution: The paper unifies rational and boundedly rational approaches within a framework for studying misspecified views across several economic settings.It extends learning-in-games results to agents whose models remain misspecified even in steady state.

2 The framework

The framework separates the objective game from players’ subjective models, which may be misspecified, and defines equilibrium through optimal strategies and beliefs that best fit endogenously generated consequences. It establishes existence and illustrates how the framework produces Berk-Nash outcomes in applications.

  • The environment: The objective game specifies states, signals, actions, consequences, feedback functions, and payoffs, while subjective models describe players’ perceived consequence distributions.The separation between objective and subjective models permits misspecifications beyond uncertainty about structural primitives.
  • Definition of equilibrium: A Berk-Nash equilibrium requires each strategy to be optimal under some subjective belief concentrated on distributions minimizing weighted Kullback-Leibler divergence.The true distribution used for fitting is determined by the objective game and the actual strategy profile.
  • Existence: Every game has at least one Berk-Nash equilibrium, established through perturbed games and a fixed-point argument for a belief correspondence.The standard Nash existence proof does not apply because the analogous best-response correspondence need not be convex-valued.
  • Examples: In the monopolist example, the unique Berk-Nash equilibrium mixes prices as σ∗ = (35/36, 1/36), supported by belief θσ∗ = (40, 10/3).The supporting belief lies at the intersection of the indifference line and the relevant boundary segment, but its belief about the mean is not correct.
  • Examples: Misspecified models can generate distinct behavioral implications: model B yields optimal effort, whereas model A yields effort above the optimum.In model B, the agent has correct beliefs about the expected marginal tax rate at the equilibrium effort despite incorrectly treating that rate as constant.

3 Relationship to other solution concepts

Berk-Nash equilibrium encompasses Nash, self-confirming, and other boundedly rational concepts, with equivalence determined by correct specification, identification, and feedback. The framework also connects analogy-based games to ABEE and fully cursed equilibrium.

  • Berk-Nash equilibrium includes standard and boundedly rational solution concepts as special cases.
  • 3.2 Relationship to Nash and self-confirming equilibrium: Correct specification and strong identification make Berk-Nash equilibrium equivalent to Nash equilibrium.
  • 3.2 Relationship to Nash and self-confirming equilibrium: Correct specification alone implies that every Berk-Nash equilibrium is a self-confirming equilibrium, even without strong identification.
  • 3.2 Relationship to Nash and self-confirming equilibrium: When the game is misspecified, Berk-Nash beliefs can be incorrect on the equilibrium path, so the equilibrium need not be Nash or self-confirming.
  • 3.3 Relationship to fully cursed and ABEE: In analogy-based games, Berk-Nash equilibrium is equivalent to analogy-based expectation equilibrium.
  • 3.3 Relationship to fully cursed and ABEE: With a single analogy class, Berk-Nash coincides with fully cursed equilibrium, while under partial feedback it coincides with naive behavioral equilibrium.

4 Equilibrium foundation

The section develops the learning foundation for Berk-Nash equilibrium using perturbed repeated games with Bayesian updating. It shows that stabilized optimizing behavior must be Berk-Nash, while convergence to arbitrary equilibria requires asymptotically vanishing optimization mistakes under additional conditions.

  • Perturbed games: Payoff perturbations make behavior continuous in beliefs, ensuring that convergent beliefs imply convergent behavior.Players privately observe own-action payoff perturbations before acting; optimal mixed behavior equals the probability that each action is optimal across perturbation realizations.
  • Learning setup: Players repeatedly play a stationary objective game, update subjective-model beliefs by Bayes’ rule, and myopically maximize current expected payoffs.Beliefs use past signals, actions, and consequences, while intended strategies depend on current beliefs and therefore on history.
  • Learning setup: If intended behavior stabilizes, posterior beliefs increasingly concentrate on the subjective distributions closest to the limiting strategy profile.This extends misspecified-learning results to non-i.i.d. data generated endogenously by players’ own actions, although posteriors themselves need not converge.
  • Equilibrium characterization: Any strategy profile stable under an optimal policy in a perturbed game is a Berk-Nash equilibrium.The characterization follows because limiting behavior is statically optimal and beliefs concentrate on the best-fitting subjective distributions.
  • Convergence: The theorem rules out non-equilibrium limits of optimizing behavior, but it does not guarantee convergence because Nash equilibrium, a special case, need not be reached by learning.The paper therefore relaxes exact optimality to allow vanishing optimization mistakes when proving converse results.
  • Convergence: For any weakly identified Berk-Nash equilibrium and any a > 0, an asymptotically optimal policy can converge to it with probability greater than 1 − a.The result applies for priors whose conditional beliefs on best-fitting distributions equal a supporting belief profile.

5 Discussion

The discussion clarifies the framework’s scope and its relation to bounded-rationality models. It identifies settings requiring additional assumptions or extensions, including dynamic interaction, large populations, and misspecified extensive forms.

  • Scope and assumptions: The framework can fail to deliver existence or posterior concentration when the model’s regularity assumptions are violated.In the example, a misspecified two-parameter model makes one action incompatible with an observed outcome, preventing a Berk-Nash equilibrium and allowing posteriors to remain concentrated on the wrong parameter after experimentation.
  • Forward-looking agents: For forward-looking players, the extension requires weak identification to eliminate steady-state incentives to experiment.The main dynamic model assumes myopic players; non-myopic behavior may differ in the steady state when experimentation remains valuable.
  • Large population models: The fixed-player stationary framework rules out repeated-game motives to influence others’ future play.An appendix adapts the concept to large populations, where each agent has negligible influence over others’ actions.
  • Extensive-form games: The equilibrium concept applies to extensive-form settings when players choose contingent action plans and know the extensive form.The extension is less clear when players misunderstand the extensive form or move sequentially, and those cases remain future work.
  • Bounded rationality: Misspecified endogenous learning offers a way to derive beliefs from a plausible model restriction rather than fixing observed beliefs at the outset.The discussion contrasts this approach with singleton subjective models that leave no room for learning about parameter values.
  • Related models: Several bounded-rationality models fall outside the paper’s scope because they involve mediated interactions, dynamic problems, or alternative restrictions such as limited forecasting and similarity classes.The paper focuses on repetition of a static problem with stationary subjective models.

Appendix

The appendix supplies formal definitions and proof ingredients for equilibrium existence and learning convergence. It uses Bayesian operators, best-fit sets, continuity, compactness, and convergence arguments to establish the main results.

  • Existence: The proof of equilibrium existence constructs a correspondence from belief profiles to distributions over best-fitting subjective parameters and applies a fixed-point theorem.Compactness and upper hemicontinuity of the best-fit correspondence provide the key regularity conditions.

A Example: Trading with adverse selection

The trading example combines adverse selection with alternative feedback structures and misspecified beliefs. It derives the best-fitting parameters and perceived-profit expressions for independent, cursed, analogy-based, and related behavioral models.

  • Trading environment: The trading environment has a true distribution over buyer signals and valuations, with payoffs determined by whether the buyer’s price covers the signal-dependent trading condition.Partial feedback reveals valuation only when trade occurs, whereas full feedback reveals both signal and valuation.
  • Comparison: The example compares four cases formed by combining partial or full feedback with independent or analogy-based parameter sets.For each case, the appendix specifies the subjective model and weighted Kullback-Leibler divergence.
  • Independent beliefs: Independent-belief misspecification separates the signal and valuation distributions, so the best-fitting parameters use the marginal distributions pA and pV.This parameterization underlies the behavioral and cursed-equilibrium cases with independent beliefs.
  • Behavioral equilibrium: Under partial feedback, the behavioral-equilibrium best fit conditions valuation beliefs on observed signals and the chosen price.For each price x, the valuation component is pV |A(v | A ≤ x), while the signal component remains pA.
  • Analogy-based beliefs: Analogy-based beliefs assign separate signal distributions to analogy classes while retaining a common valuation distribution.The best-fitting signal parameter for class j is pA|Vj, and the valuation parameter is pV.
  • Behavioral equilibrium with analogy classes: With partial feedback and analogy classes, stabilized behavior can generate multiple beliefs about deviations because valuation beliefs may vary across classes and prices.The buyer may therefore infer that changing price affects perceived valuation through correlations across analogy classes.

B Proof of converse result: Theorem 3

The proof constructs perturbed policies that become asymptotically optimal while beliefs converge to distributions supported on the model’s best-fitting set. It then shows the limiting strategy profile satisfies the conditions for a Berk-Nash equilibrium.

  • Asymptotic optimality: Vanishing perturbations make the constructed policies asymptotically optimal relative to the evolving beliefs.The proof uses ε_t ≥ 0 with ε_t → 0 and bounds utility differences by 0.5ε_t.
  • Policy construction: The stationary policy profile maximizes utility under a fixed belief profile and provides the benchmark for constructing perturbed policies.The construction defines each policy as optimal when beliefs remain fixed at the supporting belief.
  • Perturbation sequence: The perturbation sequence is chosen from stopping times so that ε_t eventually equals 1/N(t) and converges to zero.Because the stopping times diverge, N(t) grows without bound and ε_t vanishes.
  • Convergence: Under the constructed policy, the strategy process converges almost surely to the target strategy profile.The proof establishes probability one for convergence of σ_t(h∞) to σ.
  • Belief convergence: Belief supports converge toward the distributions consistent with the limiting strategy, allowing the target profile to satisfy the equilibrium characterization.The argument uses weak identification, closed-set bounds, the portmanteau lemma, and continuity of the induced distributions.

C Non-myopic players

The non-myopic extension lets players maximize discounted expected payoffs and experiment while maintaining a stationary-environment belief. Under weak identification, stable optimal behavior still yields a Berk-Nash equilibrium.

  • Dynamic optimization: Players solve a discounted dynamic optimization problem and may experiment, although they believe the environment is stationary.The recursive problem uses a value function and Bayesian belief updating.
  • Dynamic optimization: The Bellman equation has a unique bounded continuous value function, and the optimal policy correspondence is single-valued and continuous almost surely.These regularity properties support continuity of dynamic behavior in beliefs.
  • Equilibrium characterization: Stable strategy profiles are dynamically optimal for a convergent posterior, while weak identification makes the limiting posterior a fixed point that provides no new information.Nonnegative experimentation value then implies myopic optimality at the stable profile.
  • Equilibrium characterization: A strategy profile stable under an optimal policy in a perturbed, weakly identified game is a Berk-Nash equilibrium.The theorem extends the learning foundation from myopic to forward-looking players.

D Population models

Population models generate heterogeneous Berk-Nash equilibria under different matching and feedback structures. The relevant beliefs may depend on individual experiences or aggregate population strategies.

  • Model variants: The appropriate population-model variant depends on the application because matching technology and feedback differ across settings.The paper considers random matching, single-pair, and population-feedback environments.
  • Single-pair model: In the single-pair model, steady-state behavior corresponds exactly to the paper’s Berk-Nash equilibrium.Each period selects one pair from each population, with population histories publicly revealed.
  • Random matching: Under random matching with private match feedback, agents in the same population may hold different beliefs and play different steady-state strategies.Different experiences motivate a heterogeneous extension analogous to heterogeneous self-confirming equilibrium.
  • Scope: Some population variants may be unrealistic when agents cannot observe previous generations’ private signals.The paper notes that such settings may instead be better modeled with public information.
  • Population feedback: With population feedback, beliefs depend on aggregate population strategies rather than players’ own strategies.This is the main behavioral distinction introduced by population feedback.

D.1 Equilibrium foundation

The population learning foundation links individual limiting behavior to heterogeneous Berk-Nash equilibrium. When individual strategies stabilize, each agent’s strategy is optimal for a best-fitting belief, and the aggregate strategy lies in the corresponding convex hull.

  • Equilibrium foundation: A population strategy is the aggregate of individual strategies, and convergence of these aggregates makes the limiting belief support depend on population behavior.The population is modeled as a continuum of agents indexed by k ∈ [0,1].
  • Random matching: Individual stabilization is required; aggregate stabilization alone is insufficient for the random-matching foundation.The assumption is motivated by unstable agents eventually recognizing model misspecification.
  • Equilibrium foundation: The heterogeneous formulation uses convex combinations because the best-response set may contain mixed strategies without corresponding pure strategies.The paper cites this as a reason not to support each action with a separate belief.
  • Equilibrium foundation: Each agent’s convergent belief is supported on the distributions consistent with that agent’s limiting strategy and the opponents’ aggregate strategy.Continuity of behavior in beliefs makes each individual strategy optimal for its limiting belief.
  • Equilibrium foundation: The aggregate strategy belongs to the convex hull of the best-response set when individual strategies stabilize.This establishes the heterogeneous Berk-Nash equilibrium condition for the population.

E Lack of payoff feedback

The paper considers settings where players receive no or only partial payoff feedback and adjusts optimality and identification accordingly. With partial feedback, observed consequences may leave the underlying distribution only partially identified.

  • Players may observe no payoff feedback or only partial feedback about payoffs.
  • No payoff feedback: With random payoff functions and no payoff feedback, payoffs provide nothing new to learn from observed consequences.The payoff function is modeled as an independent random variable.
  • No payoff feedback: Under no payoff feedback, optimality is equivalent to using the expected payoff function.
  • Partial payoff feedback: With partial payoff feedback, observing consequences can provide information about other players’ actions and the state, and therefore about payoffs.
  • Partial payoff feedback: Partial payoff feedback can leave several distributions over other players’ actions and states consistent with the same consequence distribution.
  • Partial payoff feedback: Identification must require both a unique consequence distribution matching observed data and a unique expected utility function implied by it.

F Global stability: Example 2.1 (monopoly with unknown demand).

The monopoly example models learning with perturbed payoffs and an unknown demand parameter, then analyzes the resulting stochastic dynamics. It establishes a unique steady state and global convergence to the unique equilibrium under the stated assumptions.

  • Global convergence: The unique equilibrium is globally asymptotically stable, so behavior converges to it with probability 1 in the example.
  • Perturbed game: The example perturbs monopoly payoffs before analyzing learning dynamics and then takes the payoff perturbations to zero.
  • Subjective model: The monopolist learns only the demand parameter b while beliefs about a remain fixed at a value different from its true value.
  • Decision rule: Only the expected value of B affects the monopolist’s optimal strategy.
  • Bayesian updating: The Gaussian prior remains Gaussian after Bayesian updating, with its mean and variance evolving recursively.
  • Stochastic approximation: The learning process is represented as a nonlinear stochastic difference equation and studied using stochastic approximation and an associated ODE.
  • Steady states: A unique steady state exists because the relevant function is continuous, crosses zero, and is decreasing on the relevant domain.
Loading 1411.1152v4…