Next Article in Journal
Profiling Winegrowers’ Attitudes Towards Organic and Sustainable Viticulture in Western Macedonia
Previous Article in Journal
Economic Aspects of the Circular Food Economy: The Case of Olive Oil
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Proceeding Paper

Allais–Ellsberg Convergent Markov–Network Game †

by
Adil Ahmad Mughal
1,2
1
Department of Economics, Forman Christian College University, Lahore 54600, Pakistan
2
Department of Commerce & Finance, Government College University, Lahore 54000, Pakistan
Presented at the 1st International Electronic Conference on Games (IECGA 2025), 15–16 October 2025; Available online: https://sciforum.net/event/IECGA2025.
Proceedings 2026, 135(1), 2; https://doi.org/10.3390/proceedings2026135002
Published: 19 January 2026
(This article belongs to the Proceedings of The 1st International Electronic Conference on Games (IECGA 2025))

Abstract

Behavioral deviations from subjective expected utility theory, most famously captured by the Allais paradox and the Ellsberg paradox, have inspired extensive theoretical and experimental research into risk and ambiguity preferences. While the existing analyze these paradoxes independently, little work explores how such heterogeneously biased agents interact in networked strategic environments. Our paper fills this gap by modeling a convergent Markov–network game between Allais-type and Ellsberg-type players, each endowed with fully enriched loss matrices that reflect their distinct probabilistic and ambiguity attitudes. We define convergent priors as those inducing a spectral radius of <1 in iterated enriched matrices, ensuring iterative convergence under a matrix-based update rule. Players minimize their losses under these priors in each iteration, converging to an equilibrium where no further updates are feasible. We analyze this convergence under three learning regimes—homophily, heterophily, and type-neutral randomness—each defined via distinct neighborhood learning dynamics. To validate the equilibrium, we construct a risk-neutral measure by transforming losses into payoffs and derive a riskless rate of return representing players’ subjective indifference to risk. This applies risk-neutral pricing logic to behavioral matrices, which is novel. This framework unifies paradox-type decision makers within a networked Markovian environment (stochastic adjacency matrix), extending models of dynamic learning and providing a novel equilibrium characterization for heterogeneous, ambiguity-averse agents in structured interactions.

1. Game Setup

Behavioral violations of subjective expected utility theory, most notably the Allais paradox [1] and the Ellsberg paradox [2], have generated a large literature on risk and ambiguity preferences. Foundational models formalizing these deviations include rank-dependent and non-expected utility approaches [3] and ambiguity-averse preferences based on multiple priors [4]. Separately, a rich body of work studies learning and belief updating in networked environments [5,6,7]. However, little work integrates heterogeneous risk- and ambiguity-biased agents into strategic network settings. This paper contributes to this gap.
There are five players, three Ellsberg-type E1, E2, and E3 and two Allais-type A1 and A2. Each player i has an enriched loss matrix Li ∈ ℝ2×2 whose rows correspond to states, θ ∈ {θ1(good) and  θ2 (bad)}, and whose columns correspond to actions a ∈ {a1 (risky),  a2 (safe)}. The baseline risk-neutral loss matrix is L = [l11 l12 l21 l22] = [−1 −1 1 1]. This corresponds to the loss of −1 in the good state under either action and the loss of +1 in the bad state under either action. The game introduces behavioral distortions through multiplicative weights. Machina’s [3] smooth ambiguity model modifies utility by probability-separable curvature. We introduce a matrix-superadditive analog that has multiplicative state-action weights—to be detailed in the full paper. For behavioral enrichment of losses, let W = (wij) and collect the multiplicative weights applied to each entry of L. Enriched losses are defined elementwise by L∗ = W ∘ L. Ellsberg-type players overweight ambiguity in the bad state. The weight matrix is WE = [1 1 3 2], LE∗ = [−1 −1 3 2]. Allais-type players overweight certainty in the good state. Their weight matrix is WA = [2 2 1 1] and LA∗ = [−2 −2  1  1]. Thus, E-type players become ambiguity averse in the bad state and A-type agents become certainty lovers in the good state.
The loss enrichment has implications for prior choice also. Let π∗(θ1)max denote the highest convergent prior for a player’s type—i.e., the largest prior probability assigned to state θ1 for which the player’s loss matrix L yields spectral radius ρ(L) < 1. A loss matrix is thus fully enriched if the chosen prior is only the highest convergent prior (only decile priors are used here for simplicity). This makes sure, as a second-order condition on multiplicative state-action weights, that all of the subjective or empirical probability weights have been loaded upon, and transformed into, loss magnitudes. In theory, it must leave only purely unbiased priors to be chosen, such that any hesitation in choosing the highest convergent prior must force a revelation of some hidden probability or preference. So, the loss matrix for A has π∗(θ1)max ≈ 0.5, where π∗(θ1)A ∈ [0.1, 0.5] and for E π∗(θ1)max ≈ 0.9, where π∗(θ1)E ∈ [0.5, 0.9]. One can see how overweighting of good states (bad states) by Allais-type (Ellsberg-type) players is exactly penalized in the convergent prior set {π∗(θ1)A} ({π∗(θ1)E}). The fully enriched convergent expected loss vector is defined as
EL∗: = minπ∗(θ1)max {π∗(θ1)max[l11, l12] + (1 − π∗(θ1)max)[l21, l22]}.
Equivalently, EL∗ = EL(π∗(θ1)max), where ρ(L) < 1. This expresses the best-case (most favorable) convergent expected loss available to a player under its own feasible prior set. A player i ∈ {A, E} is declared the winner of the game M if and only if it most minimizes its fully enriched convergent expected loss vector: Winner = arg min i ∈ {A,E} ELi∗. Because losses are negative in our normalization, this means that the winner is the player whose enriched convergent expected loss is most negative.
The network structure of M is a row-stochastic Markovian adjacency matrix X ∈ ℝ5×5. There are three scenarios built around neighborhood strength, or learning influence, r between any two players or two network nodes. For simplicity, there is no self-learning or self-neighborhood so {rii = rjj = 0} ⇒ diag(X) = 0.
Scenario 1, Homophily:X is homophilous if rA,A > rA,E and rE,E > rA,E.
Scenario 2, Heterophily: X is heterophilous if rA,A < rA,E and rE,E < rA,E.
Scenario 3, Type-Neutral Random (TNR): X is TNR if each r is random.
Each iteration of X updates loss matrices by the linear Markov rule,
Lit+1 = ∑j Xij Ljt 
This step does not use priors at all (it only needs the L matrices and X). Once a new iterated matrix is formed, the player selects a new highest convergent prior for reporting by scanning deciles and minimizing their EL* score—this is a post-processing selection based on the current Lit. That is, the transmission rule (how chosen π affects the next iteration) is as follows: once player i selects πi(t)max, they transmit to neighbors the row-scaled matrix Lit~  =  diag(πi(t)max, 1 −  πi(t)max) Lit or equivalently, Lit~ = D(πi(t)max) ∘ Lit in Hadamard form. The network update then uses L~ as the input:
Lit+1 = ∑j Xij Ljt~ 
Repeat: at iteration t + 1, each player computes its candidate π from the newly computed Lit+1, picks the highest convergent decile, transmits the scaled matrix to neighbors, and so on. The process stops when no player changes π (and/or when L converges).

2. Convergent Risk-Neutral Measure and Nash Equilibrium

Unique and Positive Risk-Neutral Measure: The following is only intended to show that under the unique and positive risk-neutral measure π⩛(θ) = {θ1, θ2} there is,
π⩛(θ1)Ainf Δ*(Ω)A 
π⩛(θ1)Esup Δ*(Ω)E,
where Δ*(Ω) denotes the set of convergent priors.
The formal implications of this measure will be detailed in the full paper. Let there be two assets, a stock s, and a bond z, with a risk-free return of 0.05. There are two times t and u, where u has two states, θ1 and θ2 ∊ Ω, with number of assets = 2. At u, s has an empirical return probability of 10% and −5% for θ1 and θ2, respectively; so, the risky investment of USD 100 at t becomes USD 110 and USD 95 against a riskless USD 105 at u. The loss matrices are, for actions a1 = invest in s and a2 = invest in z. Take the risk-neutral baseline matrix L = [−1 −1 1 1] and incorporate the 10% and −5% for θ1 and θ2, respectively; we obtain L → E = [−1.1 −1.05 0.95 1.05] → E = [−1.1 −1.05 3 1], where E tends to a canonical fully enriched Ellsberg-type = [−1 −1 3 2]. So, the payoff matrix (of the loss matrix E) [1.1 1.05 0.35 1.5], where riskless payoff a2 = [1.05 1.5] is uninteresting. There is 1.1/0.35 = 3.14 in the second equation below for the case of E, which corresponds to the risky asset column a1 = [−1, 3] ⇒ {−⅓ = −3 = (good state/bad state loss ratio) ≈ 1.1/0.35 = 3.14 (good state/bad state payoff ratio)} ⇒ risky payoff column a1 = [1.1 0.35] for the loss matrix E. For L → A = [−1.1 −1.05 0.95 1.05] → A = [−2 −2 1 1] ⇒ a1 = [−2 1] ⇒ {−2/1 = (good state/bad state loss ratio) ≈ 2/1 = 2 (good state/bad state payoff ratio)} ⇒ [2 1], so goes the second equation θ12 + θ2 = 1.05 of the A case below. So, the payoff matrix (of the loss matrix A) [2 1.3 1 1.05], where riskless payoff a2 = [1.3 1.05] is again uninteresting.
θ11.05 + θ21.05 = 1.05    —E case
θ11.1 + θ20.35 = 1.05
1 = 0.93 ≈ 0.9 ≈ sup Δ*(Ω)E, θ2 = 0.06}| π⩛(θ)E.
θ11.05 + θ21.05 = 1.05    —A case
θ12 + θ2 = 1.05
1 = 0.05 ≈ 0.1 ≈ inf Δ*(Ω)A, θ2 = 0.95}| π⩛(θ)A.
θ11.05 + θ21.05 = 1.05     —L case
θ1 + θ2 = 1.05
There is no measure for the risk-neutral loss matrix L, but every decile prior is convergent for L.
Under the calibrated enrichment weight, matrices WE and WA are used in the paper; the intersection between convergent and risk-neutral priors exists near the extremes of the decile grid used for the selection: Ellsberg-type preferences admit high admissible convergent priors (top decile), while Allais-type preferences admit low admissible convergent priors (bottom decile). Thus, when a unique positive risk-neutral measure exists and is used as the behavioral valuation kernel, it naturally maps to the top convergent decile for E-types (empirically close to 0.9, punishing undue risk aversion) and to the bottom convergent decile for A-types (empirically close to 0.1, penalizing undue love of certainty). The risk-neutral measure thus seals the lower convergent half for A-types into its limit case of the bottom decile and mutatis mutandis for E-types. So, this contextualizes iterative convergence under the spectral radius condition.
This is not a claim that 0.9 and 0.1 are universal; they are the model-consistent risk-neutral priors that emerge from our chosen W-calibration and the spectral stability criterion. Different enrichment weights, W, produce different risk-neutral priors. Note that this result—the risk-neutral priors’ smile from Allais to Ellsberg convergent priors—is owed to the second-order condition of choosing the highest convergent prior decile for the full enrichment of loss matrices.
Theorem 1 
(Nash Equilibrium). Let the game be M and let each player i hold a fully enriched loss matrix Li∗ and updates by the linear network rule Lt+1 = X Lt with fixed row-stochastic network X, and at each iteration, each player selects a prior π only from the finite decile grid Δ(Ω) ≡ Π = {0.1,…,0.9}. Define for each player i the convergent set, such that
ELi∗ = EL(πi∗(θ1)max) where ρ(Li*) < 1, πi∗(θ1)max ∈ Δ*(Ω)i
Suppose that W is positive and bounded so that the operator T(L) = X(W ∘ L) has a unique fixed point L∗. Then, the profile (πi∗){i =1, 5} is a Nash equilibrium, such that no player can unilaterally choose a different prior other than πi(t)max and obtain a strictly smaller fully enriched expected loss ELi.
Proof: 
Under the assumptions, X is row-stochastic and W is positive and bounded; the map T(L) = X(W ∘ L) is a (uniform) contraction on a suitable compact set of bounded 2 × 2 matrices (one may normalize by a uniform bound on |W| and |L|). Hence, there exists a unique fixed point L∗ with L∗ = T(L∗). The intervening lemma on contraction necessity and sufficiency will be provided in the full paper. Feasible prior sets are closed intervals. For each player i, its convergent priors set Δ*(Ω)i ⊂ Δ(Ω) is non-empty by construction (we assume that at least one decile satisfies ρ < 1), and in the decile grid it takes the form of an initial segment {0.1,…,πi max}. So, πi max* = maxΔ*(Ω)i is well-defined.
Expected loss is affine in π (hence, there are extremal minimizers). That is, for a fixed Li, each action’s expected loss ELi(a)(π) is affine in π. Consequently, the vector of action losses is minimized (or the criterion used to select EL*, e.g., the minimum across actions or the mean) over the discrete interval (with practically no loss of generality in the continuous case; more detail will be provided in the full version). Δ*(Ω)i attains its minimum at an endpoint of that interval. By the game’s selection rule (highest convergent decile), the chosen endpoint is πi max*.
For the essential Nash property of no profitable unilateral deviation, see the following. Suppose player i deviates to some other admissible prior πi*~ ∈ Δ*(Ω)i with πi*~ ≠ πi max*. Because the objective (EL*) is affine in π and we minimized it over Δ*(Ω)i by selecting the endpoint πi max*, we have ELi∗(πi max*) ≤ ELi∗(πi*~). Thus, the deviation cannot strictly decrease player i’s fully enriched convergent expected loss. (If multiple deciles tie at the minimum, the tie-breaking rule picks the highest decile; thus, any deviation to a lower decile cannot improve the objective.) Therefore, no player can strictly improve by unilaterally changing its prior from πi max*, which means that (πi max*) is a Nash equilibrium.
Corollary ⇒ Uniqueness under Strict Monotonicity: If, in addition, for each player i the function ELi(π) is strictly monotone on the admissible interval, as is the case (no ties at the minimizing endpoint), then the Nash equilibrium (πi max*) is unique. □

3. Computational Results

The following are computations of the iterations for the above-mentioned three scenarios of the game.
By the third iteration of homophily in Figure 1 and Figure 2, E-type players (green) win identically by minimizing the convergent expected loss the most as compared to A-type players (purple).
Heterophilous iterates in Figure 3 and Figure 4 produce fluctuations but converge quickly toward near-equal EL*s; E-type players (green) identically edge out slightly by iteration 3 as compared to A-type players (purple).
Type-neutral randomness in Figure 5 provides a full taste of the five-player stochastic network dynamics. Player A2 wins uniquely by the third iteration, and the TNR scenario shows the fastest completion of learning with within-type differences.
Epistemic Value of the Game M: Consider a game ¬M with only one A and one E player. In the case of heterophily, after one iteration the matrices stop changing with same loss matrix {−1.5 −1.5 2 1.5}; in the Allais–Ellsberg convex combination case, under the transition matrix T = {0.5 0.5 0.5 0.5}, EL*E = EL*A = {−1.08 −1.14} with π*(θ1)max = 0.88.
An Application (Inverse Problems in Mean-Field Games): Let the mean-field operator be
T ( L ) :   = X ( ξ ) [ W ( ξ )   D ( π * L ) L ( ξ ) ] d ξ .
If ρ ( T ) < 1 , then the mean-field game admits a unique equilibrium L * with π * L * , revealing true utility. X(ξ) is the mean-field interaction kernel (limit of the Markov network), aggregating how a player’s updated losses influence the game population; L(ξ) is the loss matrix of a representative player indexed by type ξ (Allais or Ellsberg, possibly with heterogeneity).
Sketch of Proof. 
ρ ( T ) < 1 implies that T is a contraction; hence, L t L * . Since expected loss is affine in π , minimization over the convergent set selects its extreme (top for E, bottom for A), so any biased prior worsens convergence, and no deviation improves the inverse fit.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The raw data (originally created) supporting the conclusions of this article will be made available by the author on request.

Conflicts of Interest

The author declares no conflicts of interest.

References

  1. Allais, M. Le comportement de l’homme rationnel devant le risque: Critique des postulats et axiomes de l’école Américaine. Econometrica 1953, 21, 503–546. [Google Scholar] [CrossRef]
  2. Ellsberg, D. Risk, ambiguity, and the Savage axioms. Q. J. Econ. 1961, 75, 643–669. [Google Scholar] [CrossRef]
  3. Machina, M.J. Risk, ambiguity, and the rank-dependence axioms. Am. Econ. Rev. 2009, 99, 385–392. [Google Scholar] [CrossRef]
  4. Gilboa, I.; Schmeidler, D. Maxmin expected utility with non-unique prior. J. Math. Econ. 1989, 18, 141–153. [Google Scholar] [CrossRef]
  5. Bala, V.; Goyal, S. Learning from neighbors. Rev. Econ. Stud. 1998, 65, 595–621. [Google Scholar] [CrossRef]
  6. Golub, B.; Jackson, M.O. How homophily affects the speed of learning and best-response dynamics. Q. J. Econ. 2012, 127, 1287–1338. [Google Scholar] [CrossRef]
  7. Gale, D.; Kariv, S. Bayesian learning in social networks. Games Econ. Behav. 2003, 45, 329–346. [Google Scholar] [CrossRef]
Figure 1. Homophily priors’ evolution across iterations.
Figure 1. Homophily priors’ evolution across iterations.
Proceedings 135 00002 g001
Figure 2. Homophily.
Figure 2. Homophily.
Proceedings 135 00002 g002
Figure 3. Heterophily priors’ evolution across iterations.
Figure 3. Heterophily priors’ evolution across iterations.
Proceedings 135 00002 g003
Figure 4. Heterophily.
Figure 4. Heterophily.
Proceedings 135 00002 g004
Figure 5. Type-neutral randomness.
Figure 5. Type-neutral randomness.
Proceedings 135 00002 g005
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Mughal, A.A. Allais–Ellsberg Convergent Markov–Network Game. Proceedings 2026, 135, 2. https://doi.org/10.3390/proceedings2026135002

AMA Style

Mughal AA. Allais–Ellsberg Convergent Markov–Network Game. Proceedings. 2026; 135(1):2. https://doi.org/10.3390/proceedings2026135002

Chicago/Turabian Style

Mughal, Adil Ahmad. 2026. "Allais–Ellsberg Convergent Markov–Network Game" Proceedings 135, no. 1: 2. https://doi.org/10.3390/proceedings2026135002

APA Style

Mughal, A. A. (2026). Allais–Ellsberg Convergent Markov–Network Game. Proceedings, 135(1), 2. https://doi.org/10.3390/proceedings2026135002

Article Metrics

Back to TopTop