1. Introduction
To meet the rapidly growing capacity demands of next-generation wireless systems, advanced multi-antenna techniques such as massive MIMO have become essential. Massive MIMO leverages large-scale antenna arrays to enable highly directional beamforming, significantly enhancing spectral efficiency and interference suppression. By coherently coordinating a large number of antenna elements, it provides a practical path toward realizing theoretical capacity gains in real deployments [
1,
2]. However, fully digital beamforming architectures require a dedicated RF chain per antenna, leading to high power consumption, hardware complexity, and increased implementation cost as the array size scales [
3]. These limitations motivate scalable and distributed beamforming approaches in which antennas or subarrays cooperate rather than operate independently.More generally, MIMO systems enable simultaneous transmission and reception of multiple data streams by exploiting multipath propagation. Unlike SISO systems, which are restricted to a single spatial stream [
4], MIMO improves data rates, reliability, and coverage by converting multipath propagation into a performance advantage. Through spatial diversity and coherent signal combination, system throughput increases with the number of antennas [
5,
6].
In the context of future 6G networks, MIMO is expected to play an even more critical role. Massive MIMO improves spectral efficiency by supporting multi-user transmission, while cooperative full-duplex MIMO can potentially double spectral efficiency by enabling simultaneous transmission and reception [
7,
8]. Additionally, multiband MIMO enables operation across multiple frequency bands through multi-resonant antenna structures, reducing interference and improving hardware flexibility [
9,
10,
11]. Looking forward, 6G networks will support extremely dense device connectivity, enabling applications such as IoT, smart cities, and large-scale data-driven services [
12]. This increased density will require advanced antenna architectures capable of managing massive connectivity while maintaining performance and reliability. In this context, massive MIMO, dynamic beamforming, and spatial division techniques will be essential for efficient spectrum utilization and interference management [
13]. Achieving these goals will require dense, scalable, and energy-efficient network infrastructures, making antenna design a key enabler of future wireless evolution.
In this paper, we develop a sociotropy-inspired distributed beamforming framework for MIMO systems, in which antenna elements are modeled as interdependent strategic agents rather than isolated optimizers. The proposed formulation introduces inter-agent coupling into the beamforming process, allowing each antenna to adapt its transmission strategy by jointly accounting for individual signal contribution and local alignment with neighboring elements. Within this context, classical sociotropic [
14] notions are translated into engineering constructs, where sensitivity to collective feedback is captured through interference and misalignment-aware adaptation mechanisms that encourage coordinated behavior.
Moreover, we design utility functions that explicitly promote cooperative alignment while discouraging overly aggressive unilateral optimization. By implicitly incentivizing antennas to conform to network-level objectives—such as interference mitigation and coherent beam pattern formation—the resulting game-theoretic model, using a potential game [
15], permits stable equilibrium solutions with improved convergence properties. This approach enables a scalable and robust distributed beamforming architecture that effectively balances local performance with global system efficiency, making it well suited for large-scale MIMO deployments. For performance evaluation, the proposed sociotropic beamforming framework is benchmarked against three representative schemes: a non-cooperative selfish strategy, a conventional selfish baseline, and a learning-based MARL-GNN coordination approach. The comparative analysis highlights that sociotropic coordination consistently enhances interference management and system robustness, particularly under imperfect channel state information. At the same time, it achieves performance that is comparable to the learning-based approach, while maintaining significantly lower computational and implementation complexity, making it well-suited for practical distributed MIMO deployments. Classical beamforming approaches such as MRT, ZF, and RZF provide closed-form solutions but lack adaptability to evolving interference and channel conditions in dynamic wireless environments. In contrast, the proposed game-theoretic and learning-based framework enables distributed coordination among antennas, allowing the system to continuously refine beamforming strategies in response to temporal channel variations and interaction effects.
4. System Model
We consider a MIMO downlink scenario with a single base station (BS) equipped with
transmit antennas serving
K single-antenna users. The transmitted signal vector is
where
is the information symbol for user
k with
, and
is the beamforming vector.
The received signal at user
k is
where
is the channel vector and
is additive noise.
The SINR for user
k is
with achievable rate
To ensure mathematical completeness of the proposed optimization framework, we explicitly provide the structure of the gradient of the sum-rate with respect to each antenna-wise beamforming vector. The sum-rate is defined as
The gradient with respect to each conjugate beamforming vector
is computed using Wirtinger calculus as
where each SINR derivative captures both the desired signal contribution and the multi-user interference coupling induced by antenna
i. This expression directly follows from the chain rule applied to the logarithmic rate function under complex-valued differentiation. The explicit closed-form expressions of
are consistent with standard multi-user MIMO beamforming derivations and are omitted here for compactness but are used directly in the projected gradient and KKT computations in subsequent sections.
All gradient computations in this work are performed using Wirtinger calculus, which is the standard approach for optimization over complex-valued variables in wireless beamforming problems. Specifically, for a real-valued objective function defined over complex matrices, gradients are taken with respect to the conjugate variables , i.e., , ensuring mathematically consistent descent/ascent directions in the complex domain.
We assume a centralized CSI acquisition at the base station, consistent with massive MIMO architectures. However, the optimization is implemented in a distributed fashion at the antenna level, where each antenna updates its local beamforming row using locally accessible gradient components derived from the globally computed sum-rate. Thus, while CSI acquisition is centralized, the decision-making process is decomposed into antenna-wise updates, making the scheme a semi-distributed optimization framework rather than a fully decentralized system.
To avoid ambiguity, we explicitly define the beamforming matrix
We also define the antenna-wise representation
with the explicit mapping:
Thus, both representations are equivalent views of the same matrix W.
Each antenna
is modeled as a strategic agent controlling its row vector
. The full beamforming matrix is defined as above. Each antenna satisfies the constraint
Sociotropic beamforming refers to a coordination mechanism in which antenna agents jointly optimize system-level performance while incorporating explicit regularization that encourages structured similarity in their beamforming decisions.
Unlike conventional regularized sum-rate maximization or standard potential-game formulations, the proposed framework introduces a sociotropic beamforming structure, where antenna agents are not only coupled through interference-aware sum-rate maximization but also through an explicit cooperative alignment mechanism in the antenna-wise parameter space. Specifically, the novelty lies in three aspects: (i) antenna-level game decomposition with row-wise strategies in the beamforming matrix, (ii) integration of a structured coordination term that couples antenna strategies beyond classical power regularization, and (iii) a unified potential formulation that preserves exact-game properties while embedding cooperative spatial alignment constraints. This combination enables a hybrid design that bridges distributed game-theoretic beamforming and structured inter-antenna coordination beyond conventional regularized or standard potential-game beamforming models.
The utility of antenna
i is defined as
where
,
, and
.
The term introduces a controlled alignment between antenna-wise beamforming vectors. While this promotes structured cooperation and improves stability of distributed updates, we acknowledge that excessive penalization may reduce spatial diversity in highly heterogeneous channel conditions. To mitigate this, the parameter acts as a tunable coordination strength that allows the system to interpolate between fully independent antenna behavior () and fully aligned antenna strategies ( large). This ensures that the model does not enforce strict similarity but rather regulates inter-antenna coherence depending on channel correlation regimes.
5. Potential Game Formulation and Distributed Updates
The relationship between sociotropic coordination and interference management is inherently embedded in the SINR structure of the system model. Specifically, each antenna-wise update affects all users through the interference terms in
The sociotropic coordination term introduces an additional coupling among antenna strategies, which indirectly regulates interference generation by discouraging uncoordinated beamforming updates that amplify cross-user interference. To make this relationship explicit, the potential function can be interpreted as a weighted combination of (i) sum-rate maximization capturing interference suppression, and (ii) coordination regularization enforcing structured beamforming behavior across antennas. This dual interpretation directly links the proposed framework to interference-aware distributed optimization.
A game with utility functions
is an exact potential game if there exists a function
such that, for any antenna
i and any unilateral deviation
The potential function is given by
Consider a unilateral deviation of antenna
i, changing
. The change in utility is
The change in the potential function is
Using symmetry
, we obtain
so the game is an exact potential game.
Proposition 1. The proposed sociotropic beamforming game admits at least one pure-strategy Nash equilibrium.
Proof. Each feasible set is compact and convex. The potential function is continuous and bounded above on this compact set. Hence, it attains a maximum. Any maximizer of is a pure-strategy Nash equilibrium. □
Although the existence of a pure-strategy Nash equilibrium is guaranteed by the exact potential game structure, this equilibrium is not necessarily globally optimal with respect to the original sum-rate maximization problem. The inclusion of regularization and coordination terms in the potential function modifies the optimization landscape, meaning that the maximizer of may differ from the sum-rate maximizer . Therefore, the equilibrium characterizes a trade-off between communication efficiency, power regularization, and inter-antenna coordination rather than a strict global optimum of the sum-rate objective.
Each antenna performs projected gradient ascent
where gradients are taken using Wirtinger calculus with respect to
. Under the assumptions that
is continuously differentiable,
is Lipschitz continuous on the feasible set, and
,
, the iterates converge to a stationary point of
, which corresponds to a Nash equilibrium of the exact potential game.
9. Results
To resolve potential ambiguity between the methodological description and the numerical implementation, we clarify that the MARL-GNN framework should be interpreted as a unified modeling abstraction rather than two separate algorithmic entities. At the methodological level, MARL-GNN is formulated as a multi-agent reinforcement learning framework incorporating policy gradients, advantage functions, and graph neural network-based message passing to capture antenna interactions in a stochastic optimization setting. However, in the numerical evaluation, this framework is implemented in its deterministic equivalent form as a graph-regularized distributed gradient dynamics scheme. In this implementation, the policy-gradient update reduces to a structured beamforming ascent step augmented by graph-induced coordination terms derived from antenna similarity relations. Therefore, the MARL-GNN baseline does not correspond to an independently trained reinforcement learning model separate from the proposed dynamics. Instead, it represents a graph-structured parameterization of the same underlying update rule, consistent with the exact potential game formulation.
We evaluate DSB in a MIMO downlink with antennas and users. Channels follow Rayleigh fading , with spatial correlation and additive perturbations . Per-antenna power is , noise variance , and DSB uses step-size over 80 iterations. Parameters are , , and ; is the selfish baseline.
The numerical evaluation is performed for a moderate-scale MIMO configuration with antennas and users. This choice is consistent with standard controlled experimental setups in game-theoretic beamforming literature, where the focus is on isolating the effect of coordination mechanisms under tractable simulation conditions. Importantly, the proposed formulation is not limited to this dimensionality. The sociotropic beamforming model and associated potential game structure are defined for arbitrary values of and K, as reflected in the system model formulation. The distributed antenna-wise update rule scales linearly with the number of antennas, while interference coupling scales with the number of users, making the approach inherently extensible to larger systems.
A selfish beamforming scheme is considered as a reference baseline. In this approach, each transmit antenna aligns its beamforming vector with the corresponding user channel in order to maximize the received signal power, without explicitly accounting for inter-user interference or coordination among antennas. This method serves as a widely used reference in multi-user MIMO systems, providing a simple yet effective strategy for signal enhancement, against which the performance of the proposed distributed and coordinated beamforming schemes can be evaluated.
In addition, a MARL-GNN-based coordination scheme is also considered for comparison. In this approach, a multi-agent reinforcement learning framework is combined with a graph neural network to model the interactions among transmit antennas, enabling adaptive and context-aware coordination. Each antenna is treated as an agent, while the GNN captures the underlying interaction topology based on the system state, allowing antennas to exchange information and adjust their beamforming strategies accordingly. Importantly, this scheme operates on top of the same system model and constraints, serving as an enhanced coordination mechanism.
The numerical results are obtained through a stochastic simulation that directly implements the analytical system model and game-theoretic formulation. Specifically, the MIMO downlink channel model, SINR definition, and achievable rate expressions are evaluated at each iteration to compute and the aggregate sum-rate as defined in the theoretical model. The sociotropic utility function is implemented as a composite objective consisting of the sum-rate term, per-antenna transmit power regularization, and pairwise beamforming alignment penalty, which together reproduce the structure of the potential function . The baseline selfish and sociotropic strategies are realized through projected gradient ascent updates, which act as first-order approximations of the corresponding KKT stationarity conditions; accordingly, the reported KKT residual is computed as the norm of the gradient mapping and serves as a practical measure of convergence toward stationary points. The MARL-GNN extension is implemented as a graph-regularized interaction mechanism, where a similarity-induced adjacency matrix over beamforming vectors defines a message-passing operator that aggregates neighboring antenna states. This operator modifies the local gradient dynamics by incorporating structured coordination effects, consistent with the alignment term in the utility, rather than introducing a separate learning architecture. Overall, the simulation should be interpreted as a stochastic projected-dynamics realization of the exact potential game under centralized CSI, where convergence trends reflect empirical stationarity of the induced dynamical system rather than exact closed-form Nash equilibrium computation.
Figure 1 illustrates the evolution of the system sum-rate over the iterations for the considered distributed beamforming strategies with eight transmit antennas and four users. At the beginning of the optimization process, all methods start from similar randomly initialized beamforming vectors, yielding an initial sum-rate close to 2 bit/s/Hz. As the iterations progress, the sum-rate increases steadily as each antenna updates its beamforming vector while satisfying the per-antenna power constraint. The selfish strategy exhibits the slowest improvement since each antenna optimizes its transmission independently without considering the impact on the other antennas. Around iteration 100, the selfish approach achieves approximately 8 bit/s/Hz, and by iteration 200 the sum-rate increases to about 10.5 bit/s/Hz. The improvement continues gradually and approaches a steady value near 12.4 bit/s/Hz by iteration 500. In contrast, the sociotropic strategy demonstrates faster growth because each antenna accounts for the collective system performance during the update process. At iteration 100, the sociotropic method reaches roughly 8.7 bit/s/Hz and improves further to around 11 bit/s/Hz by iteration 200. The sum-rate continues to increase smoothly and stabilizes close to 12.5 bit/s/Hz toward the end of the iterations. The MARL-GNN approach follows a similar trajectory but with slightly improved convergence behavior due to the graph-based coordination among antennas. At iteration 100, the MARL-GNN method achieves approximately 8.9 bit/s/Hz, which is higher than both the selfish and sociotropic strategies. By iteration 200, the sum-rate reaches about 11.1 bit/s/Hz and maintains a small performance advantage during the subsequent iterations. As the algorithm approaches convergence, the difference between the coordinated strategies becomes smaller because all methods tend toward a similar equilibrium determined by the system model. At iteration 500, the MARL-GNN scheme attains roughly 12.55 bit/s/Hz, the sociotropic strategy converges near 12.5 bit/s/Hz, and the selfish method stabilizes around 12.4 bit/s/Hz.
Figure 2 depicts the evolution of the potential function over the iterations for the distributed beamforming strategies under the considered system configuration with eight transmit antennas and four users. At the initial stage of the algorithm, the potential values differ significantly among the three strategies due to the different coordination behaviors of the antennas. The sociotropic and MARL-GNN strategies start with a potential value close to
, reflecting the large misalignment among the randomly initialized beamforming vectors and the corresponding penalty introduced by the coordination term. In contrast, the selfish strategy begins at a much higher value of approximately
, since it does not incorporate any coordination penalty and therefore does not account for the misalignment between antennas. As the iterations progress, the potential values increase steadily for all strategies as the beamforming vectors adapt to improve the system performance. For the sociotropic approach, the potential increases rapidly during the first 100 iterations, moving from about
to approximately
, and then continues to grow more gradually, reaching roughly 5 around iteration 200. The growth remains smooth throughout the remaining iterations and eventually stabilizes close to
by iteration 500. The MARL-GNN strategy follows a very similar trajectory, starting near
and increasing to around
at iteration 100, then reaching approximately
near iteration 200. Toward the end of the optimization process, the potential converges to about
, which is slightly higher than the sociotropic strategy. The selfish approach, conversely, begins at about
and increases more quickly during the early iterations, reaching nearly
around iteration 100 and approximately 10 at iteration 200. The growth then slows down as the algorithm approaches convergence, stabilizing close to
by iteration 500. These results indicate that although the selfish strategy achieves a larger potential value due to the absence of coordination penalties, the coordinated approaches progressively reduce the misalignment among antennas and achieve stable convergence. Furthermore, the MARL-GNN method exhibits a marginal improvement over the sociotropic strategy, suggesting that the graph-based coordination mechanism can slightly enhance the cooperative behavior of the antennas while maintaining similar convergence characteristics.
Figure 3 illustrates the evolution of the KKT residual across the iterations for the considered distributed beamforming system composed of eight transmit antennas serving four users. The KKT residual measures how close the iterative beamforming updates are to satisfying the optimality conditions of the formulated optimization problem under the adopted system model. At the beginning of the iterative process, the residual values are relatively high for all strategies due to the random initialization of the beamforming vectors and the resulting mismatch with the optimality conditions. Specifically, the MARL-GNN strategy starts with the largest residual of approximately 3.35, while the sociotropic approach begins near 3.05 and the selfish strategy slightly below at around 2.95. This difference reflects the stronger exploratory behavior of the learning-based method during the early stage of the optimization. During the first phase of the iterations, a rapid decrease in the residual is observed for all strategies, indicating that the beamforming vectors quickly move toward a feasible stationary region of the optimization landscape. By iteration 50, the residual values have already dropped to roughly 1.7 for the sociotropic approach, 1.75 for the selfish strategy, and about 1.6 for the MARL-GNN approach. The decline continues as the distributed updates adapt to the channel conditions defined in the system model. Around iteration 100, the residual values decrease further to approximately 0.95 for the sociotropic strategy, 1.0 for the selfish case, and about 0.85 for the MARL-GNN method, suggesting that the cooperative and learning-based strategies guide the beamforming updates slightly more efficiently toward the optimality region. Beyond this point, the reduction becomes more gradual, reflecting the typical convergence behavior of iterative optimization algorithms approaching a stationary point. At iteration 200, the residuals are already below 0.5, with approximate values of 0.46 for the sociotropic strategy, 0.42 for the selfish approach, and 0.48 for the MARL-GNN scheme. As the iterations continue, the residual values keep decreasing slowly while stabilizing, demonstrating that the algorithms are approaching the KKT conditions of the formulated beamforming optimization problem. Near iteration 300, the residuals reach approximately 0.34 for the sociotropic strategy, 0.29 for the selfish strategy, and 0.36 for the MARL-GNN approach. Toward the final stage of the process, the curves become almost flat, indicating convergence. At iteration 500, the residual values settle around 0.24 for the sociotropic strategy, approximately 0.21 for the selfish approach, and close to 0.23 for the MARL-GNN method. These results confirm that all three strategies converge to solutions that satisfy the KKT optimality conditions of the system model, with only minor differences in the convergence speed.
Figure 4 illustrates the gain relative to the selfish strategy over the iterations for both the sociotropic and the MARL-GNN approaches in the considered. At the initial stage, both coordinated strategies already provide a positive gain compared to the selfish case, with the sociotropic method starting at approximately
bit/s/Hz and the MARL-GNN approach slightly lower at around
bit/s/Hz. As the iterations progress, the gain increases rapidly due to the coordinated adjustment of the beamforming vectors. The MARL-GNN strategy exhibits a sharp rise during the first 50 iterations, reaching a peak gain of approximately
bit/s/Hz around iteration 40, while the sociotropic method increases more gradually to about
bit/s/Hz in the same region. Between iterations 50 and 100, both methods maintain high gains, with the MARL-GNN approach stabilizing near
to
bit/s/Hz and the sociotropic strategy fluctuating slightly around
to
bit/s/Hz. After this peak region, the gain for both strategies begins to decrease gradually as all methods move toward a similar equilibrium. Around iteration 200, the sociotropic gain reduces to approximately
bit/s/Hz, while the MARL-GNN gain remains higher at about
bit/s/Hz. This decreasing trend continues steadily, and by iteration 300 the gains drop to roughly
bit/s/Hz for the sociotropic approach and about
bit/s/Hz for the MARL-GNN method. In the later stages of the optimization, the gains continue to diminish as the difference between the coordinated and selfish strategies becomes smaller. Around iteration 400, the sociotropic gain approaches
bit/s/Hz, while the MARL-GNN gain is still slightly higher at approximately
bit/s/Hz. Finally, near iteration 500, both curves approach zero, with the sociotropic strategy slightly underperforming the selfish approach at about
bit/s/Hz and the MARL-GNN method close to
bit/s/Hz. These results indicate that coordinated strategies provide significant performance gains during the transient phase of the optimization, particularly in the early and mid iterations, with the MARL-GNN approach achieving the highest peak improvement. However, as the system converges to a common equilibrium dictated by the system model and constraints, the relative advantage diminishes and eventually vanishes.
The sociotropic and MARL-GNN methods primarily provide gains in the transient regime, while the performance gap with the selfish baseline diminishes as the system converges. This behavior is consistent with the theoretical formulation of the proposed potential game, where all strategies evolve toward stationary points of the same underlying wireless system model under identical SINR and power constraints. As a result, the selfish strategy can partially recover performance at convergence due to the absence of coordination penalties, which reduces the gap in the final iterations. Importantly, the objective of the proposed sociotropic coordination is not to guarantee strict dominance at convergence, but to improve convergence behavior, stability, and transient system efficiency under interference coupling. Therefore, the slight performance convergence or crossover near steady state does not weaken the contribution, but rather highlights the inherent trade-off between coordination-induced regularization and asymptotic sum-rate optimality.
Figure 5 presents the evolution of the network sum-rate performance over 500 iterations for the proposed learning-based approaches (sociotropic, selfish, and MARL-GNN) and the benchmark schemes (MRT, ZF, and RZF). The learning-based methods start from approximately 2 bit/s/Hz and progressively improve their performance, converging to stable values between
and
bit/s/Hz. Specifically, MARL-GNN achieves the highest final sum-rate of approximately
bit/s/Hz, followed by sociotropic (
bit/s/Hz) and selfish (
bit/s/Hz), corresponding to an improvement compared to their initial performance. In contrast, the conventional benchmark schemes exhibit a significant performance degradation over time. MRT decreases from approximately
bit/s/Hz to
bit/s/Hz, while ZF and RZF decline from approximately 26 bit/s/Hz to below 1 bit/s/Hz, corresponding to performance losses exceeding
. These results demonstrate the superior adaptability and long-term performance of the proposed learning-based approaches, with MARL-GNN consistently providing the best overall throughput among all evaluated methods.
10. Conclusions
This work proposed a sociotropy-inspired potential game for distributed MIMO antenna design, where each antenna is modeled as an adaptive agent whose utility integrates system throughput, interference regulation, and behavioral alignment with neighboring antennas. By embedding sociotropic interaction principles into the beamforming optimization process, the framework establishes a conceptual bridge between social cooperation models and electromagnetic system design. We derived the associated potential function, characterized the equilibrium structure, and developed update dynamics with convergence toward a stable operating point under standard regularity conditions. Numerical investigations demonstrated stable behavior, highlighting the benefits of cooperative regularization in distributed antenna optimization.
The numerical results confirm that incorporating sociotropic interactions into distributed beamforming leads to a more stable and structured optimization process compared to purely selfish updates. The observed sum-rate evolution demonstrates that coordination among antennas significantly improves transient performance, particularly in the early and intermediate iterations, where alignment-aware updates accelerate system-wide throughput growth. This behavior is consistent with the underlying mechanism of the sociotropic formulation, which couples antenna decisions through a collective objective, effectively reducing destructive interference and improving beamforming coherence. The potential function evolution further shows that this coordination acts as a stabilizing mechanism, guiding the system from highly misaligned initial states toward a well-structured equilibrium, while avoiding the irregular dynamics often associated with uncoordinated optimization. In addition, the monotonic reduction of the KKT residual across all strategies confirms that the proposed distributed updates converge reliably to stationary solutions satisfying the optimality conditions of the underlying problem. Although the MARL-GNN approach achieves a marginal advantage in certain regimes, the sociotropic formulation maintains competitive performance while offering a more interpretable and analytically grounded coordination structure. Importantly, the diminishing performance gap in later iterations indicates convergence toward a common equilibrium, where individual and collective objectives become aligned. These findings validate the effectiveness of sociotropic design as a scalable and energy-aware coordination mechanism for distributed MIMO systems, ensuring both convergence stability and improved transient efficiency without requiring centralized control.
Moreover, the results demonstrate the effectiveness of the proposed multi-agent learning framework in maximizing network throughput under dynamic operating conditions. While conventional beamforming schemes such as MRT, ZF, and RZF exhibit substantial performance degradation, the learning-based approaches consistently improve their performance and converge to stable high-throughput operating points. Among the evaluated methods, MARL-GNN achieves the highest final sum-rate, reaching approximately bit/s/Hz and outperforming the benchmark techniques by a significant margin. These findings highlight the ability of graph-based multi-agent reinforcement learning to capture complex interactions among network entities and to adapt efficiently to changing environments. Future work will focus on extending the framework to larger-scale deployments, investigating robustness under varying channel conditions, and incorporating additional constraints related to energy efficiency and fairness.
Future research directions include reinforcement learning–based adaptive strategies capable of learning sociotropic interaction parameters online, as well as experimental validation on reconfigurable intelligent surface (RIS) platforms and large-scale distributed MIMO testbeds. Extending the framework to dynamic channel environments and partially observable settings also constitutes a promising avenue toward practical 6G implementations.