1. Introduction
The ongoing trend toward decentralized electricity production is transforming how power systems are planned and operated, and it is creating a need for completely new techniques and approaches to manage increasingly complex energy networks [
1]. As distributed energy resources (DERs) continue to expand, relying solely on traditional supply-side control in no longer sufficient to maintain a reliable balance between generation and consumption [
2,
3]. This challenge is amplified by the growing share of renewable sources such as solar, wind and hydropower. Although these technologies are central to long-term sustainability objectives but their output can fluctuate significantly because of dependence on weather and other uncontrollable intermittent effects [
4]. Such variability places additional stress on grid operation and motivates the adoption of management strategies that can respond quickly and intelligently to changing conditions [
5].
Demand-side management (DSM) has therefore become an essential pillar in modern gird operation [
6]. State-of-the-art (SoTA) DSM frameworks extend beyond simple consumer response to price signals as they aim to support system-level reliability by shaping demand in ways that complement variable and every changing renewable supply. In this context, intelligent energy management system provide monitoring and control capabilities that can reduce peak demand along with improvement of load flexibility and increase the effective utilization of renewable generation [
7,
8]. The deployment of energy storage systems (ESSs) also strengthens system robustness by buffering the mismatches between supply and demand. Storage enables surplus energy to be retained when generation is high and dispatched later when demand increases or renewable output declines [
9].
These considerations are especially important at the district scale as well where multiple buildings coexist with diverse usage patterns, storage capacities along with on-site generation profiles and where they are connected through shared infrastructures or through the wider grid [
10]. Managing buildings independently can prevent the system from exploiting coordination benefits across the district. In contrast, coordinated DSM across multitude of buildings can improve the overall efficiency an d lead to better district-level outcomes [
11]. When buildings operate within shared structures such as microgrids, communication and coordination can support improved energy allocation and more effective use of local resources [
12]. By jointly accounting for building-level requirements and broader grid constraints, coordinated district operation can increase efficiency and facilitate renewable integration and also contribute to improved system stability [
13,
14].
A convenient way to formalize coordinated control at the district scale is to model the setting as a multi-agent system (MAS), where each building is treated as an agent that makes decisions using local measurements and, when available, shared district information. MAS designs are commonly grouped into three architectural families which are classified as decentralized, centralized and cooperative. In a decentralized architecture, agent select actions independently which typically improves scalability because computations and decision-making are distributed in nature. But a key drawback for this is that learning and control can become non-stationary when multiple agents adapts simultaneously. This effect is particularly prevalent if each agent observes only local state rather than district level context. this limited system awareness can lead to sub-optimal coordination and reduced overall performance [
15]. In a centralized architecture, there is a single decision-making entity which controls the entire system and has access to global information across all agents. This full observability can support more informed coordination since the controller can optimize actions with respect to district-level objectives. However, centralized control becomes difficult to scale as the number of agents grows which is typically the case. The dimensionality of the joint observation–action space expands rapidly with each additional agent, increasing computational burden and making it challenging for a single controller to produce timely decisions in large deployments. Cooperative designs aim to combine the strengths of both the centralized and decentralized schemes. Each agent retains its own controller fr execution but policy learning and evaluation are performed in a way that accounts for the behavior of other agents and this enables coordination without collapsing the entire problem into a single monolithic controller. This structure supports more consistent collective behavior by allowing agents to reason about how their actions interact [
16]. Cooperative multi-agent control has demonstrated strong potential in energy applications which includes coordination of microgrids with solar generation and EV charging [
17,
18], as well as strategies that reduce peak demand while increasing renewable utilization through the joint operation of distributed local controllers [
19].
A closely related and active subfield in the domain of multi-agent systems that is gaining traction is multi-agent reinforcement learning (MARL) which offers a principled way to formulate and solve decision-making which is sequential and involves multiple interacting agents. In MARL, agents improve their behavior and this done through interactions that are repeated in nature. This is done through feedback signals to refine policies that can simultaneously reflect local objectives and system-level goals [
20]. A growing body of work has shown that this learning paradigm is very will suited to energy-management settings where decisions are coupled by shared infrastructure and where uncertainty and variability are intrinsic in nature. An example of this in the literature is the implementation of scalable MARL actor-critic framework with the core objective of residential energy flexibility. The study findings report a 47.2% reduction in under-voltage events while also reducing the electricity costs and all of this is done without the requirement of exchange of private user data [
21]. Addressing scalability from another angle, subsequent work proposed a distributed coordination method that isolates marginal reward contributions, allowing prosumers to quantify how their actions influence district objectives while maintaining privacy. By utilizing this formulation, the authors reported the improvements across several operational and economic indicators and these include reduced import costs, lower distribution losses and decreased battery degradation and also the reduction of greenhouse gas emissions [
22].
Recently, many researchers have reported much success with cooperative MARL systems for energy management of microgrids. In these cooperative frameworks, distributed agents coordinate to balance local energy demands with global grid stability [
23,
24]. This coordination is particularly important in interconnected microgrids where local renewable generation and battery storage must be jointly optimized to prevent grid congestion and reduce reliance on centralized power plants. Researchers have increasingly turned to the use of attention-based reinforcement learning architectures to manage the increased complexity of these cooperative systems [
25]. Many traditional multi-agent models have difficulty determining the impact of one agent on the larger system; however, attention mechanisms allow agents to assign different weights to information from other agents to improve the identification of association between agents and help to solve the problems of credit assignment, and thus provide more stable training and updates for agents.
Furthermore, the field has seen a recent shift toward transformer-based control strategies for energy applications [
26,
27]. By treating energy management as a sequence modeling problem, transformer architectures excel at capturing long-term temporal dependencies in weather patterns, building loads, and dynamic pricing data. Our proposed AttentionKAN framework builds directly upon these recent advancements.
MARL has also been explored in microgrid operation where coordination among storage units and higher-level aggregators can improve renewable energy utilization and cost efficiency. In this context, agents controlling energy storage and microgrid coordination have been trained using deep reinforcement learning such as deep deterministic policy gradient (DDPG) and multi-agent extensions like Multi-Agent Multi-Agent Deep Deterministic Policy Gradient (MADDPG), demonstrating improved cooperation and performance relative to single-agent alternatives [
28]. MARL has been used in residential community settings which is used to manage renewable uncertainty and avoid rebound peaks by coordinating household schedules. In such formulations each building is modeled as an agent that seeks to reduce electricity cost while on the other hand also preserve and respect user comfort and this coodination leads to lower community-wide cost and better peak mitigation through coordinated use of renewable generatio [
29]. Related studies have also emphasized the value of coordination in multi-building energy systems even when not framed strictly as MARL. Work combining surrogate models with deep reinforcement learning highlighted the role of coordinated control in reducing both operating costs and peak demand in district-scale management [
30]. Additional comparisons between coordinated and cooperative DRL controllers in district energy systems have similarly shown consistent improvements over rule-based strategies which include lower costs, reduced demand peaks and increased self-consumption, further supporting the case for intelligent, learning-based coordination in multi-building environments [
31].
Table 1 are some of the studies related to MARDL for energy management and depicts the scale of studies, along with the settings and assets which are controlled along with KPIs measured and limitations in them. Beyond residential MARL-based BESS control, recent energy-management studies have also addressed risk-aware equipment scheduling, heterogeneous hydrogen-storage planning, state-driven coordination of flexible loads, and AI-assisted optimization of thermal storage systems [
32,
33,
34,
35].
Multi-agent reinforcement learning is increasingly used for district energy control under centralized training with decentralized execution (CTDE). When the control objective is specifically the operation of building battery energy storage systems (BESS), CTDE still leaves unresolved issues. Critics can access global information and joint actions during training, but at deployment each building must act from local observations, which creates a training–execution mismatch and encourages policies that behave as if agents are conditionally independent. At the same time, learning an accurate centralized action-value function becomes difficult as the number of buildings grows because the joint state–action space expands rapidly; factorized critics reduce complexity but continue to rely on global signals that are not available at execution and remain sensitive to non-stationarity due to simultaneous updates across agents. Multi-head attention offers a principled way to learn which inter-building interactions matter by weighting messages in a query–key–value form, but attention-based DRL for energy systems is still typically demonstrated at small scale, is often tied to continuous policies and is rarely examined across different reward regimes (cooperative, competitive and mixed) under consistent benchmarking. In addition, most existing attention-based controllers retain multilayer percetron (MLP) function approximators, which can be restrictive for BESS control where nonlinearities are sharp (SoC saturation, power limits, efficiency losses) and strongly coupled with tariff discontinuities, PV intermittency and time-varying carbon intensity. This motivates a controller that targets effective BESS scheduling, retains CTDE scalability through selective attention and improves approximation quality through a more expressive function class.
This work is relevant not only for control performance but also for sustainability. By coordinating battery operation, household demand can be shifted away from periods when grid electricity is more carbon intensive, while excess on-site photovoltaic generation can be stored and used later. This approach can also reduce electricity costs under time-of-use tariffs, making residential electrification more affordable. At the district level, smoother demand profiles with lower peaks and less ramping can provide greater grid flexibility and help accommodate variable renewable generation more reliably. Overall, the study contributes to sustainability through reduced emissions, lower energy costs, and residential demand patterns that better support grid operation.
Contributions of Paper
The main contributions of the study are:
- 1.
We propose an attention-driven centralized critic within a centralized-training decentralized-execution (CTDE) framework (
Figure 1) to learn coordinated control policies for building battery energy storage systems (BESS). The critic evaluates each agent-specific soft action-value using a multi-head query–key–value mechanism over per-agent state–action embeddings, enabling selective cross-building credit assignment without explicit concatenation of the full joint state–action vector. This yields a value-learning structure whose per-head aggregation scales linearly with the number of agents, improving both stability and scalability of CTDE value estimation for district-level BESS dispatch.
- 2.
To improve approximation fidelity under nonlinear BESS dynamics and discontinuous objective signals, we replace the MLP components used for policy parametrization, state/state–action embedding and Q-value evaluation with Kolmogorov–Arnold Networks (KANs). Each KAN layer is expressed as a sum of learnable univariate spline functions (with linear skip terms) acting on individual input dimensions, which increases representational capacity for sharp regime changes induced by SoC saturation, charge/discharge limits, round-trip efficiency, PV intermittency and time-varying tariff and carbon-intensity profiles.
- 3.
Our analysis examines five core metrics, specifically cost, emissions, average daily peak demand, ramping and load factor, assessed across both individual building and district-wide scales. To ensure meaningful comparison, all outcomes are standardized relative to a scenario in which no storage is present and the findings are benchmarked against established centralized control approaches, such as a standard soft actor-critic (SAC), Tabular Q-Learning and Rule-Based Controller (RBC).
- 4.
We explicitly demonstrated the sustainability value of coordinated residential BESS scheduling in which the proposed controller reduces reduces carbon-intensive electricity imports along with improvement of the temporal utilization of local renewable generation and on top of that lowers household operating cost and enhances distrcit-level grid flexibilty through peak and rampin reduction.
The paper is organized in to following sections.
Section 2 described the overall methodology and this includes the introduction of residential dataset and simulation setup in CityLearn, the definition of key performance parameters or evaluation metrics, the observation and action design for BESS control, the reward function formulation and the details of benchmark controllers to be used in the simulation.
Section 3 presents the proposed AttentionKAN-based multi-agent actor–critic framework which details the attention critic, the KAN-based function approximation which is used in both actor and critic components.
Section 4 provides the results which are implemented in CityLearn and provides a detailed discuss of cost, emissions and grid related performance indicators across baselines and the proposed controller. In the end,
Section 5 provides the limitations along with future work and concludes the study.
2. Methodology
To validate the proposed algorithm in the study which will be introduced in the subsequent sections. Through the use of this framework, direct performance comparison can be made with the existing standards (SAC, Rule Based Controllers) in district energy management. This structured testing protocol allows us to comprehensively measure the effectiveness of the proposed solution in all aspects of adaptability and scalability to energy systems that are complex in nature. Each phase of this methodology is described below.
2.1. Selection of Dataset
Our evaluation uses a dataset consisting of seventeen single-family homes in Fontana, California. The range of building types included in the case study is illustrated in
Figure 2, which highlights the different exterior designs represented in the dataset. Differences in home layout and floor area contribute substantially to variation in electricity use, which is the kind of diversity our decentralized control approach is designed to manage. Data for these homes were collected as part of the California Solar Initiative, which examined how communities with widespread rooftop solar adoption and behind-the-meter battery systems influence distribution-grid operation [
48].
The specific parameters for battery capacity, power limits and operational efficiency were selected in such a wat that it mirrors the realistic physical constraints [
49,
50]. The 5 kW charge and discharge limit coupled with the 6.4 kWh capacity matches with the specifications of commercially available residential batteries. Furthermore, the 90% round-trip efficiency is also a widely accepted industry standard as well.The operational constraints of the battery control also require specific justification. The simulation allows the reinforcement learning agent unrestricted operation across the full depth of discharge. While physical lithium-ion batteries require strict upper and lower state of charge boundaries to prevent accelerated degradation, the 6.4 kWh parameter used in this study represents the strictly usable capacity rather than the gross hardware capacity [
51,
52].
The dataset comprises 17 prototype single-family residences, with conditioned floor space varying between 177 and 269 m
2. As detailed in
Figure 3, these structures were designed to minimize baseline energy consumption through the use of superior insulation, high-grade glazing and premium-efficiency appliances. The homes are fully electrified, utilizing electric systems for both space conditioning and water heating. Usage data is recorded by circuit-level meters at sub-minute intervals, while system operations are handled via a platform that supports both manual and automated inputs. Additionally, eight of the homes are equipped with 5 kW lithium-ion storage units. Based on a 75% depth of discharge and 90% round-trip efficiency, these batteries provide a nominal capacity of 6.4 kWh.
The community under study has been simulated in the CirtyLearn Environment and it is evaluated to test and check diverse control strategies for BESS. CityLearn provides the associated dataset as a privacy-preserving adaptation of the original as-built building data and it is developed to correct limitations in the data quality along with support to open-source dissemination. While preparing the adapted dataset, the raw time series data was first transformed into hourly energy quantities and expressed in kilowatt-hours. After transforming, the data irregularities were then addressed after incorporation of a systematic procedure in which an inter-quartile-range-based screen detected and identified atypical spikes, brief missing segments were reconstructed by employing linear interpolation. In addition to that, longer segments were imputed by employing a compact supervised learning model which is informed by neighboring observations. The buildings under study utilized a lithium-ion battery with the capacity of 6.4 kWh and it has discharge power of 5 kW while the round-trip efficiency is around 90%. This value is for a full charge and subsequent discharge. The study also assumes an unrestricted operation over a full depth of discharge. Furthermore, plug loads, heating and space cooling were considered and consolidated as single non-shiftable load demand category using the aggregated measurement captured at main service meter.
To support simulation, Los Angeles International Airport weather records were acquired along with an hourly carbon-intensity trajectory and it is measured in kilograms of
equivalent per kilowatt-hour. Electricity prices follow TOU-D-PRIME time-of-use structure, intended for households with behind-the-meter batteries and their tariff values summarized in
Table 2. As shown in the table, the lowest unit prices occur briefly before sunrise and again after midnight. From October through May and when regional electricity demand is comparatively lower the prices remain close to their average levels for much of the day. Weekend pricing is further moderated, thereby creating an incentive to shift or store energy during weekday peaks while still permitting greater evening consumption on Saturdays and Sundays.
2.2. CityLearn Environment
CityLearn in this study will serve as a reinforcement learning platform to investigate the effects of demand-response control under both centralized and decentralized settings. In centralized environment, there is a single agent while in decentralized environment, there are multiple agents per building. The CityLearn environment is compatible with the OpenAI Gym API and it adopts the established multi-agent interaction patterns described in the study [
53]. Through the usage of pre-generated building profiles together with internal component models, CityLearn avoids the computational overhead which is typically associated with co-simulation tools such as EnergyPlus and other buildings engines of same capability. The architecture for CityLearn is summarized in
Figure 4. As evident in the figure, the platform is symbolic of an execution layer for RL methods and thus it enables building energy simulation by drawing on datasets feed through several external sources integrated into CityLearn. These data streams parameterize the core environment which then mediates the exchange between the simulation and the control agents. The operation proceeds through typical RL cycle in which at each step the environment provides agents with the current state and along with that an associated reward and after that the agents copute control decision and the resultant actions are returned for implementation within the environment. After training the policy performance is assessed and reported back using load-shaping indicators that capture outcomes such as peak-demand mitigation and reductions in energy expenditure.
Within the environment of CityLearn, a set of controllable subsystems is represented through explicit device models and the notable of them are electric heaters and air-to-water heat pumps. In addition to that, the dataset provides hourly time series inputs for other end-use categories and it includes space cooling, plug and appliance loads. Control agents operate at an hourly resolution and the observations are updated along with updation of corresponding reward feedback with each time interval. Operational limits are embedded in the environment to ensure that actions remain feasible and that available device capacities are adequate to meet even minute loads. To safeguard occupant well-being, there is supervisory fallback controller present which prioritizes the thermal comfort by enforcing cooling and heating constraints regardless of the any energy-storage-driven objectives. Taken together, these mechanisms enable the investigation of sophisticated control strategies while maintaining assurance that comfort constraints are respected at every simulation step.
CityLearn is designed to support ongoing demand-response behavior without relying on direct and real-time dispatch commands from the electricity grid. Instead, it promotes temporal load shifting by managing controllable storage resources so that energy use is redistributed across time. To represent this process, the platform includes physics-informed and data-driven models for key components such as buildings, electric resistance heating, heat pumps, thermal energy storage (TES), PV generation and BESS. In this work, the environment was instantiated with energy models for two buildings; however, the same configuration approach can be applied to larger communities and the reported procedures remain reproducible for any number of buildings and we have extended the number of buildings to 7 as a test case. Each simulated building includes a PV array and a battery system consistent with the specifications of the assets installed in the corresponding real community.
2.3. Performance Evaluation Metrics
To assess the proposed multi-agent control framework in a rigorous way, we utilize five distinct performance indicators adapted from the standardized grid-interactive building evaluation framework proposed by Vazquez-Canteli et al. [
54]. These metrics, which mathematically quantify load shaping and economic performance, are categorized into grid stability indices calculated based on the aggregated district load, and building-level efficiency indices computed individually.
Let t denote the time step over the simulation horizon H and let represent the aggregated net load of the district at time t.
2.3.1. Grid Stability Indicators
The following three metrics evaluate the collective impact of the batteries on the distribution network:
Ramping (
): This metric quantifies the temporal volatility of the demand profile. It is calculated as the summation of the absolute differences in aggregated net load between consecutive time intervals, representing the effort required by the grid to match fluctuating demand:
Average Daily Peak (
): To measure peak-shaving performance, we compute the mean of the maximum daily loads. Let
D be the total number of days in the simulation and
be the set of time steps belonging to day
d:
Load Factor Deviation (
): The load factor represents the uniformity of energy usage. To frame this as a minimization problem (where lower is better), we calculate the complement of the load factor (
). This is computed on a monthly basis to account for seasonal variances. For month
m, let
be the average load and
be the peak load:
2.3.2. Economic and Environmental Indicators
The remaining two metrics focus on the operational efficiency of individual buildings. Let be the net load of building b at time t. Costs and emissions are attributed only when the building imports energy from the grid (i.e., ).
Operational Cost (
): The total electricity expenditure is derived by integrating the positive net load against the time-of-use tariff vector, denoted as
:
Carbon Footprint (
): The environmental impact is measured by the total mass of CO
2 equivalent emitted. This is calculated by multiplying the imported energy by the dynamic grid carbon intensity signal,
(kg CO
2e/kWh):
2.3.3. Metric Normalization
To facilitate a unified comparison across different controllers and seasons, all Key Performance Indicators are normalized against a baseline scenario where no battery storage is present. The normalized metric
is defined as:
where
represents the value achieved by the control algorithm and
represents the value derived from the raw building loads. Consequently, a value of
indicates a performance improvement over the no-storage reference case.
2.4. Observation and Action Space
A well-chosen definition of the observation and action space is paramount for learning of a high-quality control policy. The observations must capture enough information to characterize the current operating context while the actions must provide the agent with enough authority to influence system behavior. In our setup, each controller is supplied with (i) a detailed view of its own BESS, (ii) external signals that are reflective of grid conditions and the (iii) limited look-ahead information that supports anticipatory decisions. Accordingly, the observation vector combines community or district level inputs with building-specific states (
Table 3). Time information is represented through two temporal indicators included across all trajectories. Meteorological variables include direct solar irradiance together with forecast values drawn from the auxiliary weather records. Non-shiftable load corresponds to the building demand before accounting for PV production or operation of battery. Consumption of net electricity corresponds to addition of PV output, battery and non-shiftable load. In addition, carbon intensity and the effective electricity price quantify the environmental and economic implications of grid imports. For stable learning, cyclical encoding are applied to periodic variables and categorical quantities are represented via one-hot vectors and along with that continuous features are scaled using min-max normalization. At each hourly step, the agent outputs a scalar command
that specifies the requested charge or discharge level as a fraction of the battery capacity. Negative values correspond to discharging, whereas positive values correspond to charging.
2.5. Reward Function
The control objective is to jointly minimize the operating cost of electricity and carbon footprint associated with grid electricity. In parallel, we seek to improve gird-facing load-shaping behavior by suppressing peak demands and reducing steep hour-to-hour changes in the net load and promoting a higher load factor. A practical strategy to achieve these goals is to encourage charging during the periods when the tariffs are low which are typically after 21:00 and during the off-peak window preceding the afternoon peak as these intervals often align with comparatively lower grid carbon intensity. However, each building also has on-site PV generation which introduces an additional opportunity which is that during periods of strong solar production (roughly late morning through afternoon), the battery can be charged using locally generated energy rather than importing from the grid. The stored energy can then be dispatched later to offset demand during higher-cost and higher-emission periods, thereby reducing both cost and emissions while simultaneously relieving morning and evening peaks and improving peak-related and load-factor performance. The policy should also further learn to avoid curtailing renewable production by giving priority to battery charging whenever PV output is available and storage headroom remains. Conversely, when the building experiences net import condition and sufficient energy is still stored, the controller should discharge to reduce reliance on the grid. These behavioral targets motivate the reward design which is constructed to align the learning signal with economic cost, emissions impact and load-shaping priorities.
Considering the above mentioned learning objectives, the reward function can be formulated to align the agent’s learning updated with the desired control outcomes.
Based on the above equations, reward signal is constructed to drive down the electricity cost . It is evaluated separately for each building b and then aggregated across the full set of buildings so that the learning signal reflects community-level operation. The formulation promotes net-zero behavior through a state-dependent penalty factor which means that importing from the grid is discouraged when the battery still contains usable energy and exporting is also discouraged when the battery has remaining headroom (i.e., it is not yet fully charged). When export occurs with a battery at full charge, the shaping term contributes neutrally (no additional reward or penalty). In contrast, the most severe penalty arises when the battery is at its maximum state of charge while the building is simultaneously importing power from the grid.
As the reward design established in Equations (
7) and (
8) strongly affects and influences the final policy outcomes, the selection of weights, impact of shaping and the evaluation of alternative formulations require strong justifications.
The framework avoids complex and manually tuned scaling coefficients between and . Both of the terms are uniformly weighted and this design choice make sure that the primary objective of cost reduction act as a dominant gradient signal. This is due to the fact that shaping term is mathematically bounded by the sign function and it serves as a supplementary guide rather than an overwhelming scalar that could drown out the true economic objective. On the other hand, while the reward shaping inherently directs agent behaviour, is explicitly constructed to prevent adverse policy bias. Rather than forcing the agent to imitate a predefined schedule, an approach that would defeat the purpose of using reinforcement learning, strictly penalizes operationally illogical states. Furthermore, alternative evaluations were also tested to validate the approach. First, a purely economic sparse reward was designed which omit . This formulation slowed down the convergence and cause frequent battery underutilziation because the agent struggled to learn the delayed correlation between charging during the morning and saving money during the evening peak. Second, a heavily parameterized multi-objective reward was tested which combined cost, emissions and peak demand as separately weighted linear terms. This approach was proved fragile as the relative weights require continuous manual retuning whenever building characteristic differ. So, the chosen formulation was ultimately selected because it provided the most reliable and stable convergence across multitude of buildings without requiring manual weight adjustments.
2.6. Benchmarking
For benchmarking, the proposed multi-agent controller is evaluated against the following reference approaches:
Rule-based control (RBC), where battery charging and discharging are governed by predefined heuristic rules and threshold settings [
55].
Adaptive tabular Q-learning, which requires a discrete state-action formulation; therefore, the original observation and action spaces are discretized before training and evaluation [
56].
Soft Actor-Critic (SAC), in which a single agent jointly manages the operation of two buildings [
57].
The detailed methodological framework for the proposed study is given in
Figure 5. It sums up all the detailed steps of the study including the proposed architecture, control problem and KPIs along with benchmarking controller implementation against which our proposed controller will be evaluated.
2.7. Validation Framework
Due to the safety and logistical constraints of deploying untested reinforcement learning algorithms directly onto physical residential battery systems, this study employs a rigorous simulation-based validation framework. In line with established protocols for DRL in energy systems, the validation of the proposed AttentionKAN-based multi-agent controller is structured across three dimensions:
Empirical Data and Environmental Fidelity: The validation environment (CityLearn) is driven by a full year of real-world operational data from the California Solar Initiative. By utilizing actual 8760-h profiles for non-shiftable load, PV generation, and local weather, the controller is forced to operate under realistic seasonal variations, weather anomalies, and distinct behavioral differences between buildings. This ensures the learned policy is not overfitted to idealized mathematical models but is validated against real-world stochasticity.
Comparative and Ablative Benchmarking: The methodology is validated through direct comparison against established control paradigms. The implementation of a standard multi-agent Soft Actor-Critic (SAC) baseline serves as an implicit ablation study. By comparing the standard SAC with the proposed AttentionKAN-SAC, we isolate and validate the performance contributions of the KAN function approximators and the multi-head attention mechanism. Furthermore, the Rule-Based Controller (RBC) serves as an industry-standard validation threshold, ensuring the DRL agent outperforms highly tuned, deterministic logic.
Algorithmic and Statistical Robustness: DRL algorithms are inherently stochastic. To validate that the performance of the AttentionKAN controller is stable and not the result of a favorable random initialization, the hyperparameter tuning and final policy evaluations were conducted across multiple random training seeds. The aggregated key performance indicators (KPIs) reflect the consistent convergence of the algorithm.
4. Results and Discussion
This section provides that experiemental results acquired by applying the AttentionKAN based multi agent controller. We start by analyzing the operational data of the selected buildings (solar generation, carbon-intensity signals, non-shitable load etc). We subsequently report the no-control baseline, which is used as the anchor for all later results. We then assess a set of strong benchmark controllers, including RBC, SAC and Tabular Q-learning, in order to place our method within the broader landscape of existing solutions.
4.1. Data Selection, Preprocessing and Visualizations
In our experimental setup, we randomly selected Buildings 2 and 7 and applied the multi-agent SAC controller to both. The computational cost of training the agents for one year is quite hight and thus the evaluation is limited to two buildings. The goal of this setup is to demonstrate the proof of concept and if enough computational resources are present, then the methodology can be expanded to larger number of buildings.
Load and solar irradiance with hourly measurements for building 2 is shown in
Figure 7. There is clear evidence of non-shiftable load which fluctuates daily and with the reasonably constant baseline and is mostly in the 1–4 kWH range with occasional surges that go to around 6–7 kWH. For the PV profile, typical diurnal cycle and seasonal fluctuations are evident. The PV profile increases throughout the day and reaches its peak around midyear and then it gradually decreases after late summer.
The carbon-intensity profile is shown in
Figure 8 and the graphs depicts that it varies around 0.09 to 0.25
per kWh over the course of a year. Several multi-day stretches with reduced intensity emerge in the mid-to-late part of the year, while higher-intensity regimes are more prevalent early on and again near the end of the timeline. An important finding is that there is a weak relation between PV output and carbon intensity. Carbon intensity is only slightly below average even during peak solar production and it indicates that the most efficient way for carbon-aware control is through usage of intensity signal rather than relying directly on solar output.
Building 7 (
Figure 9) has a more pronounced and more consistent non-shiftable load when compared with Building 2 and it usually ranges from 3.3 to 4.9 kWh per hour and there are occasional excursions exceeding 5.8 kWh. The PV production shows a typical seasonal pattern with a wide plateau in the summer and some dips in the earlier months, which might be due to inverter limitations or temporary weather changes.
Figure 10 shows the carbon intensity for Building 7 which is almost similar to Building 2 and it varies within the same range as Building 2 and shows a general trend of decreasing intensity around mid-year. The two buildings’ load and PV profiles are different, thus controllers working in both environments should have different decision thresholds based on the seasons. They should also have some spare capacity that may be utilized strategically on nights when carbon intensity is high.
The hourly meteorological data for a whole year is summarized in
Figure 11 which depicts circumstances characteristic of Southern California’s coast. The outdoor dry-bulb temperature is around 15 to 25 degree Celsius and there is a visible cooling phase in the midst of the series when temperatures level off at around 8 to 10 degree Celsius and then there is a steady rise again at the end of the series. While relative humidity is quite high (from 60 to 90%) throughout the year but one exception is that it drops significantly during the dry season.
Because Buildings 2 and 7 experience the same meteorological inputs, the observed differences in demand and PV generation are more reasonably explained by building-specific properties and usage patterns than by weather. But nonetheless, these weather trends are still crucial fo optimal designing of controller because in summer there is quite strong direct solar irradiance and it creates extended midday opportunities for charging of the battery whereas in winter, operation relies more on diffuse radiation and careful use of stored energy. In addition to that, mild winter climate indicates that the comfort constraints can generally be maintained without deep battery depletion although there are some short periods of high humidity which may increase latent cooling requirements and thus it can lead to temporary adjustments in cooling setpoints.
4.2. Baseline Controller Implementation
The first step for checking the performance of the proposed controller is to establish the baseline performance metrics within the CityLearn environment. In the baseline configuration, the agent will operate without managing the BESS and the battery charge and discharge remains inactive or disabled. This is a crucial and necessary step as it will provide a standardized reference point and other control algorithms will be measured against this reference.
Figure 12 shows that the normalized cost and emission KPIs are precisely at 1.0 which means that battery control is inactive in this configuration.
In order to better understand the temporal dynamics,
Figure 13 zoom in to a single week of electrical demand for the two buildings. The plot is restricted to seven days so that visual trend can be seen in a clear and more comprehensible manner. Through out the timeseries, Building 2 usage stays mostly the same while conversely, Building 7 shows a noticeable dip in power consumption on the final recorded day. Such variations are a result of unexpected independent factors such as changing occupant schedules, different thermal comfort preferences and family size.
Figure 14 provides averaged 24-h load curves that further demonstrate this building-to-building behavioral divergence. Building 2’s trend reveals that grid demand almost flat-lined overnight. By mid-afternoon, the net load begins to decline as the building’s internal consumption is momentarily eclipsed by localized rooftop solar power. After 4:00 PM, solar output begins to decline, causing the grid dependency to increase significantly. By 10:00 PM or 11:00 PM, the building reaches its maximum capacity. Building 7 has a completely different trend. Its electrical consumption increases continuously between 10:00 AM and noon, after relatively little activity throughout the night. From 4:00 PM to 10:00 PM, the home’s power use shows a steady, linear rise before gradually falling down later in the day.
In the baseline simulation, SoC of all battery units will naturally be zero as the BESS interventions are purposefully disregarded.
Figure 15 confirms that the grid dynamics are unaffected by the battery architecture during baseline control implementation, proving that the battery is permanently depleted.
The district-level kPIs shown in
Figure 16 remain unchanged at the value of 1.0 when we shift our focus from individual building data to an aggregated neighborhood view. Finally, this statistic only confirms that the whole community network is a good representation of the original, uncontrolled electrical baseline before optimization tactics were used.
4.3. Benchmark Controllers Implementation
After establishing the no-control baseline, we implemented the RBC, TQL and SAC controller. All evaluated benchmark controllers operate in a centralized manner, meaning a single shared control policy is applied uniformly to the batteries of both buildings.
Figure 17 reports the building-level cost and emissions KPIs. Among the implemented controllers, the RBC achieves the strongest performance and, notably, provides a highly competitive reference point for data-driven approaches. This outcome is largely explained by the alignment between the RBC’s simple heuristic and the pronounced daily structure of solar availability and time-varying electricity tariffs.
The full simulation load trajectories and their average daily profiles are shown in
Figure 18 and
Figure 19. These plots suggest that, under the RBC, net grid consumption is moderated during periods of strong solar availability, consistent with systematic charging behavior around midday. Moreover, grid draw is reduced during the evening peak window, which aligns with battery discharge being used to mitigate peaks and relieve grid stress. Overall, the RBC exhibits effective peak shaving and valley filling—two core objectives of demand response. A key limitation of rule-based strategies, however, is their lack of adaptability: because each building has unique operating characteristics and usage patterns, designing and tuning hand-crafted rules to achieve near-optimal charge–discharge schedules becomes cumbersome as the number of buildings increases.
Figure 20 illustrates the state-of-charge trajectories, providing additional insight into charging and discharging behavior. The reason for TQL’s poor performance is apparent in this visualization: the charging and discharging decisions appear weakly structured and do not converge to a consistent operational pattern. By comparison, the RBC exhibits a clear and repeatable daily cycle—charging earlier in the day and discharging later—which explains its stronger performance across the evaluated metrics.
District-level KPIs are summarized in
Figure 21, where the RBC again delivers the best overall performance across the reported indicators. The objective is to achieve improvements relative to the baseline by reducing all KPI values; in this regard, RBC and SAC succeed on multiple measures, whereas TQL remains consistently weaker.
4.4. Implementation of AttentionKAN-MADRL
This section is dedicated to deployment of AttentionKAN-MADRL controller and the hyperparamters associated with the implementation to obtain the strong performance. The implementation of the controller was carried out by employing RLlib which is a widely used framework for multi-agent reinforcement learning. In our formulation, AttentionKAN-MADRL is used as the control backbone for battery operation in which one learning agent assigned to each building so that charging and discharging decisions can be optimized at the building level. The controller is a model-free and off-policy RL learning algorithm and its off-policy nature allows the controller to reuse previously collected transitions through a replay buffer which can improve sample efficiency by reducing the amount of new environment interaction required to learn effective behavior. The approach couples an actor and critic trained via off-policy updated and it includes and entropy-regulairized objective that promotes exploration and oten improves training stability. Concretely, the proposed controller maintains and updates three components which are (i) a stochastic policy representing the actor, a soft state-action value function for evaluating actions and a state value function that serves as additional baseline in the learning process.
The controller is trained under centralized training and decentralized execution, with one SAC agent per building. Since replacing MLP blocks with KAN modules increases function expressivity through spline-parameterized univariate edge functions, the primary tuning objective is to obtain stable learning dynamics while preserving the ability to adapt to heterogeneous building behavior. We therefore separate hyperparameters into three groups as shown in
Table 4: (i) SAC optimization and replay settings, which govern sample efficiency and stability; (ii) attention critic capacity, which controls how strongly cross-agent interactions are modeled; and (iii) KAN-specific capacity and regularization, which determines the smoothness and complexity of learned mappings.
For SAC optimization, we tune actor and critic learning rates, entropy temperature control, discounting and target-network mixing. In practice, KAN parameterization generally exhibit sharper gradient responses as compared to MLPs and thus search emphasizes conservative learning rates along with explicit gradient clipping and sufficient replay warm-up before updates. Entropy temperature is considered as learnable parameter with the target entropy defined as function of the action dimensionality. Replay and update cadence are tuned through the replay buffer size along with number of warm-up steps and training batch size. These settings are particularly important for yearly-long hourly simulation because temporal correlation can otherwise lead to unstable value estimation.
For the attention critic, we tune the number of heads and the key and value dimensions per head along with the embedding widths for the per-agent-state-action encoders and the final critic evaluator. The search is designed to avoid over-parameterization in small districts while maintaining scalability for larger districts. When the number of buildings is small as in our case fewer heads and smaller embeddings dimensions typically suffice the purpose but as the number of building increases more heads can improve the selectivity of cross-agent credit assignment.
For the KAN modules, the search explicitly controls spline degree along with grid resolution and regularization on spline coefficients. We enable the linear skip term in all KAN layers so as to improve conditioning and to allow the network to represent near-linear control laws without the requirement of dense spline activations. Regularization includes and penalty on spline coefficients to encourage sparse functional structure along with an optional smoothness penalty to prevent oscillatory edge functions. All scalar inputs to KAN layers are normalized to a bounded range so that spline bases operate on a consistent domain across features and buildings. The method was rigorously validated within the CityLearn environment using a full year of real-world empirical data. The validation results show that the proposed controller consistently improves both building-level and district-level outcomes when compared against established industry and academic baselines.
The results of the proposed controller (
Figure 22) summarizes the cost and emissions outcomes. The proposed controller setup in green bar yields the strongest overall performance and delivered an approximately 50% reduction in cost relative to reference configuration and there is also a marked decrease in
emissions. The corresponding load trajectories in
Figure 23 and
Figure 24 further indicate a smoother net demand profile and this is desirable behavior from both operational and grid-integration perspectives. In this regime, surplus daytime solar generation is preferentially stored and subsequently released during evening peak periods. This cycle of charge-discharge scheduling mitigates peak demand along with alleviation of grid stress and it supports both demand response objectives by shifting consumption towards more favorable hours.
Figure 25 depicts the SoC trajectories for both buildings and shows that learned behavior closely resembles the charging-discharging cadence encouraged by the RBC but this time with more clearer timing and improved effectiveness. The agents converge to a stable and repeatable strategy in which they raise the batteries to near-full capacity during midday hours which coincide with the highest PV output and this indicates that the controller reliably captures available solar surplus and stores it for later use. he SoC curves then decrease sharply during the evening window with the most pronounced discharge occurring between approximately 16:00 and 21:00. This period aligns with the elevated demand on the grid and higher electricity prices and this demonstrates that store energy is released when it is most valuable from both economic and operational perspectives. One important observation is that these outcomes exceed RBC benchmark while avoiding the need to hand-design charging and discharging rules.
The district-level evaluation in
Figure 26 confirms that these improvement extend beyond individual buildings and substantial gains are observed for emissions, cost, daily peak demand and ramping. Overall, the results indicate that the proposed controller outperforms simple SAC configuration. Replacing MLP blocks with KAN layers improves the critic and policy’s ability to model sharp, nonlinear relationships among load, PV, prices and carbon intensity, which translates into more accurate value estimates and more reliable control decisions. This results to the controller schedules charging and discharging more effectively across varying conditions and improving costs and emissions while producing smoother net-load behavior than conventional SAC, tabular and fixed rule RBC.
4.5. Comparison with Other SoTA MARL Architectures
To directly evaluate the scalability of the proposed approach, we introduce five new buildings into the environment for a total of seven buildings. The setup of this experiment is thus seven buildings forming a cluster. In this experiment, we compare our AttentionKAN controller with two SoTA multiagent reinforcement learning baseline algorithms, namely Multi-Agent Deep Deterministic Policy Gradient (MAPPO) and MADDPG. We use KPIs at the district level for the purpose of comparison. The existing agents’ multi-head attention mechanism is assessed for its correctness in credit assignment as well as for learning in the enlarged joint state-action space without any performance deterioration. This indicates that our controller remains the best against popular multi-agent benchmark algorithms. The results are depicted in
Figure 27.
Under the economic and environmental metric, MADDPG can’t find a stable coordinated policy as the state-action space increases and has the largest relative cost and carbon intensity in comparison to the learning-based controllers. The MAPPO showed improved performance compared to MADDPG’s performance by taking advantage of the stronger convergence properties of on-policy algorithms. However, the savings achieved are still below the proposed approach. AttentionKAN results in the lowest district-level electricity price and -intensity. From these experiments, we can see that as the scale gets even larger, the performance does not degrade, which shows that the KAN-based function approximators retain their strong representation power and ability to capture the sharp non-linearities of the ToU tariffs and carbon intensity signals compared to standard MLPs even as the environment gets more complex.
Grid stability metrics such as daily average peak, ramping, and load factor further illustrate the advantage of using attention mechanisms for multi-agent systems. For a cluster of 7 buildings, if multiple agents simultaneously discharge their batteries without coordination, secondary demand peaks (rebound effects) can be generated easily. The results indicate that, while both MAPPO and MADDPG have exhibited higher ramping values, meaning that their actors have occasionally synchronized their battery efforts resulting in a more extreme variation of total net load and also had higher values of average peak.
4.6. Ablation Studies
To justify the architectural choice of replacing KANs with MLPs, an ablation study was conducted. The objective of this ablation is to isolate the the specific performance contribution of KAN modules. TO ensure a fair comparison the baseline configuration was built using the exact same framework, attention mechanism and reward formulation of the proposed model. The only difference was the actor and critic networks utilized standard fully connected MLP layers with ReLU activations.
The results of the ablation study was shown in
Table 5 where it can be seen that for all the KPIs related to district level, KAN variant outperforms MLP variant. Furthermore, it can be seen that computational time is also reduced in the case of KAN variant which justifies the replacement of MLP with KANs. This direct comparison confirms that replacing MLPs with KANs provides a structural and mathematical advantage for energy storage scheduling.
4.7. Discussion
The results show that the AttentionKAN controller learns an effective BESS dispatch strategy and consistently improves performance at both the building and district levels compared with the learning-based benchmarks. The largest gains appear in the cost and emissions KPIs, indicating that the controller reduces grid consumption during high-impact periods and uses storage to shift demand across time. This is consistent with the KPI summaries and the net-load trajectories, which exhibit smoother demand and reduced exposure to peaks. The learned behavior goes beyond a direct imitation of rule-based scheduling: charging is more tightly synchronized with periods of high renewable availability, while discharging is more consistently concentrated in hours associated with higher grid stress.
The benchmarking results also highlight that simple controllers can be strong when the environment is dominated by regular diurnal patterns. The rule-based controller benefits from predictable PV production and time-of-use tariffs and it achieves competitive peak shaving and valley filling without training. Its limitations are practical rather than conceptual: the rule set must be engineered and re-tuned as building characteristics change and it cannot respond systematically to misalignment between PV availability, electricity prices and carbon intensity. As the number of buildings increases and heterogeneity becomes more pronounced—different occupancy profiles, different PV surplus regimes and different storage operating ranges—the effort required to maintain an effective set of rules becomes substantial and performance can degrade if the rule logic is not adapted. These observations support the need for learning-based approaches that can generalize across diverse buildings without manual redesign.
Within a decentralized execution setting, the proposed approach improves upon a standard SAC configuration by addressing two common failure modes in multi-agent energy control: unstable credit assignment under partial observability and approximation error in the value function. The attention critic directly targets the credit-assignment issue by learning which cross-building information is relevant when evaluating each agent’s actions. Instead of treating all other agents as equally informative through full concatenation, the query–key–value mechanism constructs an agent-specific context representation that emphasizes influential interactions and downweights redundant signals. This is particularly important when rewards include district-level terms, because agents must learn how local BESS actions affect shared objectives. By providing a more selective and structured training signal, attention can reduce variance in value estimates and improve the quality of policy gradients, even in small districts where coordination demands are intermittent.
A second factor is the use of KAN modules in place of MLPs within the actor and critic mappings. BESS control is inherently nonlinear and often piecewise: actions saturate at power limits, SoC dynamics impose asymmetric feasibility constraints and the optimal policy can change abruptly around tariff boundaries or when PV surplus transitions between availability and scarcity. While MLPs can represent these effects, doing so reliably often requires increased width, careful regularization and substantial data to avoid value overestimation and unstable learning. KAN layers provide a more structured function class through spline-parameterized univariate components combined with linear skip paths, which can represent steep nonlinear transitions more directly. This is most relevant in operating regimes where small changes in state lead to qualitatively different actions—for example near SoC saturation (where additional charging yields little benefit) or near depletion (where stored energy becomes scarce and highly valuable). The observed reductions in peaks and ramping are consistent with improved value accuracy and, consequently, better dispatch timing.
The training dynamics also reflect these architectural choices. Spline-based parametrizations can yield sharper gradients than standard MLPs, particularly when grid resolution is high or regularization is insufficient. For this reason, the selected hyperparameter strategy—more conservative learning rates, explicit gradient clipping, adequate replay warm-up and coefficient regularization—serves as a stability mechanism rather than a tuning preference. Attention capacity interacts with this stability: overly large embeddings or excessive heads can lead to overfitting in small districts, whereas insufficient capacity can collapse the attention module into near-uniform weighting and remove any coordination benefit. The final configuration suggests a workable balance in which attention improves inter-agent evaluation without destabilizing critic learning.
Several limitations should be considered when interpreting the results. First, the evaluation is conducted on a limited number of buildings due to the computational cost of year-long training, so scalability should be verified on larger districts where correlated behavior can produce more complex peak interactions but we have somewhat tried to mitigate this issue and also provide the analysis for 7 buildings as well. Second, the observation design includes look-ahead signals for tariffs and irradiance, which provides ideal short-horizon information; operational deployments would rely on forecasts and therefore introduce uncertainty. Third, the current formulation focuses on BESS dispatch and does not explicitly account for degradation or cycling costs, which may be important when frequent cycling trades operational savings against asset lifetime. Finally, the discrete sampling policy mapped to continuous setpoints introduces a discretization design choice; the granularity can influence both learning and control quality, especially under heterogeneous actuator capabilities.
These limitations point to several concrete extensions. Scaling experiments with more buildings would clarify how attention selectivity and KAN capacity should grow with agent count. Replacing look-ahead inputs with realistic forecasts and introducing forecast-noise augmentation would allow robustness to uncertainty to be quantified. Incorporating battery degradation terms would support a multi-objective formulation that jointly optimizes cost, emissions and wear. Finally, targeted ablations that remove attention and/or replace KAN with MLPs would strengthen causal interpretation of the observed gains and identify which component contributes most under different reward designs and district compositions.
The practical value of the proposed controller should be interpreted in sustainability terms rather than only in algorithmic terms. First, the reduction in CO2 emissions indicates improved environmental performance through lower dependence on carbon-intensive grid electricity. Second, the reduction in electricity cost supports the economic sustainability and affordability of residential energy management. Third, the improvements in peak demand, ramping, and load shape indicate stronger grid-supportive flexibility, which is important for districts with growing shares of distributed renewable energy. In addition, the learned strategy stores daytime solar surplus and discharges during higher-impact evening periods, which improves the temporal matching between local renewable availability and demand. Therefore, the proposed method contributes to environmentally and economically sustainable building operation while also supporting a more stable and renewable-ready electricity system.
5. Conclusions
This work presented an AttentionKAN-based multi-agent reinforcement learning controller for coordinated BESS operation in multi-building residential demand response. The design proposed in this work aggregates centralized training with the execution that is decentralized in nature where each building is governed by its own actor whereas a centralized critic uses multi-head attention to learn which cross-building interactions are most relevant for value estimation and credit assignment. Through replacement of MLP blocks with KAN function approximators in both actor ad critic pathways, the controller is able to better represents the strongly nonlinear and piecewise structure of battery scheduling under practical constraints, photovoltaic intermittency, time-of-use pricing and time-varying carbon intensity. The experimental results are performed in CityLearn and it show that the proposed controller consistently improves both building-level and district-level outcomes when compared with the implemented baselines. In particular, it learns a stable and interpretable operating pattern that stores daytime solar surplus and dispatches energy during the late-afternoon and evening peak window, which reduces high-impact grid imports. The outcome translates in to quite susbtanital gains both in terms of economic and enviromental aspects and it includes an approximate 50% reduction in cost relative to the reference configuration and a clear reduction in carbon emissions. All of this is achieved while simultaneously smoothing out aggregatged demand and improving peak-related and raming indicators. The findings show that the proposed controller is not only an effective reinforcement-learning architecture for residential demand response, but also a sustainability-oriented energy-management strategy. By reducing electricity cost and carbon emissions while improving peak-related and ramping indicators, the method supports more affordable, low-carbon, and grid-compatible operation of residential districts with distributed PV and battery storage. In this sense, the study contributes to sustainable building electrification and renewable-energy integration, while also identifying the remaining challenges of scalability, forecast uncertainty, and battery lifetime effects.
There are some limitation as well that demand future work. The evaluation is carried out in smaller subset of buildings due to the computational cost of year-long training and the scalability should be validated on larger districts where coordination structure is richer. In addition to that, incorporating battery degradation terms and performing systematic ablations (attention-only, KAN-only and full AttentionKAN) would further clarify the contribution of each architectural component and support deployment-oriented trade-off analysis.