1. Introduction
The swift progression and extensive utilization of artificial intelligence (AI) in various industries have markedly heightened the demand for large-scale computer infrastructure [
1,
2]. Modern AI-driven systems necessitate significant computational power for data processing, model training, and real-time decision-making across cloud data center, hybrid, and edge contexts. Although these capabilities have facilitated transformative innovations in sectors like manufacturing, healthcare, finance, and logistics, they have concurrently presented significant challenges concerning energy consumption and environmental sustainability [
3,
4,
5]. Data centers are widely acknowledged as significant contributors to global electricity consumption and carbon emissions, prompting urgent worries on the long-term ecological consequences of AI technology [
6].
In reaction to these difficulties, the notion of Green AI has arisen as a transformative approach that prioritizes the creation and implementation of energy-efficient and ecologically friendly computing systems [
7]. In contrast to conventional AI methodologies that emphasize accuracy and performance regardless of computational expense, Green AI promotes the optimization of models and infrastructures to minimize energy usage, carbon emissions, and operating expenditures. This transition is especially significant in industrial systems, where extensive, ongoing workloads exacerbate the environmental consequences of inefficient resource use. Achieving sustainability in these environments is intricate due to fluctuating workload patterns, diverse infrastructure, and stringent SLA standards [
8,
9].
Current methodologies for energy-efficient computing have investigated strategies like virtual machine consolidation, dynamic resource allocation, and hardware-level improvements [
10]. Recent research in carbon-aware computing has proposed scheduling workloads according to the carbon intensity of energy sources, allowing for systems to diminish emissions by utilizing cleaner energy availability over time and location [
11]. Although these methods demonstrate potential, they are frequently executed in isolation, lacking a cohesive framework that concurrently tackles energy efficiency, carbon consciousness, and system performance. Moreover, numerous current solutions fail to fully leverage the capabilities of intelligent learning-based methodologies for adaptation to dynamic industrial contexts [
12,
13,
14].
The rapid growth of AI, cloud computing, and digital transformation in industry has led to increased energy usage and carbon emissions in computing infrastructures. As the use of large scale, AI-enabled applications continue to increase and become dependent upon data centers and distributed cloud environments for support, sustainability is now recognized as a key issue in computing research. Most existing methods for workload scheduling and resource allocation have mainly focused on improving performance while neglecting environmental factors, including energy use, greenhouse gas emissions, and carbon intensity. Thus, there is an increasing demand for intelligent and sustainable computing frameworks that balance operational performance with environmental sustainability. This research addresses these issues by presenting a new “Green AI” framework called the HC-DQNCAPS that combines energy-efficient hardware architectures with carbon-aware workload scheduling to facilitate sustainable industrial cloud computing.
The contributions of this research are as follows:
We introduce a novel HC-DQNCAPS framework to combine HAC with the DQN reinforcement learning algorithm to achieve carbon-aware and energy-efficient workload scheduling.
We constructed a multi-objective optimization model to minimize energy consumption, carbon emissions, and SLA violations while maximizing resource utilization for industrial cloud environments.
Through experimental evaluation and statistical validation, we showed that our proposed HC-DQNCAPS framework performed better than existing methods such as FCFS, Energy-Aware VM, Carbon-Unaware RL, PPO, DDQN, and MADRL.
The remainder of this document is structured as follows:
Section 2 provides a detailed review of the existing literature on concepts related to Green AI, sustainable cloud computing, efficient use of resources, carbon-aware scheduling, and reinforcement learning for solving problems;
Section 3 presents the proposed HC-DQNCAPS framework, including the system architecture, mathematical formulation, and optimization constraints;
Section 4 presents the experimental design, implementation details, simulations, evaluation criteria, and statistical methods used to validate results;
Section 5 presents and compares the results against baselines and current state of the art; and
Section 6 presents conclusions from a summary of findings along with contributions, limitations, and ideas for future research directions within the sustainable industrial computing based on AI.
2. Related Literature
The growing environmental consequences of extensive AI systems have given rise to Green AI, a research framework that prioritizes computational efficiency in conjunction with model performance [
15,
16]. Initial AI research predominantly concentrated on optimizing accuracy, frequently sacrificing substantial computational and energy resources [
17,
18,
19]. Recent studies have indicated that the training and deployment of complicated models, especially deep learning architectures, can lead to significant carbon emissions. Consequently, researchers have initiated investigations into methodologies including model compression, pruning, quantization, and optimized neural architecture design to diminish the energy consumption of AI systems while preserving satisfactory performance standards [
20,
21]. These methodologies establish the basis for energy-efficient AI; nevertheless, they frequently focus just on model-level optimizations, neglecting comprehensive system-wide deployment techniques [
18,
19,
22].
The domain of energy-efficient computing inside cloud and data center infrastructures has garnered significant interest [
23]. Methods like dynamic voltage and frequency scaling, virtual machine (VM) consolidation, and workload migration have been extensively researched to diminish energy consumption in distributed systems [
24,
25,
26]. VM consolidation, for example, reduces the quantity of operational physical servers by optimizing workload distribution, thus diminishing idle power consumption [
27]. Likewise, adaptive resource provisioning enables systems to adjust resources according to demand, hence enhancing overall efficiency [
28]. Although these strategies significantly diminish energy usage, they generally emphasize infrastructure-level optimization and do not explicitly account for the carbon intensity of energy sources, which fluctuates by area and time [
29].
Recently, carbon-aware computing has emerged as a promising strategy to promote sustainability in computing systems. This research presents the notion of scheduling workloads according to the carbon intensity of power, allowing for systems to relocate computational jobs to areas or times when cleaner energy is accessible [
30]. Research has shown that carbon-aware scheduling can substantially decrease emissions without necessitating extensive alterations to hardware infrastructure. Geographically distributed data centers can optimize renewable energy utilization by dynamically reallocating workloads [
31]. Notwithstanding these developments, numerous current solutions regard carbon awareness as a supplementary feature instead of embedding it thoroughly within resource management and decision-making frameworks [
7].
Recent years have demonstrated significant potential for the application of machine learning and reinforcement learning in resource optimization [
32]. Reinforcement learning methodologies, especially DQN, have been employed to tackle dynamic resource allocation challenges in cloud environments. These strategies facilitate systems in acquiring optimal policies via interaction with the environment, hence adjusting to fluctuating workloads and system states [
33,
34]. Furthermore, Hierarchical Agglomerative Clustering approaches have been utilized to categorize analogous workloads, hence diminishing complexity and enhancing scalability in resource management [
35]. Nevertheless, the majority of reinforcement learning-based methodologies prioritize performance and energy efficiency, frequently overlooking carbon-aware targets, and seldom integrate clustering with reinforcement learning within a cohesive framework [
36,
37].
Although there has been a lot of advancement in these different areas, there is still a large gap when it comes to developing an integrated model for industrial systems that combines energy efficiency and carbon awareness with intelligent decision-making. The current methods for addressing these issues may address each of them individually (and therefore are not very effective at all), but they do not help to create greener computing environments for the overall sustainability of industrial computing. This paper builds on and expands the research conducted in previous research papers to propose a hybrid solution that merges the components of energy efficient architectures with carbon-aware scheduling and reinforcement learning-based optimizations with clustering algorithms. The purpose of this hybrid model is to provide a more integrated and scalable solution for Green AI and sustainable computing in industrial settings.
3. Proposed Model
3.1. Overview
The HC-DQNCAPS (Hierarchical Clustering Deep Q-Network Carbon-Aware Placement System) framework is being evaluated as a potential industrial computing resource management solution based on its ability to create an intelligent, adaptive, and sustainable physical resource management system. This proposed design, as shown in
Figure 1, integrates both carbon-aware scheduling of workloads, as well as deep reinforcement learning-based optimizations and energy efficient computing architectures aimed at reducing energy consumption, carbon emissions, and achieving Service Level Agreement (SLA) compliance within distributed cloud infrastructure.
The HC-DQNCAPS is composed of three tightly coupled layers that provide a cohesive framework for the placement of workloads within geographically distributed cloud data centers, hybrid cloud environments, and edge computing platforms to support distributed computing. The Energy-Efficient Infrastructure Layer utilizes adaptive resource provisioning, dynamic scaling of resources, virtual machine (VM) consolidation and intelligent workload balancing to reduce wasteful energy use and improve the efficiency of physical computing resources. By eliminating the idle state of servers and maximizing the effective use of physical resources, the HC-DQNCAPS will reduce the operational power use of physical computing equipment, and therefore reduce the cost of cooling those resources. In contrast to traditional static resource allocation systems, the HC-DQNCAPS will continuously monitor the infrastructure utilization and dynamically allocate physical computing resources to respond to actual workload demand.
The Carbon-Aware Scheduling Layer is a mechanism that uses both real-time and predictive data about carbon intensity in order to make decisions on workload placement. By utilizing a different approach than scheduling workloads based merely on the availability of resources or performance requirements, this approach prioritizes the execution of workloads in geographic areas and timeframes associated with lower carbon intensity for energy generation. This environmentally focused mechanism for scheduling contributes to significant reductions in greenhouse gas emissions while preserving operational efficiencies. Additionally, this framework incorporates a short-term carbon forecast that supports proactive deferring and migrating workload decisions.
The Intelligent Decision-Making Layer combines Hierarchical Agglomerative Clustering (HAC) technology with Deep Q-Network (DQN) reinforcement learning capitalizing upon the relative strengths of each technology to create highly efficient and scalable intelligent systems. HAC is a clustering methodology that groups workloads that exhibit similar resource utilization characteristics such as CPU utilization, memory demand, network bandwidth consumption, or I/O activity; this aids in decreasing complexity of decision-making and increasing scalability. DQNs utilize their continuous feedback loops in the form of reward and re-punishment to learn the best possible policies for allocating workloads to different areas of an emperor. A DQN’s performance as an agent continually learns optimal workload allocation policies through interaction with its environment under reward-based criteria and therefore has dynamic capabilities for adapting to differing workload patterns, varying carbon intensities, and various SLA conditions, enabling it to make timely decisions.
The HC-DQNCAPS framework is intended to accommodate a diverse range of industrial computing infrastructures and has been designed for large scale deployments. The HC-DQNCAPS framework will feature a modular architectural design, allowing for it to be integrated with cloud or edge computing or hybrid systems, while also being able to achieve computing objectives from a sustainability perspective.
3.2. System Architecture
The design framework of a HAC-DQNCAPS is based on a layered intelligent computing architecture that consists of four distinct operational layers: Data Collection Layer, Processing Layer, Decision Layer, and Execution Layer. This design framework includes the capabilities for real-time monitoring, intelligent analysis, adaptive learning, and sustainable workload management.
The Data Collection Layer collects operational and environmental data from distributed industrial infrastructures in real time. Examples of input parameters included in the Data Collection Layer are CPU utilization, memory usage, network bandwidth used, storage I/O statistics, workload arrival rates, carbon intensity data, SLA status, and energy consumption records. When these input parameters are aggregated together, they provide a complete picture of the operational status of the computing environment.
The Processing Layer processes incoming raw workload data and generates a standard feature vector to represent the state of the system (processing phase). The raw workload input data are cleaned, transformed, and aggregated into a standard set of feature vectors that will be suitable for machine learning preprocessing. Workload clustering is then performed using HAC, which clusters workloads according to similar behavioral characteristics. By utilizing cluster profiling, the quantity of computational resources required to make a decision is reduced, and more scalable decisions can be made based on these workloads. The processed state data (i.e., outcome of HAC) is processed into an n-dimensional state vector, which is used for reinforcement learning optimization.
The Decision Layer is the core intelligence element of the HAC-DQNCAPS framework and is driven by a DQN agent. The DQN receives the processed state data of the system as input and determines what the best workload to execute is.
The Execution Layer utilizes Cloud Orchestration Services and Resource Controllers to implement its chosen actions, such as VM Deployment, Workload Migration, Adaptive Scaling, and SLA Monitoring. Continuous Monitoring provides feedback to the Data Collection Layer so that the architecture can facilitate Iterative Learning and Adaptive Optimization.
Within this architecture, there is an ability to intelligently optimize in a closed loop fashion while balancing the sustainability and operational performance required by organizations.
3.3. Mathematical Formulation
The optimization goal of the proposed framework is to ensure that sustainability and performance are balanced through minimizing total carbon emissions and energy consumption and ensuring that SLA requirements will not be violated. This has been formulated as a multi-objective optimization problem, where multiple system objectives have been combined into one weighted, objective function. The trade-offs between energy efficiency, environmental impact, and quality of service will all be considered through defined objective function parameters in the formulation of the mathematical equations.
In this formulation, F defines the entire cost function to minimized. E defines the total amount of energy consumed by computing resources (servers, networking devices, and cooling). C defines the total quantity of carbon emissions produced from energy used (which is based upon the carbon intensity associated with where and when the energy is consumed). Lastly, S defines the penalty associated with SLA violations (e.g., delay of task completion, missing deadlines, or degraded service).
α, β, and γ are weighting factors that control the relative importance of these objective components. A higher value of the weighting factor α implies greater emphasis on energy efficiency; similarly, higher values for the weighting factor β indicate a greater emphasis on carbon reduction. Higher values for the weighting factor γ results in a stricter adherence to SLA requirements. The weights associated with the objective components can be adjusted to reflect the operational priorities of the industrial system, thereby allowing for the overall sustainability and performance goals of the model to be varied based on the needs of the user.
The total energy consumption is calculated as
where
Carbon emissions are computed using carbon intensity data:
where
SLA violation penalties are expressed as:
where
The framework additionally satisfies the following constraints:
The DQN updates Q-values using the Bellman optimization equation:
3.4. Hybrid Optimization Algorithm
The HC-DQNCAPS is a new optimization algorithm, as shown in
Figure 2, that utilizes a combination of HAC and DQN reinforcement learning (RL) techniques for managing workloads in an environmentally friendly and dynamic manner.
Initially, the HAC groups workloads into clusters based on shared characteristics (i.e., CPU utilization, memory consumption, storage access patterns, and network characteristics), which enables the state-space to be significantly reduced in dimensions, and thus large-scale optimizations to be conducted.
Once the workloads have been clustered, the workload profile for each cluster is fed into the DQN agent, and the agent can then observe the state of the system and choose appropriate actions according to an ε-greedy selection policy; examples of actions that can be taken include VM/resource allocation, workload migration, workload procrastination, dynamic resource adjustment, and server consolidation.
To ensure that sustainability and performance metrics are taken into account when assigning rewards to the agent based on the actions it takes, the reward function contains factors for energy consumption, carbon emissions, and SLA violations.
| Algorithm 1: HC-DQNCAPS optimization algorithm |
![Computers 15 00339 i001 Computers 15 00339 i001]() |
4. Experiment and Implementation
4.1. Experimental Setup
A simulated multi-cloud computing environment was developed and used for testing the proposed HC-DQNCAPS framework and simulating real industrial cloud infrastructure in order to evaluate the ability of the framework to optimize energy efficiency, reduce carbon emissions, increase SLA compliance, and maximize resource utilization under dynamic operational conditions. The simulation used multiple disadvantages data centers with different types of hardware and software resources, different operating characteristics such as energy usage, and regionally limited carbon dioxide emissions to create a realistic representation of geographically dispersed data centers using virtualized resources in real time. The experimental environment used simulated multi-cloud computing systems connected by various types of network links with different types of CPU/memory/storage/network components, as well as different types of data center sites. The simulation also attempted to simulate actual workloads placed on actual industrial clouds using actual carbon profiles for the actual services being provided. As such, the experimental environment provides an accurate model of the current state of the multi-data-center industrial cloud computing infrastructures using actual virtualization technology, on-demand processing capacity, and elastic resource scalability in a virtual environment.
Within this simulation environment, multiple intelligent operational capabilities were integrated including dynamic provisioning of virtual machines, adaptive migration of workloads, monitoring of carbon intensity in real time, scheduling policies considering SLAs, and managing workloads with variable characteristics. Therefore, the framework it is based upon is able to dynamically adjust to varying levels of workload demand, energy supply and environmental conditions while sustaining reliable service and operational efficiency.
A wide variety of workload types were created within the simulation environment to evaluate how well the framework could adapt to these workloads. The types of workloads included CPU-bound workloads (scientific computation and training AI), memory-bound workloads (large databases that were stored on disk), I/O-bound (large amounts of network traffic and disk storage) when used with industrial applications, and real-time streaming workloads (strict latency and SLA requirements). The framework was evaluated based on its responsiveness to steady-state and burst workloads in stable and highly variable operational conditions.
Table 1 provides an overview of the HC-DQNCAPS architecture, the comparative approaches, the evaluation metrics, and the statistical validation methods employed.
4.2. Implementation
A Deep Q-Network structure was used to implement the HC-DQNCAPS framework, which incorporated hierarchical clustering for adaptive workload optimization.
The DQN neural network consisted of three hidden layers of fully connected ReLU-activated units. The Adam optimization method was used to train the network and minimize mean squared error loss.
In order to support stable training and to help diminish temporal correlation between successive experiences, experience replay was implemented. Experience replay maintained a replay memory of 100,000 transitions, from which mini-batches of 64 randomly sampled transitions were selected during the training phase.
The use of a separate target network is critical to improving the convergence stability and to minimizing the oscillatory behavior caused by the learning of the DQN network. Periodic updates to the target network were made every 1000 iterations.
An ε-greedy policy was used to balance between exploration and exploitation. Initially, exploration dominated through ε = 1.0, which gave the agent ample opportunity to develop a variety of experiences in the environment over time. As training progressed, ε decreased steadily towards 0.01, allowing for the agent to use optimal policies to exploit previous training.
Hierarchical Agglomerative Clustering was accomplished through the use of Ward linkage criteria and Euclidean distance metrics. The implementation of hierarchical clustering also allowed for clustering with respect to workload behavior prior to reinforcement learning, which in turn lowered the computational complexity of the reinforcement learning process. The computational complexity of HAC is expressed as
The DQN training complexity is given by
where
training episodes,
batch size,
network weights.
The framework achieved convergence after approximately 450 training episodes, where reward variance stabilized below 2%, demonstrating efficient learning and stable policy optimization.
5. Results and Discussion
5.1. Energy Consumption
In
Figure 3 we can see that our HC-DQNCAPS method had the lowest amount of energy use compared with all of the other solutions tested. The traditional First-Come-First-Serve scheduling method used the highest amount of energy because it allocated workloads sequentially without ever looking at how efficient resources were or how optimally the system was actually using its resources. The Energy-Aware VM Allocation solution decreased energy use through consolidating (reducing the number of) virtual machines, but it was not an effective method due to the lack of intelligent adaptive learning capabilities.
Carbon-Unaware RL used a method based on reinforcement learning to provide more efficient energy use when allocating workloads; however, its effectiveness in actually using less energy was reduced due to the lack of carbon-aware optimization methods. Methods such as PPO, DDQN and MADRL optimized their use of energy better than these methods due to their use of more complex reinforcement learning models, but they were not as efficient as our HC-DQNCAPS framework in terms of energy consumption. The HC-DQNCAPS framework achieved an average of approximately 30 to 35% less energy consumption than the FCFS method through the combination of workload clustering, adaptation of VM migrations, carbon-aware workload scheduling and dynamic resource allocation. In addition, the hierarchical clustering approach made it possible to significantly reduce the number of idle resources used across the infrastructure, which minimized any unnecessary power consumed.
5.2. Carbon Emissions
The carbon emissions performance of all investigated methodologies is depicted in
Figure 4. The framework HC-DQNCAPS performed the best in terms of carbon emissions with its carbon-aware workload scheduling methodology. The methodology of the framework combined real-time carbon intensity data with the workload placement decision-making process, which allowed for the shifting of workloads to data centers that were powered by cleaner energy sources as carbon intensity decreased, unlike traditional methodologies where performance or energy reduction were primary concerns.
The scheduling method of FCFS resulted in the greatest amount of carbon emission because workload allocation was not dependent on environmental conditions at all. The Energy-Aware VM was able to reduce emissions moderately due to savings from using energy efficiently; however, it also did not possess explicit carbon-aware scheduling capabilities. Although the Carbon-Unaware RL application improved the efficiency of allocations, it did not optimize placement of workloads according to regional carbon intensity variations. Advanced DRL methods, such as PPO, DDQN, and MADRL, improved environmental performance due to the scheduling capabilities being more intelligent; however, the HC-DQNCAPS framework was superior due to optimizing energy efficiency and being carbon-aware. Overall, the HC-DQNCAPS framework reduced approximately 25–30% of the carbon emission compared to traditional scheduling methodologies.
5.3. SLA Compliance
Figure 5 shows that the HC-DQNCAPS framework had the lowest rate of SLA violations across all tested methodologies. The highest number of SLA violations occurred with the FCFS scheduling because it processed tasks chronologically without prioritizing workloads or optimizing adaptively. The Energy-Aware VM approach had some SLA violations due to overly aggressive resource consolidation to save energy.
The Carbon-Unaware RL and the more progressive reinforcement learning methodologies like PPO, DDQN and MADRL all showed improved SLA performance because of their more intelligent scheduling policies and adaptive resource management; nonetheless, they were all still incapable of incorporating the integrated clustering-based workload optimization that the proposed framework contained. The HC-DQNCAPS also fulfills the sustainability objectives of reducing carbon emissions and use of natural resources while ensuring 100% reliability of service by introducing SLA penalties directly into the reward function of the reinforcement learning process, enabling the framework to sustain SLA violations of <5%, thus demonstrating that optimizing the environment is possible without sacrificing quality of service.
5.4. Resource Utilization
As can be seen in
Figure 6, the HC-DQNCAPS framework had the highest resource utilization efficiency when compared to all other evaluated methods. The FCFS algorithm used inefficient resource utilization because there was no optimization or load balancing in the allocation process. However, Energy-Aware VM Allocation achieved higher utilization through consolidation of servers, but static optimization mechanisms limited the scalability and adaptability of the solution.
Through the use of reinforcement learning-based techniques such as PPO, DDQN and MADRL, there was some increase in the quantity of resources used by the systems due to dynamic allocation strategies learned from their respective environments; however, these methods processed workloads on a one-to-one basis without optimizing for workload clustering.
The HC-DQNCAPS framework had much better resource utilization compared to all other methods as a result of the Hybrid Clustering and DRL techniques being utilized. Workloads with the same type of resource characteristics can be clustered together to allow for better resource allocations while also reducing the idle capacity of the infrastructure. The HC-DQNCAPS framework had approximately 20% higher utilization than all other baseline methods, while also reducing energy consumption and carbon emissions.
6. Conclusions and Future Work
The study introduced a new framework for sustainable computing, called the HC-DQNCAPS, that will enable intelligent, environmentally friendly workload management in industrial cloud computing. The proposed framework uses principles from Green AI, hierarchical workload clustering, and deep reinforcement learning to optimize energy consumption, reduce carbon emissions, comply with Service Level Agreements (SLAs), and utilize resources in a geographically distributed cloud computing environment.
The combination of HAC with DQN RL allows for adaptive, real-time, and carbon-aware workload scheduling that is capable of responding in an adaptive manner to dynamic workload demand, changing infrastructure situations, and variable environmental factors.
The proposed HC-DQNCAPS framework was designed to maximize resource utilization while minimizing total energy consumption and GHG emissions and strictly complying with SLA requirements by using a multi-objective optimization model. The framework utilized real-time resource monitoring, carbon intensity awareness, adaptive workload migration, and intelligent VM allocation to support the operation of sustainable cloud computing. The utilization of clustering techniques also eliminated much of the complexity of decision-making and improved the scalability and efficiency of the reinforcement learning process.
We used simulated/industrial cloud environments for an experimental evaluation to show the performance and reliability of the proposed framework. Comparing the results with the FCFS scheduler, Energy-Aware VM Allocation, Carbon-Unaware RL (PPO), DDQN, and MADRL, shows how effective and efficient the HC-DQNCAPS was in all cases compared to other schedulers. Current techniques provide savings from the HC-DQNCAPS of approximately 30–35% in energy savings, 25–30% in carbon emission savings and +20% in resource utilization savings and SLA violations are less than 5%. As evidenced by Convergence Analysis, the DQN agent reached convergence after about 450 training episodes, proving the stability and efficiency of the policy. Learning evidence demonstrated by statistical validity using ANOVA and Wilcoxon signed-rank tests shows the results were reliable and statistically significant (p < 0.05) at 95% confidence intervals.
By providing the capability to deploy sustainable and carbon-aware AI solutions without an impact on operational efficiency (performance) or the operational capability of the HC-DQNCAPS framework’s modular and scalable architecture makes it appropriate for many different industrial computing use cases such as data centers operating in the cloud, hybrid cloud businesses, and distributed edge computing environments. The use of environmental awareness in reinforcement learning-based scheduling demonstrates how intelligent Green AI systems have the potential to address many current sustainability challenges.
Despite promising results from the proposed framework, there are still limitations that can be addressed in future work. The scope of this study was limited to a simulated cloud environment, thus requiring future work to implement and validate the framework on production-scale industrial infrastructures in the real world. More research is needed to test the framework for heterogeneous hardware configurations, variable network delays, and noisy sensing environments typically found within operational industrial systems.
The direction for future research is to expand the HC-DQNCAPS framework to include renewable energy forecasting and carbon intensity forecasting models through the use of advanced time series learning techniques, such as Long Short-Term Memory (LSTM) networks and Transformer-based forecasting architectures. The predictive features of these forecasting techniques will permit proactive scheduling decisions for workloads based on the anticipated availability of clean energy. In addition, the framework will be developed to work in edge computing and Internet of Things (IoT) environments, where there are challenges related to resource constraints, latency sensitivity, and distributed intelligence optimization.
Further research will focus on using Multi-Agent Reinforcement Learning (MARL) and Federated Reinforcement Learning for decentralized workload optimization across large cloud–edge ecosystems, thus providing new ways to achieve scalability and distributed decision-making. Moreover, considering carbon pricing policies and energy market complexities will also yield new avenues for research.