Next Article in Journal
Hybrid Knowledge Distillation for Edge-Efficient Video Action Recognition: Improving Lightweight 3D CNNs via Joint Distillation
Next Article in Special Issue
Heuristic Cross-Temporal Reconciliation Approaches Applied to Heterogeneous Models in Photovoltaic Forecasting
Previous Article in Journal
Investigation of Augmented Datasets for Security in Internet of Medical Things (IoMT) Ecosystems
Previous Article in Special Issue
Editorial: Machine Learning and Statistical Learning with Applications 2025
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Systematic Review

Traffic Congestion Prediction Algorithms in Urban Environments: A Survey

1
Department of Computer Science, Tshwane University of Technology, Pretoria 0001, South Africa
2
School of Agriculture and Science, Westville Campus, University of KwaZulu-Natal, Durban 4000, South Africa
*
Authors to whom correspondence should be addressed.
Computers 2026, 15(6), 370; https://doi.org/10.3390/computers15060370
Submission received: 24 March 2026 / Revised: 29 May 2026 / Accepted: 1 June 2026 / Published: 5 June 2026

Abstract

Traffic congestion poses a significant challenge in urban environments. The use of digital techniques has emerged as a pivotal trend, as it offers substantial safety to and mitigates stress and frustration for road users. The purpose of this survey was to explore the current approaches and digital techniques for managing traffic congestion. We address this through a systematic literature review (SLR) approach by adopting PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines. We began by exploring the key techniques of topological data analysis (TDA), machine learning (ML) and deep learning (DL) for modeling urban traffic prediction. We evaluated the robustness of the topological data analysis technique (Persistent Homology (PH)) against deep learning frameworks (Graph Convolutional Neural Networks (GCNNs)). We found that each framework has its own strengths and weaknesses, and neither of the frameworks independently provides a complete solution. PH may offer richer structural insights and robustness to noise but may struggle with direct predictive implementation, while deep learning models do better at extracting dynamic predictive patterns but are assumed to lack interpretability and generalizability. Therefore, the integration of multiple techniques, either PH with stacking ensemble methods or deep learning with stacking ensemble methods, can improve prediction and generalization of the model while at the same time reducing over-reliance on local graph assumptions. Future research should focus not only on performance metrics or methods but also on explainability, transferability, adaptability across heterogeneous road environments and computational cost.

1. Introduction

Traffic congestion is a condition characterized by slower speeds of vehicles, longer travel times, and increased queuing of vehicles on the roadways. This condition poses a significant challenge in many urban environments worldwide [1]. In [1], the authors presented two approaches of traffic congestion management, namely dynamic and static approaches. The dynamic approach is described as a method of managing traffic congestion in real-time, while static is managing traffic congestion offline. The problem of traffic congestion is presented as two groups: recurring and non-recurring congestion [2]. The former defines congestion caused by bottlenecks, signal timings, or capacity, while the latter describes congestion caused by incidents, work zones, weather, or extraordinary events. Several authors presented road traffic congestion as an ubiquitous problem, caused by multiple factors, such as rapid urbanization, an increase in vehicle ownership, poor infrastructure planning, mixed traffic, and a limited sensing infrastructure [2,3,4].
Existing approaches have been finding it difficult to manage the tremendous growth of vehicles on the roadways [5], leading to traffic congestion on urban roads. Consequently, traffic congestion has slowed down economic progress [6] and increased travel time, which has led to high fuel consumption. Several authors consider traditional statistical models as commonly used methods for prediction and managing traffic congestion [7,8,9]. Despite the effectiveness of state-of-the-art traditional methods, they still rely on single-model approaches that fail to address complex spatiotemporal [10,11] aspects and explore hidden patterns in the datasets. Table 1 presents the quantitative performance results that have been reported for traditional traffic prediction approaches from the reviewed literature. However, it is important to understand that Table 1 only serves as a descriptive synthesis of the reported performance metrics and does not present a comparative evaluation of the traditional prediction methods. Therefore, thoughtfulness should be applied when interpreting the reported findings in this table, as dissimilarities in datasets may influence reported performance results.
Therefore, this systematic review is motivated by the fundamental problem of traffic congestion, particularly on urban roads, and whether the integration of an ensemble with persistent homology (PH) or hybrid deep learning frameworks can be used as an alternative solution to the problem of traffic congestion and to improve prediction. For this reason, this study aimed to critically examine whether the integration of ensemble frameworks with deep learning frameworks can enable a model to capture complex spatial or structural characteristics of urban roads and improve its efficiency prediction. This survey was guided by two hypotheses: H1—The integration of ensemble learning models with PH significantly improves the precision of traffic prediction compared to traditional prediction methods; H0—The integration of ensemble learning with persistent homology does not significantly improve traffic prediction precision when compared to traditional prediction methods.

2. Background

This review focuses on the integration of PH with an ensemble method to predict traffic congestion on urban roads. Topological data analysis (TDA) and machine learning (ML) techniques [13,14] have been extensively examined in the literature as viable strategies for mitigating traffic congestion issues within intelligent transport systems (ITS). In this context, PH offers a robust framework for modeling traffic congestion prediction by uncovering hidden traffic patterns and intricating structural relationships within traffic data, especially when used with point cloud representations. Recently, deep learning architectures have also appeared as a strong contender for modeling urban traffic congestion prediction [15], as they have the ability to extract complex spatial temporal dependencies within road networks. Although several studies have identified TDA and deep learning techniques as potential tools for modeling traffic congestion prediction, not enough has been done to resolve the problem of traffic congestion in urban areas. Several cities in the world still depend on traditional approaches to manage traffic congestion on urban roads. Consequently, it is essential to evaluate the robustness of PH and deep learning (DL) frameworks in modeling traffic congestion prediction. This study fulfils this requirement by analyzing the advantages and disadvantages of these two frameworks and exploring their potential synergy to enhance traffic congestion forecast efficacy.

2.1. Persistent Homology

Recent advances in the use of novel techniques to model traffic prediction have seen an increase in the adoption of TDA techniques, particularly PH. This technique is one of the emerging methods in the TDA framework that has arisen as one of the strongest techniques to address issues pertaining to hidden traffic patterns and is typically applied to point cloud data. This tool can compute topological features, like Betti numbers β 0 , β 1 , β 2 , persistent diagrams (PDs) and persistence barcodes. When modeling with PH, features are seen as a connectivity of features, clusters, loops, or global structures [16,17,18], while without PH, features are seen independently. The traditional ML models, like random forest, decision tree, and support vector machines (SVM) [18,19,20,21,22,23], fail to explore hidden topological patterns. The connectivity of these features in the dataset presents the likelihood of bottlenecks, which gradually lead to traffic congestion. The study [24] investigated the critical speed method, which can identify traffic congestion periods in the bottleneck sections. The study incorporated a stationary density–flow relationship at the key bottleneck areas where vehicles move slowly to identify congestion.
Several studies present PH as a rigorous method for computing robust topological features in discrete experimental observations that often contain various sources of uncertainties [25,26,27]. In the study conducted by Turkeš and Montúfar [28], they considered PH as an extension of homology. However, one of the popular techniques in PH used to identify topological features is Ripser. In [1] used Ripser.py version 0.6.15 to commutate persistent homology features. It is used to compute persistent features and manage memory in a parallel distributed environment [29,30]. It is a powerful tool as it can analyze and understand complex datasets through the lens of topology. In many TDA-related studies, Ripser has been presented as an effective tool, particularly for computing persistent homology on large, high dimensions [20,31]. Its use has inspired the development of several extensions and wrappers that further enhance its application in topological data analysis. Figure 1 illustrates the filtration process used to compute barcodes and persistent diagrams using Ripser.
Figure 1 [32] presents how topological features evolve from birth to death across a filtration parameter. It also demonstrates how evolution is encoded into persistence barcodes and persistent diagrams. The dashed line in the middle represents the diagonal line. It is a reference that indicates features with zero persistence. The points are very close to diagonal has low persistence and short lifetime. The points that are closer to diagonal have low persistence and short lifetime. While the points far from the diagonal have high persistence and long lifetime. The high persistence features represent more significant topological structures in the dataset.

2.2. Topological Features in Traffic Data

Road networks have several topological features such as connectivity, which is the connection of a number of paths within the road network. A highly connected road network often offers redundancy and alternate routes, while less connected road networks are more susceptible to disruptions [33]. Therefore, a clustering coefficient is used to measure the tendency of nodes to cluster together to form a local community [34]. Densely clustered areas are segments of the road network that usually experience heavier traffic. At these segments of the road, vehicles form a cluster when moving closer to each other, particularly when approaching an intersection or traffic cycle, while sparser areas indicate underutilized segments [35,36]. Therefore, PH techniques can unveil and compute these intricate features, which individual models do not identify in a traffic dataset.

2.3. Ensemble Stacking Approach

In recent years, the use of ML ensemble approaches has appeared as a powerful approach for transportation systems, which can be integrated with other techniques. An ensemble is also known as a meta-learning technique where multiple base models are trained, and the outcomes from their predictions are used as inputs to a higher-level meta-model. There are several examples of ensemble methods, including Bagging, Gradient Boosting, AdaBoost, and Stacking [37,38]. Several studies [37,38] consider ensemble methods as techniques to enhance the prediction of the meta-model. In this approach, base models are combined to reduce the weakness of individual models [39]. The learning of these approaches is deployed in three steps: (1) base models are fitted to the data of the first fold; (2) base models predict observations in the second fold; (3) the meta-model is trained from the output predictions from the base models. The ensemble process for the trained model is iterated for a number of folds (n-folds). This process delivers two types of models: the Tier 1 model includes base models and a combination of multiple base models, and the Tier 2 model includes a meta-model trained from the outputs from Tier 1. The assessment of the performance of the trained models is performed using the Pearson correlation coefficient, root mean squared error (RMSE), and confusion matrix. The confusion matrix approach compares predicted and actual labels to find the levels of congestion; however, cross-validation is also used to evaluate the model’s generalization ability and mitigate issues of model overfitting [40,41,42]. With the ensemble stacking framework, models can overcome the limitations of individual models and enhance the prediction of the meta-model [37,43,44,45]. The studies [46,47] presented details of designing an ensemble model using labeled data corresponding to the distinct levels of traffic congestion (Figure 2).
Figure 2 presents multiple layers of an ensemble stacking framework for traffic predictive modeling. In this figure, multiple machine learning algorithms have been used in the ensemble model to generate predictions aggregated from four base models. The outputs from the base models train the meta-model. This hybrid architecture can improve predictive accuracy and generalization (note: Figure 1 was created by the authors based on concepts discussed in [47]). It is important to note that this article is a survey, and its purpose is to contribute to the development of a hybrid model in the near future that can combine PH and a stacking ensemble approach.

2.4. Real-World Applications of ML, PH and DL Architectures

Machine learning algorithms have been widely used for modeling prediction of traffic using historical traffic data to discern traffic behavior patterns and relationships. These methods have been used along with other techniques to extract hidden meaningful insights from datasets [48]. In the study [49], the authors proposed the development of a model that integrates persistent homology with machine learning to predict traffic congestion. TDA components extracted spatial and temporal features from diverse data sources, which were used as inputs for ML algorithms to predict traffic congestion. The model demonstrated accuracy in predicting congestion on the roads of New York. In the study [50], the authors considered the use of TDA techniques to model traffic congestion prediction. The proposed model utilized traffic data from traffic sensors, global positioning systems (GPS) probe points, and weather data and exhibited accurate predictions up to 24 h in advance. Nguyen et al. [51] and Wu et al. [52] explored the integration of the TDA technique, principally PH, with ML to predict traffic congestion. The proposed model achieved high accuracy in the prediction of traffic congestion. The work by [53] identified bottlenecks and meaningful patterns pace structures using PH algorithms. The proposed model identified bottlenecks in the road network, which could not be identified by conventional machine learning. In [54], the authors confirmed the incorporation of PH and ML to extract topological structures. The study used a dataset collected from the urban roads of Shenzhen, China, including geographical and temporal information. The dataset provided an ideal sample for urban residents’ behavioral patterns and successfully identified five patterns of residents’ activities. The understanding of patterns in a dataset can assist with the formulation of dynamic traffic management strategies, like adjustment of traffic signals and the rerouting of traffic from congested areas [25,55].
As traffic congestion prediction continues to advance, the adaption of deep learning architectures is also increasing, enabling the extraction and modeling of complex spatial–temporal dependencies within a road network. The study by Qi and Cheng [56] proposed the integration of TDA with deep learning (DL) frameworks to model traffic congestion prediction. In the study, the TDA component was used to extract and engineer topological features, and the DL algorithm, using features from TDA, managed to predict traffic congestion. The proposed model utilized data aggregated from road traffic sensors, global positioning systems (GPS) powered vehicles, and weather data. Kumar et al. [57] and Wang et al. [58] recommended the use of advanced deep learning techniques, particularly convolutional neural networks (CNNs) [59], long short-term memory (LSTM) networks, and support vector regressions, to analyze and predict traffic congestion in Bengaluru, India. The proposed model provided both correct congestion detection and routing recommendations. Therefore, it is important to explore the ability and potential synergy of TDA and DL techniques to enhance traffic congestion prediction effectiveness.

2.4.1. Graph Neural Networks

In recent advancements in traffic congestion prediction, deep learning architectures like GCNNs [60] and transformers have been increasingly recommended by many scholars. These architectures have gained popularity because they consistently outperform conventional models in capturing complex patterns and improving predictive performance [61,62]. Their ability to extract complex spatial temporal dependencies within road networks has made them powerful tools for traffic congestion prediction. Several studies have considered graph convolutional neural networks (GCNNs) [15,60] and transformers to effectively capture hidden and dynamic relationships, particularly when road segments and intersections are represented as nodes and edges. The significance of adopting graph convolutional neural network (GCNN) architectures depends on their capability to handle non-Euclidean traffic data, exploit spatial dependencies, and construct graph topologies, typically represented through a fixed adjacency matrix [63,64]. Although GCNN architectures often demonstrate high predictive performance compared to traditional convolutional neural networks (CNNs) [59], their effectiveness remains dependent on assumptions related to data representation and network structure, raising concerns regarding robustness and generalizability. It is assumed that GCNN’s powerful performance is often achieved under controlled experimental settings. Empirical evidence suggests that GCNN models benefit significantly from large-scale datasets having substantial numbers of locations and long-term temporal observations [64,65]. Nevertheless, insufficient data can negatively affect their performance. This dependence on large data raises practical concerns regarding scalability and real-world deployment, particularly in resource constrained environments where a comprehensive traffic dataset may be unavailable. These limitations introduce methodological challenges and raise questions about whether the superiority of GCNN models reflects applicability or is primarily dependent on ideal conditions. Further, such limitations suggest a potential mismatch between graph representation assumptions and real-world traffic behavior.

2.4.2. PH-Based Model

In contrast to GCNN architectures, persistent homology (PH) has also emerged as a powerful topological data analysis (TDA) tool by adopting a different analytical perspective. Instead of focusing only on local relationships or node connectivity, it explores and finds global geometric and topological structures within a dataset. Through multi-scale analysis, PH can detect higher-order structures, such as connected components, loops and voids. Unlike graph-based deep learning methods, PH does not rely on predefined assumptions about network adjacency and therefore offers a more flexible framework for capturing intrinsic data structures. It can distinguish persistent topological patterns from transient noise through persistence diagrams and barcodes. This property is valuable in traffic congestion prediction because it enables the identification of structural patterns and uncertainties [25,26,27].
Despite its strengths, PH cannot be directly integrated into conventional machine learning pipelines without transformation [66]. This creates a representational challenge where valuable topological information risks being distorted during feature conversion. Using PH alone may cause the model to capture only structural characteristics while failing to fully model the temporal dependencies required for real-time prediction, which can significantly affect prediction accuracy. These limitations suggest that PH alone may not be enough for comprehensive congestion modeling and may require integration with other techniques. PH also has a high computational complexity, particularly when analyzing higher-order structures [28,67]. Balancing the strengths and weaknesses of PH and deep learning models shows that neither approach independently provides a complete solution. Deep learning architectures excel at extracting dynamic predictive patterns but often face challenges related to interpretability and generalizability. PH offers richer structural insights and robustness to noise but struggles with direct predictive implementation. Therefore, integrating PH with either ensemble or hybrid deep learning frameworks may improve structural interpretability while enhancing predictive performance [41]. It is important to note that Table 2 only serves as a presentation of the experimental set up and does not compare deep learning and TDA frameworks. Therefore, thoughtfulness should be applied when interpreting this table and any reported dissimilarities may influence the results. Figure 3 presents the proposed integration of PH with the stacking ensemble method. Table 2 shows the experimental set up of GCNNs and PH.
Figure 3 demonstrates the process of the proposed ensemble model. The blue circle represents the online data source; the red dashed arrow stands for the data flow necessary for training of the model. While the white rectangular boxes represent the multiple machine learning algorithms as well as the prediction outputs aggregated from the four base models. Lastly the meta-metal model is also represented by white rectangular box, and it uses the outputs from the ensembles to make a final prediction.

2.4.3. Critical Evaluation of Hybrid Models Using Ensemble Methods

In [68], the authors confirmed that the ensemble model can enhance interpretability and elimination of hyper-parameter tuning. Study [45] suggested that ensemble stacking can improve classification performance through the integration of multiple base learners. It enhances prediction accuracy by leveraging the strength of multiple base models while improving model generalization [69]. Despite the powerful performance of stacking ensemble models, they often rely heavily on the variety and variance of the base classifiers, which can negatively affect overall model performance if not effectively managed [69]. In addition, stacking models may struggle with large datasets due to increased computational and training time [70] requirements, which may lead to suboptimal performance [42]. The complex of ensemble models can potentially also lead to overfitting, which is a notable limitation compared to traditional homogeneous ensemble methods. Integrating an ensemble with frameworks such as PH or GCNN architectures is a complex task that may require high computing costs compared to individual models [69], but it can yield improved performance.

3. Materials and Methods

3.1. Materials

Table 3 presents the materials used in this survey. It also presents the details of the manufacturer and the location where the product was manufactured, including the city, state and the country. However, it is important to understand that Table 3 only serves as a descriptive synthesis of the materials used in this study. Therefore, thoughtfulness should be applied when interpreting the reported findings.

3.2. Aim and Objectives

The aim of this survey was to provide the scope to determine relevant studies. We did this by finding a set of research questions in the context of finding the current approaches and the challenges faced by urban environments when trying to implement solutions to mitigate traffic congestion. The following are details of the research questions: 1. What challenges affect the accuracy, implementation, and application of predictive models in urban road traffic systems? (RQ1); 2. What challenges affect the accuracy, implementation, and application of predictive models in urban road traffic systems? (RQ2); 3. To what extent can the integration of persistent homology and ensemble stacking framework techniques improve the development and predictive performance of traffic congestion models in urban road networks? (RQ3). This systematic review was conducted in strict accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidelines [71] to ensure a comprehensive, transparent, and reproducible research synthesis process. The method incorporates a structured multiphase workflow, which includes identification, screening, eligibility, and inclusion to systematically analyze TDA, ML and DL techniques for prediction of traffic congestion in urban environment roads. To enhance methodological transparency and rigor, the full PRISMA 2020 checklist [2] has been adopted, and a consensus-based protocol was implemented to resolve any disagreements during the study selection phase. This rigorous approach provides a structured, evidence-based foundation in this survey, which could be valuable for guiding future research, informing critical decision-making, and facilitating real-world practical applications in areas of traffic management, congestion prediction solutions and algorithms, prevention, and formulation of traffic management strategies. It also provides guidelines that enable experts to conduct a formal and impartial survey to recognize, evaluate, and refine research questions.

3.3. Data Source

A systematic literature search was performed across six major academic databases: ScienceDirect, ACM digital Library, MDPI, Springer Nature, IEEE Xplore and Google Scholar. These six platforms enabled the retrieval of a wide range of journal and conference publications, resulting in a comprehensive and representative collection of recent studies related to prediction of traffic congestion in urban roads.

3.4. Search Strategy

We adopted a systematic and transparent search strategy to ensure the comprehensive identification of relevant studies, minimize potential biases, and set up a solid foundation for the analysis. The search phrases and keywords were based on information extracted from the titles, abstracts, theoretical frameworks, and curated collections of related studies. The search for this survey only focused on the journal articles and conference papers that were published between 2017 and 2026, as this timeframe is presumed to represent the most active phase of research on topological data analysis (TDA) and machine learning (ML) techniques for predicting traffic congestion in urban roads. It is understood that the literature published within this period is highly relevant to the focus of this survey. To ensure quality, the review primarily targeted peer-reviewed journal articles and conference papers, as these sources provide rigorous, original, and in-depth contributions to the field. Studies that were not peer-reviewed or appeared non-original despite their relevance were excluded from the survey to uphold the methodological robustness of the review. Table 3 outlines the specific search strategy that was employed in this survey. The search that was employed in this study combined keywords and Boolean operators focusing on three core themes: (1) Prediction methods: “Traditional methods” and “Hybrid techniques”); (2) Implementation and application context challenges: “Algorithmic bias”, “accuracy of prediction”, “target type”, “prediction zone” and “urban roads”, “Urban setting,” “GPS and infrastructure”, “Real-time data source”; and (3) Model architectures: “Deep learning architecture”, “CNNs”, “GCNNs”, “TDA based model”, “Hybrid based models”,” Persistent homology with Ensemble approaches”. The search was limited to publications between January 2017 and June 2026 to capture the most recent advancements and insights in ML, Deep learning and topological data analysis and machine learning-based congestion prediction.
The initial search yielded 1857 records, which were later screened according to the PRISMA guidelines [2], as illustrated in Figure 4. Table 4 presents the search strategy and data sources selected for this survey. It also further presents the keyword strategies that were used during the literature search process. Finally, it presents the conceptual scope, disciplinary coverage, and methodological direction of this survey review. The selection of the right database for this study and combination of the keyword strategies offered the insight of the selected studies related to persistent homology (PH), deep learning, topological data analysis (TDA), machine learning (ML), ensemble learning, and traffic congestion prediction. The search strategy employed also has important implications for comprehensiveness and potential biases affecting traffic modeling, particularly for urban roads. Furthermore, the database search strategy employed in this study revealed the multidisciplinary nature of traffic congestion prediction research.
This shows that there is a growing transition of predicting traffic congestion toward graph-based (GCNNs), ensemble, and PH frameworks. The common discussion of concepts such as spatial dependencies, adjacency matrices, and non-Euclidean traffic data confirms that traditional Euclidean representations are inadequate to capture urban traffic complexity. Therefore, ensemble learning can also support the development of integrated predictive frameworks and use topological features for dynamic urban traffic prediction.

3.5. Inclusion and Exclusion Criteria

Using relevant Creative Commons (CC)-based keywords, a total of 1857 records were initially identified and screened. Of these, 800 records were removed as duplicates, leaving 1057 records for further screening. An additional 600 records were excluded because they did not originate from peer-reviewed journals. This resulted in 457 records sought for full retrieval, of which 150 could not be retrieved and were excluded because they were just reports rather than journal articles. Following the retrieval stage, 307 records were assessed for eligibility. However, not all of these met the inclusion criteria, as several studies did not adequately address the core themes of this survey. At the end, we only selected sixty-six (66) journal articles to be included in the final review. Table 5 presents the inclusion criteria that were used to select eligible journal articles. It also presents the period of the distribution of the selected articles. The most included studies were those published between 2020 and 2024, followed by those published between 2025 and 2026. Table 6 presents the exclusion criteria used for ineligible journal articles.

3.6. Quality Assessment

We assessed the rigorous methodological quality of the included studies by adopting the standard of modified version of the AMSTAR2 tool [72]. The assessment focused on several critical dimensions related to performance of traditional methods in urban roads: the TDA and ML research, including experimental rigor of traffic congestion predictive models (80% of studies), performance (70%), dataset description (90%), and statistical validity (54.5%). Furthermore, a strong majority of the included studies (90%) provided meaningful comparative analysis against established traditional ML models/approaches. Table 7 presents the comprehensive results of this quality assessment aligned to modified AMSTAR2 criteria and principles. In Table 7 quality assessment results for the relevant studies are presented. It includes the methodological quality assessment of the selected 66 studies for this survey review. The method used for quality assessment was to evaluate robustness, transparency, and comprehensiveness of the selected studies related to traffic congestion prediction, deep learning, persistent homology (PH), topological data analysis (TDA), and machine learning (ML).
This study initially found 1857 studies, which were reduced to only 66 eligible studies. This proved that there is scarcity of research about the integration of urban traffic congestion prediction with PH and the stacking ensemble method. The exclusion strategies applied in Table 6 reveal a significant gap in urban traffic congestion prediction focusing on predictive frameworks, supporting the need for using PH and ensemble-based methods to model complex spatial dynamics of urban traffic systems. The quality assessment strategy used in this study demonstrated a general strong methodological rigor and dataset transparency among selected studies; however, notable weaknesses remain in the comparative evaluation and reporting practices on selected techniques, particularly a low proportion of studies identified for comparative analyses between PH, CNN, and GCNN models. This suggests that there is still a wide research landscape gap and a lack of standardized benchmarking procedures. These findings also signify that future research should focus on developing integrated and comparative predictive frameworks capable of leveraging both PH and deep learning architectures to improve prediction of traffic congestion in urban environments. Table 7 presents the quality assessment result process, while Table 8 presents the study selection process of how the studies for this survey were progressively identified/selected, screened, and excluded. It presents the complete process from the identification of studies to final inclusion. It also includes the number of records identified and exclusion reasons.

3.7. Data Extraction and Synthesis

Table 8 and Table 9 present a standardized data extraction procedure implemented to six categories of information: (1) model architecture specifications; (2) prediction techniques; (3) evaluation metrics; (4) dataset descriptions; and (5) prediction approaches. Table 9 details of the high completeness for model architecture details (100%) and prediction techniques (90%), although performance metric details were consistently reported a bit low (70%). Table 10 provides a comprehensive summary of data extraction protocols. distribution and reporting completeness from the 66 included studies. The extracted data for this study indicated that there is a growth in prediction of traffic congestion from the use of a traditional single model to deep learning architectures and hybrid topological machine learning frameworks (PH frameworks). Although deep learning architectures remain dominant, increasing research on PH-based models suggests that traditional methods alone may not be able to capture the complex dynamics of urban traffic; therefore, they need to be integrated with other powerful frameworks. Integrating topological features with ensemble or graph-based learning methods may improve representations of spatial and structural traffic patterns.

3.8. Synthesis Outcomes

The extracted data was synthesized through qualitative assessment to find prominent trends, performance patterns, and critical research gaps. The qualitative assessment was concerning prediction of traffic congestion using traditional methods, a topological machine learning (PH)-based approach and deep learning. The key findings include the dominance of deep learning architectures, ensemble hybrid models, and PH approaches. From the observed patterns, the integration of deep learning architectures with PH or ML techniques can improve the accuracy of prediction efficiency. Table 11 presents the synthesis outcomes derived from the reviewed studies. The synthesis focused on the relationships between model characteristics, datasets, and predictive performance of various models. The studies also reviewed the emerging future directions in traffic congestion prediction and provide important implications on the suitability of persistent homology (PH), deep learning architectures (CNN, GNN and GCNN), ensemble hybrid approaches and data strategies in modeling urban traffic prediction.

4. Results

4.1. Traditional Methods to Manage Traffic Congestion (RQ1)

Empirical evidence reveals that the current approaches fall short in capturing complex traffic data. It revealed that induce loops show weaknesses associated with poor infrastructure, inconsistent sensor coverage, and unreliable communication systems and fail to recognize the informal transport network infrastructure [73]. Traditional statistical models are challenged with limited sample sizes, measurement errors, non-independent observations, and missing data values [9,74,75,76]. The empirical evidence has also revealed that most current approaches rely on single-model approaches that fail to understand the complexity of the spatial structure of data [10,11] and also suffer from unreliable modeling of correct traffic congestion prediction. It is important to note that Table 12 only presents qualitative insight into the prediction methods. Therefore, thoughtfulness should be applied when interpreting the reported findings in this table, and any reported dissimilarities may influence the results.

4.2. Implementation and Application Context Challenges (RQ2)

Empirical evidence reveals that the prediction of traffic congestion in urban roads has been impeded by several challenges. The following is a comparative analysis between traditional methods and hybrid models. First, the traditional methods do not extract meaningful spatiotemporal features from a dataset [77] as well as do not give a full representation of traffic complexity [7,8], while the hybrid models do. Second, there is inadequate quality of traffic data, which results in missing important data values. This is the result of a lack of proper infrastructure to collect traffic data [1,78]. Hybrid models are robust in terms of computing power and can be linked to an online data source. Third, there is limited integration of real-time data streams [79] to traditional models in low-resource environments. This can hinder the effectiveness of monitoring and enhancing learning outcomes. A real-time data stream can be integrated with a hybrid model to resolve the issue of inadequate and poor data quality because of its high computation power. Fourth, most traditional methods have been evaluated within their single-city scenario and do not guarantee enough evidence to perform better in an environment of different topologies, driving behavior patterns, and data regimes [1,80]. Hybrid models are cross-functional models that can detect geometric features in a limited data environment and outperform several baseline models. The results have proved how hybrid models can outperform traditional models. It has also been revealed that a lack of experts in advanced analytics and topology can restrict the adoption of TDA-based models [79]. The challenges addressed in this section have made prediction of traffic congestion difficult, unreliable, and inaccurate and have impeded the development of predictive models for urban road environments.

4.3. Predicting Traffic Congestion Using PH with Ensemble Stacking Algorithms (RQ3)

This survey is a combined, comparative, and forward-looking analysis of current literature. The empirical evidence provided insightful information on how to solve the problem of urban traffic congestion by integrating multiple frameworks. The comparative analysis of multiple frameworks on the performance, strength, and weakness of the frameworks provided in the literature helped to figure out the best way of modeling urban traffic congestion prediction. The integration of multiple techniques can robustly improve the accuracy of the prediction of traffic congestion [28]. This survey proposes to integrate a stacking ensemble with a persistent homology technique, particularly using Ripser, to capture both local temporal dynamics and global structural patterns in urban traffic networks, which could not be found by traditional ML models. Integrating PH with ensemble techniques aids the development of context-sensitive models for roads in urban environments. The ensemble technique may offer high predictive power by learning temporal patterns and congestion cycles computed by PH [4, 5], and PH may have the ability to explore hidden patterns in the dataset.

5. Conclusions

This systematic analysis provided in-depth and insightful findings on the weaknesses of the current approaches and how these weaknesses are affecting the accuracy of the prediction of traffic congestion. It also revealed distinct performance patterns across the architectural approaches. It found that traditional methods are inconsistent and inaccurate, and they overlook factors like weather, topologies, and complex spatiotemporal dependencies [10,11]. We also found that, regardless of the superiority of the stacking ensemble method, PH and deep learning architectures, independently, none of these algorithms can provide a complete solution. An ensemble technique may offer high predictive power by learning temporal patterns and congestion cycles but increases computational cost. Deep learning architecture excels at extracting dynamic predictive patterns but are assumed to lack interpretability and generalizability. While PH offers richer structural insights and robustness to noise, it struggles with direct predictive implementation. Despite the limitations of these architectures, this survey advocates the integration of PH with either stacking ensemble methods or with a deep learning framework to model prediction of urban traffic congestion. This integration may improve the model’s ability for structural interpretability as well as enhance the model’s predictive ability [44]. Despite the limitations of the reviewed architectures, this survey decides to integrate multiple frameworks to model prediction of traffic congestion in urban roads:
  • First, we proposed architectural innovations that can use novel technologies to enable possible deployment of a predictive model to effectively bridge the accuracy efficiency gap in prediction of traffic congestion [79]. The proposed solution should offer more advantages in terms of transparency, strong prediction, and ease of operation, making it a valuable tool in scenarios where model interpretability is critical, particularly in prediction of traffic congestion.
  • Second, the proposed solution should use representational power, enabling it to capture complex spatiotemporal dependencies and nonlinear interactions in the high-dimensional traffic datasets, specifically on urban roads, an aspect which was often overlooked by traditional methods. The novelty of integrating multiple frameworks defines the strength, originality and methodological rigor.
  • Third, although the transition from theoretical validation to real-world deployment is still a critical challenge, the model demonstrates practical feasibility and effectiveness.
  • Lastly, we recommend that future research focus not only on performance metrics or methods but also on the model’s explainability, transferability, adaptability across heterogeneous road environments as well as computational cost.

Author Contributions

Conceptualization, S.F.N.; methodology, S.F.N.; validation, O.P.K.; formal analysis, R.H.; writing—original draft preparation, S.F.N.; writing—review and editing, S.F.N., O.P.K. and R.H.; and supervision, O.P.K. and R.H.; project administration, O.P.K. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding authors.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
DLDeep Learning
TDATopological Data Analysis
MLMachine Learning
PHPersistent Homology
SLRSystematic Literature Review
SVMSupport Vector Machine
LRLogistic Regression
RFRadom Forest
GPSGlobal Positioning Systems
GNNGraph Neural Network
GCNNGraph Convolutional Neural Network
CNNConvolutional Neural Network
LSTMLong Short-Term Memory
ITSIntelligent Transport Systems

References

  1. Munga, J.N.; Kasongo, R. Dynamic Management of Traffic Congestion-Case Study in Developing Countries. J. Transp. Eng. 2023, 12, 41–48. [Google Scholar]
  2. Faheem, H.B.; El Shorbagy, A.M.; Gabr, M.E. Impact of Traffic Congestion on Transportation Systems: Challenges and Remediation—A Review. Mansoura Eng. J. 2024, 49, 18. [Google Scholar] [CrossRef]
  3. Subair, S.O.; Ibitoye, B.A.; Kuranga, A.T. Evaluation of Traffic Congestion in an Urban Roads: A Review. ABUAD J. Eng. Appl. Sci. 2024, 2, 1–7. [Google Scholar] [CrossRef]
  4. Duan, L.; Song, L.; Wang, W.; Jian, X.; Heijungs, R.; Chen, W.-Q. Urbanization inequality: Evidence from vehicle ownership in Chinese cities. Humanit. Soc. Sci. Commun. 2024, 11, 703. [Google Scholar] [CrossRef]
  5. Wen, T.H.; Chin, W.; Lai, P. Understanding the topological characteristic and flow complexity of urban traffic congestion. Phys. A Stat. Mech. Its Appl. 2017, 473, 166–177. [Google Scholar] [CrossRef]
  6. Zhao, J.; Liu, Z.; Lin, J. The impact of urban traffic congestion on residents’ quality of life: A case study of Beijing. Sustainability 2018, 10, 1001. [Google Scholar]
  7. Guo, X. Research on Deep Learning Models for Traffic Flow prediction. Appl. Comput. Eng. 2024, 111, 87–96. [Google Scholar] [CrossRef]
  8. Tripathi, N.; Sharma, B. Evaluation of a Probabilistic Framework for Traffic Volume Forecasting Using Deep Learning and Traditional Models. Int. J. Exp. Res. Rev. (IJERR) 2024, 45, 237–250. [Google Scholar] [CrossRef]
  9. Ojo, O.; Blessing, K. Statistical challenges and solutions in multidisciplinary clinical research: Bridging the gap between. World J. Biol. Pharm. Health Sci. 2024, 19, 246–258. [Google Scholar]
  10. Xue, M. Comprehensive Approaches to Traffic Flow Prediction. Appl. Comput. Eng. 2024, 111, 60–65. [Google Scholar] [CrossRef]
  11. Samonte, M.J.; Balan, G.A.F.; Gaviño, P.P.; Monasterial, J.A.S.; Reforsado, R.A.C.; Samonte, D.C. Deep Learning in Traffic Flow Control and Prediction for Traffic Management. In Proceedings of the International Conference on Industrial Engineering and Operations Management, Istanbul, Turkey, 23–26 June 2023. [Google Scholar]
  12. Azizi, F.; Zhang, W.; Malik, A.; Shen, Z. Exploring the Potential of Crowdsourced Traffic Data for Improved Traffic Predictions: A Big Data Approach. In Research Square; Springer Nature: Durham, NC, USA, 2024. [Google Scholar]
  13. Ravishanker, N.; Chen, R. Topological Data Analysis (TDA) for Time Series. arXiv 2019, arXiv:1909.10604. [Google Scholar] [CrossRef]
  14. Cornell, F. Using topological autoencoders as a filtering function for global and local topology. arXiv 2020, arXiv:2012.03383. [Google Scholar]
  15. Patil, M.; Ahmed, Q.; Midlam-Mohler, S. Urban Traffic Forecasting with Integrated Travel Time and Data. In Proceedings of the 2024 IEEE 27th International Conference on Intelligent Transportation Systems (ITSC), Edmonton, AB, Canada, 24–27 September 2024. [Google Scholar]
  16. Feng, M.; Porter, M.A. Spatial applications of topological data analysis: Cities, Snowflakes, random structures, and spinning under the influence. Phys. Rev. Res. 2020, 2, 033426. [Google Scholar] [CrossRef]
  17. Nair, A. opological Methods in Data Analysis: Applications in Machine Learning. In Modern Dynamics: Mathematical Progressions; Modern Dynamics: Gurugram, India, 2024; Volume 1. [Google Scholar]
  18. Kumar, T.G.; Babu, G.S.; Murthy, D. Topological data analysis: Theory, methods, and practical applications. Int. J. Comput. Program. Database Manag. 2025, 6, 28–36. [Google Scholar] [CrossRef]
  19. Otter, N.; Porter, M.; Tillmann, U.; Grindrod, P.; Harrington, H.A. A roadmap for the computation of persistent homology. EPJ Data Sci. 2017, 6, 17. [Google Scholar] [CrossRef] [PubMed]
  20. Corcoran, P.; Deng, B. Regularization of Persistent Homology Gradient computation. arXiv 2020, arXiv:2011.05804. [Google Scholar] [CrossRef]
  21. Edelsbrunner, H.; Harer, J. Computational Topology: An Introduction; American Mathematical Society: Providence, RI, USA, 2010. [Google Scholar]
  22. Pun, C.S.; Lee, S.X.; Xia, K. Persistent-homology-based machine learning: A survey and a comparative study. Artif. Intell. Rev. 2022, 55, 5169–5213. [Google Scholar] [CrossRef]
  23. Pan, Y.A.; Hu, X.; Zhuo, X.S. A fundamental diagram-consistent fluid queue model for dynamic throughput under heavy traffic congestion. Transp. Res. Part C Emerg. Technol. 2026, 184, 105533. [Google Scholar]
  24. Rocks, J.W.; Liu, A.J.; Katifori, E. A revealing structure-functionn relationships in functional flow networks via persistent homology. Phys. Rev. Res. 2020, 2, 033234. [Google Scholar] [CrossRef]
  25. Ghorbanchian, R.; Restrepo, J.G.; Torres, J.J.; Bianconi, G.; Bianconi, G. Higher order simplical synchronization of coupled topolocal signals. Commun. Phys. 2021, 4, 120. [Google Scholar] [CrossRef]
  26. Aggarwal, M.; Periwal, V. Tight basis cycles representatives for persistent homology of large dataset. PLoS Comput. Biol. 2023, 19, e1010341. [Google Scholar]
  27. Turkeš, R.; Montúfar, G. On the Effectiveness of Persistent Homology. arXiv 2022, arXiv:2206.10551. [Google Scholar]
  28. Turner, K. Rips fitration for quasimetric spaces and asymmetric function with stability results. Algebr. Geom. Topol. 2019, 19, 1135–1170. [Google Scholar] [CrossRef]
  29. Kaji, S.; Sudo, T.; Ahara, K. Cubical Ripser: Software for computing persistent homology of image and volume data. arXiv 2020, arXiv:2005.12692. [Google Scholar] [CrossRef]
  30. Bauer, U. Ripser: Efficient computation of Vietoris–Rips persistence barcodes. J. Appl. Comput. Topol. 2021, 5, 391–423. [Google Scholar] [CrossRef]
  31. Nguyen, V.T.; Pham, D.A.; Le, A.T.; Peter, J.; Gust, G. Persistent Homology-induced Graph Ensembles for Time Series Regressions. arXiv 2025, arXiv:2503.14240. [Google Scholar] [CrossRef]
  32. Daniel, C.B.; Saravanan, S.; Mathew, S. GIS Based Road Connectivity Evaluation Using Graph Theory. In Transportation Research: Proceedings of CTRG 2017; Springer: Singapore, 2019; Volume 45, pp. 213–226. [Google Scholar]
  33. Povaliaev, N.D.; Krylatov, A.Y. Methods of cluster analysis of road networks for bottlenecks detection and traffc optimization. T-Comm 2025, 19, 34–40. [Google Scholar] [CrossRef]
  34. Košanin, M.; Macek, N. A Clustering-Based Approach to Detecting Critical Traffic Road. Axioms 2023, 12, 509. [Google Scholar] [CrossRef]
  35. Luo, J.; Zhang, Q. Subdivision of Urban Traffic Area Based on the Combination of Static Zoning and Dynamic Zoning. Discret. Dyn. Nat. Soc. 2021, 2021, 9954267. [Google Scholar] [CrossRef]
  36. Andrade-Girón, D.C.; Sandivar-Rosas, J.; Marin-Rodriguez, W.J. Comparison of Ensemble and Meta-Ensemble Models for Early Risk Prediction of Acute Myocardial Infarction. Informatics 2025, 12, 109. [Google Scholar] [CrossRef]
  37. Tavana, P.; Akraminia, M.; Koochari, A.; Bagherifard, A. An efficient ensemble method for detecting spinal curvature type using deep transfer learning and soft voting classifier. Expert Syst. Appl. 2023, 213, 119290. [Google Scholar] [CrossRef]
  38. Artin, J.; Valizadeh, A.; Ahmadi, M.; Kumar, S.A.; Sharifi, A. Presentation of a Novel Method for Prediction of Traffic with Climate Condition Based on Ensemble Learning of Neural Architecture Search (NAS) and Linear Regression. Complexity 2021, 2021, 8500572. [Google Scholar] [CrossRef]
  39. He, R.; Xu, Y. Overfitting Identification in Machine Learning Models with the Person-Fit Indicator; IEEE: New York, NY, USA, 2023; pp. 520–524. [Google Scholar]
  40. Zhu, Z. Systematic Optimization of Overfitting Problem in Machine Learning. Highlights Sci. Eng. Technol. 2024, 111, 353–359. [Google Scholar] [CrossRef]
  41. Barreñada, L.; Dhiman, P.; Timmerman, D.; Boulesteix, A.; Calster, B.V. Understanding overfitting in random forest for probability estimation: A visualization and simulation study. BMC Diagn. Progn. Res. 2024, 8, 6–14. [Google Scholar] [CrossRef]
  42. Mattei, P.-A.; Garreau, D. Are Ensembles Getting Better All the Time? J. Mach. Learn. Res. 2025, 26, 1–46. [Google Scholar]
  43. Gul, G.; Korejo, I.A.; Hakro, D.N.; Alqahtani, H.; Abbasi, A.; Babar, M.; Rahbi, O.A.; Ali, N.I. Machine Learning and Ensemble Methods for Cardiovascular Disease Prediction: A Systematic Review of Approaches, Performance Trends, and Research Challenges. Computers 2026, 15, 25. [Google Scholar] [CrossRef]
  44. Wang, H.; Ma, Z.; Qi, W.; Zhang, N.; Zhuang, H. A Research Review of the Stacking Classification Model. In Proceedings of the 2024 6th International Conference on Machine Learning, Big Data and Business Intelligence (MLBDBI), Hangzhou, China, 1–3 November 2024. [Google Scholar]
  45. Taiwo, E.O.; Ogunsanwo, G.O.; Alaba, O.B.; Ogu, V.I. A Comparative Study of Ensemble Methods for Predicting Road Traffic Congestion. Int. J. Traffic Transp. Eng. (IJTTE) 2021, 11, 1013–1027. [Google Scholar]
  46. Taiwo, E.O.; Ogunsanwo, G.O.; Alaba, O.B.; Ogunbanwo, A.S. Traffic Congestion Prediction using Supervised Machine Learning Algorithms. TASUED J. Pure Appl. Sci. 2023, 2, 110–116. [Google Scholar]
  47. Indah, D.; Mwakalonge, J.; Comert, G.; Siuhi, S.; Masau, H.; Osei, E.; Omulokoli, P.; Sulle, M.; Ruganuza, D.; Gyimah, N.K. Topological data analysis for driver behavior classification driven by vehicle trajectory data. J. Transp. Res. Board 2025, 21, 100719. [Google Scholar] [CrossRef]
  48. Wu, J.; Zhang, K.; Deng, K.; Chen, G.; Li, X. A machine learning approach for traffic congestion prediction in developing countries. Transp. Res. Part C Emerg. Technol. 2019, 99, 200–215. [Google Scholar]
  49. Song, W. Data Analysis and Congestion Prediction Model Intelligent Transportation Systems. J. Prog. Eng. Phiysical Sci. 2024, 3, 1–8. [Google Scholar] [CrossRef]
  50. Carmody, D.; Sowers, R. Topological Analysis of traffic pace via persistent homology. J. Phys. Complex. 2021, 2, 025007. [Google Scholar] [CrossRef]
  51. Huang, N.; Wu, Y. Unveiling activity-travel pattern through topological data analysis. arXiv 2024, arXiv:2406.16742. [Google Scholar] [CrossRef]
  52. Chazal, F.; Michel, B. An introduction to Topological Data Analysis: Fundamental and practical aspects for data scientists. Front. Artif. Intell. 2021, 4, 667963. [Google Scholar] [CrossRef] [PubMed]
  53. Qi, Y.; Cheng, Z. Research on Traffic Congestion Forecast Based on Deep Learning. Information 2023, 14, 108. [Google Scholar] [CrossRef]
  54. Kumar, K.D.; Anitha, M.L.; Veena, M.N. Real-Time Bengaluru City Traffic Congestion Prediction Using Deep Learning Models. Int. J. Transp. Dev. Integr. 2025, 9, 619–628. [Google Scholar] [CrossRef]
  55. Wang, J.; Chen, R.; He, Z. Traffic speed prediction for urban expressways using a deep learning approach. Transp. Res. Part C Emerg. Technol. 2019, 105, 372–385. [Google Scholar] [CrossRef]
  56. Lin, L.; Li, W.; Zhu, L. Data-Driven Graph Filter-Based Graph Convolutional Neural Network Approach for Network-Level Multi-Step Traffic Prediction. Sustainability 2022, 14, 16701. [Google Scholar] [CrossRef]
  57. Diao, Z.; Wang, X.; Zhang, D.; Liu, Y.; Xie, K.; He, S. Dynamic Spatial-Temporal Graph Convolutional Neural Networks for Traffic Forecasting. In Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence (AAAI-19), Honolulu, HI, USA, 27 January–1 February 2019. [Google Scholar]
  58. Zheng, G.; Chai, W.K.; Zhang, J.; Katos, V. Graph Convolution Neural Network and Transformer based Traffic Prediction. In Knowledge Based Systems; Elsevier: Amsterdam, The Netherlands, 2023. [Google Scholar]
  59. Yan, F.; Wang, J.; Zhang, Y. Traffic Flow Prediction Based on Pivotal Graph Convolutional Network and Transformer. In Proceedings of the 2025 6th International Conference on Computer Information and Big Data Applications (CIBDA 2025), Wuhan, China, 14–16 March 2025. [Google Scholar]
  60. Patil, M.; Ahmed, Q.; Midlam-Mohler, S. Travel Time and Weather-Aware Traffic Forecasting in a Conformal Graph Neural Network Framewor. IEEE Trans. Intell. Transp. Syst. 2025, 26, 21734–21744. [Google Scholar] [CrossRef]
  61. Neyipapula, B.S. Rsearch Square; Springer Nature: Durham, NC, USA, 2023. [Google Scholar]
  62. Ragiri, R.; Raza, Z. Traffic Congestion Prediction using graph convolutional Networks. Int. Res. J. Adv. Eng. Manag. 2024, 2, 2117–2122. [Google Scholar]
  63. Ma, G. Using Topological Data Analysis to Process Time-series Data: A Persistent Homology way. J. Phys. Conf. Ser. 2020, 1550, 032082. [Google Scholar] [CrossRef]
  64. Buffelli, D.; Soleymani, F.; Rieck, B. Clique PH: Higher-Order Information for Graph Neural Networks through Persistent Homology on Clique Graphs. arXiv 2024, arXiv:2409.08217. [Google Scholar] [CrossRef]
  65. Wu, W.; Tang, L.; Zhao, Z.; Teo, C.-P. Enhancing binary classification: A new stacking method via leveraging computational geometry. arXiv 2024, arXiv:2410.22722. [Google Scholar] [CrossRef]
  66. Attipoe, E.K.; Yussiff, A.-S.; Asante-Mensah, M.G.; Tetteh, E.D. An ensemble learning approach for diabetes prediction using the stacking method. Comput. Sci. Inf. Technol. 2025, 6, 102–111. [Google Scholar] [CrossRef]
  67. Proskura, P.; Zaytsev, A. Effective training-time stacking for ensembling of deep neural networks. In Proceedings of the 2022 5th International Conference on Artificial Intelligence and Pattern Recognition, Xiamen, China, 23–25 September 2022. [Google Scholar]
  68. Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Moher, D. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef]
  69. Shea, B.J.; Reeves, B.; Wells, G.A.; Thuku, M. AMSTAR 2: A Critical appraisal tool for systematic reviews that include radomise or non-randomised studies of healthcare interventions, or both. BMJ 2017, 358, j4008. [Google Scholar]
  70. Guo, S.; Qian, X. Optimal Drive-By Sensing in Urban Road Networks with Large-Scale Ridesourcing Vehicles. IEEE Trans. Intell. Transp. Syst. 2024, 25, 14389–14400. [Google Scholar] [CrossRef]
  71. Lavelle-Hill, R.; Smith, G.; Murayama, K. Bridging Traditional Statistics and Machine Learning Approaches in Psychology: Navigating Small Samples, Measurement Error, Non-independent Observations and Missing Data. Adv. Methods Pract. Psychol. Sci. 2025, 8, 25152459251345696. [Google Scholar]
  72. Lavelle-Hill, R.; Smith, G.; Murayama, K. Machine Learning Meets Traditional Statistical Methods in Pysychology: Challenges and Future Directions; OSF Preprints: Charlottesville, VA, USA, 2023. [Google Scholar]
  73. Yanpei, C.; Archana, G. Challenges and Opportunities for Managing Data Systems Using Statistical Models. IEEE Data Eng. Bull. 2011, 34, 53–60. [Google Scholar]
  74. Wasserman, L. A Topological data analysis. In Annual Review of Statistics and Its Application Topological Data Analysis; Annual Reviews: San Mateo, CA, USA, 2018; Volume 5, pp. 501–532. [Google Scholar]
  75. Solodkij, A.; Gorev, A. System Approach to Elimination of Traffic Jams in Large Cities in Russia. Int. J. Traffic Transp. Eng. 2013, 23, 1112–1117. [Google Scholar]
  76. Suryadevara, G.; Pachipulusu, P. Integrating Real-Time Data Streams. In Advances in Computational Intelligence and Robotics Book Series; IGI Global: Hershey, PA, USA, 2025; pp. 67–90. [Google Scholar]
  77. Bochenina1, K.; Agriesti, S.; Roncoli, C.; Ruotsalainen, L. From Urban Data to City-ScaleModels: A Review of Traffic Simulation Case Studies. IET Intell. Transp. Syst. 2025, 19, e70021. [Google Scholar] [CrossRef]
  78. Murtagh, F. Data Science Foundations: Geometry and Topology of Complex Hierarchic Systems and Big Data Analytics; CRC Press: Boca Raton, FL, USA, 2017. [Google Scholar]
  79. Wu, Y.; Shindnes, G.; Karve, V.; Yager, D.; Work, D.B.; Chakraborty, A.; Sowers, R.B. Congestion Barcodes:Exploring the topology of Urban congestion using persistent homology. In Proceedings of the IEEE 20th International Conference on Intelligent Transportation Systems 2017, Yokohama, Japan, 16–19 October 2017. [Google Scholar]
  80. Pereira, C.M.; de Mello, R. Persistent homology for time series and spatial data clustering. Expert Syst. Appl. 2015, 42, 6026–6038. [Google Scholar] [CrossRef]
Figure 1. Persistent homology.
Figure 1. Persistent homology.
Computers 15 00370 g001
Figure 2. Ensemble stacking framework [46,47].
Figure 2. Ensemble stacking framework [46,47].
Computers 15 00370 g002
Figure 3. Proposed PH with stacking ensemble method.
Figure 3. Proposed PH with stacking ensemble method.
Computers 15 00370 g003
Figure 4. Complete PRISMA diagram. * Denotes databases that were searched between January 2014, and December 2026 included ACM Digital Library, IEEE Xplore, ScienceDirect, Springer Nature, MDPI and Google Scholar (see in Table 4). ** Denotes records that were excluded after title and abstract screening because they were not addressing urban traffic congestion predictions, did not consider use of TDA or machine learning techniques, or records were outside the study period.
Figure 4. Complete PRISMA diagram. * Denotes databases that were searched between January 2014, and December 2026 included ACM Digital Library, IEEE Xplore, ScienceDirect, Springer Nature, MDPI and Google Scholar (see in Table 4). ** Denotes records that were excluded after title and abstract screening because they were not addressing urban traffic congestion predictions, did not consider use of TDA or machine learning techniques, or records were outside the study period.
Computers 15 00370 g004
Table 1. Quantitative results for the traditional methods.
Table 1. Quantitative results for the traditional methods.
Model TypeDatasetPrediction TargetRMSEMAESources
Traditional statistical model Delhi-NCR UTF and CMP Traffic Speed25.418.9[8]
Crowdsourced modelUTF and CMPTraffic Speed18.112.1[12]
GPS-based modelUTF and CMPTraffic Speed21.314.8[12]
Social media modelUTF and CMPTraffic Speed19.7 13.5[12]
Table 2. Experimental set up for DL, TDA and ML frameworks.
Table 2. Experimental set up for DL, TDA and ML frameworks.
MethodAlgorithms/FrameworkDatasetInput
Features
Training
Data
Hyper-Parameter TuningImplementationEvaluation
Metric
Prediction
Target
DLGCNNUrban:
Traffic flow
Traffic speedTrain: 80%
Test: 20%
Grid search TensorFlow 2.21.0 RMSE
Accuracy
Future speed
TDAPH
(Ripser 0.6.15)
Urban:
Traffic flow
Traffic speedTrain: 80%
Test: 20%
Grid searchPyTorch (CUDA 13.2)/
TensorFlow 2.21.0
RMSE
Accuracy
Persistence diagrams
MLEnsemble (DT, RF, SVM, LR)Urban:
Traffic flow
Traffic speedTrain: 80%
Test: 20%
Grid searchPyTorch (CUDA 13.2)/
TensorFlow 2.21.0
RMSE,
Accuracy
Congestion levels
Table 3. Material used in the study.
Table 3. Material used in the study.
ItemOrganization/ManufacturerCityState/ProvinceCountry
AMSTAR 2McMaster University-led groupHamiltonONCanada
PRISMA 2020PRISMA Group Ottawa/OxfordON/OxfordshireCanada/UK
PyTorchMeta → PyTorch FoundationMenlo ParkCAUSA
TensorFlowGoogleMountain ViewCAUSA
CUDANVIDIASanta ClaraCAUSA
Ripser (PH)TUM (Ulrich Bauer)MunichBYGermany
Ensemble MethodsStatistics communityBerkeleyCAUSA
GCNNAcademic Research CommunityAmsterdamNorth HollandNetherlands
StackingClassifierScikit-learnParisÎle-de-FranceFrance
TomTom dataset (Online)TomTom N.V.AmsterdamNHNetherlands
Table 4. Search strategy and data source.
Table 4. Search strategy and data source.
DatabasesRecordsSearch Keywords Used
ScienceDirect900“TDA tools” OR “Persistent homology algorithms”, “Benefits” OR “Challenges of TDA techniques”, “Machine learning” OR “Ensemble approach”, “Target type” OR “Prediction”,
ACM500“ML models” OR “Traditional methods”, “Benefits” OR” Challenges”, “Traffic Congestion” OR “Prediction horizon”, “Hidden patterns” OR “Fix adjacency matrices”
Springer
Nature
59“Current Approaches Challenges OR “Prediction of Traffic congestion”
“Traffic Prediction models” OR “Real-Time monitoring”, “Local data source” OR “Real-Time data Stream”, “non-Euclidean traffic data” OR “exploiting spatial dependencies”
IEEE Xplore302“Intelligent Transport Systems” OR “Urban Road Networks”, “Ensemble learning” OR “Predictive model”, “Stacking strategies” OR “Boosting, Bagging” OR “Accuracy speed measurement”
MDPI36“Ensemble approach” OR “Traditional ML models”, “Benefits” OR “Challenges”
Google Scholar20“Traffic management infrastructure” OR “Machine learning models”
“Urban roads” OR “Urban environment roads”, “GCNNs models” OR “deep learning models”
Table 5. Inclusion criteria.
Table 5. Inclusion criteria.
CategoryCriteriaApplication
Inclusion Criteria
I1Published date between January 2014 and February 2026Applied during database search
I2Peer-reviewed journals, conference proceedings in EnglishApplied to all 1857 initial records
I3Focus on traffic congestion prediction of urban roads in urban roads Excluded 142 studies from databases and website search
I4Explicit discussion of machine learning and topological data analysis toolsExcluded 60 studies from database and website search
Table 6. Exclusion criteria.
Table 6. Exclusion criteria.
CategoryCriteriaApplication
Exclusion Criteria
E1Review articles, surveys, theses, or
non-peer-reviewed works
Excluded during initial screening
E2Models without using topological data analysis or machine learning techniquesExcluded 5 studies during full-text
review
E3Not addressing challenges current approach Excluded 26 studies during full-text review
E4Not addressing using benefits TDA and ML techniquesExcluded 6 studies during full-text review
Table 7. Quality assessment result.
Table 7. Quality assessment result.
Quality DimensionAssessment CriteriaStudies Meeting
Criteria (n = 66)
Percentage
Experimental
Rigor
Clear methods for predicting traffic congestion, clear TDA and ML prediction techniques, focus of road environment setting, adequate dataset; a total of 52 (80%) studies out of 66 met selection criteria, and 20% of the studies did not meet the criteria of experimental rigor. 5280
PerformancePerformance metrics description4670
DatasetDetailed traffic dataset and feature information5990
Comparative AnalysisComparison of PH, CNN, GCNN models 2538
Limitations DiscussionClear acknowledgment of study limitations4365
Table 8. Study selection process outcomes.
Table 8. Study selection process outcomes.
Selection PhaseRecordsExclusion Reasons
Initial Identification1857 (6 databases)(800) Duplicate removal
After Duplicate Removal1057(600) records irrelevant (350), Review article (150), non-English (100)
Title/Abstract Screening457150 (Record not retrieved)
Full-Text Assessment307241 records excluded irrelevant scope, Older than 2014 (62), Not conducted in urban roads (50), Not TDA or ML (50), Not Prediction of congestion (40), Not Focus urban roads (39)
Final Inclusion66 studiesNone
Table 9. Extracted data distribution from 66 studies.
Table 9. Extracted data distribution from 66 studies.
Data CategoryStudies Completeness RateKey Findings
Architecture Details66100%TDA (PH) (16 studies) and ML (ensemble (15 studies)), Hybrid model (TDA and ML based model (10 studies)), Deep learning (GCNN and CNN (25 studies))
Traffic Prediction5990%Urban traffic flow prediction (39 studies)
Evaluation Metrics4670%Average accuracy rate score: 90.8
Dataset Information5990%Traffic congestion
Prediction Methods5685%Traditional ML model (26 studies)
Table 10. Data extraction protocol.
Table 10. Data extraction protocol.
Extraction CategorySpecific Data PointsExtraction Methods
Model ArchitectureHybrid model PH with ensemble method, GPS and real-time traffic data source (online data source).Direct extraction from method sections
Prediction MethodsUrban roads traffic flow prediction, category type: peak-hours traffic flow (low, moderate and congestion), non-peak-hours traffic flow (low, moderate and congestion) and unexpected (road works and accidents), short-time prediction of congestion between 15 min and 60 min (one hour).Direct extraction from method sections
Performance
Metrics
RMSE, MAE, confusion matrix (accuracy, precision, recall, and F1-score) and cross-validation approach (avoid overfitting).Numerical extraction from results sections
Dataset DescriptionUrban roads real-time traffic dataset, topological features (covered): different roads, various road features, different time points of the day) and dataset size.Systematic
categorization
Table 11. Synthesis outcomes.
Table 11. Synthesis outcomes.
Synthesis FocusAnalysis ApproachIdentified Pattern
AccuracyCorrelation analysis between identification of hidden features and accuracy predictionIntegration of PH in a model improves performance and accuracy prediction, captures local temporal dynamics and global structural patterns. PH alone cannot fully model temporal dependencies necessary for real-time prediction like deep learning models. Deep learning (CNN, GNN and GCNN) models require predefined assumptions.
Congestion predictionPerformance analysisTraffic authorities and road users can receive help from
prediction of traffic congestion and make travel decisions.
Architecture efficiencyTraditional methods vs. hybrid model analysisIntegrating PH with ensemble or deep learning frameworks may be an efficient way of modeling urban traffic prediction.
DatasetReal-time dataset Vs. local datasetReal-time datasets may improve accuracy of prediction and address data shortage and infrastructure gap.
Table 12. Qualitative insight of the prediction methods.
Table 12. Qualitative insight of the prediction methods.
Method TypeTemporal ModelingSpatial ModelingComplexityAccuracy
Traditional Statistical ModelYesNoLowModerate
ML ModelPartialNoMediumModerate–High
GCNN ModelYesYesVery HighVery High
Ensemble with PHYesYesVery HighVery High
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Nyalugwe, S.F.; Kogeda, O.P.; Hans, R. Traffic Congestion Prediction Algorithms in Urban Environments: A Survey. Computers 2026, 15, 370. https://doi.org/10.3390/computers15060370

AMA Style

Nyalugwe SF, Kogeda OP, Hans R. Traffic Congestion Prediction Algorithms in Urban Environments: A Survey. Computers. 2026; 15(6):370. https://doi.org/10.3390/computers15060370

Chicago/Turabian Style

Nyalugwe, Symon Fumu, Okuthe P. Kogeda, and Robert Hans. 2026. "Traffic Congestion Prediction Algorithms in Urban Environments: A Survey" Computers 15, no. 6: 370. https://doi.org/10.3390/computers15060370

APA Style

Nyalugwe, S. F., Kogeda, O. P., & Hans, R. (2026). Traffic Congestion Prediction Algorithms in Urban Environments: A Survey. Computers, 15(6), 370. https://doi.org/10.3390/computers15060370

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop