Artificial intelligence (AI) and machine learning (ML) are increasingly applied to road traffic congestion prediction, but heterogeneous outcomes, models, horizons, and evaluation practices limit comparability. The objective of this study was to synthesize methods, applications, validation, explainability, and reproducibility in AI/ML-based road traffic congestion prediction and forecasting. Following PRISMA 2020, Scopus, Web of Science Core Collection, and IEEE Xplore were searched through 5 July 2026 for English-language journal articles and full conference papers published from 2000 to 2026. Two external reviewers independently screened 734 unique records and assessed the retrieved full texts, while the author resolved disagreements against the predefined eligibility criteria. Study characteristics, prediction tasks, congestion indicators, model families, metrics, explainability, validation, and data/code availability were synthesized descriptively and narratively. Of 1131 records identified, 397 duplicates were removed and 734 records were screened. Full-text retrieval was sought for 339 reports; 195 could not be retrieved, 144 were assessed for eligibility, and 129 were included. Congestion level was the main prediction task, while traffic flow and speed were the most frequent indicators. Heterogeneity and the absence of verified numerical performance values precluded meta-analysis or model ranking. Explainability was limited, and external validation, transferability, and reproducibility were insufficiently documented. Progress requires standardized outcomes, transparent validation, reproducible workflows, explainable models, and independent testing across networks and cities.
Road traffic congestion remains one of the most persistent challenges affecting contemporary urban transportation systems. Rapid urbanization, increasing travel demand, growing motorization, and the limited capacity of existing road infrastructure have intensified the frequency and severity of congestion in many cities. Beyond increasing travel times, recurrent and non-recurrent congestion can reduce network reliability, increase fuel consumption and emissions, constrain economic productivity, and negatively affect the quality of urban life. These impacts have made the timely identification and forecasting of congestion a central priority for intelligent transportation systems and data-informed mobility management. Consequently, transportation agencies increasingly require predictive tools capable of anticipating critical traffic conditions before network performance deteriorates substantially [
1,
2].
The growing availability of traffic sensors, connected vehicles, geographic information, mobile devices, cameras, weather records, and other digital data sources has expanded the possibilities for congestion prediction. Traditional statistical models remain useful for representing relatively stable and interpretable traffic patterns, but they may have difficulty capturing the nonlinear, dynamic, and interdependent behavior of complex road networks. Artificial intelligence and machine learning methods offer alternative mechanisms for learning relationships among traffic flow, speed, occupancy, weather, incidents, temporal patterns, and network characteristics. In particular, conventional machine learning algorithms, artificial neural networks, convolutional and recurrent architectures, graph-based models, reinforcement learning, transformers, and hybrid approaches have progressively diversified the methodological landscape. This development has positioned congestion prediction at the intersection of transportation engineering, urban computing, data science, and intelligent decision support [
2].
Despite these advances, the available evidence remains fragmented across prediction tasks, operational definitions of congestion, data sources, spatial scales, forecasting horizons, model architectures, and evaluation strategies. Some studies predict congestion occurrence or severity directly, whereas others estimate traffic flow, speed, density, occupancy, or travel time as proxies for future congestion conditions. Model performance is also reported through heterogeneous classification and regression metrics, often using different datasets, prediction horizons, targets, validation schemes, and evaluation subsets. In addition, concerns regarding interpretability, external validation, transferability, computational requirements, data availability, and code accessibility complicate the assessment of whether technically accurate models can be reproduced or deployed in contexts other than those in which they were developed. Therefore, comparing algorithms exclusively through isolated accuracy measures may produce misleading conclusions when the underlying prediction problems and evaluation conditions are not equivalent [
2,
3].
The importance of this study lies in its attempt to move beyond a model-centered description of the literature and towards an evidence-based assessment of how congestion prediction studies are designed, validated, interpreted, and reported. Such a synthesis can help researchers select modelling approaches that are consistent with the target variable, data structure, spatial scale, and intended forecasting horizon rather than selecting algorithms solely because of their popularity. It can also support practitioners and transportation authorities in evaluating whether reported models provide sufficient transparency, external validity, and operational relevance for decision-support applications. Moreover, examining data and code availability contributes to identifying barriers to reproducibility and independent verification within this rapidly developing field. The review is thus relevant not only for measuring methodological progress but also for determining whether that progress can support reliable, explainable, and transferable congestion-management solutions.
Previous reviews have examined artificial intelligence applications in traffic congestion and traffic-state prediction from different methodological perspectives. Attioui and Lahby [
4] conducted a PRISMA-based systematic review covering the 2010–2024 period and included 115 studies from 9695 initially identified records, while a subsequent review by the same authors examined 100 peer-reviewed publications published between 2014 and 2024. Other contributions have focused on AI-based congestion detection, broader traffic prediction methods, or the scientometric evolution of traffic forecasting research. However, these reviews differ substantially in their definitions, analytical units and methodological objectives. Some address direct congestion forecasting, whereas others examine congestion detection or use traffic flow, speed and travel time as proxy outcomes.
Table 1 summarizes the scope and methodological coverage of the most relevant previous reviews. Collectively, these studies provide valuable taxonomies of algorithms, data sources and application scenarios. Nevertheless, the available evidence does not show a consistent review-level assessment of prediction horizons, external spatial and temporal validation, cross-city transferability, model explainability, data and code availability, and computational reproducibility. Moreover, bibliometric records, systematically included studies and studies discussed in narrative surveys represent different units of analysis and therefore should not be compared as equivalent evidence volumes.
Against this background, the present systematic review focuses specifically on AI-, machine learning- and deep learning-based road traffic congestion prediction. Beyond identifying model families and performance metrics, it evaluates the operational and methodological maturity of the evidence by examining prediction horizons, validation design, transferability, explainability, data and code accessibility, and reproducibility. This scope enables a distinction between models that perform well within a particular dataset and models supported by evidence of generalizability, transparency and potential real-world deployment.
Accordingly, this systematic review aims to synthesize and critically examine the application of artificial intelligence and machine learning methods to road traffic congestion prediction and forecasting. Specifically, it characterizes the geographical and methodological distribution of the evidence, congestion definitions and indicators, prediction tasks, input data, model families, forecasting horizons, performance metrics, explainability techniques, external validation practices, and data and code availability. The review followed the PRISMA 2020 reporting framework and applied a structured process of record identification, screening, full-text eligibility assessment, study linkage, data extraction, and human verification [
10]. Extracted evidence was organized into study-characteristic, model, performance, and quality or risk-of-bias datasets, followed by descriptive and narrative synthesis; quantitative performance comparisons were restricted to results with explicit numerical values and sufficiently comparable outcomes, units, datasets, horizons, targets, and evaluation subsets. Through this approach, the study provides a structured account of the current evidence while distinguishing between the frequency with which methods are reported and the strength, comparability, reproducibility, and operational relevance of their reported performance.
This entry is adapted from the peer-reviewed paper https://doi.org/10.3390/encyclopedia6090205