Next Article in Journal
Sustainable Regional Development Under Demographic Transition: Labor Market Integration and Export Quality Enhancement in the Beijing-Tianjin-Hebei Region
Next Article in Special Issue
Exploring Factors Conditioning Urban Cyclist Road Safety Under a Macro-Level Approach: The Spanish Municipalities’ Case Study
Previous Article in Journal
Analysis of Spatiotemporal Variability and Drivers of Soil Moisture in the Ziwuling Region
Previous Article in Special Issue
Credible Variable Speed Limits for Improving Road Safety: A Case Study Based on Italian Two-Lane Rural Roads
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Towards Sustainable Road Safety: Feature-Level Interpretation of Injury Severity in Poland (2015–2024) Using SHAP and XGBoost

by
Artur Budzyński
1,* and
Andrzej Czerepicki
2
1
Institute of Quality Science and Product Management, Krakow University of Economics, 31-510 Krakow, Poland
2
Faculty of Transport, Warsaw University of Technology, 00-661 Warszawa, Poland
*
Author to whom correspondence should be addressed.
Sustainability 2025, 17(17), 8026; https://doi.org/10.3390/su17178026
Submission received: 9 July 2025 / Revised: 23 August 2025 / Accepted: 4 September 2025 / Published: 5 September 2025
(This article belongs to the Special Issue New Trends in Sustainable Transportation)

Abstract

This study investigates the severity of injuries sustained by over seven million participants involved in road traffic incidents in Poland between 2015 and 2024, with a view to supporting sustainable mobility and the United Nations Sustainable Development Goals. Road safety is a crucial dimension of sustainable development, directly linked to public health, urban liveability, and the socio-economic costs of transportation systems. Using a harmonised participant-level dataset, this research identifies key demographic, behavioural, and environmental factors associated with injury outcomes. A novel five-level injury severity variable was developed by integrating inconsistent records on fatalities and injuries. Descriptive analyses revealed clear seasonal and weekly patterns, as well as substantial differences by participant type and driving licence status. Pedestrians and passengers faced the highest risk, with fatality rates more than five times higher than those of drivers. An XGBoost classifier was trained to predict injury severity, and SHAP analysis was applied to interpret the model’s outputs at the feature level. Participant role emerged as the most important predictor, followed by driving licence status, vehicle type, lighting conditions, and road geometry. These findings provide actionable insights for sustainable road safety interventions, including stronger protection for pedestrians and passengers, stricter enforcement against unlicensed driving, and infrastructural improvements such as better lighting and safer road design. By combining machine learning with interpretability tools, this study offers an analytical framework that can inform evidence-based policies aimed at reducing crash-related harm and advancing sustainable transport development.

1. Introduction

1.1. Background and Motivation

Road traffic crashes remain an important public health and transportation policy concern, contributing significantly to the global burden of disease and long-term disability [1]. While fatal outcomes are widely reported, non-fatal injuries account for the majority of healthcare burden, insurance processing, and long-term societal impact. In Poland, as in other countries, there has been a decrease in fatalities in recent years, yet the broader picture of crash-related harm remains complex. Slight and serious injuries often dominate the total number of road casualties, especially among vulnerable users such as pedestrians and cyclists [2].
The severity of injuries is shaped by various individual and contextual factors, including the role of the participant, use of safety equipment, and driving licence status. However, many national statistics and prior studies aggregate these effects or omit participant-level resolution. This limits the ability to explore patterns across roles, age groups, and temporal trends. Moreover, predictive modelling of injury severity still faces key challenges such as missing data, ambiguous injury definitions, and the need for consistent target variable construction [3].
This study utilises a large-scale dataset covering road crash participants in Poland between 2015 and 2024. The objective is to construct a harmonised injury severity variable and analyse its association with demographic, behavioural, and environmental features. By combining descriptive analysis with machine learning models, this study aims to improve empirical understanding of injury outcomes and support more data-informed road safety strategies.

1.2. Literature Overview

1.2.1. Empirical Studies on Road Traffic Injuries

In the Polish legal and statistical framework, the terminology of road events is formally defined by the Regulation No. 31 of the Chief of Police of 22 October 2015 on the methods and forms of keeping statistics of road traffic events (Dz. Urz. KGP 2015, item 70) [4]. The regulation introduces the following definitions:
  • Road accident (in Polish: ‘wypadek drogowy’)—a road event that results in at least one fatality or injury;
  • Road collision (in Polish: ‘kolizja drogowa’)—a road event that causes only material damage without casualties;
  • Road event (in Polish: ‘zdarzenie drogowe’)—either a road accident or a road collision;
  • Fatal victim—a person who died at the scene or within 30 days due to injuries sustained;
  • Seriously injured person—an individual suffering from permanent disability or a health disorder lasting more than seven days;
  • Slightly injured person—an individual with other, less severe injuries. In this article, the terms accident and collision are used consistently in line with these official definitions, ensuring full alignment with the national SEWiK database and the legal framework governing road safety statistics in Poland.
Road traffic events, including accidents and collisions, constitute a complex and multifactorial phenomenon with considerable implications for public health, transportation policy, and urban safety. Their analysis has been a recurring theme in empirical research, motivated by the scale of the problem and the potential for targeted prevention. Depending on the context, traffic crashes may be studied from epidemiological, behavioural, legal, or infrastructural perspectives, each emphasising different determinants and outcomes [5].
Empirical studies of road traffic events typically aim to understand the relationships between various crash features, such as time, location, vehicle type, and participant characteristics, and its consequences. These consequences may include injury severity, fatalities, liability, or secondary effects such as traffic congestion or legal outcomes. In practice, the granularity and availability of data often shape the scope of analysis, with many studies relying on administrative records from police or insurance systems.
Participant-related attributes, such as age, sex, driving licence status, and behaviour (e.g., alcohol use, seatbelt usage), have proven to be key explanatory variables in injury analysis [3]. In addition, environmental and temporal factors—road geometry, weather conditions, day of the week, and time of day—are often incorporated to improve explanatory power. While vehicle-specific information (e.g., production year, technical condition, or safety equipment) is frequently available, it is sometimes underutilised due to missing data or inconsistencies in reporting [6].
Despite the methodological variety, several limitations persist. Studies focus only on drivers, excluding passengers, pedestrians, or other vulnerable road users. Injury classifications are often inconsistent across jurisdictions or datasets, hindering comparability. Finally, seasonality and behavioural dynamics are still rarely addressed, even though time-based trends may substantially influence crash risk and injury outcomes.
Road traffic injuries remain a significant public health issue, with over 1.3 million deaths and approximately 50 million injuries annually, according to the World Health Organisation [7]. More than 90% of global traffic fatalities occur in low- and middle-income countries, which account for only about 60% of the world’s vehicle fleet [8]. Vulnerable road users—such as pedestrians, cyclists, and motorcyclists—are disproportionately affected by severe or fatal outcomes [9].
Numerous evaluations have shown that road infrastructure measures, such as public lighting, can substantially reduce crash risks, particularly in urban areas and for pedestrian safety [10]. In Iran, a study identified the highest fatality rates among individuals over 65 years of age in both “type 2” traffic accidents and non-traffic incidents [11]. While the mortality rate from type 2 traffic events remained relatively stable between 2013 and 2018, deaths from non-traffic incidents increased over time. These findings highlight the importance of region-specific analyses and targeted safety interventions, particularly in low- and middle-income countries, to address distinct patterns of road-related mortality within the broader global context. The Safe System approach emphasises that human error is inevitable and therefore road infrastructure and policies should be designed to minimise the consequences of crashes. This paradigm shift, already implemented in countries such as Sweden and the Netherlands, has led to substantial reductions in road fatalities. In recent years, the approach has been increasingly promoted by international organisations (OECD, WHO) as the foundation of national road safety strategies [12].
Earlier studies on crash injury severity predominantly relied on econometric approaches. Ordered and random parameters probit models have been applied to capture the influence of weather and lighting conditions on injury outcomes, explicitly accounting for unobserved heterogeneity across crashes [13]. Similarly, Tobit specifications and zero-inflated count models were employed to jointly analyse frequency and severity measures, offering a flexible statistical framework for correlated outcomes [14]. More recently, comparative research has combined statistical techniques with machine learning, demonstrating both the relative strengths of traditional econometric methods and the added value of data-driven classifiers in predicting injury severity [15]. These contributions form an important methodological background that complements the growing body of machine learning-based studies discussed later in this Section.
While traditional research has predominantly relied on statistical and epidemiological approaches to identify determinants of injury severity, recent advances in machine learning have opened new possibilities for modelling complex, nonlinear relationships and enhancing predictive accuracy. Accordingly, the next Subsection reviews applications of machine learning in road safety research.

1.2.2. Machine Learning Studies

Meanwhile, data infrastructure and analytics advances have enabled more precise modelling of injury severity, allowing researchers to evaluate both predictive performance and underlying causal mechanisms. As Mannering et al. emphasise, the emergence of big data in highway safety research introduces fundamental trade-offs between prediction and causal inference [16].
Recent years have seen a growing interest in applying machine learning (ML) techniques to road safety research, motivated by the increasing availability of detailed traffic datasets and the limitations of traditional statistical models in capturing complex, nonlinear relationships. The use of machine learning in crash injury severity modelling has expanded rapidly, with systematic reviews showing consistent improvements over traditional statistical techniques [17]. ML approaches allow for flexible modelling of high-dimensional data and have been applied to various predictive tasks, such as crash occurrence, accident type, or injury severity [18].
Previous studies have focused mainly on modelling the occurrence or type of road events using historical crash data. For example, in earlier research covering Polish accident data from 2015 to 2021, a Gradient Boosting model was used to classify events as accidents or collisions, achieving a high accuracy of 93.8% and identifying key predictors such as event type and road characteristics [19]. Similarly, a more recent study extended the analysis to include a wider set of environmental and infrastructural features, using a Decision Tree classifier to distinguish between collisions and accidents in a dataset comprising over three million records [20]. A complementary example from the Polish transport context is the work of Budzyński and Cieśla, who applied machine learning methods to predict truck parking occupancy at highway rest areas using geospatial and temporal features [21].
The present study builds on these efforts but differs significantly in scope and target definition. Instead of predicting the type of crash, it focuses on modelling the consequences for individual participants, namely injury severity. This shift introduces new methodological challenges such as target variable harmonisation, high class imbalance, and the need for participant-level feature integration. Applications of machine learning in transport analytics have also extended beyond safety-related modelling. For instance, ML models have successfully forecast freight rates in the Polish road transport market, leveraging spatial and temporal features to capture economic dynamics [22]. Furthermore, the temporal scope has been extended to 2024, allowing the analysis of recent trends and long-term seasonality effects [23]. A parallel line of research has explored real-time crash detection using video surveillance, where end-to-end deep learning models such as DETR have demonstrated high responsiveness and reduced false alarm rates [24].

1.3. Research Gap and Objective

Despite extensive research on road traffic safety, most prior studies have focused on the types and causes of crashes rather than their consequences for individual participants. In particular, the analysis of injury severity at the participant level remains underdeveloped. A key limitation in the literature is the lack of harmonised target variables and consistent definitions of injury outcomes. Moreover, existing research often relies on short timeframes or aggregated data, limiting the detection of long-term trends and participant-specific risk factors.
This study addresses these limitations by analysing a large-scale dataset covering over seven million crash participants recorded nationwide during the last decade. A new five-level injury severity variable was constructed through rule-based harmonisation of inconsistent fatality and injury records. The main objective is to assess how demographic, behavioural, and environmental characteristics influence injury severity. The analysis combines descriptive statistics with machine learning models and SHAP interpretation techniques to identify key predictors and support evidence-based traffic safety strategies.

2. Materials and Methods

2.1. Data Description and Preparation

The source data were provided by the Traffic Department of the General Police Headquarters of Poland in the form of annual CSV files, covering crash events, vehicles, participants, and assigned causes from 2015 to 2024. The dataset is not publicly available through the “SEWiK” system and was made accessible to the author upon special request for scientific research purposes. To construct a participant-level dataset, these files were merged using relational identifiers (Event ID, Vehicle ID, and Participant ID). Due to performance limitations of the CSV format—especially in terms of memory usage and read speed—the data were converted to the Apache Parquet format [25]. This columnar storage format supports selective column reads, efficient compression, and metadata-aware filtering, which have been empirically validated in benchmark studies comparing Parquet with other columnar formats [26]. All variable names were translated from Polish to English to ensure clarity and consistency during further processing. The process is shown in the diagram in Figure 1.
The source data is organised annually from 2015 to 2024. For each year, four structured files were available: incidents, vehicles, participants, and assigned causes. These were linked using unique identifiers (Incident ID, Vehicle ID, Participant ID), enabling reconstruction of complete event scenarios. As shown in Figure 2, each participant may or may not be linked to a vehicle (e.g., pedestrians), and some, though not all, were identified as perpetrators of the incident. Causes were assigned only to participants marked as responsible. This relational structure supported integration into a participant-level analytical dataset.
While CSV format ensures broad compatibility and ease of use, it presents significant limitations when working with large-scale datasets, especially regarding memory consumption and processing speed [27]. This issue became particularly relevant given the size of the integrated dataset, which consisted of over seven million participant-level records. The preparation steps addressing these challenges are described in the following sections.
The final dataset used for modelling consists of 7,137,217 rows and 55 columns, integrating information from crash events, involved vehicles, participants, and assigned causes. The data includes a wide range of variable types, with 48 columns stored as categorical or string-type (object), five columns as floating-point numerical values (float64), and two columns as integer identifiers (int64). The dataset occupies approximately 2.9 GB of memory in its in-memory representation.
This structure reflects real-world crash data’s inherent complexity and heterogeneity, which often integrates multiple interdependent outcomes across demographic, vehicular, environmental, and infrastructural dimensions [2]. A detailed description of the available features is provided in the following Sections.
The dataset includes a detailed description of each crash event’s circumstances and spatial characteristics. Event-related attributes span multiple categories, including administrative location (e.g., county, municipality, voivodeship, street name, and house number), road identifiers (road number and kilometre mark), and infrastructure layout (e.g., intersection, roundabout, or road geometry). Additional variables capture environmental and situational context, such as weather conditions, surface type, road lighting, speed limits, and the presence of traffic signals.
Geolocation is represented through GPS coordinates (x and y), although stored as strings, and should be interpreted carefully due to potential formatting inconsistencies across yearly files. The dataset also includes categorical indicators such as crash type (collision vs. accident), road type, and surface condition. Finally, crash-specific metadata such as the time and date of the event, as well as possible contributing factors marked as “Other causes,” are also present.
Including administrative, infrastructural, and environmental attributes is a common and recommended practice in traffic accident modelling to capture the multifactorial nature of crash occurrences better [5].
The dataset also contains detailed information about vehicles involved in the recorded crash events. Vehicle attributes include the brand and model, the production year, and the recorded condition of the vehicle at the time of the incident. An important safety-related feature is whether the vehicle was equipped with protective equipment such as airbags. Additional attributes specify if the vehicle belongs to a special category (e.g., emergency vehicles, agricultural machinery) and provide the date of the last technical inspection. However, the latter is recorded as a string and may require careful preprocessing for time-sensitive analyses. The year of production is available as a numeric variable. It can serve as an important proxy for vehicle age and safety technology level, which are known to influence crash severity outcomes [3]. A complete overview of the vehicle-related attributes is presented in Figure 3.
Participant-related attributes in the dataset provide essential insights into the demographic and behavioural characteristics of individuals involved in crash events. Key features include the type of participant (e.g., driver, passenger, pedestrian) and their demographic profile, such as gender and date of birth. From the date of birth, additional age-related metrics can be derived to analyse risk profiles across different age groups. Driving-related features include information on licence possession status and the number of years of driving experience, which can be critical in modelling driver behaviour and accident responsibility. Behavioural indicators include whether the participant was using safety equipment (e.g., seatbelt, helmet), position within the vehicle, and whether they were the identified cause of the crash. Furthermore, the dataset captures the severity of participant injuries through two separate variables: one indicating whether the participant was killed (either at the scene or within 30 days) and another specifying the level of injury sustained (slight or serious). Behavioural infractions, such as driving under the influence, and supplementary descriptive information about participant actions are recorded.
The detailed representation of participant characteristics is crucial for understanding human factors in traffic accidents and for accurately modelling injury severity outcomes [6].
To enable consistent modelling of crash outcomes, a new target variable named Injury severity was constructed based on two original fields: Killed (including deaths at the scene and within 30 days) and Injured (including serious and slight injuries). These two fields were combined according to a fixed priority: fatal outcomes were ranked above injuries, and injury types were distinguished where no fatality was recorded. The final variable consists of five mutually exclusive categories:
  • Killed on the scene;
  • Killed within 30 days;
  • Seriously injured;
  • Slightly injured;
  • Uninjured.
This construction ensured a unique classification for each participant based on available source data.
The descriptive analysis examined the structure and temporal variation in the target variable, Injury severity. As noted in [28], descriptive statistics are a vital first step in organising and summarising large-scale datasets. Frequency distributions and proportions were computed to assess class imbalance.
Temporal aggregations were performed for calendar year, month, and day of the week to capture potential seasonal and weekly patterns. In addition, group-wise breakdowns of Injury severity were calculated by Participant type and, for drivers, by driving licence status, following the practical use of descriptive summaries in applied organisational contexts [29].

2.2. Tools and Environment

Data processing, feature engineering, and modelling stages were conducted using Python (version 3.13). Python offers a comprehensive ecosystem for structured data analysis and is widely used in empirical traffic safety research due to its flexibility, reproducibility, and open-source nature.
The pandas (version 2.2.2) [30] and numpy (version 1.26.4) [31] libraries were used for tabular data manipulation. Data visualisation was implemented using matplotlib (version 3.9.2) [32] and seaborn (version 0.13.2) [33], enabling exploratory plotting and publication-quality figures.
Machine learning tasks were performed using XGBoost (version 2.1.3) [34]. Recent benchmarking studies have shown that gradient boosting classifiers outperform traditional methods in predicting crash severity due to their robustness to multicollinearity and ability to model nonlinear effects. These libraries were selected due to their efficiency in handling large-scale, high-cardinality datasets and built-in support for categorical features and missing value handling.
Code execution and experimentation were performed using Jupyter Notebooks (version 4.2.5) in a local environment. Reproducibility was ensured through script versioning and consistent random seeds.

3. Results

3.1. Exploratory Data Analysis

3.1.1. Empty Values

The initial results of the data inspection process are presented in Figure 4, which summarises the percentage of missing values across all features in the dataset. Over 15 variables exhibit more than 70% missingness, including no safety equipment, special vehicles, vehicle condition, and protective equipment presence. These high-missingness features pertain primarily to participant and vehicle attributes and reflect either systemic underreporting or structural absence of information in police records. Since the original variable names were provided in Polish, all feature names have been translated into English for clarity and broader accessibility of the findings. Identifying and quantifying missing data was a necessary first step in preparing the dataset for modelling. In the next stage, each affected variable will be analysed individually, and a dedicated imputation strategy will be proposed to maximise data retention while minimising bias.
The feature ‘No safety equipment’ contained 99.84% missing values. However, this absence of data does not indicate a data quality problem. Based on expert knowledge and domain-specific practices, it is known that police officers typically leave this field blank when no violation is observed. In this case, missing values imply that the participant used all required safety equipment. Therefore, all missing entries were replaced with the value ‘none’. Table 1. presents the final distribution of this feature. As expected, the vast majority of entries (99.84%) are labelled as ‘none’, while a small minority represent recorded violations, such as ‘seatbelt’ (0.10%), ‘helmet’ (0.05%), or ‘child seat’ (0.01%). This confirms that safety equipment misuse is rare among registered participants and that the feature, after appropriate interpretation, can be used in further analysis.
The feature ‘Special vehicles’ contained 99.77% missing values. This missingness does not indicate a quality issue but reflects the registration practice: when a vehicle does not fall into any special category, such as an emergency vehicle or one transporting hazardous materials, the recording officer typically leaves the field blank. Therefore, all missing entries were imputed with the value ‘none’, indicating the absence of any special classification. The final distribution is shown in Table 2. As expected, the ‘none’ category is dominant (99.77%), while other types such as ‘emergency vehicle-police’ (0.10%), ‘emergency vehicle-other’ (0.07%), and ‘hazardous materials vehicle’ (0.04%) occur infrequently.
The variable ‘Vehicle condition’ contained 7,110,186 missing values, representing 99.62% of all entries. Due to the extremely high proportion of missing data and the absence of a dedicated category for a technically sound vehicle, all missing values were imputed with no visible defect. This solution reflects the data recording practice, where officers typically note nothing in the field when no defects are observed. Despite over 300 unique textual combinations of defects, further cleaning and standardisation of these entries will be addressed in the next stage of the analysis.
The variable ‘Protective equipment present’ exhibited a high proportion of missing values, with 99.32% of entries left empty. An initial analysis revealed three distinct categories: T, N, and None. Based on domain interpretation, T was mapped to ‘yes’, and N to ‘no’. The missing values were interpreted as situations where the police officers did not report any irregularities and were thus imputed with the label ‘unknown’. This imputation strategy is summarised in Table 3, which shows the post-processed value distribution.
The dataset contains a high proportion of missing values (99.08%) in the ‘Intersection name’ feature. Upon inspection, it became apparent that these missing values most likely indicate that the incident did not occur at an intersection or roundabout. Rather than treating these cases as data errors, we interpreted them as meaningful absences of an intersection context. Therefore, all missing entries were imputed with the value ‘not at intersection’. Given the extremely high cardinality of unique intersection names (5237), I opted not to include a full value distribution table in this paper.
The variable ‘Seat position’, indicating the occupant’s place in the vehicle, exhibited 98.62% missing data. Only 0.75% of the records included the value ‘P’, and 0.63% the value ‘T’. Due to the ambiguous nature of these abbreviations and the lack of documentation explaining the distinction, the missing entries were imputed with the placeholder value ‘unknown’. This conservative approach preserves the original semantics and avoids introducing assumptions without further clarification.
The variable ‘Under influence’ is missing in 96.1% of cases, with most values labelled as Unknown. Among known cases, the most frequent entry is Not tested (2.23%), followed by confirmed presence of Alcohol (1.63%). Other values, such as Other substance or combinations, represent marginal shares. Given the high proportion of unknowns, the variable was imputed with Unknown, while its potential as a predictor will be evaluated later. The detailed distribution is presented in Table 4.
The variable ‘Distance to intersection’ exhibits a substantial degree of missing data, with 6,749,903 absent entries, corresponding to 94.57% of all records. This may indicate that, in most cases, the incidents occurred far from intersections, and the information was therefore not recorded by officers. Given this uncertainty, the variable was provisionally imputed using the placeholder Unknown, without further assumptions at this stage.
The feature ‘Other causes’ includes many missing values, with 91.83% of entries labelled as Unknown. As shown in Table 5, the most frequently recorded causes were Objects or animals on the road (3.72%), Undetermined (1.59%), and Road condition (1.25%). Less frequent entries, such as Technical malfunction, Passenger’s fault, or Improper road work protection, accounted for below 1% of cases each. This feature’s diversity and lack of structure highlight the need for future standardisation and grouping of infrequent values.
The variable ‘Last technical inspection’ exhibited a high proportion of missing values, with approximately 78.6% of entries unavailable in the original dataset. Due to the limited availability of this information, no direct imputation was applied at this stage. However, as part of the preprocessing, the non-missing values were standardised: the original strings containing both date and time (DD.MM.YYYY HH:MM) were parsed and reformatted into a simplified YYYY-MM-DD string format. This change was implemented to reduce memory usage and eliminate unnecessary timestamp precision, as all recorded times were fixed at midnight (00:00).
The variable ‘Intersection’ with road contained nearly 84% missing values and over 47,000 unique non-null entries, most of which appeared to be technical identifiers. Due to the lack of documentation and the high sparsity, missing values were left unchanged, and no imputation was applied at this stage.
The variable ‘Distance marker (KM/HM)’ exhibited a high proportion of missing values (77.5%) and over 7600 unique non-null entries. Due to the absence of a clear schema or usage context for this numeric string field, missing values were left unchanged, and no imputation was applied. The variable Production year contained missing values in approximately 75.9% of cases. Due to the scale of missingness and the absence of a reliable basis for reconstruction, no imputation was applied. The variable Intersection with street exhibited a high rate of missing data, with approximately 75% of entries unavailable. Given the free-text nature of the values and the absence of structured mapping, missing values were not imputed. The variable Intersection with street exhibited a high rate of missing data, with approximately 75% of entries unavailable. Given the free-text nature of the values and the absence of structured mapping, missing values were not imputed.
Table 6 presents the distribution of the variable Intersection, which classifies the type of road intersection. Most records (73.1%) contain no information in this field. Among the available data, most cases involved intersections with a priority road (22.7%), followed by roundabouts (3.5%) and equal-priority intersections (0.7%).
Table 7 shows the distribution of the variable ‘Additional info’, which captures supplementary behavioural or situational factors associated with crash participants. Nearly half of the records (49.4%) contain no entry in this field. Among the non-missing data, the most frequent entries relate to failure to yield right of way (9.4%), failure to maintain safe distance (8.7%), and excessive speed relative to conditions (8.3%). All values were translated into English for clarity and consistency. Missing values were left unchanged.
The variable ‘House’ number was missing in approximately 43.7% of cases. Due to the nature of the data and the lack of contextual address information, no imputation was applied.
Table 8 presents the distribution of the variable ‘Road geometry’, which describes the physical layout of the road at the crash location. Approximately 28.2% of records lacked this information. Among the specified values, the most common were straight sections, curves, and combined features such as descent or crest. All labels were translated into English to ensure consistency across the dataset.
The variable ‘Street’ was missing in approximately 22.6% of records. Due to its free-text nature and the lack of a controlled vocabulary, no imputation was applied. The variable ‘Years of driving experience’ was missing in approximately 15% of cases. Given the numerical nature of the variable and the absence of reliable proxies, no imputation was performed. The variable ‘Road number’ was missing in approximately 12.4% of records. No imputation was applied as the values represent technical identifiers and no reliable structure or mapping was available. The variable ‘Event ID (KSIP)’ was missing in approximately 9.4% of cases. As it serves only as a technical identifier and is irrelevant for modelling, missing values were left unchanged, and no imputation was performed. The variable ‘Town/city’ was missing in approximately 4.7% of records. Due to the highly granular and non-standardised nature of the entries, no imputation was applied. The variable ‘Vehicle model’ was missing in approximately 3.3% of records. Due to the large number of unique free-text entries and implicit missing indicators such as “UNSPECIFIED”, no imputation or standardisation was applied. The variables ‘GPS x’ and ‘GPS y’ were both missing in exactly 3.2% of cases, always jointly. The values are stored in degree-minute format; no imputation or conversion was applied at this stage. The variable ‘Vehicle brand’ was missing in approximately 3.1% of records. Due to the large number of free-text entries and the lack of a reliable mapping or reference list, no imputation or standardisation was applied. The variable ‘Vehicle type’ was missing in approximately 1.8% of records. As the values were already grouped into interpretable vehicle categories and the missing rate was low, no further processing or imputation was applied. A group of variables related to road infrastructure and environmental conditions was analysed jointly due to their similar nature and minimal proportion of missing values. This group included Vehicle ID, Road markings, Speed limit, Weather conditions, Road surface condition, Crash site characteristics, Road surface, Area type, Road type, and Traffic signal. The percentage of missing data in these variables was exceptionally low. For all variables except Vehicle ID, the missing rate was below 0.01%, which was considered negligible in the dataset’s overall size. No imputation or removal was applied, and the features were retained in their original form. The variable Vehicle ID had a slightly higher missing rate, equal to 1.79%. No corrective action was taken since this field is primarily used for identification or referencing purposes and is not predictive. The variable was left unchanged to preserve its relational structure with other data tables.

3.1.2. Target Value

In the study’s next step, a new target variable was created by combining the existing features ‘Killed’ and ‘Injured’. These two fields, recorded independently in the source data, often lacked mutual exclusivity and, in rare cases, led to logically inconsistent combinations, such as the simultaneous presence of a serious injury and post-crash death within 30 days. A rule-based approach was applied to derive a single, interpretable outcome per participant to address this.
The resulting feature, ‘Injury severity’, assigns each participant to exactly one of five categories: ‘Killed on scene’, ‘Killed within 30 days’, ‘Seriously injured’, ‘Slightly injured’, or ‘Uninjured’. Fatal outcomes were given priority in classification: if the value of ‘Killed’ indicated immediate or delayed death, the injury information was ignored. Without a fatal outcome, classification was based on the content of the ‘Injured’ field. Participants with no recorded death or injury were classified as ‘Uninjured’.
The final distribution was as follows: 6,799,743 participants (95.27%) were classified as ‘Uninjured’, 215,198 (3.02%) as ‘Slightly injured’, 95,123 (1.33%) as ‘Seriously injured’, 18,905 (0.26%) as ‘Killed on scene’, and 8248 (0.12%) as ‘Killed within 30 days’. The distribution of non-zero injury outcomes is visualised in Figure 5, which clearly shows that non-fatal injuries dominate the dataset, while fatal outcomes are rare. This imbalance highlights the importance of stratified evaluation methods in subsequent predictive modelling.
Figure 6 presents annual trends in the distribution of injury severity among crash participants between 2015 and 2022. The most frequent outcome was ‘Slightly injured’, with yearly counts ranging from nearly 30,000 in 2015–2017 to around 17,000 in 2022. A similar downward trend is observable for the ‘Seriously injured’ category, which declined from over 12,000 in 2015–2016 to approximately 7500 in 2022.
The number of fatal outcomes remained relatively stable between 2015 and 2019, with minor year-to-year fluctuations. Participants classified as ‘Killed on scene’ consistently ranged between 2200 and 2400 annually before dropping to 1317 in 2022. Likewise, those marked as ‘Killed within 30 days’ decreased from over 1000 annually to 579 in 2022. A notable decline in all categories occurred in 2020, which coincides with the COVID-19 pandemic and related reductions in mobility. Although 2021 showed partial recovery in injury figures, 2022 registered the lowest values across all severity levels.
These trends are likely driven by external systemic factors such as traffic volume reductions, road usage behaviour changes, and possibly vehicle or infrastructure safety improvements. They highlight the temporal dynamics of crash consequences and reinforce the importance of accounting for time effects in predictive modelling.
Seasonal variation in crash outcomes was further analysed by aggregating all cases by calendar month, regardless of year. The resulting monthly distribution is shown in Figure 7. The number of participants sustaining injuries or fatal outcomes exhibited a clear seasonal pattern, with consistently lower counts in winter months (January–March) and higher counts in summer (June–September).
The highest number of cases in every injury severity category was recorded between May and October. The category ‘Slightly injured’ peaked in August (20,461 participants), while ‘Seriously injured’ reached its maximum in July (8956) and remained elevated during June and August. Fatalities also followed this seasonal pattern: the number of participants ‘Killed on scene’ increased from 1143 in January to 1852 in October, while ‘Killed within 30 days’ rose from 483 in February to 765 in July and remained high through autumn.
These results suggest that the overall risk of severe crash consequences rises during higher road activity, better weather, and potentially riskier behaviour such as speeding or fatigue during longer journeys. Similar seasonal patterns have been documented in the U.S., Korea, and Europe, where summer months consistently correlate with increased crash frequency and severity [35].
The observed seasonality reinforces the relevance of temporal factors in modelling crash outcomes and may inform the design of targeted preventive measures.
A further dimension of temporal analysis is presented in Figure 8, which shows the distribution of injury severity across days of the week. While weekday values remain relatively stable, a marked increase in cases is observed on Fridays and Saturdays. The peak occurs on Friday, with 32,725 participants classified as ‘Slightly injured’, 14,054 as ‘Seriously injured’, and a combined 4007 fatalities. Saturday follows closely, particularly in terms of fatal outcomes, with 3001 participants ‘Killed on scene’ and 1231 ‘Killed within 30 days’—the highest fatality count among all days.
These elevated figures at the end of the week likely reflect increased long-distance travel, social activities, and potential driver fatigue or riskier behaviour associated with weekend plans. Previous studies also highlight Friday and Saturday as high-risk periods due to increased evening mobility and alcohol-related crashes [36].
By contrast, Sunday exhibits a noticeable decline in total crash-related injuries and fatalities, suggesting lower traffic volumes or more cautious driving behaviour.
This weekday pattern highlights the need to incorporate day-of-week effects into crash modelling frameworks. It also points to potential opportunities for targeted public safety interventions, particularly focused on late-week traffic behaviour and high-risk time windows.

3.2. Model Performance and Feature Ranking

An XGBoost classifier was trained on the full dataset after removing all identifier columns and variables directly related to the outcome to evaluate the predictive utility of available features. The model was assessed using 5-fold cross-validation, resulting in a mean macro F1-score of 0.2872. This relatively low value reflects both the severe class imbalance and the complexity of the injury severity classification task. As illustrated in Figure 9, ‘Participant type’ emerged as the most important predictor, significantly outperforming all other variables. Other features with notable importance included ‘Driving licence status’, ‘Sex’, and ‘Vehicle type’. In contrast, many features demonstrated minimal or zero contribution to the model, raising questions about their data quality, variability, or actual relevance to injury severity. These variables will be examined more closely in the next analysis phase.
Figure 10 presents a heatmap showing the relationship between participant type and injury severity. The proportions are normalised within each participant category to highlight the distribution of outcomes across groups. The differences are striking: drivers are overwhelmingly uninjured (97.4%), while passengers and pedestrians exhibit a much higher proportion of injuries and fatalities. For example, only 9.1% of passengers and 52% of pedestrians remain uninjured after a crash, with 22.4% and 16.8% seriously injured, respectively. The highest relative fatality rate is observed among passengers (4.9%), followed by pedestrians (4.9%). In contrast, while more vulnerable than drivers, individuals using mobility devices still experience most uninjured outcomes (86%).
The analysis was restricted to participants identified as drivers to isolate the effect of driving licence status on injury severity. As shown in Figure 11, drivers without a valid licence experienced substantially worse outcomes than their licenced counterparts. Specifically, 1.2% of unlicensed drivers were killed on scene and 3.9% were seriously injured, compared to just 0.1% and 0.5%, respectively, among licenced drivers.

3.3. SHAP Analysis Results

The SHAP analysis (Figure 12) highlights the critical role of several variables in predicting immediate fatalities (“Killed on scene”). Participant type emerges as a leading predictor, reflecting higher vulnerability among pedestrians and passengers due to lower physical protection. Moreover, high SHAP values associated with the absence of valid driving licences underscore increased risk-taking behaviours or lack of driving proficiency among unlicensed individuals. Additional influential factors include lighting conditions and speed limits, indicating environmental and infrastructural elements significantly impact the likelihood of immediate death, consistent with existing literature that emphasises these variables as critical determinants of crash severity.
The SHAP analysis (Figure 13) for participants who died within 30 days after the crash indicates that the street location emerges as the strongest predictor, likely reflecting different post-crash emergency response capabilities depending on location. Lighting conditions and the month of occurrence also appear influential, underscoring the role of environmental factors in determining delayed fatalities.
The SHAP analysis (Figure 14) for the “Seriously injured” class emphasises vehicle type and street as dominant predictors, suggesting that vehicle characteristics and crash location significantly affect the risk of serious injury. Furthermore, driving experience and road number emerge as influential, highlighting the importance of driver skill and specific road environments in determining injury severity.
The SHAP analysis (Figure 15) for the “Slightly injured” class highlights participant type and vehicle type as key predictive factors, indicating significant differences in injury risk based on participant vulnerability and vehicle characteristics. Additionally, the street location, geographical coordinates (GPS), and years of driving experience notably influence outcomes, emphasising the combined role of spatial and individual factors.
The SHAP analysis (Figure 16) for the “Uninjured” class reveals legal resolution and participant type as the primary predictors, reflecting the legal outcomes of incidents and the varying vulnerability among road users. Additionally, vehicle type significantly influences the likelihood of participants emerging without injury, highlighting protective differences associated with different vehicles.

4. Discussion

These results confirm that Participant type is not only highly predictive (as shown in the feature importance analysis) but also strongly associated with injury severity, likely due to differing levels of protection and exposure among the groups. Similar findings were reported in Canadian data, where pedestrians and passengers consistently exhibited higher fatality and injury rates than drivers [37].
These disparities suggest a strong association between lack of a valid driving licence and increased crash severity, potentially reflecting legal ineligibility and broader differences in driving behaviour, experience, or risk exposure. Previous evaluations found that unlicensed drivers were overrepresented in fatal crashes and exhibited higher risk-taking behaviour [38].
The obtained results align with contemporary findings in traffic safety research that emphasise feature-level differences in injury outcomes. For instance, in [39], used LightGBM and SHAP on highway crash data in Pakistan and identified driver age, collision type, and time of year as pivotal predictors of injury severity. Similarly, in [40] applied XGBoost with SHAP to pedestrian crash data in South Korea, showing that built environment variables significantly affect fatality risk. These examples confirm the added value of interpretable machine learning in understanding crash dynamics. Our study builds on this by deploying SHAP to interpret an XGBoost model trained on a much larger, harmonised dataset from Poland (2015–2024), offering transparent, feature-level insights that enhance policy relevance and deployment readiness.
The relevance of environmental and infrastructural features such as lighting conditions, road type, and built environment is echoed in multiple international studies. Authors in [41] for instance, identified casualty class, time of accident, and vehicle type as top contributors to injury severity using SHAP-augmented ensemble methods. These insights reinforce our findings, particularly that lighting and road geometry were among the most influential predictors in the Polish dataset.
Beyond feature identification, our study demonstrates the potential of SHAP for uncovering local instance-level explanations, enabling a nuanced understanding of why certain conditions led to higher injury severity for specific cases. This capability aligns with the methodological shift towards explainable AI in road safety. For example, an AutoML–SHAP approach [42] applied to U.S. pedestrian crash data (2024) revealed location- and weather-specific predictors challenging blanket intervention strategies. Similarly, providing participant- and case-specific insights, our XGBoost–SHAP framework supports targeted policy-making, such as improving street lighting at specific hotspots or tailoring interventions to vulnerable user groups.
Despite the strengths of this approach, several limitations should be acknowledged. First, administrative data, while large-scale and systematically collected, are prone to underreporting of less severe incidents and lack important contextual details, such as vehicle speed, driver distraction, or psychological state. Second, many variables exhibit high rates of missingness, especially those related to vehicle condition or behavioural violations. While imputation strategies were applied, the underlying sparsity introduces uncertainty. Third, the analysis is geographically limited to Poland, which may constrain the generalizability of the findings to other national contexts with different legal frameworks, infrastructure, or reporting practices.
The results of this study provide several actionable insights for road safety interventions. First, the higher vulnerability of pedestrians and passengers compared to drivers underscores the need to prioritise protective infrastructure, such as well-lit crosswalks, traffic calming measures in urban areas, and public campaigns reinforcing seatbelt use among passengers. Second, the substantially worse outcomes for unlicensed drivers highlight the importance of enforcement and monitoring of driving licence regulations, alongside targeted educational programmes for high-risk groups. Third, the influence of lighting conditions and road geometry on severe outcomes points to infrastructural interventions, including the improvement of street illumination and redesign of hazardous intersections or curved road sections. Together, these recommendations link the empirical findings directly to potential preventive measures that can be implemented by policymakers and transport authorities.

5. Conclusions

This study presented a comprehensive analysis of injury severity among participants of road traffic events in Poland between 2015 and 2024. Based on data covering over seven million participants, a harmonised target variable was constructed, and both descriptive and predictive analyses were carried out using machine learning techniques.
The results revealed substantial variation in injury outcomes depending on participant type, driving licence status, and temporal patterns. Pedestrians and passengers were found to be at significantly higher risk of severe outcomes compared to drivers. In addition, participants without a valid driving licence showed markedly worse injury outcomes in the event of a crash.
From a sustainability perspective, these findings highlight the importance of road safety as a key element of sustainable mobility and urban liveability. Reducing road traffic injuries directly contributes to Sustainable Development Goal 3 (Good Health and Well-Being) and Goal 11 (Sustainable Cities and Communities) by lowering health burdens, improving quality of life, and enhancing the resilience of transport systems.
The proposed analytical framework and target variable construction procedure may serve as a basis for future research and practical applications in traffic risk management and prevention strategies. By providing interpretable, data-driven evidence, this study supports the design of targeted interventions—such as improved pedestrian infrastructure, enforcement against unlicensed driving, and safer road environments—that contribute not only to immediate crash prevention but also to the long-term objectives of sustainable transport development.

Author Contributions

Conceptualisation, A.B.; methodology, A.B.; software, A.B.; validation, A.C.; formal analysis, A.B.; investigation, A.B.; resources, A.B.; data curation, A.B.; writing—original draft preparation, A.B.; writing—review and editing, A.C.; visualisation, A.B. and A.C.; supervision, A.B. and A.C.; project administration, A.B.; funding acquisition, A.B. and A.C. All authors have read and agreed to the published version of the manuscript.

Funding

The publication is co-financed from the subsidy granted to the Krakow University of Economics within the Support for Publishing Activities 2025 (Wsparcie Aktywności Publikacyjnej 2025) programme. This paper is co-financed under the research grant of the Warsaw University of Technology supporting the scientific activity in the discipline of Civil Engineering, Geodesy and Transport.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

All scripts, data schemas, and analysis outputs used in this study are available at https://github.com/BudzynskiA/road_crash_severity_pl_2015_2024 (accessed on 3 September 2025).

Acknowledgments

The author would like to thank the Road Traffic Bureau of the General Police Headquarters of Poland for providing access to the data used in this study. The dataset was made available for scientific research purposes. The authors also gratefully acknowledge the editor and anonymous reviewers for their constructive comments and valuable suggestions, which have significantly contributed to improving the quality of this article.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Haagsma, J.A.; Graetz, N.; Bolliger, I.; Naghavi, M.; Higashi, H.; Mullany, E.C.; Abera, S.F.; Abraham, J.P.; Adofo, K.; Alsharif, U.; et al. The global burden of injury: Incidence, mortality, disability-adjusted life years and time trends from the Global Burden of Disease study. Inj. Prev. 2013, 22, 3–18. [Google Scholar] [CrossRef]
  2. Heydari, S.; Fu, L.; Miranda-Moreno, L.F.; Jopseph, L. Using a flexible multivariate latent class approach to model correlated outcomes: A joint analysis of pedestrian and cyclist injuries. Anal. Methods Accid. Res. 2017, 13, 16–27. [Google Scholar] [CrossRef]
  3. Elvik, R. Risk of road accident associated with the use of drugs: A systematic review and meta-analysis of evidence from epidemiological studies. Accid. Anal. Prev. 2013, 60, 254–267. [Google Scholar] [CrossRef] [PubMed]
  4. Zarządzenie Nr 31 Komendanta Głównego Policji z Dnia 22 Października 2015 r. w Sprawie Metod i form Prowadzenia Przez Policję Statystyki Zdarzeń Drogowych. 2015. Available online: https://edziennik.policja.gov.pl/legalact/2015/85/ (accessed on 3 September 2025).
  5. Mannering, F.L.; Bhat, C.R. Analytic methods in accident research: Methodological frontier and future directions. Anal. Methods Accid. Res. 2014, 1, 1–22. [Google Scholar] [CrossRef]
  6. Lord, D.; Mannering, F. The statistical analysis of crash-frequency data: A review and assessment of methodological alternatives. Transp. Res. Part A Policy Pract. 2010, 44, 291–305. [Google Scholar] [CrossRef]
  7. World Health Organization. Global Status Report on Road Safety 2018; World Health Organization: Geneva, Switzerland, 2018. Available online: https://iris.who.int/handle/10665/276462 (accessed on 8 June 2025).
  8. Ameratunga, S.; Hijar, M.; Norton, R. Road-traffic injuries: Confronting disparities to address a global-health problem. Lancet 2006, 367, 1533–1540. [Google Scholar] [CrossRef]
  9. Yang, T.; Fan, W.; Song, L. Modeling pedestrian injury severity in pedestrian-vehicle crashes considering different land use patterns: Mixed logit approach. Traffic Inj. Prev. 2023, 24, 114–120. [Google Scholar] [CrossRef]
  10. Beyer, F.R.; Ker, K. Street lighting for preventing road traffic injuries. Cochrane Database Syst. Rev. 2009, 2009, CD004728. [Google Scholar] [CrossRef]
  11. Hasani, J.; Erfanpoor, S.; Nazari, S.S.H. Investigation of Death Rate due to Type 2 Traffic Accidents and Non-Traffic Accidents in Iran during 2013–2018. Iran. J. Public Health 2022, 51, 1172–1179. [Google Scholar] [CrossRef]
  12. International Transport Forum. Zero Road Deaths and Serious Injuries: Leading a Paradigm Shift to a Safe System; OECD: Paris, France, 2016. [CrossRef]
  13. Fountas, G.; Anastasopoulos, P.C.; Kim, J. The joint effect of weather and lighting conditions on injury severities of single-vehicle accidents. Anal. Methods Accid. Res. 2020, 27, 100124. [Google Scholar] [CrossRef]
  14. Anastasopoulos, P.C.; Mannering, F.L.; Shankar, V.N.; Haddock, J.E. A random-parameters multivariate tobit model of highway accident-injury-severities. Anal. Methods Accid. Res. 2016, 11, 17–32. [Google Scholar] [CrossRef]
  15. Elhenawy, M. Comparison of Statistical and Machine-Learning Models on Road Traffic Accident Severity Classification. Computers 2022, 11, 80. [Google Scholar] [CrossRef]
  16. Mannering, F.; Bhat, C.R.; Shankar, V.; Abdel-Aty, M. Big data, traditional data and the tradeoffs between prediction and causality in highway-safety analysis. Anal. Methods Accid. Res. 2020, 25, 100113. [Google Scholar] [CrossRef]
  17. Santos, K.; Dias, J.P.; Amado, C. A literature review of machine learning algorithms for crash injury severity prediction. J. Saf. Res. 2022, 80, 254–269. [Google Scholar] [CrossRef]
  18. Mehta, K.; Jain, S.; Agarwal, A.; Bomnale, A. Road Accident Prediction Using Xgboost. In Proceedings of the 2022 International Conference on Emerging Techniques in Computational Intelligence (ICETCI), Hyderabad, India, 25–27 August 2022; pp. 50–56. [Google Scholar] [CrossRef]
  19. Budzyński, A.; Sładkowski, A.; Długosz, B.; Sikora, K.; Hasan, W.; Nikitishyn, T. The analysis of factors affecting the occurence of road incidents in Poland in 2015–2021. In Transport Problems 2023, XV INTERNATIONAL Scientific Conference, XII International Symposium of Young Researchers; Silesian University of Technology: Gliwice, Poland, 2023. [Google Scholar]
  20. Budzyński, A.; Federowicz, M.; Jabłoński, A.; Hasan, W.; Gorszanów, J.; Nikitishyn, T. A Machine Learning Approach for Predicting Road Accidents. Saf. Def. 2025, 10, 60–70. [Google Scholar] [CrossRef]
  21. Budzyński, A.; Cieśla, M. Highway Rest Area Truck Parking Occupancy Prediction Using Machine Learning: A Case Study from Poland. Infrastructures 2025, 10, 151. [Google Scholar] [CrossRef]
  22. Budzyński, A.; Cieśla, M. Application of a machine learning model for forecasting freight rate in road transport. Sci. J. Silesian Univ. Technology. Ser. Transp. 2025, 126, 23–48. [Google Scholar] [CrossRef]
  23. Wu, D.; Wang, S. Comparison of road traffic accident prediction effects based on SVR and BP neural network. In Proceedings of the 2020 IEEE International Conference on Information Technology, Big Data and Artificial Intelligence (ICIBA), Chongqing, China, 6–8 November 2020; pp. 1150–1154. [Google Scholar] [CrossRef]
  24. Muttaqin, H.T.; Rachmawati, E.; Ferdian, E. Toward End-to-End Detection for Traffic Accidents from CCTV Footage. In Proceedings of the 2024 10th International Conference on Computing and Artificial Intelligence, Bali Island, Indonesia, 26–29 April 2024; pp. 74–83. [Google Scholar] [CrossRef]
  25. Apache Parquet Format Documentation. Available online: https://parquet.apache.org/docs/ (accessed on 6 July 2025).
  26. Zeng, X.; Hui, Y.; Shen, J.; Pavlo, A.; McKinney, W.; Zhang, H. An Empirical Evaluation of Columnar Storage Formats. Proc. VLDB Endow. 2023, 17, 148–161. [Google Scholar] [CrossRef]
  27. Vohra, D. Practical Hadoop Ecosystem; Apress: Berkeley, CA, USA, 2016. [Google Scholar] [CrossRef]
  28. Kaur, P.; Stoltzfus, J.; Yellapu, V. Descriptive statistics. Int. J. Acad. Med. 2018, 4, 60. [Google Scholar] [CrossRef]
  29. Acosta, J.D.; Brooks, S. Descriptive statistics are powerful tools for organizational research practitioners. Ind. Organ. Psychol. 2021, 14, 481–485. [Google Scholar] [CrossRef]
  30. McKinney, W. Data Structures for Statistical Computing in Python. In Proceedings of the Python in Science Conference, Austin, TX, USA, 28 June–3 July 2010; pp. 56–61. [Google Scholar] [CrossRef]
  31. Harris, C.R.; Millman, K.J.; van der Walt, S.J.; Gommers, R.; Virtanen, P.; Cournapeau, D.; Wieser, E.; Taylor, J.; Berg, S.; Smith, N.J.; et al. Array programming with NumPy. Nature 2020, 585, 357–362. [Google Scholar] [CrossRef] [PubMed]
  32. Hunter, J.D. Matplotlib: A 2D Graphics Environment. Comput. Sci. Eng. 2007, 9, 90–95. [Google Scholar] [CrossRef]
  33. Waskom, M. Seaborn: Statistical data visualization. JOSS 2021, 6, 3021. [Google Scholar] [CrossRef]
  34. Chen, T.; Guestrin, C. XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 13–17 August 2016; pp. 785–794. [Google Scholar] [CrossRef]
  35. Chen, H.; Du, W.; Li, N.; Chen, G.; Zheng, X. The socioeconomic inequality in traffic-related disability among Chinese adults: The application of concentration index. Accid. Anal. Prev. 2013, 55, 101–106. [Google Scholar] [CrossRef] [PubMed]
  36. Jensen, O.C.; Laursen, L.H. Reduction of slips, trips and falls and better comfort by using new anti-slipping boots in fishing. Int. J. Inj. Control. Saf. Promot. 2011, 18, 85–87. [Google Scholar] [CrossRef]
  37. Abdel-Aty, M.A.; Abdelwahab, H.T. Predicting Injury Severity Levels in Traffic Crashes: A Modeling Comparison. J. Transp. Eng. 2004, 130, 204–210. [Google Scholar] [CrossRef]
  38. Kim, H.S.; Kim, H.J.; Son, B. Factors associated with automobile accidents and survival. Accid. Anal. Prev. 2006, 38, 981–987. [Google Scholar] [CrossRef]
  39. Dong, S.; Khattak, A.; Ullah, I.; Zhou, J.; Hussain, A. Predicting and Analyzing Road Traffic Injury Severity Using Boosting-Based Ensemble Learning Models with SHAPley Additive exPlanations. Int. J. Environ. Res. Public Health 2022, 19, 2925. [Google Scholar] [CrossRef]
  40. Chang, I.; Park, H.; Hong, E.; Lee, J.; Kwon, N. Predicting effects of built environment on fatal pedestrian accidents at location-specific level: Application of XGBoost and SHAP. Accid. Anal. Prev. 2022, 166, 106545. [Google Scholar] [CrossRef]
  41. Rifat, M.A.K.; Kabir, A.; Huq, A. An Explainable Machine Learning Approach to Traffic Accident Fatality Prediction. Procedia Comput. Sci. 2024, 246, 1905–1914. [Google Scholar] [CrossRef]
  42. Rafe, A.; Singleton, P.A. Exploring the Determinants of Pedestrian Crash Severity Using an AutoML Approach. arXiv 2024, arXiv:2406.06624. [Google Scholar] [CrossRef]
Figure 1. Overview of the data processing and analysis.
Figure 1. Overview of the data processing and analysis.
Sustainability 17 08026 g001
Figure 2. Relational structure of the crash dataset.
Figure 2. Relational structure of the crash dataset.
Sustainability 17 08026 g002
Figure 3. Vehicle data structure.
Figure 3. Vehicle data structure.
Sustainability 17 08026 g003
Figure 4. Percentage of missing values for features with >20% missingness.
Figure 4. Percentage of missing values for features with >20% missingness.
Sustainability 17 08026 g004
Figure 5. Distribution of Injury Severity (excluding Uninjured).
Figure 5. Distribution of Injury Severity (excluding Uninjured).
Sustainability 17 08026 g005
Figure 6. Annual trends in injury severity (excluding “uninjured”).
Figure 6. Annual trends in injury severity (excluding “uninjured”).
Sustainability 17 08026 g006
Figure 7. Injury severity by month (all years combined, excluding uninjured).
Figure 7. Injury severity by month (all years combined, excluding uninjured).
Sustainability 17 08026 g007
Figure 8. Injury severity by day of the week (excluding uninjured).
Figure 8. Injury severity by day of the week (excluding uninjured).
Sustainability 17 08026 g008
Figure 9. XGBoost Feature Importance (Log Scale).
Figure 9. XGBoost Feature Importance (Log Scale).
Sustainability 17 08026 g009
Figure 10. Heatmap of injury severity by participant type.
Figure 10. Heatmap of injury severity by participant type.
Sustainability 17 08026 g010
Figure 11. Injury severity distribution among drivers by licence status.
Figure 11. Injury severity distribution among drivers by licence status.
Sustainability 17 08026 g011
Figure 12. SHAP values analysis for the “Killed on scene” class.
Figure 12. SHAP values analysis for the “Killed on scene” class.
Sustainability 17 08026 g012
Figure 13. SHAP values for the “Killed within 30 days” class.
Figure 13. SHAP values for the “Killed within 30 days” class.
Sustainability 17 08026 g013
Figure 14. SHAP values for the “Seriously injured” class.
Figure 14. SHAP values for the “Seriously injured” class.
Sustainability 17 08026 g014
Figure 15. SHAP values for the “Slightly injured” class.
Figure 15. SHAP values for the “Slightly injured” class.
Sustainability 17 08026 g015
Figure 16. SHAP values for the “Uninjured” class.
Figure 16. SHAP values for the “Uninjured” class.
Sustainability 17 08026 g016
Table 1. Distribution of values in the No safety equipment feature after missing value imputation.
Table 1. Distribution of values in the No safety equipment feature after missing value imputation.
ValueCountShare [%]
none = ‘full safety equipment’7,125,82599.84
seatbelt71460.1
helmet35590.05
child seat6750.01
child seat; seatbelt120
Table 2. Distribution of values in the ‘Special vehicles’ feature after missing value imputation.
Table 2. Distribution of values in the ‘Special vehicles’ feature after missing value imputation.
ValueCountShare [%]
none = ‘non-special vehicles’7,120,44699.77
emergency vehicle-police68080.1
emergency vehicle-other47290.07
hazardous materials vehicle28280.04
right-hand drive vehicle20220.03
emergency vehicle3840.01
Table 3. Distribution of values in the ‘Protective equipment present’ variable.
Table 3. Distribution of values in the ‘Protective equipment present’ variable.
ValueCountShare [%]
unknown7,088,77899.32
yes39,6900.56
no87490.12
Table 4. Distribution of values for ‘Under influence’.
Table 4. Distribution of values for ‘Under influence’.
ValueCountShare [%]
Unknown6,858,62896.1
Not tested159,3322.23
Alcohol116,2851.63
Other substance25990.04
Alcohol; Other substance3720.01
Other substance; Not tested10
Table 5. Distribution of values in the feature ‘Other causes’.
Table 5. Distribution of values in the feature ‘Other causes’.
ValueCountShare [%]
Unknown6,553,99291.83
Objects or animals on the road265,4763.72
Undetermined113,3321.59
Road condition89,0401.25
Other64,7440.91
Passenger’s fault12,1970.17
Poor road condition10,8850.15
Technical malfunction81580.11
Loss of consciousness/driver death44110.06
Road work protection39870.06
Unintentional technical malfunction28170.04
Glare from another vehicle or sun24680.03
Vehicle fire21850.03
Driver sudden illness9600.01
Technical malfunction (since November 2015)8720.01
Improper road work protection6430.01
Traffic organisation4680.01
Loss of consciousness/death (since November 2015)2960
Malfunctioning railway barrier960
Traffic light operation870
Improper traffic organisation590
Malfunctioning barrier270
Malfunctioning traffic light170
Table 6. Distribution of values in the variable Intersection.
Table 6. Distribution of values in the variable Intersection.
ValueCountShare [%]
Nan5,217,85873.11
With priority road1,620,11222.7
Roundabout250,1163.5
Equal priority49,1310.69
Table 7. Distribution of values in the variable ‘Additional info’.
Table 7. Distribution of values in the variable ‘Additional info’.
ValueCountShare [%]
No additional info3,528,29749.44
Failure to yield right of way673,5579.44
Failure to maintain safe distance621,5388.71
Speed not adapted to traffic conditions592,7988.31
Improper reversing395,1975.54
Improper lane change256,3353.59
Improper overtaking252,3643.54
Improper passing244,5673.43
Other reasons193,8612.72
Improper turning145,4532.04
Failure to yield to pedestrian at crosswalk44,5950.62
Disregard of traffic lights35,1650.49
Fatigue or falling asleep28,2530.4
Careless entry onto the road: in front of a moving vehicle19,1800.27
Disregard of signs and other signals15,8070.22
Improper U-turn12,8830.18
Improper stopping or parking11,6640.16
Sudden braking11,1170.16
Failure to yield to pedestrian in other situations95040.13
Improper crossing of bicycle crossing81120.11
Crossing the road in prohibited area41520.06
Careless entry onto the road: from behind a vehicle or obstacle40430.06
Failure to yield to pedestrian39580.06
Failure to yield to pedestrian when turning36880.05
Entering the roadway at red light34260.05
Driving through red light30750.04
Disregard of other signals29340.04
Driving on the wrong side of the road24130.03
Improper pedestrian crossing manoeuvre24100.03
Walking on the wrong side of the road18770.03
Lying, sitting, kneeling or standing on the roadway12400.02
Passing vehicle before pedestrian crossing10580.01
Driving without required lights8900.01
Overtaking vehicle before pedestrian crossing4950.01
Improper bicycle crossing manoeuvre4180.01
Standing or lying on the roadway3970.01
Driving without required lighting2970
Stopping or reversing1990
Table 8. Distribution of values in the variable ‘Road geometry’.
Table 8. Distribution of values in the variable ‘Road geometry’.
ValueCountShare [%]
Straight section4,341,28860.83
Unknown2,012,62828.2
Curve494,8696.93
Straight section; Descent91,2731.28
Straight section; Ascent62,3880.87
Descent38,8950.54
Ascent32,9650.46
Curve; Descent32,3510.45
Curve; Ascent21,0100.29
Straight section; Crest60510.08
Curve; Crest18280.03
Crest16670.02
Straight section; Curve40
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Budzyński, A.; Czerepicki, A. Towards Sustainable Road Safety: Feature-Level Interpretation of Injury Severity in Poland (2015–2024) Using SHAP and XGBoost. Sustainability 2025, 17, 8026. https://doi.org/10.3390/su17178026

AMA Style

Budzyński A, Czerepicki A. Towards Sustainable Road Safety: Feature-Level Interpretation of Injury Severity in Poland (2015–2024) Using SHAP and XGBoost. Sustainability. 2025; 17(17):8026. https://doi.org/10.3390/su17178026

Chicago/Turabian Style

Budzyński, Artur, and Andrzej Czerepicki. 2025. "Towards Sustainable Road Safety: Feature-Level Interpretation of Injury Severity in Poland (2015–2024) Using SHAP and XGBoost" Sustainability 17, no. 17: 8026. https://doi.org/10.3390/su17178026

APA Style

Budzyński, A., & Czerepicki, A. (2025). Towards Sustainable Road Safety: Feature-Level Interpretation of Injury Severity in Poland (2015–2024) Using SHAP and XGBoost. Sustainability, 17(17), 8026. https://doi.org/10.3390/su17178026

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop