Next Article in Journal
Personalized Canine Diet Generation Using Machine Learning and Constraint Optimization
Previous Article in Journal
A Model-Driven Engineering Approach to AI-Powered Healthcare Platforms
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Equity, Function, and Data: A Review of Social and Functional Representation in AI Datasets for Traumatic Brain Injury

Department of Communication Sciences and Disorders, North Carolina Central University, Durham, NC 27707, USA
*
Author to whom correspondence should be addressed.
Informatics 2026, 13(2), 33; https://doi.org/10.3390/informatics13020033
Submission received: 22 October 2025 / Revised: 29 January 2026 / Accepted: 11 February 2026 / Published: 13 February 2026
(This article belongs to the Section Health Informatics)

Abstract

Traumatic brain injury (TBI) is a leading cause of long-term disability worldwide, and each person’s recovery looks different. Artificial intelligence (AI) offers promising tools to project individual outcomes. However, these models are impacted by the quality and inclusiveness of the dataset on which they are trained, having major implications for clinical value. This scoping review evaluated publicly available datasets that use AI modeling to predict outcomes from TBI. It examined how the literature derived from these datasets captures functional and social variables. Following PRISMA guidelines, 24 studies were identified, yielding 19 distinct datasets. While most datasets emphasized biomedical and injury severity metrics, few incorporated communication, cognition, and relevant social determinants of health. Nearly all studies included age and sex, but fewer than half reported race or ethnicity, and only a small subset integrated broader contextual indicators. Results suggest that outcome modeling continues to rely heavily on global scales, with limited use of domain-specific measurements. Another limiting factor is poor use of longitudinal measures, often not extending follow-up past the six-month post-injury time. These findings point to a need for inclusive, functionally rich, and ethically transparent data practices to aid AI systems in promoting equitable and clinically meaningful care.

1. Introduction

Traumatic brain injury continues to be a leading cause of long-term disability and early mortality for individuals of all ages [1]. Following their injury, TBI survivors may continue to be impacted by cognitive-communicative, behavioral, or emotional difficulties ranging from mild to severe. Traumatic brain injury occurs when an external mechanical force disrupts normal brain function, resulting in outcomes that extend far beyond the immediate physical injury. Cognitive and psychosocial consequences are common and may interfere with new learning, employment, and community participation. These challenges create lasting demands on individuals, families, communities, and rehabilitation systems, highlighting the need to better understand how TBI recovery unfolds over time.
Predicting an individual’s anticipated recovery progression after TBI is the first step in designing effective intervention. Clinicians rely on these predictions to set treatment priorities, create appropriate and functional goals, and anticipate barriers that could limit independence and community reintegration. Researchers likewise continue to investigate the factors that influence why two individuals with similar injuries may have very different outcomes. Reliable modeling of these patterns is a laborious process that has the potential to improve clinical decision-making and help allocate resources more efficiently.
Developments in artificial intelligence (AI) may streamline this work, allowing clinicians and researchers to advance recovery predictions with greater precision and efficiency. When using large, well-documented, comprehensive datasets, AI-based modeling tools can potentially identify complex relationships among clinical and behavioral presentations, as well as community and social-driven factors that may be difficult to determine through traditional methods alone [2]. In short, these systems have the potential to help clinicians target best practices in patient-centered care. However, the strength of such tools depends on the quality and representativeness of the data that drives them. When datasets lack diversity or are incomplete, the resulting models risk bias or poor reliability, which reduces their clinical usefulness.

1.1. Artificial Intelligence in Traumatic Brain Injury Research

Neuroinformatics research has shown that many AI datasets currently used to predict recovery from TBI remain limited in scope, often placing greater emphasis on the medical variables of the injury and secondary damages associated with that, while omitting social and functional dimensions of outcome [3,4]. Subsequently, predictive models may not currently be able to capture the wide-ranging intricacies characteristic of TBI that is often seen in real-world settings. For example, Ashrafi et al. (2025) [5], used the MIMIC-III database to train algorithms to predict ventilator-associated pneumonia in individuals with TBI. While the models performed well statistically, they failed to include individualized variables such as education, social support, or socioeconomic status, and subsequently restricted the model’s generalizability to clinical applications. Similarly, Osong et al. (2025) [6] developed a decision-tree model to predict hospital discharge outcomes for older adults with fall-related TBI. This group found that recovery was shaped by multiple factors, many of which were not categorized or represented within the most used trauma registries.
Artificial intelligence has also expanded the ability to examine biological and neurophysiological mechanisms underlying TBI recovery. Jia et al. (2025) [7] and Brazdzionis et al. (2023) [8] used AI-based analytic frameworks to investigate molecular processes related to cerebral edema and neural repair, illustrating how multimodal data, including cellular, imaging, and behavioral sources, can reveal relationships across levels of function. Similarly, Zhang et al. (2025) [9] combined radiomic and clinical variables in a ResNet-based model to predict outcomes following spinal cord injury, demonstrating that integrating imaging with functional data improved the validity of clinical applicability. These studies highlight a contrast between advances in mechanistic AI modeling and more limited attention to functionally meaningful clinical outcomes. In contrast, Jojczuk et al. (2023) [10] trained a neural network using only ICD-10 injury codes to predict mortality in more than 6000 trauma cases. Although the model achieved statistical accuracy, it offered little insight into individual variability or underlying mechanisms. These examples emphasize that prediction alone is not sufficient; meaningful clinical value depends on models that are both interpretable and inclusive.
Recent multimodal studies have demonstrated that the cause and context of injury influence distinct biological and functional pathways. Su et al. (2023) [11] and Alosco et al. (2023) [12], working within the DIAGNOSE CTE project, identified tau-deposition patterns among athletes with repeated head impacts that differed from those typically observed in general TBI populations. In a related study, Hong et al. (2024) [13] reported that intracranial pressure thresholds developed from uniform samples did not translate well to other demographic or clinical groups. These findings show that TBI is not a single condition, but a range of experiences influenced by biological, environmental, and social factors. To promote equitable care, AI models need to capture this variability if they are to support and inform equitable care.

1.2. Cognitive Justice Framework

The main challenge in using AI models to predict TBI recovery is not so much the design of the algorithms themselves but rather ensuring that the data used for training accurately represents the diverse populations and experiences of individuals affected. Kwak et al. (2024) [14] demonstrated that neighborhood- and individual-level social determinants (e.g., race, insurance status, and local vulnerability) sometimes influenced neurocritical care decisions more strongly than injury severity. When these contextual factors are missing or misclassified, AI systems risk reproducing existing inequities in healthcare delivery. Ibrahim et al. (2021) [15] referred to this pattern as health-data poverty, describing the systematic exclusion of marginalized populations from the datasets that shape technological development. Even datasets that seem demographically balanced can reflect hidden biases, such as differences in assessment procedures, inconsistent access to services based on insurance or location, and limited tracking beyond one-year post-injury.
The cognitive justice framework offers a path for addressing these gaps and biases. It recognizes that recovery and health can be understood from multiple perspectives, not only through biomedical measures [16]. In TBI research, applying a cognitive justice lens encourages the creation of datasets that capture individual and community-level differences in cognition, communication, environment, and participation. These key elements are the foundation for providing person-centered rehabilitation. Lukersmith and colleagues [17] used this approach in their evaluation of case management after severe injury, showing that interventions tailored to a person’s environment, motivation, and social support led to better long-term outcomes than administrative or standardized models. When applied to AI in TBI research, the cognitive justice framework highlights that equity depends not only on who is represented in data, but on how fully their lived experiences are reflected.
Understanding equity in this way means recognizing that recovery after TBI unfolds within social systems that extend beyond the hospital or rehabilitation setting. Carlozzi et al. (2025) [18] found that caregiver engagement and social context strongly influenced participation and long-term outcomes. Mavroudis et al. (2025) [19] and Venkatesan and Lucke-Wold (2025) [20] also showed that behavioral and emotional health, along with medical comorbidities, shaped cognitive consequences and quality of life. These results show the importance of including data that captures psychosocial and communication-related factors, as well as community-related requirements that support recovery outside of clinical settings.
Recent studies suggest that the next phase of AI-based TBI research should focus on collecting data using inclusive and transparent practices that are grounded in real-world contexts. Alsaigh et al. (2024) [21] and Hegde et al. (2024) [22] state the importance of ethical oversight and shared data standards in building responsible model development. By building AI systems that are methodologically sound with datasets that are inclusive, we move toward a more equitable approach to neurorehabilitation, in which care and research outcomes reflect the diversity of individuals living with TBI.

1.3. Purpose of the Study

The purpose of this study was to investigate the composition and scope of publicly available AI datasets used to model recovery after TBI. Accordingly, this study was guided by two research questions: (1) What functional-recovery variables are included in publicly available AI datasets used to model post-TBI outcomes? (2) To what extent do these datasets represent a range of individual backgrounds and follow-up trajectories needed to support accurate, personalized recovery modeling?

1.4. Significance and Innovation

This review makes two novel contributions to the field of health informatics. First, it provides a systematic analysis of functional outcome measures in AI datasets used in TBI research by emphasizing practical domains such as communication, cognition, and participation. These critical domains are often underrepresented in existing data, yet are central to understanding recovery and guiding researchers in designing more inclusive and representative datasets.
Second, this study applies a client-centered and cognitive-justice framework to dataset inclusivity, moving beyond simple metrics of diversity toward real representation of lived experience. This approach supports AI model development that is scientifically sound, equitable, and clinically useful in TBI management.

2. Materials and Methods

This study used the widely adopted scoping review framework developed by design as set forth by Arskey & O’Malley (2005) [23] and expanded by Levac, Colquhoun, & O’Brien (2010) [24]. A two-phase methodology was used to identify publicly available datasets used in AI research on TBI outcomes and examined how well these datasets included functional and contextual recovery variables.

2.1. Phase 1: Systematic Literature Review

A systematic literature review identified AI studies that modeled TBI recovery using publicly available datasets. The review followed PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) [25] guidelines; the completed PRISMA-ScR Checklist can be found in Supplementary File S1. This review covered publications indexed between January 2018 and June 2024. The search was conducted in August 2025, using the following databases: PubMed (National Library of Medicine, Bethesda, MD, USA), ProQuest (ProQuest LLC, Ann Arbor, MI, USA), and Web of Science (Clarivate Analytics, London, UK). The databases were selected to ensure broad coverage of biomedical, interdisciplinary, and applied health sciences literature relevant to TBI, AI, and health informatics. Consistent with PRISMA guidance, the goal of this review was to map the scope and characteristics of existing research rather than to achieve exhaustive retrieval across all possible datasets. A combination of keywords related to “traumatic brain injury,” “artificial intelligence,” “machine learning,” “functional outcome,” and “public dataset” were used. Search terms were customized for each database, and Boolean operators were used to refine results. Advanced search methods were used to limit the year of publication, type of publication to journal article only, and articles written in English. Full database search strategies, including Boolean operators and applied filters, are provided in Appendix A to support transparency and reproducibility. Reference lists of relevant papers were also screened to identify additional sources.
Articles were eligible if they described the use of AI or machine learning methods to model post-TBI recovery or outcomes using publicly available data. Studies were excluded that developed algorithms without applying them to functional outcomes, relied on private datasets, or used animal models.
Screening and eligibility review were completed independently by two authors. Titles and abstracts were first screened for relevance, followed by a full-text review to confirm inclusion criteria. Disagreements were resolved through discussion until consensus was reached.
Information from the studies included was compiled in a standardized data collection spreadsheet using Microsoft Excel (Microsoft Corporation, Redmond, WA, USA).
Extraction fields included study design, type of AI method, dataset name, sample size, inclusion of sociodemographic variables, outcome domain, and longitudinal follow-up time frame. When studies cited the same dataset, details were cross-referenced to confirm dataset characteristics and variable availability. Appendix B lists the citations of the studies included in the final analysis.

2.2. Phase 2: Indirect Dataset Characterization

Phase 2 involved the inspection and categorization of the datasets identified in Phase 1. The goal was to determine the extent to which publicly available AI datasets contained variables that reflected real-world TBI recovery, including demographic, functional, and contextual measures. Rather than auditing the datasets directly, this study evaluated how the dataset characteristics were described, accessed, and applied by the authors of the included studies. Information on dataset scope, accessibility, and variable representation was extracted from study methods, appendices, and referenced documentation (e.g., repository descriptions, when cited).
Because multiple studies often used the same dataset but selected different variable subsets or analysis versions, each article was coded independently to reflect how the dataset was reported and applied in that specific context. This approach captured how existing datasets were operationalized in the published AI-TBI literature, acknowledging that variable inclusion, reporting details and access conditions varied across studies.
This secondary synthesis of inferring dataset content based on reviewing the published literature allowed for each dataset to be evaluated on the following characteristics: (1) Accessibility, (2) Study Design, (3) AI Methods; (4) Functional Outcome (5) Demographic and SDOH Representation; and (6) Temporal Structure.
These variables were coded within the data collection chart to allow for consistent comparison across articles that differed in research design, methods, and reported data. The accessibility feature was coded as public or by request. Design features were documented as sample size (using n), TBI severity (based on reported initial Glasgow Coma Scale [26] score), and scope of participant recruitment (coded as institutional or multicenter). Artificial intelligence methods were categorized as either traditional machine learning, hybrid/ensemble modeling, deep machine learning, and unsupervised or clustering models.
Functional outcomes, for the purpose of this review, were defined as measures that capture one’s ability to perform activities of daily living, or cognitive-communication functioning following TBI. Functional outcome domains considered in this review included global functional status, physical independence and disability, cognitive function, communication ability, and participation-related measures. Standardized assessments reflecting functional outcomes that were used by study authors were recorded and tallied for each included study.
Demographic and SDOH variables were initially discretely documented, then coded using a structured classification framework to support consistency of coding across studies. Studies were rated as limited when only 2–3 SDOH factors, such as basic demographic information related to age and sex, or occasionally race/ethnicity were reported, and not typically incorporated into data analysis. Studies coded as partial contained at least 3 SDOH variables but had inconsistent analytic integration of these variables (for example, the study may have included multiple SDOH domains, but reported on these inconsistently (e.g., missing data, not applied to all participants, not clearly analyzed)). Moderate rated studies demonstrated more systematic inclusion of 3 or more SDOH domains, included some socioeconomic indicators (e.g., education, employment, marital status) in addition to age, sex, race/ethnicity, but did not examine them comprehensively. Studies rated as strong demonstrated comprehensive or systematic integration of multiple SDOH, such as socioeconomic status, geographic context, insurance type, and access to care. All studies were independently coded by two reviewers, with discrepancies resolved through discussion and consensus. Finally, temporal structure was coded into the following follow-up periods: at hospitalization discharge, at 3 months post-injury, at 6 months post-injury, and at 9 months or longer post-injury. Table 1 summarizes the datasets reviewed and their corresponding characteristics.

2.3. Data Transparency

All analyses were based on publicly available, de-identified data sources. No new data were collected, and no individual-level identifiers were accessed. This study was exempt from Institutional Review Board because it involved only secondary analysis of published materials using open-access data.
Finally, study search methods, data collection templates, and coding schemes are available in the Supplementary Materials to promote transparency and reproducibility. Data repositories and documentation are also referenced to allow others to replicate or build on this study.

3. Results

The results are organized according to the two-phase methodological framework. Phase 1 identified the AI datasets used in TBI outcome prediction studies through a systematic literature review, and Phase 2 indirectly audited the content and structure of those datasets.

3.1. Phase 1: Literature Review Findings

The PRISMA [25] (Figure 1) summarizes the identification, screening, eligibility, and inclusion phases for the systematic review. The search identified 115 studies meeting the criteria: 46 from PubMed, 49 from Web of Science, and 20 from ProQuest. Four duplicates were removed, leaving 111 studies for screening. Twenty-three records were excluded after abstract review, leaving 88 full-text articles for eligibility assessment.
Sixty-eight full-text studies were excluded for one or more of the following reasons: not targeting the TBI population (n = 12), not using AI or ML methods (n = 14), not focusing on functional outcomes (n = 17), being a conference abstract or editorial (n = 6), or using privately maintained datasets not available to the public (n = 19). This process yielded 20 studies meeting all inclusion criteria. Four additional studies were identified through reference searches, resulting in 24 articles included in the final qualitative synthesis.
Of the 24 studies, two (8.3%) were published between 2017 and 2019, nine (37.5%) between 2020 and 2022, and 13 (54.2%) between 2023 and 2025. A total of 19 unique datasets were identified from the 24 studies, each varying in accessibility, scale, and scope.

3.2. Publication Trends

The datasets most frequently cited were CENTER-TBI [27] (n = 5), TRACK-TBI [28] (n = 4), the Second and First Affiliated Hospitals of Anhui Medical University [29] (n = 2), and COBRIT [30] (n = 2). Additional datasets were custom-built hospital or rehabilitation registries with documented variable access. These 19 datasets were included in the Phase 2 audit.

3.3. Phase 2: Indirect Dataset Audit Findings

Each of the 19 identified datasets was audited for accessibility, inclusion of functional recovery variables, demographic representation, and longitudinal structure. Table 1 summarizes their characteristics.

3.4. Overview of the Datasets

The 24 reviewed studies drew from 19 distinct datasets that differed in accessibility and geographic scope. Most datasets were multicenter collaborations. CENTER-TBI [27] was cited in five studies (20.8%), making it the most used dataset. TRACK-TBI [28] followed with four studies (16.7%). The COBRIT [30] dataset appeared in two studies. Other datasets included IMPACT-II [31], AURORA Consortium [32], CINTER-TBI [33], GAIN Consortium [34], ProTECT III [35], TBI-PBE [36], Uppsala TBI Registry [37], and Leuven TBI Registry [38], each represented once.
Institutional datasets accounted for about one-third of the total and were drawn from hospitals and rehabilitation centers in China, Iran, Korea, Tanzania, and the United States. Examples included datasets found in published works are from the Second and First Affiliated Hospitals of Anhui Medical University [29] (n = 2; 8.3%), Rajaee Trauma Hospital [39], Xijing Hospital [40], Kilimanjaro Christian Medical Centre [41], and two U.S. institutions, Casa Colina Acute Rehabilitation Unit [42] and the University of Pittsburgh Medical Center [43].
Slightly more than half of the datasets were publicly accessible, while the remaining datasets required data request procedures. Sample sizes ranged from 147 to 12,576 participants (mean = 1741.9; median = 1001).

3.5. Types of Artificial Intelligence or Machine Learning Used

The 24 included studies employed a range of AI and machine learning techniques. Traditional models such as logistic regression, random forests, and support vector machines were most common, appearing in nine studies (37.5%). Hybrid or ensemble models were used in seven studies (29.2%), and deep learning methods, including neural networks, in six (25.0%). Two studies (8.3%) used unsupervised or clustering approaches.

3.6. Injury Severity, Functional Outcomes, and Follow-Up Periods

The included studies represented varying injury severities and outcome measures. Over half (n = 13; 54.2%) included participants across all severity levels. Six studies (29.2%) focused on moderate to severe TBI, two (8.3%) on mild cases, and two (8.3%) on severe TBI only.
Outcome measurement primarily relied on global recovery scales. The Glasgow Outcome Scale—Extended (GOS-E) [44] appeared in 13 studies (54.2%) and the original Glasgow Outcome Scale (GOS) [45] in five (20.8%). Four studies (16.7%) used physical or pain outcome measures, and three (12.5%) included mental health–related questionnaires. Cognitive and functional assessments, including the Functional Independence Measure (FIM) [46], WAIS subtests (Processing Speed, Digit Span, PSI) [47], CVLT/CVLT-II [48], and COWAT [49], were reported in several studies. Figure 2 illustrates the distribution of functional outcome measures.
Follow-up intervals differed across studies. Eleven (45.8%) analyzed outcomes at six months post-injury, five (20.8%) at three months, and six (25%) at hospital discharge. Two studies (8.3%) included nine months or longer follow-up data.

3.7. Representation of Social Determinants of Health

Across the 24 studies, the inclusion of SDOH varied substantially. Only three studies (12.5%) were rated as having strong SDOH representation. These studies reported multiple social or contextual variables such as education, income, insurance, geographic context, and social support. Seven (29.2%) studies were rated as moderate and incorporated a smaller range of socioeconomic indicators, while eleven (45.8%) were rated as limited, and relied only on basic demographic variables. Partial representation was noted in three studies (12.5%) that reported select contextual indicators for a subset of participants.
Analysis of individual SDOH variables showed that age (100%) and sex or gender (95.8%) were nearly universal. Race or ethnicity appeared in nine studies (37.5%), and an equal proportion reported mechanism of injury. Education was included in eight (33.3%), medical comorbidities in seven (29.2%), and employment status in five (20.8%). Socioeconomic indicators such as income or insurance were documented in four studies (16.7%), mental health history in three (12.5%), and marital status in one (4.2%). Figure 3 displays these results.

4. Discussion

This review analyzed publicly available datasets used in AI modeling of TBI outcomes to identify important trends and limitations in the existing literature. While research in this area has expanded rapidly over the past five years, the datasets behind the evidence remain uneven in structure, accessibility, and representativeness.

4.1. Discussion of AI-TBI Datasets

The steady increase in publications from 2020 onward shows a growing awareness and use of AI tools in predicting TBI recovery. This increase also mirrors advances in computing and expanded access to shared datasets. Yet even current datasets are narrowly focused on biomedical measures, providing little information about social context or functional aspects of recovery. The heavy reliance on large multicenter registries, such as CENTER-TBI and TRACK-TBI, points to the need for international coordination in TBI research and treatment. However, it seems to overlook the fact that smaller datasets may be important in capturing diverse regional or cultural perspectives. Large, multicenter datasets can strengthen research quality and consistency, but depending too much on a handful of them may inadvertently constrain external validity and hide regional or healthcare system differences in how people recover [50].
Institutional datasets, though fewer in number, contributed key information about specific clinical populations. These smaller data sources may offer greater details about rehabilitation, daily functioning, and community participation. This information serves as a strong complement to the broader multicenter registries that dominate AI research. Broader access to both dataset types would allow for more accurate modeling of outcomes beyond mortality or global functional status.
The regional distribution of publicly available AI-TBI datasets also has important implications for external validity. Most datasets identified in this review originate from high-income regions, particularly North America and Europe, reflecting the historical concentration of large-scale neurotrauma research infrastructure. While these datasets offer substantial sample sizes and standardized data collection, their geographic concentration may limit the generalizability of AI models to healthcare systems, cultural contexts, and resource settings that differ substantially from those in which the data were collected. As a result, model performance and clinical relevance may be reduced when applied to underrepresented regions or populations.
To address this limitation, future efforts may benefit from hybrid data strategies that combine large, multicenter registries with smaller, locally developed datasets that include richer functional, social, and contextual information. Additionally, differences in sample size, data organization, and accessibility point to the need for greater collaboration and integration across datasets. Establishing consistent data standards would allow researchers to integrate information from multiple sources, thereby improving both reproducibility, external validity, and ecological validity. Because coding systems and outcome measures are not used consistently across datasets, it is currently difficult to compare studies or build comprehensive predictive models. Collaboration between large scale research centers and smaller scale, regionally sensitive centers, may help to identify key injury patterns, care pathways, and recovery outcomes.
The analysis of AI and machine learning methods showed that the field is both growing and using a broader range of approaches. However, traditional approaches such as logistic regression and random forests remain popular because they are easier to interpret and perform well with limited data (e.g., Balaji et al., 2023 [51]). The growing use of hybrid and deep learning models shows a shift toward more advanced techniques capable of uncovering complex, nonlinear patterns in diverse data, much like the wide variation seen among people with TBI. For instance, Fang et al. (2022) [52] combined multiple algorithms to predict GOS scores, while Gravesteijn et al. (2020) [53] used data from multiple centers across Europe to validate prognostic models. As the field continues to grow, future studies will need to balance model complexity with transparency and clinical interpretability.
The dominance of global outcome scales, such as the GOS-E, along with limited domain-specific measures, shows that TBI recovery is still often modeled as one-dimensional, rather than multifaceted. These metrics help describe overall function but miss the finer details of cognitive-communication recovery and everyday participation, like reconnecting with friends, returning to work, or reengaging in the community. Including communication, cognition, and quality-of-life variables in future datasets would make AI-based predictions more clinically relevant and more aligned with rehabilitation goals. Additionally, the emotional and behavioral changes often seen after moderate to severe TBI are factors that are highly relevant in tailoring individual prognostic statements and planning for long-term rehabilitation. Orenuga et al. (2025) [54] noted that bringing together data from wearables, behavior, and neuroimaging could lead to truly personalized and flexible rehabilitation algorithms, something that has not yet been fully achieved by current predictive TBI research.
Tracking patients for longer periods is important to see how recovery continues after the early stages of injury. Most studies focused on short- to mid-term recovery, often measuring outcomes at six months (e.g., Kals et al., 2022 [34]). Only two studies collected data past the six-month mark (Balaji et al., 2023 [51]; Bark et al., 2024 [55]). This pattern follows established clinical benchmarks but makes it harder for AI models to capture chronic or late-emerging effects of TBI. Uparela-Reyes and colleagues (2024) [56] found that most published AI studies terminate follow-up before one year, resulting in missed opportunities to model long-term rehabilitation outcomes or secondary complications. Without extending longitudinal measurement, AI models are limited in their ability to capture chronic recovery trajectories, late-emerging impairments, and long-term participation outcomes. Addressing this gap should be a priority for future AI-TBI dataset development.
There was a noted gap in the literature related to the representation of individualized social determinants of health. Few studies looked at socioeconomic, environmental, or support-related variables, which limit the ability of AI models to fully understand how recovery happens in real life. Variables such as education, employment, and income were reported inconsistently, and indicators of social support or neighborhood context were rarely included. Missing this kind of data not only makes models less accurate, but also less able to identify and respond to health disparities in TBI outcomes. A few datasets did incorporate multidimensional SDOH variables. For example, Beaudoin et al. (2023) [57] included socioeconomic status, education, and employment in predicting post-concussive symptoms, offering a strong example of meaningful integration. Several multicenter datasets (Bark et al., 2024 [55]; Bhattacharyay et al., 2022 [58], 2025 [59]) contained data diverse enough to allow for SDOH analysis, but these variables were not fully explored. This underrepresentation lends support to concerns raised by other researchers who argue for transparent documentation of how data are collected, who is represented, and where information is missing [60].

4.2. Methodological Limitations of This Review

An important limitation of this review is that dataset characteristics were inferred indirectly based on how datasets were described and operationalized in published studies, rather than through direct inspection of the raw data. As a result, it is possible that some datasets may indeed contain certain SDOH variables that were not reported or utilized by authors due to differing research aims. Accordingly, the classifications presented here reflect the functional availability and analytic use of variables within AI-based outcome models, rather than the fully contained capacity of each dataset. This distinction is methodologically intentional, as the review seeks to evaluate how existing datasets are currently shaping published AI prediction practices in TBI research.

4.3. Within the Context of Cognitive Justice

Cultivating TBI research that is responsible and equitable will require consistent efforts to capture social and contextual factors that shape recovery. Accounting for variables such as education, socioeconomic status, and community support can help AI tools reflect the lived realities of the people they aim to serve. This approach builds on the principle of cognitive justice, which holds that every perspective and lived experience adds value to a better understanding of recovery. Because no single dataset or worldview can capture that complexity, cognitive justice encourages data practices that uphold both scientific integrity and the diversity of the human experience. For example, when datasets rely solely on a global outcome score, such as the GOS-E, but leave out related to specific clinical abilities, such as communication ability, employment status, or social participation, AI models may classify two individuals with identical GOS-E scores as having equivalent recovery trajectories, when there may be in fact meaningful differences in the individual’s daily functioning and community reintegration. As a result, AI-driven predictions may systematically undervalue recovery domains that matter most to patients and families, reinforcing inequities in how outcomes are defined and interpreted. As Makoelle (2014) [61] argues, cognitive justice calls for the allowance of multiple knowledge systems, an idea that, within neurorehabilitation, translates to valuing patient stories, drawing on community perspectives, and recognizing that indicators of recovery or outcome vary across cultures.
Figure 4 illustrates the components of achieving cognitive justice, as indicated by the centered star-shape embedded in the Figure, through meaningful collaboration among clinicians and researchers. Data inclusivity, ethical transparency, and functional relevance to all work to make cognitive justice a reality in TBI research and rehabilitation. Ethical and transparent data governance remains central to this process. Progress will depend on moving towards harmonized, integrated data practices that are supported by strong ethical oversight. Equally important is maintaining transparency in how data are collected, models are developed, and results are validated. This approach can help close the gap between the research laboratory and the therapy room by helping to make sure that AI-assisted tools remain valid, reliable, and grounded in person-centered care. Framing these findings through a cognitive justice lens underscores that equitable AI-driven care depends not only on who is represented in datasets, but on how fully recovery is measured and contextualized. Advancing inclusive, functionally meaningful, and ethically transparent data practices will be essential for ensuring that AI tools support person-centered, clinically relevant, and socially just TBI rehabilitation.

5. Conclusions

This review analyzed publicly available datasets that use AI modeling of TBI outcomes. The analysis revealed that most datasets lack the social, functional, and environmental variables necessary for accurate and equitable prediction of recovery. Current models effectively recognize biomedical profiles well, but they often miss the contextual factors that shape long-term recovery.
For AI to add meaningful value to clinical practice, future data collection must move beyond traditional biomedical indicators. Including variables that describe communication, social participation, and daily functioning will allow AI systems to produce predictions that are more consistent with real-world recovery. Research that promotes greater dataset inclusivity, supported by ethical oversight and consistent data standards, can help AI promote both scientific progress and person-centered care. To accomplish this, future AI-TBI datasets should work to include and report at least one domain-specific functional outcome measurement, beyond a global outcome scale of recovery. Assessments of cognition, communication, or participation can be used to better reflect real-world recovery. Second, consistent inclusion and reporting of core SDOH—including education, insurance status, geographic context—should be prioritized, as their absence limits both model interpretability and equity. Third, longitudinal follow-up extending beyond six to twelve months post-injury should be considered a minimum benchmark for capturing chronic recovery trajectories. Finally, AI studies should explicitly document which functional and contextual variables are available in the dataset versus those selected for modeling, enabling clearer interpretation of model scope and limitations.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/informatics13020033/s1, Supplementary File S1: PRISMA-ScR Checklist.

Author Contributions

Both authors, L.W.J. and K.D.H. contributed to all aspects of this project, including conceptualization, methodology, data analysis, and manuscript preparation. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

No new data were generated for this study. All analyses were based on information extracted from publicly available datasets and published articles. The data extraction templates and coding materials used in this review are available from the authors on reasonable request.

Acknowledgments

During the preparation of this manuscript, the authors used Grammarly [version 14.1267.0, Grammarly Inc., San Francisco, CA, USA] for the purposes of editing clarity of sentence structure and reference formatting. The authors have reviewed and edited the output and take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AIArtificial Intelligence
MLMachine Learning
TBITraumatic Brain Injury
SDOHSocial Determinants of Health
PRISMAPreferred Reporting Items for Systematic Reviews and Meta-Analyses
CENTER-TBICollaborative European NeuroTrauma Effectiveness Research in TBI
TRACK-TBITransforming Research and Clinical Knowledge in TBI
COBRITCiticoline Brain Injury Treatment Trial
IMPACT-IIInternational Mission for Prognosis and Analysis of Clinical Trials in TBI (Phase II)
AURORAUnderstanding Recovery After Trauma Consortium
CINTER-TBICentral Interdisciplinary Neurotrauma Research in TBI
GAINGenetic and Imaging Network
ProTECT IIIProgesterone for the Treatment of Traumatic Brain Injury, Phase III
TBI-PBETraumatic Brain Injury Practice-Based Evidence Project
GOSGlasgow Outcome Scale
GOS-EGlasgow Outcome Scale—Extended
FIMFunctional Independence Measure
WAISWechsler Adult Intelligence Scale
PSIProcessing Speed Index (WAIS subtest)
CVLT/CVLT-IICalifornia Verbal Learning Test/California Verbal Learning Test—Second Edition
COWATControlled Oral Word Association Test
IRBInstitutional Review Board

Appendix A

Database Search Strategies

Selected Databases: PubMed, ProQuest, Web of Science
Search: This search string represents a generalized summary of the concepts and Boolean logic applied across these three databases. Because each database requires a database-specific field tag, controlled vocabulary, and wildcard syntax, the exact query differed slightly by platform. These adaptations were made in accordance with each database’s indexing structure while maintaining consistent, conceptual coverage across all searches.
Search string with Boolean operators: ((traumatic brain injury) OR (TBI) OR (brain injury)) AND ((artificial intelligence) OR (AI) OR (machine learning)) AND ((public dataset) OR (dataset)) AND ((outcome) OR (recovery) OR (functional outcome) OR (functional recovery))
Advanced filters: Custom publication date range 1 January 2018–6 January 2025; Type of publication limited to journal article only, Language limited to English, Human studies only.

Appendix B

Citations of the Studies Included in This Review

Balaji, N. N. A., Beaulieu, C. L., Bogner, J., & Ning, X. (2023). Traumatic brain injury rehabilitation outcome prediction using machine learning methods. Archives of Rehabilitation Research and Clinical Translation, 5, 100295. https://doi.org/10.1016/j.arrct.2023.100295
Bark, D., Boman, M., Depreitere, B., Wright, D. W., Lewen, A., Enblad, P., Hanell, A., & Rostami, E. (2024). Refining outcome prediction after traumatic brain injury with machine learning algorithms. Scientific Reports 14, 8036. https://doi.org/10.1038/s41598-024-58527-4
Beaudoin, F. L., An, X., Basu, A., Ji, Y., Liu, M., Kessler, R. C., … McLean, S. A. (2023). Use of serial smartphone-based assessments to characterize diverse neuropsychiatric symptom trajectories in a large trauma survivor cohort. Translational Psychiatry 13(1), 4. https://doi.org/10.1038/s41398-022-02289-y
Bhattacharyay, S., Milosevic, I., Wilson, L., Menon, D., Stevens, R., Steyerberg, E., … Participants, C. T. I. (2022). The leap to ordinal: Detailed functional prognosis after traumatic brain injury with a flexible modelling approach. PLOS ONE 17(7), Article e0270973. https://doi.org/10.1371/journal.pone.0270973
Bhattacharyay, S., van Leeuwen, F., Beqiri, E., Åkerlund, C., Wilson, L., Steyerberg, E., … Participants, C.-T. I. (2025). TILTomorrow today: Dynamic factors predicting changes in intracranial pressure treatment intensity after traumatic brain injury. Scientific Reports 15(1). https://doi.org/10.1038/s41598-024-83862-x
Fang, C., Pan, Y., Zhao, L., Niu, Z., Guo, Q., & Zhao, B. (2022). A machine learning-based approach to predict prognosis and length of hospital stay in adults and children with traumatic brain injury: Retrospective cohort study. Journal of Medical Internet Research 24(12), Article e41819. https://doi.org/10.2196/41819
Folweiler, K. A., Sandsmark, D. K., Diaz-Arrastia, R., Cohen, A. S., & Masino, A. J. (2020). Unsupervised machine learning reveals novel traumatic brain injury patient phenotypes with distinct acute injury profiles and long-term outcomes. Journal of Neurotrauma 37(12), 1431–1444. https://doi.org/10.1089/neu.2019.6705
Gravesteijn, B. Y., Nieboer, D., Ercole, A., Lingsma, H. F., Nelson, D., van Calster, B., & Steyerberg, E. W. (2020). Machine learning algorithms performed no better than regression models for prognostication in traumatic brain injury. Journal of Clinical Epidemiology 122, 95–107. https://doi.org/10.1016/j.jclinepi.2020.03.005
Hibi, A., Cusimano, M. D., Bilbily, A., Krishnan, R. G., & Tyrrell, P. N. (2024). Development of a multimodal machine learning-based prognostication model for traumatic brain injury using clinical data and computed tomography scans: A CENTER-TBI and CINTER-TBI study. Journal of Neurotrauma 41(11–12), 1323–1336. https://doi.org/10.1089/neu.2023.0446
Kals, M., Kunzmann, K., Parodi, L., Radmanesh, F., Wilson, L., Izzy, S., … Menon, D. K. (2022). A genome-wide association study of outcome from traumatic brain injury. EBioMedicine 77, 103933. https://doi.org/10.1016/j.ebiom.2022.103933
Kasprowicz, M., Mataczyński, C., Uryga, A., Pelah, A. I., Schmidt, E., Czosnyka, M., & Kazimierska, A. (2025). Impact of age and mean intracranial pressure on the morphology of intracranial pressure waveform and its association with mortality in traumatic brain injury. Critical Care 29(1), 78. https://doi.org/10.1186/s13054-025-05295-w
Nielson, J. L., Cooper, S. R., Yue, J. K., Sorani, M. D., Inoue, T., Yuh, E. L., … Ferguson, A. R. (2017). Uncovering precision phenotype-biomarker associations in traumatic brain injury using topological data analysis. PLoS ONE 12(3), e0169490. https://doi.org/10.1371/journal.pone.0169490
Nourelahi, M., Dadboud, F., Khalili, H., Niakan, A., & Parsaei, H. (2022). A machine learning model for predicting favorable outcome in severe traumatic brain injury patients after 6 months. Acute and Critical Care 37(1), 45–52. https://doi.org/10.4266/acc.2021.00486
Pan, Y., Fang, C., Zhu, X., & Wan, J. (2023). Construction of a predictive model based on MIV-SVR for prognosis and length of stay in patients with traumatic brain injury: Retrospective cohort study. Digital Health 9, Article 20552076231217814. https://doi.org/10.1177/20552076231217814
Pease, M., Arefan, D., Barber, J., Yuh, E., Puccio, A., Hochberger, K., … Wu, S. (2022). Outcome prediction in patients with severe traumatic brain injury using deep learning from head CT scans. Radiology 304(2), 385–394. https://doi.org/10.1148/radiol.212181
Pirracchio, R., Yue, J. K., Manley, G. T., van der Laan, M. J., & Hubbard, A. E. (2018). Collaborative targeted maximum likelihood estimation for variable importance measure: Illustration for functional outcome prediction in mild traumatic brain injuries. Statistical Methods in Medical Research 27(1), 286–297. https://doi.org/10.1177/0962280215627335
Say, I., Chen, Y. E., Sun, M. Z., Li, J. J., & Lu, D. C. (2022). Machine learning predicts improvement of functional outcomes in traumatic brain injury patients after inpatient rehabilitation. Frontiers in Rehabilitation Sciences 3, 1005168. https://doi.org/10.3389/fresc.2022.1005168
Snider, S. B., Temkin, N. R., Sun, X., Stubbs, J. L., Rademaker, Q. J., Markowitz, A. J., … Edlow, B. L. (2024). Automated measurement of cerebral hemorrhagic contusions and outcomes after traumatic brain injury in the TRACK-TBI study. JAMA Network Open 7(8), e2427772. https://doi.org/10.1001/jamanetworkopen.2024.27772
Tritt, A., Yue, J. K., Ferguson, A. R., Torres Espin, A., Nelson, L. D., Yuh, E. L., … Zafonte, R. (2023). Data-driven distillation and precision prognosis in traumatic brain injury with interpretable machine learning. Scientific Reports 13(1), 21200. https://doi.org/10.1038/s41598-023-48054-z
Uryga, A., Mataczyński, C., Pelah, A. I., Burzyńska, M., Robba, C., & Czosnyka, M. (2024). Exploration of simultaneous transients between cerebral hemodynamics and the autonomic nervous system using windowed time-lagged cross-correlation matrices: A CENTER-TBI study. Acta Neurochirurgica (Wien) 166(1), 504. https://doi.org/10.1007/s00701-024-06375-6
Yeboah, D., Steinmeister, L., Hier, D., Hadi, B., Wunsch, D., Olbricht, G., & Obafemi-Ajayi, T. (2020). An explainable and statistically validated ensemble clustering model applied to the identification of traumatic brain injury subgroups. IEEE Access 8, 180690–180705. https://doi.org/10.1109/ACCESS.2020.3027453
Yin, A.-a., Zhang, X., He, Y.-l., Zhao, J.-j., Zhang, X., Fei, Z., … Song, B.-q. (2024). Machine learning prediction models for in-hospital postoperative functional outcome after moderate-to-severe traumatic brain injury. European Journal of Trauma and Emergency Surgery 50(4), 1219–1228. https://doi.org/10.1007/s00068-023-02434-2
Zhang, M., Guo, M., Wang, Z., Liu, H., Bai, X., Cui, S., Guo, X., Gao, L., Gao, L., Liao, A., Xing, B., & Wang, Y. (2023). Predictive model for early functional outcomes following acute care after traumatic brain injuries: A machine learning-based development and validation study. Injury 54(3), 896–903. https://doi.org/10.1016/j.injury.2023.01.004
Zimmerman, A., Elahi, C., Rocha, T. A. H., Saklta, F., Mmbaga, B. T., Staton, C. A., Ricardo, J., & Vissoci, N. (2023). Machine learning models to predict traumatic brain injury outcomes in Tanzania: Using delays to emergency care as predictors. PLOS Global Public Health 3(10), e0002153. https://doi.org/10.1371/journal.pgph.0002156

References

  1. Wilson, L.; Stewart, W.S.; Dams-O’Connor, K.; Diaz-Arrastia, R.; Horton, L.; Menon, D.K.; Polinder, S. The chronic and evolving neurological consequences of traumatic brain injury. Lancet Neurol. 2017, 16, 813–825. [Google Scholar] [CrossRef]
  2. Alshami, A.; Nashwan, A.; AlDardour, A.; Qusini, A. Artificial intelligence in rehabilitation: A narrative review on advancing patient care. Rehabilitación 2025, 59, 100911. [Google Scholar] [CrossRef] [PubMed]
  3. Cabitza, F.; Campagner, A.; Albano, D.; Aliprandi, A.; Bruno, A.; Chianca, V.; Corazza, A.; DiPietto, F.; Gambino, A.; Gitto, S.; et al. The elephant in the machine: Proposing a new metric of data reliability and its application to a medical case to assess classification reliably. Appl. Sci. 2020, 10, 4041. [Google Scholar] [CrossRef]
  4. Esteva, A.; Robicquet, A.; Ramsundar, B.; Kuleshov, V.; DePristo, M.; Chou, K.; Cui, C.; Corrado, G.; Thru, S.; Dean, J. A guide to deep learning in healthcare. Nat. Med. 2019, 25, 24–29. [Google Scholar] [CrossRef]
  5. Ashrafi, N.; Abdollahi, A.; Alaei, K.; Pishgar, M. Enhanced prediction of ventilator-associated pneumonia in patients with traumatic brain injury using advanced machine learning techniques. Sci. Rep. 2025, 15, 11363. [Google Scholar] [CrossRef] [PubMed]
  6. Osong, B.; Sribnick, E.; Groner, J.; Stanley, R.; Schulz, L.; Lu, B.; Cook, L.; Xiang, H. Development of clinical decision support for patients older than 65 years with fall-related TBI using artificial intelligence modeling. PLoS ONE 2025, 20, e0316462. [Google Scholar] [CrossRef] [PubMed]
  7. Jia, L.; Yin, Y.; Zhang, B.Y.; Li, Q.Z.; Jia, L.N.; Lai, D.T.; Yu, L.; Luo, Q.C.; Guo, E.W. Mechanisms of zhenwu decoction in targeting ATP2A2 and ATP2C1 for treating traumatic cerebral edema. Tradit. Med. Res. 2025, 10, 34. [Google Scholar] [CrossRef]
  8. Brazdzionis, J.; Radwan, M.M.; Thankan, F.; Mari, Y.M.; Baron, D.; Connett, D.; Agrawal, D.K.; Miulli, D.E. A swine model of neural circuit electromagnetic fields: Effects of immediate electromagnetic field stimulation on cortical injury. Cureus 2023, 15, e43774. [Google Scholar] [CrossRef]
  9. Zhang, Y.; Li, N.; Sun, H.; Cheng, H. Establishment and validation of a ResNet-based radiomics model for predicting prognosis in cervical spinal cord injury patients. Sci. Rep. 2025, 15, 9163. [Google Scholar] [CrossRef]
  10. Jojczuk, M.; Kaminski, P.; Gajewski, J.; Karpinski, R.; Krakowski, P.; Jonak, J.; Nogalisk, A.; Gluchowski, D. Use of neural networks based on international classification ICD-10 in patients with head and neck injuries in Lublin Province, Poland between 2006-2018, as a predictive value of the outcomes of injury sustained. Ann. Agric. Environ. Med. 2023, 30, 281–286. [Google Scholar] [CrossRef]
  11. Su, Y.; Protas, H.; Luo, J.; Chen, K.; Alosco, M.L.; Adler, C.H.; Balcer, L.J.; Bernick, C.; Au, R.; Banks, S.J.; et al. Flortaucipir tau PET findings from former professional and college american football players in the DIAGNOSE CTE research project. Alzheimer’s Dement. 2023, 20, 1827–1838. [Google Scholar] [CrossRef]
  12. Alosco, M.L.; Tripodis, Y.; Baucom, Z.H.; Adler, C.H.; Balcher, L.J.; Bernick, C.; Mariani, M.L.; Au, R.; Banks, S.J.; Barr, W.B.; et al. White matter hyperintensities in former football players. Alzheimer’s Dement. 2023, 19, 1260–1273. [Google Scholar] [CrossRef] [PubMed]
  13. Hong, Y.; Froese, L.; Ponten, E.; Fletcher-Sandersjoo, A.; Tatter, C.; Hammarlund, C.; Akerlund, C.A.I.; Tjerkaski, J.; Alpkvist, P.; Bartek, J., Jr.; et al. Critical thresholds of long-pressure reactivity index and impact of intracranial pressure monitoring methods in traumatic brain injury. Crit. Care 2024, 28, 256. [Google Scholar] [CrossRef] [PubMed]
  14. Kwak, G.H.; Kamdar, H.A.; Douglas, M.J.; Ack, S.E.; Lissak, I.A.; Williams, A.E.; Yechoor, N.; Rosenthal, E.S. Social determinants of health and limitation of life-sustaining therapy in neurocritical care: A CHoRUS pilot project. Neurocrit. Care 2024, 41, 866–879. [Google Scholar] [CrossRef] [PubMed]
  15. Ibrahim, H.; Liu, X.; Math, N.Z.; Morris, A.D.; Dennistron, A.K. Health data poverty: An assailable barrier to equitable digital health care. Lancet Digit. Health 2021, 3, e260–e265. [Google Scholar] [CrossRef]
  16. Visvanathan, S. The Search for Cognitive Justice. Knowledge in the Question Symposium, May 2009. Available online: http://www.india-seminar.com/2009/597/597_shiv_visvanathan.htm (accessed on 10 October 2025).
  17. Lukersmith, S.; Salvador-Carulla, L.; Chung, Y.; Du, W.; Sarkissian, A.; Millington, M. A realist evaluation of case-management models for people with complex health conditions using novel analytic tools. Int. J. Environ. Res. Public Health 2023, 20, 4362. [Google Scholar] [CrossRef]
  18. Carlozzi, N.E.; Troost, J.; Lombard, W.L.; Miner, J.A.; Graves, C.M.; Choi, S.W.; Wu, Z.; Sen, S.; Sander, A.M. Completion and compliance rates for an intensive mHealth study design to promote self-awareness and self-care among care partners of individuals with traumatic brain injury: Secondary analysis of a randomized controlled trial. JMIR Mhealth Uhealth 2025, 13, e54028. [Google Scholar] [CrossRef]
  19. Mavroudis, I.; Chatzikonstantinou, S.; Petridis, F.; Kazis, D.; Kamal, F.Z.; Diaconu, S.; Visternicu, M.; Ciobica, A.; Iordache, A.; Stadoleanu, C.; et al. Unveiling the functional component in post-concussive cognitive impairment: A novel modelling approach. BRAIN Broad Res. Artif. Intell. Neurosci. 2025, 16, 83–94. [Google Scholar] [CrossRef]
  20. Venkatesan, M.; Lucke-Wold, B. Mind the gut: Navigating the complex landscape of gastroprotection in neurosurgical patients. World J. Gastroenterol. 2025, 31, 102959. [Google Scholar] [CrossRef]
  21. Alsaigh, R.; Mehmood, R.; Katib, I.; Liang, X.; Alshanqiti, A.; Corchado, J.M.; See, S. Harmonizing AI governance regulations and neuroinformatics: Perspectives on privacy and data sharing. Front. Neuroinform. 2024, 18, 1472653. [Google Scholar] [CrossRef]
  22. Hegde, A.; Vijaysenan, D.; Mandava, P.; Menon, G. The use of cloud based machine learning to predict outcome in intracerebral haemorrhage without explicit programming expertise. Neurosurg. Rev. 2024, 47, 883. [Google Scholar] [CrossRef] [PubMed]
  23. Arksey, H.; O’Malley, L. Scoping studies: Towards a methodological framework. Int. J. Soc. Res. Methodol. 2005, 8, 19–32. [Google Scholar] [CrossRef]
  24. Levac, D.; Colquhoun, H.; O’Brien, K.K. Scoping studies: Advancing the methodology. Implement. Sci. 2010, 5, 69. [Google Scholar] [CrossRef]
  25. Tricco, A.C.; Lillie, E.; Zarin, W.; O’Brien, K.K.; Colquhoun, H.; Levac, D.; Moher, D.; Peters, M.D.; Horsley, T.; Weeks, L.; et al. PRISMA extension for scoping reviews (PRISMAScR): Checklist and explanation. Ann. Intern. Med. 2018, 169, 467–473. [Google Scholar] [CrossRef]
  26. Teasdale, G.; Jennett, B. Assessment of coma and impaired consciousness: A practical scale. Lancet 1974, 304, 81–84. [Google Scholar] [CrossRef] [PubMed]
  27. Maas, A.I.R.; Menon, D.K.; Steyerberg, E.W.; Citerio, G.; Lecky, F.; Manley, G.T.; Hill, S.; Legrand, V.; Sorgner, A. Collaborative european neurotrauma effectiveness research in traumatic brain injury (CENTER-TBI): A prospective longitudinal observational study. Neurosurgery 2015, 76, 67–80. [Google Scholar] [CrossRef]
  28. Yue, J.K.; Vassar, M.J.; Lingsma, H.F.; Cooper, S.R.; Okonkwo, D.O.; Valadka, A.B.; Gordon, W.A.; Maas, A.I.R.; Mukherjee, P.; Yuh, E.L.; et al. Transforming research and clinical knowledge in traumatic brain injury pilot: Multicenter implementation of the common data elements for traumatic brain injury. J. Neurotrauma 2013, 30, 1831–1844. [Google Scholar] [CrossRef]
  29. Pan, Y.; Fang, C.; Zhu, X.; Wan, J. Construction of a predictive model based on MIV-SVR for prognosis of traumatic brain injury. Digit. Health 2023, 9, 20552076231217814. [Google Scholar] [CrossRef]
  30. Zafonte, R.; Friedewald, W.T.; Lee, S.M.; Levin, B.; Diaz-Arrastia, R.; Ansel, B.; Eisenberg, H.; Timmons, S.D.; Temkin, N.R.; Novack, T.; et al. The citicoline brain injury treatment trial (COBRIT): Design and methods. J. Neurotrauma 2009, 26, 1993–2000. [Google Scholar] [CrossRef]
  31. Marmarou, A.; Lu, J.; Butcher, I.; McHugh, G.S.; Mushkudiani, N.A.; Murray, G.D.; Steyerberg, E.W.; Maas, A.I.R. IMPACT database of traumatic brain injury: Design and Description. J. Neurotrauma 2007, 24, 239–250. [Google Scholar] [CrossRef]
  32. McLean, S.A.; Ressler, K.J.; Koenen, K.C.; Neylan, T.; Germine, L.; Jovanovic, T.; Clifford, G.D.; Zeldin, D.C.; Rasmusson, A.M.; Baker, D.G.; et al. The AURORA study: A longitudinal, multimodal library of brain biology and function after traumatic stress exposure. Mol. Psychiatry 2019, 25, 283–296. [Google Scholar] [CrossRef] [PubMed]
  33. Hibi, A.; Cusimano, M.D.; Bilbily, A.; Krishnan, R.G.; Tyrrell, P.N. Development of a multimodal machine learning-based prognostication model for traumatic brain injury using clinical data and computed tomography scans: A CENTER-TBI and CINTER-TBI study. J. Neurotrauma 2024, 41, 1323–1336. [Google Scholar] [CrossRef] [PubMed]
  34. Kals, M.; Kunzmann, K.; Parodi, L.; Radmanesh, F.; Wilson, L.; Izzy, S.; Anderson, C.D.; Puccio, A.M.; Okonkwo, D.O.; Temkin, N.; et al. A genome-wide association study of outcome from traumatic brain injury. eBioMedicine 2022, 77, 103933. [Google Scholar] [CrossRef]
  35. Wright, D.W.; Yeatts, S.D.; Silbergleit, R.; Palesch, Y.Y.; Hertzberg, V.S.; Frankel, M.; Goldstein, F.C.; Caveney, A.F.; Howlett-Smith, H.; Bengelink, E.M.; et al. Very early administration of progesterone for acute traumatic brain injury. N. Engl. J. Med. 2014, 371, 2457–2466. [Google Scholar] [CrossRef]
  36. Horn, S.D.; Corrigan, J.; Bogner, J.; Hammond, F.; Seel, R.; Smout, R.; Barrett, R.; Dijkers, M.; Whiteneck, G. Traumatic brain injury—Practice based evidence study: Design and patients, centers, treatments, and outcomes. Arch. Phys. Med. Rehabil. 2015, 96, S178–S196. [Google Scholar] [CrossRef]
  37. Nyholm, L.; Howells, T.; Enblad, P.; Lewen, A. Introduction of the Uppsala traumatic brain injury register for regular surveillance of patient characteristics and neurointensive care management including secondary insult quantification and clinical outcome. Upsala J. Med. Sci. 2013, 118, 169–180. [Google Scholar] [CrossRef]
  38. Laic, R.A.G.; Sloten, J.V.; Depreitere, B. Traumatic brain injury in the elderly population: A 20-year experience in a tertiary neurosurgery center in Belgium. Acta Neurochir. 2022, 164, 1407–1419. [Google Scholar] [CrossRef]
  39. Farrokhi, A.; Jalali, M.; Jabal, M.S.; Abdollahifard, S.; Taheri, R.; Yousefi, O.; Niakan, A.; Khalili, H. A practical approach to predicting long-term outcomes in traumatic brain injury: Enhancing clinical decision-making with machine learning. Comput. Biol. Med. 2025, 196, 110827. [Google Scholar] [CrossRef]
  40. Yin, A.-A.; Zhang, X.; He, Y.-L.; Zhao, J.-J.; Zhang, X.; Fei, Z.; Lin, W.; Song, B.-Q. Machine learning prediction models for in-hospital postoperative functional outcome after moderate-to-severe traumatic brain injury. Eur. J. Trauma Emerg. Surg. 2024, 50, 1219–1228. [Google Scholar] [CrossRef]
  41. Zimmerman, A.; Elahi, C.; Rocha, T.A.H.; Saklta, F.; Mmbaga, B.T.; Staton, C.A.; Ricardo, J.; Vissoci, N. Machine learning models to predict traumatic brain injury outcomes in Tanzania: Using delays to emergency care as predictors. PLoS Glob. Public Health 2023, 3, e0002153. [Google Scholar] [CrossRef] [PubMed]
  42. Say, I.; Chen, Y.E.; Sun, M.Z.; Li, J.J.; Lu, D.C. Machine learning predicts improvement of functional outcomes in traumatic brain injury patients after inpatient rehabilitation. Front. Rehabil. Sci. 2022, 3, 1005168. [Google Scholar] [CrossRef]
  43. Pease, M.; Arefan, D.; Barber, J.; Yuh, E.; Puccio, A.; Hochberger, K.; Nwachuku, E.; Roy, S.; Casillo, S.; Temkin, N.; et al. Outcome prediction in patients with severe traumatic brain injury using deep learning from head CT scans. Radiology 2022, 304, 385–394. [Google Scholar] [CrossRef]
  44. Wilson, J.T.L.; Pettigrew, L.E.L.; Teasdale, G.M. Structured interviews for the glasgow outcome scale and the extended glasgow outcome scale: Guidelines for their use. J. Neurotrauma 1998, 15, 573–585. [Google Scholar] [CrossRef]
  45. Jennett, B.; Bond, M. Assessment of outcome after severe brain damage: A practical scale. Lancet 1975, 305, 480–484. [Google Scholar] [CrossRef] [PubMed]
  46. Keith, R.A.; Granger, C.V.; Hamilton, B.B.; Sherwin, F.S. The Functional independence measure: A new tool for rehabilitation. Adv. Clin. Rehabil. 1987, 1, 6–18. [Google Scholar] [PubMed]
  47. Wechsler, D. Wechsler Adult Intelligence Scale, 4th ed.; (WAIS-IV) Manual; Pearson: London, UK, 2008; Available online: https://psycnet.apa.org/doi/10.1037/t15169-000 (accessed on 10 October 2025).
  48. Delis, D.C.; Kramer, J.H.; Kaplan, E.; Ober, B.A. California Verbal Learning Test, 2nd ed.; CVLT-II; Psychological Corporation: San Antonio, TX, USA, 2000; Available online: https://psycnet.apa.org/doi/10.1037/t15072-000 (accessed on 10 October 2025).
  49. Benton, A.L.; Hamsher, K.; Sivan, A.B. Multilingual Aphasia Examination, 3rd ed.; Manual; AJA Associates: Iowa City, IA, USA, 1994. [Google Scholar]
  50. Beard, K.; Pennington, A.M.; Gauff, A.K.; Mitchell, K.; Smith, J.; Marion, D.W. Potential applications and ethical considerations for artificial intelligence in traumatic brain injury management. Biomedicines 2024, 12, 2459. [Google Scholar] [CrossRef]
  51. Balaji, N.N.A.; Beaulieu, C.L.; Bogner, J.; Ning, X. Traumatic brain injury rehabilitation outcome prediction using machine learning methods. Arch. Rehabil. Res. Clin. Transl. 2023, 5, 100295. [Google Scholar] [CrossRef]
  52. Fang, C.; Pan, Y.; Zhao, L.; Niu, Z.; Guo, Q.; Zhao, B. A machine learning-based approach to predict prognosis and length of hospital stay in adults and children with traumatic brain injury: Retrospective cohort study. J. Med. Internet Res. 2022, 24, e41819. [Google Scholar] [CrossRef]
  53. Gravesteijn, B.Y.; Nieboer, D.; Ercole, A.; Lingsma, H.F.; Nelson, D.; van Calster, B.; Steyerberg, E.W. Machine learning algorithms performed no better than regression models for prognostication in traumatic brain injury. J. Clin. Epidemiol. 2020, 122, 95–107. [Google Scholar] [CrossRef]
  54. Orenuga, S.; Jordache, P.; Mirzai, D.; Monteros, T.; Gonzalez, E.; Madkoor, A.; Hirani, R.; Tiwari, R.K.; Etienne, M. Traumatic brain injury and artificial intelligence: Shaping the future of neurorehabilitation—A review. Life 2025, 15, 424. [Google Scholar] [CrossRef]
  55. Bark, D.; Boman, M.; Depreitere, B.; Wright, D.W.; Lewen, A.; Enblad, P.; Hanell, A.; Rostami, E. Refining outcome prediction after traumatic brain injury with machine learning algorithms. Sci. Rep. 2024, 14, 8036. [Google Scholar] [CrossRef]
  56. Uparela-Reyes, M.J.; Villegas-Trujillo, L.M.; Cespedes, J.; Velásquez-Vera, M.; Rubiano, A.M. Usefulness of artificial intelligence in traumatic brain injury: A bibliometric analysis and mini-review. World Neurosurg. 2024, 188, 83–92. [Google Scholar] [CrossRef]
  57. Beaudoin, F.L.; An, X.; Basu, A.; Ji, Y.; Liu, M.; Kessler, R.C.; Doughtery, R.F.; Zeng, D.; Bollen, K.A.; House, S.L.; et al. Use of serial smartphone-based assessments to characterize diverse neuropsychiatric symptom trajectories in a large trauma survivor cohort. Transl. Psychiatry 2023, 13, 4. [Google Scholar] [CrossRef]
  58. Bhattacharyay, S.; Milosevic, I.; Wilson, L.; Menon, D.K.; Stevens, R.D.; Steyerberg, E.W.; Nelson, D.W.; Ercole, A. CENTER-TBI Investigators Participants. The leap to ordinal: Detailed functional prognosis after traumatic brain injury with a flexible modelling approach. PLoS ONE 2022, 17, e0270973. [Google Scholar] [CrossRef] [PubMed]
  59. Bhattacharyay, S.; van Leeuwen, F.D.; Beqiri, E.; Åkerlund, C.A.I.; Wilson, L.; Steyerberg, E.W.; Nelson, D.W.; Maas, A.I.R.; Menon, D.K.; Ercole, A.; et al. TILTomorrow today: Dynamic factors predicting changes in intracranial pressure treatment intensity after traumatic brain injury. Sci. Rep. 2025, 15, 95. [Google Scholar] [CrossRef]
  60. Arora, A.; Alderman, J.E.; Palmer, J.; Ganapathi, S.; Laws, E.; McCradden, M.D.; Oakden-Rayner, L.; Pfohl, S.R.; Ghassemi, M.; McKay, F.; et al. The value of standards for health datasets in artificial intelligence-based applications. Nat. Med. 2023, 29, 2929–2938. [Google Scholar] [CrossRef] [PubMed]
  61. Makoelle, T.M. Cognitive justice: A road map for equitable inclusive learning environments. Int. J. Educ. Res. 2014, 2, 505–518. Available online: http://www.ijern.com/journal/July-2014/39.pdf (accessed on 11 September 2025).
Figure 1. PRISMA 2020 flow diagram for new scoping reviews showing the process of identifying, screening, and determining eligibility for article inclusion within the review.
Figure 1. PRISMA 2020 flow diagram for new scoping reviews showing the process of identifying, screening, and determining eligibility for article inclusion within the review.
Informatics 13 00033 g001
Figure 2. Functional outcome measures used across included studies.
Figure 2. Functional outcome measures used across included studies.
Informatics 13 00033 g002
Figure 3. Representation of social determinants of health (SDOH) across included studies.
Figure 3. Representation of social determinants of health (SDOH) across included studies.
Informatics 13 00033 g003
Figure 4. Conceptual framework linking data inclusivity, functional relevance, and ethical transparency as overlapping principles that enable cognitive justice (inclusive, interpretable, and equitable) in AI-based TBI research and rehabilitation.
Figure 4. Conceptual framework linking data inclusivity, functional relevance, and ethical transparency as overlapping principles that enable cognitive justice (inclusive, interpretable, and equitable) in AI-based TBI research and rehabilitation.
Informatics 13 00033 g004
Table 1. TBI Dataset Summary.
Table 1. TBI Dataset Summary.
Dataset Name, Scope, RegionAI UsedSDOH RatingFunctional Outcome MeasureFollow-Up Time PeriodAccess
CENTER-TBI (n = 5)
Multicenter, Europe and Israel
Deep LearningPartialGOS-EAt 6 months post-injuryPublic
Traditional MLLimitedGOS-EAt 6 months post-injury
Deep LearningLimitedGOS-EAt 6 months post-injury
Deep LearningModerateGOS-E; Therapy Intensity LevelDaily predictions of next day therapy intensity; 6 months post-injury
Deep LearningLimitedGOS-EAt 6 months post-injury
TRACK-TBI (n = 4)
Multicenter, US
Hybrid/EnsembleStrongGOS-E; Symptom ChecklistAt 3 and 6 months post-injuryPublic
Unsupervised/ClusteringModerateGOS-EAt 3 and 6 months post-injury
Unsupervised/ClusteringPartialGOS-EAt 3 and 6 months post-injury
Deep LearningLimitedGOSAt 6 months post-injury
Second and First Affiliated Hospital of Anhui Medical University (n = 2)
Institutional, China
Hybrid/EnsembleLimitedGOSAt discharge from hospitalRequest
Hybrid/EnsembleLimited At discharge from hospital
COBRIT (n = 2)
Multicenter, US
Hybrid/EnsembleModerateGOS-E; CVLT-II, WAIS-III, COWAT, BSI-18At 1 month, 3 months and 6 months post-injuryPublic
Unsupervised/EnsemblePartialGOS-EAt 3 and 6 months post-injury
Uppsala TBI Registry (n = 1)
Multicenter, Europe
Traditional MLPartialGOS-EWithin 6–12 months post-injuryRequest
Leuven TBI Registry (n = 1)
Multicenter, Europe
Traditional MLPartialGOS-EWithin 6–12 months post-injuryRequest
ProTECT III (n = 1)
Multicenter, US
Traditional MLPartialGOS-EWithin 6–12 months post-injuryRequest
AURORA Consortium (n = 1)
Multicenter, US
Traditional MLStrongSymptom ChecklistMultiple time points before 8 weeks post-injury, and again at 8 weeks post-injuryPublic
TBI-PBE (n = 1)
Multicenter, US
Traditional MLModerateFIMAt discharge from acute rehab, and at 9 months post-dischargeRequest
IMPACT-II (n = 1)
Multicenter, Europe and Israel
Traditional MLLimitedGOS-EAt 6 months post-injuryPublic
CINTER-TBI (n = 1)
Multicenter, Europe, Israel and India
Deep LearningLimitedGOS-EAt 6 months post-injuryPublic
GAIN Consortium (n = 1)
Multicenter, US and Europe
Traditional MLModerateGOS-EAt 6 months post-injuryPublic
Rajaee (Emtiaz) Trauma Hospital (n = 1), Institutional, IranTraditional MLModerateGOS-EAt 3 and 6 months post-injuryRequest
University of Pittsburgh Medical Center (n = 1), Institutional, USDeep LearningLimitedGOSAt hospital dischargePublic
Casa Colina Acute Rehab Unit (n = 1)
Institutional, US
Traditional MLLimitedFIMAt acute rehab dischargeRequest
Wroclow University Hospital (n = 1)
Institutional, Europe
Deep LearningLimitedGOS; GOS-EAt 3 or 6 months post-injuryPublic
Xijing Hospital (n = 1)
Institutional, China
Hybrid/EnsembleLimitedGOS-EAt hospital dischargeRequest
Korea Neuro-Trauma Data Bank System (n = 1), Institutional, KoreaTraditional MLLimitedGOSAt hospital dischargeRequest
Kilimanjaro (n = 1)
Institutional, Tanzania
Traditional MLModerateGOSAt hospital dischargeRequest
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Johnson, L.W.; Hall, K.D. Equity, Function, and Data: A Review of Social and Functional Representation in AI Datasets for Traumatic Brain Injury. Informatics 2026, 13, 33. https://doi.org/10.3390/informatics13020033

AMA Style

Johnson LW, Hall KD. Equity, Function, and Data: A Review of Social and Functional Representation in AI Datasets for Traumatic Brain Injury. Informatics. 2026; 13(2):33. https://doi.org/10.3390/informatics13020033

Chicago/Turabian Style

Johnson, Leslie W., and Kellyn D. Hall. 2026. "Equity, Function, and Data: A Review of Social and Functional Representation in AI Datasets for Traumatic Brain Injury" Informatics 13, no. 2: 33. https://doi.org/10.3390/informatics13020033

APA Style

Johnson, L. W., & Hall, K. D. (2026). Equity, Function, and Data: A Review of Social and Functional Representation in AI Datasets for Traumatic Brain Injury. Informatics, 13(2), 33. https://doi.org/10.3390/informatics13020033

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop