Structuring Variability in Human Gait Datasets: A Covariate-Centered Taxonomy and Systematic Review of Image- and Depth-Based Collections
Abstract
1. Introduction
Contributions
- We propose a covariate-centered and modality-agnostic taxonomy for publicly available image- and depth-based gait datasets, organizing dataset variability into scene-level, user-level, and sensor-level dimensions.
- We conduct a PRISMA-guided systematic review of 47 publicly available gait datasets across healthcare-oriented, biometric-oriented, and attribute-recognition domains, documenting the dataset selection process, eligibility criteria, and extracted metadata to support reproducibility.
- We instantiate the proposed taxonomy across the reviewed datasets, providing a structured comparative analysis of dataset scale, acquisition settings, sensing modalities, sequence characteristics, annotation richness, covariate coverage, and documentation quality.
- We synthesize the main limitations of the current dataset landscape and show how the proposed taxonomy can support dataset selection, covariate-aware benchmark design, model evaluation, and future dataset reporting practices.
2. Review Methodology
2.1. Research Questions
- RQ1. Which publicly accessible image- and depth-based datasets have been made available for human gait analysis, gait recognition, healthcare-oriented gait assessment, and gait-related attribute recognition?
- RQ2. What scene-level, user-level, and sensor-level sources of variability are explicitly represented or reported in these datasets?
- RQ3. How are gait datasets distributed across application domains, namely healthcare-oriented, biometric-oriented, and attribute-recognition contexts?
- RQ4. Which covariates are most frequently represented, under-represented, or insufficiently documented across the included datasets?
- RQ5. How can a covariate-centered taxonomy support a more reproducible and comparable characterization of heterogeneous human gait datasets?
2.2. Search Strategy
(“gait dataset” OR “gait database” OR “human gait recognition dataset” OR “gait analysis dataset” OR “walking dataset”) AND (RGB OR depth OR RGB-D OR silhouette OR video OR vision OR image)
(“gait recognition” OR “gait analysis” OR “human walking”) AND (dataset OR database OR benchmark) AND (view OR clothing OR carrying OR covariate OR depth OR RGB-D OR silhouette)
2.3. Eligibility Criteria
- Inclusion criteria:
- Visual modality: the dataset must include RGB, video, silhouette-based, depth, RGB-D, infrared, skeleton, pose/keypoint, or human-parsing data derived from visual sensing;
- Task relevance: the dataset must contain human gait or walking sequences as a primary component;
- Public availability: the dataset must be publicly accessible through direct download, institutional repository, project webpage, GitHub repository, or request-based access under standard data-use conditions;
- Dataset documentation: sufficient information must be available to extract at least the core dataset descriptors, including acquisition modality, number of subjects or sequences, acquisition setting, and relevant covariates;
- Human subjects: the dataset must involve human participants rather than simulated, synthetic, or non-human motion data.
- Exclusion criteria:
- Datasets based exclusively on non-visual modalities, such as inertial, pressure, force-plate, electromyography, or wearable-sensor data;
- Datasets in which gait or walking is only incidental and not a primary component of the acquisition protocol;
- Datasets with inaccessible, discontinued, private, or non-reproducible access conditions;
- Datasets lacking essential acquisition, modality, or annotation metadata required for taxonomy coding;
- Duplicate records, derivative reports, or subsets that did not introduce additional subjects, modalities, covariates, annotations, or acquisition conditions beyond a previously included dataset;
- Datasets containing only synthetic, simulated, or avatar-based walking data.
2.4. Screening and Selection Process
2.5. Data Extraction and Taxonomy Coding
2.6. Dataset Quality, Access, and Documentation Assessment
3. Systematic Review by Main Area of Use
3.1. Healthcare-Oriented Datasets
3.2. Biometric-Oriented Datasets
3.3. Attribute-Recognition Datasets
4. Covariate Taxonomy Framework
4.1. Motivation and Covariate-Centered Design Rationale
4.2. Formal Definition
4.3. Design Properties of the Taxonomy
- Bounded coverage: The taxonomy retains dominant and reproducibly documentable sources of variability, avoiding excessive fragmentation into dataset-specific or weakly reported attributes.
- Origin-based coding: Each covariate is assigned according to its primary source of variability in the acquisition and observation process. Environmental and contextual conditions are encoded as scene-level attributes, participant-related factors as user-level attributes, and acquisition or signal-formation properties as sensor-level attributes.
- Non-redundancy: No covariate is intentionally duplicated across dimensions. When a factor may affect multiple aspects of the observed gait signal, it is coded according to its primary origin rather than all downstream effects.
- Controlled granularity: Descriptors support both coarse comparison and refined specification when documentation allows. For example, a dataset may be described broadly as multi-view while also specifying the number, angular distribution, or height of the views when reported.
- Modality-agnostic applicability: The taxonomy is designed to apply across RGB, depth, RGB-D, silhouette, infrared, skeleton, pose, parsing, point-cloud, and multimodal gait datasets. Covariates are therefore defined at the level of dataset structure and acquisition variability rather than at the level of a single representation.
- Protocol awareness: The selected attributes map directly to common evaluation paradigms, including cross-view, cross-clothes, cross-speed, cross-time, cross-scene, cross-carrying, and cross-sensor evaluation. This supports a clearer connection between dataset design and benchmark interpretation.
- Conservative coding: Covariate values are assigned only when supported by dataset papers, official documentation, repositories, or access pages. Missing or ambiguous information is coded as not reported rather than inferred from unstated assumptions.
4.4. Scene-Level Covariates ()
- (A) Acquisition environment identifies the general physical context in which data collection takes place. The values observed across the reviewed datasets range from controlled indoor environments, such as laboratories, corridors, clinical rooms, studios, gyms, museums, and stair setups, to outdoor environments, multi-scene outdoor settings, and in-the-wild locations. Some datasets combine indoor and outdoor acquisition. This attribute is intended to capture the broad environmental setting, while more specific aspects such as illumination control, background complexity, and camera arrangement are encoded separately.
- (B) Background and scene complexity describes the visual structure of the scene behind and around the walking subject. Common values include static backgrounds, green chroma-key backgrounds, real-world or varied backgrounds, dynamic backgrounds, cluttered scenes, crowded scenes, static obstacles, and scene-induced occlusion. This attribute is particularly relevant because background complexity and occlusion directly affect silhouette extraction, segmentation, pose estimation, and parsing-based representations.
- (C) Walking surface or terrain captures the physical support on which gait is performed. Most datasets use planar overground walking, but several introduce treadmill-constrained walking, treadmill incline, stairs, ramps, bumpy or soft surfaces, curved roads, or mixed terrain. This attribute is distinct from gait condition: the surface defines the physical constraint imposed by the environment, whereas walking style, pathology, speed, or activity condition are treated as user-level or protocol-related locomotion variables.
- (D) Environmental control specifies whether the acquisition conditions are controlled or uncontrolled. Controlled settings include laboratory, clinical, studio, or corridor-based acquisitions with stable lighting and constrained backgrounds. Uncontrolled settings include natural illumination, illumination changes, night-time acquisition, day/night variation, real-world outdoor conditions, or mixed indoor/outdoor protocols. This attribute provides a compact indication of environmental predictability and is useful for distinguishing laboratory-style benchmarks from more realistic surveillance or in-the-wild datasets.
- (E) Viewpoint configuration describes the perspective from which gait is observed. Values include single-view setups, single side-view acquisition, frontal-view acquisition, elevated side views, limited multi-view configurations, full 360-degree protocols, and large multi-camera or surveillance-style view distributions. Although viewpoint depends on camera placement, it is treated here as a scene-level covariate because it describes the geometric relation between the walking subject and the observation scene. The physical hardware arrangement itself, including camera type, number, placement, height, calibration, and synchronization, is encoded separately under sensor-level covariates.
- (F) Walking path or trajectory characterizes the spatial form of the walking route. Common values include straight trajectories, straight bidirectional walking paths, treadmill-constrained walking, curved trajectories, circular routes, figure-eight trajectories, square walking routes, cross-scene routes, round-trip paths with turns, assisted walking paths, mixed-task routes, and unconstrained trajectories. This attribute is important because trajectory structure affects body orientation, view transitions, occlusion patterns, and the temporal continuity of gait cycles.
- (G) Temporal acquisition structure captures how data collection is organized over time. Many datasets are single-session collections, whereas others involve multi-session acquisition, time gaps between recordings, long-run exhibition-based collection, seasonal variation, day/night variation, or collection periods spanning several months. When the temporal organization is not documented, the attribute should be coded as not reported. This covariate is particularly relevant for evaluating robustness to temporal change, re-acquisition effects, clothing seasonality, and long-term variability.
4.5. User-Level Covariates ()
- (H) Activity or gait condition identifies the locomotor task, walking pattern, or clinically relevant gait state represented in the dataset. The most common condition across the reviewed datasets is ordinary walking, but several datasets include additional activity or gait variants, such as running, stair ascent, speed-transition walking, stop/non-stop walking, prosthetic or pathological walking, simulated abnormal gait, asymmetric gait, Parkinsonian gait, knee osteoarthritis, scoliosis screening classes, freezing, limp, rigidity, or attribute-defined gait patterns. This attribute captures what the participant is doing or how gait is performed. It is distinct from scene-level terrain: for example, stairs as a physical surface are encoded as a scene-level covariate, whereas stair ascent as a locomotor task is encoded as a user-level gait condition.
- (I) Speed or pace condition describes how walking speed is determined, constrained, or varied. Common values include self-selected pace, slow walking, fast walking, speed-controlled walking, acceleration or deceleration protocols, stationary conditions, and treadmill-based constant-speed acquisition. Some datasets report explicit speed ranges or treadmill speeds, whereas others only indicate whether participants walked naturally or under imposed pace constraints. When speed is not described with sufficient detail, the attribute should be coded as not reported. This covariate is important because pace affects cadence, stride length, silhouette dynamics, pose trajectories, and temporal gait representations.
- (J) Clothing or appearance condition captures worn appearance factors that may modify the visible body shape or affect the extracted gait representation. Values observed across the reviewed datasets include natural clothing, explicit clothing variation, whole-body or partial clothing changes, coat or jacket conditions, lab coats, down jackets, footwear variation, shoes condition, and hat condition. This covariate is especially relevant for silhouette-, parsing-, and appearance-based recognition, since clothing can alter body contour, limb visibility, and segmentation quality. Clothing and worn accessories should be distinguished from external objects carried by the participant.
- (K) Carrying/load and upper-limb condition describes external objects carried by the participant, load-related conditions, or configurations that affect arm motion and body-part visibility. Common values include no load, bag, backpack, shoulder bag, handbag, travelling bag, luggage, book, box, heavy box, trolley, umbrella, ball, small or large carried items, object carrying, phone use, lift-stuff conditions, and hands-in-pockets. Although not all of these conditions correspond to physical load, they are grouped together because they can modify natural arm swing, occlude body regions, change silhouette shape, or introduce asymmetric motion. This attribute therefore supports cross-carrying and upper-limb-occlusion analysis.
- (L) Participant descriptors capture demographic, anthropometric, clinical, or contextual metadata associated with the subjects. Values observed across the reviewed datasets include age, sex, height, weight, body mass, ethnicity, nationality, profession, age group, physical activity level, and clinical severity. These descriptors do not always play the same role across tasks: they may function as target labels in attribute-recognition or healthcare-oriented datasets, as descriptive metadata in biometric datasets, or as covariates for stratified robustness and fairness analysis. When such information is unavailable or insufficiently specified, the attribute should be coded as not reported.
4.6. Sensor-Level Covariates ()
- (M) Data modality or representation identifies the sensor data and derived representations made available by each dataset. Values observed across the reviewed datasets include RGB video, silhouettes, depth maps, infrared images, gait templates, 2D and 3D keypoints, pose representations, parsing maps, optical flow, point clouds, 3D volumes, 3D meshes, and gait descriptors. This attribute is intentionally broader than sensor type: it captures not only the raw acquisition modality but also the representations distributed or used as part of the dataset. This distinction is important because many gait benchmarks are evaluated on derived representations, such as silhouettes or GEI, rather than on raw video.
- (N) Frame rate or temporal sampling describes the temporal frequency at which gait data are acquired or represented. Common values include 10, 15, 20, 25, 30, 50, and 60 fps, although some datasets report mixed temporal rates across modalities, such as RGB and LiDAR streams, or include simulated low-frame-rate variants. When the acquisition or representation rate is not specified, the attribute should be coded as not reported. This covariate is particularly relevant because temporal sampling affects gait-cycle reconstruction, cadence estimation, event detection, motion smoothness, and the comparability of sequence-based models.
- (O) Spatial resolution captures the spatial size of the available visual or derived data. This may refer to RGB frames, depth maps, infrared images, silhouettes, GEI templates, optical-flow maps, parsing outputs, or other representations, depending on what is distributed with the dataset. Some datasets report a single image resolution, whereas others report different resolutions for different modalities or representations. When resolution is unavailable or only partially documented, the attribute should be coded conservatively as not reported. This covariate is important because spatial resolution influences segmentation quality, body-shape detail, keypoint localization, and the level of anatomical information preserved in the representation.
- (P) Acquisition setup or sensor placement describes the physical sensing arrangement used to capture gait. Values include fixed single-camera setups, multi-camera arcs, circular camera configurations, surveillance camera networks, Kinect-based RGB-D setups, thermal cameras, ASUS Xtion sensors, ceiling cameras, mobile robot-mounted RGB–LiDAR systems, and cameras placed at specified heights, distances, or angular layouts. This attribute differs from viewpoint configuration: viewpoint describes the observational perspective available in the dataset, whereas acquisition setup describes the physical hardware arrangement that produces those views.
- (Q) Synchronization or alignment specifies whether multiple streams, sensors, or views are temporally or geometrically coordinated. Values observed across the reviewed datasets include synchronized and calibrated multi-camera setups, timestamped multimodal streams, RGB–depth calibration or alignment, camera synchronization information, GPS-clock synchronization, external reference systems, not synchronized, not applicable, and not reported. This distinction is especially important for multi-view, multimodal, RGB-D, depth, skeleton, LiDAR, and reference-instrument datasets, where temporal or spatial misalignment can affect reconstruction, feature extraction, and cross-modal evaluation. For single-camera datasets without multiple streams, the attribute may be coded as not applicable.
- (R) Sensor-dependent acquisition limitations records limitations or artefacts that are directly attributable to the sensing device, recording hardware, or acquisition stream, as explicitly reported in the dataset documentation. This attribute is restricted to sensor-dependent issues such as frame loss, dropped frames, temporal gaps in the recorded stream, sensor noise, depth noise, LiDAR measurement noise, infrared/thermal sensor limitations, motion blur, rolling-shutter effects, sensor saturation, limited sensing range, or compression artefacts introduced at acquisition or recording time. This attribute should not be interpreted as a general quality score; it records only documented acquisition-level limitations that may affect reproducibility, temporal continuity, or signal reliability.
4.7. Boundary Cases and Coding Rules
4.8. Covariate Overview
5. Taxonomy-Aligned Dataset Tables
6. Analytical Insights Enabled by the Taxonomy
6.1. Scene-Level Patterns
6.2. User-Level Patterns
6.3. Sensor-Level Patterns
6.4. Covariate Value Distributions
6.5. Synthesis
- Explicit quantification of covariate coverage across heterogeneous datasets;
- Identification of domain-specific biases in dataset design;
- separation of scene-level, user-level, and sensor-level sources of variability;
- Distinction between explicitly varied, controlled, not reported, and not applicable covariates;
- Identification of underrepresented dimensions, particularly temporal acquisition structure (G), synchronization or alignment (Q), and sensor-dependent acquisition limitations (R);
- Analysis of normalized covariate value distributions, enabling concrete comparison of how datasets differ within each taxonomy attribute.
7. Structural Gaps and Research Implications
7.1. Application-Dependent Ecological Diversity in Acquisition Settings (A, B, D)
7.2. Uneven Viewpoint Coverage Across Domains (E, P)
7.3. Scarcity of Temporal Acquisition Structure and Longitudinal Protocols (G)
7.4. Domain-Specific Coverage of User-Related Covariates (H, J, K)
7.5. Incomplete Demographic Characterization (L)
7.6. Modality Concentration Around Silhouette-Based Representations (M)
7.7. Incomplete Reporting of Synchronization and Alignment (Q)
7.8. Underreported Sensor-Dependent Acquisition Limitations (R)
7.9. Dataset Access, Annotation Traceability, and Documentation Quality
7.10. Lack of Standardized Reporting Across Datasets
7.11. Synthesis
8. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Appendix A. List of Included Datasets
| Dataset | Year | URL | Access Date |
|---|---|---|---|
| Gait-A Database [43] | 2016 | http://hdl.handle.net/10045/70567 | 1 June 2026 |
| GAIT-IST [14] | 2020 | http://www.img.lx.it.pt/GAIT-IST/ | 1 June 2026 |
| GAIT-IT [17] | 2021 | http://www.img.lx.it.pt/GAIT-IT/ | 1 June 2026 |
| Health&Gait [19] | 2025 | https://zenodo.org/records/14039922 | 1 June 2026 |
| INIT Gait Database [15] | 2016 | https://www.vision.uji.es/gaitDB/ | 1 June 2026 |
| KOA-PD-NM [18] | 2020 | https://data.mendeley.com/datasets/44pfnysy89/1 | 1 June 2026 |
| MMGS [44] | 2019 | https://github.com/margokhokhlova/LSTM_gait_model | 1 June 2026 |
| ProGait [45] | 2025 | https://huggingface.co/datasets/ericyxy98/ProGait | 1 June 2026 |
| Scoliosis1K Dataset [20] | 2024 | https://zhouzi180.github.io/Scoliosis1K/ | 1 June 2026 |
| SPHERE Walking Dataset [16] | 2014 | https://doi.org/10.5523/bris.bgresiy3olk41nilo7k6xpkqf | 1 June 2026 |
| Walking Gait Dataset [46] | 2018 | https://www-labs.iro.umontreal.ca/~labimage/GaitDataset/ | 1 June 2026 |
| Dataset | Year | URL | Access Date |
|---|---|---|---|
| Multi-Attribute Gait [38] | 2022 | https://doi.org/10.1109/TIFS.2023.3318934 | 1 June 2026 |
| OU-ISIR LP Age [40] | 2017 | http://www.am.sanken.osaka-u.ac.jp/BiometricDB/GaitLPAge.html | 1 June 2026 |
| RA-GAR [39] | 2025 | https://github.com/BNU-IVC/RA-GAR | 1 June 2026 |
References
- dos Santos, C.F.G.; de Souza Oliveira, D.; Passos, L.A.; Pires, R.G.; Santos, D.F.S.; Valem, L.P.; Moreira, T.P.; Santana, M.C.S.; Roder, M.; Papa, J.P.; et al. Gait Recognition Based on Deep Learning: A Survey. ACM Comput. Surv. 2022, 55, 34. [Google Scholar] [CrossRef] [Scilit]
- Sepas-Moghaddam, A.; Etemad, A. Deep Gait Recognition: A Survey. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 264–284. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shen, C.; Yu, S.; Wang, J.; Huang, G.Q.; Wang, L. A Comprehensive Survey on Deep Gait Recognition: Algorithms, Datasets, and Challenges. IEEE Trans. Biom. Behav. Identity Sci. 2025, 7, 270–292. [Google Scholar] [CrossRef] [Scilit]
- Nambiar, A.; Bernardino, A.; Nascimento, J.C. Gait-based Person Re-identification: A Survey. ACM Comput. Surv. 2019, 52, 33. [Google Scholar] [CrossRef] [Scilit]
- Han, X.; Guffanti, D.; Brunete, A. A Comprehensive Review of Vision-Based Sensor Systems for Human Gait Analysis. Sensors 2025, 25, 498. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nunes, J.F.; Moreira, P.M.; Tavares, J.M.R.S. Benchmark RGB-D Gait Datasets: A Systematic Review. In VipIMAGE 2019; Springer International Publishing: Cham, Switzerland, 2019; pp. 366–372. [Google Scholar] [CrossRef] [Scilit]
- Bouchrika, I.; Goffredo, M.; Carter, J.N.; Nixon, M.S. Covariate Analysis for View-Point Independent Gait Recognition. In Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2009; pp. 990–999. [Google Scholar] [CrossRef] [Scilit]
- Zou, S.; Fan, C.; Xiong, J.; Shen, C.; Yu, S.; Tang, J. Cross-Covariate Gait Recognition: A Benchmark. Proc. AAAI Conf. Artif. Intell. 2024, 38, 7855–7863. [Google Scholar] [CrossRef] [Scilit]
- Parashar, A.; Shekhawat, R.S.; Ding, W.; Rida, I. Intra-class variations with deep learning-based gait analysis: A comprehensive survey of covariates and methods. Neurocomputing 2022, 505, 315–338. [Google Scholar] [CrossRef] [Scilit]
- Bukhari, M.; Bajwa, K.B.; Gillani, S.; Maqsood, M.; Durrani, M.Y.; Mehmood, I.; Ugail, H.; Rho, S. An Efficient Gait Recognition Method for Known and Unknown Covariate Conditions. IEEE Access 2021, 9, 6465–6477. [Google Scholar] [CrossRef] [Scilit]
- Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nunes, J.F.; Moreira, P.M.; Tavares, J.M.R.S. Structuring Variability in Human Gait Datasets: A Covariate-Centered Taxonomy and Systematic Review of Image- and Depth-Based Collections. OSF 2026. [Google Scholar] [CrossRef]
- Nunes, J.F.; Moreira, P.M.; Tavares, J.M.R.S. GRIDDS—A Gait Recognition Image and Depth Dataset. In VipIMAGE 2019; Springer International Publishing: Cham, Switzerland, 2019; pp. 343–352. [Google Scholar] [CrossRef] [Scilit]
- Loureiro, J.; Correia, P.L. Using a Skeleton Gait Energy Image for Pathological Gait Classification. In Proceedings of the 15th IEEE International Conference on Automatic Face and Gesture Recognition, Buenos Aires, Argentina, 16–20 November 2020; pp. 410–414. [Google Scholar] [CrossRef] [Scilit]
- Ortells, J.; Herrero-Ezquerro, M.T.; Mollineda, R.A. Vision-based gait impairment analysis for aided diagnosis. Med. Biol. Eng. Comput. 2018, 56, 1553–1564. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Paiement, A.; Tao, L.; Camplani, M.; Hannuna, S.; Damen, D.; Mirmehdi, M. Online quality assessment of human motion from skeleton data. In Proceedings of the British Machine Vision Conference; British Machine Vision Association: Durham, UK, 2014; pp. 79.1–79.12. [Google Scholar] [CrossRef] [Scilit]
- Albuquerque, P.; Machado, J.P.; Verlekar, T.T.; Correia, P.L.; Soares, L.D. Remote Gait Type Classification System Using Markerless 2D Video. Diagnostics 2021, 11, 1824. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kour, N.; Sunanda; Arora, S. A Vision-Based Gait Dataset for Knee Osteoarthritis and Parkinson’s Disease Analysis with Severity Levels. In International Conference on Innovative Computing and Communications; Springer: Singapore, 2021; pp. 303–317. [Google Scholar] [CrossRef] [Scilit]
- Zafra-Palma, J.; Marín-Jiménez, N.; Castro-Piñero, J.; Cuenca-García, M.; Muñoz-Salinas, R.; Marín-Jiménez, M.J. Health & Gait: A dataset for gait-based analysis. Sci. Data 2025, 12, 44. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhou, Z.; Liang, J.; Peng, Z.; Fan, C.; An, F.; Yu, S. Gait Patterns as Biomarkers: A Video-Based Approach for Classifying Scoliosis. arXiv 2024, arXiv:2407.05726. [Google Scholar] [CrossRef] [Scilit]
- Wang, L.; Tan, T.; Ning, H.; Hu, W. Silhouette Analysis-Based Gait Recognition for Human Identification. IEEE Trans. Pattern Anal. Mach. Intell. 2003, 25, 1505–1518. [Google Scholar] [CrossRef] [Scilit]
- Gross, R.; Shi, J. The CMU Motion of Body (MoBo) Database; Technical Report CMU-RI-TR-01-18; Carnegie Mellon University: Pittsburgh, PA, USA, 2001. [Google Scholar]
- Iwashita, Y.; Baba, R.; Ogawara, K.; Kurazume, R. Person Identification from Spatio-temporal 3D Gait. In Proceedings of the International Conference on Emerging Security Technologies; IEEE: New York, NY, USA, 2010. [Google Scholar] [CrossRef] [Scilit]
- Iwashita, Y.; Ogawara, K.; Kurazume, R. Identification of people walking along curved trajectories. Pattern Recognit. Lett. 2014, 48, 60–69. [Google Scholar] [CrossRef] [Scilit]
- Yu, S.; Tan, D.; Tan, T. A Framework for Evaluating the Effect of View Angle, Clothing and Carrying Condition on Gait Recognition. In Proceedings of the 18th International Conference on Pattern Recognition, Hong Kong, China, 20–24 August 2006; pp. 441–444. [Google Scholar] [CrossRef] [Scilit]
- Song, C.; Huang, Y.; Wang, W.; Wang, L. CASIA-E: A Large Comprehensive Dataset for Gait Recognition. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 2801–2815. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shen, C.; Chao, F.; Wu, W.; Wang, R.; Huang, G.Q.; Yu, S. LidarGait: Benchmarking 3D Gait Recognition with Point Clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2023; pp. 1054–1063. [Google Scholar] [CrossRef] [Scilit]
- Zhu, Z.; Guo, X.; Yang, T.; Huang, J.; Deng, J.; Huang, G.; Du, D.; Lu, J.; Zhou, J. Gait Recognition in the Wild: A Benchmark. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; pp. 14789–14799. [Google Scholar]
- Uddin, M.Z.; Ngo, T.T.; Makihara, Y.; Takemura, N.; Li, X.; Muramatsu, D.; Yagi, Y. The OU-ISIR Large Population Gait Database with Real-Life Carried Object and its Performance Evaluation. IPSJ Trans. Comput. Vis. Appl. 2018, 10, 5. [Google Scholar] [CrossRef] [Scilit]
- Zheng, J.; Liu, X.; Wang, S.; Wang, L.; Yan, C.; Liu, W. Parsing is All You Need for Accurate Gait Recognition in the Wild. In Proceedings of the 31st ACM International Conference on Multimedia, MM’23; ACM: New York, NY, USA, 2023; pp. 116–124. [Google Scholar] [CrossRef] [Scilit]
- Takemura, N.; Makihara, Y.; Muramatsu, D.; Echigo, T.; Yagi, Y. Multi-view large population gait dataset and its performance evaluation for cross-view gait recognition. IPSJ Trans. Comput. Vis. Appl. 2018, 10, 4. [Google Scholar] [CrossRef] [Scilit]
- An, W.; Yu, S.; Makihara, Y.; Wu, X.; Xu, C.; Yu, Y.; Liao, R.; Yagi, Y. Performance Evaluation of Model-Based Gait on Multi-View Very Large Population Database With Pose Sequences. IEEE Trans. Biom. Behav. Identity Sci. 2020, 2, 421–430. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Makihara, Y.; Xu, C.; Yagi, Y. Multi-View Large Population Gait Database With Human Meshes and Its Performance Evaluation. IEEE Trans. Biom. Behav. Identity Sci. 2022, 4, 234–248. [Google Scholar] [CrossRef] [Scilit]
- Shehata, A.; Castro, F.M.; Guil, N.; Marín-Jiménez, M.J.; Yagi, Y. OUMVLP-OF: Multi-View Large Population Gait Database With Dense Optical Flow and Its Performance Evaluation. IEEE Access 2025, 13, 87100–87111. [Google Scholar] [CrossRef] [Scilit]
- Tsuji, A.; Makihara, Y.; Yagi, Y. Silhouette transformation based on walking speed for gait identification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2010; pp. 717–722. [Google Scholar] [CrossRef] [Scilit]
- Altab Hossain, M.; Makihara, Y.; Wang, J.; Yagi, Y. Clothing-invariant gait identification using part-based clothing categorization and adaptive weight control. Pattern Recognit. 2010, 43, 2281–2291. [Google Scholar] [CrossRef] [Scilit]
- Mori, A.; Makihara, Y.; Yagi, Y. Gait Recognition using Period-Based Phase Synchronization for Low Frame-Rate Videos. In Proceedings of the 20th International Conference on Pattern Recognition; IEEE: New York, NY, USA, 2010; pp. 2194–2197. [Google Scholar] [CrossRef] [Scilit]
- Song, X.; Hou, S.; Huang, Y.; Cao, C.; Liu, X.; Huang, Y.; Shan, C. Gait Attribute Recognition: A New Benchmark for Learning Richer Attributes from Human Gait Patterns. IEEE Trans. Inf. Forensics Secur. 2024, 19, 1–14. [Google Scholar] [CrossRef] [Scilit]
- Wang, C.; Hou, S.; Li, A.; Cai, Q.; Huang, Y. RA-GAR: A Richly Annotated Benchmark for Gait Attribute Recognition. Proc. AAAI Conf. Artif. Intell. 2025, 39, 7591–7599. [Google Scholar] [CrossRef] [Scilit]
- Xu, C.; Makihara, Y.; Ogi, G.; Li, X.; Yagi, Y.; Lu, J. The OU-ISIR Gait Database Comprising the Large Population Dataset with Age and Performance Evaluation of Age Estimation. IPSJ Trans. Comput. Vis. Appl. 2017, 9, 24. [Google Scholar] [CrossRef] [Scilit]
- Chaquet, J.M.; Carmona, E.J.; Fernández-Caballero, A. A Survey of Video Datasets for Human Action and Activity Recognition. Comput. Vis. Image Underst. 2013, 117, 633–659. [Google Scholar] [CrossRef] [Scilit]
- Singh, J.P.; Jain, S.; Arora, S.; Singh, U.P. Vision-Based Gait Recognition: A Survey. IEEE Access 2018, 6, 70497–70527. [Google Scholar] [CrossRef] [Scilit]
- Nieto-Hidalgo, M.; Ferrández-Pastor, F.J.; Valdivieso-Sarabia, R.J.; Mora-Pascual, J.; García-Chamizo, J.M. A vision based proposal for classification of normal and abnormal gait using RGB camera. J. Biomed. Inform. 2016, 63, 82–89. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Khokhlova, M.; Migniot, C.; Morozov, A.; Sushkova, O.; Dipanda, A. Normal and pathological gait classification LSTM model. Artif. Intell. Med. 2019, 94, 54–66. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yin, X.; Yang, B.; Liu, W.; Xue, Q.; Alamri, A.; Fiedler, G.; Gao, W. ProGait: A Multi-Purpose Video Dataset and Benchmark for Transfemoral Prosthesis Users. arXiv 2025, arXiv:2507.10223. [Google Scholar] [CrossRef] [Scilit]
- Nguyen, T.N.; Huynh, H.H.; Meunier, J. 3D Reconstruction with Time-of-Flight Depth Camera and Multiple Mirrors. IEEE Access 2018, 6, 38106–38114. [Google Scholar] [CrossRef] [Scilit]
- Fang, H.S.; Li, J.; Tang, H.; Xu, C.; Zhu, H.; Xiu, Y.; Li, Y.L.; Lu, C. AlphaPose: Whole-Body Regional Multi-Person Pose Estimation and Tracking in Real-Time. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 7157–7173. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shotton, J.; Fitzgibbon, A.; Cook, M.; Sharp, T.; Finocchio, M.; Moore, R.; Kipman, A.; Blake, A. Real-time human pose recognition in parts from single depth images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; CVPR; IEEE Computer Society: New York, NY, USA, 2011; pp. 1297–1304. [Google Scholar] [CrossRef] [Scilit]
- Cao, Z.; Hidalgo, G.; Simon, T.; Wei, S.E.; Sheikh, Y. OpenPose: Realtime Multi-Person 2D Pose Estimation Using Part Affinity Fields. IEEE Trans. Pattern Anal. Mach. Intell. 2021, 43, 172–186. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Jiang, T.; Lu, P.; Zhang, L.; Ma, N.; Han, R.; Lyu, C.; Li, Y.; Chen, K. RTMPose: Real-Time Multi-Person Pose Estimation based on MMPose. arXiv 2023, arXiv:2303.07399. [Google Scholar] [CrossRef] [Scilit]
- Topham, L.K.; Khan, W.; Al-Jumeily, D.; Waraich, A.; Hussain, A.J. A diverse and multi-modal gait dataset of indoor and outdoor walks acquired using multiple cameras and sensors. Sci. Data 2023, 10, 320. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- López-Fernández, D.; Madrid-Cuevas, F.J.; Carmona-Poyato, Á.; Marín-Jiménez, M.J.; noz Salinas, R.M. The AVA Multi-View Dataset for Gait Recognition. In Activity Monitoring by Multiple Distributed Sensing; Mazzeo, P.L., Spagnolo, P., Moeslund, T.B., Eds.; Lecture Notes in Computer Science; Springer International Publishing: Cham, Switzerland, 2014; pp. 26–39. [Google Scholar] [CrossRef] [Scilit]
- Tan, D.; Huang, K.; Yu, S.; Tan, T. Efficient Night Gait Recognition Based on Template Matching. In Proceedings of the 18th International Conference on Pattern Recognition; IEEE: New York, NY, USA, 2006. [Google Scholar] [CrossRef] [Scilit]
- Li, W.; Hou, S.; Zhang, C.; Cao, C.; Liu, X.; Huang, Y.; Zhao, Y. An In-Depth Exploration of Person Re-Identification and Gait Recognition in Cloth-Changing Conditions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2023; pp. 13824–13833. [Google Scholar] [CrossRef] [Scilit]
- Chattopadhyay, P.; Sural, S.; Mukherjee, J. Frontal gait recognition from occluded scenes. Pattern Recognit. Lett. 2015, 63, 9–15. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Z.; Tran, L.; Yin, X.; Atoum, Y.; Liu, X.; Wan, J.; Wang, N. Gait Recognition via Disentangled Representation Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2019; pp. 4705–4714. [Google Scholar] [CrossRef] [Scilit]
- Iwashita, Y.; Kurazume, R.; Stoica, A. Gait Identification Using Invisible Shadows: Robustness to Appearance Changes. In Proceedings of the 5th International Conference on Emerging Security Technologies; IEEE: New York, NY, USA, 2014; pp. 34–39. [Google Scholar] [CrossRef] [Scilit]
- Hou, S.; Fan, C.; Cao, C.; Liu, X.; Huang, Y. A Comprehensive Study on the Evaluation of Silhouette-Based Gait Recognition. IEEE Trans. Biom. Behav. Identity Sci. 2023, 5, 196–208. [Google Scholar] [CrossRef] [Scilit]
- Huang, P.; Peng, Y.; Hou, S.; Cao, C.; Liu, X.; He, Z.; Huang, Y. Occluded Gait Recognition with Mixture of Experts: An Action Detection Perspective. In Computer Vision—ECCV 2024; Springer Nature: Cham, Switzerland, 2024; pp. 380–397. [Google Scholar] [CrossRef] [Scilit]
- Mansur, A.; Makihara, Y.; Aqmar, R.; Yagi, Y. Gait recognition under speed transition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2014; pp. 2521–2528. [Google Scholar] [CrossRef] [Scilit]
- Mu, Z.; Castro, F.M.; Marin-Jimenez, M.J.; Guil, N.; Li, Y.R.; Yu, S. ReSGait: The Real-Scene Gait Dataset. In Proceedings of the IEEE International Joint Conference on Biometrics (IJCB); IEEE: New York, NY, USA, 2021. [Google Scholar] [CrossRef] [Scilit]
- Sivapalan, S.; Chen, D.; Denman, S.; Sridharan, S.; Fookes, C. Gait energy volumes and frontal gait recognition using depth images. In Proceedings of the International Joint Conference on Biometrics; IEEE: New York, NY, USA, 2011; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- Yin, Y.; Liu, L.; Sun, X. SDUMLA-HMT: A Multimodal Biometric Database. In Biometric Recognition; Springer: Berlin/Heidelberg, Germany, 2011; pp. 260–268. [Google Scholar] [CrossRef] [Scilit]
- Singh, J.P.; Arora, S.; Jain, S.; Singh SoM, U.P. A Multi-Gait Dataset for Human Recognition under Occlusion Scenario. In Proceedings of the International Conference on Issues and Challenges in Intelligent Computing Techniques (ICICT); IEEE: New York, NY, USA, 2019; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- Shutler, J.D.; Grant, M.G.; Nixon, M.S.; Carter, J.N. On a Large Sequence-Based Human Gait Database. In Applications and Science in Soft Computing; Springer: Berlin/Heidelberg, Germany, 2004; pp. 339–346. [Google Scholar] [CrossRef] [Scilit]
- Yu, S.; Wang, Q.; Huang, Y. A Large RGB-D Gait Dataset and the Baseline Algorithm. In Biometric Recognition; Springer International Publishing: Cham, Switzerland, 2013; pp. 417–424. [Google Scholar] [CrossRef] [Scilit]
- Hofmann, M.; Sural, S.; Rigoll, G. Gait recognition in the presence of occlusion: A new dataset and baseline algorithms. In Proceedings of the 19th International Conferences on Computer Graphics, Visualization and Computer Vision (WSCG), Plzeň, Czech Republic, 31 January–3 February 2011; pp. 99–104. [Google Scholar]
- Sun, K.; Xiao, B.; Liu, D.; Wang, J. Deep High-Resolution Representation Learning for Human Pose Estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2019; pp. 5686–5696. [Google Scholar] [CrossRef] [Scilit]
- Martinez, J.; Hossain, R.; Romero, J.; Little, J.J. A Simple Yet Effective Baseline for 3d Human Pose Estimation. In Proceedings of the IEEE International Conference on Computer Vision (ICCV); IEEE: New York, NY, USA, 2017; pp. 2659–2668. [Google Scholar] [CrossRef] [Scilit]
- Zhu, W.; Ma, X.; Liu, Z.; Liu, L.; Wu, W.; Wang, Y. MotionBERT: A Unified Perspective on Learning Human Motion Representations. arXiv 2022, arXiv:2210.06551. [Google Scholar] [CrossRef] [Scilit]





| Work | Primary Organizing Principle | Relation to the Present Work |
|---|---|---|
| Sepas-Moghaddam and Etemad [2] | Body, temporal, feature, and architectural representations | Focuses on deep recognition models rather than dataset-level covariate coverage. |
| Shen et al. [3] | Deep representations, architectures, benchmarks, and challenges | Surveys deep gait recognition, treating datasets mainly as evaluation benchmarks. |
| Santos et al. [1] | Deep learning methods, datasets, architectures, and limitations | Emphasizes recognition pipelines and performance, not dataset variability or reporting consistency. |
| Nambiar et al. [4] | Gait-based person re-identification methods and evaluation | Focuses on re-identification rather than cross-domain dataset covariate characterization. |
| Parashar et al. [9] | Covariates and deep learning strategies for handling them | Addresses covariates mainly as robustness challenges for recognition algorithms. |
| Han et al. [5] | Vision sensors, machine learning methods, and applications | Reviews sensor systems and applications, not dataset-level covariate taxonomy. |
| Nunes et al. [6] | Depth/RGB-D gait dataset properties and availability | Reviews RGB-D gait datasets, while the present work covers broader modalities and domains. |
| Present work | Scene-level, user-level, and sensor-level covariate coverage | Treats dataset variability as the primary object of analysis and provides a structured taxonomy for dataset characterization, comparison, documentation, and benchmark design. |
| Dimension | Code | Taxonomy Attribute |
|---|---|---|
| Scene-level | A–G | Acquisition environment; background and scene complexity; walking surface or terrain; environmental control; viewpoint configuration; walking path or trajectory; temporal acquisition structure. |
| User-level | H–L | Activity or gait condition; speed or pace condition; clothing or appearance condition; carrying/load condition; participant descriptors. |
| Sensor-level | M–R | Data modality or representation; frame rate or temporal sampling; spatial resolution; acquisition setup or sensor placement; synchronization or alignment; sensor-dependent acquisition limitations. |
| Dataset | Year | Subjects | Scene-Level | User-Level | Sensor-Level |
|---|---|---|---|---|---|
| Gait-A Database [43] | 2016 | 5 (73 samples: 38 normal; 35 abnormal) | (A) Indoor (corridor) (B) Static background (C) Planar (overground) (D) Controlled (E) Multi-view (2 views: side; front) (F) Straight trajectory (G) Single-session | (H) Walking (Normal; Simulated abnormal gait) (I) Self-selected (J) Natural clothing (K) No load (L) Not reported | (M) SIL (N) 30 fps (O) 1920 × 1080 (P) Fixed single camera (Q) Not applicable (R) Not reported |
| GAIT-IST [14] | 2020 | 10 (8 ♂–2 ♀) | (A) Indoor (laboratory setting) (B) Static background (C) Planar (overground) (D) Controlled (E) Single side view (F) Straight bidirectional walking path (G) Single-session | (H) Walking (Normal; Diplegic; Hemiplegic; Neuropathic; Parkinsonian) (I) Self-selected (J) Natural clothing (K) No load (L) Age; Sex | (M) SIL; GEI; SEI; 2D-KPT (OpenPose) (N) 25 fps (O) 224 × 224 (P) Fixed single camera (1.5 m height; ≈4 m distance) (Q) Not applicable (R) Not reported |
| GAIT-IT [17] | 2021 | 21 (19 ♂–2 ♀) | (A) Indoor (laboratory setting) (B) Static background (green chroma-key) (C) Planar (overground) (D) Controlled (E) Multi-view (2 views: side; front) (F) Straight bidirectional walking path (G) Single-session | (H) Walking (Normal; Scissor; Spastic; Steppage; Propulsive) (I) Self-selected (J) Natural clothing (K) No load (L) Age; Sex | (M) SIL; GEI; SEI (N) 10 fps (O) 224 × 224 (P) Fixed 2 cameras (1.75 m height) (Q) Synchronized calibrated multi-camera setup (R) Not reported |
| Health&Gait [19] | 2025 | 398 (199 ♂–199 ♀) | (A) Indoor (laboratory setting) (B) Static background (C) Planar (overground) (D) Controlled (E) Single side view (F) Straight bidirectional walking path (G) Eight-month collection | (H) Walking (I) Self-selected; Fast (J) Natural clothing; Jacket condition (K) No load (L) Age; Sex; Height; Weight; Physical activity level | (M) SIL; 2D-KPT (AlphaPose); 2D-PRS; OF (N) 30 fps (O) SIL: 960 × 540; OF: 480 × 270 (P) Fixed single camera (placement varied by session) (Q) OptoGait/MuscleLAB reference data; DSU 1 ms for photocells (R) Not reported |
| INIT Gait Database [15] | 2016 | 10 (9 ♂–1 ♀) | (A) Indoor (laboratory setting) (B) Static background (green chroma-key) (C) Planar (overground) (D) Controlled (E) Single side view (F) Straight trajectory (G) Single-session | (H) Walking (Normal; 7 simulated impaired gait styles) (I) Self-selected (J) Natural clothing (K) No load (L) Sex | (M) SIL (N) 15 fps (O) 800 × 400 (P) Fixed single camera (Q) Not applicable (R) Not reported |
| KOA-PD-NM [18] | 2020 | 96 (50 KOA; 16 PD; 30 NM) | (A) Indoor (clinical setting) (B) Static background (green chroma-key) (C) Planar (overground) (D) Controlled (E) Single side view (F) Straight bidirectional walking path (G) 2018–2019 collection | (H) Walking (Normal; Knee Osteoarthritis; Parkinsonian) (I) Self-selected (J) Natural clothing (K) No load (L) Age; Sex; Height; Clinical severity | (M) RGB (N) 50 fps (O) 1920 × 1080 (P) Fixed single camera (8 m from walking mat) (Q) Not applicable (R) Not reported |
| MMGS [44] | 2019 | 27 (19 ♂–8 ♀) | (A) Indoor (laboratory setting) (B) Static background (C) Planar (overground) (D) Controlled (E) Single frontal view (F) Straight trajectory (G) Single-session | (H) Walking (Normal; Simulated limp; Simulated knee rigidity) (I) Self-selected (J) Natural clothing; Shoes condition (K) No load (L) Age; Sex; Height; Weight | (M) DPT; SIL; 3D-KPT (Kinect SDK) (N) 30 fps (O) DEP: 512 × 424 (P) Fixed single Kinect v2 (≈2.0 m height) (Q) Timestamped multimodal streams (R) Not reported |
| ProGait [45] | 2025 | 4 (4 ♂) | (A) Indoor (clinical setting) (B) Dynamic background (healthcare staff present) (C) Planar (overground) (D) Controlled (mixed indoor lighting) (E) Multi-view (2 views: frontal; sagittal) (F) Straight trajectory (assisted walking) (G) Single-session | (H) Walking (Prosthetic/Pathological) (I) Self-selected (J) Natural clothing (K) No load (L) Age; Sex | (M) RGB; 2D-KPT (Assisted Manual Annotation); 2D-PRS (N) 30 fps (O) 1920 × 1080 (P) Fixed 2 cameras (Q) Synchronized calibrated multi-camera setup (R) Not reported |
| Scoliosis1K [20] | 2024 | 1050 (641 ♀–409 ♂) | (A) Indoor (corridor) (B) Static background (C) Planar (overground) (D) Controlled (E) Single-view (F) Straight trajectory (G) Single-session | (H) Walking (Scoliosis screening classes: positive; neutral; negative) (I) Self-selected (J) Natural clothing (K) No load (L) Age; Sex; Height; Weight | (M) SIL; 2D-KPT (RTMPose) (N) 15 fps (O) 1280 × 720 (P) Fixed single camera (1.4–4.2 m from participants) (Q) Not applicable (R) Not reported |
| SPHERE Walking Dataset [16] | 2014 | 12 | (A) Indoor (stairs setup) (B) Static background (C) Stairs (D) Controlled (E) Single frontal view (F) Straight trajectory (G) Single-session | (H) Stair ascent (Normal; Left-leg lead; Right-leg lead; Freezing) (I) Self-selected (J) Natural clothing (K) No load (L) Not reported | (M) DPT; 3D-KPT (OpenNI) (N) 30 fps (O) 480 × 320 (P) Fixed single Kinect v1 (Q) Timestamped multimodal streams (R) Not reported |
| Walking Gait Dataset [46] | 2018 | 9 | (A) Indoor (laboratory setting) (B) Static background (C) Treadmill (planar) (D) Controlled (E) Single frontal view (F) Treadmill-constrained (G) Single-session | (H) Walking (Normal; 8 simulated asymmetric gait conditions) (I) Speed-controlled (J) Natural clothing (K) No load (L) Not reported | (M) SIL; 3D-KPT (Kinect SDK); 3D-PCL (N) 30 fps (O) 512 × 424 (P) Fixed single Kinect v2 (Q) Timestamped multimodal streams (R) Not reported |
| Dataset | Year | Subjects | Scene-Level | User-Level | Sensor-Level |
|---|---|---|---|---|---|
| 360 Degree Gait Capture [51] | 2022 | 65 (38 ♂–27 ♀) | (A) Indoor (laboratory setting); Outdoor (B) Static background (C) Planar (overground) (D) Uncontrolled (natural illumination; illumination changes) (E) Multi-view (8 views: 0°–360°; 45° increments) (F) Straight bidirectional walking path (G) Collection period reported | (H) Walking (I) Self-selected (J) Clothing variation (K) No load (L) Age; Sex; Height; Mass; Ethnicity | (M) RGB; 2D-KPT (OpenPose); GD (N) 30 fps (O) RGB: 1280 × 720 (P) Fixed 2 cameras (layouts varied across experiments) (Q) Synchronized calibrated multi-camera setup (R) Not reported |
| AVA Multi-View Dataset [52] | 2013 | 20 (16 ♂–4 ♀) | (A) Indoor (studio-controlled) (B) Static background (C) Planar (overground) (D) Controlled (E) Multi-view (6 views) (F) Straight; Curved; Figure-eight trajectories (G) Single-session | (H) Walking (I) Self-selected (J) Natural clothing (K) No load (L) Sex | (M) RGB; SIL (N) 25 fps (O) 640 × 480 (P) Fixed 6 cameras (Q) Synchronized calibrated multi-camera setup (R) Not reported |
| CASIA-A [21] | 2001 | 20 | (A) Outdoor (B) Static background (C) Planar (overground) (D) Uncontrolled (natural illumination; illumination changes) (E) Single-view (F) Straight trajectories (frontal; oblique; lateral) (G) Single-session | (H) Walking (I) Self-selected (J) Natural clothing (K) No load (L) Not reported | (M) RGB (N) 25 fps (O) 352 × 240 (P) Fixed single camera (Q) Not applicable (R) Not reported |
| CASIA-B [25] | 2005 | 124 (93 ♂–31 ♀) | (A) Indoor (laboratory setting) (B) Static background (C) Planar (overground) (D) Controlled (E) Multi-view (11 views: 0°–180°; 18° steps) (F) Straight trajectory (G) Single-session | (H) Walking (I) Self-selected (J) Natural clothing; Coat condition (K) No load; Bag (L) Sex; Height | (M) RGB; SIL (N) 25 fps (O) 320 × 240 (P) Fixed 11-camera arc (18° spacing) (Q) Not synchronized (R) Frame loss |
| CASIA-C [53] | 2005 | 153 (130 ♂–23 ♀) | (A) Outdoor (night-time) (B) Static background (C) Planar (overground) (D) Uncontrolled (night-time thermal) (E) Single side view (F) Straight trajectory (G) Single-session | (H) Walking (I) Self-selected; Slow; Fast (J) Natural clothing (K) No load; Bag (L) Sex | (M) SIL; GEI; IR (N) 25 fps (O) SIL: 129 × 130; IR: 320 × 240 (P) Fixed single camera (thermal) (Q) Not applicable (R) Not reported |
| CASIA-E [26] | 2022 | 1014 (507 ♂–507 ♀) | (A) Outdoor (multiple scenes) (B) Static background; Dynamic background (C) Planar (overground) (D) Uncontrolled (natural illumination; illumination changes) (E) Multi-view (26 views: 13 horizontal × 2 vertical) (F) Straight trajectories (multiple walking lines) (G) Five-month collection; seasonal variation | (H) Walking (stop/non-stop walking style) (I) Self-selected (J) Natural clothing; Coat condition (K) No load; Bag (L) Age; Sex; Height; Weight;Nationality | (M) SIL; GEI; IR (N) 25 fps (O) SIL: 1920 × 1080; IR: 640 × 480 (P) Fixed 8 cameras (2 heights: 1.2/3.5 m) (Q) Not reported (R) Not reported |
| CCGR [8] | 2024 | 970 | (A) Indoor (laboratory setting) (B) Static background (C) Planar; stairs; ramp; bumpy; soft; curved road (D) Controlled (E) Multi-view (33 views; 5 pitch layers) (F) Straight; Curved; Stairs; Ramps; Mixed task routes (G) 20-month collection | (H) Walking (I) Self-selected; Fast; Stationary (J) Natural clothing; Coat condition (K) No load; Book; Bag; Heavy bag; Box; Heavy box; Trolley; Umbrella (L) Age; Sex | (M) SIL; 2D-KPT (HRNet); 2D-PRS (N) 25 fps (O) 1280 × 720 (P) Fixed 33 cameras (Q) Not reported (R) Not reported |
| CCPG: Cloth-Changing benchmark for Person re-identification and Gait recognition [54] | 2023 | 200 (122 ♂–78 ♀) | (A) Indoor; Outdoor (B) Dynamic background (static obstacles) (C) Planar (overground) (D) Uncontrolled (illumination changes) (E) Multi-view (10 views) (F) Cross-scene route; Square walking route (G) Single-session | (H) Walking (I) Self-selected (J) Clothing variation (whole-, upper-, lower-body changes) (K) No load; Bag (L) Sex | (M) RGB; SIL (N) 25 fps (O) RGB: 256 × 128; SIL: 128 × 88 (P) Fixed 10 cameras (2.7–3.0 m height) (Q) Not reported (R) Not reported |
| CMU Motion of Body (MoBo) [22] | 2001 | 25 (23 ♂–2 ♀) | (A) Indoor (laboratory setting) (B) Static background (C) Treadmill (planar and incline) (D) Controlled (E) Multi-view (6 views) (F) Treadmill-constrained (G) Single-session | (H) Walking (I) Self-selected; Slow; Fast (J) Natural clothing (K) No load; Ball (L) Age; Sex; Weight | (M) RGB; SIL (N) 30 fps (O) 640 × 480 (P) Fixed 6 cameras (Q) Synchronized calibrated multi-camera setup (R) Not reported |
| Depth-Based Gait Dataset [55] | 2015 | 29 | (A) Indoor (laboratory setting) (B) Static background (person-induced occlusion) (C) Planar (overground) (D) Controlled (E) Multi-view (2 views: front, back) (F) Straight trajectory (G) Single-session | (H) Walking (I) Self-selected; Fast (J) Natural clothing (K) No load (L) Not reported | (M) SIL; DPT; 2D-KPT (Kinect SDK); 3D-KPT (Kinect SDK) (N) 30 fps; 15 fps (O) SIL: 320 × 240 (P) Fixed 2 Kinect v1 (2.5 m height; −27° tilt) (Q) Timestamped multimodal streams (R) Not reported |
| Frontal-View Gait (FVG-B) [56] | 2018 | 226 | (A) Outdoor (B) Static background (person-induced occlusion) (C) Planar (overground) (D) Uncontrolled (natural illumination) (E) Multi-view (3 views: −45°; 0°; +45°) (F) Straight trajectory (G) Multi-session/time gap | (H) Walking (I) Self-selected; Slow; Fast (J) Natural clothing; Footwear variation; Hat condition (K) No load; Bag (L) Not reported | (M) RGB (N) 15 fps (O) 1920 × 1080 (P) Fixed single camera (≈1.5 m height) (Q) Not applicable (R) Not reported |
| Gait3D-Parsing [30] | 2023 | 4000 | (A) Indoor (in-the-wild supermarket) (B) Dynamic background (scene-induced occlusion) (C) Planar (overground) (D) Uncontrolled (natural illumination) (E) Multi-view (39 views) (F) Unconstrained trajectories (G) Not reported | (H) Walking (I) Self-selected (J) Natural clothing (K) No load; Bag; Accessories (L) Not reported | (M) SIL; 2D-KPT (HRNet); 3D-KPT; 3D-MSH; 2D-PRS (N) 25 fps (O) 1920 × 1080 (P) Fixed 39 surveillance cameras (Q) Not reported (R) Not reported |
| GREW [28] | 2020 | 26,345 | (A) Outdoor (in-the-wild) (B) Dynamic background (scene-induced occlusion) (C) Planar (overground) (D) Uncontrolled (natural illumination) (E) Multi-view (882 views) (F) Unconstrained trajectories (G) Single-session | (H) Walking (I) Self-selected (J) Clothing variation (K) No load; Backpack; Shoulder bag; Handbag; Lift-stuff (L) Age; Sex | (M) SIL; GEI; OF; 2D-KPT (HRNet); 3D-KPT (Lifting 3D) (N) 25 fps (O) 1920 × 1080 (P) Fixed 882 surveillance cameras (Q) Not reported (R) Not reported |
| GRIDDS [13] | 2018 | 35 (11 ♂–24 ♀) | (A) Indoor (laboratory setting) (B) Static background (C) Planar (overground) (D) Controlled (E) Single side view (F) Straight bidirectional walking path (G) Single-session | (H) Walking (I) Self-selected (J) Natural clothing (K) No load (L) Age; Sex; Height | (M) RGB; SIL; DPT; IR; 2D-KPT (Kinect SDK); 3D-KPT (Kinect SDK) (N) 30 fps (O) RGB: 1920 × 1080; DPT/IR: 512 × 424; SIL: 120 × 80 (P) Fixed single Kinect v2 (1.8 m height) (Q) Timestamped multimodal streams (R) Not reported |
| KY4D (A+B) [23,24] | 2010–2014 | 42 | (A) Indoor (laboratory setting) (B) Static background (C) Planar (overground) (D) Controlled (E) Multi-view (16 views) (F) Straight; Curved circular trajectories (G) Single-session | (H) Walking (I) Self-selected (J) Natural clothing (K) No load (L) Not reported | (M) RGB; SIL; 3D-VOL (N) 15 fps (O) 1032 × 776 (P) Fixed 16 circular camera setup (Q) Synchronized calibrated multi-camera setup (R) Not reported |
| KY IR Shadow Gait Database [57] | 2014 | 54 | (A) Indoor (laboratory setting) (B) Static background (C) Planar (overground) (D) Controlled (infrared illumination) (E) Multi-view (2 views: side-view; ceiling view) (F) Straight trajectory (G) Single-session | (H) Walking (I) Self-selected (J) Natural clothing; Down jacket; Lab coat; Coat (K) No load; Backpack; Hand bag; Travelling bag (L) Not reported | (M) SIL (N) 30 fps (O) 1600 × 1200 (P) Fixed single ceiling camera (Q) Not applicable (R) Not reported |
| Multi-Height Gait (MHG) [58] | 2023 | 200 | (A) Indoor (college gym) (B) Static background (C) Planar (overground) (D) Controlled (E) Multi-view (9 views at 45° intervals) (F) Straight trajectory (G) Single-session | (H) Walking (I) Self-selected (J) Natural clothing; Clothing variation (K) No load; Bag (L) Age group | (M) SIL (N) 25 fps (O) 128 × 88 (P) Fixed multi-camera setup (3 heights × 3 angles) (Q) Synchronized calibrated multi-camera setup (R) Not reported |
| OccGait [59] | 2025 | 101 | (A) Indoor (laboratory setting) (B) Static background (scene-induced occlusion) (C) Planar (overground) (D) Controlled (E) Multi-view (8 views) (F) Straight segments; Square walking route (G) Single-session | (H) Walking (I) Self-selected (J) Natural clothing (K) No load; Umbrella; Luggage (L) Not reported | (M) SIL (N) 25 fps (O) 1920 × 1080 (P) Fixed 3 cameras (0°; 45°; 315°) (Q) Synchronized calibrated multi-camera setup (R) Not reported |
| OU-ISIR Treadmill A [35] | 2010 | 34 (26 ♂–8 ♀) | (A) Indoor (laboratory setting) (B) Static background (C) Treadmill (planar) (D) Controlled (E) Single side view (F) Treadmill-constrained (G) Single-session | (H) Walking (I) Speed-controlled (2–10 km/h;1 km/h steps) (J) Natural clothing (K) No load (L) Sex | (M) SIL (N) 60 fps (O) 128 × 88 (P) Fixed single camera (Q) Not applicable (R) Not reported |
| OU-ISIR Treadmill B [36] | 2010 | 68 (31 ♂–37 ♀) | (A) Indoor (laboratory setting) (B) Static background (C) Treadmill (planar) (D) Controlled (E) Single side view (F) Treadmill-constrained (G) Single-session | (H) Walking (I) Not reported (J) Clothing variation (up to 32 clothing combinations) (K) No load (L) Age; Sex | (M) SIL (N) 60 fps (O) 128 × 88 (P) Fixed single camera (Q) Not applicable (R) Not reported |
| OU-ISIR Treadmill D [37] | 2010 | 185 | (A) Indoor (laboratory setting) (B) Static background (C) Treadmill (planar) (D) Controlled (E) Single side view (F) Treadmill-constrained (G) Single-session | (H) Walking (I) Speed-controlled (J) Natural clothing (K) No load (L) Not reported | (M) SIL (N) 60 fps (simulated low-FPS variants) (O) 128 × 88 (P) Fixed single camera (Q) Not applicable (R) Not reported |
| OU-ISIR Gait Speed Transition (GaitST) [60] | 2014 | 179 | (A) Indoor (laboratory setting) (B) Static background (C) Planar (overground); Treadmill (planar) (D) Controlled (E) Single side view (F) Treadmill-constrained (G) Single-session | (H) Walking (speed-transition) (I) Speed-controlled (Acceleration/De- celeration: 1–5 km/h) (J) Natural clothing (K) No load (L) Not reported | (M) SIL; GEI (N) 60 fps (O) 32 × 22 (P) Fixed single camera (Q) Not applicable (R) Not reported |
| OU-LP-Bag [29] | 2018 | 62,528 | (A) Indoor (museum exhibition setup) (B) Static background (green chroma-key) (C) Planar (overground) (D) Controlled (E) Single elevated side view (F) Straight trajectory (G) Long-run collection | (H) Walking (I) Self-selected (J) Natural clothing (K) No load; Bag (L) Age; Sex | (M) SIL; GEI (N) 25 fps (O) SIL: 1280 × 980; GEI: 128 × 88 (P) Fixed single camera (≈8 m distance; 5 m height) (Q) Not applicable (R) Not reported |
| OU-MVLP family [31,32,33,34] | 2018–2025 | 10,307 (5114 ♂–5193 ♀) | (A) Indoor (laboratory setting) (B) Static background (green chroma-key) (C) Planar (overground) (D) Controlled (E) Multi-view (14 views: 0°–90°; 180°–270°; 15° intervals) (F) Straight bidirectional walking path (G) Long-run collection; seasonal clothing variation | (H) Walking (I) Self-selected (J) Natural clothing (K) No load (L) Age; Sex | (M) SIL; GEI; 2D-KPT (OpenPose; AlphaPose); 3D-KPT; 3D-MSH; OF (N) 25 fps (O) SIL/GEI: 128 × 88; OF: 256 × 256 (P) Fixed multi-camera setup (≈8 m radius; 5 m height) (Q) Synchronized calibrated multi-camera setup (R) Not reported |
| ReSGait [61] | 2019 | 172 (134 ♂–38 ♀) | (A) Indoor (corridor) (B) Static background (C) Planar (overground) (D) Uncontrolled (indoor illumination) (E) Single-view (F) Straight; Curved trajectories (G) 15-month time span | (H) Walking (I) Self-selected (J) Natural clothing; Clothing variation (K) No load; Small item; Large item; Bag; Phone use (L) Sex | (M) SIL; 2D-KPT (OpenPose) (N) 20 fps (O) 224 × 224 (P) Fixed single camera (≈1.2 m height) (Q) Not applicable (R) Not reported |
| SAIVT-DGD [62] | 2011 | 35 | (A) Indoor (laboratory setting) (B) Static background (C) Planar (overground) (D) Controlled (E) Single frontal view (F) Straight trajectory (G) Single-session | (H) Walking (I) Self-selected; Fast (J) Natural clothing (K) Not reported (L) Not reported | (M) DPT; SIL; 3D-VOL (N) 30 fps (O) 640 × 480 (P) Fixed single Kinect v1 (Q) Timestamped multimodal streams (R) Not reported |
| SDUMLA-HMT [63] | 2010 | 106 (61 ♂–45 ♀) | (A) Indoor (laboratory setting) (B) Static background (C) Planar (overground) (D) Controlled (E) Multi-view (6 views: 0°; 45°; 67.5°; 90°; 112.5°; 135°) (F) Straight trajectory (G) Single-session | (H) Walking; Running (I) Self-selected (J) Natural clothing (K) No load; Bag (L) Age; Sex | (M) RGB (N) 25 fps (O) 320 × 240 (P) Fixed 6-camera arc (radius 6 m) (Q) Not reported (R) XviD-compressed videos |
| SMVDU-Gait (Single/Multi) [64] | 2019 | 20 (11 ♂–9 ♀) | (A) Outdoor (B) Static background (C) Planar (overground) (D) Uncontrolled (natural illumination) (E) Multi-view (3 views: lateral; oblique; frontal/back) (F) Straight trajectories (diagonal; front/back) (G) Single-session | (H) Walking (I) Self-selected (J) Natural clothing (K) No load (L) Age; Sex; Height; Profession | (M) RGB (N) 50 fps (O) 1920 × 1080 (P) Fixed single camera (12 m from background) (Q) Not applicable (R) Not reported |
| SOTON HumanID—Small DB [65] | 2002 | 12 | (A) Indoor (laboratory setting) (B) Static background (C) Planar (overground) (D) Controlled (E) Multi-view (4 views: side; oblique; elevated; frontal) (F) Straight trajectory (G) Single-session | (H) Walking (I) Self-selected (J) Clothing variation (K) Carrying items (L) Not reported | (M) RGB; SIL (N) 25 fps (O) 720 × 576 (P) Fixed multi-camera setup (Q) Not reported (R) Not reported |
| SOTON HumanID—Large DB [65] | 2002 | 115 (100 ♂–15 ♀) | (A) Indoor; Outdoor (B) Static background (green chroma-key); Dynamic background (C) Planar (overground); Treadmill (planar) (D) Controlled; Uncontrolled (natural outdoor illumination) (E) Multi-view (6 views: fronto- parallel and oblique per scenario) (F) Straight bidirectional walking path; Treadmill-constrained (G) Single-session | (H) Walking (I) Overground: self-selected; Treadmill: constant speed (J) Natural clothing; Footwear (K) Carrying items (L) Sex | (M) RGB; SIL; OF (N) 25 fps (O) 720 × 576 (P) Fixed multi-camera setup (Q) Camera sync information available (R) Not reported |
| SUSTech1K [27] | 2023 | 1050 | (A) Outdoor (multiple scenes) (B) Dynamic background (scene-induced occlusion) (C) Planar (overground) (D) Uncontrolled (natural illumination;day/night variation) (E) Multi-view (12 views) (F) Straight; Round-trip with turns (G) Multi-session (day/night variation) | (H) Walking (I) Self-selected (J) Natural clothing; Clothing variation (K) Bag; Object carrying; Umbrella (L) Not reported | (M) RGB; SIL; 3D-PCL (N) RGB: 30 fps; LiDAR: 10 fps (O) RGB: 1280 × 980; SIL: 64 × 64 (P) Mobile single robot-mounted RGB–LiDAR system (Q) Timestamped multimodal frames; GPS-clock synchronized (R) LiDAR noise |
| SZU RGB-D Gait Dataset [66] | 2013 | 99 | (A) Indoor (B) Static background (C) Planar (overground) (D) Controlled (E) Multi-view (2 views: 90° side; ≈60° oblique) (F) Straight trajectories (side; diagonal) (G) Single-session | (H) Walking (I) Self-selected (J) Natural clothing (K) No load (L) Not reported | (M) RGB; DPT; 3D-PCL; SIL; GEI (N) 30 fps (O) 640 × 480 (P) Fixed single ASUS Xtion PRO LIVE (≈80 cm height) (Q) RGB–depth calibrated/aligned (R) Not reported |
| TUM-IITKGP [67] | 2010 | 35 | (A) Indoor (corridor) (B) Static background (static and dynamic inter-object occlusion) (C) Planar (overground) (D) Controlled (indoor illumination) (E) Single side view (F) Straight bidirectional walking path (G) Single-session | (H) Walking (I) Self-selected (J) Natural clothing; Clothing variation (K) No load; Backpack (L) Not reported | (M) SIL (N) 30 fps (O) 640 × 480 (P) Fixed single camera (≈1.85 m height) (Q) Not applicable (R) Not reported |
| Dataset | Year | Subjects | Scene-Level | User-Level | Sensor-Level |
|---|---|---|---|---|---|
| MA-Gait [38] | 2022 | 95 (65 ♂–30 ♀) | (A) Indoor (B) Static background (C) Planar (overground) (D) Controlled (E) Multi-view (12 views: 0°–360°; 30° intervals) (F) Straight bidirectional walking path (G) Single-session | (H) Walking (12 patterns; 16 attribute labels) (I) Self-selected (J) Natural clothing (K) No load (L) Age; Sex | (M) SIL; 2D-KPT (HRNet) (N) 25 fps (O) RGB: 1920 × 1080; SIL: 64 × 44 (P) Fixed 6 cameras (2.2 m height) (Q) Not reported (R) Not reported |
| OU-LP-Age [40] | 2017 | 63,846 (31,093 ♂–32,753 ♀) | (A) Indoor (B) Static background (green chroma-key) (C) Planar (overground) (D) Controlled (E) Single side view (F) Straight trajectory (G) Long-run exhibition collection | (H) Walking (I) Self-selected (J) Natural clothing (K) No load (L) Age; Sex | (M) GEI (N) 30 fps (O) 128 × 88 (P) Fixed single camera (≈4 m from walking course) (Q) Not applicable (R) Not reported |
| RA-GAR [39] | 2025 | 533 | (A) Outdoor (B) Dynamic background (C) Planar (overground) (D) Uncontrolled (natural illumination; illumination changes) (E) Multi-view (10 views: 0°–360°; 36° intervals) (F) Straight trajectory (G) Single-session | (H) Walking (15 gait attributes) (I) Self-selected (J) Clothing variation (K) No load; Backpack (L) Age; Sex; Height; Body mass | (M) SIL; 2D-KPT (RTMPose); 3D-KPT (MotionBERT) (N) 30 fps (O) RGB: 1920 × 1080; SIL: 64 × 44 (P) Fixed 5 cameras (1.3 m height) (Q) Not reported (R) Not reported |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Nunes, J.F.; Moreira, P.M.; Tavares, J.M.R.S. Structuring Variability in Human Gait Datasets: A Covariate-Centered Taxonomy and Systematic Review of Image- and Depth-Based Collections. J. Imaging 2026, 12, 334. https://doi.org/10.3390/jimaging12070334
Nunes JF, Moreira PM, Tavares JMRS. Structuring Variability in Human Gait Datasets: A Covariate-Centered Taxonomy and Systematic Review of Image- and Depth-Based Collections. Journal of Imaging. 2026; 12(7):334. https://doi.org/10.3390/jimaging12070334
Chicago/Turabian StyleNunes, João Ferreira, Pedro Miguel Moreira, and João Manuel R. S. Tavares. 2026. "Structuring Variability in Human Gait Datasets: A Covariate-Centered Taxonomy and Systematic Review of Image- and Depth-Based Collections" Journal of Imaging 12, no. 7: 334. https://doi.org/10.3390/jimaging12070334
APA StyleNunes, J. F., Moreira, P. M., & Tavares, J. M. R. S. (2026). Structuring Variability in Human Gait Datasets: A Covariate-Centered Taxonomy and Systematic Review of Image- and Depth-Based Collections. Journal of Imaging, 12(7), 334. https://doi.org/10.3390/jimaging12070334

