Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (29)

Search Parameters:
Keywords = MPJPE

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
26 pages, 4343 KB  
Article
DiMMPose: A Diffusion-Mamba Hybrid Framework with Multi-Prompt for Efficient and Robust 3D Human Pose Estimation
by Xu Li, Xuefeng Guan, Chang Liu, Zengjie Wang, Xiaoyu Chen, Qingyang Xu, Shuyang Hou, Xiaopu Zhang and Huayi Wu
Sensors 2026, 26(17), 5559; https://doi.org/10.3390/s26175559 - 1 Sep 2026
Abstract
Monocular 3D Human Pose Estimation (3D HPE) typically adopts a two-stage approach: estimating 2D joint positions from images and then lifting them to 3D coordinates, effectively reducing dataset bias inherent in direct methods. However, current lifting techniques face two key challenges: many Transformer-based [...] Read more.
Monocular 3D Human Pose Estimation (3D HPE) typically adopts a two-stage approach: estimating 2D joint positions from images and then lifting them to 3D coordinates, effectively reducing dataset bias inherent in direct methods. However, current lifting techniques face two key challenges: many Transformer-based methods rely on attention-based or staged spatial–temporal modeling, which can limit efficient long-range frame-joint reasoning, while diffusion models support probabilistic modeling of pose uncertainty but remain sensitive to joint-coordinate noise. We propose DiMMPose, a diffusion-based framework enhanced by Mamba’s state-space model for robust and efficient 3D pose estimation. Its denoising process consists of two coordinated modules. The Spatiotemporal Mamba Block (STMB) serves as the core feature extraction module, employing internal Pose Mamba components with bidirectional state propagation and linear complexity to efficiently model long-range frame-joint dependencies. STMB further refines these features through Spatiotemporal Scan and Merge, which traverses the same skeleton tokens in complementary frame-joint orders and fuses the resulting representations. The Multi-Prompt Mamba Denoiser (MPMD) combines structured prompts encoded by LongCLIP with learnable prompt representations to provide anatomical and motion-related guidance during denoising. DiMMPose achieves an average MPJPE of 28.9 mm on Human3.6M under the DET setting, with action-specific errors of 21.2 mm for Walking and 22.0 mm for WalkTogether. It improves over FinePOSE by 3.0 mm, reduces inference latency by 57.1%, and achieves 23.0 mm MPJPE on MPI-INF-3DHP (N = 243). Full article
(This article belongs to the Section Sensing and Imaging)
34 pages, 4691 KB  
Article
Evaluation Protocols and Validation for Cameras in Indoor Healthcare Monitoring
by Amirhossein Dadashzadeh, Jingjing Liu, Qianhui Men, Qiushuo Cheng, Kirsty Scott, Lisa Alcock, Ian Craddock and Majid Mirmehdi
Sensors 2026, 26(17), 5460; https://doi.org/10.3390/s26175460 - 28 Aug 2026
Viewed by 153
Abstract
Camera-based monitoring systems are increasingly adopted in healthcare settings for the continuous assessment of patient movement and activities. However, their technical performance under real-world indoor conditions remains insufficiently characterised, preventing appropriate selection when choosing cameras for clinical or home adoption and reproducibility. Existing [...] Read more.
Camera-based monitoring systems are increasingly adopted in healthcare settings for the continuous assessment of patient movement and activities. However, their technical performance under real-world indoor conditions remains insufficiently characterised, preventing appropriate selection when choosing cameras for clinical or home adoption and reproducibility. Existing validation studies typically assess either device metrological performance or algorithm accuracy in isolation, and often do not systematically account for practical deployment factors, such as lighting variability, occlusions, and camera positioning. To address this, we present two technical validation protocols that evaluate the same cameras at both the metrological and pose-estimation levels under systematically controlled deployment conditions rarely addressed together in prior work: the first evaluates the metrological performance of RGB and RGBD cameras, and the second assesses their use in supporting human pose estimation, validated using state-of-the-art pose estimators. The proposed protocols systematically assess five cameras (four RGBD and one RGB) under controlled variations in lighting, camera height, viewing angle, and occlusion level, within representative indoor scenarios. The experimental results show that metrological performance varies substantially across cameras, with depth bias at 5 m ranging from ∼10 mm to over 1400 mm depending on the device. For 2D pose estimation, all cameras achieve broadly comparable accuracy (mean mAP between ~78% and ~90%) across cameras and estimators, whereas 3D reconstruction error differs markedly across devices (MPJPE ranging from 104 mm to 365 mm), closely reflecting underlying depth sensing quality. Environmental factors have a camera- and estimator-dependent effect on 3D performance, while camera mounting height has minimal influence within the evaluated range. This work provides evidence-based guidance for the selection and deployment of cameras in healthcare monitoring applications, addressing an important gap in current technical validation practice. Full article
(This article belongs to the Special Issue AI-Based Sensing and Imaging Applications)
Show Figures

Figure 1

36 pages, 1522 KB  
Review
Human Pose Estimation in 2D and 3D: A Survey of Analytical Methods, Benchmarking Frameworks, and Engineering Applications
by Rojan Shrestha, Aroudra Syamantak Thakur and Chenxi Wang
J. Exp. Theor. Anal. 2026, 4(3), 28; https://doi.org/10.3390/jeta4030028 - 5 Aug 2026
Viewed by 513
Abstract
This survey presents a comprehensive review of Human Pose Estimation spanning 2D and 3D settings, unifying prior work through a taxonomy of body representations (2D keypoints, 3D skeletons, dense meshes), processing flows (top-down vs. bottom-up), problem formulations (regression vs. detection/heatmaps), and modern learning [...] Read more.
This survey presents a comprehensive review of Human Pose Estimation spanning 2D and 3D settings, unifying prior work through a taxonomy of body representations (2D keypoints, 3D skeletons, dense meshes), processing flows (top-down vs. bottom-up), problem formulations (regression vs. detection/heatmaps), and modern learning architectures (CNNs, Transformers, GCNs). We compare reported benchmark results of representative methods across widely used datasets (e.g., COCO, MPII, Human3.6M, 3DPW) and evaluation metrics (AP/OKS, PCK/AUC, MPJPE/PA-MPJPE, PVE), highlighting trade-offs between accuracy, robustness, and efficiency. Despite substantial progress driven by deep learning and temporal modeling, we identify persistent challenges, including costly and biased annotations, domain shift, occlusion, depth ambiguity, multi-person association, and real-time constraints on edge devices. We synthesize emerging directions that target these gaps, data-centric learning, stronger temporal and kinematic priors, and whole-body modeling, and outline deployment-oriented frontiers including generative motion priors, model compression, and on-device inference, framing their implications for engineering systems that demand reliable, low-latency human motion analysis. Full article
Show Figures

Figure 1

55 pages, 5372 KB  
Article
Text-to-Korean Sign Language Pose Sequence Generation Using Non-Manual Signal Conditioning and Multi-Scale Temporal Refinement
by Seungju Lee and Gooman Park
Sensors 2026, 26(13), 4245; https://doi.org/10.3390/s26134245 - 4 Jul 2026
Viewed by 309
Abstract
Automatic sign language generation has the potential to support information accessibility for deaf and hard-of-hearing individuals. Generating sign language pose sequences from natural language text can serve as an intermediate representation for avatar-based sign language expression and sign language video synthesis. However, text-to-sign [...] Read more.
Automatic sign language generation has the potential to support information accessibility for deaf and hard-of-hearing individuals. Generating sign language pose sequences from natural language text can serve as an intermediate representation for avatar-based sign language expression and sign language video synthesis. However, text-to-sign pose generation is challenging because sign language conveys meaning through both manual movements and non-manual signals, while requiring temporally coherent motion over local and sentence-level contexts. In addition, text length does not directly correspond to the number of pose frames required for sign language expression. To address these issues, this study proposes a text-to-Korean Sign Language (KSL) pose generation model based on non-manual signal conditioning and multi-scale temporal refinement. The proposed framework integrates a text encoder, pose decoder, non-manual signal conditioning, multi-scale temporal refinement, and length prediction/blending. The model generates normalized 58-joint KSL keypoint sequences from morpheme-level text inputs and jointly optimizes pose reconstruction, motion continuity, bone consistency, PCK-aware precision, non-manual signal prediction, and length consistency. Experimental results on a KSL text–pose dataset show that the proposed model outperforms text-only and Transformer-based baselines. Compared with the Transformer text-to-pose baseline, the proposed model reduced MPJPE from 0.408236 to 0.316366 and Pose MAE from 0.165473 to 0.128570. It also improved PCK@0.05 from 0.136090 to 0.163928 and reduced the length relative error from 0.221455 to 0.127152. In particular, the best-threshold non-manual F1 substantially increased from 0.010859 to 0.494566. These results suggest that text-based KSL pose generation should jointly consider non-manual expressions, length consistency, and long-term temporal motion structure rather than relying only on frame-wise keypoint prediction. However, the reported improvements should be interpreted as coordinate- and label-level evidence, not as a complete validation of linguistic meaningfulness or real-world accessibility. Full article
(This article belongs to the Section Intelligent Sensors)
Show Figures

Figure 1

23 pages, 981 KB  
Review
From Optical to AI-Driven Markerless Motion Capture in Motor Learning and Rehabilitation
by Panagiotis Georganakis, Konstantinos Spinthiropoulos, Konstantinos Panitsidis, Dimitrios Parris and Vasiliki Gerodimou
Bioengineering 2026, 13(7), 776; https://doi.org/10.3390/bioengineering13070776 - 3 Jul 2026
Viewed by 1269
Abstract
Traditional biomechanical analysis is constrained by high capital costs and the physical limitations imposed by markers, posing significant barriers to clinical adoption. This review evaluates the emergence of artificial intelligence (AI)-based markerless motion capture (MMC) as a transformative approach for democratizing movement science [...] Read more.
Traditional biomechanical analysis is constrained by high capital costs and the physical limitations imposed by markers, posing significant barriers to clinical adoption. This review evaluates the emergence of artificial intelligence (AI)-based markerless motion capture (MMC) as a transformative approach for democratizing movement science in clinical rehabilitation. The discussion outlines the progression from legacy geometric visual hulls to advanced deep learning architectures, with particular focus on YOLO-based two-dimensional detection and spatio-temporal transformer models for three-dimensional pose estimation. Evidence indicates that multi-camera MMC frameworks achieve research-grade positional accuracy (16–34 mm Mean Per-Joint Position Error—MPJPE), while monocular systems provide sufficient sensitivity (82–88%) for longitudinal monitoring of geriatric fall risk and stroke recovery. While challenges persist in achieving precise axial rotation measurement, integrating real-time signal refinement enables objective and ecologically valid assessments in community-based healthcare settings. This technological advancement redefines movement analysis, shifting it from a laboratory-bound procedure to a widely accessible and interoperable diagnostic tool. Full article
Show Figures

Graphical abstract

23 pages, 3576 KB  
Article
3D Pose Estimation Using Virtual Projection Based on 3D Reconstructed Model
by Jung-Woo Kim, Sol Lee, Byung-Seo Park, Hak-Bum Lee, Dong-Ho Kang and Young-Ho Seo
Sensors 2026, 26(11), 3302; https://doi.org/10.3390/s26113302 - 22 May 2026
Viewed by 501
Abstract
In this paper, we estimate and refine 3D human pose using the 3D point cloud or mesh model reconstructed from RGB-D cameras or volumetric capture systems. We first reconstruct the 3D model using the multi-view cameras to estimate a highly accurate skeleton. To [...] Read more.
In this paper, we estimate and refine 3D human pose using the 3D point cloud or mesh model reconstructed from RGB-D cameras or volumetric capture systems. We first reconstruct the 3D model using the multi-view cameras to estimate a highly accurate skeleton. To obtain a 2D skeleton with low error, the reconstructed 3D model is projected to four virtual planes after decidi ng the direction of the 3D model. Four 2D skeletons are estimated from four images projected in the virtual plane. Afterward, the refinement process selects candidate joints based on the distribution of local vertices and the DBSCAN algorithm. It applies a sphere fitting to ensure that the final joints are located within the body volume. The joints are combined at the intersection through the back-projection of the joints, including those in the 2D skeleton on the virtual plane. The joints in the intersection are refined using the spatial distribution of the 3D information. Through the proposed method, we estimated a stable and geometrically consistent 3D human pose from reconstructed volumetric data. Using models with ground truth, we calculated the MPJPE between the skeletons of the proposed and the ground truth. The 3D pose estimation was evaluated through a visual assessment of the captured image, and the results were quantitatively compared with the 3D joint positions acquired by the motion capture device. Full article
(This article belongs to the Section Sensing and Imaging)
Show Figures

Figure 1

31 pages, 13501 KB  
Article
Adaptive 3D Human Pose Estimation Based on Spatial–Temporal Complexity Awareness
by Wensi Zhang, Ziyan Yang, Chengfeng Hu, Jing Sun and Jie Li
Electronics 2026, 15(10), 2076; https://doi.org/10.3390/electronics15102076 - 13 May 2026
Viewed by 685
Abstract
Existing 3D human pose estimation methods use fixed computation strategies processing diverse action sequences, leading to computational redundancy for simple actions, insufficient high-frequency information capture for complex actions, and low long-sequence processing efficiency. To address these issues, this paper proposes a Spatial–Temporal Complexity-Aware [...] Read more.
Existing 3D human pose estimation methods use fixed computation strategies processing diverse action sequences, leading to computational redundancy for simple actions, insufficient high-frequency information capture for complex actions, and low long-sequence processing efficiency. To address these issues, this paper proposes a Spatial–Temporal Complexity-Aware Adaptive Computation Framework (CAAPoseFormer). First, a spatial–temporal coupled complexity quantification module is built to integrate spatial dispersion and temporal motion variance for graded action complexity quantification. On this basis, a time–frequency dual-domain adaptive pruning strategy is proposed to dynamically allocate temporal window length and frequency-domain DCT coefficients on demand. Furthermore, a mask-guided sparse interaction encoding mechanism is designed to enable efficient parallel computation of variable-length features by shielding invalid padding regions. Experiments on the Human3.6M dataset show that, versus the baseline PoseFormerV2, the proposed method cuts parameters by 85.3% and computational cost by 64.8% while retaining comparable accuracy (MPJPE 44.2 mm), boosting unit computational efficiency 2.8×. Moreover, compared with state-of-the-art (SOTA) methods like MHFormer and MotionBERT, our method reduces computational costs (MACs) by 97.4% and nearly three orders of magnitude, respectively. This framework effectively breaks the inference bottleneck of high-precision models on low-power hardware, suiting latency-sensitive real-time applications well. Full article
(This article belongs to the Special Issue Advances in Real-Time Object Detection and Tracking)
Show Figures

Figure 1

14 pages, 18688 KB  
Article
Outdoor Motion Capture at Scale
by Michael Zwölfer, Martin Mössner, Helge Rhodin and Werner Nachbauer
Sensors 2026, 26(6), 1951; https://doi.org/10.3390/s26061951 - 20 Mar 2026
Viewed by 829
Abstract
Capturing kinematic data in outdoor sports is challenging, as motions span large capture volumes and occur under difficult environmental conditions. Video-based approaches, particularly with pan–tilt–zoom cameras, offer a practical solution, but the extensive manual post-processing required limits their use to short sequences and [...] Read more.
Capturing kinematic data in outdoor sports is challenging, as motions span large capture volumes and occur under difficult environmental conditions. Video-based approaches, particularly with pan–tilt–zoom cameras, offer a practical solution, but the extensive manual post-processing required limits their use to short sequences and few athletes. This study presents a motion capture pipeline that automates the detection of both reference points and sport-specific keypoints to overcome this limitation. The field test employed eight cameras covering a 250×80×30 m capture volume with nearly 300 reference points. Ten state-certified ski instructors performed eight standardized maneuvers. Reference points were localized through a hybrid approach combining YOLO object detection and ArUco marker identification. AlphaPose was fine-tuned on a new manually annotated dataset to detect skier-specific keypoints (e.g., skis, poles) alongside anatomical landmarks. Continuous frame-wise calibration and 3D reconstruction were performed using Direct Linear Transformation. Evaluation compared automated detections with manual annotations. Automated reference point detection achieved a mean localization error of 4.1 pixels (0.1% of 4K width) and reduced 3D segment-length variation by 23%. The skier-specific keypoint model reached 98% PCK, mAP of 0.97, and an MPJPE of 10.3 pixels while lowering 3D segment-length variation by 0.5 cm compared to manual digitization and 0.6 cm relative to a pretrained model. Replacing manual digitization with automated detection improves accuracy and facilitates kinematic data collection in large outdoor fields with many athletes and trials. The approach also enables the creation of sport-specific datasets valuable for biomechanical research and training next-generation 3D pose estimation models. Full article
(This article belongs to the Special Issue Advanced Sensors in Biomechanics and Rehabilitation—2nd Edition)
Show Figures

Graphical abstract

20 pages, 7825 KB  
Article
STAG-Net: A Lightweight Spatial–Temporal Attention GCN for Real-Time 6D Human Pose Estimation in Human–Robot Collaboration Scenarios
by Chunxin Yang, Ruoyu Jia, Qitong Guo, Xiaohang Shi, Masahiro Hirano and Yuji Yamakawa
Robotics 2026, 15(3), 54; https://doi.org/10.3390/robotics15030054 - 4 Mar 2026
Viewed by 1710
Abstract
Most existing research in human pose estimation focuses on predicting joint positions, paying limited attention to recovering the full 6D human pose, which comprises both 3D joint positions and bone orientations. Position-only methods treat joints as independent points, often resulting in structurally implausible [...] Read more.
Most existing research in human pose estimation focuses on predicting joint positions, paying limited attention to recovering the full 6D human pose, which comprises both 3D joint positions and bone orientations. Position-only methods treat joints as independent points, often resulting in structurally implausible poses and increased sensitivity to depth ambiguities—cases where poses share nearly identical joint positions but differ significantly in limb orientations. Incorporating bone orientation information helps enforce geometric consistency, yielding more anatomically plausible skeletal structures. Additionally, many state-of-the-art methods rely on large, computationally expensive models, which limit their applicability in real-time scenarios, such as human–robot collaboration. In this work, we propose STAG-Net, a novel 2D-to-6D lifting network that integrates Graph Convolutional Networks (GCNs), attention mechanisms, and Temporal Convolutional Networks (TCNs). By simultaneously learning joint positions and bone orientations, STAG-Net promotes geometrically consistent skeletal structures while remaining lightweight and computationally efficient. On the Human3.6M benchmark, STAG-Net achieves an MPJPE of 41.8 mm using 243 input frames. In addition, we introduce a lightweight single-frame variant, STG-Net, which achieves 50.8 mm MPJPE while operating in real time at 60 FPS using a single RGB camera. Extensive experiments on multiple large-scale datasets demonstrate the effectiveness and efficiency of the proposed approach. Full article
(This article belongs to the Special Issue Human–Robot Collaboration in Industry 5.0)
Show Figures

Figure 1

29 pages, 3921 KB  
Article
A Semantic Priors-Based Non-Euclidean Topological Enhancement Method for 3D Human Pose Estimation in Multi-Class Complex Human Actions
by Xiaowei Han, Chaolong Fei, Yibo Feng, Wenbao Si and Guilin Yao
Electronics 2026, 15(1), 155; https://doi.org/10.3390/electronics15010155 - 29 Dec 2025
Viewed by 691
Abstract
Three-dimensional human pose estimation (3D HPE) aims to recover the three-dimensional coordinates of human joints from 2D images or videos to achieve precise quantification of human movement. In 3D HPE tasks based on multi-class complex human action datasets, the performance of existing Graph [...] Read more.
Three-dimensional human pose estimation (3D HPE) aims to recover the three-dimensional coordinates of human joints from 2D images or videos to achieve precise quantification of human movement. In 3D HPE tasks based on multi-class complex human action datasets, the performance of existing Graph Convolutional Network (GCN) and Transformer fusion models is constrained by the fixed physical connections of the skeleton, which impedes the modeling of cross-joint long-range semantic dependencies and hinders further performance gains. To address this issue, this study proposes a semantic prior-based non-Euclidean topology enhancement method for multi-class complex human actions, built upon a GCN–Transformer fusion model. The proposed method retains the original physical connections while introducing semantic prior edges; by constructing a hybrid topology structure, it explicitly models long-range semantic dependencies between non-adjacent joints, thereby facilitating the extraction of cross-joint semantic information. Experimental results on the Human3.6M and HumanEva-I datasets surpass those of SOTA baseline models. On the Human3.6M dataset, MPJPE and P-MPJPE are reduced by 1.25% and 0.63%, respectively. For the Walk and Jog actions on the HumanEva-I dataset, MPJPE is reduced by approximately 6.5%. These results demonstrate that the proposed method offers significant advantages for 3D HPE tasks based on multi-class complex human action data. Full article
(This article belongs to the Section Artificial Intelligence)
Show Figures

Figure 1

15 pages, 2618 KB  
Article
Multi-Agent Collaboration for 3D Human Pose Estimation and Its Potential in Passenger-Gathering Behavior Early Warning
by Xirong Chen, Hongxia Lv, Lei Yin and Jie Fang
Electronics 2026, 15(1), 78; https://doi.org/10.3390/electronics15010078 - 24 Dec 2025
Cited by 2 | Viewed by 982
Abstract
Passenger-gathering behavior often triggers safety incidents such as stampedes due to overcrowding, posing significant challenges to public order maintenance and passenger safety. Traditional early warning algorithms for passenger-gathering behavior typically perform only global modeling of image appearance, neglecting the analysis of individual passenger [...] Read more.
Passenger-gathering behavior often triggers safety incidents such as stampedes due to overcrowding, posing significant challenges to public order maintenance and passenger safety. Traditional early warning algorithms for passenger-gathering behavior typically perform only global modeling of image appearance, neglecting the analysis of individual passenger actions in practical 3D physical space, leading to high false-alarm and missed-alarm rates. To address this issue, we decompose the modeling process into two stages: human pose estimation and gathering behavior recognition. Specifically, the pose of each individual in 3D space is first estimated from images, and then fused with global features to complete the early warning. This work focuses on the former stage and aims to develop an accurate and efficient human pose estimation model capable of real-time inference on resource-constrained devices. To this end, we propose a 3D human pose estimation framework that integrates a hybrid spatio-temporal Transformer with three collaborative agents. First, a reinforcement learning-based architecture search agent is designed to adaptively select among Global Self-Attention, Window Attention, and External Attention for each block to optimize the model structure. Second, a feedback optimization agent is developed to dynamically adjust the search process, balancing exploration and convergence. Third, a quantization agent is employed that leverages quantization-aware training (QAT) to generate an INT8 deployment-ready model with minimal loss in accuracy. Experiments conducted on the Human3.6M dataset demonstrate that the proposed method achieves a mean per joint position error (MPJPE) of 42.15 mm with only 4.38 M parameters and 19.39 GFLOPs under FP32 precision, indicating substantial potential for subsequent gathering behavior recognition tasks. Full article
(This article belongs to the Section Computer Science & Engineering)
Show Figures

Figure 1

15 pages, 2201 KB  
Article
CGFusionFormer: Exploring Compact Spatial Representation for Robust 3D Human Pose Estimation with Low Computation Complexity
by Tao Lu, Hongtao Wang and Degui Xiao
Sensors 2025, 25(19), 6052; https://doi.org/10.3390/s25196052 - 1 Oct 2025
Viewed by 1302
Abstract
Transformer-based 2D-to-3D lifting methods have demonstrated outstanding performance in 3D human pose estimation from 2D pose sequences. However, they still encounter challenges with the relatively poor quality of 2D joints and substantial computational costs. In this paper, we propose a CGFusionFormer to address [...] Read more.
Transformer-based 2D-to-3D lifting methods have demonstrated outstanding performance in 3D human pose estimation from 2D pose sequences. However, they still encounter challenges with the relatively poor quality of 2D joints and substantial computational costs. In this paper, we propose a CGFusionFormer to address these problems. We propose a compact spatial representation (CSR) to robustly generate local spatial multihypothesis features from part of the 2D pose sequence. Specifically, CSR models spatial constraints based on body parts and incorporates 2D Gaussian filters and nonparametric reduction to improve spatial features against low-quality 2D poses and reduce the computational cost of subsequent temporal encoding. We design a residual-based Hybrid Adaptive Fusion module that combines multihypothesis features with global frequency domain features to accurately estimate the 3D human pose with minimal computational cost. We realize CGFusionFormer with a PoseFormer-like transformer backbone. Extensive experiments on the challenging Human3.6M and MPI-INF-3DHP benchmarks show that our method outperforms prior transformer-based variants in short receptive fields and achieves a superior accuracy–efficiency trade-off. On Human3.6M (sequence length 27, 3 input frames), it achieves 47.6 mm Mean Per Joint Position Error (MPJPE) at only 71.3 MFLOPs, representing about a 40 percent reduction in computation compared with PoseFormerV2 while attaining better accuracy. On MPI-INF-3DHP (81-frame sequences), it reaches 97.9 Percentage of Correct Keypoints (PCK), 78.5 Area Under the Curve (AUC), and 27.2 mm MPJPE, matching the best PCK and achieving the lowest MPJPE among the compared methods under the same setting. Full article
Show Figures

Figure 1

22 pages, 8860 KB  
Article
Generating Multi-View Action Data from a Monocular Camera Video by Fusing Human Mesh Recovery and 3D Scene Reconstruction
by Hyunsu Kim and Yunsik Son
Appl. Sci. 2025, 15(19), 10372; https://doi.org/10.3390/app151910372 - 24 Sep 2025
Cited by 2 | Viewed by 3284
Abstract
Multi-view data, captured from various perspectives, is crucial for training view-invariant human action recognition models, yet its acquisition is hindered by spatio-temporal constraints and high costs. This study aims to develop the Pose Scene EveryWhere (PSEW) framework, which automatically generates temporally consistent, multi-view [...] Read more.
Multi-view data, captured from various perspectives, is crucial for training view-invariant human action recognition models, yet its acquisition is hindered by spatio-temporal constraints and high costs. This study aims to develop the Pose Scene EveryWhere (PSEW) framework, which automatically generates temporally consistent, multi-view 3D human action data from a single monocular video. The proposed framework first predicts 3D human parameters from each video frame using a deep learning-based Human Mesh Recovery (HMR) model. Subsequently, it applies tracking, linear interpolation, and Kalman filtering to refine temporal consistency and produce naturalistic motion. The refined human meshes are then reconstructed into a virtual 3D scene by estimating a stable floor plane for alignment, and finally, novel-view videos are rendered using user-defined virtual cameras. As a result, the framework successfully generated multi-view data with realistic, jitter-free motion from a single video input. To assess fidelity to the original motion, we used Root Mean Square Error (RMSE) and Mean Per Joint Position Error (MPJPE) as metrics, achieving low average errors in both 2D (RMSE: 0.172; MPJPE: 0.202) and 3D (RMSE: 0.145; MPJPE: 0.206) space. PSEW provides an efficient, scalable, and low-cost solution that overcomes the limitations of traditional data collection methods, offering a remedy for the scarcity of training data for action recognition models. Full article
(This article belongs to the Special Issue Advanced Technologies Applied for Object Detection and Tracking)
Show Figures

Figure 1

25 pages, 9990 KB  
Article
Bidirectional Mamba-Enhanced 3D Human Pose Estimation for Accurate Clinical Gait Analysis
by Chengjun Wang, Wenhang Su, Jiabao Li and Jiahang Xu
Fractal Fract. 2025, 9(9), 603; https://doi.org/10.3390/fractalfract9090603 - 17 Sep 2025
Cited by 3 | Viewed by 4353
Abstract
Three-dimensional human pose estimation from monocular video remains challenging for clinical gait analysis due to high computational cost and the need for temporal consistency. We present Pose3DM, a bidirectional Mamba-based state-space framework that models intra-frame joint relations and inter-frame dynamics with linear computational [...] Read more.
Three-dimensional human pose estimation from monocular video remains challenging for clinical gait analysis due to high computational cost and the need for temporal consistency. We present Pose3DM, a bidirectional Mamba-based state-space framework that models intra-frame joint relations and inter-frame dynamics with linear computational complexity. Replacing transformer self-attention with state-space modeling improves efficiency without sacrificing accuracy. We further incorporate fractional-order total-variation regularization to capture long-range dependencies and memory effects, enhancing temporal and spatial coherence in gait dynamics. On Human3.6M, Pose3DM-L achieves 37.9 mm MPJPE under Protocol 1 (P1) and 32.1 mm P-MPJPE under Protocol 2 (P2), with 127 M MACs per frame and 30.8 G MACs in total. Relative to MotionBERT, P1 and P2 errors decrease by 3.3% and 2.4%, respectively, with 82.5% fewer parameters and 82.3% fewer MACs per frame. Compared with MotionAGFormer-L, Pose3DM-L improves P1 by 0.5 mm and P2 by 0.4 mm while using 60.6% less computation: 30.8 G vs. 78.3 G total MACs and 127 M vs. 322 M per frame. On AUST-VisGait across six gait patterns, Pose3DM consistently yields lower MPJPE, standard error, and maximum error, enabling reliable extraction of key gait parameters from monocular video. These results highlight state-space models as a cost-effective route to real-time gait assessment using a single RGB camera. Full article
Show Figures

Figure 1

17 pages, 939 KB  
Article
Whole-Body 3D Pose Estimation Based on Body Mass Distribution and Center of Gravity Constraints
by Fan Wei, Guanghua Xu, Qingqiang Wu, Penglin Qin, Leijun Pan and Yihua Zhao
Sensors 2025, 25(13), 3944; https://doi.org/10.3390/s25133944 - 25 Jun 2025
Cited by 2 | Viewed by 2397
Abstract
Estimating the 3D pose of a human body from monocular images is crucial for computer vision applications, but the technique remains challenging due to depth ambiguity and self-occlusion. Traditional methods often suffer from insufficient prior knowledge and weak constraints, resulting in inaccurate 3D [...] Read more.
Estimating the 3D pose of a human body from monocular images is crucial for computer vision applications, but the technique remains challenging due to depth ambiguity and self-occlusion. Traditional methods often suffer from insufficient prior knowledge and weak constraints, resulting in inaccurate 3D keypoint estimation. In this paper, we propose a method for whole-body 3D pose estimation based on a Transformer architecture, integrating body mass distribution and center of gravity constraints. The method maps the pose to the center of gravity position using the anatomical mass ratio of the human body and computes the segment-level center of gravity using the moment synthesis method. A combined loss function is designed to enforce consistency between the predicted keypoints and the center of gravity position, as well as the invariance of limb length. Extensive experiments on the Human 3.6M WholeBody dataset demonstrate that the proposed method achieves state-of-the-art performance, with a whole-body mean joint position error (MPJPE) of 44.49 mm, which is 60.4% lower than the previous Large Simple Baseline method. Notably, it reduces the body part keypoints’ MPJPE from 112.6 to 40.41, showcasing the enhanced robustness and effectiveness to occluded scenes. This study highlights the effectiveness of integrating physical constraints into deep learning frameworks for accurate 3D pose estimation. Full article
Show Figures

Figure 1

Back to TopTop