Next Article in Journal
A2S2C-Det: Dual-Path Adaptive Aggregation with Spatial-Semantic Compensation for Strip Steel Surface Defect Detection
Previous Article in Journal
Echo Model Analysis and Frequency-Domain Imaging Algorithm for Geosynchronous Spaceborne–Airborne FMCW Bistatic SAR with High-Maneuvering Receiver
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Hitting the Gym with Fit3D: Benchmarking and Improving Monocular 3D Human Reconstruction on Extreme Fitness Motions

Institute of Mathematics of the Romanian Academy, 010702 Bucharest, Romania
J. Imaging 2026, 12(7), 311; https://doi.org/10.3390/jimaging12070311
Submission received: 12 June 2026 / Revised: 2 July 2026 / Accepted: 6 July 2026 / Published: 9 July 2026
(This article belongs to the Section Computer Vision and Pattern Recognition)

Abstract

Fitness motions present some of the most challenging cases for monocular 3D human reconstruction: extreme articulations, heavy self-occlusion, and frequent self-contact. The Fit3D dataset captures these motions at large scale via a 12-camera VICON motion capture system synchronized with 4 RGB cameras, but in its original release provides only 3D skeletons. This paper contributes three studies built on top of Fit3D. First, we design and validate a methodology for constructing dense GHUM and SMPL-X pseudo-ground-truth shape and pose annotations on top of the raw MoCap: an optimization-based fitting pipeline that combines markers, multi-view 2D keypoints, separate body and hand normalizing-flow priors, and a self-collision loss, which we show improves on the marker-only MoSh++ baseline on the hands and extremities. Second, we define a standardized evaluation protocol—metric set, frame sampling, and coordinate conventions—for monocular 3D human reconstruction on Fit3D, served through the IMAR-hosted Fit3D evaluation resource, and use it to conduct a comparative benchmark study of 19 representative methods spanning optimization-based, single-frame, and video-based families; the two trained on Fit3D (NLF and SMPLest-X) lead the position and orientation metrics, respectively. Third, a controlled fine-tuning experiment shows that adding Fit3D to the training mixture of a strong baseline (HMR2.0) sharply lowers error on the hardest fitness poses without degrading out-of-domain generalization. The Fit3D dataset and the GHUM/SMPL-X annotations are available, under a non-commercial research license, through the IMAR Fit3D resource; as of June 2026, 1088 academics have registered for access.

1. Introduction

Monocular 3D human pose and shape reconstruction has matured rapidly: methods such as 4DHumans [1], NLF [2], CameraHMR [3], SMPLest-X [4] and SAM 3D Body [5] now produce convincing meshes on standard benchmarks like 3DPW [6] and EMDB [7]. Fitness motions, however, remain a regime in which these methods struggle. Deep squats, lunges, bridges, handstands, kicks, and weight-loaded poses generate extreme articulations, severe self-occlusion, and frequent self-contact—failure modes that the dominant training corpora (Human3.6M [8], 3DPW [6], BEDLAM [9], AGORA [10]) do not adequately cover. Progress in this regime is further bottlenecked by two missing ingredients: there is no parametric (GHUM, SMPL-X) pseudo-ground-truth body annotation on a representative fitness corpus, and no standardized evaluation protocol that researchers can compare against.
The Fit3D dataset, introduced by Fieraru et al. [11], addresses the data side of the problem: it captures 13 subjects (one licensed fitness instructor and 12 trainees) performing 47 simple and compound exercises that cover all the major muscle groups, using a 12-camera VICON motion capture system synchronized with 4 calibrated RGB cameras, and contains approximately 3.0 million RGB frames together with the corresponding 3D MoCap skeletons and per-subject 3D body scans. However, in its original release, Fit3D provides only the 3D skeletons, which limits its direct use as supervision for modern parametric-body-model regressors.
In this paper we make three contributions on top of the existing Fit3D dataset. First, we design and validate the methodology used to construct dense GHUM [12] and SMPL-X [13] pseudo-ground-truth shape and pose annotations on top of the raw MoCap, by an optimization-based fitting pipeline that combines VICON markers, multi-view 2D keypoints, separate body and hand normalizing-flow priors, and a self-collision loss in order to accommodate the extreme articulations and self-occlusions characteristic of fitness routines. These annotations are available through the IMAR Fit3D resource [11] and have since been adopted in the training corpora of large-scale 3D human reconstruction methods such as NLF [2] and SMPLest-X [4]. Second, we define a standardized evaluation protocol—metric set, frame sampling, and coordinate conventions—and conduct a comparative benchmark study of monocular 3D human reconstruction on Fit3D; the protocol is served through the IMAR-hosted Fit3D evaluation server [11] at https://fit3d.imar.ro/ (accessed on 5 July 2026), which accepts GHUM, SMPL-X, and 3D-joint submissions and reports a standardized set of metrics. Using it, we benchmark 19 representative methods spanning optimization-based, single-frame, and video-based families (2019–2026) and analyze the impact of including the dataset in their training corpus. Third, we run a controlled fine-tuning experiment that isolates the effect of Fit3D supervision, showing that it improves a strong baseline most on exactly the hard, out-of-distribution poses where current methods fail, without regressing out-of-domain performance. Figure 1 gives an overview of the three contributions.

2. Related Work

Visual human sensing has been extensively studied [13,14,15,16,17,18,19,20,21,22,23,24,25,26]. Applications exist in many domains such as the automotive industry [27,28], fashion industry [29,30], activity recognition [31] and many others. One particular area that has been less considered by research is high-quality 3D human reconstruction for fitness, the focus of the present work.
  • 3D Human Pose and Shape Datasets.
Large-scale 3D human datasets have been the main driver of monocular 3D human reconstruction. Human3.6M [8] captures daily activities of 11 subjects with a MoCap rig, MPI-INF-3DHP [18] extends the setting to outdoor sequences, and 3DPW [6] pioneers in-the-wild capture via IMUs. More recent datasets target specific failure modes: EMDB [7] provides accurate world-grounded motion via electromagnetic sensors, RICH [32] adds dense scene contact, MOYO [33] captures yoga and a wide range of extreme articulations, and BEDLAM [9] and AGORA [10] provide synthetic ground truth at scale. None of these datasets are specifically targeted at fitness training, where extreme articulation, weight loading, self-contact, and tight gym-style camera setups co-occur. Fit3D fills this gap by combining high-fidelity MoCap with parametric body-model annotations on a corpus of fitness exercises; Table 1 summarizes how it relates to these datasets along the axes of scale, capture modality, and the body-model ground truth available.
  • Monocular 3D Human Reconstruction.
Methods for monocular 3D human reconstruction range from optimization-based fitting (SMPLify-X [13]) to single-frame regressors that estimate SMPL, SMPL-X, or GHUM parameters from an image (PARE [34], HybrIK [35], PyMAF-X [36], CLIFF [37], 4DHumans/HMR2 [1], TokenHMR [38], Multi-HMR [39], CameraHMR [3], SMPLest-X [4], BLADE [40], NLF [2], SAM 3D Body [5]). Video-based regressors (WHAM [41], GVHMR [42], PromptHMR [43]) and refinement-style approaches (ScoreHMR [44]) additionally exploit temporal context or post-hoc consistency. We benchmark a representative cross-section of these families on Fit3D in Section 4.
  • Evaluation metrics and benchmarks.
Monocular 3D human reconstruction is most commonly scored with the mean per-joint position error (MPJPE) and its Procrustes-aligned form on Human3.6M [8] and 3DPW [6]. For parametric body models, these are complemented by the mean per-vertex error (MPVPE) on expressive SMPL-X benchmarks such as EHF [13] and RICH [32], and by the mean per-joint angle error (MPJAE) introduced with 3DPW [6], which measures orientation accuracy independently of limb length. To isolate body shape from pose, we additionally adopt the shape-only T-pose vertex error (MPVPE-T), following the PVE-T metric of Sengupta et al. [45]. Our protocol (Section 4) brings this metric set together on Fit3D, evaluating GHUM, SMPL-X, and 3D-joint predictions against the released annotations.

3. Fit3D

This section briefly summarizes the Fit3D dataset [11] that the present work builds on, and then describes the methodology we use to extend its annotations with dense GHUM and SMPL-X pseudo-ground-truth.

3.1. Dataset

The Fit3D dataset, introduced by Fieraru et al. [11], features 13 human subjects performing fitness exercises. Among them, there is one licensed fitness instructor (considered the reference for correctness in exercise execution), while the rest are considered trainees of various levels of skill.
The dataset was captured with a VICON motion capture system consisting of 12 motion cameras, synchronized with 4 RGB cameras. The capture process involved reflective markers affixed to either the subject’s skin or clothing. All subjects were dressed in gym-like attire that fits tightly on the body. During the recordings, the subjects used several typical gym objects: 2 dumbbells, a barbell, and a rubber band. A low-height, footless table was used to ease the difficulty of the exercises involving lifting the barbell.
The fitness exercises target the major body muscle groups: arms, legs, back, and abdomen. They are split into two groups: simple (involving basic repetitions such as push-up, squat, or dumbbell biceps raise) and compound (assuming more complex routines involving multiple body regions, such as burpees, entailing a push-up and a jump, or clean and press, assuming certain trajectories of the arms). Each subject was asked to perform each type of exercise for a minimum of 5 repetitions.
Subject heights range from 1.55 to 1.90 m and weights from 60 to 110 kg; the dataset covers a mix of fit subjects (persons who exert a high degree of physical activity) as well as less trained ones. Fit3D consists of approximately 3.0 million unique MoCap 3D skeletons synchronized with RGB images. Each subject was also 3D scanned. The dataset is split by subject into training (8 subjects), validation (2 subjects), and testing (3 subjects); the 10 training and validation subjects together comprise approximately 2.3 million images and the 3 test subjects approximately 0.7 million, with all exercise types present in every split. Unless noted otherwise, the marker-fitting comparison (Section 3.3) and the fine-tuning experiment (Section 6) use the 8 training subjects, while the benchmark (Section 4) evaluates on the 3 test subjects. Each video is also manually segmented into repetitions, with a total of 2964 annotated timestamps.
In its original release [11], Fit3D provides only the 3D skeletons synchronized with RGB. In Section 3.2, we describe the methodology used to augment each frame with dense GHUM [12] (and, by conversion, SMPL-X [13]) shape and pose pseudo-ground-truth (pGT), obtained by fitting the body model to the markers and multi-view RGB evidence. We use the term pseudo-ground-truth to emphasize that these annotations are produced by optimization-based fitting rather than by direct measurement, following the terminology adopted for comparable optimization-fitted datasets such as RICH [32]. Throughout, we reserve the unqualified term “ground truth” for the directly measured VICON marker positions, and refer to the optimized GHUM/SMPL-X parameters as pseudo-ground-truth or, where unambiguous, the reference annotations.

3.2. Motion and Shape Capture

By default, the VICON motion capture system is designed to track only the 3D position of markers attached to the body of a person and to regress a 3D skeleton underneath. Yet, we are interested in tracking the entire surface of the body, as well as the articulation of the two hands. To achieve our goals, we leverage any additional information that is available to us: scans of each subject, the four synchronized calibrated cameras, keypoint detectors, personalized marker placement on the body surface, pose priors, and self-intersection losses.
As fitness motions produce extreme articulations and significant self-occlusions, it can happen that the VICON system outputs erroneous marker positions or marker types. Therefore, we first manually fix these errors in the VICON Nexus software by switching marker identities and/or interpolating between correct tracks.
To capture the motion and shape of each subject, we use GHUM [12], a parametric 3D human body model. The model is learned from a large dataset of human shape and poses and is able to generate a mesh consisting of 10,168 vertices and 20,332 facets. Its inputs are body pose parameters θ represented as 6D rotations between the major limbs of the body (including hand articulations), body shape parameters β represented as the latent of a variational autoencoder, a global rotation R , and a global translation T . Estimating the motion and shape of a subject reduces to estimating the body shape parameter β from the 3D body scan, once per subject, and then optimizing θ , R , and  T at each frame.
We first estimate the body shape β of each subject by fitting the GHUM body model to the point cloud generated by the 3D body scanner, following the methodology described in [12].
With the body shape β fixed, we focus on fitting the pose θ , global rotation R , and global translation T at each frame. The first frame of a sequence is initialized by running THUNDR [46] in each of the four RGB views and averaging the resulting GHUM parameters; subsequent frames are warm-started from the optimum of the previous one, exploiting temporal continuity of the motion.
We minimize a weighted sum of five losses. The marker alignment term
L 3 D = 1 K i = 1 K m i 3 D M i 3 D 2 2
measures the squared error between the K = 53 GHUM-regressed markers m i 3 D and the corresponding (post-cleanup) VICON markers M i 3 D . Since the marker layout does not densely cover the hands, we additionally use a 2D keypoint alignment loss over the four synchronized RGB views,
L 2 D = 1 4 J c = 1 4 n = 1 J π c ( j n 3 D ) k ^ c , n 2 D 2 2 ,
where j n 3 D is the n-th of the J joints regressed from the GHUM mesh, π c projects it into camera c, and  k ^ c , n 2 D are detections from a 2D keypoint predictor [47]. L 2 D acts primarily as a constraint on the hand joints, which are not covered by VICON markers.
Fitness motions induce large articulations and frequent self-occlusions; to keep the reconstruction physically plausible, we add a self-collision penalty. We approximate the GHUM mesh by a coarser watertight proxy mesh M ^ with vertices V ^ , computed with the PHUM body model [48], and, following [48], apply the generalized winding number test of [49] to label each proxy vertex as inside ( L + ) or outside the remaining body surface. The penalty
L sc = l L + V ^ l NN ( V ^ l , M ^ ) 2 2
pulls each inside vertex V ^ l toward its nearest neighbor NN ( V ^ l , M ^ ) on the outside of the proxy mesh. The body and hand poses are regularized by two separate normalizing-flow priors [50], NF b for the body joints and NF h for the hand articulation:
L θ b = NF b ( θ b ) 2 2 , L θ h = NF h ( θ h ) 2 2 ,
which penalize implausible articulations. Finally, a temporal smoothness term
L s = θ t θ t 1 2 2 + R t R t 1 2 2 + T t T t 1 2 2
couples consecutive frames and damps jitter, where the pose θ and the global rotation R are differenced in their 6D rotation representation and T in metric coordinates.
We minimize the weighted sum L = λ 3 D L 3 D + λ 2 D L 2 D + λ sc L sc + λ θ b L θ b + λ θ h L θ h + λ s L s per frame with BFGS; the weights λ were hand-tuned by visual inspection of the reconstructions. Some Fit3D motions lie outside the distribution the body prior was trained on—specifically mule kick, warm-up 18, and warm-up 19— causing L θ b to fight the data terms; for these three sequences, we lower λ θ b so the body regularizer does not bias the reconstruction away from the markers. Because marker placement on the skin varies by a few millimeters between subjects, we manually adjust each marker’s position on the GHUM template at two points: once per subject at the start of the capture, before optimization, and again whenever a marker detaches and is reattached, using a dedicated T-pose sequence recorded afterwards to re-register that marker. All reconstructions in Fit3D were visually inspected from multiple viewpoints during the development of this pipeline, and the loss list above is the outcome of that iterative refinement. Figure 2 shows example fits of the GHUM model, overlaid on the four synchronized RGB views, for two exercise sequences.

3.3. Comparison with MoSh++

We also fit each frame with MoSh++ [51], the de facto standard for marker-driven body fitting, and find it insufficient for fitness motions: because MoSh++ consumes only VICON marker positions and ignores the synchronized RGB views, it has no signal on the fingers and leaves the hands in a default rest pose regardless of what the subject is gripping. Our pipeline addresses this gap directly through the multi-view 2D keypoint term L 2 D , which propagates the visible finger configuration in the RGB streams into the GHUM hand joints. Figure 3 shows the contrast on a single frame seen from the four synchronized RGB cameras of our capture setup: our reconstruction (left two columns) versus MoSh++ (right two columns), with a zoomed inset on the hand region of each panel.
  • Quantitative comparison.
We quantify both effects on all eight Fit3D training subjects (Table 2). To measure hand articulation, we reproject the 21 hand joints of each fit into the four RGB views and compute their distance to the corresponding 2D hand keypoints from a strong whole-body keypoint detector [52], different from the 2D keypoint predictor used to fit the GHUM annotations (Section 3.2); aggregated across the four cameras, our fits reproject 13 % closer than MoSh++ ( 8.72 vs. 9.83  px), and MoSh++ is the worse of the two on 70 % of all hands. As a body-wide fitting residual, we additionally report, for every (cleaned) VICON marker, the shortest distance to the fitted body surface—the GHUM mesh for our method, and the SMPL-X mesh that MoSh++ outputs for MoSh++. Our surface lies 8.49  mm from the markers on average (median 6.0 ) against 10.53  mm for MoSh++ (median 9.9 ), with the largest differences at the extremities: our fits place the ankles and toes within 3–5 mm of the markers, versus 12–14 mm for MoSh++. MoSh++ models each marker as sitting at a fixed learned offset off the body surface, so part of this gap reflects that modeling choice rather than fitting error; the extremity differences, however, exceed any uniform offset and follow from the hands and feet being left in less accurate configurations.

4. Evaluation Protocol and Benchmark

To help improve the state of the art in 3D human reconstruction for fitness training, we introduce a public challenge on the test set of Fit3D. We standardize the evaluation protocol and metrics to facilitate the comparison of different research methods.
Our evaluation server accepts multiple formats for the predicted reconstructions: GHUM [12] parameters, SMPL-X [13] parameters, and 3D joints in the Human3.6M format with 17 joints. The released test set contains only one random camera viewpoint per fitness sequence, so multi-view triangulation in reconstruction is not possible. Each prediction has to be provided in the coordinate system of the released camera. Predictions are required only on a subset of frames, equally spaced through each test sequence.
We report the following metrics, averaged over the sampled test frames:
  • MPJPE: the mean per-joint position error, with and without Procrustes alignment, computed on the Human3.6M 17-joint configuration.
  • MPVPE: the mean per-vertex position error, with and without Procrustes alignment, computed on the SMPL-X mesh when the submission provides SMPL-X parameters and on the GHUM mesh when it provides GHUM parameters.
  • MPVPE-T (Mean Per-Vertex Position Error, T-pose): the predicted and reference bodies are posed in the same neutral T-pose using only their shape ( β ) parameters, and the distance between corresponding mesh vertices is averaged. Measures body-shape accuracy with pose removed.
  • MPJAE (Mean Per-Joint Angle Error) [6]: for each SMPL-X body joint, the geodesic angle between the predicted and reference orientation (the smallest rotation aligning one to the other), averaged over joints. Measures rotation/pose accuracy, independent of limb lengths. We report it both without alignment (including overall body facing) and in its Procrustes-aligned form (PA-MPJAE), which first rotates the whole predicted body to best align it with the reference and so scores the articulated pose alone.
  • Translation error: the L 2 distance between the predicted and reference root translation in the camera coordinate system.
Table 3 reports the performance of several recent monocular methods on the Fit3D test set. Every number in the table comes from our own runs: none are copied from the originating papers. For each of the 19 methods, we obtain its publicly available open-source implementation, run it ourselves on the Fit3D test frames under a single common protocol, and submit the resulting predictions to the evaluation server, which computes the reported metrics. Running every method ourselves—rather than collating self-reported numbers—is what makes the comparison controlled: all methods receive identical input frames, are evaluated on the same sampled test set, and are scored under the same metric definitions. When a method outputs only SMPL predictions, we convert them to SMPL-X before computing the vertex metrics; this conversion carries a mean vertex-to-vertex residual of approximately 3–5 mm, which slightly penalizes SMPL-output methods on MPVPE relative to methods that predict SMPL-X natively. REMIPS predicts GHUM natively; we submit its predictions in both GHUM and SMPL-X formats and, for comparability with the other rows, report its SMPL-X evaluation in the table.
Newer learned methods substantially improve over earlier optimization- and regression-based approaches: NLF [2] lowers the MPJPE from 119.45 mm (REMIPS) to 59.81  mm and the SMPL-X MPVPE-PA to 25.78  mm. The Fit3D training set has since been adopted by several recent large-scale methods: as reflected in the Trained on Fit3D column, the two methods that include Fit3D in their training data lead most reported metrics: NLF attains the lowest joint-position and vertex-position errors (and the lowest translation error among the methods evaluated under the true Fit3D intrinsics), while SMPLest-X attains the lowest joint-orientation errors; only the body-shape metric (MPVPE-T) is led by a method not trained on Fit3D (CameraHMR, by a 2 mm margin over SMPLest-X). This correlation should be read with care, however, NLF and SMPLest-X also differ from the other entries in architecture and in the remainder of their training corpora—NLF in particular leads several benchmarks it was not trained on—so the benchmark alone cannot attribute their lead to Fit3D. The controlled fine-tuning experiment of Section 6 isolates the contribution of Fit3D supervision from these confounding factors. Regardless of attribution, reconstructing fitness motions remains challenging, with the best non-aligned per-joint error still close to 60 mm, reflecting the out-of-distribution poses present in Fit3D.
Translation error varies by orders of magnitude across methods. For most of the benchmarked methods we obtain the root translation under the true Fit3D calibration—passing the intrinsics to the method where its interface accepts them (e.g., NLF, CameraHMR, GVHMR, SAM 3D Body), or re-solving the translation from the method’s image-aligned outputs under the true camera where it does not (e.g., 4DHumans, PARE, HybrIK); the resulting errors cluster at 185–350 mm (SMPLify-X excepted), reflecting the inherent monocular depth ambiguity, which grows on bent and ground-level poses. The remaining methods recover depth under their own camera model—a fixed 5000-pixel crop focal length for SMPLest-X, a self-estimated focal length for BLADE and PromptHMR, an image-diagonal focal length for WHAM and CLIFF—so their translation error measures the mismatch between assumed and true intrinsics rather than reconstruction quality: SMPLest-X places the body roughly an order of magnitude too deep, while CRMH predicts a near-zero root translation, so its error is dominated by the subject’s distance from the camera.
The benchmark, data, and submission instructions are publicly available at https://fit3d.imar.ro/ (accessed on 5 July 2026).

5. Analysis: Where Current Methods Succeed and Fail on Fit3D

Beyond the aggregate ranking of Table 3, the complementary metrics and a per-sequence breakdown over the 141 test sequences (3 subjects × 47 exercises) reveal where current methods break and what a fitness-capable reconstructor must improve.
  • No single method wins on all axes.
Position, orientation, and shape accuracy decouple. NLF attains the lowest joint- and vertex-position errors (MPJPE 59.81 , PA-MPJPE 34.44 , MPVPE-PA 25.78  mm), yet SMPLest-X is markedly the most accurate in joint orientation (MPJAE 9 . 32 ° , PA-MPJAE 8 . 77 ° , versus 12– 13 ° for NLF), and CameraHMR produces the most accurate body shape (MPVPE-T 23.61  mm). The ranking, therefore, reorders with the chosen axis, confirming that the orientation (MPJAE, PA-MPJAE) and shape (MPVPE-T) metrics capture failure modes that MPJPE alone does not: a method can localize joints well while misorienting limbs or defaulting to a generic body.
  • Most of the ranking is a statistical tie.
The differences near the top of Table 3 are smaller than the spread across sequences. A paired Wilcoxon signed-rank test on the per-sequence MPJPE shows that the six methods spanning 70.3 71.9  mm—CameraHMR, SAM 3D Body, TokenHMR, SMPLest-X, ScoreHMR, and GVHMR—are statistically indistinguishable from their immediate neighbors ( p 0.05 for each successive pair), whereas NLF is clearly separated from the field ( p < 10 3 ). Sub-millimeter differences in aggregate MPJPE on Fit3D are, therefore, not meaningful in isolation; improvements should be demonstrated on the hard regime, where the spread is large.
  • Difficulty is strongly action-conditioned.
Averaged over all methods, the hardest exercises incur roughly double the error of the easiest (Table 4). The hardest are dynamic, ground-level, and inverted poses with heavy self-occlusion—mule kick, man maker, burpees, push-up, and the forward-bend warm-up 1 (118–155 mm). The easiest are upright, low-articulation exercises such as standing dumbbell curls and most other warm-up routines (∼67 mm). This gap persists even for the strongest method, and it is this out-of-distribution tail—not the near-upright exercises, which are largely solved—that future work should target. These are precisely the poses that the corpora dominating current training pipelines (Human3.6M, 3DPW, BEDLAM, AGORA) cover poorly. Figure 4 makes this concrete on an inverted mule kick frame: the Fit3D-trained NLF and SMPLest-X recover the handstand, whereas 4DHumans exhibits a large global-orientation error and the earlier PARE regressor fails to recover the pose.
  • Body shape is barely personalized.
The T-pose metric exposes that most methods cluster at MPVPE-T 30  mm: they recover a near-mean body rather than a subject-specific one, and even the best (CameraHMR, 23.6  mm) leaves a sizeable residual. This is an under-addressed axis. As Section 6 shows, adding Fit3D to the training mixture lowers HMR2.0’s MPVPE-T from 47.4 to 38.4  mm (Table 5), i.e., the network begins to predict the released per-subject shapes instead of defaulting to the mean. This gain comes less from additional pose data than from the in-domain camera intrinsics used in our supervision: because the Fit3D annotations are metric under the true focal length, training on them removes the depth–scale ambiguity and lets the network estimate translation and body shape more accurately.
  • The residual error is genuine, not only a camera artifact.
As noted above, absolute root translation is comparable only among the methods that consume the true intrinsics and is dominated by internal camera assumptions for the rest. Crucially, however, the error does not vanish once placement is removed: across the learned methods, PA-MPJPE remains 34–68 mm and PA-MPJAE 9– 19 ° , so even after a global rigid alignment, a substantial articulation-and-shape gap persists. Fit3D releases per-sequence camera calibration, and a clear direction is for methods to consume it and produce metrically placed reconstructions rather than presuming a fixed focal length.
  • Takeaways for future methods.
Three priorities emerge. (i) Personalized body shape: most regressors default to a mean body, so subject-specific shape estimation is largely open. (ii) The hard-pose tail: ground-level, inverted, weight-loaded, and self-occluded articulations, where errors roughly double. (iii) Metric placement under known intrinsics. We further note that the classical optimization baseline (SMPLify-X, MPJPE 193 mm) trails the learned regressors by a wide margin under these out-of-distribution poses, and that the two methods trained with Fit3D (NLF and SMPLest-X) lead the position and orientation axes, respectively—underscoring in-domain data as the strongest available lever. Finally, because the released protocol samples sparse, equally spaced frames, it does not reward temporal consistency; this is why the video-based methods (GVHMR, WHAM, PromptHMR) are not advantaged here, and a dense-frame temporal track is a natural extension of the benchmark.

6. Impact of Fit3D Training: Fine-Tuning HMR2.0

The benchmark in Section 4 suggests that methods that include Fit3D in their training corpus achieve the lowest position and orientation errors. To isolate the effect of adding Fit3D from confounding factors such as differences in architecture and other training data, we run a controlled fine-tuning experiment, evaluating the resulting checkpoints both on the Fit3D test set (Table 5) and on AIST++ [54] as an out-of-domain generalization test on street dance motions (Table 6).
  • Fine-tuning setup.
We initialize from the public 4DHumans/HMR2.0 [1] model (ViT-H backbone) and fine-tune it on Fit3D-train mixed with the original HMR2.0 training corpus. To prevent catastrophic forgetting, Fit3D is given weight 0.3 in the sampling mixture and the original datasets (Human3.6M, COCO, MPII, AI-Challenger, AVA, MPI-INF-3DHP, InstaVariety) share the remaining 0.7 ; the adversarial body-pose prior of HMR2.0 is retained. We optimize with AdamW (learning rate 10 5 , weight decay 10 4 , mixed precision) and an effective batch of 256 (64 per GPU on 4 GPUs), evaluating the checkpoints at 30k, 60k, and 90k steps. Since Fit3D provides SMPL-X pseudo-ground-truth whereas HMR2.0 is trained on SMPL, we convert each frame’s SMPL-X fit to SMPL with the optimization-based transfer code provided by SMPL-X [13]—the same procedure we use to convert our GHUM fits to SMPL-X when building the annotations, after registering the topology of the two meshes; Fit3D-train is sub-sampled to 5 fps, yielding ∼178k images over its 8 subjects × 47 actions × 4 camera views.
On the Fit3D test set, every metric improves substantially after fine-tuning. MPVPE-T drops from 47.4 to 38.4 mm, indicating that the network learns to predict per-subject body shape parameters β that match the reference shapes released with our annotations, rather than defaulting to the mean shape produced by the released checkpoint. MPJAE drops from 16 . 1 ° to 10 . 4 ° and PA-MPJAE from 15 . 2 ° to 10 . 1 ° , a relative reduction of approximately 33 % , tracking the gains in joint and vertex position errors. The improvements largely saturate by 30k–60k steps, with the 90k checkpoint offering only marginal additional gains.
  • The gains concentrate on the hard poses.
Section 5 showed that errors on Fit3D are strongly action-conditioned. Breaking the fine-tuning improvement down per exercise (Figure 5) shows that it is exactly the hard actions that benefit most: the per-action MPJPE reduction is positively correlated with the released model’s per-action error (Pearson r = 0.70 ), and the ten hardest exercises improve by 12.3 mm on average against only 4.2 mm for the ten easiest. In other words, Fit3D supervision helps precisely where current methods fail—on the out-of-distribution, heavily articulated poses that the standard training corpora do not cover—rather than uniformly shifting an already-well-reconstructed regime.
On AIST++, all eight metrics remain within noise across fine-tuning checkpoints—MPJAE stays around 23 ° and PA-MPJAE around 21 . 5 ° throughout—confirming that fine-tuning on Fit3D does not degrade out-of-domain generalization. Two cross-dataset caveats apply: (†) MPVPE-T on AIST++ measures only the deviation of the predicted shape from neutral, because AIST++ ground truth uses β = 0 , so the column is not directly comparable with the Fit3D MPVPE-T column and MPVPE / MPVPE-PA on AIST++ are computed against SMPL ground truth, whereas on Fit3D they are computed against SMPL-X.

7. Discussion

The benchmark and the fine-tuning study together point to in-domain supervision as the strongest available lever for fitness-domain reconstruction, with personalized body shape, the hard-pose tail, and metric placement under known intrinsics as the main open axes (Section 5). At the same time, four caveats bound the contributions of this work.
First, the released GHUM/SMPL-X annotations are pseudo-ground-truth obtained by optimization rather than by direct measurement: the loss weights λ were hand-tuned, the body prior was down-weighted per sequence on out-of-distribution motions, and the fits were validated by multi-view visual inspection together with the marker-to-surface residual ( 8.49 mm mean, Section 3.3) rather than against an independent sensor, so a small residual error on the most extreme poses cannot be excluded. Second, the hand-articulation comparison of Section 3.3 scores each fit against 2D keypoint detections, the same modality our pipeline optimizes through L 2 D (whereas MoSh++ ignores it); it should, therefore, be read as a consistency check that favors multi-view-supervised fits over marker-only fits, not as fully independent validation. Third, Fit3D is a single-person, indoor laboratory capture of 13 subjects in tight gym attire under controlled studio lighting and fixed calibrated cameras; multi-person interaction, outdoor and truly in-the-wild imagery, loose or varied clothing, uncontrolled capture conditions, and a broader demographic range lie outside its scope and outside what the benchmark currently measures, so the generalization of models trained on these annotations to real-world fitness videos remains to be established. Because the train/validation/test split is subject-disjoint, our benchmark and fine-tuning results already measure generalization to unseen subjects, but not across these other axes of variation. Fourth, the benchmark protocol scores sparsely sampled frames independently, which is fair to image-based methods but does not credit the temporal consistency that video-based methods can provide; adding temporal-consistency metrics (e.g., acceleration error) and dense per-sequence evaluation is a natural extension for future versions of the benchmark.

8. Conclusions

Building on the Fit3D fitness dataset of Fieraru et al. [11], this paper has contributed three extensions for the monocular 3D human reconstruction community. First, we have described the methodology used to construct dense GHUM and SMPL-X pseudo-ground-truth shape and pose annotations on top of the raw Fit3D MoCap, a marker- and multi-view-keypoint-driven fitting pipeline regularized by separate body and hand normalizing-flow priors and a self-collision loss, designed to remain stable under the extreme articulations and self-occlusions characteristic of fitness motions. Second, we have defined a standardized evaluation protocol—metric set, frame sampling, and coordinate conventions, served through the IMAR-hosted Fit3D evaluation server—and conducted a comparative benchmark study of 19 representative methods spanning optimization-based, single-frame, and video-based families, finding that the two methods that include Fit3D in their training corpus lead the position and orientation metrics, respectively. Third, a controlled fine-tuning experiment showed that adding Fit3D to the training mixture of a strong baseline improves it most on exactly the hard, out-of-distribution poses where current methods fail, without regressing out-of-domain performance—direct evidence that in-domain supervision is the strongest available lever. The Fit3D dataset and the GHUM/SMPL-X annotations are available, under a non-commercial research license, through the IMAR Fit3D resource [11] at https://fit3d.imar.ro/ (accessed on 5 July 2026); to date (June 2026), 1088 academics have registered for access, indicating broad interest in the dataset and benchmark.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Informed consent was obtained from all subjects participating in the Fit3D motion capture sessions.

Data Availability Statement

Restrictions apply to the availability of these data. Data were obtained from the Institute of Mathematics of the Romanian Academy (IMAR) and are available at https://fit3d.imar.ro/ (accessed on 5 July 2026) [11] with the permission of IMAR. The Fit3D dataset is the property of the Institute of Mathematics of the Romanian Academy (IMAR) and is available under a non-commercial research license. The GHUM/SMPL-X annotations and the evaluation protocol described in this paper are accessible through that same resource, subject to that license. Users of the dataset and annotations are required to cite the original Fit3D publication [11].

Acknowledgments

The author would like to thank the creators of the original Fit3D dataset, Fieraru et al. [11], and the Institute of Mathematics of the Romanian Academy (IMAR) for providing the dataset and the raw motion-capture data under a non-commercial research license, without which this work would not have been possible. The author also thanks the subjects who participated in the capture sessions.

Conflicts of Interest

The author declares no conflicts of interest.

References

  1. Goel, S.; Pavlakos, G.; Rajasegaran, J.; Kanazawa, A.; Malik, J. Humans in 4D: Reconstructing and Tracking Humans with Transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2–6 October 2023; pp. 14783–14794. [Google Scholar]
  2. Sárándi, I.; Pons-Moll, G. Neural Localizer Fields for Continuous 3D Human Pose and Shape Estimation. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 10–15 December 2024. [Google Scholar]
  3. Patel, P.; Black, M.J. Camerahmr: Aligning people with perspective. In Proceedings of the 2025 International Conference on 3D Vision (3DV); IEEE: Piscataway, NJ, USA, 2025; pp. 1562–1571. [Google Scholar]
  4. Yin, W.; Cai, Z.; Wang, R.; Zeng, A.; Wei, C.; Sun, Q.; Mei, H.; Wang, Y.; Pang, H.E.; Zhang, M.; et al. SMPLest-X: Ultimate Scaling for Expressive Human Pose and Shape Estimation. arXiv 2025, arXiv:2501.09782. [Google Scholar]
  5. Yang, X.; Kukreja, D.; Pinkus, D.; Fan, T.; Park, J.; Shin, S.; Cao, J.; Liu, J.W.; Ugrinovic, N.; Sagar, A.; et al. Sam 3d body: Robust full-body human mesh recovery. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Denver, CO, USA, 3–7 June 2026; pp. 7209–7219. [Google Scholar]
  6. von Marcard, T.; Henschel, R.; Black, M.; Rosenhahn, B.; Pons-Moll, G. Recovering Accurate 3D Human Pose in The Wild Using IMUs and a Moving Camera. In Proceedings of the ECCV, Munich, Germany, 8–14 September 2018. [Google Scholar]
  7. Kaufmann, M.; Song, J.; Guo, C.; Shen, K.; Jiang, T.; Tang, C.; Zárate, J.J.; Hilliges, O. EMDB: The Electromagnetic Database of Global 3D Human Pose and Shape in the Wild. In Proceedings of the International Conference on Computer Vision (ICCV), Paris, France, 2–6 October 2023. [Google Scholar]
  8. Ionescu, C.; Papava, D.; Olaru, V.; Sminchisescu, C. Human3.6M: Large Scale Datasets and Predictive Methods for 3D Human Sensing in Natural Environments. IEEE Trans. Pattern Anal. Mach. Intell. 2014, 36, 1325–1339. [Google Scholar] [CrossRef] [Scilit]
  9. Black, M.J.; Patel, P.; Tesch, J.; Yang, J. BEDLAM: A Synthetic Dataset of Bodies Exhibiting Detailed Lifelike Animated Motion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 17–24 June 2023. [Google Scholar]
  10. Patel, P.; Huang, C.H.P.; Tesch, J.; Hoffmann, D.T.; Tripathi, S.; Black, M.J. AGORA: Avatars in Geography Optimized for Regression Analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021. [Google Scholar]
  11. Fieraru, M.; Zanfir, M.; Pirlea, S.C.; Olaru, V.; Sminchisescu, C. AIFit: Automatic 3D Human-Interpretable Feedback Models for Fitness Training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021. [Google Scholar]
  12. Xu, H.; Bazavan, E.G.; Zanfir, A.; Freeman, W.T.; Sukthankar, R.; Sminchisescu, C. GHUM & GHUML: Generative 3D Human Shape and Articulated Pose Models. In Proceedings of the CVPR, Seattle, WA, USA, 13–19 June 2020. [Google Scholar]
  13. Pavlakos, G.; Choutas, V.; Ghorbani, N.; Bolkart, T.; Osman, A.A.A.; Tzionas, D.; Black, M.J. Expressive Body Capture: 3D Hands, Face, and Body From a Single Image. In Proceedings of the CVPR, Long Beach, CA, USA, 15–20 June 2019. [Google Scholar]
  14. Popa, A.I.; Zanfir, M.; Sminchisescu, C. Deep Multitask Architecture for Integrated 2D and 3D Human Sensing. In Proceedings of the CVPR, Honolulu, HI, USA, 21–26 July 2017. [Google Scholar]
  15. Martinez, J.; Hossain, R.; Romero, J.; Little, J.J. A simple yet effective baseline for 3d human pose estimation. In Proceedings of the ICCV, Venice, Italy, 22–29 October 2017. [Google Scholar]
  16. Luvizon, D.C.; Picard, D.; Tabia, H. 2d/3d pose estimation and action recognition using multitask deep learning. In Proceedings of the CVPR, Salt Lake City, UT, USA, 18–23 June 2018; pp. 5137–5146. [Google Scholar]
  17. Yang, W.; Ouyang, W.; Wang, X.; Ren, J.; Li, H.; Wang, X. 3D Human Pose Estimation in the Wild by Adversarial Learning. In Proceedings of the CVPR, Salt Lake City, UT, USA, 18–23 June 2018. [Google Scholar]
  18. Mehta, D.; Sotnychenko, O.; Mueller, F.; Xu, W.; Sridhar, S.; Pons-Moll, G.; Theobalt, C. Single-Shot Multi-person 3D Pose Estimation from Monocular RGB. In Proceedings of the 3DV, Verona, Italy, 5–8 September 2018. [Google Scholar]
  19. Zanfir, A.; Marinoiu, E.; Sminchisescu, C. Monocular 3D Pose and Shape Estimation of Multiple People in Natural Scenes—The Importance of Multiple Scene Constraints. In Proceedings of the CVPR, Salt Lake City, UT, USA, 18–23 June 2018. [Google Scholar]
  20. Zanfir, A.; Marinoiu, E.; Zanfir, M.; Popa, A.I.; Sminchisescu, C. Deep Network for the Integrated 3D Sensing of Multiple People in Natural Images. In Proceedings of the NIPS; Curran Associates Inc.: Red Hook, NY, USA, 2018. [Google Scholar]
  21. Li, J.; Wang, C.; Zhu, H.; Mao, Y.; Fang, H.S.; Lu, C. Crowdpose: Efficient crowded scenes pose estimation and a new benchmark. In Proceedings of the CVPR, Long Beach, CA, USA, 15–20 June 2019. [Google Scholar]
  22. Su, K.; Yu, D.; Xu, Z.; Geng, X.; Wang, C. Multi-Person Pose Estimation with Enhanced Channel-wise and Spatial Information. In Proceedings of the CVPR, Long Beach, CA, USA, 15–20 June 2019. [Google Scholar]
  23. Benzine, A.; Luvison, B.; Pham, Q.C.; Achard, C. Deep, Robust and Single Shot 3D Multi-Person Human Pose Estimation from Monocular Images. In Proceedings of the ICIP; IEEE: Piscataway, NJ, USA, 2019. [Google Scholar]
  24. Kolotouros, N.; Pavlakos, G.; Black, M.J.; Daniilidis, K. Learning to Reconstruct 3D Human Pose and Shape via Model-fitting in the Loop. In Proceedings of the ICCV, Seoul, Republic of Korea, 27 October–2 November 2019. [Google Scholar]
  25. Fieraru, M.; Zanfir, M.; Oneata, E.; Popa, A.I.; Olaru, V.; Sminchisescu, C. Three-Dimensional Reconstruction of Human Interactions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020. [Google Scholar]
  26. Fieraru, M.; Zanfir, M.; Oneata, E.; Popa, A.I.; Olaru, V.; Sminchisescu, C. Learning complex 3d human self-contact. In Proceedings of the AAAI Conference on Artificial Intelligence, Virtual, 2–9 February 2021; Volume 35, pp. 1343–1351. [Google Scholar]
  27. Pang, Y.; Xie, J.; Khan, M.H.; Anwer, R.M.; Khan, F.S.; Shao, L. Mask-guided attention network for occluded pedestrian detection. In Proceedings of the ICCV, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 4967–4975. [Google Scholar]
  28. Rasouli, A.; Kotseruba, I.; Kunic, T.; Tsotsos, J.K. PIE: A Large-Scale Dataset and Models for Pedestrian Intention Estimation and Trajectory Prediction. In Proceedings of the ICCV, Seoul, Republic of Korea, 27 October–2 November 2019. [Google Scholar]
  29. Ge, Y.; Zhang, R.; Wang, X.; Tang, X.; Luo, P. DeepFashion2: A Versatile Benchmark for Detection, Pose Estimation, Segmentation and Re-Identification of Clothing Images. In Proceedings of the CVPR, Long Beach, CA, USA, 15–20 June 2019. [Google Scholar]
  30. Sidnev, A.; Trushkov, A.; Kazakov, M.; Korolev, I.; Sorokin, V. DeepMark: One-Shot Clothing Detection. In Proceedings of the ICCV Workshops, Seoul, Republic of Korea, 27 October–2 November 2019. [Google Scholar]
  31. Shahroudy, A.; Liu, J.; Ng, T.T.; Wang, G. Ntu rgb+ d: A large scale dataset for 3d human activity analysis. In Proceedings of the CVPR, Las Vegas, NV, USA, 27–30 June 2016; pp. 1010–1019. [Google Scholar]
  32. Huang, C.H.P.; Yi, H.; Höschle, M.; Safroshkin, M.; Alexiadis, T.; Polikovsky, S.; Scharstein, D.; Black, M.J. Capturing and Inferring Dense Full-Body Human-Scene Contact. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022. [Google Scholar]
  33. Tripathi, S.; Müller, L.; Huang, C.H.P.; Taheri, O.; Black, M.J.; Tzionas, D. 3D Human Pose Estimation via Intuitive Physics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 17–24 June 2023. [Google Scholar]
  34. Kocabas, M.; Huang, C.H.P.; Hilliges, O.; Black, M.J. PARE: Part attention regressor for 3D human body estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 11–17 October 2021; pp. 11127–11137. [Google Scholar]
  35. Li, J.; Xu, C.; Chen, Z.; Bian, S.; Yang, L.; Lu, C. Hybrik: A hybrid analytical-neural inverse kinematics solution for 3d human pose and shape estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 3383–3393. [Google Scholar]
  36. Zhang, H.; Tian, Y.; Zhang, Y.; Li, M.; An, L.; Sun, Z.; Liu, Y. Pymaf-x: Towards well-aligned full-body model regression from monocular images. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 12287–12303. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  37. Li, Z.; Liu, J.; Zhang, Z.; Xu, S.; Yan, Y. Cliff: Carrying location information in full frames into human pose and shape estimation. In Proceedings of the Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, 23–27 October 2022; Proceedings, Part V; Springer: Cham, Switzerland, 2022; pp. 590–606. [Google Scholar]
  38. Dwivedi, S.K.; Sun, Y.; Patel, P.; Feng, Y.; Black, M.J. Tokenhmr: Advancing human mesh recovery with a tokenized pose representation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 1323–1333. [Google Scholar]
  39. Baradel, F.; Armando, M.; Galaaoui, S.; Brégier, R.; Weinzaepfel, P.; Rogez, G.; Lucas, T. Multi-hmr: Multi-person whole-body human mesh recovery in a single shot. In Proceedings of the European Conference on Computer Vision; Springer: Cham, Switzerland, 2024; pp. 202–218. [Google Scholar]
  40. Wang, S.; Li, J.; Li, T.; Yuan, Y.; Fuchs, H.; Nagano, K.; De Mello, S.; Stengel, M. Blade: Single-view body mesh estimation through accurate depth estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 11–15 June 2025; pp. 21991–22000. [Google Scholar]
  41. Shin, S.; Kim, J.; Halilaj, E.; Black, M.J. WHAM: Reconstructing World-grounded Humans with Accurate 3D Motion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024. [Google Scholar]
  42. Shen, Z.; Pi, H.; Xia, Y.; Cen, Z.; Peng, S.; Hu, Z.; Bao, H.; Hu, R.; Zhou, X. World-grounded human motion recovery via gravity-view coordinates. In Proceedings of the SIGGRAPH Asia 2024 Conference Papers; Association for Computing Machinery: New York, NY, USA, 2024; pp. 1–11. [Google Scholar]
  43. Wang, Y.; Sun, Y.; Patel, P.; Daniilidis, K.; Black, M.J.; Kocabas, M. PromptHMR: Promptable Human Mesh Recovery. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 11–15 June 2025. [Google Scholar]
  44. Stathopoulos, A.; Han, L.; Metaxas, D. Score-guided diffusion for 3d human recovery. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 906–915. [Google Scholar]
  45. Sengupta, A.; Budvytis, I.; Cipolla, R. Synthetic Training for Accurate 3D Human Pose and Shape Estimation. In Proceedings of the British Machine Vision Conference (BMVC), Virtual, 7–10 September 2020. [Google Scholar]
  46. Zanfir, M.; Zanfir, A.; Bazavan, E.G.; Freeman, W.T.; Sukthankar, R.; Sminchisescu, C. Thundr: Transformer-based 3d human reconstruction with markers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 11–17 October 2021; pp. 12971–12980. [Google Scholar]
  47. Zanfir, A.; Bazavan, E.G.; Zanfir, M.; Freeman, W.T.; Sukthankar, R.; Sminchisescu, C. Neural Descent for Visual 3D Human Pose and Shape. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; pp. 14484–14493. [Google Scholar]
  48. Fieraru, M.; Zanfir, M.; Szente, T.; Bazavan, E.; Olaru, V.; Sminchisescu, C. REMIPS: Physically Consistent 3D Reconstruction of Multiple Interacting People under Weak Supervision. In Proceedings of the Advances in Neural Information Processing Systems; Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., Vaughan, J.W., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2021; Volume 34, pp. 19385–19397. [Google Scholar]
  49. Jacobson, A.; Kavan, L.; Sorkine-Hornung, O. Robust inside-outside segmentation using generalized winding numbers. ACM Trans. Graph. (TOG) 2013, 32, 33. [Google Scholar] [CrossRef] [Scilit]
  50. Zanfir, A.; Bazavan, E.G.; Xu, H.; Freeman, W.T.; Sukthankar, R.; Sminchisescu, C. Weakly supervised 3d human pose and shape reconstruction with normalizing flows. In Proceedings of the European Conference on Computer Vision; Springer: Cham, Switzerland, 2020; pp. 465–481. [Google Scholar]
  51. Mahmood, N.; Ghorbani, N.; Troje, N.F.; Pons-Moll, G.; Black, M.J. AMASS: Archive of Motion Capture as Surface Shapes. In Proceedings of the ICCV, Seoul, Republic of Korea, 27 October–2 November 2019. [Google Scholar]
  52. Yang, Z.; Zeng, A.; Yuan, C.; Li, Y. Effective Whole-body Pose Estimation with Two-stages Distillation. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Paris, France, 2–6 October 2023; pp. 4210–4220. [Google Scholar]
  53. Jiang, W.; Kolotouros, N.; Pavlakos, G.; Zhou, X.; Daniilidis, K. Coherent reconstruction of multiple humans from a single image. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 5579–5588. [Google Scholar]
  54. Li, R.; Yang, S.; Ross, D.A.; Kanazawa, A. AI Choreographer: Music Conditioned 3D Dance Generation with AIST++. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 11–17 October 2021; pp. 13401–13412. [Google Scholar]
Figure 1. Overview of the contributions. From left: an input frame from the Fit3D test set (subject s02, man maker—a dynamic, ground-level compound exercise); the GHUM/SMPL-X pseudo-ground-truth annotation for the frame, obtained by the multi-view marker- and keypoint-driven optimization of Section 3.2; predictions of two of the 19 methods evaluated under our benchmark protocol (Section 4)—NLF, the strongest method, trained on Fit3D, and PARE, which fails on this pose and the effect of fine-tuning HMR2.0 with Fit3D supervision (Section 6), which lowers the error on the ten hardest exercises by 12.3  mm on average, versus 4.2  mm on the ten easiest. The two predicted meshes are rendered at a common root depth, as in the qualitative comparison of Section 5.
Figure 1. Overview of the contributions. From left: an input frame from the Fit3D test set (subject s02, man maker—a dynamic, ground-level compound exercise); the GHUM/SMPL-X pseudo-ground-truth annotation for the frame, obtained by the multi-view marker- and keypoint-driven optimization of Section 3.2; predictions of two of the 19 methods evaluated under our benchmark protocol (Section 4)—NLF, the strongest method, trained on Fit3D, and PARE, which fails on this pose and the effect of fine-tuning HMR2.0 with Fit3D supervision (Section 6), which lowers the error on the ten hardest exercises by 12.3  mm on average, versus 4.2  mm on the ten easiest. The two predicted meshes are rendered at a common root depth, as in the qualitative comparison of Section 5.
Jimaging 12 00311 g001
Figure 2. Motion and shape capture results. The fitted GHUM mesh is overlaid on each of the four synchronized RGB views (shown as a 2 × 2 grid per frame), for three frames of two exercise sequences: burpees (top) and mule kick (bottom). The latter illustrates the robustness of the fitting pipeline to extreme articulations and significant self-occlusions.
Figure 2. Motion and shape capture results. The fitted GHUM mesh is overlaid on each of the four synchronized RGB views (shown as a 2 × 2 grid per frame), for three frames of two exercise sequences: burpees (top) and mule kick (bottom). The latter illustrates the robustness of the fitting pipeline to extreme articulations and significant self-occlusions.
Jimaging 12 00311 g002
Figure 3. Hand articulation. Our multi-view motion and shape capture pipeline (left two columns) vs. MoSh++ [51] (right two columns), shown on a single frame seen from the four synchronized RGB cameras of our capture setup, arranged 2 × 2 per method. Our fits resolve realistic finger poses around the grip because L 2 D ties the GHUM hand joints to multi-view RGB keypoint detections; MoSh++ relies on VICON marker positions only and leaves the hands in a default rest pose. The bottom-left inset of each panel is a zoom on the corresponding pair of hands.
Figure 3. Hand articulation. Our multi-view motion and shape capture pipeline (left two columns) vs. MoSh++ [51] (right two columns), shown on a single frame seen from the four synchronized RGB cameras of our capture setup, arranged 2 × 2 per method. Our fits resolve realistic finger poses around the grip because L 2 D ties the GHUM hand joints to multi-view RGB keypoint detections; MoSh++ relies on VICON marker positions only and leaves the hands in a default rest pose. The bottom-left inset of each panel is a zoom on the corresponding pair of hands.
Jimaging 12 00311 g003
Figure 4. Qualitative reconstruction on a hard Fit3D pose (subject s02, mule kick—an inverted handstand-kick). From left: the input frame; the pseudo-ground-truth annotation and four benchmarked methods, each predicted SMPL-X mesh overlaid on the subject. To compare pose and shape independently of camera placement, all meshes are rendered at a common root depth: absolute placement is not comparable across methods because of the monocular depth–focal-length ambiguity, and is reported separately as the translation error. The Fit3D-trained methods (NLF, SMPLest-X) recover the inverted pose, while 4DHumans exhibits a large global-orientation error and PARE fails to recover the pose—the action-conditioned difficulty quantified in Table 4. The close agreement of the pseudo-ground-truth mesh with the subject also illustrates the quality of the released annotations.
Figure 4. Qualitative reconstruction on a hard Fit3D pose (subject s02, mule kick—an inverted handstand-kick). From left: the input frame; the pseudo-ground-truth annotation and four benchmarked methods, each predicted SMPL-X mesh overlaid on the subject. To compare pose and shape independently of camera placement, all meshes are rendered at a common root depth: absolute placement is not comparable across methods because of the monocular depth–focal-length ambiguity, and is reported separately as the translation error. The Fit3D-trained methods (NLF, SMPLest-X) recover the inverted pose, while 4DHumans exhibits a large global-orientation error and PARE fails to recover the pose—the action-conditioned difficulty quantified in Table 4. The close agreement of the pseudo-ground-truth mesh with the subject also illustrates the quality of the released annotations.
Jimaging 12 00311 g004
Figure 5. Fine-tuning gains concentrate on the hard actions. Each point is one of the 47 Fit3D exercises; the x-axis is the released HMR2.0 per-action MPJPE (a proxy for difficulty) and the y-axis is the per-action MPJPE reduction after fine-tuning on Fit3D-train (60k steps). The improvement grows with difficulty (Pearson r = 0.70 ): the hardest exercises gain the most, mirroring the action-conditioned difficulty of Section 5.
Figure 5. Fine-tuning gains concentrate on the hard actions. Each point is one of the 47 Fit3D exercises; the x-axis is the released HMR2.0 per-action MPJPE (a proxy for difficulty) and the y-axis is the per-action MPJPE reduction after fine-tuning on Fit3D-train (60k steps). The improvement grows with difficulty (Pearson r = 0.70 ): the hardest exercises gain the most, mirroring the action-conditioned difficulty of Section 5.
Jimaging 12 00311 g005
Table 1. Fit3D in the context of 3D human pose and shape datasets. With released GHUM/SMPL-X annotations on a fitness corpus, Fit3D is the only dataset combining high-fidelity marker MoCap, multi-view RGB, parametric whole-body pseudo-ground-truth, and a fitness focus. In the Hands column,  ✓ indicates that articulated hand annotations are provided and – that they are not.
Table 1. Fit3D in the context of 3D human pose and shape datasets. With released GHUM/SMPL-X annotations on a fitness corpus, Fit3D is the only dataset combining high-fidelity marker MoCap, multi-view RGB, parametric whole-body pseudo-ground-truth, and a fitness focus. In the Hands column,  ✓ indicates that articulated hand annotations are provided and – that they are not.
DatasetSubjectsFramesGT SourceCamerasBody-Model GTHandsSetting/Focus
Human3.6M [8]11 3.6 Moptical marker MoCap4 (+10 mocap)3D jointsindoor lab/daily activities
MPI-INF-3DHP [18]8 1.3 Mmarkerless multi-view MoCap143D jointsstudio + outdoor/general poses
3DPW [6]751 kIMU + handheld video1 movingSMPLin-the-wild/daily motion
EMDB [7]10∼100 kelectromagnetic sensors1 movingSMPLin-the-wild/world-grounded motion
RICH [32]22577 kmarkerless MoCap + scene scans6–8SMPL-Xindoor + outdoor/ human–scene contact
MOYO [33]1 1.75 Mmarker MoCap + pressure mat8SMPL-Xindoor lab/yoga, extreme poses
BEDLAM [9]271380 ksynthetic rendersyntheticSMPL-Xsynthetic/appearance diversity
AGORA [10]424017 ksynthetic render of 3D scanssyntheticSMPL & SMPL-Xsynthetic/ multi-person, occlusion
Fit3D (ours) [11]13 3.0 Mmarker MoCap + multi-view RGB + scans4 (+12 mocap)GHUM & SMPL-Xindoor lab/fitness exercises
Table 2. Quantitative comparison with MoSh++ on the eight Fit3D training subjects. Hand reproj.: mean distance (px) between projected 3D hand joints and detected 2D hand keypoints over the four RGB views. Marker–surf.: mean shortest distance (mm) from each (cleaned) VICON marker to the fitted body mesh. Lower is better (indicated by ↓); best in bold.
Table 2. Quantitative comparison with MoSh++ on the eight Fit3D training subjects. Hand reproj.: mean distance (px) between projected 3D hand joints and detected 2D hand keypoints over the four RGB views. Marker–surf.: mean shortest distance (mm) from each (cleaned) VICON marker to the fitted body mesh. Lower is better (indicated by ↓); best in bold.
Hand Reproj. (px) ↓Marker–Surf. (mm) ↓
MoSh++ [51] 9.83 10.53
Ours 8.72 8.49
Table 3. Benchmark of monocular 3D human reconstruction methods on the Fit3D test set, sorted by MPJPE (best first). All position errors are in millimeters and angle errors (MPJAE, PA-MPJAE) in degrees.  Translation reported under the method’s own camera model (a fixed or self-estimated focal length) rather than the true Fit3D intrinsics; these values measure camera-assumption mismatch and are not comparable with the other rows (see text). * NLF predicts a non-parametric Neural Localizer Field; we evaluate using its SMPL-X regression head. Lower is better for all metrics (indicated by ↓); the best value in each column is in bold.
Table 3. Benchmark of monocular 3D human reconstruction methods on the Fit3D test set, sorted by MPJPE (best first). All position errors are in millimeters and angle errors (MPJAE, PA-MPJAE) in degrees.  Translation reported under the method’s own camera model (a fixed or self-estimated focal length) rather than the true Fit3D intrinsics; these values measure camera-assumption mismatch and are not comparable with the other rows (see text). * NLF predicts a non-parametric Neural Localizer Field; we evaluate using its SMPL-X regression head. Lower is better for all metrics (indicated by ↓); the best value in each column is in bold.
MethodYearTrained
on Fit3D
InputOutputTransl. Err.
3D JointsSMPL-X
MPJPE
MPJPE-PA
MPVPE
MPVPE-PA
MPVPE-T
MPJAE
PA-MPJAE
NLF [2]2024YesImageSMPL-X * 185.16 59.81 34.44 51.41 25.78 31.31 12.83 12.59
CameraHMR [3]2025NoImageSMPL 212.29 70.31 39.69 59.80 32.64 23.61 14.87 14.31
SAM 3D Body [5]2026NoImageMHR 187.67 70.65 37.89 56.74 28.89 29.13 12.62 12.34
TokenHMR [38]2024NoImageSMPL 301.33 71.21 37.05 72.28 36.91 37.50 15.80 14.46
SMPLest-X [4]2025YesImageSMPL-X 34,668.77 71.51 40.35 54.60 26.86 25.66 9.32 8.77
ScoreHMR [44]2024NoImageSMPL 332.48 71.61 37.00 73.98 34.18 46.94 14.71 13.90
GVHMR [42]2024NoVideoSMPL-X 249.72 71.94 40.07 64.62 35.57 35.22 14.76 14.25
4DHumans [1]2023NoImageSMPL 337.33 74.48 37.56 77.26 38.37 47.43 16.14 15.16
PromptHMR [43]2025NoVideoSMPL-X 1078.62 75.09 42.51 62.17 32.79 23.71 13.51 12.85
WHAM [41]2024NoVideoSMPL 931.43 85.85 46.74 84.99 41.76 42.48 18.58 17.26
Multi-HMR [39]2024NoImageSMPL-X 342.65 89.16 50.07 77.71 39.82 34.42 15.66 14.66
PyMAF-X [36]2023NoImageSMPL-X 350.05 89.60 48.61 74.79 38.74 42.42 15.82 14.82
PARE [34]2021NoImageSMPL 296.86 89.90 53.74 85.10 46.84 35.65 18.33 16.47
BLADE [40]2025NoImageSMPL-X 2374.31 95.19 55.34 83.05 46.67 34.28 17.06 15.53
HybrIK [35]2021NoImageSMPL-X 322.85 97.56 50.76 91.45 41.60 36.12 18.44 17.16
CLIFF [37]2022NoImageSMPL 896.49 105.42 55.17 107.73 51.54 48.25 20.80 16.24
CRMH [53]2020NoImageSMPL 3283.89 116.19 62.41 120.12 59.86 56.52 23.53 18.52
REMIPS [48]2021NoImageGHUM 667.01 119.45 67.51 112.53 58.68 39.52 19.38 16.77
SMPLify-X [13]2019NoImageSMPL-X 1388.61 193.16 94.88 197.09 88.88 92.02 33.49 25.41
Table 4. Hardest and easiest Fit3D exercises, mean error over all 19 benchmarked methods and the 3 test subjects. Position errors in mm, MPJAE in degrees. The italicized rows are group labels separating the five hardest from the five easiest exercises.
Table 4. Hardest and easiest Fit3D exercises, mean error over all 19 benchmarked methods and the 3 test subjects. Position errors in mm, MPJAE in degrees. The italicized rows are group labels separating the five hardest from the five easiest exercises.
ExerciseMPJPEMPJPE-PAMPJAE
Hardest
mule kick 155.5 71.2 28.7
man maker 135.3 73.1 22.6
burpees 121.9 62.1 20.7
push-up 118.7 61.2 21.6
warm-up 1 118.0 61.6 19.2
Easiest
dumbbell biceps curls 71.5 39.3 15.6
warm-up 6 69.9 35.2 12.1
warm-up 14 69.1 39.1 14.9
warm-up 8 67.6 35.5 12.0
dumbbell hammer curls 67.0 35.6 12.5
Table 5. Fit3D test set: released 4D-Humans (HMR2.0) vs. fine-tuning on Fit3D-train (30k/60k/90k steps), all 141 test sequences. MPJPE/MPJPE-PA on joints3d (H36M-17); MPVPE/MPVPE-PA on posed SMPL-X meshes; MPVPE-T on the T-pose (shape only); MPJAE/PA-MPJAE on SMPL-X body-joint orientations. Lower is better; best in bold.
Table 5. Fit3D test set: released 4D-Humans (HMR2.0) vs. fine-tuning on Fit3D-train (30k/60k/90k steps), all 141 test sequences. MPJPE/MPJPE-PA on joints3d (H36M-17); MPVPE/MPVPE-PA on posed SMPL-X meshes; MPVPE-T on the T-pose (shape only); MPJAE/PA-MPJAE on SMPL-X body-joint orientations. Lower is better; best in bold.
Transl.
(mm)
MPJPE
(mm)
MPJPE-PA
(mm)
MPVPE
(mm)
MPVPE-PA
(mm)
MPVPE-T
(mm)
MPJAE
(°)
PA-MPJAE
(°)
Released (HMR2.0) 337.3 74.5 37.6 77.3 38.4 47.4 16.1 15.2
   + FT (30k) 243.5 66.2 35.9 53.3 24.3 38.4 10.8 10.4
   + FT (60k) 236.9 66.5 36.2 55.9 25.2 37.8 10.6 10.2
   + FT (90k) 249.7 66.7 36.0 56.0 24.3 39.4 10.4 10.1
Table 6. AIST++ (out-of-domain street dance): generalization after Fit3D fine-tuning (30k/60k/90k steps). Joint metrics over 19 test sequences; mesh and angle metrics over the eight sequences with ground-truth SMPL. All differences are within noise—fine-tuning does not regress out-of-domain performance. Lower is better. MPVPE-T on AIST++ measures only the deviation of the predicted shape from the neutral body, because the AIST++ ground truth uses β = 0 ; it is therefore not comparable with the Fit3D MPVPE-T column (see text).
Table 6. AIST++ (out-of-domain street dance): generalization after Fit3D fine-tuning (30k/60k/90k steps). Joint metrics over 19 test sequences; mesh and angle metrics over the eight sequences with ground-truth SMPL. All differences are within noise—fine-tuning does not regress out-of-domain performance. Lower is better. MPVPE-T on AIST++ measures only the deviation of the predicted shape from the neutral body, because the AIST++ ground truth uses β = 0 ; it is therefore not comparable with the Fit3D MPVPE-T column (see text).
Transl.
(mm)
MPJPE
(mm)
MPJPE-PA
(mm)
MPVPE
(mm)
MPVPE-PA
(mm)
MPVPE-T
(mm)
MPJAE
(°)
PA-MPJAE
(°)
Released (HMR2.0) 501.2 164.1 108.9 109.0 69.2 22.6 23.0 21.6
   + FT (30k) 516.0 165.1 107.4 110.2 68.3 24.0 22.5 21.0
   + FT (60k) 511.0 163.6 107.7 110.8 69.1 20.9 22.6 21.1
   + FT (90k) 496.3 164.0 107.9 110.7 68.8 21.9 22.7 21.2
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Fieraru, M. Hitting the Gym with Fit3D: Benchmarking and Improving Monocular 3D Human Reconstruction on Extreme Fitness Motions. J. Imaging 2026, 12, 311. https://doi.org/10.3390/jimaging12070311

AMA Style

Fieraru M. Hitting the Gym with Fit3D: Benchmarking and Improving Monocular 3D Human Reconstruction on Extreme Fitness Motions. Journal of Imaging. 2026; 12(7):311. https://doi.org/10.3390/jimaging12070311

Chicago/Turabian Style

Fieraru, Mihai. 2026. "Hitting the Gym with Fit3D: Benchmarking and Improving Monocular 3D Human Reconstruction on Extreme Fitness Motions" Journal of Imaging 12, no. 7: 311. https://doi.org/10.3390/jimaging12070311

APA Style

Fieraru, M. (2026). Hitting the Gym with Fit3D: Benchmarking and Improving Monocular 3D Human Reconstruction on Extreme Fitness Motions. Journal of Imaging, 12(7), 311. https://doi.org/10.3390/jimaging12070311

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop