Fitness motions present some of the most challenging cases for monocular 3D human reconstruction: extreme articulations, heavy self-occlusion, and frequent self-contact. The
Fit3D dataset captures these motions at large scale via a 12-camera VICON motion capture system synchronized with 4 RGB cameras, but
[...] Read more.
Fitness motions present some of the most challenging cases for monocular 3D human reconstruction: extreme articulations, heavy self-occlusion, and frequent self-contact. The
Fit3D dataset captures these motions at large scale via a 12-camera VICON motion capture system synchronized with 4 RGB cameras, but in its original release provides only 3D skeletons. This paper contributes three studies built on top of
Fit3D. First, we design and validate a methodology for constructing dense GHUM and SMPL-X pseudo-ground-truth shape and pose annotations on top of the raw MoCap: an optimization-based fitting pipeline that combines markers, multi-view 2D keypoints, separate body and hand normalizing-flow priors, and a self-collision loss, which we show improves on the marker-only MoSh++ baseline on the hands and extremities. Second, we define a standardized evaluation protocol—metric set, frame sampling, and coordinate conventions—for monocular 3D human reconstruction on
Fit3D, served through the IMAR-hosted Fit3D evaluation resource, and use it to conduct a comparative benchmark study of 19 representative methods spanning optimization-based, single-frame, and video-based families; the two trained on
Fit3D (NLF and SMPLest-X) lead the position and orientation metrics, respectively. Third, a controlled fine-tuning experiment shows that adding
Fit3D to the training mixture of a strong baseline (HMR2.0) sharply lowers error on the hardest fitness poses without degrading out-of-domain generalization. The
Fit3D dataset and the GHUM/SMPL-X annotations are available, under a non-commercial research license, through the IMAR Fit3D resource; as of June 2026, 1088 academics have registered for access.
Full article