Hierarchical Whole-Body Control for Tendon-Cable-Driven Humanoids via Reference-Residual Policy and Offline-Learned Tendon Mapping
Abstract
1. Introduction
- A transmission-aware hierarchical control architecture. A simulation-trained reference-residual policy and an independently trained state-conditioned mapper are connected in series at deployment. The mapper converts desired joint positions and the measured plant state into motor-position commands, while the original fixed static calibration is retained as the Mapping-OFF baseline.
- A globally anchored, multi-source, single-stage training pipeline. Heterogeneous motions are retargeted to a common robot-space representation. Global-anchor rewards, motion- and segment-level adaptive sampling, and tendon-oriented domain randomization are integrated into one PPO stage that directly produces the deployable actor.
- Randomized repeated physical validation. Mapping OFF and ON are evaluated in randomized repeated trials on two nominally identical Droid X3 units. Within each robot–motion block, the frozen PPO checkpoint, reference motion, controller settings, safety bounds, and frozen mapper weights are held fixed; complete trials are used as the statistical units. Walk, squat, and dance provide multiple motion conditions, while the two separately assembled cable transmissions provide a cross-unit hardware check. The aggregate complete-trial action-completion rates are 68% for Mapping OFF and 79% for Mapping ON.
2. Related Work
2.1. Humanoid Motion Tracking
2.2. Multi-Source Motion Data and Global Consistency
2.3. Sim-to-Real Transfer via Domain Randomization
2.4. Tendon-Cable Actuation
3. Platform and Problem Formulation
3.1. The Droid X3 Platform
3.2. System Overview
4. Method
4.1. Reference Motion Representation
4.2. Observations, Policy, and Residual Action
4.3. Tendon Mapping Network
4.4. Reward Design and Termination
4.5. Adaptive Sampling and Domain Randomization
4.6. Single-Stage Training and Deployment
5. Experiments
5.1. Experimental Setup
5.2. Complex Whole-Body Motion Tracking
5.3. Tendon Mapping Validation and Transmission Characterization
5.4. Policy-Side Design Scope
5.5. Single-Stage Training Diagnostics and Adaptive Sampling
6. Conclusions and Discussion
Scope and Limitations
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Peng, X.B.; Abbeel, P.; Levine, S.; Van de Panne, M. DeepMimic: Example-guided deep reinforcement learning of physics-based character skills. ACM Trans. Graph. (TOG) 2018, 37, 143. [Google Scholar]
- Luo, Z.; Cao, J.; Merel, J.; Winkler, A.; Huang, J.; Kitani, K.; Xu, W. Universal humanoid motion representations for physics-based control. In Proceedings of the International Conference on Learning Representations (ICLR), Vienna, Austria, 7–11 May 2024. [Google Scholar]
- He, T.; Xiao, W.; Lin, T.; Luo, Z.; Xu, Z.; Jiang, Z.; Kautz, J.; Liu, C.; Shi, G.; Wang, X.; et al. HOVER: Versatile neural whole-body controller for humanoid robots. In Proceedings of the 2025 IEEE International Conference on Robotics and Automation (ICRA), Atlanta, GA, USA, 19–23 May 2025; pp. 9989–9996. [Google Scholar]
- Chen, Z.; Ji, M.; Cheng, X.; Peng, X.; Peng, X.B.; Wang, X. GMT: General motion tracking for humanoid whole-body control. arXiv 2025, arXiv:2506.14770. [Google Scholar]
- Yang, C.; Sun, Y.; Ye, P.; Chen, X.; Yu, C.; Chen, T. Efficiently learning general motion tracking policy for high dynamic humanoid whole-body control. arXiv 2025, arXiv:2512.19043. [Google Scholar]
- Luo, Z.; Yuan, Y.; Wang, T.; Li, C.; Chen, S.; Castañeda, F.; Cao, Z.A.; Li, J.; Minor, D.; Ben, Q.; et al. SONIC: Supersizing motion tracking for natural humanoid whole-body control. arXiv 2025, arXiv:2511.07820. [Google Scholar]
- Yao, Y.; Luo, C.; Du, J.; He, W.; Lu, J.G. GBC: Generalized behavior-cloning framework for whole-body humanoid imitation. arXiv 2025, arXiv:2508.09960. [Google Scholar]
- Wang, Y.; Yang, M.; Zeng, W.; Zhang, Y.; Xu, X.; Jiang, H.; Ding, Z.; Lu, Z. From experts to a generalist: Toward general whole-body control for humanoid robots. arXiv 2025, arXiv:2506.12779. [Google Scholar]
- He, T.; Luo, Z.; He, X.; Xiao, W.; Zhang, C.; Zhang, W.; Kitani, K.; Liu, C.; Shi, G. OmniH2O: Universal and dexterous human-to-humanoid whole-body teleoperation and learning. arXiv 2024, arXiv:2406.08858. [Google Scholar]
- Ze, Y.; Chen, Z.; Araújo, J.P.; Cao, Z.A.; Peng, X.B.; Wu, J.; Liu, C.K. TWIST: Teleoperated whole-body imitation system. arXiv 2025, arXiv:2505.02833. [Google Scholar]
- Li, Y.; Lin, Y.; Cui, J.; Liu, T.; Liang, W.; Zhu, Y.; Huang, S. CLONE: Closed-loop whole-body humanoid teleoperation for long-horizon tasks. arXiv 2025, arXiv:2506.08931. [Google Scholar]
- Truong, T.E.; Liao, Q.; Huang, X.; Tevet, G.; Liu, C.K.; Sreenath, K. BeyondMimic: From motion tracking to versatile humanoid control via guided diffusion. arXiv 2025, arXiv:2508.08241. [Google Scholar]
- Sun, Z.; Huang, B.S.; Peng, Y.; Li, X.; Ma, J.; Sun, Y.; Li, Z.; Jiang, H.; Gao, B.; Bing, Z.; et al. MOSAIC: Bridging the sim-to-real gap in generalist humanoid motion tracking and teleoperation with rapid residual adaptation. arXiv 2026, arXiv:2602.08594. [Google Scholar]
- Zhao, S.; Ze, Y.; Wang, Y.; Liu, C.K.; Abbeel, P.; Shi, G.; Duan, R. ResMimic: From general motion tracking to humanoid whole-body loco-manipulation via residual learning. arXiv 2025, arXiv:2510.05070. [Google Scholar]
- Araújo, J.P.; Ze, Y.; Xu, P.; Wu, J.; Liu, C.K. Retargeting matters: General motion retargeting for humanoid motion tracking. arXiv 2025, arXiv:2510.02252. [Google Scholar]
- Pan, Y.; Qiao, R.; Chen, L.; Chitta, K.; Pan, L.; Mai, H.; Bu, Q.; Zhao, H.; Zheng, C.; Luo, P.; et al. Agility meets stability: Versatile humanoid control with heterogeneous data. arXiv 2025, arXiv:2511.17373. [Google Scholar]
- Xue, Y.; Lin, Y.; Dong, W.; Tang, Y.; Wang, J.; Pang, J.; Zhou, M.; Liu, M.; Zhang, W. Scalable and general whole-body control for cross-humanoid locomotion. arXiv 2026, arXiv:2602.05791. [Google Scholar]
- Zhang, Z.; Wen, K.; Xu, M.; He, J.; Li, C.; Miki, T.; Schwarke, C.; Zhang, C.; Peng, X.B.; Hutter, M. Learning whole-body humanoid locomotion via motion generation and motion tracking. arXiv 2026, arXiv:2604.17335. [Google Scholar]
- Li, J.; Tang, B.; Wu, F. TeleGate: Whole-body humanoid teleoperation via gated expert selection with motion prior. arXiv 2026, arXiv:2602.09628. [Google Scholar]
- Zhu, T.; Cai, G.; Yang, Z.; Ren, G.; Xie, H.; Wang, Z.; Wu, J.; Wang, J.; Yang, X.; Mu, Y.; et al. CLOT: Closed-loop global motion tracking for whole-body humanoid teleoperation. arXiv 2026, arXiv:2602.15060. [Google Scholar]
- Li, Y.; Ma, L.; Lin, Y.; Du, Y.; Liu, M.; Hu, K.; Cui, J.; Zhu, Y.; Liang, W.; Jia, B.; et al. OmniClone: Engineering a robust, all-rounder whole-body humanoid teleoperation system. arXiv 2026, arXiv:2603.14327. [Google Scholar]
- Wang, Y.; Zhao, Q.; Lau, Y.F.; Yu, R.; Tsui, H.W.; Chen, Q.; Wang, J.; Pang, J.; Tan, P. HumanX: Toward agile and generalizable humanoid interaction skills from human videos. arXiv 2026, arXiv:2602.02473. [Google Scholar]
- Zhang, Z.; Lu, H.; Lian, Y.; Chen, Z.; Liu, Y.; Lin, C.; Xue, H.; Zeng, Z.; Qi, Z.; Zheng, S.; et al. Learning athletic humanoid tennis skills from imperfect human motion data. arXiv 2026, arXiv:2603.12686. [Google Scholar]
- Shanghai DroidUp Co., Ltd. Bipedal Humanoid Robot Product Series. Available online: https://www.droidup.com/product (accessed on 27 July 2026).
- Peng, X.B.; Ma, Z.; Abbeel, P.; Levine, S.; Kanazawa, A. AMP: Adversarial motion priors for stylized physics-based character control. ACM Trans. Graph. (TOG) 2021, 40, 144. [Google Scholar]
- Lu, Q.; Feng, Y.; Shi, B.; Piseno, M.; Bao, Z.; Liu, C.K. GentleHumanoid: Learning upper-body compliance for contact-rich human and object interaction. arXiv 2025, arXiv:2511.04679. [Google Scholar]
- Xiong, Z.; Fang, L.; Huang, J.; Yamazaki, K.; Zhang, H.; Gan, C. ExtremControl: Low-latency humanoid teleoperation with direct extremity control. arXiv 2026, arXiv:2602.11321. [Google Scholar]
- Lee, K.; Park, S.; Park, G.; Kim, M.J.; Park, J. Safety-critical whole-body control for humanoid robots via input-to-state safe control barrier functions. arXiv 2026, arXiv:2605.25546. [Google Scholar]
- Mao, Z.; Wang, J.; Zhang, J.; Ohgi, J.; Zheng, Y.; Peng, Y.; Zhao, L.; Su, Q.; Huang, W.; Xu, B. Fine-tuned multimodal large language model for autonomous state cognition system of shape-recognition 6-bar tensegrity integrated with flexible sensors. Microsyst. Nanoeng. 2026, 12, 228. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Todorov, E.; Erez, T.; Tassa, Y. MuJoCo: A physics engine for model-based control. In Proceedings of the 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Vilamoura, Portugal, 7–12 October 2012; pp. 5026–5033. [Google Scholar]
- Mahmood, N.; Ghorbani, N.; Troje, N.F.; Pons-Moll, G.; Black, M.J. AMASS: Archive of motion capture as surface shapes. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 5442–5451. [Google Scholar]
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal policy optimization algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar]
- Schulman, J.; Moritz, P.; Levine, S.; Jordan, M.; Abbeel, P. High-dimensional continuous control using generalized advantage estimation. In Proceedings of the International Conference on Learning Representations (ICLR), San Juan, PR, USA, 2–4 May 2016. [Google Scholar]






| Term | Definition | Weight |
|---|---|---|
| Global-anchor position | 0.5 | |
| Global-anchor orientation | 0.5 | |
| Global-anchor linear velocity | 1.0 | |
| Key-body position | 1.0 | |
| Key-body orientation | 1.0 | |
| Key-body linear velocity | 1.0 | |
| Key-body angular velocity | 1.0 | |
| Feet position | 0.8 | |
| Base orientation | 0.5 | |
| Base linear velocity | 0.5 | |
| Base angular velocity | 0.5 | |
| Action smoothness | −0.3 | |
| Joint limit | −10 | |
| Self-collision | −10 |
| Layer | Parameter | Range/Unit |
|---|---|---|
| Reference input | Root position () | ±0.05/±0.05/±0.01 m |
| Root orientation (r/p/y) | ±0.1/±0.1/±0.2 rad | |
| Joint reference offset | ±0.1 rad | |
| Robot body | Torso CoM offset () | ±0.025/±0.05/±0.05 m |
| Encoder bias | ±0.01 rad | |
| Contact | Foot friction coefficient | |
| Push velocity (every 1–3 s) | ±0.5 m/s, ±0.78 rad/s | |
| Observation noise | Projected gravity | ±0.05 |
| Base angular velocity | ±0.2 rad/s | |
| Joint position | ±0.01 rad | |
| Joint velocity | ±0.5 rad/s | |
| Tendon execution | Response delay, | 0–2 steps (0–40 ms) |
| Equiv. stiffness/damping, | ×[0.8, 1.2] | |
| Mapping residual, | rad |
| Parameter | Value | Parameter | Value |
|---|---|---|---|
| Parallel environments | 8192 | Discount factor, | 0.99 |
| Rollout length | 24 | GAE factor, | 0.95 |
| Learning epochs | 5 | Clip range, | 0.2 |
| Mini-batches | 4 | Value loss coef., | 1.0 |
| Learning rate | (adaptive) | Entropy coef., | 0.005 |
| Desired KL | 0.01 | Max gradient norm | 1.0 |
| Hidden layers | Activation | ELU | |
| Initial action std | 1.0 | Control frequency | 50 Hz |
| Source Group | Sequences | Frames | Duration (h) | Windows |
|---|---|---|---|---|
| LAFAN1 | 57 | 608,818 | 3.382 | 1386 |
| Pico-record | 9 | 93,951 | 0.522 | 213 |
| SEED | 63,235 | 22,945,495 | 127.475 | 77,623 |
| TWIST2 | 31,058 | 15,833,851 | 87.966 | 47,869 |
| Total | 94,359 | 39,482,115 | 219.345 | 127,091 |
| Unit | Motion | Condition | Frames | Duration (s) | Joint MAE (Rad/Joint) | Torso Error (Deg) |
|---|---|---|---|---|---|---|
| 1 | Walk | OFF | 768 | 15.36 | 0.1296 | 5.369 |
| 1 | Walk | ON | 790 | 15.80 | 0.1139 | 5.197 |
| 1 | Squat | OFF | 833 | 16.66 | 0.1300 | 6.657 |
| 1 | Squat | ON | 738 | 14.76 | 0.1109 | 4.261 |
| 1 | Dance | OFF | 1041 | 20.82 | 0.1327 | 6.510 |
| 1 | Dance | ON | 1077 | 21.54 | 0.1089 | 5.000 |
| 2 | Squat | ON | 833 | 16.66 | 0.1088 | 4.538 |
| 2 | Dance | ON | 1135 | 22.70 | 0.1109 | 3.795 |
| Record | Iter. ↓ | Reward Score (%) ↑ | Episode-Length Proxy (%) ↑ | Entropy |
|---|---|---|---|---|
| Archive A (run 0609) | 10,551 | 52.3 | 31.2 | 0.743 |
| Archive B (run 0610) | 1412 | 70.9 | 61.8 | 0.709 |
| Archive C (run 0608) | 1604 | 78.1 | 72.3 | 0.665 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Gan, W.; Chen, J.; Li, Q.; Mou, H.; Zhang, J. Hierarchical Whole-Body Control for Tendon-Cable-Driven Humanoids via Reference-Residual Policy and Offline-Learned Tendon Mapping. Biomimetics 2026, 11, 607. https://doi.org/10.3390/biomimetics11090607
Gan W, Chen J, Li Q, Mou H, Zhang J. Hierarchical Whole-Body Control for Tendon-Cable-Driven Humanoids via Reference-Residual Policy and Offline-Learned Tendon Mapping. Biomimetics. 2026; 11(9):607. https://doi.org/10.3390/biomimetics11090607
Chicago/Turabian StyleGan, Wencong, Jiehui Chen, Qingdu Li, Haiming Mou, and Jianwei Zhang. 2026. "Hierarchical Whole-Body Control for Tendon-Cable-Driven Humanoids via Reference-Residual Policy and Offline-Learned Tendon Mapping" Biomimetics 11, no. 9: 607. https://doi.org/10.3390/biomimetics11090607
APA StyleGan, W., Chen, J., Li, Q., Mou, H., & Zhang, J. (2026). Hierarchical Whole-Body Control for Tendon-Cable-Driven Humanoids via Reference-Residual Policy and Offline-Learned Tendon Mapping. Biomimetics, 11(9), 607. https://doi.org/10.3390/biomimetics11090607
