Learning Dexterous Humanoid Loco-Manipulation
from Human Demonstrations

1Department of Artificial Intelligence, School of Engineering, Westlake University

2School of Artificial Intelligence, Shanghai Jiao Tong University

*Equal contribution†Corresponding author

Unitree G1 humanoids performing diverse whole-body motions and dexterous interactions with everyday objects.

TL;DR DexWeave combines interaction-consistent motion retargeting with anatomy-aware reinforcement learning to enable dexterous humanoid loco-manipulation from human demonstrations.

Abstract

Learning dexterous humanoid loco-manipulation from human demonstrations requires transferring not only human motion, but also the coordinated interaction structure underlying the demonstrated behavior. This is challenging because embodiment differences distort the coupling among body motion, wrist placement, finger articulation, and object interaction, while kinematically accurate references may still be difficult to realize under robot dynamics. We present DexWeave, a unified framework that connects interaction-consistent motion retargeting with anatomy-aware whole-body policy learning. DexWeave first employs a two-stage retargeting procedure that initializes body and hand motions with specialized solvers and subsequently performs coupled refinement over the upper-body interaction chain while preserving lower-body support. The resulting references are tracked by an anatomy-aware Transformer policy that represents anatomical regions as structured tokens and uses directed masked attention to model their dependencies, with object information selectively conditioning the upper-body pathway for dexterous interaction. The policy jointly outputs body and dexterous-hand actions and is trained directly with reinforcement learning, without pretrained tracking policies, teacher–student distillation, or subsequent residual refinement. Across multiple dexterous loco-manipulation motions, DexWeave improves retargeting fidelity and interaction consistency while achieving higher manipulation performance and faster policy convergence than MLP baselines. We further deploy the learned policies on a physical Unitree G1 humanoid equipped with Inspire dexterous hands, demonstrating dexterous whole-body loco-manipulation in the real world.

Method overview

Figure 2: Human-object demonstrations flow through interaction-consistent motion retargeting and an anatomy-aware policy to produce dexterous whole-body motion.

Given a human body–hand–object demonstration, DexWeave first constructs an interaction-consistent robot reference through specialized body–hand initialization and coupled refinement of the upper-body interaction chain. An anatomy-aware whole-body loco-manipulation policy then realizes the reference under robot dynamics using regional anatomical tokens, directed masked attention, and selective object conditioning, producing executable dexterous humanoid loco-manipulation skills.

Real-world experiments

Loco-Manipulation

Locomotion

Retargeting results

Video results

Loco-Manipulation from HUMOTO Dataset

Locomotion from LAFAN1 Dataset

Quantitative results

Body-only motion retargeting
DatasetMethodPenetrationFoot skatingContact preservation
Duration ↓Max. depth
(cm) ↓
Duration ↓Max. velocity
(m/s) ↓
Duration ↑Contact distance
(cm) ↓
LAFAN1OmniRetarget0.1252.9340.0910.416N/AN/A
GMR0.1834.3140.0100.548N/AN/A
SOMA0.9686.3790.1150.565N/AN/A
DexWeave (Ours)≈02.57500N/AN/A
OMOMOOmniRetarget0.1312.735≈01.4550.64410.289
GMR0.9163.1170.0080.8020.8797.677
SOMA0.7464.1630.0340.4670.57716.205
DexWeave (Ours)0.0022.117≈00.3550.9992.944

Bold marks the best value and underlining the second-best within each dataset. Durations are frame fractions; ≈0 denotes a near-zero value.

Dexterous whole-body retargeting
DatasetMethodPenetrationHand alignment
Duration ↓Max. depth
(cm) ↓
Primary error
(mm) ↓
Secondary error
(mm) ↓
Palm error
(°) ↓
GRABOmniRetarget + DexPilot + IK0.2122.26916.58513.36811.555
OmniRetarget + SBR + IK0.1902.36214.92514.6599.530
DexWeave (Ours)0.0011.2634.34214.4763.652
HUMOTOOmniRetarget + DexPilot + IK0.5223.59732.38931.01814.444
OmniRetarget + SBR + IK0.5113.58229.99928.43612.476
DexWeave (Ours)0.0072.8547.19811.7024.439

Bold marks the best value and underlining the second-best within each dataset. Primary: thumb and index fingertips. Secondary: remaining fingertips. Palm: palm-normal angular error.

Sim-to-Sim results

Deploy an IsaacLab-trained policy in MuJoCo

Loco-Manipulation from HUMOTO Dataset

Locomotion from LAFAN1 Dataset

Body-only tracking performance
MethodCompletion ratio
(%) ↑
Body-position error
(cm) ↓
Anchor-position error
(cm) ↓
Anchor-rotation error
(deg) ↓
Any2Track72.826.0944.8832.07
GMT69.865.8536.0621.92
BeyondMimic1002.825.192.48
DexWeave (Ours)1002.624.822.34

MuJoCo sim-to-sim evaluation, averaged equally over five tracking motions. DexWeave trains a separate policy for each reference sequence; Any2Track and GMT are included as general motion-tracking baselines. Completion ratio is the percentage of evaluation episodes reaching the reference end or time limit. Best values are in bold.

Dexterous loco-manipulation performance
MethodSuccess rate
(%) ↑
Body-position error
(cm) ↓
Object-position error
(cm) ↓
Object-rotation error
(deg) ↓
InterMimic57.746.513.587.71
Object MLP85.005.693.846.16
DexWeave (Ours)97.54.933.436.83

MuJoCo sim-to-sim evaluation, averaged equally over five loco-manipulation references. InterMimic is trained and evaluated using the SMPL model. Best values are in bold.

Citation

@article{sun2026dexweave,
  author  = {Sun, Naichuan and Shen, Haotian and Zhang, Yizhang and Feng, Luying and Wang, Haoze and Xiangli, Yuanbo and Jin, Yaochu and Liu, Peidong},
  title   = {DexWeave: Learning Dexterous Humanoid Loco-Manipulation from Human Demonstrations},
  journal = {arXiv},
  year    = {2026},
}