Haoran (Cris) Wang
<- All projects

Humanoid Teleoperation and Motion Retargeting

Role -
Motion Control Algorithm Engineering Intern, UBTECH
Date -
January 1, 2026
Stack —
Humanoid Robotics, Teleoperation, GMR, Redis, ROS 2

A modular pipeline unifying PICO, Xsens, and camera motion inputs for humanoid retargeting, replay, and robot deployment.

Overview

MULTITELEOP is a modular teleoperation system that connects heterogeneous human-motion devices to one humanoid trajectory interface. I built the integration around three acquisition paths: PICO VR tracking, Xsens MVN inertial motion capture, and an experimental RGB-camera path based on PromptHMR.

The same downstream pipeline handles live inputs and recorded-session replay. This made it possible to debug device adapters without a human operator, compare retargeting changes on repeatable motion, and generate robot trajectories for imitation-learning experiments.

My contribution

  • Extended the open-source General Motion Retargeting pipeline to generate large-scale robot motion trajectories for imitation learning.
  • Built a unified receiver, Redis transport, retargeting, replay, and ROS 2 publishing framework on top of TWIST2.
  • Integrated PICO and Xsens as real-time inputs, including timestamp-aware JSONL recording and offline replay.
  • Added Walker S2 joint mapping, optional ground-height correction, and shoulder-yaw offsets for robot-specific post-processing.
  • Prototyped an RGB/OAK camera input that converts PromptHMR SMPL-X predictions into the same retargeting interface.
  • Supported simulation-to-real validation on UBTECH’s Walker S2 humanoid robot.

Pipeline

01 / Capture PICO, Xsens, Camera Live devices or recorded replay
02 / Normalize Unified Human Pose Joint positions and quaternions in Redis
03 / Retarget GMR + Post-process Robot mapping, offsets, and smoothing
04 / Execute ROS 2 / Simulation / Robot Walker S2, G1, and dataset recording

System structure

Multi-source acquisition

Each device has different coordinates, skeleton definitions, update rates, and transport protocols. Source-specific receivers convert PICO tracking frames, Xsens UDP datagrams, or PromptHMR SMPL-X output into a shared representation of named joints with position and quaternion data.

PICO 4 Ultra input running through the unified teleoperation and retargeting pipeline.

Shared pose bus

Redis separates acquisition from retargeting. Device receivers publish source-specific keys, while translators consume them independently. Offline senders replay recorded JSONL with original timing or a controlled playback rate, so the same processing path can be tested without reconnecting hardware.

Robot-space retargeting

GMR maps the normalized human pose to the selected humanoid model. Post-processing applies optional ground alignment, Walker S2 shoulder corrections, hand/controller states, and temporal smoothing before publishing a consistent robot mimic observation.

Xsens MVN motion retargeting evaluated in simulation before physical deployment.

Simulation and deployment

A ROS 2 bridge converts the Redis output into the joint-state layout expected by downstream sim-to-sim or sim-to-real controllers. Recording utilities preserve processed trajectories and episode data for imitation-learning workflows. Physical-robot tests follow simulation checks and conservative safety limits.

Xsens-driven motion validated on the physical humanoid platform after simulation checks.

Implementation status

Input path Live stream Offline replay Status
PICO 4 Ultra XRoboToolkit JSONL Integrated and used in teleoperation experiments
Xsens MVN UDP, 9763 JSONL Integrated and used in teleoperation experiments
RGB / OAK camera PromptHMR Video Experimental; not yet validated end to end

Engineering decisions

  • One interface, multiple sensors. Normalization keeps device logic out of robot-specific translators.
  • Replay as a first-class input. Timestamp-aware playback turns hardware sessions into repeatable regression tests.
  • Robot-specific corrections stay optional. Walker S2 offsets are explicit flags rather than hidden assumptions in the common pipeline.
  • Experimental paths are isolated. Camera inference can evolve independently without destabilizing the PICO and Xsens workflows.

Open-source scope

The public repository includes the receiver, translator, replay, ROS 2 bridge, deployment utilities, architecture documentation, and short rendered demonstrations. Raw motion-capture sessions, model weights, vendor SDKs, SMPL-X assets, and internal deployment notes are excluded. The project builds on TWIST2 and GMR, with upstream authorship and licensing retained in the repository.