Build a usable digital twin
Combine image-based asset generation with metric scene alignment and arm–hand calibration.
A REAL-TO-SIM-TO-REAL FRAMEWORK
From a single image.
To real robot behavior.
From a Single Image to Deployable
Dexterous Manipulation via Real-to-Sim-to-Real
1 The University of Texas at Austin2 University of Florida3 BrainCo
DEXTERA reconstructs a workspace from one image, enabling policy learning in simulation and deployment on real robots.
REAL TO SIM · INTERACTIVE
Choose a layout and background, compare the real reference, and look around in 3D.
Three captured object layouts, paired with their real reference images. The 3D view combines textured object meshes and independently selectable reconstructed backgrounds.
OVERVIEW
Reconstruction, calibration, experience, and deployment in one framework.
Dexterous robot learning needs large amounts of experience, but collecting demonstrations on hardware and manually building matching simulators are expensive. DEXTERA connects these steps in a unified workflow.
A single RGB image provides the visual input for a static Gaussian background and interactive rigid or articulated objects. Calibration aligns the reconstructed scene with known robot geometry. Task primitives, VR demonstrations, and object-centric trajectory synthesis then support imitation and reinforcement learning through a shared policy interface.
Experiments examine visual reconstruction, physical trajectory replay, and learned-policy deployment across two robot platforms. Simulation-only policies transfer to hardware, while adding limited real demonstrations improves physical success.
Combine image-based asset generation with metric scene alignment and arm–hand calibration.
Turn a small set of VR demonstrations into trajectories for randomized object configurations.
Evaluate six policy architectures through a consistent observation and action interface.
THE FRAMEWORK
Generate the scene, align it, define the task, and learn a deployable policy.

01 / RECONSTRUCT
Factor the input image into a static background and individual interactive objects. Complete hidden geometry, generate textured meshes, and assign collision geometry and physical defaults.
OUTPUTVisual assets and simulation geometry.
Link geometry, joint axes, and motion limits.
02 / CALIBRATE
Use a calibration capture to recover metric scale and align scene, object, camera, arm, and hand geometry in a common coordinate system.
OUTPUTA calibrated robot workspace.
03 / CONSTRUCT
Task primitives define rewards, success tolerances, termination, and goal conditions. Randomized initial configurations provide variation for demonstrations and policy training.
OUTPUTTraining tasks and evaluation conditions.
04 / LEARN & DEPLOY
A common multimodal task interface supports imitation learning and reinforcement learning. Deployment preserves the policy's observation representations and action space.
OUTPUTPolicies evaluated on physical robots.
VR demonstrations and synthesized trajectories, with simulation–real co-training.
A privileged simulation teacher supervises a student with deployable observations.
The learned policy produces actions for the physical robot.
Input image for asset generation; calibration observations and the known robot model for metric registration.
GENERATING EXPERIENCE
Adapt demonstrated motion to new object configurations using task-relative poses.
Collect source motions for contact-rich manipulation in the reconstructed workspace.
Object-centric synthesis adapts end-effector motion to new task frames, interpolates waypoints, and adds action noise. Collision checks and task success criteria filter invalid simulated trajectories.
Across 13 task–embodiment pairs. The physical trajectories are collected or replayed on hardware.
EXPERIMENTS
Three evaluations connect visual quality to physically executable motion and learned robot behavior.
01 / RECONSTRUCTION
18 RGB-D captures across 10 tabletop scenes.
| Method | Mask IoU ↑ | RGB MAE ↓ | Centroid error (mm) ↓ | F-score @ 1 cm ↑ |
|---|---|---|---|---|
| DEXTERA | 0.731 ± 0.136 | 0.182 ± 0.062 | 11.5 ± 16.0 | 0.906 ± 0.132 |
| Hunyuan3D-2.1 | 0.661 ± 0.180 | 0.223 ± 0.063 | 14.3 ± 17.3 | 0.750 ± 0.250 |
| TRELLIS.2 | 0.559 ± 0.196 | 0.250 ± 0.063 | 25.8 ± 14.4 | 0.346 ± 0.202 |
| SAM3D + FP | 0.604 ± 0.143 | 0.407 ± 0.118 | 18.3 ± 17.3 | 0.622 ± 0.235 |
| SAM3D only | 0.488 ± 0.285 | 0.391 ± 0.111 | 29.9 ± 35.8 | 0.538 ± 0.339 |
RGB values are normalized to [0, 1]. Centroid errors are converted from meters to millimeters. DEXTERA has the best mean values for these selected metrics.
02 / PHYSICAL CONSISTENCY
343 successful physical replays / 390 trials
For each task pair, 30 trajectories that succeed in simulation are executed on the real robot. This measures the physical consistency of the retained trajectories.
Learned-policy success is measured in separate experiments.
03 / SIMULATION–REAL CO-TRAINING
Mean physical success across three policies on six selected task pairs.
+32.7 percentage points in mean physical success.
20 physical trials per condition. Co-training improves 17 of 18 policy–task comparisons, with one tie.
Sim only 35.0%
Sim + Real 62.5%
Sim only 15.8%
Sim + Real 55.8%
Sim only 36.7%
Sim + Real 67.5%
Success rate · all bars use a 0–100% scale
FULL POLICY BENCHMARK
Mean success across all 13 task–embodiment pairs, weighted equally. Imitation policies use simulation and real training data; PPO uses teacher–student distillation.
Physical evaluation: 20 trials per policy–task pair. Values from Table V.
| Task / robot platform | ACT | DP | BC-RNN | OpenVLA | π0.5 | PPO T–S |
|---|---|---|---|---|---|---|
| OpenArmLift toy bus | 55% | 50% | 0% | 5% | 65% | 20% |
| OpenArmLift blue mug | 100% | 90% | 5% | 10% | 100% | 10% |
| OpenArmPlace toy bus in basket | 15% | 0% | 0% | 10% | 10% | 0% |
| OpenArmHand over gray bowl | 65% | 0% | 0% | 0% | 40% | 10% |
| OpenArmLift basket with both hands | 80% | 45% | 5% | 0% | 90% | 30% |
| OpenArmReorient thermometer | 55% | 15% | 0% | 0% | 50% | 0% |
| OpenArmLift box with both hands | 35% | 50% | 0% | 5% | 60% | 0% |
| OpenArmPush blue cube | 50% | 35% | 10% | 20% | 55% | 10% |
| OpenArmPlace box in basket with both hands | 10% | 5% | 0% | 0% | 20% | 0% |
| OpenArmClose laptop | 80% | 100% | 45% | 45% | 100% | 30% |
| KUKAPlace coconut-water carton in basket | 25% | 0% | 0% | 0% | 5% | 15% |
| KUKALift mustard bottle | 35% | 20% | 0% | 20% | 40% | 25% |
| KUKAClose laptop | 100% | 75% | 80% | 65% | 100% | 75% |
| Mean · all 13 pairs | 54.2% | 37.3% | 11.2% | 13.8% | 56.5% | 17.3% |
Same action spaces and success criteria across methods. Simulation uses 100 trials per pair; hardware uses 20. The platform selector filters the table; the chart above always covers all 13 pairs.
ON THE ROBOT
Physical ACT policy rollouts. Both demonstration panels use 3× playback.
Laptop closing, mug lifting, thermometer reorientation, and toy-bus placement.
Laptop closing, mustard-bottle lifting, and coconut-water placement.
PROJECT VIDEO
2 min 59 sec · English narration
Captions available
@misc{wu2026dextera,
title = {DEXTERA: From a Single Image to Deployable Dexterous Manipulation via Real-to-Sim-to-Real},
author = {Jin Wu and Lianjie Yuan and Zeyan Sun and Yuanyuan Lei and Disi A and Bicheng Han and Fangzhou Xia},
year = {2026},
eprint = {2609.21045},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2609.21045}
}