DEXTERA

A REAL-TO-SIM-TO-REAL FRAMEWORK

DEXTERA

From a single image.
To real robot behavior.

REAL SIM REAL

Scroll to explore
Nine tasks · top viewTask demonstrations
REALSIMREAL

DEXTERA

From a Single Image to Deployable
Dexterous Manipulation via Real-to-Sim-to-Real

Jin Wu1, Lianjie Yuan1,†, Zeyan Sun1,†, Yuanyuan Lei2, Disi A3, Bicheng Han3, Fangzhou Xia1,*

1 The University of Texas at Austin2 University of Florida3 BrainCo

† Lianjie Yuan and Zeyan Sun contributed equally.* Corresponding author.

DEXTERA reconstructs a workspace from one image, enabling policy learning in simulation and deployment on real robots.

REAL TO SIM · INTERACTIVE

Explore the reconstructed workspace.

Choose a layout and background, compare the real reference, and look around in 3D.

Saved preview
Saved reconstruction preview for Layout 1

Step inside the reconstruction.

Load once, then switch layouts and backgrounds.

Saved image · Load 3D to apply your background.Drag to orbit · Right-drag to pan · Scroll to zoom

Three captured object layouts, paired with their real reference images. The 3D view combines textured object meshes and independently selectable reconstructed backgrounds.

13
task–embodiment pairs
2
dexterous robot platforms
6
policy architectures
86
object instances evaluated

OVERVIEW

From a workspace image
to real robot behavior.

Reconstruction, calibration, experience, and deployment in one framework.

Dexterous robot learning needs large amounts of experience, but collecting demonstrations on hardware and manually building matching simulators are expensive. DEXTERA connects these steps in a unified workflow.

A single RGB image provides the visual input for a static Gaussian background and interactive rigid or articulated objects. Calibration aligns the reconstructed scene with known robot geometry. Task primitives, VR demonstrations, and object-centric trajectory synthesis then support imitation and reinforcement learning through a shared policy interface.

Experiments examine visual reconstruction, physical trajectory replay, and learned-policy deployment across two robot platforms. Simulation-only policies transfer to hardware, while adding limited real demonstrations improves physical success.

01

Build a usable digital twin

Combine image-based asset generation with metric scene alignment and arm–hand calibration.

02

Scale robot experience

Turn a small set of VR demonstrations into trajectories for randomized object configurations.

03

Connect learning to hardware

Evaluate six policy architectures through a consistent observation and action interface.

ON REAL ROBOTSOpenArm + BrainCo Revo1 / KUKA + LEAP Hand
Selected physical ACT policy rollouts. Playback is shown at 3×, as labeled in the footage.

THE FRAMEWORK

Four stages, from image to action.

Generate the scene, align it, define the task, and learn a deployable policy.

DEXTERA framework: scene factorization, metric alignment, task construction, policy learning and deployment
The complete real-to-sim-to-real pipeline.

01 / RECONSTRUCT

Separate appearance from interaction.

Factor the input image into a static background and individual interactive objects. Complete hidden geometry, generate textured meshes, and assign collision geometry and physical defaults.

  • Static Gaussian scene for visual context.
  • Rigid meshes and structured articulated assets.
  • VLM-inferred physical parameters for initialization.

OUTPUTVisual assets and simulation geometry.

Image-based reconstruction of interactive scenes.
See articulated asset examples

Link geometry, joint axes, and motion limits.

Input image for asset generation; calibration observations and the known robot model for metric registration.

GENERATING EXPERIENCE

A small demonstration set goes further.

Adapt demonstrated motion to new object configurations using task-relative poses.

VR teleoperation

Collect source motions for contact-rich manipulation in the reconstructed workspace.

10VR demos per task pair
300 + 30simulated + physical trajectories

Preserve the interaction. Vary the starting point.

Object-centric synthesis adapts end-effector motion to new task frames, interpolates waypoints, and adds action noise. Collision checks and task success criteria filter invalid simulated trajectories.

3,900
simulated trajectories
390
physical trajectories

Across 13 task–embodiment pairs. The physical trajectories are collected or replayed on hardware.

EXPERIMENTS

Reconstruction. Interaction. Policy transfer.

Three evaluations connect visual quality to physically executable motion and learned robot behavior.

01 / RECONSTRUCTION

Better geometry and appearance.

18 RGB-D captures across 10 tabletop scenes.

Selected object-instance metrics from Table II. Mean ± standard deviation over 86 instances.
MethodMask IoU ↑RGB MAE ↓Centroid error (mm) ↓F-score @ 1 cm ↑
DEXTERA0.731 ± 0.1360.182 ± 0.06211.5 ± 16.00.906 ± 0.132
Hunyuan3D-2.10.661 ± 0.1800.223 ± 0.06314.3 ± 17.30.750 ± 0.250
TRELLIS.20.559 ± 0.1960.250 ± 0.06325.8 ± 14.40.346 ± 0.202
SAM3D + FP0.604 ± 0.1430.407 ± 0.11818.3 ± 17.30.622 ± 0.235
SAM3D only0.488 ± 0.2850.391 ± 0.11129.9 ± 35.80.538 ± 0.339

RGB values are normalized to [0, 1]. Centroid errors are converted from meters to millimeters. DEXTERA has the best mean values for these selected metrics.

Paired simulated and physical executions. Playback speeds are labeled in the video.

02 / PHYSICAL CONSISTENCY

Simulated motion transfers to hardware.

87.95%

343 successful physical replays / 390 trials

For each task pair, 30 trajectories that succeed in simulation are executed on the real robot. This measures the physical consistency of the retained trajectories.

Learned-policy success is measured in separate experiments.

03 / SIMULATION–REAL CO-TRAINING

A little real data improves policy transfer.

Mean physical success across three policies on six selected task pairs.

29.2%SIM ONLY
61.9%SIM + REAL

+32.7 percentage points in mean physical success.

20 physical trials per condition. Co-training improves 17 of 18 policy–task comparisons, with one tie.

FULL POLICY BENCHMARK

Six architectures. Thirteen task pairs.

Mean success across all 13 task–embodiment pairs, weighted equally. Imitation policies use simulation and real training data; PPO uses teacher–student distillation.

0%25%50%75%100%
ACT
54.2%
Diffusion Policy
37.3%
BC-RNN
11.2%
OpenVLA
13.8%
π0.5
56.5%
PPO teacher–student
17.3%

Physical evaluation: 20 trials per policy–task pair. Values from Table V.

Explore task-level results 13 task pairs

13 task pairs · real robot success

Table V · Success (%) · physical robot
Task / robot platformACTDPBC-RNNOpenVLAπ0.5PPO T–S
OpenArmLift toy bus55%50%0%5%65%20%
OpenArmLift blue mug100%90%5%10%100%10%
OpenArmPlace toy bus in basket15%0%0%10%10%0%
OpenArmHand over gray bowl65%0%0%0%40%10%
OpenArmLift basket with both hands80%45%5%0%90%30%
OpenArmReorient thermometer55%15%0%0%50%0%
OpenArmLift box with both hands35%50%0%5%60%0%
OpenArmPush blue cube50%35%10%20%55%10%
OpenArmPlace box in basket with both hands10%5%0%0%20%0%
OpenArmClose laptop80%100%45%45%100%30%
KUKAPlace coconut-water carton in basket25%0%0%0%5%15%
KUKALift mustard bottle35%20%0%20%40%25%
KUKAClose laptop100%75%80%65%100%75%
Mean · all 13 pairs54.2%37.3%11.2%13.8%56.5%17.3%

Same action spaces and success criteria across methods. Simulation uses 100 trials per pair; hardware uses 20. The platform selector filters the table; the chart above always covers all 13 pairs.

ON THE ROBOT

Dexterous behavior on two platforms.

Physical ACT policy rollouts. Both demonstration panels use 3× playback.

01

OpenArm + BrainCo Revo1

Laptop closing, mug lifting, thermometer reorientation, and toy-bus placement.

ArticulationLiftingReorientationPlacement
02

KUKA + LEAP Hand

Laptop closing, mustard-bottle lifting, and coconut-water placement.

ArticulationLiftingPlacement

PROJECT VIDEO

The complete DEXTERA walkthrough.

2 min 59 sec · English narration
Captions available

Playback speeds are labeled in the footage; factors are relative to the source clips.

BibTeX

@misc{wu2026dextera,
  title = {DEXTERA: From a Single Image to Deployable Dexterous Manipulation via Real-to-Sim-to-Real},
  author = {Jin Wu and Lianjie Yuan and Zeyan Sun and Yuanyuan Lei and Disi A and Bicheng Han and Fangzhou Xia},
  year = {2026},
  eprint = {2609.21045},
  archivePrefix = {arXiv},
  primaryClass = {cs.RO},
  url = {https://arxiv.org/abs/2609.21045}
}