A unified benchmark for robots that see and touch

OpenViTac

Learning and Benchmarking Visuo-Tactile Policies
in a Unified Sim-and-Real Framework

Yifan Wu1,*Qin Li3,*Nan Min1,*Guojin Zhong1,†Haoyu Zhao1,†Zhiyuan Li1Houze Xu1Shengqi Xu1Xingyao Lin1Zijie Diao1,2Zhaoxiang Liu6Shiguo Lian6Shunlin Lu5Shihao Zhao5Ziyi Ye1Zuxuan Wu1,2,5Yu-Gang Jiang1
1 Fudan University2 Shanghai Innovation Institute3 Hefei University of Technology5 Neote AI6 China Unicom

* Equal contribution † Project leaders

OpenViTac overview: paired simulated and real tasks, four tactile capabilities, two gripper setups, three tactile sensors, and policy evaluation.
One benchmark. Paired simulation and reality. Four dimensions of touch.

Beyond what vision can tell.

Weight, texture, fragility, and contact are often hidden from a camera. OpenViTac brings these physical challenges into a unified benchmark, with corresponding simulation and real-world tasks for evaluating vision-language-action, world-action, and vision-tactile-language-action policies.

Alongside the benchmark, we introduce OpenVTLA: a practical way to adapt pretrained VLA models to touch, and study how simulated experience can improve real-world manipulation.

Simulation tasks
11
Paired real-world tasks
8
Tactile capabilities
4
Tactile sensor types
3

Put touch to the test.

From recognizing an object's physical properties to making a precise insertion, each task isolates a reason to feel.

See the complete capability taxonomy Taxonomy organizing benchmark tasks by physical-property perception, fragility-aware, contact-rich, and precision capabilities.

From the real world, into simulation.

Aligned objects, scenes, and task configurations connect virtual evaluation with physical manipulation.

Agent-assisted real-to-sim pipeline reconstructing assets from multi-view images and building corresponding tasks and trajectories.
Agent-assisted asset reconstruction and task construction preserve task-relevant structure across domains.

Corresponding tasks

Eight shared tasks support comparison under matched manipulation objectives in simulation and on real hardware.

Diverse environments

Backgrounds, workspace appearance, and surrounding context vary while the underlying task configuration stays consistent.

The same manipulation scene across a range of simulated backgrounds and workspace appearances.
Scene diversification broadens the training distribution.

Introducing OpenVTLA

Give a pretrained VLA a sense of touch.

A temporally aware tactile encoder. Early token-level integration. The original vision-language and action pathway.

Design space comparing four tactile representations and five integration strategies for adapting pretrained vision-language-action models.
We study what tactile information to represent, and where to introduce it into a pretrained policy.

Temporal touch, integrated early.

OpenVTLA uses AnyTouch2 to encode tactile history into compact tokens. These tokens are projected to the VLM hidden dimension and prepended to the visual sequence of π0.5.

Token-level concatenation lets touch participate from the VLM input, without an additional fusion module or tactile-specific expert.

Tactile representations studied
4
Integration strategies studied
5
Radar plot comparing tactile representations under a fixed token-level concatenation strategy.
Representation comparison with the same integration strategy.

Evaluate in sim. Validate in reality.

OpenVTLA achieves the highest average success rate among the evaluated policies in both domains.

68.7%

OpenVTLA in simulation

54.6%

OpenVTLA on real hardware

0.913

Sim-real Pearson correlation

Success rates (%). Average is the unweighted mean across tasks. All policies use single-task training. Full task names are available in the column headings.

Simulation makes real data go further.

With 100 real demonstrations per task, adding 500 simulated demonstrations raises OpenVTLA's average success by 15 percentage points across USB insertion and gear assembly.

OpenVTLA real-world success (%)
Training data Insert USB Gear assembly
100 real 20 50
100 real + 500 sim 40 60

Citation

If you find OpenViTac useful for your research, please consider citing this work.

Preliminary citation. Publication details will be added with the paper release.

BibTeX
@misc{wu2026openvitac,
  title={OpenViTac: Learning and Benchmarking Visuo-Tactile
         Policies in a Unified Sim-and-Real Framework},
  author={Wu, Yifan and Li, Qin and Min, Nan and Zhong, Guojin
          and Zhao, Haoyu and Li, Zhiyuan and Xu, Houze
          and Xu, Shengqi and Lin, Xingyao and Diao, Zijie
          and Liu, Zhaoxiang and Lian, Shiguo and Lu, Shunlin
          and Zhao, Shihao and Ye, Ziyi and Wu, Zuxuan
          and Jiang, Yu-Gang},
  year={2026}
}