Building a full stack tactile intelligence ecosystem from tactile perception, data collection to embodied intelligence
Qianjue Robot has officially released its first VTLA embodied tactile model X-TouchMind V1, as well as the first visual tactile multimodal embodied dataset TacVerse 1k for real physical interaction.
This release marks the completion of the complete technical ecosystem construction of Qianjue Robot from tactile hardware system, data acquisition system, to data assets and embodied models, further promoting the evolution of tactile intelligence from perception ability to robot intelligence ability.
The next level of embodied intelligence: from modal post-processing to tactile native
In the process of robots moving towards real-world applications, relying solely on visual technology is no longer sufficient to meet complex task requirements. However, the additional tactile information often needs to go through a splicing path to integrate into the existing embodied model, which limits the upper limit of physical interaction ability.
Faced with this industry bottleneck, Qianjue believes that touch is the key mode for robots to truly enter the physical world, and it must be based on the native ecology and built around real strong contact scenarios on the full stack.
So that robots can not only perceive the world, but also understand and stably complete complex physical tasks.
02 X-TouchMind V1: Qianjue's first VTLA embodied tactile model
X-TouchMind V1 is a VTLA embodied tactile model trained on multimodal visual and tactile data by Qianjue.
X-TouchMind V1 focuses more on the physical state after contact occurs, modeling vision, language, touch, action, and robot body state in a unified manner, achieving collaborative modeling from task understanding, action generation to contact feedback control.
This model adopts a hierarchical architecture, corresponding to different levels in real robot operations:
System 2: Task Understanding and Semantic Reasoning Layer, responsible for understanding the environment, task objectives, and language instructions, and answering 'what to do'.
System 1: Action generation and trajectory planning layer, generates action paths based on vision, task semantics, and robot state, and answers' how to do it '.
System 0: Tactile ontology interaction control layer, using tactile and ontology feedback for high-frequency correction in real contact, answering "How to make stable and accurate after contact"
Through this architecture, X-TouchMind V1 can not only generate action sequences, but also understand slip, force, deformation, texture, and contact stability in real operations, and dynamically correct actions, thereby improving the stability of the robot in tasks such as fine operation, flexible material processing, and intelligent quality inspection.
The goal of X-TouchMind V1 for Contact Rich Contact tasks such as opening, pulling, pressing, inserting, aligning, and restoring is not to enable the robot to complete a single action demonstration, but to enable the robot to perform tasks stably, reliably, and without damage in complex physical interactions.
03TacVerse 1k: First 1000 hour high-quality visual tactile multimodal embodied dataset
The TacVerse 1k, released simultaneously with X-TouchMind V1, is the first 1000 hour high-quality visual tactile multimodal embodied dataset built by Qianjue for real physical interaction.
On the data collection side, TacVerse 1k relies on its self-developed collection hardware system to decouple data production from the robot body, achieving capacity decoupling, scenario decoupling, and ontology reuse. Based on this system, the cost of data collection can be reduced by over 80%, and efficiency can be improved by 3 to 5 times, providing a foundation for large-scale real interactive data precipitation.
On the data quality side, Qianjue's self-developed XTacFlow is responsible for unified scheduling of the acquisition link, supporting unified access of multiple devices, millisecond level cross modal time alignment, and achieving over 90% of data feedback and post-processing automation, converting raw multimodal data into trainable, traceable, and reusable data samples and training assets.
The TacVerse 1k data structure is natively compatible with mainstream robot data formats such as LeRobot, MCAP, HDF5, ROS2, and standardizes tactile modal extension fields, making the data not only available internally but also has a foundation for industry sharing, reuse, and benchmarking.
According to the release plan, the first batch of 100 hour datasets for TacVerse 1k will be open for application on July 30th, covering over 40 real-life scenarios; Starting from August, applications for 1000 hour datasets will be opened in batches, and open cooperation will be promoted through platforms such as Alibaba Tower and Hugging Face.
Targeting real industrial tasks, promoting the large-scale implementation of physical intelligence
The simultaneous release of X-TouchMind V1 and TacVerse 1k marks the official entry of Qianjue into a new stage of industrial landing.
In the future, Qianjue will focus on conducting in-depth industry cooperation in various scenarios such as industrial precision assembly, consumer electronics assembly, flexible logistics, intelligent quality inspection, consumer product authenticity verification, and data center construction.
The person in charge of Qianjue Robotics stated that we hope the release of X-TouchMind V1 and TacVerse 1k will continue to provide value to the industry, opening a new stage of industry landing with "tactile native" technology. ”
In the future, Qianjue will continue to promote technological iteration and theoretical exploration around "tactile native" and "real scenes", and make progress together with the industry.
Original title: Heavy Release | Qianjue Robot releases first VTLA embodied tactile model X-TouchMind V1 and first batch of visual tactile multimodal dataset TacVerse 1k