HomeArticle

Early robotics education calls for a kindergarten that allows "making mistakes".

周谣2026-05-26 18:29
How does the "kindergarten" co-built by Sutton and Tashan Technology enable robots to learn autonomously and continuously?

In 2024, Richard Sutton, the founding father of reinforcement learning, and his mentor Andrew Barto jointly won the Turing Award.

This award did not come early. Over the past three decades, Sutton's theories have supported the evolution of systems including AlphaGo and ChatGPT, but the theories he put forward 30 years ago have only been truly understood by the embodied intelligence industry today:

Agents must learn from trial and error, and evolve from real experience.

In 2023, Sutton co-founded Openmind, a non-profit research institution. In April 2025, in his co-published article *Welcome to the Era of Experience*, he once again pointedly pointed out:

"The new generation of agents must possess a stream of experience that advances continuously over long timescales like humans, and achieve self-evolution in real physical feedback."

This time, beyond theories, Sutton has set his sights on a farther horizon.

In May this year, Sutton officially signed a cooperation agreement with Tashan Technology in Canada, to jointly promote a project called "Robot Kindergarten" in the form of long-term collaboration.

A Turing Award winner and a Chinese tactile technology company hit it off immediately, jointly making a judgment in advance for the next decade of embodied intelligence: a brand-new path for training robots may lie precisely in real touch and trial and error.

What Embodied Intelligence Lacks is "First-Person Experience"

Ma Yang, CEO of Tashan Technology, put forward a very straightforward judgment. For robots to perform practical work, they only need to solve two core problems: one is the movement of the robot itself in the physical world, which can be realized through biped, quadruped, wheeled and other forms, and many companies are already working on this direction.

The other is to operate target objects, to grasp, place and twist objects with hands, with smooth actions that will not be interrupted by the deviation of the previous movement. The combination of these two capabilities can basically cover 90% to 95% of the work that humans currently expect robots to complete.

From the very beginning, Tashan Technology has been focusing on the tactile track to make the second capability well implemented.

When Tashan Technology was first established in 2017, most robot manufacturers were developing mobile platforms, showcasing their capabilities of running, jumping and rolling. However, more than 90% of human physical interactions are actually completed through fingers. Unlike legs, fingers need to keep in contact with different target objects all the time, to perceive, make decisions and adjust, which is a difficult and continuous process.

To properly solve the "finger function" of embodied intelligence, tactile perception capability is a core variable, and also the underlying methodology to "enable robots to work". Tashan Technology has been sticking to this track for nearly a decade.

The mainstream training direction of embodied intelligence relies on end-to-end imitation based on static datasets, which is like memorizing question banks. The data demonstrated by humans is essentially second-person experience. When robots learn human practices without touching the objects by themselves, they cannot understand the operation laws of the physical world.

Tashan Technology realized the problems of this route very early: just as human children need to grow up from imitation and practice in their early years, the "enlightenment" training of robots requires not only imitation, but also their own first-person experience.

The training method of perceiving consequences in actions and adjusting behaviors in feedback may be the methodology closest to enabling embodied intelligence to carry out "self-training".

This judgment coincides with Sutton's idea.

The concept of "experience stream" proposed by Sutton requires the learning process and behavior process of agents to be completely integrated: every action is data collection, and every feedback is a training signal. Therefore, a real environment that can provide first-person experience is the key to the implementation of this concept.

However, this concept has remained at the theoretical level for a long time, precisely because the real physical environment cannot provide low-cost, high-frequency and standardized interactive feedback. For a long time, the embodied intelligence industry has been committed to solving the problems of the brain and eyes, lacking a channel that can accurately perceive physical contact.

Tactile sense is the most core perception channel in physical interaction. When a robot touches an object, the tactile sensor can feedback the 3D force distribution of the contact point, the local deformation of the object and the slip trend in real time. With this information, the robot can quickly adjust its force and angle, and decide to tighten or relax its grip.

The continuous emergence of high-precision tactile perception technologies has completed the "afferent nerve" that robots once lacked, and theoretical pioneers represented by Sutton have also begun to focus on this field. In November 2025, Sutton visited China, and Tashan Technology was one of the two embodied intelligence companies he took the initiative to contact for a visit.

Sutton visiting Tashan Technology

Tashan Technology is the company with the most complete technical reserves in the tactile perception track.

The self-developed tactile sensor of Tashan Technology has a force resolution of 0.01N, which is "similar to the force of a hair falling on a finger". With years of R&D in AI tactile perception technology and full-stack tactile solutions, it has overcome the global technical difficulty of simultaneous parsing of multi-dimensional tactile perception signals, and built a complete technical system covering "chips - sensors - algorithm models - scenario applications".

While most tactile sensor manufacturers are still staying at the single-dimensional force measurement or simple capacitance change level, Tashan Technology has realized the simultaneous parsing of 3D force, material recognition, proximity perception and collaborative perception.

More importantly, Tashan Technology has realized mass production of tactile perception capabilities. In the past two years, its products have entered the commercialization stage, and have begun to deliver in batches to mainstream dexterous hand manufacturers. In 2025, Tashan Technology accounted for more than 80% of the market share in the humanoid robot tactile sensor track.

TS-VT Visual-Tactile Fusion Training Platform

After Sutton visited Tashan Technology, the two sides quickly promoted the cooperation. In addition to the matching of methodologies, he also saw a team that had pushed tactile perception from the laboratory to industrial implementation in the building of Tashan Technology.

Thus, 30 years after the release of the reinforcement learning theory, the theory and technology have achieved two-way integration in the field of embodied intelligence: the academic titan has found an ally who can engineer the theory, and Tashan Technology has supplemented the theoretical puzzle of using tactile sense to accelerate robot training.

Robot Kindergarten: "Enlightenment" in the Real Environment

The foothold of the cooperation between the two sides is the "Robot Kindergarten" in specific form.

At Tashan Technology, Sutton saw Chinese primary school students taking robot courses, and was amazed that the domestic embodied enlightenment environment was so open, where humans and robots could coexist more naturally. The idea of Robot Kindergarten came into being from this experience.

Robot Kindergarten is a tactile and multi-modal experience training platform oriented to the continuous learning of robots. It integrates real physical environments, simulation environments, multiple robot bodies, tactile and multi-modal perception devices, task courses, data collection and evaluation mechanisms, so that robots can form trainable experience through repeated contact, attempts, failures and corrections.

Why is it called Kindergarten? Ma Yang said that the current embodied intelligence is very similar to a baby in the 0-3 age stage. We see that robots can do all kinds of things in videos and think they are very capable, but in fact the success rate is not high, and the robot itself does not even know whether it has succeeded or failed. "It just completed the action, and people will applaud for it."

It is actually very difficult for robots to understand what they did correctly through human correct demonstrations, because the definition of "correctness" is very vague and covers a very wide range. Only errors have clear boundaries. Sufficient error experiments can help a robot know where the boundary of a task is, and how to adjust itself in the next operation.

"The safety boundary of embodied intelligence is not defined by everyone drawing a line together, but gradually explored by the robot in objective interactions."

Ma Yang firmly believes that just as human safety instincts are not only obtained by reading manuals, but also developed through repeated contact, falling and adjustment, robots are the same. Only through sufficient real trial and error can they understand what is unsafe. If a robot can draw its own safe operation boundary, it will not only protect itself, but also demonstrate safety for human beings.

After Sutton visited Tashan Technology, the two sides quickly promoted the cooperation matters, and completed the signing on May 11, 2026.

At the signing ceremony, Sutton talked about the significance of the cooperation: "As early as when we were graduate students, someone proposed to create a robot like a baby, let it interact with the world and grow through experience. This idea was almost impossible to realize at that time. Now we have sufficient computing power and sufficient robot experience, but I think the key missing factor has always been the clear recognition of the value of this ideal. What it needs is not just capital, but more importantly, time and persistence."

Sutton stated that when he visited Tashan Technology, he was pleasantly surprised to find that this Chinese company has understood this point. The entire cooperation plan has a five-year cycle, with the goal of finding the most suitable learning methodology for embodied intelligence.

Scene of the Signing Ceremony

Next, "Robot Kindergarten" will build a real environment, and place robot bodies in it to complete training. Although the initial stage is training with isomorphic robot bodies, Ma Yang believes that under the exploration of continuous learning, heterogeneous robots will not become a major obstacle to learning in the later stage. Because if an agent understands the underlying logic of a task, the difference in body form will not hinder the transfer of learning and experience.

In contrast, it is more important to face the real environmental variables directly at present.

Ma Yang said frankly that the hardware of the embodied intelligence industry has reached a 60-point level, and what is lacking is reasoning ability and continuous learning ability. Without these two capabilities, it is impossible to achieve better generalization and deduction, and the whole industry will be forced to compete on parameters, failing to find a broader application space.

Therefore, early learning must interact with the real environment continuously. The built training environment can no longer deliberately avoid the variables and unfavorable factors in real scenarios, otherwise the experience ceiling that robots can learn will be very low, and it will be difficult for them to make further progress.

The cooperation between Tashan Technology and Sutton is also aimed at finding a new path. "There is no sophisticated high-tech in this matter, only the choice of methodology."

The Premise of Commercialization is "Learning While Working"

Methodologies ultimately need to be tested in application scenarios. Ma Yang also has a very pragmatic judgment on commercial implementation: in the next 3 to 5 years, the scenarios that embodied intelligence will most likely enter first will not be those with high logic and high timeliness requirements.

It is more suitable to replace a specific type of work: work that humans do not want to do, and cannot have too low fault tolerance.

This type of work has three characteristics: the tasks are repetitive, but not completely fixed assembly line work; the requirements for success rate are very high, a single failure may directly interrupt the whole process and require strong manual intervention; the timeliness requirement for a single task is relatively loose, and second-level response is not needed.

Ma Yang gave several examples: one is the service industry scenario, the dishwasher in North American restaurants. Their work is to rinse the dishes, put them into the dishwasher, with very simple actions but boring and heavy work. At present, millions of people in the United States are working in this position. If robots can achieve a sufficiently high success rate for this action, they can release huge commercial value. At the same time, the dishwashing task does not have high timeliness requirements, as long as all dishes are washed before the end of the night. But the success rate requirement is very high: if one bowl is broken, the whole process has to stop.

There is a more specific case in the agricultural processing field. In the crayfish processing factories in Qianjiang, the step of "removing the head of crayfish" has always been completed manually. Because crayfish are different in size, and the hardness of their shells varies with seasons, the tactile perception technology of the equipment is required to be very high. The labor cost of one factory on this process alone is as high as tens of millions of yuan every year, and 1,000 to 2,000 people are working on the production line during peak hours.

Tashan Technology spent half a year doing imitation learning and simulation training first, and then let the robot practice autonomously repeatedly through reinforcement learning on the real production line. Finally, the success rate of shelling crayfish was increased to more than 95%. While efficiently removing the crayfish head, the crayfish roe is completely retained, improving the value structure of the product. At present, the intelligent crayfish shelling equipment of Tashan Technology has reached cooperation with leading crayfish processing enterprises, with the first batch of 100 units signed.

Intelligent Crayfish Shelling Equipment of Tashan Technology

The logic for selecting these scenarios is very clear. At present, robots cannot compete with humans in reasoning speed, but they are very suitable to fill the gaps that cannot be automated and that humans are unwilling to do. Tactile perception is the key to unlock these scenarios. Because it provides real-time feedback, robots can flexibly adjust their force and angle during the execution process, without requiring a perfectly pre-set trajectory.

If most of the energy in the industry is focused on training robots to imitate humans, the ceiling of embodied intelligence will be human beings themselves. To break through this ceiling, the whole industry needs to explore together.

Ma Yang has always emphasized that compared with Tashan Technology's own moat, he hopes to see more peers join in and push forward in the correct direction together. Tashan Technology and Sutton hope to build an open and shared R&D infrastructure, attracting global academia and industry to jointly explore the methodology of continuous learning for embodied intelligence.

At this stage, Tashan Technology and Sutton, as the initiators, will focus on building the platform. In the future, the whole system will be gradually open to the whole industry, and the upstream and downstream of Tashan Technology's industrial chain, global universities and scientific research institutions may all become ecological partners in this cooperation project.

The combination of tactile perception and continuous learning is paving the way for the next decade of embodied intelligence.

Sutton's answer has already been written in the vision of the real experience stream. Tashan Technology will soon turn this answer into an executable engineering solution with a Robot Kindergarten, enabling embodied