AI has launched research on Physical AI: the FSD-level team unveils its first version model Simate-beta, making a surprise debut in RoboDojo.
AI is now conducting AI research, and this time it has delivered an answer sheet for the physical world.
A company established just three months ago has unveiled its first general-purpose physical fast system, Simate-beta: demonstrating capabilities in memory, long-horizon tasks, fine manipulation and task adaptation, and according to the company's disclosure, it has topped the RoboDojo leaderboard.
The leaderboard certainly cannot represent all the capabilities of the model, but it provides a perspective to observe the R&D speed and strength.
According to the team, the base model submitted for evaluation this time has not undergone any special optimization for this leaderboard.
Delivering such results with the first version of the model just three months after establishment is well worth attention.
The team is called Simate (Silicon Mate).
It also has an impressive background. According to the company, the core team once advanced the capability of the one-stage end-to-end autonomous driving model to a level comparable to Tesla FSD, and completed mass production and deployment.
Across the globe, very few teams have crossed this threshold; it is even rarer to carry forward this accumulation to continue advancing general physical AI.
However, the most interesting part this time is the R&D approach they set on the very first day of founding:
AI for Physical AI, letting AI participate in the research and iteration of physical intelligence itself.
The company organizes models, data, computing power and experimental workflows around this roadmap. Simate-beta is exactly the first practical product of this AI-native R&D system.
Moreover, the research tools have already been adopted by external researchers.
According to the company, researchers from universities including MIT, Caltech, Tsinghua University and Peking University have participated in the internal beta test of the AutoResearch platform. Its supporting Sinfra infrastructure covers training, simulation and inference, and specific available resources will be confirmed upon application.
In the same period, Simate has successively completed multiple rounds of financing at the level of hundreds of millions of RMB.
In just three months, the model, research platform, external trial use and financing have been advanced simultaneously. This speed really makes people eager to see what comes next.
The team also released a preview: It has formed technical ideas for zero-shot generalization of complex tasks, and plans to launch phased results by the end of this year.
In their view, it is an expected progress that GPT-6 begins to participate in physical world manipulation, and the breakthrough of Physical AI may come faster than the outside world imagines. They hope to further break through the dependence on task-by-task adaptation and post-training, and push physical AI to usher in its own "GPT-3 moment".
Three research works including models and automated research will be released successively through papers and technical reports, and relevant achievements are planned to be open-sourced in phases.
This R&D speed, and whether it can be sustained, has become the most noteworthy thing for Simate next.
So what exactly has the first model achieved? How does AI participate in the research?
Simate-beta: General-purpose physical fast system for zero-shot generalization
What Simate ultimately wants to approach is zero-shot and few-shot general physical AI manipulation.
In a completely new and unfamiliar environment, the system should not need to re-collect a batch of data, train the model once, and re-adjust the engineering every time a new task is added.
This is exactly the most popular and most challenging track in current physical AI - hundreds of schools of thought are contending and showing their own strengths.
The current solution given by Simate is called: General-purpose physical fast system.
That is, System 1.
Let's take humans as an example first:
Facing complex instructions like "help me tidy up the table", we need "System 2" to understand the intention and break down the steps;
But when the hand is already reaching for the cup, the millisecond-level action response of whether to adjust a few millimeters to the left or deviate a little to the right cannot rely on deep thinking for every step.
The same goes for physical AI.
A general "slow system" is responsible for high-level planning and complex reasoning, while a sufficiently powerful "fast system" is responsible for real-time perception of the constantly changing physical world in front of you, and outputs actions at extreme speed.
Simate chose the latter ——
Expand the parameter scale of the fast system, explore the emergence of zero-shot generalization capabilities, and challenge the upper limit of the scale and capability of edge-deployable models.
A general-purpose physical fast system.
However, making it both large and fast is a very difficult trade-off.
Simate chose to start with two things:
1. 4D Physical Perception
To understand the physical world, it is necessary to capture the dynamic changes of spatial structure and time dimension at the same time, so that physical representations can truly serve actions.
2. Hierarchical Temporal Memory
Long-horizon tasks require remembering what happened before, continuous manipulation requires tracking status, and immediate response cannot be slowed down by huge amounts of historical information.
4D physical perception and hierarchical temporal memory jointly point to the alignment with the way the human fast system processes the continuous world.
This roadmap also continues the team's accumulation in large-scale vision models, temporal modeling, real-time inference and mass production deployment. The specific architecture and model scale will be disclosed in subsequent technical reports.
The final expected state is: the model clearly knows what is happening in front of it, also remembers what happened just now, and can decide the next step in time.
Then cooperate with the general reasoning capabilities represented by cutting-edge models such as GPT-6 to advance towards zero-shot execution of complex tasks.
Judging from the real machine demonstration of Simate-beta currently provided by the company, the test mainly focuses on demonstration-driven task adaptation, memory, complex long-horizon execution and fine manipulation.
According to the RoboDojo leaderboard updated by the company on September 23, 2026, Simate-beta ranks first with an average Score of 33.95, corresponding to an SR of 27.96%. This result belongs to the leaderboard evaluation, and the real machine performance will be displayed separately.
Behind this model answer sheet is another set of systems that are also worth expanding on.
Make the flywheel spin
In recent years, people have become more and more accustomed to AI writing code.
But in fact, in AI research, writing code is only a very small part.
Most of the time, iteration efficiency is the key to winning. Whoever can verify more research hypotheses per unit time can run ahead first.
Take the simplest example.
Robot long-horizon tasks always get stuck at a certain link, and the reasons are nothing more than several: distorted data distribution, model structure bottleneck, unbalanced trajectory proportion, or invalid training strategy.
Behind each direction, there are more than a dozen or even dozens of groups of experiments.
If all rely on researchers to manually modify configurations, start tasks, monitor training, run evaluations, and then sort out the results, the opportunity will be completely missed.
So Simate built AutoResearch to accelerate this process ——
Human researchers: Propose hypotheses, set goals, and define constraints.
AutoResearch engine: Automatically takes over, and automates the whole process of disassembling, running and feeding back the remaining experiments.
The final result is also revealed, they relied on this method to take the first place on RoboDojo.
The more critical point is the speed.
It only took Simate three months from its establishment to topping the RoboDojo leaderboard.
It's almost unbelievable.
So the question arises:
How exactly should AI "research AI"?
Obviously, just throwing an Agent into the process is far from enough.
Real robot R&D penetrates at the bottom layer through model code, data pipeline, training cluster, evaluation environment, real machine feedback, and experimental records of past months or even years.
In the final analysis, this is still a problem of "context". The Agent needs a super interface that can schedule everything.
To this end, Simate has built a three-layer architecture: SiPAI + AutoResearch + AI-native Infra.
SiPAI: Make the model architecture "AI-researchable"
This is a set of pluggable model frameworks natively oriented to "automated research".
In simple terms, it builds the robot model into a structure that AI can easily understand — the module boundaries, system interfaces, configuration items and verification processes are all defined extremely clearly.
Only in this way, after the Agent gets the research topic, can it accurately know: which module to modify? What will the modification affect? And how to verify it in the end?
At present, Simate has completed the construction of this set of infrastructure, and has successfully accessed a variety of mainstream architectures such as world model, world action model, VLA and VLM.
Both human researchers and Agents can freely combine components, adjust local implementations, and follow a unified training and evaluation pipeline.
AutoResearch: Provide real-time refreshed "research context"
If SiPAI provides a static instruction manual, then AutoResearch solves the dynamic information gap:
Where is the research progress? What pitfalls have been stepped on before? What is the latest evidence?
It connects all external cutting-edge papers, internal team discussion materials, model code, data ratio, historical experimental logs and the latest evaluation feedback, and integrates them into the same dynamic research context.
No matter a new paper is published, a new version of internal code is submitted, or the latest experimental data overturns the old hypothesis, this Context will be refreshed in real time.
In this way, every research suggestion put forward by the Agent can be accurately anchored to the specific code branch and experimental version:
The most interesting part lies in the deep coupling of "human experience" and "Agent experiments":
Human researchers can inject constraints and adjust optimization directions at any time; once an experiment verifies an effective component, it will immediately precipitate as a new starting point for subsequent experiments; the methodology summarized from a certain branch can also be seamlessly reused by the next Agent.
Research results have truly become the means of production for the next round of research.
AI-native Infra: Make the 7×24 hour R&D flywheel roll
Of course, it is not enough for the Agent to only write code, the experiments must run efficiently.
Simate connects the whole process of training, inference and evaluation to its self-developed Infra, and simultaneously advances dozens of independent research routes in parallel through extreme task orchestration and resource scheduling.
In order to prevent computing power waste, Simate adopts a high-frequency screening mechanism: all routes first go through the first round of rapid filtering via "world model + simulation environment" to eliminate most invalid directions;
Only solutions with real potential will enter the real machine test.
When new problems are exposed on the real machine, they flow back to the research Context, and the closed loop is thus formed:
The final presented effect is: Humans only need to put forward an exploration direction, and the underlying system can automatically and concurrently drive dozens of rounds of experiments to advance rapidly.
This is also the key magic weapon for Simate to reach the top in three months.
However, the most unexpected thing is that they chose to directly open source and open such an extremely core internal productivity tool!
The web page looks like this, interested friends can get it at the end of the article.
At present, researchers from top universities such as Tsinghua University, MIT, and the Hong Kong University of Science and Technology have taken the lead in entering the internal beta test.
Weak RSI is just the starting point
The founder Zhang Ying, former core technical person in charge of a leading autonomous driving company.
He has very rich engineering practical experience, participated in the research of three generations of autonomous driving solutions including map-based, mapless, and end-to-end, and has done mass production for many years.
According to the introduction, the autonomous driving system he promoted to build is the autonomous driving system closest to Tesla FSD in China.
In addition, there are two other key figures in the team.
1. Zhan Fangneng.
Assistant Professor at the Hong Kong University of Science and Technology, Head of World Mind Lab, long-term