HomeArticle

Sun Peng, former core R&D member at ByteDance and Tencent AI, has joined Stardust Intelligence to take charge of the post-training work for robot reinforcement learning | Frontline

黄 楠2026-09-02 18:25
Improve the full-stack technology layout of Physical AI.

Author|Huang Nan

Editor|Yuan Silai

36Kr learned that on September 2, Dr. Sun Peng, former reinforcement learning expert at ByteDance and former head of the Agent Center of Tencent Robotics X, officially joined Stardust Intelligence, where he will focus on technical R&D and application exploration in the direction of robot reinforcement learning.

Dr. Sun Peng graduated from Tsinghua University, and then successively engaged in postdoctoral research at Cornell University and Rutgers University. He has long been focused on the research of Reinforcement Learning (RL), Agent and Multi-Agent Reinforcement Learning (MARL), with research and industrial practice experience spanning robot reinforcement learning, large-scale distributed RL systems and large model reinforcement learning, as well as technical accumulation in both algorithm research and system engineering.

During his early tenure at Tencent AI Lab and Robotics X Robotics Laboratory, Sun Peng served as the head of the Agent Center, conducting research on deep reinforcement learning and robot control. He once used deep reinforcement learning and adversarial games to train wheeled robots to achieve end-to-end active target following; at the same time, he developed the StarCraft AI agents TStarBots and TStarBotX, making important achievements in the research of multi-agent reinforcement learning (MARL) complex game environments.

After joining ByteDance, Sun Peng fully expanded his technical practice from algorithms to large-scale RL system engineering. He led the development of ByteRL, the core reinforcement learning infrastructure, which supports efficient training of multiple self-developed game AIs. His team once won the championship of the IEEE CoG 2023 Strategy Card AI Competition, and the Hearthstone AI they developed has the competitive level to beat top industry players.

With the rapid iteration of large language model technology, Sun Peng smoothly extended his years of accumulated reinforcement learning technology to the field of post-training of large models. During his tenure at Byte AI Lab/ByteResearch, he led the R&D and implementation of the reinforcement fine-tuning method ReFT and the mathematical reasoning agent DeltaProver as the project leader; he deeply participated in model pre-training and RLHF alignment work in the Seed team, and has accumulated solid technical reserves and implementation experience in cutting-edge directions such as model alignment, reasoning capability enhancement, and agent R&D.

Dr. Sun Peng (Source/Enterprise)

From robot control to large-scale RL system engineering, and then to post-training of large models, over the past ten years, Sun Peng's technical accumulation has always revolved around one core issue: how to enable agents to continuously learn through environmental interaction and feedback, and continuously optimize their own strategies.

Nowadays, AI capabilities are further moving from the digital world into the real physical world, and the reinforcement learning logic of "interaction-feedback-iteration" has also found a new foothold in the field of embodied intelligence. In the past, the industry competed on whether robots "can complete tasks", and verified model capabilities through simulation leaderboards; but when robots move out of laboratories to real deployment, the evaluation criteria are changing from "can it do it" to "can it do it stably and repeatedly". The long-tail interference in the real environment, the robustness of long-term autonomous decision-making, and the ability to recover from faults and exceptions are the keys that determine commercial value.

A repeatedly verified consensus is taking shape: imitation learning solves the problem of "how to do it", while reinforcement learning solves the problem of "how to correct it when you make a mistake". After the task success rate reaches 60%, the marginal contribution of continuing to optimize action reproduction decreases sharply, and the real bottleneck lies in the self-adaptation and error correction capabilities in an open environment. Post-training with reinforcement learning has become an unavoidable core link in embodied intelligence at this node.

This is exactly the key proposition that Stardust Intelligence must solve to promote the construction of a full-stack technical system. Since its establishment, Stardust Intelligence has adhered to the concept of Design for AI, and believes that the robot's body, data and AI models should not be designed in isolation. This idea has gradually evolved into a complete technical closed loop: the base model, front-end Agent and reinforcement learning post-training operate in synergy.

The Lumo series of base models are responsible for the robot's understanding, prediction and generation of the environment and actions. Lumo-1 enables robots to understand the intent logic behind actions, not just "how to do it", but also pays more attention to "why do it this way"; Lumo-2, as the first implicit world action model for households, introduces Latent World Dynamics, which first deduces the future state in the latent space before generating actions, which is equivalent to letting the robot "think first, then act". The front-end Agent Philia undertakes user-oriented intelligent capabilities such as long-term memory, task management, multi-robot collaboration and natural interaction.

Reinforcement learning is exactly the key link connecting "model capability" and "real deployment" in this system. It solves the problem of how robots can continuously learn and optimize their own behavior strategies through continuous interaction and task feedback in the real environment.

Based on the full-stack technical system of "AI Model + Embodied OS + Cable-driven Ontology", after Sun Peng joins Stardust Intelligence, he will further strengthen Stardust Intelligence's layout in the direction of robot reinforcement learning, and form synergy with the AI model, embodied OS and cable-driven robot ontology, so as to realize the continuous learning and continuous evolution of robots in the real world. In addition, the company will continue to expand the robot reinforcement learning team to promote the improvement of relevant technical capabilities and application deployment.

Regarding this joining Stardust Intelligence, Sun Peng said: "Over the past ten years, I have been working on reinforcement learning and agents, and have personally experienced the process of reinforcement learning moving from robot control and complex games to post-training of large models. This logic has been verified in the digital world, and now I want to bring these accumulations back to the physical world. Stardust Intelligence already has a complete technical foundation from robot ontology, embodied OS to AI models. What we need to do next is to truly integrate reinforcement learning into the training and deployment of real robots together, so that robots can not only 'learn a task', but also keep doing the task better and better in every real execution and feedback."