HomeArticle

Tong Xin, the legendary master of computer graphics, has joined Meshy, aiming to be the leading player of "AI for Fun".

量子位2026-09-07 15:24
Let's jointly find the global optimal solution for "concretizing imagination".

The next high ground of AI innovation is converging towards 3D generation.

From multimodal generation and world models to spatial intelligence and embodied intelligent robots...

3D is becoming the core engine of the next-generation AI.

Now, the grandmaster-level figure in the field of computer graphics has joined hands with the most sought-after star company on the AI 3D track.

Latest news:

Xin Tong has officially joined Meshy, the company founded by Yuanming Hu.

The Historic Partnership Between Yuanming Hu and Xin Tong

According to the latest news, Xin Tong will serve as the chief scientist of AI 3D company Meshy, responsible for formulating Meshy's long-term scientific research strategy.

This may mean that the two most representative generations of researchers in the Chinese computer graphics community — Yuanming Hu, founder of Meshy, and Xin Tong — will jointly promote breakthroughs in multimodal world models and real-time interactive systems.

Xin Tong previously worked at Microsoft Research Asia for 25 years, during which he led the Internet Graphics Group and served as a Global Research Partner.

His research interests cover computer graphics and 3D computer vision, specifically including but not limited to material capture and modeling, texture synthesis, 3D geometry processing and modeling, light transport analysis and simulation, photorealistic rendering, and 3D face animation.

In simple terms, computer graphics focuses on "how to create, manipulate and present the 3D world in computers".

This field occupies a fundamental position in games, film and television, and industrial simulation, and is increasingly becoming the infrastructure for AI perception and expression at present.

After all, without understanding 3D, AI can hardly truly understand and enter the physical world.

And Xin Tong is one of the scholars who has taken root in this direction for the longest time and accumulated the deepest attainments.

In 1999, right after Xin Tong graduated from Tsinghua University with a doctorate degree, he joined Microsoft Research Asia (then named Microsoft Research China), which was founded less than a year and later known as the "Huangpu Military Academy of China's technology sector" Microsoft Research Asia.

Over the past two decades, countless core figures that have influenced the pattern of computer research in China and even the whole world have emerged from here — but that is another story.

As one of the first batch of researchers at MSRA, Xin Tong stayed here for 25 years, growing from a researcher all the way to Global Research Partner and head of the Internet Graphics Group.

He has almost witnessed all the important moments of the development of computer graphics in China.

During this period, Xin Tong has published a large number of papers in top journals and conferences, and his Google Scholar citations have exceeded 21,000 times.

These technical researches have left many valuable achievements for Microsoft, including Xbox game development APIs, Xbox compatible software, Windows 3D print drivers, and Direct3D graphics development kits, etc.

In addition, Xin Tong left another "asset" to the Internet Graphics Group of MSRA: a group of backbones who are fully capable of supporting the current Chinese computer graphics community.

Over 25 years, he has cooperated, communicated with, and guided batches of outstanding researchers here.

Kun Zhou from Zhejiang University, ACM Fellow, IEEE Fellow, Director of the State Key Lab of CAD&CG, has cooperated with Xin Tong for many years, exploring from 3D printing and computational design all the way to NeRF and neural rendering;

Kun Xu from Tsinghua University, tenured associate professor of the Department of Computer Science and Technology of Tsinghua University, National Distinguished Young Scholar, has cooperated with Xin Tong to publish multiple SIGGRAPH papers during his PhD study;

Pengshuai Wang from Peking University, assistant professor of the Wang Xuan Institute of Computer Technology of Peking University, has been working under the guidance of Xin Tong from an intern to a senior researcher, and is now a representative scholar in the field of 3D geometric learning...

These people who have worked with him often have surprisingly consistent evaluations of Xin Tong:

He is approachable, always devotes himself to front-line work, can conduct rigorous discussions on extremely detailed technical issues, and can also explain complex technical issues in plain language.

Because in Xin Tong's view, computer graphics is an applied discipline of computer science. For this reason, all research problems come from practical needs and the emergence of new devices and new data.

In addition, Xin Tong is also very humorous. For example, he described his research in this way:

Apart from recognizing the delicious food I cook, my family often jokes that I am obsessed with 3D teapots and rabbits on the screen and enjoy myself all day long, but what I do seems to be neither able to change the world situation at a macro level nor improve people's lives at a micro level, and they really don't understand where my sense of achievement comes from.

This is the well-known "Old Tong" in the computer graphics circle, who has extremely profound academic skills but maintains the childlike kindness and curiosity.

The description is very appropriate.

"AI for Fun"

It can be said that at this stage of Xin Tong's career in the computer graphics industry, only a star unicorn like Meshy can match his influence in the field.

Meshy is the first company founded by Yuanming Hu after he graduated from MIT with a doctorate degree. Its positioning is very simple, that is, to turn text and images into 3D models.

At present, Meshy has gathered 12 million users, and its customers cover half of the top 10 enterprises in the world by market value and valuation.

In July this year, Meshy completed a nearly 400 million US dollar Series B financing, with a post-money valuation of over 10 billion RMB.

It has set two global records of single-round financing scale and valuation in the AI 3D field in one fell swoop, becoming the company with the highest valuation in this field.

However, no matter the faster growth rate or higher valuation, they are too specific compared with Meshy's vision.

Yuanming Hu, founder of Meshy.ai, champion of 400m race at Yangzhou High School Sports Meeting, second prize winner of the singing competition in Tsinghua University's Yao Class, PhD in computer graphics at MIT, SIGGRAPH Best Doctoral Dissertation Nominee, author of the Taichi programming language and compiler.

He put forward a rather unique goal that sounds a little special:

AI for Fun.

His logical chain goes like this.

In the future, when AGI solves a large number of productivity problems, there will only be one core problem left for humanity.

"How do we create, express and connect? The most important thing is, how do we gain a sense of meaning?"

Reading, traveling, playing games, watching short dramas... these can indeed create happiness for us, but in an era when free time is so abundant, we "will definitely pursue more novel and more effective ways to generate happiness and a sense of meaning."

Such methods are destined to be driven by AI.

Therefore, "using AI to bring happiness and a sense of meaning to humanity will be one of the most important issues in the next five years".

And Meshy wants to become a crucial piece of the puzzle in this vision.

In the words of Yuanming Hu, what they want to do is "reshape computer graphics based on generative AI, and find the global optimal solution to materialize imagination."

What does that mean?

In other words, today, what AI 3D needs to do is not only to allow AI to complete modeling, texturing and rendering more efficiently under the guidance of computer graphics.

In the long run, the ultimate destination of AI 3D may be the materialization of imagination.

This forward-looking perspective is quite consistent with Xin Tong's philosophy.

In 2016, Xin Tong proposed to solve the "everyone and everywhere" problem of graphics content production, so that anyone anywhere can create visual media content.

Ten years later, "everyone creates visual media content everywhere" has evolved into "allowing anyone to turn their imagination into an explorable and interactive world with just one sentence".

This is of course a very ambitious vision.

To achieve this goal, Meshy is currently carrying out basic technical research work at three levels:

First, and the most important — to reinvent computer graphics.

To this end, it is necessary to break away from the local optimal solution of the graphics rendering pipeline in the traditional network computer graphics, and use AI to re-create a set of global optimal solutions that do not rely on "triangles", where "triangles" refer to the smallest unit for GPUs to render 3D objects.

Secondly, to build a director system.

If the new computer graphics is more like a film crew, then a supporting director system must be built to realize AI for Fun.

This director system is responsible for formulating the operation mechanism of the world and the 3D skeleton, and realizing the real-time writing of scripts.

Finally, to build the supporting infrastructure.

A 3D generation and rendering infrastructure with low latency, high image quality and low cost is essential, because even a momentary lag will be magnified hundreds of times in the AI for Fun experience.

To this end, it is necessary to redesign a set of high-performance infra, which is in the same line as the foundation of the Taichi programming language that Yuanming Hu worked on in his doctoral dissertation and at the early stage of Meshy's founding.

In addition to these basic researches, Meshy's new models are also overcoming the specific bottlenecks in the current AI 3D field.

For example, the "controllability" problem, that is, whether the generated result is faithful to the input image, which is manifested in the accuracy of the overall proportion, the rationality of spatial distribution, and the restoration degree of surface details.

On August 10 this year, Meshy released Meshy-7, which has taken a big step forward in the direction of geometry alignment. Specifically:

In terms of organic characters, the model can restore details such as facial expressions, anatomical structures and skin wrinkles;

In terms of hard surfaces and mechanical parts, all components can be accurately located at the positions given in the input image, and adjacent components will not stick to each other;

In terms of text and patterns, the intaglio text and embossed patterns are clear and clean, and can be directly used in industrial scenarios such as laser engraving.

It can be said that Meshy's exploration in these directions is coupled with Xin Tong's research focus, and his recent research direction is at the intersection of 3D and video generation.

Just as the classic question he raised in 2024: "Is 3D just a special case of video generation?"

That is, if a video can already "see" an object from any angle, do we still need to explicitly construct its 3D model?

The answer is not yet clear.

But what is certain at the moment is that Xin Tong's cutting-edge research and engineering experience in AI multimodal training can take Meshy's exploration to a further level.

"Mora: Beyond the World Model"

When this article was being finalized, Yuanming Hu released Meshy team's new work, Mora, on his personal WeChat official account.

Mora is the abbreviation of "Multimodal Open-world Real-time Architecture".

This technology is quite interesting, and you can experience the interactive game demo generated by it.

In simple terms, it consists of three pieces of puzzle:

Coding agent, used to generate the game world skeleton and running code;

3D generation, which is the most important puzzle of Meshy, can be used to enrich the skeleton and output control signals to the video model;

Video model, which receives 3D scene signals and outputs the final images and sound effects.

Yuanming Hu calls this a technology that "goes beyond the world model".

It's a bold claim, what exactly does it transcend?

Remember in January this year, Google Genie 3 launched a playable demo, claiming that "it is a general world model that allows users to generate explorable real environments using text".

For this reason, Unity's stock price fell by 24% in one day, and Roblox fell by 27%. The market was extremely excited: In the future, there will be no need for game engines to make games, one video model will be enough.

However, more than half a year has passed, after Yuanming Hu tried all the world model demos on the market, he pointed out eight major drawbacks of almost all current interactive video models.

These