Exclusive | Hu Han, Head of Hunyuan Multimodal Understanding, Resigned to Start His Own Business, The Original Team May Focus on World Models
Text by ZHOU Xinyu
Edited by ZHANG Yuxin
"Intelligent Emergence" has exclusively learned that recently, HU Han, who leads the multimodal understanding team of Tencent Hunyuan, has submitted his resignation.
Prior to this, he served as a principal researcher in the Visual Computing Group at Microsoft Research Asia. After joining Tencent in early 2025, he was responsible for the research of large vision models. In subsequent adjustments, he joined the "Frontier" advanced technology research group under the Large Language Model Department to lead research on multimodal understanding, reporting to YAO Shunyu.
It is understood that HU Han was also responsible for the research and development of world models.
At the same time, YAO Shunyu, the head of Tencent's Large Language Model Department, is currently comprehensively reorganizing his teams, the research group where HU Han previously worked may focus on cutting-edge research of world models.
As of press time, Tencent has not responded to the above information.
Large language models remain the top priority for Hunyuan
Since mid-2026, personnel adjustments around the Hunyuan system have been ongoing. In early July, TIAN Yonglong, former researcher at OpenAI and PhD from the Massachusetts Institute of Technology, joined the Large Language Model Department — we learned that he will replace HU Han to lead the R&D of Vision-Language Model (VLM) and report to YAO Shunyu.
Since YAO Shunyu joined Tencent, making up for shortcomings and rapidly elevating the capabilities of Hunyuan's foundational model to the first tier has become the core line of Tencent AI.
Reorganizing resource allocation and concentrating talents and computing power on the foundational model R&D under the Large Language Model Department is one of the core measures after YAO Shunyu took office.
For example, after Tencent disbanded AI Lab, all core R&D personnel were integrated into the Large Language Model Department. The recruitment of large language model talents has also been intensified with YAO Shunyu's arrival, including several members from ByteDance's Seed Infra team and post-training team.
Later, at the Tencent Cloud AI Industry Application Conference on June 5, TANG Daosheng, Senior Executive Vice President of Tencent Group and CEO of Cloud and Smart Industries Group, asked a sharp question for the public to YAO Shunyu to some extent: Many people say Tencent is lagging behind in AI, do you think we are really behind?
But in fact, "Shunyu is more anxious than Tencent's senior management," a researcher from Hunyuan told "Intelligent Emergence". He mentioned that YAO Shunyu has repeatedly fought for more computing resources for Hunyuan from Tencent's executives.
Another detail is that "LatePost" once reported that one month before the release of Hy3, a batch of data had problems. YAO Shunyu warned the team with rarely severe words: "Data is extremely important. If such a situation occurs again, the responsible person will leave directly."
Reassessing the returns of multimodal understanding
When none of the AI business lines have taken a leading position, Tencent is bound to reassess the returns of various AI technology directions to achieve a breakthrough.
On the one hand, the multimodal understanding research that HU Han is engaged in has become increasingly mature, and the benefits brought by continuing to invest resources are very limited.
"In multimodal understanding research, the recognition accuracy of text, images, and videos can basically reach more than 85%," a multimodal researcher told us, "The real difficulty lies in visual reasoning, which is to make the model understand what images and videos mean. This step cannot be achieved by multimodal research alone, but requires improving the reasoning ability of language models."
With the gradual fading of technological dividends, the commercial benefits that multimodal understanding can bring are also unclear.
Previously, multimodal research once occupied a core position in Tencent's AI system. The core reason is that computer vision technology can be strongly integrated with "cash cow" businesses such as advertising, games, and audio-visual content to generate revenue for Tencent.
However, compared with generative technology, scenarios corresponding to multimodal understanding such as "image recognition" and "describing images with text" cannot be directly commercialized.
"No user is willing to pay for image recognition, and there are plenty of free alternative products on the market," a product manager of Yuanbao mentioned. What users are really willing to pay for are office scenarios such as document processing, PPT production, and research report generation — which correspond to the model's capabilities of reasoning, coding, and agentic operation.
On the other hand, Tencent's computing resources are limited, and it does not have advantages compared with ByteDance and Alibaba.
Tencent's 2025 financial report shows that Tencent's annual capital expenditure was 79.2 billion yuan; in contrast, as of March 2026, Alibaba's fiscal year capital expenditure reached 126 billion yuan; and according to Bloomberg, ByteDance plans to push its AI infrastructure investment up to $700 billion in 2026.
With limited resources, it is inevitable that there is a shortage of supply within Hunyuan. At this point, suspending the multimodal understanding research with limited technological dividends and focusing on more future-oriented cutting-edge research directions is Tencent's choice after weighing the returns.
On July 6, YAO Shunyu delivered his first official achievement after joining Tencent, the large model Hy3. Among medium-sized models with the same parameter size, Hy3 can already compete with GLM-5.2 and DeepSeek V4 Pro.
On Tencent AI's own coordinate system, the organizational restructuring, reconstruction of the infrastructure system (data, Infra, evaluation), and the strategy of returning to foundational models since YAO Shunyu joined, have allowed Hunyuan to gain a seat at the AI Tier 1 table in just half a year.
(Author LI Zhaofeng of "Intelligent Emergence" also contributed to this article.)