China's data generation team has appeared in Nature for the first time.
A Hong Kong-based startup team has unexpectedly come into our sight.
Not long ago, a paper on AI-assisted renal cancer surgical decision-making was published in Nature Communications (https://www.nature.com/articles/s41467-026-73813-7). With this paper, Wiener Intelligence has become the first data generation sci-tech innovation company in China and the fourth in the world to be featured in a major Nature journal (with an impact factor of over 10 in the past three years) — the previous Chinese large model companies that published papers in such journals are DeepSeek and MiniMax.
Behind the company is its founder Professor LIU Qifeng, who once built the world's first 1000-node H800 SuperPod cluster at the Hong Kong University of Science and Technology, pre-trained China's third large model with hundred-billion-level parameters, and managed project R&D funds exceeding 100 million USD. Later, he focused on how to improve the "question-asking" capability of AI to generate high-quality reasoning question-and-answer data, which is one of the key points of the upcoming boom of AI autonomous learning, so he founded Wiener Intelligence in Hong Kong.
Out of curiosity, the Investment Community had an in-depth conversation with LIU Qifeng for nearly three hours. The discussion started from this paper, went all the way to large models and embodied intelligence, and covered his understanding of the next stage of AI development.
Starting from a Paper Published in Nature Communications
Others "start from papers and end up in papers", but LIU Qifeng "starts from real problems and solves real problems".
In early 2025, a relative of LIU Qifeng was diagnosed with renal cancer, and the attending physician was Director ZHANG Zhiling from Sun Yat-sen University Cancer Center. Like all other renal cancer surgeries, doctors have long faced a clinical challenge — between partial nephrectomy and radical nephrectomy, is there a more quantitative and intelligent judgment basis available?
The essence of this challenge is: whether AI can predict complex choices in the real world.
Right in the hospital ward, a collaboration spanning medicine and AI was launched: ZHANG Zhiling was responsible for the medical work and completed data collection in cooperation with multiple hospitals, while Wiener Intelligence took charge of AI and data processing work. WANG Yatian, the co-first author of the paper, is a PhD student at the Hong Kong University of Science and Technology and an intern at Wiener Intelligence, co-supervised by LIU Qifeng and Professor LUO Wenhan.
To address the challenges of multi-source heterogeneous sparse data, the team proposed the RDPM model, which integrates 3D imaging and clinical variables/indicators into the same prediction framework. The model was trained and validated on a cohort of 1621 patients, with the AUC of external multi-center test reaching 0.788 to 0.873. The paper predicts the long-term renal failure risk of patients, providing quantifiable support for surgical decisions that highly rely on experience.
Wiener Intelligence has completed a public verification of AI prediction in the medical scenario, which is one of the most error-intolerant fields. This also points to the other side of AI prediction that LIU Qifeng will elaborate next — prediction is essentially the underlying mechanism of large models, which generates answers by predicting the next Token, and is naturally "good at answering". The focus of Wiener Intelligence goes a step further — to make AI not only good at answering, but also "good at asking". To endow AI with "knowledge and inquisitiveness", it must be able to "learn" effectively as well as "ask" properly.
Entrepreneurship of a Professor from HKUST
Lenovo Capital Led the First Round of Financing
"Others follow the trend and do whatever is popular. But he makes whatever he does become popular." This is the impression LIU Qifeng's friends have of his past experience.
This statement is no exaggeration. Back in 2001, LIU Qifeng joined the National Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences, studying under Academician TAN Tieniu — winner of the 2022 J.K. Aggarwal Prize, the highest award in the field of international pattern recognition. Later, he successively served as a researcher at Samsung Lab, data scientist at Yahoo! Lab, director of Gamma AI Lab of Ping An Group, AI director of the Hong Kong Institute of Chinese Academy of Sciences for Innovation and Technology, and other positions. In 2018, he co-founded the Hong Kong Artificial Intelligence and Robotics Society with Academician YANG Qiang. In 2021, he was highly forward-looking to write the proposal for "Hong Kong Cloud Brain" and "Hong Kong Foundational Large Model" for the Hong Kong SAR Government, becoming an early promoter of AI supercomputing construction and large model training in Hong Kong.
All these seemingly scattered experiences point to the same goal: to enable machines to identify patterns from complex information and make judgments accordingly.
The real turning point came in 2023 when ChatGPT became popular worldwide. At that time, with the strong support of the SAR government and school leaders, LIU Qifeng, together with 6 other universities in Hong Kong and Academician GUO Yike, co-founded the Hong Kong Generative AI R&D Center, led the team to build the world's first 1000-node H800 SuperPod AI supercomputing cluster. In 2024, he completed the pre-training and post-training of China's third hundred-billion-parameter MoE large model. This is a key milestone for the development of AI in Hong Kong.
It was precisely during this experience that he spotted the next gap: the further large models develop, the more they rely on high-quality data, and "data is king" always holds true, especially the reasoning question-and-answer data covering all industries. Therefore, enabling large models to "ask questions" with high quality has become the top priority.
LIU Qifeng divided the development of large models into three stages: the first stage is "from data to model", using massive internet data for pre-training; the second stage is "from model to Token", where large models start to output Tokens to generate content or execute tasks; the next stage is "from Token to data" — enabling the large model system to actively ask questions, think step by step, and verify answers, that is, reasoning question-and-answer data generation. In this way, a large feedback closed loop of "data → model → Token → data" is formed, so that AI can acquire autonomous learning capability.
The purpose of AI autonomous learning is to acquire "knowledge and inquisitiveness", which combines "learning" training and "question-asking" practice. The Qing Dynasty scholar LIU Kai wrote in *On Inquiry*: "A man of virtue in learning must be fond of asking. Inquiry and learning complement each other. No learning can lead to no doubts to raise, no inquiry can lead to no knowledge to expand."
In July 2024, Hong Kong Wiener Intelligence was officially established. The company is named after Norbert Wiener, the founder of cybernetics. What LIU Qifeng values most is exactly the feedback closed loop in cybernetics. The mission of Wiener Intelligence is to enable AI to "ask" accurately and "answer" correctly, so as to realize the large closed loop of "data → model → Token → data", and further enable Agentic AI to evolve autonomously in professional fields.
Wiener Intelligence aims to solve a counter-intuitive problem: on the one hand, the development of large models is advancing by leaps and bounds, on the other hand, the implementation of large models in enterprises is still very difficult. The reason is very simple: the accuracy is low. Using students preparing for exams as an analogy — with only textbooks (professional documents) and no exercise sets (reasoning question-and-answer data), the exam results can never be good (the system accuracy is low). Because memorizing textbooks only brings rigid knowledge, while doing exercise sets trains flexible problem-solving capabilities. What Wiener Intelligence is doing is to help all industries complete this "exercise set", so that AI can not only "study textbooks" but also "do exercises", thus solving the bottleneck problems of inaccurate measurement, difficult optimization and incorrect answers faced by the current overflow of Agents.
There is a popular view nowadays: large model Q&A is outdated, and task execution is the real key. This view is rather superficial. Execution capability depends on two pillars: the accuracy of a single agent in professional fields, and the collaboration capability between multiple agents. But in reality, the current execution capability is far from reliable. One of the core problems is that the single-agent Q&A accuracy is often less than 70% — it can't even reach the threshold of "reliability", let alone achieve "collaboration".
The specific definition of "exercise" is cQrA: context, Question, reasoning, Answer. Context refers to the task scenario, Question refers to the generated question, reasoning refers to the reasoning process, and Answer refers to the verified answer. In other words, Wiener Intelligence enables the model to generate questions, answers and reasoning processes simultaneously in the specific context of a certain industry.
This also makes it completely different from traditional data labeling. Traditional data labeling relies heavily on manual work even experts, with high cost and difficulty in scaling, and only provides answers without reasoning, which wastes expert experience in repetitive work. Wiener Intelligence enables Agentic AI to act as a tireless team of intelligent experts, automatically generating cQrA data with a complete chain of thought, completely breaking through the manpower bottleneck. More critical than cost saving is that the closed-loop mechanism enables the data generated in each round to feed back to the generation and evaluation models, driving the next round of iteration to continuously improve in accuracy and logic — thus realizing the qualitative change from a "manual workshop" to a "self-evolving knowledge factory".
Wiener Intelligence soon caught the attention of the industry and investors. Shortly after its establishment, the company completed a seed round financing of 50 million Hong Kong dollars, led by Lenovo Capital. Lenovo Capital has long been investing in the three core tracks of AI: it has invested in computing power companies such as Muxi and Cambricon, model companies such as Zhipu AI and Stepfun, and the data track investment falls on Wiener Intelligence. Meanwhile, Muxi and Wiener Intelligence have reached in-depth cooperation. In the upcoming era of the large "data → model → Token → data" closed loop, one has designed a computing power platform for the future paradigm in advance, and the other has defined the workload for the future paradigm in advance.
The Next Stage of AI Development
"Let's Generate the World"
Commercial verification starts with two core questions.
Question 1: Without large-scale expert labeling, will professional institutions pay for the generated data?
Question 2: Can the solution be cross-industry and replicable?
To answer these two questions, Wiener Intelligence withstood the pressure and broke the "depth-first" principle of traditional 2B sci-tech innovation companies, which requires enterprises to first "penetrate a certain industry". Instead, it adopted the "breadth-first" principle, deliberately selecting four seemingly unrelated industries that all have high requirements for accuracy: value safety, government affairs, insurance, and horse racing, and has secured leading clients in each of the four industries.
"We have proved that we can achieve cross-industry replication with a small team, no industry experts, and low cost." LIU Qifeng said, now that we have completed the verification from "0 to 4", the next step is to promote the model from "1 to M x N" (M industries, with N leading clients in each industry).
Behind this is the long-term judgment on the value of data.
In LIU Qifeng's view, the gap between AI development in China and the United States largely stems from the cognitive difference on data — data has long been regarded as "tedious and unglamorous work", and the salary of data engineers is generally lower than that of algorithm and model engineers.
But the situation is changing. As data production evolves from manual labeling to reasoning, interaction and closed-loop feedback, large model companies are continuously increasing their investment in the data field. It has become a consensus in the industry that reasoning interactive data generation determines the upper limit of large model capabilities.
What is the most critical in the future is not the model, not even the data itself, but the "large closed loop". Just like the key to evolution is neither men nor women, but mating and natural selection — the mechanism of chromosome replication, crossover, mutation and survival of the fittest. Data distillation is only one path that leverages external resources. The real moat is to build an autonomous learning closed loop where model training and data generation drive each other, and continuously generate high-quality data through model collaboration and feedback mechanisms.
This judgment also extends to the currently hottest field of embodied intelligence.
The traditional way to train embodied intelligence is based on imitating human behaviors. But real intelligence should be like toddlers learning to walk by "tumbling around": autonomously generating motion data in continuous falls and attempts, and then iteratively optimizing the decision-making model through closed-loop feedback.
LIU Qifeng said that the closed-loop training logic in the digital world has extended to the physical world. Whether it is Agents entering industries or robots moving to real work scenarios, massive high-quality reasoning interactive data generated autonomously in advance is indispensable. Correspondingly, cQrA evolves into cTrA — context, Task, reasoning, Action, which is not only the new fuel for training, but also the new ruler for evaluation.
The journey has just begun. LIU Qifeng's answer points to the near future: "Let's generate the world!"
This article is from the WeChat official account "Investment Community" (ID: pedaily2012), author: WANG Lu, published with authorization from 36Kr.