Kong Lingpeng, a professor at the University of Hong Kong, founded his own startup and secured the world's largest round of financing in the dLLM model sector.
On October 9, exclusive information from ChinaVenture reveals that DiffuSpace (hereinafter referred to as "DiffuSpace"), a diffusion AI startup, has successfully closed two consecutive rounds of financing, co-led by Matrix Partners China, Shunwei Capital and Legend Capital, with follow-on investments from CAS Star, Huawei Hubble, Horizon Robotics and other institutions. The total financing amount reaches several hundred million RMB, setting the new global record for the largest financing scale in the diffusion large language model (dLLM) sector.
According to public data, the Dream 7B developed by the DiffuSpace team is at the leading position among global diffusion large language models, with partial capabilities comparable to DeepSeek V3 (671B). In relevant speeches at ICML, Dream 7B, Google Gemini Diffusion and Inception Mercury are listed as three representative diffusion large language models.
It is understood that DiffuSpace is currently training a new generation of diffusion large language models with larger parameter scale, and plans to release and open source it in the near future.
DiffuSpace was jointly founded in Shenzhen in May 2026 by Professor KONG Lingpeng, co-director of the NLP Lab of the University of Hong Kong, together with GONG Shansan and YE Jiacheng, PhD candidates at the University of Hong Kong, and other team members. As early as 2022, the core team of the company began to conduct research on diffusion large language models, making it one of the world's earliest teams to systematically explore this technical route.
From the first day of its establishment, DiffuSpace clearly chose a technical path different from the mainstream autoregressive models such as GPT and Claude, that is, the diffusion large language model. The generation method of autoregressive models is like "dictating word by word", generating Tokens one by one from left to right; the diffusion large language model processes the whole first, then gradually completes the details, can generate content at multiple positions at the same time, and repeatedly corrects the existing content.
Since 2026, the development of diffusion large language models has been accelerating: in the first 8 months, the number of relevant papers has reached 2.4 times that of the whole year of 2025; Google further launched the open model DiffusionGemma on the basis of Gemini Diffusion, Inception continued to iterate the Mercury series of models, and Ant Group also launched LLaDA 2.2.
Backed by Huawei and Horizon Robotics, a new large model player emerges in Shenzhen
Essentially, what DiffuSpace is doing is challenging the underlying generation method that has been used in the large model industry for many years.
This is not a random idea, but a choice after long-term research. When mainstream large language models have widely adopted the autoregressive architecture, this mode of generating one Token after another backwards also has unavoidable shortcomings.
The limitations of autoregressive models are not difficult to understand. First, autoregressive models can only generate from left to right. When encountering tasks that require global consideration, they tend to focus only on the current step and fall into local optimum. Second, generating each Token needs to wait for all the previous content to be written, which also slows down the overall generation speed of autoregressive models.
In the process of looking for solutions, GONG Shansan, who is interested in scientific research, got acquainted with KONG Lingpeng, which eventually laid the groundwork for the two to embark on the road of entrepreneurship in the future.
After graduating from Shanghai Jiao Tong University with bachelor's and master's degrees, GONG Shansan joined Shanghai AI Lab and joined the NLP team led by KONG Lingpeng, who once worked at DeepMind, and began to study diffusion large language models in 2022. After that, GONG Shansan became a student of KONG Lingpeng, and is now in the final year of his PhD study at the University of Hong Kong.
In their view, diffusion models have advantages that autoregressive models do not have: the generated content can be modified, more complete context can be referenced when generating any position at the same time, and computing resources can be allocated more to parts with higher difficulty.
KONG Lingpeng uses a simple example to explain the difference between the two training logics: "After the autoregressive model sees '1, 2', it only needs to predict the next number '3', and can only predict the subsequent single content one by one; the diffusion model is more like getting a sequence with multiple missing positions, filling in these gaps at the same time, and making mutual corrections between different positions."
At that time, the first problem facing the team was not how large the model could be, but whether the diffusion model could be truly applied to language. Images can be processed as continuous signals, but language is composed of discrete Tokens, and a sentence must also meet grammar, semantics and contextual relations at the same time.
The methods that have been proven effective in the image field cannot be directly copied, which also means that if you want to take another technical route different from the mainstream, you first need to prove that the diffusion large language model itself can work stably.
With the release of papers related to diffusion models, they verified the feasibility of the technical path of diffusion large language models for the first time. By 2023, the team had built a complete diffusion model training paradigm adapted to discrete language signals, laying a foundation for subsequent model scaling and iteration.
The Dream 7B released in April 2025 became a watershed on this route. With 7 billion parameters, it means that the team's research has moved from early experiments to the magnitude of large models. At that time, Dream 7B outperformed autoregressive models of the same parameter scale in tasks such as mathematics, code and planning, and some indicators were even comparable to DeepSeek V3 (671B).
With Dream 7B gaining global recognition, DiffuSpace came into being. Their idea is very clear: focus on diffusion large language models, and build DiffuSpace into an AI company that explores the next-generation basic model architecture after autoregressive models.
There are very few entrepreneurial players in this track. The core reason is that the track has extremely high technical thresholds and great implementation difficulties. The diffusion large language model is not a partial optimization on the existing large model, but involves a series of changes in the generation paradigm, training methods, reasoning systems and even the underlying infrastructure. To move from papers to large-scale model training, algorithms, data, training and reasoning capabilities need to be advanced simultaneously.
At present, the core members of DiffuSpace, including KONG Lingpeng and GONG Shansan, mainly come from well-known universities such as the University of Hong Kong, Peking University, Tsinghua University and front-line model companies. This team that has been deeply engaged in diffusion large language models since 2022 has built a complete diffusion large language model system covering "algorithm-data-training-reasoning".
Moving from the University of Hong Kong laboratory to entrepreneurship, DiffuSpace also set up its company in Shenzhen. In their view, Shenzhen's complete industrial chain of chips, automobiles, robots and intelligent terminals provides fertile industrial soil for the development of this model route.
Their strength has also attracted many investors. As a result, we can see that less than half a year after its establishment, DiffuSpace has successfully closed two consecutive rounds of financing, receiving hundreds of millions of RMB from Matrix Partners China, Shunwei Capital, Legend Capital, CAS Star, Huawei Hubble, Horizon Robotics and other institutions.
Exploring the boundary of AGI, dLLM enters the scaling era
From the release of Dream 7B, to the exploration of application scenarios, and then to the successive completion of two rounds of financing, DiffuSpace's technology iteration and capital rhythm advance almost simultaneously, and the speed is also accelerating.
After its release and open source, the cumulative download volume of the Dream 7B model on HuggingFace has exceeded 2.5 million. The new model they plan to release within this year will further expand the parameter scale, and at the same time, it will be specially adapted to complex tasks such as code Agent and intelligent reasoning.
Promoting the model to a larger parameter scale is not simply pursuing a larger number. Models of different scales correspond to different deployment conditions: smaller models are more suitable for delay-sensitive and local deployment scenarios such as end-side devices and embodied intelligence, while larger models are oriented to office, code development and more complex scientific computing.
With the continuous expansion of the model parameter scale, DiffuSpace's goal is to find the most reasonable and efficient general intelligent generation logic, so as to break through the capability boundary of existing large models.
However, whether this goal can be achieved, the model effect is a key indicator. The general evaluation criteria are: efficiency depends on the generation speed and TPS (number of Tokens generated per second), while accuracy depends on the precision and comprehensive performance of various downstream tasks.
"On the premise of ensuring high precision, diffusion large language models can achieve efficient generation." According to GONG Shansan, "The test results show that when comparing TPS under the conditions of the same graphics card, the same generation length, and a single user's single query, the reasoning speed of DiffuSpace's diffusion large language model is about 10 times that of mainstream autoregressive models."
Combined with the advantages of the model, code generation, embodied intelligence, autonomous driving and other directions have become the priority layout areas of DiffuSpace. The three types of tasks seem different, but they have one thing in common: the order of solving problems is not always naturally from left to right, and it is often necessary to consider longer-range constraints at the same time.
Code generation is the top priority. The real programming process is rarely written sequentially from the first line to the last line. Developers will first build the structure, then fill in the details, and repeatedly modify between different positions. The global generation and parallel generation capabilities of the diffusion large language model can just process code in a more flexible order.
The second is embodied intelligence, where the diffusion large language model can globally control a complete set of action sequences; the last is autonomous driving, a scenario most sensitive to latency, where the fast generation advantage of the diffusion large language model can be directly applied.
In September 2026, DiffuSpace reached a cooperation with Acrab, an Asian agent computing platform company. The two parties carried out system-level collaborative adaptation for end-side Agents. Tests show that under specific conditions, the operation speed of end-side Agents can be increased by 5 times through the parallel generation and global modeling capabilities of dLLM.
For DiffuSpace, its own strength still needs further testing by the market: in the code direction, the model's ability to repeatedly modify between different positions needs to be stably applied in actual use; in the direction of embodied intelligence and autonomous driving, global planning and dynamic correction in longer action sequences also need to be further verified.
It cannot be ignored that autoregressive models are not standing still. Mainstream models are still raising their capability ceilings by expanding parameters, improving reasoning and training methods.
In response to this, DiffuSpace does not appear to be overly worried. In their view, every technology has a marginal effect. For tasks that are not suitable for strict left-to-right modeling, if autoregressive models want to achieve similar effects to diffusion large language models, they need to pay more computing and data costs.
"From the perspective of underlying principles, diffusion large language models are a generalized upgraded version of autoregressive models, which fully contain all the capabilities of autoregressive models. Diffusion large language models can be degenerated into autoregressive models through rule constraints. When the technical system and algorithm iteration of diffusion large language models are sufficiently complete, there is no need to retain the technical system of autoregressive models separately." GONG Shansan believes that there are most likely two situations in the future: one is that diffusion large language models fully replace autoregressive models and become the mainstream paradigm of the industry; the other is that the two types of models, relying on their respective advantages, coexist for a long time and adapt to different scenarios.
In the longer-term plan, DiffuSpace defines itself as an "intelligence producer". By continuously achieving AI technological breakthroughs, exploring the boundary of general artificial intelligence, and empowering various industries with top AI capabilities. "We hope to become a top, kind and technically aesthetic intelligence production enterprise."
This article is from the WeChat official account "ChinaVenture", author: LU Zhigao, editor: WANG Qingwu, published with authorization from 36Kr.