HomeArticle

A top-performing female student from Fudan University has raised 10 billion yuan in financing.

36氪的朋友们2026-07-22 16:52
Token factories that are being built across the world.

A Fudan University alumna majoring in Computer Science founded a company in the United States, which Jensen Huang called "the TSMC of AI factories." Top-tier VCs including HSG, NVIDIA, Index, TCV, and Lightspeed have placed consecutive bets on it. Only four years after its establishment, its valuation has reached 17.5 billion US dollars (approximately 1.185 trillion yuan).

This company is Fireworks AI. In today's popular terms, Fireworks AI is a Token factory. It neither trains cutting-edge large models nor develops consumer-facing AI applications, but focuses solely on inference, helping enterprises fine-tune and host open-source models, and then charges based on usage volume. At this year's GTC conference, Jensen Huang had a conversation with Lin Qiao, during which he stated bluntly: "In a lot of ways, you're the TSMC of AI factories."

The company has just completed a Series D round of 1.505 billion US dollars (approximately 10.2 billion yuan), with a post-money valuation of 17.5 billion US dollars. This round was led by Atreides Management, Index Ventures and TCV, with NVIDIA participating as a follow-on investor, while existing shareholders including Lightspeed, Bessemer, Menlo, HSG, and Benchmark Capital continued to increase their stakes.

In 2022, Lin Qiao, the Fudan alumna, left Meta and assembled a seven-person team to found Fireworks AI in Redwood City, California, one month before ChatGPT was launched.

In 2023, HSG, NVIDIA, and AMD invested in its Series A round; in 2024, its valuation reached approximately 552 million US dollars after Series B; in October 2025, it secured 250 million US dollars in Series C, pushing its valuation to 40 billion US dollars. Now, it has arrived at this Series D round with a 17.5 billion US dollar valuation.

In less than two years, its valuation has surged from 552 million to 17.5 billion US dollars, representing an increase of over 30 times. Its ARR has exceeded 1 billion US dollars, 5 times that of the same period last year. The number of tokens processed by the platform daily has risen from 15 trillion to over 40 trillion. More critically, over 95% of the tokens on the platform come from models customized by clients using their own data, rather than off-the-shelf general-purpose models.

With the open-sourcing of Kimi K3 and the upcoming launch of the full-powered version of DeepSeek V4, the valuation and barriers of foundational models have been impacted. The inference layer is the part of the AI technology stack that is being quietly revalued by capital, with funding shifting from "who has the strongest model" to "who can run models stably and cost-effectively."

Founder Lin Qiao is also a figure who is easily overlooked. Having worked her way up from the underlying framework PyTorch to building a company valued at 17.5 billion US dollars, her public profile is far less prominent than many founders of cutting-edge models. What's more notable is that this business is being pursued globally. In the US there are players like Together and Baseten, while in China there are firms including Silicon Flow, Wujing Xinqiong, Juzhen Data Science, and Qujing Tech. Putting the two groups of players side by side makes it possible to clearly assess whether the "Token factory" is truly a viable business.

The TSMC of AI

Essentially, Fireworks' product is a dedicated inference cloud platform. Through an OpenAI-compatible API, developers can access more than 200 open-source models on the platform, covering text, image, and multimodal scenarios.

Enterprise clients can perform fine-tuning on their own proprietary data, supporting SFT (Supervised Fine-Tuning), DPO (Direct Preference Optimization), and RFT (Reinforcement Fine-Tuning). The platform also allows attaching up to 100 LoRA adapters simultaneously without additional charges. On the compliance front, it has obtained enterprise-level certifications including SOC2, HIPAA, and GDPR. In March 2026, it further integrated with Microsoft's Foundry, enabling clients on Azure to directly use its inference services.

Speed is the core value proposition that Fireworks emphasizes to the public. It has independently developed a CUDA kernel stack called Fire Attention, which can achieve 167 tokens per second on DeepSeekV4 Pro, claiming to be 5 times faster than competitors at the same price point. Internally, it adopts decoupled services, semantic caching, and speculative decoding to control long-tail latency.

This engineering capability is the fundamental difference between Fireworks and "simply renting a few GPUs to run open-source models." What it sells is not just computing power, but the software efficiency that maximizes throughput from the available computing resources.

Currently, the starting price of Fireworks' serverless inference service is approximately 0.2 US dollars per million tokens (for 8B-level models) to 0.9 US dollars (for 70B-level models). It also offers GPU rentals on an hourly basis. A significant signal of change is that 95% of the tokens processed on the platform now come from customized models.

In other words, to effectively leverage AI, basic large models are far from sufficient to meet enterprise demands. Companies must have their own dedicated models trained on their unique business data. As Fireworks stated plainly in its financing announcement, enterprises all possess unique assets that others do not: customer relationships, workflows, data, and their own definitions of quality. What Fireworks does is to turn these assets into proprietary intelligence that clients can own and continuously improve.

Fireworks' clients include well-known companies such as Cursor, Notion, Quora, Sourcegraph, Uber, Shopify, and Perplexity. There is also a company named Sentient that relied on Fireworks to support 1.8 million users in the waiting queue 24/7. A pivotal detail is that while roughly half of Fireworks' revenue previously came from the AI programming tool Cursor, its client base has now expanded to a wider range of enterprises, which aligns with its strategic shift from "early adopter use cases" to "production-level penetration."

One of Fireworks' investors is Atreides, whose leader Gavin Baker made early significant investments in Meta and NVIDIA. Gavin Baker believes that cutting-edge models and open-source models will become increasingly complementary, and Fireworks is one of the few trusted platforms for enterprises to train and host open-source models in production environments.

Index is an existing shareholder that has participated in multiple rounds of financing, while TCV has invested in infrastructure-type businesses such as Netflix and Spotify. Coupled with NVIDIA's role as "both an investor and a supplier," the entire shareholder base is almost entirely composed of long-term capital focused on infrastructure and enterprise software. These investors are betting on foundational models on one hand, and using Fireworks' enterprise-level inference services as a hedge on the other.

With this round of financing, Fireworks' primary goal is to expand its computing power. Currently, Fireworks procures GPUs from more than 20 suppliers, including Microsoft. In addition, Fireworks plans to expand its engineering team from approximately 200 people to 600 by the end of the year, and deepen its cooperation with Microsoft and NVIDIA. In short, it aims to elevate the scale and reliability of model operation, as well as the scale of client services, to a new level.

Leader Lin Qiao

Lin Qiao grew up in China. Her father was a senior mechanical engineer at a shipyard who built cargo ships from scratch. As a child, she would study the precise angles and dimensions on ship blueprints, fascinated by the feeling of breaking down complex objects into calculable components. From junior high school, she focused on mathematics, physics, and chemistry. In high school, she wrote her first Snake game in BASIC, and at that point she knew computer science would be her future.

In 1995, Lin Qiao was admitted to Fudan University to study computer science, completing both her bachelor's and master's degrees there, where she was first introduced to the world of AI. Later, she went to the United States and earned her PhD in computer science at the University of California, Santa Barbara, with a focus on distributed systems and database management.

After obtaining her doctorate, she first joined IBM as a researcher at the Silicon Valley Lab and Almaden Research Center, leading the development of an ultra-high-speed data analysis optimizer. She then spent four years at LinkedIn working on distributed data services and participating in the Gobblin data ingestion framework.

A turning point in Lin Qiao's career came in July 2015, when she joined Meta and stayed for seven years, rising from Senior Manager of Data Infrastructure to Senior Director of Engineering, leading a team that grew from 5 to over 300 people.

At Meta, she accomplished something that influenced the entire industry: she co-founded and led the development of PyTorch.

She initially thought it was a "six-month project," but it turned into a five-year underlying reconstruction that rewrote all of Meta's AI workloads, from data loading and distributed inference to training across global data centers and billions of devices. By the time she left, the system was supporting over 5 trillion inferences per day. She later made an analogy: at Meta she built the engine, and now she wants to build the roads.

The timing of Lin Qiao's departure from Meta was very strategic. In October 2022, one month before ChatGPT was launched, she left Meta to found Fireworks.

At that time, her judgment was that enterprises wanted to adopt AI, but lacked the infrastructure, resources, and talent to deploy models into production environments — exactly the problem she had spent five years solving at Meta. Lin Qiao's vision was to compress the 5-year process Meta took into just five weeks or even five days. Six of the seven founding team members came from Meta's PyTorch team, and one from Google's Vertex AI; three of them are Chinese, including Lin Qiao, Benny Chen (UCLA alumnus and former head of Meta's advertising infrastructure), and Chenyu Zhao (UC Berkeley alumnus and former technical manager of Google Vertex AI).

Initially, the concept she promoted publicly was called "compound AI," which advocates using hundreds of small expert models to solve narrow problems respectively, rather than focusing on a single large model. Applied to Fireworks, this means helping enterprises turn open-source models into their own proprietary intelligence.

"There are two paths for AI. One path is that intelligence belongs to a few large labs, and everyone else rents access. The other path is that every company in the world builds its own proprietary intelligence, shaped by the domain that only it understands. We are building the second path."

According to CNBC, Fireworks is 5 to 10 times cheaper than closed-source models of equivalent quality. This is precisely the engine of its growth: when the bills for the latest models make business owners and CFOs uneasy, and when the capability gap between open-source and closed-source models narrows to a certain extent, enterprises will start to seriously consider open-source alternatives.

Token Factories Being Built Globally

Currently, many companies are claiming to be "Token factories," which refers to packaging underlying computing power, various open-source large models, and inference engines into standardized APIs. Developers do not need to purchase GPUs or set up environments themselves; they can use AI by writing a few lines of code and paying on demand. Whoever controls this service layer can collect the tolls.

This business is currently booming on both sides of the Pacific.

In the US, Fireworks is not an isolated case. Together AI just closed an 800 million US dollar Series C round on July 1, with a valuation of 8.3 billion US dollars. This round was led by Aramco Ventures (a subsidiary of Saudi Aramco), with NVIDIA and Salesforce participating as follow-on investors. Its annual booking amount is reportedly over 1.15 billion US dollars, and it accelerates inference through an adaptive speculative decoding engine called ATLAS.

Baseten is even more aggressive: on June 23, it announced a 1.5 billion US dollar Series F round with a valuation of 13 billion US dollars. It has completed four rounds of financing in 18 months, with revenue increasing by about 20 times year-on-year. Its client list even includes Open Evidence, Harvey, Cursor, and Notion, and NVIDIA is also one of its shareholders.

Other players in the space include Anyscale, Lepton founded by Yangqing Jia, Modal, Replicate, Groq, and Cerebras. Their slogans are largely similar: "own your intelligence."

In China, the narrative is different, but the core focuses on heterogeneous architectures. Simply put, it is about how to combine NVIDIA chips and domestic chips to make model inference more efficient.

For example, Silicon Flow, founded by Yuan Jinhui (a Tsinghua PhD who previously created the deep learning framework OneFlow), restarted its business in August 2023 after going through the twists and turns of being acquired and then spun off.

This company has submitted its listing application to the Hong Kong Stock Exchange this year, aiming to become the "first AI Token factory stock." In less than three years since its establishment, it has completed seven rounds of financing, with a post-money valuation of 7.74 billion yuan (approximately 1.1 billion US dollars). One of its Series B rounds, exceeding 2 billion yuan, set a record for the largest single financing in China's third-party MaaS sector. Alibaba is its largest institutional shareholder, holding 7.42% of the shares, while Hubble Technology (a subsidiary of Huawei) holds 4.07%.

Silicon Flow's cross-chip and multi-model adaptation capabilities are reportedly the best in the world. Its platform has cumulatively supported more than 170 mainstream models. When DeepSeek became a hit in 2025 and its official website was overwhelmed with traffic, it was the first company in the industry to successfully deploy the full-powered version of DeepSeek on domestic Ascend chips.

Another example is Wujing Xinqiong, initiated by Professor Wang Yu from the Department of Electronic Engineering of Tsinghua University, with Xia Lixue serving as CEO. It recently secured over 700 million yuan in financing, with state-owned capital and industrial capital participating. Its Agentic MaaS platform has launched more than 160 models, with the daily average token usage increasing by over 20 times compared to the end of 2025. The company has proposed a formula: "AI Productivity = Intelligence Scale × Token Production Efficiency × Token Value Conversion", positioning itself as a refinery in the petrochemical industry chain that converts energy into digital oil.

Established players such as Juzhen Data Science, founded by Fang Lei (a PhD from the Department of Electronic Engineering of Tsinghua University), has completed multiple rounds of strategic financing led jointly by two major state-owned AI industrial funds in Beijing, with state-owned capital and leading industrial capital continuing to increase their stakes. It recently held a launch event for its "Token factory" initiative, and its new generation Alaya NeW Intelligent Computing Cloud AI factory system integrates over 1000 large models and industry-specific models. The platform's target daily average Token processing capacity is 10 trillion, representing a dozens-of-times increase compared to the end of 2025. It pioneered the DCU "one unit of computing power" measurement unit, enabling computing power to be traded, priced, and accessed like electricity.

In addition, the biggest difference between the US and Chinese players lies in their positioning. Capital in the US is more aggressive with higher valuations, and clients are global enterprise software companies, with NVIDIA directly investing in many of these firms. In China, the fragmented chip ecosystem has created space for these intermediate-layer Token factories. However, another challenge is that major cloud vendors such as Volcano Engine, Alibaba Cloud, and Baidu Cloud are already in this space, and model providers like DeepSeek may also develop their own inference services.

What the future holds depends on how the industry evolves, but for now, the demand is enormous.

Since the explosion of agent applications at the beginning of the year, the number of tokens per single request has surged. The quality of open-source models is approaching that of closed-source models, providing enterprises with viable alternatives. Moreover, enterprises have finally started to carefully calculate their costs: training is a one-time expense, while inference is a continuous, draining cost — the more popular the product and the more frequent the usage, the greater the potential losses.

For example, some investors have mentioned that manufacturers of AI glasses and AI toys would prefer users to use fewer AI features and treat the devices as ornaments, because "Tokens are simply too costly to afford."

As a result, "whether the cost of each inference can be reduced" has become a matter of survival, and the Token factories positioned between models, GPUs, cloud vendors, and applications are exactly the ones that can solve this problem.

However, for Token-related businesses, gross margin is of utmost importance. Computing power