DeepSeek V4.1 Flash is officially launched, Cui Tianyi joins the "Godfather Battle"
Just two days after DeepSeek tested the intermediate version of V4.1 Flash, V4.1 Flash was officially launched.
This MoE model with a total of 552B parameters adopts a brand new asymmetric Causal-Encoder-Decoder structure. It only has 8B input activation and 16B output activation, with inference cost far lower than models of the same size. It natively supports multi-modal visual understanding without the need for an external visual module.
Through KV Cache compression, its VRAM requirement is reduced to 1/4 of the previous generation, and SSD requirement to 1/8. Community real-world tests show its output speed reaches 300 to 500 tokens per second.
Generally speaking, smaller and faster models tend to have lower performance. However, V4.1 Flash is different: it outperforms V4 Pro across all benchmarks, with a GPQA Diamond score of 90.9.
What's more, its price has been cut. At 12:00 on September 10, the off-peak pricing for Flash is 0.02 yuan for cached hit input, 1 yuan for cache miss, and 4 yuan per million tokens for output, with prices doubled during peak hours.
The previous pricing was 0.05 yuan for cache hit, 1.5 yuan for cache miss, and 4.5 yuan per million tokens for output.
A day earlier, Cui Tianyi posted on social media that all future calls to V4 Pro will be redirected to the higher-performance, faster V4.1 Flash, and billed at Flash's price points.
The original price of Pro was 3 times that of Flash. With this change, I can now add two more chicken legs to my takeout order. Looks like I've got another cyber benefactor to recognize.
I'm always a big spender at the start of the month, only to realize by the end that I'll probably be living on instant noodles until the next payday.
Model credit works the same way. On the first day of your subscription, you dump an entire large file into the context window, even though you only need 20 lines of logic from the file, you let the agent read 2000 lines of source code thoroughly. Then it reads through the entire repository's indexes, build artifacts and lock files. A problem that could be solved with a few hundred tokens ends up costing tens of thousands of tokens.
No way around it, having enough credit makes you reckless.
But after a few days, your credit is almost used up. You want to buy a higher-tier subscription, but you don't have enough money, and you can't finish your work without AI.
Just as you're stuck in a dilemma, a stroke of good luck comes: you get a chance to reset your credit quota.
That's why internet users call these people who distribute extra credit "cyber benefactors".
The most well-known ones include Tibo from OpenAI, Logan from Google, and now Cui Tianyi has also joined the ranks.
The "cyber benefactor" economy is becoming a key competitive edge for major AI companies. Especially when model performances are almost on par, credit quotas have become the gentlest weapon in the hands of AI vendors.
What's the point of having better model parameters when my cyber benefactor is better than yours!
Origin of Cyber Benefactors
In the large model circle, the concept of "cyber benefactor" originated in the West, following the "benevolence distribution" path. They hand out small perks like quota resets from time to time. The master of this playbook is Tibo from OpenAI.
Tibo is nicknamed the God of Quota Resets, his real name is Thibault Sottiaux, a Belgian who graduated from the Catholic University of Leuven with a degree in applied mathematics.
He was an "atypical prodigy" from childhood, writing his first program at the age of 7. But Tibo says what he really wanted back then was a machine that could talk to him. Later he first joined Google Maps, then went to DeepMind to work on the research infrastructure behind AlphaGo, before joining OpenAI in 2024 to take over Codex, a project few people were optimistic about at the time.
What truly made him a legend was November 2025.
During that period, Codex suffered frequent system errors, and Tibo compensated affected users with extra credit. No one expected that this simple compensation move would turn into an internet ritual that continues to this day.
In April 2026, Codex's weekly active users exceeded 3 million. OpenAI announced that they would reset all user quotas for every 1 million new users added, up to 10 million total users. So resets happened at 5 million, 6 million, 7 million, 8 million users.
On July 15, the active users of Codex and ChatGPT Work reached 8 million. Tibo announced on social media that he would perform another full reset for all users, and incidentally "keep the 5-hour rate limit lifted". By the end of August, he mentioned in an interview that Codex now has more than 20 million users.
Codex has bugs? Reset the quota to make it right. Celebrate after fixing the bugs? Reset the quota again. Launch a new model? Definitely reset the quota.
Tibo was like Wang Sicong back in the day, giving out free iPhones to people on Weibo just based on his own mood.
The community even built a dedicated website called codex-reset.com to track how many times Tibo has reset the quota.
According to data from the site, Tibo has performed nearly 40 quota resets.
10 resets happened in the 12 days following the release of GPT-5.6. When OpenAI launched GPT-6 Astra, only authorized cybersecurity institutions and Plus users could access it due to cybersecurity restrictions. So Tibo "saved one reset per day" for those paid users who couldn't access Astra temporarily, and eventually gave everyone a full reset anyway.
Tibo joked about his schedule: rest, reset, rest, reset.
Tibo is also great at playing along with internet memes. When reports broke out that OpenAI was facing computing power shortages, he joked in the comment section: "I really should stop pressing that huge Codex reset button on my desk."
He even got a physical reset button made, with his famous line: "If I feel like resetting, I can reset whenever I want." Once a user called him "babe", he replied: "No need to call me babe, the reset will happen this afternoon."
These memes even spawned a website called TiboGPT. The site only has one line of text: "Did Tibo reset today?" and the answer is "Most likely!"
This Belgian even speaks Mandarin. On his resume under "Languages Spoken", he wrote: French, Dutch, English, Mandarin.
At first everyone thought it was a joke, but when a Chinese fan left a message on social media telling Tibo that he hadn't opened Claude Code ever since he started using Codex, Tibo replied directly in Chinese: "You're the best!"
While Claude was constantly banning Chinese users on its platform, Tibo replied to Chinese-speaking fans in Mandarin, making many domestic developers exclaim: This benefactor is totally worth recognizing.
On June 11, 2026, Codex launched the Banked Reset feature.
Just like a card wallet, OpenAI stores reset coupons directly in user accounts, which can be used any time within 30 days. One full redemption resets both the 5-hour window and the weekly window, and pushes the weekly reset date back by 7 days.
If Tibo is the "God of Quotas" at OpenAI, his Google counterpart is Logan Kilpatrick, the developer relations lead for AI Studio and Gemini API under DeepMind.
He used to be the head of developer relations at OpenAI, so he and Tibo basically swapped jobs between the two companies.
Earlier in his career, Logan worked as a machine learning engineer at Apple, served as an open source policy advisor for NASA, and helped build the developer community for the Julia language. When he switched jobs in 2024 at the age of 27, the media once described him as "Sam Altman's secret weapon".
At Google, foreign media directly called him the "face" of Gemini, and internal reviews at Google say "he does 90% of all the marketing work all by himself".
During his tenure, the free quota for AI Studio was so generous that developers couldn't believe it.
2.5 Pro can be used about 50 times per day for free, with a maximum context window of 1 million tokens, and a single call to 2.5 Pro costs roughly 1 cent.
In November 2025, Gemini 3 Pro was released. Feeling threatened by this model, Sam Altman sounded the first internal red alert at OpenAI, and the company put aside all other work to fully focus on model development to avoid being overtaken by Google.
Logan went even further: not only could developers use Gemini 3 Pro for free on AI Studio, he also raised the daily free quota to 250 calls at one point. Calculated at 1.5 cents per single call to Gemini 3 Pro, that's about 4 dollars per day, and 120 dollars per month. For reference, the monthly relief provided by the UNHCR is only 100 dollars.
Under Logan's oversight, the free quota for Gemini 3 Pro alone is worth more than 800 RMB per month.
If you think that's all, you're still underestimating the "free access ambassador" title. AI Studio is open to everyone, but it still has a small learning curve. What truly made Logan a legend is the 12 months of free Google AI Pro subscriptions he gave out to students. This even led to a flood of paid edu email registration services on Xianyu priced between 3 to 5 RMB.
According to US regional pricing, a one-year Google AI Pro subscription costs 240 dollars. No extra requirements, as long as you get an edu email somehow, you can claim this big gift from Logan for free.
Later, Logan made another move: new users who register and bind a payment card get 300 dollars in Google Cloud balance valid for 3 months. Even if you cancel the card the next day, you still get the credit.
On top of that, Logan revised Google's in-app direct payment process, removing the obstructive step that redirects users to Cloud Console to bind a credit card.
Just as there are cyber benefactors, there are also cyber benefactresses, such as Lydia Hallie from Anthropic.
Lydia is a Member of Technical Staff (MTS) on the Claude Code team, in charge of developer experience and education. She previously worked as a DX engineer at Vercel, and later served as Head of DX at Bun.
It all started at the end of March, when Anthropic was mocked across the internet for tightening its quota limits. A Max 20x user posted that he burned through his 5-hour quota in just 19 minutes.
Lydia replied to him: "You're not using it the right way."
On May 20, Lydia released the full free course *Claude Code*, which officially launched on Frontend Masters and master.dev, and can be accessed even without binding a credit card.
On September 5, she suddenly posted: "We've just reset the weekly quota for all Claude Max users. With Fable 5.1 coming soon and many of you about to start a long weekend, we want you to keep coding."
The common trait of Western cyber benefactors is that they can't change the official pricing of the models, but they find all kinds of creative ways to help users save money.
Cyber Benefactors in China
If the Western model follows the "benevolence distribution" path, the Chinese path is the "pricing disruption" path. No fancy rituals, no reset buttons, just slash the prices to rock bottom directly.
The typical representative is "Saint Liang" Liang Wenfeng.
Liang Wenfeng never gives out coupons, and doesn't even do community operations. Instead, he open-sources the model weights under the MIT license, and releases them to the market at a token price far lower than the industry average.
Liang Wenfeng's logic is rational and restrained: DeepSeek's pricing principle is to achieve payback in 10 months, corresponding to 6 times profit. Liang believes that under this target, open sourcing will not have any negative impact on the business model.
In April this year, DeepSeek cut the cache hit price of Flash to 0.02 RMB per million tokens, setting a new global record for the lowest large model price.
But on August 6, DeepSeek suddenly announced a full price hike. The new prices took effect on August 17: the peak-time cache hit price for V4 Pro rose 11 times, and the output price rose 3.5 times.
Overnight, Liang Wenfeng went from "Saint Liang" to "Uncle Liang", then to "Xiao Liang", "Convict Liang" and "Little Liang".
The call volume of DeepSeek on OpenCode also dropped accordingly: Flash's usage fell by 59.1% in two days, and Pro's usage fell by 55.7%.
In fact, back in April, DeepSeek officially announced that "after the mass launch of super nodes in the second half of the year, the price of V4 Pro will be significantly reduced".
However, when V4.1 Flash was released, everyone slapped their foreheads and realized: "Oh! We totally misunderstood Saint Liang!"
A day later, as mentioned earlier, Cui Tianyi, head of DSH, posted: "Given that the DS V4.1 Flash model comprehensively outperforms V4 Pro in all indicators including performance, cost, speed and total latency, it would be inappropriate to continue providing the worse-performing V4 Pro model to DS users at higher prices, slower speeds and more computing power consumption. After V4.1 Flash is officially launched and before V4.1 Pro goes online, all requests to the V4 Pro model will be routed to V4.1 Flash, and billed at Flash's price points."
My cyber benefactor never abandons me after all.
Another domestic representative of the pricing disruption path is Tang Jie, chief scientist of Zhipu AI, and CEO Zhang Peng.
While DeepSeek's models still cost money, Zhipu AI offers its models for free directly.
The first free large model API from Zhipu AI is GLM-4 Flash, launched in August 2024, which started the trend of "free Flash tier".
On January 20, 2026, GLM-4.7 Flash was released. The official announcement states that it is permanently free, with no token cap, 200K context window, around 30 QPS concurrency limit, and even free cache access.
In multiple "Free Access Guides" circulated in Chinese developer communities, GLM-4.7 Flash has always been the top choice.
On August 20, an anonymous model called Ox Alpha landed on OpenRouter.
With a 1 million token context window, native multi-modality, and 7 days of free access, its usage once surged to more than twice that of DeepSeek, topping the charts and ending DeepSeek's 56-day monopoly on OpenCode.
Since the model name contains "Ox", and the movie *Ox is Coming* was a huge hit at the time, the community named the model "Ox is Coming".
To find out what "Ox is Coming" really is, developers across the country played detective for six days, guessing from tokenizer features to model architectures, trying to figure out which company developed it.
On August 26, the mystery was solved. Zhipu AI posted that "Ox is Coming" is GLM-5.3-Flash, with 320B-A18B parameters, an Artificial Analysis intelligence score of 57, on par with Claude Opus 4.8, and its performance even exceeds DeepSeek V