Microsoft has halted Tokenmaxxing, with budgets strictly capped, and any overspending beyond the set limit shall be borne by the responsible party itself.
Wait, even Microsoft is starting to care about token costs??
According to internal emails, starting from July this year, Microsoft has set "AI Token Budget Targets" for all its business divisions.
It has also designated GPT-5.6 as the default internal model, for the reason that it is more affordable.
Microsoft, you clearly haven't seen the mysterious power from the East.
If you're chasing low costs, why not just use DeepSeek directly…
Jay Parikh, Executive Vice President of Microsoft, also personally stated internally:
Tokenmaxxing is not our optimization goal.
At Microsoft, AI Usage Also Needs to Be Tightened to Cut Costs
According to Parikh's internal email and the updated Copilot usage guidelines, Microsoft's new policy can be broken down into several points:
First, set clear budgets. Starting from July 2026, all business divisions of Microsoft have their own "AI token budget targets".
However, the specific figures for the spending targets have not been announced yet.
Second, full visibility into spending. Employees can view their personal AI token consumption.
Internal data shows that many engineers spend hundreds to thousands of US dollars on tokens every month.
Third, switch to a new default model. The internal default model has been changed from the previous options to GPT-5.6, as it is cheaper than other models.
That's not all. Microsoft stated that further restrictions may be imposed later based on monitoring results.
Parikh wrote in the internal email, "Manage token spending just like we manage all other critical resources."
He required employees to "be fully aware of how they consume tokens" while using GitHub Copilot to accelerate business operations.
In fact, this move did not come out of nowhere. Back in June, Satya Nadella mentioned in a podcast episode that there was widespread Tokenmaxxing inside Microsoft.
He even admitted that he is a Tokenmaxxer himself, who is obsessed with burning tokens and finds the high numbers on the usage dashboard extremely addictive.
But as the head of the company, he still has to keep an eye on the books. So he came back to his senses: The marginal benefit of productivity improvement must match the marginal cost of tokens.
Not every problem deserves to call the most powerful and expensive cutting-edge model. Letting the model run longer, stuffing more content into the context window, or launching more Agents does not necessarily lead to better final outputs.
Two months later, this judgment turned from a CEO's podcast remark into an internal policy of Microsoft.
However, Microsoft employees are very dissatisfied with this new rule. One employee commented:
A company that has poured huge sums of money into AI and subsidized massive inference workloads is now actually guiding its own employees to save money??
This employee raised a very sharp question: If even a company that hosts AI infrastructure needs to worry about whether it can afford its own internal AI products, what should ordinary enterprises that purchase these services do?
Tokenmaxxing: No One Can Keep Up the Frenzy Anymore…
The concept of Tokenmaxxing has only really become popular in the past few months.
The term combines "token" and "maxxing" (pushing usage to the extreme), following the same internet neologism logic as gymmaxxing (working out to the extreme) and looksmaxxing (modifying one's appearance to the extreme).
However, it does not describe a type of self-improvement, but a distorted incentive mechanism that treats AI usage as an internal KPI for enterprises.
Back in March this year, employees at companies including Meta and OpenAI had already started competing for token consumption on internal leaderboards.
Some companies even regarded higher token budgets as a welfare for engineers, hoping that employees could run multiple Agents at the same time and assign more work to AI.
By May, this trend had spread to Amazon.
Amazon launched an internal leaderboard called KiroRank, which scores employees based on their AI usage on Kiro, Amazon's in-house AI development platform.
In fact, the original intention of the company was probably to encourage engineers to use AI tools more frequently, as more usage would lead to a higher ranking.
As a result, some employees started "gaming the leaderboard": they made AI Agents perform unnecessary or low-value tasks just to boost their usage volume. A large amount of tokens were burned, but the actual code output did not grow accordingly.
At the end of that month, the leaderboard was shut down. Senior Vice President Dave Treadwell specifically reminded employees:
Please do not use AI just for the sake of using AI.
But similar token usage spamming competitions have swept across Silicon Valley in just a few months:
One OpenAI employee once processed 210 billion tokens in a single week, enough to fill the entire Wikipedia 33 times.
Shopify and Meta also included AI usage in their performance assessments, rewarding employees who use AI heavily and criticizing those who do not.
There was also a self-built leaderboard for Meta employees called Claudeonomics, which tracked the AI usage of more than 85,000 employees across the company.
It not only listed the top 250 users with the highest token consumption, but also awarded virtual titles such as "Token Legend" and "Cache Wizard".
Wow, this token version of "League of Legends" has even appeared…
Jensen Huang also added fuel to this trend. Back in March this year, he publicly stated on the All-In Podcast:
If an engineer with an annual salary of 500,000 US dollars consumes less than 250,000 US dollars worth of tokens every year, I will be deeply worried.
Later, at GTC 2026, Jensen Huang added:
"I am willing to provide a token budget equivalent to half of an engineer's annual salary, on top of their six-figure annual pay, because these tokens have the potential to amplify the engineer's productivity by 10 times."
Of course, as a supplier of shovels in the AI boom, every extra token burned by engineers means extra demand for AI infrastructure.
Jensen Huang's public endorsement of Tokenmaxxing is obviously driven by his straightforward commercial stance.
But now major tech companies are looking back and realizing that something is wrong.
This is not sustainable: you get exactly what you incentivize. Incentivize token usage, and you end up with a token bubble.
An investigation released by 404 Media shows that Atlassian's monthly AI spending tripled in less than a year, exceeding 15 million US dollars. The company then started to set spending limits and required employees to control model usage costs.
Uber's situation is even more extreme: The entire company burned through its full-year 2026 AI programming budget as early as April, and the new quotation provided by Cursor to Uber rose to 4 to 5 times the previous level.
Uber then stipulated that each employee cannot spend more than 1,500 US dollars per month on each Agent programming tool.
From Amazon, Adobe, Atlassian to Citi, all companies are starting to tighten the previously nearly unlimited AI access for their employees. The common solution is simple: Stop using the most powerful model by default for simple tasks, and downgrade to a cheaper model whenever possible.
According to related sources, Citi once required employees to select models based on task difficulty, instead of using the most expensive version for all tasks; access to some high-cost models was temporarily suspended.
Adobe is also revoking employees' unlimited access to Claude.
However, cost is not the only problem left by Tokenmaxxing.
Many developers said that after relying on AI for a long time, they feel their programming skills are degrading.
Some people find it increasingly difficult to judge whether a piece of code is correct; others worry that when hundreds of engineers submit a large amount of AI-generated code at the same time, no one has the capacity to check all the security and quality issues line by line.
So now that bosses no longer focus on token consumption, what will they focus on instead?
The answer is: Actual work outcomes.
For example, Meta has listed "AI-driven impact" as a core work requirement for its employees starting from this year. It no longer emphasizes how many tokens are consumed, but whether employees can use AI to generate real, tangible impact.
Of course, this set of metrics may also become formalized or even abused. But compared to token leaderboards, it at least tries to answer a more fundamental question about work:
What exactly have you accomplished?
In this industry-wide shift, Microsoft is even considered one of the last batches of companies to catch up.
OMT
Speaking of which, if Silicon Valley wants to engage in Tokenmaxxing, they should just use DeepSeek directly!!
Why use GPT-5.6 or Fable 5, which are ridiculously expensive? No wonder Uber burned through its full-year budget in just 4 months…
A sincere recommendation to American programmers:
Just go for DeepSeek V4 Flash. You can use it as much as you want, and your company will never go bankrupt because of it.
References:
[1]https://www.404media.co/microsoft-tells-engineers-tokenmaxxing-is-not-what-we-are-optimizing-for/
[2]https://www.404media.co/the-tokenpocalypse-is-here-companies-are-scrambling-to-stop-spending-so-much-on-ai/
[3]https://www.404media.co/software-developers-say-ai-is-rotting-their-brains/
[4]https://www.404media.co/companies-are-throttling-employees-ai-use-because-its-too-expensive/
[5]https://www.nytimes.com/events/hardforklive
This article is from the WeChat public account "QbitAI", written by Yu Ting, and published with authorization from 36Kr.