HomeArticle

AI Weekly Observation 2026W36: Astra Expands the Boundaries of Agents, Google Unexpectedly Benefits from Tokenmaxxing

36氪的朋友们2026-09-07 20:25
Weekly Summary of AI Industry Dynamics, Research Progress and Policy Events

01. Industry

Astra did not deliver AGI, but expanded the capability boundary of Agents

On September 3, OpenAI released GPT-6 Astra in its official announcement, listing computer operation and specialized training for professional scenarios as key priorities. The tasks demonstrated by the company include updating customer management systems, organizing calendars, conducting online research and writing summaries in document editors, and operating Excel, Power BI and engineering software.

This expansion targets work scattered across different software. Taking the production of a market report as an example, an Agent needs to look up information in a browser, import data into a spreadsheet for calculation, then arrange charts and conclusions into a document. The computer operation capability allows it to switch between these applications, read interfaces, input content, execute actions, and proceed to the next step based on the results. Operations that users previously had to complete manually, such as downloading, filling out forms, copying and typesetting, are now within the scope that Agents can take over.

Official tests show that this path has yielded quantifiable progress. In the delayed simulation of OSWorld 2.0, Astra scored 72.6%, taking about 40 minutes per task; Sol scored 65.7%, taking about 75 minutes. OpenAI also updated the computer operation framework of Codex. When the model works with the framework, the task completion speed in the Mind2Web test is about 1.9 times that of the original Sol experience.

This also aligns with the recent boom in work trajectory data. Programming involves code repositories, operating environments and tests, making it relatively easy to turn a single piece of work into a task that can be practiced repeatedly. When expanding outward, operations in spreadsheets, documents, web pages and professional software are a space that can be easily implemented first: materials have been digitized, many intermediate states can be saved, and numbers and files can be verified. Model companies can first train these links with clear outputs, then try to connect them into complete workflows.

Cross-application execution further expands the use of such training. Learning to create spreadsheets and write documents independently can only take over a few links in the work; learning to call different software to complete consecutive steps gives the opportunity to take over a whole assignment. The market that model companies are targeting has accordingly expanded from programming tools for developers to software used daily in positions such as analysis, operations, and design.

Independent evaluations show that this expansion still has obvious shortcomings. In the published evaluation by Artificial Analysis, Astra improved by about 80 Elo points in AA-Briefcase, which contains multiple associated tasks and a large number of source files, with higher analysis quality but lower presentation quality; in GDPval-AA v2 covering 44 categories of professional tasks, its score dropped by about 80 Elo points instead. On September 4, after AA incorporated more complex work tasks and more proprietary questions into the composite index, Astra ranked second, 4 points higher than Sol.

Costs have not decreased along with execution efficiency. The unit prices for both input and output of Astra are 2.5 times the current promotional price of Sol. In the comprehensive intelligence test on September 3, AA measured that for the same max tier, the cost per task of Astra is 75% higher than that of Sol, about 1.75 times; in the programming Agent test, lower token consumption brought the total cost roughly on par with Sol.

In the Codex discussions on Reddit, the price increase and quota consumption caused dissatisfaction, while some users hoped that reduced rework could offset the price difference. Positive feedback has been received for partial capabilities: in Simon Willison's SVG actual measurement, the image quality of the low-tier Astra is better than all tiers of Sol he tested.

However, how much rework is still required after handing over a task directly affects whether users are willing to pay this premium.

NVIDIA acquires Hugging Face to complete the model distribution entry for AI cloud

On September 3, NVIDIA announced that it has agreed to acquire Hugging Face for approximately 12.93 billion U.S. dollars. According to the company's announcement, more than 18 million developers, researchers and creators use this platform, and more than 200,000 companies find, evaluate and deploy models on it. The transaction is pending completion.

Hugging Face gathers the developer demands that NVIDIA's cloud business needs. After developers find models here, they need to decide what tools to run them with and where to rent computing power. By acquiring this platform, NVIDIA gets the opportunity to access demands before developers select cloud service providers, connecting model selection with subsequent computing power procurement.

This path is supported by actual existing businesses. In June 2025, Hugging Face and NVIDIA launched a training cluster service: users submit demands on Hugging Face, and the two parties assist with matching, quotation and GPU cluster configuration, while DGX Cloud Lepton provides resource access and training scheduling. This July, NVIDIA launched a credit support and revenue sharing model, where cooperating cloud vendors sell cloud services based on NVIDIA hardware, and NVIDIA can share in the related cloud revenue in addition to selling equipment. The development demands gathered by Hugging Face can be directly connected to this batch of computing power.

Therefore, this acquisition also means that NVIDIA is extending to the client side of its cloud business. In the past, it mainly sold equipment to cloud vendors, who then competed for clients using computing power; now, while participating in computing power supply and financing, it is approaching the places where developers select models and launch projects. Computer rooms can be operated by partners, and NVIDIA gets the opportunity to participate in both front-end client acquisition and back-end usage revenue.

Stripe is also competing for model distribution entries. On August 19, Stripe announced the acquisition of this model routing platform, which can allocate requests between different model suppliers based on tasks, prices and speeds. Hugging Face influences what models developers discover and how they deploy them; OpenRouter is closer to actual invocation and settlement. The two transactions show that large companies are competing for the model distribution layer. Even if a certain model is replaced by a better one, users may still stay on the same platform and continue to select and purchase model services through it.

This also explains why Jensen Huang promised to continue supporting other chips and cloud service providers. The value of the entry comes from enough models and developers being willing to stay; once it becomes a sales channel that can only use NVIDIA products, its attractiveness to the entire ecosystem will decline instead. However, maintaining openness and influencing purchase choices can coexist. Which deployment solution is more convenient and which hardware is better adapted will affect where users spend their money.

Fable 5.1 cache price cut makes caching a core link in Agent pricing

On September 1, Anthropic released Fable 5.1 and Mythos 5.1 in its product announcement. The input and output prices of Fable remain at 10 USD and 50 USD per million tokens respectively, while cache reading prices have dropped from 1 USD to 0.25 USD. The company estimates that typical tasks can be about 25% cheaper, and tasks highly dependent on Agents can save about 45%.

Artificial Analysis found in tests of the same max tier that the output tokens of 5.1 are about 1.7 times that of Fable 5, and the cost per task in the comprehensive intelligence test rose from 3.14 USD to 3.76 USD, an increase of about 20%. The cache price cut has saved it about 1.40 USD, which still does not offset the new consumption. Meanwhile, its composite score increased by 4 points. What users get is a stronger, more expensive top-tier model.

Different usage patterns can lead to opposite results. Developer paddo recalculated his Claude Code records based on the API listed price: each invocation of 5.1 costs about 0.26 USD, compared with about 0.42 USD for the old version. Most of his original expenses were for cache reading, so he benefited significantly.

Anthropic's choice remains clear: to make work that repeatedly reads long contexts more affordable, and win over customers who previously thought Fable was too expensive. The official website cites feedback from Cognition that tasks such as code review, which previously could only be assigned to Opus, are now cost-effective to use Fable for.

This pricing adjustment transfers part of the budget from reading old contexts to more inference. Tasks with a high cache proportion can save money more easily, while upgrading to the top tier may spend the discount again.

Atlas world model released, taking a different path from Seedance

On September 1, World Labs introduced the new model Atlas in an article on its official website, and opened applications for early access. Atlas takes a different direction from video models like Seedance: one is more like building scenes, the other is more like making movies. Both deal with space and time, but their training focuses are different.

Seedance starts from videos, learning how consecutive frames change. How a person turns around, how clothes swing with movements, and how characters remain consistent after switching shots are all problems it needs to solve. Taking the public technical report of Seedance 1.0 as an example, the model processes the information within a single frame and the relationship between frames separately, and training is directly centered on movement, shot connection and visual quality. Videos first pursue the credibility of consecutive images. Making the scene of a cup shattering look realistic, and ensuring every fragment stays in place after changing the perspective, are two different requirements.

World Labs starts from spatial modeling. In its first-generation demo in 2024, the focus was already on generating three-dimensional worlds that can persist and be explored by users. This path requires scenes to remain consistent after perspective switching. When the camera moves behind the table, the positions of the table legs, doors and windows should match. The space can be viewed repeatedly, making it convenient for further editing, building game scenes, or being used in robot simulation systems.

Atlas continues the spatial modeling direction, training a multimodal autoregressive diffusion Transformer from scratch, putting text, images, camera positions and depth into the same spatial context. The value of the new base model lies in processing these clues uniformly, and using mature training and inference technologies to continue scaling up. The autoregression here can expand along the camera position: generating the next item means changing the viewing angle for the same scene. How an object moves to the next second belongs to dynamics prediction.

Public demos mainly demonstrate this spatial capability. Atlas's one-minute video is generated along an artificially designed camera path; "bullet time" changes the perspective for already recorded events. In the robot demo, Atlas helps build simulation, and public materials have not yet specified which part predicts object movement. The robot simulation system previously introduced by World Labs combines different representations and modeling technologies according to tasks.

The two paths will eventually meet in robot tasks. When a robot pushes a cup off the table, the model needs to predict how the cup will land; when the robot turns around and comes back, the model must still retain the changed scene. Models like Seedance start from changes in consecutive images, while Atlas starts from persistent space, each accumulating part of the capabilities required to complete this task. Atlas's commercial uses are therefore closer to game scene construction, spatial editing and simulation environment generation.

Gemini 3.8 released, Google discloses the automatic improvement loop in model R&D

On September 2, Google released Gemini 3.8 Flash and Flash Cyber. According to the product announcement, Flash has been updated three times within six weeks. The regular version continues the promotional price of 3.7, at 0.75 USD per million input tokens and 3.75 USD per million output tokens, with the discount lasting until the end of this year.

In the published evaluation by Artificial Analysis, the high-inference tier of 3.8 Flash scored 59, 3 points higher than 3.7, on par with the non-highest inference tier of GPT-5.6 Sol. The above scores use the old version of the index evaluated on September 2. Google still hasn't regained the lead in the frontier capability competition, but Flash has already accumulated a large number of actual invocations.

Google disclosed in its review on September 1 that the monthly active users of the Gemini app as a whole have exceeded 1 billion. 3.7 Flash also has a large number of invocations in the third-party developer market: when checking Google's model page on OpenRouter on September 6, its usage snapshot showed 2.6 trillion tokens.

Combined with these adoption signs, Google is very likely to have unexpectedly benefited from the rise of routing after the failure of tokenmaxxing. The most difficult judgments are still worthy of being assigned to strong models, while a large number of steps with clear requirements and easily verifiable results leave room for cheaper models to play, and the Gemini Flash series is a core choice in routing.

Google also disclosed the R&D method of 3.8, which uses a continuously running Agent loop in development to repeatedly evaluate and improve the underlying model. This is the automated research link disclosed by Google, and the degree of human participation in the complete training process has not yet been made public.

This is consistent with the automated training route we analyzed earlier. Research Agents propose improvement plans based on the model's failures, generate new tasks or data, call the training system to conduct experiments, and decide what to test in the next round based on the results. Repeatedly effective experiences are screened and incorporated into the next round of weight training. What the model participates in improving includes both training materials and the process of finding these materials and training methods.

Following this path, cheap, sufficient-capability models have another potential value: laboratories can afford more explorations. Constructing tasks in batches, analyzing failures, and screening candidate solutions all consume inference computing power; if every step calls the most expensive model, automated research itself will become a cost bottleneck. If Flash can complete these steps, laboratories can test more solutions with the same budget. Even if a certain generation does not get the highest score, it can still help the next generation find more problems worthy of training.

If the improved model can better organize the next round of experiments, it will have the compound interest of Recursive Self-Improvement (RSI). This loop screens solutions based on experimental results: whether the code runs successfully, whether the task in the environment is completed, and whether independent evaluations improve can all determine what to retain in the next round. Cheap models expand the number of attempts, and reliable verification selects effective improvements from them. The two together determine the efficiency of automated research.

The cost per task of 3.8 has also risen along with improved capabilities. Artificial Analysis measured that the cost per task of the 3.8 high-inference tier is about 0.58 USD, about 40% higher than that of 3.7, due to increased output tokens and more Agent interaction rounds. Google also retains 3.7 for tasks that prioritize efficiency.

Gemini 3.8 released, Google discloses the automatic improvement loop in model R&D

On August 31, OpenAI announced in a corporate statement that ChatGPT Ads had achieved an annualized revenue run rate of 1 billion U.S. dollars in less than 200 days since its launch, with tens of thousands of advertisers already using the service, covering more than 40 countries, and opening self-serve ad placement to Europe, India, the Middle East and North Africa. The 1 billion dollar figure is the full-year amount converted based on the current revenue rate.

These ads appear below ChatGPT's responses, displayed as ad units with a "sponsored" label, separated from the main text of the response. In regions where ads are available, Free and Go users may see them, while tiers such as Plus and Pro remain ad-free. OpenAI states that ads are served by an independent system, and merchants cannot pay to modify the model's responses.

Advertisers are matched primarily based on what users are chatting about, with display order determined by a combination of relevance and merchant bids. For example, if a user is comparing office software suitable for small teams, ads for related software may appear below the response. When permitted by region and user settings, the system can also refer to historical chats and memories for personalized matching. Advertisers see aggregated performance data such as impressions and clicks, and cannot read users' private conversations.

Merchants set budgets and bids, upload ad creatives through Ads Manager, and purchase ads on a cost-per-click or cost-per-thousand-impressions basis. OpenAI also provides the website tracking code Pixel and the Conversions API, allowing merchants to report back whether users registered, inquired or purchased after clicking, to help the system optimize ad delivery. The company says that by the end of August, ad campaigns that bid on clicks and optimize around conversion targets have become the majority.

Its business is closer to search advertising: users actively express demands, and ads bring people who are interested in learning about products to merchant websites. ChatGPT adds a continuous conversation, where users supplement budgets, use cases and preferences, and this information can help narrow the scope of ad matching. For OpenAI, this also gives Free and low-priced tiers a revenue source beyond subscriptions.

Significant differentiation has emerged in ad performance. OpenAI disclosed that an e-commerce advertiser achieved an ad spend return of about 3 times within 28 days. However, the early small-scale test reported by MediaPost on August 28 was not ideal: Canadian agency Choice OMG spent 415 USD in June, got 60 clicks, but did not attract customers with purchase intentions. In addition, Cleverly reported that the platform recorded 57 clicks, while Google Analytics recorded less than 20 visits, with no conversions.

But revenue growth shows that ChatGPT has already attracted real advertising budgets. As the self-service backend opens up to more small and medium-sized merchants, this business will be more directly tested by customer acquisition costs.

02. Research

Claude completes the Lean formal proof of Fermat's Last Theorem in 11 days

On September 4, Anthropic published the research report Formalizing Fermat’s Last Theorem, disclosing that Claude, with the help of the Prove2Me platform developed by Columbia University's team,