HomeArticle

Is Meta bouncing back strongly? It claimed five gold medals in the olympiad competitions with pure reasoning, two of which achieved perfect scores.

机器之心2026-08-07 12:50
It also corrected the wrong official answer!

In the early hours of this morning, OpenAI released GPT-5.6 Luna for free. Almost at the same time, another company that once created the Llama series and once fell behind in the large model competition also spoke out — Meta announced on 𝕏 that it had sent its in-house AI model to five STEM discipline Olympiad competitions, and all achieved gold medals or gold-medal-level results, with full marks directly obtained in the theoretical exams of two physics competitions.

What is even more striking is that Meta stated that its AI achieved these results entirely through reasoning capabilities: "We prohibited the use of all tools, which means no search, no coding, and no calculators."

In response, Lucas Beyer, a Meta researcher who previously worked at OpenAI, DeepMind and Google Brain, also teased Gary Marcus and Yann LeCun on 𝕏. However, the two renowned researchers who have long criticized autoregressive LLMs did not respond directly.

It seems that Meta is back in full swing again?

However, Meta has not yet published a relevant technical blog or research report, but Shuchao Bi, a Meta researcher involved in the project, also posted two tweets adding more details, including that the model is an internally trained version from the Muse Spark series, and uses multi-agent orchestration and parallel reasoning.

Not only that, Meta's model also correctly answered questions for which the official answers were wrong. Shuchao Bi said: "This discovery ultimately prevented the answer sheets of human participants from being incorrectly scored. The IPhO Committee kindly sent us a physical gold medal."

Result Sheet

The five results listed by Meta are: full marks in the theoretical exam of the Asian Physics Olympiad (APhO), full marks in the theoretical exam of the International Physics Olympiad (IPhO), gold medal in the International Mathematical Olympiad (IMO), gold-medal-level performance in the International Chemistry Olympiad (IChO), and gold-medal-level performance in the Romanian Master of Mathematics (RMM).

First, let's talk about the weight of these events. The 67th IMO 2026 is held in Shanghai, where contestants from more than 100 countries solve six questions in two 4.5-hour sessions, with a full score of 42 points. The RMM is held in Bucharest every February, inviting only more than a dozen of the strongest national teams, and its questions are widely recognized as no less difficult than those of the IMO. IPhO and APhO are respectively the highest-specification high school physics competitions in the world and in Asia.

In terms of results, the two physics events are marked as "full marks", the IMO is marked as "gold medal", while IChO and RMM are marked as "gold-medal-level performance". The latter wording in this field usually means that the results have not been officially assessed by the competition organizing committee, but are obtained by self-checking against the gold medal score line of the current year.

Meta also expressed its gratitude to the contestants and organizing committees of these competitions for "supporting our participation", which implies that the participation in at least some of the events was coordinated at the organizing committee level — this is more formal than OpenAI's approach in 2025, which found three former IMO medalists to score the results on their own.

Pure Reasoning: All Tools Disabled

Meta specifically emphasized that in order to test the pure reasoning ability, all tools were disabled for the model: no search, no coding, no calculators.

This constraint is very important. Over the past year, many top results on the Olympiad leaderboard were actually obtained through the agent pipeline of "strong model + verifier + multi-round refinement". Some researchers built an agent with a fairly simple structure using Gemini 3.1 Pro, which ran five times on the 2025 IPhO theoretical questions and got full marks for all five attempts; but the authors themselves also noted that the model was released later than the competition, so data contamination cannot be ruled out.

https://arxiv.org/pdf/2603.03352

Meta turned off all tools, which is equivalent to voluntarily giving up the easiest path to get high scores.

However, there is another unmentioned point about physics competitions: both IPhO and APhO include practical exams, and the results reported by Meta should only cover the theoretical part.

From Llama to Muse

2025 was the most difficult year for Meta AI. Llama 4 Maverick only scored 18 points on the Artificial Analysis Intelligence Index, and Behemoth was postponed due to internal dissatisfaction with its capabilities. Mark Zuckerberg then personally stepped in to restructure the team, invested heavily to bring in Alexandr Wang and the Scale AI team to establish Meta Superintelligence Labs (MSL), and completely rebuilt the pre-training architecture, data pipeline and RL post-training system.

On April 8, 2026, MSL launched its first model, Muse Spark, which is natively multimodal with 262k context window, and its intelligence index jumped from 18 to 52. On July 9, Muse Spark 1.1 was launched, focusing on agentic capabilities and programming, and it is also Meta's first paid API model.

Just this week, Muse Code, a terminal programming agent based on Muse Spark 1.2, entered beta, and it is said to have quite outstanding capabilities.

At this pace, this Olympiad result sheet is more like a trailer for the next-generation Muse model.

Is Meta Back in Full Swing?

This is the most noteworthy question. In July 2025, when both OpenAI and Google DeepMind won IMO gold medals (35/42, solving 5 out of 6 questions), OpenAI researchers called it the "moon landing moment" of the industry. A year later, the entire landscape has completely changed.

After the 2026 Shanghai IMO ended in July, Huawei's Celia and Xiaohongshu's dots-note-3.0 successively announced that they got a perfect 42/42 score. Xiaohongshu stated that no large model had ever obtained a full score under the official IMO review process before. Refer to the report from Machine Heart: "AI Obtains First Official Full-Score Gold Medal at IMO, With a Built-In Self-Correction Mechanism".

In addition, there are also reports that Anthropic's Claude Opus 5 passed all six questions in one attempt without using the agent framework and tools, with all four independent solutions being correct.

In other words, in the IMO event alone, Meta's "gold medal" is not particularly prominent in the 2026 rankings. What is truly valuable is the two full marks in theoretical physics — provided that these results can be independently verified.

What is more worth noting is that the Olympiad, as a measurement of reasoning ability, is becoming invalid at a visibly rapid speed. Meta's stated reason is "to understand whether we have made real progress in reasoning", but when the same set of competitions changed from "unattainable" to "multiple laboratories can get full marks" within a year, the amount of information this ruler can provide is sharply decreasing.

The mathematics community has long issued a reminder: Terence Tao pointed out in 2025 that what AI can do depends to a large extent on the test method itself. Olympiad questions have standard answers, clear scoring criteria, and a large number of historical question banks available for training, which are exactly the conditions that real scientific research problems do not have.

So, is Meta back in full swing? From Muse Spark to Muse Code to this result sheet, it is at least no longer the company that can only maintain its presence through open source narratives. But to say that it has returned to the first tier, it still needs to deliver more, especially a product that can be truly put into practical use.

Reference Links

https://x.com/prz_chojecki/status/2085444920983298494

https://x.com/AIatMeta/status/2085388945148297322

https://x.com/shuchaobi/status/2085386404805374100

This article is from the WeChat public account "Machine Heart" (ID: almosthuman2014), author: Panda, authorized for release by 36Kr.