OpenAI has poached the latest Fields Medal winner. Who is ByteDance's newly launched scientist program targeting?
This year's Fields Medal is widely regarded as the "last purely traditional" edition — it is likely that future breakthrough research will be completed in combination with AI.
This transformation process may be faster than we imagined: At the Fields Medal press conference, when someone asked Canadian mathematician Jacob Tsimerman, one of the winners, about his next work plan, he stated that he had already shifted his focus to AI safety research and would soon join the safety team at OpenAI.
On the other side, Mark Chen, Chief Research Officer of OpenAI, immediately extended a warm welcome.
Leading scholars at the cutting edge now view top AI companies as platforms where they can fulfill their ambitions. Recently, a growing number of top scientists, including mathematicians, physicists, biologists and others, have been moving to leading large model companies. In the first half of this year alone, Nobel Prize in Chemistry laureate John Jumper, economist Chad Jones, philosopher Harvey Lederman and others have all joined Anthropic.
Behind this trend, top AI companies have all turned their attention to the same direction: they are inviting scientists to the front lines of AI R&D. The deep integration of AI and cutting-edge science is evolving from an optional path to a necessary one.
Domestic major AI companies are also following this trend. Just this Thursday, ByteDance's Seed Edge team launched the Seed STEM Scientist Program, opening 100 collaboration positions for scientists and doctoral students in cutting-edge scientific fields such as mathematics, physics, chemistry and biology, inviting them to work with Seed to explore unknown problems in fundamental science. According to public information, this is the first large-scale program in China dedicated to investing in AI to accelerate cutting-edge scientific discoveries.
In the second half of the large model era, the competition is no longer about benchmark scores
In the current fast-paced AI industry where competition is measured in days, delving into long-term fundamental AI research may seem far less intuitive than polishing coding skills and chasing higher benchmark scores: improvements in the latter can be quickly translated into product experiences and commercial revenues, while the feedback cycle for fundamental research and AI-accelerated scientific discovery is much longer, and the exploration path is full of uncertainties.
But on the other hand, once a real breakthrough is made on this path, the paradigm-shifting value it brings will far exceed ordinary product iterations.
Yao Shunyu once shared a viewpoint with his team: There is no magic in large model training. The real challenge lies in getting the most basic and definitely doable things right. This judgment echoes the long-held view of Anthropic CEO Dario Amodei — the core elements driving AI progress are always computing power, data quality and scale, training duration, and scalable objective functions, and "all clever methods and tricks are actually not that important".
Looking back at truly SOTA models such as Claude and Seedance, there are no shortcuts behind them. Their core success all comes from locking in the direction early and doing a great deal of foundational work on data quality and underlying architecture.
Anthropic highly focused its resources on improving the model's coding capabilities. Through reinforcement learning training, the model can independently understand and generate high-quality code through repeated trials, and form a self-accelerating evolution flywheel based on user feedback.
In the field of video generation, which many large companies find difficult to sustain progress in, Seedance has turned AI video generation from a "random draw" style experiment into a more controllable and creative industrial-grade production tool through innovative model architecture and meticulous data work.
It is clear that what truly determines the upper limit of a model's capabilities is never the accumulation of tricks, but solid data, stable foundational structures, a long-term vision, and sustained strategic focus and sufficient patience.
This is no easy task in the current environment, but we can see that on the domestic large model track, more and more companies are willing to settle down and take this difficult but correct path.
This year, DeepSeek launched the mHC (Manifold-Constrained Hyper-Connection) architecture, which is regarded as a breakthrough in the field of large model foundational architecture. It is also one of the important architectural upgrades of DeepSeek V4, replacing traditional residual connections and running through the entire model design. It is worth mentioning that the mHC research cites the Hyper-Connections (HC) proposed by ByteDance's Seed team.
Before that, residual connections, as the basic skeleton of Transformer, were widely regarded as the default configuration in the industry. HC was the first to break this single-stream paradigm, greatly improving the model's feature expression capability while barely increasing the amount of computation. From the pioneering proposal of HC to the engineering implementation of mHC, it has become a microcosm of technological innovation by Chinese AI teams in the underlying architecture field, which is rare in the past development of large model technologies.
Early last year, ByteDance Seed specially established a team called Edge, hoping to work on research topics with low certainty that may not yield results in the short term. To provide researchers with a stable environment, they even adjusted the performance assessment cycle — no assessment is set in the middle, and a unified assessment is conducted after results are produced.
This team has been relatively low-key since its establishment. Not long ago, they released the ultra-long-horizon evaluation dataset EdgeBench on X, which was the first in the industry to systematically define the learning laws of Agents in real environments, and discovered a new Scaling Law in the environmental learning dimension, triggering discussions among many industry leaders.
It is said that OpenAI is also using this evaluation dataset to test the long-horizon capabilities of its own models.
In the AI field, people often focus on the large models themselves. Building a Benchmark actually requires a very long cycle and a large amount of manpower. Designing and open-sourcing benchmarks for advanced AI capabilities is not only to measure one's own progress, but also to guide the entire field toward a more transparent and reliable direction by establishing industry standards and exposing model weaknesses. Top AI labs invest heavily in building benchmarks.
OpenAI open-sourced the Evals framework, which has become a standardized tool in the industry for evaluating large language models and LLM-based systems. For specific cutting-edge capabilities, OpenAI has also designed a series of dedicated benchmarks, including SimpleQA, BrowseComp, MLE-bench and more.
Anthropic has done a lot of work around SWE-bench, clearly highlighting Claude's strengths in coding capabilities; they also released Terminal-Bench to measure the ability of Agents to complete long-link tasks in a real command-line environment.
The relationship between these evaluation efforts and the strengths of various models is no accident.
It can be seen that the judgment of leading AI teams on the development path of AI has been changing. The capability evolution of large models has moved beyond the pre-training dividend period that simply relies on stacking parameters and data. The core growth momentum in the next stage will inevitably come from the continuous interaction, feedback and autonomous evolution of models in real open environments.
The cutting edge of fundamental science is the highest form of open and unknown scenarios. The Seed STEM Scientist Program, newly launched by ByteDance, is designed to identify and break through research-level problems together with scientists, and explore the possibilities of AI accelerating scientific discovery. Scientists not only provide research topics, but also problem awareness, judgment criteria and feedback methods.
Why are top AI companies
all competing for domain scientists?
Why cooperate with scientists? The answer is that the remaining "problems" are no longer challenging enough for large models.
Most of the current mainstream model evaluation systems focus on closed-ended tasks with standard answers. The problems have definite solutions, the evaluations have fixed criteria, and the path to improving model capabilities is clear. But real scientific research follows a completely different logic: it has no clear problem-solving path, requiring researchers to independently propose hypotheses, design verification schemes, process abnormal and noisy data, and derive underlying mechanisms. The entire process is full of trials and errors, repetitions and uncertainties.
Tasks in the real world are the ultimate standard for testing the capabilities of AI models. Following pre-training Scaling and test-time Scaling, the industry is generally looking for the next Scaling direction, and "continuous interactive learning in a real open environment" is one of the most promising paths.
Therefore, when developing cutting-edge models, top global AI teams no longer only rely on their AI teams to polish models behind closed doors. Instead, they introduce top researchers from various disciplines to let real domain problems become the driving force for model capability evolution.
In North America, many university researchers have recently joined large model companies one after another. In early July, Jelani Nelson, head of the EECS department at UC Berkeley, announced his academic leave and joined Anthropic. Since the beginning of this year, at least 22 professors and researchers have temporarily left or reduced their work at universities such as Stanford, UC Berkeley and Harvard, and moved to OpenAI, Anthropic, DeepMind and Meta. These four companies currently employ at least more than 80 current or former university professors.
Behind the successive arrivals of top experts, there are also systematic projects.
At OpenAI, the Residency program has become a regular high-end talent recruitment channel for OpenAI every year. It is clearly targeted at talents "who do not currently take AI as their main research direction", allowing participants to directly join OpenAI's cutting-edge application and research teams for practical collaboration from day one. This program is equivalent to an "accelerated doctoral training", which can help cross-domain experts quickly complete interdisciplinary integration. Many milestone projects such as ImageGPT and Codex have seen deep participation from resident researchers. The breakthroughs in scientific research capabilities of the GPT series models in mathematical reasoning and biological mechanism derivation are also inseparable from the continuous input of such interdisciplinary talents.
At Anthropic, the STEM Fellows Program, which is specifically designed for experts in science, technology, engineering and mathematics fields, was systematically and massively expanded in 2026. This program lasts for several months, specially inviting top non-AI experts such as physicists, biologists, mathematicians and economists to join Anthropic. They can bring unsolved problems or complex workflows in their own fields, and directly use the unreleased Claude models and internal evaluation tools.
In such programs, material scientists will build a dedicated evaluation process for Claude's phase stability reasoning, and climate scientists will deeply integrate professional atmospheric modeling tools with the model. During the residency period, more than 80% of the researchers have produced high-quality academic results. This feedback from real disciplines directly becomes the training signal for model iteration, and also promotes Claude's evolution in scientific research reasoning and complex tool invocation.
The value of such projects has long gone beyond scientific research itself. Top AI companies are all placing the competition for domain scientists on an equal footing with the competition for computing power.
Of course, for the scientific community, such cooperation is also of great value. Large models are becoming new scientific research infrastructure: from formula derivation and numerical simulation to analysis of massive literature and experimental data, AI is reshaping the paradigm of scientific research, making it possible to advance research topics that were previously limited by manpower and computing power.
Being invited to join a large model team for resident research is as meaningful as astronomers getting exclusive access to the next-generation super space telescope. Such programs bring many advantages to scientists, not only providing access to massive AI computing power and cutting-edge model experience, but also allowing them to use AI as an "external brain" to focus their energy on defining problems themselves.
For these scientists pouring into Anthropic and OpenAI, such in-depth collaboration is not only an amplifier of scientific research efficiency, but also a lever to leverage new discoveries. Now, looking back at China, we find that the gap in the intersection of AI and fundamental scientific research is being filled first by ByteDance Seed.
The endgame of AI
is a competition for long-termism
When a Fields Medal winner decides to join OpenAI, and more and more STEM scientists start to work with model teams, the next stage of AI competition is shifting from short-term product iterations, benchmark rankings and single-point task breakthroughs to longer-term, more complex and more cutting-edge fundamental propositions.
Scientific research is critical because it stands at the forefront of human knowledge. If AI can gain a firm foothold in such high-complexity open environments, it means that the next generation of AI truly has the ability to participate in real complex tasks.
This path cannot produce eye-catching results quickly, nor can it form a clear commercial closed loop in the short term. It is slow, resource-intensive, and full of uncertainties. But precisely because of this, it can best test whether a team is willing to truly settle down and invest in fundamental and cutting-edge research.
The final answer that large models need to deliver is whether they can work with humans to advance those unsolved problems.
Scientific progress requires long-termism, and so does the realization of AGI.
This article is from the WeChat public account "Machine Heart" (ID: almosthuman2014), author: Zenan, published with authorization from 36Kr.