HomeArticle

People like Jeff Dean, jump ship before the arrival of Bayesian AI.

36氪的朋友们2026-08-07 14:56
Research has pointed out the direction, and leading tech giants are also reinforcing the cornerstone of Agent. The present moment marks the very beginning of the Bayesian AI era, as well as the era when models will witness the fastest development.

On August 5, Jeff Dean announced his departure from Google, where he had worked for 27 years, to found a company called Discovery Loop. The ambition behind this name is to enable AI to enter a continuously running scientific discovery loop.

But last week, a Google DeepMind paper titled *LLM can't jump* also went viral.

The literal translation of the paper's title is "Large Language Models Cannot Jump". The "jump" here refers to a link in the scientific discovery process that Einstein mentioned in a letter to a friend.

The entire scientific discovery process can be regarded as:

Sensory experience E → Obtain axiom A through jumping → Deduce conclusion S from axiom → Return to experience for verification.

Scientists start from empirical facts, and form new basic principles through a leap that cannot be fully derived by existing logic. Then they deduce predictions from these principles and test the predictions in the empirical world.

This leap is the key to scientific discovery.

Zahavy's judgment seems to pour a basin of cold water on the grand narrative of Discovery Loop. Models may automate experiments, continuously improve along existing problems, and even complete complex mathematical deductions, but they lack the most critical step in scientific discovery: inventing a set of previously non-existent basic principles from limited experience.

A popular optimistic interpretation holds that scientific creativity is essentially a process of seeking better data compression methods. When old theories cannot explain observations concisely, the system will look for more unified new laws to reduce prediction errors.

But Zahavy does not agree that compression can cover the whole process of scientific discovery. He believes that this is only one of the three modes for humans to form knowledge. Data compression is essentially induction, which is responsible for finding patterns from data. In addition, deduction is responsible for deriving results from premises. But the real leap that changes the scientific framework is to invent new explanatory principles from limited facts, that is, Abduction.

From the author's point of view, current models are structurally incapable of doing this.

The abduction done by Einstein relied on translating simulated experiences into formal axioms. He first discovered unexplainable phenomena in the system, then used thought experiments such as free fall and elevators to simulate, and finally refined new axioms in this repeated simulation process.

However, current LLMs mainly process symbolized content that has been converted into language, and they do not have an action-controllable world model (a model that can change the state of the environment through actions) to complete those simulated thought experiments.

So, does it mean that the path Jeff Dean and several core Google AI figures have bet their careers on, which attempts to enable models to continuously evolve through long-term interaction, experimentation and self-modification, is fundamentally unfeasible?

In fact, there are very many layers between the capabilities of current models and the level of creating a completely new problem domain at Einstein's level, such as general researchers, and researchers who can propose partial theories and make revisions.

Current AI does not even have the ability to become a real scientist.

*LLMs Can't Jump* asks whether AI can complete the final leap. What Jeff Dean's entrepreneurial bet is on is whether AI finally has the conditions to reach the edge of that cliff.

Before answering whether a model can become Einstein, we may have to answer a more realistic question first: how far are today's large models from an ordinary but qualified scientific researcher?

01. The ceiling of scientific discovery for large models has already collapsed twice before the leap

Before a research enters the stage of experimentation and verification, two steps need to be completed: first, confirm "where is a problem worthy of research?", and then confirm "in what way to explain or solve this problem?".

Current studies on the scientific research taste of LLMs have targeted both of these two issues. They can prove that today's models do not suddenly lose their creativity when they reach the leap step, but have undergone two consecutive contractions long before that.

The collapse of the problem domain stems from the lack of skepticism and conceptual games

Scientific discovery is often described as finding answers. In fact, in many important studies, what is scarce is not the answer, but the ability to identify something as a problem.

The starting point of current AI scientific research systems is usually quite late. The models receive paper titles, abstracts, related work, datasets and task descriptions. The research objects have been named, the goals have been specified, and the advantages and disadvantages of existing methods have been organized into a mature narrative by the paper authors.

What the model really needs to do is often not to discover problems from the world, but to continue writing the next step in the problem space that has been organized by others.

Under this model, how will the model's ability to discover problems change?

The Hong Kong University of Science and Technology published the paper *AI Research Agents Narrow Scientific Exploration* in May, which conducted a large-scale study on this topic. The authors used four mainstream scientific research agent frameworks including AI Scientist, ResearchAgent, Co-Scientist, and Agent Laboratory, as well as five large models, to generate 219,655 research ideas from 155 research areas, and then compared them with the real human scientific research choices.

All four agent frameworks are set up to encourage innovation. AI Scientist conducts multiple rounds of self-reflection, ResearchAgent adds phased planning and review, Agent Laboratory allows multiple agents to discuss, and Co-Scientist even introduces competitive hypothesis evolution.

But even if all agents are explicitly required to find novel and high-impact directions and can actively retrieve additional literature, the problem awareness of the models is still relatively conservative.

The authors split each solution into "what to research" and "how to research", and then compared them with the seed papers. Only 10.5% of the AI solutions proposed new research problems that did not exist in the seed materials, but 90.4% introduced new methods that did not appear in the seed papers.

In other words, most of the novelty of the models does not come from proposing new problems based on the existing situation, but only from changing the solution locally.

For example, existing work is studying X, but current methods have deficiencies in efficiency, robustness or generalization. We can introduce method Y, add module Z, and then design a set of experiments to improve X.

It usually does not first ask: Why did X become a problem? Does X mix two completely different phenomena? Is the current way of measuring X correct? Is there a more important problem that is not even seen in the description of existing papers?

This is of course legitimate scientific research. The vast majority of papers in the real world are not Einstein-style theoretical revolutions. The problem is that when this structure becomes the overwhelming default route, the problem space itself will not be fully opened.

It tends to accept the presuppositions contained in the materials it reads, that X is a research object worthy of study, current indicators represent progress, and existing task boundaries are reasonable. Subsequently, innovation is limited to how to use new technologies to continue advancing X.

Retrieval and complex Agent frameworks can slightly help the model move away from the initial materials, but they have not fundamentally changed this local exploration mode.

The authors call this phenomenon local elaboration. It shows that models are good at expanding along existing research, but not good at broadening the territory of scientific exploration.

This is actually the opposite of in-depth human research. Human research often starts from conflicts and doubts.

Another paper from the University of Chicago published in June, *Contemporary AI Lacks the Imagination to Diverge or Negate in Science*, also proves this point. Their study invited the authors of 121,640 recent preprints to let the models generate follow-up hypotheses based on the background and scientific puzzles in the researchers' own papers. In the end, 6,749 scientists returned 25,139 sets of evaluations.

The study found that no model category will spontaneously propose a null hypothesis, that is, explicitly assuming that the relationship in the topic may not exist.

Moreover, the research directions they proposed are highly convergent. Although non-reasoning models can generate multiple hypotheses, multiple hypotheses often share the same premise, causal direction and problem framework.

Another paper goes a step further. The University of Chicago's paper *Measuring the Gap Between Human and LLM Research Ideas* published in July analyzed what kind of knowledge deficiency the models would define as a research opportunity.

The authors extracted references from 11,683 human papers that had been published in journals, and then let different models propose new research solutions based on these references.

They divided the research questions proposed in scientific research into seven categories: contradiction or puzzle, explanatory gap, mismatch of application scope, evidence gap, bridging opportunity, failure risk and resource bottleneck.

The distribution of human papers on these seven entrances is relatively scattered. According to the charts of the paper, humans most often enter from explanatory gaps, accounting for about 30%, failure risk, application scope, evidence gap and bridging opportunities each account for more than 10%, and the number of contradictions and resource bottlenecks is slightly less.

However, the model side shows a significantly different pattern. Among the nine model settings, 47.1% - 64.2% of the research motivations are classified as bridging opportunities, while the proportion for humans is only 12.1%.

In other words, when given several adjacent papers, the models often transform scientific research conception into a relational judgment, that is, re-connecting separated literature, methods, theories or evidence streams.

The reasoning-type models that we all expect to dig deeper into the structure and form deep cognition, on the contrary, seem to be more likely to take this splicing path. Taking Qwen3-8B as an example, under Thinking Mod, its bridging motivation rises from about half to more than 70%.

The University of Chicago's paper also found this phenomenon. They believe that models usually understand scientific tasks as "proposing a relationship between existing variables", rather than redesigning the way we observe the world.

Why do models make such choices? The first reason is that AI regards measurement methods and experimental conditions as external conditions. What it can see and be familiar with are all texts, so it only generates hypotheses based on published literature, and can only finally carry out recombination.

Moreover, bridging and recombination are the conclusions that can be most easily obtained directly from the literature. Paper A already has one method, Paper B already has another set of mechanisms, and the two have not been used together, so connecting them can directly be written as a research motivation. The models read the literature, and this method can even propose new directions without in-depth understanding of the literature.

As a result, the models not only propose fewer new problems, but even when looking for gaps in existing problems, the types of gaps they see are significantly narrower.

Current AI is like a researcher who dares not deny the premise, and only has connections between concepts in his hands and mind.

The collapse of the solution domain lies with superficial exploration and narrow imagination

Once the problem is determined, the model enters its more familiar area.

It can quickly call on a huge knowledge base and come up with more than a dozen technical routes for the same problem in one go. But "a large number of candidates" does not equal "a large number of solution structures". A large number of solutions may only perform the same type of cognitive action with different materials.

The first feature of these innovations is that the solutions are relatively close to the origin.

In *AI Research Agents Narrow Scientific Exploration*, researchers measured the distance of AI solutions relative to the starting literature. They compared each solution with the semantic center of five seed papers, and found that the average distance between AI ideas and seed materials is 0.322, while the average distance of subsequent human studies that actually cited these seed papers is 0.410.

Whether using single-round generation, or adding literature retrieval, reflection, multi-agent discussion and hypothesis competition, the overall AI solutions are closer to the starting point. Even Agent Laboratory, which is the farthest from the seeds, has an average distance of 0.382, which is still far from the level of subsequent human research.

This shows that models obviously tend to proximal scientific research search. Instead of looking for value in all possible problems, they look for the nearest reasonable next step in the visible space defined by the current materials.

This shallowness of the search domain does not mean that the solutions are naive. A technically quite complex solution may also be very close to the starting point. The solution may involve complex architecture and complete experiments, but what it changes is still only a section of technical configuration on the existing route.

In addition to the "shallow" search domain, it is also "narrow", and it is particularly easy to be attracted by a few "novel combinations".

This paper also tested the exploration breadth of the models, that is, how different a large number of model ideas are from each other in the same research field. The average breadth of AI ideas is 0.554, and that of human papers is 0.599, the overall model is about 7.5% lower. This difference does not mean that all the model's solutions are the same, but it means that when the model is generated repeatedly, the semantic regions it occupies are more concentrated.

But what exactly is this concentrated pattern?

When talking about the problem domain earlier, the paper *Measuring the Gap Between Human and LLM Research Ideas* mentioned that models like to look for research problems starting from the possibility of bridging. Correspondingly, its solution is also obviously biased towards Synthesis / Unification.

In this approach, what the model does is to connect, integrate, reconcile or unify different literature, theories, evidence, mechanisms and methods. This type of method only accounts for 5.1% in human papers, but reaches 22.5% - 38.7% in models. In the subset of machine learning, the deviation is even more exaggerated. The comprehensive solutions of different models reach 35.5% - 70.7%, and that of humans is 6.6%.

Subsequent prototype clustering restores this preference to a more specific research operator. The most common operation of the model is integration, which accounts for 34.2% of the model's output, but only 2.35% of human ideas. Humans, on the other hand, more often first check whether a certain part should be replaced, whether two things should not be mixed together, or whether this problem has not been clarified from the beginning.

The most common AI scientific research solution path you can see is to put two different solutions into a unified framework.

The novelty of this kind of solution mainly comes from the novelty of operator combination. Existing objects have not been matched in this way, existing modules have not been put into the same system, and existing mechanisms have not been applied to this scenario. Combination can of course produce real breakthroughs, but humans are not dominated by this method.

It is just application and combination, which is the easiest way to use the literature itself.

But even with this set of application and combination, models prefer to use novel concepts rather than delve into the literature itself to find mechanisms, hypotheses and evidence.

The authors of this paper found that when using the integration method, although the models frequently mention "integrating multiple works" in language, after putting the solution and its 4-8 preceding papers into the same embedding space, human solutions instead use multiple papers more evenly. In the comprehensive index that considers the central position of the literature, coverage uniformity and whether it is dominated by a single paper, the human score is 1.4662, Qwen is 1.3345, and DeepSeek is 1.4237.