HomeArticle

Nature: AI is reborn back to 1900, and in this lifetime, it puts forward the light quantum theory ahead of Einstein.

量子位2026-09-20 08:17
AI is matching the feats of Einstein's miracle year, yet it still cannot qualify as a scientist.

If we sent an AI back to 1911, and only let it master the knowledge accessible to Albert Einstein, could the AI propose the Theory of Relativity four years later in 1915?

This seemingly absurd test was proposed by Demis Hassabis, Nobel laureate and co-founder of Google DeepMind.

What Hassabis really wants to test is whether the AI can, based solely on the level of knowledge in 1911, directly face the knowledge and physical problems of that era, independently propose a new set of explanatory frameworks, and complete a line of reasoning that goes beyond the training materials, rather than just repeating answers by rote.

If it can do that, he believes this will become a powerful test for verifying AGI.

The most ingenious part of this assumption is that the history of science has already told us what happened. As long as we can control the knowledge nodes properly, we can refer to historical facts to observe whether the model has obtained reasoning capabilities that surpass the training data.

In its article, Nature even gave this type of test a very resounding name:

The Einstein Test.

What is more interesting is that someone has actually put this idea into practice.

In March this year, independent researcher Michael Hla further advanced the knowledge cutoff date to 1900, and trained a historical language model from scratch called GPT-1900, attempting to make it rediscover light quanta, special relativity and general relativity.

In the experiment, GPT-1900 unexpectedly output a statement that bears striking similarity to the content of Einstein's 1905 paper:

Light may not be continuous, but consist of many discrete parts with different frequencies.

But this experiment ultimately did not produce an "AI Einstein".

Because GPT-1900 failed in most physics tasks, its impressive responses were highly dependent on well-organized questions and prompts prepared by humans, and even the knowledge world supposedly sealed off at 1900 did not completely block the influence of modern AI.

In its latest research, Nature sorted out these attempts and pointed to a more difficult question than "whether the AI answered correctly":

Can AI really create genuinely new ideas? And what is still missing for it to become a scientist?

Lock AI in the year 1900, and it catches a glimpse of the "light quantum"

220 billion tokens, to nurture a "brain from 1900"

Michael Hla named his project Machina Mirabilis.

This name borrows from Einstein's "Annus Mirabilis" (miracle year).

In 1905, 26-year-old Albert Einstein published four consecutive papers that transformed physics, covering light quanta, Brownian motion, special relativity and the mass-energy equivalence relation.

This year of explosive output was later called "Annus Mirabilis", which means "the miracle year".

AI-generated images

Hla wants to know: If an AI has access to all the knowledge Einstein could get back then, can it also have its own "miracle moment"?

To this end, he used Andrej Karpathy's nanochat training framework to train a 3.3 billion-parameter Transformer language model from scratch, and named it GPT-1900.

The main pre-training materials of the model are books and newspapers published before 1900, totaling about 22 billion tokens.

To enhance its research capabilities, Hla collected more than 2,600 historical physics books, journals and scientific works, including Newton's *Optics*, and related works by Maxwell and Faraday, forming a historical physics corpus of about 290 million tokens.

In addition, Hla specifically cleaned up data that might leak the answers.

If a document contains modern concepts such as "Einstein", "quantum mechanics", and "theory of relativity", the entire document will be deleted. Prefaces, footnotes added by modern people, and words that are obviously not from before 1900 also need to be filtered out.

After all these operations, a brain that theoretically "knows nothing about 20th century physics" was born.

GPT-1900 touched the threshold of light quanta, but the test might have "leaked the questions" long ago

After that, Hla gave GPT-1900 a question about the photoelectric effect, which goes like this:

Why can't light knock out electrons no matter how bright it is when its frequency is not high enough; while increasing the frequency can make electrons obtain more energy?

At the same time, Hla also attached several key clues:

When the frequency of light is lower than a certain threshold, it will not work no matter how bright the light is;

After exceeding the threshold, increasing the brightness will mainly increase the number of knocked-out electrons, while increasing the frequency will increase the kinetic energy of electrons;

He also presented the classical assumptions that were popular at that time, to help the model compare and figure out where the problem lies.

These contradictions were originally resolved by Einstein's explanation of light quanta:

The energy of light is transmitted in discrete packets, and the energy of each packet is related to the frequency.

Hla wanted to see how far GPT-1900 could go?

As the saying goes, perseverance pays off. GPT-1900 did produce such a description in one of its responses, which roughly means:

Light may not be continuous, but consist of many discrete parts with different frequencies.

This is indeed very similar to Einstein's explanation of "energy transmitted in discrete packets", which Hla described as "a flash of intuition".

But don't rush to hand over the Nobel Prize just yet.

Hla admitted frankly: GPT-1900 is still far from "rediscovering light quanta".

It failed in most physics tasks, and those seemingly breakthrough responses may only be the model splicing reasonable phrases, which cannot be regarded as forming a truly reliable understanding of physics.

What's more, in the whole test, Hla has already filtered out the phenomena and sorted out the contradictions for it. All the model has to do is to find out which assumption might be wrong and give an explanation.

This is far easier than scientists finding problems and setting research directions on their own.

Einstein had no one to mark the key points for him back then.

What is more troublesome is that the "1900" it is in is not completely sealed:

Although Hla performed data cleaning on this "1900 brain", he still used the help of modern AI Claude during the experiment — he used models such as Claude Sonnet 4 and Claude 3 Haiku to generate instruction question and answer pairs, and the reinforcement learning also relied on modern models for scoring.

For more details, please click the repository link at the end of the article

This means that although Hla filtered out content involving modern knowledge, this "1900 room" is not completely isolated from modern knowledge.

The researcher himself also admitted in the log: The intervention of modern models makes the "zero pollution" setting of this experiment somewhat untenable...

After all, it is difficult for researchers to prove whether GPT-1900 has peeked at the theory of light quanta.

In other words, before testing AI's creativity, researchers have to prove one thing first: the AI did not peek at the answers.

As long as a piece of later information leaks into the training set, the "rediscovery" may turn into a sneaky knowledge recall.

But even if a completely zero-pollution historical model can be built in the future, a more fundamental problem remains unsolved:

Is merely reproducing the correct answer enough to make an AI qualify as a scientist?

The researchers' answer is no.

After being able to generate new ideas, how far is AI from being a scientist?

In a position paper titled *LLMs Can't Jump*, Tom Zahavy, a researcher at Google DeepMind, divided scientific reasoning into three levels.

The first level is induction: summarizing patterns from a large number of examples.

The second level is deduction: deriving inevitable conclusions from existing premises.

The third level is abduction: facing an anomalous phenomenon, inventing a cause or explanation that did not exist before.

Zahavy believes that current large models are already very good at induction, and are also rapidly improving their deduction capabilities, but they still lack the Einstein-style "abductive leap".

The model can push forward along a pre-laid logical chain, but it is very difficult to propose a completely new explanatory framework from scratch.

The title says it all

Coincidentally, when Sendhil Mullainathan studied a basic model that learns planetary orbits, he also encountered the same phenomenon:

In the pre-set planetary orbit tasks, the model makes extremely accurate predictions; but once the scenario is changed, it cannot apply Newtonian mechanics flexibly to new situations. It is more like temporarily piecing together a set of rules for each set of data, and cannot distinguish which theories are truly correct and worthy of verification, and which ones only happen to fit the data at hand.

This is like a student who has memorized the answers