HomeArticle

We are still a long way away from AGI.

王智远2026-09-20 19:46
It has to grow out of every computer.

Recently, many industry leaders have come out for interviews to talk about AI, and I have gone through all of them.

Everyone talks about industry cycles, investment hotspots, and the office agent battle. After reading all these contents, I felt I learned a lot, but still couldn't live a smooth life after the arrival of AI.

Until I read the 20,000-word long exclusive interview with Chen Danian, co-founder of Shanda Networks, from cover to cover, there is one sentence that feels more and more correct the more I think about it:

Let everyone have their own intelligence, instead of letting intelligence only belong to several companies. In other words, human wisdom exists among the masses of people, and so does AI.

Nowadays, people discussing AGI are roughly divided into two groups: one group fears its arrival, the other is waiting for it. The absurd point is that the two groups actually believe in the same thing: the advent of a god.

My point is: I neither fear its arrival, nor passively wait for it.

.......

The "god" that the two groups believe in has shown up very frequently in the past half month. If you connect all the news about AGI together, it is just like the classic novel Investiture of the Gods.

On September 8, OpenAI released a big news: ten thousand AI agents worked continuously for 88 hours, and submitted an answer to a millennium-level math problem that had been unsolved for 90 years, with 2.7 million messages exchanged back and forth and tens of millions of dollars consumed.

As soon as this incident broke out, the whole internet deified it directly, with headlines getting more and more exaggerated, but the mathematics community poured cold water faster than anyone else. Twenty-five Fields Medal winners issued a joint statement, cooling down the overheated AI "problem-solving competition".

The current status of this problem is "announced, pending verification", and it is still far from being "solved" before getting the full recognition of the entire mathematics community.

The drama on the other side is even more intense.

27-year-old Jakob Coxson resigned from Anthropic, without even waiting for the equity that was about to be credited to his account, and one of his resignation posts got 170 million views.

He had experience in pre-training work at both OpenAI and Anthropic, and he did not show any mercy in his resignation post, saying that the two companies are rushing towards superintelligence at full speed, which is equivalent to "gambling with all human lives".

The people who responded to his remarks are no trivial figures: the alignment lead of Anthropic gave a statistic that the probability of human extinction within ten years exceeds 10%; a few days later, the company's CEO Dario wrote an article calling for stepping on the brakes, and Sam Altman from OpenAI also expressed his approval.

Looking at the entire industry in the past six months, there is another wave of more frequent bustling phenomena.

New product launches are held one after another, large models are updated one after another, and the benchmark scores are refreshed as frequently as the calendar, with every other claim of "the strongest in history".

But if you move your sight away from the conference venue, the scenarios where AI actually works now are basically only a few categories: writing documents, searching for information, making spreadsheets, and making PPTs, just like an intern with fairly fast hands and feet.

As for the life of ordinary people? It goes on as usual. Going to work and off work, family trivialities, daily social interactions, AI has not been involved in any of these at all.

After observing for a long time, all these bustling phenomena show the same essence: chasing easy gains. Deification is easy: push a tiny achievement to the altar immediately; panic is also easy: link a tiny risk directly to the extinction of humanity; as for the benchmark scores? The numbers look amazing, but real life does not change at all.

Going around in circles, all the problems in the past half month point to the same direction: when will AGI arrive? But a more fundamental problem has been ignored: What on earth should the real AGI look like?

.......

Zhiyuan believes that being able to complete existing tasks is not enough. The real AGI must be able to invent completely new things.

Why? Let's talk about the criteria first. In that exclusive interview, Chen Danian said: The real AGI must be able to step into the unknown areas and break through the unknown barriers, solving the problems that humans cannot solve.

I quite agree with this point. I checked and found that Mark Zuckerberg also agrees with this view.

This criterion sounds harsh, but it is very fair when you think about it carefully. Being able to do existing tasks well only makes known things more and more skilled. Apart from causing some unemployment and making several AI giants make money, has it created real social value? The answer is no.

Being able to invent is completely different, which means writing the first stroke in the blank field that no one has ever touched; the word "creativity", when landed in reality, means creating something that never existed before.

On both sides of this dividing line, one side is just a more handy tool, and only the other side can be called real intelligence.

Measured by this ruler, the answer is not very optimistic. The number of unknown problems that AI has actually conquered in recent years is very small, almost all concentrated in the field of mathematics; mathematics has the advantage of clear rules and clear right and wrong results, which is the most suitable field for AI.

This is not "missing the carriage in the era when the automobile is emerging", but a reasonable doubt that AI is far from reaching the mature "automobile era".

Stanford has an AI Index Report, which contains several sets of data that are quite thought-provoking when combined together:

AI has reached the level of winning a gold medal in the International Mathematical Olympiad; but for a very simple task of reading an analog clock, the accuracy rate of the top large model is only 50.6%, barely better than tossing a coin, while the accuracy rate of human beings is 90.1%.

When you let AI operate the computer independently, the task completion rate is 66.3%, which means about one-third of the work will still go wrong. The success rate of robots in a fully controlled environment is 89.4%, but once they enter a real ordinary family, the success rate drops to only 12%.

The report named this uneven performance "jagged intelligence".

There is an original sentence in the report that is most vivid: it can win the IMO gold medal, but cannot reliably read the time on an analog clock.

Demis Hassabis, Chairman of Google DeepMind, the core AI institution of Google, put it more plainly: AI is a student with extremely unbalanced subjects, with super high talent in programming and poetry writing, but it keeps making mistakes when facing simple logic and basic physical common sense.

Fast progress does not mean getting close to the goal. The more amazing its test scores are, the more partial its ability distribution becomes.

Note:

Looking deeper, the current large language models are the most stubborn systems on this planet. How do they get into this state? It is completely different from the way human beings think stubbornly:

The model is generated by calculating the probability of each word, and at each step it chooses the path with the highest probability. The more it calculates, the narrower the path becomes, and finally it is completely locked in the old routine.

This description is very interesting: the lock is not outside the model, but in its own algorithm. It has walked this path countless times, so familiar that it can't be more familiar, but also so narrow that it can't break out of it.

For example, to beautify a web page, AI can make it loved by most people, because what the public thinks is good-looking can be unified; but if you ask it to be Picasso and create its own unique work, it cannot do it at all.

Not to mention being Picasso, if you ask it to come up with a slogan, people in the industry all know that the result is completely unusable, and you can tell it is AI-generated at first glance.

Its cleverness has a fixed direction, which is to convert all "possibility" into "absolute certainty". The contrast comes from this: tasks with known answers are done more and more skillfully, but it has never stepped into the door of tasks with no ready answers, while new things are exactly born in the place with the lowest probability.

"The real AGI moment should be when it finds the path that all human beings think is wrong, but it insists that it is correct."

Room temperature superconductivity is a ready example. Human beings have spent decades researching on it without breakthrough. If you throw this problem to AI, it will still go back to the old path that everyone thinks is correct.

A system without creativity cannot truly solve unknown problems, so creativity is the first threshold for AGI. If this threshold cannot be crossed, AGI is just three meaningless letters.

.......

So how to cross this threshold?

Recently, the AGI solutions given by many people and many enterprises all point to the same thing: local models.

What is a local model? It is a large model whose data is stored on your local computer, mobile phone and various personal devices, instead of a model that calls cloud resources every time you query or run tasks, which is what most of the AI large models we use now are.

The trend is changing faster than expected. The main development line in the past two years was still "who has the largest number of parameters", and this line has loosened this year.

Ilya Sutskever, former co-founder and chief scientist of OpenAI, put it more directly: the era of "blindly making models larger and larger" has come to an end; the evaluation criteria have also changed accordingly, from "how good the answer is" to "whether the task is successfully completed".

Moreover, it is not just one or two companies that are taking actions. On the other side of the ocean, Mark Zuckerberg published a 6500-word manifesto The Future Belongs to Everyone, stating that superintelligence should belong to everyone, not be held in the hands of a few institutions; on the same day, Meta open-sourced its 30B Muse Glimmer model.

Perplexity even directly entered this track, releasing Portable Computer, a localized digital employee, in late August. China is not lagging behind either. Lenovo released a forecast that 80% of token consumption in the future will occur on personal devices.

Alibaba open-sourced Qwen3.8-27B, which was downloaded more than 1 million times in two days; BigBang-V1 from Shanghai Jiao Tong University, Yuankong Intelligence incubated by Peking University, universities are also joining this track, and Yuankong even cooperated with HP to pre-install the model directly on new computers.

Chen Danian, whom I mentioned earlier, has a more sharp judgment: Within three years, local models will match the strongest cloud models today and capture 80% of the large model market.

Note:

Why are all parties gathering on this track? Because the economic and practical accounts are very clear.

Cost saving is the first point. Every reasoning in the cloud consumes tokens, and the more you use it, the more cost you generate; for the local model, the main cost is only electricity, so you can get real token freedom.

Security is a more important point. Continuously feeding core data to other people's closed-source models is equivalent to training for your competitors. Once the most competitive data of every person and every industry is uploaded to the cloud model and leaked, the core competitiveness of an enterprise will disappear completely.

From a broader perspective, if the cloud model development path goes to the end, it may lead to the largest monopoly in human history: a super intelligent partner that never betrays and keeps evolving is held in the hands of only a few companies. This is a matter related to the lifeline of the whole industry.

There is also the privacy consideration. The photos of several terabytes in your album, the plan you have revised eight versions and never shown to others, these things that you are reluctant to upload to the public cloud are exactly the most needed data for AI training.

You see, uploading your private data out or storing the model locally, the difference is whether you give away your lifeline to others, or let your own ability grow in your own hands.

Once the accounts are settled clearly, people will dare to use AI freely. Only when people dare to use AI freely, the vision that 8 billion people around the world can run their own exclusive models can be realized.

"I don't believe that a supreme god-like AGI can be realized. What I believe in is the collective power of 8 billion people around the world." A single model may not be smart enough, but the advantage lies in the infinite diversity, each model has its own unique personality and thinking pattern.

Then what about the cloud models that everyone is using now?

There is a very vivid description for it: "passing god". We still need to use them when appropriate, respect them when we should, and learn from them. Without the development of cloud models, the AI industry cannot thrive; they are just passing by on this road, changing from the only standard answer to one of many optional answers.

.......

Everyone actually knows the general direction in their hearts. But what about the usable products? Open source models are iterated batch after batch, and more and more teams are working on local deployment. It looks very lively, but there is still no product that can be easily used by ordinary people.

Installation is the first barrier that blocks users. Tools like Ollama cannot be set up smoothly without a certain technical background.

Running the model has higher requirements for the machine. Even if you finally get it running, it will fail when you run a long task. According to the aforementioned computer operation evaluation, only 18% of the models can complete tasks with more than 500 consecutive steps, and most of the tasks get stuck halfway.

There are also a number of "AI health products" mixed in the market, which can run smoothly in demo shows but expose all defects in real business scenarios.

The government and enterprise sector has taken the lead in landing related applications, and AI all-in-one machines have become standard procurement items. The demand has jumped from 150,000 units last year to 390,000 units this year; administrative reconsideration cases across Beijing have adopted local deployment, cutting the review time by 70%.

But these are all public service scenarios. For ordinary consumers standing in front of the counter, there is still no ready-to-use product for them.

This scene is not new, it happened 30 years ago. In the 1990s, the core of computing power was Sun's minicomputers, which cost hundreds of thousands of dollars each, and the most critical tasks in the whole industry had to run on them.

In 2000, Sun's market value soared to 200 billion US dollars, showing a scene of great prosperity.

As everyone has seen later, cheap, sufficient, and scalable PC servers spread all over the market. Sun fell from the peak and was acquired by Oracle for 7.4 billion US dollars, while Java, which was incubated by Sun itself, has become the global industry standard.

History never creates new stories out of nothing:

Today's cloud large models are sitting in the exact chair that Sun sat in back then. The truth remains unchanged: the cheap, open, and accessible product for all users will always defeat the expensive, closed, and tightly controlled one.

Is there a local model that ordinary people can use