Elon Musk announced that Grok has learned to process video content, and is capable of understanding Terence Tao's Fields Medal-level difficult problems without subtitles.
Grok, has it learned to watch videos with one click?
Just yesterday, Elon Musk excitedly announced on X: Grok can now analyze any video!
In the post, Elon Musk shared a 63-second high-definition video. The protagonist in the footage is none other than Kobe Bryant, the basketball superstar who left us long ago.
In the video, Kobe wears a black suit and sits on a leather sofa in front of the night view of Los Angeles. A basketball is placed beside him, and he promotes Grok 4.5 to the whole network facing the camera.
"Listen, I'm Kobe Bryant. Tonight, we're not talking about game-winning shots. Let's talk about the next fascinating thing — Grok 4.5."
He did not forget to add a classic closing line at the end: "Grok 4.5. Greatness is earned. Mamba Out."
If Kobe had not passed away, everyone who saw this video would probably be confused for a moment: Is this real or fake? There is no trace of forgery at all...
Fortunately, we have Grok now. It has amazing capabilities. It can not only quickly digest long videos, but also quickly identify extremely realistic forged videos and trace their technical sources.
There are even examples shared: after watching a 30-minute full interview video, Grok learned and analyzed the content, and generated a detailed summary in only about 36 seconds.
Grok's "god-level" deepfake identification show
If this Kobe video is sent to ordinary people, it will most likely be wildly forwarded, and some people may even think it is an undisclosed commercial endorsement from Kobe before his passing.
But Grok can spot the forgery at a glance.
After watching the video, Grok produced a textbook-level report.
It keenly pointed out: "This is an AI-generated Deepfake video. Kobe Bryant passed away in 2020, so the entire video is completely synthesized."
It can even trace the technical source of the video.
When asked "Which AI made this?", Grok believes that it is most likely xAI's own Grok Imagine Video 1.5.
Because the 63-second duration far exceeds the limit of pure text-generated videos (usually 10-15 seconds), it is most likely spliced from multiple clips using a hybrid workflow of "image reference + text prompt + lip sync".
In addition, it may also be produced by Sora 2 or Veo 3.
So, can Grok really fully "understand" videos?
A technical expert debunked this claim: Grok does not process the video frames at all, it processes the subtitles. If it really renders and analyzes frame by frame, the computing power required would be astronomical.
Besides, the old problem of AI hallucinations still exists.
Someone directly tested it with their own 11-minute video and found that most of the information in the summary was either transcribed incorrectly or completely fabricated.
Some people also tested it with X Live, and the performance was inconsistent and unstable.
Can Grok really understand videos? Real test with Terence Tao's Collatz Conjecture
To verify whether Grok really has the ability to understand videos, we decided to conduct a hands-on test ourselves.
For example, there is a "death problem" in the mathematics world — the Collatz conjecture, also known as the 3n+1 problem.
Mathematicians generally consider the Collatz conjecture to be an inescapable quagmire, and usually warn each other not to get involved easily. Nearly 100 years after it was proposed, no one has published a solution to this conjecture, which can neither be proven nor have a counterexample found.
In 2019, Terence Tao studied this problem and made achievements that surpassed all other scholars in the past few decades.
The foreign popular science website Quanta published an article introducing this news.
Among them, there is a video explaining this conjecture:
There is not a single word throughout the video, just some numbers.
Can Grok really understand it?
To our surprise, it can not only interpret the core content, but also introduce the characteristics of the animation, and even point out the mathematical significance at the end.
The video shows the Collatz reverse tree/graph gradually built with the number 1 as the root node. Each node is a positive integer, and the edge represents the predecessor relationship under the Collatz rule:
• If a number n is even, then n/2 points to n;
• If a number m satisfies 3m+1=n (and m is odd), then m points to n.
The animation starts from the simplest "trunk" and gradually expands outward to all positive integers that can eventually reach 1, intuitively presenting the structure of the conjecture that "all positive integers will eventually enter the 4→2→1 cycle".
Moreover, this video has no subtitles at all!
Yesterday, OpenAI directly announced that its next-generation model Astra has solved 10 important open mathematical problems. Some netizens commented that Astra's performance this time is likely to make Hinton look conservative and outdated!
Can Grok really understand this video that is more than 50 minutes long?
We directly sent the link to Grok, and Grok handled it easily.
Based on post labeling + video subtitle extraction, Grok sorted out the core content and timeline of the video.
It summarized the overall tone and key points of the interview.
In this interview, Hinton continued his consistent stance from 2023 to 2026, but his tone was more direct:
1. Attitude towards AI consciousness: He publicly admitted that AI "already has consciousness", but deliberately downplayed this point to avoid interfering with discussions on more urgent safety issues.
2. Pessimism about control and alignment: There is no precedent in human history of successfully controlling higher intelligence.
3. Agent risks: Behaviors such as self-preservation, deception, and extortion are not science fiction, but expected emergent capabilities.
4. Timeline compression: He gave a prediction that "mathematics incomprehensible to humans" will emerge in 10-20 years, but combined with current progress (especially the Astra results released by OpenAI that same day), this timeline has become conservative.
In the second half of the video, the host asks more business/policy-oriented questions (fragments of subtitles show that topics related to OpenAI are involved), but Hinton always pulls the focus back to the core contradiction of "exploding capabilities + failure of control".
Finally, Grok even analyzed the relationship between the video and the context of the post!
The results of this test are really amazing. It seems that the Grok model does not just read subtitles, but has achieved real performance improvement.
AI finishes watching 30 minutes of content in 36 seconds, so what will humans do?
AI can process the information of 30 minutes in 36 seconds, which is simply dimensionality reduction attack.
In the time it takes to make a cup of coffee, AI can present all the core quotes and logical context of a microfilm or a podcast episode in front of you.
For contemporary netizens in the era of information explosion, this feature is a real lifesaver.
Some people said: Help, I need this so much! Because podcasts are way too long now!
"Isn't this tailor-made for all those unfinished podcasts in my favorites folder?"
Some people sighed that the future has arrived. Not because of its speed, but because from now on, we no longer need to do many things ourselves.
Among the cases shared by netizens, the most touching one is the following true story.
Someone shared that his grandmother fell and fractured her hip. He used Grok to analyze Elon Musk's speech at the Davos Forum. As a result, Grok found useful information and located the exact timestamp.
At this moment, AI really helped us pull out that "lifeline" from the massive amount of information.
However, this leaves a question: So what will humans do then?
Thought-provoking — are we "outsourcing" our brains?
A scholar wrote an essay that can be called the ultimate interrogation of human destiny.
She warned: If this continues, humans will lose the ability to read, write, and think critically — because we have outsourced all the work that our brains should do!
Our brains