AI has crushed the Mensa IQ test, scoring a full mark of 151 to claim the top spot, outperforming 99.97% of all humans.
AI has completely blown human IQ tests out of the water.
The Mensa Norway IQ test includes 35 absurdly difficult graphic reasoning questions with a 25-minute time limit.
Claude Fable 5.1, paired with GPT-6 Astra that supports image-based answering, sat for the exam, got every question correct, and directly hit the theoretical maximum score of 151 for this test paper.
What is more, it has achieved full marks for 7 consecutive recent attempts. 151 is only the ceiling of this test paper, and no one knows yet where their own upper limit lies.
Placed among human beings, nearly 98% of people around the world cannot even reach the 130-point membership threshold of Mensa, and you and I are most likely among the vast group of people who are completely outperformed.
If you refuse to be convinced, try Question 33 on this test paper first, which is recognized as hell-level difficulty. The answer is placed at the end of the first section.
Question 33 of Mensa Norway Test
Don't lose heart even if you are totally confused. This question does not only torture humans, the vast majority of AIs have also stumbled here.
The cruelty of this test paper lies in that it does not examine any knowledge reserve, it just throws a bunch of confusing messy graphics at you, to see if you can dig out the hidden pattern from the total chaos.
Psychologists call this ability to find patterns on the spot "fluid intelligence", which was also the last line of defense that humans took pride in for their intelligence and used to outperform machines.
Now this line of defense has been blown to nothing left by AI.
Humans Only Reached 145, AI Scored 151
Open the ranking list of TrackingAI, and the first thing you see is a bell curve.
That is the distribution of human IQ, the vast majority of people gather at the hilltop near the 100-point mark.
In the past, the avatars of AIs were scattered in the low-score area on the left side of the curve.
Now, they have squeezed one after another onto the lonely long tail at the far right end of the curve, so crowded that the avatars are stacked on top of each other. Next to them stands a vertical line marked "Highest possible score for this test (temporary)".
Mensa Norway scores of various AIs on TrackingAI, the vertical line on the far right is the full score line of this test paper
This ranking list was compiled by American journalist Maxim Lott.
Just like a strict invigilator, he administers the Mensa Norway test to more than 30 mainstream AIs every week. Some models answer questions based on text descriptions, while others answer directly by looking at the original images. The score takes the average of the last 7 attempts, so there is no chance to get on the ranking list by sheer luck once.
This set of Mensa Norway tests he used was originally prepared for humans. When a human takes this test, the system only reports a maximum score of 145, and no further subdivision will be made for scores higher than that.
However, when TrackingAI scores AIs, it continues to convert scores upwards according to the number of correct answers. Getting all 35 questions right gives a score of 151, which has exceeded the measuring range set by Mensa for humans.
Nevertheless, the 145-point cap is only the limit of this test paper, and the IQ of the human population does not have such a cap.
We estimated according to the standard distribution of IQ, a score of 151 only appears in about 1 out of 3000 people.
Among the more than 8 billion people around the world, there are only about 2.7 million people with a score above 151, and there are only about 3000 such people in a big city with a population of 10 million.
The membership threshold of the Mensa Club is 130 points, and these two full-mark AIs are a full 21 points higher than that.
And 151 is still only the ceiling of this test paper. If the test paper can continue to test higher, at the 160-point level, there are only more than 200,000 people in the whole world.
The position where AIs stand is already the spot occupied by the tiny top group of people on the human IQ pyramid.
Moreover, far more than these two "monster" AIs have squeezed into this top tier.
On the TrackingAI ranking list, GPT-5.6 Sol and Gemini 3.1 Pro scored 145, a long line of models have broken through 140 points on public tests, and domestic models have also entered the first echelon.
With so many models crowded in the high-score section, the gap is fully reflected in the last few final challenging questions.
TrackingAI has counted that almost every AI can get points for the first 10 questions. At the end of the test paper, only 5 AIs got Question 35 correct, and only 2 AIs got Question 33 correct.
According to Lott's conversion, getting 1 wrong out of 35 questions only gives a score of 148, so the 2 AIs that got Question 33 correct are basically these two full-mark contestants.
The answer to the question mentioned at the beginning is announced here, the correct option is E.
Number of AIs that got each question correct, the orange bar represents the Mensa Norway test paper, only 2 AIs got Question 33 correct
A full score itself is scary enough, but you have to know that just a few years ago, AI was a student with severe partial subjects.
From 64 Points to 151 Points, AI Only Spent Two and a Half Years
In March 2023, a Finnish psychological assessment expert gave ChatGPT a formal Wechsler Adult Intelligence Scale, which only tested language questions such as vocabulary and general knowledge.
ChatGPT directly got a verbal IQ of 155, outperforming 99.9% of humans.
But when encountering brain teasers that require a little detour like "What is the name of the father of Sebastian's child", it crashed on the spot.
The non-verbal part was even untestable. In the words of this expert, ChatGPT "has no eyes, ears or hands".
It is obvious that language and knowledge have always been the subject that AI is best at, graphic reasoning is its biggest weak point.
What the Mensa Norway test paper examines is exactly graphic reasoning.
This kind of matrix question was first designed by British psychologist John Raven in 1936, and psychologists generally regard it as the question type that can best measure a person's overall intelligence level.
AI has been gnawing at this hard bone for 60 years, counting from the first geometric analogy problem-solving program developed by MIT in 1962.
Timeline of AI solving IQ test questions sorted out by us based on public materials
Until the beginning of 2024, when Gemini Advanced did the easiest Question 1 on the Mensa test paper, it could even seriously make up a series of "A+B=C" equations, while there was no arithmetic expression in the question at all.
In February 2024, Gemini Advanced did Question 1 of Mensa, making up a series of equations out of thin air
Lott translated the questions into text one by one and read them to AI, the result was not much better. Some AIs only guessed 6 questions correctly, got a shameful 64 points, which was almost the same as scribbling on the answer sheet with eyes closed, and the best model Claude-3 had just barely crossed the human average line.
So Lott asserted at that time that AI did not have the kind of general intelligence that humans have. Unexpectedly, more than half a year later, the turning point came.
In September 2024, OpenAI's o1 learned to think carefully before acting, it conducts internal reasoning before outputting, trying errors and checking step by step, and the score immediately rushed past 120.
Lott, who said AI was no good just more than half a year ago, wrote the title this time as "Huge Breakthrough in AI Intelligence".
Since then, thinking before answering has become the standard configuration of all flagship models, and the scores have been rising all the way. In April this year, the first 151-point score appeared on the ranking list for the first time.
Sorted out by us based on public data from TrackingAI and Maxim Lott, each point is the new highest record refreshed at that time
In this way, a student who had partial subjects for 60 years only spent two and a half years, evolving from a guesser with 64 points to a full-mark scorer with 151 points.
According to Lott's statistics, from May 2024 to October 2025, the IQ scores of top AIs soared by an average of 2.5 points per month.
In contrast, there are ready-made data in psychology showing how slowly humans improve.
Throughout the 20th century, with the improvement of nutrition and education, the human IQ test score increased by about 3 points every ten years, which is the famous Flynn Effect.
By comparison, it takes humans more than 30 years to gain 10 points, while AI only takes 4 months to achieve the same progress.
Growing so fast, the first reaction of any normal person is: Did it secretly memorize the answers?
Even with a Never-Seen-Before Test Paper Not Available Online, AI Still Nearly Hits Full Marks
Lott also doubted this back then.
After all, the Mensa Norway test has been posted online for many years, and the questions and answers have long been dug up completely by netizens.
Even TrackingAI itself made the text version of the 35 questions public together with the correct answers, so AI most likely brushed these questions during training.
Text version of Mensa Norway questions published by TrackingAI, the right column is the correct answers
To rule out this possibility, in May 2024, he specifically asked a Mensa member to create a brand new set of questions from scratch.
This set of questions has only been tried once by readers to set the scoring standard, and has never appeared on the Internet, so it is impossible for any AI's training data to contain these questions.
After changing to the new test paper, the memorization effect did show up. Until today, some models reveal their true colors immediately when switched to this confidential test paper, with scores dropping by more than 20 points compared with the public test paper.
However, the top models, after switching to a test paper they have never seen before, still nearly hit full marks.
The confidential test paper has only 16 questions, and the highest possible score is about 136. The image-based version of OpenAI's GPT-5.6 Terra has already reached 135.
Fable 5