HomeArticle

Who is more creative after all? 100,000 people and AI completed the same test question.

开智学堂2026-08-25 08:18
On a creativity test, GPT-4 has already scored higher than the average human level. However, the most creative group of people still stays ahead of all the tested models. This four-minute quiz designed by a magician, which has been taken by 100,000 people, is now also presented to AI. Guess what top percentage of all participants you can rank into?

First, let's try a quick quiz

The rule is simple: write down 10 nouns, and the more unrelated their meanings are, the better.

"Cat, dog, table" won't work — these three are far too closely related, clearly belonging to the same category. You should aim for something like "galaxy, fork, freedom, algae, harmonica", where each word comes from a completely distinct, unconnected domain.

You can actually grab a piece of paper and try this out, it only takes four minutes. Don't read further until you finish — because in the research we're covering next, 100,000 people alongside models like GPT-4 and Claude 3 completed exactly this same task.

Figure: The quiz page on the official test website datcreativity.com. There are 6 rules: only English words are allowed, only nouns, no proper nouns, no professional jargon, you have to think of the words yourself (don't copy items you see around the room), and you need to be creative.

This task has an official name, the Divergent Association Task, which was published in *Proceedings of the National Academy of Sciences* in 2021. What it measures is very straightforward: how far apart the meanings of the words you come up with can be. Psychology has long held a view that people who can generate new ideas often spot connections between two things that seem completely unrelated.

The task designer Jay Olson started doing magic tricks at the age of seven, and later brought this skill into the psychology lab. He once ran an experiment on the street, asking 118 passersby to "randomly" pick a card: 98% of them picked exactly the card he wanted them to choose, and 90% of the participants were fully convinced they had made the choice of their own free will afterwards. Magic relies on manipulating attention, and the kind of difficult problems he picks for research are of the same type: how do you measure something intangible that exists only inside the human mind?

Creativity happens to be one of the hardest traits to measure. Traditional testing methods require judges to grade answers manually, which is tiring for the judges and often leads to inconsistent scoring standards across different people. The clever part of this task is that it leaves the grading to algorithms. The meanings of words can be mapped onto a "semantic map", where words with similar meanings are placed close to each other, and unrelated words are far apart. The average distance between every pair of your 10 words is your final score. It measures only one aspect of creativity called "divergent association", not the full scope of creativity — we will revisit this point later in the article. But precisely because the grading is done by algorithms, the scoring is consistent for every participant, no matter if the test taker is a human or an AI model.

This Time, the Task Was Given to AI

In January 2026, *Scientific Reports* published the largest human-AI creativity comparison study to date, led by Karim Jerbi, a cognitive neuroscientist at the University of Montreal whose lab normally researches EEG and consciousness.

All the collaborators on the study have connections to creativity. The first author Antoine Bellemare-Pépin composes music in the music department at Concordia University; Olson, the magician we mentioned earlier, is also part of the team; Kory Mathewson from Google DeepMind, a recipient of the Canadian Comedy Award, is one of the first people to bring AI onto the improv comedy stage; the list of authors also includes deep learning pioneer and Turing Award winner Yoshua Bengio. Composers, magicians, comedians, and AI pioneers came together to study this very simple 10-word task.

Figure: One of the authors, Turing Award winner Yoshua Bengio. Photo by Maryse Boyce, from Wikimedia Commons, CC BY 4.0.

Before this study, the task had already collected 100,000 responses, with an even split of men and women, ranging from 18 years old to over 60, with 20% of participants in each of the five age groups. The researchers had nine models — GPT-3.5, GPT-4, GPT-4-turbo, Claude 3, GeminiPro, plus four small open-source models including Vicuna — complete the same task, then ranked the AI models' scores alongside the scores from the 100,000 human participants.

The Average Score Falls Behind, But the Highest Score Still Belongs to Humans

When answering the task with default settings, GPT-4 achieved the highest score, with an average score significantly higher than the human average; GeminiPro's average score was on par with the human average; the remaining seven models, including Claude 3, all scored below the human average. If you line up all the scores on a bar chart, the human average ranks second, only lower than GPT-4.

Figure: Average DAT creativity scores, sorted from lowest to highest. The grey bar represents 100,000 human participants, the green bar represents GPT-4; GPT-3 in the figure refers to GPT-3.5. Source: Bellemare-Pepin et al., *Scientific Reports*, 2026 (CC BY 4.0).

But if you sort the 100,000 human participants by their scores and look at the top 50% of scorers, their average score is still higher than all the tested models; if you only look at the top 10% of human scorers, the gap between humans and machines becomes even larger. The distribution chart below makes this even clearer: the scores of the 100,000 people range from 30 all the way up to over 95, and no model's score can reach the highest range at the far right end; the scores of each individual model are all clustered in a very narrow range.

Figure: Score distribution of all tested models and 100,000 human participants (grey area). The dashed line represents the mean value. Source: same as above (CC BY 4.0).

The researchers then selected three models: GPT-3.5, Vicuna and GPT-4, and gave them three more tasks closer to real creative work — writing haikus (Japanese three-line short poems), writing movie plot summaries, and writing micro-fiction. The trend remained the same: the work produced by the machines looked polished and well-formed, but still could not match the output of the most talented human writers.

There was another unexpected finding in the results: GPT-4's successor version GPT-4-turbo scored far lower than its predecessor, even falling behind much smaller models like Vicuna and GPT-3.5. At least for this set of tested models, model size and version update level do not directly translate to higher scores.

Temperature and Prompts Can Both Improve AI Performance

So are these scores fixed and unchangeable? The researchers tried two methods to adjust them.

The first method is adjusting the temperature parameter. Temperature is a model parameter that determines how "uninhibited" its responses are: a low temperature means the model picks the most safe, conventional words; a high temperature means it is willing to pick more rare, unconventional words. Regular users cannot modify this parameter by default, but developers can. The researchers raised GPT-4's temperature from 0.5 to 1.5, and its score increased accordingly. At the highest temperature setting, its average score reached 85.6, surpassing 72% of all human responses.

The second method is modifying the prompt. Adding the instruction "start with the origin (etymology) of the words to find unrelated words" improved the performance of both GPT-3.5 and GPT-4 compared to their performance with the original instruction; switching to the instruction "consult a thesaurus" did not lead to significant score changes.

The score a given model gets depends partially on how people prompt it. The researchers noted in their paper that prompting and interaction have already become an integral part of the creative process.

Similar Studies Have Been Conducted Multiple Times in Three Years

Back in September 2023, two researchers from Finland and Norway had ChatGPT (versions 3.5 and 4) and copywriting tool Copy.Ai complete the "Unusual Uses Task" with 256 human participants — the task asked people to think of new uses for everyday objects like ropes and cardboard boxes. The AI models had higher average scores, but no AI model's average score exceeded the highest score achieved by a human. In the same year, a team from the University of Montana mixed GPT-4's responses in with student submissions and sent them to the official scoring agency of the Torrance Tests of Creative Thinking, a classic creativity test that has been used in the United States for decades. The scorers did not know AI responses were mixed in, and the result showed that the AI's originality score ranked in the top 1% of all submissions.

For three consecutive years, three independent studies all found that under their respective test setups, the strongest AI models had average creativity scores comparable to the human average. The three studies had different tasks, participant samples and tested models, so none of their conclusions can replace each other — but they all point to the same trend. This study, which uses 100,000 humans as the reference group and adds writing tasks, is the largest-scale study of its kind.

However, doubts have persisted, coming from several different perspectives. The first set of objections targets the task itself. The textbook definition of creativity has two core requirements: original and useful. This task can measure how far apart the associated words are, but it cannot measure whether these associations are practically useful. The task's original designer also acknowledged this point in the paper. Other researchers ran a correlation test comparing this task with other established creativity tests, and reported that the correlation was weaker than the original study claimed, though those results were only presented at an academic conference and have not yet passed the peer review process for formal journal publication.

The second set of objections is hard to refute: while a single AI model can get a high score in one test, the outputs from different AI models are surprisingly similar. This trend is already visible in the earlier distribution chart: every model's scores are clustered in a very narrow range. A 2025 study looked deeper into this phenomenon: when multiple models repeatedly completed this type of task, the similarity between different AI outputs was far higher than the similarity between outputs from different humans. Exaggeratedly speaking, models developed by different companies produce responses that look like they came from people who lived together in the same dormitory. Another group of researchers ran two pre-registered experiments with a total of 1100 participants (pre-registration means publishing the full research plan publicly before the experiment starts, to avoid only reporting positive results afterwards): when humans used AI to help them generate ideas, their immediate performance was indeed better; but in the final independent task that did not allow AI assistance, the researchers observed no lasting creative benefits from previous AI use, and the ideas produced by participants showed signs of converging to each other. The rich, unpatterned diversity among human creators, where no two people produce identical ideas, is a dimension that individual standardized tests cannot measure.

The third set of objections points to the well-known data contamination problem. These tasks are publicly available online, so there is no way to rule out that the models' training datasets already contain these tasks or highly similar content. The authors of this study themselves noted in the paper that high scores on standardized tests can be achieved by AI through completely different paths than humans use. The Torrance Tests task is kept confidential by the scoring agency and never published publicly, so this risk is relatively lower; but for open tasks like this 10-word association task, the contamination risk is much harder to eliminate.

The fourth set of objections argues that the test conditions themselves put humans at a disadvantage. In 2025, another team of researchers replicated the earlier 256-participant experiment: as long as humans were given more time and the task requirements were clearly explained, human participants could match the performance of AI models. Part of the original gap between humans and AI came from the unfair test setup.

Back to the 10 Words You Wrote Down

These four sets of doubts challenge the extended claim that "test scores equal real creativity"; they do not invalidate the actual test results themselves.

When it comes to "coming up with a set of unconventional associations within a time limit", AI performance has already exceeded the human average — both the DAT task and the Unusual Uses Task confirm this trend, and GPT-4 also ranked in the top 1% for originality in the Torrance Tests. Real creative work is of course far more complex than a four-minute quiz, and test scores do not equate to evidence of better job performance; but if the core requirement of a job is exactly this kind of standardized, fast divergent thinking ability, these scores are at least a meaningful reminder — this is an extrapolation beyond the test results, not a conclusion from the study. There are also reassuring findings in the data: the top 10% of human participants outperform all tested AI models, and the same trend holds for the writing tasks; another study found that at least on these creativity tests, the diversity of outputs from human participants is far higher than the diversity between different AI models — no two humans produce the same ideas, each has their own unique quirks.

In this study, the scores AI models get are also tied to how they are prompted — both temperature parameters and prompt wording can change their performance, though not every arbitrary prompt adjustment will work. As for whether the highly similar outputs from different AI models can be made more diverse through human guidance, existing studies have only raised the question, and have not yet found a clear answer. Karim Jerbi put it very straightforwardly in the press release: instead of obsessing over whether AI will replace humans, it is better to treat generative AI as a creative tool. How it will reshape the creative process depends on future research, and more importantly, depends on the people who use it.

The 10 words you wrote down for the opening quiz are still in your mind, right? The official test website (datcreativity.com) is still open today, and you can get your score in four minutes. 100,000 people have taken this exact same quiz — guess what percentage of all participants your performance would rank in?

References

Bellemare-Pepin, A., Lespinasse, F., Thölke, P., Harel, Y., Mathewson, K., Olson, J. A., Bengio, Y., & Jerbi, K. (2026). Divergent creativity in humans and large language models. Scientific Reports, 16, 1279. https://doi.org/10.1038/s41598-025-25157-3

Carolus, A., Koch, M. J., & Feng, S. (2025). Time