HomeArticle

Nature Sub-journal: When 800 million people are using AI for writing, linguistic diversity is disappearing rapidly

账号已注销2026-08-26 17:40
After the advent of ChatGPT, our writing styles have become more and more similar.

"We are shaping language into its final form, the form it will take when no one speaks any other language." That's what British writer George Orwell said in his dystopian novel 1984.

Nowadays, a similar trend may be unfolding in the real world.

This week, a research paper published in Nature Human Behaviour, a sub-journal of Nature, by a team from the University of Southern California (USC) and their collaborators reveals that the widespread use of Large Language Models (LLMs) as writing assistants is linked to a decline in linguistic diversity.

Put simply, LLMs are making the texts produced by different people "more and more alike".

Paper link: https://www.nature.com/articles/s41562-026-02550-0

Language, everyone's unique "fingerprint"

Language is far more than a communication tool. What a person says and how they say it often reveals their identity, psychological state and social background. From regional dialects to subtle expressions in daily communication, language is always closely tied to individuals and the groups they belong to.

Accordingly, language has become a key clue for observing people in many fields. Psychologists track mental health conditions through linguistic analysis, social scientists use it to study moral values and social biases, the marketing industry leverages language for large-scale personalization, and clinical research even attempts to use language to assist in identifying diseases such as depression.

But all these applications have a prerequisite: when people express the same meaning, they do not use exactly the same language. Everyone has their own expression habits, and it is precisely these differences that allow language to reflect personal identity, personality and psychological state.

Then here comes the question: the training goal of LLMs is to generate "the statistically most probable continuation", a mechanism that naturally favors mainstream language patterns. When tools like ChatGPT and Claude have more than 800 million users, and everything from emails and codes to papers and copywriting is being "polished" by AI assistants, will the diversity of human language continue to exist?

In this study, the research team analyzed more than 880,000 texts covering different types including news articles, academic papers, argumentative essays, social media posts and speeches, and through three studies found that: the use of LLMs may indeed weaken the linguistic clues related to personal identity, personality and mental health, making it harder to identify these traits solely through language.

Study 1: The variance in writing complexity drops significantly after the release of ChatGPT

This sub-study examines whether the growing use of LLMs is associated with the homogenization of writing styles. The study is divided into two steps.

The first step is to observe real-world data. The research team analyzed three large-scale longitudinal datasets: arXiv paper abstracts (about 80,000), Patch News articles (about 379,600), and Reddit creative stories (about 318,000), with a time span from 2018 to 2024. They used the high-precision detection tool Binoculars to determine whether a text was AI-generated, and tracked the variance changes of the "writing complexity" indicator. The results are highly consistent:

  • arXiv: After the release of ChatGPT, the variance in writing complexity saw a significant and continuous decline, and changes in AI usage can significantly predict subsequent changes in variance. In other words, the more frequently AI is used, the more convergent the styles of academic writing become.
  • Patch News: There was a sharp drop in variance when ChatGPT was launched, but the Granger causality test was not significant. The research team believes that professional editing processes and institutional norms may have buffered the homogenization effect of AI, yet subtle impacts still exist — it may not be direct usage, but AI-generated content plays an indirect role by influencing the overall writing norms.
  • Reddit: Similarly, a continuous decline in variance appeared after the release of ChatGPT, and AI usage also has a significant predictive effect on the decline in variance, but with a longer lag (about 20 months). They speculate that this may be because the Reddit user group is more extensive and diverse, and adopts AI tools more slowly than the research community.

The second step is controlled experiment. They randomly selected 1000 original human texts published before the release of ChatGPT from Reddit and arXiv, and used GPT-3.5, Llama 3 70B, and Gemini Pro to polish them with neutral prompts (such as "rewrite" and "improve grammar"), then compared the variance of writing complexity before and after polishing. The results show that while retaining the original meaning (87% of the texts have a similarity of over 0.95), LLMs significantly compress the variance of writing styles.

Study 2: It has become harder to "read a person" from their writing

If Study 1 discusses the phenomenon that "everyone's writing is getting more and more similar", Study 2 raises a more pointed question: Can we still recognize the author from their writing?

The research team first collected text corpora paired with authors' psychometric data, including age, gender, Big Five personality, empathy and moral values, then trained classifiers to predict these personal traits based on text features, and compared the prediction accuracy of original texts and LLM-rewritten texts. They found that:

First, the prediction accuracy drops by about 6% on average (absolute F1 value). The prediction performance of all six types of traits decreases significantly, among which age is the most affected: the F1 value falls from 0.351 to 0.260. However, the performance of the classifier is still above the random level. This indicates that AI has not completely erased the identity signals in the text, but only weakened these signals.

Second, this change is not random, but systematically biased towards specific profiles. Texts rewritten by LLMs are more often judged as: from older people, with a stronger sense of morality, lower level of empathy, lower extraversion, but higher openness and agreeableness.

Third, this trend holds across different models and prompts. No matter which LLM is replaced, or how the rewriting instructions are adjusted, the offset direction of identity traits remains consistent.

In other words, LLMs do not simply "erase" identity signals, but introduce new biases, gradually pushing the texts of different authors towards a certain similar "default persona".

Study 3: The classic "word-person" associations are being washed away

Sociolinguistics and psychology have accumulated a large number of well-established conclusions that "certain word categories are associated with certain traits". Study 3 aims to test: do these associations still exist after being rewritten by LLMs?

In the original texts, the research team successfully reproduced many classic findings, such as "extroverts use more positive emotional words" and "openness is associated with complex vocabulary". However, after being rewritten by LLMs, many of these associations are "washed away". For example:

  • The association between extraversion and pronoun usage: disappears significantly;
  • The association between loyalty and friend-related vocabulary: disappears significantly;
  • The association between age and future-oriented vocabulary: disappears significantly.

Some other associations, such as "neuroticism and negative emotional words", are retained. This shows that the impact brought by LLMs is closer to a kind of selective erosion, rather than one-size-fits-all noise.

For disciplines that rely on lexical clues to conduct research, this may be more troublesome than overall distortion — it is necessary to test one by one which associations are still reliable and which have failed, otherwise researchers may draw biased conclusions about identity, group differences or cultural patterns.

What does this mean?

Since the large-scale adoption of LLMs has already brought these changes, what potential impacts will they have on us? The research team gave some examples as follows:

1. Clinical and mental health: After writing styles become highly homogenized, it may become more difficult to identify key markers such as depression through texts, hindering early detection and intervention;

2. Personalization and marketing: Targeted analysis that relies on subtle linguistic differences may fail;

3. Recruitment: Applicants whose resumes are polished by AI and meet mainstream style expectations may have an advantage over applicants with unique but "less polished" writing styles, exacerbating unfairness;

4. Cultural preservation: Cultural markers carried by language are smoothed out, putting linguistic heritage at risk.

What is more alarming is a feedback loop: texts modified by LLMs continuously enter online corpora, and these corpora will be used to train the next generation of models, which may make the loss of linguistic diversity gradually self-reinforcing.

Finally, the research team also raised a far-reaching question: Given that language and thinking are inseparable, the homogenization of language may eventually limit cognitive flexibility and creativity. In extreme cases, as Orwell warned, this may "narrow the scope of thought".

Of course, this study is not intended to deny the value of LLMs. They do make writing clearer and make it easier for more people to express themselves.

However, when the texts from hundreds of millions of people all flow through the same set of statistical models, the uniqueness of individual expression is being quietly diluted. It does not stem from censorship or deletion, but is gently smoothed away in every request of "polish this for me".

How to make AI tools a positive force that amplifies the diversity of human expression, while avoiding that everyone gradually writes the same "AI tone", will become a question that must be answered in the AI era.

This article is from the WeChat Official Account "Academic Headlines" (ID: SciTouTiao), author: Academic Jun, published with authorization from 36Kr.