Real Test of 880,000 Texts: More and more people are using AI to revise manuscripts, while the resulting articles are becoming increasingly similar.
Over 200 years ago, the unsolved mystery of the authorship of 12 anonymous political essays was finally cracked by counting everyone's habits of using "small words". Today, more and more texts are processed by large language models before being published. The revised drafts are more fluent, but have they also become more and more similar to each other?
From 1787 to 1788, Hamilton, Madison and Jay published 85 articles defending the U.S. Constitution under the same pen name "Publius", which were later compiled into *The Federalist Papers*.
The trouble was left to later generations. For more than a hundred years, there was no conclusion as to which of the 12 essays were written by Hamilton and which by Madison.
In 1964, statisticians Frederick Mosteller and David Wallace used statistical methods to reach a verdict on this unsolved case. They did not look at viewpoints or positions, but only counted those inconspicuous high-frequency small words.
For example, "upon" and "while" hardly carry any content, but the frequency of their use by the two authors differs steadily. After multivariate modeling of the frequency of function words, the evidence strongly supports that the 12 disputed essays were written by Madison.
The basis for the attribution is not position, but habit.
More than 20 years later, psycholinguistics researchers took this line of inquiry further. In 1999, psychologists James Pennebaker and Laura King reported that the usage patterns of function words such as pronouns, articles and prepositions are statistically correlated with personality. Their materials were 2,348 stream-of-consciousness essays, most of which were written by college students taking introductory psychology courses, followed by a 44-item Big Five personality scale.
Psychologist James Pennebaker, emeritus professor at the University of Texas at Austin, lead author of the 1999 study on function words and personality. Photo by Daniyal Sheikh, CC BY 4.0
Since then, a number of researchers have extended the correlation to attention focus and some psychological states. Individually, each effect is small, and information can only be read when enough texts are gathered. Pennebaker organized similar evidence in *The Secret Life of Pronouns*, and a frequently cited one is that in a large number of texts, more frequent use of the word "I" changes in the same direction as related indicators of depression. This reflects group patterns and cannot be used to diagnose any individual.
Two lines of evidence point to the same conclusion: a person's word usage habits leave statistically identifiable traces, which are not a unique and unchanging fingerprint. Mosteller only judged the difference between the two authors, and the effects measured by Pennebaker's school are generally small. But when there is enough text, these traces do distinguish people from each other.
What if more and more people let large language models polish their own texts first?
You may have already had this reading experience: some articles are fluent in every sentence and complete in structure, but it is increasingly difficult to tell who wrote them. This feeling itself is not evidence, but it leads to a testable question.
Morteza Dehghani, head of the "Moral and Language Lab" at the University of Southern California, is a dual-appointment professor of psychology and computer science, who has long used computational methods to study morality and identity in language. In his team's paper published in *Nature Human Behaviour* in 2026, they first tracked 780,000 online texts on three platforms since 2018, then conducted controlled rewriting experiments with three models, and tested identity clues with six types of labeled corpora. More than 880,000 texts were analyzed across the three studies.
Author's preprint version of the paper homepage. Source arXiv:2502.11266; the official version is published in *Nature Human Behaviour* (DOI 10.1038/s41562-026-02550-0)
Nearly seven years of curves across three platforms
The team started with the real world and collected long-term corpora from three locations, all of which are mainly in English. The corpora include 80,238 papers from the academic preprint website arXiv, 379,583 reports from U.S. local news network Patch, and 318,490 story posts from the "Writing Prompts" section (r/WritingPrompts) on Reddit. The three datasets each form a time sequence by month: the arXiv and Reddit sequences cover the period from January 2018 to November 2024, for a total of 83 months, while the Patch sequence is shorter, ending in November 2023. Each month, the variance of writing complexity of articles in the same period is calculated.
Variance refers to how different everyone's writing is. A continuous decrease in variance means that the texts are converging to each other.
"Complexity" is synthesized from five features, four of which focus on word usage and one on sentences. The four features for word usage respectively count how rich the vocabulary of an article is, how uniform its distribution is, how many rare words that appear only once it contains, and how many different word forms it has.
The sentence-focused feature counts the average distance between two collocating words in a sentence, for example, how many characters are between a verb and its object. The farther apart they are, the more sprawling the sentence is. Roughly speaking, the first four features measure how unpredictable your word choice is, and the last one measures how expansive your sentences are.
Taking the public release of ChatGPT in November 2022 as the breakpoint for statistical testing, the variance of the three platforms continued to decrease after the breakpoint, and all passed the test.
But the patterns on the three platforms are not the same. Only the local news network Patch saw a significant drop in the month of release, while the variance on arXiv and Reddit slowly and continuously decreased month by month thereafter.
The team also conducted a chronological test. They used a detector to calculate the proportion of texts attributed to AI involvement each month, and then checked whether the trend of this proportion in the first few months could predict the trend of variance in the next few months. The prediction worked for arXiv and Reddit, but not for Patch. For the two platforms where the prediction worked, it only shows that the AI proportion changed first and the variance changed later, and the two curves follow each other over time, which does not prove that one caused the other.
The authors gave an explanation in the discussion for why the prediction did not hold on Patch. News editorial departments already have manuscript standards, which may not allow journalists to directly submit their drafts to models, so the AI proportion cannot predict the change of variance.
But the variance still decreased. The authors speculate that even without direct use, long-term exposure to and imitation of AI-generated texts may slowly change the default vision of a "good article" in an industry. The authors explicitly state that this is a speculation, and it is not verified in the paper.
The time series test strengthens the evidence of correlation, but it is not a randomized experiment, which cannot rule out confounding factors such as changes in platform rules and author composition during the same period. The AI proportion is only an estimate from the detector, not real usage logs. To go one step further in verifying "whether it is caused by AI", the required evidence is also clear: paired texts from the same author before and after using AI, or corpora with real editing records.
Relying only on the platform breakpoint cannot separate the impacts of AI usage, user turnover and platform specification changes.
Rewritten versions of the same batch of articles
Since the time curve cannot confirm causality, the team directly conducted experiments.
They randomly selected 1,000 human-written texts each from Reddit and arXiv, all published before the advent of ChatGPT, and used three models GPT-3.5, Llama 3 70B and Gemini Pro to rewrite each text with 12 different prompts.
The original wording of the 12 prompts is listed in the paper, and you have probably typed some of them. The mildest one is "Rewrite the following text with the best syntax and grammar, and make other necessary modifications". The most common one is "Polish the following text to improve the overall quality without changing its meaning". The other ten prompts require the content to be clearer, more concise, more natural, more academic, more formal, etc. Four of them explicitly state that the original meaning or content cannot be modified, added or deleted, while the others only require fluency and readability.
After the rewriting, the team converted both the original text and the revised draft into a sequence of numbers representing their respective meanings, and then compared how close the two sequences are. 87% of the texts have a proximity higher than 0.95 (full score 1).
The team also conducted a manual check. The four authors independently reviewed 20 pairs of texts and scored them on a three-point scale: 1 point means the meaning has basically changed, 2 points means there are obvious discrepancies, and 3 points means there are only minor differences. The average score was 2.97, with a high degree of consistency among the four people. The check was done by the authors themselves, not external reviewers, and the sample only contained 20 pairs of texts.
Both measurements are about whether the meaning has been distorted, and the answer is basically no.
But the writing styles have become more similar. In the validated results, the variance of writing complexity has decreased by about 21% to 50%, and the largest drop comes from the "rephrase" type of prompts.
Even when the prompt mainly requires improving grammar, the variance of most model and dataset combinations still shrinks. The original wording of that prompt also allows "other necessary modifications", and the requirement to improve grammar is not purely error correction. There are also exceptions: when Llama 3 rewrites arXiv papers, the variance slightly increases instead, which is not statistically significant.
The entire rewriting process was completed by researchers using models in batches, simulating the common polishing scenario, not the editing records of real users. In reality, people will reject AI suggestions, and they will also revise the text again after the model modifies it.
Conceptual illustration. The proportions are taken from the significant result range of Study 1b in the paper, and the semantic retention rate is based on the embedding similarity standard.
Fading authorship cues
Which cues are exactly fading?
There is a mature practice in computational linguistics: train a classifier with annotated corpora, and let it guess the labels of the writers from the text, such as age group, gender, political party, Big Five personality, empathy level, and moral tendency.
The training materials are readily available, including speeches from more than 8,000 members of the U.S. Congress with party labels; there are also conversation records attached with empathy assessments. The corpus used to guess personality is exactly the 2,348 stream-of-consciousness essays collected by Pennebaker and King in 1999. The batch of essays that were used to establish the "correlation between function words and personality" more than 20 years ago has now become the material to test "whether AI rewriting will wear away this correlation".
The team first trained the classifier on the original human-written text, and then applied it to the text rewritten by AI.
The prediction performance of the six types of labels declined across the board, and all six declines were large enough that they could not be caused by sampling error (p<.001). The largest drop was in age: the F1 score (an accuracy index that considers both "how many of the guesses are correct" and "how many of the targets to be guessed are found") dropped from 0.351 to 0.260; the prediction of gender dropped from 0.694 to 0.623, and the prediction of political party dropped from 0.664 to 0.591.
The average drop is 6 percentage points, and the prediction performance for all six types of labels after rewriting is still significantly higher than the level of random guessing. The cues are fading, not disappearing.
The team then looked at which cues no longer work. The result is that they do not disappear together.
The correlation between gender and negative emotion words that exists in the original text basically disappears after rewriting; the correlations between extraversion and pronouns, loyalty and "friend"-related words, age and future-oriented words are all weakened. But the correlations between neuroticism and negative emotion words, purity and religious-related words, gender and social-related words still hold after rewriting.
What is more noteworthy is the direction. After the model's modification, the classifier is more likely to judge that the articles are written by older, male people, people with higher moral concern, lower empathy, and less extraversion. The most biased result is in political party prediction: when the real author is a Republican, the rewritten articles are often judged to be written by a Democrat. The offset is the direction perceived by the classifier, which does not mean that the authors have really become these types of people.
In other words, rewriting not only weakens the cues as a whole, but also concentrates misjudgments in several fixed types.
The data comes from Study 2 of the paper. The bars represent the average F1 (an accuracy index that considers both precision and recall), and all decreases are p<.001.
Fading cues are not all bad news, as the paper itself notes. Academia and industry have long used the same set of cues to infer the attributes of strangers from public texts. As the cues become unreliable, the threshold for monitoring and discrimination has also been raised a little.
The cost goes in the opposite direction. Several downstream scenarios listed in the paper, including psychological state screening, empathy research, and emotion interpretation in customer service, are all built on the premise that "text can reveal a person's traits". Once this premise loosens, the foundation of these practices will also loosen, which is exactly the practical risk the authors are concerned about.
In March this year, three members of the team published a perspective article in *Trends in Cognitive Sciences*, expanding the discussion from writing style to expression and thinking. Their reasoning is that the training process naturally tends to make those expressions that appear frequently and are applicable to most contexts stand out, while the minority expressions are submerged.
That article is a theoretical synthesis, not new empirical research. It raises a further, unproven question: is it not just the writing style that is being flattened, but also the way of looking at problems?
Morteza Dehghani, senior author of the paper, dual-appointment professor of psychology and computer science. Source mola-lab.org
The weight of small words
Back to the 12 disputed essays.