HomeArticle

AI ghostwriting of academic papers exposed? When the editor-in-chief of a top journal asked questions, the author could not answer any of them.

机器之心2026-09-26 12:39
Did you write the paper? The editor-in-chief of TMLR quizzed 10 authors in person.

Edited by | Panda

You have written a paper, and now you are asked three questions: What is the problem setting of this paper? What does this symbol stand for? Which part of the main text supports the conclusion mentioned in the abstract?

If you did write the paper yourself, these questions must be very easy, and you can even answer them without any preparation.

However, in a recent small-scale experiment conducted by TMLR (Transactions on Machine Learning Research), a journal focused on machine learning, the authors of three papers failed to answer even these basic questions.

On September 16, TMLR published an article titled "Asking Authors About Their Own Papers" on its official blog, written by Nihar B. Shah, one of the journal's co-editors-in-chief from Carnegie Mellon University (CMU). During his two-week rotating tenure as editor-in-chief, he selected 10 submissions that were originally supposed to receive desk rejection. Instead of rejecting them directly, he arranged to "chat" with each of the authors one by one. As a result, all 10 papers were still rejected.

https://medium.com/@TmlrOrg/asking-authors-about-their-own-papers-3d2e04e5dee0

Gautam Kamath, who is also an editor-in-chief of TMLR, reposted the news on X and called it a "heroic experiment", saying it confirmed the long-held suspicion of many people: some submitters actually do not know what is written in their own papers.

Spot-check of 10 papers pending desk rejection, all rejected

According to Shah's description, this experiment took place during his tenure as rotating editor-in-chief from August 14 to 28, 2026. He informally selected 10 papers from the pool of submissions pending direct desk rejection, and sent a short message to the authors via the peer review platform OpenReview: One of the editors-in-chief hopes to chat with you before the paper is sent for review to better understand this work. If it is convenient for you, please send an email to inform the available time for the meeting.

After the message was sent, things quickly diverged. The author of one paper directly withdrew the submission; one author replied that he had too many things to handle and had no time for the call; the remaining 8 papers arranged the meeting, but one of the authors did not show up at the appointed time.

In the end, the authors of 7 papers met with Shah. They had different identities, including undergraduates, postgraduates, doctoral students, university teachers and independent researchers. Most of these papers were single-author works, but not all of them.

Shah's questions were divided into two categories: the first category was basic questions about the problem setting, symbols and the results claimed by the paper, and the second category was detailed questions about specific technical formulas, theoretical results and experimental design choices.

Among the 7 meetings, only the author of 1 paper answered all the questions. The authors of the other 3 papers could clarify the high-level ideas, but ran into difficulties when pressed for technical details. The worst case was the other 3 papers, whose authors could not even answer the basic questions, and all 3 of these papers were single-author works.

Shah wrote that two of these authors seemed to have almost no substantial understanding of the content of their papers, and the other could not find where several key results claimed in the abstract were presented or supported in the main text.

The only paper whose author answered all questions correctly still failed to pass the review. During the review, Shah found that there was a major error in one of the main conclusions of the paper, which the author later admitted. TMLR directly rejected this paper, but allowed the author to resubmit after correcting the error or narrowing down the conclusion. The remaining 9 papers were directly rejected and no resubmission was allowed.

Two episodes after the meetings

Shah also recorded two incidents in the article, which are rather darkly humorous to read.

The first incident: The authors of two papers could not answer the basic questions during the meeting, but sent him written replies after the meeting. Shah submitted these two emails to the AI text detection tool Pangram, and both were judged as "100% AI".

The second incident: In another meeting, an author tried to introduce the analysis method he used to Shah, but while explaining, he inadvertently described a complete p-hacking process, that is, the practice of repeatedly adjusting the analysis method until the data becomes "statistically significant".

It needs to be clarified that Shah himself does not reject AI. He confessed in the article that in order to finish reading 8 papers within two weeks, he also used LLM to assist his understanding, and even learned some concepts he was not familiar with before. What he cares about is not whether the authors use AI, but whether the authors can take responsibility for the content that bears their names.

Why do this?

TMLR was founded in 2022 and operated by the team behind JMLR. It is known for its peer review standard that "only focuses on whether the conclusion is supported by evidence, and does not take novelty and SOTA as reasons for rejection". All its reviewers, Action Editors and editors-in-chief are unpaid volunteers, which makes it particularly difficult to cope with the surge of submissions this year.

According to the announcement released by TMLR in June, the number of submissions in the past year has tripled, among which single-author submissions have even skyrocketed to 13 times the original number. The editorial department has even seen someone submit 5 papers in one day. Another set of figures given by Shah in the latest article is more intuitive: in 2023, the direct desk rejection rate of TMLR was about 6%, and now it has risen to about 53%, which means that more than half of the submissions cannot enter the external review process.

Facing this situation, TMLR has launched a series of intensive measures this summer.

The first measure is setting an annual submission quota for each author. Different from the fixed upper limit of "maximum N papers per person" adopted by many conferences, TMLR uses a "harmonic" quota rule: the quota consumed by each author for one paper decreases as the number of co-authors increases. The specific parameters are: authors who only submit single-author papers can submit up to 2 papers per year; if all submissions are co-authored by 9 people, they can submit up to 9 papers per year; active reviewers and action editors have double the quota. This rule has been implemented since July 1. TMLR explained in the announcement that the reason for not splitting the quota evenly according to the number of authors is to prevent people from "adding nominal co-authors" to obtain more submission quotas.

The second measure is introducing AI peer review. TMLR announced in July that each submission will be accompanied by an AI-generated review opinion in addition to the regular manual review, which only evaluates the "reliability" of the paper, that is, whether the conclusion is supported by accurate, clear and convincing evidence, without making subjective judgments or giving acceptance suggestions. The final decision is still made by the action editor. After evaluation, TMLR selected the AI reviewer from CSPaper.

The third measure is incorporating "clear writing" into the acceptance criteria. On August 28, TMLR revised the original "audience interest" standard, explicitly requiring papers to clearly convey their findings to readers. The announcement stated directly that current AI writing is often verbose, full of jargon and difficult to understand; if a paper is almost entirely generated by AI with little human participation, at least with the capability of existing AI systems, it is very likely that it cannot meet this standard.

Shah concluded at the end of the article that this series of interviews has made the editorial department more confident in the existing direct desk rejection process, and the quota system and the emphasis on clear writing have also proved to be very helpful.

Not only TMLR

The situation encountered by TMLR is not unique. Since the beginning of this year, several important publishing venues in the AI field have been grappling with the same problem.

In early June, the official blog of NeurIPS disclosed that among the 971 submissions to the NeurIPS 2026 Position Paper track, 273 (28.2%) were judged by Pangram to have an AI score of 100%.

https://blog.neurips.cc/2026/06/02/ai-generated-papers-in-the-neurips-2026-position-paper-track/

This track explicitly requires that papers must be "substantially written by humans", and AI can only be used for peripheral modifications such as text polishing. NeurIPS also pointed out that the growth of AI writing is comprehensive: in the "Evaluation and Datasets" track, the number of papers with a Pangram score of no less than 90% has increased by more than ten times from 2025 to 2026. However, NeurIPS also admitted that the detection results are sensitive to parameters. After using a medium-sized detection window, the proportion of papers with an AI score between 90% and 100% will drop from 42.7% to 12.7%.

Earlier in mid-May, Thomas G. Dietterich, head of the Computer Science section of arXiv, announced a clear penalty rule on social media: if there is "conclusive evidence" that the author has not checked the output of large language models in a submission, such as hallucinated references or chatbot dialogue sentences left in the main text, all authors will be banned from submitting to arXiv for one year, and subsequent submissions must be accepted by a formal peer review venue before they can be published on the platform. Dietterich emphasized that signing on a paper means that every author is responsible for the entire content, no matter how the content is generated.

The relevant data is equally striking. According to a *Lancet* study by Columbia University, among the biomedical papers included in PubMed Central, the proportion of papers containing at least one fictitious citation has risen from about 0.04% in 2023 to about 0.57% in early 2026.

https://www.nursing.columbia.edu/news/nearly-3-000-peer-reviewed-medical-papers-have-fake-citations-columbia-nursing-ai-assisted-audit-finds

From "checking the text" to "checking the author"

Shah admitted in the article that this experiment took him 20 to 25 hours, and he only processed 8 papers. Facing the current scale of submissions, this approach is difficult to promote widely.

He also mentioned that his team is researching scalable solutions.

Just a few days ago, Shah, together with Justin Payan, Bálint Gyevnár and Atoosa Kasirzadeh from CMU, published a paper proposing an evaluation method called greCAPTCHA. Half of the name is taken from GRE, the American graduate admission test, and the other half is taken from CAPTCHA, the verification code that distinguishes humans from machines.

https://www.cs.cmu.edu/~nihars/preprints/greCAPTCHA.pdf