HomeArticle

The use of AI to write academic papers is no longer a novelty, and AI dedicated to scientific research has begun to compete for the "next ship ticket".

韦韦_wiwi2026-09-10 19:52
Scientific research AI is evolving from merely "helping you write" to deeper engagement in the in-depth research process.

On August 20, 2026, *Nature* reported a study: after analyzing biomedical papers indexed in PubMed Central, the authors estimated that by December 2025, nearly 90% of papers showed traces of AI-assisted writing. The estimated proportion for the whole year of 2025 was about 77%, 52% in 2024, and only about 20% in 2023.

There are two points to clarify first to avoid confusion: First, "77% for the full year of 2025" and "nearly 90% by December 2025" are not contradictory. The former is the annual average, while the latter is the peak of the single month at the end of the year — the proportion of AI-assisted writing climbed steadily throughout the year, rising to around 90% by the end of the year. Second, although this study was published in August 2026, its data only covers up to December 2025. This is not because the data is outdated, but because such text mining studies rely on the full archiving of papers by PubMed Central, and archiving itself has a lag of several months to more than half a year. Researchers also prefer to use "full calendar years" for year-on-year comparison, so there is no credible official statistic for 2026 data yet — as of press time, this is still the latest version available.

This news has been recounted many times in recent weeks, and the numbers themselves are indeed easy to misinterpret, so it is worth clarifying a few more points.

First, this study is a preprint posted on arXiv on August 12, 2026, and has not yet undergone peer review. It detects changes in word frequency in the text of papers — which words and expressions suddenly became more common after the emergence of ChatGPT — rather than judging one by one "whether this paper was written by AI". When broken down by different sections, the signal intensity varies greatly: 68% for the Discussion section, 67% for the abstract, 59% for the introduction, 46% for the results section, and only 32% for the methods section. In other words, AI appears more in the "linguistic layer" of the paper, rather than in the data and methods themselves.

Second, a peer-reviewed study (by Kyle Siler) in the same period used different methods and journal samples, and obtained much more conservative figures: 12% in 2023 and 57% in 2025. Different methods can lead to a difference of 20 percentage points in the conclusion, which itself shows that the answer to the question "how many papers use AI" depends to a large extent on how the question is asked and how the detection is carried out.

Therefore, "AI participating in academic writing" and "letting AI complete the research for you" are two different things. The former is already an established fact, while the latter touches the boundary of research authenticity and academic integrity.

But what this news really makes the industry think about is another layer: when most researchers are already using AI to write papers, what else can companies that specifically sell "AI paper writing" services sell? The answer may not lie in the act of "writing" itself, but in the longer process behind "writing" — the complete process of a research from raising a question to finalizing the manuscript. Whoever can truly embed into this process will have the opportunity to retain users.

Three years ago, as long as a product could upload PDFs, summarize literature, polish English, and generate a structurally complete first draft according to the theme, it was enough to make researchers feel fresh. If you take these functions out to sell today, what users think of is mostly not paying, but directly opening ChatGPT, Claude or Gemini. File reading, web search, long context, code execution, data analysis, LaTeX generation — these functions that originally required vertical scientific research AI to package separately are gradually becoming native capabilities of general-purpose models.

More people are using AI to assist scientific research than ever before, but it is increasingly difficult to make money solely by "generating papers". Scientific research AI is starting to look for its next ticket to ride.

"Helping you write papers" is becoming a difficult business to run

Scientific research was originally one of the most suitable scenarios for vertical AI tools to grow.

A researcher deals with many digital tools every day: finding papers in academic databases, managing literature with Zotero, processing data with Excel, R or Python, writing the main text with Word or Overleaf, and finally submitting to the publisher's system. These softwares are not naturally connected to each other, and all information is manually transferred by researchers themselves.

In the first two years after the emergence of large models, opportunities were everywhere in this gap: if you couldn't write Prompts, there were products that turned Prompts into buttons; if the model couldn't read dozens of papers at a time, there were products that specially parsed PDFs and built indexes; if the model didn't know how to handle academic citations reliably, an extra layer of literature retrieval was added outside; if you didn't know LaTeX, templates, format conversion and typesetting were supplemented.

The first batch of scientific research AI businesses were largely supported by a "capability gap" between general models and real scientific research needs: models can theoretically do many things, but ordinary researchers can't use them, or they are too troublesome to use. Vertical products wrap the capabilities in a shell, and put on an interface that is closer to the habits of scientific research, then the business is established.

In the past few years, these products have basically found their positions along different links: Consensus serves as the entry for academic search, Elicit cuts in from literature review and evidence extraction, Scite has long operated the citation relationship between papers, and products such as Paperpal and Writefull enter the daily workflow of researchers from language polishing.

These demands themselves are not problematic. The trouble is that this capability gap between vertical products and general models is narrowing rapidly.

Foundation models natively start to support files, web search, longer context, and code execution and data analysis have gradually become standard configurations. The emergence of Agent has taken the model a step further: from "answering a question" to "continuously completing a set of tasks". The capabilities that used to be packaged by scientific research AI companies themselves are being eaten away layer by layer by model vendors.

This is not the first time this has happened in the AI industry. Summarization used to be an independent product, and later became a button; translation, rewriting, PDF conversation, web search have all gone through similar paths. Every time the model is upgraded, the application layer has to re-answer the question "why should I still exist".

Scientific research AI has now reached this point — moreover, the figure of "nearly 90% of papers showing signs of AI-assisted writing" is not necessarily good news for scientific research AI companies. The more accustomed researchers are to using AI, the lower the cost of educating users will naturally be; but a person who is already used to ChatGPT, Claude or Gemini will be more difficult to be persuaded to pay by another chat box that "understands papers better".

The real problem is no longer whether researchers will use AI, but: apart from the model itself, what else can scientific research AI companies provide?

From search, citation to Agent, scientific research AI is starting to move away from "paper writing"

In the past few months, changes in several products have provided some clues.

Elicit is a typical example. It was first known to researchers for literature retrieval, evidence extraction and systematic review; but since 2026, it has clearly moved beyond "finding papers". In May, it further aligned its systematic review function with specifications such as PRISMA 2020, and took "reproducible, traceable, auditable" as its key capability; in June, its Research Agent began to process research materials such as gene sequences, omics data, and microscopic images, and perform data analysis and plotting; in July, APIs and MCP were opened, allowing other Agents and scientific research workflows to directly call its capabilities; in August, a Research Agent for complex scientific decision-making was launched, plus collaborative sessions, Artifact Editing and shared projects — these steps happened almost in the same month as the *Nature* report, which is not a coincidence, but rather the same reality that "AI has widely entered scientific writing" emerging simultaneously on the product side and public opinion side.

The meaning of this product trajectory is clear: it is no longer satisfied with helping researchers find a few papers, but wants to enter the ongoing process of a research itself.

Scite takes another path. Its value does not lie in making another chatbot, but in the paper citation relationships accumulated over many years — whether subsequent papers are supporting, refuting, or just mentioning a certain study. The stronger the general model is, the more valuable this structured academic data becomes, because the model may be replaced at any time, but a reliable data source will not depreciate just because GPT is upgraded once.

Similar changes are spreading across the entire market: some are seizing academic search, some are seizing citations and evidence, and some are trying to take over the complete workflow.

Some younger domestic products are also exploring in this direction, and Textira is a sample worth expanding on.

Its positioning is very straightforward — the sentence on the homepage of its official website is "Turn the experiment into a paper you can submit" — but the real product design is not "enter a topic and generate an article", but make the four stages of Topic, Experiment, Results, and Draft into a continuous pipeline: first lock the research methods, variables and evaluation indicators, then move forward, the original experimental records will be automatically organized into materials for tables, charts and discussion paragraphs, instead of being manually transferred when the paper is almost finished; no matter which step you are writing at, you can directly trace back the current sentence to the corresponding literature or experimental record, instead of adding citations after finishing writing. At the finalization stage, Textira will automatically adjust the column width, anonymous review requirements and typesetting specifications according to the target journal (conference formats such as NeurIPS and ICML are all supported), and finally export four formats of Word, LaTeX, PDF and PPT at one time, all of which can be further edited.

This route is quite different from Elicit and Scite: the latter two are more like "digging deep into a certain link", while Textira is closer to putting the complete link of a researcher from topic selection to submission into the same workbench. It is too early to draw a conclusion on whether this path can work, but at least it provides a different entry angle from literature retrieval and citation relationships.

Textira official website homepage (September 2026, source: textira.com)

What really deserves a closer look here is not that "AI can write several more chapters". On the contrary: general models are already good enough at writing, and it makes little sense to continue stacking generated word counts. The more difficult problem is whether we can connect the research processes that were previously scattered in different softwares.

For example, when a new result comes out of an experiment, a researcher may have to throw the data into Excel or Python to redraw the graph first, then copy the numbers back to Word; once the results in the paper change, they have to go back and check whether the abstract and Discussion have been modified accordingly; to add citations, they have to open Zotero again; when preparing for submission, they have to go to Overleaf or the publisher's system to readjust the format. These tools do not know what happened to each other, and it is the researchers themselves who maintain the relationship between them. In a sense, researchers have long served as the most expensive API between scientific research softwares.

This is also where Agent is truly attractive to scientific research: if AI can understand that an experiment has changed, and know which charts, numbers, and text paragraphs need to be updated accordingly, it will save not only the time of helping to write two paragraphs of Introduction, but also a large amount of mechanical labor throughout the several-month research cycle.

But once scientific research AI reaches this step, a more intractable problem than efficiency emerges: AI can help researchers do more and more things, but who can prove that these things are done correctly?

The real difficulty of scientific research AI is not automation, but whether people can still trust it after automation

This is the most fundamental difference between scientific research Agents and most office Agents.

When programmers ask AI to modify a piece of code, at least there are compilation, testing and running results to help judge whether the system has been broken. Scientific research is not that simple: a successful LaTeX compilation only means that the paper can generate a PDF, not that the numbers inside are real; when AI finds a real existing paper, it does not mean that this paper really supports the sentence it is writing; when a set of statistical programs runs successfully, it does not mean that the experimental design itself is tenable. The model may even write a paragraph of Discussion that is full of professionalism and logically coherent, but some causal explanation in it has long exceeded the scope that the experimental data itself can support.

Moreover, the more smoothly AI writes, the more difficult it is to discover such errors sometimes.

Therefore, "scientific research Agent" cannot be simply understood as a more automatic paper generator. If the product only expands the original one-time generation into dozens of automatically executed steps, the higher the degree of automation, the wider the spread of errors after they occur — a wrong citation mixed into the research context will then affect the literature review; a wrong number entering the Results may further appear in the abstract, Discussion and charts. When an Agent can autonomously complete more than a dozen steps of tasks, researchers may already find it hard to tell exactly at which step the error was introduced.

This is also why the next round of competition for scientific research AI may have an easily overlooked premise: automation must bring stronger verifiability at the same time. This is why products like Elicit are increasingly emphasizing "reproducible, traceable, auditable" — a truly valuable scientific research system cannot only answer "what the answer is", but also keep track of "where the answer comes from".

This is also a problem that all AI products trying to enter the scientific research workflow will face sooner or later: after putting experiments, literature, main text and LaTeX into a continuous environment, what really matters is not how much content the Agent can modify, but whether researchers can clearly see what it has modified, what data it has used, what the citation basis is, and whether the research facts have undergone undesired changes before and after the modification. If this relationship cannot be retained, the so-called scientific research Agent is actually just spreading the risk of "AI generating papers" across the entire workflow.

Textira's approach to this point gives a specific answer, rather than staying at the principle level. Its logic is "results take precedence over generation": real records such as experimental logs and data tables are always the primary source, and AI is responsible for organizing them into tables, charts and discussion texts. Once the content generated by the model does not match the original data, the system will directly mark the conflict and return it to the author for judgment, instead of letting the model "cover it up" by itself. Methods, variables and evaluation indicators are also locked before the first run, and no matter how the text is rewritten later, the underlying experimental design itself will not be quietly changed by the model. In other words, what it wants to do is not "let AI write more boldly", but "make the content written by AI always verifiable" — which is exactly the implementation of the aforementioned reproducible, traceable, auditable in a specific product.

Once evidence traceability, data provenance, and modifiability check are required, the product difficulty of scientific research AI is completely different — it is no longer just a model invocation problem, but requires understanding the relationship between papers, data, experiments, citations and charts, knowing which content can be rewritten by AI and which facts cannot be freely generated by the model, while leaving the final judgment right to the person who is really responsible for the paper.

This is also where academic integrity really needs to be discussed in the AI era, and the problem should not be simplified as "allowing or not allowing the use of AI". Retrieval, language editing, coding, data analysis — AI has been embedded in most links of scientific research, and it is difficult to describe the reality with just two words "used" or "not used". What is more worth asking is: whether researchers have really completed this research, what exactly AI has done in it, whether these operations can be explained and checked, and whether the signatory can still be responsible for every key conclusion in the paper.

AI can be a scientific research tool, but it cannot become the bearer of academic responsibility. This boundary will not disappear because the Agent becomes stronger; on the contrary, the more things AI can do, the more important this boundary becomes.

The next ticket to ride may be far more valuable than "AI writing papers"

From this perspective, the context of this round of changes in scientific research AI is actually very clear.

The opportunities in the first stage came from the sudden scarcity of generative capabilities: whoever could make AI better at reading papers, writing abstracts, and polishing English would have the chance to win users. But this scarcity is disappearing rapidly — it is no longer new that models can write, it has become common that they can search, and capabilities such as analyzing files, running code, and adjusting tools have gradually become basic capabilities.

The truly scarce things have shifted to several other aspects: high-quality scientific research data, reliable evidence