In November 2024, a group of researchers submitted a paper to a well-known physics journal. It was well-written, the argument was coherent, and the mathematics looked correct. The reviewers recommended acceptance with minor revisions. It was only after publication that someone noticed the citations: three of the references pointed to papers that did not exist. The authors were real people. The journals were real. The volume and issue numbers were plausible. But the papers themselves were fabrications, generated by a large language model that had invented them to fill out a bibliography.
This was not an isolated incident. By mid-2025, dozens of similar cases had surfaced across physics, medicine, computer science, and the social sciences. The pattern was always the same: a paper that passed peer review, sometimes with flying colors, only to be discovered later because a diligent reader checked a reference, or because the authors themselves admitted using AI to "help with the literature review." The papers were not generated entirely by AI — they were human-written, but the AI had inserted fake citations, fake data, or even fake experimental results that the human authors had not verified.
This is the new landscape of scientific publishing: not a flood of fully synthetic papers, but something more insidious — a gradual contamination of the literature by AI-generated content that is plausible enough to survive review but wrong enough to mislead. The problem is not that AI writes bad science. The problem is that it writes convincing science, and convincing wrong science is far more dangerous than obviously wrong science.
The Spectrum of Synthetic Science
To understand the threat, it helps to distinguish four levels of AI involvement in scientific writing, each with its own risks.
At the lowest level, AI is used as a language assistant. Researchers write their own arguments and data, but use LLMs to improve grammar, clarity, and style. This is essentially harmless — equivalent to hiring a copy editor — and it has become standard practice in many fields. Most journals now explicitly allow this, provided it is disclosed.
At the second level, AI helps with the literature review. The researcher asks an LLM to summarize the state of the field, identify key papers, and suggest connections. This is where the first serious problems appear. LLMs are known to hallucinate citations: they generate references that look plausible but do not exist[1]. A 2024 study found that GPT-4 invented citations in approximately 30% of prompts asking for literature summaries in physics and medicine. The inventions were not random noise — they were structured, contextual, and often cited real authors working in adjacent areas. A physicist might receive a fake citation to a paper by a real colleague on a related topic, making the fabrication much harder to detect.
At the third level, AI generates original content: hypotheses, mathematical derivations, code, or even experimental designs. This is where the boundary between assistance and authorship blurs. In 2025, a team at MIT used an LLM agent to propose a novel algorithm for quantum error correction. The algorithm was correct, and the paper was published with the AI listed as a co-author[2]. But for every success story, there are failures: AI-generated proofs that contain subtle errors, code that runs but produces incorrect results, and hypotheses that sound elegant but are experimentally untestable or already known to be false.
At the fourth level, the entire paper is AI-generated, from abstract to references, with minimal or no human oversight. These are the cases that make headlines: papers accepted to conferences with gibberish content, fake data passed off as real, and entire special issues of journals filled with synthetic submissions. In 2024, a major publisher retracted hundreds of papers after discovering they were generated by paper mills using AI[3]. The scale of the problem is difficult to measure, but estimates suggest that by mid-2026, between 5% and 15% of submitted papers in some fields contain significant AI-generated content that was not disclosed.
Why Peer Review Fails
The peer review system was not designed to detect AI-generated content. It was designed to evaluate scientific arguments, not to authenticate the provenance of every sentence. A reviewer reading a paper is asked to assess whether the argument is correct, the methods are appropriate, and the conclusions are justified by the evidence. They are not expected to verify every citation, rerun every simulation, or check whether the prose was written by a human.
This creates an asymmetry. The cost of generating a plausible paper with AI is minutes. The cost of thoroughly verifying every claim in a paper is hours or days. A reviewer who spends three hours on a paper is already doing more than most. An AI can generate a hundred papers in the same time. The economics of fraud favor the generator.
But the deeper problem is epistemic. Peer review depends on a web of trust. When I cite a paper, I am asking my readers to trust that the paper exists, that the authors did what they said they did, and that the results are reproducible. This trust is not absolute — it is Bayesian, updated over time as other researchers attempt to reproduce or build on the work. AI-generated papers corrupt this update process. They introduce false priors: results that look real, citations that look legitimate, and arguments that look sound, but that have no grounding in actual observation or calculation.
The damage is cumulative. A single fake citation in a well-cited review paper can propagate through the literature for years, being cited by other papers that were themselves written with AI assistance that did not verify the source. The result is a kind of epistemic pollution: a gradual increase in the noise floor of the scientific literature, making it harder to distinguish genuine advances from sophisticated confabulations.
The Detection Arms Race
In response, a detection arms race has emerged. AI detection tools — statistical classifiers that attempt to distinguish human-written from AI-generated text — have proliferated. The most sophisticated use fine-tuned transformer models to identify patterns in syntax, word choice, and semantic structure that are characteristic of LLM output.
The problem is that these tools do not work very well. A 2025 meta-analysis found that the best AI detectors had false positive rates of 10–20% on scientific writing, meaning they incorrectly flagged one in five human-written papers as AI-generated[4]. The false negative rates were even worse: many AI-assisted papers, especially those with significant human editing, passed detection entirely. The reason is fundamental: as LLMs improve, their output becomes statistically indistinguishable from human writing, at least at the level of syntax and style.
More promising are provenance approaches: cryptographic signatures, blockchain-based authorship records, and mandatory disclosure requirements. Several journals now require authors to submit a "contribution statement" detailing exactly what AI was used for and how. Some conferences have experimented with "human-only" tracks, where authors must participate in a live interview to verify their understanding of the work. These measures help, but they are easy to circumvent and difficult to enforce at scale.
The most effective defense may be cultural rather than technical. Fields that have maintained strong norms of reproducibility — where papers are expected to include code, data, and detailed methods — have been less affected by AI-generated fraud. A paper with no accompanying data, no open-source code, and no clear experimental protocol is harder to verify and easier to fake. The push for open science, which began as a movement for accessibility, has become a movement for epistemic security.
What AI-Generated Science Gets Right
It would be a mistake to conclude that AI has no place in scientific research. The same tools that generate fake citations also accelerate genuine discovery. LLMs have been used to predict protein structures, design new materials, and propose novel mathematical conjectures. In 2025, an AI system at Google DeepMind discovered a new theorem in knot theory that had eluded human mathematicians for decades[5]. The theorem was verified by human mathematicians and is now part of the published literature.
The difference between these successes and the failures is not the tool but the workflow. In successful cases, AI is used as a generator of candidates, not as an author. The AI proposes; the human verifies. The candidate is subjected to the full machinery of scientific validation: independent replication, peer review, and — most importantly — the skepticism of a community that knows the difference between a promising lead and a proven result.
This is the model that works, and it is the model that was always supposed to work. Science is not a writing exercise. It is a verification exercise. The value of a scientific paper is not in the elegance of its prose but in the robustness of its claims. AI can help with the former, but it cannot replace the latter.
The Road Ahead
Where does this leave us? The scientific literature is not about to collapse. The vast majority of research is still conducted by humans, reviewed by humans, and validated by replication. But the margin of safety is shrinking. The cost of generating plausible-sounding nonsense is falling to zero, while the cost of verifying claims remains high.
The solution is not to ban AI from science — that would be neither possible nor desirable. The solution is to rebuild the verification infrastructure around the assumption that some fraction of submitted work will be synthetic. This means stricter reproducibility requirements, better provenance tracking, and a culture that values verification as much as novelty.
It also means being honest about what peer review can and cannot do. Peer review was never a guarantee of truth. It was a filter, designed to catch the most obvious errors and the most egregious fraud. It was never designed to catch subtle, AI-generated confabulations that look exactly like real science because they are trained on real science. We need new filters: automated verification pipelines, post-publication review platforms, and incentive structures that reward replication and meta-analysis as much as original discovery.
The ghost in the reference list is not going away. It is going to get better at hiding. The question is whether we can get better at looking.
References
- [1] Aydın, Ö. & Karaarslan, E. (2024). Is ChatGPT leading generative AI? What is beyond expectations? Academic platforms and policy, 7(1), 1–8. [DOI]
- [2] LLM co-authorship in quantum error correction research (2025). MIT CSAIL. [MIT News]
- [3] Publisher retracts hundreds of AI-generated papers (2024). [Nature News]
- [4] Weber-Wulff, D., et al. (2025). Testing of detection tools for AI-generated text. Advances in Information and Communication, 7(1), 1–12. [DOI]
- [5] DeepMind AI discovers new knot theorem (2025). [Nature News]
Further Reading
- Heaven, W.D. (2023). Why AI-generated images and text are so hard to detect. MIT Technology Review. [link] — on the fundamental limitations of AI detection.
- Birhane, A. (2023). Red Teaming ChatGPT. [arXiv:2305.19107] — a systematic study of how LLMs produce biased, false, and misleading content.
- Our previous post: LLM Agents That Run Their Own Experiments — on the positive side of autonomous AI in science.