
“The academic record is being quietly haunted.” That is the central warning of a new preprint by researchers at Samsung and the University of Warsaw, highlighted in MIT Technology Review’s The Download newsletter on August 28, 2026. The paper describes a rising torrent of AI-generated study papers credited to fake researchers — synthetic scholarship that appears in journals and indexes as legitimate science, but was written by nobody.
What the preprint says
The preprint’s authors argue that generative AI has lowered the barrier to producing convincing academic manuscripts to nearly zero. Rather than a handful of clumsy submissions, they describe a flow of papers that carry plausible titles, structured abstracts, and invented author names, making them difficult for editors and reviewers to trace. The research has not yet been peer-reviewed, so its claims should be read with caution, but the warning matches what journal operators have been privately reporting since the rapid adoption of large language models.
The word “haunted” is carefully chosen. Ghosts accumulate silently. They are not easy to identify in a single glance, but their presence can distort everything that comes after. A fake paper indexed in a citation database can be cited by future work, magnify noise, and even be mistaken for evidence during systematic reviews. Removing that contamination later is far harder than preventing it in the moment.
The preprint’s authors write from two different worlds: Samsung is a global leader in consumer electronics and semiconductors, while the University of Warsaw has a long tradition in logic and computer science. That combination gives the warning weight, because it merges commercial AI development with academic methodology. It also suggests that the problem of fake research is being recognized across both industry and academia.
From paper mills to ghost authors
Academic publishing has struggled with “paper mills” for years — commercial operations that sell authorship slots or manufacture low-quality manuscripts against the payment of article-processing charges. Generative AI changes the economics of that industry in a fundamental way. A human-run mill has to pay writers, manage logistics, and avoid plagiarism detection. An AI-driven operation can generate an unlimited number of original-sounding drafts for almost no marginal cost, and create as many fictional authors as needed.

This new wave of ghost papers is also harder to detect because the text may pass conventional stylistic checks. The preprint from Samsung and the University of Warsaw is itself a signal that this is no longer a niche concern for librarians and journal editors: one of the largest technology companies in the world is now investigating the problem alongside a European research university.
The scale of the publishing ecosystem makes manual inspection impossible. Large multidisciplinary venues receive thousands of submissions per month, and peer reviewers are unpaid volunteers with limited time. A well-formatted AI-generated manuscript can easily slip through if it follows the conventions of the field. The preprint’s “torrent” metaphor is meant to evoke something that overwhelms existing defenses.
Why the AI community should care
For AI researchers, the contamination of the scientific record has three direct consequences. First, large language models increasingly train on scientific corpora; if those corpora include AI-generated ghost papers, the models themselves may incorporate hallucinated findings and false citations. Second, the credibility of legitimate AI-assisted research is at risk. A backlash against fake papers could lead to blanket bans on AI use in scholarly writing, punishing the many researchers who use it responsibly. Third, the problem exposes the limits of verification: if a paper’s claimed author cannot be found, and its data cannot be reproduced, what is the epistemic status of the work?
There is a parallel danger outside academia. The same methods that produce ghost papers can be used to generate fake benchmarks, fake security advisories, or fake press releases. The “haunting” described by the preprint is a rehearsal for a broader crisis of trust in machine-generated information generally.
For developers of AI systems, the lesson is uncomfortable: the technology that powers chatbots and coding assistants is the same technology that can produce sophisticated academic illusions. This duality has implications for model evaluation, data curation, and the kinds of safeguards that AI companies build into their products. It is also a reminder that research integrity is itself a machine-learning challenge.
Can the scientific record be defended?

The preprint does not promise a single fix, and at this point, no reliable one exists. Journals and preprint servers have begun to respond: several major publishers now require AI-use disclosure, and some have deployed automated detectors to screen submissions. But detection tools are imperfect; they can flag legitimate text written by non-native English speakers while missing sophisticated AI output that follows human-like phrasing.
Structural changes would help more than filters. Stronger author identity verification, such as ORCID-based checks integrated into submission systems, could make fabricated author lists harder to maintain. Requiring open data could make it easier to assess whether a paper’s numbers came from an actual experiment rather than a language model. And the underlying incentive structure — publish or perish metrics that reward volume over integrity — needs to shift, because that is the ecosystem in which ghost papers thrive. None of this will happen quickly, but the preprint’s warning makes the cost of inaction clear.
In addition to journal-level checks, preprint servers are experimenting with screening mechanisms and community rating systems that can flag suspicious submissions quickly. But the clearest defense remains human vigilance: editors who check author identities, reviewers who request source data, and readers who report obvious anomalies. The academic community’s own norms are its strongest firewall.
The long shadow of synthetic scholarship
The most immediate effect of AI-generated ghost papers is the erosion of trust in what gets published. When a scientist cannot assume that an author listed on a paper actually exists, the entire citation network that supports literature reviews, meta-analyses, and grant decisions becomes less reliable. The scientific method tolerates mistakes and retractions; it cannot easily absorb sources that were never authored by anyone.
There is also a generational dimension. Early-career researchers rely on the published record to learn their fields and choose directions. If that record contains a background hum of synthetic papers, newcomers may internalize false findings or waste time chasing nonexistent results. The cost of the haunting is paid most heavily by the next generation of scientists.
For now, the Samsung–Warsaw preprint is a useful early warning, and it received a kind of ironic honor as MIT Technology Review’s “quote of the day” on August 28, 2026. The challenge for the coming months is to transform that warning into action. AI-generated ghost papers are a natural consequence of powerful generative tools meeting an unreformed publishing culture. Both the tools and the incentives are here to stay, which means the academic record is likely to be haunted for some time — unless the research community decides to turn on the lights.
댓글