The AI ‘Ghosts’ Contaminating Academic Publishing

404 Media · collected 2026-08-28 · by Emanuel Maiberg
Read the original at 404 Media ↗

Summary

A new preprint research paper from Samsung and the University of Warsaw has identified hundreds of names that large language models repeatedly produce when generating experts in certain fields, including Elena Vasquez and Marcus Chen. These "ghost authors" appear as co-authors on AI-generated academic papers, articles, and books, with some being used to backdate fake publications. The researchers found 1,655 records claiming nonexistent journals with fabricated publication dates on the Zenodo repository. According to Michał Brzozowskim, lead author of the paper, searching for these names online revealed other instances of AI-generated personalities, including a fake "executive director Dr. Elena Vasquez" mentioned in a Facebook rumor about a nursing job incident.
Written by the local model on 2026-08-28, using this article's own text rather than the other coverage of the same event (that is the story summary below).

Signals How these are calculated →

Claims extracted
28
claim-shaped sentences
Uncertain
7%
2 of 28 hedged
Leaning
not political
takes no side on a contested political question
Publisher trust
94.8
red-flag proxy, not a credibility rating
Outlets on this story
1
Technology
Narrative spread
1
articles carrying this framing
Analyzed 2026-08-28 · how these are computed

AI analysis (generated at analysis time, not now)

Story summary

Elena Vasquez and Marcus Chen, two names that do not belong to real people, have been found to be prolific authors of AI-generated academic papers, articles, and books. According to a new research paper from Samsung and the University of Warsaw, these names have co-authored hundreds of documents across various fields, including volcanology, astronomy, and computer science. The researchers discovered that this is not an isolated phenomenon, but rather a result of large language models (LLMs) like ChatGPT, Gemini, and Claude relying on pre-programmed name patterns when generating experts in specific domains. For example, users have noticed that LLMs often default to using the name Marcus Chen for software developers. The researchers' findings suggest that AI-generated content is contaminating academic publishing, raising concerns about the validity of research and the potential for fake or misleading information to spread.

Written for “Academic Paper Integrity Threats” on 2026-08-31, grounded in this article and the 0 other(s) covering the same event.
Why this leaning score
This article does not take a side on a contested political question, so it has no leaning score. That is an answer rather than a gap: a match report or a rescue can be warmly or critically written without being left or right, and scoring it anyway is how approval of a subject gets recorded as a political position.
No political leaning scored for article 2944 · logged 2026-08-28

Story

📰 Academic Paper Integrity Threats
Technology · 1 article(s) covering the same event. This is the one the site leads with.

How this is being covered How these are calculated →

Article leaning vs. publisher reliability
Source leaning vs. consistency

Compared with similar articles

This article reads unscored and hedges 7% of its claims. Each row says how that neighbour differs.
AI Freezes The Scholarly Voice different event · 100%
Reason.com
⚖️ Leans strongly right 🔴 5% hedged 1 of 21 📰 publisher trust 94
“Article A describes a workshop about law professors using AI to write in their own style, while Article B discusses a research paper about 'AI ghosts' contaminating academic publishing with fake co-author names generated by large language models”
Beware the Ready-Made Op-Ed different event · 100%
The Dispatch
⚖️ leaning not scored 🔴 0% hedged 0 of 4 📰 publisher trust 96
“Article A discusses a research paper on AI-generated documents, while Article B talks about an op-ed piece written using AI by Stanley Druckenmiller”
Semafor
⚖️ leaning not scored 🔴 11% hedged 5 of 47 📰 publisher trust 96
“Article B describes a separate research paper on AI-generated academic publications, while Article A discusses the presence of AI-written articles in opinion pages”

Publisher

404 Media · 26 article(s) · 0 correction(s) detected
SignalValueWeight
Correction rate 0.000 0.4
Uncertainty density 0.104 0.25
Assertive mismatch rate 0.000 0.35
No corrections detected for this publisher. That may mean careful reporting, or simply that nothing has been checked.

Who wrote this

Emanuel Maiberg
3 article(s) here · 1 carrying a prediction
🔮 The paper, titled “The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing,” utilized a known phenomenon where certain LLMs will keep coming up with the same names in certain contexts.
2026-08-28 · assertive framing · The AI ‘Ghosts’ Contaminating Academic Publishing
🔮 Where I worked, primarily they would bring us boxes of books, or sometimes there's shuttles of books. What kind of books? All kinds of books.
🔮 But for the past few months I’ve been looking at a new and strange type of AI generated nonconsensual image on X that I think will fool a lot of people.
Also by Emanuel Maiberg
Nothing else under this byline is closely related to this article, so these are simply their most recent.

Topics

ChatGPT Claude Gemini Research Gold Zenodo

Subjects

Elena Vasquez PERSON · 5× ChatGPT ORG · 3× Zenodo ORG · 3× Brzozowskim PERSON · 2× Google Scholar ORG · 2× Marcus Chen PERSON · 2× Research Gold ORG · 2× Sam PERSON · 1× Samsung ORG · 1× the University of Warsaw ORG · 1×

Narrative

“Elena Vasquez and Marcus Chen have appeared as volcano experts, astronauts, thriller protagonists, podcast hosts, and academic co-authors across hundreds of independently produced AI-generated documents, never having lived,” a new preprint research paper from Samsung and the University of Warsaw said.
framing: assertive · carried by 1 article(s) · first seen 2026-08-28
🔮 The paper, titled “The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing,” utilized a known phenomenon where certain LLMs will keep coming up with the same names in certain contexts.
2026-08-28 · 404 Media
The AI ‘Ghosts’ Contaminating Academic Publishing · assertive framing

Claims (28 extracted, 2 hedged)

“Elena Vasquez and Marcus Chen have appeared as volcano experts, astronauts, thriller protagonists, podcast hosts, and academic co-authors across hundreds of independently produced AI-generated documents, never having lived,” a new preprint research paper from Samsung and the University of Warsaw said. asserted
paper → appear → Warsaw
The paper identified a number of names that co-authored hundreds of AI-generated academic papers, articles, and books. asserted
that → identify → papers
The authors don’t actually exist, but are instead names that large language models repeatedly produce when tasked with generating experts in certain fields. asserted
models → exist → fields
The paper, titled “The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing,” utilized a known phenomenon where certain LLMs will keep coming up with the same names in certain contexts. asserted
LLMs → title → contexts
For example, In June, Sam wrote about how ChatGPT, Gemini, and Claude were likely to use the name Elias Thorne in fiction they generated, and that the character was often a lighthouse keeper. asserted
character → write → fiction
Similarly, users noticed that if they ask ChatGPT to generate a software developer, their name will often be Marcus Chen. asserted
name → notice → developer
The researchers were able to show not only that LLMs often default to the same names, but that they produce “correlated character ensembles,” meaning some names were more likely to appear together. asserted
names → show → ensembles
Other names that were consistently generated by AI models include Elena Amara Okafor from Claude, Aris Thorne and Lena Petrova from Gemini, and Elara Voss from ChatGPT. asserted
that → generate → ChatGPT
Earlier this month, I reported a story about Research Gold, a company that offered what it claimed was human medical research, but that was in fact entirely AI generated. asserted
that → report → fact
The founder and lead methodologist for that company was named Elena Vasquez. asserted
founder → name → company
Research Gold removed Elena Vasquez from its site after I published the story. asserted
I → remove → story
Michał Brzozowskim, the lead author of the paper, told me that searching for these names on Google turned up other instances of AI generated personalties. asserted
searching → tell → personalties
For example, following the killing of Alex Pretti at the hands of U.S. Border Patrol agents in January, a rumor spread on Facebook that he was fired from his nursing job for misconduct allegations. asserted
he → follow → allegations
Snopes reported that the false statement was attributed to “executive director Dr. Elena Vasquez,” who does not exist. asserted
who → report → Vasquez
Brzozowskim was able to search databases of academic papers for the names they knew LLMs often generated. asserted
LLMs → search → names
“On Zenodo, a CERN operated repository that mints real DataCite DOIs, we identify 1,655 ghost-authored records claiming nonexistent journals with fabricated publication dates.” asserted
we → operate → dates
A DOI, or a Digital Object Identifier, is a string of characters and numbers used to identify academic papers. asserted
DOI → use → papers
The researchers saw that many of the papers authored by these AI names were backdated, meaning their publication dates were different from the date they were uploaded to Zenodo. asserted
they → see → Zenodo
Anyone with a free account can create a DOI on Zenodo, but the existence of AI generated papers with DOIs has impacts on other parts of the web and academic publishing. asserted
existence → create → web
“These [AI generated papers] carry real DOIs harvestable by any scholarly aggregator; the infrastructure for large-scale scholarly record contamination is already in place. asserted
infrastructure → generate → place
Ghost names additionally appear on ResearchGate, forming synthetic research groups with collaborators drawn from multiple model families, and are indexed without verification by Google Scholar and Semantic Scholar [...] The academic record is being quietly haunted.” asserted
record → appear → Scholar
Research Gate and Google Scholar are both aggregators of academic publishing that are likely to come up in search results. asserted
that → come → results
The researchers say that different LLMs and versions of those LLMs generate certain names so consistently, they think they can use them to determine the provenance of AI slop. asserted
they → say → slop
For example, the name Elena Vasquez was particularly common in content generated by Claude Sonnet 4, so papers that list her as an author were likely generated by that LLM. asserted
that → generate → LLM
However, Brzozowskim said that this level of accuracy might not hold for long now that AI generated content with these names is flooding the internet and feeding back into all AI models that are scraping the internet for training data. uncertain
that → say → data
The fact that AI academic publishing is struggling to deal with the load of AI generated content isn’t new. asserted
publishing → struggle → content
We’ve previously reported that scientific journals have published AI generated text, that AI is impacting the peer review process, and that the open-access repository for preprint academic research Arxiv will now ban authors for a year if they are caught submitting AI generated work. asserted
they → report → work
The potential upside of this research is that these names might allow us to detect this AI generated content more easily. uncertain
us → allow → content
💬 Give feedback
🕘 History 🎫 Support