The Rise of the Undead: How AI-Generated "Ghost Names" Are Polluting Academic Publishing
The Rise of the Undead: How AI-Generated "Ghost Names" Are Polluting Academic Publishing
Picture this: Elena Vasquez and Marcus Chen have appeared as volcano experts, astronauts, thriller protagonists, podcast hosts, and academic co-authors across hundreds of documents. There's just one problem — these people have never existed. They're not hidden figures in history or obscure researchers. They're ghost names, conjured repeatedly by AI models that have never met them.
A fascinating new research paper (arXiv:2606.02184) has pulled back the curtain on one of the more unsettling side effects of the generative AI boom: our AI systems have developed distinct "name priors" — preferences for certain fictional names that appear with suspicious frequency across their outputs. And this isn't just a curiosity. It's creating real problems for academic publishing, scholarly repositories, and anyone trying to verify that content on the internet was actually written by a human being.
When AI Gets Nostalgic About Made-Up People
Here's what's happening: when you ask an AI model to generate a fictional expert, a podcast host, or a research paper co-author, it doesn't pull from a neutral pool of names. Each model family has its own favorites. Claude apparently loves Elena Vasquez and Marcus Chen as a duo, often adding Amara Okafor for a trio. GPT tends toward Elara Voss but hasn't settled on a consistent partner. Gemini prefers Aris Thorne and Lena Petrova.
These aren't random selections. The researchers found that the co-occurrence rates of these name pairs far exceed what you'd expect by chance. It's as if each AI model has been reading the same fictional novels and keeps casting the same characters in new roles.
What's particularly interesting is that these priors are model-family-specific and even version-specific. When a company like Anthropic, Google, or OpenAI updates their model, the ghost name preferences often shift. This leaves what the researchers call "dateable behavioral fingerprints" — temporal markers that can help date when a piece of AI-generated content was created based on which names it uses.
The Academic Pollution Problem
But here's where this crosses from quirky to concerning. The researchers identified a massive contamination problem on Zenodo, the CERN-operated repository that mints real DataCite DOIs. They found 1,655 ghost-authored records claiming to be published in journals that don't exist, with fabricated publication dates.
This isn't just about fake papers sitting in a dusty corner of the internet. These records carry real DOIs registered in DataCite, making them harvestable by any scholarly aggregator that ingests DOI metadata. That's Google Scholar, ResearchGate, academic databases, citation indexes — the entire infrastructure that researchers rely on to find legitimate scholarship.
The evidence suggests deliberate backdating. Server-side DataCite timestamps prove someone registered these DOIs intentionally to make the content appear older than it actually is. And here's the smoking gun: 991 of these records were registered in a single month. That's not organic growth. That's a coordinated contamination event.
The researchers also found that ghost names appear on ResearchGate forming synthetic research groups with collaborators drawn from multiple model families. The publication dates on these records provide a reliable temporal proxy for model deployment windows — essentially creating a timeline of when different AI models were released based on when their "favorite" ghost names started appearing in published research.
Why This Matters for Developers and Tech Leaders
You might be thinking: "I'm not in academia, so why should I care?" Here's why this matters even if you've never written a research paper:
First, it's a canary in the coal mine for content authenticity. If academic repositories — which presumably have more rigorous standards than most corners of the internet — can be this easily polluted, imagine what's happening with product descriptions, marketing copy, news articles, and technical documentation. We're entering an era where "AI-generated" is no longer the exception; it's becoming the default.
Second, it reveals the hidden patterns in AI output. Understanding that AI models have these biases and preferences isn't just academically interesting — it's practically useful. If you're building applications that rely on AI-generated content, knowing that models have "favorite" names, phrases, and character archetypes can help you build better detection and quality assurance systems.
Third, it highlights the need for AI-assisted development practices that account for these quirks. At NameOcean, we're thinking a lot about how AI is changing web development, hosting, and digital presence. This research is a reminder that AI tools are powerful but not neutral. They carry their own fingerprints, their own biases, their own tendencies. Smart developers need to understand those patterns, not just use the tools blindly.
The DOI Infrastructure Problem
Perhaps the most troubling aspect of this research is what it reveals about DOI infrastructure. DOIs (Digital Object Identifiers) are supposed to be the backbone of scholarly communication — permanent, reliable links to academic works. They're trusted because they're assumed to point to real content with real authors.
But the current system has no mechanism to verify that a DOI registration corresponds to actual human authors who actually conducted the research. As long as you can pay the registration fee and provide metadata, you get a legitimate DOI. The system trusts that registrants are acting in good faith.
This creates an interesting parallel to domain registration. At NameOcean, we know that domain names are often the first line of defense for brand identity and trust. Similarly, DOIs are supposed to signal trustworthiness in the academic world. When that trust is abused, it undermines the entire system — just like domain spoofing or DNS-based attacks can undermine trust in web infrastructure.
What Can We Do About It?
The researchers suggest several approaches to addressing this problem:
DOI registries could implement verification checks that flag suspicious patterns — like hundreds of papers appearing from journals that don't exist, or author names that appear in impossible frequency across unrelated publications.
AI model developers could track and audit their name generation patterns more carefully, potentially reducing the memorization of specific name combinations.
Academic institutions and publishers could implement more rigorous verification before accepting papers, including author identity verification.
Detection tools could be developed to identify ghost-authored content based on the telltale signs the researchers identified.
For developers building applications that interact with AI-generated content, the lesson is clear: treat AI output with appropriate skepticism and build verification into your workflows. Don't assume that content that looks professional was necessarily created by a professional.
The Bigger Picture
This research is a fascinating glimpse into the emerging field of AI forensics — the study of how AI systems leave traces of their behavior that can be detected and analyzed. As AI-generated content becomes ubiquitous, these traces become increasingly important for maintaining trust and authenticity.
For the tech community, it's a reminder that we're still in the early days of understanding how these systems behave at scale. The ghost names phenomenon is just one example of the unexpected patterns that emerge when millions of people start using the same AI tools. There will be more surprises ahead.
What do you think about this ghost authorship phenomenon? Are you concerned about AI-generated content polluting academic and digital infrastructure? We'd love to hear your perspective.
Have questions about AI-assisted development or how to build trustworthy digital presence? Check out NameOcean's vibe hosting solutions designed for the AI era.