Why AI Invents Sources in the First Place
Ask an AI for a statistic and a citation, and it will hand you both — confidently, formatted correctly, with a plausible author and journal name. The problem is that language models don’t retrieve facts from a filing cabinet. They predict what a citation usually looks like based on patterns in their training data, then fill in the blanks. When the model doesn’t actually know the source, it generates one that fits the shape of a real one.
That’s why a fabricated citation is so hard to catch on sight. It isn’t sloppy — it’s engineered to pass a glance. A made-up study might cite a real journal, attach a real-sounding author name, and land on a plausible year. AI tools have been known to invent a Harvard Business Review article complete with a page range, for a claim HBR never published. The formatting was perfect. The content didn’t exist.
For client work, this is where trust breaks. You’re not just handing over prose — you’re handing over claims your name is attached to. A wrong verb tense is forgivable. A citation that doesn’t exist, discovered by your client’s own fact-checker, is not.
The Four-Step Verification Routine
Treat every AI-generated citation as a claim to test, not a fact to trust. Before anything goes to a client, run it through four checks.
Open every link. Not most of them — every one. A broken link is the easiest fabrication to catch, and AI models produce them constantly: URLs that look structurally correct but lead nowhere, or redirect to a publisher’s homepage instead of the actual article.
Confirm the quote exists in the source. Search the exact phrase inside the page rather than skimming for the general idea. AI models paraphrase real sources into quotes that sound right but were never actually said — a different failure mode from an invented source, and a more common one.
Check the date. A statistic from “a 2024 study” attached to a source published in 2019 is a mismatch worth catching before your client does. Dates also tell you whether a claim has since been updated or retracted.
Check who published it. A legitimate-sounding domain isn’t the same as a legitimate publisher. Look at the About page, check whether the outlet has an editorial process, and be wary of sites that exist mainly to host SEO content with citations bolted on for credibility.
None of these steps takes more than a minute or two per source. Skipping them is the expensive option. It just doesn’t feel that way until the client emails back.
Red Flags Worth Slowing Down For
A few patterns show up often enough in AI output that they’re worth flagging on sight, before you even start the verification routine.
A stat with no source attached at all is the most obvious one — “73% of freelancers report burnout,” with nothing behind it. If the AI won’t tell you where a number came from when asked directly, that’s your answer.
Broken DOIs are another. A real academic citation includes a Digital Object Identifier, and pasting it into doi.org should take you straight to the paper. When it returns an error, the citation was either mistyped or never existed — a check that catches an otherwise convincing fabrication in under ten seconds.
Authors you can’t find anywhere — not on LinkedIn, not in a Google search, not on the publication’s own author page — are the third pattern. A real writer, even an obscure one, leaves some trace. A name invented to sound credible usually doesn’t.
Tools That Make This Faster
Google Scholar is the fastest way to check whether an academic-sounding claim actually exists in the literature. Search the author’s name plus a keyword from the claim, and a real paper usually surfaces on the first page of results.
The original publisher’s site matters more than the search-engine snippet. Aggregators and content farms often reproduce, or mangle, statistics from real studies, stripped of context. Going to the source — the journal itself, the government agency, the original report — tells you whether the number still means what the AI claimed it meant.
Searching the exact quote in quotation marks is the fastest gut-check for text citations. If a supposedly direct quote returns zero results, it was never said, at least not in those words.
When You Can’t Verify a Claim
Sometimes the trail runs cold. The source doesn’t answer, the archive doesn’t have it, or the claim is plausible but simply untraceable in the time you have. You have three options, and which one you pick should depend on how load-bearing the claim is to the piece.
Remove it. If the point still stands without that particular statistic, cut it. A client report is stronger with four verified numbers than five, one of which might be fiction.
Soften it. Turn “73% of freelancers report burnout” into “burnout is a common complaint among freelancers, based on informal surveys and anecdotal reports.” You lose some punch, but you stop making a specific claim you can’t back up.
Flag it to the client. For internal drafts or working documents, a bracketed note — [unverified, could not confirm this stat, recommend removing or finding a primary source] — keeps the work moving without quietly passing along something shaky.
What you shouldn’t do is publish it anyway because it sounds right. “Sounds right” is exactly the failure mode this whole process exists to catch.
Keep a Verification Log
For anything client-facing, keep a running log of what you checked and how. It doesn’t need to be elaborate. A simple table works: the claim, the source you verified it against, the date you checked, and a link.
This does two things. First, it forces the actual verification step instead of a vague sense that you “looked into it.” Second, it gives you an answer when a client asks where a number came from six months later, which happens more often than you’d expect once you’re citing data regularly.
A five-minute habit per report is a small price for never having to say “I’m not sure where that came from” to a client.