Why a General AI Chatbot Isn't Enough for Scientific Literature Review
Why general chatbots invent citations, what agents do differently, and why the useful question isn't how smart the AI is but what it's connected to.

These days, when you want to understand what's known about a topic, the first thing you do is ask an AI chatbot. In seconds, it gives you a fluent, confident summary, complete with a tidy list of citations. It looks perfect. Then you go to pull the papers, and something's wrong. One citation leads nowhere. Another has the right authors but a title that doesn't exist.
If this sounds like an isolated annoyance, it isn't. Peer-reviewed studies that have systematically audited AI-generated literature reviews have found that a substantial share of the references did not correspond to any real publication. And the consequences are no longer hypothetical: Deloitte was recently placed under formal investigation over a $1.6-million government report built on citations to studies that did not exist. Fabricated citations are not the exception. They are a predictable feature of how these systems work.
The reasons are still debated, but part of it comes down to how these systems work. A general chatbot doesn't know things the way a researcher does. It predicts the next most likely word based on patterns in its training data. When it has seen enough relevant examples, that prediction is often right. When it hasn't, it doesn't stop and say "I'm not sure." It produces a fluent, confident guess instead. One argument, put forward by researchers at OpenAI in 2025, is that models hallucinate partly because training and evaluation reward confident guessing over admitting uncertainty.
This limitation is one of the reasons agents developed. An agent isn't a different kind of intelligence. It is usually the same kind of model, placed inside a system designed to work around that weakness: instead of answering from memory, it can go and look. Think of it this way. Asking a chatbot is like asking a brilliant colleague who has read everything but is now speaking entirely from memory, with no books in front of them. Fluent, often right, but unable to show you the source. The same colleague, walking to the library, pulling the paper off the shelf, and pointing to the exact line that supports the claim, is doing something different.
What an Agent Actually Does
In practice, this comes down to three moves: how a question gets broken up, where the answer comes from, and what happens before it reaches you.
Working through a question in stages. A single-pass system takes your question, runs one search, and answers. That is fine for simple questions, but it breaks down for real ones, where what you find changes what you should be looking for next. An agent splits a question into smaller ones, searches, and uses what comes back to decide where to look next.
Going to the source rather than to memory. Instead of relying on what it picked up during training, the system searches live databases and builds its answer from papers it has actually pulled up. If it can only cite what it retrieved, making up a paper becomes far less likely. And the difference shows. In one study, a general model made up most of its citations. The same model, connected to a real library of papers, cited as accurately as human experts. The model hadn't changed. What it was connected to had.
Checking the answer before handing it over. This is the part that matters most for research. Instead of just producing a conclusion, the system can go back over it against the sources it pulled, tracing each claim to the passage behind it. Some systems do this in a loop: draft an answer, look for gaps, retrieve more, revise. The idea is that grounding should work like a contract. Every claim in the answer should be traceable, line by line, to a source you can open yourself. In research, an answer you can't check isn't much of an answer.
Where the Limits Still Are
So does building a system this way solve the problem? Not entirely. Grounding an answer in retrieved sources reduces hallucination, but it doesn't remove it. Even with good retrieval, a model can misread a passage, over-generalise, or claim something the evidence doesn't quite support. Working through a question in stages also takes longer. And when a search comes back empty, a poorly built system may quietly fill the gap and guess anyway. Worse, some hide that failure entirely, handing you a confident answer with no sign that something broke along the way.
Which is why the honest framing isn't AI instead of experts. The more credible view of where this is heading describes it not as a robot scientist, but as a tireless lab partner working inside a human-verified loop. The point isn't automation for its own sake. It's a system whose steps you can actually see, so that when something goes wrong, someone can catch it.
It's Not About Smarter AI. It's About What It's Connected To.
Which brings the question back to where we started. It was never "should I use AI for this?" The more useful question is what the system is connected to, whether you can see how it got there, and whether an expert is still in a position to catch what it got wrong. A smarter model doesn't answer any of those. Better architecture does.
That's the thinking behind Clarisyn. Several agents work the same question independently, and their outputs are cross-checked against each other and against the sources they came from, then returned as a three-layer report. Human oversight is built in throughout, and when you have your answer, expert consultation is one click away to verify it with someone who works in that field.
Accelerate Your Translational Research
Experience how Clarisyn transforms complex scientific problems into structured, evidence-based reports.
Get Started with Clarisyn