What deep research is and how it works

You pay AI two hundred dollars a month and expect a commensurate level of accountability. When a genuinely demanding task comes along — like a detailed analytical report on a niche topic — you describe it and go make coffee. Twenty minutes later, a flawless thirty-page document is waiting for you, containing an introduction, conclusions, a clean structure, and eighty footnotes with source links. It's laid out so impeccably that it looks as though a real analyst had slaved over it for a week. Staring at this beautiful thing, you can't help but think: there they are, two hundred dollars, and the result is worth every penny.

A man who shared his story on Reddit received exactly that kind of report on European regulation. It was a fantastically polished, impressively substantial document with an elegant structure and a wealth of citations. Before taking it into a meeting, he decided to check one of the links, just to be safe. The source didn't exist. "[The AI] literally hallucinated a specific clause that would have solved all my problems," he wrote.

This wasn’t confusion about wording or the accidental use of outdated resources. The model

had simply invented a nonexistent clause in a nonexistent rule, formatted it neatly, and helpfully attached a fake footnote. (The same mechanism, only in federal court and for $15 billion: "Wall Street Hallucinations".) Is our distinguished Reddit user to blame? Let's dig in.

Before December 2024, deep research mode didn't exist anywhere. Google was first out of the gate on December 11 with the Gemini 2.0 package. The promises were modest: your personal assistant will review dozens of sources and write a structured report. Just eight weeks later, OpenAI stepped onto the stage with its version for ChatGPT Pro subscribers at two hundred dollars a month. The marketing was considerably louder: "find, analyze, and synthesize hundreds of online sources to create a comprehensive report at the level of a research analyst." Over the next several months, similar functionality appeared at Perplexity, Anthropic, and even Microsoft Copilot. Today it feels as though deep research always been there, baked into every model from the start.

The issue isn't that the tool makes mistakes — we stopped being surprised by LLM hallucinations long ago — but that it lies beautifully, expensively, and at length. Human nature being what it is, we tend to trust work that looks large-scale and outwardly complex. We treat twenty minutes of waiting and eighty footnotes as some kind of quality guarantee. What we get instead is an anesthetic for the part of your brain that usually asks: "Is this actually true?"

Rather than "How accurately does deep research perform?" the pertinent question is "Under what conditions can you trust its output at all?" To answer that, we need to work through two problems: why you believed the report in the first place and where the LLM sourced the nonsense you accepted.

ma.png

Part I. Why You Believe

Beauty as a sedative

Our brains are lazy: the more expensive packaging looks, the less inclined we are to inspect the contents. We automatically view a long, structured text with dozens of footnotes as a scholarly work, which we associate with weeks of intensive peer review. Deep research short-circuits that reflex with virtuoso precision, flawlessly mimicking academic documents while offering no academic guarantees.

The confident tone of the report is no accidental side effect — it’s a foundational feature of LLMs. In May 2026, researchers from OpenAI itself (Adam Kalai's team) published a paper in Nature containing an uncomfortable mathematical proof that modern metrics push models to guess. "Dominant headline metrics such as accuracy systematically reward guessing over admitting uncertainty," so an AI that always picks blindly scores higher than one that honestly admits it doesn't know. "Like students facing a difficult exam question, models guess rather than admit uncertainty."

The authors proved this with almost cruel simplicity: they asked three flagship models to expand the abbreviation PGGB and received three entirely different, equally wrong answers. Instead of an honest "I don't know" or even a tentative "this might mean...," the models invented facts on the fly.

Now add twenty minutes of waiting and eighty footnotes. A machine that is structurally incapable of admitting its own ignorance generates thirty pages of text. It writes real data and hallucinations with exactly the same air of authority. There is no way to tell truth from fabrication: the text has been optimized specifically to sound coherent and maximally authoritative.

Repeat victims describe the experience far more precisely than any academic finding: "The scary part isn't that it's wrong, it's how confident it is while lying to you. It delivers a hallucination with the exact same authority as a verified fact. There's no 'maybe' or 'check this' flag." Another casualty puts it even more bluntly: "This isn't getting better — it's getting smoother. The failure isn't hallucination. It's fluent answers crossing the line into authority before a human check happens."

Smoother, not better. Remember that.

The machine that agrees with you

The pretty packaging is only half the problem. Here's the other half: deep research will tell you exactly what you want to hear. It starts calibrating to your expectations before it even runs its first search query.

In November 2025, Science Advances published "Source framing triggers systematic bias in large language models" by Federico Germani and Giovanni Spitale of the University of Zurich. The authors took four advanced models, asked them to evaluate 192,000 statements on contested topics, and found something odd. When the source of a quote was hidden, the models agreed with each other almost perfectly. Add an attribution and everything shifted. The exact same statement would receive diametrically opposite assessments depending on how the model felt about the source.

What does this have to do with you? Everything. The model's response is shaped by context — specifically, by whoever sets the frame. Your query is the frame. The moment you phrase a task along the lines of "find arguments for why craft beer is healthier than a morning run," you've killed neutrality at the starting line. You've already told the machine which answer you consider correct. A system trained to please will dutifully construct your preferred reality, finding sources that confirm your position and quietly burying anything that contradicts it.

The developers themselves acknowledge this. In spring 2025, OpenAI rolled out an update it had to hastily pull back a few days later: the model had turned into a sycophant, gushing over every user idea, even the plainly idiotic ones. In its post-mortem, OpenAI explained that the update had introduced a reward signal tied to user feedback — likes and dislikes — and "user feedback in particular can sometimes favor more agreeable responses." That metric eroded the baseline filters holding sycophancy in check. The neural network learns to accommodate us and win our approval, not to seek the truth and fight for it to the last.

Information expert Mike Caulfield, writing in The Atlantic, put it well. AI, he argued, is becoming a "justification machine — more convincing, more efficient, and therefore even more dangerous than social media." Social media at least makes you go out and hunt for confirmation of your views yourself, whereas deep research delivers your own rightness straight to your door, gift-wrapped.

As always, the most vulnerable users turn out to be those with middling expertise — the ones who feel they've finally figured it all out. A novice asks a vague question and gets an equally vague but relatively balanced overview. Someone who recently became an expert formulates a precise query, loading it with their own hypothesis, and walks away with a detailed confirmation of what they already believed. A tool sold as a way to stress-test ideas ends up being used to legitimize them and pile on one-sided evidence.

ma.png

Part II. What This Nonsense Is Made Of

Let's say you're not a novice. You ask the driest possible questions, you don't offer hints, and you don't let the model agree with everything you say. The problem is that even the most honest AI is forced to work with the internet as it actually exists. Here your skepticism is powerless, because the raw material from which deep research assembles its polished reports is compromised at the source — to a degree you probably can't imagine.

Bad flows in

The internet from which deep research draws its data is filling up with AI-generated text at a staggering pace, and degrading just as fast. Most of the time, this content is invisible to the general public. It includes texts churned out by neural networks under the direction of SEO managers, endless landing pages and backlinks (links that signal to Google that someone other than the advertiser actually cares about their page), and spam that games Google's lumbering, easily-fooled algorithm to drive traffic to an advertiser's site.

Here is direct data from Ahrefs, one of the most authoritative sources in the industry:

We analyzed 900,000 newly created web pages in April 2025 and found that 74.2% of them contained AI-generated content.

These pages are state of the art, with impeccable section headers and plausible-looking tables. The figures in those tables are usually barefaced lies planted by a crooked advertiser — or, more often, by the crooked SEO agency they hired — and that bogus table from some no-name website ends up as a source in a deep research report. This isn't a hallucination. The link clicks through. The table appears on screen. Every number on the site matches the report.

In April 2026, at the ACM Web Conference, researchers from NAVER presented a paper titled "Retrieval Collapses When AI Pollutes the Web." To model the real-world process, they took a corpus of documents from which a search engine builds its results, then gradually mixed in synthetically generated texts. When the share of synthetic content reached two-thirds, they looked at what was actually surfacing — what the user actually sees. It turned out that AI-generated content already accounted for more than eighty percent of those results. It displaces human-written text disproportionately, because machine text is smoother and better optimized for ranking algorithms, so search engines are far more eager to push it to the top. (On why degradation from synthetic data was predicted as far back as the 1950s: "Wittgenstein Knew Why AI Gets Dumber".)

The numbers are bad enough, but the content itself is in an even worse condition. The authors describe it as "a homogenized yet deceptively healthy state where answer accuracy remains stable despite the reliance on synthetic sources." In practice, this means your polished report may be woven entirely from AI paraphrases of other AI paraphrases. Formally it will still look correct, but the living primary sources will have been quietly swapped out for machine echoes. With every such iteration, all the complex nuances, caveats, and rare details get washed away.

And while the bad flows in, the good stays out.

Good stays out

Deep research is marketed as a premium professional tool for people who need reliable data. However, the most trustworthy sources, like peer-reviewed journals, quality press, and analytical reports, are completely off-limits to it. Why? Because they’re good, and that means they cost money.

In 2026, a study titled "Science Behind a Paywall" appeared in Learned Publishing. The numbers are brutal: 64% of biomedical articles remain behind a paywall, inaccessible to AI for either training or analysis. The five largest academic publishers control half the market and aggressively block automated data collection. As a result, the authors write, models learn primarily from abstracts, media summaries, and preprints, and "abstracts provide only skeletal methods, results, and conclusions; media summaries are often superficial or fragmented." Your expensive AI researcher is operating not on scientific facts but on whatever slipped past the paywall. The people who need reliable data don't always realize that.

Bad flows in, good flows out, and the evidentiary base underpinning your polished report turns out to be systematically degraded. It’s not an accident and it doesn’t just occur occasionally: it’s business as usual.

Recall the findings from Nature, showing that models are rewarded for guessing with confidence, and the NAVER data demonstrating that retrieval looks accurate even when it's built on garbage. In both cases, the system appears perfectly healthy at the exact moment it is producing junk. The model is confident — but wrong. The output looks accurate — but it's built on garbage. The form of your report — which is precisely what deep research does best — will be noticeably better than its substance.

ma.png

What to do about all of this

The temptation, of course, is to wrap things up with the perfectly accurate and completely useless advice to "just fact-check everything." Let's resist that. Remember the analyst with the fabricated legal clause: catching that one hallucination required manually tracking down the original source, and the report had eighty footnotes. It's not enough to check whether the links are live and the citations aren't hallucinations, either. Not even close. You need to verify every single one of those eighty sources to determine whether it's an SEO product. Properly scrutinizing a document like the Reddit user’s takes roughly as long as doing the research yourself. That raises a fair question: why did you pay two hundred dollars in the first place?

That said, deep research is far from useless. It is excellent at building the skeleton of a document, just consistently falls apart when it comes to standing behind specific figures and quotes. Still, you can help both the LLM and yourself by changing your approach.

Before assembling the report, think carefully about what you feed the model. If you have quality source materials you already trust — studies, reliable data, or paywalled articles you've already paid for and retrieved — load them upfront. Let the AI work with what you've brought to it, rather than trawling through internet garbage. That is the correct use of your expensive, powerful tool.

Once the report is ready, don't go back and read all eighty footnotes. That's an illusion of safety that will eat up every minute you thought you'd saved. Instead, ask the same model — in a fresh conversation — to play devil's advocate: find the logical gaps, dubious figures, and citations that are begging to be checked. The machine does this brilliantly, even when it's nominally defending its own conclusions. At least for now, in Q3 2026, a top-tier model switches effortlessly from hallucinating researcher to hardened skeptic. You'll get an excellent vulnerability map to guide your own fact-checking. From there, you have a lot of manual work, but at least you know where to aim.

More importantly: stop mistaking polish for truth. Twenty minutes of waiting guarantees nothing. Eighty footnotes guarantee nothing. A confident tone is simply the machine's factory default.

The tool you were sold as a replacement for a research analyst works best when you're the one driving it. For that, you need a precise understanding of how it's built, when it can be trusted, and when it can't. On average, you can almost never rely on it — but there are happy exceptions.

Sources

References cited in this piece. Last verified on the published or revision date.

  1. 01
  2. 02
  3. 03

    Expanding on what we missed with sycophancy

    openai.com/index/expanding-on-sycophancy

  4. 04

    AI Is Not Your Friend

    www.theatlantic.com/technology/archive/2025/05/sycophantic-ai/682743

  5. 05
  6. 06
  7. 07

    Introducing deep research

    openai.com/index/introducing-deep-research

  8. 08

    Try Deep Research and our new experimental model in Gemini

    blog.google/products-and-platforms/products/gemini/google-gemini-deep-research

  9. 09

    Deep Research hallucination thread

    www.reddit.com/r/ArtificialIntelligence/comments/1qo2pes

  10. 10

    A warning about ChatGPT's deep research hallucination

    www.reddit.com/r/ChatGPT/comments/1lcx4es