Using AI for Competitor Research (Without Publishing a Hallucinated Number)

Using AI for Competitor Research (Without Publishing a Hallucinated Number)
The synthesis is instant. The verification is the part you have to schedule.

AI competitor research tools compress days of manual analysis into minutes, and they fabricate data often enough to matter. On Vectara's hallucination leaderboard, even top models invent unsupported claims in roughly 2 to 3 percent of document summaries, and some reasoning models run above 17 percent. The fix is procedural: no AI-sourced number enters a strategy deck until a human traces it to a primary source.

That opening stat deserves a second look, because the benchmark it comes from is the easy case. Vectara measures whether a model stays faithful to a document it was literally handed. Competitor research is the hard case: you're asking about pricing pages that change weekly, market share nobody publishes, and funding rounds the model half-remembers. The error rate on those questions is not 2 percent. Nobody knows exactly what it is, which is sort of the problem.

Deloitte's AU$440,000 proofreading failure

In July 2025, Deloitte delivered an assurance review of a welfare compliance system to Australia's Department of Employment and Workplace Relations. Price tag: AU$440,000. A University of Sydney academic, Chris Rudge, then found more than a dozen fabrications in it: made-up academic papers, a quote attributed to a federal court judgment that doesn't appear in the judgment, footnotes pointing at nothing. Deloitte later disclosed it had used Azure OpenAI GPT-4o in drafting and agreed to repay the final instalment of the contract.

This is the most instructive AI research failure I've seen, and the reason is who it happened to. Deloitte has review layers most marketing teams can only dream about. Partner sign-off, QA processes, a brand built on rigor. The fabricated citations sailed through all of it, because hallucinated citations look exactly like real ones. Correct formatting, plausible authors, confident phrasing. Every error arrived dressed as a fact.

Your competitor battlecard gets less scrutiny than a government deliverable. If the fabrications survived Deloitte's process, they will survive yours.

Where AI competitor tools actually earn their subscription

None of this means skip the tools. Klue's roundup of the category reports that 60 percent of competitive intelligence teams now use AI daily, with teams citing around a 45 percent reduction in data-processing time. Those gains are real, and from what I've seen they cluster in specific jobs:

  • Change detection. Tools like Visualping watching a competitor's pricing page and diffing it. The tool reports what it observed. Very little room to hallucinate.
  • Traffic and search data. Similarweb, Semrush, Ahrefs. These are estimates with known error bars, which is a different thing from fabrication. An estimate labeled as an estimate is honest.
  • Summarizing material you supply. Feeding 40 G2 reviews into a model and asking for complaint themes. This is the Vectara scenario, where error rates are low single digits for good models.
  • Generating the question list. Asking a model what you should investigate about a competitor is close to risk-free, because questions can't be false.

The risk concentrates in one place: asking a general-purpose chatbot open questions about the world. "What is Competitor X's pricing?" "How big is their sales team?" "What's their market share?" These are exactly the questions where public data is thin, and thin data is what the model papers over with plausible invention. It behaves like a junior analyst who fills in every blank on the template whether or not the data exists, because leaving a cell empty feels like failing the assignment.

A number without a source link is the model's best guess, formatted as a fact.

Why hallucinated numbers look so trustworthy

Language models are trained to produce likely text, and a specific number is more likely-sounding than a hedge. "Competitor X charges $49 per seat" reads better than "pricing information was not publicly available," so under uncertainty the model drifts toward the version that reads better. The Vectara leaderboard adds an uncomfortable wrinkle: some reasoning-focused models hallucinate more than their simpler siblings, with one scoring above 23 percent on a task as constrained as summarization. More thinking steps seem to create more chances to drift from the source. Seems backwards, but the data keeps showing it.

And to be fair, the models have gotten better. The best current models are dramatically more grounded than what was shipping two years ago. But "better" in competitor research is a trap, because the improvement shows up as fewer wrong numbers wrapped in identical confidence. A tool that's wrong 30 percent of the time gets double-checked. A tool that's wrong 5 percent of the time gets trusted, and the 5 percent goes straight into the QBR deck. On paper, lower error rates sound like the problem solving itself. In practice they mostly relocate the risk.

I'll make the prediction concrete: before the end of 2027, at least one Fortune 500 company will publicly retract a competitive claim that traces back to an AI research tool, Deloitte-style, with lawyers involved. The ingredients are all sitting out on the counter.

The 15-minute verification pass

Here's the workflow, adapted from the citation-checking process Editage recommends for academic references, trimmed for marketing speed. Run it on any AI-generated competitor report before it touches a deck.

  1. Extract every number and named fact into a flat list. Pricing figures, headcounts, dates, quotes, feature claims. Two minutes. Most reports have 8 to 15 of these.
  2. Sort by consequence. Which of these would embarrass you in front of a VP if wrong? Verify those first.
  3. Check for a source link on each. No URL attached means it came from the model's head. Flag it.
  4. Open the URLs that exist and search the page for the exact figure. Ctrl-F the number. A real citation contains the claim. A hallucinated one links somewhere vaguely topical, and this pattern (real page, wrong claim) is the one that fools people most.
  5. For flagged claims, search the exact phrase in quotes. If the first page of results doesn't corroborate it, the claim dies. Reword or cut.

The benchmark: if more than 2 of 10 claims fail verification, don't patch the report. Discard it and rerun the research with a browsing-enabled mode and tighter constraints, because a one-in-five fabrication rate means the whole synthesis is contaminated, not just the flagged lines. We follow a similar rule on this site, and honestly it's the least fun part of publishing. It's also why the system we use for scaling AI content treats source verification as a separate pass rather than trusting the drafting step. The one time you skip it is the time a number was invented.

Prompts that cut the cleanup roughly in half

You can't prompt hallucination away, but you can make it visible. Three constraints do most of the work:

Require the model to say "no public data." Add: "Only include figures you can attach a specific URL to. For anything else, write 'no public data found' instead of estimating." You're giving the model permission to leave the blank blank. Expect the report to get noticeably thinner. The thinness is the point; those missing numbers were never real.

Split observed from inferred. Ask for two labeled sections: things directly stated by a source, and the model's inferences. Models are decent at this separation when forced, and the inference section is often the genuinely useful part, as long as it's labeled as thinking rather than fact.

Use browsing or deep-research modes for anything factual, and still check. Retrieval-backed answers hallucinate less, since the model quotes pages instead of reconstructing them from memory. Less is doing a lot of work in that sentence. The Deloitte report had access to real source material and still invented citations. We saw the same dynamic in the early ChatGPT ads data: the tooling improves faster than the trust it deserves.

A 10-minute test worth running today: take the last competitor question you asked an AI tool, rerun it with the "no public data" constraint, and diff the two answers. Count how many confident numbers from round one disappear in round two. In my experience the count lands between a quarter and half, and every one of them was a number you almost believed.

Questions people keep asking

Can ChatGPT do competitor analysis? Yes, for structure, question generation, and summarizing material you paste in. For factual claims, use a browsing-enabled mode, require URLs, and verify anything consequential by hand. Never accept a market share or pricing figure generated without a source attached.

Which AI competitor research tool is most accurate? Tools that report observed data (Visualping for page changes, Similarweb for traffic estimates, Semrush for search visibility) are structurally more reliable than general chatbots answering open questions, because their numbers come from measurement rather than generation. Purpose-built CI platforms like Klue or Crayon sit in between: strong at organizing verified intel, only as good as the sourcing feeding them.

How do I stop AI tools from making up statistics? You reduce it rather than stop it. Constrain prompts to sourced claims only, allow the model to answer "no data," prefer retrieval-backed modes, and run a verification pass on every output. Deloitte's AU$440,000 refund is the going rate for skipping that last step.

The boring 15 minutes nobody budgets for

AI competitor research is a first draft. Deloitte shipped the first draft and repaid part of AU$440,000 for the privilege, with the story covered by every outlet in the country. The verification pass costs 15 minutes and has no cool demo, which is probably why almost nobody does it. I'd budget for it the way you budget for the tool itself. The subscription gets you speed. The 15 minutes is what makes the speed safe to use.

By Notice Me Senpai Editorial