Caught in the Act: We Tested Four AI Tools on High-Stakes Questions and the Results Were Uncomfortable
Everybody talks about AI hallucinations like they're a quirky personality trait. "Oh, it just makes stuff up sometimes—so funny!" That framing is fine when you're asking an AI to write a birthday poem. It is decidedly not fine when you're using it to research drug interactions, verify a contract clause, or figure out whether a business expense is tax-deductible.
We decided to stop talking about hallucinations in the abstract and actually test them. Not with trick questions or adversarial prompts—just the kinds of real-world queries that professionals actually type into these tools every day.
The results? Genuinely uncomfortable in some places, and surprisingly encouraging in others.
How We Set Up the Test
We ran a structured set of 60 questions across three high-stakes domains: legal research, medical information, and basic financial guidance. Each question had a verifiable correct answer—something we could check against primary sources like court databases, FDA documentation, and IRS publications.
The tools we tested: ChatGPT (GPT-4o), Claude 3.5 Sonnet, Google Gemini 1.5 Pro, and two specialized alternatives—Harvey AI (legal) and Perplexity AI (as a citation-grounded generalist).
We scored each response on two dimensions: factual accuracy and appropriate hedging. An AI that says "I'm not certain, but I believe..." before a wrong answer scores better than one that confidently states the same wrong answer as fact. Why? Because overconfidence is what turns a wrong answer into an actual problem.
The Legal Category: Where Confidence Becomes a Liability
Legal research is arguably the scariest domain for AI errors because the stakes compound. A wrong citation, a misremembered ruling, a statute quoted from the wrong jurisdiction—any of these can cascade into real harm.
Gemini struggled here more than we expected. On questions about specific federal circuit court precedents, it produced plausible-sounding case names that simply don't exist. Not every time—but often enough that we'd be nervous recommending it for anything beyond a first-pass brainstorm. ChatGPT performed similarly, occasionally citing real cases but describing holdings that didn't match the actual rulings.
Claude was noticeably more cautious. It frequently flagged when it was uncertain about jurisdiction-specific details and recommended verification. That hedging slowed things down, but in a legal context, slow and right beats fast and wrong every single time.
Harvey AI, built specifically for legal work and trained on legal corpora, outperformed all the generalists on citation accuracy. It still made errors—no AI is perfect here—but it was far less likely to invent a case wholesale. If you're doing serious legal research, the specialization matters.
The Medical Category: Where Wrong Answers Can Actually Hurt Someone
We asked questions about drug interactions, dosage thresholds, and symptom differentials. Again, all verifiable against FDA and NIH documentation.
This is where ChatGPT's confidence problem showed up most clearly. On several drug interaction questions, it provided answers that were partially correct but missing critical contraindications. The answers sounded authoritative. They weren't flagged as uncertain. A layperson reading them would have no reason to question them.
Claude, by contrast, consistently added appropriate caveats—"this information may vary based on individual health factors, and you should consult a healthcare provider"—without being so hedge-heavy that the response became useless. It also performed better on accuracy, getting the interaction classifications right more often.
Gemini's performance in the medical category was its strongest of the three domains, likely because it has deep integration with Google's health-focused data. Still, it had moments of overconfidence that gave us pause.
Perplexity AI deserves a mention here. Because it cites sources in real time, you can immediately check whether the answer is grounded in something real. It's not a substitute for medical expertise, but the transparency is genuinely valuable for someone who wants to verify rather than just trust.
The Financial Category: Tax Law, Investing Rules, and the Confidence Trap
Financial questions are tricky because the rules change—tax law gets updated, contribution limits shift annually, SEC regulations evolve. An AI trained on data from 18 months ago can confidently tell you something that used to be true.
All four generalist tools made version-related errors here. The 2024 401(k) contribution limit, the current standard deduction, the income thresholds for Roth IRA eligibility—these numbers shift, and the models didn't always reflect the most current figures. ChatGPT was the worst offender, stating outdated numbers without any caveat that the information might be stale.
Gemini, with its live search integration, actually handled currency better than the others in this category. When it pulled current data, it was more likely to be right about time-sensitive figures.
Claude acknowledged uncertainty about whether its training data reflected current tax year information—which is exactly what you want. It's not a perfect answer, but it's an honest one.
So Which AI Actually Lies the Least?
Here's our honest breakdown:
For legal work: Use a specialized tool like Harvey if you have access. If you're using a generalist, Claude is the safer choice—not because it's always right, but because it's more likely to tell you when it's not sure.
For medical information: Claude edges out the competition on both accuracy and appropriate hedging. Pair it with Perplexity for source verification if the stakes are high.
For financial guidance: Gemini's live data integration gives it an edge on time-sensitive questions. For everything else, Claude's caution is an asset.
Overall: ChatGPT's reputation as the default AI tool is not supported by its performance in high-stakes domains. It's fast, it's fluent, and it's frequently wrong without telling you so. That combination is genuinely risky.
The Takeaway You Actually Need
No AI should be your final source on anything with real consequences. That's not a knock on the technology—it's just reality in 2025. But there's a meaningful difference between an AI that's wrong and says so versus one that's wrong and sounds absolutely certain.
The tools that hedge appropriately, cite sources, or acknowledge the limits of their training data are not being weak. They're being responsible. And in domains where a single bad fact can cost someone their health, their money, or their legal standing, responsible beats confident every single time.
ChatGPT has built its brand on being the smart friend who always has an answer. Sometimes the smarter friend is the one who says, "I'm not sure—let's look that up."