Google AI Overviews 90% Accuracy, 10% Error Rate

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- The New York Times analysis, conducted with startup Oumi, used the SimpleQA benchmark to assess Google AI Overviews and reported a 90% accuracy rate (10% incorrect answers).
- Oumi previously measured AI Overviews with Gemini 2.5 at 85% accuracy; after the Gemini 3 update, the accuracy rose to 91%.
- Google spokesperson Ned Adriance said SimpleQA contains incorrect information and that Google relies on a smaller, vetted benchmark called SimpleQA Verified, suggesting the study may not reflect typical user searches.
- AI Overviews dynamically selects among Gemini models, using the faster Gemini Flash for most queries and the more accurate Gemini 3.1 Pro when needed, balancing speed and cost.
- Google’s internal benchmarks for new models show factuality of 60‑80% without web grounding, while grounding with web data improves accuracy, highlighting a gap between internal and real‑world performance.
- AI Overviews has produced notable errors, such as citing wrong dates for Bob Marley’s museum and incorrectly stating that the Classical Music Hall of Fame does not exist for Yo‑Yo Ma’s induction.
Why it matters: Users who rely on AI Overviews risk being misinformed by the estimated 10% error rate, amounting to tens of millions of false answers each day, while Google faces criticism that its internal benchmarks (60‑80% factuality without web grounding) understate the problem, potentially eroding trust in its search product.


