Google DeepMind is tired of AI hallucinations and OpenAI’s messy homework

Google DeepMind is apparently tired of your favorite AI confidently lying to your face. In a move that feels like a seasoned teacher catching a student using a broken answer key, the research giant just dropped "SimpleQA Verified." It’s a new 1,000-prompt benchmark designed to actually track whether Large Language Models (LLMs) know what they’re talking about, or if they’re just guessing with style.
It’s a direct, if somewhat snarky, response to the original SimpleQA benchmark released by the folks at OpenAI earlier this year. DeepMind claims that the previous version—designed by Wei et al. (2024)—was, well, a bit of a mess. According to the research team, the original test was riddled with "noisy and incorrect labels," topical biases, and enough question redundancy to make any serious data scientist wince.
Essentially, the very test we were using to see if AI was smart was itself hallucinating. DeepMind is stepping in to provide a "more precise instrument" to track genuine progress. By cleaning up the noise and removing the "benchmark artifacts," they’re trying to stop the industry from "overfitting"—which is just a fancy way of saying "cheating on the exam."
While one half of the lab is busy being the industry’s fact-checker, the other is playing digital Indiana Jones. DeepMind’s latest research highlights a system designed to contextualize ancient inscriptions, helping historians restore and attribute fragmentary texts that have been rotting in the dirt for centuries. It’s a fascinating contrast: using AI to bridge the gaps in human history while simultaneously trying to stop AI from inventing its own version of modern history.
We have to ask the real question here: If an LLM can help us understand a piece of 2,000-year-old stone, why does it still struggle to tell us basic facts without breaking into a cold sweat? The answer lies in "parametric knowledge"—the stuff baked into the model's brain—which is exactly what SimpleQA Verified is trying to measure. It focuses on short-form factuality, stripping away the flowery prose to see if the machine actually knows its stuff.
DeepMind says this is all about "trustworthy AI," which is corporate-speak for "we want you to stop being afraid that our products will give you a recipe for poisonous mushrooms." It’s a noble goal, but as long as these models are built on probabilistic guesses, the fight for 100% factuality is going to be an uphill battle.
Let's plan for the long game: Will the industry actually adopt a harder, "verified" test, or will they just find a new way to game the system? Either way, DeepMind is making it much harder for competitors to hide behind a flashy, but ultimately hollow, leaderboard score.
In the end, it’s refreshing to see a tech giant admit that the metrics we’ve been using are flawed. Whether they’re deciphering ancient Greek or checking OpenAI’s math, Google is clearly positioning itself as the adult in the room—at least until the next hallucination hits.
Sources: Google DeepMind Research, SimpleQA Verified Evals.


