AI Models Can’t Agree on Basic Facts Most of the Time, Study Shows

by shayaan

In short

  • Five groundbreaking AI models disagreed on 67% of 1,000 actual fact-check claims.
  • Unanimous agreement was reached on only 328 claims.
  • With a Krippendorff alpha of 0.639, the models fall below the reliability threshold of 0.8.

Ask five of the world’s most advanced AI systems if a statement is true, and two-thirds of the time, at least one will give a different answer. That is the finding of A new study published this month by researcher Kosta Jordanov of Lenz Research.

The survey found GPT-5.4, Claude Opus 4.7, Gemini 3 Pro, Gemini 3 Pro with Search, and Sonar Pro the same 1,000 real-world fact-check claims submitted by actual users. The models had to choose one of four labels: true, mostly true, misleading, or false.

In 672 of the 1,000 claims, at least one model broke away from the majority. In 34% of cases, the disagreement was serious: one model called a statement true, while the other called it false.

“These are not benchmark items with public answer keys; they are claims that real users have been submitted to a fact-checking platform for verification,” the study reads. “Only one judgment bucket can be correct per claim, so any disagreement between the panel means that at least one model’s judgment is inconsistent under this four-bucket heading.”

Previous studies on AI hallucination have shown that chatbots make up facts. That’s one problem. This is another one. The models don’t necessarily make things up, they just can’t agree on basic factual judgments about the same material.

The research used a design that makes it more difficult for the AI ​​companies to explain anything away. Instead of pulling claims from standard test sets — the kind that often leak into training data — the researchers used claims submitted by real people to Lenz’s fact-checking platform. “Most of these statements are unlikely to appear in a training corpus with a gold label attached – there is no canonical answer key to match patterns against, no benchmark leaderboard to anchor to,” the paper notes.

See also  UK Panel Calls Crypto Donations 'High Risk,' Seeks Immediate Ban

The statistical measure of agreement, called Krippendorff’s alpha, came to 0.639 on a scale where 1.0 means perfect agreement and 0 means random chance. The study says this indicates “non-trivial but limited agreement.” “The models’ rulings are structured rather than random, but not consistent enough to treat the panel as one interchangeable judge,” the researchers note. Researchers generally consider anything below 0.8 as weak.

When all five models agreed – which happened on only 328 of 1,000 claims – they almost never agreed that something was misleading or mostly true. Only four claims received a unanimous “misleading” verdict. Zero received a unanimous ‘mostly true’ rating.

The researchers provided example claims where the AI ​​models showed the most differences, including “The World Bank’s active portfolio in Nigeria will exceed $16.4 billion by 2025.” ChatGPT 5.4 said it was “mostly true,” while Gemini 3 Pro called it “false” and its sister model Gemini 3 Pro + Search rated it “misleading.”

In another example, the models included the claim: “Donald Trump said an attack on Iran was postponed at the request of Gulf allies.” GPT-5.4 said it was false, Claude Opus 4.7 called it mostly true, Gemini 3 Pro said false, and Gemini 3 Pro + Search rated it true.

“The panel comes together to make final statements; the middle of the rubric is where it breaks,” the researchers found. There was only unanimity at the extremes: the statement was either definitely true or definitely false.

This matters as people increasingly turn to AI systems for fact-checking. If you paste a claim from a news article into ChatGPT, Claude, or Gemini, you may get three different responses. Which one do you trust?

See also  Trump Filing Shows $1.4 Billion In 2025 Crypto-Linked Earnings

AI companies are happy to tell you that their models are becoming increasingly accurate. They publish benchmark scores that show steady improvement. But Lenz’s research tested these models on the kinds of erratic, ambiguous claims that real people actually argue about—and found that the models do just that.

The newspaper is careful to point this out. “A majority of boundary models is not ground truth. The majority’s judgment is sometimes wrong; an individually deviating model is sometimes right. We use the majority as a structural reference point for measuring disagreement, not as a substitute for correctness.”

There is a deeper problem hidden in the numbers. If models disagree, at least one of them must be wrong. The study calls one model’s judgment “label inconsistent under this four-bucket rubric.” There is no tiebreaker mechanism, no appeals court. Recent reporting about the reliability of AI has raised similar alarms.

Of the 328 statements that all five models agreed on, zero was unanimously rated “mostly true.” The nuance bucket was completely empty. If AI models can only find consensus at the extremes, can they even be trusted as fact-checkers?

Daily debriefing Newsletter

Start every day with today’s top news stories, plus original articles, a podcast, videos and more.

Source link

You may also like

Latest News

Copyright © Sovereign Wealth Signals