AI Still Can’t Beat the On-Call Engineer: Here’s Why

by shayaan
AI Traffic to US Retailers Jumps 393% in Q1 as Agentic Shoppers Outspend Humans

In short

  • ARFBench is the first AI benchmark built entirely from real production incidents.
  • GPT-5 leads all existing AI models with 62.7% accuracy, but lags behind domain experts at 72.7%.
  • A theoretical model-expert oracle – which combines AI and human judgment – ​​achieves an accuracy of 87.2%, setting the ceiling for what collaborative AI-human teams could achieve.

AI companies continue to pitch autonomously site reliability engineer agents—AI that investigates production incidents instead of humans. Datadog ran the actual benchmark on real failures, and the best AI models can’t yet beat the engineers they’re supposed to replace.

The benchmark is ARFBench (Anomaly Reasoning Framework Benchmark), a joint project between Datadog and Carnegie Mellon. Built from 63 real production incidents extracted from engineers’ own Slack threads during live emergencies: 750 multiple-choice questions with 142 monitoring metrics and 5.38 million data points, with each question verified by hand. No synthetic data. No textbook scenarios.

“Trillions of dollars are lost every year due to system failures,” the researchers write. The benchmark tests whether AI can actually change this.

“Despite the central role of such demand-driven analysis in incident response, it remains unclear whether modern baseline models can reliably answer the kinds of time-series questions that engineers ask in practice,” the paper said.

Questions come in three levels. Level I: Is there an anomaly in this diagram? Level II: When did it start, how serious is it, what type?

The Tier III – the most difficult – requires cross-metric reasoning: is this diagram causing the problem in that other diagram? That’s where AI falls apart. GPT-5 scores just 47.5% F1 on Tier III questions, a metric that penalizes gaming answer models by choosing the most common class.

See also  Quantum Computing Threat 'Mostly a Coordination Issue' for Bitcoin: Fireblocks CEO

“Despite the central role of such demand-driven analysis in incident response, it remains unclear whether modern baseline models can reliably answer the kinds of time-series questions that engineers ask in practice,” the researchers write.

How each model stacked up

GPT-5 gave all existing models an accuracy of 62.7% – in a test where random guessing is 24.5%. Gemini 3 Pro scored 58.1%. Claude Opus 4.6: 54.8%. Claude Sonnet 4.5: 47.2%.

Domain experts scored an accuracy of 72.7%. Non-domain experts – time series researchers at Datadog without extensive observational experience – still get 69.7%.

No AI model could beat the human baseline.

Image built by Decrypt based on the ARFBench leaderboard CSV

The model that actually topped the leaderboard was Datadog’s own hybrid: Toto – their in-house time series forecasting model – combined with Qwen3-VL 32B. Toto-1.0-QA-Experimental scored an accuracy of 63.9%, surpassing GPT-5 while using a fraction of the parameters. Specifically in the area of ​​anomaly identification, the model outperformed all other models in F1 by at least 8.8 percentage points.

The expected result is a purpose-built domain model, trained on observability data, that outperforms a general-purpose state-of-the-art system on this specific task. That’s the point.

The most valuable finding is not which model scored highest.

“We observe substantially different error profiles between leading models and human experts, suggesting that their strengths are complementary,” the researchers write. Models hallucinate, miss metadata and lose domain context. People misinterpret precise timestamps and sometimes fail to execute complex instructions. The errors hardly overlap.

Model a theoretical ‘Model-Expert Oracle’ – a perfect judge that always chooses the right answer between the AI ​​and the human – and you get 87.2% accuracy and 82.8% F1. Way above either alone.

See also  SpaceX Revenue Beat Overshadowed By $540 Million Bitcoin Impairment As Public Company Era Begins

That’s not a product. It’s a documented goal – built from real emergencies, not curated data sets – that quantifies exactly how much better human-AI collaboration could perform. The rankings are live on Hugging Face. GPT-5 is at 62.7%. The ceiling is 87.2%.

Daily debriefing Newsletter

Start every day with today’s top news stories, plus original articles, a podcast, videos and more.

Source link

You may also like

Latest News

Copyright © Sovereign Wealth Signals