✦ For YouGeopoliticsTechFinanceHealthEnergySportsCulture◆ SN Last Week★ Saved

New SOB Benchmark Exposes 17-26 Point LLM Value Gap

By Hacker News · Summarized & edited by · 2026-04-29
New SOB Benchmark Exposes 17-26 Point LLM Value Gap

Get the Tech newsletter

Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.

Why it matters: For engineering teams wiring LLMs into data pipelines, SOB's headline finding inverts the cost calculus: the most expensive frontier models lead Value Accuracy by roughly 2 points, but open-weight models like Qwen3.5-35B and GLM-4.7 cost 6-11x less per correct field on the same API tier. The 17-26 point gap between schema validity and value grounding means existing JSON-compliance benchmarks have been telling teams their extraction pipelines are reliable when they may silently hallucinate one in every five leaf values.

Share this story

Ask SkimNews
More tech → Read original →

Get the Tech newsletter

Curated tech stories, every morning. Free.

No spam. Unsubscribe anytime.