AI News
The day's most important AI stories — researched and written for professionals. Updated daily.
As frontier models saturate static benchmarks, the field is shifting toward dynamic, adversarial, and agentic evaluation—measuring judgment, tool use, and robustness rather than recall.
1 Sep 2026
INFORMS issues a call for contributions to build a shared, rigorous benchmark for evaluating how well large language models translate real-world problems into mathematical optimization models.
1 Sep 2026
A dynamic benchmark shows large language models improving steadily at predicting future events, while central banks test transformer-based models for macroeconomic forecasting.
1 Sep 2026
AI-driven autonomy overtakes electrification as the industry's core narrative, with production timelines, open-source platforms, and explainable AI reshaping the road ahead.
1 Sep 2026
Anthropic leads the BenchLM leaderboard, but open-weight models from Moonshot AI and Tencent are closing the gap. Here's how the scores are calculated and what they mean for deployment.
1 Sep 2026
Public leaderboards drive AI procurement, but contamination, saturation, and a gap between test scores and real-world performance demand a more critical look at how models are evaluated.
1 Sep 2026
Google DeepMind's latest model card shows strong agentic coding and reasoning scores, but the missing successor has turned Gemini 3.1 Pro into an extended flagship under intense scrutiny.
1 Sep 2026
Open-source reasoning models like Alpamayo let vehicles explain their decisions in plain language, a shift that could ease regulators, insurers, and public distrust.
1 Sep 2026
The chipmaker projects $108 billion in quarterly revenue and signals that supply constraints, not demand, are the main limit on growth.
1 Sep 2026
The AI company’s planned $2 trillion listing would dwarf SpaceX’s record debut, but accounting scrutiny, capacity limits and regulatory friction cloud the outlook.
1 Sep 2026
A widely shared benchmark of 40 new AI models became inaccessible, exposing the fragility of community-driven evaluation in an industry desperate for trustworthy comparisons.
31 Aug 2026
A cited preprint on VLM architectural evolution cannot be verified, highlighting the opacity of frontier AI research even as architectural choices become multi-million-dollar decisions.
31 Aug 2026Newsletter
One short brief with the day's most important AI stories — written for professionals.
We send a confirmation link. No spam. Unsubscribe anytime.