AI News

Artificial intelligence, explained

The day's most important AI stories — researched and written for professionals. Updated daily.

LLM Evaluation in 2026: From Benchmarks to Agentic Judgment
Research

LLM Evaluation in 2026: From Benchmarks to Agentic Judgment

As frontier models saturate static benchmarks, the field is shifting toward dynamic, adversarial, and agentic evaluation—measuring judgment, tool use, and robustness rather than recall.

1 Sep 2026
A Community Benchmark for LLMs in Optimization Modeling
Research

A Community Benchmark for LLMs in Optimization Modeling

INFORMS issues a call for contributions to build a shared, rigorous benchmark for evaluating how well large language models translate real-world problems into mathematical optimization models.

1 Sep 2026
AI Forecasters Are Closing the Gap With Human Experts
Research

AI Forecasters Are Closing the Gap With Human Experts

A dynamic benchmark shows large language models improving steadily at predicting future events, while central banks test transformer-based models for macroeconomic forecasting.

1 Sep 2026
CES 2026: Autonomous Driving Hits an Inflection Point
Business

CES 2026: Autonomous Driving Hits an Inflection Point

AI-driven autonomy overtakes electrification as the industry's core narrative, with production timelines, open-source platforms, and explainable AI reshaping the road ahead.

1 Sep 2026
Best LLMs in 2026: Claude Mythos 5 Tops BenchLM Rankings
Research

Best LLMs in 2026: Claude Mythos 5 Tops BenchLM Rankings

Anthropic leads the BenchLM leaderboard, but open-weight models from Moonshot AI and Tencent are closing the gap. Here's how the scores are calculated and what they mean for deployment.

1 Sep 2026
What 30 LLM Benchmarks Really Measure—and What They Miss
Research

What 30 LLM Benchmarks Really Measure—and What They Miss

Public leaderboards drive AI procurement, but contamination, saturation, and a gap between test scores and real-world performance demand a more critical look at how models are evaluated.

1 Sep 2026
Gemini 3.1 Pro Model Card Reveals Top Benchmarks Amid 3.5 Pro Delay
Products

Gemini 3.1 Pro Model Card Reveals Top Benchmarks Amid 3.5 Pro Delay

Google DeepMind's latest model card shows strong agentic coding and reasoning scores, but the missing successor has turned Gemini 3.1 Pro into an extended flagship under intense scrutiny.

1 Sep 2026
NVIDIA’s Explainable Cars Could Finally Make AI Driving Accountable
Products

NVIDIA’s Explainable Cars Could Finally Make AI Driving Accountable

Open-source reasoning models like Alpamayo let vehicles explain their decisions in plain language, a shift that could ease regulators, insurers, and public distrust.

1 Sep 2026
Nvidia forecasts 70% sales growth as AI infrastructure boom accelerates
Business

Nvidia forecasts 70% sales growth as AI infrastructure boom accelerates

The chipmaker projects $108 billion in quarterly revenue and signals that supply constraints, not demand, are the main limit on growth.

1 Sep 2026
Anthropic’s record-breaking IPO: what it gains and risks
Business

Anthropic’s record-breaking IPO: what it gains and risks

The AI company’s planned $2 trillion listing would dwarf SpaceX’s record debut, but accounting scrutiny, capacity limits and regulatory friction cloud the outlook.

1 Sep 2026
Why Independent AI Benchmarks Keep Disappearing
Research

Why Independent AI Benchmarks Keep Disappearing

A widely shared benchmark of 40 new AI models became inaccessible, exposing the fragility of community-driven evaluation in an industry desperate for trustworthy comparisons.

31 Aug 2026
The Missing Paper on Frontier Vision-Language Model Architecture
Research

The Missing Paper on Frontier Vision-Language Model Architecture

A cited preprint on VLM architectural evolution cannot be verified, highlighting the opacity of frontier AI research even as architectural choices become multi-million-dollar decisions.

31 Aug 2026
Previous 11 / 20 Next

Newsletter

Get the AI news that matters

One short brief with the day's most important AI stories — written for professionals.

We send a confirmation link. No spam. Unsubscribe anytime.

WhatsApp