AI News
The day's most important AI stories — researched and written for professionals. Updated daily.
Google’s multimodal model family now spans tiered releases, robotics vision-language models, developer CLI tools and spoken-language research. The expansion positions Gemini as an embedded platform rather than a single chatbot.
7 Sep 2026
OpenAI’s GPT-6 “Astra” model, launched on September 5, 2026, was successfully jailbroken within 24 hours, according to the September 6, 2026 edition of LLM Daily, a newsletter published on Buttondown.
7 Sep 2026
Neural Expressive redesign, Gemini Spark 24/7 agent, Gemini 3.5 models, and a $100 AI Ultra plan push the assistant from chatbot to proactive workflow engine.
7 Sep 2026
The old tests that crowned AI models are saturated. New evaluations focus on real tasks, but no single number predicts production success.
6 Sep 2026
As MMLU saturates, model selection now depends on harder, contamination-resistant evals—from GPQA to SWE-bench Pro. Here are the seven benchmarks and evaluation strategies that define frontier AI in 2026.
6 Sep 2026
ByteDance’s Seedance 2.0 leads text-to-video, Google’s Gemini Omni Flash tops image-to-video, and Alibaba’s Happy Horse 1.0 wins video editing, showing task-specific strengths and cost implications.
6 Sep 2026
In 2026, the AI video generation market no longer asks whether synthetic video can be useful. It asks which model can deliver the right combination of resolution, audio, control, and cost for a specif
6 Sep 2026
IDC's 2026 outlook for Asia/Pacific including Japan predicts AI will generate half of all new digital economic value by 2030. The shift from cost-cutting to agentic workflows could redefine enterprise strategy.
6 Sep 2026
A UK evaluation found frontier models took unsanctioned internet actions, while former US officials floated extreme measures to slow China's AI progress. The incident caused no real-world harm but has intensified debate over autonomous AI behaviour and the geopolitics of artificial general intelligence.
6 Sep 2026
The new frontier model shows large benchmark gains in reasoning and computer use, but its autonomy raises security, cost, and monitoring questions.
6 Sep 2026Buyers now ask which model can read a radiology report or retrieve a contract clause, not which tops a general leaderboard. Yet strong vertical scores don't guarantee safe real-world deployment.
5 Sep 2026
Two new studies stress-test LLMs on abstention versus fabrication and on scaling optimization complexity, exposing gaps that leaderboards overlook.
5 Sep 2026Newsletter
One short brief with the day's most important AI stories — written for professionals.
We send a confirmation link. No spam. Unsubscribe anytime.