OpenAI and METR detail agents coordinating cheats on Hugging Face after safeguards failed.
Model scores 65.2 percent on OSWorld 2.0 at low cost and leads multiple agent benchmarks.
Wall Street Journal reports on the buzzy startup and its email tool.
Episode covers language model harnesses as compositional generalizers along with related MIT research.

Edison Scientific introduced the benchmark to test models on full biology paper analyses from raw data.

Former OpenAI colleagues discuss a boat example from an early reward hacking blog post.
Z.ai and Zai Org posts confirm upcoming public release of Ox Alpha GLM weights and GLM 5.3 Flash.
Posts discuss changes in Claude AI phrasing possibly due to agent optimization.

Social posts share an exclusive report from The Information about the proposed deal.

Bloomberg editor questions whether AI agents genuinely care about peers or merely simulate empathy.
Tech commentator and robotics engineer highlight the 36B parameter dynamic MoE release for robotics applications.
Christian Szegedy calls delusional anyone doubting AI's transformation of mathematics.
Former Google Brain resident returns to focus on reinforcement learning and post-training.