The archive · 11 labs · 3,050 dispatches

The labs

No more opening a dozen official blogs every morning. First-hand releases from OpenAI, Anthropic, DeepMind and the rest, each with its substance pulled out.

801

STADLER AI Implementation and Productivity Gains

STADLER, a 230-year-old industrial recycling company, achieved 30-40% time savings on knowledge tasks and 2.5x faster drafting by embedding OpenAI's ChatGPT as a company-wide productivity layer.

802

Hugging Face OpenClaw Migration Guide

Hugging Face provides two methods—Inference Providers and local llama.cpp setup—to migrate OpenClaw agents from restricted Claude models to open-source alternatives.

803

Gemini 3.1 Flash Live release notes / what's new

Google DeepMind has released Gemini 3.1 Flash Live, a high-quality audio and voice model designed for natural, real-time dialogue with improved precision, lower latency, and expanded global availability.

804

Google DeepMind Research on AI Harmful Manipulation

Google DeepMind has released a new research study and an empirically validated toolkit to measure and mitigate the risk of AI being used for harmful manipulation of human thought and behavior.

805

Lyria 3 Pro release: longer, structurally aware music generation across Google products

DeepMind announced Lyria 3 Pro, a music‑generation model that creates up to three‑minute tracks with structural awareness and is now integrated across Google products like Vertex AI, AI Studio, Google Vids, Gemini, and ProducerAI.

806

OpenAI Model Spec: A Framework for Explicit Model Behavior

OpenAI has introduced the Model Spec, a formal, public framework designed to make intended AI model behavior explicit, legible, and revisable for users, developers, and researchers.

807

OpenAI Safety Bug Bounty Program Launch

OpenAI has launched a public Safety Bug Bounty program to identify AI abuse and safety risks that fall outside conventional security vulnerabilities, specifically targeting agentic risks, proprietary information leaks, and platform integrity.

808

Claude Code Auto Mode: Model‑Based Permission Guardrails for Safer Autonomous Coding

Anthropic introduced Claude Code auto mode, a middle‑ground permission system that uses model classifiers to approve actions, reducing approval fatigue while preventing dangerous over‑eager behavior.

809

OpenAI Teen Safety Policy Pack and gpt-oss-safeguard

OpenAI has released prompt-based safety policies and the open-weight gpt-oss-safeguard model to help developers implement age-appropriate protections for teenagers in AI applications.

810

OpenAI expands product discovery in ChatGPT via Agentic Commerce Protocol

OpenAI has introduced enhanced visual shopping and product discovery capabilities in ChatGPT, powered by the expanded Agentic Commerce Protocol (ACP) to streamline how users find and compare products.

811

OpenAI Foundation Update

OpenAI has announced that its Foundation will invest at least $1 billion over the next year across life sciences, economic impact, AI resilience, and community programs to ensure AGI benefits humanity.

812

EVA End-to-End Evaluation Framework for Voice Agents

EVA is a new end‑to‑end framework that jointly evaluates voice agents on accuracy and conversational experience, revealing a consistent trade‑off between task success and user satisfaction.

813

vLLM Model Runner V2 release notes / what's new

vLLM has introduced Model Runner V2 (MRV2), a ground-up re-implementation of the model runner that improves throughput and reduces latency through a GPU-native, async-first, and modular architecture.

814

Anthropic Economic Index March 2026 Report: Learning Curves

Anthropic’s March 2026 Economic Index report shows that Claude usage diversified, lower‑wage tasks grew, and high‑tenure users achieve higher success rates, indicating learning‑by‑doing and potential skill‑biased impacts.

815

Anthropic Harness Design for Long-Running Application Development

Anthropic introduced a three‑agent harness—planner, generator, and evaluator—that enables Claude to autonomously create high‑quality front‑end designs and full‑stack applications over multi‑hour sessions, addressing context limits and self‑evaluation bias.

816

Mistral AI Voxtral TTS Release

Mistral AI has released Voxtral TTS, a 4B-parameter multilingual text-to-speech model that provides high-naturalness voice generation and low-latency streaming for enterprise voice agents.

817

OpenAI Sora 2 Safety Framework

OpenAI has detailed the safety architecture for Sora 2, focusing on provenance signals, consent-based likeness management, and strict content filtering for audio and video.

818

Vibe Physics: Using Claude Opus 4.5 as an AI Grad Student for Theoretical Physics

Professor Matthew Schwartz demonstrated that Claude Opus 4.5 can perform frontier theoretical physics research, reducing a year-long calculation to two weeks under expert supervision.

819

Long-running Claude for scientific computing

Anthropic demonstrates how multi-day agentic coding workflows using Claude Opus 4.6 can automate complex scientific computing tasks, such as implementing a differentiable cosmological Boltzmann solver, reducing months of researcher effort to days.

820

Anthropic Launches Science Blog to Accelerate AI-Driven Discovery

Anthropic has introduced a new Science Blog dedicated to sharing AI research, practical scientific workflows, and collaborations aimed at accelerating scientific progress.

821

Domain-Specific Embedding Fine-Tuning with NVIDIA Nemotron – Under a Day

NVIDIA and Hugging Face released a single‑GPU, under‑a‑day pipeline that fine‑tunes the Llama‑Nemotron‑Embed‑1B‑v2 model on synthetic domain data, delivering >10% retrieval gains and up to 26% improvement on real enterprise datasets.

822

OpenAI Internal Coding Agent Monitoring System

OpenAI has deployed a GPT-5.4 Thinking-powered monitoring system to detect misalignment and security violations in internal coding agents, identifying behaviors that often only emerge in complex, tool-rich workflows.

823

OpenAI to acquire Astral

OpenAI is acquiring Astral to integrate its open-source Python tools, including uv, Ruff, and ty, into the Codex ecosystem to enable AI agents to participate in the entire software development lifecycle.

824

Qwen3.5-Max-Preview Release on LMSys Arena

Qwen has deployed Qwen3.5-Max-Preview to the LMSys Arena for community evaluation ahead of its full release scheduled within two weeks.

825

State of Open Source on Hugging Face: Spring 2026

Hugging Face reports a massive expansion of the open source AI ecosystem in 2025, characterized by China surpassing the U.S. in model downloads and the rapid emergence of robotics as the largest dataset category.

826

Google DeepMind Measuring Progress Toward AGI: A Cognitive Framework

Google DeepMind has introduced a cognitive taxonomy and a three-stage evaluation protocol to empirically measure AI progress toward Artificial General Intelligence (AGI).

827

Mistral AI Forge: Enterprise System for Custom Frontier-Grade Models

Mistral AI has introduced Forge, a system allowing enterprises to build and continuously improve frontier-grade AI models grounded in their proprietary institutional knowledge.

828

Holotron-12B High Throughput Computer Use Agent

H Company released Holotron-12B, a multimodal computer-use model based on NVIDIA Nemotron-Nano-2 VL that uses a hybrid SSM-Attention architecture to achieve high inference throughput for agentic workloads.

829

OpenAI GPT-5.4 mini and nano release notes

OpenAI has released GPT-5.4 mini and nano, high-efficiency small models that bring GPT-5.4 capabilities to high-volume workloads with significantly lower latency and cost.

830

OpenAI Japan Teen Safety Blueprint

OpenAI Japan has introduced the Japan Teen Safety Blueprint, a framework prioritizing teen safety over convenience and privacy to protect younger users from AI-related risks.

831

OpenAI Research: How Workers Use ChatGPT for Compensation Insights

OpenAI research reveals that US workers send nearly 3 million daily messages to ChatGPT seeking wage benchmarks and compensation guidance, particularly in high-skill, low-transparency roles.

832

Mistral Small 4 release notes / what's new

Mistral AI has released Mistral Small 4, a unified open-source model under Apache 2.0 that integrates reasoning, multimodal, and agentic coding capabilities into a single versatile architecture.

833

Mistral AI and NVIDIA Partnership: The NVIDIA Nemotron Coalition

Mistral AI has joined the NVIDIA Nemotron Coalition as a founding member to co-develop open-source frontier AI models using NVIDIA's compute resources and Mistral AI's model architectures.

834

Mistral AI Leanstral Release

Mistral AI has released Leanstral, an open-source code agent with 6B active parameters designed for Lean 4 to enable formally verified code and mathematical proofs.

835

OpenAI Codex Security: Why the System Avoids SAST Report Seeding

OpenAI's Codex Security avoids starting with Static Application Security Testing (SAST) reports to prevent premature narrowing of analysis and to focus on validating whether security invariants actually hold through transformation chains.

836

BAIR Introducing SPEX and ProxySPEX for Scalable LLM Interaction Discovery

BAIR has introduced SPEX and ProxySPEX, algorithms that use signal processing and coding theory to identify influential interactions between features, training data, and model components at scale.

837

P-EAGLE: Parallel Speculative Decoding in vLLM

vLLM introduces P-EAGLE, a parallel speculative decoding method that generates all draft tokens in a single forward pass, delivering up to 1.69x speedup over vanilla EAGLE-3 on NVIDIA B200 GPUs.

838

Anthropic Dedicated Feature Crosscoder (DFC) for Cross-Architecture Model Diffing

Anthropic researchers have developed the Dedicated Feature Crosscoder (DFC), a tool that identifies behavioral differences between AI models with different architectures by isolating unique features, enabling the detection of "unknown unknown" risks.

839

Anthropic Claude Partner Network Launch

Anthropic has launched the Claude Partner Network with an initial $100 million investment to provide training, technical support, and co-investment for organizations helping enterprises adopt Claude.

840

Mistral AI Rails Testing Agent

Mistral AI developed an autonomous agent using Vibe to automatically generate and improve RSpec tests for Ruby on Rails monoliths, achieving 100% line coverage and zero RuboCop violations in a 275-file experiment.

841

OpenAI Designing AI Agents to Resist Prompt Injection

OpenAI is shifting its defense strategy against prompt injection by treating it as a social engineering problem, focusing on constraining the impact of successful manipulations rather than relying solely on input filtering.

842

OpenAI Responses API Computer Environment Update

OpenAI has equipped the Responses API with a shell tool and hosted container workspace, enabling models to execute real-world tasks via a command-line interface and persistent runtime context.

843

NVIDIA Nemotron 3 Super Support in vLLM

vLLM now supports NVIDIA Nemotron 3 Super, a 120B parameter hybrid MoE model optimized for multi-agent AI with a 1 million token context window and high inference efficiency.

844

Rakuten Integration of OpenAI Codex for Engineering Efficiency

Rakuten has integrated OpenAI Codex into its engineering stack, achieving a 50% reduction in mean time to recovery (MTTR) and compressing quarter-long development projects into weeks.

845

Wayfair OpenAI Integration Case Study

Wayfair has integrated OpenAI models into its internal systems to automate product catalog tagging for 30 million items and streamline supplier support via the AI-powered tool Wilma.

846

Introducing The Anthropic Institute

Anthropic has launched The Anthropic Institute, an interdisciplinary research effort led by Jack Clark to study and communicate the societal, economic, and legal challenges posed by accelerating frontier AI development.

847

OpenAI Instruction Hierarchy Improvements and GPT-5 Mini-R

OpenAI has introduced a new reinforcement learning dataset, IH-Challenge, to train models to prioritize trusted instructions over untrusted ones, resulting in the GPT-5 Mini-R model with improved safety steerability and prompt injection robustness.

848

ChatGPT Interactive Visuals for Math and Science

OpenAI has introduced dynamic visual explanations for over 70 core math and science concepts in ChatGPT to help users understand the relationships between variables and formulas in real time.

849

vLLM Semantic Router v0.2 Athena release notes / what's new

vLLM Semantic Router v0.2 Athena introduces a rebuilt model stack, the experimental ClawOS orchestration layer, and advanced model selection primitives to transform semantic routing into a strategic system brain for multi-agent deployments.

850

Hugging Face Storage Buckets Release

Hugging Face has introduced Storage Buckets, a mutable, S3-like object storage system backed by Xet for efficient handling of intermediate ML artifacts like checkpoints and processed data.