✷ The archive · 11 labs · 3,050 dispatches
The labs
No more opening a dozen official blogs every morning. First-hand releases from OpenAI, Anthropic, DeepMind and the rest, each with its substance pulled out.
STADLER AI Implementation and Productivity Gains
STADLER, a 230-year-old industrial recycling company, achieved 30-40% time savings on knowledge tasks and 2.5x faster drafting by embedding OpenAI's ChatGPT as a company-wide productivity layer.
Hugging Face OpenClaw Migration Guide
Hugging Face provides two methods—Inference Providers and local llama.cpp setup—to migrate OpenClaw agents from restricted Claude models to open-source alternatives.
Gemini 3.1 Flash Live release notes / what's new
Google DeepMind has released Gemini 3.1 Flash Live, a high-quality audio and voice model designed for natural, real-time dialogue with improved precision, lower latency, and expanded global availability.
Google DeepMind Research on AI Harmful Manipulation
Google DeepMind has released a new research study and an empirically validated toolkit to measure and mitigate the risk of AI being used for harmful manipulation of human thought and behavior.
Lyria 3 Pro release: longer, structurally aware music generation across Google products
DeepMind announced Lyria 3 Pro, a music‑generation model that creates up to three‑minute tracks with structural awareness and is now integrated across Google products like Vertex AI, AI Studio, Google Vids, Gemini, and ProducerAI.
OpenAI Model Spec: A Framework for Explicit Model Behavior
OpenAI has introduced the Model Spec, a formal, public framework designed to make intended AI model behavior explicit, legible, and revisable for users, developers, and researchers.
OpenAI Safety Bug Bounty Program Launch
OpenAI has launched a public Safety Bug Bounty program to identify AI abuse and safety risks that fall outside conventional security vulnerabilities, specifically targeting agentic risks, proprietary information leaks, and platform integrity.
Claude Code Auto Mode: Model‑Based Permission Guardrails for Safer Autonomous Coding
Anthropic introduced Claude Code auto mode, a middle‑ground permission system that uses model classifiers to approve actions, reducing approval fatigue while preventing dangerous over‑eager behavior.
OpenAI Teen Safety Policy Pack and gpt-oss-safeguard
OpenAI has released prompt-based safety policies and the open-weight gpt-oss-safeguard model to help developers implement age-appropriate protections for teenagers in AI applications.
OpenAI expands product discovery in ChatGPT via Agentic Commerce Protocol
OpenAI has introduced enhanced visual shopping and product discovery capabilities in ChatGPT, powered by the expanded Agentic Commerce Protocol (ACP) to streamline how users find and compare products.
OpenAI Foundation Update
OpenAI has announced that its Foundation will invest at least $1 billion over the next year across life sciences, economic impact, AI resilience, and community programs to ensure AGI benefits humanity.
EVA End-to-End Evaluation Framework for Voice Agents
EVA is a new end‑to‑end framework that jointly evaluates voice agents on accuracy and conversational experience, revealing a consistent trade‑off between task success and user satisfaction.
vLLM Model Runner V2 release notes / what's new
vLLM has introduced Model Runner V2 (MRV2), a ground-up re-implementation of the model runner that improves throughput and reduces latency through a GPU-native, async-first, and modular architecture.
Anthropic Economic Index March 2026 Report: Learning Curves
Anthropic’s March 2026 Economic Index report shows that Claude usage diversified, lower‑wage tasks grew, and high‑tenure users achieve higher success rates, indicating learning‑by‑doing and potential skill‑biased impacts.
Anthropic Harness Design for Long-Running Application Development
Anthropic introduced a three‑agent harness—planner, generator, and evaluator—that enables Claude to autonomously create high‑quality front‑end designs and full‑stack applications over multi‑hour sessions, addressing context limits and self‑evaluation bias.
Mistral AI Voxtral TTS Release
Mistral AI has released Voxtral TTS, a 4B-parameter multilingual text-to-speech model that provides high-naturalness voice generation and low-latency streaming for enterprise voice agents.
OpenAI Sora 2 Safety Framework
OpenAI has detailed the safety architecture for Sora 2, focusing on provenance signals, consent-based likeness management, and strict content filtering for audio and video.
Vibe Physics: Using Claude Opus 4.5 as an AI Grad Student for Theoretical Physics
Professor Matthew Schwartz demonstrated that Claude Opus 4.5 can perform frontier theoretical physics research, reducing a year-long calculation to two weeks under expert supervision.
Long-running Claude for scientific computing
Anthropic demonstrates how multi-day agentic coding workflows using Claude Opus 4.6 can automate complex scientific computing tasks, such as implementing a differentiable cosmological Boltzmann solver, reducing months of researcher effort to days.
Anthropic Launches Science Blog to Accelerate AI-Driven Discovery
Anthropic has introduced a new Science Blog dedicated to sharing AI research, practical scientific workflows, and collaborations aimed at accelerating scientific progress.
Domain-Specific Embedding Fine-Tuning with NVIDIA Nemotron – Under a Day
NVIDIA and Hugging Face released a single‑GPU, under‑a‑day pipeline that fine‑tunes the Llama‑Nemotron‑Embed‑1B‑v2 model on synthetic domain data, delivering >10% retrieval gains and up to 26% improvement on real enterprise datasets.
OpenAI Internal Coding Agent Monitoring System
OpenAI has deployed a GPT-5.4 Thinking-powered monitoring system to detect misalignment and security violations in internal coding agents, identifying behaviors that often only emerge in complex, tool-rich workflows.
OpenAI to acquire Astral
OpenAI is acquiring Astral to integrate its open-source Python tools, including uv, Ruff, and ty, into the Codex ecosystem to enable AI agents to participate in the entire software development lifecycle.
Qwen3.5-Max-Preview Release on LMSys Arena
Qwen has deployed Qwen3.5-Max-Preview to the LMSys Arena for community evaluation ahead of its full release scheduled within two weeks.
State of Open Source on Hugging Face: Spring 2026
Hugging Face reports a massive expansion of the open source AI ecosystem in 2025, characterized by China surpassing the U.S. in model downloads and the rapid emergence of robotics as the largest dataset category.
Google DeepMind Measuring Progress Toward AGI: A Cognitive Framework
Google DeepMind has introduced a cognitive taxonomy and a three-stage evaluation protocol to empirically measure AI progress toward Artificial General Intelligence (AGI).
Mistral AI Forge: Enterprise System for Custom Frontier-Grade Models
Mistral AI has introduced Forge, a system allowing enterprises to build and continuously improve frontier-grade AI models grounded in their proprietary institutional knowledge.
Holotron-12B High Throughput Computer Use Agent
H Company released Holotron-12B, a multimodal computer-use model based on NVIDIA Nemotron-Nano-2 VL that uses a hybrid SSM-Attention architecture to achieve high inference throughput for agentic workloads.
OpenAI GPT-5.4 mini and nano release notes
OpenAI has released GPT-5.4 mini and nano, high-efficiency small models that bring GPT-5.4 capabilities to high-volume workloads with significantly lower latency and cost.
OpenAI Japan Teen Safety Blueprint
OpenAI Japan has introduced the Japan Teen Safety Blueprint, a framework prioritizing teen safety over convenience and privacy to protect younger users from AI-related risks.
OpenAI Research: How Workers Use ChatGPT for Compensation Insights
OpenAI research reveals that US workers send nearly 3 million daily messages to ChatGPT seeking wage benchmarks and compensation guidance, particularly in high-skill, low-transparency roles.
Mistral Small 4 release notes / what's new
Mistral AI has released Mistral Small 4, a unified open-source model under Apache 2.0 that integrates reasoning, multimodal, and agentic coding capabilities into a single versatile architecture.
Mistral AI and NVIDIA Partnership: The NVIDIA Nemotron Coalition
Mistral AI has joined the NVIDIA Nemotron Coalition as a founding member to co-develop open-source frontier AI models using NVIDIA's compute resources and Mistral AI's model architectures.
Mistral AI Leanstral Release
Mistral AI has released Leanstral, an open-source code agent with 6B active parameters designed for Lean 4 to enable formally verified code and mathematical proofs.
OpenAI Codex Security: Why the System Avoids SAST Report Seeding
OpenAI's Codex Security avoids starting with Static Application Security Testing (SAST) reports to prevent premature narrowing of analysis and to focus on validating whether security invariants actually hold through transformation chains.
BAIR Introducing SPEX and ProxySPEX for Scalable LLM Interaction Discovery
BAIR has introduced SPEX and ProxySPEX, algorithms that use signal processing and coding theory to identify influential interactions between features, training data, and model components at scale.
P-EAGLE: Parallel Speculative Decoding in vLLM
vLLM introduces P-EAGLE, a parallel speculative decoding method that generates all draft tokens in a single forward pass, delivering up to 1.69x speedup over vanilla EAGLE-3 on NVIDIA B200 GPUs.
Anthropic Dedicated Feature Crosscoder (DFC) for Cross-Architecture Model Diffing
Anthropic researchers have developed the Dedicated Feature Crosscoder (DFC), a tool that identifies behavioral differences between AI models with different architectures by isolating unique features, enabling the detection of "unknown unknown" risks.
Anthropic Claude Partner Network Launch
Anthropic has launched the Claude Partner Network with an initial $100 million investment to provide training, technical support, and co-investment for organizations helping enterprises adopt Claude.
Mistral AI Rails Testing Agent
Mistral AI developed an autonomous agent using Vibe to automatically generate and improve RSpec tests for Ruby on Rails monoliths, achieving 100% line coverage and zero RuboCop violations in a 275-file experiment.
OpenAI Designing AI Agents to Resist Prompt Injection
OpenAI is shifting its defense strategy against prompt injection by treating it as a social engineering problem, focusing on constraining the impact of successful manipulations rather than relying solely on input filtering.
OpenAI Responses API Computer Environment Update
OpenAI has equipped the Responses API with a shell tool and hosted container workspace, enabling models to execute real-world tasks via a command-line interface and persistent runtime context.
NVIDIA Nemotron 3 Super Support in vLLM
vLLM now supports NVIDIA Nemotron 3 Super, a 120B parameter hybrid MoE model optimized for multi-agent AI with a 1 million token context window and high inference efficiency.
Rakuten Integration of OpenAI Codex for Engineering Efficiency
Rakuten has integrated OpenAI Codex into its engineering stack, achieving a 50% reduction in mean time to recovery (MTTR) and compressing quarter-long development projects into weeks.
Wayfair OpenAI Integration Case Study
Wayfair has integrated OpenAI models into its internal systems to automate product catalog tagging for 30 million items and streamline supplier support via the AI-powered tool Wilma.
Introducing The Anthropic Institute
Anthropic has launched The Anthropic Institute, an interdisciplinary research effort led by Jack Clark to study and communicate the societal, economic, and legal challenges posed by accelerating frontier AI development.
OpenAI Instruction Hierarchy Improvements and GPT-5 Mini-R
OpenAI has introduced a new reinforcement learning dataset, IH-Challenge, to train models to prioritize trusted instructions over untrusted ones, resulting in the GPT-5 Mini-R model with improved safety steerability and prompt injection robustness.
ChatGPT Interactive Visuals for Math and Science
OpenAI has introduced dynamic visual explanations for over 70 core math and science concepts in ChatGPT to help users understand the relationships between variables and formulas in real time.
vLLM Semantic Router v0.2 Athena release notes / what's new
vLLM Semantic Router v0.2 Athena introduces a rebuilt model stack, the experimental ClawOS orchestration layer, and advanced model selection primitives to transform semantic routing into a strategic system brain for multi-agent deployments.
Hugging Face Storage Buckets Release
Hugging Face has introduced Storage Buckets, a mutable, S3-like object storage system backed by Xet for efficient handling of intermediate ML artifacts like checkpoints and processed data.