The archive · 11 labs · 3,050 dispatches

The labs

No more opening a dozen official blogs every morning. First-hand releases from OpenAI, Anthropic, DeepMind and the rest, each with its substance pulled out.

351

Anthropic Constitutional Classifiers: Defending against universal jailbreaks

Anthropic announced Constitutional Classifiers, a new input‑output filtering system that reduces universal jailbreak success to under 5% while adding only modest compute overhead and a 0.38% increase in harmless refusals.

352

OpenAI National Security Principles and Government Partnerships

OpenAI has published its National Security Principles to provide transparency into how the company handles government partnerships and the use of its technology in national security and law enforcement.

353

OpenAI Analysis of SWE-Bench Pro Coding Evaluation Flaws

OpenAI has retracted its recommendation for SWE-Bench Pro after an audit revealed that approximately 30% of the benchmark's tasks are broken due to issues like overly strict tests and underspecified prompts.

354

Robostral Navigate Release Notes

Mistral AI has introduced Robostral Navigate, an 8B model that enables robots to autonomously navigate complex environments using only a single RGB camera, achieving 76.6% success on unseen R2R-CE benchmarks.

355

OpenAI Academy AI Skills Jam for K–12 Educators

OpenAI Academy, in partnership with the Walton Family Foundation, is launching an AI Skills Jam to provide over 1,600 K–12 educators with hands-on training to integrate AI into teaching and administration.

356

Hugging Face Native-speed vLLM Transformers Modeling Backend

Hugging Face has updated the transformers vLLM backend to match or exceed the throughput of custom vLLM implementations by dynamically applying inference-specific layer fusions at runtime.

357

Introducing GPT‑Live‑1 and GPT‑Live‑1 mini: OpenAI’s New Full‑Duplex Voice Models

OpenAI announces GPT‑Live‑1 and GPT‑Live‑1 mini, full‑duplex voice models that enable natural, interruptible conversations while delegating complex reasoning to GPT‑5.5, rolling out globally to ChatGPT users.

358

Anthropic and AE Studio Introduce GRAM for Dual-Use Knowledge Control

Anthropic and AE Studio have developed Gradient-Routed Auxiliary Modules (GRAM), a method to isolate dual-use knowledge into removable compartments to prevent misuse without sacrificing general model performance.

359

Hugging Face and Amazon SageMaker Studio Integration

Hugging Face has introduced a deep-link integration with Amazon SageMaker AI, allowing developers to move from model discovery to fine-tuning or deployment in SageMaker Studio with a single click.

360

Hugging Face Models on Foundry Managed Compute

Microsoft Foundry now integrates a curated, weekly-refreshed catalog of Hugging Face open-weight models that can be deployed in one click onto Foundry Managed Compute for enterprise-grade operationalization.

361

Intelligence is Free, Now What? Data Systems for, of, and by Agents

BAIR researchers propose a new framework for data systems redesigned for agentic workloads, agentic state management, and agent-driven system synthesis as AI inference costs approach zero.

362

Hugging Face and SkyPilot Integration for Zero-Egress AI Storage

Hugging Face and SkyPilot have integrated to allow AI workloads to run on any cloud provider with zero-egress costs for reading models and datasets stored on the Hugging Face Hub.

363

MUFG AI-Native Transformation with OpenAI

MUFG is partnering with OpenAI to implement ChatGPT Enterprise for 35,000 employees and develop AI-driven customer experiences to transition into an AI-native financial group.

364

Australian Payments Plus Integration of ChatGPT Enterprise and Codex

Australian Payments Plus (AP+) has deployed ChatGPT Enterprise and Codex to accelerate technical investigations, streamline complex knowledge work, and reduce product prototyping time from weeks to a single day.

365

LeRobot v0.6.0 release notes / what's new

LeRobot v0.6.0 introduces world model policies, a expanded VLA model zoo, a unified reward models API, and a new deployment CLI with DAgger-style human-in-the-loop corrections.

366

PRX Data Strategy: Scaling Pre-training with VLM Re-captioning and Mosaic Streaming

Photoroom details the data pipeline for PRX, emphasizing the use of long VLM-generated captions, a hybrid Lance and Mosaic Data Shards storage strategy, and high-quality JPEG encoding to optimize a 7B parameter model.

367

vLLM × HPC-Ops: High-Performance Attention and MoE Backends from Tencent Hunyuan

vLLM now includes HPC-Ops attention and MoE backends optimized for NVIDIA Hopper H20, delivering up to 2.95× attention speedup and 1.59× MoE speedup, cutting TTFT by about 24% and TPOT by about 17% on Hy3.

368

Hugging Face Kernels Major Updates

Hugging Face announced major updates to its Kernels project, introducing a new kernel repository type, trusted publishers and code signing, revamped CLIs, broader framework support, and foundations for agentic kernel development.

369

Anthropic discovers a global workspace (J-space) in Claude language models

Anthropic announced the discovery of a emergent “J-space” in Claude that functions as a global workspace for internal, reportable reasoning, enabling silent thought monitoring and new safety tools.

370

Alberta Government Uses Claude Code to Scan and Fix 466M Lines of Code in Hours

The Government of Alberta deployed Anthropic's Claude Code (Opus and Sonnet) to scan 466 million lines of code across 3,400 repositories in 20 hours, automatically generate fixes, and establish continuous AI‑driven security reviews, dramatically reducing a task that would have taken years.

371

xAI Grok Voice Update: 21 New Flagship Voices

xAI has released 21 new multilingual flagship voices for Grok Voice, alongside improvements to the original five voices and the introduction of the Grok Voice Agent Builder.

372

Google DeepMind and A24 Research Partnership

Google DeepMind and A24 have entered a research partnership to integrate AI innovation directly into the creative process to develop new filmmaking workflows and techniques.

373

Leanstral 1.5 Release: State‑of‑the‑art Formal Verification Model

Mistral AI released Leanstral 1.5, a 119 B‑parameter, Apache‑2.0 model that saturates miniF2F, solves 587 of 672 PutnamBench problems, and discovers previously unknown bugs in open‑source code.

374

Claude Fable 5 Cybersecurity Safeguards and Jailbreak Framework

Anthropic has detailed the safety classifiers used to prevent cybersecurity misuse in Claude Fable 5 and proposed a new Cyber Jailbreak Severity (CJS) framework to standardize how AI jailbreaks are assessed.

375

Anthropic Restores Access to Claude Fable 5 and Mythos 5

Anthropic announced that Claude Fable 5 and Mythos 5 are back online after U.S. export controls were lifted, detailing new safety classifiers, an industry jailbreak framework, and expanded government collaboration.

376

Claude Fable 5 and Claude Mythos 5 launch – state-of-the-art Mythos-class model with safety safeguards and tiered access

Anthropic released Claude Fable 5, a general‑use Mythos‑class 1 model with record performance across software engineering, vision, and science, and Claude Mythos 5, a lifted‑safeguard version for trusted cyber‑defenders and future biology partners.

377

BAIR 2026 Graduate Showcase

The Berkeley Artificial Intelligence Research (BAIR) Lab announced its class of 2026 Ph.D. graduates, highlighting research across robotics, large language models, AI safety, and healthcare.

378

Hugging Face and Cerebras Real-Time Voice AI with Gemma 4

Hugging Face and Cerebras have developed an open, cascaded speech-to-speech pipeline using Gemma 4 31B to enable natural, low-latency voice AI interactions.

379

vLLM-Omni Optimizations for Qwen3-Omni-30B-A3B-Instruct Serving

vLLM-Omni serves Qwen3-Omni-30B-A3B-Instruct via a three-stage pipeline (Thinker, Talker, Code2Wav) and improves throughput and latency using stage decomposition, CUDA Graphs, async chunk handoffs, async output, stage replicas, and hot‑path cleanup.

380

xAI Voice Agent Builder Beta Release

xAI has launched the Voice Agent Builder in beta, a no-code platform for creating production-ready voice agents powered by Grok Voice.

381

ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration

IBM Research introduces ScarfBench, an open benchmark to evaluate AI agents' ability to migrate enterprise Java applications across Spring, Jakarta EE, and Quarkus frameworks.

382

Nano Banana 2 Lite and Gemini Omni Flash Release

Google DeepMind has released Nano Banana 2 Lite for high-speed, cost-efficient image generation and Gemini Omni Flash for high-quality video generation and conversational editing.

383

Why Specialization Is Inevitable in AI Systems

Based on the 2026 research by Goldfeder et al., specialization is a structural necessity for AI performance because finite resources make concentrated capacity more effective than broad generality.

384

OpenAI Signals: Analysis of Global ChatGPT Adoption Trends

OpenAI Signals data reveals that ChatGPT usage is deepening and expanding globally, with users increasing their daily message volume by 50% and doubling the variety of tasks performed six months after signup.

385

Anthropic Frontier Red Team: Stress-Testing AI Capabilities

Anthropic's Frontier Red Team conducts evidence-based analysis on AI's implications for cybersecurity, national security, and autonomous systems to anticipate future risks.

386

Hugging Face Community Evals and Every Eval Ever (EEE) Integration

Hugging Face has integrated Community Evals with the Every Eval Ever (EEE) project to enable cross-posting of standardized evaluation results and direct linking between model pages and detailed technical records.

387

OpenAI Core Dump Epidemiology: Fixing an 18-Year-Old libunwind Bug

OpenAI identified and resolved two distinct crash populations in its Rockset data infrastructure, including a silent hardware failure and a rare 18-year-old race condition in GNU libunwind.

388

OpenAI Introduces GeneBench-Pro for Computational Biology Reasoning

OpenAI has released GeneBench-Pro, a research-level benchmark designed to evaluate the high-level scientific judgment and iterative analysis capabilities of AI agents in computational biology.

389

Inside Genebench-Pro: Case Studies in Genomic AI Benchmarking

OpenAI's Genebench-Pro provides a specialized benchmark for evaluating AI models on complex genomic tasks, featuring ten detailed case studies across somatic oncology, functional genomics, and population genetics.

390

Claude Science AI Workbench Release

Anthropic has launched Claude Science, an AI workbench that integrates scientific tools, manages scalable compute, and provides auditable research artifacts to accelerate scientific discovery.

391

DiScoFormer: One transformer for density and score, across distributions

DiScoFormer is a new transformer-based model that estimates both the density and score of a distribution from a set of data points in a single forward pass without requiring retraining for new distributions.

392

OpenAI Mapping Europe’s AI Workforce Opportunity Report 2026

OpenAI’s 2026 report extends its AI Jobs Transition Framework to the EU, finding that about 12% of EU employment may grow with AI, 14% faces higher near‑term automation potential, 27% is likely to reorganize, and 47% sees less immediate change.

393

vLLM Semantic Router: Enhancing Model Performance via Micro-Agent Collaboration

vLLM introduces the Semantic Router, a serving-layer primitive that transforms a single model API call into a bounded collaboration of micro-agents to outperform frontier models on complex benchmarks.

394

Ollama 0.31: Faster Gemma 4 Performance via Multi-Token Prediction

Ollama 0.31 introduces multi-token prediction for Gemma 4 on Apple Silicon, increasing token generation speed by nearly 90% on coding-agent benchmarks without altering model output.

395

HP Inc. and OpenAI Strategic Partnership

HP Inc. is scaling its strategic partnership with OpenAI using the OpenAI Frontier platform to deploy AI agents and workflows across customer support, security, and software development.

396

GPT-5.6 Sol preview: capabilities, safeguards, and availability

OpenAI previewed GPT-5.6 Sol, its flagship next‑generation model, alongside Terra and Luna, highlighting stronger coding, biology, and cybersecurity capabilities with enhanced safeguards and a limited release to trusted partners.

397

Hugging Face Jobs: Deploying vLLM Servers with a Single Command

Hugging Face now allows users to deploy private, OpenAI-compatible vLLM endpoints on its infrastructure using a single command via HF Jobs, providing a pay-per-second alternative for testing, evaluations, and batch generation.

398

Anthropic Economic Index Cadences Report – June 2026

Anthropic’s June 2026 Economic Index report reveals that Claude usage now follows daily and weekly work cadences, produces higher‑value artifacts that consume more compute, and that users who automate more tasks are surprisingly more optimistic about AI’s impact on their jobs.

399

OpenAI Codex Agentic AI Adoption and Impact

OpenAI reports a shift in knowledge work from short chatbot interactions to long-horizon agentic tasks via Codex, with non-developer adoption growing up to 189x since August 2025.

400

Gemini 3.5 Flash Computer Use Integration

Google DeepMind has integrated computer use as a native tool in Gemini 3.5 Flash, enabling agents to see, reason, and take action across browser, mobile, and desktop environments.