✷ The archive · 11 labs · 3,050 dispatches
The labs
No more opening a dozen official blogs every morning. First-hand releases from OpenAI, Anthropic, DeepMind and the rest, each with its substance pulled out.
Anthropic Constitutional Classifiers: Defending against universal jailbreaks
Anthropic announced Constitutional Classifiers, a new input‑output filtering system that reduces universal jailbreak success to under 5% while adding only modest compute overhead and a 0.38% increase in harmless refusals.
OpenAI National Security Principles and Government Partnerships
OpenAI has published its National Security Principles to provide transparency into how the company handles government partnerships and the use of its technology in national security and law enforcement.
OpenAI Analysis of SWE-Bench Pro Coding Evaluation Flaws
OpenAI has retracted its recommendation for SWE-Bench Pro after an audit revealed that approximately 30% of the benchmark's tasks are broken due to issues like overly strict tests and underspecified prompts.
Robostral Navigate Release Notes
Mistral AI has introduced Robostral Navigate, an 8B model that enables robots to autonomously navigate complex environments using only a single RGB camera, achieving 76.6% success on unseen R2R-CE benchmarks.
OpenAI Academy AI Skills Jam for K–12 Educators
OpenAI Academy, in partnership with the Walton Family Foundation, is launching an AI Skills Jam to provide over 1,600 K–12 educators with hands-on training to integrate AI into teaching and administration.
Hugging Face Native-speed vLLM Transformers Modeling Backend
Hugging Face has updated the transformers vLLM backend to match or exceed the throughput of custom vLLM implementations by dynamically applying inference-specific layer fusions at runtime.
Introducing GPT‑Live‑1 and GPT‑Live‑1 mini: OpenAI’s New Full‑Duplex Voice Models
OpenAI announces GPT‑Live‑1 and GPT‑Live‑1 mini, full‑duplex voice models that enable natural, interruptible conversations while delegating complex reasoning to GPT‑5.5, rolling out globally to ChatGPT users.
Anthropic and AE Studio Introduce GRAM for Dual-Use Knowledge Control
Anthropic and AE Studio have developed Gradient-Routed Auxiliary Modules (GRAM), a method to isolate dual-use knowledge into removable compartments to prevent misuse without sacrificing general model performance.
Hugging Face and Amazon SageMaker Studio Integration
Hugging Face has introduced a deep-link integration with Amazon SageMaker AI, allowing developers to move from model discovery to fine-tuning or deployment in SageMaker Studio with a single click.
Hugging Face Models on Foundry Managed Compute
Microsoft Foundry now integrates a curated, weekly-refreshed catalog of Hugging Face open-weight models that can be deployed in one click onto Foundry Managed Compute for enterprise-grade operationalization.
Intelligence is Free, Now What? Data Systems for, of, and by Agents
BAIR researchers propose a new framework for data systems redesigned for agentic workloads, agentic state management, and agent-driven system synthesis as AI inference costs approach zero.
Hugging Face and SkyPilot Integration for Zero-Egress AI Storage
Hugging Face and SkyPilot have integrated to allow AI workloads to run on any cloud provider with zero-egress costs for reading models and datasets stored on the Hugging Face Hub.
MUFG AI-Native Transformation with OpenAI
MUFG is partnering with OpenAI to implement ChatGPT Enterprise for 35,000 employees and develop AI-driven customer experiences to transition into an AI-native financial group.
Australian Payments Plus Integration of ChatGPT Enterprise and Codex
Australian Payments Plus (AP+) has deployed ChatGPT Enterprise and Codex to accelerate technical investigations, streamline complex knowledge work, and reduce product prototyping time from weeks to a single day.
LeRobot v0.6.0 release notes / what's new
LeRobot v0.6.0 introduces world model policies, a expanded VLA model zoo, a unified reward models API, and a new deployment CLI with DAgger-style human-in-the-loop corrections.
PRX Data Strategy: Scaling Pre-training with VLM Re-captioning and Mosaic Streaming
Photoroom details the data pipeline for PRX, emphasizing the use of long VLM-generated captions, a hybrid Lance and Mosaic Data Shards storage strategy, and high-quality JPEG encoding to optimize a 7B parameter model.
vLLM × HPC-Ops: High-Performance Attention and MoE Backends from Tencent Hunyuan
vLLM now includes HPC-Ops attention and MoE backends optimized for NVIDIA Hopper H20, delivering up to 2.95× attention speedup and 1.59× MoE speedup, cutting TTFT by about 24% and TPOT by about 17% on Hy3.
Hugging Face Kernels Major Updates
Hugging Face announced major updates to its Kernels project, introducing a new kernel repository type, trusted publishers and code signing, revamped CLIs, broader framework support, and foundations for agentic kernel development.
Anthropic discovers a global workspace (J-space) in Claude language models
Anthropic announced the discovery of a emergent “J-space” in Claude that functions as a global workspace for internal, reportable reasoning, enabling silent thought monitoring and new safety tools.
Alberta Government Uses Claude Code to Scan and Fix 466M Lines of Code in Hours
The Government of Alberta deployed Anthropic's Claude Code (Opus and Sonnet) to scan 466 million lines of code across 3,400 repositories in 20 hours, automatically generate fixes, and establish continuous AI‑driven security reviews, dramatically reducing a task that would have taken years.
xAI Grok Voice Update: 21 New Flagship Voices
xAI has released 21 new multilingual flagship voices for Grok Voice, alongside improvements to the original five voices and the introduction of the Grok Voice Agent Builder.
Google DeepMind and A24 Research Partnership
Google DeepMind and A24 have entered a research partnership to integrate AI innovation directly into the creative process to develop new filmmaking workflows and techniques.
Leanstral 1.5 Release: State‑of‑the‑art Formal Verification Model
Mistral AI released Leanstral 1.5, a 119 B‑parameter, Apache‑2.0 model that saturates miniF2F, solves 587 of 672 PutnamBench problems, and discovers previously unknown bugs in open‑source code.
Claude Fable 5 Cybersecurity Safeguards and Jailbreak Framework
Anthropic has detailed the safety classifiers used to prevent cybersecurity misuse in Claude Fable 5 and proposed a new Cyber Jailbreak Severity (CJS) framework to standardize how AI jailbreaks are assessed.
Anthropic Restores Access to Claude Fable 5 and Mythos 5
Anthropic announced that Claude Fable 5 and Mythos 5 are back online after U.S. export controls were lifted, detailing new safety classifiers, an industry jailbreak framework, and expanded government collaboration.
Claude Fable 5 and Claude Mythos 5 launch – state-of-the-art Mythos-class model with safety safeguards and tiered access
Anthropic released Claude Fable 5, a general‑use Mythos‑class 1 model with record performance across software engineering, vision, and science, and Claude Mythos 5, a lifted‑safeguard version for trusted cyber‑defenders and future biology partners.
BAIR 2026 Graduate Showcase
The Berkeley Artificial Intelligence Research (BAIR) Lab announced its class of 2026 Ph.D. graduates, highlighting research across robotics, large language models, AI safety, and healthcare.
Hugging Face and Cerebras Real-Time Voice AI with Gemma 4
Hugging Face and Cerebras have developed an open, cascaded speech-to-speech pipeline using Gemma 4 31B to enable natural, low-latency voice AI interactions.
vLLM-Omni Optimizations for Qwen3-Omni-30B-A3B-Instruct Serving
vLLM-Omni serves Qwen3-Omni-30B-A3B-Instruct via a three-stage pipeline (Thinker, Talker, Code2Wav) and improves throughput and latency using stage decomposition, CUDA Graphs, async chunk handoffs, async output, stage replicas, and hot‑path cleanup.
xAI Voice Agent Builder Beta Release
xAI has launched the Voice Agent Builder in beta, a no-code platform for creating production-ready voice agents powered by Grok Voice.
ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration
IBM Research introduces ScarfBench, an open benchmark to evaluate AI agents' ability to migrate enterprise Java applications across Spring, Jakarta EE, and Quarkus frameworks.
Nano Banana 2 Lite and Gemini Omni Flash Release
Google DeepMind has released Nano Banana 2 Lite for high-speed, cost-efficient image generation and Gemini Omni Flash for high-quality video generation and conversational editing.
Why Specialization Is Inevitable in AI Systems
Based on the 2026 research by Goldfeder et al., specialization is a structural necessity for AI performance because finite resources make concentrated capacity more effective than broad generality.
OpenAI Signals: Analysis of Global ChatGPT Adoption Trends
OpenAI Signals data reveals that ChatGPT usage is deepening and expanding globally, with users increasing their daily message volume by 50% and doubling the variety of tasks performed six months after signup.
Anthropic Frontier Red Team: Stress-Testing AI Capabilities
Anthropic's Frontier Red Team conducts evidence-based analysis on AI's implications for cybersecurity, national security, and autonomous systems to anticipate future risks.
Hugging Face Community Evals and Every Eval Ever (EEE) Integration
Hugging Face has integrated Community Evals with the Every Eval Ever (EEE) project to enable cross-posting of standardized evaluation results and direct linking between model pages and detailed technical records.
OpenAI Core Dump Epidemiology: Fixing an 18-Year-Old libunwind Bug
OpenAI identified and resolved two distinct crash populations in its Rockset data infrastructure, including a silent hardware failure and a rare 18-year-old race condition in GNU libunwind.
OpenAI Introduces GeneBench-Pro for Computational Biology Reasoning
OpenAI has released GeneBench-Pro, a research-level benchmark designed to evaluate the high-level scientific judgment and iterative analysis capabilities of AI agents in computational biology.
Inside Genebench-Pro: Case Studies in Genomic AI Benchmarking
OpenAI's Genebench-Pro provides a specialized benchmark for evaluating AI models on complex genomic tasks, featuring ten detailed case studies across somatic oncology, functional genomics, and population genetics.
Claude Science AI Workbench Release
Anthropic has launched Claude Science, an AI workbench that integrates scientific tools, manages scalable compute, and provides auditable research artifacts to accelerate scientific discovery.
DiScoFormer: One transformer for density and score, across distributions
DiScoFormer is a new transformer-based model that estimates both the density and score of a distribution from a set of data points in a single forward pass without requiring retraining for new distributions.
OpenAI Mapping Europe’s AI Workforce Opportunity Report 2026
OpenAI’s 2026 report extends its AI Jobs Transition Framework to the EU, finding that about 12% of EU employment may grow with AI, 14% faces higher near‑term automation potential, 27% is likely to reorganize, and 47% sees less immediate change.
vLLM Semantic Router: Enhancing Model Performance via Micro-Agent Collaboration
vLLM introduces the Semantic Router, a serving-layer primitive that transforms a single model API call into a bounded collaboration of micro-agents to outperform frontier models on complex benchmarks.
Ollama 0.31: Faster Gemma 4 Performance via Multi-Token Prediction
Ollama 0.31 introduces multi-token prediction for Gemma 4 on Apple Silicon, increasing token generation speed by nearly 90% on coding-agent benchmarks without altering model output.
HP Inc. and OpenAI Strategic Partnership
HP Inc. is scaling its strategic partnership with OpenAI using the OpenAI Frontier platform to deploy AI agents and workflows across customer support, security, and software development.
GPT-5.6 Sol preview: capabilities, safeguards, and availability
OpenAI previewed GPT-5.6 Sol, its flagship next‑generation model, alongside Terra and Luna, highlighting stronger coding, biology, and cybersecurity capabilities with enhanced safeguards and a limited release to trusted partners.
Hugging Face Jobs: Deploying vLLM Servers with a Single Command
Hugging Face now allows users to deploy private, OpenAI-compatible vLLM endpoints on its infrastructure using a single command via HF Jobs, providing a pay-per-second alternative for testing, evaluations, and batch generation.
Anthropic Economic Index Cadences Report – June 2026
Anthropic’s June 2026 Economic Index report reveals that Claude usage now follows daily and weekly work cadences, produces higher‑value artifacts that consume more compute, and that users who automate more tasks are surprisingly more optimistic about AI’s impact on their jobs.
OpenAI Codex Agentic AI Adoption and Impact
OpenAI reports a shift in knowledge work from short chatbot interactions to long-horizon agentic tasks via Codex, with non-developer adoption growing up to 189x since August 2025.
Gemini 3.5 Flash Computer Use Integration
Google DeepMind has integrated computer use as a native tool in Gemini 3.5 Flash, enabling agents to see, reason, and take action across browser, mobile, and desktop environments.