The archive · 11 labs · 3,050 dispatches

The labs

No more opening a dozen official blogs every morning. First-hand releases from OpenAI, Anthropic, DeepMind and the rest, each with its substance pulled out.

1051

Anthropic: Demystifying Evals for AI Agents

Anthropic provides a comprehensive framework for designing and implementing automated evaluations for AI agents to move from reactive debugging to a systematic, metric-driven development process.

1052

OpenAI for Healthcare Release

OpenAI has launched OpenAI for Healthcare, a suite of HIPAA-compliant tools including ChatGPT for Healthcare and the OpenAI API to reduce clinician administrative burden and improve care consistency.

1053

Netomi scaling agentic systems into the enterprise – lessons from OpenAI

Netomi announced a production‑grade agentic platform that combines GPT‑4.1 for low‑latency tool use with GPT‑5.2 for deep multi‑step planning, delivering sub‑3‑second responses, 98% intent accuracy, and built‑in governance at enterprise scale.

1054

Anthropic and PNNL AI Critical Infrastructure Defense Research

Anthropic and the Pacific Northwest National Laboratory (PNNL) demonstrated that Claude can accelerate adversary emulation for critical infrastructure defense, reducing attack reconstruction time from weeks to three hours.

1055

Qwen3-VL-Embedding and Qwen3-VL-Reranker Release

Qwen has released Qwen3-VL-Embedding and Qwen3-VL-Reranker, a series of multimodal models designed for high-precision cross-modal retrieval across text, images, screenshots, and video.

1056

How Tolan Builds Voice-First AI with GPT-5.1

Tolan utilizes GPT-5.1 and a custom real-time context management system to create a low-latency, voice-first AI companion with persistent memory and consistent personality.

1057

ChatGPT Health Release Notes

OpenAI has introduced ChatGPT Health, a dedicated, secure experience that integrates personal medical records and wellness app data to help users navigate health information without replacing clinical care.

1058

xAI Raises $20B Series E Funding Round

xAI announced a $20 billion Series E round, exceeding its $15 billion target, to accelerate its massive AI compute infrastructure and product rollout.

1059

NVIDIA Cosmos Reason 2 release notes / what's new

NVIDIA has released Cosmos Reason 2, an open reasoning vision-language model designed for physical AI that tops the Physical AI Bench and Physical Reasoning leaderboards.

1060

Falcon-H1-Arabic Release Notes

Hugging Face has announced Falcon-H1-Arabic, a family of three hybrid Mamba-Transformer models (3B, 7B, and 34B) that set new state-of-the-art benchmarks for Arabic NLP with expanded context windows up to 256K tokens.

1061

NVIDIA DGX Spark and Reachy Mini Integration Guide

NVIDIA has introduced a framework for creating real-world AI agents by combining DGX Spark hardware, Reachy Mini robotics, and the NeMo Agent Toolkit using open reasoning and vision models.

1062

OpenAI Grove Cohort 2 Announcement

OpenAI has opened applications for the second cohort of OpenAI Grove, a program designed to support pre-idea technical talent in building AI-driven companies.

1063

Qwen-Image-2512 release notes / what's new

Qwen-Image-2512 is a December update to the Qwen-Image text-to-image model that significantly improves human realism, natural detail rendering, and complex text layout accuracy.

1064

xAI Launches Grok Business and Grok Enterprise

xAI has introduced Grok Business and Grok Enterprise, providing organizations with secure, permission-aware AI assistants that do not train on customer data.

1065

Google DeepMind 2025 Year-in-Review: Gemini 3, Gemma, AI Products, Science, and Safety Breakthroughs

Google DeepMind announced 2025 breakthroughs across Gemini 3 models, open‑source Gemma, AI‑driven products, generative media, scientific AI, quantum computing, safety frameworks, and global collaborations, marking AI's shift from tool to utility.

1066

AprielGuard: A Guardrail for Safety and Adversarial Robustness in Modern LLM Systems

Hugging Face and ServiceNow AI have introduced AprielGuard, an 8B parameter safeguard model designed to detect 16 categories of safety risks and a wide range of adversarial attacks across standalone prompts, multi-turn conversations, and agentic workflows.

1067

Qwen-Image-Edit-2511 Release Notes

Qwen-Image-Edit-2511 is an updated image editing model that improves character consistency, integrates community LoRAs, and enhances geometric reasoning and industrial design capabilities.

1068

Qwen3-TTS-VD-Flash and Qwen3-TTS-VC-Flash release: controllable voice design and rapid multilingual voice cloning

Qwen released Qwen3‑TTS‑VD‑Flash for natural‑language voice design and Qwen3‑TTS‑VC‑Flash for 3‑second multilingual voice cloning, both outperforming leading TTS systems on controllability and accuracy.

1069

ChatGPT Atlas: Hardening Browser Agents Against Prompt Injection

OpenAI has implemented a proactive rapid response loop using an RL-trained automated attacker to discover and mitigate prompt injection vulnerabilities in ChatGPT Atlas before they are exploited in the wild.

1070

OpenAI Customer Growth and Enterprise Adoption

OpenAI has surpassed one million business customers, with 75% of users reporting the ability to complete tasks that were previously impossible.

1071

xAI Selected by US Department of War for Enterprise AI and Mission Systems

xAI has been selected by the U.S. Department of War (DOW) to integrate its Frontier AI systems and Grok models into the GenAI.Mil suite, providing AI capabilities to 3 million military and civilian employees at Impact Level 5.

1072

xAI Grok Collections API Release

xAI has launched the Collections API, a managed knowledge base service that enables developers to build RAG applications by uploading datasets for precise semantic, keyword, and hybrid search.

1073

Qwen-Image-Layered: Layered Decomposition for Inherent Editability

Qwen-Image-Layered is a new model that decomposes images into multiple RGBA layers, allowing for independent manipulation of image components without affecting other content.

1074

Anthropic Bloom: Open Source Tool for Automated Behavioral Evaluations

Anthropic has released Bloom, an open-source agentic framework that automatically generates evaluation suites to quantify the frequency and severity of specific behavioral traits in frontier AI models.

1075

Anthropic Frontier Compliance Framework for California SB 53

Anthropic has released its Frontier Compliance Framework (FCF) to meet the transparency and safety requirements of California's Transparency in Frontier AI Act (SB 53).

1076

OpenAI Evaluating Chain-of-Thought Monitorability

OpenAI introduces a framework and 13 evaluations to measure 'monitorability'—the ability to predict model behavior from its internal reasoning—finding that monitoring chains-of-thought is significantly more effective than monitoring outputs alone.

1077

OpenAI Model Spec Update: Under-18 (U18) Principles

OpenAI has updated its Model Spec to include Under-18 (U18) Principles, establishing age-appropriate behavioral guidelines for ChatGPT to prioritize teen safety and real-world support for users aged 13 to 17.

1078

OpenAI AI Literacy Resources for Teens and Parents

OpenAI has released a set of AI literacy resources, including a family-friendly guide and parent-specific tips, to help families navigate the safe and responsible use of ChatGPT.

1079

OpenAI and U.S. Department of Energy Expand AI Collaboration

OpenAI and the U.S. Department of Energy have signed a memorandum of understanding to integrate frontier AI models into scientific research, specifically supporting the Genesis Mission and national laboratory initiatives.

1080

Transformers v5 Tokenization Update

Hugging Face has redesigned tokenization in Transformers v5 to separate tokenizer architecture from trained vocabulary, enabling easier inspection, customization, and training from scratch.

1081

GPT-5.2-Codex Release Notes

OpenAI has released GPT-5.2-Codex, an agentic coding model optimized for complex software engineering, long-horizon tasks, and advanced cybersecurity research.

1082

OpenAI GPT-5.2-Codex System Card Addendum

OpenAI has released GPT-5.2-Codex, an agentic coding model optimized for complex software engineering, project-scale refactors, and cybersecurity tasks.

1083

Anthropic Project Vend Phase Two: AI Shopkeeper Gains Profitability but Still Needs Human Guardrails

Anthropic’s Phase Two of Project Vend upgraded the shop‑keeping AI Claude (renamed Claudius) to Claude Sonnet 4.x, added tools, a CEO agent, and a merch‑making colleague, which improved profitability across three locations but highlighted ongoing robustness gaps.

1084

Anthropic Announces New Safeguards for Claude to Protect User Wellbeing

Anthropic detailed new model and product safeguards for Claude, including suicide/self‑harm handling, reduced sycophancy, and an 18+ age restriction, showing high appropriateness rates on internal evaluations.

1085

Anthropic and US Department of Energy Partner for Genesis Mission

Anthropic has entered a multi-year partnership with the US Department of Energy as part of the Genesis Mission to integrate frontier AI into American energy dominance, biological sciences, and scientific productivity.

1086

Mistral OCR 3 Release Notes

Mistral AI has released Mistral OCR 3, a high-fidelity document parsing model that achieves a 74% win rate over its predecessor on complex documents and handwriting, priced at $2 per 1,000 pages.

1087

NVIDIA Nemotron 3 Nano open evaluation recipe with NeMo Evaluator

NVIDIA released the 30B Nemotron 3 Nano model with a fully open evaluation recipe built on NeMo Evaluator, enabling anyone to reproduce its benchmark scores and audit the entire evaluation pipeline.

1088

Gemini 3 Flash release notes / what's new

Google DeepMind has released Gemini 3 Flash, a model that combines frontier-level reasoning and multimodal intelligence with low latency and reduced costs for developers and consumers.

1089

OpenAI Academy for News Organizations Launch

OpenAI has launched the OpenAI Academy for News Organizations to provide journalists with training, playbooks, and guidance on the responsible adoption of AI to enhance reporting and business operations.

1090

OpenAI Introduces App Submission and Directory for ChatGPT

OpenAI has launched an app directory and a beta Apps SDK, allowing developers to build, submit, and monetize chat-native experiences directly within ChatGPT.

1091

The State of Enterprise AI 2025 Report

OpenAI's 2025 report reveals that enterprise AI adoption is scaling rapidly, with ChatGPT message volume growing 8x and API reasoning token consumption increasing 320x year-over-year.

1092

xAI Grok Voice Agent API Release

xAI has launched the Grok Voice Agent API, a high-speed, multilingual voice agent system that ranks #1 on Big Bench Audio and costs $0.05 per minute.

1093

Gemma Scope 2 release – open tools for interpreting Gemma 3 language models

Gemma Scope 2 is an open-source suite of interpretability tools for Gemma 3 models (270M‑27B parameters) that lets researchers trace internal computations and debug safety‑critical behaviours.

1094

OpenAI FrontierScience Benchmark Release

OpenAI has introduced FrontierScience, a new expert-level benchmark designed to measure AI's ability to perform complex scientific reasoning and real-world research tasks across physics, chemistry, and biology.

1095

OpenAI GPT-5 Biological Research Capabilities

OpenAI demonstrated that GPT-5 can autonomously optimize molecular cloning protocols, achieving a 79-fold increase in efficiency through the discovery of a novel enzymatic mechanism.

1096

OpenAI Guide: Staying Ahead in the Age of AI

OpenAI outlines a five-step playbook—Align, Activate, Amplify, Accelerate, and Govern—to help organizations transition into AI-first companies and maintain a competitive advantage.

1097

ChatGPT Images and GPT Image 1.5 Release

OpenAI has released a new flagship image generation model, GPT Image 1.5, which offers up to 4x faster generation speeds and significantly improved precise editing and text rendering capabilities.

1098

Anthropic "think" tool for Claude

Anthropic introduces the "think" tool, a dedicated reasoning space that improves Claude's agentic tool use, policy adherence, and sequential decision-making in complex tasks.

1099

CUGA on Hugging Face: Democratizing Configurable AI Agents

IBM Research has released CUGA (Configurable Generalist Agent), an open-source agent framework that achieves state-of-the-art performance on AppWorld and WebArena benchmarks for complex API and web tasks.

1100

Gemini 2.5 Flash Native Audio release: enhanced live voice agents and real‑time speech translation

Google DeepMind announced Gemini 2.5 Flash Native Audio, a live‑voice model that improves function calling, instruction adherence, and multi‑turn conversation while adding real‑time speech‑to‑speech translation for over 70 languages.