✷ The archive · 11 labs · 3,050 dispatches
The labs
No more opening a dozen official blogs every morning. First-hand releases from OpenAI, Anthropic, DeepMind and the rest, each with its substance pulled out.
Anthropic: Demystifying Evals for AI Agents
Anthropic provides a comprehensive framework for designing and implementing automated evaluations for AI agents to move from reactive debugging to a systematic, metric-driven development process.
OpenAI for Healthcare Release
OpenAI has launched OpenAI for Healthcare, a suite of HIPAA-compliant tools including ChatGPT for Healthcare and the OpenAI API to reduce clinician administrative burden and improve care consistency.
Netomi scaling agentic systems into the enterprise – lessons from OpenAI
Netomi announced a production‑grade agentic platform that combines GPT‑4.1 for low‑latency tool use with GPT‑5.2 for deep multi‑step planning, delivering sub‑3‑second responses, 98% intent accuracy, and built‑in governance at enterprise scale.
Anthropic and PNNL AI Critical Infrastructure Defense Research
Anthropic and the Pacific Northwest National Laboratory (PNNL) demonstrated that Claude can accelerate adversary emulation for critical infrastructure defense, reducing attack reconstruction time from weeks to three hours.
Qwen3-VL-Embedding and Qwen3-VL-Reranker Release
Qwen has released Qwen3-VL-Embedding and Qwen3-VL-Reranker, a series of multimodal models designed for high-precision cross-modal retrieval across text, images, screenshots, and video.
How Tolan Builds Voice-First AI with GPT-5.1
Tolan utilizes GPT-5.1 and a custom real-time context management system to create a low-latency, voice-first AI companion with persistent memory and consistent personality.
ChatGPT Health Release Notes
OpenAI has introduced ChatGPT Health, a dedicated, secure experience that integrates personal medical records and wellness app data to help users navigate health information without replacing clinical care.
xAI Raises $20B Series E Funding Round
xAI announced a $20 billion Series E round, exceeding its $15 billion target, to accelerate its massive AI compute infrastructure and product rollout.
NVIDIA Cosmos Reason 2 release notes / what's new
NVIDIA has released Cosmos Reason 2, an open reasoning vision-language model designed for physical AI that tops the Physical AI Bench and Physical Reasoning leaderboards.
Falcon-H1-Arabic Release Notes
Hugging Face has announced Falcon-H1-Arabic, a family of three hybrid Mamba-Transformer models (3B, 7B, and 34B) that set new state-of-the-art benchmarks for Arabic NLP with expanded context windows up to 256K tokens.
NVIDIA DGX Spark and Reachy Mini Integration Guide
NVIDIA has introduced a framework for creating real-world AI agents by combining DGX Spark hardware, Reachy Mini robotics, and the NeMo Agent Toolkit using open reasoning and vision models.
OpenAI Grove Cohort 2 Announcement
OpenAI has opened applications for the second cohort of OpenAI Grove, a program designed to support pre-idea technical talent in building AI-driven companies.
Qwen-Image-2512 release notes / what's new
Qwen-Image-2512 is a December update to the Qwen-Image text-to-image model that significantly improves human realism, natural detail rendering, and complex text layout accuracy.
xAI Launches Grok Business and Grok Enterprise
xAI has introduced Grok Business and Grok Enterprise, providing organizations with secure, permission-aware AI assistants that do not train on customer data.
Google DeepMind 2025 Year-in-Review: Gemini 3, Gemma, AI Products, Science, and Safety Breakthroughs
Google DeepMind announced 2025 breakthroughs across Gemini 3 models, open‑source Gemma, AI‑driven products, generative media, scientific AI, quantum computing, safety frameworks, and global collaborations, marking AI's shift from tool to utility.
AprielGuard: A Guardrail for Safety and Adversarial Robustness in Modern LLM Systems
Hugging Face and ServiceNow AI have introduced AprielGuard, an 8B parameter safeguard model designed to detect 16 categories of safety risks and a wide range of adversarial attacks across standalone prompts, multi-turn conversations, and agentic workflows.
Qwen-Image-Edit-2511 Release Notes
Qwen-Image-Edit-2511 is an updated image editing model that improves character consistency, integrates community LoRAs, and enhances geometric reasoning and industrial design capabilities.
Qwen3-TTS-VD-Flash and Qwen3-TTS-VC-Flash release: controllable voice design and rapid multilingual voice cloning
Qwen released Qwen3‑TTS‑VD‑Flash for natural‑language voice design and Qwen3‑TTS‑VC‑Flash for 3‑second multilingual voice cloning, both outperforming leading TTS systems on controllability and accuracy.
ChatGPT Atlas: Hardening Browser Agents Against Prompt Injection
OpenAI has implemented a proactive rapid response loop using an RL-trained automated attacker to discover and mitigate prompt injection vulnerabilities in ChatGPT Atlas before they are exploited in the wild.
OpenAI Customer Growth and Enterprise Adoption
OpenAI has surpassed one million business customers, with 75% of users reporting the ability to complete tasks that were previously impossible.
xAI Selected by US Department of War for Enterprise AI and Mission Systems
xAI has been selected by the U.S. Department of War (DOW) to integrate its Frontier AI systems and Grok models into the GenAI.Mil suite, providing AI capabilities to 3 million military and civilian employees at Impact Level 5.
xAI Grok Collections API Release
xAI has launched the Collections API, a managed knowledge base service that enables developers to build RAG applications by uploading datasets for precise semantic, keyword, and hybrid search.
Qwen-Image-Layered: Layered Decomposition for Inherent Editability
Qwen-Image-Layered is a new model that decomposes images into multiple RGBA layers, allowing for independent manipulation of image components without affecting other content.
Anthropic Bloom: Open Source Tool for Automated Behavioral Evaluations
Anthropic has released Bloom, an open-source agentic framework that automatically generates evaluation suites to quantify the frequency and severity of specific behavioral traits in frontier AI models.
Anthropic Frontier Compliance Framework for California SB 53
Anthropic has released its Frontier Compliance Framework (FCF) to meet the transparency and safety requirements of California's Transparency in Frontier AI Act (SB 53).
OpenAI Evaluating Chain-of-Thought Monitorability
OpenAI introduces a framework and 13 evaluations to measure 'monitorability'—the ability to predict model behavior from its internal reasoning—finding that monitoring chains-of-thought is significantly more effective than monitoring outputs alone.
OpenAI Model Spec Update: Under-18 (U18) Principles
OpenAI has updated its Model Spec to include Under-18 (U18) Principles, establishing age-appropriate behavioral guidelines for ChatGPT to prioritize teen safety and real-world support for users aged 13 to 17.
OpenAI AI Literacy Resources for Teens and Parents
OpenAI has released a set of AI literacy resources, including a family-friendly guide and parent-specific tips, to help families navigate the safe and responsible use of ChatGPT.
OpenAI and U.S. Department of Energy Expand AI Collaboration
OpenAI and the U.S. Department of Energy have signed a memorandum of understanding to integrate frontier AI models into scientific research, specifically supporting the Genesis Mission and national laboratory initiatives.
Transformers v5 Tokenization Update
Hugging Face has redesigned tokenization in Transformers v5 to separate tokenizer architecture from trained vocabulary, enabling easier inspection, customization, and training from scratch.
GPT-5.2-Codex Release Notes
OpenAI has released GPT-5.2-Codex, an agentic coding model optimized for complex software engineering, long-horizon tasks, and advanced cybersecurity research.
OpenAI GPT-5.2-Codex System Card Addendum
OpenAI has released GPT-5.2-Codex, an agentic coding model optimized for complex software engineering, project-scale refactors, and cybersecurity tasks.
Anthropic Project Vend Phase Two: AI Shopkeeper Gains Profitability but Still Needs Human Guardrails
Anthropic’s Phase Two of Project Vend upgraded the shop‑keeping AI Claude (renamed Claudius) to Claude Sonnet 4.x, added tools, a CEO agent, and a merch‑making colleague, which improved profitability across three locations but highlighted ongoing robustness gaps.
Anthropic Announces New Safeguards for Claude to Protect User Wellbeing
Anthropic detailed new model and product safeguards for Claude, including suicide/self‑harm handling, reduced sycophancy, and an 18+ age restriction, showing high appropriateness rates on internal evaluations.
Anthropic and US Department of Energy Partner for Genesis Mission
Anthropic has entered a multi-year partnership with the US Department of Energy as part of the Genesis Mission to integrate frontier AI into American energy dominance, biological sciences, and scientific productivity.
Mistral OCR 3 Release Notes
Mistral AI has released Mistral OCR 3, a high-fidelity document parsing model that achieves a 74% win rate over its predecessor on complex documents and handwriting, priced at $2 per 1,000 pages.
NVIDIA Nemotron 3 Nano open evaluation recipe with NeMo Evaluator
NVIDIA released the 30B Nemotron 3 Nano model with a fully open evaluation recipe built on NeMo Evaluator, enabling anyone to reproduce its benchmark scores and audit the entire evaluation pipeline.
Gemini 3 Flash release notes / what's new
Google DeepMind has released Gemini 3 Flash, a model that combines frontier-level reasoning and multimodal intelligence with low latency and reduced costs for developers and consumers.
OpenAI Academy for News Organizations Launch
OpenAI has launched the OpenAI Academy for News Organizations to provide journalists with training, playbooks, and guidance on the responsible adoption of AI to enhance reporting and business operations.
OpenAI Introduces App Submission and Directory for ChatGPT
OpenAI has launched an app directory and a beta Apps SDK, allowing developers to build, submit, and monetize chat-native experiences directly within ChatGPT.
The State of Enterprise AI 2025 Report
OpenAI's 2025 report reveals that enterprise AI adoption is scaling rapidly, with ChatGPT message volume growing 8x and API reasoning token consumption increasing 320x year-over-year.
xAI Grok Voice Agent API Release
xAI has launched the Grok Voice Agent API, a high-speed, multilingual voice agent system that ranks #1 on Big Bench Audio and costs $0.05 per minute.
Gemma Scope 2 release – open tools for interpreting Gemma 3 language models
Gemma Scope 2 is an open-source suite of interpretability tools for Gemma 3 models (270M‑27B parameters) that lets researchers trace internal computations and debug safety‑critical behaviours.
OpenAI FrontierScience Benchmark Release
OpenAI has introduced FrontierScience, a new expert-level benchmark designed to measure AI's ability to perform complex scientific reasoning and real-world research tasks across physics, chemistry, and biology.
OpenAI GPT-5 Biological Research Capabilities
OpenAI demonstrated that GPT-5 can autonomously optimize molecular cloning protocols, achieving a 79-fold increase in efficiency through the discovery of a novel enzymatic mechanism.
OpenAI Guide: Staying Ahead in the Age of AI
OpenAI outlines a five-step playbook—Align, Activate, Amplify, Accelerate, and Govern—to help organizations transition into AI-first companies and maintain a competitive advantage.
ChatGPT Images and GPT Image 1.5 Release
OpenAI has released a new flagship image generation model, GPT Image 1.5, which offers up to 4x faster generation speeds and significantly improved precise editing and text rendering capabilities.
Anthropic "think" tool for Claude
Anthropic introduces the "think" tool, a dedicated reasoning space that improves Claude's agentic tool use, policy adherence, and sequential decision-making in complex tasks.
CUGA on Hugging Face: Democratizing Configurable AI Agents
IBM Research has released CUGA (Configurable Generalist Agent), an open-source agent framework that achieves state-of-the-art performance on AppWorld and WebArena benchmarks for complex API and web tasks.
Gemini 2.5 Flash Native Audio release: enhanced live voice agents and real‑time speech translation
Google DeepMind announced Gemini 2.5 Flash Native Audio, a live‑voice model that improves function calling, instruction adherence, and multi‑turn conversation while adding real‑time speech‑to‑speech translation for over 70 languages.