Gemma 4 Release Notes: High-Intelligence Open Models for Reasoning and Agents
Google DeepMind has released Gemma 4, a family of open models under an Apache 2.0 license designed for advanced reasoning, agentic workflows, and efficient on-device deployment.
OpenAI acquires TBPN – strategic media integration and editorial independence
OpenAI announced the acquisition of TBPN, a media platform known for AI‑focused conversations, to embed editorial independence within OpenAI’s Strategy organization and leverage TBPN’s communications expertise.
OpenAI Codex Flexible Pricing for Teams
OpenAI has introduced pay-as-you-go pricing for Codex-only seats in ChatGPT Business and Enterprise workspaces to lower the barrier for team adoption and pilot programs.
Gemma 4 on vLLM: Advanced Reasoning and Multimodal Capabilities
vLLM introduces Day 0 support for Gemma 4, a family of open models from Google featuring advanced reasoning, multimodal inputs, and broad hardware compatibility across TPUs, GPUs, and XPUs.
Gemma 4 release: open‑source multimodal models with on‑device support
Gemma 4, a new open‑source multimodal family from Google DeepMind, is released on Hugging Face with image, audio, video support, up to 256 K context, and day‑0 compatibility with transformers, llama.cpp, MLX, Rust, and WebGPU.
Sandakan Central Market sign vertical text in 2024 – “2006"
The vertical text on the right side of Sandakan Central Market’s main sign in 2024 reads “2006”.
Falcon Perception and Falcon OCR Release
Hugging Face and TII introduce Falcon Perception, a 0.6B-parameter early-fusion Transformer for open-vocabulary grounding, and Falcon OCR, a 0.3B-parameter high-throughput document understanding model.
Gradient Labs AI Account Managers for Banking
Gradient Labs is deploying AI agents powered by OpenAI's GPT-5.4 mini and nano models to provide bank customers with dedicated AI account managers capable of handling complex, compliant financial procedures.
Gradio Server: Integrating Custom Frontends with Gradio Backend
Hugging Face introduces gradio.Server, a FastAPI extension that allows developers to use any custom frontend framework while retaining Gradio's queuing, API infrastructure, and ZeroGPU support.
Granite 4.0 3B Vision release notes / what's new
IBM has released Granite 4.0 3B Vision, a compact multimodal model optimized for enterprise document understanding, featuring high-accuracy table extraction, chart reasoning, and key-value pair extraction.
OpenAI Funding and Infrastructure Strategy Update
OpenAI has raised $122 billion in committed capital at an $852 billion post-money valuation to expand its compute infrastructure and develop a unified AI superapp.
OpenMed CodonRoBERTa multi-species mRNA language models release
OpenMed released an end-to-end protein engineering pipeline with CodonRoBERTa-large-v2 (perplexity 4.10, CAI 0.404) and a 25-species codon‑optimization model suite trained in 55 GPU‑hours for $165.
TRL v1.0 release notes / what's new
Hugging Face releases TRL v1.0, a stable post-training library implementing over 75 methods, featuring a dual-track stability model to balance rapid experimental iteration with production-grade reliability.
vLLM Hidden States Extraction System
vLLM v0.18.0 introduces a native hidden states extraction system that enables high-performance retrieval of internal model representations for training speculative decoding draft models.
OpenAI AI Jam for Disaster Management in Asia
OpenAI, in partnership with the Gates Foundation, APDC, and DataKind, hosted an AI Jam in Bangkok to help disaster management professionals from 13 Asian countries develop practical AI workflows for emergency response.
Qwen3.5-Omni release notes
Qwen announced Qwen3.5-Omni, a new omnimodal LLM that handles text, images, audio, and video, supports 256k context, 113-language speech recognition, 36-language synthesis, and adds real-time features like semantic interruption, websearch, voice control, and voice cloning.
Google DeepMind AI-Enabled Pointer
Google DeepMind is developing an AI-enabled pointer powered by Gemini that replaces complex text prompts with intuitive pointing and voice commands to interact with digital content across applications.
STADLER AI Implementation and Productivity Gains
STADLER, a 230-year-old industrial recycling company, achieved 30-40% time savings on knowledge tasks and 2.5x faster drafting by embedding OpenAI's ChatGPT as a company-wide productivity layer.
Hugging Face OpenClaw Migration Guide
Hugging Face provides two methods—Inference Providers and local llama.cpp setup—to migrate OpenClaw agents from restricted Claude models to open-source alternatives.
Gemini 3.1 Flash Live release notes / what's new
Google DeepMind has released Gemini 3.1 Flash Live, a high-quality audio and voice model designed for natural, real-time dialogue with improved precision, lower latency, and expanded global availability.
Google DeepMind Research on AI Harmful Manipulation
Google DeepMind has released a new research study and an empirically validated toolkit to measure and mitigate the risk of AI being used for harmful manipulation of human thought and behavior.
Lyria 3 Pro release: longer, structurally aware music generation across Google products
DeepMind announced Lyria 3 Pro, a music‑generation model that creates up to three‑minute tracks with structural awareness and is now integrated across Google products like Vertex AI, AI Studio, Google Vids, Gemini, and ProducerAI.
OpenAI Model Spec: A Framework for Explicit Model Behavior
OpenAI has introduced the Model Spec, a formal, public framework designed to make intended AI model behavior explicit, legible, and revisable for users, developers, and researchers.
OpenAI Safety Bug Bounty Program Launch
OpenAI has launched a public Safety Bug Bounty program to identify AI abuse and safety risks that fall outside conventional security vulnerabilities, specifically targeting agentic risks, proprietary information leaks, and platform integrity.
OpenAI Teen Safety Policy Pack and gpt-oss-safeguard
OpenAI has released prompt-based safety policies and the open-weight gpt-oss-safeguard model to help developers implement age-appropriate protections for teenagers in AI applications.
OpenAI expands product discovery in ChatGPT via Agentic Commerce Protocol
OpenAI has introduced enhanced visual shopping and product discovery capabilities in ChatGPT, powered by the expanded Agentic Commerce Protocol (ACP) to streamline how users find and compare products.
OpenAI Foundation Update
OpenAI has announced that its Foundation will invest at least $1 billion over the next year across life sciences, economic impact, AI resilience, and community programs to ensure AGI benefits humanity.
EVA End-to-End Evaluation Framework for Voice Agents
EVA is a new end‑to‑end framework that jointly evaluates voice agents on accuracy and conversational experience, revealing a consistent trade‑off between task success and user satisfaction.
vLLM Model Runner V2 release notes / what's new
vLLM has introduced Model Runner V2 (MRV2), a ground-up re-implementation of the model runner that improves throughput and reduces latency through a GPU-native, async-first, and modular architecture.
OpenAI Sora 2 Safety Framework
OpenAI has detailed the safety architecture for Sora 2, focusing on provenance signals, consent-based likeness management, and strict content filtering for audio and video.
Domain-Specific Embedding Fine-Tuning with NVIDIA Nemotron – Under a Day
NVIDIA and Hugging Face released a single‑GPU, under‑a‑day pipeline that fine‑tunes the Llama‑Nemotron‑Embed‑1B‑v2 model on synthetic domain data, delivering >10% retrieval gains and up to 26% improvement on real enterprise datasets.
OpenAI Internal Coding Agent Monitoring System
OpenAI has deployed a GPT-5.4 Thinking-powered monitoring system to detect misalignment and security violations in internal coding agents, identifying behaviors that often only emerge in complex, tool-rich workflows.
OpenAI to acquire Astral
OpenAI is acquiring Astral to integrate its open-source Python tools, including uv, Ruff, and ty, into the Codex ecosystem to enable AI agents to participate in the entire software development lifecycle.
Qwen3.5-Max-Preview Release on LMSys Arena
Qwen has deployed Qwen3.5-Max-Preview to the LMSys Arena for community evaluation ahead of its full release scheduled within two weeks.
State of Open Source on Hugging Face: Spring 2026
Hugging Face reports a massive expansion of the open source AI ecosystem in 2025, characterized by China surpassing the U.S. in model downloads and the rapid emergence of robotics as the largest dataset category.
Google DeepMind Measuring Progress Toward AGI: A Cognitive Framework
Google DeepMind has introduced a cognitive taxonomy and a three-stage evaluation protocol to empirically measure AI progress toward Artificial General Intelligence (AGI).
Holotron-12B High Throughput Computer Use Agent
H Company released Holotron-12B, a multimodal computer-use model based on NVIDIA Nemotron-Nano-2 VL that uses a hybrid SSM-Attention architecture to achieve high inference throughput for agentic workloads.
OpenAI GPT-5.4 mini and nano release notes
OpenAI has released GPT-5.4 mini and nano, high-efficiency small models that bring GPT-5.4 capabilities to high-volume workloads with significantly lower latency and cost.
OpenAI Japan Teen Safety Blueprint
OpenAI Japan has introduced the Japan Teen Safety Blueprint, a framework prioritizing teen safety over convenience and privacy to protect younger users from AI-related risks.
OpenAI Research: How Workers Use ChatGPT for Compensation Insights
OpenAI research reveals that US workers send nearly 3 million daily messages to ChatGPT seeking wage benchmarks and compensation guidance, particularly in high-skill, low-transparency roles.
OpenAI Codex Security: Why the System Avoids SAST Report Seeding
OpenAI's Codex Security avoids starting with Static Application Security Testing (SAST) reports to prevent premature narrowing of analysis and to focus on validating whether security invariants actually hold through transformation chains.
BAIR Introducing SPEX and ProxySPEX for Scalable LLM Interaction Discovery
BAIR has introduced SPEX and ProxySPEX, algorithms that use signal processing and coding theory to identify influential interactions between features, training data, and model components at scale.
P-EAGLE: Parallel Speculative Decoding in vLLM
vLLM introduces P-EAGLE, a parallel speculative decoding method that generates all draft tokens in a single forward pass, delivering up to 1.69x speedup over vanilla EAGLE-3 on NVIDIA B200 GPUs.
OpenAI Designing AI Agents to Resist Prompt Injection
OpenAI is shifting its defense strategy against prompt injection by treating it as a social engineering problem, focusing on constraining the impact of successful manipulations rather than relying solely on input filtering.
OpenAI Responses API Computer Environment Update
OpenAI has equipped the Responses API with a shell tool and hosted container workspace, enabling models to execute real-world tasks via a command-line interface and persistent runtime context.
NVIDIA Nemotron 3 Super Support in vLLM
vLLM now supports NVIDIA Nemotron 3 Super, a 120B parameter hybrid MoE model optimized for multi-agent AI with a 1 million token context window and high inference efficiency.
Rakuten Integration of OpenAI Codex for Engineering Efficiency
Rakuten has integrated OpenAI Codex into its engineering stack, achieving a 50% reduction in mean time to recovery (MTTR) and compressing quarter-long development projects into weeks.
Wayfair OpenAI Integration Case Study
Wayfair has integrated OpenAI models into its internal systems to automate product catalog tagging for 30 million items and streamline supplier support via the AI-powered tool Wilma.
OpenAI Instruction Hierarchy Improvements and GPT-5 Mini-R
OpenAI has introduced a new reinforcement learning dataset, IH-Challenge, to train models to prioritize trusted instructions over untrusted ones, resulting in the GPT-5 Mini-R model with improved safety steerability and prompt injection robustness.
ChatGPT Interactive Visuals for Math and Science
OpenAI has introduced dynamic visual explanations for over 70 core math and science concepts in ChatGPT to help users understand the relationships between variables and formulas in real time.