501

IBM and UC Berkeley Diagnose Enterprise Agent Failures Using IT-Bench and MAST

IBM Research and UC Berkeley introduced MAST (Multi-Agent System Failure Taxonomy) to diagnose why enterprise IT agents fail, revealing that frontier models suffer from isolated verification errors while open models face cascading systemic collapses.

502

Gemini Music Generation with Lyria 3

Google DeepMind has integrated the Lyria 3 generative music model into the Gemini app, enabling users to create 30-second AI-generated tracks from text prompts or image and video uploads.

503

Gradio 6 gr.HTML: One-Shot Web App Development

Gradio 6 introduces enhanced gr.HTML support for custom templates, scoped CSS, and JavaScript interactivity, enabling the creation of complex web components within a single Python file.

504

OpenAI Introducing EVMbench

OpenAI and Paradigm have released EVMbench, a benchmark designed to evaluate AI agents' ability to detect, patch, and exploit high-severity smart contract vulnerabilities.

505

Google DeepMind National Partnerships for AI: India Expansion

Google DeepMind is establishing new partnerships with Indian government bodies and institutions to broaden access to frontier AI models for science, education, agriculture, and energy security.

506

Qwen 3.5‑397B‑A17B release: hybrid linear‑attention MoE model with 1 M token context and state‑of‑the‑art multimodal performance

Qwen 3.5‑397B‑A17B is a 397 billion‑parameter multimodal model that activates only 17 billion parameters per token, delivering state‑of‑the‑art performance on language, coding, reasoning and vision tasks while being up to 19× faster than its predecessor.

507

GPT-5.2 Derives New Result in Theoretical Physics

OpenAI has announced that GPT-5.2 Pro and a scaffolded version of GPT-5.2 were used to derive a new result in theoretical physics regarding non-zero gluon tree amplitudes in the half-collinear regime.

508

ChatGPT Lockdown Mode and Elevated Risk Labels

OpenAI has introduced Lockdown Mode and Elevated Risk labels to mitigate prompt injection attacks and provide users with greater control over data exfiltration risks.

509

OpenAI Scaling Access to Codex and Sora via Hybrid Credit System

OpenAI has implemented a real-time access engine that combines rate limits with a purchasable credit system to prevent hard stops for Codex and Sora users.

510

OpenAI GABRIEL: Scaling Social Science Research with GPT

OpenAI has released GABRIEL, an open-source Python toolkit that uses GPT to transform unstructured text and images into quantitative measurements for social science research.

511

Hugging Face CUDA Kernels Agent Skill

Hugging Face has introduced an agent skill that enables coding agents like Claude and Codex to write, benchmark, and integrate production-ready CUDA kernels for transformers and diffusers libraries.

512

Gemini 3 Deep Think Update

Google DeepMind has updated Gemini 3 Deep Think, a specialized reasoning mode that achieves gold-medal level performance in mathematics, physics, and chemistry olympiads and is now available via the Gemini API for select users.

513

GPT-5.3-Codex-Spark Release Notes

OpenAI has released GPT-5.3-Codex-Spark, a small, low-latency model optimized for real-time coding collaboration and delivering over 1,000 tokens per second on Cerebras hardware.

514

OpenEnv: Evaluating Tool-Using Agents in Real-World Environments

Hugging Face and Meta introduce OpenEnv, an open-source framework that evaluates AI agents against real systems and production-grade environments like the Calendar Gym to bridge the gap between research and production reliability.

515

OpenAI Harness Engineering: Leveraging Codex for Zero-Manual-Code Development

OpenAI developed an internal software product with zero lines of manually-written code using Codex, reducing development time by approximately 10x through a shift from manual coding to environment design and agent orchestration.

516

Qwen-Image-2.0 Release: Professional Infographics and Photorealism

Qwen-Image-2.0 is a unified image generation and editing model that supports 1k-token instructions for professional infographics and native 2K resolution for high-fidelity photorealism.

517

Gemini Deep Think enables autonomous research across mathematics, physics, and computer science

Gemini Deep Think powers autonomous and collaborative agents that solve research‑level math, physics, and computer‑science problems, producing publishable results and resolving open conjectures.

518

OpenAI Integrates ChatGPT into GenAI.mil for Department of War

OpenAI is deploying a custom version of ChatGPT to GenAI.mil, providing 3 million Department of War personnel with secure, unclassified generative AI capabilities for administrative and operational support.

519

Transformers.js v4 release notes / what's new

Hugging Face has released Transformers.js v4, introducing a new C++ rewritten WebGPU runtime for hardware acceleration across browsers and server-side runtimes, alongside a standalone tokenizers library.

520

OpenAI Localization Approach and OpenAI for Countries Initiative

OpenAI has introduced a framework for localized AI systems through the OpenAI for Countries initiative, allowing nations to adapt frontier models to local languages, laws, and cultural norms while adhering to global safety red-lines.

521

SyGra 2.0.0 Studio Release

SyGra 2.0.0 introduces Studio, a visual interactive environment for designing and executing synthetic data generation workflows without needing to manually edit YAML files.

522

GPT-5 Lowers the Cost of Cell-Free Protein Synthesis

OpenAI and Ginkgo Bioworks used GPT-5 in a closed-loop autonomous lab system to reduce cell-free protein synthesis costs by 40% and reagent costs by 57%.

523

OpenAI Trusted Access for Cyber

OpenAI has launched Trusted Access for Cyber, an identity-based framework to provide security professionals with prioritized access to GPT-5.3-Codex to accelerate cyber defense and vulnerability remediation.

524

OpenAI Frontier Platform Release

OpenAI has introduced Frontier, an end-to-end platform designed to help enterprises build, deploy, and manage AI agents with shared business context, secure execution environments, and integrated governance.

525

GPT-5.3-Codex System Card

OpenAI has released GPT-5.3-Codex, an agentic coding model that integrates GPT-5.2-Codex performance with GPT-5.2 reasoning to handle complex, long-running technical tasks.

526

GPT-5.3-Codex release notes / what's new

OpenAI has released GPT-5.3-Codex, an agentic coding model that improves upon GPT-5.2-Codex in performance, reasoning, and speed, enabling it to execute complex, long-running technical tasks autonomously.

527

OpenAI Codex App Server Architecture and Integration

OpenAI has introduced the Codex App Server, a JSON-RPC based protocol and process that exposes the Codex agent harness to various clients, enabling a consistent agent experience across IDEs, web runtimes, and terminal interfaces.

528

Hugging Face Community Evals

Hugging Face has introduced Community Evals, a decentralized system for reporting and aggregating model benchmark scores directly on the Hub to increase transparency and reproducibility.

529

H Company Holo2-235B-A22B Preview Release

H Company has released Holo2-235B-A22B Preview, a UI localization model that achieves state-of-the-art performance on Screenspot-Pro and OSWorld G benchmarks.

530

The Future of the Global Open-Source AI Ecosystem: From DeepSeek to AI+

Hugging Face analyzes how open source has become the dominant strategy for Chinese AI organizations, shifting from isolated model breakthroughs to a scalable, integrated ecosystem of models, hardware, and infrastructure.

531

Training Design for Text-to-Image Models: Lessons from Ablations

The provided source material for the Hugging Face post on text-to-image model training design is unavailable due to a 429 Too Many Requests error.

532

OpenAI Sora Feed Philosophy

OpenAI has detailed the design principles and safety frameworks for the Sora feed, focusing on creativity-driven ranking, personalized recommendations, and multi-layered safety guardrails.

533

Qwen3-Coder-Next Release: High-Efficiency Agentic Coding Model

Qwen3-Coder-Next is an open-weight model based on a hybrid attention and MoE architecture that achieves over 70% on SWE-Bench Verified, offering performance comparable to models 10-20x larger.

534

Snowflake and OpenAI Partnership for Enterprise Intelligence

OpenAI and Snowflake have entered a $200 million agreement to integrate OpenAI frontier models, including GPT-5.2, directly into the Snowflake AI Data Cloud to enable the creation of secure, data-grounded AI agents and applications.

535

OpenAI Codex app release – multi‑agent desktop interface for macOS and Windows

OpenAI launched the Codex desktop app for macOS (with Windows support added in March 2026), a unified interface that lets developers run, supervise, and collaborate with multiple AI agents in parallel, extend them with reusable skills, and automate repetitive tasks.

536

Project Genie: Google DeepMind's Interactive World Model Prototype

Google DeepMind has launched Project Genie, an experimental research prototype powered by Genie 3 that allows users to create, explore, and remix interactive, real-time generated environments.

537

Inside OpenAI's In-House Data Agent

OpenAI has developed a bespoke internal AI data agent that enables employees to perform complex data analysis across 600 petabytes of data using natural language, utilizing a multi-layered context system and a self-learning reasoning loop.

538

Introducing Daggr: Chain AI Apps Programmatically with Visual Inspection

Hugging Face has released Daggr, an open-source Python library that allows developers to programmatically chain Gradio apps, ML models, and custom functions into workflows with an automatically generated visual canvas for debugging and state management.

539

OpenAI Retires GPT-4o, GPT-4.1, and o4-mini in ChatGPT

OpenAI is retiring GPT-4o, GPT-4.1, GPT-4.1 mini, and OpenAI o4-mini from ChatGPT on February 13, 2026, as usage has shifted to GPT-5.2.

540

Taisei Corporation ChatGPT Enterprise Implementation

Taisei Corporation has deployed ChatGPT Enterprise as a core component of its talent development strategy to expand human potential and reshape workforce capabilities in the construction industry.

541

Qwen3-ASR and Qwen3-ForcedAligner Release

Qwen has open-sourced Qwen3-ASR (1.7B and 0.6B) and Qwen3-ForcedAligner-0.6B, providing state-of-the-art multilingual speech recognition and non-autoregressive timestamp prediction under the Apache 2.0 license.

542

OpenAI EU Economic Blueprint 2.0

OpenAI has launched the EU Economic Blueprint 2.0, introducing a program to train 20,000 SMEs, a youth safety grant, and expanded government partnerships to close Europe's AI capability overhang.

543

OpenAI EMEA Youth & Wellbeing Grant 2026

OpenAI launched a €500,000 EMEA Youth & Wellbeing Grant to fund NGOs and researchers working on AI safety, wellbeing, and development for young people across Europe, the Middle East, and Africa.

544

Hugging Face Upskill: Transferring Expert Capabilities to Smaller Models via Agent Skills

Hugging Face introduced upskill, a tool that uses high-capability models like Claude Opus 4.5 to generate validated 'agent skills' that improve the performance and token efficiency of smaller or open-source models on complex tasks such as CUDA kernel development.

545

OpenAI AI Agent Link Safety

OpenAI has implemented a system to prevent URL-based data exfiltration by only allowing AI agents to automatically fetch URLs that have been previously verified as public via an independent web index.

546

Architectural Choices in China's Open‑Source AI Ecosystem: From DeepSeek R1 to a Hardware‑First, MoE‑Driven Landscape

One year after DeepSeek R1’s open‑source release, China’s AI community shifted from chasing the biggest single‑model performance to building flexible, cost‑effective, and hardware‑aware AI systems. Mixture‑of‑Experts (MoE) became the default architecture, enabling huge models to run affordably by activating only a subset of experts per request. Multimodal races exploded, with open releases for text‑to‑image, video, audio, 3‑D, and agents, each bundled with full toolchains. Small models (≤30 B) surged in popularity for local deployment and fine‑tuning, while large MoE models serve as teacher nets for distillation. Apache 2.0 and MIT licenses now dominate, removing legal friction and accelerating commercial adoption. A hardware‑first mindset emerged: releases ship with quantization, inference, and serving stacks tuned for domestic chips (Huawei Ascend, Cambricon, Kunlun), and training pipelines are openly documented. The competitive edge now lies in system design, deployment efficiency, and open‑source ecosystem integration rather than raw model size.

547

Alyah: Emirati Dialect Benchmark for Arabic LLMs

Hugging Face and partners introduced Alyah, a manually curated benchmark of 1,173 samples designed to evaluate the linguistic and cultural capabilities of Arabic LLMs in the Emirati dialect.

548

PVH and OpenAI Fashion Partnership

The provided source material for the PVH and OpenAI partnership is unavailable due to a server error, and no technical details could be retrieved.

549

GPT-OSS Agentic RL Training: A Practical Retrospective

Hugging Face and LinkedIn researchers detailed the engineering fixes required to enable stable agentic reinforcement learning for the GPT-OSS model, focusing on MoE routing, attention sinks, and memory efficiency.

550

OpenAI Prism Release

OpenAI has launched Prism, a free, AI-native LaTeX workspace powered by GPT-5.2 designed to integrate scientific writing, collaboration, and reasoning into a single environment.