The archive · 11 labs · 871 dispatches

The labs

No more opening a dozen official blogs every morning. First-hand releases from OpenAI, Anthropic, DeepMind and the rest, each with its substance pulled out.

201

SOTA OCR with Core ML and dots.ocr

Hugging Face demonstrates the conversion of the 3B parameter dots.ocr model to run on-device using Core ML and MLX, achieving performance that surpasses Gemini 2.5 Pro on OmniDocBench.

202

Hugging Face Introduces RTEB for Retrieval Embedding Evaluation

Hugging Face has launched the beta version of the Retrieval Embedding Benchmark (RTEB), a new standard designed to measure the true generalization and retrieval accuracy of embedding models using a hybrid of open and private datasets.

203

VibeGame: A High-Level Declarative Engine for AI-Assisted Game Development

Hugging Face researcher Dylan Ebert introduces VibeGame, a declarative game engine designed to enable 'vibe coding' by combining high-level abstractions with a web-based stack optimized for AI proficiency.

204

Qwen3-8B acceleration on Intel Core Ultra with depth‑pruned draft models

Accelerated Qwen3-8B on Intel Core Ultra using speculative decoding and a depth‑pruned Qwen3‑0.6B draft achieves ~1.4× speedup, enabling fast local AI agents via Hugging Face smolagents.

205

Nemotron-Personas-Japan Release

NVIDIA has released Nemotron-Personas-Japan, an open synthetic dataset of 6 million Japanese personas designed to support the development of Sovereign AI and culturally aware LLMs.

206

Swift Transformers 1.0 release notes / what's new

Hugging Face has released Swift Transformers 1.0, a stable library providing tokenizers, hub access, and model wrappers to simplify local LLM integration for Apple Silicon developers.

207

Smol2Operator release: turning a small VLM into an open‑source agentic GUI coder

Hugging Face released Smol2Operator, a post‑training recipe that converts the 2.2 B‑parameter SmolVLM2‑2.2B‑Instruct model into an open‑source agentic GUI coder, with all code, datasets, and the resulting model publicly available.

208

SyGra: A Low-Code Framework for LLM and SLM Data Generation

SyGra is a low-code/no-code framework developed by ServiceNow AI to simplify the creation, transformation, and alignment of datasets for Large and Small Language Models.

209

Hugging Face Gaia2 and Meta Agents Research Environments (ARE) Release

Hugging Face has released Gaia2, a read-and-write agentic benchmark, and the Meta Agents Research Environments (ARE) framework to evaluate and debug AI agents in complex, real-world simulated conditions.

210

Scaleway added as Hugging Face Inference Provider – capabilities, usage, and billing

Scaleway is now an official Inference Provider on the Hugging Face Hub, offering low‑latency, European‑hosted serverless access to frontier models with flexible billing options.

211

Hugging Face RiskRubric.ai Announcement

Hugging Face has announced RiskRubric.ai, a standardized risk assessment platform that evaluates AI models across six pillars of transparency, reliability, security, privacy, safety, and reputation to provide comparable risk scores and letter grades.

212

Hugging Face Integrates Public AI as an Inference Provider

Hugging Face has added Public AI as a supported Inference Provider, enabling serverless access to sovereign models from institutions like the Swiss AI Initiative and AI Singapore.

213

LeRobotDataset v3.0 release notes / what's new

Hugging Face has released LeRobotDataset v3.0, a standardized robotics dataset format that enables large-scale data storage and streaming to support millions of episodes.

214

Visible Watermarking with Gradio

Hugging Face has introduced a simple way to add visible watermarks to AI-generated images, videos, and text using the Gradio library to improve synthetic content transparency.

215

Writer Palmyra-mini Family Release

Writer has released the Palmyra-mini family, a set of lightweight open models (1.5B to 1.7B parameters) featuring a base model and two specialized reasoning variants optimized for logic and mathematics.

216

Transformers 4.40 performance upgrades for OpenAI GPT‑OSS: zero‑build kernels, MXFP4 quantization, parallelism, and faster loading

Hugging Face added zero-build kernels, MXFP4 4-bit quantization, tensor and expert parallelism, dynamic sliding-window cache, continuous batching, and faster model loading to the transformers library for OpenAI's GPT‑OSS models, dramatically improving efficiency and scalability.

217

Fine-tuning Hugging Face Hub LLMs with Together AI

Together AI and Hugging Face have integrated their platforms, allowing developers to fine-tune any compatible LLM from the Hugging Face Hub using Together AI's infrastructure.

218

Jupyter Agents: Training LLMs to Reason with Notebooks

Hugging Face introduces a pipeline for generating high-quality synthetic data and optimized scaffolding to transform small LLMs like Qwen3-4B into state-of-the-art data science agents.

219

mmBERT: ModernBERT goes Multilingual

Hugging Face introduces mmBERT, a massively multilingual encoder model trained on 3T+ tokens across 1,800+ languages that outperforms XLM-R in both performance and inference speed.

220

EmbeddingGemma 300M release notes

Google released EmbeddingGemma, a 308M‑parameter multilingual embedding model that tops the MTEB benchmark for sub‑500M models and is optimized for on‑device use.

221

SandboxAQ SAIR Dataset Release

SandboxAQ has released SAIR, the largest dataset of co-folded 3D protein-ligand structures paired with experimental IC50 labels, providing over 5 million AI-generated structures to accelerate AI-powered drug discovery.

222

Hugging Face ZeroGPU Ahead-of-Time Compilation Guide

Hugging Face introduces ahead-of-time (AoT) compilation for ZeroGPU Spaces, enabling speedups of 1.3x to 1.8x for generative models like Flux, Wan, and LTX by eliminating just-in-time compilation overhead.

223

NVIDIA Nemotron Post-Training Dataset v2 and Nemotron Nano 2 9B Release

NVIDIA has released a 6-million sample multilingual reasoning dataset and the Nemotron Nano 2 9B model, which utilizes a hybrid Transformer-Mamba architecture to optimize reasoning costs and throughput.

224

Generate Images with Claude and Hugging Face

Hugging Face enables image generation within Claude by connecting the AI to Hugging Face Spaces via the Model Context Protocol (MCP) server.

225

Hugging Face MCP for Research: Connecting AI to Research Tools

Hugging Face introduces the Research Tracker MCP, enabling AI agents to automate research discovery by integrating arXiv, GitHub, and Hugging Face through the Model Context Protocol.

226

Hugging Face kernel-builder: A Guide to Building and Scaling Production-Ready CUDA Kernels

Hugging Face introduces the kernel-builder library to simplify the development, multi-architecture compilation, and distribution of production-ready CUDA kernels via the Hugging Face Hub.

227

Kimina-Prover-RL Release

Hugging Face has released Kimina-Prover-RL, an open-source RL training pipeline and two SOTA models (0.6B and 1.7B) for formal theorem proving in Lean 4.

228

Arm and ExecuTorch 0.7: Expanding Generative AI to Billions of Devices

Arm and the ExecuTorch 0.7 beta enable automatic AI acceleration via KleidiAI, leveraging the SDOT instruction to bring LLMs like Llama 3.2 to billions of existing Arm-based devices.

229

Arm Neural Super Sampling (NSS) Release

Arm has released Neural Super Sampling (NSS), an AI-powered upscaling solution designed to reduce GPU workloads and enable high-resolution rendering on mobile devices.

230

FilBench: Evaluating LLM Capabilities in Philippine Languages

Hugging Face has introduced FilBench, a comprehensive evaluation suite designed to assess the fluency, linguistic abilities, and cultural knowledge of LLMs in Tagalog, Filipino, and Cebuano.

231

TextQuests: Evaluating LLM Agentic Reasoning in Text-Based Video Games

Hugging Face introduces TextQuests, a benchmark using 25 classic Infocom interactive fiction games to evaluate the long-context reasoning and exploratory capabilities of LLM agents.

232

Hugging Face AI Sheets Release

Hugging Face has released AI Sheets, an open-source no-code tool for building, transforming, and enriching datasets using open AI models.

233

Accelerate ND-Parallel: Efficient Multi-GPU Training Guide

Hugging Face has integrated ND-Parallelism into Accelerate and Axolotl, allowing users to combine Data, Fully Sharded Data, Tensor, and Context parallelism strategies to optimize multi-GPU training for massive models.

234

Vision Language Model Alignment in TRL

Hugging Face has expanded the TRL library to support advanced alignment methods for Vision Language Models, including MPO, GRPO, and GSPO, alongside native SFT support and vLLM integration.

235

NVIDIA AI-Q Blueprint: Top-Ranking Open Deep Research Agent on DeepResearch Bench

NVIDIA's AI-Q Blueprint achieves the top spot for open-licensed stacks on the Hugging Face DeepResearch Bench, utilizing a combination of Llama 3.3-70B Instruct and Llama-3.3-Nemotron-Super-49B-v1.5.

236

3LM: A Benchmark for Arabic LLMs in STEM and Code

Hugging Face and TII UAE introduce 3LM, the first comprehensive benchmark designed to evaluate Arabic Large Language Models on STEM subjects and code generation.

237

Implementing MCP Servers in Python with Gradio

Hugging Face introduces a method for Python developers to quickly build Model Context Protocol (MCP) servers using Gradio, enabling LLMs to integrate with thousands of AI models and Spaces on the Hugging Face Hub.

238

Hugging Face Trackio Release

Hugging Face has released Trackio, a free, open-source, lightweight experiment tracking library that serves as a drop-in replacement for wandb with native integration for Hugging Face Spaces and Datasets.

239

Hugging Face CLI Update: Transition to hf Command

Hugging Face has renamed the huggingface-cli to hf, introducing a more ergonomic resource-action command structure and a new hf jobs service for running scripts on HF infrastructure.

240

Parquet Content-Defined Chunking for Efficient Data Deduplication

Hugging Face introduces Parquet Content-Defined Chunking (CDC), now available in PyArrow and Pandas, to enable efficient deduplication of Parquet files on the Xet storage layer, significantly reducing upload and download times.

241

Fast LoRA Inference for Flux with Diffusers and PEFT

Hugging Face introduces an optimization recipe for Flux.1-Dev that achieves up to 2.23x speedup in LoRA inference by combining hotswapping, torch.compile, and FP8 quantization.

242

TimeScope: A New Benchmark for Long-Video Large Multimodal Model Understanding

Hugging Face introduces TimeScope, an open-source benchmark that evaluates the temporal comprehension of vision-language models using video needles inserted into content ranging from 1 minute to 8 hours.

243

NVIDIA NIM Integration with Hugging Face

NVIDIA has expanded NVIDIA NIM to support over 100,000 LLMs on Hugging Face, providing a single Docker container that automates model analysis, backend selection, and performance optimization for rapid deployment.

244

Arc Virtual Cell Challenge Primer

The Arc Virtual Cell Challenge tasks participants with training models to predict the effects of gene silencing via CRISPR on cell transcriptomes, utilizing a dataset of 300k single-cell RNA sequencing profiles.

245

Gradio 5.38.0 MCP Server Improvements

Gradio version 5.38.0 introduces five key updates to its Model Context Protocol (MCP) server capabilities, including seamless local file support, real-time progress notifications, and automated OpenAPI spec transformation.

246

Hugging Face FutureBench: Evaluating AI Agents on Future Event Prediction

Hugging Face introduced FutureBench, a benchmark that evaluates AI agents' reasoning and synthesis capabilities by requiring them to predict future real-world events, effectively eliminating data contamination.

247

Consilium: Multi-LLM Collaboration Platform

Consilium is a multi-LLM platform that enables multiple AI models to reach consensus through structured debate and research, integrating as both a Gradio interface and an MCP server.

248

Ettin Suite: SoTA Paired Encoders and Decoders

Hugging Face introduces Ettin, a suite of paired encoder-only and decoder-only models (17M-1B parameters) trained on identical data and recipes to provide a controlled comparison of architectural performance.

249

Hugging Face Hub Migration from Git LFS to Xet

Hugging Face has migrated 500,000 repositories and 20 PB of data from Git LFS to Xet, a content-addressed storage system designed to scale with AI workloads.

250

Kimina-Prover: Applying Test-time RL Search on Large Formal Reasoning Models

Hugging Face announces Kimina-Prover-72B, a state-of-the-art theorem proving model for Lean 4 that achieves a 92.2% pass rate on the miniF2F benchmark using a novel Test-Time Reinforcement Learning (TTRL) search framework.