✷ The archive · 11 labs · 871 dispatches
The labs
No more opening a dozen official blogs every morning. First-hand releases from OpenAI, Anthropic, DeepMind and the rest, each with its substance pulled out.
SOTA OCR with Core ML and dots.ocr
Hugging Face demonstrates the conversion of the 3B parameter dots.ocr model to run on-device using Core ML and MLX, achieving performance that surpasses Gemini 2.5 Pro on OmniDocBench.
Hugging Face Introduces RTEB for Retrieval Embedding Evaluation
Hugging Face has launched the beta version of the Retrieval Embedding Benchmark (RTEB), a new standard designed to measure the true generalization and retrieval accuracy of embedding models using a hybrid of open and private datasets.
VibeGame: A High-Level Declarative Engine for AI-Assisted Game Development
Hugging Face researcher Dylan Ebert introduces VibeGame, a declarative game engine designed to enable 'vibe coding' by combining high-level abstractions with a web-based stack optimized for AI proficiency.
Qwen3-8B acceleration on Intel Core Ultra with depth‑pruned draft models
Accelerated Qwen3-8B on Intel Core Ultra using speculative decoding and a depth‑pruned Qwen3‑0.6B draft achieves ~1.4× speedup, enabling fast local AI agents via Hugging Face smolagents.
Nemotron-Personas-Japan Release
NVIDIA has released Nemotron-Personas-Japan, an open synthetic dataset of 6 million Japanese personas designed to support the development of Sovereign AI and culturally aware LLMs.
Swift Transformers 1.0 release notes / what's new
Hugging Face has released Swift Transformers 1.0, a stable library providing tokenizers, hub access, and model wrappers to simplify local LLM integration for Apple Silicon developers.
Smol2Operator release: turning a small VLM into an open‑source agentic GUI coder
Hugging Face released Smol2Operator, a post‑training recipe that converts the 2.2 B‑parameter SmolVLM2‑2.2B‑Instruct model into an open‑source agentic GUI coder, with all code, datasets, and the resulting model publicly available.
SyGra: A Low-Code Framework for LLM and SLM Data Generation
SyGra is a low-code/no-code framework developed by ServiceNow AI to simplify the creation, transformation, and alignment of datasets for Large and Small Language Models.
Hugging Face Gaia2 and Meta Agents Research Environments (ARE) Release
Hugging Face has released Gaia2, a read-and-write agentic benchmark, and the Meta Agents Research Environments (ARE) framework to evaluate and debug AI agents in complex, real-world simulated conditions.
Scaleway added as Hugging Face Inference Provider – capabilities, usage, and billing
Scaleway is now an official Inference Provider on the Hugging Face Hub, offering low‑latency, European‑hosted serverless access to frontier models with flexible billing options.
Hugging Face RiskRubric.ai Announcement
Hugging Face has announced RiskRubric.ai, a standardized risk assessment platform that evaluates AI models across six pillars of transparency, reliability, security, privacy, safety, and reputation to provide comparable risk scores and letter grades.
Hugging Face Integrates Public AI as an Inference Provider
Hugging Face has added Public AI as a supported Inference Provider, enabling serverless access to sovereign models from institutions like the Swiss AI Initiative and AI Singapore.
LeRobotDataset v3.0 release notes / what's new
Hugging Face has released LeRobotDataset v3.0, a standardized robotics dataset format that enables large-scale data storage and streaming to support millions of episodes.
Visible Watermarking with Gradio
Hugging Face has introduced a simple way to add visible watermarks to AI-generated images, videos, and text using the Gradio library to improve synthetic content transparency.
Writer Palmyra-mini Family Release
Writer has released the Palmyra-mini family, a set of lightweight open models (1.5B to 1.7B parameters) featuring a base model and two specialized reasoning variants optimized for logic and mathematics.
Transformers 4.40 performance upgrades for OpenAI GPT‑OSS: zero‑build kernels, MXFP4 quantization, parallelism, and faster loading
Hugging Face added zero-build kernels, MXFP4 4-bit quantization, tensor and expert parallelism, dynamic sliding-window cache, continuous batching, and faster model loading to the transformers library for OpenAI's GPT‑OSS models, dramatically improving efficiency and scalability.
Fine-tuning Hugging Face Hub LLMs with Together AI
Together AI and Hugging Face have integrated their platforms, allowing developers to fine-tune any compatible LLM from the Hugging Face Hub using Together AI's infrastructure.
Jupyter Agents: Training LLMs to Reason with Notebooks
Hugging Face introduces a pipeline for generating high-quality synthetic data and optimized scaffolding to transform small LLMs like Qwen3-4B into state-of-the-art data science agents.
mmBERT: ModernBERT goes Multilingual
Hugging Face introduces mmBERT, a massively multilingual encoder model trained on 3T+ tokens across 1,800+ languages that outperforms XLM-R in both performance and inference speed.
EmbeddingGemma 300M release notes
Google released EmbeddingGemma, a 308M‑parameter multilingual embedding model that tops the MTEB benchmark for sub‑500M models and is optimized for on‑device use.
SandboxAQ SAIR Dataset Release
SandboxAQ has released SAIR, the largest dataset of co-folded 3D protein-ligand structures paired with experimental IC50 labels, providing over 5 million AI-generated structures to accelerate AI-powered drug discovery.
Hugging Face ZeroGPU Ahead-of-Time Compilation Guide
Hugging Face introduces ahead-of-time (AoT) compilation for ZeroGPU Spaces, enabling speedups of 1.3x to 1.8x for generative models like Flux, Wan, and LTX by eliminating just-in-time compilation overhead.
NVIDIA Nemotron Post-Training Dataset v2 and Nemotron Nano 2 9B Release
NVIDIA has released a 6-million sample multilingual reasoning dataset and the Nemotron Nano 2 9B model, which utilizes a hybrid Transformer-Mamba architecture to optimize reasoning costs and throughput.
Generate Images with Claude and Hugging Face
Hugging Face enables image generation within Claude by connecting the AI to Hugging Face Spaces via the Model Context Protocol (MCP) server.
Hugging Face MCP for Research: Connecting AI to Research Tools
Hugging Face introduces the Research Tracker MCP, enabling AI agents to automate research discovery by integrating arXiv, GitHub, and Hugging Face through the Model Context Protocol.
Hugging Face kernel-builder: A Guide to Building and Scaling Production-Ready CUDA Kernels
Hugging Face introduces the kernel-builder library to simplify the development, multi-architecture compilation, and distribution of production-ready CUDA kernels via the Hugging Face Hub.
Kimina-Prover-RL Release
Hugging Face has released Kimina-Prover-RL, an open-source RL training pipeline and two SOTA models (0.6B and 1.7B) for formal theorem proving in Lean 4.
Arm and ExecuTorch 0.7: Expanding Generative AI to Billions of Devices
Arm and the ExecuTorch 0.7 beta enable automatic AI acceleration via KleidiAI, leveraging the SDOT instruction to bring LLMs like Llama 3.2 to billions of existing Arm-based devices.
Arm Neural Super Sampling (NSS) Release
Arm has released Neural Super Sampling (NSS), an AI-powered upscaling solution designed to reduce GPU workloads and enable high-resolution rendering on mobile devices.
FilBench: Evaluating LLM Capabilities in Philippine Languages
Hugging Face has introduced FilBench, a comprehensive evaluation suite designed to assess the fluency, linguistic abilities, and cultural knowledge of LLMs in Tagalog, Filipino, and Cebuano.
TextQuests: Evaluating LLM Agentic Reasoning in Text-Based Video Games
Hugging Face introduces TextQuests, a benchmark using 25 classic Infocom interactive fiction games to evaluate the long-context reasoning and exploratory capabilities of LLM agents.
Hugging Face AI Sheets Release
Hugging Face has released AI Sheets, an open-source no-code tool for building, transforming, and enriching datasets using open AI models.
Accelerate ND-Parallel: Efficient Multi-GPU Training Guide
Hugging Face has integrated ND-Parallelism into Accelerate and Axolotl, allowing users to combine Data, Fully Sharded Data, Tensor, and Context parallelism strategies to optimize multi-GPU training for massive models.
Vision Language Model Alignment in TRL
Hugging Face has expanded the TRL library to support advanced alignment methods for Vision Language Models, including MPO, GRPO, and GSPO, alongside native SFT support and vLLM integration.
NVIDIA AI-Q Blueprint: Top-Ranking Open Deep Research Agent on DeepResearch Bench
NVIDIA's AI-Q Blueprint achieves the top spot for open-licensed stacks on the Hugging Face DeepResearch Bench, utilizing a combination of Llama 3.3-70B Instruct and Llama-3.3-Nemotron-Super-49B-v1.5.
3LM: A Benchmark for Arabic LLMs in STEM and Code
Hugging Face and TII UAE introduce 3LM, the first comprehensive benchmark designed to evaluate Arabic Large Language Models on STEM subjects and code generation.
Implementing MCP Servers in Python with Gradio
Hugging Face introduces a method for Python developers to quickly build Model Context Protocol (MCP) servers using Gradio, enabling LLMs to integrate with thousands of AI models and Spaces on the Hugging Face Hub.
Hugging Face Trackio Release
Hugging Face has released Trackio, a free, open-source, lightweight experiment tracking library that serves as a drop-in replacement for wandb with native integration for Hugging Face Spaces and Datasets.
Hugging Face CLI Update: Transition to hf Command
Hugging Face has renamed the huggingface-cli to hf, introducing a more ergonomic resource-action command structure and a new hf jobs service for running scripts on HF infrastructure.
Parquet Content-Defined Chunking for Efficient Data Deduplication
Hugging Face introduces Parquet Content-Defined Chunking (CDC), now available in PyArrow and Pandas, to enable efficient deduplication of Parquet files on the Xet storage layer, significantly reducing upload and download times.
Fast LoRA Inference for Flux with Diffusers and PEFT
Hugging Face introduces an optimization recipe for Flux.1-Dev that achieves up to 2.23x speedup in LoRA inference by combining hotswapping, torch.compile, and FP8 quantization.
TimeScope: A New Benchmark for Long-Video Large Multimodal Model Understanding
Hugging Face introduces TimeScope, an open-source benchmark that evaluates the temporal comprehension of vision-language models using video needles inserted into content ranging from 1 minute to 8 hours.
NVIDIA NIM Integration with Hugging Face
NVIDIA has expanded NVIDIA NIM to support over 100,000 LLMs on Hugging Face, providing a single Docker container that automates model analysis, backend selection, and performance optimization for rapid deployment.
Arc Virtual Cell Challenge Primer
The Arc Virtual Cell Challenge tasks participants with training models to predict the effects of gene silencing via CRISPR on cell transcriptomes, utilizing a dataset of 300k single-cell RNA sequencing profiles.
Gradio 5.38.0 MCP Server Improvements
Gradio version 5.38.0 introduces five key updates to its Model Context Protocol (MCP) server capabilities, including seamless local file support, real-time progress notifications, and automated OpenAPI spec transformation.
Hugging Face FutureBench: Evaluating AI Agents on Future Event Prediction
Hugging Face introduced FutureBench, a benchmark that evaluates AI agents' reasoning and synthesis capabilities by requiring them to predict future real-world events, effectively eliminating data contamination.
Consilium: Multi-LLM Collaboration Platform
Consilium is a multi-LLM platform that enables multiple AI models to reach consensus through structured debate and research, integrating as both a Gradio interface and an MCP server.
Ettin Suite: SoTA Paired Encoders and Decoders
Hugging Face introduces Ettin, a suite of paired encoder-only and decoder-only models (17M-1B parameters) trained on identical data and recipes to provide a controlled comparison of architectural performance.
Hugging Face Hub Migration from Git LFS to Xet
Hugging Face has migrated 500,000 repositories and 20 PB of data from Git LFS to Xet, a content-addressed storage system designed to scale with AI workloads.
Kimina-Prover: Applying Test-time RL Search on Large Formal Reasoning Models
Hugging Face announces Kimina-Prover-72B, a state-of-the-art theorem proving model for Lean 4 that achieves a 92.2% pass rate on the miniF2F benchmark using a novel Test-Time Reinforcement Learning (TTRL) search framework.