✷ The archive · 11 labs · 871 dispatches
The labs
No more opening a dozen official blogs every morning. First-hand releases from OpenAI, Anthropic, DeepMind and the rest, each with its substance pulled out.
Mask2Former and OneFormer: Universal Image Segmentation Models in 🤗 Transformers
Hugging Face released Mask2Former and OneFormer in the Transformers library, providing universal architectures that handle instance, semantic, and panoptic segmentation with a single model.
PaddlePaddle Integration with Hugging Face Hub
Hugging Face has partnered with PaddlePaddle to integrate its deep learning platform and libraries, starting with PaddleNLP, into the Hugging Face Hub for improved accessibility and sharing.
Image Similarity with Hugging Face Datasets and Transformers
Hugging Face demonstrates how to build an image similarity system using the Transformers and Datasets libraries by computing dense vector embeddings and measuring cosine similarity.
AI for Game Development: Using LLMs for Game Design
Hugging Face demonstrates how Large Language Models like ChatGPT can be used as brainstorming and acceleration tools for game design, specifically for defining core features of a farming game.
Introduction to Graph Machine Learning
Hugging Face provides a comprehensive overview of Graph Machine Learning, detailing how graphs are represented and the evolution from pre-neural features to Graph Neural Networks and Graph Transformers.
AI for Game Development: Creating a Farming Game in 5 Days (Part 1)
Hugging Face demonstrates how to use Stable Diffusion to establish a visual art style and concept art for a farming game, which is then implemented in Unity.
Accelerating PyTorch Transformers with Intel Sapphire Rapids – part 1
Hugging Face shows how to train PyTorch Transformers on a cluster of Intel Sapphire Rapids CPUs using IPEX and oneCCL, achieving up to 8× speed‑up over Ice Lake and near‑linear scaling across four nodes.
Zero-shot image segmentation with CLIPSeg
Hugging Face introduces CLIPSeg, a zero-shot image segmentation model that uses CLIP embeddings to create segmentation masks from either text or image prompts without requiring category-specific training.
Hugging Face Model Cards Documentation Framework
Hugging Face has released a suite of tools and resources, including a GUI-based creator tool and a standardized template, to improve the accessibility and standardization of machine learning model documentation.
Hugging Face Audio Datasets Guide – How to Load, Process, and Stream Audio Data
Hugging Face announced a comprehensive guide showing that the 🤗 Datasets library can load, preprocess, and stream any audio dataset from the Hub with just a few lines of Python code, enabling efficient research on speech and audio tasks.
Hugging Face Ethics and Society Newsletter #2: Addressing Bias in Machine Learning
Hugging Face outlines a sociotechnical framework for mitigating machine learning bias by treating biases as risk factors that must be addressed across task definition, dataset curation, and model training.
Habana Gaudi2 vs Nvidia A100 80GB Performance Benchmarks
Hugging Face benchmarks show that Habana Gaudi2 provides approximately twice the throughput of Nvidia A100 80GB for both training and inference across BERT, Stable Diffusion, and T5-3B models.
Hugging Face Announces Bumblebee: Transformers and Stable Diffusion in Pure Elixir
Hugging Face released Bumblebee, a pure‑Elixir implementation of Transformers that brings models from GPT‑2 to Stable Diffusion to the Elixir ecosystem, enabling native CPU/GPU inference without external dependencies.
Illustrating Reinforcement Learning from Human Feedback (RLHF)
Hugging Face explains Reinforcement Learning from Human Feedback (RLHF), a three-step process used to align large language models with complex human values by optimizing them using a reward model based on human preferences.
Deep Learning with Proteins – Hugging Face guide to protein language models and folding
Hugging Face announced a tutorial series showing how to fine‑tune protein language models and use ESMFold for protein folding, demonstrating that transfer learning techniques from NLP can be applied directly to protein sequence tasks.
Time Series Transformer probabilistic forecasting with 🤗 Transformers
Hugging Face released a vanilla Transformer model for global probabilistic time‑series forecasting, demonstrating state‑of‑the‑art performance on the Tourism Monthly benchmark.
Stable Diffusion Core ML on Apple Silicon – How to Run and Optimize
Hugging Face released Core ML‑converted Stable Diffusion checkpoints for Apple Silicon, enabling on‑device image generation in Python or Swift with up to 18 seconds per image on an M1 Max.
VQ-Diffusion: Conditional Latent Diffusion in Discrete Space
VQ-Diffusion is a conditional latent diffusion model that operates on a quantized discrete latent space, offering faster inference and higher image quality than traditional autoregressive models.
Hugging Face 2023 Internship Program
Hugging Face has announced its 2023 internship program, offering roles across Open Source, Science, and Social Impact teams to democratize responsible machine learning.
Hugging Face Diffusion Models Class and Community Event
Hugging Face announced a free Diffusion Models Class launching November 28, 2022, accompanied by a live community event on November 30 featuring researchers from Stability AI, Meta, and Runway.
Hugging Face Director of Machine Learning Insights Part 4
Four Machine Learning Directors share industry-specific insights on the impact, challenges, and integration pitfalls of ML in e-commerce, engineering, education, and SaaS.
Hugging Face Inference Solutions Overview November 2022
Hugging Face announced a suite of free and paid inference options—including a widget, API, Inference Endpoints, and Spaces—to simplify model testing, deployment, and production scaling.
Hugging Face Accelerating Document AI
Hugging Face provides a comprehensive guide to using open-source multimodal models to automate document classification, parsing, and visual question answering for enterprise workflows.
Sentiment Analysis on Encrypted Data with Homomorphic Encryption
Hugging Face demonstrates how to use the Concrete-ML library to perform sentiment analysis on encrypted data using a combination of BERT transformers and XGBoost with Fully Homomorphic Encryption (FHE).
Hugging Face and arXiv Integration for Machine Learning Demos
Hugging Face has integrated Hugging Face Spaces with arXivLabs to provide interactive machine learning demos directly on arXiv paper abstract pages.
Hugging Face Pricing Update November 2022
Hugging Face has transitioned to a compute-based monetization model, sunsetting the Paid tier of the Inference API in favor of Inference Endpoints and hardware upgrades for Spaces.
Contrastive Search for Human-Level Text Generation in Transformers
Hugging Face has integrated Contrastive Search into the transformers library, a decoding method that prevents model degeneration and maintains semantic coherence across 16 languages using off-the-shelf models.
Dreambooth Stable Diffusion fine‑tuning guide with Diffusers
Hugging Face released detailed recommendations for training Stable Diffusion with Dreambooth using the Diffusers library, showing that low learning rates, enough steps, prior preservation for faces, and text‑encoder fine‑tuning yield the highest quality results.
Fine-Tuning Whisper for Multilingual ASR with Hugging Face Transformers
Hugging Face provides a comprehensive guide on fine-tuning OpenAI's Whisper model for multilingual automatic speech recognition (ASR), demonstrating a 31.5% absolute WER improvement on Hindi using only 8 hours of data.
Hugging Face Optimum Intel and OpenVINO Integration
Hugging Face has integrated Intel OpenVINO into Optimum Intel, enabling accelerated inference and quantization for Transformer models on Intel hardware.
Evaluating Language Model Bias with 🤗 Evaluate
Hugging Face added bias metrics—toxicity, language polarity, and HONEST—to the 🤗 Evaluate library, enabling systematic measurement of harmful language in causal language models.
Distributed Training with PyTorch DDP, Accelerate, and Transformers Trainer
Hugging Face explains how to implement distributed training using three levels of abstraction: native PyTorch DDP, the Accelerate library, and the high-level Transformers Trainer API.
MTEB: Massive Text Embedding Benchmark
Hugging Face introduced MTEB, a massive and multilingual benchmark consisting of 56 datasets across 8 tasks to evaluate the performance of text embedding models.
Hugging Face Inference Endpoints
Hugging Face Inference Endpoints is a managed service that allows users to deploy machine learning models from the Hugging Face Hub to scalable, secure cloud infrastructure with a few clicks.
Stable Diffusion JAX and Flax Integration
Hugging Face Diffusers version 0.5.1 introduces support for Flax, enabling high-speed Stable Diffusion inference on Google TPUs via JAX.
Hugging Face BLOOM Inference Optimization
Hugging Face achieved a 5x reduction in latency and a 50x increase in throughput for the BLOOM model by transitioning from Pipeline Parallelism to Tensor Parallelism and implementing custom CUDA kernels.
Hugging Face Introduces DOI Support for Models and Datasets
Hugging Face now lets users generate Digital Object Identifiers (DOIs) for Hub models and datasets, providing permanent, citable links that persist across versions.
Japanese Stable Diffusion release by rinna
rinna released Japanese Stable Diffusion, a Japanese‑language fine‑tuned version of Stable Diffusion that generates culturally appropriate images from Japanese prompts.
Hugging Face Zero-Shot Evaluation on the Hub
Hugging Face has introduced zero-shot evaluation for causal language models on the Hub, enabling users to benchmark models up to 66 billion parameters without writing code.
Hugging Face AutoTrain Image Classification
Hugging Face has added Image Classification to AutoTrain, enabling users to train custom image categorization models without writing code or configuring hyperparameters.
Hugging Face Accelerate: Running Large Models with PyTorch
Hugging Face Accelerate enables the execution of massive AI models on consumer hardware by leveraging PyTorch's meta device and sharded checkpoints to manage memory across GPUs, CPU RAM, and disk.
SetFit: Efficient Few-Shot Learning Without Prompts
Hugging Face introduces SetFit, a prompt-free framework for few-shot fine-tuning of Sentence Transformers that achieves high accuracy with minimal labeled data.
Hugging Face Ethics and Society Newsletter #1
Hugging Face introduces its Ethics and Society newsletter and outlines a decentralized, value-driven approach to operationalizing AI ethics through collaboration, transparency, and responsibility.
Incredibly Fast BLOOM Inference with DeepSpeed and Accelerate
Hugging Face demonstrates sub‑millisecond per‑token generation for the 176B‑parameter BLOOM model using DeepSpeed‑Inference tensor parallelism and Accelerate pipeline parallelism on 8×80 GB A100 GPUs.
Diffusers 0.3 release adds image‑to‑image, textual inversion, inpainting, GPU optimizations, Mac MPS, ONNX support and new docs
Hugging Face announced Diffusers version 0.3, introducing image‑to‑image, textual inversion, experimental inpainting, smaller‑GPU optimizations, Mac MPS support, an ONNX exporter, expanded documentation, and a wave of community projects.
Training Decision Transformers for Offline Reinforcement Learning
Hugging Face provides a guide and implementation for training an offline Decision Transformer from scratch to solve the HalfCheetah environment using the transformers Trainer and a custom data collator.
Training Language Models with Megatron-LM
Hugging Face provides a guide on using NVIDIA's Megatron-LM framework to efficiently pre-train large language models on GPUs, including integration with the Transformers library.
OpenRAIL: Towards open and responsible AI licensing frameworks
Hugging Face announced OpenRAIL, a set of AI‑specific licenses that combine open access with use‑based restrictions to promote responsible deployment of machine‑learning models.
Visualize proteins on Hugging Face Spaces
Hugging Face provides a guide on integrating 3Dmol.js into Hugging Face Spaces via Gradio to enable 3D protein structure visualization in the browser.
Pre-training BERT with Hugging Face Transformers and Habana Gaudi
Hugging Face demonstrates how to pre-train BERT-base from scratch using Habana Gaudi DL1 instances on AWS, achieving a 25% cost reduction compared to NVIDIA V100 GPU-based training.