The archive · 11 labs · 871 dispatches

The labs

No more opening a dozen official blogs every morning. First-hand releases from OpenAI, Anthropic, DeepMind and the rest, each with its substance pulled out.

551

Hugging Face Leaderboard Templates: Implementing the Vectara HHEM Leaderboard

Hugging Face has released open-source leaderboard templates that enable developers to build dynamic LLM evaluation boards, as demonstrated by Vectara's new Hughes Hallucination Evaluation Model (HHEM) leaderboard.

552

Unsloth and Hugging Face TRL Integration for Faster LLM Fine-tuning

Unsloth is a lightweight library that accelerates LLM fine-tuning by up to 2.7x and reduces memory usage by up to 74% with 0% accuracy degradation compared to QLoRA.

553

aMUSEd: Efficient Text-to-Image Generation

Hugging Face has released aMUSEd, an efficient non-diffusion text-to-image model based on Masked Image Modeling (MIM) and an open reproduction of Google's MUSE.

554

Hugging Face SDXL Dreambooth LoRA Advanced Training Guide

Hugging Face introduces an advanced training script for SDXL Dreambooth LoRAs, combining Pivotal Tuning and the Prodigy optimizer to improve concept capture and image quality.

555

Speculative Decoding Enables 2× Faster Whisper Inference

Hugging Face demonstrates that speculative decoding halves Whisper transcription latency while preserving identical outputs and accuracy.

556

2023 Year of Open LLMs Review

Hugging Face’s 2023 recap shows a surge of open‑source LLM releases, smaller high‑performing models, and new fine‑tuning techniques that dramatically broaden access and community participation.

557

Mixture of Experts (MoE) Explained

Mixture of Experts (MoE) allows transformer models to scale parameters while maintaining efficient pretraining and faster inference by activating only a subset of neural network experts per token.

558

Mixtral 8x7B Release Notes

Mistral AI has released Mixtral 8x7B, a Mixture of Experts (MoE) model that outperforms Llama 2 70B and matches GPT-3.5 performance on most benchmarks while remaining commercially permissive under Apache 2.0.

559

SetFitABSA: Few-Shot Aspect Based Sentiment Analysis

Hugging Face and Intel Labs introduced SetFitABSA, a prompt-less, few-shot framework for Aspect-Based Sentiment Analysis that outperforms larger generative models like Llama 2 and T5 in low-data scenarios.

560

Optimum-NVIDIA Release Notes

Hugging Face has released Optimum-NVIDIA, an inference library that accelerates LLM inference on NVIDIA platforms by up to 28x using FP8 quantization and TensorRT-LLM.

561

Hugging Face LoRA dynamic loading speeds inference 300% and cuts latency

Hugging Face announced a dynamic LoRA loading system that reduces warm‑up time from 25 s to 3 s, delivering up to 300 % faster LoRA inference and cutting total response time from 35 s to 13 s.

562

Hugging Face and AMD GPU Acceleration for LLMs

Hugging Face and AMD have integrated out-of-the-box support for AMD Instinct GPUs into the Transformers library and Text Generation Inference, enabling high-performance LLM execution without code changes.

563

Hugging Face Open LLM Leaderboard DROP Benchmark Analysis

Hugging Face has removed the DROP benchmark from the Open LLM Leaderboard after discovering that flawed normalization and stop-token configurations caused most models to score incorrectly low.

564

SDXL and Stable Diffusion Fast Inference with Latent Consistency LoRAs

Hugging Face introduces LCM LoRAs, a method to enable high-quality image generation in 4 to 8 steps for SDXL and Stable Diffusion models, significantly reducing inference time.

565

Prodigy-HF Integration Release Notes

Explosion has released Prodigy-HF, a plugin that enables direct fine-tuning of Hugging Face transformer models on annotated data and the ability to upload datasets directly to the Hugging Face Hub.

566

Deploying Llama 2 on AWS Inferentia2 with optimum-neuron

Hugging Face has integrated optimum-neuron with the AWS Neuron SDK to enable the deployment of Llama 2 models on AWS Inferentia2 accelerators for high-performance text generation.

567

Comparing RoBERTa, Llama 2, and Mistral for Disaster Tweet Classification with LoRA

A comparative study reveals that the smaller RoBERTa model outperforms Llama 2 and Mistral 7B in binary classification of disaster tweets when fine-tuned using Low-Rank Adaptation (LoRA).

568

Hugging Face Hub Storage Regions

Hugging Face has introduced Storage Regions for Enterprise Hub customers, allowing organizations to select where their models and datasets are stored to improve regulatory compliance and data transfer performance.

569

Personal Copilot: Train Your Own Coding Assistant

Hugging Face demonstrates how to create a personalized coding assistant, HugCoder, by fine-tuning StarCoder on a specific codebase using QLoRA and full fine-tuning techniques.

570

Hugging Face and Renumics Spotlight Integration for Scalable Data Inspection

Hugging Face has integrated with Renumics Spotlight to enable interactive, one-line-of-code visualization and inspection of ML datasets, including support for multimodal data and model results.

571

Optimizing Stable Diffusion XL (SDXL) for Inference Speed and Memory

Hugging Face explores several optimization techniques for Stable Diffusion XL (SDXL), demonstrating how to reduce memory usage from 28GB to as low as 11.47GB and decrease inference latency from 72.2 seconds to approximately 10.3 seconds.

572

Deploying Embedding Models with Hugging Face Inference Endpoints

Hugging Face introduces Text Embeddings Inference (TEI) via Inference Endpoints, providing a high-performance, cost-efficient way to deploy open-source embedding models for RAG and semantic search.

573

The N Implementation Details of RLHF with PPO – Hugging Face Blog Summary

The Hugging Face blog post reproduces OpenAI’s 2019 RLHF codebase, matches its learning curves, and details N implementation specifics, including a key PyTorch Adam optimizer difference that causes more aggressive updates.

574

Gradio-Lite: Serverless Gradio Running Entirely in Your Browser

Hugging Face introduces Gradio-Lite (@gradio/lite), a JavaScript library that uses Pyodide to run Gradio applications directly in the web browser, eliminating the need for server-side infrastructure.

575

Accelerating Hugging Face Models with ONNX Runtime

Hugging Face and ONNX Runtime enable performance acceleration for over 130,000 models, including a latency reduction of up to 74.30% for the whisper-tiny model compared to PyTorch.

576

Hugging Face Chat Templates

Hugging Face introduced chat templates as a Jinja-based system to ensure chat models receive inputs formatted exactly as they were during training, preventing silent performance degradation.

577

Accelerating Stable Diffusion XL Inference with JAX on Cloud TPU v5e

Hugging Face Diffusers now supports serving Stable Diffusion XL (SDXL) using JAX on Cloud TPU v5e, delivering up to 2.4x greater performance per dollar compared to TPU v4.

578

Deploying AI Comic Factory via Hugging Face Inference API

Hugging Face provides a guide on deploying a private instance of the AI Comic Factory using the Inference API, leveraging Llama-2 and SDXL 1.0 models.

579

Finetuning Stable Diffusion with DDPO via TRL

Hugging Face has integrated Denoising Diffusion Policy Optimization (DDPO) into the TRL library, enabling the alignment of Stable Diffusion models with human preferences using reinforcement learning.

580

Hugging Face Ethics and Society Update Summer 2023

Hugging Face detailed its Summer 2023 efforts to influence AI regulation in the US, EU, and UK, while advancing open-source ethics through public advocacy and technical research.

581

Hugging Face Guide: Training a LLaMA 2 Chatbot Without Code

Hugging Face provides a no-code workflow using Spaces, AutoTrain, and ChatUI to allow non-engineers to fine-tune LLaMA 2 and deploy it as a functional chat application.

582

Llama 2 on Amazon SageMaker Benchmark

Hugging Face analyzed 60 deployment configurations for Llama 2 on Amazon SageMaker to identify optimal setups for cost, throughput, and latency.

583

Hugging Face Inference for PROs

Hugging Face has introduced Inference for PRO users, providing accelerated API endpoints for curated state-of-the-art models and increased rate limits for the free Inference API.

584

Rocket Money x Hugging Face: Scaling Volatile ML Models in Production

Rocket Money scaled its transaction classification system to over a billion transactions per month using Hugging Face's Inference API to replace a legacy regex-based system.

585

Introduction to 3D Gaussian Splatting

3D Gaussian Splatting is a rasterization technique that enables real-time rendering of photorealistic 3D scenes learned from a small set of images.

586

Hugging Face Object Detection Leaderboard

Hugging Face released an Object Detection Leaderboard that ranks open-source models using COCO-style metrics and published a blog explaining how Average Precision and Average Recall are computed and what factors can influence the results.

587

Optimizing LLMs in Production: Precision, Attention, and Architecture

Hugging Face outlines key techniques for efficient LLM deployment, focusing on lower precision quantization, Flash Attention for memory efficiency, and architectural optimizations like RoPE, ALiBi, MQA, and GQA.

588

Fine-tuning Llama 2 70B using PyTorch FSDP

Hugging Face demonstrates how to fine-tune Llama 2 70B using PyTorch Fully Sharded Data Parallelism (FSDP) and Accelerate to overcome CPU RAM bottlenecks and optimize VRAM usage.

589

Introducing Würstchen: Fast Diffusion for Image Generation

Würstchen is a fast and efficient text-to-image diffusion model that achieves 42x spatial compression to significantly reduce training and inference costs.

590

Hugging Face Transformers Quantization Overview

Hugging Face provides native support for bitsandbytes and auto-gptq quantization schemes to enable large model inference on smaller devices and efficient adapter fine-tuning.

591

SafeCoder vs. Closed-source Code Assistants

Hugging Face introduces SafeCoder, an enterprise-grade code assistant based on the open-source StarCoder models that prioritizes transparency, customization, and data privacy over closed-source alternatives.

592

Efficient Controllable Generation for SDXL with T2I-Adapters

Hugging Face and TencentARC introduce T2I-Adapter-SDXL, a lightweight plug-and-play model that enables precise control over Stable Diffusion XL (SDXL) generation using external signals like sketches and depth maps with significantly lower computational overhead than ControlNet.

593

Falcon 180B Release Notes

TII has released Falcon 180B, the largest openly available language model with 180 billion parameters, trained on 3.5 trillion tokens to rival proprietary models like PaLM-2.

594

Fetch Case Study: Reducing ML Processing Latency by 50% with Amazon SageMaker and Hugging Face

Fetch reduced ML processing latency for receipt scans by 50% and increased document-understanding model accuracy by 200% by migrating its ML pipeline to Amazon SageMaker and Hugging Face.

595

AudioLDM 2 Optimization Guide: Reducing Inference Time with Hugging Face Diffusers

Hugging Face demonstrates how to reduce AudioLDM 2 inference time by over 10x, bringing generation of a 10-second audio sample down to under 1 second using code and model optimizations.

596

Hugging Face Hub Git Authentication Changes

Hugging Face deprecated password-based Git authentication on October 1, 2023, requiring users to switch to personal access tokens or SSH keys for improved security.

597

Code Llama Release Notes

Code Llama is a family of open-access models based on Llama 2, specialized for code tasks with support for infilling and long-context windows up to 100,000 tokens.

598

Hugging Face AutoGPTQ and Transformers Integration

Hugging Face has integrated the AutoGPTQ library into Transformers, enabling the quantization of LLMs to 8, 4, 3, or 2-bit precision to reduce memory requirements with negligible accuracy loss at 4-bit.

599

IDEFICS: An Open Reproduction of State-of-the-art Visual Language Model

Hugging Face has released IDEFICS, an open-access visual language model based on the Flamingo architecture that supports interleaved image and text inputs in 9B and 80B parameter sizes.

600

Hugging Face SafeCoder Announcement

Hugging Face has introduced SafeCoder, a self-hosted, enterprise-grade code assistant solution that allows companies to build and deploy proprietary Code LLMs within their own secure infrastructure.