✷ The archive · 11 labs · 871 dispatches
The labs
No more opening a dozen official blogs every morning. First-hand releases from OpenAI, Anthropic, DeepMind and the rest, each with its substance pulled out.
Banque des Territoires, Polyconseil, and Hugging Face Deploy Sovereign RAG Solution for EduRénov Program
Hugging Face, Polyconseil, and Banque des Territoires launched a sovereign, open‑source RAG system to automate email support for France's EduRénov school‑renovation program, ensuring data residency while scaling generative AI for public policy.
Hugging Face Dataset Hub Search Features Update
Hugging Face has introduced four new search filters—modality, size, format, and library compatibility—to improve the discoverability of over 180,000 public datasets on the Dataset Hub.
Accelerating Protein Language Model ProtST on Intel Gaudi 2
Intel and MILA have optimized the ProtST multi-modal protein language model for Intel Gaudi 2 accelerators, achieving up to 2.92x faster fine-tuning compared to NVIDIA A100 GPUs.
Hugging Face Transformers Code Agent GAIA Benchmark Results
Hugging Face achieved a top ranking on the GAIA benchmark using a Code Agent built with the transformers.agents library, demonstrating that code-based actions are more efficient and effective than JSON-based tool calling.
Google Gemma 2 Release Notes
Google has released Gemma 2, a family of open-weight LLMs available in 9B and 27B parameter sizes, featuring technical advances in distillation and attention mechanisms to improve performance.
XLSCOUT ParaEmbed 2.0 Release
XLSCOUT has released ParaEmbed 2.0, a proprietary embedding model fine-tuned on expert-curated patent data that achieves a 23% increase in accuracy over ParaEmbed 1.0.
Hugging Face Ethics and Society Newsletter #6: The Importance of Data Quality
Hugging Face outlines a holistic, responsible approach to data quality, emphasizing that high-quality data must be fit for its intended purpose to ensure AI model performance, fairness, and scientific reproducibility.
Fine-tuning Microsoft Florence-2 for DocVQA
Hugging Face demonstrates how to fine-tune Microsoft's Florence-2 vision-language model on the DocVQA dataset, improving validation similarity from 0 to 57.0 after seven epochs.
Hugging Face Data Is Better Together Initiative
Hugging Face and Argilla launched the Data Is Better Together (DIBT) initiative to empower the open-source community to collectively create high-quality, diverse, and inclusive datasets for machine learning.
Prezi Case Study: Accelerating ML Roadmap with Hugging Face Expert Support
Prezi is leveraging the Hugging Face Expert Support Program and Inference Endpoints to integrate efficient open-source multimodal models into its Prezi AI presentation generation tool.
BigCodeBench: A New Benchmark for Complex Python Code Generation
Hugging Face has released BigCodeBench, a benchmark of 1,140 function-level tasks designed to evaluate LLMs on practical programming and diverse library usage, addressing the simplicity and contamination issues of HumanEval.
Hugging Face Accelerate: Harmonizing DeepSpeed and FSDP Precision
Hugging Face Accelerate 0.30.0 introduces automatic upcasting for PyTorch FSDP to align its precision handling with DeepSpeed, enabling seamless switching between the two backends without loss of convergence.
Stable Diffusion 3 Medium Integration with Diffusers
Hugging Face has integrated Stable Diffusion 3 Medium (2B parameters) into the Diffusers library, introducing a Multimodal Diffusion Transformer (MMDiT) and rectified flow-matching for improved text-to-image synthesis.
Hugging Face TRL RLOO Trainer Release
Hugging Face has introduced the RLOO (REINFORCE Leave One-Out) Trainer in TRL, an online RLHF algorithm that uses 50-70% less vRAM and converges up to 3x faster than PPO while remaining competitive in performance.
Hugging Face Embedding Container for Amazon SageMaker Release
Hugging Face has released a general availability (GA) Embedding Container for Amazon SageMaker, powered by Text Embedding Inference (TEI) for high-performance deployment of open embedding models.
Hugging Face Transformers Documentation Redesign
Hugging Face is redesigning the Transformers documentation to move from a rigid, incremental structure to a code-first, integrated experience tailored for product developers.
Artificial Analysis Text to Image Leaderboard and Arena Launch
Hugging Face has launched the Artificial Analysis Text to Image Leaderboard and Arena, using human preference data from over 45,000 votes to rank open-source and proprietary image generation models.
Hugging Face NPC-Playground: Integrating LLM-Powered NPCs in 3D Environments
Hugging Face introduces NPC-Playground, a 3D demo utilizing Cubzh and Gigax to enable realistic, LLM-driven NPC interactions through function calling and Lua scripting.
Intel Gaudi Assisted Generation Support
Hugging Face has integrated assisted generation (speculative sampling) into Optimum Habana, enabling up to 2x speedups for large transformer-based models on Intel Gaudi processors.
Hugging Face Spaces Secrets Security Update
Hugging Face disclosed a security incident involving unauthorized access to a subset of Spaces secrets, leading to the revocation of affected tokens and a transition to fine-grained access tokens.
Benchmarking Text Generation Inference
Hugging Face introduces a benchmarking tool for Text Generation Inference (TGI) to help developers profile throughput and latency trade-offs to optimize LLM deployment costs and performance.
Training and Finetuning Embedding Models with Sentence Transformers
Hugging Face provides a comprehensive guide on using the Sentence Transformers library to train and finetune embedding models for tasks like semantic search and RAG, introducing the new SentenceTransformerTrainer.
Falcon 2 11B Release Notes
TII has released Falcon 2 11B, a series of smaller, high-performance models including a pretrained LLM and a Vision-Language Model (VLM) trained on over 5,000 billion tokens.
CyberSecEval 2: Evaluating Cybersecurity Risks in Large Language Models
Hugging Face and Meta have introduced CyberSecEval 2, a comprehensive framework to evaluate LLM susceptibility to code interpreter abuse, offensive cybersecurity capabilities, and prompt injection attacks.
Hugging Face AWS Inferentia2 Integration
Hugging Face has expanded support for AWS Inferentia2, enabling the deployment of over 100,000 models via Amazon SageMaker and introducing new Inf2 instance options for Hugging Face Inference Endpoints.
Hugging Face and Microsoft Strategic Collaboration Update May 2024
Hugging Face and Microsoft have expanded their strategic partnership to integrate popular open models into the Azure Model Catalog, enable AMD MI300X GPU support on Azure, and introduce a VS Code integration for Hugging Face Spaces.
Hugging Face Integration for AMD Instinct MI300 GPU
Hugging Face has enabled first-class integration for the AMD Instinct MI300 GPU, delivering up to 3x faster inference and 2x faster fine-tuning for Llama 3 70B compared to the MI250.
Dell Enterprise Hub: On-Premise AI Deployment and Training
Hugging Face and Dell have launched the Dell Enterprise Hub, a platform that simplifies the on-premise training and deployment of open-source large language models on Dell hardware.
Hugging Face Spaces Dev Mode Release
Hugging Face has introduced Spaces Dev Mode, allowing PRO subscribers to connect via VS Code or SSH for real-time code editing without requiring git pushes for every change.
Hugging Face KV Cache Quantization
Hugging Face has introduced KV Cache Quantization in the Transformers library to reduce memory usage for long-context text generation with minimal impact on model quality.
Introducing the Open Arabic LLM Leaderboard
Hugging Face and the Technology Innovation Institute have launched the Open Arabic LLM Leaderboard (OALL) to provide a specialized, open evaluation platform for Arabic Large Language Models.
PaliGemma Vision Language Model Release
Google has released PaliGemma, a family of open vision-language models combining SigLIP-So400m and Gemma-2B to enable tasks like image captioning, visual question answering, and object detection.
Hugging Face and LangChain Launch langchain_huggingface Partner Package
Hugging Face and LangChain have released langchain_huggingface, a jointly maintained partner package to streamline the integration of Hugging Face models into the LangChain ecosystem.
Transformers Agents 2.0 release notes / what's new
Hugging Face introduces Transformers Agents 2.0, a modular agent framework featuring iterative problem-solving agents and high performance on the GAIA benchmark using Llama-3-70B-Instruct.
Building Cost-Efficient Enterprise RAG Applications with Intel Gaudi 2 and Intel Xeon
Hugging Face demonstrates how combining Intel Gaudi 2 accelerators and Intel Xeon CPUs via the OPEA framework allows enterprises to build RAG applications with superior performance-per-dollar compared to H100-based systems.
Hugging Face Enterprise Hub now available on AWS Marketplace
Hugging Face has integrated Enterprise Hub with the AWS Marketplace, allowing organizations to upgrade their accounts and manage billing directly through their AWS accounts.
Hugging Face Open Leaderboard for Hebrew LLMs
Hugging Face has launched an open leaderboard specifically designed to evaluate and improve Large Language Models (LLMs) in Hebrew, addressing the challenges of low-resource and morphologically complex languages.
Artificial Analysis LLM Performance Leaderboard on Hugging Face
Hugging Face has integrated the Artificial Analysis LLM Performance Leaderboard, providing AI engineers with a unified metric system for comparing the quality, price, and speed of over 100 serverless LLM API endpoints.
Hugging Face Inference Endpoints ASR and Diarization Pipeline
Hugging Face introduces a custom inference handler for deploying a modular pipeline combining Whisper ASR, Pyannote diarization, and speculative decoding on Inference Endpoints.
Improving Prompt Consistency with Structured Generations
Hugging Face and Dottxt research demonstrates that using structured generation to constrain LLM outputs reduces performance variance and improves ranking consistency across different prompt formats and shot orders.
StarCoder2-15B-Instruct-v0.1 release notes / what's new
Hugging Face introduces StarCoder2-15B-Instruct-v0.1, the first entirely self-aligned code LLM trained with a fully transparent and permissive pipeline that outperforms CodeLlama-70B-Instruct on HumanEval.
Hugging Face Open Chain of Thought Leaderboard
Hugging Face has introduced the Open Chain of Thought Leaderboard to measure the specific accuracy gain provided by chain-of-thought prompting across various LLMs on challenging reasoning tasks.
Jack of All Trades (JAT) Multi-Purpose Transformer Agent
Hugging Face introduces Jack of All Trades (JAT), a single transformer-based agent capable of performing diverse sequential decision-making tasks across Atari, BabyAI, Meta-World, and MuJoCo environments.
The Open Medical-LLM Leaderboard: Benchmarking Large Language Models in Healthcare
Hugging Face has introduced the Open Medical-LLM Leaderboard, a standardized platform to evaluate and compare the performance of LLMs across diverse medical datasets to improve reliability and patient safety.
Meta Llama 3 Release Notes
Meta has released Llama 3, an open-access LLM family featuring 8B and 70B parameter models with improved tokenization and training on 15 trillion tokens.
Ryght Case Study: Building a Life Sciences Generative AI Platform with Hugging Face
Ryght has launched Ryght Preview, an enterprise-grade generative AI platform for healthcare and life sciences that leverages Hugging Face's Expert Support, TGI, and TEI to provide secure, flexible, and high-performance AI copilots.
Running Privacy-Preserving Inferences on Hugging Face Endpoints
Hugging Face and Zama have enabled the deployment of Fully Homomorphic Encryption (FHE) models via Hugging Face Endpoints, allowing users to perform machine learning inferences on encrypted data without decrypting it.
LiveCodeBench Leaderboard: Contamination-Free Evaluation for Code LLMs
Hugging Face has introduced the LiveCodeBench leaderboard, a new benchmark developed by researchers from UC Berkeley, MIT, and Cornell to evaluate LLM code generation and reasoning capabilities while preventing benchmark contamination using time-windowed problem sets.
Gradio Reload Mode for Faster AI App Development
Gradio's reload mode enables developers to apply source code changes to AI applications instantly without restarting the server, significantly reducing development latency.
Idefics2 8B Vision-Language Model Release – Architecture, Data, and Performance
Hugging Face released Idefics2, an 8B open‑source vision‑language model that outperforms other 8‑B models on VQA and OCR benchmarks and is ready for fine‑tuning via 🤗 Transformers.