Hugging Face Open Leaderboard for Hebrew LLMs
Hugging Face has launched an open leaderboard specifically designed to evaluate and improve Large Language Models (LLMs) in Hebrew, addressing the challenges of low-resource and morphologically complex languages.
Artificial Analysis LLM Performance Leaderboard on Hugging Face
Hugging Face has integrated the Artificial Analysis LLM Performance Leaderboard, providing AI engineers with a unified metric system for comparing the quality, price, and speed of over 100 serverless LLM API endpoints.
Hugging Face Inference Endpoints ASR and Diarization Pipeline
Hugging Face introduces a custom inference handler for deploying a modular pipeline combining Whisper ASR, Pyannote diarization, and speculative decoding on Inference Endpoints.
Improving Prompt Consistency with Structured Generations
Hugging Face and Dottxt research demonstrates that using structured generation to constrain LLM outputs reduces performance variance and improves ranking consistency across different prompt formats and shot orders.
StarCoder2-15B-Instruct-v0.1 release notes / what's new
Hugging Face introduces StarCoder2-15B-Instruct-v0.1, the first entirely self-aligned code LLM trained with a fully transparent and permissive pipeline that outperforms CodeLlama-70B-Instruct on HumanEval.
Hugging Face Open Chain of Thought Leaderboard
Hugging Face has introduced the Open Chain of Thought Leaderboard to measure the specific accuracy gain provided by chain-of-thought prompting across various LLMs on challenging reasoning tasks.
Jack of All Trades (JAT) Multi-Purpose Transformer Agent
Hugging Face introduces Jack of All Trades (JAT), a single transformer-based agent capable of performing diverse sequential decision-making tasks across Atari, BabyAI, Meta-World, and MuJoCo environments.
The Open Medical-LLM Leaderboard: Benchmarking Large Language Models in Healthcare
Hugging Face has introduced the Open Medical-LLM Leaderboard, a standardized platform to evaluate and compare the performance of LLMs across diverse medical datasets to improve reliability and patient safety.
Meta Llama 3 Release Notes
Meta has released Llama 3, an open-access LLM family featuring 8B and 70B parameter models with improved tokenization and training on 15 trillion tokens.
Ryght Case Study: Building a Life Sciences Generative AI Platform with Hugging Face
Ryght has launched Ryght Preview, an enterprise-grade generative AI platform for healthcare and life sciences that leverages Hugging Face's Expert Support, TGI, and TEI to provide secure, flexible, and high-performance AI copilots.
Running Privacy-Preserving Inferences on Hugging Face Endpoints
Hugging Face and Zama have enabled the deployment of Fully Homomorphic Encryption (FHE) models via Hugging Face Endpoints, allowing users to perform machine learning inferences on encrypted data without decrypting it.
LiveCodeBench Leaderboard: Contamination-Free Evaluation for Code LLMs
Hugging Face has introduced the LiveCodeBench leaderboard, a new benchmark developed by researchers from UC Berkeley, MIT, and Cornell to evaluate LLM code generation and reasoning capabilities while preventing benchmark contamination using time-windowed problem sets.
Gradio Reload Mode for Faster AI App Development
Gradio's reload mode enables developers to apply source code changes to AI applications instantly without restarting the server, significantly reducing development latency.
Idefics2 8B Vision-Language Model Release – Architecture, Data, and Performance
Hugging Face released Idefics2, an 8B open‑source vision‑language model that outperforms other 8‑B models on VQA and OCR benchmarks and is ready for fine‑tuning via 🤗 Transformers.
Vision Language Models Explained
Hugging Face provides a comprehensive guide to Vision Language Models (VLMs), detailing their architecture, open-source options, evaluation benchmarks, and new experimental support for fine-tuning via the TRL library.
Hugging Face and Google Cloud Vertex AI Model Garden Integration
Hugging Face has launched 'Deploy on Google Cloud,' enabling users to deploy thousands of open foundation models to Vertex AI or Google Kubernetes Engine (GKE) via the Hugging Face Hub or Vertex Model Garden.
CodeGemma Release Notes / What's New
Google has released CodeGemma, a family of open-access code-specialist LLMs based on Gemma, trained on 500 billion additional tokens of code, mathematics, and English language data.
Hugging Face Public Policy Program Overview and Submitted Materials
Hugging Face announced a comprehensive public‑policy program that provides U.S., EU, and U.K. policymakers with detailed position papers, testimony, and comment letters, reflecting its cross‑functional commitment to responsible openness and shaping AI regulation.
Hugging Face and Wiz Research Partnership for AI Security
Hugging Face has partnered with Wiz to integrate advanced vulnerability management and cloud security posture management to protect its platform and the broader AI/ML ecosystem.
Text2SQL with Hugging Face Dataset Viewer API and DuckDB-NSQL-7B
Hugging Face demonstrates how to use the DuckDB-NSQL-7B model and the Dataset Viewer API to convert natural language questions into SQL queries for analyzing over 120,000 open datasets.
SetFit Inference Acceleration with 🤗 Optimum Intel on Xeon
Hugging Face demonstrates how to achieve up to 7.8x faster inference throughput for SetFit models on Intel Xeon CPUs using post-training static quantization via the 🤗 Optimum Intel library.
Hugging Face and Cloudflare Workers AI Integration
Hugging Face integrated Cloudflare Workers AI to provide serverless GPU inference for popular open models, allowing developers to deploy AI applications with a pay-per-request pricing model.
Pollen-Vision: Unified Interface for Zero-Shot Vision Models in Robotics
Hugging Face and the Pollen Robotics team have released pollen-vision, an open-source library that integrates zero-shot vision models to enable robots to detect and localize unknown objects in 3D space.
Hugging Face Transformers: A Beginner's Guide to Open-Source ML
Hugging Face provides a comprehensive introductory guide to using the Transformers library and Hub to deploy and run open-source machine learning models like Microsoft's Phi-2.
Embedding Quantization: Binary and Scalar Techniques for Faster, Cheaper Retrieval
Hugging Face announced binary and int8 embedding quantization, cutting memory by 32× or 4× and speeding up retrieval up to 45× while keeping 96%–99% of original performance.
Hugging Face and Lighthouz AI Introduce Chatbot Guardrails Arena
Hugging Face and Lighthouz AI have launched the Chatbot Guardrails Arena, a community-driven stress-testing platform designed to evaluate the data privacy and security of LLMs and their guardrails.
GaLore: Advancing Large Model Training on Consumer-grade Hardware
GaLore enables the training of billion-parameter models on consumer-grade GPUs by reducing optimizer state memory requirements by over 82.5% through low-rank gradient projection.
Cosmopedia: Large-Scale Synthetic Data for LLM Pre-training
Hugging Face introduces Cosmopedia, the largest open synthetic dataset for LLM pre-training, containing 30 million files and 25 billion tokens generated by Mixtral-8x7B-Instruct-v0.1.
Phi-2 on Intel Meteor Lake: Local LLM Inference
Hugging Face demonstrates how to run the Microsoft Phi-2 model locally on Intel Meteor Lake (Core Ultra) processors using 4-bit quantization via OpenVINO and Optimum Intel.
Quanto: a PyTorch quantization backend for Optimum
Hugging Face introduces Quanto, a versatile and device-agnostic PyTorch quantization backend for Optimum designed to simplify low-precision model deployment across any modality.
Hugging Face Train on DGX Cloud
Hugging Face launched Train on DGX Cloud, a no-code service for Enterprise Hub organizations to fine-tune open models using NVIDIA H100 and L40S GPUs.
WebSight Dataset and Sightseer Model
Hugging Face introduced WebSight, a synthetic dataset of 2 million screenshot-to-HTML pairs, and Sightseer, a vision-language model capable of converting web screenshots into functional HTML code.
CPU Optimized Embeddings with Optimum Intel and fastRAG
Hugging Face and Intel have introduced a method to accelerate embedding models on Xeon CPUs using Optimum Intel and fastRAG, achieving up to 4.5x latency reduction and 4x throughput improvement via int8 quantization.
ConTextual: Benchmarking Multimodal Reasoning in Text-Rich Scenes
Hugging Face and UCLA researchers have introduced ConTextual, a dataset and leaderboard designed to evaluate how Large Multimodal Models (LMMs) jointly reason over text and visual cues in complex, text-rich images.
Hugging Face and Argilla Enable Collective Community Dataset Building
Hugging Face and Argilla have introduced a streamlined workflow using Hugging Face Spaces and Argilla to allow communities to collectively build high-quality, open-source datasets.
Text-Generation Pipeline on Intel Gaudi 2 AI Accelerator
Hugging Face introduces a custom text-generation pipeline for Intel Gaudi 2 AI accelerators, enabling streamlined deployment of Llama 2 models (7b, 13b, and 70b) via Optimum Habana.
StarCoder2 and The Stack v2 Release
BigCode has released StarCoder2, a family of transparently trained open code LLMs in 3B, 7B, and 15B parameter sizes, powered by the massive new Stack v2 dataset.
Hugging Face TTS Arena: Benchmarking Text-to-Speech Models
Hugging Face has launched the TTS Arena, a crowdsourced, side-by-side comparison tool and leaderboard using an Elo rating system to objectively measure text-to-speech model quality.
AI Watermarking 101: Tools and Techniques
Hugging Face provides a comprehensive overview of AI watermarking techniques across images, text, and audio to combat deepfakes and ensure content provenance.
Introduction to Matryoshka Embedding Models
Matryoshka Embedding models allow for variable-size embeddings that can be truncated without significant performance loss, enabling a flexible trade-off between storage, speed, and accuracy.
Fine-Tuning Gemma Models in Hugging Face
Hugging Face provides a guide on using Parameter-Efficient Fine-Tuning (PEFT) and Low-Rank Adaptation (LoRA) to customize Google's Gemma models on GPUs and Cloud TPUs.
Hugging Face and Haize Labs Introduce Red-Teaming Resistance Leaderboard
Hugging Face and Haize Labs have launched the Red-Teaming Resistance (RTR) Benchmark to evaluate LLM robustness against high-quality, human-like adversarial prompts across specific safety violation categories.
Google Gemma Open LLM Release
Google has released Gemma, a family of open-access large language models based on Gemini, available in 2B and 7B parameter sizes with base and instruction-tuned variants.
Open Ko-LLM Leaderboard
Hugging Face and Upstage have launched the Open Ko-LLM Leaderboard to provide a fair, transparent evaluation ecosystem for Korean Large Language Models using private test sets to prevent contamination.
Hugging Face PEFT New LoRA Merging Methods
Hugging Face has introduced several new merging methods to the PEFT library, enabling users to combine multiple LoRA adapters from the same base model on the fly to synthesize new capabilities.
Synthetic Data with Open-Source LLMs Cuts Cost, Latency, and Carbon for Custom Models
Hugging Face shows how using open‑source LLMs to generate synthetic data and fine‑tune a small RoBERTa model reduces inference cost from $3061 to $2.7, latency from seconds to 0.13 s, and CO₂ emissions from ~1 t to 0.12 kg while matching GPT‑4 accuracy on financial sentiment classification.
AMD Pervasive AI Developer Contest
AMD and Hugging Face have partnered to launch the Pervasive AI Developer Contest, offering developers free access to AMD hardware and cash prizes to build AI applications in Generative AI, Robotics AI, and PC AI.
Hugging Face TGI Messages API Release
Hugging Face has introduced a Messages API for Text Generation Inference (TGI) starting with version 1.4.0, enabling OpenAI Chat Completion API compatibility for open LLMs.
SegMoE: Segmind Mixture of Diffusion Experts
SegMoE is a framework for creating Mixture-of-Experts (MoE) Diffusion models by replacing Feed-Forward layers in Stable Diffusion architectures with sparse MoE layers to improve prompt understanding.
NPHardEval Leaderboard: Evaluating LLM Reasoning via Computational Complexity
Hugging Face introduces the NPHardEval leaderboard, a dynamic benchmark that uses computational complexity classes to quantitatively measure the logical reasoning abilities of Large Language Models.