Hugging Face Transformers GGUF Support
Hugging Face has integrated support for GGUF (GPT-Generated Unified Format) models into the transformers library, allowing users to load and run quantized checkpoints—originally designed for llama.cpp—using familiar PyTorch and Transformers APIs. This integration enables efficient local inference on Apple Silicon by leveraging underlying ggml kernels to maintain performance close to that of dedicated local inference engines.
GGUF Format and Quantization
GGUF is a file format developed by the llama.cpp team that packages model weights, tokenizer information, and optional chat templates into a single file. Its primary advantage is support for various quantization levels, which allow users to reduce a model's memory footprint by trading off some precision.
For example, using Unsloth's Qwen3.5-4B model, the memory requirements vary by quantization variant:
| GGUF variant | File size | Tradeoff |
|---|---|---|
BF16 |
8.42 GB | Unquantized reference |
Q6_K |
3.53 GB | High precision |
Q5_K_M |
3.14 GB | Balanced size and precision |
Q4_K_M |
2.74 GB | Practical starting point for local inference |
Technical Implementation: ggml Kernels and Generation Loop
To achieve performance levels comparable to llama.cpp, Hugging Face implemented two primary technical optimizations: the reuse of ggml kernels and the streamlining of the generate loop.
Integration of ggml Metal Kernels
Rather than replacing the model with a separate runtime, transformers now uses the kernels library to call compatible ggml Metal kernels directly from PyTorch. This allows the model to remain in Python while the heavy computation is handled by specialized GPU programs:
ggml-quantization: Reads packed quantized weights for matrix operations, avoiding the need to expand the weight matrix before each decode operation.ggml-norm: Fuses normalization operations, including zero-centered RMSNorm used in Qwen3.5 and Qwen3.8.ggml-attn: Implementsggml's Metal flash attention for both prompt processing and token decoding.ggml-gated-delta-net: Accelerates linear-attention layers in Qwen3.5 and Qwen3.8 hybrid architectures.topk: A custom Metal implementation to optimize expert selection in Mixture-of-Experts (MoE) models.
Generation Loop Optimizations
Hugging Face reduced CPU-GPU synchronization overhead in the generate function to ensure the GPU remains saturated. Two key changes were introduced:
- Attention Mask Optimization: For decoder-only inputs without padding, the all-ones padding mask is removed at the start of generation, preventing the attention code from repeatedly inspecting the mask.
- Deferred Stopping Checks: The stopping decision is copied asynchronously, allowing the CPU to continue scheduling work while the GPU executes the current step.
Performance and Benchmarking
In benchmarks conducted on a MacBook Pro M2 Max (32 GB unified memory), transformers demonstrated throughput close to llama.cpp across small dense models, larger dense models, and MoE models. Notably, the transformers measurements included prefill, whereas the llama-bench results for llama.cpp focused on decode-only throughput.
Usage and Deployment
Loading GGUF Models
Users can load GGUF models from the Hugging Face Hub using the gguf_file parameter in from_pretrained:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "unsloth/Qwen3.5-4B-GGUF"
filename = "Qwen3.5-4B-Q4_K_M.gguf"
tokenizer = AutoTokenizer.from_pretrained(model_id, gguf_file=filename)
model = AutoModelForCausalLM.from_pretrained(model_id, gguf_file=filename)
Serving via OpenAI-Compatible API
Models can be served using the transformers serve CLI, which exposes an OpenAI-compatible API for integration with clients like Jan or Pi:
pip install -U "transformers[serving] @ git+https://github.com/huggingface/transformers.git" kernels
transformers serve "unsloth/Qwen3.5-4B-GGUF:Qwen3.5-4B-Q4_K_M.gguf"
Strategic Implications and Limitations
While llama.cpp remains the recommended engine for pure local inference efficiency, this integration allows developers to use GGUF checkpoints within the PyTorch ecosystem for tasks such as:
- Prototyping: Inspecting intermediate activations with hooks or modifying forward passes.
- Evaluation: Using existing
transformersworkflows to measure quantized model quality. - Validation: Comparing original checkpoints against GGUF conversions to check for quantization error.
- Fine-tuning: Dequantizing weights via
GgufConfig(dequantize=True)to continue standard training workflows.
Current Limitations
- Hardware: The packed inference path is currently limited to MPS (Apple Silicon).
- Batching: Padded batches currently have lower performance; optimization for
generate_batchon MPS is ongoing. - Architecture: Initial support is limited to Qwen3.5 dense and MoE architectures (including Qwen3.8).
Future Outlook
Hugging Face aims to bring ggml's performance to architectures not currently supported by llama.cpp. By integrating ggml kernels into PyTorch, transformers can accelerate new research models and custom variants without requiring a full llama.cpp implementation for every new architecture.