Hugging Face Transformers GGUF Support

Hugging Face has integrated support for GGUF (GPT-Generated Unified Format) models into the transformers library, allowing users to load and run quantized checkpoints—originally designed for llama.cpp—using familiar PyTorch and Transformers APIs. This integration enables efficient local inference on Apple Silicon by leveraging underlying ggml kernels to maintain performance close to that of dedicated local inference engines.

GGUF Format and Quantization

GGUF is a file format developed by the llama.cpp team that packages model weights, tokenizer information, and optional chat templates into a single file. Its primary advantage is support for various quantization levels, which allow users to reduce a model's memory footprint by trading off some precision.

For example, using Unsloth's Qwen3.5-4B model, the memory requirements vary by quantization variant:

GGUF variant File size Tradeoff
BF16 8.42 GB Unquantized reference
Q6_K 3.53 GB High precision
Q5_K_M 3.14 GB Balanced size and precision
Q4_K_M 2.74 GB Practical starting point for local inference

Technical Implementation: ggml Kernels and Generation Loop

To achieve performance levels comparable to llama.cpp, Hugging Face implemented two primary technical optimizations: the reuse of ggml kernels and the streamlining of the generate loop.

Integration of ggml Metal Kernels

Rather than replacing the model with a separate runtime, transformers now uses the kernels library to call compatible ggml Metal kernels directly from PyTorch. This allows the model to remain in Python while the heavy computation is handled by specialized GPU programs:

  • ggml-quantization: Reads packed quantized weights for matrix operations, avoiding the need to expand the weight matrix before each decode operation.
  • ggml-norm: Fuses normalization operations, including zero-centered RMSNorm used in Qwen3.5 and Qwen3.8.
  • ggml-attn: Implements ggml's Metal flash attention for both prompt processing and token decoding.
  • ggml-gated-delta-net: Accelerates linear-attention layers in Qwen3.5 and Qwen3.8 hybrid architectures.
  • topk: A custom Metal implementation to optimize expert selection in Mixture-of-Experts (MoE) models.

Generation Loop Optimizations

Hugging Face reduced CPU-GPU synchronization overhead in the generate function to ensure the GPU remains saturated. Two key changes were introduced:

  1. Attention Mask Optimization: For decoder-only inputs without padding, the all-ones padding mask is removed at the start of generation, preventing the attention code from repeatedly inspecting the mask.
  2. Deferred Stopping Checks: The stopping decision is copied asynchronously, allowing the CPU to continue scheduling work while the GPU executes the current step.

Performance and Benchmarking

In benchmarks conducted on a MacBook Pro M2 Max (32 GB unified memory), transformers demonstrated throughput close to llama.cpp across small dense models, larger dense models, and MoE models. Notably, the transformers measurements included prefill, whereas the llama-bench results for llama.cpp focused on decode-only throughput.

Usage and Deployment

Loading GGUF Models

Users can load GGUF models from the Hugging Face Hub using the gguf_file parameter in from_pretrained:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "unsloth/Qwen3.5-4B-GGUF"
filename = "Qwen3.5-4B-Q4_K_M.gguf"

tokenizer = AutoTokenizer.from_pretrained(model_id, gguf_file=filename)
model = AutoModelForCausalLM.from_pretrained(model_id, gguf_file=filename)

Serving via OpenAI-Compatible API

Models can be served using the transformers serve CLI, which exposes an OpenAI-compatible API for integration with clients like Jan or Pi:

pip install -U "transformers[serving] @ git+https://github.com/huggingface/transformers.git" kernels

transformers serve "unsloth/Qwen3.5-4B-GGUF:Qwen3.5-4B-Q4_K_M.gguf"

Strategic Implications and Limitations

While llama.cpp remains the recommended engine for pure local inference efficiency, this integration allows developers to use GGUF checkpoints within the PyTorch ecosystem for tasks such as:

  • Prototyping: Inspecting intermediate activations with hooks or modifying forward passes.
  • Evaluation: Using existing transformers workflows to measure quantized model quality.
  • Validation: Comparing original checkpoints against GGUF conversions to check for quantization error.
  • Fine-tuning: Dequantizing weights via GgufConfig(dequantize=True) to continue standard training workflows.

Current Limitations

  • Hardware: The packed inference path is currently limited to MPS (Apple Silicon).
  • Batching: Padded batches currently have lower performance; optimization for generate_batch on MPS is ongoing.
  • Architecture: Initial support is limited to Qwen3.5 dense and MoE architectures (including Qwen3.8).

Future Outlook

Hugging Face aims to bring ggml's performance to architectures not currently supported by llama.cpp. By integrating ggml kernels into PyTorch, transformers can accelerate new research models and custom variants without requiring a full llama.cpp implementation for every new architecture.

Sources