argilla-io/distilabel

Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.

What it solves

Distilabel provides a programmatic framework for generating synthetic data and AI feedback to improve the quality of AI models. It addresses the challenge of creating high-quality, diverse datasets for fine-tuning LLMs, predictive NLP tasks (like classification and extraction), and generative scenarios (such as instruction following and dialogue generation) without relying solely on expensive human labeling.

How it works

The framework allows engineers to build scalable, fault-tolerant pipelines that synthesize and judge data based on verified research methodologies. It uses a unified API to integrate AI feedback from any LLM provider (e.g., OpenAI, Anthropic, Cohere, Mistral, Vertex AI, and local models via vLLM or Ollama). It also supports structured generation via tools like Outlines and Instructor, and can scale distributed pipelines using Ray.

Who it’s for

AI engineers and researchers who need to create large-scale synthetic datasets for training or refining LLMs and other NLP models, specifically those focusing on data quality to improve model performance.

Highlights

  • Broad LLM Integration: Supports a wide array of providers including proprietary APIs and local inference engines.
  • Scalable Pipelines: Built-in support for Ray to distribute data generation and AI feedback tasks.
  • Scalable Synthesis: Capable of synthesizing data on an immense scale, as demonstrated by the 1M OpenHermesPreference dataset.
  • Research-Driven: Designed to implement methodologies from verified research papers for generating and judging data.
  • Data Filtering: Enables improving model performance by using AI feedback to filter out low-quality data from existing datasets.

Related

  • Project
  • Project
  • Project
  • Project
  • Project