PacificAI/langtest

Deliver safe & effective language models

What it solves

LangTest is designed to help data scientists ensure that language models are safe, robust, and fair. It addresses the lack of-the-box tools for evaluating model quality beyond simple accuracy, specifically targeting issues like demographic bias, toxicity, and inconsistency in real-world conditions.

How it works

The library provides a Harness object that allows users to generate, execute, and report on over 60 distinct types of tests. It supports a wide range of NLP tasks (NER, Translation, Text Classification) and integrates with popular frameworks like Hugging Face, Spark NLP, and various LLM providers (OpenAI, Cohere, AI21, Azure-OpenAI).

Who it’s for

It is built for NLP practitioners and data scientists who need to rigorously test their models for robustness, fairness, and safety before deploying them into production systems.

Highlights

  • Comprehensive Testing: Covers robustness, bias, representation, fairness, and accuracy.
  • Automated Data Augmentation: Can automatically augment training data based on test results for certain models.
  • LLM Support: Specialized tests for LLMs including factuality, sycophancy, and toxicity.
  • Broad Integration: Works with multiple NLP frameworks and LLM APIs.

Related

  • Project
  • Project
  • Project
  • Project
  • Project