promptfoo/promptfoo

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.

What it solves

It eliminates the trial-and-error approach to developing LLM applications by providing a systematic way to evaluate, test, and secure prompts and models. It helps developers ensure their AI apps are reliable, secure, and compliant before shipping to production.

How it works

Promptfoo is a CLI and library that allows developers to run automated evaluations and red-teaming exercises. It enables side-by-side comparison of different LLM providers (such as OpenAI, Anthropic, Azure, Bedrock, and Ollama) and integrates into CI/CD pipelines for automated checks. It also includes code scanning to review pull requests for security and compliance issues.

Who it’s for

It is designed for developers building LLM-powered applications who need to move from gut-feel decisions to data-driven metrics for prompt engineering and security testing.

Highlights

  • Automated Evaluations: Test prompts and models with a structured approach.
  • Red Teaming: Scan for vulnerabilities to secure LLM applications.
  • Model Comparison: Compare multiple LLM providers side-by-side.
  • CI/CD Integration: Automate quality and security checks in the development pipeline.
  • Local Execution: Evals run locally to keep prompts private.
  • Code Scanning: Review pull requests for LLM-related security and compliance.

Related

  • Project
  • Project
  • Project
  • Project
  • Project