promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
What it solves
它取代了開發 LLM 應用程式時的試錯方式,提供系統化的方式來評估提示詞效能,並在上線前識別安全漏洞。
How it works
Promptfoo 是一個 CLI 與函式庫,讓開發者能執行自動化評估與紅隊測試。它支援不同提示詞與模型(如 OpenAI、Anthropic、Azure、Bedrock、Ollama)的並排比較,且可整合至 CI/CD 流程,進行自動檢查與程式碼掃描,以偵測安全與合規問題。
Who it’s for
建構 LLM 驅動應用程式的開發者,需要確保 AI 應用在模型選擇與提示詞工程上具備安全、可靠且以資料為導向的特性。
Highlights
- Automated Evaluations: 系統化測試提示詞與模型。
- Red Teaming: 掃描漏洞以保護 LLM 應用。
- Model Comparison: 並排比較多家 LLM 供應商。
- Developer-First: 包含即時重新載入、快取與本地執行等隱私保護功能。
- CI/CD Integration: 透過程式碼掃描自動化檢查與 Pull Request 審核。