Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
At a glance
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
People who want an AI system to plan or carry out multi-step tasks.
The repository reports a MIT license. Setup difficulty and device support still need to be checked upstream.
Evidence & freshness
This record comes from the wider automated directory. It has not yet passed the stronger catalog verification checks.
Strong automated signals across recent activity, community, and project upkeep.
The health score is an automated maintenance signal, not a security audit or a guarantee that the tool fits your needs.
Embed a live health badge in a README or docs page.
[](https://ai-tools-scout.com/projects/promptfoo)Suggested automatically from the same category and shared topics, then ordered by current maintenance and adoption signals.
Get the fastest-growing projects, useful MCP servers, and technical reads in one weekly email.
Catalog-verified records use a timestamped GitHub snapshot that passed source and completeness checks. Basic listings come from the wider automated sync and have not yet passed that stronger verification path.
Health combines recent activity, commit frequency, issue handling, contributor depth, release cadence, documentation, and community signals. Missing evidence lowers how much confidence you should place in the number.