Projects in Awesome Lists tagged with prompt-testing
A curated list of projects in awesome lists tagged with prompt-testing .
https://github.com/linshenkx/prompt-optimizer
An AI prompt optimizer for writing better prompts and getting better AI results.
ai-prompts ai-tools llm prompt prompt-engineering prompt-optimization prompt-optimizer prompt-testing prompt-toolkit prompt-tuning
Last synced: 25 Apr 2026
https://github.com/promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming, pentesting, and vulnerability scanning for LLMs. Compare performance of GPT, Claude, Gemini, Llama, and more. Simple declarative configs with command line and CI/CD integration.
ci ci-cd cicd evaluation evaluation-framework llm llm-eval llm-evaluation llm-evaluation-framework llmops pentesting prompt-engineering prompt-testing prompts rag red-teaming testing vulnerability-scanners
Last synced: 03 Mar 2026
https://github.com/msoedov/agentic_security
Agentic LLM Vulnerability Scanner / AI red teaming kit 🧪
agent-framework agent-security ai-red-team llm-evaluation llm-evaluation-framework llm-fuzzer llm-fuzzer-aggregator llm-fuzzing llm-guardrails llm-jailbreaks llm-scanner llm-security llm-vulnerabilities prompt-testing
Last synced: 02 Sep 2026
https://github.com/lizhiyao/oh-my-knowledge
Evaluation framework for LLM knowledge inputs — prompts, RAG corpora, skills, agent workflows. Fix the model, vary the artifact. Built-in statistical rigor: bootstrap CI, Krippendorff α, length-debias, saturation curves.
agent-evaluation ai benchmark bootstrap-ci claude claude-code evaluation-as-code evaluation-framework knowledge-engineering krippendorff-alpha llm llm-evaluation llm-judge multi-judge-ensemble prompt-engineering prompt-testing rag-evaluation skill-evaluation
Last synced: 15 Jun 2026
https://github.com/jhd3197/prompture
Prompture is an API-first library for requesting structured JSON output from LLMs (or any structure), validating it against a schema, and running comparative tests between models.
ai-testing json-validation llm openai prompt-engineering prompt-testing prompture pydantic structured-output toon
Last synced: 24 May 2026
https://github.com/prompt-foundry/typescript-sdk
The prompt engineering, prompt management, and prompt evaluation tool for TypeScript, JavaScript, and NodeJS.
gpt gpt-3 gpt-4 llm llm-eval llm-evaluation llm-ops llm-test llmops open-ai prompt-engineering prompt-evaluation prompt-management prompt-manager prompt-testing typescript
Last synced: 16 Jul 2026
https://github.com/syamsasi99/prompt-evaluator
prompt-evaluator is an open-source toolkit for evaluating, testing, and comparing LLM prompts. It provides a GUI-driven workflow for running prompt tests, tracking token usage, visualizing results, and ensuring reliability across models like OpenAI, Claude, and Gemini.
ai-evaluation ai-evaluation-framework ai-evaluation-metrics ai-evaluation-tools datascience developer-tools electron llm prompt-engineering prompt-testing promptfoo react typescript
Last synced: 13 Jan 2026
https://github.com/yukinagae/genkitx-promptfoo
Community Plugin for Genkit to use Promptfoo
ai evaluation evaluation-framework firebase genkit genkit-plugin genkitx llm llm-eval llm-evaluation llm-evaluation-framework llmops plugin prompt prompt-testing promptfoo prompts testing
Last synced: 27 Jul 2025
https://github.com/gaoyechen/sharpinput
AI最强嘴替-智能输入优化框架,让你的任何输入(提问、陈述、方案、想法、需求)都被打磨成能让 AI 输出有洞察力观点的高质量版本。
agent ai ai-tools chatgpt claude input-optimization llm openai prompt prompt-engineering prompt-enhancement prompt-optimizer prompt-testing prompt-tuning skill
Last synced: 06 Jul 2026
https://github.com/homemade-software-inc/completion-kit
Your prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and compare runs to see what got better.
anthropic evaluation-framework evaluation-metrics llm llm-as-judge llm-eval llm-evaluation llm-evaluation-framework llm-evaluation-metrics llmops mcp ollama openai prompt-engineering prompt-testing rails rails-engine ruby ruby-on-rails
Last synced: 09 Jul 2026
https://github.com/radoslaw-sz/maia
A pytest-based framework for testing multi AI agents systems. It provides a flexible and extensible platform for complex multi-agent simulations. Supports many integrations like LiteLLM, CrewAI, LangChain etc.
agentic agents ai ai-testing ai-testing-tool framework llm maia prompt-engineering prompt-testing python test
Last synced: 25 Sep 2025
https://github.com/tolom/prompt-xray-skills
Reusable SKILL.md workflows for writing, auditing, testing, and hardening LLM prompts.
agentic-workflows ai-agents llm prompt-audit prompt-engineering prompt-testing rag skills structured-outputs system-prompts
Last synced: 10 Jul 2026
https://github.com/yukinagae/promptfoo-sample
Sample project demonstrates how to use Promptfoo, a test framework for evaluating the output of generative AI models
evaluation evaluation-framework llm llm-eval llm-evaluation llm-evaluation-framework llmops prompt-testing promptfoo prompts testing
Last synced: 25 Feb 2026
https://github.com/piyushgupta344/llm-test-harness
Deterministic testing framework for LLM-powered apps — record/replay cassettes, eval scoring, regression testing
anthropic cassette deterministic eval llm openai prompt-testing testing vcr
Last synced: 03 Apr 2026
https://github.com/sigmakib2/openai-prompt-testing-playground
A dynamic and interactive playground for testing and refining prompts with OpenAI's language models. Includes customizable inputs for prompts, advanced model settings, and live response streaming for seamless experimentation.
ai chatgpt openai playground prompt prompt-engineering prompt-testing
Last synced: 21 May 2026
https://github.com/taimoorkhan10/replayd
Turn failed AI agent runs into replayable regression tests. Catch regressions before you ship.
agent-ops agent-testing ai-agents ai-infrastructure ai-reliability llm-ops llm-testing open-source prompt-testing python regression-testing release-control replay-testing sdk
Last synced: 14 Jun 2026
https://github.com/yukinagae/genkit-promptfoo-sample
Sample implementation demonstrating how to use Firebase Genkit with Promptfoo
evaluation evaluation-framework genkit llm llm-eval llm-evaluation llm-evaluation-framework llmops prompt-testing promptfoo prompts testing
Last synced: 15 Aug 2025
https://github.com/ac12644/prompt-diff
The regression-test gate for AI-generated prompts. Auto-improve prompts and verify the rewrite against a YAML test suite, or diff any two prompts. Supports OpenAI, Anthropic, Google Gemini.
anthropic claude cli evaluation gemini llm openai prompt-engineering prompt-optimization prompt-testing regression-testing typescript
Last synced: 13 Aug 2026