An open API service indexing awesome lists of open source software.

Projects in Awesome Lists tagged with prompt-testing

A curated list of projects in awesome lists tagged with prompt-testing .

https://github.com/linshenkx/prompt-optimizer

An AI prompt optimizer for writing better prompts and getting better AI results.

ai-prompts ai-tools llm prompt prompt-engineering prompt-optimization prompt-optimizer prompt-testing prompt-toolkit prompt-tuning

Last synced: 25 Apr 2026

https://github.com/promptfoo/promptfoo

Test your prompts, agents, and RAGs. Red teaming, pentesting, and vulnerability scanning for LLMs. Compare performance of GPT, Claude, Gemini, Llama, and more. Simple declarative configs with command line and CI/CD integration.

ci ci-cd cicd evaluation evaluation-framework llm llm-eval llm-evaluation llm-evaluation-framework llmops pentesting prompt-engineering prompt-testing prompts rag red-teaming testing vulnerability-scanners

Last synced: 03 Mar 2026

https://github.com/lizhiyao/oh-my-knowledge

Evaluation framework for LLM knowledge inputs — prompts, RAG corpora, skills, agent workflows. Fix the model, vary the artifact. Built-in statistical rigor: bootstrap CI, Krippendorff α, length-debias, saturation curves.

agent-evaluation ai benchmark bootstrap-ci claude claude-code evaluation-as-code evaluation-framework knowledge-engineering krippendorff-alpha llm llm-evaluation llm-judge multi-judge-ensemble prompt-engineering prompt-testing rag-evaluation skill-evaluation

Last synced: 15 Jun 2026

https://github.com/jhd3197/prompture

Prompture is an API-first library for requesting structured JSON output from LLMs (or any structure), validating it against a schema, and running comparative tests between models.

ai-testing json-validation llm openai prompt-engineering prompt-testing prompture pydantic structured-output toon

Last synced: 24 May 2026

https://github.com/prompt-foundry/typescript-sdk

The prompt engineering, prompt management, and prompt evaluation tool for TypeScript, JavaScript, and NodeJS.

gpt gpt-3 gpt-4 llm llm-eval llm-evaluation llm-ops llm-test llmops open-ai prompt-engineering prompt-evaluation prompt-management prompt-manager prompt-testing typescript

Last synced: 16 Jul 2026

https://github.com/syamsasi99/prompt-evaluator

prompt-evaluator is an open-source toolkit for evaluating, testing, and comparing LLM prompts. It provides a GUI-driven workflow for running prompt tests, tracking token usage, visualizing results, and ensuring reliability across models like OpenAI, Claude, and Gemini.

ai-evaluation ai-evaluation-framework ai-evaluation-metrics ai-evaluation-tools datascience developer-tools electron llm prompt-engineering prompt-testing promptfoo react typescript

Last synced: 13 Jan 2026

https://github.com/gaoyechen/sharpinput

AI最强嘴替-智能输入优化框架,让你的任何输入(提问、陈述、方案、想法、需求)都被打磨成能让 AI 输出有洞察力观点的高质量版本。

agent ai ai-tools chatgpt claude input-optimization llm openai prompt prompt-engineering prompt-enhancement prompt-optimizer prompt-testing prompt-tuning skill

Last synced: 06 Jul 2026

https://github.com/homemade-software-inc/completion-kit

Your prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and compare runs to see what got better.

anthropic evaluation-framework evaluation-metrics llm llm-as-judge llm-eval llm-evaluation llm-evaluation-framework llm-evaluation-metrics llmops mcp ollama openai prompt-engineering prompt-testing rails rails-engine ruby ruby-on-rails

Last synced: 09 Jul 2026

https://github.com/radoslaw-sz/maia

A pytest-based framework for testing multi AI agents systems. It provides a flexible and extensible platform for complex multi-agent simulations. Supports many integrations like LiteLLM, CrewAI, LangChain etc.

agentic agents ai ai-testing ai-testing-tool framework llm maia prompt-engineering prompt-testing python test

Last synced: 25 Sep 2025

https://github.com/tolom/prompt-xray-skills

Reusable SKILL.md workflows for writing, auditing, testing, and hardening LLM prompts.

agentic-workflows ai-agents llm prompt-audit prompt-engineering prompt-testing rag skills structured-outputs system-prompts

Last synced: 10 Jul 2026

https://github.com/yukinagae/promptfoo-sample

Sample project demonstrates how to use Promptfoo, a test framework for evaluating the output of generative AI models

evaluation evaluation-framework llm llm-eval llm-evaluation llm-evaluation-framework llmops prompt-testing promptfoo prompts testing

Last synced: 25 Feb 2026

https://github.com/piyushgupta344/llm-test-harness

Deterministic testing framework for LLM-powered apps — record/replay cassettes, eval scoring, regression testing

anthropic cassette deterministic eval llm openai prompt-testing testing vcr

Last synced: 03 Apr 2026

https://github.com/sigmakib2/openai-prompt-testing-playground

A dynamic and interactive playground for testing and refining prompts with OpenAI's language models. Includes customizable inputs for prompts, advanced model settings, and live response streaming for seamless experimentation.

ai chatgpt openai playground prompt prompt-engineering prompt-testing

Last synced: 21 May 2026

https://github.com/taimoorkhan10/replayd

Turn failed AI agent runs into replayable regression tests. Catch regressions before you ship.

agent-ops agent-testing ai-agents ai-infrastructure ai-reliability llm-ops llm-testing open-source prompt-testing python regression-testing release-control replay-testing sdk

Last synced: 14 Jun 2026

https://github.com/ac12644/prompt-diff

The regression-test gate for AI-generated prompts. Auto-improve prompts and verify the rewrite against a YAML test suite, or diff any two prompts. Supports OpenAI, Anthropic, Google Gemini.

anthropic claude cli evaluation gemini llm openai prompt-engineering prompt-optimization prompt-testing regression-testing typescript

Last synced: 13 Aug 2026