From 0c696135127f369eb09b0ac2354a80de42d469be Mon Sep 17 00:00:00 2001 From: Pranshu Chittora Date: Fri, 3 Jul 2026 03:33:46 +0530 Subject: [PATCH] Add agent-qa --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index 88d649a..be5006f 100644 --- a/README.md +++ b/README.md @@ -155,6 +155,7 @@ Most "awesome" lists are link dumps. This one is **annotated and verified**: eve - **[simple-evals](https://github.com/openai/simple-evals)** — OpenAI — — minimal zero-shot/CoT scripts (MMLU, HumanEval, SimpleQA, HealthBench); the numbers OpenAI publishes. ⚠️ not actively maintained. - **[OpenAI Evals](https://github.com/openai/evals)** — — the `completion_fn` abstraction = swap the system-under-test. (Best-practices: ) - **[promptfoo](https://github.com/promptfoo/promptfoo)** — — MIT eval + red-teaming CLI; git-diffable YAML configs. **(MUST)** +- **[agent-qa](https://github.com/vostride/agent-qa)** — Vostride — · *tool/repo* — 🆕 Natural-language QA harness for web and mobile apps: memory-backed self-healing runs, dashboard/CLI, MCP/skills support, and sandboxed hooks for regression checks. - **[DeepEval / Confident AI](https://github.com/confident-ai/deepeval)** — — "pytest for LLMs," 40+ metrics (G-Eval, RAG, hallucination) + red-team; ~2M evals/day; hosted cloud. 🆕 - **[pydantic-evals](https://github.com/pydantic/pydantic-ai)** — (`ai.pydantic.dev/evals`) — 🆕 type-safe Datasets/Cases/Evaluators with OTel tracing, from the Pydantic AI team. - **[openevals](https://github.com/langchain-ai/openevals)** — LangChain — — 🆕 prebuilt evaluators + `create_llm_as_judge` (incl. multimodal); general-purpose companion to **agentevals** (, trajectory match). @@ -567,4 +568,3 @@ To the extent possible under law, [BenchFlow](https://benchflow.ai) and contribu -