#testing
GitHub repositories that have self-applied the topic "testing" — a creator-tagged metadata that surfaces how AI projects describe themselves.
REPOS Repos for #testing (top 31 by stars)
AI Agent Toolkit 2026: Smart Device Control for iOS & Android
paultyng/testagentDeterministic fake of the claude and codex CLIs. Test hooks and orchestrator integrations locally and in CI without an API key.
quangdang46/livekit_agent_simulatorBlack-box LiveKit voice agent testing: real WebRTC/SIP calls with an AI caller. Catch barge-in, noise, quiet-mic, and latency failures that text pytest misses. Forensic reports + MCP + lks CLI — no agent code changes.
TheActualTwinkle/AiAssertionsAI-powered codebase assertions. Write natural-language requirements in tests and let an AI agent inspect your project.
sjh9714/red-handedAudits what your Claude Code agent did against what it said — reads session logs and git, cites evidence, calls no model
dicnunz/agentproofFree local proof reports for AI-built apps and public PRs; $149 async outside proof audit.
teshi-org/teshiAI Testing Agent
kos322/testrail-skillAI agent skill for TestRail REST API - zero dependencies, security-friendly alternative to MCP servers
prekshajain22/AI-Test-OrchestratorAI Quality Engineering framework for evaluating LLM and RAG applications using automated prompt testing, hallucination detection, and AI evaluation metrics.
Strangelight-Merser/agentic-testopsTurn failing pytest runs into structured repair reports: failure parsing, root-cause diagnosis, flaky detection, dry-run fix diffs, and an optional LLM layer. 把失败的 pytest 运行转化为结构化修复报告。
alialavia/pbt-skillsClaude Code skills for genuine property-based testing: discovery workflow, anti-patterns, Hypothesis & beyond.
digsarab112/taskwitnessEvidence-first verification for AI coding agents — prove what changed, what was tested, and what passed.
ohm41321/lucidoneVerification-first discipline for coding agents (Claude Code + Codex CLI): 9-rule doctrine, six skills, adversarial reviewer agent, fail-open enforcement hooks. Done is proven by a command, not by opinion.
CodewithJha/mutinyBehavioral fuzz-testing engine for AI agents — policy violations, tool-call verification, regression tests. Adapter #1: OpenAI Agents SDK.
1dg618/pytest-fakellmPytest fixtures for the fakellm mock OpenAI/Anthropic server
HAMZA-SQA-ai/Bizz-Card-Scanner-QA_ReportQA testing performed for the AI-powered Biz Card Scanner mobile application covering functional testing, UI validation, navigation flow, and bug reporting.
EffortlessMetrics/riprStatic Mutation Exposure Analysis
faisibash-oss/claude-hooks-practicePractice project migrating prose-based AI enforcement rules into real, blocking Claude Code hooks — covers PreToolUse guards, Stop-hook content verification, LibreOffice hang prevention, and robust pass/fail parsing, each with a working test suite.
silindokuhleL/ai-agent-project-playbookShared engineering rules, agent instructions, API standards, Laravel/Next.js patterns, testing rules, and reusable project checklists for all AI-powered portfolio systems.
nderman/agent-harnessA test & eval harness that makes an AI agent deterministic, testable, and observable — record/replay, trajectory evals, guardrails, traces
shridhar-mca24/rails-rspec-arsenalExpert Ruby Rails RSpec Skill Files 2026 for AI Coding Agents - Best GitHub About
venkatkrishnan060-cell/AI-Test-IntelligenceAI-powered QA Automation platform for log analysis, AI test case generation, and professional bug report creation using FastAPI, React, and Google Gemini AI.
dvnc-labs/agent-ready-uiMake web UIs controllable by browser agents without brittle selectors.
263311487-ux/dsh-verifyIndependent browser acceptance testing for agent deliverables. Agents self-test and pass; real browsers tell the truth.
skyswordw/skillmatrixBehavioral cross-agent testing for agent skills — run a skill on Claude Code and Codex and assert deterministic outcomes.
jleonceo/skill-detector-control-negativoSkill de Claude Code que audita si tus pruebas sabrían ponerse rojas. Instalable como plugin desde el marketplace. La misma herramienta que el repositorio hermano, empaquetada para que la use tu asistente.
Jott2121/sabotDo your agent pipeline's own checks catch planted faults? Measured on LangGraph, CrewAI and AutoGen: median 16.7%. One prompt-level change takes it to 55.0%. Pre-registered spec, Apache-2.0 harness, every raw trace published.
abhay23-AI/raggateA thin, CI-gated evaluation gate for RAG & LLM systems — golden set, band-based pass/warn/fail gates, LLM-judge or heuristic scorers. pip install raggate
bispeklolik/debug-by-usingDebug any product by USING it (living the full user cycle, all features, like a picky user) instead of green tests. Start interview (Goal+Benchmark), two modes: Solo and Agentic+ (delegate fixes to agents). Claude Code skill.
CypherMorgan/cypherpilotAI-powered Quality Engineering Platform built with FastAPI, React, PostgreSQL, and Docker for requirement analysis, API test generation, and automation failure analysis.
jmlkin/vm-script-benchAgent-drivable, cross-platform harness that tests PowerShell & Python scripts in throwaway VMs over SSH and asserts deterministic pass/fail. Validated live on Hyper-V + WSL. CI included.
RELATED Other topics · full topics ranking →
#claude-code
1,564#ai-agents
1,160#llm
1,071#claude
943#python
806#ai
741#developer-tools
728#mcp
727#codex
521Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology