AI Dev Impact Lab JA
← Topics ranking · 2026-08
GITHUB TOPIC

#testing

GitHub repositories that have self-applied the topic "testing" — a creator-tagged metadata that surfaces how AI projects describe themselves.

31
tagged repos
188
top 31 stars
10
with tool sigs
31
shown

REPOS Repos for #testing (top 31 by stars)

Brenonunesx/agent-pilot

AI Agent Toolkit 2026: Smart Device Control for iOS & Android

HTML 155 AI 100
paultyng/testagent

Deterministic fake of the claude and codex CLIs. Test hooks and orchestrator integrations locally and in CI without an API key.

Go 6 AI 70 2 sig
quangdang46/livekit_agent_simulator

Black-box LiveKit voice agent testing: real WebRTC/SIP calls with an AI caller. Catch barge-in, noise, quiet-mic, and latency failures that text pytest misses. Forensic reports + MCP + lks CLI — no agent code changes.

Python 5 AI 70 1 sig
TheActualTwinkle/AiAssertions

AI-powered codebase assertions. Write natural-language requirements in tests and let an AI agent inspect your project.

C# 4 AI 100
sjh9714/red-handed

Audits what your Claude Code agent did against what it said — reads session logs and git, cites evidence, calls no model

TypeScript 4 AI 70 Solo live ↗
dicnunz/agentproof

Free local proof reports for AI-built apps and public PRs; $149 async outside proof audit.

JavaScript 2 AI 70 Solo 1 sig live ↗
teshi-org/teshi

AI Testing Agent

Rust 2 AI 70 3 sig live ↗
kos322/testrail-skill

AI agent skill for TestRail REST API - zero dependencies, security-friendly alternative to MCP servers

Shell 1 AI 100
prekshajain22/AI-Test-Orchestrator

AI Quality Engineering framework for evaluating LLM and RAG applications using automated prompt testing, hallucination detection, and AI evaluation metrics.

Python 1 AI 100
Strangelight-Merser/agentic-testops

Turn failing pytest runs into structured repair reports: failure parsing, root-cause diagnosis, flaky detection, dry-run fix diffs, and an optional LLM layer. 把失败的 pytest 运行转化为结构化修复报告。

Python 1 AI 70
alialavia/pbt-skills

Claude Code skills for genuine property-based testing: discovery workflow, anti-patterns, Hypothesis & beyond.

Python 1 AI 70
digsarab112/taskwitness

Evidence-first verification for AI coding agents — prove what changed, what was tested, and what passed.

TypeScript 1 AI 70
ohm41321/lucidone

Verification-first discipline for coding agents (Claude Code + Codex CLI): 9-rule doctrine, six skills, adversarial reviewer agent, fail-open enforcement hooks. Done is proven by a command, not by opinion.

Shell 1 AI 70 Solo live ↗
CodewithJha/mutiny

Behavioral fuzz-testing engine for AI agents — policy violations, tool-call verification, regression tests. Adapter #1: OpenAI Agents SDK.

Python 1 AI 60 Solo live ↗
1dg618/pytest-fakellm

Pytest fixtures for the fakellm mock OpenAI/Anthropic server

Python 1 AI 50 Solo live ↗
HAMZA-SQA-ai/Bizz-Card-Scanner-QA_Report

QA testing performed for the AI-powered Biz Card Scanner mobile application covering functional testing, UI validation, navigation flow, and bug reporting.

1 AI 45
EffortlessMetrics/ripr

Static Mutation Exposure Analysis

Rust 1 AI 20 2 sig live ↗
faisibash-oss/claude-hooks-practice

Practice project migrating prose-based AI enforcement rules into real, blocking Claude Code hooks — covers PreToolUse guards, Stop-hook content verification, LibreOffice hang prevention, and robust pass/fail parsing, each with a working test suite.

Python 0 AI 100 1 sig
silindokuhleL/ai-agent-project-playbook

Shared engineering rules, agent instructions, API standards, Laravel/Next.js patterns, testing rules, and reusable project checklists for all AI-powered portfolio systems.

0 AI 100 1 sig
nderman/agent-harness

A test & eval harness that makes an AI agent deterministic, testable, and observable — record/replay, trajectory evals, guardrails, traces

TypeScript 0 AI 100 Solo 2 sig live ↗
shridhar-mca24/rails-rspec-arsenal

Expert Ruby Rails RSpec Skill Files 2026 for AI Coding Agents - Best GitHub About

HTML 0 AI 90
venkatkrishnan060-cell/AI-Test-Intelligence

AI-powered QA Automation platform for log analysis, AI test case generation, and professional bug report creation using FastAPI, React, and Google Gemini AI.

JavaScript 0 AI 90 Solo live ↗
dvnc-labs/agent-ready-ui

Make web UIs controllable by browser agents without brittle selectors.

0 AI 75 1 sig live ↗
263311487-ux/dsh-verify

Independent browser acceptance testing for agent deliverables. Agents self-test and pass; real browsers tell the truth.

JavaScript 0 AI 70
skyswordw/skillmatrix

Behavioral cross-agent testing for agent skills — run a skill on Claude Code and Codex and assert deterministic outcomes.

TypeScript 0 AI 70 Solo live ↗
jleonceo/skill-detector-control-negativo

Skill de Claude Code que audita si tus pruebas sabrían ponerse rojas. Instalable como plugin desde el marketplace. La misma herramienta que el repositorio hermano, empaquetada para que la use tu asistente.

Python 0 AI 70
Jott2121/sabot

Do your agent pipeline's own checks catch planted faults? Measured on LangGraph, CrewAI and AutoGen: median 16.7%. One prompt-level change takes it to 55.0%. Pre-registered spec, Apache-2.0 harness, every raw trace published.

Python 0 AI 70 Solo live ↗
abhay23-AI/raggate

A thin, CI-gated evaluation gate for RAG & LLM systems — golden set, band-based pass/warn/fail gates, LLM-judge or heuristic scorers. pip install raggate

Python 0 AI 70 Solo live ↗
bispeklolik/debug-by-using

Debug any product by USING it (living the full user cycle, all features, like a picky user) instead of green tests. Start interview (Goal+Benchmark), two modes: Solo and Agentic+ (delegate fixes to agents). Claude Code skill.

0 AI 70
CypherMorgan/cypherpilot

AI-powered Quality Engineering Platform built with FastAPI, React, PostgreSQL, and Docker for requirement analysis, API test generation, and automation failure analysis.

Python 0 AI 70 Solo live ↗
jmlkin/vm-script-bench

Agent-drivable, cross-platform harness that tests PowerShell & Python scripts in throwaway VMs over SSH and asserts deterministic pass/fail. Validated live on Hyper-V + WSL. CI included.

Python 0 AI 45 2 sig

RELATED Other topics · full topics ranking →

#claude-code

1,564

#ai-agents

1,160

#llm

1,071

#claude

943

#python

806

#ai

741

#developer-tools

728

#mcp

727

#codex

521

Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology