AI開発影響研究所 EN
← Topicランキング · 2026-08
GitHub TOPIC

#testing

GitHub Topic「testing」がついているAI関連リポジトリの集計。Topicはリポジトリ作者が自己申告するメタタグで、AI関連の文脈で何が「testing」と分類されているかを可視化します。

31
タグ付Repo
188
TOP31合計★
10
AIツール痕跡あり
31
TOP表示数

REPOS #testing のRepo (TOP 31 / Stars降順)

Brenonunesx/agent-pilot

AI Agent Toolkit 2026: Smart Device Control for iOS & Android

HTML 155 AI 100
paultyng/testagent

Deterministic fake of the claude and codex CLIs. Test hooks and orchestrator integrations locally and in CI without an API key.

Go 6 AI 70 2 sig
quangdang46/livekit_agent_simulator

Black-box LiveKit voice agent testing: real WebRTC/SIP calls with an AI caller. Catch barge-in, noise, quiet-mic, and latency failures that text pytest misses. Forensic reports + MCP + lks CLI — no agent code changes.

Python 5 AI 70 1 sig
TheActualTwinkle/AiAssertions

AI-powered codebase assertions. Write natural-language requirements in tests and let an AI agent inspect your project.

C# 4 AI 100
sjh9714/red-handed

Audits what your Claude Code agent did against what it said — reads session logs and git, cites evidence, calls no model

TypeScript 4 AI 70 個人 公開済 ↗
dicnunz/agentproof

Free local proof reports for AI-built apps and public PRs; $149 async outside proof audit.

JavaScript 2 AI 70 個人 1 sig 公開済 ↗
teshi-org/teshi

AI Testing Agent

Rust 2 AI 70 3 sig 公開済 ↗
kos322/testrail-skill

AI agent skill for TestRail REST API - zero dependencies, security-friendly alternative to MCP servers

Shell 1 AI 100
prekshajain22/AI-Test-Orchestrator

AI Quality Engineering framework for evaluating LLM and RAG applications using automated prompt testing, hallucination detection, and AI evaluation metrics.

Python 1 AI 100
Strangelight-Merser/agentic-testops

Turn failing pytest runs into structured repair reports: failure parsing, root-cause diagnosis, flaky detection, dry-run fix diffs, and an optional LLM layer. 把失败的 pytest 运行转化为结构化修复报告。

Python 1 AI 70
alialavia/pbt-skills

Claude Code skills for genuine property-based testing: discovery workflow, anti-patterns, Hypothesis & beyond.

Python 1 AI 70
digsarab112/taskwitness

Evidence-first verification for AI coding agents — prove what changed, what was tested, and what passed.

TypeScript 1 AI 70
ohm41321/lucidone

Verification-first discipline for coding agents (Claude Code + Codex CLI): 9-rule doctrine, six skills, adversarial reviewer agent, fail-open enforcement hooks. Done is proven by a command, not by opinion.

Shell 1 AI 70 個人 公開済 ↗
CodewithJha/mutiny

Behavioral fuzz-testing engine for AI agents — policy violations, tool-call verification, regression tests. Adapter #1: OpenAI Agents SDK.

Python 1 AI 60 個人 公開済 ↗
1dg618/pytest-fakellm

Pytest fixtures for the fakellm mock OpenAI/Anthropic server

Python 1 AI 50 個人 公開済 ↗
HAMZA-SQA-ai/Bizz-Card-Scanner-QA_Report

QA testing performed for the AI-powered Biz Card Scanner mobile application covering functional testing, UI validation, navigation flow, and bug reporting.

1 AI 45
EffortlessMetrics/ripr

Static Mutation Exposure Analysis

Rust 1 AI 20 2 sig 公開済 ↗
faisibash-oss/claude-hooks-practice

Practice project migrating prose-based AI enforcement rules into real, blocking Claude Code hooks — covers PreToolUse guards, Stop-hook content verification, LibreOffice hang prevention, and robust pass/fail parsing, each with a working test suite.

Python 0 AI 100 1 sig
silindokuhleL/ai-agent-project-playbook

Shared engineering rules, agent instructions, API standards, Laravel/Next.js patterns, testing rules, and reusable project checklists for all AI-powered portfolio systems.

0 AI 100 1 sig
nderman/agent-harness

A test & eval harness that makes an AI agent deterministic, testable, and observable — record/replay, trajectory evals, guardrails, traces

TypeScript 0 AI 100 個人 2 sig 公開済 ↗
shridhar-mca24/rails-rspec-arsenal

Expert Ruby Rails RSpec Skill Files 2026 for AI Coding Agents - Best GitHub About

HTML 0 AI 90
venkatkrishnan060-cell/AI-Test-Intelligence

AI-powered QA Automation platform for log analysis, AI test case generation, and professional bug report creation using FastAPI, React, and Google Gemini AI.

JavaScript 0 AI 90 個人 公開済 ↗
dvnc-labs/agent-ready-ui

Make web UIs controllable by browser agents without brittle selectors.

0 AI 75 1 sig 公開済 ↗
263311487-ux/dsh-verify

Independent browser acceptance testing for agent deliverables. Agents self-test and pass; real browsers tell the truth.

JavaScript 0 AI 70
skyswordw/skillmatrix

Behavioral cross-agent testing for agent skills — run a skill on Claude Code and Codex and assert deterministic outcomes.

TypeScript 0 AI 70 個人 公開済 ↗
jleonceo/skill-detector-control-negativo

Skill de Claude Code que audita si tus pruebas sabrían ponerse rojas. Instalable como plugin desde el marketplace. La misma herramienta que el repositorio hermano, empaquetada para que la use tu asistente.

Python 0 AI 70
Jott2121/sabot

Do your agent pipeline's own checks catch planted faults? Measured on LangGraph, CrewAI and AutoGen: median 16.7%. One prompt-level change takes it to 55.0%. Pre-registered spec, Apache-2.0 harness, every raw trace published.

Python 0 AI 70 個人 公開済 ↗
abhay23-AI/raggate

A thin, CI-gated evaluation gate for RAG & LLM systems — golden set, band-based pass/warn/fail gates, LLM-judge or heuristic scorers. pip install raggate

Python 0 AI 70 個人 公開済 ↗
bispeklolik/debug-by-using

Debug any product by USING it (living the full user cycle, all features, like a picky user) instead of green tests. Start interview (Goal+Benchmark), two modes: Solo and Agentic+ (delegate fixes to agents). Claude Code skill.

0 AI 70
CypherMorgan/cypherpilot

AI-powered Quality Engineering Platform built with FastAPI, React, PostgreSQL, and Docker for requirement analysis, API test generation, and automation failure analysis.

Python 0 AI 70 個人 公開済 ↗
jmlkin/vm-script-bench

Agent-drivable, cross-platform harness that tests PowerShell & Python scripts in throwaway VMs over SSH and asserts deterministic pass/fail. Validated live on Hyper-V + WSL. CI included.

Python 0 AI 45 2 sig

RELATED 他のTopicも見る · 全Topicランキング →

#claude-code

1,564

#ai-agents

1,160

#llm

1,071

#claude

943

#python

806

#ai

741

#developer-tools

728

#mcp

727

#codex

521

集計対象: 各Repoの最新contentスナップショットの topics_json に小文字一致でマッチしたもの。 算出方法