AI Dev Impact Lab JA
← Topics ranking · 2026-08
GITHUB TOPIC

#agent-evaluation

GitHub repositories that have self-applied the topic "agent-evaluation" — a creator-tagged metadata that surfaces how AI projects describe themselves.

15
tagged repos
64
top 15 stars
3
with tool sigs
15
shown

REPOS Repos for #agent-evaluation (top 15 by stars)

ChaoYue0307/awesome-loop-engineering

🔁 Build reliable recurring AI-agent systems: 968 resources, 22 operational patterns, 22 loop contracts, 8 runtime starters, an interactive atlas, and a structured dataset.

Python 50 AI 70 Solo 1 sig live ↗
linny006/agent-eval-harness

Live, open-source benchmark for comparing AI coding agents on real GitHub issues

Python 6 AI 90
MontrealAI/proof-gradient

Proof Gradient is the agent evolution protocol where every run leaves proof, every proof selects intelligence, and every selected artifact evolves the network.

HTML 4 AI 70 Solo live ↗
FU-max-boop/statebind-guard

Catch visible-but-unbound coding-agent handoffs: CLI + GitHub Action with proof, policy gates, SARIF, HTML, and benchmark cards.

Python 2 AI 60 Solo live ↗
kizz-tech/agentic-evidence-lab

Tests exact agent interventions—skills, prompts, models, tools, and workflows—and publishes what changed, what failed, and what decision follows.

Python 1 AI 70 1 sig
xltzsoft/dsh-verifier

DSH 的全自动 LLM-as-a-Verifier 模式:支持并行/串行候选、PPT 择优与实时可视化

JavaScript 1 AI 70
cuiqi5656/agent-quality-benchmark

Evidence-first, reproducible quality benchmarks for AI agents.

Python 0 AI 100
justhandledlabs/justhandled-agent-client

Guarded CLI, TypeScript client, and MCP server for JustHandled x402 agent utilities

TypeScript 0 AI 100 live ↗
maximizeGPT/claude-eval-harness

Regression-diff eval harness for Anthropic tool-use agents that surfaces LLM-judge reasoning drift, not just pass/fail flips

Python 0 AI 90
DiogoRibeiro7/agentic-qa-lab

Autonomous UI/game-testing agent: vision-language reasoning, browser control, action planning, failure recovery, and evaluation.

Python 0 AI 70 live ↗
Samuelmartinezduran/agenteval

Framework open source para evaluar agentes LLM con function calling / tool use. Scoring por dimensión (tool accuracy, response quality, safety) + CLI, API y dashboard.

Python 0 AI 70
Jott2121/sabot

Do your agent pipeline's own checks catch planted faults? Measured on LangGraph, CrewAI and AutoGen: median 16.7%. One prompt-level change takes it to 55.0%. Pre-registered spec, Apache-2.0 harness, every raw trace published.

Python 0 AI 70 Solo live ↗
ReverseZoom2151/craftonomous

Craftonomous is an agent-agnostic Minecraft embodiment and evaluation substrate: an MCP-native body with a declared perception budget and reliability-tracked skills.

TypeScript 0 AI 70
satwiksps/scaffoldscope

Controlled coding-agent harness ablations with auditable traces, reproducible evidence bundles, and SWE-bench interoperability.

Python 0 AI 70 Solo 1 sig live ↗
roy-tong/AgentMeasure

Open measurement infrastructure for the Agent Capability Economy — how AI agents discover, choose, use, and derive value from software capabilities (Reach → Choice → Use → Utility → Value). A proposed open measurement standard.

Python 0 AI 70 Solo live ↗

RELATED Other topics · full topics ranking →

#claude-code

1,555

#ai-agents

1,156

#llm

1,066

#claude

936

#python

802

#ai

737

#developer-tools

723

#mcp

719

#codex

517

Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology