#agent-evaluation
GitHub repositories that have self-applied the topic "agent-evaluation" — a creator-tagged metadata that surfaces how AI projects describe themselves.
REPOS Repos for #agent-evaluation (top 15 by stars)
🔁 Build reliable recurring AI-agent systems: 968 resources, 22 operational patterns, 22 loop contracts, 8 runtime starters, an interactive atlas, and a structured dataset.
linny006/agent-eval-harnessLive, open-source benchmark for comparing AI coding agents on real GitHub issues
MontrealAI/proof-gradientProof Gradient is the agent evolution protocol where every run leaves proof, every proof selects intelligence, and every selected artifact evolves the network.
FU-max-boop/statebind-guardCatch visible-but-unbound coding-agent handoffs: CLI + GitHub Action with proof, policy gates, SARIF, HTML, and benchmark cards.
kizz-tech/agentic-evidence-labTests exact agent interventions—skills, prompts, models, tools, and workflows—and publishes what changed, what failed, and what decision follows.
xltzsoft/dsh-verifierDSH 的全自动 LLM-as-a-Verifier 模式:支持并行/串行候选、PPT 择优与实时可视化
cuiqi5656/agent-quality-benchmarkEvidence-first, reproducible quality benchmarks for AI agents.
justhandledlabs/justhandled-agent-clientGuarded CLI, TypeScript client, and MCP server for JustHandled x402 agent utilities
maximizeGPT/claude-eval-harnessRegression-diff eval harness for Anthropic tool-use agents that surfaces LLM-judge reasoning drift, not just pass/fail flips
DiogoRibeiro7/agentic-qa-labAutonomous UI/game-testing agent: vision-language reasoning, browser control, action planning, failure recovery, and evaluation.
Samuelmartinezduran/agentevalFramework open source para evaluar agentes LLM con function calling / tool use. Scoring por dimensión (tool accuracy, response quality, safety) + CLI, API y dashboard.
Jott2121/sabotDo your agent pipeline's own checks catch planted faults? Measured on LangGraph, CrewAI and AutoGen: median 16.7%. One prompt-level change takes it to 55.0%. Pre-registered spec, Apache-2.0 harness, every raw trace published.
ReverseZoom2151/craftonomousCraftonomous is an agent-agnostic Minecraft embodiment and evaluation substrate: an MCP-native body with a declared perception budget and reliability-tracked skills.
satwiksps/scaffoldscopeControlled coding-agent harness ablations with auditable traces, reproducible evidence bundles, and SWE-bench interoperability.
roy-tong/AgentMeasureOpen measurement infrastructure for the Agent Capability Economy — how AI agents discover, choose, use, and derive value from software capabilities (Reach → Choice → Use → Utility → Value). A proposed open measurement standard.
RELATED Other topics · full topics ranking →
#claude-code
1,555#ai-agents
1,156#llm
1,066#claude
936#python
802#ai
737#developer-tools
723#mcp
719#codex
517Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology