#evals
GitHub Topic「evals」がついているAI関連リポジトリの集計。Topicはリポジトリ作者が自己申告するメタタグで、AI関連の文脈で何が「evals」と分類されているかを可視化します。
REPOS #evals のRepo (TOP 15 / Stars降順)
Proof Gradient is the agent evolution protocol where every run leaves proof, every proof selects intelligence, and every selected artifact evolves the network.
zwright8/flight-recorderEvidence, governance, and static reports for agentic runs
zaidazmi/AI-PM-PLAYBOOKPlaybook for PMs shipping AI products with PRDs, evals, HITL, launch gates, cost, and observability.
Johna2an/critical-thinkingOpen-source Agent Skill for Claude Code that makes Claude reason with explicit probabilities, base rates, and falsifiers. Distilled from 51 books; first in 18/18 blind passes vs baseline Claude and GPT-5.5, and benchmarked against 14 rival skills and prompt techniques with full cost telemetry (round 5).
Nazim22/leadlineEvidence router & policy engine for coding agents — enforces source choice and proof-of-use through Claude Code hooks, including routes to MCP tools. Local-first; no LLM in the hook loop.
lignos-ai/lignos-labsJudgment frameworks for AI agents — Canvas → Govern → Scope → Score. Free Labs skills for Claude Code, Cursor, and eval handoff.
Dafenxz0/skillproofProve your Agent Skill works before you publish it.
kizz-tech/agentic-evidence-labTests exact agent interventions—skills, prompts, models, tools, and workflows—and publishes what changed, what failed, and what decision follows.
Gowrav-M/agentops-watchtowerLocal-first AgentOps flight recorder and capability firewall for MCP-based coding agents with OpenTelemetry traces.
Taz33m/codex-demo-day-arenaProduct-foundry control plane for AI coding agents: validator gates, buyer scoring, compute allocation, and Demo Day-style investment decisions across startup product candidates. Next step would be in automating outreach to get real PMF validation and iterate from there.
alexcao11/ai-llm-automation-systemsAI/LLM automation reference: RAG, tool calling, human review, evals, observability, safety gates, and local deterministic samples.
tonydzi/agent-runtime-integrity-benchIntegrity scenarios for agent runtimes, distilled from real fleet incidents — deterministic fault-injection checks against real SDKs
multivon-ai/multivon-mcpMCP server exposing multivon-eval + pdfhell as agent-callable tools. Drop into Claude Desktop, Cursor, Cline.
justuseapen/polarity-benchDoes following a negated instruction come free with the positive one? A reproducible eval — two pre-registered nulls for frontier Claude.
Samuelmartinezduran/agentevalFramework open source para evaluar agentes LLM con function calling / tool use. Scoring por dimensión (tool accuracy, response quality, safety) + CLI, API y dashboard.
RELATED 他のTopicも見る · 全Topicランキング →
#claude-code
1,564#ai-agents
1,160#llm
1,071#claude
943#python
806#ai
741#developer-tools
728#mcp
727#codex
521集計対象: 各Repoの最新contentスナップショットの topics_json に小文字一致でマッチしたもの。 算出方法