AI開発影響研究所 EN
← Topicランキング · 2026-08
GitHub TOPIC

#evals

GitHub Topic「evals」がついているAI関連リポジトリの集計。Topicはリポジトリ作者が自己申告するメタタグで、AI関連の文脈で何が「evals」と分類されているかを可視化します。

15
タグ付Repo
17
TOP15合計★
5
AIツール痕跡あり
15
TOP表示数

REPOS #evals のRepo (TOP 15 / Stars降順)

MontrealAI/proof-gradient

Proof Gradient is the agent evolution protocol where every run leaves proof, every proof selects intelligence, and every selected artifact evolves the network.

HTML 4 AI 70 個人 公開済 ↗
zwright8/flight-recorder

Evidence, governance, and static reports for agentic runs

Python 3 AI 45 1 sig
zaidazmi/AI-PM-PLAYBOOK

Playbook for PMs shipping AI products with PRDs, evals, HITL, launch gates, cost, and observability.

TypeScript 2 AI 100 個人 公開済 ↗
Johna2an/critical-thinking

Open-source Agent Skill for Claude Code that makes Claude reason with explicit probabilities, base rates, and falsifiers. Distilled from 51 books; first in 18/18 blind passes vs baseline Claude and GPT-5.5, and benchmarked against 14 rival skills and prompt techniques with full cost telemetry (round 5).

HTML 2 AI 70 2 sig
Nazim22/leadline

Evidence router & policy engine for coding agents — enforces source choice and proof-of-use through Claude Code hooks, including routes to MCP tools. Local-first; no LLM in the hook loop.

JavaScript 2 AI 70 1 sig
lignos-ai/lignos-labs

Judgment frameworks for AI agents — Canvas → Govern → Scope → Score. Free Labs skills for Claude Code, Cursor, and eval handoff.

1 AI 70
Dafenxz0/skillproof

Prove your Agent Skill works before you publish it.

HTML 1 AI 70
kizz-tech/agentic-evidence-lab

Tests exact agent interventions—skills, prompts, models, tools, and workflows—and publishes what changed, what failed, and what decision follows.

Python 1 AI 70 1 sig
Gowrav-M/agentops-watchtower

Local-first AgentOps flight recorder and capability firewall for MCP-based coding agents with OpenTelemetry traces.

TypeScript 1 AI 45
Taz33m/codex-demo-day-arena

Product-foundry control plane for AI coding agents: validator gates, buyer scoring, compute allocation, and Demo Day-style investment decisions across startup product candidates. Next step would be in automating outreach to get real PMF validation and iterate from there.

Python 0 AI 100
alexcao11/ai-llm-automation-systems

AI/LLM automation reference: RAG, tool calling, human review, evals, observability, safety gates, and local deterministic samples.

Python 0 AI 100
tonydzi/agent-runtime-integrity-bench

Integrity scenarios for agent runtimes, distilled from real fleet incidents — deterministic fault-injection checks against real SDKs

Python 0 AI 100 1 sig
multivon-ai/multivon-mcp

MCP server exposing multivon-eval + pdfhell as agent-callable tools. Drop into Claude Desktop, Cursor, Cline.

Python 0 AI 70 公開済 ↗
justuseapen/polarity-bench

Does following a negated instruction come free with the positive one? A reproducible eval — two pre-registered nulls for frontier Claude.

Python 0 AI 70
Samuelmartinezduran/agenteval

Framework open source para evaluar agentes LLM con function calling / tool use. Scoring por dimensión (tool accuracy, response quality, safety) + CLI, API y dashboard.

Python 0 AI 70

RELATED 他のTopicも見る · 全Topicランキング →

#claude-code

1,564

#ai-agents

1,160

#llm

1,071

#claude

943

#python

806

#ai

741

#developer-tools

728

#mcp

727

#codex

521

集計対象: 各Repoの最新contentスナップショットの topics_json に小文字一致でマッチしたもの。 算出方法