AI Dev Impact Lab JA
← Topics ranking · 2026-08
GITHUB TOPIC

#evals

GitHub repositories that have self-applied the topic "evals" — a creator-tagged metadata that surfaces how AI projects describe themselves.

15
tagged repos
17
top 15 stars
5
with tool sigs
15
shown

REPOS Repos for #evals (top 15 by stars)

MontrealAI/proof-gradient

Proof Gradient is the agent evolution protocol where every run leaves proof, every proof selects intelligence, and every selected artifact evolves the network.

HTML 4 AI 70 Solo live ↗
zwright8/flight-recorder

Evidence, governance, and static reports for agentic runs

Python 3 AI 45 1 sig
zaidazmi/AI-PM-PLAYBOOK

Playbook for PMs shipping AI products with PRDs, evals, HITL, launch gates, cost, and observability.

TypeScript 2 AI 100 Solo live ↗
Johna2an/critical-thinking

Open-source Agent Skill for Claude Code that makes Claude reason with explicit probabilities, base rates, and falsifiers. Distilled from 51 books; first in 18/18 blind passes vs baseline Claude and GPT-5.5, and benchmarked against 14 rival skills and prompt techniques with full cost telemetry (round 5).

HTML 2 AI 70 2 sig
Nazim22/leadline

Evidence router & policy engine for coding agents — enforces source choice and proof-of-use through Claude Code hooks, including routes to MCP tools. Local-first; no LLM in the hook loop.

JavaScript 2 AI 70 1 sig
lignos-ai/lignos-labs

Judgment frameworks for AI agents — Canvas → Govern → Scope → Score. Free Labs skills for Claude Code, Cursor, and eval handoff.

1 AI 70
Dafenxz0/skillproof

Prove your Agent Skill works before you publish it.

HTML 1 AI 70
kizz-tech/agentic-evidence-lab

Tests exact agent interventions—skills, prompts, models, tools, and workflows—and publishes what changed, what failed, and what decision follows.

Python 1 AI 70 1 sig
Gowrav-M/agentops-watchtower

Local-first AgentOps flight recorder and capability firewall for MCP-based coding agents with OpenTelemetry traces.

TypeScript 1 AI 45
Taz33m/codex-demo-day-arena

Product-foundry control plane for AI coding agents: validator gates, buyer scoring, compute allocation, and Demo Day-style investment decisions across startup product candidates. Next step would be in automating outreach to get real PMF validation and iterate from there.

Python 0 AI 100
alexcao11/ai-llm-automation-systems

AI/LLM automation reference: RAG, tool calling, human review, evals, observability, safety gates, and local deterministic samples.

Python 0 AI 100
tonydzi/agent-runtime-integrity-bench

Integrity scenarios for agent runtimes, distilled from real fleet incidents — deterministic fault-injection checks against real SDKs

Python 0 AI 100 1 sig
multivon-ai/multivon-mcp

MCP server exposing multivon-eval + pdfhell as agent-callable tools. Drop into Claude Desktop, Cursor, Cline.

Python 0 AI 70 live ↗
justuseapen/polarity-bench

Does following a negated instruction come free with the positive one? A reproducible eval — two pre-registered nulls for frontier Claude.

Python 0 AI 70
Samuelmartinezduran/agenteval

Framework open source para evaluar agentes LLM con function calling / tool use. Scoring por dimensión (tool accuracy, response quality, safety) + CLI, API y dashboard.

Python 0 AI 70

RELATED Other topics · full topics ranking →

#claude-code

1,564

#ai-agents

1,160

#llm

1,071

#claude

943

#python

806

#ai

741

#developer-tools

728

#mcp

727

#codex

521

Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology