#ai-evaluation
GitHub Topic「ai-evaluation」がついているAI関連リポジトリの集計。Topicはリポジトリ作者が自己申告するメタタグで、AI関連の文脈で何が「ai-evaluation」と分類されているかを可視化します。
REPOS #ai-evaluation のRepo (TOP 9 / Stars降順)
Live, open-source benchmark for comparing AI coding agents on real GitHub issues
NavidBroumandfar/agent-behavior-evals-labPolicy-mapped evaluation lab for AI assistant behavior: approval gates, refusal boundaries, uncertainty handling, tool-use grounding, traces, and quality gates.
Nazim22/leadlineEvidence router & policy engine for coding agents — enforces source choice and proof-of-use through Claude Code hooks, including routes to MCP tools. Local-first; no LLM in the hook loop.
jawwad-ali/self-healing-ragSelf-healing RAG on n8n with LLM-as-judge eval, drift detection, investigator agent, and canary A/B deploys.
kwakusei1m-tech/AI-Evaluation-of-E-commerce-Chatbot-Responses-Using-Rubric-Based-Quality-ScoringAI Evaluation project analysing chatbot response quality in e-commerce using rubric scoring, error taxonomy, and semi-automated evaluation workflows.
api-evangelist/luminosaiLuminos.AI is an AI governance and evaluation platform that tests AI systems — classical machine learning, generative AI, and autonomous agents — for legal, regulatory, and reputational risk.
multivon-ai/multivon-mcpMCP server exposing multivon-eval + pdfhell as agent-callable tools. Drop into Claude Desktop, Cursor, Cline.
thangldw/ragopsOffline regression tests and explainable release gates for RAG systems and AI agents.
RudraO2/tokenbrawlA latency-fair LLM-vs-LLM fighting game benchmark. The engine blocks until an agent responds, so inference speed can't win a match — what's scarce is a per-match token bank that drains as a model thinks.
RELATED 他のTopicも見る · 全Topicランキング →
#claude-code
1,564#ai-agents
1,160#llm
1,071#claude
943#python
806#ai
741#developer-tools
728#mcp
727#codex
521集計対象: 各Repoの最新contentスナップショットの topics_json に小文字一致でマッチしたもの。 算出方法