AI開発影響研究所 EN
← Topicランキング · 2026-08
GitHub TOPIC

#ai-evaluation

GitHub Topic「ai-evaluation」がついているAI関連リポジトリの集計。Topicはリポジトリ作者が自己申告するメタタグで、AI関連の文脈で何が「ai-evaluation」と分類されているかを可視化します。

9
タグ付Repo
11
TOP9合計★
4
AIツール痕跡あり
9
TOP表示数

REPOS #ai-evaluation のRepo (TOP 9 / Stars降順)

linny006/agent-eval-harness

Live, open-source benchmark for comparing AI coding agents on real GitHub issues

Python 6 AI 90
NavidBroumandfar/agent-behavior-evals-lab

Policy-mapped evaluation lab for AI assistant behavior: approval gates, refusal boundaries, uncertainty handling, tool-use grounding, traces, and quality gates.

Python 2 AI 100 個人 2 sig 公開済 ↗
Nazim22/leadline

Evidence router & policy engine for coding agents — enforces source choice and proof-of-use through Claude Code hooks, including routes to MCP tools. Local-first; no LLM in the hook loop.

JavaScript 2 AI 70 1 sig
jawwad-ali/self-healing-rag

Self-healing RAG on n8n with LLM-as-judge eval, drift detection, investigator agent, and canary A/B deploys.

TypeScript 1 AI 100 個人 2 sig 公開済 ↗
kwakusei1m-tech/AI-Evaluation-of-E-commerce-Chatbot-Responses-Using-Rubric-Based-Quality-Scoring

AI Evaluation project analysing chatbot response quality in e-commerce using rubric scoring, error taxonomy, and semi-automated evaluation workflows.

Jupyter Notebook 0 AI 100
api-evangelist/luminosai

Luminos.AI is an AI governance and evaluation platform that tests AI systems — classical machine learning, generative AI, and autonomous agents — for legal, regulatory, and reputational risk.

0 AI 100
multivon-ai/multivon-mcp

MCP server exposing multivon-eval + pdfhell as agent-callable tools. Drop into Claude Desktop, Cursor, Cline.

Python 0 AI 70 公開済 ↗
thangldw/ragops

Offline regression tests and explainable release gates for RAG systems and AI agents.

Python 0 AI 70
RudraO2/tokenbrawl

A latency-fair LLM-vs-LLM fighting game benchmark. The engine blocks until an agent responds, so inference speed can't win a match — what's scarce is a per-match token bank that drains as a model thinks.

TypeScript 0 AI 70 1 sig

RELATED 他のTopicも見る · 全Topicランキング →

#claude-code

1,564

#ai-agents

1,160

#llm

1,071

#claude

943

#python

806

#ai

741

#developer-tools

728

#mcp

727

#codex

521

集計対象: 各Repoの最新contentスナップショットの topics_json に小文字一致でマッチしたもの。 算出方法