AI開発影響研究所 EN
← Topicランキング · 2026-08
GitHub TOPIC

#evaluation

GitHub Topic「evaluation」がついているAI関連リポジトリの集計。Topicはリポジトリ作者が自己申告するメタタグで、AI関連の文脈で何が「evaluation」と分類されているかを可視化します。

26
タグ付Repo
19
TOP26合計★
3
AIツール痕跡あり
26
TOP表示数

REPOS #evaluation のRepo (TOP 26 / Stars降順)

AgentBenchAudit/evidence-bounds

Release repository for agent benchmark evidence-reporting artifacts and reproduction workflows.

HTML 5 AI 70 公開済 ↗
forrestbthomas/pi-harness

Provider-agnostic coding-agent harness + DeepEval evaluation suite

Go 3 AI 70 個人 1 sig 公開済 ↗
vahit19/longhaul-bench

Long-horizon reliability benchmark for industrial edge agents - do self-improving agents get better or corrupt over 1000+ episodes on constrained hardware?

Python 2 AI 70
zhanghanbo668-art/embodied-eval-loop

Software-only embodied AI evaluation stack for benchmark trajectories, ROS-compatible logs, replay, failure analysis, and comparison reports.

Python 2 AI 70
DmitryDmitriadi/llm-observability-evaluation

Case study: unified observability + evaluation for automation and conversational AI agents. Versioned error catalog, multi-signal confidence, cost meters.

1 AI 100
maddog22coder/agent-eval

Open-source multilingual evaluation toolkit for conversational AI agents.

Python 1 AI 100
iiTzAK/accounting-copilot

Fine-tuning an open 7B LLM into a GST question-answering assistant, with a reproducible eval harness measuring the gain over the base model.

Python 1 AI 100
auraoneai/rubric-studio-open

Author, validate, preview, calibrate, diff, and export criterion-level AI evaluation rubrics.

TypeScript 1 AI 70 公開済 ↗
beausome/bullshit-bench

An AI leaderboard scoring chatbots on how much they bullshit vs. give honest answers to unanswerable questions. Works with paid APIs and free/local models.

Python 1 AI 70
wwjorker/devmind

AI knowledge base & RAG QA system with gold-label retrieval evaluation (Spring Boot 3 + Vue 3): hybrid keyword/vector search, RRF fusion, rerank, four-way Hit@3/MRR comparison

Java 1 AI 70
clouatre-labs/reversibility-benchmark

Supplementary materials for Reversibility as a Safety Gate for Agentic Actions

Python 1 AI 45 1 sig
chohyerinn/mini-agent-harness

Evaluation harness for coding agents - repeated runs, pass@k, bootstrap/McNemar significance, tamper detection, cost. Single vs Planner-Coder-Reviewer multi-agent on CLOVA.

Python 0 AI 100
Hal-Hanami/incident-triage-agent

Read-only first-pass incident triage on the Claude Agent SDK + MCP: classifies an alert, retrieves the matching runbook, and proposes a cited first response — or abstains to a human. Measured over four runs: 100% abstention with 0 missed escalations, ~$0.014 per incident.

Python 0 AI 100
FTD-minds/copilot-platform

Enterprise AI Copilot Platform — modular, model-agnostic enterprise AI copilots with RAG, multi-agent orchestration, MCP, and evaluation-gated CI/CD. Module 1: ERP Sync Reconciliation Copilot.

HTML 0 AI 100
KiwiMaddog2020/evaluating-claude

How to know if your AI is actually working

0 AI 100 個人 公開済 ↗
KeigoShimadaCC/llm-from-scratch

Mac-local LLM-from-scratch lab for building and evaluating a small decoder-only language model on Apple Silicon.

Python 0 AI 90 2 sig
zakahadi/llm-eval-cli

A CLI developer tool for running custom evaluation suites to OpenAI-compatible APIs (MiMo, Claude, GPT, etc.), outputting Markdown research reports + JSON traces. Useful for developers who want to benchmark models before production. Suitable for Data/Research + Dev tools, and a natural fit using the MiMo API + Hermes Agent workflow.

Python 0 AI 90
DiogoRibeiro7/agentic-qa-lab

Autonomous UI/game-testing agent: vision-language reasoning, browser control, action planning, failure recovery, and evaluation.

Python 0 AI 70 公開済 ↗
takehiro177/skill-evaluator

Evaluate Claude Code skills with cost-weighted A/B performance testing for Claude Code skills — one skill, or several combined. It measures real token cost and blind-judged output quality from live runs.

HTML 0 AI 70
UditBuilds/agentgrade

Open-source evaluation framework for AI customer support agents. 4-dimension LLM-as-judge (tone, accuracy, compliance, resolution), provider-agnostic, configurable rubric. Dogfooded on GlamShelf Twin.

Python 0 AI 70
CornElDos/ragas-lingua

Multilingual, RAGAS-compatible RAG evaluation with language-native judge prompts and Claude as the judge.

Python 0 AI 70
abhay23-AI/raggate

A thin, CI-gated evaluation gate for RAG & LLM systems — golden set, band-based pass/warn/fail gates, LLM-judge or heuristic scorers. pip install raggate

Python 0 AI 70 個人 公開済 ↗
Coinupbtc/zwell-bench

What: local LLM bakeoff harness (coding/vision/tools/agentic). For: objective model comparisons. How: ./setup.sh then ZWELL_BASE=… python bench_zwell.py --tag …

Python 0 AI 70 個人 公開済 ↗
ahmedEid1/forgejudge

Open, always-on leaderboard + CI gate for autonomous coding agents — every patch sandboxed, every run traced, every regression fails the build. $0 stack.

Python 0 AI 70 個人 公開済 ↗
josephruocco/posthog-mcp-mini

An MCP server that serves PostHog's JS/Web SDK docs as agent-optimized context, with an eval harness that measures it against a naive baseline.

Python 0 AI 70
rosscyking1115/redteam-foundry

LLM red-team evaluation harness — prompt injection, refusal, leakage and staleness, with cross-judge validation and attack-corpus audits.

Python 0 AI 70

RELATED 他のTopicも見る · 全Topicランキング →

#claude-code

1,564

#ai-agents

1,160

#llm

1,071

#claude

943

#python

806

#ai

741

#developer-tools

728

#mcp

727

#codex

521

集計対象: 各Repoの最新contentスナップショットの topics_json に小文字一致でマッチしたもの。 算出方法