#evaluation
GitHub repositories that have self-applied the topic "evaluation" — a creator-tagged metadata that surfaces how AI projects describe themselves.
REPOS Repos for #evaluation (top 26 by stars)
Release repository for agent benchmark evidence-reporting artifacts and reproduction workflows.
forrestbthomas/pi-harnessProvider-agnostic coding-agent harness + DeepEval evaluation suite
vahit19/longhaul-benchLong-horizon reliability benchmark for industrial edge agents - do self-improving agents get better or corrupt over 1000+ episodes on constrained hardware?
zhanghanbo668-art/embodied-eval-loopSoftware-only embodied AI evaluation stack for benchmark trajectories, ROS-compatible logs, replay, failure analysis, and comparison reports.
DmitryDmitriadi/llm-observability-evaluationCase study: unified observability + evaluation for automation and conversational AI agents. Versioned error catalog, multi-signal confidence, cost meters.
maddog22coder/agent-evalOpen-source multilingual evaluation toolkit for conversational AI agents.
iiTzAK/accounting-copilotFine-tuning an open 7B LLM into a GST question-answering assistant, with a reproducible eval harness measuring the gain over the base model.
auraoneai/rubric-studio-openAuthor, validate, preview, calibrate, diff, and export criterion-level AI evaluation rubrics.
beausome/bullshit-benchAn AI leaderboard scoring chatbots on how much they bullshit vs. give honest answers to unanswerable questions. Works with paid APIs and free/local models.
wwjorker/devmindAI knowledge base & RAG QA system with gold-label retrieval evaluation (Spring Boot 3 + Vue 3): hybrid keyword/vector search, RRF fusion, rerank, four-way Hit@3/MRR comparison
clouatre-labs/reversibility-benchmarkSupplementary materials for Reversibility as a Safety Gate for Agentic Actions
chohyerinn/mini-agent-harnessEvaluation harness for coding agents - repeated runs, pass@k, bootstrap/McNemar significance, tamper detection, cost. Single vs Planner-Coder-Reviewer multi-agent on CLOVA.
Hal-Hanami/incident-triage-agentRead-only first-pass incident triage on the Claude Agent SDK + MCP: classifies an alert, retrieves the matching runbook, and proposes a cited first response — or abstains to a human. Measured over four runs: 100% abstention with 0 missed escalations, ~$0.014 per incident.
FTD-minds/copilot-platformEnterprise AI Copilot Platform — modular, model-agnostic enterprise AI copilots with RAG, multi-agent orchestration, MCP, and evaluation-gated CI/CD. Module 1: ERP Sync Reconciliation Copilot.
KiwiMaddog2020/evaluating-claudeHow to know if your AI is actually working
KeigoShimadaCC/llm-from-scratchMac-local LLM-from-scratch lab for building and evaluating a small decoder-only language model on Apple Silicon.
zakahadi/llm-eval-cliA CLI developer tool for running custom evaluation suites to OpenAI-compatible APIs (MiMo, Claude, GPT, etc.), outputting Markdown research reports + JSON traces. Useful for developers who want to benchmark models before production. Suitable for Data/Research + Dev tools, and a natural fit using the MiMo API + Hermes Agent workflow.
DiogoRibeiro7/agentic-qa-labAutonomous UI/game-testing agent: vision-language reasoning, browser control, action planning, failure recovery, and evaluation.
takehiro177/skill-evaluatorEvaluate Claude Code skills with cost-weighted A/B performance testing for Claude Code skills — one skill, or several combined. It measures real token cost and blind-judged output quality from live runs.
UditBuilds/agentgradeOpen-source evaluation framework for AI customer support agents. 4-dimension LLM-as-judge (tone, accuracy, compliance, resolution), provider-agnostic, configurable rubric. Dogfooded on GlamShelf Twin.
CornElDos/ragas-linguaMultilingual, RAGAS-compatible RAG evaluation with language-native judge prompts and Claude as the judge.
abhay23-AI/raggateA thin, CI-gated evaluation gate for RAG & LLM systems — golden set, band-based pass/warn/fail gates, LLM-judge or heuristic scorers. pip install raggate
Coinupbtc/zwell-benchWhat: local LLM bakeoff harness (coding/vision/tools/agentic). For: objective model comparisons. How: ./setup.sh then ZWELL_BASE=… python bench_zwell.py --tag …
ahmedEid1/forgejudgeOpen, always-on leaderboard + CI gate for autonomous coding agents — every patch sandboxed, every run traced, every regression fails the build. $0 stack.
josephruocco/posthog-mcp-miniAn MCP server that serves PostHog's JS/Web SDK docs as agent-optimized context, with an eval harness that measures it against a naive baseline.
rosscyking1115/redteam-foundryLLM red-team evaluation harness — prompt injection, refusal, leakage and staleness, with cross-judge validation and attack-corpus audits.
RELATED Other topics · full topics ranking →
#claude-code
1,564#ai-agents
1,160#llm
1,071#claude
943#python
806#ai
741#developer-tools
728#mcp
727#codex
521Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology