AI Dev Impact Lab JA
← Topics ranking · 2026-08
GITHUB TOPIC

#ai-evaluation

GitHub repositories that have self-applied the topic "ai-evaluation" — a creator-tagged metadata that surfaces how AI projects describe themselves.

9
tagged repos
11
top 9 stars
4
with tool sigs
9
shown

REPOS Repos for #ai-evaluation (top 9 by stars)

linny006/agent-eval-harness

Live, open-source benchmark for comparing AI coding agents on real GitHub issues

Python 6 AI 90
NavidBroumandfar/agent-behavior-evals-lab

Policy-mapped evaluation lab for AI assistant behavior: approval gates, refusal boundaries, uncertainty handling, tool-use grounding, traces, and quality gates.

Python 2 AI 100 Solo 2 sig live ↗
Nazim22/leadline

Evidence router & policy engine for coding agents — enforces source choice and proof-of-use through Claude Code hooks, including routes to MCP tools. Local-first; no LLM in the hook loop.

JavaScript 2 AI 70 1 sig
jawwad-ali/self-healing-rag

Self-healing RAG on n8n with LLM-as-judge eval, drift detection, investigator agent, and canary A/B deploys.

TypeScript 1 AI 100 Solo 2 sig live ↗
kwakusei1m-tech/AI-Evaluation-of-E-commerce-Chatbot-Responses-Using-Rubric-Based-Quality-Scoring

AI Evaluation project analysing chatbot response quality in e-commerce using rubric scoring, error taxonomy, and semi-automated evaluation workflows.

Jupyter Notebook 0 AI 100
api-evangelist/luminosai

Luminos.AI is an AI governance and evaluation platform that tests AI systems — classical machine learning, generative AI, and autonomous agents — for legal, regulatory, and reputational risk.

0 AI 100
multivon-ai/multivon-mcp

MCP server exposing multivon-eval + pdfhell as agent-callable tools. Drop into Claude Desktop, Cursor, Cline.

Python 0 AI 70 live ↗
thangldw/ragops

Offline regression tests and explainable release gates for RAG systems and AI agents.

Python 0 AI 70
RudraO2/tokenbrawl

A latency-fair LLM-vs-LLM fighting game benchmark. The engine blocks until an agent responds, so inference speed can't win a match — what's scarce is a per-match token bank that drains as a model thinks.

TypeScript 0 AI 70 1 sig

RELATED Other topics · full topics ranking →

#claude-code

1,555

#ai-agents

1,156

#llm

1,066

#claude

936

#python

802

#ai

737

#developer-tools

723

#mcp

719

#codex

517

Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology