#ai-evaluation
GitHub repositories that have self-applied the topic "ai-evaluation" — a creator-tagged metadata that surfaces how AI projects describe themselves.
REPOS Repos for #ai-evaluation (top 9 by stars)
Live, open-source benchmark for comparing AI coding agents on real GitHub issues
NavidBroumandfar/agent-behavior-evals-labPolicy-mapped evaluation lab for AI assistant behavior: approval gates, refusal boundaries, uncertainty handling, tool-use grounding, traces, and quality gates.
Nazim22/leadlineEvidence router & policy engine for coding agents — enforces source choice and proof-of-use through Claude Code hooks, including routes to MCP tools. Local-first; no LLM in the hook loop.
jawwad-ali/self-healing-ragSelf-healing RAG on n8n with LLM-as-judge eval, drift detection, investigator agent, and canary A/B deploys.
kwakusei1m-tech/AI-Evaluation-of-E-commerce-Chatbot-Responses-Using-Rubric-Based-Quality-ScoringAI Evaluation project analysing chatbot response quality in e-commerce using rubric scoring, error taxonomy, and semi-automated evaluation workflows.
api-evangelist/luminosaiLuminos.AI is an AI governance and evaluation platform that tests AI systems — classical machine learning, generative AI, and autonomous agents — for legal, regulatory, and reputational risk.
multivon-ai/multivon-mcpMCP server exposing multivon-eval + pdfhell as agent-callable tools. Drop into Claude Desktop, Cursor, Cline.
thangldw/ragopsOffline regression tests and explainable release gates for RAG systems and AI agents.
RudraO2/tokenbrawlA latency-fair LLM-vs-LLM fighting game benchmark. The engine blocks until an agent responds, so inference speed can't win a match — what's scarce is a per-match token bank that drains as a model thinks.
RELATED Other topics · full topics ranking →
#claude-code
1,555#ai-agents
1,156#llm
1,066#claude
936#python
802#ai
737#developer-tools
723#mcp
719#codex
517Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology