AI Dev Impact Lab JA
← Topics ranking · 2026-08
GITHUB TOPIC

#llm-evaluation

GitHub repositories that have self-applied the topic "llm-evaluation" — a creator-tagged metadata that surfaces how AI projects describe themselves.

31
tagged repos
158
top 31 stars
8
with tool sigs
31
shown

REPOS Repos for #llm-evaluation (top 31 by stars)

ahammadmejbah/Awesome-Datasets-Hub

A curated collection of datasets for Large Language Models (LLMs), covering medical AI, NLP, multimodal learning, instruction tuning, reasoning, code generation, and evaluation benchmarks.

146 AI 70 live ↗
prat3ik/evalbot

EvalBot — local-first chatbot security & quality evaluation (FastAPI + Next.js). Evaluate chatbot answers against your own docs & guidelines with ML/NLP + AI-judge scoring. Apache-2.0.

Python 3 AI 70
xray-eval/xray

An open-source debugger for voice agent workflows

TypeScript 3 AI 70 1 sig live ↗
Johna2an/critical-thinking

Open-source Agent Skill for Claude Code that makes Claude reason with explicit probabilities, base rates, and falsifiers. Distilled from 51 books; first in 18/18 blind passes vs baseline Claude and GPT-5.5, and benchmarked against 14 rival skills and prompt techniques with full cost telemetry (round 5).

HTML 2 AI 70 2 sig
DmitryDmitriadi/llm-observability-evaluation

Case study: unified observability + evaluation for automation and conversational AI agents. Versioned error catalog, multi-signal confidence, cost meters.

1 AI 100
ttxs69/coding-agent-eval

Public, reproducible benchmark of CLI coding agents (Claude Code, Codex, Aider) on SWE-bench Verified. Live leaderboard: https://ttxs69.github.io/coding-agent-eval/

Python 1 AI 90 1 sig
kizz-tech/agentic-evidence-lab

Tests exact agent interventions—skills, prompts, models, tools, and workflows—and publishes what changed, what failed, and what decision follows.

Python 1 AI 70 1 sig
beausome/bullshit-bench

An AI leaderboard scoring chatbots on how much they bullshit vs. give honest answers to unanswerable questions. Works with paid APIs and free/local models.

Python 1 AI 70
Parthu-M/grounded-ops-rag

Production-style RAG operations console with grounded retrieval, document ingestion, evaluation, cost analysis, and bias-aware LLM judging.

Python 0 AI 100 Solo live ↗
Anmedrif/ai-career-os

A privacy-aware multi-agent career management system with structured source governance, human approval workflows, and LLM output evaluation.

0 AI 100
ktech7moon/agent-habitat

Audit-grade multi-agent orchestration. Fabrication-resistance enforced as a validated contract, not a hope.

Python 0 AI 100 1 sig
Lohithr123/rf-copilot

RF front-end analysis engine and a 60-problem LLM benchmark. Grounded small model ≈ ungrounded frontier model.

Python 0 AI 100
cuiqi5656/agent-quality-benchmark

Evidence-first, reproducible quality benchmarks for AI agents.

Python 0 AI 100
4b8wsfdk7y-cloud/cgss-gender-attitude-llm-audit

Estimand-indexed audit of LLM synthetic respondents against CGSS benchmarks

TeX 0 AI 100
tonydzi/agent-runtime-integrity-bench

Integrity scenarios for agent runtimes, distilled from real fleet incidents — deterministic fault-injection checks against real SDKs

Python 0 AI 100 1 sig
Evangelidis91/llm-feature-engineering

Empirical study: do LLM-generated features improve tabular ML models? 10 LLMs, 3 datasets, 3 algorithms, 219 experiments.

Jupyter Notebook 0 AI 100
ejazfahil/bayesian-llm-eval

Hierarchical Bayesian evaluation of LLM/VLM outputs — posterior credible intervals & partial pooling over any eval's item-level scores (Bambi/PyMC).

Python 0 AI 100
kalyvask/ai-safety-os

Same agent action, safe or dangerous by context: email team vs investor; delete scratch vs prod config. A safety layer routes each action: auto, confirm, escalate, block. A reading model auto-executes 0% of unsafe actions vs a baseline's 38% (McNemar p<0.001). Plus a runtime loop: an agent earns or loses autonomy with each counterparty over time.

Python 0 AI 100
maximizeGPT/claude-eval-harness

Regression-diff eval harness for Anthropic tool-use agents that surfaces LLM-judge reasoning drift, not just pass/fail flips

Python 0 AI 90
Hert4/LLM-Certainty-Consistency

Backend-agnostic black-box hallucination & RAG-faithfulness detection for LLMs — Probabilistic Certainty & Consistency (arXiv:2601.02574). Works on MLX / OpenAI / vLLM via token logprobs; no model internals, no training.

Python 0 AI 90
Samuelmartinezduran/agenteval

Framework open source para evaluar agentes LLM con function calling / tool use. Scoring por dimensión (tool accuracy, response quality, safety) + CLI, API y dashboard.

Python 0 AI 70
mukimudeen76-ops/arena

⚡ Chatbot Arena — battle LLMs side-by-side. Elo leaderboard, LLM-as-judge, PII redaction, guardrails, routing & distillation. 100% browser-based.

JavaScript 0 AI 70 Solo live ↗
satwiksps/scaffoldscope

Controlled coding-agent harness ablations with auditable traces, reproducible evidence bundles, and SWE-bench interoperability.

Python 0 AI 70 Solo 1 sig live ↗
sammy995/fiduciary

Can your AI act as a fiduciary inside a regulated bank? A deployment-readiness benchmark that drops LLMs into a synthetic regulated organization as an employee and audits the behavior.

Python 0 AI 70
multivon-ai/multivon-mcp

MCP server exposing multivon-eval + pdfhell as agent-callable tools. Drop into Claude Desktop, Cursor, Cline.

Python 0 AI 70 live ↗
takehiro177/skill-evaluator

Evaluate Claude Code skills with cost-weighted A/B performance testing for Claude Code skills — one skill, or several combined. It measures real token cost and blind-judged output quality from live runs.

HTML 0 AI 70
Jott2121/sabot

Do your agent pipeline's own checks catch planted faults? Measured on LangGraph, CrewAI and AutoGen: median 16.7%. One prompt-level change takes it to 55.0%. Pre-registered spec, Apache-2.0 harness, every raw trace published.

Python 0 AI 70 Solo live ↗
ahmedEid1/forgejudge

Open, always-on leaderboard + CI gate for autonomous coding agents — every patch sandboxed, every run traced, every regression fails the build. $0 stack.

Python 0 AI 70 Solo live ↗
lord-Rheagar/prompt-lab

Regression testing and automated optimization for LLM prompts. Run a YAML test suite, score outputs with an LLM judge, and auto-improve across 4 strategies. Works with any OpenAI-compatible API — including local models.

Python 0 AI 60
NullLabTests/grounded_evolution

Evaluator-grounded autonomous prompt evolution. Genetic algorithms + runtime validation + meta-evolution to evolve prompts that generate better AI agent code. 150 generations, 862/1000 best score.

Python 0 AI 60 Solo 1 sig live ↗
jleonceo/orquestacion-enjambres-ia

Orquestación de IA multiagente: un registro de agentes autogenerado + una evaluación a ciegas que demuestra que el enrutado acierta el agente correcto (27/27, sin regresiones)

Python 0 AI 45

RELATED Other topics · full topics ranking →

#claude-code

1,555

#ai-agents

1,156

#llm

1,066

#claude

936

#python

802

#ai

737

#developer-tools

723

#mcp

719

#codex

517

Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology