#llm-evaluation
GitHub repositories that have self-applied the topic "llm-evaluation" — a creator-tagged metadata that surfaces how AI projects describe themselves.
REPOS Repos for #llm-evaluation (top 31 by stars)
A curated collection of datasets for Large Language Models (LLMs), covering medical AI, NLP, multimodal learning, instruction tuning, reasoning, code generation, and evaluation benchmarks.
prat3ik/evalbotEvalBot — local-first chatbot security & quality evaluation (FastAPI + Next.js). Evaluate chatbot answers against your own docs & guidelines with ML/NLP + AI-judge scoring. Apache-2.0.
xray-eval/xrayAn open-source debugger for voice agent workflows
Johna2an/critical-thinkingOpen-source Agent Skill for Claude Code that makes Claude reason with explicit probabilities, base rates, and falsifiers. Distilled from 51 books; first in 18/18 blind passes vs baseline Claude and GPT-5.5, and benchmarked against 14 rival skills and prompt techniques with full cost telemetry (round 5).
DmitryDmitriadi/llm-observability-evaluationCase study: unified observability + evaluation for automation and conversational AI agents. Versioned error catalog, multi-signal confidence, cost meters.
ttxs69/coding-agent-evalPublic, reproducible benchmark of CLI coding agents (Claude Code, Codex, Aider) on SWE-bench Verified. Live leaderboard: https://ttxs69.github.io/coding-agent-eval/
kizz-tech/agentic-evidence-labTests exact agent interventions—skills, prompts, models, tools, and workflows—and publishes what changed, what failed, and what decision follows.
beausome/bullshit-benchAn AI leaderboard scoring chatbots on how much they bullshit vs. give honest answers to unanswerable questions. Works with paid APIs and free/local models.
Parthu-M/grounded-ops-ragProduction-style RAG operations console with grounded retrieval, document ingestion, evaluation, cost analysis, and bias-aware LLM judging.
Anmedrif/ai-career-osA privacy-aware multi-agent career management system with structured source governance, human approval workflows, and LLM output evaluation.
ktech7moon/agent-habitatAudit-grade multi-agent orchestration. Fabrication-resistance enforced as a validated contract, not a hope.
Lohithr123/rf-copilotRF front-end analysis engine and a 60-problem LLM benchmark. Grounded small model ≈ ungrounded frontier model.
cuiqi5656/agent-quality-benchmarkEvidence-first, reproducible quality benchmarks for AI agents.
4b8wsfdk7y-cloud/cgss-gender-attitude-llm-auditEstimand-indexed audit of LLM synthetic respondents against CGSS benchmarks
tonydzi/agent-runtime-integrity-benchIntegrity scenarios for agent runtimes, distilled from real fleet incidents — deterministic fault-injection checks against real SDKs
Evangelidis91/llm-feature-engineeringEmpirical study: do LLM-generated features improve tabular ML models? 10 LLMs, 3 datasets, 3 algorithms, 219 experiments.
ejazfahil/bayesian-llm-evalHierarchical Bayesian evaluation of LLM/VLM outputs — posterior credible intervals & partial pooling over any eval's item-level scores (Bambi/PyMC).
kalyvask/ai-safety-osSame agent action, safe or dangerous by context: email team vs investor; delete scratch vs prod config. A safety layer routes each action: auto, confirm, escalate, block. A reading model auto-executes 0% of unsafe actions vs a baseline's 38% (McNemar p<0.001). Plus a runtime loop: an agent earns or loses autonomy with each counterparty over time.
maximizeGPT/claude-eval-harnessRegression-diff eval harness for Anthropic tool-use agents that surfaces LLM-judge reasoning drift, not just pass/fail flips
Hert4/LLM-Certainty-ConsistencyBackend-agnostic black-box hallucination & RAG-faithfulness detection for LLMs — Probabilistic Certainty & Consistency (arXiv:2601.02574). Works on MLX / OpenAI / vLLM via token logprobs; no model internals, no training.
Samuelmartinezduran/agentevalFramework open source para evaluar agentes LLM con function calling / tool use. Scoring por dimensión (tool accuracy, response quality, safety) + CLI, API y dashboard.
mukimudeen76-ops/arena⚡ Chatbot Arena — battle LLMs side-by-side. Elo leaderboard, LLM-as-judge, PII redaction, guardrails, routing & distillation. 100% browser-based.
satwiksps/scaffoldscopeControlled coding-agent harness ablations with auditable traces, reproducible evidence bundles, and SWE-bench interoperability.
sammy995/fiduciaryCan your AI act as a fiduciary inside a regulated bank? A deployment-readiness benchmark that drops LLMs into a synthetic regulated organization as an employee and audits the behavior.
multivon-ai/multivon-mcpMCP server exposing multivon-eval + pdfhell as agent-callable tools. Drop into Claude Desktop, Cursor, Cline.
takehiro177/skill-evaluatorEvaluate Claude Code skills with cost-weighted A/B performance testing for Claude Code skills — one skill, or several combined. It measures real token cost and blind-judged output quality from live runs.
Jott2121/sabotDo your agent pipeline's own checks catch planted faults? Measured on LangGraph, CrewAI and AutoGen: median 16.7%. One prompt-level change takes it to 55.0%. Pre-registered spec, Apache-2.0 harness, every raw trace published.
ahmedEid1/forgejudgeOpen, always-on leaderboard + CI gate for autonomous coding agents — every patch sandboxed, every run traced, every regression fails the build. $0 stack.
lord-Rheagar/prompt-labRegression testing and automated optimization for LLM prompts. Run a YAML test suite, score outputs with an LLM judge, and auto-improve across 4 strategies. Works with any OpenAI-compatible API — including local models.
NullLabTests/grounded_evolutionEvaluator-grounded autonomous prompt evolution. Genetic algorithms + runtime validation + meta-evolution to evolve prompts that generate better AI agent code. 150 generations, 862/1000 best score.
jleonceo/orquestacion-enjambres-iaOrquestación de IA multiagente: un registro de agentes autogenerado + una evaluación a ciegas que demuestra que el enrutado acierta el agente correcto (27/27, sin regresiones)
RELATED Other topics · full topics ranking →
#claude-code
1,555#ai-agents
1,156#llm
1,066#claude
936#python
802#ai
737#developer-tools
723#mcp
719#codex
517Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology