AI Dev Impact Lab JA
← Topics ranking · 2026-08
GITHUB TOPIC

#llm-as-judge

GitHub repositories that have self-applied the topic "llm-as-judge" — a creator-tagged metadata that surfaces how AI projects describe themselves.

10
tagged repos
5
top 10 stars
3
with tool sigs
10
shown

REPOS Repos for #llm-as-judge (top 10 by stars)

Anant-pentester/adversarial-gatekeeper

AI Content Guard 2026: Adversarial Fact-Check Gate for Error-Free Code & Copy

HTML 2 AI 70
Johna2an/critical-thinking

Open-source Agent Skill for Claude Code that makes Claude reason with explicit probabilities, base rates, and falsifiers. Distilled from 51 books; first in 18/18 blind passes vs baseline Claude and GPT-5.5, and benchmarked against 14 rival skills and prompt techniques with full cost telemetry (round 5).

HTML 2 AI 70 2 sig
jawwad-ali/self-healing-rag

Self-healing RAG on n8n with LLM-as-judge eval, drift detection, investigator agent, and canary A/B deploys.

TypeScript 1 AI 100 Solo 2 sig live ↗
dbystrova26/aparthotel-ai-strategy

LangChain agent on 119k real hotel bookings — cancellation prediction, dynamic pricing & guest automation for a pan-European aparthotel chain. Tableau · n8n · LLM-as-judge eval 4.6/5

Python 0 AI 100
KevinLimias/eval-driven-metrics-playbook

The Ultimate Eval-Driven Claude Plugin Suite for Product Teams 2026 - Verified Toolkit

HTML 0 AI 70
takehiro177/skill-evaluator

Evaluate Claude Code skills with cost-weighted A/B performance testing for Claude Code skills — one skill, or several combined. It measures real token cost and blind-judged output quality from live runs.

HTML 0 AI 70
markudevelop/outcome-fusion-principia

Model fusion for Claude Code — one model builds, a second model (DeepSeek) judges the results. First-principles missions, a proof ledger, and a release gate that won't let your agent stop early.

Python 0 AI 70
KranzL/jje

A generator-critic harness for Claude Code: Planner / Executor / a parallel panel of tool-backed jurors / a routing Judge, with candidate isolation and CI as the final gate.

Shell 0 AI 70 1 sig
sammy995/fiduciary

Can your AI act as a fiduciary inside a regulated bank? A deployment-readiness benchmark that drops LLMs into a synthetic regulated organization as an employee and audits the behavior.

Python 0 AI 70
guimitestai/sdk

🧪 Guimí Test AI — SDK Python para testes, observabilidade e conformidade de LLMs. LLM-as-Judge, Red-Teaming, LGPD/EU AI Act.

Python 0 AI 70

RELATED Other topics · full topics ranking →

#claude-code

1,555

#ai-agents

1,156

#llm

1,066

#claude

936

#python

802

#ai

737

#developer-tools

723

#mcp

719

#codex

517

Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology