#llm-as-judge
GitHub repositories that have self-applied the topic "llm-as-judge" — a creator-tagged metadata that surfaces how AI projects describe themselves.
REPOS Repos for #llm-as-judge (top 10 by stars)
AI Content Guard 2026: Adversarial Fact-Check Gate for Error-Free Code & Copy
Johna2an/critical-thinkingOpen-source Agent Skill for Claude Code that makes Claude reason with explicit probabilities, base rates, and falsifiers. Distilled from 51 books; first in 18/18 blind passes vs baseline Claude and GPT-5.5, and benchmarked against 14 rival skills and prompt techniques with full cost telemetry (round 5).
jawwad-ali/self-healing-ragSelf-healing RAG on n8n with LLM-as-judge eval, drift detection, investigator agent, and canary A/B deploys.
dbystrova26/aparthotel-ai-strategyLangChain agent on 119k real hotel bookings — cancellation prediction, dynamic pricing & guest automation for a pan-European aparthotel chain. Tableau · n8n · LLM-as-judge eval 4.6/5
KevinLimias/eval-driven-metrics-playbookThe Ultimate Eval-Driven Claude Plugin Suite for Product Teams 2026 - Verified Toolkit
takehiro177/skill-evaluatorEvaluate Claude Code skills with cost-weighted A/B performance testing for Claude Code skills — one skill, or several combined. It measures real token cost and blind-judged output quality from live runs.
markudevelop/outcome-fusion-principiaModel fusion for Claude Code — one model builds, a second model (DeepSeek) judges the results. First-principles missions, a proof ledger, and a release gate that won't let your agent stop early.
KranzL/jjeA generator-critic harness for Claude Code: Planner / Executor / a parallel panel of tool-backed jurors / a routing Judge, with candidate isolation and CI as the final gate.
sammy995/fiduciaryCan your AI act as a fiduciary inside a regulated bank? A deployment-readiness benchmark that drops LLMs into a synthetic regulated organization as an employee and audits the behavior.
guimitestai/sdk🧪 Guimí Test AI — SDK Python para testes, observabilidade e conformidade de LLMs. LLM-as-Judge, Red-Teaming, LGPD/EU AI Act.
RELATED Other topics · full topics ranking →
#claude-code
1,555#ai-agents
1,156#llm
1,066#claude
936#python
802#ai
737#developer-tools
723#mcp
719#codex
517Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology