AI開発影響研究所 EN
← Topicランキング · 2026-08
GitHub TOPIC

#llm-as-judge

GitHub Topic「llm-as-judge」がついているAI関連リポジトリの集計。Topicはリポジトリ作者が自己申告するメタタグで、AI関連の文脈で何が「llm-as-judge」と分類されているかを可視化します。

10
タグ付Repo
5
TOP10合計★
3
AIツール痕跡あり
10
TOP表示数

REPOS #llm-as-judge のRepo (TOP 10 / Stars降順)

Anant-pentester/adversarial-gatekeeper

AI Content Guard 2026: Adversarial Fact-Check Gate for Error-Free Code & Copy

HTML 2 AI 70
Johna2an/critical-thinking

Open-source Agent Skill for Claude Code that makes Claude reason with explicit probabilities, base rates, and falsifiers. Distilled from 51 books; first in 18/18 blind passes vs baseline Claude and GPT-5.5, and benchmarked against 14 rival skills and prompt techniques with full cost telemetry (round 5).

HTML 2 AI 70 2 sig
jawwad-ali/self-healing-rag

Self-healing RAG on n8n with LLM-as-judge eval, drift detection, investigator agent, and canary A/B deploys.

TypeScript 1 AI 100 個人 2 sig 公開済 ↗
dbystrova26/aparthotel-ai-strategy

LangChain agent on 119k real hotel bookings — cancellation prediction, dynamic pricing & guest automation for a pan-European aparthotel chain. Tableau · n8n · LLM-as-judge eval 4.6/5

Python 0 AI 100
KevinLimias/eval-driven-metrics-playbook

The Ultimate Eval-Driven Claude Plugin Suite for Product Teams 2026 - Verified Toolkit

HTML 0 AI 70
takehiro177/skill-evaluator

Evaluate Claude Code skills with cost-weighted A/B performance testing for Claude Code skills — one skill, or several combined. It measures real token cost and blind-judged output quality from live runs.

HTML 0 AI 70
markudevelop/outcome-fusion-principia

Model fusion for Claude Code — one model builds, a second model (DeepSeek) judges the results. First-principles missions, a proof ledger, and a release gate that won't let your agent stop early.

Python 0 AI 70
KranzL/jje

A generator-critic harness for Claude Code: Planner / Executor / a parallel panel of tool-backed jurors / a routing Judge, with candidate isolation and CI as the final gate.

Shell 0 AI 70 1 sig
sammy995/fiduciary

Can your AI act as a fiduciary inside a regulated bank? A deployment-readiness benchmark that drops LLMs into a synthetic regulated organization as an employee and audits the behavior.

Python 0 AI 70
guimitestai/sdk

🧪 Guimí Test AI — SDK Python para testes, observabilidade e conformidade de LLMs. LLM-as-Judge, Red-Teaming, LGPD/EU AI Act.

Python 0 AI 70

RELATED 他のTopicも見る · 全Topicランキング →

#claude-code

1,564

#ai-agents

1,160

#llm

1,071

#claude

943

#python

806

#ai

741

#developer-tools

728

#mcp

727

#codex

521

集計対象: 各Repoの最新contentスナップショットの topics_json に小文字一致でマッチしたもの。 算出方法