#llm-evals
GitHub Topic「llm-evals」がついているAI関連リポジトリの集計。Topicはリポジトリ作者が自己申告するメタタグで、AI関連の文脈で何が「llm-evals」と分類されているかを可視化します。
REPOS #llm-evals のRepo (TOP 5 / Stars降順)
Playbook for PMs shipping AI products with PRDs, evals, HITL, launch gates, cost, and observability.
AnonZ7/agentic-rag-eval-harnessProduction-shaped agentic RAG: LangGraph plan->act->verify agent + hybrid retrieval + guardrails + an eval gate in CI. Provider-agnostic; runs offline with no keys.
ch040602/agent-prompt-injection-zooSource-backed archive of agent prompt-injection incidents, research records, trust-boundary patterns, schemas, and sanitized defensive summaries.
Redsf/rag-internal-knowledge-chatbotSlack-native internal knowledge chatbot with nightly-reindexed RAG (Pinecone) and source-cited answers. Reference build behind a 50% onboarding-time-reduction case study.
mborges-dev/extraction-evalsReproducible benchmark for LLM-based structured extraction from documents. Compare Claude / GPT / Gemini / open-weight on the same task with cost + latency tracking.
RELATED 他のTopicも見る · 全Topicランキング →
#claude-code
1,555#ai-agents
1,156#llm
1,066#claude
936#python
802#ai
737#developer-tools
723#mcp
719#codex
517集計対象: 各Repoの最新contentスナップショットの topics_json に小文字一致でマッチしたもの。 算出方法