#benchmarking
GitHub Topic「benchmarking」がついているAI関連リポジトリの集計。Topicはリポジトリ作者が自己申告するメタタグで、AI関連の文脈で何が「benchmarking」と分類されているかを可視化します。
REPOS #benchmarking のRepo (TOP 12 / Stars降順)
A curated collection of datasets for Large Language Models (LLMs), covering medical AI, NLP, multimodal learning, instruction tuning, reasoning, code generation, and evaluation benchmarks.
duclamvan/agent-hard-processA Hermes Agent skill for turning painful fixes into replayable benchmarked workflows
Dafenxz0/skillproofProve your Agent Skill works before you publish it.
kizz-tech/agentic-evidence-labTests exact agent interventions—skills, prompts, models, tools, and workflows—and publishes what changed, what failed, and what decision follows.
manishklach/k3-inference-platformProduction-oriented Kimi K3 inference control plane with checkpoint release gates, MoE capacity planning, OpenAI gateway, NVL72 deployment, benchmarks, and observability.
cuiqi5656/agent-quality-benchmarkEvidence-first, reproducible quality benchmarks for AI agents.
Animesh352/llm-agent-benchMulti-agent LLM orchestration and benchmarking toolkit -- planner-solver-reviewer workflows with provider abstraction and scoring
msradam/xk6-llmLoad test LLM inference servers with k6. TTFT, ITL, TPOT, goodput, cost, and energy metrics for any OpenAI-compatible server. Ships to Prometheus and Grafana.
lamb356/photonic-benchTransparent benchmark cards, JSON artifacts, and visual tools for photonic AI accelerator energy/noise claims.
jazzonaut/agentbenchCross-platform diagnostics and live LLM performance benchmarking for Claude Code, Headroom, and system bottlenecks
320exh/prompt-diffA fast, git-native CLI & local Web UI to version, diff tokens/costs, and benchmark LLM system prompts across local & cloud models.
satwiksps/scaffoldscopeControlled coding-agent harness ablations with auditable traces, reproducible evidence bundles, and SWE-bench interoperability.
RELATED 他のTopicも見る · 全Topicランキング →
#claude-code
1,564#ai-agents
1,160#llm
1,071#claude
943#python
806#ai
741#developer-tools
728#mcp
727#codex
521集計対象: 各Repoの最新contentスナップショットの topics_json に小文字一致でマッチしたもの。 算出方法