AI開発影響研究所 EN
← Topicランキング · 2026-08
GitHub TOPIC

#benchmarking

GitHub Topic「benchmarking」がついているAI関連リポジトリの集計。Topicはリポジトリ作者が自己申告するメタタグで、AI関連の文脈で何が「benchmarking」と分類されているかを可視化します。

12
タグ付Repo
150
TOP12合計★
2
AIツール痕跡あり
12
TOP表示数

REPOS #benchmarking のRepo (TOP 12 / Stars降順)

ahammadmejbah/Awesome-Datasets-Hub

A curated collection of datasets for Large Language Models (LLMs), covering medical AI, NLP, multimodal learning, instruction tuning, reasoning, code generation, and evaluation benchmarks.

146 AI 70 公開済 ↗
duclamvan/agent-hard-process

A Hermes Agent skill for turning painful fixes into replayable benchmarked workflows

Python 1 AI 100 個人 公開済 ↗
Dafenxz0/skillproof

Prove your Agent Skill works before you publish it.

HTML 1 AI 70
kizz-tech/agentic-evidence-lab

Tests exact agent interventions—skills, prompts, models, tools, and workflows—and publishes what changed, what failed, and what decision follows.

Python 1 AI 70 1 sig
manishklach/k3-inference-platform

Production-oriented Kimi K3 inference control plane with checkpoint release gates, MoE capacity planning, OpenAI gateway, NVL72 deployment, benchmarks, and observability.

Python 1 AI 60
cuiqi5656/agent-quality-benchmark

Evidence-first, reproducible quality benchmarks for AI agents.

Python 0 AI 100
Animesh352/llm-agent-bench

Multi-agent LLM orchestration and benchmarking toolkit -- planner-solver-reviewer workflows with provider abstraction and scoring

Python 0 AI 100
msradam/xk6-llm

Load test LLM inference servers with k6. TTFT, ITL, TPOT, goodput, cost, and energy metrics for any OpenAI-compatible server. Ships to Prometheus and Grafana.

Go 0 AI 90 公開済 ↗
lamb356/photonic-bench

Transparent benchmark cards, JSON artifacts, and visual tools for photonic AI accelerator energy/noise claims.

JavaScript 0 AI 70
jazzonaut/agentbench

Cross-platform diagnostics and live LLM performance benchmarking for Claude Code, Headroom, and system bottlenecks

Rust 0 AI 70
320exh/prompt-diff

A fast, git-native CLI & local Web UI to version, diff tokens/costs, and benchmark LLM system prompts across local & cloud models.

Go 0 AI 70
satwiksps/scaffoldscope

Controlled coding-agent harness ablations with auditable traces, reproducible evidence bundles, and SWE-bench interoperability.

Python 0 AI 70 個人 1 sig 公開済 ↗

RELATED 他のTopicも見る · 全Topicランキング →

#claude-code

1,564

#ai-agents

1,160

#llm

1,071

#claude

943

#python

806

#ai

741

#developer-tools

728

#mcp

727

#codex

521

集計対象: 各Repoの最新contentスナップショットの topics_json に小文字一致でマッチしたもの。 算出方法