AI開発影響研究所 EN
← Topicランキング · 2026-08
GitHub TOPIC

#benchmark

GitHub Topic「benchmark」がついているAI関連リポジトリの集計。Topicはリポジトリ作者が自己申告するメタタグで、AI関連の文脈で何が「benchmark」と分類されているかを可視化します。

34
タグ付Repo
187
TOP34合計★
6
AIツール痕跡あり
34
TOP表示数

REPOS #benchmark のRepo (TOP 34 / Stars降順)

ahammadmejbah/Awesome-Datasets-Hub

A curated collection of datasets for Large Language Models (LLMs), covering medical AI, NLP, multimodal learning, instruction tuning, reasoning, code generation, and evaluation benchmarks.

146 AI 70 公開済 ↗
protorikis/rikis

Lightweight agent that pulls and runs LLM benchmarks from Protorikis Bench

Python 15 AI 70 個人 公開済 ↗
linny006/agent-eval-harness

Live, open-source benchmark for comparing AI coding agents on real GitHub issues

Python 6 AI 90
AgentBenchAudit/evidence-bounds

Release repository for agent benchmark evidence-reporting artifacts and reproduction workflows.

HTML 5 AI 70 公開済 ↗
kevinpeckham/barkup-bench

Pre-registered benchmark: HTML dialect + whole-tree rewrite (barkup) vs JSON, granular mutation tools, and two patch dialects, for LLM agents editing typed trees. 9,600 scored runs across 4 models. Findings in REPORT.md.

TypeScript 2 AI 70 個人 1 sig 公開済 ↗
vahit19/longhaul-bench

Long-horizon reliability benchmark for industrial edge agents - do self-improving agents get better or corrupt over 1000+ episodes on constrained hardware?

Python 2 AI 70
bmendonca3/authzbench-saas

Benchmark for AI agents proving multi-tenant SaaS authorization bugs

Python 2 AI 70 1 sig
Mazha0309/the-evil-repository

An evidence-hostile, container-isolated benchmark and behavior analysis platform for long-horizon AI software agents.

Python 2 AI 70
FU-max-boop/statebind-guard

Catch visible-but-unbound coding-agent handoffs: CLI + GitHub Action with proof, policy gates, SARIF, HTML, and benchmark cards.

Python 2 AI 60 個人 公開済 ↗
awdemos/toks-bench

Reproducible token-throughput benchmark for OpenAI-compatible LLM servers, tuned for NVIDIA Spark and GB10 inference.

Python 2 AI 50
ttxs69/coding-agent-eval

Public, reproducible benchmark of CLI coding agents (Claude Code, Codex, Aider) on SWE-bench Verified. Live leaderboard: https://ttxs69.github.io/coding-agent-eval/

Python 1 AI 90 1 sig
beausome/bullshit-bench

An AI leaderboard scoring chatbots on how much they bullshit vs. give honest answers to unanswerable questions. Works with paid APIs and free/local models.

Python 1 AI 70
clouatre-labs/reversibility-benchmark

Supplementary materials for Reversibility as a Safety Gate for Agentic Actions

Python 1 AI 45 1 sig
Vedansh5545/llm-shieldbench

Trustworthy AI evaluation tool for testing chatbot safety, reliability, hallucination behavior, privacy risk, and instruction-following quality.

Python 0 AI 100
Lohithr123/rf-copilot

RF front-end analysis engine and a 60-problem LLM benchmark. Grounded small model ≈ ungrounded frontier model.

Python 0 AI 100
tonydzi/agent-runtime-integrity-bench

Integrity scenarios for agent runtimes, distilled from real fleet incidents — deterministic fault-injection checks against real SDKs

Python 0 AI 100 1 sig
achenna-source/sagin-agent-bench

Measuring and verifying LLM decision agents under orbital dynamics, delay, and intermittency: a reproducible testbed, failure-mode taxonomy, and in-chain verification.

Python 0 AI 100
ZihanZheng2000/reservoir-multi-agent-deliberation

Synthetic benchmark for role-limited multi-agent deliberation in reservoir construction decision support

Python 0 AI 100
chohyerinn/mini-agent-harness

Evaluation harness for coding agents - repeated runs, pass@k, bootstrap/McNemar significance, tamper detection, cost. Single vs Planner-Coder-Reviewer multi-agent on CLOVA.

Python 0 AI 100
BELYAGOUBIABDELILAH/llm-from-scratch

Production-ready Transformer implementation from scratch with multi-scale benchmarking and Arabic language support

Jupyter Notebook 0 AI 100 個人 公開済 ↗
yogevat/LLM-TeamGym

Professional multi-agent benchmark library for evaluating LLMs in 23 strategy games — grid, board, social deduction, cards & game theory

Python 0 AI 100
luongnv89/dgx-spark-llm-lab

Benchmark local coding LLMs on an OpenAI-compatible endpoint, then keep the serving config that won. Hidden executable tests, reference-validated tasks, mermaid reports.

Python 0 AI 90
zakahadi/llm-eval-cli

A CLI developer tool for running custom evaluation suites to OpenAI-compatible APIs (MiMo, Claude, GPT, etc.), outputting Markdown research reports + JSON traces. Useful for developers who want to benchmark models before production. Suitable for Data/Research + Dev tools, and a natural fit using the MiMo API + Hermes Agent workflow.

Python 0 AI 90
apinode-pro/ai-api-gateway-benchmark

Reproducible OpenAI-compatible AI API gateway benchmark with API NODE defaults, latency metrics, success rate, and GitHub Actions results.

JavaScript 0 AI 80 公開済 ↗
guangxiangdebizi/tool-output-spoofing-lab

Benchmarking schema-valid false tool observations and defense baselines for tool-using LLM agents.

Python 0 AI 70
sammy995/fiduciary

Can your AI act as a fiduciary inside a regulated bank? A deployment-readiness benchmark that drops LLMs into a synthetic regulated organization as an employee and audits the behavior.

Python 0 AI 70
Jott2121/sabot

Do your agent pipeline's own checks catch planted faults? Measured on LangGraph, CrewAI and AutoGen: median 16.7%. One prompt-level change takes it to 55.0%. Pre-registered spec, Apache-2.0 harness, every raw trace published.

Python 0 AI 70 個人 公開済 ↗
Coinupbtc/zwell-bench

What: local LLM bakeoff harness (coding/vision/tools/agentic). For: objective model comparisons. How: ./setup.sh then ZWELL_BASE=… python bench_zwell.py --tag …

Python 0 AI 70 個人 公開済 ↗
mborges-dev/extraction-evals

Reproducible benchmark for LLM-based structured extraction from documents. Compare Claude / GPT / Gemini / open-weight on the same task with cost + latency tracking.

Python 0 AI 70
RudraO2/tokenbrawl

A latency-fair LLM-vs-LLM fighting game benchmark. The engine blocks until an agent responds, so inference speed can't win a match — what's scarce is a per-match token bank that drains as a model thinks.

TypeScript 0 AI 70 1 sig
Eqqinox/Ollamancer

Fully-local terminal AI agent for Ollama. 34 tools, MCP, local RAG, deterministic hallucination checks. No cloud, no API keys, nothing leaves your machine.

Python 0 AI 70
rosscyking1115/redteam-foundry

LLM red-team evaluation harness — prompt injection, refusal, leakage and staleness, with cross-judge validation and attack-corpus audits.

Python 0 AI 70
justuseapen/polarity-bench

Does following a negated instruction come free with the positive one? A reproducible eval — two pre-registered nulls for frontier Claude.

Python 0 AI 70
coroot/rca-lab

A Kubernetes lab that reproduces real production incidents on a live, instrumented microservice stack — for testing root-cause-analysis tools and agents.

Go 0 AI 45

RELATED 他のTopicも見る · 全Topicランキング →

#claude-code

1,564

#ai-agents

1,160

#llm

1,071

#claude

943

#python

806

#ai

741

#developer-tools

728

#mcp

727

#codex

521

集計対象: 各Repoの最新contentスナップショットの topics_json に小文字一致でマッチしたもの。 算出方法