#benchmark
GitHub Topic「benchmark」がついているAI関連リポジトリの集計。Topicはリポジトリ作者が自己申告するメタタグで、AI関連の文脈で何が「benchmark」と分類されているかを可視化します。
REPOS #benchmark のRepo (TOP 34 / Stars降順)
A curated collection of datasets for Large Language Models (LLMs), covering medical AI, NLP, multimodal learning, instruction tuning, reasoning, code generation, and evaluation benchmarks.
protorikis/rikisLightweight agent that pulls and runs LLM benchmarks from Protorikis Bench
linny006/agent-eval-harnessLive, open-source benchmark for comparing AI coding agents on real GitHub issues
AgentBenchAudit/evidence-boundsRelease repository for agent benchmark evidence-reporting artifacts and reproduction workflows.
kevinpeckham/barkup-benchPre-registered benchmark: HTML dialect + whole-tree rewrite (barkup) vs JSON, granular mutation tools, and two patch dialects, for LLM agents editing typed trees. 9,600 scored runs across 4 models. Findings in REPORT.md.
vahit19/longhaul-benchLong-horizon reliability benchmark for industrial edge agents - do self-improving agents get better or corrupt over 1000+ episodes on constrained hardware?
bmendonca3/authzbench-saasBenchmark for AI agents proving multi-tenant SaaS authorization bugs
Mazha0309/the-evil-repositoryAn evidence-hostile, container-isolated benchmark and behavior analysis platform for long-horizon AI software agents.
FU-max-boop/statebind-guardCatch visible-but-unbound coding-agent handoffs: CLI + GitHub Action with proof, policy gates, SARIF, HTML, and benchmark cards.
awdemos/toks-benchReproducible token-throughput benchmark for OpenAI-compatible LLM servers, tuned for NVIDIA Spark and GB10 inference.
ttxs69/coding-agent-evalPublic, reproducible benchmark of CLI coding agents (Claude Code, Codex, Aider) on SWE-bench Verified. Live leaderboard: https://ttxs69.github.io/coding-agent-eval/
beausome/bullshit-benchAn AI leaderboard scoring chatbots on how much they bullshit vs. give honest answers to unanswerable questions. Works with paid APIs and free/local models.
clouatre-labs/reversibility-benchmarkSupplementary materials for Reversibility as a Safety Gate for Agentic Actions
Vedansh5545/llm-shieldbenchTrustworthy AI evaluation tool for testing chatbot safety, reliability, hallucination behavior, privacy risk, and instruction-following quality.
Lohithr123/rf-copilotRF front-end analysis engine and a 60-problem LLM benchmark. Grounded small model ≈ ungrounded frontier model.
tonydzi/agent-runtime-integrity-benchIntegrity scenarios for agent runtimes, distilled from real fleet incidents — deterministic fault-injection checks against real SDKs
achenna-source/sagin-agent-benchMeasuring and verifying LLM decision agents under orbital dynamics, delay, and intermittency: a reproducible testbed, failure-mode taxonomy, and in-chain verification.
ZihanZheng2000/reservoir-multi-agent-deliberationSynthetic benchmark for role-limited multi-agent deliberation in reservoir construction decision support
chohyerinn/mini-agent-harnessEvaluation harness for coding agents - repeated runs, pass@k, bootstrap/McNemar significance, tamper detection, cost. Single vs Planner-Coder-Reviewer multi-agent on CLOVA.
BELYAGOUBIABDELILAH/llm-from-scratchProduction-ready Transformer implementation from scratch with multi-scale benchmarking and Arabic language support
yogevat/LLM-TeamGymProfessional multi-agent benchmark library for evaluating LLMs in 23 strategy games — grid, board, social deduction, cards & game theory
luongnv89/dgx-spark-llm-labBenchmark local coding LLMs on an OpenAI-compatible endpoint, then keep the serving config that won. Hidden executable tests, reference-validated tasks, mermaid reports.
zakahadi/llm-eval-cliA CLI developer tool for running custom evaluation suites to OpenAI-compatible APIs (MiMo, Claude, GPT, etc.), outputting Markdown research reports + JSON traces. Useful for developers who want to benchmark models before production. Suitable for Data/Research + Dev tools, and a natural fit using the MiMo API + Hermes Agent workflow.
apinode-pro/ai-api-gateway-benchmarkReproducible OpenAI-compatible AI API gateway benchmark with API NODE defaults, latency metrics, success rate, and GitHub Actions results.
guangxiangdebizi/tool-output-spoofing-labBenchmarking schema-valid false tool observations and defense baselines for tool-using LLM agents.
sammy995/fiduciaryCan your AI act as a fiduciary inside a regulated bank? A deployment-readiness benchmark that drops LLMs into a synthetic regulated organization as an employee and audits the behavior.
Jott2121/sabotDo your agent pipeline's own checks catch planted faults? Measured on LangGraph, CrewAI and AutoGen: median 16.7%. One prompt-level change takes it to 55.0%. Pre-registered spec, Apache-2.0 harness, every raw trace published.
Coinupbtc/zwell-benchWhat: local LLM bakeoff harness (coding/vision/tools/agentic). For: objective model comparisons. How: ./setup.sh then ZWELL_BASE=… python bench_zwell.py --tag …
mborges-dev/extraction-evalsReproducible benchmark for LLM-based structured extraction from documents. Compare Claude / GPT / Gemini / open-weight on the same task with cost + latency tracking.
RudraO2/tokenbrawlA latency-fair LLM-vs-LLM fighting game benchmark. The engine blocks until an agent responds, so inference speed can't win a match — what's scarce is a per-match token bank that drains as a model thinks.
Eqqinox/OllamancerFully-local terminal AI agent for Ollama. 34 tools, MCP, local RAG, deterministic hallucination checks. No cloud, no API keys, nothing leaves your machine.
rosscyking1115/redteam-foundryLLM red-team evaluation harness — prompt injection, refusal, leakage and staleness, with cross-judge validation and attack-corpus audits.
justuseapen/polarity-benchDoes following a negated instruction come free with the positive one? A reproducible eval — two pre-registered nulls for frontier Claude.
coroot/rca-labA Kubernetes lab that reproduces real production incidents on a live, instrumented microservice stack — for testing root-cause-analysis tools and agents.
RELATED 他のTopicも見る · 全Topicランキング →
#claude-code
1,564#ai-agents
1,160#llm
1,071#claude
943#python
806#ai
741#developer-tools
728#mcp
727#codex
521集計対象: 各Repoの最新contentスナップショットの topics_json に小文字一致でマッチしたもの。 算出方法