#benchmarking
GitHub repositories that have self-applied the topic "benchmarking" — a creator-tagged metadata that surfaces how AI projects describe themselves.
REPOS Repos for #benchmarking (top 12 by stars)
A curated collection of datasets for Large Language Models (LLMs), covering medical AI, NLP, multimodal learning, instruction tuning, reasoning, code generation, and evaluation benchmarks.
duclamvan/agent-hard-processA Hermes Agent skill for turning painful fixes into replayable benchmarked workflows
Dafenxz0/skillproofProve your Agent Skill works before you publish it.
kizz-tech/agentic-evidence-labTests exact agent interventions—skills, prompts, models, tools, and workflows—and publishes what changed, what failed, and what decision follows.
manishklach/k3-inference-platformProduction-oriented Kimi K3 inference control plane with checkpoint release gates, MoE capacity planning, OpenAI gateway, NVL72 deployment, benchmarks, and observability.
cuiqi5656/agent-quality-benchmarkEvidence-first, reproducible quality benchmarks for AI agents.
Animesh352/llm-agent-benchMulti-agent LLM orchestration and benchmarking toolkit -- planner-solver-reviewer workflows with provider abstraction and scoring
msradam/xk6-llmLoad test LLM inference servers with k6. TTFT, ITL, TPOT, goodput, cost, and energy metrics for any OpenAI-compatible server. Ships to Prometheus and Grafana.
lamb356/photonic-benchTransparent benchmark cards, JSON artifacts, and visual tools for photonic AI accelerator energy/noise claims.
jazzonaut/agentbenchCross-platform diagnostics and live LLM performance benchmarking for Claude Code, Headroom, and system bottlenecks
320exh/prompt-diffA fast, git-native CLI & local Web UI to version, diff tokens/costs, and benchmark LLM system prompts across local & cloud models.
satwiksps/scaffoldscopeControlled coding-agent harness ablations with auditable traces, reproducible evidence bundles, and SWE-bench interoperability.
RELATED Other topics · full topics ranking →
#claude-code
1,555#ai-agents
1,156#llm
1,066#claude
936#python
802#ai
737#developer-tools
723#mcp
719#codex
517Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology