#swe-bench
GitHub Topic「swe-bench」がついているAI関連リポジトリの集計。Topicはリポジトリ作者が自己申告するメタタグで、AI関連の文脈で何が「swe-bench」と分類されているかを可視化します。
REPOS #swe-bench のRepo (TOP 8 / Stars降順)
Fable-style spec + evidence gate for Claude Code + Codex. Makes Opus/Codex work under Fable-like discipline: blocks every edit until a deterministic spec passes, and there is no "done" without live acceptance evidence. Spec-first, verification-gated, forbidden-paths enforced.
linny006/agent-eval-harnessLive, open-source benchmark for comparing AI coding agents on real GitHub issues
ttxs69/coding-agent-evalPublic, reproducible benchmark of CLI coding agents (Claude Code, Codex, Aider) on SWE-bench Verified. Live leaderboard: https://ttxs69.github.io/coding-agent-eval/
ziyilam3999/local-first-agent-harnessA local-first coding agent: runs the heavy executor on your local model and escalates to the cloud only when stuck. Out-resolves single-shot Opus/Sonnet by planning, running the project's real tests, and retrying — at ~half the cost of an all-cloud chain. Graded by SWE-bench, not an LLM.
lilfry09/Awesome-coding-agent-paperCurated papers, benchmarks, datasets, environments, and engineering notes for repository-level coding agents.
ahmedEid1/forgejudgeOpen, always-on leaderboard + CI gate for autonomous coding agents — every patch sandboxed, every run traced, every regression fails the build. $0 stack.
manfromnowhere143/telosEvidence protocol and benchmark harness for verifying autonomous agent task completion.
satwiksps/scaffoldscopeControlled coding-agent harness ablations with auditable traces, reproducible evidence bundles, and SWE-bench interoperability.
RELATED 他のTopicも見る · 全Topicランキング →
#claude-code
1,555#ai-agents
1,156#llm
1,066#claude
936#python
802#ai
737#developer-tools
723#mcp
719#codex
517集計対象: 各Repoの最新contentスナップショットの topics_json に小文字一致でマッチしたもの。 算出方法