#swe-bench
GitHub repositories that have self-applied the topic "swe-bench" — a creator-tagged metadata that surfaces how AI projects describe themselves.
REPOS Repos for #swe-bench (top 8 by stars)
Fable-style spec + evidence gate for Claude Code + Codex. Makes Opus/Codex work under Fable-like discipline: blocks every edit until a deterministic spec passes, and there is no "done" without live acceptance evidence. Spec-first, verification-gated, forbidden-paths enforced.
linny006/agent-eval-harnessLive, open-source benchmark for comparing AI coding agents on real GitHub issues
ttxs69/coding-agent-evalPublic, reproducible benchmark of CLI coding agents (Claude Code, Codex, Aider) on SWE-bench Verified. Live leaderboard: https://ttxs69.github.io/coding-agent-eval/
ziyilam3999/local-first-agent-harnessA local-first coding agent: runs the heavy executor on your local model and escalates to the cloud only when stuck. Out-resolves single-shot Opus/Sonnet by planning, running the project's real tests, and retrying — at ~half the cost of an all-cloud chain. Graded by SWE-bench, not an LLM.
lilfry09/Awesome-coding-agent-paperCurated papers, benchmarks, datasets, environments, and engineering notes for repository-level coding agents.
ahmedEid1/forgejudgeOpen, always-on leaderboard + CI gate for autonomous coding agents — every patch sandboxed, every run traced, every regression fails the build. $0 stack.
manfromnowhere143/telosEvidence protocol and benchmark harness for verifying autonomous agent task completion.
satwiksps/scaffoldscopeControlled coding-agent harness ablations with auditable traces, reproducible evidence bundles, and SWE-bench interoperability.
RELATED Other topics · full topics ranking →
#claude-code
1,564#ai-agents
1,160#llm
1,071#claude
943#python
806#ai
741#developer-tools
728#mcp
727#codex
521Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology