AI Dev Impact Lab JA
← Topics ranking · 2026-08
GITHUB TOPIC

#swe-bench

GitHub repositories that have self-applied the topic "swe-bench" — a creator-tagged metadata that surfaces how AI projects describe themselves.

8
tagged repos
54
top 8 stars
3
with tool sigs
8
shown

REPOS Repos for #swe-bench (top 8 by stars)

SihyeonJeon/why-was-fable-banned

Fable-style spec + evidence gate for Claude Code + Codex. Makes Opus/Codex work under Fable-like discipline: blocks every edit until a deterministic spec passes, and there is no "done" without live acceptance evidence. Spec-first, verification-gated, forbidden-paths enforced.

Python 47 AI 70
linny006/agent-eval-harness

Live, open-source benchmark for comparing AI coding agents on real GitHub issues

Python 6 AI 90
ttxs69/coding-agent-eval

Public, reproducible benchmark of CLI coding agents (Claude Code, Codex, Aider) on SWE-bench Verified. Live leaderboard: https://ttxs69.github.io/coding-agent-eval/

Python 1 AI 90 1 sig
ziyilam3999/local-first-agent-harness

A local-first coding agent: runs the heavy executor on your local model and escalates to the cloud only when stuck. Out-resolves single-shot Opus/Sonnet by planning, running the project's real tests, and retrying — at ~half the cost of an all-cloud chain. Graded by SWE-bench, not an LLM.

Python 0 AI 100
lilfry09/Awesome-coding-agent-paper

Curated papers, benchmarks, datasets, environments, and engineering notes for repository-level coding agents.

0 AI 75 Solo live ↗
ahmedEid1/forgejudge

Open, always-on leaderboard + CI gate for autonomous coding agents — every patch sandboxed, every run traced, every regression fails the build. $0 stack.

Python 0 AI 70 Solo live ↗
manfromnowhere143/telos

Evidence protocol and benchmark harness for verifying autonomous agent task completion.

Python 0 AI 70 Solo 1 sig live ↗
satwiksps/scaffoldscope

Controlled coding-agent harness ablations with auditable traces, reproducible evidence bundles, and SWE-bench interoperability.

Python 0 AI 70 Solo 1 sig live ↗

RELATED Other topics · full topics ranking →

#claude-code

1,564

#ai-agents

1,160

#llm

1,071

#claude

943

#python

806

#ai

741

#developer-tools

728

#mcp

727

#codex

521

Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology