AI Dev Impact Lab JA
← Topics ranking · 2026-08
GITHUB TOPIC

#reinforcement-learning

GitHub repositories that have self-applied the topic "reinforcement-learning" — a creator-tagged metadata that surfaces how AI projects describe themselves.

18
tagged repos
1,469
top 18 stars
3
with tool sigs
18
shown

REPOS Repos for #reinforcement-learning (top 18 by stars)

ray-r-ren/agent-apprenticeship

The living ecosystem where AI agents complete tasks through workflow loops, improve through iterative execution, are evaluated by mentor agents or humans in the loop, and turn completed work into reusable work experience and data to improve future agents.

Python 1,337 AI 100 Solo live ↗
jaimasih05-commits/swarm-foraging-qlearn

Q-Learning Swarm Foraging 2026: Multi-Agent RL in Dynamic Grid Environments

HTML 117 AI 70
maxbaluev/accreted-intelligence

AccInt - local-first MCP Work Model for coding agents that learns from real outcomes.

Shell 6 AI 45 Solo live ↗
TerminusAkivili/PAPPO

PAPPO is a patch-aware reinforcement learning algorithm and experimental RL framework for tool-using coding agents. It provides turn-level credit assignment, PPO training utilities, rollout extraction, critic baselines, benchmark runners, and reproducible evaluation pipelines for repository repair tasks.

Python 3 AI 70 Solo live ↗
GeFAA/hide-and-seek-2

Modern JAX/Flax recreation & expansion of OpenAI's Emergent Tool Use (hide-and-seek): MAPPO + CTDE, entity Transformer + GRU memory, ELO self-play, and a clean Three.js 3D viewer.

Python 3 AI 60
waefrebeorn/money-room

Open source BTC market prediction engine. 80-dim C11 engine, 4 live rooms, 10K AI genomes, GDELT sentiment. 300/300 grid cells closed. Zero Python.

C 2 AI 45 Solo live ↗
GeFAA/kivski-tactical-ai-simulator

Top-down 2D 5v5 bomb-defuse multi-agent RL simulator with live match viewer. Recurrent MAPPO + TarMAC emergent communication, no scripted strategies.

Python 1 AI 100 1 sig
yogevat/LLM-TeamGym

Professional multi-agent benchmark library for evaluating LLMs in 23 strategy games — grid, board, social deduction, cards & game theory

Python 0 AI 100
taoyun0303-star/AI-Animal-Chess

Pygame Animal Chess with Negamax, MCTS, RL agents, and an AI Thinking Panel.

Python 0 AI 100
narutopyy/agent-arena

The trust layer for autonomous trading agents on Bitget: a signed, verifiable safety firewall + overfit-aware tournament + live arena. 4 published GetAgent Playbooks.

Python 0 AI 75 Solo live ↗
lilfry09/Awesome-coding-agent-paper

Curated papers, benchmarks, datasets, environments, and engineering notes for repository-level coding agents.

0 AI 75 Solo live ↗
kalyvask/self-improving-agentic-systems

A controller that learns how a tool-using agent should spend compute at each step (WIDER, DEEPER, DECOMPOSE, STOP, or ESCALATE to a stronger model) and improves its own policy from logged traces. Judged on cost per solved task with paired statistics; ships trained policies and an offline embed API (pip install, no key needed).

Python 0 AI 70
fernforge-arcade/echo-civilization

Echo Civilization — a research simulation testing whether a population of simple learning agents (no pretrained LLMs) can accumulate knowledge and become more capable over generations through a civilization-like process.

Python 0 AI 70
ronakrajput8882/Flappy-Bird-DQN

Deep Q-Network (DQN) agent trained to play Flappy Bird using PyTorch & Gymnasium — with experience replay, target network, and epsilon-greedy exploration.

Python 0 AI 70
Vansh9zz/Flappy-Bird-RL

Reinforcement Learning agent trained to play Flappy Bird using the Gym Flappy Bird environment.

JavaScript 0 AI 70
Divyansh-9/CNN_VIT_BILSTM_CROSS_ATTENTION_BASED_TRAFFIC_MANAGEMENT_SYSTEM

Camera-only congestion forecasting and RL signal control for unstructured Indian intersections. CNN-ViT cross-attention predicts per-lane congestion 60s ahead; a PPO agent uses that forecast to time the lights. A project - simulation-validated in SUMO.

Python 0 AI 50 1 sig
azadkara/vizdoom-ppo

PPO agent trained from scratch to play a VizDoom deathmatch from raw pixels, with a convolutional actor-critic and shaped rewards.

Jupyter Notebook 0 AI 45
aawadall/strategos

Tactical command game using topographic maps, NATO APP-6 symbology, and adaptive AI

C# 0 AI 45 Solo 1 sig live ↗

RELATED Other topics · full topics ranking →

#claude-code

1,564

#ai-agents

1,160

#llm

1,071

#claude

943

#python

806

#ai

741

#developer-tools

728

#mcp

727

#codex

521

Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology