#reinforcement-learning
GitHub repositories that have self-applied the topic "reinforcement-learning" — a creator-tagged metadata that surfaces how AI projects describe themselves.
REPOS Repos for #reinforcement-learning (top 18 by stars)
The living ecosystem where AI agents complete tasks through workflow loops, improve through iterative execution, are evaluated by mentor agents or humans in the loop, and turn completed work into reusable work experience and data to improve future agents.
jaimasih05-commits/swarm-foraging-qlearnQ-Learning Swarm Foraging 2026: Multi-Agent RL in Dynamic Grid Environments
maxbaluev/accreted-intelligenceAccInt - local-first MCP Work Model for coding agents that learns from real outcomes.
TerminusAkivili/PAPPOPAPPO is a patch-aware reinforcement learning algorithm and experimental RL framework for tool-using coding agents. It provides turn-level credit assignment, PPO training utilities, rollout extraction, critic baselines, benchmark runners, and reproducible evaluation pipelines for repository repair tasks.
GeFAA/hide-and-seek-2Modern JAX/Flax recreation & expansion of OpenAI's Emergent Tool Use (hide-and-seek): MAPPO + CTDE, entity Transformer + GRU memory, ELO self-play, and a clean Three.js 3D viewer.
waefrebeorn/money-roomOpen source BTC market prediction engine. 80-dim C11 engine, 4 live rooms, 10K AI genomes, GDELT sentiment. 300/300 grid cells closed. Zero Python.
GeFAA/kivski-tactical-ai-simulatorTop-down 2D 5v5 bomb-defuse multi-agent RL simulator with live match viewer. Recurrent MAPPO + TarMAC emergent communication, no scripted strategies.
yogevat/LLM-TeamGymProfessional multi-agent benchmark library for evaluating LLMs in 23 strategy games — grid, board, social deduction, cards & game theory
taoyun0303-star/AI-Animal-ChessPygame Animal Chess with Negamax, MCTS, RL agents, and an AI Thinking Panel.
narutopyy/agent-arenaThe trust layer for autonomous trading agents on Bitget: a signed, verifiable safety firewall + overfit-aware tournament + live arena. 4 published GetAgent Playbooks.
lilfry09/Awesome-coding-agent-paperCurated papers, benchmarks, datasets, environments, and engineering notes for repository-level coding agents.
kalyvask/self-improving-agentic-systemsA controller that learns how a tool-using agent should spend compute at each step (WIDER, DEEPER, DECOMPOSE, STOP, or ESCALATE to a stronger model) and improves its own policy from logged traces. Judged on cost per solved task with paired statistics; ships trained policies and an offline embed API (pip install, no key needed).
fernforge-arcade/echo-civilizationEcho Civilization — a research simulation testing whether a population of simple learning agents (no pretrained LLMs) can accumulate knowledge and become more capable over generations through a civilization-like process.
ronakrajput8882/Flappy-Bird-DQNDeep Q-Network (DQN) agent trained to play Flappy Bird using PyTorch & Gymnasium — with experience replay, target network, and epsilon-greedy exploration.
Vansh9zz/Flappy-Bird-RLReinforcement Learning agent trained to play Flappy Bird using the Gym Flappy Bird environment.
Divyansh-9/CNN_VIT_BILSTM_CROSS_ATTENTION_BASED_TRAFFIC_MANAGEMENT_SYSTEMCamera-only congestion forecasting and RL signal control for unstructured Indian intersections. CNN-ViT cross-attention predicts per-lane congestion 60s ahead; a PPO agent uses that forecast to time the lights. A project - simulation-validated in SUMO.
azadkara/vizdoom-ppoPPO agent trained from scratch to play a VizDoom deathmatch from raw pixels, with a convolutional actor-critic and shaped rewards.
aawadall/strategosTactical command game using topographic maps, NATO APP-6 symbology, and adaptive AI
RELATED Other topics · full topics ranking →
#claude-code
1,564#ai-agents
1,160#llm
1,071#claude
943#python
806#ai
741#developer-tools
728#mcp
727#codex
521Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology