#reinforcement-learning
GitHub Topic「reinforcement-learning」がついているAI関連リポジトリの集計。Topicはリポジトリ作者が自己申告するメタタグで、AI関連の文脈で何が「reinforcement-learning」と分類されているかを可視化します。
REPOS #reinforcement-learning のRepo (TOP 18 / Stars降順)
The living ecosystem where AI agents complete tasks through workflow loops, improve through iterative execution, are evaluated by mentor agents or humans in the loop, and turn completed work into reusable work experience and data to improve future agents.
jaimasih05-commits/swarm-foraging-qlearnQ-Learning Swarm Foraging 2026: Multi-Agent RL in Dynamic Grid Environments
maxbaluev/accreted-intelligenceAccInt - local-first MCP Work Model for coding agents that learns from real outcomes.
TerminusAkivili/PAPPOPAPPO is a patch-aware reinforcement learning algorithm and experimental RL framework for tool-using coding agents. It provides turn-level credit assignment, PPO training utilities, rollout extraction, critic baselines, benchmark runners, and reproducible evaluation pipelines for repository repair tasks.
GeFAA/hide-and-seek-2Modern JAX/Flax recreation & expansion of OpenAI's Emergent Tool Use (hide-and-seek): MAPPO + CTDE, entity Transformer + GRU memory, ELO self-play, and a clean Three.js 3D viewer.
waefrebeorn/money-roomOpen source BTC market prediction engine. 80-dim C11 engine, 4 live rooms, 10K AI genomes, GDELT sentiment. 300/300 grid cells closed. Zero Python.
GeFAA/kivski-tactical-ai-simulatorTop-down 2D 5v5 bomb-defuse multi-agent RL simulator with live match viewer. Recurrent MAPPO + TarMAC emergent communication, no scripted strategies.
yogevat/LLM-TeamGymProfessional multi-agent benchmark library for evaluating LLMs in 23 strategy games — grid, board, social deduction, cards & game theory
taoyun0303-star/AI-Animal-ChessPygame Animal Chess with Negamax, MCTS, RL agents, and an AI Thinking Panel.
narutopyy/agent-arenaThe trust layer for autonomous trading agents on Bitget: a signed, verifiable safety firewall + overfit-aware tournament + live arena. 4 published GetAgent Playbooks.
lilfry09/Awesome-coding-agent-paperCurated papers, benchmarks, datasets, environments, and engineering notes for repository-level coding agents.
kalyvask/self-improving-agentic-systemsA controller that learns how a tool-using agent should spend compute at each step (WIDER, DEEPER, DECOMPOSE, STOP, or ESCALATE to a stronger model) and improves its own policy from logged traces. Judged on cost per solved task with paired statistics; ships trained policies and an offline embed API (pip install, no key needed).
fernforge-arcade/echo-civilizationEcho Civilization — a research simulation testing whether a population of simple learning agents (no pretrained LLMs) can accumulate knowledge and become more capable over generations through a civilization-like process.
ronakrajput8882/Flappy-Bird-DQNDeep Q-Network (DQN) agent trained to play Flappy Bird using PyTorch & Gymnasium — with experience replay, target network, and epsilon-greedy exploration.
Vansh9zz/Flappy-Bird-RLReinforcement Learning agent trained to play Flappy Bird using the Gym Flappy Bird environment.
Divyansh-9/CNN_VIT_BILSTM_CROSS_ATTENTION_BASED_TRAFFIC_MANAGEMENT_SYSTEMCamera-only congestion forecasting and RL signal control for unstructured Indian intersections. CNN-ViT cross-attention predicts per-lane congestion 60s ahead; a PPO agent uses that forecast to time the lights. A project - simulation-validated in SUMO.
azadkara/vizdoom-ppoPPO agent trained from scratch to play a VizDoom deathmatch from raw pixels, with a convolutional actor-critic and shaped rewards.
aawadall/strategosTactical command game using topographic maps, NATO APP-6 symbology, and adaptive AI
RELATED 他のTopicも見る · 全Topicランキング →
#claude-code
1,564#ai-agents
1,160#llm
1,071#claude
943#python
806#ai
741#developer-tools
728#mcp
727#codex
521集計対象: 各Repoの最新contentスナップショットの topics_json に小文字一致でマッチしたもの。 算出方法