All explorations
Exploration

Reinforcement Learning · Multi-agent · Simulation

Multi-agent RL

Exploring reinforcement learning in a multi-agent environment.

Status
Academic project
Stack
Python · Gym-style environment · Pygame · PPO · TD3 · DQN · Stable-Baselines3 · PyTorch · Reward Shaping · Curriculum Learning · Fine-tuning · Rule-based baselines

01

Context

Shepherd-RL is an academic reinforcement learning project devoted to a shepherding problem: one or more “shepherd” agents learn to guide one or more “sheep” toward a target zone, in a simulation environment whose complexity increases progressively.

Simulation environment for the shepherding problem, with shepherds, sheep, target zone, and obstacle.
Simulation environment for the shepherding problem.

02

Environment progression

The environment is introduced through increasing difficulty levels. Simple agents (Lazy, Tipsy, Rule-Based) serve as baselines to separate the environment’s difficulty from the limitations specific to the learning algorithms.

  • one shepherd, one sheep, deterministic environment
  • adding stochastic sheep movement
  • adding an obstacle
  • moving to several sheep, then to two shepherds

03

Algorithms

Three algorithm families were compared: PPO, TD3, and DQN. On the simplest configuration, TD3 and PPO converge markedly better than DQN. The performance of all three algorithms degrades as the environment’s complexity increases.

04

Curriculum and reward shaping

Moving to several sheep sharply increases the difficulty: flock dispersion, stochasticity, a larger observation space, and a rarer success reward. The first reinforcement learning approaches fail on the noisy three-sheep configuration, while the Rule-Based baseline confirms the environment remains solvable. Work then shifts toward reward shaping, increasing network capacity, curriculum learning, and fine-tuning: the problem isn’t solved by algorithm choice alone.

05

Multi-agent coordination

The final step introduces two shepherds. The first attempts reveal an observation issue: the second agent lacks a frame of reference properly centered on its own position. The environment is then fixed with symmetric observations and a proximity reward, followed by curriculum learning on a simplified configuration before returning to the full flock.

  • intermediate step (two shepherds, one sheep): TD3 reaches 96% success
  • final configuration (two shepherds, three sheep): TD3 reaches 37% success over 100 episodes, a partial result in a difficult environment

06

Limits

The results correspond to a simulated environment and a limited training budget. The final configuration doesn’t reach full convergence and remains heavily dependent on reward shaping, at a significant computational cost. Training-time comparisons remain imperfect, since the hardware used changed over the course of the project.

Back

All explorations