01
Context
Shepherd-RL is an academic reinforcement learning project devoted to a shepherding problem: one or more “shepherd” agents learn to guide one or more “sheep” toward a target zone, in a simulation environment whose complexity increases progressively.

02
Environment progression
The environment is introduced through increasing difficulty levels. Simple agents (Lazy, Tipsy, Rule-Based) serve as baselines to separate the environment’s difficulty from the limitations specific to the learning algorithms.
- one shepherd, one sheep, deterministic environment
- adding stochastic sheep movement
- adding an obstacle
- moving to several sheep, then to two shepherds
03
Algorithms
Three algorithm families were compared: PPO, TD3, and DQN. On the simplest configuration, TD3 and PPO converge markedly better than DQN. The performance of all three algorithms degrades as the environment’s complexity increases.
04
Curriculum and reward shaping
Moving to several sheep sharply increases the difficulty: flock dispersion, stochasticity, a larger observation space, and a rarer success reward. The first reinforcement learning approaches fail on the noisy three-sheep configuration, while the Rule-Based baseline confirms the environment remains solvable. Work then shifts toward reward shaping, increasing network capacity, curriculum learning, and fine-tuning: the problem isn’t solved by algorithm choice alone.
05
Multi-agent coordination
The final step introduces two shepherds. The first attempts reveal an observation issue: the second agent lacks a frame of reference properly centered on its own position. The environment is then fixed with symmetric observations and a proximity reward, followed by curriculum learning on a simplified configuration before returning to the full flock.
- intermediate step (two shepherds, one sheep): TD3 reaches 96% success
- final configuration (two shepherds, three sheep): TD3 reaches 37% success over 100 episodes, a partial result in a difficult environment
06
Limits
The results correspond to a simulated environment and a limited training budget. The final configuration doesn’t reach full convergence and remains heavily dependent on reward shaping, at a significant computational cost. Training-time comparisons remain imperfect, since the hardware used changed over the course of the project.