STAGE 1-2
PROTOTYPE

Reinforcement Learning on Flappy Bird

Role ML / Gameplay Engineer
Timeline 2025
Platform Unity + Python (ML-Agents CLI)
Team Size 1 (Solo)
Unity ML-Agents 4.0.1PPOC#PythonRigidbody2DNew Input System
STAGE 1-2 // BOTH REINFORCEMENT LEARNING ON FLAPPY BIRD Unity ML-Agents 4.0.1 PPO C# Python Rigidbody2D New Input System SYSTEM_INIT: OK | MEMORY: ALLOCATED | RENDERING: ONLINE

Mission Brief

Transform the physics-based Flappy Bird game into a rigorous reinforcement-learning training environment without altering the underlying physics or giving the AI agent mechanical capabilities unavailable to a human player.

Key Systems Engineered

Minimal 4-Continuous-Observation & 2-Discrete-Action Specification (BirdAgent.cs)

CollectObservations() feeds exactly 4 continuous floats per step: (1) Player Y position, (2) Player vertical velocity, (3) vertical distance from player to the next pipe's gap center, and (4) next pipe's X position, guarded by a defensive null-check during singleton initialization. OnActionReceived() maps a single discrete branch (0 = no-op, 1 = jump) to the exact same Jump() method used by human input.

Event-Driven Sparse + Dense Reward Shaping

Subscribes directly to the existing PipesManager.OnScoreIncrease event to grant a sparse +1.0 reward whenever a pipe is cleared, paired with a continuous +0.1 * Time.fixedDeltaTime survival shaping bonus every physics step so early-stage policies receive a gradient signal before clearing their first obstacle.

Dual Heuristic / Autonomous Policy Architecture

Implements Heuristic() using Unity's New Input System to map the Spacebar directly into the discrete action buffer, allowing instant switching between human baseline play and trained neural-network inference via the Inspector Behavior Type dropdown without maintaining two codebases.

Boss Fight: Cold-Start Sparse Reward Collapse & Episode Reset Race Conditions

With only a +1.0 reward on passing a pipe, early random policies crashed into the floor or ceiling before ever reaching the first gap, yielding zero learning signal. Adding the +0.1 * Time.fixedDeltaTime per-step survival shaping reward stabilized early exploration, while adding defensive null checks in OnEpisodeBegin() and CollectObservations() prevented first-frame singleton order crashes across 500,000 training steps.

Implementation Details

Architecture Diagram

Architecture & Deep Dive

The RL pipeline layers a BirdAgent MonoBehaviour on top of the existing Flappy Bird game without modifying physics or movement code. The agent observes the environment through 4 floats and outputs a single discrete action.

Observation Space

The 4 continuous observations are carefully chosen to give the agent minimal but sufficient information: its own Y position and velocity (for trajectory prediction) and the next pipe’s gap center Y and X position (for timing the jump). No pipe velocity observation is needed because pipes move at a fixed constant speed.

Reward Architecture

The reward function combines sparse task rewards (+1.0 per pipe cleared via event subscription) with dense survival shaping (+0.1 × fixedDeltaTime per physics step). This two-tier design solved the cold-start problem where random policies would die before ever reaching a pipe.

Training Configuration

Training used PPO with a batch size of 1024, buffer size of 10,240, and a linearly decaying learning rate starting at 3.0e-4. The policy network uses 2 hidden layers of 128 units each, trained for 500,000 total steps across two preserved runs.

Media Gallery

Loot & Rewards

  • Completed two full training runs preserved in results/ (TestRun01 and FlappyBird_Run1) trained to 500,000 max steps.
  • Verified PPO Hyperparameters (FlappyBirdConfig.yaml): Batch Size = 1024, Buffer Size = 10,240, Learning Rate = 3.0e-4 (linear decay), Beta (entropy) = 1.0e-2, Epsilon (clip) = 0.2, Lambda (GAE) = 0.95, Gamma = 0.99, Time Horizon = 64, 2 hidden layers x 128 units (normalize: false).

Select Player Class

Your class determines which skills and projects are highlighted throughout the site.

Achievements (0/11)

Press Start
Skipped or clicked through the boot sequence.
Class Act
Selected a Player Class to customize project ordering.
Efficient
Opened the 60-second Recruiter Mode summary.
Level Up
Opened your first engineering project Case File.
Completionist
Cleared all 12 project Case Files.
Coin Collector
Found and collected all 7 hidden coins across the site.
Old School
Entered the Konami Code or activated Arcade Mode.
Snake Charmer
Scored 10 or more points in the 404 Game Over Snake mini-game.
Hello, World
Copied the contact email address to your clipboard.
Curious
Opened the Keyboard Controls (?) overlay.
Bonus Round
Entered the hidden Bonus Stage development lab.

System Settings

Theme
Motion
Audio
Sound Effects
Advanced
Keyboard Shortcuts
Show Hidden Coins
Arcade Mode (CRT)

Keyboard Shortcuts

Navigate Sections
1 2 3 4
Open Achievements
A
Open Settings
S
Toggle Audio
M
Close Modal
Esc
Konami Code
↑ ↑ ↓ ↓ ← → ← → B A