Built in the open · lab experiment
In-browser machine learning

Skyline Run AI Lab

Train, watch, and compare seven learning methods without leaving the browser.

Start with the basics

PPO starts with the recommended baseline: train on District 1 and validate on Districts 2 and 3. Start training beside the live simulation, or open the advanced sections to customize worlds, manage policies, and inspect diagnostics.

Training setup

All seven methods train with the same unhandicapped controls, and switching algorithms keeps your chosen worlds. Demonstrations remain PPO-only and reward control remains plain-NEAT-only.
Reproducibility

Saved in this browser

No models saved in this browser yet.

Run & inspect

Idle Environment steps: 0
More run statistics
Evaluation steps: 0 0 ticks/s
Speed
Turn off rendering to train at full speed.
Overlay the sensed tiles and inspect the exact network inputs.

Save and watch controls become available after training produces an evaluated policy.

Live network

Waiting for the first network decision…

Training worlds

The recommended baseline trains on the primary district and validates on two separate campaign districts. Customize these worlds explicitly when you want to test generalization.

Training districts
Validation worlds
Method-specific options
Human demos Behavior-cloning demonstrations are implemented only for PPO.
Trainer AI Reward-controller training is implemented only for plain NEAT.

Human demos

PPO can behavior-clone eligible runs from campaign districts in the current training set before reinforcement learning.

0 of 0 saved demos are eligible for the selected training districts.

No saved demos yet. Clear a district in the playable game and save the run.

Policies & evaluation

Active policy

No evaluated active policy is ready.

Test policy

Test the selected policy on all campaign districts and five fresh generated worlds. Results include clean and 25% sticky-action passes and never change policy selection.

Comparison roster

No policies in the comparison roster.

Diagnostics Show algorithm-specific training charts and network diagnostics
Iteration 0 Episode 0 Best progress: 0 px Agents alive: 0 0 ticks/s Idle Environment steps: 0 Evaluation steps: 0
Episode reward

Metrics appear once training starts.

Normalized episode progress

Metrics appear once training starts.

Mean episode progress

Metrics appear once training starts.

Training loss

Metrics appear once training starts.

Policy entropy

Metrics appear once training starts.

Completions per generation

Metrics appear once training starts.