Training Lab
Phase 2
Reward Events
Training Track
Evaluation track is always fixed.
Reward History
Metrics
Press Train to start learning. Metrics update in real-time.
Episode
Current Reward
Best Reward
Avg Reward
Distance
Collisions
Q-Updates
States Explored
Exploration ε
Best Episodes
No episodes yet.
Current Discrete State
Action Values
Run training or simulation to see Q-values.
About Q-Values
Each number shows how much total reward the agent expects from choosing that action in this state. Higher = the agent has learned that action works well here.
Policy History
Train to see how the policy changes over time.
Watch a saved episode
Before vs After Training

Use Replay First Episode to see untrained behaviour, then Replay Best Episode after training to compare improvement.

Speed: Slow Fast Medium
Connect and click Train to begin.