Current Discrete State
—
Action Values
Run training or simulation to see Q-values.
About Q-Values
Each number shows how much total reward the agent expects from choosing that action in this state.
Higher = the agent has learned that action works well here.
Policy History
Train to see how the policy changes over time.