ME102B cable robot · simulated

Q-Learning Air Hockey

Two learned players, PPO and tabular Q-learning, trained in a 2D sim built on the real table calibration, mallet workspace and 10 ms tick. The robot (orange) defends the left goal against a scripted shooter that fires straight shots, bank shots and slow drifters. Switch to the hand-coded goalie to compare.

Robot0
Shooter0
action – puck 0 mm/s t 0.0 s
robot mallet puck (with trail) dashed box = safe mallet workspace dashed line = defense line
Evaluation

Where each shot ends up

Training

Learning curve

Tuning

Hyperparameter sweep