Pendulum DQN

Can one frozen DQN recipe balance a pendulum when gravity changes?

Context
Deep Learning coursework, with Mohamad Aniq
My part
Evaluation design and analysis; Aniq led the DQN implementation
Stack
Python, TensorFlow, Keras, DQN, Gym
Year
2026
Pendulum control rollouts across gravity conditions

4

Fresh policies at g = 10, 0, −10 and 15, one frozen recipe

What it is
Continuous torque discretised to seven actions; Vanilla, Double and Dueling DQN compared; Dueling kept and its recipe frozen; then four separate policies trained, one per gravity.
What I did
I designed the fixed development states and the 24 protected held-out start states, the seed-matched comparison rules and the gravity hypotheses, then ran the final four-gravity analysis and verified the saved weights.
Result
Mean held-out return was −135 at normal gravity, −54 at zero and anti-gravity, and −203 at 15. Selection mattered: the supergravity policy fell to −326 by its last episode, and choosing checkpoints before touching held-out states kept that honest.