My brother swore the Q-learning tutorial would click in a week, 5 weeks later I'm still stuck on the same grid world
My brother has been doing machine learning stuff for a while now and when I told him I wanted to get into reinforcement learning he said just do the Sutton and Barto book and that Q-learning chapter, you'll get it in like a week. That was 5 weeks ago. I'm still on the same 5x5 grid world example, watching my agent run into the same wall over and over, and my reward curve looks like a flat line with some noise. I've messed with the learning rate, the discount factor, epsilon decay, everything, and I still can't get it to find the goal reliably in under 200 episodes. Meanwhile the tutorial says it should converge in like 20. He keeps saying just be patient but he did this stuff in grad school so I think he forgot what it's like to not already know it. The part that really gets me is I understand the math on paper, I can write out the update rule, but getting it to actually work in code is a different animal. Anyone else hit a wall like this with RL and did something other than just grinding the same tutorial?