PPolySim OS
For Researchers · Q-Learning Gridworld

Q-Learning Gridworld for a pathfinding AI

Built for researchers prototyping or validating an idea. Prototype fast, reproduce exactly, and share a citable, interactive version of your model. Simulate a pathfinding AI live below — adjust the inputs and watch it respond, right in your browser.

Q-Learning GridworldLive

Controls

Presets

Reinforcement learning finds a policy — an arrow in every cell — that maximizes long-term reward. A high discount γ makes the agent value the distant goal; a costlier step reward pushes it to take the shortest path. The colors show each cell's learned value. Educational tool.

▶ Run in Python

Data Inspector

Discount γ0.90
Start-cell value0.09
Behaviorcautious path

Governing equation

Reading this result: With γ=0.9 and a mild step reward, the agent balances path length against reaching the goal, favoring a steady route.

Runs locally in your browser — free forever. Scale to the cloud when reality gets heavy.

or unlock everything with Pro →

More with Q-Learning Gridworld

Frequently asked questions

Is this good for researchers?
Yes — this version of "Q-Learning Gridworld for a pathfinding AI" is framed for researchers prototyping or validating an idea. Prototype fast, reproduce exactly, and share a citable, interactive version of your model.
Do I need to install anything?
No. It runs in any modern browser, free, with no account required.