Main Article Content
Abstract
Deep reinforcement learning (DRL) has been widely adopted for mobile robot path planning in unknown and complex environments. However, existing DRL-based approaches suffer from two main limitations: sparse reward, which provides meaningful feedback only at terminal states, and experience heterogeneity, which reduces the efficiency of uniform sampling from the standard experience replay buffer due to the diversity of transition qualities stored within it. To tackle both issues at once, we propose SS-PER-DQN, a novel path-planning approach in which stratified sampling and a prioritized experience replay mechanism are incorporated into the Double Deep Q-Network. The memory buffer is divided into three priority layers based on temporal-difference error levels, with a dynamically varying sampling ratio between layers, and importance-sampling compensation to balance the sampling-distribution bias. A densely shaped reward function is additionally employed to expedite early exploration in the presence of sparsely rewarded actions in the environment. Experiments performed on Grid World and ROS Gazebo simulations show that the SS-PER-DQN is able to successfully plan paths at success rates of 94.5% and 89.0%, with a 24.5% and 21.8% reduction in path lengths, respectively, when compared to the traditional DQN algorithm, after convergence in approximately 53% fewer training episodes. The analysis conducted during the ablation study reveals the specific roles of each contribution: stratified sampling helps the algorithm converge early by ensuring that informative state transitions guide the first phase of learning, and importance weighting influences the quality of gradient updates in later stages. It is also observed that the absolute planning time per decision step ranges from 5.0 ms (Grid World) to 5.3 ms (Gazebo) for SS-PER-DQN, representing an incremental overhead of approximately 1.9–2.2 ms per step relative to the classic DQN baseline, which remains well within the real-time requirements for robot navigation.
Keywords
Article Details
References
- B. Singh, R. Kumar, V.P. Singh, Reinforcement learning in robotic applications: a comprehensive survey, Artificial Intelligence Review 55(2) (2022) 945-990.
- K. Zhu, T. Zhang, Deep reinforcement learning based mobile robot navigation: A review, Tsinghua Science and Technology 26(5) (2021) 674-691.
- V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, M. Riedmiller, Playing atari with deep reinforcement learning, arXiv preprint arXiv:1312.5602 (2013). https://doi.org/10.48550/arXiv.1312.5602
- V. Mnih, K. Kavukcuoglu, D. Silver, A.A. Rusu, J. Veness, M.G. Bellemare, A. Graves, M. Riedmiller, A.K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, D. Hassabis, Human-level control through deep reinforcement learning, Nature 518(7540) (2015) 529-533.
- Y. Zhang, W. Zhao, J. Wang, Y. Yuan, Recent progress, challenges and future prospects of applied deep reinforcement learning : A practical perspective in path planning, Neurocomputing 608 (2024) 128423.
- L. Liu, X. Wang, X. Yang, H. Liu, J. Li, P. Wang, Path planning techniques for mobile robots: Review and prospect, Expert Systems with Applications 227 (2023) 120254.
- H. Han, J. Wang, L. Kuang, X. Han, H. Xue, Improved Robot Path Planning Method Based on Deep Reinforcement Learning, Sensors 23(12) (2023) 5622.
- L. Chen, Q. Wang, C. Deng, B. Xie, X. Tuo, G. Jiang, Improved Double Deep Q-Network Algorithm Applied to Multi-Dimensional Environment Path Planning of Hexapod Robots, Sensors 24(7) (2024) 2061.
- Z. Wang, S. Song, S. Cheng, Path planning of mobile robot based on improved double deep Q-network algorithm, Frontiers in Neurorobotics 19 (2025) 1512953. https://doi.org/10.3389/fnbot.2025.1512953
- T. Schaul, J. Quan, I. Antonoglou, D. Silver, Prioritized experience replay, arXiv preprint arXiv:1511.05952 (2015). https://doi.org/10.48550/arXiv.1511.05952
- N. Cheng, P. Wang, G. Zhang, C. Ni, E. Nematov, Prioritized experience replay in path planning via multi-dimensional transition priority fusion, Frontiers in Neurorobotics 17 (2023) 1281166. https://doi.org/10.3389/fnbot.2023.1281166
- W. Hu, Y. Zhou, H.W. Ho, Mobile Robot Navigation Based on Noisy N-Step Dueling Double Deep Q-Network and Prioritized Experience Replay, Electronics 13(12) (2024) 2423.
- D.A. Deguale, L. Yu, M.L. Sinishaw, K. Li, Enhancing Stability and Performance in Mobile Robot Path Planning with PMR-Dueling DQN Algorithm, Sensors 24(5) (2024) 1523.
- E. Massi, J. Barthélemy, J. Mailly, R. Dromnelle, J. Canitrot, E. Poniatowski, B. Girard, M. Khamassi, Model-Based and Model-Free Replay Mechanisms for Reinforcement Learning in Neurorobotics, Frontiers in Neurorobotics 16 (2022) 864380. https://doi.org/10.3389/fnbot.2022.864380
- R. Raj, A. Kos, Intelligent mobile robot navigation in unknown and complex environment using reinforcement learning technique, Scientific Reports 14(1) (2024) 22852.
- H. Liu, Y. Shen, S. Yu, Z. Gao, T. Wu, Deep reinforcement learning for mobile robot path planning, arXiv preprint arXiv:2404.06974 (2024). https://doi.org/10.48550/arXiv.2404.06974
- R. Yang, J. Lyu, Y. Yang, J. Yan, F. Luo, D. Luo, L. Li, X. Li, Bias-reduced multi-step hindsight experience replay for efficient multi-goal reinforcement learning, arXiv preprint arXiv:2102.12962 (2021). https://doi.org/10.48550/arXiv.2102.12962
- J. Wang, H. Han, X. Han, L. Kuang, X. Yang, Reinforcement learning path planning method incorporating multi-step Hindsight Experience Replay for lightweight robots, Displays 84 (2024) 102796.
- M. Quinones-Ramirez, J. Rios-Martinez, V. Uc-Cetina, Robot path planning using deep reinforcement learning, arXiv preprint arXiv:2302.09120 (2023). https://doi.org/10.48550/arXiv.2302.09120
- J. Tan, A Method to Plan the Path of a Robot Utilizing Deep Reinforcement Learning and Multi-Sensory Information Fusion, Applied Artificial Intelligence 37(1) (2023) 2224996.
- Z. Zhang, H. Fu, J. Yang, Y. Lin, Deep reinforcement learning for path planning of autonomous mobile robots in complicated environments, Complex & Intelligent Systems 11(6) (2025) 277.
- Y. Yin, Z. Chen, G. Liu, J. Yin, J. Guo, Autonomous navigation of mobile robots in unknown environments using off-policy reinforcement learning with curriculum learning, Expert Systems with Applications 247 (2024) 123202.
- E. Erkan, M.A. Arserim, Mobile Robot Application with Hierarchical Start Position DQN, Computational Intelligence and Neuroscience 2022(1) (2022) 4115767.
- Y. Yin, Z. Chen, G. Liu, J. Guo, A Mapless Local Path Planning Approach Using Deep Reinforcement Learning Framework, Sensors 23(4) (2023) 2036.
- C. Li, X. Yue, Z. Liu, G. Ma, H. Zhang, Y. Zhou, J. Zhu, A modified dueling DQN algorithm for robot path planning incorporating priority experience replay and artificial potential fields, Applied Intelligence 55(6) (2025) 366.
- M. Gök, Dynamic path planning via Dueling Double Deep Q-Network (D3QN) with prioritized experience replay, Applied Soft Computing 158 (2024) 111503.
- B. Daley, C. Hickert, C. Amato, Stratified experience replay: Correcting multiplicity bias in off-policy reinforcement learning, arXiv preprint arXiv:2102.11319 (2021). https://doi.org/10.48550/arXiv.2102.11319
- Y. Wang, Y. Jia, S. Fan, J. Xiao, Deep reinforcement learning based on balanced stratified prioritized experience replay for customer credit scoring in peer-to-peer lending, Artificial Intelligence Review 57(4) (2024) 93.
- H. Zhong, Z. Wang, TD3 Algorithm of Dynamic Classification Replay Buffer Based PID Parameter Optimization, International Journal of Control, Automation and Systems 22(10) (2024) 3068-3082.
- Y. Zhu, W.Z. Wan Hasan, H.R. Harun Ramli, N.M.H. Norsahperi, M.S. Mohd Kassim, Y. Yao, Deep Reinforcement Learning of Mobile Robot Navigation in Dynamic Environment: A Review, Sensors 25(11) (2025) 3394.
- A. Jonnarth, O. Johansson, J. Zhao, M. Felsberg, Sim-to-real transfer of deep reinforcement learning agents for online coverage path planning, IEEE Access 13 (2025) 106883-106905.
References
B. Singh, R. Kumar, V.P. Singh, Reinforcement learning in robotic applications: a comprehensive survey, Artificial Intelligence Review 55(2) (2022) 945-990.
K. Zhu, T. Zhang, Deep reinforcement learning based mobile robot navigation: A review, Tsinghua Science and Technology 26(5) (2021) 674-691.
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, M. Riedmiller, Playing atari with deep reinforcement learning, arXiv preprint arXiv:1312.5602 (2013). https://doi.org/10.48550/arXiv.1312.5602
V. Mnih, K. Kavukcuoglu, D. Silver, A.A. Rusu, J. Veness, M.G. Bellemare, A. Graves, M. Riedmiller, A.K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, D. Hassabis, Human-level control through deep reinforcement learning, Nature 518(7540) (2015) 529-533.
Y. Zhang, W. Zhao, J. Wang, Y. Yuan, Recent progress, challenges and future prospects of applied deep reinforcement learning : A practical perspective in path planning, Neurocomputing 608 (2024) 128423.
L. Liu, X. Wang, X. Yang, H. Liu, J. Li, P. Wang, Path planning techniques for mobile robots: Review and prospect, Expert Systems with Applications 227 (2023) 120254.
H. Han, J. Wang, L. Kuang, X. Han, H. Xue, Improved Robot Path Planning Method Based on Deep Reinforcement Learning, Sensors 23(12) (2023) 5622.
L. Chen, Q. Wang, C. Deng, B. Xie, X. Tuo, G. Jiang, Improved Double Deep Q-Network Algorithm Applied to Multi-Dimensional Environment Path Planning of Hexapod Robots, Sensors 24(7) (2024) 2061.
Z. Wang, S. Song, S. Cheng, Path planning of mobile robot based on improved double deep Q-network algorithm, Frontiers in Neurorobotics 19 (2025) 1512953. https://doi.org/10.3389/fnbot.2025.1512953
T. Schaul, J. Quan, I. Antonoglou, D. Silver, Prioritized experience replay, arXiv preprint arXiv:1511.05952 (2015). https://doi.org/10.48550/arXiv.1511.05952
N. Cheng, P. Wang, G. Zhang, C. Ni, E. Nematov, Prioritized experience replay in path planning via multi-dimensional transition priority fusion, Frontiers in Neurorobotics 17 (2023) 1281166. https://doi.org/10.3389/fnbot.2023.1281166
W. Hu, Y. Zhou, H.W. Ho, Mobile Robot Navigation Based on Noisy N-Step Dueling Double Deep Q-Network and Prioritized Experience Replay, Electronics 13(12) (2024) 2423.
D.A. Deguale, L. Yu, M.L. Sinishaw, K. Li, Enhancing Stability and Performance in Mobile Robot Path Planning with PMR-Dueling DQN Algorithm, Sensors 24(5) (2024) 1523.
E. Massi, J. Barthélemy, J. Mailly, R. Dromnelle, J. Canitrot, E. Poniatowski, B. Girard, M. Khamassi, Model-Based and Model-Free Replay Mechanisms for Reinforcement Learning in Neurorobotics, Frontiers in Neurorobotics 16 (2022) 864380. https://doi.org/10.3389/fnbot.2022.864380
R. Raj, A. Kos, Intelligent mobile robot navigation in unknown and complex environment using reinforcement learning technique, Scientific Reports 14(1) (2024) 22852.
H. Liu, Y. Shen, S. Yu, Z. Gao, T. Wu, Deep reinforcement learning for mobile robot path planning, arXiv preprint arXiv:2404.06974 (2024). https://doi.org/10.48550/arXiv.2404.06974
R. Yang, J. Lyu, Y. Yang, J. Yan, F. Luo, D. Luo, L. Li, X. Li, Bias-reduced multi-step hindsight experience replay for efficient multi-goal reinforcement learning, arXiv preprint arXiv:2102.12962 (2021). https://doi.org/10.48550/arXiv.2102.12962
J. Wang, H. Han, X. Han, L. Kuang, X. Yang, Reinforcement learning path planning method incorporating multi-step Hindsight Experience Replay for lightweight robots, Displays 84 (2024) 102796.
M. Quinones-Ramirez, J. Rios-Martinez, V. Uc-Cetina, Robot path planning using deep reinforcement learning, arXiv preprint arXiv:2302.09120 (2023). https://doi.org/10.48550/arXiv.2302.09120
J. Tan, A Method to Plan the Path of a Robot Utilizing Deep Reinforcement Learning and Multi-Sensory Information Fusion, Applied Artificial Intelligence 37(1) (2023) 2224996.
Z. Zhang, H. Fu, J. Yang, Y. Lin, Deep reinforcement learning for path planning of autonomous mobile robots in complicated environments, Complex & Intelligent Systems 11(6) (2025) 277.
Y. Yin, Z. Chen, G. Liu, J. Yin, J. Guo, Autonomous navigation of mobile robots in unknown environments using off-policy reinforcement learning with curriculum learning, Expert Systems with Applications 247 (2024) 123202.
E. Erkan, M.A. Arserim, Mobile Robot Application with Hierarchical Start Position DQN, Computational Intelligence and Neuroscience 2022(1) (2022) 4115767.
Y. Yin, Z. Chen, G. Liu, J. Guo, A Mapless Local Path Planning Approach Using Deep Reinforcement Learning Framework, Sensors 23(4) (2023) 2036.
C. Li, X. Yue, Z. Liu, G. Ma, H. Zhang, Y. Zhou, J. Zhu, A modified dueling DQN algorithm for robot path planning incorporating priority experience replay and artificial potential fields, Applied Intelligence 55(6) (2025) 366.
M. Gök, Dynamic path planning via Dueling Double Deep Q-Network (D3QN) with prioritized experience replay, Applied Soft Computing 158 (2024) 111503.
B. Daley, C. Hickert, C. Amato, Stratified experience replay: Correcting multiplicity bias in off-policy reinforcement learning, arXiv preprint arXiv:2102.11319 (2021). https://doi.org/10.48550/arXiv.2102.11319
Y. Wang, Y. Jia, S. Fan, J. Xiao, Deep reinforcement learning based on balanced stratified prioritized experience replay for customer credit scoring in peer-to-peer lending, Artificial Intelligence Review 57(4) (2024) 93.
H. Zhong, Z. Wang, TD3 Algorithm of Dynamic Classification Replay Buffer Based PID Parameter Optimization, International Journal of Control, Automation and Systems 22(10) (2024) 3068-3082.
Y. Zhu, W.Z. Wan Hasan, H.R. Harun Ramli, N.M.H. Norsahperi, M.S. Mohd Kassim, Y. Yao, Deep Reinforcement Learning of Mobile Robot Navigation in Dynamic Environment: A Review, Sensors 25(11) (2025) 3394.
A. Jonnarth, O. Johansson, J. Zhao, M. Felsberg, Sim-to-real transfer of deep reinforcement learning agents for online coverage path planning, IEEE Access 13 (2025) 106883-106905.