Multi-Agent Path Planning and Optimization Using Q-learning

Publication Date

10-2025

Document Type

Conference Proceeding

Exhibition/Performance Dates

18-21 August 2025

Publication Title

IEEE 8th International Conference on Robotics and Control Engineering (IRCE)

Conference Location

Kunming, China

DOI

https://doi.org/10.1109/IRCE66030.2025.11203139

First Page

486

Last Page

495

Abstract

Path planning and optimization play an important role in robotics and autonomous multi-agent systems, especially in cluttered 2D environments with strict smoothness and clearance requirements. This paper introduces a novel twophase hybrid framework combining reinforcement learning and convex optimization to generate diverse, smooth, and collision-free paths. The first phase, which constitutes the core contribution, employs a Q-learning-based method on a visibility graph to learn an adaptive policy that generalizes across multiple start-goal pairs and varying environmental complexities. This approach enables path diversity, robust obstacle avoidance, and environment-specific adaptation without relying on hand-crafted heuristics. The second phase refines these initial trajectories using Multi-agent Interleaving Convex Optimization (MICO) to improve curvature, length, and clearance. Experimental results demonstrate that MICO outperforms baseline convex optimization methods and provides flexible tuning of path characteristics through multiobjective optimization. The paths are represented as piecewise Bézier curves, ensuring smooth and feasible motion for multi-agent navigation.

Keywords

Q-learning, Navigation, Trajectory tracking, Diversity reception, Convex functions, Hybrid power systems, Collision avoidance, Robots, Tuning, Trajectory optimization

Department

Computer Science

Share

COinS