Sarvan Gill
-
BASc (University of British Columbia, 2022)
Topic
Robust Policy Learning for Robotic Systems
Department of Mechanical Engineering
Date & location
-
Monday, July 13, 2026
-
2:00 P.M.
-
Virtual Defence
Reviewers
Supervisory Committee
-
Dr. Daniela Constantinescu, Department of Mechanical Engineering, University of Victoria (Supervisor)
-
Dr. Yang Shi, Department of Mechanical Engineering, UVic (Member)
External Examiner
- Dr. Pan Agathoklis, Department of Electrical and Computer Engineering, University of Victoria
Chair of Oral Examination
-
Dr. Avigail Eisenberg, Department of Political Science, UVic
Abstract
This thesis tackles the problem of stability in reinforcement learning (RL) based controllers for robotic systems. The thesis proposes the state-action Lyapunov function to provide stability guarantees and sample efficiency improvements to reinforcement learning algorithms. The thesis further explores how the learned Lyapunov function can not only provide stability certificates but also improve the performance of the reinforcement learning policies for robots.
Reinforcement learning is widely used to find control policies for complex robotic problems. Traditional reinforcement learning lacks the ability to provide stability guarantees and physical systems such as robots require these guarantees due to physical safety concerns. Recent methods investigate learning Lyapunov functions alongside RL algorithms but the current self-learned Lyapunov functions are sample inefficient due to their on-policy nature. This thesis provides the first stable RL algorithm that learns Lyapunov functions using off-policy data to certify stability and improving the sample efficiency of existing stable RL methods. The proposed framework extends the traditional Lyapunov function to a function of both state and action, hence the name state-action Lyapunov function. The thesis provides a proof displaying that the state-action Lyapuov guarantees stability in the sense of Lyapunov. The state action Lyapunov is then combined with modern RL algorithms using an off-policy Lie derivative estimator to provide sample efficiency improvements to stable RL.
While the goal of stable RL is to certify the stability of learned RL policies, this work argues that stable RL can also provide performance improvements to RL policies as well. There is a general discrepancy between the scale of learned neural Lyapunov functions and RL value functions. The thesis proposes 3 scaling methods claiming that scaling the learned Lyapunov functions will increase the performance of the RL policies. 3 robot simulations are employed to validate this claim, testing the proposed scaling method with the aforementioned state-action Lyapunov against other RL and stable RL methods. Results illustrate that the scaling methods combined with the state-action Lyapunov function provide both sample-efficiency and performance improvements compared to other comparable methods.