I develop learning agents — AI systems that learn from interactive experience, reason over complex environments, and continually self-improve. My research bridges Reinforcement Learning and Foundation Models to develop autonomous decision-makers that are efficient, adaptive, and grounded in safe and responsible AI principles.
I received my Ph.D. from UTS, advised by Prof. Dacheng Tao, and then worked as a postdoc with Prof. Trevor Cohn at the University of Melbourne. I held research and internship positions at Tencent (Robotics X and AI Lab), CSIRO, and Microsoft Research Asia before.
Building intelligent systems capable of human-like learning, reasoning, and planning in complex environments.
Advancing Large Language Models (LLMs) through data-centric training, robust data selection, and enhancing mathematical and scientific reasoning.
Advancing foundational algorithms for decision-making, tackling core challenges in offline RL, multi-goal learning, and sparse rewards.
Investigating multilateral fairness and social bias in LLMs, alongside developing Safe RL algorithms to guarantee safety constraints in deployment.