Rethinking Ratio-Based Trust Regions for Policy Optimization in Multi-Agent Reinforcement Learning
Rethinking policy updates to make cooperative multi-agent reinforcement learning more stable.

About me
I’m CK, a PhD student at Northeastern University and a research engineer at STR.
My research focuses on multi-agent reinforcement learning, learning under partial observability, and policy optimization. At Northeastern, I work with Christopher Amato in the Lab for Learning and Planning in Robotics. At STR, I work on applied research projects spanning reinforcement learning and computer vision.
Outside of work, I enjoy lifting weights, playing video games, reading, photography, and improving my iced latte game.
01 / Publications
Rethinking policy updates to make cooperative multi-agent reinforcement learning more stable.
Using state information during training to improve offline reinforcement learning with incomplete observations.
02 / Work
Here are a few highlights from my work at STR, where I apply reinforcement learning to a variety of problems.
Algorithm leadership
I led the algorithm development that earned STR a Phase 2 award in DARPA’s Artificial Intelligence Reinforcements (AIR) program.
About DARPA AIRSimulation transfer
I developed and demonstrated techniques for transferring RL agents trained in lower-fidelity simulations to higher-fidelity models of real-world systems without additional training.
From PhD to programs
I transitioned JAX-based multi-agent RL research and software infrastructure from my PhD into active STR programs. This work now supports a range of programs across the company.
Adversarial learning
I developed RL-based human surrogate agents for mixed-reality environments using VLM perception, alongside adversarial agents trained to discover novel visual attacks and expose weaknesses in surrogate perception and robustness
03 / Software
Frameworks and environments for training and evaluating RL agents.
Training framework
A JAX framework for multi-agent policy-gradient reinforcement learning research.
In developmentSimulation & benchmarks
JAX-native aerial environments for studying decentralized control and coordination in continuous three-dimensional space.
Related paper