My research goal is to make creative exploration reliable in real environments, so that an agent can reach behavior it did not begin with. My work so far centers on multi-turn RL, on credit assignment and stable optimization over long agent-environment interactions, and I am equally interested in the reward models and world models that shape the feedback an agent receives and the experience it learns from. Beyond these, I am drawn to how the training process itself is designed, from the architecture of the learning algorithm to the role a capable model can play in directing training. What ties these together is efficiency, where I expect gains in degree to accumulate into differences in kind.
I am looking for research internship and collaboration opportunities in reinforcement learning, LLM agents. Feel free to contact me via email.