Yang Wan

万扬
Ph.D. Student
College of Computer Science and Technology, Zhejiang University

About

I am a Ph.D. candidate in the College of Computer Science and Technology at Zhejiang University, advised by Prof. Linchao Zhu. I received my B.Eng. (Honors) in Computer Science and Technology from Hangzhou Dianzi University.

My research goal is to make creative exploration reliable in real environments, so that an agent can reach behavior it did not begin with. My work so far centers on multi-turn RL, on credit assignment and stable optimization over long agent-environment interactions, and I am equally interested in the reward models and world models that shape the feedback an agent receives and the experience it learns from. Beyond these, I am drawn to how the training process itself is designed, from the architecture of the learning algorithm to the role a capable model can play in directing training. What ties these together is efficiency, where I expect gains in degree to accumulate into differences in kind.

I am looking for research internship and collaboration opportunities in reinforcement learning, LLM agents. Feel free to contact me via email.

Selected Papers

SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents
Yang Wan, Zhenhao Zhang, Jierui Wang, Linchao Zhu
Preprint
Mitigating Conversational Inertia in Multi-Turn Agents
Yang Wan, Zheng Cao, Zhenhao Zhang, Zhengwen Zeng, Shuheng Shen, Changhua Meng, Linchao Zhu
ICML 2026 (Accepted)

Competitive Programming

I have been involved in competitive programming (ICPC and CCPC) for several years, earning a regional gold medal and several silver medals as a contestant, and I serve as a judge for CCPC regional contests.