Publications

*: indicating equal contribution or alphabetic ordering.

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence RL Theory

Visored: A Controlled-Natural-Language Prover for LLM-Generated Mathematics LLM

Unregularized Linear Convergence in Zero-Sum Game from Preference Feedback RL Theory

RLAX: Large-Scale, Distributed Reinforcement Learning for Large Language Models on TPUs RL LLM

RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments RL LLM

ICML 2026

The Ramón Llull’s Thinking Machine for Automated Ideation LLM

Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs RL Theory

NeurIPS 2025

Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO RL Theory LLM

ICML 2026

CASCADE Your Datasets for Cross-Mode Knowledge Retrieval of Language Models LLM

COLM 2025

Extragradient Preference Optimization (EGPO): Beyond Last-Iterate Convergence for Nash Learning from Human Feedback RL Theory LLM

COLM 2025

Transformers are Efficient Compilers, Provably Theory LLM

COLM 2025

The Crucial Role of Samplers in Online Direct Preference Optimization RL Theory LLM

ICLR 2025

Multi-Agent Reinforcement Learning from Human Feedback: Data Coverage and Algorithmic Techniques RL Theory

Reflect-RL: Two-Player Online RL Fine-Tuning for LMs RL LLM

ACL 2024

Free from Bellman Completeness: Trajectory Stitching via Model-based Return-conditioned Supervised Learning RL Theory

ICLR 2024

Sharp Variance-Dependent Bounds in Reinforcement Learning: Best of Both Worlds in Stochastic and Deterministic Environments RL Theory

ICML 2023

Horizon-Free and Variance-Dependent Reinforcement Learning for Latent Markov Decision Processes RL Theory

ICML 2023

Understanding Curriculum Learning in Policy Optimization for Online Combinatorial Optimization RL Theory

TMLR

Stochastic Shortest Path: Minimax, Parameter-Free and Towards Horizon-Free Regret RL Theory

NeurIPS 2021 (Spotlight, 3% acceptance rate)