Research Focus

Reinforcement Learning

RL
  • Fundamental hardness of multi-step planning Theory: For problems with long planning horizons, I design algorithms with regrets independent of the horizon, showing that multi-step problems can be no harder than single-step problems.
  • Instance-dependent regrets Theory: I design algorithms with regrets depending on quantities that characterize each problem instance, while still recovering worst-case guarantees.

Related Publications

Large Language Models

LLM
  • Knowledge acquisition and manipulation: I study failure patterns during knowledge retrieval in LLMs and design theory-inspired approaches to improve performance.

Related Publications

RL × LLM

  • Reasoning: I research infrastructure, data, and algorithmic perspectives for building accurate and efficient RL systems for LLMs.
  • Reinforcement learning from human feedback Theory: I design RLHF algorithms with provably faster convergence rates and better benchmark performance, and study the separation between reward feedback and preference feedback.
  • Agent: I train LLMs to autonomously explore environments using agentic scaffolds.

Related Publications