Research Focus
- Fundamental hardness of multi-step planning Theory: For problems with long planning horizons, I design algorithms with regrets independent of the horizon, showing that multi-step problems can be no harder than single-step problems.
- Instance-dependent regrets Theory: I design algorithms with regrets depending on quantities that characterize each problem instance, while still recovering worst-case guarantees.
Related Publications
- Knowledge acquisition and manipulation: I study failure patterns during knowledge retrieval in LLMs and design theory-inspired approaches to improve performance.
Related Publications
- Reasoning: I research infrastructure, data, and algorithmic perspectives for building accurate and efficient RL systems for LLMs.
- Reinforcement learning from human feedback Theory: I design RLHF algorithms with provably faster convergence rates and better benchmark performance, and study the separation between reward feedback and preference feedback.
- Agent: I train LLMs to autonomously explore environments using agentic scaffolds.
Related Publications