Experience
About Me
Currently, my research focuses on LLM post-training techniques and algorithms, with a specific emphasis on search agent, tool use, reasoning, and instruction following. My expertise spans the full LLM pipeline—from training and inference to evaluation—covering key areas such as RLVR, reward modeling, data synthesis, and multi-domain training. I have extensive hands-on experience with large-scale GPU clusters (~15M GPU hours) and end-to-end training of multiple 100B+ parameter foundation models. Prior to this, I completed my Ph.D. with a focus on applying deep reinforcement learning (DRL) to cloud environment scheduling challenges, supervised by Prof. Lin Zhang. My doctoral work aimed at utilizing neural network generalization for decision-making to optimize solution flexibility and reduce response times, drawing extensively on MDP modeling, uplift modeling, and online/offline DRL.
Recent Publications
EMNLP 2026
- ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay Co-First Corresponding
- HiDiffTIR: Hierarchical Difficulty-Aware Policy Optimization for Multi-Turn Tool-Integrated Reasoning Co-First
- MoE Sparsity Begets Monosemanticity: Insights into Understanding and Fine-Tuning Co-First
- When Not to Imitate: Boundary-Aware Skill Memory for Reliable Tool-Use LLM Agents Corresponding
- Joint Training of Multi-Token Prediction in Reinforcement Learning via Optimal Coefficient Calibration
- ToolForge: A Data Synthesis Pipeline for Multi-Hop, Multi-Turn, and Self-Reflective Tool-Use Data
- UniMem: Complementary Episodic-to-Parametric Memory for Boundary-Agnostic Task Streams
- π-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data
KDD 2026
ICML 2026
- ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning (Link) Co-First Corresponding
- Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards (Link)
- SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning (Link)
- GRASP: Graph Reasoning via Agentic Solving and Probing of LLMs
ACL 2026
- AutoSearch: Adaptive Search Depth for Efficient Agentic RAG via Reinforcement Learning (Link)
ICLR 2026
- ResT: Reshaping Token-Level Policy Gradients for Tool-Use Large Language Models (Link) Co-First
- LogiConBench: Benchmarking Logical Consistencies of LLMs
- MAD-Logic: Multi-Agent Debate Enhances Symbolic Translation and Reasoning
- SAE as a Crystal Ball: Interpretable Features Predict Cross-domain Transferability of LLMs without Training (Link)
ICLR 2026 Workshop
- MTIR-SQL: Multi-turn Tool-Integrated Reasoning Reinforcement Learning for Text-to-SQL (Link)
→ Full list on Google Scholar