Xiaohan Wang
XW

Xiaohan Wang 王霄汉

LLM Post-Training / Reinforcement Learning  ·  Meituan

Experience

2024 – Present
LLM Post-Training Engineer Meituan 北斗人才计划
2023 – 2024
Joint Ph.D. KTH Royal Institute of Technology 瑞典皇家理工学院
2019 – 2024
Ph.D. Beihang University 北京航空航天大学
2015 – 2019
B.Eng. Beihang University 北京航空航天大学

About Me

Currently, my research focuses on LLM post-training techniques and algorithms, with a specific emphasis on search agent, tool use, reasoning, and instruction following. My expertise spans the full LLM pipeline—from training and inference to evaluation—covering key areas such as RLVR, reward modeling, data synthesis, and multi-domain training. I have extensive hands-on experience with large-scale GPU clusters (~15M GPU hours) and end-to-end training of multiple 100B+ parameter foundation models. Prior to this, I completed my Ph.D. with a focus on applying deep reinforcement learning (DRL) to cloud environment scheduling challenges, supervised by Prof. Lin Zhang. My doctoral work aimed at utilizing neural network generalization for decision-making to optimize solution flexibility and reduce response times, drawing extensively on MDP modeling, uplift modeling, and online/offline DRL.

RLVR Reward Modeling Agentic RL Data Synthesis Multi-domain Training Search Agent Tool Use Reasoning Instruction Following

Recent Publications

EMNLP 2026

  • ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay Co-First Corresponding
  • HiDiffTIR: Hierarchical Difficulty-Aware Policy Optimization for Multi-Turn Tool-Integrated Reasoning Co-First
  • MoE Sparsity Begets Monosemanticity: Insights into Understanding and Fine-Tuning Co-First
  • When Not to Imitate: Boundary-Aware Skill Memory for Reliable Tool-Use LLM Agents Corresponding
  • Joint Training of Multi-Token Prediction in Reinforcement Learning via Optimal Coefficient Calibration
  • ToolForge: A Data Synthesis Pipeline for Multi-Hop, Multi-Turn, and Self-Reflective Tool-Use Data
  • UniMem: Complementary Episodic-to-Parametric Memory for Boundary-Agnostic Task Streams
  • π-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data

KDD 2026

  • CDRRM: Contrast-Driven Rubric Generation for Reliable and Interpretable Reward Modeling (Link) Corresponding
  • LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services (Link)

ICML 2026

  • ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning (Link) Co-First Corresponding
  • Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards (Link)
  • SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning (Link)
  • GRASP: Graph Reasoning via Agentic Solving and Probing of LLMs

ACL 2026

  • AutoSearch: Adaptive Search Depth for Efficient Agentic RAG via Reinforcement Learning (Link)

ICLR 2026

  • ResT: Reshaping Token-Level Policy Gradients for Tool-Use Large Language Models (Link) Co-First
  • LogiConBench: Benchmarking Logical Consistencies of LLMs
  • MAD-Logic: Multi-Agent Debate Enhances Symbolic Translation and Reasoning
  • SAE as a Crystal Ball: Interpretable Features Predict Cross-domain Transferability of LLMs without Training (Link)

ICLR 2026 Workshop

  • MTIR-SQL: Multi-turn Tool-Integrated Reasoning Reinforcement Learning for Text-to-SQL (Link)