Xiaoyun Zhang张啸云
M.S. student at the Institute of Computing Technology, CAS, working on LLM Reinforcement Learning and Code / Search Agentic Intelligence. Currently building Agent Swarm RL at Moonshot AI; previously StepFun & Meituan. First-author papers at ACL · EMNLP ×3 · AAAI (Oral).
Training agents
to think, code & search.
I work on LLM Reinforcement Learning and Code / Search Agentic Intelligence — currently building Agent Swarm RL training & infrastructure for general agents on the RL Team at Moonshot AI.
Before Moonshot, I was an AgentRL researcher at StepFun and an LLM algorithm intern at Meituan (LongCat). My first-author work appears at ACL / EMNLP ×3 / AAAI (Oral), and I was a core participant in the Kimi K3 · K2.5 and Step 3.5 Flash · Step-DeepResearch technical reports.
Selected papers.
FIRST-AUTHOR / CO-FIRST-AUTHOR ONLY — FULL LIST ON SCHOLAR ↗
Revisiting Entropy Regularization: Adaptive Coefficient Unlocks Its Potential for LLM Reinforcement Learning
RLVR training suffers from policy entropy collapse — an adaptive entropy coefficient unlocks stable exploration for LLM RL.
TRUST-SQL: Tool-Integrated Multi-Turn Reinforcement Learning for Text-to-SQL over Unknown Schemas
Real enterprise databases break the Full-Schema Assumption — TRUST-SQL explores unknown schemas with tool-integrated multi-turn RL.
Safety Alignment of Large Language Models via Contrasting Safe and Harmful Distributions
Adversarial Contrastive Decoding — a training-free framework that contrasts safe and harmful distributions via opposite soft system prompts.
When to Continue Thinking: Adaptive Thinking Mode Switching for Efficient Reasoning
Don't pay for overthinking — quantifying reasoning upper bounds and switching thinking modes adaptively for efficient inference.
A Reasoner for Real-World Event Detection: Scaling RL via Adaptive Perplexity-Aware Sampling Strategy
APARL — dual-loop curriculum RL with perplexity-aware sampling; +17.19% F1 on real-world industrial dialogues.
Where I've trained.
PRESENT
Moonshot AI 月之暗面
Agent Swarm RL Intern · RL TeamReinforcement learning training & infrastructure for general agents — scaling swarms of agents that learn together.
2026.02
StepFun 阶跃星辰
AgentRL Researcher · Foundation Model AlgorithmsAgent RL training and infrastructure for frontier foundation models.
2025.08
Meituan 美团
LLM Algorithm Intern · LongCat InteractionGenerative analysis & detection with LLMs, and RL framework construction — shipped three EMNLP papers along the way.
炼丹之余,
也写笔记。
白天炼丹,晚上复盘。我在小红书上以「新世纪炼丹炉战士」的身份记录大模型训练日常 —— 模型发布、论文解读、实习踩坑,还有和 Bad Case 的日常搏斗。
小红书FOLLOW 新世纪炼丹炉战士