AGENT SWARM RL INTERN @ MOONSHOT AI — BEIJING

Xiaoyun Zhang张啸云

RESEARCHING /

M.S. student at the Institute of Computing Technology, CAS, working on LLM Reinforcement Learning and Code / Search Agentic Intelligence. Currently building Agent Swarm RL at Moonshot AI; previously StepFun & Meituan. First-author papers at ACL · EMNLP ×3 · AAAI (Oral).

0+CITATIONS
0H-INDEX
01ST-AUTHOR PAPERS
0FRONTIER MODEL REPORTS
Xiaoyun Zhang 张啸云
ICT · CAS → MOONSHOT AI BEIJING 39.9°N
SCROLL
01ABOUT

Training agents
to think, code & search.

I work on LLM Reinforcement Learning and Code / Search Agentic Intelligence — currently building Agent Swarm RL training & infrastructure for general agents on the RL Team at Moonshot AI.

Before Moonshot, I was an AgentRL researcher at StepFun and an LLM algorithm intern at Meituan (LongCat). My first-author work appears at ACL / EMNLP ×3 / AAAI (Oral), and I was a core participant in the Kimi K3 · K2.5 and Step 3.5 Flash · Step-DeepResearch technical reports.

02RESEARCH

Selected papers.

FIRST-AUTHOR / CO-FIRST-AUTHOR ONLY — FULL LIST ON SCHOLAR ↗

/01
ACL 20261ST AUTHOR

Revisiting Entropy Regularization: Adaptive Coefficient Unlocks Its Potential for LLM Reinforcement Learning

RLVR training suffers from policy entropy collapse — an adaptive entropy coefficient unlocks stable exploration for LLM RL.

Xiaoyun Zhang, Xiaojian Yuan, Di Huang, Wang You, Chen Hu, Jingqing Ruan, Ai Jian, Kejiang Chen, Xing Hu
/02
EMNLP 2026CO-1ST AUTHOR

TRUST-SQL: Tool-Integrated Multi-Turn Reinforcement Learning for Text-to-SQL over Unknown Schemas

Real enterprise databases break the Full-Schema Assumption — TRUST-SQL explores unknown schemas with tool-integrated multi-turn RL.

Ai Jian*, Xiaoyun Zhang*, Wanrou Du, Jingqing Ruan, Jiangbo Pei, Weipeng Zhang, Ke Zeng, Xunliang Cai
/03
AAAI 2026 · ORALCO-1ST AUTHOR

Safety Alignment of Large Language Models via Contrasting Safe and Harmful Distributions

Adversarial Contrastive Decoding — a training-free framework that contrasts safe and harmful distributions via opposite soft system prompts.

Xiaoyun Zhang*, Zhengyue Zhao*, Wenxuan Shi, Kaidi Xu, Di Huang, Xing Hu
/04
EMNLP 20251ST AUTHOR

When to Continue Thinking: Adaptive Thinking Mode Switching for Efficient Reasoning

Don't pay for overthinking — quantifying reasoning upper bounds and switching thinking modes adaptively for efficient inference.

Xiaoyun Zhang, Jingqing Ruan, Xing Ma, Yawen Zhu, Haodong Zhao, Hao Li, Jiansong Chen, Ke Zeng, Xunliang Cai
/05
EMNLP 2025 · INDUSTRY1ST AUTHOR

A Reasoner for Real-World Event Detection: Scaling RL via Adaptive Perplexity-Aware Sampling Strategy

APARL — dual-loop curriculum RL with perplexity-aware sampling; +17.19% F1 on real-world industrial dialogues.

Xiaoyun Zhang, Jingqing Ruan, Xing Ma, Yawen Zhu, Jiansong Chen, Ke Zeng, Xunliang Cai
TECHNICAL REPORTS — CORE PARTICIPANT / COLLABORATION
03EXPERIENCE

Where I've trained.

2026.03 —
PRESENT

Moonshot AI 月之暗面

Agent Swarm RL Intern · RL Team

Reinforcement learning training & infrastructure for general agents — scaling swarms of agents that learn together.

Kimi K3 ↗ Kimi K2.5 ↗ Agent Swarm RL
2025.08 —
2026.02

StepFun 阶跃星辰

AgentRL Researcher · Foundation Model Algorithms

Agent RL training and infrastructure for frontier foundation models.

2025.01 —
2025.08

Meituan 美团

LLM Algorithm Intern · LongCat Interaction

Generative analysis & detection with LLMs, and RL framework construction — shipped three EMNLP papers along the way.

04OFF THE CLOCK

炼丹之余,
也写笔记

白天炼丹,晚上复盘。我在小红书上以「新世纪炼丹炉战士」的身份记录大模型训练日常 —— 模型发布、论文解读、实习踩坑,还有和 Bad Case 的日常搏斗。

FOLLOW 新世纪炼丹炉战士