Zhengyang Tang

Ph.D. Candidate, The Chinese University of Hong Kong, Shenzhen

I build executable and verifiable agents for reasoning, coding, and phone-use environments.

About

I am a Ph.D. candidate at The Chinese University of Hong Kong, Shenzhen, advised by Prof. Benyou Wang and co-supervised by Ming Yan. My research focuses on executable and verifiable agentic systems built with large language models: systems that can reason, use tools, write code, and act in realistic environments.

I am especially interested in scalable post-training, environment construction, harness engineering, and verifiable feedback for reasoning agents, executable coding agents, and phone-use agents.

I have had the opportunity to work with teams including Tencent Hunyuan, Moonshot AI (Kimi), Qwen, and Microsoft Research Asia.

Research Themes

Reasoning Agents and Self-Improvement
Scalable data generation, post-training, critique, tool use, and capability acquisition for reasoning models, including Qwen3, MathScale, GLAN, ALAN, SCRIT, and CoRT.
Executable Coding and Tool-Use Agents
Agents and benchmarks where success is checked by execution, artifacts, solver traces, and interactive behavior, including Kimi K2.5, ORLM, CALM/STORM, and GameCraft-Bench.
Phone-Use Agents and Mobile Environments
Training environments, mixed-action harnesses, and safety/privacy benchmarks for reliable phone-use agents, including PhoneBuddy, PhoneWorld, PhoneHarness, PhonePrivacy, and PhoneSafety.

Selected Publications

For a full list, please see Google Scholar. (* denotes equal contribution.)

Kimi K2.5: Visual Agentic Intelligence
Kimi Team (including Zhengyang Tang)
Technical Report, 2026
Qwen3 Technical Report
Qwen Team (including Zhengyang Tang)
Technical Report, 2025
PhoneBuddy: Training Open Models for Agentic Phone Use
Zhengyang Tang*, Xin Lai*, Pengyuan Lyu*, Xinyuan Wang*, Tianyi Bai*, Chenxin Li*, et al.
Preprint, 2026
GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?
Tongxu Luo*, Rongsheng Wang*, Jiaxi Bi*, Chenming Xu*, Zhengyang Tang*, et al.
arXiv Preprint, 2026
PhoneWorld: Scaling Phone-Use Agent Environments
Zhengyang Tang*, Yuxuan Liu*, Xin Lai*, Junyi Li*, et al.
arXiv Preprint, 2026
Do Phone-Use Agents Respect Your Privacy?
Zhengyang Tang*, Ke Ji*, Xidong Wang*, Zihan Ye*, Xinyuan Wang, Yiduo Guo, et al.
arXiv Preprint, 2026
Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents
Zhengyang Tang*, Yi Zhang*, Chenxin Li, Xin Lai, et al.
arXiv Preprint, 2026
PhoneHarness: A Mixed-Action Orchestration Harness and Benchmark for Phone Agents across CLI, GUI, and MCP Tools
Jason*, Zhengyao Fang*, Zhengyang Tang*, Pengyuan Lyu*, et al.
Preprint, 2026
CALM Before the STORM: Unlocking Native Reasoning for Optimization Modeling
Zhengyang Tang*, Zihan Ye*, Chenyu Huang*, Xuhan Huang, Chengpeng Li, Sihang Li, et al.
ICML, 2026
CoRT: Code-integrated Reasoning within Thinking
Chengpeng Li*, Zhengyang Tang*, Ziniu Li*, Mingfeng Xue, Keqin Bao, et al.
NeurIPS, 2025
Self-Evolving Critique Abilities in Large Language Models (SCRIT)
Zhengyang Tang*, Ziniu Li*, Zhenyang Xiao*, Tian Ding, Ruoyu Sun, Benyou Wang, et al.
COLM, 2025
ORLM: A Customizable Framework in Training Large Models for Automated Optimization Modeling
Chenyu Huang*, Zhengyang Tang*, Shixi Hu, Ruoqing Jiang, Xin Zheng, Dongdong Ge, Benyou Wang, Zizhuo Wang
Operations Research, 2025
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
Zhengyang Tang, Xingxing Zhang, Benyou Wang, Furu Wei
ICML, 2024
Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models (GLAN)
Haoran Li*, Qingxiu Dong*, Zhengyang Tang*, Chaojun Wang*, Xingxing Zhang, et al.
TMLR, 2025
Second Language (Arabic) Acquisition of LLMs via Progressive Vocabulary Expansion
Jianqing Zhu*, Huang Huang*, Zhihang Lin*, Juhao Liang*, Zhengyang Tang*, Khalid Almubarak, et al.
ACL 2025 (Oral & Panel)
DPTDR: Deep Prompt Tuning for Dense Passage Retrieval
Zhengyang Tang, Benyou Wang, Ting Yao
COLING, 2022

Talks

Invited Talk on MathScale, ICML, Jul 2024. [Slides]
Invited Talk on Dense Passage Retrieval, Baidu Search Department, 2022.

Selected Patents

Service

Reviewer for ACL Rolling Review (since 2024), COLM, NeurIPS, ICLR, and ICML.

Awards / Honors

Runner-up, LIC 2022 Passage-Ranking Competition (2nd place / 793 teams), Baidu & CCF, 2022.
Top-3 Ranking, MS MARCO Passage Ranking Leaderboard, Microsoft Research, 2022.