Zhengyang Tang

Research Scientist, Tencent Hunyuan

About

I am a Research Scientist in the Qingyun Talent Program at Tencent Hunyuan. I received my Ph.D. in Computer and Information Engineering from The Chinese University of Hong Kong, Shenzhen, where I was advised by Prof. Benyou Wang.

Research

My current interests lie in multimodal intelligence and agents that can understand, create, and act across digital and physical environments.

I study how such agents learn from interaction—how to construct environments and tasks at scale, verify outcomes, and turn experience into training signals. Rather than targeting a single interface or domain, I am interested in capabilities that transfer across tools, applications, and embodiments.

Selected Projects

Kimi K2.5 explores visual agentic intelligence through reasoning, coding, and tool use. [Paper]
GameCraft-Bench evaluates whether agents can build playable games end-to-end in a real game engine. [Code]
PhoneBuddy and PhoneWorld develop open models and scalable interactive environments for agentic phone use. [PhoneBuddy Code] [PhoneWorld Code]
ORLM pioneers open-source LLMs for automated optimization modeling, introducing OR-Instruct and the first industrial benchmark, IndustryOR; the work became Operations Research's first paper on LLMs. [Paper] [Code & Models]
MathScale and GLAN scale synthetic instruction generation: MathScale creates two million mathematical question-answer pairs through concept graphs, while GLAN spans diverse disciplines without task-specific seed examples. [MathScale Code]
懂模帝 is a reproducible evaluation and public leaderboard platform for AI models, launched with friends. [Launch Post]

Openings & Collaboration

Our team is looking for interns and full-time members interested in multimodal models and agents. Please feel free to reach out if you would like to work with us.

I am also open to research collaborations in related areas. I would be glad to hear from faculty members and researchers with shared interests.

Selected Publications

For a full list, please see Google Scholar. (* denotes equal contribution.)

Kimi K2.5: Visual Agentic Intelligence
Kimi Team (including Zhengyang Tang)
Technical Report, 2026
Qwen3 Technical Report
Qwen Team (including Zhengyang Tang)
Technical Report, 2025
PhoneBuddy: Training Open Models for Agentic Phone Use
Zhengyang Tang*, Xin Lai*, Pengyuan Lyu*, Xinyuan Wang*, Tianyi Bai*, Chenxin Li*, et al.
Preprint, 2026
GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?
Tongxu Luo*, Rongsheng Wang*, Jiaxi Bi*, Chenming Xu*, Zhengyang Tang*, et al.
arXiv Preprint, 2026
PhoneWorld: Scaling Phone-Use Agent Environments
Zhengyang Tang*, Yuxuan Liu*, Xin Lai*, Junyi Li*, et al.
arXiv Preprint, 2026
Do Phone-Use Agents Respect Your Privacy?
Zhengyang Tang*, Ke Ji*, Xidong Wang*, Zihan Ye*, Xinyuan Wang, Yiduo Guo, et al.
arXiv Preprint, 2026
Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents
Zhengyang Tang*, Yi Zhang*, Chenxin Li, Xin Lai, et al.
arXiv Preprint, 2026
PhoneHarness: A Mixed-Action Orchestration Harness and Benchmark for Phone Agents across CLI, GUI, and MCP Tools
Jason*, Zhengyao Fang*, Zhengyang Tang*, Pengyuan Lyu*, et al.
Preprint, 2026
CALM Before the STORM: Unlocking Native Reasoning for Optimization Modeling
Zhengyang Tang*, Zihan Ye*, Chenyu Huang*, Xuhan Huang, Chengpeng Li, Sihang Li, et al.
ICML, 2026
CoRT: Code-integrated Reasoning within Thinking
Chengpeng Li*, Zhengyang Tang*, Ziniu Li*, Mingfeng Xue, Keqin Bao, et al.
NeurIPS, 2025
Self-Evolving Critique Abilities in Large Language Models (SCRIT)
Zhengyang Tang*, Ziniu Li*, Zhenyang Xiao*, Tian Ding, Ruoyu Sun, Benyou Wang, et al.
COLM, 2025
ORLM: A Customizable Framework in Training Large Models for Automated Optimization Modeling
Chenyu Huang*, Zhengyang Tang*, Shixi Hu, Ruoqing Jiang, Xin Zheng, Dongdong Ge, Benyou Wang, Zizhuo Wang
Operations Research, 2025
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
Zhengyang Tang, Xingxing Zhang, Benyou Wang, Furu Wei
ICML, 2024
Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models (GLAN)
Haoran Li*, Qingxiu Dong*, Zhengyang Tang*, Chaojun Wang*, Xingxing Zhang, et al.
TMLR, 2025
Second Language (Arabic) Acquisition of LLMs via Progressive Vocabulary Expansion
Jianqing Zhu*, Huang Huang*, Zhihang Lin*, Juhao Liang*, Zhengyang Tang*, Khalid Almubarak, et al.
ACL 2025 (Oral & Panel)
DPTDR: Deep Prompt Tuning for Dense Passage Retrieval
Zhengyang Tang, Benyou Wang, Ting Yao
COLING, 2022

Talks

Invited Talk on MathScale, ICML, Jul 2024. [Slides]
Invited Talk on Dense Passage Retrieval, Baidu Search Department, 2022.

Selected Patents

Service

Reviewer for ACL Rolling Review (since 2024), COLM, NeurIPS, ICLR, and ICML.

Awards / Honors

Runner-up, LIC 2022 Passage-Ranking Competition (2nd place / 793 teams), Baidu & CCF, 2022.
Top-3 Ranking, MS MARCO Passage Ranking Leaderboard, Microsoft Research, 2022.