Qizhen Weng 翁祈桢
Large Model System Researcher. Ph.D. in CSE from HKUST.
My research interests encompass AI Infrastructure, Machine Learning Systems, and Cloud Computing, with a particular emphasis on enhancing GPU cluster efficiency and optimizing training performance for large-scale generative models, such as large language models (LLMs), multimodal LLMs (MLLMs), and diffusion transformers (DiTs).
(1) From 2024 to May 2026, I led the AI Infrastructure Research Center at the Institute of Artificial Intelligence (TeleAI), China Telecom, where I oversaw initiatives to advance AI system capabilities. (2) Prior to this, I joined the Shanghai AI Laboratory in 2022 as a Systems Researcher, contributing to the systems for large language model training and inference. (3) Earlier, I gained valuable experience as a Research Intern at Alibaba Cloud & Alibaba Group, where I focused on GPU cluster management and AI job scheduling for over two years, beginning in 2020.
I received my Ph.D. in Computer Science and Engineering from The Hong Kong University of Science and Technology in 2022, under the guidance of Prof. Wei Wang. I also hold a B.Eng. degree from Shanghai Jiao Tong University in 2017 and enriched my academic journey with a study period at UC Berkeley in 2015.
Awards
- Young Elite Scientists Sponsorship Program, CAST, 2025: for AI development tools and infrastructure
- Hong Kong PhD Fellowship Scheme, RGC of HK, 2017: awarded to 231 top students worldwide
- Shanghai Outstanding Graduates, SH Gov., 2017: awarded to top 3% students in the college
- Cyber-Security Scholarship, CIDF, 2016: awarded to 1% students in the major
News & Highlights
| Aug 21, 2026 | 📜EMNLP 2026: Beyond Binary, which turns partial success into dense verifiable rewards for reinforcement learning in code generation, is accepted to EMNLP 2026! See our paper “Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation” for details. |
|---|---|
| Jul 02, 2026 | 📜SC 2026: Arachne, a training framework that orchestrates cascades for efficient text-to-video model training at scale, is accepted to SC 2026! See our paper “Arachne: Orchestrating Cascades for Efficient Text-to-Video Model Training” for details. |
| Apr 29, 2026 | 📜ICML 2026: Swift-SVD, a theoretically optimal and practically efficient method for low-rank compression of LLM weights, is accepted to ICML 2026! See our paper “Swift-SVD: Theoretical Optimality Meets Practical Efficiency in Low-Rank LLM Compression” for details. |
| Jan 31, 2026 | 📜EuroSys 2026: Suika, a cluster training system that supports efficient and high-quality rescheduling for 3D-parallelized LLM training jobs, is accepted to EuroSys 2026! See our paper “Suika: Efficient and High-quality Rescheduling of 3D-parallelized LLM Training Jobs in Shared Clusters” for details. |
| Aug 23, 2025 | 📜EuroSys 2026: GRouter, a GPU-centric data plane system designed for serverless inference workflows, is accepted to EuroSys 2026! See our paper “Efficient Data Passing for Serverless Inference Workflows: A GPU-Centric Approach” for details. |
| Jun 15, 2025 | ♻️AI for Good Global Summit: I was invited as a keynote speaker at the AI for Good Global Summit (8–11 July 2025, Geneva, ITU). I presented “China Telecom drives ubiquitous intelligence through AI Flow”, on bridging devices, edge, and cloud for ubiquitous intelligence. |