News & Highlights

Aug 21, 2026 📜EMNLP 2026: Beyond Binary, which turns partial success into dense verifiable rewards for reinforcement learning in code generation, is accepted to EMNLP 2026! See our paper “Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation” for details.
Jul 02, 2026 📜SC 2026: Arachne, a training framework that orchestrates cascades for efficient text-to-video model training at scale, is accepted to SC 2026! See our paper “Arachne: Orchestrating Cascades for Efficient Text-to-Video Model Training” for details.
Apr 29, 2026 📜ICML 2026: Swift-SVD, a theoretically optimal and practically efficient method for low-rank compression of LLM weights, is accepted to ICML 2026! See our paper “Swift-SVD: Theoretical Optimality Meets Practical Efficiency in Low-Rank LLM Compression” for details.
Jan 31, 2026 📜EuroSys 2026: Suika, a cluster training system that supports efficient and high-quality rescheduling for 3D-parallelized LLM training jobs, is accepted to EuroSys 2026! See our paper “Suika: Efficient and High-quality Rescheduling of 3D-parallelized LLM Training Jobs in Shared Clusters” for details.
Aug 23, 2025 📜EuroSys 2026: GRouter, a GPU-centric data plane system designed for serverless inference workflows, is accepted to EuroSys 2026! See our paper “Efficient Data Passing for Serverless Inference Workflows: A GPU-Centric Approach” for details.
Jun 15, 2025 ♻️AI for Good Global Summit: I was invited as a keynote speaker at the AI for Good Global Summit (8–11 July 2025, Geneva, ITU). I presented “China Telecom drives ubiquitous intelligence through AI Flow”, on bridging devices, edge, and cloud for ubiquitous intelligence.
Apr 01, 2025 📜USENIX ATC 2025: Toppings, an efficient multi-tenant system that serves many LoRA adapters with a common base LLM, is accepted to USENIX ATC 2025! See our paper “Toppings: CPU-Assisted, Rank-Aware Adapter Serving for LLM Inference” for details.
Apr 01, 2025 📜USENIX NSDI 2025: Prism, a production DLRM serving system that eliminates GPU fragmentation by means of resource disaggregation, is accepted to USENIX NSDI 2025! See our paper “GPU-Disaggregated Serving for Deep Learning Recommendation Models at Scale” for details.
Sep 01, 2024 💡Openings: Calling highly motivated students interning in Shanghai for 3+ months. Passionate about large AI models (e.g., LLM, VLM, DiT)? Please send your CV to my email. Experience with DL frameworks, distributed systems, or CUDA programming is a plus but not required.