About
I am working on building efficient AI systems.
I received my Ph.D. from MMLab, CUHK, advised by Prof. Dahua Lin. My research interests lie in the broad area of machine learning systems, especially efficient LLM training and inference. During my Ph.D. studies, I also worked closely with Prof. Zhihao Jia at CMU and Prof. Hao Zhang at UC San Diego. Before joining CUHK, I received my Bachelor's degree in Computer Science from University of Chinese Academy of Sciences, advised by Prof. Shiguang Shan.
News
- [July 2026] I will attend OSDI '26.
older news
- [Aug. 2024] Proteus is finally accepted by TPDS!
- [Aug. 2024] Start my internship at AWS, Santa Clara!
- [July 2024] We announce a survey about LLM training systems and infrastructure; check arXiv!
- [July 2024] SKVQ is accepted by COLM 2024. Congratulations to Duanmu!
- [May 2024] MuxServe is accepted by ICML 2024!
- [Apr. 2024] I will attend NSDI '24 in person at Santa Clara, CA. See you there!
Education
The Chinese University of Hong Kong
Aug. 2021 - July 2025
Ph.D. in Information Engineering
Thesis: Efficient and Affordable Large Language Systems
University of Chinese Academy of Sciences
Sep. 2016 - July 2020
B.E. in Computer Science and Technology
Experience
-
Tongyi Lab July 2025 - May 2026
Alibaba Cloud PAI -> Tongyi Lab.
Built infra for Qwen, including agentic RL infra and attention-FFN disaggregation inference.
Internships & Visiting Research
-
ByteDance Seed Mar. 2025 - June 2025
Hosts: Jianyu Jiang, Yanghua Peng
Built multimodal pretraining infra.
-
Amazon Web Services (AWS AI) Aug. 2024 - Jan. 2025
Hosts: Liangfu Chen, Zhen Jia, Zhihao Jia, Yida Wang
Designed sequence sharding and padding-efficient attention for LLM inference on Amazon's in-house AI accelerator, improving efficiency under static-shape constraints and reducing end-to-end latency.
-
Hao AI Lab, UC San Diego Aug. 2023 - Aug. 2024
Advisor: Hao Zhang
Initiated and led MuxServe, the first system enabling efficient multi-LLM serving via spatial-temporal multiplexing, cutting resource cost while maintaining service quality.
-
Catalyst, Carnegie Mellon University Apr. 2022 - May 2023
Advisors: Zhihao Jia, Minjia Zhang, Xupeng Miao
Led Parcae for cost-efficient DNN training and built SpotServe for distributed LLM serving on preemptible instances.
-
MMLab, CUHK Aug. 2020 - Aug. 2021
Advisors: Dahua Lin, Shengen Yan, Xiuhong Li
Explored automated DNN parallelization and led Proteus, a performance model for optimizing parallelization strategies.
Publications
2026
-
TrainMover: An Interruption-Resilient Runtime for ML Training
In Proceedings of the 20th USENIX Symposium on Operating Systems Design and Implementation (OSDI), July 2026.
[Paper] -
CoCoQuant: Breaking the Bandwidth Wall via Co-Optimized Communication and Computation Quantization
In Proceedings of the International Conference on Machine Learning (ICML), 2026.
[Paper] -
NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding
arXiv preprint, May 2026.
[Paper] -
Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
In Proceedings of the 21st European Conference on Computer Systems (EuroSys), April 2026.
[Paper]
2025
-
MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design
In Proceedings of the International Conference on Machine Learning (ICML), July 2025.
[Paper], [Code] -
Tropical: Enhancing SLO Attainment in Disaggregated LLM Serving via SLO-Aware Multiplexing
In Proceedings of the 62nd ACM/IEEE Design Automation Conference (DAC), June 2025.
[Paper] -
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
In Proceedings of the Conference on Machine Learning and Systems (MLSys), May 2025.
[Paper]
2024
-
SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models
In Proceedings of the Conference on Language Modeling (COLM, Spotlight), October 2024.
[Paper] -
Proteus: Simulating the Performance of Distributed DNN Training
IEEE Transactions on Parallel and Distributed Systems (TPDS), August 2024.
[Paper], [Code] -
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
In Proceedings of the International Conference on Machine Learning (ICML), July 2024.
[Paper], [Code], [Blog], [Video (Chinese)] -
Centauri: Enabling Efficient Scheduling for Communication-Computation Overlap in Large Model Training via Communication Partitioning
In Proceedings of the ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS), April 2024.
Best Paper Award
[Paper], [Video (Chinese)] -
SpotServe: Serving Generative Large Language Models on Preemptible Instances
In Proceedings of the ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS), April 2024.
Distinguished Artifact Award
IEEE Micro Top Picks Honorable Mention
[Paper], [Code] -
Parcae: Proactive, Liveput-Optimized DNN Training on Preemptible Instances
In Proceedings of the Symposium on Networked Systems Design and Implementation (NSDI), April 2024.
[Paper], [Code]
Survey
Teaching
- IERG3050: Simulation and Statistical Analysis Teaching Assistant · Fall 2021 · CUHK
- CSCI2100: Data Structure Teaching Assistant · Spring 2022 · CUHK
Services
- Program Committee Member
- IJCAI 2025, AAAI 2026
- Reviewer
- TPDS 2024, ACM TIST 2024, IJCNN 2025, ICML 2025, ICLR 2026, ICML 2026
- Artifact Evaluation Committee Member
- MLSys 2023, OSDI 2024, ATC 2024, ASPLOS 2025
Awards
- IEEE Micro Top Picks Honorable Mention
- Best Paper Award, ASPLOS
- Distinguished Artifact Award, ASPLOS
- Outstanding Graduate of Beijing
- Outstanding Graduate of University of Chinese Academy of Sciences
- Tang Lixin Scholarship
- First-class Academic Scholarship, UCAS (top 5%)