Robust and efficient runtime optimization
We study KV-cache management, model compression, and SLO-aware scheduling for large language and diffusion models under dynamic and heterogeneous workloads.
ZJU 100 Young Professor
I study performance optimization for intelligent computing systems, with current interests in agentic services, generative AI runtime, and cloud-edge intelligence.
Armor, a missing-aware multimodal fusion system for unified microservice incident management, is accepted at ASE ’26.
Agentic Services Computing presents a service-computing perspective on systems centered on autonomous agents.
Towards Optimal Robustness in Learning-Augmented Paging receives an ICML ’26 Spotlight.
I work across online optimization, systems, and machine learning for the deployment, coordination, and operation of intelligent services.
We study KV-cache management, model compression, and SLO-aware scheduling for large language and diffusion models under dynamic and heterogeneous workloads.
We design algorithms and systems for heterogeneous cloud-edge platforms that jointly improve performance, reliability, and efficiency.
We study runtime, incident-management, and resource-control mechanisms for agent-centered service platforms.
PFSUM studies the Bahncard problem, a classical online rent-or-buy problem. Guard and RPB study robust learning-augmented caching and paging with imperfect predictions.
SegQuant studies semantic-aware quantization. Shiva-DiT develops differentiable selection for efficient diffusion transformers.
We study agent-centered service platforms and system support for microservice incident management. Recent work includes Agentic Services Computing and Armor.
One PhD position is available for the 2027 cohort. Please email a CV and a brief introduction.
Applicants working on AI systems, systems optimization, or related areas are welcome.
Contact meResearch interests include agentic services, generative AI systems, cloud-edge intelligence, and optimization.
Ask about supervision“†” denotes equal contribution; “*” denotes corresponding author. For the complete and most up-to-date record, see Google Scholar ↗.
ARMOR: Missing-Aware Multimodal Fusion for Unified Microservice Incident Management
Wenzhuo Qian, Hailiang Zhao*, Ziqi Wang, Zhipeng Gao, Jiayi Chen, Zhiwei Ling, and Shuiguang Deng*.
Towards Optimal Robustness in Learning-Augmented Paging
Peng Chen, Hailiang Zhao*, Xueyan Tang, Yixuan Wang, and Shuiguang Deng*.
SegQuant: A Semantics-Aware and Generalizable Quantization Framework for Diffusion Models
Jiaji Zhang, Ruichao Sun, Hailiang Zhao*, Jiaju Wu, Peng Chen, Hao Li, Yuying Liu, Kingsum Chow, Gang Xiong, and Shuiguang Deng*. Project ↗
PeerSync: Accelerating Containerized Model Inference at the Network Edge
Yinuo Deng†, Hailiang Zhao†*, Dongjing Wang, Peng Chen, Wenzhuo Qian, Jianwei Yin, Schahram Dustdar, and Shuiguang Deng*.
Robustifying Learning-Augmented Caching Efficiently without Compromising 1-Consistency
Peng Chen, Hailiang Zhao*, Jiaji Zhang, Xueyan Tang, Yixuan Wang, and Shuiguang Deng*. Poster ↗
CADRef: Robust Out-of-Distribution Detection via Class-Aware Decoupled Relative Feature Leveraging
Zhiwei Ling, Yachen Chang, Hailiang Zhao*, Xinkui Zhao, Kingsum Chow*, and Shuiguang Deng. Code ↗
Learning-Augmented Algorithms for the Bahncard Problem
Hailiang Zhao, Xueyan Tang*, Peng Chen, and Shuiguang Deng. Code ↗ · Slides ↗
Online Workload Scheduling for Social Welfare Maximization in the Computing Continuum
Hailiang Zhao, Ziqi Wang, Guanjie Cheng*, Wenzhuo Qian, Peng Chen, Jianwei Yin, Schahram Dustdar, and Albert Y. Zomaya.
CATScaler: A Convolution-Augmented Transformer Scaling Framework for Cloud-Native Applications
Fan'an Meng, Hongjun Dai*, Guoqing Cong, Bo Zhu, and Hailiang Zhao.
Tail-Learning: Adaptive Learning Method for Mitigating Tail Latency in Autonomous Edge Systems
Cheng Zhang, Yinuo Deng, Hailiang Zhao*, Tianlv Chen, and Shuiguang Deng*.
Cloud-Native Computing: A Survey from the Perspective of Services
Shuiguang Deng*, Hailiang Zhao*, Bingbing Huang, Cheng Zhang, Feiyi Chen, Yinuo Deng, Jianwei Yin, Schahram Dustdar, and Albert Y. Zomaya. Cover Paper
Scheduling Multi-Server Jobs with Sublinear Regrets via Online Learning
Hailiang Zhao, Xueyan Tang*, Peng Chen, Jianwei Yin, and Shuiguang Deng.
Learning to Schedule Multi-Server Jobs with Fluctuated Processing Speeds
Hailiang Zhao, Shuiguang Deng*, Zhengzhe Xiang, Xueqiang Yan, Jianwei Yin, Schahram Dustdar, and Albert Y. Zomaya.
Dependent Function Embedding for Distributed Serverless Edge Computing
Shuiguang Deng, Hailiang Zhao, Zhengzhe Xiang, Cheng Zhang, Rong Jiang, Ying Li*, Jianwei Yin, Schahram Dustdar, and Albert Y. Zomaya.
DPoS: Decentralized, Privacy-Preserving, and Low-Complexity Online Slicing for Multi-Tenant Networks
Hailiang Zhao, Shuiguang Deng*, Zijie Liu, Zhengzhe Xiang, Jianwei Yin, Schahram Dustdar, and Albert Y. Zomaya. Trending Paper
Distributed Redundant Placement for Microservice-based Applications at the Edge
Hailiang Zhao, Shuiguang Deng*, Zijie Liu, Zhengzhe Xiang, Jianwei Yin, Schahram Dustdar, and Albert Y. Zomaya. ESI Highly Cited Paper
Edge Intelligence: The Confluence of Edge Computing and Artificial Intelligence
Shuiguang Deng, Hailiang Zhao, Weijia Fang*, Jianwei Yin, Schahram Dustdar, and Albert Y. Zomaya. ESI Hot & Highly Cited Paper
A Mobility-aware Cross-edge Computation Offloading Framework for Partitionable Applications
Hailiang Zhao, Shuiguang Deng*, Cheng Zhang, Wei Du, Qiang He, and Jianwei Yin. Best Student Paper
Shuiguang Deng*, Hailiang Zhao*, Ziqi Wang, Guanjie Cheng, Peng Chen, Wenzhuo Qian, Zhiwei Ling, Jianwei Yin, Albert Y. Zomaya, and Schahram Dustdar. arXiv preprint.
Shiva-DiT: Residual-Based Differentiable Top-k Selection for Efficient Diffusion Transformers
Jiaji Zhang, Hailiang Zhao*, Guoxuan Zhu, Ruichao Sun, Jiaju Wu, Xinkui Zhao, Hanlin Tang, Weiyi Lu, Kan Liu, Tao Lan, Lin Qu, and Shuiguang Deng*. arXiv preprint.
Adaptive Dual-Weighting Framework for Federated Learning via Out-of-Distribution Detection
Zhiwei Ling, Hailiang Zhao*, Chao Zhang, Xiang Ao, Ziqi Wang, Cheng Zhang, Zhen Qin, Xinkui Zhao, Kingsum Chow*, Yuanqing Wu, and MengChu Zhou. arXiv preprint.
Morphis: SLO-Aware Resource Scheduling for Microservices with Time-Varying Call Graphs
Yu Tang†, Hailiang Zhao†*, Chuansheng Lu, Yifei Zhang, Kingsum Chow*, Shuiguang Deng, and Rui Shi. arXiv preprint.
Hailiang Zhao, Ziqi Wang, Daojiang Hu, Zhiwei Ling, Wenzhuo Qian, Jiahui Zhai, Yuhao Yang, Zhipeng Gao, Mingyi Liu, Kai Di, Xinkui Zhao, Zhongjie Wang, Jianwei Yin, MengChu Zhou, and Shuiguang Deng. arXiv preprint.
TrackTeller: Temporal Multimodal 3D Grounding for Behavior-Dependent Object References
Jiahong Yu†, Ziqi Wang†, Hailiang Zhao*, Wei Zhai, Xueqiang Yan, and Shuiguang Deng. arXiv preprint.
OmniFuser: Adaptive Multimodal Fusion for Service-Oriented Predictive Maintenance
Ziqi Wang, Hailiang Zhao*, Yuhao Yang, Daojiang Hu, Cheng Bao, Mingyi Liu, Kai Di, Schahram Dustdar, Zhongjie Wang, and Shuiguang Deng. arXiv preprint.
Toward Robust and Efficient ML-Based GPU Caching for Modern Inference
Peng Chen, Jiaji Zhang, Hailiang Zhao*, Yirong Zhang, Jiahong Yu, Xueyan Tang, Yixuan Wang, Hao Li, Jianping Zou, Gang Xiong, Kingsum Chow, Shuibing He, and Shuiguang Deng*. arXiv preprint.
Ziqi Wang, Hailiang Zhao*, Cheng Bao, Wenzhuo Qian, Yuhao Yang, Xueqiang Sun, and Shuiguang Deng*. arXiv preprint.
Learning Unified System Representations for Microservice Tail Latency Prediction
Wenzhuo Qian, Hailiang Zhao*, Tianlv Chen, Jiayi Chen, Ziqi Wang, Kingsum Chow, and Shuiguang Deng*. arXiv preprint.
Young Editorial Board Member, Computer Engineering & Science; PC member and reviewer for NeurIPS, IJCAI-ECAI, AAAI, PAKDD, and related venues; reviewer for TSC, TMC, FGCS, TKDD, IoTJ, and others.
Selected lecture notes, explanatory slides, and systems analyses.