Robust and efficient runtime optimization
We study KV-cache management, model compression, and SLO-aware scheduling for large language and diffusion models under dynamic and heterogeneous workloads.
ZJU 100 Young Professor
I study performance optimization for intelligent computing systems, with current interests in agentic services, generative AI runtime, and cloud-edge intelligence.
ARMOR, a robust self-supervised framework for microservice root-cause analysis under missing observability modalities, is accepted at ASE ’26.
Agentic Services Computing presents a service-computing perspective on systems centered on autonomous agents.
Towards Optimal Robustness in Learning-Augmented Paging receives an ICML ’26 Spotlight.
I work across online optimization, systems, and machine learning for the deployment, coordination, and operation of intelligent services.
We study KV-cache management, model compression, and SLO-aware scheduling for large language and diffusion models under dynamic and heterogeneous workloads.
We design algorithms and systems for heterogeneous cloud-edge platforms that jointly improve performance, reliability, and efficiency.
We study runtime, incident-management, and resource-control mechanisms for agent-centered service platforms.
PFSUM studies the Bahncard problem, a classical online rent-or-buy problem. Guard and RPB study robust learning-augmented caching and paging with imperfect predictions.
SegQuant studies semantic-aware quantization. Shiva-DiT develops differentiable selection for efficient diffusion transformers.
We study agent-centered service platforms and system support for microservice incident management. Recent work includes Agentic Services Computing and Armor.
One PhD position is available for the 2027 cohort. Please email a CV and a brief introduction.
Applicants working on AI systems, systems optimization, or related areas are welcome.
Contact meResearch interests include agentic services, generative AI systems, cloud-edge intelligence, and optimization.
Ask about supervision“†” denotes equal contribution; “*” denotes corresponding author. For the complete and most up-to-date record, see Google Scholar ↗.
Wenzhuo Qian, Hailiang Zhao*, Ziqi Wang, Zhipeng Gao, Jiayi Chen, Zhiwei Ling, and Shuiguang Deng*. Research note
Towards Optimal Robustness in Learning-Augmented Paging
Peng Chen, Hailiang Zhao*, Xueyan Tang, Yixuan Wang, and Shuiguang Deng*. Research note
SegQuant: A Semantics-Aware and Generalizable Quantization Framework for Diffusion Models
Jiaji Zhang, Ruichao Sun, Hailiang Zhao*, Jiaju Wu, Peng Chen, Hao Li, Yuying Liu, Kingsum Chow, Gang Xiong, and Shuiguang Deng*. Project ↗Research note
PeerSync: Accelerating Containerized Model Inference at the Network Edge
Yinuo Deng†, Hailiang Zhao†*, Dongjing Wang, Peng Chen, Wenzhuo Qian, Jianwei Yin, Schahram Dustdar, and Shuiguang Deng*.
Robustifying Learning-Augmented Caching Efficiently without Compromising 1-Consistency
Peng Chen, Hailiang Zhao*, Jiaji Zhang, Xueyan Tang, Yixuan Wang, and Shuiguang Deng*. Poster ↗
CADRef: Robust Out-of-Distribution Detection via Class-Aware Decoupled Relative Feature Leveraging
Zhiwei Ling, Yachen Chang, Hailiang Zhao*, Xinkui Zhao, Kingsum Chow*, and Shuiguang Deng. Code ↗
Learning-Augmented Algorithms for the Bahncard Problem
Hailiang Zhao, Xueyan Tang*, Peng Chen, and Shuiguang Deng. Code ↗ · Slides ↗
Online Workload Scheduling for Social Welfare Maximization in the Computing Continuum
Hailiang Zhao, Ziqi Wang, Guanjie Cheng*, Wenzhuo Qian, Peng Chen, Jianwei Yin, Schahram Dustdar, and Shuiguang Deng*.
CATScaler: A Convolution-Augmented Transformer Scaling Framework for Cloud-Native Applications
Fan'an Meng, Hongjun Dai*, Guoqing Cong, Bo Zhu, and Hailiang Zhao.
Tail-Learning: Adaptive Learning Method for Mitigating Tail Latency in Autonomous Edge Systems
Cheng Zhang, Yinuo Deng, Hailiang Zhao*, Tianlv Chen, and Shuiguang Deng*.
Cloud-Native Computing: A Survey From the Perspective of Services
Shuiguang Deng*, Hailiang Zhao*, Bingbing Huang, Cheng Zhang, Feiyi Chen, Yinuo Deng, Jianwei Yin, Schahram Dustdar, and Albert Y. Zomaya. Cover Paper
Scheduling Multi-Server Jobs with Sublinear Regrets via Online Learning
Hailiang Zhao, Shuiguang Deng*, Zhengzhe Xiang, Xueqiang Yan, Jianwei Yin, Schahram Dustdar, and Albert Y. Zomaya.
Learning to Schedule Multi-Server Jobs with Fluctuated Processing Speeds
Hailiang Zhao, Shuiguang Deng*, Feiyi Chen, Jianwei Yin, Schahram Dustdar, and Albert Y. Zomaya.
Dependent Function Embedding for Distributed Serverless Edge Computing
Shuiguang Deng, Hailiang Zhao, Zhengzhe Xiang, Cheng Zhang, Rong Jiang, Ying Li*, Jianwei Yin, Schahram Dustdar, and Albert Y. Zomaya.
DPoS: Decentralized, Privacy-Preserving, and Low-Complexity Online Slicing for Multi-Tenant Networks
Hailiang Zhao, Shuiguang Deng*, Zijie Liu, Zhengzhe Xiang, Jianwei Yin, Schahram Dustdar, and Albert Y. Zomaya. Trending Paper
Distributed Redundant Placement for Microservice-based Applications at the Edge
Hailiang Zhao, Shuiguang Deng*, Zijie Liu, Jianwei Yin, and Schahram Dustdar. ESI Highly Cited Paper
Edge Intelligence: The Confluence of Edge Computing and Artificial Intelligence
Shuiguang Deng, Hailiang Zhao, Weijia Fang*, Jianwei Yin, Schahram Dustdar, and Albert Y. Zomaya. ESI Hot & Highly Cited Paper
A Mobility-aware Cross-edge Computation Offloading Framework for Partitionable Applications
Hailiang Zhao, Shuiguang Deng*, Cheng Zhang, Wei Du, Qiang He, and Jianwei Yin. Best Student Paper
Prediction-Robust Service Deployment with Capacity-Aware Edge Admission
Hailiang Zhao, Ziqi Wang, Yifei Zhang, Mingyi Liu, Xinkui Zhao, Kingsum Chow, and Shuiguang Deng. arXiv preprint.
Shuiguang Deng*, Hailiang Zhao*, Ziqi Wang, Wenzhuo Qian, Xiang Ao, Guanjie Cheng, Jianwei Yin, Albert Y. Zomaya, and Schahram Dustdar. arXiv preprint.
Shiva-DiT: Residual-Based Differentiable Top-k Selection for Efficient Diffusion Transformers
Jiaji Zhang, Hailiang Zhao*, Jiaju Wu, Ruichao Sun, Xinkui Zhao, and Shuiguang Deng*. arXiv preprint.
Adaptive Dual-Weighting Framework for Federated Learning via Out-of-Distribution Detection
Zhiwei Ling, Hailiang Zhao*, Chao Zhang, Xiang Ao, Ziqi Wang, Cheng Zhang, Zhen Qin, Xinkui Zhao, Kingsum Chow*, Yuanqing Wu, and MengChu Zhou. arXiv preprint.
Morphis: SLO-Aware Resource Scheduling for Microservices with Time-Varying Call Graphs
Yu Tang†, Hailiang Zhao†*, Chuansheng Lu, Yifei Zhang, Kingsum Chow*, Shuiguang Deng, and Rui Shi. arXiv preprint.
Industrial Data-Service-Knowledge Governance: Toward Integrated and Trusted Intelligence
Hailiang Zhao, Ziqi Wang, Daojiang Hu, Mingyi Liu, Jiahui Zhai, Kai Di, Xinkui Zhao, Zhongjie Wang, Jianwei Yin, Albert Y. Zomaya, MengChu Zhou, and Shuiguang Deng. arXiv preprint.
TrackTeller: Temporal Multimodal 3D Grounding for Behavior-Dependent Object References
Jiahong Yu†, Ziqi Wang†, Hailiang Zhao*, Wei Zhai, Xueqiang Yan, and Shuiguang Deng. arXiv preprint.
OmniFuser: Adaptive Multimodal Fusion for Service-Oriented Predictive Maintenance
Ziqi Wang, Hailiang Zhao*, Yuhao Yang, Daojiang Hu, Cheng Bao, Mingyi Liu, Kai Di, Schahram Dustdar, Zhongjie Wang, and Shuiguang Deng. arXiv preprint.
Toward Robust and Efficient ML-Based GPU Caching for Modern Inference
Peng Chen, Jiaji Zhang, Hailiang Zhao*, Yirong Zhang, Shenyao Chen, Jiahong Yu, Xueyan Tang, Yixuan Wang, Hao Li, Jianping Zou, Gang Xiong, Kingsum Chow, Shuibing He, and Shuiguang Deng*. arXiv preprint.
Ziqi Wang, Hailiang Zhao*, Cheng Bao, Wenzhuo Qian, Yuhao Yang, Xueqiang Sun, and Shuiguang Deng*. arXiv preprint.
Wenzhuo Qian, Hailiang Zhao*, Jiayi Chen, Ziqi Wang, Tianlv Chen, Zhiwei Ling, Xinkui Zhao, Kingsum Chow, Albert Y. Zomaya, and Shuiguang Deng*. arXiv preprint.
Young Editorial Board Member, Computer Engineering & Science; PC member and reviewer for NeurIPS, IJCAI-ECAI, AAAI, PAKDD, and related venues; reviewer for TSC, TMC, FGCS, TKDD, IoTJ, and others.
Selected lecture notes, explanatory slides, and systems analyses.