Search DistillSys

Find a concept

Type at least two characters to search lessons, designs, papers, and interview prep.

Research area · 2024–2026

Cloud Computing

Cloud-native platforms, serverless systems, scheduling, resource management, and edge computing.

31 curated papers

31 papers

2026

2 papers
arXiv

Serverless Platform-Driven CPU Load Balancing

Abdul Rehman

Lets a serverless control plane influence Linux CPU scheduling through SchedExt domains, reducing loaded-request latency and energy use.

serverlessschedulingenergy
arXiv

Hierarchical Server Architecture for Agentic Science

Vanessa Sochat · Daniel Milroy

Designs a hierarchical discovery and negotiation architecture for dispatching agentic scientific workloads across cloud, edge, and HPC resources.

resource discoveryagentsHPC

2025

9 papers
arXiv

AI-Driven Cloud Resource Optimization for Multi-Cluster Environments

Vinoth Punniyamoorthy · Akash Kumar Agarwal · Bikesh Kumar · Abhirup Mazumder · Kabilan Kannan · Sumit Saha

Coordinates predictive, policy-aware resource decisions across clusters to improve utilization, adaptation speed, and workload stability.

multi-clusterresource managementprediction
arXiv

AdaptiFlow: An Extensible Framework for Event-Driven Autonomy in Cloud Microservices

Brice Arléon Zemtsop Ndadji · Simon Bliudze · Clément Quinton

Provides event-driven abstractions for decentralized runtime adaptation without imposing a centralized cloud controller.

cloud-nativemicroservicesadaptation
IEEE Micro
Industry researchHuawei Technologies (China)11 citations

UB-Mesh: A Hierarchically Localized nD-FullMesh Data Center Network Architecture

Heng Liao · Bingyang Liu · Xianping Chen · Zhigang Guo · Chuanning Cheng · et al.

The scaling of Large-scale Language Models (LLMs) demands unprecedented computational power and bandwidth. We present UB-Mesh, an innovative AI datacenter network architecture that enhances scalability, performance, and cost-efficiency through a hierarchical nD-FullMesh topology.

Cloud ComputingLLM
Open-access preprint
Industry researchHuawei Technologies (China)23 citations

BurstGPT: A Real-World Workload Dataset to Optimize LLM Serving Systems

Yuxin Wang · Yuhan Chen · Zeyu Li · Xueze Kang · Yuchu Fang · et al.

Despite efforts to improve the quality of service (QoS) and throughput in Large Language Model (LLM) serving systems, progress is often limited by the lack of publicly available real-world workloads. Consequently, evaluations usually depend on synthetic or oversimplified load patterns, and systems that appear promising in testing frequently underperform…

Cloud ComputingLLM
IEEE Transactions on Cloud Computing
Industry researchIBM (Canada)8 citations

A Reference Architecture for Governance of Cloud Native Applications

William Pourmajidi · Lei Zhang · John Steinbacher · Tony Erwin · Andriy Miranskyy

The evolution of cloud computing has given rise to Cloud Native Applications (CNAs), presenting new challenges in governance, particularly when faced with strict compliance requirements. This work explores the unique characteristics of CNAs and their impact on governance.

Cloud Computingcloud
Open-access preprint
Industry researchMicrosoft (United States)23 citations

TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms

Jovan Stojkovic · Chaojie Zhang · Íñigo Goiri · Esha Choukse · Haoran Qiu · et al.

The rising demand for generative large language models (LLMs) poses challenges for thermal and power management in cloud datacenters. Traditional techniques are often inadequate for LLM inference due to the fine-grained, millisecond-scale execution phases, each with distinct performance, thermal, and power profiles.

Cloud ComputingcloudschedulingLLM
IEEE Transactions on Computers
Industry researchTencent (China)11 citations

Enabling Consistent Sensing Data Sharing Among IoT Edge Servers via Lightweight Consensus

Xiulong Liu · Zhiyuan Zheng · Hao Xu · Zhelin Liang · Gaowei Shi · et al.

Blockchain offers distinct advantages in terms of data credibility and provenance certification, and its fusion with Internet of Things (IoT) technology holds great promise. Nevertheless, IoT environments are marked by extensive node networks and intricate communication patterns, especially the sensing environment.

Cloud Computingconsensusedge
IEEE Transactions on Services Computing
Industry researchEricsson (Sweden)17 citations

A Multi-Domain Survey on Time-Criticality in Cloud Computing

Remo Andreoli · Raquel A. F. Mini · Per Skarin · Harald Gustafsson · J. Harmatos · et al.

Conventional cloud services and infrastructures are mainly designed to maximize utilization of resources and provide best-effort Quality-of-Service levels. However, many emerging use cases in both public and private cloud computing scenarios are time-critical in nature.

Cloud Computingcloud
IEEE Transactions on Sustainable Computing
Industry researchAlibaba Group (China)10 citations

Adaptive Capacity Provisioning for Carbon-Aware Data Centers: A Digital Twin-Based Approach

Zhiwei Cao · Ruihang Wang · Xin Zhou · Rui Tan · Wen Yonggang · et al.

This paper considers the carbon-aware data center (DC) capacity provisioning problem under uncertain green energy availability and computing demand. To address it, accurate carbon emissions estimation and robust capacity provisioning are necessary.

Cloud Computing

2024

20 papers
arXiv

Adaptive Heuristics for Scheduling DNN Inferencing on Edge and Cloud for Personalized UAV Fleets

Suman Raj · Radhika Mittal · Harshil Gupta · Yogesh Simmhan

Schedules deadline-sensitive inference across drones, edge devices, and serverless cloud functions using dropping, stealing, migration, and adaptation.

edge-cloudschedulinginference
arXiv

Stimpack: An Adaptive Rendering Optimization System for Scalable Cloud Gaming

Jin Heo · Vic Wang · Ketan Bhardwaj · Ada Gavrilovska

Adapts server-side rendering quality to network compression so constrained edge servers spend compute where users can perceive it.

edge computingcloud gamingadaptation
IEEE Transactions on Computers
Industry researchAlibaba Group (United States)12 citations

Humas: A Heterogeneity- and Upgrade-Aware Microservice Auto-Scaling Framework in Large-Scale Data Centers

Hua Qin · Dingyu Yang · Shiyou Qian · Jian Cao · Guangtao Xue · et al.

An effective auto-scaling framework is essential for microservices to ensure performance stability and resource efficiency under dynamic workloads. As revealed by many prior studies, the key to efficient auto-scaling lies in accurately learning performance patterns, i.

Cloud Computing
IEEE Transactions on Services Computing
Industry researchHuawei Technologies (China)13 citations

Flexible Computing: A New Framework for Improving Resource Allocation and Scheduling in Elastic Computing

Weipeng Cao · Jiongjiong Gu · Zhong Ming · Zhiyuan Cai · Yuzhao Wang · et al.

Since the advent of cloud computing, Elastic Computing (EC) has become the standard architecture for resource allocation and scheduling. EC typically allocates computing resources based on predefined specifications, such as virtual machine or container flavors.

Cloud Computingscheduling
IEEE Transactions on Services Computing
Industry researchAlibaba Group (China) · Huawei Technologies (China)10 citations

No More Data Silos: Unified Microservice Failure Diagnosis With Temporal Knowledge Graph

Shenglin Zhang · Yongxin Zhao · Sibo Xia · Shirui Wei · Yongqian Sun · et al.

Microservices improve the scalability and flexibility of monolithic architectures to accommodate the evolution of software systems, but the complexity and dynamics of microservices challenge system reliability. Ensuring microservice quality requires efficient failure diagnosis, including detection and triage.

Cloud Computingedge
IEEE/ACM Transactions on Networking
Industry researchZTE (China)54 citations

Fluid-Shuttle: Efficient Cloud Data Transmission Based on Serverless Computing Compression

Rong Gu · Shulin Wang · Haipeng Dai · Xiaofei Chen · Zhaokang Wang · et al.

Nowadays, there exists a lot of cross-region data transmission demand on the cloud. It is promising to use serverless computing for data compressing to save the total data size.

Cloud Computingserverlesscloud
IEEE Transactions on Parallel and Distributed Systems
Industry researchHuawei Technologies (China)18 citations

ComboFunc: Joint Resource Combination and Container Placement for Serverless Function Scaling With Heterogeneous Container

Zhaojie Wen · Qiong Chen · Quanfeng Deng · Yipei Niu · Zhen Song · et al.

Serverless computing provides developers with a maintenance-free approach to resource usage, but it also transfers resource management responsibility to the cloud platform. However, the fine granularity of serverless function resources can lead to performance bottlenecks and resource fragmentation on nodes when creating many function containers.

Cloud Computingserverless
IEEE Transactions on Services Computing
Industry researchTencent (China)9 citations

Tetris: Proactive Container Scheduling for Long-Term Load Balancing in Shared Clusters

Fei Xu · Xiyue Shen · Shuo-Hao Lin · Li Chen · Zhi Zhou · et al.

Long-running containerized workloads (e. g.

Cloud Computingscheduling
Open-access preprint
Industry researchAlibaba Group (China) · Alibaba Group (United States)126 citations

Alibaba HPN: A Data Center Network for Large Language Model Training

Kun Qian · Yongqing Xi · Jiamin Cao · Jiaqi Gao · Yichi Xu · et al.

This paper presents HPN, Alibaba Cloud's data center network for large language model (LLM) training. Due to the differences between LLMs and general cloud computing (e.

Cloud Computing
Open-access preprint
Industry researchAlibaba Group (China)33 citations

Crux: GPU-Efficient Communication Scheduling for Deep Learning Training

Jiamin Cao · Yu Guan · Kun Qian · Jiaqi Gao · Wencong Xiao · et al.

Deep learning training (DLT), e. g.

Cloud Computingscheduling
Open-access preprint
Industry researchHuawei Technologies (China)25 citations

YuanRong: A Production General-purpose Serverless System for Distributed Applications in the Cloud

Qiong Chen · Jianmin Qian · Yulin Che · Ziqi Lin · Jianfeng Wang · et al.

We design, implement, and evaluate YuanRong, the first production general-purpose serverless platform with a unified programming interface, multi-language runtime, and a distributed computing kernel for cloud-based applications. YuanRong addresses many limitations of existing Function-as-a-Service (FaaS) systems, particularly in performance and lack of…

Cloud Computingserverlesscloud
IEEE Micro
Industry researchMicrosoft (United States)13 citations

Data Center Power and Energy Management: Past, Present, and Future

Ricardo Bianchini · Christian Belady · Anand Sivasubramaniam

This article overviews some of the key past developments in cloud data center power and energy management, where we are today, and what the future could be. This topic is gaining enormous renewed interest in the context of the conflicting needs of the AI revolution and the climate crisis.

Cloud Computing
Open-access preprint
Industry researchMicrosoft Research (United Kingdom)41 citations

Designing Cloud Servers for Lower Carbon

Jaylen Wang · Daniel S. Berger · Fiodar Kazhamiaka · Celine Irvene · Chaojie Zhang · et al.

To mitigate climate change, we must reduce carbon emissions from hyperscale cloud computing. We find that cloud compute servers cause the majority of emissions in a general-purpose cloud.

Cloud Computingcloud
IEEE Internet of Things Journal
Industry researchAlibaba Group (China)9 citations

Data-Driven Flexibility Capability Modeling of Internet Data Center Considering Task Dependency

Jiahao Ma · Ruiyang Yao · Bochao Zhang · Zhaoyang Wang · Yuejun Yan

The power consumption flexibility provided by the energy-intensive Internet data centers (IDCs) has been extensively studied as a potential solution for enhancing the flexibility of power systems. In IDCs, computational workloads are further divided into potentially interdependent tasks.

Cloud Computing
Open-access preprint
Industry researchMicrosoft (United States)69 citations

Characterizing Power Management Opportunities for LLMs in the Cloud

Pratyush Patel · Esha Choukse · Chaojie Zhang · Íñigo Goiri · Brijesh Warrier · et al.

Recent innovation in large language models (LLMs), and their myriad use cases have rapidly driven up the compute demand for datacenter GPUs. Several cloud providers and other enterprises plan to substantially grow their datacenter capacity to support these new workloads.

Cloud ComputingcloudLLM
IEEE Transactions on Cloud Computing
Industry researchHuawei Technologies (China)9 citations

FLAIR: A Fast and Low-Redundancy Failure Recovery Framework for Inter Data Center Network

Yuchao Zhang · Haoqiang Huang · Ahmed M. Abdelmoniem · Gaoxiong Zeng · Chenyue Zheng · et al.

Due to the fast developments of 5G and IoT technologies, Inter-Datacenter (Inter-DC) networks are facing unprecedented pressure to duplicate large volumes of geographically distributed user data in a real-time manner. Meanwhile, with the expansion of Inter-DC networks scale, link/node failures also become increasingly frequent, negatively affecting the…

Cloud Computing
IEEE Internet of Things Journal
Industry researchBaidu (China)73 citations

A2C-DRL: Dynamic Scheduling for Stochastic Edge–Cloud Environments Using A2C and Deep Reinforcement Learning

Jialin Lu · Jing Yang · Shaobo Li · Yijun Li · Jiang Wu · et al.

Resource management challenges frequently manifest in systems and networks as tough online decision tasks, for which the proper solution is dependent on an understanding of the workload and environment and facilitates smooth use of mobile edge and cloud resources. Due to the geographical dispersion of resources, constrained resource capacity,…

Cloud Computingcloudschedulingedge
IEEE Transactions on Parallel and Distributed Systems
Industry researchHuawei Technologies (China)10 citations

Joint Optimization of Parallelism and Resource Configuration for Serverless Function Steps

Zhaojie Wen · Qiong Chen · Yipei Niu · Zhen Song · Quanfeng Deng · et al.

Function-as-a-Service (FaaS) offers a fine-grained resource provision model, enabling developers to build highly elastic cloud applications. User requests are handled by a series of serverless functions step by step, which forms a multi-step workflow.

Cloud Computingserverless
ACM Transactions on Storage
Industry researchAlibaba Group (China)13 citations

An End-to-end High-performance Deduplication Scheme for Docker Registries and Docker Container Storage Systems

Nannan Zhao · Muhui Lin · Hadeel Albahar · Arnab K. Paul · Zhijie Huan · et al.

The wide adoption of Docker containers for supporting agile and elastic enterprise applications has led to a broad proliferation of container images. The associated storage performance and capacity requirements place a high pressure on the infrastructure of container registries that store and distribute images and container storage systems on the Docker…

Cloud Computingstorage
Journal of Artificial intelligence and Machine Learning
Industry researchOracle (United States)25 citations

Oracle OIPA Cloud Migration Analysis: Machine Learning Models for Predicting Resource Utilization and Success Outcomes

Tirumala Gundala

This study examines Oracle Insurance Policy Administration (OIPA) Coud Migration projects, analyzing 30 implementations that migrated from SQL Server to Oracle Cloud Infrastructure (OCI) environments. The research focuses on Universal Life Insurance systems migrating from AWS-hosted environments to Oracle’s cloud platform, including site upgrades from…

Cloud Computingcloud