Search DistillSys

Find a concept

Type at least two characters to search lessons, designs, papers, and interview prep.

Research area · 2024–2026

Distributed Systems

Consensus, replication, decentralization, fault tolerance, and distributed coordination.

31 curated papers

31 papers

2026

2 papers
arXiv

PLB: Priority-Aware Load Balancing for Replicated Databases under Constrained Resources

Belkis Djeffal · Pierre Bourhis · Romain Rouvoy

Introduces replica assignment that protects high-priority sessions while allowing capacity borrowing in constrained replicated database clusters.

replicationload balancingQoS
arXiv

Diagnosing High-Performance BFT Consensus via Mixture Modeling of Block Time Distributions

Hongru He · Akihiro Fujihara

Models quorum-formation latency and multimodal block-time distributions to diagnose heterogeneous network behavior in HotStuff-based BFT systems.

BFTconsensusobservability

2025

6 papers
arXiv

AdaptiFlow: An Extensible Framework for Event-Driven Autonomy in Cloud Microservices

Brice Arléon Zemtsop Ndadji · Simon Bliudze · Clément Quinton

Uses decentralized monitoring and declarative actions to add self-healing, self-protection, and self-optimization to microservices.

microservicesautonomic computingresilience
arXiv

Local Rendezvous Hashing: Bounded Loads and Minimal Churn via Cache-Local Candidates

Yongjie Guan

Combines a token ring with cache-local rendezvous candidates to improve balance and lookup throughput while preserving minimal churn.

consistent hashingload balancepartitioning
IEEE Micro
Industry researchHuawei Technologies (China)11 citations

UB-Mesh: A Hierarchically Localized nD-FullMesh Data Center Network Architecture

Heng Liao · Bingyang Liu · Xianping Chen · Zhigang Guo · Chuanning Cheng · et al.

The scaling of Large-scale Language Models (LLMs) demands unprecedented computational power and bandwidth. We present UB-Mesh, an innovative AI datacenter network architecture that enhances scalability, performance, and cost-efficiency through a hierarchical nD-FullMesh topology.

Distributed SystemsLLM
Open-access preprint
Industry researchAlibaba Group (United States) · Alibaba Group (China)11 citations

Unlocking the Potential of CXL for Disaggregated Memory in Cloud-Native Databases

Xinjun Yang · Yingqiang Zhang · Hao Chen · Feifei Li · Gerry Fan · et al.

Memory disaggregation has become a major trend in cloud-native databases. However, most existing memory disaggregation solutions suffer from read/write amplification, limited bandwidth, inefficient recovery, and challenges in data sharing.

Distributed SystemsclouddatabaseCXL
IEEE Internet of Things Journal
Industry researchHuawei Technologies (China)5 citations

BiTrust: Hybrid Trust Management for Secure Data Transmission in LEO Satellite Networks

Xinghai Wei · Jie Yuan · Runshan Hu · Haiguang Wang · Xiang Liu · et al.

Due to characteristics such as the openness and exposure of inter-satellite links, Low-Earth Orbit (LEO) satellite networks are subject to heightened vulnerability to malicious attacks compared to ground-based networks. Given various security risks, implementing trust management in LEO satellite networks becomes imperative.

Distributed Systems
ACM Transactions on Database Systems
Industry researchGoogle (United States)3 citations

Synchronizing Disaggregated Data Structures with One-Sided RDMA: Pitfalls, Experiments and Design Guidelines

Matthias Jasny · Tobias Ziegler · Jacob Nelson-Slivon · Viktor Leis · Carsten Binnig

Remote data structures built with one-sided Remote Direct Memory Access (RDMA) are at the heart of many disaggregated database management systems today. Concurrent access to these data structures by thousands of remote workers necessitates a highly efficient synchronization scheme.

Distributed SystemsRDMA

2024

23 papers
arXiv

OciorMVBA: Near-Optimal Error-Free Asynchronous MVBA

Jinyuan Chen

Presents error-free, information-theoretically secure asynchronous multi-valued Byzantine agreement with near-optimal communication.

Byzantine agreementasynchronyconsensus
arXiv

Blockchain-Empowered Cyber-Secure Federated Learning for Trustworthy Edge Computing

Ervin Moore · Ahmed Imteaj · Md Zarif Hossain · Shabnam Rezapour · M. Hadi Amini

Combines reputation, outlier detection, and blockchain-backed identity to defend cross-device federated learning against poisoned updates.

federated learningedgetrust
IEEE Transactions on Communications
Industry researchAlibaba Group (Cayman Islands)8 citations

Joint Deployment and Resource Allocation for Multi-AeBS Networks: A Two-Timescale Optimization Framework Using MADRL

Yikun Zhao · Fanqin Zhou · Lei Feng · Yao Sun · Wenjing Li · et al.

As an important component of the space-air-ground integrated network, aerial base station (AeBS) systems have gained significant attention for their flexibility in mobility and cost-effective construction. Nevertheless, the scarce spectrum resources and difficulty in accessing global information bring necessity and challenges to the deployment and…

Distributed Systems
Open-access preprint
Industry researchAlibaba Group (China)16 citations

TrEnv: Transparently Share Serverless Execution Environments Across Different Functions and Nodes

Jianda Huang · Mingxing Zhang · Teng Ma · Zheng Liu · Sixing Lin · et al.

Serverless computing is renowned for its computation elasticity, yet its full potential is often constrained by the requirement for functions to operate within local and dedicated background environments, resulting in limited memory elasticity. To address this limitation, this paper introduces TrEnv, a co-designed integration of the serverless platform…

Distributed Systemsserverless
ACM Transactions on Storage
Industry researchHuawei Technologies (China)4 citations

Efficiently Enlarging RDMA-Attached Memory with SSD

Zhe Yang · Qing Wang · Xiaojian Liao · Youyou Lu · Keji Huang · et al.

RDMA-based in-memory storage systems offer high performance but are restricted by the capacity of physical memory. In this article, we propose TeRM to extend RDMA-attached memory with SSD.

Distributed SystemsRDMASSD
IEEE/ACM Transactions on Networking
Industry researchHuawei Technologies (China)3 citations

Stable Byzantine Fault Tolerance in Wide Area Networks With Unreliable Links

Sitong Ling · Zhuotao Liu · Qi Li · Xinle Du · Jing Chen · et al.

With the increasing demand for blockchain technology in various industry sectors, there has been a growing interest in the Byzantine Fault Tolerance (BFT) consensus that is the backbone of most of these blockchains. However, many state-of-the-art algorithms that require reliable connections can only offer limited throughput in wide-area networks (WANs),…

Distributed Systemsfault tolerance
IEEE Transactions on Services Computing
Industry researchTencent (China)9 citations

Tetris: Proactive Container Scheduling for Long-Term Load Balancing in Shared Clusters

Fei Xu · Xiyue Shen · Shuo-Hao Lin · Li Chen · Zhi Zhou · et al.

Long-running containerized workloads (e. g.

Distributed Systemsscheduling
Proceedings of the VLDB Endowment
Industry researchTencent (China)12 citations

TDSQL: Tencent Distributed Database System

Yuxing Chen · Anqun Pan · Hailin Lei · Anda Ye · S. Han · et al.

Distributed databases have become indispensable in contemporary computing and data processing, owing to their pivotal role in ensuring high availability and scalability. They effectively cater to the requirements of data management and high-concurrency access.

Distributed Systemsdatabase
Open-access preprint
Industry researchAlibaba Group (United States) · Alibaba Group (China)11 citations

Relational Network Verification

Xieyang Xu · Yifei Yuan · Zachary Kincaid · Arvind Krishnamurthy · Ratul Mahajan · et al.

Relational network verification is a new approach for validating network changes. In contrast to traditional network verification, which analyzes specifications for a single network snapshot, it analyzes specifications that capture similarities and differences between two network snapshots (e.

Distributed Systems
Proceedings of the ACM on software engineering.
Industry researchHuawei Technologies (China)7 citations

TraStrainer: Adaptive Sampling for Distributed Traces with System Runtime State

Haiyu Huang · Xiaoyu Zhang · Pengfei Chen · Zilong He · Chen Zhi-ming · et al.

Distributed tracing has been widely adopted in many microservice systems and plays an important role in monitoring and analyzing the system. However, trace data often come in large volumes, incurring substantial computational and storage costs.

Distributed Systems
Proceedings of the VLDB Endowment
Industry researchZTE (United States)2 citations

Fast Commitment for Geo-Distributed Transactions via Decentralized Co-Coordinators

Zihao Zhang · Huiqi Hu · Xuan Zhou · Yaofeng Tu · Weining Qian · et al.

In a geo-distributed database, data shards and their respective replicas are deployed in distinct datacenters across multiple regions, enabling regional-level disaster recovery and the ability to serve global users locally. However, transaction processing in geo-distributed databases requires multiple cross-region communications, especially during the…

Distributed Systemstransactions
ACM Transactions on Architecture and Code Optimization
Industry researchHuawei Technologies (China)9 citations

Scythe: A Low-latency RDMA-enabled Distributed Transaction System for Disaggregated Memory

Kai Lü · Siqi Zhao · Haikang Shan · Qiang Wei · Guokuan Li · et al.

Disaggregated memory separates compute and memory resources into independent pools connected by RDMA (Remote Direct Memory Access) networks, which can improve memory utilization, reduce cost, and enable elastic scaling of compute and memory resources. However, existing RDMA-based distributed transactions on disaggregated memory suffer from severe…

Distributed SystemsRDMAtransactions
Open-access preprint
Industry researchBaidu (China)18 citations

FileDES: A Secure, Scalable and Succinct Decentralized Encrypted Storage Network

Minghui Xu · Jiahao Zhang · Hechuan Guo · Xiuzhen Cheng · Dongxiao Yu · et al.

Decentralized Storage Network (DSN) is an emerging technology that challenges traditional cloud-based storage systems by consolidating storage capacities from independent providers and coordinating to provide decentralized storage and retrieval services. However, current DSNs face several challenges associated with data privacy and efficiency of the…

Distributed Systemsstorage
IEEE Transactions on Services Computing
Industry researchBaidu (China)18 citations

MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services

Dianhai Yu · Liang Shen · Hongxiang Hao · Weibao Gong · Huachao Wu · et al.

While modern internet services, such as chatbots, search engines, and online advertising, demand the use of large-scale deep neural networks (DNNs), distributed training and inference over heterogeneous computing systems are desired to facilitate these DNN models. Mixture-of-Experts (MoE) is one the most common strategies to lower the cost of training…

Distributed Systems
Open-access preprint
Industry researchMicrosoft Research (United Kingdom)18 citations

Just-In-Time Checkpointing: Low Cost Error Recovery from Deep Learning Training Failures

Tanmaey Gupta · Sanjeev Krishnan · Rituraj Kumar · Abhishek Vijeev · Bhargav S. Gulavani · et al.

Deep Learning training jobs process large amounts of training data using many GPU devices, often running for weeks or months. When hardware or software failures happen, these jobs need to restart, losing the memory state for the Deep Neural Network (DNN) model trained so far, unless checkpointing mechanisms are used to save training state periodically.

Distributed Systemscheckpointing
IEEE Network
Industry researchHuawei Technologies (China)67 citations

NetGPT: An AI-Native Network Architecture for Provisioning Beyond Personalized Generative Services

Yuxuan Chen · Rongpeng Li · Zhifeng Zhao · Chenghui Peng · Jianjun Wu · et al.

Large language models (LLMs) have triggered tremendous success to empower our daily life by generative information. The personalization of LLMs could further contribute to their applications due to better alignment with human intents.

Distributed Systems
IEEE Transactions on Dependable and Secure Computing
Industry researchHuawei Technologies (United Kingdom)10 citations

Parallel Byzantine Consensus Based on Hierarchical Architecture and Trusted Hardware

Xiao Chen · Tiejun Ma · Btissam Er-Rahmadi · Jane Hillston · Guanxu Yuan

Byzantine fault-tolerant (BFT) state machine replication (SMR) is adopted to support blockchain consensus by tolerating arbitrarily faulty behaviours. However, the inherent complexity of BFT protocols makes existing BFT protocols hard to adapt to large-scale applications that require high scalability and performance.

Distributed Systemsconsensus
Proceedings of the ACM on Management of Data
Industry researchMicrosoft Research (United Kingdom)6 citations

Optimizing Distributed Protocols with Query Rewrites

David Chu · Rithvik Panchapakesan · Shadaj Laddad · Lucky E. Katahanas · Chris Liu · et al.

Distributed protocols such as 2PC and Paxos lie at the core of many systems in the cloud, but standard implementations do not scale. New scalable distributed protocols are developed through careful analysis and rewrites, but this process is ad hoc and error-prone.

Distributed Systemsquery optimization
IEEE Transactions on Computers
Industry researchHuawei Technologies (China)14 citations

BFT-DSN: A Byzantine Fault-Tolerant Decentralized Storage Network

Hechuan Guo · Minghui Xu · Jiahao Zhang · Chunchi Liu · Rajiv Ranjan · et al.

With the rapid development of blockchain and its applications, the amount of data stored on decentralized storage networks (DSNs) has grown exponentially. DSNs bring together affordable storage resources from around the world to provide robust, decentralized storage services for tens of thousands of decentralized applications (dApps).

Distributed Systemsfault tolerancestorage
IEEE Transactions on Parallel and Distributed Systems
Industry researchHuawei Technologies (China)10 citations

Joint Optimization of Parallelism and Resource Configuration for Serverless Function Steps

Zhaojie Wen · Qiong Chen · Yipei Niu · Zhen Song · Quanfeng Deng · et al.

Function-as-a-Service (FaaS) offers a fine-grained resource provision model, enabling developers to build highly elastic cloud applications. User requests are handled by a series of serverless functions step by step, which forms a multi-step workflow.

Distributed Systemsserverless
IEEE Journal on Selected Areas in Communications
Industry researchGoogle (United States)10 citations

Space System Internetworking: The Foundational Role of Delay and Disruption-Tolerant Networking

Jordan L. Torgerson · Vinton G. Cerf · Sky U. DeBaun · Larissa Suzuki

As humanity’s reach in space exploration extends beyond Earth’s orbit, the demand for robust and efficient communication systems intensifies. The traditional paradigm of isolated missions has given way to a complex network of diverse spacecraft engaged in terrestrial research, lunar exploration, and the provision of essential satellite services.

Distributed Systems
ACM Transactions on Architecture and Code Optimization
Industry researchHuawei Technologies (China)20 citations

Rcmp: Reconstructing RDMA-Based Memory Disaggregation via CXL

Zhonghua Wang · Y. M. Guo · Kai Lü · Jiguang Wan · Daohui Wang · et al.

Memory disaggregation is a promising architecture for modern datacenters that separates compute and memory resources into independent pools connected by ultra-fast networks, which can improve memory utilization, reduce cost, and enable elastic scaling of compute and memory resources. However, existing memory disaggregation solutions based on remote…

Distributed SystemsRDMACXL
IEEE Access
Industry researchGoogle (United States)24 citations

Logical Synchrony Networks: A Formal Model for Deterministic Distribution

Logan Kenwright · Partha S. Roop · Nathan Allen · Sanjay Lall · Călin Caşcaval · et al.

In the modelling of distributed systems, most Model of Computations (MoCs) rely on blocking communication to preserve determinism. A prominent example is Kahn Process Networks (KPNs), which supports non-blocking writes and blocking reads, and its implementable variant Finite FIFO Platforms (FFPs) which enforces boundedness using blocking writes.

Distributed Systems