Search DistillSys

Find a concept

Type at least two characters to search lessons, designs, papers, and interview prep.

Research index · 2024–2026

Track the ideas
shaping systems.

A curated starting point for recent work in distributed systems, cloud computing, databases, storage systems, and artificial intelligence—with a dedicated emphasis on research from industry labs. Every entry links to its source abstract and paper.

155papers125industry papers5research areas

155 papers

01

Distributed Systems

Consensus, replication, decentralization, fault tolerance, and distributed coordination.

2026

2 papers
arXiv

PLB: Priority-Aware Load Balancing for Replicated Databases under Constrained Resources

Belkis Djeffal · Pierre Bourhis · Romain Rouvoy

Introduces replica assignment that protects high-priority sessions while allowing capacity borrowing in constrained replicated database clusters.

replicationload balancingQoS
arXiv

Diagnosing High-Performance BFT Consensus via Mixture Modeling of Block Time Distributions

Hongru He · Akihiro Fujihara

Models quorum-formation latency and multimodal block-time distributions to diagnose heterogeneous network behavior in HotStuff-based BFT systems.

BFTconsensusobservability

2025

6 papers
arXiv

AdaptiFlow: An Extensible Framework for Event-Driven Autonomy in Cloud Microservices

Brice Arléon Zemtsop Ndadji · Simon Bliudze · Clément Quinton

Uses decentralized monitoring and declarative actions to add self-healing, self-protection, and self-optimization to microservices.

microservicesautonomic computingresilience
arXiv

Local Rendezvous Hashing: Bounded Loads and Minimal Churn via Cache-Local Candidates

Yongjie Guan

Combines a token ring with cache-local rendezvous candidates to improve balance and lookup throughput while preserving minimal churn.

consistent hashingload balancepartitioning
IEEE Micro
Industry researchHuawei Technologies (China)11 citations

UB-Mesh: A Hierarchically Localized nD-FullMesh Data Center Network Architecture

Heng Liao · Bingyang Liu · Xianping Chen · Zhigang Guo · Chuanning Cheng · et al.

The scaling of Large-scale Language Models (LLMs) demands unprecedented computational power and bandwidth. We present UB-Mesh, an innovative AI datacenter network architecture that enhances scalability, performance, and cost-efficiency through a hierarchical nD-FullMesh topology.

Distributed SystemsLLM
Open-access preprint
Industry researchAlibaba Group (United States) · Alibaba Group (China)11 citations

Unlocking the Potential of CXL for Disaggregated Memory in Cloud-Native Databases

Xinjun Yang · Yingqiang Zhang · Hao Chen · Feifei Li · Gerry Fan · et al.

Memory disaggregation has become a major trend in cloud-native databases. However, most existing memory disaggregation solutions suffer from read/write amplification, limited bandwidth, inefficient recovery, and challenges in data sharing.

Distributed SystemsclouddatabaseCXL
IEEE Internet of Things Journal
Industry researchHuawei Technologies (China)5 citations

BiTrust: Hybrid Trust Management for Secure Data Transmission in LEO Satellite Networks

Xinghai Wei · Jie Yuan · Runshan Hu · Haiguang Wang · Xiang Liu · et al.

Due to characteristics such as the openness and exposure of inter-satellite links, Low-Earth Orbit (LEO) satellite networks are subject to heightened vulnerability to malicious attacks compared to ground-based networks. Given various security risks, implementing trust management in LEO satellite networks becomes imperative.

Distributed Systems
ACM Transactions on Database Systems
Industry researchGoogle (United States)3 citations

Synchronizing Disaggregated Data Structures with One-Sided RDMA: Pitfalls, Experiments and Design Guidelines

Matthias Jasny · Tobias Ziegler · Jacob Nelson-Slivon · Viktor Leis · Carsten Binnig

Remote data structures built with one-sided Remote Direct Memory Access (RDMA) are at the heart of many disaggregated database management systems today. Concurrent access to these data structures by thousands of remote workers necessitates a highly efficient synchronization scheme.

Distributed SystemsRDMA

2024

23 papers
arXiv

OciorMVBA: Near-Optimal Error-Free Asynchronous MVBA

Jinyuan Chen

Presents error-free, information-theoretically secure asynchronous multi-valued Byzantine agreement with near-optimal communication.

Byzantine agreementasynchronyconsensus
arXiv

Blockchain-Empowered Cyber-Secure Federated Learning for Trustworthy Edge Computing

Ervin Moore · Ahmed Imteaj · Md Zarif Hossain · Shabnam Rezapour · M. Hadi Amini

Combines reputation, outlier detection, and blockchain-backed identity to defend cross-device federated learning against poisoned updates.

federated learningedgetrust
IEEE Transactions on Communications
Industry researchAlibaba Group (Cayman Islands)8 citations

Joint Deployment and Resource Allocation for Multi-AeBS Networks: A Two-Timescale Optimization Framework Using MADRL

Yikun Zhao · Fanqin Zhou · Lei Feng · Yao Sun · Wenjing Li · et al.

As an important component of the space-air-ground integrated network, aerial base station (AeBS) systems have gained significant attention for their flexibility in mobility and cost-effective construction. Nevertheless, the scarce spectrum resources and difficulty in accessing global information bring necessity and challenges to the deployment and…

Distributed Systems
Open-access preprint
Industry researchAlibaba Group (China)16 citations

TrEnv: Transparently Share Serverless Execution Environments Across Different Functions and Nodes

Jianda Huang · Mingxing Zhang · Teng Ma · Zheng Liu · Sixing Lin · et al.

Serverless computing is renowned for its computation elasticity, yet its full potential is often constrained by the requirement for functions to operate within local and dedicated background environments, resulting in limited memory elasticity. To address this limitation, this paper introduces TrEnv, a co-designed integration of the serverless platform…

Distributed Systemsserverless
ACM Transactions on Storage
Industry researchHuawei Technologies (China)4 citations

Efficiently Enlarging RDMA-Attached Memory with SSD

Zhe Yang · Qing Wang · Xiaojian Liao · Youyou Lu · Keji Huang · et al.

RDMA-based in-memory storage systems offer high performance but are restricted by the capacity of physical memory. In this article, we propose TeRM to extend RDMA-attached memory with SSD.

Distributed SystemsRDMASSD
IEEE/ACM Transactions on Networking
Industry researchHuawei Technologies (China)3 citations

Stable Byzantine Fault Tolerance in Wide Area Networks With Unreliable Links

Sitong Ling · Zhuotao Liu · Qi Li · Xinle Du · Jing Chen · et al.

With the increasing demand for blockchain technology in various industry sectors, there has been a growing interest in the Byzantine Fault Tolerance (BFT) consensus that is the backbone of most of these blockchains. However, many state-of-the-art algorithms that require reliable connections can only offer limited throughput in wide-area networks (WANs),…

Distributed Systemsfault tolerance
IEEE Transactions on Services Computing
Industry researchTencent (China)9 citations

Tetris: Proactive Container Scheduling for Long-Term Load Balancing in Shared Clusters

Fei Xu · Xiyue Shen · Shuo-Hao Lin · Li Chen · Zhi Zhou · et al.

Long-running containerized workloads (e. g.

Distributed Systemsscheduling
Proceedings of the VLDB Endowment
Industry researchTencent (China)12 citations

TDSQL: Tencent Distributed Database System

Yuxing Chen · Anqun Pan · Hailin Lei · Anda Ye · S. Han · et al.

Distributed databases have become indispensable in contemporary computing and data processing, owing to their pivotal role in ensuring high availability and scalability. They effectively cater to the requirements of data management and high-concurrency access.

Distributed Systemsdatabase
Open-access preprint
Industry researchAlibaba Group (United States) · Alibaba Group (China)11 citations

Relational Network Verification

Xieyang Xu · Yifei Yuan · Zachary Kincaid · Arvind Krishnamurthy · Ratul Mahajan · et al.

Relational network verification is a new approach for validating network changes. In contrast to traditional network verification, which analyzes specifications for a single network snapshot, it analyzes specifications that capture similarities and differences between two network snapshots (e.

Distributed Systems
Proceedings of the ACM on software engineering.
Industry researchHuawei Technologies (China)7 citations

TraStrainer: Adaptive Sampling for Distributed Traces with System Runtime State

Haiyu Huang · Xiaoyu Zhang · Pengfei Chen · Zilong He · Chen Zhi-ming · et al.

Distributed tracing has been widely adopted in many microservice systems and plays an important role in monitoring and analyzing the system. However, trace data often come in large volumes, incurring substantial computational and storage costs.

Distributed Systems
Proceedings of the VLDB Endowment
Industry researchZTE (United States)2 citations

Fast Commitment for Geo-Distributed Transactions via Decentralized Co-Coordinators

Zihao Zhang · Huiqi Hu · Xuan Zhou · Yaofeng Tu · Weining Qian · et al.

In a geo-distributed database, data shards and their respective replicas are deployed in distinct datacenters across multiple regions, enabling regional-level disaster recovery and the ability to serve global users locally. However, transaction processing in geo-distributed databases requires multiple cross-region communications, especially during the…

Distributed Systemstransactions
ACM Transactions on Architecture and Code Optimization
Industry researchHuawei Technologies (China)9 citations

Scythe: A Low-latency RDMA-enabled Distributed Transaction System for Disaggregated Memory

Kai Lü · Siqi Zhao · Haikang Shan · Qiang Wei · Guokuan Li · et al.

Disaggregated memory separates compute and memory resources into independent pools connected by RDMA (Remote Direct Memory Access) networks, which can improve memory utilization, reduce cost, and enable elastic scaling of compute and memory resources. However, existing RDMA-based distributed transactions on disaggregated memory suffer from severe…

Distributed SystemsRDMAtransactions
Open-access preprint
Industry researchBaidu (China)18 citations

FileDES: A Secure, Scalable and Succinct Decentralized Encrypted Storage Network

Minghui Xu · Jiahao Zhang · Hechuan Guo · Xiuzhen Cheng · Dongxiao Yu · et al.

Decentralized Storage Network (DSN) is an emerging technology that challenges traditional cloud-based storage systems by consolidating storage capacities from independent providers and coordinating to provide decentralized storage and retrieval services. However, current DSNs face several challenges associated with data privacy and efficiency of the…

Distributed Systemsstorage
IEEE Transactions on Services Computing
Industry researchBaidu (China)18 citations

MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services

Dianhai Yu · Liang Shen · Hongxiang Hao · Weibao Gong · Huachao Wu · et al.

While modern internet services, such as chatbots, search engines, and online advertising, demand the use of large-scale deep neural networks (DNNs), distributed training and inference over heterogeneous computing systems are desired to facilitate these DNN models. Mixture-of-Experts (MoE) is one the most common strategies to lower the cost of training…

Distributed Systems
Open-access preprint
Industry researchMicrosoft Research (United Kingdom)18 citations

Just-In-Time Checkpointing: Low Cost Error Recovery from Deep Learning Training Failures

Tanmaey Gupta · Sanjeev Krishnan · Rituraj Kumar · Abhishek Vijeev · Bhargav S. Gulavani · et al.

Deep Learning training jobs process large amounts of training data using many GPU devices, often running for weeks or months. When hardware or software failures happen, these jobs need to restart, losing the memory state for the Deep Neural Network (DNN) model trained so far, unless checkpointing mechanisms are used to save training state periodically.

Distributed Systemscheckpointing
IEEE Network
Industry researchHuawei Technologies (China)67 citations

NetGPT: An AI-Native Network Architecture for Provisioning Beyond Personalized Generative Services

Yuxuan Chen · Rongpeng Li · Zhifeng Zhao · Chenghui Peng · Jianjun Wu · et al.

Large language models (LLMs) have triggered tremendous success to empower our daily life by generative information. The personalization of LLMs could further contribute to their applications due to better alignment with human intents.

Distributed Systems
IEEE Transactions on Dependable and Secure Computing
Industry researchHuawei Technologies (United Kingdom)10 citations

Parallel Byzantine Consensus Based on Hierarchical Architecture and Trusted Hardware

Xiao Chen · Tiejun Ma · Btissam Er-Rahmadi · Jane Hillston · Guanxu Yuan

Byzantine fault-tolerant (BFT) state machine replication (SMR) is adopted to support blockchain consensus by tolerating arbitrarily faulty behaviours. However, the inherent complexity of BFT protocols makes existing BFT protocols hard to adapt to large-scale applications that require high scalability and performance.

Distributed Systemsconsensus
Proceedings of the ACM on Management of Data
Industry researchMicrosoft Research (United Kingdom)6 citations

Optimizing Distributed Protocols with Query Rewrites

David Chu · Rithvik Panchapakesan · Shadaj Laddad · Lucky E. Katahanas · Chris Liu · et al.

Distributed protocols such as 2PC and Paxos lie at the core of many systems in the cloud, but standard implementations do not scale. New scalable distributed protocols are developed through careful analysis and rewrites, but this process is ad hoc and error-prone.

Distributed Systemsquery optimization
IEEE Transactions on Computers
Industry researchHuawei Technologies (China)14 citations

BFT-DSN: A Byzantine Fault-Tolerant Decentralized Storage Network

Hechuan Guo · Minghui Xu · Jiahao Zhang · Chunchi Liu · Rajiv Ranjan · et al.

With the rapid development of blockchain and its applications, the amount of data stored on decentralized storage networks (DSNs) has grown exponentially. DSNs bring together affordable storage resources from around the world to provide robust, decentralized storage services for tens of thousands of decentralized applications (dApps).

Distributed Systemsfault tolerancestorage
IEEE Transactions on Parallel and Distributed Systems
Industry researchHuawei Technologies (China)10 citations

Joint Optimization of Parallelism and Resource Configuration for Serverless Function Steps

Zhaojie Wen · Qiong Chen · Yipei Niu · Zhen Song · Quanfeng Deng · et al.

Function-as-a-Service (FaaS) offers a fine-grained resource provision model, enabling developers to build highly elastic cloud applications. User requests are handled by a series of serverless functions step by step, which forms a multi-step workflow.

Distributed Systemsserverless
IEEE Journal on Selected Areas in Communications
Industry researchGoogle (United States)10 citations

Space System Internetworking: The Foundational Role of Delay and Disruption-Tolerant Networking

Jordan L. Torgerson · Vinton G. Cerf · Sky U. DeBaun · Larissa Suzuki

As humanity’s reach in space exploration extends beyond Earth’s orbit, the demand for robust and efficient communication systems intensifies. The traditional paradigm of isolated missions has given way to a complex network of diverse spacecraft engaged in terrestrial research, lunar exploration, and the provision of essential satellite services.

Distributed Systems
ACM Transactions on Architecture and Code Optimization
Industry researchHuawei Technologies (China)20 citations

Rcmp: Reconstructing RDMA-Based Memory Disaggregation via CXL

Zhonghua Wang · Y. M. Guo · Kai Lü · Jiguang Wan · Daohui Wang · et al.

Memory disaggregation is a promising architecture for modern datacenters that separates compute and memory resources into independent pools connected by ultra-fast networks, which can improve memory utilization, reduce cost, and enable elastic scaling of compute and memory resources. However, existing memory disaggregation solutions based on remote…

Distributed SystemsRDMACXL
IEEE Access
Industry researchGoogle (United States)24 citations

Logical Synchrony Networks: A Formal Model for Deterministic Distribution

Logan Kenwright · Partha S. Roop · Nathan Allen · Sanjay Lall · Călin Caşcaval · et al.

In the modelling of distributed systems, most Model of Computations (MoCs) rely on blocking communication to preserve determinism. A prominent example is Kahn Process Networks (KPNs), which supports non-blocking writes and blocking reads, and its implementable variant Finite FIFO Platforms (FFPs) which enforces boundedness using blocking writes.

Distributed Systems
02

Cloud Computing

Cloud-native platforms, serverless systems, scheduling, resource management, and edge computing.

2026

2 papers
arXiv

Serverless Platform-Driven CPU Load Balancing

Abdul Rehman

Lets a serverless control plane influence Linux CPU scheduling through SchedExt domains, reducing loaded-request latency and energy use.

serverlessschedulingenergy
arXiv

Hierarchical Server Architecture for Agentic Science

Vanessa Sochat · Daniel Milroy

Designs a hierarchical discovery and negotiation architecture for dispatching agentic scientific workloads across cloud, edge, and HPC resources.

resource discoveryagentsHPC

2025

9 papers
arXiv

AI-Driven Cloud Resource Optimization for Multi-Cluster Environments

Vinoth Punniyamoorthy · Akash Kumar Agarwal · Bikesh Kumar · Abhirup Mazumder · Kabilan Kannan · Sumit Saha

Coordinates predictive, policy-aware resource decisions across clusters to improve utilization, adaptation speed, and workload stability.

multi-clusterresource managementprediction
arXiv

AdaptiFlow: An Extensible Framework for Event-Driven Autonomy in Cloud Microservices

Brice Arléon Zemtsop Ndadji · Simon Bliudze · Clément Quinton

Provides event-driven abstractions for decentralized runtime adaptation without imposing a centralized cloud controller.

cloud-nativemicroservicesadaptation
IEEE Micro
Industry researchHuawei Technologies (China)11 citations

UB-Mesh: A Hierarchically Localized nD-FullMesh Data Center Network Architecture

Heng Liao · Bingyang Liu · Xianping Chen · Zhigang Guo · Chuanning Cheng · et al.

The scaling of Large-scale Language Models (LLMs) demands unprecedented computational power and bandwidth. We present UB-Mesh, an innovative AI datacenter network architecture that enhances scalability, performance, and cost-efficiency through a hierarchical nD-FullMesh topology.

Cloud ComputingLLM
Open-access preprint
Industry researchHuawei Technologies (China)23 citations

BurstGPT: A Real-World Workload Dataset to Optimize LLM Serving Systems

Yuxin Wang · Yuhan Chen · Zeyu Li · Xueze Kang · Yuchu Fang · et al.

Despite efforts to improve the quality of service (QoS) and throughput in Large Language Model (LLM) serving systems, progress is often limited by the lack of publicly available real-world workloads. Consequently, evaluations usually depend on synthetic or oversimplified load patterns, and systems that appear promising in testing frequently underperform…

Cloud ComputingLLM
IEEE Transactions on Cloud Computing
Industry researchIBM (Canada)8 citations

A Reference Architecture for Governance of Cloud Native Applications

William Pourmajidi · Lei Zhang · John Steinbacher · Tony Erwin · Andriy Miranskyy

The evolution of cloud computing has given rise to Cloud Native Applications (CNAs), presenting new challenges in governance, particularly when faced with strict compliance requirements. This work explores the unique characteristics of CNAs and their impact on governance.

Cloud Computingcloud
Open-access preprint
Industry researchMicrosoft (United States)23 citations

TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms

Jovan Stojkovic · Chaojie Zhang · Íñigo Goiri · Esha Choukse · Haoran Qiu · et al.

The rising demand for generative large language models (LLMs) poses challenges for thermal and power management in cloud datacenters. Traditional techniques are often inadequate for LLM inference due to the fine-grained, millisecond-scale execution phases, each with distinct performance, thermal, and power profiles.

Cloud ComputingcloudschedulingLLM
IEEE Transactions on Computers
Industry researchTencent (China)11 citations

Enabling Consistent Sensing Data Sharing Among IoT Edge Servers via Lightweight Consensus

Xiulong Liu · Zhiyuan Zheng · Hao Xu · Zhelin Liang · Gaowei Shi · et al.

Blockchain offers distinct advantages in terms of data credibility and provenance certification, and its fusion with Internet of Things (IoT) technology holds great promise. Nevertheless, IoT environments are marked by extensive node networks and intricate communication patterns, especially the sensing environment.

Cloud Computingconsensusedge
IEEE Transactions on Services Computing
Industry researchEricsson (Sweden)17 citations

A Multi-Domain Survey on Time-Criticality in Cloud Computing

Remo Andreoli · Raquel A. F. Mini · Per Skarin · Harald Gustafsson · J. Harmatos · et al.

Conventional cloud services and infrastructures are mainly designed to maximize utilization of resources and provide best-effort Quality-of-Service levels. However, many emerging use cases in both public and private cloud computing scenarios are time-critical in nature.

Cloud Computingcloud
IEEE Transactions on Sustainable Computing
Industry researchAlibaba Group (China)10 citations

Adaptive Capacity Provisioning for Carbon-Aware Data Centers: A Digital Twin-Based Approach

Zhiwei Cao · Ruihang Wang · Xin Zhou · Rui Tan · Wen Yonggang · et al.

This paper considers the carbon-aware data center (DC) capacity provisioning problem under uncertain green energy availability and computing demand. To address it, accurate carbon emissions estimation and robust capacity provisioning are necessary.

Cloud Computing

2024

20 papers
arXiv

Adaptive Heuristics for Scheduling DNN Inferencing on Edge and Cloud for Personalized UAV Fleets

Suman Raj · Radhika Mittal · Harshil Gupta · Yogesh Simmhan

Schedules deadline-sensitive inference across drones, edge devices, and serverless cloud functions using dropping, stealing, migration, and adaptation.

edge-cloudschedulinginference
arXiv

Stimpack: An Adaptive Rendering Optimization System for Scalable Cloud Gaming

Jin Heo · Vic Wang · Ketan Bhardwaj · Ada Gavrilovska

Adapts server-side rendering quality to network compression so constrained edge servers spend compute where users can perceive it.

edge computingcloud gamingadaptation
IEEE Transactions on Computers
Industry researchAlibaba Group (United States)12 citations

Humas: A Heterogeneity- and Upgrade-Aware Microservice Auto-Scaling Framework in Large-Scale Data Centers

Hua Qin · Dingyu Yang · Shiyou Qian · Jian Cao · Guangtao Xue · et al.

An effective auto-scaling framework is essential for microservices to ensure performance stability and resource efficiency under dynamic workloads. As revealed by many prior studies, the key to efficient auto-scaling lies in accurately learning performance patterns, i.

Cloud Computing
IEEE Transactions on Services Computing
Industry researchHuawei Technologies (China)13 citations

Flexible Computing: A New Framework for Improving Resource Allocation and Scheduling in Elastic Computing

Weipeng Cao · Jiongjiong Gu · Zhong Ming · Zhiyuan Cai · Yuzhao Wang · et al.

Since the advent of cloud computing, Elastic Computing (EC) has become the standard architecture for resource allocation and scheduling. EC typically allocates computing resources based on predefined specifications, such as virtual machine or container flavors.

Cloud Computingscheduling
IEEE Transactions on Services Computing
Industry researchAlibaba Group (China) · Huawei Technologies (China)10 citations

No More Data Silos: Unified Microservice Failure Diagnosis With Temporal Knowledge Graph

Shenglin Zhang · Yongxin Zhao · Sibo Xia · Shirui Wei · Yongqian Sun · et al.

Microservices improve the scalability and flexibility of monolithic architectures to accommodate the evolution of software systems, but the complexity and dynamics of microservices challenge system reliability. Ensuring microservice quality requires efficient failure diagnosis, including detection and triage.

Cloud Computingedge
IEEE/ACM Transactions on Networking
Industry researchZTE (China)54 citations

Fluid-Shuttle: Efficient Cloud Data Transmission Based on Serverless Computing Compression

Rong Gu · Shulin Wang · Haipeng Dai · Xiaofei Chen · Zhaokang Wang · et al.

Nowadays, there exists a lot of cross-region data transmission demand on the cloud. It is promising to use serverless computing for data compressing to save the total data size.

Cloud Computingserverlesscloud
IEEE Transactions on Parallel and Distributed Systems
Industry researchHuawei Technologies (China)18 citations

ComboFunc: Joint Resource Combination and Container Placement for Serverless Function Scaling With Heterogeneous Container

Zhaojie Wen · Qiong Chen · Quanfeng Deng · Yipei Niu · Zhen Song · et al.

Serverless computing provides developers with a maintenance-free approach to resource usage, but it also transfers resource management responsibility to the cloud platform. However, the fine granularity of serverless function resources can lead to performance bottlenecks and resource fragmentation on nodes when creating many function containers.

Cloud Computingserverless
IEEE Transactions on Services Computing
Industry researchTencent (China)9 citations

Tetris: Proactive Container Scheduling for Long-Term Load Balancing in Shared Clusters

Fei Xu · Xiyue Shen · Shuo-Hao Lin · Li Chen · Zhi Zhou · et al.

Long-running containerized workloads (e. g.

Cloud Computingscheduling
Open-access preprint
Industry researchAlibaba Group (China) · Alibaba Group (United States)126 citations

Alibaba HPN: A Data Center Network for Large Language Model Training

Kun Qian · Yongqing Xi · Jiamin Cao · Jiaqi Gao · Yichi Xu · et al.

This paper presents HPN, Alibaba Cloud's data center network for large language model (LLM) training. Due to the differences between LLMs and general cloud computing (e.

Cloud Computing
Open-access preprint
Industry researchAlibaba Group (China)33 citations

Crux: GPU-Efficient Communication Scheduling for Deep Learning Training

Jiamin Cao · Yu Guan · Kun Qian · Jiaqi Gao · Wencong Xiao · et al.

Deep learning training (DLT), e. g.

Cloud Computingscheduling
Open-access preprint
Industry researchHuawei Technologies (China)25 citations

YuanRong: A Production General-purpose Serverless System for Distributed Applications in the Cloud

Qiong Chen · Jianmin Qian · Yulin Che · Ziqi Lin · Jianfeng Wang · et al.

We design, implement, and evaluate YuanRong, the first production general-purpose serverless platform with a unified programming interface, multi-language runtime, and a distributed computing kernel for cloud-based applications. YuanRong addresses many limitations of existing Function-as-a-Service (FaaS) systems, particularly in performance and lack of…

Cloud Computingserverlesscloud
IEEE Micro
Industry researchMicrosoft (United States)13 citations

Data Center Power and Energy Management: Past, Present, and Future

Ricardo Bianchini · Christian Belady · Anand Sivasubramaniam

This article overviews some of the key past developments in cloud data center power and energy management, where we are today, and what the future could be. This topic is gaining enormous renewed interest in the context of the conflicting needs of the AI revolution and the climate crisis.

Cloud Computing
Open-access preprint
Industry researchMicrosoft Research (United Kingdom)41 citations

Designing Cloud Servers for Lower Carbon

Jaylen Wang · Daniel S. Berger · Fiodar Kazhamiaka · Celine Irvene · Chaojie Zhang · et al.

To mitigate climate change, we must reduce carbon emissions from hyperscale cloud computing. We find that cloud compute servers cause the majority of emissions in a general-purpose cloud.

Cloud Computingcloud
IEEE Internet of Things Journal
Industry researchAlibaba Group (China)9 citations

Data-Driven Flexibility Capability Modeling of Internet Data Center Considering Task Dependency

Jiahao Ma · Ruiyang Yao · Bochao Zhang · Zhaoyang Wang · Yuejun Yan

The power consumption flexibility provided by the energy-intensive Internet data centers (IDCs) has been extensively studied as a potential solution for enhancing the flexibility of power systems. In IDCs, computational workloads are further divided into potentially interdependent tasks.

Cloud Computing
Open-access preprint
Industry researchMicrosoft (United States)69 citations

Characterizing Power Management Opportunities for LLMs in the Cloud

Pratyush Patel · Esha Choukse · Chaojie Zhang · Íñigo Goiri · Brijesh Warrier · et al.

Recent innovation in large language models (LLMs), and their myriad use cases have rapidly driven up the compute demand for datacenter GPUs. Several cloud providers and other enterprises plan to substantially grow their datacenter capacity to support these new workloads.

Cloud ComputingcloudLLM
IEEE Transactions on Cloud Computing
Industry researchHuawei Technologies (China)9 citations

FLAIR: A Fast and Low-Redundancy Failure Recovery Framework for Inter Data Center Network

Yuchao Zhang · Haoqiang Huang · Ahmed M. Abdelmoniem · Gaoxiong Zeng · Chenyue Zheng · et al.

Due to the fast developments of 5G and IoT technologies, Inter-Datacenter (Inter-DC) networks are facing unprecedented pressure to duplicate large volumes of geographically distributed user data in a real-time manner. Meanwhile, with the expansion of Inter-DC networks scale, link/node failures also become increasingly frequent, negatively affecting the…

Cloud Computing
IEEE Internet of Things Journal
Industry researchBaidu (China)73 citations

A2C-DRL: Dynamic Scheduling for Stochastic Edge–Cloud Environments Using A2C and Deep Reinforcement Learning

Jialin Lu · Jing Yang · Shaobo Li · Yijun Li · Jiang Wu · et al.

Resource management challenges frequently manifest in systems and networks as tough online decision tasks, for which the proper solution is dependent on an understanding of the workload and environment and facilitates smooth use of mobile edge and cloud resources. Due to the geographical dispersion of resources, constrained resource capacity,…

Cloud Computingcloudschedulingedge
IEEE Transactions on Parallel and Distributed Systems
Industry researchHuawei Technologies (China)10 citations

Joint Optimization of Parallelism and Resource Configuration for Serverless Function Steps

Zhaojie Wen · Qiong Chen · Yipei Niu · Zhen Song · Quanfeng Deng · et al.

Function-as-a-Service (FaaS) offers a fine-grained resource provision model, enabling developers to build highly elastic cloud applications. User requests are handled by a series of serverless functions step by step, which forms a multi-step workflow.

Cloud Computingserverless
ACM Transactions on Storage
Industry researchAlibaba Group (China)13 citations

An End-to-end High-performance Deduplication Scheme for Docker Registries and Docker Container Storage Systems

Nannan Zhao · Muhui Lin · Hadeel Albahar · Arnab K. Paul · Zhijie Huan · et al.

The wide adoption of Docker containers for supporting agile and elastic enterprise applications has led to a broad proliferation of container images. The associated storage performance and capacity requirements place a high pressure on the infrastructure of container registries that store and distribute images and container storage systems on the Docker…

Cloud Computingstorage
Journal of Artificial intelligence and Machine Learning
Industry researchOracle (United States)25 citations

Oracle OIPA Cloud Migration Analysis: Machine Learning Models for Predicting Resource Utilization and Success Outcomes

Tirumala Gundala

This study examines Oracle Insurance Policy Administration (OIPA) Coud Migration projects, analyzing 30 implementations that migrated from SQL Server to Oracle Cloud Infrastructure (OCI) environments. The research focuses on Universal Life Insurance systems migrating from AWS-hosted environments to Oracle’s cloud platform, including site upgrades from…

Cloud Computingcloud
03

Databases

Query processing, indexing, transactions, data quality, and modern database architectures.

2026

3 papers
arXiv

PLB: Priority-Aware Load Balancing for Replicated Databases under Constrained Resources

Belkis Djeffal · Pierre Bourhis · Romain Rouvoy

Enforces priority differentiation through replica assignment while retaining high utilization through controlled capacity borrowing.

replicationOLAPload balancing
arXiv

Tytan: Interactive Neurosymbolic Construction of Analytic Semantic Schemas from Relational Data

Donna Hooshmand · Shubham Shahi · Cameron Barrie · Abhratanu Dutta · Marko Sterbentz · Harper Pack · Kristian J. Hammond

Combines symbolic database analysis, language-model inference, and targeted user questions to construct usable analytic semantic schemas.

semantic layerschemaneurosymbolic
Proceedings of the VLDB Endowment
Industry researchAlibaba Group (United States)3 citations

Why Database Manuals are Not Enough: Efficient and Reliable Configuration Tuning for DBMSs via Code-Driven LLM Agents

Xinyi Zhang · Tiantian Chen · Zhentao Han · Zhaoyan Hong · Wei Lu · et al.

Modern database management systems (DBMSs) expose hundreds of configuration knobs that critically influence performance. Existing automated tuning methods either adopt a data-driven paradigm, which incurs substantial overhead, or rely on manual-driven heuristics extracted from database documentation, which are often limited and overly generic.

DatabasesdatabaseLLMagents

2025

6 papers
arXiv

Database Theory in Action: Yannakakis' Algorithm

Paraschos Koutris · Stijn Vansummeren · Qichen Wang · Yisu Remy Wang · Xiangyao Yu

Reviews recent work that makes the optimal algorithm for acyclic joins practical in modern query engines.

joinsquery processingdatabase theory
arXiv

LMG Index: A Robust and Efficient Learned Index Framework for Multi-Dimensional Performance Balance

Yuzhen Chen · Bin Yao

Balances lookup, range-query, update, stability, and space objectives in one learned-index framework.

learned indexesquery performanceindexing
Proceedings of the VLDB Endowment
Industry researchTencent (China)4 citations

SiriusBI: A Comprehensive LLM-Powered Solution for Data Analytics in Business Intelligence

Jie Jiang · Haining Xie · Siqi Shen · Yu Shen · Zihan Zhang · et al.

With the proliferation of Large Language Models (LLMs) in Business Intelligence (BI), existing solutions face critical challenges in industrial deployments: functionality deficiencies from legacy systems failing to meet evolving LLM-era user demands, interaction limitations from single-round SQL generation paradigms inadequate for multi-round…

DatabasesanalyticsLLM
Proceedings of the VLDB Endowment
Industry researchMicrosoft Research (United Kingdom)8 citations

Scaling GPU-Accelerated Databases Beyond GPU Memory Size

Yinan Li · Bailu Ding · Ziyun Wei · Lukas M. Maas · Momin Al-Ghosien · et al.

There has been considerable interest in leveraging GPUs' computational power and high memory bandwidth for analytical database workloads. However, their limited memory capacity remains a fundamental limitation for databases whose sizes far exceed the GPU memory size.

Databasesdatabase
Proceedings of the ACM on Management of Data
Industry researchAlibaba Group (China)5 citations

Yannakakis+: Practical Acyclic Query Evaluation with Theoretical Guarantees

Qichen Wang · Bingnan Chen · Binyang Dai · Ke Yi · Feifei Li · et al.

Acyclic conjunctive queries form the backbone of most analytical workloads, and have been extensively studied in the literature from both theoretical and practical angles. However, there is still a large divide between theory and practice.

Databasesquery optimization
ACM SIGMOD Record
Industry researchMicrosoft (United States) · Alibaba Group (China)4 citations

A Roadmap to Graph Analytics

Angela Bonifati · M. TAMER ÖZSU · Yuanyuan Tian · Hannes Voigt · Wenyuan Yu · et al.

Graphs are ubiquitous data structures used in a large spectrum of applications, spanning from transportation networks, financial networks, social networks, product-order transactions and biomedical applications [33]. A recent survey on the usage of graph applications from real users has highlighted the fact that analytics is the most time-consuming task…

Databasesanalytics

2024

22 papers
Proceedings of the ACM on Management of Data
Industry researchAlibaba Group (China)4 citations

Towards a Converged Relational-Graph Optimization Framework

Yunkai Lou · Longbin Lai · Bingqing Lyu · Y Yang · Xiaoli Zhou · et al.

The recent ISO SQL:2023 standard adopts SQL/PGQ (Property Graph Queries), facilitating graph-like querying within relational databases. This advancement, however, underscores a significant gap in how to effectively optimize SQL/PGQ queries within relational database systems.

Databases
Proceedings of the ACM on Management of Data
Industry researchMicrosoft (United States)3 citations

Output-sensitive Conjunctive Query Evaluation

Shaleen Deep · Hangdong Zhao · Austen Z. Fan · Paraschos Koutris

Join evaluation is one of the most fundamental operations performed by database systems and arguably the most well-studied problem in the Database community. A staggering number of join algorithms have been developed, and commercial database engines use finely tuned join heuristics that take into account many factors including the selectivity of…

Databasesquery optimization
Proceedings of the VLDB Endowment
Industry researchAlibaba Group (United States)10 citations

PRICE: A Pretrained Model for Cross-Database Cardinality Estimation

Tianjing Zeng · Junwei Lan · Jiahong Ma · Wenqing Wei · Rong Zhu · et al.

Cardinality estimation (CardEst) is essential for optimizing query execution plans. Recent ML-based CardEst methods achieve high accuracy but face deployment challenges due to high preparation costs and lack of transferability across databases.

Databasesdatabase
Proceedings of the VLDB Endowment
Industry researchDatabricks (United States)12 citations

Adaptive and Robust Query Execution for Lakehouses at Scale

Maryann Xue · Yingyi Bu · Abhishek Somani · Wen Fan · Ziqi Liu · et al.

Many organizations have embraced the "Lakehouse" data management paradigm, which involves constructing structured data warehouses on top of open, unstructured data lakes. This approach stands in stark contrast to traditional, closed, relational databases and introduces challenges for performance and stability of distributed query processors.

Databasesquery optimization
Proceedings of the VLDB Endowment
Industry researchIntel (Germany) · Intel (India)21 citations

An Examination of CXL Memory Use Cases for In-Memory Database Management Systems Using SAP HANA

Minseon Ahn · Thomas Willhalm · Norman May · Donghun Lee · Suprasad Mutalik Desai · et al.

CXL-based disaggregated memory systems offer options to expand the memory beyond the limits of a single server via cache-coherent memory expansion cards or memory pools. Especially, In-Memory Database Management Systems (IMDBMSs) can benefit from alleviating two critical constraints: (1) limited memory capacity in a server and (2) long restart time…

DatabasesdatabaseCXL
Proceedings of the VLDB Endowment
Industry researchGoogle (United States)9 citations

SQL Has Problems. We Can Fix Them: Pipe Syntax In SQL

Jeff Shute · Shannon Bales · Matthew W. Brown · Jean-Daniel Browne · Brandon Dolphin · et al.

SQL has been extremely successful as the de facto standard language for working with data. Virtually all mainstream database-like systems use SQL as their primary query language.

Databases
Proceedings of the VLDB Endowment
Industry researchTencent (China)12 citations

TDSQL: Tencent Distributed Database System

Yuxing Chen · Anqun Pan · Hailin Lei · Anda Ye · S. Han · et al.

Distributed databases have become indispensable in contemporary computing and data processing, owing to their pivotal role in ensuring high availability and scalability. They effectively cater to the requirements of data management and high-concurrency access.

Databasesdatabase
Proceedings of the VLDB Endowment
Industry researchTencent (China)5 citations

X-Stor: A Cloud-Native NoSQL Database Service with Multi-Model Support

Hongyu Lei · Chunhua Li · Ke Zhou · Jianping Zhu · Kezhou Yan · et al.

In recent years at Tencent, we have observed that the use of multiple NoSQL databases for storing business data with diverse models has led to increased programming and deployment costs, as well as inefficient maintenance and underutilized resources. In this paper, we report X-Stor, a cloud-native NoSQL database system that supports multiple data models…

Databasesclouddatabase
Proceedings of the VLDB Endowment
Industry researchGoogle (United States)6 citations

Towards Optimal Transaction Scheduling

Audrey Cheng · Aaron Kabcenell · Jason Chan · Xiao Shi · Peter Bailis · et al.

Maximizing transaction throughput is key to high-performance database systems, which focus on minimizing data access conflicts to improve performance. However, finding efficient schedules that reduce conflicts remains an open problem.

Databasesschedulingtransactions
Proceedings of the VLDB Endowment
Industry researchAmazon (United States)35 citations

Why TPC is Not Enough: An Analysis of the Amazon Redshift Fleet

Alexander van Renen · Dominik Horn · Pascal Pfeil · Kapil Vaidya · Wenjian Dong · et al.

Database research and development is heavily influenced by benchmarks, such as the industry-standard TPC-H and TPC-DS for analytical systems. However, these twenty-year-old benchmarks neither capture how databases are deployed nor what workloads modern cloud data warehouse systems face these days.

Databases
The VLDB Journal
Industry researchAlibaba Group (China)14 citations

A survey on hybrid transactional and analytical processing

Haoze Song · Wenchao Zhou · Heming Cui · X. Peng · Feifei Li

Abstract To provide applications with the ability to analyze fresh data and eliminate the time-consuming ETL workflow, hybrid transactional and analytical (HTAP) systems have been developed to serve online transaction processing and online analytical processing workloads in a single system. In recent years, HTAP systems have attracted considerable…

Databasestransactions
Proceedings of the VLDB Endowment
Industry researchMicrosoft Research (United Kingdom)14 citations

DEX: Scalable Range Indexing on Disaggregated Memory

Baotong Lu · Kaisong Huang · Chieh-Jan Mike Liang · Tianzheng Wang · Eric Lo

Memory disaggregation can potentially allow memory-optimized range indexes such as B+-trees to scale beyond one machine while attaining high hardware utilization and low cost. Designing scalable indexes on disaggregated memory, however, is challenging due to rudimentary caching, unprincipled offloading and excessive inconsistency among servers.

Databasesindexing
Proceedings of the ACM on Management of Data
Industry researchAlibaba Group (China)4 citations

Relational Algorithms for Top-k Query Evaluation

Qichen Wang · Qiyao Luo · Yilei Wang

The evaluation of top-k conjunctive queries, a staple in business analysis, often requires evaluating the conjunctive query prior to filtering the top-k results, leading to a significant computational overhead within Database Management Systems (DBMSs). While efficient algorithms have been proposed, their integration into DBMSs remains arduous.

Databasesquery optimization
Proceedings of the ACM on Management of Data
Industry researchMicrosoft (United States)5 citations

Wii: Dynamic Budget Reallocation In Index Tuning

Xiaoying Wang · Wentao Wu · Chi Wang · Vivek Narasayya · Surajit Chaudhuri

Index tuning aims to find the optimal index configuration for an input workload. It is often a time-consuming and resource-intensive process, largely attributed to the huge amount of "what-if" calls made to the query optimizer during configuration enumeration.

Databasesindexing
Open-access preprint
Industry researchAmazon (United States) · Amazon (Germany)20 citations

Stage: Query Execution Time Prediction in Amazon Redshift

Ziniu Wu · Ryan Marcus · Zhengchun Liu · Parimarjan Negi · Vikram Nathan · et al.

Query performance (e. g.

Databasesquery optimization
ACM SIGMOD Record
Industry researchMicrosoft (United States)3 citations

Auto-Tables: Relationalize Tables without Using Examples

Peng Li · Yeye He · Cong Yan · Yue Wang · Surajit Chaudhuri

Relational tables, where each row corresponds to an entity and each column corresponds to an attribute, have been the standard for tables in relational databases. However, such a standard cannot be taken for granted when dealing with tables "in the wild".

Databases
arXiv

HTAP Databases: A Survey

Chao Zhang · Guoliang Li · Jintao Zhang · Xinning Zhang · Jianhua Feng

Classifies hybrid transactional/analytical systems by storage architecture and reviews synchronization, optimization, and scheduling techniques.

HTAPtransactionsanalytics
ICSE 2024 / arXiv

CERT: Finding Performance Issues in Database Systems Through the Lens of Cardinality Estimation

Jiale Chen · Yu Liang · Yuheng Shen · Yuting Chen · Jiachi Chen · Yinxing Xue

Uses cardinality-estimation consistency rules to find confirmed performance bugs in widely used database systems.

cardinality estimationtestingquery optimization
Proceedings of the ACM on Management of Data
Industry researchMicrosoft Research (United Kingdom)6 citations

Optimizing Distributed Protocols with Query Rewrites

David Chu · Rithvik Panchapakesan · Shadaj Laddad · Lucky E. Katahanas · Chris Liu · et al.

Distributed protocols such as 2PC and Paxos lie at the core of many systems in the cloud, but standard implementations do not scale. New scalable distributed protocols are developed through careful analysis and rewrites, but this process is ad hoc and error-prone.

Databasesquery optimization
Proceedings of the ACM on Management of Data
Industry researchMicrosoft (United States)9 citations

Sibyl: Forecasting Time-Evolving Query Workloads

Hanxian Huang · Tarique Siddiqui · Rana Alotaibi · Carlo Curino · Jyoti Leeka · et al.

Database systems often rely on historical query traces to perform workload-based performance tuning. However, real production workloads are time-evolving, making historical queries ineffective for optimizing future workloads.

Databasesquery optimization
Proceedings of the VLDB Endowment
Industry researchGoogle (United States)5 citations

POLAR: Adaptive and Non-invasive Join Order Selection via Plans of Least Resistance

David Justen · Daniel P. Ritter · Campbell Fraser · Andrew Lamb · Allison Lee · et al.

Join ordering and query optimization are crucial for query performance but remain challenging due to unknown or changing characteristics of query intermediates, especially for complex queries with many joins. Over the past two decades, a spectrum of techniques for adaptive query processing (AQP)---including inter-/intra-operator adaptivity and tuple…

Databases
Proceedings of the VLDB Endowment
Industry researchAlibaba Group (China)19 citations

Eraser: Eliminating Performance Regression on Learned Query Optimizer

Lianggui Weng · Rong Zhu · Di Wu · Bolin Ding · Bolong Zheng · et al.

Efficient query optimization is crucial for database management systems. Recently, machine learning models have been applied in query optimizers to generate better plans, but the unpredictable performance regressions prevent them from being truly applicable.

Databasesquery optimization
04

Storage Systems

Caching, checkpointing, tiered memory, I/O, and distributed data management.

2026

2 papers
arXiv

Operating Multi-Node Full Fine-Tuning on NVIDIA B300

Seon Ho Kim · Ui Jeong Jeon · Su Hyeon Kim · Min Tae Hwang

Reports checkpoint, NFS, page-cache, telemetry, and deadlock lessons from full fine-tuning a 32.76B model across two B300 nodes.

checkpointingdistributed trainingI/O
arXiv

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure

Yuhan Zhou · Yuchu Luo · Hao Nie · Wangrunze Lv · Yu Zhou · Yibo Zhu · Daxin Jiang · Chenren Xu

Introduces Tensor-as-a-Service and a programmable distributed layer for weights, KV caches, checkpoints, and tensor movement.

tensor storageKV cacheLLM infrastructure

2025

5 papers
arXiv

Vulcan: Instance-Specialized, Verifiable Systems Heuristics Through LLM-Driven Search

Rohit Dwivedula · Divyanshu Saxena · Sujay Yadalam · Eric Hayden Campbell · Daehyeok Kim · Aditya Akella

Synthesizes safe, deployment-specific heuristics for cache eviction, tiered memory, and spot-VM scheduling using restricted generated code.

cache evictiontiered memorysystems optimization
arXiv

Understanding LLM Checkpoint/Restore I/O Strategies and Patterns

Mikaila J. Gossman · Avinash Maurya · Bogdan Nicolae · Jon C. Calhoun

Measures aggregation, alignment, coalescing, buffered I/O, direct I/O, and io_uring for large-model checkpoint/restore workloads.

checkpointingio_uringparallel file systems
ACM Transactions on Architecture and Code Optimization
Industry researchHuawei Technologies (China)16 citations

ShuffleInfer: Disaggregate LLM Inference for Mixed Downstream Workloads

C.-C. Hu · Heyang Huang · Liangliang Xu · Xusheng Chen · Chenxi Wang · et al.

Transformer-based large language model (LLM) inference serving is now the backbone of many cloud services. LLM inference consists of a prefill phase and a decode phase.

Storage SystemsLLM
IEEE Transactions on Electron Devices
Industry researchSK Group (South Korea)9 citations

Investigation of Cell Variation Effect on Z-Interference in Charge-Trap-Based 3-D NAND Flash Memory

Sangmin Ahn · Hyungjun Jo · Sechun Park · Jongwoo Kim · Hyungcheol Shin

In this article, we investigated the effects of cell variations, specifically the variations in gate length (${L}_{\text {g}}$), spacer length (${L}_{\text {s}}$), filler oxide thickness (${T}_{\text {f}}$), channel thickness (${T}_{\text {ch}}$), tunneling oxide thickness (${T}_{\text {tox}}$), charge trap nitride thickness (${T}_{\text {ctn}}$), and…

Storage Systems
IEEE Transactions on Emerging Topics in Computing
Industry researchIntel (United States) · IBM (United States)9 citations

FHEmem: A Processing In-Memory Accelerator for Fully Homomorphic Encryption

Minxuan Zhou · Yujin Nam · Pranav Gangwar · Weihong Xu · Arpan Dutta · et al.

Fully Homomorphic Encryption (FHE) is a technique that allows arbitrary computations to be performed on encrypted data without the need for decryption, making it ideal for secure computation outsourcing. However, computation on FHE-encrypted data is significantly slower than that on plain data, primarily due to the explosive increases in data size and…

Storage Systems

2024

24 papers
ACM Transactions on Storage
Industry researchMicrosoft Research (United Kingdom) · Microsoft (United States)9 citations

Project Silica: Towards Sustainable Cloud Archival Storage in Glass

Patrick Anderson · Erika Aranas · Youssef Assaf · R. Behrendt · Richard Black · et al.

Sustainable and cost-effective long-term storage remains an unsolved problem. The most widely used storage technologies today are magnetic (hard disk drives and tape).

Storage Systemscloudstorage
ACM Transactions on Storage
Industry researchSamsung (South Korea) · Samsung (United States)10 citations

Storage Abstractions for SSDs: The Past, Present, and Future

Xiang-Qun Zhang · Janki Bhimani · Shuyi Pei · Eunji Lee · Sungjin Lee · et al.

This article traces the evolution of SSD (solid-state drive) interfaces, examining the transition from the block storage paradigm inherited from hard disk drives to SSD-specific standards customized to flash memory. Early SSDs conformed to the block abstraction for compatibility with the existing software storage stack, but studies and deployments show…

Storage SystemsstorageSSD
arXiv

Dynamic Optimization of Storage Systems Using Reinforcement Learning Techniques

Chiyu Cheng · Chang Zhou · Yang Zhao

Applies reinforcement learning to adapt cache size, queue depth, and readahead settings to changing I/O workloads.

reinforcement learningI/Oadaptive storage
arXiv

Optimizing SSD Caches for Cloud Block Storage Systems Using Machine Learning Approaches

Chiyu Cheng · Chang Zhou · Yang Zhao · Jin Cao

Predicts write-only data so cloud block-storage caches can avoid low-value SSD writes and improve utilization.

SSDcachingcloud block storage
ACM Transactions on Storage
Industry researchMicrosoft Research (United Kingdom) · Microsoft (Ireland) · Amazon (United Kingdom)9 citations

Holographic Storage for the Cloud: advances and challenges

Nathanaël Cheriere · Jiaqi Chu · Grace Brennan · Pashmina Cameron · Pedro F. da Costa · et al.

Holographic Storage is an old idea that has always promised high density and fast random access, but has never been commercially competitive with Hard Disk Drives (HDDs) and Solid State Devices (SSDs). In Project HSD at Microsoft Research we asked the question: “Does holographic storage finally make sense for cloud storage?

Storage Systemscloudstorage
Proceedings of the ACM on Management of Data
Industry researchAlibaba Group (China) · Alibaba Group (United States)22 citations

LSMGraph: A High-Performance Dynamic Graph Storage System with Multi-Level CSR

Song Yu · Shufeng Gong · Tao Qian · Sijie Shen · Yanfeng Zhang · et al.

The growing volume of graph data may exhaust the main memory. It is crucial to design a disk-based graph storage system to ingest updates and analyze graphs efficiently.

Storage Systemsstorage
ACM SIGEnergy Energy Informatics Review
Industry researchMicrosoft (United States)12 citations

A Call for Research on Storage Emissions

Sara McAllister · Fiodar Kazhamiaka · Daniel S. Berger · Rodrigo Fonseca · Kali Frost · et al.

Major cloud providers have committed to lowering carbon emissions by 2030 across their datacenters, and research has contributed many ideas on how this may be achieved. However, a major contributor to datacenter emissions has not received enough attention: storage.

Storage Systemsstorage
IEEE Electron Device Letters
Industry researchSK Group (South Korea)13 citations

Retention Improvement in Vertical NAND Flash Memory Using Electron Back-Tunneling

Sungho Park · Ho-Nam Yoo · Jong-Won Back · Joonhyung Cho · Yeongheon Yang · et al.

We propose an electron back-tunneling (EBT) method to enhance the retention characteristics of vertical NAND (V-NAND) flash memory. The storage of back-tunneled electrons in the spacer region between adjacent cells is facilitated by the synergistic effect of the fringing electric field between adjacent word-lines.

Storage Systems
IEEE Electron Device Letters
Industry researchSamsung (South Korea)13 citations

A Large Window Nonvolatile Transistor Memory for High-Density and Low-Power Vertical NAND Storage Enabled by Ferroelectric Charge Pumping

Zijian Zhao · Yixin Qin · Jiahui Duan · Yushan Lee · Suhwan Lim · et al.

In this work, we have developed a large memory window (MW) ferroelectric field effect transistor (FeFET) memory for vertical NAND storage. We demonstrate that: 1) by inserting a top functional layer above the ferroelectric, gate side injection pumped by ferroelectric switching event can be enhanced, thus increasing the MW; 2) inspired by the charge trap…

Storage Systemsstorage
IEEE Electron Device Letters
Industry researchSamsung (South Korea)12 citations

Disturb and its Mitigation in Ferroelectric Field-Effect Transistors With Large Memory Window for NAND Flash Applications

Prasanna Venkatesan · Chinsung Park · Taeyoung Song · Lance Fernandes · Dipjyoti Das · et al.

We study the disturb characteristics of ferroelectric field-effect transistors (FEFETs) with band-engineered gate stacks. We demonstrate that integrating a dielectric Al2O3 layer within the ferroelectric (FE) Hf$_{{0}.

Storage Systems
IEEE Electron Device Letters
Industry researchSK Group (South Korea)15 citations

Optimization of Programming Pulse Shape for Vertical NAND Flash Memory Using Neural Networks

Sungho Park · Jaehyeon Kim · Jonghyun Ko · Jiseong Im · Yeongheon Yang · et al.

We optimize the shape of the pulse to maximally increase threshold voltage (${V}_{\text {th}}\text {)}$during the incremental step pulse programming (ISPP) of vertical NAND (V-NAND) flash memory using neural networks (NNs). NN is trained using data on the increase in${V}_{\text {th}}$of commercial V-NAND flash memory in response to randomly shaped…

Storage Systems
Proceedings of the VLDB Endowment
Industry researchIntel (Germany) · Intel (India)21 citations

An Examination of CXL Memory Use Cases for In-Memory Database Management Systems Using SAP HANA

Minseon Ahn · Thomas Willhalm · Norman May · Donghun Lee · Suprasad Mutalik Desai · et al.

CXL-based disaggregated memory systems offer options to expand the memory beyond the limits of a single server via cache-coherent memory expansion cards or memory pools. Especially, In-Memory Database Management Systems (IMDBMSs) can benefit from alleviating two critical constraints: (1) limited memory capacity in a server and (2) long restart time…

Storage SystemsdatabaseCXL
IEEE Micro
Industry researchMicrosoft (United States)13 citations

Data Center Power and Energy Management: Past, Present, and Future

Ricardo Bianchini · Christian Belady · Anand Sivasubramaniam

This article overviews some of the key past developments in cloud data center power and energy management, where we are today, and what the future could be. This topic is gaining enormous renewed interest in the context of the conflicting needs of the AI revolution and the climate crisis.

Storage Systems
Proceedings of the VLDB Endowment
Industry researchMicrosoft Research (United Kingdom)12 citations

Bf-Tree: A Modern Read-Write-Optimized Concurrent Larger-Than-Memory Range Index

Xiangpeng Hao · Badrish Chandramouli

A B-Tree is the most widely used range index for larger-than-memory data systems. It organizes data in pages (usually 4 KB) that efficiently align with disk IO operations, fully utilizing each IO operation to narrow down the search space.

Storage Systemsindexing
Proceedings of the VLDB Endowment
Industry researchMicrosoft Research (United Kingdom)9 citations

DDS: DPU-Optimized Disaggregated Storage

Qizhen Zhang · Philip A. Bernstein · Badrish Chandramouli · Jiasheng Hu · Yiming Zheng

This paper presents DDS, a novel disaggregated storage architecture enabled by emerging networking hardware, namely DPUs (Data Processing Units). DPUs can optimize the latency and CPU consumption of disaggregated storage servers.

Storage Systemsstorage
ACM Computing Surveys
Industry researchIntel (United States) · Microsoft (United States)104 citations

An Introduction to the Compute Express Link (CXL) Interconnect

Debendra Das Sharma · Robert Blankenship · Daniel S. Berger

The Compute Express Link (CXL) is an open industry-standard interconnect between processors and devices such as accelerators, memory buffers, smart network interfaces, persistent memory, and solid-state drives. CXL offers coherency and memory semantics with bandwidth that scales with PCIe bandwidth while achieving significantly lower latency than PCIe.

Storage SystemsCXL
Proceedings of the ACM on Management of Data
Industry researchHuawei Technologies (China)10 citations

Hyper: A High-Performance and Memory-Efficient Learned Index via Hybrid Construction

Shunkang Zhang · Ji Qi · Xin Yao · André Brinkmann

Learned indexes use machine learning techniques to improve index construction. However, they often face a fundamental trade-off between performance and memory consumption, especially in dynamic environments with frequent insert and delete operations.

Storage Systemsindexing
Applied Thermal Engineering
Industry researchIntel (United States)200 citations

Liquid cooling of data centers: A necessity facing challenges

Mohammad Azarifar · Mehmet Arık · Je-Young Chang

The evolving data generation landscape requires faster and more efficient microprocessors, prompting innovative manufacturing methods for smaller and faster transistors. Transistor congestion and rising demand for parallel processing are pushing the thermal design power of microprocessors well beyond 280 W, a limit for air cooling, and are expected to…

Storage Systems
IEEE Transactions on Knowledge and Data Engineering
Industry researchTencent (China)11 citations

Data-Aware Adaptive Compression for Stream Processing

Yu Zhang · Feng Zhang · Hourun Li · Shuhao Zhang · Xiaoguang Guo · et al.

Stream processing has been in widespread use, and one of the most common application scenarios is SQL query on streams. By 2021, the global deployment of IoT endpoints reached 12.

Storage Systems
IEEE Transactions on Computers
Industry researchHuawei Technologies (China)14 citations

BFT-DSN: A Byzantine Fault-Tolerant Decentralized Storage Network

Hechuan Guo · Minghui Xu · Jiahao Zhang · Chunchi Liu · Rajiv Ranjan · et al.

With the rapid development of blockchain and its applications, the amount of data stored on decentralized storage networks (DSNs) has grown exponentially. DSNs bring together affordable storage resources from around the world to provide robust, decentralized storage services for tens of thousands of decentralized applications (dApps).

Storage Systemsfault tolerancestorage
IEEE Micro
Industry researchSK Group (South Korea)18 citations

Improving Key-Value Cache Performance With Heterogeneous Memory Tiering: A Case Study of Compute-Express-Link-Based Memory Expansion

KyungSoo Lee · Sohyun Kim · Joohee Lee · Donguk Moon · R. Kim · et al.

CXL memory brings extra bandwidth and capacity via PCIe-based memory expansion beyond DDR-based DRAM. This paper introduces the CXL 2.

Storage Systemscaching
IEEE Transactions on Electron Devices
Industry researchSK Group (South Korea)12 citations

Investigation of Endurance Characteristics in 3-D NAND Flash Memory With Trap Profile Analysis

Hyungjun Jo · Jongwoo Kim · Yonggyu Cho · Hyunyoung Shim · Jaesung Sim · et al.

In this article, the endurance characteristic of 3-D NAND Flash memory is investigated. The bitline (BL) current and the word-line (WL) current are measured for erase/write (EW) cycling up to 50k.

Storage Systems
ACM Transactions on Architecture and Code Optimization
Industry researchHuawei Technologies (China)20 citations

Rcmp: Reconstructing RDMA-Based Memory Disaggregation via CXL

Zhonghua Wang · Y. M. Guo · Kai Lü · Jiguang Wan · Daohui Wang · et al.

Memory disaggregation is a promising architecture for modern datacenters that separates compute and memory resources into independent pools connected by ultra-fast networks, which can improve memory utilization, reduce cost, and enable elastic scaling of compute and memory resources. However, existing memory disaggregation solutions based on remote…

Storage SystemsRDMACXL
IEEE Transactions on Electron Devices
Industry researchSK Group (South Korea)25 citations

Reliability Improvement in Vertical NAND Flash Cells Using Adaptive Incremental Step Pulse Programming (A-ISPP) and Incremental Step Pulse Erasing (ISPE)

Sungho Park · Ho-Nam Yoo · Yeongheon Yang · Jae‐Joon Kim · Jong‐Ho Lee

In order to improve the reliability of vertical NAND (V-NAND) flash memory cells, a scheme using adaptive incremental step pulse programming (A-ISPP) and incremental step pulse erasing (ISPE) is proposed. Incremental step pulse programming (ISPP) with adaptive step voltage is used to precisely adjust${V}_{\text {th}}$to a low target value while rapidly…

Storage Systems
05

Artificial Intelligence

Foundation models, reasoning, agents, evaluation, reliability, and model architectures.

2026

2 papers
arXiv

Learning When to Trust via Selective Context Preference Optimization

Xian Sun · Wei Chow · Yingshuo Wang · Junhao Liu · Wei Gao · Qing Wu · Lingdong Kong

Frames robustness as selective trust and trains language models to resist misleading signals without ignoring useful context.

reliabilitypreference optimizationevaluation
arXiv

Agentic Artificial Intelligence: Architectures, Taxonomies, and Evaluation of Large Language Model Agents

Arunkumar V · Gangadharan G. R. · Rajkumar Buyya

Proposes a unified taxonomy for agent perception, planning, action, tool use, collaboration, environments, and evaluation.

agentstaxonomyevaluation

2025

12 papers
ACM Transactions on Intelligent Systems and Technology
Industry researchAmazon (United States)62 citations

A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness

Fali Wang · Zhiwei Zhang · Xianren Zhang · Zongyu Wu · Tzuhao Mo · et al.

Large language models (LLMs) have demonstrated emergent abilities in text generation, question answering, and reasoning, facilitating various tasks and domains. Despite their proficiency in various tasks, LLMs like PaLM 540B and Llama-3.

Artificial IntelligenceLLM
ACM Computing Surveys
Industry researchMicrosoft Research (United Kingdom)55 citations

Domain Specialization as the Key to Make Large Language Models Disruptive: A Comprehensive Survey

Chen Ling · Xujiang Zhao · Jiaying Lu · Chengyuan Deng · Can Zheng · et al.

Large language models (LLMs) have significantly advanced the field of natural language processing (NLP), providing a highly useful, task-agnostic foundation for a wide range of applications. However, directly applying LLMs to solve sophisticated problems in specific domains meets many hurdles, caused by the heterogeneity of domain data, the…

Artificial Intelligence
ACM Transactions on Information Systems
Industry researchHuawei Technologies (China)90 citations

A Survey on the Memory Mechanism of Large Language Model-based Agents

Zeyu Zhang · Quanyu Dai · Xiaohe Bo · Chen Ma · Rui Li · et al.

Large language model (LLM)-based agents have recently attracted much attention from the research and industry communities. Compared with original LLMs, LLM-based agents are featured in their self-evolving capability, which is the basis for solving real-world problems that need long-term and complex agent-environment interactions.

Artificial Intelligenceagents
ACM Computing Surveys
Industry researchAmazon (United States) · Amazon (Germany) · Microsoft (United States) · Microsoft Research (United Kingdom)27 citations

Survey on Factuality in Large Language Models

Cunxiang Wang · Xiaoze Liu · Yuanhao Yue · Qipeng Guo · Xiangkun Hu · et al.

This survey addresses the crucial issue of factuality in Large Language Models (LLMs). As LLMs find applications across diverse domains, the reliability and accuracy of their outputs become vital.

Artificial Intelligence
Proceedings of the AAAI Conference on Artificial Intelligence
Industry researchAlibaba Group (Cayman Islands)40 citations

Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference

Han Zhao · M. Zhang · Wei Zhao · Pengxiang Ding · Siteng Huang · et al.

In recent years, applying multi-modal large language models (MLLMs) in various fields has achieved remarkable success. However, as the foundation model for many downstream tasks, MLLMs comprise the well-known Transformer network, which has a less efficient quadratic computation complexity.

Artificial Intelligence
arXiv

Large Language Model Agent: A Survey on Methodology, Applications and Challenges

Junyu Luo et al.

Organizes LLM-agent architecture, collaboration, evolution, evaluation, applications, and open research challenges.

agentssurveymulti-agent systems
IEEE Transactions on Pattern Analysis and Machine Intelligence
Industry researchHuawei Technologies (China) · Huawei Technologies (Sweden)37 citations

NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning

Bingqian Lin · Yunshuang Nie · Ziming Wei · Jiaqi Chen · Shikui Ma · et al.

Vision-and-Language Navigation (VLN), as a crucial research problem of Embodied AI, requires an embodied agent to navigate through complex 3D environments following natural language instructions. Recent research has highlighted the promising capacity of large language models (LLMs) in VLN by improving navigational reasoning accuracy and interpretability.

Artificial IntelligenceLLMreasoning
ACM Transactions on Software Engineering and Methodology
Industry researchHuawei Technologies (China)27 citations

An Empirical Study of Retrieval-Augmented Code Generation: Challenges and Opportunities

Zezhou Yang · Sirong Chen · Cuiyun Gao · Zhenhao Li · Xing Hu · et al.

Code generation aims to automatically generate code snippets of specific programming language according to natural language descriptions. The continuous advancements in deep learning, particularly pre-trained models, have empowered the code generation task to achieve remarkable performance.

Artificial Intelligenceretrieval
IEEE Transactions on Pattern Analysis and Machine Intelligence
Industry researchAlibaba Group (China)43 citations

Uni-MoE: Scaling Unified Multimodal LLMs With Mixture of Experts

Yunxin Li · Shenyuan Jiang · Baotian Hu · Longyue Wang · Wanqi Zhong · et al.

Recent advancements in Multimodal Large Language Models (MLLMs) underscore the significance of scalable models and data to boost performance, yet this often incurs substantial computational costs. Although the Mixture of Experts (MoE) architecture has been employed to scale large language or visual-language models efficiently, these efforts typically…

Artificial IntelligenceLLMmultimodal
arXiv

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

DeepSeek-AI

Studies large-scale reinforcement learning for emergent reasoning and presents a multi-stage path to capable, readable reasoning models.

reasoningreinforcement learningLLM
IEEE Transactions on Cognitive Communications and Networking
Industry researchTencent (China)38 citations

AutoHMA-LLM: Efficient Task Coordination and Execution in Heterogeneous Multi-Agent Systems Using Hybrid Large Language Models

Tingting Yang · Ping Feng · Qixin Guo · Jindi Zhang · Xiufeng Zhang · et al.

Heterogeneous multi-agent systems (HMAS) comprise various intelligent agents with specialized functions, such as drones, ground robots, and automated devices, working in coordinated settings. This paper presents AutoHMA-LLM, a novel framework that combines Large Language Models (LLMs) with classical control algorithms to address the challenges of task…

Artificial IntelligenceLLMagents
IEEE Transactions on Audio Speech and Language Processing
Industry researchMicrosoft (United States)103 citations

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Sanyuan Chen · Chengyi Wang · Yu Wu · Ziqiang Zhang · Long Zhou · et al.

We introduce a language modeling approach for text to speech synthesis (TTS). Specifically, we train aneural codec language model(calledVALL-E) using discrete codes derived from an off-the-shelf neural audio codec model, and regard TTS as a conditional language modeling task rather than continuous signal regression as in previous work.

Artificial Intelligence

2024

17 papers
arXiv

Titans: Learning to Memorize at Test Time

Ali Behrouz · Peilin Zhong · Vahab Mirrokni

Adds a neural long-term memory module that learns at inference time and complements attention's short-term context.

memoryarchitecturelong context
arXiv

DeepSeek-V3 Technical Report

DeepSeek-AI

Describes a 671B-parameter mixture-of-experts model using latent attention, auxiliary-loss-free load balancing, and multi-token prediction.

mixture of expertsLLMtraining systems
National Science Review
Industry researchTencent (China)634 citations

A survey on multimodal large language models

Shukang Yin · Chaoyou Fu · Sirui Zhao · Ke Li · Xing Sun · et al.

Recently, the multimodal large language model (MLLM) represented by GPT-4V has been a new rising research hotspot, which uses powerful large language models (LLMs) as a brain to perform multimodal tasks. The surprising emergent capabilities of the MLLM, such as writing stories based on images and optical character recognition-free math reasoning, are…

Artificial Intelligencemultimodal
Frontiers of Computer Science
Industry researchTencent (China)248 citations

Large language models for generative information extraction: a survey

Derong Xu · Wei Chen · Wenjun Peng · Chao Zhang · Tong Xu · et al.

Abstract Information Extraction (IE) aims to extract structural knowledge from plain natural language texts. Recently, generative Large Language Models (LLMs) have demonstrated remarkable capabilities in text understanding and generation.

Artificial Intelligence
ACM Transactions on Software Engineering and Methodology
Industry researchTencent (China)47 citations

On the Effectiveness of Large Language Models in Domain-Specific Code Generation

Xiaodong Gu · Meng Chen · Yalan Lin · Yuhan Hu · Hongyu Zhang · et al.

Large language models (LLMs) such as ChatGPT have shown remarkable capabilities in code generation. Despite significant achievements, they rely on enormous training data to acquire a broad spectrum of open-domain knowledge.

Artificial Intelligence
Open-access preprint
Industry researchBaidu (China)649 citations

A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models

Wenqi Fan · Yujuan Ding · Liangbo Ning · Shijie Wang · Hengyun Li · et al.

As one of the most advanced techniques in AI, Retrieval-Augmented Generation (RAG) can offer reliable and up-to-date external knowledge, providing huge convenience for numerous tasks. Particularly in the era of AI-Generated Content (AIGC), the powerful capacity of retrieval in providing additional knowledge enables RAG to assist existing generative AI…

Artificial IntelligenceLLMretrieval
Open-access preprint
Industry researchMicrosoft Research (United Kingdom)192 citations

Splitwise: Efficient Generative LLM Inference Using Phase Splitting

Pratyush Patel · Esha Choukse · Chaojie Zhang · Aashaka Shah · Íñigo Goiri · et al.

Generative large language model (LLM) applications are growing rapidly, leading to large-scale deployments of expensive and power-hungry GPUs. Our characterization of LLM inference shows that each inference request undergoes two phases: a compute-intensive prompt computation phase and a memory intensive token generation phase, each with distinct…

Artificial IntelligenceLLM
ACM SIGKDD Explorations Newsletter
Industry researchBaidu (China)177 citations

Exploring the Potential of Large Language Models (LLMs)in Learning on Graphs

Zhikai Chen · Haitao Mao · Hang Li · Wei Jin · Hongzhi Wen · et al.

Learning on Graphs has attracted immense attention due to its wide real-world applications. The most popular pipeline for learning on graphs with textual node attributes primarily relies on Graph Neural Networks (GNNs), and utilizes shallow text embedding as initial node representations, which has limitations in general knowledge and profound semantic…

Artificial IntelligenceLLM
Proceedings of the AAAI Conference on Artificial Intelligence
Industry researchAmazon (United States)36 citations

CLIPSyntel: CLIP and LLM Synergy for Multimodal Question Summarization in Healthcare

Akash Ghosh · A. Seetharama Acharya · Raghav Jain · Sriparna Saha · Aman Chadha · et al.

In the era of modern healthcare, swiftly generating medical question summaries is crucial for informed and timely patient care. Despite the increasing complexity and volume of medical data, existing studies have focused solely on text-based summarization, neglecting the integration of visual information.

Artificial IntelligenceLLMmultimodal
Proceedings of the AAAI Conference on Artificial Intelligence
Industry researchGoogle (United States) · Nvidia (United Kingdom)39 citations

Directed Diffusion: Direct Control of Object Placement through Attention Guidance

Wan-Duo Kurt · Avisek Lahiri · Jonathan Lewis · Thomas Leung · W. Bastiaan Kleijn

Text-guided diffusion models such as DALLE-2, Imagen, and Stable Diffusion are able to generate an effectively endless variety of images given only a short text prompt describing the desired image content. In many cases the images are of very high quality.

Artificial Intelligencediffusion
Proceedings of the AAAI Conference on Artificial Intelligence
Industry researchAlibaba Group (United States)65 citations

FontDiffuser: One-Shot Font Generation via Denoising Diffusion with Multi-Scale Content Aggregation and Style Contrastive Learning

Zhenhua Yang · Dezhi Peng · Yuxin Kong · Yuyi Zhang · Cong Yao · et al.

Automatic font generation is an imitation task, which aims to create a font library that mimics the style of reference images while preserving the content from source images. Although existing font generation methods have achieved satisfactory performance, they still struggle with complex characters and large style variations.

Artificial Intelligencediffusion
Proceedings of the AAAI Conference on Artificial Intelligence
Industry researchTencent (China)763 citations

T2I-Adapter: Learning Adapters to Dig Out More Controllable Ability for Text-to-Image Diffusion Models

Chong Mou · Xintao Wang · Liangbin Xie · Yanze Wu · Jian Zhang · et al.

The incredible generative ability of large-scale text-to-image (T2I) models has demonstrated strong power of learning complex structures and meaningful semantics. However, relying solely on text prompts cannot fully take advantage of the knowledge learned by the model, especially when flexible and accurate controlling (e.

Artificial Intelligencediffusion
ACM Transactions on Asian and Low-Resource Language Information Processing
Industry researchAlibaba Group (China)63 citations

CodeKGC: Code Language Model for Generative Knowledge Graph Construction

Zhen Bi · Jing Chen · Yinuo Jiang · Feiyu Xiong · Wei Guo · et al.

Current generative knowledge graph construction approaches usually fail to capture structural knowledge by simply flattening natural language into serialized texts or a specification language. However, large generative language model trained on structured data such as code has demonstrated impressive capability in understanding natural language for…

Artificial Intelligenceedge
ACM Transactions on Intelligent Systems and Technology
Industry researchMicrosoft Research Asia (China)2604 citations

A Survey on Evaluation of Large Language Models

Yupeng Chang · Xu Wang · Jindong Wang · Yuan Wu · Linyi Yang · et al.

Large language models (LLMs) are gaining increasing popularity in both academia and industry, owing to their unprecedented performance in various applications. As LLMs continue to play a vital role in both research and daily use, their evaluation becomes increasingly critical, not only at the task level, but also at the society level for better…

Artificial Intelligence
Open-access preprint
Industry researchTencent (China)286 citations

A Survey on Multimodal Large Language Models for Autonomous Driving

Can Cui · Yunsheng Ma · Xu Cao · Wenqian Ye · Yang Zhou · et al.

With the emergence of Large Language Models (LLMs) and Vision Foundation Models (VFMs), multimodal AI systems benefiting from large models have the potential to equally perceive the real world, make decisions, and control tools as humans. In recent months, LLMs have shown widespread attention in autonomous driving and map systems.

Artificial Intelligencemultimodal
Computational Linguistics
Industry researchAdobe Systems (United States) · Intel (United States)570 citations

Bias and Fairness in Large Language Models: A Survey

Isabel O. Gallegos · Ryan A. Rossi · Joe Barrow · Md Mehrab Tanjim · Sungchul Kim · et al.

Abstract Rapid advancements of large language models (LLMs) have enabled the processing, understanding, and generation of human-like text, with increasing integration into systems that touch our social sphere. Despite this success, these models can learn, perpetuate, and amplify harmful social biases.

Artificial Intelligence
Proceedings of the VLDB Endowment
Industry researchAlibaba Group (United States)244 citations

Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation

Dawei Gao · Haibin Wang · Yaliang Li · Xiuyu Sun · Yichen Qian · et al.

Large language models (LLMs) have emerged as a new paradigm for Text-to-SQL task. However, the absence of a systematical benchmark inhibits the development of designing effective, efficient and economic LLM-based Text-to-SQL solutions.

Artificial Intelligence