Search DistillSys

Find a concept

Type at least two characters to search lessons, designs, papers, and interview prep.

Research area · 2024–2026

Databases

Query processing, indexing, transactions, data quality, and modern database architectures.

31 curated papers

31 papers

2026

3 papers
arXiv

PLB: Priority-Aware Load Balancing for Replicated Databases under Constrained Resources

Belkis Djeffal · Pierre Bourhis · Romain Rouvoy

Enforces priority differentiation through replica assignment while retaining high utilization through controlled capacity borrowing.

replicationOLAPload balancing
arXiv

Tytan: Interactive Neurosymbolic Construction of Analytic Semantic Schemas from Relational Data

Donna Hooshmand · Shubham Shahi · Cameron Barrie · Abhratanu Dutta · Marko Sterbentz · Harper Pack · Kristian J. Hammond

Combines symbolic database analysis, language-model inference, and targeted user questions to construct usable analytic semantic schemas.

semantic layerschemaneurosymbolic
Proceedings of the VLDB Endowment
Industry researchAlibaba Group (United States)3 citations

Why Database Manuals are Not Enough: Efficient and Reliable Configuration Tuning for DBMSs via Code-Driven LLM Agents

Xinyi Zhang · Tiantian Chen · Zhentao Han · Zhaoyan Hong · Wei Lu · et al.

Modern database management systems (DBMSs) expose hundreds of configuration knobs that critically influence performance. Existing automated tuning methods either adopt a data-driven paradigm, which incurs substantial overhead, or rely on manual-driven heuristics extracted from database documentation, which are often limited and overly generic.

DatabasesdatabaseLLMagents

2025

6 papers
arXiv

Database Theory in Action: Yannakakis' Algorithm

Paraschos Koutris · Stijn Vansummeren · Qichen Wang · Yisu Remy Wang · Xiangyao Yu

Reviews recent work that makes the optimal algorithm for acyclic joins practical in modern query engines.

joinsquery processingdatabase theory
arXiv

LMG Index: A Robust and Efficient Learned Index Framework for Multi-Dimensional Performance Balance

Yuzhen Chen · Bin Yao

Balances lookup, range-query, update, stability, and space objectives in one learned-index framework.

learned indexesquery performanceindexing
Proceedings of the VLDB Endowment
Industry researchTencent (China)4 citations

SiriusBI: A Comprehensive LLM-Powered Solution for Data Analytics in Business Intelligence

Jie Jiang · Haining Xie · Siqi Shen · Yu Shen · Zihan Zhang · et al.

With the proliferation of Large Language Models (LLMs) in Business Intelligence (BI), existing solutions face critical challenges in industrial deployments: functionality deficiencies from legacy systems failing to meet evolving LLM-era user demands, interaction limitations from single-round SQL generation paradigms inadequate for multi-round…

DatabasesanalyticsLLM
Proceedings of the VLDB Endowment
Industry researchMicrosoft Research (United Kingdom)8 citations

Scaling GPU-Accelerated Databases Beyond GPU Memory Size

Yinan Li · Bailu Ding · Ziyun Wei · Lukas M. Maas · Momin Al-Ghosien · et al.

There has been considerable interest in leveraging GPUs' computational power and high memory bandwidth for analytical database workloads. However, their limited memory capacity remains a fundamental limitation for databases whose sizes far exceed the GPU memory size.

Databasesdatabase
Proceedings of the ACM on Management of Data
Industry researchAlibaba Group (China)5 citations

Yannakakis+: Practical Acyclic Query Evaluation with Theoretical Guarantees

Qichen Wang · Bingnan Chen · Binyang Dai · Ke Yi · Feifei Li · et al.

Acyclic conjunctive queries form the backbone of most analytical workloads, and have been extensively studied in the literature from both theoretical and practical angles. However, there is still a large divide between theory and practice.

Databasesquery optimization
ACM SIGMOD Record
Industry researchMicrosoft (United States) · Alibaba Group (China)4 citations

A Roadmap to Graph Analytics

Angela Bonifati · M. TAMER ÖZSU · Yuanyuan Tian · Hannes Voigt · Wenyuan Yu · et al.

Graphs are ubiquitous data structures used in a large spectrum of applications, spanning from transportation networks, financial networks, social networks, product-order transactions and biomedical applications [33]. A recent survey on the usage of graph applications from real users has highlighted the fact that analytics is the most time-consuming task…

Databasesanalytics

2024

22 papers
Proceedings of the ACM on Management of Data
Industry researchAlibaba Group (China)4 citations

Towards a Converged Relational-Graph Optimization Framework

Yunkai Lou · Longbin Lai · Bingqing Lyu · Y Yang · Xiaoli Zhou · et al.

The recent ISO SQL:2023 standard adopts SQL/PGQ (Property Graph Queries), facilitating graph-like querying within relational databases. This advancement, however, underscores a significant gap in how to effectively optimize SQL/PGQ queries within relational database systems.

Databases
Proceedings of the ACM on Management of Data
Industry researchMicrosoft (United States)3 citations

Output-sensitive Conjunctive Query Evaluation

Shaleen Deep · Hangdong Zhao · Austen Z. Fan · Paraschos Koutris

Join evaluation is one of the most fundamental operations performed by database systems and arguably the most well-studied problem in the Database community. A staggering number of join algorithms have been developed, and commercial database engines use finely tuned join heuristics that take into account many factors including the selectivity of…

Databasesquery optimization
Proceedings of the VLDB Endowment
Industry researchAlibaba Group (United States)10 citations

PRICE: A Pretrained Model for Cross-Database Cardinality Estimation

Tianjing Zeng · Junwei Lan · Jiahong Ma · Wenqing Wei · Rong Zhu · et al.

Cardinality estimation (CardEst) is essential for optimizing query execution plans. Recent ML-based CardEst methods achieve high accuracy but face deployment challenges due to high preparation costs and lack of transferability across databases.

Databasesdatabase
Proceedings of the VLDB Endowment
Industry researchDatabricks (United States)12 citations

Adaptive and Robust Query Execution for Lakehouses at Scale

Maryann Xue · Yingyi Bu · Abhishek Somani · Wen Fan · Ziqi Liu · et al.

Many organizations have embraced the "Lakehouse" data management paradigm, which involves constructing structured data warehouses on top of open, unstructured data lakes. This approach stands in stark contrast to traditional, closed, relational databases and introduces challenges for performance and stability of distributed query processors.

Databasesquery optimization
Proceedings of the VLDB Endowment
Industry researchIntel (Germany) · Intel (India)21 citations

An Examination of CXL Memory Use Cases for In-Memory Database Management Systems Using SAP HANA

Minseon Ahn · Thomas Willhalm · Norman May · Donghun Lee · Suprasad Mutalik Desai · et al.

CXL-based disaggregated memory systems offer options to expand the memory beyond the limits of a single server via cache-coherent memory expansion cards or memory pools. Especially, In-Memory Database Management Systems (IMDBMSs) can benefit from alleviating two critical constraints: (1) limited memory capacity in a server and (2) long restart time…

DatabasesdatabaseCXL
Proceedings of the VLDB Endowment
Industry researchGoogle (United States)9 citations

SQL Has Problems. We Can Fix Them: Pipe Syntax In SQL

Jeff Shute · Shannon Bales · Matthew W. Brown · Jean-Daniel Browne · Brandon Dolphin · et al.

SQL has been extremely successful as the de facto standard language for working with data. Virtually all mainstream database-like systems use SQL as their primary query language.

Databases
Proceedings of the VLDB Endowment
Industry researchTencent (China)12 citations

TDSQL: Tencent Distributed Database System

Yuxing Chen · Anqun Pan · Hailin Lei · Anda Ye · S. Han · et al.

Distributed databases have become indispensable in contemporary computing and data processing, owing to their pivotal role in ensuring high availability and scalability. They effectively cater to the requirements of data management and high-concurrency access.

Databasesdatabase
Proceedings of the VLDB Endowment
Industry researchTencent (China)5 citations

X-Stor: A Cloud-Native NoSQL Database Service with Multi-Model Support

Hongyu Lei · Chunhua Li · Ke Zhou · Jianping Zhu · Kezhou Yan · et al.

In recent years at Tencent, we have observed that the use of multiple NoSQL databases for storing business data with diverse models has led to increased programming and deployment costs, as well as inefficient maintenance and underutilized resources. In this paper, we report X-Stor, a cloud-native NoSQL database system that supports multiple data models…

Databasesclouddatabase
Proceedings of the VLDB Endowment
Industry researchGoogle (United States)6 citations

Towards Optimal Transaction Scheduling

Audrey Cheng · Aaron Kabcenell · Jason Chan · Xiao Shi · Peter Bailis · et al.

Maximizing transaction throughput is key to high-performance database systems, which focus on minimizing data access conflicts to improve performance. However, finding efficient schedules that reduce conflicts remains an open problem.

Databasesschedulingtransactions
Proceedings of the VLDB Endowment
Industry researchAmazon (United States)35 citations

Why TPC is Not Enough: An Analysis of the Amazon Redshift Fleet

Alexander van Renen · Dominik Horn · Pascal Pfeil · Kapil Vaidya · Wenjian Dong · et al.

Database research and development is heavily influenced by benchmarks, such as the industry-standard TPC-H and TPC-DS for analytical systems. However, these twenty-year-old benchmarks neither capture how databases are deployed nor what workloads modern cloud data warehouse systems face these days.

Databases
The VLDB Journal
Industry researchAlibaba Group (China)14 citations

A survey on hybrid transactional and analytical processing

Haoze Song · Wenchao Zhou · Heming Cui · X. Peng · Feifei Li

Abstract To provide applications with the ability to analyze fresh data and eliminate the time-consuming ETL workflow, hybrid transactional and analytical (HTAP) systems have been developed to serve online transaction processing and online analytical processing workloads in a single system. In recent years, HTAP systems have attracted considerable…

Databasestransactions
Proceedings of the VLDB Endowment
Industry researchMicrosoft Research (United Kingdom)14 citations

DEX: Scalable Range Indexing on Disaggregated Memory

Baotong Lu · Kaisong Huang · Chieh-Jan Mike Liang · Tianzheng Wang · Eric Lo

Memory disaggregation can potentially allow memory-optimized range indexes such as B+-trees to scale beyond one machine while attaining high hardware utilization and low cost. Designing scalable indexes on disaggregated memory, however, is challenging due to rudimentary caching, unprincipled offloading and excessive inconsistency among servers.

Databasesindexing
Proceedings of the ACM on Management of Data
Industry researchAlibaba Group (China)4 citations

Relational Algorithms for Top-k Query Evaluation

Qichen Wang · Qiyao Luo · Yilei Wang

The evaluation of top-k conjunctive queries, a staple in business analysis, often requires evaluating the conjunctive query prior to filtering the top-k results, leading to a significant computational overhead within Database Management Systems (DBMSs). While efficient algorithms have been proposed, their integration into DBMSs remains arduous.

Databasesquery optimization
Proceedings of the ACM on Management of Data
Industry researchMicrosoft (United States)5 citations

Wii: Dynamic Budget Reallocation In Index Tuning

Xiaoying Wang · Wentao Wu · Chi Wang · Vivek Narasayya · Surajit Chaudhuri

Index tuning aims to find the optimal index configuration for an input workload. It is often a time-consuming and resource-intensive process, largely attributed to the huge amount of "what-if" calls made to the query optimizer during configuration enumeration.

Databasesindexing
Open-access preprint
Industry researchAmazon (United States) · Amazon (Germany)20 citations

Stage: Query Execution Time Prediction in Amazon Redshift

Ziniu Wu · Ryan Marcus · Zhengchun Liu · Parimarjan Negi · Vikram Nathan · et al.

Query performance (e. g.

Databasesquery optimization
ACM SIGMOD Record
Industry researchMicrosoft (United States)3 citations

Auto-Tables: Relationalize Tables without Using Examples

Peng Li · Yeye He · Cong Yan · Yue Wang · Surajit Chaudhuri

Relational tables, where each row corresponds to an entity and each column corresponds to an attribute, have been the standard for tables in relational databases. However, such a standard cannot be taken for granted when dealing with tables "in the wild".

Databases
arXiv

HTAP Databases: A Survey

Chao Zhang · Guoliang Li · Jintao Zhang · Xinning Zhang · Jianhua Feng

Classifies hybrid transactional/analytical systems by storage architecture and reviews synchronization, optimization, and scheduling techniques.

HTAPtransactionsanalytics
ICSE 2024 / arXiv

CERT: Finding Performance Issues in Database Systems Through the Lens of Cardinality Estimation

Jiale Chen · Yu Liang · Yuheng Shen · Yuting Chen · Jiachi Chen · Yinxing Xue

Uses cardinality-estimation consistency rules to find confirmed performance bugs in widely used database systems.

cardinality estimationtestingquery optimization
Proceedings of the ACM on Management of Data
Industry researchMicrosoft Research (United Kingdom)6 citations

Optimizing Distributed Protocols with Query Rewrites

David Chu · Rithvik Panchapakesan · Shadaj Laddad · Lucky E. Katahanas · Chris Liu · et al.

Distributed protocols such as 2PC and Paxos lie at the core of many systems in the cloud, but standard implementations do not scale. New scalable distributed protocols are developed through careful analysis and rewrites, but this process is ad hoc and error-prone.

Databasesquery optimization
Proceedings of the ACM on Management of Data
Industry researchMicrosoft (United States)9 citations

Sibyl: Forecasting Time-Evolving Query Workloads

Hanxian Huang · Tarique Siddiqui · Rana Alotaibi · Carlo Curino · Jyoti Leeka · et al.

Database systems often rely on historical query traces to perform workload-based performance tuning. However, real production workloads are time-evolving, making historical queries ineffective for optimizing future workloads.

Databasesquery optimization
Proceedings of the VLDB Endowment
Industry researchGoogle (United States)5 citations

POLAR: Adaptive and Non-invasive Join Order Selection via Plans of Least Resistance

David Justen · Daniel P. Ritter · Campbell Fraser · Andrew Lamb · Allison Lee · et al.

Join ordering and query optimization are crucial for query performance but remain challenging due to unknown or changing characteristics of query intermediates, especially for complex queries with many joins. Over the past two decades, a spectrum of techniques for adaptive query processing (AQP)---including inter-/intra-operator adaptivity and tuple…

Databases
Proceedings of the VLDB Endowment
Industry researchAlibaba Group (China)19 citations

Eraser: Eliminating Performance Regression on Learned Query Optimizer

Lianggui Weng · Rong Zhu · Di Wu · Bolin Ding · Bolong Zheng · et al.

Efficient query optimization is crucial for database management systems. Recently, machine learning models have been applied in query optimizers to generate better plans, but the unpredictable performance regressions prevent them from being truly applicable.

Databasesquery optimization