Search arXiv⌕ Search

arXiv · 2610.01646

DIADA: Automatic Data Composition in Data Lakes

Abstract

Data lakes contain a plethora of attributes scattered across many tables that, when combined, provide enhanced assets for data analysis. Nonetheless, deciding which attributes belong together in meaningful relations remains a manual, per-task effort. Merging by joinability alone provides no guarantees regarding attribute relevance, while selecting features against a single target discards attributes useful to other tasks. To address this gap, we introduce the data composition problem: organizing a fragmented, heterogeneous lake into meaningful relations, agnostic of any particular analytical task so that the resulting organization can serve as a common foundation for diverse downstream analyses. We propose DIADA, a composition system that employs multivariate dependence as the criterion for assessing the meaningfulness of a relation and approximates it by hypothesizing independence among attributes and identifying those sets that violate this hypothesis. To do so, we map the attributes to a predicate space, forming a lattice under inclusion and mining those predicate sets that exhibit dependence among their constituents. We contribute a dedicated and scalable algorithm to effectively explore this space, outscaling classical algorithms for mining relationships, thus discovering dependencies that would otherwise be impractical to identify. We demonstrate that applying a single data composition process benefits diverse potential downstream tasks. This is the result of providing a subset of low-noise, statistically relevant attributes that increases the confidence that detected patterns are grounded in real relationships, thus preventing common modeling issues in large-scale environments.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Marc Maynou, Albert Martin, Sergi Nadal, Anna Queralt, Oscar Romero. 2026-10-01. DIADA: Automatic Data Composition in Data Lakes. https://arxiv.org/abs/2610.01646

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

VectorMaton: Efficient Vector Search with Pattern Constraints via an Enhanced Suffix Automaton

Approximate nearest neighbor search (ANNS) has become a cornerstone in modern vector database systems. Given a query vector, ANNS retrieves the closest vectors from a set of base vectors. In real-world applications, vectors are often accompanied by additional information, such as sequences or structured attributes, motivating the need for fine-grained vector search with constraints on this auxiliary data. Existing methods support attribute-based filtering or range-based filtering on categorical and numerical attributes, but they do not support pattern predicates over sequence attributes. In relational databases, predicates such as LIKE and CONTAINS are fundamental operators for filtering records based on substring patterns. As vector databases increasingly adopt SQL-style query interfaces, enabling pattern predicates over sequence attributes (e.g., texts and biological sequences) alongside vector similarity search becomes essential. In this paper, we formulate a novel problem: given a set of vectors each associated with a sequence, retrieve the nearest vectors whose sequences contain a given query pattern. To address this challenge, we propose VectorMaton, an automaton-based index that integrates pattern filtering with efficient vector search, while maintaining an index size comparable to the dataset size. Extensive experiments on real-world datasets demonstrate that VectorMaton consistently outperforms all baselines, achieving up to 10x higher query throughput at the same accuracy and up to 18x reduction in index size.

cs.DB↗

Recommendation Systems for Exploratory Data Tasks

A large class of data-centric tasks is exploratory, where users iteratively steer workflows, refining subjective goals as new insights emerge. These Exploratory Data Tasks (EDTs) are performed by millions of users with varying levels of expertise to understand unfamiliar data, discover trends, and identify evidence that informs critical decision-making. However, a key challenge in EDTs is the enormous space of possible actions that one can take at each step: users struggle to choose among thousands of joins, transformations, and aggregations, causing "exploration paralysis". Because EDT workflows are interconnected, each choice impacts subsequent exploration, and suboptimal choices can lead to inefficiency, missed insights, confirmation bias, and incomplete coverage. This calls for intelligent recommendations that efficiently guide users toward optimal EDT actions. We envision recommendation as a core capability of data systems, proactively guiding users toward promising actions and thereby lowering the barrier to exploratory data tasks. EDT recommendation is challenging because the action space is combinatorial and actions are data-dependent, which require costly materialization. Moreover, recommendation often involves bundles or sequences of actions across interdependent tasks, requiring coordination across tasks. In this paper, we present our vision of EDT recommendation systems along two axes: single-task vs. multi-task settings and single-action vs. multi-action recommendations. We outline a research agenda that progresses from recommending individual EDT actions to constrained bundles and sequences of actions, and ultimately to coordinated recommendations across interconnected EDTs. We identify research directions for incorporating various contexts (user, data, task, and ecosystem), addressing efficiency challenges, and coordinating across tasks.

cs.DB↗

HakiCC: LLM-Driven Multi-Agent Design and Optimization of Concurrency Control Protocols

Large language models (LLMs) have recently been applied in systems research as a tool to reduce human-intensive engineering effort through cost-efficient automation. Decades of research have produced a rich landscape of concurrency control (CC) protocols, each encoding distinct trade-offs in correctness, throughput, and abort behavior. However, most applications in practice default to 2PL or OCC, because selecting and adapting a protocol to a specific application requires expert knowledge that is rarely available to application designers. This is a wasted opportunity, as an application-specific CC protocol can yield significant performance advantages over a generic baseline, but designing one requires deep expertise in CC protocol design. In this paper, we propose HakiCC, an LLM-driven multi-agent pipeline that automatically designs, verifies, and optimizes concurrency control protocols tailored to a given target application. HakiCC provides a two-stage pipeline. In Stage 1, a multi-agent system takes a workload description as input and generates an application-specific CC protocol implementation, which is iteratively repaired and verified for conflict-serializability. In Stage 2, the verified protocol is further optimized for that application through an LLM-driven evolutionary loop targeting correctness and throughput. We evaluate HakiCC on TPC-C and AuctionMark as target workloads, producing and reporting ten application-specific CC protocols. All ten are conflict-serializable after Stage 1; Stage 2 improves throughput for every protocol, with average gains of +50.6% for TPC-C protocols and +92.2% for AuctionMark protocols.

cs.DB↗