Search arXiv⌕ Search

arXiv · 2610.08189

Strength Is Not Value: Rethinking Approximate Functional Dependency Measures for Error Detection

Abstract

Approximate functional dependency (AFD) strength measures are commonly used to prioritize candidate dependencies, including when selected dependencies support downstream tasks such as error detection. This practice implicitly assumes that dependencies ranked higher by strength are also better candidates for the task they are intended to support. We test this assumption directly for error detection. We evaluate detection value through precision and recall without imposing a single weighting on different error types, and introduce an explicit false-positive/false-negative cost only when a single candidate-selection objective is required. Across three real-world relations from an established AFD benchmark, $1{,}155$ corruption-and-detection runs, and $513{,}480$ candidate-level evaluations, we find a clear asymmetry. Strength measures tend to track precision positively, but their associations with recall are substantially less stable, with both magnitude and direction varying across measures and experimental settings. Strength-based prioritization is also unreliable at the ranking level: top-$k$ sets chosen by strength often have little overlap with those favored by actual detection performance, and strength-based selections can incur substantial regret. An important part of the recall mismatch is associated with structural detection opportunity under the candidate equivalence-class partition, which tracks recall more consistently overall than the strength measures we evaluate. Corruption mechanism and several static structural properties account for selected effects but do not provide a common explanation across settings, leaving part of the mismatch unresolved. Overall, AFD strength and detection value are distinct and should be evaluated separately.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Xiaolong Wan, Xixian Han. 2026-10-06. Strength Is Not Value: Rethinking Approximate Functional Dependency Measures for Error Detection. https://arxiv.org/abs/2610.08189

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Crane: An Accurate and Scalable Neural Sketch for Graph Stream Summarization

Graph streams are rapidly evolving sequences of edges that convey continuously changing relationships among entities, playing a crucial role in domains such as networking, finance, and cybersecurity. Their massive scale and high dynamism make obtaining accurate statistics challenging with limited memory constraints. Traditional methods summarize graph streams through hand-crafted sketches, while recent studies have begun to replace these sketches with neural counterparts to improve adaptability and accuracy. However, this shift faces a major challenge: under limited memory, dominant frequent items tend to overshadow rare ones, hindering the neural network's ability to recover accurate statistics. To address this, we propose Crane, a hierarchical neural sketch architecture for graph stream summarization. Crane uses a hierarchical carry mechanism that automatically elevates frequent items to higher memory layers, reducing interference between frequent and infrequent items within the same layer. To better accommodate real-world deployment, Crane further adopts an adaptive memory expansion strategy that dynamically adds new layers once the occupancy of the top layer exceeds a threshold, enabling scalability across diverse data magnitudes. Experimental results show that Crane significantly reduces estimation error compared to state-of-the-art (SOTA) methods, while also offering competitive throughput.

cs.DB↗

Hierarchical Decomposition of Separable Workflow-Nets

The Partially Ordered Workflow Language (POWL) has recently emerged as a process modeling notation, offering strong quality guarantees and high expressiveness. While early versions of POWL relied on strict block-structured operators for choices and loops, the language has recently evolved into POWL 2.0, introducing choice graphs to enable the modeling of non-block-structured decisions and cycles. To bridge the gap between the theoretical advantages of POWL and the practical need for compatibility with established notations, robust model transformations are required. This paper presents a novel algorithm for transforming safe and sound workflow nets (WF-nets) into equivalent POWL 2.0 models. The algorithm recursively identifies structural patterns within the WF-net and translates them into their POWL representation. Unlike the previous approach that required separate detection strategies for exclusive choices and loops, our new algorithm utilizes choice graphs to capture generalized decision and cyclic patterns. We formally prove the correctness of our approach, showing that the generated POWL model preserves the language of the input WF-net. Furthermore, we prove the completeness of our algorithm on the class of separable WF-nets, which corresponds to nets constructed via the hierarchical nesting of state machines and marked graphs. We evaluate our algorithm on large-scale process models to demonstrate its high scalability. Furthermore, to test its practical expressiveness, we applied it to a benchmark of 1,493 industrial and synthetic process models. Our algorithm successfully transformed all models in this benchmark, suggesting that POWL 2.0's expressive power is generally sufficient to capture the complex logic found in real-world business processes. This work paves the way for broader adoption of POWL in practical process analysis and improvement applications.

cs.DB↗

Disk-Resident Graph ANN Search: An Experimental Evaluation

As data volumes grow while memory capacity remains limited, disk-resident graph-based approximate nearest neighbor (ANN) methods have become a practical alternative to memory-resident designs, shifting the bottleneck from computation to disk I/O. However, since their technical designs diverge widely across storage, layout, and execution paradigms, a systematic understanding of their fundamental performance trade-offs remains elusive. This paper presents a comprehensive experimental study of disk-resident graph-based ANN methods. First, we decompose such systems into five key technical components, i.e., storage strategy, disk layout, cache management, query execution, and update mechanism, and build a unified taxonomy of existing designs across these components. Second, we conduct fine-grained evaluations of representative strategies for each technical component to analyze the trade-offs in throughput, recall, and resource utilization. Third, we perform comprehensive end-to-end experiments and parameter-sensitivity analyses to evaluate overall system performance under diverse configurations. Fourth, our study reveals several non-obvious findings: (1) vector dimensionality fundamentally reshapes component effectiveness, necessitating dimension-aware design; (2) existing layout strategies exhibit surprisingly low I/O utilization (less than or equal to 15%); (3) page size critically affects feasibility and efficiency, with smaller pages preferred when layouts are carefully optimized; and (4) update strategies present clear workload-dependent trade-offs between in-place and out-of-place designs. Based on these findings, we derive practical guidelines for system design and configuration, and outline promising directions for future research.

cs.DB↗