Search arXiv⌕ Search

arXiv · 2610.02603

Look Here or Look Across: Unified Cardinality Constraints for N-ary Relationships

Abstract

A cardinality constraint written on an edge of an entity-relationship diagram admits two opposite readings. Under the reading used by UML and by Chen's original model, the label is read with the entity set on the same side of the relationship; under the reading used by standard database textbooks, it is read with the entity set on the opposite side. The two readings are exact opposites, so a reader who assumes the wrong one takes away the opposite of what the designer meant. For binary relationships the difficulty is confined to interpretation, because the two constraints a binary relationship set admits can both be drawn. For relationships among three or more entity sets the picture is worse: a ternary relationship set admits twelve cardinality constraints, and a diagram with three edges can carry at most three of them. This paper presents a notation, Card(R; p; q) = (lower, upper), that removes the ambiguity without taking a side in it, and that is not limited to one constraint per edge. We give the number of constraints an n-ary relationship set admits, show which of them a design determines without ever writing them down, and state two inference rules, decomposition and augmentation, that derive one constraint from another, together with the side conditions under which each is sound. We close with a worked case study that turns three business requirements into three explicit constraints and nine more that the design decides on its own.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Huanyi Chen. 2026-10-01. Look Here or Look Across: Unified Cardinality Constraints for N-ary Relationships. https://arxiv.org/abs/2610.02603

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

MaDI-Bench: An End-to-End Data Integration Benchmark

Data integration is the process of combining data from multiple, heterogeneous sources into a consistent, unified representation. Data integration involves a sequence of interdependent tasks including schema matching, value normalization, blocking, entity matching, and data fusion. Existing table-based benchmarks either evaluate these steps in isolation or cover only incomplete versions of the data integration pipeline, omitting specific steps. The lack of public end-to-end data integration benchmarks hinders research on data integration methods that address the integration process as a whole and account for the interdependencies among the different tasks. This paper fills this gap by introducing the Mannheim Data Integration Benchmark (MaDI-Bench), the first benchmark for the end-to-end integration of relational tables covering all steps of the integration process. MaDI-Bench contributes (i) a set of end-to-end data integration tasks spanning several application domains, each requiring the full schema matching, value normalization, entity matching, and data fusion pipeline, and (ii) a generic method for deriving task variants that mitigates rapid benchmark saturation as data integration systems advance. We validate the benchmark using human-engineered pipelines, a best-of-breed pipeline, an LLM workflow, and a pipeline written by a coding agent. The validation demonstrates the utility of the benchmark for measuring the step-wise as well as the end-to-end performance of data integration pipelines. All benchmark artifacts are available for public download.

cs.DB↗

Generalized DBLog: A Verified Contract for Interleaving Copied Rows with a Change Log

Change-data capture (CDC) feeds downstream systems like caches, search indexes, and data warehouses from a database's log of committed row changes. When bootstrapping, adding a table, or repairing downstream data, a pipeline must also copy existing rows. Merging this copy with the active log introduces the copy-to-log handoff problem. Changes must not fall through a gap, and older copied state must not overwrite a newer logged update or resurrect a deleted row. DBLog, developed at Netflix, addressed this problem by reading tables in chunks and interleaving those reads with the live log. Watermarks identify the changes that overlap each read, and the log wins when a copied row is stale. Debezium and Flink CDC have since adapted this design. Earlier work proved that applying the original algorithm's copied rows and logged changes in their emitted order reconstructs the source's rows, including the effect of every logged insert, update, and delete processed. Generalized DBLog asks when the same result holds for variants of that design. We state the conditions the source and capture implementation must satisfy. Once copying and reconciliation are complete, we prove that the result holds across all selected tables and key ranges even when their rows were read at different times. A single database snapshot is not required for the copy. Further logged changes advance the reconstructed state one event at a time. We establish these guarantees for classic watermarking, Debezium's signal-table and read-only modes, Flink CDC's parallel chunks, reads and dumps tied to exact log positions, and engine-consistent backups whose log position lies within known bounds. The complete theory is machine-checked in Isabelle/HOL, its core independently verified in Lean 4, and the protocols are also examined by bounded model checking in TLA+.

cs.DB↗

Protocol-Sensitive Evaluation of Log Anomaly Detection: Component Costs and Target-Access Sensitivity on HDFS and BGL

Protocol choices can change the conclusions drawn from log anomaly detection benchmarks even when detector settings are fixed. We present a joint empirical study of split construction, representation visibility, and component costs using six fixed count, sequence, and semantic configurations on Hadoop Distributed File System (HDFS) and Blue Gene/L (BGL) logs. Random splits place several configurations near the average-precision ceiling, whereas group-disjoint HDFS and chronological BGL evaluation produce lower scores and different observed orderings. At a fixed BGL cutoff, parser choice spans 0.124 in semantic XGBoost mean average precision while preserving its lead over count XGBoost; the earliest rolling period reverses that ordering. A two-factor cross-system ablation contrasts source-only representations with offline transductive access to unlabeled target templates through the representation corpus and inverse document frequency: HDFS-to-BGL mean average precision moves from 0.191 with source-only access to 0.325 with union-corpus, target-IDF access, and the intermediate conditions reveal direction-dependent interactions in average precision and retrieval at fixed review budgets. Component-level profiling separates parsing and representation costs from classifier training, prediction, and storage. Together, these findings connect detector comparisons to the test population, preprocessing state, visible information, and measured pipeline stages, and identify the protocol fields needed alongside a score to support interpretable comparisons of log anomaly detection accuracy and resource use.

cs.DB↗