Search arXiv⌕ Search

arXiv · 2609.36512

SAIVE: Selecting AI Valuable Entities

Abstract

Data lakes store large amounts of telemetry, with logs from network sensors, hosts, and applications containing possibly hundreds of fields for every event. Large enterprises are then left with data lakes that cannot be analyzed efficiently with AI. Aggregate analysis looks at persistent shifts in behavior over time. Many of the fields and columns in data lakes are not useful as they do not contain information that is sufficiently diverse or concentrated to support AI analysis. SAIVE is a simple method for examining a few rows in a large table and applies a histogram of histograms filtering criterion to select the fields that for AI analysis is more likely to yield useful results. This paper provides a principled foundation for the SAIVE heuristics by assuming of a Zipf-Mandelbrot power-law distribution of the underlying data. Constraining the Zipf-Mandelbrot exponent alpha to a reasonable range provides a a practical, cheap, expert-free filter for selecting AI valuable entities in large data sets.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Inna Voloshchuk, Hayden Jananthan, Jeremy Kepner. 2026-09-29. SAIVE: Selecting AI Valuable Entities. https://arxiv.org/abs/2609.36512

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

X-DigCheck: Co-Evolving Application Profiles and Knowledge Graphs, Demonstrated on the RTI Documentation of Rupe Magna

We demonstrate X-DigCheck, a domain-independent environment for building and maintaining application profiles as they co-evolve with the data they describe. Profiles developed against a fixed ontology quickly drift from the schema they were meant to capture. X-DigCheck treats profile construction as a continuous ontology-data co-evolution loop: data are lifted into RDF against the profile, checked through competency questions and SHACL, and the resulting reports jointly drive revisions of the ontology, mappings, constraints, and graph. The loop is agnostic to the domain and to the pipeline that produces the graph. We validate and demonstrate the tool in the cultural heritage domain, on the construction of RupeMagna-RTI, the first Reflectance Transformation Imaging (RTI) specialisation of the Cultural Heritage Survey ODP (CHS-ODP), aligned with CIDOC-CRM/CRMdig, ArCo, CHAD-KG, and Getty AAT, with semRTI as the lifting pipeline of this use case. The demonstration lets visitors run one full turn of the loop -on the shipped Rupe Magna (Grosio, Italy) RTI survey, or on a profile and graph of their own -executing the competency-question and SHACL checks live and reading the bidirectional coverage report that flags modelling gaps and stale assumptions. The result is a portable co-evolution environment for profile engineering, together with a reusable RTI application profile produced through it. A screencast of the demonstration is available at https://zenodo.org/records/22210609.

cs.DB↗

Transformations for Evolving Property Graph Schemas

Property graph databases are widely used to represent complex and evolving data; yet, systematic support for property graph schema evolution remains limited. In practice, schema transformations are typically defined manually, coupled to specific application contexts, and are difficult to reuse across schemas or evolution scenarios. We present GRAFT, a logic-based framework that models prop- erty graph schema evolution as reusable, order-constrained meta- transformations derived from atomic edits. Schema evolution is formulated as exploration of a finite meta-graph with schemas as nodes and grounded meta-transformations as edges. To ensure tractability, GRAFT combines similarity-guided search and pruning, guaranteeing duplication-freeness, termination and correctness. An experimental evaluation on four benchmark and real-world property graph schema evolution scenarios shows that GRAFT effi- ciently computes high-quality schema transformation sequences. Using greedy exploration, GRAFT reaches the exact target schema on most datasets, producing stable transformation sequences while keeping runtimes low. A qualitative study on both real-world and a synthetic large-scale dataset further shows the quality and robust- ness of the obtained reusable meta-transformations.

cs.DB↗

Epistemic Typing as a PostgreSQL Table Access Method: Adversarial Conflict Resolution Under Confidence Forgery and Sybil Coordination

We describe KNDB, a PostgreSQL 18 table access method (TAM) that types every row with an engine-assigned epistemic kind (MEASURED, INFERRED, or DERIVED) and resolves per-slot conflicts inside every write-time heapam callback. Rows land as ordinary heap tuples; seven of the 44 TAM callbacks are overridden (tuple_insert, multi_insert, tuple_update, tuple_delete, tuple_insert_speculative, tuple_complete_speculative, relation_toast_am), the other 37 delegate to heap; we provide a completeness argument over the interface as a paper artefact. This paper reports the engineering behind that decision and the adversarial evaluation that motivated it. On a confidence-forgery workload where an attacker asserts INFERRED writes with confidence in [0.95,1.0] against honest MEASURED writes with confidence in [0.5,0.9], KNDB beats a confidence-only baseline by 63 percentage points on the Book-Author fusion dataset and 92.7 points on the Zheng crowdsourcing dataset. Both wins are proven load-bearing on the kind axis by a source-rebuild disable-and-test in which the lattice is neutralised and the win vanishes. Against four truth-discovery baselines (TruthFinder, CRH, CATD, ACCU) reimplemented from the original equations and validated to within 0.3 percentage points of the published numbers, KNDB is competitive below a per-dataset density-saturation cell and dominant at or above it. We formalise the cell as k* ~ rho_alg * h_top, where h_top is per-slot top honest surface-form support, and validate the prediction within +/-20% on Book-Author and +/-30% on Zheng. Because the kind axis is assigned by the engine from independent metadata and cannot be forged at write time, KNDB's k* is unbounded. The paper is honest about where KNDB loses: CRH and ACCU outperform KNDB below saturation on Zheng, and KNDB scores zero on three temporal knowledge-editing benchmarks whose ground truth is last-writer-wins.

cs.DB↗