Search arXiv⌕ Search

arXiv · 2204.04164

Uncertain Case Identifiers in Process Mining: A User Study of the Event-Case Correlation Problem on Click Data

Abstract

Among the many sources of event data available today, a prominent one is user interaction data. User activity may be recorded during the use of an application or website, resulting in a type of user interaction data often called click data. An obstacle to the analysis of click data using process mining is the lack of a case identifier in the data. In this paper, we show a case and user study for event-case correlation on click data, in the context of user interaction events from a mobility sharing company. To reconstruct the case notion of the process, we apply a novel method to aggregate user interaction data in separate user sessions-interpreted as cases-based on neural networks. To validate our findings, we qualitatively discuss the impact of process mining analyses on the resulting well-formed event log through interviews with process experts.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Marco Pegoraro, Merih Seran Uysal, Tom-Hendrik Hülsmann, Wil M. P. van der Aalst. 2022-04-08. Uncertain Case Identifiers in Process Mining: A User Study of the Event-Case Correlation Problem on Click Data. https://arxiv.org/abs/2204.04164

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

IBLTs Measure Before They Decode: Self-Sizing Set Reconciliation for Database Consistency Verification

Cross-system data replication pipelines cannot confirm end-to-end consistency from the local guarantees of each hop, so the two endpoints must be compared directly on a periodic basis. Once the rows of a fixed snapshot are normalized into fingerprints, the task reduces to finding the symmetric difference of the two sets. An Invertible Bloom Lookup Table (IBLT) reconciles the sets with communication that grows only with the difference cardinality $d$, independent of table size, but its capacity must be fixed while $d$ is still unknown. Across 41,603 production reconciliations over 90 days, nonzero $d$ spans about seven orders of magnitude, and no reliable empirical constant exists. We show that the count array of an IBLT has already measured $d$ before decoding. The measurement is in-band: it is carried by the recovery sketch itself and adds no bytes dedicated to estimation. A mapping-aware theorem extends the construction to Irregular, Rateless, and MET IBLTs. The protocol reads the estimate only after a decoding failure; we prove that the failure-conditioned lower quantile bounds the risk of underestimation, which gives the second-round capacity a configurable success-probability guarantee. The resulting self-sizing protocol attempts recovery with a small first-round sketch and stops on success; on failure it reads $d$, sizes the second round, and completes reconciliation in at most two rounds. Against a controlled oracle, communication is 1.29-1.47 times that of a scheme given $d$ in advance. Production workload characterization, relational-database replay, and a cross-city KV deployment confirm the end-to-end mechanism. In production on an Oracle-MySQL link, all completed runs succeeded within two rounds, over 90% on the 1-RTT fast path with a single 16 KB sketch.

cs.DB↗

Transformations for Evolving Property Graph Schemas

Property graph databases are widely used to represent complex and evolving data, yet systematic support for property graph schema evolution remains limited. In practice, schema transformations are typically defined manually, coupled to specific application contexts, and difficult to reuse across schemas or evolution scenarios. We present GRAFT, a logic-based framework that models property graph schema evolution as reusable, order-constrained meta-transformations derived from atomic edits. Schema evolution is formulated as the exploration of a finite meta-graph with schemas as nodes and grounded meta-transformations as edges. To ensure tractability, GRAFT combines similarity-guided search and pruning, guaranteeing duplication-freeness, termination, and correctness. An experimental evaluation on four benchmark and real-world property graph schema evolution scenarios shows that GRAFT efficiently computes high-quality schema transformation sequences. Using greedy exploration, GRAFT reaches the exact target schema on most datasets, producing stable transformation sequences while keeping runtimes low. A qualitative study on both real-world and synthetic large-scale datasets further shows the quality and robustness of the obtained reusable meta-transformations.

cs.DB↗

VADER: Filtered Vector Search with Declarative Recall

Approximate filtered vector search (FVS), a core operation in many data management tasks that combine structured data with vector embeddings, exhibits increased complexity due to the characteristics of filtering predicates. Each predicate is defined by selectivity (i.e., the fraction of vectors that satisfy the predicate) and correlation (i.e., the relationship between the filter and the vector space), which can significantly affect search difficulty even for the same query vector. This poses a key challenge for users aiming to integrate vector search with structured data, as efficient execution often requires extensive manual tuning of algorithm parameters. In this paper, we present VADER, the first approach that eliminates hyperparameter tuning by introducing declarative recall for approximate filtered vector search. With declarative recall, users specify a desired recall target, and VADER executes FVS queries to meet this target without requiring manual configuration. VADER achieves this by employing a filter-aware recall predictor that generalizes across varying selectivities and correlations without explicit tuning, and by performing early termination once the predicted recall reaches the user-defined target. Through extensive experimental evaluation, we show that VADER achieves near-optimal early termination, while providing significant speedups of up to 53% faster and improved result quality of 28% compared to the best-performing baseline.

cs.DB↗