arXiv · 2609.34349
Fuzzy Distribution Modeling for Synthetic Tabular Data Generation with Causality Preservation
Abstract
Synthetic tabular data generation provides an effective alternative for the training of machine learning models when real-world data is limited or inaccessible. However, the heterogeneous, non-smooth, and incomplete nature of tabular data poses fundamental challenges to conventional probabilistic and deep generative models, where their interpretability remains limited. This paper proposes a novel fuzzy distribution modeling methodology for synthetic tabular data generation based on fuzzy sets theory. Feature distributions are represented using fuzzy sets and feature dependencies are modeled through Fuzzy Cognitive Maps, resulting in a low-parameter, and an interpretable data representation. Synthetic samples are generated by sampling fuzzy concepts rather than raw values, enabling native support for mixed data types, missing values, and domain constraints. The methodology further supports linguistic queries and IF-THEN reasoning, facilitating transparent simulation of decision-making processes. Experimental results on benchmark datasets demonstrate competitive performance with respect to utility, fidelity and privacy compared to state-of-the-art methods, while offering substantially improved interpretability. These results establish fuzzy distribution modeling as a principled and effective approach for synthetic tabular data generation in fuzzy systems and decision support applications.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Michael Vasilakakis, Dimitris K. Iakovidis. 2026-09-28. Fuzzy Distribution Modeling for Synthetic Tabular Data Generation with Causality Preservation. https://doi.org/10.1109/fuzz69877.2026.11626474
Cite the original work for its findings. Save a collection to share your selection of sources.