Generating High-Diversity Synthetic Tabular Data via Less-Constrained Prior
Sanghun Park, Jaesung Lim, Jong-June Jeon, Seunghwan An
Abstract
Generating high-quality synthetic tabular data, regarding fidelity, diversity, and utility, is crucial for many practical purposes. Recent tabular data synthesis methods, including two-stage generative modeling approaches, have achieved this with nearly perfect fidelity. However, we observe that there remains room for improving diversity, and we show that this limitation arises from an additional constraint imposed on the latent support. Our main contribution is that we effectively eliminate this redundant constraint by directly deriving the objective function from the KL-divergence between the ground-truth density and the generative model used for synthetic sample generation. We empirically demonstrate that our model, which relies on a single prior distribution, significantly improves the quality of the synthetic data, especially in terms of diversity. Our implementation code is available at https://anonymous.4open.science/r/SPT-1EE2/.
BibTeX
@inproceedings{ijcai2026_generatinghighdi,
title = {Generating High-Diversity Synthetic Tabular Data via Less-Constrained Prior},
author = {Sanghun Park and Jaesung Lim and Jong-June Jeon and Seunghwan An},
booktitle = {IJCAI 2026},
year = {2026}
}