← Search

Seunghwan An

5 accepted papers

2026

TabularBERT: Binning-Based Self-Supervised Learning for Tabular Representation

ICML 2026poster

Tabular data is one of the most fundamental and widely used formats for representing structured information. Classical machine learning algorithms continue to achieve substantial success in extracting predictive patterns and constructing accurate models from structured data; however, representation …

Cited by 0SourceScholar
2025

Masked Language Modeling Becomes Conditional Density Estimation for Tabular Data Synthesis

AAAI 2025technical

In this paper, our goal is to generate synthetic data for heterogeneous (mixed-type) tabular datasets with high machine learning utility (MLu). Since the MLu performance depends on accurately approximating the conditional distributions, we focus on devising a synthetic data generation method based o…

2023

Distributional Learning of Variational AutoEncoder: Application to Synthetic Data Generation

NeurIPS 2023poster

The Gaussianity assumption has been consistently criticized as a main limitation of the Variational Autoencoder (VAE) despite its efficiency in computational modeling. In this paper, we propose a new approach that expands the model capacity (i.e., expressive power of distributional family) without s…