Modeling the probability distribution of rows in tabular data and generating realistic synthetic data is a non-trivial task. Tabular data usually contains a mix of discrete and continuous columns. Continuous columns may have multiple modes whereas discrete columns are sometimes imbalanced making the modeling difficult. Existing statistical and deep neural network models fail to properly model this type of data. We design CTGAN, which uses a conditional generative adversarial network to address these challenges. To aid in a fair and thorough comparison, we design a benchmark with 7 simulated and 8 real datasets and several Bayesian network baselines. CTGAN outperforms Bayesian methods on most of the real datasets whereas other deep learning methods could not.
Lei Xu (MIT)
Maria Skoularidou (University of Cambridge)
I hold a 4-year Bachelor of Science in Informatics and 2-year Master of Science in Statistical Science both from Athens University of Economics and Business and now I am a second year PhD student at MRC-BSU, University of Cambridge. During my undergraduate studies I was delighted to explore the essentials of information theory and theoretical computer science (complexity, computability, asymptotic theory, algorithmic game theory). Later, as postgraduate student I focused on Bayesian theory and applications, under the insightful supervision of Professor Petros Dellaportas. My fields of interest lie in Bayesian high-dimensional problems, mixture models and probabilistic machine learning.
Alfredo Cuesta-Infante (Universidad Rey Juan Carlos)
Kalyan Veeramachaneni (Massachusetts Institute of Technology)
More from the Same Authors
2022 : Invited Talk #2, Kalyan Veeramachaneni, SDMetrics: Evaluating Synthetic Data »
2020 Workshop: Resistance AI Workshop »
Suzanne Kite · Mattie Tesfaldet · J Khadijah Abdurahman · William Agnew · Elliot Creager · Agata Foryciarz · Raphael Gontijo Lopes · Pratyusha Kalluri · Marie-Therese Png · Manuel Sabin · Maria Skoularidou · Ramon Vilarino · Rose Wang · Sayash Kapoor · Micah Carroll
2013 Workshop: Data Driven Education »
Jonathan Huang · Sumit Basu · Kalyan Veeramachaneni