DISCRETIZING BINARY VARIABLES FROM CONTINUOUS VARIABLES FOR SIMULATION DATA GENERATION: A CRITICAL LITERATURE REVIEW
DOI:
https://doi.org/10.35631/JISTM.1144033Keywords:
Continuous-To-Binary, Discretization, Mixed-Variable Simulation, ReviewAbstract
Generating simulation datasets with both continuous and categorical variables is essential for evaluation of statistical methods involving mixed variables. Continuous variables can be generated directly from specified distribution, however, generating categorical variables requires procedures, as they cannot be generated directly in the same manner as continuous variables. Several reviews have addressed general data discretization or categorical encoding as procedures for generating categorical variables, but there are limited studies synthesised continuous-to-binary discretization methods for simulation data generation. Therefore, this review aims to synthesise the literature on continuous-to-binary discretization methods for simulation data generation, comparing their strengths, limitations, and suitability across simulation contexts. This literature review was searching using Scopus AI with no restriction on publication year, in which the searching completed in August 2026 yielded 18 peer-reviewed journal articles retained as representative of the major methodological approaches identified. Four methodological themes were identified which are threshold-based discretization, latent-variable and probabilistic modelling, dependence-preserving simulation frameworks, and considerations relating to threshold selection, information loss, and model compatibility. Threshold-based methods are simple and efficient but prone to information loss and biased estimation when thresholds are poorly justified, whereas latent-variable, Gaussian copula, and power-polynomial approaches better preserve dependence structures at the cost of stronger assumptions and greater computational complexity. This review offers a useful methodological reference for selecting discretization methods suited to specific simulation objectives, while identifying limitations and priorities for future research.
Downloads
References
Agresti, A. (2007). An introduction to categorical data analysis (2nd ed.). John Wiley & Sons. https://doi.org/10.1002/0470114754
Ahlemeyer-Stubbe, A., & Müller, A. (2021). The importance of domain knowledge for successful and robust predictive modelling. Applied Marketing Analytics, 6(4), 344–352.
Alazaidah, R. (2023). A comparative analysis of discretization techniques in machine learning. In 2023 24th International Arab Conference on Information Technology (ACIT). IEEE. https://doi.org/10.1109/ACIT58888.2023.10453749
Amatya, A., & Demirtas, H. (2016). Concurrent generation of multivariate mixed data with variables of dissimilar types. Journal of Statistical Computation and Simulation, 86(18), 3595–3607. https://doi.org/10.1080/00949655.2016.1177530
Demirtas, H. (2017). Concurrent generation of binary and nonnormal continuous data through fifth-order power polynomials. Communications in Statistics: Simulation and Computation, 46(1), 344–357. https://doi.org/10.1080/03610918.2014.963613
Demirtas, H., & Hedeker, D. (2011). A practical way for computing approximate lower and upper correlation bounds. The American Statistician, 65(2), 104–109. https://doi.org/10.1198/tast.2011.10090
Demirtas, H., Hedeker, D., & Mermelstein, R. J. (2012). Simulation of massive public health data by power polynomials. Statistics in Medicine, 31(27), 3337–3346. https://doi.org/10.1002/sim.5362
DeYoreo, M., & Kottas, A. (2015). A fully nonparametric modeling approach to binary regression. Bayesian Analysis, 10(4), 821–847. https://doi.org/10.1214/15-BA963SI
Emrich, L. J., & Piedmonte, M. R. (1991). A method for generating high-dimensional multivariate binary variates. The American Statistician, 45(4), 302–304. https://doi.org/10.1080/00031305.1991.10475828
Fialkowski, A., & Tiwari, H. (2019). SimCorrMix: Simulation of correlated data with multiple variable types including continuous and count mixture distributions. The R Journal, 11(1), 248–264. https://doi.org/10.32614/RJ-2019-022
Fleishman, A. I. (1978). A method for simulating non-normal distributions. Psychometrika, 43(4), 521–532. https://doi.org/10.1007/BF02293811
Grobler, A. C., & Lee, K. (2020). Multiple imputation in the presence of an incomplete binary variable created from an underlying continuous variable. Biometrical Journal, 62(2), 467–478. https://doi.org/10.1002/bimj.201900011
Hadjicostas, P. (2006). Maximizing proportions of correct classifications in binary logistic regression. Journal of Applied Statistics, 33(6), 629–640. https://doi.org/10.1080/02664760600723367
Higham, N. J. (2002). Computing the nearest correlation matrix—A problem from finance. IMA Journal of Numerical Analysis, 22(3), 329–343. https://doi.org/10.1093/imanum/22.3.329
Jiryaie, F., Withanage, N., Wu, B., & de Leon, A. R. (2016). Gaussian copula distributions for mixed data, with application in discrimination. Journal of Statistical Computation and Simulation, 86(9), 1643–1659. https://doi.org/10.1080/00949655.2015.1077386
Liquet, B., & Riou, J. (2019). CPMCGLM: An R package for p-value adjustment when looking for an optimal transformation of a single explanatory variable in generalized linear models. BMC Medical Research Methodology, 19, Article 79. https://doi.org/10.1186/s12874-019-0711-2
Lustgarten, J. L., Gopalakrishnan, V., Grover, H., & Visweswaran, S. (2008). Improving classification performance with discretization on biomedical datasets. AMIA Annual Symposium Proceedings, 445–449.
MacCallum, R. C., Zhang, S., Preacher, K. J., & Rucker, D. D. (2002). On the practice of dichotomization of quantitative variables. Psychological Methods, 7(1), 19–40. https://doi.org/10.1037/1082-989X.7.1.19
Morlini, I. (2011). Mixed mode data clustering: An approach based on tetrachoric correlations. In F. Palumbo et al. (Eds.), Studies in Classification, Data Analysis, and Knowledge Organization (pp. 95–103). Springer. https://doi.org/10.1007/978-3-642-13312-1_9
Morlini, I. (2012). A latent variables approach for clustering mixed binary and continuous variables within a Gaussian mixture model. Advances in Data Analysis and Classification, 6(1), 5–28. https://doi.org/10.1007/s11634-011-0101-z
Nelson, S. L. P., Ramakrishnan, V., Nietert, P. J., Wolf, B. J., & Taylor, M. H. (2017). An evaluation of common methods for dichotomization of continuous variables to discriminate disease status. Communications in Statistics—Theory and Methods, 46(21), 10823–10834. https://doi.org/10.1080/03610926.2016.1202274
Nojavan, F. A., Qian, S. S., & Stow, C. A. (2017). Comparative analysis of discretization methods in Bayesian networks. Environmental Modelling & Software, 87, 64–71. https://doi.org/10.1016/j.envsoft.2016.10.007
Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., Shamseer, L., Tetzlaff, J. M., Akl, E. A., Brennan, S. E., Chou, R., Glanville, J., Grimshaw, J. M., Hróbjartsson, A., Lalu, M. M., Li, T., Loder, E. W., Mayo-Wilson, E., McDonald, S., ... Moher, D. (2021). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ, 372, n71. https://doi.org/10.1136/bmj.n71
Royston, P., Altman, D. G., & Sauerbrei, W. (2006). Dichotomizing continuous predictors in multiple regression: A bad idea. Statistics in Medicine, 25(1), 127–141. https://doi.org/10.1002/sim.2331
Rucker, D. D., McShane, B. B., & Preacher, K. J. (2015). A researcher's guide to regression, discretization, and median splits of continuous variables. Journal of Consumer Psychology, 25(4), 666–678. https://doi.org/10.1016/j.jcps.2015.04.004
Wang, J.-H., & Liu, B. (2005). Method of the discretization of continuous probability distribution for risk analysis. Journal of Xi'an Shiyou University (Natural Sciences Edition), 20(2), 83–85.
Zhao, J., Liu, X., Du, B., & Liu, Y. (2025). Approximation error from discretizations and its applications. The Annals of Statistics, 53(2), 589–614. https://doi.org/10.1214/24-AOS2470
