DISCRETIZING BINARY VARIABLES FROM CONTINUOUS VARIABLES FOR SIMULATION DATA GENERATION: A CRITICAL LITERATURE REVIEW

Authors

DOI:

https://doi.org/10.35631/JISTM.1144033

Keywords:

Continuous-To-Binary, Discretization, Mixed-Variable Simulation, Review

Abstract

Generating simulation datasets with both continuous and categorical variables is essential for evaluation of statistical methods involving mixed variables. Continuous variables can be generated directly from specified distribution, however, generating categorical variables requires procedures, as they cannot be generated directly in the same manner as continuous variables. Several reviews have addressed general data discretization or categorical encoding as procedures for generating categorical variables, but there are limited studies synthesised continuous-to-binary discretization methods for simulation data generation. Therefore, this review aims to synthesise the literature on continuous-to-binary discretization methods for simulation data generation, comparing their strengths, limitations, and suitability across simulation contexts. This literature review was searching using Scopus AI with no restriction on publication year, in which the searching completed in August 2026 yielded 18 peer-reviewed journal articles retained as representative of the major methodological approaches identified. Four methodological themes were identified which are threshold-based discretization, latent-variable and probabilistic modelling, dependence-preserving simulation frameworks, and considerations relating to threshold selection, information loss, and model compatibility. Threshold-based methods are simple and efficient but prone to information loss and biased estimation when thresholds are poorly justified, whereas latent-variable, Gaussian copula, and power-polynomial approaches better preserve dependence structures at the cost of stronger assumptions and greater computational complexity. This review offers a useful methodological reference for selecting discretization methods suited to specific simulation objectives, while identifying limitations and priorities for future research.

Downloads

Download data is not yet available.

References

Agresti, A. (2007). An introduction to categorical data analysis (2nd ed.). John Wiley & Sons. https://doi.org/10.1002/0470114754

Ahlemeyer-Stubbe, A., & Müller, A. (2021). The importance of domain knowledge for successful and robust predictive modelling. Applied Marketing Analytics, 6(4), 344–352.

Alazaidah, R. (2023). A comparative analysis of discretization techniques in machine learning. In 2023 24th International Arab Conference on Information Technology (ACIT). IEEE. https://doi.org/10.1109/ACIT58888.2023.10453749

Amatya, A., & Demirtas, H. (2016). Concurrent generation of multivariate mixed data with variables of dissimilar types. Journal of Statistical Computation and Simulation, 86(18), 3595–3607. https://doi.org/10.1080/00949655.2016.1177530

Demirtas, H. (2017). Concurrent generation of binary and nonnormal continuous data through fifth-order power polynomials. Communications in Statistics: Simulation and Computation, 46(1), 344–357. https://doi.org/10.1080/03610918.2014.963613

Demirtas, H., & Hedeker, D. (2011). A practical way for computing approximate lower and upper correlation bounds. The American Statistician, 65(2), 104–109. https://doi.org/10.1198/tast.2011.10090

Demirtas, H., Hedeker, D., & Mermelstein, R. J. (2012). Simulation of massive public health data by power polynomials. Statistics in Medicine, 31(27), 3337–3346. https://doi.org/10.1002/sim.5362

DeYoreo, M., & Kottas, A. (2015). A fully nonparametric modeling approach to binary regression. Bayesian Analysis, 10(4), 821–847. https://doi.org/10.1214/15-BA963SI

Emrich, L. J., & Piedmonte, M. R. (1991). A method for generating high-dimensional multivariate binary variates. The American Statistician, 45(4), 302–304. https://doi.org/10.1080/00031305.1991.10475828

Fialkowski, A., & Tiwari, H. (2019). SimCorrMix: Simulation of correlated data with multiple variable types including continuous and count mixture distributions. The R Journal, 11(1), 248–264. https://doi.org/10.32614/RJ-2019-022

Fleishman, A. I. (1978). A method for simulating non-normal distributions. Psychometrika, 43(4), 521–532. https://doi.org/10.1007/BF02293811

Grobler, A. C., & Lee, K. (2020). Multiple imputation in the presence of an incomplete binary variable created from an underlying continuous variable. Biometrical Journal, 62(2), 467–478. https://doi.org/10.1002/bimj.201900011

Hadjicostas, P. (2006). Maximizing proportions of correct classifications in binary logistic regression. Journal of Applied Statistics, 33(6), 629–640. https://doi.org/10.1080/02664760600723367

Higham, N. J. (2002). Computing the nearest correlation matrix—A problem from finance. IMA Journal of Numerical Analysis, 22(3), 329–343. https://doi.org/10.1093/imanum/22.3.329

Jiryaie, F., Withanage, N., Wu, B., & de Leon, A. R. (2016). Gaussian copula distributions for mixed data, with application in discrimination. Journal of Statistical Computation and Simulation, 86(9), 1643–1659. https://doi.org/10.1080/00949655.2015.1077386

Liquet, B., & Riou, J. (2019). CPMCGLM: An R package for p-value adjustment when looking for an optimal transformation of a single explanatory variable in generalized linear models. BMC Medical Research Methodology, 19, Article 79. https://doi.org/10.1186/s12874-019-0711-2

Lustgarten, J. L., Gopalakrishnan, V., Grover, H., & Visweswaran, S. (2008). Improving classification performance with discretization on biomedical datasets. AMIA Annual Symposium Proceedings, 445–449.

MacCallum, R. C., Zhang, S., Preacher, K. J., & Rucker, D. D. (2002). On the practice of dichotomization of quantitative variables. Psychological Methods, 7(1), 19–40. https://doi.org/10.1037/1082-989X.7.1.19

Morlini, I. (2011). Mixed mode data clustering: An approach based on tetrachoric correlations. In F. Palumbo et al. (Eds.), Studies in Classification, Data Analysis, and Knowledge Organization (pp. 95–103). Springer. https://doi.org/10.1007/978-3-642-13312-1_9

Morlini, I. (2012). A latent variables approach for clustering mixed binary and continuous variables within a Gaussian mixture model. Advances in Data Analysis and Classification, 6(1), 5–28. https://doi.org/10.1007/s11634-011-0101-z

Nelson, S. L. P., Ramakrishnan, V., Nietert, P. J., Wolf, B. J., & Taylor, M. H. (2017). An evaluation of common methods for dichotomization of continuous variables to discriminate disease status. Communications in Statistics—Theory and Methods, 46(21), 10823–10834. https://doi.org/10.1080/03610926.2016.1202274

Nojavan, F. A., Qian, S. S., & Stow, C. A. (2017). Comparative analysis of discretization methods in Bayesian networks. Environmental Modelling & Software, 87, 64–71. https://doi.org/10.1016/j.envsoft.2016.10.007

Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., Shamseer, L., Tetzlaff, J. M., Akl, E. A., Brennan, S. E., Chou, R., Glanville, J., Grimshaw, J. M., Hróbjartsson, A., Lalu, M. M., Li, T., Loder, E. W., Mayo-Wilson, E., McDonald, S., ... Moher, D. (2021). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ, 372, n71. https://doi.org/10.1136/bmj.n71

Royston, P., Altman, D. G., & Sauerbrei, W. (2006). Dichotomizing continuous predictors in multiple regression: A bad idea. Statistics in Medicine, 25(1), 127–141. https://doi.org/10.1002/sim.2331

Rucker, D. D., McShane, B. B., & Preacher, K. J. (2015). A researcher's guide to regression, discretization, and median splits of continuous variables. Journal of Consumer Psychology, 25(4), 666–678. https://doi.org/10.1016/j.jcps.2015.04.004

Wang, J.-H., & Liu, B. (2005). Method of the discretization of continuous probability distribution for risk analysis. Journal of Xi'an Shiyou University (Natural Sciences Edition), 20(2), 83–85.

Zhao, J., Liu, X., Du, B., & Liu, Y. (2025). Approximation error from discretizations and its applications. The Annals of Statistics, 53(2), 589–614. https://doi.org/10.1214/24-AOS2470

Downloads

Published

2026-09-28

How to Cite

Kasim, K., Hamid, H., & Abdul-Rahman, A. (2026). DISCRETIZING BINARY VARIABLES FROM CONTINUOUS VARIABLES FOR SIMULATION DATA GENERATION: A CRITICAL LITERATURE REVIEW. JOURNAL INFORMATION AND TECHNOLOGY MANAGEMENT (JISTM), 11(44), 546–562. https://doi.org/10.35631/JISTM.1144033