A new fuzzy-Bayesian multiple imputation approach for missing data
Applied Soft Computing, vol.202, 2026 (SCI-Expanded, Scopus)
- Publication Type: Article / Article
- Volume: 202
- Publication Date: 2026
- Doi Number: 10.1016/j.asoc.2026.115833
- Journal Name: Applied Soft Computing
- Journal Indexes: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Applied Science & Technology Source, Compendex, INSPEC
- Keywords: Fuzzy numbers, JAGS, Linear regression, Missing values, Model uncertainty, Nonlinear relationships
- Hacettepe University Affiliated: Yes
Abstract
Missing values pose a fundamental challenge and must be imputed accurately in data-driven applications. Imputation and model uncertainties, as well as the impact of data characteristics such as non-linear relationships, skewness, and outliers, must be addressed through multiple imputation. Mainly, we use statistical, machine learning, and deep learning methods for this purpose. Available methods struggle to simultaneously handle the multiple imputation challenges. Statistical methods often rely on strong distributional assumptions or struggle with data challenges. Although machine learning and deep learning methods improve upon the imputation uncertainty, they remain sensitive to data irregularities and do not quantify model uncertainty. To address these limitations, we introduce the fuzzy-Bayesian multiple imputation (FBMI) framework, which integrates Bayesian modeling with fuzzy operations, employing Gaussian fuzzy numbers. FBMI handles imputation and model uncertainty while capturing challenging data characteristics. Moreover, we conduct an extensive benchmarking study against ten widely used statistical, machine learning, and deep learning imputation methods across twenty five datasets, under varying levels of outlier contamination and distributional characteristics. The results demonstrate that while FBMI consistently outperforms benchmark methods across different missingness settings, improving imputation accuracy on average by between 26% and 68% across multiple performance indicators, it is also computationally efficient and supported by openly available software.