A new fuzzy-Bayesian multiple imputation approach for missing data
Applied Soft Computing, cilt.202, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 202
- Basım Tarihi: 2026
- Doi Numarası: 10.1016/j.asoc.2026.115833
- Dergi Adı: Applied Soft Computing
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Applied Science & Technology Source, Compendex, INSPEC
- Anahtar Kelimeler: Fuzzy numbers, JAGS, Linear regression, Missing values, Model uncertainty, Nonlinear relationships
- Hacettepe Üniversitesi Adresli: Evet
Özet
Missing values pose a fundamental challenge and must be imputed accurately in data-driven applications. Imputation and model uncertainties, as well as the impact of data characteristics such as non-linear relationships, skewness, and outliers, must be addressed through multiple imputation. Mainly, we use statistical, machine learning, and deep learning methods for this purpose. Available methods struggle to simultaneously handle the multiple imputation challenges. Statistical methods often rely on strong distributional assumptions or struggle with data challenges. Although machine learning and deep learning methods improve upon the imputation uncertainty, they remain sensitive to data irregularities and do not quantify model uncertainty. To address these limitations, we introduce the fuzzy-Bayesian multiple imputation (FBMI) framework, which integrates Bayesian modeling with fuzzy operations, employing Gaussian fuzzy numbers. FBMI handles imputation and model uncertainty while capturing challenging data characteristics. Moreover, we conduct an extensive benchmarking study against ten widely used statistical, machine learning, and deep learning imputation methods across twenty five datasets, under varying levels of outlier contamination and distributional characteristics. The results demonstrate that while FBMI consistently outperforms benchmark methods across different missingness settings, improving imputation accuracy on average by between 26% and 68% across multiple performance indicators, it is also computationally efficient and supported by openly available software.