Indonesian Journal of Statistics and Its Applications
Vol 10 No 1 (2026): Vol 10 Issue 1 June 2026

A Comparative Study of Sequential Biclustering and Fuzzy C-Means, K-Nearest Neighbors, and Mean Imputation for Missing Value Estimation in Gene Expression Data

Yolanda Azzahra (Department of Mathematics, Faculty of Mathematics and Natural Science, University of Indonesia, Indonesia)
Titin Siswantining (Department of Mathematics, University of Indonesia, Indonesia)
Setia Pramana (Department of Statistics, Politeknik Statistika STIS, Indonesia)
Mogana Darshini Ganggayah (Department of Econometrics and Business Statistics, School of Business, Monash University, Malaysia)
Alhadi Bustamam (Department of Mathematics, University of Indonesia, Indonesia)



Article Info

Publish Date
30 Jun 2026

Abstract

Missing values are a common issue in gene expression data and can significantly affect downstream analysis. This study aims to compare the performance of a hybrid sequential biclustering and centroid-based clustering method with conventional imputation approaches for handling missing values. The proposed method integrates sequential biclustering based on mean squared residue to identify coherent submatrices, followed by centroid-based clustering to estimate missing entries. The dataset used in this study consists of gene expression data of patients with type 2 diabetes mellitus, with missing values introduced under various proportions ranging from 5% to 55%. The performance of the proposed method is evaluated and compared with mean imputation and nearest neighbor imputation using mean squared error, root mean squared error, and mean absolute error. The experimental results show that the proposed method consistently produces lower errors across all missing rates than the baseline methods. This indicates that incorporating local pattern structures through biclustering improves the accuracy of missing value estimation. The findings suggest that the proposed hybrid framework is more effective in preserving the underlying structure of gene expression data and provides a reliable approach for handling missing data in high-dimensional biological datasets.

Copyrights © 2026






Journal Info

Abbrev

ijsa

Publisher

Subject

Computer Science & IT Mathematics Other

Description

Indonesian Journal of Statistics and Its Applications (eISSN:2599-0802) (formerly named Forum Statistika dan Komputasi), established since 2017, publishes scientific papers in the area of statistical science and the applications. The published papers should be research papers with, but not limited ...