Missing values are a common issue in gene expression data and can significantly affect downstream analysis. This study aims to compare the performance of a hybrid sequential biclustering and centroid-based clustering method with conventional imputation approaches for handling missing values. The proposed method integrates sequential biclustering based on mean squared residue to identify coherent submatrices, followed by centroid-based clustering to estimate missing entries. The dataset used in this study consists of gene expression data of patients with type 2 diabetes mellitus, with missing values introduced under various proportions ranging from 5% to 55%. The performance of the proposed method is evaluated and compared with mean imputation and nearest neighbor imputation using mean squared error, root mean squared error, and mean absolute error. The experimental results show that the proposed method consistently produces lower errors across all missing rates than the baseline methods. This indicates that incorporating local pattern structures through biclustering improves the accuracy of missing value estimation. The findings suggest that the proposed hybrid framework is more effective in preserving the underlying structure of gene expression data and provides a reliable approach for handling missing data in high-dimensional biological datasets.
Copyrights © 2026