Titin Siswantining
Department of Mathematics, University of Indonesia, Indonesia

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

A Comparative Study of Sequential Biclustering and Fuzzy C-Means, K-Nearest Neighbors, and Mean Imputation for Missing Value Estimation in Gene Expression Data Yolanda Azzahra; Titin Siswantining; Setia Pramana; Mogana Darshini Ganggayah; Alhadi Bustamam
Indonesian Journal of Statistics and Applications Vol 10 No 1 (2026): Vol 10 Issue 1 June 2026
Publisher : Statistics and Data Science Program Study, SSMI, IPB University, in collaboration with the Forum Pendidikan Tinggi Statistika Indonesia (FORSTAT) and the Ikatan Statistisi Indonesia (ISI)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29244/ijsa.v10i1p37-46

Abstract

Missing values are a common issue in gene expression data and can significantly affect downstream analysis. This study aims to compare the performance of a hybrid sequential biclustering and centroid-based clustering method with conventional imputation approaches for handling missing values. The proposed method integrates sequential biclustering based on mean squared residue to identify coherent submatrices, followed by centroid-based clustering to estimate missing entries. The dataset used in this study consists of gene expression data of patients with type 2 diabetes mellitus, with missing values introduced under various proportions ranging from 5% to 55%. The performance of the proposed method is evaluated and compared with mean imputation and nearest neighbor imputation using mean squared error, root mean squared error, and mean absolute error. The experimental results show that the proposed method consistently produces lower errors across all missing rates than the baseline methods. This indicates that incorporating local pattern structures through biclustering improves the accuracy of missing value estimation. The findings suggest that the proposed hybrid framework is more effective in preserving the underlying structure of gene expression data and provides a reliable approach for handling missing data in high-dimensional biological datasets.