Background: The increasing complexity of Indonesian educational texts across print and digital platforms raises concerns about mismatches with students’ reading capacities, while existing readability formulas remain largely English-centric. Objective: This study develops and evaluates an automatic readability assessment framework for Indonesian texts integrating surface metrics, linguistic features, and machine learning. Method: A stratified corpus of textbooks and digital materials was analyzed using readability indices, lexical–morphological–syntactic features, and supervised models with cross-validation. Results: Surface complexity rises across levels but shows overlap, indicating limits of traditional metrics; linguistic features such as lexical density, nominalization, and morphological complexity strongly predict readability, while hybrid ensemble models achieve the highest accuracy and lowest misclassification. Implication: These findings support the need for language-specific, multidimensional readability tools to improve text design and educational alignment in Indonesian contexts. Novelty: This study proposes a linguistically grounded hybrid NLP framework that reconceptualizes readability and enables scalable, automated evaluation of Indonesian educational texts.
Copyrights © 2026