Background: Data-driven requirement elicitation has been increasingly used in modern software engineering due to the growing availability of online user-generated textual data. However, existing approaches mostly rely on single-source data, which are limited in handling the diversity of characteristics found in online textual sources. Objective: This study proposes an automated natural language processing (NLP)-based framework for requirement elicitation that integrates heterogeneous online sources, such as app reviews, online news, and tweets, for process innovation in the requirements engineering phase. Methods: The proposed framework combines rule-based and AI-based extraction methods, semantic clustering, and diagram generation. Data were collected from application reviews, Twitter/X, and online news across six domains. The framework was evaluated using expert-annotated ground truth to measure extraction performance and expert-based assessments to examine clustering quality and artifact usefulness. Result: AI-based extraction outperforms rule-based methods for requirements extraction, achieving F1 scores of 0.92, 0.80, and 0.67 on app reviews, 0.80 on Twitter, and 0.67 on online news. In the expert evaluation, the proposed system demonstrates high topic coherence, reduces elicitation time, and helps identify potential system requirements that may be overlooked in manual processes. Conclusion: Multisource integration enhances the completeness and contextual richness of automated requirement elicitation. The proposed framework effectively transforms heterogeneous textual data into actionable requirement artifacts, providing a scalable and practical solution for early-stage software development. Keywords: Requirement Elicitation, Natural Language Processing, Process Innovation, Multisource Data
Copyrights © 2026