Siti Zulaikha Mohd Jamaludin
School of Mathematical Sciences, Universiti Sains Malaysia, Malaysia

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

INTEGRATING STRUCTURED AND UNSTRUCTURED FEATURES FOR E-TICKETING CLASSIFICATION: A MACHINE LEARNING AND ENSEMBLE-BASED APPROACH Siti Zulaikha Mohd Jamaludin; Majid Khan Majahar Ali; Eric Wong Vun Shiung; Mohd Tahir Ismail; Noor Farizah Ibrahim; Nur Ezlin Zamri
BAREKENG: Jurnal Ilmu Matematika dan Terapan Vol 20 No 4 (2026): BAREKENG: Journal of Mathematics and Its Application
Publisher : PATTIMURA UNIVERSITY

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30598/barekengvol20iss4pp2839-2850

Abstract

E-ticketing systems (ETS) generate mixed information from categorical ticket attributes and free-text descriptions, yet many classification studies still treat these sources separately, which can limit routing accuracy. This study aims (i) to develop an effective preprocessing pipeline for hybrid data, (ii) to develop a two-stage feature selection (2-FS) pipeline for hybrid data, (iii) to design an ensemble classification framework that improves hybrid data performance, and (iv) to benchmark all implemented and proposed models using real-world ETS data. The methodology builds a unified preprocessing workflow by combining Natural Language Processing for textual feature with categorical encoding for categorical features, followed by feature concatenation and vector-space integration to form hybrid representations. Five baseline classifiers (LR, SVM, MNB, RF, and KNN) are evaluated and extended with majority-vote ensembles that pair LR with other classifiers (LR-S, LR-M, LR-R, LR-K, and LR-ALL). Model performance is assessed using accuracy, precision, recall, and F1-score, and the Friedman test is applied to examine statistical consistency across datasets. Results show that hybrid data consistently outperforms single-type datasets, while categorical-only data yields the lowest scores. The best hybrid ensemble of LR-K achieves up to 93% classification accuracy, with strong overall performance also observed for RF on hybrid data. The Friedman test indicates minimal rank differences across models, suggesting that several classifiers are competitive under this ETS setting. Limitations include possible noise and overfitting effects that reduce separation between dataset types. Future work will explore feature ranking-based selection, data-driven segmentation during preprocessing, multi-level classification, imbalance handling, and deeper ensemble strategies to strengthen robustness and generalization.