E-ticketing systems (ETS) generate mixed information from categorical ticket attributes and free-text descriptions, yet many classification studies still treat these sources separately, which can limit routing accuracy. This study aims (i) to develop an effective preprocessing pipeline for hybrid data, (ii) to develop a two-stage feature selection (2-FS) pipeline for hybrid data, (iii) to design an ensemble classification framework that improves hybrid data performance, and (iv) to benchmark all implemented and proposed models using real-world ETS data. The methodology builds a unified preprocessing workflow by combining Natural Language Processing for textual feature with categorical encoding for categorical features, followed by feature concatenation and vector-space integration to form hybrid representations. Five baseline classifiers (LR, SVM, MNB, RF, and KNN) are evaluated and extended with majority-vote ensembles that pair LR with other classifiers (LR-S, LR-M, LR-R, LR-K, and LR-ALL). Model performance is assessed using accuracy, precision, recall, and F1-score, and the Friedman test is applied to examine statistical consistency across datasets. Results show that hybrid data consistently outperforms single-type datasets, while categorical-only data yields the lowest scores. The best hybrid ensemble of LR-K achieves up to 93% classification accuracy, with strong overall performance also observed for RF on hybrid data. The Friedman test indicates minimal rank differences across models, suggesting that several classifiers are competitive under this ETS setting. Limitations include possible noise and overfitting effects that reduce separation between dataset types. Future work will explore feature ranking-based selection, data-driven segmentation during preprocessing, multi-level classification, imbalance handling, and deeper ensemble strategies to strengthen robustness and generalization.
Copyrights © 2026