Abdul Saboor Hamedi
Pamulang University

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Empirical Benchmarking of Hybrid Retrieval in Educational Conversational AI: Accuracy‑Latency Trade‑offs and Robustness Abdul Saboor Hamedi; Iqbal Hussain Alamyar; A.A. Waskita
Journal of Intelligent Systems Technology and Informatics Vol 2 No 2 (2026): JISTICS, July 2026
Publisher : Aliansi Peneliti Informatika

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.64878/jistics.v2i2.245

Abstract

This study investigates the comparative performance of lexical, semantic, and hybrid retrieval strategies in educational conversational AI, with a focus on accuracy-latency trade‑offs and robustness across diverse query types. A controlled experimental framework was implemented using PostgreSQL’s ts_rank for lexical retrieval, pgvector embeddings for semantic retrieval, and two fusion strategies: Linear Weighted Fusion and Reciprocal Rank Fusion. The evaluation corpus consisted of approximately 50,000 text chunks extracted from 2025 arXiv AI/ML papers, and a benchmark of 100 queries spanning conceptual, factual, procedural, comparative, and miscellaneous categories was executed. Effectiveness was measured using NDCG@10, Precision@5, and MRR, while efficiency was quantified via end‑to‑end latency. Relevance judgments were generated through an AI‑as‑a‑Judge pipeline to ensure scalability and reproducibility. Results showed that semantic and hybrid methods achieved a high accuracy mean NDCG@10 ≈ 0.91 but incurred latency costs between 227-505 ms. Lexical retrieval was fastest, 88 ms, but substantially less accurate, 0.346. Hybrid‑Linear fusion emerged as the most robust strategy, winning 66% of queries in the Winner‑Take‑All analysis, while semantic search excelled in conceptual queries and lexical search in acronym‑based factual lookups. Reciprocal Rank Fusion achieved comparable mean accuracy but failed to dominate in any category. The findings highlight a clear quality–speed dichotomy and establish Hybrid‑Linear fusion as the most dependable retrieval method for educational chatbots. For latency‑sensitive applications, semantic search offers the best balance of responsiveness and accuracy. The study provides actionable design guidelines and identifies future directions, including corpus generalization, human evaluation calibration, intelligent query routing, and latency optimization.