Speaking skills of Madrasah Ibtidaiyah (MI) students, particularly fluency, pronunciation, and self-confidence, remain low due to limited communicative practice, high speaking anxiety, and the absence of continuous individual feedback in conventional learning. Collaborative learning and short conversation videos have been reported to increase motivation and participation, yet they have not been systematically integrated with an objective, adaptive, and contextual automatic feedback mechanism suited to MI learners. This study aims to design an integrative conceptual model that combines short conversation videos with a deep learning-based automatic feedback system to improve MI students' English speaking fluency while strengthening their motivation and confidence. A Design-Based Research (DBR) approach was employed, covering needs-based design, conceptual model development, and expert validation, limited to Technology Readiness Level (TRL) 1-3. At TRL 1, a needs analysis and literature review identified the gap between conventional collaborative speaking instruction and systematic automatic feedback. At TRL 2, a conceptual model and system blueprint were formulated with fluency, pronunciation, and coherence as feedback indicators. At TRL 3, the model was validated by experts in English Language Teaching (ELT), MI Teacher Education (PGMI), and instructional technology to ensure pedagogical, linguistic, and technological feasibility. The results present a validated conceptual model consisting of three integrated components: short conversation video stimuli, collaborative learning activities, and a real-time deep learning feedback framework. The model offers a foundation for developing adaptive, AI-supported speaking instruction relevant to MI learners' characteristics and twenty-first century learning demands.