The rapid development of agentic IDEs calls for evaluation approaches that assess not only final outputs but also the agentic workflow enacted during software development. This study comparatively evaluates the workflow capabilities of four AI IDE agents, namely Cursor, Windsurf, Trae, and Antigravity, within the five-stage benchmark of developing a web-based learning media application, Next-Gen SPLDV. A descriptive comparative evaluation was conducted using five agentic maturity metrics: task decomposition (DC), tool-use effectiveness (TSR), autonomous recovery capability (ARC), human intervention cost (HIC), and time completion efficiency (TCT), complemented by interaction logs and internal artifacts. The findings indicate distinct performance trade-off profiles across systems. Antigravity appeared relatively more stable descriptively (TSR 96.0%; HIC 3; ARC 8), whereas the other systems exhibited context-dependent strengths: Cursor showed more selective tool use, Trae was efficient in several stages but more vulnerable during database integration, and Windsurf was more exploratory but required higher intervention and recovery effort. Qualitative evidence further suggests that these differences were associated with variations in plan-execute-verify strategies and error-response behavior. Overall, the evaluation of AI IDE agents is better interpreted as a contextual map of workflow trade-offs rather than the identification of a single winner across all settings.
Copyrights © 2026