International Journal of Engineering, Science and Information Technology
Vol 6, No 2 (2026)

Scalable AI-Driven Web Data Extraction Systems: Design and Implementation of an Enterprise Market Analytics Scraping Assistant

Arun Mallur Chandrashekar (Lamar University)



Article Info

Publish Date
25 Apr 2026

Abstract

Maintaining high-quality and continuously updated web data feeds remains a major operational challenge for enterprise market analytics platforms. Conventional rule-based web scrapers are increasingly brittle when confronted with dynamic website layouts, anti-bot countermeasures, and heterogeneous data sources at production scale. This paper presents the design and implementation of a scalable AI-driven web data extraction system deployed in production at RealPage, a leading enterprise real estate technology company. The proposed architecture integrates large language model (LLM)-guided DOM interpretation, a hybrid template-and-AI extraction pipeline, a self-healing failure recovery subsystem, and serverless orchestration using Microsoft Azure Functions to support continuous extraction across more than 50,000 property listing websites daily. The system reduces manual maintenance overhead through adaptive selector regeneration and automated failure classification, while a mandatory JSON Schema validation gate mitigates hallucination risks associated with LLM-generated extraction logic. Evaluation under sustained production conditions demonstrates high extraction accuracy across established domains, scalable throughput under bursty workloads, and significant cost efficiencies compared with equivalent virtual-machine-based infrastructure. The system also demonstrates effective automated recovery from the predominant failure mode, namely layout drift caused by website redesigns. The findings indicate that combining LLM-based semantic interpretation with deterministic validation and automated recovery can improve the reliability, scalability, and maintainability of enterprise web extraction systems. The broader contribution is a validated reference architecture for enterprise-grade AI-augmented web extraction that can be adapted to other domains requiring continuous, large-scale collection of structured information from heterogeneous, dynamic, and frequently changing web sources while reducing operational intervention and improving long-term system resilience in production environments

Copyrights © 2026