Digital advertising design is often judged with delivery or response metrics such as click-through rate (CTR), video-completion rate (VCR), and viewability, although these measures do not describe whether a banner communicates through a clear visual hierarchy. This study evaluates whether large language model (LLM)-assisted scores and lightweight image features can support graphic advertising review while keeping design-quality prediction distinct from measured visual attention. GraphicDesignEvaluation provided 1,200 principle–image rating rows for alignment, overlap, and white space. Under grouped five-fold cross-validation by base design, the GPT-only Ridge model reached R2 = 0.408, RMSE = 1.566, MAE = 1.257, Pearson = 0.639, and Spearman = 0.634. Direct GPT–human agreement was uneven: overlap was strong (Spearman = 0.769), white space was moderate (0.637), and alignment was weaker (0.519). An external construct analysis used 1,000 advertisements from ADD1000 with aggregate human fixation-density maps and matched EAID ratings. A combined gradient and local-contrast proxy modestly exceeded a center prior (CC = 0.457, SIM = 0.562, KL divergence = 0.637), while lower fixation entropy was associated with higher aesthetic ratings (Spearman = -0.319). BannerRequest400 was used only to describe brief, CTA-language, logo, and format constraints because it contains no human design-quality or gaze labels. The findings support a preliminary, human-supervised design-review framework: LLM scores are useful screening evidence, especially for overlap, but they do not replace human creative judgment, eye tracking, or advertising-performance measures.