Poetry concentrates compositional demand in the pre-writing stage, where imagery, mood, and diction must be generated before a single line is committed. This study asks not whether the Picture and Picture cooperative model raises poetry-writing attainment but where in the composing process it acts. Drawing on capacity and generative-learning accounts, it advances an offloading hypothesis: arranging an ordered image sequence externalises the generating and organising sub-processes of planning, freeing verbal working memory for translation and review. Two action research cycles were conducted with an intact Grade X class of 21 students in Tolitoli, Central Sulawesi. The design isolates the mechanism by holding every procedural element constant across cycles except the sequencing and justification step, extended in Cycle II. Attainment against the school criterion of 75 rose from 28.6% at baseline to 71.4% in Cycle I and 90.5% in Cycle II, and the class mean rose from 70.5 to 80.5 between cycles, T = 186.0, p < .001, dz = 1.22. Two findings bear on mechanism rather than magnitude. Learners furthest below the criterion gained more than twice as much as those already above it (17.5 versus 7.0 points, U = 74.0, p = .023), the pattern predicted if the scaffold relieves a generative bottleneck rather than adding motivation. The Cycle I distribution was also discontinuous, with no scores across a ten-point interval below the criterion, consistent with a threshold effect. The study repositions visual sequencing as a targeted pre-writing scaffold and specifies the evidence a mechanism claim in this literature must meet.