Turnitin similarity reports contain rich spatial, color, and textual information, yet highlighted passages are usually inspected visually rather than converted into structured data. This study develops a report-aware pipeline that transforms color-coded Turnitin highlights into analyzable records. The method combines PDF rendering, page-type triage, hue–saturation–value (HSV) segmentation, morphological processing, connected-component analysis, geometric filtering, and GPU-enabled optical character recognition (OCR), with PDF text-layer inspection as supporting information. The evaluation corpus comprised 106 reports and 3,057 pages. Page triage identified 1,315 highlight-bearing pages and routed 56.98% of pages away from detailed processing. HSV segmentation produced eight color classes and 11,471 pre-filter components. After filtering and text extraction, 7,530 structured records were retained; 76.75% satisfied the predefined good-readability criterion, 20.04% were very short fragments, and 3.21% were noisy, garbled, or empty. Document workload was strongly right-skewed. The results establish systems-level feasibility for auditable extraction without treating textual similarity as automatic evidence of plagiarism or the readability indicator as OCR accuracy.
Copyrights © 2026