Operational copilots need traceable evidence and leakage-resistant evaluation. We evaluate an evidence-gated AIOps pipeline on the 2,632-record LogLM corpus. The audit found 622 parsing repetitions and aligned Apache overlap, so exact-input and connected-case groups governed every split. The nested parser reached 0.827 exact-template accuracy, versus 0.544 for regex and 0.574 for hybrid nearest-neighbor retrieval; row stratification inflated the latter to 0.908 because 47.7% of test messages had an identical training neighbor. On grouped BGL folds, word–character logistic regression reached AUROC 0.870 ± 0.026, while a fault lexicon led abnormal F1 at 0.687. Nested Platt scaling reduced Brier loss from 0.128 to 0.095 and calibration error from 0.188 to 0.044; learned models produced a 0.058 false-positive rate on negative-only Spirit. Case-grouped Apache retrieval reached overall ROUGE-L 0.261, increasing to 0.364 at 20% coverage. The findings support selective, evidence-linked assistance and human review for sensitive actions.
Copyrights © 2026