Objectives: This study compares five AI platforms (ChatGPT 5.2, Claude Sonnet 4.5, Gemini 3 Pro, GLM 4.7, DeepSeek V3.2) with manual HR assessment in team performance evaluation, examining accuracy and predictive validity by correlating with program success rates.Methodology: Using a within-subjects design, 32 team members managing 619 locations were evaluated by six methods via standardized prompts. Analysis included comparative tests, accuracy metrics, and correlation between performance scores and graduation rates.Finding: All AI platforms scored higher than the HR baseline. Correlation analysis revealed substantial differences in predictive validity. Claude Sonnet 4.5 demonstrated the strongest correlation with graduation rates (ρ = 0.70), followed by GLM 4.7 (ρ = 0.67), accounting for approximately 50-55% of the outcome variability. The HR baseline unexpectedly showed a negative correlation with program success (ρ = -0.49), indicating misalignment between the evaluation criteria and actual effectiveness. DeepSeek V3.2 most closely matched HR scores but showed only moderate predictive validity, while ChatGPT 5.2 showed no meaningful relationship with outcomes.Conclusion: AI platforms differ substantially in predictive validity. Organizations should prioritize outcome-based validation over human agreement when selecting AI for HR decisions and consider recalibrating manual evaluation systems.
Copyrights © 2026