This Author published in this journals
All Journal Julia Jurnal
Erika Dwi Saputra
Prodi Ilmu Komputer, Universitas An Nuur

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

SENTINEL-CAI: Secure Ensemble of Neural Threat Intelligence with Noise-resilient Extraction Learning dan Constitutional AI Integration untuk Comprehensive Claude AI Security Framework Indri Atun Nafisah; Erika Dwi Saputra; Kartika Imam Santoso; Eko Supriyadi
Julia: Jurnal Ilmu Komputer An Nuur Vol 6 No 2 (2026): juliajurnal
Publisher : LPPM Universitas An Nuur

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35720/julia.v6i2.29

Abstract

The massive adoption of Claude AI in mission-critical applications has exposed a variety of sophisticated vulnerabilities, including prompt injection attacks, model inversion, adversarial perturbations, and multi-modal exploitation. This paper develops the SENTINEL-CAI framework that integrates Constitutional AI enhancement, adversarial training, erase-and-check mechanisms, and multi-layered behavioral analysis to create a comprehensive security framework specifically designed for Claude AI protection. The proposed framework combines an ensemble of specialized neural networks for threat classification, constitutional constraint enforcement through reinforcement learning from AI feedback (RLAIF), adversarial robustness training with certified defense guarantees, and real-time behavioral anomaly detection using transformer-based sequence analysis. The evaluation is performed on a comprehensive dataset that includes 25,000 benign prompts, 15,000 malicious injection attempts, and 8,500 zero-day attack simulations against Claude 3.5 Sonnet and Claude Computer Use. Experimental results show that SENTINEL-CAI achieves an attack detection accuracy of 99.2%, a false positive rate of 0.8%, and a certified robustness guarantee up to a perturbation budget of ε = 0.15. This framework successfully detects 96.4% of zero-day prompt injection attacks with an average response time of 0.12 seconds. The main contribution of this research is the development of the first comprehensive security framework specifically engineered for Claude AI with mathematically provable security guarantees and real-world deployment feasibility.