cover
Contact Name
Chairul Anwar
Contact Email
info@jci.co.id
Phone
+6285128002422
Journal Mail Official
info@jci.co.id
Editorial Address
Jl Bunga Melati GG H Yahya RT 04 RW 02 No. 52, Cipete Selatan, Cilandak, Jakarta Selatan, DKI Jakarta
Location
Kota adm. jakarta selatan,
Dki jakarta
INDONESIA
Journal of Information Systems and Business Technology
ISSN : -     EISSN : 31098886     DOI : -
Core Subject : Science,
Journal of Information Systems and Business Technology (JISBT) adalah jurnal ilmiah yang didedikasikan khusus untuk pengembangan keilmuan di bidang Sistem Informasi. Jurnal ini menjadi wadah untuk penyebaran hasil penelitian, inovasi teknologi, serta pemikiran kritis yang berfokus pada penerapan dan pengembangan sistem informasi dalam berbagai konteks bisnis dan organisasi. Jurnal ini bertujuan untuk menjadi referensi utama bagi akademisi, peneliti, dan praktisi yang bergerak di bidang Sistem Informasi, serta menjadi sarana kontribusi terhadap peningkatan literasi teknologi dan manajemen informasi di era digital. Journal of Information Systems and Business Technology (JISBT) diterbitkan sebanyak empat kali dalam setahun, yaitu pada bulan Juni, Agustus, Oktober, dan Desember. Semua publikasi dalam JISBT bersifat akses terbuka, memastikan bahwa setiap artikel tersedia secara daring tanpa biaya berlangganan. JISBT menerima artikel ilmiah orisinal yang relevan dengan topik-topik berikut dalam bidang Sistem Informasi: Manajemen Sistem Informasi Strategi Sistem Informasi dalam Bisnis Sistem Informasi Keuangan dan Akuntansi Sistem Informasi Pemasaran Sistem Informasi Sumber Daya Manusia E-Commerce dan Sistem Transaksi Digital Customer Relationship Management (CRM) Business Intelligence dan Analitik Data Big Data dan Sistem Pendukung Keputusan Keamanan Informasi dan Keamanan Siber Integrasi Sistem dan Interoperabilitas Audit Sistem Informasi dan Tata Kelola TI Rekayasa Perangkat Lunak untuk Sistem Informasi Desain UI/UX dan Pengalaman Pengguna Data Mining dan Machine Learning untuk Sistem Informasi Inovasi dan Transformasi Digital Cloud Computing dalam Implementasi Sistem Informasi Internet of Things (IoT) dalam Sistem Informasi Enterprise Resource Planning (ERP) Topik lain yang relevan dengan Sistem Informasi Catatan: Artikel yang dikirimkan harus bersifat orisinal, memiliki sitasi yang benar, dan belum pernah dipublikasikan sebelumnya, baik secara cetak maupun digital.
Articles 218 Documents
Evidence-Gated Multi-Domain AIOps Copilot on LogLM: Leakage-Controlled Parsing, Template-Level Anomaly Classification, Root-Cause Retrieval, and Confidence-Gated Remediation Brandon Wright; Ling Yun; Sophia Martinez
Journal of Information Systems and Business Technology Vol 2 No 4 (2026): Journal of Information Systems and Business Technology
Publisher : PT Jurnal Cendekia Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Operational copilots need traceable evidence and leakage-resistant evaluation. We evaluate an evidence-gated AIOps pipeline on the 2,632-record LogLM corpus. The audit found 622 parsing repetitions and aligned Apache overlap, so exact-input and connected-case groups governed every split. The nested parser reached 0.827 exact-template accuracy, versus 0.544 for regex and 0.574 for hybrid nearest-neighbor retrieval; row stratification inflated the latter to 0.908 because 47.7% of test messages had an identical training neighbor. On grouped BGL folds, word–character logistic regression reached AUROC 0.870 ± 0.026, while a fault lexicon led abnormal F1 at 0.687. Nested Platt scaling reduced Brier loss from 0.128 to 0.095 and calibration error from 0.188 to 0.044; learned models produced a 0.058 false-positive rate on negative-only Spirit. Case-grouped Apache retrieval reached overall ROUGE-L 0.261, increasing to 0.364 at 20% coverage. The findings support selective, evidence-linked assistance and human review for sensitive actions.
Fair Incrementality Learning and Conservative Policy Selection without Persistent Identifiers: Calibrated Counterfactual Evaluation for Ads and Job Ranking Chen Yang; Derek Peterson; Lin Feng; Rachel Adams
Journal of Information Systems and Business Technology Vol 2 No 4 (2026): Journal of Information Systems and Business Technology
Publisher : PT Jurnal Cendekia Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Identifier loss complicates advertising measurement, while job recommenders must balance utility and opportunity. We linked two separate tracks through conservative selection: randomized Criteo Uplift v2.1 for intention-to-treat estimation and FairJob for identifier-free ranking and proxy-group audit. S-HGB achieved test Qini 4.689×10⁻³, top-decile uplift 6.746 points, and top-20% gain 0.974 points; a lower-confidence-bound rule treated 30% and gained 1.007 points. FairJob's best PR-AUC was 0.00833. Proxy penalization raised MRR from 0.574 to 0.585 but widened group disparity from 0.023 to 0.071, so deployment was rejected. Evidence cards quantified local faithfulness without interpreting anonymized fields. Calibration and gating supported identifier-free decisions, but predictive parity did not ensure fair ranking.
Review-Grounded Explainable Recommendation under Extreme User Sparsity: Self-Supervised Reconstruction, Behavior Retrieval, Client-Partitioned Preference Learning, and Evidence-Support Evaluation Joshua Baker; Yan Wang; Bradley Cook; Xiaoyu Liu
Journal of Information Systems and Business Technology Vol 2 No 4 (2026): Journal of Information Systems and Business Technology
Publisher : PT Jurnal Cendekia Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Amazon Reviews'23 Gift Cards tests whether review text helps when users are sparse and products popularity-dominated. Its 152,410 reviews cover 132,732 users, 1,137 parent products, and complete metadata. Collapsing repeat user-product pairs yielded 125,704 verified positives rated at least four stars. Leave-two-out retained 120,154 training interactions and evaluated 2,769 active users over all 1,009 items with one frozen pre-validation history. Ten recommenders covered popularity, transitions, item/user retrieval, low-rank and corrupted-view reconstruction, BPR, graph propagation, review retrieval, and gated fusion. Separate branches tested first-observation cold start, clipped-noise client learning, and post-hoc evidence attribution. Markov achieved NDCG@10 0.2346 and HR@10 0.4056; gating gave it all weight, so reviews neither improved warm ranking nor caused recommendations. In the 2021 cold-start proxy, metadata TF-IDF reached micro NDCG@10 0.4613 over 113 candidates, but item-macro and dominant-target-excluded NDCG@10 scores were 0.0627 and 0.0605. Across 500 cases, the cited template reached 0.998 lexical support and 1.000 source localization; shuffling preserved support but reduced user-profile cosine from 0.1683 to 0.0392. Thus ranking, support, alignment, and causal faithfulness diverge.
Conformal GPU Demand Envelopes and Power-Aware Scheduling for Heterogeneous AI Clusters under Cold Start and Cross-Cloud Shift Sijia Chen; Marcus Reed; Hong Zhang
Journal of Information Systems and Business Technology Vol 2 No 4 (2026): Journal of Information Systems and Business Technology
Publisher : PT Jurnal Cendekia Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Production LLM traffic is bursty, and capacity decisions couple GPU availability with power and carbon. We convert probabilistic demand forecasts into risk-controlled admission, heterogeneous allocation, calibrated overbooking, and cross-site dispatch. The evaluation uses all 1,429,737 requests in the 61-day BurstGPT trace and SustainDC carbon, weather, workload, configuration, and physical power model. Requests form a complete five-minute grid with a weighted-token target. A chronological split assigns 42 days to fitting, nine to conformal calibration, and ten to testing. Point models include persistence, seasonality, Ridge, histogram gradient boosting, and Extra Trees; interval models include direct quantiles, split, rolling, regime-conditioned, and conformalized quantile regression. Histogram gradient boosting attained 49.68% WAPE versus 115.81% for daily seasonality and reduced mean absolute error by 31,545.90 weighted tokens/bin (95% block-bootstrap interval, 23,152.07–37,746.37). Rolling conformal achieved 90.00% coverage at mean width 86,728.29. With a calibration-selected 0.80 overbooking factor, the proposed policy recorded 5.11% bin-level SLO violations and reduced emissions from 52.51 to 12.57 tCO2e through SustainDC-aware dispatch. The remaining 8.81% workload shortfall quantifies burst risk; GPU type, site, and price remain scenario inputs.
Retrieval-Augmented Triage for Software Engineering Agents: Source-Path and Pre-Fix Hunk Localization with Calibrated Repair-Effort Prediction on SWE-bench Verified Eric Sullivan; Fang Wei; Chloe Bennett; Jun Ma
Journal of Information Systems and Business Technology Vol 2 No 4 (2026): Journal of Information Systems and Business Technology
Publisher : PT Jurnal Cendekia Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Repository-scale software agents need constrained context before reasoning about a defect. This study evaluates retrieval and triage on all 500 human-validated SWE-bench Verified tasks from 12 Python repositories. Unified-diff parsing recovered 373 modified source paths and 1,220 source hunks. In the primary issue-only closed-world experiment, a leave-one-repository-out reranker achieved 0.500 Hit@1, 0.747 Recall@5, and 0.625 MRR; its MRR gain over character TF-IDF was 0.065 (95% CI 0.046-0.087; adjusted p < 0.001) and remained significant with equal repository weighting. Failing-test identifiers increased post-failure RRF MRR from 0.590 to 0.685, while issue-only word TF-IDF reached 0.684 MRR on gold-selected pre-fix hunks. Nested cross-repository calibration yielded 0.139 PR-AUC, 0.632 ROC-AUC, 0.082 Brier score, and 0.028 ECE for high-effort triage. Lexical localization was effective, but pre-execution effort forecasting remained limited; retrieval evidence, escalation, and executable repair should therefore remain separate stages.
Continual Guardrail Learning for Tool-Using LLM Agents: Cross-Benchmark Jailbreak Detection, Indirect Prompt-Injection Filtering, and Utility Preservation Hao Ran; Natalie Foster; Bo Liang; Justin Meyer
Journal of Information Systems and Business Technology Vol 2 No 4 (2026): Journal of Information Systems and Business Technology
Publisher : PT Jurnal Cendekia Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Tool-using large language model agents process untrusted text while holding permissions to communicate, book services, and move funds. This study evaluates lightweight guardrails under changing attacks while preserving authorized task utility. The text study uses all 6,899 saved runs in one internally consistent AgentDojo v1 GPT-4o pipeline and 200 JailbreakBench prompts covering 100 behavior groups. Deduplication produced 5,280 texts: 1,450 attacks and 3,830 benign records. Injection-goal grouping assigned 3,169, 1,056, and 1,055 records to training, validation, and testing without sharing an AgentDojo attack target. A normalized word-character classifier achieved 85.43% F1, 90.31% recall, 7.96% benign false-positive rate, and 0.9613 ROC-AUC. A 1,000-round group bootstrap gave a wide 59.71-97.68% F1 interval. Cross-source recall fell to 33.00% from AgentDojo to JailbreakBench and 0.16% in the reverse direction. Sequential tests compared frozen, benign-anchored, hard-replay, and soft-label-replay learners. Reanalysis of 629 matched AgentDojo attacks showed that tool filtering reduced targeted attack success from 47.69% to 6.84% while increasing safe utility from 29.73% to 52.62%. Layered guardrails are therefore necessary: replay maintains coverage under shift, while tool-level enforcement provided the strongest observed end-to-end security-utility balance.
Risk-Controlled Adaptive Multi-Hop RAG with Evidence-Chain Retrieval, Word-Level Hallucination Detection, Attribution, and Calibrated Abstention Megan Walters; Qing Song; Austin Jiang
Journal of Information Systems and Business Technology Vol 2 No 4 (2026): Journal of Information Systems and Business Technology
Publisher : PT Jurnal Cendekia Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Reliable RAG fails when retrieval omits evidence, generation is unsupported, or confidence is misplaced. We evaluated all 2,556 MultiHop-RAG questions and 2,700 official RAGTruth test responses (2,675 good-quality in primary analysis). On chain-group-disjoint MultiHop-RAG, adaptive retrieval achieved 0.8810 Recall@k and 0.6925 complete-chain rate with 10.85 documents per answerable query; the nearest lower-cost comparator reached 0.8463 and 0.6273. Answer token F1 was 0.5562 and strict URL attribution 0.7753. On RAGTruth, the calibrated detector reached 0.3254 word F1, 0.0986 overlap span F1, 0.00582 ECE, and 0.03550 Brier score. Selective release cut risk from 0.3525 at full coverage to 0.2140 at 80%. Evidence, grounding, provenance, and release risk require separate coordinated measurement.
From Enterprise UI Screenshots to Trustworthy Front-End Prototypes: Structural–Visual Consistency, Accessibility, Progressive Rendering, and an Evidence Layer for LLM-Assisted Design Critique Wei Dong; Sierra Campbell; Lei Wu; Gabriel Ross
Journal of Information Systems and Business Technology Vol 2 No 4 (2026): Journal of Information Systems and Business Technology
Publisher : PT Jurnal Cendekia Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Screenshot-to-code systems are often judged by visual resemblance, yet deployable prototypes also require coherent DOM structure, accessible semantics, and predictable rendering. We evaluated these dimensions on 484 Design2Code screenshot–HTML pairs using deterministic screenshot features, static markup audits, and CSS-hash-grouped five-fold cross-validation. Random Forest and Extra Trees predicted DOM size with R² values of 0.215 and 0.244, respectively, but accessibility-burden R² remained below zero, confirming that visual fidelity cannot substitute for markup inspection. Train-fold-only visual retrieval scored 78.26 on a composite consistency measure; structural reranking scored 79.59, and a deterministic evidence reranker combining predicted structure, accessibility, and static critical-render-path profiles scored 79.63. Its 1.37-point gain over visual retrieval was significant (95% bootstrap confidence interval 0.92–1.86; one-sided Wilcoxon p < 0.001) and cost 0.36 visual-similarity points. Conservative remediation removed 218 of 1,981 detected issues, raised zero-issue pages from 59 to 79, and preserved normalized body text and embedded CSS on all pages. On a fixed 25-page diagnostic subset, mean SSIM relative to each page's complete-CSS render increased from 0.693 for semantic HTML to 0.784 with typography; the complete renders independently reached median SSIM 0.996 against the stored screenshots. These findings establish an auditable evidence layer for subsequent LLM-assisted critique rather than an evaluated language-model critic.