Journal of Technology Informatics and Engineering
Vol. 5 No. 2 (2026): AUGUST | JTIE : Journal of Technology Informatics and Engineering

Token-Burst-Aware Capacity Planning for LLM Inference Services: Request Arrival, Token Demand, and Failure Risk Modeling from BurstGPT Traces

Jiayi Nie (Operations Research, Columbia University, NY, USA)
Yinchen Shi (Computer Science, New York University, NY, USA)
Lucas Zhao (Computer Science, Georgia Institute of Technology, GA, USA)



Article Info

Publish Date
03 Aug 2026

Abstract

Large language model (LLM) inference services convert request arrivals into coupled input-token prefill and output-token decoding workloads. Capacity planning therefore depends on token volume, temporal bursts, service mix, queueing behavior, and reliability signals rather than request counts alone. This study evaluates an integrated planning pipeline on BurstGPT v1.1, comprising 5,288,173 raw requests over 121 trace days and 5,188,507 completed requests. Strictly chronological experiments aggregate demand, forecast hourly completed tokens, detect minute-level burst pressure, estimate zero-response risk, simulate capacity policies, and replay representative test hours in Vidur. Random forest, selected on the validation interval, achieved 64.73% weighted absolute percentage error (WAPE) on the locked test interval; XGBoost achieved the lowest test WAPE (64.57%), while the last-hour baseline reached 67.42%, indicating limited forecastability under a pronounced level shift. A seasonal-residual burst detector achieved F1 = 0.653, although burst prevalence was sensitive to rolling-horizon and quantile settings. For minute-level zero-response risk, raw XGBoost achieved ROC-AUC = 0.813 and average precision = 0.116; isotonic calibration improved the Brier score (0.0230–0.0204) and 10-bin expected calibration error (0.0262–0.0122), despite low F1. Static P90/P95 capacity eliminated under-provisioned test hours at cost indices of 13.10 and 18.09. More economical dynamic baselines achieved 12.65% under-provisioned hours at a cost index of 2.37 and 13.63% at 2.43. The validation-selected random-forest policy was cheaper but less reliable (34.31% at 1.23). Vidur replay linked normalized demand to A100/H100 GPU counts, latency, batching, and memory pressure. The results support conservative interpretation of point forecasts and validation of reserve rules under distribution shift, rare-event calibration, and serving-stack constraints.

Copyrights © 2026






Journal Info

Abbrev

jtie

Publisher

Subject

Computer Science & IT

Description

Power Engineering Telecommunication Engineering Computer Engineering Control and Computer Systems Electronics Information technology Informatics Data and Software engineering Biomedical ...