Fifth generation (5G) networks enable radio access network (RAN) slicing, allowing a shared physical infrastructure to be partitioned into multiple logical slices, each tailored to meet heterogeneous service requirements. Slice tenants specify service-level agreements (SLAs) that must be satisfied; however, inefficient resource allocation mechanisms frequently result in SLA violations. Recent studies have adopted constrained multi-agent reinforcement learning (MARL) to dynamically allocate RAN resources, yet these approaches often remain inefficient due to reward misalignment with SLA objectives and unstable exploration in high-dimensional state–action spaces. This paper presents a reward-function–centric analysis of constrained MARL for 5G RAN slicing, examining how different reward designs affect learning efficiency and SLA compliance. Based on this analysis, we propose heuristic-guided reward shaping strategies that explicitly align the learning objective with SLA satisfaction and resource utilization efficiency. In addition, we introduce a state-space optimization technique that aggregates local agent states into a weighted global state representation, thereby reducing state dimensionality while preserving critical SLA-related information. Extensive experimental evaluations demonstrate that the proposed approach achieves a 60% reduction in SLA violations compared to an existing constrained MARL model and a 46% reduction compared to a state-of-the-art model-based reinforcement learning approach. Furthermore, the proposed solution converges faster and requires less training time and computational resources, highlighting its effectiveness and practical applicability for efficient 5G RAN slicing.
Copyrights © 2026