Adhitya Ronnie Effendie
Department of Mathematics, Faculty of Mathematics and Natural Sciences, Gadjah Mada University, Yogyakarta

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

From Risk-Neutral to Risk-Sensitive Reinforcement Learning: Actor–Critic vs REINFORCE with Tail-Based Risk Measures Aprida Siska Lestia; Adhitya Ronnie Effendie; Made Tantrawan; Muhammad Rafli Azrarsyah
CAUCHY: Jurnal Matematika Murni dan Aplikasi Vol 11, No 1 (2026): CAUCHY: JURNAL MATEMATIKA MURNI DAN APLIKASI
Publisher : Mathematics Department, Maulana Malik Ibrahim State Islamic University of Malang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.18860/cauchy.v11i1.40309

Abstract

This study investigates risk-sensitive reinforcement learning (RL) for portfolio decision-making under empirically heavy-tailed return distributions. We compare two policy-gradient architectures—REINFORCE with baseline (REINFORCE-BL) and batched Advantage Actor–Critic (A2C-B)—and examine how tail-based risk measures modify learning dynamics and robustness. Quantitative diagnostics confirm substantial excess kurtosis and strong rejection of normality in daily NASDAQ returns, motivating the integration of tail-sensitive objectives. Risk sensitivity is introduced at the episodic level through penalties based on Value at Risk (VaR), Conditional Value at Risk (CVaR), and Entropic Value at Risk (EVaR) at the 95% confidence level. Experiments are conducted in a multi-asset portfolio exposure-control environment, with performance evaluated across multiple random seeds using both training dynamics and out-of-sample financial metrics (CAGR, volatility, Sharpe ratio, drawdown, and realized tail risk). Results show that while both architectures perform comparably under the risk-neutral objective, actor–critic learning exhibits greater stability and lower dispersion under coherent tail penalties. In particular, CVaR and EVaR objectives lead to smoother convergence and reduced instability compared to VaR, especially for A2C-B. Statistical tests indicate that performance differences become more pronounced under coherent tail-risk objectives. These findings highlight the interaction between heavy-tailed environments, coherent risk measures, and algorithmic architecture, suggesting that actor–critic methods provide a more robust foundation for risk-sensitive RL in financial settings exposed to extreme events.