Claim Missing Document
Check
Articles

Found 1 Documents
Search

Exploration Is a State, Not a Setting: A Markov-Switching Reinforcement-Learning Model of Strategy Transitions in Sequential Choice Fathimah Al-Ma'shumah; Feneta Fidi Kirani; Nita Ratnawaty
Journal of Computing Theories and Applications Vol. 4 No. 1 (2026): JCTA 4(1) 2026
Publisher : Universitas Dian Nuswantoro

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.62411/jcta.16818

Abstract

Computational accounts of human exploration usually assume a stationary policy, in which a single set of parameters for value sensitivity and uncertainty seeking generates every choice in a task, so that all within-task variation is treated as decision noise. Neuroscience instead treats exploration and exploitation as dissociable modes between which the brain switches, which a stationary model cannot represent. We introduce the Markov-Switching Reinforcement-Learning (MS-RL) model, in which a first-order hidden Markov chain governs transitions among a small number of latent decision regimes, each with its own softmax policy over a shared value-learning process. The three regimes are exploitation, directed exploration, and random exploration. The model contains the standard stationary account (one regime) and a temporally unstructured mixture (memoryless transitions) as nested special cases, making the stationarity assumption testable. We estimate the model using Expectation-Maximization with Viterbi decoding and select the number of regimes using the Bayesian information criterion. A parameter- and state-recovery study confirmed that the generating parameters and latent regime paths are recoverable (parameter correlations 0.88–0.94; state accuracy 86 percent; Cohen’s kappa ≈ 0.79). Applied to an openly available two-armed bandit dataset (46 adults, 13,800 choices), a three-regime MS-RL model was preferred over the nested baselines, two- and four-regime variants, and a single-regime model with smoothly time-varying value sensitivity. An ablation analysis attributes the largest gains to the latent regimes and temporal switching. The decoded regime path revealed a systematic within-game shift from directed exploration toward exploitation. The estimated transition matrix provides a per-participant measure of strategy change that has no counterpart in stationary models. We conclude that exploration is better described as a dynamic state than as a fixed trait, and that modeling it as a switching process is both more accurate and useful.