Journal of Computing Theories and Applications
Vol. 4 No. 1 (2026): JCTA 4(1) 2026

Exploration Is a State, Not a Setting: A Markov-Switching Reinforcement-Learning Model of Strategy Transitions in Sequential Choice

Fathimah Al-Ma'shumah (Bina Nusantara University)
Feneta Fidi Kirani (Bina Nusantara University)
Nita Ratnawaty (Uttama Analytics Consulting)



Article Info

Publish Date
21 Aug 2026

Abstract

Computational accounts of human exploration usually assume a stationary policy, in which a single set of parameters for value sensitivity and uncertainty seeking generates every choice in a task, so that all within-task variation is treated as decision noise. Neuroscience instead treats exploration and exploitation as dissociable modes between which the brain switches, which a stationary model cannot represent. We introduce the Markov-Switching Reinforcement-Learning (MS-RL) model, in which a first-order hidden Markov chain governs transitions among a small number of latent decision regimes, each with its own softmax policy over a shared value-learning process. The three regimes are exploitation, directed exploration, and random exploration. The model contains the standard stationary account (one regime) and a temporally unstructured mixture (memoryless transitions) as nested special cases, making the stationarity assumption testable. We estimate the model using Expectation-Maximization with Viterbi decoding and select the number of regimes using the Bayesian information criterion. A parameter- and state-recovery study confirmed that the generating parameters and latent regime paths are recoverable (parameter correlations 0.88–0.94; state accuracy 86 percent; Cohen’s kappa ≈ 0.79). Applied to an openly available two-armed bandit dataset (46 adults, 13,800 choices), a three-regime MS-RL model was preferred over the nested baselines, two- and four-regime variants, and a single-regime model with smoothly time-varying value sensitivity. An ablation analysis attributes the largest gains to the latent regimes and temporal switching. The decoded regime path revealed a systematic within-game shift from directed exploration toward exploitation. The estimated transition matrix provides a per-participant measure of strategy change that has no counterpart in stationary models. We conclude that exploration is better described as a dynamic state than as a fixed trait, and that modeling it as a switching process is both more accurate and useful.

Copyrights © 2026






Journal Info

Abbrev

jcta

Publisher

Subject

Computer Science & IT Decision Sciences, Operations Research & Management

Description

Journal of Computing Theories and Applications (JCTA) is a refereed, international journal that covers all aspects of foundations, theories and the practical applications of computer science. FREE OF CHARGE for submission and publication. All accepted articles will be published online and accessed ...