Nonconvex Regularization for Feature Selection in Reinforcement Learning

  • 2025-09-19 06:21:20
  • Kyohei Suzuki, Konstantinos Slavakis
  • 0

Abstract

This work proposes an efficient batch algorithm for feature selection inreinforcement learning (RL) with theoretical convergence guarantees. Tomitigate the estimation bias inherent in conventional regularization schemes,the first contribution extends policy evaluation within the classicalleast-squares temporal-difference (LSTD) framework by formulating aBellman-residual objective regularized with the sparsity-inducing, nonconvexprojected minimax concave (PMC) penalty. Owing to the weak convexity of the PMCpenalty, this formulation can be interpreted as a special instance of a generalnonmonotone-inclusion problem. The second contribution establishes novelconvergence conditions for the forward-reflected-backward splitting (FRBS)algorithm to solve this class of problems. Numerical experiments on benchmarkdatasets demonstrate that the proposed approach substantially outperformsstate-of-the-art feature-selection methods, particularly in scenarios with manynoisy features.

 

Quick Read (beta)

loading the full paper ...