Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models

  • 2025-08-21 03:31:11
  • Yuanchen Zhou, Shuo Jiang, Jie Zhu, Junhui Li, Lifan Guo, Feng Chen, Chi Zhang
  • 0

Abstract

Process Reward Models (PRMs) have emerged as a promising framework forsupervising intermediate reasoning in large language models (LLMs), yetexisting PRMs are primarily trained on general or Science, Technology,Engineering, and Mathematics (STEM) domains and fall short in domain-specificcontexts such as finance, where reasoning is more structured, symbolic, andsensitive to factual and regulatory correctness. We introduce \textbf{Fin-PRM},a domain-specialized, trajectory-aware PRM tailored to evaluate intermediatereasoning steps in financial tasks. Fin-PRM integrates step-level andtrajectory-level reward supervision, enabling fine-grained evaluation ofreasoning traces aligned with financial logic. We apply Fin-PRM in both offlineand online reward learning settings, supporting three key applications: (i)selecting high-quality reasoning trajectories for distillation-based supervisedfine-tuning, (ii) providing dense process-level rewards for reinforcementlearning, and (iii) guiding reward-informed Best-of-N inference at test time.Experimental results on financial reasoning benchmarks, including CFLUE andFinQA, demonstrate that Fin-PRM consistently outperforms general-purpose PRMsand strong domain baselines in trajectory selection quality. Downstream modelstrained with Fin-PRM yield substantial improvements with baselines, with gainsof 12.9\% in supervised learning, 5.2\% in reinforcement learning, and 5.1\% intest-time performance. These findings highlight the value of domain-specializedreward modeling for aligning LLMs with expert-level financial reasoning. Ourproject resources will be available at https://github.com/aliyun/qwen-dianjin.

 

Quick Read (beta)

loading the full paper ...