Abstract
The ability to recognize one's own limitations and decide whether to solve a problem or seek help is fundamental to reliable intelligent systems. Yet modern large language models tend to overestimate their competence and attempt queries they cannot solve. We study how models can learn Capability Self-Assessment (CSA) while preserving their problem-solving ability. Because the same model both assesses and solves queries, self-assessment training can change the capability it must recognize. We systematically compare supervised fine-tuning (SFT) and reinforcement learning (RL), evaluating each trained model's self-assessment against its updated capability and measuring how well it retains its solving ability. Across model families, scales, and domains, RL provides the strongest assessment-retention trade-off. We also find that CSA transfers to the out-of-distribution settings we evaluate, and demonstrate its practical value for local-cloud routing at inference time and targeted data selection during training.