Abstract
Recent developments in reasoning capabilities have enabled large language models to solve increasingly complex mathematical, symbolic, and logical tasks. Interestingly, while reasoning models are often trained to generate monolingual text, these models have also been observed to code-switch (i.e., mix languages). Prior works have either viewed code-switching as an undesirable error, attempted to control code-switching through modifications to input prompts or the output decoding process, or focus on narrow subsets of languages, domains, tasks, and models. We address these gaps by introducing the first linguistically and behaviorally motivated fine-tuning framework for identifying beneficial code-switched reasoning behaviors in large language models and teaching these models to code-switch more effectively for reasoning. We create the Code-Switched Reasoning (CoRe) corpus, consisting of (1) 7k reasoning traces from 15 models, 18 languages, 10 scripts, and diverse reasoning domains, providing insights into potentially helpful code-switching behaviors, and (2) 40 carefully curated datasets for training and evaluating six interventions for improving code-switching in reasoning across three models and seven languages, totaling 120 fine-tuning conditions. Across 80k+ reasoning traces from both language/culture-agnostic and -specific evaluations, English-dominated reasoning, semantically accurate code-switching, and a more even mix of languages are positively associated with correct answers, whereas dense switching is associated with incorrect answers. Moreover, we are able to elicit positive behaviors through fine-tuning tasks that do not directly demonstrate code-switching. Our work suggests that small but well-curated datasets can change how reasoning models code-switch, allowing us to reap the benefits of reasoning in the many languages that lack large-scale reasoning data.