Beyond Fair Answers: Learning Fair and Unlearning Biased Reasoning in LLMs
Abstract
Large Language Models (LLMs) have demonstrated strong reasoning capabilities on tasks with well-defined solutions, such as mathematical and logical reasoning. However, reasoning capability does not necessarily imply unbiased reasoning. In socially ambiguous contexts, where the available evidence is insufficient to support one conclusion over another, models may rely on demographic stereotypes or unsupported social assumptions. Recent work shows that chain-of-thought (CoT) reasoning can contain social stereotypes and that biased reasoning steps are associated with stereotype expression and incorrect predictions. This motivates evaluating fairness in the reasoning process, not only in the final answer. We therefore ask: Can we train LLMs not only to produce fair answers, but to reason fairly? We study this question using a fairness-oriented adaptation of BBQ, where ambiguous cases test whether demographic alternatives receive comparable probability when evidence is insufficient, while disambiguated cases test evidence-grounded accuracy. Following reasoning-trace fine-tuning approaches such as STaR and ReGiFT, we fine-tune instruction-tuned LLMs on reasoning traces generated by stronger models and compare base and fine-tuned models under direct and CoT inference. We evaluate both final decisions and generated reasoning traces, and test generalization on SocioEconomicQA and CrowS-Pairs. Building on this analysis, we propose a reasoning-level fairness framework that jointly learns fair reasoning and unlearns biased reasoning. Evidence-grounded traces form the positive learning set, while traces containing unsupported demographic or stereotype-based inferences form the forget set. Our objective combines supervised fine-tuning on fair traces, a negative objective on biased traces, and KL regularization on unrelated reasoning data to preserve general reasoning ability. We evaluate whether this process improves selection of the fair answer and, more importantly, reduces the probability gap between gendered options when the context provides no evidence favoring either one. We separately measure accuracy in disambiguated contexts to ensure that the model still uses available evidence correctly. At the reasoning level, we evaluate stereotype-dependent inferences and the relative likelihood of fair versus biased reasoning traces. We compare against ReGiFT, standard CoT, and instruction-based reasoning baselines, and evaluate generalization across additional bias benchmarks while monitoring general reasoning performance for capability loss. This work moves fairness intervention from the final decision to the reasoning process that produces it. By combining fair-reasoning supervision, biased-reasoning unlearning, and capability preservation, we investigate whether LLMs can learn not only to give fair answers, but to reason without relying on unsupported demographic stereotypes.