Safety-Aware Latent Space Reasoning in Large Language Models
Abstract
Latent space reasoning improves inference efficiency by compressing chain-of-thought (CoT) reasoning into continuous latent states, but its safety alignment remains poorly understood. In this work, we conduct the first systematic safety study of latent space reasoning models under jailbreak attacks. Our results reveal that most latent reasoning methods exhibit higher vulnerability to jailbreak attacks than both the base model and token-space CoT reasoning counterparts, suggesting that moving reasoning into latent states can weaken safety alignment. To address this issue, we propose SaLR, a Safety-aware Latent space Reasoning framework that injects safety supervision into latent space reasoning. SaLR converts long-form safety reasoning into a compact four-block safe-chain and transfers this safety signal through teacher-student hidden state distillation. For safety instances, SaLR aligns an early safe response prefix to guide the model toward a safe response trajectory; for reasoning instances, it retains single position distillation to preserve latent reasoning ability. Experiments across model scales and harmful benchmarks show that SaLR substantially reduces jailbreak attack success while preserving reasoning accuracy and token efficiency. Beyond standard jailbreaks, SaLR also mitigates reasoning specific attacks such as token-inflation and CoT-hijacking attacks by keeping reasoning in fixed latent states and avoiding exposed textual reasoning traces. Overall, SaLR provides a stronger safety utility trade-off for latent space reasoning. Our anonymous code repository is available at: https://anonymous.4open.science/r/SaLR-3C19