Attributing Gains in Paired Training: Objective or Data Direction?
Abstract
Fine-tuning a model on one-way code obfuscation destroys its ability to perform the inverse operation of deobfuscation. Recent work attributes the recovery of this inverse capability to Contrastive Fine-Tuning (CFT), an auxiliary equivalence-judgement objective that purports to work without reverse generation supervision. However, auxiliary objectives introduce both a new loss term and additional paired data distributions simultaneously, confounding the true source of recovery. Using a controlled factorial design on code transformations, we isolate the loss formulation from the data distribution. We find that the auxiliary contrastive objective provides no meaningful retention on its own. Instead, a small amount of reverse-direction supervision prevents the collapse; the recovery is attributable to data direction rather than the contrastive objective. Simply exposing the model to reversed input-output pairs under standard fine-tuning restores the inverse capability without adding compute or modifying the training loss. While demonstrated on code deobfuscation, these findings establish bidirectional data formatting as an essential, zero-overhead baseline that extends naturally to evaluate and train models across reversible paired domains.