Learning Where It Matters: Geometric Anchoring for Robust Preference Alignment
Abstract
Preference optimization aligns large language models from pairwise preferences by increasing the margin between preferred and dispreferred responses. However, margin-based losses alone do not test whether each pair's preference signal is locally stable enough to trust with full update strength. We propose Geometric Anchor Preference Optimization (GAPO), a geometry-aware objective that introduces a batch-conditioned stress test for preference learning. For each mini-batch, GAPO constructs a pessimistic anchor by perturbing the current policy in the first-order direction that decreases the batch-average preference margin. The resulting Anchor Gap measures how much each pair's margin degrades under this shared pessimistic probe and converts this degradation into an instance-wise update weight. Pairs with larger Anchor Gap receive smaller update weights, while non-brittle pairs retain their preference-gradient direction. Under smoothness assumptions, we characterize the Anchor Gap as a batch-directional proxy for local margin degradation. Across multiple open-weight model families, GAPO matches or improves strong preference-optimization baselines on instruction-following and reasoning benchmarks. It also improves robustness under random and structured preference noise without explicitly modeling label corruption. Mechanistic diagnostics show that the pairs receiving the smallest GAPO weights are statistically enriched in corrupted supervision, suggesting that GAPO improves robustness by reducing the cumulative influence of brittle preference signals.