C2FT: Enhancing Fine-Grained Perception in MLLMs via Confuse-then-Contrast Fine-Tuning
Abstract
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities, yet they frequently struggle with fine-grained visual perception, often suffering from hallucinations when faced with visually similar but semantically distinct instances. A major underlying cause is the conventional single-sample Supervised Fine-Tuning (SFT) paradigm, which isolates training instances and limits the model's ability to explicitly learn subtle discriminative boundaries. To address this, we propose C2FT, a novel Confuse-then-Contrast Fine-Tuning framework that enhances fine-grained perception by shifting from independent supervision to joint supervision. Specifically, C2FT assesses the MLLM's internal uncertainty to dynamically mine model-specific semantic confusions, subsequently constructing challenging multi-image groups comprising hard positive and hard negative samples. By interleaving these samples into a unified prompt, our approach forces the MLLM's internal attention mechanisms to cross-reference images and explicitly capture localized visual discrepancies. Extensive experiments on both Fine-Grained Visual Classification (FGVC) and visual question answering (VQA) benchmarks demonstrate the effectiveness of C2FT. For instance, on the Qwen3-VL-4B model, our approach achieves an average improvements of 3.87\% and 4.35\% points over standard SFT and GRPO baselines, respectively.