Cross-Modal Forgetting and Transfer in Unified Multimodal Post-Training
Abstract
Unified multimodal models answer questions about images and also generate images with one shared set of weights. This creates a post-training risk that single-purpose models do not have: fine-tuning one capability can damage another that works in a different modality. We measure this cross-modal forgetting for LoRA post-training of Janus-Pro-7B. Every target is chosen by how poorly the base model already handles it, from a poetic register to a language it barely speaks, so that retention is tested against a real intervention; we then evaluate the capabilities we never trained, first after a single run and then after chaining three. A single run looks entirely safe: nothing we did not train degrades on four public benchmarks or in free-form English, even when the adapter cuts held-out Welsh perplexity by 76% while moving under 2% of the weights. That safety belongs to the run, not to the recipe. Chaining three runs destroys the capability learned one run earlier and raises general-text perplexity by more than two orders of magnitude above the never-adapted model, on the very measures that stayed flat before. The cause is in the objective, which constrains only the modality being trained and leaves the other head unconstrained: cutting the learning rate to one fifth does not prevent the collapse, while anchoring the untrained head removes it at no cost to what is learned. What is learned also stays inside its own modality, with Welsh acquired from text closing only 3% of the English-Welsh gap in image generation.