Learning What to Predict: Downstream-Guided Task Design for Continued Pretraining
Shuqi Ke ⋅ Giulia Fanti
Abstract
Continued pretraining is optimized on a fixed self-supervised task but selected by downstream performance. This creates a coarse feedback loop: practitioners evaluate checkpoints, revise data mixtures or objectives, and rerun pretraining runs, while individual pretraining updates receive no signal about whether they help the target capability. We ask whether a small set of verifiable downstream examples can provide step-level feedback during continued pretraining without becoming learner supervision. We introduce V-pretraining, which separates a learner trained only by a self-supervised loss from a lightweight task designer that constructs targets or views for unlabeled batches. Given the current learner and an unlabeled batch, V-pretraining estimates the downstream value of a candidate target or view construction by the first-order predicted decrease in downstream loss after the self-supervised update it induces. The designer is trained to increase this value; the learner then applies the resulting self-supervised update with targets or views detached, so downstream labels never directly update learner parameters. V-pretraining can be used to learn adaptive top-$K$ soft targets for language modeling and learned views for self-supervised vision. Under wall-clock-matched continued pretraining, V-pretraining improves GSM8K Pass@1 for Qwen models using 1,024 GSM8K examples only as feedback, including a +7.4 point single-run gain for Qwen2.5-0.5B. In vision, V-pretraining improves DINOv3 transfer to ADE20K semantic segmentation and NYUv2 depth estimation while preserving ImageNet linear accuracy, indicating that feedback-guided task construction improves target downstream capabilities without collapsing general-purpose representations.
Chat is not available.
Successful Page Load