On the Effect of Token Correlations on Semantic Transfer in Transformers
Abstract
We study issues which may arise when using transformer models to formally reason about problems stated informally using natural language. In particular, we focus on transfer learning between tasks that are semantically equivalent but phrased differently. We assume that examples with a particular target label are only available for a subset of tasks, and aim to transfer zero-shot the ability to predict this label to other tasks. Using controlled experiments on propositional logic problems, we show that transformers are capable of such transfer, but that it can be impeded by the presence of task-specific tokens, due to the model picking up on correlations between these tokens and the label early in the training. As a simple solution, we show that introducing the target label later in the training dramatically improves transfer.