Mixing Matters: Evidence Position Bias Across Sequence Mixers in Long-Context Question Answering
Abstract
Long-context models often use answer-bearing evidence less effectively when it appears away from a prompt boundary. We compare this position bias across Transformer, state-space, and hybrid Mamba architectures by moving one gold document through a fixed set of nine distractors. The evaluated Transformers show a beginning advantage that the evaluated state-space models lack, while both families favor evidence near the end. A matched comparison of pure and hybrid state-space architectures points in the same direction but remains statistically uncertain. The family difference appears among models that can answer the task. Transformer primacy reappears on synthetic retrieval, where state-space ceiling performance prevents a family comparison. A corpus swap within one state-space architecture changes overall accuracy without a detectable change in curve shape. Prompt changes can create or nearly remove a beginning advantage. Position remains decodable across architectures even when their response curves differ, while attention concentration near initial tokens covaries with beginning sensitivity in the tested Transformer family. These exploratory findings show that effective context use depends on architecture family, capability, task, and prompt design, without identifying the sequence mixer as the sole cause.