A Harmonic Basis for Auditing Learned Sequential Biases in LLMs
Ievgen Redko
Abstract
We show that the next-token statistics of any autoregressive language model, observed at a chosen context order $k$, induce a weighted random walk on the de Bruijn graph $G_k$. Using recently derived closed-form eigenvectors of the de Bruijn graph Laplacian, we develop a harmonic analysis of this walk in which every deviation from uniform, unbiased generation is captured, with no residual, by temperature-independent coefficients -- the model's harmonic spectrum -- computable directly from the model's centred logits. Because the decomposition is exact rather than sampled, it offers a lightweight audit for latent regularities that are hard to catch by prompting alone: unintended associations with sensitive tokens, and behavioural predictability an adversary could exploit. We validate both uses. First, the harmonic coordinates show that LLMs process gender tokens through the same sequential statistics as abstract symbols, exposing a quantifiable, non-uniform bias toward one gender token in several models. Second, in Rock-Paper-Scissors the harmonic spectrum decomposes an LLM's exploitable strategy into interpretable components and predicts susceptibility to a deterministic adversary (Pearson $r=-0.89$ across 19 models), with instruction-tuned models consistently more exploitable than their base counterparts. We propose the harmonic spectrum as a cheap, exact, falsifiable audit for unwanted structure in model behaviour.
Chat is not available.
Successful Page Load