A Tiny Qualitative Coder with a Large Evaluation Problem
Abstract
AI-assisted qualitative coding often depends on cloud services or GPUs. We study whether a 15.6M-parameter model can instead generate open-code suggestions locally on a CPU. Its 61.7~MiB checkpoint processes one passage in .021 seconds on an M1 Max and achieves nearly the same similarity to human reference codes as the best of nine prompt-only models between 0.27B and 4.7B (.362 vs.\ .364). In a matched five-seed comparison, retaining human-authored rather than machine-generated training codes provides no clear advantage in reference agreement. Yet the two models fail differently: the machine-code model repeats content words in 35\% of its outputs, compared with 1\% for the human-code model, even though repetition is almost absent from its training targets. Experts rate these repetitive outputs lower, whereas embedding similarity does not penalize them. Tiny models can therefore provide fast, local coding suggestions, but adapted larger models remain stronger, and reference agreement alone can hide substantial generation failures.