When Are Predictions Enough? An Evaluation Protocol for Frozen Expert Composition
Abstract
Practitioners often combine independently trained models as a fixed pool of frozen experts. Standard combination methods, such as voting, averaging, or stacking, rely on final predictions, amounting to a few scalars per expert, and discard potentially richer cross-expert information found in the penultimate representations. An alternative is to combine these models in feature space. Whether this yields a meaningful improvement depends heavily on the evaluation design. Informal assessments on the same data can support materially different conclusions depending on baseline strength, overlap control, reporting granularity, and pool construction. We argue that auditing the sensitivity of evaluation outcomes to these design choices is critical for the effective composition of expert models. To address these sensitivities, we propose a four-part evaluation protocol: (1) a baseline ladder for prediction-space (PS) models; (2) a deterministic feature-space (FS) anchor plus a higher-capacity cross-check; (3) axis-isolated pools to disentangle performance gains; and (4) grouped evaluation with overlap-aware controls. Applying our protocol to 21 frozen deepfake detectors from DF40 (Yan et al., 2024), we find that a 10K-parameter logistic regression on raw concatenated features outperforms the strongest PS baseline by +0.123 AUROC on an architecture-diverse pool (generator-level bootstrap 95% CI [+0.078, +0.174]). Crucially, the higher-capacity cross-check (a 3.5M-parameter feature-space MLP) adds only +0.004 AUROC, suggesting that the gain arises from the richer feature interface rather than from combiner complexity. Prediction-space models suffice when individual expert performance is high but fall short in architecture-diverse, high-bottleneck pools. Our primary contribution is an evaluation-design audit showing that, on this benchmark and expert pool, relaxing any of these controls risks overstating the apparent FS advantage or creating the illusion of a gain that fails to hold across data subgroups. Our findings demonstrate that, without rigorous controls, the same data can support multiple, sometimes spurious, conclusions.