Discarding Answer Ranks Changes the Case for Mixing Biomedical LLMs
Odysseus Kaloidas
Abstract
Should a biomedical reasoning system spend eight calls on one model or several? An aggregation rule can make mixing look better by worsening the one-model reference. A 288-question, four-model study compares mixing with repeated Gemma calls. Counting every shortlisted answer rather than only the first shifts the mixed-minus-Gemma accuracy gap by $9.72$ percentage points (95% confidence interval $[5.29,14.21]$). Using ranks to resolve ties accounts for $7.62$ points; the remaining contrast is unresolved. A favorable rule contrast need not imply a better policy. In a separate prespecified 720-question MedXpertQA test, mixing Gemma and Llama3.3-70B loses $8.95$ points $[-11.16,-6.75]$ against development-selected Gemma first-choice voting. The mixture already has lower per-call accuracy. On 256 MMLU-Pro Health questions, mixing improves inclusion-vote accuracy by $5.00$ points, yet remains far below the selected first-choice policy. On Health, first-choice voting also exceeds a ceiling for selectors restricted to listed answers that use only which calls mention them and favor no answer label. These results show how discarding ranks changes the apparent benefit of mixing, without isolating quality-independent diversity.
Chat is not available.
Successful Page Load