An Accuracy and Cost Analysis of Prompting and Fine-Tuning for Small Language Models
Abstract
When small language models (SLMs) are deployed for narrow tasks, adapting them with a small number of labeled samples can substantially improve performance. Three adaptation families compete in this regime: in-context learning (ICL), parameter-efficient fine-tuning, and automatic prompt optimization (APO), yet no prior study compares all three at matched conditions. We compare three ICL selection strategies, fine-tuning using LoRA with per-task hyperparameter optimization, and the APO method GEPA across 17 narrow tasks, three Qwen3 model sizes, and data budgets up to 200 labeled samples. ICL sample efficiency largely saturates by 25–50 examples regardless of selection strategy, while fine-tuning keeps improving across the full budget range. Approximately 50 labeled samples suffice for fine-tuning to overtake all ICL strategies, with the crossover occurring earlier on 0.6B and 1.7B than on 4B. APO is not competitive. A throughput-calibrated two-component cost analysis shows that data availability, not deployment volume, drives the method transition, reducing the adaptation decision to a single variable.