In STeP: Speculative Tensor Parallelism for Concurrent Heterogeneous Inference of LLMs
Viren Luke Radhakrishnan ⋅ Dhruva Kashyap ⋅ Pranav K Nayak ⋅ Chiranjib Bhattacharyya ⋅ Prakash Raghavendra
Abstract
Running dense, open-weight large language models on commercially accessible workstations or single-GPU cloud setups is increasingly desirable due to cost and privacy constraints, particularly on tasks like coding and reasoning. Starting at around 24B parameters, however, they exceed the available GPU VRAM on these setups, forcing reliance on offloading methods that limit throughput, while leaving substantial CPU compute underutilized. Heterogeneous CPU–GPU speculative decoding is the dominant lossless-acceleration framework in this regime, but existing methods are limited by what we refer to as the **heterogeneity gap**: their CPU and GPU do not co-execute, their draft is a separate set of weights rather than a subnetwork of the verifier, and their VRAM footprint cannot scale to fill the available budget. We observe that channel-saliency methods induce an ordering on the transformer’s FFN channels so that the top-$k$ prefix closely approximates the full output, with graceful degradation as $k$ decreases. Building on this observation, we introduce **Speculative Tensor Parallelism (STeP)**, a training-free, self-speculative method that extracts a subnetwork of the verifier as a GPU-resident draft, places the remainder on the CPU, and verifies via concurrent CPU and GPU computation, provably preserving the verifier’s sampling distribution. Across dense models from 24B to 123B, GPU memory budgets from 32 to 192 GB, and five benchmarks, STeP outperforms SpecExec and SubSpec by up to $1.4 \times$ and $1.8 \times$ respectively, surpasses either tensor parallelism or speculative decoding alone, maximally utilizes the available VRAM, and closes the heterogeneity gap.
Chat is not available.
Successful Page Load