Dancing in Fetters: Pareto-Optimal On-Device LLMs under Hardware Constraints
Luoyang Sun ⋅ Jiwen Jiang ⋅ Yifeng Ding ⋅ Fengfa Li ⋅ Yan Song ⋅ Haifeng Zhang ⋅ Lei Ren ⋅ Kun Zhan ⋅ Chen Wei ⋅ Xie Yan ⋅ Jun Wang ⋅ Cheng Deng
Abstract
Recent embodied AI systems increasingly rely on large language models (LLMs) for high-level planning, semantic reasoning, and long-horizon decision making. However, deploying such capabilities on edge platforms such as autonomous vehicles and mobile robots requires balancing model quality against strict latency and hardware constraints. Existing LLM architectures are primarily designed for cloud-scale accelerators and often become inefficient in on-device settings. We present $\textbf{PLAS}$ ($\textbf{P}$areto-optimal $\textbf{L}$LM $\textbf{A}$rchitecture $\textbf{S}$earch), a hardware-aware framework for identifying Pareto-optimal on-device LLM architectures under deployment constraints. Our approach jointly models training loss and inference latency as functions of architectural design choices. We fit an empirical architecture-to-loss model from 170 LLMs trained on 10B tokens each, and estimate latency using roofline-based hardware modeling, enabling efficient exploration of the accuracy-latency Pareto frontier without exhaustively training or profiling every candidate architecture. Using PLAS, we evaluate 1942 candidate architectures on NVIDIA Jetson Orin. At the same measured latency as Qwen2.5-0.5B, our co-designed architecture achieves 19.42\% lower WikiText-2 perplexity. Our results suggest that effective edge deployment requires explicit hardware--model co-design, and that Pareto-based architecture modeling provides a practical alternative to exhaustive neural architecture search for on-device LLMs.
Chat is not available.
Successful Page Load