REINS: A Self-Evolving Agent Harness for Real-Time Trajectory Planning
Zhihong Cui ⋅ Hengyu Liu ⋅ Haoran Tang ⋅ shijun liu ⋅ Amir Taherkordi ⋅ Tor Skeie
Abstract
Agent harnesses—runtime scaffolds that wrap an LLM's reasoning loop with tool dispatch, verification, and persistent memory—have become the standard substrate for deploying LLMs as autonomous agents, as demonstrated by Claude Code and Codex in software engineering. In trajectory planning, however, this paradigm has not taken hold: existing LLM-based trajectory planning methods either keep the LLM in the decision loop, exceeding the millisecond-level control budget, or offload it to offline rules or training-time supervision—leaving the runtime planner non-evolving. Intermediate hybrids still lack the verification-and-consolidation loop that makes harnesses reliable. We ask: *what runtime harness would let an LLM planner meet real-time constraints while self-evolving into a deployable one?* We propose **REINS**, a harness built on three LLM–planner couplings: LLM outputs are (i) **grounded** to a learned variable-level dynamics model, (ii) **verified** via calibrated predicates and forward rollout, and (iii) **consolidated** into an indexed skill memory that serves future scenes in $\mathcal{O}(\log n)$ without re-invoking the LLM. The harness closes a self-evolution loop: familiar scenes resolve by memory lookup; novel scenes invoke the LLM, whose verified outputs enrich the memory. Without modifying LLM weights, REINS attains 98%+ compliance, 50%+ collision reduction, and 2–5 ms latency on simulated and real-world benchmarks, with LLM invocation rate monotonically declining as memory matures. Code: https://anonymous.4open.science/r/REINS-1FE3/
Chat is not available.
Successful Page Load