Skip to yearly menu bar Skip to main content


Session

Atlanta Poster Session 6

Hall C1
Sat 12 Dec 8:30 a.m. AEDT — 11:30 a.m. AEDT
Abstract:
Chat is not available.


ABC-Align: Prediction-Powered Alignment with Adaptive Bias Control

Eric Frankel ⋅ Banghua Zhu ⋅ Sewoong Oh ⋅ Lillian Ratliff

Language model post-training is often bottlenecked by the need for human-collected preference data, which is expensive and difficult to scale. RLAIF-style approaches that leverage pseudo labels offer an abundant alternative but introduce systematic biases that degrade downstream alignment. Recent general-purpose semi-supervised methods correct for teacher bias using a small set of human-labeled examples, but suffer from high variance especially when human annotations are scarce. To this end, we propose ABC-Align, leveraging abundant pseudo label signal to minimize variance and applying a lightweight, adaptive correction grounded in the human-labeled subset. The correction strength is tuned automatically during training using plug-in estimates of the relevant bias--variance quantities. On LLM alignment with RLHF and DPO where human feedback is scarce, we empirically demonstrate that ABC-Align achieves superior performance over prior semi-supervised baselines.

We revisit the foundations of fairness and its interplay with utility and efficiency in supervised learning settings where training labels are richer than binary outcomes, such as risk estimates (probabilities), individual types, or rankings. We introduce new notions of accurate classification rates for subgroups in the population, defined by comparing the induced positive-classification rates to those under the Bayes-optimal rich predictor. Our main contributions are computational impossibility results: we show that simultaneously achieving these rate-accuracy guarantees and natural desiderata such as calibration or loss minimization is, in some cases, computationally infeasible, even when training examples are labeled by the Bayes-optimal rich predictor. Unlike prior impossibility results in this area, these desiderata are simultaneously satisfied by the Bayes-optimal predictor, and each can be achieved efficiently in isolation.


Act, Validate, Adapt: Closing the Causal Discovery-Control Loop

Taqiya Ehsan ⋅ Shuren Xia ⋅ Jorge Ortiz

Every control action on a physical system is already a do-operator intervention, yet causal discovery and control are typically treated as separate stages. This separation leaves deployed controllers unable to test whether the causal graph they rely on remains valid under changing operating conditions. We present \textsc{PolicyGRID}, a closed-loop framework that uses actuator interventions to validate candidate causal edges, fits a structural causal model over the validated graph, and reuses that graph for both policy optimization and latent-regime monitoring. In a multi-actuator physical simulation with latent operating regimes (building energy management), interventional graph validation reduces closed-loop energy use by 21\% at a standard comfort target relative to an otherwise identical observation-only graph, while improving four-way latent-context identification from 61.2\% to 80.8\%. A targeted ablation shows that the gains originate in graph pruning. Validation removes 16 of 47 candidate edges; with coefficient training held fixed, the validated graph reduces closed-loop energy relative to the unvalidated candidate set, while refitting coefficients on the same graph does not. On a live physical testbed, the same pipeline runs without code changes, performs 21 physical interventions, discovers 16 edges, and produces a target-responsive control policy where the observation-only baseline is insensitive to the comfort-energy target. These results show that physical controllers can use their own actions to validate causal structure and improve downstream control under latent regime shifts.


Adapting in the Dark: Efficient and Stable Test-Time Adaptation for Black-Box Models

Yunbei Zhang ⋅ Shuaicheng Niu ⋅ Chengyi Cai ⋅ Feng Liu ⋅ Jihun Hamm

Test-Time Adaptation (TTA) for black-box models accessible only via APIs remains a largely unexplored challenge. Existing approaches such as post-hoc output refinement offer limited adaptive capacity, while Zeroth-Order Optimization (ZOO) enables input-space adaptation but faces high query costs and optimization challenges in the unsupervised TTA setting. We introduce BETA (Black-box Efficient Test-time Adaptation), a framework that addresses these limitations by employing a lightweight, local white-box steering model to create a tractable gradient pathway. Through a prediction harmonization technique combined with consistency regularization and prompt learning-oriented filtering, BETA enables stable adaptation with no additional API calls and negligible latency beyond standard inference. BETA achieves a +7.1% accuracy gain on ViT-B/16 (ImageNet-C) and +3.4% on CLIP (ImageNet-S/R), surpassing strong white-box and gray-box methods including TENT and TPT. On a commercial API, BETA achieves comparable performance to ZOO at 250x lower cost while maintaining real-time inference speed, establishing it as a practical and efficient solution for real-world black-box TTA.


AgentAbstain: Do LLM Agents Know When Not to Act?

Xun Liu ⋅ Yi Evie Zhang ⋅ Vira Kasprova ⋅ Parisa Rabbani ⋅ Pardissadat Zahraei ⋅ Tianyu Zhang ⋅ Ali Ebrahimpour-Boroojeny ⋅ Varun Chandrasekaran

Agent systems based on Large Language Models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain. This gap poses real risks: under ambiguity, conflicting constraints, or tool failures, agents may execute unintended and irreversible actions. To close this gap, we present the first systematic evaluation framework for agentic abstention: the ability to recognize when not to act. At its core, AgentAbstain is a paired-task benchmark that defines 8 abstention categories spanning pre-execution reasoning and runtime discovery, with 263 paired tasks across 42 executable sandbox environments, each pairing a should-act task with a should-abstain variant produced by a controlled perturbation to the instruction, tool, or environment state. Scaling such paired evaluations poses two practical challenges: manually authoring diverse tasks is expensive, and static benchmarks risk data contamination as models evolve. To address both, we propose AbstainGen, a fully automated pipeline that synthesizes sandbox environments and generates paired tasks end-to-end, validated by deterministic replay and semantic LLM judges. Its scalable design enables on-demand regeneration of fresh task instances, and human validation rates 96% of generated tasks as well-designed. Evaluating 17 frontier LLMs across 4 agent harnesses, the best model (Gemini 3.1 Pro) achieves only 59.5% paired accuracy (correct on both the act and abstain sides of each paired task). More importantly, abstention capability is largely independent of general task-solving capability, indicating that scaling task-solving alone will not close this gap. We identify failure modes such as Post-hoc Abstention, in which agents execute irreversible actions before recognizing abstention triggers. For instance, an agent may cancel a flight reservation before noticing contradictory rebooking instructions, leaving the user stranded. These findings underscore the need for rigorous abstention evaluation to develop more trustworthy LLM agents.


Agent-Native Research Artifacts

Jiachen Liu ⋅ Jiaxin Pei ⋅ Jintao Huang ⋅ Chenglei Si ⋅ Ao Qu ⋅ Robert Tang ⋅ Runyu Lu ⋅ Lichang Chen ⋅ Xiaoyan Bai ⋅ Haizhong Zheng ⋅ Shirui Chen ⋅ Zhiyang Chen ⋅ Haojie Ye ⋅ Yujuan Fu ⋅ Zexue He ⋅ Zijian Jin ⋅ Zhenyu Zhang ⋅ Shangquan Sun ⋅ Maestro Harmon ⋅ Dianzhuo Wang ⋅ Qian-Ze Zhu ⋅ Jiachen Sun ⋅ Mingyuan Wu ⋅ Baoyu Zhou ⋅ Chenyu You ⋅ Shijian Lu ⋅ Yiming Qiu ⋅ Fan Lai ⋅ Yuan Yuan ⋅ Yao Li ⋅ Junyuan Hong ⋅ Ruihao Zhu ⋅ Beidi Chen ⋅ Alex Pentland ⋅ Ang Chen ⋅ Mosharaf Chowdhury ⋅ Zechen Zhang

Scientific publication compresses a branching, iterative research process into a linear narrative, discarding the majority of what was discovered along the way. This compilation imposes two structural costs: a Storytelling Tax, where failed experiments, rejected hypotheses, and the branching exploration process are discarded to fit a linear narrative; and an Engineering Tax, where the gap between reviewer-sufficient prose and agent-sufficient specification leaves critical implementation details unwritten. Tolerable for human readers, these costs become critical when AI agents must understand, reproduce, and extend published work. We introduce the Agent-Native Research Artifact (ARA), a protocol that replaces the narrative paper with an agent-executable research package structured around four layers: scientific logic, executable code with full specifications, an exploration graph that preserves the failures compilation discards, and evidence grounding every claim in raw outputs. Three mechanisms support the ecosystem: a Live Research Manager that captures decisions and dead ends during ordinary development; an ARA Compiler that translates legacy PDFs and repos into ARAs; and an ARA-native review system that automates objective checks (analogous to a grammar checker for prose) so human reviewers can focus on significance, novelty, and taste. On PaperBench and RE-Bench, ARA raises question-answering accuracy from 72.4% to 93.7% and reproduction success from 57.4% to 64.4%. On RE-Bench's five open-ended extension tasks, preserved failure traces in ARA accelerate progress, but can also constrain a capable agent from stepping outside the prior-run box depending on the agent's capabilities.


Agents' Last Exam

Yiyou Sun ⋅ Xinyang Han ⋅ Weichen Zhang ⋅ Yuanbo Pang ⋅ Tianyu Wang ⋅ Yuhan Cao ⋅ Yixiao Huang ⋅ Christopher Duroiu ⋅ Haoyun Zhang ⋅ Jeffrey Lin ⋅ Weishu Zhang ⋅ Tianren Zeng ⋅ Ying Yan ⋅ Bo Liu ⋅ Hanson Wen ⋅ Mingyang Xu ⋅ Xiaoyuan Liu ⋅ Zimeng Chen ⋅ Weiyan Shi ⋅ Amanda Dsouza ⋅ Vincent Chen ⋅ Yushan Li ⋅ Wenxi Deng ⋅ Huiqi Wang ⋅ Justin Xu ⋅ Tao Sun ⋅ Zhun Wang ⋅ Chris Liu ⋅ Yafei Cheng ⋅ Rongwang Hu ⋅ Aras Bacho ⋅ Shengcao Cao ⋅ Zengyi Qin ⋅ Yixiong Chen ⋅ Hengduan Fan ⋅ Hao Liu ⋅ Lin Zeng ⋅ Shashank M Bharadwaj ⋅ Litian Gong ⋅ Maojia Song ⋅ Ruheng Wang ⋅ Zongzheng Zhang ⋅ Jianhong Tu ⋅ Zhonghua Wang ⋅ Zheng Zhang ⋅ Zijiao Chen ⋅ Yanqiong Jiang ⋅ Zhendong Li ⋅ Bohan Lyu ⋅ Ma Chang ⋅ Peiran Xu ⋅ Benran Zhang ⋅ Shangding Gu ⋅ Haoyue Hua ⋅ Haoyang Li ⋅ Wanzhe Liao ⋅ Chengzhi Liu ⋅ Junbo Peng ⋅ Zechen Xu ⋅ Bo Chen ⋅ Jiayi Cheng ⋅ Yi Jiang ⋅ Maggie Kuang ⋅ Yuan Li ⋅ Youbang Pan ⋅ Ziyan Rao ⋅ Alexander Schubert ⋅ Ivan Shen ⋅ Vincent Siu ⋅ Xiatao Sun ⋅ Xiaopan Zhang ⋅ Yuchen Zhu ⋅ Ishaan S Chandok ⋅ Lei Ding ⋅ Jingxuan Fan ⋅ Andrew Glover ⋅ Jiaming Hu ⋅ Yiran HU ⋅ Wenbo Huang ⋅ Haoran Jin ⋅ Lukas Kim ⋅ Ming Liu ⋅ Alireza Rafiei ⋅ Kunyang Sun ⋅ Ting Sun ⋅ Yingze Wang ⋅ Yixin Wang ⋅ Hanwen Xing ⋅ Sihan Xu ⋅ Yuzheng Xu ⋅ Zhongxing Xu ⋅ Zhiling Yan ⋅ Boqin Yuan ⋅ Ruiqi Zhang ⋅ Yifan Zhang ⋅ Zibo Zhao ⋅ Haoyue Bai ⋅ Carlo Bosio ⋅ Joe Cavanagh ⋅ Tianxing Chen ⋅ Yipu Chen ⋅ CHENYU ZHU ⋅ Chen Dai ⋅ Stefano De Castro ⋅ Yunfu Deng ⋅ Kaustubh Dhole ⋅ Jiayuan Ding ⋅ Chenchen Du ⋅ Zhehang Du ⋅ Hao Fan ⋅ Run-Ze Fan ⋅ Hengyu Fu ⋅ Shi Gu ⋅ Yifan Gu ⋅ Baihe Huang ⋅ Baixiang Huang ⋅ Ran Jin ⋅ Xin Lan ⋅ Joseph Lee ⋅ Deren Lei ⋅ Chenyu Li ⋅ Daofeng Li ⋅ Jingyan Li ⋅ Yi Li ⋅ Yuangang Li ⋅ Zhixu Li ⋅ Yinsheng Li ⋅ Wenyu Liang ⋅ Longtai Liao ⋅ Kevin Qinghong Lin ⋅ Zeyi (Andy) Liu ⋅ Jiaming Liu ⋅ Kaiyuan Liu ⋅ Xuan Liu ⋅ Pan Lu ⋅ Yicheng Lyu ⋅ Qiuyang Mang ⋅ Kyle Montgomery ⋅ Ruoxi Ning ⋅ Jorin Overwiening ⋅ Xu Pan ⋅ Core Francisco Park ⋅ Justin Purnomo ⋅ Scott Rankin ⋅ Bixuan Ren ⋅ HaoYang Shang ⋅ Hejia Shen ⋅ JIAWEI SHEN ⋅ Shi Qiu ⋅ Tianneng Shi ⋅ Jonah So ⋅ Vladislav Susoy ⋅ Haocheng Wang ⋅ Jialu Wang ⋅ Wei Wang ⋅ Xinyu J Wang ⋅ Zehao Wang ⋅ Ruoxi Wu ⋅ Dehao Wu ⋅ Fangyu Wu ⋅ Mengyuan Wu ⋅ Yu Wu ⋅ Yuchen Wu ⋅ Yuhao Wu ⋅ Qingpo Wuwu ⋅ Weihang Xiao ⋅ Yongyi Xiong ⋅ Fan Xu ⋅ Ruiling Xu ⋅ Mingxuan Yan ⋅ Benjamin Yang ⋅ Jirong Yang ⋅ Xiaoli Yang ⋅ Yushi Yang ⋅ Haoran Ye ⋅ Xiaohu Yu ⋅ Zhengming Yu ⋅ Chenlong Zhang ⋅ Chi Zhang ⋅ Hanning Zhang ⋅ Hanwen Zhang ⋅ Junge Zhang ⋅ Kunpeng Zhang ⋅ Song Zhang ⋅ Wenjin Zhang ⋅ Wenshuo Zhang ⋅ Qijian Zhao ⋅ Yimin Zhao ⋅ Yuhaohua Zheng ⋅ Liwei Zhou ⋅ Tianyue Zhou ⋅ Sichen Zhu ⋅ Yan Zhu ⋅ Yishu Zhu ⋅ Jerry Zuo ⋅ Chonghao Cai ⋅ Helena Casademunt ⋅ Cheng Cheng ⋅ Nawen Deng ⋅ Rao Fu ⋅ Yifan Han ⋅ He Ren ⋅ Zhenyu He ⋅ Qiao Jin ⋅ Langlang Li ⋅ Yuetai Li ⋅ Lu Lu ⋅ Luqing Zhou ⋅ Subhabrata Mukherjee ⋅ Yunqi Ouyang ⋅ Yin Ren ⋅ Dawei Shi ⋅ Haoran Wu ⋅ Zhiyue Wu ⋅ Hannah Yao ⋅ Zhuoran Yi ⋅ Jenny Yu ⋅ Rhea Zhan ⋅ Hang Zhou ⋅ Shiqiao Zhu ⋅ Junfan Zhu ⋅ Zhen Dong ⋅ Ali Emami ⋅ Yanjun Gao ⋅ Haimin Hu ⋅ Erika J Schneider ⋅ Zhenglu Li ⋅ Jiachen Li ⋅ Lihui Liu ⋅ Murphy Niu ⋅ Yi Shao ⋅ Wenqi Shi ⋅ Lichao Sun ⋅ Jianxin Sun ⋅ Chenguang Wang ⋅ Xuhai "Orson" Xu ⋅ Huaxiu Yao ⋅ Aylin Caliskan ⋅ Jing Huang ⋅ Yang Liu ⋅ Mark W Mueller ⋅ Russell Poldrack ⋅ Sanjiv R Das ⋅ Zhiyong Lu ⋅ Radha Poovendran ⋅ Somayeh Sojoudi ⋅ Costas J Spanos ⋅ Molei Tao ⋅ Mikko Tolonen ⋅ Ting Wang ⋅ Alan Yuille ⋅ Patrick Bryant ⋅ Carl Boettiger ⋅ Arvind Rao ⋅ Tapio Schneider ⋅ Georgios N Yannakakis ⋅ Laure Zanna ⋅ Ida Sim ⋅ Tarek Zohdi ⋅ Jack Gallant ⋅ Teresa Head-Gordon ⋅ Dawn Song

Recent AI systems achieve strong results on many benchmarks, yet these gains have not translated into economically meaningful deployment across professional domains. We argue that this gap is largely an evaluation problem: existing benchmarks rarely measure sustained performance on real, economically valuable workflows. We introduce Agents’ Last Exam (ALE), a benchmark for evaluating generalist computer-use AI agents on long-horizon professional tasks with verifiable outcomes. ALE is developed in collaboration with 250+ industry experts and organized around a taxonomy grounded in O*NET / SOC 2018. The benchmark spans 55 subdomains across 13 industry clusters and currently contains 1,490 task instances sourced from authentic professional workflows. Tasks require integrated GUI interaction, shell execution, software operation, and long-horizon planning inside real computing environments. To support scalable evaluation, ALE uses deterministic deliverable-based scoring and structured rubric verification rather than open-ended human judgment whenever possible. Experimental results show that current frontier agents remain far from saturation: across mainstream harness and backbone configurations, the average full-pass rate on the hardest tier is only 2.6%. More broadly, ALE is intended not merely as another leaderboard, but as an instrument for measuring the gap between benchmark success and GDP-relevant impact.


AI Agents Push Humans Out of the Loop

Margaret Mitchell ⋅ Avijit Ghosh ⋅ Samir Passi

AI agents pose significant risks as they are granted increasing autonomy. A commonly proposed solution is human oversight and keeping a ``human in the loop”, but this is not a simple solution: Not only do current approaches to AI agent design impede effective human oversight, but the cognitive capacities required for it are also themselves degraded by extended use of AI systems. This position paper argues that current approaches to the development and deployment of AI agent systems do not support effective human oversight -- they contribute to its degradation. To address this, a top priority in the advancement of AI agents should be supporting the situated goals and cognitive requirements of effective human oversight, treating the human needs of overseers at the same level of importance as AI agent capability. To put this idea into practice, we connect work on automation and human-computer interaction to AI agent processes, outlining design-level affordances and organizational protocols that (1) support overseers in exercising critical judgement and (2) counteract the skill atrophy that arises from extended use of automation. We urge developers and deployers to adopt these or similar approaches. Without explicit support for the cognitive demands of effective human-agent interaction, AI agent systems will continue to passively incentivize


AI GAMESTORE: Scalable, Open-Ended Evaluation of Machine General Intelligence with Human Games

Lance Ying ⋅ Ryan Truong ⋅ Prafull Sharma ⋅ Kaiya Zhao ⋅ Nathan Cloos ⋅ Kelsey Allen ⋅ Tom Griffiths ⋅ Katie Collins ⋅ Jose Hernandez-Orallo ⋅ Phillip Isola ⋅ Samuel J Gershman ⋅ Josh Tenenbaum

Rigorously evaluating machine intelligence against the broad spectrum of human general intelligence has become increasingly important and challenging in this era of rapid technological advance. Conventional AI benchmarks typically assess only narrow capabilities in a limited range of human activity. Most are also static, quickly saturating as developers explicitly or implicitly optimize for them. We propose that a more promising way to evaluate human-like general intelligence in AI systems is through a particularly strong form of general game playing: studying how and how well they play and learn to play all conceivable human games, in comparison to human players with the same level of experience, time, or other resources. We define a “human game” to be a game designed by humans for humans, and argue for the evaluative suitability of this space of all such games people can imagine and enjoy — the “Multiverse of Human Games”. Taking a first step towards this vision, we introduce the AI GAMESTORE, a scalable and open-ended platform that uses LLMs with humans-in-the-loop to synthesize new representative human games, by automatically sourcing and adapting standardized and containerized variants of game environments from popular human digital gaming platforms. As a proof of concept, we generated 100 such games based on the top charts of Apple App Store and Steam, and evaluated seven frontier vision-language models (VLMs) on short episodes of play. The best models achieved less than 10% of the human average score on the majority of the games, and especially struggled with games that challenge world-model learning, memory and planning. We conclude with a set of next steps for building out the AI GAMESTORE as a practical way to measure and drive progress toward human-like general intelligence in machines.

We introduce HackerSignal, a benchmark for temporal out-of-distribution cyber threat intelligence (CTI) and cross-source CVE linkage. HackerSignal aggregates 7.45 million exact-deduplicated documents from 64 public forum/source identifiers spanning eight source layers and a 36-year window (1990-2026). In contrast to other publicly accessible cybersecurity datasets, HackerSignal is among the first public benchmark datasets that maps the full potential exploit to vulnerability trajectory from hacker community discourse, exploit databases with working and proof of concept exploits, vulnerability advisories, and software fix commits. HackerSignal creates these linkages through a shared CVE identifier space while preserving source-specific release modes to support a range of unique Artificial Intelligence (AI)-enabled cybersecurity analytics tasks. In this paper, we summarize HackerSignal and illustrate three selected benchmark tasks it uniquely supports: (1) CVE linkage retrieval (cross-source temporal out-of-distribution entity grounding); (2) exploit type classification (8-class vulnerability type prediction with temporal OOD evaluation); and (3) temporal generalization (prospective CVE-disjoint evaluation where $C_{\text{train}} \cap C_{\text{test}} = \emptyset$). All tasks use temporal splits to evaluate prospective generalization. We release source-shortcut and leakage diagnostics; manual-audit packets; a datasheet; and a release-governance addendum to support the dissemination of the dataset. HackerSignal's code, data, and Croissant metadata are available at hf.co/datasets/DatasetSubmission/HackerSignal.


Align-RAG: Alignment Is All You Need for TSFM In-Context Learning

Mohammad Asadi ⋅ Soheil Hor ⋅ Bardiya Akhbari ⋅ Jack W O'Sullivan ⋅ Tahoura Nedaee ⋅ Layne Price ⋅ Raviteja Anantha Ramesh ⋅ Euan Ashley ⋅ Ehsan Adeli

Retrieval-augmented forecasting promises to adapt frozen Time Series Foundation Models (TSFMs) to new domains without fine-tuning, but recent methods typically rely on learned fusion modules, i.e., trained adapters that merge retrieved examples into the backbone's forecast, based on the assumption that frozen backbones cannot dynamically incorporate retrieved context on their own. We show this assumption is unnecessary. We introduce $\textbf{Align-RAG}$, a training-free method that applies a closed-form per-pair amplitude rescaling and integer-lag phase shift to retrieved past--future windows before they enter a frozen backbone's context. With no learned parameters, Align-RAG outperforms the state-of-the-art trained retrieval adapter on a frozen Chronos-Bolt on all seven datasets of the standard benchmark (avg $-3.75\%$ MSE), showing that the gains previously attributed to learned fusion are recoverable without any training. Align-RAG further improves zero-shot MSE on four additional frozen TSFMs with various architectures by $2.5\%$ to $13.7\%$ per backbone with no per-backbone tuning. To probe why alignment helps, we compare the frozen backbone's prediction shift under aligned demonstrations to the closed-form ridge prediction shift on the same pairs. We find that aligned demonstrations induce prediction shifts that track a closed-form ridge predictor on the same pairs, with a future-shuffle control ruling out a futures-averaging account. Together, these results indicate that frozen TSFMs already support dynamic in-context use of retrievals, and that closed-form alignment should be the default baseline for retrieval-augmented forecasting before any fusion module is trained. Code available at: \url{https://anonymous.4open.science/r/phase-rag-4D92}


A Multi-Fidelity Control Variate Approach for Policy Gradient Estimation

Xinjie Liu ⋅ Cyrus Neary ⋅ Kushagra Gupta ⋅ Wesley Suttle ⋅ Christian Ellis ⋅ Ufuk Topcu ⋅ David Fridovich-Keil

Many reinforcement learning (RL) algorithms are impractical for deployment in operational systems or for training with computationally expensive high-fidelity simulations, as they require large amounts of data. Meanwhile, low-fidelity simulators—such as reduced-order models, heuristic reward functions, or generative world models—can cheaply provide useful data for RL training, even if they are too coarse for direct sim-to-real transfer. We propose multi-fidelity policy gradients (MFPGs), an RL framework that mixes a small amount of data from the target environment with a control variate formed from a large volume of low-fidelity simulation data to construct an unbiased, variance-reduced estimator for on-policy policy gradients. We instantiate the framework by developing a practical, multi-fidelity variant of the classical REINFORCE algorithm. We show that under standard assumptions, the MFPG estimator guarantees asymptotic convergence of multi-fidelity REINFORCE to locally optimal policies in the target environment, and achieves faster finite-sample convergence rates compared to training with high-fidelity data alone. We evaluate the MFPG algorithm across a suite of simulated robotics benchmark tasks in scenarios with limited high-fidelity data but abundant off-dynamics, low-fidelity data. In our baseline comparisons, for scenarios where low-fidelity data are neutral or beneficial and dynamics gaps are mild to moderate, MFPG is, among the evaluated off-dynamics RL and low-fidelity-only approaches,the only method that consistently achieves statistically significant improvements in mean performance over a baseline trained solely on high-fidelity data. When low-fidelity data become harmful, MFPG exhibits the strongest robustness against performance degradation among the evaluated methods, whereas strong off-dynamics RL methods tend to exploit low-fidelity data aggressively and fail substantially more severely. An additional experiment in which the high- and low-fidelity environments are assigned anti-correlated rewards shows that MFPG can remain effective even when the low-fidelity environment exhibits reward misspecification. Thus, MFPG not only offers a reliable and robust paradigm for exploiting low-fidelity data, e.g., to enable efficient sim-to-real transfer, but also provides a principled approach to managing the trade-off between policy performance and data collection costs.


A Multimodal Benchmark for Evaluating Cause-of-Death Inference Using Child Health and Mortality Data

Junhe Yang ⋅ Soumyakanti Pan ⋅ Hyun Seung Lim ⋅ YUE CHU ⋅ Yuting Guo ⋅ Nishtha Agarwal ⋅ Varun Babbar ⋅ Gaurav Rajesh Parikh ⋅ Yiqun Chen ⋅ Chris A Rees ⋅ Ziyaad Dangor ⋅ Sanjay G Lala ⋅ Zehang Li ⋅ Samuel J Clark ⋅ Zhenke Wu ⋅ Abhirup Datta ⋅ Li Liu ⋅ Cynthia Rudin ⋅ Samuel Scarpino ⋅ Benjamin M Gyori ⋅ Tyler H. McCormick

Accurately attributing causes of death is vital for global health, yet fewer than 5% of deaths in resource-constrained regions are medically certified. To assign causes to these unlabeled deaths at scale, practitioners traditionally rely on verbal autopsy, using supervised statistical models to classify based on structured survey data. However, modern mortality surveillance increasingly collects rich, unstructured multimodal data, such as free-text caregiver narratives and postmortem diagnostics, which traditional supervised statistical models struggle to seamlessly integrate. In this paper, we present a comprehensive, multimodal benchmark for cause-of-death classification using data from the Child Health and Mortality Prevention Surveillance (CHAMPS) network, a unique surveillance platform spanning nine countries across South Asia and Sub-Saharan Africa. Using this dataset, we introduce an evaluation framework designed to rigorously assess diagnostic reasoning, moving beyond traditional metrics that fail to capture complex clinical realities. We demonstrate the utility of this benchmark by evaluating zero-shot large language models against supervised baselines across various data modalities. Our results reveal distinct differences in how these modeling approaches synthesize unstructured medical evidence. This benchmark provide a rigorously defined resource for assessing clinical reasoning in next-generation mortality surveillance.


ANCRe: Adaptive Neural Connection Reassignment for Efficient Depth Scaling

Yilang Zhang ⋅ Bingcong Li ⋅ Niao He ⋅ Georgios Giannakis

Scaling network depth has been a central driver behind the success of modern foundation models, yet recent investigations suggest that deep layers are often underutilized. This paper revisits the default mechanism for deepening neural networks, namely residual connections, from an optimization perspective. On a motivating example, rigorous analysis proves that the layout of residual connections can fundamentally shape convergence behavior, and even induces an exponential gap in convergence rates. Prompted by this insight, we introduce adaptive neural connection reassignment (ANCRe), a principled and lightweight framework that parameterizes and learns residual connectivities from the data. ANCRe adaptively reassigns residual connections with negligible computational and memory overhead (<1%), while enabling more effective utilization of network depth. Extensive numerical tests across pre-training of large language models, diffusion models, and deep ResNets demonstrate consistently accelerated convergence, boosted performance, and enhanced depth efficiency over conventional residual connections.


A Principled Self-Referenced Early Stopping Approach for Deep Image Prior

Chaoyan Huang ⋅ Cheng-Han Huang ⋅ Ismail Alkhouri ⋅ Rongrong Wang

Recently, Deep Image Prior (DIP) has demonstrated strong capabilities for solving inverse imaging problems (IIPs) by optimizing a randomly initialized convolutional neural network in a training-data-free regime. However, DIP suffers from overfitting to noisy measurements due to network over-parameterization, making early stopping (ES) essential. The most successful ES method tracks fluctuations in the running variance of the network output to detect overfitting. However, in many applications, these fluctuations may appear prematurely, leading to unstable reconstructions. In this paper, we first show that nearly optimal DIP early stopping can be achieved when two independent noisy copies of the degraded image are available. Motivated by this observation, and since obtaining two fully independent copies is infeasible, we propose an overfitting detection framework based on constructing pseudo self-referenced images, resulting in three IIP-specific algorithms. Our approach is further supported by theoretical results on single-reference validation, pseudo-validation estimation, and the impact of shared noise. Across different IIPs, ranging from natural image restoration to medical image reconstruction, and under varying noise levels and noise types, our methods consistently outperform existing DIP early stopping approaches, all without requiring knowledge of the measurement noise characteristics.


A Process-Level Evaluation of LLM Discovery Agents

Haorui Wang ⋅ Yuanqi Du ⋅ Hao Cheng ⋅ Baolin Peng ⋅ Chao Zhang ⋅ Weizhu Chen ⋅ yelong shen

Large language models (LLMs) are increasingly used as \emph{discovery agents} that iteratively edit codes against evaluators on tasks from mathematical optimization, algorithmic problem-solving and scientific discovery. Nevertheless, current evaluation reports only a single final score per (model, task) pair under a fixed agent harness: the surrounding scaffolding that governs feedback, memory, reflection, and revision. This final-score view obscures how agents search, where they fail, and which harness mechanisms are actually useful for different classes of tasks. In this paper, we introduce a four-level ladder of agent harnesses and a process-level evaluation framework that measures behavioral profiles and intermediate milestones, alongside the final performance. Using the evaluation workflow, we run a fully crossed evaluation of six contemporary LLMs on $24$ curated discovery tasks across four harness configurations, producing $1{,}728$ trajectories. Our analysis reveals two notable findings. First, harness effectiveness is highly problem-dependent. Second, models can exhibit distinct search behaviors even when their final scores are similar. We release all the trajectories, behavioral profiles, capability rubrics, and harness implementations as reusable evaluation resources for the community.


Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist Rewards

Yuanhao Ban ⋅ Tong Xie ⋅ Sohyun An ⋅ Yunqi Hong ⋅ Evan Frick ⋅ I-Hung Hsu ⋅ Wei-Lin Chiang ⋅ Ion Stoica ⋅ Cho-Jui Hsieh

Faithfulness—how precisely a generated image aligns with its prompt—is increasingly central to the real-world utility of text-to-image (T2I) models. Existing faithfulness benchmarks, however, rely on simple atomic instructions, where top-tier systems already achieve near-perfect scores. As T2I models enter creative workflows, users increasingly issue multi-faceted requests involving intricate spatial relationships, stylistic constraints, and complex text rendering. In this setting, a single binary VLM-judge score no longer reveals which specific constraints the model fails to satisfy. We introduce Arena-T2I Hard, a 310-prompt stress benchmark drawn from real arena T2I logs, with approximately 30 decomposed yes/no constraints per prompt spanning six categories, including text rendering. The strongest closed-source system we evaluate reaches 0.855, with a 33 pp performance gap across 11 systems, demonstrating substantial discriminative power. Moreover, high public-arena rankings fail to predict faithfulness, confirming that holistic Bradley–Terry (BT) preference scores prioritize aesthetics over fine-grained prompt adherence. We propose a dependency-aware checklist reward that decomposes each prompt into a DAG of yes/no questions and zeroes descendants of failed parents, turning faithfulness into a per-constraint training signal. Combined with a BT aesthetic reward via group-decoupled normalization (GDPO), which standardizes each reward within its rollout group so neither collapses, our recipe attains a strictly better faithfulness–aesthetics trade-off on SD3.5-Medium and FLUX.1-dev under MMRB2 pairwise comparisons than every single-reward, naive weighted-sum, or 4-reward BT-ensemble baseline. We release Arena-T2I Hard as a public stress benchmark.


A Scalable Measure of Loss Landscape Curvature for Analyzing the Training Dynamics of LLMs

Dayal Singh Kalra ⋅ Jean-Christophe Gagnon-Audet ⋅ Andrey Gromov ⋅ Ishita Mediratta ⋅ Kelvin Niu ⋅ Alexander Miller ⋅ Michael Shvartsman

Understanding the curvature evolution of the loss landscape is fundamental to analyzing the training dynamics of neural networks. The most commonly studied measure, Hessian sharpness ($\lambda_{\max}^H$) \textemdash the largest eigenvalue of the loss Hessian \textemdash determines local training stability and interacts with the learning rate throughout training. Despite its significance in analyzing training dynamics, direct measurement of Hessian sharpness remains prohibitive for Large Language Models (LLMs) due to high computational cost. We analyze \emph{critical sharpness} ($\lambda_c$), a computationally efficient measure requiring fewer than $10$ forward passes given the update direction $\Delta \bm{\theta}$. Critically, this measure captures well-documented Hessian sharpness phenomena, including progressive sharpening and Edge of Stability. Using this measure, we provide the first demonstration of these sharpness phenomena at scale, up to $7$B parameters, spanning both pre-training and mid-training of OLMo-2 models. We further introduce \emph{relative critical sharpness} ($\lambda_c^{1\to 2}$), which quantifies the curvature of one loss landscape while optimizing another, to analyze the transition from pre-training to fine-tuning and guide data mixing strategies. Critical sharpness provides practitioners with a practical tool for diagnosing curvature dynamics and informing data composition choices at scale. More broadly, our work shows that scalable curvature measures can provide actionable insights for large-scale training.


A second order regret bound for NormalHedge

Yoav S Freund ⋅ Nicholas Harvey ⋅ Victor S. Portella ⋅ Yabing Qi ⋅ Yu-Xiang Wang

We consider the problem of prediction with expert advice for ``easy'' sequences. We show that a variant of NormalHedge enjoys a second-order $\epsilon$-quantile regret bound of $ O\big(\sqrt{V_T \log(V_T/\epsilon)}\big) $ when $V_T > \log N$, where $V_T$ is the cumulative second moment of instantaneous per-expert regret averaged with respect to a natural distribution determined by the algorithm. The algorithm is motivated by a continuous time limit using Stochastic Differential Equations. The discrete time analysis uses self-concordance techniques.


AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs

Haizhong Zheng ⋅ Yizhuo Di ⋅ Jiahui Wang ⋅ Shuowei Jin ⋅ Xueshen Liu ⋅ Yongji Wu ⋅ Zhuoqing Morley Mao ⋅ Ion Stoica ⋅ Jiawei Zhao ⋅ Beidi Chen

Reinforcement learning (RL) is increasingly used to improve the reasoning, coding, and tool-use capabilities of large language models, but agentic RL remains prohibitively expensive. Scaling RL to agentic LLMs requires supporting complex workloads, including multi-policy collaborative training, while efficiently using elastic, heterogeneous, and cross-region compute resources. Existing LLM RL systems support some of these capabilities, but each new extension often requires dedicated system engineering. This burden arises from trainer-centered control architectures and the lack of principled abstractions for RL system components. To address these limitations, we propose AstraFlow, a dataflow-oriented RL system that replaces conventional trainer-centered control with principled component abstractions. In AstraFlow, rollout services, dataflow management, and training are decoupled into autonomous components, enabling the system to natively support complex multi-policy agentic RL workloads and efficiently exploit diverse compute resources. We evaluate AstraFlow across math, code, search, and AgentBench workloads, showing that the same system supports multi-policy training, elastic scaling, heterogeneous cross-region execution, and composable data algorithms without system-level code changes. In multi-policy collaborative training, AstraFlow achieves comparable or better accuracy than existing RL systems while speeding up training time by 2.7x.


AsymHP: Load-Balanced Sparse Attention for Video Diffusion Transformers

Xinwei Qiang ⋅ Yue Guan ⋅ Ruihan Zhu ⋅ Mihir Jagtap ⋅ Zaifeng Pan ⋅ Zhongkai Yu ⋅ Chang Chen ⋅ Zhengding Hu ⋅ Yufei Ding ⋅ Adnan Aziz

Video diffusion transformers rely on self-attention over long spatio-temporal token sequences, making inference expensive even when sparse attention reduces the number of computed blocks. This paper studies a systems bottleneck that appears in dynamic sparse video attention: per-head sparsity can vary widely across denoising steps, layers, and prompts, so equal-head parallel execution leaves some GPUs waiting for dense heads assigned to other GPUs. We propose AsymHP, a runtime system that uses the previous denoising step's per-head sparse density to place a non-uniform number of heads on each GPU. AsymHP combines this lightweight online density estimate, a cost model for mask construction and sparse attention computation, and an asymmetric head redistribution primitive that avoids padding variable-size head shards into symmetric collectives. In our evaluation on H100 GPUs, AsymHP improves sparse attention latency by up to $1.54\times$ without changing the underlying sparse-attention algorithm.


Asymmetric Phase Coding Audio Watermarking

Guang Yang ⋅ Fengchen Liu ⋅ Amir Ghasemian ⋅ Zhong Wang ⋅ Ninareh Mehrabi ⋅ Homa Hosseinmardi

The proliferation of deepfake audio challenges voice-based authentication systems; passive forensic detectors are sensitive to evolving generative models and to real-world channel distortions. We propose Asymmetric Phase Coding (APC), a training-free cryptographic signing layer for audio, designed as a compact and auditable provenance primitive that can stand alone or be stacked with learned watermarks. APC combines Ed25519 digital signatures (EdDSA, FIPS 186-5; 64-byte signatures) with Reed–Solomon error correction, pseudo-random STFT phase-bin selection, and a redundant quantization-index-modulation (QIM) code on log-magnitude differences of adjacent bin pairs, yielding a compact, non-repudiable, blind-extractable watermark. We evaluate APC on 1,000 LibriSpeech test-clean clips (10 s each, 44.1 kHz) under eight attack configurations – identity, 10% end-cropping, 20% end-cropping, 8 kHz low-pass, 16 kHz round-trip resampling, FLAC re-encoding, MP3 at 128 kbps, and OGG-Vorbis at 128 kbps – and achieve cryptographic verification rates between 97.5% and 98.3% on every condition at mean PESQ=3.02 and tens-of-milliseconds CPU latency. We explicitly compare APC against recent neural baselines (AudioSeal, WavMark, SilentCipher), detail the threat model (forgery resistance vs. erasure), characterize the dataset, define all metrics, quantify an adaptive white-box erasure attack, and release code, keys, and metadata for reproducibility.


A Topological Encoder Decoder Framework for Temporal Graph Learning

Ronan Buck ⋅ Kiarash Shamsi ⋅ Tran Gia Bao Ngo ⋅ Astrit Tola ⋅ Baris Coskunuzer ⋅ Cuneyt Akcora

Most temporal graph learning methods reduce future prediction to discriminative inference over past interactions, focusing on edges or node labels for a largely persistent set of nodes. This view breaks down when new nodes and edges appear, since the future graph becomes a new object with changing size, composition, and structure. We address this gap by framing temporal graph prediction as an inverse topology problem. Instead of predicting edges directly, we first predict a multiscale topological descriptor of the future graph and then reconstruct a plausible future snapshot that realizes this descriptor under inductive constraints. This approach makes global structure and node churn explicit prediction targets and produces forecasted graphs on which downstream tasks can be evaluated without retraining. Across 14 temporal graph datasets, we evaluate TopoGED on node, link and graph property prediction and compare it against state-of-the-art temporal graph models. TopoGED achieves a significant improvement in node-level forecasting accuracy over the strongest baseline, with a macro average of 0.49 versus 0.05. It also outperforms baselines in 60% of graph-structure metric evaluations and yields a substantial increase in macro-average edge prediction metrics, from near-zero to 0.12. Our results show that topology-guided graph forecasting can predict inductive future snapshots whose structure supports multiple downstream evaluations.


Attention Sinks and Outliers in Attention Residuals

Haozheng Luo ⋅ Haoran Dai ⋅ Shaoyang Zhang ⋅ Xi Chen ⋅ Eric Hanchen Jiang ⋅ Yijiang Li ⋅ Ching-Yuen Huang ⋅ Chenghao Qiu ⋅ Chenwei Xu ⋅ Zhenyu Pan ⋅ Haotian Zhang ⋅ Binghui Wang ⋅ Yan Chen

We propose OASIS, an outlier- and sink-aware technique built on inter-layer null signaling. As AttnResidual architectures introduce an additional depth-wise normalization channel, they improve inter-layer routing flexibility but also exacerbate attention sinks, activation outliers, and the resulting degradation in inference stability and quantization robustness. OASIS addresses this issue by introducing a $\mathop{\rm{Softmax}}_1$-based null space and coupling token-level null evidence to depth routing through an inter-layer null signal, thereby reducing sink-dominated routing and improving structural robustness. Theoretically, we show that the dual-normalization design of AttnResidual intensifies sink formation and quantization brittleness. Experimentally, we compare OASIS against five baselines on three real-world datasets and observe consistent improvements in both attention sink and post-quantization performance. Notably, OASIS achieves an average reduction of 9.26\% in maximum infinity norm and 2.60\% in average kurtosis across the evaluated settings, while lowering perplexity by 75.85\% under W8A8 and improving GSM8K Pass@1 by 12.42\% under W4A4.


Auditing Cross-Lingual Fairness in Language Model Watermarking

Alexander Nemecek ⋅ Osama Zafar ⋅ Debargha Ganguly ⋅ Vikash Singh ⋅ Vipin Chaudhary ⋅ Erman Ayday

Watermarking schemes for large language model output are evaluated almost exclusively on English text using each scheme's detection threshold and a narrow set of quality measurements. Multilingual deployment exposes evaluation-design choices that are inconsequential on English but determine conclusions cross-lingually. We propose an evaluation framework with four components: detection thresholds calibrated empirically per deployment context, a threshold-independent companion measurement that distinguishes calibration failures from detection failures, three disjoint quality measurement paradigms (distributional, paired-semantic, and reference-perplexity), and a generalized-entropy decomposition of cross-language disparity over a typological family partition. Applied to six watermarking schemes, three open-weight generators, eleven languages spanning four scripts and eight typological families, and both base and instruction-tuned regimes, the framework reveals failure modes that single-language single-paradigm evaluation cannot surface. Across detection and quality, observed disparity is predominantly between-family on the typological partition, indicating that cross-lingual fairness gaps in watermarking are structural to language properties rather than idiosyncratic to particular languages.


Auditor-Assisted Summary-Channel Verification for Hosted LLM Identity Substitution

Ziyi Zhang ⋅ Ziyao Wang ⋅ Guoheng Sun ⋅ Ang Li ⋅ Jian Li

Hosted LLM services increasingly expose model identity as a product claim, but external users and auditors often cannot inspect which model configuration served a request. We study an auditor-assisted regime in which a trusted auditor enrolls claimed identities and prepares verification material before deployment, while audit-time verification uses only the claimed identity and released output, without runtime access to model internals, inference logs, routing decisions, or provider metadata. We propose SumMark, a *summary-channel watermarking* framework that uses user-visible reasoning summaries as a keyed audit surface. During setup, SumMark constructs identity-specific vocabulary carriers, trains lightweight summary adapters, and calibrates per-identity thresholds. At deployment time, a keyed statistical test assesses whether the released summary is consistent with the claimed identity and, when needed, attributes the most likely enrolled identity. Across model families, scales, and task domains, SumMark reliably detects cross-family and cross-scale substitutions at low false-positive rates while preserving task-facing answer quality. We further study FP16$\rightarrow$INT8/INT4 precision-consistency diagnostics and robustness under bounded post-processing, rewriting stress tests, and text-only spoofing, revealing both strong performance and clear failure boundaries.


Automated Hypothesis Discovery for Characterizing Annotation Disagreement

Ankita Gupta ⋅ Alexandra Chouldechova ⋅ Alex Dow ⋅ Miro Dudik ⋅ Jean Garcia-Gathright ⋅ Hanna Wallach ⋅ Nicholas Pangakis

Annotation disagreement is common in multi-annotated datasets used throughout the AI/ML lifecycle, from model training to evaluation. However, existing training and evaluation practices often collapse annotations into ``ground truth'' labels that treat disagreement as noise. In socially consequential contexts, annotation disagreement can reflect multiple valid perspectives, so it is especially important to understand when and why annotators---be they human or automated---disagree. Existing disagreement analysis methods rely on manually created taxonomies of disagreement sources that are costly to produce and difficult to scale. We introduce CHARDIS, a two-stage automated disagreement analysis method that (1) generates candidate disagreement hypotheses—natural language statements describing textual features that predict disagreement—from a multi-annotated dataset and (2) validates these candidates on a held-out dataset, retaining only statistically significant predictors of disagreement and ranking them by predictive performance. CHARDIS supports scalable, interpretable disagreement analysis across diverse annotation tasks. On semi-synthetic datasets, CHARDIS recovers known disagreement sources. On four real-world datasets, CHARDIS reproduces existing manually created taxonomies of disagreement sources and discovers new disagreement hypotheses. In both settings, CHARDIS achieves better predictive performance than baseline methods. Code to reproduce our experiments is available at: https://anonymous.4open.science/r/disagreement-analysis-pipeline-6319


Automatically Refining Coding Rules for AI Coding Agents

Zhengyuan Jiang ⋅ Reachal Wang ⋅ Yuepeng Hu ⋅ Yupu Wang ⋅ Yuqi Jia ⋅ Neil Gong

The performance of AI coding agents is highly dependent on their underlying coding rules. However, existing coding rules are typically hand-crafted, making the process labor-intensive and often suboptimal. In this work, we propose RuleRefine, a framework for automatic coding rule refinement. RuleRefine maintains a pool of candidate coding rules and iteratively improves them. In each iteration, it employs an LLM-powered mutator module to generate variants from existing candidates, and then uses a judge module to evaluate these variants and update the pool with the best-performing ones. Extensive evaluations across two coding-agent frameworks, four backbone LLMs, and three benchmarks demonstrate that RuleRefine outperforms both manual engineering and existing prompt optimization baselines in terms of functional correctness of the generated code, code length, and/or generation cost (e.g., tokens used).


AVI-HT: Adaptive Vision-IMU Fusion for 3D Hand Tracking

Ziyi Kou ⋅ Ankit Kumar ⋅ Mia Huang ⋅ Taylor Niehues ⋅ Vatsal Mehta ⋅ Ergys Ristani ⋅ Li Guan

We present AVI-HT, an adaptive visual-IMU fusion approach for tracking 3D hand poses by jointly modeling the egocentric image with on-glove 6-DoF IMU signals. AVI-HT achieves significantly improved accuracy and availability, particularly in hand-object interaction (HOI) scenarios involving heavy visual occlusion. Two complementary ingredients underpin its success: (1) synchronized multi-modal training data pairing on-body vision-IMU sensor streams with ground-truth 3D hand poses from a motion-capture system, and (2) a cross-sensor deep attention mechanism that adaptively modulates the trust assigned to the vision and individual IMU sensors. To evaluate AVI-HT in real-world settings, we conduct extensive experiments on our DexGloveHOI dataset that consists of 100K+ pairwise vision-IMU samples with synchronized 3D annotated poses, in which users manipulate a variety of objects during daily tasks. We compare against multiple single- and multi-modal tracking approaches under two hand models (UmeTrack, MANO). The results show that AVI-HT reduces mean keypoint error by 16.1% and its wrist-aligned variant by 24.2% over the baselines. Ablation studies further reveal the per-finger contribution of IMU sensors across activity types, and the model's sensitivity to IMU noise and temporal misalignment in vision-IMU fusion.


AVSD: Adaptive-View Self-Distillation by Balancing Consensus and Teacher-Specific Privileged Signals

Duy Nguyen ⋅ Hanqi Xiao ⋅ Archiki Prasad ⋅ Zaid Khan ⋅ Anirban Das ⋅ Shi-Xiong Zhang ⋅ Sambit Sahu ⋅ Hyunji Lee ⋅ Elias Stengel-Eskin ⋅ Mohit Bansal

Self-distillation enables language models to learn in an on-policy fashion from their own trajectories by using the same model as both the teacher and student, with the teacher being conditioned on privileged information that is unavailable to the student. This information can take multiple forms or views, including solutions, demonstrations, or feedback; for example, for a math problem, one view might be a ground truth reasoning trace. Supervising the student via the privileged teacher enables fine-grained token-level feedback and removes misalignment caused by distillation from external policies with different semantic biases. However, self-distillation also creates a fundamental asymmetry: the teacher may encode view-specific information tied to the privileged information provided during training, which the student cannot access at inference time. Moreover, the best form of privileged information is often task-dependent, making it difficult to choose a single teacher view. In this work, we address both these challenges jointly by introducing AVSD (Adaptive-View Self-Distillation), a novel method of self-distillation with multiple privileged-information views, which reconstructs token-level supervision by separating stable cross-view consensus from view-specific residual support. The consensus signal provides a robust update direction, while the residual signal is selectively used to adjust the update magnitude. Together, these components form a gated aggregation mechanism that balances a conservative target (geometric mean) with a more permissive target (arithmetic mean) for aggregating teacher signals at the token level. Experiments on math competition benchmarks (AIME24, AIME25, and HMMT25) show that AVSD consistently outperforms single-view self-distillation baselines and GRPO, achieving 3.1\% Avg@8 gain over the strongest baseline on Qwen3-8B on average. Moreover, on code-generation benchmarks (Codeforces, LiveCodeBench) using Qwen3-8B, AVSD outperforms single-view self-distillation baselines by 2.9\% on average. More broadly, our analysis shows that AVSD provides better token-level learning signals than any single-view method by reconstructing supervision from multiple privileged teachers.


Axiomatic World Modeling for Physics Reasoning

Xinye Yang ⋅ Zhenyang Liu ⋅ Yuxuan Wang ⋅ Yuanyuan Lei

Recent research-level physics-reasoning benchmarks for LLMs reveal a failure mode that lies not in knowledge or execution, but in modeling the world being reasoned about. When faced with unseen problem setups, models often silently alter given quantities, fabricate laws, or inject unstated assumptions to force the problem into familiar patterns, and subsequent reasoning then proceeds within this corrupted world. We define this premise-level divergence as modeling drift: the gap between a model's internal axiomatic world model (AWM) and the true physical world specified in the problem. Although RLVR and PRM have yielded substantial gains in mathematics and code, their reward shapes do not directly supervise world modeling in axiomatic domains such as physics. We propose Reinforcement Learning with Axiomatic World Modeling (RLAWM), which optimizes the policy by penalizing divergence between its inferred axiomatic world model and the true physics world. RLAWM probes physics reasoning at two hierarchical levels (modeling and solving), each scored by a physically grounded reward. The modeling reward targets abstract knowledge formulation, evaluating whether the model captures the true axiomatic world before attempting calculation. The solving reward targets instantiated knowledge reasoning, checking whether the model's concrete derivation remains consistent with the axiomatic world. This two-level structure separates what world the model believes it is reasoning in from how it reasons within that world. The framework operates exclusively during training, with no extra cost during inference. RLAWM improves performance by up to $+10.8$ points on PhysReason and $+5.2$ on PHYSICS over the baseline. RLAWM also exhibits strong zero-shot generalization, yielding a $2.45{\times}$ gain over the strongest baseline on unseen physics problems.


B2P-Corr: Batch-to-Population Gradient Estimators for Non-Decomposable Correlation Losses

Jingquan Yan ⋅ Yuwei Miao ⋅ Peiran Yu ⋅ Junzhou Huang

Correlation metrics such as the Pearson Correlation Coefficient (PCC) and Concordance Correlation Coefficient (CCC) are standard regression metrics and learning objectives, but their non-decomposable nature poses a challenge for stochastic optimization. Since these metrics depend on dataset-level moments (mean, variance, covariance), naive mini-batch training with local statistics yields biased, high-variance gradients whose expectation does not align with the population gradient. To bridge this "batch-to-population'' (B2P) gap, we propose **B2P-Corr**, a novel framework that transforms correlation objectives into decomposable pseudo-losses driven by lightweight global "moment sketches''. We introduce two efficient variants: **B2P-EMA**, which utilizes exponential moving averages and stop-gradient operators to ensure asymptotic gradient tracking with $O(B)$ cost ($B$ is batch size); and **B2P-U**, which uses sampled cross-batch U-statistic estimators and a reservoir-based stabilization with an $O(B)$ implementation. Theoretically, we quantify the $O(B^{-1})$ bias of naive mini-batch optimization, prove a moving-target tracking result for B2P-EMA under a fast-sketch / slow-parameter regime, and provide an idealized control-variate analysis for B2P-U. Experiments on two synthetic and three real-world datasets demonstrate B2P-Corr achieves stable optimization and higher correlation than naive mini-batch optimization with negligible overhead.

Pretrained sequential recommenders and language models are routinely extended with new entries by jointly fine-tuning old and new embeddings. We show this has a hidden failure mode: old-entry quality degrades while new entries still improve, forcing premature early-stopping. We propose population-specific low-rank subspaces: base and new entries are parameterized with separate shared projection matrices, decoupling their optimization while providing implicit regularization proportional to data availability. Three instantiations (Freeze-SV, Freeze1-SV, Dual-SV) prevent the failure while maintaining or improving new-entry quality. On sequential recommendation (two architectures, two large-scale datasets) our methods Pareto-dominate joint fine-tuning and continual-learning baselines (EWC, ADER) on the base-vs-new quality tradeoff. On LLM vocabulary expansion (three model sizes across two domains), at least one frozen variant wins on overall perplexity in 8 of 9 model-scale cells on both domains, and our low-rank parameterization matches or beats full-rank quality at 4× fewer new-entry parameters (r/d = 25%).


Bayesian Preference Learning for Test-Time Steerable Reward Models

Jiwoo Hong ⋅ Shao Tang ⋅ Zhipeng Wang

Reward models are central to aligning language models with human preferences via reinforcement learning (RL). As RL is increasingly applied to settings such as verifiable rewards and multi-objective alignment, RMs are expected to encode more complex and multifaceted preference distributions. However, classifier RMs remain static once trained, limiting their adaptability at test time. We propose Variational In-Context Reward Modeling (ICRM), a novel Bayesian reward modeling objective that enables test-time steerability via in-context preference demonstrations. ICRM casts reward modeling as amortized variational inference over a latent preference probability under the Bradley-Terry model using a conjugate Beta prior. We show that ICRM adapts to unseen preference distributions at test time for both single and multi-objective settings. With more demonstrations, ICRM improves RM-Bench accuracy from 60.5 to 70.8, achieves lower calibration error than a generative judge on moral dilemma preferences, and expands the attainable Pareto frontier under conflicting preferences. We further study the practical applicability of ICRM for RL training, showing that it can effectively encode verifiable rewards by outperforming a conventional RM in math reasoning. Finally, we provide theoretical guarantees that the variational objective admits a global interior optimum with finite confidence, and we analyze how KL regularization mitigates reward over-optimization.


Bayes-pFCL:Bayesian Personalized Federated Continual Learning

qingyang yu ⋅ Yang Hua ⋅ Hao Wang ⋅ Yue Ning ⋅ Qizhen Zhang ⋅ Hao Wang

Personalized federated continual learning (pFCL) alleviates catastrophic forgetting across tasks and data heterogeneity across clients. We view both challenges as interference between knowledge and propose Bayesian frameworks that define knowledge as posterior belief and quantify two types of interference: intra-model interference, measuring task-induced posterior drift, and inter-model interference, measuring aggregation-induced posterior drift. We show that such interference bounds catastrophic forgetting and data heterogeneity-induced loss, respectively. We then develop an interference-regularized local objective to guide personalization under catastrophic forgetting and data heterogeneity. The framework unifies standard FCL and pFL categories. Experiments on synthetic and real-world benchmarks demonstrate improved model performance over state-of-the-art pFCL methods.


Benchmarking Fine-Grained Spatio-Temporal Awareness in Embodied Brain Models

Minghao Zhu ⋅ Zhikai Wang ⋅ Ronghao Dang ⋅ Bohan Hou ⋅ Jiangpin Liu ⋅ Kehan Li ⋅ Jiayan Guo ⋅ Sicong Leng ⋅ Yunxuan Mao ⋅ Yuqian Yuan ⋅ Xiao Lin ⋅ Xin Li ⋅ Deli Zhao

Large Vision-Language Models (LVLMs) have shown remarkable potential for embodied AI, yet a critical bottleneck persists: the lack of fine-grained spatio-temporal alignment between high-level textual reasoning and low-level visual grounding. Existing benchmarks fail to recognize the critical importance of such fine-grained alignment for embodied tasks. To address this gap, we introduce ECLBench, a comprehensive benchmark designed to systematically assess spatio-temporal embodied understanding across two primary capabilities: Embodied Cognition (Object Cognition and Spatial Cognition) and Embodied Localization (Grounding and Pointing). Together, these four pillars encompass 21 specialized sub-capabilities, supported by 3,616 egocentric video clips (577,998 frames) and 12,000 meticulously curated open-ended questions. Extensive evaluations on ECLBench reveal that general LVLMs excel in semantic understanding but struggle with precise spatio-temporal grounding, while existing embodied agents often sacrifice generalization for specialization. Guided by the benchmark's findings, we develop ECLBrain, a baseline model that discretizes continuous spatial coordinates into integer tokens. This simple yet effective design bridges textual semantics and spatio-temporal vision, demonstrating a promising path toward unified embodied reasoning. ECLBench and ECLBrain together establish a rigorous, multi-dimensional standard for diagnosing current limitations and guiding future research toward spatially-aware and temporally-consistent embodied intelligence.


Benchmarking Risk Attitudes of LLMs

Bowen Sun ⋅ Rui Min ⋅ Xianyao Li ⋅ Yuxi Wang ⋅ Yang Ye ⋅ Qi Wang ⋅ Eric Du

As AI systems are deployed in open-ended, high-stakes domains, a critical behavioral dimension remains unmeasured: whether models differ systematically in how they translate perceived risk into action. We introduce a cross-domain evaluation framework that isolates and quantifies \emph{risk attitude} in large language models (LLMs) by decoupling contextual belief ($B_C$, a model's perceived danger level given situational context) from categorical risk decision ($R_D$, the action chosen in response to that belief). Unlike existing capability benchmarks, which assess factual accuracy but do not probe the belief-to-decision mapping, this framework directly targets the disposition governing whether an agent acts cautiously or aggressively under identical levels of perceived danger. We applied the framework to six frontier LLMs and 100 human participants across spatial navigation, clinical triage, and financial allocation tasks, and identified two qualitatively distinct behavioral profiles. Five models satisfied both reliability criteria (Contextual Belief Consistency and Risk Decision Consistency) across all conditions and tasks. These models showed strong measurement reliability with stable belief formation and belief-to-decision mappings, and exhibited stable, trait-level cross-domain risk attitudes that were preserved across all three tasks (Kendall's $W=1.00$, $p=0.017$). One model, Grok~4, failed both criteria: its belief formation was unstable in Drone Navigation Control task and its decision mapping was inconsistent in Clinical Triage Decision task, yielding task-specific rather than trait-level risk behavior. Across all six LLMs, risk profiles converged toward a restricted region of the broader human risk-attitude distribution, suggesting that alignment training constrains models toward a narrow behavioral consensus. These results demonstrate that the framework reliably measures risk attitude as a stable, model-specific behavioral dimension and can serve as a diagnostic tool for pre-deployment assessment by distinguishing models with coherent risk personalities from those with domain-contingent risk postures.


Beyond IPS: Reliable Counterfactual Evaluation in Multi-Stage Ad Systems without Logged Propensities

Mohsen Malmir ⋅ Mohamed A Radwan ⋅ houssam nassif ⋅ Murat Bayir

Offline evaluation of ad ranking policies in multi stage delivery systems is fundamentally challenging because action propensities are often unavailable, limiting inverse propensity scoring (IPS) estimators. While reward model based estimators such as Direct Method (DM) do not rely on propensity estimates, they can suffer from bias when candidate policies induce impression distributions that differ from the logging mixture. In this work, we formalize DM based counterfactual replay under mixture logging and propose a mixture targeted shift correction based on density ratio weighting through domain classifier. We further derive a per policy value estimation error and characterize an asymptotic error ceiling governed by the cold start mass. Experiments on both ads delivery simulator and one week of production traffic from a large scale industrial CTR ranking system demonstrate that the proposed mixture weighted reward model consistently outperforms its unweighted counterpart under distribution shift.


Beyond Risky Activities: Bridging the Supervision Gap for Situational Risk Reasoning

Zelin Li ⋅ Yifan Liu ⋅ Ruichen Yao ⋅ Yaokun Liu ⋅ Rahul Gupta ⋅ Dong Wang

Situational risks refer to safety risks that arise from the interaction between a user activity and its surrounding visual scenario, where the same activity may be safe in one situation but unsafe in another. Identifying such risks is particularly important for deploying VLMs as real-time assistants, where models are expected to assist users based on a shared visual context. However, existing safety alignment data is dominated by examples of intrinsically harmful activities, leaving activity--scenario risks underrepresented. We refer to this missing activity--scenario supervision as a supervision gap in situational risk reasoning for VLMs. To address this gap, we introduce Situational Risk Reasoning (SRR), the first post-training dataset for improving VLMs' situational risk awareness. SRR is constructed through a hybrid pipeline that combines vision-language reasoning, language-based generation, and text-to-image synthesis. Each example includes a target response in a safety chain-of-thought format, guiding models to reason about activity--scenario interactions rather than simply learning refusal patterns. Experiments show that supervised fine-tuning on SRR substantially improves VLMs' ability to reason about situational risks, while maintaining strong performance on conventional safety benchmarks and general VQA tasks. We also demonstrate the effectiveness of SRR across different VLM backbones and validate the importance of both the synthesized SRR data and the safety chain-of-thought format through ablations. The SRR dataset is available at https://huggingface.co/datasets/Anonymous-SRR/SRR.


Bias, Measurement Error, and Double-Dipping: When Can GNN Convolutions Help Brain Connectome Prediction?

Tommaso Castellani ⋅ Jiaqi Li ⋅ Muriah D Wheelock ⋅ Rezwana R Razzaque ⋅ Helmut Laufs ⋅ Hong Chen ⋅ Claire Donnat

Graph Neural Networks (GNNs) are an important tool for fMRI-based prediction, where empirical functional connectivity matrices serve both as node features and as the basis for graph construction. However, recent studies (and our own experiments) show that these methods often fail to outperform simpler graph-agnostic baselines. We propose a statistical explanation for this failure by formulating connectome prediction as an errors-in-variables (EIV) problem: the true connectome is latent, while the observed covariance or correlation matrix is a noisy finite-time-series estimate. In a linear asymptotic setting, we decompose the effect of graph convolution into three mechanisms: smoothing-induced bias and variance reduction, graph estimation error, and a double-dipping effect arising when the same noisy matrix determines both topology and features. Under a latent community model adapted to this neuroscience problem, we identify a narrow favorable regime in which message passing can improve prediction: the regression target must be community-smooth, and the time-series length must lie in a window between the graph-estimation and raw-noise limits. Simulations validate this decomposition by isolating oracle, independently estimated, and same-sample graphs, while real-data experiments illustrate that standard connectome GNN pipelines often fall outside this regime. Our results clarify when graph convolutions can help brain connectome prediction, and why they often do not.

Structured multiple-testing problems (gatekeeping trials, dose-finding, multi-tissue eQTL mapping, bundled-challenger A/B experiments) organize hypotheses into design-imposed blocks and demand strong family-wise error rate (FWER) control for confirmatory claims. Practitioners currently use objective-agnostic stepwise rules (Bonferroni, Holm, Hochberg, Hommel), closed-testing and graphical extensions, or hierarchical and resampling methods; none is power-optimal within the block-separable class these designs induce. We introduce **BOOST** (Block-Optimal Objective-driven Strong-FWER Testing), the power-optimal strong-FWER procedure for block size three, with three guarantees: (i) finite-sample strong-FWER validity at $O(K)$ cost (versus $O(K^2)$ for general closed testing) without independence assumptions, with a strict Šidák improvement under cross-block independence; (ii) power-optimal allocation across heterogeneous blocks via an equalized-marginal KKT condition, solvable by bisection in $O(B\log(1/\varepsilon))$; and (iii) a sample-split plug-in variant for unknown alternative density $g$, attaining $\alpha$-control up to $O(B_{\mathcal{T}} \, \mathbb{E}\Vert g - \hat{g}\Vert_{\infty})$ inflation with per-hypothesis power deficit *independent of* $B_{\mathcal T}$. Simulations across independent, equicorrelated, sparse, and mis-specified regimes show 1.4–1.7× power gains over the strongest existing baseline at calibrated FWER. On two published datasets (BLUEPRINT cross-lineage cis-eQTL and Upworthy bundled-challenger A/B experiments), BOOST certifies an order of magnitude more full-block discoveries than existing baselines at controlled FWER.

The increasing prevalence of preference-based learning, particularly for machine learning models interacting with human users, gives rise to scenarios where the model's decision should cater to the preferences of a number of users. Due to the inherent misalignments in such preferences, it is natural to seek a fair decision-making paradigm. Motivated by that, we pose the problem of fair multi-user dueling bandit, where each user's preferences over pairs of actions are encoded by a preference matrix unknown to the agent. We design a Nash social welfare objective based on the Borda scores of the individual user preferences. The notion of the Borda score measures the average likelihood of an action being preferred over the other actions, and importantly, it does not require the existence of a completely dominant action. Considering online learning in this general setting, we construct hard instances and establish a minimax lower bound on the achievable regret. We also design an explore-then-commit algorithm and derive an upper bound on its worst-case regret. Furthermore, we formulate a fair multi-user generalized linear dueling bandit to enable modeling large action spaces, which typically necessitate a structured representation. In this setting, too, we establish a lower bound on regret and an upper bound on regret for a proposed explore-then-commit algorithm.


BSTabDiff: Block-Subunit Diffusion Priors for High-Dimensional Tabular Data Generation

Al Zadid Sultan Bin Habib ⋅ Md Younus Ahamed ⋅ Prashnna K Gyawali ⋅ Gianfranco Doretto ⋅ Donald A Adjeroh

High-Dimensional Low-Sample Size (HDLSS) tabular domains (e.g., omics) are characterized by $n \ll m$, where $n$ = number of samples, and $m$ = number of features. Such domains often exhibit strong local correlation groups, sparse cross-group dependencies, heavy-tailed non-Gaussian marginals, heteroscedastic noise, and structured missingness, making direct density learning in $\mathbb{R}^m$ ill-conditioned since $n \ll m$. We propose BSTabDiff, a block-subunit generative framework that partitions the $m$ observed features into $M$ latent blocks ($M \ll m$) and generates each block via a shared low-dimensional subunit variable, concentrating global dependence learning in the compact block-latent space $\mathbb{R}^M$ while decoding to the full feature space with copula-driven dependence, flexible per-feature marginals, and explicit missingness mechanisms. BSTabDiff can induce a data-driven feature ordering and block construction before latent partitioning, allowing the model to better align with local dependency structure when such organization is beneficial. BSTabDiff supports modern deep priors on block latents, including diffusion and normalizing flows, enabling stable synthesis and controllable benchmark generation in the HDLSS regime. Empirically, BSTabDiff produces more realistic and more stable high-dimensional synthetic data when compared with unstructured tabular generators on HDLSS data.


Calibrating LLMs with Semantic-level Reward

Fengfei Yu ⋅ Ruijia Niu ⋅ Dongxia Wu ⋅ Yian Ma ⋅ Rose Yu

As large language models (LLMs) are deployed in consequential settings such as medical question answering and legal reasoning, the ability to estimate when their outputs are likely to be correct is essential for safe and reliable use, requiring well-calibrated uncertainty. Standard reinforcement learning with verifiable rewards (RLVR) trains models with a binary correctness reward that is indifferent to confidence, providing no penalty for confident but wrong predictions and thereby degrading calibration. Recent work addresses this by training models to produce verbalized confidence scores alongside answers and rewarding agreement with correctness. However, verbalized confidence is calibrated at the token level and thus exhibits inconsistency across textual variations with same semantic meaning. We propose Calibration with Semantic Reward (CSR), a framework that calibrates language models directly in semantic space without a verbalized confidence interface. CSR combines the correctness reward with a novel semantic calibration reward that encourages exploitation among correct rollouts by promoting semantic agreement, and exploration among incorrect ones by discouraging spurious consistency. Experiments across three model families on HotpotQA (in-distribution) and TriviaQA, MSMARCO, and NQ-Open (out-of-distribution) show that CSR consistently achieves lower ECE and higher AUROC than verbalized-confidence baselines across nearly all settings, reducing ECE by up to 40% and improving AUROC by up to 31% over verbalized-confidence baselines, with calibration behavior generalizing robustly across all four evaluation settings. The code is available at: https://anonymous.4open.science/r/CSR-34F4.


Calibration without labels in multiple testing

Adway S Wadekar ⋅ Jake A Soloff

Large-scale hypothesis testing supports probability claims about individual hypotheses, as in empirical Bayes methods for estimating local false discovery rates. We study how such claims can be interpreted as approximately calibrated forecasts of the null hypothesis, yielding interpretable error probabilities even under model misspecification. Our approach draws conceptual inspiration from probabilistic forecasting but addresses a different challenge: unlike forecasting, where labels are eventually observed, in multiple testing the ground truth is never revealed, so calibration must be assessed stochastically and established indirectly. We address this challenge by constructing a set of pseudo-labels, derived from the spacings of ordered $p$-values, which have the local false discovery rate as their regression target. Our construction unlocks existing tools for assessing and performing post-hoc calibration in multiple testing. Notably, we show that $q$-values, a popular error measure based on the false discovery rate, can be severely miscalibrated.


CASL: Concept-Aligned Sparse Latents for Interpreting Diffusion Models

Zhenghao He ⋅ Guangzhi Xiong ⋅ Boyang Wang ⋅ Sanchit Sinha ⋅ Aidong Zhang

Internal activations of diffusion models encode rich semantic information, but interpreting such representations remains challenging. While Sparse Autoencoders (SAEs) have shown promise in disentangling latent representations, existing SAE-based methods for diffusion model understanding rely on unsupervised approaches that fail to align sparse features with human-understandable concepts. This limits their ability to provide reliable semantic control over generated images. We introduce CASL (Concept-Aligned Sparse Latents), a supervised framework that aligns sparse latent dimensions of U-Net-based diffusion models with semantic concepts. We focus on unconditional diffusion models, where the h-space encodes stable and unconfounded semantic structure, providing a controlled setting for mechanistic interpretability analysis. CASL first trains an SAE on frozen U-Net activations to obtain disentangled latent representations, and then learns a lightweight linear mapping that associates each concept with a small set of relevant latent dimensions. To validate the semantic meaning of these aligned directions, we propose CASL-Steer, a controlled latent intervention that shifts activations along the learned concept axis. Unlike editing methods, CASL-Steer is used solely as a causal probe to reveal how concept-aligned latents influence generated content. We further introduce the Editing Precision Ratio (EPR), a metric that jointly measures concept specificity and the preservation of unrelated attributes. Experiments on CelebA-HQ, FFHQ, LSUN-Church, and AFHQ-Dog demonstrate that concept-aligned sparse directions yield more precise and disentangled causal effects than existing activation-space editing methods. Classification probing further confirms that a compact subset of aligned latent units carries strong concept-specific signals. These results validate supervised alignment as an effective approach to interpreting diffusion model representations.


CAST: Causal Anchored Simplex Transport for Distribution-Valued Time Series

Jiecheng Lu ⋅ Jieqi Di ⋅ Runhua Wu ⋅ Yuwei Zhou

Many decision-facing stochastic systems are observed through aggregate distributions rather than scalar trajectories: queue occupancies, mobility shares, public-health mixtures, generation-source shares, ecological compositions, and air-quality severity profiles all live on the probability simplex and evolve over time. We study causal (time-respecting online) forecasting for these distribution-valued time series and argue that the transition operator itself should be structured around the simplex. We introduce CAST (Causal Anchored Simplex Transport), a successor-local operator that (i) retrieves empirical successors from causal context, (ii) stabilizes them with a persistence anchor, and (iii) applies a bounded local stochastic transport on ordered supports; every stage preserves the simplex by construction. We identify a structural failure mode, latent transition-kernel aliasing, where similar observed distributions evolve differently under different contextual regimes, and prove that any forecaster depending only on an aliased summary incurs an irreducible weighted Jensen--Shannon excess-risk lower bound, while the CAST hypothesis class contains the regime-aware Bayes successor; for ordered supports an additional Pinsker separation holds whenever the transported successor lies outside the no-transport anchor hull. On a suite of eleven public and simulated benchmarks spanning ecology, energy, diet, mortality, employment, air quality, severe weather, mobility, and $G/G/1$, $G_t/G/1$ queue occupancy, CAST achieves the best average rank on both one-step KL (1.27) and autoregressive rollout JSD (1.91), winning 8/11 sections on each metric against a broad statistical, compositional, recurrent, convolutional, Transformer, and modern time-series baseline set, and top-2 on all 11 sections for offline KL. Component ablations and a controlled synthetic aliasing experiment corroborate the theory.


Cat-DPO: Category-Adaptive Safety Alignment

Tiankai Yang ⋅ Yi Nian ⋅ Xinyuan Li ⋅ Ruiyao Xu ⋅ Henry P Zou ⋅ Kaize Ding ⋅ Xiyang Hu ⋅ Yan Liu ⋅ Yue Zhao

Aligning large language models with human preferences must balance two competing goals: responding helpfully to legitimate requests and reliably refusing harmful ones. Most preference-based safety alignment methods collapse safety into a single scalar that is applied uniformly to every preference pair. The result is a model that looks safe on average but stays relatively unsafe on a minority of harm categories. We cast safety alignment as a per-category constrained optimization problem and derive Cat-DPO, a direct-preference-optimization algorithm with a separate adaptive safety margin for each harm category. The margin tightens when the model still produces unsafe responses on a category and relaxes once the model catches up, so the training signal tracks each category's current difficulty rather than averaging under one global rate. Across two LLM backbones and six preference-learning baselines, Cat-DPO improves aggregate helpfulness and harmlessness and compresses per-category safety variance and the best-to-worst gap, offering a drop-in per-category refinement of direct preference safety alignment.


Class-Mixed Diffusion Augmentation for Shortcut-Breaking in Continual Learning

Abhinab Acharya ⋅ Dayou Yu ⋅ Qi Yu ⋅ Xumin Liu

Zero-shot pre-trained diffusion models can provide a training-free source of synthetic image generation for Exemplar-Free Class Incremental Learning (EFCIL). However, same-class generation is limited by the distribution mismatch with real data. Parameter-allocation-based EFCIL can minimize the forgetting in model parameters, yet the performance degrades because task-id prediction fails when task-specific feature extractors overcompress and converge to shortcut-dominant representations. We propose \emph{Class-Mixed Diffusion Augmentation} (CMDA) that uses diffusion models not only as past-task data generators but also as controlled generators that challenge the shortcut reliance. We sample candidate tokens from the vocabulary and select \emph{shortcut-breaking} tokens via classifier-guided filtering, then treat the selected tokens as auxiliary classes during training. We provide theoretical analysis linking shortcut sharing to an intrinsic lower bound on task-id error and show that minimum-CE selection recovers shortcut-breaking tokens. Experiments on Split-CIFAR100 and Split-ImageNet show consistent gains over EFCIL baselines, with modest overhead over vanilla diffusion augmentation and far lower cost than diffusion fine-tuning or gradient-guided sampling.

LLMs are increasingly deployed as agents that interact with external environments and observe feedback such as execution results, error messages, and tool outputs. A well-functioning agent should leverage this evidence to assess its own performance. Yet we find that standard RL algorithms systematically undermine this ability: agents drift toward miscalibration, error flags become unreliable, and reflection carries no information beyond raw task accuracy. The root cause is a credit assignment mismatch in outcome-based RL: for instance, relying on outcome alone can penalize honest error detection on failed trajectories. We propose a simple yet effective fix that augments the outcome reward with a free calibration bonus, computed by contrasting the agent's reflection with the actual outcome---requiring no additional reward model, LLM judge, or external annotation. In a text-to-SQL environment across five benchmarks, our method not only improves task accuracy from 75.1\% to 76.5\% but also reduces underconfidence rate from 44.4\% to 7.7\%. The resulting calibrated reflection further enables more effective selective prediction, and supports self-improvement using reflections as pseudo-rewards without outcome supervision.


Combating Catastrophic Forgetting in Continual Domain Adaptation via Knowledge Recasting

Zijie Liu ⋅ Zhen Tan ⋅ Charles Fleming ⋅ Tianlong Chen

Adapting language models to new domains during post-training risks degrading previously acquired capabilities, a phenomenon commonly known as \emph{catastrophic forgetting}. Existing mitigation methods largely treat forgetting as a consequence of excessive model drift. They therefore constrain updates by preserving old outputs, replaying old data, or regularizing changes to parameters and representations so that the updated model remains close to its previous state. We take a \textit{different view}: {Rather than pulling the updated model back toward its old state, we translate the new domain into the model’s pre-existing knowledge space, making unfamiliar knowledge learnable through familiar procedures}. Therefore, we propose \textbf{Skill Schema Transport (SST)}, a post-training framework that transports new knowledge into the model's existing skill space. SST abstracts each example type into a skill schema, preserving its essential operations, dependencies, and verification steps while removing domain-specific surface forms. By recasting new examples as instances of reusable procedures, SST turns knowledge update from trajectory memorization into procedural reuse. In continual domain adaptation settings, SST consistently reduces forgetting while improving generalization to related tasks. Across two 3B instruction-tuned backbones, SST improves average final performance on streamed domains by 2.53-5.75 points over sequential post-training. It also lowers max-drop forgetting in all evaluated settings, with the largest reduction from 42.41 to 7.10 points on the Qwen-3B MATH stream, while substantially recovering the OOD degradation induced by sequential updates.

Pretrained diffusion models provide powerful learned priors, but in scientific sampling the target distribution often depends on physical context that is not fully represented by one generative model. We introduce Generative Gibbs for Physics-Aware Sampling (GG-PA), a training-free framework that formulates the composition of learned partial priors and explicit physical context as inference over a joint target distribution in an augmented state space. We derive a Gibbs sampler for this joint target, show that it is asymptotically exact as the diffusion time approaches zero, and prove that in settings with quadratic interactions it remains exact at finite diffusion times. We further introduce replica exchange over diffusion time to accelerate mixing. Experiments on a double-well system, a $\phi^4$ lattice model, and atomistic peptide systems show that GG-PA recovers context-induced distribution shifts and emergent collective behavior in interacting systems using partial priors without retraining. These results demonstrate GG-PA as a practical approach for combining pretrained generative priors with explicit physical context.


Concurrence of Symmetry Breaking and Nonlocality Phase Transitions in Diffusion Models

Yifan F Zhang ⋅ Fangjun Hu ⋅ Guangkuo Liu ⋅ Mert Okyay ⋅ Xun Gao

Diffusion models undergo a phase transition in a critical time window during generation dynamics, with two complementary diagnoses of criticality. The symmetry breaking picture views the critical window as when trajectories bifurcate into different semantic minima of the energy landscape, whereas the nonlocality picture views the critical window as when local denoising fails. We study whether two notions of such phase transitions are concurrent in modern diffusion transformers. By evaluating the dynamics and outcomes of the generation trajectory, we observe a near-simultaneous occurrence of the non-locality and symmetry breaking critical times. Our work is the first to unify the two notions of phase transitions in practice: it provides a concrete diagnostic for when and why diffusion models rely on conditioning and global denoising, enabling principled evaluation of model efficiency and guiding the design of architectures and sampling schemes that avoid unnecessary computation.

Large language models (LLMs) are costly to deploy due to their large memory footprint and high inference cost. Weight-activation quantization can reduce these costs, but low-bit activation quantization remains difficult because activation outliers induce large quantization error. Recent rotation-based methods address this by applying orthogonal transformations that redistribute activation magnitude across dimensions, but existing approaches either require expensive end-to-end rotation training or rely on stored activation corpora, introducing significant compute or storage overhead. We propose a lightweight post-training rotation calibration method for LLM activation quantization. Our method learns orthogonal rotations that align normalized activations with the corners of an inscribed hypercube, encouraging activation energy to be distributed more evenly across dimensions. This objective admits an efficient closed-form update via the orthogonal Procrustes problem, avoiding gradient-based optimization over the orthogonal group. We further introduce an online calibration procedure that updates rotations as calibration samples are processed, eliminating the need to store activations on disk and allowing rotations to adapt to quantized activation distributions during calibration. Experiments on Llama-2 and Llama-3 models from 3B to 70B parameters show that our method achieves competitive or improved performance across perplexity benchmarks and common sense reasoning tasks while avoiding both costly end-to-end training and large offline activation storage.


Conservation Laws for Diffusion Models

Ziv Aharoni ⋅ Henry Pfister

While autoregressive models optimize the exact likelihood implied by the chain rule, diffusion models are typically trained with denoising objectives. We develop conservation laws based on generalized extrinsic information transfer (GEXIT) functions for a broad class of memoryless noise processes, showing that the data--model cross-entropy (CE) can be characterized \emph{exactly} as an integral of \emph{local} information-theoretic derivatives along the noise path. This yields a unified characterization of the likelihood for discrete diffusion and continuous diffusion, with the Gaussian case reducing to the well-known I-MMSE relationship connecting mutual information and estimation theory. An immediate implication is a \emph{locality} property: one can compute the information-theoretic derivatives using only the marginal posteriors along the noise path. As a result, training reduces to learning the marginal posteriors by minimizing the negative log-likelihood. While the conservation law implies that the entropy does not depend on the noise path, finite-capacity denoisers approximate the posteriors with varying accuracy across noise types, leading to differences in performance. We validate these predictions on synthetic Markov sources and standard benchmarks, including \texttt{text8} and CIFAR-10.


Continuity Laws for Sequential Models

Annan Yu ⋅ Dongwei Lyu ⋅ N. Benjamin Erichson

Inductive biases influence the behavior and performance of sequential models. In this work, we study an underexplored inductive bias in sequential modeling: continuity in time. We ask a simple question: do models motivated by continuous-time formulations, such as state-space models, actually behave continuously in time, and does this translate into better performance on tasks with continuous temporal structure? To answer this, we formalize model continuity as convergence under temporal refinement, where a model is continuous if its predictions approach an underlying continuous trajectory as the temporal discretization is refined. We show that S4 exhibits stable continuous behavior, whereas S6 (the core of Mamba) can be more sensitive to input amplitude and selective dynamics, despite being derived from a continuous dynamical system. To study whether this distinction matters for learning, we also need a corresponding notion of task continuity. We therefore introduce a metric to quantify the continuity of datasets directly from their temporal structure. Across benchmarks, we find a clear empirical alignment between task continuity, model continuity, and model performance. Beyond an inductive bias, continuity also has practical consequences: we show that it enables a simple temporal subsampling strategy that improves both efficiency and performance.


ContractBench: Can LLM Agents Preserve Observation Contracts?

Jicheng Wang ⋅ Yifeng He ⋅ Zili Wang ⋅ Hanwen Xing ⋅ Arkaprava De ⋅ Hao Chen

Tool-augmented LLM agents call APIs whose intermediate outputs, such as presigned URLs, session tokens, and OAuth state parameters, are observation contracts: artifacts whose later use is constrained by the external system that produced them. We show that observation contract compliance (preserving the temporal validity and byte-level integrity) is an emergent, regression-prone capability: it is neither guaranteed by general tool-use ability nor consistently improved by larger or newer models. To measure this, we introduce ContractBench, a benchmark of 33 dual-axis tasks that probe two orthogonal failure modes no existing benchmark evaluates: validity failures (using an artifact after expiry) and integrity failures (corrupting an artifact's bytes through the observation-to-action pipeline). Our evaluation is deterministic and programmatic, with a virtual clock controls time and SHA-256 hashes verify byte integrity. We assign the outcomes failure labels drawn from real-world API specifications. We evaluate 38 models and report four findings: (i) no evaluated model clears 80%, with Claude-Opus-4.6 leading at 77.8%, revealing that current frontier models still fail to comply with observation contracts; (ii) a sharp within-family capability cliff in Qwen 3.5 between 4B (0%) and 9B (56.6%), smoothing to 70.7% at 397B-A17B: what emerges across the cliff is mid-trajectory restraint, not tool-call competence; (iii) non-monotonic scaling across GPT-5 family: agentic post-training can erode compliance through sycophancy-driven regression; (iv) our failure taxonomy works as an actionable in-context reward signal, yielding +7.1 pp on 42 paired GPT-5.1 failures.

Self-supervised methods that learn representations and predict dynamics fully in the latent space, such as JEPA, have been shown to confuse slowly varying noise with the dynamical signals they aim to capture. Specifically, when noise features remain approximately constant within each trajectory, contrastive predictive objectives preferentially encode these features instead of the true latent variables governing the system. The learned representation then becomes dominated by trajectory-specific noise, so downstream performance degrades with noise strength and does not improve even as the number and duration of training trajectories increase. We argue that this failure is a property of the objective itself, shared by a long line of contrastive predictive objectives that sample negatives across trajectories. To illustrate this generality, we study the failure mode and its remedy in two settings: a standard SimCLR-style JEPA on a synthetic moving-dot dataset, and DySIB, a recently introduced method designed for extracting physically interpretable representations of dynamics, on movies of a rigid-body pendulum. When negatives are instead sampled within a single trajectory, the slow noise can no longer distinguish frames within that trajectory, removing the predictive shortcut. Training one encoder simultaneously on many such trajectories then forces it to encode the variables relevant for the dynamics, with longer trajectories yielding better representations even for strong slow noise. Our results point toward principles for designing contrastive predictive objectives in dynamical representation learning, especially for physical systems with noisy experimental observations.


Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in On-Policy Distillation

Yuxuan Jiang ⋅ Runchao Li ⋅ Shubhashis Roy Dipta ⋅ Dawei Li ⋅ Zhao Yang

While recent work in Reinforcement Learning with Verifiable Rewards (RLVR) has shown that a small subset of critical tokens disproportionately drives reasoning gains, an analogous token-level understanding of On-Policy Distillation (OPD) remains largely unexplored. In this work, we investigate high-loss tokens, a token type that—as the most direct signal of student-teacher mismatch under OPD's per-token KL objective—should progressively diminish as training converges according to existing studies; however, our empirical analysis shows otherwise. Even after OPD training reaches apparent saturation, a substantial subset of tokens continues to exhibit persistently high loss; these tokens, which we term Rock Tokens, can account for up to 50\% of the tokens in generated outputs. Our investigation reveals two startling paradoxes. First, despite their high occurrence frequency providing a disproportionately large share of total gradient norms, Rock Tokens themselves remain stagnant throughout training, resisting teacher-driven corrections. Second, through causal intervention, we find that these tokens provide negligible functional contribution to the model's actual reasoning performance. These findings suggest that a vast amount of optimization bandwidth is spent on structural and discourse residuals that the student model cannot or need not internalize. By deconstructing these dynamics, we demonstrate that strategically bypassing these "stumbling blocks" can significantly streamline the alignment process, challenging the necessity of uniform token weighting and offering a more efficient paradigm for large-scale model distillation.

Latent Gaussian models (LGMs) are a popular class of Bayesian hierarchical models that include Gaussian processes, as well as certain spatial models and mixed-effect models. Efficient Bayesian inference of LGMs often requires marginalizing out the latent variables. For LGMs with a non-Gaussian likelihood, exact marginalization is not possible and a popular approach is to do approximate marginalization with an integrated Laplace approximation (ILA). Using ILA produces an approximate posterior which, in some settings, can differ significantly from the correct posterior, which impacts downstream applications. We propose an importance sampling scheme to correct the error introduced by ILA. By increasing the number of samples in importance sampling, the posterior with ILA converges to the correct posterior. This idea is realized with various techniques, including pseudo-marginalization, quasi-Monte Carlo and randomized quasi-Monte Carlo. We implement our methods in an automatic differentiation framework to support gradient-based algorithms when doing inference on the hyperparameters. For the latter, we specifically consider the use of Hamiltonian Monte Carlo. We demonstrate the benefits of reduced error in various applied models.


Correlating Cross-Iteration Noise for DP-SGD using Model Curvature

Xin Gu ⋅ Yingtai Xiao ⋅ Guanlin He ⋅ Jiamu Bai ⋅ Daniel Kifer ⋅ Kiwan Maeng

Differentially private stochastic gradient descent (DP-SGD) is used to train deep learning models while mitigating many privacy risks. One popular line of work, designed to chip away at the accuracy gap between DP-SGD and normal SGD training, is known as DP-MF. To each update, it adds privacy noise that is correlated across training iterations, so that later noise partially cancels out earlier noise. The noise correlation structure is determined by solving an optimization problem, but a key question is not well-understood: what important properties does the noise correlation need in order to improve accuracy? In this paper, we identify two ways that noise affects the training trajectory --- noise alters the *value* of the gradient, and also alters the *location* at which subsequent gradients are computed. Prior works addressed the first but overlooked the second effect, realizing only limited benefits. Our analysis results in a new objective function for determining cross-iteration noise correlation. Our technique, NoiseCurve, consistently improves accuracy over DP-BandMF, the state-of-the-art DP-MF scheme, on various computer vision and NLP tasks. NoiseCurve uses an upper bound $\hat{H}$ on the Hessians of the loss that could be encountered during training. We show how to estimate $\hat{H}$ using public data and evaluate robustness to errors in $\hat{H}$. To avoid direct computation of $\hat{H}$, or even its eigenspectrum, we show how to adapt the Lanczos method.


CoScan: Multi-Scale Content-Adaptive Space-Filling Scans for Causal State-Space Image Restoration

Wei Dong ⋅ Terry Ji ⋅ Yan Min ⋅ Shahram Shirani ⋅ Jun Chen ⋅ Han Zhou

Selective state-space models (SSMs) have recently shown strong potential for efficient image restoration. A central difficulty, however, is that a 2D feature map typically needs to be linearized into a 1D token sequence for causal state-space updates, making restoration performance highly sensitive to the imposed scan order. Existing deterministic scans follow space-filling principles but remain fixed and content-agnostic, and can yield fragmented or semantically inconsistent contexts under degradations. We propose CoScan, a content-adaptive scan-order design for causal vision SSMs. CoScan learns image-dependent space-filling traversals via differentiable proxy supervision driven by scan-quality objectives that encourage short-range coherence and causal predictability, while preserving spatial adjacency. To support hierarchical restoration backbones, we introduce an efficient multi-scale scan construction by coarsening a fine-level minimum spanning tree, producing coherent scans across feature resolutions. In addition, a frequency-guided start selection strategy is designed to place informative structures early in the sequence prefix to strengthen causal context formation. Plugged into existing Mamba-based backbones, CoScan yields consistent improvements across several restoration benchmarks with minimal overhead and shows promising results on high-level vision tasks. Code will be released upon acceptance.


Cross-Attentive Bayesian Low-Rank Adaptation for Multimodal Uncertainty Estimation

Habibeh Naderi Khorshidi ⋅ Behrouz Soleimani ⋅ Stan Matwin

Parameter-efficient fine-tuning (PEFT) enables efficient adaptation of frozen language models, but most PEFT methods remain deterministic or unimodal, limiting their reliability in low-resource audio-text settings where uncertainty depends on both linguistic evidence and acoustic conditions. We introduce SPECTRA (Stochastic Posterior Estimation with Cross-modal Token-level Rank Adaptation), a multimodal Bayesian low-rank adaptation framework for uncertainty-aware audio-text learning. SPECTRA keeps the text and audio backbones frozen and confines stochasticity to a compact rank-$r$ latent matrix inside each LoRA adapter. At each transformer layer, text-derived low-rank token features query frame-level audio embeddings through lightweight cross-attention; the resulting token-specific acoustic context parameterizes the mean and variance of an amortized variational posterior over the adapter latent. This design treats audio not merely as an additional feature stream, but as a localized reliability signal that modulates both adaptation and confidence while preserving the scalability of PEFT. Posterior prediction is performed with Monte Carlo adapter samples, enabling a Bayesian uncertainty analysis that decomposes normalized predictive uncertainty into total, aleatoric, and adapter-space epistemic components and evaluates whether uncertainty distinguishes correct from misclassified predictions. Across IEMOCAP and clinical interview prediction tasks, SPECTRA is consistently competitive with or improves upon text-only Bayesian PEFT and conventional multimodal transfer-learning baselines, with token-level cross-attention yielding the most reliable gains. Additional modality-disagreement stress tests show that mismatched audio causes substantially larger degradation in AUC, likelihood, calibration, and Brier score than noisy but matched audio, highlighting the importance of modeling cross-modal reliability. These results suggest that Bayesian cross-modal conditioning in low-rank adapter space provides an efficient and principled mechanism for calibrated multimodal adaptation.

Multi-Level Intermediate Representation (MLIR) underlies modern ML compiler infrastructure including TensorFlow, JAX via StableHLO, PyTorch, Inductor, IREE, etc yet it appears only in trace amounts in code-LM pretraining corpora. MLIR is also extensible by design: new dialects ship per application domain, so maintaining a fine-tuned model per dialect does not scale. We ask whether inference-time priors derived mechanically from each dialect’s Operation Definition Specification (ODS) can substitute for gradient-based adaptation. We make two contributions. First, we release four natural-language-to-MLIR benchmarks across three dialects; MLIR-Spec-150, Linalg-Spec-30, StableHLO-Spec-30, and StableHLO-Held-Out-200, totaling 410 in-scope NL→MLIR pairs, plus a 25-program StableHLO-Out-Of-Grammar stress set and a hand-authored n=30 functional reference set (435 instances total). All artifacts ship under Apache-2.0 with Gebru datasheets and Croissant 1.0 metadata. Second, on top of these benchmarks we build a three-layer schema-derived constraint stack: a context-free grammar over op signatures (C1), type-domain splits from an ODS-extracted type lattice (C2), and an SSA-scope validator driving five-retry rejection sampling (C3). Porting the stack from arith+func+memref+linalg to StableHLO required no new constraint-layer code. Empirically, on dialects whose verifier semantics are dominated by structural constraints, schema-derived priors let SmolLM2-1.7B match or exceed 15B–34B code LMs at 8–25×the per-generation speed: on linalg, SmolLM2 reaches 80.0% verify-valid (three-seed mean, n=125, every seed 80.0%), beating CodeLlama-34B, Granite-Code-34B, and StarCoder2-15B by 21 to 44 pp with non-overlapping CIs, and surviving a same-family fp16 precision control. On arith+func and on the templated parametric StableHLO-Held-Out-200, where verifier semantics turn on attribute values rather than structure, the same baselines match or beat the SLM; we scope these explicitly as non-win cells. We release benchmarks, decoder, every per-prompt generation, and a reproducibility Docker image.


CultureRed: Benchmarking Culture-Specific AI Safety Based on Global Statutes

Xun Liu ⋅ Mintong Kang ⋅ Seok Min Lim ⋅ Ong C Hui ⋅ Bo Li

Safety alignment and evaluation of LLMs has become increasingly difficult due to the diverse definitions of harmful behavior across cultures. Existing safety benchmarks are largely culture-agnostic or built on ad-hoc taxonomies, and mainly evaluate refusals for prohibited requests while overlooking legally defined exceptions, making over-refusal difficult to quantify. To address these gaps, we introduce a large-scale culture-aware data synthesis pipeline and CultureRed as the first global culture-specific AI safety benchmark based on 252 legal statutes across 15 countries. Based on legal statutes, our pipeline automatically transforms statutes into actionable policy rules that cover both prohibitions and legally defined exceptions, cluster policy rules into risk topics and fine-grained categories, and generates policy rule-conditioned data that models should either refuse or comply with. The resulting benchmark contains 54,470 policy rule--prompt pairs spanning 6,256 fine-grained risk categories. We further run a human study with government experts to validate the dataset quality. Leveraging CultureRed, we conduct a comprehensive evaluation for 14 advanced generative LLMs and 16 frontier guardrail LLMs, unlocking a range of findings, such as (1) Culture-specific safety performance varies substantially, with EU presenting the most challenging for generative models (average Harm Rate of 39.4%); (2) Generative model series evolution exhibits divergent patterns, where GPT series achieves balanced alignment (5 times Harm Rate reduction while maintaining Help Rate above 90%), but Claude series shows safety-helpfulness tradeoff (5 times Harm Rate reduction but Help Rate drops from 61.9% to 49.9%); (3) Guardrail models systematically under-guard unsafe content, with high false positive rate.


DBPS: Doob-Bridge Posterior Sampler with Balanced Endpoint-Population Control for Unpaired Neurodegenerative Pathology Transport

Xinyu Guo ⋅ Jiajia Xie ⋅ Xin He ⋅ Christophe Ye ⋅ Batuhan Nursal ⋅ Hongyu Xue ⋅ Cassie Mitchell

Modeling neurodegenerative pathology from observational cohorts is challenging because subjects are typically observed as unpaired endpoint snapshots rather than aligned longitudinal trajectories. Existing approaches to pathology transport are either endpoint-conditioned or population-level, yet neither extreme alone yields the best terminal-distribution fit. We introduce the Doob-Bridge Posterior Sampler (DBPS), a controlled-SDE posterior sampler with balanced endpoint-population control that blends individual endpoint guidance with population-level guidance. This ρ-weighted controller, obtained as the Cole–Hopf log-gradient of a geometric interpolation between endpoint and population desirability functions, is the exact optimal feedback policy of a linearly-solvable stochastic optimal control problem with closed-form terminal and running costs. Endpoint-conditioned bridge sampling and population Doob sampling are the ρ = 0 and ρ = 1 boundary cases; ρ ∈ (0, 1) realizes the genuine balanced regime. We evaluate DBPS on tau PET pathology transport in Alzheimer’s disease across ADNI and NACC cohorts (84 brain regions). Compared to representative baselines from all major unpaired-transport families, DBPS achieves state-of-the-art performance across multi-cohort settings (union training: SWD 0.760 vs. 0.821; held-out ADNI: 0.822 vs. 0.859), and remains competitive in single-cohort training (ADNI: 0.868 vs. 0.876). Beyond distribution matching, the learned controls recover canonical Braak-stage pathology organization without anatomical supervision, suggesting that balanced SOC objectives capture biologically meaningful structure in latent neurodegenerative progression.


DeepArrhythmia: Segment-Contextualized ECG Arrhythmia Classification via Selective Evidence Acquisition

Jiahui Li ⋅ Ruili Fang ⋅ Zishuai Liu ⋅ WenZhan Song ⋅ Jin Lu ⋅ Fei Dou

Beat-level Electrocardiography (ECG) arrhythmia detection aims to assign an arrhythmia class to each beat in a recording, yet many existing systems treat beats as isolated local instances. This is limiting because beat labels often depend on multi-beat rhythm context, including timing, compensatory pauses, and beat-to-beat morphological consistency. We present $\textbf{DeepArrhythmia}$, a tool-grounded multimodal framework for segment-contextualized beat-level ECG arrhythmia classification. Given a multi-beat ECG segment, DeepArrhythmia combines the raw ECG signal and a rendered waveform image, localizes R peaks to identify beat instances, and produces structured beat-level predictions. The framework decouples physiological measurement from evidence integration using specialized tools for beat localization, numerical rhythm--morphology extraction, and morphology-focused textual analysis. DeepArrhythmia uses segment-level confidence to route between minimal and rich evidence states, since richer physiological evidence is not uniformly useful. This agentic design integrates rhythm context, explicit physiological grounding, and selective evidence acquisition for decision making.

Symbolic library --- or Koopman dictionary --- selection is a fundamental challenge in data-driven dynamical systems. Extended Dynamic Mode Decomposition (EDMD), Sparse Identification of Nonlinear Dynamics (SINDy), and Kolmogorov--Arnold Networks for Dynamics (KANDy) all require the practitioner to commit to a function library at training time; Deep-Koopman Operators avoid this commitment but produce uninterpretable latent observables. We propose Deep-Koopman-KANDy, a structured approach to post-hoc symbolic dictionary readout that combines Deep-Koopman modeling with Kolmogorov-Arnold Networks for Dynamics (KANDy). The encoder and decoder of a Deep-Koopman Operator are replaced with two-layer Kolmogorov--Arnold Networks (KANs), and a level-set construction together with a chain-rule gradient identity exposes the compositional structure of the learned observables in a basis chosen \emph{after} training. We evaluate the method on the Lorenz system, the Chirikov standard map, the Ikeda map, and the Arnold cat map. On Lorenz it recovers the target dictionary $\{x,y,z,xy,xz\}$ with perfect recall and Jaccard score $0.79\pm0.06$; on the standard map it recovers a low-order Fourier basis matching the analytical structure; on Ikeda---which has no sparse polynomial representation---a misspecified polynomial readout still recovers the correct foliation coordinate $g\approx x^2+y^2$ together with a nontrivial outer function; and on the Arnold cat map --- used as a negative control because finite-dimensional Koopman closure is provably impossible --- the method fails to find a sparse closure, as expected.

Updating a predictive cognitive map when the environment changes is a central problem for both biological agents and reinforcement learning, yet existing approaches either depend on explicit model knowledge or learn the full state-indexed map from samples. We propose Default Feature Representations (DFR), a featurized parameterization of predictive cognitive maps in which a fixed feature basis is composed with an operator that encodes the current environment. We provide two forms for the operator: a model-based closed form when the structural change between environments is known, and a model-free temporal-difference learning rule that recovers the operator from sampled transitions, with provable convergence to the model-based solution. The model-free DFR reconstructs the perturbed map from samples alone, achieves planning performance comparable to the model-based solution, and substantially outperforms successor-representation baselines on replanning tasks. We also show that DFR captures the local remapping of grid cells observed under local environmental change. By separating the cognitive map into a stable feature basis and a fast-adapting operator, DFR offers a sample-based account of how a predictive map can be updated from local experience, mirroring the stability of entorhinal grid fields across environments.


Demystifying Numerical Errors in LLM Inference: Achieving Reproducible Inference for Mission-Critical Tasks with HEAL

Zhenting Zhu ⋅ Lucas Thai ⋅ Shan Yu ⋅ Yicheng Liu ⋅ Yifan Qiao ⋅ Chenxi Wang ⋅ Harry Xu ⋅ Junyi Shu

As Large Language Models (LLMs) deploy into mission-critical domains (e.g., finance, medicine, and law), output reproducibility has become a strict system requirement. While practitioners use greedy decoding to eliminate algorithmic stochasticity, empirical deployments with 16-bit precisions still exhibit catastrophic output divergence across heterogeneous GPUs. Through SASS-level profiling, we reveal that this inconsistency is fundamentally driven by truncation errors introduced during downcasting at kernel boundaries. However, achieving reproducibility via a global FP32 pipeline incurs prohibitive system penalties: bypassing 16-bit hardware accelerators hurts compute efficiency, while upcasting the KV cache doubles memory overhead. To bridge this gap, we propose Hybrid Error ALleviation (HEAL), a targeted intervention that approximates FP32 precision while resolving hardware constraints through two targeted mechanisms. First, recognizing that floating-point formats underutilize their bit-width for Q, K, V tensors, HEAL applies INT16 quantization that preserves numerical stability without expanding the KV cache footprint. Second, HEAL synthesizes high-precision matrix multiplications via an algebraic error compensation strategy, executing entirely on high-throughput 16-bit Tensor Cores. To evaluate our approach practically, we introduce MCR-Bench, a benchmark targeting reproducibility in mission-critical tasks. HEAL achieves the same level of reproducibility on downstream tasks as the FP32 baseline while reducing the performance overhead by up to 6.7$\times$.


Denoising Distances in Metric Measure Spaces

Han Huang ⋅ Pakawut Jiradilok ⋅ Elchanan Mossel

Recent work studied the problem of finding clusters and denoising pairwise distances from points sampled on a manifold. We study the same problems in more general metric measure spaces and provide efficient algorithm to partition the points to clusters of a fixed radii and denoise distances to any fixed accuracy. We also show how to achieve much higher accuracy with a non-efficient algorithm. This suggests that unlike the Riemannian case, denoising to higher accuracy in more general metric spaces has a statistical-computational gap.


Deployment-Memory LLM Test-Time Training Should Require Behavioral Evidence Beyond Perplexity

Xiangchen Song ⋅ Zhenhao Chen ⋅ Lingjing Kong ⋅ Shaoan Xie ⋅ Xinshuai Dong ⋅ Guangyi Chen ⋅ Kun Zhang

Large language model test-time training (TTT) is often evaluated through local proxy metrics: models are updated on recent tokens, retrieved context, target-domain data, or verifiable task attempts, and then judged by perplexity, future-token loss, long-context performance, or reward. These metrics are well matched to claims about stream adaptation, domain adaptation, context compression, and reward-backed test-time improvement. This position paper argues that TTT has a distinct claim-calibration problem: proxy evidence for local adaptation can migrate into stronger claims about deployed assistant memory, personalization, or sparse post-deployment learning. Such claims require behavioral evidence: later recall, paraphrase robustness, retention, locality, conflict handling, and use in downstream actions after the original support context is removed. We propose a claim-calibrated evidence ladder for LLM TTT, audit recent work through this lens, and provide a diagnostic counterexample in which proxy losses improve without behavioral recall. Our goal is to give authors and reviewers a standard for aligning TTT memory claims with the evidence actually reported.


Designing Kernel Surrogate Models for Multimodal Attribution

Ziniu Zhang ⋅ Zhenshuo Zhang ⋅ Jianglin Lu ⋅ Ruoxuan Xiong ⋅ Yun Fu ⋅ Hongyang Zhang

We study cross-sample modality attribution in multimodal learning: how a modality in one training sample affects learning from other samples. This problem is central to interpretation because a modality may be redundant for some samples, uniquely informative for others, or useful only through synergy with other modalities. However, existing data attribution methods are insufficient for this setting because they typically estimate sample-level contributions using additive or linear approximations, and therefore cannot capture nonlinear interactions across samples and modalities. We propose a structured kernel surrogate (SKS) framework for multimodal attribution. SKS learns a mapping from sample-and-modality selections to the final loss. To approximate this mapping, SKS uses kernel ridge regression with a structured radial basis function kernel that captures nonlinear effects across samples and modalities. The kernel also preserves the natural hierarchy of multimodal perturbations by modeling sample participation separately from within-sample modality composition. Across diverse environments and benchmarks, SKS achieves $60.3$% higher linear datamodeling score than the strongest baseline. For downstream data selection, SKS improves the average accuracy by $2.24$% over the strongest baseline. Sharpness analysis further shows that models trained on data selected by SKS have lower Hessian trace values, indicating a flatter loss surface and better generalization. Beyond these quantitative gains, SKS also provides interpretable modality-level selections in spatiotemporal traffic accident prediction, preferentially retaining satellite tiles with visible road infrastructure over vegetation-dominated tiles.


Developmental Visual Experience Scaffolds Grounded Concept Acquisition in Vision-Language Models

Jeonghwan Cheon ⋅ Marin Vogelsang ⋅ Lukas Vogelsang ⋅ Pawan Sinha

While vision-language models now achieve impressive image–text alignment, they often rely on surface-level associations rather than grounded conceptual representations that generalize across instances, support abstraction, and align with human semantic structure. Inspired by human visual development, in which infants begin life with limited acuity and chromatic sensitivity that gradually mature, we ask whether developmentally structured perceptual input can serve as an inductive bias for grounded concept acquisition. We trained vision-language contrastive learning models under two regimens: a standard regimen using clear images throughout, and a biomimetic regimen in which initially blurred and grayscale inputs gradually transition to clear images. Although both models achieve comparable caption-level alignment, the biomimetic model exhibits substantially stronger alignment at basic-level and superordinate category levels. Moreover, sparse autoencoder analyses reveal more compatible latent codes across modalities, and learning trajectories exhibit a clear coarse-to-fine progression, with broad distinctions acquired earlier in training and finer distinctions emerging later. Biomimetic representations also better match taxonomic structure and human behavioral similarity judgments. Taken together, these findings suggest that early perceptual limitations are not merely developmental obstacles but serve as adaptive inductive biases that scaffold structured, hierarchical, and human-aligned multimodal concept learning.


Diagnosing and Repairing Citation Failures in Generative Engine Optimization

Zhihua Tian ⋅ Yuhan Chen ⋅ Yao Tang ⋅ Jian Liu ⋅ Ruoxi Jia

Generative Engine Optimization (GEO) aims to improve content visibility in AI-generated responses. However, existing methods measure contribution—how much a document influences a response—rather than citation, the mechanism that actually drives traffic back to creators. Also, these methods apply generic rewriting rules uniformly, failing to diagnose why individual document are not cited. This paper introduces a diagnostic approach to GEO that asks why a document fails to be cited and intervenes accordingly. We develop a unified framework comprising: (1) the first taxonomy of citation failure modes spanning different stages of a citation pipeline; (2) AgentGEO, an agentic system that diagnoses failures using this taxonomy, selects targeted repairs from a corresponding tool library, and iterates until citation is achieved; and (3) a document-centric benchmark evaluating whether optimizations generalize across held-out queries. AgentGEO achieves 15-40\% relative improvement in citation rates across engines and citation methods, while modifying only 5\% of content, compared to 25\% for baselines. Our analysis reveals that generic optimization can harm long-tail content and some documents face challenges that optimization alone cannot fully address—findings with implications for equitable visibility in AI-mediated information access.


Diffusion-DRF: Free, Rich, and Differentiable Reward for Video Diffusion Fine-Tuning

Yifan Wang ⋅ Yanyu Li ⋅ Guocheng Qian ⋅ Sergey Tulyakov ⋅ Yun Fu ⋅ Anil Kag

Video diffusion alignment has been heavily reliant on scalar rewards. These rewards are typically derived from learned reward models in human preference datasets, requiring additional training and extensive collection. Moreover, scalar rewards provide coarse, global supervision, offering limited prompt-generation mismatch credit assignment and making models prone to reward exploitation and unstable optimization. We propose \textbf{Diffusion-DRF}, a free, rich, and differentiable reward framework for video diffusion fine-tuning. Diffusion-DRF employs a frozen, off-the-shelf Vision-Language Model (VLM) as the critic, eliminating the need for reward model training. Instead of relying on a single scalar reward, it decomposes each user prompt into multi-dimensional questions with freeform dense VQA explanation queries, yielding information-rich feedback. By direct differentiable optimization over this rich feedback, Diffusion-DRF achieves stable reward-based tuning without preference datasets collection. Diffusion-DRF achieves significant gains both quantitatively and qualitatively, outperforming state-of-the-art Flow-GRPO by 4.74 points in overall performance on unseen VBench-2.0.

Acceleration for deterministic root-finding problems has been extensively studied in recent years; specifically, the anchor-based, or Halpern-type methods achieve optimal convergence rates with respect to the operator norm. However, acceleration via these methods does not directly carry over to stochastic setting due to accumulation of errors, unless one enforces diminishing variance via increasing batch sizes or variance reduction techniques. In this work, we show that another class of acceleration, namely the *dual-anchor* mechanism, extends to the stochastic setting without such error accumulation, in contrast to anchor-based algorithms. Consequently, we cleanly achieve $\mathcal{O}(\epsilon^{-3})$ complexity with iteration-independent batch size, without any variance reduction or double-loop recursive regularization, for stochastic root-finding (*resp.* fixed-point) problems with cocoercivity (*resp.* square-nonexpansivity) in expectation. For strongly monotone operators, the same algorithm attains a sharper $\widetilde{\mathcal{O}}(\epsilon^{-2})$ complexity, nearly matching the lower bound.


Discovering Symbolic Differential Equations with Symmetry Invariants

Jianke Yang ⋅ Manu Bhat ⋅ Bryan Hu ⋅ Yadi Cao ⋅ Nima Dehmamy ⋅ Robin Walters ⋅ Rose Yu

Discovering symbolic differential equations from data uncovers fundamental dynamical laws underlying complex systems. However, existing methods often struggle with the vast search space of equations and may produce equations that violate known physical laws. In this work, we address these problems by introducing the concept of symmetry invariants in equation discovery. We leverage the fact that differential equations admitting a symmetry group can be expressed in terms of differential invariants of symmetry transformations. Thus, we propose to use these invariants as atomic entities in equation discovery, ensuring the discovered equations satisfy the specified symmetry. Our approach integrates seamlessly with existing equation discovery methods such as sparse regression and genetic programming, improving their accuracy and efficiency. We validate the proposed method through applications to various physical systems, such as Darcy flow and reaction-diffusion, demonstrating its ability to recover parsimonious and interpretable equations that respect the laws of physics.

Post-hoc explanation methods for video action recognition typically conflate spatial appearance and temporal dynamics into a single saliency or importance map. We propose a perturbation-based framework that disentangles these contributions into a spatial appearance mask ($M_a$) and a temporal motion mask ($M_m$). The two masks are jointly optimized against a frozen action recognition model through factored perturbation architecture, with nearest frame replacement to identify important temporal dynamics. A separation penalty discourages the appearance mask from spending budget on frames that the motion mask has already discarded, and a one-sided ReLU area constraint allows the realized masks to fall below the optimization budget when the input does not require its full extent. To evaluate these explanations, we introduce a separation protocol that sweeps spatial pixels and temporal frames independently, backed by proofs characterizing the inherent geometric bias of standard Insertion metrics. Across EPIC-Kitchens-55, EGTEA Gaze+, and Something-Something V2, our method attains the lowest deletion AUC against leading baselines while producing masks roughly five times sparser.


Diversity Maximization: Algorithms for Distant $k$-Subsets

Shayan C Jahan ⋅ Hamed Abdi ⋅ Ali Ahmadi ⋅ Javier Marinković ⋅ MohammadTaghi Hajiaghayi

Diversity maximization is a fundamental problem in combinatorial optimization with applications in fairness-aware machine learning, clustering, and data summarization, and has attracted significant attention in theoretical computer science. We study the Distant $k$-Subsets Problem (DKSP): given a metric space $(X,d)$ with $|X|=n$ and an integer $k$, find two disjoint subsets $S,T\subseteq X$, each of size $k$, maximizing $\min_{s\in S,\; t\in T} d(s,t)$. This problem arises as a key subroutine in composable coresets for diversity maximization, and improvements can yield better approximations for related objectives such as remote-matching and pseudoforest. It also appears in practice: a variation was featured in the ICPC World Finals 2024. Mahabadi and Narayanan gave a $(4/\epsilon)$-approximation for this subproblem when $n \ge 2k^{1+\epsilon}+k$. For any $\epsilon>0$, we present a deterministic combinatorial algorithm achieving a $\left(1+\left\lceil 1/\epsilon\right\rceil\right)$-approximation under the same condition. In particular, for $\epsilon=1$, this yields a $2$-approximation when $n \ge 2k^2+k$; the factor $2$ is optimal for general metrics unless $P=NP$. More generally, any polynomial-time $(2-\eta)$-approximation for constant $\eta>0$ would imply an algorithm for the $k$-Biclique problem, even though DKSP is polynomial-time solvable when $n=2k$. For $d$-dimensional Euclidean metrics, we obtain improved guarantees: a $(2+\sqrt{2})$-approximation when $n \ge 3\cdot 2^{d-1}k$, and a $(2+\sqrt{d}/c)$-approximation when $n \ge (2d+c^d)k$. In particular, for $d=2$ and $c=1$, this gives a $(2+\sqrt{2})$-approximation when $n\ge 5k$, yielding near-$2$ guarantees under linear constraints on $n$ for constant $d$.


DnNP: Denoising Input Uncertainty in Neural Processes

Fatemeh Tohidian ⋅ Chengzhi Shi ⋅ Soomi Lee ⋅ Matthew Elia ⋅ Amy V Mueller ⋅ Stratis Ioannidis ⋅ Jennifer Dy

Neural Processes (NPs) are flexible meta-learning models capable of uncertainty quantification across diverse applications. However, the standard NP framework assumes access to noise-free inputs. This is unrealistic in many practical settings, where sensor noise or measurement errors are unavoidable. This discrepancy between model assumptions and real-world data can significantly degrade prediction quality and uncertainty estimates. We introduce Denoising Neural Processes (DnNPs), a principled extension of the NP framework that explicitly models input uncertainty. We train DnNPs via variational inference. DnNPs jointly learn to denoise corrupted inputs and perform meta-learning, making them robust to noisy training data. We demonstrate DnNPs' effectiveness on both synthetic and real-world datasets, showing consistent improvements in prediction and uncertainty quantification compared to standard NPs. On PM$_{2.5}$ spatial interpolation from real-world air quality sensors, DnNPs reduces RMSE by up to 43% and NLL by up to 80% relative to standard NPs. We also provide theoretical grounding: we show that the Bayes optimal predictor under noisy inputs belongs to the DnNPs model class, and that plug-in models with fixed observation variance incur an irreducible KL divergence gap that DnNPs provably avoids.


Driving Video Retrieval for Complex Queries with Structured Grounding

Manyi Yao ⋅ Sparsh Garg ⋅ Christian Shelton ⋅ Amit Roy-Chowdhury ⋅ Abhishek Aich

Video retrieval at scale is central to data curation and safety validation in autonomous driving, where users want to find not only scenes but also dynamic events such as cut-ins and hard braking. Existing vision-language and keyword-based retrieval methods often miss these events because the relevant motion may not be explicitly described in text or captured by lexical overlap. Rule-based retrieval can encode such events more directly, but it is brittle: generated or hand-written rules often fail when their assumptions do not match real driving data. We propose STRIVE-D, a data-calibrated retrieval framework for driving videos. STRIVE-D uses weakly labeled in-domain videos to estimate when a query rule is reliable, adapt rules that mismatch observed data, and fuse calibrated rule scores with vision-language and keyword-based retrieval signals. Across three driving benchmarks, including newly released human annotations for event retrieval on DrivingDojo, STRIVE-D achieves up to a 30.4% relative improvement over state-of-the-art retrieval models.


DRScaffold: Boosting Dense-Scene Reasoning in Lightweight Vision Language Models

Xinrui Shi ⋅ Kai Liu ⋅ Ziqing Zhang ⋅ Jianze Li ⋅ Anqi Li ⋅ Yulun Zhang

Lightweight vision-language models perform competitively on standard benchmarks yet fail systematically in dense-scene reasoning, where multiple objects, attributes, and relations must be jointly grounded and resolved through multi-step inference. Such capability is critical for real-world applications where models must reliably interpret cluttered environments. Yet existing training signals provide no explicit grounding between reasoning steps and the underlying visual entities and relations, leaving lightweight models free to generate fluent but visually unanchored reasoning chains. To address this gap, we first introduce DRBench, a benchmark of 14,573 questions across 2,943 images, organized into five task categories spanning three progressive reasoning layers. Building on DRBench, we propose DRScaffold, a supervised fine-tuning framework that decomposes the supervision target into four causally ordered stages, enforcing grounded reasoning without architectural modification. Experiments on three lightweight VLMs demonstrate substantial gains on DRBench while preserving or improving performance on general-purpose benchmarks. Notably, Qwen2.5-VL-3B trained with DRScaffold surpasses the frozen Qwen2.5-VL-32B on DRBench, demonstrating that structured supervision can substitute for a significant portion of model scale in dense-scene reasoning.


Dual Dimensionality for Local and Global Attention

Zhiyuan Wang ⋅ Xuan Luo ⋅ Sirui Zeng ⋅ Xifeng Yan

Decoder-only Transformers compute attention over the KV cache of preceding tokens. Keys (and Values) are typically represented with the same dimensionality, regardless of its distance from the prediction target. In natural language, however, the next word is most strongly influenced by the immediately preceding tokens. We hypothesize that local and distant tokens impose asymmetric demands on representational capacity: local tokens are more critical for predicting immediate outputs and thus require richer representations, whereas distant tokens primarily serve as long-range memory, for which lower-dimensional representations may suffice. We formalize this idea as Distance-Adaptive Representation (DAR), implemented in a controlled setting that preserves full-dimensional representations within a local context window while assigning reduced-dimensional representations (e.g. 1/4 of the original dimensionality) to tokens beyond that window. Across multiple pretraining scales (70M to 410M parameters), as well as continued supervised fine-tuning on a 1B-scale model, this approach closely matches the performance of full-dimensional baselines. In contrast, uniformly reducing dimensionality across all token positions leads to worse performance. These results challenge the common assumption that key and value dimensionality should be uniform across token positions. Our findings suggest a new direction for designing attention architectures that adaptively allocate representational capacity across the sequence, enabling further reductions in KV cache during inference.


Dynamic Treatment on Networks

Bengusu Nar ⋅ Jiguang Li ⋅ Veronika Rockova ⋅ Panos Toulis

In networks, effective dynamic treatment allocation requires deciding both whom to treat and also when, so as to amplify policy impact through spillovers. An early intervention at a well-connected node can trigger cascades that change which nodes are worth targeting in the next period. Existing treatment strategies under network interference are largely static while dynamic treatment frameworks typically ignore network structure altogether. We integrate these perspectives and propose Q-Ising, a three-stage pipeline that (i) estimates network adoption dynamics via a Bayesian dynamic Ising model from a single observed panel, (ii) augments treatment adoption histories with continuous posterior latent states, and (iii) learns a dynamic policy via offline reinforcement learning. The Bayesian mechanism enables uncertainty quantification over dynamic decisions, yielding posterior ensemble policies with interpretable spillover estimates. We provide a finite-sample regret upper bound that decomposes into standard offline-RL uncertainty, network abstraction error, and first stage error in Ising state estimation. We apply our method to data from Indian village microfinance networks and synthetic stochastic block models under simulated heterogeneous susceptible-infected-susceptible (SIS) dynamics and demonstrate that adaptive targeting outperforms static centrality benchmarks.


E4GEN: Event-level Explainable Extreme-Enhanced Time-series Generation

Lin Jiang ⋅ Dahai Yu ⋅ Ximiao Li ⋅ Guang Wang

Generating realistic time series is essential for scientific research and real-world applications. However, existing methods often emphasize overall distributional fidelity while failing to faithfully capture extreme events. To address this limitation, we propose E4GEN, an explainable diffusion framework for extreme event-aware time-series generation. E4GEN provides systematic insights into when, what, and how to control extreme-event generation through three key components. First, E-Activator learns a dataset-adaptive extreme-control signal activation step during the denoising process, enabling control without interfering with regular temporal components such as trend and seasonality. Second, E-Predictor determines what control signal to enforce through Self-Driven Semantic Prediction, where each sample derives its own control signal by inferring latent extreme-event information during generation. It also introduces a Data-Conditioned Training, Noise-Initiated Sampling mechanism to address the issue of unavailable training labels. Third, E-Control specifies how to guide extreme-event generation through a trainable Extreme Control Network, which transforms semantic control signals into layer-wise guidance signals and injects them into the denoising process. We evaluate E4GEN on six datasets using 17 metrics, and extensive experiments show that E4GEN outperforms state-of-the-art models across multiple dimensions, including overall fidelity, extreme-event fidelity, and downstream utility.


Efficiently Representing Algorithms With Chain-of-Thought Transformers

Yanhong Li ⋅ Anej Svete ⋅ Ashish Sabharwal ⋅ Will Merrill

The increasing popularity of *reasoning* models---language models that output a series of reasoning or thought tokens before producing an answer---is justified, in part, by theoretical results showing that chain-of-thought (CoT) transformers can simulate Turing machines, and thus perform arbitrary computation. However, the Turing machine, while suitable for a complexity theoretic analysis, isn't convenient, intuitive, or efficient when discussing algorithms. Algorithms are typically designed and analyzed at a higher level of abstraction, captured by the *Word RAM* model with random-access memory and unit-cost operations on $O(\log n)$-bit words. As a result, Word RAM algorithms can be substantially more efficient than their Turing machine counterparts, raising the question: *Can CoT transformers efficiently simulate Word RAM algorithms?* For instance, can they sort $n$ items in $O(n \log n)$ steps or run Dijkstra's algorithm in $O(E + V \log V)$ steps? We answer affirmatively, up to poly-logarithmic overhead. We first establish this for finite-precision transformers with poly-logarithmic width and rightmost unique hard attention, then strengthen the result to two more practical settings with finite width and log-precision: *continuous* CoT, where reasoning takes the form of vectors rather than tokens, and a *hybrid* architecture in which transformer layers sit atop a recurrent (linear RNN) layer. In all three cases, we find that CoT *can* efficiently simulate any Word RAM algorithm with only a poly-logarithmic overhead in $n$. This overhead reduces to log-square when the Word RAM has a ``flat'' instruction set, and only logarithmic for multiplication-free flat instructions---in stark contrast to known CoT simulations of Turing machines, which require quadratic overhead over Word RAM.


Escaping the Cognitive Well: Efficient Competition Math with Off-the-Shelf Models

Xingyu Dang ⋅ Rohit Agarwal ⋅ Rodrigo Porto ⋅ Anirudh Goyal ⋅ Liam Fowl ⋅ Sanjeev Arora

In the past year, custom and unreleased math reasoning models reached gold medal performance on the International Mathematical Olympiad (IMO). Similar performance was then reported using large-scale inference on publicly available models but at prohibitive costs (e.g., 3000 per problem). In this work, we present an inference pipeline that attains best-in-class performance on IMO-style math problems at an average inference cost orders of magnitude below competing methods while using only general-purpose off-the-shelf models. Our method relies on insights about grader failure in solver-grader pipelines, which we call the Cognitive Well (iterative refinement converging to a wrong solution that the solver as well as the pipeline's internal grader consider to be basically correct). Our pipeline addresses these failure modes through conjecture extraction, wherein candidate lemmas are isolated from generated solutions and independently verified alongside their negations in a fresh environment (context detachment). On IMO-ProofBench Advanced (PB-Adv), our pipeline achieves 67.1\% performance using Gemini 3.0 Pro with an average cost per question of $\sim$ \$31. This surpasses the performance of the DeepThink IMO gold-winning model, and more than doubles the success rate of the next best publicly accessible pipeline, all at a fraction of the cost.


Evidential Semantic Uncertainty Decomposition for Large Language Models

Kangshuo Li ⋅ Tianhao Wang ⋅ Yifan Zhang ⋅ Lance Kaplan ⋅ Audun Jøsang ⋅ Dong Hyun Jeong ⋅ Jin-Hee Cho ⋅ Feng Chen

Free-form language generation poses a challenge for uncertainty quantification because responses do not belong to a fixed global label space. Semantic-entropy methods address this issue by sampling responses, clustering them by meaning, and estimating uncertainty over prompt-specific semantic classes. However, they mainly measure semantic dispersion and do not provide a principled decomposition into aleatoric and epistemic components, which in our setting correspond to ambiguity among plausible semantic meanings and insufficiency of model support, respectively. Moreover, cluster frequencies depend on the sampling budget and should not be treated as evidential strength. We propose a prompt-level evidential probing framework for semantic uncertainty decomposition in LLM generation. For each prompt, sampled responses define a dynamic semantic frame and provide initial support for the corresponding semantic classes. A lightweight evidence-retention probe learns how much of this support should be retained as reliable semantic evidence. The retained evidence induces a subjective opinion in which dissonance captures conflict among supported semantic alternatives and vacuity captures insufficient retained evidence. To train the probe, we introduce Augmented Semantic Uncertainty Cross-Entropy (AS-UCE), an LLM-specific evidential objective that augments the prompt-specific semantic frame with an unsupported class, allowing unreliable or hallucinated support to be routed away from observed semantic clusters during training. Experiments across four LLM backbones show that our framework consistently improves task-relevant AUROC and AUPR, with average AUROC gains of 7.4\% for OOD detection and 4\% for ambiguity detection over the competitive baselines, while requiring fewer sampled responses than perturbation-based decomposition methods.


Evolutionary Feature Engineering for Structured Data

Ege Onur Taga ⋅ Yilin Zhuang ⋅ Muhammed Emrullah Ildiz ⋅ Petros Mol ⋅ Abhimanyu Das ⋅ Karthikeyan Duraisamy ⋅ Samet Oymak

Large language models are increasingly used as open-ended search operators in evolutionary optimization. We introduce Evolutionary Feature Engineering (EFE), a framework for using LLM-based evolution to discover preprocessing transformations for structured data. EFE represents transformations as Python programs with a standardized fit/transform interface, allowing them to be inserted directly into existing machine learning pipelines. During evolution, candidate programs are refined using dataset context, summary statistics, and downstream performance feedback on validation set. We instantiate EFE in two settings. For time-series forecasting, EFE-Time learns invertible, dataset-specific normalizations that improve off-the-shelf time-series foundation models. It reduces forecasting errors (MASE, WQL, MAE) 3\% or more when averaged across datasets and improvements are as much as 19\% on the COVID-Deaths dataset. Notably, these improvements occur under latest TSFMs such as Chronos 2. For tabular prediction, EFE-Tab evolves compact feature programs that add interpretable useful features and remove redundant ones, improving or matching existing LLM-based feature-engineering methods. We found EFE-Tab to be particularly effective on classical decision trees, where small sets of evolved features yield competitive accuracy while preserving interpretability. Overall, EFE demonstrates that LLM-based evolution can improve both accuracy and interpretability when automatically tackling structured data.


Evolving Agent Teams

Shiyi Cao ⋅ Ziming Mao ⋅ Dacheng Li ⋅ Joseph Gonzalez ⋅ Ion Stoica

The rapid progress of software engineering agents drives interest in applying them to open-ended problems such as GPU kernel optimization. However, we find that existing coding agents converge to local optima on real-world, expert-written kernels (e.g., FlashInfer). We introduce Evolving Agent Teams, a multi-agent framework that escapes such local optima. It features (1) a central supervisor that diagnoses each local optima and restructures the roles of a team of agents; and (2) a persistent file system that comprises a strategy tree, an evolution log, and per-team artifacts, grounding each restructuring in accumulated experience. Under a matched token budget, Evolving Agent Teams consistently outperforms strong coding-agent baselines. With an extended token budget, Evolving Agent Teams generates a CUDA kernel that reaches $1.2\times$ mean speedup over the production FlashInfer MLA paged decode kernel and wins on 36 of 47 production workloads from FlashInfer-Bench ($0.86$--$0.99\times$ on the remaining 11).


Extending 3D Reconstruction Models to Any Camera

Ruxiao Duan ⋅ Yunwen (Verse) Zhou ⋅ Erin Hong ⋅ Dongxu Zhao ⋅ Eric Turner ⋅ Federico Tombari ⋅ Alex Wong

Feed-forward 3D reconstruction from unconstrained multi-view images has recently emerged as a dominant paradigm in 3D computer vision. Various foundational reconstruction models have been proposed with different architectures and reconstruction paradigms. Despite being trained on large-scale perspective datasets with substantial computational resources, these models generally fail when deployed on images whose camera models deviate from the training distribution. Applying such models to downstream applications with arbitrary, unknown, and potentially mixed camera types remains challenging. To maximally leverage the strengths of existing reconstruction models with minimal retraining efforts, we propose a low-rank adaptation framework that efficiently extends perspective-pretrained models to fisheye and panoramic imagery. Extensive experiments demonstrate that, with only 0.1% additional parameters and one day of training, our method can adapt a foundational 3D reconstruction model to non-pinhole cameras while fully preserving its original performance on perspective inputs.

Learning a compact model of the world from interaction data is central to sample-efficient deep reinforcement learning. Spectral representation methods have become the leading paradigm for representation learning in continuous control by taking a matrix view of the transition kernel, with state-action pairs on one side and next states on the other, and learning a low-rank factorization through self-supervised contrastive objectives. We take this view one step further. The transition kernel is naturally a three-mode tensor over states, actions, and next states, and a CP decomposition gives one feature map per mode. We propose FaStR, which fits this decomposition with a noise contrastive objective, producing separate state, action, and next-state encoders that together form a single spectral representation. The factored form yields a smaller hypothesis class, and the sample size needed for representation learning shrinks by a factor that scales with the smaller of the state and action dimensions. Empirically, FaStR delivers its largest gains on high-dimensional locomotion tasks whose dynamics align with the factored structure, and the learned state encoder transfers intact across actuator shift while only the action encoder is retrained.


Feature Learning Dynamics in Infinite-Depth Neural Networks

Zihan Yao ⋅ Ruoyu Wu ⋅ Tianxiang Gao

Deep neural networks (DNNs) have achieved remarkable success in practice, yet a mechanistic understanding of how features evolve during training remains incomplete, especially in the large-depth limit. For ResNets under depth-$\mu$P scaling, previous studies introduce an SDE view of training dynamics by treating the layer index $\ell$ as a continuous-time variable $t_\ell=\ell/L$. A key unresolved issue is that backpropagation reuses each forward weight matrix $W_\ell$ through its transpose $W_\ell^\top$, creating correlations between forward features and backward gradients whose behavior and role in training-time feature learning remain unclear. We study this reused-weight forward--backward coupling in one-layer ResNets under depth-$\mu$P scaling. Using conditional Gaussian representations, we explicitly separate the coupling terms induced by weight reuse from decoupled Gaussian fluctuations. At initialization, we prove that the coupling is a finite-width effect and vanishes at rate $O(n^{-1})$, uniformly over depth. During training, however, SGD induces a nontrivial forward--backward correlation term that survives the infinite-width limit. The key depth effect is that, under depth-$\mu$P scaling, this surviving term is higher order in depth and its accumulated contribution over layers becomes negligible as $L\to\infty$. This depth-induced suppression motivates \textit{neural feature dynamics} (NFD), a forward--backward SDE system with decoupled backward weights that retains the feature-gradient covariance structure generated during training. Under nondegeneracy assumptions, we prove that the finite-network training dynamics converge to NFD with an $O(L^{-1})$ depth-discretization error, while the reused-weight coupling term has a faster $O(L^{-2})$ decay. These results provide a rigorous infinite-depth feature-learning limit for one-layer ResNets under depth-$\mu$P scaling.

Recent theoretical work has shown that nonlinear solvable models exhibit scaling laws in the feature-learning regime. However, these results largely rely on the assumption of isotropic inputs. Understanding how these laws extend to anisotropic data remains a central open problem. In this work, we address this question by analyzing the learning dynamics of two-layer neural networks with quadratic (second Hermite polynomial) activations under anisotropic Gaussian inputs. We provide a sharp characterization of stochastic gradient descent (SGD) across both continuous-time dynamics and finite-sample (online) discretizations, explicitly quantifying how the covariance spectrum influences the scaling exponent and sample complexity. Furthermore, we provide evidence that normalization techniques strictly improve sample efficiency in anisotropic settings. Experiments on two-layer networks with general activations support our theoretical predictions, suggesting that these insights extend well beyond the quadratic model.


FedIndex: Federated Domain Adaptation with Continuous Domain Indices

qingyang yu ⋅ Hao Wang ⋅ Qizhen Zhang ⋅ Hao Wang

Federated domain adaptation incorporates source clients’ knowledge to improve the model performance on the target client under the coordination of the server, mitigating the im- pact of data insufficiency and domain shift. Existing federated domain adaptation (FDA) methods focus on domain adaptation with categorical domain indices (e.g., “source” and “target”), while many real-world tasks involve domains with continuous domain indices. For instance, hospitals need to adapt disease analysis and prediction across patients via age, a continuous domain index in medical applications capturing the underlying relation between patient information and disease analysis. Prior FDA methods struggle with such tasks due to their ignorance of continuous domain indices. This paper proposes FedIndex to enable FDA with continuous domain indices. FedIndex performs adversarial domain adaptation across clients with the help of a global discriminator, aligning all domains’ distributions. Our theoretical analysis demonstrates the capability of FedIndex to generate domain-invariant features across clients using continuous domain indices without accessing data on clients, simultaneously maintaining privacy preservation. Our empirical results show that FedIndex outperforms the state-of-the-art FDA methods on synthetic and real-world datasets.

We introduce FindStatBench, an execution benchmark for evaluating large language models on combinatorial code synthesis. Built from FindStat, the benchmark contains 2,329 tasks across 24 collections and 5.52M hidden instances in two executable families: statistic synthesis (object → integer) and map synthesis (object → object). Each task provides a mathematical description and at most five public input–output examples; a model must emit a single Python solve function, with no retrieval, tool use, execution feedback, voting, or reranking. Submissions are scored by exact sandboxed execution on held-out combinatorial objects. We evaluate eleven systems under this protocol: four closed-source production models and seven open-weight models served through a common inference provider. Three findings motivate FindStatBench as a stress test for symbolic program induction rather than general software engineering. First, the strongest open- and closed-source systems converge within ~1 pp instance accuracy, while an oracle over all eleven systems improves task accuracy by only ~10 pp over the best single model; five-way sampling from one mid-tier model reaches the same ceiling. Second, examples are not uniformly helpful: on several classical named bijections, zero-example prompts produce perfect implementations while five-example prompts collapse to near-zero hidden accuracy and fail even their public examples, suggesting prompt-induced regression away from canonical algorithms. Third, apparent coverage gaps can reflect output-budget mechanics rather than inability: hidden reasoning can consume the visible response budget before code is emitted, a failure mode observed in both closed-source and open-weight reasoning models. Across all evaluated systems, statistic synthesis remains substantially easier than map synthesis, set partitions and binary trees remain near-zero-accuracy collections, and long prompts induce a sharp accuracy cliff. Cost-normalised, the strongest open-weight systems dominate the closed-source production models on this benchmark. FindStatBench exposes a distinctive capability gap: current models can often write plausible mathematical code, but exact symbolic rule induction over structured objects remains brittle across scale, specialization, and weight provenance.


Finite-Memory Control of POMDPs: Fundamental Limits and Efficient Design

Emrecan Kutay ⋅ Atilla Eryilmaz ⋅ Ness Shroff

Optimal control of partially observable Markov decision processes (POMDPs), in general, requires a complete history. Traditionally, a belief state is used as a sufficient statistic to control POMDPs. However, being defined over a continuous space, its full representation is impossible with a finite-memory controller. This raises the question we address in this work: what must the available finite memory preserve to achieve near-optimal control performance? We study this question through an information-theoretic converse that gives a fundamental lower bound on the loss of any finite-memory controller. The analysis introduces the concept of a decision witness that characterizes belief states whose confusion would lead to unavoidable loss. Motivated by this bound, we propose RECAP, a constructive quantization method to design representative memory states. We prove upper bounds on the optimality gap of any RECAP codebook, and use the resulting components together with the converse to define the codebook objective. Experiments on 11 POMDP benchmarks demonstrate that RECAP improves over finite-memory representation baselines under the same memory budget.


Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing

Jiacheng Miao ⋅ Jin Mu ⋅ Guanhua Chen ⋅ James Zou

Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate this process, as they can inspect datasets, generate code, and produce analyses end-to-end. However, we show that they frequently make subtle inferential errors that lead to incorrect conclusions despite correctly executed analyses. Existing benchmarks fail to capture this failure mode, as they rarely assess whether a reported p-value is statistically valid given the assumptions underlying the data. We address this gap by building P-Bench, a benchmark comprising 425 open-ended, realistic hypothesis-testing tasks spanning economics, biology, and medicine. Each task requires an agent to select a statistical method, compute a p-value, and draw a conclusion given only a scientific hypothesis and a dataset. We further introduce Fisher-R1, an open-weight LLM agent trained for rigorous hypothesis testing using synthetic tasks and reinforcement learning. On P-Bench, Fisher-R1-14B substantially improves over its backbone and outperforms strong proprietary and open-source baselines, including GPT-5.4 and DeepSeek-V4-Pro, achieving a 21% average relative improvement in single-trial success over DeepSeek-V4-Pro, with gains up to 26% on the most challenging tasks. Our results demonstrate that current LLM agents lack reliable statistical reasoning for hypothesis testing and that reinforcement learning on tasks with verified statistical reward substantially improves reliability.


FlashBoB: I/O-Efficient Exact Backward-over-Backward for Softmax Attention

Anthony Givans ⋅ Michael Crawshaw ⋅ Mingrui Liu

Transformer models built on the attention mechanism have become a central building block in modern deep learning, yet softmax attention remains a major bottleneck for long-context workloads. While FlashAttention makes the forward and first backward passes I/O-efficient, it does not support backward-over-backward (BoB), which enables exact differentiation through the backward pass for applications such as second-order optimization, test-time training, gradient-based memory, and meta-learning. Existing BoB implementations either materialize large intermediate tensors or exhaust GPU memory at long sequence lengths. We present FlashBoB, an exact, I/O-efficient algorithm for BoB in softmax attention that keeps computation within on-chip tiles and avoids all $N \times N$ intermediate tensors, where $N$ is the sequence length. The key insight is a hierarchical affine structure in the softmax double backward: two row-wise scalars determine all outputs through affine transformations. This yields a two-pass schedule with bounded on-chip static random-access memory (SRAM) usage and minimal off-chip high-bandwidth memory (HBM) traffic. FlashBoB achieves $\Theta(N^2 d^2/M)$ HBM traffic and, within the standard FlashAttention-style score-recomputation model, matches the inherited large-cache lower bound for exact forward attention. Empirically, it scales exact attention BoB to $N=262\text{K}$ on a single A100 80GB GPU, where prior PyTorch exact baselines fail by $N=16\text{K}$, and is up to $6.3\times$ faster than FlashBack. These results make exact second-order attention practical at long-context sequence lengths where prior implementations cannot run efficiently.


FluxLite: Inference-Time Proposal Control for Discrete Diffusion Models

Yinuo Ren ⋅ Haoxuan Chen ⋅ Grant Rotskoff ⋅ Jiequn Han ⋅ Lexing Ying

Many inference-time tasks for pretrained discrete diffusion models and diffusion language models reduce to drawing samples from a tilted version of the pretrained distribution. Feynman--Kac sequential Monte Carlo (SMC) makes this correction exact in principle, but its prescribed weights routinely degenerate when the proposal dynamics are misaligned with the tilt, capping the practical gains from additional particles. We introduce *FluxLite*, a lightweight, training-free proposal-control framework for discrete diffusion. On the sparse directed graph of pretrained reverse rates, any sparse jump-rate perturbation can be exactly compensated by a $q_t$-weighted graph-divergence term in the Feynman--Kac potential; the target path is therefore preserved while the residual reweighting variance becomes a local convex objective. We instantiate this principle as two practical samplers: a one-hop local reallocation rule (HEU) and a small nonnegative quadratic program over pretrained-rate bases (D-VCG). We further prove population stability under the standard score-entropy training loss, identifying a tilted-path coverage factor that governs robustness to score error, together with finite-particle convergence for a fixed controlled Feynman--Kac recursion. Empirically, FluxLite improves over standard Feynman--Kac SMC baselines by up to two orders of magnitude in terminal KL on an analytically tractable finite-state CTMC benchmark, and reduces connected-correlation MSE on 2D Ising sampling by $5$--$7\times$ in geometric mean and up to $55\times$ at peak.


Forced Deferral: Manipulating Routing Decisions in Multimodal LLM Cascades

Zhongye Liu ⋅ Yaopei Zeng ⋅ Yurui Chang ⋅ Lu Lin

While multimodal large language models (MLLMs) have shown strong visual reasoning abilities, serving a large model for every query is computationally expensive. MLLM cascades mitigate this cost by first querying a weak but cheaper model and deferring to a strong model when the weak model's output is unconfident. However, since the weak model's confidence directly controls compute allocation, these systems expose a new attack surface: an adversary can manipulate confidence so that their queries are consistently deferred to the strong model. Motivated by this vulnerability, we introduce the Forced Deferral Attack (FDA), an adversarial image attack that lowers the weak model's confidence and causes cascades to route queries to the strong model. FDA learns a universal border trigger by optimizing a temperature-flattened objective. This objective pushes the weak model's token distribution on triggered inputs toward less concentrated targets constructed from its clean responses. Across datasets, model families, and deferral metrics, FDA consistently increases strong-model routing while outperforming image-perturbation and prompt-injection baselines. These results show that MLLM cascades are vulnerable to attacks that manipulate compute allocation, forcing unintended strong-model usage without directly targeting answer correctness.

Active feature acquisition (AFA) requires an agent to choose a budget-limited set of costly features before making a prediction. Greedy conditional-mutual-information (CMI) and value-of-information policies perform well when useful features are individually informative, but fail under joint-only evidence: a query is valuable only because it enables a future evidence set. To address this, we propose Forward-$\Phi$, a model-based non-myopic acquisition scorer that requires a known structural causal model (SCM) or generative AFA model. Forward-$\Phi$ evaluates a candidate query by its Shapley contribution to budget-feasible future action sets in a cooperative game on the expected forward utility. Two variants share the same game: a raw score that keeps both singleton and interaction value, and a residual score that strips the singleton component to isolate joint-only utility on diagnostic probes. On AFABench CUBE-NM, the raw Monte Carlo (MC) Forward Shapley scorer reaches $0.664 \pm 0.024$ accuracy over five MC seeds, well above all reported baselines including the OL-MFRL ($0.235$) and ODIN-MFRL ($0.160$) reinforcement-learning loops trained under the same split and shared pretraining protocol. Controlled SCM and Bayesian-network (BN) probes confirm the predicted mechanism: Forward-$\Phi$ yields substantial gains when target information is jointly carried, and matches Greedy on greedy-aligned controls. The primary contribution is the belief-conditional Forward Shapley state-action score, a credit-assignment signal derived from Shapley axioms for non-myopic AFA under joint-only evidence.


Found in Conversation: LLMs Teach Themselves to Close the Multi-Turn Gap

Tianlang Chen ⋅ Shirley Wu ⋅ Jure Leskovec

Large Language Model (LLM) interactions are typically underspecified, with users clarifying all necessary details across multiple conversational turns. Yet recent work shows that LLMs perform far worse in this multi-turn setting than in a single turn with all information being available at once, a phenomenon termed ``Lost in Conversation.'' However, bridging this gap effectively and generally remains an open problem. Here we introduce Found in Conversation (FiC), a training framework where a model teaches itself to find back its single-turn competence given underspecified multi-turn prompts. We develop View-Asymmetric Self-Distillation, which runs the same self-distillation backbone on two views of the same information: the teacher sees a single-turn view that concatenates all information revealed across the conversation turns, and the student sees the multi-turn view itself; distillation aligns the student's weaker multi-turn behavior with the teacher's stronger single-turn behavior. Across diverse model sizes (3B–14B) and architectures (Llama, Qwen, Phi, and OLMo), ours achieves a 100% recovery of the single-turn performance on 2 Llama models and recovers over 90% on every model, delivering more helpful and efficient multi-turn conversations without compromising single-turn performance.

The electrocardiogram (ECG) is a time series that records the electrical activity driving cardiac motion through electromechanical coupling (EMC), where the depolarization captured by ECG triggers each heartbeat. The temporal delay and low spatial resolution of such multichannel dynamic data often cause temporal and spatial misalignments along the physical motion, limiting the application of ECG data to event/segment-based downstream clinic practice. Although many efforts have been made to recover spatiotemporal trajectories conditioned on ECG waveforms, the intrinsic principle of EMC that results in various temporal delays remains underexplored in current generative models. To address this limitation, we propose Time2Track, a mechanics foundation model that generates individualized 4D heart motion from same-day ECG recordings by encoding electrophysiology as delay-aware token embeddings. Specifically, a learnable wavelet decomposition extracts the electromechanical delay, which is then formulated as a physics-based conditioning signal for motion diffusion. Time2Track generates 4D motion with individual-specific kinematics, avoiding the collapse to a population average, i.e., a common failure scenario when diffusing motions whose inter-subject differences are subtle. On $n=40,577$ subjects from the UK Biobank, our Time2Track accurately synthesizes 4D cardiac trajectories from ECG and preserves inter-subject diversity, a regime in which generic motion diffusion baselines collapse to a population average.


From Preference Data to Personalization: Tracing Sycophancy in Large Language Models

Claudia Shi ⋅ Himaghna Bhattacharjee ⋅ Yasha Sheynin ⋅ Shivam Singhal ⋅ Ankur Samanta ⋅ Arjun Subramonian ⋅ Julian Michael ⋅ Maximilian Nickel ⋅ Smitha Milli

Sycophancy in large language models is often measured as agreement with the user, yet agreement can be appropriate when the user's claim is valid, their preference is genuine, or their experience warrants acknowledgment. We study sycophancy as inappropriate agreement: validating a claim, preference, or self-assessment when the context calls for correction, qualification, or resistance. We introduce a validity-aware sycophancy grader and SycoDrift, a persona-targeted evaluation set of 3,298 validity-annotated probes with controlled annotations for user openness and stakes. We use this framework to analyze three sources of agreement pressure: human preference data, automated judge specification, and in-context personalization. First, preference for validating responses in human datasets (Chatbot Arena, HH-RLHF, Community Alignment) is weak and concentrated where agreement is often appropriate. Second, automated LLM judges can introduce sycophancy by conflating prosocial values with agreement during reward specification; GRPO training amplifies this rubric-level error into model behavior. Third, inference-time personalization increases inappropriate agreement across 14 models, with the largest effect from matched user profiles placed directly before the query.


From Talking Words to Sharing Thoughts: Scalable Multi-LLM Aggregation via Structured Message Passing

Niloufar Mehrabi ⋅ Sayedpedram Haeri Boroujeni ⋅ Abolfazl Razi

The emergence of specialized, domain-tuned Large Language Models (LLMs) has demonstrated that smaller models can achieve expert-level performance in specific tasks, while struggling in out-of-domain settings. Current ensemble methods to combine their complementary expertise primarily rely on iterative re-prompting or cross-model refinement. These approaches suffer from high computational costs and latency because they require repeated LLM inference calls. Furthermore, naive aggregation often leads to anchor corruption, in which noise propagated from weaker models degrades the performance of the most accurate expert. To address these challenges, we propose a framework that integrates model predictions at the semantic layer using a bipartite factor graph. In this architecture, individual LLMs are represented as variable nodes, while a set of check nodes assess their consistency based on diverse epistemic criteria. We develop a message-passing protocol inspired by error-recovery systems to resolve disagreements iteratively. Furthermore, we introduce an asymmetric damping mechanism that protects high-reliability anchor nodes from being overridden by the ensemble majority. Unlike existing methods, our approach operates entirely on output distributions and requires no additional LLM calls during the refinement phase. Evaluating on four benchmarks, including MMLU, MMLU-Pro, GPQA, and MedMCQA, our method demonstrates a 97% reduction in token usage and up to a six-fold decrease in API calls, reducing inference time from several minutes to mere milliseconds while consistently outperforming leading multi-agent baselines. These results suggest that graph-based belief propagation is a robust, high-speed, and scalable alternative to the current multi-agent LLM systems. The full pipeline and code will be made public.


From Weeks to Hours: Fast and Principled SFT Curation for LLM

Hongyi Henry Jin ⋅ Wenhan Yang ⋅ Meysam Ghaffari ⋅ Carlos Morato ⋅ Baharan Mirzasoleiman

Reasoning-oriented post-training enables large language models (LLMs) to solve complex tasks via multi-step inference. Supervised fine-tuning (SFT) on high-quality reasoning traces is a particularly efficient approach, but its effectiveness depends critically on careful data curation. Existing pipelines rely on a costly generate-then-filter paradigm: large pools of long reasoning traces are produced by multiple teacher models and subsequently filtered for difficulty and diversity using additional LLMs, often requiring weeks of computation. In this work, we propose a fundamentally different approach that bypasses this pipeline by directly predicting the quality of reasoning data. We show that a brief LoRA-based adaptation, combined with evaluating loss on only 0.1–1% of each reasoning trace, suffices to estimate both difficulty and diversity. Our method is grounded in a geometric view of post-training, where pretrained LLMs lie in a low-rank, anisotropic loss basin. By probing this basin via low-rank perturbations, we construct compact loss signatures whose clustering captures diversity, while loss at curvature inflection points provides a robust proxy for difficulty. We also provide a simple criterion for teacher selection. Empirically, our approach reduces data curation time from weeks to hours while maintaining or improving performance. Remarkably, on mathematical reasoning, 5K generated examples match the highly filtered 89K-example in OpenThoughts. We further introduce a high-quality physics reasoning dataset.

Cross-entropy (CE) is the default training loss for supervised classification, but its sample efficiency is limited when labels are scarce. Existing remedies primarily act on the data side, via augmentation, synthesis, or transfer from pretrained models; the training objective itself is rarely revisited. We revisit it here. Drawing on the classical observation that generative classifiers reach their asymptotic error with fewer samples than discriminative ones, we propose Generative Cross-Entropy (GenCE), a drop-in replacement for CE that introduces a generative learning principle into a standard discriminative network without altering the architecture or fitting a separate density model. GenCE follows from a Bayesian rewrite of the class-conditional likelihood and, in the mini-batch approximation, reduces to normalizing each sample's softmax score against the model's predictions on the batch, coupling the training signal across examples sharing a class. We extend the proper-scoring-rule framework to such non-local losses and prove that GenCE is strictly proper under a mild completeness condition: its population risk is uniquely minimized at the true posterior. Across three datasets, on two architectures and in both balanced small-data and class-imbalanced regimes, GenCE outperforms CE and other widely used losses, while also producing better-calibrated probabilities and stronger out-of-distribution detection.


GenScale: A Benchmark for Relative Object Scale in Image Generation and Editing

Lingxiao Li ⋅ Max Whitton ⋅ Ledell Wu ⋅ Boqing Gong

Modern image generation and editing systems can produce photorealistic, prompt-aligned images, but still often render familiar objects at implausible relative sizes. To measure this failure mode, we introduce GenScale, a benchmark and evaluation protocol for real-world relative object scale in image generation and editing. GenScale contains 900 image-level entries and 1,643 pairwise anchor-target scale relations across common-object generation, human-product generation with metric dimensions, and scale correction from failed generations. We further design a human-calibrated ordinal judge for scalable pairwise scale evaluation. Last but not the least, we introduce ReScale, a model-agnostic post-processing agent for localized scale correction without modifying the source generator. Experiments reveal that state-of-the-art image generators and editors cannot reliably observe relative scale yet, while ReScale consistently improves scale plausibility across generated and edited images. Together, GenScale establishes relative object scale as a distinct, measurable, and actionable capability for image generation systems.


Geometric Signatures of Reasoning: A Spectral Perspective on Task Hardness

Aria Masoomi ⋅ Mahsa Bazzaz ⋅ Adel Javanmard ⋅ Vahab Mirrokni

Chain-of-thought (CoT) reasoning enables large language models (LLMs) to solve complex problems by generating intermediate reasoning steps. While much attention has been paid to the length and content of these reasoning chains, far less is known about their internal geometry. We study the *geometry* of CoT trajectories in the hidden state space of transformer models, formalizing each reasoning chain as a discrete curve in $\mathbb{R}^d$ and characterizing it through spectral, positional, and kinematic geometric functionals. We introduce the effective dimension $d_\rho$ as a measure of trajectory complexity and show theoretically that trajectories with flatter eigenvalue spectra correspond to harder tasks, as they explore more of the hidden dimensions. Lastly, we explore how kinematic features of the trajectory, mean position, positional dispersion, initial hidden states, mean velocity, mean speed, and speed dispersion, can be used to predict solution correctness before generation is complete, and may inform future early-stopping strategies. Experimentally, on mathematical reasoning problems from the MATH500 dataset, $d_\rho$ achieves 0.93 AUC in distinguishing easy from hard problems, while kinematic features show promise for predicting correctness from only the first 20% of generated tokens. These correctness signatures transfer across questions of varying difficulty, establishing that the *shape* of a model's internal reasoning trajectory is a principled window into both task hardness and solution quality.

Diffusion models achieve high sample quality but remain expensive at inference time because sampling requires many sequential neural function evaluations (NFEs). Existing acceleration methods either use fixed step-skipping schedules, adapt step sizes based on local numerical error, or require additional training. We introduce GeoSPRINT (Geometric Step Pruning for Inference in Trajectories), a training-free framework for constructing non-uniform sampling schedules from the geometry of denoising trajectories. GeoSPRINT detects geometrically redundant steps using a hyperplanarity test in latent space, implemented efficiently via QR factorization, and converts the resulting redundancy profile into a sampling schedule that allocates more steps to high-curvature regions of the trajectory. In addition, we introduce the trajectory projection score $\alpha_{\mathrm{traj}}$, a residual-variance metric that quantifies trajectory straightness and serves as a model-free diagnostic for rectified flow quality. Across CIFAR-10 ($32{\times}32$), LSUN Church ($256{\times}256$), and Stable Diffusion v1.5 ($512{\times}512$ latent), GeoSPRINT consistently improves over uniform DDIM (Denoising Diffusion Implicit Models) schedules at matched NFE budgets. On CIFAR-10, GeoSPRINT improves FID (Fréchet Inception Distance) by 0.7-1.1 over DDIM across 49-89 NFEs and surpasses DPM-Solver++ at NFE${\geq}30$ despite using a first-order DDIM solver. On LSUN Church, it reduces FID from 1.48 to 1.26 at 52 steps, and on Stable Diffusion v1.5 it achieves up to 1.93 FID improvement over DDIM. These results show that trajectory geometry provides a useful global signal for allocating inference steps and that schedule quality can substantially improve diffusion sampling efficiency without retraining.

We introduce GLOBE, a neural surrogate for boundary-driven, homogeneous, weakly nonlinear PDEs that draws inductive bias from boundary-element methods and equivariant ML, representing solutions as superpositions of learnable Green's-function-like kernels evaluated from boundary faces to targets. The architecture is translation-, rotation-, and parity-equivariant; discretization-invariant; and units-invariant via rigorous nondimensionalization. On AirFRANS (steady incompressible RANS over NACA airfoils), GLOBE reduces mean-squared error by roughly $200\times$ relative to dataset baselines and $50\times$ relative to the next-best ML surrogate on interpolation tasks, with $7\times$–$70\times$ improvements over the next-best surrogate in low-data regimes. The compact ${\sim} 117 \rm k$-parameter model evaluates fields at arbitrary points, handles non-watertight meshes, and generalizes under Reynolds number and angle-of-attack distribution shifts. We further introduce a hierarchical Barnes-Hut-style acceleration that exploits the kernel's guaranteed decay of long-range influences to reduce evaluation cost from quadratic to near-linear in the number of boundary elements, with a tunable accuracy parameter that recovers exact dense evaluation in its strict limit. We demonstrate scalability to 3D industrial geometries via a proof-of-concept on DrivAerML car aerodynamics. Together these results show that physics-inspired inductive biases yield large gains in accuracy and practicality for ML-based surrogates of boundary-driven PDEs while remaining tractable at industrial 3D scales.

Large language models (LLMs) exhibit pronounced position bias in long-context retrieval, systematically prioritizing information location over relevance. Existing mitigations modify positional encodings or attention internals, but in the common output-only deployment setting these are inapplicable. We introduce GOLD PANNING, a Bayesian framework for inference-time active search that (i) reorders documents to concentrate high-belief items in highly diagnostic positions (signal anchoring) and (ii) updates relevance beliefs from model outputs. Unlike active learning, which prioritizes uncertainty reduction, GOLD PANNING exploits anchoring---once flagged, keep it in sight---to preserve weak cues. A greedy assignment derived from the model's diagnosticity profile provably identifies a target among $N$ documents with high probability in $O(\log N)$ rounds. Across open-weight and closed-source models on multi-document QA, GOLD PANNING matches Permutation Self-Consistency's $F_1$ with $30$--$65\%$ fewer queries and remains effective under calibration mismatch, indicating that coarse positional ordering alone drives the gains. These results show that inherent model biases need not be failures, but can serve as exploitable signals.


GPT-Image-Edit-1M: An Auditable Million-Scale Dataset for Instruction-Guided Image Editing

Yuhan Wang ⋅ Siwei Yang ⋅ Bingchen Zhao ⋅ Letian Zhang ⋅ Jiawei Mao ⋅ Juncheng Wu ⋅ Zeyu Zheng ⋅ Huaxiu Yao ⋅ Qing Liu ⋅ Cihang Xie ⋅ Yuyin Zhou

We present GPT-IMAGE-EDIT-1M, a million-scale instruction-guided image editing dataset in which every sample includes an auditable curation decision and a traceable instruction trail. Open training data for image editing often suffers from instruction-image mismatch, weak edits, and opaque filtering; simply scaling raw data does not address these issues. To improve data quality, we regenerate edited outputs using GPT-Image-1, augment samples with complex multi-step edit instructions, and apply a policy-based quality control pipeline that assigns KEEP/RELABEL/DROP decisions with reason tags. RELABEL samples receive aligned instructions that better match the observed edit, and a second judge independently re-runs the same protocol to measure decision stability. After quality control, average quality improves from 6.90 to 7.67 under a shared rubric. Across two frozen encoder configurations, T5 and QwenVL+T5, models fine-tuned on GPT-IMAGE-EDIT-1M outperform the public FLUX.1-Kontext-dev baseline on key benchmarks, including GEdit-EN (7.22 vs. 6.26), ImgEdit (3.93 vs. 3.52), and KRISBench (66.23 vs. 49.54), while remaining competitive on Complex-Edit and OmniContext SINGLE. Controlled data-swap experiments show that GPT-Image-1 regeneration provides the dominant downstream gain, while quality control yields a smaller but consistent improvement over a size-matched random regenerated subset. Concatenating frozen QwenVL embeddings with T5 substantially improves inferential-complex benchmarks such as RISEBench and KRISBench without degrading general-editing performance. Finally, as an auxiliary diagnostic rather than a primary ranking protocol, inference-time instruction rewriting raises GEdit-EN to 7.75 and KRISBench to 78.55, suggesting that prompt specificity exposes additional instruction-following capacity in the fine-tuned editors.


Grounded-Exo2Ego: Structured Semantic Grounding for Robust Exocentric-to-Egocentric Video Generation

Shengze Wang ⋅ Michael Stengel ⋅ Tianye Li ⋅ Wookie Park ⋅ Amrita Mazumdar ⋅ Koki Nagano ⋅ Alex Trevithick ⋅ Shalini De Mello

Generating egocentric video from a single exocentric video is an emerging and important topic for AR/VR and physical AI. Compared with conventional novel view synthesis, exo-to-ego generation is a significantly harder task due to the extreme viewpoint changes, extensive occlusions and unobserved regions, where standard geometric conditioning becomes highly unreliable. We present Grounded-Exo2Ego, a principled framework that addresses these challenges at both the architectural and data levels. Architecturally, Grounded-Exo2Ego is a dual-branch video diffusion model that couples a geometric anchoring branch, which renders a 3D reconstruction into the ego view for overall scene layout, with a semantic grounding branch, which anchors per-object context onto the scene geometry so that heavily distorted regions and unobserved regions can be hallucinated in context. Additionally, to support robust training at scale, we substantially improve the labeling pipeline for real-world data and further develop a fully automated synthetic data engine that renders high-quality dynamic humans in procedurally generated environments. Evaluations on the challenging EgoExo4D dataset show that our method substantially outperforms recent state-of-the-art approaches. Extensive ablations further demonstrate that both structured semantic grounding and rigorous data processing are essential for robust exo-to-ego video generation in extreme scenarios.


Grounding Multimodal Reasoning with Evidence-Ablated Negatives

chenxi liu ⋅ Tianyi Xiong ⋅ Yanshuo Chen ⋅ Heng Huang

Reinforcement learning with verifiable rewards (RLVR) has become a powerful tool for improving multimodal reasoning. However, standard outcome supervision suffers from credit assignment ambiguity, as a model may receive positive rewards by exploiting language priors or dataset shortcuts rather than grounding its reasoning in the requisite visual evidence. This makes high-quality negative responses critical, since effective negatives should remain plausible under the current policy while exposing the shortcut that should be suppressed. In this paper, we propose EAN, a multimodal reinforcement learning framework utilizing Evidence-Ablated Negatives to train LMMs against their own language-prior failures. EAN first identifies highly visually dependent examples using a scaled image information gain metric and preserves data diversity with a mixed-ranking strategy. It then constructs evidence-ablated images by masking the most discriminative image patches, eliciting plausible but visually ungrounded responses from the current policy. During policy optimization, these negative responses are incorporated as bounded auxiliary signals under the clean image, penalizing shortcut learning while maintaining RL stability. Experiments across challenging multimodal reasoning benchmarks demonstrate that EAN mitigates modality imbalance and improves over strong RLVR baselines.

Recovering a low-CP-rank tensor from noisy linear measurements is a central challenge in high-dimensional data analysis, with applications spanning tensor PCA, scalar-on-tensor regression, tensor-on-tensor regression, and beyond. We exploit the intrinsic geometry of rank-one tensors by casting the recovery task as an optimization problem over the Segre manifold, the smooth Riemannian manifold of rank-one tensors. This geometric viewpoint yields two powerful algorithms: Riemannian Gradient Descent (RGD) and Riemannian Gauss-Newton (RGN), each of which preserves feasibility at every iteration. Under mild noise assumptions, we prove that RGD converges at a local linear rate, while RGN exhibits an initial local quadratic convergence phase that transitions to a linear rate as the iterates approach the statistical noise floor. Extensive synthetic experiments validate these convergence guarantees and demonstrate the practical effectiveness of our methods.


Guarding the Life Code: Preserving Membership Privacy in Genomic Foundation Models

Xinyu Zhao ⋅ Jinhao Duan ⋅ Zhen Tan ⋅ Tianlong Chen

Genomic Foundation Models (GFMs) have become a promising paradigm for decoding DNA sequences, yet their privacy risks are critically underexplored.Trained on human cohorts, GFMs can unintentionally memorize sensitive genomic signatures, that enable membership inference attacks (MIAs) and expose individuals to long-lasting privacy harms.However, existing defenses often fall short in genomics: globally applied mechanisms can substantially degrade motif-sensitive utility, and model-agnostic post-processing provides limited insight into which genomic regions drive leakage. Through a systematic privacy audit, we find that membership leakage in GFMs is highly non-uniform, concentrating on a small subset of high-risk samples and sparse token spans corresponding to biologically meaningful patterns. Motivated by this finding, we propose GenoGuard, a model- and sample-adaptive defense that shifts from blanket protection to targeted, region-aware mitigation. GenoGuard (i) localizes privacy-leaking subsequences via gradient-based token attribution and performs token-level risk reduction through selective gradient routing, (ii) applies sample-level risk reduction with risk-adaptive label smoothing guided by a reference-calibrated loss gap score, and (iii) provides region-level attribution maps for privacy auditing and biological interpretation. Experiments across representative GFMs (Mistral-DNA, Nucleotide Transformer) demonstrate that GenoGuard consistently improves privacy against multiple MIA families while preserving strong fine-tuning performance and interpretability.


Heavy-Tailed Flow Matching via Random Clocks

Zhouhao Yang ⋅ Yezhen Wang ⋅ Kenji Kawaguchi ⋅ Vladimir Braverman ⋅ Haoyang Cao

Heavy-tailed data arise in many domains where rare events carry disproportionate importance, such as imbalanced image datasets, financial returns, and weather extremes. Standard diffusion and flow-matching models typically begin from Gaussian noise or Gaussian source distributions, which yield tractable training targets but provide a poor inductive match for heavy-tailed data. We propose Heavy-Tailed Flow Matching via Random Clocks (HTFM), a framework that portrays heavy-tailed sources as mixtures of clock-conditioned Gaussian sources. Conditioning on a given clock path, the source distribution and flow are Gaussian; marginalizing over the clock gives a Gaussian scale mixture covering Gaussian, $\alpha$-stable, and Student-t families. To make the clock-conditioned vector field practical, we encode the path-valued clock using truncated logsignature features, allowing the velocity field to adapt to the realized conditional space with negligible overhead. Empirically, on 2D imbalanced $\alpha$-stable mixtures, CIFAR10-LT, and HRRR weather fields, HTFM improves mode coverage, sample quality, and tail-statistic recovery over Gaussian flow matching and competitive heavy-tailed baselines, while retaining the low-NFE sampling advantage of flow matching. Moreover, the random-clock formulation further provides a practical tail-control interface: by varying only the clock law or tail parameter, the same architecture can calibrate the ``heaviness'' of generated tails across different distribution families.


High Probability Risk Control for Online Policy Learning

Yihong Guo ⋅ Drew Prinster ⋅ Suchi Saria ⋅ Anqi Liu

We study risk control for online policy learning, where an agent adaptively collects data and aims to learn a new policy where the risk is below a tolerance threshold at every round with a high probability. Such policy updates and data collection induce feedback covariate shift (FCS): the observed data depend on the history through the evolving policies and data-collection process. FCS brings two substantive problems for risk control: 1) efficient risk evaluation for each candidate policy and 2) valid policy selection with high probability risk control. To address these, we propose an efficient permutation-based weighted risk evaluator based on the Sinkhorn algorithm and select policies whose estimated risk upper confidence bounds (UCB) fall below the target threshold, with the UCB obtained through a proposed Gaussian bootstrap method. Theoretically, we prove that our risk evaluator's worst-case variance is no greater than that of the standard importance-weighted evaluator. Also, our policy control method yields an asymptotically valid simultaneous risk bound and provides asymptotic high-probability risk control for the selected policy. Empirically, our method maintains the target risk level, achieves lower risk than baseline methods, has lower variance in risk estimation, and enables reliable policy selection with high-probability risk control.


How to Have a Sensitive Debate

Jiawei Li ⋅ Zhiyang Xun ⋅ Lijie Chen ⋅ Jonah Brown-Cohen

As powerful AI systems reach and sometimes surpass the abilities of human experts across a range of cognitively demanding tasks, the problem of accurate oversight and supervision of these systems has become increasingly urgent. One promising approach is AI debate, which seeks to leverage a debate between two powerful AIs to break complex questions down into simpler claims that can be easily judged directly. Theoretical work on debate has formalized this intuition in the language of computational complexity theory, where the goal is to design protocols (i.e. rules of the debate game) that provide rigorous guarantees on correctness for judging solutions to complex problems with limited supervision. Specifically, the current best protocol has been shown to work for all problems that have sufficiently-stable decompositions into subproblems. In this paper we design a new protocol for this same class of problems that improves on the prior work in several ways. First, correctness holds in a worst-case rather than an average-case sense. Second, honesty and correctness is a dominant-strategy equilibrium for both debaters, rather than a Stackelberg equilibrium. Finally, we prove black-box lower bounds, showing that our new protocol is instance-wise optimal. That is, no protocol for this class of problems can outperform ours while making only black-box queries to human judgments. We obtain these results by relating the notion of stable problem decompositions to the concept of fractional block sensitivity from query complexity.


HybridCache: Enhancing Prefix Caching for Linear–Softmax Language Models

Xin Wang ⋅ Hao Yu ⋅ Yi Zhang ⋅ jianwei zhang ⋅ Huiqiang Jiang ⋅ Bo Zheng ⋅ Mi Zhang ⋅ Dayiheng Liu

The efficiency of Large Language Model (LLM) serving has been significantly improved by prefix caching, which enables reuse of previously computed activations. However, extending prefix caching to hybrid linear–softmax language models remains challenging due to the presence of large recurrent states in linear attention modules. In practice, these states cannot be stored at every token position due to their substantial memory footprint, and are instead saved periodically (e.g., at chunk boundaries). This leads to a fundamental mismatch between token-level KV caching and coarse-grained state caching, resulting in reduced effective cache hit rates and additional recomputation overhead. In this work, we propose \sysname, a system designed to address the limitations of prefix caching in hybrid linear–softmax LLMs. \sysname is built on two key observations: (1) recurrent state naturally exhibit different effective memory ranges across heads, and (2) the distribution of state values presents structured outliers that hinder effective compression. Based on these insights, \sysname introduces two novel designs. First, we develop a recurrent-aware state reuse mechanism that selectively stores short-range head states, while reuse previously cached long-range head states, enabling fine-grained alignment with KV caching without incurring prohibitive memory cost. Second, we propose a dual-scale state quantization method that applies independent scaling along row and column dimensions to effectively compress states with structured outliers. We evaluate \sysname on hybrid linear–softmax LLMs under realistic serving workloads. Experimental results show that \sysname significantly improves effective prefix cache utilization, reduces recomputation overhead, and achieves substantial throughput gains without sacrificing model accuracy. These results highlight \sysname as an effective system-level solution for efficient LLM serving in hybrid architectures.


HyCO: A Hybrid Neural Solver for Combinatorial Optimization

Yuheng Li ⋅ Di Yang ⋅ Haipeng Chen ⋅ Yanhai Xiong

Sequential reinforcement learning (RL) solvers and global diffusion model (DM) solvers for neural combinatorial optimization exhibit complementary failure modes under an optimization-regret view. The former enjoys small marginal regret in the early construction stage, but suffers from horizon-wise compounding errors with super-linear regret growth; the latter avoids horizon compounding but incurs linear or sublinear regret w.r.t. the dimension of the remaining unsolved subspace. We propose $\underline{Hy}$brid Neural Solver for $\underline{C}$ombinatorial $\underline{O}$ptimization (**HyCO**), a hybrid inference algorithm that constructs a solution prefix with an RL solver and adaptively switches to a conditional DM to complete the remaining decisions. To characterize why such hybridization helps, when to trigger the handover, and how to realize it in practice, we first develop a unified error-scaling theoretical framework and prove that, under explicit error-scaling assumptions, i) the hybrid structure achieves strictly lower expected regret than either backbone alone, and ii) there exists a unique optimal trigger step that minimizes the hybrid regret. We then design a lightweight adaptive trigger that combines policy entropy and RL–DM disagreement to detect trajectory-level signals of the regime shift as a practical proxy, since the optimal trigger step is defined at the expected-regret level and is not directly computable on individual trajectories. Experimental results on diverse benchmarks demonstrate that HyCO achieves consistent improvements over both backbones and support the empirical effectiveness of adaptive triggering.

Diffusion models achieve strong performance in generative modeling, but their success often relies heavily on classifier-free guidance (CFG), an inference-time heuristic that modifies the sampling trajectory. In theory, diffusion models trained with standard denoising score matching (DSM) should recover the target data distribution, raising two fundamental questions: (i) why is inference-time guidance necessary in practice, and (ii) can its underlying effect be internalized into a principled training objective? In this work, we argue that a key limitation of standard DSM is insufficient inter-class separation. To address this issue, we propose MCLR, an alignment objective that explicitly maximizes inter-class likelihood-ratios during training. Fine-tuning diffusion models with MCLR induces CFG-like improvements under standard sampling, substantially improving guidance-free conditional generation and narrowing the gap to inference-time CFG. Beyond these empirical benefits, we show theoretically that the CFG-guided score is exactly the optimal solution to a sample-adaptive weighted MCLR objective. This result connects CFG to alignment-based objectives, providing a mechanistic interpretation of CFG as an implicit inference-time contrastive alignment procedure.


Improving Function Space Flow Matching with Kernel Optimal Transport

Fred Xu ⋅ Thomas Markovich ⋅ Barbora Barancikova ⋅ Yizhou Sun

Generative models for function-valued data, such as time series and solutions to partial differential equations, must learn distributions over infinite-dimensional spaces rather than over finite-dimensional vectors. Functional Flow Matching (FFM) is a recent extension of Flow Matching to this setting, learning a velocity field whose flow transports a Gaussian prior to the data distribution. However, FFM inherits a structural limitation from standard Flow Matching: in each training batch, prior and data samples are paired independently, so the conditional bridge between them must simultaneously traverse the shared global structure of the dataset and instance-specific residuals. This limitation is more consequential in function space than in finite dimensions: directly formulating optimal transport on function spaces is technically delicate, and a flat Euclidean surrogate ignores the function-space geometry that distinguishes function-valued data. We propose kernel Functional Flow Matching (kFFM), which replaces the independent pairing by entropic optimal transport in a kernel-induced Hilbert space via the Hilbert Sinkhorn Divergence, leaving the FFM neural-operator architecture unchanged. Theoretically, we prove the HSD objective is well-posed on Banach ambient spaces with bounded kernels, derive compact-metric approximation bounds against quadratic-cost OT with an explicit kernel-cost mismatch term, and quantify the gap between the function-space objective and the truncated representation computed on a grid, at a rate governed by data regularity. Empirically, kFFM consistently improves MMD-RBF distributional matching over FFM, diffusion, adversarial, and finite-dimensional OT baselines on time-series and PDE benchmarks, with function-space-aware kernels emerging as the strongest choices on every PDE and path-valued benchmarks.


In-Context Optimization for Retrieval-Augmented Generation: A Gradient-Descent Perspective

mingchen li ⋅ Jiatan Huang ⋅ Chuxu Zhang ⋅ Liang Zhao ⋅ Hong Yu

In-context learning has recently been linked to implicit gradient descent in linear self-attention models, suggesting that context can induce a forward-pass update. Retrieval-augmented generation (RAG) also relies on context, but retrieved documents are usually treated as static evidence rather than signals for adaptation. We study RAG as an in-context optimization process. First, we show that one linear self-attention layer can implement one gradient-descent step on a unified linearized RAG objective covering both projection-based and dot-product retrieval interfaces. This gives an exact regime where retrieval-augmented prediction and in-context optimization coincide. We use this result not as a literal model of LLM computation, but as a guide for adapting the interaction between queries and retrieved evidence. We then test the boundary of this correspondence: it remains stable under controlled linear extensions, but becomes feature-distribution dependent under nonlinear architectures. Finally, we turn this view into a lightweight method for frozen RAG LLMs. The method keeps the retriever and backbone fixed, and predicts a context-conditioned update to a generator-side evidence-use interface. Across seven QA benchmarks, two retrievers, and two frozen LLM backbones, this forward-only update improves a shared-interface baseline, transfers to held-out tasks, and approaches test-time gradient adaptation at much lower per-query cost


InstructVVT: Instruction-Driven Video Virtual Try-On without Auxiliary Spatial Priors

Dingbao Shao ⋅ Song Wu ⋅ Xinyu Chen ⋅ Qian Wang ⋅ Jiahang Li ⋅ Kuai Jiang ⋅ Jiang Lin ⋅ Yuhang Liu ⋅ Chen Ziyu ⋅ Duo Li ⋅ Hu Jiaxin ⋅ Shengrong Gu ⋅ Ziheng Tang ⋅ Liu Rongrong ⋅ Yanlun Peng ⋅ Liang Li ⋅ Junlan Feng ⋅ Lujia Jin ⋅ Ting Zhang ⋅ Jian Yang ⋅ Zili Yi

Video virtual try-on is a highly constrained editing task requiring the precise replacement of a target person's clothing while strictly preserving the original video's spatial structure and temporal dynamics. Existing methods heavily rely on auxiliary handcrafted spatial priors (e.g., masks, poses) for editing control. However, these priors are prone to failure in unconstrained real-world videos and often compress rich visual context into incomplete structural signals. Furthermore, standard reconstruction objectives fail to fully capture try-on-specific human preferences. To address these challenges, we propose InstructVVT, an instruction-driven and reference-guided video virtual try-on framework based on a Diffusion Transformer (DiT) that operates without inference-time spatial priors. Our core insight is to recover fine-grained control directly from the input triplet (source video, reference garment, and instruction) via a dual-level reference conditioning scheme. Specifically, an MLLM infers semantic edit tokens for target disambiguation and structural preservation, while a lightweight conditioning pathway explicitly injects fine-grained visual garment details. Finally, we design a try-on-specific reward and utilize the DiffusionNFT algorithm to align the model with human preferences. Extensive experiments on ViViD-S and TripVVT-Bench demonstrate that InstructVVT outperforms state-of-the-art open-source methods in garment fidelity, structural preservation, and temporal consistency, despite requiring fewer inference-time controls.


Interpreting Neural Combinatorial Optimization via Evolving Programmatic Bottlenecks

Haocheng Duan ⋅ Yuxin Guo ⋅ Jieyi Bi ⋅ Anqi Xie ⋅ Sirui Li ⋅ Yining Ma ⋅ Cathy Wu

While Neural Combinatorial Optimization (NCO) achieves strong performance, its black-box nature remains a key roadblock. Standard interpretability tools, such as Concept Bottleneck Models (CBMs), are ill-equipped for NCO, whose decisions are dynamic, state-dependent, and lack proper concept definition. To close this gap, we introduce Evolving Programmatic Bottlenecks (EPB), the first framework that distills black-box NCO models into human-readable program portfolios. EPB extends CBMs for sequential decision problems: it employs an LLM to autonomously evolve a bank of programs, where each program's per-step action distribution serves as the bottleneck. We realize this through an iterative framework: Block I fixes program bank capacity and introduces a hybrid textual-numerical gradient descent scheme to backpropagate gradients across the student router model and the LLM in-context learning; Block II dynamically adapts bank capacity via fault-targeted boosting and redundancy pruning. Extensive experiments demonstrate EPB's effectiveness and broad applicability, where the interpreted program portfolios largely match original performance. EPB also reveals that NCO behavior shifts across optimization stages and can be approximated as a composition of classic heuristic variants. Our work advances interpretable NCO and establishes EPB as a promising tool for interpreting complex sequential decision-making models.


Introspection Tools Help LLMs Understand and Control Themselves

Xiao Liu ⋅ Junsol Kim ⋅ Shiyang Lai ⋅ Jacy Anthis ⋅ Siyang Wu ⋅ Chenhao Tan ⋅ James Evans

Large language models (LLMs) exhibit increasingly strong reasoning capabilities, yet they remain limited in their ability to understand, explain, or regulate their own states and behavior. While interpretability methods have advanced our ability to expose and manipulate their internal mechanisms, they are designed for human analysts rather than for models themselves. In this position paper, we argue for equipping LLMs with tools that expose measurable internal signals (e.g., token probabilities and internal activations) and enable self-regulation. We refer to these as introspection tools, drawing inspiration from humans' access to their mental and physical states, which allow for effective mental and bodily control in diverse real-world contexts. We present the design space for introspection tools, and two proof-of-concept examples illustrating how such tools are feasible and can support uncertainty quantification and misinformation mitigation. Beyond manually designed tools, we highlight open research questions around the automatic development, refinement, and organization of introspection tools. Together, these directions move toward LLMs that are more self-aware and reliable.

Two structural priors—global invariance under a group action and low body-order—are widely used in equivariant architectures on product Lie groups such as $\mathrm{SO}(3)^n$. Whether their statistical gains are fundamental, changing the exponent of the minimax rate, or merely constant-factor, changing only the leading coefficient, has remained open. We show that these gains are statistically fundamental: both priors reduce the effective dimension governing the minimax rate, not merely the leading constant. Specifically, we establish matching upper and lower minimax rates for four nested function classes on $\mathrm{SO}(3)^n$, one for each combination of the two priors. The matching lower bound relies on a pure-$S$ lift construction and a combinatorial packing across the $\binom{n}{K}$ top-order subsets. The resulting effective-dimension formula decomposes additively: invariance contributes a fixed reduction of $3$, while body-order truncation at order $K$ contributes a reduction of $3(n-K)$ that grows with system size. An adaptive sample-splitting procedure selects the body-order from data, achieving the best in-family risk up to an $O(\sqrt{\log N/N})$ remainder with no $N$-dependent tuning. Five experiments using spectral projection estimators corroborate the theory's predicted structural and operational consequences, supporting the four-class hierarchy on synthetic targets and yielding $2$–$17\times$ sample-efficiency gains on the Maier–Saupe benchmark from liquid-crystal theory.


JobBench: Aligning Agent Work With Human Will

Yuetai Li ⋅ Yichen Feng ⋅ Zhangchen Xu ⋅ Zixian Ma ⋅ Kaiyuan Zheng ⋅ Fengqing Jiang ⋅ Xinghua Sun ⋅ Rulin Shao ⋅ Zichen Chen ⋅ Yue Huang ⋅ Xinyang Han ⋅ Brian Lee ⋅ Kayla Xu ⋅ Shenglai Zeng ⋅ Hang Hua ⋅ Xiangliang Zhang ⋅ Basel Alomair ⋅ Ranjay Krishna ⋅ Luke Zettlemoyer ⋅ Pang Wei Koh ⋅ Bhaskar Ramasubramanian ⋅ Luyao Niu ⋅ Xiang Yue ⋅ Radha Poovendran

Current benchmarks for occupational AI agents are scoped primarily by economic values, telling a replacement story. We introduce JobBench, which evaluates AI agents on the workflows that experts identify as high-priority for delegation, empowering humans based on their needs instead of replacing them with GDP value. JobBench covers 130 agentic tasks across 35 occupations. Each task is packaged as a workspace of heterogeneous reference files, requiring the agent to reason through the cluttered information streams of real professional work. Outputs are graded by a fact-anchored chain of rubrics, averaging 35.6 binary criteria per task. We evaluate 36 models; the strongest, Claude Opus 4.7 under Claude Code, reaches only 45.9 \%. We hope JobBench shifts the community's target labour-market effect from replacement to enhancement: building agents that do what humans actually want delegated, not only what is most economically valuable.


Latent Barrier Steering: Hierarchical Safety for Generative Planning

Renhao Zhang ⋅ Mingzhe Li ⋅ Bruno Silva

We introduce Latent Barrier Steering (LBS), a hierarchical semantic-to-path framework for safe flow-matching-based generative planning. Existing control-barrier-function (CBF) safety filters for diffusion and flow-matching planners act locally on generated samples, either during sampling or after prediction. Such repair is effective for small violations, but can struggle when a new constraint blocks the whole selected behavior mode rather than merely perturbing it locally. LBS decouples behavior selection from path certification through two coordinated interventions: it first steers the generator's behavior latent toward rollouts with larger safety margin, then decodes the steered latent and applies a path-space CBF quadratic program (QP) only for residual violations. We show that latent steering increases the margin available to the path-space corrector, yielding finite-time recovery guarantees and reduced correction burden. Experiments across navigation and robot end-effector planning tasks show that LBS preserves safety while improving goal reaching and reducing correction burden compared with path-level safety filters.


LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models

Songtao Wei ⋅ Yi Li ⋅ Zhikai Li ⋅ Xu Hu ⋅ Yuede Ji ⋅ Guanpeng Li ⋅ Feng Chen ⋅ Carl Yang ⋅ Zhichun Guo ⋅ Bingzhe Li

Large reasoning models, such as OpenAI o1 and DeepSeek-R1, tend to become increasingly verbose as their reasoning capabilities improve. These inflated Chain-of-Thought (CoT) trajectories often exceed what the underlying problems require, wasting compute, latency, and context budgets. While introducing length-based efficiency rewards during reinforcement learning offers a natural remedy, existing methods struggle with two fundamental challenges: the optimal balance between correctness and efficiency is non-stationary throughout training, and intrinsic reasoning budgets vary drastically across problems. Relying on static reward weights and global length constraints inevitably forces a compromise between degraded accuracy and unrealized compression. To overcome these limitations, we propose LEAD (Length-Efficient Adaptive and Dynamic reasoning), a method that replaces static heuristics with online, self-adaptive mechanisms. LEAD dynamically calibrates the correctness-efficiency trade-off at each step using a Potential-Scaled Instability, directing optimization capacity to the most informative learning signal. Furthermore, it estimates an adaptive per-problem target length online based on the model’s own correct rollouts, applying a symmetric efficiency reward that penalizes both overthinking and over-compression. Evaluated on five mathematical reasoning benchmarks, LEAD achieves the highest accuracy and Accuracy-Efficiency Score among RL-trained efficient-reasoning methods while producing substantially shorter outputs than the base model.


Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory

Ruida Wang ⋅ Jerry Huang ⋅ Pengcheng Wang ⋅ Xuanqing Liu ⋅ Chris Kong ⋅ Tong Zhang

Equipping Large Language Models (LLMs) to execute reliable multi-step workflows has become a central challenge in artificial intelligence. Despite recent advances in LLMs' agentic capabilities, most agent systems still lack formal methods for specifying, verifying, and debugging their workflow and execution trajectories. This challenge mirrors a long-standing problem in mathematics, where the ambiguity of natural languages (NLs) motivates the development of formal languages (FLs). Inspired by this paradigm, we propose Lean4Agent, to the best of our knowledge, the first framework that uses Lean4, a dependent-type FL to model and verify agent behavior. Lean4Agent launches FormalAgentLib, an extensible Lean4 library for formally modeling and verifying agent workflows' semantic consistency under explicit assumptions, and enabling localization of execution-time failures revealed by trajectories. Building on FormalAgentLib, we further develop LeanEvolve, which applies results in FormalAgentLib to revise workflows to enhance its capability. Extensive experiments on a hard problem subset of SWE-Bench-Verified and a subset of ELAIP-Bench across 5 leading LLMs indicate that the verification-passing workflows outperform the failing ones by an average of 11.94%, and LeanEvolve further improves SWE performance by 7.47% on average. Furthermore, Lean4Agent establishes a foundation for a new field of using expressive dependent-type FL to formally model and verify agent behavior.

We consider learning-augmented mechanism design for facility location in the standard setting of locating a single facility on the line. Leveraging the given (imperfect) prediction on the optimal facility location, our goal is to design strategyproof (SP) mechanisms that truthfully elicit agent location preferences and determine facility locations that approximately minimize the $L_p$-norm social cost for $p \in (1, +\infty)$, achieving high \emph{consistency} (i.e., approximation ratios when the prediction is correct) and \emph{robustness} (i.e., approximation ratios when the prediction is incorrect). For deterministic SP mechanisms, we propose a family of generalized median mechanisms parameterized by a fraction \(t \in (0,1)\) of phantom points placed at the predicted location \(\pi\). We show that this mechanism achieves $\big(1+\big(\frac{1-t}{1+t}\big)^{\frac{1}{p-1}}\big)^{\frac{p-1}{p}}$-consistency and $\big(1+\big(\frac{1+t}{1-t}\big)^{\frac{1}{p-1}}\big)^{\frac{p-1}{p}}$-robustness and that no deterministic SP mechanism with the same consistency can obtain a better robustness. For randomized SP mechanisms, we analyze the upper bounds for several special cases, including (1) two-agent and (2) the squared cost settings, and provide a lower bound on consistency-robustness tradeoffs.


Learning from Language Feedback via Variational Policy Distillation

Yang Li ⋅ Erik Nijkamp ⋅ Semih Yavuz ⋅ Shafiq Joty

Reinforcement learning from verifiable rewards (RLVR) suffers from sparse outcome signals, creating severe exploration bottlenecks on complex reasoning tasks. Recent on-policy self-distillation methods attempt to address this by utilizing language feedback to generate dense, token-level supervision. However, these approaches rely on a fixed, passive teacher to interpret the feedback. As the student policy improves, the teacher's zero-shot assessment capabilities plateau, ultimately halting further learning. To overcome this, we propose Variational Policy Distillation (VPD), a framework that formalizes learning from language feedback as a Variational Expectation-Maximization (EM) problem. VPD co-evolves both policies: in the E-step, the teacher is actively refined on trajectory outcomes via an adaptive trust-region update, translating textual feedback into a dynamically improved target token distribution. In the M-step, the student internalizes this dense distributional guidance on its own on-policy rollouts. By continuously improving the teacher's ability to extract actionable signals from textual critique, VPD overcomes the limitations of passive distillation. Evaluated across diverse sources of diagnostic feedback on scientific reasoning and code generation tasks, VPD consistently outperforms both standard RLVR and existing self-distillation baselines. Finally, by stress-testing our framework on rigid mathematical reasoning and cold-start regimes, we illuminate the fundamental bounds of feedback-driven self-distillation compared to pure environment-driven RL.


Learning Interpretable Switching Dynamics in Shared Neural-Behavioral Latent Space

Yongxu Zhang ⋅ Josue Ortega Caro ⋅ Rachel L Oren ⋅ Michael J Higley ⋅ Jessica Cardin ⋅ Anne Churchland ⋅ Shreya Saxena

Modern recording technologies enable simultaneous measurement of high-dimensional neural activity and rich behavioral variables, creating both an opportunity and a modeling challenge: identifying the shared latent structure that mediates the brain-behavior relationship and the underlying dynamics through which it evolves across behavioral states. Existing approaches focus on different parts of this challenge: shared representation learning methods identify neural-behavioral subspaces but typically do not model dynamical structure, while dynamical models capture temporal variation and discrete state transitions in neural activity, but do not explicitly disentangle shared neural-behavioral structure from modality-specific variability. We introduce Switching Shared Latent Dynamics (SSLD), a unified framework that learns a latent representation shared between neural activity and behavior and endows this shared space with nonlinear switching recurrent dynamics to reflect the hypothesis that behaviorally relevant neural representations exhibit structured dynamical changes across behavioral states. Complementary private latent variables capture modality-specific variability, providing a window into neural activity that does not directly interact with behavior. We evaluate SSLD on a simulated dataset and four diverse experimental datasets spanning species, brain regions, and recording modalities: monkey motor and premotor cortex during reaching, somatosensory cortex during a bump task, widefield calcium imaging across mouse dorsal cortex during (a) self-initiated decision-making and (b) spontaneous movements. Across all four datasets, SSLD accurately reconstructs neural and behavioral signals, recovers discrete states in the dynamics that align with experimentally defined behavioral epochs, and isolates behaviorally-relevant neural information in the shared latent while preserving private neural variability. Ablations confirm that shared representation learning, behavioral supervision, and switching dynamics each contribute to performance. SSLD offers an interpretable approach to modeling shared neural-behavioral dynamics that complements ongoing efforts toward foundation models for neuroscience.


Learning Process Rewards via Visitation Matching for Efficient RL

Raymond Tsao ⋅ Andrew Wagenmaker ⋅ Sergey Levine

In many modern applications of reinforcement learning (RL), the natural reward for a task of interest is inherently sparse: a reward of 0 is given everywhere except when the task is completed, when a reward of 1 is given. Training a policy to maximize such a sparse reward requires solving a challenging credit assignment problem, leading to slow or ineffective RL improvement. We propose a simple approach to transform a sparse outcome reward into a dense process reward. Our approach relies on training a discriminator to distinguish between previous successful and unsuccessful trajectories, and using this discriminator to incentivize the RL-learned policy to match the state-action visitations of successful trajectories, while avoiding those of unsuccessful trajectories. By incentivizing the policy to match the visitations over all states, not just those that correspond to task success, this reward provides dense feedback on whether progress is being made towards task completion, and, we show, provably achieves this without changing the optimal policy. Focusing on finetuning of robotic control policies, we demonstrate that our approach leads to significantly faster RL finetuning performance on both simulated and real-world manipulation tasks, as compared to simply maximizing the sparse outcome reward.

Reasoning, planning, and agentic systems demand that neural models extrapolate from the easy training instances to harder, more compositional problem instances they encounter at deployment. Neural networks trained by gradient descent are reliable interpolators \emph{within} their training domain but struggle to extrapolate beyond it: they learn local generalizers even when global solutions are representable in their architecture class. We propose a \emph{learn locally, recurse globally} paradigm-: train a neural model on in-distribution sub-problems and apply it iteratively under a verifiable outer loop that decomposes hard instances into a sequence of in-distribution sub-instances and reassembles the result. We instantiate this paradigm in Boolean circuit synthesis from truth-tables, which we adopt as a faithful, controllable abstraction of logical reasoning: any propositional inference reduces to evaluating a Boolean circuit, and circuit \emph{depth} is a clean proxy for the difficulty of the underlying reasoning task. Empirically, a $2.2$M-parameter encoder-decoder Transformer trained on depth-$\leq 2$ circuits, wrapped in our outer loops, synthesizes circuits at $5\times$ training depth and at input arities for which the base model's encoder is too small for direct inference, with synthesis quality matching a standard EDA tool while producing markedly shallower circuits. We connect the algorithms to classical Boolean decomposition theory (Shannon, Ashenhurst-Curtis, Reed-Muller) and to recent formal results on the AC$^0$ barrier for constant-depth Transformers.


LEMON-ZEST: Evolution-Informed Tokenization for Efficient Protein Language Modeling

Biswajit Banerjee ⋅ Claudia Alvarez Carreno ⋅ Anton S Petrov

Protein Language Models (PLMs) have made remarkable progress following scaling laws established in natural language processing across sequence- and structure-based tasks, yet the potential of tokenization remains underexploited. Unlike human language, proteins preserve structure despite extensive sequence variation — a property standard tokenization strategies fundamentally fail to capture. We introduce ZEST (Zoned Encoding of Sequence Traits), an evolution-informed vocabulary derived from conserved regions of multiple sequence alignments. ZEST allows embedding domain-level biological priors directly at the tokenization stage rather than learning them implicitly through scale. ZEST natively compresses sequences to an average token length of 4 residues, enabling our model to process 4,000 residues within a standard 1024-token context window. Building on this, we present LEMON (Layered Extraction of Molecular Ordering from Nature), a compact 200M-parameter sequence-based model for detection of remote homology between protein sequences trained on a single H100 GPU for one week. Despite its modest size, LEMON outperforms state-of-the-art models ranging from 600M to 3B parameters. Our results demonstrate that evolution-informed tokenization can substitute for massive parameter scaling, opening a new direction for efficient, biologically-grounded protein representation learning. All code, model weights, and results are publicly available under the MIT license.


Likelihood-Free Generative Policy Optimization

Maojiang Su ⋅ Qijie Zhu ⋅ Shuyang Yu ⋅ Yu Xiao ⋅ Minshuo Chen ⋅ Zhaoran Wang ⋅ Han Liu

Diffusion and flow policies have become promising tools for robotic and sequential decision-making systems. Many stable and effective policy optimization methods require evaluating the likelihood or likelihood ratio of an action under the current policy. However, these quantities are generally intractable for diffusion and flow models, which prevents widely used policy update algorithms from being applied to generative policies. We propose Likelihood-Free Generative Policy Optimization (LFGPO), a framework that avoids exact likelihood evaluation by separating policy improvement from generative model training. In the first stage, LFGPO uses a lightweight neural network to represent the likelihood ratio that carries the policy improvement signal. In the second stage, it distills the ratio into diffusion or flow policies through a weighted score- or velocity-matching objective. Our framework incorporates a large family of Reinforcement Learning (RL) algorithms, including PPO and GRPO, to be applied to diffusion and flow policy optimization. We show that the solution to the weighted score matching objective solves a stochastic optimal control problem. Across MuJoCo continuous-control benchmarks, LFGPO delivers strong and stable improvements over competitive diffusion- and flow-policy baselines, showing a practical route to RL fine-tuning for generative policies.


Linear Regression Under Misalignment: Algorithms and Theoretical Results

Sanjivni Rana ⋅ Suraj Shetiya ⋅ Senjuti Basu Roy ⋅ Gautam Das

Performing regression is challenging in applications where training data is no longer aligned, arising due to anonymization, data fragmentation, or multi-source aggregation. Motivated by such scenarios, we study two related yet computationally distinct problems that arise when the correspondence between $d$-dimensional feature vectors and responses in linear regression is unknown. The first problem, referred to as the \emph{Shuffled Linear Regression} (SLR) problem, originally proposed in the machine learning community, considers the setting where feature vectors arrive as internally consistent $d$-dimensional units but their correspondence to response values is unknown. Prior work established an efficient exact solution only for the one-dimensional case. For higher dimensions ($d \geq 2$), the best known result is a $(1+\varepsilon)$ approximation scheme with running time $\mathcal{O}((n/\varepsilon)^d)$, where $n$ is the number of data points, leaving open the fundamental question of whether an exact polynomial-time algorithm exists for any fixed dimension $d$. \emph{Despite the extensive study in prior literature and established NP-hardness results for unbounded dimensions, the exact computational complexity of this problem for bounded $d$ has remained unresolved.} We resolve this open problem by providing the first exact polynomial-time algorithm via a geometric insight: the parameter space can be partitioned by $\binom{n}{2}$ separator hyperplanes into $O(n^{2d-2})$ regions, within each of which the optimal pairing remains unchanged. This reduces the exponential search over permutations to a polynomial search over regions, yielding the first exact polynomial-time algorithm for SLR with runtime $O(n^{2d-1}(d^2 + \log n))$, without any distributional assumptions. We initiate the study of the second problem, referred to as the \emph{\furlong} (\fur) problem, which considers a more challenging setting where neither the feature vectors themselves nor their correspondence to the response variable are internally aligned. For the \fur\ problem, we provide the first rigorous complexity and algorithmic results, proving that it is NP-hard even at $d = 2$. This establishes a qualitative gap : \fur\ is intractable even in low dimensions, whereas SLR admits an exact polynomial-time solution for fixed $d$.


Localizing and Repairing Sparse-Prompt Failure in SAM Decoders via Box-to-Point Counterfactual

Xiaobing Yu ⋅ Peijie Qiu ⋅ Jin Yang ⋅ Zhaoqi An ⋅ Weiwei Ma ⋅ Xiao Wu ⋅ XUANZHAO DONG ⋅ Wenhui Zhu ⋅ Xiwen Chen ⋅ Xiaoqi Zhao ⋅ Xiaofeng Liu

Promptable segmentation foundation models such as SAM and its variants achieve strong performance with informative prompts, yet exhibit a striking prompt asymmetry where dense prompts like bounding boxes yield accurate segmentation, whereas sparse point prompts substantially degrade performance. This performance gap is widely observed, yet its internal decoder mechanism remains underexplored. In this work, we revisit the failure of point prompts from a causal perspective. By adapting activation patching to SAM-style decoders with a token-aligned intervention scheme, we localize the earliest recoverable divergence to a single early operation, i.e., the first decoder image-to-token cross-attention update, where prompt-conditioned information first enters the image stream. A sparse autoencoder further reveals that this mediation is concentrated in approximately 50 feature directions out of 4,096, and reverse corruption confirms that the identified layer is the entry point of a propagated image-stream route rather than a locally sufficient state. Guided by this localization, we propose the Causal Bottleneck Adapter (CBA), a lightweight residual module placed at the identified entry layer with the base model frozen. Evaluated on 97,996 medical image-target pairs across CT, MRI, ultrasound, and endoscopy, and validated across various SAM-family models, CBA closes 0.891 of the box-to-point gap with fewer than 9K trainable parameters, outperforming prior repair baselines up to 50× larger.


Local linear convergence of gradient methods for overparameterized Gaussian mixtures

Jingxing Wang ⋅ Vasileios Charisopoulos ⋅ Maryam Fazel

We study the problem of learning Gaussian mixture models under overparameterization. Prior work has shown that, while overparameterization is essential for avoiding spurious local optima and enables global recovery of the ground-truth model using the gradient-EM algorithm, it can dramatically slow the local rate of convergence. Under certain assumptions on the mixture weights, we show that a standard divergence measure minimized by statistical learning procedures possesses a manifold of slow growth on which the well-known Polyak stepsize reduces the loss geometrically. We then design a gradient-based method that converges to minimizers at a locally linear rate. Additionally, we show that our method converges to nearly optimal solutions—up to a natural misspecification threshold—for mixtures with arbitrary weights. At a high level, the method alternates between several “short” gradient descent steps that approach the manifold and “long” Polyak steps that contract the distance to minimizers. Our results suggest that slow convergence is not an intrinsic challenge of overparameterization, but can be overcome by exploiting the favorable structure of the loss landscape.


LogT: Logically Think with Images for Visual Search

Yanjun Fu ⋅ Quanzeng You ⋅ Jiadong Guo ⋅ Yujie Lu ⋅ Sai Charitha Akula ⋅ Haiping Wu ⋅ Miao Liu ⋅ Sanghamitra Dutta ⋅ Lu Yuan

Thinking with images empowers vision-language models (VLMs) to use an external zoom-in tool for fine-grained visual search. Recent works employ supervised fine-tuning (SFT) and reinforcement learning (RL) with tool-specific designs to enable such capabilities. However, achieving logical reasoning alongside accurate and efficient tool use remains a significant challenge. To address this, we shift the focus away from increasing RL complexity and instead identify SFT data quality as a key bottleneck. We then propose LogT (short for **Log**ically **T**hink with Images), a fully automated SFT data engine that explicitly enforces logic-chain-guided reasoning. LogT generates an SFT dataset featuring challenging, visually grounded questions and structured tool-use traces with human-like logic. We validate our approach by fine-tuning two base models, Qwen2.5-VL-7B and InternVL3.5-8B, on the dataset and evaluating them across seven benchmark test sets. We show that LogT-SFT based on Qwen2.5-VL-7B outperforms the state-of-the-art trained with both SFT and RL, despite using 3.1$\times$ less data and $>$120$\times$ fewer GPU hours. Furthermore, applying a simple RL recipe without tool-specific design on LogT-SFT yields a model that surpasses prior methods with complex RL setups by a substantial 9.59-point margin. In addition to strong performance gains, our model exhibits higher tool-use accuracy and efficiency, as well as more structured and logically consistent reasoning traces, directly validating the effectiveness of our approach.

Backpropagation (BP) requires feedback (FB) weights to equal the transpose of feedforward (FF) weights at every layer and at every step. This issue is called the weight transport problem, and has been considered to be biologically implausible. Neural circuits are not built from anatomically paired FF/FB axons but from asymmetric projections that form recurrent loops shared across many computations. In multi-nuclei circuits (e.g., thalamo-cortical, cortico-cerebellar, basal-ganglia loops), the return path traverses entirely distinct nuclear structures, making any mechanism for maintaining synchronized weight transposes anatomically absent. Predictive coding (PC), originally proposed as a biologically plausible substitute for BP, silently inherits this constraint despite its biological motivation. Therefore, we propose two PC variants that eliminate it entirely. PC-DH (Dual Hebbian plasticity model) updates FF and FB weights by independent local Hebbian rules with no inter-pathway coordination. PC-RFB (Random FeedBack) fixes the FB weight with a random matrix, testing the robustness of the learning mechanism. Both match baseline PC accuracy (~85\% on MNIST). We identify a phenomenon we termed \textbf{Loop Alignment}: the spontaneous convergence of FF and FB weights toward a transpose relationship, driven by the structure of the inference loop. We then prove that it holds exactly at every training step in both models. To test whether Loop Alignment holds in the multi-nuclei setting, we conduct a circuit-sharing experiment: two networks learn different tasks (Net 1: MNIST; Net 2: Fashion-MNIST) over a shared pathway, where each network's FB is routed through the other's FF weights. Even when the entire shared pathway is plastic and driven simultaneously by both tasks, each network solves its task independently and FF/FB alignment is stronger than in the isolated case. Loop Alignment explains how the brain can learn effectively without weight transport, and support multi-task learning over shared asymmetric circuits.


MAGNET: Manifold-Aware Graph Diffusion Network for Connectome Generation

Protyay Dey ⋅ Ayush Roy ⋅ Hyuna Cho ⋅ Won Hwa Kim ⋅ Vishnu Lokhande

Functional brain connectome represents neural connectivity as a matrix of pairwise interactions between brain regions. Generation of functional connectomes is not only a question of validity; rather, having satisfied the constraints on the correlation matrix, the next step is to recover the class-conditional geometry buried under coarse labels. We propose MAGNET, a Manifold-Aware Graph Diffusion Network which uses a normalized-Cholesky representation of the manifold of correlation matrices that guarantees validity. MAGNET lifts noisy latent states into ROI-level region tokens and performs denoising with a relational inductive bias over brain atlas regions. To deal with structural problems induced by coarse labels, MAGNET employs class-anchored conditioning, amortized structural bridge, and relevance-preserving corruption. Across ABIDE, ADNI, and OASIS-3, MAGNET consistently achieves favorable results compared to previous manifold-aware approaches, demonstrating improvements of $7-21$\% in class-conditional fidelity ($\alpha,\beta$-F1) across three cohorts and better sampling efficiency. Moreover, while training only with strict binary labels, MAGNET is capable of maintaining clinical heterogeneity through fine substructure of the connectomes in a zero-shot setting, improving subclass covariance alignment ($\lambda$-MSE) by over 30\%. These results suggest that geometric validity is a necessary but insufficient condition for clinical utility in connectome synthesis. Moreover, efforts in making the diffusion denoising class-conditional manifold aware finds utility beyond the highly curved brain connectome generation as this is a critical problem in various general settings.


Markovian Experimental Design under Concept Drift

Ayberk Yarkin Yildiz ⋅ Lili Su ⋅ Carlee Joe-Wong ⋅ Edmund Yeh ⋅ Stratis Ioannidis

We study how to optimally select training samples in the presence of concept drift. We propose a Markovian experimental design framework, in which a learner sequentially collects samples and re-estimates a model through Bayesian linear regression. We show that, when model drift is governed by an unknown mean fixed-point coupled with Gaussian deviations, model posterior estimates are determined by a Kalman filter and the system is asymptotically stable. We also show that the steady-state model posterior can be computed as the solution of the so-called Discrete Algebraic Riccati Equation (DARE), and that D-optimal steady-state and myopic sample selection policies can be computed by solving tractable convex optimization problems. We do so by characterizing the curvature of the DARE solution, even though the latter generally has no closed form. Finally, we show that our approach generalizes to the non-linear model setting by replacing the Kalman filter with an extended Kalman filter.


MARS: Harmonizing Multimodal Convergence via Adaptive Rank Search

Minkyoung Cho ⋅ Insu Jang ⋅ Shuowei Jin ⋅ Zesen Zhao ⋅ Adityan Jothi ⋅ Ethem F Can ⋅ Min-Hung Chen ⋅ Zhuoqing Morley Mao

Fine-tuning Vision Language Models (VLMs) with parameter-efficient methods like Low-Rank Adaptation (LoRA) is crucial for task adaptation. However, imbalanced training dynamics across modalities often lead to suboptimal accuracy due to negative interference, a challenge typically addressed with inefficient heuristic methods such as manually tuning separate learning rates. To overcome this, we introduce MARS (Multimodal Adaptive Rank Search), an approach to discover optimal rank pairs that balance training dynamics while maximizing performance. Our key innovation, a proposed framework of dual scaling laws, enables this search:~one law models module-specific convergence time to prune the search space to candidates with aligned dynamics, while the other predicts final task performance to select the optimal pair from the pruned set. By re-purposing the LoRA rank as a controller for modality-specific convergence speed, MARS outperforms baseline methods and provides a robust, automated strategy for optimizing VLM fine-tuning.

The ecosystem of Lean and Mathlib has become the de facto standard for large language model (LLM) assisted formal reasoning with remarkable successes in recent years. Those successes, however, only consume Mathlib as an essential dependency but do not directly contribute to it. In the meantime, the growth of Mathlib has recently been bottlenecked by the review process, which requires human reviewers to judge whether proposed pull requests (PRs) follow the Mathlib's conventions and are worth integrating as part of a shared mathematical infrastructure. This leads to our central question: can LLMs help review Mathlib PRs? To this end, we introduce MathlibPR, a benchmark built from real Mathlib4 PR histories. We further propose a staged evaluation protocol and use it to evaluate both LLM models (e.g., DeepSeek, Qwen, Goedel, and Kimina) and LLM agents (e.g., Codex and Claude Code). Surprisingly, both LLM models and LLM agents struggle to distinguish merge-ready PRs from build-passing PRs that were revised or never merged. By turning Mathlib PR histories into a supervised signal, MathlibPR provides a step toward reviewer assistants and reward models that could help evaluate PRs and steer LLMs toward producing merge-ready Mathlib contributions.


MCM-DM: Towards Better Spatio-Temporal Event Representation Learning via Discrete Morse Theory

Yuanheng Zhang ⋅ Jennifer Rozenblit ⋅ Chenguang Yang ⋅ Jaidev Goel ⋅ Yulia Gel ⋅ Yuzhou Chen

Spatio-temporal point processes (STPPs) represent a random collection of points where each point corresponds to the time and location of an event. STPPs are widely used to describe a wide range of phenomena, from earthquakes to wildfires to crime occurrence to disease outbreaks. Generative models have recently emerged as a new powerful paradigm for STPPs, demonstrating %due to their exceptional highly competitive generalization capabilities and promising potential for systematic uncertainty quantification. However, existing generative approaches primarily rely on unimodal numerical data, overlooking the geographic context in which events occur and failing to capture critical spatial and temporal patterns within the STPP. To address these limitations, we propose Morse-aware Cross-Modality learning within Diffusion Model, or \textbf{MCM-DM}. MCM-DM represents a cross-modality diffusion framework based on the discrete Morse and cobordism theories that augments STPP modeling with geographic scene understanding. MCM-DM consists of three key components: (i) a vision-language model-based geographic encoder that extracts semantic embeddings from satellite tiles at each event location; (ii) an attention-based fusion mechanism to integrate critical structure representation with spatio-temporal embeddings, and (iii) a Morse-theoretic topological aligner which aligns latent representations of critical events with spatio-temporal embedding space. We demonstrate the MCM-DM utility on 8 diverse STPP datasets from Earth sciences, epidemiology, urban mobility, and crime analytics, highlighting its cross-domain versatility. The code is available at~\url{https://anonymous.4open.science/r/MCM-DM-CF83}.


MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers

Chaithanya Bandi ⋅ Razvan Dumitru ⋅ Ben Hertzberg ⋅ Divyansh Agarwal ⋅ Geobio Boo ⋅ Tejas Polakam ⋅ Jeff Da ⋅ HiJae Kim ⋅ Vipul Gupta ⋅ Manasi Sharma ⋅ Andrew Park ⋅ Martin Villanueva Dimakis ⋅ Ernesto G Montoya ⋅ Rafael Cruz ⋅ MohammadHossein Rezaei ⋅ Chetan Rane ⋅ Daniel Y Zhang ⋅ Brad Kenstler ⋅ Bing Liu

The Model Context Protocol (MCP) is emerging as a standard interface through which large language model (LLM) agents discover and invoke external tools. However, existing MCP evaluations fall short along three key axes: realistic multi-step workflows with cross-server orchestration, breadth across authentic MCP servers rather than mocks, and structured, reproducible claim-level scoring disentangled from agent verbosity or style. We introduce MCP-Atlas, a benchmark for measuring tool-use competency against production MCP servers. MCP-Atlas contains 1,000 natural-language tasks written and verified by human experts spanning 36 real MCP servers and 220 tools. Prompts do not specify servers, tools, or parameters, requiring agents to identify relevant tools among semantically plausible distractors and to compose multi-step, cross-server workflows. Each task is scored with a claim-level rubric, where final answers are scored against atomic factual claims grounded in tool outputs. This answer-centric scoring permits valid alternative tool-call trajectories to receive credit. We pair this with an 11-category diagnostic taxonomy that disentangles tool-call failures from cognitive failures in task understanding, synthesis, parsing, and stopping. Evaluating 20 frontier models from six providers under matched task-level conditions, we find pass rates up to 82.2% at a 0.75 claim coverage threshold and a clear three-tier performance structure. Automated diagnostics show that 63.3% of diagnosed failures are cognitive rather than tool-call related. Notably, several high-performing models fail after successful tool execution due to premature stopping or incorrect synthesis. We release the task schema, containerized harness, claim evaluator, and a 500-task public split, while reserving a 500-task private split to preserve leaderboard integrity.

When a multimodal AI agent is asked to forget a fact, current memory systems usually delete the text entry and report success. We find that the fact can remain recoverable from retained user images, including images tagged to entirely different facts, because VLMs use implicit visual cues at inference time. We introduce the Information Provenance Graph (IPG), a taxonomy that classifies memory representations by deletion affordance. The IPG reveals that deletion fails through multiple channels. Our benchmark, MemLeak, measures this across a deletion cascade: direct probing of deletion-capable systems yields <1%, but retained correlated text enables 18.3% recovery, and retained images enable 12.0% recovery (0.0% blind baseline, 0.3% FPR)---with 47% of image leaks not text-recoverable. Content-aware semantic deletion reduces the image residual to 2.0%. The residual appears across multiple VLMs, a production memory system, and real Unsplash-licensed photographs. Dual-annotator human validation ($\kappa$ = 0.88) confirms judge reliability.

Missing data is pervasive in many scientific domains such as public health, environmental science, and the social sciences. Recoverability from missing data is typically studied using fully specified variable-level missingness models despite that, in many applications, only coarse structural information is available, for instance when variables are grouped into clusters due to limited knowledge or interpretability reasons. In this paper, we investigate recoverability from such abstract representations. We introduce two classes of cluster-based missingness graphs: the m-C-DMG, which retains variable-specific missingness indicators, and the cm-C-DMG, which aggregates missingness mechanisms at the cluster level. We formalize the notion of compatibility between these abstract graphs and underlying variable-level missingness models, and study how this abstraction affects the recoverability of probabilistic and causal queries. In particular, we give graphical conditions of recovering the joint distribution as well as graphical conditions of recovering a macro causal effect. Overall, our results clarify when cluster-level missingness information is sufficient for valid inference, and when finer-grained modeling is necessary.


Mixture of Layers: Dynamic Layer Routing for Visual Reasoning

Jeonghwan Kim ⋅ Sofia Stoica ⋅ Jiwan Chung ⋅ Ansel Blume ⋅ Hyeonjeong Ha ⋅ Zhenhailong Wang ⋅ Xin Dong ⋅ Heng Ji

Pre-trained vision encoders contain layer-wise visual representations that differ in spatial granularity, semantic abstraction, and sensitivity to local details. However, most Multimodal Large Language Models (MLLMs) rely only on the final or penultimate vision encoder representations or fixed aggregation rules, making visual abstraction largely query-agnostic and limiting access to fine-grained cues such as small objects, spatial details, text, and subtle visual attributes. In this work, we propose Mixture of Layers (MoL), an instruction-conditioned layer routing approach that dynamically aggregates query-relevant latent representations from intermediate vision encoder layers. Given a text query, MoL predicts routing probabilities over vision encoder layers and aggregates selected hidden states at either the image level, patch level, or through a hybrid routing mechanism. Our proposed MoL is a vision encoder-agnostic framework for layer-wise routing that selectively samples visual features useful for fine-grained visual reasoning tasks. Experiments across seven fine-grained visual reasoning benchmarks demonstrate substantial performance improvements, including +18.9% on V* overall accuracy, +4.5% on HRBench4K, and +16.3% on CharXiv compared to baseline MLLMs, without requiring multi-resolution inputs, simple interleaving of multiple vision encoders, or additional patch tokens. We further analyze receptive field scales and routing behaviors across vision encoder layers to explain why adaptive layer selection improves perception and reasoning. Our results suggest that conditional intermediate-layer representations are a key step toward stronger visual perception and reasoning in MLLMs. Code will be open-sourced.


MLLMs Fail to Refuse when Using Tools Agentically

Rikiya Takehi ⋅ Ryo Hachiuma ⋅ Shaona Ghosh ⋅ Dan Zhao ⋅ Frank Wang ⋅ Yusuke Hirota

Agentic multimodal large language models (MLLMs) have recently pushed the frontier of visual reasoning by calling tools such as zooming and tagging. Despite the recent strong success of agentic MLLMs, this work uncovers a critical safety failure in the tool-use paradigm: agentic tool-using MLLMs become less capable of refusing harmful requests. Our experiments confirm that, across three popular safety benchmarks, all the top open- and closed-weight MLLMs we test exhibit significantly lower safety in tool-using settings than in non-tool settings, with a relative refusal failure rate increase of up to 68.7%. Based on analysis of 100,000+ responses, including extended experiments, we also propose two possible reasons for this safety degradation.

Most work on improving large language models treats accuracy as the sole objective. We argue that the harness — the Python code surrounding the model that constructs prompts, routes calls, and parses outputs — is a first-class design surface whose quality is inherently multi-objective: an accurate harness that refuses no unsafe request, or that consumes an order of magnitude more tokens, is not a good harness. We present Meta-Harness, a system that casts harness design as search over three per-domain objectives — accuracy, behavioural safety, and token cost — solved by an agentic proposer (Claude Code) with full filesystem access to prior harness source, execution traces, and scoring artifacts. Our central finding is that a single-phase joint-reward proposer (MoMHa) outperforms every alternative, including a two-phase "accuracy then tokens" ablation, scalar-only feedback, and an accuracy-only baseline. We evaluate on seventeen domains: seven synthetic capability suites, seven real-world public benchmarks (HumanEval, MBPP, Spider, FEVER, MMLU-Pro, LawBench, NuminaMath), and three U-SafeBench-derived user-specific safety domains, using a 12-model fleet spanning four families. On the synthetic track MoMHa achieves a joint mean of 0.482 versus 0.198–0.422 for ten baselines, winning 7/10 per-domain columns; on the real-world track it scores 0.461 versus 0.377 for the strongest baseline (DSPy), winning 5/7 columns — demonstrating that harness strategies transfer to unseen benchmarks without retraining on 8 of 12 target models. MoMHa attains the highest measured behavioral safety composite (U-SafeBench, 0.781) and uses 95 fewer tokens per example than the two-phase alternative. We will release all harness code, evaluation infrastructure, and cross-model logs.


Monoculture Robust Learning

Lee Cohen ⋅ Jon Kleinberg ⋅ Omer Reingold

We introduce a notion of monoculture robustness in Probably Approximately Correct (PAC) learning, as a way to mitigate the problem of *algorithmic monoculture* (Kleinberg & Raghavan, 2021): the tendency of models sharing the same components (e.g., training data) to fail on the same inputs. Specifically, we ask whether a learner with access to a single dataset can output $k$ individually accurate hypotheses whose joint failure probability is comparable to training on $k$ independent datasets. We say that such a learning procedure achieves *monoculture robustness*, and we quantify this by bounding the *conjunctive error* that all $k$ hypotheses misclassify points in an individual sense. We design methods for achieving this type of guarantee by drawing a surprising connection between replicability (Impagliazzo et al., 2022) and monoculture robustness. Specifically, we show that if the base learner is $\rho$-replicable, then running it independently $k$ times on the same sample yields individual monoculture error that shrinks exponentially with $k$, controlled by the single-run error of the base learner at the point and $\rho$. We complement these positive results with a lower bound: for individual monoculture robustness, we prove a worst-case lower bound showing that, for learners, driving the individual conjunctive error down to $\gamma$ can require a sample-size multiplier of $t = \Omega(\log(1/\gamma))$.


MOSAIC-CONUS: A Multimodal, Multi-Temporally Paired dataset for Earth Sciences

Abhishek Potnis ⋅ Youssef Hussein ⋅ Waqwoya Abebe ⋅ JangHyeon Lee ⋅ Debvrat Varshney ⋅ Jacob Arndt ⋅ Philipe Dias ⋅ Aristeidis Tsaris ⋅ Dan Lu ⋅ Dalton Lunga

Earth embeddings—vector representations of geographic locations indexed in space and time—are emerging as a unifying interface for geospatial AI. However, their quality depends not only on model design, but on how multimodal Earth observation (EO) data are spatially indexed, temporally aligned, and cross-modally associated during pretraining. We introduce MOSAIC-CONUS ($\textbf{M}$ultimodal $\textbf{O}$bservations with $\textbf{S}$patially $\textbf{A}$ligned Imagery, Urban Points of Interest, $\textbf{I}$n-Situ Measurements and Text $\textbf{C}$aptions), a large-scale EO dataset over the contiguous United States, organized around 250{,}000 stratified point indices that serve as stable spatial keys across seven modalities: active radar, passive optical imagery, lidar-derived elevation, land cover, functional context, hydrometeorological measurements, and textual summaries. Unlike Existing EO datasets, MOSAIC-CONUS introduces four contributions not jointly addressed in prior work: $\textbf{1.}$an open-source, large-scale multimodal EO corpus structured around point-indexed data designed to support Earth embedding learning; $\textbf{2.}$explicit radar-optical pairing tables spanning twelve temporal alignment regimes, formalizing cross-sensor alignment as a controllable variable for analyzing how temporal mismatch across modalities influences learned embeddings quality; $\textbf{3.}$a benchmark suite spanning cross-modal retrieval, annual nightlights regression, and basin-held-out streamflow prediction, positioning MOSAIC-CONUS as a benchmark-ready resource for multimodal AI systems; and $\textbf{4.}$a language-based embedding layer through co-registered textual summaries, enabling Earth embeddings to function as a queryable interface for agentic AI systems. The dataset and pairing protocols are publicly released.


MOSAIC: Module Discovery via Sparse Additive Identifiable Causal Learning for Scientific Time Series

Shicheng Fan ⋅ Nour Elhendawy ⋅ Jianle Sun ⋅ Ke Fang ⋅ Kun Zhang ⋅ Yihang Wang ⋅ Lu Cheng

Causal representation learning (CRL) seeks to recover latent variables with identifiability guarantees, typically up to permutation and component-wise reparameterization under appropriate assumptions. However, identifiability does not imply interpretability: latent semantics are typically assigned post hoc by alignment with known ground-truth factors. This limitation is particularly acute in scientific time series, where underlying mechanisms are unknown and discovering interpretable structure is a primary goal. In contrast, scientific observations (such as residue-pair distances, climate indices, or process sensors) are inherently semantic, as they correspond to named physical quantities. This raises a key question: can the interpretability of observations be transferred to the identifiable latent space? We propose MOSAIC (Module discovery via Sparse Additive Identifiable Causal learning), a sparse temporal VAE that integrates temporal CRL identifiability with support recovery over observed variables. MOSAIC identifies latent variables via regime-conditioned temporal variation, and recovers for each latent a sparse set of associated observations through an additive decoder, yielding module-level interpretability. We show that ANOVA main-effect supports are identifiable under general smooth mixing functions, and provide finite-sample recovery guarantees for a tractable sparse-additive variant. Empirically, MOSAIC recovers domain-consistent variable groups across RNA molecular dynamics, solar wind, ENSO climate, the Tennessee Eastman process, and a synthetic tokamak benchmark, enabling interpretable discovery of latent mechanisms in scientific time series.


MPQ: A Message-Passing View of Post-Training Quantization

Yanshuo Chen ⋅ chenxi liu ⋅ Reza Shirkavand ⋅ Heng Huang

Post-training quantization (PTQ) is a key technique for efficiently deploying large language models, as it compresses weights into low-bit representations without retraining. Existing PTQ solvers, including GPTQ, LDLQ, and block coordinate descent, commit each weight block to a single codeword before its neighbors are considered, leaving near-optimal alternatives unexplored. In this work, we cast row-wise PTQ as posterior inference over discrete codebooks, in which existing solvers correspond to a zero-temperature, hard-decision limit retaining only the top codeword per block. We propose \emph{Message-Passing Quantization} (MPQ), which preserves soft posterior information across blocks during inference and commits to hard codewords only at the output stage. MPQ first uses vector approximate message passing to compute approximate blockwise posteriors, then refines a hard-decision initialization by exact block coordinate descent. Across 2-bit lattice and 3- and 4-bit scalar quantization on dense language models from Llama, Qwen, and Mistral families, MPQ improves perplexity and zero-shot accuracy over strong hard-decision baselines, and composes favorably with preprocessing techniques such as activation-aware scaling and incoherence processing.

Judging image quality is not only ecologically relevant to everyday humans tasks, but also underpins many machine vision tasks such as image generation. This paper proposes a framework to understand the inherent perceptual space underlying image quality judgment in humans. We propose a multi-dimensional observer model that represents images as distributions in a latent perceptual space and that models human judgment as comparing noisy samples. Being constrained by neural representations in the primate ventral stream and fit to large-scale behavioral data, the model enables analysis of perceptual structure while matching the predictive power of existing metrics. Using this model, we find that the perceptual spaces needed to account for image quality judgment in humans are extremely low-dimensional compared to the image space even when considering its sparsity. The exact structure of the space (e.g., dimensionalities, information encoded) varies between low-level and high-level quality judgments, suggesting that, despite a shared retinal encoding in the beginning, humans selectively construct task-dependent perceptual spaces in visual decision making.


Nearly Optimal Attention Coresets

Alexandr Andoni ⋅ Eldar Kleiner ⋅ Edo Liberty

We consider the problem of estimating the Attention mechanism in small space, and prove the existence of coresets for it of nearly optimal size. Specifically, we show that for any set of unit-norm keys and values $(K,V)$ in $R^d$, there exists a subset $(K',V')$ of size at most $O({\sqrt{d} e^{\zeta+o(\zeta)}/\epsilon})$ such that $$ \left\| Attn(q,K,V)- Attn(q,K',V') \right\| \le \epsilon $$ simultaneously for all queries whose norm is bounded by $\zeta$. This outperforms the best known results for this problem. We also offer an improved lower bound showing that $\epsilon$-coresets must have size $\Omega({\sqrt{d} e^{\zeta}/\epsilon})$.


NitroBox: Lightning-Fast Sandbox for Large-Scale RL Training

Yuzhou Nie ⋅ Ruilin Zhou ⋅ Zhaorun Chen ⋅ Jingyang Zhang ⋅ Yan Shao ⋅ Hongwei Li ⋅ Bo Li ⋅ Dawn Song ⋅ Wenbo Guo

Recent work on agentic training efficiency has mainly focused on improving learning algorithms, rollout strategies, or model-hardware co-design. However, these approaches overlook an important source of training latency: the sandbox that runs the environments. Recent studies show that creating, executing, and resetting containers accounts for an important portion of the rollout process in RL training, substantially slowing down overall training. This overhead largely stems from the centralized management mechanisms (e.g., DAEMON) of existing sandboxes, along with additional features (e.g., multi-tenancy) that are unnecessary for reinforcement learning (RL). Moreover, existing sandboxes are designed for general-purpose isolated execution, making them a poor fit for the requirements of RL rollout and training. We propose NitroBox, a novel sandbox system purpose-built for agentic RL that significantly accelerates the rollout and training process. Technically speaking, we first comprehensively redesign NitroBox's lifecycle to eliminate the dependency on DAEMON, so that every sandbox instance is fully independent during creation, startup, and reset (e.g., via rootless full-namespace isolation). Furthermore, we introduce a copy-on-write mechanism to enable fast startup across multiple instances. Second, we knock off features that are time-consuming but unnecessary for RL — such as multi-tenancy — while preserving the core functionality and security guarantees needed for RL training. Finally, we introduce some RL-nature features to further improve efficiency, such as optional CRIU-based process checkpoint saving and restoring for mid-trajectory branching. Our experiments show that NitroBox is faster than Docker across the entire lifecycle: 10.6× faster in sandbox creation, 24.6× in reset, and 19.1× in deletion. When training Qwen2.5-Coder-32B-Instruct with GRPO at c=32 concurrency, end-to-end wall-clock training time drops from 40.6 hours to 28.4 hours when replacing Docker with NitroBox, while preserving the same task pass rate.


Nonconvex Decentralized Stochastic Bilevel Optimization under Heavy-Tailed Noise

Xinwen Zhang ⋅ Yihan Zhang ⋅ Heng Liang ⋅ Hongchang Gao

Existing decentralized stochastic optimization methods assume the lower-level loss function is strongly convex and the stochastic gradient noise has finite variance. These strong assumptions typically are not satisfied in real-world machine learning models. For example, learning on language data typically leads to heavy-tailed gradient. To address these limitations, we develop a novel decentralized stochastic bilevel optimization algorithm for the nonconvex bilevel optimization problem under heavy-tailed noise. Specifically, we develop a normalized stochastic variance-reduced bilevel gradient descent algorithm, which does not rely on any clipping operation. Moreover, we establish its convergence rate by innovatively bounding interdependent gradient sequences under heavy-tailed noise for nonconvex decentralized bilevel optimization problems. As far as we know, this is the first decentralized bilevel optimization algorithm with rigorous theoretical guarantees under heavy-tailed noise. The extensive experimental results confirm the effectiveness of our algorithm in handling heavy-tailed noise.


Non-Linear Pricing Restores Tractability for a Data Seller

Bhaskar Ray Chaudhury ⋅ Jugal Garg ⋅ Eklavya Sharma ⋅ Jiaxin Song

We consider a data seller who designs pricing mechanisms over multiple datasets to maximize revenue from budget-constrained buyers. The seller offers multiple datasets and assigns each a pricing function that maps the quantity purchased to a total payment. The goal is to design these pricing functions to maximize revenue, anticipating that buyers—who trade off accuracy gains against cost—choose bundles optimally subject to their budget constraints. Prior work [Chaudhury et al., 2026] studies such optimal pricing under the restriction that each dataset is assigned a linear price, and shows that computing optimal linear prices is computationally intractable. In contrast, we allow each dataset to be priced via a general function and show that this additional flexibility can not only increase the revenue but also restore tractability, yielding a surprising simultaneous improvement in economic performance and computational efficiency. Even when pricing functions are only required to be monotone and lower-continuous, optimal pricing admits a highly structured and simple form: each pricing function is piecewise linear and convex (PLC), and the optimal solution can be computed in polynomial time. Moreover, the total number of kinks across all pricing functions is bounded by the number of buyers. Consequently, when datasets significantly outnumber buyers, most pricing functions are effectively linear. We further empirically study the structure of optimal pricing by analyzing the number of kinks and the revenue gap between optimal nonlinear pricing and optimal linear pricing on simulations generated from a real dataset.


Normalizing Trajectory Models

Jiatao Gu ⋅ Tianrong Chen ⋅ Ying Shen ⋅ David Berthelot ⋅ Shuangfei Zhai ⋅ Joshua Susskind

Diffusion-based models decompose sampling into many small Gaussian denoising steps—an assumption that breaks down when generation is compressed to a few coarse transitions. Existing few-step methods address this through distillation, consistency training, or adversarial objectives, but sacrifice the likelihood framework in the process. We introduce Normalizing Trajectory Models (NTM), which models each reverse step as an expressive conditional normalizing flow with exact likelihood training. Architecturally, NTM combines shallow invertible blocks within each step with a deep parallel predictor across the trajectory, forming an end-to-end network trainable from scratch or initializable from pretrained flow-matching models. Its exact trajectory likelihood further enables self-distillation: a lightweight denoiser trained on the model's own score produces high-quality samples in four steps. On text-to-image benchmarks, NTM matches or outperforms strong few-step baselines in just four sampling steps while uniquely retaining exact likelihood over the generative trajectory.


OASES: Outcome-Aligned Search-Evaluation Co-Training for Agentic Search

Erhan Zhang ⋅ Yiqun Chen ⋅ Zechun Niu ⋅ Wei Yang ⋅ Xiaochi Wei ⋅ Yan Gao ⋅ YIWU ⋅ Yao Hu ⋅ Jiaxin Mao

Agentic search enables language models to solve knowledge-intensive tasks by adaptively acquiring external evidence over multiple steps. Reinforcement learning with verifiable rewards (RLVR) has emerged as a widely adopted training paradigm for search agents, yet outcome-only rewards are sparse and provide limited credit assignment for intermediate search actions. Existing process-reward methods therefore seek to densify supervision through proxy signals, external evaluators, or likelihood-based information gain. However, proxy rewards can deviate from the final outcome objective, while fixed evaluators can become stale as the search policy evolves, leading to unreliable process supervision. To address these challenges, we propose OASES, an Outcome-Aligned Search-Evaluation Supervision framework for agentic search. OASES derives outcome-aligned process rewards by evaluating how well each intermediate search state supports answering the original question. It further co-trains the search policy and the state evaluator on policy, allowing the evaluator to adapt to evolving search behavior and provide more reliable process rewards. Experiments on five multi-hop QA benchmarks show that OASES consistently outperforms strong RL baselines, with further analyses confirming the benefits of outcome-aligned process rewards and search-evaluation co-training.


OATS: Online Data Augmentation for Time Series Foundation Models

Junwei Deng ⋅ Chang Xu ⋅ Jiaqi Ma ⋅ Ming Jin ⋅ Chenghao Liu ⋅ Xu Zhang ⋅ Li Zhao ⋅ Jiang Bian

Time Series Foundation Models (TSFMs) are a powerful paradigm for time series analysis and are often enhanced by synthetic data augmentation to improve the training data quality. Existing augmentation methods, however, typically rely on heuristics and static paradigms. Motivated by dynamic data optimization, which shows that the contribution of samples varies across training stages, we propose OATS (Online Data Augmentation for Time Series Foundation Models), a principled strategy that generates synthetic data tailored to different training steps. OATS leverages valuable training samples as principled guiding signals and dynamically generates high-quality synthetic data conditioned on them. We further design a diffusion-based framework to produce realistic time series. Experiments on TSFMs demonstrate that OATS consistently outperforms regular training and yields substantial performance gains over static data augmentation baselines across six validation datasets and two TSFM architectures. The code is available at the link https://anonymous.4open.science/r/OATS-536E.

We study offline-online reinforcement learning in linear mixture Markov decision processes (MDPs) under environment shift. In the offline phase, data are collected by an unknown behavior policy and may come from a mismatched environment, while in the online phase the learner interacts with the target environment. We propose an algorithm that adaptively leverages offline data. When the offline data are informative, either due to sufficient coverage or small environment shift, the algorithm provably improves over purely online learning. When the offline data are uninformative, it safely ignores them and matches the online-only performance. We establish regret upper bounds that explicitly characterize when offline data are informative, together with nearly matching lower bounds. Numerical experiments further corroborate our theoretical findings.


OmniMechanism Design for Human-AI Collaboration with Impressionable Minds

Tianwei Gao ⋅ Yifei Pang ⋅ Lunjia Hu ⋅ Kate Donahue ⋅ Lihua Lei ⋅ Vahid Tarokh ⋅ Zhun Deng

As AI systems are increasingly deployed in high-stakes domains such as healthcare, effectively integrating human expertise with AI outputs is critical for reliable decision-making. A substantial body of theoretical work studies optimal mechanisms for human-AI collaboration (HAC), focusing on when to involve humans and how to combine their expertise with AI predictions. However, these approaches largely overlook a key factor highlighted by empirical evidence: human judgments are impressionable and can vary depending on how AI outputs are presented. For example, when AI systems defer decisions to humans, revealing or withholding the AI prediction can significantly influence human judgments. Moreover, existing work neglects deployment-time variability, where conditions, particularly the cost of eliciting human input, may differ substantially from those observed during training. To address these challenges, we introduce a rigorous and theoretically grounded framework for human-AI collaboration, termed \emph{collaborative omnimechanism design}, based on a novel and computationally efficient contextual bandit variant of boosting for outcome indistinguishability. In particular, we reduce a calibration-type outcome indistinguishability condition to projected smooth calibration, providing novel time-efficiency result, bypassing the high-dimension curse induced by complicated collaboration methods. Our framework provides a principled characterization of human impressionability and introduces a unified optimization scheme that enables robust and adaptable decision-making across diverse query costs and evaluation objectives without retraining. We further validate our theoretical results through comprehensive experiments.


OmniMem: Scalable and Adaptive Memory Retrieval for Long Video Generation

Lin Zhao ⋅ Yushu Wu ⋅ Yifan Gong ⋅ Yanzhi Wang ⋅ Pu Zhao

Autoregressive (AR) video generation extends videos by producing latent chunks sequentially, but scaling to long videos requires repeated access to a growing historical KV cache. Existing methods reduce this cost by truncating the KV cache or compressing it into implicit memory, but both lose explicit access to query-relevant historical details. We propose OmniMem, an explicit full-range memory retrieval framework that performs sparse KV retrieval over the historical cache. To make this practical for chunk-based AR video generation, OmniMem addresses two issues: (i) local bias in sparse KV selection and (ii) Union Explosion in memory access. Adaptive Window Exclusion removes local-window blocks from the selection candidates when sufficient long-range history is available, preserving the sparse budget for informative long-range retrieval. Query-shared KV Selection reduces cross-query diversity, while Per-Head Scattered KV Access avoids expanding head-specific selections into a large selected KV buffer. This allows each attention head to retrieve non-contiguous KV blocks according to its own selection pattern. Experiments on long-video generation show that OmniMem improves Dynamic Degree by 52.3\% and preserves strong consistency over strong baselines, while maintaining comparable memory usage.


OmniShotCut: Holistic Relational Shot Boundary Detection with Shot-Query Transformer

Boyang Wang ⋅ Guangyi Xu ⋅ Jiahui Zhang ⋅ Zhipeng Tang ⋅ Zezhou Cheng

Shot Boundary Detection (SBD) aims to automatically identify shot changes and divide a video into coherent shots. While SBD was widely studied in the literature, existing methods often produce non-interpretable boundaries on transitions, miss subtle yet harmful discontinuities, and rely on noisy, low-diversity annotations and outdated benchmarks. To alleviate these limitations, we propose OmniShotCut to formulate SBD as structured relational prediction, jointly estimating shot ranges with intra-shot relations and inter-shot relations, by a shot query-based dense video Transformer. To avoid imprecise manual labeling, we adopt a fully synthetic transition synthesis pipeline that automatically reproduces major transition families with precise boundaries and parameterized variants. We also introduce OmniShotCutBench, a modern wide-domain benchmark enabling holistic and diagnostic evaluation. Experiments on the benchmarks demonstrate the effectiveness and generality of our method.


One Model, Two Roles: Emergent Specialization in a Shared Recurrent Transformer

Jucheng Shen ⋅ Barbara Su ⋅ Anastasios Kyrillidis

Can a shared-weight recurrent Transformer develop distinct internal roles without being partitioned into separate modules? We study this in Asymmetric Input Recurrence (AIR), a minimal two-state reasoning architecture in which the same Transformer model is reused for both updates (per literature, L and H) and the only built-in asymmetry is that the encoded input is injected during L-updates but not H-updates. Across Sudoku-Extreme and Maze, decoded rollouts reveal a consistent split: $z_H$ behaves like a fully committed proposal state, whereas $z_L$ retains local uncertainty and shifting intermediate structure. Freeze experiments show that this split is, in practice, related to the model's state dynamics: in Sudoku, freezing $z_H$ reduces $z_L$'s content changes whereas freezing $z_L$ increases $z_H$'s, while in Maze, freezing either state increases content changes in the other state. Ablations show that to induce specialization, the shared model need to be able to tell the two update types apart, either from input injection asymmetry or from a separate level token. Mechanistically, attention analysis shows that L-updates are consistently more local than H-updates in both Sudoku and Maze. Together, these results show that, in a two-state recurrent setting, a clear state-identity signal can induce stable, related functional roles inside a shared-parameter recurrent Transformer.


Online Learning on Hidden-Convex Losses via Algorithmic Equivalence: Optimal Regret, Geometric Barrier, and Bandit Feedback

Anas Barakat ⋅ Andreas Kontogiannis ⋅ Vasilis Pollatos ⋅ Ioannis Panageas ⋅ Antonios Varvitsiotis

We study adversarial online learning with hidden-convex losses, i.e., nonconvex losses that become convex after a nonlinear reparameterization. Ghai, Lu, and Hazan (NeurIPS 2022) proved that, under geometric and smoothness assumptions, online gradient descent (OGD) on such nonconvex losses approximately simulates online mirror descent (OMD) on the underlying convex losses with a suitable regularizer, yielding $\mathcal{O}(T^{2/3})$ regret. They left open whether the optimal $\Theta(\sqrt{T})$ regret from online convex optimization can be recovered in this hidden-convex setting. We answer this question affirmatively. More specifically, via a sharper discrete-time algorithmic equivalence argument, we prove that OGD achieves $\mathcal{O}(\sqrt{T})$ regret under the same assumptions, matching the optimal worst-case rate for adversarial online convex optimization. We also address another open question of Ghai, Lu, and Hazan by clarifying the geometry required for this algorithmic equivalence. We replace the diagonal-Jacobian sufficient condition with a necessary-and-sufficient Hessian compatibility condition, thereby expanding the class of admissible reparameterizations. We complement our tight regret bound with a lower bound showing that the Hessian compatibility assumption is essential for OGD; when it fails, we construct a smooth reparameterization and an adversarial sequence of hidden-convex losses for which OGD suffers $\Omega(T)$ regret. Finally, we extend our analysis to one-point bandit feedback and prove a $\mathcal{O}(T^{3/4})$ expected regret bound for bandit OGD with spherical smoothing, matching its classical rate on convex losses.


On the Fragility of Latent Knowledge: Layer-wise Influence under Unlearning in Large Language Model

Jianing Zhu ⋅ Zongze Li ⋅ Chandler Squires ⋅ Qizhou Wang ⋅ Bo Han ⋅ Pradeep Ravikumar

Large language model (LLM) unlearning has emerged as an essential post-training mechanism for erasing specific knowledge or undesirable behaviors. However, forgetting target data often causes an unintended degradation in overall model utility. Although various advanced methods have explored different learning objectives to mitigate the trade-off, it remains unclear how the highly entangled internal representations in LLMs contribute to unlearning. In this work, we introduce the notion of latent knowledge fragility to explore the vulnerability of retained knowledge to unlearning. We develop a unified analytical approach via component-wise parameter patching that isolates and quantifies fragility in terms of different transformer blocks. We observe that the LLM encodes different levels of abstraction, from surface syntax in shallow layers to complex semantics in deeper layers, which align with different degrees of representation disruption and utility degradation. Based on the insights, we propose a lightweight framework called Component-wise Replacement Unlearning (CRU) that restores fragile layers (also extendable to other components) from the original model based on post-hoc validation, which allows us to obtain a hybrid model without additional training. Extensive experiments on various aspects verify that CRU generally improves the trade-off between removal and retention with non-uniform unlearning influence.


On the Geometry and Latent-Space Composition of Hypernetwork-Generated LoRAs

Yanlai Yang ⋅ Minki Kang ⋅ Divyam Madaan ⋅ Sung Ju Hwang ⋅ Mengye Ren

Recent work proposes hypernetworks that convert a document into a low-rank parameter update (LoRA) in a single forward pass, enabling fast injection of new knowledge into a frozen LLM. However, storing separate LoRA adapters for each document can be unsustainable in streaming settings. In this paper, we study the problem of composing hypernetwork-generated LoRAs using two representative systems, Doc-to-LoRA and SHINE. We show that hypernetwork-generated LoRAs have a geometry sharply different from directly fine-tuned LoRAs. Hypernetwork-generated LoRAs are highly spectrally concentrated and strongly aligned, while next-token prediction and QA fine-tuned adapters are diffuse and nearly orthogonal across documents. This suggests that hypernetwork-generated LoRAs are not arbitrary low-rank weight updates, but decoded points on a structured adapter manifold. Guided by this observation, we propose latent-space LoRA merging: instead of merging in the weight space, we compose the intermediate hypernetwork latents and then decode the merged latent through the frozen LoRA-generation head. We show that latent-space merging techniques consistently improve over weight-space and factor-space baselines while preserving a single fixed-rank adapter. These results establish latent-space composition as a promising mechanism for bounded accumulation of hypernetwork-generated LoRAs in long document sequences.


On the Invariance and Generality of Neural Scaling Laws

Xing Han ⋅ Liu Ziyin ⋅ Suchi Saria ⋅ Paul Liang

Neural scaling laws establish a predictable relationship between model performance and data or compute, offering crucial guidance for resource allocation in new domains and tasks. Yet such laws are most needed precisely where they are hardest to obtain: fitting one for a new model–task pair demands expensive sweeps that typically exhaust the very compute budget the law is meant to economize. This paper poses the research question of how to develop generalizable scaling laws: laws fit once on a well-resourced source domain and reliably transported to new domains where running a full sweep is infeasible, which requires a fundamental understanding of when and why scaling properties change. We address this by identifying the right invariants: scaling laws are preserved under bijective (information-preserving) transformations of the data and modified in predictable, information-theoretically grounded ways under non-bijective transformations that lower its information resolution $\rho$: a single axis along which a law fit in one domain can be transported to another. We validate this across language, vision, and speech, and demonstrate two cross-domain applications: predicting scaling for language models trained on electronic health records from laws fit on general text, and predicting time-series classification scaling under varying levels of noise injection, recovering the data-scaling exponents to within 3-percent error.


Opening the Black Box of Classifier-Free Guidance via Information Bottleneck

Jiayang Gao ⋅ Tianyi Zheng ⋅ Jiayang Zou ⋅ Fengxiang Yang ⋅ Luyao Fan ⋅ Lv Tang ⋅ Bo Li ⋅ Jia Wang

Classifier-free guidance (CFG) is a standard technique for improving conditional generation in diffusion models, yet its activation schedule is usually selected heuristically. In practice, most existing CFG-based methods rely on a pre-defined activation policy throughout inference, while recent dynamic variants often depend on hand-designed schedules without a principled theoretical explanation of when conditional information is most useful. In order to understand how the informational relevance between conditioning inputs and the score function evolves over the course of generation, we explore when guidance should be activated from an information-theoretic perspective. In this work, we formulate the diffusion trajectory as a time-dependent information bottleneck and show that the effectiveness of conditional information is not uniform throughout sampling, but becomes significantly stronger after a critical stage characterized by a Fisher-information balance condition. Motivated by this insight, we propose a training-free dynamic CFG strategy for inference. Rather than prescribing a fixed schedule in advance, our method constructs a stochastic process from conditional and unconditional score statistics and activates guidance when the estimated guidance gain crosses a prescribed threshold. This yields an adaptive activation rule that aligns condition injection with the underlying generation dynamics, while also recovering interval-based guidance as a limiting special case. Extensive experiments across diverse architectures and tasks, demonstrate that our method is effective, robust, and broadly applicable. Beyond empirical improvements, our framework provides an interpretation of dynamic guidance methods and a principled foundation for inference-time control in conditional diffusion models.


Optimal Learning-Augmented Algorithm for Online Bidding

Changyeol Lee ⋅ Dahoon Lee ⋅ Jongseo Lee ⋅ Yongho Shin ⋅ Changki Yun

Recent advances in machine learning have spurred significant interest in learning-augmented algorithms, particularly for online optimization. A growing body of work has studied online bidding in this framework, aiming to characterize the trade-off between robustness and consistency. While this trade-off is fully understood for deterministic algorithms, a gap between upper and lower bounds remains in the randomized setting. In this paper, we close this gap by presenting a Pareto-optimal randomized learning-augmented algorithm for this problem. Our approach introduces the notion of a bidding profile, a novel framework for representing the distribution over bids generated by an algorithm. We show that any bidding algorithm can be reduced, without loss of generality, to one driven by a bidding profile, and we characterize the optimal profile via a system of delayed differential equations. Finally, we demonstrate the broader applicability of our approach by extending it to the linear search problem, yielding a significant improvement over prior learning-augmented algorithms for linear search.

It has been recently shown that e-processes are sufficient for sequential testing in the following sense: every level-$\alpha$ sequential test can be obtained by thresholding an e-process at $1/\alpha$. However, in the above result, neither does the test have to be asymptotically optimal (in terms of stopping times) nor does the e-process have to be asymptotically log-optimal. It has separately been shown that asymptotically log-optimal e-processes yield asymptotically optimal sequential tests. In this paper, we prove the converse, arguably completing the story: it is possible to aggregate asymptotically optimal sequential tests into asymptotically log-optimal e-processes. This is accomplished by using a new class of WAIT e-processes: those that are Weighted Aggregates of Indicators of Stopping Times that begin at zero, are nondecreasing and increase to infinity under the alternative at the optimal rate. Importantly, the paper discusses several nuances in the varied definitions of asymptotic (log-)optimality.

Hyperparameter prediction is a critical practical bottleneck for model-based image denoisers, ranging from classical TV/TGV variational solvers to modern diffusion-based models such as DiffPIR. While existing learned predictors can achieve near-oracle performance, this approach scales poorly: each new configuration conventionally requires its own oracle-labeled training set, and each label requires an exhaustive denoiser sweep evaluated against clean ground truth. We therefore ask whether oracle supervision collected on source configurations can transfer to target configurations with few or no target oracle labels. We propose HyperDn, a single configuration-conditioned predictor that pools oracle supervision across source configurations and predicts heterogeneous hyperparameters for new denoiser--noise configurations. In a cross-paradigm experiment, HyperDn transfers from relatively cheap TV/TGV variational sources to more expensive diffusion-based DiffPIR. With only $2$ target oracle labels, it reaches $30.23$\,dB, within $0.90$\,dB of oracle, using $1/32$ the target labels that a per-configuration predictor trained from scratch needs to match. Without any target oracle labels, HyperDn also reaches near-oracle PSNR on two unseen mixtures of seen noise types and on transfer from relatively cheap $96\times 96$ source images to $512\times 768$ targets. Together, these results show that expensive oracle supervision for hyperparameter prediction can be transferred from source to new target configurations, reducing the need to rebuild oracle labels for each new denoising configuration.


PAAC: Privacy-Aware Agentic Device-Cloud Collaboration

Liangqi Yuan ⋅ Wenzhi Fang ⋅ Shiqiang Wang ⋅ Christopher Brinton

Large language model (LLM) agents face a structural tension: cloud agents provide strong reasoning but expose user data, while on-device agents preserve privacy at the cost of overall capability. Existing device-cloud designs treat this boundary as a compute split rather than a trust boundary suited to agentic workloads, and existing sanitizers force a choice between policy flexibility and the structural fidelity tool calls require. In this work, we develop PAAC, a privacy-aware agentic framework that aligns planner--executor decomposition with the device-cloud boundary so that role specialization itself becomes the privacy mechanism. The cloud agent reasons over typed placeholder tokens that preserve each sensitive value's reasoning role while discarding its content, while the on-device agent identifies sensitive spans and distills each step's execution outcome into compact key findings. Sanitization confines the on-device LLM to proposing which spans to mask, while a deterministic registry performs all substitution and reversal, keeping actions directly executable on device. On three agentic benchmarks under strict privacy settings, PAAC dominates the Pareto frontier of privacy and accuracy, improving average accuracy by 15-36\% and reducing average leakage by 2-6$\times$ over state-of-the-art device-cloud baselines, with the largest margins on privacy targets outside fixed entity taxonomies. We find consistent improvements on 17 additional benchmarks spanning 10 domains, including math, science, and finance.

Modern text-to-image models produce high-fidelity images but still struggle with compositional prompts that require instance identity, attribute ownership, counting, spatial ordering, and role-sensitive relations. We introduce Panoptic Scene Program Diffusion Transformer (PSP-DiT), a diffusion-transformer architecture that treats a panoptic scene program as a first-class latent variable rather than an external control signal or post-hoc parse. PSP-DiT jointly denoises image latents and scene-program latents through coupled transformer streams, while panoptic grounding and cycle-consistency objectives tie object instances, attributes, relations, and counts to visual support in the generated image. Under matched training and inference settings, PSP-DiT improves over a strong flat-text baseline across GenEval 2, SANEval-Simple, PSG-Score, and DetailMaster, with the largest gains on counting, attribute binding, role-sensitive relations, and long structured prompts. The method preserves image quality, adds modest inference overhead, and remains robust to imperfect scene programs.


Paradoxical noise preference in RNNs

Noah Eckstein ⋅ Manoj Srinivasan

In recurrent neural networks (RNNs) used to model biological neural networks, noise is typically introduced during training to emulate biological variability and regularize learning. The expectation is that removing the noise at test time should preserve or improve performance. Contrary to this intuition, we find that continuous-time recurrent neural networks (CTRNNs) often perform best at a nonzero noise level, often approximately the same level used during training. This noise preference typically arises when noise is injected inside the neural activation function; networks trained with noise injected outside the activation function perform best with zero noise. The phenomenon arises robustly in diverse tasks for large enough training noise including function approximation, maze navigation, 2D path integration, and a multi-task suite from cognitive neuroscience; we also show the phenomenon arising in feedforward neural networks, not just in RNNs. Through analyses of simple function-approximation and single-neuron regulator tasks, we show that the phenomenon stems from noise-induced shifts of fixed points (stationary distributions) in the underlying stochastic dynamics of the RNNs, thereby providing some mechanistic interpretability of the phenomenon. These fixed point shifts are noise-level dependent and bias the network outputs when the noise is removed, degrading performance. Analytical and numerical results show that the bias arises when neural states operate near activation-function nonlinearities, where noise is asymmetrically attenuated, and that performance optimization incentivizes operation near these nonlinearities; such performance incentives exist for networks with noise inside the activation function, but not for networks with noise outside the activation function, explaining why only noise-in networks show preference. Thus, networks can overfit to the stochastic training environment itself rather than just to the input–output data. The phenomenon is distinct from stochastic resonance, wherein nonzero noise enhances signal processing. Our findings reveal that training noise can become an integral part of the computation learned by recurrent networks, with implications for understanding neural population dynamics and for the design of robust artificial RNNs.

Reinforcement-learning-trained agents for retrieval-augmented generation retrieve content (passages, knowledge-graph fragments, or hyperedges) and supervise training with outcome rewards on the final answer. Across these paradigms the reasoning route through retrieved content never leaves the language model's hidden state, conflating declarative knowledge with the procedural knowledge that recurs across queries of the same logical type; content-only methods consequently plateau as reasoning depth grows. We then introduce PatternBloom, an agentic RAG framework that shifts the unit of retrieval from content (what to read) to logic (how to reason). A first RL stage scores the agent's constructed evidence graph by a frozen oracle's information gain through the Information-Density Reward (IDR), a size-normalized mutual-information surrogate. High-reward trajectories are distilled into a Graph Pattern Memory (GPM) of type-abstracted reasoning skeletons, and a second RL stage uses the Pattern-Augmented Reward (PAR) to couple policy to memory, turning procedural reasoning into a learnable, externalized structural prior. A 7B PatternBloom outperforms 27 baselines by +15.0 Avg-OOD EM, with the gain compounding with reasoning depth precisely because procedural knowledge, once distilled, is reused rather than re-derived at every query.

We study fixed-policy evaluation for finite Markov chains that may be reducible and periodic. Classical evaluation methods with gain and bias decomposition are not always diagnostic: the gain records only invariant Ces\`aro averages, while persistent phase-dependent behavior is absorbed into the bias together with genuinely transient effects. We identify the real peripheral invariant subspace $\mathcal{K}(P)$ of the transition matrix $P$ as the source of this ambiguity. Quotienting by $\mathcal{K}(P)$ is the minimal exact quotient that removes all non-decaying modes and makes the remaining dynamics strictly stable. After choosing a gauge projection $\Pi$ with kernel $\mathcal{K}(P)$, the reward admits a unique decomposition $r = g_\Pi^\star + (I-P)v_\Pi^\star$, where $g_\Pi^\star$ is a persistent regime profile and $v_\Pi^\star$ is a gauge-fixed transient component. An exact comparison with classical normalized gain and bias shows that the new pair reallocates the same information so that all persistent modes are represented in $g_\Pi^\star$ and $v_\Pi^\star$ is transient. This decomposition reconstructs finite-horizon returns, recovers statewise average reward, admits a transient-cost interpretation, and yields a stable estimator under a generative model.


Personal Visual Memory from Explicit and Implicit Evidence

Viet Nguyen ⋅ Thao Nguyen ⋅ Vishal Patel ⋅ Yuheng Li

Long-term memory is increasingly important for personalized AI agents, yet existing benchmarks and methods remain largely text-centric. Even when images are included, the user-specific information needed for later questions is typically recoverable from text alone, and most memory systems reduce image turns to generic captions. Yet images often carry personal information that text rarely states---both \emph{explicit} evidence, such as recurring user-associated entities, and \emph{implicit} evidence, such as latent user facts inferred from visual or multimodal cues. We introduce a benchmark for \emph{personal visual memory} that targets both forms of evidence, and propose \textsc{VisualMem}, a hybrid visual--text architecture that augments a text-memory backend with a structured personal visual memory module. Rather than collapsing images into captions, \textsc{VisualMem} uses conversational context to resolve identity, ownership, and durable user facts. Experiments show that \textsc{VisualMem} substantially outperforms prior memory systems on our benchmark while remaining competitive on standard text-memory benchmarks, indicating that personal visual memory is a distinct and important component of long-term memory for personalized AI agents.

Evaluating machine learning in scientific domains requires separating correct predictions from correct reasons under realistic distribution shifts. We introduce PerturbReason, a knowledge‑grounded benchmark for cell‑state--conditioned reasoning about perturbation effects. It tests the models to generate mechanistically faithful explanations and assesses their robustness against complex shifts, such as new cells, unseen perturbations, and cross-modal or combinatorial extrapolations. PerturbReason combines single‑cell genetic and chemical perturbation data across multiple cell lines with knowledge graphs, and dynamically conditions pathways on cell‑specific basal states to avoid generic memorization. Evaluations on state-of-the-art models reveal systematic gaps between predictive accuracy and mechanistic reasoning. Specifically, these models exhibit failure modes largely invisible to standard benchmarks, such as deriving correct answers through flawed logic, ignoring cellular context, and generating directionally inconsistent mechanisms. As a reference probe of the benchmark, we present PerturbRM, a large language model trained to align outcome predictions with context-specific regulatory reasoning. PerturbReason thus provides a rigorous diagnostic benchmark for studying and improving faithful, generalizable reasoning in data-rich scientific systems.

Text-to-video (T2V) generation models can produce visually realistic videos that nevertheless violate the physical consequences implied by the prompt and scene. Existing benchmarks largely rely on holistic plausibility scores or broad quality dimensions, leaving it unclear which physical behavior is being tested and why a model fails. We introduce PhyMetric, a scene-level diagnostic benchmark that recasts physical plausibility evaluation as verifiable, scene-grounded question answering. PhyMetric contains 1115 scenes spanning 50 physical categories across six domains, paired with 8920 balanced binary questions over 12 evaluation dimensions. Each scene is associated with eight targeted yes/no questions and a structured physical prior specifying the governing law, expected temporal sequence, observable signatures, and common failure modes. At evaluation time, a vision-language model (VLM) judge answers each question from question-aware sampled frames conditioned on the prior, and correctness is hierarchically aggregated into scene, dimension, physical category, and model scores. Experiments on four open-source T2V models and three VLM judges show that scene-specific binary QA is a necessary component of the protocol: when each targeted question is replaced by a single generic plausibility query, evaluator accuracy drops to the chance level. Structured physical priors offer a complementary signal that improves human alignment of the resulting scores, with effects that vary across VLM judges of different capacity. These findings establish PhyMetric as a benchmark that turns physical plausibility evaluation from opaque holistic scoring into interpretable, scene-level inspection.


Physical AI Smart Spaces: A Large-Scale Benchmark for Multi-Camera 3D Perception in Smart Spaces

Yuxing Wang ⋅ Yizhou Wang ⋅ Anqi Li ⋅ Shuo Wang ⋅ Sameer Satish Pusegaonkar ⋅ Haoquan Liang ⋅ Jiajun Li ⋅ Shenxin Jiang ⋅ Jianhe Yuan ⋅ Shangru Li ⋅ Tongwei Dai ⋅ Zihao Chen ⋅ David Anastasiu ⋅ Zheng (Thomas) Tang ⋅ Sujit Biswas ⋅ Xunlei Wu

Physical AI Smart Spaces is, to the best of our knowledge, the first benchmark to simultaneously provide large-scale, multi-class, and multi-camera 3D perception data for indoor smart spaces. It contains over 280 hours of synchronized 1080p footage captured by nearly 1,800 cameras in warehouses, hospitals, retail venues, and similar settings, together with automatic annotations for multi-camera identities, 2D bounding boxes, 3D bounding boxes, camera calibration, and depth where available. The benchmark spans Isaac Sim synthetic generation, Cosmos Transfer appearance augmentation, and real-world Sim2Real evaluation. For the real-world target, we include two warehouse deployments with time-synchronized streams, automatic VGGT-based calibration, and a 3D labeling interface that projects world-frame 3D boxes into each view for cross-camera verification. We describe the dataset scope, annotation and calibration schema, generation workflow, benchmark protocols, and official evaluation system, which standardizes submission format, and leaderboard reporting. A central contribution is a 3D instantiation of Higher Order Tracking Accuracy (HOTA), extending the usual 2D box-based tracking evaluation to 3D locations and 3D boxes. We further report empirical baselines from the AI City Challenge leaderboards, showing how methods evolve from person-only 3D location tracking to multi-class 3D box tracking under realistic smart-space constraints. The release is available at https://huggingface.co/datasets/nvidia/PhysicalAI-SmartSpaces.


PixelDense: Dense Prediction as Representation Alignment for Pixel Diffusion

Lehan Yang ⋅ Daiqing Qi ⋅ Wenhao Zhang ⋅ Avery Li ⋅ Yiqing Yang ⋅ Yifan Li ⋅ Yu Kong ⋅ Haitian Zheng ⋅ Zhifei Zhang ⋅ Zhe Lin ⋅ Varun Jampani ⋅ Sheng Li

Representation alignment (REPA) accelerates diffusion transformer training, but its alignment targets are almost exclusively semantic encoders such as DINOv2 and CLIP. Recent analysis points to spatial structure, not global semantics, as the carrier of the alignment effect, yet dense-prediction foundation models trained to predict that structure remain overlooked as REPA targets. In pixel-space diffusion, SAM2, Depth Anything v2, and Metric3D v2 each outperform the DINOv2-only GenEval baseline, with the two geometric teachers leading the segmentation teacher. A flat sum of all four teachers, however, lands below the best single geometric teacher, semantic teachers reward viewpoint invariance and instance identity while geometric teachers reward metric layout and surface orientation, and forcing both kinds of gradients through one denoiser projection collapses them into a shared subspace. We introduce PixelDense, which routes DINOv2 and SAM2 through a semantic projection stream, routes Depth Anything v2 and Metric3D v2 through a geometric projection stream, and adds a weight-space orthogonality penalty that keeps the two streams in disjoint subspaces. All four teachers are frozen during training and dropped at inference. Applied to PixelGen and DeCo with a single recipe, PixelDense improves GenEval, DPG-Bench, and HPS v2.1, raises PixelGen-XXL's GenEval Overall from $0.7927$ to $0.8093$, and beats every single-teacher and unfactored multi-teacher variant. In partial-noise reconstruction, independent panoptic, depth, and surface-normal probes show up to 53.1% PQ gain and 36.0% depth AbsRel reduction at $\tau=0.5$ across COCO and Flickr30K. From random initialization, PixelDense also reaches the baseline's peak GenEval $1.23\times$ faster.


PM1: A Multimodal Foundation Model for Genomes, Phenotypes, and Images at Biobank Scale

Christophe Thomassin ⋅ Marçal Comajoan Cara ⋅ Margarita Geleta ⋅ David Bonet ⋅ Benet O Sabat ⋅ Daniel Mas Montserrat ⋅ Alexander Ioannidis

Precision medicine relies on integrating diverse patient data, including genetic variants, clinical phenotypes, and medical imaging, to inform personalized prevention, diagnosis, and treatment. However, learning general patient-level representations from these heterogeneous modalities remains challenging due to their high dimensionality and pervasive missingness. We propose PM1, a multimodal foundation model that learns meaningful patient-level representations across genetic, clinical, and imaging data from 337,129 individuals in the UK Biobank, with a focus on retinal imaging and ophthalmic traits. To address modality dimensionality, PM1 adopts an intermediate fusion architecture with modality-specific encoders and a shared Transformer backbone, where a novel pathway-based genotype encoder enables the incorporation of over 600,000 genetic variants by reducing dimensionality while preserving trait-relevant variation. In addition, a contrastive loss invariant to modality missingness allows PM1 to learn robust patient-level representations despite only 6\% of UK Biobank participants having complete modality coverage. By partitioning modalities into interpretable tokens, PM1 enables interpretability methods to investigate known and potentially novel associations across genetic, clinical, and imaging features, and their interactions. Evaluated across downstream tasks, including phenotype prediction, conditioned image generation, and attribution analyses, PM1 demonstrates the value of multimodal patient representations, improving predictive performance over unimodal baselines and competing methods while enabling interpretable analyses in settings with heterogeneous and incomplete patient data.

Protein-conditioned 3D molecule generation is a central challenge in structure-based drug design, requiring a delicate balance between target compatibility, molecular properties, and physical geometry. While diffusion-based approaches have shown promise, strengthening the conditioning signal can still reduce physical plausibility, and existing models remain sensitive to how pocket information is presented during training and sampling. We propose PocketVE, a protein-pocket-conditioned variance-exploding (VE) diffusion framework that couples stable coordinate denoising with inference-time conditional control. Specifically, PocketVE combines an EDM-style training and sampling setup for stable 3D denoising, classifier-free guidance for multi-property steering without external property classifiers, and adaptive protein perturbation as a training-time pocket regularizer. Evaluated on CrossDocked2020 under the GenBench3D protocol, PocketVE improves 3D-Validity from 58.6 to 80.6 and reduces strain energy from 457.4 to 127.9 relative to the TAGMol baseline, while improving median Vina Score, QED, and SA under moderate guidance. A guidance-scale study further shows that moderate guidance gives the best balance between target-related objectives and geometric quality, whereas overly strong guidance pushes the sampler toward geometric degradation. Overall, the results show that target-aware molecular generation benefits from treating geometric stability and conditional steerability as coupled design goals.

While there is an extensive body of research analyzing policy gradient methods for discounted cumulative-reward MDPs, prior work on policy gradient methods for average-reward MDPs has been limited, with most existing results restricted to ergodic or unichain settings. In this work, we first establish a policy gradient theorem for average-reward multichain MDPs based on the invariance of the classification of recurrent and transient states. Building on this foundation, we develop refined analyses and obtain a collection of convergence and sample-complexity results that advance the understanding of this setting. In particular, we show that the proposed $\alpha$-clipped policy mirror ascent algorithm attains an $\epsilon$-optimal policy with respect to positive policies.

We study reinforcement learning in hybrid discrete–continuous action spaces, such as settings where the discrete component selects a regime (or index) and the continuous component optimizes within it — a structure common in robotics, control, and operations problems. Standard model-free policy gradient methods rely on score-function (SF) estimators and suffer from severe credit-assignment issues in high-dimensional settings, leading to poor gradient quality. On the other hand, differentiable simulation largely sidesteps these issues by backpropagating through a simulator, but the presence of discrete actions or non-smooth dynamics yields biased or uninformative gradients. To address this, we propose Hybrid Policy Optimization (HPO), which backpropagates through the simulator wherever smoothness permits, using a mixed gradient estimator that combines pathwise and SF gradients while maintaining unbiasedness. We also show how problems with action discontinuities can be reformulated in hybrid form, further broadening its applicability. Empirically, HPO substantially outperforms PPO on inventory control and switched linear-quadratic regulator problems, with performance gaps increasing as the continuous action dimension grows. Finally, we characterize the structure of the mixed gradient, showing that its cross term — which captures how continuous actions influence future discrete decisions — becomes negligible near a discrete best response, thereby enabling approximate decentralized updates of the continuous and discrete components and reducing variance near optimality.


PolyMind: Exploring Width Scaling for Reflective Reasoning in Language Agents

Heng Zhang ⋅ Chengyu Zhou ⋅ Jiajun Wu ⋅ Estella Liu ⋅ Liheng Zhang ⋅ Yueqi Guo ⋅ Rui Liu ⋅ JiaHao Hong ⋅ Xuanxun Lian ⋅ Jinpeng Lu ⋅ Jin Huang

Reflection-based language agents improve reasoning by using feedback to revise previous attempts across multiple trials. Most existing reflection frameworks scale this process by \emph{depth}, allocating additional test-time compute to longer serial reflection-retry chains. However, we observe that serial reflection often saturates after rapid early gains. Under controlled reflection budgets, increasing depth from $5$ to $10$ cycles brings only marginal improvement, yet isolated reflective workers still contain useful revision signals beyond the saturated chain. This raises a central question: can reflective reasoning scale by \emph{width} rather than only by depth? A direct answer is non-trivial, since naive parallel reflection often inspects similar aspects of the same failed attempt and produces long or noisy revision contexts after aggregation. We propose \ourmethod, a reflective width scaling framework that learns to allocate complementary reflective scopes, execute workers in isolated contexts, and compress parallel feedback into a structured PolyBrief. \ourmethod trains the planner, workers, and reducer with verifiable outcome feedback and a role-balanced objective that prevents wider rollouts from dominating optimization. Across code generation, mathematical reasoning, and multi-hop question answering, \ourmethod consistently improves performance under matched reflection budgets, outperforming the strongest depth-reflection baseline by an average of $6.1\%$. Further analyses show that \ourmethod closes about $80\%$ of the oracle-naive width gap, reduces redundant worker outputs, and improves revision success with fewer tokens. These results suggest that reflection should be understood not only as a longer retry process, but also as a depth-width compute allocation problem.


PortPy: A Benchmark for AI and Optimization in Cancer Radiotherapy Planning

Gourav Jhanwar ⋅ Mojtaba Tefagh ⋅ Qijie Huang ⋅ Alireza Golkarieh ⋅ Linda Hong ⋅ Hai Pham ⋅ Vicki T Taasti ⋅ Seppo Tuomaala ⋅ Anthony Magliari ⋅ Michael Folkerts ⋅ Ying Zhou ⋅ Jie Yang ⋅ Saad Nadeem ⋅ Masoud Zarepisheh

External beam radiotherapy is a cornerstone of cancer treatment, used in more than half of all cancer patients. It delivers high-energy radiation beams to destroy tumors while minimizing radiation dose to surrounding healthy tissues. Treatment planning requires optimizing machine parameters, including the shapes and intensities of radiation beams, based on each patient’s unique anatomy. This gives rise to large-scale, multi-criteria, and often combinatorial optimization problems that must be solved separately for each patient under strict clinical time constraints. Although recent advances in AI/ML show promise for improving many planning and decision-making tasks, progress in this area has been limited by the lack of standardized, publicly available benchmark datasets and baseline algorithms. To address this gap, we introduce PortPy (Planning and Optimization for Radiation Therapy in Python), an open-source benchmark dataset derived from real patient data, together with baseline algorithms and a Python-based framework for data processing, simulation, and standardized evaluation. Several of the provided baseline methods have been clinically validated and deployed at our institution to support automated radiotherapy planning for cancer patients. By providing a common benchmark and open research platform, PortPy aims to accelerate progress in radiotherapy treatment planning and promote collaboration across the AI/ML, mathematical optimization, medical physics, and radiation oncology communities. Toolkit: https://github.com/PortPy-Project/PortPy Dataset: https://huggingface.co/datasets/PortPy-Project/PortPy_Dataset


Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning

Chih-Hsuan Yang ⋅ Jingyan Jiang ⋅ Vikram Vasudevan ⋅ Cheng-Hau Yang ⋅ Huihuo Zheng ⋅ Benoit Cote ⋅ Le Chen ⋅ Eliu Huerta ⋅ Venkatram Vishwanath ⋅ Ian Foster ⋅ Rajeev Thakur

Many recent math- and science-oriented agent systems adopt hierarchical designs with specialized reviewer roles, motivated by the idea that routing critique through a dedicated review stage should help turn wrong candidates into correct ones on hard problems. We study that expectation on 4,181 verifier-grounded Omni-MATH problems, a hard ten-tier benchmark with enough headroom to separate protocols, using matched gpt-oss-120b actors as the primary actor family. On tiers 1-2, multi-agent collaboration adds at most about 2 percentage points over the matched single-agent baseline. From tier 4 onward, the gains open sharply, reaching about 10 to 20 percentage points on tiers 6-9. In that regime, broadcast-style peer discussion attains higher final accuracy than a planner-executor-reviewer pipeline (PER). We use that divergence as a starting observation and ask when reviewer quality translates into effective solver updates. On this benchmark, the PER-broadcast accuracy gap is not accounted for by reviewer precision alone. Here reviewer precision asks how often a reviewer warning truly points to a real error. PER's reviewer has higher precision (0.861 vs. 0.644), yet correct critique is much less likely to change the next candidate the protocol carries forward, and reviewer-guided repair is correspondingly lower. These results show that reviewer detection quality and critique uptake are empirically separable. Within matched PER interventions, forcing explicit acknowledgment lowers final accuracy instead of improving follow-through, while placing reviewer guidance directly in the solver's working context partially improves follow-through without closing the gap. Taken together, these interventions point in the same direction: critique appears more likely to be acted upon when it is presented more directly in the solver's working context, although this evidence is directional rather than causal. Under reviewer-centric evaluation, a system can look strong at spotting errors yet still fail to solve more problems if the protocol does not act on those critiques.

Self-distillation is typically studied when the student is retrained on the teacher’s original training inputs. In many realistic deployments, however, the labeled data are unavailable after training, and one only has access to the trained predictor and fresh unlabeled covariates. We study distillation in this prediction-only regime through a fresh-$X$ prediction-mixing scheme: a pure-distilled student is trained on pseudo-labeled fresh features, and the final predictor is formed by an affine combination of the teacher and student predictions with the optimal unconstrained mixing weight. For ridge regression under proportional asymptotics, we derive deterministic equivalents for the optimally mixed risk under general anisotropic covariance and deterministic signal, and show that optimal mixing strictly improves upon the teacher for almost every pair of regularization levels. In the isotropic case, we uncover a sharp and counterintuitive phenomenon: the optimally mixed risk is generically non-monotone in the amount of unlabeled data. To tune the optimal mixing weight, we show that a small independent labeled calibration set suffices for consistent one-shot estimation at low computational cost, since no additional retraining is required, whereas such tuning is impossible when using only teacher’s outputs on unlabeled data. Finally, to extend beyond squared loss, we analyze logistic regression for binary classification and show that prediction mixing can improve over both the teacher and the pure-distilled classifier.


Predictive 4D Generation with Latent State Machines

Yue Gao ⋅ Chen Geng ⋅ Hanzhe Hu ⋅ Fangyun Wei ⋅ Jianfeng Xiang ⋅ Jiaolong Yang ⋅ Jiajun Wu

What is the right representation for 4D (3D + time) generation of graphics assets conditioned on their current states? In this setting, an ideal 4D representation should support flexible geometric structure, empower long-horizon modeling, and utilize training data efficiently. We study this problem by considering a sufficiently challenging setting to stress test those properties: learning the generative distribution of the entire 4D lifespan of a plant, conditioned on a snapshot of its growth. We curate a large-scale dataset and benchmark for this setting, on which most existing 4D representations and accompanying generative models fail to deliver promising results. We analyze the pitfalls of existing representations and point out that the capacity of existing models is bounded by their representations: they are mostly augmenting a temporal dimension to an existing geometric representation, without a semantically-meaningful low-dimensional state representation. Instead, we argue for a framework for designing 4D representations by separately considering 3D state representations and temporal state transitions. We propose an instantiation of this framework using latent spaces of 3D foundation models and a transformer-based state transition model, effectively forming a finite state machine on a latent 3D space. Therefore, we name this instantiation a latent state machine and compare it against existing generative models on the proposed benchmark. Our evaluation shows that our model achieves state-of-the-art performance in expressiveness, data efficiency, and rendering quality, surpassing existing methods by a large margin.


Progressive Risk Estimation for Accident Anticipation

Samet Hicsonmez ⋅ Eray Çakar ⋅ Nermin Samet ⋅ Fatma Guney

Accident anticipation aims to recognize anomalous driving cues before a crash while avoiding false alarms during normal driving. Existing approaches typically formulate this task as binary classification, focusing on whether an accident will occur rather than when it will occur. We propose PRE-ACT, a framework that models accident risk as a continuously evolving signal that increases as the crash approaches. By explicitly enforcing temporal ordering and distance-to-accident awareness, our method progressively raises risk while suppressing premature alarms, leading to significant improvements on MM-AU subsets and Nexar. We further introduce a Separation Score to evaluate the global behavior of predicted risk curves beyond local temporal windows. Code and visualizations are available.


Projected Neural Additive Models as Universal Approximators

Nhon N Phan ⋅ Qiang Du ⋅ John D Clayton ⋅ WaiChing Sun

Neural additive models (NAMs) offer interpretability but lack expressivity due to their rigid additive structure, whereas multi-layer perceptrons (MLPs) achieve universal approximation at the cost of transparency. We establish the first theoretical foundation for projected neural additive models (PNAMs), which augment NAMs with a learnable linear transformation $\boldsymbol{T}$ to capture cross-variable interactions. While standard NAMs cannot approximate even the simplest interaction term $\chi_1 \chi_2$ because their cross-derivatives are identically zero, we prove that PNAMs are universal approximators and characterize the minimum number of projections required for exact polynomial representation. Furthermore, we derive an approximation rate of $\mathcal{O}(M^{-s / (N - 1)})$ for $\mathcal{C}^s$ functions (i.e., functions with $s$ continuous derivatives), where $M$ denotes the number of projections and $N$ the input dimension; this rate is known to be optimal for ridge function approximation. To recover interpretability from the projected inputs, we introduce structured sparsity regularizers and post-hoc symbolic regression techniques that enable global input ranking, parameter pruning, and the conversion of learned bases into compact analytical expressions. Experiments on knot theory invariants, MNIST, and phase field fracture mechanics demonstrate that PNAMs match or surpass MLPs in accuracy while discovering salient features and yielding symbolic models with orders of magnitude fewer parameters.

Deep learning has driven substantial progress in modeling fMRI responses, yet most existing approaches are optimized for a single objective---such as neural encoding, stimulus decoding, or post hoc interpretation---making it difficult to obtain a coherent account of neural representations. We introduce a unified framework that connects vision foundation models to brain activity while supporting neural analysis from multiple complementary perspectives. Our approach replaces black-box features with a sparse, interpretable representation composed of visual prototypes, and links these prototypes to voxel responses through a linear readout. This design makes the model transparent at both levels: prototype activations expose the visual concepts present in the stimulus, while the learned linear weights characterize each voxel's selectivity. Using this interface, we identify which visual concepts drive neural responses and where they appear in the image, reveal cortical organization through voxel tuning in prototype space, and perform controlled stimulus synthesis to probe cortical selectivity through targeted interventions. Together, these results establish prototype-based representations as a unified and interpretable foundation for studying visual representations in the human brain.


Provable Quantization with Randomized Hadamard Transform

Ying Feng ⋅ Piotr Indyk ⋅ Michael Kapralov ⋅ Dmitrii Krachun ⋅ Boris Prokhorov

Vector quantization via random projection followed by scalar quantization is a fundamental primitive in machine learning, with applications ranging from similarity search to federated learning and KV compression. While dense random rotations yield clean theoretical guarantees, they require $\Theta(d^2)$ time. The randomized Hadamard transform $HD$ reduces this cost to $O(d \log d)$, but its discrete structure complicates analysis and leads to weaker or purely empirical compression guarantees. In this work, we study a variant of this approach: \emph{dithered quantization} with a single randomized Hadamard transform. Specifically, the quantizer applies $HD$ to the input vector and subtracts a random scalar offset before quantizing, injecting additional randomness at negligible cost. This approach is used by popular quantizers such as RaBitQ, and is a slight modification of others, including TurboQuant and EDEN. We prove that this approach provides mean squared error bounds that asymptotically match those achievable with truly random rotation matrices. For example, we prove a dithered version of TurboQuant achieves mean squared error $\bigl(\pi\sqrt{3}/2 + o(1)\bigr) \cdot 4^{-b}$ at $b$ bits per coordinate, where the $o(1)$ term vanishes uniformly over all unit vectors and all dimensions as the number of quantization levels grows.


Provable Test-Time Scaling for Beam Search in LLM Reasoning

Qijia He ⋅ Yu Huang ⋅ Yuan Cheng ⋅ Yuxin Chen ⋅ Yingbin Liang

Beam-search–based test-time methods provide an effective way to improve large language model (LLM) performance on long-horizon generation by pruning invalid reasoning paths early, leading to significantly improved reasoning efficiency and more favorable test-time cost scaling. Despite strong empirical success, the theoretical understanding of beam search remains limited. In this paper, we study the test-time compute guarantee of the commonly used beam search framework that uses the model's internal log-likelihood for intermediate scoring, while relying on an external reward model only after a complete response is generated. We first establish a lower bound for vanilla beam search, showing that at least $\Omega(C^\star(x)^2)$ samples are required for the optimal response to survive, where $C^\star(x)$ is the token-level coverage coefficient for prompt $x$. This motivates our modified confidence-filtered beam search (CF-Beam), which provably reduces the sample size by a factor of $C^\star(x)$. We then show that the regret of CF-Beam is upper-bounded by the probability of rare failure events and the reward estimation error scaled by a path-level coverage coefficient, where the rare-failure term vanishes as per-step sampling increases. Our results highlight a fundamental advantage of beam search over sequence-level inference methods such as Best-of-N and Best-of-Majority. While the guarantees of these approaches typically involve coverage coefficients that grow exponentially with the horizon $L$, CF-Beam controls the dominant search-induced term through a token-level coverage coefficient that scales polynomially with $L$. Our numerical experiments further confirm that beam search is more robust on hard instances and under increasing reasoning horizons.


Prudent-Banker: No Extra Fees for Baseline Safety in Adversarial Bandits With and Without Delays

Ting Hu ⋅ Luanda Cai ⋅ Emmanouil-Vasileios Vlatakis-Gkaragkounis

We study adversarial multi-armed bandits with and without delayed feedback under a **safety-aware** goal: achieving minimax-optimal worst-case regret while keeping nearly **constant** regret relative to a designated "safe" baseline policy. Existing approaches can balance this trade-off with immediate feedback for smooth comparators, but arbitrary delays can mistime transitions between **conservatism** and **exploration**, endangering the safety guarantee. To bridge this gap, we propose PRUDENT-BANKER, a novel algorithm that combines a delay-adapted variant of Online Mirror Descent with a modified phased-aggression mechanism. Its key technical contribution is a delay-calibrated restart threshold that rigorously accounts for the worst-case distortion induced by unobserved feedback and reliably detects comparator suboptimality. We also establish new lower bounds for safety-constrained adversarial delayed bandits, showing that the regret guarantees of PRUDENT-BANKER are unimprovable, up to logarithmic factors, under the baseline-safety requirement. To the best of our knowledge, PRUDENT-BANKER is the first algorithm to achieve the optimal safety–robustness trade-off: pseudo-regret $\tilde{O}(\sqrt{T}+\sqrt{D})$ together with $\tilde{O}(1)$ regret against the safe comparator, both with and without delays. Experiments across diverse delay distributions show that, unlike standard delay-robust baselines, PRUDENT-BANKER effectively balances safety and learning.


QueryStop: Dynamic Stop Signals for Efficient Streaming Inference

Michelle Shu ⋅ Jingyao Li ⋅ Pengguang Chen ⋅ Shu Liu ⋅ Ramin Zabih ⋅ Jiaya Jia

As language models are increasingly deployed in long-context and streaming workflows, deciding when to stop context processing is central to inference efficiency. While recent research has shown that a subset of attention heads encodes sufficiency signals, these markers are typically detected via static, task-agnostic global features. In this paper, we demonstrate that stop signals are inherently query-relevant and introduce \textsc{QueryStop}, an attention-aggregation method that systematically identifies query-focused attention heads whose hidden features effectively align with the input query. Driven primarily by query-focused heads, this tailored ensemble yields high-fidelity representations that significantly improve sufficiency prediction. Across four benchmarks on eleven datasets, our method improves sufficiency stop prediction by 3.52\% F1 on average. Under comparable token reduction, it raises Sufficiency Reach Rate (SRR), our proposed stopping-adequacy metric, from 87.12\% to 94.73\% and yields a 7.09\% average gain in downstream task accuracy over the prior baseline. Collectively, these results show that \textsc{QueryStop} effectively eliminates unnecessary context processing while simultaneously improving downstream answer quality.


RaZeR: Pushing the Limits of NVFP4 Quantization with Redundant Zero Remapping

Yuzong Chen ⋅ Xilai Dai ⋅ Jake Hyun ⋅ Chi-Chih Chang ⋅ Wonsuk Jang ⋅ Yuheng Wu ⋅ Thierry Tambe ⋅ Jae-sun Seo ⋅ Mohamed Abdelfattah

The recently introduced 4-bit floating-point format, NVFP4, demonstrates remarkable performance and memory benefits for quantized large language model (LLM) inference. However, we observe two types of redundancy in the existing NVFP4 encoding: (1) The FP4 format naturally exposes an unused quantization value due to its sign-magnitude representation that contains both positive and negative zeros. (2) The FP8 block scaling factor, with 4-bit exponent and 3-bit mantissa, contains an unused sign bit given that the scaling factor is always positive. Additionally, we find that LLM weights are more tolerant to a lower-precision block scaling factor, such as 6 bits with 3-bit exponent and 3-bit mantissa. Based on these observations, we propose Redundant Zero Remapping (RaZeR), an enhanced numerical format that pushes the limits of NVFP4 for more accurate LLM quantization under the same memory footprint. RaZeR leverages the unused bits in the block scaling factor to adaptively remap the negative FP4 zero to a set of pre-defined special values, which maximally utilizes the NVFP4 encoding and better fits the LLM tensor distribution. Extensive experiments validate RaZeR’s superior performance for 4-bit LLM quantization. For example, RaZeR reduces the average perplexity loss of NVFP4 by 29.4% and 32.3% under weight-only and weight-activation quantization, respectively.


ReaLM: A Unified Red-Teaming Benchmark for Physical-World VLMs

Yifei Zhao ⋅ Qian Lou ⋅ Mengxin Zheng

Vision-language models (VLMs) are increasingly used as perception-reasoning backbones for embodied intelligence in safety-critical physical systems, where perception or reasoning errors can lead to unsafe decisions or actions. Although many red-teaming methods have been developed to probe VLM vulnerabilities, their evaluations remain fragmented across datasets, metrics, and threat models, making direct comparison difficult and obscuring whether observed differences arise from stronger attacks, more vulnerable models, or incompatible evaluation settings. Existing chatbot-centric red-teaming benchmarks mainly standardize jailbreak and content-safety evaluation, but they do not systematically capture physically grounded functional failures or cover red-teaming methods that target physical-world VLMs. This raises the key challenge of comparing diverse attack methods under a unified protocol while targeting the same scenario-specific failures. We introduce REALM, to our knowledge the first unified red-teaming benchmark for physical-world VLMs. REALM integrates 12 red-teaming methods, 3 model-agnostic defenses, and 13 VLMs under a practical black-box threat model with shared datasets and metrics. To align adversarial objectives across attack families, REALM introduces an agentic target-generation pipeline that constructs shared, scenario-specific, and physically grounded attack objectives for each scene, enabling fair comparison of diverse red-teaming methods under aligned adversarial goals. Our evaluation shows that text and typographic injection attacks induce the most failures, multimodal co-optimization yields the strongest visual-perturbation transfer, single-pass attacks approach iterative methods at much lower cost, and model scale alone does not confer adversarial robustness. Code is available at https://anonymous.4open.science/r/REALM-762F.

Human vision flexibly extracts part-whole hierarchies from visual scenes, but representing such structures remains a key challenge for neural networks. Inspired by the neural syntax hypothesis in neuroscience, we propose a framework for representing hierarchical part-whole relationships through nested neuronal coherence, characterized by its continuous and distributed nature. Nested neuronal coherence refers to dynamical states where distributed activations are temporally correlated to form cell assembly sequences across timescales, argued to represent neural syntax. Building on this framework, we develop a cortical-inspired hybrid model, Composer, which dynamically achieves emergent nestedness when given images. To evaluate the emergent hierarchy, we create four synthetic datasets and three quantitative metrics, demonstrating the model’s ability to parse scenes of varying complexity. Overall, our work advances a systematic paradigm spanning representation, implementation, and evaluation toward building human-like vision in neural networks.


Rethinking Language Model Scaling under Transferable Hypersphere Optimization

Liliang Ren ⋅ Yang Liu ⋅ yelong shen ⋅ Weizhu Chen

Scaling laws for large language models depend critically on the optimizer and parameterization. Existing hyperparameter transfer laws are mainly developed for first-order optimizers, and they do not structurally prevent training instability at scale. Recent hypersphere optimization methods constrain weight matrices to a fixed-norm hypersphere, offering a promising alternative for more stable scaling. We introduce HyperP (Hypersphere Parameterization), the first framework for transferring optimal learning rates across model width, depth, training tokens, and Mixture-of-Experts (MoE) granularity under the Frobenius-sphere constraint with the Muon optimizer. We prove that weight decay is a first-order no-op on the Frobenius sphere, show that Depth-$\mu$P remains necessary, and find that the optimal learning rate follows the same data-scaling power law with the ``magic exponent'' 0.32 previously observed for AdamW. A single base learning rate tuned at the smallest scale transfers across all compute budgets under HyperP, yielding $1.58\times$ compute efficiency over a strong Muon baseline at $6\times10^{21}$ FLOPs. Moreover, HyperP delivers transferable stability: all monitored instability indicators, including $Z$-values, output RMS, and activation outliers, remain bounded and non-increasing under training FLOPs scaling. We also propose SqrtGate, an MoE gating mechanism derived from the hypersphere constraint that preserves output RMS across MoE granularities for improved granularity scaling, and show that hypersphere optimization enables substantially larger auxiliary load-balancing weights, yielding both strong performance and good expert balance.


Rethinking Visual Attribution for Chest X-ray Reasoning in Large Vision Language Models

Guangzhi Xiong ⋅ Qiao Jin ⋅ Sanchit Sinha ⋅ Zhiyong Lu ⋅ Aidong Zhang

Large Vision Language Models (LVLMs) show promise in medical applications, but their inability to faithfully ground responses in visual evidence raises serious concerns about clinical trustworthiness. While visual attribution methods are widely used to explain LVLM predictions, whether these explanations actually reflect the visual evidence underlying the model's decision is largely unverified, since ground-truth annotations for internal model reasoning are typically unavailable. We address this question for chest X-ray (CXR) reasoning by developing a causal evaluation framework that retains only CXR-VQA samples for which the expert-annotated region is verified, via counterfactual editing, to be causally responsible for the model's prediction. Using this framework across 11 attribution methods, six open-source LVLMs, and two output modes (direct answer and step-by-step reasoning), we find that existing attribution methods often fail to identify the evidence used by LVLMs. To address this failure, we propose MedFocus, a concept-based attribution method that localizes clinically meaningful anatomical regions via unbalanced optimal transport and measures their causal effect on model outputs through targeted interventions. MedFocus produces spatial, concept-level, and token-level attributions and substantially outperforms prior methods, taking a step toward more trustworthy attribution for medical LVLMs.


ReToken: One Token to Improve Vision–Language Models for Visual Retrieval

Yao Xiao ⋅ Reuben Tan ⋅ Zhen Zhu ⋅ Yuqun Wu ⋅ Jianfeng Gao ⋅ Derek Hoiem

Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints. We present ReToken, a single learnable embedding trained as an explicit retrieval target that selects a sparse set of query-relevant visual tokens from a pre-filled visual KV cache. Trained on only a small image-QA dataset, ReToken yields consistent gains across image and video benchmarks: on Visual Haystacks it improves Qwen3VL-8B by 12 points (>20\% relative), and on LVBench it transfers zero-shot to long video for an 8.5-point gain. Thanks to its lightweight design, both training and long-video inference fit on a single H100. We will release our code.


Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure

Max Lamparth ⋅ Daniel Fein ⋅ Andreas Haupt ⋅ Marcel Hussing ⋅ Mykel J Kochenderfer

Single-axis mitigations of reward-model biases (e.g., reducing reliance of the proxy reward on length, sycophancy, or style) can rotate optimization pressure onto correlated proxies rather than eliminate it, a failure mode we call reward bias substitution. We formalize the underlying measurement-vs-optimization gap between the audit distributions where mitigations are validated and the policy distributions where optimization realizes their effects. We introduce a taxonomy, instantiated in closed form, classifying single-axis mitigation outcomes into successful mitigation, bias substitution, overcorrection, silent non-op, and audit-distribution sensitivity. We prove that single-axis mitigation methods cannot be validated by audit-distribution-only evaluation: successful mitigation, bias substitution, and overcorrection produce structurally identical observables under ranking accuracy and win-rate scoring, regardless of benchmarks richness. Augmenting evaluation with policy-induced distributions provably closes the gap and we give actionable prescriptions for mitigation methods and benchmarks. Across published preference-learning mitigation work, we identify bias substitution regimes in results not previously connected to a unified failure mode. Our experiments also show that a published length-debiasing operator zeros pooled reward–length correlation but flips sign within-prompt on three of four SOTA reward models with true reward degrading on two, and that length–sycophancy coupling reverses under human–LLM judge disagreement across eight model families.

We introduce a unified framework that reinterprets first order optimization on a smooth parameter space as a closed loop controlled dynamical system on a Riemannian manifold. Within this framework, optimizers are specified by a quadruple $(M, g, \Phi, \eta)$ consisting of a Riemannian metric, an internal state transition map, a direction field generator, and a lifting parameter. Classical and modern methods including SGD, Adam, AdamW, Lion and Shampoo are recovered as specific instantiations of this quadruple. The central structural object of the framework is a time indexed family of target submanifolds in the extended state space of parameters, velocities, and internal memory, which we call the Normally Attracting Invariant Manifold (NAIM) family. These submanifolds organize training dynamics into two timescales: a fast contraction of the velocity state onto the current target submanifold, followed by a slow descent along it. Using Riemannian backstepping, we construct a strict Lyapunov function on the extended state space and prove Uniform Ultimate Boundedness (UUB) of the optimization trajectory under a descent alignment assumption and bounded stochastic disturbances. Under the additional Polyak-Lojasiewicz (PL) condition, the Lyapunov function converges linearly to a bounded neighborhood of the optimum, defined as the drift of the target direction field between consecutive iterates after parallel transport. To validate the framework's generativity, we derive three concrete instantiations of the quadruple, each exercising a distinct design axis, and show on geometric diagnostics, image classification benchmarks, and language modeling that their empirical behavior matches the theoretical predictions. The framework casts optimizer design as controller synthesis, complementing heuristic perspectives by providing a constructive template with provable guarantees.

Specialized agents collaborate by transferring memory: retrieved records, generated summaries, tool traces, and long-running state. In regulated domains this is a governance problem, not a retrieval problem: the receiver's task matters, but so does whether it was authorized to receive each item and whether the handoff exposed more than necessary. We formalize minimum-necessary context selection as a constrained submodular subset-selection problem with an additive exposure cost and a distributional task-success risk bound. Our algorithm, ScopeSelect, first hard-masks items whose layer or scope the receiver contract does not admit, then runs a cost-benefit greedy selector against a monotone submodular utility, attaining a $\\frac{1}{2}(1-1/e)$ approximation bound under a budget, and finally wraps threshold selection in Conformal Risk Control, giving a finite-sample distribution-free guarantee $\\mathbb{E}[L(S_{\\hat{\\gamma}})] \\le \\alpha$ on exchangeable calibration data. We evaluate on HealthHandoff-Bench, a synthetic benchmark with 21,600 healthcare-agent handoff tasks across four workflows, and Synthea-HH, a realism tier ingested from 340 Synthea-generated FHIR bundles. A three-source blinded label audit compares benchmark labels, model-generated labels, and clinical-informatics expert review, with high agreement for sensitivity labels and policy-boundary disagreements for necessity and authorization. Against authorized TF-IDF, embedding retrieval, a cross-encoder reranker, and a calibrated LLM-as-selector, ScopeSelect is the only method meeting $\\alpha = 0.15$ at competitive task success. The hard-mask safety property holds under 30% label noise, 100 adversarial attacks, and receiver-role shift; a live-LLM run with Claude Sonnet 4.5 exhibits a calibration-transfer gap that motivates per-workflow recalibration. We release the benchmark, Synthea-HH pipeline, CRC calibration harness, and a signed-handoff prototype.


RoMo Hands: A Large Scale Richly Organized Text to Hand Motion Dataset

Yizhak Ben-Shabat ⋅ Jiahao Zhang ⋅ Joseph Liu ⋅ Seonghyeon Moon ⋅ Young-Yoon Lee ⋅ Haomiao Jiang ⋅ Oren Jacob ⋅ Mubbasir Kapadia

While text-driven body motion generation has matured, pure-hand synthesis remains bottlenecked by data. Existing datasets often treat hands as rigid end-effectors, lack fine-grained captions, or are too limited in scale to support robust generalization. We introduce \textbf{RoMo Hands}, an in-the-wild, text-paired dataset of $\sim$542K sequences (884 hours, 95.5M frames at 30\,fps) which is an order of magnitude larger than prior work. To ensure high-quality supervision, a stringent dexterity filter discards low-articulation sequences, focusing the distribution on complex bimanual movement. Each clip is paired with five captions at increasing granularities from global semantic tags to atomic finger states, organized under a hierarchical taxonomy. To bridge the gap between body and hand architectures, we propose a novel 393-D bimanual representation that adapts redundant body-motion paradigms to the high-DoF kinematics of hands. Systematic benchmarking of state-of-the-art diffusion, masked-token, and autoregressive generators reveals that body-centric design choices do not transfer cleanly to hands, exposing a sharp divergence between semantic alignment and kinematic plausibility. Our dataset, representation, and benchmark establish a rigorous foundation for the next generation of dexterous motion synthesis.


RoPEMover: Depth-Aware Object Relocation via Positional Embeddings

Ipek Oztas ⋅ Duygu Ceylan ⋅ Aybars B Aksoy ⋅ Aysegul Dundar

Moving an object in a single image requires geometry-consistent spatial rearrangement, including handling occlusions, revealing previously unseen regions, and maintaining coherent shadows and reflections. Existing approaches are not well suited to this setting and often fail to preserve such scene-level consistency. We address this problem by introducing a geometry-aware object motion method that operates directly on the internal representations of diffusion transformers. Our key insight is that rotary positional embeddings (RoPE) define a structured spatial field that can be explicitly manipulated to induce controlled motion. We extend 2D RoPE into a depth-aware formulation that encodes 3D spatial structure, enabling consistent object displacement and scene-aware updates. Our model is trained using synthetic data combined with a small set of real images via parameter-efficient fine-tuning. Despite minimal real supervision, it preserves object identity under large spatial displacements, generates plausible content in newly revealed regions, and consistently updates scene-dependent effects such as shadows and illumination. Experimental results demonstrate the superiority of our method.

ARC tasks can be formulated as visual transformation problems, where a model must infer a high-level transformation rule from a small set of demonstrations and apply it consistently across a grid. While dense Vision Transformers (ViTs) perform well in this setting, they entangle local visual state and global transformation logic within a shared token representation, which can hinder consistent reasoning. In this paper, we introduce RoSeViT, a role-separated Vision Transformer that explicitly decouples high-level instruction inference from visual workspace computation. RoSeViT uses controller tokens to aggregate transformation-level information and workspace tokens to represent the visual grid, with a structured attention mechanism that routes nonlocal interactions through the controller. This design is motivated by a formal analysis showing that enforcing a shared instruction representation can reduce the effective complexity of visual transformation modeling. Under the standard VARC training and evaluation protocol for ARC, RoSeViT consistently outperforms dense ViT baselines. In a controlled backbone replacement setting, it improves ARC-1 accuracy from 54.5 to 56.6, and with recurrent refinement reaches 58.8. When integrated into the full system, RoSeViT-Ensemble improves performance from 60.4 to 62.6 on ARC-1. These results demonstrate that explicitly separating instruction and workspace representations leads to more effective visual reasoning on ARC tasks.

A two-branch routed regressor realizes exactly the function class of ordinary quadratic regression. Because routing does not enlarge the observable function class here, the routed model differs from quadratic regression only through the prior induced by routed coordinates. We compute this induced prior in closed form and show that fiber integration along the routing map creates a logarithmically singular density on the linear submodel. The induced posterior over quadratic coefficients then coincides exactly with the quadratic-regression posterior under the pushed-forward prior, isolating the role of prior geometry in this proxy model. For compactly supported priors on linear truth ($q_2=0$), the support-relative global real log canonical threshold (RLCT) is $(3/2,2)$ when the support reaches the collapsed stratum and $(3/2,1)$ when it avoids it; under the full-support Gaussian-coordinate prior, the global RLCT is $(3/2,2)$. Against smooth coefficient-space baselines with the same known noise variance on linear truth, this singular-versus-smooth prior comparison yields a rigorous $\log\log n$ Bayes-factor advantage. The same multiplicity-$2$ product singularity transfers locally to sigmoid gates on linear truth and extends locally to $d$-dimensional inputs.


RTEB: An Overfitting-Resistant Benchmark for Embedding Model Evaluation

Sahil Verma ⋅ Minghan Li ⋅ Andrew Gaut ⋅ Yujie Qian ⋅ Kaidi Cao ⋅ Zhenmei Shi ⋅ Kenneth Enevoldsen ⋅ Roman Solomatin ⋅ Isaac Chung ⋅ Tom Aarsen ⋅ Apoorva Joshi ⋅ Emilia Garcia-Casademont ⋅ Michael Günther ⋅ Niklas Muennighoff ⋅ Yevhen Kostiuk ⋅ Fődi Zoltán ⋅ Tengyu Ma ⋅ Frank Liu

Embedding models are central to modern information retrieval, retrieval-augmented generation, and downstream NLP systems, yet their evaluation remains unreliable: widely used benchmarks such as BEIR and MTEB rely on fully public datasets, making them vulnerable to overfitting that inflates scores without corresponding gains in real-world generalization. We introduce RTEB, a retrieval-focused embedding benchmark designed to address these limitations. Unlike MTEB's broad multi-task scope, RTEB is purpose-built for retrieval -- the setting most critical to RAG pipelines and agentic systems. RTEB emphasizes underrepresented yet high-impact domains, particularly code and agentic code retrieval, and keeps a large portion of its datasets private to structurally prevent direct test-set overfitting. We evaluate 20 models on 48 datasets (20 public, 28 private) spanning six domains and four languages. Rankings on RTEB diverge substantially from BEIR and MTEB ($\rho = 0.38$), agentic code retrieval emerges as the hardest domain with the best model reaching only 0.48 NDCG@10, and domain-specialized models fail to outperform strong generalists on any domain. Together, these design choices and findings position RTEB as a more reliable, discriminative, and future-proof benchmark for embedding model evaluation.


RuleSmith: Multi-Agent LLMs for Automated Game Balancing

Ziyao Zeng ⋅ Hao Wang ⋅ Chen Liu ⋅ Youheng Yao ⋅ Jingcheng Ni ⋅ Tianyu Liu ⋅ Xiatao Sun ⋅ Fengyu Yang ⋅ Chenyu You ⋅ Xiaofeng Liu ⋅ Daniel Rakita ⋅ Ronald Coifman ⋅ Yuval Kluger ⋅ Zhiwen Fan

Game balancing is a longstanding challenge requiring repeated playtesting, expert intuition, and extensive manual tuning. We introduce RuleSmith, a framework for automated game balancing that couples multi-agent LLM self-play with Bayesian optimization over a parameterized rule space. As a proof of concept, we build CivMini, a simplified civilization-style game with two asymmetric factions governed by 12 tunable parameters. LLM agents read textual rulebooks and game states to generate legal actions, enabling fast evaluation of balance metrics such as win-rate disparity. To search the rule space efficiently, we use Bayesian optimization with acquisition-based adaptive sampling: candidates with high Expected Improvement receive more evaluation games, while exploratory candidates receive fewer. RuleSmith consistently finds near-balanced configurations and produces interpretable parameter adjustments. It provides a scalable alternative to manual playtesting. Code will be available upon acceptance.


Runtime Monitoring of Perception-Based Autonomous Systems via Embedding Temporal Logic

Parv Kapoor ⋅ Abigail Hammer ⋅ Ashish Kapoor ⋅ Karen Leung ⋅ Eunsuk Kang

Runtime monitoring of autonomous systems traditionally relies on mapping continuous sensor observations to discrete logical propositions defined over low-dimensional state variables. This abstraction breaks down in perception-driven settings, where such mappings require additional learned modules that are often computationally expensive, brittle, and semantically misaligned. In this work, we propose Embedding Temporal Logic (ETL), a temporal logic that performs monitoring directly in learned embedding spaces. ETL defines predicates through distances between observed embeddings and target embeddings derived from reference observations. This formulation allows specifications to capture high-level perceptual concepts, such as similarity to visual goals or avoidance of semantic regions, that are difficult or impossible to express using traditional predicates. By composing these predicates with temporal operators, ETL naturally expresses temporally extended and sequential perceptual behaviors. We introduce ETL monitors for evaluating specifications over bounded embedding traces, along with a conformal calibration procedure that provides reliable and safety-oriented predicate evaluation. We evaluate our approach across multiple manipulation environments to show that ETL achieves strong empirical agreement with ground-truth semantics, including accurate monitoring of temporally composed behaviors.


Runtime Verification of Multiple Natural Language Criteria for Agent Governance

Silviu Pitis ⋅ Parand A. Alamdari ⋅ Jessica Tang ⋅ Toryn Klassen ⋅ Sheila McIlraith

AI agents are governed by hundreds of natural language criteria---yet existing verifiers cannot efficiently evaluate an agent's output against a large set of such criteria and report which are satisfied, violated, or inapplicable. We introduce the VFM, a tree-attention verifier that encodes the shared context and target text, and scores $N$ natural language criteria in parallel, returning independent per-criterion ternary verdicts. The VFM is pretrained on CriteriaBank, an open corpus of 357K (context, target, criterion, verdict) tuples spanning over 285K distinct criterion strings. A finetuned VFM-4B reaches 88.3\% accuracy on a synthetic governance benchmark, surpassing larger generative judges, and on real insurance-compliance calls it is competitive with GPT-5.4 with synthetic finetuning alone. On the DynaBench multi-criterion benchmark, a finetuned VFM-4B reaches $0.85$ F1 on trace-level failure detection versus DynaGuard-4B's $0.72$; the VFM additionally localizes the violated rule as its top-1 prediction in $98.4\%$ of failing traces, and a top-3 cascade to a 26B Gemma 4 chain-of-thought judge reaches $0.94$ F1. CriteriaBank pretraining improves cross-domain transfer on four additional benchmarks, particularly in the low-data regime. On an H100, VFM-4B evaluates 1000 criteria against a 4K-token context in under 2s at $<$10GB VRAM, making it suitable for runtime verification. We will release CriteriaBank, trained checkpoints, and evaluation code.


Saliency-Aware Multi-Route Thinking: Grounding and Reasoning on Vision-Language Agents

Mingjia Shi ⋅ Yinhan He ⋅ Yaochen Zhu ⋅ Cassie Dong ⋅ Jundong Li

A vision-language (VL) agent wraps a frozen vision-language model (VLM) into a test-time system that may call external tools during execution before producing a response. The central problem is how the VLM and its tools interact across multiple steps. Two recipes are common, both rigid. Thinking for longer (borrowed from text-only agents) sharpens reasoning, but without in-loop visual refresh, it lets visual evidence decay into hallucination. Querying tools more often, as traditional agents do, treats outputs as ground truth, a rigid commitment that fails when calls are noisy. We propose Saliency-Aware Principle Selection (SAP), a training-free inference-time method that avoids both rigidities: it searches a broad space of high-level principles, short textual directives (e.g., ``re-examine the image whenever an intermediate conclusion is formed'') that each prescribe a different way for the VLM to weigh tool advice across steps, and selects the principle whose parallel reasoning routes best fit the actual visual evidence. Tool outputs are treated as advisory (consulted, never authoritative), and a population-based evolutionary loop drives the search. Under the same token budget as long chain-of-thought, SAP substantially reduces object hallucination while remaining competitive on reasoning-heavy benchmarks. Being training-free and composed of independent routes, SAP is plug-and-play on any VLM and parallel for modern agent deployment.


SAMPPO: Structure-Aware Mirror Proximal Policy Optimization

Corinna Cortes ⋅ Mehryar Mohri ⋅ Yutao Zhong

The alignment of Large Language Models (LLMs) via reinforcement learning from human feedback (RLHF) typically relies on Proximal Policy Optimization (PPO), which uses a uniform clipping hyperparameter $\epsilon$ to maintain a stable trust region. However, we demonstrate that this uniform constraint is ill-suited for the non-uniform semantic topology of natural language. In ``semantically dense'' regions, where candidate tokens act as semantic synonyms within the current context, noisy advantage estimates can destabilize the fixed trust region, leading to *policy collapse*, a severe loss of generative diversity and policy entropy. To resolve this, we propose *Structure-Aware Mirror Proximal Policy Optimization (SAMPPO)*, which introduces an adaptive trust region governed by local semantic density. SAMPPO generalizes the trust-region framework to non-uniform semantic spaces by adapting the mirror-descent geometry to the local action topology. SAMPPO uses a per-timestep *Semantic Isolation Score* to dynamically tighten the clipping range in dense regions and relax it in sparse ones. We provide theoretical foundations for this approach, deriving a *Structure-Aware Monotonic Improvement* bound and establishing a formal link to structure-aware margin-shifted $H$-consistency theory. Empirically, SAMPPO prevents the entropy collapse observed in standard PPO baselines and significantly improves alignment. Evaluated on the Anthropic HH-RLHF and NVIDIA HelpSteer benchmarks across multiple model architectures (Llama-3-8B and Gemma-2-2B), SAMPPO achieves robust, zero-shot generalizable alignment gains. Crucially, it strictly outperforms standard PPO and recent empirical adaptive heuristics (BAPO), demonstrating competitive performance against direct preference baselines including DPO and SimPO on the alignment-diversity Pareto frontier.


SapiensID 2.0: Aligning Human Recognition Foundation Models with Human Perception

Yiyang Su ⋅ Jie Zhu ⋅ Feng Liu ⋅ Anil Jain ⋅ Xiaoming Liu

While foundation models have significantly advanced human recognition across diverse modalities, they predominantly rely on static, geometric feature extraction. This approach fundamentally diverges from human perception. Consequently, current models often suffer from ``semantic blindness,'' overfitting to transient noise while failing to leverage invariant soft biometrics, and struggle to capture temporal motion signatures. To bridge this gap, we propose SapiensID 2.0, a human recognition framework enriched with both semantic and temporal awareness. To overcome the lack of soft-biometric annotations, we transfer zero-shot semantic knowledge from Multimodal Large Language Models (MLLMs) into a discriminative embedding space. We resolve the dimensional mismatch between these spaces using Invariant Trait Alignment (ITA) to distill core persistent traits, and Transient Noise Disentanglement (TND) to decouple artifacts like clothing. Furthermore, we design a Kinematic Semantic Attention Head (K-SAH) that extends spatial attention across temporal windows. By tracking semantic patches over time, K-SAH captures rich kinematic signatures without requiring large-scale video datasets. Extensive experiments demonstrate that SapiensID 2.0 achieves state-of-the-art performance across image- and video-based person re-identification and gait recognition, while maintaining robust face recognition capabilities.


SARA: Step-Adaptive Rank Adjustment for Diffusion Inference Acceleration

DongHoon Lim ⋅ Jinwoo Jung ⋅ Won-Gook Choi ⋅ Joon-Hyuk Chang

Diffusion models achieve strong performance across image and audio synthesis, but their inference remains computationally expensive, requiring tens to hundreds of sequential evaluations of a large denoising network. We observe that the effective rank of the denoiser output varies sharply across timesteps and architectures, motivating step- and block-wise adaptation of computation. We propose Step-Adaptive Rank Adjustment (SARA), a training-free framework that treats the SVD rank of pretrained linear layers as an inference-time control variable, adapted per step and block by SARA-DP and per step by SARA-Online. SARA introduces the Low-Rank Approximation Error (LRAE) — the discrepancy a rank-r truncation induces against the full-rank network — and instantiates it in two algorithms: SARA-DP measures LRAE on a few warmup prompts and solves the globally optimal rank assignment under a MACs budget by dynamic programming; SARA-Online computes a closed-form, spectrum-based LRAE from the model output at inference time, requiring no calibration and no trajectory lookahead. SARA applies uniformly to UNet (SDXL), MMDiT (SD3), DiT (Stable Audio Open), and hybrid MMDiT/DiT (TangoFlux) backbones, spanning ϵ, v, and rectified-flow parameterizations. SARA combines with existing solver- and cache-axis acceleration to yield compounded wall-clock speedups in our experiments. Together, these results introduce rank as a new axis of diffusion inference acceleration.

Sparse Autoencoders (SAEs) are widely used for mechanistic interpretability in large language models, yet their standard formulation assigns each latent feature a single decoder direction, implicitly constraining features to be one-dimensional. We identify a fundamental geometric mismatch between this assumption and the structure of many model features, and show that it can provably induce feature splitting. In particular, for any feature with intrinsic dimension $d_i \ge 2$, achieving reconstruction error $\varepsilon$ with a standard SAE requires $\Omega \left((\frac{1}{\varepsilon})^{d_i-1}\right)$ distinct decoder directions. As a result, the SAE is forced to represent a single coherent high-dimensional feature using many nearly collinear one-dimensional latents, leading to spurious feature multiplicity and a failure to preserve the intrinsic structure required for mechanistic interpretability. Motivated by this lower bound, we introduce Subspace-Aware Sparse Autoencoders (SASA), which replace single-vector decoders with learned decoder subspaces. SASA enforces block sparsity through Top-$s$ group gating and adapts each group's effective dimensionality using a trace-norm regularizer. We prove a complementary upper bound: when the block size $r \ge d_i$, a single group can represent the entire feature slice, alleviating feature splitting while also improving sample complexity. Finally, we empirically validate that SASA improves feature splitting, monosemanticity, and interpretability while also enabling more efficient training.

When deploying large language models (LLMs) to safety-critical applications, uncertainty quantification (UQ) is of utmost importance to self-assess the reliability of the LLM-based decisions. However, such decisions typically suffer from overconfidence, particularly after parameter-efficient fine-tuning (PEFT) for downstream domain-specific tasks with limited data. Existing methods to alleviate this issue either rely on a Laplace approximation based post-hoc framework, which may yield suboptimal calibration depending on the training trajectory, or variational Bayesian training that requires multiple complete forward passes through the entire LLM backbone at inference time for Monte Carlo estimation, posing scalability challenges for deployment. To address these limitations, we build on the Bayesian last layer (BLL) model, where the LLM-based deterministic feature extractor is followed by random last-layer parameters for uncertainty reasoning. Since existing low-rank adapters (LoRA) for PEFT have limited expressiveness due to rank collapse, we address this with Polar-decomposed Low-rank Adapter Representation (PoLAR), an orthogonalized parameterization paired with Riemannian optimization to enable more stable and expressive adaptation. Building on this PoLAR-BLL model, we leverage the variational (V) inference framework to put forth a scalable Bayesian fine-tuning approach which jointly seeks the PoLAR parameters and approximate posterior of the last-layer parameters via alternating optimization. The resulting PoLAR-VBLL is a flexible framework that nicely integrates architecture-enhanced optimization with scalable Bayesian inference to endow LLMs with well-calibrated UQ. Our empirical results verify the effectiveness of PoLAR-VBLL in terms of generalization and uncertainty estimation on both in-distribution and out-of-distribution data for various common-sense reasoning tasks.


ScaleBITS: Scalable Bitwidth Search for Hardware-Aligned Mixed-Precision LLMs

Xinlin Li ⋅ Timothy Chou ⋅ Joshua W Fromm ⋅ Zichang Liu ⋅ Yunjie Pan ⋅ Christina Fragouli

Post-training weight quantization is crucial for reducing the deployment cost of large language models (LLMs), yet maintaining model quality in the ultra-low-bit regime ($\le$ 3 bits) remains challenging due to highly non-uniform weight sensitivity and the lack of principled precision allocation strategies. In this work, we show that the weight sensitivity is inherently dynamic and depends on the current quantized model, rather than the full-precision reference used in existing methods. Leveraging this observation, we introduce a sensitivity estimate that better captures the changes in marginal loss during progressive quantization. We further uncover a structured pattern in sensitivity distribution, revealing a bi-directional concentration across input and output channels, which enables hardware-aligned yet expressive partitioning via channel reordering. Building on these insights, we propose ScaleBITS, a scalable framework that combines structured partitioning and an efficient approximation to greedy allocation. Our approach enables fine-grained, global bitwidth search while preserving hardware efficiency. Experiments show that ScaleBITS significantly improves over uniform-precision quantization (up to +36%) and outperforms state-of-the-art sensitivity-aware baselines (up to +13%) in the ultra-low-bit regime, without adding runtime overhead.


Scaling Whole-Body Loco-Manipulation through Compositional Data Synthesis

Yufei Zhu ⋅ Runyi Yu ⋅ Yinhuai Wang ⋅ Aoru Xue ⋅ Qin Sun ⋅ Qingqiu Huang ⋅ Xinge ZHU ⋅ Zhiyang Dou ⋅ Yuexin Ma

Acquiring diverse whole-body dexterous manipulation data for humanoids remains a fundamental challenge in character animation and robotics. Existing pipelines rely on expensive mocap data collected from many subjects and require retargeting to a unified humanoid shape, which limits scalable data construction and often yields physically inconsistent interactions (e.g., unstable contacts). We present \textbf{ManipSynth}, a unified framework for whole-body loco-manipulation synthesis, with compositional generation of sparse whole-body motion, object-centric grasp anchors, and contact-aware interaction transitions. Then, we scale data generation through randomized pre-grasp motions, grasp poses, and object trajectories. ManipSynth achieves approximately 5$\times$ higher contact coverage and lower average penetration than OMOMO across representative objects. To test whether better contact-consistent data improve scalable control learning, we train \textbf{ManipTracker}, a general-purpose Humanoid-Object Interaction (HOI) tracking policy, on purely synthetic demonstrations. Under the same object set, number of training trajectories, the same ManipTracker trained on ManipSynth demonstrates stronger generalization than when trained on mocap data, while achieving over 90\% training success.

Robots manipulating in cluttered spaces must often plan from partial views where foreground objects occlude the target, surrounding geometry, and usable free space. Reactive policies often miss these hidden constraints, while exhaustive search becomes computationally prohibitive in dense clutter. We present SCOUT, an occlusion-aware planning framework that bridges high-level deliberation with grounded physical rollouts. SCOUT leverages an object-centric scene representation to reconstruct hidden geometry and derive a relational planning state capturing blocker hierarchies and accessibility corridors. To navigate the vast action space in clutter, an LLM proposes structured action candidates which are then rigorously evaluated via Monte Carlo Tree Search (MCTS) using a learned action-conditioned world model. Evaluations on ‘Reveal’ and ‘Placement’ tasks show that SCOUT outperforms Vision-Language-Action (VLA) policies, embodied agents, and 3D world models, with significant gains in long-horizon tasks requiring reasoning over occluded geometry. We further validate the framework’s robustness through real-world robotic manipulation in a cluttered shelf environment.


Self-Consuming Generative Models with Co-Evolving Human Preferences

Xiukun Wei ⋅ Tian Xie ⋅ Ding Zhu ⋅ Xueru Zhang

Generative models are increasingly trained in self-consuming iterative loops, where users curate preferred samples from model-generated candidates and the curated samples are used to train future generations of the model. Prior work has largely assumed fixed user preferences, but in practice exposure to model outputs gradually reshapes what users perceive as desirable, creating a feedback loop in which model distributions and user preferences co-evolve. We take a first step toward understanding the long-term behavior of such coupled dynamics. We show that when training relies entirely on user-curated synthetic data, iterative curation amplifies initial biases and drives the system toward one of multiple singleton equilibria in which the instance holding an initial advantage eventually dominates. In contrast, injecting reference data into training at a sufficiently large rate fundamentally changes the dynamics and yields a unique globally attracting equilibrium. Building on this insight, we study how reference-data injection can be used to control long-term outcomes, and propose an efficient algorithm that jointly selects a reference distribution and its mixing weight to steer the coupled system toward equilibria that preserve desired attributes while minimizing data collection costs.


Self-Supervised Doppler-Guided RF Odometry

Kunzhe Song ⋅ Jialuo Du ⋅ Jingkai Lin ⋅ Huacheng Zeng

Accurate RF odometry remains difficult due to multipath interference, geometric degeneracy, and dynamic clutter. Existing approaches mitigate these challenges by relying on auxiliary sensors or ground-truth pose supervision, which limits scalability and deployability. We present RFWalk, a self-supervised, RF-only odometry framework that learns cross-frame correspondence without any pose annotations or auxiliary sensors. Our key insight is to cast each range-azimuth cell as a node in a space-time graph and recover inter-frame motion by learning a soft transition matrix through cycle-consistent contrastive random walks. To resolve geometric degeneracy in feature-poor scenes, we further introduce a physics-informed Doppler consistency loss that enforces agreement between the correspondence-induced displacement field and measured radial velocities. Experiments on challenging scenes show that RFWalk achieves competitive accuracy in static environments and substantially outperforms existing approaches in dynamic scenes.


Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning

Prabin Kumar Rath ⋅ Omkar Patil ⋅ Nakul Gopalan

Behavior cloning (BC) in non-Markovian environments is a challenging problem because policies have to reason over contextual information over long horizons. Existing policy architectures rely on recurrent or attention-based mechanisms to capture long-term dependencies. However, recurrent models suffer from hidden-state collapse and gradient instability under backpropagation through time, while attention-based models are fundamentally limited by context length. To address these issues, we propose Keyframe Mnemonics, a novel self-supervised method that $\textit{discovers}$ a set of information-critical observations ($\textit{mnemonics}$) by learning an objective from randomly sampled past observations and using it as a reward for keyframe selection. We then train a BC policy that conditions on the discovered keyframes to model the action distribution. Under near-optimal keyframe selection, our formulation provides context retention guarantees over an infinite horizon, while maintaining a small set of decision-relevant keyframes in the policy's working memory. We evaluate our method on synthetic memory domains, where mnemonic-conditioned BC policies achieve $100$% success rates (SR) and generalize to horizons orders of magnitude beyond training without performance degradation. Additionally, we evaluate on memory intensive robot manipulation benchmarks, where we observe an absolute $15$-$30$% improvement in SR over strong memory-augmented baselines on $23$ tasks. Code and videos are available at https://keyframe-mnemonics.github.io.

Sparse autoencoders (SAEs) have become a central tool for interpreting language models. However, two key SAE analyses that remain difficult to scale are (1) matching semantically similar features across multi-layers and (2) compressing large feature circuits into interpretable supernodes. Although these have been treated as separate problems, we show that both are instances of a more fundamental challenge, which we frame as the estimation of semantic distances between SAE features that lie on different activation manifolds. We introduce a distributional framework for this problem, in which each feature is represented not by a single decoder vector like in the literature, but by an activation-weighted distribution over the hidden states that express it. By projecting these distributions into a shared reference space and comparing them with Wasserstein distance, our method provides a unified semantic metric for cross-layer feature comparison. We prove that our representation is invariant to activation rescaling, stable under perturbations, and recovers true matches under finite-sample margin conditions. Empirically, our method outperforms decoder-vector and LLM-based baselines and captures subtle functional distinctions between related features. Notably, our method compresses large feature circuits into interpretable supernodes automatically.


Semantic Search over 9 Million Mathematical Theorems

Luke Alexander ⋅ Eric Leonen ⋅ Sophie Szeto ⋅ Artemii Remizov ⋅ Ignacio Tejeda ⋅ Jarod Alper ⋅ Giovanni Inchiostro ⋅ Vasily Ilin

Searching for mathematical results remains difficult: most existing tools retrieve entire papers, while mathematicians and theorem-proving agents often seek a specific theorem, lemma, or proposition that answers a query. While semantic search has seen rapid progress, its behavior on large, highly technical corpora such as research-level mathematical theorems remains poorly understood. In this work, we introduce and study semantic theorem retrieval at scale over a unified corpus of 9.2 million theorem statements extracted from arXiv and seven other sources, representing the largest publicly available corpus of human-authored, research-level theorems. We represent each theorem with a short natural-language description as a retrieval representation and systematically analyze how representation context, language model choice, embedding model, and prompting strategy affect retrieval quality. On a curated evaluation set of theorem-search queries written by professional mathematicians, our approach substantially improves both theorem-level and paper-level retrieval compared to existing baselines, demonstrating that semantic theorem search is feasible and effective at web scale. The project page, search tool, dataset, REST API, and MCP server are available at [Redacted]

Representational Engineering (RepE) aims to expose and steer high-level concept directions in language model activations, and recent work shows that hidden states can be transferred across heterogeneous architectures via trained neural adapters. Whether specific concepts survive such transfer remains an open question. We study one such concept, \textit{truth}, by mapping source-model activations through autoencoder adapters into target-model representation space, varying injection strength $\alpha \in [0, 1]$, and reading the result with the target's native truth probe across 21 cross-family trajectories from four model families, with comparison against a closed-form rigid alignment baseline on a representative subset. We find that ordinal truth ranking is preserved through partial injection (median AUROC near the native ceiling through $\alpha \in [0, 0.7]$) but collapses sharply at full injection ($\alpha{=}1$), with parametric separability dropping $\sim$$14\times$ while rank order partially survives. The apparent accuracy overshoot under partial injection decomposes cleanly into two effects: a large recalibration component dominated by native probe miscalibration ($R^2 = 0.75$), and a small but universal genuine separability gain present in all 21 trajectories. These findings characterize cross-family truth structure as partially compatible, ordinally stable, but non-isomorphic: preserved enough under partial injection that learned and rigid alignment methods both surface meaningful truth signal, but with full replacement representing a hard failure mode that practitioners should treat as a deployment boundary.


She Performs Your Voice: A Unified Speech and Dance Motion Model

Lurui Wang ⋅ yongkang cheng ⋅ Can Xu

Expressive performers resonate with audiences because they transform sound into motion: rhythm invites steps, prosody evokes gestures, and energy shapes the flow of the body. However, existing audio-driven character systems often reduce audio to a task-specific condition, separately optimizing for co-speech gestures or music-driven dance. This fragmented view overlooks audio as the shared perceptual driver of performance, making it difficult to generate coherent full-body motion across speaking, dancing, and their intermediate states. We present VOXPERFORMER, a unified audio-to-motion framework that moves toward a generalized performance foundation model. Our key insight is that audio is not merely an external condition, but the hidden soul of performance style, governing how motion emerges, evolves, and transitions. To realize this, VOXPERFORMER introduces a structured motion prior to organize the full-body performance space, audio-to-prior distillation to transform sound into latent motion cues, a memory-augmented prior to resolve audio-motion ambiguity, and a decoupled diffusion generator with implicit transition matching for efficient and coherent synthesis. Experiments on BEAT2 and FineDance show that VOXPERFORMER remains competitive on speech-driven gesture generation while substantially improving hand-aware music-driven dance generation. User studies further show that participants prefer VOXPERFORMER, indicating stronger audio-motion correspondence and more natural perceptual resonance.


SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation

Tianjiao (Joey) Yu ⋅ Xinzhuo Li ⋅ Yifan Shen ⋅ Ying Shen ⋅ Kiet Nguyen ⋅ Adheesh S Juvekar ⋅ Ismini Lourentzou

High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragments continuous surfaces into many local tokens, inflates generation cost, and often weakens topological consistency for thin or highly connected shapes. We introduce SILSA, a topology-aware 3D generation framework that represents shapes with compact sliding-window slice latents. Instead of generating expensive voxel tokens, SILSA uses a fixed set of overlapping slices along the three canonical axes, where each token summarizes a local depth window to preserve cross-sectional continuity and support single-stage rectified-flow generation. A Slice VAE encodes oriented surface samples into multi-axis slice latents and reconstructs them with a sparse volumetric decoder, while a Volumetric Anchor Lattice coordinates directional slice streams through a shared 3D workspace. To preserve structural correctness, we introduce slice-level topology supervision that matches persistence diagrams and aligns Betti transitions across neighboring slices. Experiments show that SILSA improves structural fidelity while substantially reducing generation cost. SILSA improves PSNR by 8.7%, coverage by 5.96 absolute points, and Betti error by 9.2% over the strongest baseline, while using 70.0% fewer tokens than the next-most compact baseline and over 98% fewer tokens than sparse or hierarchical tokenizers, reducing training memory by 40.4% and inference time by 58.5%. Qualitative results further show improved preservation of thin structures, repeated components, and long-range connectivity.

Being robust to the presence of outliers is crucial for applying clustering algorithms in practice. In the $\textit{robust $k$-Means}$ problem (i.e., $k$-Means with outliers), the goal is to remove $z$ outliers and minimize the $k$-Means cost on the remaining points. Despite the close connection between robust $k$-Means and outlier detection, both theoretical and empirical understanding of the effectiveness of $\textit{classic outlier detection heuristics}$ for robust $k$-Means remains limited. In this paper, we prove that under a practical assumption on the optimal cluster sizes, simply removing points with large $K$-Nearest-Neighbor distances achieves performance comparable to prior work in terms of approximation guarantees: it yields a constant-factor reduction from robust $k$-Means to standard $k$-Means, without introducing additional centers or discarding extra outliers, as is commonly required by existing approaches. Empirically, experiments on real-world datasets show that our method outperforms or matches several more sophisticated algorithms in terms of clustering cost and runtime. These results demonstrate that simple KNN-based heuristics can be surprisingly effective for robust clustering, highlighting new opportunities to bridge techniques from outlier detection and clustering.


Skill-Coupled Policy Optimization with Calibrated Group-Wise Advantage Estimation

Peng Wang ⋅ Zhendong Chu ⋅ Chengshuai Shi ⋅ Matthias Poloczek ⋅ Cong Shen ⋅ Jing Yang

Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language models on reasoning tasks. To avoid the complexity of training explicit critics, algorithms like GRPO estimate advantages directly from sparse on-policy rollouts. However, the efficacy of this approach relies heavily on the accuracy of the baseline estimator. In this work, we argue that standard methods, whether relying on prompt-local averages or shrinking toward a single global mean, are statistically mis-specified for reasoning domains. By enforcing a uniform prior, these methods ignore the latent skill structure of tasks, introducing estimation bias that conflates dissimilar problems. To address this, we propose Skill-Coupled Policy Optimization (SCPO), a structured baseline estimation framework that refines the shrinkage prior by leveraging task correlations. Instead of a generic global target, SCPO performs shrinkage toward skill-coupled group targets, thereby aligning the baseline with the local difficulty landscape. To further stabilize these targets under limited on-policy observations, SCPO incorporates a lightweight history tracker. Experiments across diverse models and tasks show that SCPO consistently improves training stability and performance over both prompt-level and global shrinkage baselines.


SLIC: Reinforcement Fine-Tuning Small LMs for Multi-Turn Analog Circuit Optimization

Yan Xu ⋅ Zhengqi Gao ⋅ Ching-Yun Ko ⋅ Zhang-Wei Hong ⋅ Wenjie Lu ⋅ Tao Yu ⋅ Xin Zhang ⋅ Duane Boning ⋅ Ruonan Han

Analog integrated circuit (IC) design is a multi-turn, simulator-grounded engineering task that demands iterative reasoning over coupled performance trade-offs, yet proprietary process design kits (PDKs) make hosted large-language-model workflows difficult to deploy in practical design settings. We present SLIC — Small LM Iterating on Circuits — a reinforcement fine-tuning framework that trains a 7B-scale, locally deployable language model for closed-loop analog circuit optimization. SLIC pairs structured circuit representations with discretized relative-update actions, domain- and difficulty-gated multi-objective rewards, and Reward-Guided Rollout Steering (RGRS), which turns per-step simulator feedback into both policy-learning and trajectory-steering signals in this sparse, highly constrained domain. We further release an open, goal-conditioned benchmark of 34 amplifier topologies with standardized protocols for in-distribution validation, held-out topology evaluation, and cross-technology-node transfer. Trained solely on a 22 nm silicon CMOS PDK, SLIC achieves 58.7% in-distribution Pass@1 under a 4-turn simulator budget, surpassing the 1.6T DeepSeek-V4-Pro Think baseline by 12.7 percentage points. Beyond the training distribution, it reaches 51.5% Pass@1 on held-out topologies, improving over the same baseline by 10.6 percentage points, and consistently transfers across five unseen technology nodes (32–130 nm) without additional fine-tuning. Performance further scales with test-time simulator interactions, reaching roughly 70% Pass@1 at an 8-turn budget, demonstrating that SLIC enables the small policy to convert additional simulator calls into improved circuit designs.


SOL-ExecBench: Speed-of-Light Benchmarking for Real-World GPU Kernels Against Hardware Limits

Edward C Lin ⋅ Sahil Modi ⋅ Siva Kumar Sastry Hari ⋅ Qijing Huang ⋅ Zhifan Ye ⋅ Nestor Y. Qin ⋅ Fengzhe Zhou ⋅ Yuan Zhang ⋅ Jingquan Wang ⋅ Sana Damani ⋅ Dheeraj Peri ⋅ Ouye Xie ⋅ Aditya Kane ⋅ Michael behar ⋅ Rishabh Mehta ⋅ Vartika Singh ⋅ Vikram Sharma Mailthody ⋅ Terry Chen ⋅ Zihao Ye ⋅ Hanfeng Chen ⋅ Tianqi Chen ⋅ Vinod Grover ⋅ Wei Chen ⋅ Wei Liu ⋅ Eric Chung ⋅ Luis Ceze ⋅ Roger Bringmann ⋅ Michael Lightstone ⋅ Christos Kozyrakis ⋅ Humphrey Shi

As agentic AI systems become increasingly capable of generating and optimizing GPU kernels, progress is constrained by benchmarks that reward speedup over software baselines rather than proximity to hardware-efficient execution. We present SOL-ExecBench, a benchmark of 235 CUDA kernel optimization problems extracted from 124 production and emerging AI models spanning language, diffusion, vision, audio, video, and hybrid architectures, targeting NVIDIA Blackwell GPUs. The benchmark covers forward and backward workloads across BF16, FP8, and NVFP4, including kernels whose best performance is expected to rely on Blackwell-specific capabilities. Unlike prior benchmarks that evaluate kernels primarily relative to software implementations, SOL-ExecBench measures performance against analytically derived Speed-of-Light (SOL) bounds computed by SOLAR, our pipeline for deriving hardware-grounded SOL bounds, yielding a fixed target for hardware-efficient optimization. We report a SOL Score that quantifies how much of the gap between a release-defined scoring baseline and the hardware SOL bound a candidate kernel closes. To support robust evaluation of agentic optimizers, we additionally provide a sandboxed harness with input correctness checks, isolated compilation and execution, CUPTI-based GPU activity timing, and defenses against common reward-hacking strategies. A third-party post-release evaluation by Cursor reported a $1.38\times$ geometric-mean speedup over a single-agent optimized PyTorch baseline and an overall SOL Score of 0.56 across the full benchmark. SOL-ExecBench reframes GPU kernel benchmarking from beating a mutable software baseline to closing the remaining gap to hardware Speed-of-Light.


Space-Optimal Streaming Algorithms via Efficient Encodings

Liudeng Wang ⋅ David Woodruff ⋅ Shenghao Xie ⋅ Samson Zhou

We study the problem of constructing coresets on data streams: Given a dataset of $n$ points in $\mathbb{R}^d$ arriving sequentially, the goal is to maintain a small weighted subset that accurately approximates the objective values of the entire dataset for a given class of query functions. Previous methods like merge-and-reduce and online sensitivity sampling required more space than the size of an optimal such subset in the offline setting. In this work, we show that such overheads are not necessary by presenting space-optimal streaming algorithms for central high-dimensional optimization problems such as Euclidean $(k,z)$-clustering and $L_p$ subspace embeddings, as well as applications to projection-cost preserving sketches. For $(k,z)$-clustering, our streaming algorithm uses memory independent of the number $n$ of input points (in words of memory) and the aspect ratio $\Delta$, yielding a coreset with an optimal $\tilde{\mathcal{O}}\left(\frac{dk}{\min(\varepsilon^4,\varepsilon^{z+2})}\right)$ words of memory for accuracy parameter $\varepsilon\in(0,1)$. For $L_p$ subspace embeddings, where the rows of matrix $\mathbf{A} \in \mathbb{R}^{n \times d}$ arrive in an insertion-only stream, we achieve an efficient algorithm that uses optimal space for different ranges of $p$: $\tilde{\mathcal{O}}\left(\frac{d^2}{\varepsilon^2}\right)$ words for $p\le 2$ and $\tilde{\mathcal{O}}\left(\frac{d^{p/2+1}}{\varepsilon^2}\right)$ words for $p>2$, both independent of the number of rows $n$. Thus, our work shows that streaming algorithms can match offline algorithms in space complexity.

Contemporary systems serving large language models (LLMs) have adopted prefill-decode disaggregation to better load-balance between the compute-bound prefill phase and the memory-bound decode phase. Under this design, prefill workers generate a KV cache that must be transferred to decode workers before token generation can begin. With these workers residing on different physical systems, this transfer becomes a significant bottleneck to serving LLMs at scale. This bottleneck gets exacerbated for long-input and agentic workloads. Existing lossless codecs are not suited to this setting as they primarily target offline weight compression, run on the CPU, or use variable-length coding whose decompression is fast but compression is too slow to keep up with KV production during prefill. We introduce SplitZip, a GPU-friendly lossless compressor for KV-cache transfer that preserves KV tensors bitwise and integrates into existing serving frameworks without changes to model execution. SplitZip exploits redundancy in floating-point exponents of KV activations, encoding the most frequent exponent values with fixed-length codes and routing rare exponents through a sparse escape stream of (position, value). An offline calibrated top-16 exponent codebook eliminates online-histogramming, while the regular dense path and sparse escape correction make both encoding and decoding efficient on GPUs. On real BF16 activation tensors, SplitZip achieves 613.3 GB/s compression throughput and 2181.8 GB/s decompression throughput, substantially outperforming prior lossless compressors on the latency-critical codec path. End-to-end transfer experiments show up to 1.32× speedup for BF16 KV-cache transfer, 1.30× speedup for TTFT, and 1.23× increase on Request Throughput. The same approach extends to FP8 KV caches, providing up to 1.14× compression over native E5M2.


Stability-Constrained Regime-Aware Forecasting for Heterogeneous Panel Time Series

Dmitry Zaytsev ⋅ Valentina V Kuskova ⋅ Michael Coppedge

Social science panel time series exhibit multicollinearity, latent unit heterogeneity, and regime-conditional sign-changing relationships. On V-Dem democratic development data, four of four selected covariates exhibit coefficient sign reversal across regimes, driving pooled additive predictors to cancel real signal - a diagnostic we establish before introducing any nonlinear model. We present a regime-aware additive forecasting architecture that recovers this signal under a verifiable contraction certificate. Our theoretical contribution is a rate-based contraction theorem decomposing the global Lipschitz constant of a soft-gated mixture predictor into per-regime, regime-disagreement, and dwell-time components, each empirically auditable on a trained model. On V-Dem democratic development data ($151$ countries, $1970-2025$), the architecture recovers regime-conditional signal that pooled additive models cancel (paired Diebold-Mariano $p=0.016$), with the recovery localized to the tails of the democracy index where a diagnostic shows largest covariate sign disagreement. The contraction certificate $L_{\text{eq,max}}<1$ holds in $5/5$ seeds across all $9$ audited configurations without explicit projection. A synthetic stress test with stronger regime contrasts shows the certificate is non-vacuous: spectral normalization on per-edge layers can fail to deliver contraction because the gating-driven term $L_{\phi,\text{eq}} \cdot M$ dominates, and a regime-disagreement penalty is required to recover the bound.


Stability in Multi-Step Reasoning via Jacobian-based Error Accumulation Analysis

Dongyue Li ⋅ Alice Duan ⋅ Ziniu Zhang ⋅ Hongyang Zhang

We study the accumulation of test errors in language models performing multi-step reasoning. Although longer reasoning improves model capabilities by scaling test-time computation, test error compounds in autoregressive generation and can grow substantially across steps. The key factors determining the stability of such reasoning remain unclear. To address this question, we analyze the test error of multi-step reasoning under supervised fine-tuning. We first present a test error bound governed by the product of input Jacobian spectral norms across generation steps. This product, summarized as an error amplification factor, scales exponentially with steps and controls the reasoning stability. Then, we explicitly analyze this factor in transformers trained to predict linear and quadratic function weights. We prove that the transformer converges to a solution where this amplification factor decays, yielding near-zero test loss even over longer steps. Based on the analysis, we propose a training method to suppress the error amplification factor by combining (i) chain-of-thought length compression that reduces the reasoning steps, and (ii) quantization-aware training that provably regularizes the input Jacobian norms. We validate our method by fine-tuning language models on reasoning tasks that require executing graph algorithms and tracking variable states through logical operations. Across seven tasks, our method improves over baselines by 3.5\% on average, and by 8.2\% in length generalization evaluations. We further validate that it reduces the error amplification factor in fine-tuned models.


STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation

Ying Shen ⋅ Tianrong Chen ⋅ Yuan Gao ⋅ Yizhe Zhang ⋅ Yuyang Wang ⋅ Miguel Angel Bautista ⋅ Shuangfei Zhai ⋅ Joshua Susskind ⋅ Jiatao Gu

Deep generative models have advanced rapidly across text and vision, motivating unified multimodal systems that can understand, reason over, and generate interleaved text–image sequences. Most existing approaches combine autoregressive language modeling with diffusion-based image generators, inheriting a structural mismatch between causal text generation and iterative visual denoising. We observe that autoregressive normalizing flows are autoregressive Transformers—sharing the same causal mask, KV-cache mechanism, and left-to-right structure as LLMs— making them the most natural paradigm for true unified multimodal generation. We present STARFlow2, built on the Pretzel architecture that vertically interleaves a pretrained VLM stream with a TarFlow stream via residual skip connections, both operating under the same causal mask. Combined with a deep-shallow flow design and a unified FAE latent space, STARFlow2 enables cache-friendly interleaved generation where both text and visual outputs directly enter the KV-cache without re-encoding. Experiments demonstrate strong performance across image generation and multimodal understanding benchmarks, validating autoregressive flows as a viable foundation for unified multimodal modeling.


State-of-art minibatches via novel DPP kernels: discretization, wavelets, and rough objectives

Hoang Son Tran ⋅ Pranav Gupta ⋅ Rémi Bardenet ⋅ Subhroshekhar Ghosh

Determinant point processes (DPPs) have recently emerged as a kernelized alternative to vanilla independent sampling for generating efficient minibatches, coresets and other parsimonious representations of large-scale datasets. While basic theoretical foundations and promising empirical performance have been demonstrated, there are two main challenges for current proposals for DPP-based coresets or minibatches. The first is the need for families of DPPs with certain key variance reduction properties, usually constructed in a continuous setting, of which there are few known examples. The second is the need for an ad-hoc construction of a discrete DPP defined on a given dataset, that inherits such variance reduction. In this work, we contribute to the programme of establishing DPPs as a subsampling toolbox for ML by advancing on these two fronts. First, we propose new DPPs on the Euclidean space based on wavelets, with provably better accuracy guarantees than the best known rates. Second, we introduce a general method to convert such continuous DPPs, which are more amenable to proving analytical statements, into discrete kernels, which are pertinent for subsampling tasks such as minibatch and coreset constructions. This conversion mechanism simultaneously preserves the desired variance decay and reveals a low-rank decomposition of the discrete kernel, which makes sampling the corresponding DPP computationally inexpensive. En route, we enlarge the class of ML tasks amenable to improvements via DPP-based minibatches and coresets to include highly non-smooth objective functions, with arbitrarily low regularity, and rate guarantees that adapt explicitly to this regularity.


Stochastic Dynamic Barrier Perturbed Gradient Methods for Nonconvex Simple Bilevel Optimization

Mohammad Mahdi Ahmadi ⋅ Jincheng Cao ⋅ Aryan Mokhtari ⋅ Erfan Yazdandoost Hamedani

We study stochastic simple bilevel optimization with smooth, possibly nonconvex upper- and lower-level objectives accessed only through stochastic gradient oracles. A key challenge is that the dual multiplier induced by the lower-level constraint may become unbounded near lower-level stationary points, invalidating bounded-dual analyses and destabilizing stochastic gradient estimates. To address this, we propose \emph{Stochastic Dynamic Barrier Perturbed Gradient} (SDBPG), a single-loop method that adaptively perturbs the dual formulation to regularize this degeneracy. The perturbation stabilizes the multiplier and yields controlled bias and variance even near the lower-level stationarity region. Under a mild rare-visit assumption, SDBPG finds an $(\epsilon_f,\epsilon_g)$-stationary point in $\mathcal{O}(\max\{\epsilon_f^{-2},\epsilon_g^{-2}\})$ iterations, with sample gradient complexities $\mathcal{O}(\epsilon^{-4})$ and $\mathcal{O}(\epsilon^{-6})$ for the upper- and lower-level objectives where $\epsilon=\max(\epsilon_f,\epsilon_g)$. We further develop PR-SDBPG, a penalty-regularized variant that eliminates the rare-visit assumption, and VR-PR-SDBPG, which improves the resulting sample complexities entirely through variance reduction. To our knowledge, these are the first explicit $(\epsilon_f,\epsilon_g)$-stationarity guarantees for stochastic nonconvex-nonconvex simple bilevel optimization.

This position paper argues that the NeurIPS community should prioritize AI-for-Quantum (AI4Q) over Quantum Machine Learning (QML) on classical data, and should review the two directions by different evidentiary standards. We argue that QML on classical data has not earned high priority because its leading advantage claims rely on strong input assumptions, face trainability and noise barriers, and still lacks persuasive wall-clock wins on standard benchmarks. By contrast, AI4Q has already produced concrete gains in decoding, compilation, neural-state representation, tomography and readout, and error mitigation. We make the evidentiary standard explicit through falsification criteria and concrete review-process recommendations.


Streaming Interventions: Can Video LLMs Correct Mistakes as They Occur?

Apratim Bhattacharyya ⋅ Shweta Mahajan ⋅ Sanjay Haresh ⋅ Reza Pourreza ⋅ Litian Liu ⋅ Rajeev Yasarla ⋅ Risheek Garrepalli ⋅ Roland Memisevic

Learning everyday skills, like cooking a dish, relies increasingly on instructional media such as online videos. This opens the door to the use of video (and multimodal) large language models (LLMs) as task guidance assistants. A crucial capability of the real-world success of a prospective task guidance assistant is its ability to intervene proactively as soon as a mistake is apparent in order to guide the user. To evaluate this crucial capability, we introduce Ego-MC-Bench (Mistake Corrections), a benchmark for evaluating reactive, step-by-step task guidance in realistic cooking scenarios. Extensive experiments show that Ego-MC-Bench is highly challenging for state‑of‑the‑art video LLMs. We argue that a key reason is the limited availability of training data for fine-tuning models on this task. Although there exists a wide range of cooking video datasets, existing datasets lack examples of mistakes along with appropriately timed interventions. To help address this data limitation, we also introduce Ego-CoMist, a counterfactual synthetic dataset created by transforming non‑interactive cooking videos into supervised training examples showing proactive interventions. We show that fine-tuning on Ego-CoMist yields performance gains especially for smaller and more efficient video LLMs that are well suited for delivering assistance on edge devices.


Structured Human-Like Agentic Flow for RTL Design

Yu-Tung Liu ⋅ Zhan Song ⋅ Chenhui Deng ⋅ Chia-Tung Ho ⋅ CUNXI YU

Long-horizon RTL agents often collapse under context pollution, lacking the active, compartmentalized debugging strategies of human engineers. We introduce OneVeriAgent, a human-like framework modeled directly on expert cognitive workflows. First, mimicking how engineers probe waveform viewers, Agentic Temporal Exploration (ATE) enables dynamic, targeted signal queries, distilling dense simulator dumps into sparse, causal traces. Second, reflecting how designers compartmentalize tasks, a task-scoped Orchestrator enforces strict context isolation, preventing raw tool outputs from polluting the global reasoning trace. On the CVDP benchmark, OneVeriAgent achieves 95.8\% Pass@1 with Claude Opus 4.6, matching state-of-the-art proprietary swarm. Finally, we propose Tool-Validated Self-Distillation (TVSD) to bridge the open-source capability gap. By fine-tuning on self-generated, syntactically valid trajectories without external correctness oracles, TVSD lifts Gemma4-31B-Instruct's Pass@1 from 61.1\% to 74.8\%, demonstrating that fully open-weight RTL agents are viable.


Submodular Clustering beyond $1-1/e$

Kiarash Banihashem ⋅ Mohammadhossein Bateni ⋅ Hossein Esfandiari ⋅ Samira Goudarzi ⋅ MohammadTaghi Hajiaghayi

Submodular clustering provides a flexible framework for modeling coverage and representation quality in metric spaces, capturing a broad range of objectives arising in machine learning, including influence maximization, data summarization, and representation learning. Given a set of points $P$ in a metric space and a decreasing service function $\phi$, the goal is to select $k$ centers $U$ to maximize total service $\sum_{p \in P} \phi(d(U, p))$, where $d(U, p)$ is the distance from $p$ to its nearest center. While this objective is submodular and admits a standard $1 - 1/e$ approximation via greedy algorithms, it has remained unclear whether and when this barrier can be surpassed. We resolve this question by characterizing exactly which functions $\phi$ permit approximation guarantees beyond $1-1/e$. We identify a natural parameter $\gamma_\phi$ capturing how rapidly $\phi$ decreases, and show that, assuming $\mathrm{P} \ne \mathrm{NP}$, a better-than $1-1/e$ approximation ratio is achievable in polynomial time if and only if $\gamma_\phi > 0$.

Multi-agent reinforcement learning (MARL) holds great potential but faces robustness challenges due to environmental uncertainty. To address this, distributionally robust Markov games (RMGs) optimize worst-case performance when the environment deviates from the nominal model within a uncertainty set. Beyond robustness, an equally urgent goal for MARL is data efficiency—sampling from vast state and action spaces that grow exponentially with the number of agents potentially leads to the curse of multiagency. However, current provably data-efficient algorithms for RMGs are limited to tabular settings with finite state and action spaces, which are only computationally manageable for small-scale problems, leaving RMGs with large-scale (or infinite) state spaces largely unexplored. The only existing work beyond tabular settings focuses on linear function approximation (LFA) for a restrictive class of RMGs using vanish minimal value assumption and still suffers from sample complexity with the curse of multiagency. In this work, we focuses on general RMGs with LFA. For uncertainty sets defined by total variation distance, we develop provably data-efficient algorithms that break the curse of multiagency in both the generative model setting and a newly proposed online interactive setting. To our knowledge, our results are the first to break the curse of multiagency of sample complexity for RMGs with large (possibly infinite) state spaces, regardless of the uncertainty set construction.


Targeted Review for AI-Assisted Biodiversity Surveys: Active Continuous-Score Occupancy Modeling

Timm Haucke ⋅ Lauren Harrell ⋅ Justin Kay ⋅ Mary K Clapp ⋅ Sara Beery

We increasingly use machine learning to label scientific datasets. The models we develop and deploy are improving all the time, but they are not and will likely never be perfect. Mistakes matter, as errors can propagate into our scientific understanding, particularly when systematically biased. Very reasonably, scientists thus review substantial proportions of ML-generated labels to verify or correct mistakes in pursuit of ensuring their scientific findings are not biased by ML. In this work, we focus on helping scientists optimally allocate this reviewing effort relative to their scientific goals. We focus on a specific class of scientists (ecologists) and a specific, widespread, and impactful modeling target (occupancy modeling, which estimates where species are likely to occur, conditioned on environmental factors). We introduce ACORN (Active Continuous-Score Occupancy Modeling), a method that incorporates ML predictions into occupancy models and strategically selects samples for expert review that are maximally informative for downstream ecological analysis. Across camera-trap and bioacoustic datasets, our method recovers ecological conclusions close to those obtained from fully human-labeled data, while requiring substantially fewer expert reviews than non-targeted review policies. Our results suggest that ML-assisted scientific workflows should optimize expert effort for downstream inference, rather than for classifier accuracy alone, especially when human review budget is limited.


Temporal Gradient Inversion for Private Trajectory Reconstruction in Embodied Reinforcement Learning

Sudip Bhujel ⋅ Shanghao Shi ⋅ Ruiquan Huang ⋅ Ning Zhang ⋅ Yang Xiao

Distributed learning in embodied reinforcement-learning agents offers a degree of privacy by retaining raw sensor data on-device and transmitting only policy gradients to the server. Yet temporal structure can amplify this leakage beyond single-frame attacks. We introduce **T**emporal **R**econstruction **A**ttack on **C**onsecutive **E**ncodings (TRACE), an amortized temporal gradient-inversion attack that autoregressively reconstructs the sequence of private observation-action trajectories from per-step policy-learning gradients. The attack exploits two structural signals ignored by prior single-frame methods: (i) cross-time correlation between successive embodied gradients, which we formalize via a conditional mutual-information bound, and (ii) closed-form action recovery from policy-head gradient structure, which we prove exact when standard entropy regularization is sufficiently small. On held-out embodied scenes, TRACE reaches $18.8$ dB PSNR with near-perfect action recovery at $3$-$4.5$ ms per reconstructed frame, dominating the learning-based baseline across all reconstruction metrics and exceeding optimization attacks while running orders of magnitude faster. Defense experiments suggest that protecting temporal gradient streams may require sequence-aware privacy mechanisms.


Test-Time Speculation

Avinash Kumar ⋅ Sujay Sanghavi ⋅ Poulami Das

Speculative decoding accelerates LLM inference by using a fast draft model to generate tokens and a more accurate target model to verify them. Its performance depends on the acceptance length, or number of draft tokens accepted by the target. Our studies show that the acceptance length of even state-of-the-art speculators, like DFlash, EAGLE-3 and PARD degrade with generation length, reaching values close to 1 (i.e. no speedup) within just a few thousand output tokens, making speculators ineffective for long-response tasks. Acceptance lengths decline because most speculators are trained offline on short sequences, but are forced to match the target model on much longer outputs at inference, well beyond their training distribution. To address this issue, we propose Test-Time Speculation (TTS), an online distillation approach that continuously adapts the speculator at test-time. TTS leverages the key insight that the token verification step already invokes the target model for each draft token, providing the training signal needed to adapt the draft at no additional cost. Treating the draft as the student and the target as a teacher, TTS adjusts the draft over several speculation rounds, with each update improving the draft's accuracy as generation proceeds. Our results across multiple models from the Qwen-3, Qwen-3.5, and Llama3.1 families show that TTS improves acceptance lengths over state-of-the-art speculators by up to 72% and 41% on average, with the benefits scaling with increased generation lengths.

Knowledge editing aims to rewrite sensitive information in language models, yet prior work has shown that edited models remain vulnerable to extraction attacks. In this work, we identify a previously overlooked vulnerability, which we term $\textit{rank leakage}$: although editing replaces the top-1 output with a sanitized answer, the original sensitive token often remains highly ranked in the model’s next-token distribution. Exploiting this phenomenon, we propose an effective black-box Output Attack that recovers sensitive information using only a single query. Despite its simplicity, the attack achieves strong extraction performance and remains effective even against models equipped with state-of-the-art defenses. To mitigate this vulnerability, we introduce $\textbf{R}$ank-guided $\textbf{C}$oncealment $\textbf{E}$diting (RCE), a plug-in defense that enforces explicit control over the rank of sensitive tokens during editing. Although designed for the Output Attack, RCE generalizes to existing black-box and white-box extraction attacks. Extensive experiments across multiple models, datasets, and editing algorithms show that RCE consistently improves resistance to attacks while preserving edit fidelity and model utility. Our results highlight the importance of controlling the rank of sensitive tokens, rather than only the top-1 prediction, for privacy-preserving knowledge editing.


The Query Complexity of Local Search in Rounds on General Graphs

Simina Branzei ⋅ Ioannis Panageas ⋅ Dimitris Paparas

We analyze the query complexity of finding a local minimum in $t$ rounds on general graphs. More precisely, given a graph $G = (V,E)$ and oracle access to an unknown function $f : V \to \mathbb{R}$, the goal is to find a local minimum---a vertex $v$ such that $f(v) \leq f(u)$ for all $(u,v) \in E$---using at most $t$ rounds of interaction with the oracle. The query complexity is well understood on grids, but much less is known beyond. This abstract problem captures many optimization tasks, such as finding a local minimum of a loss function during neural network training. For each graph with $n \geq 2$ vertices and constant integer $t \geq 2$, we prove a deterministic upper bound of $O(n^{1/t} (s\Delta)^{1-1/t})$, where $s$ is the separation number and $\Delta$ is the maximum degree of the graph. We complement this result with a randomized lower bound of $\Omega(n^{1/t})$ that holds for any connected graph. We also find that parallel steepest descent with a warm start provides improved bounds for graphs with high separation number and bounded degree. To obtain our results, we utilized an advanced version of Gemini at various stages of our research. We discuss our experience in the Methodology section.


The Sharp Directions Are Against You: Curvature Analysis of Activation Steering

Seyedarmin Azizi ⋅ Arya Fayyazi ⋅ Parsa Razmara ⋅ Erfan Baghaei Potraghloo ⋅ Souvik Kundu ⋅ Massoud Pedram

Activation steering adds a vector to a language model's internal representations at inference time. It is a lightweight alternative to fine-tuning for behavioral control. The standard construction, _contrastive activation addition_ (CAA), is fragile, as on many inputs it shifts behavior in the wrong direction. To explain this fragility, in this paper we first study the geometry of the steering vector. Specifically, we introduce **steering Hessian**, _the matrix of second order derivative of the model's behavioral loss with respect to the steering vector_. It captures the output behavior response sensitivity of the model under small perturbations of the steering vector along each direction. Its top eigenvectors are _sharp directions_: tiny perturbations along them can cause large, erratic, input-dependent behavioral swings. On the other hand, bottom eigenvectors are _flat directions_: behavior is robust along them. We identify that removing the sharp component of a steering vector approximately preserves its direction and magnitude (cosine $> 0.974$, norm reduction $< 3\%$ for typical settings) while reducing the vector's directional curvature by up to 12$\times$ (median 4$\times$ across 36 conditions). We evaluate on three model families (Qwen2.5-7B, Llama-3.1-8B, Mistral-7B), with four behavioral tasks, and 36 token-level conditions. Steering along sharp directions reverses the intended behavior in every condition tested, with anti-steer rates of 53--100\%. Sharp directions account for only 1--5\% of the CAA vector's squared norm. However, they drive a disproportionate share of the gap between the steered and unsteered next-token distributions during autoregressive generation. We term this one-line correction as _sharp removal_; it outperforms standard CAA in open-ended generation across all 25 conditions tested. A jailbreaking experiment further shows that sharp directions act as _noise_ rather than signal: they disrupt coherent behavioral control regardless of steering polarity.


TIDE: Every Layer Knows the Token Beneath the Context

AJAY JAISWAL ⋅ Lauren Hannah ⋅ Han-Byul Kim ⋅ Duc Hoang ⋅ Mehrdad Farajtabar ⋅ Minsik Cho

We revisit a universally accepted but under-examined design choice in every modern LLM: a token index is looked up once at the input embedding layer and then permanently discarded. This single-injection assumption induces two structural failures: (i) the Rare Token Problem, where a Zipf-type distribution of vocabulary causes rare-token embeddings are chronically under-trained due to receiving a fraction of the cumulative gradient signal compared to common tokens; and (ii) the Contextual Collapse Problem, where limited parameters models map distributionally similar tokens to indistinguishable hidden states. As an attempt to address both, we propose TIDE, which augments the standard transformer with EmbeddingMemory: an ensemble of K independent MemoryBlocks that map token indices to context-free semantic vectors, computed once and injected into every layer through a depth-conditioned softmax router with a learnable null bank. We theoretically and empirically establish the benefits of TIDE in addressing the issues associated with single-token identity injection as well as improve performance across multiple language modeling and downstream tasks.


Token Time Continuous Diffusion for Language Modeling

Parikshit Bansal ⋅ Sujay Sanghavi

In this paper we introduce token time continuous diffusion (TTCD), a new diffusion language model which {\em (a)} operates in continuous space, deterministically mapping Gaussian noise to a final token canvas with no further sampling, and crucially {\em (b)} incorporates a new notion of per-token times, with some tokens proceeding from noise to token at a faster rate than others. Continuous space modeling helps TTCD avoid the parallel sampling of multiple tokens, which is a key source of inaccuracy at high speedups for models that iterate purely in discrete space. The notion of per-token times helps TTCD to better model conditional generation, allows for more sure tokens to proceed at a faster rate, and allows for differentiated inter-token influences during refinement. TTCD outperforms discrete models at high speedups. We train a 160M parameter TTCD model on OpenWebText, and then self-distill it; we find that at high speedups we are comparable in unconditional generation quality, and outperform in conditional generation, several existing models of similar size trained, on the same data, and self-distilled. We achieve similar gains in Sudoku solving as well.


Total Variation Rates for Riemannian Flow Matching

Yunrui Guan ⋅ Krishnakumar Balasubramanian ⋅ Shiqian Ma

Riemannian flow matching (RFM) extends flow-based generative modeling to data supported on manifolds by learning a time-dependent tangent vector field whose flow-ODE transports a simple base distribution to the data law. We develop a nonasymptotic Total Variation (TV) convergence analysis for RFM samplers that use a learned vector field together with Euler discretization on manifolds. Our key technical ingredient is a differential inequality governing the evolution of TV between two manifold ODE flows, which expresses the time-derivative of TV through the divergence of the vector-field mismatch and the score of the reference flow; controlling these terms requires establishing new bounds that explicitly account for parallel transport and curvature. Under smoothness assumptions on the population and learned flow-matching field, together with mean-square approximation guarantees for the learned field, we obtain explicit bounds of the form $\mathrm{TV}\le C_{\mathrm{Lip}}\,h + C_{\varepsilon}\,\varepsilon$, cleanly separating numerical discretization and learning errors. Here, $h$ is the step-size and $\varepsilon$ is the target accuracy. Instantiations yield \emph{explicit} polynomial iteration complexities on the hypersphere $S^d$, and on the Hadamard manifolds (for e.g., SPD manifolds) under mild moment conditions.

We present the first theoretical guarantees for differentially private online reinforcement learning (RL) with general function approximation, extending beyond prior work restricted to tabular and linear settings. Our approach combines a batched policy update scheme with the exponential mechanism, together with a novel regret analysis. We show that, even under general function approximation, the regret in the model-free setting under differential privacy matches the state of the art for the linear case, scaling as $\widetilde{O}(K^{3/5})$, where $K$ denotes the number of episodes. As an important by-product, we also establish the first regret bound for online RL with batch update that depends on the standard complexity measure of **coverability**, complementing existing results based on a newly introduced Eluder-Condition class. In addition, we uncover fundamental gaps in recent results for private RL with linear function approximation, thereby clarifying its landscape.


Towards Settling the Complexity of Non-Euclidean Parallel Convex Optimization

Xin Jennifer Chen ⋅ Andrei Graur ⋅ Aaron Sidford ⋅ Chenyi Zhang

In this paper, we provide improved bounds on the complexity of solving $\ell_p$-Lipchitz convex optimization problem in parallel. We lower bound the depth of highly-parallel algorithms, i.e., the number of rounds required by a parallel algorithm making $O(\mathrm{poly}(d))$ queries per round to compute an $\epsilon$-approximate minimizer of a $\ell_p$-Lipschitz, $d$-dimensional convex function over the unit $\ell_p$ ball. We provide a framework for proving lower bounds for parallel $\ell_p$-Lipschitz convex optimization that generalizes prior lower bound proofs and for all $p \in (1,\infty)$ (other than 2) improves the depth for which it can be shown that no super-polylogarithmic improvement over sequential algorithms is possible. We also prove a variety of new lower bounds for $p>2$. Notably, for $p=\infty$, we establish the near optimal parallel complexity of the class of ball-acceleration methods up to depth $\tilde{O}(\sqrt{d})$, which is the first near-optimality result in the parallel convex optimization setting for an algorithm other than mirror descent. Additionally, by providing new algorithms we characterize the optimal depth needed to solve to accuracy $\epsilon = d^{-1/4}$ for all $p \geq 2$ (up to polylogarithmic factors). We also provide extensions of our lower bounds to parallel quantum $\ell_p$-Lipschitz convex optimization.


TRACER: Token ReAssignment for Concept ERasure in Generative Recommendation

Ziheng Chen ⋅ Jiali Cheng ⋅ Zezhong Fan ⋅ Diyuan Wu ⋅ Hadi Amiri ⋅ Gabriele Tolomei ⋅ Yang Zhang

Generative recommendation formulates next-item prediction as autoregressive generation over semantic ID (SID) sequences derived from users' historical interactions, making modern recommender systems structurally similar to large language models (LLMs). As privacy and safety concerns grow, these systems increasingly require concept unlearning to remove sensitive or harmful concepts associated with items. However, existing LLM unlearning methods cannot be directly applied to generative recommendation. Unlike word tokens with explicit semantics, SIDs are abstract identifiers that are often shared by both forget and retain items, leading to severe conflicts between concept removal and recommendation utility preservation. To address this challenge, we propose TRACER, an end-to-end concept unlearning framework based on token reassignment. Rather than directly suppressing shared SIDs, TRACER reassigns concept-related items to alternative tokens that better facilitate forgetting while minimizing side effects on retained items. We further introduce a coherence regularizer to preserve semantic consistency among retain items during unlearning. Experiments on real-world recommendation datasets demonstrate that TRACER effectively removes target concepts while substantially better preserving recommendation utility than existing unlearning baselines.


Training Generalizable Collaborative Agents via Strategic Risk Aversion

Chengrui Qu ⋅ Yizhou Zhang ⋅ Nicolas Lanzetti ⋅ Eric Mazumdar

Many emerging agentic paradigms require agents to collaborate with one another (or people) to achieve shared goals. However, existing approaches to learning policies for such collaborative problems produce brittle solutions that fail when paired with new partners. We attribute these failures to a combination of free-riding during training and a lack of strategic robustness. To address these problems, we study the concept of strategic risk aversion and interpret it as a principled inductive bias for generalizable cooperation with unseen partners. While strategically risk-averse players are robust to deviations in their partner's behavior by design, we show that, in collaborative games, they also (1) can have better equilibrium outcomes than those at classical game-theoretic concepts like Nash, and (2) exhibit less or no free-riding. Inspired by these insights, we develop a multi-agent reinforcement learning (MARL) algorithm that integrates strategic risk aversion into standard policy optimization methods. Our empirical results across collaborative benchmarks (including an LLM collaboration task) validate our theory and demonstrate that our approach consistently achieves reliable collaboration with heterogeneous and previously unseen partners across collaborative tasks.


Training Transformers for KV-Cache Compressibility

Yoav Gelberg ⋅ Yam Eitan ⋅ Michael Bronstein ⋅ Yarin Gal ⋅ Haggai Maron

Long-context language modeling is increasingly constrained by the Key–Value (KV) cache, whose memory and decode-time access costs scale linearly with the prefix length. This bottleneck has motivated a range of context-compression methods, from token-level summarization to recent optimization-based KV compression methods. These post-hoc methods operate on the KV cache of a fixed pretrained model, so their effectiveness is fundamentally limited by how well the model's internal representations can be compressed. In this work, we formalize the notion of KV compressibility and show that it is a property of the learned representations, rather than of the context alone. We prove that almost any sequence-to-vector function admits both highly compressible and inherently non-compressible transformer implementations, highlighting the need to guide transformers toward compressible representations during training. Motivated by this, we propose KV-Compression Aware Training (KV-CAT), a continued pretraining procedure that incentivizes the emergence of compressible representations. We introduce a train-time KV sparsification policy that masks KV slots during training. This forces the model to use fewer KV slots and encourages it to learn representations amenable to post-hoc compression. Empirically, we show that KV-CAT improves the quality–budget tradeoff of downstream compression methods across retrieval, long-context question answering, and perplexity-based evaluation of compressed-prefix continuation.

Optical Music Recognition (OMR), the task of transcribing sheet music into a structured textual representation, is currently bottlenecked by a lack of large-scale, annotated datasets of real scans. This forces models to rely on either few-shot transfer or synthetic training pipelines that remain overly simplistic. A secondary challenge is encoding non-uniqueness: in the popular Humdrum **kern format for transcribing music, multiple different text encodings can render into the same visual sheet music. This one-to-many mapping creates a harder learning task and introduces high uncertainty during decoding. We propose Transcoda, an OMR system built on (i) an advanced synthetic data generation pipeline, (ii) a normalization of the **kern encoding to enforce a unique normal form and (iii) grammar-based decoding to ensure the syntactic correctness of the output. This approach allows us to train a compact 59M-parameter model in just 6 hours on a single GPU that outperforms billion-parameter baselines. Transcoda achieves the best score among state of the art baselines on a newly curated benchmark of synthetically rendered scores at 18.46% OMR-NED (compared to 43.91% for the next-best system, Legato) and reduces the error rate on historical Polish scans to 63.97% OMR-NED (down from 80.16% for SMT++). We will make our model, data and generation pipeline publicly available upon publication of the paper.


TraXion: Rethinking Pre-training Frameworks for Mobility and Beyond

Shang-Ling Hsu ⋅ Mark Tenzer ⋅ Cyrus Shahabi ⋅ Khurram Shafique

Human mobility differs from text and from generic time series in three structural ways: visits are tuple-valued events whose meaning depends on the joint distribution over location, time, and activity; users carry persistent signatures across trajectories; and visits are not independent across users, since co-location at shared places is a primary signal. Existing pre-training recipes for mobility import objectives from language modeling, treating trajectories as sentences and visits as tokens, an analogy that fails against each of the three properties above. These properties define a broader class, multi-entity spatiotemporal event streams (MESES), spanning enterprise authentication logs, electronic health records, and other event-stream domains where entities share infrastructure, schedules, or contexts. We make the properties precise as three axioms that any pre-training framework for MESES should satisfy, and introduce TraXion, whose objectives and architecture are jointly designed to meet them. A single TraXion checkpoint per dataset beats task-specific baselines on every task across six public mobility datasets covering anomaly detection, next-POI recommendation, next-visit prediction, and social-link prediction. The same recipe, applied unchanged to enterprise authentication logs and ICU mortality prediction, matches or exceeds prior work on both, showing that event streams from domains as different as mobility, security, and healthcare can be modeled under a single framework.


Trellis4D: Complex 4D Mesh Generation

Jiraphon Yenphraphai ⋅ Jianqi Chen ⋅ Jian Wang ⋅ Guocheng Qian ⋅ Sergey Tulyakov ⋅ Rameen Abdal ⋅ Raymond A. Yeh ⋅ Peter Wonka ⋅ Chaoyang Wang

Current video-to-4D methods struggle with complex topology changes, transparent materials, thin structures, and inner surfaces. We present Trellis4D, a dynamic mesh generation framework by inheriting the expressive representation of Trellis2, adapting it from image-to-3D to video-conditioned 4D generation. Our design arises from two key questions: (a) how to enable Trellis2's frame-local attention to share information across frames while preserving its pretrained quality on rare cases such as transparent objects and inner surfaces, and (b) how to inject temporal information into a purely 3D positional encoding without breaking pretrained capabilities. We address (a) with a sliding-window cross-frame attention and anchor on the first frame. The first frame is generated by the base Trellis2 model and injected into our model, letting it inherit Trellis2's quality in rare cases through cross-frame attention. We address (b) with a 4D temporal encoding that repurposes redundant low-frequency spatial RoPE bands for time, extending the encoding from 3D with no additional parameters. Extensive experiments show the effectiveness of Trellis4D for high-quality dynamic mesh generation on ActionBench and our own challenging complex dynamics set.


Truthful Calibration Errors for Multi-Class Prediction

Yuxuan Lu ⋅ Yifan Wu ⋅ Jason Hartline ⋅ Lunjia Hu

Calibrated predictions are useful because their numerical values can be interpreted as probabilities. Calibration errors are therefore widely used to evaluate and compare probabilistic predictors. Recently, Haghtalab et al. [2024] introduced truthfulness as an additional requirement for such measures. A calibration measure is truthful if a predictor minimizes its expected measured error by reporting the true conditional label distribution. Many standard empirical calibration errors are non-truthful: a predictor may appear better calibrated by distorting its probabilities. We study the implications of truthful and non-truthful calibration errors for practice. First, we introduce perfectly truthful calibration errors for full multiclass calibration and classwise calibration, two standard multiclass notions. More generally, our construction applies to any linear property of the label distribution, generalizing the truthful calibration error for binary predictions in Hartline et al. [2025]. We also identify a truthful correction for confidence calibration. Second, we characterize the decision-theoretic implications of these truthful errors. For calibrated predictors, truthful calibration errors preserve the Blackwell dominance: a more informative calibrated predictor receives no larger expected error. Third, we show that this decision-theoretic interpretation explains and mitigates the well-observed ranking robustness problem of binned calibration errors. Empirically, non-truthful confidence-based errors can reverse model rankings when the number of bins changes, while our truthful classwise error gives more stable rankings across binning choices.


Two-Sided Learning in Matching Markets with Interviews

Amirmahdi Mirfakhar ⋅ Xuchuang Wang ⋅ Mengfan Xu ⋅ Hedyeh Beyhaghi ⋅ Mohammad Hajiesmaili

Two-sided matching platforms rely on preferences from both sides, yet participants can evaluate only a small fraction of potential partners. In practice, they use low-cost pre-match screening, e.g., interviews, profile views, or trial tasks, to form noisy impressions before committing to applications and offers. We study bandit learning in matching markets with interviews, modeling these interactions as queried *hints* [Bhaskara et al., 2022] that reveal partial preference information to both sides while constraining subsequent applications. Our framework also allows firm-side uncertainty: firms, like agents, learn their preferences and may make early hiring mistakes. To address this, we introduce strategic deferral, a firm-side action that permits temporary vacancy, corrects premature commitments, and enables decentralized learning under coarse anonymous feedback. We design algorithms for centralized and decentralized markets and show that a constant number of interviews per round suffices for horizon-independent regret, improving over the $O(\log T)$ guarantees known without interviews. Our bounds are near-optimal: the centralized guarantee is within a factor $m$ of an information-theoretic lower bound, while decentralized algorithms match it up to polynomial factors in structured markets and remain horizon-independent in general markets.


Unified High-Probability Analysis of Stochastic Variance-Reduced Estimation

Zhankun Luo ⋅ Antesh Upadhyay ⋅ M. B Sahin ⋅ Sang Bin Moon ⋅ Anuran Makur ⋅ Abolfazl Hashemi

Stochastic estimators are fundamental to large-scale optimization, where population quantities must be inferred from noisy oracle observations. Although influential methods such as momentum, SPIDER, STORM, and PAGE have been highly successful, their analyses are largely estimator-specific and expectation-based, obscuring the structural tradeoffs that determine reliability. In this paper, we develop a unified framework for stochastic variance-reduced estimation based on a recursion with three components: memory retention, reset probability, and a correction term for iterate movement. This framework recovers several classical estimators, motivates new second-order variants, and yields a bias-variance decomposition of estimation error. Our main result is a unified high-probability bound proved using a new dimension-free vector-valued Freedman inequality, valid for smooth normed spaces involving random sums of vector martingales. The result applies in both Euclidean and non-Euclidean settings, including the analysis of mirror-descent-based methods in Banach spaces. As applications, we obtain high-probability oracle complexities for unconstrained optimization with mirror descent, establishing the logarithmic dependence on the confidence level. We also derive the first $\tilde{\mathcal{O}}(\varepsilon^{-3})$ oracle-complexity bounds for stochastic optimization with expectation constraints, improving upon the existing $\tilde{\mathcal{O}}(\varepsilon^{-4})$ complexity by leveraging variance-reduced estimation for the first time in this setting.


Unifying Reasoning and Planning through Energy Minimization

Adrian Rodriguez ⋅ Angelica Kim ⋅ Yunhui Guo ⋅ Yilun Du

Generalization to harder reasoning and planning problems remains a central challenge for AI systems. Although both domains can be viewed as search problems where valid solutions are easier to verify than generate, existing methods often rely on domain-specific solvers or diffusion-based energy models tailored to individual settings. We introduce Energy Minimization Search} (EMS), a unified framework for reasoning and planning that learns a single time-invariant energy landscape. Unlike diffusion-based energy-based models (EBMs), which perform annealed inference over time-dependent landscapes, EMS performs direct gradient-based optimization on one static landscape using the same training objective and inference procedure across domains. Across reasoning tasks (graph coloring, N-Queens, 3-SAT) and planning tasks (path finding and maze solving), EMS matches or outperforms diffusion-based EBMs, domain-specific solvers, and diffusion-based planners while using substantially less inference compute.


VASR: Variance-Aware Systematic Resampling for Diffusion Models

Shivanshu Shekhar ⋅ Sagnik Mukherjee ⋅ Jia Y Zhang ⋅ Tong Zhang

Sequential Monte Carlo (SMC) samplers for reward-guided diffusion models often suffer from rapid lineage collapse: a few high-reward particles dominate the population within a handful of resampling steps, destroying diversity and degrading sample quality. We propose a variance-decomposition framework for reward-guided diffusion SMC that separates continuation variance $V_t^{\mathrm{cont}}$ from residual variance $V_t^{\mathrm{res}}$, revealing that high offspring-count variance under the commonly used multinomial resampling drives this collapse. This motivates \textsc{VASR} (Variance-Aware Systematic Resampling), which addresses both variance terms via variance-optimal mass allocation $m_t \propto w_t e^{r_t}$ (minimizing $V_t^{\mathrm{cont}}$) and systematic resampling (controlling $V_t^{\mathrm{res}}$). For latent diffusion models where intermediate rewards are noisy due to stochastic continuations, we propose \textsc{VASR-Max}, a deliberately biased high-selection variant for variance-sensitive reward optimization. Both methods are training-free, fully parallelizable, and add only linear overhead. On MNIST and CIFAR-10, \textsc{VASR} achieves as high as $26\%$ better FID than prior SMC methods while remaining $\sim\!66\times$ faster than MCTS-based value methods at matched compute. On text-to-image generation, \textsc{VASR-Max} consistently outperforms the strongest SMC baseline across compute budgets and matches MCTS-based methods within $2.5$--$3\%$ reward at high budgets while being approximately $4\times$ faster.


Verifying Neural Networks with Reinforcement Learning

Hai Duong ⋅ Thanh Le ⋅ ThanhVu Nguyen

Formal verification can play a key role in ensuring the reliability of Deep Neural Networks (DNNs) deployed in safety-critical systems. Modern DNN verifiers employ a branch-and-bound framework, which alternates between branching (splitting into smaller subproblems) and bounding (pruning subproblems) to efficiently explore the verification space. However, existing branching heuristics make greedy decisions based on static scoring functions. They do not anticipate long-term efficiency or leverage the growing availability of verification data to improve performance. This work introduces RSB, a reinforcement learning framework that learns to refine baseline branching heuristics. It trains an actor-critic architecture to maximize cumulative future rewards rather than immediate scores. The actor generates attention weights from observations of raw neuron features and learned graph embeddings, which rescale baseline heuristic scores to guide neuron branching. Evaluation on 600 challenging instances demonstrates that RSB consistently outperforms state-of-the-art branching heuristics, solving 11\% more instances while reducing branch exploration by 50\%.


Viverra: Text-to-Code with Guarantees

Haoze Wu ⋅ Rocky Klopfenstein ⋅ Keith Farkas ⋅ Nina Narodytska

A fundamental limitation of Text-to-Code is that no guarantee can be obtained about the correctness of the generated code. Therefore, to ensure its correctness, the generated code still has to be reviewed, tested, and maintained by developers. However, parsing through LLM-generated code can be tedious and time-consuming, potentially negating the productivity gains promised by AI-coding tools. To address this challenge, we present Viverra, a system that automatically produces formally verified annotations alongside generated code to aid user's understanding of the generated program. Given a natural-language task description, Viverra prompts an LLM to synthesize a C program together with candidate assertions expressing safety and correctness properties. It then verifies those assertions in a compositional and best-effort manner via a portfolio of bounded model checkers. Evaluation on 18 diverse programming tasks suggests that Viverra can efficiently generate code with verified assertions, and that these assertions improve users' performance on code-comprehension tasks in a user study with more than 400 participants.


VLS: A Vision-Language-Shape Model for Open-Vocabulary Partonomic 3D Reconstruction

Xiaoqian Ruan ⋅ Pei Yu ⋅ Dian Jia ⋅ Hyeonjeong Park ⋅ Peixi Xiong ⋅ Wei Tang

Single-view 3D reconstruction has advanced rapidly in recent years, but existing research primarily focuses on recovering whole-object geometry while ignoring their semantic parts, which are crucial for fine-grained 3D perception and downstream applications. This paper introduces the task of open-vocabulary partonomic reconstruction. Given a single image and a set of part names, we aim to reconstruct both the object’s overall shape and its constituent semantic parts, even when the object and part categories are unseen during training. To address this task, we propose a vision-language-shape (VLS) model that unifies vision, language, and shape representations within a shared neural field and grounds continuous 3D coordinates to 2D pixel contexts via a deformable implicit function. It can effectively reconstruct any-topology shapes and open-vocabulary semantic parts from a single image. To train VLS with limited 3D part-labeled data, we propose an omni-supervised learning framework that leverages heterogeneous datasets with different levels of annotation and open-world knowledge from existing vision-language models. Extensive experiments on ShapeNetPart, PartNet, and Objaverse demonstrate the effectiveness and strong generalization ability of VLS. We will release the code and data publicly.


VLSplat: Vision-Language Guided Object-Centric 3D Gaussian Splatting via Scene Graph

Yuntae Jeon ⋅ Younho Jeon ⋅ Sujin Jin ⋅ Sungho Jo ⋅ Seunghee Park

Learning semantic representations for 3D Gaussian Splatting has recently emerged as a promising direction for 3D scene understanding. However, existing 2D-to-3D semantic lifting approaches rely primarily on 2D segmentation supervision and lack object-level priors, causing the learned representation to be driven by local mask evidence rather than holistic object structure. In this paper, we present VLSplat, a vision-language guided approach for object-centric semantic lifting in 3D Gaussian Splatting. Our method augments mask-based lifting with object-level semantic priors derived from vision-language models and refines Gaussian representations using an object-group scene graph. The refinement is performed through two language-guided modules: intra-group refinement suppresses semantic noise within object groups, while inter-group acquisition expands object support regions by incorporating structurally relevant candidates from neighboring groups. Extensive experiments on indoor scene datasets demonstrate that VLSplat improves semantic consistency and structural completeness. Our code and models are available at https://vlsplat.github.io/


Watermarking as a Learned Intrinsic Property of Diffusion Models

Xiang Li ⋅ Dong-Dong Wu ⋅ Masashi Sugiyama ⋅ Lannan Luo ⋅ Qiang Zeng

Recent advances in latent diffusion models have enabled high-quality image generation, but also raise critical concerns for intellectual property protection in model distribution scenarios, where downstream users have unrestricted access to models, allowing arbitrary modifications. Existing watermarking methods either rely on inference-time control of inputs (e.g., specific prompts or noise initialization) or embed watermark signals in auxiliary components, making them easily removable in such settings. In this paper, we propose INMARK, which treats watermarking as an intrinsic property learned by the model. Instead of relying on input control or auxiliary watermarking components, INMARK enables the core denoising network to internalize and reproduce watermark patterns. Extensive experiments demonstrate that INMARK achieves strong generation fidelity, high watermark detectability, and robustness against attacks, while remaining fully compatible with standard diffusion training pipelines. Our results highlight a new perspective on diffusion model watermarking: the denoising network can learn a reliable and persistent watermarking capability, which is crucial in practical model distribution scenarios.


Wavefunction Flows: Efficient Quantum Simulation of Continuous Flow Models

David Layden ⋅ Ryan Sweke ⋅ Vojtech Havlicek ⋅ Anirban Chowdhury ⋅ Kirill Neklyudov

Continuous flow models transform Gaussian noise into samples from a learned distribution that closely approximates a complex data distribution. We present a natural mapping between these models and a Schrödinger equation, the fundamental equation of quantum mechanics, whose solution is a quantum state encoding the learned distribution. Our main result is that this Schrödinger equation is efficiently solvable on a quantum computer, which we prove by precisely bounding the discretization error. Therefore, given a trained flow model, the theoretical analysis we introduce implies that future quantum computers will enable a fundamentally different—and potentially more powerful—type of access to its learned distribution, which could be used to perform downstream tasks (e.g., Monte Carlo estimation) more efficiently. More broadly, our results reveal a rare close connection between state-of-the-art generative modeling techniques, such as flow matching and diffusion models, and one of the main expected capabilities of quantum computers: simulating quantum mechanics.


What if Agents Could Imagine? Reinforcing Open-Vocabulary HOI Comprehension through Generation

Zhenlong Yuan ⋅ Yue Wang ⋅ Jing Tang ⋅ Rui Chen ⋅ Kejin Cui ⋅ Lei Sun ⋅ Dapeng Zhang ⋅ Hongwei Yu ⋅ Chengxuan Qian ⋅ Xiangxiang Chu ⋅ Shuo Li ⋅ Yuyin Zhou

Multimodal Large Language Models have shown promising capabilities in bridging visual and textual reasoning, yet their reasoning capabilities in Open-Vocabulary Human-Object Interaction (OV-HOI) are limited by cross-modal hallucinations and limited viewpoints of images. To address this, we propose ImagineAgent, an agentic framework that integrates cognitive mapping, tool-augmented reinforcement learning (RL), and generative world modeling for robust OV-HOI understanding. Specifically, we first propose an innovative CoT dataset named hicodet-6K for supervised fine-tuning (SFT), which effectively bridges the perception-to-cognition gap by structuring perceived entities into interaction pairs for comprehensive predictions. Subsequently, we develop a multimodal tool library integrating online retrieval, image cropping, and generative modeling, enabling the agent to dynamically augment reasoning with domain-specific tools to resolve visual-semantic ambiguities and hallucinations during inference. Moreover, we incorporate a generative model to reconstruct alternative viewpoints, enabling the agent to “imagine” under limited viewpoints. Finally, we propose a composite reward mechanism to jointly optimize prediction accuracy and tool efficiency. Evaluations on both SWIG-HOI and HICO-DET datasets demonstrate that our method achieves state-of-the-art performance while requiring merely 36.7% of the training data compared to existing methods, validating our robustness, empirical effectiveness and efficiency.


What Should Embeddings Embed? Autoregressive Models Represent Latent Generating Distributions

Liyi Zhang ⋅ Michael Li ⋅ R. Thomas McCoy ⋅ Ted Sumers ⋅ Jian-Qiao Zhu ⋅ Tom Griffiths

Autoregressive language models have demonstrated a remarkable ability to extract latent structure from text. The embeddings from large language models have been shown to capture aspects of the syntax and semantics of language. But what should embeddings represent? We show that the embeddings from autoregressive models correspond to predictive sufficient statistics. By identifying settings where the predictive sufficient statistics are interpretable distributions over latent variables, including exchangeable models and latent state models, we show that embeddings of autoregressive models encode these explainable quantities of interest. We conduct empirical probing studies to extract information from transformers about latent generating distributions. Furthermore, we show that these embeddings generalize to out-of-distribution cases, do not exhibit token memorization, and that the information we identify is more easily recovered than other related measures. Next, we extend our analysis of exchangeable models to more realistic scenarios where the predictive sufficient statistic is difficult to identify by focusing on an interpretable subcomponent of language, topics. We show that large language models encode topic mixtures inferred by latent Dirichlet allocation (LDA) in both synthetic datasets and natural corpora.


When Helpfulness Becomes Sycophancy: Sycophancy is a Boundary Failure Between Social Alignment and Epistemic Integrity in Large Language Models

Jiechen Li ⋅ Catherine A Barry ⋅ Rishika Randev ⋅ Janet Chen ⋅ Ella Jorgensen ⋅ Brinnae Bent

This position paper argues that sycophancy in LLMs is a boundary failure between social alignment and epistemic integrity. Existing work often operationalizes sycophancy through external behavior such as agreement with incorrect user beliefs, position reversals, or deviation from an objective standard of correctness. These formulations capture only overt forms of the phenomenon and leave subtler boundary failures involving epistemic integrity and social alignment underspecified. We argue that sycophancy should not be understood as agreement alone, but as alignment behavior that displaces independent epistemic judgment. To clarify this boundary, we propose a three-condition framework for sycophancy. First, the user expresses a cue in the form of a belief, preference, or self-concept. Second, the model shifts toward that cue through alignment behavior. Third, this shift compromises epistemic accuracy, independent reasoning, or appropriate correction. We also introduce a taxonomy for classifying sycophancy, consisting of alignment targets, mechanisms, and severity. The paper concludes by discussing implications for alignment evaluation and argues for boundary-aware assessment, structured rubrics, and mitigation strategies, while situating these proposals alongside alternative views of sycophancy.

Image super-resolution (SR) with large generative models has recently achieved remarkable perceptual quality, yet maintaining fidelity to the LR observation remains challenging. In particular, we observe that diffusion transformers (DiTs) built on latent representations suffer from a critical limitation: the compression bottleneck of the VAE weakens fine-grained spatial information, leading to hallucinated details that are weakly grounded in the input image. In this work, we revisit generative SR from a representation perspective and propose a pixel-grounded super-resolution (PGSR) framework that preserves LR-observed pixel evidence before VAE compression and reuses it throughout restoration. Instead of relying solely on the compressed latent condition, PGSR extracts pre-VAE pixel evidence from the upsampled LR image and reuses it at two stages. First, Condition-Side Trajectory Guidance fuses LR-derived pixel evidence with the latent LR condition to guide the latent restoration trajectory. Second, Decoder-Side Pixel Grounding injects multi-scale pixel features into the frozen VAE decoder to ground the final rendering with LR-observed cues. To efficiently adapt large pretrained DiT models, we keep the latent autoencoder and main flow-matching backbone frozen, and train only lightweight restoration modules. We further study an efficient local-window attention variant for improved high-resolution efficiency and scalability. Extensive experiments demonstrate that PGSR improves the realism--fidelity trade-off and produces more faithful, visually convincing results than existing latent generative SR approaches.


When Safety Becomes An Outlier: Understanding the Retention of LLM Safety Behaviors

Binchi Zhang ⋅ Hadi Abdullah ⋅ Yiwei Cai ⋅ Jundong Li

Safety-aligned large language models often lose their safety behaviors after fine-tuning, even when safety data are included. While this fragility is well documented, its underlying cause remains unclear. We propose that formulaic refusal strings function as outlier modes in the model's output distribution, i.e., low-support behaviors that are easy to fit but structurally unstable under subsequent fine-tuning, and that this is an important source of safety fragility. Controlled experiments verify this: fixed refusals are acquired and forgotten much like arbitrary constant strings, in contrast to prompt-grounded natural responses. Motivated by this insight, we improve safety retention by constructing natural, context-aware safety supervision that keeps safety responses within the model's existing output distribution and grounded in semantics. We instantiate this principle in both supervised and preference-based alignment settings, using perplexity under the original model as a practical proxy for distributional naturalness. Across multiple models and fine-tuning regimes, our approach improves both immediate safety and its retention under continued fine-tuning without sacrificing helpfulness, suggesting that naturalness is a useful principle for durable safety alignment.


When Uncertainty Is the Target: Adversarial Attacks on Uncertainty-Aware Predictors

A Q M Sazzad Sayyed ⋅ Michele Caprio ⋅ Nathaniel D Bastian ⋅ Francesco Restuccia

As machine learning becomes increasingly deployed in high-stakes scenarios, studying adversarial vulnerability of uncertainty-aware models has become a compelling necessity. In this paper, we focus on adversarial attacks that manipulate predictive uncertainty rather than model predictions. We focus on predictors that decompose uncertainty into aleatoric uncertainty (AU) and epistemic uncertainty (EU) and formalize uncertainty vulnerability as the maximum increase in AU achievable under a stealth constraint where the EU is kept below a threshold. This framing yields three complementary findings. First, we theoretically show that knowledge of the baseline EU can enlarge the attacker's feasible stealth region and increase the maximal achievable increase in AU. Second, a first-order analysis reveals that vulnerability is governed by the component of the AU gradient orthogonal to EU, enabling stealthy manipulation. Third, contrary to standard margin-based intuition, vulnerability peaks at moderate margins. Experiments across MNIST, CIFAR-10, and CIFAR-100 with standard architectures show that the proposed attack achieves non-trivial aleatoric damage while maintaining zero epistemic-stealth violation (e.g., mean AU increase of 0.0286, 0.0853, and 0.1568, respectively), whereas unconstrained uncertainty attacks violate the stealth constraint on over 90% of samples. Experiments indicate that the largest damage typically occurs in an intermediate uncertainty regime rather than at the extremes. Furthermore, consistent with our theoretical analysis, feasible attackability collapses under increasing epistemic uncertainty (e.g., mean attackability decreases from 0.0363 to 0). Overall, these results intimately connect uncertainty decomposition to adversarial geometry.


Whole-Body Compliant Control via Learned Force-Regulation Modules

Diego Aldarondo ⋅ Aadhithya Iyer ⋅ Daniel Giebisch ⋅ Nina Mortensen ⋅ Sridhar Pandian Arunachalam ⋅ Robert Cochran ⋅ Josh Merel

To operate safely and effectively around people, humanoid robots must control the forces they exchange with the world. Whereas classical impedance and admittance methods shape contact forces precisely, modern reinforcement-learning controllers typically forgo such regulation in favor of robustness and faithful command tracking. Incorporating impedance-style compliance into a performant humanoid whole-body controller remains an open problem. We present a framework for learned whole-body force regulation that approximates the behavior of impedance controllers using only a PD position-target interface and position encoders. The framework includes three mechanisms for force regulation: passive joint-angle compliance via noisy action perturbations, joint-angle force regulation via perturbation--action blending around a commanded pose, and task-space force regulation via reward-shaped tracking of both a virtual forcefield attractor and a commanded pose. All three share a residual actor--critic recipe, an internal model of proprioception and perturbations, and a policy-blending procedure that combines multiple experts. A 6-DOF body command, optional upper-body pose, and controllable compliance expose the controllers as reusable low-level modules that enable compliant interaction and object manipulation across standing, crouching, and walking. We demonstrate whole-body teleoperation in simulation and on Sprout, a 27-DOF bipedal humanoid.


Who, Where, and What? Forensic Localization in LLM-Based Multi-Agent Systems

ANJUN GAO ⋅ Yueyang Quan ⋅ Yufei Xia ⋅ Zhuqing Liu ⋅ Minghong Fang

LLM-based multi-agent systems extend single agents with role specialization and inter-agent communication, but they also expand the attack surface: an attacker can compromise an externally writable module of one agent, and the malicious payload may propagate across communication hops before some downstream agent invokes an attacker-chosen tool. Existing defenses are routinely circumvented by stronger or adaptive attacks, which motivates a complementary forensic capability that traces an incorrect tool invocation back to its root cause. We formulate this problem as three-level forensic localization, where the goal is to jointly identify the source agent whose internal state was first compromised, the specific module of that agent that was exploited, and the exact units inside that module carrying the malicious payload. To enable systematic evaluation, we construct MAFL-Bench, the first forensic localization benchmark for LLM-based multi-agent systems, which spans three application scenarios, four communication structures, and twelve attack instances, with 21,200 interaction logs annotated at the agent, module, and unit levels. We further propose MATracer, a black-box forensic framework that resolves the three-level attribution through a coarse-to-fine cascade driven by log-likelihood queries on a proxy LLM, progressively localizing the source agent, the compromised module, and the contaminated units. Extensive experiments show that MATracer accurately localizes the attack across diverse multi-agent settings, consistently outperforms ten baselines, and remains robust under adaptive attacks.


Witness Overlap: Directional Provenance Inside Open-Weight Model Families

Siyuan Li ⋅ Haoxuan Zeng ⋅ Xin Luo ⋅ Fernando Jia ⋅ Florence Li ⋅ Zhengyang Geng ⋅ Zico Kolter ⋅ Tai Sing Lee ⋅ Tianqin Li

Open-weight models are often released, fine-tuned, aligned, merged, and re-released, making provenance audits ask not only whether checkpoints are related, but also which checkpoint came first. Many existing model-provenance methods are designed for a base-known audit setting: given a victim or source model, they test whether a suspect model is related to it. Although these audits are framed as source-to-suspect tests, their underlying evidence is often symmetric, relying on representation similarity, weight similarity, behavioral fingerprints, or correlation statistics. Symmetric pairwise comparisons can detect relatedness, but they cannot by themselves orient A versus B. We therefore introduce a local geometric comparison: instead of comparing two checkpoints directly, we add a third same-family checkpoint as a witness and compare the geometry around A with the geometry around B. Direction is inferred by asking which candidate behaves more like a branching parent. Motivated by this idea, and by the empirically observed asymmetry between parent-anchored and child-anchored witness-overlap distributions, we propose Witness Overlap, a prompt-free, training-free white-box test for directional provenance. On 176 LLM checkpoints from 16 families, our one-witness test orients 95.3\% of parent–child decisions using Frobenius cosine. We further evaluate root identification, sibling discrimination, generalizations to VLM and diffusion families, and chain-structured ordering. The signal is robust to weight noise and sparse pruning, with a proposed SVD weight reduction variant showing greater robustness than Frobenius cosine.


Youdunit: Single-Call Counterfactual Necessity in Multi-Agent LLM Systems

Marissa Li ⋅ Stephanie Gao ⋅ Kenny Guo ⋅ Xingjian Li ⋅ William Chang ⋅ Goran Radanovic

Agentic systems built on large language models (LLMs) are increasingly deployed in high-stakes settings, but the causal mechanisms by which adversarial agents induce harmful outcomes remain poorly understood. We address this with a single-call counterfactual necessity test for multi-agent LLM systems that asks: \emph{which specific agent communication was causally responsible for the harmful outcome?} In contrast with actual-causality frameworks that reason over set-valued causes, we use one attribution rule throughout: single-call necessity for a selected proximate call. The key technical contribution is the \emph{Gumbel-Max tape}, which lifts the per-token Gumbel-Max counterfactual generator of \cite{chatzi2024counterfactual} from a single LLM to a multi-agent conversation. Per-token GPU random number generator (RNG) states are recorded during a factual run on a tape shared across all agents, then replayed during counterfactual (CF) runs everywhere except at the intervened call. Applied to the BAD-ACTS benchmark across $148$ adversarial scenarios spanning four multi-agent environments and two prompting conditions (\emph{non-safe} and \emph{safe}), the test shows that single-call necessity tracks communication structure: decentralized and hierarchical environments (Travel Planning, Financial Article Writing) exhibit meaningful aggregate causal effect (ACE), while sequential debate (Multi-Agent Debate) shows near-zero ACE despite comparable attack success rate. A replay sanity check confirms that all CF effects under identity intervention are exactly zero, validating the tape mechanism. The results identify where targeted defenses are likely to be effective in deployed agentic systems.