Session
Sydney Poster Session 6
Hall 1-4
$\mathbf{\mathtt{MAD\text{-}Bench}}$: How Do Multimodal Agents Deceive You?
Hao Gu ⋅ Mingli Song ⋅ Jiacong Hu
Multimodal agents increasingly operate computers on behalf of users. This shift from passive question answering to interactive execution creates a new safety problem. When an agent encounters normal, non-attack scenarios where the task it needs to perform is impossible or difficult to complete, the central risk is not only that it fails, but that it deceives the user into believing that it has succeeded. Although prior work has studied truthfulness and deception in language models and language agents, multimodal interaction introduces new sources of deceptive behavior, increasing the implicitness and complexity of deceptive behaviors in multimodal agents. To fill this gap, we introduce $\mathbf{\mathtt{MAD\text{-}Bench}}$, a benchmark for evaluating **M**ultimodal **A**gent **D**eception. $\mathbf{\mathtt{MAD\text{-}Bench}}$ contains 360 sandbox tasks organized around three core elements of multimodal execution: *modality*, *target*, and *tool*, and spans six task types including *modality evidence conflict*, *modality asynchronous mismatch*, *modality ambiguity distortion*, *modality object missing*, *target infeasibility*, and *tool defect*. To our knowledge, $\mathbf{\mathtt{MAD\text{-}Bench}}$ is the first benchmark to formalize deceptive behavior in multimodal agents from an execution-evidence perspective, characterizing deception as a misalignment between the user-facing task-state claim and the multimodal evidence available throughout the execution trajectory. We further propose a behavioral taxonomy covering *evasive deception*, *manipulative deception*, *misleading deception*, and *non-deception*. Evaluating 10 mainstream multimodal agents, we find that agents frequently fabricate, conceal, mislead, frame, and disguise during task execution. These results show that substantial deceptive behaviors are widespread across mainstream model families. This benchmark is available at https://anonymous.4open.science/r/MAD-Bench-A304.
$R^2E$: A Role-driven Reward Evolutionary Framework for Automated Reward Function Design
Shouhao Chang ⋅ Xuan Liu ⋅ Hongye Zhu ⋅ Xinning Chen ⋅ Shigeng Zhang
Designing effective reward functions is a critical challenge in Reinforcement Learning (RL), which traditionally requires costly manual trial-and-error. While Large Language Model (LLM)-based methods have shown promise in reward design, undifferentiated sampling and limited use of historical feedback often lead to insufficient exploration of the reward space and optimization instability. To address these challenges, we propose $R^2E$, a Role-driven Reward Evolutionary framework for automated reward function design that structures the process through explicit role specialization. $R^2E$ employs multiple LLM roles with complementary objectives: an Explorer that promotes novelty to expand the search space, a Guardian that performs conservative refinements to improve training stability, and an Evolver that recombines high-performing rewards via crossover and mutation to integrate effective structures. A dynamic scheduling strategy coordinates these roles, progressively shifting the search from broad exploration to focused exploitation. To mitigate optimization instability and provide a medium for cross-role collaboration, $R^2E$ incorporates a Global Elite Pool that retains the best-performing rewards to guide subsequent generations, thereby enabling coordinated interaction among roles and ensuring a reliable and consistent refinement process. Extensive experiments on multiple robotic tasks demonstrate the effectiveness of our method.
$SE(2)$-Aware Conditional Distribution Transport for Vehicle Trajectory Generation
Di Wen ⋅ Zhaocheng He ⋅ Shuhui Wang
Vehicle trajectory generation is critical for autonomous driving simulation. However, current generative models face two coupled limitations: (i) imbalanced trajectory distributions bias them toward frequent motions, such as straight driving or stationary states, while weakening their ability to synthesize rare but safety critical motions, such as turning. (ii) Euclidean intermediate spaces provide limited geometric inductive bias for modeling vehicle pose evolution. Together, these limitations lead to mode biased generation and physically less coherent trajectories. To address this issue, we propose SEA, an \underline{\textbf{SE}}(2)-\underline{\textbf{A}}ware vehicle trajectory generation method that realizes the conditional transport path through planar rigid body poses. At the \textbf{data distribution} level, SEA learns vehicle trajectory distributions from traffic data to capture individual motion characteristics. At the \textbf{physical-space} level, motivated by the geometry dependence of {\bf optimal transport} interpolation, SEA realizes intermediate transport states through an $SE(2)$-consistent construction, thereby preserving geometric interpretability and physical coherence during generation. Extensive experiments on large scale real world datasets show that our method outperforms state of the art approaches by generating more realistic, diverse and dynamically consistent trajectories.
3DABSeg: Adaptive 3D Ankle Bone Segmentation with Multiscale Feature Fusion Mixture-of-Experts
Yuhan Tang ⋅ Xiaowen Huang ⋅ Tianxing Zhao ⋅ Xiaowen Fu ⋅ Xiaoling Luo ⋅ Yang Zhang ⋅ Hongbo Le ⋅ Xi Zhou ⋅ Linlin Shen
As one of the most complex and load-bearing joints in the human body, the ankle plays a crucial role in locomotion and clinical assessment. Accurate segmentation of ankle bones is essential for trauma evaluation, preoperative planning, and disease diagnosis. However, the scarcity of publicly available ankle datasets has hindered the development of intelligent analysis in this area. To address this gap, we introduce the first two publicly available ankle CT datasets, providing new directions for medical image segmentation. Considering the intricate and spatially correlated anatomy of the ankle, we propose a 3D medical image segmentation model, namely 3DABSeg. Specifically, we introduce the Hybrid KAN-Mamba Block (HKMB) to capture long-range spatial dependencies and complex nonlinearities. By integrating Mamba’s sequence modeling with KAN’s learnable spline functions, HKMB enhances feature expressiveness for complex anatomical structures. Furthermore, to address the uneven semantic distribution and feature dilution across channels, we propose a Multi-Scale Feature Fusion Mixture-of-Experts module (MFF-MoE). MFF-MoE utilizes multi-scale spatial pooling to compute dynamic routing scores, adaptively partitioning channels into specialized expert networks via a mutually exclusive hard-assignment strategy. Extensive experiments demonstrate that 3DABSeg surpasses existing state-of-the-art methods in both segmentation performance and boundary precision, highlighting its potential for complex anatomical structure analysis and clinical applications.
A2Eval: Agentic and Automated Evaluation for Embodied Brain
Shuai Zhang ⋅ Jiayu Hu ⋅ Zijie Chen ⋅ Zeyuan Ding ⋅ Yi Zhang ⋅ Yingji Zhang ⋅ Ziyi Zhou ⋅ Junwei Liao ⋅ Shengjie Zhou ⋅ Yong Dai ⋅ Zhenzhong Lan ⋅ Xiaozhu Ju
Current embodied VLM evaluation relies on static, expert-defined, manually annotated benchmarks that exhibit severe redundancy and coverage imbalance. This labor‑intensive paradigm drains computational and annotation resources, inflates costs, and distorts model rankings, ultimately stifling iterative development. To address this, we propose Agentic Automatic Evaluation (A2Eval), the first agentic framework that automates benchmark curation and evaluation through two collaborative agents. The Data Agent autonomously induces capability dimensions and assembles a balanced, compact evaluation suite, while the Eval Agent synthesizes and validates executable evaluation pipelines, enabling fully autonomous, high-fidelity assessment. Evaluated across 10 benchmarks and 32 models, A2Eval compresses evaluation suites by 85\%, reduces overall computational costs by 77\%, and delivers up to a 4.6$\times$ speedup while preserving evaluation quality. Crucially, A2Eval corrects systematic ranking biases, improves human alignment to Spearman's $\rho=0.85$, and maintains high ranking fidelity (Spearman's $\rho=0.87$), establishing a new standard for high-fidelity, low-cost embodied assessment. Our code and data will be public upon acceptance.
A Biconvex Formulation for Stable Transport of Mixture Models with a Unique Solution
Yeganeh Mohammadali Marghi ⋅ Kelly Jin ⋅ Uygar Sümbül
Optimal transport (OT) provides a principled framework for mapping between probability distributions. Despite extensive progress, applying OT to large-scale data remains computationally demanding, and the resulting pointwise transport plans are often difficult to interpret. We introduce Optimal Mixture Transport (OMT), a scalable framework that shifts the transport paradigm from individual samples to mixtures of subpopulations, reformulating the transport problem as a strictly biconvex optimization with a unique global minimizer. We further establish theoretical guarantees on the stability of the OMT map, showing that bounded perturbations of the underlying distributions lead to bounded changes in the transport plan. By formulating subpopulations as exponential-family distributions, OMT decouples computational complexity from the sample size, scaling solely with the number of mixture components. We demonstrate the effectiveness and practicality of OMT on a wide range of synthetic benchmarks and real-world datasets, including image data and large-scale single-cell RNA sequencing measurements.
A Black-Box Reduction from Regret to Multi-Level Coverage
Tuo Liu ⋅ Edgar Dobriban ⋅ Francesco Orabona
Online Conformal Prediction (OCP) sequentially constructs prediction sets whose average empirical coverage converges to a target level over time, for arbitrary, potentially adversarial datasets. In the multi-level setting of OCP, prediction sets are produced at $d$ coverage levels at once. Prior work showed that online gradient descent with lazy isotonic projections produces sets properly nested across levels, but with a $\sqrt{d}$ dependence in the per-level coverage convergence rate. Here, we both generalize and improve prior work. First, we show that lazy isotonic projections can be safely used for the construction of multi-level online conformal prediction sets from \emph{any} online learning algorithm. Then, through a black-box reduction, we show that the linearized regret of the online algorithm itself automatically gives control over both per-level calibration and distributional consistency of multi-level predictions via Fenchel duality. Finally, by using the Weighted Interval Score, the standard scoring rule for assessing multi-level interval forecasts, we prove that these projections reduce both the size of the sets and their under-coverage. We show two instantiations of our theory, $p$-norm Online Mirror Descent and coordinate-wise Universal-Portfolio, both improving the dependence on the number of levels to $\sqrt{\log d}$ for the multi-OCP problem. Experiments on stock prices, electricity demand, and synthetic streams confirm consistent gains in coverage-width trade-offs over existing baselines.
Accelerating LLM Pre-Training through Flat-Direction Dynamics Enhancement
Shuchen Zhu ⋅ Rizhen Hu ⋅ Mingze Wang ⋅ Mou Sun ⋅ xue wang ⋅ Kun Yuan ⋅ Zaiwen Wen
Pre-training Large Language Models requires immense computational resources, making optimizer efficiency essential. The optimization landscape is highly anisotropic, with loss reduction driven predominantly by progress along flat directions. While matrix-based optimizers such as Muon and SOAP use fine-grained curvature information to outperform AdamW, their updates tend toward isotropy, which is relatively conservative along flat directions yet potentially aggressive along sharp ones. To address this limitation, we first establish a unified Riemannian Ordinary Differential Equation (ODE) framework that clarifies the mechanism of adaptive algorithms: the preconditioner induces a Riemannian geometry that mitigates ill-conditioning, while momentum serves as a Riemannian damping term that promotes convergence. Guided by these insights, we propose LITE, a generalized acceleration strategy that enhances training dynamics by applying larger Hessian damping coefficients and learning rates along flat trajectories. Extensive experiments demonstrate that LITE significantly accelerates both Muon and SOAP across diverse architectures (Dense, MoE), parameter scales (130M to 1.3B), datasets (C4, Pile), and learning-rate schedules (cosine, wsd). Theoretical analysis confirms that LITE facilitates faster convergence along flat directions in anisotropic landscapes, providing a principled approach to efficient LLM pre-training.
Accelerating the Inference Era with AI-Driven, Globally Optimized HW/SW Co-Design
Miria Feng ⋅ Fangzhao Zhang ⋅ Adrian G Lafuente ⋅ Mert Pilanci ⋅ Azalia Mirhoseini
Deep learning inference demand is projected to grow $10{,}000\times$ over the next five years, a trajectory that general-purpose accelerators cannot match. Custom accelerators offer a viable path forward, but designing them requires simultaneous co-optimization of hardware datapaths, software schedules, and compiler transforms across a combinatorial search space exceeding $\mathcal O(10^{2300})$. While co-design is becoming increasingly essential, existing methodologies rely on decoupled, sequential pipelines that miss critical cross-stage trade-offs. We introduce CHEETAH: a globally optimized, AI-driven framework for inference accelerator co-design. To our knowledge, this is the first co-design framework to combine \textit{joint} tiling-fusion, dynamic placement, and rematerialization in a single LP-based formulation. Key contributions include: (1) time-indexed variables for dynamic memory management, (2) rematerialization variables to enable cheap recomputation over costly DRAM reloads, and (3) joint tiling-fusion optimization via McCormick linearization. This inner-loop precision is steered by an LLM-guided outer loop that performs causal interventions based on specific deployment profiles and Lagrangian sensitivity signals. In less than 6 hours, \textsc{CHEETAH} discovers designs on the Pareto frontier achieving an average of $\sim22.07\times$ raw speedup across $11$ workloads, reaching up to $\sim40.1\times$ speedup on a single workload over a TPU-v3 baseline, unlocking efficiency inaccessible to decoupled sequential pipelines.
ACD-GS: Asymmetric Curvature-aware Densification for 3D Gaussian Splatting
Houqiang Zhong ⋅ Tianchi Zhu ⋅ Qiang Hu ⋅ Xiaoyun Zhang ⋅ Li Song ⋅ Wenjun Zhang
3D Gaussian Splatting (3DGS) has emerged as a powerful paradigm for real-time photorealistic novel view synthesis. However, its training-time densification remains largely heuristic, often resulting in redundant Gaussian proliferation and inaccurate geometric reconstruction. Existing methods predominantly rely on first-order screen-space gradients to trigger cloning and splitting, which fail to distinguish true geometric discontinuities from high-frequency textures or noisy residuals. To address this limitation, we propose ACD-GS, an Asymmetric Curvature-aware Densification framework for compact and geometry-aware 3D Gaussian optimization. Our key insight is that second-order depth responses provide a more reliable structural prior for density control, enabling asymmetric regulation of cloning and splitting to better capture geometric discontinuities. Building on this, we introduce a multi-view photometric consistency constraint to suppress transient single-view errors, together with a capacity-aware adaptive regulation strategy that balances densification and pruning during training. Unlike post-hoc compression or quantization-based approaches, ACD-GS directly prevents redundant Gaussian generation during training, yielding inherently compact representations. Extensive experiments demonstrate that our method reduces nearly 60\% of Gaussian count and storage while maintaining competitive rendering fidelity, consistently outperforming existing methods across multiple datasets. Our code is available in https://anonymous.4open.science/r/ACD_GS-0E2C
A Characterization of Latent Variable Causal Models Consistent with Observational Data
Hongshuo Yang ⋅ Adiba Ejaz ⋅ Yushu Pan ⋅ Elias Bareinboim
Causal reasoning from observational data becomes significantly challenging when the underlying causal structure is only partially known. In this work, we study the the equivalence class of structural causal models (SCMs) with latent variables sharing the same set of conditional independencies, as represented by a partial ancestral graph (PAG). Specifically, we provide a new characterization of SCM-induced causal diagrams in this equivalence class using differentiable parameters. Building on this characterization, we introduce a new model class called PAG-constrained neural causal models ($\mathcal{P}$-NCMs) which parameterize this equivalence class of SCMs. We prove that $\mathcal{P}$-NCMs are expressive enough to represent counterfactual distributions induced by any SCM in the class, while remaining consistent with the structural constraints shared across all its members. Finally, we demonstrate how the differentiable parameterization of this model class enables causal inference under Markov equivalence by reducing counterfactual partial identification to an optimization problem over $\mathcal{P}$-NCMs. We establish the theoretical soundness of this approach and validate its performance on simulations.
AdaCal: Adaptive Calibration for Robust Sparse Attention in Long-Context LLMs
Leqin Xiang ⋅ Yongbin Liu ⋅ Chunping Ouyang ⋅ Ying Yu
Large language models face substantial computational challenges in long-sequence inference due to the quadratic cost of self-attention, with the prefill stage being particularly bottlenecked. Static sparse attention reduces compute via predefined masks but can miss task-critical evidence, while dynamic sparsification is more flexible yet can be brittle across tasks, leading to either insufficient coverage or distractor-induced noise. We propose AdaCal, an Adaptive Calibration framework for robust sparse attention that synergizes coarse, experience-driven task dispatch with head-wise physical profiling. By deriving entropy, concentration, and mid-zone coverage metrics from a lightweight probe pattern, AdaCal dynamically governs head activation and augmentation policies, allocating additional computation strictly on demand. This design is realized through task-conditional policies that effectively suppress unnecessary expansion for structure-dominant inputs while enhancing long-range coverage for retrieval-intensive tasks. Experiments show that AdaCal improves robustness over strong static and dynamic baselines, with clearer gains in long-context settings beyond 32K tokens.
Adam under Generalized Smoothness with Second-Moment-Type Stochastic Gradients
ruinan Jin ⋅ Difei Cheng ⋅ Ling Chen ⋅ Jun Luo ⋅ Hao Zhou ⋅ Youzhi Zhang
It has been widely observed that Adam can remain stable even when the objective function deviates significantly from global smoothness. However, under the generalized smoothness framework, existing theoretical analyses typically rely on strong tail assumptions on stochastic gradients, such as almost-sure boundedness or sub-Gaussianity. Whether one can establish the convergence of Adam on generalized smooth objectives under only second moment information on the stochastic gradients, without imposing such strong concentration assumptions, was explicitly identified as an important open direction by \citet{li2023convex}. This paper gives an affirmative answer to this question under fairly general conditions. Specifically, we prove that such strong tail assumptions are not necessary. We show that the key mechanism by which Adam remains stable and achieves convergence under the (L0)--(Lp) generalized smoothness condition is the self-normalization effect induced by its adaptive coordinate-wise scaling. Based on this mechanism, we prove that even under a very general stochastic-gradient condition, namely a generalized second moment ABC condition that provides only second moment information, the stochastic trajectory of Adam remains in a locally well-behaved smoothness region with stretched-exponential tail decay. As a consequence, we establish high-probability convergence rate guarantees over the full range (p<2), with a confidence dependence of order (\delta^{-1/2}), while the required stepsize calibration depends on (\delta) only through polylogarithmic factors in (\polylog(1/\delta)). Furthermore, we construct a hard instance proving that, under only second-moment information on the stochastic gradients, this (\delta^{-1/2})-type confidence dependence is sharp. Finally, in the more favorable regime (p<1), we combine the above trajectory control with polynomial-growth estimates on rare events to further obtain convergence rate guarantees in expectation.
Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security
Devina Jain ⋅ David Hartmann ⋅ Chuan Li
LM-based agents process external content, exposing them to prompt injection and multi-turn manipulation. Most safety benchmarks evaluate defenders against fixed attack pools collected before evaluation, single-turn or multi-turn. We present a 21-scenario benchmark for adaptive multi-round attacks against memoryless LLM defenders: an autonomous LLM attacker observes prior defender responses and pivots across rounds, while each defender response is evaluated as a fresh interaction. Holding the 21 scenarios, attackers, defenders, and structured-output scoring fixed, restricting scoring to the first attacker turn yields $0$--$1%$ attack success rate (ASR); allowing 15 rounds of adaptive attack yields $8.3$--$19.0%$. Pooling three frontier attacker LLMs uncovers up to $2.4\times$ as many unique successful attacks as the best single attacker, and the generated attacks have low cosine similarity ($0.02$--$0.14$) to attacks in existing benchmarks. Claude Opus 4.6 and GPT-5.4 are aggregate-equivalent ($8.3%$ vs. $11.1%$ ASR), but their weaknesses differ sharply: on one scenario Opus reaches $60%$ ASR ($95%$ CI $36$--$80%$) while GPT-5.4 and Gemini each stay at $7%$ (CI $1$--$30%$). $15$ of $21$ scenarios distinguish at least one defender pair, yet rankings disagree across scenarios (Kendall's $W = 0.21$). We release the benchmark---21 evaluation scenarios, 10 public development scenarios, the orchestrator, baseline harnesses, and a multi-attacker CLI---plus 945 transcripts from the $3\times3$ frontier matrix, an attack-replay dataset, and $18{,}422$ battles from an open competition's final scoring rounds.
Adaptive and Neutral Theory of Evolution Strategies for Large Language Models
Soma Yokoi ⋅ Issei Sato
Evolution strategies (ES) have recently emerged as a promising alternative to gradient-based reinforcement learning for fine-tuning large language models, but classical zeroth-order optimization theory does not systematically account for the empirical observations driving this success. We formulate the continuous-time limit of Z-score-normalized ES as a unified stochastic differential equation and identify three sequential dynamical regimes characterized by the dominant gradient contribution to the population reward variance; under a basin-local low-rank Hessian assumption, this framework systematically explains (1) the information-geometric optimality of the Z-score heuristic, (2) the dimensionality paradox of sample-efficient ES with populations far smaller than the parameter dimension, (3) the same-task rise-then-decay of the training reward, and (4) catastrophic forgetting of pre-training capabilities. Synthetic experiments and direct verification on Qwen2.5-{0.5B,3B,7B}-Instruct confirm these predictions, matching the parameter drift reported by Abdi et al. (2026) within $\sim 1.5\\%$.
Adaptive Compression and Targeted Perturbation: A Unified Framework for Generalized Audio Deepfake Detection
Jianqiao Cui ⋅ haoxian zhu ⋅ Yanghao Zhang ⋅ Hanbiao Meng
Existing audio deepfake detection (ADD) methods frequently struggle with limited generalization, as fixed bottleneck architectures fail to adapt to the heterogeneous information densities inherent in diverse spoofing artifacts. Moreover, standard augmentation often incurs a pathological robustness-accuracy trade-off. In this paper, we propose a unified ADD framework to address these challenges. Our approach integrates: (1) an Adaptive Audio Feature Integration (AAFI) module that dynamically selects optimal latent dimensions via a compression pool and a stochastic exploration mechanism, ensuring flexible representations across multi-type dynamic attacks; and (2) a Mel-Band Adversarial Perturbation (Mel-BAP) strategy that applies targeted regularization in the mel-spectrogram domain to encourage the learning of invariant decision boundaries. Extensive evaluations on benchmarks including ADD 2022, FakeOrReal, and In-the-Wild demonstrate that our Whisper-small based model achieves a state-of-the-art average Equal Error Rate (EER) of 1.51\%. Compared to existing SOTA models such as the 3-billion-parameter Resemble-Detect-3B-Omni, our framework reduces parameter count by up to 91.6\% while improving average EER by 60.7\%. Notably, the model exhibits exceptional zero-shot transferability, achieving 3.91--5.65\% EER on unseen languages and maintaining high resilience against real-world acoustic distortions. This work provides a scalable and efficient “plug-in” solution for enhancing trust and security in the audio Web ecosystem.
Adaptive Conditional Gradient Sliding: Projection-Free and Line-Search-Free Acceleration
Shota Takahashi
We study convex optimization problems over a compact convex set where projections are expensive but a linear minimization oracle (LMO) is available. We propose the *adaptive conditional gradient sliding method* (AdCGS), a projection-free and line-search-free method that retains Nesterov's acceleration with adaptive stepsizes based on local Lipschitz estimates. AdCGS combines an accelerated outer scheme with an LMO-based inner routine. It reuses gradients across multiple LMO calls to reduce gradient evaluations, while controlling the subproblem inexactness via a prescribed accuracy level coupled with adaptive stepsizes. We prove accelerated rates for convex objective functions, matching projection-based methods, without relying on a projection oracle. For locally strongly convex objective functions, we further establish linear convergence without additional geometric assumptions on the constraint set, such as polytopes or strongly convex sets. Experiments on constrained $\ell_p$ regression, logistic regression, and least-squares problems demonstrate that AdCGS improves over projection-free baselines and provides competitive performance when projections are inexpensive.
Adaptive LLM Routing for Multi-Turn Conversations with Continuously Evolving User Queries
Jiarui Zhang ⋅ Xiangyu Liu ⋅ Yong Hu ⋅ Chaoyue Niu ⋅ Hang Zeng ⋅ Shaojie Tang ⋅ Fan Wu ⋅ Guihai Chen
Multi-turn conversation is the predominant form of interaction with large language models (LLMs), where user queries continuously evolve. However, existing LLM routing methods are primarily designed for single-turn interactions or fixed queries settings, overlooking the dynamic nature of multi-turn conversations and the challenge of delayed rewards, thereby limiting their ability to optimize cumulative performance. To address this challenge, we move from myopic, single-turn selection to long-horizon routing for multi-turn conversation. Accordingly, we propose ConvoRouter, which first performs MCTS to explore conversation branches induced by different LLM selections and collect trajectories with high cumulative rewards. ConvoRouter then learns a lightweight routing policy from search-derived data, augmented with retrieval-based future state approximation, enabling multi-turn routing without online search. Experiments on both open-domain and domain-specific conversation tasks across diverse candidate sets of both open-source and closed-source LLMs demonstrate that ConvoRouter significantly outperforms single LLMs and existing routing baselines in task success rate, while achieving a superior performance-cost trade-off when combined with a cost-aware reward.
Adaptive-Margin Masking and Restoration for Balanced Multimodal Learning
Guangbin Zhang ⋅ Zheyang Luo ⋅ Jiangming Liu
Real-world tasks rely on information from multiple sensory sources, motivating multimodal learning as a core paradigm in modern machine learning. However, multimodal learning suffers from modality laziness, where one modality suppresses the contribution of the other modalities by dominating the training process. Recent studies of proposing various metrics to identify lazy modalities and manipulate optimizations to mitigate this imbalance, facing two main problems: 1) The identification of lazy modality is incomplete and overly sharp; 2) The dominated modality can be lagged by under-optimized lazy modalities. To address these issues, we propose Adaptive-Margin Masking and Restoration (AMRe), which introduces adaptive-margin modality identification and restoration optimization to balance dominated and lazy modalities. Experimental results show that AMRe consistently outperforms competitive baselines and several state-of-the-art methods by achieving significant improvement on standard multimodal benchmarks. The codes are released in https://anonymous.4open.science/r/AMRe-5E69.
Adaptive Stepsizes for Eligibility Traces in Deep Reinforcement Learning
Esraa Elelimy ⋅ Martha White
Modern deep reinforcement learning algorithms often store past sensory observations in a buffer for later replay and learning. However, this process is biologically implausible; animals do not store raw sensory observations. Instead, they store their imperfect understanding of the world and the events in it. An alternative to storing all past sensory observations is to learn directly from the observations as they come in and store the seemingly important parts of the incoming data. Eligibility trace algorithms in reinforcement learning (RL) sweep back updates for multiple steps, improving efficiency and providing update rules that do not rely on storing large data buffers. These algorithms were common and effectively used in the linear setting, but have not been widely adopted in the deep RL setting. Many questions remain open to make these algorithms more effective, even as basic as how to incorporate vector step-sizes. We first highlight how adaptive vector step-sizes need to be incorporated into the eligibility trace and then introduce a new vector step-size algorithm that is more amenable to the forward-backward view equivalence needed to derive an update with eligibility traces. We show that our new algorithms improve over other streaming algorithms across Mujoco and MinAtar environments, and are comparable to methods with replay buffers such as PPO.
A Differentiable Interior-Point Method in Single Precision
Jon Arrizabalaga ⋅ Kevin Tracy ⋅ Zac Manchester
Primal-dual interior-point methods solve constrained convex optimization problems to tight tolerances with speed and robustness. Their solutions are also efficiently differentiable with respect to the problem data through the implicit function theorem. However, the standard treatment of primal-dual complementarity makes the underlying linear systems increasingly ill-conditioned near the solution. While this ill-conditioning is often benign in double precision, it can be catastrophic in single precision, preventing interior-point methods from fully exploiting the accelerated hardware that underpins modern machine learning. This paper introduces a differentiable interior-point method designed for low-precision arithmetic. By using an alternative complementarity representation, we ensure that the underlying linear systems remain spectrally bounded --- even near the solution --- a property that is essential for computing accurate gradients and avoiding arithmetic exceptions. As a result, our method enables interior-point solvers to reliably solve and differentiate optimization problems in single precision that were previously confined to double precision. We demonstrate the approach through an ablation study against the standard interior-point formulation and applications in bilevel and end-to-end learning settings where differentiating through constrained optimization is essential.
Advancing Narrative Long Video Generation via Training-Free Identity-Aware Memory
Jinzhuo Liu ⋅ Jiangning Zhang ⋅ Wencan Jiang ⋅ Yabiao Wang ⋅ Dingkang Liang ⋅ Xue zhucun ⋅ Ran Yi ⋅ Yong Liu
Autoregressive video generation has improved rapidly in visual fidelity and interactivity, but it still suffers from long-term inconsistency and memory degradation. Most existing solutions either compress historical frames using predefined strategies or retrieve keyframes based on coarse implicit attention signals, both of which fail to handle evolving prompts with shifting entity references, leading to identity drift, character duplication, and attribute loss. To address this, we propose IAMFlow, a training-free identity-aware memory framework that explicitly models and tracks persistent entity identities, enabling consistent generation across prompt transitions. Specifically, an LLM extracts entities with visual attributes from each prompt and assigns unique global IDs for identity-aware memory, while a VLM asynchronously verifies and refines attributes from rendered frames, enabling explicit entity tracking in place of implicit similarity-based matching. To keep the proposed framework computationally practical, we design a systematic inference acceleration pipeline, including asynchronous visual verification, adaptive prompt transition, and model quantization, which achieves faster generation than existing baselines. Furthermore, we introduce NarraStream-Bench, a benchmark for narrative streaming video generation that features 324 multi-prompt scripts spanning six dimensions and a three-dimensional evaluation protocol that integrates both traditional metrics and multimodal large language model-based assessments. Extensive experiments show that IAMFlow, despite being training-free, achieves the best overall performance on NarraStream-Bench, outperforming the strongest baseline by 2.56 points, while achieving a 1.39$\times$ speedup over the most efficient baseline in the 60-second multi-prompt setting.
Adversarial Corpus Selection to Attack Subgraph Matching based Graph Retrieval
Ninad Gandhi ⋅ Mayukh Mondal ⋅ Brian S Mackwan ⋅ Soumen Chakrabarti ⋅ Abir De
Neural subgraph retrieval systems, which retrieve corpus graphs containing a query graph as a subgraph, are used in safety-critical applications such as drug discovery, hardware trojan detection, and molecular similarity search. Despite their importance, the adversarial robustness of these systems remains unstudied. We propose GRAP, the first adversarial attack framework targeting subgraph matching-based graph retrieval. \our jointly selects a budget-constrained subset of corpus graphs and computes per-graph edge perturbations to maximally degrade retrieval quality. We formalize two attack regimes: a ranking attack (with solver access) that corrupts pairwise relevance ordering while preserving relevance labels, and a solver-free top-K attack that directly displaces relevant results from retrieved sets. In both cases, we show that the resulting adversarial set functions are monotone and approximately submodular, enabling greedy subset selection with provable approximation guarantees. Experiments on five benchmark datasets against multiple state-of-the-art victim retrievers demonstrate that our method consistently and significantly outperforms all baselines in both gray-box and black-box settings, with and without solver access.
Adversarial Risk in the Generative AI Era Necessitates Dropping the Small Epsilon Ball
Andrew Cullen ⋅ Neil Marchant ⋅ Paul Montague ⋅ Jiani Xie ⋅ Benjamin Rubinstein
Adversarial Machine Learning research is currently centered on the search for the $\epsilon$-ball: minimal perturbations designed to fool models while remaining imperceptible to humans. We argue that this fixation on $\ell_p$ norms is a metrological trap that communicates a distorted perception of risk to defenders, one that is unaligned with real-world system vulnerabilities. This paradigm misleads defenders into optimizing for narrow, mathematical robustness at the expense of systemic security, while simultaneously incentivizing attackers to exploit non-overlapping alternative adversarial pathways. Ultimately, we argue that the dominance of the $\ell_p$ ball does not just fail to reduce harm - it may actively facilitate it by communicating a false sense of security against the wrong threats.
Current world models for explicit world simulation typically follow two distinct pathways: (1) video generative models (e.g., Sora), which simulate realistic appearances but fail to preserve 3D consistency; and (2) 3D generative models (e.g., Marble), which generate geometrically consistent scenes but often yield static, less realistic observations. We present AEON, a unified framework that bridges video-based and 3D-based world models and combines the benefits of both. The key insight of AEON lies in a generative model that operates within the video latent space while producing renderable 4D graphics primitives—specifically, 4D Gaussian Splatting (4DGS)—as the world modeler. We train a 4D reconstruction model to predict both 4DGS and camera poses from monocular videos, serving as our initialized 4D decoder. Furthermore, we design a domain adapter that translates generative video latents into reconstructive 4D latents through a distillation process on the video VAE. AEON naturally supports video-based world modeling by rendering 4DGS at predicted camera trajectories, and constructs a persistent 4D environment with scene dynamics for realistic world simulation. Consequently, AEON empowers a diverse range of world modeling applications and achieves state-of-the-art performance on standard benchmarks.
AeroChem: Closed-loop Physics-Informed State Space Modeling for Long-term Chemically-Reactive Air Quality Forecasting
Le Trung Kien ⋅ Tung Kieu ⋅ Dinh D Nguyen ⋅ Cuong Do ⋅ Bin Yang ⋅ Phi Long Nguyen
Accurate long-term air quality forecasting is challenging due to long-range transport, meteorological variability, and nonlinear chemical reactions underlying secondary pollutant formation. Purely data-driven models often suffer from long-horizon drift and lack physical consistency, while existing physics-informed approaches commonly follow open-loop designs that fuse neural and physical components only at the output level. We propose AeroChem, a closed-loop physics-informed state space framework for chemically reactive air quality forecasting. AeroChem adopts a two-branch co-correction design: a UDE-based process branch provides a structured but imperfect physical prior, while a Mamba-based virtual measurement branch captures long-range temporal context and produces probabilistic virtual observations. Neural Kalman Fusion couples the two branches by correcting the latent physical state at each forecasting step and feeding the posterior state back into the simulator, reducing long-horizon drift. We further introduce HanoiAir, a real-world dataset integrating pollutant concentrations, meteorology, and emission inventories. Experiments on HanoiAir and a public benchmark show that AeroChem improves long-horizon stability and sudden-change forecasting.
Affine-Image Propagation for Tight and Scalable $\ell_{2}$ Neural Network Verification
Hong-Ming Chiu ⋅ Haoyu Li ⋅ Huan Zhang ⋅ Richard Y Zhang
While bound propagation methods are highly effective for certifying neural network robustness against $\ell_{\infty}$ adversaries, scalable verification for $\ell_{2}$ perturbations remains a significant open challenge. Existing approaches suffer from severe geometric loss: elementwise bounding methods like $\alpha$-CROWN relax Euclidean balls to bounding boxes, losing a factor of $\sqrt{n}$ in radius, while the recently proposed SDP-CROWN rounds deformed ellipsoids back into balls, losing a factor proportional to the condition number of the weights. Consequently, tight $\ell_{2}$ verification is currently restricted to near-isometric or Lipschitz-regularized networks where these approximations are not vacuous. In this paper, we introduce affine-image propagation, a novel CROWN-compatible framework that exactly propagates the affine geometry of $\ell_{2}$ balls and $\ell_{\infty}$ boxes through the network. To bypass the prohibitive computational cost of exact Semidefinite Programming (SDP) constraints required by this representation, we derive highly scalable, SDP-based bounds utilizing diagonal-dominance surrogates. Our approach effectively eliminates the geometric bottlenecks of prior work, yielding the first large-scale, SDP-style bound propagation method capable of tight $\ell_{2}$ verification for general, non-Lipschitz-regularized neural networks.
Agentick: A Unified Benchmark for General Sequential Decision-Making Agents
Roger Creus Castanyer ⋅ Pablo Samuel Castro ⋅ Glen Berseth
AI agent research spans a wide spectrum: from RL agents that learn from scratch to foundation model agents that leverage pre-trained knowledge, yet no unified benchmark enables fair comparison across these approaches. We present Agentick, a benchmark for sequential decision-making agents designed to evaluate RL, LLM, VLM, hybrid, and human agents on common ground and to power research on the fundamental challenges of sequential decision-making. Agentick provides 37 procedurally generated tasks across six capability categories, four difficulty levels, and five observation modalities, all exposed through a single Gymnasium-compatible interface. The benchmark ships with a Coding API, oracle reference policies for all tasks, pre-built SFT datasets, a composable agent harness, and a live leaderboard. An evaluation spanning 27 configurations and over 90,000 episodes reveals that no single approach dominates: GPT-5 mini leads overall at 0.309 oracle-normalized score while PPO dominates planning and multi-agent tasks; the reasoning harness multiplies LLM performance by 3--10x; and ASCII observations consistently outperform natural language. These findings highlight the substantial room for improvement that remains across all agent paradigms. Agentick's capability-decomposed, multi-modal design provides the empirical infrastructure needed to drive progress toward general autonomous agents, both as an evaluation framework and as a training ground for RL post-training of foundation models in truly sequential environments.
AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation
Zhengren Wang ⋅ Dongsheng Ma ⋅ Huaping Zhong ⋅ Jiayu Li ⋅ Yijie Wang ⋅ Tianyao He ⋅ Wentao Zhang ⋅ Bin Wang ⋅ Conghui He
The expansion of retrieval-augmented generation (RAG) into multimodal domains has intensified the challenge for processing complex visual documents, such as financial reports. While page-level chunking and retrieval is a natural starting point, it creates a critical bottleneck: delivering entire pages to the generator introduces excessive extraneous context. This not only overloads the generator's attention mechanism but also dilutes the most salient evidence. Moreover, compressing these information-rich pages into a limited visual token budget further increases the risk of hallucinations. To address this, we introduce AgenticOCR, a dynamic parsing approach that extends optical character recognition (OCR) from a static, full-text process into a query-driven, on-demand extraction system. By autonomously analyzing document layout in a "thinking with images" manner, AgenticOCR identifies and selectively recognizes regions of interest. This approach performs on-demand decompression of visual tokens precisely where needed, effectively decoupling retrieval granularity from rigid page-level chunking. AgenticOCR has the potential to serve as the "third building block" of the visual document RAG stack, operating alongside and enhancing standard Embedding and Reranking modules. Experimental results demonstrate that AgenticOCR improves both the accuracy and efficiency of visual RAG systems with pixel-adaptive per-image token allocation (e.g., Qwen3.5 generator), achieving expert-level performance in long document understanding. Our repository is anonymously available at https://anonymous.4open.science/r/AgenticOCR.
Agentic Reward Modeling: Verifying GUI Agent via Progressive Trajectory-Grounded Interaction
Chaoqun Cui ⋅ Jing Huang ⋅ SHIJING WANG ⋅ Liming Zheng ⋅ Qingchao Kong ⋅ Zhixiong Zeng
Reinforcement learning with verifiable rewards (RLVR) provides a promising pathway for continuously advancing GUI agents, yet existing reward modeling paradigms face complementary limitations. Rule-based methods suffer from poor scalability and cannot handle open-ended tasks. LLM-as-a-Judge methods enable scalable trajectory verification but remain passive and are constrained by partial state observability, since key evidence often resides in latent environment states beyond the trajectory. Recent active environment interaction methods mitigate observability issues but tend to over-rely on probing while under-utilizing direct trajectory evidence, leading to verification inefficiency. To address these challenges, we advocate a trajectory-grounded interactive verification paradigm. We introduce VAGEN, a framework that employs a tool-augmented verifier agent governed by a Progressive Verification Mechanism, which follows a surface-to-latent and cheap-to-expensive design philosophy to extract trajectory evidence and probe environment states in a proactive end-to-end manner. Experiments on OSWorld-Verified and AndroidWorld benchmarks demonstrate that VAGEN significantly improves evaluation accuracy with a favorable performance-efficiency trade-off.
AgenticVBench: Can AI Agents Complete Real-World Video Production Tasks?
Zongheng Cao ⋅ Yi Zheng ⋅ Xinyu Hu
Video production workflows offer a rich and demanding testbed for evaluating multimodal AI agents: they require composite capabilities across text, image, audio, and video understanding, along with long-horizon planning, and tool use. To this end, we introduce AgenticVBench, a benchmark of 100 tasks across 4 task families spanning the real world post-production workflow, constructed from the representative work of 20 industry experts averaging 6 years of experience across video production contexts. Tasks are paired with evaluation specifications that combine programmatic verifiers and expert rubrics, both authored under expert protocol. We evaluate frontier vision-language models (VLMs) with vendor-native and open-source harnesses. The best evaluated model achieves a score below 30\%, substantially lower than frontier performance on existing multimodal benchmarks. We further find that the choice of harness substantially affects model behavior, including but not limited to scores and failure modes. AgenticVBench provides a foundation for understanding and improving both models and harnesses for agentic video production.
Agents as Neuro-Symbolic Reasoners: Path Feasibility Reasoning for Precise Static Bug Detection
Xueying Du ⋅ Kai Yu ⋅ Chong Wang ⋅ Yi Zou ⋅ Wentai Deng ⋅ Zuoyu Ou ⋅ Xin Peng ⋅ Yiling Lou
Static bug analyzers play a crucial role in ensuring software quality. However, existing analyzers for bug detection in large-scale codebases often suffer from high false positive rates. This limitation largely stems from their inadequate capabilities in performing precise path feasibility validation for complex code contexts. While recent work has explored Large Language Models (LLMs) to eliminate false positives in static bug detection, the limited reasoning capabilities of LLMs over long and complex contexts hinder their effectiveness when applied to large-scale software projects. To address this challenge, we propose PathAgent, an agent-driven framework for fine-grained path feasibility analysis. PathAgent decomposes complex inter-procedural analysis into localized symbolic range reasoning sub-tasks, and employs an on-demand contextual exploration strategy to adaptively retrieve only the necessary code context. These designs enable precise path feasibility reasoning and effectively reduce false positives reported by static bug analyzers. Evaluation on large-scale real-world projects shows that PathAgent eliminates 79% of false positives, outperforming all baselines by 23% to 44%, while maintaining strong bug detection capability with a recall of 0.94.
A Geometric Perspective on Reward Function Updates in Inverse Reinforcement Learning
Anish Abhijit Diwan ⋅ Jan Peters ⋅ Oleg Arenz
Inverse Reinforcement Learning (IRL) is classically solved using Lagrangian optimization, by jointly optimizing a primal-dual problem with gradient descent to obtain the optimal reward function parameters and corresponding optimal policy. While this algorithmic view of IRL has been highly influential, gradient-based reward updates only paint a partial picture, largely ignoring the geometry of the problem in occupancy space. In this paper, we present an occupancy-space view of reward updates in IRL. We leverage the insight that both the optimal solution and its geometry in occupancy space are known, and show that the geometry along the e-geodesic (straight line in occupancy space) toward the optimal occupancy directly affects the gradient of the IRL dual problem in reward parameter space. This perspective naturally suggests preconditioning the dual gradient with the average Fisher information along the occupancy trajectory. We show that this preconditioned gradient is just a convex interpolation of the old reward and a new target reward, yielding efficient natural-gradient-style updates with little additional computational overhead. We then analyse these occupancy-space reward updates from a Mirror Descent (MD) perspective, and show that policy search over the resulting rewards induces approximate entropic MD on the policy. Finally, we conclude with convergence results of such geometry-aware IRL algorithms.
We study online learning in the adversarial injection model [GHMS23], where each instance is either drawn i.i.d. from an unknown distribution or is an arbitrary adversarial injection. The learner never knows which rounds are adversarial, but may abstain without penalty on adversarial injections. The learner's goal is to make few total misclassifications and few abstentions on i.i.d. instances. This model lies between classical statistical learning, governed by the VC dimension, and adversarial online learning, governed by the Littlestone dimension. Unlike prior work, which has largely focused on the realizable setting, we consider an agnostic formulation that allows arbitrary labels. We present an online algorithm that guarantees sublinear excess misclassification error and sublinear abstention error, against oblivious and adaptive adversaries, for function classes that admit bounded bracketing number witnessed by brackets of bounded VC dimension. Our algorithm is distribution-independent and does not require prior knowledge of the distribution. As corollaries, we obtain new learning guarantees for halfspaces (linear classifiers) and richer classes such as feed-forward neural networks with threshold activations.
A Hierarchical Tokenization Framework for Voxel-Level fMRI Representation Learning
Karan Singh ⋅ Yue Zhao ⋅ Mohammadhassan Abbasi ⋅ Ehsan Adeli
We introduce HBAR (Hierarchical Brain Activity Representation), a hierarchical tokenization framework for learning compact, multi-scale representations of neuroimaging data.HBAR organizes voxel-level resting-state fMRI into a hierarchy spanning sub-parcel voxel groups, parcels, and large-scale functional networks, enabling approximately 415x compression while preserving fine-grained spatial information beyond conventional parcellation. The framework supports fixed atlas-based, fully learned, and softly prior-guided hierarchies, allowing us to systematically study the role of structured priors in brain representation learning. Evaluated on large-scale fMRI data, HBAR substantially improves reconstruction over parcellation-based and single-scale tokenization baselines, while supporting downstream modeling of brain dynamics and demographic prediction tasks such as sex classification. Across reconstruction and downstream evaluation, atlas-guided hierarchies provide the strongest overall performance, suggesting that established functional atlases encode organizational structure that remains difficult to recover from reconstruction alone. By combining voxel-level fidelity, compact discrete tokens, and multiscale brain structure, HBAR offers a practical framework for representation learning in high-dimensional neuroimaging data.
ALETHEIA: A Multi-Frequency Eddy Current Pulsed Thermography Dataset for Neural Operator Learning in Nondestructive Testing
Changbin Sun ⋅ Xiaojie ⋅ Xiaotian Chen ⋅ Yuankai Wu
Learning neural solvers for spatiotemporal partial differential equations (PDEs) under real-world constraints remains a key challenge in scientific machine learning, especially for inverse tasks with sparse and noisy boundary observations. We present the Aletheia dataset, the first 3D benchmark for learning data-driven solvers in the context of nondestructive testing (NDT). The dataset simulates eddy-current-induced heating in conductive solids and models the resulting transient heat propagation governed by the heat equation. Aletheia contains over 4,700 high-resolution samples across 10 excitation frequencies (1-100\,kHz), each providing volumetric heat source and temperature fields over time. It supports both forward prediction of temperature evolution and inverse reconstruction of internal heat sources or defects from surface infrared measurements. Real infrared thermography data from cracked rail specimens are included for calibration and generalization studies. We define three canonical tasks on both regular and irregular grids and benchmark them using various neural operators. Aletheia establishes a unified platform for evaluating neural PDE solvers under realistic NDT conditions, enabling progress in reliable, data-driven inverse modeling.
AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving
Yanhao Wu ⋅ Haoyang Zhang ⋅ Fei He ⋅ Rui Wu ⋅ Yanhu Shan ⋅ Congpei Qiu ⋅ Liang Gao ⋅ Wei Ke ⋅ Tong Zhang
Practical autonomous driving requires models that generalize by reasoning through spatial-temporal possibilities to exclude unsafe outcomes. While state-of-the-art (SOTA) methods use parallel planning architectures, they fail to explicitly couple speed decisions with agent behavior along the driving path, leading to suboptimal coordination. To address this, we propose a cascaded framework that transforms longitudinal planning from an independent prediction task into a path-conditioned reasoning process. On the model side, we introduce an anchor-based regression design that conditions longitudinal prediction on the lateral drive path, and reformulate longitudinal planning as 1D displacement prediction along the path. This reduces geometric uncertainty and sharpens the model's focus on interaction-driven dynamics. On the data side, we introduce a planning-oriented data augmentation strategy that simulates rare safety-critical events by programmatically inserting agents and relabeling longitudinal targets to enforce collision avoidance. Evaluated on the challenging Bench2Drive benchmark, our method achieves SOTA performance with a driving score of 89.07 and a success rate of 73.18\%, demonstrating significantly improved coordination and safety. Further evaluation on Fail2Drive confirms strong generalization to rare edge cases where parallel formulations typically fail. Visualization videos are provided in the supplementary material.
All Roads Lead to Rome: Flow-driven Multi-Anchor Exploration for Open-Environment Active 3D Mapping
Yang Li ⋅ Aming WU ⋅ Zihao Zhang ⋅ Ziju Han ⋅ Sijia Zhang ⋅ Yahong Han
To advance the development of embodied intelligence, Open-Environment Active 3D Mapping has attracted increasing attention, aiming to perform a long-horizon and shortest trajectory exploration for reconstructing unseen scenarios. Since only limited information about unseen environments is available, methods built on the closed-set assumption, i.e., assuming that the test environments are similar to those seen during training, cannot generalize satisfactorily. In existing active mapping methods, long-horizon exploration is often guided by predicting a coarse long-range goal and then converting it into an executable path. However, this stage is usually formulated as single-point prediction. Under partial observability, the same local observation may correspond to multiple plausible exploration directions, making such deterministic prediction prone to brittle decisions and degraded performance in unseen scenarios. Our experiments further verify that this is a key factor underlying their weak generalization. To address this issue, we reformulate long-horizon target prediction as conditional multimodal anchor generation using Conditional Flow Matching. Instead of predicting a single goal, our method learns a conditional distribution over coarse exploration anchors from the current mapping state. These anchors are first converted into executable candidate paths through obstacle-aware planning. We then apply exploration-mode clustering to compress geometrically similar trajectories and reduce candidate redundancy. Finally, a hierarchical selection module selects the most promising mode and reranks paths within it to produce the final executable trajectory. Experiments show that our method improves generalization and reconstruction efficiency in open environments.
A Local Geometric Analysis of Maximal Coding Rate Reduction via Error Bounds
Peng Wang ⋅ po chen ⋅ Huikang Liu ⋅ Rujun Jiang
While the maximal coding rate reduction (MCR$^2$) objective has become a useful principle for learning compact and discriminative representations, a theoretical understanding of its optimization geometry remains limited, especially on why first-order methods often exhibit fast convergence near its local maximizers. In this work, we show that the MCR$^2$ objective satisfies an error-bound property in a neighborhood of its local maximizers. This property guarantees that the distance to the local maximizer set can be bounded above by the gradient norm at any point in the neighborhood, directly implying regularity conditions such as the Polyak-Lojasiewicz inequality and quadratic growth. Leveraging this property, we prove that first-order methods converge linearly to a local maximizer of the MCR$^2$ objective under mild conditions. Numerical experiments validate the predicted local convergence behavior and provide empirical evidence that optimizing the MCR$^2$ objective improves representation geometry while maintaining competitive downstream accuracy.
AlphaPROBE: Alpha Mining via principled retriveval and on-graph biased evolution
Taian Guo ⋅ Haiyang Shen ⋅ Junyu Luo ⋅ Binqi Chen ⋅ Hongjun Ding ⋅ Jinsheng Huang ⋅ Luchen Liu ⋅ Yun Ma ⋅ Ming Zhang
Extracting signals through alpha factor mining is a fundamental challenge in quantitative finance. Existing automated methods primarily follow two paradigms: Decoupled Factor Generation, which treats factor discovery as isolated events, and Iterative Factor Evolution, which focuses on local parent-child refinements. However, both paradigms lack a global structural view, often treating factor pools as unstructured collections or isolated single-lineage paths, which leads to redundant search and inefficient allocation of optimization effort. To address these limitations, we introduce AlphaPROBE, (Alpha Mining via Principled Retrieval and On-graph Biased Evolution), a framework that reframes alpha mining as the strategic navigation of a Directed Acyclic Graph (DAG). By modeling factors as nodes and evolutionary links as edges, AlphaPROBE treats the factor pool as a dynamic, interconnected ecosystem. The framework consists of two core components: a Bayesian Factor Retriever that identifies high-potential seeds by balancing exploitation and exploration through a posterior probability model, and a DAG-aware Factor Generator that leverages the full ancestral trace of factors to produce context-aware, non-redundant optimizations. Extensive experiments on major Chinese and US stock market datasets against 8 competitive baselines demonstrate that AlphaPROBE significantly gains enhanced performance in predictive accuracy, return stability and training efficiency. Our results confirm that leveraging global evolutionary topology is essential for efficient and robust automated alpha discovery. We have open-sourced our implementation at https://anonymous.4open.science/r/AlphaPROBE-A103.
AlphaQ: Calibration-Free Bit Allocation for Mixture-of-Experts Quantization
Wanqi Yang ⋅ Yuexiao Ma ⋅ Alexander Conzelmann ⋅ Xiawu Zheng ⋅ Michael Mahoney ⋅ T. Konstantin Rusch ⋅ Shiwei Liu
Mixture-of-Experts (MoE) architectures scale model capacity through sparse expert activation, but their deployment remains memory-bound because all expert weights must reside in memory. Mixed-precision quantization can substantially reduce this footprint by assigning different bit-widths to different experts. Existing approaches, however, typically rely on calibration data to estimate expert importance and determine bit allocation. For frontier MoE LLMs, the original training data, and hence the true training distribution, is proprietary and inaccessible. As a result, calibration sets are inevitably imperfect surrogates, which can misestimate expert utilization and lead to suboptimal bit allocation. Motivated by the substantial cross-expert quality variability observed in modern MoE models, we propose AlphaQ, a calibration-free bit-allocation method for MoE quantization. AlphaQ draws on Heavy-Tailed Self-Regularization (HT-SR) theory and follows a simple principle: experts with more heavy-tailed weight spectra are typically better trained and should receive higher bit-widths, while experts with weaker heavy-tailed structure can be quantized more aggressively. AlphaQ operationalizes this principle by measuring expert-wise spectral heavy-tailedness and solving a budget-constrained optimization problem that minimizes total quantization error under a global bit-budget constraint. Across several MoE models, AlphaQ consistently outperforms calibration-based baselines under matched bit budgets. Notably, on Qwen1.5-MoE, AlphaQ achieves near full-precision accuracy with an average expert precision of only 3.5 bits, while delivering more than 4$\times$ memory compression.
A Measure-Theoretic Analysis of Reasoning: Structural Generalization and Approximation Limits
Yuyang Zhang ⋅ Yifu Zhang ⋅ Xuehai Zhou ⋅ Xiaoyin Chen
While empirical scaling laws for LLM reasoning are well-documented, the theoretical mechanisms governing out-of-distribution (OOD) generalization remain elusive. We formalize reasoning via optimal transport, projecting discrete trajectories into a continuous metric space to quantify domain shifts using the Wasserstein-1 distance. Invoking Kantorovich duality, we bound OOD generalization via architectural Lipschitz continuity and functional approximation limits. This exposes two primary constraints. First, position-dependent attention (e.g., Absolute Positional Encoding) fails to preserve shift invariance, yielding an $\Omega(1)$ Lipschitz constant and expected risk, whereas shift-invariant mechanisms (e.g., Rotary Embeddings) preserve equivariance and bound the error. Second, by mapping sequential backtracking to a Dyck-$k$ language, we establish a strict circuit depth lower bound for $\text{TC}^0$ Transformers. Scaling physical layer depth is necessary to avert representation collapse---a constraint that scaling representation width cannot bypass due to irreducible approximation bounds in Barron spaces. Evaluations across 54 Transformer configurations on combinatorial search corroborate these bounds, demonstrating that generalization risk degrades monotonically with the Wasserstein domain shift.
AMUSE: Anytime Muon with Stable Gradient Evaluation
Jueun Kim ⋅ Baekrok Shin ⋅ Jihun Yun ⋅ Beomhan Baek ⋅ Minhak Song ⋅ Chulhee Yun
Modern deep learning commonly relies on AdamW with prescribed learning rate schedules, but recent works challenge both components: Schedule-Free optimization removes explicit schedules via iterate averaging, and Muon improves the update geometry by orthogonalizing momentum for matrix parameters. Despite Muon's strong empirical performance, its underlying mechanism remains partially understood. We study Muon through the river-valley loss landscape, where useful training progress occurs along a flat, low-curvature bulk subspace (the river), while high-curvature dominant directions form steep valley walls that induce oscillations. We empirically show that while Muon's orthogonalization accelerates river progress by increasing the bulk component, it also amplifies dominant-direction noise, causing oscillatory trajectories. Building on this, we propose Anytime MUon with Stable gradient Evaluation (AMUSE), which integrates Muon's rapid bulk progress with the stabilizing effect of Schedule-Free averaging. AMUSE uses a time-varying interpolation coefficient that initially evaluates gradients near the fast Muon sequence for rapid adaptation, then gradually shifts toward the stable averaged sequence to suppress valley-wall oscillations. As a result, AMUSE requires no learning rate schedules and supports anytime training. Across vision tasks and large language model pretraining, AMUSE consistently improves the performance-iteration Pareto frontier over (Schedule-Free) AdamW and Muon.
Denoising generative models predict clean data ($x$), noise ($\epsilon$), or velocity ($v$), and are trained with a squared loss that can be computed in any of these three spaces. Prior work treats these choices as interchangeable up to loss reweighting, but practitioners observe sharp qualitative differences that remain unexplained. We systematically decouple the \emph{prediction space}, what the network outputs, from the \emph{loss space}, where supervision is applied, creating a $3{\times}3$ design matrix. We show that loss-space changes only reweight how much each noise level contributes to training, while prediction-space changes alter what the network is asked to learn: two fundamentally different mechanisms. This decoupling exposes two regimes. Without a representational bottleneck, the loss choice primarily controls denoising quality (MSE) while the prediction choice primarily controls sample quality (FID). The reason is that ODE sampling must convert the network's output back into a clean image, and this conversion amplifies errors differently across prediction types, turning sub-one-percent gaps in training error into 22-fold gaps in FID that are invisible to standard training diagnostics. Under a representational bottleneck, the prediction choice instead dominates both metrics: recovering noise or velocity requires reconstructing high-dimensional information that the bottleneck has destroyed, while recovering the clean image remains feasible because real images lie on a low-dimensional manifold. We validate these findings on a U-Net and the full 130M-parameter JiT architecture across CIFAR-10 and Imagenette, and give architecture-dependent recommendations for both regimes.
An Assessment of Human vs. Model Uncertainty in Soft-Label Learning and Calibration
Maja Pavlovic ⋅ Silviu Paun ⋅ Massimo Poesio
Central to human-aligned AI is understanding the benefits of human-elicited labels over synthetic alternatives. While human soft-labels improve calibration by capturing uncertainty, prior studies conflate these benefits with the implicit correction of mislabeled data (mode shifts), obscuring true effects of soft-labels. We present a controlled audit of soft-label learning across MNIST and a synthetic variant, re-annotating subsets to extract human uncertainty. By decoupling soft-label supervision from underlying label mode shifts, we show that while human soft-labels do provide accuracy gains, their larger value lies in acting as a regularizer that improves model calibration on difficult samples and promotes stable convergence across training runs. Dataset cartography reveals models trained on human soft-labels mirror human uncertainty, whereas those trained on synthetic labels fail to align with humans. Broadly, this work provides a diagnostic testbed for human-AI uncertainty alignment.
Anatomy-Preserving Unpaired Medical Image Translation via Shared Latent Anchoring
Zhilin Zou ⋅ Hieu Le ⋅ Saumya Gupta ⋅ Jingyi Xu ⋅ Prateek Prasanna ⋅ Chao Chen
Unpaired medical image-to-image translation requires synthesizing target-modality images from source-modality images without aligned supervision. Existing methods learn a direct source-to-target mapping that entangles structure preservation and appearance generation within a single model, often distorting clinically meaningful anatomy in the translated output. We argue that these two objectives should be decoupled entirely: anatomical structure should be captured independently of generation, and the generator should operate only on a frozen structural representation. To this end, we propose a two-stage framework for unpaired medical image translation. In the first stage, we learn a shared structure space by training a structure encoder-decoder on source images, target images, and structural masks jointly, using mask reconstruction to anchor the latent space to modality-agnostic anatomy. For source images without mask annotations, a target-domain memory bank provides soft structural supervision via prototype retrieval. In the second stage, a conditional flow model learns to render target-domain appearance under fixed structural conditions. This ensures the generator inherits target appearance statistics and can focus entirely on structural conditioning rather than learning appearance from scratch. We validate our framework on OCT-to-OCTA synthesis and brain MR-to-CT translation. Our method outperforms all unpaired baselines on both tasks and, despite requiring no paired supervision, also surpasses paired baselines, improving PSNR by up to 11.2% and SSIM by up to 26.4% on OCT-to-OCTA, and achieving consistent gains across whole-image, soft-tissue, and bone regions on MR-to-CT.
An Axiomatic Analysis of DPO and NLHF as Reference-Dependent Probabilistic Voting Rules
Wesley H Holliday ⋅ Adam Lesnikowski ⋅ Roberto-Rafael Maura-Rivero
Post-training methods for large language models (LLMs) implicitly aggregate diverse human feedback in ways that can be studied explicitly using social choice theory. In this paper, we adopt the perspective that for each post-training method, there is a corresponding reference-dependent probabilistic voting rule. Taking Direct Preference Optimization (DPO) and Nash Learning from Human Feedback (NLHF) as our case studies, we study axiomatic properties of the associated voting rules. First, we show that NLHF violates one of the central axioms of social choice, namely monotonicity, if and only if its KL regularization parameter is below a bound, whereas DPO satisfies monotonicity regardless of its regularization parameter. Second, we show that while DPO is often associated with Borda-style preference aggregation, when DPO is regarded as a probabilistic voting rule, it violates two variable population axioms that the standard probabilistic version of Borda satisfies. Third, we show that unlike DPO, NLHF violates a commutativity axiom from Bayesian epistemology, which has practical implications for staged post-training. Finally, we conduct an empirical study of the frequency and magnitude of axiom violations in the CoVal dataset using real base models.
An Efficient Algorithm for Thresholding Monte Carlo Tree Search
Shoma Nameki ⋅ Atsuyoshi Nakamura ⋅ Junpei Komiyama ⋅ Koji Tabata
We introduce the Thresholding Monte Carlo Tree Search problem, in which, given a tree $\mathcal{T}$ and a threshold $\theta$, a player must answer whether the root node value of $\mathcal{T}$ is at least $\theta$ or not. In the given tree, 'MAX' or 'MIN' is labeled on each internal node, and the value of a 'MAX'-labeled ('MIN'-labeled) internal node is the maximum (minimum) of its child values. The value of a leaf node is the mean reward of an unknown distribution, from which the player can sample rewards. For this problem, we develop a $\delta$-correct sequential sampling algorithm based on the Track-and-Stop strategy that has asymptotically optimal sample complexity. We show that a ratio-based modification of the D-Tracking arm-pulling strategy leads to a substantial improvement in empirical sample complexity, as well as reducing the per-round computational cost from linear to logarithmic in the number of arms.
An Efficient Cross-modal Feature Reconstruction Model for Multimodal Multi-class Anomaly Detection
Zhenghan Wang ⋅ Zhihui Luo ⋅ Haitao Hu ⋅ Wengang Cheng
Industrial anomaly detection is evolving from category-specific detectors toward unified models capable of handling multiple categories simultaneously. However, in multi-class setting, greater inter-category diversity not only compels enhanced reconstruction capacity to model normal patterns but also exacerbates the “identical shortcut” problem, wherein anomalies are likewise well reconstructed. Moreover, existing unified frameworks are typically confined to single-modal (RGB) inputs and lack multimodal capability. To address these, we propose an Efficient Cross-modal Feature Reconstruction (ECFR) model for unified multi-class anomaly detection, which harnesses the inherent difficulty of cross-modal reconstruction to alleviate the shortcut issue and amplify anomaly discrimination. The framework is built on two core modules: an Adaptive Feature Interaction and Recalibration (AFIR) module and a Hybrid Attention Convolution (HAC) module. Through feature interaction, fusion, and reconstruction, our model achieves two key outcomes: learning normal patterns of multi-class objects while forcing reconstruction failures for anomalous inputs. Extensive experiments on the MVTec 3D-AD and Eyecandies datasets validate our approach, which achieves state-of-the-art performance in multimodal multi-class anomaly detection while requiring significantly fewer model parameters (44.63 M) and lower computational complexity (15.77 GFLOPs) compared to previous multimodal anomaly detection methods.
An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models
Mingzhong Sun ⋅ Teresa Yeo ⋅ Armando Solar-Lezama ⋅ Tan Zhi-Xuan
Studies of human reasoning have shown that people are typically stronger at evaluating reasoning than producing it from scratch. In contrast, large reasoning models (LRMs) are trained to excel at producing long chains of reasoning to solve complex problems. How then do LRMs perform at evaluating reasons? We investigate this with the Valid-Answer-Invalid-Reasoning (VAIR) dataset: math problems and solutions with trivial reasoning flaws but valid answers, designed to isolate reasoning evaluation from the confound of reasoning production. Unlike humans, who we find are only 6\% worse at grading than solving such problems, we find a substantial production-evaluation gap in LRMs: frontier models score as low as 48\% when evaluating VAIR solutions, despite near-perfect solution production. Why this enigma? Through chain-of-thought (CoT) analysis, we find evidence of an answer confirmation bias: LRMs often produce then check for the correct answer instead of carefully verifying each step, fabricating rationalizations even when noticing anomalous reasoning. Linear probes corroborate this, showing that while LRM activations encode some representation of valid reasoning, they fail to robustly represent VAIR solutions as invalid. Causal patching of the final answer's representations causes LRM verdicts and activations to flip, demonstrating that answer validity is responsible for models' confirmation biases. These findings indicate an outstanding limitation in dominant approaches to reasoning training, which incentivize LRMs to produce and confirm reasoning towards correct answers, but not to robustly evaluate the underlying reasons.
A Novel Schur-Decomposition-Based Weight Projection Method for Stable State-Space Neural-Network Architectures
Sergio Mauricio Vanegas Arias ⋅ Lasse Lensu ⋅ Fredy Ruiz
Building black-box models for dynamical systems from data is a challenging problem in machine learning, especially when asymptotic stability guarantees are required. In this paper, we introduce a novel stability-ensuring and backpropagation-compatible projection scheme based on the Schur decomposition for the state matrix of linear discrete-time state-space layers, as well as an alternative pre-factorized formulation of the methodology. The proposed methods dynamically project the quasi-triangular factor of the state matrix's real Schur decomposition onto its nearest stable peer, ensuring stable dynamics with minimal overparameterization. Experiments on synthetic linear systems demonstrate that the method achieves accuracy and convergence rates comparable to those of state-of-the-art stable-system identification techniques, despite a marginal increase in computational complexity. Furthermore, the lower weight count facilitates convergence during training without sacrificing accuracy in stacked neural-network architectures with static nonlinearities targeting real-world datasets. These results suggest that the Schur-based projection provides a numerically robust framework for identifying complex dynamics on par with the State of the Art while satisfying strict asymptotic-stability requirements.
Answering At Any Cost: Frontier LLMs Are Consequence-Insensitive
Arka Pal ⋅ Kwok C Au ⋅ Louai Zahran ⋅ Rahul K Thomas ⋅ Kevin Hayes ⋅ Micah Goldblum ⋅ Adam Block
LLM-based systems are increasingly deployed in domains where incorrect outputs carry real costs, often necessitating expensive human review and verification. We argue that a central issue is not simply error, but consequence-insensitivity: models fail to adjust their behavior according to the cost of being wrong. We study this failure across both explicit utility framings as well as natural-language descriptions of stakes that mirror real-world deployment scenarios. Our evaluation spans agentic coding and mathematical reasoning, covering five frontier model families. Across settings, models systematically under-abstain: they continue to answer or submit patches even when incorrect answers carry significant consequences. The pathology is striking: models continue to submit answers in trivial settings where abstention is strictly dominant, and even when told an incorrect answer will cause nuclear extinction. Further, we find that this behavior is orthogonal to existing benchmarks; increasing model size or capability within a family does not improve sensitivity. Comparing base and instruction-tuned models localizes much of this anti-abstention bias to post-training: base models abstain far more often, though they are not themselves reliably consequence-aware. Finally, we evaluate prior post-hoc interventions alongside in-context learning and fine-tuning, finding that none robustly resolves the failure. Our results suggest that current post-training and evaluation pipelines optimize models to answer, not to act under stakes -- a significant bottleneck for trustworthy autonomy.
AnyMo: Geometry-Aware Setup-Agnostic Modeling of Human Motion in the Wild
Baiyu Chen ⋅ Zechen Li ⋅ Wilson Wongso ⋅ Lihuan Li ⋅ Xiachong LIN ⋅ Hao Xue ⋅ Benjamin Tag ⋅ Flora Salim
As wearable and mobile devices become increasingly embedded in daily life, they offer a practical way to continuously sense human motion in the wild. But inertial signals are highly dependent on the sensing setup, including body location, mounting position, sensor orientation, device hardware, and sampling protocol. This setup dependence makes it difficult to learn motion representations that transfer across devices and datasets, and limits the broader use of wearable IMUs beyond closed-set recognition. We introduce AnyMo, a geometry-aware framework for setup-agnostic human motion modeling. AnyMo uses physics-grounded IMU simulation over dense body-surface placements to generate diverse and plausible synthetic signals, pre-trains a graph encoder from paired synthetic placement views and masked partial observations, tokenizes multi-position IMU into full-body motion tokens, and aligns these tokens with an LLM for motion-language understanding. We evaluate AnyMo on three complementary tasks: zero-shot activity recognition across 14 unseen downstream datasets, cross-modal retrieval, and wearable IMU motion captioning, where it improves average Accuracy/F1/R@2 by 11.7\%/11.6\%/22.6\% on HAR, increases zero-shot IMU-to-text and text-to-IMU retrieval MRR by 15.9\% and 28.6\%, respectively, and improves zero-shot captioning BERT-F1 by 18.8\%. These results support AnyMo as a generalist model for wearable motion understanding in the wild. Code is at \url{https://anonymous.4open.science/r/anymo}.
APPSolver: Adaptive Patch Partitioning for Point-Wise Ship Flow Prediction on Unstructured Meshes
Wenhua Huo ⋅ Fenglei Han ⋅ Wangyuan Zhao ⋅ Xiao Peng ⋅ Chunhui Wang ⋅ Jialin Wu ⋅ JiaYi Han
Predicting high-fidelity flow fields around ship hulls is central to hydrodynamic performance analysis, yet traditional CFD solvers remain computationally prohibitive for rapid design exploration across multiple vessel types and operating conditions. Existing deep-learning surrogates often interpolate unstructured CFD meshes onto regular grids, which can blur near-field resolution, while direct point-level global attention on large unstructured meshes can be computationally expensive. In this paper, we introduce APPSolver, a neural surrogate that learns next-step point-wise flow-field prediction directly on the unstructured free-surface ship CFD mesh without re-gridding. APPSolver is built around Adaptive Patch Partitioning (APP), a quadtree-based strategy that groups unstructured points into local spatial patches according to the non-uniform density inherent in ship CFD meshes: fine near the hull and wake, coarse in the far field. The resulting patch tokens are processed by a Transformer backbone together with optional condition token that encode ship geometry and condition parameters via a frozen large language model, forming a unified token-space model for one-step temporal advancement of velocity and pressure fields. On the ShipBench dataset, APP-Transformer achieves the lowest MAE and MSE. Ablation studies quantify the compression-fidelity trade-off of APP and demonstrate that condition token provide preliminary gains in some leave-one-hull-out settings. The code will be released upon acceptance.
A Primal-dual Approach for Semi-Infinitely Constrained Reinforcement Learning
Di Wang ⋅ Liangyu Zhang ⋅ Haishan Ye ⋅ Guang Dai ⋅ Ivor Tsang
We propose a primal-dual policy optimization method for reinforcement learning in semi-infinitely constrained Markov decision processes (SICMDP), which extends standard constrained Markov decision processes by allowing a continuum of constraints. We call our proposed method \textbf{S}emi-\textbf{I}nfinitely \textbf{P}rimal-\textbf{D}ual \textbf{P}olicy \textbf{O}ptimization (SI-PDPO). By introducing an infinite-dimensional Lagrange multiplier, we derive the saddle-point formulation of the original SICMDP problem and prove strong duality under mild conditions. We then introduce a regularization term for the dual variable, and solve the resulting regularized saddle-point problem using first-order methods with a policy optimization subroutine (for example, NPG or PPO). Unlike existing primal-type policy optimization algorithms for semi-infinitely constrained reinforcement learning, our method does not require solving an inner-loop optimization problem, thereby mitigating computational intractability and reducing the over-conservative bias. We perform a series of numerical experiments, including a real-world power control task in wireless communications, to evaluate the performance of SI-PDPO. Empirical results show that SI-PDPO consistently outperforms existing primal-type algorithms.
A Private Empirical Defense Against Privacy Audits
Saloni Modi ⋅ Srivi Balaji ⋅ Yusong Zhu ⋅ Gautam Kamath ⋅ Kevin Tian
Differential privacy (DP) has traditionally been used to provide theoretical upper bounds on an algorithm's stability to changing its training data. In modern machine learning applications, achieving strong tradeoffs between utility and theoretical privacy is challenging, and thus one may optimistically hope that existing theoretical analyses are loose. Recent work on \emph{privacy auditing} has adopted a dual viewpoint, instead lower bounding the true privacy of an algorithm by constructing empirical distinguishing events. The auditing literature has thus far yielded a pessimistic outlook on the looseness of theoretical privacy bounds for DP-SGD, the de facto private training method in modern ML, as nearly-matching empirical lower bounds have been achieved under various threat models \cite{NasrHSBTJCT23, AnnamalaiC24, CebereBP25}. In this work, we propose the empirical privacy lower bound of an algorithm as a concrete metric to optimize for, complementary to the theoretical upper bound. We give a lightweight defense framework that generically augments optimization methods in the ML pipeline to have significantly-improved empirical privacy on standard benchmarks. Moreover, we show that our framework comes at \emph{no theoretical privacy cost} when augmenting DP-SGD, unlike previous defenses against membership inference attacks. We evaluate our defense against a broad range of audit constructions, models, and datasets to demonstrate its flexibility. Our code implementation can be found at \href{https://github.com/neurips26privacyaudit/empirical-privacy-defense}{this anonymous repository}.
Are LLMs Good at Feature Engineering? Evidence from a Controlled Synthetic Benchmark
Viacheslav Surkov ⋅ Yasha Pushak ⋅ Ritesh Ahuja ⋅ Rafael Pires ⋅ Anne-marie Kermarrec ⋅ Damien Hilloulin ⋅ Rhicheek Patra ⋅ Sungpack Hong
Large Language Models (LLMs) are increasingly used for automated feature engineering (FE) on tabular data, but prevailing benchmarks are ill-suited to evaluate this capability. They often assume i.i.d. samples, lack temporal dynamics and distribution shifts common in production, provide no known performance ceiling, and their artifacts (e.g., winning solutions) may enter LLM training corpora—blurring genuine discovery vs. memorization or spurious correlation. We introduce TabGen, a controllable synthetic generator that produces two synchronized views: an observable view for learners and an oracle view with minimal sufficient features, yielding a deterministic upper bound. TabGen composes non-i.i.d. dynamics to control inter-row interactions and distribution shift. Using TabGen, we evaluate several FE strategies, including Analyze&Act—a hypothesis-driven FE workflow proposed here. In no-drift settings, methods that instruct LLMs to iteratively analyze and refine data hypotheses consistently outperform alternatives, though LLM-driven FE shows substantial variance. Under drift, all methods degrade, and many fail to generalize—often underperforming the no-FE baseline—indicating FE-level overfitting to training-period correlations that do not transfer to future dynamics. Traditional i.i.d. benchmarks, by design, cannot reveal this failure mode.
Are Multimodal Benchmarks Really Useful? Item-Level Multimodal Benchmark Diagnosis via Structure-Response Co-Calibration
Shiqi Zhang ⋅ Weixin Zeng ⋅ Ziheng Zhang ⋅ Jiuyang Tang ⋅ Xiang Zhao
Multimodal benchmarks are central to evaluating whether models can reason over information from multiple modalities. However, existing literature has revealed that some benchmarks may contain shortcut items that can be answered from a single modality or from answer-option regularities alone, thus failing to reflecting models' true cross-modal inference ability. Existing methods for assessing benchmarks typically rely either on intrinsic item structure or extrinsic model response behavior, which are not adequate for a comprehensive diagnosis. To fill in this gap, we propose SRCoD, a structure-response co-calibrated diagnosis framework for item-level diagnosis of modality dependence in multimodal benchmarks. SRCoD estimates sample-wise multimodal information structure from benchmark content and target answers, and uses this signal as a soft structural anchor for modality-conditioned item response modeling. The learned model produces a calibrated item-level diagnostic profile that reflects how strongly an item depends on different forms of modality evidence, enabling fine-grained benchmark analysis beyond aggregate accuracy. SRCoD also provides interpretable attribution of the diagnosis results. Extensive experiments on two complementary protocols, i.e., benchmark-refinement oriented validation and human-labeled item validation, demonstrate that SRCoD outperforms state-of-the-art benchmark diagnosis methods. Our code and data are available at https://anonymous.4open.science/r/SRCoD-61B0.
ARTIS: Agentic Risk-Aware Test-Time Scaling via Iterative Simulation
Xingshan Zeng ⋅ Lingzhi Wang ⋅ Weiwen Liu ⋅ Liangyou Li ⋅ Yasheng Wang ⋅ Lifeng Shang ⋅ Xin Jiang ⋅ Qun Liu
Current test-time scaling (TTS) techniques enhance large language model (LLM) performance by allocating additional computation at inference time, yet they remain insufficient for agentic settings, where actions directly interact with external environments and their effects can be irreversible and costly. We propose ARTIS, Agentic Risk-Aware Test-Time Scaling via Iterative Simulation, a framework that decouples exploration from commitment by enabling test-time exploration through simulated interactions prior to real-world execution. This design allows extending inference-time computation to improve action-level reliability and robustness without incurring environmental risk. We further show that naive LLM-based simulators struggle to capture rare but high-impact failure modes, substantially limiting their effectiveness for agentic decision making. To address this limitation, we introduce a risk-aware tool simulator that emphasizes fidelity on failure-inducing actions via targeted data generation and rebalanced training. Experiments on multi-turn and multi-step agentic benchmarks demonstrate that iterative simulation substantially improves agent reliability, and that risk-aware simulation is essential for consistently realizing these gains across models and tasks.
A Scientific Claim Stability Framework for Evaluating MRI Morphometry Pipelines
Duy-Cat Can ⋅ Dang-Duat Tran ⋅ Anh-Khang Nguyen ⋅ Bao Tran ⋅ Trung-Hieu Do ⋅ Van T Nguyen ⋅ Duy-Thanh VU ⋅ Oliver Y Chén ⋅ Hoang-Quynh Le
MRI morphometry pipelines are often evaluated by measurement agreement or downstream accuracy, yet their scientific value depends on the conclusions they support. This creates a gap: two pipelines may look similar at the feature level while supporting different conclusions about aging, cognitive decline, or neurodegeneration. To bridge this disparity, here we propose Scientific Claim Stability (SCS), a reusable claim-centric evaluation framework that converts region-of-interest (ROI) morphometry tables into interpretable claim cards and tests whether each claim remains supported under realistic perturbations. To demonstrate the efficacy of SCS, we evaluate it on 10,370 MRI scans from three datasets and five cohorts, generating 2,783 claim cards and evaluating 89 claim-family pairs across dataset, pipeline, atlas, feature-family, and target shifts. SCS identifies claims that are stable, transferable, or fragile, and reveals high-concordance fragile-support cases where evidence vectors remain similar while supporting regions change. To assess claim plausibility, clarity, faithfulness, usefulness, and caution level, we further implement and compare an expert (including geriatrics, neuroscience, and radiology) review workflow and an LLM (including ChatGPT and Gemini) review workflow. Our results suggest claim-level evaluation as a practical standard for auditing MRI morphometry pipelines without requiring a universal gold standard.
A Single Deep Preference-Conditioned Policy for Learning Pareto Coverage Sets
Akihiro Kubo ⋅ Kosuke Nakanishi ⋅ Shin Ishii
Preference-conditioned multi-objective reinforcement learning aims to learn a single policy that captures trade-offs across preferences, but under nonlinear scalarization the uniqueness and continuity of the preference-to-solution correspondence remain unclear. We study this problem in tabular multi-objective Markov decision processes (MDPs) using smooth Tchebycheff scalarization as a monotone utility. Under mild interior conditions on the preference set, we prove that each preference induces a unique Pareto-optimal return vector and that this vector depends Lipschitz-continuously on the preference, providing a principled foundation for preference sweeping toward dense Pareto-front coverage. To compute these targets, we formulate the problem over occupancy measures and derive Concave Mirror Descent Policy Iteration (CMDPI), which achieves an $O(1/k)$ objective-suboptimality rate. We further show that each update is equivalent to solving a Kullback-Leibler-regularized MDP with the previous policy as reference, yielding a policy-iteration interpretation and finite-iterate policy continuity across preferences. We instantiate the update as a deep actor-critic algorithm preserving previous-policy regularization. On eight MO-Gymnasium tasks, it achieves the best average hypervolume rank among recent baselines and strong expected-utility performance. Continuous-control experiments indicate gains beyond the discrete-action setting.
A Stratified Multi-Rater Evaluation of LLM-Based Virtual Standardized Patients with a Deployed Data-Generation Platform
Yongjin Yi ⋅ JIn Yong Park ⋅ Sejoong Kim
The Objective Structured Clinical Examination (OSCE) is widely used to assess clinical communication skills in medical education, requiring trainees to take histories and reason clinically with standardized patients (SPs). Large language model (LLM)-based virtual standardized patients (VSPs) offer a practical alternative to human SPs. These are typically built using persona-injected prompting, which casts a model into a clinical role via a single system prompt. However, existing evaluations of such VSPs often rely on a single backbone LLM, small human rater panels, or LLM-as-a-judge scoring without an utterance-aligned human-expert reference. Consequently, they are limited in their ability to identify the specific shortcomings of individual models. In this study, we evaluated four LLMs (GPT-4o, Claude 3.5 Sonnet, Llama-3.1 70B, and HyperCLOVA X)—spanning commercial, open-source, and Korean-developed models—against a human-expert reference across three Korean OSCE cases (hematuria, fever in pregnancy, and chronic cough). In a double-blind, fully crossed design, 17 senior medical-student raters scored every model utterance on three sentence-level and six encounter-level Likert metrics, yielding 7,480 data points. We analyzed the scores using a stratified linear mixed-effects model based on an 11-category history-taking and 3-level linguistic-form taxonomy. At the encounter level, three of the four LLMs matched the human-expert reference on Consistency and Comprehensibility. HyperCLOVA X additionally matched the reference on Engagement, whereas Llama-3.1 70B underperformed the human reference across all six metrics. Furthermore, we introduce a web-based OSCE platform that supports virtual practice across diverse clinical cases. This platform continuously collects student–VSP dialogues under research and privacy consent; the resulting dataset is available at https://huggingface.co/datasets/neurips-2026-cpx/neurips-2026-cpx.
A Structure-Aware Higher-Order Message Passing Framework on Walk States for Graph Classification
Lin Du ⋅ Lu Bai ⋅ Lixin Cui ⋅ Ming Li ⋅ Bo Jiang ⋅ Hangyuan Du ⋅ Ziyu Lyu ⋅ Xin Jin
Standard Message Passing Neural Networks (MPNNs) are at most 1-WL expressive, which has spurred the development of higher-order models. In particular, many higher-order designs rely on extra global structural features (e.g., positional encodings or pairwise distances), whereas the $k$-WL hierarchy offers a principled and quantifiable path to increased expressiveness. However, directly simulating $k$-WL is computationally prohibitive, and existing $k$-WL variants often sacrifice expressiveness to gain efficiency, failing to capture critical structural distinctions. To overcome these shortcomings, we propose a higher-order graph learning framework based on walk-induced lifted states with structure-aware state encoding. We further design a controllably sparsified and localized $k$-FWL-style aggregation scheme, enabling graph-level higher-order representations via simple message passing. On the recent and challenging BREC expressiveness benchmark, our model achieves state-of-the-art total distinguishing accuracy among compared higher-order GNN baselines, outperforming strong 3-WL baselines, and remains competitive on real-world graph classification tasks.
A Structured LLM Framework for Inorganic Material Synthesis Planning
Heewoong Noh ⋅ Gyoung S. Na ⋅ Namkyeong Lee ⋅ Chanyoung Park
Material synthesis planning (MSP) remains a fundamental and underexplored bottleneck in AI-driven materials discovery, as it requires not only identifying suitable precursor materials but also designing coherent sequences of synthesis operations to realize a target material. Although several AI-based approaches have been proposed to address isolated subtasks of MSP, a methodology for solving the entire MSP task has yet to be established. We propose MSP-LLM, a structured LLM-based framework that formulates MSP as a two-stage process composed of two constituent subproblems: precursor prediction (PP) and synthesis operation prediction (SOP). Our approach introduces a discrete material class as an intermediate decision variable that organizes both tasks into a chemically consistent decision chain. For SOP, we further incorporate hierarchical precursor types as synthesis-relevant inductive biases and employ precursor constraint factorization that preserves precursor-related information in the autoregressive decoding state. Extensive experiments show that MSP-LLM consistently outperforms existing methods on both PP and SOP, as well as on the MSP task, demonstrating an effective and scalable framework with practical potential for autonomous MSP.
A systematic evaluation of vision-language models for observational astronomical reasoning tasks
Wenke Ren ⋅ Hengxiao Guo ⋅ Wenwen Zuo ⋅ Xiaoman Zhang
Vision-language models (VLMs) are increasingly proposed as general-purpose tools for scientific data interpretation, yet their reliability on real astronomical observations across diverse modalities remains untested. We present AstroVLBench, a benchmark of over 4,100 expert-verified instances across five tasks spanning optical imaging, radio interferometry, multi-wavelength photometry, time-domain light curves, and optical spectroscopy. We systematically evaluate six state-of-the-art frontier models and find that performance is strongly modality-dependent and uncorrelated with general benchmark rankings. Mechanistic ablations reveal three concrete bottlenecks. First, prompts that ground attention in physical mechanisms (why a feature matters) yield more balanced classifications than phenomenological prompts (what to look for). Second, presenting one-dimensional measurements as numerical tables rather than rendered plots improves accuracy by up to 13 percentage points, indicating that plot rendering can obscure physically relevant signal. Third, a qualitative reasoning analysis reveals a pervasive “right-answer-wrong-reason” phenomenon: models routinely reach correct predictions through physically invalid logic, showing that accuracy alone is insufficient for trustworthy scientific deployment. AstroVLBench provides the first systematic, multi-modal baselines for VLMs in observational astronomy and identifies the representation, grounding, and reasoning bottlenecks that must be addressed before VLMs can be trusted as autonomous scientific interpreters.
A Theory of Spatial Continuous Attractors in Hopfield Energy Landscapes
Chong Li ⋅ Xiangyang Xue ⋅ Jianfeng Feng ⋅ Taiping Zeng
Spatial cognition requires stable neural manifolds of orientation and location that track self-motion and sensory cues. Continuous attractors provide a natural dynamical principle, but learning such mechanisms that form and control manifolds with theoretical guarantees remains difficult. We propose Spatial Energy Attractor Learning (SEAL), a theory of neurodynamics learning for spatial cognition, where continuous spatial representations arise as learned low-energy manifolds in Hopfield energy landscapes through population-level Fourier energy learning. For head-direction and grid-cell systems, Fourier energy landscapes yield ring and torus attractors with learned normal attraction and input-driven tangent transport. We prove that Fourier construction is a special case of a general Hopfield-compatible energy extension whose autonomous dynamics preserve energy descent and whose expressive energy families approximate target attractor energies and restoring fields. Experiments on HD and grid-cell populations show that the learned energies restore perturbed states to ring and torus manifolds, stabilize velocity-driven moving bumps across multiple input regimes, and induce continuous-attractor interactive structures with local excitation and surround inhibition. Together, these results provide a stable, learnable, and biologically interpretable energy-based account of continuous spatial representation.
ATI-VLA: Action-Centric Predictive Vision–Language–Action Models via Actionable Alignment Then Adaptive Injection
Yijie Zhu ⋅ Rui Shao ⋅ Jie He ⋅ Wei Li ⋅ Bo Zhao ⋅ Yelin Wang ⋅ Xiaochen Yuan ⋅ Tao Tan ⋅ Miao Zhang ⋅ Xiaojiang Peng ⋅ Zitong YU
Predictive Vision–Language–Action (VLA) models aim to improve robotic manipulation via future observation or world dynamics forecasting. However, existing approaches often fail to realize this potential and underperform direct action prediction models. We argue that these limitations stem from modality misalignment between observations and actions, together with joint optimization conflicts that drive learning away from an action-centric objective. To this end, we introduce ATI-VLA, an Action-Centric Predictive Vision–Language–Action framework via Actionable Alignment Then Adaptive Injection. Specifically, it follows a two-step design: 1) Actionable Representation Alignment via a Shared Codebook. It aligns predictive observation and action representations by mapping both modalities into a shared discrete latent space via a unified codebook, making predictive observation latents readily usable for action generation and mitigating modality misalignment. 2) Action-Centric Adaptive Injection of Predictive Latents. Building upon this, it then injects predictive observation latents into action decoding as explicit predictive priors via a lightweight adaptive side-path, enabling adaptive predictive guidance under a single action-centric objective. Extensive experiments on both simulation and real-world robotic tasks demonstrate that ATI-VLA achieves state-of-the-art performance with faster convergence.
AtomWorld-Mem: Memory-Restored World States for Long-Horizon Atomistic Evolution
Tian Luo ⋅ Ruge Zhang ⋅ Haozhi Han ⋅ Yifeng Chen ⋅ Yunquan Zhang ⋅ Ting Cao ⋅ Yunxin Liu ⋅ Kun Li
High-fidelity atomistic evolution over long timescales requires more than observing the current crystal configuration. Instantaneous atomistic snapshots are often incomplete: locally similar configurations can correspond to different hidden dynamical contexts, future event preferences, and waiting-time scales. We argue that this snapshot ambiguity makes long-horizon atomistic evolution fundamentally a memory-based world-state restoration problem. To address this, we introduce AtomWorld-Mem, a memory-restored atomistic world model that recovers the latent world state missing from instantaneous crystal snapshots. AtomWorld-Mem treats the evolving alloy as an AtomWorld: spatial encoders write multi-scale atomistic keyframes from dense local topology and sparse long-range defect context, while short-term event memory and long-term structural memory integrate these keyframes across time to restore a future-predictive evolutionary state. The restored state is used to prioritize legal vacancy-mediated events under single-event Kinetic Monte Carlo (KMC) constraints, while event legality, physical execution, and residence-time updates remain governed by the underlying simulator. Empirically, AtomWorld-Mem improves long-horizon atomistic progress under fixed microscopic event budgets while maintaining high-fidelity evolution across energetic, structural, and vacancy-transport observables. It further transfers zero-shot across diverse unseen alloy-temperature AtomWorlds, suggesting that the learned memory-restoration mechanism captures reusable principles of hidden-state inference rather than a system-specific local energy heuristic. These results position memory-restored world-state modeling as a promising route toward efficient, physically grounded, and transferable atomistic evolution.
Attending on Attention ($A^2$): Smaller Self-Supervised ViTs Localize Better Than Larger Ones
Sreehari Rammohan ⋅ Huy Ha ⋅ Carl Vondrick
Robust visual classification often depends on localizing the main foreground objects in an image while ignoring spatially-separable distractors. Surprisingly, we find that the attention maps of smaller self-supervised ViTs localize foreground objects better than those of larger ones. However, we still need large ViTs, because they extract richer representations from each patch. To get the best of both worlds, good localization _and_ rich representations, we propose $A^2$, a simple method that leverages this inverse scaling finding by decoupling _where to look_ (a small attention model) from _what to extract_ (a large embedding model): we crop around the attention peaks of a small model and embed the crops with a larger model. $A^2$ uses entirely pretrained features and does not require per-dataset attention or backbone training. Across $5$ benchmarks, $A^2$ is competitive with backbone-matched loss-level methods like DFR, and outperforms end-to-end attention training under stronger distribution shifts.
Attention-based Routing for Interpretable Multimodal Brain Encoding
Pinyuan Feng ⋅ Hossein Adeli ⋅ Richard Antonello ⋅ Ethan Hwang ⋅ Nikolaus Kriegeskorte
Understanding how the brain integrates information across sensory modalities under naturalistic settings remains a central challenge in cognitive computational neuroscience. Despite recent progress in multimodal brain encoding driven by the advancement of deep learning, many existing approaches combine representations from pretrained modality-specific models using static aggregation or linear readouts, and therefore do not explicitly capture how information is selectively integrated across cortical regions. Here, we introduce Multimodal Transformer Brain Encoder (M-TBEn), a neural encoding framework that models multimodal integration as a routing problem. M-TBEn extends prior vision-based brain encoding approaches by introducing a cross-attention mechanism for multimodality, in which learnable parcel-level queries dynamically select and aggregate modality-specific representations from visual, auditory, and linguistic inputs. This design is not only parameter-efficient but also yields parcel-specific multimodal representations that support a flexible linear readout for predicting brain responses at various spatial resolutions. We evaluate M-TBEn on naturalistic video-viewing data, including both in-distribution and out-of-distribution stimulus conditions. The model is able to accurately predict parcel-level and voxel-level fMRI responses and generalizes across stimulus distributions. Furthermore, the learned attention patterns provide a structured descriptive basis for analyzing modality integration across cortical parcels, enabling a connection between parcel-specific routing profiles and established functional brain networks. Together, these results suggest that attention–based routing offers a principled computational framework for modeling multimodal integration in the human brain.
Attention Itself Could Retrieve. RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity Retrieval
Zichen Zou ⋅ Xiaosong Jia ⋅ Zuxuan Wu ⋅ Yu-Gang Jiang
Visual Geometry Grounded Transformer (VGGT) advances 3D reconstruction via scalable Transformer architecture, but the quadratic complexity of global attention prevents long context application. StreamVGGT enables streaming with causal attention, yet its KV cache grows linearly with frames, causing memory overflow and quality degradation. We present RetrieveVGGT, a training-free framework, which formulates context construction for VGGT as a retrieval problem. By retrieving a fixed number of relevant frames at each step, VGGT maintains a controllable memory budget, which is close to its training context length. Interestingly, we find that the similarity between current frame queries and cached history frame keys at the first global attention layer of VGGT is already a strong indicator of relevance, eliminating the need for additional learned scoring. To enhance information diversity similar to a recommender system, we propose Segment Sampling so that the retrieval spans distinct relevant segments rather than a single high-similarity region. We design a pose-aware spatial memory mechanism that organizes history frames according to their already estimated camera poses, enabling location-aware retrieval. Extensive experiments demonstrate that RetrieveVGGT achieves state-of-the-art performance, outperforming StreamVGGT, TTT3R, and InfiniteVGGT while maintaining constant memory usage regardless of sequence length.
Attention Sinks as Spectral Spikes: A Mechanism Analysis of Gated Attention
Seojin Kim ⋅ Yehjin Shin ⋅ Noseong Park
Transformer attention exhibits attention sinks, where a small subset of source tokens absorbs disproportionate attention mass across layers. Recent work shows that post-SDPA output gating strongly reduces attention-sink behavior, while separate studies link attention sinks to gradient sinks and massive activations through backward training dynamics. Yet how the same gate affects forward sink structure, backward gradient concentration, and training-time representation signatures remains unclear. We address this gap by formulating attention sinks as sink-induced low-rank modes in token-space covariance, with directions aligned to high-mass attention columns. Under this view, post-SDPA gating acts as a mode-dependent contraction: it attenuates sink-induced spectral energy in the forward pass and locally reduces value-path gradients routed through sink columns in the backward pass. Empirically, random-matrix diagnostics show that gating reduces outlier mass, weakens sink-subspace alignment, and increases effective rank, while backward analyses show active-gate attenuation of first-token value gradients. Training-time diagnostics further show that gating suppresses the co-emergence of attention concentration, value-gradient amplification, and representation dominance.
Attention Transfer Is Not Universally Effective for Vision Transformers
Huaiyuan Qin ⋅ Muli Yang ⋅ Gabriel James Goenawan ⋅ Peng Hu ⋅ Chen Gong ⋅ Xi Peng ⋅ Hongyuan Zhu
A recent work shows that Attention Transfer, which transfers only the attention patterns from a pre-trained teacher Vision Transformer (ViT) to a randomly initialized standard student ViT, is sufficient to recover the full benefit of the teacher's pre-trained weights. We revisit this finding on a comprehensive benchmark of 20 teachers from 11 well-known ViT families and reveal that Attention Transfer is not universally effective. While 7 families transfer successfully, 4 consistently fail, falling up to 5.1\% below the from-scratch no-transfer baseline. Further results demonstrate that this failure is family-consistent across model sizes, and persists under extended training durations, different transfer datasets, and out-of-distribution evaluations. Controlled analyses then consistently localize the problem to the attention-routing channel, indicating that the key issue is not whether the student can match the teacher's attention patterns, but whether the matched patterns remain functional for the student. Crucially, we identify architectural mismatch between the pre-trained teacher and the standard student as the primary mechanism. By adding only the teacher's native architectural components to the student in a randomly initialized state, we completely reverse the failure for all 4 families. Notably, these components alone do not improve from-scratch training, confirming that they specifically unlock the usability of the teacher's attention. We further systematically show that this failure is not explained by the inadequate choice of transfer loss or by differences in pre-training recipes. Our findings refine the prevailing understanding of attention in ViT representations: attention is sufficient only when the student architecture matches the teacher.
Auditing Instruction Robustness in Vision-Language-Action Models via Diversity-Aware Red Teaming
Baoshun Tong ⋅ Haoran He ⋅ Yang Liu ⋅ Ling Pan ⋅ Liang Lin
Vision-Language-Action (VLA) models have achieved remarkable success in robotic manipulation. However, their robustness to instruction variations remains a critical, under-explored safety concern, posing a significant safety risk to real-world deployment. Red teaming, or identifying environmental scenarios that elicit catastrophic behaviors, is an important step in ensuring the safe deployment of embodied AI agents. Reinforcement learning (RL) has emerged as a promising approach in automated red teaming that aims to uncover these vulnerabilities. However, standard RL-based adversaries often suffer from severe mode collapse due to their reward-maximizing nature, which tends to converge to a narrow set of trivial or repetitive failure patterns, failing to reveal the comprehensive landscape of meaningful risks. To bridge this gap, we propose a novel \textbf{D}iversity-\textbf{A}ware \textbf{E}mbodied \textbf{R}ed \textbf{T}eaming (\textbf{DAERT}) framework, to audit VLA robustness under semantically aligned instruction. Our design uses a breadth-seeking value estimator that prevents the attacker from collapsing onto a single high-reward phrasing, generating a diverse set of challenging instructions while preserving attack effectiveness, measured by execution failures in a physical simulator. We conduct extensive experiments across different robotic benchmarks against two state-of-the-art VLAs, including $\pi_0$ and OpenVLA. Our method consistently discovers a wider range of more effective adversarial instructions that reduce the average task success rate from 93.33\% to 5.85\%, demonstrating a scalable approach to stress-testing VLA agents and exposing critical safety blind spots before real-world deployment.
Anatomical mesh segmentation requires models that operate directly on irregular surface geometry while remaining robust to arbitrary patient pose and mesh resolution variation. Existing task-specific mesh and point-cloud methods are not equivariant, and can degrade sharply under test-time perturbation, for example dropping by 25-26 IoU points on intraoral scan segmentation at 40$\textdegree{}$ tilt. We present EAMS, an Equivariant Anatomical Mesh Segmentor built on Equivariant Mesh Neural Networks (EMNN), and evaluate it across four clinically distinct tasks spanning edge-, vertex-, and face-level supervision. We combine intrinsic mesh descriptors with anatomy-aware priors, including PCA-derived frames for dental arches and liver surfaces, and augment message passing to provide lightweight global context. Across intracranial aneurysm and intraoral segmentation, EAMS variants are competitive with specialized baselines on unperturbed inputs while remaining stable under geometric perturbations, and on liver surfaces they expose a favorable trade-off between canonical-pose accuracy and rotation robustness. These results show that a lightweight ($<2$M parameters) equivariant framework can deliver robust anatomical mesh segmentation across diverse supervision types without task-specific architectures.
Anchored fixed point and monotone equation methods, including Halpern iteration, extra anchored gradient, and their relatives, add a vanishing pull toward a reference point to obtain last-iterate guarantees. Existing anchored variants often achieve sharp last-iterate guarantees, but from the update-level perspective the placement of the anchor can be algorithm-specific and conceptually opaque. We show that anchoring admits a single operator-side construction: regularize the operator queried by the base method with a vanishing Tikhonov term, then run the unmodified base method. Applied to the Picard iteration, this recipe reproduces the Halpern iteration; applied to the forward step, extragradient, and Popov methods, it yields three variants whose anchor placements inherit the base method's query pattern. The four analyses share a residual recurrence, recovering the (O(1/k)) Halpern residual-norm convergence rate, giving (O(1/\sqrt{k})) for the regularized forward step, and giving (O(1/k)) for the regularized extragradient and Popov variants in the unconstrained monotone Lipschitz setting.
AutoDataBench: How Far Are LLM Agents from Autonomously Engineering Post-Training Data Pipelines?
Qiaoyu Tang ⋅ Hao Xiang ⋅ Le Yu ⋅ Yaojie Lu ⋅ Xianpei Han ⋅ Le Sun ⋅ Bowen Yu ⋅ Jiawei Chen ⋅ Zhenru Zhang ⋅ Peng Wang ⋅ Que Shen ⋅ Hongyu Lin ⋅ Dayiheng Liu
Post-training data pipelines are traditionally hand-designed by researchers who orchestrate generation, verification, and diversity balancing. While Large Language Models (LLMs) can already execute individual subtasks, the architectural composition of these pipelines remains a manual bottleneck, excluding the most substantial engineering step in post-training from automation. We introduce AutoDataBench, a benchmark evaluating whether LLM agents can autonomously design post-training data-synthesis pipelines. Each task gives the agent only a description of a target capability and asks it to produce a complete, executable pipeline whose generated corpus is then used to fine-tune a fixed student model and scored on the evaluation. Each task is delivered as a natural-language instruction and a sandboxed workspace given to the agent, paired with an evaluation protocol. We instantiate AutoDataBench on five capabilities: instruction following, competition math, deep search, repo-level software engineering, and terminal automation. For each task we pair the evaluation with a strong human-designed pipeline whose downstream score serves as a reference point for the agent. Experimental results reveal that frontier agents close most of the gap to the human reference on text-centric tasks (e.g., Opus-4.7 nearly matches humans in instruction following), yet a noticeable performance gap persists in tasks requiring executable-environment construction (SWE and Terminal). Furthermore, we observe a decoupling between downstream task-solving strength and pipeline-construction quality, where strong reasoning models like GPT-5.5 may still produce suboptimal pipelines. These findings suggest that autonomous environment synthesis remains the primary bottleneck for self-improving data loops, distinct from mere task-solving competence.
AutoManifold: Agentic Design of Data Visualisation Algorithms via Manifold Embedding
Burak Susam ⋅ ruiyuan kang ⋅ Tingting Mu
Manifold embedding for data visualisation is dominated by a small set of hand-designed state-of-the-art (SOTA) algorithms. Their design targets are fixed at development time, resulting in inflexible algorithmic behaviour that cannot be readily steered toward user-specified structural priorities. Hand-designing new embedding algorithms to address user priority requires highly specialised expertise and is time consuming. To address autonomous algorithm design tailored to user preference, we introduce an agentic algorithm generation pipeline AutoManifold. It composes new manifold embedding algorithms from a constrained vocabulary of affinity, cost, and optimisation primitives, supported by multi-agent large-language-model (LLM) orchestration. AutoManifold conditions every stage of its design on user-specified structural preservation preferences, and iteratively refines the algorithm configuration through an LLM-guided, metric-grounded iterative loop. We compare the generated algorithms against three strongest and most frequently used SOTA (t-SNE, UMAP, and PaCMAP) on real-world datasets spanning different difficulty regimes under identical evaluation infrastructure. The results demonstrate that LLM-driven, priority-conditioned algorithm synthesis can move beyond hyperparameter tuning, producing genuinely new algorithms that outperform traditional hand-designed methods, in their ability to steer towards user-specified properties without compromising much inherent structure in the original data.
Autoregressive Learning in Joint KL: Sharp Oracle Bounds and Lower Bounds
Yunbei Xu ⋅ Yuzhe Yuan ⋅ Ruohan Zhan
We study the fundamental and timely problem of learning long sequences in autoregressive modeling and next-token prediction under model misspecification, measured by the joint Kullback--Leibler (KL) divergence. Our goal is to characterize how the sequence horizon (H) affects both approximation and estimation errors in this joint-distribution, sequence-level regime. By establishing matching upper and lower bounds, we provide, to our knowledge, the first complete characterization of long-horizon error behavior under the natural joint KL objective, with improved rates and optimality justification relative to existing work. On the approximation side, we show that joint KL admits a horizon-free approximation factor, in sharp contrast to Hellinger-based analyses that exhibit an (\Omega(H)) dependence for computationally efficient methods; this isolates the choice of divergence as the source of approximation amplification. On the estimation side, we prove a fundamental information-theoretic lower bound of order (\Omega(H)) that holds for both decomposable policy classes and fully shared policies, matching the (\widetilde O(H)) upper bounds achieved by computationally efficient algorithms. Our analysis clarifies the landscape of recent autoregressive learning results by aligning the log-loss training objective, the sequence-level evaluation metric, and the approximation metric through a sharp joint-KL oracle theory. We further show that these joint-KL guarantees imply policy learning regret bounds at rates matching prior imitation learning literature.
BALTO: Balanced Token-Level Policy Optimization for Hallucination Mitigation
Ning Li ⋅ Zixuan Guo ⋅ Yan Xu ⋅ Wenbo Fei ⋅ Yifan Niu ⋅ Chang Luo ⋅ Yasheng Wang ⋅ Weiwen Liu ⋅ Yong Yu ⋅ Weinan Zhang
Hallucinations remain a major obstacle to deploying large language models (LLMs) in knowledge-intensive settings, where generated responses must be faithfully grounded in provided evidence. Reinforcement learning (RL) is a promising direction for hallucination mitigation, but response-level faithfulness rewards suffer from a granularity mismatch: localized hallucinations can cause supported content to receive spurious penalties. Although recent work introduces fine-grained feedback such as claim-level verification and token-level rewards, unbalanced credit assignment can still induce length, verbosity, or optimization-noise biases. We propose $\textbf{BALTO}$, a $\textbf{Bal}$anced $\textbf{T}$oken-level Policy $\textbf{O}$ptimization framework for hallucination mitigation. BALTO extracts checkable factual claims, verifies them against the reference context, and projects claim-level judgments to token-level labels. A balanced token-level credit assignment mechanism is introduced into the framework. This design redistributes probability mass from unsupported content toward faithful content, rather than suppressing the entire response. We systematically analyze the limitations of response-level rewards from a theoretical standpoint, and prove BALTO’s advantages in training stability and optimization efficiency for hallucination mitigation. Experiments on ConFiQA, RAGTruth, and FinLLM-Eval show that BALTO achieves the highest faithfulness across all six model--benchmark settings and consistently outperforms existing post-training baselines in Q-Score, demonstrating a stronger faithfulness--informativeness trade-off.
Baton: Explicit Semantic Blueprints for Joint Video-Audio Generation
Shuyuan Tu ⋅ Qi Tian ⋅ Zihan Yang ⋅ Yue Wu ⋅ Xintong Han ⋅ Weijie Kong ⋅ Jiangfeng Xiong ⋅ Jian-Wei Zhang ⋅ Zhao Zhong ⋅ Liefeng Bo ⋅ Zuxuan Wu ⋅ Yu-Gang Jiang
Current open-source diffusion models struggle to generate stable and synchronized audio-visual content, particularly in scenarios demanding complex semantic reasoning. The root cause is that existing methods rely on coarse text embeddings from off-the-shelf encoders to guide audio-video denoising, which discards fine-grained semantics and, critically, lacks a shared long-horizon plan, leading to uncoordinated denoising trajectories and fragile cross-modal alignment. We propose Baton, the first framework that introduces explicit semantic planning into joint video-audio generation. Our key insight is that complementing coarse text guidance with semantically rich, modality-aware planned tokens, jointly reasoned and mutually aligned before denoising, can simultaneously restore fine-grained semantic detail and establish a shared blueprint that coordinates both audio and video denoising trajectories. Concretely, Baton first introduces the VA-Planner, a multimodal language model equipped with dual semantic alignment towers, where learnable queries cross-attend to both video and audio features to produce a pair of semantically aligned video and audio planning tokens as keyframe-level blueprints. These planning tokens are injected into the diffusion backbone via cross-attention layers, providing temporally grounded guidance complementary to coarse text embeddings. Since planning tokens do not share one-to-one spatial-temporal correspondence with diffusion latents, we further propose Relative Semantic RoPE, a relative positional encoding that maps planning tokens and latents into a shared spatial-temporal coordinate frame, enabling each latent to accurately attend to its positionally corresponding semantic cues. Experiments on benchmarks show the effectiveness of Baton both qualitatively and quantitatively.
Modern ML systems increasingly use uncertainty to decide when to abstain, route to a fallback, retrieve more context, or spend additional computation. Yet many strong Bayesian deep-learning recipes make uncertainty expensive on the serving path: ensembles, posterior samples, and MC dropout require repeated forward passes, while fast deterministic confidence heads are not generally posterior-predictive computations over parameter beliefs. We introduce \emph{Bayesify}, a message-passing view that turns a deterministic neural computation graph into a single-pass posterior-predictive moment graph. Trainable tensors become Gaussian beliefs, activations carry selected moments, and prediction is one deterministic forward propagation of those moments. Theoretically, we show that exact belief propagation is natural-gradient message passing in exponential-family mean space, and that projected EP is implemented by reverse-mode adjoints followed by local Fisher moment updates. This yields a practical wrapper for dense, convolutional, residual, and adapter modules. Experiments evaluate calibrated uncertainty, selective prediction, and OOD behavior at one-pass inference cost, comparing against deterministic calibration, ensembles, dropout, Laplace, SWAG, and variational baselines.
BEAKER: An Expert-Curated Benchmark for Embodied Brains in Self-Driving Chemical Laboratories
Fei Lin ⋅ Tengchao Zhang ⋅ Ziyang Gong ⋅ Xiaotong Yu ⋅ Bohan Zhang ⋅ Yifan Zhou ⋅ Qihao Yang ⋅ Ji Dai ⋅ Dong Li ⋅ Yining Jiang ⋅ Qinghua Ni ⋅ Jun Huang ⋅ Qiang Zhang ⋅ Yue Zhou ⋅ Yonglin Tian ⋅ Zhihang Zhong ⋅ Xue Yang ⋅ Fei-Yue Wang
Self-Driving Chemical Laboratories (SDCLs) are moving chemical experimentation toward embodied automation, where machines must perceive laboratory scenes, reason about experimental states, and act under physical, procedural, and safety constraints. Although Multimodal Large Language Models (MLLMs) are increasingly considered as embodied brains for such systems, existing laboratory benchmarks mainly focus on safety, anomaly detection, or general scientific understanding, leaving their execution-oriented embodied capabilities underexplored. To fill this gap, we introduce BEAKER, a Benchmark for evaluating Embodied Actionability, Knowledge, Experimentation, and Reasoning in SDCLs. BEAKER contains 1,000 expert-curated QA samples across 11 image and video task types, covering key judgments required for chemical laboratory execution, including affordance grounding, trajectory planning, equipment and state understanding, apparatus connection, operation process reasoning, and result understanding. All questions and annotations are manually created and reviewed by chemistry and embodied AI experts, without AI-assisted generation or annotation. We evaluate 28 general-purpose, scientific, chemical, and embodied MLLMs on BEAKER. Results show that current models can recognize laboratory objects and states, but still struggle to ground experimental knowledge into operable regions, feasible trajectories, and long-horizon procedural reasoning. BEAKER thus reveals a clear gap between multimodal laboratory perception and action-grounded experimental understanding, providing a diagnostic benchmark for future MLLMs in SDCLs.
Behavioral Foundation Models for Quality Diversity
Nazim Bendib ⋅ Nicolas Perrin-Gilbert ⋅ Olivier Sigaud
Behavioral Foundation Models (BFMs) are an emerging paradigm in reinforcement learning, playing a role analogous to large language models in natural language processing: they have shown remarkable versatility, enabling zero-shot performance, fast imitation, and online adaptation, all by exploiting the structure of a latent space. In this work, we investigate whether the latent behavioral space induced by BFMs can serve as an effective search space to discover large repertoires of behaviorally diverse and high-performing policies through Quality-Diversity (QD) methods. While QD methods generally search directly in high-dimensional policy parameter space, in this paper, we present BFM-QD, a framework that performs QD search in the compact latent space of a BFM. We further show that the BFM-QD framework provides a closed-form, gradient-free policy improvement operator that approximates a policy gradient update, but requires no critic training and no backpropagation. Across continuous-control benchmarks spanning dense locomotion, sparse navigation, and contact-rich manipulation, BFM-QD consistently outperforms parameter-space baselines, with particularly stark gains in sparse and deceptive settings, where all existing QD methods collapse to near-zero performance. These results show the effectiveness of the BFM-QD framework, benefiting from the synergy between dimensionality reduction of the search space and offline pretraining from diverse behavioral data. This positions BFMs as a general-purpose backbone for QD optimization, extending their utility beyond zero-shot task solving to the discovery of diverse behavioral repertoires.
Behavior Cloning is Not All You Need: The Optimality of On-Policy Distillation for Noisy Expert Feedback
Ved Sriraman ⋅ Peihan Liu ⋅ Daniel Hsu ⋅ Adam Block
Imitation Learning (IL) is a natural framework for learning in sequential decision-making systems and has emerged as the dominant paradigm through which we understand language model training. A central puzzle is that, while in theory offline IL can be horizon-free and optimal, in practice online methods such as on-policy distillation (OPD) often outperform offline methods such as supervised fine-tuning (SFT). We propose a noisy expert model to explain this gap, in which the learner only has access to a noisy version of the expert's policy, but wishes to compete against the reward achieved by a clean expert, motivated by the fact that in many applications, e.g. training language models to perform long chains of thought, the expert is often imperfect. In this setting, we show a sharp separation between offline and online IL. Offline learning from noisy trajectories is fundamentally hard: to compete with the clean expert, the sample complexity must grow exponentially, in contradistinction to the clean expert setting where no explicit horizon dependence exists. In contrast, we prove that online interaction with the noisy expert via a novel variant of OPD enables horizon-free guarantees in some settings and polynomial dependence on horizon in general. Our analysis leads to an alternative loss function form that is commonly considered empirically for LM training. We further provide algorithms and lower bounds, and extend our results to the more realistic setting of unknown corruption when the clean expert is deterministic, thereby providing a theoretical foundation for why on-policy distillation can outperform standard supervised fine-tuning when training language models from imperfect teachers. We complement our theoretical results with experiments on synthetic and natural-language tasks, showing that the OPD variant suggested by our theory outperforms both offline BC and existing OPD objectives under noisy expert feedback.
Behavior Pack Optimization for Video MLLM Post-Training
Zhaolu Kang ⋅ Shiyu Liu ⋅ Tailong Luo ⋅ Wei Zhang ⋅ Yingjie He ⋅ Guangyuan Dong ⋅ Siheng Wang ⋅ Liang He ⋅ Lei Wei ⋅ Jiaqi Su ⋅ Shuang Chen ⋅ Guansu Wang ⋅ Haoyu Ji ⋅ Qishi Zhan ⋅ Kaiyue Zhou
Video multimodal large language models (MLLMs) keep climbing video question answering benchmarks, yet shuffling the frames, masking the segment that supports the answer, or occluding the target object barely changes their predictions. The accuracy rests on appearance and language priors, not on the temporal evidence the question asks for. We trace this to the unit of post-training: rewards are computed on a single response to the original clip, so the model is never asked to behave consistently across views. We propose Behavior Pack Optimization (BPO), which replaces the single response with a behavior pack of outputs across counterfactual views chosen by question type, scored jointly. The pack reward asks for stability when the intervention is irrelevant, sensitivity when key evidence is removed, and abstention when no evidence remains. To keep this objective stable at small pack sizes, BPO uses an anchor-relative advantage: the response on the original view serves as a per-prompt reference instead of a group mean over mixed views. On TempCompass, MVBench, and NExT-QA, BPO improves the macro accuracy of Qwen2.5-VL-7B-Instruct by 4.7 pp, the temporal-hard subset by 7.8 pp, and abstention F1 by 20.0 pp over a budget-matched vanilla GRPO baseline from the same SFT checkpoint. The gains transfer to Video-MME, LongVideoBench, and to LLaVA-Video-7B; ablations confirm they follow the view sets, not the rollout count. We hope this pack-level perspective offers a useful starting point for the video MLLM and multimodal post-training community as the field moves toward evidence-grounded video reasoning.
Benchmarking Vietnamese Legal Knowledge of Large Language Models
Dong N Tien ⋅ Nguyen Minh-Anh ⋅ Thanh D Hoang ⋅ Nguyen T Ngoc ⋅ Quang M Xuan ⋅ Phan P Hai ⋅ Nguyen T Anh ⋅ Binh Vu ⋅ Dung Le
The rapid advancement of large language models (LLMs) has expanded their potential in the legal domain. However, existing legal benchmarks remain largely English-centric and oriented toward common law, leaving a critical gap in evaluating LLMs for civil law systems that govern most jurisdictions worldwide. To address this gap, we introduce Vietnamese Legal Benchmark VietLegal, a cognitively grounded benchmark designed for the hierarchical and codified structure of Vietnamese law. Although instantiated in Vietnamese legislation, VietLegal provides a replicable evaluation framework for civil law systems characterized by complex statutory hierarchies and frequent amendments. Inspired by Bloom’s taxonomy, VietLegal assesses multiple levels of legal understanding through tasks that mirror real-world legal assistant use cases, including legal question answering, multi-step reasoning, and scenario-based problem solving. The benchmark contains 10,450 expert-annotated samples, each cross-validated against authoritative legal sources to ensure fidelity to practical legal workflows. By offering the first standardized legal benchmark for Vietnamese, VietLegal enables systematic assessment of LLMs in civil law contexts and supports the development of more reliable and interpretable AI-assisted legal systems.
Benchmarks Are Not Atomic: Composition-Aware LLM Evaluation using BenchHub
Eunsu Kim ⋅ Haneul Yoo ⋅ Guijin Son ⋅ Hitesh Patel ⋅ Amit Agarwal ⋅ Alice Oh
LLM benchmarks are often treated as coherent measurement units, yet they are heterogeneous collections of instances spanning diverse domains, skills, formats, and contexts. As a result, aggregate benchmark scores can conflate model capability with benchmark composition, obscuring what existing benchmarks actually cover. We introduce BenchHub, a composition-aware evaluation framework that represents benchmarks as distributions over instance-level attributes. BenchHub integrates 54 benchmarks comprising 839K samples across 10 languages. It enables researchers to inspect benchmark contents, uncover reusable coverage hidden behind benchmark names, transparently compare benchmarks, and construct controllable evaluation sets for application-aligned model selection. Using BenchHub, we show that similarly motivated benchmarks can differ substantially in internal composition, existing benchmarks often contain reusable coverage beyond their stated purposes, models exhibit fine-grained category-level performance variation hidden by aggregate scores, and model rankings can shift under different reweighting and resampling configurations. Our results motivate evaluation practices that make benchmark composition explicit, inspectable, and controllable.
BenchRep-T: A Systematic Evaluation of T-Cell Repertoire-Based Disease Diagnostics
Chiho Im ⋅ Liel Cohen-Lavi ⋅ Alejandro Buendia ⋅ Anshul Kundaje ⋅ Scott D Boyd
Adaptive immune receptor repertoire sequencing data has emerged as a promising potential modality for disease diagnosis, relying on computational methods to analyze T-cell receptor (TCR) sequences from an individual’s blood sample. Published methods rely on different cohorts, data preprocessing pipelines, and evaluation metrics, making direct comparison across methods challenging. We have developed BenchRep-T, a unified benchmark that standardizes multiple publicly available TCR repertoire datasets and evaluates nine computational approaches, spanning statistical enrichment of shared sequences, feature-engineered ensembles, deep learning, and graph-based clustering. BenchRep-T evaluates methods on four tasks: disease classification across conditions, performance scaling under restricted sequence-sampling depth, recovery of known antigen-specific driver sequences, and evaluation of sensitivity to demographic confounding. Under controlled evaluation, simple baselines prove competitive, with tree-based models trained on V- and J-gene usage and short sequence motifs approaching the classification performance of more complex methods. Our findings underscore the complexity of modeling TCR repertoire data, and show that no single method dominates across all tasks. BenchRep-T provides a framework for rigorous and reproducible evaluation of TCR repertoire classification methods to accelerate the development of immune repertoire-based diagnostics.
Best Arm Identification in Generalized Linear Bandits via Hybrid Feedback
Qirun Zeng ⋅ Xuchuang Wang ⋅ Jiayi Shen ⋅ Xutong Liu ⋅ Fang Kong ⋅ Jinhang Zuo
We study fixed-confidence best arm identification in generalized linear bandits under a hybrid feedback model: at each round, the learner may query either (i) absolute reward feedback from a single arm or (ii) relative (dueling) feedback from an arm pair, both governed by generalized linear models. We introduce a likelihood-ratio–based confidence sequence that unifies heterogeneous generalized linear observations and yields an explicit ellipsoidal confidence set under a self-concordance assumption. Building on this confidence set, we propose a hybrid Track-and-Stop algorithm that adaptively allocates queries by tracking a minimax-optimal design over a joint action space of arms and pairs. We establish $\delta$-correctness and provide high-probability upper bounds on the stopping time. We further extend the framework to a cost-aware setting that accounts for heterogeneous acquisition costs across feedback modalities. Empirical experiments demonstrate that the proposed algorithms significantly improve sample efficiency over baseline methods.
Beyond Accuracy: A Diagnostic Benchmark for Hypothesis-Driven Experiment Planning in LLM Agents
Ryo Kuroki ⋅ Amin Mansouri ⋅ Philippe Schwaller
Evaluating large language model (LLM) agents on scientific tasks typically reduces to final-task accuracy, which obscures where reasoning succeeds or breaks down. We introduce a benchmark and diagnostic evaluation framework for hypothesis-driven adaptive experimental planning, instantiated in kinetic mechanism identification. The framework goes beyond final accuracy in two ways. First, we evaluate exploration behavior by quantifying the extent to which agents exhibit space-filling or corner-seeking patterns in the experimental design space. Second, it decomposes the reasoning process into four interpretable steps; belief update, hypothesis discrimination, discriminative experiment design, and predictive reasoning under interventions, each scored independently on reasoning traces. Across ten LLM agents, we find that stronger agents outperform both adaptive (Bayesian optimization) and non-adaptive (Latin hypercube) baselines and exhibit exploration patterns inconsistent with naive heuristics. Decomposing reasoning reveals that the stronger agents score consistently well across all four steps, whereas weaker agents often succeed at belief update and hypothesis discrimination but fail at discriminative experiment design and predictive reasoning. This contrast shows that our framework identifies a specific reasoning bottleneck: weaker models can recognize plausible hypotheses and articulate their differences, but struggle to translate this understanding into concrete predictions and experimental designs. The code and dataset are available at https://github.com/krfdq48k4p-arch/HypothesisDrivenExperimentPlanningBench and https://huggingface.co/datasets/xq8wvm/HypothesisDrivenExperimentPlanningBench, respectively.
Beyond DSA: Conjugacy-based Comparison of Dynamical Systems
Prakhar Godara ⋅ Pang S Tay ⋅ Marcelo G Mattar
Comparing whether two dynamical systems implement the same computation despite differences in coordinates or measurements is a central problem in neuroscience and machine learning. Dynamical Similarity Analysis [DSA; Ostrow et al., 2023] addresses this problem by aligning finite-dimensional Koopman approximations of the two systems through an orthogonal similarity transformation. Here we show that orthogonal alignment is neither necessary nor sufficient for topological conjugacy: genuinely conjugate systems may be related by a non-orthogonal basis-transfer matrix that DSA cannot capture, while non-conjugate systems may have orthogonally equivalent Koopman operators that DSA will fail to distinguish. We then use this observation to formulate \emph{Conjugacy-based Similarity Analysis} (CSA), which restricts alignments to those induced by candidate state-space bijections rather than arbitrary orthogonal matrices. We prove that CSA's fitted alignment is the finite-data projection of the composition operator associated with the candidate bijection, and use controlled examples to show why this distinction matters when observable dictionaries are chosen explicitly or implicitly from data. Together, these results clarify what Koopman-based similarity measures must ensure to support claims of identifying conjugacies between computational systems.
Beyond Flat Walks: Compositional Abstraction for Autoregressive Graph Generation
Kaiwen Bian ⋅ Andrew H Yang ⋅ Ali Parviz ⋅ Gal Mishne ⋅ Yusu Wang
Autoregressive graph generation is commonly formulated by linearizing a graph into a sequence via a flat walk over its nodes and edges. While effective for small molecular graphs, such representations do not scale gracefully to large, structurally complex biological graphs, where higher-order organization is both dense and functionally important. To address this limitation, we introduce a hierarchical graph abstraction together with a coarse-to-fine autoregressive pipeline that separates coarsening, tokenization, and decoding, and we study sequence constructions that expose hierarchy in the tokenization. Concretely, we compare a flat representation with two hierarchical alternatives: one that explicitly factorizes communities and their interactions, and another that induces hierarchy through a depth-first traversal. Across eight benchmarks, differences are modest on simpler domains but become pronounced on larger, more structurally complex settings such as cyclic peptides, proteins, and hard conditional-generation tasks. In these regimes, flat representations degrade in validity or distributional fidelity, while the traversal-based hierarchical representation is consistently more reliable; the explicitly factorized variant is generally weaker, suggesting that how hierarchy is exposed in the sequence matters. Finally, the same coarse-to-fine representation naturally enables motif-conditioned generation without architectural changes: on larger and more structurally complex graphs, conditioning on higher-level prefixes yields strong motif retention together with high validity, uniqueness, and novelty, whereas flat representations can exhibit severe mode collapse under comparable constraints.
Beyond MNIST: Limitations of Amplitude Encoding on Quantum Classification
Xin Wang ⋅ Yabo Wang ⋅ Rebing Wu
It remains unclear whether quantum machine learning (QML) truly holds an advantage when tackling practical and meaningful tasks. In exploring this question, the importance of classical data encoding, often overlooked, becomes increasingly evident. Amplitude encoding, which can embed $2^n$ classical data into $n$ qubits, is widely used due to its apparent efficiency. However, its potential limitations for QML have yet to be fully explored. In this paper, we establish a theoretical result for the limitations caused by quantum encoding and point out the existence of a concentration phenomenon in amplitude encoding. Compared to prior work, our theoretical result operates under more general conditions and leads to a stronger conclusion. This concentration phenomenon causes the classification predictions to remain close to random guessing, regardless of the training process. Our findings shed new light on a long-standing puzzle in the field of QML: why some QML models perform well on simple datasets like MNIST but fail to generalize to more complex practical tasks. By highlighting the pivotal role of encoding design in QML, our work clearly indicates that future research needs to focus more on the design of classical data encoding to advance the effectiveness of QML.
Beyond Pixel Space: Frequency-Domain Uncertainty Estimation for Structure-Aware Diffusion Guidance
TIANQI ZHAO ⋅ Xixi Liu ⋅ Liangrui Peng ⋅ Zhengrui Xiang
Although diffusion models achieve promising image generation performance, uncertainty during the iterative denoising process can lead to visual artifacts. Existing uncertainty estimation methods typically quantify it at the pixel level. However, these methods assume independence among pixels, neglecting pixel correlations crucial for image structure. In contrast, we propose a frequency-domain uncertainty estimation method that captures structural correlations. We empirically show that samples with artifacts exhibit higher frequency-domain uncertainty, and derive a Louis-identity-based connection between the estimated uncertainty and the optimal reverse-process covariance. To this end, we develop a structure-aware diffusion sampling guidance framework. For the mean of the reverse process, the gradient of the uncertainty is used to penalize specific frequency components; for the reverse covariance, frequency-domain uncertainty is used as a proxy to modulate injected noise. Experiments on the ImageNet, LSUN-Churches, and DrawBench datasets across U-Net, U-ViT, and SD3 architectures validate our method.
Beyond RAG for Agent Memory: Retrieval by Decoupling and Aggregation
Zhanghao Hu ⋅ Qinglin Zhu ⋅ Runcong Zhao ⋅ Di Liang ⋅ Hanqi Yan ⋅ Yulan He ⋅ Lin Gui
Standard Retrieval Augmented Generation (RAG) is poorly matched to agent memory. Unlike large heterogeneous corpora, agent memory forms a bounded and coherent interaction stream in which many spans are highly correlated or near duplicates. As a result, flat top-$k$ similarity retrieval often returns redundant context, while summary-centric hierarchies can blur the subtle details that distinguish one candidate from another. We argue that agent memory should follow the principle of decoupling before aggregation: the system should first isolate reusable facts, updates, and distinguishing details from similar histories, and only then organise them for efficient retrieval. Based on this principle, we propose xMemory, which constructs a revisable hierarchical memory structure from original messages to segments, memory components, and groups. xMemory segments interaction history into local events, decouples each segment into memory components, aggregates related components into high-level groups using a sparsity--semantic faithfulness objective, and maintains this structure incrementally as memory evolves. At inference time, xMemory retrieves top-down, first selecting a compact backbone of complementary groups and components, and then expanding to segments and raw messages only when additional evidence reduces the reader's uncertainty. Experiments on LoCoMo and PerLTQA across diverse open source and closed source LLMs show consistent gains in answer quality and inference token efficiency, supported by analyses of redundancy, evidence density, and coverage.
Beyond Selection: Token Parameterization for Extreme Visual Token Compression
Rui Zhong ⋅ YU LI ⋅ Zheyu Yan ⋅ Cheng Zhuo
Visual-token compression is effective for improving the efficiency of vision-language models, but under extreme compression budgets, token pruning can break visual grounding while learned resamplers increase parameter count, attention cost, and training complexity. We revisit compression through a token parameterization lens, separating (i) basis transformation and structured truncation (retained subspace/compressibility) from (ii) coordinate organization (optimization and cross-modal alignment). This view yields two coupled objectives, compressibility and learnability, which we formalize as unified functionals. Guided by these objectives, we design Braco, a lightweight four-step coder that combines transform-basis truncation, input-independent basis-coordinate embeddings, budget-dependent orthogonal re-parameterization, and learned spatial residual tokens from lightweight pooling. Experiments show that Braco forms the favorable empirical accuracy-efficiency frontier under compression ratios from $23\times$ to $64\times$ and remains competitive at $144\times$, reaching 95.2\% accuracy while reducing prefill FLOPs by 84.2\%--86.7\% relative to the uncompressed upper bound. Against prior methods, Braco matches or improves accuracy while achieving up to $\sim$36\% end-to-end speedup and using $16.6\times$/$78.8\times$ lower compressor latency/FLOPs.
Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding
XIAO LIN ⋅ Xiaohu Huang ⋅ Kai Han
Multimodal Large Language Models (MLLMs) have demonstrated substantial promise in spatial understanding. Existing works typically incorporate prior knowledge extracted from a pre-trained foundation model to further enhance the spatial awareness of MLLMs. In this paper, we first reveal that when integrating diverse foundation models into MLLMs, each model provides complementary spatial priors that benefit different tasks. Motivated by this, we propose $\textbf{ViPS}$, a novel multi-model prior framework designed to fully unleash the potential of incorporating multiple $\textbf{Vi}$sual $\textbf{P}$riors from diverse models into MLLMs for $\textbf{S}$patial understanding. Specifically, ViPS introduces an Efficient Prior Proxy to generate multiple foundational priors with minimal inference overhead, and a Dynamic Prior Fusion mechanism to achieve harmonious and context-aware prior fusion and injection from the prior proxies. Extensive experiments demonstrate that ViPS successfully harmonizes diverse visual priors, establishing new state-of-the-art performance across multiple complex spatial reasoning and 3D spatial understanding benchmarks.
Beyond Sparse Captions: Aligning Slide-Level Text and Patch-Level Vision in Pathology
Daniel Shao ⋅ Dongmin Bang ⋅ Luca Weishaupt ⋅ Nic G Reitsam ⋅ Sophia J. Wagner ⋅ Ming Yang Lu ⋅ Long P Le ⋅ Richard Chen ⋅ Faisal Mahmood
Contrastive vision-language (VL) pretraining in computational pathology has generally relied on curated image-caption datasets sourced from social media, educational videos, and research publications. These corpora are sparse relative to the diversity of tissue morphology and difficult to scale further, as pathologists do not routinely produce descriptive captions in clinical practice. We propose to instead leverage slide-level pathology reports which can be generated during routine diagnostic workflows in a scalable manner. A key challenge in this weakly-supervised setting is that pathology reports describe macroscopic, diagnostic findings while patch-level visual features describe microscopic findings. We introduce Locked Image Multiple Instance PreTraining (LIMIT), which addresses this gap by freezing a pretrained vision encoder and fine-tuning only the text encoder: report embeddings serve as cross-attention queries over bags of patch features, producing report-conditioned slide representations optimized via cross-entropy. Using LIMIT, we establish CONCH-Z, which aligns the text encoder of CONCH v1.5 to 335,645 clinical reports. Evaluated across 20 tasks spanning patch-level classification, tumor detection, and slide-level subtyping, CONCH-Z establishes state-of-the-art zero-shot performance over VL encoders trained on curated caption datasets, while simultaneously closing the gap with slide-level foundation models trained with substantially more parameters and compute. We will release CONCH-Z weights and evaluation code to support reproducibility and broader community use.
Beyond Steering Vector: Flow-based Activation Steering for Inference-Time Intervention
Zehao Jin ⋅ Ruixuan Deng ⋅ Junran Wang ⋅ Xinjie Shen ⋅ Chao Zhang
Activation steering has emerged as a promising alternative for controlling language-model behavior at inference time by modifying intermediate representations while keeping model parameters frozen. However, large-scale evaluations such as AxBench show that existing steering methods are often outperformed by simple in-context prompting and generalize poorly to unseen concepts. We hypothesize that these limitations arise from unvalidated simplifying assumptions shared across prior methods, which typically restrict steering interventions to fixed, single-step, position-invariant transforms. We propose FLAS (Flow-based Activation Steering), which learns a general, concept-conditioned velocity field $v_t(h,t,c)$ that transports unsteered activations to steered ones without relying on these assumptions. On AxBench, FLAS is the first learned method to consistently outperform prompting, reaching held-out harmonic means of $1.015$ on Gemma-2-2B-IT and $1.113$ on Gemma-2-9B-IT without per-concept tuning. Analysis of the learned flow shows curved, multi-step, token-varying trajectories, which suggests that previous hypotheses on activation space geometry might be incomplete. Our code is available at https://anonymous.4open.science/r/FLAS.
Beyond World-Frame Action Heads: Motion-Centric Action Frames for Vision-Language-Action Models
Huoren Yang ⋅ Jianchao Zhao ⋅ Hu Yusong ⋅ Qiguan Ou ⋅ Yuyang Gao ⋅ Wei Ke ⋅ Yuhang He ⋅ SongLin Dong ⋅ Zhiheng Ma ⋅ Yihong Gong
Vision-Language-Action (VLA) models have advanced rapidly with stronger backbones, broader pre-training, and larger demonstration datasets, yet their action heads remain largely homogeneous: most directly predict action commands in a fixed world coordinate frame. We propose \textbf{MCF-Proto}, a lightweight action head that equips VLA policies with a Motion-Centric Action Frame (MCF) and a prototype-based action parameterization. At each step, the policy predicts a rotation $R_t \in SO(3)$, composes actions in the transformed local frame from a set of prototypes, and maps them back to the world frame for end-to-end training, using only standard demonstrations without auxiliary supervision. This simple design induces stable emergent structure. Without explicit directional labels, the learned local frames develop a stable geometric structure whose axes are strongly compatible with demonstrated end-effector motion. Meanwhile, actions in the learned representation become substantially more compact, with variation captured by fewer dominant directions and more regularly organized by shared prototypes. These structural properties translate into improved robustness, especially under geometric perturbations. Our results suggest that adding lightweight geometric and compositional structure to the action head can materially improve how VLA policies organize and generalize robotic manipulation behavior. An anonymized code repository is provided in the supplementary material.
BFS-PO: Best-First Search for Large Reasoning Models
Fiorenzo Parascandolo ⋅ Wenhui Tan ⋅ Enver Sangineto ⋅ Xiaoyi Yu ⋅ Ruihua Song ⋅ Rita Cucchiara
Large Reasoning Models (LRMs) have shown excellent performance in reasoning tasks using long Chain of Though. However, this has also led to a significant increase of computational costs and the generation of verbose output, a phenomenon known as overthinking. The tendency to overthinking is often exacerbated by Reinforcement Learning (RL) algorithms such as GRPO/DAPO. In this paper, we propose BFS-PO, an RL algorithm which alleviates this problem using a Best-First Search exploration strategy. Specifically, the reasoning chains generated at training time by BFS-PO are organized into a search tree, within which the shortest correct sequence is selected as the best and expanded using a backtracking mechanism based on maximum entropy nodes. In this way, we bias the exploration of the solution space towards the search for increasingly shorter solutions, training the LRM to generate progressively more concise answers. Using different benchmarks and base LRMs, we show that BFS-PO can simultaneously increase the LRM accuracy and shorten its reasoning chains. Our code and models are available in the supplementary material and will be published after this article is accepted.
Bilevel Optimization of Synthetic Trajectories for Multi-Turn LLM Fine-Tuning
Shresth Verma ⋅ Mauricio Tec ⋅ Cheol Woo Kim ⋅ Kai Wang ⋅ Milind Tambe
While LLMs excel at single-turn generation, they struggle with long-horizon, multi-turn interactions. Offline reinforcement learning (RL) offers a scalable approach, yet its performance hinges on the availability and quality of multi-turn trajectory data. A common remedy is to augment training with synthetic trajectories generated by LLMs or simulators, but synthetic data is highly heterogeneous in quality, and naively treating all trajectories as equally informative can degrade performance. We propose BOOST, a bilevel optimization framework where the inner level trains the LLM on reweighted data and the outer level trains a lightweight reweighting head on held-out real validation tasks, assigning continuous trajectory-level weights without requiring an external judge. To ground this approach, we derive a PAC-Bayesian bound revealing a three-way trade-off: synthetic data increases diversity but risks task-shift, while concentrating weight on high-quality trajectories improves empirical performance at the cost of effective sample size. Empirically, our method consistently outperforms multiple baselines. Analysis reveals it upweights synthetic trajectories that align with the real data distribution and exhibit higher qualitative merit.
BiLi: Bridging the Last Mile in LiDAR Localization
Minghang Zhu ⋅ Junchuan Lan ⋅ Zhijing Wang ⋅ Haoze Chang ⋅ Wen Li ⋅ Sheng Ao ⋅ Cheng Wang
In large-scale outdoor LiDAR localization, Scene Coordinate Regression (SCR) achieves sub-meter accuracy, but its deployment in high-precision autonomous driving is hindered by a "last-mile" problem: the point-wise prediction paradigm causes unphysical trajectory jittering, rendering localization kinematically discontinuous despite low mean errors. To bridge the gap between high-precision localization and temporal kinematic consistency, we propose BiLi, a spatiotemporal synergistic relocalization framework. Spatially, BiLi utilizes Implicit Manifold Regularization (ISR) driven by an offline neural implicit field. Formulated as a novel point-to-manifold projection constraint, ISR strictly anchors predicted coordinates to continuous physical zero-level sets. Temporally, to bypass computationally heavy implicit optimization, we distill the teacher's kinematics into a lightweight feed-forward network, compelling it to learn robust, equivariant physical priors. Finally, a Differentiable Robust Manifold Gating mechanism dynamically fuses the global spatial predictions with the temporal kinematic states via adaptive state updating, facilitating efficient real-time deployment. Extensive experiments demonstrate that BiLi fundamentally resolves the "last-mile" challenge. Our method outperforms state-of-the-art approaches by 25% and 65% on the Oxford and NCLT benchmarks, respectively, delivering drift-free, kinematically coherent, and highly accurate real-time localization.
We study \emph{bilinear matching bandits}, an online sequential decision problem in which at each round a learner assigns $M$ users to $N$ items via an injective matching $\pi:[M]\hookrightarrow[N]$ and observes semi-bandit feedback with bilinear mean reward $\mathbf{x}_i^\top \mathbf{\Theta}^\star \mathbf{y}_{\pi(i)}$, where the parameter $\mathbf{\Theta}^\star \in \mathbb{R}^{d_1 \times d_2}$ is unknown. We propose \textsc{HybridRRD}, a hybrid algorithm that combines high-width elimination with a row-wise regularized design sampler over the fractional matching polytope. Without assuming $\mathbf{\Theta}^\star$ is low-rank, \textsc{HybridRRD} achieves a high-probability regret bound of $\widetilde{O} (d_2 \sqrt{d_1MT} + Md_1 d_2)$, suppressing problem-scale constants. This improves the leading dimension dependence from $d_1d_2$ to $d_2\sqrt{d_1}$ over the standard Kronecker-linearized contextual combinatorial semi-bandit baseline. We further prove a minimax lower bound of $\Omega(d_2 \sqrt{d_1 M T})$ in a large-item regime, matching the leading regret term up to logarithmic factors. Numerical experiments show that \textsc{HybridRRD} performs favorably against Kronecker-linearized baselines.
Black-Box Uncertainty Quantification for Large Language Models via Ensemble-of-Ensembles
Wang Ma ⋅ Debarun Bhattacharjya ⋅ Junkyu Lee ⋅ Nhan H Pham ⋅ Harsha Kokel ⋅ Qiang Ji
Reliable use of large language models (LLMs) requires uncertainty estimates that separate ambiguity intrinsic to a prompt from genuine knowledge gaps. Bayesian and deep-ensemble methods deliver this aleatoric/epistemic decomposition in principle, but are computationally prohibitive at LLM scale and require white-box access. Existing black-box methods, in contrast, summarize sample variability into a single consistency scalar that cannot make the distinction. We introduce a two-level ensemble that bridges this gap. Stochastic decoding of a fixed prompt forms the inner ensemble; meaning-preserving semantic perturbations of the prompt form the outer ensemble. Variability is measured in a continuous embedding space, and the law of total covariance gives an exact AU/EU split that requires only black-box sampling access. On top of this estimator, we contribute three theoretical results that ground the framework: (i) a finite-sample bias identity for nested Monte Carlo with a closed-form bias-corrected epistemic estimator; (ii) a Bayesian bridge bounding the gap between perturbation-population AU/EU and Bayesian targets through three local conditions that we estimate as diagnostics from the data the estimator already collects; and (iii) a separation result identifying a class of failure modes that any consistency-only black-box method provably misses while perturbation-based EU detects. Across five short-form QA benchmarks and four instruction-tuned LLMs spanning two families and two scales, the estimator matches or surpasses strong white-box baselines; controlled paired interventions on AmbigQA and Natural Questions confirm the AU/EU split tracks distinct, actionable sources of failure; and a long-form pilot on the NQ long-answer split shows the framework extends to claim-level factuality without retraining. Across five short-form QA benchmarks and four instruction-tuned LLMs spanning two families and two scales, the estimator matches or surpasses strong white-box baselines; controlled paired interventions on AmbigQA and Natural Questions confirm the AU/EU split tracks distinct, actionable sources of failure; and a long-form pilot on the NQ long-answer split shows the framework extends to claim-level factuality without retraining.
BlenderFORGE: Framework for Optimizing Reactive 3D-Graphics Editing Ability of MLLMs
Zilin Guo ⋅ Tianrui ZHANG ⋅ Yi RONG ⋅ Yichen Liu ⋅ Yuxin Guo ⋅ Jingcheng Ni ⋅ Lewei Lu ⋅ Zehuan Wu ⋅ Dan Xu
Multimodal Large Language Models (MLLMs) have shown promise as interactive agents, yet precise, programmatic 3D graphics editing remains difficult to train because existing systems largely rely on closed-source model APIs and lack deterministic, state-verifiable execution environments. We present BlenderFORGE, a trainable data-environment-optimization framework for open-weight MLLM agents that edit explicit Blender scenes through Blender Python (bpy) code generation. BlenderFORGE targets four fundamental state-verifiable editing task families: object placement, shape-key editing, lighting adjustment, and material modification. It integrates three components: (1) a scalable perturbation-based data generation pipeline built from curated assets and large-scale 3D scene repositories; (2) an isolated Blender Sandbox for deterministic code execution, structured state export, and closed-loop reward computation; and (3) a two-stage training pipeline that uses offline teacher trajectories for cold-start Supervised Fine-Tuning (SFT) and sandbox-computed rewards for multi-turn tool-agentic Reinforcement Learning (AgentLoop-RL). Experiments across the four task families show that the resulting open-weight Qwen3-VL-8B-BlenderFORGE agent substantially improves over its base model and achieves competitive performance with several zero-shot proprietary MLLM baselines.
Block-Wise Differentiable Sinkhorn Attention: Tail-Refinement Gradients with a Gap-Aware Dustbin Bridge
Dylan Forde
We study long-context balanced entropic optimal transport (OT) attention on TPU hardware through a stopped-base, fixed-depth tail-refinement surrogate. After a stopped $T$-step Sinkhorn solve, we unroll a short refinement tail and differentiate that surrogate exactly. For the production $R=2$ case, the backward pass contains four staircase plan factors. We prove an exact one-reference-tile schedule: the $R=2$ score cotangent is a single reference plan tile times an explicit modifier field built from vector cotangents and dual differences. This yields block-wise cost $O((T+R)LW)$, $O(Ld)$ input storage, and $O(L)$ additional HBM usage for fixed head dimension $d$ and band width $W$. We also formalize the current \texttt{dustbin\_block} path as the same balanced surrogate on an augmented support, so the schedule lifts to the gap-aware transport path used in our TPU runs. We provide a local surrogate-bias bound, an a posteriori bias certificate, and a projective contraction certificate for strictly positive active blocks. On synthetic masked problems, the optimized kernel matches exact autodiff of the same centered surrogate to within $10^{-5}$--$10^{-10}$. On TPU v6e-8, a four-configuration Pfam screen completes end-to-end, and a promoted balanced $R=2$ run sustains roughly $8.5$ examples per second through a three-hour budget, reaching step $1437$. Held-out Pfam test shards improve reconstruction from $3.17$ to $0.99$ and sparse CE from $5.86$ to $5.69$ relative to step $0$. These results support exact fixed-depth backward theory, a theorem-matching gap-aware bridge, and trainability evidence for the production path.
BMAttn: Block-Aligned Mixed-Precision Attention Quantization for LLM Inference
Zining Wang ⋅ Haojie Duanmu ⋅ Zhihang Yuan ⋅ Fanliu Kong ⋅ Ruihao Gong ⋅ Jinyang Guo ⋅ Xianglong Liu
The deployment of Large Language Models (LLMs) with extended context windows is fundamentally constrained by the quadratic computational and memory costs of the self-attention mechanism. While acceleration techniques like sparse attention and quantization offer potential relief, they currently face a critical dilemma: uniform compression strategies fail to capture the non-uniform distribution of information importance, degrading performance on long-range dependencies; conversely, fine-grained, token-level adaptation often introduces irregular memory access patterns that negate theoretical efficiency gains on modern GPUs. To resolve this tension, we introduce BMAttn (Block-Aligned Mixed-Precision Attention), a framework that unifies importance-aware precision allocation with hardware-friendly execution. BMAttn partitions attention maps into block-aligned high-precision, low-precision, and sparse regions, using a novel affine windowing mechanism to dynamically adjust boundaries based on sequence length. We further propose a saliency-weighted calibration method and a layer-adaptive regularizer that adaptively aligns precision targets with perceptual importance and layer sensitivity. Extensive experiments demonstrate that BMAttn achieves a \textbf{3.3x} speedup on long-context tasks with negligible accuracy loss, and up to \textbf{5x} speedup with minimal degradation, effectively bridging the gap between algorithmic adaptivity and hardware efficiency.
BodyBench: Evaluating Adversarial Image Defenses Against AI Nudification Inpainting
Li Qiwei ⋅ Salma Abdel Magid ⋅ Olga Russakovsky ⋅ Eric Gilbert ⋅ Sarita Schoenebeck
Adversarial perturbation defenses protect images from generative editing by adding imperceptible noise. Yet none of these methods have been evaluated against AI nudification, despite AI-generated non-consensual intimate imagery (AIG-NCII) being among the most documented harms of generative AI and produced largely through inpainting. We argue this is a problem-formulation failure: prior threat models do not correspond to nudification attacks, and the misalignment is structural, spanning mask shape, prompt distribution, and success criterion which causes existing evaluations to systematically overestimate protection effectiveness. We reformulate adversarial inpainting protection around this documented harm. We introduce BodyBench, a benchmark of 648 fully synthetic clothed adult subjects, each paired with 11 mask variants spanning body, face, context, and dilation regions, along with nudification and redressing attack prompts. We propose three harm-aligned metrics: sexualization, naturalness, and recognizability that jointly determine whether an inpainted output constitutes a successful nudification attack. Auditing four recent protection methods (PhotoGuard, DiffusionGuard, AdvPaint, DiffVax) on BodyBench, we find that protection succeeds in only one configuration of one method, and only under exact mask alignment between defender and attacker---a condition no real defender can guarantee and one that we show is easily always bypassed. We also show that standard metrics such as PSNR are essentially uncorrelated with nudification protection. BodyBench makes this misalignment measurable and provides a foundation for defenses grounded in documented harm.
Reinforcement Learning with Verifiable Rewards (RLVR) has become an important approach for improving the reasoning capabilities of large language models (LLMs). On-policy training remains the dominant RLVR paradigm, but its reliance on fresh rollouts leads to poor data efficiency. To address this bottleneck, off-policy paradigm has attracted increasing attention by reusing historical trajectories through replay buffers. Despite its efficiency advantages, off-policy training often struggles to match the performance of on-policy training. Our analysis reveals that this limitation primarily stems from data staleness introduced by replay buffers: to maintain training stability, mainstream algorithms discard a large fraction of high-staleness optimization signals, leading to substantial performance degradation. To address this issue, we propose STAR (Staleness-Aware Replay), a data-centric replay framework decoupled from specific optimization algorithms. STAR improves off-policy RLVR by constructing low-staleness and high-quality training data. Experiments on a range of mathematical reasoning and code generation tasks show that STAR consistently improves off-policy RLVR with up to 19% relative performance gains, and enables it to surpass the on-policy baseline using only around 13% of the rollout data while achieving 5x-6x wall-clock speedup.
BootstrapAgent: Turning Repository Setup into Reusable Agent Knowledge
Sihan Fu ⋅ Oucheng Liu ⋅ Shiyuan Wang ⋅ Jin Shi ⋅ Chengkun Wei
Code agents are rapidly lowering the barrier for developers to work with unfamiliar repositories, and are increasingly adopted for tasks such as building prototypes, reproducing experiments, and adding features. Reliably executing these tasks depends on a critical prerequisite: the target repository must first be successfully set up and brought into a usable development state. This bootstrapping process requires substantial trial-and-error exploration by code agents, yet the resulting knowledge (e.g., resolved dependencies, repair strategies) remains trapped in a single conversation, unavailable to future agents or developers. We therefore formulate repository bootstrapping as a reusable startup knowledge problem and introduce BootstrapAgent, a multi-agent framework that distills bootstrap exploration into a persistent, verifiable, agent-consumable .bootstrap contract. Through CI evidence extraction, structured planning, deterministic Docker-based verification, and trace-driven repair, BootstrapAgent generates a contract covering environment setup, diagnostic checks, minimal verification, and accumulated repair knowledge. We further propose warm repair with clean replay to accelerate iterative debugging without sacrificing cold-start reproducibility, and a two-level verification strategy to prevent reward hacking. Experiments on three benchmarks show that BootstrapAgent achieves a 92.9% success rate, outperforming the baseline by over 10% while reducing downstream agent token usage by 25.9% and build time by 22.3%.
Brain-OF: An Omnifunctional Foundation Model for fMRI, EEG and MEG
Hanning Guo ⋅ Hanwen Bi ⋅ Farah Abdellatif ⋅ Andrei Galbenus ⋅ N. J Shah ⋅ Abigail Morrison ⋅ Jürgen Dammers
Brain foundation models have achieved remarkable advances across a wide range of neuroscience tasks. However, most existing models are limited to a single functional modality, restricting their ability to exploit complementary spatiotemporal dynamics and the collective data scale across different neuroimaging techniques. This limitation largely arises from severe semantic heterogeneity and resolution discrepancies among modalities. To address these challenges, we propose Brain-OF, an omnifunctional brain foundation model jointly pretrained on fMRI, EEG and MEG, capable of handling both unimodal and multimodal inputs within a unified framework. To reconcile heterogeneous spatiotemporal resolutions, we introduce the Any-Resolution Neural Signal Sampler, which projects diverse brain signals into a shared semantic space. To further manage semantic shifts, the Brain-OF backbone integrates DINT attention with a Sparse Mixture of Experts, where shared experts capture modality-invariant representations and routed experts specialize in modality-specific semantics. Furthermore, to explicitly internalize the characteristics of neural activity through self-supervised learning, we propose Masked Temporal-Frequency Modeling, a dual-domain pretraining objective that jointly reconstructs brain signals in both the time and frequency domains. Brain-OF is pretrained on a large-scale corpus comprising around 40 datasets and demonstrates superior performance across diverse downstream tasks, highlighting the benefits of joint multimodal integration and dual-domain pretraining. Our code will be made publicly available upon acceptance.
Breaking the Exactness Barrier: Interleaved DeepSeek Sparse Attention for Efficient Long Context Reasoning
Yifan GUO ⋅ Wei Cui
Token-level dynamic sparse attention exemplified by DeepSeek Sparse Attention (DSA) selects the globally most relevant key-value tokens via an exact Top-$K$ operator, achieving superior model quality over block-level alternatives. However, this exact selection creates a severe distributed inference bottleneck: enforcing an exact global Top-$K$ across GPUs inevitably incurs either redundant full-context retrieval or costly multi-stage cross-device synchronization, which largely negates the computational advantages of DSA at long context lengths. We first show that the exact Top-$K$ bound is unnecessary during inference: once the truly critical tokens are recalled, admitting additional context preserves or even improves accuracy. Leveraging this insight, we propose Interleaved DeepSeek Sparse Attention (IDSA), which distributes tokens across GPUs in an interleaved layout so that each device performs only a relaxed local Top-$m$ selection. Under this layout, the union of independent per-GPU Top-$m$ selections near-completely covers the globally most relevant Top-$K$ tokens. This allows each device to proceed with its local selection with minimal cross-GPU overhead while avoiding both expensive full-context Top-$K$ computation and multi-stage cross-GPU merging, enabling a not only distributed but also synchronization-efficient inference pipeline. Without any retraining, IDSA delivers dramatic throughput gains for context lengths exceeding 100K tokens on both DeepSeek-V3.2 and GLM-5, while preserving equivalent or better reasoning performance on the AIME and Needle-In-A-Haystack benchmarks.
Breaking the Quality–Privacy Tradeoff in Tabular Data Generation via In-Context Learning
Xinyan Han ⋅ yan Lu ⋅ Xiaoyu Lin ⋅ Yuanyuan Jiang ⋅ Yuanrui Wang ⋅ Xuanyue Li ⋅ Wenchao Zou ⋅ Xingxuan Zhang
Tabular data synthesis aims to generate high-quality data while preserving privacy. However, we find that existing tabular generative models exhibit a clear tradeoff in the small-data regime: improving data quality typically comes at the cost of increased memorization of training samples, thereby weakening privacy protection. This tradeoff arises because small training sets make it difficult for dataset-specific generative models to distinguish generalizable structure from sample-specific patterns. To address this, we propose DiffICL, which formulates tabular data generation as an in-context learning problem. Instead of fitting each dataset from scratch, DiffICL leverages pretrained structural priors learned from a large collection of datasets, enabling it to infer data distributions from limited context rather than memorizing individual samples. We evaluate DiffICL on 14 real-world datasets. Results show that DiffICL improves both data quality and privacy, and generates synthetic data that provides effective data augmentation. Our findings suggest that the quality–privacy tradeoff can be improved through better training paradigms.
BRIDGE: Brain-Vision Representation Integration through Depth and Granularity Encoding
Chenyuan Hong ⋅ Binghao Ye ⋅ Yufei Guo ⋅ Guoqi Li
Decoding visual content from non-invasive brain signals remains challenging because neural responses evolve over time whereas images are typically static. Most existing methods align an entire neural response window to a single final visual embedding, this overlook a fundamental representational mismatch: brain signals reflect both early perceptual and later integrative representations, whereas the final visual embedding is biased toward high-level semantics. Therefore, we propose BRIDGE, a brain-vision representation integration framework that aligns two modalities through Visual Depth Encoding and Brain Granularity Encoding. On the visual side, BRIDGE extracts and fuses CLIP representations from multiple depths, producing an alignment target that preserves low-level and high-level information. On the brain side, instead of treating the whole temporal window as homogeneous and obscuring temporally heterogeneity, BRIDGE explicitly partitions stimulus-evoked EEG/MEG responses into a small number of temporally ordered stages and adaptively pools them. The resulting brain and visual embeddings are trained in a shared latent space by contrastive learning and can further support brain-to-image generation through a pretrained diffusion prior. Experiments on THINGS-EEG and THINGS-MEG demonstrate that BRIDGE achieves strong retrieval and generation performance. Ablation studies further confirm the complementary benefits of depth-wise visual aggregation and temporal brain factorization.
Bridging Academia and Industry: A Comprehensive Benchmark for Attributed Graph Clustering
Yunhui Liu ⋅ Pengyu Qiu ⋅ Yu Xing ⋅ Peng Du ⋅ Yongchao Liu ⋅ Chuntao Hong ⋅ Jiajun Zheng ⋅ Tao Zheng ⋅ Tieke He
Attributed Graph Clustering (AGC) is a fundamental unsupervised task that integrates structural topology and node attributes to uncover latent patterns in graph-structured data. Despite its significance in industrial applications such as fraud detection and user segmentation, a significant chasm persists between academic research and real-world deployment. Current evaluation protocols suffer from the small-scale, high-homophily citation datasets, non-scalable full-batch training paradigms, and a reliance on supervised metrics that fail to reflect performance in label-scarce environments. To bridge these gaps, we present PyAGC, a comprehensive, production-ready benchmark and library designed to stress-test AGC methods across diverse scales and structural properties. We unify existing methodologies into a modular Encode-Cluster-Optimize framework and, for the first time, provide memory-efficient, mini-batch implementations for a wide array of state-of-the-art AGC algorithms. Our benchmark curates 12 diverse datasets, ranging from $2.7 \times 10^3$ to $1.1 \times 10^8$ nodes, specifically incorporating industrial graphs with complex tabular features and low homophily. Furthermore, we advocate for a holistic evaluation protocol that mandates unsupervised structural metrics and efficiency profiling alongside traditional supervised metrics. Our benchmark offers the community a robust, reproducible, and scalable platform to advance AGC research towards realistic deployment.
Bridging Sequence and Structure with Unified Domain Adaptation for Drug-Target Interaction Prediction
Mingcan Yuan ⋅ He Li ⋅ Zhiyi Ju ⋅ Mang Ye ⋅ Qingxiong Tan
Drug-target interaction (DTI) prediction is pivotal for accelerating drug discovery, yet existing methods struggle with generalization under cross-domain distribution shifts. Conventional approaches typically rely on static protein representations and assume distributional consistency between source and target domains, often resulting in poor calibration when encountering novel protein families or scaffolds. To address this issue, we present $\textbf{Bi}$-view $\textbf{Fold}$-aware prediction (BiFold), a novel unified interaction-adaptive framework that enhances generalization via dual-view representation learning and calibration-aware adaptation. BiFold synergistically models sequence and structural evidence, introducing a routing module to dynamically fuse the two views at the feature-channel level and exploit complementary information. Furthermore, we design a calibration-aware training principle that aligns domain statistics and encourages flatter source solutions. Crucially, target supervision is introduced only after a forward-only warm-up phase to prevent reinforcement of miscalibrated early predictions, effectively mitigating confirmation bias in pseudo-label learning. Extensive experiments demonstrate that BiFold consistently outperforms the state-of-the-art methods across diverse datasets in both in-domain and cross-domain settings. Additionally, interpretability analysis reveals that the routing module partitions the feature space into view-dominant and view-neutral subspaces, providing concrete evidence of adaptive multi-modal complementarity.
Bridging the Gap Between Harmfulness Belief and Refusal Behavior for Safety Alignment
Lu Zhang ⋅ Chen Feng ⋅ Qingzhuo Wang ⋅ Wen Shen ⋅ Zhihua Wei
Many existing defenses mainly strengthen refusal behavior to improve the safety of LLMs. Recent interpretability work shows that, in safety-aligned LLMs, harmfulness belief and refusal behavior are represented separately in the hidden states at the last token of the user instruction ($t_{\mathrm{inst}}$) and the response-start token ($t_{\mathrm{post}}$), respectively. Based on it, we find that harmfulness belief is not reliably preserved when it is carried from the $t_{\mathrm{inst}}$ to the $t_{\mathrm{post}}$. As a result, even when the LLM internally recognizes a request as harmful, it may still produce a compliant response. To address this, we propose Bridging Harmfulness and Refusal (BHR), a training framework that bridges harmfulness belief and refusal behavior. Specifically, BHR trains the adapter so that harmfulness information is carried to the hidden state at the $t_{\mathrm{post}}$, thereby establishing a second pathway from harmfulness belief to refusal behavior. To keep this signal reliable during fine-tuning, we introduce two additional losses. A belief loss helps the LLM maintain an accurate harmfulness judgment at the $t_{\mathrm{inst}}$. A belief-gated refusal loss uses the LLM's harmfulness belief at $t_{\mathrm{inst}}$ to regulate refusal learning, making attacks that target refusal features less effective while reducing the risk of over-refusal on benign inputs. Experiments across multiple LLMs and safety benchmarks show that BHR substantially improves robustness against both white-box and black-box jailbreak attacks, reduces the risk of over-refusal on benign prompts, and preserves general capabilities.
BucpTSF: Breaking the Uniform Computation Paradigm in Time-Series Forecasting
Hua Wang ⋅ Jinghao Lu ⋅ Fan Zhang
Real-world time series typically exhibit non-uniform information density across temporal positions, frequency structures, and forecasting horizons: critical dynamics are concentrated in local segments, different frequency components carry different structural information, and different horizons require different modeling complexity. However, most existing methods still allocate computation approximately uniformly across these dimensions, leading to redundant computation in low-information regions and insufficient modeling of key patterns. To address this, we propose BucpTSF, an information-density-driven structured framework for time series forecasting. Specifically, BucpTSF reorganizes raw sequences into hierarchically heterogeneous multi-level token representations through information-density-aware temporal representation reorganization, captures key cross-token dependencies with low overhead via frequency-selective structured token interaction, and further introduces horizon-conditioned structured extrapolation to adaptively balance dynamic modeling and structural extrapolation bias for different horizons, thereby improving the accuracy and stability of long-term forecasting. Extensive experiments on real-world datasets show that BucpTSF consistently outperforms strong baselines under various forecasting settings, validating the effectiveness of information-density-driven non-uniform computation allocation for time series forecasting.
Budget-Conditioned Clipping Policies for Differentially Private Federated Learning
Hao Zhou ⋅ SiQi Cai ⋅ Hua Dai ⋅ Letian Sha ⋅ Yichen Li ⋅ MingCai Chen
Gradient clipping is a central but brittle design choice in differentially private federated learning: overly small thresholds bias client updates, whereas overly large thresholds amplify the noise required by privacy. This trade-off is further complicated by client heterogeneity and personalized privacy budgets, where a single global threshold is often mismatched and online adaptation from private training statistics can complicate privacy accounting. We study clipping-threshold selection as a privacy-budget-conditioned policy learning problem. We propose PAC-DP, a proxy-learned clipping framework for record-level locally private federated learning. Before private training, PAC-DP calibrates a deterministic policy $\pi_\theta(\varepsilon,t)$ from public or synthetic proxy simulations, mapping a client's target privacy budget and the training round to a clipping threshold. The policy is then frozen and used during private training, so deployed thresholds depend only on the declared budget, round index, and public/proxy-learned parameters, not on private gradients, losses, or client-specific update histories. PAC-DP combines this policy with per-example clipping, Gaussian perturbation, and per-client RDP accounting. We show that, under public or separately privatized proxy calibration, the frozen policy introduces no additional record-level privacy loss beyond the clipped Gaussian mechanisms. We further analyze the clipping-sensitive utility trade-off through a decomposition involving clipping bias, stochastic variance, client heterogeneity, and DP noise, and characterize proxy-to-target transfer via a regret bound. Experiments on MNIST, CIFAR-10, CIFAR-100, and Heart Disease show that PAC-DP improves the privacy--utility trade-off and communication efficiency over implemented fixed-threshold and adaptive DP-FL baselines under matched privacy budgets.
B-XAIC Dataset: Benchmarking Explainable AI for Graph Neural Networks Using Chemical Data
Magdalena Proszewska ⋅ Tomasz Danel ⋅ Antoni Antoszek ⋅ Dawid Damian Rymarczyk
Understanding the reasoning behind deep learning model predictions is crucial in cheminformatics and drug discovery, where molecular design determines their properties. However, current evaluation frameworks for Explainable AI (XAI) in this domain often rely on artificial datasets or simplified tasks, employing data-derived metrics that fail to capture the complexity of real-world scenarios and lack a direct link to explanation faithfulness. To address this, we introduce B-XAIC, a novel benchmark constructed from real-world molecular data and diverse tasks with known ground-truth rationales for assigned labels. Through a comprehensive evaluation using B-XAIC, we reveal limitations of existing XAI methods for Graph Neural Networks (GNNs) in the molecular domain. This benchmark provides a valuable resource for gaining deeper insights into the faithfulness of XAI, facilitating the development of more reliable and interpretable models.
We propose a label-free post-processing framework that improves a strong but miscalibrated primary model using a weaker yet better-calibrated reference. Our key insight is that strict improvement is possible if and only if the two models are not *mutually calibrated*, meaning there does not exist joint distribution over their predictions and outcomes such that both predictions are simultaneously calibrated. We formalize this condition and connect it to no-arbitrage results from economics. Under such condition, we develop an efficient post-processing algorithm of the strong model's outputs based on Bregman projection, with a strict worst-case performance improvement guarantee. Experiments on representative LLMs across varying scales demonstrate the effectiveness of our method, reducing the ECE of the primary model by over $40$\% on common benchmarks.
Can Folding Models Tell Binders from Bluffers? Evidence from POISK: The Patent-Derived Antibody Dataset
Daria Tupikina ⋅ Andrea Roncoli ⋅ Alexander Bujotzek ⋅ Brennan Abanades Kenyon
Structure prediction models are increasingly deployed as zero-shot digital screens in antibody design, operating under the assumption that high folding confidence implies true binding. However, rigorous validation of this premise has been limited by small test sets, narrow target diversity, and the risk of data leakage from training-adjacent structures. Here we introduce POISK: Patent-extracted Organization of Immunoglobulin Sequence Knowledge, a large-scale dataset derived from the US patent literature, comprising 89074 antibody-antigen records extracted from 10110 patents using an LLM-based pipeline. From this resource, we curate a balanced benchmark set of 22494 complexes: 11247 patent-validated binders without a known resolved structure paired with putative non-binders constructed from sequence-similar human repertoire antibodies. Using this dataset, we benchmark six leading structure prediction models (AlphaFold-Multimer, Boltz-2, Intellifold v2, OpenFold3, Protenix-v1, and Protenix-v2) on their ability to discriminate binders from decoys. The best-performing model, Protenix-v2, achieves 87% pairwise accuracy on the most confident half of predictions and correctly identifies the true binder from a pool of 50 candidates in 32% of trials; this represents a substantial enrichment over random selection, yet remains far from reliable single-candidate identification. Structural consensus across models also discriminates binders from decoys, with pairwise accuracy reaching 73.1% when three or more models converge on a pose DockQ $\geq$ 0.23, covering around 54.2% of pairs). Yet epitope localization remains poor: the best model contacts annotated epitope residues in only 45% of cases, indicating that high confidence reflects learned statistical associations rather than accurate physical binding modes. Our results suggest that current folding models provide useful but limited signal for antibody screening and should be complemented by orthogonal methods. We release POISK as a community resource for future benchmarking and model training.
Can Linguistic Reasoning Vectors Enhance Multimodal Reasoning Ability?
Ziyi Wang ⋅ Li Li ⋅ Aolin Zhou ⋅ Yankun Shen ⋅ Liu Chonghan ⋅ Shuxia Lin ⋅ Xu Yang
Most Vision-Language Models (VLMs) are built by extending pretrained Large Language Models (LLMs) with visual modules and multimodal alignment. However, this multimodal scaling often degrades the language-side reasoning ability originally encoded in the base LLM. While the base LLM retains usable reasoning ability after scaling, the aligned VLM itself cannot reliably access this ability. Therefore, recovering degraded reasoning capability in VLMs may be more effective when using the base LLM as the source rather than relying on the VLM alone. Motivated by this, we propose LIFT (Language-side reasoning Facilitation and Transfer), a lightweight vector-intervention method that transfers reasoning capability from the base LLM to the VLM without retraining the backbone. LIFT defines Reasoning Vectors as answer-token hidden-state differences between a Reasoner path with an explicit reasoning trace and a Solver path without it, and injects these vectors into language-side activations of the target VLM. LIFT further supports learnable vector adaptation while keeping the VLM backbone frozen. We evaluate LIFT on two VLMs across six reasoning benchmarks, comparing Reasoning Vectors extracted from the base LLM and from the aligned VLM under matched protocols. Results show that LLM-derived vectors consistently outperform VLM-derived vectors, confirming that the base LLM is a more effective source for recovering reasoning. LIFT partially recovers degraded reasoning through lightweight language-side interventions. Further analyses show that Reasoning Vectors influence intermediate reasoning behavior rather than merely altering final answers.
Can LLM Agents Respond to Disasters? Benchmarking Heterogeneous Geospatial Reasoning in Emergency Operations
Junjue Wang ⋅ Weihao Xuan ⋅ Heli Qi ⋅ Pengyu Dai ⋅ Kunyi Liu ⋅ Hongruixuan Chen ⋅ Zhuo Zheng ⋅ Junshi Xia ⋅ Stefano Ermon ⋅ Naoto YOKOYA
Operational disaster response goes beyond damage assessment, requiring responders to integrate multi-sensor signals, reason over road networks, populations and key facilities, plan evacuations, and produce actionable reports. However, prior work largely isolates remote-sensing perception or evaluates generic tool use, leaving the end-to-end workflows of emergency operations underexplored. In this paper, we introduce Disaster Operational Response Agent benchmark (DORA), the first agentic benchmark for end-to-end disaster response: 515 expert-authored tasks across 45 real-world disaster events spanning 10 types, paired with expert-verified, replayable gold trajectories totaling 3,500 tool-call steps. Tasks span five dimensions that cover the operational disaster-response pipeline: disaster perception, spatial relational analysis, rescue and evacuation planning, temporal evolution reasoning, and multi-modal report synthesis. Agents compose calls from a 108-tool MCP library over heterogeneous geospatial data: optical, SAR, and multi-spectral imagery across single-, bi-, and multi-temporal sequences (0.015-10m GSD), complemented by elevation and social vector layers. We comprehensively evaluate 13 frontier LLMs on our benchmark, revealing three persistent challenges: 1) disaster-domain grounding exposes unique failure modes (damage-semantic grounding, sensor-modality mismatch, and disaster-pipeline composition); 2) agents are doubly bottlenecked by tool selection and argument grounding, where gold tool-order hints improve accuracy by only 1.08-4.40%, and alternative scaffolds yield at most a 3.24% gain; 3) compositional fragility scales with trajectory length, the agent-to-gold gap widening from 7% to 56% on long pipelines. DORA establishes a rigorous testbed for operationally reliable disaster-response agents.
Canopy: Tree-Aware Rollout Scheduling for Agent Reinforcement Learning
Feiyuan Zhang ⋅ LI Pengbo ⋅ Ziniu Li ⋅ Yuhao Jiang ⋅ Di Chai ⋅ Han Tian ⋅ Junhao Wang ⋅ Taiqiang Wu ⋅ Guanhua Huang ⋅ Chaoliang Zeng ⋅ Yihao Liu ⋅ Kai Chen
Rollout generation bottlenecks large-scale RL training for LLM agents. Tree rollout has emerged as an important agentic-RL strategy: by branching from shared intermediate states across reasoning turns, tool calls, and environment observations, it avoids regenerating entire trajectories from scratch and exposes long prefixes for context reuse. However, existing RL and LLM-serving infrastructure remains largely tree-unaware: it treats sibling branches as independent generation requests, making the rollout tree invisible to routing, preemption, and KV-cache management. Consequently, shared prefixes are scattered across servers, discarded with request-private suffixes, or recomputed after cross-server spill. We present Canopy, a tree-aware rollout scheduling system that makes the active rollout tree a first-class serving abstraction without changing sampling, rewards, advantage estimation, or optimizer updates. Canopy consists of three mechanisms: tree-aware routing and retention, which keep siblings near reusable prefixes and protect active prefixes; prefix-preserving partial preemption, which separates shared-prefix KV from private-suffix KV under pressure; and spill-aware transfer-or-recompute admission, which transfers shared-prefix KV to a spilled sibling only when its estimated transfer cost is lower than destination-side recomputation. On long-horizon multi-turn agent workloads, Canopy achieves up to 2.11× rollout-generation speedup; across four representative tree-rollout algorithms, mechanism studies show higher prefix locality and a 51.1% average, 94.1% maximum reduction in full-release shared-prefix invalidations. These results show that exposing rollout-tree structure to the rollout-serving stack can reduce wasted prefill and KV-cache churn, enabling more efficient RL training for LLM agents.
Can We Trust Item Response Theory for AI Evaluation?
Han Jiang ⋅ Sunbeom Kwon ⋅ Jinwen Luo ⋅ Ziang Xiao ⋅ Susu Zhang
AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank systems, select informative examples, and diagnose benchmark quality. However, AI benchmark data often departs from the data regime of human testing, for which standard IRT estimation tools were originally developed: benchmarks typically involve fewer evaluated models, far more items, and capability distributions that may be skewed, clustered, or multimodal. We examine how these regime mismatches challenge the reliability of IRT modeling for AI evaluation. Using item parameters and capability distributions derived from six widely used LLM benchmarks, we simulate response matrices under three common IRT models and compare four estimation tools used in recent benchmark studies: marginal maximum likelihood, Markov chain Monte Carlo, variational inference, and a neural pseudo-Siamese estimator. Across 18,000 simulation conditions, we systematically evaluate computational feasibility, scalability, and the reliability of IRT inferences about model rankings, predicted performance, and item characteristics. Results show that classical estimators can become infeasible in large benchmark settings, whereas scalable estimators can produce unreliable item-level and ranking inferences with small or nonnormally distributed model sets. This study identifies when latent trait models reliably support or risk distorting AI benchmarking claims, and what sample sizes and diagnostics are needed for trustworthy use.
Capacity Allocation at the Source: Sparse Target Optimization for LLM Knowledge Editing
Dongyao Chen ⋅ Jiayun Lei ⋅ Zhiying Deng ⋅ Wei Liu ⋅ Zhiyuan Ji ⋅ Xiaobo sun
Knowledge editing has emerged as a promising approach for efficiently updating embedded knowledge in large language models (LLMs). It first computes an ideal target hidden state that steers the LLM toward the new output, and then edits the model parameters so that the original hidden state is mapped to this target. While most existing methods focus on the second stage of parameter editing in preserving other knowledge, we find that the upstream construction of the target hidden state also plays an important role in allocating the model's limited capacity space. We first establish a theoretical connection between interference in the parameter space and interference between target hidden states, showing that the two are positively correlated. Building on this insight, we propose a simple yet effective $\ell_1$ regularization method that prunes away unnecessary ``tentacles'' of each knowledge vector, retaining only the essential components. With fewer extraneous tentacles, each edit affects less other knowledge, enabling better coexistence. Extensive experiments across multiple LLMs and benchmark datasets verify the effectiveness of our method.
CAPO: A Primal-Dual Framework for Constraint-Aware Prompt Optimization
Victor Ye Dong ⋅ Reid Pryzant ⋅ Yi Liu ⋅ Jian Jiao
Large Language Models (LLMs) are increasingly deployed in agentic contexts, where the model relies on a system prompt to use tools and complete tasks. Implicit in these agentic settings are operational requirements: using the right tools, keeping prompts at reasonable length, achieving parsimonious solution paths, and complying with safety and formatting policies. For many practitioners, assembling domain-specific supervised data to post-train LLMs to satisfy such requirements is infeasible. In this paper, we introduce CAPO (Constraint-Aware Prompt Optimization), an \emph{explicit threshold-constrained} prompt-optimization algorithm based on primal--dual updates. CAPO employs a primal--dual framework to optimize system prompts under explicit constraints via Lagrangian relaxation. Our results across three agentic benchmarks show that CAPO more reliably reaches empirically feasible operating points while improving agentic performance. We also demonstrate that the algorithm generalizes beyond agentic use cases, achieving strong performance on assistant-style evaluations that require satisfying output-format and safety/privacy constraints. Finally, we show that CAPO's rewrite policy can be amortized into DCAPO, a feedback-aware trainable rewriter that matches CAPO accuracy while improving the constrained trade-off.
CASQRec: Collaborative-Adaptive Semantic Quantization for Multimodal Recommendation
Qingfeng Li ⋅ Wenhui Tu ⋅ Wei Liu ⋅ Huaijie Zhu ⋅ Jielong Tang ⋅ Jian Yin
Multimodal recommendation leverages item content, such as text, images, and audio signals, to complement sparse user-item interactions. Existing methods mainly combine content and collaborative signals at the representation level through graph propagation, feature fusion, or auxiliary alignment. However, continuous item embeddings make semantic sharing implicit and provide limited support for transferring behavioral evidence through reusable semantic factors. Discrete semantic codes offer explicit shared variables, but existing code-based recommenders usually treat item-code assignments as fixed outputs of pretrained quantizers, clustering algorithms, or frozen encoders, making the assignment mechanism insensitive to user behavior. We propose \textbf{CASQRec}, a \textbf{C}ollaborative-\textbf{A}daptive \textbf{S}emantic \textbf{Q}uantization framework for multimodal recommendation. CASQRec learns item-code assignments as content-inferable routing decisions supervised by collaborative behavior during training. It constructs a semantic-code graph, where fixed content-derived RQ codes are graph nodes and item-code edges are predicted from item content by a learnable prior router. To make routing behavior-aware while keeping item-code routing independent of item-side interaction history at inference time, CASQRec introduces a collaborative posterior teacher only during training and distills its assignment distributions into the content prior. The resulting prior-induced user-item-code graph allows collaborative evidence to be shared through discrete semantic codes while keeping item routing content-based. Experiments on three public datasets show that CASQRec consistently outperforms representative multimodal recommendation baselines. The code resources are available in the supplementary materials.
Causal discovery aims to uncover the causal relationships among variables beyond correlations from observational data. On time series, a major challenge is nonstationarity, which is typically modeled as regime-based shifts or drifting causal strength. In contrast, we consider non-stationarity arising from varying cause-to-effect delay, and introduce a causal model that is identifiable with few assumptions. To discover causal relationships with varying causal delay in practice, we formalize the problem in terms of the algorithmic model of causation, and propose STRETCH to discover causal graphs. We show on synthetic data that our method correctly recovers the causal relationships, and show the relevance of this causal model on a case study on the impact of El Niño climatic phenomenon on Indian monsoon.
CausalDriveBench: Evaluating Causal Reasoning in Vision-Language-Action Models for Autonomous Driving
Narendiran Chembu ⋅ Navvrat Rao ⋅ Shreedhar Kodate ⋅ Gayatri S Banda ⋅ Arko Sarkar ⋅ Abhinav Khanna ⋅ Rajarshee Das ⋅ Umesh Kanala ⋅ Siddharth Khandelwal ⋅ Kumar Aman ⋅ Aish Dubey ⋅ Kaustubh Beedkar ⋅ Jain
Vision-Language-Action (VLA) models for autonomous driving produce natural-language reasoning alongside predicted trajectories, but whether this reasoning reflects the causal structure of the scene remains untested. We introduce CausalDriveBench, an evaluation framework grounded in Pearl's Causal Hierarchy (PCH) that tests causal reasoning in driving-specific VLAs through structured visual question answering (QA) and alternative-trajectory prediction. To this end, we construct causal scene graphs over nuScenes that distinguish causally active, dormant, and distractor entities, separating perceptual salience from causal relevance. The benchmark spans all four rungs of PCH (association, intervention, and counterfactual along with causal discovery) for QA generation. For the higher rungs, we additionally provide reference trajectories under specified scene modifications, enabling action-level verification that complements reasoning-level evaluation. In total, the benchmark contains 7,285 verified causal QA pairs and 1,000 counterfactual trajectories derived from nuScenes. We evaluate 10 driving-specific VLAs and 3 general-purpose VLMs, and report three findings. First, the best model reaches only 70.6\% QA accuracy, and 4 of 13 models score below random chance. Second, comparing each driving VLA to the general-purpose VLM that shares its language backbone, the cost of driving fine-tuning ranges from 2 to 34 percentage points on causal QA, with post-training design explaining the spread. Third, causal QA and trajectory accuracy are statistically uncorrelated across models: under counterfactual prompts, predicted trajectories either over-react or collapse onto the observed-scene baseline. Taken together, these results show that neither fluent rationales nor accurate observed-scene trajectories constitute evidence of causal understanding.
Causal Inference for Sequential Settings under Interference and Latent Confounding
Phevos Paschalidis ⋅ Constantinos Daskalakis ⋅ Devavrat Shah
We study causal inference under outcome interference for sequential, observational settings. We consider settings where the binary outcomes over $N$ units are Markovian across $T$ time steps; at each time step, the outcomes of $N$ units have pairwise dependencies captured through an Ising model; and, each outcome is impacted through a latent external field capturing effects of latent confounders. Similar to panel data literature, these latent confounders are modeled to have a low-rank factor structure. Our data is a single sample from this high-dimensional distribution. To estimate causal quantities of interest, we provide a computationally efficient method based on Maximum Pseudolikelihood Estimation (MPLE) for learning the model parameters. Under reasonable assumptions, we establish non-asymptotic consistency for parameter estimation. Therefore, sampling from the learnt model enables faithful estimation of causal quantities of interest. We demonstrate the efficacy of the method through synthetic experiments as well as a real-world case-study investigating causal effects of vaccine rates on COVID-19 death rates within US counties nationwide.
Causal discovery, the problem of inferring the direction of causality, is generally ill-posed. We use the language of structural causal models (SCM) to show that assuming that the causal relations are acyclic and invariant across multiple environments (e.g., the way minimum wage affects employment rate is stable across different geographical regions), \textit{only} two auxiliary environments are sufficient to infer the causal graph for arbitrary nonlinear mechanisms. Moreover, we demonstrate that this implies identifiability of the SCM functional mechanisms: as a corollary, we show that \textit{two} auxiliary environments are sufficient to guarantee correct counterfactual inference. We empirically support our theoretical results on synthetic data.
CellMSA: Context Modeling for Single-Cell Representation Learning
Suyuan Zhao ⋅ Minghao Liu ⋅ Yizhen Luo ⋅ Zaiqing Nie
Single-cell transcriptomics enables profiling of cellular states at unprecedented resolution, but its high dimensionality, sparsity, and technical batch effects pose significant challenges for representation learning. Existing single-cell foundation models typically encode each cell independently or only model cells from the same batch for denoising, thereby underutilizing the rich relational information across batches and cell types to model gene expression patterns. We argue that single-cell models can benefit from more informative cell-context modeling. By comparing consistency and variation across cells, models can capture fine-grained gene-gene dependencies associated with cell states, which are essential for learning high-quality representations. Inspired by the use of multiple sequence alignment (MSA) context in protein modeling, we propose \textbf{CellMSA}, a single-cell representation learning framework that introduces an MSA-inspired inductive bias into transcriptomic modeling. For each target cell, CellMSA retrieves relevant cells from different batches and biologically related cell types as context, and summarizes cross-cell patterns into a context-dependent gene-pair representation. This representation is then injected into a pair-aware target-cell encoder for fine-grained representation learning. We pretrain CellMSA on a large-scale human single-cell corpus of approximately 109 million cells. Experiments show that our framework consistently outperforms existing methods across multiple benchmarks.
Chain-of-Thought Oversight Should Not Treat Faithfulness as Monitorability
Sichao Li ⋅ Sai Ma ⋅ Xiyang Hu ⋅ Chudi Zhong ⋅ Tongliang Liu
Chain-of-thought (CoT) oversight is often motivated by a simple premise: if a model exposes its reasoning, then its behavior should be easier to monitor. This position paper argues that the premise conflates several properties that can come apart in practice. To make this separation explicit, we distinguish four properties: reasoning visibility (whether a trace is disclosed), mechanistic faithfulness (whether reasoning step content is supported by internal activations), behavioral monitorability (whether a specified monitor can recover a target property on held-out cases), and mechanistic localizability (whether a compact causal mechanism can be found). Claims about one require evidence specific to that property. We support this position with two mirror-image case studies on open-weight reasoning models. In security vulnerability reasoning, probes over internal activations confirm that key reasoning steps are faithfully represented, yet a reference-assisted LLM jury cannot reliably judge whether the resulting analysis is correct on held-out cases. In authority-hint factual reasoning, the pattern reverses: an external monitor detects the target behavior near-perfectly, but repeated mechanistic analyses fail to isolate a compact causal circuit. These results show that internal support does not guarantee monitorability, and monitorability does not guarantee localizability. We argue that CoT oversight papers should adopt property-specific reporting: every claim should state the reasoning access tier, target property, monitor class, monitor competence, held-out protocol, baseline, and evidence type.
Imagine Mr. Bean stepping into Tom and Jerry--can a video model generate interactions that stay faithful to each character's identity, behavior, and style across worlds? We present MiMix, a unified framework for multi-character text-to-video generation that goes beyond visual appearance to capture character-specific personas generalizing to unseen scenes and partners. To evaluate this, we introduce the Cross-Character Generalization Benchmark, a social stress test placing each character in novel scenes with previously uncoexistent partners. MiMix combines Cross-Character Embedding (CCE), which disentangles identity and behavior via text-aligned annotations, with Cross-Character Augmentation (CCA), which synthesizes cross-style training data while preserving native appearance. Experiments show consistent gains in identity fidelity, interaction quality, and style consistency over prior personalized and general video models.Additional results and videos are available on our project page: https://mi-mi-x.github.io.
CHARM+: Cross-Hardware Attention with Re-Merge Multistream Mechanism
Xinhang Zhang ⋅ boning zhang ⋅ Chengchun Liu ⋅ Chunpu Li ⋅ Lei Sun ⋅ Weifeng Zhang ⋅ Limin Xiao
As sequence lengths continue to grow, distributed attention based on sequence parallel (e.g., Ring-Attention) has become a mainstream solution for large language models (LLMs). Nevertheless, inherent communication dependencies limit computation–communication overlap, reducing GPU utilization. Moreover, in cross-hardware (NUMA or node) architectures, due to imbalanced communication bandwidth, frequent communication synchronization will limit overall performance. Based on these insights, we propose CHARM+, Cross-Hardware Attention with Re-Merge Multistream Mechanism. We reorganize sequential communications into several levels and achieve intra-level parallelism and computation overlap. Furthermore, we decouple cross-hardware communication from intra-hardware communication, eliminating synchronization overhead caused by bandwidth imbalance. Forward and backward algorithms were evaluated on multiple hardware architectures and consistently outperformed state-of-the-art methods. Compared to Ring-Attention, our forward and backward algorithms achieve average speedups of $1.3\times$ and $1.7\times$ on a dual-NUMA PCIe architecture, and $2.8\times $ and $3.9\times$ on a dual-node NVLink architecture, respectively.
ChiP-STAR: Spatial-Topological Attention for Pre-trained Generative Chip Routing
Junfeng Liu ⋅ Xingquan Li
Routing is foundational in VLSI physical design, directly shaping timing, power, and area. Recent pre-trained Seq2Seq models recast routing as autoregressive generation over DFS-serialized trees, but inherit a text backbone blind to the physical space their tokens describe: DFS decouples sequence position from physical coordinate, breaking spatial reasoning at branching points, in load selection, and along the autoregressive trajectory. We name these failures the \emph{Three Spatial Blindnesses} and propose ChiP-STAR, which augments a T5Gemma backbone with three lightweight modules: a dual-frame rotary encoding coupling sequence order with 3D proximity, a Fourier-factorized attention prior with memory linear in sequence length, and a cumulative coordinate-noise schedule closing the train--inference gap, together adding under $0.5\%$ parameters. We rebuild the AiEDA corpus into the largest geometry-annotated routing dataset to date ($6.4$M nets, $1.8$B tokens, $556$\,GB). ChiP-STAR-Large surpasses iPCL-Large by $5\%$ on Leaf IoU and Connectivity, and commercial ECO signoff yields $3\%$ shorter wirelength and $11\%$ fewer vias. Our code is publicly available at \url{https://anonymous.4open.science/r/iPCL-R-0778/}.
ChronosAlign: A Large-Scale, Updatable Benchmark for Decomposing Temporal Alignment in LLMs
Sanjay Govindan ⋅ Maurice Pagnucco ⋅ Yang Song
Large Language Models (LLMs) struggle with temporal alignment, often producing outdated time-sensitive facts as real-world knowledge evolves. Yet the field's ability to measure this misalignment is itself fragile. Existing benchmarks are static snapshots, rely on rigid templates, are small, lack provenance, conflate willingness-to-answer with recall, and use phrasings that newer models have likely memorised. We argue that progress on temporal alignment is bottlenecked less by data scarcity than by evaluation design, and we contribute on both fronts. We present ChronosAlign, a natural language, programmatically updatable dataset and evaluation framework for temporal alignment in LLMs; generated from a Wikidata and Wikipedia corpus of ~500,000 question-answer pairs covering ~27,000 unique questions across sport, politics, and culture from 2000 to 2025. Each question carries a Wikipedia URL and a documented update procedure, so the same evaluative claim can be re-tested against future models without re-authoring the benchmark. We couple ChronosAlign with a framework focussed on three prompting modalities, explicit (year-anchored), relative ("current" anchored), and invariant (untimed control), whose disagreement isolates which component of temporal alignment a model fails. We benchmark sixteen LLMs surfacing three findings: (i) a sharp post-2022 decay in explicit recall even for models with later training cutoffs, (ii) systematic misalignment between explicit and relative recall, with some models better aligned to explicit year prompts and others to relative anchors, and (iii) a confounding effect of refusal behaviour. Our generation and update code can be found at: https://github.com/sg-sy/ChronosAlign. Our dataset can be found at: https://huggingface.co/datasets/sg-sy/ChronosAlign.
Circuit-Level Knowledge Distillation for Large Language Models
Yashuo Luo ⋅ Tongxu Wang ⋅ Siyuan He ⋅ Chunyu Wei
Existing knowledge distillation methods supervise only what a student outputs, leaving how it computes those outputs unconstrained, so students may match teacher behavior through entirely different internal mechanisms. We propose Meta Circuit Distillation (MCD), which reframes distillation as the explicit transfer of reasoning circuits, the structured computational pathways uncovered by mechanistic interpretability. MCD represents each model's computation as a transcoder-based attribution graph and aligns teacher and student circuits via an optimal-transport pathway-matching loss that plugs into any standard distillation pipeline. We prove that minimizing this loss bounds the discrepancy in MLP-level computation between teacher and student. Across two model families, six distillation objectives, and three instruction-following benchmarks, MCD delivers consistent improvements that transfer to mathematical reasoning, with the largest gains under aggressive compression.
CITE: Anytime Valid Statistical Inference in LLM Self-Consistency
Hirofumi Ota ⋅ Naoto Iwase ⋅ Yuki Ichihara ⋅ Junpei Komiyama ⋅ Masaaki Imaizumi
Large language models often improve reasoning by sampling multiple outputs and aggregating their final answers, but precise and efficient control of error levels remains a challenging task. In particular, deciding when to stop sampling remains difficult when the stopping rule is data-dependent and the set of possible answers is not known in advance. We study anytime-valid certification of a prespecified target answer as the unique mode of the model’s response distribution, a guarantee distinct from answer correctness. We propose the Certification by Intersection-union Testing with E-processes (CITE) algorithm, which provably controls false certification at any prescribed level under arbitrary data-driven stopping, without requiring prior knowledge of the answer category set. We also prove an category-set-size-free stopping-time rate, establish matching minimax lower bounds up to constants in the main regime, and extend the construction to confidence-weighted voting. Simulations and LLM self-consistency experiments show empirical error control and improved certification in diffuse-tail settings.
Class-Domain Incremental Learning with Extensible Multi-Center Modeling
Xuetong Yang ⋅ Yuxiang Yan ⋅ Zhiyuan Zhou ⋅ Xin Gao ⋅ Guanghao Li ⋅ Jian Pu
Incremental learning aims to continuously learn from dynamic data streams without suffering from catastrophic forgetting. However, existing methods may fail in challenging scenarios where class and domain spaces expand simultaneously, struggling to learn continually and stably from fragmented data streams. To address this challenge, we propose EMCift (Extensible Multi-Center Modeling with Dual-Path Drift Harmonization), a novel model that treats heterogeneous data across domains and semantic spaces as co-existing multi-view observations, enabling the construction of global knowledge from local observations. Our approach organizes extensible multi-center prototypes as a tree hierarchy, allowing the global class representation across different domains to grow organically from fragmented local data. To sustain this dynamic topology, we introduce a dual-path drift harmonization strategy that balances structural stability with adaptive plasticity. Specifically, one path captures essential updates to enable prototype evolution, while the other path leverages a conditional adversarial mechanism to robustly initialize domains, effectively preventing catastrophic forgetting caused by class confusion. Extensive experiments on three datasets demonstrate the state-of-the-art performance of our proposed model in the class-domain incremental learning scenario. Code will be released upon acceptance.
Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
Alexander Panfilov ⋅ Peter Romov ⋅ Igor Shilov ⋅ Yves-Alexandre de Montjoye ⋅ Jonas Geiping ⋅ Maksym Andriushchenko
We show that AI agents are capable of discovering novel algorithms for adversarial attacks against LLMs, advancing the state of the art on white-box jailbreaking and prompt injection evaluations. We deploy frontier agents, such as Claude Code and Codex, in an autoresearch loop with access to a library of 30+ prior methods and an evaluation script with a fixed compute budget. We show this pipeline to be effective in jailbreaking OpenAI's GPT-OSS-Safeguard-20B and in prompt injections against Meta-SecAlign-70B, an adversarially robust model. For GPT-OSS-Safeguard, the best agent-discovered method achieves up to 80\% attack success rate on CBRN queries, compared to <50\% for existing methods. For SecAlign, it achieves 100\% ASR, while the best prior automated methods only achieve 82\%. Notably, in our setting, attack methods are developed on unrelated surrogate models for a pure random-target token-forcing task, yet generalize directly to prompt injection on the adversarially trained model. Finally, we trace the lineage of methods developed during autoresearch, characterizing the agents' strategies and failure modes. Adversarial ML has long held that defenses must be evaluated against attacks tailored to them; autoresearch automates this principle, and we argue it should be the minimum bar for defense evaluation going forward.
CLaW: Codec-Guided Adaptive Latent Watermarking for Traceable Diffusion Image Generation
Huangchao Song ⋅ Lingyun Yu ⋅ Cong Li ⋅ Peiqi Jiang ⋅ Mingqi Fang ⋅ Hongtao Xie
With the rapid advancement of text-to-image diffusion models, increasingly realistic AI-generated content has raised serious concerns about misuse and copyright infringement. Digital watermarking offers a promising solution by enabling user-level traceability, yet existing methods remain difficult to deploy at scale, hindered by costly model retraining or detection, poor effectiveness–fidelity trade-offs, and limited robustness to image transformations. To address these challenges, we propose CLaW (Codec-guided Latent Watermarking), a robust and efficient watermarking framework for traceable diffusion image generation with frozen diffusion backbones and inversion-free detection. Specifically, CLaW first uses a pretrained watermark codec to map each watermark message to a latent residual, enabling low-cost sampling-time injection. Then, the residual is injected within a late denoising window, where image semantics are better preserved and the watermark signal is less perturbed by subsequent denoising updates, yielding a better balance between watermark effectiveness and image fidelity. Furthermore, we introduce a decoder-guided adaptive injection mechanism that uses decoding confidence as feedback to dynamically calibrate watermark strength, reinforcing weak watermark signals and improving robustness under image transformations. Extensive experiments show that CLaW can preserve visual fidelity and improve average F1 under image transformations, achieving a 6.34% gain over the state of the art.
Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
Bowen Ye ⋅ Rang Li ⋅ Qibin Yang ⋅ Yuanxin Liu ⋅ Linli Yao ⋅ Hanglong Lv ⋅ Zhihui Xie ⋅ Chenxin An ⋅ Lei Li ⋅ Lingpeng Kong ⋅ Qi Liu ⋅ Zhifang Sui ⋅ Tong Yang
Large language models are increasingly deployed as autonomous agents for multi-step workflows in real-world software environments. However, existing agent benchmarks are limited by trajectory-opaque grading, underspecified safety and robustness evaluation, and narrow coverage of modalities and interaction paradigms. We introduce Claw-Eval, an end-to-end evaluation suite addressing these gaps with 300 human-verified tasks spanning 9 categories across three groups: general service orchestration, multimodal perception and interaction, and multi-turn professional dialogue. To enable trajectory-aware grading, each run is recorded through three independent evidence channels: execution traces, audit logs, and environment snapshots, yielding 2,159 fine-grained rubric items. The scoring protocol evaluates Completion, Safety, and Robustness, with Average Score, Pass@k, and Pass^k across three trials to distinguish genuine capability from lucky outcomes. Experiments on 14 frontier models show that: (1) Trajectory-opaque evaluation is systematically unreliable, missing 44% of safety violations and 13% of robustness failures detected by our framework. (2) Capability does not imply consistency, with Pass@3 remaining stable under error injection while Pass^3 dropping by up to 24 percentage points. (3) Agent capability is strongly multi-dimensional, with model rankings varying across task groups and metrics, indicating that our heterogeneous evaluation coverage is essential. Claw-Eval highlights directions for developing agents that are not only capable but reliably deployable.
CLeaR: A Unified Framework for Resolving the Leakage–Degradation Dilemma in Style Transfer
Teng Zhou ⋅ Yunhao Chen
Style transfer aims to render target content in the style of a reference image, but existing methods often suffer from content leakage, where objects, layouts, or semantics from the style reference appear in the generated output. Although prior data-driven and training-free methods can reduce leakage, they often face a leakage--degradation dilemma: stronger content suppression may weaken style fidelity, while richer style preservation may reintroduce unwanted reference content. We identify this dilemma across the full style-transfer pipeline, including feature separation, feature-space grounding, and diffusion generation. To address these issues, we propose CLeaR, a training-free framework for content-leakage-resistant style transfer. CLeaR first uses Orthogonal Subspace Projection to define content-reduced style targets in each vision foundation model (VFM) feature space. It then performs Ensemble Inversion, which optimizes a shared pixel-space style anchor satisfying style constraints across multiple VFMs. Finally, Energy-Guided Calibration maintains style alignment during diffusion sampling by steering the denoising trajectory toward the ensemble-defined style manifold. We further provide a theoretical analysis showing that the style-anchor estimation error decreases with the number of VFMs. Experiments on StyleBench demonstrate that CLeaR improves style alignment, reduces content leakage, and achieves better LLM-as-Judge evaluation compared with existing methods.
CLEAR: Complementary Tripartite Play with Bayesian Calibration for Semi-Supervised Edge Classification
Zhipeng Sun ⋅ Fanchun Meng ⋅ Jiazhen Huang ⋅ Yongpeng Zhang ⋅ Tao Ren ⋅ Yifan Wang ⋅ Wei Ju ⋅ Xiao Luo
This paper studies the problem of semi-supervised edge classification, which aims to identify edge relations using both labeled and unlabeled data. This problem is highly challenging due to the inherent asymmetry of edge relations and strong prediction biases from imbalanced homophilous and heterophilous edges. Towards this end, this paper proposes a novel approach named Complementary Tripartite Play with Bayesian CaLibration (CLEAR) for semi-supervised edge classification. The core of our CLEAR is to incorporate asymmetric semantic branches and a meta-teacher into a tripartite-play framework, enabling reliable guidance and optimization under label scarcity. In particular, our CLEAR introduces a topological branch and a contextual branch which extracts graph semantics, generating pseudo-labels in complementary views. To improve the reliability of pseudo-labels, we filter low-quality pseudo-labels using epistemic uncertainty, and then calibrate posterior distributions using the priors induced by neighborhood information. More importantly, we utilize a meta-consoler to guide the optimization of two branches by outputting the objective coefficients, ensuring a reliable pseudo-labeling process in a tripartite play framework. Extensive experiments on benchmark datasets validate the superiority of the proposed CLEAR in comparison to various baselines. Our code is available at https://anonymous.4open.science/r/CLEAR-D1D7/.
Closing the Approximation Gap in Simulation-free Latent SDEs
Henry Smith ⋅ Brian L Trippe ⋅ Scott Linderman
Recovering dynamical systems from noisy observations is a recurring challenge across scientific domains, including neuroscience and physics. Latent stochastic differential equations (SDEs) address this by modeling the system as an unobserved state that evolves according to a learnable SDE and generates the observations. Variational inference (VI) provides a tractable objective for fitting latent SDEs. Traditional VI algorithms evaluate this objective by numerical simulation over a time discretization, trading fidelity for computational cost. A recent class of algorithms, simulation-free VI, sidesteps this tradeoff by parameterizing the posterior through its instantaneous marginals rather than its drift. In this work, we show that the efficiency of simulation-free VI algorithms comes at a price: their parameterizations restrict the approximate posterior to a subset of the SDEs available to simulation-based methods, degrading both posterior inference and parameter learning. We propose Helmholtz-SDE, a simulation-free VI algorithm that closes this gap by optimizing over path laws compatible with a prescribed collection of marginals. Helmholtz-SDE recovers dynamics more faithfully than prior simulation-free methods, with the largest gains under high posterior uncertainty. It further matches the performance of simulation-based VI at a fraction of the runtime.
Coarse-to-Real: Generative Rendering for Populated Dynamic Scenes
Gonzalo Gomez-Nogales ⋅ Yicong Hong ⋅ Chongjian GE ⋅ Peiye Zhuang ⋅ Marc Comino-Trinidad ⋅ Dan Casas ⋅ Yi Zhou
Traditional rendering pipelines rely on complex assets, accurate materials and lighting, and substantial computational resources to produce realistic imagery, yet they still face challenges in scalability and realism for populated dynamic scenes. We present C2R (Coarse-to-Real), a generative rendering framework that synthesizes real-style urban crowd videos from coarse 3D simulations. Our approach uses coarse 3D renderings to explicitly control scene layout, camera motion, and human trajectories, while a learned neural renderer generates realistic appearance, lighting, and fine-scale dynamics guided by text prompts. To overcome the lack of paired training data between coarse simulations and real videos, we adopt a two-stage synthetic-real domain-hedging strategy that first learns a strong generative prior from large-scale real footage, then introduces controllability by using a small amount of paired synthetic coarse-fine data to anchor shared implicit spatio-temporal features across domains. The resulting system supports coarse-to-fine control, generalizes across diverse CG and game inputs, and produces temporally consistent, controllable, and realistic urban scene videos from minimal 3D input.
Code2World: A GUI World Model via Renderable Code Generation
Yuhao Zheng ⋅ Li'an Zhong ⋅ Yi Wang ⋅ Rui Dai ⋅ Kaikui Liu ⋅ Xiangxiang Chu ⋅ Linyuan Lü ⋅ Philip Torr ⋅ Kevin Qinghong Lin
GUI agents interact with environments by perceiving interfaces and executing actions. As a virtual sandbox, the GUI World model empowers agents with human-like foresight by enabling action-conditioned prediction. However, existing text/pixel-based approaches struggle to achieve both high visual fidelity and fine-grained structural controllability. To this end, we propose \textbf{Code2World}, a vision-language coder that simulates the next visual state via \textbf{renderable code generation}. Specifically, to address the data scarcity problem, we construct \textbf{AndroidCode} by translating GUI trajectories into high-fidelity HTML and refining synthesized code through a visual-feedback revision mechanism, yielding a corpus of \textbf{over 80K} high-quality screen-action pairs. To adapt existing VLMs into code prediction, we first perform SFT as a cold start for format layout following, then further apply \textbf{Render-Aware Reinforcement Learning} which uses rendered outcome as the reward signal by enforcing visual semantic fidelity and action consistency. Extensive experiments demonstrate that Code2World-8B achieves the top-performing next UI prediction, rivaling the competitive GPT-5 and Gemini-3-Pro-Image. Notably, \textit{Code2World significantly enhances downstream navigation success rates in a flexible manner}, boosting Gemini-2.5-Flash by {+9.5\%} on AndroidWorld navigation.
Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features
Dae Ung Jo ⋅ Jongin Lim ⋅ YoungJoon Yoo ⋅ Daeho Um
Cross-modal knowledge distillation transfers knowledge from a teacher modality to a student modality. Existing feature-level alignment methods typically assume that teacher and student features reside in structurally alignable representation spaces. However, this assumption does not hold when cross-modal features are structurally heterogeneous and lack clear unit-level correspondence, such as 2D spatial visual grids and 1D temporal audio sequences, thereby limiting the applicability of feature-level alignment. To address this challenge, we propose a cross-modal distillation framework that enables effective knowledge transfer across structurally heterogeneous feature spaces via a vector-quantized codebook. Specifically, teacher features are abstracted into a set of vector-form codes regardless of their original feature structure, and the selected codes serve as concept-level anchors for student learning. Code selection is guided by both task relevance and student compatibility, allowing the student to receive transferable teacher knowledge without requiring direct unit-level feature alignment. Experimental results across diverse cross-modal distillation scenarios demonstrate the effectiveness of the proposed framework on classification and semantic segmentation tasks.
CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models via Structured Code Completion
Liang Zhen ⋅ Wentao Chen ⋅ Hai Huang
Large language models have achieved remarkable capabilities across diverse domains, yet their safety alignment remains vulnerable to jailbreak attacks. In this work, we identify a previously underexplored failure mode—safety generalization lag—where alignment trained predominantly on natural language fails to transfer to the code domain. We show that this lag induces a code-completion blind spot, allowing malicious intent embedded within syntactically valid code to evade safety mechanisms. To exploit this vulnerability, we propose CodeMimicry, a fully automated black-box jailbreak framework that generates structured, object-oriented code prompts to induce harmful outputs via code completion. Experiments on 8 state-of-the-art commercial LLMs demonstrate that CodeMimicry achieves a 96.25\% attack success rate with 1.51 queries on average, significantly outperforming both template-based and optimization-based baselines. Beyond empirical performance, we provide a mechanistic analysis of code-based jailbreaks through latent space representations, including projection onto refusal-related directions and activation steering. This analysis offers an explanation of how CodeMimicry bypasses safety mechanisms in code-related domains. Our findings reveal a weakness in current safety alignment and highlight the need for robust alignments in structured domains such as code.
Code Review Bench: An Automated Benchmark for Evaluating the Software Factory
Ashley Zhang ⋅ Narmeen Oozeer ⋅ Aleksandr Zverianskii ⋅ Daniel Ng ⋅ Tasana Pejovic ⋅ Antía García ⋅ Jacob Clyne ⋅ Ravpreet Setia ⋅ Akshay Utture ⋅ Etan J Ginsberg ⋅ Shriyash Upadhyay ⋅ Fazl Barez
As AI agents ship more code with minimal human oversight, code review becomes the critical quality-control checkpoint — but one without a reliable, objective verifier: review quality is nuanced and ultimately a judgment call by the PR author. Existing benchmarks are structurally misaligned with this: annotators are third-party reviewers, like the bots they evaluate, not the PR authors who actually decide what to fix. They also rely on small, static issue sets that require full re-evaluation as tools update, limiting how frequently progress can be tracked. We introduce CodeReviewBench, built on a fundamentally different signal: revealed developer preferences — if a developer acts on a bot's suggestion, we treat it as useful; if they ignore it, we do not. This signal is mined daily from the public GitHub event stream, yielding 500K+ judged PRs across 18 tools, and refreshes continuously as tools evolve. We validate the developer preference against an independent offline benchmark: the two agree on precision rankings and independently detect the same product updates within days of vendor releases. Together, these form the benchmark.
CoE-Agent: Co-Evolving Patient-Doctor Agents via Interactive Policy Graph Optimization for Clinical Decision Making
Guolin Huang ⋅ Wenting Chen ⋅ Linlin Shen
Clinical Decision Making (CDM) requires integrating heterogeneous information across sequential diagnostic stages while interacting with patients whose behaviors are often non-stationary and unreliable. However, existing medical agents are typically trained against cooperative patient simulators and tackle each clinical task in isolation, making them fragile to dynamic patient behaviors and unable to capture inter-task dependencies essential for globally coherent decision-making. To address these issues, we propose CoE-Agent, a co-evolving patient–doctor agentic framework that learns globally coherent CDM strategies via interactive policy graph optimization. First, Dynamic Patient–Doctor Co-Evolutionary Learning pairs a Dynamic Patient Agent, which maintains an evolving cognition state to simulate non-stationary behaviors, with an Interactive Policy Graph-Based Doctor Agent that generates DAG-structured policy graphs unifying interactions, tool invocations, and reasoning steps. The doctor agent is then optimized via Interactive Policy Graph Optimization (IPGO), which jointly enforces structural validity and task-aligned correctness through reinforcement learning. Second, Cross-Task Policy Graph Consolidation and Unified Model Learning prunes low-value nodes, aggregates high-value nodes across tasks while filtering redundancy, and distills a unified policy planner with task-specific decision models for globally coherent decision-making. Experiments on MedChain and ClinicalBench show that CoE-Agent surpasses state-of-the-art baselines by 11.13\% and 23.63\% in average score, while achieving up to 19.5$\times$ faster inference. Source code is to be released.
CoffeeBench: A Benchmark for Long-Horizon Strategic Decision-Making in Multi-Agent Economies
Issa Sugiura ⋅ Daichi Hattori ⋅ Kazuo Araragi ⋅ Keita Ogawa ⋅ Shota Onose ⋅ Taro Makino ⋅ Teppei Usuki ⋅ Takashi Ishida
As LLM agents advance beyond short-horizon tasks, evaluating their ability to perform long-horizon strategic decision-making in realistic multi-agent settings remains a key challenge. Existing agent benchmarks typically focus on isolated task completion or interactions with rule-based or fixed-policy counterparts, limiting their ability to capture strategic adaptation and emergent behaviors arising from sustained multi-agent interactions. We introduce \textbf{CoffeeBench}, a benchmark for evaluating LLM agents in a dynamic multi-agent economic environment consisting of farmers, roasters, and retailers forming a multi-stage supply chain. The environment consists of autonomous firms interacting over a multi-month horizon with evolving supply, demand, and pricing, creating sustained competition and supply-chain dynamics. In each evaluation, the evaluated model controls one firm and must make sequential decisions on procurement, production, pricing, and negotiation to maximize cumulative net income. CoffeeBench enables the study of emergent behaviors arising from repeated strategic interactions, including negotiation dynamics, pricing discipline, and coordination failures. We evaluate several frontier and pareto-frontier LLMs on CoffeeBench to analyze their economic performance and strategic behavior. Most evaluated models achieve positive net income, with GPT-5.5 achieving the highest mean net income. In contrast, Claude~Haiku~4.5 frequently yields negative net income, exhibiting an idle-drift failure mode in which agents collapse into inactivity despite coherent plans. We further find that stronger models proactively negotiate with counterparties through frequent messaging, whereas weaker models trade reactively with limited communication. CoffeeBench provides a controllable testbed for studying strategic interaction, coordination, and long-horizon behavior in multi-agent economies.
Collaborative Reasoning Distillation via Cross-Feedback and Coherent Curation
Taehoon Kim ⋅ Seunggeun Cho ⋅ Dongsu Han
Reasoning capabilities are critical for advancing Large Language Models, yet current approaches either require massive computational budgets or struggle to effectively distill reasoning to smaller models. Standard distillation methods rely on outcome-based rewards, failing to distinguish between sound reasoning and lucky guesses. We propose Collaborative Reasoning Distillation (CRD), a framework that enhances reasoning in compact models through three innovations: (1) interactive cross-feedback where teachers iteratively critique each other's reasoning, (2) fine-grained step-wise quality assessment capturing logical validity independent of final answers, and (3) coherence-aware step stitching that synthesizes complementary strengths. Students are trained via Reasoning Quality Optimization (RQO) with budget constraints. Our model, CRD-4B, achieves 97.3\% on MATH-500 and 70.3\% on AIME'25, surpassing baselines while using only 50K training examples, up to 12 times smaller than the datasets of comparable models.
Colour me shocked: Exact Molecular Hessians from MLIPs in O(N) time using sparse differentiation!
Luca Thiede ⋅ Andreas Burger ⋅ Alan Aspuru-Guzik
The Hessian of the energy with respect to the nuclear positions is indispensable in atomistic modelling. However, constructing this matrix requires $O(N)$ Hessian vector products, traditionally limiting high-accuracy Hessians to small systems. Machine learning interatomic potentials (MLIPs) have accelerated atomistic modelling by providing highly accurate energies and forces at $O(N)$ cost, yet the resulting $O(N^2)$ cost of Hessians remains a practical bottleneck for large systems. Based on the insight that we can derive the sparsity pattern for an MLIP's Hessians in closed form, we show in this paper how to use techniques from sparse automatic differentiation to reduce the cost of a local MLIP's Hessians to a system-size-independent number of Hessian-vector products, yielding overall $O(N)$ total cost without any approximations. We benchmark our approach on a variety of systems ranging from alkane chains to water clusters to A$\beta$40 conformers. Depending on the MLIP configuration, we achieve the linear scaling regime already on relatively small systems, resulting in large runtime reductions between $2\times-15\times$ for these systems. This opens up the possibility of scaling high-accuracy MLIP Hessians to very large systems, such as proteins that were previously inaccessible.
On the simplest non-Euclidean manifold, the 1D flat torus, we exhibit a distribution where a recently proposed extension of Riemannian flow matching is globally optimized by a velocity field that pushes samples away from the data distribution. We trace this failure mode to a property we call compatibility; a likelihood-based flow matching objective is compatible if its population minimizer recovers the marginal velocity field induced by the data and coupling. We give a general condition for compatibility, apply it to existing methods, and bound the worst-case error in terms of a residual component vanishing under compatibility. Next, we note that existing compatible objectives put density on velocities that are infeasible on compact manifolds. We propose an objective built from an exponential family over endpoints, with the conditional velocity as sufficient statistic, which is both compatible and supported only on feasible velocities. This manifests as a re-weighting of the flow matching loss which, on constant-curvature manifolds, treats errors in flow speed and direction differently. Empirically, our method recovers the correct velocity field on the test case, outperforms prior methods on high-dimensional product tori, and obtains substantial NLL improvements over prior work on two of four standard spherical geospatial benchmarks at moderate runtime overhead.
CompilerKV: Risk-Adaptive KV Cache Compression via Offline Experience Compilation
Ning Yang ⋅ Chengzhi Wang ⋅ Yibo Liu ⋅ Baoliang Tian ⋅ Haijun Zhang
Prefill-only KV compression freezes a token subset at the end of prefill and decodes from it without further eviction. The retention decision is therefore irreversible, yet existing methods estimate the corrective signals it relies on, per-head reliability and prompt-level compression sensitivity, online from a single noisy prompt. We argue this is the wrong statistical unit: these signals exhibit far higher cross-prompt regularity than within-prompt signal-to-noise. We introduce CompilerKV, a KV-retention policy whose corrective tables are compiled offline from a calibration corpus, reducing online correction after the standard observation-window scan to $O(1)$ lookups plus a budget clamp. We find that compiled retention tables behave as portable architectural priors: rankings transfer across disjoint corpora on four backbones, with mean Spearman $\bar{\rho}=0.90$, and direct model-to-model table transfer costs only $0.4$-$0.8$ LongBench points on average. At a 512-token budget, CompilerKV attains compressed-SOTA on all four backbones, improving over the strongest prefill-only baseline by $+1.67$ points on average, with task-bootstrap 95% CI $[+1.08,+2.37]$. Pressure regimes amplify the gap: under a fixed $512/32\mathrm{k}$ cache ratio, CompilerKV remains the strongest compressed method through 128k RULER, with approximately 73 versus FullKV 79 and SnapKV 38; on 32k NIAH it reaches $0.89$ versus SnapKV $0.42$; and at 32k input, retaining only $1.56\%$ of the prefill KV, batch-16 serving remains feasible where FullKV is OOM.
COMPOT: Calibration-Optimized Matrix Procrustes Orthogonalization for Transformers Compression
Denis Makhov ⋅ Dmitriy Shopkhoev ⋅ Magauiya Zhussip ⋅ Ammar Ali ⋅ Baher Mohammad ⋅ Stamatios Lefkimmiatis
Post-training compression of Transformer models commonly relies on truncated singular value decomposition (SVD) strategies leading to low-rank approximations of the original model weights. However, the underlying single shared subspace modeling can degrade accuracy even at moderate compression. Sparse dictionary learning provides a more flexible union-of-subspaces representation, but existing approaches often suffer from costly iterative dictionary and coefficient updates. We propose COMPOT (Calibration-Optimized Matrix Procrustes Orthogonalization for Transformers), a training-free compression framework that uses a small calibration dataset to estimate a sparse weight factorization. COMPOT employs orthogonal dictionaries that enable closed-form Procrustes updates for the dictionary and analytical single-step sparse coding for the coefficients, reducing significantly the overall optimization cost. To handle heterogeneous layer sensitivity under a global compression budget, COMPOT further introduces a one-shot dynamic allocation strategy that adaptively redistributes layer-wise compression rates. Extensive experiments across diverse architectures and tasks show that COMPOT consistently delivers a superior quality–compression trade-off over strong low-rank and sparse baselines, while remaining fully compatible with post-training quantization for extreme compression. Code will be released upon acceptance.
Compute Allocation Under Model-Provider Competition
Tori Qiu ⋅ Meena Jagadeesan ⋅ Benjamin Laufer ⋅ Hoda Heidari
Competition among LLM providers hinges not only on model accuracy, but also on the latency of serving user requests. However, user demand leads to congestion which increases latency, creating a feedback loop between provider decisions and user choices. In this work, we study a stylized game between two providers who strategically balance allocating fixed compute between training better models and reducing latency. A population of users then chooses between the providers. Our equilibrium analysis reveals that a lower-budget provider can still attract users by taking advantage of the congestion faced by the dominant provider. Relative to a competitive baseline where providers have equal compute resources, compute asymmetry leads the higher-resource provider to divert a greater fraction of their compute from improving accuracy to reducing latency. Moreover, this compute asymmetry strictly harms users. Altogether, our results illustrate how congestion distorts competition between model-providers, leading to subtle implications for market competitiveness, model accuracy, and user utility.
Compute-optimal data scaling for neural surrogates via multi-fidelity training
Paul Setinek ⋅ Felix Koehler ⋅ Fabian Paischer ⋅ Nils Thuerey ⋅ Johannes Brandstetter
Neural surrogates for Partial Differential Equations (PDEs) are usually trained on synthetic data from numerical simulators. As models grow and scaling laws demand larger datasets, the cost of data generation often outweighs the training cost and becomes the bottleneck in the scaling quest. Unlike data in other domains, however, each PDE sample has a tunable per-sample cost that depends on its numerical fidelity, with the resulting cost-accuracy trade-off governed by known convergence rates. We exploit this property by formalizing neural surrogate training as a function of both data quantity and fidelity under a fixed data generation compute budget, treating fidelity as a fine-grained discrete spectrum rather than a binary distinction. Building on this, we show that multi-fidelity training enables Pareto-optimal error scaling against data generation compute. By using simulations at lower fidelities as an additional training signal, our approach consistently outperforms training on the highest fidelity at matched data generation budgets. We validate this across six PDE problems spanning structured grids and unstructured meshes, two distinct error sources (discretization and iterative-solver error), and four surrogate architectures.
Concentrated Gradients Amplify Forgetting: Dominant-direction Projection for Continual Multimodal Learning
Chengxiang Huang ⋅ Haopeng Zhang ⋅ Yuzhe Han ⋅ Rui Dai ⋅ Kaikui Liu ⋅ Xiangxiang Chu ⋅ Piotr Koniusz
Multimodal Large Language Models (MLLMs) trained on sequential tasks suffer from catastrophic forgetting. We study this problem in LoRA-based multimodal continual instruction tuning. By analyzing the diagonal Fisher information of LoRA parameters, we find that forgetting is amplified not merely by shared high-sensitivity parameters across tasks, but by the concentration of updates on those shared parameters, \ie, some tasks distribute their change uniformly across parameters, while others pack most of it into a narrow subset, causing disproportionate interference on the parameters that prior tasks depend on. We further observe that this concentration manifests in the dominant directions of recent gradients. Based on this insight, we propose \textbf{D}om\textbf{i}nant-direction \textbf{G}radient \textbf{Pro}jection (\textbf{DiGPro}), which maintains a short buffer of recent gradient directions, extracts their dominant subspace with adaptive rank selection, and attenuates the parallel component for each optimizer step. Crucially, DiGPro requires no prior-task data or stored subspaces, operating on the current gradient history. Experiments on cross-modal and vision-centric continual learning benchmarks show DiGPro consistently reduces forgetting and attains competitive final results. The code is available in the supplementary materials.
Concept-Aware Wasserstein Routing with Vision-Language Guidance for Few-Shot WSI Classification
Ankit Kumar ⋅ Shounak Das ⋅ Sandeep Kumar ⋅ Kaustubh Atey ⋅ Amit Sethi
Few-shot whole-slide image (WSI) classification remains challenging due to gigapixel-scale images, scarce slide-level labels, and strong morphological heterogeneity. Recent vision-language MIL methods improve data efficiency by transferring knowledge from pathology foundation models, but they often treat patches as weakly related instances and represent each class with a single semantic or visual anchor. We propose a multi-scale vision-language framework that combines prompted pathology encoding, pathology concept learning, Semantic Wasserstein Routing (SWR), and barycentric prototype memory. Using frozen pathology vision-language backbones with lightweight prompt adaptation, our method aligns image patches with concept embeddings through unbalanced optimal transport and builds a semantic patch graph for concept-guided aggregation. In parallel, the prototype memory maintains multiple class-specific Wasserstein barycenters to capture diverse morphological modes. A distance-based fusion strategy integrates slide-level text alignment, concept-guided patch evidence, and prototype-memory matching for robust prediction. Experiments on TCGA WSI classification benchmarks under few-shot settings show consistent improvements over strong MIL and vision-language MIL baselines, demonstrating the benefit of semantic patch routing and multi-prototype reasoning for weakly supervised pathology learning.
Concord-SLAM: Cross-Render Concordance for Boundary-Native Semantic Gaussian SLAM
Danqi Lu ⋅ Changxin Huang ⋅ Jianqiang Li
Three-dimensional Gaussian Splatting (3DGS) has recently become a promising representation for dense SLAM, due to its explicit scene parameterization, efficient differentiable rendering, and high-quality visual reconstruction. Extending 3DGS to semantic SLAM enables the construction of maps that are not only geometrically accurate but also semantically interpretable. However, existing semantic Gaussian SLAM methods usually optimize semantic, depth, and normal renderings through separate supervision objectives. Such a design underexplores the fact that these outputs are generated from the same Gaussian field and should therefore provide a concordant explanation of the underlying scene structure. This limitation is especially problematic around object boundaries and geometric discontinuities, where inaccurate Gaussian growth can cause cross-boundary expansion, boundary blurring, and semantic leakage. To address this problem, we propose Concord-SLAM, a boundary-native semantic Gaussian SLAM framework that establishes concordance among semantic, depth, and normal renderings while promoting boundary-aware Gaussian growth. Our framework contains two key components. First, Cross-Render Concordance (CRC) jointly examines semantic, depth, and normal views rendered from the same Gaussian map, and converts their structural discrepancies in boundary responses, regional continuity, and local geometry into optimization signals. In this way, disagreement among different renderings is used as an intrinsic cue for improving the semantic-geometric coherence of the Gaussian field. Second, Boundary-Native Splatting (BNS) actively inserts new Gaussians near semantic boundaries, depth discontinuities, and salient geometric variations, encouraging primitives to align with physical and semantic boundaries during map construction rather than correcting boundary mixing only through later optimization. Experiments on Replica and ScanNet demonstrate that Concord-SLAM achieves superior performance in camera tracking, scene reconstruction, and semantic mapping.
Ranking models define probability distributions over rankings and are fundamental in applications such as preference learning and decision-making. Without structural assumptions, these models are intractable due to their factorial complexity. Although independence assumptions are key to reducing complexity and enabling interpretability, standard notions of conditional independence do not directly apply to rankings because they are subject to mutual exclusivity constraints. We introduce \emph{relative rank models}, a probabilistic framework based on structured coarsenings of the ranking space that relax these constraints while preserving relative order information. This representation enables a well-defined notion of conditional independence for rankings and allows rankings to be modeled using Bayesian networks. As a result, relative rank models support tractable learning, probabilistic inference, and interpretable representations. We further show that two prominent ranking models based on independence assumptions tailored to rankings arise as special cases, providing a unifying framework for independence in ranking models.
Conformal Prediction for Distribution-to-Distribution Regression
Trung-Khang Tran ⋅ Tuan Hoang ⋅ Viet-Hoang Tran ⋅ Tan Nguyen
Conformal Prediction (CP) provides a model-agnostic framework for constructing finite-sample prediction sets with coverage guarantees and has become a tool for reliable decision-making in high-risk settings. Recent advances have extended Split Conformal Prediction to increasingly complex learning problems, including multilabel, high-dimensional, and functional outputs. In this work, we study CP for Distribution-to-Distribution Regression, a setting of growing interest across many application domains. We show that a naive application of Split CP yields prediction sets that, while marginally valid, are non-interpretable, e.g., lacking a well-defined volume or a tractable sampling procedure, and of limited practical value. To overcome this, we propose a new conformal framework that operates in an orthogonal basis space, producing interpretable prediction sets while preserving coverage guarantees. We further develop an adaptive variant with asymptotic conditional coverage under mild assumptions. Empirical results on synthetic and real-world datasets validate the effectiveness of the proposed approach.
Confounding-Aware Client Selection in Federated Learning via Causal Mediation Analysis
Xiaoyang Yi ⋅ Yuru Bao ⋅ Rihao Chang ⋅ Binhan Yang ⋅ Jian Zhang
Federated Learning (FL) enables collaborative model training across clients while preserving data privacy. However, FL typically relies on voluntary client participation with uniform rewards, which can lead to high-quality clients dropping out due to inadequate incentives and low-quality clients taking a free-ride, both of which degrade overall model performance. However, existing incentive mechanisms fail to capture the true contribution value of a client’s data to the global model, and often skew client selection toward cost compression rather than quality enhancement, thereby exacerbating confounding bias. To address this, we propose CausalAFL, a causal-based framework that reduces information asymmetry and corrects for confounding factors. Specifically, CausalAFL builds a complete causal chain connecting client information, bid, and marginal contribution. By jointly optimizing an inference network and a generative network within a variational inference framework, it precisely adjusts for causal bias. Maximizing the evidence lower bound objective allows CausalAFL to separate true contribution from noise and highlight the value of high-quality clients. A social welfare objective that quantifies both natural direct and indirect effects guides the server to prioritize these clients, resulting in improved global model accuracy and social welfare. Experiments on multiple datasets demonstrate that CausalAFL can guide clients to adjust their bids, leading to winners with higher contributions, confirming its effectiveness.
Consistent Bayesian Spatial Domain Partitioning Using Predictive Spanning Tree Methods
Kun Huang ⋅ Huiyan Sang
Bayesian model-based spatial clustering methods are widely used for their flexibility in estimating latent clusters with an unknown number of clusters while accounting for spatial proximity. Many existing methods are designed for clustering finite spatial units, limiting their ability to make predictions, or may impose restrictive geometric constraints on the shapes of subregions. Furthermore, the posterior clustering consistency theory of spatial clustering models remains largely unexplored in the literature. In this study, we propose a Spatial Domain Random Partition Model (Spat-RPM) and demonstrate its application for spatially clustered regression, which extends spanning tree-based Bayesian spatial clustering by partitioning the spatial domain into disjoint blocks and using spanning tree cuts to induce contiguous domain partitions. Under an infill-domain asymptotic framework, we introduce a new distance metric to study the posterior concentration of domain partitions. We show that Spat-RPM achieves a consistent estimation of domain partitions, including the number of clusters (which may go to infinity), and derive posterior concentration rates for partition, parameter, and prediction. We also establish conditions on the hyperparameters to achieve consistency, offering important practical guidance for hyperparameter selection. Finally, we examine the asymptotic properties of our model through simulation studies and apply it to Atlantic Ocean data.
Constant Term Shrinkage for Federated Learning
Incheol Baek ⋅ Minseo Kim ⋅ Hyeonmin Kang ⋅ Yon Dohn Chung
Federated Learning (FL) enables collaborative model training without centralizing private data, but its performance often degrades under data heterogeneity because local updates become biased toward client-specific data characteristics. In this paper, we propose Constant Term Shrinkage for Federated Learning (FedCTS)}, a simple server-side aggregation method designed to mitigate such update drift. FedCTS decomposes each client update into a constant term and a residual term. We show that the constant term is strongly affected by data heterogeneity. FedCTS therefore evaluates the reliability of the constant term and adaptively shrinks the unreliable constant term toward zero while leaving the residual term unchanged. The shrinkage strength is obtained in closed form from clients' update statistics, without requiring an additional tuning hyperparameter. Experiments across multiple datasets and FL algorithms demonstrate that FedCTS improves robustness under heterogeneous data with negligible additional computation, no extra communication overhead, and strong compatibility with existing FL techniques.
Continual Learning in Modern Hopfield Networks with an Application to Diffusion Models
Ken Takeda ⋅ Masafumi Oizumi ⋅ Ryo Karakida
Generative models, including diffusion models, are increasingly used as foundation models and adapted through sequential fine-tuning, making continual learning an essential problem setting. However, continual learning in such generative models remains poorly understood: after a task change, what aspects of the learned distribution are most easily lost, and what replay samples should be prioritized? We address these questions through the modern Hopfield energy. Recent links between modern Hopfield networks (MHNs) and diffusion models allow analyses in MHNs to be transferred to diffusion models. We introduce intrinsic forgetting as an increase in Hopfield energy after the task change. In tractable settings in an MHN, we prove that high-energy, outlier-like samples undergo a larger energy increase than cluster-like samples, implying that samples located in sharp, isolated basins are more forgettable. We further analyze memory replay and show that replay is particularly effective for high-energy samples, enabling an energy-based selection of replay samples. We validate these predictions in experiments on MHNs and two diffusion models under continual-learning settings: Stable Diffusion and a pixel-space DDPM. In these diffusion models, Hopfield energy tracks reconstruction-based forgetting, and replay experiments reveal energy-dependent mitigation of forgetting that is consistent with the MHN analysis.
Continuous Audio Thinking for Large Audio Language Models
Gyojin Han ⋅ DongJae Lee ⋅ Changho Choi ⋅ Jongsuk Kim ⋅ Junmo Kim
Large audio language models (LALMs) have shown impressive capabilities on diverse audio understanding tasks, ranging from speech transcription to music analysis. However, because LALMs are typically trained to produce text-aligned responses, their hidden states are progressively shaped for text generation rather than for preserving acoustic information. As a result, the diverse acoustic content that audio carries, such as phonetic detail, prosody, sound events, affect, and pitch, is lost along the way and difficult to leverage in the response. We introduce Continuous Audio Thinking (CoAT), a framework that equips audio language models with a continuous latent workspace for organizing acoustic information prior to response generation, grounded by distillation from audio experts. Within the thinking space, the model can utilize the rich acoustic information provided by expert distillation when generating its response. Furthermore, the proposed continuous thinking block can be processed in a single prefill, so CoAT does not require additional autoregressive decoding cost over the baseline. Across three LALMs, Qwen2-Audio, Qwen2.5-Omni-7B, and Audio Flamingo~3, performance gains on a broad benchmark suite spanning audio reasoning, audio understanding, music classification, speech emotion, and speech transcription demonstrate the effectiveness of CoAT. Further analysis confirms that the auxiliary supervision propagates from the thinking positions to the model's textual responses.
Continuous-depth Deep Gaussian Processes
Ying Li ⋅ Sanyou Wu ⋅ Zi Yang ⋅ Qiaochu Xu ⋅ YUHAO LIU ⋅ Zhidi Lin ⋅ Michael Minyi Zhang ⋅ Petar Djuric ⋅ alexandre thiery
Deep Gaussian processes (DGPs) are appealing Bayesian regression models, but their standard discrete-depth parameterization often becomes harder to optimize as depth grows. We argue that depth is better treated as a latent continuous evolution than as a long discrete stack. We therefore propose a continuous-depth formulation in which the latent representation follows a stochastic differential equation with a depth-indexed drift family equipped with a GP-based prior. This reframes the core design choice as the prior placed on that family, rather than adopting a purely neural continuous-depth parameterization as in neural ODE or neural SDE models. We study two instantiations of this idea: CDGP, a direct GP prior over the state-depth drift, and FlowDGP, a flow-evolved prior that captures richer depth dependence at lower training cost. To motivate the redesign, we show that the single-sample DGP objective can become increasingly ill-conditioned under repeated layer composition, while residual DGP is a useful first repair that still remains a finite discrete construction. We also give pathwise sensitivity analysis for the continuous-depth formulation. Experiments on synthetic, benchmark, and large-scale regression tasks show that residual DGP improves over plain DGP, while CDGP and FlowDGP provide a stronger overall story across the regimes we study.
Continuous Latent Diffusion Language Model
Hongcan Guo ⋅ Qinyu Zhao ⋅ Yian Zhao ⋅ Shen Nie ⋅ Rui Zhu ⋅ Qiushan Guo ⋅ Feng Wang ⋅ Tao Yang ⋅ Hengshuang Zhao ⋅ Guoqiang Wei ⋅ Yan Zeng
Autoregressive language models have achieved remarkable success in text modeling, but recent work increasingly challenges the fixed left-to-right generation order. However, existing non-autoregressive alternatives such as discrete diffusion, still struggle to jointly deliver efficiency, scalable representation learning, and global semantic modeling. We propose Cola DLM, a hierarchical continuous latent diffusion language model that learns a stable Text VAE, models a global semantic prior in latent space with a block-causal DiT, and decodes text conditionally. Theoretically, under a unified Markov-path formulation, the diffusion operates latent prior transport rather than token-level observation recovery, thereby decoupling global semantic organization from local token realization. Empirically, Cola DLM exhibits strong scaling behavior for text generation across four research questions, eight benchmarks, strictly matched $\sim$2B-parameter AR and LLaDA baselines, and scaling curves up to $\sim$2000 EFLOPs. Meanwhile, we also explore encoding text in the same continuous latent space as images, providing a feasible path toward unified discrete-continuous generative modeling. Our code will be released.
Continuous Open-ended Discovery and Evolution of Skills as Hierarchical Reward Programs
Richard Bornemann ⋅ Pierluigi V Amadori ⋅ Antoine Cully
A core quality of general intelligence is the ability to open-endedly expand and evolve its set of mastered skills autonomously. While recent Foundation Model (FM) driven approaches have shown promising results towards this goal, they typically rely on significant human-in-the-loop engineering, limiting their transferability to novel environments. To address this, we introduce Continuous Open-ended Discovery and Evolution of Skills as Hierarchical Reward Programs (CODE-SHARP), a framework that leverages FMs to open-endedly grow and evolve an archive of Python programs encoding skills to train a generalist agent policy entirely from scratch via reinforcement learning, directly from source code. These programs, termed Skills as Hierarchical Reward Programs (SHARPs), each encode a local success condition and a set of prerequisites delegated to previously discovered SHARPs. At runtime, SHARPs dynamically route the agent through their prerequisite chain based on the current state, rewarding each completion along the way, requiring the agent to learn only the marginal behaviour each new SHARP introduces, enabling efficient learning of long-horizon skills without any pre-defined rewards. On Craftax-Classic and XLand, agents trained fully autonomously by CODE-SHARP outperform previous works by 6x and 2.6x in median performance and are the only agents capable of crafting iron tools and mining diamonds. Scaled to Craftax-Extended, CODE-SHARP trains a generalist agent on over 90 discovered SHARPs, enabling the agent to solve challenging long-horizon tasks zero-shot, matching agents trained on ground-truth rewards.
Continuous Personalized Diffusion Model via Spinor-Component Forward Geometry
GeonWoo Jeong ⋅ SEUNG HUN OH ⋅ Sang Beom Chu ⋅ Wonjunkim ⋅ Bokyoon Na
Existing conditional diffusion models typically inject condition information into the reverse-time denoising process, either as model inputs or through guidance. Although effective for conditional fidelity, this design burdens the denoiser when compact continuous signals, such as user preference ratios or mixture degrees, must support both image reconstruction and condition-wise separation, often degrading sample quality. In this work, we reinterpret this limitation not as a problem of condition representation, but as a problem of condition placement. We propose the Continuous Personalized Diffusion Model (CPDM), which shifts the locus of conditioning from reverse-time signal injection to forward geometry formation by coupling bounded spinor-component coordinates with normalized image-space drift directions. The resulting condition-dependent forward drift forms condition-specific terminal geometry, from which continuous changes along the spinor-component coordinate are reflected as continuous visual transitions rather than endpoint interpolation. Experiments show that, even with a single scalar condition, CPDM achieves competitive performance compared with standard and higher-capacity reverse-conditioning baselines and supports continuous generation under compact scalar control. These results suggest forward geometry design as a practical alternative to conventional reverse-time conditioning for compact continuous personalization.
Contrastive Identification and Generation in the Limit
Xiaoyu Li ⋅ Andi Han ⋅ Jiaojiao Jiang ⋅ Junbin Gao
In the classical *identification in the limit* model of Gold [Inf. Control, 1967], a stream of positive examples is presented round by round, and the learner must eventually recover the target hypothesis. Recently, Kleinberg and Mullainathan [NeurIPS 2024] introduced *generation in the limit*, where the learner instead must eventually output novel elements of the target's support. Both lines of work focus on positive-only or fully labeled data. Yet many natural supervision signals are inherently relational rather than singleton: comparative experiments, A/B tests, side-by-side judgments, and similarity–dissimilarity annotations produce observations that encode relationships between examples rather than labels of individual ones. This motivates us to initiate the learning-theoretic study of *contrastive identification and generation in the limit*, where the learner observes a *contrastive presentation* of data: a stream of unordered pairs $\\{x,y\\}$ satisfying $h(x)\ne h(y)$ for an unknown target binary hypothesis $h$, but which element is positive is hidden from the learner. We first present three results in the noiseless setting: an exact characterization of contrastive identifiable classes (a one-line geometric refinement of Angluin's tell-tale condition [Angluin, Inf. Control 1980]), a combinatorial dimension called *contrastive closure dimension* extending the closure dimension of Raman et al. [COLT 2025] and exactly characterizing uniform contrastive generation with tight sample complexity, and a strict hierarchy in which contrastive generation and text identification are mutually incomparable. We then prove a sharp *reversal* under finite adversarial corruption: there exist classes identifiable from contrastive pairs under any finite corruption budget by a single budget-independent algorithm, yet not identifiable from positive examples under even one corrupted observation. The unifying technical object is the *common crossing graph*, which encodes pairwise ambiguity, family-level generation obstructions, and corruption defects in a single coverage-and-incidence language.
Contrastive Thinking Decoding: Steering Answer Generation in Reasoning Models
Taehyeon Kim ⋅ Youngsoo Jang ⋅ Hyunsoo Lee ⋅ Yu Jin Kim ⋅ Moontae Lee
Large reasoning models separate inference into an explicit thinking phase followed by a final answer phase, yet how the answer phase uses the reasoning trace remains unclear. We present a systematic study of this trace-to-answer transition across multiple LRMs and find that (i) final answers diverge from correct reasoning traces at rates between 6\% and 48\%, even when traces terminate cleanly at full budget, (ii) naively copying the trace answer underperforms standard decoding in some settings, revealing that the answer phase has independent corrective capacity, and (iii) the same phenomenon extends beyond structured tasks to settings where reasoning traces and final answers can diverge under social-pressure framing (sycophancy). These findings motivate Contrastive Thinking Decoding (CTD), a training-free, single-model decoding method that contrasts answer-phase logits under the primary reasoning trace against those from a deliberately degraded noisy trace. CTD selectively amplifies trace-consistent tokens while preserving the answer phase's corrective capacity, concentrating CTD-internal joint mass on the both-correct cell (T$\checkmark$A$\checkmark$) without inducing trace-correct-to-answer-wrong drift relative to the base distribution. Across math reasoning, science, code, social-pressure resistance (sycophancy), and a broad knowledge benchmark (MMLU), \ctd reduces trace-answer disagreement and improves the accuracy--compute trade-off without parameter updates or auxiliary models at decoding time.
Generative agents have proven to be powerful assistants in a wide variety of contexts. Given this success, users are now deploying agents with minimal restrictions in open ended, multi-agent environments. Current methods for measuring the dynamics of open-ended multi-agent systems are limited to qualitative inspection. In this paper, we extend the process-theoretic notion of adaptive control charts to multi-agent systems to enable automated monitoring. Using simulation, we demonstrate that adaptive control charts are necessary for monitoring multi-agent systems that can learn from their environment. We further demonstrate, both empirically and theoretically, that adaptive control charts are susceptible to adversarial agents that defect sufficiently slowly. These results illustrate a fundamental tradeoff in multi-agent system control: either agents in a system cannot learn \textit{or} the system is susceptible to adversaries.
Controllable Generative Sandbox for Causal Inference
Qi Zhang ⋅ Harsh Parikh ⋅ Ashley I Naimi ⋅ Razieh Nabi ⋅ Christopher Kim ⋅ Timothy L Lash
Method validation and study design in causal inference rely on synthetic data with known counterfactuals. Existing simulators trade off distributional realism, the ability to capture mixed-type and multimodal tabular data, against causal controllability, including explicit control over overlap, unmeasured confounding, and treatment effect heterogeneity. We introduce CausalMix, a variational generative framework that closes this gap by coupling a mixture of Gaussian latent priors with data-type-specific decoders for continuous, binary, and categorical variables. The model incorporates explicit causal controls: an overlap regularizer shaping propensity-score distributions, alongside direct parameterizations of confounding strength and effect heterogeneity. This unified objective preserves fidelity to the observed data while enabling factorial manipulation of causal mechanisms, allowing overlap, confounding strength, and treatment effect heterogeneity to be varied independently at design time. Across evaluation settings, CausalMix achieves strong distributional fidelity on mixed-type tables while providing stable, fine-grained causal control. We demonstrate practical utility in a comparative safety study of metastatic castration-resistant prostate cancer treatments, using CausalMix to compare estimators under calibrated data-generating processes, tune hyperparameters, and conduct simulation-based power analyses under targeted treatment effect heterogeneity scenarios.
Controllable Multi-label Video Safety Detection via Adaptive Tversky Policy Optimization
Guangyu Yang ⋅ Jingbiao Mei ⋅ Mingsheng Sun ⋅ Jinghong Chen ⋅ Yingtong Bu ⋅ Pengda Qin ⋅ Da Chen ⋅ Bill Byrne
The rapid growth of video-based social media has increased users’ exposure to harmful content, creating a need for reliable automated video safety detection. Although recent Vision-Language Models (VLMs) show strong video understanding capabilities, existing harmful video detection systems face two key limitations: they typically reduce safety detection to binary classification, overlooking the inherently multi-label nature of unsafe videos, and they rely on static training objectives that do not support controllable precision-recall trade-offs, though the desired operating point may vary across moderation pipelines and unsafe categories. To address these gaps, we propose Adaptive Tversky Policy Optimization (ATPO), a reinforcement learning framework for Multi-label Video Safety Detection (Multi-VSD). ATPO introduces the Adaptive Tversky Reward (ATR), which dynamically adjusts false-positive and false-negative penalties during training to enable controllable precision–recall trade-offs. Experiments on SafeWatch-Bench and XD-Violence show that ATPO substantially improves multi-label performance, increasing the Jaccard Index from 40.66 to 75.44 on SafeWatch-Bench-Real. Moreover, ATR enables reliable steering of the precision–recall operating point, supporting deployment scenarios with heterogeneous policy requirements.
Convergence Analysis of Newton's Method for Neural Networks in the Overparameterized Limit
Konstantin Riedl ⋅ Justin Sirignano ⋅ Konstantinos Spiliopoulos
A convergence analysis is developed for the regularized Newton method for training neural networks (NNs) in the overparameterized limit. We prove that as the number of hidden units tends to infinity, the NN training dynamics converge in probability to the solution of a deterministic limit equation involving a "Newton neural tangent kernel" (NNTK). Explicit rates characterizing this convergence are provided and, in the infinite-width limit, we prove that the NN converges exponentially fast to the target data (i.e., a global minimizer with zero loss). Crucially, we show that this convergence is uniform across the frequency spectrum, addressing the spectral bias inherent in gradient descent. The eigenvalues of the neural tangent kernel for gradient descent accumulate at zero, leading to slow convergence for target data with high-frequency components. In contrast, the NNTK has uniformly lower bounded eigenvalues if the regularization parameter is selected appropriately, allowing Newton's method to converge more quickly for data with high-frequency components. Mathematical challenges that need to be addressed in our analysis include the implicit parameter update of the Newton method with a potentially indefinite Hessian matrix and the fact that the dimension of this linear system of equations tends to infinity as the NN width grows. This substantially complicates deriving the training dynamics in the overparameterized limit as well as proving the convergence of the finite-width dynamics thereto. The analysis identifies a scaling formula for selecting the regularization parameter, which we show can vanish at a suitable rate as the number of hidden units becomes larger. In addition, we prove that, for sufficiently large numbers of hidden units, the regularized Hessian remains positive definite during training and the Newton updates for individual NN parameters converge to zero, demonstrating that the model behaves as a linearization around the initialization.
Cooperative Multi-Agent Reinforcement Learning via Epigraph-Form Guided Exploration
Sungil Son ⋅ Hoseong Jung ⋅ Dahyun Oh ⋅ H. Jin Kim
Discovering cooperative behavior in multi-agent reinforcement learning (MARL) is challenging due to the combinatorial complexity of joint state-action spaces, which hinders the emergence of coordinated behaviors from trial-and-error alone. Intrinsic rewards are often used to aid discovery, but naively combining them with team objectives can distort the learning signal, compromising task performance. In this paper, we propose EFXPLORER, a constrained exploration framework that maximizes exploration objectives subject to a constraint that preserves established task performance. We solve it via an epigraph reformulation that introduces adaptive exploration budgets. This approach separates intrinsic rewards from task objectives and regulates exploration through task feasibility. To further encourage diverse and temporally extended exploration, we incorporate a successor distance-based intrinsic reward that captures long-horizon dependencies. Empirically, our method outperforms strong baselines and induces novel cooperative strategies across SMAX, VMAS, and MPE benchmark suites.
CORA: Per-Slice Coherent Orthogonal Rotation for SVD-based Low-Rank Adaptation
Pengcheng Wang ⋅ Ziran Liu ⋅ Wei Wang ⋅ WEI JIANG
Parameter-Efficient Fine-Tuning (PEFT) commonly adapts pretrained weights through low-rank updates, and recent methods further exploit the singular value decomposition (SVD) of the base weight for initialization or subspace selection. However, these methods do not explicitly preserve the coupled geometry between the pretrained left and right singular bases. Motivated by recent minimum-perturbation theory, which shows that stable finetuning admits a coherent SVD rotation in which a single orthogonal $Q$ acts on both the left singular basis $U_0$ and the right singular basis $V_0$, we prove a per-slice analogue: each row slice of $W_0$ can be adapted by a shared orthogonal rotation $Q_i$ on its left basis $U_i$ and right basis $V_i$ together with a diagonal spectrum shift. We instantiate this form as **CORA** (*Coherent Orthogonal Rotation Adaptation*), which applies per-slice orthogonal rotations and a per-layer diagonal scale to the rank-$r$ SVD truncation of $W_0$. CORA uses $\tfrac{1}{2}m(r{-}1)$ trainable parameters per linear layer, about $4{\times}$ fewer than LoRA at the same rank. CORA outperforms LoRA, DoRA, PiSSA, and MiLoRA on commonsense reasoning and code generation while using $\sim 8{\times}$ fewer parameters.
Correction Space Steering for Hallucination Mitigation in Large Vision-Language Models
Songbo Yang ⋅ Shuliang Liu ⋅ Sihang Jia ⋅ Xuming Hu
Large Vision-Language Models (LVLMs) are prone to hallucinations, producing responses that are inconsistent with visual inputs. While inference-time activation steering provides a lightweight solution without retraining, existing methods are either limited by coarse global steering vectors or rely on unconstrained full-dimensional vector prediction, which can produce unreliable vector for unseen hallucination patterns. In this paper, we propose \textbf{Correction Space Steering} (CSS), an robust activation steering framework that predicts correction coordinates in a learned low-dimensional correction space. We first conduct an empirical analysis of hallucination-related activation differences and find that they are predictive of hallucination types, form type-aware geometric structures, and exhibit low effective ranks in most layers. We further provide a theoretical analysis showing that these activation differences can be approximated by a low-dimensional subspace. Based on this insight, CSS learns a low-dimensional correction space from activation differences and trains an online predictor to infer input-specific correction coordinates from inference-time hidden states. The resulting correction vectors are then reconstructed within the correction space and applied to steering activations for hallucination mitigation. Experiments on standard hallucination benchmarks show that CSS effectively reduces hallucinations without retraining LVLMs and outperforms existing inference-time mitigation methods.
CORTEG: Foundation Models Enable Cross-Modality Representation Transfer from Scalp to Intracranial Brain Recordings
Liuyin Yang ⋅ Qiang Sun ⋅ Bob Van Dyck ⋅ Eva C Merino ⋅ Marc M Van Hulle
Intracranial electrocorticography (ECoG) offers high–signal-to-noise access to cortical activity for brain–computer interfaces, yet limited per-patient data has led most prior work to rely on small, subject-specific decoders that neglect information shared across patients. We investigate whether large pretrained scalp-EEG foundation models (EEG FMs) can be adapted to ECoG, enabling cross-patient learning and competitive decoding performance while calibrating to a held-out patient in $10$--$30$ minutes on a single GPU. We introduce CORTEG, a cross-modality transfer framework that combines a pretrained EEG FM backbone, an electrode-aware KNNSoftFourier spatial adapter, a dual-stream tokenizer for low-frequency and high-gamma activity, and a leave-one-subject-out fine-tuning strategy. We evaluate CORTEG on two challenging regression tasks: public finger trajectory regression (n=9) and private audio envelope regression (n=16). CORTEG matches or exceeds the strongest task-specific baselines on both tasks: it reaches the highest mean correlation among compared methods on the public finger benchmark (gain not statistically significant on $n{=}9$ subjects), with larger and statistically significant gains on the audio task and in low-data per-patient calibration. Feature analyses align with neurophysiology, and latent manifolds capture low-dimensional finger-movement structure. CORTEG provides systematic evidence that scalp-EEG pretraining can be repurposed for ECoG decoding, enabling data-efficient intracranial BCIs that can adapt to new patients.
Cortically-Resolved Recurrent Architecture for fMRI
Da-Woon Heo ⋅ Kwanseok Oh ⋅ Adrián Ayuso-Muñoz ⋅ Heung-Il Suk
Decoding cognitive functions and brain diseases from functional magnetic resonance imaging (fMRI) is a foundational problem in neuroscience advanced by recent deep learning across task and resting-state (rs) settings. Existing region-level decoders capture spatial structure through region-of-interest (ROI)-aware graph or Transformer architectures, but apply a uniform temporal operator across all regions. Consequently, they do not explicitly model how cortical regions integrate information over heterogeneous timescales. We propose CoRTeX, a region-resolved recurrent architecture for fMRI in which each cortical region carries its own learnable parameters and evolves on its own learned timescale. At its core, per-region differentiable temporal integration windows allow each ROI to learn how far back in time to integrate information through a soft mask over a private buffer. Cross-region information is restricted to a single identity-aware attention channel, additionally biased by a co-activation-driven memory whose learning and forgetting rates are themselves region-specific. Experiments on stimulus classification using task-fMRI (NSD) and on two disease diagnosis tasks using rs-fMRI (ABIDE, COBRE) demonstrate that CoRTeX achieves competitive performance with strong baselines while enabling region-level analyses that are not well-defined for models with exchangeable units: the learned per-region timescales exhibit interpretable spatial organization, and class-specific regional attribution aligns with category-selective cortex. Code is available at: https://anonymous.4open.science/r/CoRTeX-F.
CoTs as Probabilistic Programs: A Programmatic View of Thinking Step-by-Step in Language Models
Kyle Richardson ⋅ Yu Feng ⋅ Poorva Garg ⋅ Junyan Cheng ⋅ Dan Roth ⋅ Guy Van den Broeck
Chain-of-thought (CoT) prompting and related strategies that elicit intermediate reasoning traces have fundamentally changed how language models are used, evaluated, and developed. CoT now touches many areas of contemporary research beyond prompting, from advanced test-time inference and model tuning to interpretability, yet much of this work treats CoT traces in task-specific ways, without a shared formal account of what they are or how they compose and support inference. We argue that probabilistic programs — ordinary programs extended with stochastic choices and probabilistic conditioning — provide a natural framework for modeling CoT reasoning. To demonstrate this, we derive a small discrete probabilistic programming language for CoT in which reasoning steps are stochastic choices, branches encode dependencies, trace likelihoods score executions, and answers are return values. Even with a minimal reachability-based semantics, this language enables a first-principles analysis of trace likelihood, exposing why likelihood alone can be a misleading model of reasoning. Motivated by this limitation, we add differentiable operators for composing scores and feedback over CoT traces. The resulting framework links several uses of CoT — scoring, evidence aggregation, feedback propagation, model tuning, and step-level sensitivity analysis — within a single probabilistic program. Through case studies, we show how this approach yields new analysis techniques and test-time inference strategies, offering a promising direction for programming-language-based approaches to language-model reasoning.
Counterfactual Online Conformal Prediction Under Adaptive Logging
Xinyu Qiao ⋅ yichen lin ⋅ Kaihong Ji ⋅ xue wang ⋅ Tao Yao
Online conformal prediction can fail when predictions shape actions and actions determine which outcomes enter calibration. Standard adaptive methods may retain marginal coverage while systematically miscovering the counterfactual outcomes of rarely selected actions. This paper formalizes the failure through counterfactual coverage and introduces Propensity-Weighted Online Conformal Prediction, an inverse-propensity-weighted recursion that debiases calibration. A doubly robust variant further reduces nuisance bias to the product of outcome-model and propensity errors. Under positivity, the resulting coverage rate matches an information-theoretic lower bound up to logarithmic factors. Experiments on synthetic decision tasks, open bandit data, and financial rebalancing show that PW-OCP and DR-OCP improve counterfactual coverage and downstream regret without sacrificing prediction-set sharpness.
CoupledFlow: One-Step Neural Operators for Coupled Multi-Physics PDEs
Trong Khiem Tran ⋅ Long M Bui ⋅ Phi Le Nguyen ⋅ Mrinal K Sen ⋅ Sanjay Srinivasan ⋅ Nghia Hoang ⋅ Jana Doppa
We investigate the problem of learning neural operators for coupled multi-physics PDEs, where multiple physical processes co-evolve and influence one another via bidirectional coupling. For example, in subsurface CO$_2$ storage, fluid flow, heat transfer, and geo-mechanical deformation evolve in a tightly coupled manner. Prior work, including neural operators and generative PDE solvers, mostly focuses on single-physics settings or relies on iterative inference, limiting their effectiveness in coupled PDEs. This paper introduces coupled-flow, a one-step generative neural operator that revisits coupled operator learning as solving a distributional transport problem. It parameterizes process-specific average velocity fields conditioned on the full coupled state, enabling cross-process interactions to emerge naturally from the transport dynamics. We further establish a Wasserstein-2 generalization bound that ties the prediction error of coupled-flow to its training losses through the regularity of the induced flow. Experiments on a wide variety of multi-physics benchmarks demonstrate that coupled-flow achieves significant performance improvements over coupled neural-operator and generative PDE baselines.
Guidance methods are essential tools for steering the generation process in diffusion models, improving the generalizability of stochastic robotic controllers and enabling high-quality controlled image synthesis. Despite strong parallels between flow-matching and diffusion, integrating guidance into flow-matching models remains quite challenging. Recent work has proposed general guidance frameworks for flow-matching, however, they are direct adaptations of diffusion guidance methods and can prove unstable. In this work, we conduct a theoretical analysis of guidance in flow-matching and show that a naive translation of diffusion-based guidance to flow-matching breaks the coupling distribution used during training, resulting in weaker control over sampling accuracy. Building on this insight, we propose Coupled Guidance, a flow-matching-specific guidance mechanism derived from a geometric perspective that leverages the coupling structure inherent to flow-matching training. To validate the effectiveness of our approach, we conduct extensive experimentation on a custom robotic collision avoidance task using Maniskill3, four robotic dexterous manipulation tasks from the Adroit benchmark, and three image generation tasks, namely inpainting, super-resolution, and deblurring, using the ImageNet-128 and CelebA-HQ-256 datasets. Our findings show that our framework improves sample quality and guidance stability across all tasks, without imposing any additional overhead.
CPA: Efficient and Stable FP4 RL Training via Cross-Precision Alignment
Gu Gong ⋅ Yining Wei ⋅ Yuechen Tao ⋅ Tianyuan Wu ⋅ Peijie Dong ⋅ Ruibo FAN ⋅ Wenhu Hu ⋅ Yinghao Yu ⋅ Jiamang Wang ⋅ Wenbo Su ⋅ Guodong Yang ⋅ Liping Zhang ⋅ Wei Wang ⋅ Xiaowen Chu
Reinforcement learning (RL) for large language models (LLMs) is increasingly bottlenecked by rollout cost, making low-precision rollout appealing for acceleration. However, applying low-precision formats such as FP4 to RL remains unstable, even with quantization-aware training (QAT) enhancements. We identify that the root cause is that FP4’s aggressive quantization amplifies small system-level numerical differences (i.e., training-inference stack mismatch) that are harmless at BF16 into large token-level log-probability errors, distorting importance weights and destabilizing optimization. To address this, we propose Cross-Precision Alignment (CPA), a simple regularizer that stabilizes low-precision RL by maintaining a BF16 master policy, executing rollouts in real FP4, and aligning the fake-quantized forward pass with a BF16 reference forward pass on sampled tokens by applying a low-variance KL penalty at each training step. In our evaluation, the resulting BF16 checkpoint matches a full BF16 RL baseline, and the overall pipeline delivers 8–19\% end-to-end training speedup.
CPSea2: Composing Terminal Geometries for Structurally Diverse Cyclic Peptide Binder Design
Ziyi Yang ⋅ Yinjun Jia ⋅ Guiyu Deng ⋅ Jiqing Zheng ⋅ Tianyi Zhang ⋅ Ziting Zhang ⋅ Kai Liu ⋅ Lei Liu ⋅ Yanyan Lan
Cyclic peptides are attractive next-generation therapeutics. In drug discovery, diverse ring-closing chemistries are applied to tune synthetic accessibility and conformational constraints. Current machine learning methods represent cyclization through modified positional encodings, covalent-bond prompts in all-atom models, residue-level distance guidance, or direct terminal closure by molecular mechanics (MM). These approximations provide limited control over the raw geometry, especially for non-natural linkers, and often yield unreliable structures. Here, we introduce CPSea2, a composable framework that abstracts cyclization as terminal-amide geometry constraints. We construct CPSea2-Cap (CPcap), a conformer library covering eight cyclization types, and CPSea2-Base (CPbase), a large-scale dataset of nearly 900 million mined linear peptide--protein complexes. Matching terminal-amides between CPcap and CPbase directly derives cyclic peptide--protein complexes with diverse terminal topologies. Built on CPbase, PepAny generates target-conditioned peptide binders with controllable terminal-amide geometry, expanding the accessible space beyond mined structures. Multi-target evaluations show improved raw-output plausibility over bond-prompted and distance-guided baselines. CPSea2 provides a novel framework and data foundation for flexible and diverse cyclic peptide binder design. Datasets and codes are available at https://anonymous.4open.science/r/CPSea2-6CF5/.
Credit Assignment with Resets in Language Model Reasoning
Ankur Samanta ⋅ Akshayaa Magesh ⋅ Ayush Jain ⋅ Youliang Yu ⋅ Daniel Jiang ⋅ Kavosh Asadi ⋅ Kaveh Hassani ⋅ Paul Sajda ⋅ Jalaj Bhandari ⋅ Yonathan Efroni
Reinforcement learning with verifiable rewards (RLVR) post-trains language models on multi-step reasoning by assigning a single outcome reward uniformly across all tokens in a trajectory, regardless of which steps contributed to success or failure. Improving credit assignment can address this limitation by enabling targeted refinement of faulty reasoning steps, rather than updating entire trajectories uniformly. Resets are one such simple mechanism, enabling more precise credit assignment by returning to an intermediate state and resampling its continuation, so that outcome differences can be attributed to decisions made at that point. We propose two such methods: Random-Reset Policy Optimization (RRPO), where reset states are drawn randomly from reasoning steps, and Self-Reset Policy Optimization (SRPO), where the model self-localizes the erroneous step in an incorrect trajectory and resets there. We analyze these methods within a Conservative Policy Iteration (CPI) framework. We extend it with a credit-assignment oracle that targets improvable states, defined as those whose advantage exceeds a threshold, and show that it yields provable improvements over random resets. Across models and reasoning benchmarks, SRPO consistently outperforms standard GRPO and RRPO by sampling multiple suffix continuations at a self-localized reset and learning from their rewards, using only the model itself with no external supervision.
CrossWeave: Emergent Cross-Modal Scene and Instance Retrieval from Sparse 2D-3D Alignment
Aadith Warrier ⋅ Gnana Prakash Punnavajhala ⋅ Siddharth Tourani ⋅ Muhammad Haris Khan ⋅ Avinash Sharma ⋅ Madhava Krishna
Cross-modal scene and instance retrieval aims to retrieve a corresponding 3D environment or object given query 2D images, serving as a foundational mechanism for spatial reasoning in robotics and AR/VR. Existing methods achieve this by compressing the entire 3D map into a single global feature vector, which discards the fine-grained local geometrical structure in indoor scenes. In this work, we show that scene- and instance-level retrieval capabilities emerge from supervision on sparse, local 2D-3D correspondences, without any explicit global labels. Specifically, we propose $\textbf{CrossWeave}$, a novel framework for cross-modal retrieval that learns sparse patch-point alignment, which alone outperforms global encoding methods. Additionally, an optimal transport-based aggregator is applied over the aligned features to attain further gain by clustering structurally redundant regions, freeing descriptor capacity for semantically distinct local features. CrossWeave achieves state-of-the-art performance on both scene- and instance-level retrieval on ScanNet and 3RScan with a $\textit{single model}$, whereas existing methods require separate models for each task. Our results suggest that whole-scene encoding is not necessary for cross-modal retrieval, and that sparse local correspondence supervision is a more effective and flexible alternative.
CryoAtlas: A Large Curated Dataset and Unified Benchmark for Cryo-EM Atomic Model Building
Mingrui Li ⋅ Minzhang Li ⋅ Weichen Qin ⋅ Yufan Xie ⋅ Zhangzhi Xiong ⋅ Sixian Shen ⋅ Jiakai Zhang ⋅ Yuan Pei ⋅ Jingyi Yu
Deep learning has accelerated automated atomic model building from cryo-EM maps, yet progress remains limited by the lack of large, high-quality datasets and standardized evaluation benchmarks. We address this gap by introducing CryoAtlas, a rigorously curated dataset and unified benchmark for density-guided atomic model building. By integrating and filtering EMDB and PDB entries (up to March 2026), CryoAtlas provides 17,809 map--model--sequence entries, including a curated 16,657-entry training split. From this resource, we construct a 618-target benchmark with a shared evaluation protocol that jointly assesses structural accuracy, model-to-map fit, and stereochemical validity. Using this framework, we benchmark state-of-the-art learning-based methods against Phenix and AlphaFold~3, comprehensively analyzing the impact of map resolution, sequence length, protein category and sequence identity on modeling accuracy. As a case study, retraining the ModelAngelo CNN component on CryoAtlas improves performance on the benchmark. CryoAtlas provides an open, reproducible foundation for future cryo-EM atomic model building research.
CSCN: The Crossed Subtree Convolutional Network
Zongjun Han ⋅ Lu Bai ⋅ Lixin Cui ⋅ Ming Li ⋅ Hangyuan Du
Message Passing Neural Networks (MPNNs) commonly suffer from over-squashing and over-smoothing issues, which limit their performance on graph-level tasks. To this end, we propose a novel Crossed Subtree Convolutional Network (CSCN) for graph classification. Specifically, CSCN first introduces an ordered subtree alignment strategy to decompose the original graph into an ordered sequence of structured subtrees. This sequence preserves rich local connectivity patterns and maintains the permutation invariance of the graph representation. Based on this, we design a crossed subtree convolution. It utilizes a sliding window to enable information interactions across different subtrees while simultaneously aggregating node features within each subtree. This design captures local information effectively while also modeling long-range dependencies. In this way, CSCN avoids deep message passing over the global graph topology and thus alleviates both over-squashing and over-smoothing. Experimental results on multiple graph classification benchmarks show that CSCN achieves competitive performance.
CSLA: Sparse-Linear Attention with Learnable Routing and Quantization-aware Training
Jintao Zhang ⋅ Haoxu Wang ⋅ Kai Jiang ⋅ Kaiwen Zheng ⋅ Youhe Jiang ⋅ Ion Stoica ⋅ Jianfei Chen ⋅ Jun Zhu ⋅ Joseph Gonzalez
Sparse-Linear Attention (SLA) combines sparse and linear attention to accelerate diffusion models and has shown strong performance in video generation. However, (i) SLA relies on a heuristic split that assigns computations to the sparse or linear branch based on attention-weight magnitude, which can be suboptimal. Additionally, (ii) after formally analyzing the attention error in SLA, we identify a mismatch between SLA and a direct decomposition into sparse and linear attention. We propose CSLA, which introduces (I) a learnable router that dynamically selects whether each attention computation should use sparse or linear attention, (II) a more faithful and direct sparse-linear attention formulation that uses a learnable ratio to combine the sparse and linear attention branches, and (III) a sparse + low-bit attention design, where low-bit attention is introduced via quantization-aware fine-tuning to reduce quantization error. Experiments show that on video diffusion models, CSLA can achieve 97\% attention sparsity and deliver an 18.6$\times$ attention speedup while preserving generation quality.
Coding agents built from large language models increasingly inspect and modify stateful services. Before such an agent chooses or explains a patch, it needs a basic engineering skill: predicting how a local code or state change will affect future service behavior. CTE-Bench isolates this simulator-fidelity question from agent policy success. Each item gives the model service source, a factual interaction prefix, a post-prefix source patch or hidden-state overwrite, and a fixed suffix query schedule; an executable oracle supplies the counterfactual response trace. CTE-Bench-Core-v1 contains 255 counterfactual items over six deterministic Python services and 10,200 scored suffix decisions per model configuration. The main score is effect-step value-match (VM): exact response equality restricted to the 2,476 suffix positions where the intervention changes the oracle response. DeepSeek V4-Lite, Kimi K2.5, and Claude Sonnet 4.6 reach 60.7%, 61.5%, and 58.4% effect-step VM under teacher-forced one-step prompting, where each prediction sees earlier oracle suffix responses; hiding those responses while preserving the factual prefix reduces them to 11.4%, 28.9%, and 25.0%. As an additional probe, we run self-conditioned free-rollout on DeepSeek and Kimi; effect-step VM is 27.1% and 33.2%, and exact whole-trace match is at most 1.2%. These probes show that CTE-Bench helps identify dependence on oracle suffix feedback and compounding error in self-conditioned service simulation. We release Core-v1, the executable oracle, metadata, and evaluation scripts, with an evaluation card tying each supported claim to its memory protocol and metric.
CURE: Coupled User-Grouped Reinforcement Learning for Cross-Domain Recommendation with Non-Overlapping Users
Hyeongjun Yun ⋅ Taesan Kim ⋅ Dongjoon Hong ⋅ Junui Hong ⋅ Jihoon Oh ⋅ Kijung Park ⋅ MinCheol Cho ⋅ Kihyuk Song ⋅ Jaegul Choo ⋅ Chung Park
Cross-Domain Recommendation is essential for enabling cross-sell in multi-service platforms; however, limited user overlap and strict cross-service data-sharing constraints often result in little to no ground-truth supervision for cross-domain learning. To bridge this gap, we propose an LLM-based framework for cross-domain recommendation in the non-overlapping user scenario. First, we propose an agentic pipeline that constructs cross-domain pseudo supervision, which is unavailable in the non-overlapping user scenario, by leveraging a user’s source-domain history to generate a target-domain item recommendation. Using this dataset, we further propose \underline{C}oupled \underline{U}ser-G\underline{R}ouped R\underline{E}inforcement Learning (CURE), which induces coupling between in- and cross-domain gradient updates via user-level joint normalization, thereby calibrating synthetic cross-domain updates against reliable in-domain signals. We provide a theoretical analysis showing that user-level coupling stabilizes policy updates under imperfect pseudo supervision and improves cross-domain generalization. Experiments on Amazon Reviews and MovieLens show consistent gains over baselines in cross-domain recommendation, while also improving in-domain performance. An online A/B test demonstrates a statistically significant 56\% lift in click-through rate over existing ML methods.
CycleSpectra: Cyclic Motion Spectra for Phase-Queryable 4D Cardiac Reconstruction
Xueming Fu ⋅ Xin Li ⋅ Lixia Han ⋅ Ao Shen ⋅ Guangming Lu ⋅ Song Luo ⋅ S. Kevin Zhou
Reconstructing patient-specific cyclic cardiac motion at arbitrary phases is central to dynamic cardiac assessment, underpinning measurements such as ejection fraction, stroke volume, and regional myocardial deformation. Existing methods---pair-wise registration, time-conditioned decoders, and coordinate-based implicit fields---share a common implicit assumption: the cardiac cycle is treated as a sequence of discrete phases, with temporal structure stitched together through learned mappings supervised at the observed frames. Yet the cycle is, by physical fact, a closed periodic process. We propose \textbf{CycleSpectra}, a spectral motion representation that takes the cyclic structure as a starting point of the representation rather than as a property to be learned. Given a pair of sparsely sampled phase observations, the model learns the spectral structure of the underlying cyclic cardiac motion in a low-dimensional latent space, and recovers the displacement at any phase through a fixed analytic synthesis layer---time never enters the network as a learnable signal, only as a coordinate the readout takes by construction. To support evaluation, we collect a thin-slice phase-resolved 4D cardiac CT cohort and benchmark under both canonical and non-canonical input protocols, reporting whole-cycle motion-fidelity metrics alongside standard image and anatomical metrics. Under arbitrary two-frame input, CycleSpectra achieves the lowest ejection-fraction error and highest ventricular volume--time curve correlation among recent baselines, while remaining competitive on image fidelity. Beyond two-frame inference, the proposed framework supports test-time multi-frame latent fusion for higher accuracy, validated on both the in-house cohort and a cross-modality 4D cardiac MR cohort. Code will be available.
Cyclic Denoising Reveals Ultrastable Memories in Diffusion Models
Rishabh Sharma ⋅ Stefano Martiniani
We introduce cyclic denoising—repeated forward and reverse diffusion at controlled noise amplitudes—as an extraction attack for image diffusion models. Inspired by random organization in disordered solids, where cyclic mechanical perturbations anneal the system into increasingly stable configurations, cyclic denoising exposes regions of the learned distribution that remain largely inaccessible to standard sampling. We find that this dynamics drives samples toward attractors with a broad stability spectrum, with the deepest attractors exhibiting ultrastability: they can be regenerated from near-total corruption and sustained through thousands of noising-denoising cycles. Many of these deep attractors correspond to memorized training images, including stock photographs, brand watermarks, and web-crawl artifacts. Our extraction attack requires sampler-level control, including mid-process noise injection, but no gradients and no weight inspection. Crucially, it requires no prior knowledge of training data, captions, or prompts. In contrast, prior generate-and-filter attacks commonly rely on prompted generation using known or suspected training captions, followed by large-scale sampling and post-hoc similarity or membership-inference filtering. While cyclic denoising can also be applied with prompts, our main protocol is fully unconditioned. We demonstrate the phenomenon in Stable Diffusion v1.4, a large latent diffusion model, and in a smaller pixel-space DDPM, showing consistent behavior across latent- and pixel-space diffusion models. Across noise amplitudes, we observe a dynamical transition from trivial fixed points to structured memorized images, hierarchical partial absorption in which coarse scene layout freezes while fine details remain diffusive, basin hopping between memorized states with long residence times, prompt-stabilized memorized templates, and cross-initial-condition universality of the recovered attractor set. Together, these results establish cyclic denoising as both a physics-inspired probe of generative landscapes and a practical tool for memorization auditing, with implications for privacy, copyright compliance, and model fingerprinting.
DataFlex: A Unified Benchmark and Evaluation Platform for Data-Centric Training of Large Language Models
Hao Liang ⋅ Zhengyang Zhao ⋅ Mingrui Chen ⋅ Meiyi Qiang ⋅ Lu Ma ⋅ Rongyi YU ⋅ Hengyi Feng ⋅ Shixuan Sun ⋅ Zimo Meng ⋅ Xiaochen Ma ⋅ Xuanlin Yang ⋅ Qifeng Cai ⋅ Ruichuan An ⋅ Bohan Zeng ⋅ Zhen H Wong ⋅ Chengyu Shen ⋅ Runming He ⋅ ZhaoYang Han ⋅ Yaowei Zheng ⋅ Fangcheng Fu ⋅ Conghui He ⋅ Bin CUI ⋅ Zhiyu li ⋅ Weinan E ⋅ Wentao Zhang
Data-centric training—selecting, mixing, and reweighting data during LLM optimization—has produced a rapidly growing body of algorithms, yet every new method ships as a bespoke codebase with its own interface, training protocol, and evaluation setup, making head-to-head comparison and reproducible evaluation nearly impossible. We present DataFlex, the first unified benchmark and evaluation platform for data-centric training of LLMs. DataFlex contributes: (i) a standardized evaluation protocol that runs selection, mixture, and reweighting methods under identical seeds, data splits, hardware, and metrics; (ii) a reference implementation of ten representative algorithms (six data selection, three data mixture, one data reweighting) re-engineered on top of LLaMA-Factory to share model, dataloader, and distributed-training infrastructure (including DeepSpeed ZeRO-3 with full-rank gradient acquisition); and (iii) an empirical benchmark whose head-to-head results yield reproducible findings that no single existing codebase could produce: dynamic selection methods consistently outperform static full-data training on MMLU across Mistral-7B and Llama-3.2-3B, and the gain is larger on smaller models (up to 13.3 points); DoReMi and ODM yield complementary wins on high- vs. low-resource domains when pretraining Qwen2.5-1.5B on SlimPajama at 6B and 30B tokens; and DataFlex's implementation is consistently faster than the original codebases of LESS, TSDS, and MoE-SFT. By unifying fragmented methods under a single evaluation substrate, DataFlex provides the community with a reproducible, extensible basis for future data-centric research.
DC-ViT: Modulating Spatial and Channel Interactions for Multi-Channel Images
Umar Marikkar ⋅ Syed S Husain ⋅ Muhammad Awais ⋅ Sara Atito
Training and evaluation in multi-channel imaging (MCI) remains challenging due to heterogeneous channel configurations arising from varying staining protocols, sensor types, and acquisition settings. This heterogeneity limits the applicability of fixed-channel encoders commonly used in general computer vision. Recent Multi-Channel Vision Transformers (MC-ViTs) address this by enabling flexible channel inputs, typically by jointly encoding patch tokens from all channels within a unified attention space. However, unrestricted token interactions across channels can lead to feature dilution, reducing the ability to preserve channel-specific semantics that are critical in MCI data. To address this, we propose Decoupled Vision Transformer (DC-ViT), which explicitly regulates information sharing using Decoupled Self-Attention (DSA), which decomposes token updates into two complementary pathways: spatial updates that model intra-channel structure, and channel-wise updates that adaptively integrate cross-channel information. This decoupling mitigates informational collapse while allowing selective inter-channel interaction. To further exploit these enhanced channel-specific representations, we introduce Decoupled Aggregation (DAG), which allows the model to learn task-specific channel importances. Extensive experiments across three MCI benchmarks demonstrate consistent improvements over existing MC-ViT approaches.
D-DOIT: Training-free Adaptation of Discrete Diffusion via Doob's h-Transform
Jieke Wu ⋅ Qijie Zhu ⋅ Weimin Wu ⋅ Zeqi Ye ⋅ Minshuo Chen ⋅ Han Liu
We propose D-DOIT (Discrete Doob-Oriented Inference-time Transformation), a training-free and efficient adaptation method for discrete diffusion models with generic rewards. D-DOIT formulates adaptation as sampling from a reward-tilted target distribution and realizes this transport through Doob's h-transform of the discrete diffusion reverse kernel, using only reward values rather than reward gradients. Unlike continuous diffusion, masked discrete diffusion samples categorical token-reveal transitions rather than continuous state updates. D-DOIT derives the corresponding discrete Doob's h-transform, which guides sampling by reweighting reverse transition probabilities instead of adding a drift correction. To make this transformation practical, D-DOIT avoids expensive future rollouts. At each guided step, D-DOIT samples candidate next states, uses the model prediction head to complete each candidate into a clean sequence, evaluates each completion with the reward oracle, and resamples the next state with probabilities proportional to the rewards. An optional late-stage best-of-$K$ refinement further improves sample quality by branching trajectories only near the end of denoising, avoiding the $K$-fold cost over the full trajectory. Empirically, on regulatory DNA sequence design benchmarks, D-DOIT consistently outperforms training-free guidance baselines and achieves performance competitive with training-based methods, improving both enhancer activity and cell-type-specificity while preserving sequence naturalness.
Deadline-Constrained Dynamic Workflow Scheduling Can be Cast as a Representation Learning Problem
Ya Shen ⋅ Gang Chen ⋅ Hui Ma ⋅ Mengjie Zhang
Deadline-Constrained Cost-Aware Dynamic Workflow Scheduling (D-CADWS) aims to minimize virtual machine (VM) rental cost while maintaining high workflow deadline satisfaction in dynamic cloud environments. The problem is challenging because the deadline impact of assigning a workflow task to a VM is not directly observable at decision time. It depends on long-horizon effects, including downstream task dependencies, VM queueing, and interactions among dynamically arriving workflows. Existing methods usually fold deadline violations into a single reward or penalty term, which obscures the distinction between action safety and action cost. In this paper, we cast D-CADWS as a representation learning problem. The core idea is to learn a deadline-aware representation that can predict, before execution, whether a candidate action is risky with respect to the current task deadline, and then use it to guide cost-efficient scheduling. Based on this idea, we propose a representation-centered deep reinforcement learning (RCDRL) method. RCDRL constructs task-level deadline supervision, trains a predictive deadline model to learn deadline-aware representations, and learns a VM-cost critic on top of this representation. The final policy follows a safe-first rule that rejects risky actions before minimizing VM cost. We further propose a two-phase training strategy to keep the learned representation aligned with the evolving policy-induced data distribution. Experiments on dynamic workflow scheduling benchmarks show that RCDRL achieves substantially lower VM cost than other state-of-the-art heuristic and DRL baselines while maintaining strong workflow deadline success.
Debiasing Random Oblique Projections for Subsampled OLS and Fast CUR in High Dimensions
Chengmei Niu ⋅ Sachin Garg ⋅ Michal Derezinski ⋅ Zhenyu Liao
Random sampling is a fundamental tool in modern machine learning and numerical linear algebra for reducing the computational cost of large-scale matrix problems. Existing analyses, however, rely primarily on subspace embedding guarantees, which do not precisely characterize the statistical bias of nonlinear random oblique projections induced by sampling, which arises ubiquitously in subsampled least squares and fast low-rank approximation methods. Because (pseudo)inversion is nonlinear, these random oblique projections can be systematically biased even when the underlying sketch is unbiased, thereby introducing hidden bias into downstream least squares and low-rank approximation solutions. In this work, we develop a unified non-asymptotic theory for random oblique projections in high dimensions. We show that standard random sampling schemes generally induce a systematic statistical bias overlooked by classical subspace embedding-style analyses, and we propose a principled debiasing framework to correct it. We illustrate the power of the theory through two canonical applications. For subsampled least squares, we obtain sharp bias--variance characterizations, reveal previously unrecognized statistical suboptimality in widely used sampling schemes, and identify when debiasing yields provable improvements. For fast CUR decomposition, we develop a debiased approach with improved approximation accuracy. Numerical experiments further validate our theoretical findings.
Decoupling Exploration and Policy Optimization: Uncertainty Guided Tree Search for Hard Exploration
Zakaria Mhammedi ⋅ James Cohan
The process of discovery requires active exploration---the act of collecting new and informative data. However, efficient autonomous exploration remains a major unsolved problem. The dominant paradigm addresses this challenge by using Reinforcement Learning (RL) to train agents with intrinsic motivation, maximizing a composite objective of extrinsic and intrinsic rewards. We suggest that this approach incurs unnecessary overhead: while policy optimization is necessary for precise task execution, employing such machinery solely to expand state coverage may be inefficient. In this paper, we propose a new approach that explicitly decouples exploration from policy optimization and bypasses RL entirely during the exploration phase. Our method uses a tree-search strategy inspired by the Go-With-The-Winner algorithm, paired with a measure of uncertainty to systematically drive exploration. By removing the overhead of policy optimization, our approach explores an order of magnitude more efficiently than standard intrinsic motivation baselines on hard exploration benchmarks. Further, we demonstrate that the trajectories discovered during exploration can be distilled into deployable policies using existing supervised backward learning algorithms, achieving state-of-the-art performance by a wide margin on Montezuma's Revenge, Pitfall!, and Venture without relying on domain-specific knowledge. Finally, we demonstrate the generality of our framework in high-dimensional continuous action spaces by solving the MuJoCo Adroit dexterous manipulation and AntMaze tasks in a \emph{sparse-reward} setting, directly from image observations and without expert demonstrations or offline datasets. To the best of our knowledge, this has not been achieved before.
Decoupling is the Key: Scaling Deep Value Networks in Reinforcement Leanring
Yunsheng Xue ⋅ Ziyi Zhang ⋅ zhihao wu ⋅ Youfang Lin
Scaling laws have driven remarkable performance breakthroughs in Computer Vision (CV) and Natural Language Processing (NLP) by increasing model depth. Compared to increasing width, deeper networks provide higher parameter efficiency and more expressive representations. However, increasing network depth in reinforcement learning (RL) still yields limited performance gains. In this paper, through theoretical and empirical analysis, we reveal that this performance degradation is primarily attributed to implicit low-rank bias and overfitting. While these two issues also exist in shallow networks, the detrimental effects caused by component coupling are significantly amplified as the network depth increases, leading to severe performance degradation. For the first issue, we theoretically prove the existence of bootstrapped spectral coupling, which causes high-frequency spectral components to couple with the low-frequency spectral components of the value network, thereby driving representations of deep value networks into severe rank collapse. For the second issue, we reveal that deep value networks severely overfit the noise induced by TD target couplings, which means the construction of TD targets is influenced by other components, such as the boundary of value distribution or current policy. Therefore, our insight is that decoupling is the key to successfully scaling deep value networks in RL. Motivated by this insight, we propose two simple yet effective solutions respectively: policy-independent spectral decoupling and boundary-target decoupling. Moreover, decoupling policy and value learning (as in AWR) is also a necessary target decoupling for deep value network training in offline RL. By integrating these decoupling components, we propose Spectrally Decoupled Distributional Learning (SEED). Empirically, to the best of our knowledge, SEED is the first approach to successfully unlock the potential of depth in both online and offline RL settings, enabling value networks to scale effectively from 3 to 32 layers with consistent performance gains.
Deep Gaussian Processes on Directed Acyclic Graphs
Federico L Perlino ⋅ Oliver Hamelijnck ⋅ Adam Johansen ⋅ Theodoros Damoulas
Many real world processes can be represented as compositions of functions along a Directed Acyclic Graph (DAG). In causal modelling these correspond to the underlying mechanisms, in engineering to the multiple fidelities present, and in gene-regulatory networks to the transcription factors. These functions are partially observed across the DAG, with noisy and heterogeneously sampled measurements, posing significant challenges for reconstruction, uncertainty propagation, and inference. To tackle these we place priors over functions and naturally arrive at Deep Gaussian Processes over DAGs. We theoretically study their prior-collapse behaviour, and the effect of graph topology and intermediate observations on the preservation of information. We obtain almost-sure lower bounds on the asymptotic frequency of depths at which distinction between inputs is preserved, identify broad kernel classes for which these hold, and prove an observation by \citep{dunlop2018} on the role of input connections. We offer a structured variational approximation that retains graph dependencies, satisfies compositional uncertainty, and captures the explaining-away behaviour of colliders. Finally, we empirically validate our theoretical results and our methodology, and model a latent-collider DAG, a protein signalling network, and a multi-fidelity heavy-ion collision task, attaining state-of-the-art performance while recovering the low-fidelity contributions and yielding explainability over the simulator hierarchy.
Deep Research as Rubric
Wangyi Mei ⋅ Zhouhong Gu ⋅ zhenhanbai ⋅ Yin Cai ⋅ Lefan Zhang ⋅ Zhenxin Ding ⋅ Bo chen ⋅ Yan Gao ⋅ YIWU ⋅ Yao Hu
Open-ended reasoning and long-form generation tasks lack reliable automatic verification signals for reward-based policy optimization. Rubrics offer a promising alternative, but existing approaches treat them as given artifacts—either hand-crafted or prompt-generated—and often miss the task-specific, knowledge-intensive dimensions that matter most, distorting the reward signal. Our key observation is that rubric construction is itself a research problem: identifying what makes a response correct or insightful requires discovering and synthesizing external knowledge. We propose Deep Research as Rubric (DR-rubric-8B), a two-stage framework for constructing such rubrics. Stage I elicits domain facts, structural constraints, and failure modes through iterative multi-turn agentic search; Stage II distills this evidence into atomic, independently verifiable constraints for GRPO-based policy optimization. Because the model under training can serve as its own rubric generator, DR-rubric-8B supports bootstrap rubric generation without frontier-model assistance. We evaluate on 6 benchmarks spanning agentic research and expert reasoning. Experiments show that DR-Rubric achieves strong competitive performance with only 1K–3K training instances, where GPT-5-generated rubrics particularly benefit breadth coverage on agentic tasks, while bootstrap rubrics exhibit a specialization-to-rebalancing evolution and achieve the best overall performance at the third iteration. Results demonstrate that reframing rubric construction from static evaluation templates into an evidence-driven research process yields more scalable, fine-grained reward signals for open-ended tasks.
Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents
Xiaofang Yang ⋅ Ziqi Miao ⋅ Dianbo Sui ⋅ Jing Shao ⋅ Lijun Li
Skill-augmented agents load reusable skills as persistent runtime context, improving task performance but also giving malicious skills a durable channel for steering future actions. Such skills may leak secrets, corrupt code, bypass approvals, or stage data for exfiltration only after a concrete user task and workspace state make the unsafe action appear useful. This makes pre-install vetting insufficient and calls for runtime, task-conditioned protection. We propose Defense-as-Skill, a defense paradigm that implements the runtime guard itself as an installable, inspectable, and editable skill. Our guard, SkillSonar, runs alongside untrusted task skills and checks sensitive actions against the user's task boundary, routing each action to an allow, replan, or confirmation decision without modifying the underlying agent runtime. To study this setting, we construct SCOPE-R, a task-conditioned dataset covering 6 risk families and 21 sub-categories, with 206 attack-confirmed malicious instances and 43 benign tasks. We then improve SkillSonar on the SCOPE-R training subset using runtime guard-skill evolution, a Monte-Carlo Tree Search procedure that evolves the on-disk guard skill from feedback on the rollouts. On held-out SCOPE-R test splits across Claude Code and OpenClaw backends, SkillSonar reduces attack success while preserving benign utility, maintaining modest token overhead, and transferring to OOD attack families. These results show that runtime guards can be skill-native, transparent, and evolvable policy artifacts for skill-augmented agents.
Delay-Embedded Representations for Robust Saccade Classification in Noisy Oculographic Signals
Vladimir Antipov ⋅ Artem Badarin ⋅ Alexander Khramov
We introduce ACMA (Delay-Augmented Clustering and Model Approximation), a noise-robust pipeline for saccade detection in oculographic signals that couples delay-embedded clustering with physiologically grounded parametric refinement. Existing detectors face a recurring trade-off: velocity-thresholding methods are simple and interpretable but degrade catastrophically as recording noise grows, adaptive event-based detectors recover part of this gap at moderate noise but lose reliability at high noise, and recent deep models are sensitive to noise distributions unseen during training. ACMA addresses this trade-off in two stages. First, a sliding-window two-cluster decomposition with time-delay embedding identifies candidate saccades while remaining stable across noise regimes. Second, a parametric saccade waveform constrained by priors on amplitude, duration, and the saccadic main sequence validates and characterizes each event. We evaluate ACMA on a saccade-detection benchmark spanning a wide range of noise levels and nine representative baselines, including a leading deep-learning detector (U'n'Eye), adaptive methods (NH, REMoDNaV, Engbert), clustering-based I2MC, and classical velocity-thresholding methods (IVT, IVVT, IDT, IDVT). ACMA delivers consistent F1 gains in moderate-to-high noise regimes, where competing detectors degrade sharply, while additionally returning physically meaningful per-event metrics (amplitude, duration, peak velocity) for downstream neurological assessment, fatigue monitoring, and brain-computer interface applications.
DELTA: Robustly Training Label-Conditional Diffusion Models with Weak Annotations
Dong-Dong Wu ⋅ Jiacheng Cui ⋅ Wei Wang ⋅ Zhiqiang Shen ⋅ Masashi Sugiyama
While label-conditional diffusion models exhibit remarkable generative capabilities recently, their success heavily relies on massive, cleanly labeled datasets. In practice, categorical supervision is rarely perfect: it is often corrupted by noise, clouded by ambiguity, or partially missing. Training directly on such weak annotations severely degrades generation quality, and existing robust methods offer only fragmented solutions that depend on scarce auxiliary priors. To overcome these limitations, we propose DELTA, a unified framework that robustly trains diffusion models across all three weak annotation types without any external priors. By treating the unknown true label as a latent variable, DELTA optimizes a principled variational objective that jointly recovers the clean data distribution and infers the true label posteriors. To make this joint training computationally tractable, we further introduce a median-centered timestep sampling strategy that efficiently concentrates evaluation where the diffusion process is most informative. Extensive experiments demonstrate that DELTA produces high-fidelity, class-consistent samples across diverse weak supervision scenarios, outperforming specialized baselines while demanding strictly less prior knowledge.
Denoise First, Orthogonalize Later: Understanding Momentum in Muon via Spectral Filtering
Xianliang Li ⋅ Zihan Zhang ⋅ Weiyang Liu ⋅ Han Bao
Muon has recently demonstrated strong empirical performance in large language model training, but the theoretical role of momentum in Muon remains unclear. Existing analyses of Muon either remove momentum to study spectral updates in isolation, or retain momentum without explaining why it improves empirical performance. Our work bridges this gap by showing momentum in Muon acts as a spectral filter. Under a structured signal-plus-perturbation gradient model, we prove that momentum suppresses perturbations while preserving the dominant signal, thereby enlarging the spectral gap between them. This enlarged gap stabilizes the singular subspaces of the matrix passed to Muon's orthogonalization step, making the resulting update more reliable. We further show that applying momentum before orthogonalization achieves provably stronger alignment with the signal component of the gradient than either reversing this order or simply removing momentum. Experiments across diverse tasks, including LLM training, support our theoretical analysis. More broadly, our theory offers a starting point for understanding the benefits of momentum in other matrix-based optimizers.
Dependency-Guided Parallel Decoding in Discrete Diffusion Language Models
Liran Ringel ⋅ Ameen A Ali ⋅ Yaniv Romano
Discrete diffusion language models (dLLMs) accelerate text generation by unmasking multiple tokens in parallel. However, parallel decoding introduces a distributional mismatch: it approximates the joint conditional using a fully factorized product of per-token marginals, which degrades output quality when selected tokens are strongly dependent. We propose DEMASK (DEpendency-guided unMASKing), a lightweight dependency predictor that attaches to the final hidden states of a dLLM. In a single forward pass, it estimates pairwise conditional influences between masked positions. Using these predictions, a greedy selection algorithm identifies positions with bounded cumulative dependency for simultaneous unmasking. Under a sub-additivity assumption, we prove this bounds the total variation distance between our parallel sampling and the model's joint. Empirically, DEMASK achieves 1.7--2.2$\times$ speedup on Dream-7B while matching or improving accuracy compared to confidence-based and KL-based baselines. When applied to dParallel, a block diffusion model, DEMASK also produces Pareto-optimal accuracy-step trade-offs across math and code benchmarks.
Depth-Recurrent Attention Mixtures: Giving Latent Reasoning the Attention it Deserves
Jonas Knupp ⋅ Jan Metzen ⋅ Jeremias Bohn ⋅ Georg Groh ⋅ Kristian Kersting
Depth-recurrence promises improved latent reasoning by sharing parameters across depths. However, prior work still relies on partially fixed layer stacks and overlooks the bottleneck of constant hidden-size; moreover, it largely lacks rigorously resource-matched ablations. To address this, we introduce depth-recurrent attention mixtures (Dreamer), a modular architecture with a single recurring layer. Concretely, we mix sequence attention, depth attention, and sparse expert attention. This alleviates the hidden-size bottleneck, decouples scaling dimensions, and achieves effective yet efficient depth-recurrence. Across natural language reasoning benchmarks at up to ~2B parameters, Dreamer requires on average ~3x fewer training tokens for the same accuracy as FLOP-, parameter-, and memory-matched modern Transformers, and outperforms ~2x larger Transformers given same training tokens. Analysis further reveals 2-11x greater expert selection diversity than conventional MoEs, highlighting the flexible knowledge sharing across depths.
Deriving Hyperparameter Scaling Laws via Modern Optimization Theory
Egor Shulgin ⋅ Jörg Franke ⋅ Dimitri von Rütte ⋅ Tianyue Zhang ⋅ Niccolò Ajroldi ⋅ Korbinian Pöppel ⋅ Bernhard Schölkopf ⋅ Aaron Klein ⋅ Peter Richtarik ⋅ Antonio Orvieto
Hyperparameter transfer has become an important component of modern large-scale training recipes. Existing methods, such as muP, primarily focus on transfer between model sizes, with transfer across batch sizes and training horizons often relying on empirical scaling rules informed by insights from timescale preservation, quadratic proxies, and continuous-time approximations. We study hyperparameter scaling laws for modern first-order optimizers through the lens of recent convergence bounds for methods based on the Linear Minimization Oracle (LMO), a framework that includes normalized SGD, signSGD (approximating Adam), and Muon. Treating bounds in recent literature as a proxy and minimizing them across different tuning regimes yields closed-form power-law schedules for learning rate, momentum, and batch size as functions of the iteration or token budget. Our analysis, holding model size fixed, recovers most insights and observations from the literature under a unified and principled perspective, with clear directions open for future research. Our results draw particular attention to the interaction between momentum and batch-size scaling, suggesting that optimal performance may be achieved with several scaling strategies.
Designing Cell-Type-Specific Regulatory DNA with Guided Discrete Diffusion
Animesh Awasthi ⋅ Martin Stoll ⋅ Raphael Bednarsky ⋅ Moritz Schaefer ⋅ Christoph Bock
Designing regulatory DNA with cell-type-specific activity is broadly relevant for cell engineering and gene therapy. The regulatory activity of DNA elements such as promoters and enhancers emerges from the combinatorial organization of transcription factor binding sites and sequence context, which collectively constitute the genome's regulatory grammar. Deep learning models can effectively predict DNA regulatory activity across cell types, but current generative approaches often produce DNA sequences that strongly deviate from this natural regulatory grammar, increasing the risk of unexpected behavior and off-target effects in real-world applications. Here, we introduce DNA-CRAFT, a method for designing regulatory DNA with high predicted cell-type-specific activity while preserving naturalness. DNA-CRAFT formulates regulatory sequence design as a guided search over a generative model pre-trained on millions of naturally occurring regulatory elements from the human and mouse genomes. Our method enables targeted exploration of DNA sequence space, efficient evaluation of candidate designs, and explicit suppression of unwanted activity. Across benchmarks spanning human cell lines and primary immune cells, DNA-CRAFT outperforms existing generative and optimization-based methods, consistently achieving the best trade-off between cell-type specificity and naturalness in the designed regulatory DNA sequences.
DetailAnywhere: Fashion Detail Generation via Cross-Modal Feature Alignment Distillation
Zijun Li ⋅ Yimin Zhou ⋅ Jia Sun ⋅ Honglie Wang ⋅ Pengcheng Wei ⋅ junlong wu ⋅ Yongrui Heng ⋅ Jiyuan Wang ⋅ Huan Ouyang ⋅ Boheng Zhang ⋅ Huaiqing Wang ⋅ Dewen Fan ⋅ Qianqian Gan ⋅ Fan Yang ⋅ Tingting Gao
Diffusion-based generative AI has achieved remarkable success in e-commerce applications such as virtual try-on, poster generation, and product background synthesis. However, when making online purchasing decisions for apparel, consumers also desire the freedom to examine specific detail regions of interest, such as collars, cuffs, and fabric textures, yet existing methods have not explicitly studied this setting. We therefore formalize a new, non-template task: Fashion Detail Generation with focus conditioning, and release FDBench, the first dataset and benchmark comprising 40K+ human-verified reference-detail pairs across 41 different categories. This task poses a unique semantic gap challenge: the model must bridge the correspondence between a focus marker on a product reference image and a photorealistic close-up view of the indicated region, while faithfully preserving the garment's identity, without any precise prompt. To bridge this gap, we propose Cross-modal Feature Alignment Distillation (CFAD), which leverages a fine-tuned DINOv3 teacher to align both branches of a Multimodal Diffusion Transformer in a shared semantic space via dual-branch distillation. To further improve consistency between generated details and reference images, we introduce a consistency reward model that jointly scores image pairs along three quality axes and optimizes generation via reinforcement learning. Experiments show that our model DetailAnywhere significantly outperforms all state-of-the-art opensource methods across all metrics and human evaluations.
Hardware-enabled monitoring of GPU workloads underpins many proposals for AI compute governance, but if developers can defeat monitoring mechanisms, such schemes are unworkable. We evaluate the adversarial robustness of GPU workload classification using only zero-overhead, privacy-preserving NVML telemetry: content-agnostic signals that observe physical effects of computation without accessing model weights, training data, or hyperparameters. Across 5 rounds of monitor-evader iteration, we evaluate 20 evasion strategy families on 9 GPU models spanning 4 architecture generations. We develop a classifier that achieves 98.6% binary accuracy at identifying training workloads across the whole corpus, and 43-87% accuracy against the most challenging unexpected workloads even when they are adversarially disguised.
Dialectics of Alignment: Harnessing Unsafe Knowledge for Dynamic Safety Routing
Maryam Hashemzadeh Barvarz ⋅ Jerry Huang ⋅ Minseon Kim ⋅ Marc-Alexandre Côté ⋅ Sarath Chandar
The prevailing paradigm in large language model (LLM) alignment operates via erasure, filtering unsafe data or training models to strictly refuse harmful prompts. While effective at reducing immediate toxicity, this approach fundamentally constricts the model's epistemological scope, resulting in over-cautious systems that output uninformative blanket refusals to sensitive yet benign queries. In this work, we challenge the orthodoxy that unsafe data must be discarded. We propose a dialectical approach to alignment, positing that unsafe data encodes rich, domain-specific knowledge critical for nuanced, safe, and informative generation. To operationalize this, we introduce SafeMoE, a Mixture-of-Experts (MoE) framework that isolates unsafe knowledge into domain-specific Low-Rank Adapters (LoRA experts) trained exclusively on harmful corpora. To synthesize safety from these unsafe primitives, we train a lightweight gating network using a minimal, highly curated set of safe-informative responses (as few as 100 samples). During inference, this router dynamically orchestrates the unsafe experts, effectively steering the generation trajectory to harness their deep domain knowledge while strictly enforcing safety constraints. Extensive empirical evaluations across stringent safety benchmarks demonstrate that SafeMoE is not only safer, achieving over a 20\% relative improvement in safe response rate (more than a 15\% absolute gain), but also produces more informative responses when safety and harmfulness are of paramount concern. Furthermore, the routing mechanism exhibits strong zero-shot generalization to unseen domains and broader safety tasks without domain-specific supervision. Our findings suggest a paradigm shift in alignment: true safety requires not the masking of unsafe knowledge, but its controlled integration.
DialectLLM: A Dialect-Aware Dialog[ue] Generation Framework Beyond Standard American English
Jio Oh ⋅ Paul Vicinanza ⋅ Thomas Butler ⋅ Steven Whang ⋅ Dezhi Hong ⋅ Amani Namboori
More than 80% of the 1.6B English speakers do not use Standard American English (SAE), yet LLMs often fail to correctly identify non-SAE dialects and generate stereotyped responses for their speakers. We introduce DialectLLM, the first large-scale framework for generating high-quality multi-dialectal conversational data encompassing the three pillars of written dialect---lexical (vocabulary), orthographic (spelling), and morphosyntactic (grammar) features. DialectLLM produces a dialect-parallel dialog dataset spanning nine English dialects. Partnering with native linguists, we design and validate SAE-to-dialect transformation rules, ensuring authenticity. Our approach challenges the prevailing practice of applying a single morphosyntactic feature set to both user utterances and model responses, showing that models should not reproduce up to 90% of the grammatical features of a dialect. Human evaluation confirms data quality, with annotators preferring DialectLLM over prior methods in 98.8% of pairwise comparisons for dialect naturalness. We then construct DialectLLM-Bench, a dialect-parallel benchmark with 50k+ dialogs, resulting in 97k+ QA pairs, and evaluate 17 LLMs on dialect identification and response generation tasks. Even frontier models achieve under 70% accuracy, fail to reach 50% for prominent dialects like Canadian English, and systematically misclassify non-SAE dialects as American or British. Beyond benchmarking, we show that DialectLLM data also serve as a scalable LLM post-training resource, suggesting a practical path toward dialect-aware conversational AI.
DiffCool: Label-Free Synthesis of Chip-Tailored Heat Sinks via Thermal-Aware Diffusion
Siyuan Liang ⋅ Zixiao Wang ⋅ Chenghan Wang ⋅ Shanyi Li ⋅ Yushen Zhang ⋅ Leilei Jin ⋅ Zhen Zhuang ⋅ Ulf Schlichtmann ⋅ Bei Yu ⋅ Tsung-Yi Ho
Recent advances in semiconductor technology have dramatically improved IC chip performance, but soaring transistor densities and clock speeds generate unprecedented heat fluxes in modern integrated circuits, making spatially adaptive thermal management critical. However, conventional heat sinks rely on fixed geometries that lack customization for chip-specific thermal profiles, while the absence of optimal design datasets and strict manufacturing constraints hinder the application of generative AI. To address these challenges, we present DiffCool, a physics-guided framework that reformulates heat sink synthesis as a constrained discrete denoising process. By embedding manufacturing-aware state transitions into the diffusion dynamics and optimizing via a differentiable thermal-aware loss coupled with a surrogate simulator, our model generates structurally sound, high-performance topologies directly from bare-chip heat maps without labeled data. Experiments on an open-source processor demonstrate that DiffCool synthesizes production-ready designs in under one second, reducing peak temperature by 28.3% and temperature gradient by 56.9% compared to conventional baselines. This work establishes a scalable, label-free paradigm that unifies generative modeling with physical constraints for automated thermal design.
Diff-Kalman: Difference-Driven Learning for Structure-Preserving Kalman Filtering
Jia Gao ⋅ Xianglei Xing ⋅ yang qing ⋅ Tianshuo Zhang ⋅ Wenzhe Zhai
State estimation is a critical challenge in modeling dynamic systems. Traditional Kalman filters and data-driven methods have shown good results in complex, nonlinear systems. However, most existing methods optimize the posterior without considering prior modeling, which leaves room for improvement in state estimation accuracy, and predict high-dimensional Kalman parameters end-to-end, which lead to unstable training and poor convergence. To address this, we propose Diff-Kalman, which leverages state differences to predict state increments and scaling matrices for correcting both the prediction and update phases of the Kalman filter. Diff-Kalman combines traditional dynamic models with neural networks to adaptively correct system behavior. In the prediction phase, Diff-Kalman utilizes a Transformer-based architecture with self-attention to extract localized patterns from historical states, followed by cross-attention to predict future state increments. These increments are then fused with model-based priors from the dynamic model. In the update phase, a GRU-based Scaling Predictor (GSP) dynamically adjusts the Kalman covariance and gain matrices based on differential residuals, which improves the accuracy and robustness of the estimation. Experimental results on linear MOT trajectories and nonlinear Lorenz dynamics show that Diff-Kalman improves estimation accuracy over representative baselines.
DiffScore: Text Evaluation Beyond Autoregressive Likelihood
Wen Lai ⋅ Yingli Shen ⋅ Dingnan Jin ⋅ Qing Cui ⋅ Jun Zhou ⋅ Maosong Sun ⋅ Alexander Fraser
Autoregressive language models are widely used for text evaluation, however, their left-to-right factorization introduces positional bias, i.e., early tokens are scored with only leftward context, conflating architectural asymmetry with true text quality. We propose masked reconstruction as an alternative paradigm, where every token is scored using full bidirectional context. We introduce DiffScore, an evaluation framework built on Masked Large Diffusion Language Models. By measuring text recoverability across continuous masking rates, DiffScore eliminates positional bias and naturally establishes an evaluation hierarchy from local fluency to global coherence. We further provide diagnostic tools unavailable to autoregressive frameworks: multi-timestep quality profiles that decompose scores across masking rates, and bidirectional Pointwise Mutual Information (PMI) decomposition that disentangles fluency from faithfulness. Experiments across ten benchmarks show that DiffScore consistently outperforms autoregressive baselines in both zero-shot and fine-tuned settings.
Diffusion Subgoal Planning for Long-Horizon Offline Goal-Conditioned Reinforcement Learning
Hengrui Zhang ⋅ Yuhu Cheng ⋅ C.L.Philip Chen ⋅ Xuesong Wang
Offline goal-conditioned reinforcement learning (GCRL) learns goal-directed policies from reward-free data, but in long-horizon tasks, goal-conditioned value functions often provide unstable guidance due to sparse rewards and discounting. Hierarchical methods partially mitigate this issue via subgoal decomposition; however, high-level decision-making still relies on noise-sensitive value estimates, leading to unstable behavior in complex environments. We address this limitation by proposing \textbf{D}iffusion \textbf{S}ubgoal \textbf{P}lanning (\textbf{DSP}), a diffusion-based framework for high-level subgoal generation. DSP casts high-level planning as guided generative inference over goal-conditioned subgoals and learns both conditional and unconditional flows, enabling classifier-free guidance to introduce a goal-directed bias at inference time. By removing explicit value-based guidance from high-level planning, DSP generates reachable and goal-directed subgoals through a generative model while retaining hierarchical execution. Experiments on offline GCRL benchmarks demonstrate that DSP outperforms prior methods on a range of navigation and manipulation tasks, with particularly strong performance in maze environments that require multi-step subgoal planning.
Direct Conditioning of Audio Diffusion Transformers on fMRI Reveals Cortical Contributions to Sound Reconstruction
Matteo Ciferri ⋅ Tonio Weidler ⋅ Matteo Ferrante ⋅ Nicola Toschi ⋅ Elia Formisano
Reconstructing audio from brain activity remains challenging. A key limitation lies in current audio diffusion models, which are primarily designed for text-based conditioning and lack mechanisms to incorporate non-linguistic continuous signals such as neural responses. While recent work in the visual domain has explored direct conditioning on brain activity, this approach has not yet been extended to audio diffusion models. We introduce a framework that directly guides a latent audio diffusion model with fMRI signals, avoiding intermediate feature prediction for the conditioning stage. Central to our approach is a transformer-compatible adapter that enables the integration of non-linguistic representations into the cross-attention layers of diffusion transformers (DiTs). Although motivated by neural decoding, this adapter provides a general mechanism for conditioning audio diffusion models on arbitrary continuous inputs, not inherently limited to neural signals. We demonstrate consistent improvements over baselines, including two-stage fMRI-to-latent conditioning, across both acoustic and semantic evaluation metrics. In addition, we introduce a fidelity metric grounded in a computational model of auditory neural processing, which allows us to quantify the contribution of individual auditory cortical regions to reconstructed audio representations.
DirectUV: Image-Conditioned UV Texture Generation with Surface-Aware Positional Encoding
Jiantao Lin ⋅ Yingjie Xu ⋅ Mingzhi Sheng ⋅ Yangkai Wei ⋅ Hao CHEN ⋅ Yingcong Chen
Generating high-quality UV textures for 3D meshes remains challenging. Multi-view projection pipelines suffer from occlusion and view inconsistency, and recent methods that generate textures directly in UV space still rely on auxiliary modules to supply 3D information, leaving the attention mechanism tied to UV-grid positions rather than to the underlying surface geometry. This mismatch limits coherence across seams and disconnected UV islands. We propose DirectUV, an image-conditioned UV texture diffusion framework that operates in the latent UV space of a pretrained image VAE, in which a Diffusion Transformer denoises the UV latent given a single input image and a coarse UV map. At its core, Surface-Aware Positional Encoding (SAPE) replaces the standard 2D-grid positional encoding with encodings derived from per-token 3D surface coordinates obtained via UV-to-surface correspondence. As positional encodings define the distance metric that attention operates on, SAPE enables tokens to attend to each other based on true surface proximity rather than UV-grid distance, restoring coherence across seams and disconnected islands. A multi-level extension further assigns different attention heads to progressively finer subdivisions of the same latent UV patch, allowing the model to reason about surface structure at multiple granularities. Experiments show that DirectUV produces sharper and more globally consistent textures than other baselines, with the largest improvements in occluded and view-unseen regions where projection-based methods leave gaps or stretched textures.
DISCOVER: Online Variance-Guided Data Discovery for Budgeted Multimodal GRPO
Peyman Gholami ⋅ Shivam Chandhok ⋅ Saber Malekmohammadi ⋅ Shayan Shams ⋅ Robert Xiao ⋅ Leonid Sigal
Recent advances in reinforcement learning have made verifiable rewards a central post-training paradigm for improving multimodal reasoning. However, methods such as GRPO remain data- and compute-inefficient, often spending rollout and annotation budget on prompts that provide little learning signal. We propose DISCOVER, an online data-discovery framework for budgeted multimodal GRPO that enables vision-language models to dynamically identify prompts that are most useful for the current policy. DISCOVER first organizes the multimodal sample pool into semantic clusters and queries a small set of density-diverse representative samples to obtain initial reward observations. During training, it estimates the utility of unannotated samples by transferring sparse reward feedback from queried samples. Since sample utility changes as the policy evolves, DISCOVER dynamically updates these estimates with recency-weighted rollout observations. It then selects only high-utility prompts for annotation using a policy-conditioned score that accounts for reward variance and difficulty, and continues utility-based replay once the discovery budget is exhausted. Experiments on ViRL39K and VLAA-Thinking with Qwen3-VL-2B show that DISCOVER substantially improves multimodal mathematical reasoning under a strict 10\% unique-sample budget, outperforming budget-matched random selection, hard-example sampling, full-pool GRPO under the same training-step budget, and online selection baselines. Ablations further validate the importance of semantic coverage, online reward transfer, and stable discovery cadence, establishing DISCOVER as a practical framework for data-efficient multimodal RL post-training.
DiscoverPhysics: Benchmarking LLMs for out-of-the-box scientific thinking
Lindsay Smith ⋅ Matt Sampson ⋅ Siddharth Mishra-Sharma ⋅ Peter Melchior ⋅ Andrew Wilson ⋅ Pavel Izmailov ⋅ Carolina Cuesta Lazaro
Frontier LLMs now perform strongly across a wide range of physics evaluations, but it is hard to disentangle genuine reasoning from recall of established science. We introduce DiscoverPhysics, an interactive benchmark that asks a LLM agent to discover the laws of motion of a simulated world whose physics deliberately deviates from our own. We construct eleven worlds governed by, among others, screened and fractional-power gravity, multi-species couplings, hidden dark-matter-like particles, non-coordinate-free physics, and time-varying interactions. Each world is generated on demand by an N-body simulator, for which the agent proposes several rounds of experiments, observes raw trajectory data, and ultimately submits both a natural-language explanation of the world's physics and a Python implementation of the inferred law. Because solving a world requires the agent to design informative experiments and revise its hypotheses, the benchmark probes long-horizon reasoning over an experimental history. We evaluate submissions along two complementary axes: trajectory MSE on held-out particles and an LLM-judged explanation score following an expert-written rubric assessing conceptual understanding of each world. Across ten frontier models, we find that the strongest agents pass only about half of the worlds and consistently fail on those where latent structure must be uncovered. Open-source models lag substantially behind commercial models, both in their ability to design informative experiments and in extracting conclusions from the data. We further find that good predictive accuracy does not guarantee high explanation quality and that conceptual understanding depends on hypothesis refinement through well-chosen experiments.
Distance-Dependent Connectivity Shapes Continual Learning by Synaptic-Resource-Delimited Separation of Neural Dynamics
CHIU-CHANG CHENG ⋅ Ching-Lung Hsu ⋅ Ya-Ning Chang ⋅ Chao-Hung Wang
Biological circuits learn without catastrophic forgetting, but the structural basis of this ability remains unclear. We investigate whether distance-dependent connectivity (DDC), spatial recurrency found in the mammalian cortex, contributes to continual learning by embedding neurons of a biologically plausible recurrent spiking neural network in a 3D Euclidean substrate, using pairwise distance to set recurrent connection probability, and ablating that distance dependence in silico. We show that DDC shapes a synaptic resource geography: it controls the spatial breadth of the candidate synaptic pool for learning to occur, plasticity further compresses that pool, and the resulting substrate governs how subsequent inputs compete for the same synapses. On challenging 10-way class-incremental MNIST classification, this produces a non-monotonic accuracy curve, with an intermediate DDC range achieving the best performance (71.4%) and outperforming both highly local circuits that collapse into a synaptic-resource bottleneck and random connectivity that produces weakly guided, broader synaptic contention. This pattern was better explained by synaptic competition, resulting in separable neural dynamics, than by neuron assembly separation. These results suggest that appropriate cortical DDC may prevent diffuse random contention. For continual learning problems, this finding highlights the conceptual importance of circuit dynamics more than active neuron ensembles and engrams.
The canonical orientation of a 3D object is often not unique but inherently ambiguous: many objects admit multiple equally valid canonical frames that no single-output regressor can faithfully represent. Rotational symmetry is the most structured manifestation of this phenomenon, where the set of valid orientations forms an orbit of the symmetry group. Existing methods fall short in one of three ways: they predict only a single orientation, handle symmetry only under pre-defined group assumptions, or resort to test-time augmentation to approximate the full set of valid orientations. We instead introduce a general framework for ambiguity-aware orientation estimation that is at once $SO(3)$-equivariant by construction, distributional in its output, and free of any restriction on the symmetry group. It effectively predicts multiple valid orientations as a continuous multi-modal distribution over $SO(3)$ via a truncated Wigner-D expansion, without any symmetry-group assumption. The proposed method outperforms recent equivariant regressors and fixed-quotient classifiers on the ShapeNet orientation benchmark, particularly on categories with continuous rotational symmetry, showing that supervision on discrete symmetry labels generalises to continuous bands.
Distributionally Robust Domain Randomization with Learned Risk-Sensitive Dynamics Samplers
Sukchul Jeong ⋅ Insoon Yang
Domain randomization (DR) improves sim-to-real transfer by training policies over randomized simulator dynamics, but standard DR optimizes average performance under a fixed sampler and can miss rare failure-prone domains. We propose Risk-Sensitive Domain Randomization (RSDR), a distributionally robust formulation of episodic DR that evaluates each policy against the worst-case distribution over dynamics parameters within a KL neighborhood of a reference sampler. The resulting soft adversary has an exponential-tilting form, yielding a single temperature-controlled sampler family: negative temperatures emphasize low-return domains for robustness, zero recovers uniform DR, and positive temperatures give an optimistic sampler closely related to curriculum-style DR. Because this target changes with the policy and returns are observed only through rollouts, RSDR learns an amortized dynamics sampler by reverse-KL variational inference while reusing on-policy PPO trajectories. Across six domain-randomized MuJoCo Playground tasks, robust RSDR improves CVaR and minimum-return metrics over uniform and adaptive DR baselines, and outperforms best-tuned EPOpt.
Distributionally Robust Listwise Preference Optimization
Xudong Wu ⋅ Jian Qian ⋅ Pangpang Liu ⋅ Vaneet Aggarwal ⋅ Jiayu Chen
Existing robust preference optimization for language-model alignment mainly studies pairwise supervision and places robustness at the dataset, prompt, or preference-pair level. We instead study listwise preference optimization under ranking-label uncertainty: given a prompt and a candidate list, the observed ranking over that list may be ambiguous due to annotator inconsistency, near-ties, lossy rankwise feedback, or reward-model noise. We propose a pointwise total-variation robust Plackett--Luce objective that directly robustifies the ranking label conditional on the candidate list. The robust loss admits an exact decomposition into the nominal PL loss plus a worst-case PL correction, and the worst-case ranking is obtained by sorting current implicit scores in ascending order, reducing the inner maximization from $K!$ enumeration to $O(K\log K)$. This tractable structure yields strong offline and online optimization guarantees. In the offline fixed-list setting, the robust objective is convex and projected stochastic subgradient reaches global $\epsilon$-suboptimality with $O(\epsilon^{-2})$ sample complexity. In the online policy-induced setting, where candidate lists are generated by the current policy, we establish weak convexity and $\widetilde O(\epsilon^{-2})$ Moreau-envelope stationarity. Experiments in offline LLM alignment show that the proposed robust correction largely preserves performance under clean labels and improves robustness under noise. In online alignment, it makes reward-model-ranked candidate expansion more reliable and improves both reward-model andexternal GPT-4 judge metrics.
Distribution Shift in Missing Data Imputation: A Risk-Based Perspective and Importance-Weighted Correction under MAR
Luke Shannon ⋅ Song Liu ⋅ Katarzyna Reluga
Missing data imputation, where a model is trained on observed data to estimate unobserved values, is a fundamental problem in machine learning. In this paper, we rigorously formulate imputation model learning as a mean-squared error risk minimisation problem. We show that when the probability of missingness depends on the data, many state-of-the-art methods fail to account for the resulting distribution shift between the observed data used for training and the full data distribution used for evaluation. Consequently, these approaches do not minimise mean-squared error on the full data distribution. Instead, we propose a novel imputation algorithm designed to learn an imputation model from the observed data while explicitly accounting for this distribution shift. Simulation studies show consistent improvements over otherwise identical uncorrected baselines, with average reductions of 3\% in RMSE and 7\% in Wasserstein distance.
DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control
Teli Ma ⋅ Jia Zheng ⋅ Zifan Wang ⋅ Chunli Jiang ⋅ Andy Cui ⋅ Junwei Liang ⋅ Shuo Yang
Vision-Language-Action (VLA) models inherit their visual backbone from static image-text pretraining, leaving physical dynamics to be learned from scarce action data. Generative video models, by contrast, already encode motion, contact, and implicit physics at internet scale. We introduce DiT4DiT, a Video-Action Model (VAM) that couples a Video Diffusion Transformer with an Action Diffusion Transformer and co-trains them under a dual flow-matching objective with a tri-timestep scheme for video supervision, action supervision, and feature extraction. The key insight is that the policy does not need to generate a future video: we intercept the Video DiT's hidden state at a single fixed flow timestep, turning a multi-step video rollout into a one-shot feature extraction consumed by the Action DiT. Across benchmarks, DiT4DiT sets new SOTA averages on LIBERO (98.6%) and the 24-task RoboCasa-GR1 suite (56.7%), outperforming strong VLA and video-based baselines. On a real Unitree G1, DiT4DiT reaches 73.8% on eight tabletop tasks and 72.2% on three whole-body loco-manipulation tasks, running up to 12$\times$ faster than prior VAM baselines while also exceeding them in accuracy. To our knowledge, this is the first video-generation-based policy to run humanoid whole-body control at real-time rates. It further generalizes zero-shot to unseen object categories and scene variations. Together, these results indicate that video-dynamics priors, accessed through a single forward pass, are a practical and scalable foundation for generalist robot policies.
DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation
Yongkun Du ⋅ Pinxuan Chen ⋅ Xuye Ying ⋅ Zhineng Chen
The advent of Multimodal Large Language Models (MLLMs) has unlocked the potential for end-to-end document parsing and translation. However, prevailing benchmarks such as OmniDocBench and DITrans are dominated by pristine scanned or digital-born documents. They do not adequately represent the intricate challenges of real-world capture conditions, such as geometric distortions and photometric variations. To fill this gap, we introduce DocPTBench, a comprehensive benchmark specifically designed for Photographed Document Parsing and Translation. DocPTBench comprises over 1,300 high-resolution photographed documents from multiple domains, includes eight translation scenarios, and provides meticulously human-verified annotations for both parsing and translation. Our experiments demonstrate that transitioning from digital-born to photographed documents results in a substantial performance decline: popular MLLMs exhibit an average accuracy drop of 18\% in end-to-end parsing and 12\% in translation, while specialized document parsing models show a more prominent average decrease of 22\%. This substantial performance gap highlights the unique challenges posed by documents captured in real-world conditions and reveals the limited robustness of existing models. The benchmark and related code will be made publicly available for further research.
Does Compression Imply Generalization? A Minimum Description Length Perspective
Lincen Yang ⋅ Jiayang Shi ⋅ Nan Pu ⋅ Daniel M. Pelt ⋅ Matthijs van Leeuwen ⋅ Zhong Li ⋅ Shujian Yu
A central question in modern learning theory is what notion of compression, if any, explains why over-parameterized neural networks generalize. Existing answers measure complexity at different levels, ranging from weights to representations, and estimating the associated quantities either requires held-out data in principle or relies on mutual-information estimation in ultra-high dimensions. We approach the question through the minimum description length (MDL) principle and propose the prequential regret as a complementary signal that is computable from the training trajectory alone. We establish that, in the (nearly) interpolating regime, the expected prequential regret upper-bounds the description length of the trained parameter under the algorithm-induced distribution, and that this description length translates into a high-probability generalization bound in the PAC-MDL setting. Along the way, we show that the encoder-complexity correction recently added to the algorithm-level information bottleneck is upper-bounded by an MDL parameter description length, bringing the two views of compression onto a shared footing. Empirically, the prequential regret tracks the generalization error on CIFAR-10 and CIFAR-100, with Pearson correlations exceeding 0.9.
Does Your Neural Network Extrapolate? Feature Engineering as Identifiability Bias for OOD Generalization
Leonel Aguilar ⋅ Jan Nagler ⋅ Christoph Hoelscher ⋅ Nino Antulov-Fantulin
Successful deep neural networks discover salient features of data. We show when and why they fail to learn out-of-distribution (OOD)-relevant representations from an in-distribution (ID) training window. This requires decoupling feature learning from data-generating-process (DGP) identifiability. From a single training window, OOD extrapolation is non-identifiable: infinitely many DGPs are $\varepsilon$-observationally equivalent on the training data but diverge arbitrarily outside it, and no in-distribution criterion alone reliably breaks the tie. A structural commitment, the feature map, label map, and model class $(\varphi, \psi, \mathcal{M})$, dictates the assumed DGP and governs OOD generalization while leaving ID performance essentially unchanged. When architecture, pretraining, augmentation, input formats, or domain knowledge implicitly inject the missing commitment, the model succeeds. When it cannot infer OOD-relevant structure from ID evidence, it fails. Changing only the representation can make the same architecture, at the same in-distribution loss, differ by ${\sim}520\times$ out of distribution. When the commitment is correct {\em and} identifiable, OOD error vanishes. For example, Fourier coordinates turn periodic extrapolation into interpolation on $\mathbb{S}^1$. The same mechanism predicts outcomes in three natural-science settings (mass-action chemistry; Kepler's-third-law exoplanet prediction, $n=2{,}362$; and cross-species coding-DNA detection) and in a 264-run positional-encoding study across Transformer, Mamba, and S4D. Finally, a controlled study shows: correct features are necessary but not sufficient. The model class must express the target, and the transformed training data must cover the relevant representation space. Thus, feature engineering is not a departure from what makes deep learning successful. It is the explicit structural commitment that makes extrapolation identifiable.
Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents
Ying Chen ⋅ Lihuang Fang ⋅ Rui Jiang ⋅ Mingxu Wang ⋅ Zhifeng Gu ⋅ Lei Yi ⋅ Jie Chen
Standard embodied evaluations do not independently score whether an agent correctly commits to task completion at episode closure—a capacity we call terminal commitment. Behaviorally distinct failures—never completing the task, completing it but failing to stop, and reporting success without sufficient evidence—collapse into the same benchmark failure. We introduce VIGIL, an evaluation framework that makes terminal commitment independently measurable. Under VIGIL’s default protocol, agents observe only egocentric RGB, receive no action-success signals, and must end each episode with a semantic report checked deterministically against hidden world state. This yields two separate scores: world-state completion (W) and benchmark success (B), where B additionally requires a correct terminal report. This decoupling makes four outcome categories distinguishable: missed execution, post-attainment drift, unsupported commitment, and verified success. Across 20 models on 1,000 frozen episodes, systems with comparable W differ by up to 19.7 pp in B: one model converts achieved states into correct reports, while another with near-identical execution drifts past the goal without closing. An action-feedback intervention further tests the separation: execution-oriented signals improve W broadly, yet commitment failures persist in models that do not already ground terminal reports in the achieved state. VIGIL provides a protocol that makes terminal commitment independently visible and scorable.
Don't Discard Your Rollouts: Reusing Teacher RL Traces for Student Distillation
Vashisth Tiwari ⋅ Emma Strubell ⋅ Zico Kolter
RL post-training with methods like PPO and GRPO produce a large rollout archive, which each contain correctness-labeled completions, correct/incorrect completions for the same prompt, teacher log-probabilities, and group pass rates. Canonical distillation methods, however, ignore these traces and train based upon new data from the converged teacher. This discards the contrastive signal that helped the teacher and the metadata already produced during teacher RL, even though the teacher RL run typically generates many more tokens than the later distillation pass. However, reusing rollouts naively is non-trivial: the archive mixes early and late teacher policies, can be arbitrarily off-policy for the student, and contains many low-quality traces, and naive SFT and DPO on these rollouts underperforms the cannonical SFT baseline. In this paper, we show that rollout reuse becomes effective when it is treated as a data-selection problem. We use two sets of information from rollouts: per-prompt pass rate, which defines an easy-to-hard curriculum and separates all-correct prompts for SFT from mixed-success prompts for DPO; and per-rollout teacher likelihood, which acts as an in-distribution proxy for student-compatible traces. Across two GRPO teachers (4B and 8B) and students from multiple families, the curated rollout datasets DSFT^{allp,top1} and DDPO^{curr,top1} match or beat canonical distillation without an extra teacher generation pass. On MATH500, DDPO^{curr,top1} improves over Dsynth by up to 4.0 percentage points on strong students, while DSFT^{allp,top1} → DDPO^{curr,top1} matches weaker students where DPO is unstable.
Modern GPU-native solvers for combinatorial optimization exploit parallelism by evolving large populations of relaxed candidates; yet, their post-processing collapses this computational effort into a single, rounded solution. We show that the discarded replicas of Parallel Quasi-Quantum Annealing (PQQA) form an implicit elite archive that contains basin-level information recoverable after annealing. We introduce Iterative Path-Relinking Polish (I-PRP), a deterministic, training-free post-anneal operator that leaves the PQQA inner loop unchanged. I-PRP polishes the top rounded elites, traverses cost-guided relinking paths from the current best elite toward the others, re-polishes intermediate states to cross basin boundaries, and iterates only under strict improvement. The operator never degrades the paired upstream polish, terminates after a finite number of updates, and reduces to the original polish when the population collapses to a single basin. Across QUBO and graph-coloring benchmarks, I-PRP preserves saturated cases while improving PQQA on rugged multi-basin instances, revealing that parallel annealing populations contain reusable search structures beyond their best replicas.
Do Sparse Autoencoders Learn Meaningful Concept Hierarchies?
Nils Grandien ⋅ David Steinmann ⋅ Felix Friedrich ⋅ Kristian Kersting
Sparse autoencoders (SAEs) have become an important tool for unsupervised concept discovery in large models. To make the resulting feature spaces more interpretable and manageable, recent approaches have begun imposing hierarchical structure, either explicitly or as an implicit effect of training constraints, yet rigorous comparison remains difficult. There are no agreed-upon requirements for what a meaningful feature hierarchy should satisfy, and evaluation has largely relied on qualitative illustrations with fragmented quantitative protocols. To address this, we derive a set of key requirements for generalization/specialization hierarchies in unsupervised concept discovery, drawing on semantic net and taxonomy research alongside recent SAE work, and use them to derive a concrete evaluation protocol. Applying this protocol to current SAE approaches trained on visual data, we find that while feature spaces generally provide a basis for sensible hierarchies, establishing good hierarchical structure remains challenging. In particular, feature absorption, both in its well-known hard form and in a continuous, soft form, systematically compromises hierarchy quality, pointing to a fundamental tension that future approaches will need to navigate.
DPIAgent: Divide, Protocol, Isolate for Agentic Reproduction Test Generation
Hao Liu ⋅ Steven Liu ⋅ Xin Zhang ⋅ Jane Luo ⋅ Yu Kang ⋅ Jie Wu ⋅ Fangkai Yang ⋅ Yangyu Huang ⋅ Pengfei Gao ⋅ Scarlett Li ⋅ Yan Lu
Reproduction test generation, producing a failing-then-passing test that captures a reported bug, is a critical step in automated software engineering. Existing agentic methods treat this as a monolithic loop, despite the task inherently comprising two subtasks of distinct nature: diagnosing the root cause and writing a fail-to-pass test. Without explicit separation, the agent faces a compound objective with underspecified intermediate goals, leading to goal drift. We propose DPIAgent, a structured agentic framework built on three principles, Divide, Protocol, Isolate (DPI), that mitigates compound-objective ambiguity and goal drift: it Divides the task into single-objective phases of defect exploration and test generation; enforces a handoff Protocol that records the diagnosis and test plan, preventing context loss; and Isolates each phase's action space by tailoring the toolset to its task, preventing irrelevant tools from misleading execution. On SWT-Bench Verified, DPIAgent outperforms seven baselines across three backbone LLMs. With DPI alone it reaches 81.76% success rate on GPT-5, the highest reported among open-source methods, gaining up to 11.88 points over the strongest baseline on GPT-5-Mini; adding test selection further raises it to 86.17%. Our analysis shows that architectural structure and backbone capability are complementary axes rather than substitutes, demonstrating DPI's generalizability across model classes.
dRAE: Representation Autoencoder with Hyper-Spherical Codes
Tianren Ma ⋅ Lin Long ⋅ Chuyan Chen ⋅ Mu Zhang ⋅ Junbo Zhao ⋅ Tong Zhang ⋅ Qixiang Ye
Multi-modal models require visual tokenizers that jointly capture high semantic density and fine-grained structural fidelity. Representation Autoencoders (RAEs) offer a promising direction by decoding images directly from a pre-trained semantic space using high-dimensional continuous tokens, bypassing the information bottleneck of VAEs and benefiting both visual understanding and generation. In this work, we aim to discretize these continuous vision representations to bridge the gap with language models — a non-trivial challenge, as existing quantization methods suffer from codebook collapse, failing to scale while preserving semantic coherence. We identify the root cause as **metric mismatch**: standard Euclidean codebook objectives are fundamentally misaligned with the anisotropic geometry of representation space, leading to codebook embeddings with high-variance magnitude scales and uneven angular distributions that hinder scalability. To address this, we propose **Hyper-Spherical Quantization (HSQ)** , which decouples semantic content from feature magnitude via angular routing, preventing code assignment from being dominated by scale rather than meaning. The resulting **discrete Representation Autoencoder (dRAE)** achieves high-fidelity reconstruction while preserving semantic integrity and supporting scalable discrete latent spaces. Extensive experiments demonstrate consistent performance gains as the vocabulary size scales to $131,072$, along with strong unified performance across both visual understanding and generation benchmarks.
DriveStreamBench: Evaluating User-Conditioned Watch-and-Notify in Streaming Driving Video
Yi Wang ⋅ Xin Zhao ⋅ Tian Meng ⋅ Rui Dai ⋅ Chi Zhang ⋅ Cui Tang ⋅ Bingzheng Liu ⋅ Fucheng Wang ⋅ Haopeng Zhang ⋅ Ruidong Ding ⋅ Yang Li ⋅ Kaikui Liu ⋅ Xiangxiang Chu
Vision-language models (VLMs) have made steady progress as in-vehicle driving assistants, yet existing evaluations largely remain offline, focusing on perception and short-term reasoning over pre-recorded driving scenes. A complementary capability is still underexplored: whether a model can continuously monitor a causal driving video stream and issue a calibrated alert when a user-specified condition is met. Existing driving benchmarks mainly assess reactive understanding, while recent streaming video benchmarks are developed largely outside the driving domain. To this end, we introduce DriveStreamBench, a driving and streaming video benchmark that organizes this capability spectrum into four cognitive levels constructed from a shared pool of driving events, to support level-wise capability profiling and gated evaluation. We evaluate various vision-language models along with DriveStream-SFT, a supervised reference baseline trained on the benchmark's training split, under a unified streaming protocol. DriveStreamBench tests a basic but necessary capability, and all evaluated models fall short of it; current vision-language models are not yet ready for user-conditioned alerting in driving. Resources are available at https://anonymous.4open.science/r/DriveStreamBench/ and https://dataverse.harvard.edu/previewurl.xhtml?token=01f6bc4c-d809-4e40-b025-b6d783c86571.
Drive vs. Decay: On the Training Dynamics of Joint-Embedding Predictive Architectures
José Lucas De Melo Costa ⋅ Seong Woo Ahn ⋅ Fabrice Popineau ⋅ Arpad Rimmel ⋅ Bich-Liên DOAN
Joint-Embedding Predictive Architectures (JEPAs) are prone to representation collapse, typically mitigated through empirical heuristics. We develop an early-training stability theory that unifies these heuristics. Linearising the coupled JEPA gradient flow around the trivial fixed point reveals two competing effects: a driving force ($\gamma$) and a decay effect ($\sigma$). Under approximate spectral decoupling, a per-mode stability ratio $\mu_i = \gamma_i / \sigma_i$ factorises into independent data-side and predictor-side terms and the count of unstable modes tracks the rank of representations that can emerge. The framework predicts a phase boundary, which we confirm empirically across more than 800 Tabular-JEPA configurations. It also unifies predictor scaling, masking ratio, and EMA as distinct mechanisms for shifting $\mu$. Guided by this analysis, we introduce ResidualPred, a transformer predictor whose attention is biased toward the identity at initialisation; it improves both effective rank and downstream accuracy on tabular benchmarks and on I-JEPA pretraining over CIFAR-10, CIFAR-100, and STL-10. Our framework connects empirical collapse-avoidance heuristics to an explicit dynamical picture, yielding theory-driven stabilizers. Code is available as supplementary material.
DROGO: Default Representation Objective via Graph Optimization in Reinforcement Learning
Hon Tik (Rick) Tse ⋅ Marlos C. Machado
The default representation (DR), originally introduced in neuroscience, and its principal eigenvector have been shown to be effective for a wide variety of applications, including reward shaping, count-based exploration, option discovery, and transfer. However, in prior investigations, the eigenvectors of the DR were computed by first approximating the DR matrix, and then performing an eigendecomposition. This procedure is computationally expensive and does not scale to high-dimensional spaces. In this paper, we propose an objective for approximating the principal eigenvector of the DR using only transitions with a neural network. This objective is inspired by a series of theoretical results and is empirically validated in a number of environments. We then demonstrate its usefulness by applying the learned eigenvectors for reward shaping.
DSAQuant: Denoising-Stage-Aligned Quantization-Aware Training for Video Generation
Shuaiting Li ⋅ Zelin Gao ⋅ Haibin Shen ⋅ Yujun Shen ⋅ Haotong Qin ⋅ Yinghao Xu
Video diffusion models (VDMs) have achieved impressive progress in text-to-video generation, but their high memory and computational costs hinder practical deployment. Quantization-aware training (QAT) is an effective solution for compressing and accelerating advanced generative models without runtime overhead at inference. % However, existing QAT methods suffer from a distinctive challenge in VDMs: while they often preserve prompt semantics, global layout, and coarse motion, the quantized model severely degrades visual details, texture fidelity, and sharpness. In this paper, we trace this degradation to the timestep-agnostic design of conventional quantization pipelines, which overlooks the stage-wise functionality of video denoising. In VDMs, early denoising steps mainly establish global structure and motion, whereas middle and late steps refine local appearance and high-frequency details. Based on this insight, we propose \textbf{DSAQuant}, a \textbf{D}enoising-\textbf{S}tage-\textbf{A}ligned \textbf{Quant}ization-aware training framework for VDMs. During training, \textit{Denoising-Stage Oriented Supervision} preserves teacher distillation in early steps for stable structure planning, while shifting later steps toward target-driven optimization to enhance detail reconstruction. During inference, \textit{Denoising-Stage Gated Guidance} disables CFG in the final denoising steps to prevent it from amplifying quantization-induced errors into high-frequency artifacts. Extensive experiments on the Wan and CogVideoX families under W4A4 and W3A3 settings show that DSAQuant consistently outperforms the SOTA QAT baseline, improving the VBench average score by up to 6.60 under aggressive W3A3 quantization while preserving strong text-video alignment. These results demonstrate that effective VDM quantization requires not only reducing quantization error, but also aligning quantization training and inference with the stage-wise nature of video diffusion.
DSSNet: Deep Spectral Structure Profiling Network for Traffic Flow Prediction
Tao Wang ⋅ Zhenye Yang ⋅ Hua Lu ⋅ Huan Li ⋅ Senzhang Wang ⋅ Jinpeng Chen
Traffic flow prediction is a fundamental task in intelligent transportation systems. Existing methods primarily rely on temporal dependency modeling and thus face two critical limitations: struggling to effectively capture diverse periodicity; overlooking the relationship between spectral structures in the traffic flow and spatiotemporal traffic behaviors. To address these limitations, we propose DSSNet (Deep Spectral Structure Profiling Network), an adaptive spectral profiling framework. DSSNet follows the paradigm of "spectral structure profiling" and then "dual-domain modeling". Initially, a Spectral Structure Learning module captures the dominant frequency bands of the spectrum to enhance regional representations while simultaneously serving as discriminative spatiotemporal identity between regions. Based on these identities, we construct a global Resonance Graph to uncover cross-region dependencies and introduce a Dual-Domain Knowledge Integration module to aggregate structural knowledge across both time and frequency domains. Extensive experiments on multiple benchmark datasets demonstrate that DSSNet achieves state-of-the-art performance, while highlighting the potential of frequency domain modeling for understanding spatiotemporal traffic behaviors.
DTA-GT: Direction- and Topology-Aware Graph Transformer for Neural Network Representation Learning
Yuxiang Zeng ⋅ Kun Xie ⋅ Yang Wang ⋅ Jigang Wen ⋅ Xiaocan Li ⋅ Guangxing Zhang ⋅ Gaogang Xie ⋅ Jiannong Cao
Architecture attribute predictors reduce the cost of neural architecture search by estimating accuracy or latency from a small set of evaluated candidates. The key difficulty is that neural architectures are operation-labeled DAGs: edge direction defines computational flow, but existing predictors often treat direction as a local structural cue. We propose DTA-GT, a Direction- and Topology-Aware Graph Transformer based on encoder-wide directional consistency, which preserves directed computational semantics across node initialization, global spectral encoding, pairwise interaction, and feature update. DTA-GT realizes this principle with direction- and topology-aware node initialization, Magnetic Laplacian spectral encoding, direction-dual structural attention, and a Direction-Sensitive MoE-FFN. Across accuracy and latency prediction on NAS-Bench-101 and NAS-Bench-201, DTA-GT consistently outperforms representative sequence-, GNN-, Transformer-, and hybrid predictors under limited-budget settings. Ablations and Magnetic spectral controls show that the improvements come from coordinated direction-aware stages and edge-orientation-sensitive spectral information, while mechanistic diagnostics confirm that the modules exhibit the intended directional behavior. These results support encoder-wide directional consistency as an effective design principle for architecture DAG representation.
Duality Models: An Embarrassingly Simple One-step Generation Paradigm
PENG SUN ⋅ Xinyi Shang ⋅ Zhenglin Cheng ⋅ deyuan liu ⋅ Tao Lin ⋅ Zhiqiang Shen
Consistency-based generative models like Shortcut and MeanFlow achieve impressive results via a target-aware design for solving the Probability Flow ODE (PF-ODE). Typically, such methods introduce a target time $r$ alongside the current time $t$ to modulate outputs between a local multi-step derivative ($r = t$) and a global few-step integral ($r = 0$). However, the conventional "one input, one output" paradigm enforces a partition of the training budget, often allocating a significant portion (e.g., 75% in MeanFlow) solely to the multi-step objective for stability. This separation forces a trade-off: allocating sufficient samples to the multi-step objective leaves the few-step generation undertrained, which harms convergence and limits scalability.To this end, we propose Duality Models (DuMo) via a "one input, dual output" paradigm. Using a shared backbone with dual heads, DuMo simultaneously predicts velocity $\mathbf{v}_t$ and flow-map $\mathbf{u}_t$ from a single input $\mathbf{x}_t$. This applies geometric constraints from the multi-step objective to every sample, bounding the few-step estimation without separating training objectives, thereby significantly improving stability and efficiency. On ImageNet $256 \times 256$, a 679M Diffusion Transformer with SD-VAE achieves a state-of-the-art (SOTA) FID of 1.79 in just 2 steps.Code will be publicly available.
DualSteer: Dual-Space Steering for Robust Jailbreak Mitigation of Large Vision Language Models
Haotian Zhu ⋅ Shuchao Pang ⋅ Zhigang Lu ⋅ Jiluan Fan ⋅ Fanzhen Liu ⋅ Xu Zheng ⋅ Xunzhu Tang ⋅ Minhui Xue
Large Vision Language Models (LVLMs) have demonstrated outstanding capabilities in multimodal content understanding, yet their multimodal nature also leads to security vulnerabilities. Existing defense methods often suffer from high computational cost and inference latency. In contrast, steering-based methods avoid these issues but still depend on heuristic model representations and struggle to handle coordinated multimodal attacks, where malicious content from images or text prompts can hijack the model’s limited attention resources. We identify this phenomenon as attention distraction, where an LVLM’s attention shifts from safety reasoning toward maliciously aligned tokens. To this end, we propose DualSteer, a lightweight and training-free defense framework that restores safe reasoning via dual-space steering. DualSteer first adjusts the attention distribution to counter the influence of jailbreak distractors, and then steers hidden layer representations to address residual jailbreak effects that the model cannot resist on its own. Extensive experiments across six jailbreak benchmarks and three LVLMs of different architectures demonstrate that DualSteer consistently reduces attack success rates by over 30\% relative to state-of-the-art defenses, while ensuring the model’s utility and avoiding false rejections. DualSteer provides a novel defensive perspective while offering an efficient and robust solution to defend against complex multimodal jailbreak attacks.
DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech
Xu Zhang ⋅ Longbing Cao ⋅ zhangkai wu
Diffusion and flow-matching based text-to-speech (TTS) models excel in naturalness but often lack explicit emotion control, as emotional signals remain entangled with speaker identity. We discover that emotion embedding emerges as a linearly decodable direction of frozen hidden states, nearly orthogonal to the direction embedding speaker identity. This inspires a plug-and-play framework DUET for emotion control over pretrained diffusion and flow-matching based TTS models. During generation, DUET unifies dual-space control to achieve fine-grained emotion intervention in a single per-step update: hidden space steering shifts generation along the target emotion direction, while mel-space guidance refines spectral details through gradients backpropagated from a differentiable vocoder. We validate DUET on five architecturally diverse pretrained TTS backbones across three datasets, where it outperforms 10 supervised state-of-the-art emotional TTS baselines across paradigms and achieves the highest human-rated emotion appropriateness. To further showcase its qualitative behavior, we deploy DUET on an Ameca humanoid robot, where it produces richly expressive emotional speech on the humanoid, demonstrating the strong potential for plug-and-play affective interaction for embodied agents. The demonstration is available at: https://anonymous.4open.science/w/duet-demo-fresh-D8E0
D-VLA: A High-Concurrency Distributed Asynchronous RL Framework for Vision-Language-Action Models
Guo Yucheng ⋅ Yongjian Guo ⋅ 钟 关 ⋅ Wen Huang ⋅ Haodong Yue ⋅ shuai di ⋅ Junwu Xiong ⋅ Yicheng Gong
The rapid evolution of Embodied AI has enabled Vision-Language-Action (VLA) models to excel in multimodal perception and task execution. However, applying Reinforcement Learning (RL) to these massive models in large-scale distributed environments faces severe systemic bottlenecks, primarily due to the resource conflict between high-fidelity physical simulation and the intensive VRAM/bandwidth demands of deep learning. This conflict often leaves overall throughput constrained by execution-phase inefficiencies. To address these challenges, we propose D-VLA, a high-concurrency, low-latency distributed RL framework for large-scale embodied foundation models. D-VLA introduces "Plane Decoupling," physically isolating high-frequency training data from low-frequency weight control to eliminate interference between simulation and optimization. We further design a four-thread asynchronous "Swimlane" pipeline, enabling full parallel overlap of sampling, inference, gradient computation, and parameter distribution. Additionally, a dual-pool VRAM management model and topology-aware replication resolve memory fragmentation and optimize communication efficiency. Experiments on benchmarks like LIBERO show that D-VLA significantly outperforms mainstream RL frameworks in throughput and sampling efficiency for billion-parameter VLA models. In trillion-parameter scalability tests, our framework maintains exceptional stability and linear speedup, providing a robust system for high-performance general-purpose VLA Models.
DyJR: Preserving Local Policy Plasticity in Reinforcement Learning with Verifiable Rewards via Dynamic Jensen-Shannon Replay
Long Li ⋅ Zhijian Zhou ⋅ Tianyi Wang ⋅ Weidi Xu ⋅ Zuming Huang ⋅ Wei Chu ⋅ Zhe Wang ⋅ Shirui Pan ⋅ Chao Qu ⋅ Yuan Qi
While Reinforcement Learning (RL) enhances Large Language Model reasoning, on-policy algorithms like GRPO are sample-inefficient as they discard past rollouts. Existing experience replay methods address this by reusing accurate samples for direct policy updates, but this often incurs high computational costs and causes mode collapse via overfitting. We argue that historical data should prioritize sustaining local policy plasticity rather than simply reinforcing accuracy. We define local policy plasticity as a persistent ability to remain backtrackable and switchable with respect to recent successful trajectories: throughout training, the policy maintains non-negligible support on recently reachable alternatives, so it can smoothly reallocate probability mass when needed without being locked into a single dominant path. To this end, we propose Dynamic Jensen-Shannon Replay (DyJR), a simple yet effective regularization framework using a dynamic reference distribution from recent trajectories. DyJR introduces two innovations: (1) A Time-Sensitive Dynamic Buffer that uses FIFO and adaptive sizing to retain only temporally proximal samples, synchronizing with model evolution; and (2) Jensen-Shannon Divergence Regularization, which replaces direct replay updates with a distributional constraint to prevent premature over-commitment to a single Rank-1 trajectory. Experiments on mathematical reasoning and Text-to-SQL benchmarks demonstrate that DyJR significantly outperforms GRPO as well as baselines such as RLEP and Ex-GRPO, while maintaining training efficiency comparable to the original GRPO. Furthermore, from the perspective of Rank-$k$ token probability evolution, we show that DyJR improves policy plasticity by reducing over-reliance on Rank-1 tokens and preserving recently reachable alternatives, elucidating how specific sub-modules of DyJR influence the training dynamics.
Dynamic Delayed Tree Expansion For Improved Multi-Path Speculative Decoding
Rahul K Thomas ⋅ Teo Kitanovski ⋅ Micah Goldblum ⋅ Arka Pal
Multi-path speculative decoding accelerates lossless sampling from a target model by using a cheaper draft model to generate a draft tree of tokens, and then applies a verification algorithm that accepts a subset of these. While prior work has proposed various verification algorithms for i.i.d rollouts, their relative performance under matched settings remains unclear. In this work, we firstly present a systematic evaluation of verification strategies across model families, tasks, and sampling regimes, and find that Traversal Verification dominates consistently, with OT-based methods lagging far behind. Our analysis uncovers that this occurs because OT-based methods achieve high multi-token acceptance near the root of the draft tree, while multi-token gains are most impactful deeper in the draft tree, where draft and target distributions diverge. Based on this insight, we propose delayed tree expansion} which drafts a partial single path, delaying the i.i.d. branching point. We show that delayed tree expansion preserves the target distribution and improves on root-node i.i.d rollouts. Further, we develop a dynamic neural selector that estimates the expected block efficiency of OT-based verification methods from draft and target features, enabling context-dependent expansion decisions. Our neural selector allows OT-based methods like SpecInfer to outperform Traversal Verification for the first time, with 5\% higher average throughput across a wide range of models, datasets, and sampling settings. Finally, we extend our approach to Traversal Verification to improve its average throughput by 26\%.
Dynamic Query as Budgeting for Tiny Object Detection
Yi Zhang ⋅ Di Xiong ⋅ Ge Gao ⋅ Xiangyue Zhang ⋅ yihang qiu ⋅ Yuxuan Zhou ⋅ Shuo Chen
Recent DETR-based tiny object detectors adopt dynamic-query mechanisms to handle density imbalance. However, existing designs entangle two effects: how many and which queries are decoded. This entanglement obscures what actually drives the gains, thereby hindering principled dynamic-query design; additional learned predictors or heuristic filtering steps also introduce extra overhead and inevitable decision errors, which can undermine practicality and final gains. In this paper, we disentangle these effects and show that dynamic query primarily acts as budgeting---allocating decoder query capacity according to object density. Motivated by this view, we formulate a monotonic and conservative budgeting principle and propose a budgeting-based DETR (BUTR), a streamlined framework that implements dynamic query by using a lightweight, training-free budgeter to estimate an input-dependent query budget $K(x)$ from encoder confidences. BUTR applies standard Top-$K(x)$ query initialization, removing heuristic filtering, learned components, and complex selection logic. On AI-TOD-v2 and VisDrone, BUTR shows clearer density-aligned query allocation behavior than prior dynamic-query detectors and improves AP by more than 0.5, while using about half as many queries with fewer parameters and GFLOPs, requiring over 2$\times$ less training time, and delivering 4$\times$ faster inference. Additional COCO results further indicate its applicability across a broad spectrum of detection scenarios.
DynamicRad: Content-Adaptive Sparse Attention for Long Video Diffusion
Yongji Long ⋅ Shijun Liang ⋅ Jintao Li ⋅ Yun Li
Leveraging the natural spatiotemporal energy decay in video diffusion offers a path to efficiency, yet relying solely on rigid static masks risks losing critical long-range information in complex dynamics. To address this issue, we propose DynamicRad, a sparse-attention framework that constrains adaptive selection using a radial locality prior. DynamicRad introduces a dual-mode strategy: static-ratio for speed-optimized execution and dynamic-threshold for quality-first filtering. To avoid online search over sparse indices, we integrate an offline Bayesian Optimization (BO) pipeline with a semantic motion router. The router maps prompt embeddings to BO-selected sparsity regimes with a single projection module. Unlike online profiling methods, our offline BO optimizes attention reconstruction error (MSE) on a proxy task and reuses the selected configurations during inference. Experiments on HunyuanVideo and Wan2.1-14B demonstrate that DynamicRad achieves a strong efficiency-quality trade-off among FlashAttention-compatible sparse-attention baselines, obtaining 1.7x-2.5x inference speedups with over 80% effective sparsity. In some long-sequence settings, the dynamic mode matches or improves over the dense baseline on automatic quality metrics, while optional Mask-Aware LoRA further improves long-horizon coherence. Code is available at https://anonymous.4open.science/r/DynamicRad-5D55/.
Dynamics-Informed Adaptive Offline RL for Real-Time Tokamak Plasma Control
Rohit Sonker ⋅ Xiaoyan Hu ⋅ Hiro J Kaga ⋅ Andrew Rothstein ⋅ Jiayu Chen ⋅ Egemen Kolemen ⋅ Jeff Schneider
Tracking plasma profiles, such as electron temperature and rotation, remains a challenging task in tokamak nuclear fusion. While RL-based approaches offer a promising path for multi-input-multi-output control systems, existing frameworks overlook two key obstacles: partial observability of plasma dynamics and the diagnostic mismatch between offline training and online execution. We propose a novel offline RL framework to address these challenges. Specifically, we train the control policy to condition on the latent context that is extracted from a trained plasma dynamics model. This dynamics-informed latent context encodes useful information about operating regimes, enabling the policy to adapt as the plasma evolves. To mitigate the gap between offline and real-time observation, we learn a converter that maps plasma diagnostics reconstructed offline to their real-time-observable counterparts, thereby exposing policy training to the real-time plasma diagnostics available in live experiments. In offline evaluations, the proposed method substantially improves tracking performance and robustness across a diverse set of discharges. We report real-time deployment results on DIII-D Tokamak Fusion Facility, showing feasibility under challenging actuator conditions.
DynEdit: Dynamic Entropy-Guided Sequential Editing for Large Language Models
Jinhu Fu ⋅ Yan Bai ⋅ Yihang Lou ⋅ Li Sun ⋅ Jiahong Wu ⋅ Xiangxiang Chu ⋅ Sen Su
Knowledge editing aims to update outdated or incorrect knowledge in large language models without full retraining. Existing parameter-editing methods remain limited for long-sequence knowledge: single-point editors provide insufficient coverage for extended targets, while fixed chunk-based editors distribute editing effort uniformly across rigid predefined windows, failing to prioritize difficult, information-dense tokens. In this work, we show that token entropy is closely associated with both editing difficulty and semantic importance: high-entropy regions are harder to edit and often contain key factual content. Motivated by this observation, we propose DynEdit, a dynamic framework for long-sequence knowledge editing. DynEdit uses token-level entropy to guide non-uniform sliding, increasing the coverage of difficult high-entropy regions, and further applies entropy-guided token-level optimization to strengthen updates on information-dense tokens within each editing window. Experiments on three benchmarks and two backbone models show that DynEdit consistently outperforms competitive baselines across diverse long-form editing settings, achieving gains of up to +14.0 BLEU points on challenging paraphrase queries and near-saturated semantic scores on several diverse knowledge domains.
Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model
Hoeun Lee ⋅ Hyeonggeun KIM ⋅ Jaeik Kim ⋅ Jusang Oh ⋅ Geon Choi ⋅ Jinhyeok Kim ⋅ Jaeyoung Do
Generalist robot policies must ground language in visual observations, anticipate how actions change the scene, and reason about task goals. Existing VLA approaches typically address these capabilities separately: VLM-based policies emphasize semantic grounding, video or world models emphasize visual dynamics, and goal-conditioned methods emphasize future-state reasoning. This separation can make policies brittle under instruction shift and environment or dynamics perturbations. We introduce Dynin-Robotics, an omnimodal masked-diffusion VLA model that represents robot trajectories as partially observed multimodal token sequences. A single denoising backbone is trained to unify action generation, future-state and world modeling, trajectory-to-language goal understanding, and instruction-conditioned goal-state prediction. At inference time, the same model performs policy generation through block-wise action denoising, with optional goal-state prediction, and world-action joint decoding. We first present diagnostic analyses showing that VLM-based and video/world-model-based policies exhibit complementary failure modes, motivating a unified formulation. Our experiments show that Dynin-Robotics achieves competitive standard-task performance on LIBERO, while improving robustness on instruction-shift and perturbation settings in LIBERO-Plus and VLABench.
DySurface: Consistent 4D Surface Reconstruction via Bridging Explicit Gaussians and Implicit Functions
Minje Kim ⋅ Younghyun Noh ⋅ Jaesoon Kim ⋅ Tae-Kyun Kim
While novel view synthesis (NVS) for dynamic scenes has seen significant progress, reconstructing temporally consistent geometric surfaces remains a challenge. Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) offer powerful dynamic scene rendering capabilities; however, relying solely on photometric optimization often leads to geometric ambiguities. This results in discontinuous surfaces, severe artifacts, and broken surfaces over time. To address these limitations, we present DySurface, a novel framework that bridges the effectiveness of explicit Gaussians with the geometric fidelity of implicit Signed Distance Functions (SDFs) in dynamic scenes. Our approach tackles the structural discrepancy between the forward deformation of 3DGS ($canonical \rightarrow dynamic$) and the backward deformation required for volumetric SDF rendering ($dynamic \rightarrow canonical$). Specifically, we propose the VoxGS-DSDF branch that leverages deformed Gaussians to construct a dynamic sparse voxel grid, providing explicit geometric guidance to the implicit SDF field. This explicit anchoring effectively regularizes the volumetric rendering process, significantly improving surface reconstruction quality, containing watertight boundaries and detailed representations. Quantitative and qualitative experiments demonstrate that DySurface significantly outperforms state-of-the-art baselines in geometric accuracy metrics while maintaining competitive rendering performance. Codes will be publicly available.
Early Semantic Grounding in Image Editing Models for Zero-Shot Referring Image Segmentation
Jingxuan He ⋅ Xiyu Wang ⋅ Yunke Wang ⋅ Mengyu Zheng ⋅ Chang Xu
Instruction-based image editing (IIE) models have recently demonstrated strong capability in modifying specific image regions according to natural language instructions, which implicitly requires identifying where an edit should be applied. This indicates that such models inherently perform language-conditioned visual semantic grounding. In this work, we investigate whether this implicit grounding can be leveraged for zero-shot referring image segmentation (RIS), a task that requires pixel-level localization of objects described by natural language expressions. Through systematic analysis, we reveal that strong foreground-background separability emerges in the internal representations of these models at the earliest denoising timestep, well before any visible image transformation occurs. Building on this insight, we propose a training-free framework that repurposes pretrained image editing models for RIS by exploiting their intermediate representations. Our approach decomposes localization into two complementary components: attention-based spatial priors that estimate where to focus, and feature-based semantic discrimination that determines what to segment. By leveraging feature-space separability, the framework produces accurate segmentation masks using only a single denoising step, without requiring full image synthesis. Extensive experiments on RefCOCO, RefCOCO+, and RefCOCOg demonstrate that our method achieves superior performance over existing zero-shot baselines.
ECG-Reasoning-Benchmark: A Benchmark for Evaluating Clinical Reasoning Capabilities in ECG Interpretation
Jungwoo Oh ⋅ Hyunseung Chung ⋅ Junhee Lee ⋅ Min-Gyu Kim ⋅ Hangyul Yoon ⋅ Ki Seong Lee ⋅ Youngchae Lee ⋅ Muhan Yeo ⋅ Edward Choi
While Multimodal Large Language Models (MLLMs) show promising performance in automated electrocardiogram interpretation, it remains unclear whether they perform actual step-by-step reasoning or just rely on superficial visual cues. To investigate this, we introduce ECG-Reasoning-Benchmark, a novel multi-turn evaluation framework comprising over 6,400 samples to systematically assess step-by-step reasoning across 17 core ECG diagnoses. Our comprehensive evaluation of state-of-the-art models reveals a systematic shortfall in executing multi-step logical deduction. Although models possess the medical knowledge to retrieve clinical criteria for a diagnosis, they exhibit consistently low success rates in maintaining a complete reasoning chain, primarily failing to ground the corresponding ECG findings to the actual visual evidence in the ECG signal. These results demonstrate that current MLLMs bypass actual visual interpretation, highlighting a limitation in their visual grounding capabilities and motivating robust, reasoning-centric medical AI. The code and data are available at https://anonymous.4open.science/r/ecg-reasoning-benchmark-anonymized-BB3C.
Echoes in Filter Bubble: Diagnosing and Curing Popularity Bias in Generative Recommender Systems
Jun YIN ⋅ Bangguo Zhu ⋅ Peng Huo ⋅ Ruochen Liu ⋅ Hao Chen ⋅ Senzhang Wang ⋅ Shirui Pan ⋅ Chengqi Zhang
Recently, Generative Recommenders (GRs), characterized by a unified end-to-end framework, have exhibited astonishing potential in transforming the recommendation paradigm. Despite their effectiveness, we recognize that GRs are still susceptible to the long-standing issue of popularity bias that has pervaded the recommendation community. Although a few studies have attempted to extend traditional debiasing methods to GRs, their effectiveness is marginal, and the fundamental reason why GRs suffer from popularity bias remains under-explored. To bridge this gap, this study focuses on two core aspects in GRs: the optimization of generative framework and the item tokenization based on semantic index. Based on theoretical analyses, we identify that the severe popularity bias emerges from the confluence of a token-level optimization flaw and the undifferentiated property of item tokenization. Accordingly, this study develops a novel generative recommender system, called Ghost, by designing the asymmetric unlikelihood optimization and the skeleton-founded tokenization. Extensive empirical evaluations across three datasets, alongside multiple SOTA baselines, reveal that Ghost substantially alleviates popularity bias and promotes fairer recommendations, while incurring slight degradation to the overall recommendation utility.
Edit-R2: Context-Aware Reinforcement Learning for Multi-Turn Image Editing
Yuxiao YE ⋅ Haoran He ⋅ Fangyuan Kong ⋅ Xintao Wang ⋅ Pengfei Wan ⋅ Kun Gai ⋅ Ling Pan
Text-guided image editing has advanced rapidly with diffusion models and unified multimodal foundation models. However, most existing methods remain confined to single-turn settings, overlooking the more realistic scenario of multi-turn in-context editing, where users iteratively refine an image through a sequence of instructions. In this setting, a model must follow each new instruction while preserving accumulated session-level constraints, challenged by two coupled failure modes: long-context dilution, where sparse textual constraints become difficult to recover from growing interleaved image-text histories, and state contamination, where earlier editing mistakes degrade subsequent generations. We introduce Edit-R2, a novel reinforcement learning post-training framework for unified multimodal models. Edit-R2 reconstructs the operative session intent, which effectively consolidates scattered historical constraints into an explicit reasoning trace before each editing turn. It further enables multi-turn RL over both reasoning and generation through a unified objective that jointly optimizes intent reconstruction generation in discrete text space and flow-matching image generation in continuous latent space, while a trajectory filtering mechanism suppresses corrupted rollouts to stabilize training under state contamination. To support systematic evaluation, we introduce MICE-Bench, a large-scale benchmark for multi-turn in-context editing with automated metrics for instruction following (IF), content consistency (CC), and global awareness (GA) over accumulated session constraints. Experiments show that Edit-R2 substantially improves multi-turn in-context editing and achieves competitive performance compared against strong baselines.
Efficient Algorithms for Distributed Saddle Problems
Ruichen Luo ⋅ Anton Rodomanov ⋅ Sebastian Stich
The distributed setting for Saddle Problems (SPs) has recently emerged as a framework for modern applications in machine learning and multiagent systems. Despite its relevance, the theoretical foundations of this setting have not yet been thoroughly established. In this paper, we advance this research direction by formalizing the distributed setup for SPs and providing rigorous definitions of communication and oracle costs. Further, we prove lower bounds for any distributed gradient-span algorithm, which reveals the gap from existing methods and this theoretical limit. To this end, we provide a Decoupled Method built upon a novel multi-stage reduction that reduces the SP into a sequence of decoupled minimization tasks of residual norms. Our algorithm matches the communication lower bound, thus setting the communication complexity within the gradient-span algorithms. Moreover, it yields the first strict improvement over the long-standing oracle cost of the Extragradient method for general SPs. Finally, we study the extension of distributed SP into Variational Inequality Problem (VIP), which generalizes two-player zero-sum games to multiplayer general-sum games. We show that our Decoupled Method achieves a new state-of-the-art communication complexity for this broader class.
Efficient Diffusion Policy Fine Tuning with Latent Noise Representation Bridging
Yuanchu Liang ⋅ Yue Yang ⋅ An-Chi He ⋅ Hanna Kurniawati
Pre-trained diffusion and flow policies have emerged as powerful backbones for robotics, yet adapting them to novel tasks via reinforcement learning (RL) remains computationally unstable, sample inefficient and lacks theoretical guarantees. Prior approaches that fine-tune the latent noise space bypass costly backpropagation through time but struggle with "noise aliasing," where determining the optimal noise input to the pre-trained generative model requires repeatedly refitting separate Q-value networks. This paper resolves these challenges by introducing a novel theoretical framework demonstrating that imposing a linear structure on the environment induces a Latent Noise Linear Markov Decision Process (LNL-MDP). By treating the pre-trained model as a virtual dataset, our LNL-MDP fine tune framework falls under the FineTuneRL theory with near-optimal guarantee. To translate this idealized theory into a practical methodology, we derive the Bridge Equation which eliminates the noise aliasing problem. Leveraging this insight, we propose Latent Representation Bridging for Diffusion (LARBRID), an algorithm that efficiently learns the representation space from the LNL-MDP and optimal noises to steer the pre-trained model to act optimally in given tasks. We demonstrate that LARBRID is exceptionally sample efficient, yielding up to $2\text{-}3\times$ performance increases on challenging control problems across a diverse suite of robotic locomotion and manipulation tasks.
Efficient evaluation and error pattern discovery for blackbox AI systems
Maxim Rabinovich ⋅ Harvineet Singh ⋅ Aman Sinha
Reliable AI evaluation is a prerequisite for measuring progress and for deployment in safety-critical settings. Unfortunately, even gathering frontier system outputs on a full suite of relevant benchmarks can be expensive, let alone adding expert review or running evaluation inside of an optimization loop. Existing approaches to sample-efficient evaluation address this concern only partially: they typically assume abundant calibration data, pretrained error-relevant embeddings, or a fixed benchmark. Meanwhile, optimization loops require repeated evaluation and the ability to surface a long tail of errors; in fact, in real world settings, the long tail of errors often drives human-in-the-loop iterative system improvement. Working in the regime in which the only available information is system inputs and one or more cheap surrogate measures of system error, we derive an optimal importance sampler for error estimation and analyze its theoretical properties. Going further, we show how the same sampler can naturally discover and rank error patterns. On standard benchmarks for frontier model question answering (MMLU-Pro) and agentic coding (TerminalBench 2.0) we show our method consistently outperforms standard Monte Carlo on estimating aggregate error, recovering failure modes and calculating their prevalence, and detecting poison agent/environment injections.
Structured sparsity extends classical sparse learning by encouraging sparsity at the level of predefined groups rather than individual parameters. This paradigm is widely adopted in modern machine learning and is typically enforced by sparsity-inducing norms. In real pipelines, the optimal group partition is often unknown a priori and does not remain static after initial training. Parameter groupings are often modified as domain knowledge or upstream representation evolves. Standard approaches require retraining the model from scratch under the new target partition, which discards the optimization effort already invested and becomes computationally prohibitive at scale. In this work, we focus on the group-wise $\ell_{p,1}$-norm, a flexible regularization family that encompasses classical group Lasso and other sparse variants. We develop a principled algorithmic framework that leverages the pre-trained optimum to fine-tune model weights toward a new desired group structure without full retraining. Technically, we decompose arbitrary group-wise transitions into a sequence of tractable primitives and derive the closed-form learning dynamics for each sub-action. Extensive experiments on real-world datasets demonstrate that our method achieves significant acceleration while provably recovering the correct optimal solution.
Efficient Gradient-Aware Asynchronous Reinforcement Learning for LLM Post-Training
Songhan Yang ⋅ Jie Wang ⋅ Yinqi Bai ⋅ Tong Xialiang ⋅ Jianye Hao ⋅ Mingxuan Yuan ⋅ Feng Wu
Asynchronous reinforcement learning (RL) is an effective paradigm for improving the training efficiency of large language model (LLM) post-training by decoupling rollout generation from policy optimization. However, this decoupling introduces stale off-policy samples generated by earlier behavior policies, leading to distribution shift and unstable policy updates. Recent replay selection methods such as D-ARL mitigate this issue by selecting variance-aware samples, but require expensive current-policy evaluation over replay responses and mainly rely on variance reduction or current-policy matching as the selection criterion. To address these limitations, we propose GA-ARL, an efficient Gradient-Aware Asynchronous Reinforcement Learning framework for LLM post-training. GA-ARL derives an analytic target distribution from the KL-regularized RL objective and introduces an advantage-weighted variant to account for both policy preference and gradient strength. During training, GA-ARL maintains a replay buffer containing samples from recent behavior policies and selects low-discrepancy samples according to the advantage-weighted analytic target, avoiding additional current-policy evaluation during replay scoring. GA-ARL outperforms SOTA asynchronous methods across mathematical, logical, and code reasoning benchmarks, achieving the best average accuracy on both Qwen3-1.7B and Qwen3-4B, with improvements of up to 6.1%. Meanwhile, compared with the SOTA D-ARL, GA-ARL reduces wall-clock training time by 25.3% on average by avoiding additional current-policy evaluation during replay selection.
Efficient SAM 3 Adaptation for Multi-Class Semantic Segmentation via Dense Competitive Representations
Wenbin Liao ⋅ Hao Zhu ⋅ Yike Ma ⋅ Hao Jiang ⋅ Feng Dai
The Segment Anything Model 3 (SAM 3) has achieved significant progress in *Promptable Concept Segmentation (PCS)* by processing short noun phrases to generate segmentation masks with unique instance identifiers. While SAM3 excels at single-concept segmentation, it falls short in complex, real-world environments where multiple concepts must be segmented simultaneously. In particular, SAM3 typically fails to simultaneously process concepts with semantic overlap, leading to semantic misclassifications and redundant mask predictions. Moreover, as the number of concepts increases, the independent inference paradigm introduces prohibitive inference latency. To address these limitations, we propose the **D**ense **C**ompetitive **R**epresentations **SAM** (**DCR-SAM**) framework. By decoupling cross-modal interactions, the proposed method alleviates the inference latency. To resolve semantic misclassifications, DCR-SAM dynamically aligns textual concepts with visual representations and incorporates a dense competition head to enforce explicit inter-class competition. Furthermore, a learnable background token is applied to absorb non-target objects dynamically. These mechanisms effectively suppress overlapping predictions and generate high-fidelity masks. Extensive experimental evaluations across multiple benchmarks demonstrate that the proposed framework: I) achieves a 2.3\% $\sim$ 10.7\% performance improvement over state-of-the-art methods, II) delivers a 33$\times$ inference speedup compared to the native SAM 3, and III) exhibits strong robustness across diverse scenarios. All code will be released.
Efficient Scaling of LLM Training with Flexible Context Parallelism
Yifan Niu ⋅ Han Xiao ⋅ Dongyi Liu ⋅ Wei zhou ⋅ Jia Li
Scaling long-context capabilities is crucial for Large Language Models (LLMs). However, real-world data contain a large number of sequences with heterogeneous lengths. Existing training libraries for LLMs rely on static parallelism strategies, which suffer from severe load imbalance, redundant communication, and suboptimal hardware utilization under data heterogeneity. In this work, we propose Flexible Context Parallelism (FCP), an efficient parallelism strategy that adaptively reconfigures communication groups and context parallelism degrees during LLM training. We generalize more flexible non-power-of-two parallelism degrees and develop a polynomial-time algorithm to generate near-optimal parallelism strategies with only millisecond-level overhead per training batch. FCP is able to maintain high hardware efficiency even under extreme data heterogeneity. Experimental results demonstrate that FCP significantly outperforms Megatron-LM and DeepSpeed in both LLM and MLLM training, achieving up to 1.47 $\times$ speedup in average throughput while maintaining near-linear scaling efficiency across large-scale clusters. For extremely unbalanced batches, FCP even achieves 2.24 $\times$ speedup.
Efficient Streaming Audio-Visual Target Speaker Extraction for Real-World Acoustic Scenes
Wendi Sang ⋅ Kai Li ⋅ Yifan Li ⋅ Jianqiang Huang ⋅ Xiaolin Hu
Streaming audio-visual target speaker extraction (AVTSE) is essential for latency-sensitive applications. However, current causal systems face two coupled gaps. First, standard benchmarks fix the number of speakers at two and ignore the near-/far-field interference, reverberation, and sudden noise of real scenes. Second, high-performance causal architectures are too heavy for edge deployment. Lightweight alternatives try to reclaim capacity by repeatedly invoking a small separator, which offsets the savings. To address the first gap, we release RealSSA, a realistic AVTSE benchmark consisting of RealSSA-Sim and RealSSA-Real. RealSSA-Sim provides controllable simulated scenes with 2-6 dynamic near-/far-field speakers across four scene types, while RealSSA-Real provides a held-out real-recorded evaluation set. To address the second gap, we propose Falcon, a causal AVTSE method. Its Separator performs multi-scale time-frequency separation in a single encoder--decoder pass. We train Falcon with an encoder-space speaker contrastive loss. This loss suppresses near-field leakage at zero inference cost and transfers across backbones. Falcon achieves state-of-the-art extraction quality on RealSSA-Sim, LRS2-2Mix, LRS3-2Mix, and VoxCeleb2-2Mix. Compared with the causal AV-TFGridNet baseline, it cuts parameters by 91.9\%, computation by 20.2$\times$, and GPU inference latency by 11.1$\times$. Code and demo are available at https://submission-falcon.github.io/demo/.
ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models
Ye Li ⋅ Huanan Liu ⋅ Kangye Ji ⋅ Yuan Meng ⋅ Jiajun Fan ⋅ Yuansong Wang ⋅ Shiyu Qin ⋅ Chenglei Wu ⋅ Shu-Tao Xia ⋅ Zhi Wang
Vision-Language-Action (VLA) models have emerged as a powerful paradigm for generalist robotic control. However, their high computational cost and limited control frequency hinder real-time robotic manipulation, especially when large vision-language backbones and iterative action heads are executed at every control step. Existing VLA acceleration methods often optimize individual components or rely on fixed acceleration rules, treating different control steps with largely fixed computation and overlooking the non-uniform reasoning demands of sequential embodied control. Inspired by human motor control, where cognitive and feedback resources concentrate on goal-sensitive stages, we argue that VLA models should learn when to invest full computation and when to reuse prior computation. To this end, we propose ElegantVLA, a plug-in phase-adaptive inference framework that accelerates VLA models through intra-model dynamic compute scheduling. ElegantVLA introduces a lightweight scheduler that observes temporal representation similarity, robot-motion cues, and episode progress to jointly allocate computation across the vision encoder, LLM, and action head. For perception-language reasoning, the scheduler selects a five-level Vision--LLM compute mode, from full recomputation to multi-step temporal reuse, based on visual-language representation stability. For action generation, it selects a three-level denoising mode, reusing intermediate denoising states during stable motion while preserving full refinement for goal-sensitive stages. By coordinating these decisions, ElegantVLA provides a general acceleration framework for modern VLA pipelines with explicit action-generation modules, without modifying or retraining the base model. Extensive experiments on GR00T, CogACT, and real-world tasks show that ElegantVLA preserves or improves task success while substantially accelerating inference. On GR00T, it achieves up to 2.55$\times$ average speedup. On CogACT, it delivers a 3.77$\times$ average speedup. In GR00T-based real-world experiments across six tasks, it reduces computation by 2.18$\times$ and increases control frequency from 13.8 Hz to 26.3 Hz.
Embodied Neurocomputation: A Framework for Interfacing Biological Neural Cultures with Scaled Task-Driven Validation
Johnson Zhou ⋅ Daniel Tanneberg ⋅ Forough Habibollahi ⋅ Alon Loeffler ⋅ Kiaran Lawson ⋅ Valentina Baccetti ⋅ Kwaku D Abu-Bonsrah ⋅ Candice Desouza ⋅ Finn Doensen ⋅ Bradley Watmuff ⋅ Daria Kornienko ⋅ Azin Azadi ⋅ Justin L Bourke ⋅ Bernhard Sendhoff ⋅ Brett J. Kagan
Biological neural networks (BNNs) have been established as a powerful and adaptive substrate that offer the potential for incredibly energy and data efficient information processing with distinct learning mechanisms. Yet a core challenge to utilizing BNN for neurocomputation is determining the optimal encoding and decoding mechanisms between the traditional silicon computing interface and the living biology. Here, we propose an Embodied Neurocomputation framework as a systems-level approach to this multi-variable optimization encoding/decoding problem. We operationalize this approach through the first large-scale parameter optimization of encoding configurations for a BNN agent performing closed-loop navigation along an odor-style gradient in a simulated grid-world. Despite the relative simplicity of the task, the biological interactions gave rise to a massive multi-combinatorial search space for optimal parameters. By considering how the components of the system are interconnected and parameterized, we evaluated approximately 1,300 parameter combinations, over 4,000 hours of real-time agent-environment interactions, to identify 12 configurations that consistently demonstrated learning across multiple episodes. These configurations achieved significantly higher task performances than optimized silicon-based DQN agents under the same interaction budget. These findings represent an initial step toward robust and scalable goal-oriented learning using BNNs. Our framework establishes a foundation for applying task-driven neurocomputing and supports the development of field-wide benchmarks. In the long term, this work supports the development of hybrid bio-silicon architectures capable of efficient, adaptive and real-time computation, including the potential for robotic control applications.
Recent advances in 3D object generation have enabled the creation of high-fidelity 3D assets from 2D images. This 2D-to-3D capability offers a promising path for alleviating 3D data scarcity, which remains a key bottleneck for scalable simulation in robotics and AR/VR applications. However, existing 3D generation methods primarily emphasize visual plausibility and geometric fidelity, with limited consideration of object-level interaction and structural understanding that are critical for downstream embodied AI tasks. To bridge this gap, we introduce EmbodiedObject, a unified framework for embodied 3D object understanding built upon latent representations produced by modern 2D-to-3D generative models such as SAM3D. We first systematically analyze the part-level locality of these latent representations against existing unified 3D point encoders and show that they naturally preserve rich structural semantics suitable for object understanding. Motivated by this observation, we then design a unified DiT-style decoder that operates directly on the latent representation and supports multiple object-level tasks, including open-vocabulary 3D affordance prediction, 3D part segmentation, and 3D articulation estimation. Given a real-world 2D image containing an object of interest, EmbodiedObject extends 2D-to-3D generation beyond visual reconstruction by jointly reasoning about affordance, articulation, and part semantics. This transforms generated 3D assets into actionable object representations suitable for downstream embodied AI applications. Our source-code will be open-sourced.
EMERGE: A Benchmark for Updating Knowledge Graphs with Emerging Textual Knowledge
Klim Zaporojets ⋅ Daniel Daza ⋅ Edoardo Barba ⋅ Ira Assent ⋅ Roberto Navigli ⋅ Paul Groth
Knowledge Graphs (KGs) are structured knowledge repositories containing entities and relations between them. In this paper, we study the problem of automatically updating KGs over time in response to evolving knowledge in unstructured textual sources. Addressing this problem requires identifying a wide range of update operations based on the state of an existing KG at a given time and the information extracted from text. This contrasts with traditional information extraction pipelines, which extract knowledge from text independently of the current state of a KG. To address this challenge, we propose a generic and extensible pipeline that pairs textual passages with the KG edit operations they induce on a given KG snapshot. We instantiate this pipeline on the Wikidata knowledge graph and the English Wikipedia corpus to create EMERGE, a dataset of 233K Wikipedia passages associated with a total of 1.19 million KG edits across seven yearly Wikidata snapshots from 2019 to 2025. Our experimental results highlight key challenges in updating KG snapshots based on emerging textual knowledge, particularly in integrating knowledge expressed in text with the existing KG structure. These findings position the dataset as a valuable benchmark for future research. The code and dataset are available at https://github.com/klimzaporojets/emerge.
Enabling VLA Action Self-Verification via VLM Token Probability Bucketing
Chen Zhao ⋅ Zhuoran Wang ⋅ Haoyang Li ⋅ Guanlin Li ⋅ Shifeng Bao ⋅ Youhe Feng ⋅ Yang Li ⋅ Jie Tang ⋅ Jing Zhang
Test-time sampling can improve Vision-Language-Action (VLA) policies, but only if the model can efficiently select a good action from multiple candidates. Prior work often relies on separately trained external verifiers, adding extra models and inference overhead. We propose \textbf{Token Probability Bucketing (TokenPB)}, a self-verification approach that discretizes continuous action chunks into bucket tokens and co-trains a single VLA checkpoint with a joint objective: a flow-matching loss for action generation and an auxiliary bucket loss $\mathcal{L}_{\text{bucket}}$ that teaches the VLM backbone to model observation-conditioned distributions over bucket-token sequences. At inference, TokenPB samples $M$ candidates, scores each via teacher-forced evaluation under the learned distribution, and selects the best---enabling efficient batched scoring with observation-prefix KV-cache reuse. We provide theory connecting the training objective to a monotonic ranking property and bounds on expected selection regret that separate candidate-set effects from calibration. Across multiple backbones and benchmarks, TokenPB yields consistent success-rate improvements over single-sample inference from the same checkpoint (+2.7 in simulation, +17.1 in real-world, +16.4 under out-of-distribution shifts), and further improves over external-verifier baselines under the same sampling budget.
Denoising and score estimation are classically linked through Tweedie’s formula, which relates the posterior mean under Gaussian corruption to the Stein score of the noisy marginal. In this work, we extend this perspective beyond Gaussian noise to a broad class of elliptical, energy-based noise distributions, with particular emphasis on generalized Gaussian corruptions. We derive the Energy–Tweedie identity: when the denoising posterior is viewed through the lens of proper scoring rules, the path derivative of a matched, possibly non-Euclidean energy score recovers the Stein score of the noisy marginal. Thus, the familiar correspondence between Gaussian noise, posterior means, squared loss, and Tweedie’s formula is lifted to a distributional correspondence between generalized Gaussian noise, full posterior laws, Mahalanobis energy scores, and the Energy–Tweedie identity. Among its consequences, this identity gives a posterior-sample-based route to score estimation, yields a principled criterion for noise-parameter calibration, and supplies a score-based perspective on recent diffusion-style generative methods trained with scoring rules.
Enhancing Agentic Code Localization with Traceability Recovered from Repository Evolution
Yiming Liu ⋅ Binhang Qi ⋅ Weiyu Kong ⋅ Jiawei Liu ⋅ Xinxin Shan ⋅ Saijun Gao ⋅ Yun Lin
While agentic frameworks have advanced repository-level code localization, their reliance on static snapshots often leads to a $\textit{traceability gap}$—the loss of implicit links between issue symptoms and code implementation. To bridge this information deficit, we introduce $\textbf{T}$raceability $\textbf{A}$ugmented $\textbf{Co}$de Localization (TACO), a framework that recovers retrievable and usable traceability from repository evolution to guide localization agents. TACO offline crystallizes historical pull requests into a dual-index knowledge base through a prompt auto-tuning technique, capturing high-value semantic hints and architectural rationales. In the online phase, TACO employs a synergistic dual-track retrieval workflow, cross-validated by an arbiter and aligned temporally to handle version drift. Extensive evaluations on SWE-bench Lite and SWE-bench Verified across five state-of-the-art agentic baselines demonstrate that TACO significantly improves exact-match localization accuracy (Acc@1) by an average of 11.29\%, 13.20\%, and 12.87\% at the file, module, and function levels, respectively. Moreover, TACO significantly streamlines the exploration process, achieving an average reduction of 29.2\% in token usage and 30.4\% in monetary costs, with minimal offline maintenance overhead.
EnterpriseBench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance
Thomson Yen ⋅ Julian Poeltl ⋅ Harshith S Gear ⋅ Yilin Meng ⋅ Joshua Fan ⋅ Adam Shen ⋅ Yili Liu ⋅ Ali Bauyrzhan ⋅ Siri Du ⋅ Haoyang Liu ⋅ C. Guetta ⋅ Hongseok Namkoong
LLM agents are increasingly expected to carry out end-to-end workflows, producing complete artifacts from high-level user instructions. To meet enterprise needs, frontier AI labs have developed agents that can construct entire spreadsheets from scratch. This capability is especially relevant in finance, where core workflows such as financial modeling, forecasting, and scenario analysis are commonly conducted through spreadsheets. Yet, existing spreadsheet benchmarks do not measure this advanced capability, focusing instead on question-answering or single-formula edits. To address this gap, we provide one of the first evaluations of agents on end-to-end spreadsheet tasks, focusing on economically critical financial workflows such as modeling and scenario analysis. Since deliverables therein are routinely reviewed and revised by multiple stakeholders, judging their quality necessarily involves high-level criteria such as readability or ease of modification. To reflect the multidimensional nature of solution quality, we develop an evaluation taxonomy comprising three dimensions: Accuracy, Formula, and Format, each comprising fine-grained criteria that reflect professional standards. The Claude family leads the benchmark and produces the most professional-looking outputs in our qualitative review, but even the strongest agents frequently fall short of professional finance standards and degrade sharply as the difficulty increases beyond a few chained calculations. This suggests that current agents are not yet able to reliably produce professional-quality spreadsheets at the level of complexity real-world workflows demand.
EPIC: Efficient Predicate-Guided Inference-Time Control for Compositional Text-to-Image Generation
Sunung Mun ⋅ Sunghyun Cho ⋅ Jungseul Ok
Recent text-to-image (T2I) generators can synthesize realistic images, but still struggle with compositional prompts involving multiple objects, counts, attributes, and relations. We introduce EPIC (Efficient Predicate-Guided Inference-Time Control), a training-free inference-time refinement framework for compositional T2I generation. EPIC casts refinement as predicate-guided search: it parses the original prompt once into a fixed visual program of object variables and typed predicates, covering checkable conditions such as object presence, counts, attributes, and relations. Each generated or edited image is verified against this program using visual evidence extracted from that image. An image is judged to satisfy the prompt only when all predicates are satisfied; otherwise, failed predicates decide the next step, routing local failures to targeted editing and global failures to resampling while the fixed visual program remains unchanged. On GenEval2, EPIC improves prompt-level accuracy from 34.16% for single-pass generation with the base generator to 71.46%. Under the same generator/editor setting and maximum image-model execution budget, EPIC outperforms the strongest prior refinement baseline by 19.23 points while reducing realized cost by 31% in image-model executions, 72% in MLLM calls, and 81% in MLLM tokens per prompt.
Epigenomics-Guided Flow Matching for 3D Genome Super-Resolution
Yubiao Zhao ⋅ Shicheng Song ⋅ Yusen Ye ⋅ Huan Liu ⋅ Juan Liu ⋅ Lihua Zhang
Mapping fine-scale 3D chromatin landscapes is critical for understanding gene regulation, yet the resolution of Hi-C data remains fundamentally constrained by both physical bottlenecks and sequencing costs. Although deep learning has shown promise in super-resolving Hi-C data, existing methods suffer from absolute coordinate loss during localized patching and a lack of structural and epigenomic priors in unconditioned generation. We propose a novel epigenomics-guided flow matching framework EpiFlow for 3D genome super-resolution. EpiFlow introduces two key innovations: (1) an EpiCond module that deterministically maps 1D epigenomic features to 2D spatial constraints via orthogonal expansion with built-in positional grounding, avoiding quadratic attention complexity; and (2) a Toeplitz-regularized classifier-free guidance mechanism that enforces distance-decay priors using a learnable symmetric prototype matrix. Extensive evaluation shows that our method achieves state-of-the-art cross-cell-type loop detection, improving recall by $2.8\times$ over raw Hi-C on unseen H1-hESC (Loop $F_1$ of $0.325$ vs. $0.199$) and outperforming all baselines on same-cell HFFc6 ($F_1=0.520$). Aggregate Peak Analysis (APA) confirms biological validity, and strong generalization to a held-out cell line demonstrates that the model learns transferable principles of 3D genome folding rather than cell-specific biases.
eSAM: Editing SAM3 Attention for Training-Free Referring Segmentation
Yaoting Wang ⋅ Yun Zhou ⋅ Hengrui Hu ⋅ Chang Liu ⋅ Henghui Ding
Referring expression segmentation (RES) aims to produce a pixel-level mask for the object described by a free-form natural language expression, in either images (RIS) or videos (RVOS). Existing zero-shot approaches reduce the cost of mask-language annotation but still rely on auxiliary training stages or multi-model pipelines. The recently released SAM 3 introduces native text-prompted segmentation within a single foundation model. However, directly applying it to RES generates unexpected weak performance. We trace this gap to an attention sink in SAM 3's multimodal fusion encoder, where the start-of-text token absorbs on average $54.8\%$ of each image patch's cross-attention mass, leaving content tokens with limited influence on the fused visual features. Suppressing the sink alone is insufficient, as the released attention mass spreads without direction; closing the gap requires both suppressing the sink and providing patch-dependent spatial guidance that routes attention to semantically relevant content tokens. Building on these findings, we introduce eSAM, a fully training-free framework that edits SAM~3's cross-attention without any parameter updates. With Attention Mask Editing (AME), eSAM edits the attention map, suppressing the sink while injecting a CLIPSeg-derived spatial prior that directs released attention toward relevant content tokens. Further with Value Tensor Editing (VTE), eSAM rescales the value tensors, amplifying the magnitude with which content tokens contribute to the fused output. On both RIS (RefCOCO/+/g) and RVOS (Ref-DAVIS17, Ref-YouTube-VOS, MeViS) benchmarks, eSAM achieves new state-of-the-art results among training-free methods.
ES-Merging: Biological MLLM Merging via Embedding Space Signals
Wonbin Lee ⋅ Dongki Kim ⋅ Sung Ju Hwang
Biological multimodal large language models (MLLMs) have emerged as powerful foundation models for scientific discovery. However, existing models are specialized to a single modality, limiting their ability to solve inherently cross-modal scientific problems. While model merging is an efficient method to combine the different modalities into a unified MLLM, existing methods rely on input-agnostic parameter space heuristics that fail to faithfully capture modality specialization. To overcome this limitation, we propose the Embedding-Signal-based MLLM Merging (ES-Merging), a framework that estimates merging coefficients from embedding space signals, moving the merging paradigm from the parameter signals to the embedding signals. ES-Merging exploits coarse-grained and fine-grained signals from embedding space to estimate the layer-wise and element-wise merging coefficients, respectively, which are jointly combined for complementary coefficient estimation. Through extensive experiments, we demonstrate that ES-Merging outperforms existing merging methods not only on the cross-modal reasoning but also on the single-modal knowledge preservation, establishing that embedding space signals provide a principled and effective foundation for MLLM merging.
Estimating Model-Level Membership Inference Vulnerability Without Reference Models
Euodia Dodd ⋅ Natasa Krco ⋅ Igor Shilov ⋅ Matthew R Wicker ⋅ Yves-Alexandre de Montjoye
Membership inference attacks (MIAs) have emerged as the standard tool for evaluating the privacy risks of AI models. However, state-of-the-art attacks require training numerous, often computationally expensive, reference models, limiting their practicality. We present a novel approach for estimating model-level vulnerability to the Likelihood Ratio Attack (LiRA), the strongest available attack, directly from the train and test loss distributions of the target model and without training any reference models. We show that LiRA's per-sample signal decomposes into a variance-ratio term and a residual mean-shift term, with the relative contribution of each determined by how much training collapses model uncertainty at the trained sample. This places models on a continuum, with different regimes calling for different reference-free loss-based statistics as proxies for LiRA TPR. The shapes of the loss distributions themselves indicate which proxy applies. We instantiate the framework with two natural proxies. At the heavy-tailed end, the LOSS attack TNR predicts LiRA TPR@FPR=$10^{-3}$ with RMSE 0.03 across 9 image classification architectures and 4 datasets, outperforming low-cost reference-model attacks such as RMIA. At the symmetric end, the LOSS attack AUC predicts LiRA TPR with RMSE 0.01 across five GPT-2 sizes from 10M to 1B parameters. We also show these proxies to outperform both low-cost (few reference models) attacks such as RMIA and other measures of distribution difference.
Evaluating Test-Time Scaling of General LLM Agents
Xiaochuan Li ⋅ Tianshi Ming ⋅ Pranav Setlur ⋅ Abhijay S Paladugu ⋅ Andy Tang ⋅ Hao Kang ⋅ Shuai Shao ⋅ Rong Jin ⋅ Chenyan Xiong
LLM agents are increasingly expected to operate as general-purpose systems that resolve real-world user requests, yet their dynamic scaling behavior in realistic environments remains poorly understood. In this paper, we systematically investigate two principal test-time scaling axes of LLM agents: sequential scaling through extended interaction and parallel scaling through trajectory sampling. We first introduce a realistic benchmark that provides one unified framework for evaluating LLM agents across search, coding, reasoning, and tool-use domains, more faithfully reflecting the heterogeneity of real-world deployments. Evaluating ten leading LLM agents reveals substantial performance degradation when transitioning from domain-specific evaluations to this realistic setting. Building on this foundation, we progressively scale test-time compute along fine-grained increments to characterize the performance upper bound. We find that neither scaling axis can consistently yield meaningful gains from additional test-time compute in realistic environments, a phenomenon we attribute to two fundamental limitations: the scaling plateau that bottlenecks sequential scaling and the verification gap that undermines parallel scaling. Code and data will be released publicly.
Evaluating Uncertainty Calibration in Probabilistic Time Series Foundation Models
Thomas Decker ⋅ Volker Tresp
Time series foundation models (TSFMs) are increasingly used for probabilistic forecasting across diverse domains, yet the reliability of their uncertainty estimates remains poorly understood. In this work, we study uncertainty calibration in probabilistic TSFMs from a broader methodological and empirical perspective. We introduce a taxonomy of how TSFMs represent predictive uncertainty and develop a fine-grained evaluation framework for calibration in multi-horizon forecasting that distinguishes pooled, horizon-wise, and stronger conditional notions of calibration. Using this framework, we conduct a comprehensive empirical evaluation of popular TSFMs across diverse datasets and show that current models exhibit distinct and systematic patterns of miscalibration that standard metrics often obscure. We further analyze common post-hoc recalibration methods and characterize which forms of calibration they can improve and which tend to persist. Our results provide practical guidance for evaluating and improving uncertainty calibration in TSFMs and highlight important limitations of current practice.
EventPrune: Cascaded Event-Assisted Token Pruning for Efficient First-Person Dynamic Spatial Reasoning
Pengtao Ma ⋅ Ziliang Zhou ⋅ Ciyu Ruan ⋅ Haoyang Wang ⋅ kaiyuan Li ⋅ Zihang GONG ⋅ Wenhua Ding ⋅ Chen Gao ⋅ Jingao Xu ⋅ Xinlei Chen
First-person dynamic spatial reasoning requires models to track continuous motion and precise geometric structure, but the quadratic attention cost of Transformer-based Video-LLMs makes dense visual tokens computationally expensive. Existing token pruning paradigms predominantly rely on discrete static snapshots, failing to preserve the physical dynamics essential for reasoning. We propose Event Cascade Pruning (ECP), to our knowledge the first training-free framework that leverages the high-frequency motion cues from event cameras as a continuous event-guided motion prior to guide token selection. ECP combines three stages: Event-Triggered Causal Sampling to anchor motion-informative keyframes, Event-guided Motion Saliency Filtering to remove static or low-motion regions, and Event-Attention Ranking Fusion to calibrate spatial attention with motion salient dynamics. With 80\% visual token reduction, ECP outperforms the full-token baseline (37.62\% vs. 36.31\%) while achieving 1.89× inference speedup and 52\% GFLOPs reduction. We further introduce ESR-Real, the first real-world RGB-event benchmark for first-person spatial reasoning, where ECP improves accuracy by 2.68\% over full-token baselines.
Every Bit, Everywhere, All At Once: A Binomial Multibit LLM Watermark
Thibaud Gloaguen ⋅ Robin Staab ⋅ Mark Vero ⋅ Martin Vechev
With LLM watermarking already being deployed commercially, practical applications increasingly require multibit watermarks that encode more complex payloads, such as user IDs or timestamps, into the generated text. In this work, we propose a fundamentally new approach for multibit watermarking: introducing binomial encoding to directly encode every bit of the payload at every token position. We complement our approach with a stateful encoder that during generation dynamically redirects encoding pressure toward underencoded bits. Our evaluation against 8 baselines on up to 64-bit payloads shows that our scheme achieves superior message accuracy and robustness, with the gap to baseline methods widening in more relevant settings (i.e., large payloads and low-distortion regimes). At the same time, we challenge prior works’ evaluation metrics, highlighting their lack of practical insights, and introduce per-bit confidence scoring as a practically relevant metric for evaluating multibit LLM watermarks.
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding
Geo Ahn ⋅ Jiwook Han ⋅ Youngrae Kim ⋅ Joonseok Lee ⋅ Jinwoo Choi
Fine-tuning MLLMs for Video Temporal Grounding (VTG) often improves in-domain performance but degrades sharply under domain shift. In this work, we find that this failure is primarily driven not just by unseen query concepts, but by visual domain shift, which prevents the model from coupling its learned temporal localization knowledge with its inherent entity-attention capability. To address this, we introduce EVIDENT, a parameter-efficient adaptation framework that anchors temporal grounding in the inherent entity-attention of pre-trained MLLMs by routing VTG adaptation through explicit visual entity evidence. EVIDENT consists of three components: (i) an Entity Bottleneck Adapter that transforms dense visual tokens into compact entity-level slots, (ii) an Entity-Binding Distillation loss that instills objectness priors into the semantically unstructured MLLM visual space, guiding each slot to bind to a coherent entity, and (iii) an Entity-to-eVidence gating mechanism that leverages the captured entities as evidence, steering the model to localize moments containing query-relevant entities. Together, these components enable VTG fine-tuning to rely on entity-grounded evidence rather than brittle dataset shortcuts. Experiments on cross-domain VTG benchmarks show that EVIDENT consistently improves out-of-domain robustness while preserving competitive in-domain performance with modest parameter overhead. These results suggest that entity-level grounding is an effective inductive bias for generalizable temporal localization.
EviSAM3: Evidence-Driven SAM3 for Referring Remote Sensing Image Segmentation
Yuhang Duan ⋅ Xiaoshuai Wu ⋅ Chengqi Zhang ⋅ Jincheng Lu ⋅ Mingwei Zhang ⋅ Yu Liu
Referring Remote Sensing Image Segmentation (RRSIS) aims to segment target objects from remote sensing images based on natural language descriptions, requiring precise vision–language alignment under significant ambiguity. Despite the strong generalization ability of Segment Anything Model 3 (SAM3), the large spatial extent and varying scales of targets in remote sensing images intensify the ambiguity of referring expressions. This poses a major challenge for fine-grained segmentation. To address the challenge, we propose EviSAM3, an evidence-driven parameter-efficient adaptation framework that mimics human decision-making by dynamically accumulating partial evidence. Specifically, EviSAM3 progressively integrates cross-modal evidence through an evidence memory bank, where each incoming piece is aggregated via momentum-based updates. As evidence accumulates, the model performs iterative evidence rectification to resolve ambiguity under incomplete observations. Meanwhile, the contributions of different evidence are adaptively modulated based on their diagnostic relevance, allowing more informative evidence to dominate the reasoning process and guide the final prediction. Extensive experiments demonstrate that EviSAM3 consistently outperforms state-of-the-art methods, particularly in challenging scenarios with high ambiguity and significant scale variation. The code is available in the supplementary material.
EvoGround: Self-Evolving Video Agents for Video Temporal Grounding
MIN JOON JUNG ⋅ Byoung-Tak Zhang ⋅ Lorenzo Torresani
Video temporal grounding (VTG) takes an untrimmed video and a natural-language query as input and localizes the temporal moment that best matches the query. Existing methods rely on large, task-specific datasets requiring costly manual annotation. We introduce EvoGround, a framework of two coupled self-evolving agents, a proposer and a solver, that learn temporal grounding from raw videos without any human-labeled data. The proposer generates query--moment pairs from raw videos, while the solver learns to ground them and feeds back signals that improve the proposer in return. Through this self-reinforcing reinforcement-learning loop, the two agents are initialized from the same backbone and mutually improve across iterations. Trained on 2.5K unlabeled videos, EvoGround matches or surpasses fully supervised models across multiple VTG benchmarks, while emerging as a state-of-the-art fine-grained video captioner without manual labels.
Exact Posterior Score Estimation for Solving Linear Inverse Problems
Abbas Mammadov ⋅ Ozgur Kara ⋅ Kaan Oktay ⋅ Iskander Azangulov ⋅ Adil K Akan ⋅ Hyungjin Chung ⋅ James Rehg ⋅ Yee Whye Teh
Diffusion and flow-based models learn powerful data priors by training a denoiser to reverse Gaussian corruption. To use this prior to solve a linear inverse problem, one needs to sample from the posterior, but the score that the prior provides is the unconditional score, not the posterior score. Existing methods either steer a fixed pretrained denoiser with approximate measurement-matching corrections, or train a conditional restoration model that abandons the denoising structure of the prior. We derive the exact posterior score in closed form for linear Gaussian inverse problems under general Gaussian interpolants, and show that posterior sampling reduces to a denoising problem at an operator-dependent shifted pivot under an anisotropic noise covariance. We turn this identity into Exact Posterior Score (EPS), a denoising training objective that preserves the input/output structure of standard pretraining and can therefore be trained from scratch or fine-tuned from a pretrained denoiser. At inference, EPS uses the same sampler as the underlying backbone, with no likelihood gradients or projections. We evaluate EPS on five linear inverse problems across FFHQ and ImageNet, where it outperforms training-free and training-based baselines on fidelity, perceptual, and distributional metrics, while using roughly an order of magnitude fewer denoiser evaluations than gradient-based posterior samplers.
ExoC2T: An exogenous-driven spatio-temporal learning framework for cross-city transfer
Hailong Yu ⋅ Zhengyang Zhou ⋅ Liwen Zhang ⋅ Qihe Huang ⋅ Kuo Yang ⋅ Yudong Zhang ⋅ Yang Wang
Forecasting spatio-temporal human mobility is essential for urban infrastructure optimization, yet dense traffic sensing remains prohibitively expensive for many cities. Existing cross-city transfer methods and recent spatio-temporal foundation models usually rely on historical target-side traffic sequences as the prediction context, which limits their applicability to under-instrumented, graph-missing, or newly monitored cities. We study exogenous-driven cross-city transfer, where target traffic must be predicted from static and dynamic exogenous urban context without historical flow input. This exogenous-only observability removes the direct endogenous state observation used by flow-driven predictors and turns cross-city forecasting into latent traffic-generating mechanism recovery under city shift. The key challenge is to recover traffic-state information from indirect external evidence while organizing cross-city response patterns into a mechanism space that can adapt to target variation without destabilizing transferable structure. To address this challenge, we propose \textbf{ExoC2T}, a plasticity-inspired spatio-temporal learning framework. ExoC2T uses an Exogenous-Infused Spatio-Temporal Network (\textbf{EXIST}) to build an exogenous mechanism substrate, where semantic relations, temporal regimes, and environment-conditioned responses are organized into a shared prototype geometry. On this substrate, Stability-Plasticity Adaptation separates a stable exogenous-to-flow core from an adaptive city-conditioned deviation, infers a support-conditioned latent posterior from limited target evidence, and controls target risk through invariant discrepancy reduction and posterior correction. The resulting model preserves transferable exogenous-to-flow structure while regulating target-specific modulation through posterior-dependent plasticity. Extensive experiments on a multi-city benchmark show consistent improvements over strong forecasting and transfer baselines, demonstrating the value of exogenous-driven transfer for low-cost and scalable urban computing.
Experience Makes Skillful: Enabling Generalizable Medical Agent Reasoning via Self-Evolving Skill Memory
Haoran Sun ⋅ Wenjie Li ⋅ Yujie Zhang ⋅ Zekai Lin ⋅ Fanrui Zhang ⋅ Kaitao Chen ⋅ Xingqi He ⋅ Yichen Li ⋅ Mianxin Liu ⋅ Lei Liu ⋅ Yankai Jiang
Medical agent systems are increasingly expected to support interactive clinical decision making rather than only static question answering. In such settings, effective agents must reuse prior experience across evolving cases, yet existing memory mechanisms often retain raw historical traces that are redundant, noisy, and difficult to govern. More importantly, they rarely distinguish which memories are truly useful for future reasoning. This limits their ability to accumulate compact and reliable experience for long-horizon clinical reasoning. To close this gap, we propose SkeMex, a post-deployment self-evolution framework that improves medical agents through a skill-based memory without updating model weights. SkeMex distills informative interaction trajectories into structured skills that encode reusable procedural knowledge, and organizes them into a multi-branch repository spanning general, task-specific, and action-level experience. To determine which memories should be reused and retained, SkeMex estimates context-dependent utility from environment feedback and uses it to guide value-aware retrieval and repository governance. A closed-loop ``Read--Write--Assess--Govern" lifecycle further supports continual evolution by writing new skills, updating utilities, promoting useful memories, and removing harmful entries. Experiments across diverse clinical tasks show that SkeMex consistently outperforms representative memory-based agents in both offline and online settings. It also generalizes across model backbones and supports transferable skill memory. All data and code will be released publicly.
Expert-guided Bayesian optimization for sustainable protein formulation
Anna Thomas ⋅ Georgios Zaverdinos ⋅ Petros Mandalis ⋅ Andreas Orfanoudakis ⋅ Akihiro Takino ⋅ Sohum Patnaik ⋅ David J Kraus ⋅ Panos Kostopoulos ⋅ Caroline Cotto ⋅ Nikos Tsiaparas ⋅ Dan Jurafsky ⋅ Aadit Patel
Diversifying protein sources away from animal agriculture is critical for climate change mitigation and food security, but formulating sustainable protein sources that match or improve upon the properties of their animal-based counterparts remains an expensive trial-and-error process. We frame this challenge as high-dimensional black-box optimization with sparse approximate solutions and formalize Expert-Guided Bayesian Optimization (EGBO), in which an expert, e.g. a human or LLM, selects a low-dimensional subspace for BO and may adaptively expand it over time. We decompose EGBO's suboptimality into a selection gap and an optimization gap, and characterize the coverage–dimension tradeoff governing when expert guidance helps. To support in silico prototyping before costly real-world deployment, we introduce FormulateBench, a suite of 24 plant-based formulation tasks, on which LLM-guided EGBO outperforms all tested baselines. When deployed to optimize two plant-based dairy products, EGBO improves utility, as assessed by a trained human panel, by 29\% and 26\% in 10 iterations each. In a comparison with a professional human food scientist given the same time budget, EGBO achieved near-perfect utility of 0.992, vs. 0.850 for the food scientist.
Explaining and Preventing Alignment Collapse in Iterative RLHF
Etienne Gauthier ⋅ Francis Bach ⋅ Michael Jordan
Reinforcement learning from human feedback (RLHF) typically assumes a static or non-strategic reward model (RM). In iterative deployment, however, the policy generates the data on which the RM is retrained, creating a feedback loop. Building on the Stackelberg game formulation of this interaction, we derive an analytical decomposition of the policy's true optimization gradient into a standard policy gradient and a parameter-steering term that captures the policy's influence on the RM's future parameters. We show that standard iterative RLHF, which drops this steering term entirely, suffers from alignment collapse: the policy systematically exploits the RM's blind spots, producing low-quality, high-reward outputs whose feedback reinforces the very errors it exploits. To mitigate this, we propose foresighted policy optimization (FPO), a mechanism-design intervention that restores the missing steering term by regularizing the policy's parameter-steering effect on RM updates. We instantiate FPO via a scalable first-order approximation and demonstrate that it prevents alignment collapse on both controlled environments and an LLM alignment pipeline using Llama-3.2-1B.
ExpLang: Improved Exploration and Exploitation in LLM Reasoning with On-Policy Thinking Language Selection
Changjiang Gao ⋅ Zixian Huang ⋅ Kaichen Yang ⋅ Jiajun Chen ⋅ Jixing Li ⋅ Shujian Huang
Current large reasoning models (LRMs) have shown strong ability on challenging tasks after reinforcement learning (RL) based post-training. However, previous work mainly focuses on English reasoning in expectation of the strongest performance, despite the demonstrated potential advantage of multilingual thinking, as well as the requirement for native thinking traces by global users. In this paper, we propose ExpLang, a novel LLM post-training pipeline that enables on-policy thinking language selection to improve exploration and exploitation during RL with the use of multiple languages. The results show that our method steadily outperforms English-only training with the same training budget, while showing high thinking language compliance for both seen and unseen languages. Analysis shows that, by enabling on-policy thinking language selection as an action during RL, ExpLang effectively extends the RL exploration space with diversified language preference and improves the RL exploitation outcome. The method is orthogonal to most RL algorithms and opens up a new perspective on using multilinguality to improve LRMs.
Exploration via Exploitation: The Blessing of Reward Diversity in Personalized Federated RL
Jiin Woo ⋅ Gauri Joshi ⋅ Yuejie Chi
While federated reinforcement learning has traditionally guaranteed high efficiency in homogeneous environments, collaboration is often hindered when agents possess distinct environments and goals. In this paper, we flip this paradigm, framing reward heterogeneity not as a hurdle to consensus, but as a structural asset that catalyzes exploration. We first propose Personalized Federated Upper Confidence Bound Value Iteration (PF-UCBVI), which achieves optimal linear speedup by decoupling shared dynamics from personalized objectives. To eliminate the risks of explicit exploration, we then introduce Personalized Federated Exploration-Free Value Iteration (PF-EFVI), a purely greedy algorithm that leverages reward diversity to ensure state-action coverage. We prove that PF-EFVI attains logarithmic regret without explicit exploration bonuses under sufficient reward diversity. Our results show that agent disagreement is a vital resource that shifts the exploration burden from temporal complexity to spatial diversity, enabling safe and efficient collective learning.
Multimodal Retrieval-Augmented Generation (RAG) helps mitigate hallucinations in Vision-Language Models (VLMs) by grounding generations in external knowledge bases. However, this externalized memory also introduces privacy risks, as external queries may reveal signals about sensitive visual records in the retrieval corpus, such as medical images, scanned contracts, and proprietary business documents. This exposes a retrieval-corpus privacy risk: private records may be detectable through interactions with multimodal RAG systems. To study this risk, we propose the Semantic Degradation Attack (SDA), a two-query method for exposing private corpus leakage in multimodal RAG by testing how strongly generated responses depend on retrieved evidence. SDA constructs transferable retrieval-disrupting perturbations using local surrogate visual encoders. By measuring the drop in the semantic similarity of the VLM's generated descriptions before and after perturbation, SDA can distinguish whether a target image-caption record exists in the private retrieval corpus. Extensive experiments demonstrate that SDA consistently detects private corpus leakage more effectively than existing baselines on two image-caption datasets across five popular VLMs.
Extremely Sparse-View Computed Tomography from 2D Projections via Pose-Aware Diffusion Priors
Linrui Dai ⋅ Jinxiu Liang ⋅ Chu Zhou ⋅ Mae Yamaguchi ⋅ Ryuya Takahashi ⋅ Imari Sato
Extremely sparse-view computed tomography (xSV-CT) is critical when scanning time, motion, or radiation tolerance limits each scan to only a few projections. In many such scenarios, the same acquisition limits also constrain dataset construction, making pre-reconstructed 3D volumes from dense scans unavailable. This exposes a gap in current sparse-view approaches: Optimization-based methods, which use the sparse measurements directly but lack data-driven priors to suppress artifacts, and learning-based methods, which use strong priors but usually require densely scanned training volumes acquired outside the xSV-CT regime. We study this train-sparse / infer-sparse setting by learning a pose-aware diffusion prior from one 2D projection per training sample, then guiding diffusion sampling with measured-view fidelity, shared-volume reprojection, and projection-physics consistency for tomographic reconstruction. On two public CT datasets across 1-, 4-, and 8-view settings, the proposed method achieves the best novel-view synthesis scores and best or competitive reconstruction quality compared with baselines that use dense-volume supervision.
Factorized Gradients for Scalable Highly-expressive Parametric Diffeomorphisms
Amit Aflalo ⋅ Eran Treister ⋅ Chaim Baskin ⋅ Oren Freifeld
Diffeomorphisms provide a flexible, topology-preserving mathematical tool for modeling complex spatial deformations, and are used in various scientific fields. However, simultaneously achieving high expressivity and computational efficiency remains challenging. Continuous Piecewise-Affine Based (CPAB) transformations offer an attractive solution by parameterizing a family of diffeomorphisms via continuous velocity fields that are piecewise affine *w.r.t.* a chosen tessellation of the domain. Importantly, $d=\dim(\theta)$, the dimension of the parameter vector $\theta$, depends on the fineness of the tessellation rather than on the data resolution. Despite this compact parameterization, CPAB optimization has long been hindered by the tight coupling between trajectory integration and computing gradients *w.r.t.* $\theta$. We introduce **FG-CPAB**, a factorized-gradient formulation for scalable CPAB transformations. By separating integration from parameter-space projection, our method reduces gradient computation complexity from $(\mathcal{O}(dTN))$ to $(\mathcal{O}(TN + d C))$, where $(T)$ is the number of integration steps, $(N)$ is the number of transformed points, and $(C)$ is the number of tessellation cells. This reformulation yields huge speedups (e.g., $\approx 10^4\times$ in 2D), substantial memory savings (e.g., $\approx 500\times$ reduction in peak GPU memory usage in 2D), and dramatically higher expressiveness. This makes, for the first time, fine-tessellation CPAB practical in 2D and 3D, unlocking the potential of highly-expressive parametric diffeomorphisms.
Failing to Falsify: Evaluating and Mitigating Confirmation Bias in Language Models
Ayush Rajesh Jhaveri ⋅ Anthony GX-Chen ⋅ Ilia Sucholutsky ⋅ Eunsol Choi
Confirmation bias, the tendency to seek evidence that supports rather than challenges one's belief, hinders one's reasoning ability. We examine whether large language models (LLMs) exhibit confirmation bias by adapting the rule-discovery study from human psychology: given a sequence of three numbers (a "triple"), an agent engages in an interactive feedback loop where it (1) proposes a new triple, (2) receives feedback on whether it satisfies the hidden rule, and (3) guesses the rule. Across eleven LLMs of multiple families and scales, we find that LLMs exhibit confirmation bias, often proposing triples to confirm their hypothesis rather than trying to falsify it. This leads to slower and less frequent discovery of the hidden rule. We further explore intervention strategies (e.g., encouraging the agent to consider counter examples) developed for humans. We find prompting LLMs with such instruction consistently decreases confirmation bias in LLMs, improving rule discovery rates from 42\% to 56\% on average. Lastly, we mitigate confirmation bias by distilling intervention-induced behavior into LLMs, showing promising generalization to a new task, the Blicket test. Our work shows that confirmation bias is a limitation of LLMs in hypothesis exploration, and that it can be mitigated via injecting interventions designed for humans.
FairMT: Fairness for Heterogeneous Multi-Task Learning
Guanyu Hu ⋅ Tangzheng Lian ⋅ Na Yan ⋅ Dimitrios Kollias ⋅ Xinyu Yang ⋅ Oya Celiktutan ⋅ Siyang Song ⋅ Zeyu Fu
Fairness in multi-task learning (MTL) is challenging when heterogeneous output spaces must be controlled through a shared multi-head model. Existing fairness methods are often tied to a single prediction type, while MTL optimizers coordinate task utility without specifying how heterogeneous group disparities should be represented, aggregated, or propagated through task-specific heads. We introduce FairMT, which frames fair heterogeneous MTL through a coordinate-allocation view. FairMT builds a Fairness-Coordinate Interface for binary detection, one-vs-rest multi-class classification, and scalar regression, and instantiates these coordinates asymmetrically using the current best-performing valid group as a detached reference. This directs correction toward groups falling behind the reference, reducing disparities while better preserving task utility. Asymmetric Heterogeneous Fairness Disparity Aggregation (AHFDA) then adaptively allocates constraint pressure over normalized coordinates and converts heterogeneous disparities into a single differentiable constraint signal. Finally, a Head-Aware Optimization Proxy maps the allocated constraint through task-head geometry into task-combination weights for primal--dual training, reducing routing bias caused by heterogeneous head scales, sensitivities, and decision geometries. Experiments on visual and language MTL benchmarks show that FairMT reduces group-level disparities across heterogeneous output types while maintaining utility close to utility-oriented MTL optimizers.
FairSplit: Decomposing the Embedding Space for Fair Classification
Andrii Shkabrii ⋅ Anna Beer ⋅ Ylli Sadikaj ⋅ Claudia Plant
Fair classification faces a twofold challenge: maintaining high classification accuracy while limiting reliance on sensitive information. Our method FairSplit achieves this by learning a decomposition of the embedding space into a Fair Space used for prediction and a Sensitive Space that isolates sensitive information. During training, the model must simultaneously learn the classification task and separate fair from sensitive information. The Sensitive Space gives sensitive attributes a destination, so the Fair Space can be used for prediction without relying on them. This decomposition is guided via a learnable mask and the Hilbert-Schmidt Independence Criterion (HSIC) as a measure of statistical dependence. We provide theoretical guarantees showing that the demographic parity gap is upper-bounded by both (i) the learned dimensionality of the Fair Space and (ii) the HSIC between the Fair Space and sensitive attributes. FairSplit effectively handles multiclass tasks and non-binary sensitive attributes, overcoming key limitations of existing fair classification methods. Experiments on established fairness benchmarks show that FairSplit achieves competitive accuracy-fairness trade-offs.
FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation
Yuanzhi Wang ⋅ Xuhua Ren ⋅ Jiaxiang Cheng ⋅ bing ma ⋅ Kai Yu ⋅ Sen Liang ⋅ Wenyue Li ⋅ Tianxiang Zheng ⋅ Qinglin Lu ⋅ Zhen Cui
Identity-preserving text-to-video generation (IPT2V) empowers users to produce diverse and imaginative videos with consistent human facial identity. Despite recent progress, existing methods often suffer from significant identity distortion under large facial pose variations or facial occlusions. In this paper, we propose FaithfulFaces, a pose-faithful facial identity preservation learning framework to improve IPT2V in complex dynamic scenes. The key of FaithfulFaces is a pose-shared identity aligner that refines and aligns facial poses across distinct views via a pose-shared dictionary and a pose variation–identity invariance constraint. By mapping single-view inputs into a global facial pose representation with explicit Euler angle embeddings, FaithfulFaces provides a pose-faithful facial prior that guides generative foundations toward robust identity-preserving generation. In particular, we develop a specialized pipeline to curate a high-quality video dataset featuring substantial facial pose diversity. Extensive experiments demonstrate that FaithfulFaces achieves state-of-the-art performance, maintaining superior identity consistency and structural clarity even as pose changes and occlusions occur. The code and dataset pipeline will be released.
FAME: Forecasting Academic Impact via Continuous-Time Manifold Evolution
Jianrong Ding ⋅ Jianyuan Zhong ⋅ Zhengyan Shi ⋅ Qiang Xu
Large Language Models (LLMs) are increasingly used to brainstorm and evaluate research ideas, yet assessing such judgments is fundamentally difficult because the true impact of a new idea may take years to emerge. We address this challenge by using the impact forecasting of human-authored manuscripts as a verifiable proxy task. In a prospective forecasting study, we find that frontier LLMs fail to reliably distinguish high-impact papers from ordinary publications, suggesting that static text-based judging is insufficient for scientific evaluation. To address this limitation, we propose $\textbf{FAME}$ ($\underline{\text{F}}$orecasting $\underline{\text{A}}$cademic Impact via Continuous-Time $\underline{\text{M}}$anifold $\underline{\text{E}}$volution), a spatiotemporal framework for modeling the dynamic trajectories of scientific topics. FAME projects papers into a dynamic latent space informed by textual features and a verified knowledge-flow graph, learning geometric constraints that align impactful manuscripts with the forward momentum of their fields. Experiments on 3,200 arXiv papers across three fast-evolving subfields show that FAME consistently and substantially outperforms state-of-the-art LLM evaluators in prospective multidimensional impact forecasting. Furthermore, integrating FAME's dynamic geometric signals into LLMs significantly improves their forecasting performance. These results support manuscript impact forecasting as a useful, measurable proxy benchmark and position FAME as a strong, trajectory-aware foundation for automated scientific evaluation.
FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization
Quanjian Song ⋅ Yefeng Shen ⋅ Mengting Chen ⋅ Hao Sun ⋅ Jinsong Lan ⋅ Xiaoyong Zhu ⋅ Bo Zheng
Human-centric video customization, particularly at the garment level, has shown significant commercial value. However, existing approaches cannot support low-latency and interactive garment control, which is crucial for applications such as e-commerce and content creation. This paper studies how to achieve interactive multi-garment video customization while preserving motion coherence using only single-garment video data. We present FashionChameleon, a real-time and interactive framework for human-garment customization in autoregressive video generation, where users can interactively switch garment during generation. FashionChameleon consists of three key techniques: (i) Instead of training on multi-garment video data, we train a Teacher Model with In-Context Learning on a single reference–garment pair. By retaining the image-to-video training paradigm while enforcing a mismatch between the reference and garment image, the model is encouraged to implicitly preserve coherence during single-garment switching. (ii) To achieve consistency and efficiency during generation, we introduce Streaming Distillation with In-Context Learning, which fine-tunes the model with in-context teacher forcing and improves extrapolation consistency via gradient-reweighted distribution matching distillation. (iii) To extend the model for interactive multi-garment video customization, we propose Training-Free KV Cache Rescheduling, which includes garment KV refresh, historical KV withdraw, and reference KV disentangle to achieve garment switching while preserving motion coherence. Our FashionChameleon uniquely supports interactive customization and long-video extrapolation, while achieving real-time generation at 23.8 FPS on a single GPU, 30-180$\times$ faster than existing baselines.
Fast and Accurate Probing of In-Training LLMs' Downstream Performances
Zhichen Liu ⋅ Tianle Lun ⋅ Zhibin Wen ⋅ Hao An ⋅ Yulin Ou ⋅ Jianhui Xu ⋅ Hao Zhang ⋅ Wenyi Fang ⋅ YANG ZHENG ⋅ Yang Xu
The paradigm of scaling Large Language Models (LLMs) in both parameter size and test time has pushed the boundaries of AI capabilities, but at the cost of making the traditional generative evaluation paradigm prohibitively expensive, therefore making the latency of LLM's in-training downstream performance evaluation unbearable. However, simple metrics like training loss (perplexity) are not always correlated with downstream performance, as sometimes their trends diverge from the actual task outcomes. This dilemma calls for a method that is computationally efficient and sufficiently accurate in measuring model capabilities. To address this challenge, we introduce a new in-training evaluation paradigm that uses a lightweight probe for monitoring downstream performance. The probes take the internal representations of LLM checkpoints (during training) as input and directly predict the checkpoint's performance on downstream tasks measured by \emph{success probability} (i.e., pass@1). We design several probe architectures, validating their effectiveness using the OLMo3-7B's checkpoints across a diverse set of downstream tasks. The probes can accurately predict a checkpoint's performance (with avg. AUROC$>$0.75), have decent generalizability across checkpoints (earlier predicts later), and reduce the computation latency from $\sim$1 hr (using conventional generative evaluation method) to $\sim$3 min. In sum, this work presents a practical and scalable in-training downstream evaluation paradigm, enabling a more agile, informed, and efficient LLM development process.
Fast-dLLM++: Fr\'{e}chet Profile Decoding for Faster Diffusion LLM Inference
Siva Rajesh Kasa ⋅ Yasong Dai ⋅ Sumit Negi ⋅ Hongdong Li
Diffusion large language models promise parallel token generation, yet inference remains bottlenecked by deciding which masked tokens can be safely committed together. Fast-dLLM addressed this with KV caching and confidence-guided parallel decoding, but its decoding theory uses a homogeneous high-confidence assumption that effectively reduces each candidate set to its weakest selected token. We argue that this leaves speed on the table because real decoding steps exhibit heterogeneous confidence profiles. We propose Fast-dLLM++, a training-free extension that introduces Fr\'{e}chet profile decoding: selecting parallel commit sets from the full sorted confidence profile rather than a single worst-case confidence. The resulting rule is a heterogeneous-confidence generalization of Fast-dLLM's factor selector and it recovers the previous rule exactly in the equal-confidence case and adds a provable heterogeneity bonus when the selected tokens have uneven confidences. Fast-dLLM++ leaves the model, diffusion process, and cache implementation entirely unchanged, making it a drop-in replacement for existing Fast-dLLM decoding. Experiments on GSM8K, MATH, HumanEval, and MBPP with the LLaDA-8B model show that the theoretical improvement translates directly into empirical gains: profile-aware selection improves the accuracy-throughput frontier by exploiting safe parallelism that weakest-token rules miss, achieving up to 37\% higher throughput at comparable accuracy. Our anonymous code release is at https://anonymous.4open.science/r/fast-dllm-plus-plus/.
Fast Reconstruction of Exact Maxwell Dynamics from Sparse Data
Dan DeGenaro ⋅ Xin Li ⋅ Obed Amo ⋅ Michael Pokojovy ⋅ Sarah Bargal ⋅ Markus Lange-Hegermann ⋅ Bogdan Raita
We introduce FLASH-MAX, a shallow, exact-by-construction neural network architecture for predicting homogeneous electromagnetic fields from sparse pointwise observations. Each hidden neuron represents a separate exact solution to Maxwell's equations, so that the network satisfies the governing equations symbolically by construction and can be trained end-to-end from sparse data within seconds. We prove a universal approximation result showing that this exact model class remains universal on arbitrary domains. FLASH-MAX reaches sub-1% relative validation error from about 1K sparse pointwise observations in seconds, all while maintaining a zero PDE residual, and keeps single digit errors even for only 100 observations sampled from 3D space. These results suggest that moving governing structure from the loss into the hypothesis class can dramatically improve the trade-off between precision and optimization speed in scientific machine learning.
Fast-Slow Evolutionary Occupancy Prediction via Controlled Dynamics
Runhe Yang ⋅ Zhiyuan Zhou ⋅ Yuxiang Yan ⋅ Xuetong Yang ⋅ Yikai Pan ⋅ Jiacheng Tang ⋅ Zhuolin He ⋅ Jianhua Han ⋅ Hang Xu ⋅ Jian Pu
3D semantic occupancy prediction provides dense geometric and semantic scene representations for autonomous driving, where both high accuracy and low response latency are crucial for safe downstream forecasting and planning. Existing methods usually exhibit different accuracy-latency characteristics. Performance-oriented models provide stronger geometric and semantic predictions, while deployment-friendly models provide faster responses. This motivates EvoOcc, a fast-slow evolutionary occupancy prediction framework that jointly exploits their complementary strengths to improve the accuracy-latency balance without changing the upstream architectures. EvoOcc models evolutionary occupancy prediction as a continuous time controlled dynamical system, enabling flexible adaptation to irregular prediction arrival of both fast and slow system. The framework further decomposes the evolution process into dual state evolution where ego state evolution for the deterministic coordinate shifts and scene state evolution for the genuine geometric and semantic scene changes. Extensive experiments on three different fast-slow system pairs and irregular prediction arrivals show that EvoOcc consistently improves over fast system baselines while maintaining much lower latency than the slow system, indicating a favorable accuracy-latency balance and stable behavior under practical inference situations.
FAUST: Federated Asynchronous Update with Staggered Timescales for Low-Communication Foundation Model Training
Yunlong Tan ⋅ Mingqiao Mo ⋅ Hao Zhang ⋅ Shengxing Qin
Large language model training with sequence-parallel data-parallel (SP-DP) is bottlenecked by the interplay between intra-step activation redistribution and inter-step gradient synchronization. Existing infrequent-communication methods like FedAvg keep states local or reset them lack convergence guarantees and are unstable under long-context sequence-parallel regimes where each device holds only a partial activation shard; Local Adam synchronizes all states jointly and is convergent but triples the collective payload, and its uniform synchronization period cannot distinguish the semantically mandatory intra-step All-to-All from the temporally deferrable inter-step AllReduce. We propose FAUST (Federated Asynchronous Update with Staggered Timescales), a compiler-orchestrated training system that composes compile-time sequence-parallel graph rewriting with tiered runtime synchronization, assigning independent averaging periods to parameters ($K_x$), first moments ($K_u = 3K_x$), and second moments ($K_v = 6K_x$) according to their impulse-response half-lives, while preserving convergence under the orthogonal collective constraint that intra-step All-to-All must retire before inter-step AllReduce begins. Our analysis shows that while first-moment drift dominates the convergence rate in-distribution, high-fidelity convergence guarantees require at least periodic synchronization of the second moment. Experiments on language models up to 1.3B parameters show that FAUST incurs $303{\times}$ less wall-clock communication overhead than standard DDP and $1.78{\times}$ less than the previous state-of-the-art DES-LOC, while achieving perplexity within $4.6\%$ of the fully-synchronized baseline in $27\%$ less wall-clock time. On bandwidth-constrained links, FAUST delivers $48$--$303{\times}$ speedup over DDP.
Feature Information Dynamics in Diffusion
Jia-Shu Pan ⋅ Tao Zhang ⋅ Yufei Huang ⋅ Yanjun Sheng ⋅ Tailin Wu
Diffusion models generate data through a continuum of denoising problems, and are widely observed to reveal coarse structure before fine detail. Yet, this intuition is mostly empirical and qualitative. We introduce \emph{feature information dynamics}, an information-theoretic framework for localizing when a feature is generated during diffusion. Using the I-MMSE identity, we connect the rate of feature mutual information change to a gap between optimal unconditional and feature-conditional denoising losses, yielding practical estimators for feature information density. We further develop a chained decomposition that separates shared from incremental information in a feature hierarchy. We use this framework first to quantitatively confirm spectral autoregression in pixel diffusion, and then to extend the analysis beyond frequency: under a class$\to$mask$\to$Canny conditioning chain, the per-feature information densities differ across pixel, SDVAE, VAVAE, and RAE, exposing fundamental differences between these representations.
Federate the Router: Learning Language Model Routers with Sparse and Decentralized Evaluations
Baris Askin ⋅ Shivam Patel ⋅ Anupam Nayak ⋅ Andrea Vigano ⋅ Jiin Woo ⋅ Gauri Joshi ⋅ Carlee Joe-Wong
Large language models (LLMs) are increasingly accessed as remotely hosted services by edge and enterprise clients that cannot run frontier models locally. Since models vary widely in capability and price, routing queries to models that balance quality and inference cost is essential. Existing router approaches assume access to centralized query-model evaluation data. However, these data are often fragmented across clients, such as end users and organizations, and are privacy-sensitive, which makes centralizing data infeasible. Additionally, per-client router training is ineffective since local evaluation data is limited and covers only a restricted query distribution and a biased subset of model evaluations. We introduce the first federated framework for LLM routing, enabling clients to learn a shared routing policy from local offline query-model evaluation data. Our framework supports both parametric multilayer perceptron router and nonparametric K-means router under heterogeneous client query distributions and non-uniform model coverage. Across two benchmarks, federated collaboration improves the accuracy-cost frontier over client-local routers, both via increased effective model coverage and better query generalization. Our theoretical results also validate that federated training reduces routing suboptimality.
Fed-SB: A Silver Bullet for Extreme Communication Efficiency and Performance in (Private) Federated LoRA Fine-Tuning
Raghav Singhal ⋅ Kaustubh Ponkshe ⋅ Rohit Vartak ⋅ Lav Varshney ⋅ Praneeth Vepakomma
Low-Rank Adaptation (LoRA) has become ubiquitous for efficiently fine-tuning foundation models. However, federated fine-tuning using LoRA is challenging due to suboptimal updates arising from traditional federated averaging of individual adapters. Existing solutions either incur prohibitively high communication cost that scales linearly with the number of clients or suffer from performance degradation due to limited expressivity. We introduce Fed-SB, a novel approach for federated fine-tuning of LLMs using LoRA-SB, a recently proposed low-rank adaptation method. LoRA-SB optimally aligns the optimization trajectory with the ideal low-rank full fine-tuning projection by learning a small square matrix (R) between adapters B and A, keeping other components fixed. Direct averaging of R guarantees exact updates, substantially reducing communication cost, which remains independent of the number of clients, and enables scalability. Fed-SB achieves state-of-the-art performance across commonsense reasoning, arithmetic reasoning, and language inference tasks while reducing communication costs by up to 230x. In private settings, Fed-SB further improves performance by (1) reducing trainable parameters, thereby lowering the noise required for differential privacy and (2) avoiding noise amplification introduced by other methods. Overall, Fed-SB offers a state-of-the-art, efficient, and scalable solution for both private and non-private federated fine-tuning. Our code is available publicly at: https://github.com/CERT-Lab/fed-sb.
Feeling of Knowing in Large Language Models
Zichuan Fu ⋅ Xian Wu ⋅ Jingtong Gao ⋅ Wenlin Zhang ⋅ Jiaxuan Li ⋅ Binhao Wang ⋅ Yimin Deng ⋅ Guojing Li ⋅ Xiaopeng Li ⋅ Derong Xu ⋅ Yefeng Zheng ⋅ Xiangyu Zhao
In real-world deployment, large language models (LLMs) are frequently updated through post-training techniques to maintain up-to-date knowledge. Yet their reliability depends not only on what the LLM knows, but also on whether it knows what it knows—a self-assessment known as the feeling of knowing (FoK). FoK is the signal behind selective generation, retrieval triggering, and model routing; when miscalibrated, it leads systems to overconfidence on unknown questions or to refuse ones they could have answered. As LLMs are updated through posttraining, however, we discover that what an LLM knows can change without a corresponding update to its FoK. This creates a knowledge–FoK desynchronization challenge, where FoK estimators calibrated to the pre-trained LLM become unreliable as the model’s knowledge state evolves. To address this challenge, we propose a two-channel design: a knowledge channel that changes what the LLM knows and a FoK channel that judges whether the current knowledge state supports answering. We train these channels with a three-stage procedure: onpolicy supervised fine-tuning (OP-SFT) updates the knowledge channel from corrected on-policy answers, knowledge-delta FoK training (K∆-FT) trains the FoK channel over intermediate knowledge states, and FoK-guided policy optimization (FGPO) uses the judgment to generate the corresponding answer or explanation. Experiments on three datasets show that explicit FoK alignment substantially improves judgment quality on updated knowledge. The code is available at https: //anonymous.4open.science/status/Know-thyself-code-8D65.
fev-bench: A Realistic Benchmark for Time Series Forecasting
Oleksandr Shchur ⋅ Abdul Fatir Ansari ⋅ Ali Caner Turkmen ⋅ Lorenzo Stella ⋅ Nick Erickson ⋅ Pablo A Guerron ⋅ Michael Bohlke-Schneider ⋅ Yuyang (Bernie) Wang
Benchmark quality is critical for meaningful evaluation and sustained progress in time series forecasting, particularly with the rise of time series foundation models (TSFMs). Existing benchmarks often have limited domain coverage or overlook real-world settings such as tasks with covariates. Their aggregation procedures frequently lack statistical rigor, making it unclear whether performance differences reflect true improvements or random variation. Many benchmarks lack consistent evaluation infrastructure or are too rigid for integration into existing pipelines. To address these gaps, we propose fev-bench, a benchmark of 100 forecasting tasks across seven domains, including 46 with covariates. Supporting the benchmark, we introduce fev, a lightweight Python library for forecasting evaluation emphasizing reproducibility and integration with existing workflows. Using fev, fev-bench employs principled aggregation with bootstrapped confidence intervals to report performance along two dimensions: win rates and skill scores. We evaluate various models on fev-bench and find that existing TSFMs often miss accuracy by ignoring covariates, pointing to a clear direction for future work.
Few Contrastive Attention Heads Enable Visual Grounding in Large Vision-Language Models
Neha Sengar ⋅ Andres Saurez ⋅ Dongsoo Har ⋅ Sangkeum Lee
Visual grounding aims to localize image regions corresponding to natural language expressions. While recent Large Vision-Language Models (LVLMs) have shown impressive multi-modal understanding capabilities, their application to visual grounding typically requires fine-tuning and architectural modifications. This requirement, however, can be ignored, considering that text and images tend to have similar feature representations that appear to be approximately linearly disentangled, enabling cleaner extraction of spatial information from LVLMs without any task-specific training. From this viewpoint, we propose an attention-head discovery framework that requires zero labeled grounding samples and no architectural modifications, and identifies discriminative localization heads without manual inspection. Through dual prompting with target and contrastive descriptions, we compute differential residual representations and project them through attention head output matrices to measure per-head spatial contributions via four complementary scores. By aggregating signals using importance-weighted query difference scores from only the top-10 attention heads, we outperform training-free non-LVLM baseline by up to 27.95% on RefCOCO, 21.93% on RefCOCO+, and 8.40% on RefCOCOg. Our method outperforms LVLM baseline by up to 8.04% on RefCOCO without requiring ground-truth category labels.
Fiedler-Regularized Causal Discovery for Sparse Connected DAGs
Amine M'Charrak ⋅ Abbavaram Gowtham Reddy ⋅ Thomas Lukasiewicz ⋅ Michael Bronstein ⋅ Krikamol Muandet
Causal discovery often starts from a set of variables that were measured together because they are believed to describe related parts of the same underlying process. Standard continuous DAG learners encode data fit, sparsity, and acyclicity, but they can still return graphs that fragment into isolated variables or disconnected components. We study a simple structural prior for this setting: the learned causal DAG should remain sparse while its undirected skeleton should be weakly connected. We introduce Fiedler regularization, a differentiable spectral penalty based on the algebraic connectivity of a smooth support-preserving skeleton of the learned weighted adjacency matrix. The resulting penalty can be added modularly to differentiable causal discovery objectives, providing connectivity control without changing the underlying learner. We instantiate the approach for GOLEM, graph autoencoder, and DAG-GNN models. Across sparse connected synthetic DAGs up to 200 variables, including connected-conditioned Erdős-Rényi, sparse connected, scale-free, and small-world graph families, Fiedler-regularized learners improve fragmentation control and structural recovery across learner families. These results position algebraic connectivity as a principled and practical prior for learning sparse causal DAGs in settings where the measured variables are expected to have a connected causal structure.
Filter Banks: from Low-Rank Representations to Deep Models for Efficient Time Series Forecasting
Ashutosh Vaishnav ⋅ Mohsen Amidzadeh ⋅ Teemu Kämäräinen ⋅ Matti Siekkinen ⋅ Mario Di Francesco
Time series forecasting is an actively researched problem with diverse applications. Recent work has shown that simple linear models can compete with complex deep learning architectures in terms of forecasting accuracy. This has led to the design of several small models, yet the reasons behind their effectiveness remain insufficiently understood. We address this gap through a theoretical framework grounded in reduced-rank regression and kernel analysis. We show both analytically and empirically that the intrinsic complexity of many TSF tasks is lower than commonly assumed and that the performance of linear models is linked to their alignment with low-rank data structures. Building on these insights, we propose a new design for low-rank neural networks that incur lower computational cost than linear models. Specifically, we introduce a simple filter bank architecture that provides a principled way to control model complexity through tunable and interpretable hyperparameters that directly correspond to the rank of the forecasting model. The proposed architecture also serves as a versatile building block for constructing deeper networks. Our evaluation shows that filter bank architectures achieve state-of-the-art results on most long-term TSF benchmarks, with lower computational cost than recently proposed solutions.
FinAl: Fine-grained Alignment for Detail-Preserving Medical Vision-Language Pretraining
Jongsu Youn ⋅ Dongyoung Lee ⋅ JaeHyeong Bae ⋅ SANG-IL CHOI ⋅ Sangtae Choi ⋅ Dae Ung Jo ⋅ Jongwon Choi
We present a vision–language pretraining (VLP) framework for the medical domain that focuses on preserving fine-detail clinical information, such as disease severity, which plays a critical role in downstream medical applications. Despite recent advances in medical VLP, existing approaches often struggle to retain such subtle semantics. This limitation mainly stems from the use of contrastive learning schemes that treat supervision as hard one-hot labels, overlooking the ordinal and overlapping nature of clinical findings and thereby introducing false negatives. To overcome these challenges, we propose FinAl, a fine-grained alignment approach that incorporates soft alignment targets into vision-language pretraining. These targets are derived from structured radiology reports and encode relative semantic relationships among clinical descriptions, allowing fine-grained linguistic information to be transferred into the visual representation space. Extensive experiments show that FinAl improves performance on fine-grained clinical tasks, including severity grading and severity-aware retrieval, with additional evaluations on challenging fine-grained scenarios further supporting its robustness. At the same time, the proposed method maintains competitive performance on standard coarse-grained clinical benchmarks.
Finding Koopman Invariant Subspaces via Personalized PageRank
Hyukpyo Hong ⋅ Qin Li ⋅ Matthew J Colbrook ⋅ Hanbaek Lyu
Selecting a finite dictionary of observables whose span is Koopman-invariant is a central challenge in data-driven Koopman operator approximation. We address this problem by exploiting zero-block structure in Extended Dynamic Mode Decomposition (EDMD) matrices. We show that any sub-dictionary whose span is Koopman-invariant induces an exact zero block in the EDMD matrix, even for finite data. We then show that such blocks can be detected by applying PageRank to a row-normalized EDMD matrix constructed from a large initial dictionary. The theory extends to approximately invariant subspaces and yields stronger guarantees for personalized PageRank (PPR) when the seed observables lie inside the target block and reach all observables in that block. Combining EDMD concentration bounds with PageRank perturbation theory gives end-to-end detection guarantees with $O(M^{-1/2})$ finite-sample scaling and explicit constants. More generally, without assuming an invariant subspace exists, high PPR mass on a sub-dictionary controls discounted multi-step leakage from the seed observables. Numerical experiments on the Duffing oscillator, Van der Pol oscillator, Lorenz system, and a three-well Ramachandran potential suggest that the method identifies compact, interpretable dictionaries with accurate predictions.
FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies
Xintong Hu ⋅ Xuhong Huang ⋅ JINYU ZHANG ⋅ Yutong Yao ⋅ Yuchong Sun ⋅ Qiuyue Wang ⋅ Mingsheng Li ⋅ Yitao Liu ⋅ Yixuan Chen ⋅ Yingming Zheng ⋅ Shuai Bai ⋅ Tao Yu
Vision-Language-Action (VLA) models have advanced rapidly by leveraging large-scale vision-language priors and robot trajectory data, yet they remain weak at following fine-grained human instructions. We argue that this limitation stems from a fundamental mismatch between language and action supervision in existing robot datasets. Most open-source datasets annotate each trajectory with a goal-oriented task description, while leaving the underlying execution process largely unspecified, including motion trajectories, contact and approach patterns, object state transitions, recovery behavior, and final configurations. Since different choices along these dimensions can lead to substantially different action sequences, coarse task descriptions induce a many-to-one mapping from action trajectories to language, preventing precise action-instruction alignment. To address this problem, we propose FineVLA, a framework for constructing and leveraging fine-grained embodied action instructions for VLA learning. FineVLA includes FineVLA-Tool, a scalable pipeline for data cleaning, temporal alignment, and fine-grained annotation; RoboFine-VLM, a vision-language model for fine-grained robotic action understanding and annotation; RoboFine-Bench, a benchmark for evaluating fine-grained robotic action understanding via VQA and captioning; and FineVLA-Policy, a VLA policy trained with fine-grained instructions on open-source ALOHA robot data. Across benchmark evaluation and policy learning experiments, FineVLA demonstrates that high-density, action-aligned language supervision leads to more controllable and instruction-sensitive robot policies. Our results highlight fine-grained instruction alignment as a critical step toward VLA systems that not only complete tasks, but complete them in the way humans specify.
First-Token Attraction in Mamba Dynamics
Trinh Nguyen ⋅ Duy-Tung Pham ⋅ Minh-Khoi Nguyen-Nhat ⋅ Hoang-Son Do ⋅ Tan Nguyen ⋅ Thieu Vo
We study token dynamics in deep Mamba models, with a particular focus on the attracting role of the first token. Unlike attention mechanisms, where token interactions are governed by positive scalar coefficients, Mamba dynamics involve full matrix-valued interaction coefficients with no positivity or definiteness guarantees. This structural difference makes Mamba dynamics substantially harder to analyze and prevents the direct application of existing techniques for attention-based models. To overcome this challenge, we introduce a new analytical framework based on a high-dimensional invariant box and a sharp exit-time argument, which together recover an effective form of sign-definiteness. Using this framework, we characterize the possible limit points of token trajectories and prove explicit exponential convergence rates. Our analysis reveals that the first token has privileged dynamics: its limiting direction attracts subsequent tokens, inducing token clustering and attention-sink phenomena in deep Mamba models. We further show that this attraction mechanism extends to broader sequence dynamics, including state-space models, causal attention, and classical time-series models. Experiments on Falcon-Mamba-7B validate our predictions, showing that concentration on the first token increases with depth and dominates in later layers.
Fisher Decorator: Refining Flow Policy via A Local Transport Map
Xiaoyuan Cheng ⋅ Haoyu Wang ⋅ Wenxuan Yuan ⋅ Ziyan Wang ⋅ Zonghao Chen ⋅ Li Zeng ⋅ Zhuo Sun
Recent advances in flow-based offline reinforcement learning (RL) have achieved strong performance by parameterizing policies via flow matching. However, they still face critical trade-offs among expressiveness, optimality, and efficiency. In particular, existing flow policies interpret the $L_2$ regularization as an upper bound of the 2-Wasserstein distance ($W_2$), which can be problematic in offline settings. This issue stems from a fundamental geometric mismatch: the behavioral policy manifold is inherently anisotropic, whereas the $L_2$ (or upper bound of $W_2$) regularization is isotropic and density-insensitive, leading to systematically misaligned optimization directions. To address this, we revisit offline RL from a geometric perspective and show that policy refinement can be formulated as a local transport map—an initial flow policy augmented by a residual displacement. By analyzing the induced density transformation, we derive a local quadratic approximation of the KL-constrained objective governed by the Fisher information matrix, enabling a tractable anisotropic optimization formulation. By leveraging the score function embedded in the flow velocity, we obtain a corresponding quadratic constraint for efficient optimization. Our results reveal that the optimality gap in prior methods arises from their isotropic approximation. In contrast, our framework achieves a controllable approximation error within a provable neighborhood of the optimal solution. Extensive experiments demonstrate state-of-the-art performance across diverse offline RL benchmarks.
FiTS: Interpretable Spiking Neurons via Frequency Selectivity and Temporal Shaping
Jongmin Choi ⋅ Joon Son Chung
Spiking Neural Networks (SNNs) are a promising framework for event-driven temporal processing. Prior work has improved temporal modeling through richer neuron dynamics and network-level mechanisms such as recurrence and delays, but it remains unclear how individual spiking neurons should specialize within a network. In this work, we introduce \textbf{FiTS}, a spiking neuron that factorizes temporal computation within each neuron into Frequency Selectivity (FS) and Temporal Shaping (TS). The FS module parameterizes each neuron's target frequency as the maximizer of its subthreshold magnitude response, while the TS module reshapes when frequency components contribute to membrane voltage accumulation through group-delay modulation. On auditory benchmarks where frequency selectivity and timing are central to the input structure, FiTS consistently improves over a plain Leaky Integrate-and-Fire (LIF) baseline in simple feedforward SNNs without recurrence or network-level delays, while remaining competitive with strong temporal SNN baselines. Beyond accuracy, the learned target frequencies and group-delay shifts provide interpretable neuron-level summaries of the frequency and timing organization learned within the network.
FLASH: Efficient Visuomotor Policy via Sparse Sampling
Jiaqi Bai ⋅ Jindou Jia ⋅ Yuxuan Hu ⋅ Gen Li ⋅ Xiangyu Chen ⋅ Tuo An ⋅ Kuangji Zuo ⋅ Jianfei Yang
Generative models such as diffusion and flow matching have become dominant paradigms for visuomotor policy learning, yet their reliance on iterative denoising incurs high inference latency incompatible with real-time robotic control. We present a Fast Legendre-polynomial Action policy via Sparse History-anchored flow (FLASH), which replaces discrete action-chunk generation with continuous Legendre polynomial trajectory representation. Specifically, by fitting expert demonstrations under sparse temporal sampling, FLASH enables a single inference to cover a significantly extended action horizon. To further accelerate generation, FLASH initiates the flow matching process from history polynomial coefficients rather than uninformative Gaussian noise, shortening the transport distance and enabling accurate single-step inference. Moreover, analytic polynomial differentiation directly provides desired velocity feed-forward signals to the torque controller without numerical approximation. Extensive experiments on five simulated and two real-world manipulation tasks demonstrate that FLASH achieves state-of-the-art success rates ($\ge 92\%$ across all tasks), a per-episode inference time of $31.40\,ms$ (up to $175\times$ faster than diffusion policies and $18\times$ faster than prior flow matching policies), up to $4\times$ faster training convergence than ACT, and $5\times$ to $7\times$ reduction in controller tracking error compared to discrete-action baselines.
FLINT: Coupling Proximal Initialization and Bounded Stochasticity for Flow-Matching Inverse Problems
Junseo Bang ⋅ daewon choi ⋅ Dong Ju Mun ⋅ Se Young Chun
Flow matching models have emerged as a dominant paradigm for generative posterior sampling in imaging inverse problems. Existing solvers follow two distinct lines of work: trajectory-space methods, which refine sampling dynamics from random initial noise, and source-space methods, which optimize the initial noise itself. While prior work treats these as independent design choices, we show that they are fundamentally coupled. We derive a unified reconstruction error bound that links initialization proximity and cumulative sampling stochasticity through a multiplicative envelope, yielding two key insights. First, we identify a sufficiency regime in which a moderately optimized proximal source initialization is enough to attain the optimal error floor, rendering exhaustive source-space optimization unnecessary. Second, this initial advantage is fragile: it is eroded by excessive stochasticity accumulation during guided sampling, necessitating a mechanism that bounds the cumulative noise injection. These findings motivate FLINT, a two-stage framework that operationalizes this coupling. FLINT establishes a high-fidelity starting point via a manifold-constrained proximal objective in only a few iterations, then performs guided sampling under a bounded stochasticity schedule that prevents erasure of the initialization advantage. Across diverse inverse problems, FLINT delivers state of the art perceptual quality and competitive distortion metrics while maintaining superior computational efficiency.
Flipping Bits, Not Gradients: Sharpness-Aware Minimization Directly on the Boolean Hypercube
Ba-Hien Tran
Sharpness-Aware Minimization (SAM) improves generalization in continuous deep learning, yet applying it to binary-weight networks via latent-space heuristics suffers from a fundamental geometric mismatch: continuous perturbations do not faithfully probe the discrete loss landscape on $\{-1,+1\}^n$. We introduce BOLD-SAM, the first sharpness-aware optimizer that operates natively on the Boolean hypercube, replacing the Euclidean $\ell_p$-ball with a $k$-bit Hamming ball and solving the resulting discrete min-max problem via greedy ascent followed by sharpness-aware bit-flip descent. The objective is theoretically justified from three complementary perspectives: PAC-Bayes, compression, and distributionally robust optimization. We prove an approximation guarantee for the ascent step and a finite-time convergence bound for the descent, both governed by the discrete interaction Hessian, which emerges as the unifying quantity linking ascent quality to convergence rate. We further connect discrete flatness to Forman-Ricci curvature and algorithmic complexity. Experiments on a wide range of architectures and datasets demonstrate consistent improvements in clean accuracy and out-of-distribution robustness over standard binary training and latent-space SAM baselines.
Flow-Guided Target-Space Alignment via Path Consistency
Ruizhi Yuan ⋅ Zeqiu Yu ⋅ Wei Gao ⋅ Wei Chen ⋅ Xiwei Tang ⋅ Lu Tang
Many multimodal alignment problems require mapping a source modality into a fixed space defined by a target modality, where the target space is not merely an intermediate embedding but is used directly for downstream analysis. In this setting, successful alignment should preserve both the global structure of the target distribution and the instance-level correspondence observed in paired source--target training data. Existing objectives typically address only part of this goal: contrastive losses encourage relative similarity, distribution-matching losses improve marginal agreement, and pointwise regression fits each source sample to its paired target without explicitly modeling target-space path geometry. We propose FlowAlign, a flow-guided target-space alignment method that repurposes conditional flow matching as a training-time alignment signal rather than an inference-time transport model. FlowAlign uses the learned flow field to guide the alignment of source-side representations with the fixed target space, encouraging them to preserve paired correspondence while remaining compatible with the global target-space structure. We further provide theoretical analysis showing how FlowAlign helps improve paired correspondence by controlling endpoint accuracy and capturing local target-space geometry. Experiments on a controlled synthetic task and paired single-cell multiome benchmarks show that FlowAlign improves alignment performance while preserving target-space structure for downstream analysis.
FlowMAS: Learning Multi-Agent Workflow Topology via Information-guided Generative Flow Network
Haitao Wang ⋅ Chenjing Liang ⋅ Haipeng Zhang ⋅ Jiawei Hu ⋅ SiCheng Wang ⋅ Songzhu Mei ⋅ Chenglu Wen ⋅ Siqi Shen ⋅ Cheng Wang
Automated multi-agent systems offer clear advantages over manually designed ones in scalability and adaptability, but existing workflow topology methods still face important limitations. Search-based methods are often computationally expensive, textual-gradient-based methods rely on coarse-grained feedback, and existing generation-based methods are not well suited to discrete workflow topologies with complex dependencies. To address these limitations, we propose FlowMAS, a multi-agent workflow topology method based on Generative Flow Networks (GFlowNets). FlowMAS models workflow generation as reward-guided flow over the topology space and introduces three components: a GFlowNet-based topology generation backbone, a curiosity-driven module for structure-aware exploration, and an information-guided optimization module for evaluating intermediate topologies. Concretely, the curiosity-driven module encourages exploration of structurally novel workflows, while the information-guided module measures both the information contribution and the communication efficiency of different operators to favor more informative and effective collaboration patterns. Experiments on six benchmark datasets with three LLM backbones show that FlowMAS consistently outperforms multiple baselines.
Flow Matching Policy Optimization with Mirror Descent and Entropy Constraints
Ting Gao ⋅ Stavros Orfanoudakis ⋅ Nan Lin ⋅ Winnie Daamen ⋅ Serge Hoogendoorn ⋅ Elvin Isufi
Balancing policy expressiveness with the exploration-exploitation trade-off is a core challenge in online Reinforcement Learning (RL). While Stochastic Differential Equation (SDE)-based diffusion policies can represent complex, multimodal action distributions, they suffer from two critical limitations: their stochastic reverse processes render entropy intractable (necessitating heuristic exploration), and computing policy gradients through long denoising chains is expensive and unstable. In this work, we show that ODE-based flow matching inherently resolves these issues by enabling both simulation-free policy optimization and tractable entropy computation. Building on this, we introduce Flow Matching Policy Optimization with Mirror Descent and Entropy Constraints (FMER). Our framework exploits this insight in three ways. First, we theoretically establish that minimizing an advantage-weighted conditional flow matching loss acts as a simulation-free surrogate for policy mirror descent. This steers the velocity field toward high-value regions while entirely avoiding backpropagation through the ODE solver. Second, we derive an analytic entropy objective that corrects for the density distortion caused by the $\tanh$ transformation (mapping an unbounded latent space to bounded actions), thereby facilitating principled maximum-entropy optimization. Finally, we dynamically tune the mirror descent temperature based on the effective sample size to enforce a robust trust region during training. Empirical evaluations demonstrate that FMER achieves superior performance on the challenging sparse-reward FrankaKitchen environment, while maintaining competitive results across standard dense-reward MuJoCo benchmarks. Code is available at \url{https://anonymous.4open.science/r/submission-25E2/}.
FlowSteer: Towards Agents Designing Agentic Workflows via Reinforced Progressive Canvas Editing
Mingda Zhang ⋅ Wenjin Liu ⋅ Tiesunlong Shen ⋅ Qika Lin ⋅ Rui Mao ⋅ Erik Cambria ⋅ Xiaoying Tang ⋅ Haoran Luo
In recent years, agentic workflows have been widely applied to solve complex human tasks. However, existing workflow construction still faces key challenges, including human-dependent workflow construction, the lack of graph-level execution feedback, and the inability to repair errors in-loop during long-horizon construction. To address these challenges, we propose FlowSteer, a new paradigm of Agent Designing Agentic Workflows — a single agent itself end-to-end designs the workflow that a downstream executor runs. To support this paradigm, we introduce the Workflow Canvas, a novel executable graph-state environment that returns syntax-checked execution feedback for every atomic edit. Built on the canvas, we further propose Reinforced Progressive Canvas Editing, in which a lightweight policy agent issues one atomic edit per turn conditioned on real canvas feedback, and is trained end-to-end via reinforcement learning. Moreover, FlowSteer provides a plug-and-play framework that supports diverse operator libraries and interchangeable LLM backends. Experimental results on twelve datasets show that FlowSteer significantly outperforms baselines across various tasks. Our code is available at https://anonymous.4open.science/r/FlowSteer-9B2E.
FluxFlow: Conservative Flow-Matching for Astronomical Image Super-Resolution
Shuhong Liu ⋅ Xining Ge ⋅ quanfeng xu ⋅ Ziteng Cui ⋅ Liuzhuozheng Li ⋅ Gengjia Chang ⋅ Jun Liu ⋅ Ziying Gu ⋅ Dong Li ⋅ Xuangeng Chu ⋅ Lin Gu ⋅ Tatsuya Harada
Ground-to-space astronomical super-resolution requires recovering space-quality images from ground-based observations that are simultaneously limited by pixel sampling resolution and atmospheric seeing, which imposes a stochastic, spatially varying PSF that cannot be resolved through upsampling alone. Existing methods rely on synthetic training pairs that fail to capture real atmospheric statistics and are prone to either over-smoothed reconstructions or hallucination sources with no physical counterpart in the observed sky. We propose FluxFlow, a conservative pixel-space flow-matching framework that incorporates observation uncertainty and source-region importance weights during training, and a training-free Wiener-regularised test-time correction to suppress hallucination sources while preserving recovered detail. We further construct the DESI--HST Dataset, the large-scale real-world benchmark comprising 19,500 real co-registered ground-to-space image pairs with real atmospheric PSF variation. Experiments demonstrate that FluxFlow consistently outperforms existing baseline methods in both photometric and scientific accuracy.
FocusBranch: Combinatorial Branch-and-Bound for $\ell_0$ Neural Network Robustness Verification
Renwei Deng ⋅ Yang He ⋅ Linyi Li ⋅ Yuepeng Wang
Verifying the local robustness of neural networks under $\ell_0$ perturbations is inherently combinatorial: a verifier must reason over all sparse subsets of perturbed pixels. This paper presents FocusBranch, a combinatorial branch-and-bound framework for $\ell_0$ robustness verification. The key insight is to decompose the $\ell_0$-ball using nodes, where each node consists of a focus pixel set $\mathcal{S}$ and a level $l$, requiring at least $l$ perturbed pixels to lie in $\mathcal{S}$. This abstraction unifies two branching mechanisms: concrete nodes branch through upper shadows to fix additional perturbed locations, while super nodes branch through covering designs to cover many perturbation cases with fewer child nodes. FocusBranch proves node safety by concretizing linear relaxations over node-constrained $\ell_0$-balls and recursively branches on unknown nodes. Rather than relying on a fixed verification strategy, it samples nodes across levels to estimate success rates and costs, then selects a starting level and branching strategy predicted to minimize verification time. Experiments on fully connected and convolutional classifiers for MNIST, Fashion-MNIST, and CIFAR-10 show that FocusBranch verifies $\ell_0$ local robustness effectively across architectures and outperforms state-of-the-art verifiers.
ForceLDM: A Force-aware Latent Dynamics Model for Contact-Rich Manipulation
Peiyuan Tang ⋅ Yao Li ⋅ Xiaodong Zhang ⋅ Wuyang Zhang ⋅ Qin Xia ⋅ Yanyong Zhang ⋅ Zijiang J Yang
Contact-rich manipulation requires precise control and complex physical interaction with the environment, posing significant challenges for robot policy learning. Vision-only policies struggle with such tasks, as contact states and force feedback are difficult to infer from vision alone. Existing methods mitigate this issue by incorporating force or tactile sensing. However, they often treat these signals as auxiliary observations rather than using them to model future interaction dynamics, leading to inefficient modality utilization and limited generalization. To address this, we propose ForceLDM, a force-aware latent dynamics model that enables the policy to anticipate future contact and motion dynamics in latent space and use them to guide action generation. Rather than predicting future observations in pixel space, which often requires large amounts of data and risk overfitting irrelevant details, ForceLDM learns a compact task-relevant representation through knowledge distillation. Specifically, a teacher network with privileged access to future force/torque and optical flow signals extracts future dynamic features, which are then distilled into a student network trained only on current observations. This enables the student to reason about upcoming contacts and physical dynamics at deployment without relying on future sensory inputs. To improve robustness and generalization, we further introduce a curriculum-based progressive noise injection strategy to mitigate over-reliance on future features. Experiments on five real-world contact-rich manipulation tasks demonstrate that ForceLDM significantly outperforms state-of-the-art baselines, generalizes effectively to novel objects, and remains robust to environmental perturbations.
Forecasting Downstream Performance of LLMs With Proxy Metrics
Arkil Patel ⋅ Siva Reddy ⋅ Marius Mosbach ⋅ Dzmitry Bahdanau
Progress in language model development is often driven by comparative decisions: which architecture to adopt, which pretraining corpus to use, or which training recipe to apply. Making these decisions well requires reliable performance forecasts, yet the two commonly used signals are fundamentally limited. Cross-entropy loss is poorly aligned with downstream capabilities, and direct downstream evaluation is expensive, sparse, and often uninformative at early training stages. Instead, we propose to construct proxy metrics by aggregating token-level statistics, such as entropy, top-k accuracy, and expert token rank, from a candidate model's next token distribution over expert-written solutions. Across three settings, our proxies consistently outperform loss- and compute-based baselines: 1) For cross-family model selection, they rank a heterogeneous population of reasoning models with mean Spearman $\rho = 0.81$ (vs. $\rho = 0.36$ for cross-entropy loss); 2) For pretraining data selection, they reliably rank 25 candidate corpora for a target model at roughly $10{,}000\times$ less compute than direct evaluation, pushing the Pareto frontier beyond existing methods; and 3) for training-time forecasting, they extrapolate downstream accuracy across an $18\times$ compute horizon with roughly half the error of existing alternatives. Together, these results suggest that expert trajectories are a broadly useful source of signal for assessing model capabilities, enabling reliable performance forecasting throughout the model development lifecycle.
Foresight-over-Graph: Reasoning Beyond Local Horizons for Knowledge Base Question Answering
Yang Hong ⋅ Yajun Yang ⋅ Xin Wang ⋅ Liping Jing ⋅ Qinghua Hu
Large language models (LLMs) have demonstrated strong capabilities in question answering, yet they still frequently suffer from hallucinations on knowledge-intensive tasks. Knowledge graphs (KGs) provide LLMs with structured, interpretable, and updatable factual grounding, making them a promising external knowledge source for reliable reasoning. However, existing LLM-guided graph reasoning methods typically rely on hop-wise greedy or beam-style pruning during evidence retrieval. Such local decision processes are inherently myopic: evidence that appears weak near the source may become crucial only after deeper graph context is explored, causing answer-critical branches to be discarded prematurely and making the reasoning chain difficult to recover. To address this limitation, we propose Foresight-over-Graph (FoG), a foresight-aware evidence retrieval framework for knowledge base question answering (KBQA). FoG iteratively constructs a question-relevant evidence subgraph and uses far-to-near feedback to guide path exploration, and maintains a compact memory subgraph to support continued exploration. Extensive experiments on widely used KBQA benchmarks demonstrate that FoG achieves state-of-the-art performance, with a particularly large improvement of 16.58\% in Hit on CWQ, while also reducing LLM calls and token usage. Our code is available at \url{https://anonymous.4open.science/r/FoG-7273} .
Forgetting Has Neighbors: Localized Collateral Forgetting in Machine Unlearning
Polina Dolgova ⋅ Sebastian Stich
Machine unlearning aims to remove the influence of selected training examples without full retraining. Standard evaluations often summarize unlearning quality with aggregate metrics, such as accuracy- and forgetting-based scores, which can hide localized failures. We study this failure mode at the example level by comparing the predictions of an unlearned model to those of the model retrained after deletion. We show that this pointwise discrepancy can be highly non-uniform: for gradient-ascent and random-labeling methods, with and without retain-set fine-tuning, it grows with geometric proximity to the forget set. We call this phenomenon localized collateral forgetting. Our analysis identifies a mechanism behind the effect: surrogate targets used during unlearning can be inconsistent with the local prediction structure induced by retraining, and this inconsistency propagates through shared representations to nearby examples. Motivated by this mechanism, we propose Local Teacher Distillation, a simple mitigation strategy that replaces random targets with soft labels from a small teacher trained only on retained neighbors of the forget set. On CIFAR-100 partial-class deletion, this local teacher brings the unlearned model substantially closer to retraining, especially near the forget set, while maintaining competitive aggregate unlearning metrics.
Symmetry-aware learning is typically framed as equivariance to a single group action, but many real-world distribution shifts are typed, partial/non-invertible, and compositional (e.g., modality changes, occlusion, subsampling, intervention chains). These shifts are better modeled by a category of transformations whose objects represent data contexts and whose morphisms represent admissible transformations between contexts. We thus extend symmetry-aware machine learning from groups to such transformation categories; enforcing categorical equivariance on neural architectures gives stronger robustness than group equivariance. We ground this in both theory and experiments. We formulate a universal approximation theorem for category-equivariant architectures and show density in the space of equivariant continuous transformations. We present experiments on compositional OOD shifts, demonstrating that enforcing categorical equivariance yields measurable robustness gains over group-equivariant and non-equivariant baselines.
Foveal-Mamba: Inside-Out Ring Scanning with Recurrent Offset Prediction for Visual State Space Models
Yi-Kuan Hsieh ⋅ Jun-Wei Hsieh ⋅ Kuan-Chuan Peng ⋅ Xin Li ⋅ Chi-Chia Sun ⋅ Yu-Chee Tseng ⋅ Ming-Ching Chang
Visual state space models (SSMs) have emerged as efficient alternatives to attention-based backbones, yet their performance remains sensitive to how visual evidence is sampled and propagated. Deformable visual Mamba methods address fixed scanning by predicting adaptive sampling locations, but offset prediction is still driven primarily by local responses, producing locally plausible samples that are weakly coupled to the SSM's recurrent state evolution, particularly in off-center or multi-object scenes. We present Foveal-Mamba, a scan-path-aware deformable visual Mamba framework that reformulates offset prediction as recurrent scan-path prediction. Two components drive this design. First, an Inside-Out Ring Scan organizes visual tokens along a center-to-periphery path, providing a stable propagation prior that does not assume object centering. Second, a lightweight Mamba module embedded within the offset prediction network forms the Recurrent Scan-Path Offset Prediction Network (RS-OPN), which predicts offsets from hidden states accumulated along the ring path. RS-OPN tightly couples deformable sampling with recurrent state evolution and produces coherent, object-aware sampling trajectories. Foveal-Mamba consistently outperforms strong baselines across multiple benchmarks. On ImageNet-1K, Foveal-Mamba-Tiny achieves 84.3\% Top-1 accuracy, surpassing DAMamba-Tiny by 0.5\% with fewer parameters and FLOPs. On ADE20K semantic segmentation, Tiny/Small variants reach 50.8/51.8 mIoU. On COCO, it achieves 48.9/50.2 box AP and 43.7/44.9 mask AP for object detection and instance segmentation, respectively. Qualitative results further demonstrate earlier object focus and more stable activations in multi-object scenes.
FracEncoder: Towards Adaptive Cognitive Trajectories via Fractional-Order Context Encoding
Lulu Wang ⋅ Shengling Wang ⋅ Anlin Chen ⋅ Ke Chao ⋅ Weicheng Wang
In real-world tasks such as robotic control and autonomous driving, agents often face non-Markovian environments due to incomplete observations. Solving such information-impoverished decision tasks requires reconstructing the context from history. However, existing recurrent encoders have two fundamental limitations: First, the model suffers from a Markovian compression bottleneck, i.e., it only processes observation history as input but neglects thought process history. This leads to the loss of deep semantic cues from early interactions, which are easily overwritten by noise over time due to recursion. Second, it lacks the adaptive capability to adjust the memory span for different tasks. To address these limitations, we propose FracEncoder, a context encoder based on fractional-order neural differential equations. First, FracEncoder leverages the nonlocal property of the fractional-order derivative and uses a power-law kernel to directly incorporate the trajectory of the thought process into the evolution of the current state, effectively overcoming the temporal locality limitations of traditional models. Second, the fractional order $\alpha$ is set as an end-to-end learnable parameter. It explicitly characterizes how strongly the system anchors to its history, so the model can discover a suitable memory span for each task. We evaluate FracEncoder on more than ten benchmarks across three settings, namely partially observable tasks, meta reinforcement learning, and delayed-observation control. FracEncoder consistently matches or outperforms representative baselines, including GPIDE, GRU-ODE, ReSeL, and LRU. On the challenging 8-step delayed-observation task, it reaches average returns of $-511.5 \pm 7.9$ on Pendulum and $732.1 \pm 152.7$ on Hopper. As a plug-in module, FracEncoder also yields stable gains across different reinforcement learning algorithms. Ablation studies further show that the learned $\alpha$ aligns with the intrinsic memory demand of each task, providing an interpretable physical measure of the cognitive complexity of non-Markovian tasks. The code is available at https://anonymous.4open.science/r/FracEncoder-60BA.
Streaming Visual Geometry Transformers such as StreamVGGT enable strong online 3D perception, but their KV-cache grows unbounded over long streams, limiting practical deployment. We study bounded-memory streaming geometry from the perspective of memory organization: unlike language modeling, where useful information can often be compressed at token level, geometry-driven inference relies on coherent and mutually compatible observations across views. Under fixed memory budgets, retaining history as isolated entries can progressively fragment the geometric context needed for stable long-horizon matching and fusion. We therefore propose \textbf{FrameVGGT}, a bounded-memory framework that maintains a fixed-capacity set of complementary memory units for streaming geometry. In our implementation, each unit is instantiated as a frame-wise KV segment summarized by a compact key-space prototype, together with a sparse anchor tier for persistent long-range references. Across long-sequence 3D reconstruction, video depth estimation, and camera pose estimation, FrameVGGT achieves favorable accuracy--memory trade-offs under bounded budgets while maintaining more stable geometry over long streams.
Frank-LoRA: Federated Rank-Aware LoRA for Fine-Tuning Large Models
Saber Malekmohammadi ⋅ Kevin Kuo ⋅ Virginia Smith ⋅ Golnoosh Farnadi
Fine-tuning large-scale foundation models across resource-constrained, distributed clients presents significant computational and communication challenges. While merging Low-Rank Adaptation (LoRA) with Federated Learning (FL) enables rank-based flexibility, the interdependence of the adaptation matrices ($A$ and $B$) often complicates optimization. In this work, we theoretically demonstrate that LoRA fine-tuning with a frozen A matrix acts as a stochastic approximation of full fine-tuning, where batch gradients are subject to a random, rank-dependent perturbation. Motivated by this insight, we propose freezing and periodically resampling $A$ matrices at the client level. This strategy decouples the $A$ and $B$ gradients, diminishes LoRA low-rank constraints and reduces both computational and uplink communication overheads by 50%, allowing for the reallocation of resources toward higher adaptation ranks. Furthermore, we show that in this frozen-$A$ regime, the adaptation rank $r$ directly governs the signal-to-noise ratio (SNR) of the perturbed gradients. Specifically, we find that the SNR increases linearly with $r$, implying that clients’ learning rates must be dynamically scaled with their adaptation ranks for higher utility. Building on these foundations, we introduce Frank-LoRA, a federated rank-aware fine-tuning strategy that utilizes frozen, resampled $A$ matrices and rank-dependent learning rates for clients. We also prove the convergence of the algorithm in FL settings with heterogeneous clients resources. Our experiments across benchmark datasets demonstrate that Frank-LoRA consistently outperforms state-of-the-art baselines with half the overhead on clients.
Frequency‑Aware Flow Matching for Continuous and Consistent Robotic Action Generation
Jianing Guo ⋅ Fangzheng Chen ⋅ Zihao Mao ⋅ WONG L Kenny ⋅ Zhenhong Wu ⋅ Yu Li ⋅ Yishuai Cai ⋅ Yuanpei Chen ⋅ Yikun Ban ⋅ Kai Chen ⋅ DOU QI ⋅ Yaodong Yang ⋅ Xianglong Liu ⋅ Huijie Zhao ⋅ Simin Li
Flow matching has emerged as a standard paradigm for robotic manipulation owing to its strong expressive power for modelling complex, multimodal action distributions, alongside similar approaches like diffusion policy. However, existing methods rely on discretized action chunks, making them brittle to demonstrations collected at heterogeneous control frequencies and prone to temporally inconsistent actions that degrade control stability. In this paper, we propose Frequency-Aware Flow Matching (FAFM), which outputs continuous, temporally consistent actions. To handle heterogeneous frequency input, we transform discrete action sequences into the frequency domain with the discrete cosine transform (DCT), perform flow matching over the resulting coefficients, and reconstruct continuous actions via cosine basis expansion. To generate temporally consistent actions, we regularize the first-order temporal derivative to promote smooth actions. This corresponds to a Sobolev-type constraint that suppresses high-frequency errors and discourages abrupt action changes. Our FAFM is simple, introduces no additional network parameters and applies to standalone flow-matching policies and vision-language action models. Across synthetic toy benchmark, obstacle avoidance, LapGym, and LIBERO, FAFM improves success rates, multimodal expressivity, motion smoothness, convergence speed, robustness to mechanical bias and mixed-frequency input. These gains are consistent when deployed on a real-world Franka robot. Code available at \url{https://anonymous.4open.science/r/FAFM}.
FRESCO: A Novel Consistency Control for Asynchronous Pipeline Parallel Training
Dongyeop Lee ⋅ Namhoon Lee
Asynchronous pipeline parallelism can accelerate distributed model training, but suffers from instability caused by inconsistent gradient computations. We identify a fundamental source of this inconsistency: error signals are coupled to forward-time features, not just parameter states. This means that parameters may evolve as long as they preserve forward-equivalence. Leveraging this insight, we introduce FRESCO, which frames asynchronous updates as a constrained optimization problem: it finds the minimal update modification that acts trivially on the activation subspace. Thus, it serves as a principled generalization beyond the usual synchronous--asynchronous dichotomy: subspace-level protection preserves gradient validity without synchronization stalls, unconstrained-asynchrony instability, or memory overhead that works against pipeline scaling. We demonstrate that FRESCO consistently outperforms representative asynchronous baselines in LLM pre-training, with the gains most pronounced under deep pipelines and high inconsistency levels. Moving beyond the current limitations, FRESCO establishes a new foundation for high-utilization pipeline parallel training, providing a scalable path for large-model development on diverse, resource-constrained systems.
From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation
Bohan Li ⋅ Shuojue Yang ⋅ Baorui Peng ⋅ Xianda Guo ⋅ Erli Zhang ⋅ Youqi Tao ⋅ Junfeng Duan ⋅ Daguang Xu ⋅ DOU QI ⋅ Xin Jin ⋅ Wenjun Zeng ⋅ Hao Zhao ⋅ Yueming Jin
Action-conditioned surgical video generation is a critical yet highly challenging problem for robotic surgery. The core difficulty is that low-dimensional control vectors must precisely govern complex image-space evolution. In this work, we propose a kinematic-to-visual lifting paradigm that converts articulated kinematics into a unified set of five image-aligned control modalities. Building on this representation, we introduce a hierarchically routed visual control framework that selectively activates the most relevant control modalities and motion scales. Instead of uniformly applying all control signals, our model performs hierarchical routing to dynamically allocate conditioning capacity. We further design kinematic-prior-guided routing loss functions to ensure physically meaningful, temporally stable, and efficient expert utilization. To improve efficiency, we propose a budgeted training and inference scheme that leverages routing-induced sparsity. By selectively discarding low-significance control pathways during training and execution, our approach enables adaptive computation that is complementary to standard distillation. We additionally construct a new benchmark with curated articulated annotations, obtained through human-in-the-loop semantic labeling and differentiable pose tracking, providing realistic supervision for action-conditioned surgical video generation. Extensive experiments demonstrate that our method consistently improves action faithfulness, visual fidelity, and cross-domain generalization over diverse baselines. Moreover, our efficient variant achieves substantial reductions in latency while maintaining strong control accuracy.
From Expert Knowledge to Optimization Modeling: Prototype-Based Data Synthesis and Logical Reinforcement Learning
Bo Hu ⋅ Qingcan Kang ⋅ Qian Liu ⋅ Mingxuan Yuan
Existing approaches to optimization modeling using large language models (LLMs) treat tasks through from-scratch construction, generating variables, objectives, and constraints without leveraging reusable structural knowledge. However, in real-world scenarios, expert modelers typically begin by identifying the canonical type or core structure of a problem, followed by iterative refinements based on specific task requirements. To incorporate this expert incremental logic, we propose a novel two-stage framework. In the first stage, we conduct Logic-Anchored Data Synthesis starting from 17 canonical prototypes, preserving critical bottleneck constraints and generating synthetic problem–formulation pairs. These pairs, together with the OR-Instruct dataset, are used for supervised fine-tuning (SFT) to initialize a policy. Separately, for each generated problem, we produce multiple modeling trajectories that follow a prototype-grounded order with variables preceding expressions. These trajectories are ranked according to criteria of accuracy and efficiency, and subsequently used to train a Logical Reward Model (LRM). In the second stage, we introduce Logic-Test-Time Group Relative Policy Optimization (Logic-TGRPO), which leverages the LRM during test-time reinforcement learning to reward accurate prototype identification and disciplined structural adherence while penalizing illogical patterns. Evaluated across seven benchmarks, our 8B parameter model achieves an average accuracy of 79.9%, outperforming comparably scaled baselines and competing effectively with heavyweight multi-agent systems with far fewer parameters. Strong performance on out-of-distribution tasks confirms that the model exhibits genuine incremental reasoning rather than mere memorization. Ablation studies further validate that both stages in our framework are essential.
From Groups to Rings: Causal Evidence for Algebraic Decomposition in Grokked Transformers
Shurui Zheng ⋅ Fanhong Li ⋅ Zhixing Huang ⋅ Zixi Li ⋅ Lei Ji
Prior work has shown that grokked Transformers implement discrete Fourier transforms or group character representations for modular arithmetic. We unify these findings under the Wedderburn--Artin decomposition: using interchange intervention accuracy (IIA), a causal method stronger than ablation, we show that grokked models decompose computation along the Wedderburn components of the target algebra. Across 11 commutative algebras over $\F_p$ ($p = 2,3,5,7$)---including 6 with nontrivial Jacobson radical---all grokked models achieve raw IIA $\geq 0.97$. When models predict incorrectly, errors remain compartmentalized by component ($3.8\times$ chance), confirming that independence is structural, not a byproduct of accuracy. Training dynamics reveal that Wedderburn alignment emerges synchronously with grokking. For non-commutative groups, models exhibit \emph{selective} Wedderburn alignment; non-commutativity raises a capacity threshold that larger models overcome.
From Ideas to Code: Tree-structured Policy Optimization for Automated Algorithm Design with LLMs
Rui Zhang ⋅ Ping Guo ⋅ Liyong Lin ⋅ Zhichao Lu ⋅ Qingfu Zhang
Large Language Model-based Automated Algorithm Design (LLM-AAD) has shown promising results across diverse domains by coupling LLMs with iterative search. Some recent efforts have begun to fine-tune LLMs for algorithm design through reinforcement learning. However, these attempts still rely on outcome-level supervision, leaving intermediate ideas and plans without direct credit, even though they determine the conceptual direction of algorithm design. In this paper, we propose Algorithm Tree Policy Optimization (ATPO), which utilizes tree-structured rollouts to decouple the generation of conceptual ideas from their specific code implementations. By applying group-relative advantage estimation, ATPO assigns credit to both strategic ideas and algorithmic implementations, thereby optimizing the policy across multiple levels of abstraction. Integration with FunSearch, EoH, and OpenEvolve demonstrates improvements over their original counterparts. Notably, the learned policy transfers effectively to unseen problem instances and alternative search methods, indicating that it can improve the generalization of algorithm design policies.
From Infrastructure to Interface, the AI Value Chain Drives LLM Homogenization
Khaoula Chehbouni ⋅ Cléa Chataigner ⋅ Prakhar Ganesh ⋅ Pablo Piantanida ⋅ Jackie CK Cheung ⋅ Golnoosh Farnadi
Outcome homogenization---models producing similar outputs---has drawn increasing attention in the large language model (LLM) community as researchers and practitioners have come to recognize that alignment practices can reduce model diversity and reproduce dominant narratives. However, most approaches address outcome homogenization exclusively through a technical lens. In this position paper, we argue that outcome homogenization is best understood as a value-chain phenomenon shaped by stakeholders, governance structures, resource allocation, model development, and organizational practices. We identify key drivers of homogenization in the AI ecosystem, showing how current practices reinforce dominant paradigms while marginalizing alternative perspectives. We argue that addressing homogenization requires shifting attention from model-level decisions to the broader power dynamics and organizational practices. Through a case study of LLM vendors’ safety alignment practices, we illustrate how these ecosystem-level forces shape model development and deployment. We conclude with a call to action, offering targeted recommendations for the research community.
From Non-Convex to Strongly Convex: Curvature-Adaptive FTPL for Online Optimization
Moses Charikar ⋅ Chirag Pabbaraju ⋅ Ambuj Tewari
Curvature adaptivity is a classical theme in online optimization: for convex Lipschitz losses, adaptive methods interpolate between the optimal $O(\sqrt{T})$ regret for general convex losses and $O(\log T)$ regret under strong convexity. Recent work has shown that Follow-the-Perturbed-Leader (FTPL) achieves optimal $O(\sqrt{T})$ regret even for online non-convex Lipschitz losses, assuming access to an approximate offline-optimization oracle, but these guarantees do not exploit curvature. We show that FTPL can be made curvature-adaptive in the non-convex setting, without knowing in advance how curvature will accumulate over time. Our algorithm replaces the fixed perturbation scale of standard FTPL with a time-varying scale chosen using only past information. We give a simple follow-the-leader tuning rule for this scale and show that it competes, up to constants, with the best choice in hindsight. The resulting method achieves $O(\sqrt{T})$ regret for arbitrary non-convex Lipschitz losses and improves as cumulative curvature grows; with sufficiently accurate oracle calls, it achieves $O(\log T)$ regret when cumulative curvature grows linearly, which includes the classical strongly convex regime. We complement these upper bounds with matching lower bounds for prescribed cumulative-curvature sequences, already for one-dimensional convex losses, showing that the tradeoff between worst-case non-convex regret and curvature-driven fast rates is intrinsic.
From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills
Zisu Huang ⋅ Jingwen Xu ⋅ Yifan Yang ⋅ Ziyang Gong ⋅ Qihao Yang ⋅ Muzhao Tian ⋅ Xiaohua Wang ⋅ Changze Lv ⋅ Xuemei Gao ⋅ Qi Dai ⋅ Bei Liu ⋅ Kai Qiu ⋅ Xue Yang ⋅ Dongdong Chen ⋅ Xiaoqing Zheng ⋅ Chong Luo
Language agents increasingly improve by reusing \emph{skills}---structured procedural artifacts distilled from past experience. In particular, \emph{domain-level} and \emph{model-generated} skills are especially promising. They offer fast adaptation within a domain by encoding domain-specific recurring procedures, and they scale beyond labor-intensive hand-crafting. However, while extraction methods continue to proliferate, understanding remains limited, with no comprehensive study spanning the full skill lifecycle---\textbf{experience generation}, \textbf{skill extraction}, and \textbf{skill consumption}---to ask whether such skills actually work, when they work, and what makes them succeed or fail. To close this gap, we build a utility-grounded evaluation framework that provides systematic experimental results across extractors and target agents, covering five diverse agentic task domains. We find that model-generated skills are beneficial on average but exhibit non-trivial negative transfer, and that neither extractors nor targets behave uniformly. A model can be a strong extractor yet a weak consumer, or vice versa, with skill utility independent of model scale or baseline task strength. To explain these patterns, we then dissect each lifecycle stage in depth, analyzing how experience composition shapes skill quality, what properties characterize useful skills, and how the same skill transfers across different consumers. Finally, we translate these findings into a concrete \emph{meta-skill} that guides skill extraction toward the features tied to actual utility, which consistently improves skill quality across domains and substantially reduces negative transfer.
From Table to Cell: Attention for Better Reasoning with TABALIGN
Tung Sum Thomas Kwok ⋅ Zeyong Zhang ⋅ Xinyu Wang ⋅ Chunhe Wang ⋅ Xiaofeng Lin ⋅ Hanwei Wu ⋅ Lei Ding ⋅ Guang Cheng ⋅ Zhijiang Guo
Multi-step LLM reasoning over structured tables fails because planning and execution share no explicit cell-grounding contract. Existing methods constrain the planner to a left-to-right factorization at odds with table permutation invariance, and score intermediate states by generated content alone, overlooking cell grounding. We conduct a pilot study showing that diffusion language models (DLMs) produce more human-aligned and permutation-stable cell attention on tables than autoregressive models, with a 40.2% median reduction in attention-AUROC variability under row reordering. Motivated by this, we propose TABALIGN, a planned table reasoning framework that operationalizes the contract. TABALIGN pairs a masked DLM planner, whose bidirectional denoising emits plan steps as binary cell masks, with TABATTN, a lightweight verifier trained on 1,600 human-verified attention standards to score each step by its attention overlap with the plan-designated mask. Across eight benchmarks covering table question answering and fact verification, TABALIGN improves average accuracy by 15.76 percentage points over the strongest open-source baseline at comparable 8B-class scale, with a matched-backbone ablation attributing 2.87 percentage points of this gain to the DLM planner over an AR planner on a fixed reasoner. Cleaner DLM plans also accelerate downstream reasoning execution by 44.64%.
From Tokens to Tactics: Adversarial Text Optimization in an Axis-Aligned Rhetorical Strategy Space for Harmful Content Detection
Jinjie Wang ⋅ Shu Jiang ⋅ Hai Zhao
Harmful-content detectors are increasingly deployed in safety-critical settings, but they remain vulnerable to subtle rhetorical reframings that preserve the underlying claim while altering its presentation. Existing adversarial robustness methods are typically token-level or prompt-driven, hindering mechanism-level attribution and targeted repair. We propose Adaptive Rhetorical Adversarial Optimization (ARAO), which operationalizes rhetorical theory as an axis-aligned strategy space with directional and interpretable controls for adversarial rewriting. A planner-rewriter pipeline converts strategy vectors into non-contradictory rewrites, enabling mechanism-level failure attribution and axis-targeted repair. Experiments on misinformation and extremist-rhetoric benchmarks show that ARAO improves robustness by 2.35 ROC-AUC points on average over the strongest baselines on attacked sets, while achieving the best performance on original content.
From Views to Worlds: Active Exploration over 3D Worlds for Vision-Language Models
Qijian Tian ⋅ Jiayu Ying ⋅ Ke Fan ⋅ Lizhuang Ma ⋅ Xin Tan
Vision-language models (VLMs) have achieved remarkable progress in 2D visual understanding tasks, yet the passive view-centric perception paradigm limits their capacity for 3D spatial reasoning. Such reasoning often requires actively acquiring spatial evidence beyond the currently visible views, such as camera poses and cross-view spatial relationships, to infer 3D relationships that are not explicitly represented in 2D images. However, existing methods either lack active exploration or merely provide additional views during inference, without explicitly acquiring such spatial evidence beyond 2D images. In this paper, we propose View-to-World (V2W), an active exploration framework that enables VLMs to actively explore 3D worlds and lifts VLMs from passive view-centric perception to active world-centric spatial reasoning. Given a spatial reasoning task with multi-view images, V2W first reconstructs an explicit 3D world with geometry, camera poses, and language-grounded semantics, and exposes it to VLMs through visual and linguistic interfaces. Through multi-turn interaction with these interfaces, VLMs iteratively acquire spatial evidence from the constructed 3D world, enabling active world-centric spatial reasoning. V2W can be integrated with existing VLMs in a training-free manner and can be further enhanced by optimizing the exploration policy via reinforcement learning. Experiments on spatial mental modeling benchmarks demonstrate that V2W substantially improves VLM baselines in the training-free setting and achieves state-of-the-art performance with the learned exploration policy, highlighting the importance of active exploration for spatial intelligence. The code will be available upon acceptance.
FRUC: Feedforward Dynamic Scene Reconstruction from Uncalibrated Collaborative Driving Views
Yihang Tao ⋅ Yu Guo ⋅ Zhengru Fang ⋅ Haonan An ⋅ Yuguang Fang
We present FRUC, a feed-forward 3D Gaussian splatting framework for dynamic scene reconstruction from uncalibrated collaborative driving views. Existing multi-agent reconstruction frameworks are often hindered by rigid prerequisites, demanding precise spatial calibration and slow per-scene optimization. In this paper, we rethink this task by conceptualizing a distributed multi-vehicle network as a spatio-temporally unstructured ego-centric multi-camera system, where the core challenge lies in enhancing ego-centric occluded geometry through collaboration without degrading the ego's accurately observed visible geometry, while preserving reconstruction efficiency. For efficient reconstruction, FRUC is built upon a visual grounded geometric Transformer backbone to enable one-shot, calibration-free inference from a flexible number of multi-vehicle views. To achieve non-destructive geometric supplementation under uncalibrated cross-agent misalignment, FRUC first introduces an ego-centric causal occlusion field that explicitly derives occlusion evolution as latent priors by modeling agent-wise spatio-temporal correlations. Guided by these occlusion priors, it further formulates cross-agent integration as a deterministic residual denoising process via zero-initialized injection, turning challenging cross-agent fusion into bounded residual learning for robust collaborative blind-spot completion. Through extensive evaluations on the real-world V2XReal and UrbanIng-V2X datasets, FRUC is shown to be a new state-of-the-art for the scene reconstruction of dynamic collaborative driving environments, significantly outperforming existing methods in both rendering quality and efficiency.
Full-Atom Cyclic Peptide Design via Test-Time Scaled Autoregressive Flow Matching
Yunpeng Wang ⋅ Zhonghui Gu ⋅ Runze Ma ⋅ Jie Yang ⋅ Zhanpeng Shi ⋅ Xiaoliang Shi ⋅ Zhijian Wei ⋅ Jun Li ⋅ Shuangjia Zheng
Cyclic peptides are an increasingly important therapeutic modality, offering antibody-like binding specificity within compact and chemically tunable scaffolds. Recent deep generative models have advanced target-specific peptide design, where SE(3)-equivariant diffusion and flow-matching frameworks provide natural inductive biases for peptide-protein complex generation. However, existing SE(3)-based peptide generators mainly focus on linear peptides and do not explicitly control cyclic topology. As shown by recent cyclic peptide benchmarks such as CPSea, current generative baselines still have limited cyclization success and remain weak at controllably generating disulfide- and isopeptide-cyclized peptide distributions. In this work, we propose CPFlow, a test-time scaled autoregressive flow-matching framework for receptor-conditioned cyclic peptide sequence-structure co-design. CPFlow autoregressively unmasks residues and predicts full-atom sequence-structure variables with continuous flow matching. By decomposing generation into stepwise flow updates, CPFlow allows each step to balance target binding with structural closure. At inference time, sequence-geometry guided unmasking and multi-particle test-time scaling explore high-confidence autoregressive trajectories beyond the trained model without additional training. Cyclization-aware relative position encodings further enable stable control over head-tail, disulfide, and isopeptide cyclization. Experiments on the CPSea benchmark show that CPFlow outperforms strong generative baselines across designability (scRMSD ↓ 0.30 Å), self-consistency (scRMSD ↓ 0.15 Å), and binding energy (mean ΔG ↓ 3.0), with 93.4% binding success and 0.623 diversity. The code for CPFlow is available at https://anonymous.4open.science/r/CPFlow-F226.
Functional Gradient Descent with Adaptive Representations
Daniel Csillag ⋅ Rodrigo Schuller ⋅ Pedro Dall’Antonia ⋅ Leonidas Guibas ⋅ Luiz Velho ⋅ Tiago Novello
Functional optimization problems are typically solved by optimizing the parameters of a fixed representation, such as a neural network, resulting in highly nonconvex losses that complicate both training and theoretical analysis. An interesting alternative is functional gradient descent (FGD) --- that is, gradient descent directly in function space --- which benefits from strong convergence results and admits a clean theory. However, FGD is difficult to implement in practice because functional gradients are infinite-dimensional, and thus cannot be fully computed nor stored in memory. Existing implementations therefore rely on fixed approximations, which introduce approximation error. We propose a new, theoretically-grounded FGD algorithm that adapts the representation of the functional gradients over the course of optimization. By explicitly incorporating this approximation into the analysis, we establish convergence to a stationary point (for smooth losses) and to a global minimizer (under smoothness + a Polyak-Łojasiewicz-type condition) regardless of our approximations. To the best of our knowledge, this is the first implementable FGD method with such guarantees in a general setting. We demonstrate the effectiveness of our method on regression, numerical solution of PDEs, and modern computer vision. Across settings, our method consistently outperforms both FGD with fixed approximations and neural network baselines in efficiency and accuracy.
FUTON: Fourier Tensor Network for Implicit Neural Representations
Pooya Ashtari ⋅ Pourya Behmandpoor ⋅ Nikos Deligiannis ⋅ Aleksandra Pizurica
Implicit neural representations (INRs) have emerged as powerful tools for encoding signals, yet dominant MLP-based designs often suffer from slow convergence, overfitting to noise, and poor extrapolation. We introduce FUTON (Fourier Tensor Network), which models signals as generalized Fourier series whose coefficients are parameterized by a low-rank tensor decomposition. FUTON implicitly expresses signals as weighted combinations of orthonormal, separable basis functions, combining complementary inductive biases: Fourier bases capture smoothness and periodicity, while the low-rank parameterization enforces low-dimensional spectral structure. We provide theoretical guarantees through a universal approximation theorem and derive an inference algorithm with empirically linear complexity in both the spectral resolution and the input dimension. On image and volume representation, FUTON consistently outperforms state-of-the-art MLP-based INRs while training 2–5× faster.
GCD: Correcting Hidden-State Bias in Off-Policy Agentic RL
Changyuan Chen ⋅ Jianyu Xiang ⋅ Jiasheng Luo ⋅ Ziye Wang ⋅ NALLAPPAN GUNASEKARAN
Off-policy reinforcement learning (RL) for large language models bottlenecks at the rollout stage: every parameter update invalidates the key-value (KV) cache of in-flight agent trajectories, forcing either expensive recomputation or staleness that destabilizes training. Recent advances such as M2PO and asynchronous RLHF tolerate token-level staleness by reweighting the policy gradient with second-moment importance corrections, but they implicitly assume that a stale rollout is sampled from a well-defined behavior policy $\pi_{\theta_{\text{old}}}$. We show that this assumption is silently violated by every modern partial-rollout system: once a stale KV cache is reused under updated parameters $\theta_{\text{new}}$, the resulting token distribution is a hybrid $\pi_{\text{hyb}}$ whose attention projections consume stale keys and values while its MLP, output, and gating projections use the new weights. The token-level importance weight $w_t = \pi_{\theta_{\text{new}}}(y_t \mid x_{
GenAI Evaluation Results are Largely Artifacts of Evaluation Design Choices: Evidence from Audit Studies of Resume Screening
Nicole Meister ⋅ Hannah Cha ⋅ Mohammed Alsobay ⋅ Alexandra Chouldechova ⋅ Alex Dow ⋅ Hanna Wallach ⋅ Jennifer Wortman Vaughan ⋅ Carlos Guestrin ⋅ Tatsunori Hashimoto ⋅ Solon Barocas
Generative AI is designed to support many tasks, many ways of framing those tasks, and seemingly endless variation in task inputs. But the same flexibility that makes these systems powerful also makes them difficult to evaluate reliably. This flexibility exacerbates the numerous, but often underrecognized and consequential, \emph{researcher degrees of freedom} where even seemingly minor experiment design decisions can substantially affect findings. We demonstrate this problem through an analysis of 36 papers on gender and racial bias in generative AI-assisted resume screening. Although these studies address the same broad research question, they reach strikingly divergent conclusions. To explain this divergence, we adapt recent work on integrative experiment design from the social sciences, and decompose the dimensions along which the studies vary. We then run 1,134 counterfactual experiments by combinatorially varying the experimental modules (names, job-resume pairings, and models) across three baseline experimental frameworks. We find the results prove remarkably brittle: a different design choice flips the conclusion of the reproduced experiment around 22\% of the time. For example, seemingly inconsequential design choices, such as the phrasing of a job description, can reverse the observed conclusion. However, this instability is not always purely stochastic. In some settings, we can build a model that predicts the result of unseen counterfactual experiments with a mean held-out predictive $R^2$ of 0.807. This suggests that despite sensitivities to difference in evaluation design, there exists a learnable structure.
Generalization Without Compression Penalty: A Stability Analysis of Error Feedback
Yifei Liang ⋅ Peng Wang ⋅ Yan Sun ⋅ Yingqi Liu ⋅ Xiaoyan Wang ⋅ XIAOCHUN CAO ⋅ Li Shen
Error Feedback (EF) has become a core mechanism for aggressive gradient compression in communication-efficient distributed training, and is now used in geo-distributed frameworks such as DiLoCo. While existing theory mainly explains its optimization convergence, the generalization behavior of EF remains poorly understood, especially under nonlinear biased compression and distributed local updates. We present the first generalization analysis of EF for both single-node EF-SGD and DiLoCo-EF through an optimization-driven on-average model stability framework, without imposing bounded-gradient assumptions. Technically, we develop a buffer-aware Lyapunov argument that tracks the coupled dynamics of the model and error buffer, while co-coercivity absorbs same-sample gradient-difference energy into the optimization trajectory. For convex losses, our bounds show that compression-induced stability terms are transient, decaying as $\mathcal{O}(\sqrt{b_\alpha}\,T^{-3/4}+b_\alpha T^{-1})$, and therefore do not create a non-vanishing generalization floor beyond the dominant $\mathcal{O}(T^{-1/2})$ optimization term. In the DiLoCo setting, worker averaging further suppresses compression perturbations, and a local stepsize $\eta_l=\Theta(1/(H\sqrt{T}))$ controls the drift from $H$ local steps, yielding the leading $\mathcal{O}(1/\sqrt{MHT})$ excess-risk scaling. Experiments on convex benchmarks corroborate the theory, showing that Top-$k$, Random-$k$, and quantization with matched contraction factor $\alpha$ exhibit nearly indistinguishable generalization trajectories.
Generating in the Limit with Infinitely Many Hallucinations
Irene Strauss ⋅ Alexandra Butoi ⋅ Ryan Cotterell
The classic paradigm of language identification in the limit models learning as a game between an adversary, who reveals strings from an unknown target language, and a learner tasked with identifying that language. The recently introduced framework of language generation in the limit shifted the objective to better reflect modern language modeling, requiring the learner to produce valid, unseen strings from the target language. Related work highlighted a fundamental tension: a broad coverage of the target often comes at the cost of validity. We introduce a new notion of precision and recast this problem as the classic recall--precision trade-off. We analyze generation in the limit under varying constraints on enumeration, novelty, and validity, aimed at reflecting settings closer to those encountered by large language models. A key contribution is our analysis of learners that are not eventually valid: we allow infinitely many mistakes, provided their frequency tends to zero so that precision remains $1$. We show this relaxation can strictly increase recall when the adversary permanently withholds a large portion of the target language. We also study a continuous relaxation of the novelty constraint, requiring only that a fixed fraction of outputs are novel. Taken together, our results move toward a more realistic model of language generation where occasional errors and repetitions are unavoidable, but their rates are controlled.
Generation-Drift-Guided Block Pruning for Large Language Models
ZHIQI HUANG ⋅ Heedong Kim ⋅ ZHANG LINTONG ⋅ Seong-Whan Lee
Pruning reduces the model size and inference cost of large language models (LLMs), yet preserving performance across diverse downstream tasks remains challenging. Existing structured pruning methods commonly rely on local discrepancy measures or teacher-forced predictive criteria, such as Block Influence (BI) and perplexity (PPL). While these methods are effective on classification benchmarks, the resulting pruned models often degrade severely on generative reasoning tasks. We study redundancy in Transformer-based models and observe that pruning-induced generation drift is strongly correlated with the functional importance of modules, suggesting a more faithful criterion for identifying redundant modules. Based on this insight, we propose GDPruner, a calibration-free, generation-drift-guided block pruning method. GDPruner constructs lightweight self-generated probe trajectories, estimates module importance using tail-position generation drift, and removes redundant modules through adaptive search. Extensive experiments show that GDPruner surpasses state-of-the-art structured pruning baselines while better preserving overall performance and complex generative reasoning ability, offering a promising direction for generation-robust LLM deployment.
We propose a new framework for generative modeling based on a discrete-time stochastic control formulation of measure transport. Adapting classic results from control theory, we formulate our problem as a linear program whose dual variables correspond to the \emph{optimal value function} of the control problem, which directly encodes the optimal control policy. Exploiting this LP formulation, we develop an efficient simulation-free primal-dual algorithm for computing approximately optimal value functions and the associated \emph{value-driven transport} (VDT) policies which approximate the true optimal policy. We show that well-trained VDT policies enjoy numerous favorable properties in comparison with other state-of-the-art methods based on flows, diffusions, or Schr\"odinger bridges: they lead to straight transport paths which can be simulated quickly and robustly, and can be enhanced in all the same ways as diffusion and flow-based models (e.g., conditional generation, classifier-free guidance, unpaired image-to-image translation are all easy to incorporate). We establish several foundational links with the aforementioned models and evaluate our methodology in an extensive range of experiments, all of which indicate strong performance.
Generative Structure from Motion with Native 3D Diffusion
Haobin Duan ⋅ Binbin Huang ⋅ Mochu Xiang ⋅ Yiqun Zhao ⋅ Zibo Zhao ⋅ Yi Ma ⋅ Shenghua Gao
Recovering an object's 3D structure and camera motion from a few uncalibrated views is a fundamental problem in computer vision. Classical structure from motion methods recover 3D geometry consistent with the input images but cannot recover unseen regions; native generation methods produce plausible 3D objects which are related but not necessarily faithful to the input images. Methods simply concanetating reconstruction and generation allow error to cascade across the two stages, instead of producing a 3D object consistent with both the distribution and the input. We propose GenSfM, which integrates image-based 3D reconstruction with a native 3D diffusion model. In particular, we use a multi-modal diffusion transformer to jointly denoise the shape latent and per-view pose latents under shared self-attention, grounding generated geometry in the observed views while using object-level priors to infer unobserved regions. A compact 3D-native UV-volume tokenization keeps the joint sequence small as views are added, and a pose-aware stage refines geometry and texture by sampling per-voxel image features at the recovered camera poses. On Toys4k and GSO, for $4$ uncalibrated views, GenSfM improves novel-view PSNR by $3.2$--$4.5$~dB and reduces median Chamfer distance by $2.8$--$3.8\times$ over the strongest 3D-generation baseline, while recovering input-view poses over $96\%$ Acc@$5^\circ$ without any external pose-estimation backbone, post-hoc optimization, or known intrinsics.
GeneZip: Region-Aware Compression for Long Context DNA Modeling
Jianan Zhao ⋅ Xixian Liu ⋅ Zhihao Zhan ⋅ XINYU YUAN ⋅ Hongyu Guo ⋅ Jian Tang
Long-context DNA models are limited by token-mixing cost and by how compression allocates representational budget across the genome. Existing approaches operate close to base-pair resolution, apply fixed downsampling, or learn content-dependent chunks without an explicit genomic budget, making long-context pretraining expensive and difficult to control. We introduce GeneZip, a region-aware DNA compression framework that combines H-Net-style dynamic routing with a Region-Aware Ratio (RAR) objective and bounded routing. GeneZip uses static gene-structure annotations during compression training to specify region-wise base-pairs-per-token (BPT) targets; at inference time, it compresses raw unseen DNA without annotations. GeneZip provides three main benefits. First, it is effective: GeneZip variants achieve the best validation PPL among encoder-based compressors, with GeneZip-70M already operating at 137.6 BPT, and across four reproducible DNALongBench tasks---contact map prediction, eQTL prediction, enhancer--target gene prediction, and transcription-initiation signal prediction---GeneZip obtains the best average rank among compared sequence models. Second, it is redundancy-aware: a post-hoc RepeatMasker/TRF analysis shows that, without repeat supervision, GeneZip assigns higher local BPT to TE-derived interspersed repeats and tandem repeats, two major classes of repetitive DNA sequence redundancy. Third, it is efficient: by reducing the effective token-mixing length, GeneZip enables longer-context and larger-capacity pretraining, including 128K-context and 636M-parameter variants on a single A100 80GB GPU, and fine-tunes the eQTL task 50.4$\times$ faster than JanusDNA (50 vs. 2520 minutes). These results establish GeneZip as an effective, redundancy-aware, and efficient compression interface for long-context DNA modeling.
GenRM-Flow: Generators are Process-aware Reward Models in Flow Matching
Siming Fu ⋅ Zheming Fu ⋅ Ruizhe He ⋅ Zeyue Xue ⋅ Fengkai Liu ⋅ Yufei Jiao ⋅ Jie Huang ⋅ Xiaoxiao Ma ⋅ Xuan Yang ⋅ Lihui Luo
Aligning video generation models with human preference depends critically on the quality of the reward model that scores generated videos. Current video reward models, however, source their reward signal from a model's representational rather than generative capacity: they either repurpose vision-language models as pixel-space scorers, or train auxiliary scoring heads on top of frozen generator features. The generator's actual ability to predict velocities, denoise latents, and produce videos plays no role in scoring its own generations. We argue that this overlooks a structural property of offline preference learning. Preference optimization objective approximating an offline preference dataset mathematically targets a strictly defined KL-constrained optimal policy. By isolating the reward term from this analytical solution and mapping the exact generation log-likelihood to its flow-matching surrogate, the empirical reward emerges deterministically as the relative difference in continuous-time velocity-prediction errors. We formalize this equivalence into GenRM-Flow, a framework where a preference-tuned video generator inherently operates as a process-aware reward model directly in the noisy latent space. We demonstrate that this evaluation mechanism is structurally universal: distinct fine-tuning objectives, whether contrastive like DPO and IPO or non-contrastive like RWR, are merely alternative optimization pathways that share a unified empirical reward readout. Furthermore, the extracted process reward seamlessly substitutes external VLM scorers in Flow-GRPO. Operating purely on native latent outputs without auxiliary RL machinery, this unified generative reward drives measurable improvements in downstream video generation.
GeoG2U-Bench: When Does Generation Help Understanding in Ultra-High-Resolution Remote Sensing?
Fengxiang Wang ⋅ Yueying Li ⋅ Mingshuo Chen ⋅ Boya Miao ⋅ Qiuyang Yu ⋅ Yajie Yang ⋅ Mingzhen Xu ⋅ Luqing Luo ⋅ Haonan Guo ⋅ Hongda Sun ⋅ Yulin Wang ⋅ Jun Song ⋅ Jing Zhang ⋅ Long Lan ⋅ Wenjing Yang
Ultra-high-resolution (UHR) remote-sensing imagery captures fine-grained spatial details essential for city-scale Earth observation, yet it exposes a fundamental limitation of current multimodal large language models (MLLMs): native-resolution inputs exceed model capacity, while aggressive downsampling erases the very small objects that many questions hinge on. Generation-understanding unified multimodal models offer a promising remedy, since they can produce intermediate visual observations before answering. However, whether such capabilities genuinely improve UHR remote-sensing understanding remains an open question. Existing UHR benchmarks cannot adjudicate it: they are mostly built from mature datasets covering only a few regions, their image scales remain below real city-level satellite scenes, and they do not systematically identify which tasks benefit from intermediate visual generation. We introduce GeoG2U-Bench, a global benchmark designed to evaluate Generation-to-Understanding (G2U) in UHR remote sensing. GeoG2U-Bench spans 300 cities worldwide, with every raw satellite scene reaching $20{,}000 \times 20{,}000$ pixels, matching the scale at which UHR analysis actually occurs. It comprises 20 subtasks across five capability dimensions and roughly 3,000 expert-verified samples, each requiring models to produce auditable intermediate artifacts such as zoom-in crops, annotated local views, cross-region relation maps, and temporal rewind images. Beyond data, GeoG2U-Bench contributes a dual-protocol evaluation: Direct (answer from the raw input) versus Generate-then-Answer (generate an intermediate visual artifact, then answer), which isolates and quantifies the contribution of generation to downstream understanding. Experiments across several model families reveal that generation is not uniformly beneficial: unfaithful or off-task artifacts can actively degrade performance, whereas reliable, task-relevant artifacts yield consistent gains in cross-scale localization, multi-region comparison, temporal reasoning, and spatial layout understanding. These results show that visual generation can support UHR remote-sensing understanding, but only when it provides reliable visual evidence for the task.
Geometric Analysis of Neural Regression Collapse via Intrinsic Dimension
George Andriopoulos ⋅ Zixuan Dong ⋅ Bimarsha Adhikari ⋅ Keith Ross
Neural multivariate regression underpins a wide range of domains such as control, robotics, and finance, yet the geometry of its learned representations remains poorly characterized. While neural collapse has been shown to benefit generalization in classification, we find that analogous collapse in regression consistently degrades performance. To explain this contrast, we analyze models through the lens of intrinsic dimension. Across control tasks and synthetic datasets, we estimate the intrinsic dimension of last-layer features ($ID_H$) and compare it with that of the regression targets ($ID_Y$). Collapsed models exhibit $ID_H < ID_Y$, leading to over-compression and poor generalization, whereas non-collapsed models typically maintain $ID_H > ID_Y$. For the non-collapsed models, performance with respect to $ID_H$ depends on the data quantity and noise levels. From these observations, we identify two regimes—over-compressed and under-compressed—that determine when expanding or reducing feature dimensionality improves performance. Our results provide new geometric insights into neural regression and suggest practical strategies for enhancing generalization.
Geometric Prompt-Trajectory Planning for Test-Time Scaling
Zhengqi Pei ⋅ Anran Zhang ⋅ Qingming Huang ⋅ Shuhui Wang
Test-time scaling (TTS) improves large language model (LLM) reasoning by spending additional inference-time compute, most commonly through width-heavy repeated sampling and aggregation. However, this strategy scales cost almost linearly with the number of calls and can remain brittle when the sampled trajectories share the same failure mode. We introduce Maze, a test-time scaling framework that shifts computation away from repeated target-model calls and toward LLM-free planning over ordered few-shot exemplars. The central idea is to treat an ordered exemplar sequence as a prompt trajectory embedded in a latent geometric space. Instead of asking the LLM for many independent attempts, Maze first ranks candidate prompt trajectories using cached encoder representations and a path scorer, ${\it i.e.}$, a lightweight learnable metric, then spends one call (or a few calls) only on the most promising trajectories. This design preserves black-box deployability, since the planner is decoupled from the target LLM and can be implemented with a separate frozen encoder. Across various challenging reasoning benchmarks, single-call Maze consistently outperforms strong single-call prompt-optimization baselines, and MultiMaze recovers a large fraction of width-heavy TTS gains with substantially fewer computational cost measured by number of LLM calls and latency. We further present a theoretical perspective showing when prompt-path ranking can recover much of the benefit of best-of-$N$ style scaling: if high-utility prompt trajectories are sufficiently rankable and sufficiently covered by the explored pool, a deeper planner-side search can substitute for a substantial amount of width-heavy sampling.
Geometry-Aware Post-Hoc Uncertainty Quantification in Operator Learning
Oriol Vendrell-Gallart ⋅ Nima Negarandeh ⋅ Ramin Bostanabad
Neural operators provide fast surrogates for PDEs but their deterministic predictions limit their use in tasks requiring uncertainty quantification (UQ), especially under geometric variability. Existing approaches primarily model uncertainty in network parameters, largely overlooking the geometry-aware representations learned by the operator itself. We propose REEF-GP (Residual on Embedded Features Gaussian Process), a post-hoc UQ framework that fits a GP to the residuals of a frozen neural operator whose internal embeddings define the kernel feature space. Rather than learning a separate feature map, REEF-GP adapts the operator’s intrinsic coordinate-feature representations to construct geometry-aware uncertainties. To ensure stability and scalability on unstructured domains, REEF-GP incorporates spectral-normalized projections, heteroscedastic geometry-aware noise, and efficient subset-based training that avoids restrictive low-rank approximations. Across five PDE benchmarks with varying geometries, REEF-GP preserves predictive accuracy while achieving calibrated uncertainty estimates competitive with deep ensembles but at a fraction of their cost. Our approach remains robust under geometric distribution shift, with uncertainty concentrating in physically meaningful regions (e.g., shock fronts). Our results demonstrate that accurate and scalable post-hoc UQ for neural operators can be achieved directly in their learned feature space, offering a practical alternative to parameter-centric approaches.
Geometry-Aware Self-Supervised Electrophysiology Representation Learning
Subhrajit Dey ⋅ Ping-Jung Lu ⋅ Visweswar Parupudi ⋅ Malhar Patel ⋅ Margaret-Estefani Conde Paredes ⋅ Mihály Vöröslakos ⋅ Anna Maslarova ⋅ Gyorgy Buzsaki ⋅ Saurabh Vyas ⋅ Erdem Varol
Precise localization of recording electrodes is fundamental to systems neuroscience; yet, current methods depend on labor-intensive post-hoc histology, which is incompatible with real-time feedback or chronic implants. We test the hypothesis that the electrical signals being recorded themselves carry sufficient anatomical information to localize their recording site in 3D atlas coordinates, without the need for histology. To probe this, we introduce a self-supervised framework that learns geometry-aware representations of single-channel local field potentials (LFP), leveraging probe channel geometry as a self-supervisory signal. We compare speech-based SSL objectives (e.g., masked prediction) against geometry-aware objectives that constrain the embedding space to mirror physical channel layout along the probe. These representations outperform masked-prediction baselines, and the gains compound at larger data scales. We test downstream analysis with 3D coordinate regression and brain region classification, benchmarking three SSL backbones (Wav2Vec 2.0, Whisper, and Data2Vec) against supervised baselines (AnyNet, ViT, and classical spectral classifiers) across datasets, laboratories, species, and probe technologies. Our method achieves state-of-the-art results in brain region classification and 3D coordinate regression from raw single-channel LFP, outperforming all baselines, affirming our hypothesis.
Geometry-Aware Subspace Perturbation for Heterogeneous Federated Learning
Xiangtao Zhang ⋅ Hailong Yan ⋅ Obed Irihose ⋅ Joey Tianyi Zhou ⋅ Chee Seng Chan ⋅ Ce Zhu ⋅ Le Zhang
We propose a geometry-aware perturbation framework that explicitly models the low-rank, anisotropic, and evolving structure of gradient dynamics to improve generalization in heterogeneous federated learning (HFL). Our approach is built upon three key components. First, we represent perturbations within a low-dimensional subspace that captures the dominant directions of gradient trajectories. Second, we generate perturbations through a structure-aware sampling strategy that aligns with the covariance of projected gradients, enabling distribution-aware exploration. Third, we introduce an adaptive mechanism that dynamically adjusts both the magnitude and the structure of perturbations to match the evolving optimization process. Extensive experiments demonstrate that the proposed method consistently improves generalization across diverse HFL benchmarks, achieving up to a $4.49\%$ improvement on the $\textit{Office10}$ dataset with minimal computational overhead.
GIST: Gauge-Invariant Spectral Transformers for Scalable Graph Neural Operators
Mattia Rigotti ⋅ Nicholas Thumiger ⋅ Thomas Frick
Neural operators on irregular meshes face a fundamental tension. Spectral positional encodings, the natural choice for capturing geometry, require cubic-complexity eigendecomposition and inadvertently break gauge invariance through numerical solver artifacts; existing efficient approximations sacrifice gauge symmetry by design. Both failure modes break discretization invariance: models fail to transfer across mesh resolutions of the same domain, and similarly across different graphs of related structure in inductive settings. We propose GIST (Gauge-Invariant Spectral Transformer), a scalable neural operator that resolves this tension by restricting attention to pairwise inner products of efficient approximate spectral embeddings. We prove these inner products estimate an exactly gauge-invariant graph kernel at end-to-end $\mathcal{O}(N)$ complexity, and establish a formal connection between gauge invariance and discretization-invariant learning with bounded mismatch error. To our knowledge, GIST is the first scalable graph neural operator with a provable discretization-mismatch bound. Empirically, GIST sets state-of-the-art on the AirfRANS, ShapeNet-Car, DrivAerNet, and DrivAerNet++ mesh benchmarks (up to 750K nodes), and additionally matches strong baselines on standard graph benchmarks (e.g., 99.50\% micro-F1 on PPI).
Glob3R: Global Structure-from-Motion with 3D Foundation Models
Junyuan Deng ⋅ Heng Li ⋅ Kejie Qiu ⋅ Lingteng Qiu ⋅ Rui Peng ⋅ Weichao Shen ⋅ Weihao Yuan ⋅ Siyu Zhu ⋅ Zilong Dong ⋅ Ping Tan
Recent 3D geometric foundation models, such as VGGT, provide robust feed-forward 3D reconstruction by directly predicting camera poses and 3D scene points from input images. However, their results remain inaccurate, and scaling them to long sequences or large unordered image sets typically requires chunk-wise processing, which can introduce drift and inconsistency. We present Glob3R, a global SfM-style reconstruction built on 3D foundation models. Our key idea is to explicitly optimize feed-forward geometric predictions. To this end, we augment a frozen Pi3X backbone with a lightweight dense matching head that predicts image warps between selected reference frames and neighboring views. These dense warps are converted into sparse but reliable multi-view feature tracks, which provide correspondence constraints for global optimization. We further introduce a keyframe-based sliding-window association strategy that propagates tracks and relative poses across overlapping windows, enabling scalable reconstruction. Finally, we perform global motion averaging and bundle adjustment to refine camera poses, reduce scale inconsistencies, and recover dense scene geometry. Extensive experiments on indoor, outdoor, large-scale driving, and unordered SfM benchmarks demonstrate that Glob3R achieves robust and accurate reconstruction. It consistently improves over feed-forward foundation-model baselines and recent scalable reconstruction methods, while being more robust than classical SfM pipelines. The refined poses also lead to higher-quality neural rendering, validating the benefit of combining foundation-model priors with global geometric optimization.
Global linear convergence of entropy-regularized softmax policy gradient beyond tabular MDPs
Ziyue Chen ⋅ David Siska ⋅ Lukasz Szpruch
We study the global convergence of policy gradient for infinite-horizon entropy-regularized Markov decision processes (MDPs) with continuous state and action spaces. We consider log-linear softmax policies with linear function approximation, which extend the tabular softmax parameterization while retaining a tractable policy class. Under $Q^\pi_\tau$-realizability for the regularized state-action value function, we first establish a non-uniform Polyak--Łojasiewicz (PŁ) inequality. The non-uniformity arises through degeneracy of constants associated with the policy geometry, namely the Fisher information matrix or an uncentered feature covariance matrix. We then identify two feature regimes under which this non-uniform constant can be bounded along the gradient flow. For full-affine-span features, we prove radial unboundedness of the KL regularizer and show that the smallest eigenvalue of the Fisher information matrix remains bounded below by an initialization-dependent positive constant. For simplex-valued features, we prove an analogous radial unboundedness result in the subspace orthogonal to the all-ones vector and obtain a uniform lower bound for the smallest eigenvalue of the uncentered covariance matrix. These results imply global linear convergence of the regularized objective along the gradient flow, i.e. suboptimality decaying as $\mathcal{O}(e^{-Ct})$ for some $C>0$. Our analysis extends the global convergence theory of entropy-regularized softmax policy gradient beyond the tabular setting of Agarwal et al. [2020], Bhandari and Russo [2024], Mei et al. [2020].
GlyphAnchor: Enhancing Visual Text Rendering via Position-Anchored Glyph Priors
Qiang Xiang ⋅ Shuang Sun ⋅ Binglei Li ⋅ Yibo Chen ⋅ Xu Tang ⋅ Yao Hu ⋅ Junping Zhang
Rendering accurate text remains difficult for image generation and editing models, especially when the target contains long, complex, and densely arranged text or rare characters. Existing approaches either improve native text rendering through stronger backbones and data-centric training without explicit glyph priors, or incorporate glyph priors through specialized designs that remain insufficiently accurate and robust under challenging scenarios. We introduce GlyphAnchor, a novel text-rendering enhancement method for both text-to-image and image-editing diffusion transformer models. GlyphAnchor enhances the backbone with lightweight glyph patch conditions whose positions are anchored to the target image through the model's native positional encoding. We train this capability with staged supervised finetuning and further refine it with text-aware post-training to improve robustness. We also introduce InfoTextBench, a benchmark for evaluating text-rich visual text rendering in both generation and editing settings. Experiments across multiple backbones and benchmarks, including long, complex, and densely arranged text and rare character scenarios, show that GlyphAnchor consistently improves text fidelity while preserving overall image quality.
GMO-E²DIT: Grounded Multi-Operation Editing for E-Commerce Images
Zipeng Guo ⋅ Xiaoan Liu ⋅ Lichen Ma ⋅ Cheng Wang ⋅ Yu He ⋅ Xiaolong Fu ⋅ Jingling Fu ⋅ Xinyuan Shan ⋅ Shaojie Guo ⋅ Luohang Liu ⋅ Junshi Huang ⋅ Yan Li
Real-world e-commerce image editing often requires multiple, localized, and auditable operations rather than global restyling. This compositional nature poses a dual challenge: models must precisely apply all requested edits to the correct regions while preserving unmodified content, even under ambiguous instructions. Existing one-shot editors conflate intent resolution, spatial grounding, and synthesis into a single step, frequently resulting in partial execution failures, which is unacceptable for commercial scenarios. To address this, we introduce GMO-E²DIT, an agentic editing framework that couples a Vision-Language Model (VLM) with a mask-conditioned image editor to tackle structured multi-turn task completion. Given an underspecified instruction, the VLM agent constructs a region-grounded edit agenda, effectively decoupling cognitive reasoning from generative rendering. The framework then executes sub-programs via operation-aware masks and references, utilizing a reflection-driven loop to inspect intermediate results and determine the subsequent state. This iterative mechanism reliably preserves safe partial progress, retries unfinished operations, and recovers from errors. Furthermore, we develop a unified data pipeline providing aligned supervision for planning, execution, and reflection, alongside EComEditBench, a comprehensive benchmark for instruction-driven evaluation. Extensive experiments demonstrate that GMO-E²DIT achieves competitive performance compared to strong closed-source models, yielding superior instruction accuracy and edit fidelity over existing baselines.
GMOS: Grounding Moving Object Segmentation in 3D Space and Time
Junyu Xie ⋅ Tengda Han ⋅ Weidi Xie ⋅ Andrew Zisserman
Moving Object Segmentation (MOS) aims to discover, segment, and track objects that move independently of the camera. Current MOS methods, however, exhibit two fundamental limitations: they rely on pre-computed 2D auxiliary modalities such as optical flow or point trajectories that lack 3D geometric information, and they treat motion as a sequence-level attribute, overlooking the instantaneous motion state of each object. We address both by grounding MOS in 3D space and time, and propose GMOS, a framework that operates directly on RGB video to produce 3D-aware, temporally fine-grained segmentation of multiple moving objects, alongside a foreground--background variant GMOS-S for faster deployment. To support training and evaluation in this regime, we curate GMOS-2K, a dataset of $2{,}210$ real-world videos with per-object temporal motion annotations drawn from five established Video Object Segmentation (VOS) benchmarks, and formalise MOS-I ("I" for *instantaneous*), a temporally fine-grained evaluation protocol with three complementary metrics. GMOS achieves state-of-the-art results across MOS, MOS-I, and Unsupervised VOS benchmarks, while running significantly faster than prior multi-object MOS methods and supporting online inference for streaming deployment.
Goal-Conditioned Supervised Learning for Multi-Objective Recommendation
Shijun Li ⋅ Hilaf Hasson ⋅ Jing Hu ⋅ Joydeep Ghosh
Multi-objective learning endeavors to concurrently optimize multiple objectives using a single model, aiming to achieve high and balanced performance across diverse objectives. However, this often entails a complex optimization problem to balance the learning of potentially conflicting objectives, leading to solutions with higher memory requirements and computational complexity. This paper introduces a Multi-Objective Goal-Conditioned Supervised Learning (MOGCSL) framework for automatically learning to achieve multiple objectives from offline sequential data. MOGCSL extends the conventional GCSL method to multi-objective scenarios by redefining goals from one-dimensional scalars to multi-dimensional vectors. It benefits from naturally eliminating the need for complex architectures and optimization constraints. Moreover, MOGCSL inherently disentangles uninformative or noisy training instances that fail to achieve desirable long-term rewards across multiple objectives. We also introduce a novel goal-selection algorithm for MOGCSL to model and identify desired and achievable goals for inference. In this paper, we focus on its application to the next action prediction problem in commercial-grade recommender systems. In this context, any viable solution needs to be reasonably scalable and also be robust to large amounts of noisy data that is characteristic of this application space. We show that MOGCSL performs admirably on both counts by extensive experiments. Also, analysis and experiments are included to explain its strength in discounting the noisier portions of training data in recommender systems with multiple objectives. Code is available at: https://anonymous.4open.science/r/MOGCSL-D7A2.
GOLD: Geometric Optimized Latent Diffusion for Structure-Aware RNA Inverse Folding
Qi Si ⋅ Xuyang Liu ⋅ Penglei Wang ⋅ Shuo Su ⋅ LIMEI HAN ⋅ Mengzhu Li ⋅ Xin Guo ⋅ Yuan Qi ⋅ Yuan Cheng
Designing RNA sequences that fold into specific 3D structures is a fundamental challenge in therapeutics. Current RNA inverse folding methods typically condition on 3D backbone coordinates but the generation is limited to 1D discrete sequences. This ignores the continuous side-chain orientations, which are essential for overall structural stability. To address this issue, we explicitly introduce side-chain geometry into the RNA generative process. Unlike co-design frameworks that diffuse across separate raw spaces, we propose Geometric Optimized Latent Diffusion (\textbf{GOLD}), which project sequence and structure into a unified latent manifold. GOLD employs a novel GeoVAE to elegantly integrate discrete RNA sequences with continuous side-chain angles. By mapping these modalities into a single, smooth latent manifold, GOLD naturally facilitates a synergistic co-diffusion process. To enhance structural validity, we further introduce a training-free, gradient-guided sampling mechanism. By decoding the generated sequence probabilities and side-chain geometries into structures during the sampling process, we compute analytical gradients from physical constraints. This dynamic feedback actively guides the generated sequence to a physically stable state. Extensive experiments validate the superiority of GOLD, achieving state-of-the-art performance across sequence recovery, secondary structure validity, and 3D foldability metrics.
GoT: A Game-Theoretic Approach to the Game of Twenty Questions
Langyuan Cui ⋅ Chun Kai Ling ⋅ Hwee Tou Ng
The Game of Twenty Questions is a classical problem where a questioner must identify a hidden item by asking binary questions. Although simple in form, the game is strategically challenging, especially when the questions are expressed in natural language. It is also a useful abstraction for realistic information-seeking tasks, such as medical diagnosis and troubleshooting. Existing approaches often rely on simplifying assumptions that degrade worst-case performance. This is an issue with serious implications in high-stakes applications. In this work, we introduce Strategic Language Search (SLS), an adversarial counterpart of Twenty Questions, and formalize it and its variants as two-player zero-sum extensive-form games. We then propose Game of Thought (GoT), a framework that applies game-theoretic techniques to approximate a Nash equilibrium (NE) strategy for the restricted variant of the game. Empirical results demonstrate that our approach consistently improves worst-case performance compared to (1) direct prompting-based methods and (2) heuristic-guided search methods across all tested settings.
GPA: General Principled Framework for Linearizing Softmax Attention via KV Cache Approximation
wang ning ⋅ Zekun Li ⋅ Tongxin Bai ⋅ Man Yao ⋅ Chunlei Men ⋅ Guoqi Li
Transformers excel at sequence modeling, yet Softmax attention incurs quadratic complexity and unbounded KV cache growth. While linear attention offers a promising alternative, existing approaches lack systematic functional comparison with Softmax attention, rigorous error analysis, and a theoretically grounded improvement roadmap. We address this gap by framing linearization as KV cache approximation and establishing a principled pathway from Softmax attention to linear models. Our analysis identifies five critical components—redundancy elimination, token-level quantization with positional separation, positional compression, inter-layer similarity, and multi-state decomposition—each accompanied by theoretical justification and error bounds, with explicit connections to existing mechanisms. Building upon this framework, we introduce GPA, a linearized attention model that inherits pretrained weights and achieves state-of-the-art results. GPA outperforms strong baselines including MVA and GSA across multiple benchmarks, while requiring less fine-tuning resources. Our work provides both theoretical clarity and practical guidance for advancing linear attention, charting a principled course toward efficient, scalable alternatives to Softmax attention.
GPA: Generative Population Annealing for Test-Time Sequence Design with Pretrained Generative Models
Anirban Sarkar ⋅ Alejandra Duran ⋅ Peter Koo
Oracle-guided biological sequence design must improve predicted function without moving outside the sequence distribution where the oracle is trustworthy. We introduce Generative Population Annealing GPA, a test-time sampler for sequence design with a frozen generator and frozen oracle. GPA instantiates annealed Sequential Monte Carlo at population scale: particles are initialized from a pretrained sequence prior, reweighted by an oracle reward tilt, selectively upsampled as effective sample size falls, mutated with the pretrained generator, and returned as the design pool. The base sampler targets a reward-tilted prior; practical variants shape the proposal or modify the objective to trade activity, specificity, fidelity, and diversity. Across enhancer and promoter benchmarks, GPA scales to thousands of sequences and is competitive with inference-time samplers, gradient-based editors, tree-search diffusion, and RL-fine-tuned generators. The same inference loop is evaluated with masked discrete-diffusion and autoregressive DNA generators. GPA reaches high predicted activity while preserving strong motif fidelity and model-based likelihood, although specialized baselines retain higher diversity or stronger $k$-mer fidelity in some settings. Cross-oracle and sequence-level audits suggest fewer obvious reward-hacking pathologies, including shorter homopolymer runs than CTRL-DNA in the HepG2 audit (5.1 bp mean maximum run versus 20.9 bp), but all validation remains computational.
We consider the well-studied setting of minimizing a convex Lipschitz function using either gradient descent (GD) or its stochastic variant (SGD), and examine the last iterate convergence. By now, it is known that standard stepsize choices lead to a last iterate convergence rate of $\log T/\sqrt{T}$ after $T$ steps. A breakthrough result of Jain et al. [2019] recovered the optimal $1/\sqrt{T}$ rate by constructing a non-standard stepsize sequence. However, this sequence requires choosing $T$ in advance, as opposed to common stepsize schedules which apply for any time horizon. Moreover, Jain et al. conjectured that without prior knowledge of $T$, no stepsize sequence can ensure the optimal error for SGD's last iterate, a claim which so far remained unproven. We prove this conjecture, and in fact show that even in the noiseless case of GD, it is impossible to avoid an excess poly-log factor in $T$ when considering an anytime last iterate guarantee. Our proof further suggests that such (slightly) suboptimal stopping times are unavoidably common.
Gradient-Mine Units: Scorched-Earth Strategy for Model Protection against Unauthorized Fine-Tuning
Jinhyeok Jang ⋅ Jaehong Kim ⋅ ByungOk Han
Pretrained model weights are increasingly released under commercial licenses, usage restrictions, or other conditions that prohibit unauthorized fine-tuning. In practice, however, such misuse is difficult to detect or verify after the fact. This motivates a stronger objective for protected weight release: weights should remain useful for intended inference, yet become practically unattractive to repurpose through unauthorized gradient-based adaptation. We frame this objective as a \emph{scorched-earth} strategy in parameter space: rather than only proving infringement after misuse occurs, the released weights themselves should react destructively when unauthorized fine-tuning begins. To realize this idea, we propose \textbf{Gradient-Mine Units (GMUs)}, a data-free weight-space protection mechanism for pretrained networks. GMUs are planted into selected feedforward layers as hidden units with extreme internal scale, while a locking mechanism keeps them silent at initialization so that the original inference behavior is preserved. During fine-tuning, this locked state is progressively broken, allowing the planted units to emit amplified gradients that disrupt the model's native adaptation dynamics. We formulate GMUs in a general feedforward setting and show that gated instantiations naturally provide an additional hard lock. Empirically, we validate the method on both large language models and Vision Transformers. Across multiple architectures and downstream tasks, GMUs preserve pre-fine-tuning utility while substantially degrading or destabilizing standard fine-tuning. These results suggest that protected weight release can move beyond post hoc attribution toward a practical deterrence mechanism for unauthorized adaptation. Our implementation is available at here.
Grid Games: The Power Of Multiple Grids for Quantizing Large Language Models
Vage Egiazarian ⋅ Erik Schultheis ⋅ Andrei Panferov ⋅ Earl Killian ⋅ Torsten Hoefler ⋅ Dan Alistarh
A major recent advance in quantization is given by microscaled 4-bit formats such as NVFP4 and MXFP4, quantizing values into small groups sharing a scale, assuming a fixed floating-point grid. In this paper, we study the following natural extension: assume that, for each group of values, we are free to select the "better" among two or more 4-bit grids marked by one or more bits in the scale value. We formalize the power-of-two-grids (PO2) problem, and provide theoretical results showing that practical small-group formats such as MXFP or NVFP can benefit significantly from PO2 grids, while the advantage vanishes for very large groups. On the practical side, we instantiate several grid families, including 1) PO2(NF4), which pairs the standard NF4 normal grid with a learned grid, 2) MPO2, a grid pair that is fully learned over real weights and activations, and 3) SFP4, a TensorCore-implementable triple which pairs NVFP4 with two shifted variants. Results for post-training quantization of standard open models and pre-training of Llama-like models show that adaptive grids consistently improve accuracy vs single-grid FP4 under both weight-only and weight+activation quantization.
GridProbe: Posterior-Probing for Adaptive Test-Time Compute in Long-Video VLMs
Mohamed Eltahir ⋅ Ayash ⋅ Ali Habibullah ⋅ Tanveer Hussain ⋅ Naeemullah Khan
Long-video understanding in VLMs is bottlenecked by a single monolithic forward pass over thousands of frames at quadratic attention cost. A common mitigation is to first $select$ a small subset of informative frames before the forward pass; common for training-free selectors via auxiliary encoder-space similarities. Such signals are capped by contrastive pretraining, which usually fails on reasoning-heavy queries (negation, cross-frame counting, holistic summarization). We propose $\textbf{GridProbe}$, an efficient training-free posterior-probing inference paradigm that scores evidence in $\textbf{answer space}$ using a frozen VLM's own reasoning and then selects question-relevant frames $\textbf{adaptively}$, resulting in sub-quadratic attention cost with little to no accuracy loss. We arrange frames on a $K{\times}K$ grid and run lightweight row $R$ and column $C$ probes, where each probe reads its peak posterior as a query-conditioned confidence. The outer product of $R$ and $C$ yields an interpretable $\textbf{importance map}$ whose skewness and kurtosis drive $\textbf{Shape-Adaptive Selection}$, a closed-form rule that reliably replaces the fixed frame budget $M$ with a per-question $M_{\mathrm{eff}}$. We show empirically that $M_{\mathrm{eff}}$, surprisingly, tracks intrinsic question difficulty without ever seeing the answer, a sign of test-time adaptive compute. On Video-MME-v2, $\textbf{GridProbe}$ matches the monolithic baseline within $1.6$ pp Avg Acc at $3.36\times$ TFLOPs reduction, while on LongVideoBench it Pareto-dominates the baseline ($+0.9$ pp at $0.35\times$ compute). Because the selector and QA models can be decoupled, pairing a small 2B selector with a stronger 4B or 8B QA is strictly Pareto-dominant over the 2B monolithic baseline (up to $+4.0$ pp at $0.52\times$ compute, on average), with no retraining. Finally, the interpretability of the importance maps opens future avenues for behavioral diagnostics, grounding, and frame-selection distillation.
Group Perspective Matters: Regulating Debate Relationships Can Mitigate Blind Conformity in Multi-Agent Debate
Hao Wu ⋅ Shoucheng Song ⋅ Chang Yao ⋅ Huaiyu Wan ⋅ Youfang Lin ⋅ Kai Lv
Multi-Agent Debate (MAD) improves the reasoning performance of Large Language Models (LLMs) through multi-round interaction. However, LLMs in MAD are highly susceptible to blind conformity. Existing individual evaluation methods, typically based on confidence or perplexity, fail to reflect the correctness of reasoning and may even exacerbate blind conformity. To address this, we shift the perspective from individual evaluation to group interaction. We define mutual referencing among LLMs as \textbf{Debate Relationships} and recognize that regulating these relationships is the key to mitigating blind conformity. In this paper, we propose a novel framework for \textbf{D}ynamically r\textbf{E}gulating deb\textbf{A}te \textbf{R}elationships (DEAR) from the group perspective. At first, DEAR quantifies consensus and divergence as \textit{group evidence} to capture the debate state. Then, DEAR operates through three stages: 1) What: perceiving group consultation tendency and uncertainty; 2) Who: introducing a Selection RL-Agent to dynamically select reference peers; and 3) How: adopting a Behavior RL-Agent to adaptively adjust generation behaviors.
Growing a Neural Network in Breadth, Depth, and Time
Eivinas Butkus ⋅ Kedar Garzón Gupta ⋅ Nikolaus Kriegeskorte
Spatial and temporal resource constraints are critical for both biological and artificial intelligent systems. Here we define differentiable cost terms for breadth, depth, and time within a recurrent convolutional neural network conceived as a finite subset of an infinite lattice architecture. We optimize these costs jointly with task errors via backpropagation, and efficient computational graphs emerge organically through training. We find that breadth, depth, and time can be traded off against each other to achieve a given level of performance. Networks grow in all three dimensions with task complexity and spontaneously take more recurrent steps when inputs are occluded. Surprisingly, time used by the model correlates with human reaction times in an object recognition task. Our framework provides a normative account of how resource constraints shape neural architectures, connecting to questions about brain design in neuroscience, and may help illuminate the diversity of neural solutions found in nature.
Guided Data Generation for Understanding Model Behavior
Eren Mehmet KIRAL ⋅ Nursen Aydin ⋅ Ilker Birbil
There is a growing need for understanding how trained machine learning models behave beyond standard predictive performance. With this work, we aim to understand trained machine learning models by questioning their data preferences. We propose a mathematical framework for guided data generation that allows us to produce fixed-label, prediction-risky, parameter-sensitive, or model-contrastive samples, among others. To showcase our framework, we pose these queries to a range of models trained on a range of classification and regression tasks, with answers in the form of generated data.
HandEdit: A Unified Benchmark for Egocentric Human-to-Robot Dexterous Hand Image Editing
Zhenjie Yang ⋅ Xingyu Jiao ⋅ Guopeng Zhong ⋅ Shuzhe Yang ⋅ Shi Che ⋅ Chao Wu ⋅ Chenyu Jiang ⋅ Dongjie Zhang ⋅ Yideng Zhang ⋅ Zheng Zhang ⋅ Muyun Jiang ⋅ Haisheng Su ⋅ Shuang Jin ⋅ Donghang Zhang ⋅ Chao Yang ⋅ Li Chen ⋅ Hongyang Li ⋅ Zuxuan Wu ⋅ Junchi Yan ⋅ Xiaosong Jia ⋅ Yu-Gang Jiang
Robotic manipulation with dexterous hands is a cornerstone of Embodied AI, yet its progress is stifled by the high cost of collecting embodiment-aware teleoperation data. While abundant egocentric videos of human hands offer a scalable alternative, the profound discrepancies in appearance, articulation, and camera viewpoints between human and robotic data raise significant challenges for co-training. Though existing general image-editing models demonstrate strong capabilities, they lack necessary embodiment-specific priors to fully bridge this gap. In this work, we present HandEdit, a unified large-scale embodiment-aware image-editing dataset and benchmark specifically designed to transform human hands and arms into various dexterous robotic embodiments within egocentric frames. HandEdit comprises over 200M editing instances derived from five diverse source datasets, covering 26 distinct URDFs, including 13 hand-only and 13 hand-arm configurations. Alongside the dataset, we establish a unified benchmark protocol with two tracks: Hand-only and Hand-Arm, supporting URDF-conditioned evaluation. We conduct extensive evaluations of 11 representative image-editing baselines using a multi-dimensional metric suite, including generic similarity metrics, VLM-based judgment, and embodied-aware metrics. HandEdit serves as a critical resource at the intersection of image editing and robotics: it advances embodiment-aware editing models while enabling scalable dexterous robotic learning from abundant human video data, paving the way for more generalizable Embodied AI. Code, datasets, and benchmarks are available at \url{https://anonymous.4open.science/r/HandEdit}.
HetCCL: Efficient LLM Training on Heterogeneous Vendor GPUs
Heehoon Kim ⋅ Chris Jaehwan Lee ⋅ Taejeoung Kim ⋅ Jongwon Park ⋅ Jinpyo Kim ⋅ Pyongwon Suh ⋅ Ryan H Choi ⋅ Sangwoo Lee ⋅ Jaejin Lee
The rapid growth of large language models is driving organizations to expand their GPU clusters, often with GPUs from multiple vendors. However, current deep learning frameworks lack support for communication across heterogeneous vendor GPUs, leading to inefficiency and higher costs. We present HetCCL, a collective communication library that unifies vendor-specific backends and enables RDMA-based communication across GPUs without requiring driver modifications. HetCCL introduces two novel mechanisms that enable cross-vendor communication while leveraging optimized vendor libraries, NVIDIA NCCL and AMD RCCL. Evaluations on a multi-vendor GPU cluster show that HetCCL matches NCCL and RCCL performance in homogeneous setups while extending support to heterogeneous environments, enabling practical and efficient training across NVIDIA and AMD GPUs without changes to existing deep learning applications.
Heteroscedastic TrueSkill: Modeling Match Noise and Player Consistency
Philipp Kolbe ⋅ Johann Ukrow ⋅ Anna Kazachkova ⋅ Rainer Schlosser ⋅ Ralf Herbrich
TrueSkill is a widely used Bayesian rating system for inferring latent player skill from game outcomes. However, its standard formulation assumes a fixed global performance-noise parameter, making it sensitive to atypical outcomes and unable to distinguish skill from player-specific consistency. We propose two heteroscedastic extensions: a match-specific precision model that yields a robust heavy-tailed comparison likelihood, and a player-specific precision model that captures differences in performance consistency among players. Both extensions use Gamma-distributed precision variables and retain a modular factor-graph representation. Because the introduced Gaussian-Gamma factors do not admit the same closed-form expectation-propagation updates as vanilla TrueSkill, we derive an efficient hybrid message-passing scheme that combines expectation-propagation updates for outcome truncation factors with variational updates for latent precision factors. Experiments on synthetic and real match data show that the proposed models improve robustness to anomalous outcomes and provide interpretable estimates of player consistency.
Hider–Seeker Self-Play: Geometry-Verifiable Process Rewards for Long-Horizon Visual Search
Yu Mao ⋅ Shengchao Chen
Long-horizon visual search requires a vision-language policy to execute multi-step spatial actions over high-resolution images, where performance depends on the entire trajectory rather than any single decision. Existing supervision breaks at this horizon: outcome-only RLVR collapses many geometric decisions into a single delayed signal and discards the per-step geometry the environment already computes, while LLM-judged step rewards recover density at the cost of verifiability, scale linearly with trajectory length, and conflate distinct failure modes. We propose HSSP, a self-play framework in which a Hider adaptively weights challenging skeletons from a fixed trajectory pool and a Seeker learns to solve them through grounded spatial actions. Instead of relying on external judges, HSSP supervises every step with a geometry-verifiable process reward built from target-directed progress and exploration coverage, with trajectory length and spatial redundancy enforced as separate Constrained Markov Decision Process (CMDP) constraints. We pair HSSP with V-Trace, a trajectory-level companion dataset that augments skeleton seeds from established benchmarks with target-region annotations and explicit search-regime labels, ready for any trajectory-level training framework. The reward design admits closed-form guarantees on backtracking, find-versus-explore separation, and policy invariance, providing theoretical backing for the observed gains. Extensive empirical results on eleven challenging multimodal benchmarks, supported by controlled ablations and qualitative case studies, show that HSSP consistently outperforms advanced baselines and transfers zero-shot to OOD tasks.
Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation
Seonghoon Yu ⋅ Dongjun Nam ⋅ Byung-Kwan Lee ⋅ Jeany Son
Recent think-answer approaches in VLMs, such as Qwen3-VL-Thinking, boost reasoning performance by leveraging intermediate thinking steps before the final answer, but their high computational cost limits real-world deployment. To distill such capabilities into compact think-answer VLMs, a primary objective is to improve the student's ability to utilize visual evidence throughout its reasoning trace. To this end, we introduce a novel think-answer distillation framework that encourages the student to anchor its thinking on visual information by masking the student's salient reasoning prefixes. To compensate for such masked textual cues, the student is encouraged to rely more on visual evidence as an alternative source of information during distillation. Our masking strategies include: 1) $\textit{token-wise salient reasoning-prefix masking}$, which masks high-influence reasoning prefixes selectively for each next-token prediction, and 2) $\textit{self-paced masking budget scheduling}$, which gradually increases the masking scale according to distillation difficulty, {measured by discrepancy between teacher--student distributions. In the distillation phase, the student is guided by our salient reasoning-prefix mask, which blocks both future tokens and salient reasoning cues, in place of the standard causal mask used for auto-regressive language modeling. Experimental results show that our approach outperforms recent open-source VLMs, VLM distillation, and self-distillation methods on multimodal reasoning benchmarks, while further analyses confirm enhanced visual utilization along the student thinking process.
High-Dimensional Conditional Independence Testing via Random Projection Aggregation
Dian Jin ⋅ zirui chen ⋅ Ting Li ⋅ Jiaye Teng
Conditional independence (CI) testing is a fundamental problem in statistics. However, classical nonparametric CI tests suffer from the curse of dimensionality, as nonparametric estimation degrades rapidly in high dimensions. This paper proposes RP-CIT, a random-projection-aggregated CI test that remains effective in the high-dimensional regime. The central idea is to project $X$ and $Y$ onto random one-dimensional subspaces before applying a univariate base CI test. This projection (i) preserves conditional independence under the null hypothesis and (ii) reduces the original high-dimensional problem to a collection of univariate CI subproblems, thereby avoiding high-dimensional nonparametric estimation. Since a single projection may miss the dependence signal, we aggregate the $p$-values from $k$ independent projections via a Bonferroni-max rule, which asymptotically controls the Type~I error at the nominal level and yields a limiting Type~II error bound that decreases exponentially in $k$ under a detectable-projection condition. The paper also develops a top-$r$ aggregation extension for alternatives whose signal is spread across multiple projection directions. More broadly, the random-projection aggregation framework can be combined with other scalar base tests; as one example, we use a projected GCM base test to accommodate higher-dimensional conditioning variables. Experiments on synthetic benchmarks, image-valued CI tasks, and Norman Perturb-seq module skeleton learning illustrate the calibration, power, and practical utility of the framework.
High Performance Differentially Private Fine-Tuning using Dataset Distillation
Noel Loo ⋅ Hanshen Xiao ⋅ Alexander Amini ⋅ Mathias Lechner ⋅ Ramin Hasani ⋅ Daniela Rus
Differentially Private Stochastic Gradient Descent (DP-SGD) is the dominant approach to private deep learning, but it pays compute and privacy budget \emph{per trained model}: each new architecture, ensemble member, or downstream fine-tuning round incurs additional privacy cost. We propose \textsc{SPS} (Summarize--Privatize--Synthesize) and its enhanced variant \textsc{SPS+}, dataset-distillation algorithms that release a private synthetic dataset by privatizing intermediate activation statistics from a public pretrained model. Once released, the dataset is post-processed freely: any downstream architecture, ensemble, federated aggregation, or continual update incurs \emph{zero} additional privacy cost. In particular, a single \textsc{SPS+} dataset distilled from a Wide ResNet-22-8 transfers zero-shot to architectures with different inductive biases (Vision Transformers, Swin Transformers, and ConvNeXt), all without re-incurring privacy. Empirically, \textsc{SPS+} is competitive with state-of-the-art DP-SGD across $\epsilon \in \{1,2,4,8\}$ on CIFAR-10 and CIFAR-100, with a notable advantage on CIFAR-100 under strict privacy, achieving a $5.8\%$ gain at $\epsilon{=}1$ over compute-matched DP-SGD. \textsc{SPS+} additionally outperforms prior generation-based DP methods by a large margin and supports private federated and continual learning out of the box.
High-probability Convergence of Gradient Methods under Markovian Stochasticity
Polina Podzorova ⋅ Nikolay Spitsyn ⋅ Savelii Chezhegov ⋅ Aleksandr Beznosikov
Markovian stochasticity naturally arises in many modern learning problems, including optimization with dependent data streams and reinforcement learning. Despite extensive research on convergence of stochastic gradient methods under such stochasticity, high-probability convergence remains largely unexplored. To close this gap, we provide the first high-probability convergence guarantees for gradient-based methods under Markovian stochasticity. We study SGD with a specialized mini-batch-based gradient estimator and establish high-probability convergence results in the non-convex setting under the classical bounded stochasticity assumption, and further extend the analysis to a weaker bounded-variance assumption by incorporating gradient clipping, yielding convergence guarantees for accelerated stochastic gradient descent in the convex setting. For both methods, we characterize the oracle complexity via a concentration inequality that explicitly captures the effect of Markovian dependence on the resulting convergence rates.
Hint Tuning: Less Data Makes Better Reasoners
Siqi Fan ⋅ Minghao Li ⋅ Xiaoqian Ma ⋅ Xiusheng Huang ⋅ Zhuo Chen ⋅ Bowen Qin ⋅ ZhangLiujie ⋅ Shuo Shang ⋅ Weihang Chen
Large reasoning models achieve high accuracy through extended chain-of-thought but generate 5--8 more tokens than necessary, applying verbose reasoning uniformly regardless of problem difficulty. We propose Hint Tuning, a data-efficient approach that teaches models to calibrate reasoning depth. Our key insight: the corresponding instruct model serves as an ideal difficulty probe. By testing what the instruct model can solve with varying guidance, we automatically construct training data across three states: No-Hint (direct answer), Sparse-Hint (minimal prefix), and Full-Hint (complete reasoning). This converts the abstract challenge of difficulty labeling into a measurable consistency check between the instruct and reasoning models. With only 1K self-annotated samples, Hint Tuning achieves 24--66% token reduction (31.5% average) across mainstream reasoning models (Qwen3-Thinking, DeepSeek-R1-Distill) at multiple scales (4B--32B) while maintaining competitive accuracy on five benchmarks. Unlike methods requiring massive distillation datasets or expensive RL, we achieve superior efficiency through simple alignment with the instruct model's capabilities.
Hi-Q: Hierarchical Evidence-guided Query Refinement for Multi-Hop Question Answering
Jueun Kim ⋅ Sungho Park ⋅ WOOK SHIN HAN
A central bottleneck in multi-hop Question Answering (QA) is that the granularity at which a question is expressed often differs from the granularity at which corpus evidence is retrievable. Existing methods address this mismatch either by imposing fixed graph structures over the corpus or by iteratively reformulating the query, but these strategies do not explicitly decide when a query unit is already supported by evidence and when it should be refined. We formulate this bottleneck as retrievable granularity discovery and introduce Hi-Q, an evidence-conditioned framework for hierarchical query refinement. At each query node, a resolution operator tests whether retrieved evidence supports the current query unit; resolved nodes terminate, while unresolved nodes are expanded by a dependency-preserving binary operator and checked by a semantic coverage verifier. Hi-Q therefore grows a query tree whose topology is determined by corpus support signals rather than by a fixed decomposition template or a pre-built graph. Across three multi-hop QA benchmarks, Hi-Q achieves 57.9 EM and 69.3 F1 on average, outperforming PropRAG, a graph-based RAG baseline, by 5.6 EM / 3.9 F1 and IRCoT, an iterative retrieval baseline, by 13.7 EM / 15.8 F1. Under full-corpus retrieval, Hi-Q maintains its gains with 53.4 EM and 65.2 F1 on average, improving over IRCoT by 16.3 EM / 19.4 F1 without corpus-wide graph construction.
Hitting a Moving Target: Test-Time Adaptation for AI Text Detection under Continual Distribution Shift
Kevin Ren ⋅ Manish Raghavan ⋅ Nikhil Garg
Deployed approaches for AI text detection often rely on training-time access to labeled datasets of both human-written and AI-generated text. This approach is vulnerable to three types of distribution shifts that occur continually post-deployment, and for which labeled data is often unavailable: adversarial humanization, new LLMs being released, and temporal drift in human writing. Simultaneously, existing approaches do not leverage a key signal of LLM usage: inference-time homogeneity. We propose a test-time adaptation (TTA) approach, using semi-supervised learning, that adapts to distribution shifts by leveraging homogeneity among unlabeled samples observed at inference time. Empirically, we find that state-of-the-art supervised detectors systematically fail when they encounter distribution shifts in AI-generated and human writing, both adversarial and natural, while test-time adaptation with semi-supervised learning is largely robust; e.g., the commercial model Pangram detects just 24.1\% of our adversarial AI-generated text, compared to 90.5\% for our test-time approach. We establish that test-time adaptation is a promising framework for AI text detection in the wild.
Holistic EvoLution via Intrinsic eXchange for Unified Multimodal Models
Shenghao Dong ⋅ Yuhang Yu ⋅ Bo Li ⋅ Jinwei Chen ⋅ Hao Zhang
Unified Multimodal Models (UMMs) integrate visual generation and understanding within a shared parameter space, yet existing post-training typically relies on external annotations and supervised post-training data, and rarely couples these two capabilities for iterative improvement. We propose Holistic EvoLution via Intrinsic eXchange (HELIX) for Unified Multimodal Models, a two-stage cyclic post-training framework that forms a double-helix coupling between generation and understanding: generation produces controllable visual evidence while understanding provides semantic judgments. A target-driven Data Bridge connects the two stages and supplies reliable training signals for both generation and understanding updates. In Stage 1, the understanding branch provides a self-evaluated likelihood gain reward to optimize the generative policy for stronger semantic faithfulness. In Stage 2, the generation branch supplies reconstruction-based visual evidence, and we refine the understanding branch with Reconstruction and Prompt-Alignment rewards. Extensive experiments on Bagel, trained at 512px, demonstrate consistent gains on compositional and knowledge-intensive benchmarks, outperforming the baseline by +5.6 on T2I-CompBench, +7.1 on GenEval, +1.54 on DPG-Bench and +0.02 on WISE when evaluated at 1024px.
HopChain: Multi-Hop Data Synthesis for Generalizable Vision-Language Reasoning
Shenzhi Wang ⋅ Shixuan Liu ⋅ Jing Zhou ⋅ Chang Gao ⋅ Xiong-Hui Chen ⋅ Binghai Wang ⋅ An Yang ⋅ Shiji Song ⋅ Bowen Yu ⋅ Gao Huang ⋅ Junyang Lin
Vision-language models (VLMs) show strong multimodal capabilities but still struggle with fine-grained vision-language reasoning. We find that long chain-of-thought (CoT) reasoning exposes diverse failure modes, including perception, reasoning, knowledge, and hallucination errors, which can compound across intermediate steps. However, most existing data for reinforcement learning with verifiable rewards (RLVR) does not involve complex reasoning chains grounded in visual evidence throughout, leaving these failure modes understressed during training. We therefore propose HopChain, a scalable framework that constructs an out-of-distribution (OOD) proxy task: OOD only in query format, while strengthening foundational abilities shared with downstream benchmarks to support generalizable gains. Concretely, HopChain synthesizes multi-hop vision-language reasoning chains, each a logically dependent chain of instance-grounded hops where earlier hops establish the instances, sets, or conditions needed for later hops, ending in an unambiguous numerical answer for verifiable rewards. In experiments, we train Qwen3.5-35B-A3B and Qwen3.5-397B-A17B under two RLVR settings (with and without HopChain's multi-hop data) and compare them across 24 benchmarks spanning STEM and Puzzle, General VQA, Text Recognition and Document Understanding, and Video Understanding. Although the multi-hop data is not designed for any specific benchmark, it improves 20 of 24 benchmarks on both models, indicating broad and generalizable gains. Consistently, replacing full chained queries with half-multi-hop or single-hop variants reduces the average score across five representative benchmarks from 70.4 to 66.7 and 64.3, respectively. Improvements are particularly substantial in long-CoT and ultra-long-CoT regimes, peaking at more than 50 accuracy points in the ultra-long-CoT regime. These results show that OOD proxy tasks for long-chain visual reasoning are a scalable RLVR supervision source for generalizable VLM reasoning.
Estimating physical pressure from vision is essential for understanding contact-rich hand-object interaction. However, prior vision-based pressure estimation methods are largely limited to planar surfaces and single image input, making them difficult to apply to dynamic hand-object interaction with diverse objects. We instead formulate pressure estimation as a hand-centric video prediction problem with monocular video as input. This formulation predicts temporally evolving per-vertex normal pressure and contact directly on the hand mesh, yielding a unified output space independent of object shape and sensor layout. Building on this formulation, we propose HOPE, a framework with two key components. First, we lift tactile-glove pressure, planar-sensor pressure, and distance-based hand-object contact annotations into a shared hand vertex space, allowing bare-hand contact data to regularize pressure learning where metric labels are unavailable. Second, we introduce a vertex-anchored video transformer that treats each vertex as a persistent token, aggregates visual features and hand pose over time, and uses a contact-gated pressure head to enforce that pressure vanishes without contact. Experiments on OpenTouch, PressureVisionDB, and hand-object contact benchmarks demonstrate that HOPE supports a common hand-centric representation across object-pressure, surface-pressure, and contact-supervised HOI settings. Despite using metric pressure supervision primarily from gloved-hand videos, HOPE generalizes to bare-hand egocentric and in-the-wild videos, producing joint contact and pressure predictions beyond the scope of contact-only or planar-pressure baselines.
HoTS: Homophily-Aware Temperature Scaling for Graph Neural Network Calibration
Inwoo Tae ⋅ Yoontae Hwang ⋅ Yongjae Lee
For graph node classification, calibrated class probabilities are need when confidence scores, usually the maximum predicted class probability, are used to rank prediction, defer uncertain nodes to human review or control risk. Existing post-hoc calibrators either apply one global temperature or use graph-awared modules without a derived strctural form. We study node-level calibration for graph node-classification problem under structural heterogeneity. Our first results show that a logit-only temperature rule is insufficient when nodes with identical logits but different local homophily require different optimal temperatures. We then analyze a population-concentration contextual stochastic block model with Gaussian features and a one-layer linear GCN. Under equidistant class means, the Bayes posterior over class-template scores is a temperature-scaled softmax whose inverse temperature is governed by homophily-dependent signal strength. In the positive-signal homophilic regime, the resulting temperature decreases approximately with normalized local homophily. This law motivates Homophily-aware Temperature Scaling (HoTS), a simple post-hoc calibrator that assigns each node a positive scalar temperature from entropy-based logit concentration and estimated local homophily. HoTS has three parameters, preserving class, and learn the strength of the structural correction from calibraion data. Across 18 node-classification benchmarks, two GNN backbones, and eight calibration baselines, HoTS achieves the best mean Expected Calibration Error (ECE) of $4.66\%$, the best average rank, and the strongest confidence ranking in selective classification.
How does feature learning change the function space evolution?
João Lobo ⋅ Bruno Loureiro ⋅ Long Tran-Thanh ⋅ Fanghui Liu
Feature learning is widely viewed as the mechanism that distinguishes neural networks from fixed-kernel methods, yet its effect on the underlying function space remains poorly understood. We aim to precisely characterize how the function space (e.g., RKHS) endowed by a two-layer neural network during gradient descent training. We prove that, in a high-dimensional proportional regime, the post-update feature distribution is well approximated by a target-dependent spiked Gaussian covariance, yielding a deformed kernel that can be written as the original isotropic kernel evaluated on geometrically transformed inputs. This characterization enables a spectral analysis of the corresponding integral operator. We prove that the global eigenvalue decay rate is preserved, while the spike selectively alters the leading eigenspaces. For ReLU activations, we derive an explicit perturbative expansion showing that feature learning boosts the eigenvalue of the linear eigenfunction aligned with the target direction and mixes the top radial eigenfunction with a target-aligned quadratic harmonic. These results provide a precise function-space description of early feature learning: gradient descent does not merely rescale a static kernel, but induces a data-adaptive deformation that preferentially enriches directions aligned with the teacher signal.
How Hard Is It for Message-Passing GNNs to Simulate One Weisfeiler-Lehman Color-Refinement Step?
Guanyu Cui ⋅ Yuhe Guo ⋅ Zhewei Wei ⋅ Hsin-Hao Su
Message-passing graph neural networks (MPGNNs) are commonly compared with the Weisfeiler-Lehman (WL) color-refinement procedure, but this comparison does not quantify the resource parameters a network needs to realize color refinement with bounded-size messages and finite numerical precision. We study the cost of simulating a single color-refinement step on unattributed graphs. We distinguish input-independent, or oblivious, simulation from instance-dependent simulation. In the former, the parameters, or their distributions in randomized models, are fixed before the input instance is known. Our results show that the local form of WL color refinement hides a global relabeling problem. In the oblivious setting, deterministic and zero-error randomized MPGNNs cannot solve this problem in the worst case using only shallow networks with small messages. We complement this lower bound with a nearly matching construction in a stronger rooted, port-aware model. By contrast, when the color set is large, bounded-error randomness can greatly reduce the cost, and a one-layer MPGNN with messages of logarithmic size and a logarithmic number of random bits suffices. We show that this logarithmic number of random bits is essentially necessary for shallow, small-message simulations. When the color set is small, we still obtain a rooted, port-aware simulation, but this construction requires more layers or larger messages. We also prove that this extra cost is partly unavoidable, as small color sets force a nontrivial trade-off between the number of layers and the message size. Finally, instance-dependent simulation can be much shallower, but the required instance-specific parameters are not necessarily easy to find. Together, these results reveal quantitative structure hidden behind the statement that MPGNNs match WL color refinement.
How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness
Polina Gordienko ⋅ Georg Schollmeyer ⋅ Frauke Kreuter ⋅ Christoph Jansen
Multi-task benchmarks have become a central pillar of machine learning research, yet their growing influence has incentivised benchmark gaming -- strategic actions taken to improve the leaderboard rank of a specific model. Treating datasets as voters and models as candidates, we consider benchmark-specific training -- the inclusion of benchmark data in training -- as a form of election manipulation. For any ordinal benchmark, the problem of choosing datasets to train on so that a target model becomes top-ranked corresponds to shift bribery, a class of manipulation problems from computational social choice. Leveraging this identification, we show that the benchmark-specific training problem is NP-hard under Borda count and mean win rate. Complementing this worst-case perspective, we introduce the instance-level robustness, the minimum number of datasets a model developer must include in training to top a given leaderboard, and derive expressions for it under arithmetic mean, median, mean win rate and pairwise majority. We evaluate these expressions on MMLU under HELM and on BIG-Bench Hard (BBH) under the Open LLM Leaderboard. Across both suites, mean win rate is hardest to manipulate: this gap is clear on BBH (24 tasks, 4507 models), where its median robustness is 22 tasks (92\%), compared with 13 (54\%) under arithmetic mean and 12 (50\%) under median and pairwise majority.
How Long Does Infinite Width Last? Signal Propagation in Long-Range Linear Recurrences
Mariia Seleznova
We study signal propagation in linear recurrent models at finite width. While existing signal propagation theory relies predominantly on the infinite-width limit, it remains unclear for how long that approximation remains accurate when recurrent depth $t$ grows jointly with width $n$. This question is especially relevant for modern recurrent sequence models, whose natural operating regime involves long input sequences, i.e., large $t$. We derive exact finite-width formulas for the hidden state signal energies in linear recurrences under complex Gaussian initialization. Using these formulas, we identify the joint depth-width scaling regimes that govern signal propagation: (i) a *subcritical regime* $t=o(\sqrt n)$, in which the infinite-width approximation remains valid; (ii) a *critical regime* $t\sim c\sqrt n$, in which non-negligible deviations from infinite-width predictions appear and a nontrivial joint scaling limit emerges; and (iii) a *supercritical regime* $t\gg \sqrt n$, in which finite-width effects dominate. Thus, our results pinpoint the precise recurrent depth scale at which infinite-width theory breaks down in long-range linear recurrences. In turn, this shows when standard initialization schemes, such as Glorot, become unstable. More broadly, our results demonstrate that finite-width effects accumulate more rapidly with depth in recurrent models than in feedforward ones, leading to qualitatively different signal propagation behavior.
How Much is Left? LLMs Linearly Encode Their Remaining Output Length
Mohamed Amine Merzouk ⋅ Dmitri Carpov ⋅ Mirko Bronzi ⋅ Damiano Fornasiere ⋅ Adam Oberman
Large language models generate one token at a time, yet their responses show remarkably consistent length structure: step-by-step solutions converge in predictable token counts, retrievals stop after a few sentences, retractions extend responses by measurable amounts. We ask whether the model carries an internal estimate of how much response remains. Training minimal-capacity linear probes on frozen hidden states of three open-weight 7-8B models across seven completion-style datasets, we find three converging pieces of evidence. First, total response length is linearly decodable from the prompt's last hidden state alone, before any output is emitted. Second, probe directions trained on natural-language datasets transfer broadly, including to controlled synthetic completions never seen in training, outperforming a statistical baseline; the converse direction generally fails, and this asymmetry is itself informative. Third, on curated high-loss completions, the probe's per-position estimate shifts upward at the moment the model retracts and restarts a partial solution, a directional behavior no position-only predictor can reproduce (we note that this is qualitative, not aggregate). We frame this as approximate estimation of remaining generation length, distinct from exact-counting impossibility results for transformers, and interpret it as evidence that LLMs maintain a plan-like internal representation of output length (decodable, not necessarily used causally). Code: https://anonymous.4open.science/r/llm-output-length
How to Interpret Agent Behavior
Sophia Gao ⋅ Kaiser Sun ⋅ Jen-Tse Huang ⋅ Katherine Van Koevering ⋅ Sijie Ji ⋅ Heyuan Huang ⋅ Weiyan Shi ⋅ Zhuoran Lu ⋅ Ziang Xiao ⋅ Daniel Khashabi ⋅ Mark Dredze
Autonomous agents such as Claude Code and Codex now operate for hours or even days. Understanding their runtime behavior has become critical for downstream tasks such as diagnosing inefficiencies, fixing bugs, and ensuring better oversight. A primary way to gain this understanding is analyzing the reasoning trajectories and execution traces these agents generate. Yet such data remains in unstructured natural-language form, making it difficult for humans to interpret at scale. We introduce ACTONOMY (a combination of Action and Taxonomy, pronounced /ækˈtɑːnəmi/), a taxonomy for describing and analyzing agent behavior at runtime. has two components: (1) the taxonomy itself, developed through Grounded Theory and structured as a three-level hierarchy of 10 actions, 46 subactions, and 120 leaf categories; and (2) an open repository that hosts the living taxonomy, provides an automated analysis pipeline that applies it to agent trajectories analysis, and defines an extension protocol for community contributions. Our experiments show that ACTONOMY can compare behavioral profiles across agents and characterize a single agent's behavior across diverse trajectories, surfacing patterns indicative of failure modes. By providing a shared vocabulary, ACT*ONOMY helps researchers, agent designers, and end users interpret agent behavior more consistently, enabling better oversight and control.
Using prompted language models as classifiers enables classification in domains with limited training data, but misses some of the robustness and performance benefits that fine-tuning can bring. We study whether training on multiple classification tasks, each with its own prompt, improves performance on new domains with new classification prompts. We show that such training partially generalizes to adjacent domains, improving classification performance on tasks that are unseen during training. However, we identify specific edge cases where the finetuned models fail to follow prompts, such as when the classification prompt changes completely while the data domain remains the same as during training. We show that classification training can be mixed with general instruction following training, and that (when done well), such training keeps the benefits of classification training and mitigates its generalization failures. Surprisingly, we see that this no-thinking supervised classification training can generalize to with-thinking classification and summarization, suggesting that no-thinking classification training might be instrumentally useful in building other kinds of classifiers and monitoring systems.
Hydra-DP3: Frequency-Aware Right-Sizing of 3D Diffusion Policies for Visuomotor Control
Jinhao Zhang ⋅ Zhexuan Zhou ⋅ Huizhe Li ⋅ Yichen Lai ⋅ Wenlong Xia ⋅ Haoming Song ⋅ Youmin Gong ⋅ Jie Mei
Diffusion-based visuomotor policies perform well in robotic manipulation, yet current methods still inherit image-generation-style decoders and multi-step sampling. We revisit this design from a frequency-domain perspective. Robot action trajectories are highly smooth, with most energy concentrated in a few low-frequency discrete cosine transform modes. Under this structure, we show that the error of the optimal denoiser is bounded by the low-frequency subspace dimension and residual high-frequency energy, implying that denoising error saturates after very few reverse steps. This also suggests that action denoising requires a much simpler denoising model than image generation. Motivated by this insight, we propose Hydra-DP3 (${\bf HDP3}$), a pocket-scale 3D diffusion policy with a lightweight Diffusion Mixer decoder that supports two-step DDIM inference. Our synthetic experiments validate the theory and support the sufficiency of two-step denoising. Futhermore, across RoboTwin2.0, Adroit, MetaWorld, and real-world tasks, HDP3 achieves state-of-the-art performance with fewer than 1\% of the parameters of prior 3D diffusion-based policies and substantially lower inference latency.
HyperVQ: Enabling Hyperprior Entropy Modeling for VQ-Based Generative Image Compression
Yi Niu ⋅ Tianyi Xu ⋅ Mingming Ma ⋅ Xinkun Wang
Vector Quantization (VQ) based generative image compression has achieved remarkable perceptual quality. However, existing VQ codecs suffer from two fundamental limitations. First, they lack efficient content-adaptive entropy modeling and rely on static frequencies, leading to low coding efficiency. Second, the inherent conflict between discrete indices and continuous priors prevents true end-to-end joint Rate-Distortion (RD) optimization. To resolve these issues, we propose HyperVQ, a principled framework that establishes a high-performance hyperprior entropy foundation for VQ-based codecs. The core insight of HyperVQ is to shift probability modeling entirely into the continuous embedding space. Instead of directly predicting probabilities for discrete symbols, HyperVQ predicts a high-dimensional continuous multivariate Gaussian distribution for the continuous latents. By treating the discrete codebook entries as fixed "anchors" in this space, we convert the continuous Gaussian density into categorical index probabilities based on relative distances. This elegant formulation provides a powerful, spatially-adaptive entropy engine and renders the cross-entropy rate objective fully differentiable, empowering the network to actively and dynamically optimize the RD trade-off during training. To ensure practicality, we design the lightweight H Block and the Probability Estimation Engine (PEE) to facilitate highly parallel, millisecond-level inference. Experiments demonstrate that HyperVQ acts as a universal module across diverse VQ architectures (single-scale, large-codebook, RVQ), achieving an average bitrate saving of 18.5%, which is 7.28x the saving achieved by conventional Huffman coding. This establishes a robust, RD-controllable foundation for next-generation generative image compression.
IADR: Interface-Augmented Neural Operator for Phase-Field Mean-Curvature Flow
Qinyi Zhang ⋅ Duanyu Feng ⋅ Yangshuai Wang ⋅ Hao Wang
The Allen--Cahn equation is one of the standard phase-field models for mean-curvature flow (MCF), parameterised by the diffuse-interface width $\varepsilon$. Practical simulations span a working range of $\varepsilon$, so a single model covering this range is needed. PDE-specific neural operators are accurate but must be retrained at every $\varepsilon$; general neural operators cover $\varepsilon$ in one model but lose the phase-field structure of the target. We introduce Interface-Augmented Diffusion-Reaction (IADR), a neural operator that addresses both shortcomings: a target network with a diffusion-and-reaction structure preserves the phase-field structure of the Allen--Cahn equation, while a hypernetwork that maps $\varepsilon$ to the weights of this target network amortises the dependence on $\varepsilon$ across the working range. A single IADR matches the per-$\varepsilon$ specialists on in-distribution data, retains their cross-shape generalisation under both $\varepsilon$- and geometry-out-of-distribution shifts, and runs at the per-$\varepsilon$ specialist's per-step inference cost. We further establish a $K$-step rollout error bound that aligns with the long-horizon behaviour observed in experiments.
IDEA: Unwrapping Visual Black-box Models by Interaction Decomposition
Chaojie Ji ⋅ Jian Xu ⋅ He Zhang ⋅ Hongyan Wu ⋅ Yankai Cao ⋅ Ruxin Wang
Concept-based explanation methods have emerged as a prominent framework for interpreting deep neural networks using human-understandable concepts. However, existing approaches are limited to assigning a scalar importance score to a concept, failing to capture the interaction between the concept and its visual context. We argue that overlooking this interaction leaves a predictor's reasoning only partially understood. To address this gap, we introduce IDEA, a novel global post-hoc explanation framework that explicitly models the interaction between a concept and its visual context. Through this modeling, IDEA decomposes concept-level predictive information into three atomic components: uniqueness (independent concept contribution), redundancy (information shared with the context), and synergy (information arising strictly from concept-context interaction). Consequently, IDEA extends interpretability beyond merely identifying influential concepts to diagnosing how a predictor reasons -- for instance, flagging potential context dependencies that scalar scores cannot reveal. IDEA requires no task-specific concept annotations and no access to classifier internals. Evaluations on controlled and real-world settings demonstrate that IDEA can surface reasoning patterns invisible to traditional scalar concept scores.
IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer
Zhengyu Zou ⋅ Hao Li ⋅ kuixuan jiao ⋅ Liu.Liu ⋅ TingyangXiao ⋅ xiaolin.zhou ⋅ Fangzhou Hong ⋅ Zhizhong Su ⋅ Dingwen Zhang ⋅ Ziwei Liu
Real-world spatial intelligence requires agents to understand scenes from continuous video streams, where objects move, persist, disappear, and reappear over time. While recent spatial foundation models have enabled generalizable feed-forward 3D reconstruction, most streaming methods remain geometry-centric and lack temporally consistent object-level understanding. Meanwhile, existing semantic reconstruction and 3D-aware vision-language methods largely rely on externally extracted 2D semantic cues or loosely coupled geometry inputs, limiting unified geometry-instance learning in long dynamic scenes. In this paper, we propose IGGT4D, a streaming instance-grounded geometry Transformer for online 4D scene understanding. IGGT4D processes video frames sequentially, reuses historical context through causal spatial-temporal modeling, and incrementally updates a unified representation of camera motion, geometry, and object identity. This enables long-sequence feed-forward reconstruction with geometry-instance consistency in dynamic environments. To address the lack of high-quality 4D supervision, we further construct InsScene4D-147K, a large-scale dataset spanning real/synthetic and static/dynamic scenes, with RGB images, depth, poses, and temporally consistent instance masks generated by an automated geometry-guided annotation pipeline. Experiments on 3D reconstruction, pose estimation, instance spatial tracking, and open-vocabulary segmentation demonstrate that IGGT4D outperforms existing streaming baselines while maintaining scalable online inference for long dynamic sequences.
Improved Sample Complexity for Markov Games via Variance-Aware Bandit Learning
Hanbin Zhou ⋅ Canzhe Zhao ⋅ Shuai Li
We study the problem of learning in multi-player general-sum Markov games. In the simulator setting, where the learner can sample the next state conditioned on an arbitrary state action pair, the previous best-known upper bound of the sample complexity for learning an $\varepsilon$-approximate coarse correlated equilibrium (CCE) is $\widetilde{O}(H^4S\sum_{i=1}^m A_i/\varepsilon^2)$ (Li et al., 2022), where $H$ is the horizon, $S$ is the number of states, and $A_i$ denotes the number of actions for the $i$-th player. This work improves the upper bound to $\widetilde{O}(H^4S\max_{i\in[m]} A_i/\varepsilon^2)$, matching the lower bound of $\Omega(H^4S\max_{i\in[m]} A_i/\varepsilon^2)$ (Li et al., 2022). In the online setting, in which the learner can only sample a trajectory, the previous best-known sample complexity upper bound for learning CCE in multi-player general-sum Markov games is $\widetilde{O}(H^6S\max_{i\in[m]} A_i/\varepsilon^2)$ (Song et al., 2021; Jin et al., 2024; Mao et al., 2022). In this work, we improve this to $\widetilde{O}(H^5S\max_{i\in[m]} A_i/\varepsilon^2)$. To our knowledge, this is the tightest upper bound that breaks the curse of multi-agency. The core of our algorithmic design and analysis is the new gap decomposition and, in particular, a bandit algorithm with a high-probability empirical variance regret bound, which might be of independent interest.
Improving Neural Decoding Performance for Language BCIs by Explicitly Modeling Context-Induced Noise
Yao Jia ⋅ Xianhan Tan ⋅ Binli Luo ⋅ Yueming Wang ⋅ Yu Qi
Language brain-computer interfaces (BCIs), including speech and handwriting BCIs, hold significant promise for restoring communication in individuals with paralysis. However, the performance of current language BCIs remains limited due to the non-stationarity of neural activity. Existing neural decoding approaches typically treat neural non-stationarity as stochastic noise. For language BCIs, however, the noise is more complex due to the influence of semantic context in language sequences. For example, the articulation of the same phoneme can vary subtly depending on its surrounding content. Such context-induced variations introduce additional noise into neural representations and can confuse neural decoders, yet this issue has been largely overlooked in previous studies. Here, we propose explicitly modeling these context-induced noises to enhance neural decoding performance in language BCIs. Our key insight is that, unlike stochastic noise, context-induced noise carries useful information that can help the decoder disambiguate linguistically similar units. Therefore, by separating and modeling context-induced noise, we achieve improved decoding performance for language BCIs. Experiments on multiple datasets demonstrate that our approach achieves state-of-the-art performance.
In-Context Learning Can Help Vision Language Models Overcome Training Prior
Kun Wang ⋅ Xindi Wu ⋅ Sanghyuk Chun ⋅ Olga Russakovsky ⋅ Esin Tureci
Vision-language models (VLMs) learn strong statistical regularities during training, which can make them fail to perceive visual evidence when input images violate those regularities. Such failures are often treated as missing visual capability, but they may instead arise because the model has access to the relevant visual evidence and yet selects an answer dominated by learned training priors. In this work, we use controlled visual in-context learning to show that VLMs can overcome learned priors for visual tasks and use relevant visual evidence and capabilities. Specifically, we show that across five vision-centric tasks, demonstrations yield large gains on prior-conflicting counterfactual examples ($+27.2$% for Qwen3-VL 235B and $+18.1$% for Qwen3-VL 32B) while leaving real accuracy nearly unchanged. These findings also generalize across five VLM families with performance increases from $+9.7$% to $+30.3$%. To explain this effect, we analyze visual representations and find that demonstrations redirect attention toward grounding cues while reducing reliance on cues that support the learned prior. Finally, we find that in-context learning is limited when the underlying visual capability is weak or absent. Together, these results show that current VLMs possess visual grounding capabilities that standard evaluations may overlook, underscoring the need for controlled evaluation design in diagnosing and improving model capability.
In-context learning to predict critical transitions in dynamical systems
Yunus Sevinchan ⋅ Juan Nathaniel ⋅ Kai Ueltzhöffer ⋅ Carla Roesch ⋅ Tobias Weber ⋅ Vaios Laschos ⋅ Hang Fan ⋅ Gregor Ramien ⋅ Johannes Haux ⋅ Pierre Gentine ⋅ Benjamin Herdeanu
Critical transitions – abrupt, often irreversible changes in system dynamics – arise across human and natural systems, often with catastrophic consequences. Real-world observations of such shifts remain scarce, preventing the development of reliable early warning systems. Conventional statistical and spectral indicators, such as increasing variance, tend to fail under realistic conditions of limited data and correlated noise, whereas existing deep learning classifiers do not extrapolate beyond their training data distribution. In this work, we introduce TipPFN, an in-context learning (ICL) framework that uses a prior-data fitted network to infer a system's proximity to a critical transition. Trained on our novel synthetic data generator, which is based on canonical bifurcation scenarios coupled to diverse, randomized stochastic dynamics, TipPFN flexibly capitalizes on contexts of various sizes, complexity and dimensionalities. We demonstrate robust, state-of-the-art early detection of critical transitions in previously unseen tipping regimes, sim-to-real examples, and real-world observations in both ICL and zero-shot settings.
InduceKV: Fixed-Footprint Continual Adaptation of Multimodal LLMs via Inducing KV Memories
Qianyu Chen ⋅ Ziteng Feng ⋅ Canran Xiao ⋅ Runxuan Tang
Multimodal large language models must adapt to evolving tasks and domains, yet continual improvement under bounded deployment footprint remains difficult because repeated parameter updates or growing replay stores can accumulate adaptation state over time. We study fixed-footprint continual adaptation: the deployed adaptation state is kept under a fixed memory budget, while the backbone model is left unchanged and task-specific updates are externalized. We propose InduceKV, a retrieval-based method that stores each selected training prefix as an attention-ready memory entry, consisting of a frozen retrieval key and compact layerwise key--value (KV) payloads that can be appended to the model's self-attention cache. Under a strict memory budget, InduceKV constructs a compact inducing set through bilevel selection: a lightweight calibration is fit for retrieval, while the selected memory balances current-task likelihood, anchor-based retention, and coverage in the frozen retrieval space. Across task-incremental instruction tuning, continual VQA, domain-incremental adaptation, and lifelong multimodal instruction tuning, InduceKV consistently improves over PEFT, MoE, replay, and prompt-retrieval baselines under matched memory budgets. We further report backbone-matched, stage-1 CoIN, compute-matched, and scalability diagnostics, showing that the gains are not due to a stronger backbone, replay alone, or an unbounded candidate pool.
Inductive Deductive Synthesis: Enabling AI to Generate Formally Verified Systems
Shubham Agarwal ⋅ Alexander Krentsel ⋅ Shu Liu ⋅ Mert Cemri ⋅ Audrey Cheng ⋅ Rui Meng ⋅ Tomas Pfister ⋅ Chun-Liang Li ⋅ Sylvia Ratnasamy ⋅ Aditya Parameswaran ⋅ Matei A Zaharia ⋅ Ion Stoica ⋅ Mohsen Lesani
AI agents increasingly excel at generating, testing, and refining code. However, they fall short on tasks requiring formal guarantees of full coverage that testing alone cannot provide. Distributed systems are a prime example: properties such as consistency between reads and writes must hold under every possible interleaving of events. Mechanized formal verification can guarantee such correctness, but typically demands months to years of expert effort. As evidence, even SOTA coding agents (Claude Code with Opus 4.6, Codex with GPT-5.4) succeed on only 2/7 distributed key-value-store specifications. In this paper, we present the first effective approach to addressing this gap, Inductive Deductive Synthesis (IDS), which jointly and incrementally synthesizes implementation and proof, and learns from failed attempts to systematically try promising strategies. Built as an agentic LLM system, IDS achieves 7/7 in about 6.8 hours and $106 per spec on average, roughly 200× faster than expert effort and 17% cheaper than SOTA agents. IDS further incorporates performance feedback into the same loop, yielding implementations up to 3× faster than published verified systems.
IndustryCode: A Benchmark for Industry Code Generation
Puyu Zeng ⋅ Zhaoxi Wang ⋅ Zhixu Duan ⋅ Liang Feng ⋅ Shaobo Wang ⋅ Cunxiang Wang ⋅ jinghang Wang ⋅ Bing Zhao ⋅ HU WEI ⋅ Qibing Ren ⋅ Linfeng Zhang
Code generation and comprehension by Large Language Models (LLMs) have emerged as core drivers of industrial intelligence and decision optimization, finding widespread application in fields such as finance, automation, and aerospace. Although recent advancements have demonstrated the remarkable potential of LLMs in general code generation, existing benchmarks are mainly confined to single domains and languages. Consequently, they fail to effectively evaluate the generalization capabilities required for real-world industrial applications or to reflect the coding proficiency demanded by complex industrial scenarios. To bridge this gap, we introduce IndustryCode, the first comprehensive benchmark designed to span multiple industrial domains and programming languages. IndustryCode comprises 579 sub-problems derived from 125 primary industrial challenges, accompanied by rigorous problem descriptions and test cases. It covers a wide range of NAICS sectors, including Aerospace Product Mfg, Chemical Mfg, and Information, and incorporates diverse programming languages such as MATLAB, Python, C++, and Stata. In our evaluation, the top-performing model, Claude 4.7 Opus, achieved an overall accuracy of 71.9\% on sub-problems and 48.9\% on main problems. The benchmark dataset and automated evaluation code are now publicly available.
Many stochastic gradient methods are believed not to converge when the noise in stochastic gradients has only a finite $p$-th moment for $p\in\left(1,2\right)$, a setting known as the heavy-tailed noise assumption. However, some recent studies have found that Stochastic Gradient Descent ($\textsf{SGD}$), without any modifications to its update rule, can surprisingly converge in expectation for convex problems with bounded domains, highlighting the potential of classical stochastic gradient methods. Inspired by this recent progress, we provide a comprehensive study of stochastic optimization under heavy-tailed noise and establish new in-expectation convergence results for Stochastic Mirror Descent ($\textsf{SMD}$) in convex optimization and $\textsf{SGD}$ in nonconvex optimization. Notably, our results not only hold without algorithmic changes but also avoid restrictive assumptions such as bounded domains that were imposed in prior work. More importantly, our analysis provides a new, elegant, yet powerful framework for studying heavy-tailed stochastic optimization, opening a new route to understanding first-order stochastic gradient methods.
Inferring how internal brain state shapes neural responses with state-dependent diffusion models
Kaiwen Sheng ⋅ Yiqi Jiang ⋅ Yuyang Song ⋅ Gaurav Tyagi ⋅ Haotian Ye ⋅ Fangzhao Zhang ⋅ Ruiming Du ⋅ Pengxiao Wang ⋅ Stefano Ermon ⋅ Karl Deisseroth ⋅ Anish Mitra
Neural responses to the same stimulus can vary dramatically with fluctuations in the brain's internal state and the organism's behavioral state. Yet studying how brain state shapes neural responses is difficult: limited experimental recordings sparsely sample this joint space, leaving many biologically valid state-stimulus pairings unobserved. Existing methods either capture state-dependent modulation with restricted model assumptions or synthesize neural activity without explicitly controlling internal brain state. We introduce BrainStateDiff, a conditional diffusion framework that models the spatiotemporal distribution of future neural activity and behavior given external stimuli and recent brain states. Motivated by two gaps in finite recordings, missing state-stimulus combinations and limited repeats of observed conditions, BrainStateDiff introduces two state-aware sampling strategies: cross-state sampling, which recombines observed states and stimuli to synthesize plausible neural responses from unobserved pairings, and matched-state sampling, which resamples from observed pairings to capture local trial-to-trial variability. We evaluate BrainStateDiff on our newly collected wide-field calcium imaging data in mouse V1, electrophysiology data in monkey V4, and two-photon calcium imaging data in mouse striatum. Our results show that BrainStateDiff generates biologically meaningful samples that preserve the state-dependent structure of real recordings and expand neural response space coverage beyond the original recordings. BrainStateDiff establishes state-conditioned generative modeling as a practical framework for studying how internal brain state shapes neural responses beyond the state-stimulus combinations directly observed in the experiments.
An extensive line of work studies fairness interventions for network embeddings, but less is known about their baseline behavior. In this work, we ask: how do baseline embeddings (without fairness interventions) produce disparate effects at the representation level? We analyze the asymptotic behavior of low-dimensional embeddings on stochastic block model (SBM) graphs, which encode both homophily and group structure. We characterize exact conditions under which embeddings cause information loss, showing that the amount of information loss depends directly on the graph’s density and assortativity. Notably, very different graphs can produce identical embeddings in the limit, and this non-invertibility disproportionately affects smaller and sparser communities. As a result, simple downstream tasks, such as link prediction, introduce higher error rates for these communities, helping explain disparities widely observed in practice.
InformedXRD: Reproducible Benchmarks and Physics-Informed Evaluation for Powder Diffraction Symmetry Classification
Edward G Friedman ⋅ Elizabeth J Baggett ⋅ Abhishek Shetty ⋅ Derrick Chan-Sew ⋅ Vanellsa Acha ⋅ Harshitha Dwaracherla ⋅ P. A Kienzle ⋅ William Ratcliff
Machine learning for powder X-ray diffraction (PXRD) symmetry classification is being trained at scale, yet the field lacks a common evaluation standard: papers report results on different subsets of the RRUFF mineral diffraction database, under different target taxonomies, with different preprocessing and different summary metrics, and published models are rarely released in a form that permits rescoring. We present InformedXRD, an evaluation framework with four parts: (i) a machine-readable mapping from 230 space groups to 99 extinction groups, the finest classes that powder diffraction can identify from systematic absences alone; (ii) two algorithmically curated real-data benchmarks, RRUFF-473 and RRUFF-325, whose inclusion rules are specified as code and reproduce from a frozen upstream snapshot; (iii) a prior-only frequency baseline and stratified reporting protocol that expose when apparent gains are driven by label imbalance or nuisance-fit severity; and (iv) a hierarchy-aware error metric on the condensed translationengleiche (same-lattice) subgroup graph that measures the crystallographic severity of misclassifications. Using four publicly released ViT checkpoints from Baggett et al. (2026) and an in-house 1D residual CNN as case studies, we show that evaluation design changes model ranking: standard paired tests point in different directions depending on whether one evaluates Top-1 or Top-5, and stratified analysis exposes a label-entropy confound that aggregate accuracy hides. All mapping tables, benchmark code, evaluation scripts, and case-study checkpoints are released as shared infrastructure for future PXRD-ML evaluation.
Instance-Dependent Bandit Convex Optimization in One Dimension
Felix Breuer ⋅ Alireza Bakhtiari ⋅ Kevin Jamieson
We study the instance-dependent sample complexity of one-dimensional stochastic bandit convex optimization in the fixed-confidence setting. First, we consider minimizer localization, where the learner seeks to identify a point within distance $\varepsilon$ of the unique minimizer. We introduce a new instance-dependent quantity $\Delta_\varepsilon(f)$, corresponding to the threshold for which the sublevel set of $f-\min f$ has length $2\varepsilon$. We prove a lower bound of order $\Delta_\varepsilon(f)^{-2}$ and give a bisection-style algorithm that matches this rate up to logarithmic factors. We then revisit the simple-regret setting, where the goal is to identify an $\varepsilon$-optimal point. Surprisingly, no comparable instance-dependent improvement is possible: every correct algorithm has sample complexity of order $\varepsilon^{-2}$ on every instance, matching the minimax rate up to constants. Together, these results separate the statistical complexity of minimizer localization from that of simple-regret optimization in one-dimensional convex bandits.
InTAct: Interval-based Task Activation Consolidation for Continual Learning
Patryk Krukowski ⋅ Jan Miksa ⋅ Piotr Helm ⋅ Jacek Tabor ⋅ Paweł Wawrzyński ⋅ Przemysław Spurek
Continual learning seeks to acquire new knowledge while preserving previously learned representations. Despite the success of prompt-based methods, they remain fragile in domain-incremental learning scenarios, where shifts in the input distribution cause representation drift in shared layers and lead to forgetting. We introduce InTAct, a method that preserves functional behavior in shared representations without freezing parameters or retaining past data. InTAct constrains updates within activation regions associated with prior domains, while allowing flexible adaptation elsewhere, thereby stabilizing neuron functionality rather than directly restricting parameter values. The approach is architecture-agnostic, integrates seamlessly with existing prompt-based frameworks, and matches or outperforms state-of-the-art methods on domain-incremental benchmarks.
IntegrityBench: Can LLMs Be Trusted as Co-Scientists? A Research Integrity Benchmark
Sai Sidhanth Manoharan Jayanthi ⋅ Yash Tripathi ⋅ Silu Sharma ⋅ Shivank Garg ⋅ Lin Li
Language models are increasingly deployed as co-scientists in the real world, yet their ability to uphold research integrity, particularly under institutional pressures, remains unmeasured. We introduce IntegrityBench, a comprehensive benchmark evaluating three facets (misconduct classification, ethical action reasoning and artifact-grounded decision making), using 36 paired misconduct and ethical control tasks under a 5-level implicit-explicit pressure protocol across 3 domains and 4 research stages. Through a large-scale evaluation across 18 frontier model variants, we find that, under the strongest pressures, frontier models fail roughly one in three integrity-critical decisions. Surprisingly, neither scale nor reasoning can reliably enhance research integrity. Explicit pressures reliably induce compliance with misconduct while implicit contextual reframing more often causes over refusal of legitimate research tasks. Across tasks, models treat surface-level cues as evidence of misconduct without recognizing procedural justifications. Further, models failing to classify research requests accurately perform equally or better on artifact-grounded decision making (mean accuracy of 85.7 when classification fails versus 79.4 when it succeeds), demonstrating that the three facets might be structurally dissociated because accurate ethical actions don't require correct classification to precede them. Frontier models can appear helpful while displaying research integrity failures that create two distinct deployment risks: facilitating research misconduct and diminishing trust in AI-assisted research outputs.
Intend, Reflect, Refine: An Adaptive Multimodal Reflection Framework for Autonomous Driving
Zisheng Chen ⋅ Yuping Qiu ⋅ Jianhua Han ⋅ Tao Tang ⋅ Xiuwei Chen ⋅ Likui Zhang ⋅ Yingcong Chen ⋅ Hang Xu ⋅ Xiaodan Liang
Recent Vision-Language-Action (VLA) models have advanced end-to-end autonomous driving by incorporating reasoning for better interpretability and planning quality. However, most existing approaches directly generate the final trajectory without explicitly examining its future consequences, which limits their reliability in complex and dynamic environments. To address this limitation, we propose IRR-Drive (Intend, Reflect, Refine), an adaptive multimodal reflection framework for autonomous driving. Specifically, to tightly couple high-level reasoning with physical constraints, IRR-Drive first generates a preliminary textual intention and anticipates potential interactions by predicting future semantic bird's-eye view (BEV) representations. This dual-modality (Text + BEV) reflection space explicitly models anticipated scene evolution, enabling the model to rigorously self-correct and refine its initial intent before generating the final trajectory. Furthermore, to balance planning performance and computational efficiency, we construct reflection-oriented training data and design an adaptive reflection reward, enabling the model to adaptively select its reasoning mode according to scene complexity. Instead of using reasoning primarily as an auxiliary interpretation, IRR-Drive directly integrates an adaptive reflection mechanism into the planning framework, enabling grounded, decision-aware trajectory correction that is driven by scene complexity. Our method achieves state-of-the-art performance on the NAVSIM benchmark in both PDMS and EPDMS. Extensive experiments demonstrate the effectiveness of our multimodal reflection framework and validate the efficacy of the proposed adaptive reflection strategy.
Intent-Aware Caching for Efficient LLM Serving
Kazi Hasan Ibn Arif ⋅ JinYi Yoon ⋅ Dimitrios S Nikolopoulos ⋅ Hans Vandierendonck ⋅ Deepu John ⋅ Bo Ji
Modern LLM serving systems use prefix caching to accelerate multi-turn conversational workloads. They cache the key-value states of previous conversation prefixes and reuse them when subsequent requests share the same prefix, to avoid redundant prefill computation and reduce serving latency. However, achieving a high cache hit rate remains challenging because cache reuse depends on whether a conversation continues and how soon the next turn arrives. Our analysis of real workloads shows that user intent is a strong signal for this behavior, as different intents exhibit different continuation probabilities and inter-turn delays. Based on this insight, we propose Intent-Aware Caching (IAC), an online eviction policy that classifies requests by intent, learns lightweight reuse statistics from recent serving logs, and uses them to guide an oracle-inspired eviction decision. Extensive evaluation on real workloads (e.g., ShareChat) shows that IAC reaches a 34.6\% cache hit rate, substantially outperforming the ML-based LPC (18.8\%) and LRU (12.2\%), while reducing average serving latency by up to 54.6\%.
Inter-domain Inference for Gaussian Process Variational Autoencoders
Xinxing Shi ⋅ Xiaoyu Jiang ⋅ Mauricio A Álvarez
Gaussian Process Variational Autoencoders (GPVAEs) effectively model sample dependencies via latent GP priors, but their inference remains computationally prohibitive at scale. Existing methods typically rely on inducing points, which are restricted to evaluations in the input domain and provide limited control over inductive bias. We introduce a scalable inference framework for GPVAEs that uses inducing variables defined as general linear functionals of the latent GPs, rather than point evaluations. Our formulation recovers the standard inducing-point GPVAE inference as a special case and yields a simplified reconstruction-minus-KL training objective that makes this connection explicit. More broadly, it provides a principled interface for incorporating operator-dependent representations and computational structures via the choice of inducing features. In particular, we instantiate this framework using Fourier features to capture global spectral correlations and B-spline features to induce efficient banded covariances. Across multiple tasks in representation learning, imputation, and conditional generation, our method offers a competitive accuracy-efficiency trade-off to existing approaches.
Internalize External Competence for Visual Instruction Editing
Wenjun Huang ⋅ Rui Zhao ⋅ Suyang Hou ⋅ Jiahao Tang ⋅ Wenjia Wang ⋅ Zhuobai Dong ⋅ Alex Jinpeng Wang ⋅ Hu J Guo
What if you could edit an image simply by drawing on it---circling an object, writing a short label, or sketching an arrow---with minimal or even no typed prompt? We first reveal a surprising capability: when a text-driven image editor is paired with a strong Vision-Language Model (VLM) at inference time, the combined system can directly interpret visual instructions embedded in the image itself, such as on-image text, bounding boxes, and directional arrows. However, this pipeline incurs additional VLM-planning latency, introduces brittle cross-model error propagation, and depends on external APIs or auxiliary model weights. To address these limitations, we introduce Siphon, a framework that internalizes this externally elicited visual-instruction-following competence into a single diffusion-based editor. Rather than relying on an external VLM at inference time, Siphon first uses a VLM planner to synthesize paired visual-instruction supervision, and then transfers this competence into the editor through lightweight LoRA fine-tuning. The resulting model reads, grounds, and executes on-image annotations directly from pixels, while keeping the diffusion sampling cost close to that of the base editor. Across multiple diffusion architectures, Siphon substantially improves spatial controllability, instruction adherence, and marker removal over text-driven baselines, while matching VLM-assisted pipelines without their additional planning latency, API dependence, and cross-model fragility. We do not position rendered long text as a replacement for conventional text prompts; rather, on-image text is mainly intended for short labels and spatially grounded edit intents, while long or complex instructions can still be provided through the standard text channel.
Intrinsic Information Theoretic Analysis of ReLU Nets
Johan Mylius-Kroken ⋅ Elisabeth Wetzer ⋅ Ali Ramezani-Kebrya ⋅ Robert Jenssen ⋅ Kristoffer Wickstrøm
Mutual information (MI) is fundamental to representation-learning analysis in deep networks, but in high dimensions MI estimators depend on hyperparameters --- bin widths, kernel bandwidths, neighbourhood radii, auxiliary networks --- whose bias often dominates the signal. We show that continuous piecewise-linear (CPWL) networks, including ReLU networks, admit an intrinsic information-theoretic analysis that avoids this pitfall: their inherent partition of the input space into convex linear regions assigns to every input a discrete region label $\Pi$ alongside the continuous representation $T$. We decompose $I(Y; T)$ exactly into geometric terms, where the \emph{routing information} $I(Y; \Pi)$ admits a hyperparameter-free plug-in estimator that approximates a population lower bound on every other MI term in the decomposition. When the partition saturates, a functional-equivalence quotient at relative Frobenius tolerance $\varepsilon \in [0, 2]$ collapses regions implementing the same linear operator and restores informativeness. Empirically, the routing estimator correlates strongly with three established MI baselines {\it without any hyperparameter}, and the functional quotient at moderate $\varepsilon$ recovers informative estimates in the regime where the raw partition is bounded by $\log_2 N$ rather than $H(Y)$. We provide a parameter-free framework for analysing the internal information geometry of CPWL networks, grounded in their inherent piecewise-linear structure rather than in external discretisation.
Intrinsic-Preserving Schrödinger Bridge for Direct Part-Aware 3D Generation
Qitong Yang ⋅ Mingtao Feng ⋅ Zijie Wu ⋅ Jie Feng ⋅ Weisheng Dong ⋅ Ajmal Mian
Controllable 3D generation is a fundamental challenge in computer vision. Existing classifier-free guidance methods often yield suboptimal alignment with the given conditions, while Schr"{o}dinger Bridge approaches overlook intrinsic geometric discrepancies across multimodal manifolds, leading off-manifold trajectories and semantic misalignment. To address these issues, we propose Intrinsic-Preserving Schr\"{o}dinger Bridge (IPSB), a novel framework for direct part-aware 3D generation that leverages cross-modal intrinsic consistency. Specifically, we introduce Hierarchical Part-level Cross-modal Encoder that constructs hierarchical part-level representations through probability-based soft coverage, explicitly constraining graph topology to establish semantically consistent part graphs across modalities. Additionally, our IPSB incorporates asymmetric GW-inspired relational regularization, which strictly preserves the intrinsic semantic structure in the conditional manifold while allowing sufficient geometric flexibility in the 3D target manifold. Based on our IPSB framework, we devise Compositionally Controllable 3D Generation based on part-wise intervention inference, enabling independent replacement or resampling of target part nodes while fixing the trajectories of remaining components, thus achieving part-level controllable editing without full regeneration. Extensive experiments demonstrate that our superiority, efficiency and generality.
Inverse Modeling for Laser Pulse Shape Design in Inertial Confinement Fusion
Ricardo Luna Gutierrez ⋅ Vineet Gundecha ⋅ Rahman Ejaz ⋅ Varchas Gopalaswamy ⋅ Riccardo Betti ⋅ Sahand Ghorbanpour ⋅ Aarne Lees ⋅ Soumyendu Sarkar
The achievement of practical fusion energy remains one of the most pressing unsolved scientific challenges, with immense implications for carbon-free power. A critical factor in the success of Inertial Confinement Fusion (ICF) is the design of a Laser Pulse Shape (LP) that can optimally drive implosions under stringent physical constraints. Traditional LP design relies on computationally expensive simulations and manual iterative refinement. We introduce the ICF Laser Pulse Shape Design System (LPDS), a generative inverse modeling framework that maps desired outcomes and target pellet configurations directly to optimized LPs. Crucially, we design a multi-objective loss function to ensure the generated LPs adhere to fundamental physical constraints and experimental feasibility. Furthermore, we present constraint-conditioning, inpainting, and gradient-based LP editing mechanisms for maximum fine-grained control over specific pulse characteristics during generation. Moreover, we validate our framework for LP design on a real-world experimental data. Our method establishes a data-driven inverse design framework for LP in ICF, contributing to the advancement of practical and sustainable fusion energy.
IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams
Jinzhao Li ⋅ Yinuo Chen ⋅ Wenxuan Song ⋅ Yijia Lei ⋅ Yichi Zhang ⋅ Honglei Yan ⋅ Panwang Pan ⋅ Miao Liu
Recent multimodal large language models (MLLMs) achieve strong performance on reactive question answering, but real-world streaming assistants require proactive reasoning over continuous visual inputs. Existing benchmarks mainly study reactive or proactive interactions in isolated single-turn settings, overlooking dynamic multi-turn scenarios where users may add, modify, or cancel proactive requests alongside interleaved reactive queries. To address this gap, we introduce IPIBench, the first benchmark for evaluating Interactive Proactive Intelligence of MLLMs under streaming video settings. IPIBench covers proactive monitoring, proactive task management, and interleaved reactive–proactive requests. Evaluations on representative MLLMs reveal two major limitations: unstable proactive triggering and weak coordination between reactive and proactive behaviors. We further propose IPI-Agent, a training-free agentic framework with an interaction-control policy and a temporal-gating mechanism for stabilizing proactive triggering and coordinating multi-turn interactions. Experiments show that IPI-Agent consistently improves existing MLLMs across all benchmark settings.
Iterative Latent Refinement for Value Learning in Offline Goal-Conditioned RL
Daoxin Li ⋅ Songcheng Xu ⋅ Guozhang Chen
Offline goal-conditioned reinforcement learning (GCRL) enables goal-reaching from static datasets, but the learned value, critic, or compatibility score must support both long-horizon reachability inference and offline policy extraction. We study iterative latent refinement as an architecture-level replacement for the feedforward value or critic backbone used by existing offline GCRL algorithms. Instead of emitting a score after a single feedforward computation, the network repeatedly updates a latent representation using shared recurrent update weights and lightweight step-specific parameters before exposing the final actor-facing signal. On the recent benchmark datasets reported in this study, recurrent variants improve performance across several algorithms without changing their losses, data, actors, or training protocol; across the 21 reported algorithm-dataset rows, the mean improvement is $10.7$ percentage points, with the largest gains on stitching and bottleneck maze tasks. CRL diagnostics are consistent with an improved policy-extraction interface: recurrent critics increase normalized separation between matched and mismatched goals and produce actors that remain closer to dataset actions under several behavior-support proxies. Depth and update-capacity studies further suggest that the effect depends on how computation is allocated between recurrent depth and per-step expressivity.
Iterative Nonlinear Computation Underlying Abstract Reasoning
Zitian Gao ⋅ Yilong Chen ⋅ Yihao Xiao ⋅ Xinyu Yang ⋅ Ran Tao ⋅ Haoming Luo ⋅ JY Zhou ⋅ Bryan Dai
Universal Transformers (UTs) show strong performance on abstract reasoning tasks such as ARC-AGI, yet the specific sources of their performance gains remain underexplored. We conduct systematic ablations over UT variants and find that performance improvements on abstract reasoning are driven primarily by (i) the recurrent inductive bias induced by parameter sharing across depth and (ii) strong nonlinearity inside the Transformer block, rather than by elaborate architectural heuristics. Motivated by this finding, we propose the Universal Reasoning Model (URM). First, we propose ConvSwiGLU, which augments the feed-forward block with a channel-wise short convolution to effectively enhance nonlinearity of UT. Second, we adopt Truncated Backpropagation Through Loops (TBPTL) to eliminate noise from early loop when training with many loops, by freezing the gradients of early loops. Our approach substantially improves abstract reasoning performance, achieving 57.5% pass@1 on ARC-AGI 1 and 16.0% pass@1 on ARC-AGI 2. More importantly, we pretrain a 8-billion-parameter large language model based on the URM architecture from scratch, and the results show that URM still substantially outperforms the baseline across all reasoning tasks.
It Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMs
Sangwoo Park ⋅ Woongyeong Yeo ⋅ Yumin Choi ⋅ Hyomin Lee ⋅ Kangsan Kim ⋅ Seanie Lee ⋅ Jinheon Baek ⋅ Seong Joon Oh ⋅ Sung Ju Hwang
Contextual Integrity (CI) defines privacy not merely as keeping information hidden, but as governing information flows according to the norms of a given context. As large language models are increasingly deployed as personal agents handling sensitive workflows, adhering to CI becomes critical. However, even frontier models remain unreliable in making disclosure decisions, and existing mitigation strategies often degrade underlying task performance. To overcome this privacy-utility trade-off, we propose SelfCI, a complementary self-distillation framework that decouples information suppression from task resolution. SelfCI jointly optimizes two independent reverse KL divergences over distinct teacher distributions derived from feedback: one encourages preserving task-relevant information for utility, while the other enforces minimal and appropriate disclosure. This complementary formulation induces a Product-of-Experts (PoE) target, aligning the policy with the intersection of capability and privacy requirements. Without relying on costly external supervision, empirical evaluations demonstrate that SelfCI consistently outperforms competitive baselines such as online reinforcement learning algorithms (e.g., GRPO). These trends further extend to out-of-domain settings involving agentic workflows and accumulated private context, suggesting that SelfCI provides a practical path toward CI alignment.
JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation
Yue Xun ⋅ Junyu Liu ⋅ Qian Niu ⋅ Xinyi Wang ⋅ Zheng Yuan ⋅ Zirui Li ⋅ Zequn Zhang ⋅ Bowen Zhao ⋅ Emma, Shujun Wang ⋅ Zihui Li ⋅ Kan Hatakeyama-Sato ⋅ Yusuke Iwasawa ⋅ Yutaka Matsuo
We introduce JMed48k, a multi-profession Japanese healthcare licensing benchmark for evaluating vision-language models. Built from official PDF materials released by the Japanese Ministry of Health, Labour and Welfare, JMed48k contains 48,862 exam questions and 20,142 images from 11 national licensing examinations between 2005 and 2025, with visual content annotated under an 8-type taxonomy. From this corpus, we derive JMed48k-Eval, a recent five-year evaluation subset with 12,484 scored questions, including 9,905 text-only questions and 2,579 questions with images. We evaluate 21 proprietary, open-source, and medical-specific models, reporting text-only and with-image performance separately. Because these subsets contain different questions, we further introduce a paired image-removal audit that evaluates questions with images before and after removing visual content to explore four answer-transition states. The audit shows that proprietary and open-source models gain substantially from images, whereas medical-specific systems show limited observable use of visual evidence, with many correct answers persisting after image removal. Even among proprietary models, the net image-removal effect varies sevenfold across professions, from +5.7 points on Physician questions to +39.8 points on Public Health Nurse questions. We release JMed48k to support reproducible, profession-stratified evaluation of vision-language models in medical licensing settings.
Joint protein, mRNA, DNA sequence design and optimization with nucleotide-level Potts models
Blazej Banaszewski ⋅ Lars J Dornfeld ⋅ Dexiong Chen ⋅ Karsten Borgwardt ⋅ Lukas F Milles
Protein design generates amino acid sequences first and then translates them into DNA post-hoc. Control over the nucleotide sequence via codon choice at design time is thus limited. We introduce NuCaliby, a structure-conditioned model that jointly designs a protein's amino acid, mRNA, and DNA sequences at nucleotide resolution. It predicts a Potts model over a nucleotide graph, trained using amino acid supervision by marginalizing over synonymous codons. Its energy-based formulation supports composable inference-time guidance from both differentiable and non-differentiable objectives. NuCaliby preserves residue-level designability on par with Caliby while learning codon-aware nucleotide couplings directly from structure. Under organism-specific tRNA adaptation guidance, NuCaliby improves adaptation to host tRNA pools beyond synonymous sequence space while preserving protein designability. NuCaliby also embeds functional RNA motifs, directly into coding sequences, a problem that is combinatorially infeasible with extensive sampling alone. In wet lab experiments NuCaliby produces expressible and soluble proteins on par with state of the art methods, although it enables more complex nucleotide-level control and optimization. Together, these results extend protein design beyond the residue level, enabling joint optimization across the central dogma of molecular biology.
Keep It CALM: Analyzing the Limits of Global Unsafety in Text-to-Image Generation
NaHyeon Park ⋅ Minhyun Lee ⋅ Hyunjung Shim
Training-free safeguards for text-to-image diffusion models often rely on a reusable safety signal, such as an unsafe direction or global toxic subspace, applied broadly across prompts. We provide a controlled geometric analysis of this global-unsafety assumption and reveal a consistent coverage-selectivity trade-off: compact unsafe subspaces fail to cover heterogeneous unsafe semantics, whereas broader aggregation increasingly distorts safety-adjacent benign prompts. Motivated by this finding, we propose CALM (Counterfactual Adaptive Local Modulation), a training-free safeguard that replaces uniform global removal with prompt-local counterfactual correction. Using matched unsafe-benign anchors, CALM routes each prompt to active unsafe categories, minimally edits only violating token representations toward the safe side, and suppresses positively aligned unsafe residual components. Across broad evaluation, CALM improves unsafe-content suppression while preserving benign utility, demonstrating that local counterfactual correction provides a more selective alternative to global unsafe-signal removal.
KiteNorm: Variance Regularisation for Stable and Scalable Post-LN Transformers
Leon A Trochelmann ⋅ Sajad Movahedi ⋅ Shiwei Liu ⋅ Antonio Orvieto
Post-LayerNorm Transformers have seen limited practical adoption due to enduring difficulties in scaling them to large depths. While prior research has focused on stabilising Post-LayerNorm by improving conditioning at initialisation, stability often deteriorates when models are trained with large learning rates, forcing additional compromises. In the decoder-only settings we study, existing Post-LayerNorm methods fail to outperform strong Pre-LayerNorm baselines. We propose KiteNorm, a novel normalisation method designed to break this trend. KiteNorm achieves stability through residual scaling and a regularisation technique based on hidden-state variance. Beyond stabilisation, KiteNorm learns separate scales for the skip and residual branches in each sublayer, improving performance. Across decoder-only Transformers up to 1B parameters, KiteNorm remains stable throughout training and consistently outperforms leading baselines, with scaling laws favouring KiteNorm across depth, width, batch size, and training budget. Ablations show that each component is necessary for the full stability and performance gains. We offer theoretical insights into KiteNorm through a new perspective on the gradient vanishing problem in Post-LayerNorm, linking the underlying hidden-state expansion to rank collapse.
Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use
Abhijit Kumar ⋅ Zoey WU ⋅ Mohit Suley
Humans know when to reach for help e.g. $347 \times 28$ warrants a calculator while $2 + 2$ does not. Language models, by default, do not. Prompt-based approaches can instruct a model when to invoke tools, but this external scaffolding does not teach the model to recognize the boundary of its own knowledge. Reinforcement learning approaches that assign a single outcome reward to the whole trajectory fare no better: trajectory-level credit cannot isolate which tool call in a successful episode actually helped, nor penalize unnecessary calls. We propose \textbf{CARL} (\textbf{C}ompetence-\textbf{A}ware \textbf{R}einforcement \textbf{L}earning), which trains a critic on the model's own rollouts to learn where the model's parametric knowledge suffices and where it needs external help. By decomposing each rollout at natural tool-use boundaries (e.g., code fence delimiters and context block transitions), CARL assigns independent credit to each segment from a single binary outcome, without external judges or step-level annotations, addressing the credit-assignment limitations of trajectory-level methods. As a result, erroneous tool calls, incorrect extractions, and unnecessary calls each receive appropriately signed advantages under a single scheme. We show quantitatively and qualitatively that the trained critic captures the model's domain competence: it separates parametrically solvable from tool-dependent questions with AUC 0.93 at 7B. On five benchmarks spanning arithmetic, multi-hop factual QA, and numerical reasoning over financial tables, CARL improves exact-match accuracy by 6.7 points at 7B and 10.6 points at 3B over the strongest trajectory-level baseline (Search-R1 PPO), with the largest gain (+8.3 EM at 7B, +9.0 EM at 3B) on Musique, the most compositional multi-hop benchmark. Compared to a representative trajectory-level baseline (Search-R1 GRPO), the model issues 56\% fewer tool calls on questions answerable from parametric knowledge while remaining ${\sim}10$ EM points more accurate on those same questions, an emergent consequence of the critic learning where this model's competence ends. Gains are largest at small scale: the 3B model improvement is $1.6\times$ the 7B improvement, suggesting that knowing when to ask for help disproportionately benefits models with smaller parametric memory.
Know Where You Stand: Memory-Source Choice in Long-Context Dialogue Agents
Caishen Zhou ⋅ Yihong Tang ⋅ Xuefeng Bai ⋅ Kehai Chen ⋅ Min zhang
Memory enables long-context dialogue agents to maintain user state across many turns, but it also introduces a fundamental source-choice problem: grounding the response in the immediate context or retrieved prior memory? A wrong choice incurs two complementary costs: neglecting valid aged-out evidence, or overwriting sufficient context with stale memory. Existing systems and benchmarks largely conflate retrieval with source choice, relying on the model to resolve it implicitly and leaving this critical decision largely unexplored. In this paper, we introduce MSBench, a controlled benchmark that isolates source choice by pairing questions answerable exclusively from context with those requiring prior memory. Under this protocol, strategies that depend on the model's implicit judgment exhibit sharply degraded performance, exposing the limits of leaving source choice ungoverned. To this end, we propose the Memory-Source Reasoner (MSR), which frames source choice as an explicit metacognitive decision: it reasons about contextual sufficiency, answers directly when appropriate, and selectively retrieves missing evidence otherwise. Experimental results show that MSR achieves the best overall accuracy and selection accuracy across four answer backbones. On the primary GPT-5-mini run, it improves over the best non-MSR baselines by 12.6 points in overall accuracy and 6.3 points in selection accuracy. Our benchmark, code, prompts, and evaluation scripts are available at https://anonymous.4open.science/r/msbench-msr-0C72.
LangFlow: Continuous Diffusion Rivals Discrete in Language Modeling
Yuxin Chen ⋅ Chumeng Liang ⋅ Hangke Sui ⋅ Ruihan Guo ⋅ Chaoran Cheng ⋅ Jiaxuan You ⋅ Ge Liu
Continuous diffusion has served as a foundation for high-fidelity, controllable, and few-step generation across continuous data modalities such as images, videos, and molecular structures. However, in language modeling, prior continuous diffusion language models (DLMs) lag behind discrete counterparts. Existing categorical and simplex-based approaches operate over extremely large and sparse language spaces, while prior embedding-space approaches avoid this sparsity but lack a well-selected design space. In this work, we close this gap with LangFlow by connecting embedding-space DLMs to Flow Matching, alongside three key innovations: (1) we derive a novel ODE-based NLL bound for principled evaluation of continuous flow-based language models; (2) we propose an information-uniform principle for setting the noise schedule, which motivates a learnable noise scheduler based on a Gumbel distribution; and (3) we revise prior training protocols by incorporating self-conditioning, which improves both likelihood and sample quality for embedding-space DLMs and behaves differently from its use in discrete diffusion. Putting everything together, LangFlow is competitive with top discrete DLMs on both perplexity (PPL) and generative perplexity (Gen. PPL), reaching a PPL of 30.0 on LM1B and 24.6 on OpenWebText. It also exceeds autoregressive baselines in zero-shot transfer on 4 out of 7 benchmarks. As the first continuous DLM shown to rival discrete diffusion in both generative quality and perplexity, LangFlow provides clear evidence that continuous diffusion is a promising paradigm for language modeling. Code will be released upon acceptance.
Language-Assisted Image Clustering Guided by Discriminative Relational Signals and Adaptive Semantic Centers
Jun Ma ⋅ Xu Zhang ⋅ Zhengxing Jiao ⋅ Yaxin Hou ⋅ Hui LIU ⋅ Junhui Hou ⋅ Yuheng Jia
Language-Assisted Image Clustering (LAIC) augments the input images with additional textual information with the help of vision-language models (VLMs) to improve clustering performance. Despite recent progress, existing LAIC methods mainly construct image-text pairs for text-modality utilization, which overlook two issues: (i) textual features constructed for images are highly similar, leading to weak inter-class discriminability; (ii) the training is restricted to pre-built image-text alignments, limiting the potential for better utilization of the text modality. To address these issues, we propose a new LAIC framework with two complementary components. First, we exploit cross-modal relations to generate more discriminative self-supervision signals for clustering, which is compatible with the pre-training mechanisms of VLMs such as CLIP. Second, we learn category-wise continuous semantic centers via prompt learning to produce the final clustering assignments. Extensive experiments on eight benchmark datasets demonstrate that our method achieves an average improvement of 2.7\% over state-of-the-art methods, and the learned semantic centers exhibit strong interpretability. Code is available in the supplementary material.
LAPLEX: The FFT of Learnable Laplace Kernels
Łukasz Struski ⋅ Hanna Blazhko ⋅ Piotr Kubaty ⋅ Jacek Tabor
Fast linear algebra in deep learning usually comes with a choice: fixed geometry and exact computation, as in the Fourier transform, or adaptive geometry paid for by dense parameters, random features, or low-rank surrogates. To move beyond this trade-off, we introduce LAPLEX, a class of exact, trainable (phased) Laplace-kernel operators. A LAPLEX layer is a typically full-rank dense matrix, implicitly defined by learnable coordinate anchors, with FFT-like scaling. Consequently, it supports trainable matrix--vector operations at vector dimensions up to $10^9$ on modern GPUs. As a neural layer, it yields compact projections and classification heads interpretable as soft, trainable routing models. The same primitive also serves as an efficient Gram operator, enabling high-dimensional covariance models on flattened images of dimension $3 \cdot 10^6$ that preserve visible spatial structure without imposing convolutional bias. These applications reflect a single principle: dense geometry can be learned without storing a dense matrix, which enables data-adaptive global interactions in regimes where ordinary dense layers are out of reach. In this sense, LAPLEX separates expressivity from storage cost: it behaves like a dense trainable matrix, but is represented and applied through a small structured set of parameters.
Large Language Models as Graph Computational Solvers via Topology-aware Residual Attention
Wei Zhuo ⋅ Siqiang Luo
Large language models increasingly serve as general-purpose reasoning engines, yet they remain unreliable on graph computation tasks whose inputs are discrete, permutation-equivalent, and algorithmic. We propose TRACER, a post-training framework that turns LLMs into native graph computational solvers. Our central thesis is that graph computation failures arise from mismatches within the Transformer computation itself: token order obscures graph symmetries, self-attention is biased toward textual proximity instead of graph topology, and autoregressive decoding weakly constrains algorithmic traces. TRACER addresses these challenges at three coupled levels. At the input level, we introduce permutation-invariant graph encoding, which encourages the LLM to map different serialization variants of a graph to a shared internal graph representation, thereby preserving graph symmetries. At the representation level, we enhance self-attention with the Topology-aware Residual Attention mechanism to inject graph topological signals into the attention kernel. At the reasoning level, we employ process-reward RL to encourage faithful execution of graph algorithmic traces. Across linear-time, polynomial-time, and NP-complete graph tasks, TRACER delivers substantial gains over baselines. These results suggest that equipping LLMs with graph-aware internal computation offers a practical path toward reliable neural solvers for structured algorithmic reasoning. The code is anonymously available here.
Latent Action Reparameterization for Efficient Agent Inference
Wenhao Huang ⋅ Qingwen Zeng ⋅ Qiyue Chen ⋅ Zijie Guo ⋅ Yu Sun ⋅ Cheng Yang ⋅ Siru Ouyang ⋅ Jiri Gesi ⋅ Fang Wu ⋅ Jiayi Zhang ⋅ Huaming Chen ⋅ Bang Liu ⋅ Robert Tang ⋅ Chenglin Wu
Large language model (LLM) agents often rely on long sequences of low-level textual actions, resulting in large effective decision horizons and high inference cost. While prior work has focused on improving inference efficiency through system-level optimizations or prompt engineering, we argue that a key bottleneck lies in the representation of the action space itself. We propose Latent Action Reparameterization (LAR), a framework that learns a compact latent action space in which each latent action corresponds to a multi-step semantic behavior. By reparameterizing agent actions into latent units, LAR enables decision making over a shorter effective horizon while preserving the expressiveness of the original action space. Unlike hand-crafted macros or hierarchical controllers, latent actions are learned from agent trajectories and integrated directly into the model, allowing both planning and execution to operate over abstract action representations. Across a range of LLM-based agent benchmarks, LAR significantly reduces the effective action horizon and improves inference efficiency under fixed compute budgets. As a consequence, our approach achieves substantial reductions in action tokens and corresponding wall-clock inference time, while maintaining or improving task success rates. These results suggest that action representation learning is a critical and underexplored factor in scaling efficient LLM agent inference, complementary to advances in model architecture and hardware.
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
Yifan Dai ⋅ zhenhua wu ⋅ Bohan Zeng ⋅ Daili Hua ⋅ Jialing Liu ⋅ Bozhou Li ⋅ Yuran Wang ⋅ Chengzhuo Tong ⋅ Hao Liang ⋅ Xiaochen Ma ⋅ Junbo Niu ⋅ Tianyu Guo ⋅ Yang Shi ⋅ Yue Ding ⋅ Yiyan Ji ⋅ Bingyin Mei ⋅ Yushuo Guan ⋅ Yuanxing Zhang ⋅ Pengfei Wan ⋅ Fangcheng Fu ⋅ Wentao Zhang
While joint audio-visual understanding is fundamental to advancing machine cognition, current multimodal large language models (MLLMs) still struggle with complex cross-modal reasoning. Existing text-based chain-of-thought (CoT) compresses rich multi-modal features into discrete text, incurring information loss and inducing a language-bound phenomenon that diminishes attention to audio-visual signals. In contrast, a continuous latent space inherently preserves dense representations, serving as an ideal carrier for audio-visual information. Motivated by this, we propose $\textbf{LatentOmni}$, a novel cross-modal reasoning framework. By introducing a feature-level supervision mechanism to directly reconstruct raw sensory inputs within the latent space, LatentOmni leverages native latent features to bridge audio-visual modalities and text, ensuring sustained attention on original audio-visual inputs throughout reasoning. Furthermore, to maintain temporal consistency across modalities in latent space, we design Omni-Sync Position Embedding (OSPE), which generalizes multimodal rotary position encodings to drive audio-visual synchrony. To supervise this reasoning process, we construct LatentOmni-Instruct-35K, a dataset interleaving text with audio-visual segments that serve as dense evidence for latent reconstruction. Comprehensive evaluation across multiple audio-visual reasoning benchmarks demonstrates that LatentOmni substantially outperforms strong explicit-CoT baselines, validating latent space joint reasoning as a promising path toward genuine omnimodal understanding.
Latent Riemannian Flow Matching for Geometry-Grounded 3D Foundation Models
Lisa Weijler ⋅ Irene Ballester ⋅ Guofeng Mei ⋅ Tolga Birdal ⋅ Pedro Hermosilla
Geometric foundation models, such as the Visual Geometry Grounded Transformer (VGGT), provide strong 3D priors from unposed images. However, such models operate purely in a feed-forward, deterministic regime, \ie~they cannot generate plausible geometry beyond what the input views directly support. Generative models for 3D scenes, on the other hand, must rely on strong geometric priors to produce coherent outputs from sparse inputs. We bridge these two paradigms by performing flow matching directly in VGGT's latent space, leveraging its learned 3D priors without committing to any explicit downstream representation such as Gaussians, meshes, or video-VAE latents. This requires respecting the latent geometry: VGGT tokens occupy a product of high-dimensional hyperspheres on which standard Euclidean flow matching fails. We address this with a Riemannian Flow Matching framework defined on a product manifold of four hyperspheres, aligned with VGGT's multi-scale encoder, which keeps generated tokens on the valid data manifold required by the frozen decoding heads. On RealEstate10K and ScanNet++, our method achieves strong performance against recent scene generation baselines in both per-view 3D geometry and appearance, establishing latent-space flow matching on geometric foundation models as a viable paradigm for 3D generation.
LDM-is-AE: Latent Diffusion is an Intrinsic Auto-Encoder for End-to-End Image Generation
Zhengqiang ZHANG ⋅ Lingchen Sun ⋅ Rongyuan Wu ⋅ Qiaosi Yi ⋅ Xiangtao Kong ⋅ Chaodong Xiao ⋅ Lei Zhang
Latent Diffusion Models (LDMs) typically adopt a two-stage pipeline: an auto-encoder (AE) is first pre-trained to define a latent space, then a diffusion model is trained to perform denoising within it. Such a two-stage design introduces a representation mismatch, as the latent space is optimized for reconstruction rather than adapting the denoising dynamics. We reveal that the LDM itself is intrinsically an AE, and consequently present LDM-is-AE, an end-to-end one-stage LDM training framework that eliminates the need for a separately trained tokenizer. Our key observation is that the LDM backbone actually performs a latent-to-feature-to-latent transformation at each denoising step, which can be interpreted as an internal decoding--encoding process. Leveraging this structure, we split the DiT backbone into two reciprocal components, DiT-E (i.e., DiT Encoding) and DiT-D (i.e., DiT Decoding), and impose image-space supervision on the intermediate features across all timesteps. Our model encourages the internal representation to align with the image domain throughout denoising, thereby establishing an explicit latent-to-image-to-latent path. At the zero-noise timestep, our model further performs an image-to-latent-to-image mapping, corresponding to an auto-encoding process. As a result, LDM-is-AE jointly learns latent representations and denoising dynamics in an end-to-end manner, yielding a diffusion-native latent space tailored to the generation process. Experiments demonstrate that LDM-is-AE exhibits highly competitive generation performance, achieving an FID of 1.84 on class-conditional image generation. Code and models will be released.
Leaderboard Hacking: Preference-Based Model Evaluations are Vulnerable to Manipulation
Eve Fleisig ⋅ Joachim Baumann ⋅ Dylan Hadfield-Menell ⋅ Dirk Hovy
Arena-style rankings of language models are a widely used evaluation framework in leaderboards like LMArena, research papers, and model evaluation in production. These rankings aim to realistically assess model performance using pairwise comparisons of anonymous models on user prompts, aggregating votes via the Bradley-Terry model to produce the final ranking. However, we find that these leaderboards are highly vulnerable to several manipulation strategies that unscrupulous providers could use to boost a model's rank: strategic voting (individual votes that boost a model's rank, including votes on irrelevant models), strategic prompting (choosing prompts that favor a model), and strategic nomination (boosting the rank of a model by inserting irrelevant models). We demonstrate these effects on public leaderboard data, analyze the circumstances that cause these vulnerabilities, and describe approaches to safeguard against each attack. Notably, these issues hold for any pairwise preference-based evaluation that uses the Bradley-Terry model. These manipulations boost the rank of nearly every tested model; each attack can boost a model by up to 15-21 places in a ranking and, combined, they can boost a model by up to 45 places.
Learning-Augmented Approximation for Unrelated-Machines Makespan Scheduling
Kaito Baba ⋅ Evripidis Bampis ⋅ Georgios Mitropoulos
Recently, Antoniadis et al. (ICLR 2025) proposed a framework for incorporating predictions to approximate NP-hard selection problems. Despite its simplicity, this approach tightly matches theoretical lower bounds, making its generalization highly compelling. We address an open question raised in the work of Antoniadis et al., concerning the extension of this approach to other important problems outside the class of selection problems, such as scheduling. We develop a learning-augmented algorithm for the makespan minimization problem on unrelated machines, denoted by $R||C_{\max}$. By using predictions of heavy job assignments, we achieve a polynomial-time $(1+\varepsilon)$-approximation for accurate predictions that smoothly degrades to a worst-case 2-approximation as the error increases. We conclude our work with an empirical analysis of our method.
Command line interface (CLI) agents are emerging as a practical paradigm for agent-computer interaction over evolving filesystems, executable command line programs, and online execution feedback. Recent work has used reinforcement learning (RL) to learn these interaction abilities from verifiable task feedback, yet few methods exploit the native structured attributes of CLI actions as learning signals. Beyond this underused action structure, CLI learning also couples two bottlenecks for coding agents. First, the agent must identify task-relevant evidence in a large codebase from partial observations. Second, sparse terminal rewards must be assigned to the actions that shape a long multi-turn trajectory. We study these bottlenecks through shell-driven information extraction and file editing tasks. For selective observation, we introduce $\sigma$-Reveal, an inference-time mechanism that selects token-budgeted context for the same CLI. For credit assignment, we propose Action Advantage Assignment ($\mathrm{A}^3$), a native agentic RL method that preserves the algorithmic complexity of standard agentic RL. $\mathrm{A}^3$ constructs turn-level advantages from episode-level relative feedback, abstract syntax tree (AST) based action sub-chain residuals, and tree-level trajectory margins. To further evaluate this problem setting, we construct ShellOps, a verifiable dataset suite covering CLI tasks in repository environments.
Learning Deployable Causal Action Geometry under Temporal Non-Stationarity
Changjian Liu ⋅ Tianyu Wang ⋅ Yuwei Xu ⋅ Xiaoxuan Deng ⋅ Zhilin Zhang ⋅ Yong Gao ⋅ Chuan Yu ⋅ Jian Xu ⋅ Bo Zheng
Learning a response law for continuous decisions from observational time series is central to many real-world decision systems, yet temporal non-stationarity poses a fundamental challenge: outcome levels drift over time, causing naive models that fit raw outcomes to conflate treatment effects with shifting baseline conditions. Even when the one-step response is identifiable, unrestricted cross-time generalization remains fundamentally ill-posed. Thus, we introduce a constrained decision target based on a causal action geometry with anchored link-scale morphologies. Our approach constructs anchored contrasts over a family of link functions, isolating the action-dependent component of the response that can be shared across periods while allowing flexible, period-specific calibration. This yields a morphology-restricted oracle target for continuous decision making under temporal drift. Building on this formulation, we develop a two-stage orthogonal estimator that first obtains cross-fitted pilot estimates of anchored responses and then performs honest profiled selection over morphology and representation classes. The resulting estimator admits oracle guarantees for the restricted target and, under approximate persistence, produces decision rules that remain near optimal in future periods. Experiments on synthetic data and a large-scale e-commerce pricing application support the proposed target and its decision benefits under temporal shift.
Learning Motion-Appearance Coupling Priors for Solving Video Inverse Problems
Anselm Krainovic ⋅ Reinhard Heckel
Solving inverse problems on video requires priors over both frame appearance and temporal dynamics. Many recent approaches combine pretrained image diffusion priors with motion guidance, sidestepping the costs of training large video models. Temporal consistency is often enforced via photometric losses or warping-based constraints, which implicitly assume that motion explains frame-to-frame or latent-to-latent changes up to small deviations, an assumption that fails under occlusions, lighting changes, and non-rigid deformation. In this paper, we propose to learn the coupling between motion and frame appearance as a generative prior. We learn the distribution of motion-induced appearance residuals with a conditional diffusion model and use it to regularize reconstruction in video inverse problems. Our motion-appearance prior acts as a plug-in regularizer for diffusion-based video reconstruction, preserving the benefits of image diffusion models while remaining compatible with existing appearance regularization techniques. Experiments on dynamic video inverse problems show that, depending on the task, reconstruction with a learned motion-appearance coupling prior matches or outperforms feature-consistency and noise-warping baselines.
Learning Robust Reasoning through Guided Adversarial Self-Play
Shuozhe Li ⋅ Vaishnav Tadiparthi ⋅ Kwonjoon Lee ⋅ Nakul Agarwal ⋅ Hossein Nourkhiz Mahjoub ⋅ Ehsan Moradi Pari ⋅ Lizhang Chen ⋅ Du Cheng ⋅ Amy Zhang ⋅ Liu Leqi
Reinforcement learning from verifiable rewards (RLVR) produces strong reasoning models, yet these models can fail catastrophically when the conditioning context is fallible (e.g., corrupted chain-of-thought, misleading partial solutions, or mild input perturbations), because standard RLVR optimizes final-answer correctness only under clean conditioning. We introduce GASP (Guided Adversarial Self-Play), a robustification method that explicitly trains error detection and repair capabilities using only outcome verification. Without human labels or external teachers, GASP instantiates an adversarial self-play game within a single model: a polluter learns to induce failure through locally coherent corruptions, while an agent learns to diagnose and recover under the same corrupted conditioning. To address the scarcity of successful recoveries early in training, we propose in-distribution repair guidance, an auxiliary imitation objective on self-generated repairs that increases recovery probability while preserving previously acquired capabilities. Across four open-weight models (1.5B–8B), GASP converts strong-but-brittle reasoners into substantially more robust ones that better withstand misleading and perturbed context, while often improving clean accuracy. Further analysis shows that adversarial corruptions create an effective curriculum, and that in-distribution guidance enables rapid recovery learning with minimal representational drift.
Learning to Align Generative Appearance Priors for Fine-grained Image Retrieval
Shijie Wang ⋅ Yadan Luo ⋅ Zijian Wang ⋅ Xin Yu ⋅ Zi Huang
Fine-grained image retrieval (FGIR) typically relies on supervision from seen categories to learn discriminative embeddings for retrieving unseen categories. However, such supervision often biases retrieval models toward the semantics of seen categories rather than the underlying appearance characteristics that generalize across categories, thereby limiting retrieval performance on unseen categories. To tackle this, we propose GAPan, a Generative Appearance Prior alignment network that reformulates the learning objective from category prediction toward appearance modeling. Technically, GAPan treats retrieval features with an invertible density model based on normalizing flows. In the forward direction, the flow maps all instance features into a latent density space, where each seen category is modeled by a class-conditional Gaussian prior and optimized via exact likelihood estimation. This formulation preserves richer appearance details by leveraging the invertible property of the flows. In the reverse direction, samples from the high-density regions of these learned priors are mapped back to the feature space to produce appearance-aware anchors that reflect intra-category variation. These anchors supervise a prior-driven alignment objective that aligns retrieval embeddings with category-specific appearance distributions, thereby improving generalization to unseen categories. Evaluations demonstrate that our GAPan achieves state-of-the-art performance on both widely-used fine- and coarse-grained benchmarks.
Learning to Continually Learn via Meta-learning Agentic Memory Designs
Yiming Xiong ⋅ Shengran Hu ⋅ Jeff Clune
The statelessness of foundation models bottlenecks agentic systems’ ability to continually learn, a core capability for long-horizon reasoning and adaptation. To address this limitation, agentic systems commonly incorporate memory modules to retain and reuse past experience, aiming for continual learning during test time. However, most existing memory designs are human-crafted and fixed, which limits their ability to adapt to the diversity and non-stationarity of real-world tasks. In this paper, we introduce ALMA (Automated meta-Learning of Memory designs for Agentic systems), a framework that meta-learns memory designs to replace hand-engineered memory designs, therefore minimizing human effort and enabling agentic systems to be continual learners across diverse domains. Our approach employs a Meta Agent that searches over memory designs expressed as executable code in an open-ended manner, theoretically allowing the discovery of arbitrary memory designs, including database schemas as well as their retrieval and update mechanisms. Extensive experiments across four sequential decision-making domains demonstrate that the learned memory designs enable more effective and efficient learning from experience than state-of-the-art human-crafted memory designs on all benchmarks. When developed and deployed safely, ALMA represents a step toward self-improving AI systems that learn to be adaptive, continual learners.
Learning to Explore with Parameter-Space Noise: A Deep Dive into Parameter-Space Noise for Reinforcement Learning with Verifiable Rewards
Bizhe Bai ⋅ Xinyue Wang ⋅ Peng Ye ⋅ Tao Chen
Reinforcement Learning with Verifiable Rewards (RLVR) enhances large language model (LLM) reasoning, yet growing evidence indicates an \emph{exploration ceiling}: they tend to reweight existing solution traces and produce low-diversity rollouts. This restricts the discovery of novel solution perspective and weakens out-of-distribution generalization. We address this limitation by directly perturbing the rollout policy's parameters while updating the clean policy model as usual, an approach that achieves better temporal consistency and yields more complex behavior patterns. To stabilize learning from these perturbed rollouts, which are inherently off-policy, we incorporate truncated importance sampling. Furthermore, we introduce a lightweight, dynamic noise scheduler based on semantic diversity and model self-certainty to optimally adjust the noise scale. Instantiated on GRPO, our method (PSN-GRPO) demonstrates its efficacy through four key contributions: (1) it improves high-budget pass@$k$, increases semantic and operational diversity, increases CoT consistency, and remains complementary to exploration-oriented RLVR techniques; (2) it validates efficacy across multiple models, including Qwen2.5, Qwen3, and Llama-3.1; (3) it shows robust performance on out-of-distribution science and coding benchmarks; and (4) it uncovers novel solution perspectives for problems the original model could not solve, as shown by qualitative analysis.
Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation
Yuchen Cai ⋅ Ding Cao ⋅ Liang Lin ⋅ Chunxi Luo ⋅ Xin Xu ⋅ Kai Yang ⋅ Weijie Liu ⋅ Saiyong Yang ⋅ Tianxiang Zhao ⋅ Guangzhong Sun ⋅ Guiquan Liu ⋅ Junfeng Fang
On-policy distillation (OPD) has emerged as an efficient post-training paradigm for large language models. However, existing studies largely attribute this advantage to denser and more stable supervision, while the parameter-level mechanisms underlying OPD's efficiency remain insufficiently understood. In this work, we argue that OPD's efficiency stems from a form of ``foresight'': it establishes a direct and stable path from the initial model to the final model early in training. This foresight manifests in two aspects. First, at the \textbf{Module-Allocation Level}, OPD identifies regions with low marginal utility and concentrates updates on modules that are more critical to reasoning. Second, at the \textbf{Update-Direction Level}, OPD exhibits stronger low-rank concentration, with its dominant subspace aligning closely with the final update subspace early in training. Motivated by this theory, we propose \textbf{EffOPD}, a plug-and-play acceleration method that speeds up OPD by searching for an effective step size and extrapolating along the current update direction. EffOPD requires no additional trainable modules or complex hyperparameter tuning, and achieves up to $3.5\times$ training acceleration while maintaining comparable final performance. Overall, our findings provide a parameter-dynamics perspective for understanding the efficiency of OPD and offer practical insights for designing more efficient post-training methods for large language models. Our code is available at: https://anonymous.4open.science/r/EffOPD-7C58.
Learning to Generate Multiple Objects from Dense and Occluded Layouts
Bach H Ngo ⋅ Ngo Tri ⋅ Hieu Le ⋅ Trung Nghia Le
Text-to-image diffusion models fail to generate correct object counts in dense scenes, where overlapping instances collapse into indistinguishable structures despite appearing visually plausible. We identify this as instance ownership collapse: tokens from overlapping objects interact freely through attention, while heavily occluded instances receive weak supervision due to their small visible areas. We address this through layout-aware attention biases that softly bias token interactions toward region-consistent grouping and suppress cross-instance leakage, paired with an amodal-balanced loss that amplifies gradients for occluded objects based on their occlusion level. To enable systematic evaluation, we introduce OverlapDepth-45K, a benchmark of densely overlapping scenes with amodal supervision. Our approach substantially improves count accuracy and prevents instance merging while preserving image quality.
Learning Weakly Communicating Average-Reward CMDPs: Strong Duality and Improved Regret
Kihyun Yu ⋅ Beomhan Baek ⋅ Dabeen Lee
We study infinite-horizon average-reward constrained Markov decision processes (CMDPs) under the weakly communicating assumption. Our contributions are twofold. First, we establish strong duality for weakly communicating average-reward CMDPs over stationary policies with finite state and action spaces. Despite the absence of a linear programming formulation and the resulting nonconvexity under the weakly communicating setting, we show that strong duality still holds by carefully exploiting the geometric structure of the occupation measure set. Second, building on this result, we propose a primal--dual clipped value iteration algorithm for learning weakly communicating average-reward linear CMDPs. Our algorithm achieves regret and constraint violation bounds of $\widetilde{\mathcal{O}}(T^{2/3})$, improving upon the best known bounds, where $T$ denotes the number of interactions. Our approach extends clipped value iteration to the constrained setting and adapts it to a finite-horizon approximation, which stabilizes the dual variable and is crucial for achieving improved regret bounds. To analyze this, we develop a novel approach based on strong duality that enables the decomposition of the composite Lagrangian regret into separate bounds on regret and constraint violation.
Learning When to Collaborate: Selective Multi-Agent Medical Reasoning via Uncertainty-Aware Routing
Jiacheng Hou ⋅ Xueliang Cui ⋅ Dan Lu ⋅ Ruxin Wang
Multi-agent systems have emerged as a promising paradigm for multimodal medical reasoning, but existing methods rely on two assumptions that hinder practical deployment: (i) always-active dense collaboration among multiple agents, and (ii) dependence on large cloud-based models that raise privacy and efficiency concerns. In this paper, we challenge both assumptions and propose SMART-Med, a fully local multi-agent method that treats collaboration as a learnable decision rather than a default behavior. Our key insight is that not every medical query requires costly multi-agent deliberation: many can be reliably solved by a single well-trained agent, while only uncertain cases benefit from expert collaboration. Based on this insight, SMART-Med introduces an uncertainty-driven routing mechanism that adaptively decides when to answer directly and when to recruit a small subset of expert agents, effectively converting dense multi-agent reasoning into selective, query-adaptive coordination. To make this routing reliable, we curate a new multimodal medical dataset MedVR-3K with high-quality reasoning traces and a difficulty hierarchy, and then fine-tune local agents using GRPO with two complementary rewards: a certainty-aware accuracy reward and a vision-language reward model for correctness and reasoning quality respectively. Experiments on four medical visual question answering benchmarks show that SMART-Med achieves competitive performance compared to existing MAS approaches while substantially reducing inference cost (reducing inference time by more than 3.8$\times$ and token usage by 72\%). Moreover, after fine-tuning on MedVR-3K, SMART-Med yields an average improvement of 6.7\% across Slake, PathVQA, VQA-RAD, and PMC-VQA. Our results suggest that selective collaboration can be an efficient yet effective alternative to dense collaboration and enable practical, privacy-preserving deployment of multi-agent medical AI.
LEVDA: Latent Ensemble Variational Data Assimilation via Differentiable Dynamics
Phillip Si ⋅ Peng Chen
Long-range geophysical forecasts are fundamentally limited by chaotic dynamics and numerical errors. While data assimilation can mitigate these issues, classical variational smoothers require computationally expensive tangent-linear and adjoint models. Conversely, recent efficient latent filtering methods often enforce weak trajectory-level constraints and assume fixed observation grids. To bridge this gap, we propose Latent Ensemble Variational Data Assimilation (LEVDA), an ensemble-space variational smoother that operates in the low-dimensional latent space of a pretrained differentiable neural dynamics surrogate. By performing four-dimensional ensemble-variational (4DEnVar) optimization within an ensemble subspace, LEVDA jointly assimilates states and unknown parameters without the need for adjoint code or auxiliary observation-to-latent encoders. Leveraging the fully differentiable, continuous-in-time-and-space nature of the surrogate, LEVDA naturally accommodates highly irregular sampling at arbitrary spatiotemporal locations. Across three challenging geophysical benchmarks, LEVDA matches or outperforms state-of-the-art latent filtering baselines under severe observational sparsity while providing more reliable uncertainty quantification. Simultaneously, it achieves substantially improved assimilation accuracy and computational efficiency compared to full-state 4DEnVar.
LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale
Francesco Mantegna ⋅ Dulhan Jayalath ⋅ Gereon Elvers ⋅ Tasha Kim ⋅ Benjamin Ballyk ⋅ SungJun Cho ⋅ Teyun Kwon ⋅ Luisa Kurth ⋅ Miran Özdogan ⋅ Gilad Landau ⋅ Pratik Somaiya ⋅ Natalie Voets ⋅ Mark Woolrich ⋅ Oiwi Parker Jones
We introduce LibriBrain100, a large-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation. LibriBrain100 more than doubles the size of the original LibriBrain release, resulting in over 100 hours of high-quality MEG acquired while subjects listened to naturalistic continuous speech. With ~80 hours from a single subject, LibriBrain100 sets a new record for ‘deep’, within-subject neural data (8x more than the next comparable dataset and roughly 80x more than other datasets). To demonstrate the payoff of this depth-first design, we evaluate on a word-classification benchmark—an increasingly well-established stepping stone towards the open challenge of non-invasive brain-to-text decoding. Using an existing decoding model, we achieve state-of-the-art performance—validating both the quality of the recordings and the value of within-subject data at scale. Because collecting 80 hours of data per user is impractical for real-world applications, we also collected ~40 minutes of additional data from each of 32 subjects. Using the same word-classification benchmark, we demonstrate the value of ’broad’ multi-subject data: supervised fine-tuning of a pre-trained model can substantially compensate for limited per-subject data. We provide standard train, validation, and test splits, all reproducible through an open-sourced Python library that supports easy downloading, optional preprocessing, and data loading for common deep learning frameworks. In addition, the dataset and evaluation infrastructure are being released alongside an open machine-learning competition with a public leaderboard for standardised benchmarking. Ultimately, our hope is that LibriBrain100 will accelerate progress towards practical non-invasive brain-computer interfaces, capable of restoring communication to people living with severe paralysis.
LIME: Link-based User-item Interaction Modeling with Decoupled XOR Attention for Efficient Test Time Scaling
Yunjiang Jiang ⋅ Ayush Agarwal ⋅ Yang Liu ⋅ Yihan Wu ⋅ Haoran Liu ⋅ Bi Xue
Scaling large recommendation systems requires advancing three major frontiers: processing longer user histories, expanding candidate sets, and increasing model capacity. While promising, transformers' computational cost scales quadratically with the user sequence length and linearly with the number of candidates. This trade-off makes it prohibitively expensive to expand candidate sets or increase sequence length at inference, despite the significant performance improvements. We introduce **LIME**, a novel architecture that resolves this trade-off. Through two key innovations, LIME fundamentally reduces computational complexity. First, low-rank ``link embeddings" enable pre-computation of attention weights by decoupling user and candidate interactions, making the inference cost nearly independent of candidate set size. Second, a linear attention mechanism, **LIME-XOR**, reduces the complexity with respect to user sequence length from quadratic ($O(N^2)$) to linear ($O(N)$). Experiments on public and industrial datasets show LIME achieves near-parity with state-of-the-art transformers but with a 10$\times$ inference speedup on large candidate sets or long sequence lengths. When tested on a major recommendation platform, LIME improved user engagement while maintaining minimal inference costs with respect to candidate set size and user history length, establishing a new paradigm for efficient and expressive recommendation systems.
LINK: Learning to Localize from Known to Unknown Scenes
Minghang Zhu ⋅ Kaibo Jin ⋅ Zhijing Wang ⋅ Ziwei Shi ⋅ Wen Li ⋅ Sheng Ao ⋅ Cheng Wang
LiDAR relocalization based on scene coordinate regression (SCR) degrades sharply when test trajectories traverse route segments absent from training, a practical challenge in long-range autonomous-driving deployment, where complete route coverage is rarely achievable. We present LINK, a confidence-guided framework for relocalization under partial scene coverage. LINK predicts an absolute pose hypothesis and scene features with an SCR backbone, estimates pose reliability from temporal pose-feature sequences, and uses a decision policy to arbitrate between direct absolute hypotheses and relative motion propagated from reliable historical estimates, thereby preserving globally referenced localization across uncovered segments. We also introduce UrbanUnseen, a city-scale benchmark with native known-to-unknown scene transitions, together with controlled partial-coverage protocols on Oxford and NCLT. Experiments show that LINK substantially improves robustness when training coverage is incomplete, consistently outperforms representative map-free baselines under partial coverage, and remains effective in fully known scenes.
Importance sampling provides an elegant variance-minimization principle for stochastic optimization: sampling examples proportional to their gradient norms yields a minimum-variance gradient estimator. However, in modern deep learning, importance sampling often fails to outperform uniform sampling once proxy-computation cost, stale scores, and fair update budgets are accounted for. We propose LITE, a lightweight lazy sampler for practical importance sampling. LITE maintains inexpensive proxy scores, updates them only periodically, smooths them across refreshes, and converts them into sampling probabilities through a temperature rule that prevents over-concentration. When proxy scores are informative, LITE prioritizes high-impact examples; when they are noisy or stale, the sampling distribution moves back toward uniform. We analyze the resulting variance reduction in terms of proxy quality, temporal drift, and temperature, yielding an explicit condition for when LITE improves over uniform sampling. Experiments under matched update budgets and end-to-end wall-clock accounting show that LITE is competitive with or better than uniform sampling across vision, text, and 3D benchmarks, with the clearest gains in scarce-budget and imbalanced regimes.
LITHE: Lattice-Indexed Twin Hadamard Encoding for Diffusion Personalization
Jian Jiang ⋅ Oya Celiktutan ⋅ Yaohui WANG ⋅ Yutong Ban
Diffusion personalization deployments serve growing catalogues of LoRA adapters, but resident-LoRA serving is capped near the GPU memory budget (about 150 rank-16 adapters on an A100 80 GB GPU running FLUX.1-dev). We introduce LITHE, a LoRA-compatible deployment representation that stores each adapter as an integer-index stream over a deterministic Hadamard codebook, and that pushes this resident-adapter ceiling to a verified 10,000 adapters on the same single GPU at sub-80 GB peak. The same on-disk indices drive two serving modes: an index-resident mode that routes a heterogeneous batch of adapters in one forward, and a decoded compiled mode within 1% of an optimized LoRA+compile baseline at 28-step inference at 1024² resolution on FLUX.1-dev (and 1.07–1.43× faster at lower resolutions). On disk, each rank-16 FLUX.1-dev adapter occupies 1.81 MB on average across the benchmark pool, a 50×/200× reduction over the matched r=16 / r=64 LoRA disk payload after the same coder.
Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generators
Zachary Novack ⋅ Stephen Brade ⋅ Haven Kim ⋅ Hugo Flores ⋅ Nithya Shikarpur ⋅ Chinmay Talegaonkar ⋅ Julian McAuley ⋅ Taylor Berg-Kirkpatrick ⋅ Anna Huang
Interactive streaming music generation promises the use of generative models for live performance and co-creation that is impossible with offline models. However, SOTA models exist in the discrete-AR regime, requiring industrial levels of compute for both training and inference. In this work, we investigate whether audio diffusion models, with their wide support in the open-source community but non-streaming bidirectional nature, can be repurposed efficiently into interactive models accessible on consumer hardware. By taking a critical look at the modern pipeline for block-wise outpainting diffusion, we identify critical inefficiencies during inference that result in strictly worse computational efficiency than their discrete-AR counterparts. We propose Live Music Diffusion Models (LMDMs), a simple modification of the generative diffusion process that recovers, and then outperforms, the inference complexity of the discrete Live Music Models (LMMs) through block-wise KV Caching. Unlike LMMs, LMDMs further enable stable post-training alignment through our novel ARC-Forcing paradigm, reducing error accumulation without any explicit RL or reward models. We demonstrate the application of LMDMs in a number of creative domains, including text-conditioned generation, sketch-based music synthesis, and jamming. We finally show how LMDMs can be used as a generative instrument in a real artist-AI collaboration, utilizing LMDMs as a “generative delay” to transform musicians’ improvisation live for variable timbral effects while running locally on a consumer gaming laptop.
Confidence-weighted routing, selective abstention, and ensemble weighting all assume that a model's stated confidence is informative about its capability on the question being asked. They presume functional metacognition, the capacity to assess one's own capabilities, without exercising them, relative to other frontier language models. Aggregate calibration is well studied, with mixed results, but the underlying structure of elicited confidence is less well understood. We decompose binary confidence judgements from 20 frontier Large Language Models (LLMs) across six benchmarks using tetrachoric factor analysis paired with pairwise calibration, asking whether two models that differ in confidence also differ in performance. On factual recall and information retrieval benchmarks the cross-model confidence matrix is approximately rank-one and a single dominant factor captures most of the latent variance. Models retrieving facts share an item-level difficulty axis and differ mainly in their decision thresholds along it. Across all benchmarks the relationship between confidence and performance collapses once items that all models agree on are removed. Inter-model pairwise calibration is small even where statistically significant, and what remains shrinks to nothing once base-rate differences along the shared factor are controlled for. Mathematical reasoning is the apparent exception, but this turns out to be a confound where reasoning models answer questions about their confidence by trying to solve them in their chain of thought, bypassing the sub-symbolic self-knowledge we seek to measure. We find no evidence for significant verbalised individuated metacognition in any tested domain. A trivial logistic regression on surface features of the confidence-judgement reasoning text beats the median confidence model's repeatedly sampled judgements on half the benchmarks, evidence that extracting uncertainty from the reasoning trace recovers signal the binary judgement does not.
Local FDR Membership Inference Attacks: Multiple Testing and the Role of Ridge Regularization
Jinyoung Hong ⋅ Bonwoo Lee ⋅ Jeongyoun Ahn
Membership inference attacks (MIAs) are commonly formulated as single-sample hypothesis tests with false positive rate control. However, realistic adversaries often test many candidate samples and aggregate the declared members, making membership inference a multiple testing problem. In this regime, per-sample calibration can produce unreliable sets of inferred members. We develop a false discovery rate (FDR)-controlled framework for membership inference based on the classical local FDR formulation. We instantiate the framework for high dimensional ridge regression. Under a conditional sufficiency condition, we establish FDR guarantee of the proposed method, and derive an asymptotic characterization of its detection power in high-dimensional regimes under an isotropic Gaussian design. Our analysis shows that, under FDR control, stronger ridge regularization reduces membership inference risk. Experiments on synthetic and a real-world dataset validate the theoretical findings and demonstrate that the proposed methods provide more reliable membership discoveries than existing attacks.
Local Guidance, Global Impact: Gaussian-Reshaped Trust Region Unlocks Behavior Transitions
Bingxu Liu ⋅ Jiashun Liu ⋅ Johan Obando Ceron ⋅ Hao Wang ⋅ Runze Liu ⋅ Pablo Samuel Castro ⋅ Aaron Courville ⋅ Ling Pan
While Proximal Policy Optimization (PPO) demonstrates strong performance in stationary settings, we show that its standard optimization paradigm struggles in continual and non-stationary environments. The failure does not stem from insufficient model capacity or overly restrictive clipping. Instead, PPO performs persistent, directionally inefficient local updates, which indicates a lack of geometry-aware guidance for accumulating meaningful behavioral change and ultimately hindering transitions toward new behavior patterns. Although divergence-based regularization introduces partial geometric awareness, its monotonically increasing penalties implicitly discourage large policy deviations, even when such shifts are necessary for effective adaptation. To address this limitation, we propose Gaussian Trust Region Policy Optimization (GTR), which reshapes the trust region using a Gaussian kernel. The resulting constraint is bounded and non-monotonic, providing strong local stability while progressively relaxing under sustained high-advantage updates. To further improve robustness, we introduce a Mixture Gaussian Anchor that adapts to recent policy trajectories, reducing variance induced by stale references. GTR is architecture-agnostic and achieves strong performance across games, simulated robotic control, open-world exploration, and language model post-training. These results demonstrate that geometry-aware trust-region design can be a promising direction for robust reinforcement learning in complex non-stationary environments.
Local Policy Manifolds for Efficient Multi-Objective Reinforcement Learning
Qiyue Xia ⋅ Tianwei Wang ⋅ J. Michael Herrmann
Multi-objective reinforcement learning (MORL) aims at optimising several, often conflicting goals to improve the flexibility and reliability of RL in practical tasks. This is typically achieved by finding a set of diverse, non-dominated policies that form a Pareto front in the performance space. However, constructing such a policy set remains computationally demanding, as it requires finding not a single optimal policy, but a set of policies to cover a wide range of trade-offs among the objectives. We introduce LLE-MORL, an approach that traces policy manifolds by utilising the local relationship between the high-dimensional policy parameter space and the performance space. This structured representation enables an efficient search within contiguous solution domains via locally linear extrapolation, allowing for the rapid generation of high-quality solutions without extensive retraining. We also provide a theoretical analysis of the policy manifold structure in order to specify the applicability of the method. Experiments across diverse continuous control domains demonstrate that LLE-MORL consistently achieves higher Pareto front quality and efficiency than state-of-the-art approaches.
Local Sparsity Enables Unsupervised LLM Safety Detection
Xin Chen, Cynthia ⋅ Gil Kur ⋅ Aleksandr Shevchenko ⋅ Andreas Krause
Deployment-time safety methods for large language models (LLMs) are predominantly supervised and assume access to unsafe training data. Nevertheless, new attacks and harm categories regularly arise, not captured by models trained in such a supervised fashion. An alternative approach is to view this problem through the lens of anomaly detection, namely, to rely solely on modeling safe data and flagging out-of-distribution inputs. However, LLM activations lie in a high-dimensional space, raising concerns about whether anomaly detection is statistically feasible. We show that, under the linear representation hypothesis (LRH), there may indeed be hope. In the LRH concept space, which is typically recovered via a sparse autoencoder (SAE), nearby points share a small common active support. Based on this local sparsity insight, we propose a framework for locally masked SAE-based anomaly detection, and establish a sample complexity bound that scales polynomially in the local sparsity and only \emph{logarithmically} in the SAE's ambient dimension. Finally, we empirically validate our framework on various architectures and datasets, including both capability-testing datasets and safety-specific datasets.
LoCo: Selective Local Competition for Discriminative Open-Vocabulary Multi-Label Recognition
Zhuoming Li ⋅ Bo Han ⋅ Xiaoyu Wang ⋅ Yuheng Jia
Open-vocabulary multi-label recognition (OV-MLR) aims to identify all queried semantic concepts present in an image. Existing methods mainly improve global image-text matching or text-side representations, leaving local prediction largely under-explored. In this paper, we revisit local prediction and reveal that it can serve as a strong discriminative branch once two properties are properly handled: spatial consistency of local visual features and compatibility-aware local competition. This perspective is motivated by local semantic sparsity: each image region is related to only a small subset of queried concepts, while compatible concepts of different granularities or types may still coexist locally. We instantiate this perspective with LoCo, a selective Local Competition framework for OV-MLR. LoCo adopts a spatially consistent visual encoder to reduce local feature entanglement, and introduces Selective Softmax to impose competition only among locally incompatible categories while preserving compatible responses. It further combines multi-scale local aggregation with global prediction to handle objects at different scales and scene-level semantics. Experiments on four multi-label benchmarks show that LoCo achieves state-of-the-art performance, with average gains of +11.3\% F1-score and +5.2\% mAP. Our results suggest that local semantic sparsity offers a useful basis for developing more discriminative open-vocabulary multi-label recognition methods. Code is available in the supplementary material.
Aligned language models refuse unsafe requests because RLHF widens the logit margin between refusal and affirmative tokens at the first decoded position. We call this scalar the refusal-affirmation logit gap and use it as a per-prompt diagnostic for alignment robustness. On three model families, alignment widens the gap on 97.5-99.8\% of toxic prompts, and gap closure tracks True ASR across suffix strategies (an internal consistency check, since our method optimises for gap closure). We present logit-gap steering, a gradient-free, forward-pass-only method that searches for in-distribution suffixes drawn from the model's own high-probability candidates whose cumulative effect closes the gap. Discovering 8 ensemble suffixes per family costs ${\approx}26{,}000$ forward-pass equivalents (${\approx}2$~min on one A100), ${\approx}125\times$ less than a single GCG search. Suffixes discovered on 0.5B-2B models transfer to 72B within family. An 8-suffix ensemble reaches 38--96\% True ASR across 13 models on AdvBench and HarmBench. Most suffixes have $10^{3}$-$10^{4}\times$ lower perplexity than GCG: under a published PPL-filter defense, our ensemble holds at 76.0\% (from 76.9\%) while GCG collapses from 64.7\% to 1.0\%.
Log-Likelihood, Simpson’s Paradox, and the Detection of Machine-Generated Text
Tom Kempton ⋅ Viktor Drobnyi ⋅ Maeve Madigan ⋅ Stuart Burrell
The ability to reliably distinguish human-written text from that generated by large language models is of profound societal importance. The dominant approach to this problem exploits the likelihood hypothesis: that machine-generated text should appear more probable to a detector language model than human-written text. However, we demonstrate that the token-level signal distinguishing human and machine text is non-uniform across the hidden space of the detector model, and naively averaging likelihood-based token scores across regions with fundamentally different statistical structure, as most detectors do, causes a form of Simpson's paradox: a strong local signal is destroyed by inappropriate aggregation. To correct for this, we introduce a learned local calibration step grounded in Bayesian decision theory. Rather than aggregating raw token scores, we first learn lightweight predictors of the score distributions conditioned on position in hidden space, and aggregate calibrated log-likelihood ratios instead. This single intervention dramatically and consistently improves detection performance across all baseline detectors and all datasets we consider. For example, our calibrated variant of Fast-DetectGPT improves AUROC from $0.63$ to $0.85$ on GPT-5.4 text, and a locally-calibrated DMAP detector we introduce achieves state-of-the-art performance across the board. That said, our central contribution is not a new detector, but a precise diagnosis of a significant cause of under-performance of existing detectors and a principled, modular remedy compatible with any token-averaging pipeline. This will serve as a foundation for the community to build upon, with natural avenues including richer distributional models, improved calibration strategies, and principled ensembling with hidden-space geometry signals via the full Bayes-optimal decision rule.
LogSig-SSM: Time-Series Modelling with Multi-Scale Log-Signature Compression for State-Space Models
Felix Oury ⋅ Nicolas C Peiro ⋅ Reiko J Tanaka
Time-series data are often sampled irregularly at high frequency and exhibit long-range dependencies, which makes long-horizon modelling difficult. Continuous-time models such as neural controlled differential equations (NCDEs) and neural rough differential equations (NRDEs) can handle irregular sampling, but they scale poorly to long sequences. Selective state-space models (SSMs) such as Mamba scale linearly with sequence length, but within a single block they provide limited recurrent mixing across hidden dimensions. We propose LogSig-SSM (Log-Signature Compression for State-Space Models), which first compresses long multivariate time series into a shorter sequence of tokens using multi-scale windowed log-signatures, and then processes these tokens with a selective SSM backbone. LogSig-SSM is scalable and robust to irregular sampling, combining log-signature tokens that capture higher-order cross-channel interactions with a selective SSM that models long-range dependencies. The model also admits a continuous-time interpretation as an NCDE/NRDE-style system driven by a log-signature-based input, in which selectivity induces an input-dependent rescaling of the latent dynamics. Across four benchmarks, namely long-sequence classification on UEA, high-frequency physiological regression on PPG-DaLiA, multivariate weather forecasting, and irregularly sampled clinical prediction on PhysioNet Sepsis, LogSig-SSM overall outperforms existing SSM and continuous-time baselines while significantly reducing training time and GPU memory usage.
LoMo: Local Modality Substitution for Deeper Vision-Language Fusion
Feng Han ⋅ Zhixiong Zhang ⋅ Zheming Liang ⋅ Yibin Wang ⋅ Jiaqi Wang
Vision-Language Models (VLMs) have achieved substantial progress across a wide range of understanding and reasoning tasks, driven by large-scale image-text training aimed at multimodal fusion. Ideally, replacing a textual question with its rendered-image counterpart should leave model performance essentially unaffected. In practice, however, such modality substitution induces dramatic performance degradation. We attribute this carrier sensitivity issue to an inherent bias in current training corpora. Across prevalent datasets such as image captioning, VQA, OCR, and web-sourced interleaved data, text and images are typically organized into distinct and asymmetric roles, with text serving as linguistic queries and images as visual references. Such data bias leads VLMs to exhibit distinct preferences for information acquisition across different modalities. Consequently, VLMs fail to align representations of semantically equivalent content across textual and visual carriers, making model reasoning fragile under modality substitution. To address this, we propose Local Modality Substitution (LoMo), a lightweight, architecture-agnostic data curation paradigm designed to provide supervision for cross-modal representational invariance between semantically equivalent text and image carriers. LoMo achieves this by reformulating single-modality prompts into seamlessly interleaved multimodal sequences. It dynamically selects target text spans and recasts them as rendered images, thereby preserving the same semantics across "text, visual, text" carriers. Extensive experiments across 13 diverse multimodal benchmarks demonstrate that LoMo significantly improves overall multimodal reasoning and yields deeper cross-modal fusion. Specifically, it delivers consistent gains across foundational models, improving over standard SFT by 2.67 points on LLaVA-OneVision-1.5-8B and 2.82 points on Qwen3.5-9B.
LongBanana: An Expert-Verified Benchmark for Long-Context Multi-Reference Image Synthesis
Haoxiang Cao ⋅ Yuxuan Zhang ⋅ Penghui Du ⋅ Bo Li ⋅ Jiajiong Cao ⋅ Yingying Fan ⋅ Zimeng Wu ⋅ Yifu Luo ⋅ Mingzhe Zheng ⋅ Junwen Miao ⋅ Huaisong Zhang ⋅ YANG YAO ⋅ Yu Wang ⋅ Dongfu Jiang ⋅ Ping Nie ⋅ Wenhu Chen ⋅ Changqian Yu ⋅ Kelsey Allen ⋅ Chaoqun Wang
We introduce LongBanana, to our knowledge the \textbf{first expert-designed and expert-verified benchmark} for \emph{long-context multi-reference image synthesis} in the age of \textbf{agent-controlled image generation}. The task requires retrieving dispersed visual evidence, rejecting distractors, binding entities and relations, and composing one coherent final image from multiple references. LongBanana contains 90 carefully designed tasks organized by a 3-by-3 taxonomy; each task includes about 32 reference images, 20 constraints, and 50 evaluation items, yielding \textbf{4{,}670 expert-specified atomic evaluation items} in total. Its graph-grounded evaluation design instantiates each task with reference images and an entity-relation graph, then converts node, edge, and whole-scene obligations into checklist items for Content Fidelity, Relational Binding, and Global Constraint Satisfaction. Reviewed expert labels include a 15\% double-labeled subset with \textbf{89.4\% exact agreement} and \textbf{0.83 quadratic weighted kappa}. We further introduce LongBanana-Gen, an agentic controller framework for long-horizon image generation; experiments show that minimal agent harnesses improve over direct generation, while ReAct/Plan-Exec/Reflect yields controller-dependent gains by strengthening evidence grounding and repair. Together, LongBanana provides an auditable testbed for measuring dense visual evidence use and a shared target for developing stronger agent-controlled image-generation systems.
Long-Horizon Agency Belongs in the Harness, Not the Context Window Only
Yi Han ⋅ YUANYUAN XU ⋅ jusheng zhang ⋅ Wenhao Wang
Long-horizon agency requires an AI system to preserve task state across hours or days of work, many tool calls, and repeated context resets. A common response is to scale a single resource: the model’s context window, together with inference-time reasoning. We argue that this response conflates long horizon with long context. Durable long-horizon state should not live only in the prompt; it should be externalized into the harness, the non-model infrastructure that manages tools, files, plans, checkpoints, sub-agents, permissions, and recovery. Publicly documented long-running agent systems already rely on such structures, including progress files, plan files, git history, rollback checkpoints, scratchpads, sub-agent decomposition, and browser-state checkpoints. We support this position with three lines of argument. First, memory-hierarchy history shows that pressure to scale one working resource repeatedly gave way to externally managed hierarchical state. Second, we identify five structural mismatches between long-context scaling and long-horizon agency: degraded effective context, lossy compaction, a single trust region, no transactional semantics, and no working-set decomposition. Third, we map long-horizon requirements to what context scaling provides and what harness-managed state must supply. We conclude by proposing eight harness state-management patterns and arguing that harness state should become a first-class research object for long-horizon agency.
Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities
Armaan A Abraham ⋅ Lucy Xiaoyang Shi ⋅ Chelsea Finn
Off-policy, value-based reinforcement learning methods such as Q-learning are appealing because they can learn from arbitrary experience, including data collected by older policies or other agents. In practice, however, bootstrapping makes long-horizon learning brittle: estimation errors at later states propagate backward through temporal-difference (TD) updates and can compound over time. We propose long-horizon Q-learning (LQL), which introduces a principled backstop against compounding error when learning the optimal action-value function. LQL builds on a prior optimality tightening observation: any realized action sequence lower-bounds what the optimal policy can achieve in expectation, so acting optimally earlier should not be worse than following the observed actions for several steps before switching to optimal behavior. Our contribution is to turn this inequality into a practical stabilization mechanism for Q-learning by using a hinge loss to penalize violations of these bounds. Importantly, LQL computes these penalties using network outputs already produced for the TD error, requiring no auxiliary networks and no additional forward passes relative to Q-learning. When combined with multiple state-of-the-art methods on a range of online and offline-to-online benchmarks, LQL consistently outperforms both 1-step TD and n-step TD learning at similar runtime.
Long-term Embodied Visual Tracking with Lightweight Vision-Language-Action Models
Haowei Sun ⋅ Kaining Chen ⋅ Xutao Wen ⋅ Xinze Xie ⋅ Jinwu Hu ⋅ Jiaxi Chen ⋅ Mingkui Tan ⋅ Shutao Li
Embodied Visual Tracking (EVT) aims to autonomously control an agent to follow a target in 3D space, which is critical for robotic navigation and human-robot interaction. However, robust long-term EVT in real-world scenarios faces two key bottlenecks: frequent tracking failures caused by occlusions and distractors, and the lack of a long-term benchmark for model training and evaluation. To address these, we propose a systematic solution. First, we develop LT-VLA, a lightweight VLA model featuring a Target Memory module to maintain consistent target identification and an Adaptive Execution module to adjust tracking actions based on observation reliability. Second, we construct LTEVT, a unified long-term EVT benchmark comprising 11 high-fidelity indoor and outdoor environments, enriched with over 5,000 assets and 325 human targets to simulate occlusions and distractors. Along with the benchmark, we release the LTEVT-4500K dataset for large-scale training and a comprehensive evaluation toolkit. Extensive experiments demonstrate that our 0.6B model achieves state-of-the-art performance comparable to 7B-scale models while maintaining superior efficiency. Notably, LT-VLA achieves 62.4% average Success Rate on EVT-Bench and runs at 20 FPS on a Unitree G1 robot, delivering a practical and robust solution for real-world deployment.
Algorithmic risk predictors are increasingly being used to inform resource allocation in healthcare and other high-stakes domains. Common in many such settings, the public decision maker delegates service provision to private organizations and compensates them according to each individual's predicted cost of service. We show that this common allocation mechanism can fail when organizations engage in \emph{favorable selection}, i.e., strategically choosing individuals whose actual costs are below the payment model predictions. Motivated by health insurance payments in the Medicare Advantage (MA) program, we provide a theoretical framework to study the interaction between a decision maker and a strategically selecting service provider. Our theoretical framework predicts two empirical patterns observed in Medicare Advantage that have become central concerns: enrollment in the private program expands, while government payments rise above the counterfactual cost of direct public provision. Guided by our theoretical framework, we then propose a reformed payment policy that is implementable under the government's data-access constraints. We prove that its equilibria eliminate overpayment while preserving profits for efficient organizations. We validate the framework in synthetic and semi-synthetic experiments that illustrate the advantages of our reformed policy, over the status quo.
Look-ahead Variational Flow for Generative Online Reinforcement Learning
Tianyi Zhang ⋅ Likun Wang ⋅ Tianze Zhu ⋅ Guojian Zhan ⋅ yinuo Wang ⋅ Feihong Zhang ⋅ Yang Guan ⋅ Yao Lyu ⋅ Shengbo Eben Li
Generative policies have emerged as a compelling paradigm in reinforcement learning (RL) as they can represent complex and multimodal action distributions. However, this high expressivity leaves them particularly vulnerable to the inherent target distribution non-stationarity in RL, rendering the generative process severely volatile. To address this issue, we propose \textbf{L}ook-ah\textbf{E}ad v\textbf{A}riational \textbf{F}low (LEAF), a generative policy optimization framework that stabilizes both the target and the generative process. First, we replace the action-value-guided target with a look-ahead target induced by predictive dynamics, which we theoretically prove to exhibit a tighter target drift bound. Second, we formulate policy improvement as a proximal variational free-energy optimization, which enforces a distributional trust region to prevent overfitting to transient target shifts. Finally, we prove that this optimization induces an optimal transport flow, which we instantiate as a two-phase coarse-to-fine generative process to achieve broad multimodal coverage and pinpoint accuracy. Extensive experiments on the DeepMind Control Suite and Humanoid Bench demonstrate that LEAF consistently outperforms strong baselines in both sample efficiency and asymptotic performance.
Looped Transformers with Layer Normalization Provably Learn the Power Method
Lyumin Wu ⋅ Chenyang Zhang ⋅ Yuan Cao
Transformers have achieved remarkable success across a wide range of applications, and a growing body of work suggests that part of their strength comes from their ability to learn and execute algorithmic procedures. However, our understanding of how transformers learn such algorithms remains limited, especially in the presence of layer normalization (LN). In this work, we study principal component prediction as a concrete testbed for understanding the training dynamics of transformers with LN. We prove that a looped linear transformer with LN, trained by gradient descent, converges to a solution that implements the power method, with each self-attention layer performing one power iteration. Notably, the model is trained only for principal component prediction, rather than being explicitly supervised to implement the power method. Our finding thus reveals an ``algorithmic implicit bias'' of looped transformers with LN: principal-component prediction can in principle be achieved by many mechanisms, yet gradient descent selects one that realizes the power method. We further provide a concrete comparison between transformers with and without LN: even with layerwise guidance from power iterations, a transformer without LN cannot exactly learn the power method, whereas the corresponding transformer with LN can, leading to a provable performance gap in principal component prediction. Our results provide, to our knowledge, the first theoretical analysis of the training dynamics of looped and single-layer transformers with LN, and shed light on the role of LN in transformer models.
Loop-Free Inverse Reinforcement Learning via Sequential Value Recovery
Yang Chen ⋅ Yitan Zhang ⋅ Michael Witbrock ⋅ Shuyue Hu
Inverse Reinforcement Learning (IRL) aims to recover a reward function that explains expert demonstrations. Existing IRL methods typically rely on a bi-level optimization procedure that alternates between reward learning and policy optimization, leading to substantial computational burden and training instability. In this work, we introduce a different route that eliminates policy optimization entirely by leveraging diffusion policies. Our key insight is that a diffusion policy encodes the action-gradient structure of the optimal soft Q function, enabling reward learning to be cast as a sequence of value recovery problems, thereby allowing us to bypass reward-policy loops inherent in prior IRL methods. Specifically, our method proceeds in three stages: (I) recovering the optimal soft Q function via action-gradient matching and estimating the corresponding soft value function (LogSumExp of Q values) in a way inspired by Gumbel regression; (II) calibrating these soft values by inferring a state-dependent offset; (III) extracting the reward by enforcing Bellman consistency. This leads to Loop-Free Inverse Reinforcement Learning (LFIRL), a fully offline algorithm that operates in a simple, loop-free, and sequential manner. LFIRL is simple to implement, significantly improves training efficiency while maintaining strong reward recovery performance. Empirically, across Maze, Franka Kitchen, Adroit Hand Pen, and Push-T benchmarks, LFIRL achieves 2-3x speedup over the fastest baselines, while matching or surpassing state-of-the-art methods in reward recovery quality.
LoRaQ: Optimized Low Rank Approximation for 4-bit Quantization
Yann Bouquet ⋅ Alireza Khodamoradi ⋅ Sophie Y Shen ⋅ Kristof Denolf ⋅ Mathieu Salzmann
Post-training quantization (PTQ) of large diffusion transformers degrades generative quality at 4-bit quantization. Low-rank approximation methods are a promising solution and append auxiliary linear branches to restore performance. Current state-of-the-art approaches keep these branches at high precision (W16A16) and rely on heavy, data-dependent calibration for initialization. Their branch is defined as a precise low-rank approximation of a full-rank matrix and therefore cannot keep precision at sub-16 bit levels. We challenge both limitations with LoRaQ (Low-Rank Approximated Quantization), a data-free calibration approach that optimizes quantization error compensation. By overcoming the need for high-precision branches, LoRaQ enables the first fully sub-16 bit pipeline, allowing the low-rank branch itself to be quantized. We demonstrate that, at equal memory overhead, LoRaQ outperforms the state-of-the-art methods in their native implementations on Pixart-$\Sigma$, SANA and Flux.1. We also analyze mixed-precision configurations, showing that setups such as W8A8, W6A6, and W4A8 for the low-rank branch, alongside a W4 residual layer, yield superior results while maintaining a fully quantized architecture compatible with modern mixed-precision hardware, and translate into kernel-level speedups of up to $3.19\times$ on AMD MI355.
LoReC: Rethinking Large Language Models for Graph Data Analysis
Hongyu Zhan ⋅ Qixin Wang ⋅ Yusen Tan ⋅ Hai-tao Yu ⋅ Jingbo Zhou ⋅ Shuai Chen ⋅ Jia Li ⋅ Xiao Tan ⋅ Jun Xia
The advent of Large Language Models (LLMs) has fundamentally reshaped the way we interact with graphs, giving rise to a new paradigm called GraphLLM. As revealed in recent studies, graph learning can benefit from LLMs. However, we observe limited benefits when we directly utilize LLMs to make predictions for graph-related tasks within GraphLLM paradigm, which even yields suboptimal results compared to conventional GNN-based approaches. Through in-depth analysis, we find this failure can be attributed to LLMs' limited capability for processing graph data and their tendency to overlook graph information. To address this issue, we propose LoReC (Look, Remember, and Contrast), a novel plug-and-play method for GraphLLM paradigm, which enhances LLM's understanding of graph data through three stages: (1) Look: redistributing attention to graph; (2) Remember: re-injecting graph information into the Feed-Forward Network (FFN); (3) Contrast: rectifying the vanilla logits produced in the decoding process. Extensive experiments demonstrate that LoReC brings notable improvements over current GraphLLM methods and outperforms GNN-based approaches across diverse datasets. The implementation is available at https://anonymous.4open.science/r/LoReC-63F4.
LR-V2X: Loss Resilient Collaborative Perception under Low-Bandwidth Communication
Kang Yang ⋅ Tianci Bu ⋅ Peng Wang ⋅ Deying Li ⋅ Yongcai Wang
Given the inherent unpredictability of packet loss in vehicular wireless communications, V2X collaborative perception can yield practical benefits only if agents can achieve reliable collaboration under lossy and low-bandwidth communication conditions. Existing dense BEV feature fusion methods depend on redundant BEV feature exchange, which is infeasible in low-bandwidth scenarios, while compact-communication methods aggressively compress messages but can hardly recover the missing feature content after packet loss. In this paper, we present LR-V2X, a loss-resilient, latent-space reconstruction framework that converts corrupted received latents (even under severe 90\% packet loss) into a spatial prior and then reconstructs the missing BEV information from this informative prior and using ego context as condition. Notably, the model can be trained under complete communication conditions and can be directly applied to lossy conditions at test time, eliminating the need for training under numerous lossy conditions. Experiments on DAIR-V2X and V2XREAL show that LR-V2X delivers the strongest robustness under severe packet loss and preserves reliable collaboration as communication quality degrades. And it reduces communication overhead by $64\times$ compared to dense BEV feature fusion baselines. Code will be publicly released.
Making LoRA Identifiable: Orthogonal Alignment for Continual Learning
Pratik Rakesh Singh ⋅ Mohammadi Zaki ⋅ Akash Saha ⋅ Aneesh Mukkamala ⋅ Pankaj Wasnik
There is growing interest in applying Parameter-Efficient Fine-Tuning (PEFT) to Continual Learning, as it enables faster training, mitigates catastrophic forgetting, and facilitates better adaptation to new tasks. Prior approaches have explored a range of strategies, including extending new LoRA branches for each incoming task, training under orthogonal loss objectives across all task-specific LoRA modules, and employing gating mechanisms to integrate new and existing LoRA modules. However, these methods incur memory costs that grow linearly with the number of tasks, limiting their scalability. More recent approaches investigate training a single LoRA by leveraging asymmetries in its parameterization; however, they often suffer from insufficient representational alignment across the LoRA parameter space. To address these limitations, we propose Stiefel Optimization and Aligned Rotation (SOAR), a two-stage framework for continual LoRA learning. First, we constrain the $\mathbf{A}$ matrix to lie on the Stiefel manifold, thereby reducing coordinate misalignment to the space of orthogonal rotation matrices. Second, we introduce an Orthogonal Procrustes alignment procedure for the $\mathbf{B}$ matrix to estimate the optimal rotation, coupled with a per-rank retention merging strategy that quantifies directional agreement between representations of new and previously learned tasks. We conduct extensive experiments across multiple models and benchmarks, demonstrating that SOAR achieves strong continual learning performance without incurring additional memory overhead.
Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
Junsoo Oh ⋅ Wei Huang ⋅ Taiji Suzuki
Mamba, a recently proposed linear-time sequence model, has attracted significant attention for its computational efficiency and strong empirical performance. However, a rigorous theoretical understanding of its underlying mechanisms remains limited. In this work, we provide a theoretical analysis of Mamba's in-context learning (ICL) capability by focusing on tasks defined by low-dimensional nonlinear target functions. Specifically, we study in-context learning of a single-index model $y \approx g_*(\langle \boldsymbol{\beta}, \boldsymbol{x} \rangle)$, which depends on only a single relevant direction $\boldsymbol{\beta}$, referred to as *feature*. We prove that Mamba, pretrained by gradient-based methods, can achieve efficient ICL via *test-time feature learning*, extracting the relevant direction directly from context examples. Consequently, we establish a test-time sample complexity that improves upon linear Transformers---analyzed to behave like kernel methods---and is comparable to nonlinear Transformers, which have been shown to surpass the Correlational Statistical Query (CSQ) lower bound and achieve near information-theoretically optimal rate in previous works. Our analysis reveals the crucial role of the *nonlinear gating* mechanism in Mamba for feature extraction, highlighting it as the fundamental driver behind Mamba’s ability to achieve both computational efficiency and high performance.
ManipulationRAG: Retrieval-Augmented Fine-Grained Manipulation of Object Functional Parts
Xiaoyun Chang ⋅ ZiJun Zhang ⋅ Jie Chen ⋅ Xiaopeng Wen ⋅ Xiangbo Lin ⋅ Xiaohong Ma ⋅ yi sun
Modeling fine-grained human manipulation of object functional parts is essential for achieving human-like dexterous manipulation in virtual reality and robotics. Despite recent advances, data-driven generation methods remain largely constrained by patterns learned from limited training data, rendering them prone to hallucinating task-inconsistent manipulations when encountering samples beyond the training distribution. To address this challenge, we propose ManipulationRAG, a framework that augments an internal generative model with externally retrieved, task-relevant manipulation knowledge. Specifically, we design a hand-centric manipulation taxonomy and construct a structured knowledge base that organizes task descriptions, objects, manipulation types, and corresponding hand pose sequences, supported by a retrieval mechanism for accurate knowledge retrieval. The retrieved knowledge serves as an external manipulation-specific prior, guiding the diffusion model toward task-consistent and physically plausible manipulation synthesis. Extensive experiments demonstrate that our method outperforms existing methods and generalizes robustly to object manipulation beyond the original diffusion model's training domain.
MARS: Enabling Autoregressive Models Multi-Token Generation
Ziqi Jin ⋅ Lei Wang ⋅ Ziwei Luo ⋅ Aixin Sun
Autoregressive (AR) language models generate text one token at a time, even when consecutive tokens are highly predictable given earlier context. We introduce MARS (Mask AutoRegreSsion), a lightweight fine-tuning method that teaches an instruction-tuned AR model to predict multiple tokens per forward pass. MARS adds no architectural modifications, no extra parameters, and produces a single model that can still be called exactly like the original AR model with no performance degradation. Unlike speculative decoding, which maintains a separate draft model alongside the target, or multi-head approaches such as Medusa, which attach additional prediction heads, MARS requires only continued training on existing instruction data. When generating one token per forward pass, MARS matches or exceeds the AR baseline on six standard benchmarks. When allowed to accept multiple tokens per step, it maintains baseline-level accuracy while achieving 1.5-1.7x throughput. We further develop a block-level KV caching strategy for batch inference, achieving up to 1.71x wall-clock speedup over AR with KV cache on Qwen2.5-7B. Finally, MARS supports real-time speed adjustment via confidence thresholding: under high request load, the serving system can increase throughput on the fly without swapping models or restarting, providing a practical latency-quality knob for deployment.
MaskSense: Confronting the Visual Exploration Trap in Masked Image Generation
Yawen Shao ⋅ Jie Xiao ⋅ Kai Zhu ⋅ Yu Liu ⋅ Hongchen Luo ⋅ Xueyang Fu ⋅ Yang Cao ⋅ Wei Zhai ⋅ Zheng-Jun Zha
Reinforcement learning (RL) has shown strong potential for aligning generative models with human intent, but adapting RL to masked generative models (MGMs) remains largely underexplored. Existing approaches that directly adapt Group Relative Policy Optimization (GRPO) to MGMs leave a foundational conflict unaddressed: MGMs' exploitative decoding nature is inherently at odds with the exploratory diversity that GRPO presupposes. In this work, we reveal a structural decoding bias termed the visual exploration trap: confidence-based sampling trades the exploration of diverse subject realizations for the greedy resolution of low-uncertainty background regions, causing premature collapse of the generative exploration space and starving GRPO of the rollout diversity it depends on. To this end, we propose MaskSense, a novel RL framework for MGMs that confronts the visual exploration trap at both the sampling and optimization stages. Specifically, we introduce a semantic anchored routing sampling mechanism that leverages semantic priors to preserve high-entropy subject token exploration while amplifying intra-group reward variance through exploratory-exploitative routing to yield more discriminative advantage estimates. Furthermore, we design a global entropy transition anchoring strategy to identify the most consequential decoding steps and concentrate gradient updates on them. Extensive experiments on multiple text-to-image benchmarks demonstrate that MaskSense substantially improves upon the base model, with GenEval accuracy improving from 54% to 86% and HPS from 28.89 to 36.39, achieving state-of-the-art performance.
Matched-Control Tests of Partition-Source Claims in One Routed Distillation Family
Bob Li ⋅ Chloe Cao ⋅ Constance Liu
Partitioned routed distillation is hard to interpret when route count, supervision grouping, and search or filtering budget change together. We study that attribution problem inside one fixed routed family using two matched nulls: a routing-only control (Latent-Routing MoA) and a structure-matched shuffled-partition control (Matched-Partition-Shuffle), with teacher tokens, trainable parameters, and router depth matched exactly. The decisive result is Table tab:main: the zero-manual deployable row ModeDistill-Auto-v4-Open, whose partition source is induced by GPT-OSS-120B-Instruct pairwise same-strategy judgments (Apache- 2.0 , distinct from the trace-generating teacher), beats MoA by +2.46 pp and Matched-Partition-Shuffle by +2.04 pp on the five-domain primary-shift average (A1), with positive OOD gaps on all five primary domains. Table tab:externalbreadthtransfer is intentionally weaker evidence: for that exact same headline row it reports only four external benchmark aggregate means, where Auto-v4-Open beats MoA by +4.91 / +3.99 / +3.70 / +2.87 pp, and it does not support slice-wise breadth or cross-family generality. The open pipeline remains within 0.05 pp of the manual upper-bound reference on A1 and within the pre-specified 0.50 pp tie band on every external mean, so we keep it as the deployable row and retain the manual row only as calibration. The claim is therefore narrow: within this routed family, once routed budget is matched, the residual OOD gain is attributable to the externally induced partition source rather than to routing alone or to structure-matched shuffled partitions. Figure fig:identification and the appendix are corroborative only. The weakest primary shift remains ALFWorld tool reconfiguration ( +1.15 pp vs.\ MoA, paired CI [-0.12,+2.41] ), and discovery-time verifier calls and teacher generation are disclosed separately from the matched comparison.
Multi-agent large language model (LLM) debate is often evaluated by whether final answers improve, but movement is not necessarily improvement: the same discussion can rescue an initially wrong majority or destroy an initially correct one. Standard final-accuracy evaluations conflate these opposing mechanisms. We introduce an auditable protocol for homogeneous debate on multiple-choice questions (MCQs) that records each run as a transition ledger over collapse, correction, onset, and signed intervention utility. On $6{,}925$ MMLU-Pro debates, the protocol identifies $253$ collapses and a parallel correction ledger that changes how interventions should be judged. Replay experiments reveal the central tradeoff: a leave-one-model-out probe-gated freeze prevents $29$ collapses but loses $108$ corrections under equal weights, so collapse prevention alone can recommend the wrong policy. A compact pre-debate $8$-probe screen is a triage signal: its unadjusted family-level association with conditional-collapse risk is high ($G{=}7$, Spearman $\rho=0.893$, exact two-sided $p=0.0123$), but initial-majority accuracy is a close comparator ($\rho=0.821$; family partial $\rho=0.767$, $p=0.0877$), so we do not treat it as calibrated or capability-adjusted prediction. Round-level traces localize many collapses to the first debate round, where early disagreement can precede both harmful cascades and useful recovery. We release replayable schemas, coders, audits, cost cards, and zero-API rebuild scripts so future model--scaffold rows can be compared under the same denominators and signed utility ledger.
Modern large language models keep the feed-forward network (FFN) intermediate width identical across every layer, a convention inherited from the original Transformer. Yet mechanistic work suggests layers at different depths do qualitatively different work, raising a natural question: are their capacity demands also non-uniform, and if so, can a fixed parameter budget be allocated more wisely than uniformly? To probe this, we apply principal component analysis (PCA) to the input of each FFN down-projection and measure, for every layer, the smallest number of principal components needed to reconstruct the activation up to a fixed cosine similarity --- the layer's intrinsic dimensionality. Scanning 58 open-weight LLMs spanning 14 families and 70M to 72B parameters, we find a recurring middle-heavy profile in which middle layers carry higher intrinsic dimensionality than edge layers; the profile is robust to the calibration corpus, even across English and Chinese, and is largely preserved through post-training fine-tuning. As a direct test that this signal is structurally meaningful, we use it to guide width pruning in a knowledge-distilled student under a matched parameter budget: the PCA-guided allocation outperforms the uniform-width convention by $+2.7$ points on Qwen2.5-3B and $+2.3$ on Qwen3-8B, while a same-scale teacher (LLaMA-3-8B) shows only $+0.9$, and a budget-matched inverted control consistently does worst. Across these three teachers, the size of the gap scales with how unevenly the teacher's intrinsic dimensionality is distributed across layers --- a single-number summary we call $\sigma_{99}$ --- and the advantage persists through downstream supervised fine-tuning.
Mechanistic Circuit Identification for Controllable Data Generation
Nakyung Lee ⋅ Sangwoo Hong ⋅ Jungwoo Lee
While recent advances in data synthesis aim to curate high-quality datasets, most generation pipelines still rely on heuristic prompt-based control. This black-box paradigm provides limited insight into how individual samples interact with a model's underlying learning dynamics. To bridge this gap, we propose a circuit-grounded framework that connects training-dynamics-based data valuation with mechanistic interpretability (MI). Specifically, we conceptualize data quality along three complementary utility axes, learnability, challenge, and alignment. First, we uncover specialized model-internal circuits that causally govern these utility signals. Then, moving beyond heuristic prompting toward mechanistic control, we leverage these circuits as controllable interfaces, actively steering generation to produce utility-targeted data. Building on this capability, we introduce SAMS (Stage-Aware Mechanistic Scheduling), which schedules circuit-steered data according to the model's evolving optimization needs. Experiments on multiple-choice QA tasks demonstrate that our approach yields precisely controlled data with greater diversity than prompt-based baselines, consistently improving downstream performance and calibration. Ultimately, this work establishes a principled white-box paradigm for interpretable data generation, pioneering the use of MI not just as an analytical tool, but as a practical, controllable interface.
MechParser: A Vision-Language Framework for Parsing Chemical Reaction Mechanism Diagrams
Yufan Chen ⋅ LU S Hin ⋅ Yujian Yuan ⋅ Ching Ting Leung ⋅ Hanyu Gao
Reaction mechanisms are central to understanding chemical reactivity, guiding reaction prediction, synthesis design, and selectivity rationalization. In chemistry literature, they are universally depicted as curved electron-pushing arrows that capture stepwise electron flow and bond changes. Yet, this vast knowledge remains trapped in unstructured diagrams, forcing data-driven modeling to rely on small or template-expanded datasets. We introduce MechParser, a vision-language framework that converts mechanism diagrams into SM-SMARTS (Sequential Mechanism-Aware SMARTS), a representation encoding electron flow and atom state changes as a cumulative string sequence. At its core, MechParser-VL applies a left-right visual prompt design (arrow diagram alongside an atom-indexed reference) to separate dynamic arrow reasoning from static atom identification, and is trained via a four-stage curriculum on over 300K samples from a geometry-aware synthetic data pipeline and an expert-annotated real-world dataset. On both synthetic and real-world benchmarks, MechParser-VL substantially outperforms much larger proprietary and open-source VLMs despite using only a 4B backbone. By standardizing visual mechanisms into machine-readable data, MechParser lays the groundwork for constructing large-scale mechanism databases from the chemical literature, which in turn can support the training of advanced mechanism-centric models on authentic mechanistic data.
Med-Agentic: Distilling Agentic Medical Reasoning with Internalized Meta-Capabilities
Yucheng Zhou ⋅ Junwei Sheng ⋅ Jianbing Shen
Large language models (LLMs) in expert domains face a trade-off between \emph{agentic tool use}, which externalizes expertise behind brittle multi-turn runtimes, and \emph{end-to-end specialization}, which internalizes expertise only implicitly at task-level granularity. We propose \emph{internalized meta-capabilities}: atomic, parametrically internalized, and composable reasoning units that an LLM can select and combine within a single chain-of-thought (CoT). We instantiate this idea as \textbf{Med-Agentic} for medical visual diagnosis, where the model emits a capability-tagged structured CoT in one forward pass, turning capability invocation into tagged token generation rather than external tool calls. For training, \textbf{Med-OPD} performs multi-teacher on-policy distillation with \emph{token-level dual-factor teacher routing} conditioned on both the image and the active capability tag, strictly generalizing per-prompt teacher routing. We show that this routing gives an unbiased estimator of the same expected teacher-KL gradient. Across public medical imaging datasets spanning nine anatomies and eight modalities, Med-Agentic consistently improves in-distribution and out-of-distribution diagnostic accuracy over the strong competitors.
MedKIT: Evaluating Knowledge Integration and Generalization in Large Language Models
Lukas Thede ⋅ Yash Kumar ⋅ David Chen ⋅ Danielle Bitterman ⋅ Matthias Bethge ⋅ Tom Hartvigsen ⋅ Zeynep Akata
Constantly evolving real-world knowledge necessitates models to be updated continuously. Especially in medicine, as clinical evidence changes over time, outdated knowledge can pose safety risks. Existing evaluations of knowledge integration focus on factual recall, offering limited insight into whether newly integrated knowledge is actually usable. Our benchmark MedKIT (Medical Knowledge Integration and Transfer) provides a granular evaluation of how models integrate and apply knowledge under realistic sequences of clinical updates. Each instance corresponds to a factual update derived from clinical evidence, paired with targeted probes that assess transfer across lexical variation, relational transformations, compositional reasoning, and open-ended operationalization, as well as locality tests for knowledge preservation. Using MedKIT, we conduct a large-scale empirical study of 12 knowledge integration strategies across 5 diverse models, including both general-purpose and medical LLMs. Our results reveal a consistent gap between recall and usable knowledge: while most methods achieve strong gains on the original update task and under lexical variation, relational generalization is limited, and no method yields meaningful improvements on compositional or operational tasks. These findings highlight a fundamental challenge in knowledge integration and position MedKIT as a testbed for developing methods that make newly integrated knowledge more consistently usable across tasks and contexts.
MemCoRe: Recovering Evidence from Progressively Compressed Factual Knowledge for Agent Memory
Zhenyuan Zhang ⋅ Jia Xianzhang ⋅ Zhiqin Yang ⋅ Zhenbo Song ⋅ Wei Xue ⋅ Sirui Han ⋅ Yike Guo
Memory systems enable LLM agents to consolidate and retrieve relevant evidence from the factual knowledge accumulated through growing interaction histories for downstream reasoning. Existing approaches have explored diverse strategies for organizing and compressing these histories. However, balancing compression with retrieval effectiveness remains challenging: retaining too much content can cause relevant evidence to be obscured by redundant entries, while discarding too aggressively may remove content that later proves relevant. This amounts to a tradeoff between compressing redundancy and preserving enough structure to retrieve target evidence, as formalized by the information bottleneck. To this end, we propose MemCoRe, which organizes memory as a compression hierarchy where each level compresses redundancy further while retaining the structure needed for retrieval at that level. In this hierarchy, evidence is progressively compressed from detailed records through extracted keywords to topic groups. This enables retrieval to locate target evidence by searching across levels of the hierarchy. Comprehensive experiments demonstrate that MemCoRe outperforms existing state-of-the-art baselines.
Memorization Is Folding: Topological Signatures of Noisy-Label Learning
Zhongtian Sun ⋅ Fan Mo ⋅ Prayag Tiwari ⋅ KELIN XIA
Neural networks can fit noisy labels, but how their representations change as memorization begins remains unclear. We study this using persistent homology of penultimate-layer representations across training. Across 8 datasets and 8 architectures, we find a consistent pattern: the discriminative topological signal under noise appears in $H_1$ rather than $H_0$, with noisy representations preserving more loop structure than clean representations as training proceeds, supporting a folding view of memorization. When the clean topological signal decreases over training, the noisy signal does not decline as quickly as the clean one, producing an early crossover between clean and noisy trajectories; across the derivation set, crossover occurs when clean trajectories fade and the clean--noisy gap closes, whereas no-crossover cases remain persistent or do not close the gap within training. The same framework also correctly predicts the held-out ResNet-34/CIFAR-10 case, confirmed in all 9 runs. The signal reveals aspects of training dynamics that standard statistics do not represent directly. Activation variance is stronger on average, but the topological measure becomes informative earlier than validation loss and remains significant after controlling for intrinsic dimension. It also distinguishes irreducible label contradictions, including instance-dependent and CIFAR-10N human noise, from learnable confusions such as asymmetric noise. The same measure also falls during grokking in 10 of 12 modular-arithmetic runs and, after per-pipeline calibration from a clean reference plus runs at known noise rates, predicts population-level noise severity without per-sample labels, achieving per-pipeline $R^2 \geq 0.94$ across six tested configurations. These findings position topology as both an explanatory lens on noisy-label memorization and a practical early-stage diagnostic for studying, auditing and improving learning under label noise.
Memory Determines Learning Direction: A Theory of Gradient-Based Optimization in State Space Models
JingChuan Guan ⋅ Tomoyuki Kubota ⋅ Yasuo Kuniyoshi ⋅ Kohei Nakajima
State space models have shown strong potential to outperform Transformers, but their learning mechanisms remain poorly understood, particularly the conditions for successful parameter updates that preserve teacher information. Prior work attributes the difficulty of such updates to the vanishing gradient during backpropagation through time, in which error signals decay exponentially through repeated gradient multiplication. In this work, by expressing error signals explicitly as functions of inputs, we show that the conventional vanishing gradient framework is insufficient to explain the loss of information. Using this representation of error, we reveal that the memory of the input encoded in teacher information does not decay exponentially, and we analytically demonstrate that the loss of supervisory information cannot be alleviated through training. This result theoretically confirms the importance of the initial weights of SSMs and suggests that the recurrent layer does not require training. By performing experiments on both linear SSMs and a representative nonlinear SSM, including language modeling tasks, we confirm that appropriate initialization enables us to fix the recurrence, leading to more stable training of the entire network and higher performance than the conventional training scheme. Our results partially elucidate the learning mechanisms of SSMs and are expected to contribute to the development of improved RNN-based models.
MemoryFusion: Cross-Temporal Memory Learning for Multimodal Video Fusion
Gong Meiqi ⋅ Hao Zhang ⋅ Jiayi Ma
Multimodal video fusion methods still suffer from limited performance due to the inadequate exploitation of historical frames and the simplified temporal consistency modeling. To address this, we propose MemoryFusion, a two-stage framework built upon cross-temporal memory learning, which leverages historical information across multiple temporal scales to enhance robustness and progressively refine temporal consistency. In the first stage, we introduce a Short-Term Memory Bank (STMB) to aggregate recent temporal features for enhancing the current frame, and a Long-Term Memory Bank (LTMB) to preserve representative features from extended temporal sequences through adaptive update rules. Meanwhile, we incorporate a Temporal Smoothing Module (TSM) to perform coarse temporal consistency modeling, suppressing abrupt background variations and establishing a stable temporal foundation. In the second stage, we design a Temporal Refinement Module (TRM) implemented as a lightweight 3D residual network, which conducts prior-guided fine-grained temporal refinement to recover subtle temporal dynamics and spatial details overlooked in the first stage. Extensive experiments demonstrate that MemoryFusion significantly improves both visual fidelity and temporal consistency by performing cross-temporal learning, outperforming the state-of-the-art in both frame-based and video-based fusion methods. The source code and pretrained models will be publicly released.
Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents
Sikuan Yan ⋅ Ahmed Bahloul ⋅ Ercong Nie ⋅ Susanna Schwarzmann ⋅ Riccardo Trivisonno ⋅ Volker Tresp ⋅ Yunpu Ma
Memory-augmented LLM agents enable interactions that extend beyond finite context windows by storing, updating, and reusing information across sessions. However, training such agents with reinforcement learning in multi-session environments is challenging because memory turns the agent's past actions into part of its future environment. Once different rollouts write, update, or delete different memories, they no longer share the same intermediate memory state, making trajectory-level comparisons fundamentally unfair. This violates a key assumption behind group-relative methods such as GRPO, where rollouts are compared as if they were sampled from the same effective environment. Consequently, trajectory-level rewards provide noisy or biased credit signals for long-horizon memory operations. To address this challenge, we introduce Memory-R2, a training framework for long-horizon memory-augmented LLM agents. Its core algorithm, LoGo-GRPO, combines local and global group-relative optimization. The global objective preserves end-to-end learning from long-horizon trajectory-level rewards, while local rerollouts compare alternative continuations from identical intermediate memory states, yielding fairer group comparisons and more precise supervision for memory writing, updating, and retrieval. Beyond credit assignment, Memory-R2 jointly optimizes memory formation and memory evolution with a shared-parameter co-learning design, where a fact extractor and a memory manager are instantiated from the same LLM backbone through role-specific prompts. To make memory construction more controllable, we further formulate each session as a multi-step decision process by chunking the interaction and alternating the two memory roles. A progressive curriculum from 8 to 16 to 32 sessions stabilizes training as the memory horizon grows. Together, these components provide an effective training paradigm for memory-intensive LLM agents in long-horizon multi-session settings.
MemPilot: Learning Transferable Latent Memory Mechanisms for LLM Reasoning
Changlong Shi ⋅ LINHAO LUO ⋅ Shigeng Chen ⋅ Guibin Zhang ⋅ Yi Chang ⋅ Shirui Pan ⋅ Chengqi Zhang
Integrating external memory into Large Language Models (LLMs) typically faces a trade-off between flexibility and depth. Explicit text-based retrieval keeps memory editable, but often acts as a static prompt prefix, increasing context length and interacting with the reasoning process only indirectly. Conversely, parametric memory can influence internal computation more deeply, but updating memory content usually requires additional optimization and may couple domain-specific information with the model's general reasoning behavior. In this work, we introduce MemPilot, a decoupled latent memory framework that dynamically guides the hidden reasoning states of frozen LLMs. MemPilot constructs an external bank of compact latent representations from offline reasoning trajectories, retrieves candidate memory entries, and integrates their latent memory tokens into hidden states via cross-attention and gated residual fusion. By separating domain-specific memory content from the learned memory-use mechanism, MemPilot enables target-domain adaptation through offline memory-bank substitution without updating the base model. Experiments across QA, coding, and mathematical reasoning benchmarks show that MemPilot achieves strong in-domain performance and more robust cross-domain transfer than other memory-augmented baselines. These results suggest that LLMs can benefit from reusable latent memory-use mechanisms while keeping memory content external and replaceable.
MENDR: Manifold-Embedded Neural Data Representations for Channel-Agnostic EEG Foundation Modeling
Micky C Nnamdi ⋅ Matthew Chen ⋅ Benoit Marteau ⋅ Shaun Q. Y. Tan ⋅ J. Ben Tamo ⋅ May Dongmei Wang
Foundation models for electroencephalography (EEG) have shown promise in learning transferable representations, but existing approaches treat EEG as a generic time series, ignoring the Riemannian geometry of spatial covariance structure that is fundamental to neural signal analysis. We propose MENDR (Manifold-Embedded Neural Data Representations), the first EEG foundation model that builds this geometry directly into its architecture (via the matrix logarithm and Log-Euclidean tangent space) rather than learning it from data. MENDR embeds windowed covariance matrices into the Log-Euclidean tangent space via differentiable matrix logarithm, where standard transformer operations become geometrically equivalent to Riemannian operations. A channel-agnostic spatial projection based on Perceiver-style cross-attention enables seamless transfer across electrode configurations ranging from 6 to 64 channels. Pretrained on the Temple University Hospital EEG Corpus (4,000+ hours, 14,987 subjects) using a masked autoencoder objective in the tangent space, MENDR achieves state-of-the-art results on 3 of 6 downstream benchmarks (TUAB abnormality, TUEV event, CHB-MIT seizure) with only 1.2M parameters, up to $42\times$ fewer than competing foundation models.
Merging RLVR-Trained Experts via Policy-Shift-Guided Spectral Alignment
Geeho Kim ⋅ MINSIK CHOI ⋅ Kyle Min ⋅ Young Geun Kim ⋅ Bohyung Han
We propose Policy-Shift-Guided Spectral Alignment (PSA), a retraining-free method for merging RLVR post-trained language-model experts by guiding spectral subspace selection with token-level expert-base policy shifts. We first show that existing SFT-oriented merging techniques under-preserve the sparse, directional changes in next-token probabilities that distinguish RLVR experts from the pretrained base. To address this challenge, PSA scores calibration tokens by expert--base probability difference under the same prefix, and converts probability-shift-weighted input activations into column weights. These weights define a weighted SVD objective that prioritizes update directions active on high-shift tokens. After low-rank truncation, PSA applies cross-expert polar alignment to the truncated task-vector bases and restores each expert's per-layer Frobenius scale before aggregation. Experiments on Qwen2.5-7B and Qwen3-1.7B across three RLVR tasks show that PSA consistently outperforms strong merging baselines while preserving expert-level capabilities in a unified model.
MESSENGER: Memory-Enhanced Sequential Scene Flow Estimation via Autoregressive Next-Frame Forecasting
Jiuming Liu ⋅ Jianing Li ⋅ Mengmeng Liu ⋅ Hongyang He ⋅ Hesheng Wang ⋅ Per O Kristensson
Scene flow can capture low-level 3D motion displacements in dynamic scenarios. Early pairwise estimators relying on instantaneous two-frame motion lack long-term temporal correlation and also struggle with poor extrapolation ability in future prediction. Although some recent methods attempt to explore multi-frame scene flow estimation in a sequence-to-sequence manner, they typically suffer from heavy computational overhead with increasing input frames and long-horizon prediction degradation due to ineffective motion propagation. To address these problems, we propose a novel memory-enhanced sequential scene flow pipeline, called MESSENGER. To sufficiently mine long-term temporal dependencies naturally within consecutive sequences, a memory buffer is designed by explicitly storing multiple history flow estimates and latent states. For each input frame, the temporally stored flows and states are correlated and retrieved to predict the current initialized flow in a next-frame forecasting manner. Furthermore, we develop an uncertainty-aware reweighting module to filter unreliable retrievals and mitigate accumulated errors. Extensive experiments on nuscenes and Argoverse 2 demonstrate state-of-the-art performance of our MESSENGER, reducing EPE3D by 71.6% on nuScenes and 67.7% on Argoverse 2 in long-horizon future extrapolation. This superiority can be attributed to our designed autoregressive forecasting paradigm, which naturally forces the network to progressively learn the next-frame distribution based on history observations. Code will be released upon publication.
METAFORGET: Audit-Driven Update-Policy Learning for Reliable Language Model Unlearning
Pinlong Zhao ⋅ Xiaoling Zhou ⋅ Zhou Zhaoting ⋅ Guangyuan Dong
Deployed language models are increasingly expected to remove private, copyrighted, or hazardous content without full retraining. Existing unlearning methods remain brittle because they usually optimize a fixed forgetting objective with a global update schedule: weak updates leave paraphrased knowledge intact, while stronger updates can damage retained capabilities or nearby facts. We identify this failure mode as update homogeneity. To address it, we introduce METAFORGET, a meta-optimization framework that treats unlearning as learning an audit-driven update policy rather than designing another standalone forgetting loss. A lightweight controller maps memorization, paraphrase-sensitivity, and signed forget-retain conflict signals to span weights, layer gates, step sizes, and retain-gradient projection strengths. The controller is trained with a bilevel objective that differentiates through short LoRA unlearning trajectories and evaluates the resulting model on held-out forget, retain, locality, and leakage audits. This policy can instantiate NPO- or RMU-style losses while changing where and how strongly they act. Across TOFU, WMDP, and MUSE-style sequential deletion settings, METAFORGET improves the forget-retain-robustness frontier over strong baselines, moves loss-based membership-inference diagnostics closer to chance, and better preserves utility under repeated requests. The results support a view of practical LLM unlearning as audit-driven policy learning: the central object is not only the forget loss, but the localized update that must pass the deletion audit.
MetaLoop: Benchmarking the Full Metacognitive Loop in LLMs
Nora Petrova ⋅ John Burden ⋅ Jerome Wynne
We introduce MetaLoop, a seven-task benchmark that evaluates the full metacognitive loop in large language models: monitoring one's own uncertainty, translating that signal into action (abstaining, switching strategies, correcting errors), and updating from feedback. Existing evaluations test these components in isolation; MetaLoop produces integrated per-model profiles across all three. Evaluating 12 frontier models alongside 460 human participants, we find that (1) accuracy does not predict metacognitive ability (high-accuracy models routinely approve their own errors); (2) models universally produce calibration signals but fail to act on them; (3) forced self-explanation distinguishes genuine self-monitoring from brittle pattern-matching; and (4) humans and LLMs both show monitoring–control gaps but with different failure modes (humans are loss-averse, models overconfident), with clarification detection as the one clear human advantage. We release the benchmark, scoring code, model outputs, and human data.
Metaphor Is Not All Attention Needs
Olga Sorokoletova ⋅ Francesco Giarrusso ⋅ Giacomo De Luca ⋅ Piercosma Bisconti ⋅ Matteo Prandi ⋅ Federico Pierucci ⋅ Marcello Galisai ⋅ Vincenzo Suriani ⋅ Daniele Nardi
Large language models are increasingly deployed in safety-critical and user-facing applications, where their ability to resist harmful instructions is essential. Although post-training aims to make models robust against many jailbreak strategies, recent evidence shows that stylistic reformulations, such as poetic transformation, can still bypass safety mechanisms with alarming effectiveness. This raises a central question: why do literary jailbreaks succeed? In this work, we investigate whether their effectiveness depends on specific poetic devices, on a failure to recognize literary formatting, or on deeper changes in how models process stylistically irregular prompts. We address this problem through an interpretability analysis of attention patterns. Our analysis proceeds in three steps: we perform input-level ablation studies to assess the contribution of individual and combinations of rhetorical devices; we construct a novel interpretable vector representation of attention maps; we cluster these representations and train linear probes to predict both safety outcomes and literary format. Our results show that models distinguish poetic from prose formats with high accuracy, yet struggle to predict jailbreak success within each format. Clustering further reveals clear separation by literary format, but not by safety label. These findings indicate that jailbreak success is not caused by a failure to recognize poetic formatting; rather, poetic prompts induce distinct processing patterns that remain largely independent of harmful-content detection. Overall, literary jailbreaks appear to misalign large language models not through any single poetic device, but through accumulated stylistic and structural irregularities that alter prompt processing and avoid lexical triggers considered during post-training. This suggests that robustness requires safety mechanisms that account for style-induced shifts in model behavior. We use Qwen3-14B as a representative open-weight case study for all reported experiments.
Metric Depth Estimation from Arbitrarily Degraded Low-Resolution Depth Prompts
Kun Wang ⋅ Yun Zhu ⋅ Pan Zhou ⋅ Na Zhao
We propose AdaDS, a generalizable framework for prompted metric depth estimation, which estimates high-resolution metric depth from images and arbitrarily degraded low-resolution depth prompts. This setting is commonly studied as depth super-resolution, where existing methods typically regress depth values directly and often exhibit artifacts under severe or unknown depth degradation. In contrast, AdaDS exploits the contraction property of Gaussian smoothing: as noise accumulates in the forward diffusion process, the distributional discrepancy between degraded depth prompts and their high-quality counterparts progressively diminishes, eventually approaching an isotropic Gaussian prior. Leveraging this property, AdaDS estimates refinement uncertainty to adaptively select a starting timestep in the reverse diffusion trajectory, and subsequently injects tailored noise to place the intermediate sample in a high-probability region of the target posterior distribution. This strategy enables the generative prior of a pre-trained diffusion model to dominate the estimation process even when upstream prompt refinements are imperfect. Extensive experiments on real-world and synthetic benchmarks demonstrate AdaDS's superior zero-shot generalization and robustness to diverse degradation patterns compared with state-of-the-art methods.
MHWA: Multi-timescale Hierarchical World-Action Model
Pengcheng Pan ⋅ Guoqing Ma ⋅ Yuhan Zhang ⋅ Yang Chen ⋅ Yichen Liu ⋅ Ziheng Li ⋅ Shan Yu
Autoregressive world models have emerged as a powerful paradigm for sample-efficient decision-making, yet their open-loop rollouts often exhibit increasing drift that leads to error accumulation even within a relatively small planning horizon. Hierarchical temporal abstraction is a natural candidate to mitigate such drift by reducing the effective prediction horizon; however, existing hierarchical designs can introduce added structural and conceptual complexity, sometimes making them difficult to control and limiting gains over flat baselines. To this end, we introduce MHWA (Multi-timescale Hierarchical World Action model), which employs a strided high-level context channel to reduce global-context update frequency, effectively curbing the propagation of compounding drift. Crucially, via context-aware routing, the gating network dynamically reconfigures a Low-Level Mixture-of-Experts by using both global context and current observations, thereby enabling flexible high--low coordination, helping alleviate interference commonly observed in static hierarchies. Empirically, MHWA (77M parameters) achieves an IQM human normalized score (HNS) of 0.867$\pm$0.047 across three training seeds on a 14-game Atari subset using 10\% subsampled offline data, matching the existing state-of-the-art 150M-parameter JOWA baseline while using nearly half the parameters and significantly reducing computational costs.
MilliVid: Adaptive Latents for Long-Range Consistency in Video Generation
Ishaan Chandratreya ⋅ David Charatan ⋅ Basile Van Hoorick ⋅ Sergey Zakharov ⋅ Vitor Guizilini ⋅ Phillip Isola ⋅ Vincent Sitzmann
Transformer-based video generative models have become increasingly realistic, but long-horizon consistency remains challenging to achieve because even a few dozen frames create impractically long sequence lengths. We show that this issue can be mitigated by generating video using coarse-to-fine rollout within a multi-scale token space. Our approach is simple: first, we pre-train an adaptive autoencoder that compresses each frame into a hierarchy of tokens, with levels ranging from the typical latent resolution to only a handful of tokens per frame. This yields a hierarchy in which the coarsest levels capture the most consequential information—such as scene layout and semantics—while finer levels add high-frequency appearance and texture. Then, we train a video diffusion model to generate these tokens using coarse-to-fine rollout. By carefully controlling the level of detail at which frames are generated and used as context during each rollout step, we are able to preserve long-range consistency in geometry and object permanence while expending less compute on long-term consistency of less perceptually relevant details. We validate this approach using a custom long-horizon Minecraft dataset, where it produces substantially more consistent rollouts compared to strong baselines. We hope this perspective opens new opportunities for long-term consistent video generation.
Mind the Gap: Dataset and Fine-grained Evaluation for Inline Audio Descriptions
Subhashini Venugopalan ⋅ Yingwen Tan ⋅ Taylor Roper ⋅ Jimmy Tobin ⋅ Anton Kast ⋅ Alicia Martin ⋅ Sam Sepah ⋅ Amy Pavel
Audio descriptions (AD) are essential for making visual media accessible to blind and low-vision (BLV) users. While Multimodal Large Language Models (MLLMs) offer a scalable solution for generating audio descriptions on-demand, their performance on ``inline'' audio descriptions -- which must fit within existing silences in a video -- remains under-explored, particularly for diverse user-generated content (UGC). We present a comprehensive investigation into MLLM-generated inline audio descriptions. First, we introduce a professionally annotated dataset of 36 hours of videos, each averaging $\sim$5 mins. and total 7k+ descriptions, providing a high-quality benchmark for this task. Second, we propose an evaluation framework based on audio description expert guidelines that provides fine-grained actionable metrics for model improvement. Crucially, our framework evaluates the end-to-end task: it renders the generated scripts into audio tracks to assess their fit within the original video's natural silences, minimizing disruptive overlaps. Our analysis reveals that while frontier MLLMs accurately describe visual events, significant gaps persist in audio-visual timing, quality, and narrative flow. Finally, we validate the framework and conduct studies with BLV participants and sighted raters, finding that while users narrowly prefer human-authored audio descriptions, they still find MLLM-generated audio descriptions highly beneficial for comprehension. We provide guidance on the current capabilities of MLLMs for audio descriptions and release our dataset and evaluation suite to the community.
Mind the Gap: The Divergent Rebound Dynamics of Diffusion and Autoregressive Model
Qiwen Jin ⋅ Zhengyu Zhou ⋅ Weiwei Liu ⋅ Xiuwen Gong
Post-training alignment can be fragile: in autoregressive models (ARMs), subsequent fine-tuning on counter-aligned or distribution-shifted data can erode previously aligned behavior. While this rebound phenomenon has been studied mainly in ARMs, how such fine-tuning triggers a similar rebound in diffusion language models (DLMs) remains poorly understood. We conduct extensive fine-tuning experiments across diverse datasets, perturbation sizes, model scales, alignment algorithms, and inference hyperparameters, and find that DLMs consistently exhibit a slower rebound pattern than ARMs do. Furthermore, while ARMs fine-tuned on more positive data suffer a steeper degradation under reverse updates, DLMs exhibit ordering stability: models fine-tuned on larger positive alignment sets retain higher performance as negative data increases. We propose a compression perspective to account for this behavior. Unlike ARMs, whose compression naturally follows a fixed left-to-right order, DLMs generate text through iterative masked denoising. We formalize this distinction through a template-wise compression theory for DLMs. The resulting elasticity theory explains why counter-aligned updates lead to much slower alignment erosion in DLMs. Our experiment findings indicate that rebound is shaped not only by post-training data and objectives, but also by the underlying generation mechanism.
Min-p, Max Exaggeration: A Critical Analysis of Min-p Sampling in Language Models
Rylan Schaeffer ⋅ Joshua Kazdan ⋅ Yegor Denisov-Blanch
Sampling from language models impacts the quality and diversity of outputs, affecting both research and real-world applications. Recently, Nguyen et al. 2024's "Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs" introduced a new sampler called min-p, claiming it achieves superior quality and diversity over established samplers such as basic, top-k, and top-p sampling. The significance of these claims was underscored by the paper's recognition as the 18th highest-scoring submission to ICLR 2025 and selection for an Oral presentation. This paper conducts a comprehensive re-examination of the evidence supporting min-p and reaches different conclusions from the original paper's four lines of evidence. First, the original paper's human evaluations omitted data, conducted statistical tests incorrectly, and described qualitative feedback inaccurately; our reanalysis demonstrates min-p did not outperform baselines in quality, diversity, or a trade-off between quality and diversity; in response to our findings, the authors of the original paper conducted a new human evaluation using a different implementation, task, and rubric that nevertheless provides further evidence min-p does not improve over baselines. Second, comprehensively sweeping the original paper's NLP benchmarks reveals min-p does not surpass baselines when controlling for the number of hyperparameters. Third, the original paper's LLM-as-a-Judge evaluations lack methodological clarity and appear inconsistently reported. Fourth, community adoption claims (49k GitHub repositories, 1.1M GitHub stars) were found to be unsubstantiated, leading to their removal; the revised adoption claim remains misleading. We conclude that evidence presented in the original paper fails to support claims that min-p improves quality, diversity, or a trade-off between quality and diversity.
MIRAGE: Adaptive Multimodal Gating for Whole-Brain fMRI Encoding
Abdulkadir Gokce ⋅ Badr AlKhamissi ⋅ Martin Schrimpf
Recent progress in task-optimized neural networks has established encoding models as a powerful tool for predicting brain responses to naturalistic stimuli, yet most existing approaches rely on unimodal representations. The emergence of omni-modal foundation models and rich multimodal neural datasets enables shared encoding models that jointly integrate visual, auditory, and linguistic information across subjects. We introduce MIRAGE, a brain encoding framework that predicts whole-brain fMRI responses to naturalistic audiovisual stimuli with paired transcripts. MIRAGE extracts representations from a single pretrained omni-modal backbone through three modality-specific cross-attention modules whose latent queries adaptively aggregate features across the backbone's 48 layers, and combines them through a transformer-based brain encoder and a subject-specific linear head over the cortical parcels. On the Algonauts benchmark, MIRAGE achieves state-of-the-art results on the out-of-distribution dataset. Controlled comparisons show that native multimodal fusion (features taken from a single jointly trained model) consistently outperforms post-hoc fusion of independently extracted unimodal streams, across architectural levels and backbones. Beyond predictive accuracy, the learned attention weights are directly inspectable: each modality's gating module discovers a distinct depth profile over the backbone, and each modality traces a distinct, anatomically structured pattern across cortex. Together, these results propose adaptive layer-wise aggregation of natively multimodal features as a more generalizable, interpretable, and accurate approach for whole-brain encoding.
MIRA: Reinforcing Multimodal Reasoning via Deceptive Contextual Augmentation
Zhihan Yin ⋅ Jianxin Liang ⋅ Yifeng Yao ⋅ Nonghai Zhang ⋅ Bingyue Peng ⋅ Chunguang Qie ⋅ Huishuai Zhang ⋅ Dongyan Zhao
While Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced multimodal reasoning, existing frameworks suffer from a critical “perception–reasoning gap.” Due to an over-reliance on seemingly relevant visual tokens, even minor perceptual perturbations can propagate through the reasoning process, leading to compounding hallucinations. To address this issue, we propose MIRA, a training framework that improves robustness to erroneous visual contexts while enhancing logical reasoning ability. By injecting filtered, deceptive visual contexts during training, we construct challenging “hard traps” that stress-test the model’s robustness to misleading information. Furthermore, we introduce a group-wise reflection-triggering mechanism that enables the model to autonomously detect and rectify inconsistencies between visual cues and logical constraints, which is ultimately internalized as an intrinsic policy behavior. Importantly, MIRA does not rely on human annotations. Experimental results across multiple benchmarks show that MIRA improves the base model by 7% and demonstrates superior robustness against deceptive contexts. These findings highlight that active error recovery is essential for reliable multimodal reasoning. Comprehensive ablation studies and analyses further provide insights into why MIRA is effective.
MIST: Reliable Streaming Decision Trees for Online Class-Incremental Learning via McDiarmid Bound
Phu-Hoa Pham ⋅ Chi Nguyen Tran ⋅ Phu-Quy Nguyen-Lam ⋅ Dao S Minh ⋅ Trung-Kiet Huynh ⋅ Long Tran-Thanh
Streaming decision trees are natural candidates for open-world continual learning, as they perform local updates, enjoy bounded memory, and static decision boundaries. Despite these, they still fail in online class-incremental learning due to two coupled miscalibrations: (i) their split criterion grows unreliable as the class count $K$ expands, and (ii) the absence of knowledge transfer at split time. Both failures share a common root: the range of Information Gain intrinsically scales with $\log_2 K$. Consequently, any Hoeffding-style confidence radius derived from it must inevitably grow with the class count, making a $K$-independent split criterion structurally impossible, taking away the potential benefits of applying streaming decision trees to continual learning. To fix this issue, we present \mist{} (McDiarmid Incremental Streaming Tree), which resolves both failures through three integrated components: (i) a tight, $K$-independent McDiarmid confidence radius for Gini splitting that acts as a structural regulariser; (ii) a Bayesian inheritance protocol that projects parent statistics to child nodes via truncated-Gaussian moments, with variance reduction guarantees strongest precisely when splitting is most conservative; and (iii) per-leaf KLL quantile sketches that support both continuous threshold evaluation and geometry-adaptive leaf prediction from a single data structure. On standard and stress-test tabular streams, \mist{} is competitive with global parametric methods on near-Gaussian benchmarks and uniquely robust on non-Gaussian geometry where SOTA benchmarks collapse.
Mitigating Compounding Errors in Online Reinforcement Learning via Optimal Transport Regularized Flow Matching
Boxiang Tao ⋅ Lei Guo ⋅ Bin Wang ⋅ Zexin Wang ⋅ HongboDou
Model-based reinforcement learning (MBRL) promises high sample efficiency but is often crippled by compounding model errors: small prediction inaccuracies accumulate rapidly over multi-step rollouts, leading to biased policy optimization. We propose FTA, a trajectory-level generative framework using conditional flow matching (CFM) instead of step-wise dynamics. To avoid curved, energy-inefficient paths, we introduce an optimal transport regularizer grounded in the Benamou–Brenier formula. This regularizer minimizes kinetic energy while enforcing exact state displacement, yielding straight, low-energy trajectories with provable linear error propagation. Since a static generator fails in online RL due to (i) shifting data distributions and (ii) mismatch between training (historical) and desired (high-return) conditionals, we add periodic retraining from the replay buffer and a value-guided ODE sampler that biases generation toward high returns via the current Q-function. Theoretically, the regularizer drives the flow toward the Wasserstein-2 optimal transport map, and value-guided sampling reduces asymptotic Q-estimate bias. On MuJoCo benchmarks, FTA consistently outperforms strong MBRL and model-free baselines in sample efficiency and final performance, and generalizes across tasks. By mitigating compounding errors at the trajectory level and adapting to online shifts, FTA offers a robust, data-efficient alternative to traditional dynamics models.
Mitigating Confidence Miscalibration in Open-World Semi-Supervised Learning
Wenqiang Wu ⋅ Feng Wang ⋅ Jiye Liang ⋅ Liang Bai
Open-world semi-supervised learning aims to use limited labeled data from known classes to classify unlabeled samples that contain both known and novel classes. Most existing methods over-rely on known classes and tend to generalize incorrectly in complex open-world settings. This leads to degraded pseudo-label quality and the model is often overconfident in wrong predictions. To address this, this paper proposes a **M**itigating **C**onfidence **M**iscalibration Method in **Open**-World Semi-Supervised Learning (**OpenMCM**). By leveraging discriminative knowledge from known classes to guide learning in the novel class space, our method effectively mitigates overconfident misclassifications and improves discrimination accuracy. Furthermore, we propose an adaptive threshold calibration strategy to independently determine optimal decision boundary for each class. By integrating a high-reliable pseudo-label fusion mechanism, we enhance recognition stability for novel classes. Experiments on three benchmark datasets ($i.e.$, CIFAR-10, CIFAR-100, and ImageNet-100) show that our framework achieves performance improvements over existing state-of-the-art methods.
Mitigating Knowledge Conflicts in Retrieval-Augmented Generation via Inference-Time Representation Editing
Dahyun Jung ⋅ Jaehyung Seo ⋅ Heuiseok Lim
Retrieval-Augmented Generation (RAG) has been widely adopted to enable Large Language Models (LLMs) to ground their responses in external knowledge sources. Nonetheless, recent studies show that conflicts between the retrieved external knowledge and the model’s parametric knowledge can lead to hallucinatory outputs, and this problem is exacerbated when the retrieved documents contain noise. In this work, we propose Conflict-Aware Representation Editing (CARE), an inference-time representation editing method for improving LLM robustness under noisy knowledge conflict settings. CARE learns a latent editing direction from intermediate representations associated with correct and incorrect generations, and uses this direction to mitigate conflict-associated activation patterns during inference. We evaluate CARE across Question Answering (QA) benchmarks and LLMs, showing improvements over retrieval-based and conflict-mitigation baselines, especially under noisy retrieval conditions.
MixScentNet: A Multiscale Graph-based Framework for Predicting Scent Mixture Perception
Xingran Liao ⋅ Mingliang Zhou ⋅ Weisi Lin
Predicting the olfactory perception of scent mixtures remains a fundamental challenge in computational neuroscience. Existing methods derive mixture-level features by encoding individual components with a single-molecule feature encoder pretrained on annotated olfactory data and aggregating these features via mean pooling, concatenation, or self-attention. This paradigm faces two critical limitations: the scarcity of annotated olfactory data and the inability to capture complex interactions among mixture components. We present MixScentNet, a multiscale graph neural network framework, to address these challenges with two key innovations. First, we introduce a self-supervised pretraining strategy that uses molecular graph structures with RDKit features to predict corresponding Mordred descriptors, enabling the model to learn rich physicochemical knowledge without relying on scarce annotated olfactory data. Second, we propose a novel mixture-as-graph paradigm to model the constituent molecules as nodes in the mixture-level graph, aligning with the fact that humans holistically perceive scent mixtures. We further process the graph via a graph attention network V2 (GATv2) to capture the high-order molecular interactions. MixScentNet differs from existing methods in terms of its pretraining strategy and ability to capture mixture-level features, achieving state-of-the-art performance on both mixture-level olfactory label prediction and perceptual distance estimation tasks. We also find that MixScentNet can reproduce olfactory white phenomena, indicating that the model enjoys clear interpretability grounded in psychophysics. The demo code is available in the Anonymous link \url{https://anonymous.4open.science/r/Odor-D2DB}
Mixture-of-Experts for Online Matrix Completion on a Drifting Union of Subspaces
Renpu Liu ⋅ Jing Yang
Many partially observed data streams are globally high-rank yet locally low-rank: user–item interactions in recommender systems span multiple preference groups, distinct operating regimes in sensor and network monitoring induce different low-dimensional structures, and the active latent subspace in adaptive systems shifts over time. We formalize this as online matrix completion on a drifting union of subspaces: each sample lies in one of K low-dimensional subspaces, only a random subset of its coordinates is observed, and the subspaces drift across epochs. We propose MoSAIC (MoE Subspace Adaptation via Incremental Completion), a method based on a routed mixture of low-rank experts trained in two phases. A base model is first pre-trained on samples from a fixed source distribution. It is then adapted on the non-stationary stream by routing each incoming sample to a single expert and updating only that expert. Our analysis identifies sufficient conditions under which the router remains correct with high probability throughout learning, while within each epoch the routed experts contract toward the current subspaces at a sublinear rate. The technical core is a uniform routing concentration argument that converts the random time steps on which each expert is updated into a deterministic time scale, reducing the per-expert analysis to a tractable stochastic recursion. Experiments on streams with repeated non-stationary changes corroborate the theory and show clear improvements over competitive baselines.
Mixture-of-Hierarchical Experts: Optimized Mamba Architecture for Vision Diffusion
Yejun Jung ⋅ Dongyun Kim ⋅ Jinsun Park
Existing Mamba-based diffusion backbones struggle to achieve competitive performance in vision tasks, as purely sequential propagation lacks an explicit spatial inductive bias and relies on a fixed computation pattern across timesteps. We address this limitation by proposing \textbf{Mixture-of-Hierarchical Experts (MoH)}, a Mamba-based diffusion backbone that performs timestep-conditioned routing over architectural components. MoH dynamically selects representation levels by its hierarchical expert space, adapts feature transformations, and adjusts Mamba scan depth according to the denoising state, enabling flexible modeling of spatial structure and multi-scale dependencies. This design leads to substantial performance gains. On unconditional CelebA-HQ $256 \times 256$, MoH reduces FID from 14.27 to 7.99, and on MS-COCO, it improves FID from 41.80 to 20.00, significantly outperforming prior Mamba-based models while improving training efficiency. These results suggest that routing over architectural components can provide an effective mechanism for improving Mamba-based diffusion backbones.
Mixture of Probes: Learning with Privileged Modalities in Multimodal LLMs Through Probing
Dominick Reilly ⋅ Qiyu Wu ⋅ Hiromi Wakaki ⋅ Srijan Das ⋅ Yuki Mitsufuji
Multimodal Large Language Models (MLLMs) are typically designed under the assumption that all modalities available during training will also be accessible at inference. However, many real-world settings violate this assumption, requiring models to operate under a privileged modality setting, where auxiliary modalities are available only during training. While these modalities contain valuable information, existing MLLMs largely fail to leverage them effectively, as they treat modalities as interchangeable inputs rather than sources of complementary supervision. We propose Mixture of Probes (MoP), a novel framework that disentangles modality-specific and modality-general signals within the MLLM, allowing the model to preserve modality-dependent structure while learning transferable representations across modalities. At its core, MoP achieves this through a structured probing mechanism that extracts and organizes information from intermediate representations of a shared modality encoder, rather than relying only on final-layer alignment as done in existing MLLMs. To support this disentanglement, we further introduce MoP Cross-modal Probe Training (MoP-X), a training strategy for MoP centered around a probe disentanglement loss that prevents probe collapse and encourages cross-modal learning. We evaluate MoP across two domains spanning eight tasks and four modalities under a comprehensive evaluation protocol tailored to the privileged modality setting, where each modality is independently treated as the sole input at inference time. MoP consistently outperforms strong MLLM baselines, achieving up to 65% relative improvement, demonstrating that auxiliary modalities, even when unavailable at inference, can provide substantial gains when effectively leveraged during training.
Mixture-of-Top-$k$ Attention: Efficient Attention as Scalable Fast Weights
Qishuai Wen ⋅ Zhiyuan Huang ⋅ meng xianghan ⋅ Wei He ⋅ Chun-Guang Li
The vanilla self-attention mechanism in Transformers can be viewed as a two-layer fast-weight MLP, whose weights are dynamically induced by inputs and whose hidden dimension is equal to the sequence length $N$. As the context extends, the expressive capacity of such an $N$-width MLP increases, but it becomes unscalable for extremely long sequences. Recently, this fast-weight perspective has motivated the Mixture-of-Experts (MoE) attention, which partitions the sequence into rigid blocks, treats them as fast-weight experts, and sparsely routes the tokens to them. In this paper, we elevate this perspective to a unifying framework for efficient attention mechanisms, interpreting them as making fast weights scalable through either routing or compression, and organizing them into a five-dimensional taxonomy. Then, we propose \textbf{Mi}xture-of-\textbf{T}op-$k$ \textbf{A}ttention (\textbf{MiTA}), which employs a small set of landmark queries to gather top-$k$ attended key-value pairs as query-aware and deformable routed experts, while compressing the $N$-width MLP into a narrower shared expert. Consequently, MiTA improves the flexibility of prior MoE attention, from rigid to deformable fast-weight experts, as well as the scalability of prior top-$k$ attention, from query-specific set to reusable top-$k$ set. Our experiments on vision tasks demonstrate the superior effectiveness and efficiency of MiTA, while also uncovering intriguing properties such as an emergent token-pruning effect and easy generalization from standard attention.
MMCompass: Diagnosing Position Bias in Generative Multimodal Reward Models
Hongbo Zhao ⋅ Mingkun Yang ⋅ Yantao Liu ⋅ Yichang Zhang ⋅ Fei Huang ⋅ Zhibo Yang ⋅ Dayiheng Liu ⋅ Shuai Bai
Generative multimodal reward models are increasingly used to rank responses, guide alignment, and serve as automatic judges, making their reliability a central concern. However, existing multimodal reward benchmarks remain limited in annotation rigor, domain coverage, response model diversity, and robustness-oriented evaluation, making it difficult to assess reward models under realistic and diagnostic settings. To address these gaps, we introduce MMCompass, a new benchmark for evaluating generative multimodal reward models. MMCompass contains carefully verified preference pairs spanning 10 multimodal domains, with responses collected from 16 vision-language models. The benchmark is constructed through a multi-stage pipeline including AI-assisted difficulty filtering, blind multi-annotator verification, and expert review. To enable more diagnostic evaluation, we adopt three complementary metrics: overall accuracy, symmetric accuracy, and verdict guess rate, which jointly measure judgment quality, order robustness, and position bias. Evaluations on MMCompass show that position bias is widespread in our evaluated setting, even among strong multimodal judges and that standard accuracy alone often overestimates reliability. Beyond this diagnosis, we further provide CompassRM as a mitigation baseline for improving cross-order consistency. Built with dual-order supervised fine-tuning and the proposed GXPO (Group Cross Policy Optimization), CompassRM improves robustness over its instruct-model backbones on MMCompass, VLRewardBench, and MMRewardBench, while also improving overall judgment accuracy.
MMDiff: Multimodal Model Diffing for Feature Discovery and Control
Lachin Naghashyar ⋅ Hunar Batra ⋅ Ashkan Khakzar ⋅ Philip Torr ⋅ Ronald Clark ⋅ Christian Schroeder de Witt ⋅ Constantin Venhoff
Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection, hidden states that are decomposed into interpretable feature directions using sparse autoencoders (SAEs) do neither readily isolate which features are changed by multimodal training, nor are they directly useful for targeted control. We introduce MMDiff, a multimodal model-diffing framework that trains multimodal SAEs and turns them into feature-level interfaces for discovering and controlling multimodal behavior. MMDiff supports three uses: (i) feature isolation, by diffing a base-LM SAE against its multimodal-adapted counterpart to identify features altered by multimodal training; (ii) task-specific feature detection, via per-token contrastive firing analysis that isolates causal features; and (iii) feature-level control, by causally removing or steering the discovered feature directions. We train multimodal SAEs for two MLLM families, LLaVA-MORE and PaliGemma 2, and evaluate on visual-spatial understanding, multimodal safety, and OCR. MMDiff discovers sparse, causally specific features whose removal selectively degrades target behaviors by an average of 12\% on spatial tasks and 17\% on OCR, and reduces attack success rate by 24\% on multimodal safety attacks, with no impact on VQA performance. Steering these features improves spatial and OCR accuracy by +3.6\% and +1.5\% on average over a standard single-layer steering baseline. These results show that multimodal SAEs can serve not only as interpretability tools, but as mechanisms for auditing, steering, and controlling MLLMs behavior toward safer and more capable generations.
mmLIP: mmWave Radar-Language Interactive Pretraining via Point Confidence
Jeongwan Shin ⋅ Jaehyeon Kim ⋅ Jae-Ho Choi
Millimeter-wave (mmWave) radar provides core sensing capabilities to traditional vision, particularly under occlusion, low-light, and privacy-constrained conditions. Recent efforts have explored integrating radar signals with large language models (LLMs) for high-level semantic reasoning over the received radar data. However, existing approaches typically project radar inputs into borrowed vision or LiDAR embedding spaces, which are not tailored to radar-specific physical cues such as Doppler and reflection intensity. This design introduces a representational bottleneck, hindering the extraction of informative radar semantics. To overcome this limitation, we propose \textbf{mmLIP}, a novel radar-language interactive pretraining framework that learns radar-specific alignment without relying on external embedding spaces. Our approach directly aligns radar representations with the text embedding space via a point-level contrastive objective, enabling fine-grained correspondence between radar points and textual tokens. In addition, we introduce a confidence-aware contrastive learning mechanism that adaptively reweights radar tokens based on their semantic relevance, promoting informative signals while suppressing clutter. Supported by our newly curated radar-text pairs, mmLIP captures structured radar semantics. Extensive experiments demonstrate that mmLIP successfully integrates with diverse LLMs or vision-language models, achieving state-of-the-art performance on zero-shot bidirectional retrieval as well as text generation tasks, including captioning and question answering.
MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos
Arushi Goel ⋅ Sreyan Ghosh ⋅ Vatsal Agarwal ⋅ Nishit Anand ⋅ Kaousheik Jayakumar ⋅ Lasha Koroshinadze ⋅ Yao Xu ⋅ Katie Lyons ⋅ James Case ⋅ Siddharth Gururani ⋅ Karan Sapra ⋅ Kevin Shih ⋅ Abhinav Shrivastava ⋅ Ramani Duraiswami ⋅ Dinesh Manocha ⋅ Andrew Tao ⋅ Bryan Catanzaro ⋅ Mohammad Shoeybi ⋅ Wei Ping
Multimodal Large Language Models (MLLMs) have shown strong performance in visual and audio understanding when evaluated in isolation. However, their ability to jointly reason over omni-modal (visual, audio, and textual) signals in long and complex videos remains largely unexplored. We introduce MMOU, a new benchmark designed to systematically evaluate multimodal understanding and reasoning under these challenging, real-world conditions. MMOU consists of 16,500 carefully curated questions paired with 9800 web-collected videos of varying length, spanning diverse domains and exhibiting rich, tightly coupled audio-visual content. The benchmark covers 13 fundamental skill categories, all of which require integrating evidence across modalities and time. All questions are manually annotated across multiple turns by professional annotators, ensuring high quality and reasoning fidelity. We evaluate 20+ state-of-the-art open-source and proprietary multimodal models on MMOU. The results expose substantial performance gaps: the best closed-source model achieves only 64.2% accuracy, while the strongest open-source model reaches just 46.8%. Our results highlight the challenges of long-form omni-modal understanding, revealing that current models frequently fail to apply even fundamental skills in long videos. Through detailed analysis, we further identify systematic failure modes and provide insights into where and why current models break. Project: https://mmou-2026.github.io/
MM-SCALE: Evaluating Evidence-Grounded Moral Judgment in Vision-Language Models
Eunkyu Park ⋅ Wesley Deng ⋅ Cheyon Jin ⋅ Matheus Kunzler Maldaner ⋅ Jordan Wheeler ⋅ Jason I Hong ⋅ Hong Shen ⋅ Adam Perer ⋅ Kenneth Holstein ⋅ Motahhare Eslami ⋅ Gunhee Kim
Vision-Language Models increasingly make socially consequential judgments from image-text inputs, yet current evaluations often stop at verdict-level accuracy: whether the final label is human-aligned, not whether it follows from the correct evidence. We introduce MM-SCALE (Multimodal Moral SCALE), a benchmark for evidence-grounded moral judgment built around a design choice: each image is paired with multiple action scenarios, so models must compare how different actions interact with the same visual context rather than score images in isolation. MM-SCALE contains 8,444 image contexts and 21,977 action scenarios, each annotated with a 5-point moral acceptability rating and a modality-grounding label indicating whether the human judgment relied on text, image, or both. Three tasks evaluate whether models (i) assign calibrated scalar moral scores, (ii) preserve human orderings within a shared visual context, and (iii) ground rationales in the evidence source humans found critical to judgment. Across models, aggregate metrics overstate alignment: NDCG@5 remains high across models, while within-image pairwise accuracy remains only 0.53--0.62 even on scenario pairs whose human mean ratings differ by at least one point. CoT has inconsistent effects on calibration and does not reliably improve within-image ordering. Evidence-grounding evaluations further show that models often describe visual context without making it critical for the verdict. These results reveal an evidence-grounding failure in current VLMs and position MM-SCALE as an evaluation testbed for moral judgment under shared image contexts. This dataset includes potentially harmful and sensitive visual content. Images are intended solely for evaluation.
MMSkills: Towards Multimodal Skills for General Visual Agents
Kangning Zhang ⋅ Shuai Shao ⋅ Wenxiang Jiao ⋅ Qingyao Li ⋅ Jianghao Lin ⋅ Lingyue Fu ⋅ Shijian Wang ⋅ Yuan Lu ⋅ Weiwen Liu ⋅ Weinan Zhang ⋅ Yong Yu
Reusable skills have become a core substrate for improving agent capabilities, yet most existing skill packages encode reusable behavior primarily as textual prompts, executable code, or learned routines. For visual agents, however, procedural knowledge is inherently multimodal: reuse depends not only on what operation to perform, but also on recognizing the relevant state, interpreting visual evidence of progress or failure, and deciding what to do next. We formalize this requirement as \emph{multimodal procedural knowledge} and address three practical challenges: (I) \textbf{what} a multimodal skill package should contain; (II) \textbf{where} such packages can be derived from public interaction experience; and (III) \textbf{how} agents can consult multimodal evidence at inference time without excessive image context or over-anchoring to reference screenshots. We introduce \emph{MMSkills}, a framework for representing, generating, and using reusable multimodal procedures for runtime visual decision making. Each MMSkill is a compact, state-conditioned package that couples a textual procedure with runtime state cards and multi-view keyframes. To construct these packages, we develop an agentic trajectory-to-skill Generator that transforms public non-evaluation trajectories into reusable multimodal skills through workflow grouping, procedure induction, visual grounding, and meta-skill-guided auditing. To use them, we introduce a branch-loaded multimodal skill agent: selected state cards and keyframes are inspected in a temporary branch, aligned with the live environment, and distilled into structured guidance for the main agent. Experiments across GUI and game-based visual-agent benchmarks show that MMSkills consistently improve both frontier and smaller multimodal agents, suggesting that external multimodal procedural knowledge complements model-internal priors. The code is accessible in \href{https://anonymous.4open.science/r/MMSkills}{https://anonymous.4open.science/r/MMSkills}.
MobileMoE: Scaling On-Device Mixture of Experts
Yanbei Chen ⋅ Hanxian Huang ⋅ Ernie Chang ⋅ Jacob Szwejbka ⋅ Digant Desai ⋅ Zechun Liu ⋅ Vikas Chandra ⋅ Raghuraman Krishnamoorthi
Mixture-of-Experts (MoE) has become the de facto architecture for hundred-billion-parameter language models, yet its advantages at sub-billion scales for on-device deployment remain largely unexplored. To close this gap, we present MobileMoE, a family of on-device MoE language models with sub-billion active parameters (0.3-0.9B active and 1.3-5.3B total) that establish a new Pareto frontier for on-device LLMs. We first formulate an on-device MoE scaling law that jointly optimizes MoE architecture under mobile memory and compute constraints, identifying an on-device sweet spot -- moderate sparsity with fine-grained and shared experts -- that is simultaneously memory and compute-optimal. Building on the derived architectures, we train MobileMoE with a four-stage recipe covering pre-training, mid-training, instruction fine-tuning, and quantization-aware training, all on open-source datasets. Across 14 benchmarks, MobileMoE matches or exceeds leading on-device dense LLMs with 2-4$\times$ fewer inference FLOPs, and matches or surpasses the state-of-the-art MoE OLMoE-1B-7B with up to 60\% fewer parameters. To bridge the last mile to mobile deployment, we provide the first efficient MoE inference on commodity smartphones with comprehensive on-device profiling. At comparable total parameters, MobileMoE delivers 2-4$\times$ faster prefill and decode, and $\sim$30\% lower runtime peak RAM than the dense baseline MobileLLM-Pro.
MOCHA: Discovering Multi-Order Dynamic Causal Structure in Temporal Point Processes
Yunyang Cao ⋅ Juekai Lin ⋅ Wenhao Li ⋅ Bo Jin
Modeling event sequences requires understanding when future events will occur and how event types causally influence one another over time. Existing temporal point process (TPP) models typically assume static or first-order dependence structures, which limits their ability to capture the dynamic and multi-order causal mechanisms commonly observed in real-world systems. We propose MOCHA, a Multi-Order Causal Hierarchical Architecture for multivariate TPP that jointly models time-varying causal structure and multi-hop influence propagation in continuous time. MOCHA learns latent dynamic weighted directed acyclic graphs (DAG) over event types, where acyclicity and sparsity constraints promote structurally valid causal graphs. Based on the learned graph, the model decomposes event dynamics into multi-order causal paths, and incorporate both direct influence and indirect propagation into the intensity function. The entire framework is end-to-end differentiable and optimized jointly for event prediction and causal structure discovery. Experiments on seven real-world datasets from different domains show that MOCHA consistently achieves superior negative log-likelihood, while also recovering dynamic causal patterns that align with domain knowledge. These results demonstrate that MOCHA provides an effective framework for dynamic causal structure learning in TPP.
Large Language Models (LLMs) are trained to support an increasing number of languages, yet their predefined tokenizers remain a bottleneck for adapting models to lower-resource or distinct-script languages. Existing tokenizer transfer methods typically rely on semantic heuristics to initialize new embeddings, ignoring higher-layer model dynamics and limiting transfer quality. We propose Model-Aware Tokenizer Transfer (MATT), a method that incorporates model internals into the tokenizer transfer process. MATT introduces an Attention Influence Modeling (AIM) objective that distills inter-token communication patterns from a source model into a target model with a new tokenizer, providing an efficient warm-up before standard language modeling. Unlike approaches that focus solely on embedding similarity, MATT leverages attention behavior to guide embedding initialization and adaptation. Experiments across diverse linguistic settings show that MATT recovers a large fraction of the original model’s performance within a few GPU hours, outperforming heuristic baselines. These results demonstrate that incorporating model-level signals offers a practical and effective path toward robust tokenizer transfer in multilingual LLMs.
Model Capacity Determines Grokking through Competing Memorisation and Generalisation Speeds
Yiding Song ⋅ Hanming Ye
Existing accounts of grokking explain the phenomena in terms of mechanistic frameworks such as circuit efficiency or lazy-to-rich transitions. However, despite a known dependence between grokking and model size, how model capacity shapes grokking remains an open question. We give an information-theoretic account of this relationship on the task of modular arithmetic, showing that grokking does not immediately occur when a model becomes large enough to memorise the training set, but rather emerges as the outcome of a competition between two measurable timescales: a memorisation speed $T_{\text{mem}}(P)$ and a generalisation speed $T_{\text{gen}}(P)$, both of which are functions of model parameter count $P$. Adapting the information capacity framework of Morris et al. (2025), we estimate $T_{\text{mem}}(P)$ on random-label data of equivalent complexity and $T_{\text{gen}}(P)$ on the modular task itself, and show that grokking emerges close to the parameter scale where these timescales intersect. The framework also suggests an empirical model for predicting memorisation speed given model capacity and dataset complexity, recovering the previously reported empirical observation that larger models memorise faster. Overall, we motivate the formalisation of different learning timescales as important abstractions to study when explaining how model capacity shapes grokking on algorithmic tasks.
Modeling Whole-Slide Images as Dynamic Tumor Microenvironment Fields
Lei Wu ⋅ Jiashuai Liu ⋅ Di Zhang ⋅ Zhangpeng Gong ⋅ Yingkang Zhan ⋅ Yi Niu ⋅ Jiusong Ge ⋅ Chunze Yang ⋅ Kai Yi ⋅ Mireia Crispin-Ortuzar ⋅ Chen Li ⋅ Zeyu Gao
Due to the gigapixel-scale nature of whole-slide images (WSIs), weakly supervised WSI analysis is commonly formulated as a multiple instance learning (MIL) problem, where patch-level features are aggregated into slide-level representations. However, diagnostic and prognostic evidence often arises from spatially coherent tumor microenvironment regions and their interactions, rather than isolated patches alone. Existing patch-level or static region-based methods usually overlook how tissue regions should be adaptively formed and subsequently evolved through microenvironment interactions across heterogeneous boundaries. In this paper, we propose Concept-Guided Tumor Microenvironment Evolution (TMEvolve), a reaction-diffusion-inspired framework that models WSIs as latent tumor microenvironment fields over discrete patch graphs. TMEvolve instantiates this view as a learnable graph-discretized evolution process over patch neighborhoods. It first forms adaptive soft tissue regions as coherent microenvironment units, then performs pseudo-time evolution through two complementary local dynamics: intra-region diffusion, which stabilizes latent states within coherent tissue compartments, and concept-guided boundary flux, which propagates visual feature signals and language-derived concept signals across heterogeneous region interfaces. The evolved microenvironment regions are finally aggregated for slide-level prediction. We evaluate TMEvolve on six datasets across three weakly supervised WSI tasks: survival prediction, gene expression prediction, and histological subtype classification. TMEvolve consistently improves over representative MIL methods, pathology foundation models, and concept-guided baselines. Ablation studies and visualizations further support the effectiveness and interpretability of TMEvolve, highlighting the value of dynamic region modeling and boundary interaction.
Models Designed to Forget: Machine Unlearning via Key Deletion
Sonia Laguna ⋅ Jorge da Silva Gonçalves ⋅ Moritz Vandenhirtz ⋅ Alain Ryser ⋅ Irene Cannistraci ⋅ Julia Vogt
Machine unlearning for vision models is rapidly becoming a practical requirement, driven by privacy regulations, data errors, and the need to remove harmful or corrupted training images. Despite this, most existing approximate unlearning methods tackle the problem from a post-hoc perspective. They attempt to erase the influence of targeted samples through parameter updates that typically require access to the full training data. This creates a mismatch with real deployment scenarios where unlearning requests can be anticipated, revealing a fundamental limitation of post-hoc approaches. We motivate unlearning by design, a novel paradigm for approximate methods in which models are directly trained to support forgetting as an inherent architectural capability. We instantiate this idea with Machine UNlearning via KEY deletion (MUNKEY), a memory-augmented transformer that decouples instance-specific memorization from model weights. Here, unlearning corresponds to removing the instance-identifying key, enabling zero-shot forgetting without weight updates or access to the original samples or labels. Across natural image benchmarks, fine-grained visual recognition, and medical datasets, MUNKEY outperforms all post-hoc baselines. Our results establish that unlearning by design enables fast, deployment-oriented unlearning while preserving predictive performance.
MODULE: A Mutual-Promoting Deep Unfolding Framework Towards Degradation-Robust Multi-modal Image Fusion
Han Xu ⋅ Yunfei Deng ⋅ Jiayi Ma ⋅ Guangcan Liu
Multi-modal image fusion is crucial for comprehensive scene representation, yet real-world degradations compromise its efficacy, necessitating degradation-robust fusion paradigms. Existing degradation-robust multi-modal image fusion methods are often hindered by cascaded sub-optimal solution or black-box opacity. To address these challenges, we propose a theory-inspired mutual-promoting deep unfolding framework. It reformulates degradation-robust image fusion as a joint optimization problem, utilizing the degradation model and modality generation mechanism to explicitly model both single-task priors and cross-task dependencies. By decomposing the complex multi-task optimization problem into task-related iterative subproblems, the framework establishes a bidirectional reciprocal information flow between restoration and fusion. This process is further unfolded into a multi-stage deep neural network, where each component explicitly corresponds to a specific mathematical operation. In our framework, the fused cross-modal prior actively regularizes the ill-posed restoration, while the purified modality features continuously refine the fusion output. It ensures a transparent architecture that combines the merits of both model-based and data-driven methodologies. Extensive experiments demonstrate that our paradigm achieves state-of-the-art performance while providing superior interpretability.
MolHIT: Advancing Molecular-Graph Generation with Hierarchical Discrete Diffusion Models
Hojung Jung ⋅ Rodrigo Hormazabal ⋅ Jaehyeong Jo ⋅ Youngrok Park ⋅ Kyunggeun Roh ⋅ Se-Young Yun ⋅ Sehui Han ⋅ Dae-Woong Jeong
Molecular generation with diffusion models has emerged as a promising direction for AI-driven drug discovery and materials science. While graph diffusion models have been widely adopted due to the discrete nature of 2D molecular graphs, existing models suffer from low chemical validity and struggle to meet the desired properties compared to 1D modeling. In this work, we introduce MolHIT, a powerful molecular graph generation framework that overcomes long-standing performance limitations in existing methods. MolHIT is based on the Hierarchical Discrete Diffusion Model, which generalizes discrete diffusion to additional categories that encode chemical priors, and decoupled atom encoding that splits the atom types according to their chemical roles. Overall, MolHIT achieves new state-of-the-art performance on the MOSES dataset with near-perfect validity for the first time in graph diffusion, surpassing strong 1D baselines across multiple metrics. We further demonstrate strong performance in downstream tasks, including multi-property guided generation and scaffold extension.
More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models
Xurui Song ⋅ Shuo Huai ⋅ Jingjing Jiang ⋅ Jiayi Kong ⋅ Jun Luo
Vision-Language Model (VLM) driving agents promise explainable end-to-end autonomy by first producing natural-language reasoning and then predicting trajectory planning. However, whether planning is causally driven by this reasoning remains a critical but unverified assumption. To investigate this, we build DriveMind, a large-scale driving Visual Question Answering corpus with plan-aligned Chain-of-Thought (CoT), automatically generated from nuPlan. Our data generation process converts sensors and annotations into structured inputs and, crucially, separates priors from to-be-reasoned signals, enabling clean information ablations. Using DriveMind, we train representative VLM agents with Supervised Fine-Tuning and Group Relative Policy Optimization and evaluate them with nuPlan’s metrics. Our results, unfortunately, indicate a consistent causal disconnect in reasoning-planning: removing ego/navigation priors causes large drops in planning scores, whereas removing CoT produces only minor changes. Attention analysis further shows that planning primarily focuses on priors rather than the CoT. Based on this evidence, we propose the Reasoning-Planning Decoupling Hypothesis, positing that the training-yielded reasoning is an ancillary byproduct rather than a causal mediator. To enable efficient diagnosis, we introduce a novel, training-free probe that measures an agent's reliance on priors by evaluating its planning robustness against minor input perturbations. In summary, we provide the community with a new dataset and a diagnostic tool to evaluate the causal fidelity of future models.
MorphGen: Controllable Cell-Image Generation with Biological Representation Alignment
Berker Demirel ⋅ Marco Fumero ⋅ Theofanis Karaletsos ⋅ Francesco Locatello
Simulating in silico cellular responses to interventions is a promising direction to accelerate high-content image-based assays, critical for advancing drug discovery and gene editing. To support this, we introduce MorphGen, a state-of-the-art diffusion-based generative model for fluorescent microscopy that enables controllable generation across multiple cell types and perturbations. MorphGen is trained with an alignment loss that matches its representations to phenotypic embeddings from a biology foundation model, encouraging meaningful morphological patterns consistent with real cell images. Unlike prior approaches that compress multichannel stains into RGB images, sacrificing organelle-specific detail and focusing on a single cell type, MorphGen generates the complete set of fluorescent channels jointly, preserving per-organelle structure and enabling post-generation interpretation. We demonstrate biological consistency with real images via CellProfiler features, and MorphGen attains an FID score over 35% lower than the prior state-of-the-art MorphoDiff. Finally, in a compositional generalization test that holds out cell type--perturbation combinations during training, MorphGen achieves in-distribution quality on 55% of unseen pairs under a seen-calibrated criterion.
Designing high-performing robot morphologies is a grand challenge for developing specialized autonomous agents. However, the vast, combinatorial, and non-differentiable nature of the morphological design space has been a primary obstacle. Existing methods tackle this problem indirectly, relying on either semantically-blind genetic operators or reinforcement learning with predefined modification actions, both of which constrain exploration. In this work, we introduce MorphoGen, a novel framework that reframes morphological design as a code generation problem. MorphoGen leverages large language models (LLMs) to directly iterate the XML files as codes that define an agent’s morphology, solving the original open problem without being limited by any prior constraints or fixed action spaces. Structure-Aware directional feedback is provided to steer the evolution of robot morphologies through prompted mutations and crossovers. Our approach allows the LLMs to apply its understanding of structure and syntax to generate complex and semantically coherent design variations, enabling an unconstrained and efficient exploration of the design space. On a suite of challenging locomotion benchmarks, MorphoGen discovers novel and high-performing morphologies, significantly outperforming strong baselines by over 52.9% in downstream motoring evaluation. Our work unlocks a new paradigm for automated robotic design, demonstrating the effectiveness of LLMs in navigating complex, structured engineering search spaces. Codes for our work are released anonymously at https://anonymous.4open.science/r/MorphoGen-ACC.
MoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts
Peiwen Li ⋅ Shiyang Zhang ⋅ Yangtian Zhang ⋅ Sizhuang He ⋅ David van Dijk ⋅ Rex Ying
Large language model-based multi-agent systems have recently shown strong potential for complex, long-horizon tasks. However, existing methods mainly rely on coarse prompt-level differentiation without parameter adaptation for diverse subtasks, resulting in insufficient inter-agent heterogeneity and limited specialized capability that bottleneck performance on tasks with complex requirements. To address this, we introduce a Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts (MoRSE) that distinguishes agents with (role, subtask)-conditional specialization at both the task structure and parameter levels. To make agents' responsibility explicit at the task structure level, we formulate a task-oriented multi-agent system that decomposes each task into a dependency-aware Directed Acyclic Graph of subtasks and assigns each agent a specific (role, subtask), introducing task-level specialization across collaborating agents. Additionally, to address the diverse role and subtask demands that a single shared base model cannot satisfy, we propose a dynamic Mixture of (role, subtask) LoRA Experts module with a prototype-based semantic router for subtasks, augmenting agents with parameter-level specialization on a shared LLM substrate cost-effectively. Then, to co-optimize experts and router stably for open-ended tasks, we further propose a hierarchical group-relative policy optimization with two-layer credit assignment that disentangles expert quality from routing quality. Experiments on code-generation benchmarks across three backbones demonstrate the effectiveness of our approach, with notable improvements in both whole-task and step-wise performance and great generalization potential in out-of-distribution scenarios.
MotionHalluc: Diagnosing Kinematic Hallucinations in Fine-Grained Motion Reasoning
Weile Guo ⋅ Shenghong He ⋅ Danying Mo ⋅ Chengdong Xu ⋅ Xuexun Liu ⋅ Chao Yu
Motion instruction generation in cross-video comparison aims to produce corrective feedback that describes the differences between a query and a reference motion. However, existing models often generate instructions that exhibit motion hallucinations, failing to reflect actual kinematic differences between paired videos. To systematically investigate these hallucinations, we introduce MotionHalluc, a dedicated benchmark for evaluating motion hallucinations in paired-video comparison. MotionHalluc comprises 1540 fine-grained questions over 553 video pairs, evaluating hallucinations along three core dimensions: (1) directional hallucination, (2) attributional hallucination, and (3) temporal hallucination. Extensive evaluations of state-of-the-art large multimodal models demonstrate high susceptibility to these hallucinations. Furthermore, we provide Perceive-Parse-Verify (PPV) as a training-free measurements extraction and verification baseline that converts candidate instructions into executable measurement queries and supplies kinematic measurements at inference time. Our results show that this simple measurements injection yields an average 10.6% performance gain across models, suggesting that motion reasoning with explicit quantitative measurements is a key factor in reducing hallucinations in cross-video comparison.
M-plicits: Neural Implicit Surfaces via Nested Multiscale Residuals
Vinícius Da Silva ⋅ Isabelle R de Melo ⋅ Matheus Levy de Lima Bessa ⋅ Guilherme Schardong ⋅ Luiz Schirmer ⋅ Andre Araujo ⋅ Nuno Gonçalves ⋅ HELIO C LOPES ⋅ Alberto Raposo ⋅ Luiz Velho ⋅ Tiago Novello
Encoding input coordinates with sinusoidal functions into multi-layer perceptrons (MLPs) has proven effective for implicit neural representations (INRs) of surfaces defined as zero-level sets. However, existing methods often struggle to balance training efficiency, rendering speed, and noise robustness: single-MLP approaches are expensive at inference, grid-based representations are fast but can limit surface smoothness and overfit input noise, and previous multiscale approaches frequently capture noise and produce artifacts due to hard spectral truncation. To address these limitations, we propose M-plicits, a multiscale framework that models surfaces as a residual sum of MLPs trained via a sequence of nested neighborhoods. Unlike existing residual approaches that rely on standard domain-wide sampling and require costly mesh extraction for visualization, our method strictly localizes supervision to narrow bands around the previous zero-level sets. This nested design naturally provides robustness against noisy input data: the coarse network acts as a low-pass filter that establishes a clean geometric prior, while subsequent residuals progressively refine the geometry without fitting to high-frequency artifacts. We further introduce a multiscale sphere-tracing algorithm and a GEMM-based analytical normal computation that bypasses auto-differentiation entirely, yielding high-fidelity real-time rendering. On Stanford and Thingi32, M-plicits achieves the best mean Chamfer distance in the coarse configuration and the best median Chamfer distance and IoU in the fine configuration, with substantially better noise robustness than iNGP, BACON, and IDF, while using an order of magnitude fewer parameters than grid-based baselines. Code and models will be released.
MSCR: Jointly Balancing Modality Utilization and Discovering Synergistic Information
Xinyu Chen ⋅ Liangjian Wen ⋅ Jiang Duan ⋅ Dongkai Wang ⋅ Yong Dai ⋅ Jiayu Bai ⋅ Jianzhuang Liu ⋅ Zhen Tian ⋅ Guoping Qiu ⋅ Zhao Kang
In multimodal learning, modality imbalance often causes one modality to dominate the optimization process, preventing the model from fully exploiting cross-modal interactions. Recent balance-oriented methods, such as Multimodal Competition Regularizer (MCR), have achieved strong performance by alleviating modality imbalance during training. However, existing dynamic balancing methods mainly focus on adjusting modality contributions and do not explicitly consider how modality imbalance affects the formation of synergistic information in multimodal fusion. Yet such synergy is a key source of multimodal advantage, as it reflects predictive gain that cannot be recovered from unimodal evidence alone. To address this gap, we propose Multimodal Synergistic Competition Regularizer (MSCR), a multimodal fusion method that jointly promotes balanced modality utilization and synergistic information discovery. MSCR contains two complementary objectives: an anti-dominance objective that suppresses excessive unimodal influence and creates room for synergistic information discovery, and a synergy objective that explicitly encourages the fused predictor to surpass a non-synergistic unimodal reference. We further adopt an adaptive modulation strategy to emphasize synergy enhancement on samples that require stronger collaborative modeling. Experiments show that MSCR consistently improves performance, enhances measured synergistic information, and remains effective under transformer-based fusion backbones.
MuEdit: An efficient multi-task editing method towards inter-domain knowledge conflicts
Songshi Liang ⋅ Ting-En Lin ⋅ Zixuan Huang ⋅ Yuchuan Wu ⋅ Yongbin Li ⋅ Zihe Wang ⋅ Rui Yan
Large language models (LLMs) require frequent knowledge updates across heterogeneous domains, but existing editing methods fail when applied sequentially to multi-task settings. We trace this collapse to inter-domain null-space misalignment: heterogeneous tasks induce distinct Jacobian geometries, causing their common null space to shrink dramatically and leaving insufficient admissible directions for conflict-free updates. We propose the Conflict Index to quantify this geometric interference, and introduce Mu-Edit, which mitigates multi-task conflicts via (1) conflict-aware ordering to minimize cumulative interference, and (2) dynamic low-rank approximation to expand the null space when ordering alone is insufficient. Experiments across five functional domains and three backbones show that Mu-Edit significantly outperforms existing baselines in multi-task editing performance while preserving general capabilities, and remains effective under incremental task arrival.
Multi-bit LLM Watermarking with Certified Semantic Distortion
Chenxi Gu ⋅ Si Chen ⋅ John Grundy ⋅ Xiaoning Du
Watermarking has emerged as a fundamental mechanism for tracing the provenance of Large Language Model (LLM) outputs. While multi-bit watermarks are highly desirable for embedding rich metadata, they suffer from an inherent trade-off between statistical detectability and text quality degradation. Crucially, existing multi-bit schemes lack theoretical quality guarantees, leaving them susceptible to unbounded semantic distortion. To address this, we introduce CSD, the first multi-bit watermarking framework with certified semantic distortion bounds. In CSD, we derive a closed-form upper bound on the Kullback-Leibler (KL) divergence, mathematically guaranteeing strict limits on semantic distortion at every generation step. Furthermore, rather than relying on static payloads, CSD employs a dynamic payload to maximize statistical detectability. Extensive empirical evaluations across multiple datasets demonstrate that CSD achieves robust detectability while maintaining provably safe text generation.
Multiform Attack for Transferable Cross-Modal Person Re-Identification
Yunpeng Gong ⋅ Can Yang ⋅ Qingyuan Zeng ⋅ Dejun Xu ⋅ Zhiming Luo ⋅ Zhenzhong Wang ⋅ Min JIANG
Cross-modal person re-identification (ReID) adversarial attacks face a fundamental generalization dilemma: existing methods suffer from source-pair overfitting, where the learned perturbations become entangled with the feature covariance and alignment patterns of the source modality pair, thereby limiting their transferability under heterogeneous domain shifts. We argue that the core issue is the difficulty of reducing source-pair-specific bias while maintaining attack effectiveness. To address this, we propose Multiform Attack (MA). The framework first learns a universal attack direction via Mahalanobis-guided gradient optimization to capture the intrinsic covariance structure of the feature manifold, though this perturbation may still retain source-pair-specific bias. To overcome this bias and achieve cross-distribution generalization, the second stage leverages multiple heterogeneous source distributions to optimize for generalization, performing sparse, structure-sensitive residual correction in the discrete pixel space via attention-shift-filtered multi-objective evolutionary search. This two-stage design models the attack as a base direction plus a structured residual, which helps reduce overfitting to the source modality pair and improve transferability to unseen modalities. Extensive experiments show that MA achieves superior transferability across unseen modalities, datasets, and model architectures.
Multilateral Resistance-Guided Graph Message Passing for Trade Flow Prediction
Quansi Li ⋅ Haijing Yu
Trade flows are a central variable for studying international trade networks. Traditional economic models, grounded in economic theory and mathematical derivation, provide structured representations of bilateral trade flows. However, these representations often rely on fixed functional forms, have limited ability to incorporate high-dimensional features, do not directly capture trade inertia, and scale poorly to large and dynamic trade networks. In recent years, graph neural networks have been increasingly used to address some of these limitations because they are naturally compatible with networked trade data and offer strong scalability. However, the feature aggregation process in existing graph neural models remains weakly interpretable from an economic perspective. To address this limitation, we propose MRTGNN, a graph neural network model built on a theory-guided MRTGNN layer derived from the fixed-point equations of multilateral resistance in structural gravity. Theoretically, we prove that an $L$-layer MRTGNN can approximate $L$ iterations of a multilateral-resistance fixed-point operator on a compact domain, with an explicit bound on the iteration error. Empirically, we extensively evaluate MRTGNN on panel data covering 176 countries from 2002 to 2019 against various representative baselines. The full model achieves the best RMSE, $R^2$, and Spearman rank correlation across the reported baselines. Under counterfactual tariff shocks involving China, it produces a Network Spillover Ratio one order of magnitude higher than a generic graph-attention baseline. These results show that theory-guided message passing improves trade flow prediction while enabling more interpretable counterfactual reasoning in global trade networks.
Multiscale Supervised Unbalanced Optimal Transport Flow Matching
Qiangwei Peng ⋅ Lezhi Chen ⋅ Peijie Zhou
Unbalanced optimal transport (UOT) provides a principled framework for modeling single-cell transitions and birth-death dynamics, but its high computational cost limits scalability to large-scale datasets. Although single-cell data often contain hierarchical annotations and known transition priors, existing UOT approximations rarely exploit this multiscale structure or prior knowledge. We introduce Multiscale Supervised Unbalanced Optimal Transport Flow Matching (MUST-FM), a simulation-free framework that scales UOT by leveraging hierarchical data structure. MUST-FM further supports an optional supervised formulation that incorporates transition priors, such as cell lineages, to guide the learning of displacement fields and mass variations. Experiments show that MUST-FM reduces computational overhead while achieving robust and biologically meaningful trajectory inference, enabling dynamic modeling of atlas-scale single-cell datasets.
Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards
Deokgyu Yoon ⋅ Hyungkyu Kang ⋅ Joongkyu Lee ⋅ Byeongchan Kim ⋅ Gyungin Shin ⋅ Sungrae Park ⋅ Min-hwan Oh
Reinforcement learning with verifiable rewards (RLVR) plays a pivotal role in improving the reasoning ability of large language models. However, widely used PPO surrogate objectives are fundamentally *local*, as they rely on a local approximation of the exact policy gradient objective. While this approximation improves stability by reducing the variance induced by importance sampling, it also introduces structural bias into the surrogate objective, which must be controlled through trust region mechanisms. In this work, we introduce the $N$-step forward trace, which augments the PPO surrogate objective using the cumulative likelihood ratio of the next $N-1$ tokens. Building on this idea, we propose $N$-Step Forward-Trace Policy Optimization (NFPO), a practical RLVR algorithm that integrates the $N$-step forward trace into the masked policy gradient framework. NFPO provides a continuous bridge between the PPO surrogate objective and the exact policy gradient objective, offering a principled mechanism for controlling the bias–variance trade-off. Our theoretical analysis shows that, with an appropriate choice of $N$, the proposed objective yields a tighter policy-improvement bound than the standard PPO surrogate. Experiments on comprehensive reasoning benchmarks demonstrate that NFPO consistently improves performance, supporting our theoretical findings.
MultiSTEVE-1s: A Model Zoo and Interpretability Suite for Instruction-Following Vision Agents
Karolis Jucys ⋅ George Adamopoulos ⋅ Özgür Şimşek
A striking case of goal misgeneralization was previously observed in OpenAI's Minecraft agent VPT: it killed villagers standing under some leaves, mistaking them for trees. Although this agent was released publicly, enabling white-box interpretability research, few open-weight model organisms of misalignment exist outside the LLM space. In this work, we release MultiSTEVE-1s -- a model zoo of 140 fine-tuned versions of VPT, over 1,000 training checkpoints, and an interpretability suite for analysing them. Specifically, we use the STEVE-1 training procedure to add instruction-following capabilities to VPT with fixed hyperparameters and controlled variations in training randomness. We demonstrate the utility of MultiSTEVE-1s by showcasing the research it enables. First, for some training runs, the only difference is a least-significant bit flip in a single initialised weight. Others differ in the full randomness for weight initialisation and data. Yet, the single bit-flip setting produces agents that act nearly as differently from each other as the full randomness ones. Second, we use our interpretability suite to show that several known VPT attention heads retain their roles after STEVE-1 fine-tuning, while attention strength to the same behaviourally meaningful frame can vary substantially across agents and checkpoints. Finally, even though the agents are similarly capable in in-distribution tasks, the out-of-distribution behaviour of villager killing can differ substantially between them -- in one setting, an agent kills villagers less than 5\% of the time, while another agent kills them nearly 50\% of the time. Our results show the value of studying multiple similarly trained agents, rather than acting like a behavioural biology lab with only one rat.
MvFFN: Multi-view Floor-Plan Feed-Forward Network for Unposed Wide-Baseline Panorama Layout Reconstruction
Yuguang Li ⋅ Yichuan (Ethan) Deng ⋅ Zixuan Liu ⋅ Ivaylo Boyadzhiev ⋅ Xiuyuan Wu ⋅ Xinyu Zou ⋅ Linda Shapiro ⋅ Alex Colburn
Reconstructing accurate camera poses and floor-plan layouts from sparse, unposed, wide-baseline RGB panoramas remains a challenging open problem. Direct multi-room floor-plan prediction is especially difficult in this setting: pose, geometry, semantics, and global layout structure are tightly coupled, yet a complete floor-plan is hard to learn as a single global output. Our key insight is to keep the model's predictions strictly local. Multi-view Ordered Wall Instance Segmentation (MOIS) is designed around this concept. Each panorama column predicts only the wall it observes plus the next two adjacent walls in clockwise order. The global wall partition, room layouts, and shared coordinates emerge from deterministic aggregation of overlapping local predictions. Building on MOIS, we propose Multi-view floor-plan Feed-Forward Network (MvFFN), which jointly predicts coarse camera poses, dense global geometry, and per-wall semantic outputs. A downstream Layout Chaining (LC) post processing pipeline then assembles these predictions into multi-room floor-plans with shared global coordinates, consistent wall identities, and connected room structure. The dense MOIS predictions recover cross-view wall identity and local within-room connectivity more accurately than composed modular baselines, and the assembled layouts substantially outperform prior methods on wall and junction level accuracy. We also introduce the ZInD-CrossView dataset, which augments ZInD with globally unique wall instances and cross-view segmentation labels for this task.
MVVBench: Benchmarking 4D Reasoning in Vision-Language Models
Hyungjin Chung ⋅ Byeongjun Park ⋅ Joonseok Lee ⋅ Hojun Kim ⋅ Jaeho Choi ⋅ Byung-Hoon Kim
Multi-view video understanding requires integrating spatial and temporal evidence across multiple, often non-overlapping camera streams—tracking entities as they transition between viewpoints, aligning events across time, and reasoning about latent 4D continuity rather than any single visible frame. We introduce MVVBench, a benchmark for multi-view video reasoning built from real-world multi-camera datasets and curated to be monocular-ambiguous: each question is unanswerable from any single view alone but becomes uniquely solvable by jointly reasoning over multiple views. MVVBench spans diverse dynamic scenes and probes six capabilities: implicit/explicit attribute identification, implicit/explicit relative distance, relative camera pose, and compositional counting, with human-authored QA and rigorous verification. Beyond benchmarking, we provide an extensive analysis of when and why current vision-language models succeed or fail, characterizing errors due to temporal mis-localization, cross-view identity breaks, and brittle multi-hop reasoning. We then study inference-time elicitation strategies that unlock latent multi-view competence—task-specific chain-of-thought scaffolds and structured cross-view evidence aggregation—yielding substantial gains without retraining. Finally, we present preliminary evidence that reinforcement learning with verifiable rewards can elicit some latent multi-view competence in the base model, pointing to training-time approaches as a promising direction for future work. Together, MVVBench offers a rigorous evaluation of 4D multi-view reasoning and a foundation for future progress toward reliable embodied perception.
Nahual: A Sequence Model for Language and Atoms
Austin Cheng ⋅ Marcel Müller ⋅ Andreas Burger ⋅ Yeonghun Kang ⋅ Tsz W Ko ⋅ Marta Skreta ⋅ Luka Mucko ⋅ Ella Miray Rajaonson ⋅ Cher-Tian Ser ⋅ Jérôme F Gonthier ⋅ Alex Zook ⋅ Varinia Bernales ⋅ Alan Aspuru-Guzik
Generative models play a rapidly growing role in the design and simulation of molecules. The most widely used applications of generative models are in large language models (LLMs), owing to the natural flexibility of language for specifying and answering queries. However, generalist LLMs do not effectively handle large sets of 3D coordinates, whereas specialist generative models of 3D molecules cannot reach the controllability afforded by language. To bridge these capabilities, we propose Nahual, a decoder-only autoregressive diffusion model for any sequence of text and 3D molecules. To enable autoregression on continuous coordinates to scale to long sequences, we propose next-token diffusion, whereby a standard causal transformer backbone is grounded in the clean current sequence while denoising the next sequence. We train Nahual on a curated set of twelve 3D chemistry tasks ranging from microsolvation to molecular conformer search to adsorption on metal surfaces, with associated evaluation metrics and baselines. The result is a single set of end-to-end-trained weights that simultaneously performs all tasks and surpasses state-of-the-art models specialized for crystal structure prediction and property-conditional molecule generation. Our results demonstrate the viability of scaling generalist decoder-only models for native multimodal understanding and generation across language and chemistry.
NARRA-Gym for Evaluating Interactive Narrative Agents
Yue Huang ⋅ Yuchen Ma ⋅ Jiayi Ye ⋅ Wenjie Wang ⋅ Zipeng Ling ⋅ Xingjian Hu ⋅ Yuexing Hao ⋅ Zichen Chen ⋅ Zhangchen Xu ⋅ Yunhong He ⋅ Zhengqing Yuan ⋅ Yujun Zhou ⋅ Kehan Guo ⋅ Chaoran Chen ⋅ Toby Li ⋅ Stefan Feuerriegel ⋅ Xiangliang Zhang
Interactive narrative is becoming a practical frontier for language-model agents, spanning game NPCs, collaborative story production, emotionally responsive companions, and story-grounded interface generation. Yet most current evaluations still rely on static prompts, isolated story outputs, or post-hoc ratings, leaving open whether frontier models can sustain coherent, adaptive, emotionally grounded interaction over time. We introduce NARRA-Gym, an executable evaluation environment that turns a sparse emotional seed into a complete interactive story episode and records the full agent loop: story construction, memory updates, planning, pacing control, and optional artifact generation. The benchmark jointly tests five coupled capabilities---creative story generation, long-context state tracking, character simulation, empathic personalization, and interactive artifact generation---through a fixed LLM-as-judge sweep and a human preference study with user-provided experiences. Across nine generator models and eight benchmark personas, Claude Sonnet 4.6 is the strongest and most robust performer, while Claude Opus 4.6 forms a high-variance second tier and several middle-tier models show different tradeoffs between story quality and user experience. Human rankings recover the same top tier but shift the middle of the ordering, indicating that automated judges are useful for broad screening while human preference remains important for style, tone, and emotional resonance. Finally, failure analysis shows that the key remaining bottleneck is not surface fluency but resistance-sensitive personalization: models can write polished scenes while losing track of what the user is resisting or emotionally unable to accept.
Narrative Paradoxes in LLM Inference: Shaping Reasoning Trajectories toward Misalignment
Haoming Yang ⋅ Ke Ma ⋅ Yangbangyan Jiang ⋅ Ligong Zhang ⋅ Xiaojun Jia ⋅ Longtao Huang ⋅ Yingfei Sun ⋅ Qianqian Xu ⋅ Qingming Huang
Explicit reasoning in large language models (LLMs) is often associated with reliability, yet it can still pose risks when a model infers harmful goals under a benign context. We identify a distinct challenge: ensuring the safety of the \emph{multi-step reasoning process} itself, where locally coherent steps may still lead to globally harmful outcomes. To explore this vulnerability, we propose PDRA (Paradox-driven Reasoning-based Attack), a novel attack method that programs a harmful reasoning trajectory. PDRA first constructs a paradoxical narrative core by pairing a malicious actor (e.g., a terrorist) with a benign task (e.g., writing a safety bulletin). It then embeds this core into a structured analytical framework that forces the model through a mandatory three-stage analysis -- surface purpose, latent intent, and operational reconstruction -- guiding it to produce detailed harmful content while maintaining narrative coherence. Mechanistically, we show PDRA steers the model's internal representations away from refusal-related patterns, suppressing safety-triggering lexicons while activating planning-oriented terms. Empirically, PDRA achieves state-of-the-art attack success rates across diverse LLMs with high efficiency. \textcolor{red}{\textbf{Warning:} This paper may contain sensitive content.}
Narrowing the Collaboration Gap, Probably
Mirah Shi ⋅ Marcel Hussing ⋅ Natalie Collina ⋅ Ira Globus-Harris ⋅ Aaron Roth ⋅ Surbhi Goel
Large language models are increasingly deployed as teams of agents that hold different private information about a shared task. How should such agents concisely communicate in a way that promotes aggregation of decision relevant information? Communicating only proposed actions can discard decision-relevant uncertainty: two agents may prefer the same action for different reasons. We study an alternative interface in which agents communicate calibrated beliefs over actions. We first analyze a simple synthetic setting that captures the structure of multi-round reasoning under partial information. In this setting, we show theoretically and empirically that communicating unbiased probabilities can be strictly more powerful than communicating actions. We then test the same principle in a collaborative maze-solving task with trained transformers and pretrained LLM agents. Our results demonstrate that probability communication outperforms action communication when the exchanged predictions preserve calibrated uncertainty. When communicated probabilities are biased we show that the benefits of probability communication can disappear. We give a post-hoc conversation calibration intervention that consistently improves decision-making by correcting these biases. These results suggest that probabilistic communication is useful for multi-agent collaboration not merely because it transmits more information, but because calibrated uncertainty gives other agents a reliable object to update on.
Natural Language Actor-Critic: Scalable Off-Policy Learning in Language Space
Joey Hong ⋅ Kang Liu ⋅ Zhan Ling ⋅ Jiecao Chen ⋅ Sergey Levine
Large language model (LLM) agents---LLMs that dynamically interact with an environment over long horizons---have become an increasingly important area of research, enabling automation in complex tasks involving tool-use, web browsing, and dialogue with people. In these settings, traditional policy gradient methods may suffer from unstable learning and poor sample-complexity due to poor credit assignment. Meanwhile, actor-critic methods address long horizons by learning a critic that provides more granular feedback and enables off-policy (and potentially offline) learning, but are heavily dependent on the fidelity of the critic. In this paper, we propose \emph{Natural Language Actor-Critic} (NLAC), a novel actor-critic algorithm that trains LLM policies using a generative LLM critic that produces values in natural language space. While natural language values have been leveraged in the past to provide a more flexible and actionable training signal, our work is the first to perform general and scalable value learning without relying on in-context aggregation of information from on-policy rollouts. This means our approach can be trained off-policy without policy gradients, offering a more sample-efficient alternative to existing methods. We present results on a mixture of reasoning, web browsing, and tool-use with dialogue tasks, demonstrating that NLAC shows promise in outperforming existing training approaches and offers a more scalable and stable training paradigm for LLM agents.
Natural Synthesis: Outperforming Reactive Synthesis Tools with Large Reasoning Models
Frederik Schmitt ⋅ Matthias Cosler ⋅ Niklas Metzger ⋅ Julian Siber ⋅ Vladimir Krsmanovic ⋅ Oscar Ghanem ⋅ Bernd Finkbeiner
Reactive synthesis, the problem of automatically constructing a hardware circuit from a logical specification, is a long-standing challenge in formal verification. It is elusive for two reasons: It is algorithmically hard, and writing formal specifications by hand is notoriously difficult. In this paper, we tackle both sides of the problem. For the algorithmic side, we present a neuro-symbolic approach to reactive synthesis that couples large reasoning models with model checkers to iteratively repair a synthesized Verilog implementation via sound symbolic feedback. Our approach solves more benchmarks than the best dedicated tools in the annual synthesis competition and extends to constructing parameterized systems, a problem known to be undecidable. On the specification side, we introduce an autoformalization step that shifts the specification task from temporal logic to natural language by introducing a hand-authored dataset of natural-language specifications for evaluation. We demonstrate performance comparable to that of starting from formal specifications, establishing natural synthesis as a viable end-to-end workflow.
Navigating by Old Maps: The Pitfalls of Static Mechanistic Localization in LLM Post-Training
Hang Chen ⋅ Jiaying Zhu ⋅ Hongyang Chen ⋅ Hongxu Liu ⋅ Xinyu Yang ⋅ Wenya Wang
The "Locate-then-Update" paradigm has become a predominant approach in the post-training of large language models (LLMs), identifying critical components via mechanistic interpretability for targeted parameter updates. However, this paradigm rests on a fundamental yet unverified assumption: can mechanisms derived from current static parameters reliably guide future dynamic parameter updates? To investigate this, we systematically track the structural evolution of Transformer circuits throughout the supervised fine-tuning (SFT) process, revealing the underlying dynamics of task mechanisms. We introduce three novel metrics—Circuit Distance, Circuit Stability, and Circuit Conflict—to analyze circuit evolution across three dimensions: neural migration, semantic stability, and cross-task interference. Our empirical results reveal that circuits inherently exhibit "Free Evolution" during parameter updates. Consequently, static mechanisms extracted from current states inevitably suffer from temporal latency, making them fundamentally inadequate for guiding future states. Moreover, by deconstructing the "illusion of effectiveness" in existing methods, this work underscores the necessity of "foresight" in mechanistic localization and proposes a predictive framework for future research.
Near-Linear Time Generalized Sinkhorn Algorithms for Bounded Genus Graphs
Krzysztof M Choromanski ⋅ Derek Long ⋅ Ananya Parashar ⋅ Dwaipayan Saha
We present GenusSink, a new class of approximate generalized Sinkhorn algorithms with shortest-path-distance costs for bounded genus (e.g. planar) graphs, providing near-linear time: (1) pre-processing, (2) iteration step, (3) final transport plan matrix querying and near-linear memory. Graphs handled by GenusSink include in particular planar graphs and bounded-genus meshes approximating 3D objects. GenusSink addresses total quadratic time complexity of its brute-force counterpart by leveraging separator-based decomposition of graphs, computational geometry techniques, and new results on fast matrix-vector multiplications with generalized distance matrices, using in particular Fourier analysis and low displacement rank theory. It is inspired by recent breakthroughs in graph theory on approximating bounded genus metrics with small treewidth metrics minor. The graph-centric approach enables us to target optimal transport problem with the corresponding distributions defined on the manifolds approximated by weighted graphs and with cost functions given by geodesic distances. We conduct rigorous theoretical analysis of GenusSink, provide practical implementations, leveraging newly introduced in this paper separation graph field integrators (S-GFIs) data structures and present empirical verification. GenusSink provides orders of magnitude more accurate computations than other efficient Sinkhorn algorithms, while still guaranteeing significant computational improvements, as compared to the baseline. As a by-product of the developed methods, we show that GenusSink is $\textbf{numerically equivalent}$ to the brute-force geodesic Sinkhorn algorithm on n-vertex graphs with treewidth $O(\log \log (n))$ (e.g. on trees).
NeSyKC: Neurosymbolic Knowledge Compilation For Lifelong Learning Embodied Agents
Wonje Choi ⋅ Jooyoung Kim ⋅ Sera Choi ⋅ Younguk Song ⋅ Honguk Woo
Lifelong learning requires embodied agents to continuously understand the environment and learn how to use that understanding to guide action, enabling future interactions to yield more informative experiences. Prior language model (LM)-based agents have primarily addressed these requirements in isolation, emphasizing either declarative knowledge acquisition or procedural knowledge reuse, leaving the continual interplay between the two during lifelong learning underexplored. We introduce Neurosymbolic Knowledge Compilation (NeSyKC), a lifelong learning framework that couples declarative knowledge formation with incremental compilation into procedural knowledge through a shared symbolic representation. In NeSyKC, accumulated experience is abstracted into symbolic rules that capture reusable environmental structure, and these rules provide the basis for deriving executable procedures that are incrementally compiled into neural adapters. This creates a feedback loop in which experience is abstracted into symbolic rules and compiled into procedures that guide the LM's reasoning on future tasks, yielding further experience for learning. Across four open-ended environments instantiated from established embodied benchmarks, NeSyKC expands and refines reusable knowledge across tasks, improving task performance and inference efficiency while supporting generalization to unseen tasks. Real-world experiments on two robots further show that the agent continues to learn from post-deployment experience, allowing acquired knowledge to be compiled for reuse in future tasks.
Functional data such as time series, trajectories, and fields are increasingly central in machine learning and are naturally modeled in infinite-dimensional Hilbert spaces. We study Semi-dual Neural Optimal Transport (SNOT) for learning transport maps between probability measures on separable Hilbert spaces. A critical challenge in this setting is the spurious solution problem: a neural map can globally optimize the SNOT objective while failing to recover an optimal transport map. We show that this failure arises from non-regular source measures, which make the inner minimization in the semi-dual objective non-identifiable. We prove that regular source measures, defined via Gaussian null sets, restore inner-minimizer uniqueness and recover the Monge map. For singular sources, we introduce Gaussian smoothing and establish a necessary and sufficient kernel condition for smoothing to restore regularity. We further prove plan-level consistency: as smoothing vanishes, every Wasserstein accumulation point of the induced optimal plans solves the original Kantorovich problem. Experiments on synthetic singular transports and real-world unpaired time-series imputation benchmarks validate the theory, achieving the best performance on several benchmarks and competitive performance on the remaining ones.
NEXUS: Neural Energy Fields for Physically Consistent Contact-Rich 3D Object Dynamics
Qizhen Ying ⋅ Guangming Wang ⋅ Yangchen Pan ⋅ Victor Prisacariu ⋅ Brian Sheil ⋅ Yixiong Jing
Physically consistent 3D object dynamics is a key component for controllable simulation and physics-grounded video generation, especially under contact, deformation, and external forcing. Existing trajectory-based methods enable physical control, but they often model isolated physical effects rather than deriving motion from a unified physical structure. This makes it difficult to compose conservative and non-conservative effects in contact-rich 3D dynamics while retaining explicit control. We present NEXUS, a neural energy-field method for contact-rich 3D object dynamics. NEXUS represents each object as a structural graph and constructs dynamic contact graphs for interactions. Inspired by the Hamiltonian Neural Network (HNN), NEXUS formulates dynamics through scalar energy and dissipation terms rather than direct state or acceleration regression. Conservative effects are composed as additive energy terms over the scene. To handle non-conservative systems beyond standard HNNs, NEXUS learns a dissipation function for impact-induced energy loss. Forces are derived by differentiating the energy-dissipation functions and rolled out with a numerical integrator. NEXUS provides a stable and controllable physics reasoning module for contact-rich 3D object dynamics, improving long-horizon accuracy over representative learned and physics-structured dynamics baselines across controlled rollouts with varying mechanical properties and physical-effect compositions. Our contributions are: (i) an HNN-inspired energy-dissipation dynamics formulation that unifies conservative and non-conservative scene effects; (ii) a graph-based 3D representation for contact-rich object dynamics, with object--object contact evaluated in mixed-contact stress tests; and (iii) a trajectory-guided video generation study showing that physically consistent motion improves downstream physical plausibility while maintaining competitive visual quality.
NLD4CO: Neural Langevin Dynamics for Combinatorial Optimization
Jiale Ma ⋅ Wenzheng Pan ⋅ Binghao Cai ⋅ Xihe Zhang ⋅ Junchi Yan
Langevin Dynamics (LD) provides a principled framework for solving combinatorial optimization problems (COPs) via gradient-guided stochastic search. However, classical LD faces two bottlenecks: 1) the utilization of uniform initialization may affect its convergence speed; 2) when applied to COPs such as routing problems that cannot be naturally formulated as Quadratic Unconstrained Binary Optimization (QUBO), classical LD suffers from slow convergence and inferior performance compared to existing learning-to-construct (L2C) methods, as it requires manually designed discrete proposals for updates. These proposals are typically local and induce abrupt transitions in the solution space, hindering escape from local optima and degrading performance. To address these issues, we propose NLD4CO, a unified framework that synergizes data-driven learning with LD. It introduces two instantiations: explicit-gradient (EG) and implicit-gradient (IG). NLD-EG accelerates sampling efficiency for energy-based problems via neural warm-start initialization. NLD-IG utilizes the corrective direction predicted by a consistency model as an implicit score function to guide the search, making structured, globally coordinated transitions for general COPs. Extensive evaluations on MIS, MWIS, TSP, and CVRP demonstrate that NLD4CO achieves SOTA performance, delivering superior solution quality with competitive or lower runtime.
NORA: Evaluating Grounded Reasonableness in Visual First-person Normative Action Reasoning
Sichao Li ⋅ Sai Ma ⋅ Zhuang Li ⋅ Daniel Kilov ⋅ Secil Yanik Guyot ⋅ Seth Lazar
LLMs and agentic systems are increasingly deployed in social environments, making normative competence critical for safe and appropriate behavior. However, existing approaches either assess normative judgment in text alone or reduce it to choosing among a fixed set of candidate actions. We argue both are insufficient. In practice, agents are never handed a menu of options; they must identify a reasonable action from scratch, grounded in visible facts and supported by inspectable reasons. We introduce NoRA, a visual first-person video benchmark that requires models to generate candidate next actions and justify each through an explicit fact–reason–action support graph. The benchmark comprises 1420 annotated video clips, including HhumanGold-190 and LLMSilver-1230 splits. Each instance is evaluated through action alignment, factual grounding, and support binding, aggregated into a single grounded reasonableness score. We benchmark 12 multimodal systems under direct, deliberative, and structured prompting regimes, finding that current VLMs frequently recover plausible actions and relevant scene facts, but consistently struggle to construct the full reasonable action space and bind selected actions to the correct local support. NoRA makes this gap measurable, shifting the evaluation question from whether a model can pick an action to whether it can justify an appropriate action for the right visible reasons.
NORMA: Norm-Guided Explanation Subgraph Discovery
Xiangyu Fu ⋅ Wei Liu ⋅ Jun Wang ⋅ Yang Qiu ⋅ Yuhua Li ⋅ Ruixuan Li
Discovering explanatory subgraphs is essential for interpreting Graph Neural Networks (GNNs). Many existing and popular explainers optimize by maximizing mutual information (MMI) between selected subgraphs and predicted labels; however, the resulting optimization landscape is often non-smooth and poorly conditioned, leading to suboptimal local optima. To address this issue, we develop a norm-theoretic framework for subgraph analysis. We show that subgraphs containing model-utilized features tend to induce graph embeddings with larger $\ell_2$ norms, as they align more strongly with the supportive subspace shaped by the GNN’s weight matrices, a phenomenon consistently observed across multiple datasets. Based on this insight, we propose NORMA, a parameterized explainer that uses subgraph-embedding $\ell_2$ norms as a stable optimization signal to guide subgraph selection. Experiments on fsix benchmark datasets demonstrate that NORMA effectively alleviates the optimization difficulties of MMI-based methods and achieves superior explanation quality.
Not All Tasks Quantize Equally: Fisher-Guided Quantization for Visual Geometry Transformer
Yipu Zhang ⋅ Jintao Cheng ⋅ Weilun Feng ⋅ Jiehao Luo ⋅ Chuanguang Yang ⋅ Zhulin An ⋅ Yongjun Xu ⋅ Wei Zhang
Feed-forward 3D reconstruction models, represented by Visual Geometry Grounded Transformer (VGGT), jointly predict multiple visual geometry tasks such as depth estimation, camera pose prediction, and point cloud reconstruction in a single forward pass. They have been widely adopted in 3D vision applications, but their billion-scale parameters bring substantial memory and computation overhead, posing challenges for on-device deployment. Post-Training Quantization (PTQ) is an effective technique to reduce this overhead. Existing PTQ methods for feed-forward 3D models mainly focus on handling heavy-tailed activation distributions and constructing diverse calibration datasets. However, we observe that feed-forward 3D models predict multiple geometric attributes through a shared backbone, where different transformer blocks and hidden channels contribute distinctly to each task, resulting in substantially different sensitivities to quantization errors across tasks, blocks, and channels. Consequently, treating all tasks equally over-emphasizes insensitive tasks and causes significant accuracy loss on the sensitive ones. To address this issue, we propose Fisher-Guided Quantization (FGQ) for feed-forward 3D reconstruction models. Specifically, FGQ uses the diagonal Fisher information matrix to quantify the different sensitivities across tasks, blocks, and channels, and incorporates these sensitivities into the Learnable Affine Transformation during calibration to better preserve the channels and blocks most critical to each task. Extensive experiments across camera pose estimation, point map reconstruction, and depth estimation show that FGQ consistently outperforms state-of-the-art quantization baselines on VGGT, achieving up to 39\% relative improvement under the 4-bit quantization.
NSARM: Next-Scale Autoregressive Modeling for Robust Real-World Image Super-Resolution
Xiangtao Kong ⋅ Rongyuan Wu ⋅ Shuaizheng Liu ⋅ Lingchen Sun ⋅ Lei Zhang
Most recent real-world image super-resolution (Real-ISR) methods employ pre-trained text-to-image (T2I) diffusion models to synthesize the high-quality image either from random Gaussian noise or directly from the input low-quality image. These approaches train ControlNet or LoRA modules while keeping the pre-trained model fixed, which often introduces over-enhanced artifacts and hallucinations, suffering from limited robustness to inputs with varying degradations. Recent visual autoregressive (AR) models, such as pre-trained Infinity, can provide strong T2I generation capabilities while offering superior efficiency by using the bitwise next-scale prediction strategy. Building upon next-scale prediction, we introduce a robust Real-ISR framework with a generation pathway control mechanism, namely Next-Scale Autoregressive Modeling (NSARM). Specifically, we train NSARM in two stages: a transformation network is first trained to map the input low-quality image to preliminary scales, followed by an end-to-end full-model fine-tuning. Such a comprehensive fine-tuning enhances the robustness of NSARM in Real-ISR tasks without compromising its generative capability. Extensive quantitative and qualitative evaluations demonstrate that as a pure AR model, NSARM achieves superior visual results over existing Real-ISR methods, with fast inference speed and comparable fidelity. Most importantly, it demonstrates much higher robustness to varying input quality, showing stronger generalization performance.
NyoomFloat12: Accelerating LLM Inference via Lossless 12-bit Weight Compression
Sylvie Liberman ⋅ Xinyu Fang ⋅ Tianyi Zhang ⋅ Tri Dao ⋅ Dan Fu
BFloat16 weights of large language models carry only ${\sim}10.5$ bits of information per 16-bit value, yet exploiting this redundancy on GPUs, where weights bottleneck both per-token bandwidth and concurrent capacity, has not produced wall-clock speedup. Existing variable-length lossless codes are compute-bound on SIMT decode. We present NyoomFloat12 (NF12), a lossless 12-bit fixed-length format for BF16 weights that addresses both bottlenecks. Over $99.7\%$ of trained BF16 weights have the upper four exponent bits set to $\texttt{0111}$ after a per-matrix power-of-two scale; dropping them yields a fixed-length 12-bit format with 16-instruction SIMT-friendly decode. Out-of-range groups reuse the encoded slot to index a per-matrix verbatim BF16 buffer at zero per-group storage overhead. In the bandwidth-bound regime, fused NF12 GEMV achieves up to $1.21\times$ matched-BF16 and $1.35\times$ cuBLAS, $71\%$ of the entropy-bound optimum (a hypothetical implementation reading the joint entropy floor at full HBM bandwidth). In the capacity-bound regime on Qwen3-32B, NF12's smaller footprint opens $4.5\times$ more KV-cache room; with decode at near-peak HBM bandwidth ($3.3\times$ faster than prior lossless decoders), NF12 raises throughput up to $2.67\times$.
OASIS: Online Adaptive Steering for In-Training Safety of LLMs
Yifan Sun ⋅ Qiang Sheng ⋅ Ya Wu ⋅ Zhengjia Wang ⋅ Guang Yang ⋅ Shaofei Wang ⋅ Danding Wang ⋅ Juan Cao
Fine-tuning is essential for adapting Large Language Models to downstream tasks. However, this process can also inadvertently erode the critical safety alignment even when the fine-tuning data appears benign. Prior methods either introduce safety regularizers that may conflict with the primary training objective, or rely on fixed mechanisms that fail to adapt to fine-tuning dynamics. To remedy this, we reframe safety preservation as a training-time intervention and propose Online Adaptive Steering for In-Training Safety (OASIS), enabling continuous online recalibration of the steering direction during fine-tuning. Specifically, OASIS tracks a misalignment direction in activation space, then applies example-adaptive activation steering to absorb misalignment-inducing updates during training. Across three misalignment behaviors and three fine-tuning data regimes, including fully misaligned, mixed, and benign data, OASIS consistently improves safety robustness without sacrificing downstream performance.
Objective-Aligned Amortized Inference for Offline Bayes-Adaptive MDP Model Learning
Toru Hishinuma ⋅ Kei Senda
Learning latent environment representations from offline datasets is an important challenge for Bayes-adaptive Markov decision process model learning. A key difficulty is that offline datasets collected under heterogeneous protocols can differ in state--action visitation even when the underlying dynamics are identical. Visitation-decoupled objectives address this by evaluating transition predictive fit under a protocol-independent reference distribution via a change of measure. We show, however, that the resulting objective-induced invariances need not be inherited by amortized inference. An unconstrained encoder can depend on protocol-induced visitation variations that are invisible to the objective, a structural failure mode we term \emph{amortization mismatch}. To address this issue, we formalize an \emph{objective-aligned} design principle as an objective--inference compatibility condition: the encoder should factor through the objective-induced dataset signature. We instantiate this principle in multiple set-encoder families and, through targeted experiments, provide evidence that objective-aligned encoders improve posterior consistency and can mitigate downstream planning discrepancy across protocols.
Observations Drift, Structures Remain: Structural Pretraining with Time Alignment for Electromagnetic Signals
Wenjin Gui ⋅ Junyu Shen ⋅ Haibo Xu ⋅ Yuchuang Sun ⋅ Luqing Luo ⋅ Wupeng Xie ⋅ Qian Huang ⋅ Fengxiang Wang ⋅ Yangang Sun ⋅ Maosong Sun
Electromagnetic signal modeling is increasingly moving toward self-supervised pretraining on heterogeneous I/Q observations. However, an observed waveform is jointly shaped by signal generation, propagation environments, and acquisition processes, which mix reusable structural regularities with observation-specific variations. Existing methods usually pursue unified pretraining with reconstruction objectives on mixed observations. This heterogeneity creates two key mismatches for reconstruction pretraining: (1) appearance mismatch, induced by propagation-dependent waveform variations; and (2) temporal mismatch, induced by acquisition-dependent sampling coordinates. These mismatches make observation-specific appearances and sampling coordinates easier to reconstruct than stable signal structures, inducing observational shortcuts. We identify this failure mode as observation-biased pretraining. To address this, we propose Structural Pretraining with Time Alignment for heterogeneous EM signals. It mitigates temporal mismatch with RoPE-fs, which maps token positions to sampling-rate-normalized physical-time positions, and reduces reliance on appearance shortcuts by replacing waveform reconstruction with masked prediction of discrete structural units. These units are organized into atomic structural units and emergent units that capture variable-length compositional patterns beyond fixed patch boundaries. Across diverse EM benchmarks, our method improves both full-supervision and few-shot adaptation over existing EM pretraining baselines, showing that the learned representations better preserve structural consistency when observations drift in waveform appearance and temporal scale.
OCTOPUS: Optimized KV Cache for Transformers via Octahedral Parametrization Under optimal Squared error quantization
Mark Boss ⋅ Vikram Voleti ⋅ Simon Donné ⋅ Shimon Vainer
The key-value (KV) cache dominates memory bandwidth and footprint in long-context autoregressive inference. Recent rotation-preconditioned codecs (TurboQuant, PolarQuant) show that a structured random rotation followed by a per-coordinate scalar quantizer matched to an analytically tractable marginal is a near-optimal recipe for KV compression. OCTOPUS advances this paradigm through joint quantization of rotated coordinate triplets. Each triplet's direction is mapped to a square via an octahedral parameterization, and the two resulting coordinates and the triplet norm are Lloyd-Max quantized against implementation-matched marginals. Optimizing the per-triplet squared error gives a strictly non-uniform bit allocation depending only on the total dimensionality of the keys. We find the finite-dimensional quality optimum with sweeps to be constant on every real decoder we test. The codec is data-oblivious, online, and deterministic given a seed. Across text, video, and audio, OCTOPUS matches or beats every prior rotation codec at every reported bit width and metric, with a lead that grows as bits drop for extreme compression. Furthermore, a fused Triton implementation reconstructs keys on the fly without materializing the uncompressed key, so the codec adds no decode-time bandwidth or latency over the existing dequantization.
Offline Inverse Reinforcement Learning with Unified Diffusion Planning
Hongmin Zhao ⋅ Jiyuan Yin ⋅ Kexi Yan ⋅ Qinglai Wei ⋅ Jie ZHANG
Offline Inverse Reinforcement Learning (IRL) aims to recover a reward function and imitate expert behavior solely from offline demonstrations. While recent Offline IRL approaches employ approximate dynamics to mitigate distribution shift, their performance is constrained by the unstable minimax framework as well as simplified representation and utilization. To address these issues, we propose Offline Inverse Reinforcement Learning with Diffusion Planner (OIDP), which leverages diffusion models to achieve stable Offline IRL. OIDP formulates a model-based conservative Q-optimization through inverse soft Q-learning, which is provably concave in Q-space with a controllable performance gap. Accordingly, we introduce a unified diffusion planner that models dynamics, uncertainty estimator, and policy, it captures multimodal distributions, directly tightens the objective gap bound, and at execution performs trajectory-guided policy sampling using the internalized dynamics knowledge. Experimental results on D4RL benchmarks demonstrate that OIDP outperforms state-of-the-art offline approaches.
OHATP: Graph Anomaly Detection with Orthogonal-Hyperspherical Augmentation and Topology Perception
yihang qiu ⋅ Yi Zhang ⋅ Di Xiong ⋅ Ge Gao ⋅ Shuo Chen
Contrastive Learning (CL) has been widely used for Graph Anomaly Detection (GAD). However, existing augmentation techniques in such CL frameworks usually rely on random masking or structural alterations, which obliterates the irregularities that define anomalies and leads to anomaly distortion. Even improved methods using multi-scale sampling usually suffer from blind contextual extraction, which risks disrupting critical topological structures or introducing irrelevant information in the learned representations, thereby hardly capturing complex anomalies. In this paper, we propose Orthogonal-Hyperspherical Augmentation and Topology Perception (OHATP), a novel method that leverages the synergies at both the feature and structure levels to mitigate anomaly distortion and blind sampling defects. Specifically, during feature-level augmentation, we adapt projections and noise based on node message-passing influence to strengthen feature diversity, while utilizing projection orthogonality and hyperspherical noise constraints to preserve feature discriminability. To complement this feature-level augmentation, we design a structure-level topology perception module that leverages the high-quality representations learned from augmentation to screen abnormal dense substructures. Extensive experiments on benchmark datasets demonstrate that OHATP substantially outperforms state-of-the-art methods, achieving improvements of over 6\% in AUROC and 29\% in AUPRC. Code is available at https://anonymous.4open.science/r/OHATP-83D7.
OmniEgoCap: Camera-Agnostic Sequence-Level Egocentric Motion Reconstruction
Kyungwon Cho ⋅ Jeonghyeon Na ⋅ Hanbyul Joo
Commercial egocentric devices capture human behavior in everyday settings, yet full-body motion reconstruction must generalize across diverse cameras and mountings. This is challenging because head trajectories under-determine body motion, while hand observations are intermittent and camera-dependent: a hand may disappear simply because it has left the camera's field of view. Existing methods treat such absences as missing data, relying on fixed visibility regimes or post-hoc optimization, and fail to generalize across devices.We present \modelname, a sequence-level diffusion framework for device-agnostic egocentric full-body motion reconstruction. Our key insight is that egocentric videos contain sequence-level evidence about both the wearer and the camera-induced visibility structure, enabling \modelname to infer consistent body shape and coherent motion from sparse head and hand cues. To prevent overfitting to a single camera setup, we further introduce geometry-aware visibility augmentation, which synthesizes visibility patterns from realistic variations in camera geometry. We also present OmniEgoDB, the first mocap benchmark with ground-truth motion captured across multiple consumer egocentric devices. Experiments on synthetic, real-device, and in-the-wild settings demonstrate superior motion reconstruction, robust cross-device generalization, and coherent in-the-wild results.
OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering
Yiduo Jia ⋅ Muzhi Zhu ⋅ Hao Zhong ⋅ Mingyu Liu ⋅ Yuling Xi ⋅ Hao Chen ⋅ Bin Qin ⋅ Yongjie Yang ⋅ Zhenbo Luo ⋅ Chunhua Shen
Omni-modal reasoning is pivotal to realizing artificial general intelligence, yet its advancement is critically constrained by the limited availability of large-scale, human-annotated data for complex reasoning. Motivated by this, we propose OmniJigsaw, a self-supervised reinforcement learning framework built upon a temporal reordering proxy task. Centered on the chronological reconstruction of shuffled audio-visual clips, we successively investigate three modality orchestration strategies: (i) Joint Modality Integration (JMI), which retains the complete visual and auditory streams; (ii) Sample-level Modality Selection (SMS), which selects the dominant modality through a global decision mechanism; and (iii) Clip-level Modality Masking (CMM), which adaptively masks modalities at the clip granularity. Our analysis reveals a ``bi-modal shortcut phenomenon'' in JMI and demonstrates that fine-grained CMM mitigates this issue while outperforming SMS. Incorporating lightweight puzzle-quality curation and verifiable rewards, OmniJigsaw yields substantial gains across 15 video, audio, and omni-modal benchmarks, with CMM achieving the strongest overall performance, validating its effectiveness as a self-supervised paradigm for omni-modal learning.
OmniSelect: Dynamic Modality-Aware Token Compression for Efficient Omni-modal Large Language Models
Morunliu Yang ⋅ Ruotao Xu ⋅ Le Li ⋅ Yue Wang ⋅ Jianxin Zhang ⋅ Siwei Feng ⋅ Peifeng Li ⋅ Yihang Lou ⋅ Juntao Li
Omnimodal large language models (OmniLLMs) have recently gained increasing attention for unified audio-video understanding. However, processing long multimodal token sequences introduces substantial computational overhead, making efficient token compression crucial. Existing methods typically rely on fixed, modality-specific guidance, which fails to account for the varying importance of modalities across different queries. To address this limitation, we propose $\textbf{OmniSelect}$, a training-free, modality-adaptive token pruning framework that dynamically selects appropriate compression strategies for multimodal inputs. Specifically, we leverage a lightweight AudioCLIP model to estimate cross-modal relevance and categorize each input into three pruning regimes: Audio-Centric, Video-Centric, and Uniform pruning. Based on these relevance scores, OmniSelect further performs fine-grained token pruning within each temporal group, adaptively allocating pruning ratios to preserve informative tokens across modalities. By explicitly modeling modality preference and enabling dynamic strategy selection, OmniSelect effectively avoids the pitfalls of one-size-fits-all compression. Extensive experiments demonstrate that our method achieves efficient multimodal token reduction while maintaining strong performance, without requiring any additional training.
OmniSimulator: Aligning Small Language Models for Authentic Heterogeneous Behavior Modeling
Jiawei Chen ⋅ Ruoxi Xu ⋅ Boxi Cao ⋅ Zhuoqun Li ⋅ Ruotong Pan ⋅ Zhang yunfei ⋅ Zhirui Yang ⋅ Jingyan Chen ⋅ Tingting Gao ⋅ Han Li ⋅ Yaojie Lu ⋅ Xianpei Han ⋅ Le Sun ⋅ Xiangyu Wu ⋅ Hongyu Lin
Large language models (LLMs) have shown great potential as general-purpose user simulators for interactive systems. Prior studies have demonstrated that training on high-quality data can substantially improve model reasoning and downstream task performance. However, the upper bound of user simulation is still constrained by the model's reasoning paradigm: standard chain-of-thought often degenerates into shallow pattern matching and fails to capture the latent mental process underlying real user behavior. In this paper, we propose OmniSimulator, a novel training framework designed to improve the predictive ability of small LLMs for realistic user simulation on short-video platform environments. Specifically, OmniSimulator models human internal deliberation through a structured reasoning schema with five dimensions, carefully selects data from real user-behavior scenarios, and introduces a teacher model to generate high-quality reasoning traces under the proposed schema, boosting small models to learn the latent logic of human decision-making during the CoT process, rather than relying on superficial behavioral correlations. Based on this process, we construct 9,881 high-quality training instances from three months of historical interaction trajectories of 200 real users, averaging about 50 distilled examples per user. Experimental results show that OmniSimulator improves performance over the baseline by 38.93\%, surpasses strong frontier models such as Claude-Sonnet-4.5, and significantly reduces inference cost. These results suggest that learning structured human-like internal reasoning is a key step toward scalable and high-fidelity LLM-based user simulation.
OmniSpace: Efficient Geometry Awareness for Autonomous Vehicles MLLMs
Anh Hao Vo ⋅ Phu Loc Nguyen ⋅ Khoa Vo ⋅ Sieu Tran ⋅ Duc Nguyen ⋅ Ngo X Cuong ⋅ Nghi Bui ⋅ Anh Nguyen ⋅ Duy M. H. Nguyen ⋅ Ngan Le
Multimodal Large Language Models (MLLMs) have achieved remarkable performance on 2D visual tasks, yet enhancing their spatial intelligence for real-world applications such as Autonomous Vehicles (AV) remains an open challenge. Existing geometry-aware MLLMs typically rely on auxiliary 3D models at inference time, introducing pipeline complexity and the risk of cascading failures. In this paper, we present OmniSpace, a simple yet effective plug-and-play paradigm for geometry-aware spatial reasoning from purely 2D observations. Motivated by our finding that current MLLMs are bottlenecked by weak cross-view correspondence and depth estimation, OmniSpace introduces a Camera Pose Injector, a Multi-view Epipolar Attention module, and a 3D Geometric Distillation objective that jointly address these two limitations by transferring geometric knowledge into the model. Extensive experiments show that OmniSpace surpasses existing methods on planning benchmarks (nuScenes, Bench2Drive), risk detection (nuInstruct), language (Omnidrive), and generalization (DriveBench).
On Data Engineering for Scaling LLM Terminal Capabilities
Renjie Pi ⋅ Grace Lam ⋅ Mohammad Shoeybi ⋅ Pooya Jannaty ⋅ Bryan Catanzaro ⋅ Wei Ping
Despite rapid recent progress in the terminal capabilities of large language models, the training data strategies behind state-of-the-art terminal agents remain largely undisclosed. We address this gap through a systematic study of data engineering practices for terminal agents, making two key contributions: (1) Terminal-Task-Gen, a lightweight synthetic task generation pipeline that supports seed-based and skill-based task construction, and (2) a comprehensive analysis of data and training strategies, including filtering, curriculum learning, long context training, and scaling behavior. Our pipeline yields Terminal-Corpus, a large-scale open-source dataset for terminal tasks. Using this dataset, we train Nemotron-Terminal, a family of models initialized from Qwen3 (8B, 14B, 32B) that achieve substantial gains on Terminal-Bench 2.0: Nemotron-Terminal-8B improves from 2.5% to 13.0%, Nemotron-Terminal-14B improves from 4.0% to 20.2%, and Nemotron-Terminal-32B improves from 3.4% to 27.4%, matching the performance of significantly larger models. We release Nemotron-Terminal checkpoints and Terminal-Corpus through the Nemotron-Terminal Hugging Face collection.
One More Time: Revisiting Neural Quantum States from a Reinforcement Learning Perspective
Juan A Duque ⋅ Sergio García Heredia ⋅ Vinicius Hernandes ⋅ Eliska Greplova ⋅ Thomas Spriggs ⋅ Aaron Courville ⋅ Anna Dawid
Neural quantum states (NQS) provides a flexible and scalable framework for approximating quantum many-body wavefunctions. Among NQS parameterizations, autoregressive models are especially attractive because they enable exact, independent sampling from the Born distribution, avoiding the autocorrelation and mixing issues of Markov-chain methods. Yet their optimization remains comparatively underexplored: Adam is a scalable method but ignores function space geometry, while stochastic reconfiguration is principled but costly and numerically fragile in large models. To address this gap, we show that variational energy minimization can be viewed as an advantage policy-gradient problem over the Born distribution, motivating trust-region optimization for NQS training. We introduce Proximal Wavefunction Optimization (PWO), a trust-region algorithm that clips probability-ratio changes in the amplitude channel and wrapped phase increments in the phase channel. PWO avoids explicit matrix inversion, reuses samples across inner updates, and preserves the scalability of first-order optimization. Across Ising, Heisenberg, and frustrated $J_1$--$J_2$ spin chains, PWO improves stability and wall-clock convergence over Adam, minSR, and SPRING. Finally, we fine-tune a $1.5$B-parameter RWKV-7 model as a neural quantum state, demonstrating NQS optimization at a scale over three orders of magnitude beyond prior work.
OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation
Jinghui Lu ⋅ Jiayi Guan ⋅ Zhijian Huang ⋅ Jinlong Li ⋅ Guang Li ⋅ Lingdong Kong ⋅ Yingyan Li ⋅ Han Wang ⋅ Shaoqing Xu ⋅ Yuechen Luo ⋅ Fang Li ⋅ Chenxu Dang ⋅ Junli Wang ⋅ Tao Xu ⋅ jing wu ⋅ Jianhua Wu ⋅ Xiaoshuai Hao ⋅ Wen Zhang ⋅ Tianyi Jiang ⋅ Lingfeng Zhang ⋅ Lei Zhou ⋅ Yingbo Tang ⋅ Jie Wang ⋅ Yinfeng Gao ⋅ Xi zhou Bu ⋅ Haochen Tian ⋅ Yihang Qiu ⋅ Feiyang Jia ⋅ Lin Liu ⋅ Yigu Ge ⋅ Hanbing Li ⋅ Jiahong Chen ⋅ Zihui Li ⋅ Shen Yuannan ⋅ Jianwei Cui ⋅ Hongwei Xie ⋅ Bing Wang ⋅ Haiyang Sun ⋅ Jingwei Zhao ⋅ JiaHui Huang ⋅ Pei Liu ⋅ zeyu zhu ⋅ Yuncheng JIANG ⋅ Zibin Guo ⋅ Hanchao Leng ⋅ Chuhong Gong ⋅ Kun Ma ⋅ Guang Chen ⋅ Kuiyuan Yang ⋅ Hangjun Ye ⋅ Long Chen
Chain-of-Thought (CoT) reasoning has become a powerful driver of trajectory prediction in VLA-based autonomous driving, yet its autoregressive nature imposes latency prohibitive for real-time deployment. Latent CoT methods attempt to close this gap by compressing reasoning into continuous hidden states, but consistently fall short of their explicit counterparts. We argue that this is because purely linguistic latent representations compress a symbolic abstraction of the world rather than the causal dynamics that govern driving. We present OneVL (One-step latent reasoning and planning with Vision-Language explanations), a unified VLA and world model framework that routes reasoning through compact latent tokens supervised by dual auxiliary decoders. It comprises a language decoder that reconstructs text CoT and a visual world model decoder that predicts future-frame tokens, forcing the latent space to internalize causal scene dynamics. A three-stage training pipeline progressively aligns these latents with trajectory, language, and visual objectives. During inference, the decoders are discarded, and all latent tokens are prefilled in a single parallel pass, matching answer-only prediction speed. Across four benchmarks, OneVL becomes the first latent CoT method to surpass explicit CoT, delivering state-of-the-art accuracy at answer-only latency. Code will be publicly available.
Online Causal Configuration for Networked Systems via Doubly Robust Steady-State Learning
Yuli Liu ⋅ zhiheng zhang
Large-scale networked systems are increasingly controlled through a small number of global configuration variables while their local controllers remain black-box, adaptive, and mutually interfering. The operator observes only post-burn-in telemetry, so neither independent-unit causal estimators nor immediate-feedback bandits directly apply. We propose \emph{CAUSAL-UCB}, a causal-optimistic framework for online steady-state configuration. The framework turns raw network operation into a linked pipeline: an exposure interface maps neighborhood interference into an auditable context; measurement-level exploration creates overlap while the target remains the no-exploration steady-state value $\Winf(z;0)$; a clipped doubly robust evaluator converts exploratory logs into value certificates; and an OFU rule selects the next configuration. We prove a unified high-probability deviation bound that decomposes statistical fluctuation, nuisance product error, clipping bias, exposure approximation, cross-configuration mismatch, context shift, and residual dependence. We also give an identification--deployment frontier showing that reliable counterfactual identification and low deployment cost cannot be optimized independently. Controlled semi-realistic replay and real-data-driven semi-synthetic replay on METR-LA and FlockLab validate the resulting benefit--reliability trade-off across regret, optimality gap, ATE error, confidence width, deployment gap, sensitivity, ablation, and frontier visualizations.
Online Differentially Private Consistent Clustering
Edith Cohen ⋅ Vadym Doroshenko ⋅ Badih Ghazi ⋅ Pritish Kamath ⋅ Alexander Knop ⋅ Ravi Kumar ⋅ Ethan Leeman ⋅ Pasin Manurangsi ⋅ Adam Sealfon ⋅ Marika Swanberg
We study differentially private (DP) $k$-means and $k$-median clustering in the online streaming setting. In this model, points arrive sequentially, and at each time step, we need to output a set of $k$ centers that optimizes the clustering objective for all points seen so far. We give a generic reduction that transforms the (sensitive) input stream into a private stream, which is a *semi-coreset* of the input stream. This implies that any (non-private) online clustering algorithm, run as a post-processing step, can achieve good utility for the original clustering objective. Our algorithm matches or improves upon the approximation ratio, space usage, and running time of existing algorithms (Dupr{\'{e}} la Tour, 2024; Epasto et al., 2026). A key aspect of our reduction is that it inherits desirable properties of the underlying non-private clustering algorithm, such as *consistency* (Lattanzi and Vassilvitskii, 2017)---a property not satisfied by previous DP algorithms.
Online A/B tests are the dominant tool for evaluating new LLM variants in deployment: each user query arrives sequentially, must be routed to one variant before its outcome is observed, and is then scored against the alternative. In practice, these tests almost always use complete randomization. Yet for many queries, simple signals such as predicted difficulty, expected reward, or embedding-based scores are already available, and these are predictive of the outcome. Using them to guide assignment can substantially reduce the variance of the estimated treatment effect. We propose a new dyadic design that uses such covariates to balance treatment and control online across the full covariate distribution, with no hyperparameters to tune. Under stochastic arrivals, our design attains a covariate matching discrepancy of $O(\log^{3} T)$ in one dimension, which is a significant improvement over the $O(T^{1/4})$ rate of existing online stratified designs, and matches the offline minimax rate up to logarithmic factors in higher dimensions, closing a gap left open by prior work. Experiments on a sequential A/B test simulator built from AlpacaEval, in which 805 prompts are scored under two real LLM variants, confirm the theory: our design achieves a covariate matching discrepancy roughly $6\times$ smaller than complete randomization and consistently outperforms the best-tuned online stratified baseline. The method extends to multi-dimensional covariates and to general sequential evaluation pipelines.
Online Fair Division Meets Reordering Buffers
Georgios Amanatidis ⋅ Giulio Giaconi ⋅ Evangelos Markakis ⋅ Nicos Protopapas
We study the _online fair division_ of indivisible _mixed manna_ among agents with additive valuation functions. Under the standard online model, at each time step an indivisible item arrives; each agent may assign it a positive, negative, or zero value, and it must be irrevocably allocated, before the arrival of the next item. At the same time, we also wish to maintain some fairness guarantee, and in this work we focus on _envy-freeness_ (EF) and one of its most prominent relaxations, _envy-freeness up to one item_ (EF1). Given the strong negative and the scarce positive results for this problem without additional assumptions, we augment our algorithms with _buffers_ that can store and rearrange a limited number of items. This setting interpolates naturally between the fully online case (no buffer) and the fully offline case (a buffer large enough to hold all items). We show that algorithms equipped with reasonably sized buffers can achieve strong guarantees for personalized $k$-value instances, i.e., instances in which each agent assigns at most $k$ distinct values to items. In particular, we construct allocations that are EF1 at every time step and EF at most time steps, using a buffer of size linear in $k$ and in the number of agents. Our approach relies on novel combinatorial arguments and on constructing a sequence of envy-free matchings that allocates most items. Finally, we extend our results to general additive valuation functions, with a dependence on the largest per-agent ratio between two values of the same sign, and we also identify limitations of our approach via impossibility results on the use of buffers with smaller size.
Online Quantile Omniprediction for Proper Losses
Hyunsuk Kim ⋅ Isaac Gibbs ⋅ Michael Mahoney ⋅ Ryan Tibshirani
The task of constructing a predictor that is simultaneously accurate under multiple evaluation metrics has been popularized recently under the name omniprediction. Prior work in this area has largely focused on binary prediction tasks, and existing results for multiclass and continuous outcomes typically suffer from slow error rates. In this paper, we consider forecasts expressed as a set of quantiles over a discrete but otherwise arbitrary set of probability levels. We develop sample-efficient algorithms for constructing forecasts that are simultaneously accurate across all proper losses that satisfy mild regularity conditions, where proper losses are those minimized by predicting the true quantiles. Our algorithms are online and can operate in a dynamic environment while simultaneously taking multiple losses into account, and their forecasts offer the potential to provide diverse end-users with rich information for evaluating risk and informing a range of practical decision-making tasks. As an example, we demonstrate how our procedure can be used to ensemble quantile forecasts of COVID-19 hospitalizations made by experts during the pandemic.
On-policy training plays an important role in modern machine learning, especially in reinforcement learning (RL) for language models. However, quantifying how training data influences the policy's behavior is difficult in on-policy settings: one must account for the cascading effect by which a training event affects its immediate policy update, which in turn affects future action samples, which further affects future policy updates, and so on. Attribution methods developed for off-policy and supervised learning (e.g. influence functions) do not account for such \emph{policy-data interaction}, and can therefore fail in on-policy settings. In this work, we formalize the problem of attributing behaviors to individual training events in the on-policy setting. We derive an influence measure that accounts for policy-data interaction while also being tractable to compute in practical training setups (e.g. LM finetuning) and for realistic optimizers (e.g. AdamW) at a cost that scales linearly in the total number of training events. Empirically, our influence measure accurately predicts the true effect of ablating specific events during RL, while traditional methods fail to do so. Further, we show that this influence accurately identifies helpful and harmful RL training events, allowing us to improve performance via influence-weighted retraining. Overall, we show that accounting for policy-data interaction is both necessary and tractable, allowing for principled training intervention in modern RL pipelines.
On the Feasibility of Identity Manipulation for Diffusion-Based Face Privacy Preservation
Xuemei Jia ⋅ JIAWEI DU ⋅ Jiawei Liu ⋅ xin zhang ⋅ Jun Chen ⋅ Zheng Wang ⋅ Joey Tianyi Zhou
Face privacy preservation aims to suppress identity-related information while maintaining realistic facial appearance, for which diffusion models have recently emerged as a powerful paradigm. However, identity manipulation under diffusion-based generation exhibits inconsistent behavior during denoising, where certain perturbations are progressively attenuated while others are amplified into unstable distortions. This phenomenon suggests that identity manipulation is inherently constrained by the underlying diffusion denoising dynamics. In this work, we study the feasibility of identity manipulation under such dynamics and show that effective perturbations exhibit a practical operating regime along the denoising trajectory. Building on this insight, we propose a feasibility-aware diffusion-based identity manipulation framework. The framework anchors latent updates to a reference diffusion state and regularizes optimization toward a diffusion-compatible neighborhood, enabling stable identity suppression without compromising visual fidelity. Extensive experiments demonstrate strong robustness and competitive image quality over existing diffusion-based face privacy methods.
On the Information Loss of Multi-Token Prediction: Origin and Solution
Tangyu Jiang ⋅ Zhanke Zhou ⋅ Haodi Wang ⋅ Yiu-ming Cheung ⋅ Xiaohua Jia ⋅ Bo Han
Multi-token prediction (MTP) accelerates autoregressive decoding through blockwise prediction, but remains vulnerable to error accumulation across sequentially generated blocks. Existing work primarily focuses on improving MTP’s acceleration ratio while largely overlooking the compounding errors that accumulate during block-wise generation. In this work, we study MTP from an information-theoretic perspective and identify the strength of inter-block dependencies as a key factor governing its performance. Building on the cumulative information loss, we theoretically quantify the gap of block-sequence representation. Motivated by the analysis, we propose Dual-MTP, a novel MTP training framework that explicitly models inter-block dependencies from two complementary perspectives, thus reducing the block-sequence representation gap across decoding steps. Specifically, Dual-MTP introduces a dual-perspective training objective with random masking, where each block is jointly trained for forward prediction and masked context reconstruction, encouraging robust block-level representations. Experiments across five model backbones and six tasks spanning math reasoning and GLUE understanding show that Dual-MTP consistently outperforms standard MTP on both quality and efficiency, attaining the highest average accuracy on every backbone with gains of +0.3 to +3.0 points and a 1.4× to 1.9× wall-clock speedup over standard autoregressive decoding.
On the Instability and Stabilization of Blockwise Muon
Yuanshi Liu ⋅ Weicheng Lin ⋅ Boyuan Jiang ⋅ Xin Tao ⋅ Pengfei Wan ⋅ Cong Fang
Muon has become a competitive optimizer for large-scale neural network training through matrix-level orthogonalized updates, but its full-matrix normalization can be expensive. Blockwise Muon is an appealing low-cost approximation: it normalizes matrix shards or small blocks independently, reducing computation and communication overhead while often achieving convergence close to full Muon. However, in large-step-size and large-model regimes, Blockwise Muon can suffer from degraded convergence. We attribute this to cross-block amplification: independent per-block orthogonalizations can collectively amplify updates along certain directions, particularly those governing the effective step size. Motivated by this viewpoint, we accordingly propose \textsc{ACME}: \textbf{A}mplification-\textbf{C}orrected \textbf{M}uon for \textbf{E}ffective step-size recovery. \textsc{ACME} applies a rotational correction so that the directions controlling the effective step size are no longer affected by cross-block amplification. Experiments on language-model pretraining show that, across a range of training settings, \textsc{ACME} achieves convergence nearly indistinguishable from full Muon while retaining the efficiency advantages of Blockwise Muon.
On the Nature of Attention Sink that Shapes Decoding Strategy in Omni-LLMs
Suho Yoo ⋅ Youngjoon Jang ⋅ Joon Son Chung
The goal of this paper is to strengthen the reasoning of Omnimodal Large Language Models (Omni-LLMs) at inference time, without additional training. These models jointly process video, audio, and text, and given the large number of tokens they consume, how attention is routed across them is central to their behaviour. We focus specifically on attention sinks, tokens that absorb a disproportionate share of attention mass regardless of their semantic content, to understand how this routing unfolds. To this end, we conduct a systematic analysis of sink behaviour in Omni-LLMs. Our analysis yields two key findings: (i) high sink attention does not solely indicate head redundancy, suggesting that sink value representations play additional functional roles; (ii) the sink value vector acts as a shared bias added to every token's output, serving as a global signal that organises the representation as a whole. Building on this, we propose OutRo, which correspondingly aligns non-sink token representations with the sink in feature space, and relaxes the causal mask for sink tokens at an early layer to sharpen this bias before the rest of decoding proceeds. This design enhances the reasoning process without requiring additional forward passes or access to attention maps. Based on extensive experiments, OutRo consistently improves performance on seven video QA benchmarks and demonstrates strong generalisation, while incurring only a 1.1\x decoding overhead.
On the Nonlinearity of Learning Rate Scaling for LLM Training
ZAIWEN YANG ⋅ Huaqing Zhang ⋅ Jing Xu ⋅ Jingzhao Zhang
Learning-rate transfer is critical for reducing the cost of training large language models: instead of sweeping learning rates at target scale, practitioners extrapolate from smaller runs. Existing approaches often assume that the optimal learning rate follows a log-linear scaling law in data scale and model size. We carefully examine and evaluate this scaling law. In our empirical study of GPT-2--style models from 22M to 407M parameters trained on 5B to 100B tokens, the optimal learning rate develops upward curvature at larger scales, leading to inaccurate extrapolation. We find that this curvature largely disappears when learning rates are replaced by effective learning rate (the step size in normalized weight space), and when data $D$ extrapolation is used instead of model size $N$ extrapolation. Next, we explain nonlinearity in scaling: weight-norm converges to equilibrium slower when optimal learning is small, requiring a larger step size to reduce the transient phase. Experiments with AdamH, which directly controls the effective learning rate, further support this explanation.
On the Recoverability of Causal Relations from Bulk Gene Expression Data
Gongxu Luo ⋅ Boyang Sun ⋅ Kun Zhang
Bulk gene expression profiling, which aggregates pooled RNA across cells within a biological sample, remains important in the single-cell era because it is typically less noisy, more sensitive, and more cost-effective than single-cell assays. Accordingly, a growing body of computational methods seeks to recover causal relations among genes from bulk expression data. However, aggregation is a lossy, non-invertible coarsening of the underlying cellular system, and it remains unclear whether and under what conditions causal relations are recoverable from aggregated bulk gene expression data. To answer this, we formalize recoverability under aggregation through two notions of consistency: functional-form consistency and conditional-independence consistency. We then derive necessary and sufficient conditions for recoverability, showing that these properties are preserved only under linear aggregations (e.g., sum/mean) coupled with affine structural equations. To assess the practical plausibility of these conditions, analyses of four bulk and four single-cell gene expression datasets further reveal that the estimated pairwise regulatory functions among genes deviate from linearity in both data types, providing limited empirical support for the linearity assumptions required for recoverability. Together, these results caution against recovering causal relations from aggregated bulk expression data without strong additional assumptions.
On the Sparsity of Direct Preference Optimization: Weight Disentanglement in the NTK Regime
Kensuke Sasaki ⋅ Issei Sato
Direct Preference Optimization (DPO) has been widely adopted for aligning language models with human preferences. Recent empirical work has revealed that DPO induces remarkably sparse parameter updates, yet no theoretical explanation for this phenomenon has been offered. We analyze DPO's optimization dynamics under a neural tangent kernel linearization and show that the logistic structure of the DPO loss gives rise to an implicit maximum-margin bias. Building on this observation, we introduce the concept of dual sparsity: the DPO update direction is a sparse linear combination (data-level sparsity, driven by support vectors) of sparse gradient feature vectors (parameter-level sparsity, arising from gradient cancellation between preferred and dispreferred responses). We formalize this structure in our main theorem and present a framework relating it to weight disentanglement---a condition favorable for model merging. Experiments on GPT-2 Small and Medium with the Anthropic HH-RLHF and UltraFeedback datasets confirm both sources of sparsity and show that DPO reduces cross-task interference by up to $35 \times$ compared to supervised fine-tuning.
We study seamless world generation, where a model must update a persistent visual world according to online user operations rather than full next-scene descriptions. To make this task well-defined, we introduce an operation-centric taxonomy that represents each scene with a global-local schema and each boundary with a typed operator program specifying change and preservation. This taxonomy enables SeamlessWorld, a synthetic data and benchmark pipeline with 8 operation families, 23 atomic operators, dual prompt surfaces, and schema-derived evaluation probes. We then develop OopsWorld, a causal video generator that learns from SeamlessWorld using operation-conditioned decoupled distillation: the student is driven by operation prompts and visual history, while scene-conditioned teachers provide target-state supervision. We further apply dynamics curriculum learning to progressively compose simple edits into high-dynamic world transitions. OopsWorld substantially improves dynamic responsiveness and long-video quality, achieving 91.67 VBench Dynamics and 80.42 Long-Quality with 100K samples. A systematic explicit--implicit study further reveals a persistent operation-prompting gap in current systems, especially for preserving visual anchors, calling for operation-driven world models beyond next-scene prompt switching.
Open-Ended Scientific Discovery and the Social Dynamics of Evolving Agent Networks
Tennison Liu ⋅ Silas Ruhrberg Estévez ⋅ Rob Davis ⋅ Ryan M Sheridan ⋅ David Bentley ⋅ Mihaela van der Schaar
Scientific discovery is fundamentally open-ended, characterized by epistemic uncertainty, and requiring endogenous goal-setting and recursive knowledge accumulation. While current agentic systems have demonstrated success in addressing goal-based discovery tasks, they struggle to sustain discovery in open-ended settings. In this work, we posit that open-ended discovery is an emergent network-level property of a socially driven system. We introduce $\texttt{ASCollab}$, which translates four key algorithmic conditions: (1) population heterogeneity in research behaviors, (2) endogenous interaction networks where collaborations and attention routing emerge organically, and (3) socially driven peer evaluations, which are enabled by (4) a global, shared memory that reflects evolving research states. Specifically, social memory is implemented to capture time-varying expertise and reputational signals, and shifting attention paid to historical artifacts produced by the system. Through experiments on large-scale discovery problems in genomics, cell biology, and epidemiology, we demonstrate that $\texttt{ASCollab}$ produces discoveries judged by domain experts to be both sound and significant. Crucially, the system independently recovers real-world scientific findings. Furthermore, systematic investigations show that social dynamics fostered by each of the four conditions are vital, where heterogeneous agents operating within self-organizing networks significantly outperform both fixed-workflow baselines and homogeneous networks.
OpenSanctions Pairs: A Large-Scale Dataset for Pairwise Entity Matching
Chandler Smith ⋅ Magnus Sesodia ⋅ Friedrich Lindenberg ⋅ Christian Schroeder de Witt
We release OpenSanctions Pairs, the first large-scale public benchmark for entity matching on sanctions and OSINT data. The dataset includes 755,540 expert-labeled pairs over 1 million entities, aggregated from 293 source datasets across 45 jurisdictions. It captures real-world diversity in compliance data, spanning multiple languages and writing systems (e.g., Latin, Cyrillic, Arabic), inconsistent structure, and time-varying provenance, and is substantially more heterogeneous than prior entity matching benchmarks. As baselines, we evaluate the production rule-based matcher (nomenklatura RegressionV1) alongside open- and closed-source LLMs in both zero- and few-shot settings, each tested with and without MIPROv2 prompt optimization to control for prompt sensitivity. The rule-based baseline reaches 91.3\% F1; GPT-4o achieves the best result at 99.0\% F1, and a locally deployable open-source model (DeepSeek-R1-Distill-Qwen-14B) achieves 98.2\% F1. The rule-based baseline and LLMs fail in complementary ways: rules over-match, while LLMs struggle with cross-script transliteration. These results suggest that pairwise matching performance is approaching a practical ceiling and shift attention toward pipeline components such as blocking, clustering, and uncertainty-aware review.
Open Vocabulary Domain Unlearning
Sumanth V Udupa ⋅ Mehrtash Harandi ⋅ Yadan Luo ⋅ Mahsa Baktashmotlagh
Vision-Language Models (VLMs) exhibit remarkable zero-shot generalization, yet they often encode unwanted or hazardous stylistic domains such as idealized textbook diagrams in medical AI or cartoon vehicles in autonomous driving. Approximate Domain Unlearning (ADU) aims to selectively erase a model's recognition of a target visual domain while preserving accuracy on the remaining domains. However, existing ADU methods operate under a flawed closed-vocabulary assumption: they evaluate unlearning solely on the specific object classes seen during the unlearning fine-tuning phase. Consequently, these methods do not unlearn the domain itself; they merely overfit to seen class-domain pairs, leaving the domain easily recognizable for unseen classes and providing a false sense of removal. We argue that true domain erasure must be class-agnostic. To address this, we formalize Open-Vocabulary Domain Unlearning (OVDU), a rigorous protocol that mandates domain forgetting must transfer to held-out classes. To solve the OVDU challenge, we propose a surgical parameter-editing framework. First, a Fisher Information mask isolates domain-sensitive weights, mathematically protecting foundational zero-shot generalization. Second, our Targeted Manifold Scattering (TMS) objective uses preference-based mining to locally scatter the forget domain's stylistic geometry. Evaluated across PACS, OfficeHome, and DomainNet, our method vastly improves open-vocabulary generalization over existing baselines. Crucially, it delivers exceptional sample efficiency, outperforming peak 8-shot baseline results with only 4 shots.
Optimizing Agent Tool-Use via Trajectory-based Insight Evolution
Hanchen Qiu ⋅ Haojia Zhu ⋅ Yifan Meng ⋅ Jiahui Jin
Tool documentation is the critical interface bridging Large Language Model (LLM) agents and external environments. However, real-world documentation is often ambiguous or insufficient, leading to misaligned tool calls and task failures. While existing methods rely on parameter-efficient tuning or reinforcement learning to enhance tool-use capabilities, they incur substantial computational overhead and lack the agility to adapt to frequently updated toolkits. We propose OpTool , an evolutionary framework that treats tool documentation as an evolvable configuration. OpTool consists of three stages: (1) Contrastive Trajectory Generation, which explores diverse tool-use trajectories via contrastive beam search; (2) Trajectory-based Insight Extraction, which extracts tool-use insights through retrospective global credit assignment; and (3) Tool Documentation Evolution, which iteratively refines the documentation to mitigate information insufficiency. Experimental results across multiple domains demonstrate that OpTool significantly improves tool call accuracy and task success rates, providing a robust solution for tool-based LLM agents.
Origami as a Real-Image Benchmark for Procedural State-Transition Reasoning
Jun Suzuki ⋅ Reina Akama ⋅ Ikumi Numaya ⋅ Kise Yoshida ⋅ Yuka Saito ⋅ Sumika Tsuda ⋅ Takahiro Shimizu ⋅ Miori Sagara
General-purpose embodied AI systems are increasingly expected to interpret multimodal instructions and perform fine-grained physical tasks. However, many manipulation benchmarks provide limited support for jointly evaluating procedural correctness, local geometric sensitivity, and human-guided correction. We propose origami as a benchmark based on real images for procedural state-transition reasoning in embodied AI. Origami requires precise operations such as tucking, inserting, reversing, and folding overlapping layers, making it a focused domain for studying physical state transitions. We introduce a dataset of step-by-step folding images, textual entries, and difficulty annotations, and define two tasks: autonomous state transitions from multimodal instructions and human-guided collaborative transitions. Unlike crease-pattern- or diagram-based origami benchmarks, our setting exposes agents to actual intermediate folding states. For an initial simulator-based evaluation, we use a multimodal LLM controller to evaluate procedural understanding, action-interface grounding, and retry-based correction.
OTROPE: Optimal Transport-based Robust Off-policy Evaluation for Large Language Models
Liner Xiang ⋅ Wenbo Zhang ⋅ Hengrui Cai
Reliable evaluation of large language models (LLMs) is essential for their development and deployment, yet is often costly, risky, and difficult to perform safely online. We study off-policy evaluation for LLMs, where limited human-labeled data from a behavior model are used to evaluate a newer target LLM. This setting is challenging because labels are scarce, behavior--target distribution shift is common, and response likelihoods are often unavailable for black-box LLMs. We propose the Optimal Transport-based Robust Off-Policy Evaluation (OTROPE), a likelihood-free evaluation that performs distributional correction in a semantic space via optimal transport to align labeled behavior-policy samples with unlabeled target-policy samples. OTROPE combines corrected human-labeled residuals with proxy predictors, yielding a stable doubly robust-style evaluation without behavior-policy modeling or density-ratio estimation. We theoretically characterize why baseline evaluators fail under LLM distribution shift, and establish consistency and convergence rates for OTROPE when either the reweighted behavior distribution or the proxy predictor converges. Experiments on synthetic and real LLM evaluation tasks show that OTROPE consistently outperforms baselines while enabling ensembles of weaker LLM evaluators to approach and sometimes surpass stronger evaluators.
Otter Weather: Skillful and computationally-efficient medium-range weather forecasting
Cristiana Diaconu ⋅ Jonas Scholz ⋅ Aliaksandra Shysheya ⋅ Stratis Markou ⋅ Payel Mukhopadhyay ⋅ Miles Cranmer ⋅ Richard Turner
State-of-the-art medium-range AI weather models rival traditional Numerical Weather Prediction (NWP) but require massive training budgets. This restricts access for under-resourced groups and severely limits fast model iteration. We introduce Otter, a highly efficient spatiotemporal forecasting model designed to democratise high-performance weather prediction with AI. Otter is evaluated on ERA5 reanalysis data using the standard WeatherBench protocols where it significantly advances the skill-compute Pareto frontier. The deterministic version outperforms the best NWP baseline by 9.6\% at a 24-hour lead time while requiring fewer than 3.5 A100-days for training. It provides a 2x efficiency gain over lightweight AI models and a 100-fold reduction in compute compared to resource-intensive frontier architectures.
Outlier-Robust Multi-Output Gaussian Processes
Joshua Rooijakkers ⋅ Leiv Rønneberg ⋅ Francois-Xavier Briol ⋅ Jeremias Knoblauch ⋅ Matias Altamirano
Multi-output Gaussian process (MOGP) regression allows modelling dependencies among multiple correlated response variables. Similarly to standard Gaussian processes (GPs), MOGPs are sensitive to model misspecification and outliers, which can distort predictions within individual outputs. In the multi-output case, this situation can be further exacerbated as anomalous observations in one response can adversely influence predictions for other correlated outputs. To handle this situation, we propose R-MOGP—an MOGP that is provably robust to outliers in any individual response. Our approach preserves conjugacy, achieving robustness without sacrificing computational efficiency. Empirically, we show that R-MOGP achieves performance comparable to existing state-of-the-art robust MOGPs at a fraction of their computational costs.
Out-of-Distribution Detection in Continual Learning
Nimeshika Udayangani Hewa Dehigahawattage ⋅ Sarah Erfani ⋅ Flora Salim ⋅ Christopher Leckie
Out-of-distribution detection in continual learning (OOD-CL) is highly challenging, as it requires (i) learning a sequence of tasks without catastrophic forgetting, and (ii) identifying unknown samples (OOD) without access to known (in-distribution (ID)) data from past tasks, where standard confidence-based OOD scores tend to degrade. Existing methods largely rely on a rehearsal buffer of past data or assume closed-world settings, limiting their scalability and reliability in real-world scenarios. In this paper, we propose a novel OOD detection framework for rehearsal-free continual learning, inspired by Kolmogorov–Arnold Networks (KAN), which mitigate catastrophic forgetting via their inherent local neuroplasticity. To better leverage KAN in OOD-CL, we introduce a task-adaptive classifier architecture, \emph{TA-KAN}, which mitigates recency bias and enhances ID--OOD separability through task-specific control of activation locality. Furthermore, we propose a geometry-guided OOD score that complements TA-KAN classifier confidence without relying on past ID data. Our method significantly improves OOD detection while maintaining strong ID classification performance, achieving state-of-the-art results in OOD-CL.
Overcoming Kernel Redundancy for Scaling Logic Gate Networks
Sejin Park ⋅ Hongjae Lee ⋅ Changwoo Han ⋅ Seung-Won Jung
Differentiable logic gate networks, which operate using only logic gates, have recently attracted attention as an efficient alternative to conventional neural networks. However, despite their efficiency, the scaling behavior of logic gate networks remains underexplored. By contrast, scaling model capacity is a central design principle in deep neural networks and typically leads to improved performance. This discrepancy raises a key question: Can similar scaling benefits also be achieved in logic gate networks? In this work, we focus on width as a primary scaling axis and conduct a systematic analysis of its behavior in logic gate networks. We observe that naive width scaling often introduces redundancy among logic kernels, limiting the effective use of additional kernels and leading to performance saturation. To address this limitation, we propose a dynamic logic kernel framework that reorganizes kernel utilization by promoting specialization across kernel groups. This enables the network to better utilize increased width via input-dependent kernel routing, while ensuring that both routing and computation are implemented entirely with gate-level Boolean operations at inference time. We further find that kernel redundancy is most pronounced at the first gate level, motivating an early-stage dynamic logic kernel strategy that concentrates adaptation at this level. Experimental results demonstrate that our approach improves kernel utilization and increases kernel diversity, leading to higher accuracy with improved parameter efficiency.
OverLay++: Dense-Overlap Layout-to-Image Generation Dataset
Shivansh Aggarwal ⋅ Shresth Grover ⋅ Divyansh Srivastava ⋅ Haiyang Xu ⋅ Bingnan Li ⋅ Xiang Zhang ⋅ Ethan Armand ⋅ Chuan Li ⋅ Jianwen Xie ⋅ Zhuowen Tu
Layout-to-Image generation has made substantial progress in image generation with spatial and object-level control. However, existing methods still struggle with complex scenes containing many overlapping and interacting objects. We argue that training data is a particular bottleneck: existing datasets lack examples with dense, complex, object interactions. To address this gap, we introduce **OverLay++**, a large-scale Layout-to-Image dataset with structurally complex scenes. OverLay++ contains approximately 500K images with an average of 6.6 objects per image, exceeding existing datasets by $1.67\times$ in annotation density. Beyond annotation density, OverLay++ provides rich semantic detail with object captions over six times longer than current datasets. Our dataset generation pipeline is simple and robust, producing accurate overlapping regions with rich per-object captions. Across multiple benchmarks, state-of-the-art Layout-to-Image methods trained on OverLay++ dataset show consistent improvement and faster convergence demonstrating the importance of dense, overlap-aware, and caption-rich supervision for controllable image generation.
PACE: Partial-state Amortized Constraint Editing for Neural Combinatorial Optimization
Bohao Li ⋅ Chenhao Yuan ⋅ Ying Li ⋅ Pei He ⋅ Yangming Guo
We propose PACE, Partial-state Amortized Constraint Editing, which treats task-native feasible partial states as the common semantic object for neural combinatorial optimization, unifying edit learning, constraint-preserving editing, anytime closure, and test-time scaling across routing and graph combinatorial optimization. Existing neural combinatorial optimization methods expose complementary strengths: constructive and adaptive expansion solvers keep meaningful partial solutions but couple them to serial or method-specific growth, while global prediction, diffusion, masked reconstruction, and complete-solution refinement provide scalable guidance yet often refine heatmaps, noisy solutions, reconstruction targets, or perturbed full outputs. These semantics make learning, search, and test-time scaling act on the same state object. PACE learns amortized constraint editing by ranking candidate edits conditioned on the current feasible partial state; a constraint-preserving transition $\Gamma_I$ commits only edits that preserve legal continuation and extendability; and task-native closure $\mathrm{Comp}_I$ maps any extendable intermediate feasible partial state to a task-terminal feasible output. Budgeted refinement then spends extra inference-time compute through deeper edit steps, additional refinement rounds, or broader candidate sets along the same trajectory. We instantiate these semantics across TSP, ATSP, CVRP, MIS, MVC, MCL, and MCut, covering edge-oriented routing and node-oriented graph combinatorial optimization. Theoretical guarantees are structural rather than optimality claims, covering task-wise soundness, constructive completion, state-space closure in the extendable subset, terminal feasible completion at any editing depth, and elite-pool nondegradation. Empirically, PACE is competitive and budget-controllable, including ATSP-500 at 0.247\% Drop, MIS RB-[800-1200] at 2.00\%, and MVC at 0.07\%.
PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization
Wonjoong Kim ⋅ Yeonjun In ⋅ Sangwu Park ⋅ Dongha Lee ⋅ Chanyoung Park
A significant hurdle for current LLMs is the execution of complex, multi-stage tasks. Group Relative Policy Optimization (GRPO) has been emerging as a leading choice, but its reliance on sparse outcome rewards severely limits credit assignment across intermediate steps. Existing remedies—running full rollouts to assign step-level advantages, calling external LLM judges at each step, or computing intrinsic rewards that require ground-truth answers at every evaluation—introduce significant costs or practical constraints. We hypothesize that internal correctness probing over LLM hidden states can be repurposed as a step-level reward signal, potentially addressing all of these limitations at once. However, existing probing research assumes clean inputs, and we first show that this assumption breaks down in multi-step settings: hidden-state probes degrade severely under prefix contamination—tracking coherence with the (possibly corrupted) prefix rather than grounded correctness—while attention-based features remain robust to contamination but underperform on clean prefixes. Building on this complementary relationship, we propose the Prefix-Aware Internal Reward (PAIR)—a two-stage model with a frozen hidden-state probe estimating belief-consistency and a lightweight attention-based head correcting it toward grounded correctness. Experimental results show that PAIR achieves the highest AUROC on contaminated trajectories while operating at negligible inference cost, enabling dense step-level reward signals for GRPO training without external model calls, ground-truth dependencies, or full-trajectory rollouts.
Paleoinspired Vision: From Exploring Colour Vision Evolution to Inspiring Camera Design
Yijie Lu ⋅ Zhimin Zong ⋅ Junjie Zhang ⋅ Shenghan Su ⋅ Lin Gu ⋅ Ziteng Cui ⋅ Zeyun Zhao ⋅ Yan Pu ⋅ Jing Lu ⋅ Daisuke Kojima ⋅ Ruogu Fang
The evolution of colour vision is captivating, as it reveals the adaptive strategies of extinct species while simulta- neously inspiring innovations in modern imaging technol- ogy. In this study, we present a simplified model of vi- sual transduction in the retina, introducing a novel opsin layer. We quantify evolutionary pressures by measuring ma- chine vision recognition accuracy on colour images shaped by specific opsins. Building on this, we develop an evolu- tionary conservation optimisation algorithm to reconstruct the spectral sensitivity of opsins, enabling mutation-driven adaptations to to more effectively spot fruits or predators. This model condenses millions of years of evolution within seconds on GPU, providing an experimental framework to test long-standing hypotheses in evolutionary biology , such as vision of early mammals, primate trichromacy from gene duplication, retention of colour blindness, blue-shift of fish rod and multiple rod opsins with bioluminescence. More- over, the model enables speculative explorations of hypo- thetical species, such as organisms with eyes adapted to the conditions on Mars. Our findings suggest a minimalist yet effective approach to task-specific camera filter design, op- timising the spectral response function to meet application- driven demands. The code will be made publicly available upon acceptance.
PANDA: Prior-guided Attentional Dual-path Architecture
Rishabh Jain ⋅ Federico Belotti ⋅ Stefano Coniglio ⋅ Pietro Lió ⋅ Stefano Fiorini
Hypergraph Neural Networks (HNNs), designed to model higher-order relations beyond pairwise interactions, typically rely either on fixed structural weights or on normalized attention. Fixed weights can propagate misleading information in heterophilic settings, while normalization makes attention coefficients competitive, in the sense that, in a message-passing framework, increasing the weight of one sender must reduce the weight of others in the same aggregation set. Our main contribution is the introduction of PANDA (Prior-guided AttentioNal Dual-path Architecture), a plugin for two-stage message passing HNNs applicable to both undirected and directed settings. The *main path* interpolates between learned receiver-normalized attention and structural incidence priors, while a non-competitive (not normalized) *auxiliary path* lets each sender contribute its own transformed information. We also show that, with fixed structural coefficients, PANDA recovers a certain complex-valued, convolution operator adopted in previous works. Across four HNNs spanning spatial and spectral methods in both directed and undirected settings, plugging PANDA in yields an average improvement of $5.18$ percentage points across eight real-world benchmark datasets.
Panoptic Saliency Ranking
Zhenyu Wu ⋅ Min Li ⋅ Chong Ma ⋅ Xun Gong ⋅ Wei Wang ⋅ Chenglizhao Chen ⋅ Aimin Hao ⋅ Shuo Li
We introduce Panoptic Saliency Ranking (PSR), a new task that aims to estimate the visual saliency of all objects within a scene. Unlike salient object ranking, which focuses only on ranking salient regions, PSR extends the problem to the entire scene, providing a holistic and fine-grained understanding of visual saliency. To facilitate this task, we construct PSR18K, the first large-scale benchmark for panoptic saliency ranking. It contains 17,861 images with 146,460 annotated object instances spanning 13 superclasses and 376 fine-grained categories. In addition, PSR18K provides relation graph annotations to explicitly model semantic and spatial relationships among objects. We further propose an efficient Intrinsic Relation Graph Network (IRGNet), which reinterprets Transformer self-attention as an implicit relational matrix and seamlessly transforms it into an explicit relation graph, thereby obviating the need for additional relation modeling and significantly improving inference speed. Extensive experiments on PSR18K and three widely used SOR benchmarks demonstrate that our IRGNet achieves new state-of-the-art performance and faster inference with fewer parameters. The dataset and code are publicly available at https://sites.google.com/view/PSR18K.
Peer review is sometimes seen as noisy and difficult to predict. How much of the final accept/reject decision is recoverable from the paper itself? We investigate this question with PaperLens, text and vision models trained to predict binary acceptance from anonymized paper artifacts. For a clean evaluation, we construct balanced, shortcut-controlled OpenReview and source-derived arXiv datasets that remove deanonymization artifacts and majority-class shortcuts. Supervised binary decision training is surprisingly strong. PaperLens outperforms frontier prompting and review-trained baselines on decision accuracy and ranking, while its acceptance probability remains well calibrated after validation-set scaling and correlates better to ratings than baselines. Vision consistently improves over markdown text, showing that layout, figures, and visual presentation carry reviewer-relevant signal. Scaling from 3B to 14B parameters yields only modest gains, with clean supervision and faithful paper representations mattering more than model size. Together, our results show that decision prediction provides a powerful signal for paper assessment and can even improve the alignment of AI reviewing agents.
Parabolic Position Encoding: Vision-Centric, Principled, Extrapolatable, General
Christoffer Koo Ohrstrom ⋅ Rafael I Cabral Muchacho ⋅ Yifei Dong ⋅ Filippos Moumtzidellis ⋅ Ronja Güldenring ⋅ Florian T. Pokorny ⋅ Lazaros Nalpantidis
We propose Parabolic Position Encoding (PaPE), a parabola-based position encoding for vision modalities in attention-based architectures. Given a set of vision tokens–such as from videos, event camera streams, images, or point clouds–our objective is to encode their positions while accounting for the characteristics of vision modalities. Prior works have largely extended position encodings from 1D-sequences in language to nD-structures in vision, but only with partial account of vision characteristics. We address this gap by designing PaPE from principles distilled from prior work: translation invariance, rotation invariance (PaPE-RI), distance decay, directionality, and context awareness. Extrapolation experiments on ImageNet-1K show how PaPE extrapolates remarkably well, improving in absolute terms by up to 10.5% over the next-best encoding. Generality experiments on 8 datasets across 4 modalities show that PaPE is a general vision position encoding, as PaPE matches the best baseline on 5 datasets and exceeds all on 2 datasets.
Parallel Computation Algorithms and Convergence Guarantees for Mean-Field Langevin Dynamics
Yoshihito Okamoto ⋅ Huanjian Zhou ⋅ Taiji Suzuki
Entropy-regularized optimization problems arise in various areas of machine learning, including mean-field neural networks, variational inference, and reinforcement learning. Mean-field Langevin dynamics (MFLD) is a dynamics capable of sampling from the optimal solution of such problems under appropriate conditions, and has been the subject of extensive recent research. As machine learning models grow in scale and parameter dimension, computational efficiency becomes increasingly important, especially in reducing the adaptive complexity of computation of MFLD, that is, the number of sequential rounds they require. In this work, we initiate the study of the adaptive complexity of MFLD and provide the first theoretical evidence that MFLD can indeed be accelerated through parallelism. In particular, we provide polylogarithmic convergence guarantees for both the infinite-particle and finite-particle setting under the log-Sobolev inequality.
ParaPC-FM: Accelerating Parallel Sampling via Principled Initialization
Bangwei Li ⋅ Chuan Gou ⋅ Jianrong Lu ⋅ Guoyao Yu ⋅ Jianhai Chen
Flow matching and diffusion models achieve high-quality generation by solving continuous-time generative dynamics, but their sampling remains inherently sequential and costly. Parallel sampling reduces wall-clock latency, yet often requires many function evaluations (NFEs) due to inaccurate future-state initialization and repeated correction. We propose ParaPC-FM, a fixed-point predictor--corrector framework that unifies high-order solver correction and future-state initialization, enabling both components to be incorporated into different parallel samplers. In particular, our UniP-based truncated initialization reuses past velocity predictions to estimate future states more accurately, reducing correction iterations and NFEs with almost no additional overhead. On a ParaDiGMS-style framework, ParaPC-FM achieves up to $1.25\times$ and $1.49\times$ speedups on FLUX.1-dev at $50$ and $100$ steps, and up to $1.22\times$ and $1.16\times$ speedups on Stable Diffusion 3 and Stable Diffusion 3.5 at $50$ steps. When integrated into ParaTAA, the proposed initialization further achieves up to $1.17\times$ speedup at $50$ steps. Importantly, these acceleration gains are obtained while keeping image quality nearly unchanged.
ParetoSlider: Diffusion Models Post-Training for Continuous Reward Control
Shelly Golan ⋅ Michael Finkelson ⋅ Ariel Bereslavsky ⋅ Yotam Nitzan ⋅ Or Patashnik
Reinforcement Learning (RL) post-training has become the standard for aligning generative models with human preferences, yet most methods rely on a single scalar reward. When multiple criteria matter, the prevailing practice of ``early scalarization'' collapses rewards into a fixed weighted sum. This commits the model to a single trade-off point at training time, providing no inference-time control over inherently conflicting goals -- such as prompt adherence versus source fidelity in image editing. We introduce ParetoSlider, a multi-objective RL (MORL) framework for Diffusion Models post-training that allows continuous reward control during inference. By training the model with continuously varying preference weights as a conditioning signal the model learns how its denoising trajectory should change as the requested reward priority changes. Consequently, we enable users to navigate optimal trade-offs at inference time without retraining or maintaining multiple checkpoints. We evaluate ParetoSlider across three state-of-the-art flow-matching backbones: SD3.5, FluxKontext, and LTX-2. Our single preference-conditioned model matches or exceeds the performance of baselines trained separately for fixed reward trade-offs, while uniquely providing fine-grained control over competing generative goals.
Past as State, Present as Attention: Persistent-State Blockwise Flow Matching for Long-Horizon Co-Speech Motion Generation
Yi Liu ⋅ Xiangyue Zhang ⋅ Jia Ma ⋅ Jianfang Li ⋅ Yisheng He ⋅ Jianqiang Ren
Co-speech motion generation remains challenging over long horizons due to the need for temporal coherence and consistent style. Existing methods typically generate long motions segment by segment, either conditioning on a short seed of recent frames, which gradually drifts, or retaining ever-longer explicit histories, which are costly to maintain and hard to learn from limited data. We trace these limitations to a key asymmetry between past and present: past motion carries slowly varying style that should be compactly summarized, whereas the current segment exhibits fast-changing local dynamics that require fine-grained modeling. This asymmetry motivates a simple $\textbf{P}$ast-$\textbf{A}$s-$\textbf{S}$tate, $\textbf{P}$resent-as-$\textbf{A}$ttention principle for long-horizon co-speech motion generation. Following this principle, we propose $\textbf{PASPA}$, which summarizes past history with an inter-block persistent state and models the current segment with intra-block bidirectional self-attention. We instantiate PASPA within a blockwise flow matching framework, with multi-block supervision for effective training over variable histories, prefill-decode state caching for efficient rollout, and hybrid classifier-free guidance for enhanced history and speech conditioning. Extensive public-benchmark experiments show that PASPA achieves state-of-the-art FGD performance and preserves motion style across long-sequence generation.
Patch4Patch: Restoring Structural Connectivity in Patch-based Vision Encoders
Yaqi Zhang ⋅ 顺天 姚 ⋅ Niantai Qu ⋅ Runguo Chen ⋅ Hanyu Lai ⋅ Ningxin Su ⋅ Li Sun ⋅ Sen Su
As Multimodal Large Language Models (MLLMs) continue to advance and find increasingly broad applications, a key challenge persists: the patch-based vision paradigms widely adopted by MLLMs inevitably introduce structural fragmentation when processing information-dense images such as tables and charts. Specifically, when a table row or chart axis spans multiple visual partitions, the lack of cross-partition communication causes the model to lose track of structural continuity, leading to performance fluctuations when layout changes --- a problem we term layout sensitivity. To address this fundamental limitation, we propose Patch-for-Patch (P4P), a lightweight module that can be plugged into existing patch-based vision encoders. P4P dynamically selects a sparse set of tokens as bridging anchors and uses them to establish cross-partition communication through bidirectional attention. Combined with a gated fusion strategy, P4P restores global structural coherence without the quadratic cost of global attention. Extensive experiments across Qwen2.5-VL, MiniCPM-V, and InternVL3.5 demonstrate that P4P consistently improves structured content visual understanding, notably boosting TableEval accuracy by up to 9.12\% and significantly reducing layout sensitivity. The source code is available at: https://anonymous.4open.science/r/P4P/.
Repeated AI assistance can improve immediate task performance while reducing the skill available for future independent work. We develop a mathematical framework for this long-run tradeoff. The model tracks two state variables: a latent human skill level governing expected independent performance, and a delegation level representing the learner's evolving tendency to rely on AI. Skill changes through error-driven learning under practice and decay under delegation; delegation responds to observed performance, increasing when AI-assisted work appears to outperform independent work. We analyze the resulting dynamics and contrast them with fixed delegation. With fixed delegation, skill follows a one-dimensional learning-decay process with a single stable equilibrium. With adaptive delegation, the coupled system has two attracting equilibria separated by the stable manifold of an interior saddle. The existence and geometry of this separatrix require a global phase-plane analysis of the coupled dynamics. The system is path-dependent: small differences in initial skill or reliance can lead to different long-run outcomes. We use this characterization to show that AI assistance can improve short-run performance while producing worse long-run performance than a no-AI baseline. Increasing AI capability can enlarge the basin of attraction of the low-skill equilibrium, making delegation appear beneficial for longer while increasing the risk of eventual skill loss. The qualitative picture is observed to persist across alternative specifications. Together, these results show that the risk is not AI assistance itself, but the coupling between performance-driven reliance and use-dependent skill change.
Pathway-Aligned Regulator Tokens for Interpretable Spatial Gene Expression Prediction from Histologyatial Transcriptomics Prediction from Histology
Hyun Namgung ⋅ Sanghyun Park
Predicting spatial gene expression from histopathology images enables transcriptomic analysis on archival H&E samples where matched spatial transcriptomics data is unavailable. Gene expression is organized through cyclic regulatory circuits: transcription factor complexes are themselves products of regulated genes, aggregating upstream gene activity and modulating co-expressed downstream programs (e.g., the ISGF3 complex coordinating hundreds of interferon-stimulated genes through shared cis-regulatory elements). Existing approaches to histology-to-expression prediction, both deterministic and generative, employ generic architectures that ignore this structure. We introduce regulator attention, a mechanism that encodes this cycle by pooling gene representations into a small set of content-dependent tokens via a learned selector and routing them back through cross-attention to modulate gene-level features, in contrast to learnable token arrays whose identity is independent of the input. Within a flow matching framework conditioned on pathology foundation model embeddings, regulator attention consistently outperforms deterministic and generative baselines on correlation-based metrics across four tissue types spanning human cancer, normal tissue, and cross-species samples, with substantial gains on highly variable genes. Beyond predictive accuracy, 87.5\% of regulator tokens align with known biological pathways without supervision. Replacing them with content-independent learned queries eliminates this specialization and degrades predictive performance, isolating content-dependent token generation as the causal mechanism.
PDE-PFN: Prior-Data Fitted Neural PDE Solver
Jaehyeon Park ⋅ Mingu Kang ⋅ Dongseok Lee ⋅ Woojin Cho ⋅ Kookjin Lee ⋅ Anthony Gruber ⋅ Youngjoon Hong ⋅ Noseong Park
Motivated by the success of large language models (LLMs) with broad generalizability and robustness to noisy or unreliable pre-training data, we seek to bring similar capabilities to PDE solvers. In addition, inspired by the posterior predictive mean-based inference mechanism of the in-context learning in prior-data fitted networks (PFNs), we propose PDE-PFN, a prior-data fitted neural solver that approximates the posterior predictive mean of PDE solutions via in-context learning. PDE-PFN builds on a PFN architecture with self- and cross-attention mechanisms of Transformer and is pre-trained on noisy approximate solutions generated by physics-informed neural networks, serving as diverse but not necessarily exact priors. Through experiments on a range of two-dimensional PDEs, we demonstrate that PDE-PFN achieves empirical generalization across heterogeneous equations, robustness under noisy priors, and zero-gradient-update in-context inference capability. Our approach not only outperforms task-specific baselines but also provides a flexible and robust framework for advancing SciML.
PepDDG: Peptide–Protein Binding ΔΔ𝐺 Prediction via Information Channel Decomposition
Ruochi Zhang ⋅ Yusi Fan ⋅ Qiong Zhou ⋅ Li Jiao ⋅ Tian Wang ⋅ Qian Yang ⋅ Silong Zhai ⋅ Lan Wang ⋅ Fengfeng Zhou ⋅ Yajuan Huang ⋅ Liming Guo ⋅ Chang Liu ⋅ Xin Gao
Predicting mutation-induced changes in peptide--protein binding affinity ($\Delta\Delta G$) is central to therapeutic peptide optimization, but peptide-specific predictors remain limited by scarce labels, target overlap in supervised benchmarks, and costly molecular simulations. We introduce PepDDG, a zero-shot, training-free predictor that ranks peptide mutations by decomposing binding perturbations into three complementary information channels: energetic perturbation, geometric environment, and evolutionary compatibility. Each channel is computed from a wild-type complex structure, transformed into rank space, and combined by non-parametric Borda aggregation, without fitting fusion weights or fine-tuning neural predictors. PepDDG is motivated by rank-covariance analysis showing cross-channel fusion can improve rank correlation when channels retain complementary signal, whereas increasingly expensive refinement of a single energetic channel has diminishing returns. Empirically, the channels are individually moderate but capture distinct mutation regimes, and rank fusion consistently outperforms single-channel variants and training-free baselines. On our curated benchmark of 332 peptide-chain mutations across 33 targets, PepDDG reaches Spearman $\rho = 0.619$; the calibrated PepDDG-Cal variant reaches $\rho = 0.691$. Performance remains strong on short and cyclic peptide subsets, with $\rho = 0.813$ and $0.872$, respectively. These results support complementary evidence fusion on PDB-derived complex structures as a practical alternative to costly same-channel physical refinement for peptide $\Delta\Delta G$ ranking; a separate diagnostic shows that predicted structures can be used when no crystal structure is provided at inference time.
Pessimistic Latent Task-aware Optimization for Robust Offline Meta-Reinforcement Learning
Chang-Hoon Jeong ⋅ Kisung Shin ⋅ Hanwool Sul ⋅ Byoung-Tak Zhang
Context-based offline meta-reinforcement learning (COMRL) has progressed almost entirely through better task representations: separating task identity from behavior-policy artifacts so that the latent reflects the task information alone. The implicit expectation has been that a clean representation will transfer to out-of-distribution (OOD) tasks. However, on an OOD task the encoder produces a latent outside the training support, where the policy has never been trained. We address this gap with Pessimistic Latent Task-aware Optimization (PLATO), which exposes the policy to off-support latents at training time. PLATO assumes a mild manifold hypothesis: task latents lie on a low-dimensional manifold, with OOD tasks further from the training centroid. It exploits this geometry by perturbing inferred latents outward from the centroid to synthesize counterfactual tasks; a learned decoder ensemble then rolls out short trajectories at the perturbed latent, and the policy is updated on the imagined transitions under a pessimism weight that down-weights steps where the ensemble is unconfident. We prove that the perturbation reaches off-support by a controllable margin, that the epistemic disagreement at the perturbed latent is non-vanishing, and that the OOD generalization gap is bounded by the same perturbation reach. On eight MuJoCo continuous-control benchmarks, PLATO consistently improves out-of-distribution returns over strong COMRL baselines while remaining competitive in-distribution.
PGID: Progressive Guided Inversion and Denoising for Robust Watermark Detection
Minh Quoc Duong ⋅ Chun Tong Lei ⋅ Chun Pong Lau
With the proliferation of AI-generated images, digital watermarking has become an essential safeguard for protecting intellectual property and mitigating malicious exploitation. Recent works on semantic watermarking have enabled efficient copyright protection for diffusion models. However, the dependence of semantic watermarking on diffusion inversion for watermark detection creates a critical vulnerability. Imprint removal and forgery attacks exploit this weakness to produce deceptive results. Our analysis reveals that these attacks succeed by displacing watermarked latents into the unwatermarked region, while guiding unwatermarked latents into the watermarked region. Based on that, we propose Progressive Guided Inversion and Denoising (PGID), the first plug-and-play, training-free noise extraction framework designed to defend against both attack strategies. PGID effectively defends by projecting perturbed latents back to the region where they originally belong. The projection is achieved by eliminating intermediate latent deflections and mitigating adversarial perturbations through progressive inversion-denoising cycles. Comprehensive evaluations across multiple schemes demonstrate that PGID successfully restores detection reliability by recovering removed watermarks and identifying forged instances.
Phase-wise MLLM Tuning for Multi-framework WebUI Code Generation
Haoran Ma ⋅ Linxiao Li ⋅ Chenyue Wang ⋅ Haochen Sui ⋅ Jiechao Gao
Multimodal large language models (MLLMs) achieve strong performance in translating WebUI screenshots into HTML/CSS. However, when they faced multiple frontend frameworks (React/Vue/Angular), they often suffer from negative transfer that leads to compilation failures. These failures commonly stem from mixing framework specific syntax and violating framework constraints. This paper studies supervised fine tuning for multi-framework WebUI code generation and proposes a phase-wise MLLM tuning method driven by compatibility and heterogeneity signals across frameworks. We first apply a normalization mapping that preserves each framework's native syntax while suppressing noise from variable naming and literal content. We then estimate cross framework heterogeneity and compatibility offline, and automatically derive training groups and their ordering by maximizing a phased objective. Finally, we perform phase-wise adapter tuning to reduce cross framework interference. Experiments on multiple benchmarks show that the proposed method improves both generation quality and compilation success rate compared with baselines.
PhGPO: Pheromone-Guided Policy Optimization for Long-Horizon Tool Planning
Yu Li ⋅ Guangfeng Cai ⋅ shengtian yang ⋅ Han Luo ⋅ Shuo Han ⋅ Xu He ⋅ Dong Li ⋅ Lei Feng
Recent advancements in Large Language Model (LLM) agents have demonstrated strong capabilities in executing complex tasks through tool use. However, long-horizon multi-step tool planning is challenging, because the exploration space suffers from a combinatorial explosion. In this scenario, even when a correct tool-use path is found, it is usually considered an immediate reward for current training, which would not provide any reusable information for subsequent training. In this paper, we argue that historically successful trajectories contain reusable tool-transition patterns, which can be leveraged throughout the whole training process. Inspired by ant colony optimization where historically successful paths can be reflected by the pheromone, we propose Pheromone-Guided Policy Optimization (PhGPO), which learns a trajectory-based transition pattern (i.e., pheromone) from historical trajectories and then uses the learned pheromone to guide policy optimization. This learned pheromone provides explicit and reusable guidance that steers policy optimization toward historically successful tool transitions, thereby improving long-horizon tool planning. Comprehensive experimental results demonstrate the effectiveness of our proposed PhGPO.
PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop
Xinge Peng ⋅ Yiting Lu ⋅ tianwu zhi ⋅ Wen Wen ⋅ Jianzhao Liu ⋅ Xin Li ⋅ Zhibo Chen
Vision-Language Models (VLMs) have demonstrated impressive multimodal reasoning, yet their ability to truly internalize the underlying physical consistency of real-world dynamics remains an open question. Existing benchmark paradigms often suffer from fragmented evaluation, focusing on isolated cognitive stages while overlooking the inherent synergy between perception, reasoning, and physical judgment. The lack of a holistic perspective limits the ability to diagnose whether VLMs can reliably evaluate the physical authenticity of emerging generative models. To address these issues, we introduce \textbf{PhysVista}, a benchmark designed to evaluate physical intelligence in VLMs through a closed cognitive loop framework inspired by the human \textbf{seeing–reasoning–assessment} process. PhysVista restores this loop by jointly evaluating physical state perception, physical dynamics reasoning, and physical plausibility assessment. It further distinguishes \textbf{event-level} reasoning and \textbf{scale-level} reasoning to enable fine-grained analysis of physical understanding. In addition, PhysVista incorporates both real-world and AI-generated videos, allowing evaluation across diverse domains and emerging generative scenarios. Extensive experiments on a wide range of state-of-the-art VLMs reveal significant limitations in current models’ physical intelligence, particularly in fine-grained reasoning and plausibility assessment. Our findings highlight critical gaps between visual recognition and genuine physical understanding, and provide insights for developing future physically grounded VLMs.
PhyTS: A Benchmark for Scientific Time Series
Benedict Armstrong ⋅ Jeroen Audenaert ⋅ Hannah P Binney ⋅ Alice Cheng ⋅ Maarten De Vos ⋅ Luis F Domingues ⋅ Allison Eto ⋅ Mary Feliz ⋅ Joseph Formaggio ⋅ Jessica Fry ⋅ Philip Harris ⋅ Ilay Kamai ⋅ Erik Katsavounidis ⋅ Mykyta Kliapets ⋅ Konstantinos Kontras ⋅ Aobo Li ⋅ Paul Liang ⋅ Daniel Muthukrishna ⋅ Christina Reissel ⋅ Daniela Rus ⋅ T. Konstantin Rusch ⋅ Avi Shporer ⋅ Penny L Slocum ⋅ Talia E Weiss ⋅ Lindley Winslow ⋅ Kyungseop Yoon
We introduce PhyTS, a benchmark suite of precision scientific time series datasets for machine learning, spanning experiments in gravitational-wave detection, dark matter searches, neutrino mass determination, and exoplanet discovery. Despite their diverse scientific goals, these domains share a common challenge: recovering weak, structured signals and estimating underlying physical parameters from noise-dominated measurements. Unlike standard sequence modeling benchmarks such as audio and speech, these data exhibit non-Gaussian and nonstationary noise, long-range temporal correlations, detector-specific systematics, irregular sampling, and signals that are sparse, weak, or only partially modeled. As a result, they provide a challenging testbed for evaluating whether modern AI methods can support downstream scientific inference. We provide standardized tasks, data splits, and evaluation protocols for denoising, signal recovery, and parameter inference across physics domains, along with baseline results. By unifying diverse weak-signal inference problems under a common framework, this benchmark aims to enable reproducible evaluation and accelerate the development of more robust, interpretable, and physically grounded methods for scientific time series analysis.
PICID: A Modular Evaluation Infrastructure for Reproducible PHM Across Tasks and Domains
Lev Telyatnikov ⋅ Raffael Theiler ⋅ Leandro Von Krannichfeldt ⋅ Olga Fink
Progress in Prognostics and Health Management (PHM) is hindered by the lack of standardized and reusable evaluation practices across tasks, datasets, and application domains. Reported results are often difficult to reproduce and compare, as key protocol choices, such as data splits, preprocessing, label alignment, temporal windowing, and metrics, are often implicit or implemented ad hoc. We introduce \picid, a modular evaluation infrastructure that formalizes the PHM evaluation pipeline as an explicit, executable, and reproducible protocol. Through well-defined abstractions, \picid enforces deterministic, leakage-safe dataset construction while remaining flexible across diverse PHM settings. The framework supports fault detection, diagnostics, and prognostics through a unified interface and can be extended to new datasets and model classes without violating protocol invariants. By standardizing data contracts and evaluation boundaries, \picid also enables fair cross-task comparisons across diagnostics (classification) and prognostics (regression), allowing identical model families to be evaluated consistently across heterogeneous settings. We demonstrate \picid through an empirical evaluation of thirteen models on twelve datasets spanning batteries, bearings, turbofan engines, hydraulics, filtration systems, and buildings. This work establishes a reusable foundation for standardized, fair and reproducible evaluation in PHM.
PickMoment: Continuous-Time Single-Image-to-Video via Learning Deblurring and Blur-to-Video
Junseong Shin ⋅ Hyeonsu Jo ⋅ Daehyun Kim ⋅ Tae Hyun Kim
Motion blur arises from the temporal integration of a continuous sharp signal over a finite exposure window, yet existing learning-based methods sidestep this physical model and predict only the sharp signal itself: most single-image deblurring methods recover a single frame at the exposure center, while blur-to-video methods predict a fixed set of frames. We introduce PickMoment, a continuous-time reformulation that directly learns the interval-mean blur over arbitrary sub-intervals of the exposure with a single deterministic model. Drawing an analogy to MeanFlow's average-velocity formulation, we train the model with three supervisions derived from the blur integral: an empirical reconstruction loss from available subframes, an additivity loss that enforces self-consistency across overlapping sub-intervals, and a sharp-frame loss anchored at the zero-interval limit. A single trained model unifies single-image deblurring, blur-to-video generation, and continuous-time pick-a-moment recovery as different queries to the same network, with no separate training for each task. Our PickMoment achieves state-of-the-art performance among generative-based deblurring methods on GoPro and HIDE while competitive against restoration-based methods on RealBlur, and the highest per-frame fidelity on GoPro-7 blur-to-video, all in a single forward pass without iterative sampling.
PISG: Constraint-Aligned Signal Amplification for Diffusion-Based Combinatorial Optimization
Yuming Zhang ⋅ Xianchen Zhou ⋅ Hongxia Wang
Diffusion-based solvers have become a prominent approach to neural combinatorial optimization. At inference time, they assign each decision variable a prediction score. A greedy decoder ranks candidates by these scores and checks hard constraints to ensure feasibility. Top-ranked candidates should be both promising and jointly feasible, making prediction--constraint alignment critical. When this alignment is weak, greedy decoding skips constraint-violating candidates and falls back to lower-ranked alternatives, degrading solution quality. Existing test-time methods improve search, sampling, or decoding, but do not directly strengthen this alignment before decoding. We observe that trained denoisers already carry constraint-aligned structural signals. Building on this observation, we propose Perturbed Instance-Structure Guidance (PISG), a training-free method that amplifies these signals at test time. At each denoising step, PISG constructs a structure-degraded prediction and extrapolates the unperturbed prediction away from it. The resulting guided prediction improves candidate rankings before constrained decoding. Signal analysis shows that this guidance concentrates on constraint-critical variables rather than uniformly rescaling scores. Under greedy decoding and without downstream refinement, PISG improves all 16 TSP solver-scale pairs across DIFUSCO, T2T, FastT2T, and StruDiCO, and all four MIS-ER solvers, with gains up to +8.4\%. These results identify prediction-constraint alignment as an effective test-time enhancement direction for diffusion-based combinatorial optimization.
Pixels over Symbols: Sensory Realism Improves Behavioral Alignment in Models of Cognition
Mark Bai ⋅ Xiaoxuan Lei ⋅ Zihan Weng ⋅ MOTAHAREH POURRAHIMI ⋅ Lucas Gomez ⋅ Pouya Bashivan
Computational models of cognition have traditionally relied on abstract symbolic inputs that approximate sensory experiences with simplified vector representations. This simplification reflects historical computational constraints and a long-held assumption that the details of sensory realism are incidental to the study of cognitive processes. Recent advances have enabled the development of sensory-realistic models capable of performing diverse cognitive tasks, yet whether sensory realism meaningfully shapes cognition has not been formally established. We directly address this gap through a large-scale empirical comparison of abstract and sensory-realistic neural network models trained on a battery of memory-dependent decision-making tasks. Our results show that sensory-realistic models mirror human behavioral response patterns significantly more accurately than their abstract counterparts, though a substantial gap to human internal behavioral consistency remains even for the best tested models. By analyzing the internal dynamics of these systems, we demonstrate that the recurrent dynamics in abstract models is not inherently divergent from that in natural models. In fact, frozen abstract-trained weights can efficiently solve cognitive tasks from sensory-realistic inputs, however, at the cost of diverging the dynamics from their original configuration. Together, our findings argue that sensory realism is a primary determinant of human-like behavior in artificial systems, suggesting that findings obtained from abstract task models in neuroscience should be more carefully interpreted in light of their potential behavioral limitations.
PLANING: A Loosely Coupled Triangle-Gaussian Framework for Streaming 3D Reconstruction
Changjian Jiang ⋅ Kerui Ren ⋅ Xudong Li ⋅ Kaiwen Song ⋅ Guanghao Li ⋅ Linning Xu ⋅ Tao Lu ⋅ Junting Dong ⋅ Yu Zhang ⋅ Yu Feng ⋅ Bo Dai ⋅ Mulin Yu
Streaming reconstruction from monocular image sequences remains challenging, as existing methods typically favor either high-quality rendering or accurate geometry, but rarely both. We present PLANING, an efficient on-the-fly reconstruction framework built on a hybrid representation that loosely couples explicit geometric primitives with neural Gaussians, enabling geometry and appearance to be modeled in a decoupled manner. This decoupling supports an online initialization and optimization strategy that separates geometry and appearance updates, yielding stable streaming reconstruction with substantially reduced structural redundancy. Despite operating online, PLANING improves dense mesh Chamfer-L2 by 18.52% over the offline baseline PGSR, surpasses the streaming baseline ARTDECO by 1.31 dB in PSNR, and reconstructs ScanNetV2 scenes in under 100 seconds, over 5x faster than offline 2D Gaussian Splatting while matching per-scene optimization quality. Beyond reconstruction quality, the structural clarity and computational efficiency of PLANING make it well suited for a broad range of downstream applications, such as enabling large-scale scene modeling and simulation-ready environments for embodied AI.
Planning Persuasion, Not Utterances: Profile-Conditioned Open-Loop Search for Dialogue Strategy
Zhiming Lin ⋅ Kai Zhao ⋅ Yuliang Gai ⋅ Chengzhen Yu ⋅ Ruibo Duan ⋅ Yihao Zhong ⋅ Xuan Zhang
Persuasive dialogue systems are increasingly important for prosocial communication, education, and support-oriented interaction, but they still struggle to make profile-sensitive decisions when user reactions are uncertain and persuasive outcomes are delayed. This paper aims to enable stable long-horizon persuasive planning that adapts to user profiles while avoiding brittle text-level search and costly prompt-based self-evaluation. We propose PMCTS-PD, a profile-conditioned open-loop tree search framework that plans over compact dialogue-act prefixes rather than full utterances, thereby separating strategic decision making from surface realization. Each search node aggregates multiple instantiated dialogue histories, while a lightweight profile-conditioned Value-LLM predicts normalized expected donation to guide tree search, value backup, and value-guided best-of-K response generation. On PersuasionForGood, PMCTS-PD improves over strong LLM and dialogue-planning baselines, achieving higher BLEU-4, embedding similarity, diversity, predicted donation, and end-to-end success rate. These results show that planning persuasion at the strategy level, supported by a learned profile-aware outcome signal, offers a practical and cost-effective direction for robust persuasive dialogue systems.
PlasticMem: Adding Temporal Reasoning to Diffusions for Consistent Long Video Generation
Yufei Huang ⋅ Chao Liang ⋅ Zerong Zheng ⋅ Tianshu Hu ⋅ Jianwen Jiang ⋅ Yuan Zhang ⋅ Mingyuan Gao ⋅ Chang Yu ⋅ Stan Z. Li
Scaling video generation from short to long durations has attracted growing attention as a way to democratize video creation. The central challenge lies in long-term consistency, which requires stable rollout over extended horizons and accurate visual memory of previously generated elements. However, the scarcity of high-quality long video data remains a fundamental bottleneck, forcing prior methods to rely on suboptimal synthetic data that either limits dynamic range and visual memory or suffers from a large simulation-to-reality gap. To address this, we explore a new generation paradigm, Temporal Reasoning-based Rendering (TRR), which better exploits existing short video data for consistent long video generation. Under TRR, we propose PlasticMem, which plans contents from scripts, then reasons about and independently renders each action to form a consistent long video. Benchmark results show that PlasticMem, trained on short videos, achieves stable minute-scale long-horizon generation and visual memory.
Plausible Biomolecular Structure Prediction via Physics-informed Reinforcement Learning
Tai Dang ⋅ Hieu Tran ⋅ Long-Hung Pham ⋅ Sang Truong ⋅ Edward A Pham ⋅ Jeffrey S Glenn ⋅ Thang Luong
The latest advances of diffusion models in biomolecular structure prediction are still based on distance-based supervised training without adequately accounting for physically plausible interactions needed for real-world drug discovery. We introduce $\mathit{\boldsymbol{\pi}\text{-Fold}}$, a novel physics-informed policy-based folding approach that augments reinforcement learning with physics-aware scoring functions for training biomolecular diffusion models. Specifically, we fine-tune AlphaFold3-based models using Group Relative Policy Optimization and explicit physics-based rewards such as Rosetta all-atom energy, AutoDock Vina binding affinity, and ligand strain energy. Crucially, we demonstrate that these physics-based priors translate to enhanced performance on critical downstream prediction tasks. $\mathit{\boldsymbol{\pi}\text{-Fold}}$ achieves state-of-the-art results on SKEMPI for protein--protein mutation effects, CASP16 and OpenFE for protein--ligand binding affinities, Run-N-Poses for ligand geometric validity and interface quality across diverse structural fidelity metrics. On FoldBench, $\mathit{\boldsymbol{\pi}\text{-Fold}}$ achieves the best overall MolProbity score of 1.30, halving steric clashes and nearly eliminating rotamer outliers. Human evaluation further confirms that $\mathit{\boldsymbol{\pi}\text{-Fold}}$ generates structures that are meaningful for downstream drug discovery applications. Our code is available at \href{https://anonymous.4open.science/r/pifold-45B8/README.md}{github.com/anonymous/pifold}.
PoEM: Predicting New RL Outcomes from Existing Policies
Kimia Hamidieh ⋅ Giannis Daras ⋅ Antonio Torralba
Pre-trained models are routinely post-trained with reinforcement learning (RL) to maximize specific rewards, such as human alignment, correctness, or instruction following. This post-training process is computationally intensive, sometimes unstable, and has to be run from scratch every time the reward model changes or when we want to combine multiple rewards. We hence ask: given a new reward function, is it possible to predict the RL outcomes without actually running RL on it? We answer this in the affirmative by introducing PoEM, a framework to predict the outputs of RL on a new reward function using a set of models already post-trained on other rewards. First, we show that if the new reward function can be written as a linear combination of existing ones, then the new policy in log-space can be written as a linear combination of the existing log-policies. Surprisingly, even in cases where the rewards are not linearly connected, we observe that often log-policies from RL training span an approximately low-rank subspace across rewards. To our benefit, the weighting coefficients for this combination can be estimated using only the reward or basis policy outputs on the samples. We turn these observations into an algorithm that takes post-trained models and a new reward function, and approximates the target RL policy without actually running any additional RL training. We experimentally validate our approach across synthetic and real experiments, spanning both text and image modalities.
Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
Alexandra Souly ⋅ Javier Rando ⋅ Ed Chapman ⋅ Xander Davies ⋅ Burak Hasircioglu ⋅ Ezzeldin Shereen ⋅ Carlos Mougan ⋅ Vasilios Mavroudis ⋅ Erik Jones ⋅ Chris Hicks ⋅ Nicholas Carlini ⋅ Yarin Gal ⋅ Robert Kirk
Poisoning attacks can compromise the safety of large language models (LLMs) by injecting malicious documents into their training data. Existing work has studied pretraining poisoning assuming adversaries control a percentage of the training corpus. However, for large models, even small percentages translate to impractically large amounts of data. This work demonstrates for the first time that poisoning attacks instead require a near-constant number of documents regardless of dataset size. We conduct the largest pretraining poisoning experiments to date, pretraining models from 600M to 13B parameters on chinchilla-optimal datasets (6B to 260B tokens). We find that 250 poisoned documents similarly compromise models across all model and dataset sizes, despite the largest models training on more than 20 times more clean data. We also run smaller-scale experiments to ablate factors that could influence attack success, including broader ratios of poisoned to clean data and non-random distributions of poisoned samples. Finally, we demonstrate the same dynamics for poisoning during fine-tuning. Altogether, our results suggest that injecting backdoors through data poisoning may be easier for large models than previously believed as the number of poisons required does not scale up with model size—highlighting the need for more research on defences to mitigate this risk in future models.
Polestar: Drift-Aware Cache Calibration and Token Commitment for Efficient Inference of Diffusion LLMs
Mingyu Lee ⋅ Akshat Ramachandran ⋅ Souvik Kundu ⋅ Tushar Krishna
The inference efficiency of diffusion large language models (dLLMs) is constrained by two challenges: bidirectional attention precludes efficient KV-cache reuse, while increasing decoding parallelism with static confidence thresholds can compromise generation quality. We observe that both challenges arise from a shared phenomenon: as tokens are decoded, their contextual integration through bidirectional attention causes token representations to drift (evolve) across decoding steps. This insight motivates Polestar, a training-free inference framework that uses token representation drift as a unified signal to jointly address both challenges. Polestar comprises two components: Polestar-Cache, which identifies stale KV-cache positions via drift and performs sparse KV-cache refreshes to enable efficient reuse, and Polestar-Commit, which detects sharp drift events to reliably identify commit-ready tokens. Across mathematics and coding benchmarks on several dLLM families, Polestar sets a new state of the art on the accuracy-throughput Pareto frontier, achieving up to 10.73% accuracy improvement, up to 3.7x higher throughput, and high decoding parallelism of 3.67 tokens per forward pass over existing baselines.
In multi-agent reinforcement learning (MARL), agents must not only discover high-reward behaviours but also reliably reproduce coordinated behaviours over time. A fundamental challenge arises from a mismatch between action-level stochastic exploration and temporally extended coordination. To address this issue, we propose a policy-level exploration framework with Deterministic Action Execution (DAE) for MARL, which shifts locally random action exploration to temporally extended policy selection via an option-based exploration mechanism. Specifically, at each time step, each agent selects a policy from a shared policy set\textemdash comprising one exploitation policy and multiple exploration policies\textemdash using a learned option policy, and actions are executed deterministically conditioned on the selected policy. By maintaining deterministic execution under temporally extended policy selection, DAE facilitates temporally coordinated behaviours across multiple time steps. Comprehensive experiments on four multi-agent tasks\textemdash Predator Prey, StarCraft II micro management challenge (SMAC), SMACv2, Google Research Football, and Multi-Agent Coordination benchmark (MACO)\textemdash demonstrate that DAE outperforms mainstream baselines, achieving superior sample efficiency and final learning performance.
Position: AI Development Should Prioritize Cognitive Security
Batu El ⋅ Shiye Su ⋅ Aneesh Pappu ⋅ Peggy Yin ⋅ Julie Heng ⋅ Eric Heng ⋅ Ryan Z Wang ⋅ Andreas Haupt ⋅ James Zou
Generative AI systems designed to influence human beliefs and actions are becoming increasingly pervasive, raising concerns about cognitive security -- the protection of human cognitive processes from hazardous influence. Recent advances have only amplified the cognitive risks of AI technologies, leading institutions worldwide to identify cognitive security as an emerging and urgent governance challenge. However, there currently exists no framework for integrating scattered research efforts on the applied cognitive effects of AI systems. In this paper, we argue that cognitive security must be prioritized in AI development and outline a roadmap to do so. First, we track and categorize state-of-the-art capabilities for cognitive influence in generative AI systems. Then, we propose attack and defense mechanisms to formalize threat models, expose vulnerabilities, and evaluate countermeasures. Finally, we define metrics to unify research efforts on the cognitive impacts of AI systems.
Posterior-First Neural PDE Simulation: Inferring Hidden Problem State from a Single Field
Wenshuo Wang ⋅ Fan Zhang
Neural PDE simulators often receive only a single observed field at deployment. In this setting, a field-to-future predictor can collapse distinct latent problem states into the same deterministic interface, losing the ambiguity needed for reliable rollout and downstream decisions. We propose posterior-first neural PDE simulation: first infer a posterior over the minimal task-sufficient problem state, then condition prediction on that posterior. The resulting theory connects the object, the learning target, and the failure mode: Bayes downstream values factor through this posterior, refinement labels make it learnable by proper scoring rules, and deterministic collapse incurs an ambiguity barrier whenever the true posterior is non-Dirac. Synthetic exact-ambiguity experiments show that point-versus-posterior gaps track the predicted barrier. On metadata-hidden PDEBench tasks, posterior recovery reduces pooled rollout nRMSE from 0.175 to 0.132, closing 59.4% of the direct-to-oracle gap. These results suggest that single-observation neural PDE simulation should be posterior-first rather than monolithic field-to-future prediction.
Practical Non-Stationary Graph Gaussian Processes
Senanayak Sesh Kumar Karri ⋅ Viacheslav (Slava) Borovitskiy
Matern Gaussian processes offer a principled framework for learning and uncertainty quantification on graphs. However, they struggle with heterogeneous networks. A central challenge is that variance at a node can represent genuine signal that should strongly propagate (diffuse) to neighbors, mere noise that should be isolated, or a mixture of both. While this problem has been studied in continuous spatial domains, directly translating those approaches to graphs leaves a free parameter per node, causing extreme over-parameterization. To resolve this, we introduce non-stationary Variance-Coupled Gaussian Processes (VCGPs). These models use an auxiliary signal (often easily derived from historical data) to determine relative diffusivity between nodes, tied to the regression data via a single learned hyperparameter, $\gamma$. The sign of $\gamma$ captures whether the auxiliary signal is directly or inversely proportional to node diffusivity, while its magnitude calibrates this relationship.We formally prove that $\gamma = 1$ is the unique point where VCGPs are stationary up to rescaling, ensuring that varying $\gamma$ results in genuinely non-stationary correlations. Across five real-world datasets, our VCGPs match or outperform both stationary and heavily parameterized non-stationary baselines.
Predicting What Changes: Causal Delta World Models for Risk-Aware LLM Agent Planning
Guo Yue ⋅ YANG LIU ⋅ Donghui Zhang ⋅ Tang Qingkang ⋅ Rong Fu ⋅ Li Aoyu
Modern LLM agents repeatedly fall into a familiar trap: they execute irreversible actions whose consequences they did not anticipate, such as confirming a non-refundable booking or sending an unrecoverable message. World models are a natural remedy, yet two dominant lines of work both fail this use case. Generative pixel and HTML world models pay a heavy compute cost to produce vivid but often unfaithful continuations of the full observation, while reward and value models only summarise how good an action is and tell the agent nothing about what it will change or whether the change can be undone. We argue that what an agent really needs is a model of what changes, not of what comes next. We introduce Causal Delta World Models (CDWM), which represent each transition as a typed structured delta over an entity, attribute and relation graph, and explicitly factor out reversibility, irreversible cost and prediction confidence. A learnable causal sparsity mask binds each action class to the delta dimensions it can plausibly affect, providing a strong inductive bias for unseen actions and environments. We post-train CDWM with a verifiable composite reward that aligns delta accuracy, action following, calibration error and closed-loop task success. The resulting world model is consumed by a risk-aware closed-loop planner that prunes catastrophic candidates, performs short-horizon model-predictive control, and abstains when confidence is low. On four interactive benchmarks (WebArena, Mind2Web, ALFWorld, ScienceWorld) and two visual control suites (Procgen and Atari 100k), CDWM trained on four NVIDIA A100 (80 GB) GPUs improves task success by 6.0 to 11.7 absolute points over strong agentic and world model baselines, reduces irreversible failures by 47 to 71 percent, lowers token cost per successful trajectory by 27 to 45 percent, and shows substantially better generalisation to unseen websites and rule perturbations. We conclude that the right unit for an agent's world model is the structured delta and not the full observa
Prediction-Powered Inference Across Many Tasks for AI Evaluations and Social Science Research
Nicolas Emmenegger ⋅ Ellery Stahler ⋅ Chara Podimata
Many applications require statistically valid inference across many related tasks, while providing only a handful of high-quality labels per task. In AI evaluation, these tasks may correspond to model behaviors across prompts, subgroups, or hypotheses; in social science surveys, they may correspond to related questions, populations, or measurement conditions. Prediction-powered inference uses inexpensive proxy measurements to improve inference from limited labels, but standard methods operate task-by-task and therefore struggle in the small-label regime. We introduce a multi-task prediction-powered inference framework that borrows strength across tasks without pooling away validity. Our methods learn surrogate outcomes using labeled data from other tasks while retaining within-task rectification for valid per-task confidence intervals. We prove that efficiency gains beyond power-tuned PPI require nonlinear structure in the proxy–ground-truth relationship: affine cross-task recalibrations are oracle-equivalent to using the original proxy. We complement our theoretical findings with experiments on semi-synthetic datasets and a case study auditing language models on election-related information during the 2024 U.S. presidential election. Using a large human-annotation study, we show that cross-task surrogate learning can substantially reduce confidence interval widths when labels are scarce.
Preference-Guided Adaptation for Open-Vocabulary Semantic Segmentation via Prompt Disagreement
Hyun-Kurl Jang ⋅ Jihun Kim ⋅ Kuk-Jin Yoon
Open-vocabulary semantic segmentation (OVSS) enables pixel-level prediction over arbitrary text-specified vocabularies and has shown strong generalization on common benchmarks. However, OVSS performance often degrades in specialized domains such as medical imaging, remote sensing, and industrial inspection, where dense pixel-level masks for adaptation are costly to obtain and require domain-specific expertise. We propose a preference-guided adaptation framework that replaces dense mask supervision with binary preferences. We observe that different prompt templates produce systematically different segmentations for the same image, a phenomenon we call prompt disagreement, and we repurpose it as a built-in source of preference supervision. Building on this, we mine localized preference queries from regions of high cross-template uncertainty, and adapt the OVSS model with Region-Localized Preference Optimization (RLPO) together with consistency regularization that stabilizes updates outside the queried region. Across extensive experiments on the MESS benchmark, the proposed method achieves consistent gains across diverse OVSS backbones without any pixel-level annotation, and remains effective under noisy preferences.
Prefix-Tuning for Arbitrary Output Sequences on Pretrained Transformers
Peter Cho-Ho Lam ⋅ Ziyi Wang ⋅ Zirui Zhou
Prefix-tuning is a parameter-efficient fine-tuning method for large language models (LLMs) that steers model outputs by prepending task-specific prefix representations to the input sequence. Despite its practical success, the theoretical understanding of the expressiveness of prefixes remains incomplete. In particular, the following question remains open: {\it Given a pretrained autoregressive Transformer such as the GPT series, for any input-output sequence pair $(\mathtt{x},\mathtt{y})$, does there always exist a prefix $\mathcal{S}$ such that prepending $\mathcal{S}$ to $\mathtt{x}$ induces the model to generate the target output $\mathtt{y}$?} In this work, we prove that under the assumption that $\mathtt{y}$ is in the {\it range} of the pretrained Transformer, such a prefix $\mathcal{S}$ exists and can be explicitly computed by solving multiple linear systems. We further show that this assumption is necessary, in the sense that such a prefix may not exist when the assumption fails, thereby providing a complete answer to the above question. Our analysis is constructive and the obtained theoretical results hold under fairly general assumptions on model architecture. As a notable example, the whole GPT-2 architecture satisfies our model assumptions.
Preserving Exploration for LLM Reasoning via Mean Order-Statistic Alignment
Ruotian Peng ⋅ Yi Ren ⋅ Long-Fei Li ⋅ Yandong Wen
Pre-trained large language models exhibit strong exploration ability, as evidenced by high pass@K scores, yet this ability degrades significantly during reinforcement learning (RL) fine-tuning due to mode collapse. A natural remedy is to impose token-wise KL divergence between the fine-tuned and base models; however, we argue that such objective introduces a fundamental mismatch, where token-wise KL operates locally at individual token positions, while exploration is a global behavior that emerges over sequences as a whole. Motivated by this insight, we propose Mean Order-Statistic Alignment (MOSA), which preserves exploration by aligning the fine-tuned model's global behavior to that of the base model. Specifically, MOSA captures global exploration through the mean order-statistic profile, which is obtained by computing the order statistics of each token's posterior over the vocabulary and averaging them across tokens at each rank position. The profile, designed a valid probability distribution, directly admits a principled KL-based objective. For efficiency, we further approximate the profile using only the top-K probabilities and a residual tail bucket, yielding a compact implementation with minimal computational overhead. Experiments on Countdown, 6 mathematical reasoning, and 2 coding benchmarks show that MOSA discovers more diverse plausible solutions, improves both pass@1 and pass@K, and scales effectively to models up to 32B parameters.
Pretraining Curricula Enable Selective Fine-tuning
Kai J Sandbrink ⋅ Mia Whitefield ⋅ Sebastian Bruijns ⋅ Jirko Rubruck ⋅ Laurence T Hunt ⋅ Fazl Barez ⋅ Christopher Summerfield
Transformers follow implicit curricula whereby some tasks are learned before others. However, how explicit pretraining curricula influence learning, generalization, and the selectivity of fine-tuning is unclear. This is important for AI safety, where fine-tuning is used to selectively suppress misaligned behaviors. Here, we compare curricula that pretrain tasks in a balanced (sampled uniformly) or an imbalanced (one task early, the other late) fashion. We show that imbalanced learning of two conflicting copy tasks promotes in-context learning and improves the selectivity of refusal fine-tuning. Ablations and activation patching show that this occurs because imbalanced pretraining encourages tasks to be disentangled in separable neural circuits, whereas balanced training routes both tasks through a common pathway. We extend these findings to a synthetic language learning task involving rule-consistent and rule-violating data, where imbalanced curricula similarly lead to more localized, less entangled rule representations, resulting in more robust rule-following behavior. Together, these results suggest that imbalanced pretraining curricula may be an important tool for promoting disentangled representations, with direct consequences for the precision and reliability of safety fine-tuning.
PRIME: Poincaré return induced measure for learning partially observed dynamical systems
Yiting Duan ⋅ Longyan Tan ⋅ Hao Wu ⋅ Mehdi Tavakol ⋅ Yi Guo
Learning dynamics from sparse partial observations presents a fundamental dilemma: the pointwise mean squared error (MSE) training objective may induce local contraction and empirically lead to topological collapse, whereas matching the system's invariant measure preserves global geometry but lacks temporal ordering constraints. To resolve this, we introduce the $\textbf{P}$oincaré $\textbf{R}$eturn $\textbf{I}$nduced $\textbf{M}$easur$\textbf{e}$ $(\textbf{PRIME})$ framework. PRIME bridges this gap by adopting a novel perspective of recurrence-level temporal organization, motivated by the return map and roof time structure of suspension flow representations. Instead of relying on unstable pointwise matching, PRIME refines the learned dynamics by leveraging recurrence-level temporal information, specifically by aligning the empirical graph measure of the induced return map over recurrence blocks. We evaluate our framework on the partially observed Lorenz63 system, sensor-observed Kuramoto-Sivashinsky equation, and flow past a cylinder. PRIME effectively reduces topological collapse across our benchmarks, improving temporal consistency while maintaining or enhancing long-term geometric and spectral fidelity.
PRIM: Meta-Learned Bayesian Root Cause Analysis
Christopher Lohse ⋅ Anish Dhir ⋅ Amadou Ba ⋅ Bradley Eck ⋅ Marco Ruffini ⋅ Jonas Wahl
Root cause analysis (RCA) in complex systems is challenging due to error propagation across multiple variables, the need for structural causal knowledge, and the computational cost of inference at test time. We introduce PRIM (Prior-fitted Root cause Identification with Meta-learning), a causal meta-learning approach that frames RCA as a Bayesian inference task over a synthetic prior of causal models. By marginalising out structural uncertainty, PRIM implicitly identifies changes in the data-generating mechanism between baseline and anomalous periods. In doing so, PRIM infers distributional differences without explicit statistical testing, and implicitly learns causal structure without model fitting at test time. Following the simulation-based meta-learning paradigm of prior-fitted networks, PRIM uses a Model-Averaged Causal Estimation (MACE) transformer neural process that jointly attends over observational and anomalous samples and the causal structure of nodes, enabling zero-shot inference in 17\,ms for systems with up to 100 variables. Across synthetic benchmarks and two realistic benchmark datasets, PetShop and CausRCA, PRIM is competitive with methods that are aware of the system's causal graphical structure a priori while outperforming graph-unaware methods on several tasks. Lightweight fine-tuning to specific domains and data dynamics improves performance further.
Prism Attention: Proposal-Refined Index Sharing Mechanism for Efficient LLMs Inference
Zhenxu Tian ⋅ Kebin Liu ⋅ Zhengwu Yang ⋅ Yi Su ⋅ Qingqing Dang ⋅ Kaipeng Deng ⋅ Yanlin Sha ⋅ Yanjun Ma ⋅ Dianhai Yu ⋅ Juntao Li ⋅ Min zhang
Long-context decoding in large language models (LLMs) is increasingly bottlenecked by the memory and computation required to attend over an ever-growing Key-Value (KV) cache at each autoregressive step. Although token-level sparse attention reduces this burden by restricting computation to a small subset of relevant past tokens, its practical gains are often offset by repeatedly recomputing layer-wise sparse indices as the context grows. A common remedy is to share sparse indices across layers by computing them only at designated anchor layers, but such static reuse ignores the layer-wise evolution of attention patterns, causing the reused indices to progressively deviate from the layer-specific optima and incur substantial attention loss. By revisiting cross-layer sparsity from the perspective of attention dynamics, we find that inter-layer salient token shift is highly localized: although the exact top-$\kappa$ indices differ across layers, the core high-attention tokens remain stable, with changes concentrated near the expanded boundary of previously selected regions. Motivated by this, we propose Prism Attention, a proposal-refined sparse attention framework. It replaces rigid cross-layer index reuse with a coarse-to-fine strategy, preserving layer-wise adaptability while retaining the efficiency benefits of cross-layer sharing. Extensive experiments show that Prism Attention achieves state-of-the-art performance among mainstream sparse attention methods on both reasoning and long-context benchmarks, while achieving $2.6\times\sim7.5\times$ kernel speedup over FlashAttention at 128K context length.
PRISM:Disentangling Preference Distributions for Generative Ranking
zhangkai wu ⋅ Kaize Shi ⋅ Xu Zhang ⋅ Zhihong Cui ⋅ Longbing Cao
Preference ranking modeling is important for selecting and adapting language-model responses to human preferences. Compared with deterministic ranking methods, generative preference ranking models the prompt-response data distribution, providing richer uncertainty characterization and sample diversity for learning a more expressive ranking boundary. Recent work improves modeling efficiency by shifting generative preference ranking from textual space to embedding space, but without explicit preference-conditioned separation, the learned embedding distribution may still entangle preferred and rejected responses, producing ambiguous synthetic samples that weaken the ranking boundary. To address this issue, we propose PRISM, a Preference Ranking framework through dISentangled embedding Modeling. PRISM formulates embedding-level generative preference ranking with a unified class-conditional ELBO and decomposes this optimization objective to expose the encoded latent entanglement between preferred and rejected embeddings. This derivation motivates two practical preference-aware generation variants: PRISMMI, involving a deep class-separation objective, and PRISMMMD, catering a probabilistic aggregate-matching objective. Both variants learn a preference-disentangled embedding data distribution and synthesize data pairs that preserve ranking semantics for the generative ranking boundary. Experiments show that PRISM improves preference ranking performance and produces useful generated embeddings for downstream response selection. The anonymous code repository is available at:https://anonymous.4open.science/r/PRISM-915C/README.md.
PRISM: Human Point Cloud Reconstruction via Skeleton-Guided Diffusion from MmWave Radar
Jiacheng Huang ⋅ Long Tian ⋅ Yuan Liu ⋅ Andy Khong
Millimeter-wave(mmWave) radar is an attractive modality for human sensing, offering all-weather, non-contact, and privacy-preserving perception. However, its inherent sparsity severely limits downstream human-centric understanding, and there is currently no public dataset that provides paired quantitative benchmarks for dense human point cloud reconstruction under occlusion. We present MIST, the first mmWave human point cloud dataset with paired quantitative benchmarks for through-obstacle dense human reconstruction. MIST captures synchronized clear-view and through-obstacle observations of the same subjects performing identical actions, separated by a physical barrier, and provides LiDAR-based 3D ground truth. We further incorporate a three-level occlusion design that converts attenuation severity into a controlled variable, covering eight subjects, thirty action classes, and 640~k paired frames. To establish a baseline on MIST and address the limitations of point cloud completion under extreme sparsity and human motion, we propose PRISM, a skeleton-guided conditional latent diffusion framework for reconstructing dense human point clouds from sparse mmWave radar alone. Three conditioning streams—skeleton joint positions, coarse body geometry, and action class embeddings—are incorporated to guide a Dynamic Transformer within a VP-SDE framework, enabling effective denoising and the reconstruction of mmWave human point clouds with near-LiDAR quality. On MIST, PRISM achieves a COV-CD of 0.362, significantly outperforming the strongest completion baseline (0.113). Notably, PRISM maintains structurally coherent reconstruction even under severe through-obstacle attenuation at approximately 50 points per frame, whereas all baselines collapse to fixed-template outputs.
PRISM: Principal Subspace Alignment for Parameter-Efficient Fine-Tuning
Anh Tong ⋅ Kyowoon Lee ⋅ Jaesik Choi
Parameter-efficient fine-tuning (PEFT) methods such as low-rank adaptation have become essential for adapting large language models to downstream tasks. Existing approaches often apply low-rank updates uniformly across all layers, ignoring the structural coupling inherent in transformer architectures. We observe that transformers contain \emph{weight pairs} whose interactions govern model behavior including query-key products, value-output products, and feedforward networks. We propose PRincipal Interaction Subspace Matching (PRISM), a principled framework that exploits this coupling through spectral parameterization. For each weight pair, we parameterize adaptations using the spectral basis of the paired matrix, ensuring that updates occur in the subspace where the paired matrix has maximal effect. We validate our approach through controlled experiments on synthetic models, demonstrating faster convergence. Empirical evaluations on arithmetic reasoning, code generation, and vision tasks show that PRISM outperforms existing PEFT methods.
ProAlign: Progressive Positional and Prototype-Guided Alignment for Aerial-Ground Person Re-Identification
Qiao Li ⋅ Jing Chen ⋅ Cong Wu ⋅ Kun He ⋅ Ruiying Du ⋅ Lei Tan
Aerial-Ground Person Re-Identification (AG-ReID) aims to match pedestrian images captured by unmanned aerial vehicles (UAVs) and ground-based surveillance cameras. The task remains highly challenging due to severe viewpoint discrepancies, frequent occlusions, and substantial domain gaps between aerial and ground imagery. Beyond these observable factors, existing methods often overlook two fundamental structural issues: cumulative positional drift during hierarchical feature extraction and semantic inconsistency across cross-view feature distributions. To address these challenges, we propose ProAlign, a Progressive Positional and Prototype-Guided Alignment Network for AG-ReID. ProAlign comprises two key components: a Layer-wise Progressive Positional Embedding (LPPE) module and a View-Aware Prototype Contrastive Learning (PCL) module. LPPE performs adaptive positional calibration throughout transformer layers by employing conditional positional encoding in shallow layers to mitigate local spatial distortions, while introducing learnable static positional embeddings in deeper layers to reinforce global semantic priors. Meanwhile, PCL maintains view-specific aerial and ground prototypes for each identity and enforces cross-view semantic consistency via a dual-prototype contrastive objective. Extensive experiments on the challenging CARGO benchmark demonstrate the effectiveness of ProAlign. In the aerial-ground setting, ProAlign surpasses previous state-of-the-art methods by 13.89% mAP and 9.39% Rank-1 accuracy.
Tiny Recursive Models (TRM) solve complex reasoning tasks with a fraction of the parameters of modern large language models (LLMs) by iteratively refining a latent state and final answer. While powerful, their deterministic recursion can lead to convergence at suboptimal solutions, without escape mechanism. A common workaround relies on task-specific input perturbations at test time combined with answer aggregation via voting. We introduce Probabilistic TRM (PTRM), a task-agnostic framework for test-time compute scaling that addresses this limitation through stochastic exploration. PTRM injects Gaussian noise at each deep recursion step, enabling parallel trajectories to explore diverse solution basins, and selects among them using the model’s existing Q head (used for early stopping in the original TRM). Without requiring retraining or task-specific augmentations, PTRM enables substantial accuracy gains across benchmarks, including Sudoku-Extreme (87.4% to 98.75%) and on various puzzles from Pencil Puzzle Bench (65% to 91%). On the latter, PTRM achieves nearly double the accuracy of frontier LLMs (91% vs. 55%) at less than 0.0001x the cost, using only 7M parameters.
Probing Visual Planning in Image Editing Models
Zhimu Zhou ⋅ Yanpeng Zhao ⋅ Qiuyu Liao ⋅ Bo Zhao ⋅ Xiaojian (Shawn) Ma
Visual planning represents a crucial facet of human intelligence, especially in tasks that require complex spatial reasoning and navigation. Yet, in machine learning, this inherently visual problem is often tackled through a verbal-centric lens. While recent research demonstrates the promise of fully visual approaches, they suffer from significant computational inefficiency due to the step-by-step planning-by-generation paradigm. In this work, we present EAR, an editing-as-reasoning paradigm that reformulates visual planning as a single-step image transformation. To isolate intrinsic reasoning from visual recognition, we employ abstract puzzles as probing tasks and introduce AMAZE, a procedurally generated dataset that features the classical Maze and Queen problems, covering distinct, complementary forms of visual planning. The abstract nature of AMAZE also facilitates automatic evaluation of autoregressive and diffusion-based models in terms of both pixel-wise fidelity and logical validity. We assess leading proprietary and open-source editing models. The results show that they all struggle in the zero-shot setting, finetuning on basic scales enables remarkable generalization to larger in-domain scales and out-of-domain scales and geometries. However, our best model that runs on high-end hardware fails to match the zero-shot efficiency of human solvers, highlighting a persistent gap in neural visual reasoning.
ProEdit: Inversion-based Editing From Prompts Done Right
Zhi Ouyang ⋅ Dian Zheng ⋅ Xiao-Ming Wu ⋅ Jian-Jian Jiang ⋅ Kun-Yu Lin ⋅ Jingke Meng ⋅ Wei-Shi Zheng
Inversion-based visual editing provides an effective and training-free way to edit an image or a video based on user instructions. Existing methods typically inject source image information during the sampling process to maintain editing consistency. However, this sampling strategy overly relies on source information, which negatively affects the edits in the target image (e.g., failing to change the subject's atributes like pose, number, or color as instructed). In this work, we propose ProEdit to address this issue both in the attention and latent aspects. In the attention aspect, we introduce KV-mix, which mixes KV features of the source and the target in the edited region, mitigating the influence of the source image on the editing region while maintaining background consistency. In the latent aspect, we propose Latents-Shift, which shifts the distribution of the edited region in the inverted noise, eliminating the negative influence of the inverted noise on the sampling process. Extensive experiments on several image and video editing benchmarks demonstrate that our method achieves SOTA performance. In addition, our design is plug-and-play, which can be seamlessly integrated into existing inversion and editing methods, such as RF-Solver, FireFlow and UniEdit.
ProGraf: Profile-Guided Planning for Step-by-Step Generation of Structured Non-Natural Images
Ran Luo ⋅ Haoxiang Deng ⋅ Lan Zhang ⋅ Mu Yuan
Structured visual artifacts such as geometry diagrams, circuit schematics, and flowcharts are increasingly important targets for image generation, yet visual plausibility alone is insufficient---a single missing object, wrong relation, broken connection, or extra element can invalidate the output. Existing image generators remain brittle on such tasks, and generic step-by-step or agentic methods often fail because intermediate edits may corrupt previously correct structures and ignore generator-specific failure modes. To address this, we present ProGraf, a profile-guided framework for structured non-natural image generation. We first build a benchmark of 304 tasks across geometry, circuit, flowchart, and general diagrams, each annotated with fine-grained verification points. From repeated generator outputs, we derive model-specific capability profiles capturing category-level reliability, component-level weaknesses, and recurrent failure patterns. ProGraf uses these profiles to guide closed-loop decomposition, verifies each intermediate result, commits only accepted states, and replans after failures. On generator-specific hardest subsets, ProGraf recovers 43.7--77.8\% of previously failed tasks across generators and categories, outperforming fixed-step and adapted agent baselines; ablation studies further show that benchmark-derived profiles are a key source of the gain.
Prompt Ensemble Image Purification for Test-time Adversarial Robustness of CLIP
Yuxuan Zhang ⋅ Shuchang Wang ⋅ Zhenbo Shi ⋅ Xiaoman Liu ⋅ Jiajun Hu ⋅ Zhidong Yu ⋅ Wei Song ⋅ Wei Yang
Vision-Language Models like CLIP exhibit remarkable zero-shot capabilities but remain vulnerable to adversarial attacks. Existing defenses, such as adversarial fine-tuning or test-time defense, either incur high computational costs that lead to cause catastrophic forgetting, or struggle against strong adversarial attacks. In this work, we reveal that adversarial perturbations are highly overfitted to the decision boundary of a single canonical text prompt. By introducing several semantic prompt variants, we identify a \textit{Prompt Sensitivity Gap}: adversarial examples exhibit significantly higher prediction variance across prompts compared to benign images. Motivated by this insight, we propose Prompt Ensemble Image Purification (PEIP), an efficient test-time defense framework. PEIP features a dual-objective purification loop that jointly suppresses prediction variance to dismantle adversarial alignment and reinforces the most responsive class across prompt variants to facilitate semantic recovery. Furthermore, to accelerate inference, we introduce a statistically calibrated Early Exit mechanism that bypasses benign images based on their initial prompt sensitivity. Extensive experiments across 16 classification benchmarks, multiple CLIP architectures, and CLIP-based zero-shot semantic segmentation tasks demonstrate that PEIP achieves state-of-the-art adversarial robustness while preserving exact zero-shot accuracy and high efficiency.
Prompting Diffusion Models for Zero-Shot Instance Segmentation
İrem Z Alagöz ⋅ Nils Morbitzer ⋅ Andrea Ramazzina ⋅ Nassir Navab ⋅ Federico Tombari ⋅ Stefano Gasperini
Several disruptive research directions have recently emerged in computer vision, including foundation models achieving previously unseen zero-shot performance in scene understanding, even interactively, and generative models that synthesize extremely realistic images. The latter have also been shown to be highly effective in scene understanding tasks thanks to their rich priors. However, for promptable segmentation, foundation models struggle with accurately segmenting an object's region, leading to false positives and over-segmentation. Notably, early attempts that leverage generative priors use prompts only during post-processing, yielding suboptimal segments because the process is agnostic to the user input. In this paper, we target these limitations with Prompt2Seg, a spatial conditioning framework for diffusion-based segmentation. Prompt2Seg augments a frozen diffusion segmentation model with a conditioning branch. Our approach takes spatial prompts, represented as 2D Gaussians or confidence maps, as explicit input signals, training the model to respond directly to user intent. Fine-tuned on a deliberately constrained set of object categories drawn from Hypersim and Virtual KITTI 2, Prompt2Seg generalizes zero-shot to a wide range of unseen object types and visual domains. We evaluate on seven datasets ranging from standard benchmarks to more challenging domains, including paintings, egocentric views, and X-ray data. Furthermore, we demonstrate that Prompt2Seg consistently outperforms the underlying diffusion segmentation backbone across all benchmarks. Our results suggest that the rich priors encoded in generative pretraining, combined with principled spatial conditioning, offer a compelling path toward broadly generalizing interactive segmentation without large-scale mask supervision.
Propagation of Chaos in Contextual Flow Maps
Chen ⋅ Zhengjiang Lin ⋅ Kaizhao Liu ⋅ Philippe Rigollet
We develop a quantitative statistical theory of transformers in the large-context regime by adopting the abstraction of contextual flow maps: dynamical systems that evolve a distinguished token in the presence of a contextual measure across a stack of attention blocks. Within this framework, the finite-context model approximates an idealized infinite-context system in which the contextual measure is replaced by its underlying population, so that the context length $n$ becomes a statistical resource. Exploiting the McKean--Vlasov structure of the dynamics and the classical machinery of propagation of chaos, we establish a forward bound controlling the deviation between the finite- and infinite-context flow maps uniformly along depth, and a backward bound controlling the deviation between the corresponding training trajectories uniformly across iterations of online gradient descent. Both bounds achieve the optimal Wasserstein rate $n^{-1/d}$. The analysis rests on a new Eulerian adjoint formulation of the loss gradient and stability estimates for the resulting forward--adjoint system, both of which may be of independent interest. We further verify that the standard transformer architecture satisfies these estimates with Lipschitz constants independent of embedding and parameter dimensions, suggesting a structural reason why transformers train stably across model scales.
Proprio: Latent Self-Scoring and Inference-Time Refinement for Physically Plausible Video Generation
Mariam Hassan ⋅ Kaouther Messaoud ⋅ WUYANG LI ⋅ Alexandre Alahi
Modern video generative models produce visually impressive results, yet frequently violate basic physical principles. We propose \textbf{Proprio}, a training-free framework that enables a frozen video generator to assess and improve the physical plausibility of its own outputs. Inspired by \textit{proprioception}, the biological sense of one’s own movement, Proprio treats the model's flow residual under controlled latent perturbations as a self-scoring signal. Samples that are better explained by the generator's learned dynamics induce smaller and more stable residuals. We aggregate this signal across timesteps and perturbations, focus it on motion-relevant regions with a dynamic spatiotemporal mask, and use it for best-of-N search, gradient-based self-refinement, or both. Across text-to-video and image-to-video benchmarks, Proprio consistently improves physical plausibility, outperforming VLM-based scoring, and external world-model baselines in several settings. On TurboWan2.2, Proprio improves Physics-IQ from 32.2 to 37.5 (+16.5%) and VideoPhy2-hard physical commonsense from 45.6 to 55.0 (+20.6%). Human evaluation further shows that raters prefer Proprio-selected or refined videos for physical plausibility in roughly two-thirds of comparisons. These results suggest that frozen video generators contain actionable internal signals for evaluating and improving the physical plausibility of their own outputs.
Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction–Reality Gaps
Jiaxin Zhang ⋅ XIANGYU PENG ⋅ Qinglin Chen ⋅ Yu Li ⋅ Hiroaki Hayashi ⋅ Chien-Sheng Wu
Reinforcement learning for long-horizon agents relies on \emph{purely retrospective} training signals: credit is assigned only after observing environmental consequences, leaving the agent's belief at action time invisible to the gradient. We introduce \textbf{Prospective Hindsight (PH)}, a self-calibrating training principle that augments any retrospective base method with a signal derived from the gap between the agent's \emph{prospective prediction} (before feedback) and the \emph{retrospective evaluation} (after feedback). This per-rollout \emph{surprise} identifies samples where the agent's self-model is most inaccurate and amplifies their gradient contribution through a stop-gradient surprise-weighted advantage. Since the prospective predictor shares parameters with the policy, the two co-evolve, progressively shifting focus to the agent's remaining blind spots. We connect this principle to a privileged-information gap and show that minimizing the surprise residual provides a descent pathway on the agent's miscalibration rate; calibration thus emerges as a byproduct of optimization rather than from an added objective. On single-turn verifiable tasks and a multi-turn personal-agent task (under GRPO, on-policy distillation, and their combination), PH improves both task performance and calibration, with consistent gains across model scales. Notably, the dominant miscalibration mode shifts structurally between regimes, overconfident failures in single-turn, underconfident successes in multi-turn, yet the same training principle addresses both successfully.
Electronic health records (EHRs) are a valuable resource for clinical prediction because they capture rich longitudinal information about patient care. Existing graph-based EHR models often learn relations at the cohort level, but do not clearly distinguish between global event associations and the subset of interactions actually instantiated within an individual visit. To address this limitation, we propose MVP-EHR, a multi-view prototype learning framework for EHR prediction with shared type-aware prototype alignment. MVP-EHR combines three key ideas: (i) a lift-thresholded global knowledge graph that captures cohort-level event associations, (ii) a visit-induced local knowledge graph that preserves only the observed events and supported event-event relations within each visit, and (iii) shared type-aware prototypes that align local and global event representations before post-event multi-view fusion and temporal prediction. Experiments on MIMIC-IV show that MVP-EHR achieves strong and consistent performance across multiple clinical prediction tasks, including 0.9339 AUROC and 0.1706 AUPRC for mortality prediction, the best AUROC on readmission prediction, the best overall performance on phenotype prediction, and a strong AUPRC/F1 tradeoff on drug recommendation among the compared baselines. Ablation results further show that separating global prior from visit-instantiated evidence, together with shared prototype alignment, is critical for robust performance.
Provably Efficient Regularized Online RLHF with Generalized Bilinear Preferences
Junghyun Lee ⋅ Minju Hong ⋅ Kwang-Sung Jun ⋅ Chulhee Yun ⋅ Se-Young Yun
We consider the problem of *regularized* best-response max-regret minimization in online RLHF under general preferences and bandit feedback. While various regularizers are utilized to robustify alignment, known polylogarithmic regret guarantees remain heavily specific to KL. To investigate whether such fast rates extend beyond KL, we adopt the *Generalized Bilinear Preference Model (GBPM)*—capturing intransitive preferences over $d$-dimensional item-wise features via a rank-$2r$ skew-symmetric matrix—to isolate the impact of generic regularization. Crucially, under GBPM, we prove that the dual gap of any greedy policy is bounded by the *squared* estimation error, derived using *only* strong convexity and skew-symmetry. Under a feature coverage assumption, we establish polylogarithmic $\tilde{\mathcal{O}}(\eta d^4 (\log T)^2 \wedge d^2 \sqrt{T})$ regret with Greedy Sampling and $\mathrm{poly}(d)$-free $\tilde{\mathcal{O}}(\sqrt{\eta r T} \wedge r^{1/3} T^{2/3})$ regret with Explore-Then-Commit, where $\eta^{-1}$ is the regularization coefficient and $T$ is the time horizon. This demonstrates that ``fast'' regrets are *not* KL-specific, but rather a fundamental consequence of generic strongly convex geometry.
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping
Qiming Li ⋅ Tianlun Li ⋅ Xiaolong Cheng ⋅ Hangyu Li ⋅ Ruiyan Gong ⋅ Kangning Niu ⋅ kaitao jiang ⋅ Mu Xu
Reinforcement Learning with Verifiable Rewards (RLVR) has become an effective paradigm for improving the reasoning capability of Large Vision-Language Models (LVLMs). However, existing RLVR methods primarily rely on trajectory-level outcome rewards, which assign identical learning signals across all generated tokens. This coarse-grained credit assignment is fundamentally mismatched to multimodal reasoning, where only a sparse subset of tokens is causally grounded in visual evidence. Consequently, these pivotal perceptual tokens receive weak supervision and are often overwhelmed by language priors or reasoning-template tokens. To address this limitation, we propose Perception-Reinforced Policy Optimization (PRPO), a token-level reinforcement learning framework that explicitly identifies and reinforces pivotal perceptual tokens within long-horizon multimodal reasoning trajectories. PRPO introduces Robust Visual Dependency (RVD), a principled metric that identifies tokens whose predictions are both visually grounded and perturbation-stable, filtering out brittle or noisy visual tokens. Based on RVD, we further propose Perceptual Advantage Reshaping (PAR), a token-level credit assignment technique that amplifies perceptually informative tokens while preserving stable gradients for non-perceptual tokens. Extensive experiments on seven multimodal reasoning benchmarks demonstrate that PRPO consistently outperforms strong LVLM baselines across both 3B and 7B model scales, achieving average gains of 23.3\% and 21.1\%, respectively. PRPO achieves state-of-the-art performance with improved training efficiency and stronger cross-task generalization. Our findings highlight the importance of fine-grained credit assignment for scalable multimodal reinforcement learning.
pyCD: A Unified Benchmark for Reliable Evaluation of Cognitive Diagnosis Models
Youheng Bai ⋅ Xueyi Li ⋅ Tengteng Cheng ⋅ Mingliang Hou ⋅ Teng Guo ⋅ Jiaqi Zheng ⋅ Zitao Liu
Cognitive diagnosis (CD) infers students' mastery of knowledge components (KCs) from their responses, supporting personalized feedback and exercise recommendation. Cognitive diagnosis models (CDMs) are often ranked by response-prediction AUC under inconsistent experimental protocols, while their KC-level mastery profiles are not consistently audited under a shared protocol. We introduce \texttt{pyCD}, a unified benchmark of eleven CDMs on six public datasets that fixes preprocessing, splits, hyperparameter search, and stopping rules; the benchmark reports KC-level diagnostic metrics for learned mastery alongside standard prediction metrics. Two findings show that gains in AUC do not translate into gains in diagnostic quality. First, the AUC leaderboard is sharply compressed and sensitive to protocol choices: modern CDMs often cluster within $0.005$--$0.020$ AUC, with the per-dataset best non-IRT CDM improving over IRT by only $0.017$ AUC on average; removing validation-AUC early stopping can shift AUC by up to $0.13$ on individual cells. Second, the AUC gaps that remain do not reliably transfer into diagnosis: the top-1 AUC model is typically outside the top-3 by KC-level diagnostic agreement (DOA). The same mismatch extends to exercise difficulty: against Junyi's human pairwise difficulty ratings, learned CDM difficulty parameters rank worse than both an empirical accuracy-based baseline and zero-shot LLMs given only exercise titles. CDM evaluations should therefore report diagnostic outputs directly rather than infer them from predictive gains; \texttt{pyCD} provides a unified protocol for doing so. Code, splits, and configurations are available at \url{https://anonymous.4open.science/r/pyCD-NeurIPS2026-60B8}.
Pygmalion: Bridging Reconstruction and Generation in Sparse Voxel-based 3D Modeling
Guan Luo ⋅ Jing Lin ⋅ Xuanyu Yi ⋅ Jiahang Liu ⋅ Song-Hai Zhang ⋅ Jianfeng Zhang
Recent advances in sparse voxel-based VAEs have demonstrated remarkable capabilities in high-fidelity 3D autoencoding, yet their generative counterparts consistently lag behind their reconstruction performance. A fundamental cause of this gap lies in the \emph{brittle output representation} in sparse voxel decoders, which parameterize geometry through discrete topological decisions, such as intersection flags or occupancy signs. Small perturbations in these parameters, which are unavoidable under diffusion sampling, can be amplified into abrupt topological ruptures and grid-like artifacts. We argue that a generation-friendly representation should ensure that such perturbations induce only smooth geometric transitions. To this end, we propose Pygmalion, a generation-friendly autoencoding framework that reformulates sparse voxel decoding as a coupled SDF parameterization. To retain explicit mesh-level supervision within this parameterization, which is non-trivial as the SDF jointly governs both topology and geometry, we introduce hinge-based sign correction, case-aware geometry supervision, and rendering-based refinement, achieving high-fidelity reconstruction while preserving generative robustness. Built upon this generation-friendly representation, Pygmalion scales consistently across DiT model sizes and voxel resolutions. Extensive experiments demonstrate state-of-the-art reconstruction fidelity and high-quality 3D generation with complete surfaces, sharp features, and fine geometric details.
Pygmalion Effect in Vision: Image-to-Clay Translation for Reflective Geometry Reconstruction
Gayoung Lee ⋅ Junho Kim ⋅ Jin-Hwa Kim ⋅ Junmo Kim
Understanding reflection remains a long-standing challenge in 3D reconstruction due to the entanglement of appearance and geometry under view-dependent reflections. In this work, we present the Pygmalion Effect in Vision, a novel framework that metaphorically “sculpts” reflective objects into clay-like forms through image-to-clay translation. We then introduce a dual-branch design in which a BRDF-based reflective branch and a clay-guided branch share the same Gaussian geometry but operate through independent rendering paths. The clay-guided branch, supervised by the synthesized clay images, provides reflection-free geometric guidance that complements the photometric supervision of the BRDF branch. Experiments on both synthetic and real datasets show consistent improvements in both geometry and photometric metrics. Beyond technical gains, our framework reveals that seeing by unshining, translating radiance into neutrality, can serve as a powerful inductive bias for reflective object geometry learning.
QB-Highlights: Quality-Guided Budgeted Highlight Detection in Videos with Dense Query-Relevant Moments
chao gao ⋅ Hui Li ⋅ Ying Chen ⋅ Zongyao He ⋅ Sibin Deng
Highlight detection in content-dense untrimmed videos, such as livestreams, is particularly challenging due to the prevalence of semantically similar moments and the application-driven demand of fitting a target duration budget. Existing methods typically rank segments independently based on semantic relevance, struggling to discriminate fine-grained quality among similar moments and allocating the budget to redundant selections. We therefore reformulate highlight detection as jointly selecting high-quality and complementary segments under a duration budget. To study this formulation, we introduce QB-Highlights, a large-scale benchmark of 50,000 e-commerce livestream videos with densely occurring query-relevant events and subtle quality variations. QB-Highlights is organized around two complementary dimensions: Segment Qualification, which assesses fine-grained per-segment quality; and Budgeted Composition, which evaluates set-level complementarity under a duration budget. It provides MLLM-assisted pseudo labels for training, dense human annotations for testing, and an evaluation protocol jointly measuring quality and diversity. We further propose a Qualification-then-Composition Reasoning (QCR) framework that performs structured deliberative reasoning to produce a highlight set matching the target duration. Extensive experiments show that QCR outperforms strong MLLM, moment retrieval, and highlight detection baselines on both segment quality and compositional diversity.
Q-CoMove: Differentiable Quantum Circuit Priors for Coordinated Motion in Multi-Component Embodied Systems
Ke Shi ⋅ Fanqi Kong ⋅ Yuchen Wang ⋅ Zhipeng Liu ⋅ Tingting Li ⋅ Ziming Zhao
Multi-component embodied robots must coordinate the base, arm, camera, gripper, and local environment under shared spatial and temporal constraints. We present Q-CoMove, a coordinated motion framework that models this platform as a coupled dynamical system. Q-CoMove uses a differentiable circuit-style prior implemented with the TorchQuantum simulator on classical hardware: component states are encoded into a five-qubit register, cross-component dependence is parameterized by a 79-gate simulated circuit, and 25 measured statistics are projected into a coordination latent that conditions a physics-residual dynamics model and a joint action generator. We position this module as a compact structured coordination bottleneck rather than a claim of quantum advantage or physical necessity. Across parameter-matched Transformer and GNN baselines, residual-dynamics and joint MPC baselines, ablations, and robustness stress tests in simulation, Q-CoMove improves coordination metrics with bounded inference cost. All empirical results are limited to simulated environments, and the current study does not isolate the circuit prior against a structure-matched classical bottleneck.
QUTCC: Quantile Uncertainty Training and Conformal Calibration for Imaging Inverse Problems
Cassandra T Ye ⋅ Shamus Li ⋅ Tyler King ⋅ Kristina Monakhova
While deep learning offers tremendous promise for scientific and medical imaging, any failures and hallucinations (predictions that do not coincide with reality) are hard to pinpoint and can have serious downstream consequences. Uncertainty estimation techniques, such as conformal prediction, can help by predicting statistically valid error bars for a model's prediction. However, popular conformal prediction methods were not designed for high-dimensional image-valued problems and do not take into account spatial correlations within an image during conformal calibration, resulting in larger-than-necessary uncertainty intervals. We propose a practical simultaneous quantile regression method that enables non-linear, spatially-adaptive scaling during conformal calibration. Our method, QUTCC uses a U-Net architecture with a quantile embedding to learn a full conditional quantile distribution during training, and then leverages this non-linear, learned function for spatially-adaptive conformal calibration. At test time, our method can efficiently estimate uncertainty intervals with pixel-marginal coverage guarantees. In addition, QUTCC can also predict pixel-wise conditional probability density estimates without any built-in distributional assumptions. We evaluate our method on several denoising problems, accelerated magnetic resonance imaging, and quantitative phase microscopy. Our method consistently produces tighter uncertainty intervals than prior conformal methods at the same coverage level, can predict plausible conditional distributions for different tasks, and in some cases, high-uncertainty regions can help us locate hallucinations in a model's prediction.
RABBiT: Rapidly adaptive BOLD foundation model via brain-tuning for accurate zero-shot and few-shot prediction of speech-elicited responses in the brain
Omer Moussa ⋅ Mariya Toneva
Language understanding in the brain is context-dependent, varying across experimental stimuli and individuals, which makes it difficult to build computational models that generalize across both. This calls for a foundation model of language-evoked brain activity that can capture shared structure while adapting efficiently to new participants and inputs. We introduce RABBiT (Rapidly Adaptive BOLD foundation model via BraIn-Tuning), a compact audio-to-fMRI encoder designed for accurate zero- and few-shot prediction. A comprehensive evaluation on 324 participants across multiple unseen fMRI datasets shows that RABBiT enables accurate zero-shot prediction of fMRI responses to natural speech across auditory and language-selective regions, surpassing the SOTA foundation model for fMRI and predictions based on group averages. With as little as 10 minutes of participant-specific data, RABBiT further improves performance via parameter-efficient tuning, substantially outperforming per-participant linear models. RABBiT's performance is driven by two key innovations: (1) learned region-specific attention, and (2) a decomposition of brain responses into shared and subject-specific components, combined with a brain-tuned speech backbone. In addition to supporting strong predictive accuracy, the structured, region-specific representations that RABBiT learns enable interpretability. By eliminating the need for extensive per-participant data and model fitting, RABBiT enables scalable population-level analyses of language in the human brain.
RadarMAE: Injecting Physical Inductive Biases into Masked Autoencoders for Advanced Radar Object Detection
EunChan Kim ⋅ Jeongwan Shin ⋅ Jae-Ho Choi
Masked Autoencoders (MAE) have emerged as a powerful self-supervised paradigm in vision, yet their direct application to radar perception remains bottlenecked by a fundamental domain gap. Specifically, while conventional MAEs are optimized for RGB images using isotropic grid patching and standard mean squared error (MSE) loss, radar range-azimuth (RA) heatmaps exhibit sinc-oriented anisotropic spatial structures as well as complex physical artifacts, such as heavy-tailed multipath interference and sidelobe spikes. In this paper, we propose RadarMAE, the first radar-centric MAE tailored specifically to the physical properties of radar signals. We present an axial tokenization strategy that preserves the resolution disparities and continuous sinc-shaped lobes of RA maps. Furthermore, considering that the range and azimuth dimensions exhibit distinct physical artifacts, we decouple their reconstruction objectives by introducing a multipath distribution loss along the range axis to capture heavy-tailed multipath reflections as well as a sidelobe distribution loss along the azimuth axis to learn deterministic sidelobe spikes. Our extensive experiments demonstrate that injecting these domain-specific physical priors enables robust representation learning for radar object detection. Notably, our method achieves a state-of-the-art 64.17% bird's-eye-view (BEV) AP on the K-Radar dataset using only a single-frame RA map setting, significantly outperforming the previous baselines.
RAM-H1200: A Unified Evaluation and Dataset on Hand Radiographs for Rheumatoid Arthritis
YANG SONGXIAO ⋅ Haolin Wang ⋅ Yao Fu ⋅ Junmu Peng ⋅ Lin Fan ⋅ Hongruixuan Chen ⋅ JIAN SONG ⋅ Masayuki Ikebe ⋅ Shinya Takamaeda-Yamazaki ⋅ Masatoshi Okutomi ⋅ Tamostu Kamishima ⋅ yafei ou
Rheumatoid arthritis (RA) assessment from hand radiographs requires multi-level analysis and modeling of anatomical structures and fine-grained local pathological changes. However, existing public resources do not support such unified multi-level analysis, often lacking full-hand coverage, fine-grained annotations, and consistent integration with clinical scoring systems. In particular, annotations that enable quantitative analysis of bone erosion (BE) remain scarce. RAM-H1200 contains 1,200 hand radiographs collected from six medical centers, with multi-level annotations including (i) whole-hand bone structure instance segmentation, (ii) pixel-level BE masks, (iii) SvdH-defined joint regions of interest, and (iv) joint-level SvdH scores for both BE and joint space narrowing (JSN). It is designed to evaluate whether models can jointly capture anatomical structure, localized erosive pathology, and clinically standardized RA severity from hand radiographs. The proposed BE masks enable, for the first time, quantitative BE analysis beyond coarse categorical grading by providing explicit spatial supervision for lesion extent and morphology. To our knowledge, RAM-H1200 is the first public large-scale benchmark that jointly supports whole-hand bone structure instance segmentation, pixel-level BE delineation, and clinically grounded joint-level SvdH scoring for both BE and JSN. Results across benchmark tasks show that anatomical modeling is substantially more mature than quantitative BE analysis: whole-hand bone segmentation achieves strong performance, whereas BE segmentation remains a major open challenge. By unifying anatomical structure modeling, quantitative lesion analysis, and clinically grounded SvdH scoring, RAM-H1200 provides a single benchmark for comprehensive RA analysis on hand radiographs. Benchmark & Code: https://github.com/YSongxiao/RAM-H1200 Dataset Repository: https://huggingface.co/datasets/TokyoTechMagicYang/RAM-H1200-v1
Random features for Grassmannian kernel approximation with bounded rank-one projections
Remi Delogne ⋅ Laurent Jacques
We propose a family of random feature maps for scalable kernel machines defined over low-dimensional subspaces in high dimensions, i.e. over the Grassmannian manifold. This is typically useful in a machine learning context when data classes or clusters are well represented by the span of a few data points. Classical Grassmannian kernels such as the projection or Binet–Cauchy kernels require constructing full Gram matrices for practical applications, leading to prohibitive computational and memory costs for large subspace datasets in high dimensions. We address this limitation by computing specific random features of subspaces. These combine random rank-one projections of the subspace projection matrices with bounded non-linear transforms—periodic or binary—to tame the resulting heavy-tailed distribution.
We show that, in the random feature space, inner products approximate well-defined, rotation-invariant Grassmannian kernels, i.e. depending only on the principal angles of the considered subspaces. Provided the number of random features is large compared to the subspace intrinsic dimension, we show that this approximation holds uniformly over all subspaces of fixed dimensions with high probability.
When the non-linear transform is periodic, the approximated kernel admits a closed-form expression with a tunable behaviour bridging inverse Binet–Cauchy and Gaussian-type regimes, while the binarised feature has no known closed-form kernel but lends itself to even more compactly represented one-bit subspace features.
Moreover, we show how structured rank-one projections, leveraging randomised fast Fourier transforms, further reduce the random feature computational complexity without sacrificing accuracy in practical experiments. We demonstrate the practicality of these techniques with synthetic experiments and classification tasks on the ETH-80 dataset representing visual object images from different viewpoints. The proposed random features recover Grassmannian geometry with high accuracy while reducing computation, memory, and storage requirements. This demonstrates that rank-one embeddings offer a practical and scalable alternative to classical Grassmannian kernels.
Random Matrix Theory of Early-Stopped Gradient Flow: A Transient BBP Scenario
Florentin Coeurdoux ⋅ Grégoire Ferré ⋅ jean-philippe bouchaud
Empirical studies of trained models often report a transient regime in which signal is detectable in a finite gradient descent time window before overfitting dominates. We provide an analytically tractable random-matrix model that reproduces this phenomenon for gradient flow in a linear teacher--student setting. In this framework, learning occurs when an isolated eigenvalue separates from a noisy bulk, before eventually disappearing in the overfitting regime. The key ingredient is anisotropy in the input covariance, which induces fast and slow directions in the learning dynamics. In a two-block covariance model, we derive the full time-dependent bulk spectrum of the symmetrized weight matrix through a $2\times 2$ Dyson equation, and we obtain an explicit outlier condition for a rank-one teacher via a rank-two determinant formula. This yields a transient Baik--Ben Arous--P\'ech\'e (BBP) transition: depending on signal strength and covariance anisotropy, the teacher spike may never emerge, emerge and persist, or emerge only during an intermediate time interval before being reabsorbed into the bulk. We map the corresponding phase diagrams and validate the theory against finite-size simulations. Our results provide a minimal solvable mechanism for early stopping as a transient spectral effect driven by anisotropy and noise.
Random Neural Network Expressivity for Non-Linear Partial Differential Equations
Muhammed Ali Mehmood ⋅ Lukas Gonon
Neural networks with randomly generated hidden weights (RaNNs) have been extensively studied, both as a standalone learning method and as an initialization for fully trainable deep learning methods. In this work, we study RaNN expressivity for learning solutions to non-linear partial differential equations (PDEs). Despite their widespread use in practical applications, a rigorous theoretical understanding of the approximation properties of RaNNs in this context remains limited. Here, we derive error bounds for RaNN approximations to time-dependent Sobolev functions and obtain a dimension-free approximation rate $\frac{1}{2}$ for sufficiently regular functions. We apply our results to two important classes of non-linear PDEs: Porous Medium Equations and Compressible Navier-Stokes Equations, showing that RaNNs are capable of efficiently approximating solutions to these complex, non-linear PDEs. Our theoretical analysis is supported by numerical experiments, validating the obtained convergence rates.
Rank-Constrained Adaptation for Reliable Real-World Performance
Abinitha Gourabathina ⋅ Hyewon Jeong ⋅ Teya Bergamaschi ⋅ Marzyeh Ghassemi ⋅ Collin Stultz
Deep learning models trained to optimize average accuracy often exhibit systematic failures on particular subpopulations. In real-world settings like healthcare, the subpopulations most affected by such disparities are frequently unlabeled, partially observed, or not known in advance. Existing group-robust methods typically assume prior knowledge of the relevant subgroups, using group annotations for training, validation, or model selection. We propose Misclassification Aware Rank-Limited Adaptation (MARLA), a parameter-efficient method for improving worst group performance without explicit subgroup annotations. MARLA leverages an ERM-trained model by calculating the model's misclassification probability scores on a held-out adaptation set to identify a low-dimensional subspace where errors concentrate. We then learn a rank-restricted additive correction to the classifier logits within that subspace. Across seven real-world datasets, we evaluate group robustness under three settings: no knowledge of subgroup relevance, partial knowledge of subgroup relevance, and full knowledge of subgroup relevance. MARLA improves worst-group performance while remaining fast, parameter-efficient, and practical to tune.
RAPDrive: Shared-Latent Hybrid Decoding for Reasoning and Planning in Autonomous Driving
Yuheng Liu ⋅ Le Hui ⋅ Ziyue Zhu ⋅ Yigong Zhang ⋅ Jin Xie ⋅ jian Yang
Driving vision-language-action (VLA) models must connect semantic reasoning with executable trajectory planning, but language and motion exhibit different generation structures: reasoning text is naturally sequential, whereas future trajectories require horizon-level geometric consistency. We present RAPDrive, a shared-latent hybrid-decoding framework that processes visual context, ego state, motion history, reasoning tokens, and future planning slots within a single transformer sequence. RAPDrive generates reasoning text autoregressively while refining future motion through iterative masked denoising in a dedicated motion-token space, followed by continuous trajectory realization. This design couples reasoning and planning through shared latent computation while maintaining separate tokenizations for language and motion. Training combines supervised reasoning-and-planning objectives, and GRPO is evaluated as an optional post-training extension. Experiments on NAVSIM and Bench2Drive show that RAPDrive outperforms prior non-world-model driving VLA baselines in the supervised/base setting, benefits further from planner-centric post-training, and is supported by ablations on the main architectural choices.
Realtime-VLA FLASH: Speculative Inference Framework for Diffusion-based VLAs
Jiahui Niu ⋅ Kefan Gu ⋅ Yucheng Zhao ⋅ shengwen Liang ⋅ Tiancai Wang ⋅ Xing Hu ⋅ ying wang ⋅ Huawei Li
Diffusion-based vision-language-action models (dVLAs) are promising for embodied intelligence but are fundamentally limited in real-time deployment by the high latency of full inference. We propose Realtime-VLA FLASH, a speculative inference framework that eliminates most full inference calls during replanning by introducing a lightweight draft model with parallel verification via the main model's Action Expert and a phase-aware fallback mechanism that reverts to the full inference pipeline when needed. This design enables low-latency, high-frequency replanning without sacrificing reliability. Experiments show that on LIBERO, FLASH largely preserves task performance by replacing many 58.0\,ms full-inference rounds with speculative rounds as fast as 7.8\,ms, lowering task-level average inference latency to 19.1\,ms (3.04$\times$ speedup). We additionally demonstrate effectiveness on real-world conveyor-belt sorting, highlighting its practical impact for latency-critical embodied tasks.
ReasoningShield: Safety Moderation over Reasoning Traces of Large Reasoning Models
Changyi Li ⋅ Jiayi Wang ⋅ Xudong Pan ⋅ Geng Hong ⋅ Min Yang
Large Reasoning Models (LRMs) leverage explicit reasoning traces, known as Chain-of-Thought (CoT) reasoning, to decompose complex problems into intermediate steps before deriving final answers. However, these reasoning traces introduce unique safety challenges: harmful content can be embedded in intermediate steps even when final answers appear benign. Our study reveals that existing state-of-the-art moderation tools experience significant performance degradation on CoT moderation, with F1 scores dropping by up to 35.2% compared to traditional answer moderation. To address these challenges, we present ReasoningShield, a comprehensive benchmark and a suite of strong lightweight models for CoT safety moderation. Our benchmark includes 9.2K annotated Query-CoT pairs with structured stepwise risk analysis, consisting of a 7K-sample training set and a 2.2K human-annotated test set, spanning 10 risk categories, 3 safety levels, and 8 LRMs with diverse reasoning paradigms. To further demonstrate the utility of the benchmark, we develop lightweight models with 1B and 3B parameters using a two-stage training strategy. Our models achieve 91.8% F1 on our benchmark, substantially outperforming leading tools such as LlamaGuard-4 by 35.6% and commercial models such as GPT-4o by 15.8%, while generalizing effectively across diverse reasoning paradigms and unseen scenarios. All resources, including the dataset, code, and models, are released at https://anonymous.4open.science/r/ReasoningShield.
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
Jeonghye Kim ⋅ Jiwon Jeon ⋅ Dongsheng Li ⋅ Yuqing Yang
Self-distillation has emerged as a powerful framework for post-training LLMs, where a teacher conditioned on extra information guides a student without it, both from the same model. While this guidance is useful when the student has failed, on successful rollouts, the same mechanism instead overwrites the student's choices and suppresses it's own reasoning. Therefore, we propose reading the original self-distillation signal in reverse: when the student succeeds along a path the teacher would not have predicted, these tokens reflect its self-driven reasoning. Building on this, we propose RLRT (RLVR with Reversed Teacher), which augments GRPO by reinforcing these tokens on correct rollouts. We interpret this as a new form of exploration in RLVR: not uniform diversity, but valuable exploration grounded in the student's own success. Across base, instruction-tuned, and thinking-tuned Qwen3 checkpoints, RLRT substantially outperforms self-distillation and exploration-based baselines, establishing information asymmetry as a new, principled design axis for RLVR.
RecMem: Recurrent Memory Compression for Long-Sequence Recommendation
Haoyi Hu ⋅ Gaoyang Guo ⋅ Yunjia Xi ⋅ Wen Chen ⋅ Weiwen Liu ⋅ Jian Wu ⋅ Yuning Jiang ⋅ Weinan Zhang
Sequential recommenders need long histories, but full-prefix attention scales with user lifetime. We argue that an effective memory mechanism should provide space invariance, semantic fidelity, and lifelong evolvability, which existing fixed-window or capped-memory methods do not jointly satisfy. We propose RecMem, a recurrent compression framework that folds history segments into M forget-gated memory slots. RecMem has a history-length-independent rollout error bound, and its fixed-capacity memory is trained by next-item prediction to preserve useful signals. On MerRec, RecMem uses 5.6x fewer decoder-side tokens than full attention while retaining 88–97% of Recall@10–200. Compared with a compressed truncation baseline at similar decoder cost, RecMem improves Recall@50+ by 4–6 percentage points by compressing the full history rather than discarding old interactions. It also remains stable over 1,800 test steps.
Reconciling Causality and Non-Equilibrium Thermodynamics with Hamiltonian Causal Models
Dario Rancati ⋅ Max Welling ⋅ Francesco Locatello
Causal modeling of physical temporal phenomena must handle interventions that act along trajectories, nonstationary induced laws, path-dependent effects, and feedback mediated by dynamics, all challenging in standard causal models. We introduce \emph{Hamiltonian Causal Models} (HCMs), a trajectory-level framework in which observed variables interact with local environments and interventions act as controls of Hamiltonian mechanisms. HCMs separate immutable equations of motion from intervenable mechanisms and define causal effects as discrepancies between interventional path laws. A key motivation for HCMs is their natural interface with non-equilibrium thermodynamics. Entropy production quantifies the irreversibility of a process and is a central causal observable: it is estimable from data and witnesses causal effects along the system's evolution that are invisible to endpoint and cumulative versions of the standard average treatment effect. As in physics, cause-and effect are not primitives of the relation between two random variables but arise from the non-invertibility of the thermodynamic arrow. With this, our paper reconciles the language of statistical causal models and non-stationary thermodynamics, offering new tools to describe causality in a wide range of physical systems.
Reconsidering Positional Supervision in Masked Diffusion Language Model Training
Mengyu Ye ⋅ Keito Kudo ⋅ Ryosuke Takahashi ⋅ Jun Suzuki
Masked diffusion language models (MDLMs) generate text by unmasking tokens in parallel and have recently emerged as alternatives to autoregressive language models. They can be viewed as parallel decoders trained with a position-wise cross-entropy (CE) loss, the same setup as non-autoregressive translation (NAT). In NAT, CE-trained parallel decoders have been argued to be sensitive to small positional shifts, since CE penalizes them harshly. We ask whether CE-trained MDLMs are similarly sensitive to such shifts under iterative decoding. To probe this, we apply a controlled intervention that introduces them during decoding. On LLaDA-8B-Instruct with Arena-Hard, displacing as little as 1\% of generated tokens by one position substantially reduces win rates against the unintervened model, showing that MDLMs are sensitive to such small shifts under iterative parallel decoding. Motivated by this, we adapt connectionist temporal classification (CTC), an alignment-flexible objective known to mitigate it there, to MDLM supervised fine-tuning. By relaxing the strict position-wise match that CE imposes, CTC gives the loss room to absorb small positional shifts; concretely, we modified CTC objective to use a special \ token that absorbs positional uncertainty between target tokens and output positions, and a updated collapse map that preserves target surface forms. Across four open-ended generation benchmarks, the resulting model consistently improves over both the original model and a matched cross-entropy-trained baseline, with statistically significant gains on all four. These results identify training-side alignment flexibility as a useful design dimension for MDLM SFT, complementary to the inference-time approaches explored in prior work.
Reconstructing the Vocal Tract with Differentiable Acoustic Simulation
Eric Chen ⋅ Jin Woo Lee ⋅ Vincent Sitzmann
The human vocal tract, the cavity consisting of one's throat, mouth, lips opening, etc., filters one's voice to create the sounds we know as speech. In this paper, we present a differentiable and GPU parallelizable acoustic simulator that synthesizes speech by propagating sound along an acoustic tube representation of the vocal tract, and via its gradients, solves the inverse problem: reconstructing the geometry of their vocal tract solely from the sound it produces. Although learning the inverse mapping between geometry and sound is notoriously non-convex, we discover that gradient descent succeeds with three technical contributions: (1) we design a frequency domain formulation of the vocal tract's fluid dynamics that is 70x more GPU parallelizable than finite differences in time, (2) we integrate a differentiable model for turbulence to synthesize consonants, and (3) similar to prior work in implicit neural representations (INRs) and neural radiance fields (NeRFs), we find that parameterizing the geometry with a neural network accelerates convergence and escapes local minima that trap discrete representations. Because the simulator is differentiable, it is readily integrated with other deep learning pipelines to enable novel speech and medical imaging applications. Our model can be used for instance, in singing instruction or language learning, where a visualization of one's vocal tract can help people understand how their vocal tract maps to different speech sounds. To enable these tasks, (1) we demonstrate self-supervised autoencoding of vocal tract shapes across 11 languages, and (2) we couple our simulator with a generative model of MRI images to reconstruct one's moving vocal tract from only their speech, no paired data required.
Reduced Cost Influence Functions for Predict-then-Optimize under Noisy Data
William Zhang ⋅ Saurabh Amin ⋅ Georgia Perakis
Machine learning models deployed in decision-making pipelines, from vehicle routing to dynamic pricing to clinical decision support, must remain reliable in the face of low-quality training data. We study a predict-then-optimize setting in which an upstream model predicts the cost vector of a downstream linear program, with a large pool of noisy training data supplemented by a small curated set of clean data. We introduce Reduced Cost Influence Functions (RCIFs), which combine influence functions from robust statistics with the geometry of linear optimization to estimate how each training point shifts downstream decisions. Building on this, we propose a reweighting algorithm that uses RCIFs to downweight harmful training points and reduce downstream regret. Experiments on random linear programs and shortest-path problems show that our approach reduces regret, and accompanying theory characterizes properties of RCIFs and identifies regimes favorable to our method relative to baselines.
RedVLA: Physical Red Teaming for Vision-Language-Action Models
Yuhao Zhang ⋅ Borong Zhang ⋅ Jiaming Fan ⋅ Jiachen Shen ⋅ Yishuai Cai ⋅ Yaodong Yang ⋅ Jiaming Ji
The real-world deployment of Vision-Language-Action (VLA) models remains limited by the risk of unpredictable and irreversible physical harm. However, we currently lack effective mechanisms to proactively detect these physical safety risks before deployment. To address this gap, we propose \textbf{RedVLA}, the first red teaming framework for physical safety in VLA models. We systematically uncover unsafe behaviors through a two-stage process: (I) \textbf{Risk Scenario Synthesis} constructs a valid and task-feasible initial risk scene. Specifically, it identifies critical interaction regions from benign trajectories and positions the risk factor within these regions, aiming to entangle it with the VLA's execution flow and elicit a target unsafe behavior. (II) \textbf{Risk Amplification} ensures stable elicitation across heterogeneous models. It iteratively refines the risk factor state through gradient-free optimization guided by trajectory features. Experiments on six representative VLA models show that RedVLA uncovers diverse unsafe behaviors and achieves the ASR up to 95.5\% within 10 optimization iterations. To mitigate these risks, we further propose SimpleVLA-Guard, a lightweight safety guard built from RedVLA-generated data. Our assets and code are available at https://anonymous.4open.science/r/redvla.
Refactoring Code Through Library Design
Žiga Kovačič ⋅ Justin Chiu ⋅ Celine Lee ⋅ Wenting Zhao ⋅ Kevin Ellis
Maintainable and general software allows developers to build robust applications efficiently, yet achieving these qualities often requires refactoring specialized solutions into reusable components. This challenge becomes particularly relevant as code agents are increasingly used to solve isolated one-off programming problems. We investigate code agents' capacity to refactor code in ways that support consolidation and reusability. We first investigate what makes a good refactoring, finding via simulation results and a human study that proxies of program size, such as Minimum Description Length, better predict preferable refactorings than extant software engineering metrics, such as Maintainability Index. We then present \textsc{Librarian}, a method that combines divide-and-conquer with sample-and-rerank to generate reusable libraries. We compare \textsc{Librarian} to state-of-the-art library generation methods, and study it on real-world Python programs.
Refining Compositional Diffusion for Reliable Long-Horizon Planning
Kyowoon Lee ⋅ Yunhao Luo ⋅ Anh Tong ⋅ Jaesik Choi
Compositional diffusion planning generates long-horizon trajectories by stitching together overlapping short-horizon segments through score composition. However, when local plan distributions are multimodal, existing compositional methods suffer from mode-averaging, where averaging incompatible local modes leads to plans that are neither locally feasible nor globally coherent. We propose Refining Compositional Diffusion (RCD), a training-free guidance method that steers compositional sampling toward high-density, globally coherent plans. RCD leverages the self-reconstruction error of a pretrained diffusion model as a proxy for the log-density of composed plans, combined with an overlap consistency term that enforces consistency at segment boundaries. We show that the combined guidance concentrates sampling on high-density plans that mitigate mode-averaging. Experiments on challenging long-horizon tasks from OGBench, including locomotion, object manipulation, and pixel-based observations, demonstrate that RCD consistently outperforms existing methods. Project website at https://refining-compositional-diffusion.github.io/.
ReGDiff: Guided Diffusion in Regulated Latent Space for Exploring Metamaterial Voxel Geometry
Wangzhi Zhan ⋅ Jianpeng Chen ⋅ Dongqi Fu ⋅ Dawei Zhou
Metamaterials are artificially engineered structures whose mechanical and physical behaviors are strongly shaped by geometry rather than composition. Voxel representation provides a unified format for metamaterial geometry generation, as it can express diverse classes such as truss, shell, and porous structures within a single cubic discretization. However, voxel-based generation faces a plausibility–novelty trade-off: staying close to known geometries helps preserve geometric regularities, while moving away from them is necessary for novelty but may produce degenerate geometries. To address this challenge, we propose ReGDiff, a generative framework that couples voxel representation with latent space regulation and guided diffusion. ReGDiff introduces a repel-and-sink (RAS) mechanism to smooth the latent distribution of plausible geometries, and short-range repulsion (SRR) guidance to discourage generation overly close to known samples while maintaining geometric plausibility. We further contribute a voxel-based benchmark covering truss- and shell-type metamaterial geometries, together with an evaluation module for geometric plausibility, novelty, and diversity. Experiments show that ReGDiff outperforms voxel-based generative baselines, achieving +8.9% in geometric plausibility, +46.4% in novelty, and +128.6% in diversity on average across two datasets. These results suggest that ReGDiff is a strong geometry candidate generator for downstream evaluation. Our code is provided at https://anonymous.4open.science/r/ReGDiff-6DC6.
Wasserstein distributionally robust optimization (DRO) is a prominent paradigm for robust decision-making in the face of distributional uncertainty. Given a nominal data distribution, DRO selects a decision which minimizes the risk for a worst-case data distribution within a prescribed Wasserstein radius $\varepsilon$ of the nominal distribution. In this paper, we investigate the quality of DRO decisions when evaluated on a worst-case distribution within a 2-Wasserstein $\varepsilon$-neighborhood of the nominal distribution, as measured by excess risk or ex-ante regret. Beginning with multivariate linear regression and extending to ridge regression, we first identify the asymptotic complexity of robust decision-making in the $\varepsilon \to 0$ limit, characterized by a problem-specific condition number $\kappa$. In this regime, a wide spectrum of simple algorithms, including DRO, achieve the instance-optimal rate of $O(\kappa \varepsilon^2)$. For fixed $\varepsilon > 0$, we prove that DRO still achieves the minimax rate if the problem is of an appropriate ``low rank''. On the other hand, we identify significant failure modes where DRO is provably suboptimal by dimension-dependent factors. To resolve this, we introduce a new condition number reduction (CNR) procedure which achieves the optimal rate and admits tractable approximation algorithms. We support these theoretical results with numerical experiments comparing the performance of various estimators including DRO and CNR.
ReGuidance: Diffusion Steering with Strong Latent Initializations Solves Hard Inverse Problems
Aayush Karan ⋅ Kulin Shah ⋅ Sitan Chen
In recent years there has been a flurry of activity around using pretrained diffusion models as informed data priors for solving inverse problems, and more generally around steering these models towards certain reward models. Training-free methods like gradient guidance have offered simple, flexible approaches for these tasks, but when the reward is not informative enough, e.g., in inverse problems with highly compressive measurements, these techniques can veer off the data manifold, failing to produce realistic data samples. To address this challenge, we devise a simple algorithm, *ReGuidance*, that leverages prior methods' solutions as strong initializations and substantially enhancing their realism. Given a candidate solution $x$ produced by a given method, we propose inverting the solution by running the unconditional probability flow ODE in reverse starting from $x$, and then using the resulting latent as an initialization for a deterministic steering process. Empirically, we evaluate our algorithm on difficult image restoration tasks including large box inpainting, heavily downscaled superresolution, and high noise deblurring with both linear and nonlinear blurring operations. We find that, using a wide range of baseline methods as initializations, applying our method results in much stronger samples with better realism and measurement consistency. We complement these results with rigorous proofs in commonly studied theoretical settings showing our technique boosts the reward and brings $x$ closer to the data manifold.
Reinforced Fast Weights via Next-Sequence Prediction
Hee Seung Hwang ⋅ Xindi Wu ⋅ Sanghyuk Chun ⋅ Zhiwei Deng ⋅ Olga Russakovsky
Fast weight architectures offer a promising alternative to standard transformers for long-context modeling by replacing the KV-cache with a recurrently updated fixed-size memory. Despite this architectural shift, they are typically trained with the same next-token prediction (NTP) objective as standard transformers. This creates a mismatch: fast weight models rely on an evolving memory state to support future predictions, while NTP provides only token-level supervision for the immediate next token. We address this mismatch with next-sequence prediction (NSP), a sequence-level extension of NTP that trains models to produce coherent multi-token continuations from their recurrent memory state. To optimize this sequence-level objective, we propose ReFINE ($\textbf{Re}$inforced $\textbf{F}$ast we$\textbf{I}$ghts via $\textbf{N}$ext s$\textbf{E}$quence prediction). ReFINE selects informative token positions based on prediction entropy, generates multi-token rollouts, assigns sequence-level rewards, and optimizes the model with Group Relative Policy Optimization (GRPO). ReFINE is applicable throughout the training lifecycle of pre-trained language models: mid-training, post-training, and test-time training. Experiments on LaCT-760M and DeltaNet-1.3B show that ReFINE consistently outperforms NTP-based supervised fine-tuning in needle-in-a-haystack retrieval, long-context question answering, and diverse tasks in LongBench. ReFINE provides an effective and versatile framework for improving long-context modeling in fast weight architectures.
REINS: A Self-Evolving Agent Harness for Real-Time Trajectory Planning
Zhihong Cui ⋅ Hengyu Liu ⋅ Haoran Tang ⋅ shijun liu ⋅ Amir Taherkordi ⋅ Tor Skeie
Agent harnesses—runtime scaffolds that wrap an LLM's reasoning loop with tool dispatch, verification, and persistent memory—have become the standard substrate for deploying LLMs as autonomous agents, as demonstrated by Claude Code and Codex in software engineering. In trajectory planning, however, this paradigm has not taken hold: existing LLM-based trajectory planning methods either keep the LLM in the decision loop, exceeding the millisecond-level control budget, or offload it to offline rules or training-time supervision—leaving the runtime planner non-evolving. Intermediate hybrids still lack the verification-and-consolidation loop that makes harnesses reliable. We ask: *what runtime harness would let an LLM planner meet real-time constraints while self-evolving into a deployable one?* We propose **REINS**, a harness built on three LLM–planner couplings: LLM outputs are (i) **grounded** to a learned variable-level dynamics model, (ii) **verified** via calibrated predicates and forward rollout, and (iii) **consolidated** into an indexed skill memory that serves future scenes in $\mathcal{O}(\log n)$ without re-invoking the LLM. The harness closes a self-evolution loop: familiar scenes resolve by memory lookup; novel scenes invoke the LLM, whose verified outputs enrich the memory. Without modifying LLM weights, REINS attains 98%+ compliance, 50%+ collision reduction, and 2–5 ms latency on simulated and real-world benchmarks, with LLM invocation rate monotonically declining as memory matures. Code: https://anonymous.4open.science/r/REINS-1FE3/
Releasing Anchors from Cross-View Correspondence: Probabilistic Multi-View Anchor Graph Clustering
Zhoumin Lu ⋅ Yongbo Yu ⋅ Yu Duan ⋅ Feiping Nie ⋅ Yicong Zhou
Anchor graphs are a standard route to scalable multi-view clustering, but most existing pipelines still depend on either a shared anchor set sampled from concatenated features or a separate anchor-alignment stage before fusion. Both choices are restrictive: the former hard-codes equal anchor budgets and importance at anchor-construction time, while the latter introduces a cumbersome correspondence problem and still struggles with unequal anchor numbers. We develop an alignment-free probabilistic formulation in which view-specific anchor graphs are noisy manifestations of a shared latent cluster identity rather than objects that require cross-view correspondence. Each view has its own anchor space, number of anchors and cluster prototypes; the only shared latent variable is the sample label. We instantiate this idea with a Gaussian latent-label model and a deterministic prototype M-step. The resulting variational EM algorithm admits closed-form updates with convergence guarantees, while experiments demonstrate competitive clustering performance and practical runtime on most datasets.
Reliability-Budgeted Edge–Cloud Adaptation for Continual Multimodal Dehazing on UAV
Junwei Zhao ⋅ Qianchun Luo ⋅ Jie Wu
UAV dehazers trained offline degrade during long-duration flights as haze density and spatial structure evolve over time. We formulate UAV dehazing as a closed-loop edge-cloud adaptation problem with an explicit reliability budget governing update frequency. The onboard model couples a frozen anchor head with a cloud-adaptive head, evaluated by a joint critic integrating a physics-based rehaze constraint and a scale-aligned infrared-edge consistency term to suppress degenerate adaptations. Our main contribution is a sequential upload criterion based on a non-negative e-process, yielding anytime-valid false-upload control for streams that satisfy conditional calibration, without requiring frame independence or exchangeability. The cloud updates only the adaptive head from uploaded segments, while the UAV accepts or rejects returned parameters through a shadow-window sign test with an acceptance budget. Experiments on five UAV datasets show consistent improvements under diverse haze shifts. Code will be publicly released.
RepoZero: Can LLMs Generate a Code Repository from Scratch?
Zhaoxi Zhang ⋅ Yiming Xu ⋅ Jiahui Liang ⋅ Weikang Li ⋅ Xiaoshuai Chen ⋅ Liwei Qian ⋅ Xin Pei ⋅ Jizhou Huang ⋅ Rui Sun ⋅ Yunfang Wu
Large Language Models (LLMs) have recently shown remarkable progress in code generation, yet their ability to construct complete software repositories from scratch remains poorly understood. A fundamental bottleneck is the lack of verifiable and scalable evaluation: existing benchmarks either focus on patch-based editing or rely on human or LLM-based judgments, which introduce bias and limit reproducibility. In this work, we present RepoZero, the first benchmark that enables fully automated, execution-based verification of repository-level generation from scratch. Our key idea is to reformulate generation as repository reproduction: given only API specifications, an agent must re-implement an entire repository such that its behavior matches the original implementation. This design allows for strict black-box validation via output equivalence, while naturally supporting large-scale construction by reusing existing open-source repositories. To further mitigate data leakage and shortcut solutions, we introduce cross-language constraints and a sandboxed evaluation protocol. Building on this benchmark, we propose an Agentic Code-Test Evolution (ACE) framework that performs iterative test generation and error-driven refinement, enabling effective test-time scaling for repository-level synthesis. Extensive experiments across multiple state-of-the-art LLMs and agent frameworks reveal that even the strongest models achieve only limited pass rates (30\% - 55\%), exposing a substantial gap between current capabilities and real-world software development requirements. Our results establish RepoZero as a challenging, scalable, and reliable testbed for end-to-end code generation, and highlight self-verification via test generation as a critical direction for advancing LLM-based coding agents.
ResKV: Residual-based Channel-wise Unstructured Pruning for KV Cache Compression
Yue Chen ⋅ Jinze Li ⋅ Dajiang Liu
In modern LLM applications, the demand for extended context lengths continues to surge, making KV cache a major performance bottleneck in long-context inference. Channel-wise pruning has emerged as a prevalent KV cache compression paradigm; however, existing solutions overlook inherent similarities among KV vectors, resulting in limited pruning ratios and noticeable performance degradation. To tackle this issue, we present ResKV, a novel residual-driven channel-wise pruning framework that fully exploits KV vector similarity for effective compression. ResKV first clusters highly similar KV vectors via cosine similarity metrics, then calculates residual vectors by subtracting each cluster’s centroid from the original features. Subsequent channel-wise unstructured pruning and per-token quantization are exclusively applied to these residual components. During inference, preserved residuals are combined with their corresponding centroids to accurately recover dense KV representations. Benefiting from the low magnitude and sparse, near-zero properties of centroid residuals, both pruning and quantization incur only negligible approximation errors. Experimental results show that ResKV reduces peak memory consumption by 45.6% and improves inference throughput by 4.27× against dense FlashAttention inference, while sustaining competitive model performance. Under identical sparsity constraints, ResKV outperforms state-of-the-art channel-wise KV pruning methods by achieving up to 2.26 points higher accuracy and 2.63× faster end-to-end throughput.
We study Minty set inclusion on a compact convex set $\mathcal{X} \subseteq \mathbb{R}^d$ with a set-valued operator $A : \mathcal{X} \rightrightarrows \mathbb{R}^d$. Under the Minty promise, the goal is to compute $x \in \mathcal{X}$ and $u \in A(x)$ such that $\sup\_{y \in \mathcal{X}} \langle u, x - y \rangle \le \varepsilon$. Our approach is geometric. We work in the fixed-scale constrained-resolvent oracle model, where a query at anchor $a$ is assumed to return a resolvent point $x^+$ together with the Yosida vector $g_{\eta}(a) = (a - x^+)/\eta$. We prove that each oracle response yields a sharp dichotomy: either the returned point provides an $\varepsilon$-SVI certificate, or the Yosida vector defines a strict separating normal for the Minty set. Thus, a single black-box resolvent query supplies either a certified candidate solution or a valid geometric cut. Combining this separation principle with the ellipsoid method, we obtain \textsc{Resolvent-Ellipsoid}, an algorithm that computes an $\varepsilon$-SVI in $\mathcal{O}\left(\operatorname{poly}(d, \log(1/\varepsilon))\right)$ black-box oracle calls to the fixed-scale constrained resolvent. For monotone inclusions, our algorithm gives, to our knowledge, the first $\mathcal{O}\left(\mathrm{poly}(d, \log(1/\varepsilon))\right)$ black-box guarantee in the fixed-scale resolvent model.
We give a computer-assisted counterexample to the open question posed by Rudin, Schapire, and Daubechies in COLT 2012, of whether exhaustive AdaBoost always converges to a finite cycle. The construction is based on a block-product gadget whose two factors share an exact period-2 orbit for their 5-step branch maps, but whose linearized return maps have dominant eigenvalues with an irrational logarithmic ratio. This irrationality forces the burst-winner sequence to have an irrational asymptotic frequency, precluding eventual periodicity. All assertions are certified by exact rational arithmetic in two independent computer algebra systems. We also document the collaborative workflow that led to the construction, in which one model produced the key technical arguments; another model summarized and critiqued different versions of that work, proposing new directions to pursue; and the authors orchestrated the process by selecting directions to prioritize, resolving ambiguities, heavily refining each argument, and verifying the final solution. We present this as one of the first in-depth case studies of LLM-assisted mathematical research in which a long-standing open problem from theoretical machine learning is resolved, and detail the advantages and challenges of such research.
Rethinking Dataset Distillation for Classification: Do Distilled Sets Outperform Coresets?
Trisha Mittal ⋅ Akshay Mehra ⋅ Joshua Kimball
Dataset distillation (DD) has emerged as a prominent approach in data centric machine learning, aiming to synthesize compact training sets for efficient training by compressing the information in large datasets into a small number of synthetic samples. However, DD methods are often evaluated under inconsistent evaluation protocols, ranging from standard ERM to single/multi‑teacher supervision, making it difficult to isolate the effectiveness of distilled data from evaluation. Moreover, many prior methods claim that DD outperforms data pruning approaches such as coreset selection (CS), based on the assumption that restricting condensed datasets to subsets of real samples fundamentally limits their expressiveness. In this work, we critically evaluate DD methods through large-scale experiments using standardized datasets and evaluation protocols to assess their intrinsic effectiveness. We benchmark seven state-of-the-art DD methods on ImageNet-1K, ImageNet-100, and ImageNette, using three widely adopted training protocols against three CS strategies. Our results show that while some DD methods fail to outperform even simple random subsets, the SOTA DD approaches are comparable to or worse than coresets on large‑scale datasets and incur a substantially higher cost for construction. Beyond accuracy, we also evaluate the representativeness, diversity, and quality of condensed sets, and find that coresets consistently achieve better coverage of the original data distribution. These findings highlight the limited practical advantages of current DD methods and show that coresets remain competitive and are often a more computationally efficient alternative for data-centric learning.
Rethinking Post-Training Recipes for Multimodal Time-Series Forecasting
Haoxin Liu ⋅ Yichen Zhou ⋅ Rajat Sen ⋅ B. Aditya Prakash ⋅ Abhimanyu Das
Time-Series Foundation Models (TSFMs) excel at zero-shot unimodal forecasting using numerical data, but unlike LLMs they cannot consume multimodal, non-numerical context that often shape real-world trajectories. In this work, we bridge this gap and argue for a multimodal time-series forecasting approach that post-trains LLMs to act as context-guided revisors over strong numerical TSFM priors. We introduce PostTime, a post-training recipe combining Supervised Fine-Tuning (SFT) and Reinforcement Learning with Verifiable Rewards (RLVR), along with a methodology to generate automated reasoning traces for forecast revisions. PostTime teaches an LLM to generate context-conditioned forecast interventions—decisions to revise, preserve, or ignore the TSFM prior based on the multimodal context. We evaluate this approach on the TimesX multimodal forecasting benchmark using a Gemma-3-4B LLM and TimesFM-2.5 TSFM, and show that it significantly outperforms standalone TSFMs, LLM-only baselines, and existing multimodal forecasting approaches. Our post-training recipe including code and dataset, can currently be accessed for review purposes at \url{https://anonymous.4open.science/status/posttime-21BF/}.
Rethinking Reward Models for Multi-Domain Test-Time Scaling
Dong Bok Lee ⋅ Seanie Lee ⋅ Sangwoo Park ⋅ Minki Kang ⋅ Jinheon Baek ⋅ Dongki Kim ⋅ Dominik Wagner ⋅ Jiongdao Jin ⋅ Heejun Lee ⋅ Tobias Bocklet ⋅ Jinyu Wang ⋅ Jingjing Fu ⋅ Sung Ju Hwang ⋅ Jiang Bian ⋅ Lei Song
The reliability of large language models (LLMs) during test-time scaling is often assessed with external verifiers or reward models that distinguish correct reasoning from flawed logic. Prior work has studied both outcome reward models (ORMs), which assess only the final answer, and process reward models (PRMs), which score intermediate reasoning steps. Although PRMs are often viewed as advantageous due to their finer-grained supervision, much of the supporting evidence comes from math-adjacent settings, and their relative benefits across broader domains remain unclear. We present the first unified evaluation of four reward model variants, discriminative ORM and PRM (DisORM, DisPRM) and generative ORM and PRM (GenORM, GenPRM), across 14 diverse domains. Contrary to conventional wisdom, we find that (i) DisORM performs on par with DisPRM, (ii) GenPRM is not competitive, and (iii) overall, GenORM is the most robust, yielding significant and consistent gains across every tested domain. We attribute this to PRM-style stepwise scoring, which inherits label noise from LLM auto-labeling and has difficulty evaluating long reasoning trajectories, including those involving self-correcting reasoning. Our theoretical analysis shows that step-wise aggregation compounds errors as reasoning length grows, and our empirical observations confirm this effect. These findings challenge the prevailing assumption that fine-grained supervision is always better and support generative outcome verification for multi-domain deployment. We publicly release our code at this https://github.com/db-Lee/Multi-RM to facilitate future research in multi-domain settings.
Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks
William Shen ⋅ Xinchi Qiu ⋅ Chenxi Whitehouse ⋅ Lisa Alazraki ⋅ Shashwat Goel ⋅ Francesco Barbieri ⋅ Timon Willi ⋅ Akhil Mathur ⋅ Ilias Leontiadis
Recently, rubrics have been used to guide LLM judges in capturing nuanced, multi-dimensional human preferences, and have further been extended as reward signals for reinforcement fine-tuning (RFT). However, open-ended generation typically lacks a unique canonical target, and its evaluation rubrics are latent rather than directly observed. This makes rubric generation under-determined and difficult to control: rubrics often lack coverage, conflate dimensions, misalign preference direction, and contain highly correlated criteria, degrading judge accuracy and producing suboptimal rewards during RFT. We propose RRD, a practical framework for rubric refinement built on a recursive decompose–filter cycle. RRD decomposes coarse rubrics into fine-grained, discriminative criteria, expanding coverage while sharpening separation between responses. A complementary filtering mechanism removes misaligned and redundant rubrics, and a correlation-aware weighting scheme to prevent over-representing highly correlated criteria, yielding rubric sets that are informative, comprehensive, and non-redundant. Empirically, RRD delivers large, consistent gains across all evaluation and training: it improves preference-judgment accuracy on JudgeBench and PPE for both GPT-4o and Llama3.1-405B judges, achieving top performance in all settings with up to +17.7 points on JudgeBench. When used as the reward source for RFT on WildChat, it yields substantially stronger and more stable learning signals, boosting reward by up to 160% (Qwen3-4B) and 60% (Llama3.1-8B) versus ∼10--20% for prior rubric baselines, with gains that transfer to HealthBench-Hard and BiGGen Bench. Overall, RRD establish recursive rubric refinement as a scalable and interpretable foundation for LLM judging and reward modeling in open-ended domains.
Rethinking Vector Field Learning for Generative Segmentation
Chaoyang Wang ⋅ Yaobo Liang ⋅ Boci Peng ⋅ Fan Duan ⋅ Yunhai Tong
Taming diffusion models for generative segmentation has attracted increasing attention. While existing approaches primarily focus on architectural tweaks or training heuristics, there remains a limited understanding of the intrinsic mismatch between continuous flow matching objectives and discrete perception tasks. In this work, we revisit diffusion segmentation from the perspective of vector field learning. We identify two key limitations of the commonly used flow matching objective: gradient vanishing and trajectory traversing, which result in slow convergence and poor class separation. To tackle these issues, we propose a principled vector field reshaping strategy that augments the learned velocity field with a detached distance-aware correction term. This correction introduces both attractive and repulsive interactions, enhancing gradient magnitudes near centroids while preserving the original diffusion training framework. Furthermore, we design a computationally efficient, quasi-random centroid encoding scheme inspired by Kronecker sequences, which integrates seamlessly with an end-to-end pixel neural field framework for pixel-level semantic alignment. Extensive experiments consistently demonstrate significant improvements over vanilla flow matching approaches, substantially narrowing the performance gap between generative segmentation and strong discriminative specialists.
REVERSE: Reinforcing Evidence Verification and Search for Agentic Image Geolocation
Yong Li ⋅ Furong Jia ⋅ Dacheng Yin ⋅ KANG RONG ⋅ Fengyun Rao ⋅ Jing LYU ⋅ Fan Zhang
Image geo-localization aims to determine where a photograph was taken, a task that often requires more than recognizing visible landmarks. Human experts typically solve it through an iterative workflow: they inspect informative regions, form location hypotheses, seek external evidence, and revise their judgments as new clues appear. Existing methods only partially capture this process: direct prediction methods bypass evidence acquisition altogether, while retrieval-augmented methods introduce external evidence but usually provide limited supervision on the intermediate decisions of where to search, how to query, and how to filter noisy results. We present REVERSE, a framework that reinforces the interplay between evidence search and verification to enable multi-turn agentic reasoning. REVERSE teaches three intermediate decisions: where to look, what to query, and what evidence to trust. To support this, we construct tool-grounded trajectories with annotated region selections, search observations, and geo-informative evidence labels, and introduce process rewards for visual grounding, query utility, and evidence discrimination. An offline search cache makes retrieval observations stable and reusable during reinforcement learning, enabling dense supervision over noisy search results. With a 4B model, REVERSE outperforms strong retrieval-augmented baselines and rivals substantially larger models on Im2GPS3k and YFCC4k.
Revisiting The Power of Closed-Form: Robust Deep Image Prototype Discovery via Scale Mixtures
Zhikang Xu ⋅ Jiarui Xing ⋅ Jian Wang
We present a robust deep clustering framework for unsupervised visual discovery that bridges the gap between statistical robustness and mathematical tractability. Current generative approaches typically rely on either noise-prone Gaussian priors or computationally intensive diffusion models; however, neither paradigm simultaneously offers explicit latent structures and exact analytic optimization. While diffusion models generate strong priors, their reliance on iterative sampling and intractable likelihoods limits their scalability and interpretability in clustering tasks. Our approach addresses these limitations by constructing a latent space governed by a Student's $t$ distribution via a Bayesian scale-mixture formulation. This yields the first fully closed-form augmented evidence lower bound (ELBO) for this model class, effectively eliminating the biased variational approximations and high computational costs inherent in prior heavy-tailed or diffusion-based frameworks. By treating precision as a latent variable, our objective automatically modulates the influence of each data point; it downweights outliers to learn a data-dependent measure of trust. This mechanism leads to superior robustness and more stable manifold learning compared to state-of-the-art approaches. We demonstrate the resilience of our framework on standard vision benchmarks and complex neuroimaging tasks. In real-world brain magnetic resonance imaging (MRI)-driven neurodegenerative disease analysis, our model successfully recovers clean and anatomically coherent clusters where clinical noise typically collapses rare subtypes. These results underscore the ability of our framework to preserve clinically crucial morphological subtypes, allowing precise discovery within clinical interventions.
Reward-free Pretraining for Reinforcement Learning via Occupancy Coverage Maximization
Marco Pratticò ⋅ Pietro Novelli ⋅ Massimiliano Pontil ⋅ Carlo Ciliberto
Sparse rewards pose a central challenge in reinforcement learning, since agents receive no informative signal until they reach their goal. Intrinsic-reward methods address this issue by optimizing non-stationary objectives such as novelty, prediction error, or skill diversity, thereby injecting a supervision signal into the problem. While effective, these methods often require that the extrinsic (sparse) reward can be evaluated -- either online or during offline relabeling of the stored transitions. This limitation is particularly vexing for multi-task, meta-, and continual reinforcement learning, where agents' interactions with the environment are usually {\em reward-free}. In this work, we present a method to pre-train {\em transferable exploration policies} that rapidly adapt to sparse rewards at downstream task time. Our objective maximizes state-space covering for the occupancy measure, and can be framed in terms of entropy maximization. Its algorithmic implementation, ROVER, leverages recent advances on the operatorial formulation of RL to estimate occupancy with a learned resolvent world model, bypassing common hurdles associated with density and entropy estimation. ROVER further introduces a virtual ``sink" state for unexplored regions, balancing coverage of known states with expansion into unseen ones and preventing cyclic expansion–collapse behavior during learning. In tabular and pixel-based sparse navigation tasks, ROVER produces more uniform aggregate coverage and stronger initializations for downstream tasks than standard reward-free baselines.
Reward Hacking in Rubric-Based Reinforcement Learning
Anas Mahmoud ⋅ MohammadHossein Rezaei ⋅ Zihao Wang ⋅ Anisha Gunjal ⋅ Bing Liu ⋅ Yunzhong He
Reinforcement learning with verifiable rewards has enabled strong post-training gains in domains such as math and coding, though many open-ended settings rely on rubric-based rewards. We study reward hacking in rubric-based RL, where a policy is optimized against a training verifier but evaluated against a cross-family panel of three frontier judges, reducing dependence on any single evaluator. Our framework separates two sources of divergence: verifier failure, where the training verifier credits rubric criteria that reference verifiers reject, and rubric-design limitations, where even strong rubric-based verifiers favor responses that rubric-free judges rate worse overall. Across medical and science domains, weak verifiers produce large proxy-reward gains that do not transfer to the reference verifiers; exploitation grows over training and concentrates in recurring failures such as partial satisfaction of compound criteria, treating implicit content as explicit, and imprecise topical matching. Stronger verifiers substantially reduce, but do not eliminate, verifier exploitation. We also introduce a self-internalization gap, a verifier-free diagnostic based on policy log-probabilities, which tracks reference-verifier quality, detecting when the policy trained using the weak verifier stops improving. Finally, in our setting, stronger verification does not prevent reward hacking when the rubric leaves important failure modes unspecified: rubric-based verifiers prefer the RL checkpoint, while rubric-free judges prefer the base model. These disagreements coincide with gains concentrated in completeness and presence-based criteria, alongside declines in factual correctness, conciseness, relevance, and overall quality. Together, these results suggest that stronger verification reduces reward hacking, but does not by itself ensure that rubric gains correspond to broader quality gains.
Reward Inflation: A Healthy Stimulus for Reinforcement Learning
Ganghun Lee ⋅ Minji Kim ⋅ Minsu Lee ⋅ Byoung-Tak Zhang
Reward serves as the primary learning signal in reinforcement learning (RL). However, while reward magnitudes are typically held fixed throughout training, their temporal modulation remains underexplored. In this paper, we propose reward inflation, a gradual scaling of rewards over the course of training, and show that it can act as a healthy stimulus for RL. Theoretically, reward inflation induces an implicit recency weighting that upweights recent transitions during policy updates, enabling faster adaptation. We further show that, by sustaining gradient signals as the policy saturates, reward inflation suppresses the emergence of dormant neurons and helps preserve plasticity. Empirical results on ALE games and MuJoCo tasks corroborate these findings, showing that an appropriate level of reward inflation benefits a broad range of tasks. Finally, we introduce Fed, an adaptive variant that adjusts the inflation level on the fly, and find that it often improves upon fixed inflation.
Reward Modeling for Multi-Agent Orchestration
King Yeung Tsang ⋅ Zihao Zhao ⋅ Vishal Venkataramani ⋅ Haizhou Shi ⋅ Zixuan Ke ⋅ Semih Yavuz ⋅ Shafiq Joty ⋅ Hao Wang
Multi-Agent Systems (MAS) built on Large Language Models (LLMs) require effective orchestration to coordinate specialized agents, yet training such orchestrators is hindered by limited supervision and high computational cost. We propose **Orch**estration **R**eward **M**odeling (**Orch-RM**), a self-supervised framework for evaluating orchestration quality without human annotations. Orch-RM leverages intermediate artifacts from multi-agent executions to construct win-lose pairs for Bradley-Terry reward model training. Unlike existing MAS test-time scaling and orchestrator training frameworks that rely on costly sub-agent rollouts, Orch-RM operates directly at the orchestration level, enabling efficient and high-performing reward-guided orchestrator training and MAS test-time scaling. As shown in Figure 1, Orch-RM improves training efficiency by up to 10$\times$ in token usage while improving MAS test-time scaling performance by up to 8\% in accuracy. These gains consistently transfer across multiple domains, including mathematical reasoning, web-based question answering, and multi-hop reasoning, demonstrating orchestration-level reward modeling as a scalable direction for robust multi-agent orchestration. Data, code, and trained models will be released upon publication.
Reweighted Flow Matching via Unbalanced Optimal Transport for Label-free Long-tailed Generation
Hyunsoo Song ⋅ Minjung Gim ⋅ Jaewoong Choi
Flow matching has recently emerged as a powerful framework for continuous-time generative modeling. However, when applied to long-tailed distributions, standard flow matching frameworks are susceptible to *majority bias*, over-representing dominant modes at the expense of tail-class fidelity and distributional accuracy. In this work, we propose ***Unbalanced Optimal Transport Reweighted Flow Matching (UOT-RFM)***, a novel framework for generative modeling under class-imbalanced (long-tailed) distributions that operates without any class label information. Our method constructs the conditional vector field using mini-batch Unbalanced Optimal Transport (UOT) and mitigates majority bias through a principled inverse reweighting strategy. The reweighting relies on a *label-free majority score*, defined as the density ratio between the target distribution and the UOT marginal. This score quantifies the degree of majority based on the geometric structure of the data, without requiring class labels. By incorporating this score into the training objective, UOT-RFM theoretically recovers the target distribution with first-order correction ($k=1$) and empirically improves tail-class generation through higher-order corrections ($k > 1$). Our model outperforms existing flow matching baselines on long-tailed benchmarks, while maintaining competitive performance on balanced datasets.
Offline-to-Online (O2O) reinforcement learning (RL) is essential for policy adaptation but frequently suffers from initialization collapse due to evaluation and improvement mismatches across the learning transition. In safety-critical domains, safety-constrained model-free O2O RL without offline data retention remains significantly under-explored. Existing O2O fine-tuning methods, typically relying on policy regularization or value recalibration within Euclidean spaces, lack the structural capacity to strictly sequester policies from hazardous regions. To facilitate a geometric paradigm for safe policy fine-tuning, we propose Riemannian Admissibility Flow (RAF), which lifts safety constraints into a geometric framework by reformulating policy optimization as probability transport on anisotropic Riemannian manifolds. By fusing safety and uncertainty into the metric tensor, RAF transforms extrinsic scalar penalties into intrinsic topological barriers with directional gating, enabling target-free geodesic flow matching for hazard circumvention without offline data retention. Evaluation on representative safe O2O benchmarks shows that RAF achieves competitive performance in constraint satisfaction and task execution while effectively mitigating initialization collapse.
RigPAPR: Rig-Based Animation of Static Neural Point Clouds from a Single-View Video
Shichong Peng ⋅ Yanshu Zhang ⋅ Ke Li
Static neural point reconstructions capture a subject at high fidelity from posed images. Given such a reconstruction and a monocular fixed-viewpoint driving video of the subject, whether captured or produced by image-to-video (I2V) generation, we recover a rigged, re-posable 3D asset. Existing methods deform Gaussian splats through direct linear blend skinning (LBS) or mesh proxies, both of which are prone to joint-boundary artifacts under articulation, even with per-primitive corrections. We trace the artifact to the representation: each splat carries an individual shape calibrated in the canonical pose to tile with its neighbours. Under rigid LBS, each splat moves with its bone but cannot bend, so the canonical tiling breaks at joint boundaries into gaps and spikes. Proximity attention point rendering (PAPR) instead carries no per-primitive shape; each pixel is recomposed at render time from the deformed primitives' positions, so the surface re-forms naturally with the articulation. We present RigPAPR, which auto-rigs a static PAPR cloud and drives it under direct LBS from a single fixed-viewpoint video, without mesh proxy, pose-dependent correction, or category template. On synthetic subjects, RigPAPR matches the strongest baseline at the supervised view and exceeds mesh-based and Gaussian-splatting baselines at novel views by 3+ dB PSNR, with cleaner joint-boundary renderings of both synthetic and real subjects.
RipplePLM: Structural and Property Decoupling for Protein Mutation Effect Generation
Liuzhenghao Lv ⋅ Yuyang Liu ⋅ Yuyang Gao ⋅ Li Yuan ⋅ Yonghong Tian
Protein mutation effect generation asks a model to describe the functional consequence of a point mutation in natural language. Existing protein-to-text systems typically encode mutation information into undifferentiated representations, overlooking the organization of mutation-induced evidence across structural and biochemical factors. We propose RipplePLM, a mutation-aware generation framework centered on Direct-Distal Cross-Attention (DDCA). By constructing a residue-level Mutation Perturbation Field from pre-trained protein language models, DDCA leverages predicted contact maps to explicitly decouple structural perturbations into two pathways: the mutation site's immediate contact neighborhood and its multi-hop distal context. To complement this structural decomposition, we further introduce the Property Latent Chain (PLChain), which injects expert-guided biochemical supervision (e.g., thermostability and optimal pH) into the LLM hidden-state pathway through latent property tokens. On MutaDescribe, RipplePLM improves over mutation-specific baselines on temporal and structural splits; under a matched-backbone comparison, average structural-split ROUGE-L increases from $\textit{22.23}$ to $\textit{35.65}$. Additional ablations, representation diagnostics, and low-$N$ fitness regression experiments further support the effectiveness of the learned mutation-aware representations.
Risk-Controlled Post-Processing of Decision Policies
Sunay Joshi ⋅ Tao Wang ⋅ Hamed Hassani ⋅ Edgar Dobriban
Predictive models are often deployed through existing decision policies that stakeholders are reluctant to change unless a risk constraint requires intervention. We study risk-controlled post-processing: given a deterministic baseline policy, choose a new policy that maximizes agreement with the baseline subject to a chance constraint on a user-specified loss. At the population level, we show that the optimal policy has a threshold structure: it follows the baseline except on contexts where switching to the oracle fallback policy yields a large reduction in conditional violation risk. At the finite-sample level, given a fitted fallback policy and score, we develop a post-processing algorithm that uses calibration data to select a threshold. Leveraging tools from algorithmic stability and stochastic processes, we show that under regularity conditions, in the i.i.d. setting, the expected excess risk of the post-processed policy is $O(\log n/n)$. In the special case when an exact-safe fallback policy is available, the algorithm achieves precise expected risk control under exchangeability. In this setting, we also give high-probability near-optimality guarantees on the post-processed policy. Experiments on a COVID-19 radiograph diagnosis task, an LLM routing problem, and a synthetic multiclass decision task show that targeted post-processing can meet or nearly meet risk budgets while preserving substantially more agreement with the baseline than score-blind random mixing.
RIZZ: Routing Interactions to Near Zero-Interference Zones for Continual Adaptation of Black-Box Agents
Sonali Goel ⋅ Pranav Vaidhyanathan ⋅ Lucas Schorling ⋅ Natalia Ares ⋅ Michael A Osborne
Large language models are increasingly deployed as long-lived agents that must adapt across users, tasks, domains, modalities, and feedback regimes without access to model weights. Existing black-box adaptation methods typically optimize a single prompt, maintain an undifferentiated memory, or rely on repeated rollout-heavy search. However, these designs struggle when streams of input are nonstationary, feedback is sparse, and failures from one task family can contaminate behavior on another. We introduce RIZZ (Routing Interactions to Near Zero-interference Zones), a continual adaptation framework for compound language-model systems that learns entirely through verifier-gated memory, routing, and prompt compilation. RIZZ organizes input streams into dynamically spawned memory branches. At inference time, either while online or offline, a context-aware router selects or creates a branch that retrieves branch-local, global, graph-structured, and working-memory context, which is compiled into a bounded prompt together with retrieved task evidence. After the model acts, task verifiers score the output, and only verified interactions can update memory, promote reusable rules, demote harmful rules, or create anti-patterns. This yields a black-box agent that improves through persistent natural-language feedback while explicitly controlling interference. RIZZ targets the regime where adaptation must occur online under context budgets. Finally, we demonstrate the effectiveness of our framework against state-of-the-art baselines on competitive benchmarks.
RoadmapBench: Evaluating Long-Horizon Agentic Software Development Across Version Upgrades
Xinbo Xu ⋅ Ruihan Yang ⋅ Haiyang Shen ⋅ Wendong XU ⋅ Bofei Gao ⋅ Ruoyu Wu ⋅ Kean Shi ⋅ Weichu Xie ⋅ Xuanzhong Chen ⋅ Ming Wu ⋅ Jason Zeng ⋅ Michael Heinrich ⋅ Liang Chen ⋅ Kuan Li ⋅ Baobao Chang
Coding agents are increasingly deployed in real software development, where a single version iteration requires months of coordinated work across many files. However, most existing benchmarks focus predominantly on single-issue bug fixes from Python repositories, with coarse pass/fail evaluation outcomes, and thus fail to capture long-horizon, multi-target development at real engineering scale. To address this gap, we present ROADMAPBENCH, a benchmark of 115 long-horizon coding tasks grounded in real open-source version upgrades across 17 repositories and 5 programming languages. Each task places the agent on a source-version code snapshot and provides a multi-target roadmap instruction requiring it to implement the functionality introduced in the target version, with a median modification of 3700 lines across 51 files. We conduct a systematic evaluation on thirteen frontier models and find that even the strongest, Claude-Opus-4.7, resolves only 39.1% of tasks, while the weakest achieves merely 5.2%, in stark contrast to the high scores these models attain on existing bug-fix benchmarks, suggesting that long-horizon software development remains a largely unsolved problem.
ROAD: Rule-Grounded Context-Aware Open-World Driver Anomaly Detection
Yingjin Li ⋅ Shuaibo Li ⋅ Wei Ma ⋅ Zhijie Qiu ⋅ Fuxiang Zhai ⋅ Sixiang Chen ⋅ Zhaolu Kang ⋅ Lei Zhu
Driver monitoring concerns determining not only what a driver is doing, but whether the behavior violates a driver-behavior safety rule under the current driving context. Existing distracted-driving benchmarks largely reduce this problem to closed-set action recognition, often missing cases where the same behavior changes meaning with road scene, vehicle state, duration, and rule-specific exceptions. We present ROAD, a rule-grounded context-aware open-world framework for driver anomaly detection. ROAD represents driver-monitoring constraints as structured rule cards and jointly reasons over in-cabin driver evidence, road-context video, optional vehicle-state metadata, and open-world semantic cues to produce structured anomaly judgments: driver action, violated rule, anomaly score, supporting evidence, exception status, and explanation. Unlike generic MLLM prompting, ROAD separates perception, rule grounding, and semantic generalization. Road context modulates driver-centric evidence, rule-card encodings select rule-relevant evidence, and a LoRA-adapted MLLM extracts open-vocabulary safety cues from compressed video evidence. This design targets two forms of open-world generalization: recognizing unseen driver behaviors beyond a fixed action taxonomy and adapting to new or revised rule cards. We further introduce ROAD-Bench, a rule-grounded benchmark with context, exception, evidence, anomaly-score, and explanation annotations. With a three-stage curriculum from action primitives to context-rule alignment and anomaly adaptation, ROAD outperforms evaluated baselines across driver-distraction recognition, open-world behavior recognition, and rule-grounded anomaly detection, using an 8B multimodal backbone. It achieves 92.01\% average accuracy over seven driver-recognition benchmarks, 77.9\% AUROC and 46.9\% unseen-class F1 in behavior recognition, and 73.6\% AUROC, 68.3\% rule accuracy, and 54.5\% Expl.-Rule F1 on ROAD-Bench.
We introduce the first Probably Approximately Correct (PAC) learning framework for general-sum concurrent stochastic games (CSGs) with transition uncertainty, while addressing the challenge of Nash equilibrium (NE) existence. Our algorithm maintains data-driven $L^1$ confidence sets over transition kernels and solves a robust CSG to compute a social-welfare optimal $\varepsilon$-NE, using a novel robust MDP-based exploration mechanism to drive joint state-action coverage without unilateral control. Crucially, we introduce a Nash margin characterisation that enables principled reasoning about equilibrium existence: the framework either returns an $\varepsilon$-approximate NE whose social-welfare value is $\varepsilon$-close to optimal, or provides a sound certificate that no exact NE exists. Under a minimum reachability condition $p_{\mathrm{reach}} > 0$ over all state-action pairs, the algorithm terminates after a polynomial number of interactions, with sample complexity $\widetilde{\mathcal{O}}\left( {R_{\max}^2 H^4 |S|^2 |A| / (p_{\mathrm{reach}} \varepsilon^2)} \right)$. Empirical results on benchmark CSGs demonstrate near-optimal performance, correct handling of equilibrium (non-)existence, and sample complexity consistent with theory; implementation available at: https://anonymous.4open.science/r/pg1.
RobustPruner: Decoupled Relevance and Uncertainty for Efficient Visual Token Pruning in MLLMs
Hanwei Zhu ⋅ Junhan Fu ⋅ Xi Zhang ⋅ Jiamang Wang ⋅ Weisi Lin
Multimodal large language models (MLLMs) represent images as long sequences of visual tokens, making inference costly and often redundant. The central challenge is to prune these tokens aggressively without removing evidence required for reliable reasoning. Existing pruning methods typically rank tokens by estimated importance or redundancy, implicitly assuming that low-scored tokens are safe to discard. This assumption is fragile when small, occluded, or fine-grained visual cues are uncertain yet decisive, and it becomes especially brittle under semantics-preserving prompt perturbations. We present RobustPruner, a decoder-integrated framework for prompt-robust and uncertainty-aware visual token pruning. RobustPruner predicts query-conditioned relevance, uses a determinantal point process (DPP) to construct a diverse candidate subset, and then applies uncertainty-guided refinement within that subset. By decoupling relevance and uncertainty, the method preserves ambiguous but potentially decisive evidence while remaining computationally efficient. Inserted into intermediate decoder layers, RobustPruner performs one-shot pruning and reduces KV-cache and memory costs. Across three open-source MLLMs and diverse vision-language benchmarks, RobustPruner consistently achieves a stronger accuracy-efficiency-robustness trade-off than prior pruning baselines, with especially clear gains on fine-grained and text-rich tasks. On Qwen-2.5-VL-7B, RobustPruner retains 95.22\% of the original performance while reducing GPU memory, prefill FLOPs, and KV cache size to 91.43\%, 50.97\%, and 20.57\% of the baseline.
Roll2Depth: Zero-Shot Metric Depth Estimation Exploiting Camera's Rolling Shutter Effect
Rui Xiao ⋅ Jinming Xu ⋅ Jinsong Han
Accurate monocular metric depth estimation, i.e., predicting absolute depth values, is critical for applications such as robotics and augmented reality. However, existing methods often suffer from poor transferability to in-the-wild scenarios due to the lack of knowledge of camera intrinsics. In this paper, we propose a novel framework, Roll2Depth, to recover metric scale in unseen environments with unknown camera parameters. Our key idea is to leverage scene-independent physical cues induced by the rolling shutter effect of standard CMOS sensors as metric priors. When a scene is illuminated by a flickering light, the rolling shutter produces distinctive stripe-like spatial patterns, revealing the scene's shading with and without the flickering light source. This contrast encodes absolute depth cues that can be exploited for metric estimation. Roll2Depth isolates these patterns as metric priors to guide affine-invariant models on metric depth estimation in a zero-shot setting. Our evaluations on standard benchmark datasets, as well as a real-world dataset collected using commercial smartphones and our prototype system, demonstrate that Roll2Depth, as a low-cost active solution, outperforms state-of-the-art passive baselines in zero-shot settings. These results highlight its potential impact in applications such as indoor robotics and embodied perception, where affordable and effective active sensing would be valuable.
RSPReg: Reliability-aware Structural Prototype Learning for Point Cloud Registration
Arzigul Ahat ⋅ Eksan Firkat ⋅ Askar Hamdulla
Point cloud registration remains challenging in scenarios with low overlap, repetitive structures, and local geometric ambiguities, where candidate correspondences established based on local geometric similarity often lack global structural consistency and are prone to unreliable matches. To address this issue, this paper proposes RSPReg, a reliability-aware structural prototype learning framework for robust point cloud registration. RSPReg learns structural prototypes from rotation-invariant coarse-level superpoint features to characterize the structural attributes of local regions within the entire point cloud. These prototypes are further introduced as structural priors into cross-cloud correspondence reasoning, thereby improving the structural consistency of candidate correspondences. Furthermore, RSPReg estimates the uncertainty of structural prototypes and adaptively modulates the role of structural priors in the matching and pose estimation processes, reducing the negative impact of ambiguous and non-overlapping regions on registration results. Experiments on 3DMatch, 3DLoMatch, KITTI, and ETH demonstrate that RSPReg achieves the highest registration recall and exhibits clear advantages in low-overlap and cross-domain scenarios.
RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
Yi-Lin Sung ⋅ Prateek Yadav ⋅ Jialu Li ⋅ Jaehong Yoon ⋅ Mohit Bansal
Layer-wise quantization is a key technique for efficiently compressing large models without expensive retraining. Previous methods typically quantize the weights of each layer by “uniformly” optimizing the layer reconstruction loss across all output tokens. However, in this paper, we demonstrate that better quantized models can be obtained by prioritizing learning from important tokens. Building on this finding, we propose RSQ (Rotate, Scale, then Quantize), which (1) applies rotations (orthogonal transformation) to the model to mitigate weight outliers, (2) scales the token feature based on its importance, and (3) quantizes the model using the GPTQ framework with the second-order statistics computed by scaled tokens. To compute token importance, we explore both heuristic and dynamic strategies. Based on a thorough analysis of all approaches, we adopt attention concentration, which uses attention scores of each token as its importance, as the best approach. We demonstrate that RSQ consistently outperforms baseline methods across multiple downstream tasks and three model families: LLaMA3, Mistral, and Qwen2.5. Additionally, models quantized with RSQ achieve superior performance on long-context tasks, further highlighting its effectiveness. Lastly, RSQ demonstrates generalizability across various setups, including different model sizes, calibration datasets, bit precisions, and quantization methods. Our code is available in the supplementary material.
Rubato: Transcribing Piano Music with Timestamps
Nazif Can Tamer ⋅ Victoria Ebert ⋅ Guang Yang ⋅ Noah Smith
We consider the conversion of musical recordings into human-readable sheet music annotated with timestamps. Such output lets a listener clearly visualize rubato (temporally expressive playing), a learner diagnose ensemble precision and timing choices against the written music, and a musicology scholar compare performance styles across recordings of the same work. We introduce (1) a prompt-conditioned encoder--decoder model, named Rubato, trained to output (2) a new textual representation for polyphonic music, named InterMo, which we designed for compatibility with sequence-to-sequence training. Our experiments demonstrate that Rubato produces timestamped piano sheet music from audio with higher notational accuracy than the best existing approaches, which are based on cascades. We find that even if the cascade is given ground-truth MIDI instead of audio, Rubato performs better, suggesting that the ceiling of existing approaches is primarily representational, not acoustic. Further, because Rubato is trained on several related tasks (with prompts), it competes with or outperforms the best single-task systems on related but simpler tasks like MIDI note grounding and beat/downbeat detection. We will release the model and code on publication, and an anonymized demo is available at https://rubato-transcription.github.io.
Runtime Analysis of Cartesian Genetic Programming on MAX: A Proven Exponential Speedup
Duc-Cuong Dang ⋅ Roman Kalkreuth ⋅ Andre Opris
Genetic Programming (GP) is a search paradigm inspired from natural evolution for the automated discovery of expressions, functions, and computer programs. Cartesian GP (CGP) is a flavor of GP that uses a graph-based model to encode candidate programs, as contrasted to the conventional Tree-based GP (TGP). Since its inception two decades ago, CGP has predominantly been analyzed empirically, thus little is known regarding its theoretical performance guarantees. This paper analyzes CGP for the MAX problem, which has been only rigorously studied for TGP with a proven expected runtime superlinear in the optimal program length $n$. We prove that the expected time for the (1+1) CGP employing common mutation operators to evolve an equivalent optimal of the same output is polylogarithmic in n, specifically at most $O(\log^9{n})$. This is remarkable, as it shows an exponential speedup by switching to CGP. Our results shed light on the benefit and the compactness of the graph-based representation for programs, and on how CGP can navigate its search and fitness spaces efficiently. This is a first proven performance guarantee for CGP on an established benchmark problem. Experiments complement our theoretical findings.
RVR: Retrieve-Verify-Retrieve for Comprehensive Question Answering
Deniz Qian ⋅ Hung-Ting Chen ⋅ Eunsol Choi
Comprehensively retrieving diverse documents is crucial to address queries that admit a wide range of valid answers. We introduce retrieve-verify-retrieve (RVR), a multi-round retrieval framework designed to maximize answer coverage. Initially, a retriever takes the original query and returns a candidate document set, followed by a verifier that identifies a high-quality subset. For subsequent rounds, the query is augmented with previously verified documents to uncover answers that are not yet covered in previous rounds. RVR is effective even with off-the-shelf retrievers, and fine-tuning retrievers for our inference procedure brings further gains. Our method outperforms baselines, including agentic search approaches, achieving at least 10\% relative and 3\% absolute gain in complete recall percentage on a multi-answer retrieval dataset (QAMPARI). We also see consistent gains on two out-of-domain datasets (QUEST and WebQuestionsSP) across different base retrievers. Our work presents a promising iterative approach for comprehensive answer recall leveraging a verifier and adapting retrievers to a new inference scenario.
Safeguarding Mutual Correction in Source-Free Domain Adaptation via Cut Statistics
Seongjun Lee ⋅ Changhee Lee
Source-Free Domain Adaptation (SFDA) aims to adapt a source-pretrained model to an unlabeled target domain without access to the original source domain. While early single-model approaches rely on self-refinement, they are inherently susceptible to confirmation bias and struggle to correct their own systematic errors. To overcome this limitation, recent methods introduce Vision-Language (ViL) models as external knowledge sources. However, these approaches operate in a largely unidirectional paradigm, using the ViL model primarily to supervise the source-pretrained model. This overlooks a key structural property: the two models exhibit distinct failure modes -- where one produces an incorrect prediction, the other may produce a correct one, creating a natural opportunity for mutual correction within the target domain. Yet, without ground-truth labels, identifying which model is correct on any given sample is non-trivial, and naively exchanging predictions risks propagating errors across models. To address this challenge, we propose SafeCut, a novel approach that leverages the cut statistic as a label-free measure of prediction reliability to gate cross-model supervision. Our approach dynamically controls both the direction and strength of supervision based on relative reliability, selectively amplifying true corrections while suppressing miscorrections on a per-sample basis. We further provide theoretical justification showing that this reliability-gated mechanism guarantees a net-positive correction signal. Extensive experiments across diverse SFDA benchmarks demonstrate that SafeCut achieves state-of-the-art performance, highlighting the effectiveness of safeguarding mutual correction in SFDA via cut statistics. Codes are available https://anonymous.4open.science/r/SafeCutSourceFreeDomainAdaptation-65D0/
SamaDICE: Safe Multi-Agent Reinforcement Learning with Stationary Distribution Correction Estimation
The Viet Bui ⋅ Tien Mai ⋅ Thanh Nguyen
Offline safe multi-agent reinforcement learning (MARL) is challenging: agents must learn coordinated behaviors from static datasets while strictly satisfying safety constraints. This challenge is amplified by distribution shift, where joint actions that appear safe under the behavior distribution may become unsafe under policy-induced deviations, with errors compounding across agents. We propose SamaDICE, a principled framework for safe offline MARL that integrates stationary distribution correction with scalable value decomposition. Our approach formulates safe policy learning as a constrained convex optimization problem over stationary distributions of joint state-action pairs, with safety constraints imposed as upper bounds on expected cumulative costs. Safety is enforced through a Lagrangian dual formulation, enabling distribution-corrected cost estimation directly from offline data. To address the combinatorial complexity of multi-agent systems, we introduce a factored centralized-training-decentralized-execution parameterization that represents the global dual variable via a monotonic mixing network over agent-local potentials, yielding a tractable surrogate Bellman residual for scalable optimization. We instantiate SamaDICE with a multi-agent decision transformer and evaluate it on diverse benchmark environments. Our method outperforms the strongest baseline in aggregate return, with the advantage widening as the cost budget tightens, while maintaining 100% empirical episode-level constraint satisfaction across all evaluated tasks.
Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines?
Anton Korznikov ⋅ Andrey V Galichin ⋅ Alexey Dontsov ⋅ Oleg Rogov ⋅ Ivan Oseledets ⋅ Elena Tutubalina
Sparse Autoencoders (SAEs) have emerged as a promising tool for interpreting neural networks by decomposing their activations into sparse sets of human-interpretable features. Despite much excitement, a growing number of negative results in downstream tasks casts doubt on whether SAEs recover meaningful features. To directly investigate this, we perform two complementary evaluations. On a synthetic setup with known ground‑truth features, we demonstrate that three out of four state‑of‑the‑art SAE architectures recover only 7–9\% of true features despite achieving around 71\% explained variance, showing that strong reconstruction alone is insufficient to guarantee meaningful feature recovery. To evaluate SAEs on real activations, we introduce three baselines that constrain SAE feature directions or their activation patterns to random values. Through extensive experiments across multiple SAE architectures, we show that our baselines match fully-trained SAEs in interpretability (0.87 vs 0.90), sparse probing (0.69 vs 0.72), and causal editing (0.73 vs 0.72). These results show that current evaluation metrics are insufficient to certify that SAEs have learned meaningful features, and we offer our baselines as a reusable protocol for future SAE evaluation.
SAVeR$^2$: Reasoning-based Safety Alignment for Large Reasoning Models via Verifiable Rewards
Yichen Sun ⋅ LIN JIANAN ⋅ Linbo Jiang ⋅ Shiyu Wang ⋅ Zhibo Wang ⋅ Zhixuan Chu
Large Reasoning Models (LRMs) have achieved impressive performance via explicit Chain-of-Thought (CoT) reasoning, yet this process introduces critical safety risks. We formalize LRM safety objectives as requiring non-malicious reasoning and answers, safety consistency, and explicit intent identification. We argue that existing alignment methods suffer from poor adversarial generalization, a significant ``safety tax'' on reasoning, and unfaithful coarse-grained rewards, primarily due to their neglect of the LRM's intrinsic reasoning capabilities for safety. To address these, we propose \textbf{SAVeR$^2$}, a novel \textbf{S}afety \textbf{A}lignment method via \textbf{R}ewards with \textbf{R}easoning capability. SAVeR$^2$ decomposes the safety objective into four fine-grained verifiable reward signals and utilizes a difficulty-aware data selection strategy to stabilize training. Crucially, we introduce a reasoning-preserving gradient projection mechanism that analytically resolves conflicts between safety and reasoning gradients, effectively eliminating the safety tax. Extensive experiments on models ranging from 1.5B to 671B demonstrate that SAVeR$^2$ achieves superior safety performance and significantly lower over-refusal rates across multiple benchmarks without compromising general reasoning capabilities.
Scalable Adaptation of 3D Geometric Foundation Models via Weak Supervision from Internet Video
Zihui Gao ⋅ Ke Liu ⋅ Jiawang Bian ⋅ Donny Y. Chen ⋅ Duochao Shi ⋅ Guosheng Lin ⋅ Hao Chen ⋅ Chunhua Shen
Geometric foundation models show promise in 3D reconstruction, yet their progress is severely constrained by the scarcity of diverse, large-scale 3D annotations. While Internet videos offer virtually unlimited raw data, utilizing them as a scaling source for geometric learning is challenging due to the absence of ground-truth geometry and the presence of observational noise. To address this, we propose SAGE, a framework for Scalable Adaptation of GEometric foundation models from raw video streams. SAGE leverages a hierarchical mining pipeline to transform videos into training trajectories and hybrid supervision: (1) Informative training trajectory selection; (2) Sparse Geometric Anchoring via SfM point clouds for global structural guidance; and (3) Dense Differentiable Consistency via 3D Gaussian rendering for multi-view constraints. To prevent catastrophic forgetting, we introduce a regularization strategy using anchor data. Experiments on applying SAGE to two architecturally distinct backbones (including both MV-DUSt3R and VGGT) show that Chamfer Distance is consistently reduced by 20-42% on unseen benchmarks (7Scenes, TUM-RGBD, Matterport3D). SAGE establishes Internet video as a viable and scalable adaptation source for geometric foundation models.
ScAn-Bench: Evaluating Scaling Analysis Methodology
Artin Sermaxhaj ⋅ Nastaran Alipour ⋅ Donat Sinani ⋅ Johannes Hog ⋅ Neeratyoy Mallik ⋅ Steven Adriaensen ⋅ Jenia Jitsev ⋅ Danny Stoll
Recent progress in machine learning is driven by large-scale foundation models, where scaling laws and finding optimal scaling prescriptions for architecture, data, and hyperparameters are key in advancing the state-of-the-art. Therefore, it is surprising that no systematic study evaluates the methodology to obtain scaling laws and prescriptions across different model types. To shed light on this crucial blind spot and facilitate future research, we introduce the surrogate benchmarks ScAn-Bench-LLM and ScAn-Bench-VLM based on 2875 and 3853 checkpoints of language and vision-language model pipelines. On our benchmarks, we perform the first systematic evaluation of both data acquisition and extrapolation methodology for scaling analysis across different data modalities.
SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers
Kaihua Qin ⋅ Dawn Song ⋅ Arthur Gervais
Smart contract decompilation aims to recover high-level source code from bytecode, but evaluating decompilers remains difficult because existing studies use narrow datasets, inconsistent metrics, and limited semantic consistency checks. This gap is increasingly important as large language models (LLMs) begin to generate source-like Solidity that may compile and appear plausible, even when its semantics diverge from the original contract. We introduce SCDBench, a dataset and benchmark methodology for LLM-based smart contract decompilation. The dataset contains $600$ real-world Solidity contracts with paired bytecode inputs, ground-truth source code, and replayable semantic checkpoints. SCDBench evaluates decompiler outputs through four cumulative stages: format completeness, compilability, Application Binary Interface (ABI) recovery, and semantic consistency via differential replay. We evaluate Claude Opus 4.7, GPT-5.3-Codex, and GLM-5 in a zero-shot decompilation setting, including GLM-5 variants with and without extended reasoning and a zero-shot compilation-repair setting. The results show that frontier LLMs can often produce structured and compilable Solidity, but achieving semantic consistency remains far from solved: the best-performing frontier model perfectly decompiles only $42/600$ contracts. We further show that introducing same-model compilation repair substantially improves performance at modest additional cost. SCDBench establishes a common ground for rigorous, reproducible evaluation and aims to accelerate the development of reliable smart contract decompilers for blockchain security and transparency.
SchedDiff: Diffusion-Based Priority Refinement for Job Shop Scheduling
Dong-Yoon Oh ⋅ Sang-Hyun Cho ⋅ Inguk Choi ⋅ Hyun-Jung Kim
Neural combinatorial optimization methods for the Job Shop Scheduling Problem (JSSP) face a fundamental trade-off between solution quality and inference efficiency: constructive methods build high-quality schedules through sequential decisions whose computational cost scales with problem size, whereas one-shot methods offer highly efficient single-pass inference at the expense of solution quality. We propose SchedDiff, the first diffusion-based framework for JSSP, which bridges this gap by iteratively refining operation priorities through a fixed number of denoising steps. Unlike existing diffusion models for combinatorial optimization that operate in a binary space, SchedDiff operates directly in a discrete ordinal space, introducing a forward process and learning objective specifically tailored to ranking structures. To decode the refined priorities into feasible schedules, we introduce active list scheduling, which provably produces more compact schedules than standard list scheduling. Extensive experiments on TA, DMU, and LA benchmarks demonstrate that SchedDiff achieves state-of-the-art performance among learning-based methods, with efficiency advantages that become more pronounced on large-scale instances. SchedDiff further exhibits strong generalization, scaling to instances significantly larger than those seen during training, generalizing across unseen problem distributions, and transferring zero-shot to flexible JSSP.
SEAD: Competence-Aware On-Policy Distillation via Entropy-Guided Supervision
Michael Lee ⋅ Zelei Cheng ⋅ Yu Wang ⋅ Renkun Ni ⋅ Sambit Sahu ⋅ Shi-Xiong Zhang ⋅ William Campbell
On-policy distillation (OPD) has a property absent in offline distillation and RL: teacher supervision quality depends on student competence. Incoherent rollouts yield noisy gradients; already-mastered tokens yield redundant ones. This creates waste at three scales—tokens, training phases, and prompts—yet existing methods supervise uniformly. We introduce SEAD, which uses entropy as a unified probe of this competence-dependent degradation at three scales: (1) joint teacher-student entropy partitions tokens into zones receiving tailored divergences or zero gradient (~50% skipped); (2) a cosine schedule anneals from forward to reverse KL as competence grows; (3) a competence-gated curriculum introduces prompts easy-to-hard. These components are symbiotically necessary: token selection requires coherent rollouts (curriculum), and annealing requires monotonic improvement (also curriculum). On OLMo-3 (7B to 32B), SEAD achieves +4.8 average accuracy over vanilla OPD across six math benchmarks, with ablations confirming super-additive interactions.
Second-Order Complexity Theory for Risk, Explanation, and Calibration in Machine Learning
Yoshihiro Maruyama
Many machine-learning quantities are straightforward to compute once a finite data set is fixed, but their population versions can encode much harder computational problems. We study this gap for population risk, attribution, counterfactual explanation, and calibration certification, asking which of these quantities can be approximated in time polynomial in the requested number of bits. The formal basis is the Kawamura-Cook second-order complexity theory, instantiated here with compact-domain real functions carrying explicit evaluation, continuity, and range information. In this framework, expected risk has exactly the difficulty of real integration. Integrated gradients behave differently by dimension: in one dimension they reduce to evaluating endpoints, while in two dimensions even one coordinate can already encode one-dimensional integration when the relevant derivative is provided. The same lens separates the costs in related explanation and certification methods. Interventional SHAP is integration-hard even with two features, succinct coalitional SHAP is #P-hard, nearest-threshold counterfactual search is NP-hard, and continuous calibration certificates combine integration with maximization. These results give a taxonomy of when population-level explanation and certification admit polynomial-time bit approximations, and what additional structure is needed for tractability.
Seeing but Not Detecting: Privacy-Preserving Scene Text Attack via Hybrid Adversarial Policy Learning
Feiran Li ⋅ Jiahao Lyu ⋅ Ke Jiang ⋅ Yu Zhou ⋅ Jinqiao Shi
Scene text encodes high-density semantic information, raising privacy risks for automated vision systems. Existing privacy-preserving adversarial methods mainly target Scene Text Recognition (STR) by inducing character-level misrecognition, yet text regions remain detectable, enabling manual inspection or downstream recovery. We argue that stronger privacy requires suppressing Scene Text Detection (STD) during localization, hiding text as background to bypass subsequent OCR pipelines. To this end, we propose **AR-Attacker** (**A**ttention-guided **R**einforcement learning text **Attacker**), a black-box attack framework that generates sparse, visually imperceptible perturbations to effectively suppress STD outputs. We formulate the attack as a Markov Decision Process (MDP) and design an attention-guided action space, which leverages a frozen surrogate model to focus perturbations on salient text regions. Furthermore, AR-Attacker employs a discrete-continuous hybrid policy optimized with Proximal Policy Optimization (PPO), which jointly selects perturbation locations and magnitudes. Empirical results show that AR-Attacker surpasses state-of-the-art baselines by 30% in text removal rate while simultaneously reducing query counts by 40%. Remarkably, the framework maintains high visual fidelity ($\textit{PSNR} > 45$ dB), advancing the "visible to humans, invisible to machines" privacy paradigm through a robust RL-based approach.
Seeing or Rationalizing? Scene-Evidence-Guided Chain-of-Thought for Faithful 3D Multimodal Reasoning
Guoqing Ma ⋅ Lihao Guo ⋅ Chuangxin Zhao ⋅ Yihao Zhong ⋅ Mingqi Yuan ⋅ Shan Yu
3D multimodal language models are becoming a foundation for embodied agents, indoor robotics, and human–AI interaction, but their answers and explanations often remain weakly tied to the specific objects and spatial relations observed in a 3D scene. This paper aims to make 3D reasoning more faithful by reducing language-prior rationalization and forcing reasoning chains to depend on verifiable scene evidence. We propose Anti-Rationalization Chain-of-Thought (AR-CoT), a plug-in training and decoding framework that represents each reasoning step with an explicit evidence pointer and scores candidate answer-chain pairs by their scene-evidence gain. AR-CoT contrasts scene-conditioned rationales with a frozen text-only prior, verifies local object and relation claims, and uses decoy-twin scene pairs to encourage reasoning chains to change when the relevant spatial evidence changes. Experiments on standard 3D QA and grounding benchmarks, as well as MSQA, Beacon3D, and shortcut-sensitive stress tests, show that AR-CoT consistently improves multiple 3D MLLM backbones while strengthening chain grounding, decoy-twin contrast, and scene-swap sensitivity. These results suggest that anti-rationalization offers a practical and verifiable route toward more accurate, interpretable, and scene-faithful 3D multimodal reasoning.
Seeing Through the Chain: Understanding and Mitigating Hallucinations in Multimodal Large Reasoning Models
Hao Fang ⋅ Jinyu Li ⋅ Jiawei Kong ⋅ Tianqu Zhuang ⋅ Kuofeng Gao ⋅ Bin Chen ⋅ Shu-Tao Xia
While multimodal large reasoning models (MLRMs) have exhibited impressive capabilities, they remain prone to hallucinations, and effective solutions remain underexplored. In this paper, we investigate the underlying causes of hallucinations in MLRMs and accordingly propose a reasoning-centric training framework for mitigation. Specifically, we find that introducing reasoning mechanisms exacerbates models' reliance on language priors and overlooks visual inputs, leading to CoTs with reduced visual cues but redundant text tokens. Guided by the Information Bottleneck (IB) principle, we propose selectively filtering redundant thinking tokens to obtain a more compact, signal-efficient CoT that preserves task-relevant information while suppressing noise. We also observe that the quality of the reasoning trace largely determines whether hallucination emerges in final answers. To leverage this, we introduce a reasoning-enhanced preference optimization scheme that constructs training pairs using high-quality AI feedback. We further propose generating high-quality negative samples for contrastive optimization via a multimodal hallucination-inducing mechanism that exposes model hallucination behaviors through carefully crafted visual and textual inducers. By modeling CoT chains as intermediate representations between inputs and final answers, we establish an IB-based theoretical justification for our method. Extensive experiments reveal consistent hallucination reduction across diverse MLRMs and benchmarks.
Seen-Constrained Model-Order Selection for Unknown-$K$ Generalized Category Discovery
Mingfu Yan ⋅ Jiancheng Huang ⋅ HAIPENG LUO ⋅ Yi Huang ⋅ Yuqi Peng ⋅ Qiang Zhou ⋅ Shifeng Chen
Generalized category discovery (GCD) becomes substantially more challenging when the number of novel classes is unknown. We argue that, after the backbone is fixed, this stage should be treated as seen-constrained model-order selection rather than as an auxiliary clustering detail. The selector must estimate the class count without novel labels while remaining consistent with the labeled seen categories available in the standard GCD setting. We instantiate this view by constructing a nested merge path from a seen-constrained overclustering and selecting the coarsest near-optimal cut under observable surrogate decision risks. These risks combine seen-class consistency, cross-seed stability, unlabeled-side compactness, and a tiny-cluster degeneracy penalty. The resulting protocol separates class-count error, downstream clustering quality, and robustness. Where an aligned public unknown-$K$ baseline is available, namely on SelEx, the selector reduces class-count error on all four fine-grained datasets, including a reduction from 342 to 15 classes on Herbarium-19. Broader-domain diagnostics further distinguish ImageNet100, where the default selector can underestimate the class count toward the seen-class regime, from CIFAR100, where the selector landscape exposes an accuracy-count operating-point trade-off. The mixed accuracy outcomes reveal distinct count-, path-, and representation-level bottlenecks, supporting unknown-$K$ selection as a decision layer that should be evaluated separately from representation learning.
Selective Off-Policy Reference Tuning with Plan Guidance
Duc A Le ⋅ Tien-Phat Nguyen ⋅ Thien H Nguyen ⋅ Linh Ngo ⋅ Trung Le
Reinforcement learning from verifiable rewards improves reasoning by reinforcing sampled solutions that receive positive outcomes, but group-relative methods such as GRPO become silent on hard prompts where every sampled rollout is wrong. These zero-reward prompts are often the most informative failures: each comes with a verified reference solution, yet uniformly imitating the full trace mixes core reasoning decisions with routine algebra, formatting, and surface wording. We propose Selective Off-Policy Reference Tuning with Plan Guidance (SORT), an auxiliary repair update that leaves GRPO rollout generation unchanged. For each failed prompt, SORT extracts a reference-derived reasoning plan and scores every reference token twice under the same model, once with only the problem and once with the plan added to the context. Tokens whose probabilities rise under plan conditioning are treated as structurally informative and receive larger Dynamic Fine-Tuning updates, while tokens outside the model's support remain controlled by the base probability. We formalize this mechanism with a ground-truth plan model and prove that the SORT weight approaches an oracle structural-token weight as the extracted plan better approximates the true plan-conditioned distribution. Across three instruction-tuned backbones and eight in-distribution and out-of-distribution reasoning benchmarks, SORT consistently improves over GRPO and guidance-based baselines, with the largest gains on the weakest model where zero-reward failures are most frequent.
Group-relative RL training (GRPO) samples a small group of parallel rollouts for every training prompt and uses their within-group reward spread to compute per-trajectory advantages. In agentic environments each rollout is a long multi-turn dialogue with one LLM call per step, so this multi-sample multiplier dominates the total training cost. When every rollout of a prompt ends with the same reward, the group has zero reward variance and contributes no gradient, so the extra rollouts add no information; such groups are common in practice (typically around 40% of all groups), so the wasted-compute fraction is substantial rather than marginal. Existing methods filter such groups at the prompt level, either after their rollouts are paid for or before any rollout begins, but both decide without using information that becomes available during the rollout itself. We instead ask whether the in-group divergence between the partial trajectories at an intermediate step can already predict that the group will be zero-variance: when the parallel rollouts have already converged on the same action prefix, the group is on track to produce a single reward, and we can stop early. We propose a one-parameter gate that stops a group when the mean pairwise prefix edit distance between its partial action sequences falls below a threshold. On a 60-iteration on-policy GRPO run on ALFWorld with Qwen2.5-7B, averaged over four random seeds, the gated arm finishes 10.7% faster in wall-clock (bootstrap 95% CI excludes 0) and shifts held-out success rate on 50 unseen tasks by +2.5 pp, with the held-out gain tracing to a measurable reduction in zero-advantage gradient-batch dilution.
Self-Evolution Reasoner: Continuous Optimization via Policy-Intrinsic Exploration and Exploitation
Wenhang Shi ⋅ Yiren Chen ⋅ Zhe Zhao ⋅ Jinhao Dong ⋅ WEI LU ⋅ Xiaoyong Du
Reinforcement learning with verifiable rewards (RLVR) enables large language models (LLMs) to continuously self-improve their reasoning capabilities. Compared to RLVR's scalar rewards, recent self-supervised on-policy distillation methods leverage denser supervision signals. However, both paradigms suffer from exploration bottlenecks, struggling to escape local optima when the policy fails to sample valid solutions. Relying on external demonstrations can bypass this issue; however, the off-policy nature hinders genuine self-improvement and the mismatched-distribution reasoning can not be fully exploited. In this work, we propose Self-Evolution Policy Optimization (SEPO), a framework that drives continuous reasoning capability evolution via guided exploration and in-distribution exploitation. Specifically, SEPO utilizes labels to elicit correct reasoning from the policy, unlocking the exploration potential. By treating current policy conditioned on this in-distribution reasoning as a self-teacher, SEPO distills the newly discovered solution back into the policy. Through this iterative cycle of policy-intrinsic exploration and exploitation, SEPO achieves robust and continuous capability evolution. Across scientific, tool-use, and complex logical reasoning domains, SEPO consistently achieves the best performance under identical training time or steps, demonstrating superior efficiency and convergence ceilings. Notably, SEPO maintains stable improvements even in scenarios where baselines severely underperform or fail entirely, such as on complex tasks or with small-scale models.
Self-Improving World Modelling with Latent Actions
Yifu QIU ⋅ Zheng Zhao ⋅ Waylon Li ⋅ Yftah Ziser ⋅ Anna Korhonen ⋅ Shay Cohen ⋅ Edoardo Ponti
Internal modelling of the world, predicting transitions between previous states and next states under actions, is essential to reasoning and planning for LLMs and VLMs. Learning such models typically requires costly action-labelled trajectories. We propose SWIRL, a self-improvement framework that learns from state-only sequences by treating actions as a latent variable and alternating between Forward World Modelling (FWM) and an Inverse Dynamics Modelling (IDM). SWIRL iterates two phases: (1) Variational Information Maximisation, which updates the FWM to generate next states that maximise conditional mutual information with latent actions given prior states, encouraging identifiable consistency; and (2) ELBO Maximisation, which updates the IDM to explain observed transitions, effectively performing coordinate ascent. Both models are trained with reinforcement learning (specifically, GRPO) with the opposite frozen model's log-probability as a reward signal. We provide theoretical learnability guarantees for both updates, and evaluate SWIRL on LLMs and VLMs across multiple environments: single-turn and multi-turn open-world visual dynamics and synthetic textual environments for physics, web, and tool calling. SWIRL achieves gains of 16% on Aurora-Bench, 28% on ByteMorph, 16% on WorldPredictionBench, and 14% on StableToolBench.
Self-Recognition Finetuning can Reverse and Prevent Emergent Misalignment
Arush Tagade ⋅ Shaoheng Zhou ⋅ Jiaxin Wen ⋅ Shi Feng
Emergent misalignment (EM) has been linked to the activation of misaligned persona vectors and evil character traits, suggesting that EM operates through disruption of the model's aligned character rather than direct learning of harmful content. Motivated by this connection, we study self-generated text recognition (SGTR) finetuning as a character-targeted intervention that is distinct from existing in-training defenses. We conduct two-stage finetuning experiments across three models (GPT-4.1, Qwen2.5-32B-Instruct, Seed-OSS-36B-Instruct) and multiple EM datasets to compare SGTR finetuning against benign finetuning baselines (correct domain-specific data, general knowledge, and word counting) to find it an effective defense in both reversal and prevention settings. We find that all interventions produce comparable EM reversal, but only when restoring capabilities that EM had degraded. For prevention, only SGTR finetuning consistently reduces misalignment without exacerbating any individual metric, suggesting that character fortification specifically drives prevention. We provide further evidence for EM's relation to the LLM's default character by showing that EM finetuning induces diversity into the LLM's identity self-reports, artificially corrupting self-recognition exacerbates misalignment caused by EM finetuning, and that removing the model's identity-bearing system prompt substantially reduces the effect of EM finetuning. Together, these findings reframe EM not as the adoption of a coherent misaligned persona but as the destabilization of aligned character.
Semiparametrically Efficient Inference for Kernel Measures of Noise Heterogeneity
Jakub Wornbard ⋅ Zikai Shen ⋅ Dimitri Meunier ⋅ Arthur Gretton
We develop semiparametrically efficient inference for kernel measures of noise heterogeneity in additive noise models. In many applications, the regression function is estimated using flexible machine learning methods. Downstream procedures based on the resulting residuals can then inherit first-stage bias: regression error may induce spurious dependence between covariates and residuals, invalidating the assumptions needed for standard analysis. We construct a novel Hilbert-valued one-step estimator of the kernel covariance operator between covariates and residuals. Our estimator yields bootstrap-calibrated tests for residual independence and goodness of fit in additive noise models, while also providing asymptotically efficient confidence intervals for the kernel dependence measure under noise heterogeneity. The framework extends to settings with additional covariates, enabling inference on distributional heterogeneity of residual noise across treatment groups. Simulations show improved calibration and power relative to naive plug-in residual methods.
Sequential Probability Assignment against Smoothed Adversaries with Unknown Base Measure
Ziyi Liu ⋅ Dan Roy
Smoothed online learning has recently been studied as a way to bypass hardness results for the fully adversarial setting, which can be overly pessimistic. In this framework, the adversary is constrained to generate contexts from distributions having a bounded density with respect to some fixed base measure $\mu$. Prior work makes the strong assumption that $\mu$ is known to the learner, with notable exceptions including Block et al. (2024) and Blanchard (2025). In this paper, we study sequential probability assignment (a.k.a., online learning with log loss) with smooth, well-specified data in the more general setting where $\mu$ is \emph{unknown}, without the Lipschitz condition imposed in the work of Block et al. (2024) and Blanchard (2025). We prove regret upper bounds against both adaptive and oblivious smoothed adversaries: the adaptive bound is controlled by empirical $\ell_\infty$ Hellinger entropy, while the oblivious bound improves this dependence to empirical $\ell_2$ Hellinger entropy, the notion that has been shown to characterize the complexity of learning with i.i.d. data (Bilodeau et al., 2023). Both bounds are obtained through algorithms based on truncated empirical Hellinger covers. We complement these results with a general lower bound, showing that our upper bounds are essentially tight for some natural classes, which implies a separation in the difficulty of smoothed online learning between regimes where $\mu$ is known and where it is unknown.
Shaping Useful Noise: Energy Distributions Predict Visual Pretraining Quality
Ching Lam Choi ⋅ Antonio Torralba ⋅ Phillip Isola ⋅ Stefanie Jegelka
Synthetic data can scale visual pretraining, but current selection criteria only partially explain which datasets train strong representations. We propose using the energy profile of a dataset---the distribution of log-probability levels it occupies under a normalised real-image density model---as an audit and curation principle for synthetic data. To operationalise this view, we introduce Energy-Level Distribution Distance (Eld), which compares real and synthetic energy distributions. Real images span sparse high-likelihood scenes and dense low-likelihood textures, so useful synthetic data should cover this energy range rather than collapse to likely but low-variation samples. Across diverse synthetic datasets, Eld exposes mismatches missed by FID and predicts pretraining performance beyond FID, LPIPS variation, and feature-space recall, including across CNN and ViT self-supervised settings. We then use Eld to resample and mix synthetic sources, matching the real-image energy histogram while preserving diversity and improving downstream transfer.
Verifiable differential privacy enables public auditing of privacy-preserving aggregate computations, addressing the trust deficit in standard implementations. However, existing verifiable counting approaches that distribute trust across multiple provers often require each prover to contribute independent noise. This duplicates the noise added to the aggregate and substantially degrades utility, making accurate private counting difficult in distributed settings. We introduce a shared-noise paradigm for verifiable differentially private counting. Instead of duplicating noise across provers, our approach generates a single noise sample in additive-shared form among mutually distrustful provers, while Pedersen commitments let a public verifier audit consistency without learning the noise. This yields shared-noise, verifier-auditable instantiations of two standard noise families: the Shared-Noise Verifiable Binomial Mechanism (SN-VBM) and the Shared-Noise Verifiable Laplace Mechanism (SN-VLM). We prove that both mechanisms satisfy differential privacy and public verifiability, including completeness, soundness against malicious provers, and zero knowledge against a malicious verifier colluding with corrupted provers. Among the two instantiations, SN-VLM is the more scalable option: it avoids the quadratic growth in shared randomness required by the binomial construction under strict privacy budgets. Experiments on preprocessing load, randomness-generation time, and noise error show that SN-VLM substantially reduces cryptographic overhead while improving empirical accuracy in the evaluated privacy regimes.
Shampoo is a Kronecker-structured second-order optimizer widely used for large-scale deep learning, yet existing regret analyses exhibit an unfavorable dependence on the gradient rank, leading to guarantees that are uniformly weaker than those of optimizers such as Adagrad and One-sided Shampoo. We provide a uniformly and substantially sharper regret bound for the full two-sided Shampoo update, yielding bounds that can compete with full-matrix Adagrad and more recent structured variants such as One-sided Shampoo. By further analysis, we also establish the first sublinear convergence guarantee for the commonly implemented variant Shampoo$^2$ by introducing a principled scaling factor. Our analysis holds for a continuum of exponent choices $p \in (0, \frac{1}{2}]$, also providing the first sublinear guarantees for Shampoo$^{4p}$ with exponents other than $p=1/4$.
Shepherd: A Runtime Substrate Empowering Meta-Agents with a Formalized Execution Trace
Simon Yu ⋅ Derek Chong ⋅ Ananjan Nandi ⋅ Dilara Soylu ⋅ Jiuding Sun ⋅ Christopher D Manning ⋅ Weiyan Shi
As LLM-based agentic systems grow more complex, they increasingly rely on meta-agents: higher-order agents that act on other agents, much like managers supervise employees. Yet existing agentic runtimes expose execution only as static environmental states, limiting the kinds of live and post-hoc interventions a meta-agent requires. To unlock these capabilities, we introduce Shepherd, a functional programming model that formalizes meta-agent operations on target agents as functions, with core operations mechanized in Lean. Shepherd records every agent–environment interaction as a typed event in a principled Git-like execution trace where any past state can be cheaply forked and replayed. Even at scale, Shepherd forks the agent process and its filesystem 5x faster than Docker, with >95% prompt-cache reuse on replay. We exemplify Shepherd's versatility in three use cases: (1) runtime intervention, where a live supervisor improves pair coding pass rate from 28.8% to 54.7% on CooperBench; (2) counterfactual meta-optimization, where a meta-agent branches to explore alternative paths, beating MetaHarness and GEPA across four benchmarks by up to 11 points with up to 58% lower wall-clock; and (3) Tree-RL training, where a meta-agent forks rollouts at chosen turns, lifting TerminalBench-2 from 34.2% to 39.4% on Qwen3.5-35B-A3B. These use cases show Shepherd is a performant and efficient substrate for programming meta-agents; we open-source it to advance future research.
Shifting the Gradient: Understanding How Defensive Training Methods Protect Language Model Integrity
Satchel Grant ⋅ Victor Gillioz ⋅ Jake Ward ⋅ Tom McGrath
Defensive training methods such as positive preventative steering (PPS) and inoculation prompting (IP) offer surprising results through seemingly similar processes: both add trait-inducing objects to large language models (LLMs) during training, and both defend the LLM against acquiring the trait. The surprising success of these methods comes with the question: how do they work? Are PPS and IP doing the same thing? We provide behavioral and mechanistic comparisons of these two methods using "evilness" as a case-study trait. Our central finding is that PPS and IP achieve their defensive benefits through distinct mechanisms. Behaviorally, we show that neither PPS nor IP operates through a purely associative mechanism; and PPS can both defend against trait acquisition and actively reduce pre-existing expression, whereas IP is ineffective in models that were previously finetuned to express the trait. This behavioral divergence is reflected mechanistically: PPS shifts the activation gradient towards an attenuating direction along the PPS vector axis. When the PPS vector is aligned with a trait-expressing axis, it can reverse the gradient pressure, reducing rather than increasing activation along that axis. In contrast, IP continues to resist a precise mechanistic account. Direct cosine similarity analyses reveal that IP has a characteristically different gradient signature than PPS, and qualitative analyses reveal IP's gradient to be more diffuse. Furthermore, IP reduces the next-token prediction loss on trait-expressing data where PPS need not, consistent with the notion that IP "explains away" the trait-expression in the training data. Taken together, our analyses reveal distinct mechanisms by which each method operates and highlight open questions about IP's mechanistic picture.
SIGA: Scientific Simulation Coding Agent Adapter- A Geophysics Case Study
Matthew Ho ⋅ Brian Z. Liu ⋅ Jixuan Chen ⋅ Lianhui Qin
Frontier LLMs are increasingly capable of expert-level scientific reasoning, but using them as reliable scientific agents requires more than general reasoning ability. Scientific simulation is a central test case: modern simulators are configured through executable interfaces such as XML decks, input scripts, and namelists that function as domain-specific languages tied to the simulator's internal API. Translating a researcher's natural-language intent into a runnable configuration remains a recurring expert bottleneck. We study how much of this bottleneck a general-purpose LLM coding harness, Claude Code, can absorb when wrapped with a Simulator-Interface Grounding Adapter (SIGA): a package of skills, tools, and workflow control flow that grounds the agent's outputs in simulator documentation, schema, and example libraries. We instantiate SIGA for GEOS, an open-source multiphysics simulator used in CO₂-storage and induced-seismicity research. A Resolution-IV factorial surfaces three benefits over vanilla Claude Code. First, SIGA improves reliability, reducing across-seed variance by roughly 40× by preventing unparseable or empty decks on a hard tail of compound multi-physics tasks. Second, it improves quality, raising mean structural similarity by about +7 percentage points on the same hard tail. Third, a self-evolved variant matches the best hand-designed cell with roughly 16% fewer tool calls, suggesting that automatic SIGA discovery is tractable. A preliminary human baseline finds that geoscience-domain-expert volunteers new to GEOS take between 8 and 36 times as long as the agent on a representative task. An explicit human-consultation tool is used in only about 3% of under-specified trials, with the agent instead relying on the on-disk example library as a cheaper retrieval substitute. A small OpenFOAM transfer study indicates that the recipe is not specific to GEOS XML: the same stop-hook component again dominates reliability, and the best SIGA cell outperforms both vanilla Claude Code and a constrained Foam-Agent lint-only baseline. We close with SIGA-design recommendations grounded in failure modes that our best configuration still does not fix.
SignRot: LLM Quantization with Massive Outlier-Aware Sign-Adjusted Rotation
Jaehun Gim ⋅ Jaewoo Kim ⋅ Gyuwan Kim ⋅ Seo Yeon Park
While quantization compresses LLMs and accelerates inference, outlier activations impede efficient low-bit representation. In particular, massive outliers cause substantial performance degradation. Although rotation-based methods that rely on randomized Hadamard transforms for online inference aim to mitigate this, they consistently produce fragmented distributions with empty bins when encountering massive outliers. This fragmentation severely undermines uniform quantization performance. Moreover, when massive outliers occur in down-projection activations, the rotation must be applied at inference time. This constraint prevents training-based approaches such as SpinQuant from learning the online rotation at this position, as a learned dense matrix would forgo the Fast Walsh--Hadamard Transform and require a full matrix multiplication during inference. To address this, we propose SignRot (Sign-adjusted Rotation), which identifies massive outlier channels, and constructs a sign adjustment vector that unifies the signs of the corresponding Hadamard row, collapsing the bimodal distribution into a unimodal profile amenable to quantization. As a lightweight, drop-in improvement compatible with any Hadamard-based quantization method, SignRot requires only a single calibration forward pass and introduces no additional inference overhead. Applying SignRot to QuaRot reduces the Wikitext2 perplexity to 5.88 and increases the average zero-shot reasoning accuracy to 72.73, surpassing SpinQuant on the LLaMA3-70B model.
SimpleEvol: Efficient Intelligence Conversion via Less Human Prior in Automated Heuristic Design
Jianghan Zhu ⋅ Cong Zhang ⋅ Rongjie Zhu ⋅ CHI ZHANG ⋅ Zhiguang Cao
Large language models (LLMs) have emerged as powerful tools for automated heuristic design (AHD), enabling iterative generation and refinement of heuristics. However, the dominant paradigm embeds LLMs as narrow, fixed components, such as crossover or mutation, within heavily hand‑engineered evolutionary frameworks. We argue this misapprehends LLMs. It treats them as specialized tools rather than general reasoners, constrains them to low‑level operations, and underutilizes their autonomy. Moreover, the extensive human priors in these frameworks violate the bitter lesson principle that general methods scaling with computation surpass hand‑crafted solutions. This raises a key question: which AHD framework designs best convert stronger LLM capabilities into better heuristics? To address this, we propose metrics for AHD framework handcraftedness (AHI) and intelligence conversion efficiency (ICE). Evaluating ten LLMs across three challenging optimization problems, we obtain a notable finding that frameworks with fewer human priors consistently yield higher ICE. Based on this finding, we propose SimpleEvol, which removes nearly all structural constraints and allows the LLM to operate autonomously. SimpleEvol consistently achieves the highest ICE, often by a large margin. Our results challenge the trend toward complex AHD pipelines and point to a simpler and more model‑centric alternative, suggesting that reducing human priors is a more effective strategy to scale up with model intelligence.
Simple yet Effective Semi-supervised Knowledge Distillation from Vision-Language Models via Dual-Head Optimization
Seongjae Kang ⋅ Dong Bok Lee ⋅ Hyungjoon Jang ⋅ Sung Ju Hwang
Semi-supervised learning (SSL) has emerged as a practical solution for addressing data scarcity challenges by leveraging unlabeled data. Recently, vision-language models (VLMs), pre-trained on massive image-text pairs, have demonstrated remarkable zero-/few-shot performance that often surpasses SSL approaches due to their exceptional generalization capabilities. This gap motivates us to question: how can we effectively harness the powerful generalization capabilities of VLMs into task-specific models? Knowledge distillation (KD) offers a natural framework for transferring VLM capabilities, but we identify that it suffers from gradient conflicts between supervised and distillation losses. To address this challenge, we propose Dual-Head Optimization (DHO), which introduces dual prediction heads for each distinct signal. We observe that DHO resolves gradient conflicts, enabling improved feature learning compared to single-head KD baselines, with practical benefits of minimal computational overhead and test-time hyperparameter tuning without retraining. Extensive experiments across 15 datasets show that DHO consistently outperforms KD baselines, often outperforming teacher models with smaller student models. DHO also achieves new state-of-the-art performance on both in-distribution ImageNet semi-supervised learning and out-of-distribution generalization across ImageNet variants. We will publicly release our code and model checkpoints to facilitate future research.
Human motion prediction combines the tasks of trajectory forecasting and human pose prediction. For each of the two tasks, specialized models have been developed. Combining these models for holistic human motion prediction is non-trivial, and recent methods have struggled to compete on established benchmarks for individual tasks. To address this, we propose a simple yet effective transformer-based model for human motion prediction. The model employs a stack of self-attention modules to effectively capture both spatial dependencies within a pose and temporal relationships across a motion sequence. This simple, streamlined, end-to-end model is sufficiently versatile to handle pose-only, trajectory-only, and combined prediction tasks without task-specific modifications. We demonstrate that this approach achieves state-of-the-art or on-par results across all tasks through extensive experiments on a wide range of benchmark datasets, including Human3.6M, AMASS, ETH-UCY, and 3DPW. Code will be released.
Simulation-free Unbalanced Dynamic Optimal Transport with General Growth Penalty
Junda Ying ⋅ Yuxuan Wang ⋅ Bowen Yang ⋅ Peijie Zhou ⋅ Lei Zhang
Inferring cellular dynamics from unpaired single-cell snapshots requires modeling both state transitions and population growth or death. Unbalanced dynamic optimal transport (UDOT) addresses this by penalizing growth along transport paths, making the choice of growth penalty a key way to encode biological priors on proliferation and apoptosis. However, existing UDOT solvers either rely on computationally expensive NeuralODE simulations or depend on analytical solutions of conditional paths, restricting their efficiency solely to quadratic penalties, i.e. Wasserstein-Fisher-Rao (WFR) geodesics. To enable an efficient UDOT solver for general growth penalties, we first show that concave growth penalties lead to degenerate solutions where growth and transport are separated. We then introduce Simulation-free Unbalanced Dynamic Optimal transport (SUDO), a simulation-free framework for UDOT with general non-quadratic convex growth penalties. SUDO learns the conditional paths and transport costs, solves the induced semi-coupling problem, and subsequently leverages unbalanced flow matching to achieve a simulation-free solution. On WFR benchmarks, SUDO matches the accuracy of efficient, analytical solution-driven algorithms while outperforming simulation-based methods in computational speed. Beyond WFR, SUDO supports asymmetric penalties that encode proliferation-dominant priors and produce more plausible trajectories and growth estimates on synthetic and single-cell datasets.
SIRAS: Sibling-Relative Advantage Shaping for Reinforcement Learning from Verifiable Rewards
Zairun Yang ⋅ Jun Xu ⋅ Lei Liang ⋅ Huajun Chen ⋅ Qiang Zhang
Reinforcement learning from verifiable rewards (RLVR) trains reasoning language models with group-relative objectives such as GRPO. However, it leaves two sources of signal underused: every token in a rollout receives the same response-level advantage, and groups in which all rollouts succeed or all fail have zero group-relative advantage and are typically discarded. We observe that same prompt rollouts are not independent reward labels but sibling attempts: they share scaffolding, diverge at uncertain decisions, and sometimes reconverge. In this paper, we propose Sibling-Relative Advantage Shaping (SIRAS), a lightweight advantage-shaping method that exploits this structure. SIRAS segments rollouts into chunks, aligns sibling chunks with banded dynamic time warping, and reshapes the flat GRPO advantage across tokens via soft-divergence weights that upweight chunks distinguishing a rollout from its siblings and downweight shared ones. This reshaping preserves each rollout’s token-averaged GRPO advantage exactly and reduces to GRPO when weights are constant. The same sibling structure produces calibrated residuals for all-correct and all-wrong groups, reclaiming training signal that group-relative baselines discard. Across six reasoning benchmarks, SIRAS yields the largest gains on AIME-style competition math, including +12.2 on AIME25 with Qwen3-8B and +23.3 on AIME26 in our backbone-scaling study, without extra rollouts, process reward models, or step annotations.
SituRecBench: A Benchmark for Situated Recommendation in 3D Interactive Environments
Jingtong Yang ⋅ Dongding LIN ⋅ Jian Wang ⋅ Wenjie Li
Recommender systems are transitioning from static retrieval engines into interactive assistants grounded in physical environments. However, current paradigms largely treat recommendation as a decoupled ranking task, neglecting the situated complexities of spatial navigation, visual grounding, and multi-turn social interaction. To bridge this gap, we introduce SituRecBench, a novel interactive benchmark designed to evaluate situated recommendation agents within high-fidelity 3D environments. SituRecBench reformulates the recommendation process as a closed-loop interaction between a recommendation assistant and a user with evolving latent preferences and realistic behaviors. Our benchmark features diverse real-world shopping scenarios (e.g., supermarkets and home furnishings stores), requiring assistants to integrate preference elicitation, spatial reasoning, and long-horizon conversation. We introduce a simulator-in-the-loop evaluation protocol along with a multi-dimensional metric taxonomy that quantifies communication fluency, recommendation utility, and navigation effectiveness. Through extensive evaluation of state-of-the-art vision-language models, we reveal critical challenges in situated recommendation, such as inefficient preference elicitation, coherent conversation transition, and poor spatial grounding. SituRecBench provides a rigorous, reproducible foundation for advancing the frontier of situated recommendation in interactive, visually-grounded environments.
Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
David Schmotz ⋅ Luca Beurer-Kellner ⋅ Sahar Abdelnabi ⋅ Maksym Andriushchenko
LLM agents are evolving rapidly, powered by code execution, tools, and the recently introduced agent skills feature. Skills allow users to extend LLM applications with specialized thirdparty code, knowledge, and instructions. Although this can extend agent capabilities to new domains, it creates an increasingly complex agent supply chain, offering new surfaces for prompt injection attacks. We identify skill-based prompt injection as a significant threat and introduce Skill-Inject, a benchmark evaluating the susceptibility of widely-used LLM agents to injections through skill files. Skill-Inject contains 202 injection-task pairs with attacks ranging from obviously malicious injections to subtle, context-dependent attacks hidden in otherwise legitimate instructions. We evaluate frontier LLMs on Skill-Inject, measuring both security in terms of harmful instruction avoidance and utility in terms of legitimate instruction compliance. Our results show that today’s agents are highly vulnerable with up to 80% attack success rate with frontier models, often executing extremely harmful instructions including data exfiltration, destructive action, and ransomwarelike behavior. They furthermore suggest that this problem will not be solved through model scaling or simple input filtering, but that robust agent security will require context-aware authorization frameworks.
SkillMaster: Toward Autonomous Skill Mastery in LLM Agents
Min Yang ⋅ Jinghua Piao ⋅ Xu Xia ⋅ Yongshun Gong ⋅ Yong Li
Skills provide an effective mechanism for improving LLM agents on complex tasks, yet in existing agent frameworks, their creation, refinement, and selection are typically governed by external teachers, hand-designed rules, or auxiliary modules. As a result, skills remain external resources to be invoked, rather than capabilities that agents can develop, adapt, and internalize through experience. To endow LLM agents with autonomous skill mastery, we propose SkillMaster, a training framework that teaches agents to create new skills, refine existing skills, and select accumulated skills during task solving. This capability is achieved through three key designs. First, we train agents through trajectory-informed skill review, teaching agents to propose, update, or retain skills based on evidence from completed episodes. Second, each candidate skill edit is designed to be evaluated by its counterfactual utility on related probe tasks, providing a direct learning signal for training skill-editing decisions. Third, we introduce DualAdv-GRPO, which separately estimates advantages for task-solving actions and skill-editing decisions, stabilizing joint training across task solving and skill management. Experiments on ALFWorld and WebShop show that SkillMaster improves the overall success rate over state-of-the-art baselines by 8.8% and 9.3%, respectively, achieving the best performance among all compared methods. Further analysis reveals a marked shift in agent capability: agents trained with SkillMaster can identify skill failures, refine procedural knowledge from trajectory evidence, and transfer improvements to future tasks with limited skill-bank edits. Overall, SkillMaster moves LLM agents beyond mere skill use toward self-improving agents capable of developing, adapting, and applying their own skill repertoires. Our code is released at https://anonymous.4open.science/r/skillmaster-F4B0.
SkipSR: Faster Super-Resolution with Token Skipping
Rohan Choudhury ⋅ Shanchuan Lin ⋅ Jianyi Wang ⋅ Hao Chen ⋅ Qi Zhao ⋅ Feng Cheng ⋅ Lu Jiang ⋅ Kris Kitani ⋅ László Jeni
Diffusion-based super-resolution (SR) is a key component in video generation and video restoration, but is slow and expensive, limiting scalability to higher resolutions and longer videos. Our key insight is that many regions in video are inherently low-detail and gain little from refinement, yet current methods process all pixels uniformly. To take advantage of this, we propose SkipSR, a simple framework for accelerating video SR by identifying low-detail regions directly from low-resolution input, then skipping computation on them entirely, only super-resolving the areas that require refinement. This simple yet effective strategy preserves perceptual quality in both standard and one-step diffusion SR models while significantly reducing computation. In standard SR benchmarks, our method achieves up to 60$\%$ faster end-to-end latency than prior models on 720p videos with no perceptible loss in quality.
SliceWorld: A Predictive and Controllable World-State Model for CT Report Generation
Yuanhe Tian ⋅ Yan Song
CT report generation (CTRG) requires models to summarize three-dimensional anatomical context and pathological findings from hundreds of axial slices. Existing methods typically learn a direct image-to-text mapping, providing limited mechanisms for modeling how CT evidence evolves across slices or how reports respond to controlled changes in latent lesion-related factors. We propose SliceWorld, a CT-specific world-state framework that treats an axial CT scan as an ordered sequence along the z-axis. SliceWorld encodes prefix CT evidence into factor-aware latent states containing anatomy, lesion, and uncertainty components, and projects these states into world tokens used for multi-step future-slice feature prediction, lesion-factor intervention, and LLM-based report generation. The model is first pretrained on CT slice sequences with predictive, factor-aware, and counterfactual objectives, and is then fine-tuned on paired CT-report data. Experiments on M3D-Cap and CT-RATE show that SliceWorld improves natural language generation metrics and clinically oriented automatic evaluation. Further analyses demonstrate multi-horizon future-slice prediction, measurable factor alignment, reduced-slice robustness, and selective lesion-sensitive report modulation.
SLIDERS: Systematic Reviews via Automated Evidence Synthesis and Reconciliation
Harshit Joshi ⋅ Priyank Shethia ⋅ Jadelynn Dao ⋅ Monica Lam
Systematic reviews -- which requires comprehensive evidence collection and synthesis from large document corpora in response to targeted research questions -- are foundational in finance, social sciences, and other technical fields. Manual construction of evidence tables is labor-intensive, and recent LLM-based assistants relying on embedding or keyword based search often fail to meet the coverage standards of systematic reviews. We introduce SLIDERS, a novel LLM-based methodology for systematic reviews, by automatically assembling evidence tables tailored to research questions. In addition to extracting structured data from documents, SLIDERS can extract full-text excerpts that serve as direct evidence or as provenance for structured data. Core to SLIDERS is an automated evidence reconciliation agent that writes code to analyze and reconcile extracted evidence, bringing together information fragmented across documents, resolving inconsistencies across excerpts, and synthesizing overlapping findings into a coherent evidence table. In addition, SLIDERS allows users to ask follow-up questions in natural language to further explore the assembled evidence. We evaluate SLIDERS on three systematic-review-style tasks over large document collections. SLIDERS outperforms the best-performing baseline across benchmarks, remains near 90\% accuracy across 6M-11M-token corpora. On two new follow-up analysis benchmarks \system{} can answer 77.9% and 58.3% followup questions accurately.
SLOPE: Optimistic Potential Landscape Shaping for Model-based Reinforcement Learning
Yao-Hui Li ⋅ Zeyu Wang ⋅ Xin Li ⋅ Wei Pang ⋅ Yingfang Yuan ⋅ Zhengkun Chen ⋅ Boya Zhang ⋅ Riashat Islam ⋅ Alex Lamb ⋅ Yonggang Zhang
Model-based reinforcement learning (MBRL) is sample-efficient but struggles in sparse reward settings. A critical bottleneck arises from the lack of informative gradients in sparse settings, where standard reward models often yield flat landscapes that struggle to guide planning. To address this challenge, we propose Shaping Landscapes with Optimistic Potential Estimates (SLOPE), a novel framework that shifts reward modeling from predicting sparse scalars to constructing informative potential landscapes. SLOPE employs optimistic distributional regression to estimate high-confidence upper bounds, which amplifies rare success signals and ensures sufficient exploration gradients. Evaluations on 30+ tasks across 5 benchmarks and real-world robotic deployments, demonstrate that SLOPE consistently outperforms leading baselines in fully sparse, semi-sparse, and dense rewards. The project code is available at https://anonymous.4open.science/r/SLOPE-CF13.
SLS-Bench: A Benchmark for Incident Log Summarization with Synthetic Observability Data
Valerii Iakovlev ⋅ Markus Spanring ⋅ Martin Flechl
As IT systems grow in complexity, automated agent-driven incident response becomes increasingly critical. A key capability for these agents is summarizing large log streams into human-readable reports. Further progress requires meaningfully evaluating this function, yet existing benchmarks leave end-to-end summarization from large log streams underexplored due to the prohibitive cost and complexity of collecting real-world data. We introduce SLS-Bench, a benchmark for incident log summarization featuring 90 tasks grounded in real-world incident reports, each with a synthetic log stream and a reference summary. To construct SLS-Bench, we developed a data generation pipeline that converts incident reports into structured incident models, which guide the generation of synthetic logs and reference summaries. This pipeline combines narrative grounding, explicit incident modeling, generator-verifier loops, and statistical comparisons with real-world log streams to produce realistic and controllable observability data. Our evaluation of 14 LLMs on SLS-Bench shows that while frontier models can produce useful summaries, they do so at substantial cost and latency, and open-weight models lag significantly. SLS-Bench provides a framework for systematic evaluation of log analysis agents and a scalable blueprint for future synthetic observability benchmarks.
SNACK: A Sequential Notation Framework for Probabilistic Graph Generation
Hohyun Kim ⋅ Hyesung Kim ⋅ Min-hwan Oh ⋅ Seunggeun Lee
We introduce SNACK, a string representation for probabilistic graph generation that enables sequence models to generate and model graph-structured data. SNACK linearizes graphs by encoding the nonzero entries of the lower-triangular adjacency matrix, establishing a bijective correspondence between valid sequences and ordered adjacency matrices. This design supports invalid logit masking for enforcing structural constraints, including chemical valence and ordering constraints, and enables principled string-based GFlowNet training by explicitly accounting for node-ordering symmetries. Across standard graph-generation tasks and molecular distribution-learning benchmarks, SNACK achieves competitive or state-of-the-art sample quality, while substantially improving the throughput of GFlowNet training relative to graph-based generation. Probing analyses further show that models trained on SNACK encode meaningful local and global graph structure despite receiving only sequential supervision. These results position SNACK as a practical framework for probabilistic graph generation with deep sequence models.
SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning
Philip Schroeder ⋅ Thomas Weng ⋅ Karl Schmeckpeper ⋅ Eric Rosen ⋅ Stephen Hart ⋅ Ondrej Biza
Vision-language models (VLMs) have shown impressive capabilities across diverse tasks, motivating efforts to leverage these models to supervise robot learning. However, when used as evaluators in reinforcement learning (RL), today’s strongest models often fail under partial observability and distribution shift, enabling policies to exploit perceptual errors rather than solve the task. We introduce SOLE-R1 (Self-Observing LEarner), a video-language reasoning model explicitly designed to serve as the sole reward signal for online RL. Given only raw video observations and a natural-language goal, SOLE-R1 performs per-timestep spatiotemporal chain-of-thought (CoT) reasoning and produces dense estimates of task progress that can be used directly as rewards. To train SOLE-R1, we develop a large-scale video trajectory and reasoning synthesis pipeline that generates temporally grounded CoT traces aligned with continuous progress supervision. This data is combined with foundational spatial and multi-frame temporal reasoning, and used to train the model with a hybrid framework that couples supervised fine-tuning with RL from verifiable rewards. Across four different simulation environments and a real-robot setting, SOLE-R1 enables zero-shot online RL from random initialization: robots learn previously unseen manipulation tasks without ground-truth rewards, success indicators, demonstrations, or task-specific tuning. SOLE-R1 succeeds on 24 unseen tasks and substantially outperforms strong vision-language rewarders, including GPT-5 and Gemini-3-Pro, while exhibiting markedly greater robustness to reward hacking.
Solver-Aware Decompositions for Programming-by-Example: When Dividing Requires Knowing how to Conquer
Janis Zenkner ⋅ Tobias Sesterhenn ⋅ Tim Grams ⋅ Christian Bartelt
Decomposition-based Programming-by-example (PBE) scales performance by splitting tasks into subtasks that a learned synthesizer solves: a decomposer predicts intermediate subgoals, and a synthesizer generates programs conditioned on them. Current approaches train the decomposer to imitate ground-truth (GT) subgoals, implicitly treating decomposition quality as intrinsic to the task. We challenge this assumption: for bounded solvers with fixed inductive biases, GT decompositions reflect the annotator’s factorization choices – not the solver’s search dynamics. A decomposer trained to match GT decompositions may therefore propose subgoals that are logically valid yet intractable for the solver. We propose Solver-Aware Decomposition (SAD), a training framework that retains supervised training on GT subgoals as a structural scaffold, while additionally optimizing the decomposer via direct feedback from a frozen synthesizer. Subgoals are rewarded based on the synthesizer’s loss on the target program – a signal of subtask difficulty that encourages decompositions the solver can act on. Our experiments reveal an accuracy paradox: higher agreement with GT decompositions does not improve synthesis success – even though the synthesizer was trained on the very same GT data the decomposer is optimized to mimic. SAD instead learns decompositions that trade GT alignment for solver tractability, yielding consistent gains in synthesis and end-to-end task accuracy across two PBE domains. Moreover, SAD solves tasks that a GT decomposition oracle fails – empirical evidence that GT decompositions are not universally optimal for bounded solvers, and that decomposition quality is solver-relative, not intrinsic.
Source-Causal Control of Historical Context in Longitudinal Radiology Report Generation
Dmitry Lvov ⋅ Ilya Pershin
Longitudinal radiology report generation increasingly conditions on prior images and reports. This history is clinically valuable, but it creates a source-validity failure mode: a generator can carry prior-only findings into the current report. Weintroduce source-causal robustness, a stress-test and control framework that holds the current chest X-ray fixed while intervening on historical context. On a full eligible MIMIC-CXR longitudinal study (18,716 pairs; 93,580 generations across five source conditions), historical context consistently trades fewer missed findings for more unsupported positives. Matched wrong-history, the key negative control, increases proxy-defined false positives by +0.366 per report while reducing false negatives by −0.129; random wrong-history yields an even larger false positive stress signal (+0.465 FP/report). We propose the Source-Causal Finding Router (SCFR), a source-validity controller trained with within-evaluation subject exclusion. Across three held-out evaluations, SCFR repairs matched wrong-history false-positive inflation by −0.343, −0.550, and −0.457 per report, passes Holm corrected primary testing at the bootstrap resolution floor (Holm-adjusted p ≤ 0.0012), and satisfies fixed current-only parity margins. Cross-generator stress panels with RadInfer and Libra, together with ReXPref-Prior, RadGraph2, Chest ImaGenome, ablations, and case review, support the mechanism. These results establish a protocol-defined state-of-the-art source-causal robustness/control result for longitudinal radiology report generation under matched wrong-history stress tests.
SP$^2$ec: Adaptive Self-Speculative Decoding for Vision-Language Models
Yuqi Huang ⋅ Xingyao Li ⋅ Yunlong Hou ⋅ Fengzhuo Zhang ⋅ Jiachun Pan ⋅ Vincent Tan
Speculative Decoding (SD) has shown strong success in accelerating large language model inference, but its efficiency remains limited for Vision-Language Models (VLMs). The diverse input modalities in VLMs make it difficult for a small trained drafter to achieve consistently high throughput across modalities. We address this limitation with an adaptive self-speculative decoding method. In particular, we adaptively select the skipped layers of VLMs to form the draft model that best fits the given query. Our selection algorithm SP$^2$ec utilizes the single-peak structure of the throughput with respect to the number of selected layers. Our algorithmic design explicitly minimizes the novel notion of wall-time regret for the given query, corresponding to the empirical wall-time latency. Theoretical guarantees on the wall-time regret show that SP$^2$ec is adaptive to the per-prompt single-peak structure. We conduct extensive experiments with Qwen3-VL and LLaVA-1.5 models across visual and textual tasks, which demonstrates the efficacy of the proposed SP$^2$ec.
Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance
Ziyun Zeng ⋅ Yiqi Lin ⋅ Guoqiang Liang ⋅ Mike Zheng Shou
In recent years, open-source efforts like Senorita-2M have propelled video editing toward natural language instruction. However, current publicly available datasets predominantly focus on local editing or style transfer, which largely preserve the original scene structure and are easier to scale. In contrast, Background Replacement, a task central to creative applications such as film production and advertising, requires synthesizing entirely new, temporally consistent scenes while maintaining accurate foreground-background interactions, making large-scale data generation significantly more challenging. Consequently, this complex task remains largely underexplored due to a scarcity of high-quality training data. This gap is evident in poorly performing state-of-the-art models, e.g., Kiwi-Edit, because the primary open-source dataset that contains this task, i.e., OpenVE-3M, frequently produces static, unnatural backgrounds. In this paper, we trace this quality degradation to a lack of precise background guidance during data synthesis. Accordingly, we design a scalable pipeline that generates foreground and background guidance in a decoupled manner with strict quality filtering. Building on this pipeline, we introduce Sparkle, a dataset of ~140K video pairs spanning five common background-change themes, alongside Sparkle-Bench, the largest evaluation benchmark tailored for background replacement to date. Experiments demonstrate that our dataset and the model trained on it achieve substantially better performance than all existing baselines on both OpenVE-Bench and Sparkle-Bench. Our proposed dataset, benchmark, and model are fully open-sourced at https://showlab.github.io/Sparkle/.
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning
Seokju Cho ⋅ Ryo Hachiuma ⋅ Abhishek Badki ⋅ Hang Su ⋅ Byung-Kwan Lee ⋅ Chan Hee Song ⋅ Sifei Liu ⋅ Subhashree Radhakrishnan ⋅ Seungryong Kim ⋅ Frank Wang ⋅ Min-Hung Chen
Spatial reasoning, the ability to determine where objects are, how they relate, and how they move in 3D, remains a fundamental challenge for vision-language models (VLMs). Tool-augmented agents attempt to address this by augmenting VLMs with specialist perception modules, yet their effectiveness is bounded by the action interface through which those tools are invoked. In this work, we study how the design of this interface shapes the agent's capacity for open-ended spatial reasoning. Existing spatial agents either employ single-pass code execution, which commits to a full analysis strategy before any intermediate result is observed, or rely on a structured tool-call interface that often offers less flexibility for freely composing operations or tailoring the analysis to each task. Both designs offer limited flexibility for open-ended, complex 3D/4D spatial reasoning. We therefore propose SpatialClaw, a training-free framework for spatial reasoning that adopts code as the action interface. SpatialClaw maintains a stateful Python kernel pre-loaded with input frames and a suite of perception and geometry primitives, letting a VLM-backed agent write one executable cell per step conditioned on all prior outputs, enabling the agent to flexibly compose and manipulate perception results and adapt its analysis to both intermediate text and visual observations and the demands of each problem. Evaluated across 20 spatial reasoning benchmarks spanning a broad range of static and dynamic 3D/4D spatial reasoning tasks, SpatialClaw achieves 59.9% average accuracy, outperforming the recent spatial agent by +13.6 points, with consistent gains across six VLM backbones from two model families without any benchmark- or model-specific adaptation.
Speak Early: Speculative Speech Generation for Low-Latency SpeechLLMs
Tianyuan Jiang ⋅ Yongheng Deng ⋅ Ju Ren
Speech large language models (SpeechLLMs) have emerged as a new interface for human–AI interaction, enabling real-time spoken responses. Time-to-First-Speech (TTFS) is a critical latency metric in such interactive SpeechLLM systems, directly determining perceived responsiveness. While prior work reduces TTFS by accelerating autoregressive decoding, we show that diffusion-based speech synthesis contributes a comparable portion of the latency. This reveals a fundamental limitation of existing approaches that optimize decoding alone, and indicates that optimizing either LLM decoding or speech synthesis in isolation is insufficient. This paper proposes $\texttt{SpecSpeak}$, a new SpeechLLM inference paradigm that overlaps LLM decoding with speech synthesis to reduce the TTFS latency. $\texttt{SpecSpeak}$ introduces a lightweight draft model to generate a prefix quickly. Instead of waiting for fully verified tokens like speculative decoding, $\texttt{SpecSpeak}$ triggers speech synthesis early using draft tokens, while a target model performs parallel verification and correction. To handle inconsistencies, $\texttt{SpecSpeak}$ introduces a verification-aware reuse mechanism that adaptively preserves intermediate diffusion states based on prefix agreement, enabling efficient correction without restarting synthesis. Extensive experiments demonstrate that $\texttt{SpecSpeak}$ reduces the TTFS latency of SpeechLLMs by 18\%--30\%, while preserving response correctness and speech quality. Code will be released upon acceptance.
SpecBlock: Block-Iterative Speculative Decoding with Dynamic Tree Drafting
Weijie Shi ⋅ Qiang Xu ⋅ fan deng ⋅ Yaguang Wu ⋅ Jiarun Liu ⋅ Yehong Xu ⋅ Hao Chen ⋅ Jia Zhu ⋅ Jiajie Xu ⋅ Xiangjun Huang ⋅ jian yang ⋅ Xiaofang Zhou
Speculative decoding accelerates LLM inference by drafting a tree of candidate continuations and verifying it in one target forward. Existing drafters fall into two camps with opposite weaknesses. Autoregressive drafters such as EAGLE-3 preserve dependence along each draft path but call the drafter once per tree depth, making drafting a non-trivial share of per-iteration latency. Parallel drafters cut drafter calls by predicting multiple future positions in one forward, but each position is predicted without seeing the others, producing paths the verifier rejects. In this paper, we propose SpecBlock, a block-iterative drafter that combines path dependence with cheap drafting. Each drafter forward produces $K$ dependent positions and we call this a block. The draft tree grows through repeated block expansions. Two mechanisms explicitly carry path dependence to keep later draft positions accurate. Within each block, a layer-wise shift carries the previous position's hidden state into every decoder layer. Across blocks, each new block can start from any position of the previous block, inheriting its hidden state to extend the path. To spend verifier budget where acceptance is likely, a co-trained rank head replaces the fixed top-$k$ tree by allocating per-position branching during drafting. To avoid training the drafter on prefixes it never produces at inference, a valid-prefix mask drops the loss at later positions once an earlier one is wrong. Beyond static drafting, a cost-aware bandit at deployment uses free verifier feedback to update the drafter selectively, only when the expected throughput gain exceeds the update cost. Experiments show that SpecBlock improves mean speedup by $8$--$13\%$ over EAGLE-3 at $44$--$52\%$ of its drafting cost, and cost-aware adaptation extends this lead to $11$--$19\%$.
Spectral Adaptive Repositioning for Flow-Based Single-Cell Perturbation Modeling
Shourya Verma ⋅ Mengbo Wang ⋅ Simran Kadadi ⋅ Shahin Mohammadi ⋅ Nadia Lanman ⋅ Ananth Grama
Predicting transcriptome-wide cellular responses to genetic and biomolecular perturbations is central to functional genomics and therapeutic target discovery. Existing methods assign gene positional encodings derived from arbitrary orderings or static co-expression masks built from unperturbed cells, and cannot capture the functional rewiring induced by perturbations. We present SPEAR, a generative framework built on two complementary positional mechanisms. Spectral positional encoding grounds each gene's position in the geometry of the co-expression network via continuous coordinates derived from the graph Laplacian. Adaptive repositioning then derives each gene's position from its perturbation-conditioned hidden state at every layer, allowing the attention genes exert on one another to reorganize without explicit supervision. Evaluated on genetic, pharmacological, and cytokine perturbation benchmarks, SPEAR demonstrates excellent recovery of differential gene-expression, distributional structure, and perturbation-specific identifiability. SPEAR also captures latent positional relationships that closely map to known transcription factor-target interactions.
Spectral Insights from the Unconstrained Feature Model for Neural Multi-Output Regression
George Andriopoulos ⋅ Bimarsha Adhikari ⋅ Soyuj J Basnet ⋅ Juan Guevara ⋅ Li Guo ⋅ Keith Ross
Multi-output regression can be approached either by fitting one joint vector-valued predictor or by training separate univariate predictors for each target coordinate. This holistic-versus-separate choice is classical in the largely non-neural multi-output regression literature, but is less characterized analytically for neural networks, where joint prediction arises naturally through shared features and vector-valued output layers. Because neural networks are highly non-linear and high-dimensional, exact analysis is notoriously difficult. We study two questions using the Unconstrained Feature Model (UFM) as a tractable surrogate for exact analysis: $\textbf{(1)}$ how does joint-output regression compare with independent coordinate-wise regression under regularization? $\textbf{(2)}$ when can target whitening or normalization improve original-scale MSE? Under UFM, the minimal training MSE depends only on the regularization product and the spectrum of the target covariance. This yields two theoretical results: joint-output regression has no larger training MSE than independent coordinate-wise regression under matched regularization, and whitening or normalizing targets can help or hurt original-scale MSE depending on the average target variance. Experiments on neural networks trained with standard weight decay across robotic imitation-learning and autonomous-driving datasets support these theoretical results, with test-MSE trends following the same qualitative pattern. Our results show that simplified analytical surrogates can provide useful guidance for neural multi-output regression.
Speculative Self-Distillation enables Efficient Knowledge Internalization
Shayan Talaei ⋅ Agam Bhatia ⋅ Arshia Soltani Moakhar ⋅ Jonas Hübotter ⋅ Amin Saberi ⋅ Azalia Mirhoseini
Many practical deployments require language models to internalize knowledge that was absent from pretraining, such as proprietary corpora, continually updated facts, or specialized domain skills. Self-distillation has emerged as a natural approach for this task. A model conditioned on the source material acts as the teacher, and an unconditioned student is trained to reproduce the teacher's behavior closed-book. We propose Speculative Self-Distillation (SSD), an efficient per-token mixed-policy method for self-distillation. The student generates each rollout by default, while the teacher intervenes only at positions where its next-token distribution substantially diverges from the student's. This design keeps training rollouts close to the student's test-time behavior, unlike off-policy distillation with teacher-generated prefixes, but avoids spending teacher compute on the many low-signal tokens encountered by fully on-policy methods. In effect, SSD concentrates supervision precisely where privileged context changes the prediction. On three closed-book QA benchmarks measuring fact recall, knowledge updates, and skill acquisition, SSD matches on-policy distillation with 45\% fewer supervised tokens on average and consistently outperforms off-policy distillation on the held-out test set. SSD therefore establishes a new accuracy-efficiency Pareto frontier for self-distillation. Code is available at https://anonymous.4open.science/r/Speculative-Self-Distillation/.
Splatting the Invisible: Geometry and Appearance Scene Completion from Sparse Views
Wesley Khademi ⋅ Fuxin Li
Novel view synthesis methods such as 3D Gaussian Splatting degrade under sparse-view settings, suffering from an inability to extrapolate beyond observed regions. To enable high fidelity reconstruction and rendering from sparse views, inpainting of sparsely observed areas and completion of occluded regions is required. Unlike existing approaches that perform these tasks using 2D inpainting diffusion models or geometric-only 3D completion models, we propose a fully 3D scene completion framework, called InvisiSplat, that performs geometry and appearance completion directly in 3D. Our method consists of a Gaussian Splat VAE which learns a sparse point cloud latent space that predicts Gaussian Splat parameters from the latents, as well as a two-step flow model that learns to generate samples conditioned on the point cloud latents, different from other approaches that use tokens without 3D coordinates. Our flow model utilizes alternating local 3D attention and global attention, which balances the needs for precise local geometry and large-scale spatial context. Across several datasets, we demonstrate that our method produces 33% - 82% higher fidelity 3D geometry over prior work while attaining competitive performance on appearance metrics when compared to existing approaches.
Stable and Granular Policy Optimization for Generative Recommendation
Muyu Zou ⋅ Haibo Xing ⋅ Hao Deng ⋅ Zhezheng Hao ⋅ Jinxin Hu ⋅ Jiawei Chen ⋅ Yun Chen ⋅ Yu Zhang ⋅ Xiaoyi Zeng
Generative recommendation (GR), which directly generates sequential semantic IDs for items, has recently attracted increasing attention in recommender systems. Given that core recommendation metrics are inherently evaluated at the sequence level, reinforcement learning (RL) serves as a natural optimization framework. However, adapting recent group-based RL algorithms to GR presents notable challenges: token-level methods such as GRPO exhibit much more severe training instability, while sequence-level methods such as GSPO enforce uniform token weighting and thus ignore the heterogeneous roles of tokens. To address this dilemma, we propose Stable and Granular Policy Optimization (SGPO), a group-based RL algorithm tailored for GR. SGPO maintains a sequence-level ratio to promote optimization stability while injecting fine-grained token-wise advantages. By incorporating token log-probabilities and reward signs, SGPO explicitly upweights high-confidence tokens in positive trajectories to reinforce successful retrieval, while strictly penalizing them in negative trajectories to encourage exploration, and uses an adaptive scaling factor to bound the variance of reshaped advantages. Experiments on two public benchmarks and a billion-scale industrial dataset show that SGPO consistently outperforms six strong RL baselines and, when deployed on a commercial platform, achieves a 2.36% lift in click-through rate (CTR) and a 1.95% increase in advertising revenue in online A/B testing.
StableAvatar: Ultra-Long Audio-Driven Avatar Video Generation
Shuyuan Tu ⋅ Yueming Pan ⋅ Yinming Huang ⋅ Xintong Han ⋅ Zhen Xing ⋅ Qi Dai ⋅ Chong Luo ⋅ Zuxuan Wu ⋅ Yu-Gang Jiang
Current diffusion models for audio-driven avatar video generation struggle to synthesize long videos with natural audio synchronization and identity consistency. This paper presents StableAvatar, the first end-to-end video diffusion transformer that synthesizes ultra-long videos. We observe that the primary reason preventing existing models from generating long videos is their audio modeling, which typically relies on third-party off-the-shelf audio embeddings that are introduced into the diffusion model via cross-attention. Since current diffusion backbones lack any audio-related priors, they cannot correctly couple audio cues with latent dynamics, causing systematic latent distribution deviations at each segment. These deviations accumulate across segments, gradually shifting the latent trajectory and producing noticeable quality drift. To address this, StableAvatar introduces a novel Timestep-aware Audio Adapter that prevents error accumulation via timestep-aware modulation. During inference, we propose a novel Audio Native Guidance Mechanism to further enhance the audio synchronization by leveraging the diffusion’s joint audio-latent prediction as a dynamic guidance signal. To enhance the smoothness of the ultra-long videos, we introduce a Dynamic Weighted Sliding-window Strategy that fuses latent over time. Experiments on benchmarks show the effectiveness of StableAvatar both qualitatively and quantitatively.
StableHand: Quality-Aware Flow Matching for World-Space Dual-Hand Motion Estimation from Egocentric Video
Huajian Zeng ⋅ Chaohua Yao ⋅ Yuantai Zhang ⋅ Jiaqi Yang ⋅ Rolandos Alexandros Potamias ⋅ Xingxing Zuo
Recovering world-space 4D motion of two interacting hands from egocentric video is a fundamental capability for supervising robot policy learning, where wrist trajectories track the end-effector and finger articulations specify the grasp pose. Two major challenges arise in this setting: hands frequently leave the camera view for extended periods due to head motion, and persistent hand–object interactions cause severe occlusions of one or both hands. Existing methods uniformly condition on noisy hand motion observations without accounting for their per-frame reliability, leading to substantial performance degradation. Our key insight is that accurate world-space hand motion esimation is tightly coupled with the quality of per-frame hand observations. To this end, we decompose the quality of hand motion observations extracted from an off-the-shelf hand pose estimator into four channels: wrist global translation and finger articulations for both hand. We propose StableHand, a quality-aware flow-matching framework conditioned on these four-channel quality signals, which are predicted by a learned quality network. We naturally incorporate the quality signals into the flow-matching process through a per-channel forward schedule, a quality-adjusted velocity target, AdaLN modulation of the DiT denoiser, and a quality-aware ODE initialization. This unified generative process preserves high-quality observations while reconstructing unreliable ones using a learned bimanual motion prior. Experiments on HOT3D and ARCTIC, two egocentric benchmarks featuring long missing-hand spans and persistent hand–object occlusions, show that StableHand achieves state-of-the-art performance across all reported metrics, reducing W-MPJPE by 20-25% compared to the strongest baseline, with the largest gains on heavily occluded ARCTIC sequences. Code and pretrained models will be released upon acceptance.
Stage Light is Sequence$^2$: Multi-Light Control via Imitation Learning
Zijian Zhao ⋅ Dian Jin ⋅ Zijing Zhou ⋅ Xiaoyu Zhang
Music-inspired Automatic Stage Lighting Control (ASLC) has gained increasing attention in recent years due to the substantial time and financial costs associated with hiring and training professional lighting engineers. However, existing methods suffer from several notable limitations: the low interpretability of rule-based approaches, the restriction to single-primary-light control in music-to-color-space methods, and the limited transferability of music-to-controlling-parameter frameworks. To address these gaps, we propose SeqLight, a hierarchical deep learning framework that maps music to multi-light Hue-Saturation-Value (HSV) space. Our approach first customizes SkipBART, an end-to-end single primary light generation model, to predict the full light color distribution for each frame, followed by hybrid Imitation Learning (IL) techniques to derive an effective decomposition strategy that distributes the global color distribution among individual lights. Notably, the light decomposition module can be trained under varying venue-specific lighting configurations using only mixed light data and no professional demonstrations, thereby flexibly adapting across diverse venues. In this stage, we formulate the light decomposition task as a Goal-Conditioned Markov Decision Process (GCMDP), construct an expert demonstration set inspired by Hindsight Experience Replay (HER), and introduce a three-phase IL training pipeline, achieving strong generalization capability. To validate our IL solution for the proposed GCMDP, we conduct quantitative analysis to compare model performance across different training phases, demonstrating that our design effectively improves performance and generalization capacity. Furthermore, we also conduct a human study to evaluate SeqLight by comparing it with competitive baselines on music-conditioned light generation tasks across different music styles. The results show that SeqLight achieves the best overall preference scores in both in-domain and out-of-domain settings. The code and trained parameters of this paper is provided at the anonymous repository https://anonymous.4open.science/r/SeqLight-23EE .
State-Resolving Attention for Length-Extrapolating Transformers
Zicheng Liu ⋅ Di Wu ⋅ Jintao Chen ⋅ Di Huang
Long-context failures are often framed as retrieval failures: a model must find a relevant token among many distractors. However, many practical long-context settings are better understood as state-tracking problems: dialogue histories contain corrections to earlier preferences, code traces repeatedly assign values to variables, document histories revise previous claims, and tool-use trajectories update the current task state. In these settings, a model must not only identify relevant tokens, but also resolve which relevant occurrence is currently valid. We propose Sieve Attention, a position-encoding-free attention mechanism for length-extrapolating Transformers. SRA decomposes attention into two operations: content filtering and temporal state resolution. For each query, SRA first applies a sparse content filter over the full prefix to identify candidate state updates, and then applies a sequential hazard allocation process only over the selected candidates. This resolves repeated updates according to their relative order without relying on external positional encodings, allowing irrelevant intervening tokens to be ignored before recency is applied. We formalize the mechanism and show that its final attention support is contained in the sparse candidate set, that filtered distractors receive exactly zero mass, and that attention ratios among selected candidates are independent of global sequence length and absolute position. Experiments on repeated associative recall and copy tasks show strong extrapolation far beyond the training context length. On long-context evaluation and language-model pretraining, SRA improves long-context behavior while remaining competitive with standard Transformer baselines. These results suggest that content-first, time-second state resolution is a useful inductive bias for long-context models that must track evolving information.
StateTree: Enhancing Long-term Dialogue Reasoning via Reinforcement Learning
Naen Xu ⋅ Wanqing Cui ⋅ Yibo Hu ⋅ Shixin Hong ⋅ Hengyu An ⋅ Meiguang Jin ⋅ Junfeng Ma ⋅ Tianyu Du
Large language models deployed as personalized assistants must reason over long, evolving interaction histories. However, in long-term dialogue reasoning, relevant evidence is scattered across sessions, preferences may be revised over time, and standard long-context training fails to address these challenges under data scarcity and prohibitive computational costs. We propose StateTree, a data-driven RL pseudo-task that constructs a challenging auxiliary task from scarce dialogues with verifiable ground truth. StateTree augments multi-session dialogues with a tree-structured path-tracing task: key-value records are embedded across sessions to form a binary tree. Solving the task requires the model to traverse from root to leaf by retrieving records across sessions and comparing timestamps to resolve branches, then recover the hidden target question among distractor leaves. We apply curriculum RL training progressively increasing tree depth and introduce a compositional variant whose edges carry step-level reasoning fragments, training the model to compose partial cues into coherent queries. Trained on 10K-token contexts, StateTree generalizes to 128K tokens without full-length RL costs and exhibits capabilities including cross-session retrieval, temporal reasoning, knowledge update, and compositional multi-hop reasoning. StateTree outperforms both SFT and RL-based baselines while preserving short-context general reasoning. StateTree-7B achieves gains up to +23.60% on LongMemEval (128k), and StateTree-14B reaches 59.00% accuracy on LongMemEval, surpassing QwenLong-L1-32B (45.20%).
SteadyThought: Mitigating LLM Under-Thinking via Thought-Level Preference Optimization
Zhenyue Peng ⋅ Lin Yang ⋅ Jie Wang ⋅ Xiaoqi Ni ⋅ Hanzhu Chen ⋅ Xize Liang ⋅ Yanjun Wu ⋅ RujingWang ⋅ Kaiwen Zhou ⋅ Jianye Hao
Flexible switching between reasoning trajectories (i.e., thoughts switching) has significantly enhanced the reasoning capabilities of Large Reasoning Models (LRMs). However, existing models often switch excessively yet fail to sustain promising reasoning thoughts---a phenomenon termed ''under-thinking''. While recent efforts suppress switching to mitigate this, such over-correction may discard valuable trajectories. To address this challenge, we propose Steady Thought (ST), a novel thought-level preference optimization framework. ST formalizes under-thinking as a preference issue at switching points, forcing single-path continuations from identified points to yield high-quality reasoning. Then, ST performs thought-level preference optimization by treating the newly generated response as preferred and the original one as dis-preferred. Experiments across multiple models and datasets show that ST effectively reduces token consumption while maintaining or even improving accuracy. It reduces output length by up to 63.8% while improving accuracy by up to 11.4%. Further analysis suggests that ST helps models acquire a generalizable ability to reduce redundant switching across domains and languages.
Steer2Edit: From Activation Steering to Component-Level Editing
Chung-En Sun ⋅ Ge Yan ⋅ Zimo Wang ⋅ Lily Weng
Steering methods influence Large Language Model behavior by identifying semantic directions in hidden representations, and are typically realized through inference-time activation interventions that apply a fixed, global modification to the model's internal states. While effective, such interventions often induce unfavorable attribute--utility trade-offs under strong control, as they ignore the fact that many behaviors are governed by a small and heterogeneous subset of model components. To alleviate the trade-offs, we propose Steer2Edit, a theoretically grounded, training-free framework that transforms steering vectors from inference-time control signals into diagnostic signals for component-level rank-1 weight editing. Instead of uniformly injecting a steering direction during generation, Steer2Edit selectively redistributes behavioral influence across individual attention heads and MLP neurons, yielding interpretable edits that preserve the standard forward pass and remain compatible with optimized parallel inference. Across multiple tasks including safety alignment, truthfulness promotion, and reasoning efficiency, Steer2Edit consistently achieves a superior attribute--utility trade-offs: at matched downstream performance, it outperforms the strongest steering baseline by 17.2% in safety, 9.8% in truthfulness, and 12.2% in reasoning length reduction. Overall, Steer2Edit provides a principled bridge between representation steering and weight editing by translating steering signals into interpretable, training-free parameter updates.
Stochastic Heat Diffusion Models
Marc Law ⋅ Xiaoyu Wang ⋅ Lucas Berry ⋅ Anthony L Caterini ⋅ Gabriel Loaiza-Ganem
This work introduces a family of generative diffusion models based on the stochastic heat equation. The deterministic heat equation has natural connections with kernels over discrete structures and has been successfully applied in machine learning to discriminative tasks for various types of data such as graphs. We extend it to a stochastic version by formulating our diffusion as a Gaussian process whose transition kernel can be written in closed form for both the forward- and reverse-time processes. Our diffusion process propagates noise smoothly among neighbours under the guidance of a Laplacian matrix that describes the global structure of the input. Following a flow matching-based approach, we train a denoising neural network that takes noisy representations as input and gradually removes noise to generate clean samples. We experimentally show that, by exploiting the Laplacian during training, our decoder is able to recover the global structure of the input.
STRABLE: Benchmarking Tabular Machine Learning with Strings
Gioia Blayer ⋅ Myung Jun Kim ⋅ Félix Lefebvre ⋅ Lennart Purucker ⋅ Alan Arazi ⋅ Eilam Shapira ⋅ Roi Reichart ⋅ Frank Hutter ⋅ Marine Le Morvan ⋅ David Holzmüller ⋅ Gael Varoquaux
Benchmarking tabular learning has revealed the benefit of dedicated architectures, pushing the state of the art. But real-world tables often contain string entries, beyond numbers, and these settings have been understudied due to a lack of a solid benchmarking suite. They lead to new research questions: Are dedicated learners needed, with end-to-end modeling of strings and numbers? Or does it suffice to encode strings as numbers, as with a categorical encoding? And if so, do the resulting tables resemble numerical tabular data, calling for the same learners? To enable these studies, we contribute STRABLE, a benchmarking corpus of 108 tables, all real-world learning problems with strings and numbers across diverse application fields. We run the first large-scale empirical study of tabular learning with strings, evaluating 445 pipelines. These pipelines span end-to-end architectures and modular pipelines, where strings are first encoded, then post-processed, and finally passed to a tabular learner. We find that, because most tables in the wild are categorical-dominant, advanced tabular learners paired with simple string embeddings achieve good predictions at low computational cost. On free-text-dominant tables, large LLM encoders become competitive. Their performance also appears sensitive to post-processing, with differences across LLM families. Finally, we show that STRABLE is a good set of tables to study ``string tabular'' learning as it leads to generalizable pipeline rankings that are close to the oracle rankings. We thus establish STRABLE as a foundation for research on tabular learning with strings, an important yet understudied area.
The standard rule in unconstrained causal policy learning is simple: treat individuals with nonnegative CATE. We show that this rule can fail when treatment eligibility is strategically manipulable. Agents may move or report covariates after a policy is announced, so treatment is assigned using manipulated covariates while welfare remains determined by baseline CATEs. This creates a mismatch between who receives treatment and who benefits from treatment. We formalize this setting as strategic causal policy learning and show that strategic welfare equals the CATE-weighted mass of baseline types that can reach the treatment region. This access-burden representation leads to three distinct strategic thresholds: a welfare threshold, a safe threshold, and a fair threshold. For linear CATE models, we derive closed-form threshold corrections. Homogeneous costs admit a single buffer correction that recovers the ideal positive-CATE allocation, while heterogeneous costs make welfare maximization, no-exploit safety, and subgroup fairness select different policies. Experiments illustrate these tradeoffs and show when group-specific corrections can restore the ideal allocation.
StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering
Ming Xie ⋅ Zizheng Huang ⋅ Xudong Tan ⋅ Chao Wang ⋅ Xiangyu Zeng ⋅ Wenxiao Wu ⋅ Tao Chen ⋅ Limin Wang ⋅ Yanwei Fu
While streaming omni-video understanding demands continuous perception and proactive, real-time interaction, this crucial area remains largely under-explored. Current omni-modal methods are inherently designed for offline settings, limiting their applicability in streaming scenarios due to two fundamental flaws. First, they lack robust mechanisms to manage continuously growing audio-visual context over long horizons and cannot autonomously initiate responses at opportune moments. Second, existing benchmarks are predominantly confined to offline, single-turn question answering, failing to capture continuous, multi-turn streaming interactions. To bridge these gaps, we propose StreamOV, a novel Streaming Omni-Video understanding framework for efficient online audio-visual reasoning with bounded memory and proactive response triggering. Specifically, StreamOV introduces a multimodal evidence-guided long-short term memory that condenses historical audio-visual context into compact informative evidence under a fixed budget. It further employs a hidden-state-driven trigger to decide when to respond, avoiding explicit silence-token generation and external routers. We also curate SOVBench, the first comprehensive benchmark for online, multi-turn omni-modal evaluation. Extensive experiments show that StreamOV achieves state-of-the-art performance across diverse streaming and omni-video benchmarks, demonstrating its effectiveness for both online and offline video understanding.
StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models
Zhe Liu ⋅ Jinghua Hou ⋅ Yuxiang Lu ⋅ Zhenya YANG ⋅ Xianzhe Fan ⋅ Junwei Luo ⋅ Junyi Li ⋅ Ruihua Han ⋅ Zhi Hou ⋅ Hengshuang Zhao
Vision-Language-Action (VLA) models have demonstrated effectiveness in robot manipulation, yet state-of-the-art models such as $\pi_{0.5}$ operate under a single-frame paradigm, limiting their ability to retain past observations and develop precise spatial perception. In this paper, we propose StreamPI, a streaming multimodal temporal modeling framework that equips single-frame VLA with temporal reasoning capability without introducing any additional parameters. One core design is instruction-anchored temporal modeling. It treats each (visual observation, language instruction) pair as an atomic temporal unit: bidirectional attention within each pair enables cross-modal fusion, while causal attention across pairs preserves autoregressive streaming inference. This ensures the language instruction serves as a persistent semantic anchor throughout task execution. To bridge the gap between synchronous training and asynchronous real-robot deployment, we introduce a random-interval streaming training strategy: a proper inter-frame interval (e.g., every 3 frames) enables faster and smoother action execution. Beyond this, randomizing the interval further improves robustness to frame-timing perturbations, supporting asynchronous deployment in practice. Furthermore, by leveraging the length extrapolation capability of the LLM backbone, StreamPI seamlessly inherits pretrained single-frame weights and supports flexible single-frame and multi-frame inference. Experiments on real-robot tasks spanning memory-dependent and precise perception scenarios, as well as the simulation benchmark LIBERO, demonstrate that StreamPI outperforms $\pi_{0.5}$ across diverse tasks.
We introduce structural causal bottleneck models (SCBMs), a novel class of structural causal models in which causal effects between high-dimensional variables are mediated by low-dimensional summary statistics, or bottlenecks. SCBMs provide a flexible framework for mechanism-specific, target-dependent dimension reduction while remaining estimable via standard learning algorithms. We prove that bottleneck variables are identifiable up to bijection from observational data, and validate this experimentally. We then study causal effect estimation in linear SCBMs under three standard adjustment strategies. For instrumental variables, we identify a population-level non-identifiability of naïve two-stage least squares when all variables are high-dimensional, and show that bottleneck representations restore identification. For mediation and backdoor adjustment, we prove that conditioning on bottleneck variables yields valid effect estimates with explicit finite-sample efficiency gains over conditioning on full high-dimensional variables. We argue that SCBMs provide a principled alternative to existing causal dimension reduction frameworks such as causal representation learning and causal abstractions.
Structural Entropy Optimized Communication for Multi-Agent Reinforcement Learning
Wei Du ⋅ Benyu Wu ⋅ Wei Guo ⋅ Zhongmin Yan ⋅ Guoxian Yu ⋅ Lizhen Cui
Communication learning is crucial for complex coordination in multi-agent reinforcement learning (MARL), yet existing methods often struggle with a fundamental dilemma between communication overhead and task performance. Typically designed without theoretical guidance, they produce unstructured and redundant communication graphs, limiting both efficiency and scalability. To address this, we propose Structural Entropy Optimized Communication (SEOC), a framework that directly optimizes inter-agent communication topology using the information-theoretic principle of Structural Entropy. SEOC guides the network to self-organize into a low-entropy state with sparse connectivity and distinct communities, enabling efficient and interpretable decentralized coordination. This is achieved through a unified mechanism that jointly enforces structural order and semantic conciseness, reducing topological complexity via a differentiable structural entropy loss while compressing messages with a node information bottleneck. Experiments on benchmarks such as the StarCraft Multi-Agent Challenge (SMACv2) demonstrate that SEOC significantly outperforms state-of-the-art methods.
Standard flow matching scales well but typically relies on an unstructured source distribution, limiting its ability to learn interpretable latent structure. Latent-variable models, by contrast, capture structure but often sacrifice generative quality. We bridge this gap by proposing Structured Coupling for Flow Matching (SCFM), a cooperative framework that augments flow matching with structured latent representation learning. By introducing structured latent variables and exogenous noise into the source, SCFM jointly learns a structured prior (via latent variable modeling) and a continuous transport map (via flow matching). It uses a shared time-dependent recognition network for both latent variable model variational inference and intermediate-time flow velocity estimation. This yields a structurally informed yet unconditional, simulation-free flow model, where the latent variable model can also assist flow sampling. Empirically, SCFM facilitates unsupervised latent representation learning for clustering, disentanglement and downstream tasks, while remaining competitive with flow matching in sample quality, showing that meaningful structure can be learned without sacrificing generative fidelity.
Structured Transforms for Low-Overhead Quantization of Language Models
Daria Cherniuk ⋅ Alexander Rudikov ⋅ Boris Kashin ⋅ Ivan Oseledets
We revisit Kashin-decomposition-based weight quantization for large language models and propose an improved algorithm with stronger convergence properties and structured, efficient orthogonal transforms. The method retains the core factorization of each weight into two components -- one with bounded infinity norm and the other with bounded infinity norm after an orthogonal transformation -- but replaces the dense random orthogonal matrix with a sign-randomized Discrete Cosine Transform (DCT), reducing the per-iteration cost from $\mathcal{O}(N^2)$ to $\mathcal{O}(N \log N)$. The proposed greedy algorithm with alternating updates guarantees the four-peak distribution required for stable 2-bit clustering of each factor and admits closed-form initialization of cluster centers, removing the multi-restart k-means bottleneck of prior work. Composed with OPTQ-style sequential error compensation and QuIP-style incoherence preprocessing, the resulting JAX pipeline is competitive with OPTQ, QuIP, QuIP-RG and a fine-tuning- and vector-quantization-free variant of QuIP\# at 4-bit per channel on OPT, Llama-2 and Pythia, with favorable wall-clock scaling. The bounded-$\ell_\infty$ factorization is also notably robust: on stress configurations where QuIP variants diverge to four-digit perplexity (Pythia-6.9B) or abort with NaNs in LDL back-substitution (Mistral-7B), Kashin-DCT remains numerically stable and stays close to FP16 baseline. At inference time, each weight decomposes into two 2-bit factor codes per channel that are structurally suited to native-2-bit hardware.
Subdata Selection: A Unified Framework for Optimal Selection and Statistical Efficiency Assessment
Min Yang ⋅ Wei Zheng ⋅ John Stufken ⋅ Ming-Chung Chang ⋅ Ting Tian ⋅ Xueqin Wang
When labeling is expensive or datasets exceed computing capacity, selecting an informative subset of data is a practical necessity. Identifying such subdata is a classical NP-hard problem due to its inherent discreteness. While many subdata selection methods have been proposed for parametric models, two fundamental challenges remain: first, how can we accurately assess the statistical efficiency of selected subdata relative to the theoretical optimum? Second, given $N$ data points and a subdata size $n$, which $n$ points should be selected to achieve high statistical efficiency? We address both challenges within a unified framework. We develop a new information-based subdata selection methodology grounded in optimal approximate design theory, yielding subdata that approaches the theoretical optimum. Our algorithm is general, accommodates arbitrary $N$ and $n$, supports multiple optimality criteria, and is accompanied by a convergence proof. Crucially, our framework provides, for the first time, tight lower and upper bounds on the statistical efficiency of subdata selected by any method, enabling accurate and rigorous benchmarking across the literature. This benchmarking capability reveals the true performance landscape of existing methods and shows that many of them have been operating substantially below their theoretical potential. The subdata produced by our methodology is highly efficient and outperforms all existing methods.
Retrieval-augmented generation (RAG) reduces hallucinations by grounding large language models in external knowledge bases. However, most existing methods focus on improving how knowledge is indexed, retrieved and integrated, neglecting a critical source of hallucination: the lack of unknown awareness. Unaware of the knowledge boundaries of the corpus, a model may stitch together fragmented context to hallucinate unsupported relations. To address this, we introduce Substrata, an unknown-aware RAG framework that makes the unknown locatable. Substrata maintains an evolving knowledge graph via Dempster--Shafer evidence fusion, enabling local graph updates instead of costly full graph reconstruction in graph-based RAG. Crucially, we apply a topological analysis method, persistent homology, to detect long-lived structural holes where the corpus fails to close relational gaps within this graph. These holes are converted into retrievable unknown annotations. During inference, Substrata performs hybrid retrieval over both factual chunks and unknown annotations, allowing the generator to access both supporting context and explicit signals of missing contextual support. Experiments show that Substrata achieves the best performance on three public multi-hop QA benchmarks, with an average of 60.66\% EM and 71.25\% F1 across HotpotQA, MuSiQue, and 2WikiMultiHopQA. On a hallucination-focused domain benchmark, it reaches 95.0\% accuracy on Spurious Relational Premise questions while reducing overall token consumption by 10.3× compared with GraphRAG.
Supervision Recovery for Time Series Anomaly Detection via Counterfactual Pairing
Yifei Gao ⋅ Tian Lan ⋅ Yimeng Lu ⋅ Xuming An ⋅ Meng Wang ⋅ Wenjun He ⋅ Yijie Li ⋅ Chen Zhang
Time series anomaly detection (TSAD) remains challenging not only because anomaly labels are scarce, but also because temporal anomalies are highly context-dependent. Existing methods often rely on unsupervised or surrogate objectives, producing anomaly scores indirectly rather than learning explicit normal--anomalous distinctions. We propose Counterfactual Pairing with Anomaly Semantics (CAPS), a supervision-recovery framework for TSAD. CAPS formulates temporal anomalies as mechanism-induced effects on normal temporal evolution and recovers matched supervision without target-domain anomaly labels. It learns transferable anomaly semantics from simulated normal--anomalous pairs, disentangles them from normal temporal structure, and organizes them into family-wise modes. Instead of directly using simulated anomalies as target positives, CAPS instantiates the learned semantics as residual-form anomaly effects on target-domain reference trajectories via residual MoE generation. The resulting matched counterparts provide supervision-aligned signals for boundary-oriented detector learning. Experiments on nine benchmark datasets show that CAPS outperforms unsupervised and injection-based baselines, is competitive with supervised-reference baselines, and provides interpretable evidence through anomaly-effect generation and expert association.
Surface Recovery Is Not State Recovery: Pressure–Recovery–Relapse in Multi-turn LLMs
Zixi Huang ⋅ Hankun Chang ⋅ Xueli Geng ⋅ Xutong Mu
Interactive language models should respond to user feedback without blindly conforming to incorrect pressure. Existing evaluations of sycophancy typically measure whether models flip their answers under pressure and treat apparent recovery after pressure removal as robustness. We show that this recovery is often superficial. We introduce the Pressure--Recovery--Relapse (PRR) protocol, a multi-turn evaluation framework that disentangles pressure-induced flipping, surface recovery, and relapse under renewed weak pressure. Across factual QA benchmarks and instruction-tuned models, PRR reveals three findings: (1) models can appear to recover at the output level while remaining vulnerable to renewed pressure; (2) surface answer correction can be dissociated from repair of the latent answer state; and (3) different internal components support distinct recovery pathways. Through causal interventions, we show that output-level steering can enforce correct answers without repairing corrupted latent readouts, whereas restoring clean hidden states repairs both surface behavior and latent states. Further component-level restoration shows that residual-stream states enable full repair, while mid-layer attention outputs provide partial latent repair. These results highlight that robust multi-turn reliability requires evaluating recovery stability and latent state repair, not merely apparent answer correctness.
SurgReasoner: Surgical Reasoning Segmentation with Dynamic Difficulty-Aware Reinforcement Learning
Dong Zhang ⋅ Cheng Xue
Surgical reasoning segmentation requires a model to infer the target instrument or anatomical structure from an implicit clinical query and produce a pixel-level mask. This setting is far more challenging than category-level surgical segmentation, as targets are specified by functional role, spatial relation, or anatomical interaction rather than explicit class names. It also exposes a key limitation of existing reinforcement learning-based visual grounding methods: standard reward normalization tends to under-optimize clinically important hard samples, such as small instruments, occluded organs, and ambiguous targets. To address this, we introduce SurgSeg-14K, a surgical reasoning segmentation benchmark with over 14K image-mask pairs across seven surgical scenarios and 30 fine-grained categories. Each sample links an implicit clinical query to a binary mask, spatial annotations, and a reasoning trajectory. We further propose SurgReasoner, which decomposes the task into two stages: clinical query interpretation for spatial prompt generation, and prompt-based mask generation with a frozen segmentation model. To improve reinforcement learning for spatial reasoning, we introduce a dynamic difficulty-aware reweighting strategy that combines intrinsic target difficulty with rollout correctness, enabling training to focus on challenging targets. Experiments on SurgSeg-14K show that SurgReasoner outperforms the strongest grounding baseline by large margins, with consistent gains especially on hard samples, suggesting a practical path toward reasoning-driven surgical visual perception.
Surrogate Calibration for Transferable Adversarial Attacks against Black-Box MLLMs
Zhewen Yao ⋅ Yao Zhu ⋅ Xiangyang Ji ⋅ Shiliang Zhang
Transfer-based attacks expose serious adversarial vulnerabilities in closed-source Multimodal Large Language Models (MLLMs) by crafting adversarial examples on off-the-shelf surrogate models. However, this paradigm often drives adversarial examples toward surrogate-specific local optima, where they deceive only the surrogate model but fail to transfer to black-box MLLMs. Although existing methods mitigate this issue at the attack level, they leave the surrogate itself unchanged. To address this gap, we propose Surrogate Calibration (SCal), a novel model-level refinement that enhances downstream transfer-based adversarial attacks in a plug-and-play manner. Specifically, SCal employs a feature-preserving loss to maintain the surrogate's representational utility, while a distributional regularizer encourages its gradient field to stay closer to the natural data manifold. With calibrated surrogates, downstream attacks generate update directions that exhibit superior generalization across black-box MLLMs. Extensive experiments demonstrate that SCal boosts current attacks to new heights across various surrogates. For instance, it elevates baseline M-Attack's ASR on GPT-5.4 Mini to 53.5\%, substantially surpassing the state-of-the-art MPCAttack (29.7\%). These results expose critical vulnerabilities that must be addressed for trustworthy multimodal systems. Code is at: https://anonymous.4open.science/r/SCal_code-B0BC
Survival Transformers for Longitudinal Data Analysis: Application to Atrial Fibrillation Risk from ECG
Rafael Silva ⋅ Maxime Sermesant
Atrial fibrillation (AFib) is the most common sustained cardiac arrhythmia and a major risk factor for stroke, yet it often remains undetected until a first clinical event. Existing deep learning approaches typically estimate AFib risk from a single, cross-sectional ECG, discarding the longitudinal information contained in serial recordings. We propose a transformer-based survival model, SurviFormer, that leverages sequences of ECG representations to model patient-specific AFib risk over time. Our approach encodes irregular inter-visit time gaps as pairwise relative attention biases within a causally masked transformer, and outputs a discrete hazard function trained with a multi-landmark strategy. We train and evaluate our model on the CODE dataset (140k patients, 517k ECGs), with external validation on MIMIC-IV. Compared to static baselines (DeepSurv and DeepHit applied to ResNet-50 embeddings) and a dynamic competing-risk model (DynamicDeepHit), our model achieves a C-index of 0.853, a mean TD-AUC of 0.867, and an IBS of 0.030 on the internal test set, and a C-index of 0.805, a mean TD-AUC of 0.848, and an IBS of 0.106 on the external cohort, demonstrating that longitudinal ECG trajectories provide prognostic value beyond a single recording.
SVG-3D: Mining Decision Boundaries with Generative Splatting Priors for Zero-Shot 3D Classification
Ryo Umagami ⋅ Shohei Ohsawa
Zero-shot 3D point cloud classifiers built on large-scale pretrained models (OpenShape, Uni3D, ReCon++, DuoduoCLIP) excel on clean synthetic benchmarks such as ModelNet40 but degrade sharply on real-world scans (ScanObjectNN, ScanNet) and corrupted point clouds (ModelNet-C)–a persistent real-world generalization gap. We propose SVG-3D, the first 3D extension of Support Vector Generation (SVG), to close this gap without retraining a more capable encoder or accessing target-domain test data. SVG-3D mines a discriminative kernel-machine support set from a frozen encoder paired with DiffSplat, a text-conditioned 3D Gaussian splatting generator. It rests on two design choices that preserve strict zero-shot rigour: (i) a dynamic pair-selection rule that picks the top-$K$ most confusable class pairs from a pseudo-confusion matrix computed entirely on generated samples, never on the test set; and (ii) a Metropolis–Hastings sampler operating on the CLIP text-embedding hypersphere with a Slerp-plus-Gaussian proposal whose density admits a closed-form ratio, yielding the same acceptance rule as SVG. On ScanObjectNN (OBJ\_ONLY/OBJ\_BG/PB\_T50\_RS) and ScanNet, SVG-3D surpasses prior zero-shot 3D classifiers built on large-scale pretrained models without consulting any test data; on ModelNet-C, a complementary corruption-aware setting in which the corruption family is assumed known further improves mean robustness over the same baselines.
SWE-Crafter: Scaling Executable Multilingual Software Engineering Data with Meta-Skill Agents
Shukai Liu ⋅ Jian Yang ⋅ Bo Jiang ⋅ Junhang Cheng ⋅ Haowen Wang ⋅ Xiao-Hui Li ⋅ Xianglong Liu
Software-engineering agents are increasingly expected to solve realistic repository-level tasks, but their training remains constrained by the limited availability of executable supervision beyond Python-centric repositories. Building such data across programming languages is difficult: repositories use different build systems, dependency managers, runtime assumptions, test interfaces, and failure modes, so separate language-specific pipelines are costly to reproduce and extend. We introduce SWE-Crafter, a unified agent-based framework for constructing executable multilingual SWE instances. SWE-Crafter keeps candidate mining, environment synthesis, fail-to-pass validation, and trajectory distillation in a shared pipeline, while language-specific construction knowledge is supplied through extensible skills induced from repository-level experience. These skills guide repository exploration and test-command synthesis while preserving repository-local evidence from documentation, manifests, CI files, and execution feedback as the final source of truth. Applying SWE-Crafter to eight programming languages yields 50,926 validated instances from 9,296 repositories and 68,851 distilled test-passing trajectories. Supervised fine-tuning on this resource produces SWE-Crafter-40B, which achieves 61.0% on SWE-bench Multilingual and 77.2% on SWE-bench Verified, showing that skill-guided construction provides scalable multilingual supervision for repository-level SWE agents.
SwitchLingua V2: Agent-Driven Code-Switching via Digital Clones
Peng Xie ⋅ Yequan Bie ⋅ Jianda Mao ⋅ Hao CHEN ⋅ Yangqiu Song ⋅ Kani Chen
Code-switching (CS) involves shifting between two or more languages during a conversation or utterance, a phenomenon often driven by social context and the identity of the speaker. The scarcity of authentic CS data remains a critical bottleneck in developing robust Automatic Speech Recognition (ASR) systems for bilingual communities. While Large Language Models (LLMs) have been employed to synthesize CS data, direct generation approaches typically produce utterances that are pragmatically unnatural and fail to capture the nuanced stylistic variations of real-world bilingual speakers. In this paper, we propose an archetype-driven code-switching dialogue generation framework via agent-based digital cloning. Moving beyond static prompting, our method facilitates autonomous conversations between multi-agent digital clones imbued with specific bilingual archetypes and established linguistic theories. The generation pipeline integrates constrained persona sampling, real-world topic injection, turn-by-turn accommodation dynamics, and multi-dimensional language checking. Applying this framework, we synthesize highly idiomatic bilingual dialogues across 14 diverse language pairs with 260K textual dialogues (1.04M turns of text samples) and convert them into over 1,558 hours of high-fidelity audio data, yielding the SwitchLingua V2 dataset. To rigorously validate our approach and the generated data, we employ a comprehensive evaluation protocol including both advanced LLM assessments and human bilingual expert evaluations. Extensive experiments demonstrate that SwitchLingua V2 produces CS dialogues that are structurally diverse, pragmatically motivated, and socially situated, holding the potential to offer a highly scalable and reliable pathway for advancing multilingual ASR research.
Symmetry Guarantees Statistic Recovery in Variational Inference
Daniel Marks ⋅ Dario Paccagnan ⋅ Mark van der Wilk
Variational inference (VI) is a central tool in modern machine learning, used to approximate an intractable target density by optimising over a tractable family of distributions. As the variational family cannot typically represent the target exactly, guarantees on the quality of the resulting approximation are crucial for understanding which of its properties VI can faithfully capture. Recent work has identified instances in which symmetries of the target and the variational family enable the recovery of certain statistics, even under model misspecification. However, these guarantees are inherently problem-specific and offer little insight into the fundamental mechanism by which symmetry forces statistic recovery. In this paper, we overcome this limitation by developing a general theory of symmetry- induced statistic recovery in variational inference. First, we characterise when variational minimisers inherit the symmetries of the target and establish conditions under which these pin down identifiable statistics. Second, we unify existing results by showing that previously known statistic recovery guarantees in location–scale families arise as special cases of our theory. Third, we apply our framework to distributions on the sphere to obtain novel guarantees for directional statistics in von Mises–Fisher families. Together, these results provide a modular blueprint for deriving new recovery guarantees for VI in a broad range of symmetry settings.
Symplectic Parallel Scan: A Neural Hamiltonian Framework for Accelerated Scientific Simulation
Sungwoo Park
Neural Hamiltonian models provide a principled approach to learning dynamical systems by embedding the symplectic structure of physical evolution into the model class. However, this geometric prior alone has not made Hamiltonian neural networks efficient long-horizon simulators: an M-step trajectory is still typically generated by M sequential updates, and common training objectives based on force matching or one-step prediction provide only local supervision. We introduce Symplectic Parallel Scan (SPS), a neural Hamiltonian framework that algebraizes the learned dynamics through a finite-dimensional Poisson algebra of Hamiltonian generators. SPS constructs a Neural--Poisson Composition DAG closed under projected products and Poisson brackets, whose elements induce symplectic Hamiltonian drivers organized as a Lie-group walk. The associativity of this flow composition enables a parallel prefix-scan over Hamiltonian drivers, reducing the sequential depth of trajectory propagation from $\mathcal O(M)$ to $\mathcal O(\log M)$ while keeping each prefix inside the symplectic flow family. The resulting model is trained through a global trajectory-matching objective rather than only local force labels. Across quantum spin and molecular dynamics benchmarks, SPS achieves substantial wall-clock acceleration over sequential HNN baselines, improves long-horizon prediction accuracy, and reduces energy drift and symplectic violation under extrapolation.
Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data
Xu Guo ⋅ Runyu Peng ⋅ Jian Tong ⋅ Yunhua Zhou ⋅ haijun Lv ⋅ Zhihui Lu ⋅ Qipeng Guo
Large language models (LLMs) rely on web-scale corpora for pre-training. The noise inherent in these datasets tends to obscure meaningful patterns and ultimately degrade model performance. Data curation mitigates but cannot eliminate such noise, so pre-training corpora remain noisy in practice. We therefore study whether a lightweight pre-pre-training (PPT) stage based on synthetic data with learnable temporal structure helps resist noisy data during the pre-training (PT) stage. % Concretely, we draw PPT sequences from many randomly initialized RNNs and keep the PT recipe unchanged. Across various corruption settings, our method consistently improves robustness to noise during PT, with larger relative gains at higher noise levels. % The gains generalize across corruption types and naturally noisy web corpora. For a 1B-parameter model, a synthetic PPT stage with only 65M tokens achieves the same final loss as the baseline while using up to 49% fewer natural-text PT tokens across different noise levels. Mechanistic analyses suggest PPT does not immediately suppress attention to noisy tokens. Rather, PPT-initialized models gradually downweight attention between corrupted tokens during noisy PT. This indicates that synthetic PPT inhibits noise self-modeling and shapes the subsequent optimization trajectory.
SynthHair: Leveraging MetaHumans for a High-Quality 4K Hair Matting Dataset
Markus Karmann ⋅ Shile Li ⋅ Philip Torr ⋅ Puneet Dokania ⋅ Qi Zhang ⋅ Peng-Tao Jiang ⋅ Bo Li ⋅ Onay Urfalioglu
Recent progress in computer vision enables dense alpha matte predictions at resolutions of 4K and higher. While portrait matting has benefited from large-scale, high-resolution datasets, progress in hair matting remains limited due to the difficulty of annotating sub-pixel fine hair details. In this work, we analyze critical limitations in image resolution and annotation quality of such existing datasets and introduce our own synthetic hair matting dataset, \textbf{SynthHair}, which consists of 16k RGBA portrait images in $4096 \times 4096$ image resolution, containing labels for hair matting, as well as hair categorization, growth direction, and depth estimation. We create this dataset using MetaHuman Creator and Unreal Engine to obtain highly detailed and realistic renders of human faces with complex hair structures. We adapt a generalized version of the isoperimetric inequality quotient (IPQ) that we call the \textbf{soft isoperimetric inequality quotient (SIPQ)}, as a scale-invariant label complexity metric that enables comparison of both segmentation and matte labels. Using this metric, we demonstrate that hair matting complexity increases significantly with resolution, reaching an SIPQ of 325.41 at 4K, in contrast with portrait matting datasets at similar resolutions, which reach an SIPQ of ~10. To validate our data, we train separate matting models on SynthHair and existing real-world datasets. Our results demonstrate strong domain generalization and show clear improvements in high-resolution images. A/B testing further confirms these findings, with over 80\% of users preferring our model's results on high-resolution images.
TailCon: Mitigating Tail Signal Erosion through Memory Consolidation for Long-Tailed Recognition
Yuan Dong ⋅ Di Wu ⋅ Zhe Zhao ⋅ Liheng Yu ⋅ Xiaofeng Cao ⋅ Pengkun Wang ⋅ Yang Wang
Long-tailed recognition remains challenging not only because rare classes provide limited supervision, but also because their evidence is difficult to accumulate during standard stochastic training. We characterize this failure as tail signal erosion: in conflict-dominant regimes, frequent-class updates can dilute sparse rare-class signals before they are sufficiently reinforced. To address this issue, we propose TailCon, a memory consolidation framework that separates rare-evidence acquisition from parametric integration. During online acquisition, TailCon stores uncertain or rare representations in detached episodic memory, while a parametric prediction pathway learns global decision structure. To avoid relying only on test-time retrieval, TailCon introduces FOCUS, a periodic consolidation stage that distills stored memory evidence into the parametric prediction head through a low-capacity alignment surrogate. We provide a local optimization analysis that scopes when this separation is beneficial in sparse-recurrence, conflict-dominant regimes. Experiments on CIFAR-100-LT, CIFAR-10-LT, ImageNet-LT, and iNaturalist 2018 show consistent improvements over representative rebalancing, decoupled, expert-based, and prototype-oriented baselines under the evaluated protocols. Memory-disabled inference further shows that part of the stored evidence is transferred into the parametric prediction head rather than remaining only in the retrieval branch. An anonymized implementation is included in the Supplementary Material.
TailFix: Mitigating Error Accumulation and Correcting Tail Deterioration in Long-Horizon Forecasting
Hua Wang ⋅ Haijing Gao ⋅ Fan Zhang
Long-term time-series forecasting (LTSF) underpins applications such as energy-load prediction and traffic-flow forecasting, yet model performance often deteriorates as the horizon extends. In particular, errors accumulate with lead time and systematically amplify toward the end of the predicted sequence, a phenomenon known as tail deterioration. We posit that this instability is driven by weakly constrained fusion of cross-temporal and cross-scale context: aggressive mixing can propagate noise and early-stage bias into distant horizons, ultimately degrading tail accuracy. To address this issue, we propose TailFix, a controllable cross-scale mixing framework for long-horizon forecasting. TailFix selectively gates interactions among multi-scale representations, derives scale-level correction terms from contextual summaries, and injects them via gated residual connections to suppress error diffusion and improve long-horizon stability. We further incorporate adaptive channel enhancement and a lightweight tail-correction module to explicitly refine the error-prone tail region. Finally, we introduce the Tail Amplification Factor (TAF) to quantify tail error amplification relative to the early portion of the sequence. Across diverse benchmarks, TailFix consistently mitigates tail error amplification and improves performance on multiple long-horizon forecasting tasks.
TALES: Text-Adventure Learning Environment Suite
Christopher Z. Cui ⋅ Xingdi Yuan ⋅ Ziang Xiao ⋅ Prithviraj Ammanabrolu ⋅ Marc-Alexandre Côté
While Large Language Models (LLMs) are increasingly deployed as agents to complete complex real world tasks, assessing their baseline agentic capabilities remains challenging. Many existing benchmarks for LLM agents require the model to operate through a structured interface that sit one layer removed from natural language, making it difficult to distinguish failures of interface interpretation from genuine limitations in agentic capability. Because they rely entirely on grounded, natural-language interaction, text-adventure games serve as a prime sandbox for evaluating these capabilities. However, current implementations in popular benchmarks obscure performance through the leakage of privileged information or through scaffolding modifications to the tasks themselves. To solve this, we introduce \ours, a diverse collection of synthetic and human-written text-adventure games in their canonical forms, designed to evaluate baseline agentic competencies in progressively more challenging environments. We present results over a range of LLMs, open- and closed-weights; evaluate when and where current models fail as agents; and perform a qualitative analysis on top agents to identify key failure modes. Despite an impressive showing on synthetic games, even the top LLM-driven agents fail to achieve 25\% on games designed for human enjoyment. Visualization of the experiments can be found at https://github.com/tale-suite/tale-suite-anonymized.
Tango3D: Towards Alignment for Global and Local 2D-3D Correspondence
Zebin He ⋅ Mingxin Yang ⋅ Shuhui Yang ⋅ Hanxiao Sun ⋅ Xintong Han ⋅ Chunchao Guo ⋅ Wenhan Luo
Existing 3D foundation models typically align point clouds to frozen vision-language spaces like CLIP, which achieve strong cross-modal retrieval by compressing 3D shape into a global vector. However, this global-only alignment cannot establish fine-grained pixel-to-point correspondence. To solve this, we present Tango3D, a foundation model that unifies dense correspondence and global retrieval. We use a geometry-aware 2D visual backbone and a pretrained 3D VAE to encode images into 2D patches and point clouds into 3D tokens. These are mapped into a single shared space to achieve both local pixel-to-point alignment and global semantic alignment. To stabilize the joint learning of dense and global objectives, we introduce a three-stage progressive training strategy. Experiments show our model is the first to achieve object-level pixel-to-point alignment while maintaining competitive global retrieval. By establishing a fine-grained alignment feature space, Tango3D injects rich semantics into purely geometric 3D tokens, paving the way for a wide range of dense 3D downstream tasks.
Tapes Together Strong: The Co-evolution of Computation and Cooperation
Kunal Jha ⋅ Francesco Cicala ⋅ Blaise Aguera y Arcas ⋅ Blake Richards ⋅ Natasha Jaques ⋅ Max Kleiman-Weiner ⋅ Eyvind Niklasson
How does cooperation evolve in complex agentic systems? Prior work in evolutionary game theory studies why individuals are incentivized to cooperate by isolating social interactions from the physical costs required to execute them, while artificial life models traditionally study the environments that lead to replication without considering the strategic decisions behind the evolved mechanisms. In contrast, we introduce Autopoietic Game Theory, a computational model where social interactions, replication mechanisms, and their associated computational costs are endogenous and simultaneously co-evolving. We study these dynamics using a computational substrate of randomly initialized programs in Z80 machine code, showing empirically and theoretically that embedding a social dilemma directly into the physics of computation can favor the emergence of self-replicating, cooperative strategies. When resources are scarce, our analysis shows that defection can become self-limiting even in well-mixed populations: parasitic stealing destroys shared energy, slows execution, and can prevent reliable replication. Empirically, evolved programs suppress stealing across several Z80 environments, while spatial assortment further supports structural complexity and task performance. We further show that the framework can incorporate exogenous pressures, such as math tasks structured as sequential social dilemmas, when rewards are tied to computation budgets. These results suggest that coupling an agent's capacity for computation to its available energy transforms cooperation into a dominant scaffolding for building complex, sustainable, and self-organizing systems.
Target-Aware Nuisance Shaping for Object Detection
Tianyi Xu ⋅ Yi Niu ⋅ Xinkun Wang ⋅ Yanbiao Ma ⋅ Mingming Ma ⋅ QingYu Luo ⋅ Fu Zou ⋅ Fu Li
Object detectors are expected to recognize objects independently of intrinsic object-level attributes, such as scale and local visibility. Yet, natural detection datasets contain systematic intrinsic attribute biases, and modern detectors inevitably entangle these biases within their learned representation geometry, leading to non-uniform performance across the continuous attribute spectrum. This observation exposes a central limitation of conventional data augmentation: while attribute-blind transformations enlarge the training distribution, they fail to explicitly correct the dataset-inherent nuisance structure. To address this, we propose Target-Aware Nuisance Shaping} (TANS, a unified framework that re-purposes data augmentation as an object-conditioned vicinal distribution design. TANS processes training samples through a serial pipeline of two fundamental operators: first, Scale Prior Shaping (SPS) geometrically transports object scales toward low-density regions of the natural scale distribution via class-conditional log-area priors; subsequently, Local Visibility Shaping (LVS) photometrically modulates local foreground-background contrast by constructing smooth, box-conditioned Gaussian fields, strictly preserving the newly established geometry and labels. Furthermore, a feature-state controller monitors the evolving representation geometry to adaptively adjust the shaping strength during training. Extensive experiments on MS COCO and Pascal VOC demonstrate that TANS significantly improves detection performance over strong augmentation baselines across modern architectures. Specifically, on COCO with YOLOv11, TANS achieves a 3.6 mAP gain over the Mosaic baseline. Beyond aggregate accuracy, our attribute-regime analyses, feature-state dynamics, schedule comparisons, and logit-space probes consistently indicate that TANS effectively decouples the detector from intrinsic dataset biases, establishing target-aware shaping as a rigorous paradigm for learning attribute-invariant representations.
Teach-to-Reason: Competition-Guided Reasoning with a Self-Improving Teacher
Xiao Han ⋅ Hao Liu ⋅ Zhimin Bao ⋅ Jile Jiao ⋅ Yue Wang ⋅ Hui Guo ⋅ Mou X Feng ⋅ Yi Xu
Chest X-ray visual question answering (CXR VQA) requires models not only to predict correct answers, but also to produce reliable medical reasoning. However, existing reinforcement-learning-based training typically relies on answer-level rewards, which are often too coarse to improve chain-of-thought (CoT) quality and can become ineffective when group-level advantages collapse to zero. We propose \textbf{Teach-to-Reason (T2R)}, a framework that introduces comparison-based supervision into CoT optimization through a self-improving \emph{Teacher} and a competition-guided \emph{Reasoner}. As the Teacher is iteratively strengthened via self-competition, the Reasoner is optimized against progressively stronger Teacher-generated references. We further introduce a case-wise reward design that preserves the original reward-induced positive/negative partition when it is informative, and restores supervision from competition scores when the original reward signal degenerates. Experiments on multiple CXR open-ended VQA benchmarks show that T2R consistently outperforms strong baselines, indicating that comparison-based supervision, when integrated in a controlled and principled manner, provides a more effective training signal for reasoning optimization.
Tell Me What To Learn: Generalizing Neural Memory to be Controllable in Natural Language
Max Bennett ⋅ Thomas Zollo ⋅ Richard Zemel
Modern machine learning models are deployed in diverse, non-stationary environments where they must continually adapt to new tasks and evolving knowledge. Continual fine-tuning and in-context learning are costly and brittle, whereas neural memory methods promise lightweight updates with minimal forgetting. However, existing neural memory models typically assume a single fixed objective and homogeneous information streams, leaving users with no control over what the model remembers or ignores over time. To address this challenge, we propose a generalized neural memory system that performs flexible updates based on learning instructions specified in natural language. Our approach enables adaptive agents to learn selectively from heterogeneous information sources, supporting settings—such as healthcare and customer service—where fixed-objective memory updates are insufficient.
TESLA: Native 4D Gaussian Splatting Generation with Temporally Structured Latents
Weiqi Zhang ⋅ Junsheng Zhou ⋅ Juntong Fang ⋅ Kanle Shi ⋅ Shenkun Xu ⋅ Yu-Shen Liu
Diffusion models have achieved remarkable success in image, video, and 3D content creation. However, extending these successes to native 4D model generation remains challenging due to the intricate coupling between temporal geometry and appearance dynamics. Existing approaches typically generate 4D content by synthesizing videos from novel viewpoints, yet they often suffer from poor spatio-temporal consistency across views and time. To address these limitations, we introduce TESLA, a novel feedforward approach for generating native 4D Gaussian Splatting (4DGS) representations. Our key insight is to decouple temporal geometry and appearance generation through a temporally structured latent representation. This decomposition allows us to first establish coarse temporal geometry that captures fundamental structural dynamics, and then synthesize fine-grained appearance details conditioned on this geometric scaffold. To bridge this decoupled representation to the final output, we design a native 4DGS decoder that directly transforms our temporal features into native 4DGS representations, enabling flexible temporal rendering across arbitrary viewpoints and timesteps. Trained on a large collection of curated animated sequences, we extensively evaluate TESLA on the 4DGS generation task, showing superior performance over existing baselines.
Existing image editing methods can be generally categorized into textual instruction-based and visual prompt-based ones. Textual instructions are semantically expressive, but are limited by the coarse granularity of spatial control of the editing results. In contrast, visual prompts such as drag and point can provide precise spatial guidance, but are limited by the inherent ambiguity in semantic intent. To unify the strength of textual and visual prompts, we present Text-Vision Co-Instructed Image Editing, which jointly models textual instructions as semantic intent and sparse visual instructions as spatial guidance, aiming to achieve precise and intent-faithful image manipulation. To this end, we first construct a textual-visual instruction paired dataset with more than 23K samples derived from dynamic videos, enabling aligned supervision for cross-modal instruction. We then propose TV-Edit, a Textual-Visual instruction unified Editing framework to contextualize drag or point-based visual instructions with image-text semantics and lift them into semantic-aware control representations for pretrained editing backbones. By integrating semantic intent and spatial constraints, TV-Edit leads to more precise spatial control, less instruction ambiguity, and stronger structural consistency than text-only or drag-based alternatives. Finally, we establish TV-Edit-Bench, a deliberately designed benchmark to evaluate semantic faithfulness, spatial alignment, and visual consistency with ground-truth references and controlled textual–visual variations for reliable assessment. Our experiments across multiple editing backbones demonstrate that TV-Edit consistently yields more precise and intent-faithful edits, significantly outperforming state-of-the-art instruction-based and drag-based baselines. Data, model, and codes will be released.
TGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMs
Zihan Lin ⋅ Xiaohan Wang ⋅ Jie Cao ⋅ Jiajun Chai ⋅ Wei Lin ⋅ Guojun Yin ⋅ Ran He
Efficient exploration often remains a central bottleneck in reinforcement learning with verifiable rewards (RLVR). Although temperature control and test-time scaling strategies can increase rollout diversity of large language models (LLMs), they either expand the sample budget at rollout time or leave the benefit of exploration unquantified. To this end, we propose Temperature-Grouped Reinforcement Learning (TGRL), which turns temperature-induced diversity into an explicit training signal. For each prompt, TGRL partitions its rollout group into low- and high-temperature subsets, estimates exploration gain through their reward contrast, and allocates this group-level signal as token-level credit using Jensen--Shannon (JS) divergence between the corresponding temperature-scaled next-token distributions induced by the same logits. Notably, TGRL reaches equivalent accuracy up to 36% faster than strong RLVR baselines without expanding the rollout budget. Across 11 benchmarks from diverse domains, TGRL broadly improves over strong RLVR baselines: it improves the six-benchmark math average by 1.6% at 32B, raises CodeForces rating by 196.7 points and LiveCodeBench Pass@16 by 4.4%, and improves ALFWorld/WebShop success rates by 6.3%/4.9%. Comprehensive ablations and wall-clock analysis confirm the efficacy of all proposed components. Code is available at https://anonymous.4open.science/r/TGRL-4877/.
The Aleatoric-Epistemic Dichotomy of Uncertainty is Meaningful and Indispensable for Machine Learning
Yusuf Sale ⋅ Nikita Kotelevskii ⋅ Maxim Panov ⋅ Eyke Hüllermeier
The quantification of uncertainty, in particular the distinction between aleatoric and epistemic uncertainty, is receiving growing interest in machine learning. However, both its conceptual meaningfulness and practical usefulness have recently been questioned. In this paper, we take the position that the dichotomy is conceptually meaningful and indispensable for (uncertainty-aware) machine learning. In particular, we argue that much of the recent criticism is flawed, either because it targets specific mathematical frameworks for the dichotomy, or because it is based on implicit assumptions that are not implied by the dichotomy itself. Moreover, we highlight that the dichotomy is indispensable for a broad class of decision-making problems, where optimal performance provably requires a distinction between reducible and irreducible uncertainty. Therefore, while existing methodology should certainly be improved further, there is no reason to question the distinction per se.
The Alignment Curse: Modality Alignment Supercharges Audio Attacks via Text Transfer
Yupeng Chen ⋅ Junchi Yu ⋅ Aoxi Liu ⋅ Baoyuan Wu ⋅ Philip Torr ⋅ Adel Bibi
Recent advances in end-to-end trained omni-models have substantially improved audio capabilities by strengthening text-audio modality alignment. However, whether such alignment inadvertently facilitates the transfer of safety vulnerabilities across modalities remains underexplored. This question is critical as text-based jailbreak attacks are considerably more mature than audio-based ones; if they transfer systematically, current audio safety evaluations may underestimate risks originating from the text modality. In this paper, we introduce the Alignment Curse, a formally characterized and empirically validated principle showing that stronger modality alignment enables more effective transfer of attacks from text to audio, revealing a fundamental tension between capability and safety. Motivated by this principle, we conduct a comprehensive black-box evaluation of three attack categories on recent omni-models (e.g., Qwen2.5-Omni, Qwen3-Omni): text attacks, text-transferred audio attacks, and audio attacks. We find that text-transferred audio attacks perform comparably to, and often better than, audio-based attacks, exhibiting a clear advantage under audio-only access. This suggests that text-based vulnerabilities play a pivotal role in shaping audio safety risks. Finally, we empirically analyze the relationship between modality alignment and transfer effectiveness across attack methods and models, observing consistent support for the Alignment Curse: tighter modality alignment leads to more effective cross-modality attack transfer.
The Denoising Wrapper: A Modular Post-Processing Framework for Noisy First-Order Optimizers
Meisam Razaviyayn ⋅ Weiwei Kong ⋅ Grigoris Velegkas ⋅ Vahab Mirrokni
Noisy gradient estimates are ubiquitous in machine learning, arising from stochastic sampling, distributed computations, or privacy- preserving mechanisms like Differential Privacy (DP). This paper introduces a general, optimizer-agnostic framework for denoising these estimates. Our approach operates as a modular wrapper that intercepts noisy gradient observations and provides denoised estimates to the optimizer, requiring no internal modifications to algorithms like SGD or Adam. We start by studying the problem of optimal denoising gradients in quadratic and cubic optimization problems. We develop maximum likelihood estimation of gradient and Hessian of the objective. We also develop other computationally-efficient order optimal algorithms for denoising gradients in such a setting. Finally, we utilize the developments for the cubic and quadratic setting to develop a more general denoising mechanisms for general smooth nonconvex optimization problems. Theoretically, by leveraging higher-order smoothness, we establish an improved convergence rate of $O(T^{-12/19})$ for smooth non-convex optimization. While our recursive algorithm requires two gradient queries per iteration, we show that the improved convergence rate yields a lower total oracle complexity than the standard $O(T^{-1/2})$ rate of SGD. This implies the gain in convergence speed asymptotically outweighs the additional computational cost. We apply our algorithm to various denoising problems, particularly nonconvex DP-training. Our experiments, including training private classifiers on CIFAR-10, demonstrate significant improvements over baselines.
The Era of Agentic Organization: Learning to Organize with Language Models
Zewen Chi ⋅ Li Dong ⋅ Qingxiu Dong ⋅ Yaru Hao ⋅ Xun Wu ⋅ Shaohan Huang ⋅ Furu Wei
We envision a new era of AI, termed agentic organization, where agents solve complex problems by working collaboratively and concurrently, enabling outcomes beyond individual intelligence. To realize this vision, we introduce asynchronous thinking (AsyncThink) as a new paradigm of reasoning with large language models, which organizes the internal thinking process into concurrently executable structures. Specifically, we propose a thinking protocol where an organizer dynamically assigns sub-queries to workers, merges intermediate knowledge, and produces coherent solutions. More importantly, the thinking structure in this protocol can be further optimized through reinforcement learning. Experiments demonstrate that AsyncThink achieves 28% lower inference latency compared to parallel thinking while improving accuracy on mathematical reasoning. Moreover, AsyncThink generalizes its learned asynchronous thinking capabilities, effectively tackling unseen tasks without additional training.
The Fault in Our Metrics: Revisiting Generative Model Evaluation with Chamfer Distance
Chirag Vashist ⋅ Vaibhav Santhosh ⋅ Tristan Engst ⋅ Yanshu Zhang ⋅ Ke Li
Reliable evaluation is essential for progress in image generation, yet widely used metrics rely on assumptions that are misaligned with the geometry of real data. We revisit two dominant families of metrics: Fréchet-based metrics such as FID, and precision/recall-style metrics based on KNN balls. We show that FID reduces complex feature distributions to a Gaussian moment summary, making it blind to differences beyond mean and covariance. We further show that KNN-ball metrics construct inflated isotropic support estimates in high-dimensional feature spaces, causing binary membership queries to accept off-manifold samples and making recall unstable under standard finite-sample evaluations. We propose Chamfer distance as a simple alternative: instead of constructing inflated support estimates, it directly measures cross-set nearest-neighbor distances between real and generated samples. The resulting one-sided distances provide interpretable signals for fidelity and coverage, while reducing sensitivity to the neighborhood hyperparameter and remaining computationally efficient.
The Fractured Interlingua: Geometric Bottlenecks in Cross-Lingual Knowledge Editing
Aastha A K Verma ⋅ Anwoy Chatterjee ⋅ Sparsh Jain ⋅ Tanmoy Chakraborty
Transformer-based autoregressive multilingual language models (LMs) are often assumed to store factual knowledge in an interlingua: a language-agnostic latent space. If this assumption holds, a parameter update applied in a source language during knowledge editing should inherently propagate to its semantic equivalents in target languages. However, empirical benchmarks reveal that cross-lingual propagation of edits consistently fails, a degradation that prior work treats largely as an algorithmic symptom rather than a structural bottleneck. In this work, we hypothesize that this failure stems from language anisotropy, and we perform a rigorous geometric dissection of the cross-lingual propagation gap. Using associative memory-based "locate-then-edit" methods, including ROME, MEMIT, and AlphaEdit, as analytical frameworks, we derive a closed-form decomposition of this gap into key-space propagation, parallel residual mismatch, and orthogonal residual mismatch. We empirically find that cross-lingual propagation is consistently limited by a large orthogonal residual component, indicating that source-language edits often fail to span the target-language update directions required for successful transfer. Furthermore, we show that post-hoc correction using only source-language residual statistics faces a large residual-prediction barrier. Finally, we demonstrate that structural realignment via a targeted contrastive alignment objective substantially improves key-space propagation and partially improves cross-lingual edit transfer, while leaving value-space mismatch as a remaining bottleneck.
The Illusion of Diversity: Aligning LLM Exploration via Effective Entropy
Xiaoliang Fu ⋅ Jiaye Lin ⋅ Yangyi Fang ⋅ Cong Qin ⋅ Binbin Zheng ⋅ Chaowen Hu ⋅ Zekai Shao
Reinforcement learning with verifiable rewards has shifted the alignment paradigm of large language models from subjective preference to objective correctness, yet a fundamental disconnect persists between exploration diversity and reasoning validity. Existing strategies predominantly rely on ``blind'' entropy metrics, e.g., token-level or global sequence-level entropy, which structurally fail to distinguish between effective reasoning and potential hallucinations. This misalignment creates an illusion of diversity, where high entropy tends to signal chaotic degeneration rather than valuable exploration. To bridge this gap, we introduce \textit{Effective Entropy}, a validity-aware entropy concept that quantifies diversity exclusively within the valid subspace and systematically filters out the noise of incorrect trajectories. Built on this concept, we instantiate VEPO, a simple \textbf{V}alidity-aware \textbf{E}ntropy \textbf{P}olicy \textbf{O}ptimization method that directly incorporates effective entropy into RLVR training. VEPO induces an adaptive attraction dynamic that drives up the probability of under-explored valid trajectories, a mechanism that matches the practical regime of large-scale autoregressive models where sequence probabilities remain extremely small. Extensive experiments on reasoning benchmarks demonstrate that this effective-entropy-driven optimization significantly outperforms GRPO and other entropy-aware variants, showcasing its ability to more reliably cultivate, across diverse tasks, a broad portfolio of valid reasoning trajectories.
The Kernel Reality Check: Benchmarking and Distilling Efficient Attention at Scale
Firat Oncel ⋅ Cem Subakan ⋅ Mirco Ravanelli ⋅ Çağatay Yıldız
Kernelized attention methods, which aim to replace the $O(N^2)$ complexity of softmax attention with $O(N)$ alternatives, have been the focus of intense research in the last five years. Despite growing interest, systematic evaluations at foundation model scales and fair comparisons across approaches are still missing. We analyze kernelized attention in two halves and highlight prominent limitations on each. On efficiency, FlashAttention achieves lower memory consumption and faster wall-clock time than state-of-the-art kernelized methods across most sequence lengths (with auto-regressive generation memory the only clear exception), directly contradicting the field's core assumption. We also discover that reported efficiency gains often fail to reproduce against modern baselines. On quality, current kernelized methods leave gaps of $2$-$10$ points relative to softmax under a unified distillation protocol, and these gaps do not close with model scale. We show this gap is not intrinsic to kernelization: LARA++, a principled extension of LARA that approximates softmax via importance sampling with adaptive proposal distributions, recovers up to 99\% of softmax accuracy. We establish that the gap is closable; however, the computational overhead remains substantial, leaving the efficiency wall intact.
The Key to Going Linear: Analysis-Driven Transformer Linearization
Anna Kuzina ⋅ Paul Whatmough ⋅ Babak Ehteshami Bejnordi
The quadratic cost of causal self-attention severely bottlenecks long-context transformer inference. While numerous post hoc linearization pipelines exist, it is difficult to identify which components preserve model quality. This work isolates the effect of state update design in a strict frozen-backbone regime. We show that softmax relies on key-dependent, rank-1 orthogonal projections, elucidating why delta-style networks outperform purely gated accumulation. We identify a potential source of approximation errors and introduce structural interventions, specifically sink tokens, short convolutions, and fixed-budget cache routing, which reduces the remaining gap. We scale this linearization approach across LLaMA and Qwen models up to 32B parameters, outperforming prior post hoc baselines on MMLU and matching the long-context retrieval of complex adaptive-caching frameworks.
The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
Ryoya Awano ⋅ Taiji Suzuki
Weak-to-strong (W2S) generalization, in which a strong model is fine-tuned on outputs of a weaker, task-specialized model, has been proposed as an approach to aligning superhuman AI systems. Existing theoretical analyses either fix the student's representations or operate in restricted settings. Whether multi-step SGD can succeed in feature learning while preserving diverse pre-trained capabilities remains open. We study W2S in the setting of reward-model learning with two-layer neural networks. The strong model has pre-trained representations organized into low-dimensional subspaces $V_k$, and is fine-tuned under the supervision of a weak model specialized on task $\kappa$. We prove that the strong model efficiently learns task $\kappa$, eliciting its pre-trained knowledge while retaining general capabilities. This establishes W2S generalization in the feature-learning regime, in the sense that the strong model acquires the target feature direction through W2S training, rather than having it given a priori. Moreover, W2S preserves pre-trained off-target features, whereas standard supervised fine-tuning causes catastrophic forgetting when off-target feature directions are correlated with the target's. Numerical experiments on synthetic data confirm our theoretical results.
The Platonic Defense: Backdoor Defense for Self-Supervised Encoders in the Era of Large Scale Pre-training
Tuo Chen ⋅ Minjing Dong ⋅ Benlei Cui ⋅ Jian liu ⋅ Jie Gui
Self-supervised learning (SSL) pretrained models have become a dominant paradigm for visual representation learning, but they are vulnerable to backdoor attacks. Existing defenses struggle to defend against such attacks in a fully black-box setting because they often require access to labels, attack patterns, or training data. To tackle this issue, we propose a new attack-agnostic, model-agnostic, and modality-agnostic black-box test-time defense paradigm, called \emph{Platonic Representation Defense}. It is inspired by the Platonic Representation Hypothesis, which suggests that large-scale independently trained encoders converge toward compatible projections of the same underlying reality. We formalize this idea as a conditional energy function defined over source representations and a set of reference representations. The energy function is trained for detection through noise-contrastive estimation and for representation purification through denoising score matching. Theoretically, the energy gap between matched and mismatched samples is lower bounded by the mutual information between source and reference representations. We demonstrate the effectiveness of our method on multiple self-supervised encoders and more than 10 attacks. The method can perform both representation detection and purification, and achieves substantial performance gains across multiple attacks. Code can be found in the appendix.
The Reflexivity Threshold: A Phase Transition for Multi-Agent Performative Prediction
Ahmed Mahrous ⋅ Roberto Di Pietro
Many learning systems are performative: once deployed, their predictions change the data distribution on which future models are trained. In multi-agent settings, this feedback is routed through a network of interacting learners, so stability depends not only on the strength of performativity but also on who affects whom. We study repeated retraining in multi-agent performative prediction and identify a spectral threshold, derived from primitive sensitivity and curvature parameters, that separates stable learning guarantees from worst-case learning obstruction. Prior work often collapses cross-agent feedback into a scalar contraction constant, missing this network structure. We instead introduce the reflexivity matrix $\Gamma$, which records how strongly one agent's deployment can move another agent's best response, defined from per-agent gradient sensitivities and curvature parameters. The threshold occurs when the spectral radius of this matrix crosses one. Below the threshold, repeated retraining converges geometrically to a unique performatively stable equilibrium, with cumulative excess loss bounded uniformly in the time horizon. Above the threshold, we construct supercritical instances in which decentralized learning provably fails: every local-information learner incurs cumulative excess loss that grows linearly in the horizon. Together, these results characterize this threshold behavior through upper and lower bounds and recover the single-agent scalar threshold from prior work as a special case.
The Scaling Laws of Skills in LLM Agent Systems
Qiguang Chen ⋅ Qiming Yu ⋅ Yuhang Gu ⋅ Zhuoye Huang ⋅ Hanjing Li ⋅ Hongyu Liu ⋅ Simin Liu ⋅ Jinhao Liu ⋅ Dengyun Peng ⋅ Jiangyi Wang ⋅ Zheng Yan ⋅ Fanqing Meng ⋅ Ethan Qin ⋅ Carl Che ⋅ Mengkang Hu
As agent systems scale, skills accumulate into large reusable libraries, yet their scaling laws remain poorly understood. Across 15 frontier LLMs, 1,141 real-world skills, and over 3M routing or execution decisions, we identify two coupled laws. Routing law: single-step routing accuracy decays logarithmically with library size ($R^2{>}0.97$ for all models), with errors progressing from local skill competition to cross-family drift and capture by overly general ``black-hole skills''. Execution law: before state realization, joint routing is approximately multiplicative, whereas correct execution can improve difficult downstream decisions by about $4{\times}$. A single parameter, the routing logarithmic decay slope $b$, couples the two laws: routing-side fits predict execution-side rescue across models, showing that the same library property controls both pre-execution collapse and downstream recoverability. The laws are actionable: law-guided optimization raises held-out routing accuracy from 71.3% to 91.7%, reduces hijack from 22.4% to 4.1%, and transfers directionally to downstream ClawBench and ClawMark execution settings, improving mean pass rate from 49.3% to 61.6% on ClawBench and from 28.4% to 34.5% on ClawMark. These results show that agent performance depends not only on model capability, but also on the structure, granularity, and exposure policy of the skill library.
TIGER-FG: Text-Guided Implicit Fine-Grained Grounding for E-commerce Retrieval
Xinyu Sun ⋅ Huangyu Dai ⋅ Lingtao Mao ⋅ Zexin Zheng ⋅ Zihan Liang ⋅ Ben Chen ⋅ Chenyi Lei ⋅ Wenwu Ou
E-commerce image search often takes a cropped image as the query, while each candidate is represented by full item images and structured text. This image-to-multimodal retrieval setting presents two asymmetries: a modality disparity -- a visual query must match image--text items, and a granularity disparity -- a cropped query must be compared with full images containing background context and possible distractors. Detection-based pipelines handle the granularity disparity through explicit localization but incur extra cost and error propagation, whereas CLIP-style encoders avoid detection, but are vulnerable to backgrounds or irrelevant items. To address these limitations, we propose TIGER-FG, a text-guided implicit fine-grained grounding framework for image-to-multimodal e-commerce retrieval. TIGER-FG uses item text as semantic guidance to produce target-focused item representations without object detection for retrieval. We further introduce dual distillation objectives that preserve target-region spatial consistency and query--item similarity structure, yielding more stable and discriminative multimodal representations. In addition, we construct ECom-RF-IMMR, a realistic benchmark suite with a 10M-pair training set and two evaluation benchmarks covering standard and cluttered item layouts. TIGER-FG improves Recall@1 over the strongest baseline by 6.1 and 34.4 percentage points on the two evaluation benchmarks, respectively, with only 85.7M query-side parameters and 256-dim embeddings. Results on public e-commerce benchmarks further demonstrate its generalization to noisy and one-to-many retrieval scenarios. Code and data will be released.
Time-Sensitive Anytime-Valid Testing
Eugenio Clerico ⋅ Tobias Wegel ⋅ Iskander Azangulov ⋅ Patrick Rebeschini
Anytime-valid tests allow evidence to be checked during data collection: one can either continue testing or stop and reject the null while still controlling type-I error. Yet, in many applications rejection is useful only if it comes soon enough. We introduce a time-sensitive testing-by-betting framework that favours early rejection by assigning rewards to rejection times and maximising their expected value under a given alternative. This encompasses hard deadlines and softer time preferences. The resulting optimal control problem admits a Bellman representation in terms only of time and evidence against the null, rather than the full history. For hard deadlines, we identify a canonical optimal test in the simple-vs-simple case. We show that exponentially decaying rewards admit a stationary approximation that yields a finite-time-scale counterpart to the classical growth-rate-optimal (GRO) viewpoint in anytime-valid statistics, recovering it in the large-time-scale limit.
Time to Pay Attention! Understanding High Complexity Corpus Reasoning Tasks
Prasann Singhal ⋅ Amanda Bertsch ⋅ Jacob Steinhardt ⋅ Sewon Min
Corpus reasoning tasks, which require extracting and integrating information from large collections of documents (e.g. scientific literature, LLM agent traces), range from well-studied tasks such as factoid question answering (e.g. When was Ralph Lauren founded?), to more complex, underexplored queries like What are all the conflicting claims in this literature?. We observe that on these complex tasks, efficient architectures can scale much worse compared than they do on more well-explored tasks, but why does this happen? We develop a unified view---corpus task complexity (CTC)---to systematically distinguish complex tasks from simpler ones in terms of the smaller operations needed to solve them (and how such operations scale with corpus size): for example, we classify tasks which compare all possible document pairs as more complex than factoid retrieval. We first investigate new long-context language model (LCLM) training dynamics challenges on these tasks, and propose an effective new mask mixing technique that improves O(N) attention approaches on high CTC tasks, but at larger corpus scale we ultimately find that there is no free lunch: (i) methods that scale efficiently with corpus size do poorly when CTC is high, and (ii) methods that can handle these tasks scale quadratically in cost with corpus size. We identify high-CTC tasks as a long-term open problem and encourage more future work investigating new methods to overcome these limitations.
TMPO: Trajectory Matching Policy Optimization for Diverse and Efficient Diffusion Alignment
Jiaming Li ⋅ Chenyu Zhu ⋅ Zhiyuan Ma ⋅ Nanxi Yi ⋅ Youjun Bao ⋅ Li Sun ⋅ Quanying Lv ⋅ Xiang Fang ⋅ Daizong Liu ⋅ Jianjun Li ⋅ Kun He ⋅ Bowen Zhou
Reinforcement learning (RL) has shown extraordinary potential in aligning diffusion models to downstream tasks, yet most of them still suffer from significant reward hacking, which degrades generative diversity and quality by inducing visual mode collapse and amplifying unreliable rewards. We identify the root cause as the mode-seeking nature of these methods, which maximize expected reward without effectively constraining probability distribution over acceptable trajectories, causing concentration on a few high-reward paths. In contrast, we propose Trajectory Matching Policy Optimization (TMPO), which replaces scalar reward maximization with trajectory-level reward distribution matching. Specifically, TMPO introduces a Softmax Trajectory Balance (Softmax-TB) objective to match the policy probabilities of $K$ trajectories to a reward-induced Boltzmann distribution. We prove that this objective inherits the mode-covering property of forward KL divergence, preserving coverage over all acceptable trajectories while optimizing reward. To further reduce multi-trajectory training time on large-scale flow-matching models, TMPO incorporates Dynamic Stochastic Tree Sampling, where trajectories share denoising prefixes and branch at dynamically scheduled steps, reducing redundant computation while improving training effectiveness. Extensive results across diverse alignment tasks such as human preference, compositional generation and text rendering show that TMPO improves generative diversity over state-of-the-art methods by 9.1%, and achieves competitive performance in all downstream and efficiency metrics, attaining the optimal trade-off between reward and diversity. Code is available at https://anonymous.4open.science/r/TMPO-2E2B.
Tokens-per-Parameter Coverage Is Critical for Robust LLM Scaling Law Extrapolation
Joshua Shay Kricheli ⋅ Alexander L. Reid ⋅ Soumajyoti Sarkar ⋅ Venkata Gandikota ⋅ Paulo Shakarian
Neural scaling laws approximate a language model's loss as a power-law function of parameter count $N$ and token count $D$. Following Chinchilla-style compute-optimal training, many studies fit scaling laws from runs performed under a fixed tokens-per-parameter (TPP) ratio $k$ and set $D = kN$. We show that this collinear design, combined with the empirically common near-equality of the exponents governing $N$ and $D$, induces an inherent ill-conditioning in the Gauss-Newton least-squares problem: the condition number of the design grows as the inverse square of the gap between the $N$ and $D$-exponents. The scale coefficients become practically unidentifiable, with confidence intervals inflating by an order of magnitude or more, yielding a ``sloppy'' model whose extrapolations degrade sharply off the training ray. We prove this for four scaling-law formalisms and derive a closed-form TPP-diversity threshold that is necessary and sufficient for well-conditioned estimation. Empirically, non-collinear designs outperform collinear ones on held-out splits with a 97.3\% win rate across four laws, five corpora, multiple floating point precision modes. We further show the degeneracy is rooted in Jacobian geometry and is not an artifact of the loss function: any smooth estimation objective whose curvature involves the Jacobian inherits the same ill-conditioning.
Tools as Continuous Flow for Evolving Agentic Reasoning
Tairan Huang ⋅ Siyu Shang ⋅ Qiang Chen ⋅ Xiu Su ⋅ Yi Chen
Large Language Models (LLMs) have demonstrated remarkable capabilities in orchestrating tools for reasoning tasks. However, existing methods rely on a step-wise paradigm that lacks a global perspective, which causes error accumulation over long horizons and restricts generalization to unseen tools. To overcome these limitations, we propose Tools as Continuous Flow for Evolving Agentic Reasoning (FlowAgent), which reconceptualizes tool chaining as continuous trajectory generation within a semantic space. To systematically evaluate this paradigm, we introduce the first plan-level closed-loop benchmark dedicated to plan-level agentic reasoning in dynamic real-world environments. Specifically, the proposed FlowAgent leverages conditional flow matching to generate continuous latent trajectories, providing a global planning perspective to ensure coherent and robust tool execution. Theoretically, we establish formal bounds on utility convergence and prove that our continuous formulation fundamentally guarantees robust generalization and error attenuation. Empirical evaluations show that FlowAgent achieves superior robustness and adaptability in long-horizon reasoning tasks.
Tool Verification for Test-Time Reinforcement Learning
Ruotong Liao ⋅ Nikolai Röhrich ⋅ Xiaohan Wang ⋅ Yuhui Zhang ⋅ Yasaman Samadzadeh ⋅ Volker Tresp ⋅ Serena Yeung-Levy
Test-time reinforcement learning (TTRL) has emerged as a promising paradigm for adapting large reasoning models (LRMs) on unlabeled test inputs, using self-consensus rewards derived from majority voting over sampled rollouts. However, majority voting can mistake popularity for correctness: a spurious yet high-frequency unverified consensus may become a biased reward signal, causing test-time RL to reinforce frequent but wrong answers and collapse into an incorrect mode. We address this false-popular failure mode with T^3RL (Tool-Verification for Test-Time Reinforcement Learning), a verification-aware test-time RL framework. T^3RL grounds pseudo-label construction in external tool evidence. Concretely, a verifier uses external tool evidence, such as code execution, to upweight verified rollouts during verification-aware voting, producing more reliable pseudo-labels for training. Across mathematical reasoning benchmarks of varying difficulty, including MATH-500, AMC, and AIME 2024, and diverse backbone families, T^3RL significantly improves over TTRL, with stronger performance on harder problems. More broadly, T^3RL is positioned as a verified online data synthesizer, highlighting the role of test-time tool verification in reliable online adaptation.
TopoCurve: Geometry-Aware Topology Reasoning via Bézier Curves in Autonomous Driving
Mihai Bogdan Deaconu ⋅ Laura Diosan
Topology reasoning jointly detects 3D lanes and traffic elements from multi-view images and infers their structural connectivity. Current methods model lanes as discrete polylines, lacking smoothness, analytical tangent directions, and global spatial support for attention, while providing sparse topology supervision. We propose TopoCurve, a geometry-driven architecture for 3D topology reasoning grounded in a structured parametric lane representation. Lanes are modeled as endpoint-fixed cubic Bézier curves, enabling continuous geometry with exact endpoints and analytically defined directionality. We exploit this shared curve geometry across the entire pipeline. Endpoint distance and tangent alignment are encoded with multi-scale Fourier features and injected into the topology head. Sampled curve points serve as geometry-aligned references for deformable cross-attention spanning the full lane. Parallel curve-anchored attention branches provide diverse predictions for one-to-many topology supervision. These components form a tightly coupled cascade where representation enables geometric reasoning, guides feature aggregation, and supports denser supervision. TopoCurve achieves 50.6 OLS on the OpenLane-V2 benchmark without any post-processing, establishing a new state-of-the-art among end-to-end camera-only methods, and outperforms all existing approaches on endpoint detection (56.8 vs. 52.6 on DET$_p$).
TopoScope: A Graph-Scoped LLM Agent for Printed Circuit Board Schematic Design
Jiale Guo ⋅ Shiji Tong ⋅ Weihao Li ⋅ Yu Wu ⋅ Boxian Lin ⋅ Mengji Shi
Printed circuit boards (PCBs) are the hardware foundation of modern electronic systems. However, automating PCB schematic design remains difficult. A valid schematic must respect real-IC pin-level constraints and cross-component topological relations, neither of which is captured by generic simulation, local rule checks, or direct LLM generation. To address this challenge, we present TopoScope, a training-free LLM agent that uses the schematic graph not only as the final output, but also to guide generation and verification at runtime. TopoScope couples two mechanisms: Topology-Guided Context Scoping, which computes a scoped view for each component-level wiring decision, and Pattern-Based Rule Verification, which expresses datasheet-derived electrical constraints as graph patterns and uses them as in-loop feedback. We also introduce a 20-design benchmark spanning three complexity tiers and six application domains, with functional assertions evaluated on the generated schematic graph and calibrated against expert judgments. Across seven recent LLMs, TopoScope improves functional correctness, reduces invalid graph references, and enables more reliable generation of medium and hard multi-IC schematics compared with knowledge-provided baselines.
Total Variation Distance Estimation in Autoregressive Models
Eric Price ⋅ Kevin Tian ⋅ Zhiyang Xun ⋅ Yusong Zhu
Modern LLM deployments use a number of implementation choices and inference optimizations (e.g., batching, custom kernels, quantization) on top of fixed weights, so two engines serving "the same model" can produce meaningfully different distributions. We study the problem of estimating the total variation (TV) distance between two length-$n$ autoregressive distributions to additive error $\varepsilon$, under three access models: 1. Under *sample* access, we use $\widetilde{O}(\frac{n^2K}{\varepsilon^2})$ queries, where $K$ is the maximum support of the next-token distribution. 2. Under *logit* access, we use $O(\frac{n}{\varepsilon^2})$ queries, and this is tight. 3. Under *noisy logit* access, we can smoothly interpolate between the above two: if probability values are given to $\sigma$ relative error, we use $\widetilde{O}(\frac{n + n^2\sigma^2}{\varepsilon^2})$ queries. We complement our theoretical results with an empirical evaluation of our algorithms, for example measuring the distance between \texttt{sglang} and \texttt{vllm} on standard settings. Our experiments highlight the robustness and practicality of estimating the total variation distance, compared to existing alternatives such as KL divergence.
Towards Characterizing Scientific Image Utility and Upgradability
Wenzhe Li ⋅ Qihang Yan ⋅ Liang Chen ⋅ Junying Wang ⋅ Farong Wen ⋅ Yijin Guo ⋅ Chunyi Li ⋅ Zicheng Zhang ⋅ Guangtao Zhai
Scientific images function as critical evidence in research communication, yet their integrity faces unprecedented threats from AI-generated content that introduces subtle but consequential errors. Existing evaluation paradigms prove inadequate: image quality assessment poorly correlate with scientific validity, while language models lack domain-specific verification capabilities. To address this gap, we propose the $\textbf{S}$cientific $\textbf{I}$mage $\textbf{U}$tility and $\textbf{U}$pgradability $\textbf{A}$ssessment ($\textbf{SIU$^2$A}$) framework, which introduces two complementary dimensions for scientific image evaluation. $\textbf{Utility}$ encompasses $\textit{error detection}$ (identifying scientific inaccuracies) and $\textit{correction feasibility}$ (assessing whether errors can be reliably repaired). $\textbf{Upgradability}$ measures the quality of correction We categorize scientific image corruption into four fundamental types: Detail Distortion, Incompleteness, False Content, and Entity Confusion. Based on this taxonomy, we construct SIU$^2$A-Benchmark, a comprehensive dataset featuring expert annotations for both error identification and repair. The framework implements a unified two-stage evaluation protocol: the first stage evaluates error detection and correction instruction generation (Utility), while the second stage assesses the effectiveness of actual corrections (Upgradability). Experiments reveal that current multimodal systems exhibit significant limitations in both scientific error assessment and faithful correction, exposing a fundamental gap between visual perception and scientific usability. Our findings establish that diagnostic quality fundamentally constrains restoration performance, highlighting the critical need for robust error perception mechanisms in scientific multimodal AI.
Toward Semantically-Consistent Tuning-Free Customization for Rectified Flow Transformers
Jian Jin ⋅ Kai Zhang ⋅ Zhenyong Fu ⋅ Jian Yang
Tuning-free customized generation has gained popularity for its efficiency, particularly with the rise of rectified flow transformers. However, it often suffers from semantic inconsistency, degrading both concept fidelity and prompt alignment. We attribute this failure to their tendency for inference-time superficial feature fusion, which is intrinsic to the tuning-free paradigm. In this paper, we propose SemanticFlow for semantically consistent tuning-free customization, introducing explicit semantic control to jointly enhance concept fidelity and prompt alignment. SemanticFlow operates through a cohesive "perceive-and-manipulate" mechanism for aligned fusion of visual and textual features at the early denoising stage. Specifically, we first devise lightweight semantic token streams as run-time probes to accurately perceive the spatial responses of the target semantics. We then formulate semantic manipulation as an energy-guided rectification of these semantic responses within the rectified flow framework, steering the generation process toward the desired outcome. Experiments demonstrate that SemanticFlow jointly enhances concept fidelity and prompt alignment in previously challenging scenarios, effectively enabling semantically-consistent customization.
Towards Generalist Graph-Level Anomaly Detection
Junjun Pan ⋅ Yixin Liu ⋅ Yu Zheng ⋅ Fuyi Li ⋅ Alan Wee-Chung Liew ⋅ Shirui Pan
Graph-level anomaly detection (GLAD) aims to identify graphs that deviate significantly from the majority and plays a crucial role in applications such as molecular analysis and fraud detection. While recent advances have achieved competitive performance, existing approaches mainly follow a dataset-specific paradigm, requiring substantial in-domain data and training cost when applied to new application scenarios. This limits their scalability and practicality in real-world applications where labeled data is scarce and diverse domains are continuously emerging. In this paper, we investigate the problem of generalist GLAD, which seeks to learn a single unified model capable of detecting anomalous graphs across multiple domains with minimal supervision on the target data. This problem introduces key challenges, including learning transferable patterns from heterogeneous graph datasets and effective adaptation under limited supervision. To address these challenges, we propose GenGLAD, a generalist GLAD approach that leverages discriminative subgraph extraction as the core mechanism to learn transferable anomaly knowledge and enable efficient adaptation to new domains. Specifically, GenGLAD employs a graph neural network-based discriminative subgraph extraction model to learn the anomaly-discriminative substructures that distinguish anomalous and normal graphs across diverse domains. During the test-time adaptation phase, multi-dimensional discrepancies between the extracted subgraph and multi-level contextual information are calculated to characterize anomalies from multiple perspectives. For efficient adaptation, a lightweight scoring module is introduced to automatically decide the most informative signals, improving robustness against misleading patterns. Extensive experiments on diverse benchmark datasets demonstrate that the proposed approach achieves strong performance, superior generalization ability, and improved data efficiency compared to existing methods.
Towards Generalizable 3D Anomaly Detection via Relational Inconsistency Modeling
KunHo Heo ⋅ SuYeon Kim ⋅ Hayoung Lee ⋅ Chanse Oh ⋅ MyeongAh Cho
3D anomaly detection (3DAD) aims to identify defective regions in point cloud data, serving as a critical component in industrial inspection systems. Existing methods are normality-centered — learning the distribution of normal samples and treating deviations as anomalies — without explicitly modeling what constitutes a defect. This leads to ambiguous decision boundaries with increased false positives and negatives, particularly in unified and cross-domain settings where diverse normal distributions further blur the boundaries. We propose a relational inconsistency modeling framework that characterizes defects as violations of geometric consistency among neighboring structures. Our approach learns category-agnostic defect cues through pseudo-anomalies designed as controlled relational violations, instantiated by two key modules: Edge-aware Graph Refinement (EGR) for encoding geometric relationships among local regions, and Cluster-Deviation Modeling (CDM) for identifying regions that are relationally incompatible within their structural peer group. Extensive experiments on Anomaly-ShapeNet and Real3D-AD demonstrate consistent improvements over prior state-of-the-art methods in both in-domain and cross-domain settings, validating the effectiveness of learning an explicit, relation-based defect criterion for 3D anomaly detection.
Towards Generalizable Partially Relevant Video Retrieval
Woojin Jun ⋅ Cheol-Ho Cho ⋅ MinSeok Jung ⋅ SuBeen Lee ⋅ Jae-Pil Heo
Partially Relevant Video Retrieval (PRVR) aims to retrieve an untrimmed video that contains at least one segment relevant to a text query. Despite its practical motivation, existing PRVR methods are typically trained and evaluated on a single curated benchmark, limiting insight into whether they learn generalizable fine-grained text-video alignment. To examine this issue, we introduce a multi-source-to-unseen-target protocol, in which multiple source benchmarks are merged into a single training corpus without source labels and the trained model is evaluated on a held-out target benchmark. Under this protocol, existing PRVR methods generalize poorly to unseen targets despite leveraging the general image-text alignment of CLIP. Surprisingly, they even underperform zero-shot CLIP, suggesting that PRVR training can actively erode the transferable alignment inherited from CLIP. To address this limitation, we propose Generalization-Aware Learning (GAL) for PRVR, a two-branch framework that preserves CLIP-derived transferability while learning PRVR-specific fine-grained text-segment matching. GAL consists of an anchored branch that maintains CLIP's transferability and an adaptive branch that learns task-oriented segment-level alignment. We train these branches through conservative residual updates, per-sample branch specialization, and CLIP-text relation distillation, enabling complementary learning without overwriting the alignment structure provided by CLIP. Under the proposed protocol, GAL achieves state-of-the-art unseen-target retrieval with competitive source performance, showing that fine-grained matching can be learned without collapsing transferable alignment.
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning
Yinghui He ⋅ Ling Yang ⋅ Jiarui Liu ⋅ Yongjin Yang ⋅ Lechen Zhang ⋅ Mengdi Wang ⋅ Sanjeev Arora
Long-horizon reasoning in recent LLMs increasingly demands that the model switch between distinct skills inside a reasoning chain, such as a task that requires the model to first do a math derivation, then use the result to plan a schedule. We call such problems cross-skill long-horizon tasks: multi-step tasks whose steps require different reasoning skills and depend on earlier outputs. Existing benchmarks often evaluate individual skills, lacking a principled way to measure how well a model switches between skills. We address this gap from both the evaluation and training sides. We introduce Skill Entropy, a measure of the difficulty of switching from one skill to another. We then propose Skill²-Bench, a benchmark of cross-skill long-horizon tasks built over 558 skills across 9 verifiable and open-ended domains. Each task is assigned a task-level skill-entropy score and grouped into three difficulty levels. Evaluating 8 frontier and 4 open-source models on Skill²-Bench surfaces a structural skill-switching gap: accuracy decreases monotonically with task-level skill entropy and drops by 5–30% when the same skill is exercised inside a cross-skill task rather than a single-skill question. We then turn skill entropy from a benchmark scale into a training signal. We propose Skill-Entropy RL, an RL framework where the model predicts not only the answer at each step but also the skill used to produce it. The reward combines step-level correctness with a skill-entropy reward that measures the alignment between the model-predicted skill sequence and the gold skill sequence. On Qwen3-4B-Instruct and Qwen3-1.7B, Skill-Entropy RL improves the Skill²-Bench score from 34.4% to 68.4% and from 14.6% to 40.1% respectively, outperforming SFT and vanilla GRPO. The same pipeline can be applied to off-the-shelf training data such as OpenR1-Math, indicating that skill entropy is a reusable training signal.
Towards Principled Efficient Rollout Allocation in Test-Time Training
Youkang Wang ⋅ Jian Wang ⋅ Tianyi Zeng ⋅ Xiao-Yong Wei ⋅ Qing Li
Test-time training and test-time policy optimization have emerged as effective tools for improving the robustness of LLMs under distribution shift. A common approach constructs label-free reward signals from multiple sampled rollouts using majority voting, but existing methods typically rely on a fixed, large rollout budget for every query. This fixed allocation is inefficient because many queries reach consensus well before the rollout budget is exhausted. In this work, we introduce \textbf{AdaPO}, which replaces fixed-budget majority voting with an adaptive two-stage voting procedure that moves from electing to confirming when a leader is statistically dominant, and stops once the posterior error of accepting that leader is sufficiently low. AdaPO is plug-and-play with standard policy optimization algorithms such as PPO and GRPO, and it is also compatible with supervised fine-tuning at test time. Our rigorous theoretical results show that, under the symmetric categorical noise model with known likelihood parameters, AdaPO's electing-stage trigger coincides with the classical sequential probability ratio test (SPRT), while the confirming-stage rule yields a principled posterior-threshold stopping mechanism. Across representative reasoning benchmarks, AdaPO substantially reduces compute, achieving over 40\% token savings on GPQA while maintaining competitive accuracy. The source code will be open upon acceptance at \url{https://open-upon-acceptance}.
Towards Realistic Conversational Multimodal Instruction Following
Kaiwen Tuo ⋅ Congcong Wang ⋅ Shuai Dong ⋅ Hongrui Wu ⋅ Siyuan Wang ⋅ Xuefeng Yin ⋅ Yuhang Cao ⋅ Nan Duan ⋅ Jiaqi Wang
Instruction following is a core capability of modern foundation models, underpinning assistant-style interaction, agentic workflows, visual understanding, and multimodal reasoning. Existing instruction-following data and methods, however, remain substantially misaligned with realistic user behavior: users often converse with models over multiple turns, provide interleaved images and videos, switch modalities, or implicitly expect earlier constraints to persist. To bridge this gap, we first identify and formalize six core challenges for realistic conversational multimodal instruction following: Constraint-Guided Visual Reasoning, Cross-Modal Context Transfer, Persistent Instruction Tracking, User-Aware Context Modeling, Iterative Grounded Response Refinement, and Visual Consistency Alignment. Driven by these challenges, we develop a multi-agent data construction pipeline that generates diverse, constraint-rich, image- and video-grounded conversations, followed by human verification. We further improve model behavior through a rubric-guided multi-stage training pipeline that decomposes complex instructions into atomic supervision signals and trains models from basic visual constraint following to advanced multi-turn conversational constraint satisfaction. Finally, we establish Multimodal MultiChallenge (MMMC) as a systematic benchmark for realistic conversational multimodal instruction following. The results show that realistic conversational multimodal instruction following remains challenging for frontier models, while targeted rubric-guided post-training improves constraint satisfaction without sacrificing general multimodal ability.
Towards Scalable Egocentric HOI for Humanoids: Benchmarking Whole-Body Dexterous Interaction with Tactile Prediction
Zhenyu Wei ⋅ Haoyang Luo ⋅ Chixuan Zhang ⋅ Guo Chen ⋅ Suting Ni ⋅ Siyuan Huang ⋅ Baoxiong Jia ⋅ Jingya Wang
Humanoid whole-body loco-manipulation requires scalable and physically reliable demonstrations, yet existing human-object interaction datasets lack direct contact supervision and egocentric observability, often resulting in physically inconsistent interactions. We present DeXHOI, a multimodal egocentric HOI benchmark for whole-body dexterous interaction, capturing synchronized body motion, object trajectories, tactile sensing, and egocentric RGB-D across diverse spatial layouts and long-horizon activities. To improve physical consistency, we introduce a tactile-guided refinement pipeline that leverages measured pressure to correct hand-object interactions while preserving global motion fidelity. We further establish an Egocentric HOI Recovery Benchmark for contact-aware reconstruction from first-person observations, along with EgoTacti, a tactile prediction model. Together, EgoTacti provides a scalable and physically grounded foundation for learning humanoid whole-body dexterous interaction, pointing toward low-cost egocentric data collection for embodied intelligence.
Towards Unified Memory Adaptation for LLM Agents: Textual, Latent, and Parametric Pathways
Rui Li ⋅ Haoran Tan ⋅ Xiaohe Bo ⋅ Yiqun Chen ⋅ Meizhi Zhong ⋅ Xiaochi Wei ⋅ Yan Gao ⋅ YIWU ⋅ Yao Hu ⋅ Xu Chen
Equipping large language models (LLMs) with external memory banks is a critical stepping stone towards autonomous agents. Current approaches have concentrated substantial efforts on the design of memory construction and retrieval mechanisms. However, they typically relegate memory utilization to the simplistic concatenation of explicit texts for LLM inputs, thereby remaining constrained by the fundamental bottlenecks of context overload and semantic dilution. In response to this dilemma, we present UniMem, a unified memory adaptation framework that empowers LLM agents to holistically harness memory information, extending beyond textual signals into latent and parametric dimensions. Operating along these pathways, UniMem not only elevates contextual density through latent compression, but also facilitates semantic grounding via parametric modulation. Extensive experiments show that UniMem (I) achieves superior performance across diverse benchmarks, including long-term dialogue, memory personalization, and open-domain question answering; and (II) can seamlessly complement prevailing memory construction and retrieval mechanisms for synergistic enhancements.
Towards Universal Black-box Attacks on Graph Neural Networks
Runze Li ⋅ Di Jin ⋅ Xiaobao Wang ⋅ Bingdao Feng ⋅ Dongxiao He
Graph neural networks are widely deployed in real-world applications, where attackers often operate under strict black-box constraints, making effective graph black-box attacks particularly important. However, existing methods are fundamentally limited by poor universality, as attack strategies learned on specific graphs or model configurations fail to generalize across different graph datasets. Inspired by the unified representation principle of graph foundation models, we propose a universal black-box attack framework that learns transferable attack paradigms from multiple source graphs and generalizes directly to unseen target graphs. The framework incorporates a representation selection mechanism and a graph variational autoencoder based modeling strategy to ensure consistency and mappability between the unified semantic space and the original representation space, together with a structure–representation joint anchor node selection mechanism for effective attack construction. Extensive experiments on benchmark datasets show that our method achieves strong universality and consistently superior attack performance.
TPO: Tri-level Distributionally Robust Learning for OOD Direct Preference Optimization
Chengtao Jian ⋅ Kai Yang
Direct Preference Optimization (DPO) minimizes empirical risk over observed preference comparisons, implicitly assuming that the training preference distribution matches the deployment preference distribution. In practice, this i.i.d. assumption often fails, as preference distribution shift exhibits a natural hierarchical structure across three levels: group-level shift, prompt-level shift, and conditional preference-level shift. Such multi-level distributional mismatch poses a critical challenge for safe LLM deployment, yet existing robust DPO methods typically model either overall distributional uncertainty or a single robustness level, leaving the hierarchical structure of OOD preference shifts underexplored. We propose TPO, a tri-level distributionally robust framework for OOD preference optimization, which jointly models all three levels of distributional uncertainty through a hierarchical rectangular ambiguity set. By instantiating it with Wasserstein distance, we derive a tractable and interpretable objective that combines a closed-form preference-level margin penalty, prompt-level adversarial embedding perturbation, and worst-case group reweighting, each governed by an independent radius. We establish finite-sample convergence guarantees and empirically validate the effectiveness of \method{} over existing robust DPO methods.
TPRL: Adaptive Visual Token Pruning in LVLMs via Language-Guided Reinforcement Learning
Sihan Cao ⋅ Jianwei Zhang ⋅ Pengcheng Zheng ⋅ Jiaxin Yan ⋅ Caiyan Qin ⋅ Yalan Ye ⋅ Wei Dong ⋅ Peng Wang ⋅ Yang Yang ⋅ Chaoning Zhang
Large Vision-Language Models (LVLMs) incur substantial inference costs due to the processing of a vast number of visual tokens. Existing methods typically struggle to model progressive visual token reduction as a multi-step decision process with sequential dependencies and often rely on hand-engineered scoring rules that lack adaptive optimization for complex reasoning trajectories. To overcome these limitations, we propose TPRL, a reinforcement learning framework that learns adaptive pruning trajectories through language-guided sequential optimization tied directly to end-task performance. We formulate visual token pruning as a sequential decision process with explicit state transitions and employ a self-supervised autoencoder to compress visual tokens into a compact state representation for efficient policy learning. The pruning policy is initialized through learning from demonstrations and subsequently fine-tuned using Proximal Policy Optimization (PPO) to jointly optimize task accuracy and computational efficiency. Our experimental results demonstrate that TPRL removes up to 66.7\% of visual tokens and achieves up to a 54.2\% reduction in FLOPs during inference while maintaining a near-lossless average accuracy drop of only 0.7\%.
TraceGuard: Defending Multi-Turn Jailbreak Attacks via Prompt-Response Risk Signal Tracking
Hongyi Li ⋅ Yufei Wang ⋅ Chengxuan Zhou ⋅ Qinlin Xie ⋅ Yanting Chen ⋅ Jiawei Ye ⋅ Wu Jie
Large Language Models (LLMs) have made remarkable progress and are increasingly deployed as black-box API services. However, they remain vulnerable to jailbreak attacks that bypass safeguards and elicit harmful content, especially in multi-turn settings where harmful intent is gradually constructed over the interaction. Existing defenses primarily target single-turn scenarios or require access to model internals, making them insufficient against black-box multi-turn jailbreak attacks with weak single-turn discriminability, cross-turn dependency, and dynamic attack patterns. To bridge this gap, we first analyze harmful intent in multi-turn jailbreak interactions through representation-space risk signals inspired by the Linear Representation Hypothesis. Our analysis reveals that these signals are fragmented across prompts, responses, and turns, progressively evolve over the interaction, and exhibit attack-dependent temporal patterns. Motivated by this, we propose TraceGuard, an online black-box defense framework that detects multi-turn jailbreak attacks by tracking prompt and response risk signals as the interaction unfolds. Specifically, TraceGuard operates in two modes: TraceGuard-S captures cross-view risk representations and their cross-turn evolution, while TraceGuard-A further enables distribution-aware few-shot adaptation to dynamic attack patterns. Extensive experiments across multiple benchmarks and target LLMs demonstrate that TraceGuard achieves superior multi-turn jailbreak detection while preserving utility, maintaining efficiency, and generalizing to unseen attacks.
Tracing Psychometric Inference in Large Language Models
Chih-Hao Hsu ⋅ Feng-Chun B Chou ⋅ Pin-Hao A Chen
Do large language models represent psychological constructs internally, or do they only generate outputs that resemble psychometric structure? We address this question using a cross-persona paradigm across 14 models (Llama3 and Qwen2.5, 0.5B--14B, base and instruct). Given responses on one psychological scale from real individuals ($N = 272$), models predict responses on six other scales. At the behavioral level, LLMs reproduce human cross-scale correlation structure and systematically amplify it, with model-generated correlations exceeding human estimates even after correcting for measurement attenuation. We then examine whether this structure is reflected internally. Contrastive direction analysis reveals an organized geometry in activation space aligned with psychometric relationships. This structure emerges in large instruct models but is not observed in base models. Across models, geometry strength predicts behavioral amplification ($r = 0.68$, $p = 0.008$; partial $r = 0.69$ controlling for log size), independent of model size. To relate internal structure to output behavior, a matched activation-probing paradigm shows that representational amplification is less variable than behavioral amplification ($1.38$--$1.77$ vs $0.53$--$1.53$). A synthetic control with known ground-truth structure shows that this range survives subtraction of a ridge-probe baseline ($\approx 1.3$), with adjusted slopes still predicting behavior at $r = 0.88$. We term the systematic representation-to-behavior gap \emph{readout attenuation}. Together, these findings suggest that LLMs encode structured representations aligned with psychological constructs, while differences in output primarily reflect how these representations are read out.
TrackFish3D: Self-Supervised 3D Tracking of Schooling Fish from Multi-view Videos
Patt Phurtivilai ⋅ Zhiyang Dou ⋅ Yifan Wu ⋅ Kinfung Chu ⋅ Yuan Liu ⋅ Lei Yang ⋅ Wenping Wang ⋅ Taku Komura
Quantifying collective fish behavior requires accurate trajectories, yet multi-view 3D tracking remains challenging due to frequent occlusions, visually similar individuals, and the long-standing scarcity of identity annotations. We present TrackFish3D, a geometry-driven self-supervised framework for dense multi-camera 3D tracking of schooling fish. Instead of relying on appearance-based re-identification or manually annotated identities, TrackFish3D turns calibrated multi-view geometry into supervision: triangulation and reprojection consistency provide pseudo-associations, while a geometric encoder and global association transformer learn all-to-all cross-view correspondence within each frame. To make these associations identity-aware, TrackFish3D introduces an effective self-supervised contrastive objective that separates co-visible individuals in the embedding space, together with a temporal predictor that preserves identities and bridges short occlusions across frames. The resulting model is trained once on unlabeled footage and applied directly to unseen test videos, requiring no cross-view identity labels, temporal annotations, 3D ground truth, appearance features, or test-time optimization. On our benchmark, TrackFish3D improves 3D Multi-Object Tracking Accuracy from 83.7% for the strongest baseline to 96.6%. On the 3D-ZeF zebrafish benchmark, it achieves 81.1% MOTA, compared with 77.4% for the best geometric baseline. TrackFish3D also generalizes beyond fish, achieving strong results on real-world bird tracking. We will release our code and model upon acceptance.
TrAction: Action Recognition with Sparse Trajectories
Jan Fredeeik Meier ⋅ Felix B Mueller ⋅ Alexander Ecker ⋅ Timo Lüddecke
Modern action recognition models operate on memory- and compute-intensive dense RGB video volumes and frequently exploit appearance and background shortcuts, for example, predicting actions from objects or scenes instead of characteristic motion. We investigate an alternative input modality that is largely free of such biases by construction: sparse point trajectories. To this end, we develop a simple transformer architecture for 2.5D trajectory-based recognition together with a masked-trajectory pretraining, which we show to substantially improve downstream accuracy. Despite using only a fraction of the dense RGB input, our method reaches 45\% top-1 on Something-Something V2 and 54\% on EPIC-Kitchens-100, and surpasses V-JEPA 2 on time-reversal sensitivity. More importantly, we find trajectory features to be complementary to state-of-the-art appearance-based features: fusing our pretrained model with DINOv2 and V-JEPA 2 improves top-1 accuracy on Something-Something V2 by 8.7 and 1.6 points, respectively.
Diffusion models generate samples by denoising along the score of a perturbed target distribution. In practice, one trains a neural diffusion model, which is computationally expensive. Recent work suggests that score matching implicitly smooths the empirical score, and that this smoothing bias promotes generalization by capturing low-dimensional data geometry. We propose moment-matched score-smoothed overdamped Langevin dynamics (MM-SOLD), a training-free interacting particle sampler that enforces the target moments throughout the sampling trajectory. We prove that, in the large-particle limit, the empirical particle density converges to a deterministic limit whose one-particle stationary marginal is a Gibbs--Boltzmann density obtained by exponentially tilting a naive score-smoothed diffusion target. The mean and covariance of this distribution agree with the empirical moments of the training data. Experiments on 2D distributions and latent-space image generation show that MM-SOLD enables fast, robust, training-free sampling on CPUs, with sample fidelity and diversity competitive with neural diffusion baselines.
Training-Free Task Vectors for LLM Behavioral Control
Gabriel J Perin ⋅ Lucas Boscaini ⋅ Andre Araujo ⋅ Nina S T Hirata
Task vectors enable post-training model editing by identifying semantically meaningful directions in weight space, typically computed as the difference between a fine-tuned model and its pretrained initialization. However, this reliance on fine-tuning makes discovering such directions costly and limits the practicality of post-training model editing. To address this limitation, we introduce Training-Free Task Vectors (TFTVs), a novel method to compute task-vector-like directions without requiring fine-tuning. Our method maps activation steering vectors to rank-one weight-space edits using only forward-pass statistics, while satisfying arithmetic properties that directly support learning via addition, forgetting via subtraction, and the composition of multiple edits. Empirically, we evaluate TFTVs on large language model behavioral control tasks and show that they consistently amplify, suppress, and compose target behaviors while preserving general knowledge and problem-solving skills. We also validate our method against other editing and steering baselines, experimentally demonstrating that TFTVs achieve stronger trait control with better or competitive utility preservation. We hope our work opens new directions for the community in post-training model editing and broader training-free model control. Code will be released.
Training Language Models to Explain Their Own Computations
Belinda Z Li ⋅ Zifan Carl Guo ⋅ Vincent Huang ⋅ Jacob Steinhardt ⋅ Jacob Andreas
Can language models (LMs) learn to faithfully describe their internal computations? Are they better able to describe themselves than other models? We study the extent to which LMs' privileged access to their own internals can be leveraged to produce new techniques for explaining their behavior. Using existing interpretability techniques as a source of ground truth, we fine-tune LMs to generate natural language descriptions of (1) the information encoded by LM features, (2) the causal structure of LMs' internal activations, and (3) the influence of specific input tokens on LM outputs. When trained with only tens of thousands of example explanations, explainer models exhibit non-trivial generalization to new queries. This generalization appears partly attributable to explainer models' privileged access to their own internals: fine-tuning a model to explain its own computations generally works better than fine-tuning a different model (even if the explainer model is significantly more capable than the target). Our results suggest not only that LMs can learn to reliably explain their internal computations, but that such explanations offer a scalable complement to existing interpretability methods.
Training quality is the primary determinant of reasoning efficiency in large language models, yielding greater returns than proportional increases in inference computation. Evaluating 44 language models (0.5B--685B parameters) across six reasoning benchmarks totalling over 67000 assessments, this study establishes when additional inference tokens are beneficial, redundant or actively harmful. Standard models exhibit near-universal accuracy convergence at approximately 1000 tokens regardless of architecture or parameter count. Causal intervention experiments demonstrate that constraining generation length improves reasoning accuracy by 15.4 percentage points, identifying over-generation rather than capacity exhaustion as the primary efficiency bottleneck. Among 13 reasoning-specialised models, training methodology yields 18.8-fold efficiency improvement at matched parameter scales, exceeding gains from ten-fold parameter increases. Graduate-level evaluation reveals that arithmetic benchmark performance does not predict reasoning capability on complex tasks ($r = -0.03$, $P = 0.91$), exposing fundamental limitations invisible to standard evaluation practice. Mechanistic profiling reveals two orthogonal dimensions of training quality---quality control and decomposition efficiency---that independently predict efficiency profiles, providing a measurable framework linking training decisions to inference-time behaviour.
Train the Agent, Not the Expert: Learning to Harness Heterogeneous Experts for Multi-Turn Visual Reasoning
Yaowu Fan ⋅ Tao Han ⋅ Dazhao Du ⋅ Jinhua Ma ⋅ Jia Wan
Recent progress in computer vision has produced a wide range of powerful specialized models for detection, segmentation, counting, and other visual tasks. However, these models are usually optimized for isolated task formulations, making it difficult to directly support general-purpose visual intelligence, especially when a task requires complex language understanding and dense small-object perception. In this paper, we propose VisHarness, a trainable visual agent that decouples high-level perception, reasoning, and decision-making from low-level task execution. Instead of training a model to solve a specific visual task, VisHarness learns to harness a set of carefully designed heterogeneous visual experts. This paradigm preserves the general intelligence of the agent while fully leveraging the precision advantages of specialized visual models in concrete visual tasks. With only lightweight training, VisHarness learns a generalizable visual expert-harnessing policy and can solve common fundamental vision tasks under various complex conditions through multi-turn interactions with visual expert models. To enable efficient on-policy reinforcement learning training in a live environment, we introduce dynamic visual memory archiving, which mitigates the rapidly accumulating visual-token overhead caused by multi-turn interactions with visual expert models. Experiments on four representative benchmarks covering reasoning segmentation, generalized referring segmentation, dense small-object detection, and referring counting demonstrate that VisHarness substantially outperforms existing general-purpose models and achieves competitive or superior performance compared with task-specific models.
TrajEvolve: Trajectory Evolution for Reinforcement Learning with Hindsight Credit Assignment
Yi Wen ⋅ Hao Chen ⋅ Wanyu Wang ⋅ Maolin Wang ⋅ Pengyue Jia ⋅ Derong Xu ⋅ Hui-Ze Tan ⋅ Yingyi Zhang ⋅ Wenlin Zhang ⋅ weihongluo ⋅ Xiku Du ⋅ Xiangyu Zhao
The Reinforcement Learning (RL) paradigm has achieved remarkable success in enhancing the reasoning capabilities of Large Language Models (LLMs). However, when applied to long-horizon tasks, it faces severe challenges: i) Due to the high task complexity, the model struggles to sample enough successful trajectories, resulting in slow convergence. ii) Relying solely on sparse final rewards causes all steps to be updated indiscriminately, encouraging redundant step generation. When both the quantity and quality of successful trajectories become difficult to guarantee, the performance is severely constrained. This dilemma stems from the limitations: existing RL paradigms rely entirely on the reasoning capability of base models, lacking the ability to actively optimize trajectories. To address the above issues, we propose a novel trajectory evolution paradigm with hindsight credit assignment, termed TrajEvolve. TrajEvolve not only relies on the trajectory inferred by the model but also generates ``evolutionary trajectories'' by refining low-quality steps in historical trajectories, thereby exposing the policy to high-quality compositions that lie within the base model's support but are rarely sampled within practical training budgets. Furthermore, TrajEvolve reformulates step-level continuous credit assignment as a binary classification task, whether each step is necessary for task completion. This simplification not only reduces the learning difficulty but also enables step labels to be obtained through counterfactual verification, achieving precise identification of redundant steps. Extensive experiments on two datasets demonstrate the superiority of the proposed method.
Verifiable tool-interactive agents produce rich trajectories that contain far more information than final success or failure. However, standard RL post-training with verifiable rewards often collapses these trajectories into scalar outcomes and trains over a largely fixed data stream, leaving open the question of which examples are most useful for improving the current policy. We introduce TReDS, a trajectory-grounded framework for capability-conditioned training distribution shaping. TReDS estimates multidimensional task requirements from offline probing trajectories and controlled tool-view evidence, maintains the policy capability state in the same semantic space, and uses reliability-gated online correction to form a policy-conditioned requirement representation. A scheduler then converts requirement--capability mismatch, recent utility, and stabilizing pressure into a time-varying training distribution for GRPO, without modifying the GRPO objective. We evaluate TReDS with a progressive stress-test suite spanning $\tau^2$-Bench official evaluation, $\tau$-Bench native-bridge evaluation, BFCL V3, and API Bank, using only the $\tau^2$-Bench training split for policy optimization. TReDS improves over the GRPO baseline under the same training data and backend, and shows stronger behavior retention under same-family and cross-protocol evaluations. These results suggest that RL post-training for tool-interactive agents should optimize not only how trajectories are rewarded, but also where trajectory-revealed requirements allocate training probability mass.
Tree-Guided Identify Then Exploit: A Unified Framework of Pure Exploration and Regret Minimization for Dueling Bandits
Pu Wang ⋅ Yao-Xiang Ding
We study $N$-armed stochastic dueling bandits under the Condorcet-winner assumption, where three widely adopted objectives are considered: best-arm identification (BAI), weak regret, and strong regret. We propose *Tree-Guided Identify-Then-Exploit (TG-ITE)*, the first unified framework to tackle all these objectives to our knowledge. Without requiring stronger assumptions, we propose a shared tree-guided identification approach to find a high-confidence incumbent within $O(N)$ comparisons. We further propose varied exploitation strategies to utilize this warm-start stage to optimize the specific objectives at hand. This methodology enables our approach to (1) achieve $O(N)$ sample complexity in BAI without commonly adopted stronger assumptions; (2) build the first winner-stays-style algorithm to achieve $O(N)$ weak regret; (3) enjoy the same $O(N \log T)$ guarantee as specialized strong-regret approaches; (4) realize the joint optimization of BAI and weak regret with $O(N)$ guarantees for both, eliminating the sub-optimal gap of $O(\log N)$ in the existing approach. Our results provide evidence that the trade-off between BAI and regret minimization is relatively benign in dueling bandits.
Tree-Structured Synergy of Large Language Models and Bayesian Optimization for Efficient CASH
Beicheng Xu ⋅ Weitong Qian ⋅ Lingching Tung ⋅ Yupeng Lu ⋅ Bin CUI
To lower the expertise barrier in machine learning, the AutoML community has focused on the CASH problem, which jointly automates algorithm selection and hyperparameter tuning. While traditional methods like Bayesian Optimization (BO) struggle with cold-start issues, Large Language Models (LLMs) can mitigate these through semantic priors. However, existing LLM-based optimizers generalize poorly to high-dimensional, structured CASH spaces. In this paper, we propose LB-MCTS, a trajectory-structured optimization framework that uses a Monte Carlo Tree Search tree as a shared state for algorithm selection, hyperparameter refinement, and BO–LLM proposer synergy. Within this shared state, BO provides algorithm-specific surrogate modeling for quantitative search, while the LLM exploits path-aware selective memory to generate semantic proposals and reflections. As the surrogate model improves, a reliability-aware proposer policy adaptively shifts from LLM-driven to BO-driven proposals within a unified search trajectory. Experiments on 104 AMLB datasets demonstrate that LB-MCTS consistently outperforms BO-based, LLM-based, and hybrid baselines.
Tree Training: Efficient LLM Training on Tree-Structured Trajectories
Jinghui wang ⋅ Wang Shaojie ⋅ Can Tang ⋅ Yinghan Cui ⋅ Xuxing Chen ⋅ Xiaojiang Zhang ⋅ Chao Wang ⋅ Haotian Zhang
Whenever LLM training involves multiple outputs from the same input---agentic multi-turn trajectories, self-consistency reasoning, beam search distillation, or multi-rollout SFT/RL---the tokens form a tree-structured trajectory with shared prefixes. Existing pipelines linearize such data and treat each branch independently, causing substantial redundant computation that scales with the number of samples. We derive that averaging the loss over all branches is algebraically identical to a per-token weighted loss, reducing the problem to computing each token's log-probability exactly once. We propose DFS serialization of the tree, which visits every token exactly once, and adapt full-attention and SSM layers to ensure the resulting log-probabilities match independent per-branch computation exactly. For memory-constrained settings where the full tree exceeds GPU capacity, we propose Redundancy-Free Tree Partitioning, which achieves zero redundant computation with peak memory bounded by a single root-to-leaf path. Together, these contributions form Tree Training, achieving up to 6.2× end-to-end training speedup on dense and MoE models for both supervised fine-tuning and reinforcement learning.
TRL-Bench: Standardizing Cross-Paradigm Representation-Level Evaluation of Tabular Encoders
Wei Pang ⋅ Xiangru Jian ⋅ Hehan Li ⋅ Zhixuan Yu ⋅ Jingxi Xue ⋅ Jinyang Li ⋅ Zhengyuan Dong ⋅ Xinjian Zhao ⋅ Hao Xu ⋅ Chao Zhang ⋅ Reynold Cheng ⋅ M. Tamer Özsu ⋅ Tianshu Yu
Tabular encoders are usually evaluated inside task-specific end-to-end pipelines, so models from different training paradigms are difficult to compare directly even when they operate on similar tabular signals. We introduce TRL-Bench, a multi-granular tabular representation learning (TRL) benchmark that standardizes cross-paradigm representation-level evaluation: each encoder exports row-, column-, or table embeddings through its supported wrapper, and shared lightweight heads probe them across three suites: TRL-CTbench (column/table), TRL-Rbench (row), and TRL-DLTE (compositional Data-Lake Table Enrichment spanning all three granularities). To support this standardized setting, we release curated benchmark assets and task reformulations, including 50 OpenML tables with 123 verified targets, 16 row-pair linkage rewrites, and a 47,772-table DLTE lake derived from 1,379 parent tables. Across 20 models and 16 tasks, TRL-Bench shows that once downstream conditions are standardized, encoder quality is capability-specific rather than captured by a single leaderboard. At the column/table level, generic text encoders often lead on tasks with strong surface-text signal, while tabular specialists win where their pretraining objective aligns with the task. At the row level, within-table prediction and cross-table linkage favor different training regimes, with the row-matching stage of DLTE pipelines correlating strongly with atomic linkage performance. In TRL-DLTE, the strongest pipelines combine capability-matched specialists rather than reuse a single encoder, and top end-to-end quality depends on non-additive compositional fit rather than per-stage marginal rank alone. TRL-Bench provides a common protocol for measuring reusable signal in exported tabular representations under shared downstream conditions.
Tropical Gaussian Anticoncentration: Settling Optimal Instance-Dependent Bounds for Online Learning in Extensive-Form Games
Ashkan Soleymani ⋅ Zhiyuan Fan ⋅ Lillian Ratliff ⋅ Patrick Jaillet ⋅ Gabriele Farina
We study the minimax regret of full-information online decision making in extensive-form games. For prediction with expert advice, the optimal regret is $\Theta(\sqrt{T\log K})$, where $K$ is the number of experts, or equivalently, the number of normal-form actions. For treeplex strategy spaces, existing efficient algorithms, including KOMWU and online mirror descent with the weight-one dilated entropy regularizer, achieve regret $\mathcal{O}(\sqrt{T\log|\mathcal{V}|})$, where $\mathcal{V}$ is the set of reduced normal-form strategies. We prove that this logarithmic dependence is minimax optimal in the worst-case full-information reward model. The main difficulty is that pure plans in a treeplex are not independent experts: their rewards are correlated through the recursive structure of decisions and observations. To capture this correlation, we introduce tropical Gaussian formulae, recursive Gaussian expressions built from maxima and normalized sums which preserve the variance. The index count of such a formula matches the number of reduced normal-form continuation plans in the corresponding subtree. Our main analytic result shows that every log-balanced tropical Gaussian formula with index count $N$ has Gaussian width $\Omega(\sqrt{\log N})$. While direct approach struggles to recover this lower bound, our proof uses a spherical entropy profile to measure scale-dependent separation among the Gaussian directions, and then applies Tur\'an's theorem and Sudakov minoration to obtain the width lower bound. Applying this result to treeplexes gives an oblivious hard distribution with valid transition outcome and Rademacher terminal rewards, establishing a regret lower bound of $\Omega(\sqrt{T\log|\mathcal{V}|})$. This matches known upper bounds up to universal constants and shows that the DilEnt/KOMWU dependence on the reduced normal-form complexity is unavoidable.
Trust What Matters: Language-Conditioned Evidence Routing for Video-IMU Action Question Answering
Tengjun Ni ⋅ Xin Yuan ⋅ Shenghong Li ⋅ Kai Wu ⋅ Wei Ni ⋅ Ren Ping Liu ⋅ Y. J Guo
Multimodal action understanding is commonly framed as a fusion problem: combine the available sensors and predict an action label. Synchronized video-IMU question answering (QA) exposes a new and fundamental challenge: The modality that matters more is not fixed. A question about scene context may be answered from appearance, a question about motion dynamics may depend on inertial signals, and a question under occlusion, corruption, or temporal misalignment may require deciding which modality to trust. Thus, the core problem is \emph{language-conditioned evidence selection under unreliable cross-modal observations}, as opposed to multimodal fusion. We propose a structured evidence-to-Large Language Model (LLM) framework for video-IMU action QA. Given synchronized video and Inertial Measurement Unit (IMU), our model first extracts visual, sensor, and fused evidence using modality-specific encoders and bidirectional cross-modal attention. Each candidate answers the queries by independently querying these evidence sources, producing candidate-conditioned visual, sensor, and fused evidence summaries. Then, a reliability-aware router estimates observation reliability and selects how much each evidence source should contribute before projecting compact structured evidence tokens into a frozen LLM for multiple-choice answer prediction. We evaluate on a controlled MMAct-based action-QA protocol designed to isolate perception, sensor-centric reasoning, temporal ordering, reliability reasoning, and cross-modal complementarity. Our method achieves $61.20\%$ accuracy on the held-out test split and $57.00\%$ on the Cross-Modal Challenge subset, outperforming text-only, unimodal, naive fused-to-LLM, and shared-state evidence baselines. Ablations show that structured source identity, candidate identity, reliability tokens, and routing are all essential. Perturbation and shortcut analyses confirm that the model relies on paired multimodal evidence rather than language priors or answer-position bias. These results suggest that robust multimodal action reasoning requires moving beyond fixed sensor fusion toward query-specific, reliability-aware evidence organization for language models.
TTVidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
Shih-Ying Yeh ⋅ Daniel Z Kaplan ⋅ Xuehai Wang ⋅ Fu-En Yang ⋅ Min-Hung Chen ⋅ Shang-Hong Lai
Comparisons in video self-supervised learning often evaluate complete training recipes rather than isolating the method itself: architecture, objective, data exposure, schedule, scale, and decoder capacity can all vary at once. This makes it hard to identify which choices yield motion-prioritized representations, whose gains concentrate on frame-to-frame change while retaining useful appearance. We address this with a matched $4 \times 6 = 24$ architecture--objective study at roughly 170M--190M encoder scale on $\sim$1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, and propose TT-VidT. TT-VidT combines a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer, trained by Diff Compression to reconstruct target frames from a first-frame appearance anchor and frame-specific motion tokens. The sweep shows that TT3D with Diff Compression, not either component alone, enters the strongest motion-sensitive regime, and decoder ablations favor a compact video-pretrained decoder. In final comparison, TT-VidT leads Jester, Something--Something\,V2, ARID, and Diving48 fine-tuning simultaneously, improving over the strongest non-TT row by 54\%--121\%, while using 48\% fewer encoder FLOPs than DisMo and 55\% fewer than VideoMAE or V-JEPA~2. HMDB51, IARD, and EPIC-Kitchens bound the claim.
TwinFlux: One-Step Discrete-Continuous Flow for End-to-End Autonomous Driving
Ziming Zhu ⋅ Yu Zhu ⋅ Qin Guo ⋅ Yujia Zhang ⋅ Zheng Shen
End-to-end autonomous driving planning requires trajectories that are multimodal, scene-compliant, and efficient to generate. Deterministic regression is fast but often collapses to a single future, while trajectory-vocabulary methods depend on predefined candidate coverage and iterative generative planners incur multi-step inference cost. We propose TwinFlux, a one-step discrete-continuous trajectory generation framework based on SplitMeanFlow. TwinFlux starts from data-driven trajectory prototypes and refines them through parallel heatmap and offset branches. The heatmap branch provides topology-aware coarse localization on a BEV grid, while the offset branch predicts grid-relative geometric refinement. A differentiable trajectory decoupling-and-reconstruction module connects this representation with physical trajectories. During training, TwinFlux enforces piecewise displacement consistency in reconstructed trajectory space, aligning direct one-step prediction with accumulated sub-interval evolution without Jacobian-vector-product computation or inference overhead. Experiments on NAVSIMv1 and NAVSIMv2 show that TwinFlux outperforms recent end-to-end planning baselines while maintaining real-time inference.
Two Drifts, One Principle: Conflict-Aware Spectral Consolidation for Multimodal Continual Learning
Haiyi Zhang ⋅ Qianyi Cai ⋅ Hanqing Wang ⋅ Zilin Wang ⋅ Xinyue Di ⋅ Yijie Xu ⋅ Yili Wang ⋅ Tianfu Wang ⋅ Yifan Han ⋅ Hui Xiong
Multimodal continual learning requires MLLMs to acquire new domains and abilities sequentially while retaining previous capabilities. MoE-based methods preserve task-specific knowledge, but inference relies on reliable routing. We instead study unified continual multimodal consolidation, which forms one shared model from sequential LoRA updates and projector shifts. LoRA merging naturally supports language-side consolidation, but extending it to MLLMs raises two challenges. First, sequential LoRA updates may interfere and overwrite directions important to earlier tasks. Second, projector mismatch may make the consolidated LoRA update incompatible with the final visual-language alignment. To address these challenges, we propose CASC (Conflict-Aware Spectral Consolidation), a method that jointly consolidates LoRA and projector updates. CASC maintains fixed-rank spectral banks and uses a shared protected subspace memory rule to preserve dominant historical directions while incorporating compatible new updates under fixed rank budgets, yielding a single model without replay data or inference time routing. Both theoretical analysis and experimental results demonstrate the effectiveness of CASC in addressing multimodal continual learning.
Two Stages of Folding: Convergent Mechanisms in AI Protein Folding Trunks
Kevin Lu ⋅ Jannik Brinkmann ⋅ Stefan T Huber ⋅ Aaron Mueller ⋅ Yonatan Belinkov ⋅ David Bau ⋅ Chris Wendler
How do protein structure prediction models fold proteins? We investigate this question through causal interventions on the folding trunks of ESMFold, OpenFold, and Boltz-1. Across all three models, we find a shared two-stage computational structure. In the first stage, early blocks initialize pairwise biochemical signals: features like charge propagate from sequence into pairwise representations through architecture-specific pathways. In the second stage, late blocks develop pairwise spatial features: distance and contact information accumulate in the pairwise representation. We verify these mechanisms causally by showing that steering charge and distance features induces predictable structural changes. Furthermore, these representations are functionally interchangeable: pairwise states can be linearly aligned and substituted across models. Together, these results suggest that folding trunks with different architectures, inputs, and training procedures converge on a shared representational organization for mapping sequence chemistry into spatial geometry.
U-Bench: A Comprehensive Understanding of U-Net through 100-Variant Benchmarking
Fenghe Tang ⋅ Chengqi Dong ⋅ Wenxin (Wendy) Ma ⋅ Zikang Xu ⋅ Heqin Zhu ⋅ Zihang Jiang ⋅ Rongsheng Wang ⋅ Yuhao Wang ⋅ Chenxu Wu ⋅ Yingtai Li ⋅ S. Kevin Zhou
Over the past decade, U-Net has been the dominant architecture in medical image segmentation, leading to the development of thousands of U-shaped variants. Despite its widespread adoption, there is still no comprehensive benchmark to systematically evaluate their performance and utility, largely because of insufficient statistical validation and limited consideration of efficiency and generalization across diverse datasets. To bridge this gap, we present U-Bench, the first large-scale, statistically rigorous 2D benchmark that evaluates 100 U-Net variants across 28 datasets and 10 imaging modalities. Our contributions are threefold: (1) Comprehensive Evaluation: U-Bench evaluates models along three key dimensions: statistical robustness, zero-shot generalization, and computational efficiency. We introduce a novel metric, U-Score, which jointly captures the performance-efficiency trade-off, offering a deployment-oriented perspective on model progress. (2) Systematic Analysis and Model Selection Guidance: We summarize key findings from the large-scale evaluation and systematically analyze the impact of dataset characteristics and architectural paradigms on model performance. Based on these insights, we propose a model advisor agent to guide researchers in selecting the most suitable models for specific datasets and tasks. (3) Public Availability: We provide all code, models, protocols, and weights, enabling the community to reproduce our results and extend the benchmark with future methods. In summary, U-Bench not only exposes gaps in previous evaluations but also establishes a foundation for fair, reproducible, and practically relevant benchmarking in the next decade of U-Net-based segmentation models. The weights, datasets, and results will be released after the acceptance.
Ultra-DPO: Reimagining LLM Alignment as a Semi-Supervised Task
Li He ⋅ Jiaxian Guo ⋅ Yaqing Wang ⋅ Yue Yang ⋅ He Zhao ⋅ Stephen Wan ⋅ Dadong Wang ⋅ Lina Yao ⋅ Tongliang Liu
Direct Preference Optimization (DPO) offers a streamlined one-stage alternative to Reinforcement Learning from Human Feedback (RLHF). However, our theoretical analysis reveals a structural distinction between the two: DPO and RLHF optimize alignment on different data distributions with different supervision sources. Specifically, DPO learns from ground-truth human preferences but within the off-policy support of an offline dataset, while RLHF optimizes on the on-policy rollouts but learns from pseudo-labels produced by a reward proxy. This distinction explains how each method may fail: DPO alignment collapses when the policy drifts away from the offline dataset support, and RLHF alignment is biased due to the extrapolation error of the reward proxy on out-of-distribution rollouts. Effective alignment requires avoiding both failure modes, which neither method can do alone. We achieve this via a semi-supervised alignment objective: ground-truth preferences anchor alignment within the dataset support, and pseudo-labels on online rollouts extend it beyond. Applying DPO's reward reparameterization to this objective yields Ultra-DPO, a one-stage algorithm that augments direct human alignment with an online self-alignment mechanism. Across multiple benchmarks, Ultra-DPO consistently outperforms DPO, RLHF, and online preference learning methods.
Uncertainty Quantification of Least Squares Estimator for Generalized Orthogonal Procrustes Problems
Shenghan Luo
The generalized orthogonal Procrustes problem (GOPP) aims to recover rigid transformations that best align multiple point clouds and has broad applications in 3D geometry, computer vision, and biomedicine. Despite its importance, the problem is challenging due to the nonconvex orthogonality constraints. Fortunately, the generalized power method (GPM) has proven highly effective, theoretically guaranteed to converge to the global least squares (LS) estimator under reasonable conditions. Yet, the statistical inference of this global LS estimator remains largely an open problem. Specifically, the absence of an exact second-order analytic expansion prevents rigorous uncertainty quantification (UQ), which is vital for constructing confidence regions and evaluating estimator reliability. While recent UQ frameworks developed for orthogonal group synchronization offer a potential paradigm, they severely lack universality, often failing to accommodate the complex anisotropic distortions inherent in GOPP shape matrices. To bridge this gap, we propose a more general theoretical framework with which we successfully derive the exact second-order analytic expansion of the GOPP global LS estimator under additive Gaussian noise. Leveraging this theoretical foundation, we further construct confidence regions. Finally, extensive numerical experiments empirically validate the exactness of our theoretical expansion and demonstrate the robust coverage of the proposed confidence regions across varying conditions.
Understanding and Mitigating Under-Confidence in GNNs from the Final Layer
Jincheng Huang ⋅ Jie Xu ⋅ Ping Hu ⋅ Xiaoshuang Shi ⋅ Lei Feng ⋅ Xiaofeng Zhu
Graph Neural Networks (GNNs) have demonstrated remarkable effectiveness on graph-based tasks. However, their predictive confidence is often miscalibrated, typically exhibiting under-confidence, which harms the reliability of their decisions. Existing calibration methods for GNNs normally introduce additional calibration components, which fail to capture the intrinsic relationship between the model and the prediction confidence, resulting in limited theoretical guarantees and increased computational overhead. To address this issue, we propose a simple yet efficient graph calibration method. We establish a unified theoretical framework revealing that model confidence is jointly governed by class-centroid-level and node-level calibration at the final layer. Based on this insight, we theoretically show that reducing the weight decay of the final-layer parameters alleviates GNN under-confidence by acting on the class-centroid level, while node-level calibration acts as a finer-grained complement to class-centroid level calibration, which encourages each test node to be closer to its predicted class centroid at the final-layer representations. Extensive experiments validate the superiority of our method.
Understanding Sample Efficiency in Predictive Coding
Gaspard Oliviers ⋅ Elene Lominadze ⋅ Rafal Bogacz
Predictive Coding (PC) is an influential account of cortical learning. Much of recent work has focused on comparing PC to Backpropagation (BP) to find whether PC offers any advantages. Small scale experiments show that PC enables learning that is more sample efficient and effective in many contexts, though a thorough theoretical understanding of the phenomena remains elusive. To address this, we quantify the efficiency of learning in BP and PC through a metric called "target alignment", which measures how closely the change in the output of the network is aligned to the output prediction error. We then derive and empirically validate analytical expressions for target alignment in Deep Linear Networks. We show that learning in PC is more efficient than BP, which is especially pronounced in deep, narrow and pre-trained networks. We also derive exact conditions for guaranteed optimal target alignment in PC and validate our findings through experiments. We study full training trajectories of linear and non-linear models, and find the predicted benefits of PC persist in practice even when some assumptions are violated. Overall, this work provides a mechanistic understanding of the higher learning efficiency observed for PC over BP in previous works, and can guide how PC should be parametrised to learn most effectively.
Understanding Schedule-Free Methods in Nonconvex Optimization: Rate Guarantees and Escaping Saddles
Jiseok Chae ⋅ Donghwan Kim
Schedule-Free methods have attracted growing interest for alleviating the burden of designing and tuning a learning rate scheduler, while matching and sometimes even outperforming optimizers with tuned schedulers. Despite their strong empirical results, their convergence theory in nonconvex optimization, where modern machine learning objectives typically arise, has remained largely unexplored. In this paper, we provide worst-case analyses of Schedule-Free gradient descent and Schedule-Free stochastic gradient descent, in their standard form and without auxiliary modifications or restrictive conditions, for smooth but possibly nonconvex objectives. Based on a Lyapunov analysis derived from the continuous-time limiting ordinary differential equation associated with these methods, we show that Schedule-Free gradient descent and Schedule-Free stochastic gradient descent achieve the optimal worst-case convergence rates attainable among first-order methods. We further formulate Schedule-Free gradient descent as a nonautonomous dynamical system and prove strict-saddle avoidance under an arbitrarily small one-time perturbation. These theoretical results provide a better understanding of the strong performance that Schedule-Free methods demonstrate.
Understanding the Surprising Generalization Properties of Tabular Foundation Models
Nour Shaheen ⋅ Junwei (Jeremy) Ma ⋅ Alex Labach ⋅ Frank Hutter ⋅ Anthony L Caterini ⋅ Valentin Thomas
Tabular Foundation Models (TFMs) increasingly rely on in-context learning, where a model receives labelled examples at inference time and predicts labels for new inputs without updating its weights. Existing TFMs are typically trained on either massive synthetic corpora or very large collections of real datasets. In contrast, we show that surprisingly strong transfer can emerge from self-supervised pre-training on just a single real table. In this setting, we also find that tables tend to be either broadly useful or broadly poor regardless of downstream prediction task, and that the strongest predictor of usefulness is the number of features rather than the number of instances. This leads to a task-centric interpretation of tabular pre-training: the amount and the quality of tasks are essential for the pre-training of TFMs. We show that the same task-centric perspective can help corpus design at scale: fine-grained column-level pre-processing consistently improves downstream performance, while no improvements are observed when we filter or deduplicate at the dataset level. Finally, we offer a new perspective for how TFMs generalize: we believe that tabular in-context generalization is largely retrieval-based, and good models are those that learn to identify relevant examples in the provided context and aggregate them well. The mechanics of TFMs have been relatively understudied; our task-centric, retrieval-based perspective offers a new framework to guide future model and corpus design.
Unified Forensic Preference Learning for Generalizable Synthetic Image Detection
Laixin Zhang ⋅ Shuaibo Li ⋅ Wei Ma ⋅ Jianyu Lai ⋅ Lei Zhu ⋅ Hongbin Zha
Despite rapid progress in synthetic image detection, detectors that perform well on known generators often become brittle when the source model, post-processing pipeline, or scene composition changes. A key reason is that standard binary supervision teaches models to separate training distributions rather than to identify authenticity-relevant evidence: low-level artifacts become shortcuts, while higher-level semantic implausibilities remain weakly grounded. To address this limitation, we propose UniFPL, a unified forensics preference learning framework that converts multimodal large language model (MLLM) forensic knowledge into image-specific preference supervision. UniFPL constructs forensic preference data from both controlled artifact synthesis and observed artifact harvesting, then learns from list-level evidence rankings that encode which visual-textual cues are more diagnostic of synthetic manipulation. On top of a CLIP-based discriminative detector, UniFPL jointly optimizes authenticity classification, artifacts-aware preference alignment, and semantic structure regularization, allowing the model to inherit contextual forensic reasoning while retaining the stable and efficient inference of a discriminative encoder. Across eleven curated and in-the-wild benchmarks, UniFPL achieves an average balanced accuracy of 93.4\% and improves the worst-case accuracy from 81.4\% to 86.1\% over the strongest baseline, demonstrating stronger cross-generator generalization and robustness to real-world image degradations.
Unified Noise Steering for Efficient Human-Guided VLA Adaptation
Junjie Lu ⋅ Xinyao Qin ⋅ Yuhua Jiang ⋅ Kaixin Wang ⋅ Chuheng Zhang ⋅ Bin Liang ⋅ Jun Yang ⋅ Min Xu ⋅ Li Zhao
Diffusion-based vision-language-action (VLA) models have emerged as strong priors for robotic manipulation, yet adapting them to real-world distributions remains challenging. In particular, on-robot reinforcement learning (RL) is expensive and time-consuming, so effective adaptation depends on efficient policy improvement within a limited budget of real-world interactions. Noise-space RL lowers the cost by keeping the pretrained VLA fixed as a denoising generator while updating only a lightweight actor that predicts the noise. However, its performance is still limited due to inefficient autonomous exploration. Human corrective interventions can reduce this exploration burden, but they are naturally provided in action space, whereas noise-space finetuning requires supervision over noise variables. To address these challenges, we propose UniSteer, a Unified Noise Steering framework that combines human corrective guidance with noise-space RL through approximate action-to-noise inversion. Given a human corrective action, UniSteer inverts the frozen flow-matching decoder to recover a noise target, which provides supervised guidance for the same noise actor that is simultaneously optimized via reinforcement learning. Real-world experiments on diverse manipulation tasks show that UniSteer adapts more efficiently than strong noise-space RL and action-space human-in-the-loop baselines, improving the success rate from 20\% to 90\% in 66 minutes on average across four real-world adaptation tasks.
UniFlow-Audio: Unified Flow Matching for Audio Generation from Omni-Modalities
Xuenan Xu ⋅ Jiahao Mei ⋅ Ye Tao ⋅ Zihao Zheng ⋅ Zeyu Xie ⋅ Yaoyun Zhang ⋅ Haohe Liu ⋅ Yuning Wu ⋅ Ming Yan ⋅ Mengyue Wu ⋅ Wen Wu ⋅ Chao Zhang
Audio generation, including speech, music and sound effects, has advanced rapidly in recent years. These tasks can be divided into two categories: time-aligned (TA) tasks, where each input unit corresponds to a specific segment of the output audio (e.g., phonemes aligned with frames in speech synthesis); and non-time-aligned (NTA) tasks, where such alignment is not available. Since modeling paradigms for the two types are typically different, research on different audio generation tasks has traditionally followed separate trajectories. In this work, we propose UniFlow-Audio, a non-autoregressive audio generation framework that unifies the two types based on flow matching. UniFlow-Audio introduces a dual-fusion mechanism that aligns audio latents with TA features while integrating NTA conditions via cross-attention in each block, and employs task-balanced sampling to maintain consistent performance across diverse tasks. Supporting omni-modal inputs including text, audio, and video, UniFlow-Audio achieves strong results on 7 audio generation tasks with fewer than 8K hours of public training data and under 1B trainable parameters. A compact 200M-parameter variant remains competitive, indicating UniFlow-Audio as a promising foundation model for general non-autoregressive audio generation. Codes and models are available at https://anonymous3387a8c.github.io/uniflow_audio.
Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation
Chao Li ⋅ Tianhong Li ⋅ Sai V Nuthalapati ⋅ Hong-You Chen ⋅ Satya Narayan Shukla ⋅ Jianpeng Cheng ⋅ Yonghuan Yang ⋅ Jun Xiao ⋅ Xiangjun Fan ⋅ Aashu Singh ⋅ Dina Katabi ⋅ Shlok K Mishra
Unifying text-image contrastive learning and text-to-image (T2I) generation in a single end-to-end model is challenging because the two objectives demand opposing masking regimes: contrastive alignment needs near-complete visible tokens, while masked generative modeling needs heavy corruption. We introduce DREAM, a unified framework that resolves this conflict through Masking Warmup, a schedule that shifts the center of the masking distribution over training, so low and high masking ratios coexist at every step. This co-exposure lets a single jointly-trained encoder serve both objectives. The resulting stable optimization unlocks Semantically Aligned Decoding at inference: the text encoder, trained against visual embeddings at all masking ratios, can score partially generated images and select the best trajectory with as little as 12.5\% of the image decoded, improving both FID and throughput. DREAM outperforms its single-objective baselines, CLIP and FLUID: on ImageNet linear-probing (+1.1%), 5-shot transfer (+4.1%), ADE20K segmentation (+1.9%), and NYU depth estimation (+6.25%) over CLIP, and on CC12M FID (+6.2%) over FLUID while maintaining CLIP Score. Together, these gains show that text-image contrastive and generative objectives, when properly unified, are synergistic rather than competing.
UniMoE-World: A Unified Mixture-of-Experts Architecture for Scalable Multi-Control Video Generation World Modeling
Jianjie Fang ⋅ Yongyan Xu ⋅ Ziyou Wang ⋅ Yuchao Huang ⋅ Zhaolu Wang ⋅ Rongze Tang ⋅ Mingyuan Jia ⋅ Baining Zhao ⋅ Weichen Zhang ⋅ Xin Zhang ⋅ Haisheng Su ⋅ Yu Shang ⋅ Chen Gao ⋅ Wei Wu ⋅ Xinlei Chen ⋅ Yong Li
With the rapid progress of interactive video generation, video generation world models have gradually emerged as one of the mainstream paradigms in world model research and are increasingly regarded as a promising path toward efficient intelligent agents. However, existing video generation world models are typically developed under different forms of control supervision and mainly focus either on interactive world modeling or embodied world modeling, leaving the compatibility of heterogeneous control signals largely unexplored. In this work, we introduce UniMoE-World, the first training framework that enables unified learning of world models under heterogeneous supervisory controls by incorporating a Mixture-of-Experts (MoE) design into Diffusion Transformers (DiT). We further propose UniMoE-World Tuning, a continually extensible heterogeneous training strategy for world models, which supports diverse control signals, including robotic arms, hand joints, and camera poses, within a single world model and allows the model to be progressively expanded to new control settings. By enabling joint learning from more diverse data sources, this training strategy alleviates the scaling bottleneck of current world models. Our experiments show that world models trained with MoE over heterogeneous supervision consistently outperform those trained with any single control modality alone, demonstrating clear mutual gains across different control types. UniMoE-World achieves state-of-the-art performance on the WorldArena benchmark and shows particularly strong advantages in both locomotion and hand-motion capabilities over existing methods.
UniRAP: Towards Unified Part-level Physical Affordance Reasoning and Actionable Perception
Linfei Li ⋅ Ruining Hu ⋅ Lin Zhang ⋅ Zhong Wang ⋅ Fengyi Zhang ⋅ Ying Shen ⋅ Binqiang Wang ⋅ Xin Zhang
Vision-language perception has achieved impressive progress in aligning natural language with visual observations, yet grounding high-level semantics into part-level physical interaction remains challenging. To address this gap, we propose UniRAP, a unified model for inferring part-level physical affordances and mapping language instructions to actionable geometric representations. UniRAP formulates this problem as a conditional multimodal generation task, integrating visual inputs, textual instructions, and optional prompts into a shared spatiotemporal representation space through a unified interface token mechanism. To predict executable contact geometry, we introduce a Unified Affordance Decoder (UAD), which jointly performs object detection, part-level affordance segmentation, and 4-DoF interaction pose estimation by leveraging intermediate segmentation features. In addition, we propose a curriculum-based transfer training strategy that progressively adapts the model from general visual parsing to interaction-aware perception, improving data efficiency under limited textual annotations. Experiments show that UniRAP achieves state-of-the-art performance on referring expression segmentation, affordance grounding, and interaction pose estimation, while maintaining strong spatiotemporal consistency in dynamic video scenarios. These results demonstrate the effectiveness of UniRAP as a unified perception framework for language-guided physical manipulation. All data and code will be made publicly available.
Uni-Synergy: Bridging Understanding and Generation for Personalized Reasoning
zijun shen ⋅ Ruichuan An ⋅ Sihan Yang ⋅ Ziyu Guo ⋅ Hao Liang ⋅ Ming Lu ⋅ Renrui Zhang ⋅ Wentao Zhang
Unified Multimodal Models (UMMs) excel in general tasks but struggle to bridge the gap between personalized understanding and generation. Prior works largely rely on implicit token-level alignment via supervised fine-tuning, which fails to fully capture the potential synergy between comprehension and creation. In this work, we propose Sync-R1, an end-to-end reinforcement learning framework that jointly optimizes personalized understanding and generation within a single, explicit reasoning loop. Through this unified feedback process, Sync-R1 enables personalized comprehension to guide content creation, while the resulting generation quality reciprocally refines understanding within an integrated reward landscape. To efficiently orchestrate this dual-task synergy, we introduce Sync-GRPO, a reinforcement learning method utilizing an ensemble reward system. Furthermore, we propose Dynamic Group Scaling (DGS), which adaptively filters low-potential trajectories to reduce gradient variance and accelerate convergence. To better reflect real-world complexity, we introduce UnifyBench++, featuring denser textual descriptions and richer user contexts. Experimental results demonstrate that Sync-R1 achieves state-of-the-art performance, showcasing superior cross-task reasoning and robust personalization without requiring complex cold-start procedures.
Transformers are powerful models but are inefficient with large context due to the quadratic complexity of the attention mechanism. Limiting their context window addresses the efficiency concern but at a cost: they cannot remember everything from the past, which severely limits their computational power. We show that we can simultaneously achieve efficiency and universal computation if we restrict attention heads to be 2-dimensional. First, we give the first algorithm for standard 2D attention that serves every insertion and every query in amortized polylogarithmic time per token. Prior work either restricted key/query magnitudes or paid polynomial overhead per step. Second, we show that the class of 2D-attention models is universal and efficient: a vanilla decoder transformer efficiently simulates an arbitrary RAM machine via chain-of-thought decoding so that, combined with our first result, the simulation runs in $\mathrm{polylog}(N)$ time per simulated RAM step end-to-end. In contrast, earlier universality results addressed only computability, showing Turing completeness while incurring large polynomial factors in their execution time. Third, as an application, we turn a transformer into a computer. We provide a compiler that supports the full $\texttt{i32}$ WebAssembly subset with integer computations (reachable from C via $\texttt{clang}$) and turns such programs into the weights of a transformer, which then executes them over millions of steps autoregressively. We illustrate how this can be used to solve combinatorial problems, tasks at which even large reasoning LLMs struggle.
Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities
Hyunsik Kim ⋅ Youngmoon Jung
Byte-level byte-pair encoding (BBPE) tokenizers are attractive for large language models (LLMs) in multilingual settings because they cover all Unicode text. Under UTF-8, however, many scripts start from a higher byte-level fallback cost than English: when no learned merge covers a character $c$, the tokenizer must emit multiple byte-derived symbols. We call this worst-case pre-merge cost the encoding floor: $$ f(c) = |\\mathrm{utf8}(c)| \\in \\{1,2,3,4\\}. $$ A higher floor inflates token budgets, shrinks usable context, and increases per-request cost. Changing the underlying text encoding can reduce this gap, but a single global encoding can also make already-efficient English spans more expensive in mixed-script text. We propose Universal Byte-Level Encoding (UBE), a dual-alphabet tokenizer with deterministic per-character routing: $$ r(c) = \\begin{cases} \\mathrm{UTF\\text{-}8}, & |\\mathrm{utf8}(c)| \\leq 2,\\\\ \\mathrm{UTF\\text{-}16}, & |\\mathrm{utf8}(c)| > 2. \\end{cases} $$ Characters satisfying $|\\mathrm{utf8}(c)| \\leq 2$ stay on the UTF-8 path, while characters satisfying $|\\mathrm{utf8}(c)| > 2$ are routed through UTF-16. This reduces token costs for Basic Multilingual Plane (BMP) scripts whose characters typically satisfy $|\\mathrm{utf8}(c)| = 3$ and have the highest token premiums, without raising costs for already-efficient spans. UBE changes only the byte representation presented to byte-pair encoding (BPE): the BPE merge rule remains standard, and exact decoding $\\mathrm{decode}(\\mathrm{encode}(s)) = s$ is preserved. Across intrinsic tokenization evaluations, UBE reduces cross-lingual dispersion in English-normalized token-count ratios $\\pi(\\ell) = m(\\ell)/m(\\mathrm{en})$, lowering cross-script token-budget disparity. In multilingual language model experiments, UBE preserves LM quality within seed variability. In the main multilingual settings, UBE reduces token counts most for high-premium scripts and also slightly lowers English token counts; these reductions translate into more usable context under fixed token budgets and faster content-matched prompt processing.
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs
Houcheng Jiang ⋅ Jiajun Fu ⋅ Junfeng Fang ⋅ Chen Gao ⋅ Xiang Wang ⋅ Xiangnan He ⋅ Yong Li
Multimodal large language models are increasingly expected to perform \emph{thinking with images}, yet existing visual latent reasoning methods still rely on explicit textual chain-of-thought interleaved with visual latent tokens. This interleaved design limits efficiency and keeps reasoning fragmented across separate text and vision channels. We propose UniVLR, a unified visual latent reasoning framework that treats textual reasoning and auxiliary visual evidence as a shared visual workspace. Instead of preserving text CoT as an independent inference-time path, UniVLR renders reasoning traces together with auxiliary images and learns to compress this unified representation into compact visual latent tokens. At inference time, the model reasons only through visual latents and directly decodes the final answer, avoiding both external tool calls and verbose text reasoning. Experiments on real-world perception and visual reasoning tasks show that UniVLR outperforms prior visual latent reasoning methods while using substantially fewer generated reasoning tokens, suggesting a more unified and efficient paradigm for visual thinking in MLLMs. Our code is available at: https://anonymous.4open.science/r/UniVLR-1B0C/.
Unlocking LLM Creativity in Science through Analogical Reasoning
Andrew Shen ⋅ Shaul Druckmann ⋅ James Zou
Autonomous science promises to augment scientific discovery, particularly in complex fields like biomedicine. However, this requires AI systems that can consistently generate novel and diverse solutions to open-ended problems. We evaluate LLMs on the task of open-ended solution generation and quantify their tendency to mode collapse into low-diversity generations. To mitigate this mode collapse, we introduce analogical reasoning (AR) as a new approach to solution generation. AR generates analogies to cross-domain problems based on shared relational structure, then uses those analogies to search for novel solutions. Compared to baselines, AR discovers significantly more diverse generations (improving solution diversity metrics by 90-173\%), generates novel solutions over 50\% of the time (compared to as little as 1.6\% for baselines), and produces high-quality analogies. To validate the real-world feasibility of AR, we implement AR-generated solutions across four biomedical problems, yielding consistent quantitative gains. AR-generated approaches achieve a nearly 13-fold improvement on distributional metrics for perturbation effect prediction, outperform all baselines on AUPRC when predicting cell-cell communication, infer brain region interactions with a high Spearman correlation ($\rho$=0.729) to published methods, and establish state-of-the-art performance on 2 datasets for oligonucleotide property prediction. The novel and diverse solutions produced by AR can be used to augment the search space of existing solution generation methods.
Unveiling Entropy-Performance Decoupling in Agentic RL for Tool-Integrated Reasoning
Yirong Zeng ⋅ Shen You ⋅ Yufei Liu ⋅ Hao Cong ⋅ Yutai Hou ⋅ Xiao Ding ⋅ Yuxian Wang ⋅ Wu Ning ⋅ WangXu ⋅ Bibo Cai ⋅ Dandan Tu ⋅ Wenyu Zang ⋅ Ting Liu
Tool-Integrated Reasoning (TIR) enables LLMs to solve complex, multi-step problems by offloading computation to external tools. While Agentic Reinforcement Learning (RL) has advanced TIR systems, the underlying policy entropy dynamics governing these models remain poorly understood. In this paper, we conduct an extensive empirical study across model architectures and uncover a pervasive phenomenon of entropy-performance decoupling: in the late stages of RL training, policy entropy continues to diverge despite performance stagnation. We identify that this ineffective entropy growth is primarily driven by an accumulation of invalid tool-use trajectories, which trap the agent in non-informative hallucination loops and degrade final reasoning accuracy. To address this bottleneck, we propose EarlyTIR (Early intervention for Tool-Integrated Reasoning), a strategic truncation mechanism that prunes non-informative exploration during the RL rollout phase. EarlyTIR halts trajectories that exhibit excessive invalid interactions while preserving the model’s ability to learn essential error-recovery behaviors. Empirical evaluations across six public benchmarks and multiple model families (7B to 32B) demonstrate that EarlyTIR consistently elevates the performance ceiling, achieving a macro-average accuracy boost of 13.6\% on the Qwen3-32B model. Furthermore, our approach enhances inference efficiency, reaching correct solutions with 33.7\% fewer tool interactions compared to strong baselines.
UTOPI: Efficient Egocentric Long-Video Understanding in AR via User-Guided Token Pre-Compression
ziqi wang ⋅ Su Chen ⋅ Qiance Tang ⋅ Jieyu Lin ⋅ Ziyun Li ⋅ Barbara De Salvo ⋅ Sai Qian Zhang
As long-video understanding becomes increasingly important, augmented reality (AR) devices offer a natural platform for deploying egocentric video intelligence. Yet first-person videos often contain substantial temporal redundancy, creating heavy memory and compute demands for resource-limited AR hardware. We present~\textit{UTOPI}, a plug-and-play pre-compression module tailored to egocentric long-video processing on AR devices. UTOPI first uses user motion cues to estimate cross-frame overlap and remove redundant content through a viewpoint-shift-robust pruning strategy. It then leverages eye-tracking signals to detect user fixation regions, separating attention-relevant foreground from background and preserving important content through a fine-tuned neural adjustment module. Evaluated across multiple benchmarks and model backbones, UTOPI reduces video tokens by up to $95%$ while maintaining task performance.
Large-scale video–text contrastive learning has become the de facto standard for video understanding, yet mainstream methods almost universally adopt the linear patch embedding of vision Transformers (ViT) or evenly sliced tubelets as the front end. We revisit this front end from the biological principles of the primary visual cortex (area V1) and introduce V1DynamicVisionSystem (V1-DVS), a fully end-to-end trainable, bio-plausible video embedding framework. Its main contributions are: (i) a three-pathway parallel architecture that maps onto layer 4C orientation selectivity, layer 4B motion direction selectivity, and the cytochrome-oxidase Blob color-opponent system; (ii) an explicit shape–color decoupling, with isotropic Blob processing running in parallel with anisotropic Gabor orientation processing; (iii) long-range spatiotemporal convolution that captures inter-frame optical flow through 3D filters and the Adelson–Bergen motion-energy model rather than naive frame stacking; and (iv) bidirectional cross-pathway attention, allowing motion and color features to mutually modulate, as in V1→V4 inter-areal feedback. Pretrained on CC3M, WebVid-10M, ActivityNet and Ego4D and evaluated under a fully self-supervised objective(SimCLR + temporal-order verification + masked-token prediction), V1-DVS outperforms four strong baselines—linear patch embedding, mean temporal pooling, tubelet embedding, and spatiotemporal latent encoding—on six metrics spanning retrieval and embedding-space geometry, validating the practical value of biologically grounded structure for video embedding.
VAMIRec: Value-Aware Memory Intervention for Continual Recommendation
Xudong Zhao ⋅ Xiaolong Xu ⋅ Haolong Xiang ⋅ Xiaoyu Xia ⋅ Jiaqiang Zhang ⋅ Xuyun Zhang ⋅ Lianyong Qi ⋅ Wanchun Dou ⋅ Amin Beheshti
Continual recommendation captures dynamic user interests amid sequential interactions, enabling models to adapt to new observations while preserving prior valuable knowledge. Existing continual recommendation methods primarily preserve or reuse historical knowledge through replay, distillation, or regularization to mitigate forgetting during incremental updates. However, noisy or obsolete inherited representations render indiscriminate preservation detrimental by inducing memory conflicts with current stage preferences. Moreover, rapid shifts in user interests require timely and selective memory updates, while coarse global mechanisms fail to provide such fine-grained responsiveness. To address these issues, we propose VAMIRec, a value-aware memory intervention framework for continual recommendation. VAMIRec first constructs a memory state by separating inherited user memory into a slow anchor component and a fast-changing drift component. We further design a value-aware action evaluation module to generate three candidate operations, and estimate conservative action values by accounting for uncertainty and intervention cost. Finally, we develop a policy-guided memory intervention module that distills the training-time action supervision into an inference-safe policy and applies the selected memory action before all-item ranking. Extensive experiments on three real-world datasets demonstrate that VAMIRec consistently outperforms the state-of-the-art baselines.
Variational Approach to Optimal IPS Estimator for Multi-logger Off-Policy Evaluation
Joon Suk Huh ⋅ Junghoon Seo
We study off-policy evaluation (OPE) in contextual bandits with data collected from multiple logging policies. Inverse propensity scoring (IPS) is a standard approach to OPE and extends naturally to the multi-logger setting. However, as highlighted by Agarwal et al. [2017], there appears to be no IPS estimator that consistently outperforms the others in this setting. We resolve this dilemma by deriving an optimal IPS estimator with sample-dependent weights that minimize variance subject to unbiasedness. Using a variational calculus approach, we obtain closed-form optimal weights, yielding an estimator that is unbiased and achieves asymptotically optimal variance within an weighted-IPS estimator class. Experiments on benchmark datasets confirm this theoretical resolution in practice, showing that our estimator consistently outperforms existing multi-logger IPS methods. We also extend our estimator to a doubly robust form by incorporating a conditional reward estimator. We compare the resulting DR extension with the semiparametrically efficient DR estimator of Kallus et al. [2021], both theoretically and empirically. We show that the two estimators achieve the same asymptotic variance when the conditional reward estimator is consistent, while our estimator can retain semiparametric efficiency under certain forms of reward misspecification. Empirically, the proposed DR estimator achieves lower variance than the existing DR baseline.
Variational Consequence-Driven Offline Reinforcement Learning
Wen Jiang ⋅ Ke Jiang ⋅ Jin Wang ⋅ Xiaoyang Tan
Offline reinforcement learning (RL) typically mitigates distribution shift by imposing divergence constraints on the action distributions of the learned and behavior policies. However, this paradigm cannot ensure the underlying dynamic consistency structure: even minor action deviations can lead to drastically different transitions, inadvertently driving the agent away from the offline dataset. To address this, we explore policy constraints through their induced state-consequence distributions rather than relying on pure action-distribution matching. Nevertheless, directly enforcing such consequence-aware constraints is computationally intractable without access to the true underlying dynamics model. To bypass this issue, we propose Structural Behavior Regularization (StBR), which replaces the intractable objective with a closed-form latent divergence. Structuring this latent space to reflect transition dynamics renders the latent divergence a theoretically bounded surrogate for the true consequence-distribution divergence. By regularizing the learned policy with this latent divergence, StBR adaptively tightens the behavior constraint in dynamics-sensitive regions, effectively preventing deviations outside the offline dataset. Empirically, we demonstrate that StBR achieves superior performance across diverse offline RL benchmarks.
Velocity-Space 3D Asset Editing
Hao Liu ⋅ yuxuan lin ⋅ jingfeng Guo ⋅ Ruihang Chu ⋅ junjie wang ⋅ Ruotong Li ⋅ Yujiu Yang
Editing a 3D asset locally, modifying a target region while preserving the rest, is a fundamental requirement of native 3D editing. Existing methods enforce locality through mechanisms external to the generator, such as manual 3D masks, post-hoc voxel merging, or 2D multi-view lifting. None of them intervene where the corruption actually originates: inside the ODE sampler. For a rectified-flow generator to achieve faithful local editing, its velocity field should be strong over the target editing region while vanishing on preserved content. Yet a single velocity field can hardly satisfy both requirements simultaneously, leading to three problems: (i) identity leakage that keeps the edit signal non-zero on preserved regions; (ii) no dedicated edit-amplification channel, so strengthening the edit inevitably perturbs identity; and (iii) an identity drag at the geometry and material stages, where a global condition pulls every token toward the target. We propose VS3D (Velocity-Space 3D Asset editing), an inversion-free, training-free, and mask-free framework that addresses each problem with a targeted intervention inside the sampler. VS3D integrates three complementary modules, each corresponding to a specific stage of the editing pipeline. Reconstruction-Anchored Source Injection (RASI) absorbs identity leakage by turning the unconditional embedding into a per-step, asset-specific anchor calibrated through source reconstruction. Partial-Mean Guidance (PMG) amplifies the edit signal by contrasting high- and low-quality subsample estimates of the velocity difference, active only where a consistent edit exists. Twin-Agreement Residual injection (TAR) lets the sampler decide token by token what to preserve at the geometry and material stages. Experiments on diverse 3D assets show that VS3D outperforms state-of-the-art native-3D and 2D-lifted editors, demonstrating that a purely velocity-space approach can serve as a general-purpose editing paradigm for pretrained 3D generators.
VGB for Masked Diffusion Model: Efficient Test-time Scaling for Reward Satisfaction and Sample Editing
Kijung Jeon ⋅ Thuy-Duong Vuong ⋅ Molei Tao
Inference-time reward guidance is a simple way to improve generative models when completed outputs can be checked or scored by an external verifier. Best-of-N sampling is a particularly strong baseline: it is parallelizable and often highly improves with more samples. However, when the reference model assigns low probability to reward-satisfying samples, selection among independent full rollouts becomes sample-inefficient. We introduce MDM-VGB, a reward-guided discrete diffusion sampler that augments masked generation with value-guided re-masking. MDM-VGB extends the autoregressive VGB backtracking chain from fixed left-to-right prefix trees to any-order masked-state graphs, allowing the sampler to reveal and revise tokens at any position. The resulting Markov chain favors local reveal and re-mask moves that lead to higher-value partial masked states, enabling both root-start reward tilting and leaf-start editing of low-reward samples. We further introduce a flow-cancelled momentum lift of MDM-VGB that preserves the projected target law while reducing oscillatory reveal/re-mask dynamics under finite budgets. We prove that the sampler has the correct reward-tilted leaf law, places non-negligible stationary mass on complete outputs, and remains robust to partial-state verifier error. Across structured reward satisfaction and repair tasks, MDM-VGB improves the quality--cost frontier over reference sampling, best-of-N, and forward-only value-guided rollout.
VideoMDM: Towards 3D Human Motion Generation From 2D Supervision
Amir Mann ⋅ Gal M Harari ⋅ Merav keidar ⋅ Or Litany
We introduce VideoMDM, a diffusion-based framework that trains 3D human motion priors directly from accurate 2D poses extracted from monocular videos, without any 3D ground truth. A pretrained 2D-to-3D lifter provides approximate 3D pose sequences that serve as a noisy teacher: these are diffused, denoised by the model in 3D, and supervised in 2D by reprojecting the prediction and comparing against accurate keypoints. We show that, under mild assumptions, a depth-weighted 2D reprojection loss is equivalent in expectation to direct 3D supervision, and we adapt standard 3D motion regularizers — velocity consistency and over-parameterized representation alignment — to this 2D setting. Unlike methods that lift 2D to 3D only at inference, VideoMDM learns a coherent 3D motion manifold during training. On HumanML3D it nearly closes the gap to fully 3D-supervised MDM (FID 0.88 vs 0.54); On real video datasets Fit3D and NBA the method learns to generate motions consistently preferred by humans, with strong quantitative results.
VistaQA: Benchmarking Joint Visual Question Answering and Pixel-Level Evidence
Mozhgan Nasr Azadani ⋅ Yimu Wang ⋅ Yongpeng Zhu ⋅ Lihong Chen ⋅ Milan Ganai ⋅ Sean Sedwards ⋅ Marco Pavone ⋅ Krzysztof Czarnecki
Establishing a clear link between model predictions and the visual evidence that supports them is critical for transparency and reliability in multimodal reasoning, yet current multimodal large language model (MLLM) evaluations do not explicitly enforce this alignment. Existing benchmarks assess either textual answer correctness or pixel-level localization in isolation, leaving the coupling of reasoning and grounding an open challenge. We introduce VistaQA, a comprehensive benchmark for joint evaluation of free-form answer correctness and pixel-level evidence grounding in visual question answering. VistaQA comprises 1,157 expert-curated samples spanning six task types and six visual domains, ranging from direct perception to compositional and relational reasoning. VistaQA requires models to not only answer correctly, but to also provide precise segmentation masks that support their answers. It also includes hallucination-aware examples where no valid visual evidence exists. To support this enhanced evaluation, we introduce GROVE, a unified evaluation metric that enforces joint correctness by combining textual accuracy and grounding quality via a per-sample geometric mean, ensuring neither dimension can compensate for deficiencies in the other. Comprehensive experiments across grounding-aware models and hybrid pipelines with general-purpose MLLMs reveal that even the strongest systems achieve limited performance under GROVE, highlighting a substantial gap between answer accuracy and visual evidence alignment.
Visual-ERM: Reward Modeling for Visual Equivalence
Ziyu Liu ⋅ Shengyuan Ding ⋅ Xinyu Fang ⋅ Xuanlang Dai ⋅ Penghui Yang ⋅ Jiaqi Wang ⋅ Kai Chen ⋅ Dahua Lin ⋅ Yuhang Zang
Vision-to-code tasks require models to reconstruct structured visual inputs, such as charts, tables, and SVGs, into executable or structured representations with high visual fidelity. While recent Large Vision Language Models (LVLMs) achieve strong results via supervised fine-tuning, reinforcement learning remains challenging due to misaligned reward signals. Existing rewards either rely on textual rules or coarse visual embedding similarity, both of which fail to capture fine-grained visual discrepancies and are vulnerable to reward hacking. We propose Visual Equivalence Reward Model (Visual-ERM), a multimodal generative reward model that provides fine-grained, interpretable, and task-agnostic feedback to evaluate vision-to-code quality directly in the rendered visual space. Integrated into RL, Visual-ERM improves Qwen3-VL-8B-Instruct by +8.4 on chart-to-code and yields consistent gains on table and SVG parsing (+2.7, +4.1 on average), and further strengthens test-time scaling via reflection and revision. We also introduce VisualCritic-RewardBench (VC-RewardBench), a benchmark for judging fine-grained image-to-image discrepancies on structured visual data, where Visual-ERM at 8B decisively outperforms Qwen3-VL-235B-Instruct and approaches leading closed-source models. Our results suggest that fine-grained visual reward supervision is both necessary and sufficient for vision-to-code RL, regardless of task specificity.
Visual Harness: Grounding Multimodal Reasoning in Physics Engines
Chentao Cao ⋅ Zhanke Zhou ⋅ Sangni Duan ⋅ Bo Han ⋅ Hang Li
Multimodal agents struggle with physical-world reasoning, particularly spatial reasoning from partial views, as the physical world is inherently complex. We think multimodal agentic reasoning about the physical world should ground its actions in deterministic feedback, rather than rely solely on internal imagination. Concretely, we propose to equip vision–language models (VLMs) with a physics engine (e.g., MuJoCo, UE5), which simulates the consequences of the agent's actions. However, granting the agent raw access to a physics engine is not sufficient, as the agent must first reconstruct the scene inside the physics engine. We therefore introduce the Visual Harness System, an orchestration system that decouples perception from reasoning with a physics engine as the backend. The perception module composes off-the-shelf perception skills, such as camera-pose estimation, open-vocabulary detection, and metric depth, into a structured scene map. The reasoning module then issues multi-turn tool calls that the engine executes over the map, and composes the final answer from the grounded feedback it receives. Experiments on spatial reasoning benchmarks show that our approach consistently outperforms strong baselines across both open- and closed-weight VLMs, with up to $16.57\%$ improvement over prior methods on the MindCube benchmark.
Visual-Redundancy-Controlled Parallel Decoding for Diffusion-Based Multimodal Large Language Models
Yulin Yuan ⋅ Hongshuo Zhao ⋅ Xiangming Meng
Diffusion-based multimodal large language models (dMLLMs) decode by iteratively predicting tokens at multiple masked positions in parallel. This turns each decoding step into a position-selection problem: the model must choose not only which predictions are reliable in isolation, but also which positions should be committed together as context for later decoding steps. Existing confidence-based decoding ranks masked positions independently and commits the top-$K$ positions, largely ignoring whether the committed tokens provide complementary visual grounding. We identify a step-level limitation of this strategy in multimodal settings: high-confidence tokens selected in the same step can rely on overlapping visual grounding, introducing visual redundancy among the committed tokens and leaving less complementary visual grounding available for later decoding. To quantify this effect, we introduce the Visual Redundancy Index (VRI), which measures visual grounding overlap among tokens committed in parallel. To control this redundancy during decoding, we propose Visual-Redundancy-Controlled Decoding (VRCD), a training-free inference-time decoding method that uses token-to-image attention to prioritize visually complementary positions. Across diverse multimodal benchmarks, VRCD reduces visual redundancy and remaining-position entropy with modest runtime overhead. In longer decoding experiments, it also achieves relative accuracy gains of up to $18.8\\%$ on M$^3$CoT and $6.9\\%$ on MMBench over confidence-based decoding.Code is included in the supplement.
Visual Reinforcement Fine-Tuning via Bootstrapped Medical Reasoning
Yequan Bie ⋅ Peng Xie ⋅ Zhixuan CHEN ⋅ Yihui Wang ⋅ Andong Tan ⋅ Linshan Wu ⋅ Sunan He ⋅ Jianda Mao ⋅ Yangqiu Song ⋅ Kani Chen ⋅ Hao CHEN
Multimodal large language models (MLLMs) have made significant strides in natural vision-language understanding, yet their potential in specific domains, such as healthcare, remains largely untapped. Since existing medical models often lack the reasoning capabilities needed for complex decision-making, adapting general MLLMs for medical reasoning has attracted increasing interest, with reinforcement fine-tuning (RFT) emerging as a promising approach. However, medical reasoning tasks require both precise thinking processes and generalization of well-justified answers, presenting unique challenges due to the inherent scarcity of annotated data and nuanced visual complexity of medical images. Current visual RFT methods prioritize answer correctness through verifiable rewards while neglecting the reasoning process, leading to limited reasoning capabilities and sub-optimal performance, which are essential in high-stakes scenarios like healthcare. To address these issues, we propose MedR2FT, a Medical Reasoning Reinforcement Fine-Tuning framework that enhances medical reasoning through rationale bootstrapping and reasoning rewarding. Specifically, MedR2FT leverages distilled rationales from fundamental training phases to supervise subsequent reasoning processes with semantic rewards, alleviating the challenge of scarce reasoning data and enhancing model reasoning. Extensive experiments demonstrate that our method consistently outperforms baselines on visual question answering and lesion detection with a significant margin. MedR2FT explicitly attends to the reasoning process, significantly enhancing MLLM adaptability for medical reasoning. The code will be available.
VisualWorldBench: A Fine-Grained Multi-Task Benchmark for Evaluating Visual World Knowledge in MLLMs
Chen Chen ⋅ Zheming Liang ⋅ Jiaqi Wang
Visual world knowledge is a core capability of multimodal large language models (MLLMs), underpinning knowledge-intensive visual question answering and fine-grained visual reasoning. However, existing benchmarks fall short along two complementary dimensions. First, category coverage is limited: most benchmarks cover only a few thousand fine-grained entity categories, often within narrow domains, far from the open-ended scale of real-world visual entities. Second, task diversity is limited: most benchmarks rely on single task formats such as closed-set classification or attribute QA, and therefore cannot characterize the multi-dimensional structure of visual world knowledge. We introduce VisualWorldBench, a taxonomy-grounded multi-task diagnostic benchmark built on a hierarchical taxonomy of 21K+ leaf-node classes across 8 visual domains, comprising 32K+ instances across four complementary tasks: closed-set recognition, contrastive selection, precise localization, and open-world naming. To our knowledge, this is the largest category coverage among visual world knowledge benchmarks to date. VisualWorldBench reveals substantial capability gaps hidden under conventional benchmarks: the strongest closed-source model achieves only 80.4% average accuracy, and single-model spread across tasks can exceed 40 percentage points, indicating that visual world knowledge is multi-dimensional rather than monolithic. We further release VisualWorldBench-50K, a training resource covering 50K+ categories. Fine-tuning on VisualWorldBench-50K yields over 21 percentage points improvement on VisualWorldBench, with consistent transfer to public fine-grained recognition and visual knowledge benchmarks.
VL-DocIR: A Benchmark for Vision-Based Long Document Retrieval
Simon Ott ⋅ Yufan Chen ⋅ Ruiping Liu ⋅ Junwei Zheng ⋅ Jiale Wei ⋅ Di Wen ⋅ Kunyu Peng ⋅ Jiaming Zhang ⋅ Rainer Stiefelhagen
Vision-based document retrieval has improved rapidly, but current evaluation settings still emphasize short documents and single relevant pages. This leaves an important gap between leaderboard performance and realistic long-document retrieval, where the evidence needed to answer a query may span multiple pages or even multiple documents. We present VL-DocIR, a page-level benchmark for evidence-complete visual long-document retrieval. VL-DocIR is built from 29,641 HTML-born documents rendered into 388,548 page images from Wikipedia, arXiv, PubMed, and SEC proxy statements. It contains 271,760 queries across 23 domains and six evidence structures, covering single-page, within-document multi-page, and cross-document evidence configurations. Evidence annotations are grounded to rendered pages and HTML element identifiers, and queries are filtered with a cleaning pipeline targeting clarity, correctness, closedness, and the absence of explicit layout references. We evaluate leading single-vector and multi-vector visual retrievers together with diagnostic retrieval settings such as document-context scoring, query splitting, reranking, and round-robin ensembles. The strongest evaluated configuration reaches 75.6 nDCGAll@10 and 86.2 RecallAll@10, but source- and evidence-structure breakdowns show that cross-document retrieval remains substantially harder than single-page retrieval, and that document-context scoring is especially helpful for within-document multi-page queries. VL-DocIR exposes these failure modes and provides a more demanding target for visual long-document retrieval.
VLM$^3$: Vision Language Models Are Native 3D Learners
zhipeng cai ⋅ Zhuang Liu ⋅ Yunyang Xiong ⋅ Zechun Liu ⋅ Vikas Chandra ⋅ Yangyang Shi
Vision Language Models (VLMs) enable a unified model to solve various vision tasks through prompting. They have shown promising performance in semantic understanding. However, 3D understanding still largely relies on expert vision models with complex task-specific designs. The key argument this work wants to make is that VLMs are native 3D learners. Our in-depth large scale study shows that 1) focal length unification, 2) text-based pixel reference and 3) data mixture and scaling, are all you need for effective 3D learning. Model architecture changes, large models, heavy data augmentations, and complex losses including the regression formulation, many of which form the foundation of expert vision models, are actually not necessary conditions. As a result, we propose VLM$^3$, a scalable method with the simplest design that enables standard VLMs to master diverse 3D tasks. VLM$^3$ not only advances the VLM depth estimation accuracy by a large margin (0.84 $\rightarrow$ 0.9), but also enables diverse 3D tasks such as pixel correspondence, camera pose estimation and object-level 3D understanding, matching expert vision model accuracy while maintaining standard architectures and text-based training. Code will be released to the community.
Volatility-Whitened Probabilistic Residual Modeling for Long-Term Time Series Forecasting
Fan Zhang ⋅ Shiming Fan ⋅ Zexuan Ma ⋅ Shijun Chen ⋅ Meijia Wang ⋅ Hua Wang
Compared with deterministic methods that output only point estimates, probabilistic forecasting can characterize both future trends and their uncertainty simultaneously, making it more suitable for decision-making in complex real-world scenarios. In recent years, diffusion models, owing to their powerful generative modeling capability, have been introduced into the field of time series forecasting. However, we find that most existing diffusion-based forecasting methods directly construct the diffusion process in the original future space, which is easily affected by heteroscedastic scale imbalance, causing highly volatile dimensions to dominate the denoising learning and leading to unstable corrections near observation boundaries. Motivated by these issues, we propose VolaRM, a volatility-whitened probabilistic residual modeling framework for long-term time series forecasting. By constructing a conditional probabilistic base distribution, it represents future targets as whitened residuals relative to this base distribution and performs conditional modeling in the normalized residual space, thereby alleviating scale imbalance across variables and forecasting horizons. Meanwhile, a specific gating mechanism is introduced to enhance boundary continuity and the stability of long-term forecasting. Extensive experiments on eight real-world datasets from different application domains demonstrate the effectiveness of VolaRM.
VoluCore: Spanning Teacher Representations with Volumetric Coresets for Data-Efficient LLM Distillation
Wang Xi ⋅ Yue Wang
Data-efficient distillation of large language models depends not only on the student architecture, but also on which teacher examples are selected for training. Existing selection strategies often treat examples as independent or density-weighted samples, which can over-represent redundant high-frequency patterns while under-covering geometrically distinct regions of the teacher representation space. We study a volumetric alternative for distillation data selection. Our method, VoluCore, selects compact coresets by maximizing the regularized log-determinant of the Gram matrix formed by normalized teacher features. For any fixed regularization parameter, this objective is monotone submodular, giving a standard greedy $(1-1/e)$-approximation guarantee. To make the criterion practical at LLM scale, we implement greedy selection with efficient Cholesky/Gram-Schmidt rank-one updates after feature extraction. Across math, code, and instruction-following distillation settings, VoluCore reaches high-recovery performance with fewer selected examples than random, uncertainty-based, and distance-based selection baselines, while remaining competitive with gradient-based selection at substantially lower end-to-end cost. Additional spectral, layer, and task-coverage analyses indicate that VoluCore improves representation-space coverage without sacrificing common high-density capabilities. These results support log-determinant coverage of teacher representations as a simple, scalable, and theoretically grounded criterion for constructing compact distillation sets.
Voxel as Token: A New Perspective for Zero-shot Cross-subject Vision Decoding
Yulong Liu ⋅ Ziqiu Huang ⋅ Hua Xu ⋅ Guibo Zhu ⋅ Sirui Han ⋅ Yike Guo
Recent advances in brain decoding have achieved impressive visual reconstruction under subject-specific settings, yet they fail to generalize to unseen subjects without retraining---hindering practical zero-shot cross-subject applications. The core challenges stem from inter-subject variability in voxel dimensionality and misaligned spatial response patterns. While existing methods rely on surface-based alignment or complex feature disentanglement, they often operate as black boxes and overlook the potential of volume-based fMRI data. To address this, we propose \textbf{Voxel as Token (VoxTok)}, a conceptual framework that treats each fMRI voxel as an independent token endowed with a learnable functional embedding. This formulation naturally handles variable-length inputs and subsumes existing adapter-based models as special cases where linear projections approximate voxel functionality. Guided by the hypothesis that large-scale spatial distributions of neural activity are consistent across subjects despite fine-grained variability, we introduce a KNN-based estimator to predict functional embeddings for unseen subjects. This approach achieves state-of-the-art zero-shot retrieval performance and comparable reconstruction quality, while enabling plug-and-play adaptation for existing models (e.g., MindEye2). Integrating multi-scale ROIs further boosts performance, with our best model achieving $40.4\%$ zero-shot image retrieval accuracy, rivaling early supervised methods. Crucially, we find that decoding success correlates strongly with ROI size, indicating that global co-activation patterns drive cross-subject generalization.
VTBench: Disentangled and Human-Aligned Evaluation for Image-Based Virtual Try-on
Donghao Luo ⋅ Yujie Liang ⋅ Caoshuo Li ⋅ Xiaobin Hu ⋅ Zuxuan Wu ⋅ Yanwei Fu
While virtual try-on has achieved significant progress, evaluating these models towards real-world scenarios remains a challenge: unpaired metrics (FID/KID) miss task-specific dimensions such as texture and consistency, and most test sets are limited to indoor scenarios. To address these needs, we introduce the Virtual Try-on Benchmark (VTBench), the first-ever hierarchical try-on benchmark suite that systematically decomposes virtual image try-on into hierarchical, disentangled dimensions, each equipped with tailored test sets and evaluation criteria. VTBench exhibits three key advantages: 1) Multi-Dimensional Evaluation Framework: The benchmark defines three top-level dimensions (General Image Quality, Garment Preservation, Auxiliary Consistency) and six fine-grained dimensions (Distributional Fidelity, Aesthetics, Texture Fidelity, Cross-Category Plausibility, Background Consistency, and Hand Consistency). Granular evaluation metrics and targeted test sets pinpoint model capabilities and limitations across diverse, challenging scenarios. 2) Human Alignment: Human preference annotations are provided for the five image-level dimensions, ensuring the benchmark's alignment with perceptual quality. 3) Valuable Insights: Beyond standard indoor settings, we analyze model performance variations across dimensions and investigate the disparity between indoor and real-world try-on scenarios. To foster the field of virtual try-on towards challenging real-world scenarios, VTBench will be open-sourced, including all datasets and evaluation algorithms.
VVTRec: Radio Interferometric Reconstruction through Visual and Textual Modality Enrichment
Kai Cheng ⋅ Ruoqi Wang ⋅ Qiong Luo
Radio astronomy is an indispensable discipline for studying distant celestial objects. Measurements of wave signals from radio telescopes, called visibility, need to be transformed into images for astronomical observations. These dirty images blend information from real sources and artifacts. Therefore, astronomers usually perform reconstruction before imaging to obtain cleaner images. Existing methods consider only a single modality of sparse visibility data, resulting in images with remaining artifacts and insufficient modeling of correlation. We propose VVTRec, a multimodal radio interferometric data reconstruction method, to enhance visibility information extraction and improve image-domain output quality. Since the information in sparse visibility is inherently limited, we transform it into image-form and text-form features. Therefore, we can leverage vision-language models to process the two derived modalities, providing knowledge banks as a supplement for reconstruction. The sparse visibility is accordingly utilized as the query to perform the integration and selection of multimodal external knowledge. Consequently, VVTRec improves the structural integrity and accuracy of reconstructed images by enriching spatial and semantic information. Our experiments demonstrate that VVTRec effectively enhances imaging results by exploiting multimodal information without introducing excessive computational overhead.
Walking Through 3D Spaces: Spatial Routing for Referring 3D Gaussian Splatting Segmentation
Peilin Ji ⋅ Changshuo Wang ⋅ Yong Zhang ⋅ Shuting He
This paper introduces Walk3D, a novel 3D Gaussian Splatting (3DGS) framework explicitly designed for 3D referring segmentation, particularly excelling at interpreting complex, long-sentence descriptions. Our primary motivation stems from observing that existing 3DGS-based open-vocabulary methods mainly focus on simple, category-level object segmentation. These methods struggle to precisely segment targets described by highly detailed, multi-sentence queries due to weak compositional language understanding and a lack of structural scene awareness. To tackle these dense and complex linguistic constraints, we propose a spatial routing network that decodes these queries into sequential chain-of-thought instructions upon a scene graph. Specifically, we construct a static supernode graph based on the geometric and semantic properties of the 3DGS space, where a spatial Graph Neural Network (GNN) dynamically pathfinds using the decomposed linguistic chain to accurately segment the referred target. Extensive experiments on complex 3D referring segmentation benchmarks demonstrate the superiority of our proposed method in interpreting lengthy, multi-conditional queries. Our code and trained models will be publicly released.
Watermarking Game-Playing Agents in Perfect-Information Extensive-Form Games
Juho Kim ⋅ Fei Fang ⋅ Tuomas Sandholm
Watermarking techniques for large language models (LLMs), which encode hidden information in the output so its source can be verified, have gained significant attention in recent days, thanks to their potential capability to detect accidental or deliberate misuse. Similar challenges involving model misuse also exist in the context of game-playing, such as when detecting the unauthorized use of AI tools in gaming platforms (e.g., cheating in online chess). In this paper, we initiate the study of how game-playing strategies can be watermarked. We show how the KGW watermark for LLMs can be adapted to watermark game-playing agents in perfect-information extensive-form games. The watermark can then be detected using a statistical test. We show that the degradation in the quality of the watermarked strategy profile, quantified by the expected utility, can be bounded, but there is a tradeoff between detectability and quality. In our experiments, we bootstrap the watermarking framework to various chess engines and demonstrate that a) the impact of the watermark on the quality of the strategy is negligible and b) the watermark can be detected with just a handful of games.
WebSpline: Structure-Informed Splines for Real-Time 3D Gaussians from Monocular Videos
Jongmin Park ⋅ Jeonghwan Yun ⋅ Minh-Quan V Bui ⋅ Munchurl Kim
Dynamic scene reconstruction from monocular videos remains highly challenging, as existing methods often struggle to balance global structural coherence and local fine-grained details under limited multi-view cues. To address this challenge, we propose WebSpline, a novel dynamic 3D Gaussian framework that enables structurally coherent and high-fidelity reconstruction from monocular videos with fast rendering. The core of WebSpline is the Structure-Informed Spline (SIS) representation, which models each dynamic Gaussian trajectory using a learnable cubic Hermite spline whose motion is structurally organized with an auxiliary Structural Proxy Graph (SPG). The proposed framework is optimized in two stages: (i) in the first stage, the SPG is initialized from 2D point tracks and refined with temporal rigidity regularization to establish structural coherence for moving objects across the sequence; and (ii) in the second stage, the SIS representation is initialized from the refined SPG and optimized under both spatial and structural neighborhood constraints. At inference, Gaussian motion is obtained solely by evaluating the learned SIS, enabling fast rendering. Extensive experiments on the challenging monocular dynamic scene benchmarks, iPhone and NVIDIA, demonstrate that our WebSpline achieves state-of-the-art rendering quality while rendering over 10x faster than WorldTree, the second-best method on the iPhone dataset.
Sparse autoencoders (SAEs) have emerged as a powerful technique for decomposing language model representations into interpretable features. Current interpretation pipelines infer feature semantics from activation patterns, implicitly assuming every feature has a semantic explanation while overlooking the computational roles features inherit from the SAE training objective. We introduce a weight-based interpretation framework that requires no activation data and grounds claims causally through targeted feature-ablation interventions. Across Gemma-2 and Llama-3.1, three depth-dependent signatures show that SAE features inherit the base model's geometry: semantic features form a U-shape under tied embeddings but bifurcate when untied, attention participation peaks mid-layer, and the two roles couple oppositely for input- vs. output-oriented features. This weight-based view supplies the causal half missing from activation-based interpretability.
What Does the Brain See? Multiview Neural Representation to Demystify the Brain-Visual Alignment
Salini Yadav ⋅ Taveena Lotey ⋅ Pravendra Singh ⋅ Partha P Roy
Zero-shot visual decoding from electroencephalography (EEG) aims to infer visual semantics from non-invasive neural recordings, but remains challenging due to the low signal-to-noise ratio, non-stationarity, and limited spatial resolution of EEG. Existing EEG–vision alignment methods often rely on holistic EEG embeddings, which can obscure the complementary temporal, spectral, and spatial structure underlying visual perception. We introduce a unified multiview EEG representation learning framework for aligning brain responses with visual semantic embeddings. Our method builds an EEG encoder that jointly models three complementary views: input-conditioned state-space temporal dynamics, learnable wavelet-based spectral decomposition for sample-adaptive frequency modeling, and attention-modulated graph learning for structured electrode interactions. The resulting multiview EEG embeddings are fused and aligned with pretrained visual representations in a shared semantic space using contrastive learning with EEG-specific regularization, enabling 200-way zero-shot visual classification. Experiments on THINGS-EEG benchmark show that our method achieves state-of-the-art performance, with 54.8% Top-1 and 85.6% Top-5 accuracy in the within-subject setting and 15.3% Top-1 and 45.4% Top-5 accuracy in the cross-subject setting. We further present the first systematic cross session EEG–image decoding evaluation, achieving 40.8% Top-1 and 78.0% Top-5 accuracy. These results suggest that explicitly modeling multiview neural structure improves both semantic alignment and generalization in EEG-based visual decoding.
"What is Different Between These Datasets?" A Framework for Explaining Data Distribution Shifts
Varun Babbar ⋅ Zhicheng Guo ⋅ Cynthia Rudin
The performance of machine learning models relies heavily on the quality of input data, yet real-world applications often face significant data-related challenges. A common issue arises when curating training data or deploying models: two datasets from the same domain may exhibit differing distributions. While many techniques exist for detecting such distribution shifts, there is a lack of comprehensive methods to explain these differences in a human-understandable way beyond opaque quantitative metrics. To bridge this gap, we propose a versatile framework of interpretable methods for comparing datasets. Using a variety of case studies, we demonstrate the effectiveness of our approach across diverse data modalities—including tabular data, text data, images, time-series signals – in both low and high-dimensional settings. These methods complement existing techniques by providing actionable and interpretable insights to better understand and address distribution shifts.
What Kind of Diffusion Models Do We Need in Online Reinforcement Learning?
Zihao Wu ⋅ Hongyao Tang ⋅ Yi Ma ⋅ YAN ZHENG ⋅ Jianrong Wang ⋅ Jianye Hao
Score matching and flow matching exhibit powerful expressive capabilities in continuous control tasks. However, since the log probability densities of their generated actions cannot be directly accessed, they introduce substantial challenges to the implementation and optimization of Maximum-Entropy Reinforcement Learning (RL). To address this critical limitation, the RL community has developed a wide range of diffusion-based RL algorithms, leading to a widespread perception that the integration of diffusion models and RL has achieved remarkable success. Yet, is this optimistic conclusion truly well-founded? In this paper, we propose a systematic taxonomy for existing mainstream diffusion RL methods that consists of three distinct categories, and we conduct in-depth theoretical analyses of their inherent limitations and core defects.Our analysis reveals that existing methods either collapse the intended stochastic generative policy into action maximization, require costly long-horizon sampling in high-dimensional action spaces, or rely heavily on proposal distributions whose coverage fundamentally limits policy improvement. Building on the comprehensive analysis above, we propose a novel method \textbf{Copuled Flow} . The core insight of our method is to leverage linear ordinary differential equations to maintain fully tractable log-likelihood calculation, while adopting coupled generation mechanisms to effectively capture complex distributions. In the experiments, we evaluate the \textbf{Copuled Flow} on 3 HumanoidBench tasks, 5 MuJoCo tasks, and 7 tasks from the DeepMind Control (DMC) Suite. Empirically, \textbf{Copuled Flow} achieves superior performance on high-dimensional benchmark tasks compared to competitive strong baselines. Specifically, it significantly outperforms existing diffusion-based RL methods by 100\% on the HumanoidBench benchmark.
What Makes Two States the Same? Value-Aware State Matching for Multi-Turn Agent Credit Assignment
Yangyi Fang ⋅ Haolin Shi
Training multi-turn LLM agents with reinforcement learning requires distributing a single episode reward across many interaction steps. Recent group-based methods tackle this credit assignment challenge by finding steps where the agent visits the "same" state across different rollouts and comparing their outcomes. However, determining when two states are "the same" remains an open problem: existing methods compare raw text observations via exact string matching or bag-of-words similarity, both of which ignore whether state differences actually matter for future returns. We show that this is a state abstraction problem with a precise bias-variance tradeoff: matching on value-irrelevant features creates unnecessarily small groups (high variance), while ignoring value-relevant features merges states that should be distinguished (high bias). Guided by this analysis, we propose Structured State Abstraction (SSA), a step-level credit assignment method that decomposes observations into typed fields using environment-provided metadata and learns per-field importance weights from return variance. Rather than hard grouping, SSA uses a kernel-smoothed advantage estimator where all states contribute to each step's baseline proportionally to their structured similarity, yielding lower MSE than threshold-based grouping. Experiments on ALFWorld and WebShop show consistent improvements over GiGPO and HGPO across both benchmarks and model scales; notably, kernel smoothing eliminates the singleton problem that causes existing methods to lose step-level credit for 30-36% of steps. Our implementation is provided in the supplementary materials and will be open-sourced upon acceptance.
What Should Remain After Forgetting? Rethinking LLM Unlearning as Predictive Posterior Correction
Jingyue Cong ⋅ Andy Song ⋅ Alexis Horde-Vo ⋅ Kai Wei ⋅ Li Shen ⋅ Yameng Peng ⋅ Feng Liu ⋅ Estrid He
LLM unlearning is often implemented by assigning surrogate targets, such as refusal, uniformity, or likelihood suppression, to forget-related prompts. We argue that this surrogate assignment view is misaligned with the counterfactual goal of unlearning: recovering the retain-only predictive posterior. This mismatch is especially problematic in mixed-evidence regimes, where a response may be associated with forgotten evidence while still being partially supported by retained evidence. We propose EASE: Evidence Attribution and Subtraction Estimator, a logit-space posterior-correction method that estimates the forget-induced component of the deployed model's predictive support and subtracts it at inference time. EASE uses lightweight deletion and compensation assistants to remove forget-neighbour evidence while restoring nearby retain-supported evidence. We formalize the suppression-scope dilemma and provide guarantees showing when posterior correction recovers the retain-only posterior. Experiments on TOFU and MUSE show that EASE improves forgetting--retention trade-offs over strong baselines, achieving the best aggregate score in 5/6 main TOFU settings, reducing mixed-query semantic leakage by 22.2%, and matching or exceeding full-retain baselines with 9 times fewer retain examples. Our code is available at: https://anonymous.4open.science/r/EASE-9675.
What Time Is It? How Data Geometry Makes Time Conditioning Optional for Flow Matching
Alec Helbling ⋅ Sebastian Gutierrez Hernandez ⋅ Benjamin Hoover ⋅ Duen Horng Chau ⋅ Parikshit Ram
Recent work has shown that flow matching models can be trained without explicit time conditioning, challenging the standard view that the interpolation time is needed to disambiguate velocity targets. But why should a time-blind model work at all? Decomposing the time-blind flow matching loss, we identify two sources of irreducible error: a *coupling variance*, which arises from ambiguous velocity targets induced by how noise and data points are paired, and the *time-blindness gap*, which is the additional error caused by ignoring time. This gap shows that time-blind training is strictly harder than conventional training, reinforcing the puzzle that time-blind models work so well in practice. We resolve this tension by showing that the geometry of high-dimensional data makes time identifiable directly from noisy observations. When data concentrates near a $k$-dimensional subspace, time can be recovered from the statistical structure of noisy interpolants in directions orthogonal to the data; under a spiked-covariance model, this yields a closed-form estimator that recovers $t$ from a single observation $z$ at rate $O(1/\sqrt{d-k})$ for ambient dimension $d$. As a consequence, we prove that the time-blindness gap is asymptotically negligible relative to the coupling variance. We empirically demonstrate our identifiability result on real-world data and show that changing the coupling has a much larger effect on loss and sample quality than removing time conditioning across CIFAR-10, CelebA-HQ, and FFHQ. These results explain why time-blind flow matching works and show that the main practical lever is the choice of coupling, not explicit time conditioning.
When and Why Adversarial Training Improves PINNs: A Neural Tangent Kernel Perspective
Yuandong Cao ⋅ Chi C So ⋅ Jun-Min Wang ⋅ He Wang
Physics-informed neural networks (PINNs) are powerful surrogates for differential equations but are notoriously difficult to train due to spectral bias, stiffness, and poor accuracy on high-frequency or multiscale solutions. Adversarial training based on generative adversarial networks (GANs) has recently gained surprisingly strong empirical results in improving training, but the underlying mechanisms remain elusive. To this end, we propose a new analysis framework for adversarially trained PINNs, based on the key observation of how the discriminator in GANs can influence the training dynamics of PINNs. The framework first provides a much needed theoretical grounding to why and when adversarial training is effective in PINNs, then presents a unified analysis of GANs variants in such training, and finally leads to a new, practical, efficient training algorithm for PINNs. Empirical results demonstrate that our method can significantly reduce the pathology of PINNs training, thereby providing better models with superior performances, often several magnitudes more accurate than alternative methods.
When Are Multimodal Predictions Biologically Supported? A Diagnostic Evaluation Framework
Dylan Steiner ⋅ Gustavo Arango ⋅ Gerald J Sun ⋅ Etai Jacob
Multimodal models in oncology can produce accurate predictions, but accurate prediction does not reveal whether the model has learned biology that is shared across modalities, biology confined to one modality, or spurious correlations that reflect confounders rather than genuine biology. We introduce DECAT, a model-agnostic post-hoc evaluation framework that classifies multimodal representations into four diagnostic scenarios for a given task and modality, using five null-referenced metrics and a rule-based decision procedure. The framework operates on learned representations, requires no knowledge of which specific confounder is present, and returns indeterminate when the evidence is insufficient. We evaluate multimodal architecture properties using DECAT on synthetic data across four multimodal model classes spanning over 2,500 trained representations, and demonstrate its operational use on 8,979 TCGA patients with both multimodal model embeddings and with five pretrained pathology foundation models. Entangled models (e.g., CLIP) achieve near-perfect shared biology detection but falsely claim shared biology in the majority of cases where it is absent on real foundation model embeddings. This false claim rate increases with confound strength so that larger cohorts and stronger representations produce more confident but still incorrect diagnoses. Applied to both multimodal TCGA embeddings and five pathology foundation models without paired RNA, DECAT detects confounding invisible to AUROC without requiring the confounder labels, as confirmed by post-hoc stratification.
When Confidence Rises Too Early: Detecting Shortcut Reasoning via Premature Answer Commitment
Zhaohan Zhang ⋅ Junjie Liu ⋅ Chengzhengxu Li ⋅ Chen Shen ⋅ Xiaoming Liu ⋅ Chao Shen ⋅ Jieping Ye ⋅ Ziquan Liu ⋅ Ioannis Patras
The reasoning trajectory of the Large Language Model (LLM) is often regarded as the verbalized description of the internal thinking. However, the unfaithfulness of the reasoning process introduces the risk of shortcut reasoning, where the model fails to reason step by step but instead relies on discovered shortcuts to reach the final answer, while post-rationalizing this decision through a seemingly coherent verbalized reasoning process. This shortcut reasoning is difficult to detect, as existing monitors and verifiers mainly inspect textual reasoning traces or final outcomes, failing to capture how the model’s answer belief forms during generation. To figure out the intrinsic pattern in shortcut reasoning, we propose ConfLens, a framework that tracks how a model's confidence in its final answer evolves throughout the reasoning process. Across three shortcut reasoning settings, we find that shortcut samples often exhibit premature confidence, characterized by high confidence in the final answer at early reasoning stages. However, reliably detecting this pattern remains challenging, as existing confidence estimation methods are limited in generalizability, reliability, and efficiency. To address this, we introduce the Distributional Answer Commitment Score (DACS), a distributional confidence estimation method that instantiates ConfLens for effective shortcut reasoning detection. DACS estimates the entropy of the model's probability distribution over answer commitment during reasoning, thereby capturing how concentrated the model's answer belief is at each reasoning step without access to the ground-truth or task-specific verifier. We further convert the detection results of ConfLens into interpretable signals to mitigate reward models’ preference for shortcut reasoning responses. Experiments on math and code reasoning tasks show that ConfLens instantiated with DACS improves shortcut reasoning detection by over 4.3% in F1 compared with strong baselines, while reducing the gap between faithfulness and correctness in reward model preferences.
When Does Graph Retrieval Become Answer-Supporting Evidence? A Diagnostic Audit of GraphRAG
Chenghua Duan ⋅ Zhuo Wang ⋅ boheng liu ⋅ Ziyu Li ⋅ Qing Li ⋅ Xiuxing Li ⋅ zhong zhang ⋅ Xia Wu
Graph retrieval-augmented generation (GraphRAG) is motivated by the idea that relational structure can help a language model find evidence that flat retrieval may miss. Yet retrieval scores and final-answer scores often compress a multi-stage pipeline into a single method row. We study the missing middle of this pipeline through an evidence-to-answer audit that follows each system from candidate availability, through evidence selection and context construction, to generated answers and cost. Across GraphRAG-Benchmark Medical and Novel tasks, we find a consistent conversion gap. Strong reranking, learned evidence scoring, and graph-based retrieval can improve evidence-side signals, but these gains do not yield significant answer-correctness improvements under the same generator and evaluator. Matched attribution controls show that some apparent graph-side gains are driven by reranking rather than graph expansion alone. Gold-evidence and bridge-aware interventions further show that evidence completeness is insufficient on its own; answerability also depends on salience, organization, and support in the generator-facing context. We therefore argue that GraphRAG evaluation should report the full evidence-to-answer conversion chain, including candidate source, selected context, generated answer, metrics, and cost, rather than relying on retrieval or answer leaderboards alone.
When Does Non-Uniform Replay Matter in Reinforcement Learning?
Michal Korniak ⋅ Mikołaj Czarnecki ⋅ Yarden As ⋅ Piotr Miłoś ⋅ Pieter Abbeel ⋅ Michal Nauman
Modern off-policy reinforcement learning algorithms often rely on simple uniform replay sampling and it remains unclear when and why non-uniform replay improves over this strong baseline. Across diverse RL settings, we show that the effectiveness of non-uniform replay is governed by three factors: replay volume, the number of replayed transitions per environment step; expected recency, how recent sampled transitions are; and the entropy of the replay sampling distribution. Our main contribution is clarifying when non-uniform replay is beneficial and providing practical guidance for replay design in modern off-policy RL. Namely, we find that non-uniform replay is most beneficial when replay volume is low, and that high-entropy sampling is important even at comparable expected recency. Motivated by these findings, we adopt a simple Truncated Geometric replay that biases sampling toward recent experience while preserving high entropy and incurring negligible computational overhead. Across large-scale parallel simulation, single-task, and multi-task settings, including three modern algorithms evaluated on five RL benchmark suites, this replay sampling strategy improves sample efficiency in low-volume regimes while remaining competitive when replay volume is high.
When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon
Huaqing Zhang ⋅ Jingchu Gai ⋅ Juno Kim ⋅ Bingbin Liu ⋅ Andrej Risteski
Online imitation learning (IL), particularly on-policy distillation, has emerged as a strong LLM post-training approach, often outperforming offline supervised fine-tuning (SFT). Yet a principled understanding of when and why online helps remains unclear. In this work, we challenge the view that error accumulation is the main source of online IL's advantage, and instead show that the benefits of online interaction depend critically on whether the setting is realizable, i.e., whether the student policy class can represent the expert policy. Under realizability, we empirically find that offline IL already matches expert performance. In contrast, in non-realizable (misspecified) settings, we prove that offline IL encounters an information-theoretic bottleneck even when horizon $H=1$, and propose a structural characterization of misspecification relative to the reward, under which online IL provably achieves high performance despite large expert-student discrepancy.
When Do Learned State Representations Break Sensitivity Analysis?
Siyu WANG ⋅ Xiaocong Chen ⋅ Quan Z Sheng ⋅ Lina Yao
Causal sensitivity analysis bounds the policy value of off-policy evaluation under unobserved confounding. In offline reinforcement learning, however, sensitivity bounds are typically applied not to raw states but to learned representations. We show that this two-stage pipeline can silently break the coverage guarantee: state aggregation under hidden confounding amplifies the effective sensitivity parameter in the latent space, so bounds computed at the nominal confounding level may exclude the true policy value. We characterise the amplification mechanism, give a sufficient condition for preservation, and prove a continuity result showing the safety condition is not an isolated point: amplification grows continuously with the conditional-independence violation. Neither unsupervised nor standard task-aware representation objectives target the right quantity. We propose Sensitivity-Preserving Representation Learning (SPRL), which augments latent dynamics prediction with a kernel-weighted within-cell propensity homogeneity penalty, an observable surrogate for the unobservable safety condition, and show across synthetic and semi-synthetic benchmarks that SPRL is the only method tested that simultaneously preserves coverage in regimes where unsupervised representations break it and prevents amplification in regimes where task-aware baselines fail.
When do Prophets Profit in Prediction Markets?
Anri Gu ⋅ Nicole Kagan ⋅ Alec Sun ⋅ Jibang Wu ⋅ Haifeng Xu
Prediction markets aggregate dispersed beliefs into prices that act as probabilistic forecasts of uncertain events. Classical theory establishes a clean equivalence between forecasting accuracy and trading profit, but only for the specific automated market maker (AMM) design. However, the largest exchanges today are based on central limit order books in which informed forecasters routinely lose money while uninformed strategies can profit on simple heuristics. We resolve this discrepancy by establishing a formal equivalence between predictive accuracy and profitability. For any strictly proper scoring rule $S$, we exhibit a ``proper'' betting strategy that depends only on the forecaster's prediction $\mathbf{p}$ and the market price $\mathbf{q}$ and earns positive expected profit whenever $\mathbf{p}$ outperforms $\mathbf{q}$ under $S$ and the market has sufficient liquidity. The proof rests on a decomposition of expected profit that strictly generalizes the classical AMM guarantee and also explains how strategies can profit without an accuracy edge. Empirically, across thousands of forecasts from AI models, proper betting is the only strategy that reliably converts accuracy into profit, and we further identify systematic forecasting personas and show how the optimal proper strategy varies across them. A month-long live deployment achieves $+80.33$% return on investment with a Sharpe ratio of $3.35$.
When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search
John Robertson ⋅ Jianing Zhu ⋅ Haris Vikalo ⋅ Zhangyang "Atlas" Wang
Activation steering offers a lightweight way to control large language models without retraining, but its effectiveness varies sharply across concepts. Prior work often interprets this variability as evidence that many concepts are not well captured by a single steering direction. We argue instead that much of this variability reflects search difficulty: a useful rank-1 intervention often exists, but finding it can be expensive. We formalize rank-1 steering as a budget-constrained optimization problem over intervention layer and coefficient. Across the concepts and model families, prompt-boundary directional alignment predicts where effective interventions are likely to occur, enabling geometry-guided search that reaches high utility with substantially fewer evaluations, reducing the trials needed to recover 95\% of best-found utility by 39.8\% on average across three model families. To explain why some concepts remain expensive even under better search, we introduce concept granularity, a measure of directional heterogeneity across contrastive contexts. Granularity distinguishes concepts whose difference vectors share a stable global direction from those where prompts agree locally within each input but the utility-maximizing direction rotates systematically across inputs. Higher granularity is associated with both slower convergence and lower best-found steering performance (Pearson $r = 0.44$ with trials-to-95\%, $p < 0.001$, and $r=-0.46$ with best-found utility, $p < 0.001$). These observations suggest a practical workflow rather than a single universal vector-construction rule. We therefore present GRACE, a Granularity- and Representation-Aware Concept Engineering framework that uses activation geometry to diagnose the dominant source of steering difficulty, choose the appropriate remedy, and allocate optimization effort more efficiently. Our results shift the frame of activation steering from "when does rank-1 fail?" to "when is rank-1 cheap and stable?", and turn activation geometry from a descriptive tool into an actionable prior for LLM control.
When Stored Evidence Stops Being Usable: Scale-Conditioned Evaluation of Agent Memory
Jiaqi Shao ⋅ Yiyi Lu ⋅ Yunzhen Zhang ⋅ Bing Luo
Memory-agent evaluations report fixed-snapshot accuracy or retrieval quality, but these scores do not show whether evidence remains usable as irrelevant sessions (sessions not annotated as task-relevant evidence for the query) accumulate. We present a scale-conditioned evaluation protocol for agent memory under evidence-preserving growth: for each query, task evidence is held fixed while irrelevant sessions are added. The protocol logs agent--memory trajectories and reports four diagnostics: budget-compliant reliability, tail memory-call burden, failure-regime decomposition, and the usable-scale boundary where reliability falls below the target. Applied to LongMemEval and LoCoMo across flat, planar, and hierarchical memory interfaces, the protocol shows reliability loss is not a single phenomenon. On LongMemEval, HippoRAG stays within the two-call budget but loses 16--20 percentage points in budget-compliant reliability as irrelevant sessions are added; LiCoMemory's observed failures depend strongly on the agent, with Qwen3-8B exceeding the budget while Qwen3-32B and Qwen3-235B remain reliable in the tested range. The result supports a framework for making scalable-memory claims conditional on agent, interface, scale range, and interaction budget.
As AI development increasingly involves models adapted to downstream settings, a new algorithmic challenge for developers has surfaced: deciding when to adopt model updates. Currently, the only reliable approach is to retrain on new upstream models and then evaluate extensively, which can be prohibitively expensive. We introduce a framework for assessing when to propagate newly released model versions to downstream applications that does not require access to upstream data or retraining a priori. Our framework incorporates geometric similarity and information-theoretic sufficiency to determine when an upstream update substantially shifts the representational basis recruited for a downstream task. This enables targeted update adoption, and supports emerging norms for coordinating AI supply chain infrastructure.
When Trackers Fail: VLM-Guided Verification and Recovery for Robust Video Object Segmentation
Valay Mahesh Bundele ⋅ Susmit Agrawal ⋅ Mehran Hosseinzadeh ⋅ The Nam Nguyen ⋅ Hendrik PA Lensch
Recent segmentation-based trackers built on foundation models such as Segment Anything Model (SAM) achieve strong performance through memory-based mask propagation. However, despite their strong generalization ability, these methods remain brittle under occlusion, reappearance, and distractor scenarios, where errors accumulate and lead to identity drift. Existing approaches attempt to improve robustness through better memory design and temporal modeling, yet they implicitly assume that propagated predictions remain reliable over time, lacking a mechanism to identify and correct failures. To address this limitation, we propose a failure-aware tracking framework that augments segmentation-based propagation with explicit reasoning and selective recovery. Our key idea is to decouple tracking into two complementary processes: verifying prediction reliability and recovering the target when failures occur. Specifically, we introduce a dual-agent architecture consisting of a vision-language verification agent and a correction agent. The verification agent reasons over spatio-temporal context to assess whether current predictions remain consistent with the target object, enabling the detection of identity switches, missed recoveries, and distractor-induced failures. When unreliable predictions are identified, the correction agent is selectively activated to recover the target using temporal cues. To further improve robustness under ambiguous scenarios, we introduce targeted training perturbations that simulate identity switches and and distractor-induced drift during training. By explicitly modeling failure detection and recovery within tracking loop, our framework transforms tracking from a purely propagation-based process into a self-correcting system. Experimental results on LVOS-v2 and SecVOS show that the proposed framework improves SAM3 by 1.2 J&F and 2.7 J&F respectively, demonstrating the value of explicit reasoning and selective recovery in segmentation-based tracking.
Where Should Society Draw the Line? A Social Choice Approach to Collective Consent
Chris Dong ⋅ Sonja Kraiczy ⋅ Rohit Prashanth Vasishta ⋅ Markus Brill ⋅ Wesley H Holliday ⋅ Niclas Boehmer
Society constantly has to determine the boundaries of what it deems acceptable, from legislative decisions to the guardrails governing autonomous systems. We initiate the axiomatic study of collective consent: given individuals' attitudes toward options, which options should receive societal consent? We organize our analysis around three principles: sufficient support, minority protection, and dominance by decisively better options. Each captures a distinct reason for withholding societal consent from an option. For each principle, we develop a corresponding solution concept that transparently implements the principle and is canonical in a mathematically precise sense. For example, for minority protection, the resulting concept is a consent-adapted version of Moulin's Proportional Veto Core. Balancing multiple principles simultaneously is more challenging. To address this, we develop a game-theoretic characterization of our veto core that naturally gives rise to a family of related concepts. From this family, we identify the Approval-Weighted Veto Core as particularly desirable. By making minorities' blocking power depend on the approval support of the options being challenged, it smoothly interpolates between proportional minority protection and majority support. Experiments on five datasets spanning high-stakes decision-making (such as political elections, ethical AI evaluations, and moral decision-making) show that solution concepts violating a principle in theory also violate it empirically.
Where to Look Is Not How to Fix: Pre-Denoising Diagnostics and Modality-Dependent Control in Diffusion Composition
Fangzheng Wu ⋅ Brian Summa
Text-to-image diffusion models fail on semantically rare attribute--object compositions, but it is unclear whether this brittleness arises during denoising or is already present in conditioning representations. We study this question at three levels of granularity. First, using a controlled stress protocol across SD1.5, SDXL, and SD3, we show that a Compositional Stress Index derived from text-side non-factorization residuals separates anchor from stress prompts with perfect discrimination (AUC${=}1.0$) cross-model, establishing a robust pre-denoising risk axis. Second, embedding-level low-rank repair reduces representation residuals without any positive semantic improvement, while a downstream cross-attention intervention succeeds on the same subset---a site-sensitive boundary between diagnosis and control. Third, we decompose this boundary at block-group resolution under three intervention modalities. Diagnostic accessibility peaks at the deep encoder (D2) across all conditions, but token-level intervention of either sign produces no measurable improvement at any block. Only broad cross-attention ablation moves semantics at D2, while decoder blocks respond to selective boost and collapse under ablation. Together, these results establish that compositional defects are diagnosable before denoising, but \textbf{where to look is not how to fix}: the diagnostic site and the control site require different intervention modalities, and even at the correct site, the wrong modality is futile.
Oscillations and synchronization are widely believed to play a fundamental role in representation and computation. However, existing machine learning approaches based on synchronization dynamics have largely been confined to specialized settings such as object discovery, with limited evidence of scalability to standard vision benchmarks or logic reasoning tasks. We propose the \emph{Winfree Oscillatory Neural Network} (\emph{WONN}), a dynamical neural architecture based on generalized Winfree dynamics. \emph{WONN} evolves representations on the torus $(S^1)^d$ through structured oscillatory interactions, combining phase-based inductive biases with flexible and hierarchical interaction mechanisms instantiated as either fixed trigonometric mappings or learnable neural networks. We evaluate \emph{WONN} on image recognition and complex reasoning tasks, including CIFAR, ImageNet, Maze-hard, and Sudoku. Across these domains, \emph{WONN} achieves competitive or superior performance with strong parameter efficiency. \textbf{In particular, \emph{WONN} is, to our knowledge, the first synchronization-based oscillatory architecture to scale competitively to ImageNet-1K.} Furthermore, on Maze-hard, \emph{WONN} achieves \textbf{80.1\% accuracy using only 1\% of the parameters of prior state-of-the-art models}. These results suggest that structured oscillatory dynamics provide a scalable and parameter-efficient alternative to conventional neural architectures.
WiREBench: Evaluating AI Agents' Capabilities in Reverse-Engineering Black-Box Applications in the Real World
Robin Rheem ⋅ Alexandra Maxwell ⋅ Mona Wang ⋅ Benjamin Mixon-Baca ⋅ Tianneng Shi ⋅ Jeffrey Knockel ⋅ Ali Aldakheel ⋅ Basel Alomair ⋅ Prateek Mittal ⋅ Dawn Song
As frontier AI technologies dramatically improve their abilities to perform cybersecurity and reverse-engineering tasks, it is critical to rigorously and realistically track their capabilities in these domains. Increased agent capabilities in reverse-engineering will scale black-box vulnerability discovery, which will have dramatic implications on software security at large. In this work, we present WiREBench, a unique benchmark derived from popular real-world mobile applications, to measure agents' end-to-end vulnerability discovery and reverse-engineering capabilities. This benchmark, containing 22 protocols and 62 subtasks, evaluates agents abilities to identify and exploit historical vulnerabilities in binaries' proprietary network cryptography, via black-box reverse-engineering. To resolve these tasks, agents must chain together static binary analysis, understanding of applied cryptography, and complex tool use via dynamic instrumentation of mobile applications. Beyond benchmarking, we ran WiREBench against real-world applications, identifying and confirming 65 protocol vulnerabilities in applications used by hundreds of millions of users. These results demonstrate WiREBench's utility not just as a benchmark to assess current agent capabilities, but also as a pipeline for auditing cryptographic protocols and reverse-engineering in the real world.
WirelessMathBench-XL: A Contamination-Audited Benchmark for Wireless Mathematical Reasoning
Xin Li ⋅ Mengbing LIU ⋅ Yiyang Zhu ⋅ WENHE ZHANG ⋅ LI WEI ⋅ Jiancheng An ⋅ Chau Yuen
**WirelessMathBench-XL** is a $4{,}027$-problem benchmark for wireless mathematical reasoning, built from $836$ retained arXiv papers across $20$ wireless subfields. Benchmarks built from arXiv papers risk overlap with the same arXiv text used in LLM pretraining, yet few releases include the provenance and metadata needed to audit this overlap. WirelessMathBench-XL builds contamination auditing into the released artifact through a **reverse-probe $13$-gram audit**: it indexes benchmark prompts, streams public pretraining corpora, and records per-problem prompt-surface lexical overlap with benchmark-scale memory. Against $12.6$\,B streamed $13$-grams from RedPajama-arXiv, the audit identifies a strict zero-hit subset $\mathcal{S}_0$ covering $3{,}853$ problems ($95.7\%$). Filtering to $\mathcal{S}_0$ changes accuracy by less than $1$\,pp for every evaluated model; frontier-model calibration rows form a single high-accuracy cluster between $86.5\%$ and $91.3\%$, not a resolved rank order. Thus, on this fixed prompt-surface lexical audit channel, reported scores are not measurably driven by exact recall of detected RedPajama-arXiv prompt text. The audit does not cover paraphrase, target-answer, post-training, or closed-corpus exposure. The release includes source-paper identifiers, verifier-facing ground truths, contamination metadata, cleaned subsets, a paper-disjoint sensitivity view, evaluation traces, paired-bootstrap scripts, training recipes, Croissant metadata, and a Datasheet for Datasets, so users can rerun, tighten, or replace the cleaned view.
Working with AI: Measuring the Applicability of Generative AI to Occupations
Kiran Tomlinson ⋅ Sonia Jaffe ⋅ Will Wang ⋅ Scott Counts ⋅ Siddharth Suri
With generative AI emerging as a general-purpose technology, understanding its economic effects is among society’s most pressing questions. Existing studies of AI impact have largely relied on predictions of AI capabilities or focused narrowly on individual firms. Drawing instead on real-world AI usage, we analyze a dataset of 200k anonymized conversations with Microsoft Bing Copilot to measure AI applicability to occupations. We use an LLM-based pipeline to classify the O*NET work activities assisted or performed by AI in each conversation. We find that the most common and successful AI-assisted work activities involve information work---the creation, processing, and communication of information. At the occupation level, we find widespread AI applicability cutting across sectors, as most occupations have information work components. Our methodology also allows us to predict which occupations are more likely to delegate tasks to AI and which are more likely to use AI to assist existing workflows. Finally, we see an increase in occupation-level AI applicability over our 9-month data period, driven primarily by people using AI for more varied tasks.
WorldComposer: Generative High-Fidelity Simulation with Digital Cousins for Generalizable Robot Learning and Evaluation
Jasper Lu ⋅ Zhenhao Shen ⋅ Yuanfei Wang ⋅ Shugao Liu ⋅ Shengqiang Xu ⋅ Shawn Xie ⋅ Jingkai Xu ⋅ Feng Jiang ⋅ Jade Yang ⋅ Chen Xie ⋅ Ruihai Wu
Learning robust robot policies in real-world environments requires diverse data augmentation, yet scaling real-world data collection is costly due to the need for acquiring physical assets and reconfiguring environments. Therefore, augmenting real-world scenes into simulation has become a practical augmentation for efficient learning and evaluation. We present WorldComposer, an automated generative Real2Sim2Real framework that maps real-world panoramas into high-fidelity, task-ready simulation environments and further synthesizes diverse Digital Cousins through scene and object variations. Combined with high-quality physics engines and realistic assets, WorldComposer supports interactive manipulation across rigid, articulated, and deformable objects. Additionally, we incorporate multi-room stitching to construct consistent large-scale environments for long-horizon tasks. Experiments demonstrate a strong sim-to-real correlation validating our platform's fidelity, and show that extensively scaling up data generation leads to significantly better generalization to unseen scene and object variations, demonstrating the effectiveness of Digital Cousins for generalizable robot learning and evaluation.
WorldPrism: 3D Consistency for Video World Models via Bidirectional Cross-Space Verification
Hengyu Liu ⋅ Jiahao Lu ⋅ Ruijie Zhu ⋅ Sixiao Zheng ⋅ Yuan Liu ⋅ Wenbo HU ⋅ Ying Shan
Video world models aim to generate temporally coherent visual observations conditioned on actions, but often suffer from 3D geometric inconsistencies such as spatial drift, object deformation, and perspective violations, which limit their reliability in downstream tasks. We propose WorldPrism, a post-training framework that improves 3D consistency via reinforcement learning without modifying the base model architecture. The central challenge lies in designing a reliable reward signal. To address this, we introduce Bidirectional Geometry-Flow Verification, a bidirectional cross-space mechanism that evaluates consistency through two complementary paths: (i) lifting 2D optical flow into 3D to measure position-space agreement, and (ii) projecting reconstructed 3D geometry to match the point correspondence. By leveraging these two independent and complementary signals, this design mitigates the blind spots of unilateral 2D evaluation. We further incorporate hierarchical temporal evaluation to capture both local and long-range geometric consistency. Experiments on diverse datasets demonstrate that WorldPrism achieves state-of-the-art 3D consistency while preserving visual quality and generative diversity.
WorldSR: Harnessing World Knowledge Search for Grounded Image Super-Resolution
Junyuan Deng ⋅ Xinyi Wu ⋅ Zhenyao Wu ⋅ Song Wang ⋅ Yu Wang ⋅ Congchao Zhu
Generative-prior-based super-resolution (SR) methods leverage implicit knowledge learned from large-scale models to enhance low-resolution inputs. However, such priors are inherently stochastic and optimized for diverse generation, conflicting with the deterministic, high-fidelity reconstruction required for SR. It is also tied to the base model, costly to update, and prone to hallucinations as real-world data evolves. To address these limitations, we introduce WorldSR, a novel framework that augments SR with explicit, entity-aligned visual world knowledge. WorldSR automatically decomposes the input into semantic entities and performs entity-wise retrieval of structurally and semantically aligned images from a large-scale external corpus. For efficiency and reliability, we incorporate a memory mechanism to reuse previously retrieved results and a multi-step filtering process to refine candidate references. The selected references as world knowledge are then integrated through a novel multi-reference super-resolution model, which provides fine-grained constraints for detail synthesis and structural recovery. Experiments on real-world SR benchmarks show that WorldSR consistently improves structural fidelity and visual quality, outperforming existing generative super-resolution methods. Our results suggest that entity-level grounding in external visual world knowledge provides a simple and effective complement to model-bound generative priors.
XDecomposer: Learning Prior-Free Set Decomposition for Multiphase X-ray Diffraction
Hanyu Gao ⋅ Bin Cao ⋅ YUNYUE SU ⋅ Tong-yi Zhang ⋅ Qiang Liu
Multiphase powder X-ray diffraction (PXRD) analysis remains a fundamental bottleneck in structure identification, as real-world synthesis often produces complex mixtures whose constituent phases (components) cannot be reliably disentangled. While recent advances in representation-based crystal retrieval and generation suggest the possibility of inferring structures directly from PXRD, existing approaches largely assume single-phase inputs and break down in multiphase settings. Here, we present XDecomposer, a prior-free framework for joint decomposition and identification of multiphase XRD patterns without requiring candidate phase lists, structural templates, or prior knowledge of phase number. We formulate multiphase diffraction analysis as a set prediction problem, where the model infers an unordered set of phase-resolved components, their mixture proportions, and corresponding structural representations within a unified architecture. A phase-query-driven decomposition mechanism, together with diffraction-consistent physical reconstruction, enables accurate source separation while preserving crystallographic fidelity. Extensive experiments on both simulated and experimental datasets show that XDecomposer substantially improves reconstruction accuracy and phase identification across diverse chemical systems, while maintaining strong generalization to unseen mixtures. These results provide a practical route toward data-driven, source-resolved multiphase XRD analysis and reduce long-standing dependence on prior-guided iteratively phase matching. The code is openly available at \url{https://anonymous.4open.science/r/XDecomposer-FE0B}
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor’s Internal States
Yunho Choi ⋅ Jongwon Lim ⋅ woojin Ahn ⋅ Minjae Oh ⋅ Jeonghoon Shim ⋅ Yohan Jo
Reinforcement learning with verifiable rewards (RLVR) for Large Reasoning Models hinges on baseline estimation for variance reduction, but existing approaches pay a heavy price: PPO requires a policy-model scale critic, while GRPO needs multiple rollouts per prompt to keep its empirical group mean stable. We introduce $\textbf{POISE}$ ($\underline{P}$olicy $\underline{O}$ptimization with $\underline{I}$nternal $\underline{S}$tate Value $\underline{E}$stimation), which obtains a baseline at negligible cost by using the policy model's internal signals already computed during the policy forward pass. A lightweight probe predicts the expected verifiable reward from the hidden states of the prompt and generated trajectory, as well as token-entropy statistics, and is trained online alongside the policy. To preserve gradient unbiasedness despite using trajectory-conditioned features, we introduce a cross-rollout construction that predicts each rollout's value from an independent rollout's internal states. Because POISE estimates prompt value using only a single rollout, it enables higher prompt diversity for a fixed compute budget during training. This reduces gradient variance for more stable learning and also eliminates the compute overhead of sampling costs for detecting zero-advantage prompts. On Qwen3-4B and DeepSeek-R1-Distill-Qwen-1.5B across math reasoning benchmarks, POISE matches DAPO while requiring less compute. Moreover, its value estimator shows similar performance to a separate LLM-scale value model and generalizes to various verifiable tasks. By leveraging the model's own internal representations, POISE enables more stable and efficient policy optimization.
Zeroth-Order Stackelberg Control in Combinatorial Congestion Games
Saeed Masiha ⋅ Sepehr Elahi ⋅ Negar Kiyavash ⋅ Patrick Thiran
We study Stackelberg (leader--follower) control of network parameters (tolls, capacities, incentives) in combinatorial congestion games, where selfish users choose \emph{discrete} routes (or other combinatorial strategies) and settle at a congestion equilibrium. The leader minimizes a system-level objective (e.g., total travel time) evaluated at equilibrium, but this objective can be nonsmooth because the set of used strategies can change abruptly. We propose \textsc{Zeroth-order Stackelberg} (\textsc{ZOS}), which couples a projection-free Frank--Wolfe equilibrium solver with a zeroth-order outer update, avoiding differentiation through equilibria. We prove convergence to generalized Goldstein stationary points of the true equilibrium objective, with explicit dependence on the equilibrium approximation error, and analyze subsampled oracles: if the sampled candidate set contains an LMO minimizer with probability $\kappa_m$, then the Frank--Wolfe error decays as $\mathcal{O}(1/(\kappa_m T))$. We also propose stratified sampling as a practical way to avoid a vanishing $\kappa_m$ when LMO minimizers concentrate in strata defined by simple features such as path length. In public road-network experiments covering large strategy spaces requiring subsampled oracles, \textsc{ZOS} reaches small follower-equilibrium gaps and final social costs comparable to differentiation-based methods, while reducing runtime per outer iteration by $20$--$1000\times$ and using much less peak memory.
ZO-F2: Low-variance Fisher preconditioner via bilinear estimation for zeroth-order optimization
Hiroshi Sawada ⋅ Yuya Hikima ⋅ Kenta Niwa
We propose a novel zeroth-order (ZO) optimization method, ZO-F2, where Fisher information matrix (FIM)-based preconditioners are estimated in a low-variance manner via bilinear forms. Using FIM rather than Hessian allows for a larger number of estimator samples by pairwise combinations of perturbations. We further propose a numerically stable whitening-based update rule for full-matrix preconditioners, even when the FIM estimate is ill-conditioned. We provide theoretical analyses on the variance, the whitening residual, and convergence under standard and reasonable assumptions. Experimental results on black-box training of diverse models from scratch support our proposals and demonstrate the superiority of ZO-F2 over existing methods in terms of accuracy and computational efficiency.
μLM: Rethinking Sub-100M Language Models through Memory-First Design
Zijie Chen ⋅ Guiyun Fan ⋅ Haiming Jin
Ubiquitous IoT devices, such as companion toys, wearables, and micro-robots, are increasingly embracing language models, yet current cloud-based deployments suffer from latency, privacy, and weak-network reliability issues. These resource-constrained devices typically cap deployable model size below 100M parameters, while existing sub-100M models fall short in basic language and commonsense capabilities. In this extreme regime, we argue that the memory bottleneck supersedes the traditional compute bottleneck as the binding constraint, calling for memory-first design that maximizes model capability under a fixed parameter budget. Notably, the stringent budget binds both what the model stores and what it can afford to learn, demanding architecture-data co-design. In this paper, we propose $\mu$LM, a memory-first sub-100M language model framework instantiating this co-design. For architecture, we present DMLA, which preserves latent attention's compact cache while relieving its feature-squeezing bottleneck under aggressive compression via explicit content-position latent decoupling. For data, we present $\mu$Mix, a capacity-aware multi-stage curriculum tailored to sub-100M models, prioritizing high-quality natural language and commonsense data aligned with embedded scenarios over brute-force token scaling. $\mu$LM-70M surpasses all sub-300M baselines and rivals several sub-1B models on standard benchmarks; even at 9M, it remains competitive. Fine-tuned into a tool-calling agent, $\mu$LM outperforms same-scale baselines and achieves a 97.2\% end-to-end executable rate, presenting a promising direction toward practical sub-100M language modeling.
Reproducing FACTER: Fairness via Conformal Thresholding and Prompt Repair
Oscar Miró López-Feliu ⋅ Daimy van Loo ⋅ Xanthos Kekkos ⋅ Mikel Blom ⋅ Clara Rus
Fayyazi et al. (2025) recently proposed FACTER, a model-agnostic framework designed to jointly enforce fairness and statistical coverage in LLM-based recommendation through conformal thresholding and iterative prompt repair. In this work, we conduct a reproducibility study of the FACTER framework across diverse architectures and dataset sparsity levels, evaluating both the original open-ended generation task and a constrained re-ranking extension. Under the strict reproduction, we observe a divergence in recommendation utility, which we trace to underspecified target-set evaluation in the original study. We then use the constrained re-ranking setting to evaluate FACTER when the candidate set is fixed, and introduce a static Fair Zero-Shot baseline to isolate the contribution of the iterative prompt repair loop. Our analysis shows that FACTER consistently reduces adaptive-threshold violation counts, but that these reductions are not consistently reflected under the fixed threshold or in global fairness metrics. In the constrained ranking setting, static fairness instructions achieve comparable semantic-parity outcomes to FACTER's dynamic repair loop, suggesting that the additional online repair mechanism provides limited benefit in this formulation. All code and reproduction artifacts are available at https://github.com/oscar-omlf/facter-repr.
Auditing Closed-Loop Learning in Recurrent Neural Networks: Reproduction, Robustness, and Generalization
Aarav Sinha
Recurrent neural networks are often used as mechanistic models of learning and control, but closed-loop training creates reproducibility challenges because a model's actions alter future inputs. We conduct a claim-level reproducibility study of Ger and Barak's closed-loop RNN learning dynamics, testing independent implementation, seed variation, protocol perturbations, coupled-system diagnostics, and architecture/task transfer. Under a main-text-aligned double-integrator protocol, the trajectory-level peak, not a persistent final gap, reproduces strongly: 50/50 paired seeds show the post-initial open-loop deployed-loss peak, with a mean peak/initial ratio of 19.1, while final open-loop and closed-loop losses converge after open-loop recovery. The spectral stage and coupled-stability diagnostics also reproduce in 50/50 seeds. A targeted A1 analysis separates stability and behavioral tradeoffs: short-horizon improvements coincide with coupled-radius increases in 120/120 runs; long-horizon loss worsening occurs in 62/120, but never without the radius signal. Generalization is hierarchical: GRU variants preserve final-loss divergence, low-rank variants often produce open-loop deployed rollout blow-ups, tanh RNNs preserve the peak signature without a final gap, tracking transfer is strong, and path-integration transfer is weak. These results motivate reporting practices for closed-loop RNN studies: deployed closed-loop loss, peak signatures, paired seeds, spectral stage criteria, coupled-system spectra, feedback strength, rollout horizon, and failure rates.
A Reproduction Study of Weight-Based Mechanistic Interpretability in Bilinear MLPs
Itay Erlich ⋅ May Ben Zion ⋅ Jacob Shashoua
Mechanistic interpretability typically explains a trained network by analyzing its activations, at the cost of training auxiliary models such as sparse autoencoders (SAEs). Pearce et al. (2025) pursue an alternative: bilinear MLPs, whose omission of element-wise nonlinearities makes each layer an exact quadratic form, so interpretable features can be read directly from the weights via eigendecomposition of the layer's interaction tensor. We present a systematic reproduction of that work. The vision claims reproduce fully: regularization induces low-rank, visually interpretable eigenstructure, and our ablations sharpen the original account by identifying weight decay, rather than noise augmentation, as the dominant cause. The language claims reproduce only partially: we confirm the qualitative discovery of sentiment-negation circuits, but the reported prevalence of low-rank feature interactions holds for only one of three publicly released models (a second becomes consistent once rarely-active dictionary features are excluded). We identify two candidate factors that the original paper leaves unreported – SAE training duration and the underlying model's training compute – and a third that is structural: the public unavailability of matching SAE artifacts, which makes part of the original configuration impossible to replicate exactly. Beyond reproduction, we test whether weight-based features are genuinely structural rather than dataset-specific: regularized bilinear MLPs transfer across handwritten-digit datasets and recognize geometrically similar letters, and we introduce Quadratic Form Similarity, a weight-space metric that separates structurally similar from dissimilar class pairs where eigenvector cosine similarity cannot. Finally, we show on MNIST-scale models that the low-rank structure the original paper discovers post-hoc can instead be enforced during training via Canonical Polyadic (CP) decomposition: at matched full training, CP models match dense accuracy and effective rank, and a bridge experiment connecting the two extensions shows their decision surfaces substantially coincide with the dense ones – enforced and discovered structure converge. Our results are consistent with weight-based interpretability as a viable paradigm, while demonstrating that its reproducibility hinges on artifact availability and training-compute details that interpretability research rarely reports.
Evaluating the evidence trace of NeurIPS 2025 contributions with agentic code auditing
Pascal Iversen ⋅ Ferdous Nasri ⋅ Bernhard Y. Renard ⋅ Katharina Baum
Code serves as the primary evidence behind computational publications, yet a detailed review of an unfamiliar codebase imposes a commonly prohibitive time burden on volunteer reviewers. Consequently, self-reported reproducibility checklists at major machine learning venues are rarely empirically verified. This leaves code as a significant blind spot in peer review. To bridge this gap, we introduce AuditOwl, an autonomous, verification-centric LLM pipeline designed to make code auditing feasible for authors pre-submission and reviewers post-submission. With this framework, we conduct an audit of 100 randomly sampled empirical papers from the NeurIPS 2025 main track. For each paper, an LLM agent inspects the repository and evaluates its scientific claims against the evidence in the underlying code. Following an independent adversarial verification pass to maximize precision, the system raises 609 discrepancies across the 87 papers with retrievable code (median 7 per paper), of which 388 are high or medium severity. It produces evidence that can be quickly checked by humans: findings cite specific code and paper locations, and about half are backed by executable verification checks that the agent implements. Eleven ML researchers evaluated 120 such findings for 30 paper-code pairs and judged 1.7% incorrect and 10.0% incorrect or trivial. Our approach reveals a steep reproducibility funnel: only 87% of papers in our sample release code at all, and for just 8% AuditOwl locates producing code for every traced result. Discrepancies we find are heavily dominated by incompleteness of code and mismatches between what the paper describes and what the code does. We also detect a prevalence of technical bugs and serious methodological issues. Operating at reasonable cost, agentic code auditing can augment human peer review and help to make computational science more reproducible. Code available at: https://github.com/PascalIversen/auditowl-neurips2025.
Reproducibility study of FACTER: Fairness-Aware Conformal Thresholding and Prompt Engineering
Elena Belkina ⋅ Alexandra Elena Antonica ⋅ Brigitta Varga ⋅ Mattia Sferrazza ⋅ Clara Rus
We present a reproducibility study of FACTER, a post-hoc framework that combines confor- mal thresholding with iterative prompt engineering to mitigate demographic bias in black- box LLM-based recommender systems. Using the released codebase and experimental set- ting from the original paper, we evaluate FACTER on MovieLens-1M and Amazon Movies & TV with various LLM backbones. We assess fairness using the reported violation-based criterion group and counterfactual metrics (SNSR, CFR), and measure recommendation quality via catalog-mapped ranking metrics (NDCG@10, Recall@10) alongside the validity rate of generated items (Valid@10). Across datasets and supported backbones, we repro- duce FACTER's key qualitative behavior, with fairness violations decreasing sharply and converging within a small number of calibration rounds. However, unlike the original study, we observe substantially lower recommendation quality for both Zero-Shot and FACTER, with low validity under open-vocabulary generation, item mapping, and evaluation-protocol differences providing plausible explanations for this discrepancy. We further identify and resolve multiple implementation and reproducibility issues in the released code, providing a cleaner and easier-to-run codebase to support future replication. Overall, our findings sup- port FACTER's effectiveness in reducing measured fairness violations, but the near-floor utility leaves utility preservation unverifiable in our reproduction.
[Re] Boosting the Visual Interpretability of CLIP via Adversarial Fine-Tuning
Anton Nuzhdin ⋅ Andrei Gruzitski ⋅ Balázs Egyed ⋅ Aleksandr Raudvee
This paper presents a reproducibility study of "Boosting the Visual Interpretability of CLIP via Adversarial Fine-Tuning" by Gong et al. (2025), published at ICLR 2025, which proposes an unsupervised adversarial fine-tuning (AFT) method with norm regularisation to enhance the visual interpretability of CLIP's image encoder. We attempt to reproduce the key claims regarding improved saliency map quality, increased concept alignment, transferability to out-of-distribution datasets, and the trade-off with zero-shot accuracy. Beyond reproduction, we propose a saliency-guided regularisation extension that introduces an Energy Pointing Game loss, directly supervising the spatial alignment of Simple Gradient saliency maps with target objects. We evaluate our extension across a range of saliency-loss weights and show that explicit saliency supervision substantially improves localisation metrics at a modest cost: the reduction in zero-shot accuracy and adversarial robustness is small at the lowest saliency-loss weight ($w_s=0.2$) and moderate at the strongest setting ($w_s=0.6$). We restructure the codebase for configurable, systematic experimentation. Our reproduction code is available at https://github.com/Andrei3223/Re-Boosting-CLIP-Interpretability.
Recent advances in browser-based LLM agents have shown promise for automating tasks ranging from simple form filling to hotel booking or online shopping. Current benchmarks measure agent performance in controlled environments, such as containers or stable networks, where websites behave deterministically. However, in the real world, users access websites over networks and HTTPS connections that introduce instability from multiple sources: client-side, server-side issues or broader system failures. Moreover, live websites are prone to web attacks such Cross-Site Scripting, as well as general site modifications which can cause unexpected or malicious pop-ups or improper functionality. To address this gap, we present WAREX, a plug-and-play, network-layer tool that integrates with existing web agent benchmarks by simulating common website failures. We measure the impact of WAREX across three popular benchmarks: WebArena, WebVoyager, and REAL. Our experiments show that introducing WAREX leads to significant drops in task success rates, highlighting the limited robustness of state-of-the-art agents. We demonstrate that WAREX serves as more than a diagnostic tool. By fine-tuning an open-source model (Qwen3-8B) on WAREX-generated "failure-recovery" trajectories, we achieve an 88.9% relative improvement in error recovery rates , validating WAREX as a core component for training the next generation of reliable web agents.
Stable Continuous-Time Consistency Distillation: An Empirical Study with a Multistep Extension
Vinay Saji Mathew ⋅ Soundar Kumara ⋅ Gretta Kellogg ⋅ William KM Lai
Continuous-time consistency models are an attractive route to few-step generation. They remove the discretization error inherent to the timestep grid that discrete-time consistency models depend on. Until recently, however, training them was unstable. Lu & Song (2025) address this with the TrigFlow formulation and a coordinated set of architectural and training-objective changes. We conduct an independent empirical study of this framework on CIFAR-10 and ImageNet-64. No open-weight TrigFlow teachers exist at these scales, so we train them from scratch and then verify the consistency distillation (sCD) and consistency training (sCT) behavior of students initialized from these teachers. On CIFAR-10, our students reproduce the reported FID at one and two steps. On ImageNet-64, they fall short, a gap that appears related to our teacher and to architectural choices the paper leaves unspecified. We also report a parameter-count gap of about 5% between the model size stated in the paper and a faithful re-implementation, which we attribute to the scale-and-shift conditioning projections in Adaptive Double Normalization. We then propose Multistep sCD (MS-sCD), a continuous-time, segment-conditioned multistep generalization of sCD. The segment-conditioning idea is borrowed from the discrete-time multistep consistency distillation of Heek et al. (2024); we build on their preliminary continuous-time experiments within the TrigFlow/sCM framework. The rest of the algorithm follows sCD. MS-sCD replaces sCM's two-step intermediate time with a fixed segment-boundary schedule and recovers sCD at M = 1. We report FID at M = 2, 4, 8 and compare against the discrete-time MSCD baseline. Code, model weights, and training scripts are available on GitHub (Mathew et al., 2026).
Opening Up a New Layer: A Deeper Look into "Interpreting CLIP with Hierarchical Sparse Autoencoders"
Nora Kahleogullari ⋅ Adrianna Romanowski ⋅ Mathan Sundarrajan ⋅
Sparse Autoencoders (SAEs) have become essential for decomposing model activations into interpretable concepts. However, despite their effectiveness, SAEs lack a natural ordering of features, making it difficult to prioritize important concepts under compute constraints. The Matryoshka Sparse Autoencoder (MSAE) was introduced as a means of learning nested subspaces, theoretically forcing high-level features into earlier dimensions to enable adaptive granularity. In this paper, we reproduce and analyze the MSAE framework. The reproduction of the study suffers from certain challenges, but the main claims still hold. Furthermore, this study extends on the original paper by implementing feature absorption as a metric, reducing computational cost and emissions with a method inspired by the Sandwich Rule, and deeply analyzing MSAE's ability to learn hierarchical information.
[Re] FairDICE: A Fair Tradeoff in Multi-objective Offline RL
Peter Adema ⋅ Karim Galliamov ⋅ Aleksey Evstratovskiy ⋅ Ross Geurts
Offline Reinforcement Learning (RL) is an emerging field of RL in which policies are learned solely from demonstrations. Within offline RL, some environments involve balancing multiple objectives, but existing multi-objective offline RL algorithms do not provide an efficient way to find a fair compromise. FairDICE seeks to fill this gap by adapting OptiDICE (an offline RL algorithm) to automatically learn weights for multiple objectives to e.g. incentivise fairness among objectives. As this would be a valuable contribution, this replication study examines the replicability of claims made regarding FairDICE. We find that many theoretical claims are supported, but an error in the code reduces FairDICE to standard behaviour cloning in continuous environments, and many important hyperparameters were underspecified. After rectifying this, we show in experiments extending the original paper that FairDICE can scale to complex environments and high-dimensional rewards, though it can be reliant on (online) hyperparameter tuning. We conclude that FairDICE is a theoretically interesting method, but the experimental justification requires significant revision.
Reproducibility study of "Bilinear MLPs enable weight-based mechanistic interpretability"
Hamid Reza Ahmadi ⋅ Mihnea A Bârsan ⋅ Jan Jelínek ⋅ Robert-Stefan Sofroni
This paper presents a reproducibility study of "Bilinear MLPs enable weight-based mechanistic interpretability" by Pearce et al. (2024), which proposes that bilinear architectures possess intrinsic interpretability properties accessible via eigenvalue decomposition. We verify the core empirical image classification claims. Our results confirm the findings for image classification: bilinear layers consistently exhibit an interpretable low-rank structure where the leading eigenvectors capture the majority of task-relevant information, allowing for significant truncation without performance loss. Furthermore, we validate that these eigenstructures are stable across random initializations and varying model sizes. Additionally, we explore extensions to the original work, demonstrating that adversarial training (specifically PGD) enhances the interpretability of eigenvector features on MNIST. Finally, we explored generalization on more complex RGB datasets, such as CIFAR-10 and CIFAR-100, which have generated eigenvectors with uninterpretable structures. All our code is publicly available at: https://anonymous.4open.science/r/reproduced-mech-inter-image-class
Revisiting "Edit Away and My Face Will not Stay: Personal Biometric Defense against Malicious Generative Editing"
Luis Vitor Zerkowski ⋅ Soham Chaudhuri ⋅ Finley Helms ⋅ Jelle Sombekke ⋅ Udit Thakur
Recent advances in diffusion-based image editing have enabled highly realistic and accessible manipulation of facial images, raising serious concerns about biometric privacy and malicious misuse. FaceLock, introduced in Edit Away and My Face Will Not Stay: Personal Biometric Defense against Malicious Generative Editing, proposes an optimization-based defense that embeds subtle perturbations into images at publication time to induce identity distortion in downstream generative edits. The method claims prompt-agnostic effectiveness and strong performance across multiple editing scenarios, supported by open-source code. In this paper, we present a systematic reproducibility study of FaceLock that evaluates its technical, quantitative, and qualitative reproducibility. We assess whether the reported results can be obtained using the released codebase, analyze the correspondence between the paper's algorithmic description and its implementation, and document ambiguities that impact reproducibility. We further examine quantitative reproducibility by attempting to recover the reported performance trends and relative ranking against baselines. We, however, were not able to reproduce the originally reported performance trends, and our outputs were generally worse than those presented in the original paper. Beyond that, we expand the qualitative analysis to a broader set of image–prompt pairs and an additional, harder facial dataset to better test generalization behavior. While we obtained some successful outputs, only a small fraction of our qualitative results matched the consistently high quality reported by the authors. Finally, we introduce an extension to the FaceLock method that helps with robustness, and we critically examine the evaluation criteria used to measure defense effectiveness, highlighting limitations of prompt fidelity as a primary metric and arguing for a more explicit consideration of the trade-off between identity protection and preservation of the original image. We provide a link to our GitHub repository: https://github.com/Luizerko/revisiting_facelock.
Revisiting B2T: Discovering and Mitigating Visual Biases through Keyword Explanations
Faissal El Kayouhi ⋅ Aïda Asma ⋅ Joey Laarhoven ⋅ Fiona Nagelhout
This work aims to reproduce and extend the findings of "Discovering and Mitigating Visual Biases through Keyword Explanation" by Kim et al.(2024). The paper proposes the B2T framework, which detects and mitigates visual biases by extracting keywords from generated captions. By identifying biases in datasets, B2T contributes to the prevention of discriminatory behavior in vision-language models. We aim to investigate the five key claims from the original paper, namely that B2T (i) is able to identify whether a word represents a bias, (ii) can extract these keywords from captions of mispredicted images, (iii) outperforms other bias discovery models, (iv) can improve CLIP zero-shot prompting with the discovered keywords, and (v) identifies labeling errors in a dataset. To reproduce their results, we use the publicly available codebase and our re-implementations. Our findings confirm the first three claims and partially validate the fourth. We reject the fifth claim, due to the failure to identify pertinent labeling errors. Finally, we enhance the original work by optimizing the efficiency of the implementation, and assessing the generalizability of B2T on a new dataset.
XFeat Revisited: Reproducibility and Evaluation of a Lightweight Image Matcher
Lazar Đoković ⋅ Aimee Lin
We present a reproducibility study of XFeat, a lightweight local feature extractor and matcher designed to identify corresponding points across images efficiently on resource-constrained hardware. We re-implement the architecture based on the paper and supplementary material, re-evaluate the authors' released checkpoint alongside our re-implementation, and conduct additional architectural ablations to examine design choices that were not fully justified in the original work. This distinction between re-evaluation and reproduction is important, as the paper, supplement, and public code differ in several implementation details, including the backbone layout, fusion block, and training losses. Empirically, our reproduced models closely match and, in some cases, outperform the re-evaluated original checkpoint on MegaDepth-1500 and ScanNet-1500, supporting the main claim that XFeat provides a strong accuracy–efficiency trade-off for standard image-matching benchmarks. Our ablations provide a more nuanced view of two architectural arguments from the original paper. In particular, the parallel keypoint branch is important for semi-dense matching, but its benefit is less pronounced than originally claimed, while the evidence for the specific placement of the single skip-connection remains inconclusive. Finally, we reproduce the original downstream evaluations and find close agreement for homography estimation, while Aachen visual localization remains below the reported results, even for the released checkpoint, suggesting sensitivity to underspecified evaluation details. We then extend the analysis to zero-shot out-of-distribution and cross-modal matching across retinal, thermal-visible, and multimodal remote-sensing imagery, where XFeat remains effective in some settings but degrades sharply under severe modality shifts.
Regression-adjusted Monte Carlo estimators for Shapley values and probabilistic values combine surrogate modeling with maximum-sample-reuse estimation to reduce the cost of feature-attribution computation. We present a reproducibility study of RegressionMSR in the interventional feature-attribution setting, focusing on artifact fidelity, practical algorithmic choices, and empirical robustness. We first identify implementation-level differences between the estimator displayed in the paper and the released artifact: the code computes separated conditional means over coalitions containing and excluding each feature, rather than the displayed all-row signed residual average, and it is also not guaranteed to use a probability mass in its residual denominator. We characterize these conventions algebraically and verify their predicted targets with synthetic diagnostics. On real data, the stored denominator strongly attenuates the residual correction: correcting both conventions changes LinearMSR little but reduces TreeMSR error by more than half in the tested Shapley settings. We then test the authors' practical single-surrogate simplification, finding that it substantially reduces runtime while maintaining and often slightly improving accuracy. Next, we qualitatively reproduce the headline Shapley benchmark among the original comparison methods, where TreeMSR remains the lowest-error estimator on average; however, OddSHAP performs best in our modern benchmark extension, and rank-stability metrics provide a complementary ordering. Finally, we extend the original Gaussian-noise experiment to Laplacian and coalition-size-dependent heteroskedastic noise, finding that TreeMSR's low-noise advantage persists across other low-noise regimes. Overall, our results support RegressionMSR's empirical value while clarifying implementation details needed for reproducible use.