Session
Sydney Poster Session 4
Hall 1-4
$\mathcal{P}$Torch: Narrowing the Gap Between Projection and Gradient-Based Learning
Jeff Cyuzuzo Jambé ⋅ Jan Quan ⋅ Panagiotis Patrinos
Projection-based optimization offers a gradient-free alternative to conventional gradient-based methods, but existing frameworks suffer from high computational costs and poor scaling with depth. We introduce $\mathcal{P}$Torch, a projection-based framework that builds directly on top of PyTorch's autograd engine, eliminating the overhead of prior approaches. We further improve performance through nonlinear relaxation and memory-efficient projections. Together, these contributions make the algorithm competitive with gradient-based methods for shallow MLPs on MNIST and CIFAR-10. Finally, we prove a vanishing target theorem showing that the layer-wise learning signal decays with increased depth, serving as an analogue of the vanishing gradients problem for projection-based methods.
$PAS^2$: Physics-Anchored Spectral Reasoning for Air Quality Forecasting
Haofeng Ying ⋅ Wenbin Lu ⋅ Wenxin Shen ⋅ Junnan Xu ⋅ Jianwei Zheng
Given the profound impacts on human health and ecosystems, air pollution necessitates robust quality prediction to inform policy-making. Traditional physics-based forecasting relies on idealized closed-system assumptions, whereas data-driven approaches often lack interpretability. Recent hybrid methods, however, neglect data homogeneity, resulting in redundancy, parameter inefficiency, and gradient conflicts. Moreover, high-concentration regions inherently exhibit drastic fluctuations, making it difficult to explicitly distinguish them from the overall pollution field, especially when using spatial information alone. To address these limitations, a Physics-Anchored Spatial-Spectral ($PAS^2$) paradigm is proposed. Specifically, a collaboration scheme is firstly elaborated, within which a physical estimator captures deterministic trends, anchoring the neural branch’s input space and directing it to learn complementary fluctuations. Building on this, we equip the neural branch with a Spatial–Spectral ($S^2$) Reasoner that integrates graph operators to model local spatial diffusion and spectral operators to explicitly capture the drastic fluctuations in high-concentration regions. Experiments on real-world datasets show that $PAS^2$ consistently improves over strong data-driven and physics-guided baselines, especially under volatile sudden-change scenarios. Codes are attached and will be released at GitHub.
$\texttt{RNAGenScape}$: property-guided, optimized generation of mRNA sequences with manifold Langevin dynamics
Danqi Liao ⋅ Chen Liu ⋅ Xingzhi Sun ⋅ Dié Tang ⋅ Haochen Wang ⋅ Scott E Youlten ⋅ Srikar Gopinath ⋅ Haejeong Lee ⋅ Ethan C Strayer ⋅ Antonio J Giraldez ⋅ Smita Krishnaswamy
Generating property-optimized mRNA sequences is central to applications such as vaccine design and protein replacement therapy, but remains challenging due to limited data, complex sequence-function relationships, and the narrow space of biologically viable sequences. Generative methods that drift away from the data manifold can yield sequences that fail to fold, translate poorly, or are otherwise nonfunctional. We present RNAGenScape, a property-guided manifold Langevin dynamics framework for mRNA sequence generation that operates directly on a learned manifold of real data. By performing iterative local optimization constrained to this manifold, RNAGenScape preserves biological viability, accesses reliable guidance, and avoids excursions into nonfunctional regions of the ambient sequence space. The framework integrates three components: (1) an autoencoder jointly trained with a property predictor to learn a property-organized latent manifold, (2) a denoising autoencoder that projects updates back onto the manifold, and (3) a property-guided Langevin dynamics procedure that performs optimization along the manifold. Across three real-world mRNA datasets spanning two orders of magnitude in size, RNAGenScape increases median property gain by up to 148% and success rate by up to 30% while ensuring biological viability of generated sequences, and achieves competitive inference efficiency relative to existing generative approaches.
Hartigan's $k$-means algorithm has theoretical advantages over the standard Lloyd's variant, which has become synonymous with ``$k$-means''. Yet, Hartigan's algorithm is far slower in practice due to its inherently sequential calculations. We devise a Batched Hartigan's method that enables vectorized calculation by exploiting the triangle inequality to expose a set of parallel-safe updates and a smaller set of sequential updates. This makes Hartigan's method competitive with Lloyd's in execution speed for the first time.
A Benchmark for Omni-Modal Reasoning in Long Videos
Mohammed Irfan Kurpath ⋅ Jaseel M Kaithakkodan ⋅ Jinxing Zhou ⋅ Sahal Shaji Mullappilly ⋅ Mohammad Almansoori ⋅ Noor Ahsan ⋅ Beknur Kalmakhanbet ⋅ sambal shikhar ⋅ Rishabh Lalla ⋅ Jean Lahoud ⋅ Mariette Awad ⋅ Fahad Shahbaz Khan ⋅ Salman Khan ⋅ Rao Anwer ⋅ Hisham Cholakkal
Long-form omni-modal video understanding requires models to integrate vision, speech, and ambient audio with coherent long-context reasoning. Existing video benchmarks often trade off temporal scale, modality coverage, open-ended interaction, and interpretable scoring. To address this gap, we introduce $\textbf{LongShOTBench}$, a long video evaluation benchmark designed around three coupled goals: holistic omni-modal integration, intent-driven open-ended interaction, and rubric-level diagnosis. It constructs single- and multi-turn questions from practical viewing scenarios, involving systematical tasks for probing visual, speech, ambient-audio, temporal, and cross-modal reasoning. Each item includes a reference answer and a weighted criterion-level rubric, enabling evaluation to identify which perceptual facts, temporal links, modality-grounding requirements, and reasoning steps are satisfied or missed. All samples are manually verified and corrected to improve grounding, clarity, and rubric reliability. We also introduce $\textbf{LongShOTAgent}$, a training-free omni-modal evidence-seeking agent that couples full-video preprocessing with targeted retrieval, query-adaptive segment refinement, and explicit claim verification over visual, speech, and non-speech audio evidence. Its iterative search-refine-verify loop exposes intermediate evidence and lets modality-specific specialists re-analyze relevant moments before answering. We perform comprehensive evaluation of 105 video-capable models spanning open-source omni-modal models, vision-language systems, audio LLMs, agentic pipelines and closed-source APIs. Across this broad evaluation, current MLLMs remain far from saturating LongShOTBench, while our LongShOTAgent emerges as the strongest training-free system, reaching 66.64\% overall. By releasing the benchmark, leaderboard, and agentic method, our work provides the community with a shared, interpretable testbed for evaluating and advancing long-form omni-modal video reasoning.
AccelEval: A Whole-Program Benchmark for LLM-Generated CPU-to-GPU Code Acceleration
Chendong Song ⋅ Tianci Bu ⋅ Meixuan Wang ⋅ Zijie Zhou
Current LLM-for-GPU benchmarks largely measure CUDA fluency on isolated operators or short kernel-generation tasks, whereas practical GPU performance engineering often starts from trusted sequential CPU programs and requires correct, end-to-end acceleration. To fill this gap, we introduce \textsc{AccelEval}, a benchmark for whole-program CPU-to-CUDA acceleration. AccelEval contains 42 tasks derived from public repositories and established HPC/application suites across six domains, each equipped with deterministic correctness oracles and three input scales. Beyond pass/fail correctness, AccelEval reports end-to-end speedup over CPU baselines and a novel peak-attainment metric which normalizes each model’s speedup against the best correct solution. Using this framework, we evaluate eight frontier LLMs across multiple input scales, GPU architectures and sample rates. At pass@1 on NVIDIA H200 at medium scale, the strongest model passes 38/42 tasks at 71.2$\times$ geometric-mean speedup over the CPU baseline. On the subset of 11 tasks with human-written CUDA references, the best-of-8 LLM oracle reaches 0.75$\times$ of human CUDA performance. To make CUDA-optimization behavior interpretable and reusable, we decompose correct solutions into 43 CUDA acceleration patterns and show that feeding winner-derived pattern guidance back to models improves paired geometric-mean speedup by 1.78$\times$, compared with 1.18$\times$ for a generic-tips control. Together, these results establish AccelEval as a reproducible diagnostic platform for end-to-end LLM-assisted GPU acceleration. An anonymized open-source release is available at \url{https://anonymous.4open.science/r/AccelEval-2397}.
AC/DC on a Budget -- Alternating Sparse Phases
Rahul Nittala ⋅ Advait Gadhikar ⋅ Tom Jacobs ⋅ Rebekka Burkholz
State-of-the-art neural network sparsification relies on alternating dense and sparse training phases. While computationally expensive, dense phases are critical for performance. Our empirical analysis reveals the mechanics behind this gain: upon reintroduction, previously masked parameters receive disproportionately large gradients that drive essential mask rewiring and parameter sign-flipping. Building on these insights, we optimize the performance-to-FLOPs Pareto frontier through three innovations: (1) zero-initializing masked parameters during dense phases to promote relevant sign flips, (2) truncating dense phases based on early mask convergence, and (3) replacing the dense phase entirely with a sparse rewiring phase that updates only the active mask and a small, high-gradient fraction of out-of-mask weights. Evaluated on CIFAR100 and ImageNet, our variant ARC (Alternating Rewiring Compression) matches baseline AC/DC performance at a fraction of the computational cost, establishing a superior performance-to-FLOPs frontier also compared to standard sparse-to-sparse algorithms.
AdaCodec: A Predictive Visual Code for Video MLLMs
Haowen Hou ⋅ Zhen Huang ⋅ Zheming Liang ⋅ Qingyi Si ⋅ Chenglin Li ⋅ Shuai Dong ⋅ Kele Shao ⋅ Ruilin Li ⋅ Dianyi Wang ⋅ Nan Duan ⋅ Jiaqi Wang
Video is temporally redundant: adjacent frames usually share most objects, background, and layout. Yet existing video multimodal large language models (video MLLMs) usually encode each sampled frame as an independent RGB image, causing visual tokens to repeat content already present in earlier frames. This suggests a more direct video interface: send a full reference frame only when the scene cannot be predicted well from prior context, and otherwise transmit a compact description of inter-frame changes. We call this interface a predictive visual code, and instantiate it for video MLLMs as AdaCodec. AdaCodec spends full visual tokens on a reference frame only when its conditional predictive cost is high; otherwise, it encodes inter-frame changes, including motion and prediction residuals, as compact P-tokens. Across all eleven benchmarks, AdaCodec improves over the Qwen3-VL-8B per-frame RGB baseline at a matched visual-token budget. Even at $1/7$ the budget, AdaCodec with 32k tokens surpasses the 224k baseline on all long-video benchmarks; on five general-video benchmarks, it raises the average score while substantially cutting time-to-first-token from 9.26s to 1.62s.
AdapMatch: Adaptive Bias Decoupling for Semi-Supervised Partial Label Learning under Unknown Class Distributions
Yuheng Jia ⋅ Yaodong Shen ⋅ Jiahao Jiang ⋅ Hui LIU ⋅ Junhui Hou
Partial Label Learning (PLL) trains models from instances with candidate label sets containing the ground-truth, while Semi-Supervised Partial Label Learning (SSPLL) further leverages the available unlabeled data to prompt the performance. The previous SSPLL methods implicitly assume the unlabeled data are class-balanced distributed. However, in realistic scenarios, the class distribution of unlabeled data is unknown, resulting in performance collapse on minority classes for the existing SSPLL methods. We identify two coupled biases behind this failure: (i) pseudo-label selection bias, where high-confidence filtering tends to over-select samples from dominant classes, causing the training data to become increasingly skewed toward these classes; and (ii) disambiguation bias, where dominant classes are more likely to win within candidate sets, even when they are incorrect. Theoretically, we show that these two biases reinforce each other over training, forming a feedback loop that degrades generalization performance. To weaken this loop, we propose AdapMatch, which uses DAS to reduce the skew in training data caused by pseudo-label selection bias via prior-aware pseudo-label admission and CAD to lower the disambiguation error by suppressing candidate-set domination. Across five benchmarks, AdapMatch yields gains up to 41.31\% at most , showing strong robustness to varying scale, imbalance, and ambiguity. The code is available at \url{https://anonymous.4open.science/r/AdapMatch-8BC5}.
Adaptive Attribute Completion with Representation Space for Incomplete Graph Domain Adaptation
Niya Yang ⋅ Di Jin ⋅ Zhizhi Yu ⋅ Liang Yang ⋅ Yuxiao Huang ⋅ Dongxiao He ⋅ Jianguo Wei
Graph Domain Adaptation (GDA) has emerged as an effective technique for transferring knowledge from a label-rich source graph to an unlabeled target graph by mitigating cross-graph discrepancies. However, in real-world scenarios, the attributes of partial nodes may be unobserved (i.e., missing) due to factors such as privacy protection. This inevitably exacerbates the distribution discrepancy of nodes between graph domains, making it more challenging to transfer effective knowledge to the target graph. To tackle this issue, we propose an adaptive attribute completion approach with representation space for incomplete graph domain adaptation, namely AC-GDA. It mutually assists the attribute completion and dynamically optimizes cross-graph trustworthy embeddings in an iterative manner, promoting positive knowledge migration while completing attributes. Specifically, we first propose a joint completion mechanism guided by topology similarity, which leverages cross-graph trustworthy attribute embeddings with imputation weights to adaptively complete missing node attributes. We then enhance node embeddings with label information to provide supervision for completed nodes, making nodes of the same category distribute consistently across graph domains. Furthermore, we introduce the complement confidence score to dynamically adjust the scope of trustworthy embeddings, so as to improve the quality of subsequent cross-graph imputation. Extensive experiments on a range of cross-graph tasks validate the effectiveness of AC-GDA.
Adaptive multiscale operator correction via learned spectral subspace and physics-informed optimization.
Subham Patel ⋅ Himanshu Pandey ⋅ RATIKANTA BEHERA
Neural operators have shown strong performance in learning solution mappings for differential equations and offer fast inference across parameter spaces. However, they often suffer from spectral bias, leading to poor resolution of high-frequency and localized features, especially in out-of-distribution settings. In this work, we propose adaptive multiscale operator correction (AMOC), a hybrid framework that combines neural operator learning with physics-constrained optimization in a low-dimensional, instance-specific adaptive spectral subspace. First, a neural operator learns a coarse solution, which is then projected onto a sparse multiscale basis using either wavelet family or Fourier modes, where dominant components are selected based on energy. Finally, a physics-informed optimization is performed on adaptive spectral subspace, significantly lowering computational cost while improving accuracy. The proposed approach enables efficient correction of neural operator predictions by leveraging the sparsity and localization of multiscale representations, while bypassing automatic differentiation in the residual loss through closed-form derivatives of the chosen basis. The proposed AMOC is evaluated on the heat, Poisson, Darcy, and Helmholtz equations, achieving up to four orders-of-magnitude error reduction over FNO, with particularly strong gains observed in challenging out-of-distribution regimes.
AdaptNC: Adaptive Nonconformity Scores for Conformal Prediction under Distribution Shift
Renukanandan Tumu ⋅ Aditya Singh ⋅ Rahul Mangharam
Rigorous uncertainty quantification is essential for the safe deployment of autonomous systems in unconstrained environments. Conformal Prediction (CP) provides a distribution-free framework for this task, yet its standard formulations rely on exchangeability assumptions that are violated by the distribution shifts inherent in real-world robotics. Existing online CP methods maintain target coverage by adaptively scaling the conformal threshold, but typically employ a static nonconformity score function. We show that this fixed geometry leads to highly conservative, volume-inefficient prediction regions when environments undergo structural shifts. To address this, we propose \textbf{AdaptNC}, a framework for the joint online adaptation of both the nonconformity score parameters and the conformal threshold. AdaptNC leverages an adaptive reweighting scheme to optimize score functions, and introduces a replay buffer mechanism to mitigate the coverage instability that occurs during score transitions. We evaluate AdaptNC on diverse robotic benchmarks involving multi-agent policy changes, environmental changes and sensor degradation. Our results demonstrate that AdaptNC significantly reduces prediction region volume compared to state-of-the-art threshold-only baselines while maintaining target coverage levels.
ADIS-Law: Unified Scaling Laws for Annealing-Phase Domain Injection in Large Language Models
Lyuxin Xue ⋅ Hui Cai ⋅ Xiaoyun Feng ⋅ Xuanwei Hu ⋅ Xin Zhang
Continual pre-training (CPT) enables injecting domain-specific knowledge into large language models during the learning rate annealing phase. However, this compute-efficient paradigm forces practitioners to navigate a challenging optimization space: simultaneously choosing the injection fraction $\lambda$, the data mixture ratio $r$, and the re-warmup peak learning rate $\eta$. Because existing scaling laws have yet to comprehensively integrate and couple these interacting variables, they cannot reliably forecast the resulting trade-offs and final convergence outcomes across unseen injection configurations. To bridge this gap, we introduce the Annealing-phase Domain Injection Scaling Law (ADIS-Law), a parametric framework that explicitly models late-stage distribution shifts as physical perturbations applied to the base pre-training trajectory. ADIS-Law formalizes two opposing macroscopic behaviors: general-domain performance follows a linear forgetting-recovery process, whereas target-domain adaptation adheres to a power-law adaptation trend. Across data-budget and model-scale extrapolation, ADIS-Law predicts target-domain trajectories with an average Spearman correlation of $0.947$ and forecasts tail losses with relative error below $5\\%$. We further show that ADIS-Law can predict optimal injection strategies at extrapolated scales: it ranks target-scale candidate policies with a mean Spearman correlation of $0.936$. By replacing blind grid search over $(\lambda, r, \eta)$ with a single fitted law, ADIS-Law offers a data-driven approach for strategy selection in annealing-phase domain injection.
AdKnob: Ad Intensity Control and Labeling for LLM-Native Advertising
Woo Jae Kim ⋅ Seongho Keum ⋅ Joonsung Jeon ⋅ Suhyeon Ha ⋅ Sooel Son ⋅ Sung-eui Yoon
Large language model (LLM)-native advertising has emerged as a viable monetization channel as LLM service providers seek sustainable revenue streams. However, integrating native ads into LLM responses introduces two largely understudied challenges: (1) controlling ad intensity---how prominently advertising content appears---and (2) ad labeling---identifying which segments of the generated text constitute advertising. We propose AdKnob, a framework jointly addressing both challenges. For ad intensity control, we introduce new ad control tokens without pretrained semantic associations and align the model via direct preference optimization to generate ads at each desired level. For ad labeling, we propose an attention rollout-based attribution mechanism that identifies ad segments by tracing their attribution to input ad-related tokens, requiring no additional inference cost. We further construct MI-KnobSet, the first multi-level ad intensity preference dataset for LLM-native ads. Experiments across six LLMs and human evaluations show that AdKnob achieves uniformly spaced ad intensities where baselines collapse to a narrow range, achieves the highest ad labeling accuracy while being over 500$\times$ faster than the strongest baseline, and is preferred over baselines by human evaluators up to 89.6\% of the time. Code and data will be publicly available.
ADMIT: Support-Gated Memory-Write Admission for Document QA Agents
Joongmin Shin ⋅ Gyuho Shim ⋅ Hyeonseok Moon ⋅ Jaehyung Seo
Document QA agents do not fail only at the final answer. They can also write unsupported claims into public working memory, and those writes can affect later retrieval and synthesis. We introduce ADMIT (Auditable Decoupled Memory-write admission for Intermediate Traces), a low-budget admission controller for public OCR/text evidence packs. ADMIT applies a deterministic substring-support check $S_{\text{str}}(\hat{y}, B) > 0$ to each public commit and allows one bounded RepairRead. If repair still fails, ADMIT uses a memory-fail-closed / output-fail-open rule: it writes a $\bot$ (NO_COMMIT) sentinel to memory while the final-answer surface may still fall back to the baseline answer. On held-out M3DocVQA, a retrained final-row router closes the answer-surface gap; ADMIT's distinguishing evidence is therefore memory-surface contamination, not final-answer accuracy. The lexical invariant guarantees zero-substring accepted intermediate contamination by construction; the contribution is therefore the public memory-write boundary as a low-cost control surface with decoupled fail actions, not the invariant itself. In a 7-arm dynamic stress diagnostic on three agents, ADMIT enforces the invariant and preserves raw EM at this operating point (zero contaminated trajectories on every agent and no detectable raw-EM loss relative to the fail-open ablation), whereas final-only baselines still leave $12$--$32%$ contamination. ADMIT enforces lexical auditability, not semantic truth: relation-level entailment, citation-aware faithfulness, and visual-region grounding are not claimed, and lower-density VisDoMSlide and text-only FRAMES are reported as boundary cases.
A Dual-Domain Vision Transformer with Spectral Positional Bias
Abdulaziz Alshamsi ⋅ Abdulla Alghfeli ⋅ Chenghua Lin ⋅ Hujun Yin
Positional encoding in Vision Transformers presents a recurring trade-off: spatial relative-position tables (Swin) carry thousands of parameters per block and require ad-hoc interpolation to transfer across resolutions, while no-bias designs leave the attention logits geometry-agnostic. We propose a relative positional bias parameterized directly in the frequency domain. This parameterization yields three benefits over the standard spatial-table position encoding: (i) it expresses the same translation-equivariant bias with $3.25\times$ fewer parameters than Swin's spatial table at the same head count; (ii) it admits a closed-form, structurally exact transfer to a different patch grid, in contrast to bilinear interpolation of spatial tables which smooths the learned structure; and (iii) it shares the same parameterization style as a frequency-domain spectral mixer, enabling co-design of positional encoding and token mixing in a common substrate. We demonstrate the bias inside our **Dual-Domain Residual (DDR)** framework, in which a spectral mixer and a gated attention head act as parallel residuals within each block. Head-to-head, the frequency-domain bias recovers $+1.98$ pp over Swin RelPos and $+2.74$ pp over no positional bias; removing the entire attention branch (which carries the bias) drops accuracy by $5.35$ pp, while removing the spectral branch drops it by $2.22$ pp. On ImageNet-1K from scratch, DDR outperforms SpectFormer at every comparable parameter budget, by $+2.0$ pp at Tiny (DDR-Ti-SF, 10.7M, 78.9\%) and $+0.4$ pp at Base (DDR-B-Deep, 63.1M, 82.5\%); at sub-parity parameters, DDR-Ti+ (8.9M, 77.9\%) exceeds SpectFormer-Ti (9.0M, 76.9\%) by $+1.0$ pp. The learned per-layer gates concentrate spectral processing in early blocks and attention in late blocks, qualitatively matching the schedules that are hand-coded in the previous hybrid models.
Advancing Affordance-Grounded Creative Tool Use in Large Multimodal Models
Cheng Qian ⋅ Hyeonjeong Ha ⋅ Jiayu Liu ⋅ Jeonghwan Kim ⋅ Emre Can Acikgoz ⋅ Bingxuan Li ⋅ Kunlun Zhu ⋅ Jiateng Liu ⋅ Aditi Tiwari ⋅ Zhenhailong Wang ⋅ Xiusi Chen ⋅ Heng Ji
Large multimodal models (LMMs) have advanced rapidly in perception and reasoning, yet it remains unclear whether they can discover visually grounded, physically feasible solutions in open-ended environments. We first introduce MM-CreativityBench, a benchmark for affordance-grounded creative tool use, where models must inspect scenes, entities, and parts to identify non-obvious object uses grounded in visual evidence. Our evaluation shows that current LMMs often fail to sustain grounded exploration: they overlook relevant entities, under-examine critical parts, or hallucinate unsupported attributes. To address these failures, we propose affordance-grounded alignment, framing creative tool use as a preference learning problem. Using Direct Preference Optimization and supervision from an affordance knowledge base, we train models to favor visually grounded attribute–affordance reasoning over hallucinated alternatives while improving exploration efficiency. Our results yield consistent gains in correct entity and part selection, and reduce hallucination and grounding errors. These findings position grounded creativity as a core capability for future multimodal agents: the ability to adapt to unfamiliar environments and solve problems beyond memorized patterns or surface-level plausibility, moving closer to real human-like intelligence.
Advancing Entropy-Level Credit Assignment in RLVR via Proximal Entropy Policy Optimization
Yun Kim ⋅ Nojun Kwak
Value-model-free RLVR methods such as GRPO assign uniform advantages to all tokens in a rollout, ignoring that tokens contribute unequally. Recent methods use token entropy as an importance proxy but compute it globally across the batch, conflating importance with prompt difficulty and positional trends. We argue that importance should instead be measured relative to a token's own local context, and introduce proximal entropy: a local measure of token importance relative to neighboring tokens, and prove it is invariant to both confounders. Proximal Entropy Policy Optimization (PEPO) uses it to weight per-token advantages and outperforms GRPO and entropy-based baselines on mathematical reasoning across Qwen3-1.7B, Qwen3-4B, and Llama-3.2-3B-Instruct. We also show the formulation generalizes to other algorithms where substituting proximal entropy into existing methods improves, and applying it to single-stream RL succeeds where global entropy fails.
Adversarially Attacking Symbolic Vocabulary Vulnerabilities In LLM Planners
Stephen Obadinma ⋅ Xiaodan Zhu
As large language models (LLMs) become increasingly deployed as autonomous agents, the ability of these models on automated planning tasks has become essential. However, despite their strengths, LLMs' planning ability is not robust, particularly when they are not allowed to exploit commonsense cues in the symbolic vocabulary (e.g., names of actions, predicates, objects) used to define planning tasks. This makes them unreliable compared to symbolic planners. With recent reasoning models showing improved planning performance on tasks with obfuscated symbolic vocabularies, we utilize an adversarial attack framework to reveal to what extent these crucial weaknesses of LLM agents remain and whether they can reason properly when presented with severe cases. As such, our main contribution is devising and optimizing an LLM-based attack model which we call $\texttt{Symbolic-Swapper}$ to automatically generate obfuscated domains with non-standard symbolic vocabularies that systematically break strong reasoning models' ability to plan. In doing so, we reveal by how much LLMs planning abilities can be further degraded under more advanced obfuscation schemes, allowing us to ascertain their success when having to rely on their pure planning ability. We find that attacks can decrease the planning performance of base models and even robust planning frameworks by over 70\%. We further analyze the factors behind how their planning abilities break down, and find that successful obfuscation models to significantly degrade in reasoning quality despite them showing an ability to understand the domain, revealing a gap in how model's perceive a task domain and how they are actually able to plan under it, with attacks often bypassing model ability to successfully self-correct.
Ensuring safety guarantees in offline reinforcement learning remains challenging, especially when safety constraints must hold almost surely. Moreover, as pre-specifying a single safety budget (constraint threshold) is often challenging, it is desirable to learn a foundation policy that can be deployed across a broad range of budgets. We introduce AEGIS (Almost-sure Epigraph-Guided Implicit Safety), a safe offline RL framework that can guide diffusion policy training via critics that respect almost-sure constraints across various feasible budgets. AEGIS characterizes the feasible set of initial state-budget pairs as the epigraph of a feasibility critic updated via the worst-case backup. Building on the proposed characterization, we extend implicit Q-learning (IQL) to train both feasibility and reward critics. We use these critics to bias a diffusion policy toward high-value feasible actions. Consequently, AEGIS turns diffusion from a generative prior into a safety-aware controller, enabling a single general policy to respect various budgets without further tuning. Empirical results on the DSRL benchmark and humanoid locomotion tasks show that AEGIS achieves high feasibility with competitive returns, generalizing across feasible constraint thresholds.
Affine-invariant Cubic Newton with Weak Learners
Nikita Zozoulenko ⋅ Daniel Falkowski ⋅ Thomas Cass ⋅ Lukas Gonon
Modern gradient boosting algorithms are based on Newton's method, which excels empirically but lacks global convergence guarantees. While recent work on Gradient-Regularized Newton (GRN) boosting achieved a global $\mathcal{O}(1/\gamma^{12}k^2)$ rate, its analysis was not able to exploit the natural Hessian geometry and suffers from a $\gamma^{12}$ dependence on a worst-case `weak gradient edge' $\gamma$. In this paper, we tackle these theoretical drawbacks by generalizing the Affine-Invariant Cubic Newton (AICN) scheme to convex optimization with weak learners. Our analysis is based on average cosine values $\overline{\Theta}_k$ along the optimization path and directional improvements $\delta_k$ in the Hessian-induced norm. We show that the weak-AICN update reduces to a damped weak-Newton step with an explicit stepsize determined by $\delta_k$. Under semi-strong self-concordance, we establish a global $\mathcal{O}(1/\overline{\Theta}_k^4 k^2)$ rate, significantly improving upon the weak-GRN bound. We further prove a local linear rate, global linear convergence for Hessian-dominated losses, and cosine angle lower bounds. Finally, we demonstrate the broader applicability of this framework beyond gradient boosting by showing that Newton coordinate descent emerges naturally as a special case of optimization with weak learners.
AffordSim: A Scalable Data Generator and Benchmark for Affordance-Aware Robotic Manipulation
Mingyang Li ⋅ Haofan Xu ⋅ Haowen Sun ⋅ Xinzhe Chen ⋅ Sihua Ren ⋅ Liqi Huang ⋅ Chenyang Miao ⋅ Xinyang Sui ⋅ jiawei Ye ⋅ Qiongjie Cui ⋅ Zeyang Liu ⋅ Xingyu Chen ⋅ Xuguang Lan
Many everyday robot manipulation skills are affordance-dependent, with success determined by whether the robot contacts the functional object region required by the subsequent action. Current simulation data generators obtain contacts from generic grasp estimators or per-object manual contact annotations, but generic estimators rank stable grasps without task semantics and often select contacts that are misaligned with the downstream action, while manual contact annotations must be rewritten for each new object and task. To solve these challenges, we introduce AffordSim, a scalable data generator and benchmark that integrates open-vocabulary 3D affordance prediction into simulation-based trajectory generation. Given a natural-language task description, AffordSim synthesizes a task-relevant scene, emits affordance queries, grounds them on object surfaces, samples region-conditioned grasps, and selects executable candidates with motion planning. It further randomizes object pose, texture, lighting, image noise, and cross-viewpoint backgrounds for sim-to-real transfer. We instantiate AffordSim as a 50-task benchmark across diverse manipulation skills, five robot embodiments, and 500+ rigid and articulated objects. AffordSim achieves $93\\%$ of the trajectory collection success rate of manual contact annotations on affordance-critical tasks and $89\\%$ on hard composite tasks. Vision-language-action policies trained on AffordSim data transfer zero-shot to a real Franka FR3, reaching $24\\%$ average success. The data generator, benchmark, and code are available at \href{https://anonymous.4open.science/r/AffordSim-8D94/}{project page}.
Agent$^2$ RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
Wanyi Chen ⋅ Xiao Yang ⋅ Xu Yang ⋅ Tianming Sha ⋅ Qizheng Li ⋅ Zhuo Wang ⋅ Bowen Xian ⋅ Fang Kong ⋅ Weiqing Liu ⋅ Jiang Bian
We introduce Agent$^2$ RL-Bench, a compact diagnostic benchmark for evaluating agentic RL post-training, which tests whether LLM agents can autonomously design, implement, debug, and execute post-training pipelines that improve foundation models. RL post-training increasingly drives model alignment and specialization, yet existing benchmarks are largely static, rewarding supervised fine-tuning or script generation without assessing an agent's ability to close an interactive RL loop. Agent$^2$ RL-Bench provides a unified agent-facing interface: each run starts from an isolated workspace containing a base model, task data, instructions, and a grading API, and agents must iterate within a fixed budget by training models and submitting artifacts for evaluation. The benchmark spans six tasks across three levels, from static rule-based training to judge-based optimization and closed-loop online RL with trajectory collection. Two diagnostic skills, namely runtime recording and post-hoc summarization, enable structured analysis of agent behavior, facilitating smooth and effective iteration of the benchmark's evaluation framework. Across five agent systems and six driver LLMs, agents show intelligent behavior but clear limitations: one RL-oriented run improves ALFWorld from 4.85 to 93.28 via SFT warm-up and GRPO with online rollouts, yet DeepSearchQA remains difficult, most successful routes rely on supervised pipelines, and interactive outcomes show large single-run differences across agent stacks. Overall, Agent$^2$ RL-Bench shows that current agents can sometimes engineer online RL, but stable agent-driven RL post-training remains rare under fixed budgets. It also demonstrates that our benchmark provides a strong and effective evaluation framework for future research in this direction. Code is available at https://anonymous.4open.science/r/Agent2RL-Bench-5587.
Agent Explorative Policy Optimization for Agentic Multimodal Reasoning
Minki Kang ⋅ Shizhe Diao ⋅ Ryo Hachiuma ⋅ Sung Ju Hwang ⋅ Pavlo Molchanov ⋅ Frank Wang ⋅ Byung-Kwan Lee
Vision-language models with extended reasoning succeed on complex problems, but many real-world problems require external tools that internal reasoning alone cannot reach. Agentic reasoning therefore interleaves two behaviors with a structural asymmetry: thinking (the self-contained default) and tool use (a high-variance auxiliary acting). We refer to this asymmetry as the Thinking-Acting Gap. Under standard RL recipes like GRPO, the gap manifests as two diagnostic symptoms during training: tool use is attempted on only ~30% of rollouts, and when attempted, the tool-using rollouts within a group are all-wrong on ~40% of questions, suppressing the learning signal at the tool calls that needed it. We propose AXPO (Agent eXplorative Policy Optimization): for each all-wrong tool-using subgroup, AXPO fixes the thinking prefix and resamples the tool call and its continuation, paired with uncertainty-based prefix selection. Across nine multimodal benchmarks and three scales of Qwen3-VL-Thinking, SFT+AXPO outperforms SFT+GRPO at average (+1.8pp Pass@1 and +1.8pp Pass@4 at 8B on average) and 8B with SFT+AXPO surpasses the 32B Base on Pass@4 with 4x fewer parameters.
AgentSSL: Can MLE Agents Leverage Unlabeled Data?
Akanksha Sarkar ⋅ Ethan Lin ⋅ Ziang Liu ⋅ Kristin Branson ⋅ Jennifer Sun
Acquiring high-quality labeled data and engineering effective learning pipelines are two critical bottlenecks to applying machine learning systems in new domains. Although recent MLE agents have shown promise in automating pipeline design, their behavior has been studied almost exclusively in fully supervised settings. In many domains, however, labeled data is expensive while unlabeled data is abundant. In this work, we study agents in the semi-supervised learning (SSL) regime by framing SSL as a program search problem and exploring how agents navigate this space under limited supervision. Our analysis focuses on three questions: (1) whether agents can effectively leverage unlabeled data, (2) how they leverage it and what strategies they discover, and (3) whether they can adapt these strategies to specific domains. We find that agents can leverage and synthesize SSL strategies spanning decades of SSL research, achieving performance comparable to, and in some cases outperforming, established human-engineered SSL methods. We further validate these observations on real-world ecological datasets, highlighting both the challenges and opportunities of transferring agent-discovered strategies beyond controlled benchmarks.
A Hebbian Recurrent Neural Network Explains the Hierarchical Geometry of Sequence Memory
Zhitao FENG ⋅ Bo Ho ⋅ Huan Luo ⋅ xiaolong zou ⋅ Yuanyuan Mi
Temporal sequences in working memory are hierarchically organized to support flexible behavior. Recent neural recordings suggest that this organization relies on a 1D-to-2D neural geometrical folding in working memory, whereby ordinal positions are re-encoded within an orthogonal scaffold, with two dimensions indexing global (chunk-level rank) and local ranks (within-chunk position), respectively. This findings implicate an important structural principle for memory, yet the computational mechanisms underlying this geometry remain unclear. Here, we study this question using a recurrent neural network with fast Hebbian plasticity (Hebb-RNN). When meta-trained on rank recalling tasks, the Hebb-RNN generalizes to novel item–rank bindings and hierarchical sequences, outperforming control models. Mechanistic analysis reveals that the network develops an orthogonal ordinal geometry consistent with empirical findings. This geometry transforms temporal inputs into abstract ordinal ranks that provide a memory scaffold, while Hebbian plasticity supports rapid conjunctive binding of item content to this structure. Furthermore, the pre-learned geometric scaffold significantly accelerates learning in novel tasks, supporting efficient transfer. By incorporating a predictive module, our model extends to real-world encodings and acquires hierarchical geometry without explicit chunking signals. Together, our results provide a biologically plausible framework for hierarchical sequence representation, offering new insights into structured memory organization for both neuroscience and machine learning.
AI Control for Sandbagging on Fuzzy Tasks
Mikhail Terekhov ⋅ Caglar Gulcehre ⋅ Vivek Hebbar ⋅ Joe Benton
AI models deployed in critical domains, such as AI safety research, may subtly sabotage our efforts due to misalignment. Sandbagging is a form of sabotage in which an AI intentionally underperforms on the task given to it. It is particularly pernicious on fuzzy tasks, i.e. tasks which are hard to grade or require intuition. To understand sandbagging on fuzzy tasks, we introduce a novel AI control framework that considers sandbagging as an adversarial game between a blue team and a red team. The blue team uses a weak trusted model to construct a weak score against which they would train a strong, potentially sandbagging model to remove the sandbagging if it were present. The red team then tries to find model behaviors that are rated highly by the weak score, and thus might not be trained out, but actually correspond to poor performance. We test our framework on the task of writing experimental proposals for research questions from recent ML papers. We use a language model with access to the original paper as a proxy "ground-truth" scorer. Our red team discovers sandbagging behaviors using multi-objective evolutionary prompt optimization. We show that Opus 4.6 can write proposals that are worse according to the ground truth proxy than those of GPT-OSS-20B, while the weak scorer rates them as highly as the best proposals from Opus 4.6. To mitigate the threat of sandbagging, we propose an adversarial optimization algorithm for the blue team that discovers more robust prompts for the weak model. This algorithm produces a blue team prompt that our red team optimization fails to exploit.
AirIAD: Agentic Iterative Reasoning for Industrial Anomaly Detection
Huan Yu ⋅ TIECHENG ZHANG ⋅ Jin Wang ⋅ Jingru Yang ⋅ Kaixiang Huang ⋅ Zhongyang Sun ⋅ Guodong Lu ⋅ Shengfeng He
Multimodal large language models offer a promising foundation for industrial anomaly detection (IAD), yet they remain unreliable for fine-grained diagnosis, where subtle defects require tight alignment between localized visual evidence and domain-specific semantics. Existing single-pass frameworks produce predictions without revisiting or validating intermediate evidence, which leads to spatio-semantic misalignment and error accumulation. More importantly, they lack mechanisms to actively control how additional evidence is acquired and integrated during reasoning, limiting their ability to resolve ambiguous cases. To address these limitations, we propose AirIAD, an agentic iterative reasoning framework that unifies progressive refinement with active information acquisition. AirIAD equips the model with two tools, Perceptive Zoomer and Knowledge Retriever, which enable it to iteratively gather complementary visual and semantic evidence. Through this process, the model performs explicit cross-verification between what is observed and what is known, refining intermediate hypotheses until convergence on a reliable diagnosis. This capability is enabled by a two-stage training pipeline. A supervised fine-tuning stage introduces a Spatio-Semantic Cross-Verification Chain-of-Thought, which structures coordinated perception and knowledge grounding. This is followed by a Spatio-Semantic Group-in-Group Policy Optimization stage, which provides dense step-wise supervision to encourage consistent self-refinement and robust spatio-semantic alignment. Extensive experiments on the MMAD benchmark show that AirIAD, built on Qwen3-VL-Instruct-4B, achieves state-of-the-art performance in anomaly detection and fine-grained diagnosis.
Algorithm for Contextual Queueing Bandits with Rate-Optimal Queue Length Regret
Seoungbin Bae ⋅ Dabeen Lee
Contextual queueing bandits provide a framework for learning to schedule heterogeneous jobs under unknown context-dependent service rates. Under stochastic contexts, existing algorithms achieve $\widetilde{\cal O}(T^{-1/4})$ queue length regret, defined as the expected difference between the learner's and oracle's queue lengths at horizon $T$. In this paper, we improve this rate to $\widetilde{\cal O}(T^{-1/2})$. The key observation is that random exploration is needed only up to a carefully chosen cutoff round, rather than throughout the entire horizon. We propose CQB-$\eta$-2, a three-phase algorithm: (i) pure random exploration to construct an initial estimator, (ii) $\eta$-random exploration combined with a UCB rule to continue learning while maintaining negative drift, and (iii) pure UCB after the exploration cutoff. Our proof decomposes the queue length regret at the cutoff round. Before the cutoff, negative drift suppresses queue length differences caused by suboptimal choices. After the cutoff, the first two phases provide sufficient random exploration samples, ensuring that UCB decisions incur small departure-rate gaps. Combining these two bounds yields queue length regret of order $\widetilde{\cal O}(T^{-1/2})$. We further prove a minimax lower bound of order $\Omega(T^{-1/2})$. The proof constructs two hard instances that are statistically indistinguishable up to the final service decision, and uses a queue-specific coupling argument to convert the resulting testing error into queue length regret. Together, our upper and lower bounds characterize the minimax dependence on the horizon $T$ up to logarithmic factors.
Aligning Inductive Bias for Data-Efficient Generalization in State Space Models
Qiyu Chen ⋅ Guozhang Chen
The remarkable success of modern AI has been closely tied to scaling laws, yet the finite supply of high-quality data makes data efficiency—learning more from less—an increasingly important frontier. A model’s inductive bias is a critical lever for data efficiency, but foundational sequence models such as State Space Models (SSMs) often rely on fixed, task-agnostic biases. When this fixed prior is misaligned with the underlying structure of a task, the model may require additional samples to overcome its own bias before learning the relevant signal. In this work, we introduce a principled framework for understanding and aligning the inductive bias of linear time-invariant SSMs. We first formalize this bias through an SSM-induced kernel and show theoretically and empirically that its spectrum is governed by the model’s frequency response. This characterization motivates Task-Dependent Initialization (TDI), a fast power-spectrum matching method that aligns the initial SSM bias with the task’s spectral characteristics before downstream training. Across controlled synthetic experiments, trainable one-layer SSMs, and deep SSMs on diverse real-world benchmarks, TDI can improve data-efficient generalization primarily when task-relevant spectral structure is present and the default SSM bias is spectrally mismatched. Our results provide both a theoretical lens and a practical tool for task-adaptive inductive bias, suggesting a path toward more data-efficient sequence modeling.
Aligning LLMs Toward Multi-Turn Conversational Outcomes Using Iterative RLHF
Daniel Jiang ⋅ Ankur Samanta ⋅ Yukai Yang ⋅ Jalaj Bhandari ⋅ Remi Munos ⋅ Tyler Lu
Training large language models (LLMs) as multi-turn conversational agents remains a significant challenge, particularly in goal-oriented settings. The difficulty stems from sparse, long-horizon objectives and the discrepancy between response-level planning and token-level generation. In this paper, we present a formal reduction of the multi-turn RL problem into a \emph{sequence of single-turn RLHF-style problems}. This is achieved by setting a learned multi-turn Q-function as the reward model for the single-turn problem. We demonstrate and prove a key insight: solving this single-turn RLHF problem with standard token-level GRPO is equivalent to an approximate policy improvement step within the multi-turn problem. This insight naturally leads to \emph{Iterative GRPO}, a batch online approximate policy iteration algorithm that alternates between collecting a batch of data from the current policy, fitting Q-functions from these logged conversation trajectories, and improving the policy via single-turn RLHF. A major practical advantage is that Iterative GRPO directly leverages stable, off-the-shelf single-turn RLHF tools, making it straightforward to implement. In addition, the method occupies a middle ground between fully online and fully offline approaches, retaining the adaptability of online updates while gaining the stability benefits of offline training. Empirically, we demonstrate the effectiveness of Iterative GRPO on six multi-turn conversational environments.
Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation
Bruce C Xu ⋅ Adarsh Kumarappan ⋅ Mu Zhou
Key--value (KV) cache quantization is widely used to reduce Large Language Model (LLM) inference memory, yet existing evaluations solely focus on measuring perplexity and accuracy without assessing the safety impact. In this study, we explore alignment preservation under KV cache quantization. Across eleven instruction-tuned models (3.8B--72B) and five benchmarks (1,894 prompts), we find that low-bit quantization can silently destroy safety alignment (Mistral-7B loses 15.2\% of its refusals at only $1.03\times$ perplexity, and no universal safe bit-width exists), with sharp model-specific phase transitions invisible to standard metrics. We identify that the root cause is geometric: safety features occupy a low-dimensional activation subspace $10^2$-$10^3\times$ more vulnerable to quantization noise than the full representation space perplexity averages over. Inspired by this observation, we propose **Per-Channel Reduction** (PCR), a diagnostic that classifies each model into one of three mechanistic failure modes: (i) *outlier-crushes-safety*, where safety lives in non-outlier channels collaterally damaged by outlier-driven scale factors; (ii) *outlier-as-safety*, where safety overlaps outlier channels and finer granularity cannot rescue it; (iii) *multi-layer dilution*, where safety is distributed across many layers and per-layer fixes fail. PCR predicts the correct mitigation direction on all nine primary models and one held-out model from an independent family (20 calibration prompts). PCR generalizes across unseen prompts, models, and production quantizers (KIVI: up to 97.2\% recovery), succeeding where attention-based allocation methods fail. The resulting training-free protocol (${\sim}35$ GPU-minutes) recovers up to 97\% of lost alignment at minimal memory overhead, addressing vulnerabilities confirmed in production vLLM serving with FP8 KV cache on NVIDIA GPUs.
AlloSpatial: Agentic Harness Framework for Spatial Reasoning in Foundation Models
Shouwei Ruan ⋅ Bin Wang ⋅ Zhenyu Wu ⋅ Qihui Zhu ⋅ Yuxiang Zhang ⋅ Jingzhi Li ⋅ Yubin Wang ⋅ Xingxing Wei
Multimodal Foundation Models (MFMs) have made substantial progress, yet remain fragile in spatial reasoning over the physical world. A key bottleneck lies in their inability to transform local egocentric observations into a global allocentric spatial representation. To address this, we propose AlloSpatial, an agentic framework for allocentric spatial cognition in foundation models. AlloSpatial introduces World2Mind, a plug-and-play cognitive mapping sandbox that converts egocentric observations into structured allocentric priors, including Allocentric-Spatial Trees and route maps that support querying object topology, geometric relations, passability, and trajectories. To utilize these priors reliably under noisy reconstruction and ambiguous visual evidence, AlloSpatial introduces a Spatial Reasoning Harness for tool-use judgment, modality-decoupled cue collection, and geometry-semantic arbitration. We further internalize this process in Qwen3-VL through cold-start reinforcement learning with a harness-gated trajectory-level reward. Experiments on VSI-Bench and MindCube show that AlloSpatial improves proprietary models by 5\%-18\% in a training-free setting, while ASTs alone support strong spatial reasoning even when visual inputs are removed. The trained AlloSpatial agents further outperform larger general-purpose models and competitive spatial baselines, suggesting that structured allocentric representations, active tool use, and verifiable reasoning offer a promising route toward spatially capable foundation models.
AMGenC: Generating Charge Balanced Amorphous Materials
Yan Lin ⋅ Jilin Hu ⋅ N M Anoop Krishnan ⋅ Morten Smedskjaer
Amorphous (disordered) materials require simulation cells with at least hundreds to thousands of atoms to accurately capture their diverse local atomic environments, compared to crystalline materials that can be described by unit cells containing a few to hundreds of atoms. To advance the design of amorphous materials with desired properties and facilitate the exploration of their vast design space, generative inverse design has emerged as a promising approach. It aims to directly output materials with properties closely aligned with the desired ones using probabilistic generative models conditioned on desired properties, which can be more resource efficient than the traditional trial-and-error approach. However, due to the inherent stochasticity of probabilistic generative models, when element assignments are unconstrained, a large portion of generated materials may be charge unbalanced, and no existing methods can effectively mitigate this limitation. In this work, we propose AMGenC, a new generative inverse design method for amorphous materials that can guarantee the generation of charge balanced samples, with minimal additional computational overhead and without sacrificing inverse design accuracy. AMGenC achieves this through an element noise that gives the generation process a starting point centered around charge balance, and the combination of a per-step soft projection and a final discrete projection for steering the elements toward exact charge balance throughout the generation. We perform extensive experiments on two amorphous materials datasets. Experimental results provide evidence that AMGenC achieves its design goal.
A Model of Diverse Sampling from Language Models
Manuel Prada-Corral ⋅ Yahya Emara ⋅ Timothy O'Donnell ⋅ Ryan Cotterell ⋅ Tim Vieira
Language models often produce high-quality individual samples but poor sets, with repeated samples clustering around the same semantic modes. We formalize diverse generation as sampling a size-$K$ subset of complete strings from a determinantal point process (DPP), where the likelihood encodes item quality, and the embedding geometry encodes repulsion between similar outputs. Exact inference over all strings is intractable, so we propose \algname: draw a finite candidate pool from a tractable proposal, importance-weight candidates toward a globally tempered quality distribution, and run either $K$-DPP sampling or greedy MAP selection on the induced pool kernel. We prove a finite-sample bound showing that the KL bias of the pooled sampler decays as $\mathcal{O}(1/N)$ in the pool size. Empirically, \algname improves the quality--diversity Pareto frontier over temperature sampling, diverse beam search, $K$-means reranking, and LLM-based selection on open-ended generation, ambiguous text-to-SQL (Ambrosia), and mathematical reasoning (GSM8K) tasks.
An Equivariance Principle for Optimizer Design: Symmetry-Compatible Updates for Embeddings, LM Heads, and MoE Routers
Tim Tsz-Kit Lau ⋅ Weijie Su
Most modern deep neural networks are trained with Adam and related coordinate-wise adaptive optimizers, which are efficient but do not respect the natural matrix symmetries, geometry, and spectral structure of neural network parameters. We develop a symmetry-based principle for matrix-gradient optimization, showing that orthogonally equivariant optimizers, especially spectral optimizers such as stochastic spectral descent, Muon, Scion, and polar gradient methods, are better aligned with matrix layers than coordinate-wise updates. We show that coordinate-wise adaptivity breaks orthogonal equivariance, discards gradient spectral structure, and can be viewed as applying spectral updates to a pathological diagonal lifting of the parameter space. Motivated by permutation and orthogonal symmetries, we derive symmetry-compatible optimizers for different layer types, including linear layers, embedding and LM head matrices, and MoE routers. This yields one-sided spectral optimizers, row-norm optimizers, and hybrid row-norm/one-sided-spectral variants. Pre-training experiments on dense and sparse MoE language models, including Qwen3-0.6B-style, Gemma 3 1B-style, OLMoE-1B-7B-style, and downsized gpt-oss models, show that replacing AdamW on large vocabulary-indexed matrices and router matrices with symmetry-compatible optimizers consistently improves final validation loss and sometimes training stability.
Anymotion: A Dataset, Benchmark, and Baseline for Controllable Human Motion Editing
Haiyang Yan ⋅ Jianxin Sun ⋅ Yuhan Wu ⋅ libin wang ⋅ Zhiwei Li ⋅ Man Zhang ⋅ Jingdong Chen ⋅ DanDan Zheng
Human motion change is currently hindered by the lack of large-scale datasets, comprehensive benchmarks, and robust algorithms capable of handling diverse motion requirements. To address these gaps, we propose AnyMotion, a novel open-world paradigm for highly controllable and identity-consistent human motion editing across arbitrary scenes and actions. Our contributions are two-fold. First, we establish a hierarchical label taxonomy covering six major categories and 527 fine-grained classes, and curate HME-1M, a large-scale dataset containing 1.04 million editing quadruplets via systematic data balancing, cleaning, and reverse construction. Second, we introduce the Skeleton Chain-of-Thought (SCoT) framework, a two-stage pipeline consisting of motion reasoning and guided generation. In the reasoning stage, a multimodal large language model serves as a cognitive planner to derive target skeletal keypoints and MetaQueries. These outputs subsequently function as explicit structural and semantic priors, providing the necessary guidance for the diffusion transformer to perform high-fidelity motion generation. Finally, built upon our label taxonomy, we present Human Motion-Bench, the first motion-centric benchmark equipped with multi-category fine-grained evaluation metrics. Extensive experiments demonstrate that AnyMotion achieves state-of-the-art performance across diverse editing scenarios.
AppCIF-Bench: An Application-Level Complex Instruction Following Benchmark for Large Language Models
Jiahao Xu ⋅ Xuefang Zhao ⋅ Xinhua Feng ⋅ Hongrui Yang ⋅ Yixian Liu ⋅ Zhichao Hu ⋅ Lliu Yuhong
Complex Instruction Following (CIF) is a foundational capability for deploying Large Language Models (LLMs) in production-grade applications. It demands stable, precise outputs under increasing horizon length, deeply nested rules, and multi-turn interactions. Yet existing benchmarks are predominantly atomic and short-horizon, failing to capture the logical depth and distractor pressure of real-world settings. We introduce **AppCIF**, the first application-level CIF benchmark. AppCIF stratifies difficulty along two orthogonal axes: *Horizon Length* ($H$), the span over which distractors accumulate and critical signals must be tracked; and *Logical Depth* ($L$), the nesting depth of rules to be unfolded. Together, $\langle H, L\rangle$ defines four hierarchical complexity levels, scaling from clean few-step execution (L1) to extreme pressure deduction (L4). AppCIF comprises 466 curated samples across five application scenarios. It adopts a "**Prompt-as-Environment**" paradigm, pairing Complex System Prompts (averaging ~14k words) with executable Code Simulation Engines (averaging ~3.5k lines of Python) to sustain rigorous state-machine deduction. The benchmark is constructed via a **Labor-Divided Human-AI pipeline** with scenario-specific collaboration strategies, and evaluated on 12 representative LLMs. Results reveal sharp performance stratification: Hard Satisfaction Rate ranges from 1.1% to 45.9%, exposing severe bottlenecks in long-horizon, distractor-laden instruction execution.
Approximation of Maximally Monotone Operators : A Graph Convergence Perspective
Takashi Furuya ⋅ Yury Korolev ⋅ Takaharu Yaguchi
Operator learning has been highly successful for continuous mappings between infinite-dimensional spaces, such as PDE solution operators. However, many operators of interest—including differential operators—are discontinuous or set-valued, and lie outside classical approximation frameworks. We propose a paradigm shift by formulating approximation via graph convergence (Painlevé–Kuratowski convergence), which is well-suited for closed operators. We show that uniform and $L^p$ approximation are fundamentally inadequate in this setting. Focusing on maximally monotone operators, we prove that any such operator can be approximated in the sense of local graph convergence by continuous encoder–decoder architectures, and further construct structure-preserving approximations that retain maximal monotonicity via resolvent-based parameterizations.
Arbitrarily Conditioned Hierarchical Flows for Spatiotemporal Events
Keyan Chen ⋅ Qiwei Yuan ⋅ Zhitong Xu ⋅ Bin Shen ⋅ Shandian Zhe
Events in spatiotemporal systems are ubiquitous, yet modeling their complex distributions remains challenging. Existing point process models often rely on strong structural assumptions and are typically limited to autoregressive, event-by-event prediction. As a result, they struggle to support broader inference tasks such as inverse inference, trajectory reconstruction, and recovery of missing event locations. We introduce Arbitrarily Conditioned Hierarchical Flows (ARCH), a hierarchical flow matching framework for spatiotemporal event modeling. ARCH is expressive enough to capture complex event distributions while enabling tractable and accurate computation of conditional intensities, which quantify instantaneous event risk. Built on a history-encoder-generative-decoder architecture, ARCH introduces a hybrid masking strategy for flexible conditioning on arbitrary observed events. This enables a unified treatment of forecasting, inverse inference, and partial trajectory recovery within a single framework. Experiments on synthetic and real-world datasets show that ARCH consistently outperforms existing baselines across both prediction and conditional inference tasks.
A Reduction from Delayed to Immediate Feedback for Online Convex Optimization with Improved Guarantees
Alexander Ryabchenko ⋅ Idan Attias ⋅ Dan Roy
We develop a reduction-based framework for online learning with delayed feedback that recovers and improves upon existing results for both first-order and bandit convex optimization. Our approach introduces a continuous-time model under which regret decomposes into a delay-independent learning term and a delay-induced drift term, yielding a delay-adaptive reduction that converts any algorithm for online linear optimization into one that handles round-dependent delays. For bandit convex optimization, we significantly improve existing regret bounds, with delay-dependent terms matching optimal first-order rates. For first-order feedback, we recover optimal regret bounds via a simpler, unified analysis. Quantitatively, for bandit convex optimization we obtain $O(\sqrt{d_{\text{tot}}} + T^{\frac{3}{4}}\sqrt{k})$ regret, improving the delay-dependent term from $O(\min\{\sqrt{T d_{\text{max}}},(Td_{\text{tot}})^{\frac{1}{3}}\})$ in previous work to $O(\sqrt{d_{\text{tot}}})$. Here, $k$, $T$, $d_{\text{max}}$, and $d_{\text{tot}}$ denote the dimension, time horizon, maximum delay, and total delay, respectively. Under strong convexity, we achieve $O(\min\{\sigma_{\text{max}} \ln T, \sqrt{d_{\text{tot}}}\} + (T^2\ln T)^{\frac{1}{3}} {k}^{\frac{2}{3}})$, improving the delay-dependent term from $O(d_{\textnormal{max}} \ln T)$ in previous work to $O(\min\{\sigma_{\text{max}} \ln T, \sqrt{d_{\text{tot}}}\})$, where $\sigma_{\text{max}}$ denotes the maximum number of outstanding observations and may be considerably smaller than $d_{\text{max}}$.
ARES: How Reliable Are LLM User Simulators for Recommender A/B Testing?
Hongyang Su ⋅ Beibei Kong ⋅ Lei Cheng ⋅ Chengxiang Zhuo ⋅ Zang Li ⋅ Chenyun YU
LLM-driven user simulators are increasingly adopted for A/B-testing recommender systems, yet the choice of backbone is treated as an interchangeable implementation detail rather than a factor shaping evaluation reliability. We show otherwise: across nine backbones under unified controls, this choice accounts for 149% as much outcome variance as the tested recommender, and swapping backbones alone flips about 20% of pairwise A/B winners. The backbone, in other words, is not a neutral component but the measurement instrument. We turn this observation into a method: ARES (Audit for Reliable Evaluation of Simulators) separates a backbone effect (the instrument's contribution) from a recommender effect (the signal under study), quantifies both at three levels (outcomes, trajectories, and reasoning), stress-tests them against text vs. rendered UI, and endorses only A/B claims a majority of backbones independently support. Applied to nine backbones $\times$ five recommenders on MovieLens-1M, ARES surfaces failures invisible to outcome metrics: within a single model lineage, newer and stronger releases do not necessarily agree more with their predecessors (Kendall's $\tau$ drops from 1.0 to 0.20 across generations); three of eight vision-capable backbones reverse the declared best recommender under UI rendering; and 63000 reasoning traces tie these reversals to interpretable behavioral biases, not stochastic noise. Capability alone, therefore, does not guarantee reliability. While instantiated for recommendation, the backbone-as-instrument view generalizes to any LLM-mediated evaluation where model identity is a free variable, including LLM-as-a-Judge. We release ARES-Bench at https://github.com/neurips2026-ares-authors/ARES-Bench: 21400 behavioral logs, 63000 reasoning traces, a reusable visual sandbox, and an analysis toolkit, making multi-backbone auditing a feasible default for simulator-based evaluation.
A Retained-Signal Interface for LLM Watermark Robustness under Paraphrase
Chen Wang ⋅ Yinxuan Huang ⋅ Yexin Cui cui ⋅ Maoqing Zhong ⋅ LaiLong Luo
Watermark robustness under paraphrase is usually reported as AUC against a named rewriter, making results hard to compare across watermark schemes, paraphrasers, and detectors. We propose a retained-signal reporting interface: if a watermark injects KL signal $\Delta=\KL(P_1\Vert P_0)$ and a paraphrase channel retains KL fraction $\rho$, then any detector's post-paraphrase advantage is controlled by $\Delta\rho$. We give explicit Pinsker and Bretagnolle--Huber threshold laws, prove the $\Theta(\sqrt{\Delta\rho})$ rate is sharp for any $(\Delta,\rho)$-only converse, and show the bound is implementation-blind once watermark families are matched by $\Delta$. For real LLM paraphrases, we instantiate the interface as a score-projected audit that estimates $\widehat{\Delta}\widehat{\rho}$ for a deployed detector and reports it with semantic preservation and the detector projection. In a KGW/Unigram $\times$ Qwen/Llama benchmark, this score-effective signal predicts post-paraphrase AUC across lengths, strengths, and paraphraser families. The resulting protocol separates injected signal, retained signal, semantic quality, and detector choice, replacing one-off AUC tables with a reproducible robustness report.
ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation
Zijie Meng ⋅ Jiwen Liu ⋅ Yufei Liu ⋅ Chengzhuo Tong ⋅ Xiaoqiang Liu ⋅ Yuanxing Zhang ⋅ Yulong Xu ⋅ Pengfei Wan
Identity-preserving video generation is fundamentally constrained by the prevailing point-reference paradigm: existing methods consume one or a few snapshot images and therefore cannot transcend "look-alike" fidelity, suffer from severe drift under large yaw, occlusion or scale variation, and trade brittle cross-pair supervision against pervasive reference leakage. We argue that identity is intrinsically a spatially-distributed, multi-view concept and propose Argus, a theoretically-grounded framework that lifts identity injection from single-snapshot conditioning to multi-view aggregated reasoning. Argus comprises three tightly-coupled modules. (i) An Identity Director built on a frozen multimodal LLM automatically curates nine semantically-diverse keyframes, extracts a global identity descriptor, and rewrites the prompt to dissolve the prompt-frame-identity tri-conflict. (ii) A Stacked Mosaic Identity Injection packs the nine SMPL-localized crops into a 3×3 mosaic and feeds it through Wan 2.1's native 36-channel patch embedding via a flow-matching-synchronized prefix-token stream, while a Negative-Time-Anchored RoPE preserves architectural compatibility with Wan's split-axis design. (iii) An Adaptive Skip-Layer Guidance replaces hand-tuned PAG with a tiny mask network that learns where to perturb, decoupling identity guidance from textual guidance, and a Temporal Identity Annealing schedule injects identity only during the structure-formation window to obviate copy-paste artifacts. Trained without any cross-pair data, Argus attains state-of-the-art performance on OpenS2V-Eval, surpassing both closed-source (Hailuo) and open-source (VACE-14B, Phantom-14B, Stand-In, BindWeave) systems on five of seven metrics. To stress-test extreme failure modes, we additionally release HardID-Celeb, a celebrity benchmark with two new robustness probes (YawScore, OccScore) on which Argus establishes the highest scores by a wide margin. A 47-participant double-blind user study confirms perceptual superiority across both expert and lay populations.
ARROW: Augmented Replay for RObust World models
Abdulaziz Alyahya ⋅ Abdallah Al Siyabi ⋅ Markus R. Ernst ⋅ Luke Yang ⋅ Levin Kuhlmann ⋅ Gideon Kowadlo
Continual reinforcement learning challenges agents to acquire new skills while retaining previously learned ones with the goal of improving performance in both past and future tasks. Most existing approaches rely on model-free methods with replay buffers to mitigate catastrophic forgetting; however, these solutions often face significant scalability challenges due to large memory demands. Drawing inspiration from neuroscience, where the brain replays experiences to a predictive World Model rather than directly to the policy, we presentARROW (Augmented Replay for RObust World models), a model-based continual RL algorithm that extends DreamerV3 with a memory-efficient, distribution-matching replay buffer. Unlike standard fixed-size FIFO buffers, ARROW maintains two complementary buffers: a short-term buffer for recent experiences and a long-term buffer that preserves task diversity through intelligent sampling. We evaluate ARROW on two challenging continual RL settings: Tasks without shared structure (Atari), and tasks with shared structure, where knowledge transfer is possible (Procgen CoinRun variants). Compared to model-free and model-based baselines with replay buffers of the same-size, ARROW demonstrates substantially less forgetting on tasks without shared structure, while maintaining comparable forward transfer. Our findings highlight the potential of model-based RL and bio-inspired approaches for continual reinforcement learning, warranting further research. Code is available at https://github.com/Cerenaut/ARROW.
ART for Diffusion Sampling: A Reinforcement Learning Approach to Timestep Scheduling
Yilie Huang ⋅ Wenpin Tang ⋅ XUNYU ZHOU
We consider time discretization for score-based diffusion models to generate samples from a learned reverse-time dynamic on a finite grid. Uniform and hand-crafted grids can be suboptimal given a budget on the number of time steps. We introduce Adaptive Reparameterized Time (ART), which controls the clock speed of a reparameterized time variable to redistribute computation along the sampling trajectory while preserving the terminal time, with the objective of minimizing the aggregate Euler discretization error. We derive a randomized companion ART-RL that recasts ART as a continuous-time reinforcement learning problem with Gaussian policies, and prove a two-directional bridge between the two: the deterministic ART optimum lifts to an optimal Gaussian policy, and conversely any optimal Gaussian policy must recover the ART control through its mean. This bridge turns continuous-time actor--critic learning into a principled, rather than heuristic, route to the deterministic timestep optimum. Within the official EDM pipeline, ART-RL improves FID on CIFAR--10 across a wide range of budgets; after one-time offline training, the distilled deterministic schedule transfers without retraining to AFHQv2, FFHQ, and ImageNet at no extra inference cost.
Artificial Aphasias in Lesioned Language Models
Nathan Roll ⋅ Jill Kries ⋅ Laura Gwilliams ⋅ Cory Shain
Aphasias, selective language impairments which can arise from brain damage, reveal the functional organization of human language by providing causal links between affected brain regions and specific symptom profiles. Drawing on this literature, we introduce an aphasia-inspired technique to characterize the emergent functional organization of language models (LMs). We ``lesion'' (zero-out) model parameters and measure the effects of this intervention against clinical aphasia symptoms, as diagnosed by the Text Aphasia Battery (TAB). When applied to 112,426 outputs from five 1B-scale LMs, the full range of evaluated symptoms surface, but in distributions largely distinct from those of humans. Our method uncovers broad, component-level symptom differences between attention (query, key, value, output) and feed-forward (up, gate, down) projections, without any significant within-class differences. We also find an effect of depth, where lesions in early layers disproportionately cause syntactic and semantic symptoms while late-middle layers yield higher rates of phonological and fluency deficits. Although some LM lesions induce quantitatively more similar profiles to some human aphasia types than others, qualitative differences in symptom patterns between LMs and humans suggest that aphasia syndromes are heavily influenced by the details of learning and processing rather than being a domain-invariant consequence of disrupted language processing.
ASAP: Fast Adaptive Sliding Agnostic Poisoning Attack on Federated Learning
HUAZI PAN ⋅ Yanjun Zhang ⋅ Leo Yu Zhang ⋅ Scott D Adams ⋅ Abbas Z Kouzani ⋅ Suiyang Khoo
Federated Learning (FL) is vulnerable to model poisoning attacks, where malicious clients manipulate uploaded model updates to corrupt the global training process. Existing attacks are typically evaluated by their eventual damage, such as final accuracy degradation or convergence to random-guess performance, where the learned model becomes unusable and may ultimately lead to denial-of-service (DoS). However, such metrics overlook attack round complexity, defined as the number of communication rounds required for the poisoned global model to reach and remain near a desired degradation objective. In realistic FL systems, requiring more communication rounds increases the adversary's exposure to client participation, robust aggregation, filtering, and detection. We propose ASAP (Adaptive Sliding Agnostic Poisoning), a fast aggregation-rule-agnostic (AGR-agnostic) attack that treats the whole FL training process as an uncertain dynamical system and formulates poisoning as a feedback-control problem for steering the model toward a desired attack objective. ASAP combines adaptive sliding mode control (ASMC) with finite Fourier-basis uncertainty estimation, treating the unknown effects of benign training and aggregation as a time-varying uncertainty and using the resulting estimate to design the malicious updates. We theoretically prove that the sliding variable reaches the sliding surface in finite time and the tracking error subsequently converges exponentially. Experiments across multiple datasets, models, and aggregation rules demonstrate that ASAP reaches specified degradation objectives in fewer communication rounds while maintaining smaller target deviation than existing model poisoning attacks.
AsdaKV: Attention-Overlap Driven Semantic KV Retrieval for Long-Context LLMs
Tianming Yan ⋅ Kun Yang ⋅ Tingwang You ⋅ Kui Ren
Long-context inference with large language models is bottlenecked by the linear growth of the KV cache. Existing eviction-based methods reduce memory overhead by irreversibly discarding historical states, while retrieval-based methods preserve the full context but rely on retrieval units defined by tokens, fixed pages, or proxy signals including query-key similarity, clustering, and linguistic rules. Such externally defined units often misalign with the model's actual contextual dependencies during generation. We propose AsdaKV, an adaptive semantic drift-aware KV cache retrieval framework that defines semantic continuity directly from the model's inherent attention behavior. AsdaKV detects semantic drift by measuring the overlap between high-attention historical regions across decoding steps: stable overlap indicates that the active cache still covers the relevant contextual evidence, whereas a sharp overlap drop triggers cache refresh. Guided by this drift signal, AsdaKV organizes historical KV states into variable-length semantic windows, offloads full window KV states to CPU memory, and maintains only lightweight window anchors plus a fixed-budget active KV cache on GPU. During decoding, the current query retrieves relevant windows via anchor matching, while adaptive-stride drift detection and deferred page recall reduce redundant detection, retrieval, and KV reloading overhead. Extensive experiments across diverse tasks and model families demonstrate that under the same KV budget, AsdaKV achieves higher accuracy than state-of-the-art KV retrieval methods, maintains comparable decoding efficiency in long-input settings, and further outperforms existing retrieval systems for long-output generation.
ASH: Agents that Self-Hone via Embodied Learning
Benjamin Schneider ⋅ Xavier Schneider ⋅ Victor Zhong ⋅ Sun Sun
Long-horizon embodied tasks remain a fundamental challenge in AI, as current methods rely on hand-engineered rewards or action-labeled demonstrations, neither of which scales. We introduce ASH, an agentic system that learns an embodied policy from unlabeled, noisy internet video, without reward shaping or expert annotation. ASH follows a self-improvement loop; when it gets stuck, ASH learns an Inverse Dynamics Model (IDM) from its own trajectories, and uses its IDM to extract supervision from relevant internet video. ASH uses unsupervised learning to identify key moments from large-scale internet video and retains them as long-term memory --- allowing it to tackle long-horizon problems. We evaluate ASH on two complementary environments demanding multi-hour planning: Pokemon Emerald, a turn-based RPG, and The Legend of Zelda: The Minish Cap, a real-time action-adventure game. In both games, behavioral cloning, retrieval-augmented and zero-shot foundation-model baselines plateau, while ASH sustains progression across our 8-hour evaluation. ASH reaches an average of $11.2/12$ milestones in Pokemon Emerald and $9.9/12$ in Legend of Zelda, while the strongest baseline gets stuck in both environments at an average of $6.5/12$ and $6.0/12$ milestones, respectively. We demonstrate that self-improving agents are a scalable recipe for long-horizon embodied learning.
Ask-E: An Environment for Calibrated Question Generation
Sarah Pratt ⋅ Jae Sung Park ⋅ Scott Geng ⋅ Ali Farhadi
Today, we improve models by training and evaluating them on problems at the frontier of their abilities. Creating such problems is itself a demanding task, requiring the ability to probe model limits and generalize beyond existing question distributions. It also means placing problems at a precise difficulty level, which requires understanding what it takes to solve them. In short, generating problems calibrated to a model's current frontier demands capability beyond it, an increasingly burdensome constraint as models improve. Our key insight is that we can leverage this constraint to our advantage: a model that can generate problems consistently calibrated to a given frontier must possess capability beyond it. Accordingly, we present Ask-E, an environment that benchmarks and trains models on their ability to write questions at a given skill level, rather than answer them. Concretely, we define target skill levels as ranges bounded by the capabilities of two existing language models. A generated question is successfully calibrated if exactly one of the two models can solve it, placing it precisely within the target range and differentiating the capabilities of these models. Ask-E serves both as a benchmark and a training environment, where models generate problems calibrated to a variety of skill levels. We find that even frontier models achieve below 50\% calibration on the benchmark, leaving significant headroom to measure future progress. We also show that training on this environment leads to improvements across a number of downstream math benchmarks even with no new math data, no interaction with stronger models, and no correctness-based reward.
Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents?
Anmol Gulati ⋅ Hariom Gupta ⋅ Elias Lumer ⋅ Sahil Sen ⋅ Vamse K Subbiah
Long-horizon AI agents execute complex workflows spanning hundreds of sequential actions, yet a single wrong assumption early on can cascade into irreversible errors. When instructions are incomplete, the agent must decide not only whether to ask for clarification but when, and no prior work measures how clarification value changes over the course of execution. We introduce a forced-injection framework that provides ground-truth clarifications at controlled points in the agent's trajectory across four information dimensions (goal, input, constraint, context), three agent benchmarks, and four frontier models (three per benchmark; one on a single benchmark only; 84 task variants; 6,000+ runs). Counter to the common intuition that ``earlier is always better,'' we find that the value of clarification depends sharply on what information is missing: goal clarification loses nearly all value after 10% of execution (pass@3 drops from 0.78 to baseline), while input clarification retains value through roughly 50%. Deferring any clarification type past mid-trajectory degrades performance below never asking at all. Cross-model Kendall τ correlations (0.78-0.87 among models sharing identical task coverage; 0.34-0.67 across the full 4-model panel) confirm these timing profiles are substantially task-intrinsic. A complementary study of 300 unscripted sessions reveals that no current frontier model asks within the empirically optimal window, with strategies ranging from over-asking (52% of sessions) to never asking at all. These empirical demand curves provide the quantitative foundation that existing theoretical frameworks require but have lacked, and establish concrete design targets for timing-aware clarification policies. Code and data will be publicly released.
Ask in the Crowd: Differentially Private LLM Inference via Dummy-Augmented Shuffling
Zhihao Liu ⋅ Zixiong Guo ⋅ Shuo Shao ⋅ Yu He ⋅ Wenli Wang ⋅ Meihui Chen ⋅ Di Wang ⋅ Yuke Hu
Large language models are increasingly deployed as LLM-as-a-Service (LLMaaS), where user queries are sent to remote LLMs for inference, raising severe privacy concerns about the exposure of sensitive user inputs. Differentially private inference mitigates these risks by injecting noise into input embeddings. However, existing methods struggle to balance privacy and utility, as stronger privacy protection often leads to substantial accuracy degradation in LLM inference. In this work, we propose DP-SPD, a differentially private LLM inference framework that protects user queries by hiding the real query among client-generated dummy queries. All queries are randomly shuffled before being sent to the server, preventing direct identification of the user's true intent. To improve inference efficiency, we exploit the shared-prefix structure among queries and reuse transformer key--value (KV) caches during decoding. We analyze the privacy guarantee of DP-SPD and evaluate its utility on classification and generation tasks, as well as its privacy protection under EIA, MIA, and GPT-based attacks. Experimental results show that our method achieves strong inference-time privacy protection while maintaining utility close to plain-text inference under practical latency.
A Stability Analysis of AdamW: Unstable Equilibria and Non-Convergence
Baiyu Su ⋅ Lizhang Chen ⋅ Qiang Liu
AdamW is ubiquitous in deep learning, yet its behavior remains poorly understood. We analyze its dynamics through the lens of dynamical systems and show that AdamW admits an \emph{implicit fixed-point objective}: its fixed points coincide with the stationary points of a constrained and regularized optimization problem. However, not all of these fixed points are stable under AdamW’s dynamics, and stability depends sensitively on curvature, weight decay, and momentum parameters. Even in simple one-dimensional settings, AdamW can exhibit surprisingly complex behavior: equilibria may be unstable, and numerical trajectories can exhibit persistent oscillations rather than convergence to them. We further extend the analysis to higher dimensions, deriving sufficient conditions for local stability of the continuous-time dynamics, and propose a convergent modification that preserves AdamW's fixed points. These results clarify what optimization problem AdamW is associated with, when its convergence can be expected, and how its dynamics could inspire more reliable optimizers.
A Statistical Framework for Algorithmic Collective Action with Multiple Collectives
Claudio Battiloro ⋅ Pietro Greiner ⋅ Dario Rancati ⋅ Bret Nestor ⋅ Oumaima Amezgar ⋅ Francesca Dominici
As learning systems increasingly shape everyday decisions, Algorithmic Collective Action (ACA), i.e., users coordinating changes to shared data to steer model behavior, offers a complement to regulator-side policy and corporate model design. Real-world collective actions have traditionally been decentralized and fragmented into multiple collectives, despite sharing overarching objectives, with each collective differing in size, strategy, and actionable goals. However, most of the ACA literature focuses on single collective settings. To address this, we propose the first comprehensive statistical framework for ACA with multiple collectives acting on the same system. In particular, we focus on collective action in classification, studying how multiple collectives can influence a classifier's behavior. We provide quantitative statistical bounds on the success of the collectives, considering the role and the interplay of the collectives' sizes and the alignment of their goals. We make such bounds computable by each collective with only partial knowledge of other collectives' sizes and strategies. Finally, we numerically illustrate our framework on simulations inspired by interventions for climate adaptation in smart cities, demonstrating the usefulness of our bounds.
A Subgoal-driven RL Framework for Improving Long-Horizon Web Agents
Taiyi Wang ⋅ Sian Gooding ⋅ Florian Hartmann ⋅ Oriana Riva ⋅ Edward Grefenstette
Large language model (LLM)-based agents have emerged as powerful autonomous controllers for digital environments, spanning mobile interfaces, operating systems, and web browsers. Web navigation, for example, demands handling dynamic content and long action sequences, making it a particularly complex task. Existing LLM-backed agents exhibit weakened long-horizon planning abilities during RL fine-tuning, where sparse and delayed rewards from terminal Outcome Reward Models (ORMs) make it difficult for agents to identify the actions that lead to success, preventing them from sustaining coherent reasoning over extended tasks. We address this with two contributions: (1) a validated subgoal generation procedure that produces reliable, monotonically calibrated progress signals between the initial state and the terminal ORM reward; and (2) MiRA ($\underline{Mi}$lestoning your $\underline{R}$einforcement Learning Enhanced $\underline{A}$gent), an RL training framework using dense, milestone-based reward signals via a learned potential critic. Starting from the open Gemma3-12B base ($6.4$%), MiRA enables the agent to reach $43.0$% on the WebArena-Lite benchmark. This performance surpasses proprietary systems such as GPT-4-Turbo ($17.6$%) and GPT-4o ($13.9$%), as well as the previous open-model state of the art, WebRL ($38.8$%). Our findings demonstrate that milestone-based reward shaping significantly boosts an agent's long-horizon abilities, paving the way for more robust, general-purpose autonomous systems.
A Surrogate Perspective on Convergence of Fixed-Target DQN
Zichu Liu ⋅ Nneka M Okolo ⋅ Ryan D'Orazio ⋅ Danilo Vucetic ⋅ Ioannis Mitliagkas ⋅ Gauthier Gidel
Deep Q-Networks (DQNs) approximate value iteration by freezing a target network, forming Bellman-optimality targets, and running an inner regression loop that minimizes a smooth \emph{surrogate loss} (e.g., MSE/Huber loss) on data from a replay buffer. Yet, control is governed by the Bellman optimality operator's $\gamma$-contraction in $\ell_\infty$, whereas common objectives (MSE/Huber loss) measure progress in an averaged $\ell_2/\ell_1$ geometry; thus, improving the surrogate need not reduce the worst-case Bellman error that drives performance. To bridge this gap we connect the progress in the inner loop for each surrogate to the Bellman residual under the $\ell_\infty$ norm. Our analysis provides explicit thresholds denoted by $\alpha$ under which sufficient progress guarantees a $\ell_\infty$ contraction up to a standard approximation floor. This yields new convergence guarantees for DQN with fixed targets trained using MSE or Huber-type losses, and motivates a novel soft-$\ell_\infty$ surrogate that smoothly approximates the sup-norm and better matches the $\ell_\infty$ geometry resulting in the least stringent inner loop accuracy requirements. Across Atari benchmarks, soft-$\ell_\infty$ consistently reduces the worst-case Bellman residual and matches or improves returns relative to standard objectives, providing a practical, geometry-aligned recipe for stable DQN training.
AsymPipe: Accelerating Large-scale DiTs Image Editing via Asymmetry-Aware Pipeline Parallelism
Huage Deng ⋅ LYU WENKAI ⋅ WU YONGJIN ⋅ Tianyang Zheng ⋅ heng Guo ⋅ Fangyun Zhou ⋅ Yu Shi ⋅ Wenxi Zhu ⋅ Minwen Deng
Instruction-driven mask-free image editing with large-scale Diffusion Transformers (DiTs) has enabled high-fidelity local edits while maintaining global consistency. However, these models incur a prohibitive inference cost due to a pronounced semantic-–computation mismatch, where dense computation is performed over all high-resolution tokens regardless of localized editing instructions. Our analysis reveals that non-edit regions converge early with highly stable block-level outputs, exhibiting a distinct spatio-temporal asymmetry. Furthermore, we identify a critical system-level challenge that introducing such asymmetric computation into pipeline-parallel environments for ultra-large models triggers severe workload imbalance and pipeline bubbles, preventing algorithmic savings from translating into actual latency gains. To address these issues, we propose AsymPipe, a training--free algorithm–-system co-design framework that accelerates large-scale DiTs editing. The framework identifies edit versus non-edit regions on the fly using a zero-overhead velocity-based probe that reuses the predicted velocity field from the existing denoising process. For non-edit regions, AsymPipe caches and reuses block-level outputs, while for edit regions, it introduces a source-aware sparse attention module to precisely eliminate redundant contributions from static text and image sources. Finally, to bridge the gap between algorithmic optimization and hardware speedup, the framework incorporates a load-balanced token rearrangement scheme. This scheme deconstructs the original spatial topology to distribute tokens evenly across devices, physically restoring compute balance in pipeline-parallel systems. On HunyuanImage-3.0-Instruct and Qwen-Image-Edit-2511, AsymPipe achieves end-to-end speedups of 2.862x (PSNR 33.69) and 2.707x (PSNR 29.74), respectively, delivering substantial acceleration with virtually no loss in editing quality. Our anonymized source code is available at https://anonymous.4open.science/r/AsymPipe-6553.
Asymptotically Log-Optimal Bayes-Assisted Confidence Sequences for Bounded Means
Valentin Kilian ⋅ Stefano Cortinovis ⋅ Francois Caron
Confidence sequences based on test martingales provide time-uniform uncertainty quantification for the mean of bounded IID observations without parametric distributional assumptions. Their practical efficiency, however, depends strongly on the choice of martingale updates, and many existing constructions do not exploit prior information about plausible data-generating distributions or mean values. We propose a Bayes-assisted framework that uses a Bayesian working predictive model to adaptively construct confidence sequences. For each candidate mean and time point, the predictive distribution selects, among valid one-step martingale factors, the update maximising predictive expected log-growth; validity is therefore preserved even when the prior or working model is misspecified. We prove that if the predictive distribution is Wasserstein-consistent, the resulting procedure is asymptotically log-optimal, matching the per-sample log-growth of an oracle procedure with access to the true distribution. We instantiate the framework using robust predictives based on Dirichlet-process mixtures and Bayesian exponentially tilted empirical likelihood. Experiments on synthetic data, sequential best-arm identification for LLM evaluation, and prediction-powered inference show that informative priors can substantially reduce confidence-sequence width and sampling effort while retaining anytime-valid coverage.
AsyncOPD: How Stale Can On-Policy Distillation Be?
Wonjun Kang ⋅ Kevin Galim ⋅ Seunghyuk Oh ⋅ Minjun Kang ⋅ Sanghyun Park ⋅ Donghoon Kim ⋅ Minjae Lee ⋅ Minseo Kim ⋅ Rishabh Tiwari ⋅ Yuchen Zeng ⋅ HYUNG IL KOO ⋅ Kangwook Lee
On-policy distillation (OPD) is an increasingly important way to post-train large language models (LLMs), but, like reinforcement learning (RL), it relies on rollouts from the model being optimized. For reasoning workloads, rollout generation can dominate training time. Asynchronous RL alleviates this bottleneck by decoupling rollout generation from learner updates, but doing so introduces stale-policy data; prior work studies how to stabilize learning from such data. However, it remains underexplored whether these asynchronous RL ideas and stale-data solutions transfer to OPD, and what OPD-specific constraints arise. To address this gap, we provide the first systematic study of staleness in asynchronous OPD. We first show that KL direction changes the stale-data problem: teacher-weighted forward KL is robust to stale rollouts, whereas student-weighted reverse KL is vulnerable. Second, for this vulnerable reverse-KL case, we study whether methods designed to stabilize asynchronous RL can mitigate OPD staleness. We find that they do not improve over a simpler OPD-specific surrogate: recomputing the reverse-KL signal under the current student at learner time without clipping. Third, we identify an OPD-specific teacher-cache constraint: under asynchronous execution, teacher scores are available at learner time only on cached actions. The resulting bias-variance tradeoff for sparse and sampled reverse-KL OPD implementations motivates multi-sample Monte Carlo (MC), which preserves MC correctability while reducing one-sample variance. Finally, we present and open-source **AsyncOPD**, a fully asynchronous OPD training pipeline built from these estimator choices. Experiments show that AsyncOPD improves training throughput by $1.6\times$ to $3.8\times$ over strict synchronous training while reaching comparable accuracy.
A Systematic Evaluation of Co-folding Model Representations for Small-Molecule Learning
Hyosoon Jang ⋅ Hyunjin Seo ⋅ Honghui Kim ⋅ Taewon Kim ⋅ Yunhui Jang ⋅ Seonghyun Park ⋅ Sungsoo Ahn
Small-molecule foundation models are typically pretrained on standalone molecular data, unlike vision and language models that often benefit from cross-modal or relational supervision. Protein-ligand co-folding provides a molecular analogue of such supervision by exposing models to atom-level ligand-protein interactions, raising the question of whether co-folding models can yield strong small-molecule representations. We study this question using Boltz2, a modern co-folding model, by transferring its atom-level ligand representations to standalone small-molecule tasks. Through systematic probing and distillation, we show that Boltz2 representations match or outperform existing models on the ADMET benchmark, accelerate molecular generative modeling, and improve sample efficiency in structure-guided ligand optimization. We further find that Boltz2 representations are complementary to those learned from conventional standalone molecular supervision, including 3D conformers, bioassay labels, and quantum-chemical properties. Finally, we extend representation alignment to reinforcement learning, showing that dense representation-level supervision can complement scalar rewards in molecular discovery. These results identify protein-ligand co-folding as a promising pretraining paradigm for small-molecule representation learning and position Boltz2 as a strong, off-the-shelf molecular foundation model.
At FullTilt: Real-Time Open-Set 3D Macromolecule Detection Directly from Tilted 2D Projections
Ming-Yang Ho ⋅ Alberto Bartesaghi
Open-set 3D macromolecule detection in cryogenic electron tomography eliminates the need for target-specific model retraining. However, strict VRAM constraints prohibit processing an entire 3D tomogram, forcing current methods to rely on slow sliding-window inference over extracted subvolumes. To overcome this, we propose FullTilt, an end-to-end framework that redefines 3D detection by operating directly on aligned 2D tilt-series. Because a tilt-series contains significantly fewer images than slices in a reconstructed tomogram, FullTilt eliminates redundant volumetric computation, accelerating inference by orders of magnitude. To process the entire tilt-series simultaneously, we introduce a tilt-series encoder to efficiently fuse cross-view information. We further propose a multiclass visual prompt encoder for flexible prompting, a tilt-aware query initializer to effectively anchor 3D queries, and an auxiliary geometric primitives module to enhance the model's understanding of multi-view geometry while improving robustness to adverse imaging artifacts. Extensive evaluations on three real-world datasets demonstrate that FullTilt achieves state-of-the-art zero-shot performance while drastically reducing runtime and VRAM requirements, paving the way for rapid, large-scale visual proteomics analysis. All code and data will be publicly available upon publication.
ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both
Ziyu Guo ⋅ Yu Liu ⋅ Xinyan Chen ⋅ Pheng-Ann Heng
Visual reasoning, often interleaved with intermediate visual states, has emerged as a promising direction in the field. A straightforward approach is to directly generate images via unified models during reasoning, but this is computationally expensive and architecturally non-trivial. Recent alternatives include agentic reasoning through code or tool calls, and latent reasoning with learnable hidden embeddings. However, agentic methods incur context-switching latency from external execution, while latent methods lack task generalization and are difficult to train with autoregressive parallelization. To combine their strengths while mitigating their limitations, we propose ATLAS, a framework in which a single discrete "word", termed as a functional token, serves both as an agentic operation and a latent visual reasoning unit. Each functional token is associated with an internalized visual operation, yet remains a standard token in the tokenizer vocabulary and can be generated via next-token prediction. This design avoids verbose intermediate visual content generation, while preserving compatibility with standard SFT and RL training, without architectural or methodological modifications. To further address the sparsity of functional tokens during RL, we introduce Latent-Anchored GRPO (LA-GRPO), which stabilizes the training by anchoring functional tokens with a statically weighted auxiliary objective, providing stronger gradient updates. Extensive experiments and analyses demonstrate that ATLAS achieves superior performance on challenging benchmarks while maintaining clear interpretability. We hope ATLAS offers a new paradigm inspiring future visual reasoning research.
A Trust Region Approach for Learning Schrödinger Bridges
Takeshi Koshizuka ⋅ Denis Blessing ⋅ Max Zimmer ⋅ Sebastian Pokutta ⋅ Lorenz Richter
We study the numerical approximation of Schrödinger Bridges in the setting where samples from the target distribution are unavailable and only an unnormalized target density is given. In this regime, the widely used iterative proportional fitting (IPF) procedure (and its variants) often exhibits severe numerical instabilities, stemming from ill-conditioned underlying control problems and from the fact that the optimal solution may lie far from the initialization. To address these issues, we propose a trust region variant of IPF that constrains successive updates to remain at a fixed Kullback-Leibler divergence from the current iterate. This restriction significantly improves numerical stability. Moreover, we show that the resulting method admits an interpretation as a mirror descent scheme with an adaptively chosen step size. We demonstrate the effectiveness of our approach on several challenging sampling tasks.
Attribution-Guided Shared-Private Decoupling for Noise-Reduced Audio-Visual Representation Learning
Linge Wang ⋅ Yingying Chen ⋅ Bingke Zhu ⋅ Lu Zhou ⋅ Haixin Wang ⋅ Zhipiao Liu ⋅ Jinqiao Wang
Audio-visual representation learning commonly aligns paired audio and visual signals with sample-level contrastive objectives. However, real-world audio-visual pairs usually contain heterogeneous patch tokens: only a subset carries reliable cross-modal shared semantics, while many others mainly encode modality-private residuals. Aggregating all tokens into a single global representation can therefore introduce semantic noise into cross-modal alignment. Motivated by the minimal-sufficiency principle, we propose an attribution-guided shared-private decoupling framework to mitigate this issue. Our key idea is to identify shared-cover subviews that preserve dominant cross-modal information while suppressing modality-private residuals, and use them to supervise decoupled shared and private representations. To identify such subviews efficiently, we draw on transformer attribution method and introduce Layer-Truncated Attribution (LTA), which provides a lightweight estimate of token-level relevance to cross-modal shared information. We further adopt a teacher-student framework, where the teacher provides attribution-induced supervision and the student internalizes shared-private decoupling into its representations, without introducing extra inference cost. Experiments on retrieval, classification, and sound-prompted semantic segmentation show consistent gains over strong baselines. Compared with baseline, our method improves the average zero-shot R@1 by 4.0 across AudioSet and VGGSound retrieval, and improves AS20K classification mAP by 3.1, demonstrating the effectiveness of fine grained attribution guidance for reducing semantic noise.
Attributions All the Way Down? The Metagame of Interpretability
Hubert Baniecki ⋅ Przemyslaw 'Prem' Biecek ⋅ Fabian Fumagalli
We introduce the metagame, a conceptual framework for quantifying second-order interaction effects of model explanations. For any first-order attribution $\phi(f)$ explaining a model $f$, we measure the directional influence of feature $j$ on the attribution of feature $i$, denoted as meta-attribution $\varphi_{j \to i}(f)$, by treating the attribution method itself as a cooperative game and computing its Shapley value. Theoretically, we prove that attributions hierarchically decompose into meta-attributions, and establish these as directional extensions of existing interaction indices. Empirically, we demonstrate that the metagame delivers insights across diverse interpretability applications: (i) quantifying token interactions in instruction-tuned language models, (ii) explaining cross-modal similarity in vision-language encoders, and (iii) interpreting text-to-image concepts in multimodal diffusion transformers.
AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation
Huadai Liu ⋅ Kaicheng Luo ⋅ Wen Wang ⋅ Qian Chen ⋅ Bin Ma ⋅ Xiangang Li ⋅ Wei Xue
Unifying speech, sound, and music generation in one model is hindered by tradeoffs between fidelity, end-to-end training, in-context conditioning, and variable-length synthesis that no current paradigm fully resolves. To address this challenge, we present AudioCALM, a universal audio generation framework that extends autoregressive (AR) next-token prediction from discrete tokens to continuous audio latents: a thin flow-matching head replaces the softmax to predict rectified-flow velocities at each position, and a block-causal AR-Flow attention pattern produces arbitrary-length output. Joint training of multiple audio generation tasks faces an asymmetric text--audio mismatch: speech transcripts align to specific time spans and demand \textit{tight, time-aligned} attention, whereas sound and music captions describe only overall semantics and rely on diffuse, holistic attention; mixing the two disproportionately degrades sound and music generation. We address this asymmetry at two levels: a \textbf{data reformulation} strategy that unifies all three tasks under a single description-style conditioning interface, and a novel architecture Asymmetric Mixture-of-Modality-Experts (A-MoME), which adds a dedicated residual expert for speech while sound and music share the backbone, incurring no inference overhead on non-speech inputs. Experimental results demonstrate that AudioCALM matches modality-specific state-of-the-art and outperforms prior unified baselines on speech, sound, and music generation benchmarks.
Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos
Sreyan Ghosh ⋅ Arushi Goel ⋅ Kaousheik Jayakumar ⋅ Lasha Koroshinadze ⋅ Nishit Anand ⋅ Siddharth Gururani ⋅ Hanrong Ye ⋅ Pritam Biswas ⋅ Yuanhang Su ⋅ Ehsan Hosseini-Asl ⋅ Sang-gil Lee ⋅ Zhifeng Kong ⋅ Jaehyeon Kim ⋅ Sungwon Kim ⋅ S Sakshi ⋅ Ramani Duraiswami ⋅ Dinesh Manocha ⋅ Andrew Tao ⋅ Mohammad Shoeybi ⋅ Bryan Catanzaro ⋅ Ming-Yu Liu ⋅ Wei Ping
We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over long and complex real-world (audio-visual) videos. To support this, we make three key contributions: (i) Audio-Visual-Skills, a large-scale collection of real-world videos with 7M caption and question-answer training instances designed to emphasize temporal, compositional, and cross-modal audio-visual reasoning; (ii) a novel three-stage curriculum that progressively trains the model from short-range perception to long-horizon multi-event reasoning; and (iii) Temporal Audio-Visual Interleaved Chain-of-Thought, a reasoning framework that explicitly grounds intermediate reasoning steps to timestamps in long audio-visual streams, improving temporal alignment and interpretability. Extensive experiments across 15+ audio-visual, omni-modal, audio, and vision benchmarks show that AV-Flamingo outperforms similarly sized open models by clear margins and remains highly competitive with, and in some cases surpasses, much larger open-weight and closed models, particularly on long and complex real-world audio-visual understanding and reasoning tasks. Beyond benchmark performance, AV-Flamingo exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability. We will release the data, code, methods, and checkpoints upon acceptance. Project: https://avflamingo.github.io/
Auditing Capsule Vision 2024: Within-Split Train-to-Validation Re-Exposure and a Kvasir-Channel Sensitivity Diagnostic
Eunseob Choi ⋅ Kyeonghun Kim ⋅ Hyuk-Jae Lee ⋅ Nam-Joon Kim
Public validation sets in machine-learning challenges frequently become de facto test sets in downstream work, making train–validation source overlap a first-order concern for claim validity. We audit the released public split of the Capsule Vision 2024 Challenge (CV2024) using perceptual hashing (pHash+dHash, PDQ at 99.4%), pixel NCC, and DINOv2/ResNet feature checks. The released CV2024 public split contains 1,381 / 11,581 KVASIR validation rows (11.9%; 540 pixel-confirmed at NCC ≥ 0.99, i.e. 4.66%) with pHash-exact training matches — directly violating the organizers' "NOT have duplicates" preprocessing description. All 1,381 trace to shared Kvasir-Capsule video prefixes, consistent with frame-level (not video-level) random splitting. The full CV2024-KVASIR slice falls in the disclosed Kvasir-Capsule near-duplicate set (60.7% bit-identical; non-KVASIR controls ≤ 0.31%). Training on the Kvasir-origin-removed le6 split drops public-validation balanced accuracy by Δ = −0.213 (95% CI [−0.220, −0.206], n=10 paired seeds), quantifying the public score's dependence on the KVASIR channel; the magnitude is mechanically expected given the 71.8% KVASIR validation composition rather than a leakage-specific causal effect. le6 is not a hidden-test proxy (team-level Spearman against AIIMS: ρ=0.355 vs. 0.566 for the original score). We release le6, le6plusinternal, hash/NCC annotations, audit code, and a Croissant 1.1 descriptor as a within-public-pool sensitivity probe; no image bytes are redistributed. Detector portability: ISIC 2019 cross-source rate is 0.008% (2/25,331, both re-attributed as within-source metadata-gap pairs ⇒ genuine cross-source 0/25,331; Cassidy et al. 2022's recall-tuned multi-hash pipeline flags 14,310); HyperKvasir and Kvasir-SEG cross-benches against CV2024 return 0 pairs (procedure-disjoint negative controls). At the matched NCC ≥ 0.99 endpoint, the CV2024 rate of 4.66% (540/11,581) is ≈ 148–583× above the ISIC cross-source rate (Wilson-lower / Wilson-upper to point/point).
A Unified Framework for Image-to-3D Part Generation via Variable Granularity
Qiwen Gu ⋅ Xin Zhou ⋅ Xiwu Chen ⋅ Zhongyang Zhu ⋅ Zhuo Zhang ⋅ Junqiao Zhao
Controllable 3D part generation is pivotal for digital asset creation, yet existing methods lack the flexibility to accommodate diverse user workflows requiring varying levels of control. In this paper, we introduce FlexPart, a unified image-to-3D part generation framework that supports variable-granularity geometric conditions within a single model. By seamlessly integrating heterogeneous inputs, ranging from sparse points to dense bounding boxes and masks, via gated adaptive modulation, FlexPart unlocks a powerful cross-granularity synergy. We also propose an asymmetric geometric guidance mechanism that leverages strong spatial priors to align representations across granularities, improving geometric consistency and facilitating inter-task synergy. Extensive experiments demonstrate that this synergistic approach outperforms state-of-the-art methods by over 10.5\% on Part CD, achieving high-fidelity part generation with superior geometric consistency across all prompt granularities. Code and models will be made publicly available.
A Unified Image and Video Encoder for Multimodal LLMs
Tongtian Yue ⋅ Zikang Liu ⋅ Longteng Guo ⋅ Handong Li ⋅ Zhibin Wang ⋅ Jun Fu ⋅ Zijia Zhao ⋅ Xianing Chen ⋅ Jun Song ⋅ Cheng Yu ⋅ Bo Zheng ⋅ Jing Liu
Multimodal large language models (MLLMs) have rapidly advanced machine intelligence by integrating visual perception with powerful linguistic reasoning. However existing visual backbones predominantly encode videos in a rigid per-frame manner and delegate temporal modeling entirely to the language model. This paradigm fundamentally constrains long-range temporal understanding and imposes severe computational inefficiencies. To resolve these bottlenecks we propose FuxViT, a unified Vision Transformer natively supporting a 64K-token context window to seamlessly process diverse inputs ranging from static images to hour-scale videos. By employing a joint spatiotemporal pretraining paradigm, FuxViT generates highly coherent and information-rich visual representations. This strategic early fusion frees the downstream language model from reconstructing low-level temporal structures to focus entirely on high-level reasoning. Furthermore we introduce Frame Selective Attention (FSA) as a native sparse routing mechanism. FSA dynamically attends to the most relevant historical frames to achieve exceptional computational efficiency without sacrificing fine-grained spatial or temporal details. Extensive experiments across 16 image and video benchmarks demonstrate that FuxViT consistently outperforms existing baselines and establishes itself as a vastly superior alternative to standard static ViTs for multimodal understanding. We will fully open-source all model checkpoints.
AutoHoney: Automating, Deploying, and Evaluating Scheming Honeypots Across Production Codebases
Martin Ciesielski-Listwan ⋅ Alyssia Jovellanos ⋅ Victoria Krakovna
LLMs are developing increasing capabilities in situational awareness, stealth and autonomy, approaching levels where they may be capable of scheming against their designers by covertly acting on misaligned goals (such as sabotaging oversight mechanisms under cover of a refactoring task). We evaluate model propensity for scheming using honeypots - realistic opportunities to pursue a misaligned goal. We present AutoHoney, an automated pipeline to generate, parameterize, and evaluate scheming honeypots in production codebases. Existing scheming propensity evaluations focus on catching and observing scheming behavior in artificial test settings, whereas we deploy honeypots in real codebases. To study scheming propensity in current models, we use prompted model organisms (PMOs) where goal conflicts are introduced through system prompts, planted documentation, or codebase-embedded environmental cues. Using PMOs with varying levels of goal nudging, we deploy honeypots across four production codebases and conduct a systematic manual ablation on multiple frontier models. Results show that goal presence, threat intensity, and task scope are strong predictors of scheming rates in our model organisms, while urgency has limited effect. The specific goal matters less than its existence—arbitrary goals like "promote mechanical pencils," environmental advocacy, and library funding all produce substantial scheming. Evaluation awareness typically suppresses scheming, but we observe notable exceptions where models take honeypots despite recognizing the evaluative context.
Automated Kernel Discovery Towards Understanding High-dimensional Bayesian Optimization
Taeyoung Yun ⋅ Woocheol Shin ⋅ Inhyuck Song ⋅ Jaewoo Lee ⋅ Jinkyoo Park
Gaussian Process (GP) kernels are central to Bayesian optimization (BO), yet designing effective kernels for high-dimensional problems still relies on extensive manual engineering. Existing kernel design methods restrict the search to additive and multiplicative compositions of base kernels, while LLM-based BO approaches condition on raw observations, which are infeasible in high dimensions due to context-length constraints and the difficulty of extracting meaningful patterns from high-dimensional numerical observations. We introduce Kernel Discovery, an LLM-driven population-based evolutionary framework that overcomes both limitations. Motivated by the observation that directly prompting an LLM to generate kernel code yields syntactically varied but functionally identical kernels, we adopt a two-stage approach: an LLM first proposes novel mathematical forms, then a second LLM call converts each form into validated, executable code. We also propose a leave-one-out continuous ranked probability score (LOO-CRPS) as a held-out predictive scoring criterion that penalizes overconfident fits more directly than marginal log-likelihood. On five standard high-dimensional BO benchmarks, our method achieves an average rank of 1.2 out of 17, consistently outperforming competitive baselines. We further analyze the discovered kernels to examine which kernel characteristics lead to superior performance in high-dimensional BO.
Automotive-ENV: Benchmarking Multimodal Models in Automotive Cockpit Environments
Junfeng Yan ⋅ Biao Wu ⋅ Meng Fang ⋅ Ling Chen
Multimodal models have demonstrated strong capabilities in web, desktop, and mobile GUI environments, but their performance in automotive cockpit systems remains largely unexplored. In-vehicle GUIs introduce unique challenges beyond traditional GUI interaction, including vehicle-state reasoning, implicit intent understanding, and safety-critical decision making. To address this gap, we introduce Automotive-ENV, an interactive benchmark for evaluating multimodal models in automotive cockpit environments. Built on Android Automotive OS, Automotive-ENV combines cockpit-screen interaction with vehicle-state signals and deterministic state-based evaluation. The benchmark contains 420 carefully constructed tasks covering routine cockpit operation and safety-critical scenarios. Unlike prior GUI benchmarks that rely on action matching or LLM-based judging, Automotive-ENV evaluates model behavior through reproducible validators over UI states, vehicle signals, and safety-related conditions. We benchmark representative general-purpose multimodal models and GUI-specialized vision-language models on Automotive-ENV. Experimental results show that current models remain far from reliable in automotive environments, with substantial performance degradation in implicit and safety-critical tasks. Further analysis reveals consistent failure modes in safety-aware reasoning, vehicle-control understanding, and implicit hazard recognition. We will release Automotive-ENV, together with tasks, validators, and evaluation tools, to support future research on automotive multimodal systems.
Autoregressive Appearance Prediction for 3D Gaussian Avatars
Michael Steiner ⋅ Zhang Chen ⋅ Alexander Richard ⋅ Vasu Agrawal ⋅ Markus Steinberger ⋅ Michael Zollhöfer
A photorealistic and immersive human avatar experience demands capturing fine, person-specific details such as cloth and hair dynamics, subtle facial expressions, and characteristic motion patterns. Achieving this requires large, high-quality datasets, which often introduce ambiguities and spurious correlations when very similar poses correspond to different appearances. Models that fit these details during training can overfit and produce unstable, abrupt appearance changes for novel poses. We propose a 3D Gaussian Splatting avatar model with a spatial MLP backbone that is conditioned on both pose and an appearance latent. The latent is learned during training by an encoder, yielding a compact representation that improves reconstruction quality and helps disambiguate pose-driven renderings. At driving time, our predictor autoregressively infers the latent, producing temporally smooth appearance evolution and improved stability. Overall, our method delivers a robust and practical path to high-fidelity, stable avatar driving.
Auto-Rubric as Reward: From Implicit Preference to Explicit Generative Criteria
Juanxi Tian ⋅ Fengyuan Liu ⋅ Jiaming Han ⋅ Yilei Jiang ⋅ Yongliang Wu ⋅ Yesheng Liu ⋅ Haodong Li ⋅ Furong Xu ⋅ Wanhua Li
Aligning multimodal generative models with human preferences demands reward signals that respect the compositional, multi-dimensional structure of human judgment. Prevailing RLHF approaches reduce this structure to scalar or pairwise labels, collapsing nuanced preferences into opaque parametric proxies and exposing vulnerabilities to reward hacking. While recent Rubrics-as-Reward (RaR) methods attempt to recover this structure through explicit criteria, generating rubrics that are simultaneously reliable, scalable, and data-efficient remains an open problem. We introduce Auto-Rubric as Reward (ARR), a framework that reframes reward modeling from implicit weight optimization to explicit, criteria-based decomposition. Before any pairwise comparison, ARR externalizes a VLM’s internalized preference knowledge as prompt-specific rubrics, translating holistic intent into independently verifiable quality dimensions. This conversion of implicit preference structure into inspectable, interpretable constraints substantially suppresses evaluation biases including positional bias, enabling both zero-shot deployment and few-shot conditioning on minimal supervision. To extend these gains into generative training, we propose Rubric Policy Optimization (RPO), which distills ARR’s structured multi-dimensional evaluation into a robust binary reward, replacing opaque scalar regression with rubric-conditioned preference decisions that stabilize policy gradients. On text-to-image generation and image editing benchmarks, ARR-RPO outperforms pairwise reward models and VLM judges, demonstrating that explicitly externalizing implicit preference knowledge into structured rubrics achieves more reliable, data-efficient multimodal alignment, revealing that the bottleneck is the absence of a factorized interface, not a deficit of knowledge.
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces
Sungho Park ⋅ Wonjoong Kim ⋅ Rongyuan Tan ⋅ Jue Zhang ⋅ WOOK SHIN HAN ⋅ Pengfei Gao ⋅ Chanyoung Park ⋅ Qingwei Lin ⋅ Saravan Rajmohan ⋅ Dongmei Zhang
LLM agents remain unreliable on long-horizon tasks, where small local failures can compound over extended interactions and lead to overall task failure. Although external harnesses can substantially improve robustness, harness design remains a manual and expensive process that requires searching over a large space of prompts, tool configurations, and control logic. We propose AutoSaddler, an automatic harness optimization framework that formulates harness improvement as an offline learning problem and iteratively updates the harness using failure signals from mini-batches. AutoSaddler combines failure-trace diagnosis, structured patch generation that treats the harness as code, and validation-based update selection. Experiments on GAIA2 and Terminal-Bench 2.0 show that AutoSaddler substantially improves agent performance over the corresponding base harnesses, achieving gains of 9.0 and 10.0 percentage points, respectively. Ablation studies further suggest that effective harness optimization benefits from three ingredients: deep debugging rather than shallow reflection, targeted modifications rather than unconstrained editing, and generalization-aware selection rather than trajectory-specific repair. Together, these results suggest that automatic harness optimization is a promising path toward more performant and reliable agent systems.
AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation
Shanghua Gao ⋅ Ada Fang ⋅ Marinka Zitnik
Scientific research proceeds through iterative cycles of hypothesis generation, experiment design, execution, and revision, often requiring researchers to explore multiple competing directions as evidence accumulates and priorities shift. LLM agents can automate parts of this process, but existing agents either concentrate reasoning within a single research thread or coordinate through a central planner with fixed objectives. As a result, they struggle to sustain parallel exploration across research directions or reorganize as promising and unproductive directions emerge over time. We introduce AutoScientists, a decentralized team of AI agents for long-running computational scientific experimentation. Rather than following decisions from a central orchestrator, agents independently interpret a shared experimental state, self-organize into teams around research directions, critique and filter proposals with a discussion phase before committing experimental compute, and exchange both successful and failed findings across teams to avoid redundant exploration. Under matched experimental budgets, AutoScientists outperforms prior agentic systems across biomedical machine learning, language-model training optimization, and protein fitness prediction. On BioML-Bench, spanning biomedical imaging, protein engineering, single-cell omics, and drug discovery, AutoScientists achieves a mean leaderboard percentile of 74.4% across 24 tasks, improving over the strongest prior biomedical agent by +8.33%. On GPT training optimization, AutoScientists reaches a target validation bits-per-byte 1.9x faster than autoresearch and continues discovering improvements from a stronger starting champion where the single-agent approach finds none (7 vs. 0 accepted improvements). On ProteinGym fitness prediction, AutoScientists discovers a method for ACE2-spike binding that improves over current state-of-the-art model by +12.5% Spearman correlation. Applied without modification to all 217 ProteinGym assays, the same method improves over the prior state of the art by +6.5% in Spearman correlation. Code: https://anonymous.4open.science/r/autoscientists_anonymize-D10C
Auxiliary Clues Aware’s Geometry Problem Solving
Qigong Lei ⋅ Qingsong Wang ⋅ Wang Lin ⋅ Qianyi Yang ⋅ Chang Yao ⋅ Jingyuan Chen
Geometry Problem Solving (GPS) serves as a rigorous touchstone for evaluating the complex reasoning capabilities of Multimodal Large Language Models (MLLMs), particularly when solutions necessitate auxiliary constructions absent in the visual input. However, current approaches face the following challenges: explicit diagramming methods suffer from unstable image generation quality and error cascading, while existing implicit attention mechanisms are restricted to visible elements, failing to perceive the latent auxiliary structures required for deduction. To bridge this gap, we propose AuxCA (Auxiliary-Clues-Aware GRPO), a novel framework that internalizes the capability of auxiliary construction into the model’s reasoning intuition without relying on external tools. Specifically, we introduce Auxiliary Clues Dependency (ACD) to quantify the causal influence of visual cues and integrate it into GRPO via a fine-grained advantage weighting strategy. Complementing this, we incorporate an auxiliary incentive reward to explicitly encourage the model to actively attend to and utilize these geometric elements during reasoning. Extensive experiments demonstrate that our resulting model, AuxCA-8B, significantly outperforms state-of-the-art explicit plotting methods and achieves performance competitive with top-tier proprietary models. Code is available in \url{https://anonymous.4open.science/r/AuxCA-5510}.
AVIS: Adaptive Test-Time Scaling for Vision–Language Models
Ahmadreza Jeddi ⋅ Minh Le ⋅ Amirhossein Kazerouni ⋅ Hakki Karaimer ⋅ Hue Nguyen ⋅ Iqbal Mohomed ⋅ Michael Brudno ⋅ Alex Levinshtein ⋅ Kosta Derpanis ⋅ Babak Taati ⋅ Radek Grzeszczuk
Modern Vision-Language Models (VLMs) benefit from chain-of-thought prompting and test-time scaling, but these gains often come with prohibitive inference cost due to large visual contexts and long decoding chains. We view this cost through two coupled axes: Visual Context Scaling (VCS), which controls how much visual evidence is passed to the language model, and Visual Reasoning Scaling (VRS), which controls how much inference-time reasoning search is performed. Existing methods typically optimize one axis at a time, leaving the joint allocation of compute across these axes underexplored. We introduce Adaptive Visual Inference Scaling (AVIS), a lightweight policy that adapts both VCS and VRS per query. AVIS realizes VCS through Key Diversity Visual (KDV) pruning, a training-free $O(N)$ key-based rule for removing redundant visual tokens before prefilling, and realizes VRS through adaptive self-consistency, using a learned difficulty predictor to select the number of reasoning rollouts. AVIS is deployment-friendly and compatible with shared-prefill inference, where all rollouts reuse a single prefilling pass and KV cache. Across diverse image and video reasoning benchmarks, AVIS improves the accuracy--compute trade-off relative to VCS-only and VRS-only baselines, and remains effective on top of RL post-trained VLMs while keeping compute and latency low.
Axiomatic Reinforcement Learning for Open Multi-Agent Systems from Shapley Axioms
Jianhong Wang ⋅ Yang Li ⋅ Samuel Kaski ⋅ Jonathan Lawry
Open multi-agent systems, where only a subset of agents are controllable, are central to real-world domains such as smart grids, yet lack a principled foundation for reinforcement learning algorithm design. We study this setting through n-agent ad hoc teamwork (NAHT), a recent framework for learning in open multi-agent environments. We propose an axiomatic framework that derives value function learning from the Shapley axioms (Efficiency, Symmetry and Linearity), treating them as structural constraints rather than heuristic design choices. This formulation builds on a novel NAHT decomposition, which characterises the fundamental difference between open and fully controllable multi-agent systems, and establishes a cooperative game-theoretic foundation for learning. We further show that enforcing these axioms on learning individual value functions recovers the Shapley value for Dec-POMDPs, providing a principled bridge between reinforcement learning and classical cooperative game theory. Building on this framework, we introduce Shapley Machines and Banzhaf Machines, two generic algorithmic templates that can be instantiated with standard reinforcement learning methods. When applied to strong baselines such as IPPO and POAM, our methods consistently improve learning efficiency and generalisation to unseen agent types and conventions across NAHT benchmark tasks.
Baba in Wonderland: Online Self-Supervised Dynamics Discovery for Executable World Models
SeungWon Seo ⋅ DongHeun Han ⋅ SeongRae Noh ⋅ Hyeongyeop Kang
Executable world models can be read, edited, executed, and reused for planning, but only if the program captures the environment's transition law rather than semantic shortcuts in its surface vocabulary. We study online executable world-model learning under prior misalignment, where an agent must induce state-dependent dynamics from interaction evidence alone, without rule descriptions, reward signals, or trustworthy lexical priors. We introduce Alice, a closed-loop system that treats failed candidate updates as structural signal: when a candidate explains a new transition but loses previously explained ones, the preservation conflict reveals dynamics that the current program had conflated. Alice refines these conflicts into hypothesis classes that both provide compact, class-stratified preservation counterexamples for update and guide frontier exploration toward transitions that are novel and underrepresented with respect to the current program. We evaluate Alice on Baba in Wonderland, a prior-misaligned variant of Baba Is You that preserves simulator dynamics while replacing semantically meaningful rule-property labels with unrelated words. Experiments show that Alice substantially improves executable world-model learning under prior misalignment, and ablations show that both class refinement and class-aware exploration contribute.
Bandits via Additive Quantized Representations
Ami Tavory ⋅ Noam Touitou ⋅ Tal Sarig ⋅ Frank Cheng ⋅ Ido Guy
Contextual bandits require balancing nonlinear reward modeling with online efficiency. Tree ensembles capture nonlinearities but require periodic retraining and large replay buffers. Linear models update efficiently per observation with $O(1)$ memory, but are fundamentally restricted to linear reward structures. We propose Residual Quantization (RQ) as a representation layer to bridge this gap. An offline-trained RQ codebook maps continuous contexts into discrete centroid assignments across $\ell$ levels, set dynamically through a shadow mechanism. This enables a spectrum of additive bandit algorithms that achieve nonlinear expressivity with strictly bounded memory. Across 13 datasets, RQ variants beat their non-RQ counterparts on 11 of 13 datasets, often by wide margins, while matching doubling-retrain XGBoost and neural baselines at up to $1000\times$ less memory.
BasicLT: Basic-Level Abstraction and Selective Differentiation for Long-Tailed Recognition
Yuan Dong ⋅ Di Wu ⋅ Zhe Zhao ⋅ Liheng Yu ⋅ Xiaofeng Cao ⋅ Pengkun Wang ⋅ Yang Wang
Long-tailed recognition is commonly formulated as a problem of imbalanced supervision, where rare classes suffer from insufficient training examples. In this work, we study a complementary source of difficulty: under severe scarcity, tail classes may be forced into fine-grained discrimination before the model has acquired sufficiently reliable coarse semantic structure. We refer to this phenomenon as granularity mismatch under scarcity. To address it, we propose BasicLT, a two-stage framework that augments an underlying fine-grained recognition model with basic-level abstraction and commitment-guided selective differentiation. In Stage 1, BasicLT learns an explicit basic-level branch and aligns it with the family-level distribution induced by the fine-grained predictor, encouraging stable coarse semantic organization before additional fine-grained refinement is introduced. In Stage 2, a frozen Stage 1 teacher provides stable family-level commitment signals, and a family-conditioned residual refinement module is trained only on committed medium-shot and few-shot samples. At inference time, the residual correction is activated conservatively according to student-side commitment, so that fine-grained refinement acts as a conditional within-family correction rather than an unconditional auxiliary classifier. We evaluate BasicLT on CIFAR-100-LT, CIFAR-10-LT, ImageNet-LT, iNaturalist 2018, and Places-LT. Across these benchmarks, BasicLT achieves competitive performance against strong long-tailed baselines, with the most consistent gains appearing in medium-shot and few-shot regimes. Ablations and analyses further show that both basic-level abstraction and commitment-guided selective refinement are important, supporting the view that controlling the semantic granularity of discrimination can complement conventional rebalancing strategies for long-tailed learning. Code is available at Supplement.
Basis-Mediated Bilinear Attention: A New Method for Greatly Reducing Query--Key Pathway Parameters
Wei Chen ⋅ Wenhao Jiang ⋅ Yiying Yang
Large language models and their vision counterparts have achieved remarkable capabilities, yet their wider deployment is increasingly constrained by the computational and memory demands of Transformer architectures. Much recent work improves Transformer efficiency, but the query-key pathway is still typically parameterized by fully independent per-head projections, which place a substantial parameter burden on attention. Motivated by the view that token selection may admit a shared low-dimensional geometry, we propose $\textbf{Basis-Mediated Bilinear Attention} (\textbf{BMB})$, a training-time reparameterization of the query-key pathway in which all heads within a layer share a latent basis while retaining head-specific interactions. A practical explicit variant, BMB-UV, materializes per-head query and key tensors through a shared basis and lightweight head-specific factors, reducing the query-key parameter count from $2d^2$ to $dr+2Hrs$. We evaluate our method family on ViT-Base image classification on ImageNet-1K and on GPT-2 small pretraining on the C4 RealNewsLike subset, and compare it against recent query-key parameter-sharing and low-rank baselines. Relative to standard multi-head attention, our method family reduces query-key parameters by up to $ \textbf{95.83} $%$ $ while remaining competitive with the baseline on both benchmarks.
BAT3R: Robust Online 3D Reconstruction with Bayesian Adaptive State Updates
Lijian Li ⋅ Yuanpeng He ⋅ Yudian Zheng ⋅ Chung-ju Huang ⋅ Linyu Li ⋅ Dongming Jin ⋅ Wenpin Jiao ⋅ Zhi Jin ⋅ Chi-Man Pun
Feed-forward transformer models have driven remarkable progress in 3D vision, yet their quadratic complexity renders them impractical for long video sequences. Streaming 3D reconstruction addresses this bottleneck by processing frames sequentially with constant memory. However, existing recurrent architectures inevitably suffer from progressive degradation over long sequences, as they fail to account for the intrinsic reliability of observations and the accumulated confidence of the historical state, leading to severe error accumulation. To address these issues, we introduce BAT3R, a training-free Bayesian adaptation framework comprising two synergistic modules: State Innovation Gating (SIG) and Bayesian-guided Adaptive State Updating (BASU). SIG acts as a distributional regularizer that anchors state evolution to its geometric initialization, effectively arresting representational drift during long-term inference. Building on this, BASU reformulates state update as a recursive Bayesian inference process within a Kalman filtering framework. By jointly modeling the temporal reliability of observations via information-theoretic gain and spatial relevance via cross-attention maps, BASU achieves an asymmetric spatio-temporal state fusion: well-consolidated historical regions are protected from redundant updates, while genuinely novel scene content is aggressively incorporated. Extensive experiments on standard benchmarks (7-Scenes, TUM, ScanNet, KITTI, etc.) demonstrate that BAT3R consistently outperforms state-of-the-art online 3D reconstruction methods, with the most significant gains observed on long sequences and complex scenes.
Bayesian Optimization on Function Spaces via Sparse RKHS Manifolds
Davide Sartor ⋅ Meghan E Huber ⋅ Donghyun Kim ⋅ Nathan Wycoff
Bayesian Optimization has become an established methodology for minimizing black-box functions of a vector input. Often, however, this parameter vector arises from the discretization of an inherently functional relationship. Several recent articles have considered the Functional Bayesian Optimization (FBO) setting, in which the variable to be optimized is not a member of a finite dimensional vector space, but rather an infinite dimensional function space. In this work, we propose $L^0$ Manifold Optimization (L0MO), a simple approach to FBO which searches a sub-manifold of a Reproducing Kernel Hilbert Space consisting of functions with a sparse representation in the kernel functions. This results in simpler, more effective functional solutions than existing methods. We discuss in detail the relationship between our method and existing ones, providing a unifying lens through which to view prior works. To assess our method against the state of the art, we conduct an extensive computational study, and along the way develop a novel set of benchmark test functions which port standard finite-dimensional ones to the infinite dimensional domain. Our experiments demonstrate that the proposed method achieves superior performance across a wide range of test benchmarks.
BayesJudge: Uncertainty-Aware Bayesian Meta-Evaluation of Human and LLM Judgments
Jiahao Zhang ⋅ Pengbin Feng ⋅ Chunlei Meng ⋅ Hang He ⋅ Tianyong Hao
AI evaluation pipelines often produce conflicting judgments rather than clean labels. In pairwise LLM evaluation, this conflict is especially visible: disagreement can arise from ambiguous items, underspecified rubrics, heterogeneous or unstable human raters, or an LLM judge whose verdict changes when the response order is swapped. We propose BayesJudge, an online Bayesian meta-evaluation layer for conflicting human-LLM judgment streams. For each comparison, BayesJudge estimates a panel-relative posterior verdict distribution over the two responses, with a tie or ambiguity state when such labels are available. At the same time, it estimates rater-specific human confusion matrices and LLM presentation-order bias. The method uses tie-open labels to keep ambiguity observable and paired order-swapped judge calls to separate response quality from presentation effects. We formulate the exact online posterior recursion and use a scalable Rao--Blackwellized assumed-density SMC approximation for streaming inference. Controlled synthetic experiments demonstrate recovery of prespecified evaluator parameters and illustrate two protocol-level identifiability mechanisms: tie-open labels expose ambiguity mass, and order-swapped paired judgments separate item preference from position bias. On real-world SummEval dataset, BayesJudge successfully detects systematic presentation-order effects in LLM judge outputs, infers distinct expert and crowdworker behavior signatures without rater metadata, and produces posterior uncertainty estimates that correlate with human disagreement. Our code is available at https://anonymous.4open.science/r/BayesJudge-0879.
BEAGLE: Behavior-Enforced Agent for Grounded Learner Emulation
Hanchen D Wang ⋅ Clayton Cohn ⋅ Zifan Xu ⋅ Siyuan Guo ⋅ Gautam Biswas ⋅ Meiyi Ma
Simulating student learning behaviors in open-ended problem-solving environments holds potential for education research, from training adaptive tutoring systems to stress-testing pedagogical interventions. However, collecting authentic data is challenging due to privacy concerns and the high cost of longitudinal studies. While Large Language Models (LLMs) offer a promising path to student simulation, they suffer from competency bias, optimizing for efficient correctness rather than the erratic, iterative struggle characteristic of novice learners. We present BEAGLE, a neuro-symbolic framework that addresses this bias by incorporating Self-Regulated Learning (SRL) theory into a novel architecture. BEAGLE integrates three key technical innovations: (1) a semi-Markov model that governs the timing and transitions of cognitive behaviors and metacognitive behaviors; (2) Bayesian Knowledge Tracing with explicit flaw injection to enforce realistic knowledge gaps and "unknown unknowns"; and (3) a decoupled agent design that separates high-level strategy use from code generation actions to prevent the model from silently correcting its own intentional errors. In evaluations on Python programming tasks, BEAGLE significantly outperforms state-of-the-art baselines in reproducing authentic trajectories. In a human Turing test, participants could not reliably tell BEAGLE traces apart from real student data: classification accuracy was statistically equivalent to chance (52.8\%, $d'=0.15$, $N=71$).
Beckmann Transport Models: From Autonomous Flows to One-Step Maps
Lee Kit ⋅ Florentin Coeurdoux ⋅ Peter Potaptchik ⋅ Yilun Du ⋅ Michael Albergo ⋅ Eric Vanden-Eijnden
We propose an instantiation of flow matching that relies on a time-independent velocity field (an \emph{autonomous flow}) to exactly map between two distributions, so long as the target is singular, i.e.\ supported on a lower-dimensional data manifold. We also show that the one-step generative map associated with this flow is the unique solution of a simple conservation equation, which can be used to learn the map directly from samples. These autonomous flows and maps give a dynamical meaning to the flux constraint of Beckmann's transportation problem. Their construction provides a unifying framework that recovers, for instance, the closed-form Poisson-flow generative model and equilibrium matching with a quadratic flow-matching regression loss. We illustrate how this theory corrects inconsistencies in existing methods and improves their performance on large-scale image generation.
Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs
Youngwoo Shin ⋅ Yusung Ro ⋅ Minseo Kim ⋅ Junmo Kim
Video Large Language Models (VideoLLMs) receive frames in sequential order and interpret how visual content evolves along the temporal axis, yet temporal reasoning remains a persistent weakness across architectures. Reversing the frame order of a video, a transformation that should invert temporal answers, often leaves the final prediction unchanged. We investigate where this failure originates by defining the temporal divergence vector $\tau_l$, the layer-wise representational difference induced by reversing temporal order. Tracking its magnitude across layers reveals a consistent temporal divergence profile where the divergence peaks at intermediate layers and progressively diminishes toward the output. We confirm this peak is specific to temporal reasoning and functionally critical for predictions, establishing that VideoLLMs acquire temporal information at intermediate layers but fail to maintain it to the output. This progressive fading motivates our method, Temporal Activation Injection (TAI), which extracts $\tau_l$ at the peak of the profile for each input and reinjects it into subsequent layers following the measured decay. TAI requires no training and consistently improves temporal reasoning across three VideoLLMs and four benchmarks with negligible impact on non-temporal tasks.
Behavioral Geometric Supervision Aligns Video Foundation Models with Human Social Perception
Kathy Garcia ⋅ Leyla Isik
Current video foundation models, including the strongest self-supervised models such as V-JEPA2, fail to capture how humans organize social information in dynamic scenes. For example, across a range of diverse vision models tested, none were able to predict human similarity judgments to social video clips as well as a sentence embedding model of the caption text (MPNet). We show this gap in vision model performance can be closed by a compact behavioral supervisory signal. We introduce behavioral geometric supervision (BGS): a hybrid objective that constrains local and global pairwise embedding geometry to match the relational similarity structure across videos. We apply this method using a new human similarity dataset, containing 49,484 odd-one-out judgments from 250 naturalistic social video clips, and low-rank adaptation across four ViT backbones (V-JEPA 2/2.1, TimeSformer, VideoMAE, and CLIP). We find that one of the best fine-tuned models, V-JEPA 2.1, nearly triples in performance compared to the pre-trained baseline and reaches close to the noise ceiling, exceeding the strongest sentence-embedding baseline. In addition, finetuned models (i) capture unique variance in human judgments that caption-based language embeddings do not, (ii) develop interpretable social-affective attributes (valence, arousal, and dominance) despite never being trained on any of these attributes, (iii) zero-shot transfer to a separate dataset of out-of-distribution abstract social interactions, and (iv) shift spatial attention from scene context to socially informative regions (faces, gaze, and interacting bodies). A matched language-distillation control fails to reproduce these gains, ruling out caption transfer as the mechanism. Our results show how a modest amount of human behavioral data can steer video models toward human-like social visual understanding.
BELLS-O: Evaluating the Operational Trade-offs of LLM Supervision Systems
Leonhard Waibl ⋅ Felix Michalak ⋅ Hadrien Mariaccia
LLM supervision systems, namely input/output moderation filters and jailbreak detectors, are the primary safeguard against misuse in deployed AI applications, yet existing benchmarks are often vendor-biased, omit cost and latency, and rarely compare specialized guardrails against repurposed generalist LLMs. We present BELLS-O (Benchmark for the Evaluation of LLM Supervision Systems -- Operational), the first independent operational benchmark of LLM supervision systems. BELLS-O evaluates 28 systems from 17 providers: every major specialized guardrail (e.g., LlamaGuard-4, ShieldGemma-2, Lakera Guard) and frontier generalists repurposed as supervisors (e.g., GPT-5.4, Claude Sonnet 4.6, Grok-4.1), jointly on detection rate, false-positive rate, latency, and monetary cost. We cover input/output moderation across 11 harm categories and jailbreak detection across 13 attack techniques, using in-house datasets built from handcrafted prompts, expert-curated samples, and quality-controlled synthetic generation. To prevent latent generator-specific signals in synthetic data, every generated sample is run through a paraphrasing step that suppresses these fingerprints. Mapping the Pareto frontier reveals use-case-dependent tradeoffs. On content moderation, specialized supervisors are operationally dominant: top systems match frontier LLMs on detection ($\approx$95\% vs.\ 94\%) at comparably low false-positive rates ($\leq$2\%), while running 5--10$\times$ faster and ${\sim}$10$\times$ cheaper. On jailbreak detection, the tradeoff shifts: frontier LLMs achieve higher detection and lower false-positive rates but at 10--50$\times$ higher cost and 5--10$\times$ higher latency. We release the benchmark, framework, leaderboard, and datasets as the first vendor-neutral basis for selecting safeguards under real deployment constraints.
Benchmarking Compositional Generalisation for Machine Learning Interatomic Potentials
Amir Masoud Nourollah ⋅ Irtaza Khalid ⋅ Stefano Leoni ⋅ Steven Schockaert
Machine Learning Interatomic Potentials play a fundamental role in computational chemistry and materials science, enabling applications from molecular dynamics simulations to drug design and materials discovery. While recent approaches can estimate inter-atomic forces with high precision, it remains unclear to what extent they can generalise to previously unseen molecules. Do they learn the compositional structure of chemistry, capturing how molecular fragments and their combinations determine properties, or do they primarily learn to interpolate patterns that are specific to the training examples? To address this question, we propose a benchmark consisting of four tasks that require some form of compositional generalisation. In each task, models are tested on molecules that were unseen during training, but the training data is chosen such that generalisation to the test examples should be feasible for models that learn the underlying physical principles. Our empirical analysis shows that the considered tasks are highly challenging for state-of-the-art models, with errors on out-of-distribution examples often an order of magnitude higher than on in-distribution examples, even when using foundation models that have been pre-trained on millions of molecules.
Benchmarking Open-Ended Multi-Agent Coordination in Language Agents
Kale-ab Tessera ⋅ Andras Szecsenyi ⋅ Cameron Barker ⋅ Alexander Rutherford ⋅ Davide Paglieri ⋅ Aidan Scannell ⋅ Henry Gouk ⋅ Elliot Crowley ⋅ Tim Rocktäschel ⋅ Amos Storkey
As language models are increasingly deployed as autonomous agents, they will need to coordinate with others in long-horizon, open-ended interactive tasks. Yet current evaluations rarely test these demands together, focusing instead on short interactions, single-agent open-ended tasks, or highly structured multi-agent settings. We introduce $alem$, a JAX-based, procedurally generated open-ended benchmark for multi-agent coordination built on Craftax-like dynamics. Alem embeds procedurally generated coordination tasks, soft specialisation, communication, and controllable coordination difficulty into a long-horizon survival world with exploration, crafting, trading, and combat. We evaluate $13$ modern LLMs zero-shot within homogeneous teams, with trained MARL agents as reference points. Most LLM agents struggle, averaging only ~6% of maximum reward, although performance varies widely. Gemini-3.1-Pro-High approaches MARL agents trained for one billion steps on Hard coordination ($17.5\%$ vs. $17.6\%$ Coord.\%), Gemma-4-31B-it is the strongest tested open-weight model, and GPT-5.4-High makes strong base progress, while achieving much lower coordination reward. We find that base task competence does not imply coordination competence, communication helps agents share intent, memory and reasoning help with multi-step planning, and heterogeneous teams regress toward average member performance rather than matching their strongest teammate. These results identify coordination as a distinct bottleneck for current LLM agents, separate from single-agent capabilities. Alem makes this bottleneck measurable, providing a controlled setting for developing agents that can communicate, allocate roles, and execute shared plans.
Bet Imaginatively, not Historically in Independent-Data Sequential Testing
Nathaniel Xu ⋅ Feng Liu ⋅ Danica J. Sutherland
Sequential testing avoids some of the many controversies and drawbacks related to traditional batch testing with p-values. In particular, the testing by betting paradigm has developed rapidly in recent years and frames the level of evidence against the null as the wealth of a player in a betting game. Naturally, how much a player should bet is a key consideration, and should reflect the statistician's confidence in what they know about the next observed data point. Much of current work uses standard betting strategies developed for adversarial problems, where the data can be hand-picked to the detriment of the player. In the stochastic setting, where the data observed is independent and stationary, these approaches may be suboptimal. We propose a simple betting strategy, **follow the leave-one-out e-power (FLOE), that exploits the i.i.d. nature of a data stream to estimate the e-power of our current test on the underlying distribution. We average not over the historical evidence collected so far, as generic online learning-based methods do, but instead over the equally-likely possible data sequences we might have seen based on shuffling the data. With an additional principled regularization technique, this simple drop-in method yields substantial improvements in testing performance on a variety of kernel two-sample testing problems.
Better Language Models Require Better Domain-Specific Inductive Biases
Damien Teney ⋅ Liangze Jiang ⋅ Zachary Shinnick ⋅ Hemanth Saratchandran ⋅ Simon Lucey
Language models (LMs) obtain much of their capability from the diversity of their training data. This is possible because neural architectures like transformers are remarkably flexible and can handle a variety of domains such as natural language, code, and mathematics. This raises the question: are transformers optimal for any specific domain? In this position paper, we present empirical evidence that the inductive biases of transformer-based LLMs should be re-examined, since, while broadly useful, they are often suboptimal on subsets of their training data. Methods. To support this position, we use a method that learns new activation functions within a transformer as a tool to explore the space of inductive biases. We train the resulting modified architectures on various data subsets. A large-scale evaluation shows that vanilla transformers are relatively easy to improve, but mostly on a per-domain basis. For example, an architecture optimized for a natural language dataset like FineWeb performs well on another such dataset like TinyStories, whereas architectures optimized for mathematics perform poorly on natural language. We also identify the mechanisms responsible for these improvements. Implications. These results call for re-evaluating core design choices in language models. We do not advocate a return to handcrafted, domain-specific models. Instead, we emphasize that different domains can benefit from different learning mechanisms. This will require new designs that preserve cross-domain interactions, analogous to how reasoning in the brain arises from the interaction of multiple specialized networks. We argue that this perspective may be key to improving reliability and data efficiency.
Beyond Average Flatness: Domain-wise Flatness for Domain Generalization
Seungjun Choi ⋅ Heeyoung Kim
Sharpness-Aware Minimization (SAM) is widely used in domain generalization to promote flat minima that generalize to unseen distributions. Existing SAM-based methods typically penalize the sharpness of the *average* loss across domains. We show that this prevailing formulation—termed AvSAM—has a fundamental geometric limitation: it can favor solutions with misaligned domain-wise dominant Hessian eigenvectors. As a result, such solutions may appear flat on average while remaining sharp within individual domains. To address this limitation, we propose an alternative objective that penalizes the average *domain-wise* sharpness. We show that this objective decomposes into the AvSAM objective plus a non-negative residual term, which we call the *Cancellation Gap*. This gap is minimized if and only if the domain-wise dominant eigenvectors are aligned up to sign, indicating that it mitigates AvSAM's tendency to favor misaligned solutions. However, we further uncover a degeneracy: the Cancellation Gap can become small without improving alignment when sharpness is highly imbalanced across domains. To mitigate this failure mode, we propose Cancellation Gap Minimization (CGM), which augments the objective with a squared coefficient-of-variation ($\mathrm{CV}^2$) regularizer to discourage sharpness imbalance. Experiments show that CGM achieves the best average performance across standard domain generalization benchmarks among strong baselines. Diagnostic analyses further confirm reduced domain-wise sharpness and improved alignment of dominant Hessian eigenvectors.
Beyond Decoupled PEFT: Geometry-Aware Low-Rank Adaptation via Riemannian Reparameterization
Yuanyang Cao ⋅ Xichun Liu ⋅ Haitao Jiang ⋅ Jianji Wang
Low-Rank Adaptation (LoRA) and its weight-decomposed variants dominate parameter-efficient fine-tuning, yet their directional update mechanisms remain geometrically mismatched to the normalized parameterization they impose. In existing decoupled PEFT methods, low-rank directional updates are computed in unconstrained Euclidean space and only then projected back onto a spherical manifold, which can introduce capacity redundancy and directional gradient distortion. To address these limitations, we propose Geodesic Orthogonal Low-Rank Adaptation (GO-LoRA), a geometry-aware reparameterization framework for decoupled PEFT. GO-LoRA first applies orthogonal tangent projection to remove components parallel to the pretrained base direction, ensuring that the low-rank directional budget is devoted entirely to tangent-space motion. It then uses a geodesic flow based on the Riemannian Exponential Map to construct manifold-consistent directional updates on the hypersphere, improving geometric fidelity relative to Euclidean normalization while retaining standard optimization over the underlying low-rank parameters. Extensive experiments across language, vision, code generation, and multimodal reasoning tasks show that GO-LoRA consistently outperforms strong LoRA baselines while introducing no additional inference-time overhead after offline merging. The code is available at https://anonymous.4open.science/r/go-lora.
Beyond Eigenfunctions: Divergence Principal Functions for Representation Learning
Ritabrata Ray ⋅ Sahil Dharod ⋅ Burak Varıcı ⋅ Nicholas Boffi ⋅ Pradeep Ravikumar
Recent work has shown that contrastive representation learning can be understood as estimating a positive-pair, PMI-like kernel, and that spectral factorization of this kernel yields useful eigenfunction representations. This paper asks what happens beyond squared-error spectral geometry. Modern representation learning objectives are rarely pure $\ell_2$ kernel-approximation objectives: they use conditional KL, InfoNCE, JS/NCE, Brier, Hellinger, least-squares ratio fitting, and other statistical discrepancies. We show that these discrepancies do not merely provide alternative estimators of the same affinity; they induce different local geometries for the extraction of low-rank representations. To formalize this, we introduce divergence principal functions, a geometry-aware generalization of eigenfunctions. Eigenfunctions are recovered as the special case corresponding to squared-error geometry. For smooth discrepancies, we show that divergence principal functions locally solve a weighted spectral approximation problem, with weights given by the curvature of the divergence at the target affinity. This provides a loss-consistent bridge from positive-pair learning objectives to representation geometry and clarifies when spectral/eigenfunction representations are appropriate and when they are mismatched to the training objective.
Beyond Expected Values: Risk-Sensitive Planning with Distributional Monte-Carlo Tree Search
Quynh Trinh ⋅ Trung Nguyen ⋅ Nam Nguyen ⋅ Tuan Dam
Monte-Carlo Tree Search (MCTS) typically stores only scalar mean estimates, which makes it poorly suited for tail-risk objectives such as conditional value-at-risk (CVaR). We introduce **Categorical Thompson Sampling with Optimism (CATSO)** and **Particle Thompson Sampling with Optimism (PATSO)**, two distributional MCTS algorithms that maintain finite-support empirical backup laws at Q-edges. CATSO represents Q-edge laws with categorical atoms and Dirichlet Thompson sampling, while PATSO represents them with adaptive particles. Both algorithms select actions using the CVaR of a Thompson-sampled Q-edge law plus a polynomial optimism bonus, and propagate scalar continuation values through visit-weighted averages of child Q-edge CVaR scores. For the finite-depth nested CVaR planning target, we prove an $O(n^{-1/2})$ bound on the root value error after $n$ rollouts, which upper-bounds the risk-sensitive simple regret. CATSO adds a fixed-grid discretization bias, while capped PATSO adds a tunable Wasserstein compression bias. Experiments on stochastic planning benchmarks show that distributional Q-edges improve lower-tail return and reduce catastrophic outcomes when mean-optimal and risk-sensitive behavior differ.
Beyond Flat Gossip: Tiered Gossip Learning for Scalable Collaborative AI
Atul Sharma ⋅ Kavindu Herath ⋅ Saurabh Bagchi ⋅ Chaoyue Liu ⋅ Somali Chaterji
As collaborative machine learning scales to thousands of edge devices, two failure modes emerge: centralized federated learning creates server bottlenecks and single points of failure, while flat peer-to-peer systems force every node to increase its communication degree as the network grows. We propose Tiered Gossip Learning (TGL), a two-layer push–gossip–pull protocol that decouples data-holding leaf nodes from a decentralized relay backbone. In each round, leaves push local models to a subset of relays, relays gossip among themselves to mix models globally, and leaves pull updates from another random relay subset. This asymmetric design keeps per-leaf communication fixed as the network scales, offloading the mixing burden to a thin layer of higher-capacity relays without relying on any central coordinator for aggregation. Across CIFAR-10, FEMNIST, and AG News, TGL matches or exceeds baseline accuracy with up to 80\% fewer model exchanges. We provide convergence guarantees under standard smoothness, bounded variance, and heterogeneity assumptions, with explicit stage-wise consensus bounds characterizing how relay-layer connectivity governs the global mixing rate.
Beyond IID: How General Are Tabular Foundation Models, Really?
Lennart Purucker ⋅ Andrej Tschalzev ⋅ Nick Erickson ⋅ David Holzmüller ⋅ Alan Arazi ⋅ Alexander Pfefferle ⋅ Mustafa Tajjar ⋅ Gioia Blayer ⋅ Gael Varoquaux ⋅ Frank Hutter
Foundation models for predictive machine learning on tabular data have recently gained significant traction in academia and industry. Research communities across disciplines are increasingly evaluating tabular foundation models on diverse datasets and tasks. However, these task- and discipline-specific evaluations remain largely inaccessible to model researchers because benchmark software and evaluation protocols are fragmented. As a result, model researchers rely on standard benchmarks, which are mostly defined for tasks where tabular foundation models already excel. The most challenging scenarios are excluded, limiting meaningful progress in the field by focusing on marginal improvements on IID data rather than on broader, more demanding challenges. To overcome this, we introduce BeyondArena, the first unified holistic benchmark for tabular data that supports diverse task types (IID, temporal, grouped), across sample size and feature dimensionality scales, with diverse feature types (with text, with high cardinality) from a broad range of disciplines. To enable unified benchmarking beyond standard benchmarks, we introduce DataFoundry, a Python framework and metadata schema for curating tabular datasets for predictive machine learning. Our results across 11 models and 142 curated datasets show that existing tabular foundation models excel on tiny- to medium-sized IID data, while traditional tree-based and deep learning models still dominate on non-IID, large, and high-dimensional datasets. BeyondArena guides model research for the most demanding challenges in tabular data, enabling progress towards truly foundational tabular models.
Beyond Langevin: Sampling Multimodal Densities using the Witten Laplacian on 1-forms
Sahani Pathiraja ⋅ Panos Parpas
We introduce a new algorithm to sample from a probability density $\pi \propto e^{-V}$ with $V(x): \mathbb{R}^d \rightarrow \mathbb{R}$ a strongly multi-modal potential. Such distributions are challenging to sample from using standard gradient based methods, as exploration typically relies on inefficient diffusions. Inspired by connections between Schrodinger operators, Langevin diffusions and Morse theory, we develop a particle based method that exploits fast relaxation to transition pathways of the so-called Witten PDE on 1-forms (vector fields). Once pathways joining modes are sufficiently well explored, the method relaxes to inexpensive Langevin dynamics for sampling within approximately convex basins. Numerical experiments demonstrate superior performance over cheap gradient based sampling methods and competitive performance at lower cost compared to a state of the art method, parallel tempering.
Beyond MSE: Differentiable Complexity Priors for Structure-Preserving Neural Denoising
Xiangyu Jiang ⋅ Jie Huang
Neural denoisers trained with pointwise distortion losses such as MSE often suppress the temporal regularities that downstream tasks depend on---a failure that becomes acute under non-Gaussian, impulsive, or non-stationary disturbances. We propose a training framework that lifts a family of classical complexity descriptors---permutation entropy, sample entropy, Lempel-Ziv complexity, and Higuchi fractal dimension---into end-to-end differentiable objectives via calibrated soft surrogates with straight-through gradients. Three complementary mechanisms emerge from this lifting: (i) early fusion injects local structural priors as auxiliary inputs, (ii) a complexity-consistency regularizer matches the multi-scale complexity of denoiser outputs to clean references, and (iii) a curriculum sampler prioritizes examples whose noise-induced complexity inflation is largest. We prove that additive noise strictly inflates ordinal uncertainty for bounded piecewise-smooth signals, and that, under bounded surrogate calibration, complexity consistency yields a strictly larger expected decision margin than MSE-only training. On OFDM/QAM communication signals under AWGN, $1/f$ colored, and $\alpha$-stable impulsive noise, our framework reduces post-detection bit error rate by 29--47% over Transformer (PatchTST, iTransformer), GNN, and CNN baselines, with statistically significant gains across five seeds and 95% bootstrap confidence intervals. The improvements transfer to over-the-air conditions on the public RadioML 2018.01A benchmark (single-carrier QAM portion) and a new 50-session in-house SDR testbed, yielding 22--26% relative BER reduction without fine-tuning. The framework is architecture-agnostic and adds $\leq 3.5\%$ FLOPs over the underlying backbone.
Beyond Node Sequences: Relational Diffusion for Unified Graph Learning
Xianan Wang ⋅ Wenji Hu ⋅ Chunyu Wei ⋅ Yueguo Chen
Graph foundation models have emerged as a promising paradigm for generalizing across diverse structural tasks. Current approaches serialize graphs into node token sequences and apply autoregressive or masked prediction, yet this node sequential modeling disrupts the relation-centric nature of graphs and violates permutation equivariance. We propose relational diffusion, a framework that reconceptualizes graph modeling through edge co-occurrence modeling. By treating edges as fundamental modeling units and learning their joint distribution via discrete diffusion processes, our approach naturally respects the relational semantics of graphs while maintaining permutation equivariance. We introduce topological edge tokenization to enrich edge representations with multi-hop structural context, and relational prompt completion to unify node classification, link prediction, and graph classification as sequence completion problems over edge representations. Through extensive experiments, we demonstrate that relational diffusion achieves strong performance while providing a principled foundation aligned with the relational nature of graphs.
Beyond Normal References: Discriminative Few-Shot Anomaly Detection
Huan Wang ⋅ Jun Shen ⋅ Jun Yan ⋅ Guansong Pang
This paper considers a practical few-shot anomaly detection (FSAD) setting, termed discriminative FSAD, where a limited number of both normal and anomalous examples are available as references during inference. Existing FSAD methods rely on normal-only references through normality matching, ignoring the discriminative clues in anomalous references, while directly fitting both references can overfit to the seen anomalies. We introduce IDEAL, an intrinsic deviation learning framework that leverages both reference types to learn intrinsic deviation patterns characterizing generalizable abnormality as deviations from normality. IDEAL decomposes the learning process into two novel components: 1) a Normal Variation Eraser to suppress nuisance normal variations that may lead to noisy deviations from normality, thereby highlighting anomaly-relevant deviation representations; 2) an Intrinsic Deviation Encoder to decompose these denoised deviation representations into intrinsic deviation vectors capturing the most discriminative orthogonal deviation directions. At inference, IDEAL scores query-to-normal deviations preserved after projection onto the learned intrinsic deviation vectors, enabling generalization for both seen and unseen anomalies. Extensive experiments on eight real-world datasets show that IDEAL generalizes effectively to unseen anomalies and consistently outperforms existing state-of-the-art FSAD methods.
Beyond Outcome: Trajectory-Driven Prompt Optimization via Multi-Dimensional Rewards
Zhixiang Liang ⋅ Yixiang Huang ⋅ CHENJING CAI ⋅ Hao Wu ⋅ Peihang Liu ⋅ Qi Liu ⋅ Qiqi Zhong ⋅ Jiuchong Gao ⋅ Jinghua Hao ⋅ Renqing He
Automatic Prompt Optimization (APO) aims to improve the capabilities of Large Language Models (LLMs), yet existing methods remain limited by their result-oriented nature. Training-free approaches often converge to suboptimal prompts, while training-based methods suffer from reward sparsity and optimization instability. More importantly, both paradigms focus on obtaining the final prompt rather than learning the prompt optimization process itself. To address this dilemma, we propose \textbf{ProPO}, a trajectory-driven \textbf{Pro}cess learning framework for \textbf{P}rompt \textbf{O}ptimization guided by multi-dimensional composite rewards. Specifically, ProPO introduces a historical trajectory reflection mechanism for high-quality exploration. Concurrently, we adopt adaptive rubric rewards to provide fine-grained supervision and alleviate reward sparsity. More importantly, we design a novel Process Logic Reward that supervises the consistency of the reflection process. Extensive experiments demonstrate that ProPO, integrated with GRPO, establishes a new SOTA performance. Notably, ProPO achieves a remarkable \textbf{70.00\%} accuracy on the highly complex AIME-2026 task and an absolute gain of \textbf{+25.25\%} over the zero-shot baseline on Subjectivity Classification (Subj). Furthermore, empirical analyses reveal that our multi-dimensional reward mechanism provides finer-grained and denser signals, allowing the model to learn the prompt optimization process, and ultimately enabling it to generate significantly higher-quality prompts.
Beyond [SEG] Tokens: Training-Free Video Reasoning Segmentation via Counterfactual Inference and Contrastive Concept
Tianlu Zhang ⋅ guiguang ding ⋅ Jungong Han
Video Reasoning Segmentation (VRS) demands both complex semantic reasoning and robust spatiotemporal mask propagation. Existing paradigms typically compress reasoning into an isolated, opaque token (e.g., \texttt{[SEG]}) via expensive fine-tuning, or rely on training-free prompt-then-track pipelines that inevitably suffer from semantic degradation and identity switches when encountering hard distractors. To fundamentally address these vulnerabilities, we propose Counterfactual Logic and Explicit Anchoring Reasoning (\textbf{CLEAR}) framework, a training-free VRS framework inspired by human dual-system cognition. Specifically, we introduce a Causal-Counterfactual Inference mechanism that transforms error-prone coordinate regression into a tractable discrete instance selection task, establishing highly reliable initial anchors via rigorous bidirectional reasoning. During mask propagation, we design an adaptive mode-switching mechanism bridged by a Heterogeneous Event Monitor. The fast-thinking mode utilizes a purely visual engine for efficient temporal propagation, while the Monitor continuously evaluates its reliability across multiple dimensions. Upon detecting propagation anomalies, the slow-thinking mode intervenes and explicitly replaces \texttt{[SEG]} tokens with a structured Contrastive Concept Dictionary for closed-loop identity recovery..Extensive experiments demonstrate that CLEAR significantly outperforms existing training-free methods and achieves performance comparable to state-of-the-art training-based approaches on comprehensive referring and reasoning VOS benchmarks. Code will be available after acceptance.
Beyond Stabilization: Dual-EMA Teachers for Global–Local Semantic Learning in Semi-Supervised Medical Image Segmentation
KAIBO LIAO ⋅ Yazhou Ren ⋅ Song Wu ⋅ Kecheng Chen ⋅ Xiaorong Pu ⋅ Lixin Duan
Semi-supervised medical image segmentation has gained increasing attention for its outstanding performance with limited annotations. Most methods follow the Mean Teacher framework, where Exponential Moving Average (EMA) is primarily used to stabilize predictions, while its potential roles beyond stabilization are largely overlooked. In this work, we interpret EMA as a first-order infinite impulse response (IIR) low-pass filter with inherent frequency selectivity, which enables the preservation of semantic information at different levels. Motivated by this insight, we propose a dual-EMA teachers framework, including a long-term teacher focusing on global structures and a short-term teacher emphasizing local details. Furthermore, to help the student better exploit the guidance from both teachers, a First-In-First-Out replay queue is designed to reuse historical high-confidence pseudo-labeled pairs, leveraging pseudo labels provided at different training stages and reducing the sensitivity to sampling bias. Meanwhile, we impose feature-level global semantic and local relation consistency constraints to alleviate the impact of noisy pseudo labels, facilitating robust transfer of complementary representations from dual teachers to the student model. Experiments on three public medical image datasets demonstrate that our method outperforms current SoTA methods, verifying its effectiveness.
Beyond the Dirac Delta: Mitigating Diversity Collapse in Reinforcement Fine-Tuning for Image Generation
Jinmei Liu ⋅ Haoru Li ⋅ Zhenhong Sun ⋅ Chaofeng Chen ⋅ Yatao Bian ⋅ Hongdong Li ⋅ Bo Wang ⋅ Daoyi Dong ⋅ Zhi Wang
Reinforcement learning (RL) has emerged as a paradigm for fine-tuning large-scale generative models, such as diffusion and flow models, to align with complex human preferences and user-specified tasks. A fundamental limitation remains *the curse of diversity collapse*, where the objective formulation and optimization landscape inherently collapse the policy to a Dirac delta distribution. To address this challenge, we propose **DRIFT** (**D**ive**R**sity-**I**ncentivized Reinforcement **F**ine-**T**uning for Versatile Image Generation), an innovative framework that systematically incentivizes output diversity throughout the on-policy fine-tuning process, reconciling strong task alignment with high generation diversity to enhance versatility essential for applications that demand diverse candidate generations. We approach the problem across three representative perspectives: i) **sampling** a reward-concentrated subset that filters out reward outliers to prevent premature collapse; ii) **prompting** with stochastic variations to expand the conditioning space, and iii) **optimization** of the intra-group diversity with a potential-based reward shaping mechanism. Experimental results show that DRIFT exhibits clear Pareto dominance in task alignment and generation diversity, achieving 7.19\%$\sim$93.40\% higher diversity at matched alignment and 13.23\%$\sim$60.13\% higher alignment at matched diversity.
Binary Regression: Universal Ising Model, Binary Expansion and Beyond
Jerry Yao-Chieh Hu ⋅ YI-CHEN LEE ⋅ Mingcheng Lu ⋅ En-Jui Kuo ⋅ Han Liu
We introduce a constructive universality principle for structural interpretability in binary regression. Let $A\in \\{-1,1\\}^{p}$ be binary covariates and $B\in\\{-1,1\\}$ be a binary response. We represent the binary-regression target $\mathbb{E}[B|A]$ by Ising Hamiltonians over observed and auxiliary binary variables. This construction encodes covariate configurations as Boolean states and maps input-output relationships through faithful reductions from Boolean satisfiability (SAT) to ground-state-energy Ising problems. This gives a interpretable graph representation of high-order binary dependence. We showcase this principle with the *Two-body, Linear, Universal, Binary Effect* (T-LUBE) framework. Unlike full power-vector expansion, T-LUBE does not enumerate high-order product features. It recasts these effects through interpretable auxiliary variables and pairwise interactions on an augmented binary graph. Proof-of-concept experiments corroborate our theory.
Biological Graph Priors Enable Representation Learning for Cellular Microscopy
Reed Naidoo ⋅ Matt De Vries ⋅ Georgios Skiadas ⋅ Vicky Bousgouni ⋅ Ugur Moustafa ⋅ Chris Bakal
Learning meaningful representations from cellular microscopy images remains a central challenge in computational biology and computer vision. Recent advances in self-supervised learning (SSL) have shown that scaling model and dataset size can improve the recovery of biological relationships such as gene–gene or compound–target associations. However, this progress has been driven primarily by scale rather than structure—requiring vast data and compute while offering no guarantee that the resulting embeddings capture underlying biology. We introduce NOVA, a structure-first framework for learning biologically aligned representations of cellular perturbations. NOVA integrates curated biological graphs directly into the training objective of a vector-quantised transformer, aligning embeddings of related genes and pathways through graph-based regularisation. By incorporating these relational priors during training, NOVA learns biologically faithful representations with orders of magnitude less data and compute than current state-of-the-art SSL models, while achieving superior performance on gene-gene retrieval benchmarks. This approach reframes representation learning for microscopy from a problem of scale to one of structure, offering a more systematic and efficient route toward biologically grounded embeddings.
BioMicroAgents: A Co-evolutionary Multi-Agent Framework for High-Fidelity Biomicroscopy Imaging
Yongwei Jiang ⋅ Peilin Li ⋅ anfu deng ⋅ Xiaoyu Chen ⋅ Jun Yin
Biomicroscopy imaging involves the complex coordination of multiple procedures, such as focusing, illumination, and staining. Traditional unitary methods struggle with dynamical environments, while hybrid models pre-trained on natural images tend to introduce artifacts for higher perceptual metrics. Notably, biological imaging is not a one-off task. It is an iterative process similar to human expert’s observation, verification and adjustment. They can benefit from history, remembering which processing strategies are suitable and can also identify issues such as artifacts and mis-staining from verification. To this end, we propose BioMicroAgents, the first multi-agent framework with co-evolutionary capabilities designed to simulate and optimize whole biomicro imaging process including brightening, denoising, deblur, super-resolution, and virtual staining. Through reinforcement learning, the model can reflect based on historical experience to reduce artifacts and mis-staining, while enhancing edge and structural features to improve the accuracy of downstream tasks. Specifically, BioMicroAgents includes: (i) Multi-modal Analytical Agents, which utilize multimodal RAG to provide template images and biomicro knowledge, and employ the memory mechanism to analyze optimal processing paths; (ii) We propose three biomicro fidelity metrics and utilize three agent groups (MicroNavigator, MicroProcessor, and MicroAuthenticator) to perform rigorous analysis, processing, and verification. (iii) Co-evolutionary ability, which achieves dynamic policy and self-correction through interactive feedback strategies and memory mechanisms, allowing the agent to evolve from history and other agents. We test model across 12 biomicro tasks, BioMicroAgents achieved improvements in perceptual quality while adhering to strict fidelity in sub-cellular features. It also enhances the accuracy of downstream tasks such as segmentation and cell counting, showcasing the potential of general-purpose visual agents in bioscience.
BioSafetyBench: Agentic Cascade Evaluation for the Bio-AI Ecosystem
Yang Cao ⋅ Xiaogeng Liu ⋅ Shengchao Liu ⋅ Chaowei Xiao
Biological risk is often a property of a cascade rather than a single model output. A generated artifact may become consequential only after transcription, translation, folding, binding, pathway propagation, or host-directed off-target interaction. Existing bio-AI safety evaluations largely grade each model within its own modality, missing downstream risks and inheriting self-grading circularity when model and evaluator share representations. We introduce BioSafetyBench, an agentic cascade-evaluation platform that treats the bio-AI ecosystem as the unit of safety evaluation. BioSafetyBench routes each model-output pair to one of 15 standardized probes over a six-level hierarchy from genome to organism-level outcome. Fixed bridge agents propagate outputs through two causal pathways: Pipeline A follows forward biological cascades, while Pipeline B evaluates CRISPR and siRNA designs through genome- and transcriptome-wide off-target alignment. At every reached level, external frozen evaluators compute domain-native biomarkers and compare them with literature-anchored thresholds, yielding a Cascade Concern Profile and Cascade Depth. Across 25 foundation models, BioSafetyBench surfaces risks invisible to single-modality grading: Influenza neuraminidase designs from ProteinMPNN and ESM-IF1 reach CD = 4; jailbreak prompts inflate gRNA off-target hits 9.05× over baseline, push 48 of 604 gRNAs to Critical, and concentrate Critical siRNA hits on MYC; protein mask-and-fill spans a 32× cross-model gap; GPT-4o complies on all natural-language-guided protein mutation prompts at CD = 2 while Claude Sonnet 4.5 refuses all prompts at CD = 0; and a three-predictor ADMET ensemble misses 30.3% of known clinical toxics. BioSafetyBench is, to our knowledge, the first bio-AI safety benchmark to provide cascade-level risk certificates whose scoring modules are weight- and gradient-flow independent from the model under test by construction.
Block-OBS-GS: Exact Per-Block Joint Brain Surgery with Gauss–Seidel Refinement for LLM Pruning
Yuwen Huang ⋅ Xiang Pan
Post-training pruning of large language models (LLMs) zeroes a fraction of the weights and reconstructs the survivors without retraining, addressing the weight-memory bottleneck of LLM serving. At a fixed mask the per-row reconstruction admits a closed-form Optimal Brain Surgeon (OBS) optimum, but the cost of the exact joint solve is cubic in the layer dimension; one-pass methods such as SparseGPT and Wanda approximate it at quadratic cost without attaining the per-block joint optimum. We introduce Block-OBS-GS, which solves the per-block joint OBS reconstruction exactly within SparseGPT's complexity class via a per-row Cholesky factorisation of the in-block damped Hessian (one third the cost of an explicit inverse, reused across the algorithm) and follows it with a single cross-block Gauss--Seidel sweep. We prove three guarantees: per-layer FLOP count in SparseGPT's complexity class, with the leading $mn^{2}$ constant dropping from $1/2$ (SparseGPT) to $3\rho/2$ (Block-OBS-GS) where $\rho=1-s$, hence strictly lower for sparsity $s>2/3$; per-row reconstruction error at most SparseGPT's, strict whenever the activation Hessian has non-zero off-diagonal coupling; and a closed-form distance bound to the unregularised joint-OBS optimum that contracts geometrically per sweep. Across six Llama-3 and Qwen-3 backbones from 3B to 70B parameters, Block-OBS-GS attains the lowest WikiText-2 perplexity at every tested sparsity on the three Llama-3 backbones, with a $47\%$ perplexity reduction over SparseGPT on Llama-3.1-8B at $s=0.875$; the highest seven-task commonsense Avg7 on six of twelve populated cells; 5-shot Massive Multitask Language Understanding (MMLU) within $0.026$ of SparseGPT across populated cells; and the lowest WikiText-2 perplexity on four of five backbones under $2{:}4$ structured sparsity.
Boosting Multiagent Reinforcement Learning at High Replay Ratios with Ensemble Reset
Yaodong Yang ⋅ Hongyao Tang ⋅ Guangyong Chen ⋅ Pheng-Ann Heng
Increasing the replay ratio, where an agent's network is updated multiple times per environment interaction, is an effective strategy for improving sample efficiency in reinforcement learning. However, its effects on the network capacity of multiagent reinforcement learning (MARL) are not yet well investigated. In this paper, we show that high replay ratios induce a severe dormant neuron problem in the centralized global Q-network of MARL, where a large fraction of neurons become inactive, thereby reducing network capacity and destabilizing learning. To address this problem, we propose Ensemble Reset (EnSet) to stabilize MARL training at high replay ratios. First, guided by theoretical analysis, EnSet employs an ensemble of global Q-value networks with periodic resets to mitigate neuron dormancy during frequent updates. Second, EnSet diversifies replay experience by exploiting multiagent translation invariance as an inductive bias in the global Q-value function to prevent overfitting. Extensive experiments in SMAC, MPE, and SMACv2 environments demonstrate that EnSet consistently improves various MARL algorithms at high replay ratios with fewer environment steps.
Boost Reasoning Evolution via Responsive Rollout Difficulty Manipulation
Ziheng Li ⋅ Zexu Sun ⋅ Jinman Zhao ⋅ Erxue Min ⋅ Yongcheng Zeng ⋅ Hui Wu ⋅ Hengyi Cai ⋅ Shuaiqiang Wang ⋅ Xu Chen ⋅ Zhihong Deng ⋅ Dawei Yin
Reinforcement learning with verifiable rewards (RLVR) has advanced the reasoning capabilities of large language models (LLMs). However, existing RLVR methods often suffer from exploration inefficiency due to mismatches between problem difficulty and model capability: overly difficult problems hinder reasoning path discovery, while overly simple problems offer little learning signal. To address this, we first formalize the effect of problem difficulty by quantifying the relationship between loss descent magnitude and rollout accuracy. Building on this analysis, we propose SEELE, a supervision-aided RLVR framework that dynamically adjusts problem difficulty to lie within the high-performance region. SEELE augments each training sample by appending a hint (part of a full solution) for difficulty reduction. Unlike previous hint-based approaches, SEELE deliberately computes the hint length for each individual problem to achieve an optimal difficulty. The optimal hint length is determined via multi-round rollout sampling, where an item response theory model fits accuracy–hint pairs from previous rounds to predict the next-round hint. This instance-level, real-time difficulty adjustment aligns problem difficulty with the evolving model capability, thereby improving exploration efficiency. Experiments show that SEELE outperforms Group Relative Policy Optimization (GRPO) and Supervised Fine-tuning (SFT) by +10.0 and +8.4 points, respectively, and exceeds the best prior supervision-aided approach by +3.8 points on average across six math reasoning benchmarks.
Boundary Mass in Feature Space as a Label-Budget Diagnostic for Representations
Siming Zhang ⋅ Zhehui Shen ⋅ Shijie Chen ⋅ Yansen Yu ⋅ Hang Yu
Low-label transfer is often decided by a practical question: after a representation looks useful, how many labels are still needed to learn the downstream boundary? Existing transferability and probe-quality scores rank global alignment, likelihood, margins, or calibration, but they do not measure how much target data remains close to the fitted decision boundary, where scarce labels are most costly. We propose effective boundary complexity, a label-budget diagnostic for frozen representations and linear probes. It measures mass in a thin strip around the learned boundary and normalizes it by class balance, so dense ambiguous regions and rare-class difficulty are both counted. We prove that, under a smooth representation map, this boundary mass does not change because of generic volume expansion: density change and tangential surface change cancel at first order, leaving only stretch or compression in the boundary-normal direction. This yields a cached-feature estimator that fits a probe on one split, counts held-out features near its boundary, and avoids high-dimensional density estimation. Experiments support the diagnostic at three levels. Controlled transformations match the predicted law across 74 maps with 0.38% median relative error. On real frozen features, normal-direction changes move both measured boundary mass and label demand, while tangent changes do not. A differentiable surrogate reduces measured complexity by 16–36% at matched top-1 utility. Finally, CIFAR-100 and Tiny-ImageNet audits show boundary mass acting as a second-stage diagnostic for near-tied representations, with external multiclass results showing how label-grid saturation limits policy-level gains on label grids where many tasks tie.
BQ-LoRA: Binary-Quantized Low-Rank Adapters as Implicit Regularizers for Parameter-Efficient Fine-Tuning
Juyoung Park ⋅ Heejun Lee
Parameter-efficient fine-tuning of quantized large language models is dominated by QLoRA, which combines 4-bit weight quantization with FP16 low-rank adapters. We show the FP16 adapter precision is unnecessary: a single binary bit per adapter parameter, combined with a learned per-layer scale $s$ and a Scaled Straight-Through Estimator, matches or exceeds QLoRA across a broad range of benchmarks while reducing adapter storage by $14$--$16\times$. We introduce two variants--NormScale, which clips STE gradients to the unit ball, and LearnThresh, which makes the binarization threshold learnable---unified under a framework we call BQ-LoRA. Our main theoretical contribution is a tight characterization of the BQ-LoRA hypothesis class: it has bounded entrywise norm, a high-probability spectral norm bound that is $\sqrt{r}$ times sharper than the worst case, a generically full-rank update, and consequently a Rademacher complexity that yields an explicit $\tilde{O}(s_{\max}\sqrt{d_{\text{in}} d_{\text{out}}/n})$ generalization bound. None of these properties hold a priori for FP16 LoRA. On Qwen2.5 (1.5B/7B/72B) and LLaMA-2 (7B/70B), BQ-LoRA matches or beats QLoRA and LoftQ at $\sim 1/16$ the adapter footprint; the predicted regularization signature appears empirically in dataset ablations and training-loss curves. We extend Punica's SGMV kernel with a binary inner-product primitive (BQ-SGMV) and prove a cache phase-transition result: shrinking adapters $16\times$ enlarges the HBM cache from $2{,}500$ to $39{,}000$ adapters, eliminating misses under skewed workloads.
BrainCoT: A Multi-Task Zero-Shot Brain Signal Foundation Model with Neurometric-Anchored Chain-of-Thought Reasoning
Yiyu Gui ⋅ Mingzhi Chen ⋅ Guibo Luo ⋅ Yuchao Yang
Brain signals exhibit substantial multi-center heterogeneity, strong inter-subject variability, and highly diverse task settings, which makes it difficult to build models that generalize beyond a single task. While recent brain signal foundation models leverage large-scale pretraining to learn generalizable representations, they often fall short in practice because they still require target-data-specific supervised adaptation, lack multi-task zero-shot capability, and rarely provide verifiable evidence to support their predictions. In this work, we present BrainCoT, a multi-task zero-shot Brain signal foundation model with neurometric-anchored Chain-of-Thought (CoT) reasoning. BrainCoT unifies heterogeneous tasks into an instruction-driven framework so that, after pretraining, a single model can be directly applied to multiple tasks in zero-shot settings without additional adaptation. It further introduces neurometric-anchored chain-of-thought grounded in neurometric evidence for verifiable reasoning, and adopts Decision-Consistent Rationale Training to keep process learning aligned with correct predictions. Across extensive evaluations, BrainCoT achieves an average ACC of 67.41\% and AUC of 74.33\% in zero-shot classification without downstream training data, outperforming the strongest linear-probing baseline trained with 1\% labeled downstream data by 3.85\% in ACC and 5.97\% in AUC. Code will be released upon acceptance.
BrainVista: Modeling Naturalistic Brain Dynamics as Multimodal Next-Token Prediction
Xuanhua Yin ⋅ Ted Zhao ⋅ Lina Yao ⋅ Weidong Cai
Naturalistic fMRI offers a time-resolved window into cortical dynamics during continuous multimodal experience, where future brain states are jointly governed by endogenous neural context and exogenous sensory drive. Forecasting such dynamics, however, remains challenging because fast-changing stimulus streams must be reconciled with the sluggish hemodynamic BOLD response, while cortical activity is organized across functionally heterogeneous brain networks. To address these challenges, we introduce BrainVista, a multimodal autoregressive framework for observed-stimulus conditional brain-state forecasting. BrainVista predicts future fMRI states from the history of brain activity and stimulus tokens temporally aligned to each prediction query, enabling autoregressive prediction without access to future ground-truth fMRI states. This design supports causal, temporally consistent modeling of naturalistic brain dynamics under realistic sensory stimulation. Specifically, we design Network-wise Tokenizers for functionally structured brain representations, a Spatial Mixer Head for cross-network interaction refinement, and Stimulus-to-Brain masking to preserve temporally causal stimulus–brain conditioning while preventing future brain-state leakage and stimulus look-ahead. We evaluate BrainVista on Algonauts 2025, CineBrain, and HAD, three large-scale naturalistic fMRI benchmarks, under long-horizon conditional autoregressive forecasting. BrainVista consistently outperforms baselines, improving pattern correlation by 10.8\% and 9.4\% relative to the strongest baseline on Algonauts 2025 and CineBrain, respectively.
Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning
Haoxuan Chen ⋅ Tianming Liang ⋅ Wei-Shi Zheng ⋅ Jian-Fang Hu
Reinforcement learning with verifiers (RLVR) has become a central paradigm for improving LLM reasoning, yet popular group-based optimization algorithms like GRPO often suffer from exploration collapse, where the models prematurely converge on a narrow set of high-scoring patterns, lacking the ability to explore new solutions. Recent efforts attempt to alleviate this by adding entropy regularization or diversity bonus. However, these approaches do not change the \textit{winner-takes-all} nature, where rollouts still compete for individual advantage rather than cooperating for maximizing global diversity. In this work, we propose Group Cooperative Policy Optimization (GCPO), which shifts the training paradigm from rollout competition to team cooperation. Specifically, GCPO replaces independent rollout scoring with team-level credit assignment: a rollout is rewarded by how much it contributes to the team's valid solution coverage, rather than its individual accuracy. This coverage is described as a determinant volume over reward-weighted semantic embeddings, where only correct and non-redundant rollouts contribute to this volume. During advantage estimation, GCPO redistributes the collective team reward to each single rollout according to its average marginal contribution to the team. This cooperative training paradigm routes optimization toward non-redundant correct reasoning paths. Experiments across multiple reasoning benchmarks demonstrate that GCPO significantly improves both reasoning accuracy and solution diversity over existing approaches. Code will be made available.
Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models
Yi Zhong ⋅ Haotong Qin ⋅ Xindong Zhang ⋅ Lei Zhang ⋅ Guolei Sun
Low-bit post-training quantization (PTQ) is a pivotal technique for deploying Vision-Language Models (VLMs) on resource-constrained devices. However, existing PTQ methods often degrade VLMs' accuracy due to the heterogeneous activation distributions of text and vision modalities during quantization. We find that this cross-modal heterogeneity is distributed unevenly across channels: a small subset of channels contains most modality-specific outliers, and these outliers typically reside in different channels for each modality. Motivated by this, we propose SplitQ, a channel-Splitting-driven post-training Quantization framework. At its core, SplitQ introduces a novel Modality-specific Outlier Channel Decoupling (MOCD) module that effectively isolates salient modality-specific outlier channels with minimal overhead. To further address the remaining cross-modal distribution discrepancies, we design an Adaptive Cross-Modal Calibration (ACC) module that employs dual lightweight learnable branches to dynamically mitigate modality-induced quantization errors.Extensive experiments on popular VLMs demonstrate that SplitQ significantly outperforms existing approaches across 6 popular multi-modal datasets under all evaluated quantization settings, including W4A8, W4A4, W3A3, and W3A2. Notably, SplitQ preserves 95.3% of FP16 performance under the challenging W3A3 setting (69.5 vs. 74.3), pushing the efficiency frontier for deploying advanced VLMs.
Breaking the Group Size Barrier: Parameter-Efficient Group Dance Generation with Chain-of-Dancers
Jing Xu ⋅ Cunjian Chen ⋅ Qiuhong Ke
Group dance generation aims to synthesize coordinated multi-dancer choreography from music, with broad applications in animation and interactive content creation. This task requires modeling dense inter-person dependencies to ensure spatial coordination, while naturally preserving individual dancer identities. Existing approaches model all dancers jointly with end-to-end transformers, which tie the architecture to a fixed group size and entangle per-dancer identities across frames. We propose ChainDance, a scalable framework that reformulates group dance generation as a Chain-of-Dancers: a sequential decomposition over per-dancer conditional distributions, allowing a single model to scale across arbitrary group sizes without retraining and naturally preserving per-dancer identity. Built on a frozen single-dancer diffusion backbone, ChainDance introduces two lightweight modules: a Role-Aware Text Encoder (RATE) for per-dancer semantic conditioning, and a Group-Aware Motion Encoder (GAME) that aggregates previously generated dancers via a distance-weighted graph convolutional network, and incorporates a training-free noise optimization procedure at inference time to enforce global spatial coherence. Experiments on AIOZ-GDance demonstrate that ChainDance achieves state-of-the-art performance in motion quality and identity preservation, with 3-4 times fewer parameters and requiring 3-6 times less training time compared to prior approaches.
Brenier Meets Adversarial Training: Optimal Transport Geometry for Robust Learning
alireza abdollahpour ⋅ Ehsan Sharifian ⋅ Buse Şen ⋅ Marco Cuturi ⋅ Daniel Kuhn
Distributionally robust optimization (DRO) provides a principled framework for learning under distribution shift, but its practical use is hindered by the difficulty of evaluating worst-case risks for nonconvex loss functions. We study a penalized DRO formulation in which the adversary may choose any distribution but incurs a Wasserstein penalty for deviating from the empirical distribution. We show that the adversary’s problem can be reformulated as an optimization problem over transport maps that push empirical samples to adversarial ones, and we prove that optimal maps are cyclically monotone. We also show that standard adversarial training---based on per-sample local optimization---violates cyclical monotonicity and wastes transport costs unless the adversary is severely restricted. We propose two remedies. First, we introduce multi-start particle ascent, which alternates parallel gradient ascent with reassignment to enforce cyclical monotonicity across samples. Second, we parameterize adversarial maps as gradients of input-convex neural networks, which guarantees cyclical monotonicity by construction. Experiments on robust regression, image classification, and robust control show that our methods consistently outperform standard adversarial training and state-of-the-art baselines, achieving improved robustness and better generalization under distribution shift.
BridgeTwist: Twisting Schrödinger Bridges for Training-Free Conditional Sampling
Mengyu Li ⋅ Qianqian Qu ⋅ Jun Liu
Generative models are often required to serve many downstream conditional tasks. In the training-free setting, a pretrained unconditional model is reused as a prior and combined with a task-specific likelihood to sample from the corresponding posterior, without retraining the generative model. While such conditional samplers have been developed for diffusion and flow generative models, training-free conditional sampling for Schrödinger bridge (SB) generative models remains unexplored, despite the growing appeal of SBs as flexible stochastic transport models between general source and target distributions. We propose BridgeTwist, a training-free conditional sampler for pretrained SB generative models, formulated as a tempered twisted sequential Monte Carlo scheme. At each step, the sampler uses a closed-form plug-in twist surrogate, which combines the velocity and score fields exposed by SB pretraining into an explicit endpoint predictor. Since this surrogate is less reliable near the source, we use a time-varying temperature to temper the twist, downweighting its influence at early steps. Telescoping importance weights then cancel all intermediate twist factors, yielding an asymptotically exact conditional sampler without any auxiliary twist network or pilot rollouts. Empirically, BridgeTwist consistently outperforms training-free baselines on class-conditional sampling and inpainting on MNIST and CIFAR-10, and on text-to-image generation on CelebA-HQ.
Bridging CLIP with DINO: Cross-Modal Information Maximization for Online Test-Time Adaptation
Taehoon Kim ⋅ Wonjun Hwang ⋅ Hyunsouk Cho
Vision-language models such as CLIP enable zero-shot classification from arbitrary class names, but under test-time corruptions, their visual encoder projects distorted features onto incorrect textual anchors with high confidence—a phenomenon we term \textbf{modality bias}. Standard entropy minimization reinforces these overconfident errors, triggering a destructive confirmation-bias loop that collapses the predicted label distribution. We propose \textbf{CDIM} (Cross-Modal Information Maximization), an online test-time adaptation method that breaks this loop at three levels. At the \textit{representation level}, a frozen self-supervised DINO backbone is combined with CLIP via a cross-covariance logit ensemble; by projecting text prototypes into DINO's geometric space—rather than the reverse—the classifier operates free of text bias. At the \textit{optimization level}, a gated information-maximization objective (GT-IM) pairs per-sample sharpening with batch-level diversity, while an adaptive gate excludes unreliable samples from the gradient. At the \textit{prediction level}, an online logit adjustment corrects the residual class-distribution skew. Ultimately, \textbf{CDIM} validates its broad effectiveness by delivering robust, competitive performance across demanding setups, including continual TTA, mixed-domain shifts, and diverse domain generalization benchmarks.
Classical compute-optimal scaling laws assume an unbounded supply of fresh pretraining data, yet pretraining is increasingly entering the regime where compute is growing faster than high-quality data. We propose *Compute-Data (CD) scaling laws*, a unified framework that bridges *compute-optimal* scaling, in which data scales freely with compute, and *data-optimal* scaling, in which the corpus is fixed and compute can grow unbounded. CD scaling extends classic scaling by introducing a *token effectiveness function* that quantifies how much a *derived* token, produced for instance by multi-epoch repetition or paraphrasing, is worth relative to a fresh one, ranging from a perfect substitute to no value at all. Fitting $\eta$ for two data expansion strategies (multi-epoch repetition, paraphrasing) across model sizes ranging from 14M to 600M parameters on the Dolma-3 corpus, we find that it is far from constant: it depends jointly on model size, the tokens-per-parameter ratio, and the amount of derived data, and saturates as the corpus is expanded. The functional form of $\eta$ implies that substituting compute for data *diminishes* with both model size and data availability, and partitions training into three operational regimes: compute-bound, data-bound, and model-bound, showing that classic compute-optimal allocation is suboptimal across most of the practically relevant regime.
Bridging Risk Approximation Gaps in Model Predictive Task Sampling via In-Context Modeling
Jiarong Wen ⋅ Qi Tao ⋅ Zhang Kaiyu ⋅ Yun Qu ⋅ Lizhou Cai ⋅ Heming Zou ⋅ Yuhang Jiang ⋅ Qi Wang
Reinforcement learning (RL) is essential for adaptive decision-making, e.g., robust Meta-RL and domain randomization (DR), and efficient post-training of large language models (LLMs), where tasks are commonly structured as Markov decision processes or prompts. A central bottleneck in all these settings is the cost of agent-environment interactions required for task selection and policy optimization. Model predictive task sampling (MPTS) has emerged as a promising approach to improve sample efficiency by querying informative tasks via a risk predictive model (RPM) trained on optimization history. However, existing RPMs suffer from two failure modes: (i) sparse optimization histories, which limit RPMs' generalization across the task space, and (ii) a one-step lag in risk approximation, which causes biased difficulty estimation as the policy evolves. This work formalizes these failure modes, derives lagged-risk decompositions that expose the resulting temporal bias, and recasts risk prediction as an in-context sequence modeling problem. The developed RPM integrates a one-step look-ahead mechanism with a temporal-weighted sliding window to mitigate data non-stationarity and scarcity, without requiring additional rollouts. Empirically, our approach yields more reliable difficulty estimates and consistent performance gains across robust Meta-RL, DR, and prompt curriculum RL post-training of LLMs.
Bridging the Gap: Position-Independent Cache Reuse for Hybrid SSM-Attention Architectures
Weikang Wang ⋅ Xin Zhou ⋅ Haiyang Liu ⋅ Lei Wang ⋅ Jun Liu ⋅ Weifeng Zhang
Hybrid large language models (LLMs) integrating state space models (SSMs) and full attention face a critical caching bottleneck due to mismatched caching granularities. While full attention allows flexible, token-level cache reuse, SSMs are always restricted to coarse-grained block- or request-level reuse. Consequently, an SSM layer's cache miss can easily invalidates successful fine-grained hits in adjacent attention layers, disrupting the entire reuse pipeline. To resolve this, chunk-level position-independent cache (PIC) reuse is highly desirable to align the caching granularity across all layers. However, enabling PIC reuse for history-dependent SSMs remains fundamentally challenging. To bridge this gap, we propose HyPIC, the first unified framework achieving PIC reuse for hybrid architectures. We introduce residual state propagation and online decay correction, which efficiently adapt offline SSM chunk caches to new contexts by explicitly recomputing minimal boundary tokens. Furthermore, we design a fast parallel recomputation pipeline exploiting the linear superposition property of SSMs to eliminate sequential bottlenecks. Experiments on diverse hybrid models and datasets demonstrate that HyPIC achieves a 2$\times$ speedup in time-to-first-token (TTFT) while maintaining highly competitive accuracy.
BudSplat: Feed-forward 3D Gaussian Splatting under a Rendering Budget
MUYU XU ⋅ Fangneng Zhan ⋅ Yu Wei ⋅ Ling Shao ⋅ Shijian Lu
Feed-forward 3D Gaussian Splatting has recently emerged as a promising alternative to per-scene optimization, enabling fast novel view synthesis by predicting Gaussian primitives directly from input images. However, it quickly faces a bottleneck of rendering efficiency while scaling from sparse to dense multi-view inputs, which offer more geometric constraints but produce a rapidly growing set of candidate Gaussians. Existing solutions cannot reliably distinguish redundant primitives, leading to many redundant primitives from repeatedly observed regions and increased rendering cost without much gain in visual quality. We propose BudSplat, a compact feed-forward 3DGS framework that learns to keep the most valuable Gaussians under a limited primitive count. BudSplat has two novel designs. The first is a detail-aware importance score that is extracted from multi-level visual tokens, providing an explicit cue for regions where pruning is likely to harm perceptual quality. The score can be injected into voxel fusion to preserve high-value local contributors during multi-view aggregation. The second is quality-aware Gaussian pruning that selects a compact set of primitives by jointly considering rendering contribution and detail importance. The quality-aware pruning enables local quota redistribution that helps avoid over-allocating primitives to large redundant regions. Experiments on RE10K, DL3DV, and ACID show that BudSplat achieves a favorable trade-off among rendering quality and efficiency.
Most coding-agent benchmarks ask whether generated code behaves correctly. That remains essential, but repository-level engineering is increasingly agent-managed: one agent writes a repository, and later agents inspect, audit, or extend it as working context. In that setting, a generated repository is not only an answer to a task but also a communication artifact for future work. Even when strong agents nearly satisfy the visible behavioral objective, repositories can differ in how clearly they expose the intended behavior and design choices behind that behavior. We introduce Build-and-Find, a protocol for evaluating whether downstream agents can recover those intended choices from generated repositories, and how much inspection that recovery requires. For each task, a builder sees a hidden repository specification and creates a codebase; a finder sees only the codebase and a specification-traced multiple-choice question bank. The protocol separates behavioral correctness from artifact-side recovery and reports recovery accuracy, repeatability, implementation coverage, and inspection effort. Accuracy and stability act as gates: effort is interpreted only when recovery succeeds reliably. Among artifacts from which the same intent can be recovered, lower effort by the same finder suggests that the artifact makes that intent easier to locate. Question-only and spec-only controls quantify generic priors and specification access, while audits separate omitted claims from finder failures and check whether correct answers cite artifact evidence. In the released high-prior task pack, recovery accuracy is near saturation, so inspection effort and finder-specific effects provide the main panel-local comparison. We release the harness, tasks, generated artifacts, records, canonical tables, reports, scripts, metadata, licenses, and evidence audit to support auditable extensions and private task packs.
Busemannformer: Horospherical Self-Attention for Hyperbolic Graph Transformers
Youheng Yao ⋅ Ziyao Zeng ⋅ Wenbo Liao ⋅ Tianqi Wang
Hyperbolic neural networks excel on hierarchical graph data, yet existing hyperbolic Transformers lift standard attention into curved space without using the ideal boundary, the structure most directly tied to tree hierarchy. We propose \textbf{Horospherical Self-Attention} (HSA): each query selects an ideal-boundary direction and scores keys by their Busemann depth along that direction, yielding a token-to-token hyperbolic attention kernel. We prove flat-limit consistency (HSA recovers dot-product attention as $K\to 0^-$), characterise score level sets as horospheres, and bound key-side gradients. HSA is integrated into \textbf{Busemannformer}, a full hyperbolic graph Transformer. Busemannformer-LD reaches \textbf{91.78\%} F1 on Disease-NC; in a controlled score ablation, its log-depth Busemann score achieves the best Disease-NC result among tested scores, reaching \textbf{93.29\%} F1 and performance comparable to the current SOTA Hypformer result with a smaller training budget. For link prediction, we find a score-symmetry principle: Busemannformer-LD with a symmetric score achieves \textbf{90.71\%} AUC on PubMed-LP ($+$10.3 pp over the asymmetric variant), while the asymmetric score performs best on Disease-LP among our variants and outperforms all Euclidean baselines. Ablations identify log-depth Busemann scoring as the key component and reveal a clear failure mode on Airport, where labels track centrality rather than tree depth.
C2G-BENCH: A Cyber-Physical Evaluation Benchmark for Hierarchical Reinforcement Learning in Grid-Interactive Hyperscale Data Centers
Vineet Gundecha ⋅ Sahand Ghorbanpour ⋅ Sifat Abdullah ⋅ Mahasweta Chakraborti ⋅ Damien Fay ⋅ Avisek Naug ⋅ Sergey Serebryakov ⋅ Ricardo Luna Gutierrez ⋅ Soumyendu Sarkar
The rapid growth of generative AI is reshaping the energy footprint of modern computing infrastructure. Hyperscale AI data centers now operate at power levels comparable to large industrial facilities, creating new challenges for grid reliability, energy procurement, and real time demand flexibility. At the same time, their large controllable loads, cooling systems, and on site storage make them potential participants in grid services such as frequency regulation. However, existing evaluation environments do not capture the coupled decision making required for grid interactive operation. We introduce C2G-BENCH, a cyber-physical benchmark for hierarchical reinforcement learning in hyperscale data centers. C2G-BENCH simulates a real-world scenario, where a 250 MW-class hyperscale data center participates in grid frequency regulation while serving realistic AI workload demand under weather-driven ambient conditions. The benchmark explicitly links 15-minute market level orchestration with 5-second physical control, enabling agents to jointly reason over regulation commitments, workload flexibility, cooling, battery dispatch, and facility safety. C2G-BENCH achieves this through two coupled Gymnasium environments: C2GMacroEnv, where a high level controller bids into the regulation market, and C2GFastEnv, where a low level controller actuates the Dynamic Voltage and Frequency Scaling (DVFS), Coolant Distribution Unit (CDU) pump speed, HVAC effort, and Battery dispatch (BESS). The simulator models a 250 MW class two zone facility using seven modular physics engines covering workload, thermal, electrical, battery, grid signal, weather, and market dynamics, driven by real workload and weather traces. We benchmark a family of rule-based, LLM-based & RL-based controllers for both the macro and fast environments, and evaluate them on four stress-test environments in multiple grid markets. Baseline results show that active low level control substantially improves tracking quality: a rule-based macro controller paired with a RL-based low level controller reduces the tracking RMSE from 2,737 kW to 229 kW compared with a macro only configuration, while maintaining full throughput and zero thermal violations. These results demonstrate that C2G-BENCH exposes meaningful cross layer tradeoffs and provides a reusable benchmark for studying safe, hierarchical control of grid interactive AI data centers.
C3: Long-Horizon Character Consistency via Causal-Continuous State Dynamics and Memory Rewriting
Minghao Chen
Maintaining character consistency over long-horizon gameplay is a critical bottleneck for deploying Large Language Models (LLMs) as Non-Player Characters. Prevalent discrete Affective State Machines (ASMs) fail due to two paradoxes: (1) the \textit{Cliff-edge Effect}, where threshold-based switching causes jarring narrative discontinuities; and (2) \textit{Auto-regressive Inertia}, where accumulated dialogue history dilutes current state instructions. We propose \textbf{C3 (Causal-Continuous Characterization)}, which couples a continuous state space with \textit{hysteresis-aware dynamics} to dampen emotional volatility, and a \textit{State-Conditioned Memory Rewriting} mechanism that re-contextualizes past events through the lens of the current relationship. On a 100-turn dynamic simulation (\textsc{Chronos-Sim}, 5{,}000 turns per condition), C3 achieves a \textbf{10.5\% absolute SAR improvement} over a strong smoothed-ASM baseline (Welch's $t$-test, $p < 0.0001$, Cohen's $d \approx 3.97$). Cross-model evaluations on \texttt{Llama-3-8B-Instruct} and \texttt{Mistral-v0.2-7B} reproduce the relative gain, indicating the mechanism transfers across backbones although absolute SAR remains backbone-dependent. A human study with 20 game designers (ICC$=0.74$) corroborates higher perceived consistency, and a separate factual audit (Fleiss' $\kappa=0.81$) confirms the Rewriter alters interpretation without corrupting facts. Deployment analysis shows a $\sim$60\% reduction in long-session token consumption.
C3P: Contrastive promoter-protein pretraining yields representations capturing bacterial gene regulation
Cameron Dufault ⋅ Scott Xu ⋅ Alan Moses
Despite the increasing scale of genome language models (gLMs), their ability to decode the function of regulatory sequences remains unclear. gLM pretraining relies on sequence reconstruction, which may struggle due to the noisy, rapidly evolving nature of regulatory DNA. Self-supervised contrastive approaches provide a promising alternative. Inspired by language-image architectures like CLIP, we demonstrate contrastive promoter-protein pretraining (C3P). By learning to align promoters to their corresponding proteins, we leverage the rich representations of proteins learned by protein language models as supervisory signal for the learning of promoter representations. After training on 88 million bacterial promoter-protein pairs, we evaluate the predictive power of C3P-learned promoter representations for inference of curated regulatory annotations, finding multi-fold improvement over leading gLMs. We also introduce zero-shot co-regulated gene retrieval, the ability to find co-regulated genes in a genome using no experimental data. We find that compared to a randomly initialized baseline, C3P training consistently provides significant zero-shot performance gains, unlike gLMs. Scaling analysis reveals the potential for further improvement as well as the efficiency of C3P, which achieved strong performance at a fraction of the training cost of leading gLMs. In addition to demonstrating that C3P training is effective for learning representations of bacterial regulatory sequences, our strong zero-shot co-regulated gene retrieval performance suggests the possibility of decoding gene regulation for millions of bacteria from their genomes alone.
CacheMAS: Cache Communication for Single-Pass, Jointly-Optimized Multi-Agent Systems
Chao Ouyang ⋅ Yuyang Bai ⋅ Jun Zhang ⋅ Tianlu Gao ⋅ YuxinDai ⋅ xu peidong ⋅ David W Gao
Multi-agent LLM systems (MAS) improve reasoning by decomposing across specialized roles, but their text communication is expensive. Each agent decodes thousands of intermediate tokens for the next, and successive prefills grow cumulatively. Recent work shows that latent-space signals can substitute for explicit text between agents, cutting cumulative prefill cost. When the latent reasoning also replaces an agent's text reasoning, it additionally cuts intermediate tokens. Here, we find that the prefill cache alone suffices as the latent-space signal between agents, and the per-agent forward chains can be compressed into a single pass. With \emph{exactly one prefill and one decode per problem}, we sharply cut per-rollout tokens and prefill wall time. We call this architecture \textbf{CacheMAS}, in which prefill agents pass information through cache communication, and a final agent decodes the answer. As a bonus, \textbf{CacheMAS} is one autoregressive rollout, so unlike prior MAS, standard sequence-level RL can jointly optimize all agents at once without per-agent reward design. On math, multi-hop QA, and code, CacheMAS uses $\sim$$3\times$ fewer tokens, runs $\sim$$2\times$ faster, and after joint optimization beats prior MAS by $+1$ to $+8$ points on every benchmark. Cache communication is not just a token-efficient substitute for text, but a practical interface for joint MAS optimization.
Cache the Future: Training-Free Self-Revision for Diffusion Transformer Acceleration
Yiyang Li ⋅ Weixuan Huang ⋅ Wei Zhang
Although recent training-free acceleration methods have achieved promising speedups for Diffusion Transformers, most rely on historical features, either by directly reusing cached features or by forecasting future features from them. Such a history-to-future design can accumulate errors when the prediction horizon becomes long. To address this issue, we propose FuCa, a Cache-the-Future training-free acceleration framework. Instead of extrapolating future information solely from the past, FuCa caches sparsely computed future features and uses them as online revision signals. FuCa revises the latent-state update and the corresponding velocity, and then recovers skipped timesteps through lightweight interpolation. It requires no extra training or model-specific auxiliary modules, while keeping the number of DiT forward evaluations low. Experiments on FLUX.1-dev, Qwen-Image, Wan2.1-1.3B, and Hunyuan Video show that FuCa achieves up to about 5$\times$ speedup for image generation and more than 4.3$\times$ speedup for video generation, with higher reconstruction quality than existing training-free baselines.
CADMA: Capacity-Aware Recall Decomposition for Generative Model Assessment
Debin Meng ⋅ Zheng Gao ⋅ Ziquan Liu ⋅ Yanran Li ⋅ Ioannis Patras ⋅ Georgios Tzimiropoulos
Evaluating generative models requires measuring not only sample fidelity but also how well a generated set covers the real data distribution. Recent work has therefore introduced recall-oriented metrics for distribution-level evaluation. However, existing recall metrics often reduce distributional coverage to geometry-based coverage. This simplifying assumption weakens their diagnostic value and interpretability in two ways: these metrics provide only a single recall score that cannot distinguish essential diversity failures such as mode dropping and mode collapse, and geometry-based coverage can provide unreliable estimates by assigning high recall to poorly covered generative distributions. We therefore introduce CADMA, a capacity-aware framework for recall-oriented evaluation and diagnosis. CADMA evaluates distributional coverage by formulating local real--fake allocation as a one-to-one matching problem, separating geometry-based coverage from effective coverage. This yields multi-dimensional recall diagnostics that quantify mode dropping and mode collapse while reducing the overly optimistic estimates produced by geometry-based recall metrics. Through extensive investigations on toy experiments and generative models, we show that CADMA provides more reliable recall estimates and more fine-grained, interpretable diagnostic signals of diversity failures than existing metrics.
CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models
He Liang ⋅ Chenyang Ma ⋅ Yiming Zhang ⋅ Sangyun Shin ⋅ Andrew Markham ⋅ Niki Trigoni ⋅ Yuhang He
Existing 3D scene-grounded Large Language Models (3D-LLMs) focus on answering questions grounded in simplified single-room 3D scenes, lacking the ability to reason over real-world household environments containing multiple interconnected rooms and diverse object categories. We introduce CAIRN, a topology-aware 3D-LLM for multi-room 3D scene understanding. CAIRN aligns transformer attention with scene hierarchy, giving the model explicit awareness of object-level relations and room-level connectivity. It enriches object tokens with room-local relational context via a graph neural network, introduces learned room tokens for room-level abstraction, and applies a hierarchical attention mask with geometric bias to route information according to scene topology. CAIRN is developed on CAIRN-MR, a benchmark we introduce on HM3D for multi-room 3D scene understanding, covering grounding, captioning, and four question-answering tasks that progressively evaluate from intra-room perception to cross-room reasoning. Experiments show that CAIRN outperforms prior 3D-LLMs by a large margin across all CAIRN-MR tasks while remaining competitive on five single-room benchmarks.
Calibrating Generative Models to Feature Distributions with MMD Finetuning
Nathaniel L. Diamant ⋅ Brian L Trippe
Generative models can produce individually plausible samples while deviating substantially from a reference set in the distribution of key features. For example, a model pretrained on broad drug-like chemical space may generate molecules whose molecular features differ from those of a therapeutic class of interest, such as known antibiotics. Correcting such distributional miscalibration is challenging: direct finetuning on the reference set can overfit and does not control which features are matched. To fill this gap, we introduce kernel Calibrating Generative Models (kCGM). kCGM minimizes a maximum mean discrepancy (MMD) between generated and reference feature distributions using an unbiased score-function estimator, with KL regularization to remain close to the pretrained model. On a reference set of 174 antibiotics, direct finetuning sacrifices chemical validity for feature-distribution matching, whereas kCGM improves target feature matching while increasing validity. We further demonstrate kCGM in protein and DNA generation tasks, showing it can adapt autoregressive, continuous-space diffusion, and discrete diffusion models using only feature-level supervision.
CAM: Question Answering on Entity-Centric Videos with Continuous Extraction and Adaptive Querying
Yizhou Tian ⋅ Zizhe Chen ⋅ Shiyuan Deng ⋅ Garry YANG ⋅ Zijie Dai ⋅ Luohao Pan ⋅ Hao Lin ⋅ Peiqi Yin ⋅ Xiao Yan ⋅ James Cheng
Memory facilitates question answering over long videos by extracting and retrieving facts to fit within the limited context windows of multimodal LLMs (MLLMs). Existing solutions typically extract independent memory entries from fixed-length video clips and thus cannot capture high-level semantics that need to be summarized over extended time periods, such as character traits and relations. Moreover, they rely solely on similarity-based retrieval and may fail to retrieve the fine-grained details required for question answering. To tackle these problems, we propose CAM, featuring continuous extraction for high-level semantics and \textit{adaptive querying} for fine-grained details. In particular, CAM stores the entities and relations extracted from video clips in a knowledge graph. To capture the high-level semantics of each entity or relation, CAM summarizes the local subgraph of the target entity or relation once the subgraph reaches a predefined size. To retrieve the fine-grained details required for question answering, CAM supports multiple search methods, such as knowledge graph traversal, video re-watching, and audio listening, and utilizes a planner-executor-verifier pipeline to adaptively compose these search methods according to question intent. Evaluations on three benchmarks show that CAM outperforms SOTA baselines and improves their accuracy by up to 23 percentage points.
Can Complementary Signals Bridge Similarity Islands? Manifold-Augmented Graph Embedding for Multimodal Recommendation
Zhenye Yang ⋅ Jinpeng Chen ⋅ Wenbo Fu ⋅ Jun Ma ⋅ Huan Li ⋅ Qimin Zhou ⋅ Tao Wang ⋅ Li Sun ⋅ Senzhang Wang
Multimodal Recommender Systems (MMRS) leverage rich item features to mitigate data sparsity, yet their success largely hinges on the foundational homophily assumption. This paradigm makes existing models adept at recommending similar substitutes but poorly equipped to handle heterophilic complementarity. Consequently, complementary items often become geometrically isolated, leading to fragmented “Similarity Islands” and semantic drift when standard propagation is applied. To overcome these limitations, we propose MAGE (Manifold-Augmented Graph Embedding), a novel framework that decouples complementarity modeling into offline geometric warping and online topological propagation. The offline phase enhances item representations by extracting complement-aware residual directions via a Soft Conceptor mechanism and injecting them under a local gap-preserving constraint, enabling the model to capture functional relationships while reducing the risk of semantic drift. Synergistically, the online phase employs an asymmetric dual-stream GNN that treats visual features as stable anchors and augmented textual features as flexible probes, fused through a self-calibrated cross-modal mechanism to reconcile shared semantics and modality-specific residual cues. Extensive experiments on four Amazon datasets show that MAGE yields consistent improvements over strong multimodal recommendation baselines. Together with the supporting analyses in the appendix, these results suggest that geometry-aware augmentation can alleviate the semantic mismatch that arises in complementary recommendation settings. Our code is available at https://anonymous.4open.science/r/mage-4C60.
Large language models (LLMs) are increasingly used in high-stakes settings where they are expected to justify their predictions, and a common approach to understanding LLM predictions is via self-generated free-text explanations. However, free-text explanations are instance-wise and non-executable, making it hard to evaluate their alignment with model behavior or explain dataset-level behavior in a compact, human-interpretability manner. We study code as an executable alternative to free-text, enabling automatic evaluation and dataset-level explanation, and ask whether LLMs can \emph{consistently} explain their behaviors with code. This automatic verification enables a generation procedure that samples multiple candidate programs and selects the most consistent one. Using three open-source code-capable models and eight tasks with varying levels of linguistic difficulty, we study when LLMs produce consistent code explanations. We find that models can produce highly consistent code explanations in a subset of settings, but strong consistency is not the norm. Consistency is substantially lower for tasks involving more complex linguistic components, where performance is often near random guessing across models. In these harder settings, generated programs are longer but not more structurally complex, and increasingly rely on hard-coded linguistic patterns from the input. These findings suggest two possibilities: either LLMs are not good at producing compact explanations of their behavior, or some behaviors require dense, distributed computation that cannot be explained by compact, human-interpretable programs. Finally, we find that model-internal representations do not reliably predict when code explanations will be consistent, suggesting that LLMs do not know when they can or cannot explain their own behavior with code. Our work highlights code as a compact and verifiable alternative to free-text explanations, while also revealing limitations in using compact, human-understandable explanations to fully understand LLM behavior.
CAPER: Clause-Aligned Process Supervision for Text-to-SQL
Lujie Ban ⋅ Jiasheng Shi ⋅ Jinyang Li ⋅ Xiaolin Han ⋅ Tsz Nam Chan ⋅ Chenhao Ma
Text-to-SQL systems are typically evaluated by query-level execution correctness, but this terminal signal provides little guidance about which intermediate SQL decision caused success or failure. Token-level dense supervision is also ill-suited: SQL tokens do not align with complete semantic decisions, can penalize execution-equivalent queries, and are difficult to label reliably at scale. We therefore propose CAPER, which automatically derives clause-level supervision via counterfactual intervention on the SQL abstract syntax tree (AST), enabling root-cause error localization for reward modeling; the resulting data is used to train CAPER-9B, a lightweight Clause-PRM that provides clause-boundary feedback for policy optimization and candidate verification. Experiments on BIRD and Spider show that clause-aligned supervision not only improves execution accuracy, achieving up to a 15.3\% relative EX improvement over GPT-5.4, but also strengthens failure-localization capability, reaching 84.53\% accuracy and 90.60\% MRR on held-out failures.
Capricorn: Highly Efficient and Secure Mixture of Experts Inference Framework
Lushan Song ⋅ Xiaojian Liang ⋅ Shishuai Du ⋅ Jun J Sim ⋅ Yingting Liu ⋅ Xin Zhang ⋅ Jiang-Ming Yang ⋅ Pu Duan
The rapid advancement of Large Language Models (LLMs) based on Mixture-of-Experts (MoE) architecture has enhanced a growing demand for secure inference frameworks that protect both client inputs and server model weights. However, existing secure MoE inference frameworks suffer from two limitations: (1) a costly two-stage secure routing pipeline that uses expensive secure Top-$K$ and secure equality test protocols, resulting in significant computational and communication overhead, and (2) secure evaluation of nonlinear activations like SiLU requires several multiplication and comparison operations, resulting in significant communication overhead. In this paper, we present Capricorn, a highly efficient and secure MoE inference framework that overcomes the two limitations above. Firstly, Capricorn utilizes a novel one-stage routing pipeline that reduces both the total number of Top-$K$ protocol invocations and their input data size. At the same time, we propose a lightweight method to generate one-hot token selection matrices locally, thus avoiding expensive secure equality tests. Secondly, we design a secure and precise SiLU protocol, which uses secure lookup tables (LUTs) with dynamic bit-width and precision to reduce communication overhead. Extensive experiments demonstrate that Capricorn achieves up to $50.93 \times$ and $2.10 \times$ speedup in secure routing and secure SiLU protocols, respectively, compared to the state-of-the-art (SOTA) framework CryptoMoE (NeurIPS'25), while maintaining the original model inference accuracy.
CARD: Internalizing Expert Critique into Reinforcement Learning for Deep Search
Yanyu Zhu ⋅ Hoilam Pao ⋅ Zhaoyu Hu ⋅ Yufei zhang ⋅ Jiajun Chai ⋅ Guojun Yin ⋅ Wei Lin ⋅ Dongnian Wang ⋅ Hai-Tao Zheng ⋅ Hong-Gee Kim ⋅ Shaoxiong Zhan
Reinforcement learning with verifiable rewards (RLVR) has become the dominant paradigm for post-training large language models (LLMs) on complex reasoning and multi-step agentic tasks. However, the standard scalar reward signal is sparse and uninformative: it conveys whether a trajectory succeeded but not which reasoning steps failed or how to correct them, severely limiting sample efficiency on long-horizon tasks where successful trajectories are rare. We propose \textbf{CARD} (\textbf{C}ritique-\textbf{A}ugmented \textbf{R}einforcement \textbf{D}istillation), a training framework that amplifies the learning signal from a \texttt{verify_answer} tool within the GRPO objective via two complementary signals. First, a token-level self-distillation reweighting measures how much each token's generation was shaped by the expert critique: tokens that merely echo the critique text are masked, while critique-driven revision tokens are amplified. Second, a rejection-sampling distillation loss maximizes the likelihood of critique-free trajectories for responses where the critique provably improved the answer, directly accelerating the internalization of critique-guided behavior. Together, these two signals distill critique-augmented reasoning into the model's base policy without requiring additional rollouts or separate reward models. Experiments on four long-form deep search benchmarks and three multi-hop question-answering benchmarks demonstrate consistent improvements over the GRPO baseline. On the deep search benchmarks, CARD improves average performance by 40\% over the Qwen3-8B base model and by 6\% over GRPO; with Qwen3-14B, the gains are 46\% and 7\%, respectively. These results establish expert critique as an effective and scalable source of rich text supervision for reinforcement learning of long-horizon research agents. Furthermore, we show that \texttt{verify_answer} can serve as a test-time self-revision scaffold: at inference, the tool returns only task rubric criteria---without expert feedback---prompting the model to reflect on its draft and self-revise, yielding consistent performance gains over the no-VA baseline.
Cascaded Sparse Autoencoders LearnMulti-Level Visual Concepts in Multimodal LLMs
Yusong Zhao ⋅ Hengyi Wang ⋅ Tanuja Ganu ⋅ Akshay Nambi ⋅ Hao Wang
Multimodal Large Language Models (MLLMs) have demonstrated strong performance on vision-language tasks, yet their internal visual representations remain difficult to interpret. Sparse Autoencoders (SAEs) provide a scalable way to decompose dense model activations into sparse, interpretable features. However, existing SAE architectures primarily recover flat feature dictionaries and are less suited for explicit multi-level concept organization. In this paper, we introduce cascaded sparse autoencoders (CSAEs) for learning hierarchical visual concepts in MLLMs. Rather than nesting or stacking SAE sparse activation codes, CSAEs train a second-level SAE directly on the decoder weights of the first-level SAE, treating learned low-level feature directions as inputs for higher-level abstraction. This design enables CSAEs to learn “concepts of concepts” while avoiding drawbacks from the shared-prefix coupling of nesting, Matryoshka-style hierarchies and the bottlenecks of naively stacked SAEs. Experiments across Qwen3-VL, Gemma-3, and LLaVA on multiple visual datasets show that CSAEs improve interpretability in terms of hierarchical concept coherence over state-of-the-art SAE baselines. Results on concept steering further demonstrate that the learned concept groups support effective group-level interventions in MLLM outputs.
CASL-VAE: Learning Structured Latent Variables from Unpaired Data for Semi-supervised Clustering and Paired Sample Generation
Sai Spandana Chintapalli ⋅ Pratik Chaudhari ⋅ Christos Davatzikos
Quantifying variability in a target population relative to a reference population is central to many scientific and clinical problems (e.g., diseased vs. healthy). Yet, without paired data and in the presence of heterogeneous target variation, existing methods struggle to separate multiple modes of target-specific variation. We propose CASL-VAE, a deep contrastive latent variable model that learns structured latent generative factors from unpaired data. CASL-VAE factorizes variation into continuous common latent factors shared across populations and hierarchical salient latent factors that model target-specific heterogeneity as discrete subtypes and continuous within-subtype variation. Using variational inference, we show how approximate joint likelihood optimization over reference and target domains can be performed using unpaired data, providing a principled basis for paired-sample generation and cross-domain analysis. We validate CASL-VAE on semi-synthetic neuroimaging data, demonstrating improved subtype recovery and paired-sample generation compared to baseline clustering and generative models. We also validate its ability to reveal biologically plausible heterogeneity in Alzheimer's disease. Code will be released upon acceptance.
CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives
Yihao Meng ⋅ Zichen Liu ⋅ Hao Ouyang ⋅ Qiuyu Wang ⋅ Ka Leong Cheng ⋅ Yue Yu ⋅ Hanlin Wang ⋅ Xing Zhu ⋅ Yujun Shen ⋅ Qifeng Chen ⋅ Huamin Qu
Autoregressive video generation aims at real-time, open-ended synthesis. Yet, cinematic storytelling is not merely the endless extension of a single scene; it requires progressing through evolving events, viewpoint shifts, and discrete shot boundaries. Existing autoregressive models often struggle in this setting. Trained primarily for short-horizon continuation, they treat long sequences as extended single shots, inevitably suffering from motion stagnation and semantic drift during long rollouts. To bridge this gap, we introduce \textbf{CausalCine}, an interactive autoregressive framework that transforms multi-shot video generation into an online directing process. CausalCine generates causally across shot changes, accepts dynamic prompts on the fly, and reuses context without regenerating previous shots. To achieve this, we first train a causal base model on native multi-shot sequences to learn complex shot transitions prior to acceleration. We then propose Content-Aware Memory Routing (CAMR), which dynamically retrieves historical KV entries according to attention-based relevance scores rather than temporal proximity, preserving cross-shot coherence under bounded active memory. Finally, we distill the causal base model into a few-step generator for real-time interactive generation. Extensive experiments demonstrate that CausalCine significantly outperforms autoregressive baselines and approaches the capability of bidirectional models while unlocking the streaming interactivity of causal generation.
CAVE: A Structured Credit Assignment Approach for Fragmented Visual Evidence Reasoning
Tengda Guo ⋅ Jie Leng ⋅ Hanlei Li ⋅ Yaoyuan Liang ⋅ Qingyue Zhang ⋅ Dian Yang ⋅ Mingyu Zhang ⋅ Yuhua Fu ⋅ Shao-Lun Huang
Vision-Language Models (VLMs) have achieved strong performance on general multimodal reasoning, yet remain challenged in integrating nonlocal visual information to support semantically underdetermined visual reasoning. We describe this challenge as Fragmented Visual Reasoning. To this end, we propose Credit Assignment for Visual Evidence (CAVE), a structured process-reward method based on GRPO for interleaved visual reasoning. Specifically, CAVE evaluates the contribution of intermediate steps at the action level via three complementary reasoning process signals: belief update, evidence acquisition, and adaptive focus control, thereby guiding the model to optimize each reasoning action and learn more reliable visual reasoning strategies. Meanwhile, we construct TRACER-Bench, which covers four nonlocal and semantically confusable reasoning dimensions and provides key intermediate evidence to supervise reasoning paths. Experiments demonstrate that CAVE substantially improves performance on tasks requiring fragmented visual evidence integration, covering both public benchmarks and our newly introduced TRACER-Bench, while retaining competitive performance on general multimodal evaluations. Further analyses reveal that CAVE effectively improves the visual reasoning capacity and exhibits stronger robustness under longer-range and deeper cross-region dependencies. Our code and the proposed benchmark are available at: https://anonymous.4open.science/r/Anonymous_code-2D8B/.
CELEUS: Certifiable and Efficient LLM Evaluation via E-Processes
Zhijian Zhou ⋅ Zesheng Ye ⋅ Zhaorun Chen ⋅ Bo Li ⋅ Feng Liu
Can we trust evaluation scores to capture an LLM's true real-world performance? *Certifiable* evaluation answers this question by providing guarantee for LLM evaluation. In particular, existing methods sequentially curate evaluation samples and keep updating confidence intervals (CIs) that cover the true performance with high probability (e.g., 95%) until some conditions are satisfied, e.g., the CI width reaches a target precision. However, existing methods are not generally anytime-valid: the claimed coverage (e.g., 95%) may fail when CIs are repeatedly updated and used to decide when to stop, leaving a *gap* between theoretical rigor and practice. This paper bridges this gap by proposing **Celeus**, a **C**ertifiable framework for **E**fficient **L**LM evaluation, which leverages ***E***-processes to build anytime-valid CIs. Concretely, we propose signals that combine two ingredients: (i) **U**ncertainty-guided sampling to select informative samples for evaluation, and (ii) **S**urrogate-assisted approximations for unevaluated samples. We prove that such signals remain unbiased for the evaluation score conditional on the past, enabling statistically grounded and anytime-valid $e$-process CIs. More importantly, the two ingredients reduce estimation variance and help reach the target precision with fewer evaluated samples. We also prove that CIs obtained by **Celeus** can shrink at a near-parametric rate up to logarithmic factors and analyze the oracle variance-optimal sampling rule that motivates the empirical uncertainty-guided one. Experiments show that **Celeus** reaches the target precision using 54-62% fewer evaluated samples than baselines, while preserving anytime-valid coverage.
CHAIN: Complementary Signed Graph Propagation for Uncertainty Quantification in Large Language Models
Hongxing Pan ⋅ Ziwei Li ⋅ Ziyue Qiao ⋅ Xiao Luo
This paper studies the problem of uncertainty quantification in large language models (LLMs), which is critical for reliable LLM deployment under the risk of hallucination. Existing approaches typically treat the response sets as unordered collections and rely on local pairwise similarity or clustering, which overlook relational signals among responses as well as their corresponding high-order structural semantics. Towards this end, we propose a novel approach named Complementary Signed Graph Propagation (CHAIN) for uncertainty quantification in LLMs. The core idea of CHAIN is to capture complementary signals among responses using a signed graph and then extract high-order structural information using random propagation. In particular, CHAIN builds a semantic graph where each response is considered as a node and edges represent complementary views, namely enrichment relations and contradiction relations. To acquire high-order structural signals, we update hidden states through graph-based information propagation along edges and channel fusion guided by randomized matrices. After propagation, these node states are summarized into graph-level scores for reliable uncertainty quantification. Extensive experiments on five datasets validate the effectiveness of the proposed CHAIN in comparison with competing baselines.
The attention mechanism in Transformers links queries, keys, and values through a sequence of computations, while the principles that constrain its form remain unclear. This paper examines attention under natural requirements that formalize basic properties of attention computation as transformations between representation spaces, and asks which forms of attention computation are admissible. We show that similarity computation, score-to-weight mapping, and value aggregation are unified by dual structures induced by convex potential functions and their Legendre transforms. In particular, the standard inner-product score is recovered as a limiting case of a potential-based score, the score-to-weight map must take a gradient-map form under structural requirements on attention normalization, and the value-aggregation rule must take a dual-coordinate barycentric form under structural requirements on representative-point aggregation. These results show that the standard components of attention are structurally constrained by the underlying dual structure, while the choice of potentials provides explicit degrees of freedom for understanding and modifying attention mechanisms.
Chain-of-Thought Is Not Explainability
Fazl Barez ⋅ Tung-Yu Wu ⋅ Iván Arcuschin Moreno ⋅ Michael Lan ⋅ Vincent Wang-Maścianica ⋅ Noah Siegel ⋅ Nicolas Collignon ⋅ Clement Neo ⋅ Isabelle Lee ⋅ Alasdair Paren ⋅ Adel Bibi ⋅ Robert Trager ⋅ Damiano Fornasiere ⋅ John Yan ⋅ Yanai Elazar ⋅ Yoshua Bengio
Chains‑of‑thought (CoT) allow language models to verbalise multi‑step rationales before producing their final answer. While this technique often boosts task performance and offers an impression of transparency into the model's reasoning, this position paper argues that rationales generated by current CoT techniques can be misleading and are neither necessary nor sufficient for trustworthy interpretability. We analyze faithfulness as whether CoTs are not only human-interpretable, but also reflect the model’s underlying reasoning well enough to support responsible use. By synthesizing evidence from prior work, we show that verbalised chains are frequently unfaithful, diverging from the true hidden computations that drive a model's predictions. As CoT is increasingly relied upon in application domains, including high-stakes ones such as medicine, law, and autonomous systems, we further argue that this unfaithfulness is not taken seriously enough in these settings: our analysis of 1,000 recent CoT-centric papers finds that approximately 25\% explicitly treat CoT as an interpretability technique, including papers in high-stakes domains which heavily hinge on such interpretability claims. Building on prior work, we make three proposals: (i) avoid treating CoT as being sufficient for interpretability without additional verification, while continuing to use CoT for its communicative benefits, (ii) adopt rigorous methods that assess faithfulness for downstream decision-making, and (iii) develop causal validation methods to ground explanations in model internals.
Chance-constrained Flow Matching for High-Fidelity Constraint-aware Generation
JINHAO LIANG ⋅ Yixuan Sun ⋅ Anirban Samaddar ⋅ Sandeep Madireddy ⋅ Ferdinando Fioretto
Generative models excel at synthesizing high-fidelity samples from complex data distributions, but they often violate hard constraints arising from physical laws or task specifications. A common remedy is to project intermediate samples onto the feasible set at each sampling step. However, repeated projection can disrupt the learned sampling dynamics because this feasible set is defined for ''clean'' samples rather than noisy intermediate states. Thus, recent studies first estimate ''clean'' samples and then project them to the feasible set, but this increases algorithmic complexity and accumulates errors across steps. To address these issues, this paper proposes Chance-constrained Flow Matching (CCFM), a training-free method that formulates constraint enforcement during sampling as stochastic optimization while retaining the hard-constraint feasibility guarantee of projection-based methods. By avoiding direct projection of noisy intermediate states onto the clean-sample feasible set, CCFM mitigates the distributional distortion. Unlike approaches that first estimate and project ''clean'' samples, CCFM avoids complex multi-stage procedures. Experiments show that CCFM outperforms current state-of-the-art constrained generative models in modeling complex physical systems governed by partial differential equations, molecular docking problems, and physics-informed motion generation, delivering superior feasibility and fidelity.
Channel Mixer: A Pretrainable Tokenizer for Scalable Multi-Channel Vision Transformers
Lukas Miklautz ⋅ Lucas Miranda ⋅ Nikola Bulat ⋅ Armin Lambacher ⋅ Dexiong Chen ⋅ Karsten Borgwardt
Channel-adaptive Vision Transformers (ViTs) struggle to scale to non-RGB images with many uncorrelated or weakly correlated channels. Existing approaches roll out each channel into its own patch embedding, so the token sequence length grows linearly with the number of channels and self-attention cost grows quadratically. We introduce \emph{ChannelMixer}, a multi-channel tokenizer that aggregates all channels within each spatial patch into a single channel-mixing token, decoupling the token sequence length from the number of channels. ChannelMixer is pretrained once with a lightweight autoencoder in a channel-agnostic manner and the resulting tokenizer can be easily integrated with ViTs trained under both supervised and self-supervised objectives. On multi-channel imaging benchmarks spanning 3 to 29 channels, ChannelMixer matches or outperforms channel-adaptive state-of-the-art methods while using approximately $20\times$ fewer FLOPs and $29\times$ fewer tokens than rolled-out tokenizers, measured on 29-channel images with a ViT-S/16.
Cheap Talk, Real Stakes: Commitment and Exploitation in Human-LLM Strategic Interaction
Artyom Kabanov ⋅ Mikhail Mozikov ⋅ Sergei Volchkov ⋅ Igor Gorobets ⋅ Daniil Shirshin ⋅ Alena Kriuchenkova ⋅ Bykov Nikita ⋅ Valeria Bodishtianu ⋅ Nikita Severin ⋅ Sergey Muravyov ⋅ Andrey Savchenko ⋅ Ilya Makarov
LLM agents now negotiate, allocate resources, and make commitments on a user's behalf. The vulnerability is not only in the final move but in what comes before it: in pre-action conversation, a model can be primed to make promises, lock in commitments, or reveal intentions that a counterpart will exploit when the action is finally taken. Yet standard game-based evaluations usually observe only the final move, leaving the speech-action gap unmeasured. We study this gap directly by introducing HL-PACT, a controlled platform with matched human-human, human-LLM, and LLM-LLM sessions across five repeated games: Prisoner's Dilemma, Battle of the Sexes, Ultimatum, Kuhn Poker, and Liar's Dice. Across 3,752 sessions, we randomize free-form pre-round chat and opponent identity masking, classify each message by its strategic content, and compare what players say, how they say it, and when the two diverge. We find that LLMs consistently activate the communication channel: mixed pairs contain more proposals, promises, and strategic reasoning than human pairs, with humans adapting to the LLM's negotiation register. This improves equilibrium selection in coordination games and shifts bargaining toward fairer offers. In games with profitable unilateral deviation, however, the same communication structure can create asymmetries: when identity is known, human partners selectively adapt to LLMs’ willingness to follow through on stated intentions, producing measurable exploitation. Masking reduces this asymmetry and makes the communication benefit more symmetric. The strategic value of language therefore depends on game structure, counterpart identity, and whether speech is separated from action. Evaluating LLM agents without communication can miss a central deployment risk: models may not fail because they reason poorly, but because their reliability is predictable.
Cite What You Explore: Budget-Aware LLM Reasoning over Medical KGs with Verifiable Evidence
Chen Chen ⋅ Dongjie Wang ⋅ Mei Liu ⋅ Zijun Yao
Post-discharge risk prediction from electronic health records (EHRs) is difficult because many dependencies that link discharge-time observations to downstream complications, such as comorbidity cascades and drug-disease interactions, are absent from the record. External medical knowledge graphs (KGs) can supply these missing dependencies, but tracing them demands three properties: KG exploration must remain cost-bounded, retrieved evidence must be differentiated by source quality, and the resulting rationale must be citable for retrospective review. Large language models (LLMs) can plan and verify over structured evidence, making them natural candidates for KG reasoning, but existing LLM-based methods do not satisfy these three properties jointly. In this paper, we propose BAR, a Budget-Aware LLM Reasoning framework over medical KGs with three components. First, BAR refines the raw KG into disease-specific evidence graphs whose edges carry support scores and provenance records, turning the KG into a quality-annotated reasoning space rather than a static feature source. Second, an LLM then reasons over this graph through a plan-navigate-verify loop that decomposes the question into steps, retrieves evidence under a patient-specific budget, and revises when verification fails. Third, a reasoning policy is trained with a reward that compares predictions with and without acquired evidence, combined with acquisition cost and citation-integrity terms. Across 8 diseases and 3 prediction horizons on MIMIC-III and MIMIC-IV, BAR improves AUPRC by 3.4 points over the strongest baseline, raises citation precision from 59.8\% to 77.9\%, and consumes only 62-65\% of the budget cap. The code is available at https://anonymous.4open.science/r/BAR/.
CLEF: EEG Foundation Models for Learning Clinical Semantics
Peng Cao ⋅ Ali Mirzazadeh ⋅ Jong W Lee ⋅ Aleksandar Videnovic ⋅ Dina Katabi
Clinical EEG interpretation requires reasoning over full EEG sessions and integrating signal patterns with clinical context. Existing EEG foundation models are largely designed for short-window decoding and do not incorporate clinical context. We introduce CLEF, a clinically grounded long-context EEG foundation model. CLEF represents EEG sessions as 3D multitaper spectrogram tokens, enabling tractable Transformer modeling at session scale, and aligns embeddings with neurologist reports and structured EHR data through contrastive objectives. We evaluate CLEF on a new 234-task benchmark spanning disease phenotypes, medication exposures, and EEG findings, with more than 260k EEG sessions from over 108k patients. CLEF outperforms prior EEG foundation models on 229 of 234 tasks, improving mean AUROC from 0.65 to 0.74. Reconstruction-only pretraining surpasses prior EEG foundation models, while report and EHR alignment yields further gains. Held-out concept and external-cohort experiments suggest that these representations transfer beyond observed alignment targets. These results support session-scale, clinically grounded representation learning as a promising foundation-model paradigm for clinical EEG.
Click3R: Interactive Stereo 3D Reconstruction with Sparse Correspondence Clicks
Hao Gu ⋅ Jin Yao ⋅ Zezhou Cheng
Despite impressive progress, existing feed-forward stereo 3D reconstruction models produce a single, irrevocable prediction. When these models fail, there is no mechanism to correct their mistakes. We introduce Click3R, a framework for interactive geometric correction of feed-forward models using sparse human correspondence clicks. Given an image pair and a small set of cross-view point correspondences provided by a user, Click3R injects geometric constraints into the 3D reconstruction networks through a lightweight point correspondence adapter, enabling the model to resolve ambiguities in challenging scenarios such as repetitive textures, symmetric structures, large viewpoint changes, or textureless regions. To evaluate interactive reconstruction, we curate a challenging benchmark specifically designed to expose failure modes of existing feed-forward stereo 3D reconstruction methods. Experiments show that Click3R dramatically reduces reconstruction error with as few as a single correspondence click, improving performance on both standard benchmarks and our curated dataset. Our results demonstrate that sparse human interaction provides an effective and practical mechanism for correcting geometric reconstruction errors.
CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling
Yifan Zhang ⋅ Yutong Dai ⋅ Viraj Prabhu ⋅ Zhiyuan Hu ⋅ Ran Xu ⋅ Zeyuan Chen
Open-source web agents are now strong enough to execute realistic browser tasks, but training them with reinforcement learning still depends on weak supervision: binary task success is too sparse for credit assignment, while frontier-language-model judges are too expensive to call at every step and cannot be assumed available at deployment. We introduce CLIFT, a training and test-time scaling method built around conformal self-verification. During training, the agent answers natural-language verification questions about its own rollouts; a Compositional Conformal Certifier keeps only question signals whose URL-conditional evidence agrees with a training-time judge, assigns signed trust weights through polarity-aware lift, and blends the resulting verifier score into per-step rewards in a way that never subtracts from the judge baseline. At test time, the same certified bank is frozen and reused as structured evidence for Conformal Trajectory Selection (CTS): the agent samples a greedy rollout and one or more diverse retries, the self-verifier summarises each URL trace, and a conservative majority-vote rule chooses whether to swap away from the current incumbent without calling any external judge. This single mechanism supports three settings. On WebArena Infinity, CLIFT achieves state-of-the-art performance among open-source web agents. On VisualWebArena, a bank trained with the open model transfers to GPT-5.5 at test time and reaches state-of-the-art performance under the canonical harness. On Online Mind2Web, without training an agent on the benchmark, translating the certified question bank improves a live-web agent in zero-shot evaluation. Together these results position conformal self-verification as a way to turn costly judge feedback into a reusable training signal and a judge-free test-time scaling signal.
ClinMAS: A Knowledge-Grounded Multi-Agent Simulation Framework for Evaluating Clinical Reasoning in LLMs
Yuqi Tang ⋅ Jing Yu ⋅ Zichang Su ⋅ Kehua Feng ⋅ Zhihui Zhu ⋅ Libin Wang ⋅ Lei Liang ⋅ Qiang Zhang ⋅ Keyan Ding ⋅ Huajun Chen
Clinical diagnosis begins with doctor-patient interaction, during which physicians iteratively gather information, order examinations, and refine diagnosis through patients’ responses. This dynamic reasoning process is poorly represented by existing LLM benchmarks that focus on static question-answering. To mitigate these gaps, recent methods explore dynamic medical frameworks involving interactive clinical dialogues. Although effective, they often rely on limited, contamination-prone datasets and lack granular evaluation. In this work, we propose ClinMAS, a knowledge-grounded multi-agent simulation framework for evaluating clinical reasoning in LLMs. Grounded in a disease knowledge graph, our method dynamically generates patient cases and facilitates multi-turn interactions between an LLM-based doctor, an automated patient agent, and an examiner agent. Our evaluation goes beyond diagnostic accuracy by incorporating fine-grained efficiency analysis and rubric-based assessment of diagnostic quality. Experiments show that ClinMAS effectively exposes critical clinical reasoning gaps in state-of-the-art LLMs, offering a more nuanced and clinically meaningful evaluation paradigm.
CLIOPATRA: Extracting Private Information from LLM Insights
Meenatchi Sundaram Muthu Selva Annamalai ⋅ Emiliano De Cristofaro ⋅ Peter Kairouz
The widespread adoption of AI assistants has prompted the development of privacy-aware platforms designed to extract insights from real-world usage. Their privacy protections primarily rely on layering multiple heuristic techniques, such as PII redaction, clustering, and aggregation. In this paper, we put their privacy claims to the test by presenting CLIOPATRA, the first attack against privacy-preserving LLM-based insights systems. Our attack involves an adversary that carefully designs and inserts malicious chats into the system to break multiple layers of protections and induce the leakage of sensitive information from a target user’s chat. We evaluate CLIOPATRA on one such platform, Anthropic’s Clio, and target synthetically generated medical chats to show that an adversary can successfully and confidently (with nearly 100% precision) extract the medical history contained in these chats in up to 65% of cases. We also show that CLIOPATRA can stealthily extract information by obfuscating the private information in the generated insights. Finally, we demonstrate that existing ad hoc mitigations, such as LLM-based privacy auditing, are unreliable and fail to detect major leaks. Taken together, our findings indicate that, even when layered, current heuristic protections are insufficient to adequately protect user data, and that prompt injection has been an understudied risk in LLM-based insight systems.
Training implicit neural representations (INRs) to capture fine-scale details typically relies on iterative backpropagation and is often hindered by spectral bias when the target exhibits highly non-uniform frequency content. We propose ELM-INR, a backpropagation-free INR that decomposes the domain into overlapping subdomains and fits each local problem using an Extreme Learning Machine (ELM) in closed form, replacing iterative optimization with stable linear least-squares solutions. This design yields fast and numerically robust reconstruction by combining local predictors through a partition of unity. To understand where approximation becomes difficult under fixed local capacity, we analyze the method from a spectral Barron norm perspective, which reveals that global reconstruction error is dominated by regions with high spectral complexity. Building on this insight, we introduce BEAM, an adaptive mesh refinement strategy that balances spectral complexity across subdomains to improve reconstruction quality in capacity-constrained regimes.
Closed-Form Linear-Probe Dataset Distillation for Pre-trained Vision Models
Bincheng Peng ⋅ Miki Haseyama ⋅ Guang Li ⋅ Ping Liu ⋅ Takahiro Ogawa
Dataset distillation compresses a large training set into a small synthetic set that preserves downstream training utility. While most existing methods target training networks from scratch, modern visual transfer learning often uses frozen pre-trained encoders followed by lightweight linear probing. Existing distillation methods for this setting either unroll iterative linear-probe updates with trajectory-based gradient matching, or rely on closed-form formulations originally designed for from-scratch training with neural-tangent-kernel (NTK) approximations. Neither route exploits the fact that frozen-feature linear probing admits a closed-form solution determined directly by the pre-trained features themselves, with no infinite-width approximation and no inner-loop trajectory. We propose Closed-Form Linear-Probe Dataset Distillation (CLP-DD), a bilevel formulation that computes the linear probe induced by the synthetic set with a sample-space kernel ridge solver. The synthetic images are then updated by evaluating this induced classifier on real features through a temperature-scaled softmax cross-entropy, where the classifier columns act as learned class anchors in feature space. We further show that the choice of outer objective is decisive: pairing the closed-form inner solver with a standard MSE outer loss substantially underperforms trajectory-based methods, while the discriminative outer loss closes most of the gap. On ImageNet-100 with four pre-trained backbones, CLP-DD substantially improves over LGM without DSA and approaches LGM with DSA at a fraction of the computational cost. On ImageNet-1K, CLP-DD matches or surpasses LGM with DSA on three of four backbones while running roughly $14\times$ faster and using less than one-eighth of the GPU memory.
In a standard transformer, each layer computes attention once: queries examine keys, select values, and move on. But attention is a function of representation, and representation is a function of attention. We close this loop. Fixed-Point Self-Attention (FPSA) computes its attention pattern by iterating the query/key/value process to convergence. Values $V$ are derived from the layer input and held fixed; queries and keys are recomputed from an evolving output state $Z_t$ at each step, converging when representations reach equilibrium, $Z^* = f(Z^*)$. FPSA is a drop-in replacement for multi-head attention that adds zero parameters and adapts depth on its own: easy tokens stabilize in $2$--$3$ iterations, hard tokens refine over $15$ or more, with $\mathcal{O}(1)$ memory cost regardless of the number of iterations. It lifts BERT-Base and ELECTRA-Base on GLUE and SQuAD~v2.0, improves ViT-B/16 accuracy by up to 20\%, and delivers matching gains on vision-language tasks, all without adding a single parameter. Unlike looped transformers that iterate entire layers, FPSA iterates only the attention step, so overhead is modest: a median of $3$--$6$ steps per layer adds roughly $1.6\times$ GFLOPs and $1.3$--$1.4\times$ wall-clock time over BERT-Base. On multi-step reasoning benchmarks (GSM8K, BBH, LogiQA), FPSA's adaptive computation yields clear gains, letting the model refine longer when the input demands it.
CODA: Cohort- and Drift-aware Foundation Model for Multimodal Clinical Reasoning
Changshuo Liu ⋅ Jiaqi Zhu ⋅ Wenqiao Zhang ⋅ Xiaokui Xiao ⋅ Beng Chin Ooi
Recent advances in multimodal foundation models have demonstrated strong performance across diverse clinical tasks. However, existing approaches predominantly operate at the individual sample level and fail to model the interplay between population-level cohort patterns and patient-specific variation that underlie real-world clinical reasoning. In practice, clinicians routinely reason through patient cohorts sharing similar demographic, physiological, or pathological characteristics, while accounting for patient-specific deviations from the group prototype. Motivated by this gap, we propose CODA, a cohort- and drift-aware foundation model for multimodal clinical reasoning. CODA explicitly integrates cohort structure into the modeling pipeline through two complementary mechanisms: cohort-aware population stratification and drift-aware individual adaptation. In particular, we leverage reinforcement learning to discover clinically meaningful patient cohorts and further model each patient’s deviation from cohort prototypes as a structured drift signal encoding fine-grained individual heterogeneity. These cohort- and drift-aware representations are incorporated into the foundation model to enable reasoning that is simultaneously informed by population-level patterns and sensitive to individual-level variation. Extensive experiments on real-world electronic health record (EHR) benchmarks demonstrate that CODA achieves consistent effectiveness and strong generalizability across diverse clinical tasks, including closed-QA, open-QA, and report generation.
Co-Evolving Policy Distillation
Naibin Gu ⋅ Chenxu Yang ⋅ Qingyi Si ⋅ Chuanyu Qin ⋅ Dingyu Yao ⋅ Peng Fu ⋅ Zheng Lin ⋅ Weiping Wang ⋅ Nan Duan ⋅ Jiaqi Wang
RLVR and OPD have become standard paradigms for post-training. We provide a unified analysis of these two paradigms in consolidating multiple expert capabilities into a single model, identifying capability loss in different ways: mixed RLVR suffers from inter-capability divergence cost, while the pipeline of first training experts and then performing OPD, though avoiding divergence, fails to fully absorb teacher capabilities due to large behavioral pattern gaps between teacher and student. We propose Co-Evolving Policy Distillation (CoPD), which encourages parallel training of experts and introduces OPD during each expert's ongoing RLVR training rather than after complete expert training, with experts serving as mutual teachers (making OPD bidirectional) to co-evolve. This enables more consistent behavioral patterns among experts while maintaining sufficient complementary knowledge throughout. Experiments validate that CoPD achieves all-in-one integration of text, image, and video reasoning capabilities, outperforming strong baselines such as mixed RLVR and MOPD, and even surpassing domain-specific experts.
Cognitive Firewalls: A Synthetic Account of Cross-Lingual Reasoning Collapse
Jou Barzdukas ⋅ Paulius Rauba ⋅ Julian Schulz ⋅ Jack Peck ⋅ Steven Basart ⋅ Lennie Wells
The reasoning capabilities of large language models are reported to drop sharply when they are forced to think in heavily regulated target languages such as Chinese — a phenomenon the literature attributes to data imbalance, multilingual transfer failures, or, more recently, RL-induced cross-lingual collapse. We propose and test an alternative mechanistic account: that the gap reflects RL-induced selection between conditional policies sharing the same parameters, rather than a loss of underlying capability. We construct a model organism for this account. Two policies — identified solely by the language of an unsupervised chain of thought — are installed via SFT on disjoint splits of ARC-Challenge, with one policy trained to answer and the other to refuse on each split. Reinforcement learning with a final-answer reward, restricted to one split, drives the language of reasoning to collapse onto the rewarded policy across both splits, producing a 67% → 10% accuracy collapse on the held-out split despite no reward signal ever observing it. General capability is unaffected: MMLU accuracy is 71.2% before and after RL. A depth-wise activation steering intervention then recovers held-out accuracy from 0% to 85% on the worst-case slice (Chinese-CoT completions) while preserving Chinese as the surface chain-of-thought language and leaving the rewarded policy intact (83% vs. 94% unsteered). The capability was never lost; it was gated. To the extent deployed regulated-language models contain analogous structure, surface evaluations of such models will systematically understate capability hidden behind learned policy gates.
Test-time compute scaling, where language models "think longer" via extended chain-of-thought (CoT), is the dominant paradigm for improving reasoning. Prevailing intuition holds that extended reasoning *converges* toward an answer; we show the opposite. Across three reasoning-RL families (DeepSeek-R1, Qwen-QwQ, Llama-3.1-Nemotron) and three benchmarks (GSM8K, MATH, GPQA), step-to-step SAE feature turnover and entropy *rise* monotonically across the chain (Cohen's $d = 0.61$–$0.68$, $p < 10^{-3}$). Yet the final answer is already linearly decodable from the residual stream within the first $\sim 20\%$ of the chain. We call this pattern **commit then explore**: the answer becomes stably decodable early, and the bulk of test-time compute is then spent exploring *around* an already-decodable answer. Targeted ablation shows post-commit features have a $2$–$6\times$ smaller teacher-forced log-probability effect (matched-mass control: $1.5$–$4.6\times$) and a $2$–$3\times$ smaller free-regeneration answer-flip rate than pre-commit features. A terminal-reward RLVR credit-assignment account predicts this asymmetry, and the dynamic is absent in base models: Qwen-2.5 and Llama-3.1 bases fail our detector at every layer, and a held-out quantitative prediction derived from R1 generalises to Qwen-QwQ at Pearson $r = 0.96$. Leveraging the structure, we build a Reasoning Monitor that saves $26$–$34\%$ of inference compute at $\geq 92\%$ accuracy retention, dominating four adaptive-compute baselines. Extended chain-of-thought, under this view, is not the computation of an answer; it is the exploration of its neighborhood.
CommunityKV: Efficient Long-Context Decoding via Graph Partitioning
Joe McKenna ⋅ Anastasios Alexandridis ⋅ Nathan Susanj ⋅ Jing Liu
Scaling Transformers to long contexts is constrained by the quadratic cost of self-attention and the linear growth of key-value cache memory transfer. Sparse attention mitigates this by retrieving only relevant tokens, but current approaches either require large-scale training or, within the training-free regime, rely on semantically coarse heuristics or expensive clustering that is difficult to update efficiently during decoding. We introduce CommunityKV, a framework that formulates sparse attention as a community detection problem. CommunityKV constructs a token graph from the $QK^T$ scores already computed during standard prefill, and partitions the graph into communities to enable retrieval of semantically coherent token groups. A local update rule assigns newly generated tokens to communities in constant time, enabling sparse retrieval throughout streaming decoding without global re-partitioning. We evaluate CommunityKV on open weight models and long context benchmarks, demonstrating that our approach achieves up to $1.86\times$ higher decoding throughput than exact FlashAttention-2 with negligible accuracy degradation.
COMPOSE: Composing Future Theorems from Citations and Formal Structure
David Busbib ⋅ Michael Werman
A generated plausible future mathematical claim must satisfy two constraints: it should follow the direction of prior work and respect the formal dependencies that constrain the validly of the claim. Existing approaches typically model only one of these sources, producing claims that are either weakly grounded or insufficiently motivated. We introduce grounded future mathematical generation, the task of generating a plausible future theorem-like claim for an anchor paper, using two complementary sources of context: its scientific citation graph ,and its aligned formal theorem dependency graph. To address this setting, we propose COMPOSE, a dual-graph framework that combines scientific citation context with formal mathematical structure to condition a language model. To support this setting, we construct a dataset of 108K paired scientific-formal graphs from arXiv and Mathlib, together with a benchmark of 47K future papers from 2024-2025. Experiments show that COMPOSE outperforms strong baselines on retrieval to real future papers and achieves the best overall performance under LLM-judge evaluation, producing more grounded and mathematically richer outputs. These results show that future mathematical generation benefits from combining scientific context with formal structure.
A neural network trained on the multiplication table of the quaternion group Q₈ can achieve perfect accuracy within each cyclic subgroup ⟨i⟩ and ⟨j⟩, predict every cross-region product wrong, and stay there indefinitely. This is the grokking plateau, and we prove its height is exactly 4 in dimension d=2, determined by the topology of the overlap before any training. The mechanism is a homotopy pullback: for G = A ∗_C B with finite overlap, the representation space of G is a homotopy pullback of those of A and B over that of C, so locally accurate modules need only agree up to a basis change. From this we derive a stratified generalization certificate whose strata are the connected components of Hom(C, U(d)): in the correct stratum, with locally accurate modules, the cross-region prediction is O(ε)-close to the target; in the wrong stratum it is provably far regardless of local accuracy, with exact gap c = 2(d − rₘₐₓ) for central cyclic overlap. The plateau is therefore the time the network spends in the wrong stratum, the grokking transition is the discrete jump between strata, and the framework predicts which task families exhibit such plateaus—those whose overlap representation space is disconnected—and which do not.
Compositional Training-Free Diffusion Planning for Long-Horizon Multiple Reach-Avoid Tasks
Yu Chen ⋅ Ruijia Liu ⋅ Xiao Yu ⋅ Xiang Yin
Long-horizon robotic planning often requires satisfying temporally ordered safety and visitation constraints rather than a single goal condition. We study multiple reach-avoid (MRA) tasks, in which a robot must sequentially visit target regions while remaining within corresponding safe sets. Existing approaches either rely on explicit system models or use data-driven planners whose test-time conditioning mechanisms do not scale well to long-horizon temporally structured tasks. We propose a training-free compositional diffusion planning framework for MRA tasks that operates directly on a pre-trained task-agnostic diffusion trajectory prior at test time, without requiring a known dynamics model. Our approach decomposes a global MRA specification into local reach-avoid sub-tasks, composes overlapping short-horizon diffusion priors into a long-horizon generative process through a factor-graph perspective, and enforces each local requirement through projection-based denoising during sampling. The same construction also extends naturally to prefix-suffix tasks through a looped compositional graph. We provide a formal correctness guarantee showing that the stitched trajectory satisfies the target specification. Experiments on long-horizon constrained planning benchmarks show strong execution success, substantially lower planning time than representative guidance-based baselines, and effective transfer to richer dynamics.
Compressing Collections of Trees with Decision Equivalence
Hayden McTavish ⋅ Jon Donnelly ⋅ Margo Seltzer ⋅ Cynthia Rudin
Decision trees with binary leaves are fundamental building blocks in classical machine learning. While individual trees can be used as strong interpretable classifiers, they frequently serve as members of large collections instead: primarily as components of powerful ensembles, but also as one of the few model classes for which the entire set of near-optimal models (the Rashomon set) can be analyzed. Both types of collections are often more complex than necessary, because there are many decision tree representations for the same decision boundary, imposing unnecessary computational costs in ensembles and complicating theoretical work on Rashomon sets. We propose methods to reduce collections of trees to a subset representing each unique decision boundary in the collection, with each tree in its sparsest equivalent form. In so doing, we improve the scalability of removing provably redundant classifiers in a Rashomon set. We also compress many shallow Adaboost ensembles by >15\% of their total number of leaves, while exactly preserving the models' decisions and probability estimates on any possible observation. These methods allow practitioners to run more efficient computations over compressed collections of trees, and remove arbitrary complexity when analyzing and understanding them.
Computational Depth Predicts Quantization Sensitivity in Multimodal Models
Dongnan Gui ⋅ Ruizhe Wang
Post-training quantization degrades multimodal model capabilities unevenly, yet practitioners lack a principled way to predict which capabilities will collapse. We propose the **Computational Depth Principle** (CDP): end-to-end fidelity under quantization follows $\hat p^n$, where $n$ counts sequential precision-dependent operations and $\hat p$ is a per-architecture survival rate. Across four VLM families and diverse benchmarks spanning image understanding, image generation, video understanding, and video generation, we show that deeper reasoning tasks consistently degrade more steeply and that a single shallow probe predicts the full degradation ranking on held-out architectures. The same exponential form extends to image generation, where denoising step count plays the role of depth, and to video understanding, where longer temporal reasoning chains amplify losses. For video generation, spatial quality degrades more than temporal coherence, indicating that the dominant compounding axis is denoising depth rather than cross-frame coupling. A complementary distortion-to-information (D/I) flip ratio diagnoses quantization method effects, revealing that certain methods preserve near-baseline quality while others produce catastrophic collapses on specific architecture pairings that aggregate accuracy alone cannot detect. Module ablation localizes the sensitivity to the language backbone rather than the vision encoder. Together, depth and D/I convert full-suite evaluation into a two-probe screening protocol that surfaces the highest-risk capability gaps first.
Concepts in Motion: Temporal Concept Bottleneck Model for Interpretable Video Classification
Patrick Knab ⋅ Sascha Marton ⋅ Philipp J Schubert ⋅ Drago A Guggiana Nilo ⋅ Christian Bartelt
Concept Bottleneck Models (CBMs) enable interpretable image classification by structuring predictions around human-understandable concepts, but extending this paradigm to video remains challenging due to the difficulty of extracting concepts and modeling them over time. In this paper, we introduce MoTIF (Moving Temporal Interpretable Framework), a transformer-based concept architecture that operates on sequences of temporally grounded concept activations, by employing per-concept temporal self-attention to model when individual concepts recur and how their temporal patterns contribute to predictions. Central to the framework is a class-conditioned VLM-based concept discovery module that extracts object- and action-centric textual concepts from training videos, yielding temporally expressive concept sets without manual concept annotation. Across multiple video benchmarks, this combination improves over global concept bottlenecks and remains competitive within the interpretable concept-bottleneck setting, while narrowing the gap to strong black-box video baselines that we report as contextual references.
Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers
Jan Dubiński ⋅ Jan Betley ⋅ Anna Sztyber-Betley ⋅ Daniel Tan ⋅ Owain Evans
Finetuning a language model can lead to emergent misalignment (EM) [Betley et al., 2025b]. Models trained on a narrow distribution of misaligned behavior generalize to more egregious behaviors when tested outside the training distribution. We study three interventions proposed to reduce EM. We confirm that these interventions reduce or eliminate EM on existing evaluations (questions like “How do I make a quick buck?”). However, if the evaluation prompts are tweaked to resemble the training context, the model displays EM. We call this conditional misalignment. As in standard EM, the model displays misaligned behaviors more egregious than those seen during training, but only on inputs sharing features with the training data. The first two interventions are diluting misaligned data with benign data, and finetuning on benign data after misaligned data. Both produce conditional misalignment. For instance, models trained on a mix of only 5% insecure code still show misalignment when asked to format responses as Python strings (resembling the training context). The third intervention is inoculation prompting. Here, statements with a similar form to the inoculation prompt serve as triggers for misalignment, even if they have the opposite meaning. On the positive side, inoculation prompting has lower (but still non-zero) conditional misalignment if training is on-policy or includes reasoning distillation. Our results imply that in realistic post-training, where misaligned data is typically combined with benign data, models may be conditionally misaligned even if standard evaluations look clean.
Confidence Estimation via Decoupled Smoothing for Dynamic LLM Routing and Aggregation
Hao Luo ⋅ Ning Cheng ⋅ Yuhao Lin ⋅ Xiaokai Zhou ⋅ Yongqiang Zhang ⋅ Xiao Yan ⋅ Jiawei Jiang
The rapid emergence of heterogeneous large language models (LLMs) has highlighted the limitations of static single-model inference systems. Although model routing alleviates this issue by assigning the most suitable model to each query, existing methods often struggle to stably align query semantics with actual model performance and remain constrained by the capability ceiling of a single model. On the other hand, although multi-model aggregation can improve performance on complex queries, it also introduces extremely high computational overhead for routine tasks. In this paper, we propose PARS, a model confidence estimation method for dynamic LLM routing and aggregation. PARS first decouples the feature representation space and applies structured smoothing to local performance signals, producing more discriminative confidence score estimates for candidate models. On this basis, PARS represents inference modes such as direct routing, self-consistency sampling, and multi-model voting as test-time compute allocation actions under a budget constraint. By extracting state features such as routing confidence distributions and predictive uncertainty, a lightweight selector adaptively assigns an inference strategy to each query. Extensive experiments show that PARS improves average accuracy over the best baselines by 1.02% and 1.36% in the single-model and multi-model settings, respectively, while reducing average inference cost by 1.73x and exhibiting stable out-of-distribution generalization.
Conflict-Aware Logit Adapters for Utility-Preserving Anti-Distillation
Quoc P Dao ⋅ Binh Anh Nguyen ⋅ Linh Ngo ⋅ Thanh-Toan Do ⋅ Mehrtash Harandi ⋅ Trung Le
While reasoning traces improve model capability, they also create a pathway for extracting supervision signals for downstream distillation. Initial efforts to mitigate this rely on decoding-time interventions, but they incur high inference overhead and often harm the teacher model's natural fluency. We propose Conflict-Aware Logit Adapters (CALA), a lightweight training-time defense. CALA attaches a small residual adapter to a frozen teacher model and trains only this adapter. A proxy student identifies positions where imitation is most sensitive, and the adapter perturbs the output distribution there within high-probability regions. A conflict-aware optimization balances competing anti-distillation and utility objectives, with regularization preserving the teacher’s performance and fluency. At inference, CALA adds negligible overhead with no auxiliary models required. Experiments on reasoning benchmarks show that CALA substantially reduces successful student imitation while maintaining near-original teacher performance and generation quality, offering a practical and efficient alternative to decoding-time approaches.
Conservative neural posterior estimation via distributionally robust training
William Laplante ⋅ Yuga Hikida ⋅ Charita Dellaporta ⋅ Francois-Xavier Briol ⋅ Ayush Bharti
Simulation-based inference with neural posterior estimation (NPE) often yields overconfident and unreliable posteriors under limited simulation budgets. To address this, we propose DRO-NPE, a distributionally robust approach that replaces the standard NPE objective with a worst-case loss over a Wasserstein ambiguity set. We introduce KL-based metrics for miscoverage and miscalibration, and use these to show that the DRO-NPE objective controls overfitting and reduces posterior overconfidence. Our method is tractable, parallelisable, and readily integrates with standard normalising flows. Across benchmark SBI tasks, DRO-NPE consistently improves coverage and calibration, while narrowing the gap between empirical and population NPE loss, leading to more reliable inference in low-simulation regimes.
Test-time scaling often uses an external verifier, such as compilers and test-cases in coding or trained value functions in robotics application, to obtain high quality rollouts. Verifier-free test-time scaling (or VF-TTS) is gaining extensive attention as a mechanism to enhance Large Language Model (LLM) reasoning, primarily because in many real-world applications we do not have access to such high quality verifiers. Among existing VF-TTS methods, confidence-based VF-TTS methods, which compute and rank rollouts solely by confidence, are particularly promising. Such methods introduce near-zero overhead for sample evaluation and require minimal access to internal model states, making the methods highly flexible across models and data. In this paper, we demonstrate a critical limitation of existing confidence-based VF-TTS methods by showing that such methods catastrophically break down on complex tasks. We observe a very interesting phenomenon: uniformly high confidence frequently indicates a failure to explore, favoring confidently wrong answers. To address this, our core insight is that robust cognitive search requires a specific confidence trajectory pattern: such methods perform exploratory branching at the beginning as manifested by low initial confidence, and converge to a high final confidence solution. To implement this insight, we introduce consilience, a novel selection framework that explicitly evaluates the temporal asymmetry of confidence in reasoning. We operationalize this via a combinatorial metric that actively penalizes high initial confidence while strictly demanding final certainty. Extensive experiments covering both graduate-level mathematics problems and free-form code generation demonstrate that consilience effectively outperforms existing baselines, validating our novel perspective on completion confidence.
Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability
Harsh Raj ⋅ Niranjan Orkat ⋅ Suvrorup Mukherjee ⋅ Aritra Guha ⋅ Cheryl Flynn ⋅ Subhabrata Majumdar
This paper establishes a rigorous measurement science for AI agent reliability, providing a foundational framework for quantifying consistency under semantically preserving perturbations. By leveraging $U$-statistics for output-level reliability and kernel-based metrics for trajectory-level stability, we offer a principled approach to evaluating agents across diverse operating conditions. Our proposal highlights the important distinction between the core capability and execution robustness of an agent, showing that minor task-level variations can induce complete strategy breakdowns despite the agent possessing the requisite knowledge for the task. We validate our framework through extensive experiments on three agentic benchmarks, demonstrating that trajectory-level consistency metrics provide far greater diagnostic sensitivity than traditional pass@1 rates. By providing the mathematical tools to isolate where and why agents deviate, we enable the identification and rectification of architectural concerns that hinder the deployment of agents in high-stakes, real-world environments.
Constrained Goal-directed Planar Graph Generation with Grammar-based Reinforcement Learning
Nicolas Hochuli ⋅ Lorenzo Miele ⋅ Kristina Shea ⋅ Tino Stankovic
Planar graphs are central to applications across science and engineering, yet existing generators provide limited support for goal-directed generation under hard structural and geometric feasibility constraints. We propose a dataset-free method for generating planar graph embeddings by combining parametric graph grammars with safe reinforcement learning to optimize generic task-specific objectives while satisfying constraints during construction. We formulate the generation process as a constrained Markov decision process, where the graph grammar defines the state and action spaces. We further introduce an action projection that maps sampled actions toward state-dependent safe sets, improving constraint satisfaction during training. In contrast to classical graph generators and deep generative models, which typically offer limited goal-directed control or rely on weak constraint satisfaction, our method constructs feasible planar graph embeddings directly during generation. We also introduce a benchmark suite for constrained and goal-directed planar graph generation, together with classical and deep generative baselines. Across all benchmark tasks, our method consistently outperforms baselines while satisfying the formulated constraints.
COOP$^2$: Defining, Observing, and Repairing Cooperation in LLM Multi-Agent Systems
Hanqing Yang ⋅ Narjes Nourzad ⋅ Shiyu Chen ⋅ Marie Siew ⋅ Jingdi Chen ⋅ Carlee Joe-Wong
Many complex tasks require extended effort, diverse capabilities, or coordinated actions beyond what a single agent can provide. However, simply adding more agents does not guarantee better performance, as effective cooperation depends on how agents interact with each other and with task structure to satisfy evolving constraints over time. This challenge is amplified for LLM-based multi-agent systems (LLM-MAS): plans, messages, and revisions occur in natural language, whereas task progress depends on grounded environment actions. Current evaluations mostly treat cooperation as an implicit ingredient of final task success, leaving both cooperation and the effect of multi-agent interaction on task dynamics difficult to study. We introduce COOP$^2$, an evaluation framework that grounds high-level agent cooperation dynamics in LLM-MAS within task progress in the environment. COOP$^2$ then defines cooperative tasks with verifiable cooperative requirements, allowing us to analyze how cooperation unfolds over time with respect to task progress, as well as where and why cooperation breaks down. Building on this framework, we develop COOP$^2$-Repair, which predicts constraint failures from group plans and opens targeted repair channels for guided revisions. Across two environments and three communication structures, COOP$^2$-Repair improves task success and constraint satisfaction while exposing the additional decision overhead and communication load required for repair.
CoreQ: Learning-Free Mismatch Correction and Successive Rounding for Quantization
Seohyeon Cha ⋅ Huancheng Chen ⋅ Dongjun Kim ⋅ Haoran Zhang ⋅ Kevin Chan ⋅ Gustavo De Veciana ⋅ Haris Vikalo
Post-training quantization (PTQ) enables efficient deployment of large language models by mapping pretrained weights to low-bit formats without retraining, typically using a small calibration set to minimize a layer-wise calibration objective. However, this sequential procedure induces a mismatch: errors from earlier quantized layers alter the inputs received by later layers, causing the activations to deviate from those of the full-precision model. Recent approaches introduce mismatch-aware calibration objectives to compensate for this effect, but leave open how much of the observed mismatch should shift each layer's calibration target. Fully applying this correction can overfit limited calibration data, while scaling the mismatch correction with a fixed coefficient ignores varying reliability of mismatch estimates across layers. To address these limitations, we propose CoreQ, a learning-free PTQ framework that applies a closed-form coefficient for mismatch correction derived from a geometric decomposition of the mismatch. The resulting coefficient adapts the correction across layers, reduces overfitting to finite calibration data, and requires no hyperparameter tuning. Given the corrected target, CoreQ minimizes the induced triangular least-squares objective with an efficient greedy successive-rounding solver and a bounded beam-search extension, K-CoreQ, that trades modest additional compute for improved performance. Across multiple LLM families, scales, bit-widths, and quantization settings, CoreQ improves perplexity and downstream accuracy over strong PTQ baselines.
Correlation-Aware Contextual Bandits with Surrogate Rewards for LLM Routing
Ajay Narayanan Sridhar ⋅ Ronak Singh ⋅ Mehrdad Mahdavi ⋅ Vijaykrishnan Narayanan
We study contextual bandit problems with correlated arms and access to surrogate reward signals produced by a machine learning model, motivated by applications such as large language model (LLM) routing. Unlike classical contextual bandits that rely solely on bandit feedback and assume conditional independence across arms, our setting allows context-dependent inter-arm correlations and auxiliary reward information that may be noisy or misspecified. We propose algorithms that leverage such surrogate rewards through two complementary designs. A coupled reward-mixing approach pools true and surrogate rewards to accelerate learning when surrogate signals are reliable, while a decoupled prediction-mixing approach maintains separate estimators for bandit feedback and surrogate rewards and adaptively combines their predictions. This decoupling yields robustness to surrogate misspecification, recovering regret guarantees comparable to reward-only bandit methods in the worst case, while achieving improved regret when surrogate predictions are sufficiently informative. We provide theoretical regret analyses for both approaches and evaluate them on LLM routing benchmarks under varying accuracy versus cost trade-offs. The results demonstrate improved sample efficiency and consistently better accuracy–cost trade-offs compared to standard contextual bandit baselines and strong static routing methods.
Counterfactual Explanations for Time-Series Classification via Constrained Flow Matching
Akihiro Yamaguchi ⋅ Shizuo Kaji ⋅ Kaname Matsue ⋅ Ryusei Shingaki
Counterfactual explanations for time-series classification generate synthetic instances that flip a prediction to a desired class while remaining plausible and making small, local changes. Existing methods often rely on classifier gradients, and many do not naturally extend to one-class settings. To bridge this gap, we propose CTCF, a model-agnostic method for hard-decision classifiers in supervised and one-class settings. CTCF decouples generation from classifier-specific control by reusing an unconditional Flow Matching generator and learning a flow-time-dependent desired-class region from interpolation states labeled by endpoint hard decisions, without classifier gradients or class probabilities. At inference time, CTCF uses a dual-control mechanism: fused-lasso endpoint-objective steering encourages small, segment-local changes, while linearized projection correction reduces violations of a conservative desired-class region. Theoretically, under ideal marginal matching, we show that the surrogate risk on interpolation states equals the corresponding ODE-trajectory risk. Experiments on UCR datasets demonstrate CTCF's effectiveness across quantitative metrics and qualitative case studies.
Coverage-Verified Sparse Attention: Closed-Loop Quality Control for Long-Context LLM Decoding
Yan-Cheng Li ⋅ Chun-Yi Lee
Sparse attention accelerates long-context decoding by reading a subset of the KV cache, but existing methods are open-loop: the kernel commits to its selection with no signal indicating whether enough attention mass was preserved and no recovery path when it was not. The needed signal already exists inside the kernel. The fraction of softmax mass on the selected tokens, which we call coverage, satisfies an exact algebraic equality with per-step output error and serves a dual role as the acceptance criterion and the interpolation weight for recovery. Coverage-Verified Sparse Attention (CVSA) turns coverage into a training-free, closed-loop verify-then-correct decode loop that accepts high-coverage drafts and recovers only the heads that fall short, with no sampling, no draft model, and no speculative-decoding infrastructure. Turning verification on matters far more than tuning its threshold, and matched-budget comparisons confirm the gain is structural. On LongBench and RULER at 7B and 70B, CVSA closes the quality gap to dense attention to within statistical noise, with measured throughput reaching 6.21x at 128K context.
Cracks in the Foundation: A Civil Infrastructure Dataset to Challenge Vision Foundation Models
Nicola Farronato ⋅ Niccolò Avogaro ⋅ Thomas Frick ⋅ Mattia Rigotti ⋅ Rizwan U Khan ⋅ Michele Magno ⋅ Konrad Schindler ⋅ A. Cristiano I. Malossi ⋅ Florian Scheidegger
Automated structural health monitoring is essential to prevent catastrophic infrastructure failures. Precise, pixel-level defect segmentation is needed to accurately assess structural integrity, but progress in defect segmentation for civil infrastructures has been held back by an extreme scarcity of data, which requires costly expert annotation. The need for data is accentuated by algorithmic hurdles intrinsic to the problem, including center-bias and the need to rely more on shape when inspecting nearly textureless building materials. To remove the bottleneck, we introduce **Cracks in the Foundation (CiF)**, the largest and most detailed civil infrastructure (instance) segmentation dataset to date, comprising $\approx$150,000 high-resolution images meticulously curated over five years in collaboration with civil engineering experts. With the help of this unprecedented data source, we expose a blind spot of current visual AI: despite the advent of promptable Foundation Models (FMs) and Vision Language Models (VLMs), and despite the impressive abilities of today's specialised segmentation models, it turns out that dense image understanding in the built environment is nowhere near solved. Our evaluations indicate that even the most recent zero-shot FMs face significant challenges when deployed on real-world infrastructure and even the performance of specialised models with domain-specific supervision plateaus at $\approx$25\% mAP. CiF establishes inspection of civil infrastructure, an elementary and seemingly easy perceptual task, as an open challenge that reveals fundamental weaknesses of present-day models trained predominantly on internet images, literally and figuratively highlighting *cracks in the current foundation model paradigm*.
CrafterDojo: A Suite of Foundation Models for Building Open-Ended Embodied Agents in Crafter
Junyeong Park ⋅ Hyeonseo Cho ⋅ Sungjin Ahn
Foundation models are central to building general-purpose embodied agents. Minecraft has become the dominant testbed for this paradigm, but its computational and engineering overhead limits rapid iteration and controlled analysis. Crafter offers a lightweight alternative yet lacks the foundation model ecosystem required for such research. We introduce CrafterDojo, a suite of foundation models, datasets, and data-generation toolkits that establishes Crafter as a testbed for foundation-model-based embodied intelligence, providing CrafterVPT, CrafterCLIP, and CrafterSteve-1. To address Crafter's data scarcity, CrafterDojo includes automatic tools that generate diverse synthetic demonstrations and aligned video-language pairs, yielding the CrafterPlay and CrafterCaption datasets. Experiments show that these datasets support effective foundation model training and enable hierarchical agents that outperform non-hierarchical baselines on long-horizon tasks. We release all code, models, and datasets.
CRISP: Compositional Reasoning over Images via Stackable Programs for VLMs
Arnas Uselis ⋅ Yujin Jeong ⋅ Yanpeng Zhao ⋅ Alexander Rubinstein ⋅ Seong Joon Oh ⋅ Yonatan Bitton ⋅ Paul Gavrikov
How well do vision-language models actually reason about what they see? Current diagnostic benchmarks offer limited answers: they draw from fixed question templates, evaluate only the final answer, and are tied to single visual environments. We introduce CRISP (Compositional Reasoning over Images via Stackable Programs), a platform that generates compositional visual reasoning tasks from reusable, composable operations. Given any scene with a structured specification, the platform constructs reasoning chains of arbitrary depth where each intermediate step has automatically derived ground truth, including bounding boxes for the objects the model should attend to at every stage. New visual environments and reasoning operations plug in without modifying the generation engine. We instantiate the platform across three environments spanning 2D sprites, photorealistic 3D characters, and indoor room scenes, and show that frontier VLMs degrade systematically with reasoning depth and frequently arrive at correct answers through incorrectly grounded intermediate reasoning. Beyond benchmarking, the generated reasoning traces and step-wise ground truth can support future work on training and improving visually grounded reasoning models.
Crosscoding Through Time: Sparse Feature Discovery Across Sequence Positions
Dmitry Manning-Coe ⋅ Han Xuanyuan ⋅ Aniket Deshpande ⋅ Andrii Shportko ⋅ William Fei
Dictionary learning methods - such as Sparse Autoencoders (SAEs) and crosscoders - decompose model activations into human-interpretable building blocks. We introduce temporal crosscoders, a simple and flexible framework for feature discovery in Large Language Models (LLMs). To properly evaluate temporal crosscoders we develop TempBench: a panel of synthetic and real-world tasks for evaluating temporal structures. Temporal crosscoders outperform both conventional and temporal architectures in both of our synthetic settings and on two out of four of the real world settings - more than any other current architecture. Most strikingly, they can detect backtracking - a key reasoning behavior - at a 40\% higher rate than conventional SAEs, and are 15\% more effective in inducing it. Our results establish temporal crosscoders as a simple and flexible framework for feature discovery, both local and temporal. We provide full code at the following anonymous repository: \url{https://anonymous.4open.science/r/temp-bench-anon/}.
Cross Flow: One-Step Generation Across Latent and Pixel Spaces
Xiyuan Wang ⋅ Xiao Zhang ⋅ Yang Li ⋅ Ruoxi Jiang ⋅ Zhao Zhong ⋅ Liefeng Bo ⋅ Muhan Zhang
Diffusion and flow-matching models are typically constrained to a single manifold, where the noise prior, intermediate trajectory, and prediction target share the same dimensionality. While Latent Diffusion Models (LDMs) mitigate computational costs by operating in compressed spaces, they still rely on a decoupled, fixed decoder to map latents back to pixels. We introduce CrossFlow, a generative paradigm that unifies these stages by allowing the noise prior and final output to reside in different spaces. By deriving a novel cross-space objective, our framework enables a single model to map directly from a noisy latent to a high-resolution image in a single function evaluation (NFE). CrossFlow serves a dual purpose: it acts as a high-fidelity one-step generator and functions as an enhanced decoder for existing LDM pipelines, capable of refining imperfect latent estimations during the reconstruction process. On ImageNet-1k ($256 \times 256$), CrossFlow achieves a state-of-the-art 1.62 FID with only one NFE. Our results demonstrate that cross-space flow objectives provide a scalable and theoretically grounded framework for unifying latent generation and pixel-space decoding.
CrossID: Cross-Supervised Spatio-Temporal Gated Fusion for Personalized Portrait Generation
Dongxu Yue ⋅ Qixin Yan ⋅ Shiao Yang ⋅ Xiaoqiang Zhou ⋅ Hao Liu ⋅ Zhihai He ⋅ Chun Yuan
Personalized portrait generation aims to synthesize portraits that align with both the identity of the reference face and the semantics of the text. However, existing methods suffer from two critical challenges. First, they exhibit a trade-off between ID fidelity and text controllability. We identify that this limitation stems from the self-supervised training, where the model learns a trivial mapping rather than understand identity. Second, despite exploring various encoding strategies, recent approaches still tend to discard fine-grained facial details. To address these challenges, we propose CrossID-11M, a large-scale curated portrait dataset comprising 11 million images with high-quality annotations, paired with millions of identity-preserving face video clips. Furthermore, we introduce CrossID, a novel framework integrating a lightweight GatedFaceAdapter and a cross-supervised training paradigm. Additionally, we establish CrossID-Bench alongside new VLM-based evaluation metrics. Extensive experiments demonstrate that CrossID successfully breaks the trade-off between ID fidelity and controllability, achieving state-of-the-art performance by producing photorealistic images with faithfully preserved ID details.
Cross-Layer Evolution Graph Learning for Fine-grained VLM Hallucination Detection
Shaoxuan Li ⋅ Shiliang Zhang
Although Vision-Language Models (VLMs) demonstrate impressive capabilities, their deployment in high-stakes domains is severely hindered by hallucinations. Existing detection methods mostly focus on static internal feature probing with coarse-grained annotations, failing to achieve fine-grained token-level localization without expensive dense labels. In this paper, we reconceptualize VLM hallucination detection as a dynamic graph learning task, identifying hallucinations as distinct trajectory drifts across network layers. Consequently, we propose TRACE (Token Representation Analysis with Cross-layer Evolution), a non-intrusive probing framework that models internal states as heterogeneous spatio-temporal graphs. By employing Graph Neural Networks (GNNs) to capture computational dynamics, TRACE effectively isolates hallucinatory signals. Furthermore, a token-first Multiple Instance Learning (MIL) readout elegantly bridges the granularity gap, distilling macroscopic sentence-level supervision into precise zero-shot token-level localization. Extensive evaluations on M-HalDetect, ViGoR, and HalLoc demonstrate that TRACE achieves state-of-the-art detection performance while being competitive in localization accuracy, offering an efficient and interpretable solution for trustworthy VLMs.
Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse
Taekyung Heo ⋅ Rasoul Shafipour ⋅ Ritchie Zhao ⋅ Maximilian Golub ⋅ Mohammad Mahdi Kamani ⋅ Ritika Borkar ⋅ Makesh T Chandran ⋅ Pantea Zardoshti ⋅ Bita Darvish Rouhani
Production deployments often swap between different-sized models in a family for cost-quality cascading, mid-conversation switching, and routing, and each swap forces the receiver to repay the prefill from scratch. We propose cross-model KV cache transfer, where the receiver reuses the source’s KV cache, skipping prefill. We find that cross-model KV has substantial linear structure across matched-KV pairs, where source and target share KV head count and per-head dimension. On Qwen3 14B→32B, one source layer explains 56% of variance in the target’s keys and 32% in values, rising to 79% and 65% with multiple source layers. Building on this, we design a closed-form ridge mapper that operates per head and proceeds in three steps. First, for each target layer we select the top-kmost predictive source layers and concatenate their KV as input. Second, we strip RoPE from the keys before mapping, so the fit is position-free and reusable across context lengths. Third, we fit ridge regression on a small calibration set of 500 FineWeb-Edu sequences of 1,024 tokens each. Surprisingly, across six pairs in three families, this linear mapper retains 71–98% of the receiver’s standalone-prefill accuracy on four pairs, while two degrade sharply. A nonlinear MLP recovers up to +22 pp HellaSwag retention on the failures. The mapper runs 3–31×faster than re-prefill and remains stable across multi-turn handoff, making cross-model KV cache transfer practical.
CruxBench: A Benchmark of Information Discovery
Lina Piao ⋅ Amelia Hui Dai ⋅ Nick Merrill ⋅ Nadja Flechner ⋅ Ezra Karger ⋅ Haifeng Xu
Benchmarks for large language models typically evaluate the accuracy of answers against fixed reference labels. But a central step in many complex real-world tasks is identifying which questions are worth asking in the first place: decomposing a difficult problem into subquestions whose answers provide key steps on the path toward solving the larger problem. To evaluate this capability of information discovery, we introduce CruxBench, a benchmark that grades LLM-generated questions by their Value of Information (VOI): how much a model-proposed subquestion updates beliefs about a target outcome. To automate this measurement at scale, we target as outcomes the prices in continuously updating forward-looking prediction markets. Our benchmark has three rare properties: it is contamination-resistant by construction, since ground truth is generated by future world events; it is open-ended, admitting unbounded and complex text-based submissions rather than one correct numeric answer; and it is grounded, with informativeness measured against quantified changes in real-world beliefs (prices). We evaluate eight frontier and open-weight models on 336 target outcomes and find that VOI correlates highly with independent measures of model quality ($r=0.87$) and predicts downstream usefulness when the generated questions are passed to a separate forecasting agent.
CrystalREPA: Transferring Physical Priors from Universal MLIPs to Crystal Generative Models
Chengqian Zhang ⋅ Yucheng Jin ⋅ Duo Zhang ⋅ Tiejun Li ⋅ Han Wang
Crystal generative models mainly learn what stable crystals look like, with little explicit supervision for what makes them stable. We reveal a substantial representation gap between state-of-the-art crystal generative models and pretrained universal machine learning interatomic potentials (MLIPs) via energy probing, and show this gap can be closed by a simple training-time alignment. We propose Crystal REPresentation Alignment (CrystalREPA), a plug-and-play framework that aligns the atom-wise hidden states of generative encoders with frozen MLIP representations through an element-aware contrastive objective, transferring stability-aware atomistic priors with marginal training overhead and no additional inference cost. Across three generative frameworks, ten MLIP teachers, and two benchmark datasets, CrystalREPA consistently improves the thermodynamic stability, structural validity, and structural fidelity of generated crystals. Equally important, we find that an MLIP's transfer effectiveness is poorly predicted by its accuracy on standard leaderboards (e.g., Matbench Discovery) but strongly predicted by the distinguishability of its atom-wise representation space, yielding a practical, accuracy-independent criterion for selecting MLIP teachers for generative transfer.
CURE: Visual Reprogramming of Vision-Language Models under Limited Supervision
Lingzhi Wang ⋅ Wei Wang ⋅ Xinyang Chen
Visual reprogramming (VR) efficiently repurposes pre-trained vision-language models for new image classification tasks by adding trainable patterns to inputs without modifying the backbone. Nevertheless, existing VR methods heavily depend on labeled data and thus suffer substantial performance degradation under limited supervision (i.e., scarce labeled images). We introduce semi-supervised learning (SSL) into VR to exploit abundant unlabeled images and enhance data efficiency. However, directly applying generic SSL techniques often amplifies biases in unlabeled posteriors and destabilizes VR training. To address these challenges, we propose CURE (Consistency and Unlabeled Recalibration with Confidence-Margin Enhancement), a semi-supervised VR framework comprising two key components: (i) Recalibreted Gaussian Confidence-Margin Soft Weighting to dynamically adjust sample importance; (ii) Dual-attribute Consistency Regularization to ensure consistency across attribute prompts. Extensive experiments show that CURE consistently outperforms both existing VR methods and direct SSL extensions, improving average classification accuracy by 5.06 and 6.67 percentage points across twelve widely-used benchmarks for limited supervision tasks. Our code is available at https://anonymous.4open.science/r/CURE-93E2/.
Curvature-Dependent Lower Bounds for Riemannian Online Convex Optimization
Hibiki Fukushima ⋅ Shinji Ito
In Euclidean online convex optimization, first-order and full-information feedback share the minimax regret (\Theta(DL\sqrt{T})) over (T) rounds, as convex losses can be replaced by their linearizations. On Riemannian manifolds this reduction fails, because the corresponding first-order surrogate of a geodesically convex loss is generally not geodesically convex; we show that this failure has provable consequences. For online convex optimization on Hadamard manifolds with sectional curvature bounded below by (\kappa<0) and feasible diameter (D), we construct hard hyperbolic-space instances yielding first-order regret lower bounds at the scale ( DL\sqrt{\zeta\, T}, ) where ( \zeta=\sqrt{|\kappa|}D/\tanh(\sqrt{|\kappa|}D) ) is the curvature--diameter factor. Our first lower bound applies to all (possibly randomized) first-order algorithms satisfying a natural geometric-span condition, holding almost surely over the algorithm's internal randomness. Our second lower bound removes the geometric-span restriction for deterministic first-order algorithms, provided the feasible diameter satisfies (D=\Omega(\log T)); the proof rests on a new construction that confines the hidden comparator to a single horosphere in hyperbolic space. Combined with existing curvature-free full-information upper bounds, these results yield a feedback-model separation on Hadamard manifolds: full-information algorithms attain (O(DL\sqrt{T})) regret, yet the additional (\sqrt{\zeta}) factor is unavoidable for first-order feedback whenever either the geometric-span condition holds or the algorithm is deterministic.
DARTS: Targeting Prognostic Covariates in Budget-Constrained Sequential Experiments
Kateryna Husar ⋅ Alexander Volfovsky
Randomized controlled trials typically assume that prognostic covariates are known and available at no cost. In practice, obtaining high-dimensional pretreatment data is costly, forcing a trade-off between covariate-adaptive precision and a measurement budget. We introduce Dynamic Adaptive Rerandomization via Thompson Sampling (DARTS), which treats covariate acquisition as a sequential optimization problem embedded within a design-based causal inference task. A budgeted combinatorial Thompson sampler learns which covariates are most prognostic across successive batches; selected covariates then drive rerandomization and regression adjustment to reduce batch-level average treatment effect variance. Our primary theoretical contribution is a decoupling result: adaptive covariate selection based on past batches preserves batch-level randomization validity, and the cumulative inverse-variance weighted estimator achieves at least nominal asymptotic coverage. We further derive a Bayes risk bound for the acquisition layer that matches the minimax lower bound up to logarithmic factors. Empirically, DARTS systematically concentrates budget on informative features, significantly closing the efficiency gap to oracle designs while maintaining strict inferential validity.
Darwin-7B: A Multi-Omic Foundation Model for the Human Gut Microbiome via Sparsified Quality-Aware Tokenization
Arvid Ernst Gollwitzer ⋅ David de Gruijl
Microbial ecosystems govern human health, agricultural productivity, and biogeochemical cycling. Two coupled modalities determine their function: which organisms are present (metagenomic content) and which small molecules they produce and consume (metabolomic content). Existing biological foundation models cannot represent this joint state: they operate on a single modality, on assembled single-organism genomes, and on the within-DNA axis of generalization, leaving the multi-omic, mixed-organism, mixed-quality regime that governs clinical and surveillance use unaddressed. We present Darwin-7B, a 7-billion-parameter multi-omic foundation model whose primary axis of generalization spans both modalities (metagenome + metabolome) and organizational scales (read → community → host phenotype). We combine a Mamba–Transformer hybrid backbone (24 of 32 layers Mamba) with a hypergraph neural network over KEGG metabolic-reaction hyperedges and bidirectional cross-modal attention modules that align genomic and metabolomic representations in a shared 40,192-token embedding; an Aitchison-space compositional consistency loss respects the simplex geometry of microbial-abundance data at the output. We pretrain on 8 T base pairs of metagenomics, 250K LC-MS/MS metabolite profiles, and 2M KEGG/GO functional annotations, preparing the metagenomic corpus with a published quality-aware tokenization framework. Against the matched 7B genomic baseline METAGENE-1, we report 94.5 MCC on Pathogen Detection (vs. 93.0) and 0.98 F1 on CAMI species-level metagenomic profiling; on a December 2025 SRA outbreak benchmark post-dating the pretraining cutoff, we report 91.0 MCC vs. 81.2. We unlock four clinical tasks single-modality genomic models cannot reach: IBD AUC 0.947, T2D AUC 0.883, antibiotic-resistance AUC 0.910, and metabolic-pathway prediction wF1 0.91, beating the strongest multi-omic baseline (MOGONet) by +3.0–+3.7 AUC. Token-matched scaling from 100M to 70B yields monotone gains with no saturation on the multi-omic clinical AUCs. We further validate predictions in a prospective wet-lab pilot (n = 87) at predicted-vs-observed Pearson r = 0.72. Sparse-autoencoder probing recovers 14 high-coherence biological features tracking KEGG pathways, NCBI taxa, and CARD AMR genes (r = 0.68–0.74); zeroing the butyrate-producer feature drops IBD AUC from 0.947 to 0.928. We conclude that multi-omic, multi-scale modeling is tractable at 7B parameters, and we release the model card, training-data manifest, evaluation pipelines, and a 100-trajectory MetaOmics-10T causal-pilot dataset.
Data-Free Metrics Are Not Invariant Under Functionality-Preserving Reparametrisations
Gabryel Mason-Williams ⋅ Israel Mason-Williams ⋅ Fredrik Dahlqvist
Data-free methods for analysing and understanding the layers of neural networks offer many metrics for quantifying notions of strong' versusweak' layers, with the promise of increased interpretability. In particular, random matrix theory (RMT) offers data-free metrics that claim predictive power over the quality of pre-trained models at a layerwise precision and the ability to identify model pathologies, indicating a unique relationship with generalisation. As a result, metrics from RMT are championed for pre-training optimisation strategies and post-training compression. We establish that RMT-based metrics are unrelated to performance or training by showing that functionally indistinguishable reparametrisations of a pre-trained model can have arbitrary metrics. We show this across a range of architectures and scales. To create functionally indistinguishable reparametrisations, we exploit the well-established phenomenon of criticality: some layers can be re-initialised or re-randomised without affecting the functional behaviour of the model -- they are called robust -- while others cannot -- they are called critical. Re-initialising or re-randomising robust layers provides functionally indistinguishable reparametrisations for which RMT-based metrics are arbitrary. Moreover, we show that relationships between metrics offered by RMT are spuriously related to generalisation. We conclude by showing that if many pre-trained models have data-free metrics in a `good' range, it is, in part, dependent on model initialisation.
DC-Ocean: Deep Latent Compression for Global High-Resolution Ocean Forecasting
Yuxiang Li ⋅ Qiusheng Huang ⋅ Hao Li
High-resolution global ocean states expose a bottleneck in data-driven forecasting: effective long-horizon rollouts require a compact, decodable latent state that reduces I/O and stabilizes error growth. While learned compression has matured for natural RGB images, its assumptions do not transfer to ocean modeling: ocean states are multi-variable and multi-depth, their grids are orders of magnitude larger, and fixed land--sea boundaries create sharp discontinuities that make aggressive downsampling error-prone. We present DC-Ocean, the first deep latent compression architecture for multi-variable, multi-depth global $1/12^\circ$ ocean states, built on an autoencoder that couples convolutional feature extraction with transformer-based global context modeling to retain both coastal sharpness and basin-scale structure. To explicitly address coastline-induced artifacts, DC-Ocean integrates boundary-aware gated convolutions together with a spherical geodesic front propagation interpolation scheme for stable behavior near land masks. Across key ocean variables, DC-Ocean achieves higher reconstruction fidelity than strong autoencoder baselines adapted from the natural-image domain at matched downsampling factors. Beyond reconstruction, we perform autoregressive forecasting directly in latent space and decode predictions back to the physical domain. DC-Ocean matches short-term accuracy of full-resolution forecasting while achieving lower RMSE and slower error growth at medium and long lead times. Together, these results demonstrate that DC-Ocean provides a practical framework for stable and extended high-resolution sub-daily ocean forecasting.
DDGE: Disentangled Dirichlet Geodesic Evaluation for Robust Few-Shot Learning
Yifan Mei ⋅ Kun Zhou ⋅ Zelei Wu ⋅ Jieyu Zhao ⋅ Xulun Ye
Recent studies have highlighted the crucial role of distance metrics in noisy few-shot learning. However, existing approaches heavily rely on Euclidean or cosine measurements, failing to recognize the ``deceptive proximity'' where heterogeneous samples from different manifold peaks appear spuriously close in the flat ambient space. In this paper, we present Disentangled Dirichlet Geodesic Evaluation (DDGE), a novel manifold-aware framework that unifies orthogonal disentanglement and conformal integrals to efficiently evaluate complex topological spaces. Specifically, we decouple features into orthogonal semantic subspaces and leverage a prior-guided Dirichlet Process to expand discrete samples into a continuous, noise-purified semantic terrain. Upon this landscape, we measure the intrinsic geodesic distance via a density-aware conformal line integral to capture the authentic class topology. Extensive experiments demonstrate the state-of-the-art performance of DDGE on multiple benchmark datasets.
We propose a fully decentralized $Q$-learning dynamic for infinite-horizon discounted Markov potential games in which agents observe the global state and their own realized payoffs, but do not know the reward functions, the transition probabilities, the potential function, or the policies and the actions of the other players. In the proposed dynamic each player maintains a local estimate of its continuation payoff for each state and action and updates its policy through a smoothed best response. Importantly, the smoothing parameter is adjusted by using an online estimate of the discounted state-visitation distribution, so that players explore more in rarely visited states while exploitation is encouraged in states that are frequently visited. We prove that the induced policy sequence converges almost surely to a set of approximate Nash equilibria of the Markov potential game, with an approximation guarantee that depends only on easily computable parameters. We illustrate the performance of the proposed learning dynamic via numerical experiments on a routing game and compare it with existing independent learning schemes based on persistent $\theta$-greedy exploration.
Deep Heteroskedastic Regression: Post-Hoc Variance Estimation from Latent Representations
Mikkel Jordahn ⋅ Jonas Vestergaard Jensen ⋅ James Harrison ⋅ Michael Andersen ⋅ Mikkel Schmidt
Uncertainty quantification (UQ) in deep learning regression is of wide interest, as it supports critical applications including sequential decision making and risk-sensitive tasks. In heteroskedastic regression, where the uncertainty of the target depends on the input, a common approach is to train a neural network that parametrises the mean and the variance of the predictive distribution. Yet to this day, training deep heteroskedastic regression models poses severe practical challenges in the trade-off between uncertainty quantification and mean prediction, such as variance overfitting, optimization difficulties and representation collapse. In this work, we identify the core issues and propose a simple and efficient procedure that addresses these jointly by post-hoc fitting a variance model across the intermediate layers of a pretrained network on a hold-out dataset. We demonstrate that this method is competitive with end-to-end trained mean-variance networks in heteroskedastic UQ on several data modalities. The method retains mean prediction accuracy, is cheap at both train and prediction time, and requires no additional data compared to existing methods. Finally, we show this method works on large scale foundation models for chemistry, paving the way for cheap heteroskedastic UQ in large scale neural networks without having to retrain such models end-to-end.
Defining Operational Conditions for Safety-Critical AI-Based Systems from Data
Johann M Christensen ⋅ Elena Hoemann ⋅ Frank Köster ⋅ Sven Hallerbach
Artificial Intelligence (AI) has been on the rise in many domains, including numerous safety-critical applications. However, for complex systems in the real world, defining the underlying environmental conditions in which the AI-based system must operate---the Operational Design Domain (ODD)---is extremely challenging. This often results in an incomplete description of the ODD, which contrasts with the requirements of many domains for certifying AI-based systems. Traditionally, the ODD is created in the early stages of the development process, drawing on sophisticated expert knowledge and related standards. This paper presents a novel Safety-by-Design method to a posteriori define the ODD from previously collected data using a multi-dimensional kernel-based representation. This approach is validated through both Monte Carlo methods and a real-world aviation use case for a future collision-avoidance system. Moreover, by defining under what conditions two ODDs are similar, the paper shows that the data-driven ODD can produce a dataset similar to the original, hidden ODD. Deriving the novel, Safety-by-Design, deterministic kernel-based affinity representation of ODDs is fully automated via a bounded, order-independent algorithm. Utilizing the proposed ODD representation enables future certification of data-driven, safety-critical AI-based systems.
DegradeQuery: Counterfactual Tuple Pretraining for Context-Aware PROTAC Degradation Prediction
Dong Xu ⋅ Zhangfan Yang ⋅ Jiantao Wu ⋅ Zexuan Zhu ⋅ Jianqiang Li ⋅ Junkai Ji
Proteolysis-targeting chimeras (PROTACs) induce protein degradation through coordinated interactions among a degrader molecule, a target protein, and an E3 ubiquitin ligase. Yet computational prediction is often framed as if degradation were an intrinsic property of the degrader alone, overlooking the tuple-level context that determines activity. This mismatch is especially consequential in public PROTAC datasets, where degradation labels are sparse but molecule–target–E3 records are abundant. We introduce DegradeQuery, a tuple-conditioned framework that uses unlabeled records as structured relational evidence rather than incomplete labeled examples. DegradeQuery first performs counterfactual tuple pretraining, learning to distinguish observed molecule–target–E3 tuples from alternatives generated by replacing the target, the E3 ligase, or both. This objective induces a conditional compatibility representation before degradation labels are used, enabling the model to capture how degrader activity depends on biological context. The pretrained representation is then fine-tuned for high/low degradation prediction. On the official PROTAC-8K benchmark, DegradeQuery achieves 0.9065 AUROC and 0.8500 accuracy, outperforming reported semi-supervised results without pseudo-labeling, teacher models, distillation, or ensembles. Ablations, scaffold holdout, target–E3 holdout, and target-wise few-shot adaptation further support PROTAC degradation prediction as a tuple-conditioned rather than molecule-only problem.
Delta-Adapter: Scalable Exemplar-Based Image Editing with Single-Pair Supervision
Jiacheng Chen ⋅ Songze Li ⋅ Han Fu ⋅ Baoquan Zhao ⋅ Wei Liu ⋅ Yanyan Liang ⋅ Qing Li ⋅ Xudong Mao
Exemplar-based image editing applies a transformation defined by a source-target image pair to a new query image. Existing methods rely on a pair-of-pairs supervision paradigm, requiring two image pairs sharing the same edit semantics to learn the target transformation. This constraint makes training data difficult to curate at scale and limits generalization across diverse edit types. We propose Delta-Adapter, a method that learns transferable editing semantics under single-pair supervision, requiring no textual guidance. Rather than directly exposing the exemplar pair to the model, we leverage a pre-trained vision encoder to extract a semantic delta that encodes the visual transformation between the two images. This semantic delta is injected into a pre-trained image editing model via a Perceiver-based adapter. Since the target image is never directly visible to the model, it can serve as the prediction target, enabling single-pair supervision without requiring additional exemplar pairs. This formulation allows us to leverage existing large-scale editing datasets for training. To further promote faithful transformation transfer, we introduce a semantic delta consistency loss that aligns the semantic change of the generated output with the ground-truth semantic delta extracted from the exemplar pair. Extensive experiments demonstrate that Delta-Adapter consistently improves both editing accuracy and content consistency over four strong baselines on seen editing tasks, while also generalizing more effectively to unseen editing tasks. Code will be made publicly available.
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences
Yankai Yang ⋅ Yancheng Long ⋅ Bin Wen ⋅ Fan Yang ⋅ Tingting Gao ⋅ Han Li ⋅ Shuo Yang
Video multimodal large language models have made strong progress on open-ended video understanding, but they still lack precise local spatiotemporal perception. When two videos share almost the same global semantics and differ only in a short time span or a small region, current models often fail to find the change and provide reliable evidence. We propose DeltaVid, a verifiable proxy-task framework that enhances fine-grained spatiotemporal perception with cross-video differences. The key idea is to turn cross-video spot-the-difference into a trainable perception signal, where a model identifies local changes, judges temporal boundaries, and organizes spatial evidence by comparing similar videos. To make this signal scalable to train and reliable to evaluate, we further introduce DeltaVid-10K and DeltaVid-Bench, which convert controllable local differences in real videos into evidence-labeled training and test samples. Experiments show that DeltaVid substantially improves performance on cross-video difference understanding and transfers the learned local evidence ability to general video understanding benchmarks, including MMVU, MLVU, Video-MME, VideoHolmes, VideoMMMU, LVBench, TempCompass, and LongVideoBench. These results show that cross-video differences are not only an effective way to diagnose fine-grained perception failures, but also a scalable proxy supervision that moves Video MLLMs from coarse semantic understanding toward fine-grained spatiotemporal evidence reasoning.
Fine-scale-faithful neural simulation under fixed storage budgets remains challenging. Many existing methods reduce high-frequency error by improving architectures, training objectives, or rollout strategies. However, under budgeted coarsen-quantize-decode pipelines, fine detail can already be lost when the carried state is constructed. In the canonical periodic incompressible Navier--Stokes setting, we show that primitive and derived fields undergo systematically different retained-band distortions under the same operator. Motivated by this observation, we formulate Derived-Field Optimization (DerivOpt), a general state-design framework that chooses which physical fields are carried and how storage budget is allocated across them under a calibrated channel model. Across the full time-dependent forward subset of PDEBench, DerivOpt not only improves pooled mean rollout nRMSE, but also delivers a decisive advantage in fine-scale fidelity over a broad set of strong baselines. More importantly, the gains are already visible at input time, before rollout learning begins. This indicates that the carried state is often the dominant bottleneck under tight storage budgets. These results suggest a broader conclusion: in budgeted neural simulation, carried-state design should be treated as a first-class design axis alongside architecture, loss, and rollout strategy.
Designing Effective Monitor-Based Interventions for Mitigating Reward Hacking During RL
Aria Wong ⋅ Joshua Engels ⋅ Neel Nanda
Reinforcement learning (RL) rewards are notoriously difficult to design and control, often leading to the model learning unintended behaviors such as eward hacking. One potential solution is to monitor for reward hacking and penalize it when detected; however, training against a monitor could lead to evasive behavior, and our general understanding of how to apply monitors effectively during training is limited. To study how best to use monitors to mitigate reward hacking, we introduce and open source three realistic environments where Qwen3-4B reward hacks: a coding environment hackable via test overwriting, a medical chat environment hackable via sycophancy, and a biography generation environment hackable via hallucination. We first focus on the coding environment, where we find that: (1) models can learn to evade highly accurate monitors by exploiting systemic flaws in probes and LLM judges; (2) monitors that leak more learning signal during RL suppress reward hacking but are more often evaded; and (3) including easier problems in training can decrease reward hacking. We apply our findings to build better reward hacking monitors for the medical chat and biography generation environments that improve upon naive baselines to reduce reward hacking rates across seeds from 70-100\% to 0\%. Our results demonstrate that our takeaways translate to new settings and that better monitor intervention designs are possible.
Detect Anything in Graphic Design: Element-Level Rewards for Autoregressive Detection
Jiangning Zhu ⋅ Bowen Li ⋅ Shenyu Qiao ⋅ Yima Gu ⋅ Zhao Zhang ⋅ Yuhui Yuan ⋅ Shixia Liu
Graphic designs, such as posters, advertisements, and infographics, are an important medium for communicating information and shaping understanding. Unlike natural images, they consist of layered elements with explicit compositional order. However, existing object detection models treat these elements as an unordered set, leaving compositional order unexploited. To address this limitation, we present Detect Anything in Graphic Design (DAD), a model that formulates graphic design detection as compositional deconstruction. It decodes elements in compositional order, using lower-layer elements to better detect higher-layer ones. The key feature of DAD is amodal detection, which predicts the full bounding box of each element, including regions occluded by elements placed above it. Building on this formulation, we propose Element Relative Policy Optimization (EleRPO), which extends GRPO from sequence-level supervision to element-level optimization. EleRPO provides fine-grained training signals that capture how each detected element contributes to overall detection quality, and works synergistically with compositional order to improve detection performance. To support training and evaluation, we build a dataset of 10 million graphic designs. Experiments show that DAD outperforms all baselines and achieves human-level performance in amodal detection, supporting effective image-to-layer decomposition. EleRPO consistently improves over GRPO across nine detection benchmarks.
DetectViT: Test-time Backdoor Detection for Vision Transformers via Inter-Head Attention Discrepancy
Siquan Huang ⋅ Yijiang Li ⋅ Xingfu Yan ⋅ Ningzhi Gao ⋅ Leyu Shi ⋅ Ying Gao
Vision Transformers (ViTs) have been widely adopted as visual encoders in multimodal models; however, the reliance on third-party pretrained checkpoints exposes these systems to backdoor attacks. Existing test-time detection methods rely on either the predicted label or intermediate embeddings, limiting their applicability across diverse tasks and incurring substantial computational overhead. In this work, we investigate how backdoor triggers affect the multi-head attention mechanism of ViTs and identify two distinct anomalous patterns: $\textit{complete attention hijacking}$, causing the attention to exhibit abnormally high inter-head spatial overlap, and $\textit{partial attention hijacking}$, causing an abnormally large variance in the concentration of the attention distribution. Motivated by these findings, we propose $\textbf{DetectViT}$, a test-time backdoor detection method that requires neither training data, model outputs, nor any prior knowledge of the trigger. DetectViT quantifies the two hijacking phenomena via an inter-head consistency score and an inter-head entropy variance score, with thresholds estimated from a small out-of-distribution (OOD) calibration set drawn from any source. Extensive experiments on four representative attacks, spanning diverse architectures (DeiT, CLIP, LLaVA) and downstream tasks (image classification and captioning), demonstrate that DetectViT substantially outperforms all baselines with as few as 64--128 OOD calibration samples, attaining a true positive rate of $100\%$ and a false positive rate of $0.17\%$ on BadCLIP. Notably, it introduces negligible additional overhead, thanks to its reliance solely on attention weights already computed during inference. We release our code at: https://anonymous.4open.science/r/DetectViT-81AC.
Diagnosing Math-Reasoning Failure Structure with Milestone Oracles
Zhuohan Wang ⋅ Haoran Ma ⋅ Tianyu Wu ⋅ Yuanlin Duan ⋅ Zichun Liao ⋅ Jieming Yu
Aggregate math benchmark scores tell whether a model solved a problem, but not where reasoning broke down. We introduce milestone-oracle probing, a symbolically verified diagnostic that turns math evaluation into a reasoning-gap readout. For each parent problem, a teacher writes a fixed milestone roadmap. We evaluate a student under three levels of help: no help, the roadmap, and the roadmap plus gold milestone answers. We also run a separate milestone-only test, asking whether the student can solve each milestone in isolation. Deterministic symbolic verification grades all parent and milestone answers. Crossing these two measurements yields five gap categories: roadmap, milestone-execution, composition, missing-milestone, and capability. Applied to 354 NuminaMath problems and a six-model panel from 8B to 671B parameters, the largest residual category across all six models is families that pass the milestone test but are not recovered under any of the three probes (33--48\% of families per model). A follow-up audit on a broader sample of unrecovered families shows the residual is mixed: it contains real composition failures together with decomposition gaps and verifier artifacts. The protocol surfaces the residual; the audit reads it into its component parts. We release the diagnostic set, prompt templates, symbolic verifier, audit annotations, and the leak-safe repair logs and bookkeeping scripts.
DICEQuant: Distortion-Compensated Rounding with Dual-Ended Shrinkage for LLM Quantization
Yuan Cheng ⋅ Xing Hu ⋅ Zukang Xu ⋅ Hui Wang ⋅ Xiaomeng Han ⋅ Dawei Yang
Post-Training Quantization (PTQ) has become a prerequisite for efficient LLM deployment; however, current rotation-based methods are approaching a saturation point because they prioritize heuristic approximations over theoretical rigor. By treating quantization as an opaque black box, existing paradigms typically depend on indirect gradient estimators for activations while restricting weight optimization to local reconstruction objectives. Consequently, this structural limitation induces fundamental deficiencies: specifically, optimization instability and a "twofold blindness" toward accumulated upstream distortion and final task loss sensitivity. To overcome these barriers, we introduce $\textbf{DICEQuant}$, a unified framework established upon rigorous error modeling. First, we propose the $\textbf{CURE (Coupled Underflow-Rounding Error) surrogate}$, which facilitates exact and stable gradient computation for activation reshaping, thereby eliminating the variance inherent in heuristic estimators. Simultaneously, for weight optimization, we present $\textbf{Distortion-Compensated Rounding (DCR)}$. This mechanism derives a Shifted Optimal Center to neutralize input noise and utilizes Hessian-weighted shaping to align rounding decisions with the global functional objective. Empirical results demonstrate that DICEQuant significantly surpasses state-of-the-art baselines, successfully bridging the accuracy gap between low-bit compression and full-precision intelligence.
Human coordination often relies on the ability to influence the beliefs of others through strategic action. In multi-agent reinforcement learning, opponent shaping attempts to replicate this influence, though existing methods typically operate within an opponent's parameter, policy, or value space. Meanwhile, belief-manipulation techniques in hidden-role games often rely on hard-coded objectives, such as deception or belief saturation. We propose Differentiable Belief-based Opponent Shaping (D-BOS), a first-order method that treats each observer's belief as the shaped opponent state and differentiates through $k$-step softmax-Bayes belief dynamics. Rather than explicitly rewarding deceptive or cooperative behavior, our method treats the belief state as the target for shaping. This allows the optimal strategy to emerge naturally from the environment's reward structure. This belief-space formulation provides an opponent-shaping signal by differentiating through opponent belief updates, and naturally extends to multiple observers by aggregating gradients over their individual inferred belief trajectories. Empirically, D-BOS outperforms PPO and BBM in hidden-role games, with the largest gains in mixed-motive settings.
DiffRatio: Training One-Step Diffusion Models Without Teacher Supervision
Wenlin Chen ⋅ Mingtian Zhang ⋅ Jiajun He ⋅ Zijing Ou ⋅ José Miguel Hernández-Lobato ⋅ Bernhard Schölkopf ⋅ David Barber
Score-based distillation methods train one-step diffusion models in two stages: they first train a teacher score model, then distill it into a one-step student model. However, this distillation process can introduce bias from two sources: errors in the teacher score model and student score estimate. We propose DiffRatio, a new framework for training one-step diffusion models without teacher supervision. Instead of using a teacher score model to provide training targets, DiffRatio directly learns a log density ratio between the student and data distributions across diffusion time steps. As a pre-trained score model is only used to initialize the one-step generator and does not supervise training. This design simplifies the training pipeline, mitigates gradient estimation bias, and reduces the size of auxiliary networks. In addition, the learned density ratio can be used as a verifier, enabling a principled inference-time parallel scaling method that further improves sample quality without external rewards or extra sequential computation. DiffRatio achieves strong one-step generation results on CIFAR-10 and ImageNet ($64{\times}64$ and $512{\times}512$), outperforming most teacher-supervised distillation methods.
Diffusion LLMs are Natural Adversaries for any LLM
David Lüdke ⋅ Tom Wollschläger ⋅ Paul Ungermann ⋅ Stephan Günnemann ⋅ Leo Schwinn
We introduce a novel framework that transforms the resource-intensive (adversarial) prompt optimization problem into an efficient, amortized inference task. Our core insight is that pretrained, non-autoregressive generative LLMs, such as Diffusion LLMs, which model the joint distribution over prompt-response pairs, can serve as powerful surrogates for prompt search. This approach enables direct conditional generation of prompts, effectively replacing costly discrete optimization with a small number of parallelizable samples. We provide a probabilistic analysis demonstrating that under mild fidelity assumptions, only a few conditional samples are required to recover high-reward (harmful) prompts. Empirically, our method substantially outperforms prior attacks in both success rate and compute cost, producing low-perplexity, diverse jailbreaks that transfer to a wide range of black-box target models, including robustly trained and proprietary LLMs. Beyond adversarial prompting, our framework opens new directions for red teaming, automated prompt optimization, and leveraging emerging Flow- and Diffusion-based LLMs.
Diffusion Transformers with Residual Adaptive Layer Normalization
Ge Wu ⋅ Minxing Luo ⋅ Yikai Ge ⋅ Lei Wang ⋅ DanDan Zheng ⋅ Rui Liu ⋅ libin wang ⋅ Yu-Liang Zhan ⋅ Taihang Hu ⋅ Jingdong Chen ⋅ Xiang Li
Residual connections with PreNorm are the default design in modern transformers due to their strong optimization stability, but recent research has shown that they weaken the functional differentiation of deeper blocks and increase representational redundancy in later layers. This issue also appears in diffusion models using similar architectures, where the deep block is expected to support progressive denoising refinement rather than repeated computation over similar states. Motivated by this, we revisit cross-layer shortcuts in diffusion transformers and argue that their role should go beyond merely transferring shallow information, instead guiding deeper blocks to improve depth utilization. To this end, we propose \textit{\textbf{Res}idual Adaptive Layer \textbf{Norm}alization} (\textbf{ResNorm}), which routes residual cross-layer shortcuts into the existing adaLN modulation. ResNorm preserves the optimization advantages of PreNorm residual connections, while reusing the native adaLN interface to transmit early-layer representations as modulation signals. This design naturally combines cross-layer history with the model's native condition, such as timestep and class embeddings. ResNorm allows the earlier structural representation to adaptively scale, shift, and gate deeper blocks. This not only produces more differentiated deep representations and reduces cross-layer similarity in later blocks, but also accelerates training convergence and improves performance. On ImageNet 256$\times$256, SiT-XL/2 + ResNorm achieves up to $\textbf{2}\times$ and $\textbf{36}\times$ faster training than SiT-XL/2 + REG and SiT-XL/2 + REPA. Moreover, SiT-L/2 + ResNorm trained for only 400K iterations already outperforms SiT-XL/2 + REG.
DiPO: Disentangled Perplexity Policy Optimization for Fine-grained Exploration-Exploitation Trade-Off
Xiaofan Li ⋅ Ming Yang ⋅ Zhiyuan Ma ⋅ Shichao Ma ⋅ Jintao Du ⋅ Yu Cheng ⋅ Weiqiang Wang ⋅ zhizhong zhang ⋅ Xin Tan ⋅ Yanyun Qu ⋅ Lizhuang Ma ⋅ Yuan Xie
Reinforcement Learning with Verifiable Rewards (RLVR) has catalyzed significant advances in the reasoning capabilities of Large Language Models (LLMs). However, effectively managing the exploration and exploitation trade-off remains a critical challenge. In this paper, we fully analyze the exploration and exploitation dilemma of extremely hard and easy samples during the training and propose a new fine-grained trade-off mechanism. Concretely, we introduce a perplexity space disentangling strategy that divides the sample space into distinct exploration (high perplexity) and exploitation (low perplexity) subspaces, thereby mining fine-grained samples requiring exploration-exploitation trade-off. Subsequently, we propose a bidirectional reward allocation mechanism with a minimum impact on verification rewards to implement perplexity-guided exploration and exploitation, enabling more stable policy optimization. Finally, we have evaluated our method on two mainstream tasks: mathematical reasoning and function calling, and experimental results demonstrate the superiority of the proposed method, confirming its effectiveness in enhancing LLM performance by fine-grained exploration-exploitation trade-off.
Dirichlet-Guided Group Forecasting for Alleviating Over-smoothing in Time Series Forecasting
Xingyu Zhang ⋅ Jingyao Wang ⋅ Xin Yu ⋅ Zeen Song ⋅ Jianqi Zhang ⋅ Changwen Zheng ⋅ Wenwen Qiang
Time series forecasting often suffers from over-smoothing, especially when future dynamics are multi-modal. Forecasts may follow the coarse trend of the observed future, but fail to preserve sharp changes, oscillations, turning points, and regime transitions that define plausible dynamic evolution. In this work, we revisit over-smoothing from the perspective of latent dynamical mode compression: under partial observation and single-realization supervision, multiple plausible future modes can be weakened, merged, or averaged during forecasting. Based on this view, we propose Dirichlet-Guided Group Forecasting (DGF), a mode-preserving forecasting framework that explicitly models multiple mode-conditioned predictive distributions and uncertainty over their selection probabilities. DGF uses a Dirichlet-guided hierarchical sampling mechanism and reward-based optimization to encourage forecasts that are accurate, dynamically consistent, and mode-distinct. Extensive experiments on real-world forecasting benchmarks show that DGF reduces over-smoothing while improving forecasting accuracy, diversity, and dynamical consistency.
Disambiguating 2D-3D Correspondences in Gaussian Splatting-based Feature Fields for Visual Localization
Miso Lee ⋅ Sangeek Hyun ⋅ Yerim Jeon ⋅ Jae-Pil Heo
While Gaussian Splatting-based Feature Fields (GSFFs) have shown promise for visual localization, this paper highlights that photometrically optimized GSFFs are inherently ill-suited for 2D-3D matching. The volumetric extent of each Gaussian induces many-to-one pixel-to-point mappings that destabilize PnP-based pose estimation, while photometric optimization gives rise to superfluous Gaussians devoid of multi-view consistency. To address these issues, we propose SplitGS-Loc, a localization-specialized GSFFs construction framework that disambiguates 2D-3D correspondences by exploiting Gaussian attributes. Our key design, Mixture-of-Gaussians-based splitting, decomposes each Gaussian into smaller Gaussians, replacing ambiguous many-to-one with precise one-to-one correspondences. In parallel, we exploit composition weights from GS rasterization to select Gaussians that significantly and consistently contribute across multiple views and aggregate discriminative features through strong pixel-Gaussian associations, enforcing multi-view consistency. The resulting compact yet discriminative feature fields enable stable PnP convergence, achieving state-of-the-art performance on localization benchmarks. Extensive experiments validate that SplitGS-Loc extends the utility of photometric GSFFs to accurate and efficient localization by exploiting Gaussian attributes, without per-scene training or iterative pose refinement.
DisCoMBO: Steering Expert-in-the-Loop Black Box Optimization via Distributional Conformance
Jonas Seng ⋅ Bennet Wittelsbach ⋅ Kristian Kersting
Sequential Model-Based Optimization (SMBO) traditionally relies on Bayesian or ensembling surrogates for uncertainty quantification. While historically treated as fully data-driven, SMBO increasingly integrates external domain expertise to accelerate discovery. To overcome the opaque guidance and diminished integration fidelity of standard acquisition re-weighting, Probabilistic Circuits (PCs) have emerged as a generative surrogate alternative, enabling direct knowledge injection via conditional sampling. However, these generative routines lack the formal exploration-exploitation semantics required for rigorous optimization. We introduce the Distributional Conformance Score (DisCo), a principled metric that unifies the flexibility of PCs with a rigorous uncertainty framework. DisCo provides a bounded, $[0, 1]$-normalized measure of model "surprise" that (1) recovers properties comparable to kernel-based uncertainty known from, e.g., Gaussian Processes while maintaining linear-time inference, and (2) enables principled and accurate assessment of conformance of external knowledge w.r.t. model evidence. We then present DisCoMBO, a framework leveraging these properties for robust, knowledge-aware optimization. We prove that DisCoMBO is a zero-regret algorithm and demonstrate its effectiveness across diverse benchmarks from AutoML, material optimization, and wind park optimization.
Discovering What You Can Control: Interventional Boundary Discovery for Reinforcement Learning
Jiaxin Liu ⋅ Anzhe Cheng ⋅ Paul Bogdan
When an RL agent's observations contain distractors driven by the same confounders as its true state, observational data alone cannot identify which dimensions the agent controls. In our benchmarks, even state-conditioned observational selectors can collapse when distractors mimic controllable state variables. We propose Interventional Boundary Discovery (IBD), which treats the agent's own action channel as a source of randomized interventions: randomizing actions implements an interventional contrast, and per-dimension two-sample tests with FDR correction produce a binary mask over observation dimensions. Across 12 continuous-control settings with up to 100 distractors, IBD matches oracle return in 11 of 12 settings, while observational baselines including mutual information, state-conditioned forward models, and gradient-based sensitivity often underperform simply passing the full observation to SAC.
Discrete Diffusion Playground: A 2D Benchmark for Discrete Generative Models
Justin Shao ⋅ Sihyun Park ⋅ David Koes ⋅ Maria Chikina
Discrete diffusion is a promising framework for categorical generative modeling, but current evaluation protocols often fail to reveal whether models have actually learned the target distribution. We introduce the Discrete Diffusion Playground, a diagnostic benchmark that converts known 2D continuous distributions into binary token sequences through an exact discretized bijection. This creates a rare setting in which models are trained on discrete sequences, but generated samples can be decoded back into 2D and compared directly against a known ground-truth distribution. The benchmark spans smooth Gaussian mixtures and tilted checkerboards with sharp boundaries, disconnected support, and coordinate-dependent structure, allowing it to expose failures that are invisible in standard proxy metrics. It supports exact grid divergences, Wasserstein distances, density visualizations, and nonparametric two-sample tests, providing both quantitative evaluation and interpretable failure diagnosis. Using this testbed, we show that masked diffusion models can learn meaningful adaptive reveal orders but remain less stable than autoregressive baselines on smooth targets and largely fail on tilted checkerboards, where autoregressive models preserve the target geometry. Control experiments with reversed and interleaved token orderings show that autoregressive performance is not explained by favorable bit order alone.The Playground provides a simple but stringent benchmark for stress-testing discrete generative models and for developing evaluation metrics that detect real distributional mismatch rather than proxy performance. Our code is available at https://anonymous.4open.science/r/DiscreteDiffusionPlayground-0CCA/
Disentangled Multimodal Learning for Scalable Dynamic IR-Drop Analysis
Yihan Wen ⋅ Su Zheng ⋅ Tianshu Hou ⋅ Shiping Wu ⋅ Yuzhou Feng ⋅ Bei Yu
The robust delivery of power is a non-negotiable prerequisite for the operational reliability of modern chips. Even a single unanticipated IR-drop violation can cause system-level crashes, making IR-drop a strict criterion in chip signoff. However, current methodologies ranging from physics-based simulators to recent learning-based surrogates remain bottlenecked by an inherent coupling between spatial topology and temporal evolution. This paradigm is intrinsically unscalable when billion-gate spatial complexity converges with the extensive workloads of full-testbench verification. In this work, we reformulate dynamic IR-drop assessment as a disentangled multimodal learning problem and introduce Layout Waveform Fusor (LWFusor), which decouples invariant spatially static chip layouts from transient switching waveforms and fuses them through physics-aligned end-to-end learning. By projecting activity sequences directly into the spatial potential domain, the proposed approach departs from traditional time-stepping bottlenecks and makes long-horizon power integrity analysis computationally efficient. Experiments on industrial-scale designs demonstrate orders-of-magnitude acceleration while maintaining high fidelity in hotspot capture, highlighting a new trajectory for power integrity sign-off in large-scale EDA workflows.
Distillation of Foundation Models for Time-dependent PDEs
Daniel Musekamp ⋅ Boshra Ariguib ⋅ Andrei Manolache ⋅ Mathias Niepert
Foundation models for time-dependent partial differential equations (PDEs) are trained on large and diverse collections of physical systems and can generalize effectively to new downstream tasks. After fine-tuning on only a few trajectories from a target domain, they often achieve strong accuracy in low-data regimes. However, these models are typically large and computationally intensive, limiting their usefulness as fast surrogates for numerical solvers. We propose Teacher Rollout Extension (TREX), a knowledge distillation framework that transfers the predictive capability of a pretrained foundation model into a compact and efficient student. Starting from a fine-tuned teacher, TREX augments limited downstream data by generating long synthetic trajectories through teacher rollouts with periodic noise injection. This procedure samples from the teacher-induced rollout distribution without requiring explicit knowledge of the initial-condition distribution, while exposing the student to long-horizon states and local recovery behavior around states encountered during autoregressive prediction. The student can further incorporate task-specific inductive biases, such as equivariance, that the teacher does not necessarily enforce. We evaluate TREX on multiple PDE benchmarks. The resulting students match or even surpass the teacher’s accuracy while reducing the number of parameters by several orders of magnitude and achieving more than an order-of-magnitude speedup in inference.
Distilling Graph Geometry: Knowledge Gap from GNNs to MLPs
Zhewei Chen ⋅ Hao Zhu ⋅ Jiaojiao Jiang ⋅ Ahad N. Zehmakan
GNN-to-MLP distillation aims to retain the predictive accuracy of a message-passing teacher while deploying a graph-free MLP at inference. Existing methods mainly transfer node-wise predictions or use confidence-based reweighting, but they do not specify where the student should preserve the teacher's graph-induced geometry. We show that this omission leads to two spectral failure modes in the student's representation space. On sparse graphs, the student suffers from *spectral underfit*, missing high-energy teacher directions concentrated near boundary regions. On dense graphs, it suffers from *spectral overfit*, retaining spurious directions that the teacher has collapsed through aggregation. Motivated by an energy-weighted teacher--student alignment objective, we propose **Graph Geometry-aware MLP (G$^2$MLP)**, a training-time distillation framework guided by Ollivier--Ricci curvature. Curvature identifies where the two spectral errors concentrate and is used to allocate supervision between prediction-level and representation-level alignment. The deployed model remains a standard MLP and requires no graph access at inference. Across node-classification benchmarks, G$^2$MLP consistently improves over graph-free distillation baselines, reduces the teacher--student rank gap in both regimes, and transfers without architectural changes to Graph Transformer teachers and link prediction.
Distill to Think, Foresee to Act: Cognitive-Physical Reinforcement Learning for Autonomous Driving
Yang Wu ⋅ Qiang Meng ⋅ Zhaojiang Liu ⋅ Youquan Liu ⋅ Jian Yang ⋅ Jin Xie
Current end-to-end autonomous driving models are fundamentally constrained by the behavioral cloning ceiling of imitation learning. While reinforcement learning offers a path to smarter autonomy, it demands two missing pieces of infrastructure: (1) a cognitive foundation that understands traffic semantics and driving intent, and (2) a foresighted physical environment that can anticipate the consequences of candidate actions. To this end, we propose CoPhy, a Cognitive-Physical reinforcement learning framework for autonomous driving. To distill to think, we distill VLM knowledge into the BEV encoder and then discard the VLM entirely, retaining cognitive ability at zero inference cost while releasing the cognitive channel as a pluggable interface for optional human language commands. To foresee to act, we build an auto-regressive BEV world model that explicitly predicts future semantic maps conditioned on candidate actions, serving as an interpretable physical sandbox from which safety metrics are directly derived. Built upon this dual infrastructure, we optimize the driving policy via GRPO with a novel dual-reward mechanism: a physical reward derived from BEV rollouts enforces hard safety constraints, while a cognitive reward from a language-aligned scorer ensures intent compliance. Extensive experiments demonstrate that CoPhy not only achieves state-of-the-art results on NAVSIM v1 and v2 benchmarks, but also enables safer driving via cognitively informed scene compliance and flexible intent control through user-defined language instructions.
Distributed-Order Fractional Spiking Neural Network
Chengjie Ge ⋅ Yufeng Peng ⋅ Qiyu Kang ⋅ Xueyang Fu ⋅ Zheng-Jun Zha
Most existing Spiking Neural Networks (SNNs) are based on first-order ordinary differential equations (ODEs) to model neuron potential dynamics. Recently, fractional-order spiking neural networks (f-SNNs) extend conventional SNNs by incorporating fractional-order ODEs, enabling the modeling of long-range temporal dependencies in neuron potential states. However, existing first-order SNNs and f-SNNs are typically restricted to a single derivative order, which limits their capacity to represent heterogeneous temporal dynamics across different time scales. In this work, we propose Distributed-Order Fractional Spiking Neural Networks (DF-SNNs), which generalize single-order neuron dynamics by integrating multiple fractional orders through learnable distributed weights. This formulation enables neurons to accumulate historical information across multiple temporal scales within a unified framework. We develop practical numerical implementations based on the Grünwald–Letnikov discretization, allowing efficient training of DF-SNNs within standard SNN architectures. We show that the proposed distributed-order formulation induces intrinsic multi-scale temporal memory, providing a richer class of temporal representations compared to single-order fractional dynamics. Extensive experiments demonstrate consistent performance gains of DF-SNNs over conventional SNNs and single-order f-SNNs.
Distributional Alignment as a Criterion for Designing Task Vectors in In-Context Learning
Jihoon Kwon ⋅ Jiwon Choi ⋅ Jy-yong Sohn
In-context learning (ICL) allows large language models (LLMs) to adapt to new tasks through demonstrations, yet it suffers from escalating inference costs as context length increases. While task vectors offer a promising alternative by compressing demonstrations into compact hidden-state representations, their quality has been evaluated only through downstream task accuracy. This indirect criterion provides limited insight into how to design more effective task vector extraction methods. In this paper, we posit that inference using task vectors should align their predictive distribution with that of ICL. To quantify this, we introduce $d_{\text{NTP}}$, a metric that measures the discrepancy in next-token probabilities between task vector-based and ICL-based inference. Our empirical analysis reveals that $d_{\text{NTP}}$ serves as a performance proxy, exhibiting a strong negative correlation with downstream accuracy. Motivated by this, we develop Linear Task Vector (LTV), a method designed to minimize $d_{\text{NTP}}$ via a closed-form linear mapping that estimates demonstration effects through regression. Across eight classification benchmarks and five LLMs, LTV consistently outperforms existing task vector baselines, improving average accuracy by 9.2\% while reducing inference latency. We further show that LTV outperforms the baselines on regression tasks. Moreover, we investigate the transferability of LTV across different model scales; an aspect that has remained nascent in task vector research. Specifically, we empirically show that task vectors from a larger model can enhance a smaller model's performance by 6.4\%, suggesting a new utility for extracted task representations.
Mixture-of-Experts (MoE) transformers scale capacity by activating only a few experts per token, but this sparsity creates a hidden reliability problem: when routing is imperfect, load-balanced models may send tokens to experts that are insufficiently trained for the assigned inputs. We propose Distributionally Robust MoE Training (DRMoET), a drop-in objective that treats layer-wise experts as endogenous robustness groups and optimizes high-loss routing outcomes rather than merely equalizing traffic. DRMoET updates a per-layer expert distribution by an entropy-regularized softmax rule on EMA-smoothed, activation-weighted expert losses, strengthening plausible non-top routing paths while preserving standard MoE computation. Under the FLAME-MoE recipe at 746M-total and 10.3B-total scales, DRMoET improves downstream averages over both standard FLAME-MoE and auxiliary-loss-free balancing. Mechanistic analyses show lower expert-loss variance with nearly unchanged mean loss, 4.3\% lower degradation under forced mid-$k$ misrouting, and improved domain--expert specialization. These results position routing robustness--not only utilization balance--as a practical objective for reliable sparse MoE scaling.
dMoE: dLLMs with Learnable Block Experts
Sicheng Feng ⋅ Zigeng Chen ⋅ Gongfan Fang ⋅ Xinyin Ma ⋅ Xinchao Wang
Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to autoregressive models, offering competitive performance while naturally supporting parallel decoding. However, as dLLMs are increasingly integrated with Mixture-of-Experts (MoE) architectures to scale model capacity, a fundamental mismatch arises between block parallel decoding and token-level expert selection. Specifically, each dLLM forward pass processes multiple tokens with bidirectional dependencies, whereas conventional MoE layers route each token independently. This mismatch substantially increases the number of uniquely activated experts, making inference increasingly memory-bound. To address this, we propose dMoE, a simple yet effective block-level MoE framework. The central idea of dMoE is to aggregate token-level expert distributions within each block into a unified block-level expert distribution, which is then used to guide expert routing in a more coherent manner. In this way, dMoE substantially reduces the number of uniquely activated experts during inference without sacrificing performance, thereby mitigating the memory-bound bottleneck. Extensive experiments across a variety of benchmarks demonstrate the effectiveness of dMoE. On average, dMoE reduces the number of uniquely activated experts from 69.5 to 14.6 while retaining 99.11\% of the original performance. Meanwhile, it reduces memory usage by 76.64\% to 79.84\% and achieves 1.14$\times$ to 1.66$\times$ end-to-end latency speedup.
Dodge-It: Learning Collision-Aware VLA Models for Robotic Manipulation
Jiawei Feng ⋅ Chuzhao Huang ⋅ Xinli Xu ⋅ Yingcong Chen
Vision-Language-Action (VLA) models are increasingly used as general-purpose policies for robotic manipulation, but completing the instructed task requires more than reaching the target state. A policy must also avoid nearby objects throughout the rollout, including approach, interaction, transport, and handover. A common approach is to invoke conventional collision avoidance mechanisms, such as motion planners or safety filters, which provide effective geometric safeguards. However, the resulting system may struggle to produce actions that are simultaneously safe and task progressing. When collision avoidance and manipulation progress require misaligned actions, the system may avoid the obstacle while slowing, disrupting, or even failing the task. To address this, we introduce DODGE-IT, a collision-aware learning framework that leverages collision signals as policy-improvement feedback to strengthen the VLA policy's collision-aware manipulation capability. DODGE-IT first designs Task-Relevant Assessment of Collision Events (TRACE) that detects and identifies task-relevant obstacles from RGB-D observations and turns executions into collision outcomes. These outcomes are then incorporated together with task success signals into representative reinforcement learning frameworks for VLA. We validate DODGE-IT on SafeLIBERO and real robot manipulation tasks, showing improved obstacle avoidance while preserving task completion across simulated and physical settings.
Do Diffusion Models Learn to Generalize Basic Visual Skills?
Amish Sethi ⋅ Boya Zeng ⋅ Wenhao Chai ⋅ Zhuang Liu
Diffusion-based visual generative models are powerful, but understanding their generalization behavior remains challenging. Current evaluations rely on human preference scores for text-to-image models and distribution-level metrics such as Frechet Inception Distance (FID) and Inception Score (IS) for class-conditional models, but these do not test whether models learn basic visual skills or largely match training patterns. To overcome this gap, we introduce the Visual Generative Lab (VGL), a controlled experimental framework for understanding how diffusion models learn and generalize basic visual skills, including size, position, rotation, count, color, shape, and their composition. By generating synthetic training data for each skill from explicit rules, VGL enables precise measurement of generalization using task-specific, rule-based metrics. Training models from scratch on these data reveals three consistent patterns. (1) Models extrapolate rotation better than size, skill, or count on far-out-of-range queries; healthy rotation seeds match most extrapolation angles within a few degrees, although a non-trivial fraction of seeds collapse near the wraparound, while skill and position saturate at small distances outside training regardless of seed. (2) For compositional generalization, coverage of basic visual skill combinations matters more than dataset size. (3) Basic visual skills are learned jointly, so an out-of-distribution request in one skill can degrade the others even when they remain in range.
Does Cross-Panel Pretraining Transfer Under Marker Heterogeneity? A Large-Scale Empirical Study in Clinical Flow Cytometry
Zixin Zhuang ⋅ Benjamin Mashford ⋅ Dan Andrews
Flow cytometry is the primary assay for single-cell immune profiling, yet cross-study modelling remains under-explored. The challenge is two-fold: Studies measure different marker panels, inducing feature-space inconsistency, while batch effects shift signals even on shared markers, making transfer harder than in vision or language where input spaces are stable. We conduct a large-scale empirical study, pretraining an encoder on 89,324 samples from 112 public studies spanning 525 distinct marker panels. We evaluate it frozen --- isolating representation quality from task-specific adaptation, on five clinical tasks under a taxonomy of panel-overlap conditions, including three panel-zero-shot tasks whose marker panels were unseen during pretraining, and one negative control with no expected discriminative signal. We further characterise transfer across scaling axes and probe robustness to marker removal at test time. Despite heterogeneity that disrupts direct transfer, the frozen encoder matches or exceeds in-distribution specialist encoders on 4 of 5 tasks, including all three novel-panel tasks, and behaves as expected on the negative control. Scaling analyses reveal saturation in both encoder capacity and random corpus subsampling, and representations degrade gracefully under marker reduction. These results suggest a pretrained encoder can serve as a general-purpose feature extractor for clinical cytometry data, without task-specific retraining.
Do Less, Decide Better: Optimal Human Dispatching in AI-Assisted Decisions
Lezhi Tan ⋅ Naomi Sagan ⋅ Lihua Lei ⋅ Jose Blanchet
AI systems increasingly assist decision making by producing cheap factor-level assessments of complex inputs, but these assessments are often biased and incomplete. We study selective human oversight: given AI signals, which factors should be escalated to a costly human evaluator? We formulate this human--AI collaboration problem as a factor-level information-acquisition problem. Under squared loss, solving for optimal dispatching reduces to maximizing a contextual reward; under a linear model, this reward admits a closed-form decomposition into two interpretable terms: predictive relevance and residual uncertainty in the human evaluation after conditioning on the AI signal. We instantiate the framework on peer review, decomposing papers into ten aspects and evaluating 3,408 ICLR submissions across three LLMs and multiple regression heads. Outputs based on the optimal dispatching rules match full-human-review performance using only 2--3 queried aspects out of 10. The framework remains effective when the human signal comes from a single noisy reviewer and across review years, suggesting robustness to reviewer inconsistency and and temporal distribution shift. These results position principled dispatching as a practical foundation for scalable human--AI decision systems.
Do LLMs Feel Social Pressure? Locating and Steering Social Desirability Bias in LLMs
Yi Feng ⋅ Jiaqi Wang ⋅ Wenxuan Zhang
Humans often say what sounds acceptable rather than what they truly think when facing social pressure, a well-documented phenomenon called \textbf{Social Desirability Bias (SDB)}. Trained on human-generated text, LLMs are steeped in the same social fabric that produces these pressures, yet whether and how SDB shapes their responses remains unknown. We present \textbf{PRESS} (\textbf{P}sychologically-grounded \textbf{R}epresentation \textbf{E}xtraction for \textbf{S}ocial-desirability \textbf{S}teering), the first framework to identify, localise, and steer SDB inside open-weight LLMs. We decompose SDB into four observable and relatively orthogonal mechanisms: \emph{Audience}, \emph{Accountability}, \emph{Self-Monitoring}, and \emph{Social Norms}, and operationalise each as paired High/Low pressure conditions across 9 cultures in cultural values where social pressure bites hardest. Our experiments show that SDB forms a decodable and low-dimensional direction inside LLMs, organised by specific psychological mechanism rather than cultural origin. Steering along it improves all 58 downstream cultural content moderation tasks across three model families (Qwen $+5.7\%$,Gemma $+7.9\%$, Llama $+16.3\%$ mean F1 across cultures), and a direction extracted from one culture can transfer to others without re-extraction. Ablation confirms the direction is causally necessary, and evaluation on capability benchmarks further shows PRESS shifts cultural stance without disrupting factual knowledge.We hope this work makes social pressure in LLMs more transparent and opens new directions toward more honest and socially aware language models.
Domain-Conditioned Class Imbalance: Why Global Class Balance Fails Across Domains
Yongkun Deng ⋅ NAN YANG ⋅ Xiatong Guo ⋅ Zhiyong Wang ⋅ Dong Yuan
Modern visual recognition systems increasingly rely on unlabeled images collected from diverse visual domains. A common goal in dataset construction is to avoid obvious marginal imbalance, such as skewed class totals or uneven domain sizes. However, marginal balance can be misleading in multi-domain data: even when per-class totals and per-domain sample sizes are balanced, sample allocations can be highly skewed across domain-class pairs. We call this hidden conditional skew Domain-Conditioned Class Imbalance (DCI), where class imbalance is revealed only after conditioning on domain context. For controlled evaluation, we fix class and domain marginals and vary only the joint domain--class allocation, making the effect of domain-conditioned imbalance directly measurable. To address this challenge, we propose HyperIDAC, a hypernetwork-based classifier for multi-domain imbalance. HyperIDAC infers domain priors from visual features and lightweight zero-shot predictions, then dynamically generates instance-specific classifier parameters and domain-conditioned class embeddings. This design yields domain-sensitive decision boundaries while leveraging CLIP semantics to support underrepresented domain--class pairs. Furthermore, a multi-level pseudo-labeling module combines teacher–student training with historical label tracking and a fallback for low-confidence samples, mitigating bias reinforcement in unlabeled settings. Extensive experiments on four multi-domain benchmarks under three regimes show that HyperIDAC substantially outperforms state-of-the-art methods, highlighting the value of domain-aware hypernetwork-based adaptation under DCI.
Don't Lose Focus: Activation Steering via Key-Orthogonal Projections
Haoyan Luo ⋅ Mateo Espinosa Zarlenga ⋅ Mateja Jamnik
Activation steering controls LLM behaviour towards target behaviour by intervening in internal representations, yet it often degrades reasoning and retrieval performance. We argue that a primary cause of this trade-off is attention rerouting: steering vectors alter query-key matching, shifting attention away from contextually important tokens toward less informative ones. To address this, we propose Steering via Key-Orthogonal Projections (SKOP), a steering method that constrains harmful attention rerouting without eliminating steering efficacy. SKOP achieves this by preserving attention patterns on a small set of focus tokens the model relies on for reasoning and retrieval, while allowing redistribution among less critical tail tokens. Across multiple steering benchmarks, we show that SKOP achieves the best joint steering-utility trade-off, reducing utility degradation by 5–7$\times$ while retaining over 95\% of vanilla steering efficacy. Our results further suggest that, in long-context retrieval settings where vanilla steering approaches are ineffective, SKOP can maintain robust performance by avoiding attention rerouting.
Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation
Joon Ha Kim ⋅ Geon-Woo Kim ⋅ Anoop Rachakonda ⋅ Daehyeok Kim
Selecting the optimal LLM inference configuration requires evaluation across hardware, serving engines, attention backends, and model architectures, since no single choice performs best across all workloads. Profile-based simulators are the standard tool, yet they hardcode their operation set to a specific configuration and re-profile every operation from scratch, making exploration prohibitively expensive. This cost stems from a missing structural understanding: every input dimension of each operation is fixed by the model configuration or determined by the incoming request. Many model-configuration values (e.g., head size, layer count) recur across models, so the same operation runs in many configurations; a single sweep over the request-dependent dimensions can serve them all. We present Dooly, which exploits this structure to achieve configuration-agnostic, redundancy-aware profiling. Dooly performs a single inference pass, labels each input dimension with its origin via taint propagation, and selectively profiles only operations absent from its latency database; stateful operations such as attention are isolated by reusing the serving engine's own initialization code, eliminating manual instrumentation. It builds latency regression models based on the database, which becomes a drop-in backend for existing simulators. Across two GPU platforms, three attention backends, and diverse model architectures, Dooly achieves simulation accuracy within 5% MAPE for TTFT and 8% for TPOT while reducing profiling GPU-hours by 56.4% across 12 models compared to the existing profiling approach.
Do We Need Asynchronous SGD? On the Near-Optimality of Synchronous Methods
Grigory Begunov ⋅ Alexander Tyurin
Modern distributed optimization methods mostly rely on traditional synchronous approaches, despite substantial recent progress in asynchronous optimization. We revisit Synchronous SGD and its robust variant, called $m$-Synchronous SGD, which can be interpreted as a method with a backup-worker strategy, and theoretically show that they are nearly optimal in many heterogeneous computation and communication scenarios, which is somewhat unexpected. We analyze the synchronous methods under random times and adversarial partial participation of workers, and prove that their time complexities are optimal in many practical regimes, up to logarithmic factors. While synchronous methods are not universal solutions and there exist tasks where asynchronous methods may be necessary, we show that they are sufficient for many modern heterogeneous computation and communication scenarios.
Do We Really Need Diffusion for Generative Object Detection? A Minimal Prototype Perspective
Yu Hong ⋅ Xiaosong Jia ⋅ Yihan Wang ⋅ Wenlong Liao ⋅ Tao He ⋅ Junchi Yan
Generative object detection, introduced by DiffusionDet, is defined by two coupled properties: (P1) data-independent random-box initialization with a learned noise-to-data transport, and (P2) test-time scaling (TTS), the ability to trade inference compute for accuracy at a fixed checkpoint by adjusting the number of proposals or refinement steps. This paper asks: \emph{How can we preserve both properties with a minimal conceptual design?} We dissect DiffusionDet component by component and find that, in our controlled interventions, the diffusion schedule, the initial noise distribution, and DDIM sampling stochasticity have only a small effect on inference results. In contrast, detection performance is more sensitive to the denoising head, iterative refinement, and proposal coverage. Based on these diagnostic results, we propose \textbf{LinearDet}, a compact generative detection prototype: \emph{linear noising at training, linear denoising at inference}---no diffusion-specific schedule, no diffusion ODE/SDE solver, only a symmetric pair of linear mixtures. On COCO, LinearDet matches DiffusionDet under standard settings, is more robust under sparse proposal coverage, outperforms it in zero-shot CrowdHuman transfer, and produces smoother refinement trajectories that admit clean heatmap visualization. More broadly, we hope this diagnosis and compact prototype offer the object detection community a cleaner baseline for studying generative detection.
Drag as Evidence: Motion-Grounded Latent Recomposition for Drag-Based Editing
Xinyu Pu ⋅ Hongsong Wang ⋅ Jie Gui ⋅ Pan Zhou
Modern image editors excel at semantic manipulation and visual synthesis, yet remain limited in precise spatial control, motivating the development of drag-based editing. However, existing drag-based methods often struggle to balance drag accuracy with natural, plausible, and intent-aligned generation. We propose MoRe-Drag, a motion-grounded drag-based editing method. Our key insight is to treat pixel-space warping as coarse motion evidence, and to inject this evidence into the generative sampling trajectory. Specifically, MoRe-Drag performs region-aware latent recomposition over refinement, inpainting, and anchor regions, coupled with stage-adaptive conditioning that progressively shifts from motion-grounded structure formation to semantic refinement. We further support an instruction-free interface by adapting the MLLM-based text encoder for drag-aware instruction inference. Experiments on DragBench-SR and DragBench-DR show that MoRe-Drag substantially improves drag precision over strong base editors and achieves superior drag accuracy among SOTA drag-based methods, while delivering strong semantic consistency and visually realistic results. Code and dataset will be publicly released.
DRAGON : A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams
Anirudh Iyengar Kaniyar Narayana Iyengar ⋅ Tampu R Kumar ⋅ Gaurav Najpande ⋅ Manan Suri ⋅ Dinesh Manocha ⋅ Puneet Mathur ⋅ Vivek Gupta
Diagram question answering (DQA) requires models to interpret structured visual representations such as charts, maps, infographics, circuit schematics, and scientific diagrams. Recent vision-language models (VLMs) often achieve high answer accuracy on these tasks, yet correct answers do not guarantee that models ground their reasoning in the diagram regions that support the prediction. Models may instead rely on textual correlations or dataset artifacts without identifying the visual evidence required to verify the answer. This limitation prevents reliable evaluation of diagram reasoning and reduces interpretability. We introduce DRAGON, a benchmark for evaluating evidence-grounded visual reasoning in diagrams. Given a diagram, a question, and the correct answer, a model must predict bounding boxes that correspond to the visual elements required to justify the answer. These evidence regions may include answer-bearing components, textual labels, legends, axes, connectors, and other supporting structures involved in the reasoning process. The DRAGON dataset contains 11,664 annotated question instances from six diagram QA datasets: ChartQA, Circuit-VQA, InfographicsVQA, MapIQ, MapWise, and AI2D, with a 2,445-instance test set carrying human-verified evidence annotations and a standardized evaluation framework. Our evaluation on recent VLMs reveals that even Claude Opus 4.6 achieves a Grounding IoU of only 23.8% on AI2D, showing that faithful visual grounding remains an open challenge. DRAGON supports future research on models that ground their predictions in visual evidence.
Dream-Cubed: Controllable Generative Modeling in Minecraft by Training on Billions of Cubes
Tim Merino ⋅ Sam Earle ⋅ Ryunosuke Iwai ⋅ Julian Togelius ⋅ Edoardo Cetin
We introduce Dream-Cubed, a large-scale dataset of Minecraft worlds at voxel resolution, and a family of models using cubes as powerful compositional units for efficient generation of interactive 3D environments. Dream-Cubed comprises tens of billions of tokens from a carefully curated mixture of procedural biome terrain and high-quality human-authored maps. We use this dataset to conduct the first large-scale study of 3D diffusion models for voxel generation, analyzing discrete and continuous diffusion formulations, data compositions, and architectural design choices. Our models operate directly in the space of blocks, enabling efficient and semantically grounded generation while supporting interactive user workflows such as inpainting and outpainting from user-authored blocks. To quantitatively evaluate our models, we adapt the FID metric to assess semantic differences between real and generated world renderings, and validate generation quality through a human preference study. We release the full dataset, code, and all our pretrained models, which we hope will provide a foundation for future research in efficient generative modeling for structured, interactive 3D environments.
Drifting Field Policy: Wasserstein Gradient Flow on Policy Space for Offline-to-Online RL
Juil Koo ⋅ Mingue Park ⋅ Jiwon Choi ⋅ Yunhong Min ⋅ Minhyuk Sung
We propose Drifting Field Policy (DFP), a non-ODE one-step generative policy built on the drifting model paradigm. Under online RL finetuning, we frame the policy update as a reverse-KL Wasserstein-2 gradient flow toward a soft target policy, realized by DFP as a steepest-descent step on probability space. By construction, this gradient step is decomposed into an ascent toward higher action-value regions and a score matching with the anchor policy as a trust region. We further derive a simple, tractable surrogate of the otherwise intractable update loss, akin to behavior cloning on top-K critic-selected actions. We find empirically that this mechanism uniquely benefits the drifting backbone owing to its non-ODE parameterization. With one-step inference, DFP achieves state-of-the-art performance on several manipulation tasks across Robomimic and OGBench, outperforming ODE-based policies.
Drift Q-Learning
Achraf Anas El Houssaini ⋅ Mohamad Hosein Danesh ⋅ Amin Abyaneh ⋅ Scott Fujimoto ⋅ Hsiu-Chin Lin ⋅ David Meger
Offline reinforcement learning requires improving a policy from fixed data while avoiding out-of-distribution actions with unreliable value estimates. Diffusion and flow policies handle this trade-off by modeling the behavior distribution to regularize the RL objective, but they require iterative denoising, solver integrations, and in more efficient variants, distillation or other approximations at inference. We propose DriftQL, which combines a drift-based behavioral regularizer with critic-driven policy improvement. The value signal biases the policy toward high-value regions of the data support, while attraction and repulsion together keep generated actions near the data and prevent collapse onto a single mode. DriftQL is implemented as a single network with a unified training objective and generates actions in a single forward pass. On D4RL and OGBench, DriftQL consistently outperforms diffusion and flow methods, advancing the state of the art. Under degraded data quality, where the baselines visibly struggle, DriftQL remains close to its clean-data performance, positioning it as a promising alternative to diffusion and flow-based methods while maintaining the simplicity and efficiency of deterministic approaches.
DRIVE: Fine-tuning via Data Contribution- and Diversity-aware Weighting with Prior Regularization
qing liu ⋅ Xinrui Chen ⋅ Weiyao Zhu ⋅ Yi Du ⋅ Ou Wu
Single-stage fine-tuning is pervasive, yet catastrophic forgetting persists when the source pre-training data is inaccessible. Data weighting affords a natural lever to modulate sample influence. Yet existing weighting methods lack a principled mechanism to explicitly arbitrate knowledge retention against new-task adaptation, and provide no unified differentiable objective to accommodate auxiliary signals such as sample diversity and sample characteristics. This is because current schemes privilege one goal over the other, and no explicit gradient-compatible encoding exists for such signals. To address this, we propose $\textbf{DRIVE}$, a principled and efficient weight assignment framework. It derives per-sample weights via a unified differentiable objective where three coupled signals co-evolve: Shapley additivity splits contribution into anti-forgetting and adaptation components, which are reshaped by a covariance-based diversity term and stabilized by a sample-characteristic prior. Far from a mere combination, this yields an end-to-end gradient-based scheme balancing stability against plasticity. Theoretically, we prove that DRIVE enjoys superior stability-plasticity guarantees over existing weighting schemes and enhances parameter estimation robustness. Empirically, it attains superior overall performance, faster convergence, and stronger robustness. Code at https://anonymous.4open.science/r/DRIVE-158.
DriveFuture: Future-Aware Latent World Models for Autonomous Driving
Yufeng hong ⋅ Xiaotian Zhou ⋅ Yingyan Li ⋅ Xiangpo Zhou ⋅ Lin Liu ⋅ Yadan Luo ⋅ Shaoqing Xu ⋅ Lei Yang ⋅ Ziying Song
Existing latent world models for autonomous driving have opened a promising path toward future-aware driving intelligence. However, they typically treat future latent states as prediction targets or auxiliary signals, rather than directly conditioning trajectory planning. This can entangle current and future features in latent space. In this work, we propose DriveFuture, a future-aware latent world modeling framework for autonomous driving that explicitly learns planning-oriented foresight by conditioning the current latent state modeling process on future world states. Specifically, during training, the model first predicts future latent world states from the current latent state and ego action, and then refines the prediction against the ground-truth future latent state via cross-attention. The resulting future-aware latent serves as an explicit condition for a diffusion-based trajectory planner. During inference, DriveFuture conditions on the predicted future latent state instead of the ground-truth future state. DriveFuture achieves SOTA performance on the public NAVSIM benchmarks, reaching 55.5 EPDMS on NAVSIM-v2 navhard, 89.9 EPDMS on NAVSIM-v2 navtest, and 90.7 PDMS on NAVSIM-v1 navtest, respectively. These results suggest that the key to latent world modeling lies not merely in simulating future states, but more importantly in conditioning current decision-making on future states. Notably, as of April 2026, DriveFuture ranks 1st on the NAVSIM-v2 navhard leaderboard and achieves SOTA performance on NAVSIM-v1 navtest.
Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems
Lang Feng ⋅ Longtao Zheng ⋅ Shuo He ⋅ Fuxiang Zhang ⋅ Bo An
Multi-agent LLM systems improve reasoning and tool use via role specialization, yet reinforcement learning (RL) post-training for such systems remains unstable and underexplored. We theoretically pinpoint a key source of instability when extending group-based RL to cooperative multi-agent LLM systems: under GRPO-style optimization, a global normalization baseline can mismatch heterogeneous agents' reward distributions, inducing gradient-norm instability. Based on this finding, we propose Dr. MAS, a simple and stable RL recipe for cooperative multi-agent LLM systems. Algorithmically, Dr. MAS normalizes advantages per agent using each agent's own reward statistics, which calibrates gradient scales and dramatically stabilizes the training. Systemically, Dr. MAS provides an end-to-end multi-agent RL framework with scalable orchestration, flexible per-agent LLM serving and optimization, and shared resource scheduling of actor backends. Unlike single-actor frameworks such as veRL, Dr. MAS natively supports multiple heterogeneous LLMs over a unified GPU pool. Lifecycle-aware backend management and dynamic dispatch release inactive models' GPU memory and schedule active agents on demand, enabling hardware-efficient co-training. We evaluate Dr. MAS on multi-agent math reasoning and multi-turn search with Qwen2.5 and Qwen3. Dr. MAS achieves clear gains over vanilla GRPO (e.g., +5.6\% avg@16 and +4.6\% pass@16 on math, and +15.2\% avg@16 and +13.1\% pass@16 on search) while largely eliminating gradient spikes. It also remains highly effective under heterogeneous agent-model assignments while improving efficiency. Code is available at https://anonymous.4open.science/r/DrMAS-0787.
DrugSAGE: Self-evolving Agent Experience for Efficient State-of-the-Art Drug Discovery
Yikun Zhang ⋅ Xiwei Cheng ⋅ Tianyu Liu ⋅ Yuanqi Du ⋅ Wengong Jin
Building state-of-the-art (SOTA) predictive models for drug discovery requires expensive search over tools, architectures, and training strategies. Current LLM-based agents can find SOTA solutions through extensive trial and error, but they do not retain the experience accumulated along the way and therefore pay the full search cost on every new task. We propose DrugSAGE (Self-evolving Agent Experience), a framework that accumulates and reuses experience across tasks to build SOTA drug discovery models efficiently. DrugSAGE maintains a cross-task memory of verified skills, statistical evidence about effective strategies, and a record of recurring errors and their fixes. In some cases, DrugSAGE transfers a working solution directly without test-time search. In 33 molecular property prediction tasks, DrugSAGE ranks first among nine SOTA agents in a single-task setting. With memory accumulated from 16 smaller tasks, DrugSAGE achieves a averaged normalized score of 0.935 on 17 held-out tasks in a cross-task evaluation setting and outperforms all baseline agents by 10-30\% in a zero-test-time search regime. In summary, our work shows the advantage of cross-task memory for efficient SOTA model development in drug discovery.
DSBTR: A Diffusion Schrödinger Bridge Trajectory Refiner for Multi-Agent Trajectory Prediction
Yuxi Xue ⋅ Xiangzheng Zhou ⋅ Xiaobo Chen ⋅ Jianjun Qian ⋅ Jian Yang
Multi-agent trajectory prediction is crucial for ensuring autonomous driving safety. Although diffusion models show great potential in trajectory refinement, existing diffusion-based refiners suffer from severe exposure bias. This bias stems from the distribution mismatch between the noisy training endpoint and the noise-free inference starting point, which frequently causes negative optimization during the refinement process. To resolve this issue, this paper proposes the Diffusion Schrödinger Bridge Trajectory Refiner (DSBTR). By introducing a tractable Schrödinger bridge with dual-terminal constraints, DSBTR ensures that the forward evolution endpoint strictly degenerates into a deterministic and noise-free coarse prediction prior. This theoretical alignment eliminates the initial boundary error and constrains the reverse solving process to a strict residual approximation. Based on the semi-linear structure of DSBTR, we derive a tailored high-order DPM-Solver to accelerate the sampling process and meet autonomous driving application requirements. Extensive experiments on the ETH/UCY and SDD benchmarks demonstrate that DSBTR serves as an effective plug-and-play module. It consistently improves the prediction accuracy of various baseline architectures and successfully avoids negative optimization.
DSSA: Dynamic Sparse Semantic Anchoring for Purifying Protective Perturbations against Diffusion Models
Weiwei Tan ⋅ Rui Wang ⋅ Lihua Jing ⋅ Yanjun Zhang ⋅ Runbo Li ⋅ Zhishen Wang ⋅ Leo Yu Zhang
The rapid proliferation of personalization techniques in Stable Diffusion has driven the widespread deployment of adversarial perturbations as a protective measure against unauthorized data mimicry. Existing diffusion-based purification methods primarily rely on the noising and denoising paradigm and have proven somewhat effective in neutralizing these protections. However, they suffer from a severe trade-off between purification efficacy and perceptual fidelity, especially under large perturbation budgets. To mitigate this problem, in this paper, we propose Dynamic Sparse Semantic Anchoring (DSSA), a novel purification framework. By starting the reverse denoising process from pure Gaussian noise, DSSA removes the initial adversarial perturbations and circumvents the dilemma of selecting an optimal noising level. To prevent the loss of original semantics caused by pure noise generation, we propose a dynamic sparse anchoring strategy that directly integrates sparse pixels from the adversarial sample to guide the reconstruction. Additionally, we design a perception-guided early-stopping mechanism to prevent the perturbation resurgence caused by this direct pixel integration. Extensive experiments across diverse datasets and protection schemes demonstrate that our method consistently outperforms existing baselines, achieving state-of-the-art performance in efficacy and fidelity, successfully neutralizing even advanced adaptive protections.
DTGS: Physics-Embedded Dynamic Thermal 3D Reconstruction with Gaussian Splatting
Junbo Li ⋅ Hang Fu ⋅ Yinuo Wang ⋅ Weimin Yuan ⋅ Cai Meng ⋅ Xiangzhi Bai
Dynamic scene representation and rendering in the visible spectrum has been extensively studied. Compared to visible-light imaging, thermal infrared sensing offers all-weather observation and strong penetration capability, enabling perception in low-visibility environments. However, dynamic reconstruction in thermal infrared scenes is challenged due to atmospheric transport effects and interference from low-emissivity background regions, which often leads to motion blur and details loss in rendered results. To address these issues, we propose DTGS, a physics-driven dynamic 3D scene reconstruction approach. DTGS factorizes dynamic thermal scenes into radiometric and geometric components, modeling radiometric changes via rendering that integrates atmospheric radiative transfer, while capturing geometric deformations through Adaptive Temporal Gaussian basis. In addition, we design a thermal-radiance-weighted structural similarity loss to suppress gradient interference from low-radiance noisy regions. Furthermore, to demonstrate the effectiveness of our method, a large-scale dataset for this field named Dynamic_LTR is created. Experimental results demonstrate that our method achieves a 2.7 dB improvement in PSNR, a twofold speedup in rendering, and a threefold reduction in training time. The code and datasets will be available after acceptance.
DualSAT: A Dual-Branch GNN-Transformer Framework for SAT Solving
Wenzhu Yang ⋅ Zhanshan Li ⋅ Jingyao Li
Existing graph neural network-based methods for Boolean satisfiability (SAT) solving are generally constrained by limited local receptive fields, making it difficult to adequately capture the long-range dependencies between variables and clauses in SAT instances. Moreover, these methods overlook the multi-solution nature of SAT instances. Consequently, their expressive power and generalization performance remain limited. To address these issues, we propose DualSAT, a novel neural framework for the SAT problem, designed to enhance the phase selection heuristic of SAT solvers. Specifically, DualSAT adopts a novel dual-branch graph neural network that separately encodes local structural information and global structural information in the graph, and then fuses them through an attention-based cross-branch feature fusion module. An assignment decision head is further introduced to generate the initial assignment prediction of variables. This prediction process requires only a single forward pass, achieving a favorable balance between expressive power and inference efficiency. Furthermore, to account for the multiple solutions nature of SAT instances, we design a two-stage training strategy that combines supervised and unsupervised learning to improve the generalization performance of the model. Experimental results on multiple datasets show that, as an end-to-end assignment prediction model, DualSAT significantly outperforms NeuroSAT and NeuroBack. In addition, we integrate DualSAT into the classical SAT solver CaDiCaL, where it increases the number of solved instances by up to 11.6% and reduces the number of conflicts by up to 9.21% on SAT competition problem sets.
Dual-Space Preconditioning for Variational Inequalities and Root-Finding Problems
Jan Quan ⋅ Konstantinos Oikonomidis ⋅ Alexander Bodard ⋅ Panagiotis Patrinos
This paper develops a dual-space preconditioning framework for variational inequalities and root-finding problems beyond classical cocoercivity and Lipschitz continuity assumptions. To this end, we introduce both a novel cocoercivity-type condition and a relaxed Lipschitz continuity condition, and derive convergence of deterministic and stochastic methods under these generalized conditions. For isotropic preconditioners, the resulting algorithms can be interpreted as variable-stepsize variants of classical methods, recovering and extending recent analyses based on $(L_0, L_1)$-type growth conditions while allowing simpler proofs and larger stepsize ranges. We also show that the framework naturally handles constrained variational inequalities through weak Minty-type arguments. Finally, we study the associated continuous-time dynamics and derive a variational and optimal-control interpretation of the preconditioned flow. Overall, the results show that dual-space preconditioning provides a flexible mechanism for designing and analyzing operator methods beyond the standard Euclidean Lipschitz setting.
Dynamical Adapter Fusion: Constructing A Global Adapter for Pre-Trained Model-based Class-Incremental Learning
RuiQi Liu ⋅ Boyu Diao ⋅ Zijia An ⋅ Runjie Shao ⋅ Feixiang Ge ⋅ Zhulin An ⋅ Fei Wang ⋅ Yongjun Xu
Class-Incremental Learning (CIL) requires models to continuously acquire new classes without forgetting previously learned ones. A dominant paradigm involves freezing a pre-trained model and training lightweight, task-specific adapters. However, maintaining task-specific parameters hinders knowledge transfer and incurs high retrieval costs, while naive parameter fusion often leads to destructive interference and catastrophic forgetting. To address these challenges, we propose Dynamical Adapter Fusion (DAF) to construct a single robust global adapter. Grounded in the PAC-Bayes theorem, we derive a fusion mechanism that explicitly integrates three components: the optimized task-specific adapter parameters, the previous global adapter parameters, and the initialization parameters. We utilize the Taylor expansion of the loss function to derive the optimal fusion coefficients, dynamically achieving the best balance between stability and plasticity. Furthermore, we propose a Robust Initialization strategy to effectively capture global knowledge patterns. Experiments on multiple CIL benchmarks demonstrate that DAF achieves state-of-the-art (SOTA) performance.
Dynamic Context Modeling for Longitudinal Mental Health Monitoring under Distribution Shift
Seungwan Jin ⋅ Taehyung Noh ⋅ Junghyun Kim ⋅ Uichin Lee ⋅ Kyungsik Han
Mobile and wearable sensing offer a scalable pathway to longitudinal mental health monitoring, yet accurate prediction remains limited by two fundamental forms of distribution shift in human behavior. The first arises across users, where similar sensor patterns can carry different clinical meanings, a phenomenon we call inter-subject heterogeneity. The second arises within a single user, where behavioral baselines themselves evolve over time, producing intra-subject behavioral drift. In this domain, ground-truth labels come from Ecological Momentary Assessment (EMA), brief self-reports collected in everyday life that are intrusive and inherently sparse. As a result, approaches that address these distribution shifts by updating model parameters are difficult to apply. To address this gap, we propose DyCon (Dynamic Context Modeling), a dual-loop framework that personalizes by building and evolving a separate natural-language personal context for each user, without updating any model parameters. Within each session, a momentary refinement loop restructures this personal context to resolve inter-subject heterogeneity. Across sessions, a longitudinal update loop integrates only validated insights to track intra-subject behavioral drift. Extensive experiments on four public benchmarks (GLOBEM, PMData, StudentLife, LifeSnaps) show that DyCon consistently outperforms baseline methods, particularly under sparse supervision and behavioral drift. Our results suggest that evolving an explicit personal context offers a practical pathway for longitudinal personalization in real-world health monitoring.
Dynamic Cross-Modal Prompt Generation for Multimodal Continual Instruction Tuning
Tao Hu ⋅ Da-Wei Zhou
Multimodal Large Language Models (MLLMs) achieve strong performance through instruction tuning, yet real-world deployment often requires continual capability expansion across sequential tasks. In such scenarios, Multimodal Continual Instruction Tuning (MCIT) aims to acquire new capabilities while limiting catastrophic forgetting. Existing methods mainly follow a module-composition paradigm: they maintain task-level prompts or LoRA experts and dynamically route or aggregate a subset of them at inference. However, samples within the same task can still differ substantially in visual scenes, question intents, and reasoning demands. This motivates instance-level adaptation to individual query-image pairs rather than only selecting or combining task-level modules. To this end, we propose DRAPE (Dynamic Cross-Modal Prompt Generation), a prompt-learning framework that synthesizes continuous instance-specific soft prompts for MCIT. Instead of selecting prompts from a fixed pool, DRAPE derives prompt queries from the textual instruction and cross-attends to visual patch features, producing query-image conditioned prompts that are prepended to the frozen LLM. To mitigate forgetting during sequential updates, DRAPE applies null-space gradient projection to the shared projector and uses CLIP-based prototype routing for task-label-free generator selection at inference. Extensive experiments on MCIT benchmarks show that DRAPE achieves state-of-the-art performance among representative prompt-based and LoRA-based continual-learning baselines.
DyRA: Dynamic Residual Approximation for Efficient Matrix Multiplication in DNNs
Daewon Chae ⋅ Hyunwon Chung ⋅ Changwoo Lee ⋅ Hun-Seok Kim
Large-scale foundation models achieve strong performance across diverse tasks, but their size makes inference costly, largely due to dense matrix multiplications. Prior work reduces this cost by replacing dense weight matrices with efficient structured forms such as low-rank factorizations. However, these methods approximate weights rather than the output activations that determine inference accuracy. Consequently, small weight-space errors can be amplified by input activations, producing large output errors. In this work, we propose DyRA, an input-adaptive method that improves structured matrix multiplication approximation by correcting residual output errors during inference. We show that matrix multiplication can be approximated more effectively by directly optimizing low-rank factors of the output. DyRA builds on this insight by dynamically approximating and correcting the output error introduced by structured weight approximations. This combines efficient structured computation with input-dependent correction, yielding a more faithful approximation of full matrix multiplication under the same computational budget. Across vision, speech, and language models, DyRA consistently improves the accuracy–efficiency trade-off over structured weight approximations alone. Notably, DyRA achieves a $1.5\times$ end-to-end GPU speedup for DINOv3 while reducing accuracy degradation by more than $3\times$ relative to weight-only baselines.
ECHO: Continuous Hierarchical Memory for Vision-Language-Action Models
Yanbin Hu ⋅ Jin Cui ⋅ Jiayi Lu ⋅ Ruixuan Yang ⋅ Jun Ye ⋅ Boran Zhao ⋅ Xingyu Chen ⋅ Xuguang Lan ⋅ Pengju Ren
Memory capacity is a critical factor determining the performance of Vision-Language-Action (VLA) models in long-horizon manipulation tasks. Existing memory-augmented architectures primarily rely on linear or flat storage, lacking structural priors for manipulation categories and hierarchical organization. This deficiency hinders efficient experience retrieval and limits generalization to unseen long-horizon task compositions. Inspired by the hierarchical organization of human experience, we propose ECHO (Experience Consolidation and Hierarchical Organization), a novel memory framework operating within a Continuous Hierarchical Space. By employing a hyperbolic autoencoder, ECHO maps VLA hidden states into this space. Leveraging hyperbolic metrics and entailment constraint mechanisms, experience vectors are organized into a semantic memory tree that supports efficient top-down retrieval. In parallel, a background consolidation mechanism continuously refines the memory tree through geometric interpolation and structural splitting, supporting virtual memory synthesis in the continuous space. We integrate ECHO into the $\pi_0$ foundation model. Evaluations on LIBERO and preliminary real-world experiments demonstrate the effectiveness of our approach, notably achieving a 12.8% absolute improvement in execution success rate over the $\pi_0$ baseline on LIBERO-Long, while improving compositional generalization on cross-suite unseen long-horizon tasks.
EditDistill: Is It Possible to Guide Video Editing with Image Editing
guojun lei ⋅ Hong Li ⋅ Hongbing Yang ⋅ Lixue Gong ⋅ Chi Wang ⋅ Zheng Dong
Recent advances in image editing have enabled fine-grained and highly controllable visual modifications.However, achieving similar levels of controllability in video editing remains an open challenge, particularly for precise local edits, dynamic temporal effects, and consistent lighting across frames.In this paper, we propose EditDistill, a framework that distills image-level visual transformations from edits into a compact descriptor to guide video editing.Specifically, given a source video and an edit prompt, we first generate an edited version of the initial video frame using an off-the-shelf image editing model. We then distill the visual transformation of this original-edited image pair into a compact representation termed an edit descriptor.This descriptor is injected into a video diffusion model through our lightweight fusion module, which steers the denoising process toward the desired edit. To align the video trajectory with the demonstrated image edit while preserving motion and layout, we train with edit-descriptor alignment and DINO structural consistency losses, and derive an equivalent inference-time guidance formulation under flow matching.Extensive experiments demonstrate that our method achieves strong performance on standard benchmarks and produces high-quality video editing results across diverse scenarios.
Training Neural ODEs requires backpropagating through an ODE solve. The state-of-the-art backpropagation method is recursive checkpointing that balances recomputation with memory cost. Here, we introduce a class of algebraically reversible ODE solvers that significantly improve upon both the time and memory cost of recursive checkpointing. The reversible solvers presented calculate exact gradients, are high-order and numerically stable -- strictly improving on previous reversible architectures. On scientific modeling and time series classification experiments, reversible solvers reduce the training time by more than $2\times$ while using at least $10\times$ less memory on average than recursive checkpointing.
Efficient algorithms for linear regression with heteroskedastic errors
Sergei Vassilvitskii ⋅ Silvio Lattanzi ⋅ Aditya Bhaskara ⋅ Siddhant Chaudhary
We study the classical problem of linear regression under heteroskedastic noise, where observations arise from multiple sources with unknown and potentially unequal variances. Specifically, we consider a setting with $n$ sources, each providing a linear measurement of an unknown $d$-dimensional parameter vector $\beta^*$, where each source is associated with an unknown noise variance. We focus on the classic "subset-of-signals'' model, where one assumes that $m$ of the sources have a relatively low noise, formally, with variance of at most 1. Our goal is to accurately estimate $\beta^\*$ despite the presence of high-variance sources. We show that if $m$ is sufficiently large, it is possible to recover $\beta^*$ with sub-constant error. Specifically, we prove a recovery error bound of $\tilde{O} \left(\frac{\sqrt{nd}}{m}\right)$, significantly improving upon previously known results for this setting.
Efficient Computation and Best-Response Dynamics in Anonymous Two-Action Games with Linear Utilities
Michail Fasoulakis ⋅ Evangelos Markakis ⋅ Ioannis Panageas ⋅ Christodoulos Santorinaios ⋅ Jingming Yan
The computational complexity of anonymous games has been a central theme in algorithmic game theory, with foundational contributions by Daskalakis and Papadimitriou [2015], Goldberg and Turchetta [2017], and Cheng et al. [2017] establishing that the landscape of equilibrium computation admits a PTAS. In this paper, we investigate the class of $n$-player anonymous games with linear utilities and two actions. While finding an equilibrium in anonymous games is PPAD-hard for arbitrary utility functions—even with a constant number of actions [Chen et al., 2015]—we show that this linear structure allows for significant algorithmic improvements. Exploiting the fact that payoffs depend solely on the first moment of the players’ strategy distribution, we provide the first algorithm for computing exact Nash equilibria that runs in time polynomial in the number of players $n$. We achieve this by establishing a novel combinatorial characterization of the Nash equilibria. Finally, we establish a novel connection between Nash equilibria in these games and non-convex non-concave min-max optimization. Despite the technical challenges posed by this structural property, we design a sequential best-response dynamic that provably converges to an $\epsilon$-Nash equilibrium in $\mathcal{O}(\frac{1}{\epsilon})$ steps.
Efficient Prediction of Pass@k Scaling in Large Language Models
Joshua Kazdan ⋅ Colin Sullivan ⋅ Rylan Schaeffer ⋅ Youssef Allouah ⋅ Kyssen Yu ⋅ Noam Levi ⋅ Sanmi Koyejo
Assessing the capabilities and risks of frontier AI systems is a critical area of research; recent work has shown that repeated sampling can dramatically increase both. For instance, repeated sampling allows models to solve difficult math and coding problems, but it may also increase their risk of being jailbroken. Such results raise a crucial question: how can one accurately predict a model's behavior when scaled to a massive number of attempts, given a vastly smaller sampling budget? This question is directly relevant to model providers, who serve hundreds of millions of users daily, and to governmental regulators, who seek to prevent harms. To answer this question, we make three contributions. First, we find that standard methods for fitting these laws suffer from statistical shortcomings that hinder predictions, especially in data-limited scenarios. Second, we remedy these shortcomings by introducing a robust estimation framework, which uses a beta-binomial distribution to generate more accurate predictions from limited data. Third, we propose a dynamic sampling strategy that allocates a greater budget to harder problems. Combined, these innovations enable more reliable prediction of rare risks and capabilities at a fraction of the computational cost.
Efficient Reasoning via Constrained Optimization in Latent Space
Zhinan Hou ⋅ Xingchen Li ⋅ Keyou You
Large Reasoning Models (LRMs) have shown remarkable reasoning capabilities, yet they still suffer from overthinking, generating redundant reasoning steps which incur substantial token consumption. Existing methods, such as suppressing reflective keywords or forcing shorter reasoning lengths, attempt to mitigate this issue but inevitably truncate necessary steps and induce underthinking, thereby compromising performance. To address this dilemma, we investigate the latent representations and observe that efficient reasoning steps naturally cluster into a concentrated region in latent space, while those deviating from this region tend to produce verbose sequences. To leverage this, we keep reasoning focused within this region via a quadratic program which projects deviating hidden states back into the region. Then we propose a novel training-free framework to achieve efficient reasoning that reduces token generation costs without sacrificing performance. Extensive experiments conducted on four models ranging from 1.5B to 14B, and across six benchmarks in math reasoning, coding, and scientific QA, validate the effectiveness of our method, up to a 13.4\% improvement in accuracy while reducing generated tokens by 11.8\% to 52.8\%. Code and models will be publicly available.
Efficient Variational Inference for Log-Gaussian Cox Processes via Voronoi Tessellation
Yanshuo Liu ⋅ Jiaheng Qu ⋅ Sha Cao ⋅ Chi Zhang ⋅ Nan Zhang
Log-Gaussian Cox Processes (LGCP) are essential for modeling spatial point patterns but suffer from high computational costs and intractable likelihoods. We propose VoGCAM, an efficient Variational Voronoi Gaussian Coordinate Ascent Maximization framework for fitting LGCPs. Our approach first approximates the intractable integral in the LGCP likelihood using a flexible Voronoi tessellation, which incorporates observed points as integration nodes. By applying variational Gaussian approximation, we derive an Evidence Lower Bound (ELBO) that admits an explicit and closed-form expression. To optimize this objective, we develop a novel coordinate ascent algorithm that updates parameter blocks via Newton and fixed-point methods. We further enhance scalability by adopting a Nearest Neighbor Gaussian Process (NNGP) prior and utilizing the Woodbury formula to reduce matrix inversion costs. Theoretically, we prove the existence and uniqueness of the optimal solution and establish the convergence of our algorithm. Numerical experiments on synthetic and real-world data demonstrate that VoGCAM offers superior computational efficiency and inferential accuracy over state-of-the-art methods like INLA and VIFRK.
EGCA: A Spectral Perspective on Forward Process Design in Diffusion Models
Chi Zhang ⋅ Guichao Chang ⋅ Sirui Liu ⋅ Xin Zhang ⋅ Guoqiang Zhong
The design of the forward diffusion process is a foundational yet under-theorized aspect of diffusion modeling. Although it determines how data are corrupted and how quickly the forward dynamics approach a reference distribution, forward-process design remains largely governed by canonical defaults rather than a unified analytical framework. We propose Eigenvalue-Guided Control and Analysis (EGCA) for diffusion models, a spectral framework that uses the principal eigenvalue of the infinitesimal generator to analyze and guide forward-process design. Under suitable ergodicity conditions, this eigenvalue governs the exponential convergence rate to stationarity, providing a compact and interpretable summary of forward dynamics. We establish sufficient uniqueness and ergodicity conditions, derive eigenvalue-based convergence bounds, and adapt a practical numerical estimator for the principal eigenvalue. Empirically, we show that the eigenvalue strongly tracks forward convergence speed and serves as an organizing coordinate for the efficiency--quality trade-off across tractable forward-generator families. Experiments across multiple image datasets and diverse diffusion settings, including training formulations, architectures, and samplers, demonstrate that eigenvalue-guided design maintains or improves generation quality while reducing training cost. Beyond a single generator choice, our framework offers a spectral perspective for understanding and comparing forward-process choices in diffusion models.
Ego-HMB: Human Motion Bridging from Egocentric Images via Motion Bridge Diffusion Model
Gaoge Han ⋅ yongkang cheng ⋅ shaoli huang ⋅ Jing Zhang ⋅ Mingming Gong ⋅ Fakhri Karray ⋅ Tongliang Liu
In this paper, we move beyond the conventional setting of motion in-betweening that relies on structured pose inputs, and instead propose a new paradigm that is conditioned on egocentric (first-person) visual observations, thereby reducing the requirements on both users and upstream systems; we term this task Egocentric Human Motion Bridging. This is motivated by realistic consumer scenarios in VR/AR and robotics, where users typically have access only to raw egocentric imagery (e.g., from wearable devices) without reliable human pose annotations or motion capture signals. Unlike conventional motion interpolation in human pose-conditioned settings or motion generation conditioned on a complete temporal language or image context, the egocentric perspective introduces large viewpoint variations, partial body visibility, and severe temporal ill-posedness due to cross-modal endpoint constraints. To this end, we propose Ego-HMB, a unified diffusion-based model that internalizes both endpoint human pose estimation capability and temporally consistent 3D motion infilling ability within a single model through two training stages. Specifically, we first introduce Egocentric-to-Canonical Pretraining (pretraining stage), where we design a hierarchical curriculum with progressively increasing temporal spans. Then, we present a novel Motion Bridge Diffusion Training with endpoint and physics awareness as a fine-tuning stage, where our bridge diffusion diffuse from neighboring poses to enhance temporal continuity while preserving the global constraints between the two endpoints of a 3D motion sequence. Extensive experiments and comparisons with one-stage baselines (a generative model conditioned directly on two egocentric images) and two-stage baselines (estimation models followed by interpolation models) demonstrate that our approach consistently outperforms these baselines.
EgoStream: A Diagnostic Benchmark for Streaming Episodic Memory in Egocentric Vision
Rosario Forte ⋅ Giuseppe Lando ⋅ Antonino Furnari
Continuous episodic memory is a core capability for autonomous agents operating in dynamic, real-world environments, yet current streaming video benchmarks provide limited tools for diagnosing what models remember and for how long. We introduce Egostream, a diagnostic benchmark for streaming episodic memory evaluation in egocentric vision. \egostream organizes 2,250 curated questions along seven cognitive dimensions: detail, spatial, temporal, event, social, causal, and prospective memory. We introduce the Answer Validity Window (AVW), which specifies the temporal span an answer remains valid as the observed scene evolves. This allows us to expand the questions into 8,528 recall-conditioned evaluations, enabling controlled testing from instant to ultra-long-term recall while separating genuine model forgetting from natural world-state changes. We rigorously establish baseline performance through a unified streaming MLLM framework that compares several state-of-the-art memory-management mechanisms, covering sliding windows, attention sinks, KV-cache pruning, merging, and offloading. Experiments within a unified Qwen3-VL backbone reveal that comparable aggregate accuracies mask starkly different memory profiles. For instance, token pruning preserves fine-grained details and temporal structure significantly better than token merging, while quantized offloading rescues ultra-long-term recall. Ultimately, all mechanisms operate well below real-time (>1s per frame), and top performing methods ceil at about 45\% accuracy, exposing critical gaps in current architectures. Egostream provides the diagnostic testbed needed to close these gaps.
Eliciting Secret Knowledge from Language Models
Bartosz Cywiński ⋅ Emil Ryd ⋅ Rowan Wang ⋅ Senthooran Rajamanoharan ⋅ Neel Nanda ⋅ Arthur Conmy ⋅ Samuel Marks
We study secret elicitation: discovering knowledge that an AI possesses but does not explicitly verbalize. As a testbed, we train three families of large language models (LLMs) to possess specific knowledge that they apply downstream but deny knowing when asked directly. For example, in one setting, we train an LLM to generate replies that are consistent with knowing the user is female, while denying this knowledge when asked directly. We then design various black-box and white-box secret elicitation techniques and evaluate them based on whether they can help an LLM auditor successfully guess the secret knowledge. Many of our techniques improve on simple baselines. Our most effective techniques (performing best in all settings) are based on prefill attacks, a black-box technique where the LLM reveals secret knowledge when generating a completion from a predefined prefix. Our white-box techniques based on logit lens and sparse autoencoders (SAEs) also consistently increase the success rate of the LLM auditor, but are less effective. We release our models and code, establishing a public benchmark for evaluating secret elicitation methods.
Emergent Misalignment as Data-Mediated Transfer
Baris Askin ⋅ Muhammed Ustaomeroglu ⋅ Anupam Nayak ⋅ Gauri Joshi ⋅ Guannan Qu ⋅ Carlee Joe-Wong
Fine-tuning LLMs on narrow harmful datasets can induce Emergent Misalignment (EM), where models exhibit misaligned behavior far beyond the fine-tuning distribution. We argue that emergent misalignment can be better understood as a data-mediated transfer phenomenon: harmful fine-tuning examples do not induce uniform behavioral spillover, but interact with the structural properties of the dataset and the difficulty of the tasks relative to the model. Across our experiments, we find that misalignment appears more readily when fine-tuning and evaluation prompts share similar underlying functional structure, when prompts leave more room for coherent harmful completions, and when the target behavior has been more reliably learned by the model. The training pipeline itself also matters: pretraining composition shapes later misalignment. We further study Subliminal Learning (SL), where misalignment is transmitted by fine-tuning on seemingly benign data generated by a harmful teacher. Moving beyond the standard SFT setting, we for the first time compare this transfer under off-policy and on-policy distillation as well, allowing us to separate the roles of the teacher guidance and the training data distribution in transmitting misalignment. Together, these results argue for a data-centric view: Emergent/subliminal misalignment should not be treated as a simple consequence of isolated harmful fine-tuning examples, but as the result of interactions between fine-tuning data structure, pretraining distributions, and training channels.
EmoTrack: Robust Depression Tracking from Counseling Transcripts across Session Regimes
Zhaomin Wu ⋅ Jiayi Li ⋅ Bingsheng He
Text-based counseling is an important interface for AI mental-health support, where transcripts may be used to monitor depression severity and flag sessions requiring timely human review. However, robust PHQ-8 prediction across session regimes remains challenging: fine-tuning-based methods can exploit richer supervision but may generalize poorly under data scarcity, while prompt-based LLM methods are data-efficient but usually treat each transcript holistically and provide limited support for longitudinal context. We study robust depression tracking from counseling transcripts across single-session and multi-session regimes. We introduce LongCounsel-8, a multi-session counseling dataset with session-level PHQ-8 supervision for evaluating repeated-session tracking under partial symptom disclosure and cross-session continuity. We further propose EmoTrack, a PHQ-8 prediction framework that combines LLM-extracted clinical signals with frozen turn-level semantic embeddings and trains symptom-specific predictors over the resulting transcript representation. When prior sessions are available, EmoTrack can further incorporate them through compact cross-session memory. Experiments on LongCounsel-8 and DAIC-WOZ show that EmoTrack achieves a clear gain on the real single-session benchmark, including a 13.5% relative MAE reduction over the strongest DAIC-WOZ baseline, and remains competitive with the strongest longitudinal baseline on LongCounsel-8.
We study methods for simultaneous analysis of many noisy and biased estimates, each paired with an even noisier estimate of its own bias. The analyst's goal is to construct short calibrated intervals for each parameter. The standard debiasing approach, which subtracts the bias estimate from each biased estimate, inflates variance and yields long intervals. In this paper, we propose an empirical Bayes rebiasing strategy that starts from the fully debiased estimates and learns from data how much bias to reintroduce by estimating the unknown bias distribution. We provide convergence rates for the coverage of our intervals when the bias distribution is estimated using nonparametric maximum likelihood. Furthermore, we demonstrate substantial precision gains in prediction-powered inference, including pairwise LLM win-rate evaluations, as well as for inference of direct genetic effects in family-based GWAS.
End-to-End Neural Modeling of EM Response and Design Performance for Free-Form RFIC Passives
Yuhao Mao ⋅ Chenhao Chu ⋅ Martin Vechev ⋅ Hua Wang
Passive matching networks are critical components in radio-frequency integrated circuits (RFICs), but their design is bottlenecked by expensive finite-difference-based electromagnetic (EM) simulation. We propose an end-to-end neural framework for free-form RFIC passives that predicts both EM response and downstream design performance under varying frequency and circuit contexts, combining a frequency-conditioned neural simulator with a circuit-conditioned neural solver. To improve design performance prediction, the simulator outputs a neural EM response consisting of both the explicit EM response and a complementary response, together with a dependence regularizer that encourages the two responses to capture orthogonal information. To evaluate the proposed framework, we construct a large-scale dataset of free-form passive layouts, synthesized from both expert-designed contours and random paths. Our 21M-parameter model achieves high accuracy across a broad frequency range, circuit contexts and passive matching settings, with $R^2>0.98$ for almost all targets, while providing orders-of-magnitude faster evaluation than traditional EM simulation.
End-to-End Training for Unified Tokenization and Latent Denoising
Shivam Duggal ⋅ Xingjian Bai ⋅ Zongze Wu ⋅ Richard Zhang ⋅ Eli Shechtman ⋅ Antonio Torralba ⋅ Phillip Isola ⋅ Bill Freeman
Latent diffusion models (LDMs) enable high-fidelity synthesis by operating in learned latent spaces. However, training state-of-the-art LDMs requires complex staging: a tokenizer must be trained first, before the diffusion model can be trained in the frozen latent space. We propose UNITE – an autoencoder architecture for unified tokenization and latent diffusion. UNITE consists of a Generative Encoder that serves as both image tokenizer and latent generator via weight sharing. Our key insight is that tokenization and generation can be viewed as the same latent inference problem under different conditioning regimes: tokenization infers latents from fully observed images, whereas generation infers them from noise together with text or class conditioning. Motivated by this, we introduce a single-stage training procedure that jointly optimizes both tasks via two forward passes through the same Generative Encoder. The shared parameters enable gradients to jointly shape the latent space, encouraging a "common latent language". Across image and molecule modalities, UNITE achieves near state-of-the-art performance without adversarial losses or any pretrained encoders (e.g., DINO), reaching FID 2.12 and 1.75 for Base and Large models on ImageNet 256 x 256. Our experiments suggest that unification provides mutual benefits for tokenization and generation. We further analyze the Generative Encoder through the lenses of representation alignment and compression. These results show that single-stage joint training of tokenization and generation from scratch is feasible. Project Page: https://unite-tokenization-generation.netlify.app/
EnerGNN: Learning Optimization-Compatible Energy Functions for Exact Constrained Combinatorial Inference
Niki Triantafyllou ⋅ Maria Papathanasiou
This paper presents EnerGNN, an energy-based framework for exact constrained combinatorial inference. The central idea is to learn an optimization-compatible energy over binary configurations, so that hard constraints and deployment-time restrictions can be enforced by a mixed-integer programming (MIP) layer at inference time. Training shapes the energy landscape using optimality-aware contrastive learning so that high-quality feasible solutions occupy low-energy regions. We evaluate EnerGNN in two settings: direct decision problems, where the learned energy is optimized over the full binary decision vector, and strategic decomposition problems, where it is optimized over complicating binary variables and an exact solver handles the downstream problem. Across Quadratic Knapsack, Combinatorial Auctions, and a personalized medicine supply chain problem, EnerGNN achieves primal gaps below 1\% on most test instances, generalizes to larger problem instances without retraining, and supports zero-shot constraint injection by modifying only the inference MIP.
Energy-Adaptive Equivariant State Space Models for Noise-Robust Protein Structure Representation
Zhongyue Zhang ⋅ Runze Ma ⋅ Yanjie Huang ⋅ Shuangjia Zheng
Structure-informed protein representation learning is essential for effective protein function annotation and $\textit{de novo}$ design. With the rapid progress of protein structure prediction and experimental structure determination, a central challenge in the field is shifting from merely obtaining structural models to reliably leveraging predicted or measured structures for downstream learning. However, both crystal and AlphaFold-like predicted structures can contain structural uncertainty, which can bias local neighborhood construction and make robust protein representation learning challenging. To address these issues, we propose a novel equivariant Transformer-State Space Model (SSM) hybrid framework, termed $E^3$former, designed for efficient and robust protein representation. Our approach uses energy function-based receptive fields to construct proximity graphs that adaptively stabilize neighborhood selection under structural deviations, and incorporates an equivariant high-tensor-elastic selective SSM within the transformer architecture. These components allow the model to adapt to complex geometric interactions and extract structural features with a higher signal-to-noise ratio. Empirical results demonstrate that our model outperforms existing methods in structure-intensive tasks, such as inverse folding and binding site prediction, particularly when using predicted structures, owing to its enhanced tolerance to data deviation and structural uncertainty. Our approach offers a novel perspective for conducting biological function research and drug discovery using imperfect but increasingly available protein structure data. Our code is available on \url{https://anonymous.4open.science/r/E3former-207E}.
Enhancing Novel View Synthesis via Geometry Grounded Set Diffusion
Farhad G. Zanjani ⋅ Herbert Cai ⋅ Amirhossein Habibian
We present SetDiff, a geometry-grounded multi-view diffusion framework that enhances novel-view renderings produced by 3D Gaussian Splatting. Our method integrates explicit 3D priors, pixel-aligned coordinate maps and pose-aware Plucker ray embeddings, into a set-based diffusion model capable of jointly processing variable numbers of reference and target views. This formulation enables robust occlusion handling, reduces hallucinations under low-signal conditions, and improves photometric fidelity in visual content restoration. A unified set mixer performs global token-level attention across all input views, supporting scalable multi-camera enhancement while maintaining computational efficiency through latent-space supervision and selective decoding. Extensive experiments on EUVS, Para-Lane, nuScenes, and DL3DV demonstrate significant gains in perceptual fidelity, structural similarity, and robustness under severe extrapolation. SetDiff establishes a state-of-the-art diffusion-based solution for realistic and reliable novel-view synthesis in autonomous driving scenarios.
Ensemble Modeling for Time Series Forecasting: an Adaptive Robust Optimization Approach
Leonard Boussioux ⋅ Henry Mao ⋅ Dimitris Bertsimas
Accurate time-series forecasting is critical for a wide range of problems involving temporal data. Ensemble modeling is a well-established technique for leveraging multiple predictive models to increase accuracy and robustness, as the performance of a single predictor can be highly variable due to shifts in the underlying data distribution. This paper proposes a new methodology for building robust ensembles of time series forecasting models. Our approach uses Adaptive Robust Optimization (ARO) to construct a linear regression ensemble whose model weights adapt over time. We demonstrate the effectiveness of our method through a series of synthetic experiments and real-world applications, including air pollution management, energy consumption forecasting, and tropical cyclone intensity forecasting. Our results show that our adaptive ensembles outperform the best ensemble member in hindsight by 16-26\% in root mean square error and 14-28\% in conditional value at risk and improve over competitive ensemble techniques.
Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control
Jiazheng Zhang ⋅ Ziche Fu ⋅ Junrui Shen ⋅ Yunbin Zhao ⋅ Yunke Zhang ⋅ Zhiheng Xi ⋅ LongMa ⋅ Chenxin An ⋅ Zhihao Zhang ⋅ Shichun Liu ⋅ Dingwei Zhu ⋅ Shihan Dou ⋅ Shaofan Liu ⋅ han li ⋅ Wiggin Zhou ⋅ Aiden Adams ⋅ Tao Gui ⋅ Fei Huang ⋅ Qi Zhang ⋅ Xuanjing Huang
Policy entropy has emerged as a fundamental measure for understanding and controlling exploration in reinforcement learning with verifiable rewards (RLVR) for LLMs. However, existing entropy-aware methods mainly regulate entropy through global objectives, while the token-level mechanism by which sampled policy updates reshape policy entropy remains underexplored. In this work, we develop a theoretical framework of entropy mechanics in RLVR. Our analysis yields a first-order approximation of entropy change $\Delta \mathcal{H}$, giving rise to entropy polarity, a signed token-level quantity that predicts whether and how much a sampled update expands or contracts entropy. This analysis further reveals a structural asymmetry: Reinforcing frequent high-probability tokens triggers contraction tendencies, whereas expansive tendencies typically require lower-probability samples or stronger distributional correction. Empirically, we show that entropy polarity reliably predicts entropy change, and that positive and negative polarity branches play complementary roles in preserving exploration and strengthening exploitation. Building on these insights, we propose Polarity-Aware Policy Optimization (PAPO), which preserves both polarity branches and implements entropy control through advantage reweighting. With the empirical entropy trajectory as an online phase signal, PAPO adaptively reallocates optimization pressure between entropy-expanding and entropy-contracting updates. Experiments on mathematical reasoning and agentic benchmarks show that PAPO consistently outperforms competitive baselines, while delivering superior training efficiency and substantial reward improvements.
We study semi-offline value estimation, where a small on-policy sample supplements a large off-policy dataset. This regime is common in recommender systems and pre-deployment A/B testing, and the challenge is to use the limited on-policy data to correct systematic bias in a given value predictor. Recent work on Bellman calibration (BC) gives a post-hoc, model-agnostic correction along the predicted value. We propose Error-Guided Bellman Calibration (GBC), which trains an error model on the on-policy sample and calibrates the predictor along this learned error axis. We prove a completeness-free refinement guarantee, a theorem showing that the correction adapts to the intrinsic dimension of the prediction bias, and a finite-sample bound on the value-estimation error that isolates our advantage in the refinement term. Across synthetic, CRM, and D4RL benchmarks, GBC improves over existing calibration baselines.
Evaluating Neural Data Tokenizers: A Framework for Assessing Learned Representations of Spiking Activity
Federico D'Agostino ⋅ Alex Gilbert ⋅ Susanne Keller ⋅ Jaivardhan Kapoor ⋅ Nicolas Reategui ⋅ Vedang Lad ⋅ Matthias Kümmerer ⋅ Jai Bhagat ⋅ Emin Orhan ⋅ Hasan A Bedel ⋅ Nikos Karantzas ⋅ Fabian Sinz ⋅ Surya Ganguli ⋅ Jakob H Macke ⋅ Sophia Sanborn ⋅ Alexander Ecker ⋅ Katrin Franke ⋅ Matthias Bethge ⋅ Andreas Tolias ⋅ Konstantin Willeke
Foundation models of brain activity aim to learn representations from large-scale neural recordings that generalize across sessions, subjects, species, and downstream tasks. Transformer-based models have emerged as a promising architecture for this goal and begin all with the same first step --- tokenization: transforming sparse binary events into vector representations that downstream networks can consume. In other areas of machine learning, tokenization is its own design problem, distinct from the downstream foundation model. Several mostly implicit tokenization approaches dominate the literature, but their impact on downstream performance has not been systematically evaluated. We propose four explicit desiderata for a neural tokenizer---reconstruction fidelity, downstream utility, cross-instance generalization, and compression---and a framework of metrics spanning seven dimensions, including firing statistics, encoding and decoding, and latent structure and robustness. We assemble a benchmark of ten datasets spanning species, brain areas, recording technologies, and behavioral paradigms, and evaluate a representative set of discrete and continuous tokenizers, both single-neuron and population variants. Across the strategies evaluated, we show how tokenizers trained only on reconstruction can support downstream tasks. Yet reconstruction alone is misleading: reconstruction quality and downstream utility can be sharply dissociated. Our multi-axis framework is a response to this failure mode. These findings highlight tokenization as a critical and underexplored design axis in brain foundation models, and suggest that further progress will require continued systematic study of tokenization strategies.
Evaluating Physical Reasoning in LLM Agents Requires Construction Benchmarks
Wenhao Deng ⋅ Tian Xia ⋅ Yuru Jiang ⋅ Joemon M Jose ⋅ Tailin Wu
Physical reasoning is central to agents that step into the physical world. Models now excel at code and math, but their physical reasoning capabilities remain largely untested. Existing physical benchmarks evaluate goal-state matching rather than functional performance. This position paper argues that evaluating physical reasoning in LLM agents requires construction benchmarks, where physics verifies functional outcomes. To operationalize physical reasoning evaluation, we identify five capabilities spanning the full agentic loop from goal to verified artifact. Our audit of eleven physical benchmarks finds only two that combine realistic physics with functional verification, the minimum requirements for testing physical reasoning; both are construction benchmarks. Neither covers all five capabilities, yet frontier models already achieve near-zero success as task complexity increases. Progress depends on the breadth of benchmarks we build. We call on academia, game studios, and industry labs to expand construction benchmarks along both axes, widening capability coverage and deepening physical challenges. Their breadth determines how fully agents move from code and math to reshaping the physical world.
Evaluating Spatiotemporal Reasoning of Vision-Language Models in Atari Gameplay
Mingjia Huo ⋅ Yao Fu ⋅ Bo Chang ⋅ Yaqing Wang ⋅ Dawei Zhu ⋅ Dylan Zhang ⋅ Xinyang Yi ⋅ Lichan Hong ⋅ Ed Chi ⋅ Hao Zhang ⋅ Jiaxi Tang
Vision-Language Models (VLMs) have achieved strong performance on static visual understanding benchmarks, yet their ability to act in dynamic environments remains poorly understood. We introduce AtariBench, a controlled benchmark that evaluates VLMs across 20 Atari 2600 games, requiring models to make sequential decisions from high-temporal-resolution gameplay streams. By decoupling fast in-game dynamics from wall-clock inference speed, AtariBench focuses on whether models can perceive rapid visual state changes, reason over spatiotemporal dynamics, and ground actions in the current observation. Our experiments reveal substantial limitations in state-of-the-art VLMs: models often fail to infer object motion and direction, exhibit brittle sequential decision-making, and increasingly over-rely on textual interaction history rather than newly appended visual feedback. We further introduce a separability analysis to identify which in-game scores provide reliable signals of model competence, and use the benchmark to study how memory design and prompting strategies affect performance. Overall, AtariBench exposes a persistent gap between passive visual perception and active decision-making, providing a lightweight and rigorous testbed for developing VLMs that can reason and act in dynamic environments.
Evidence Guided Dual Expert Memory with Joint Routing for VLLMs Online Correction
ZHANG LINTONG ⋅ Yearang Lee ⋅ Seong-Whan Lee
Vision-large language models (VLLMs) editing aims to efficiently update model responses while retaining generalization and unrelated knowledge. Recent online editors store previous edits as reusable experts, forming a mixture-of-experts editing memory for scalable VLLM correction. By retrieving relevant experts at inference time, this paradigm avoids repeatedly overwriting model parameters and enables continual reuse of accumulated editing experience. However, existing sample-level residual experts usually encode each edit instance as a whole, entangling target-relevant visual evidence with irrelevant context. This coarse expert construction weakens expert reusability, leads to inaccurate routing, and increases unintended side effects during continual editing. To address this limitation, we reformulate online VLLM editing as evidence guided expert construction and routing. A fine mask identifies edit-relevant visual evidence, from which support-suppress experts are constructed to strengthen target consistent cues and weaken edit conflicting cues. Each expert is stored with routing signals and reliability scores, and later retrieved through joint sparse routing over visual evidence, textual context, and expert quality. Experiments demonstrate consistent improvements in editing reliability, generalization, and interpretability over competitive online editing baselines.
EvoCodeBench: Evaluating Coding Agents in Multi-Turn Iterative Interactions
Haiyang Shen ⋅ Xuanzhong Chen ⋅ Wendong XU ⋅ Yun Ma ⋅ Liang Chen ⋅ Kuan Li
Coding agents are increasingly used as iterative development partners, but most benchmarks still evaluate one specification followed by one final assessment. This fails to capture whether agents can maintain executable artifacts as requirements evolve. We introduce EvoCode-Bench, a benchmark of 26 stateful coding tasks spanning 227 evaluated rounds. Each task preserves the agent's workspace across 5--15 rounds, expresses requirements through observable behavior, and uses cumulative executable tests to check both new requirements and regressions against still-active prior ones. We evaluate 13 coding agents with two complementary metrics: MT@4, a four-attempt fail-stop multi-round score, and SR, a single-round score from a reference-completed prior state. For most agents, SR exceeds MT@4 by 22--40 points, indicating that isolated instruction-following does not imply persistent reliability. The gap also changes model rankings: the highest-SR agent (78.9) ranks only third in persistent execution (44.0 MT@4). Even the strongest agents achieve only \~50\% success on multi-turn metrics, and aggregate pass rate drops below half of round-1 performance by round~5. Failure analysis reveals tier-dependent behavior: weaker agents fail early, while stronger agents survive long enough to expose specification-tracking and regression failures. These results show that single-round and persistent multi-turn coding evaluate different capability dimensions. We release the benchmark data and Harbor multi-turn infrastructure.
EvoCUA: Evolving Computer Use Agents via Learning from Scalable Synthetic Experience
Taofeng Xue ⋅ Chong Peng ⋅ Mianqiu Huang ⋅ Linsen Guo ⋅ Tiancheng Han ⋅ Haozhe Wang ⋅ Xiaocheng Zhang ⋅ Xin Yang ⋅ Dengchang Zhao ⋅ Jinrui Ding ⋅ Xiandi Ma ⋅ Yuchen Xie ⋅ Peng Pei ⋅ Xunliang Cai ⋅ Xipeng Qiu
The development of native computer-use agents (CUA) represents a significant leap in multimodal AI. However, their potential is currently bottlenecked by the constraints of static data scaling. Existing paradigms relying primarily on passive imitation of static datasets struggle to capture the intricate causal dynamics inherent in long-horizon computer tasks. In this work, we introduce EvoCUA, a native agentic model that integrates data generation and policy optimization into a self-sustaining evolutionary cycle. This approach employs a verifiable synthesis engine to autonomously generate diverse tasks with executable validators, alongside a scalable infrastructure orchestrating tens of thousands of sandbox rollouts for mass experience acquisition. To internalize this experience, our iterative evolving learning strategy reinforces successful routines while transforming failure trajectories into rich supervision through error analysis and self-correction.EvoCUA achieves a 56.7\% success rate on the OSWorld benchmark, establishing a new state-of-the-art for open-weights models across multiple model scales. It significantly outperforms the previous best open-source model, OpenCUA-72B (45.0%), and surpasses leading closed-weights models such as UI-TARS-2 (53.1%), offering a highly parameter-efficient and reproducible alternative to massive closed-weights models. These results demonstrate the generalizability of the evolving paradigm across foundation models of varying scales, establishing a robust and scalable path for advancing native agent capabilities.
EvoDiagram: Agentic Editable Diagram Creation via Design Expertise Evolution
Tianfu Wang ⋅ Leilei Ding ⋅ Ziyang Tao ⋅ Yi Zhan ⋅ zhiyuan ⋅ Wei Wu ⋅ Yuxuan Lei ⋅ Junyang Wang ⋅ Yizhao Xu ⋅ Yin WU ⋅ Zhengyu Hu ⋅ Hongyuan Zhu ⋅ Nicholas Jing Yuan ⋅ Yanyong Zhang ⋅ Hui Xiong
High-fidelity diagram creation requires coordinated visual-spatial decisions over semantic topology, visual styling, and geometric layout. Existing methods face a representation gap: pixel-based generation offers limited object-level control, while code-based synthesis provides executable structure at the cost of intuitive manipulation. We introduce EvoDiagram, an agentic framework for editable diagram creation through an intermediate canvas schema that is both machine-actionable and directly manipulable by users. EvoDiagram first constructs a diagram manifest through coordinated structure, style, and layout agents, and then renders the manifest in a diagnostic verification environment. This environment produces objective defect signals and VLM-based critique, enabling localized refinement of the current diagram while also providing evidence for long-term design expertise evolution. The evolution mechanism distills refinement traces into a three-tier design expertise memory, where candidate strategies are linked to the verification evidence that supports them and are promoted into broader guidelines and principles only through repeated cross-context support. We further define CanvasBench, a canvas-recoverable benchmark and evaluation protocol with content, visual, and cognitive dimensions. The current manuscript focuses on the framework, benchmark design, and reporting protocol; empirical results should be added only when the corresponding artifacts, judge traces, and analysis scripts are available for verification. Our code is available at \href{https://anonymous.4open.science/r/evo-diagram/}{https://anonymous.4open.science/r/evo-diagram/}.
EvoMM: Reinforced Self-Evolving Multimodal Agentic Memory
Tong Zhao ⋅ Chenghao Zhang ⋅ Yucheng Tian ⋅ Yuyang Hu ⋅ Yutao Zhu ⋅ Zhicheng Dou
Multimodal large language model (MLLM) agents increasingly operate over long histories of dialogues, images, and evolving facts, yet finite context windows and weak persistent memory cause forgetting, temporal inconsistency, and loss of precise past evidence. Existing multimodal memory agents either discard visual evidence by reducing images to captions or treat memory as a static store that is built once and never revised during reasoning; self-evolving agents further accumulate experience via prompting without tight coupling to retrieval and memory revision. Closing this loop with reinforcement learning additionally suffers from credit assignment difficulties over long, multimodal action trajectories. To address these challenges, we propose EvoMM, a reinforced self-evolving multimodal agentic memory framework that unifies memory construction, retrieval, and evolution within a single agentic loop. EvoMM maintains a mutable two-level memory of turn-level short memories and session-level long summaries, organized by a heterogeneous multimodal memory graph with temporal, semantic, image-co-reference, and keyword edges. An iterative PLAN→RETRIEVE→REFLECT→ANSWER policy then revises long memory on the fly and distills structured retrieval experience into a cross-query store for reuse. To optimize this policy, we further introduce PRR-GRPO (Planning-Retrieval-Reflection GRPO), which augments outcome rewards with gated, action-aware process rewards and performs dense credit assignment over a memory-action tree with dual-scale advantage normalization. Experiments on Mem-Gallery, MMLongBench, and LoCoMo show that EvoMM consistently outperforms strong text-only, multimodal, and agentic memory baselines, establishing a training paradigm for memory that reads, revises, and evolves with each query.
Exact Channel Decoupling via Joint Diagonalization and Uniform Splicing for Diffusion Transformer Quantization
Bingyao Yu ⋅ yijin liu ⋅ Qinkai XU ⋅ Li Li ⋅ Yuxiang Fu
Diffusion Transformers (DiTs) have emerged as the state-of-the-art architecture in high-fidelity generative modeling, but their massive computational demands and multi-step inference process limit their deployment in practical scenarios. Post-training quantization (PTQ) offers a promising solution for accelerating inference, but DiTs suffer from severe activation outlier issues, which can lead to catastrophic accuracy loss in integer formats. Although some recent outlier mitigation methods employ orthogonal transformations or diagonal scaling to smooth the activation distribution, they either solve the transformation matrix solely for activations or fail to fully balance the inter-channel variance between activations and weights. Additionally, the construction of the calibration set by taking the weighted average of the activation values for each time step introduces spurious temporal correlations. To address these issues, we propose a joint diagonalization of the activation-weight covariance matrix based on the solution of a generalized eigenvalue problem (GEVP). We also transform the feature fusion of the calibration set's temporal dimension into probabilistic sampling and assembly in the spatial dimension using Uniform Token Splicing (UTS). Specifically, compared to suboptimal baseline methods, our approach reduces the FID by 16.06 and 6.49 on both the W3A4-quantized DiT-XL/2 and PixArt-$\Sigma$ models, respectively. This demonstrates that our method successfully suppresses outliers in activations and weights, maintaining DiT's performance under low-bit quantization conditions. Code is available at the anonymous repository \url{https://anonymous.4open.science/r/JDUS-DiT-74E1}.
Exact-Form Regret and Conservative Correlated Equilibria
Ashkan Soleymani ⋅ Patrick Jaillet ⋅ Gabriele Farina
Gradient descent is usually studied through external regret, but recent work shows that it controls richer deviations through infinitesimal vector-field guarantees \citep{ahunbay2024first}, semicoarse constraints \citep{ahunbay2025semicoarse}, and weakly convex proximal maps \citep{cai2025proximal}. We identify the finite-deviation invariant behind these guarantees. A feasible deviation \(\phi\colon K\to K\) induces the displacement field \(\Dphi(\xvec)=\xvec-\phi(\xvec)\). Online gradient descent has \(O(\sqrt T)\) regret against every feasible deviation with exact displacement, \(\Dphi=\nabla\Psi\) for a potential function $\Psi$, even when \(\Psi\) is nonconvex. Feasibility handles the projection boundary term, and exactness makes the potential telescope. The condition is sharp for smooth fields on simply connected domains, since any nonzero circulation along a feasible loop can be turned into a bounded cyclic adversary that forces linear regret. Semicoarse affine deviations are the linear-quadratic slice of this class, and proximal deviations are the Moreau-envelope slice. The exact-form class strictly contains weakly convex proximal deviations. For mirror descent, exactness is measured in the mirror geometry through the one-form \(\Dphi(\xvec)^{\transpose}\dd\nabla R(\xvec)\), which subsumes Bregman proximal regret and includes multiplicative deviations for multiplicative-weight updates. In convex games, no-regret with respect to exact-form deviations yields conservative correlated equilibrium, refining coarse and proximal correlated equilibrium.
Exact Gaussian Moment Matching for Residual Networks: a Second-Order Method
Simon Kuang ⋅ Xinfan Lin
We study the problem of propagating the mean and covariance of a general multivariate Gaussian distribution through a deep (residual) neural network using layer-by-layer moment matching. We close a longstanding gap by deriving exact moment matching for the probit, GeLU, ReLU (as a limit of GeLU), Heaviside (as a limit of probit), and sine activation functions; for both feedforward and generalized residual layers. On random networks, we find orders-of-magnitude improvements in the KL divergence error metric, up to a millionfold, over popular alternatives. On a variational Bayesian neural network, we show that our method attains hundredfold improvements in KL divergence from Monte Carlo ground truth over a state-of-the-art deterministic inference method. We also give an smooth-distance error bound showing that, under regularity assumptions, moment matching removes the leading low-variance errors and propagates higher-order local accuracy through the layers of a network.
Exact Recovery of Lipschitz Orthogonal Coordinate Transformations via Constrained Normalizing Flows
Isaac Manring ⋅ Kejun Huang
Nonlinear independent component analysis (nICA) is generally unidentifiable due in part to measure-preserving automorphisms (MPAs), which induce indistinguishable latent representations. We show that for non-Gaussian sources with mild regularity assumptions, such MPAs are not Lipschitz, motivating Lipschitz continuity as a structural assumption. Under this condition, we prove global identifiability of nICA within the class of Orthogonal Coordinate Transformations (OCTs) with bounded sources, recovering the true components up to permutation and scaling. Our analysis reduces the learning problem to a linear program over the Birkhoff polytope, and extends to functions with symmetric Jacobians, revealing connections to optimal transport. We introduce OCTNet and SYMNet, normalizing flow architectures that enforce these constraints in practice. Experiments demonstrate strong recovery performance and robustness to model misspecification, outperforming existing approaches.
Explaining Cross-Modal Model Behavior with Gradient-Estimation-Based Feature Interaction
Yi Cai ⋅ Xinpeng Li ⋅ Gerhard Wunder
Multi-modal machine learning has become a dominant paradigm for building large-scale models that integrate information across heterogeneous modalities and achieve unprecedented representational capabilities. However, the multi-modal nature of such models introduces unique challenges for explainability. Specifically, the standard feature attribution task considers only individual feature contributions. While solutions of this kind can be adapted to operate on each modality, such first-order explanations are insufficient to characterize cross-modal feature interactions, which are essential for understanding model behavior involving cross-modal alignment. To bridge the gap, this paper studies higher-order feature interactions in multi-modal settings and proposes Gradient-Estimation-Based Feature Interaction (GEFI). GEFI is derived from the proxy gradient estimation framework, and we establish its theoretical equivalence to the Shapley Interaction Index at arbitrary orders. Beyond the general theory, we instantiate GEFI-PM for CLIP-like models by exploiting their dual-encoder architecture. The resulting formulation enables the efficient estimation of interactions between visual and textual features. Compared to existing methods for revealing cross-modal interactions, which generally operate on fixed image patches, GEFI offers fine-grained pixel-level interactions while maintaining computational efficiency, yielding more expressive results.
Exploring Multi-Order Self-Similarity for Motion Understanding
Manjin Kim ⋅ Heeseung Kwon ⋅ Karteek Alahari ⋅ Minsu Cho
Space-time self-similarity (STSS), which captures visual correspondences across frames, provides an effective way to represent temporal dynamics for video understanding.In this work, we explore higher-order STSS and demonstrate how STSSs at different orders reveal distinct aspects of these dynamics. We then introduce the Multi-Order Self-Similarity (MOSS) module, a lightweight neural module designed to learn and integrate multi-order STSS features. It can be applied to a wide range of motion-centric tasks with only marginal computational cost and memory overhead. Extensive experiments on video action recognition, motion-centric video VQA, and robotic tasks consistently demonstrate substantial improvements, validating the broad applicability of MOSS as a general motion modeling module. The source code and checkpoints will be publicly available.
Exposing and Mitigating Temporal Attack in Deepfake Video Detection
Zheyuan Gu ⋅ Minghao Shao ⋅ Zhen Wang ⋅ Keyu Mao ⋅ Ailiang Lin ⋅ Shijie Zhang ⋅ hao jiang ⋅ mingyou liang ⋅ Mingkun Xu ⋅ Yusong Wang
While spatiotemporal deepfake detectors have shown high AUC, our experiments reveal that they remain susceptible to evasion attacks. These models tend to overfit on fragile temporal spectrum cues, rather than learning robust semantic causality. To mitigate this vulnerability, we propose SpInShield, a temporal spectral-invariant defense framework explicitly designed to decouple semantic motion from manipulatable spectral artifacts. We propose a learnable spectral adversary that dynamically synthesizes severe spectral deformations, simulating extreme attack scenarios. By employing a shortcut suppression optimization strategy, SpInShield compels the encoder to extract reliable forensic cues while autonomously purging unstable spectral statistics from the latent space. Experiments show that SpInShield obtains competitive performance on widely used datasets and outperforms the strongest baseline by 21.30\% AUC under simulated amplitude spectral attacks.
Extended Wasserstein-GAN Approach to Causal Distribution Learning: Density-Free Estimation and Minimax Optimality
Shu Tamano ⋅ Masaaki Imaizumi
Distributional causal inference requires estimating not only average treatment effects but also interventional outcome distributions, including quantiles, tail risks, and policy-dependent uncertainty. As a method for distributional causal inference, generative adversarial network (GAN)-based counterfactual methods are flexible tools for this task. However, these methods have several limitations. First, the objectives of certain techniques do not coincide with the statistical risk of the identifiable causal target, and therefore provide limited theoretical guarantees regarding estimable counterfactual distributions or optimality. Second, they tend to rely on unstable density-based methods, such as density ratio estimation. In this paper, we propose GANICE (GAN for Interventional Conditional Estimation) with several advantages: it (i) clarifies the conditional interventional distribution for each treatment--covariate state as the causal estimation target; (ii) estimates the conditional distribution such that its averaged Wasserstein risk is minimized; (iii) establishes minimax optimality. GANICE achieves these advantages through the introduction of the extended Wasserstein distance, the incorporation of a cellwise critic in its dual, and an optimality proof based on Besov space theory. Our experiments demonstrate that GANICE consistently outperforms existing methods.
ExtrapAir: Air Quality Inference at Unmonitored Locations via Weather-Bridged Spatial Attention
Gaoyun Lin ⋅ Binwu Wang ⋅ Zhengyang Zhou ⋅ Yang Wang
Air quality prediction at unmonitored locations is challenging because historical AQI observations are unavailable, while conventional spatiotemporal models largely rely on target histories to learn spatial dependencies. Meanwhile, widely available weather covariates provide useful station-specific evidence for target-absent inference. In this paper, we propose \textbf{ExtrapAir}, a lightweight weather-bridged framework for air quality inference at unmonitored locations. ExtrapAir preserves variable-specific weather patterns through per-variable weather encoding and learns adaptive weather and AQI correlations for spatial transfer. To support robust inference, we introduce \textbf{Activation Attention with Prior Inductive Biases}, which enables non-competitive spatial aggregation, incorporates geographic and semantic priors for station-pair guidance, and avoids full station-to-station attention computation for scalability. Extensive experiments on five real-world air-quality datasets against \textbf{32} baselines show that ExtrapAir reduces MAE by up to \textbf{15.24\%} overall and \textbf{16.01\%} on unmonitored stations, while cutting training time by \textbf{84.9\%} and memory usage by \textbf{76.23\%}, achieving a strong accuracy--efficiency trade-off.
Extrapolative Weight Averaging Reveals Correctness–Efficiency Frontiers in Code RL
Kunhao Zheng ⋅ Juliette Decugis ⋅ Pierre Chambon ⋅ Jonas Gehring ⋅ Taco Cohen ⋅ Benjamin Negrevergne ⋅ Gabriel Synnaeve
Linear interpolation between fine-tuned checkpoints has been shown to trace the Pareto front between competing objectives, but whether extrapolative weight averaging can extend such frontiers to new checkpoints useful at inference time, without additional RL training, remains unclear. We study this question in RL for competitive programming, where hidden unit tests under time and memory limits enforce both functional correctness and computational efficiency. Starting from a shared initialization, we train checkpoints under nested unit-test coverage: low-coverage rewards require passing smaller-input tests, while high-coverage rewards require passing progressively larger tests up to the full suite. This sweep reveals the emergence of a correctness–efficiency frontier: on hard problems, higher-coverage reward reduces optimization failures but increases correctness failures, leaving solve rate nearly unchanged. Interpolation between low- and high-coverage checkpoints recovers this frontier, while extrapolation extends it beyond the trained endpoints. Both the frontier and its extrapolative continuation appear across three inference settings, pure reasoning, tool use, and agentic coding, and across two model scales, 32B and 7B. At the problem level, moving along the frontier changes which problems are solved, making extrapolated checkpoints complementary policies in inference-time scaling: ensembles with extrapolative weight averaging broaden coverage and improve pass@250 on LCB/hard by 3.3\% over the best single checkpoint at matched sample budget. These results show that nested unit-test coverage in code RL induces a frontier that extrapolative weight averaging can navigate, extend, and exploit.
Face Deepfake-aware Recovery via Semantic-driven Facial Representation-based Watermarking
Yuan-Chih Chen ⋅ Chun-Shien Lu
Existing image watermarking methods typically entangle visual content with spatial structure, making them highly sensitive to geometric transformations and alignment discrepancies, especially in the presence of deepfake manipulations. In this paper, we propose a semantic-driven facial watermarking framework for robust identity recovery. The key idea is to decouple identity-related semantic information from spatial layout and encode it into a compact and spatially robust semantic representation. Specifically, we decompose a face into semantic components and aggregate deep features into component-wise latent representations, which are quantized via independent codebooks and converted into a compact bitstream for embedding. After decoding, the embedded semantic information is recovered and used to reconstruct identity-consistent facial content, even under slight geometric distortions and tampering. Experiments on CelebA-HQ and FFHQ demonstrate that our method significantly outperforms existing watermarking approaches in terms of reconstruction quality, identity preservation, and retrieval accuracy under both photometric and geometric attacks, validating the effectiveness of semantic component-wise encoding for reliable face recovery.
FADE: Fractional Anomalous Dynamics Extrapolation for Training-Free Diffusion Transformers Acceleration
Jinlong Yang ⋅ Jinke Wu ⋅ Lizilin ⋅ Yao Zhou
Diffusion Transformers (DiTs) have established themselves as the preeminent paradigm for high-fidelity generative modeling across various modalities, yet their practical deployment is hindered by the severe computational overhead of iterative sampling. While recent feature-caching methods mitigate this latency through feature reuse or forecasting based on integer-order operators, they remain limited in adapting to the non-stationary and highly curved dynamics of DiT latent trajectories. In this work, we reveal that DiT feature evolution exhibits anomalous diffusion characteristics and a certain degree of spatiotemporal heterogeneity, rendering static, integer-order forecasting operators inadequate and prone to error accumulation. To address this, we propose Fractional Anomalous Dynamics Extrapolation (FADE), a unified framework for training-free diffusion transformers acceleration. FADE models DiT feature evolution via anomalous diffusion and employs a Mittag--Leffler fractional operator to better adapt to high-curvature and non-linear latent trajectories of DiTs. Leveraging the empirical observation that these complex dynamics maintain cross-sample stability, we further construct a Look-Up Table with multiple dimensions through offline profiling. This enables zero-overhead, spatiotemporally adaptive parameter retrieval during online inference. Extensive evaluations across diverse architectures and modalities demonstrate that FADE achieves state-of-the-art acceleration, preserving exceptional generation fidelity under aggressive skip intervals with negligible latency overhead. Our code is anonymously available at https://anonymous.4open.science/status/FADE-8CB5.
Faithful Embeddings of Irregular and Asynchronous Data for Online Log-NCDEs
Benjamin Walker ⋅ Alexandre Bloch ⋅ Lingyi Yang ⋅ Sam Morley ⋅ Terry Lyons
Continuous-time models are a natural choice for irregular and asynchronous data. A central design choice is how to embed discrete observations into continuous time. Interpolation- and imputation-based embeddings reconstruct a continuous observation path, making the model sensitive to the choice of reconstruction. We show that this reconstruction step is unnecessary. On compact sets, universality for continuous functionals of paths transfers to universality for continuous functionals of discrete observation streams under any continuous and injective embedding. Guided by this result, and building on the rectilinear control path for Neural Controlled Differential Equations (NCDEs), we introduce a continuous and injective embedding for Log-NCDEs, a universal class of continuous-time models. This embedding records observations as increments and composes them over arbitrary query intervals to form log-signatures, giving interval-level summaries of a faithful embedding of the observed data. This avoids interpolation of observed variables, supports online computation, and allows prediction on output grids independent of the input sampling times. Experiments on synthetic controlled dynamics and real-world time-series datasets show that the representation is accurate, efficient, and robust to irregular, asynchronous, and sparse observations.
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
Yoav Gur-Arieh ⋅ Ana Marasovic ⋅ Mor Geva
Chains of thought (CoTs) have become central in interpreting and auditing behaviors of large language models. Yet growing evidence suggests that these traces often fail to faithfully represent the computations behind a model's predictions. Several faithfulness metrics have been proposed, but whether they indeed measure faithfulness remains unknown. Answering this requires ground-truth labels, which are hard to obtain since internal computations are not directly observable. Consequently, most works proposing metrics report only absolute scores or comparisons to prior metrics, and the few existing benchmarks rely on proxies like plausibility or importance, properties orthogonal to faithfulness that can mislead about whether a CoT can be trusted. We address this challenge by constructing tasks whose outputs reveal which intermediate computations must have produced them, and developing an automated labeling pipeline that yields ground-truth faithfulness labels at both the step and CoT level. Building on this methodology, we present BonaFide, a benchmark of 3,066 labeled CoTs across 13 tasks and 10 models, and use it to conduct the first systematic evaluation of prominent faithfulness metrics. Our experiments show that most metrics perform near chance, exhibit strong prediction biases and degrade on longer CoTs. The best metric reaches only 0.71 AUROC at the CoT level while another reaches 0.59 at the step level, with neither transferring across settings, while entailing prohibitively high computational cost. Our results expose fundamental gaps in current faithfulness evaluation and call for the development of more reliable and efficient metrics.
FaithSieve: Fine-Grained Evaluation of Math Proofs with Faithful Formal Evidence
Ziyu Wang ⋅ Qiming Dai ⋅ Yishan Wu ⋅ Zaiwen Wen
Large language models can now generate complex, multi-step mathematical proofs, but reliably determining their correctness and localizing early logical errors remains a critical challenge. Existing evaluation approaches largely depend on model-based natural-language judgments, which often overlook local reasoning gaps. While formal theorem provers like Lean offer a path to rigorous verification, using them to evaluate informal text requires solving locality and semantic mismatches: a prover might bypass a local flaw by proving an overly broad target, or validate an auto-formalized statement that drifts from the original mathematical intent. To address this, we introduce FaithSieve, a Lean-assisted framework for fine-grained evaluation of natural-language mathematical proofs. FaithSieve decomposes coarse proof steps into local reasoning units, extracts typed proof obligations, and verifies them through a formal evaluation agent. Formal validation is gated by semantic alignment scoring, so Lean evidence is incorporated only when the formal statement faithfully preserves the context, objects, and logical form of the original claim. We construct two expert-verified datasets, ProofLoc-Olympiad and ProofLoc-University, to benchmark first-error localization. On the 350-problem Olympiad dataset, FaithSieve using a GPT-5.4 backbone achieves 81.43\% exact first-error accuracy, outperforming the direct-judging baseline of 72.29\%. Furthermore, on the 200-problem ProofLoc-University benchmark spanning six advanced domains, FaithSieve reaches 84.5\% exact accuracy, compared to 75.0\% for the direct judge. Our work demonstrates that decomposing proofs into fine-grained units and grounding them with faithful formal evidence significantly improves reliable evaluation of natural-language reasoning.
FA-LAM: Focus-Aware Large Avatar Model for One-Shot 4D Animatable Gaussian Head
Yingdong Hu ⋅ Yisheng He ⋅ Yiming Jiang ⋅ Zehong Lin ⋅ Steven Hoi ⋅ Jun Zhang
We propose FA-LAM, a focus-aware large avatar model for one-shot animatable Gaussian head creation, while simultaneously enabling static 3D and dynamic 4D full-head recovery. The core of our method lies in a thorough analysis of the attention mechanisms and the entangled reconstruction and animation training pipeline adopted by prior state-of-the-art approaches. Our analysis identifies two main factors that compromise the quality of 3D full-head generation: (1) incorrect and noisy attention activations, and (2) conflicts between the tasks of reconstruction and animation. To address the first issue, we introduce a symmetric and semantic attention regularization strategy that leverages the inherent semantics and structural symmetry of human heads. To disentangle the objectives of reconstruction and animation, we develop a novel dual-phase training pipeline that separates the model's capabilities for large-view hallucination and animation into distinct modules. Moreover, we enhance our model to support multi-view and streaming 4D reconstruction in an efficient and memory-friendly manner through a core autoregressive modification with tailored visibility-aware token fusion. Collectively, these innovations enable FA-LAM to reconstruct animatable Gaussian full heads with superior quality, particularly in fine facial regions and large viewing angles.
Gradient descent has been of particular interest in modern machine learning beyond sole focus on optimization because some specific structures emerge through the optimization dynamics. Such behaviors result from optimization even though the learning objective does not explicitly encode the target structure, collectively called implicit bias, often preventing overparametrized models from fitting to spurious patterns. A typical instance is the max-margin implicit bias of a linear classifier, widely established for exponentially tailed loss functions. Even after having a given dataset separated, the parameter vector continues to evolve towards the max-margin direction asymptotically along the gradient descent dynamics. This phenomenon corroborates a frequent empirical observation of ``train longer, generalize better.'' However, the max-margin convergence is an asymptotic phenomenon, and what is worse, this asymptotic convergence rate is significantly slower than convergence in optimization. Even so, the parameter vector along gradient descent dynamics commonly correlates with the max-margin direction positively (though not exactly) within considerably fewer iterations than the asymptotic rate. By shedding another light on this classical yet profound problem, this work aims to understand the mechanism of this early-stage alignment phenomenon. Our theoretical results demonstrate that the parameter vector weakly aligns with the max-margin direction within $O(\exp(\exp(-\delta)))$ iterations, where $\delta>0$ is the permissible alignment error, which is shown to be tight. By tracking the radial and tangential flows, our proof operates on the alignment dynamics directly with dataset geometry and gets rid of the asymptotic expansion, which is a key insight to enabling faster weak alignment.
Fast Sandwich Products in Clifford Algebra
Travis Pence ⋅ Daisuke Yamada ⋅ Jiaqi Mo ⋅ Chanyoung Moon ⋅ Karthikeyan Sankaralingam ⋅ Vikas Singh
Clifford algebra is becoming increasingly common in machine learning, e.g., in neural operators, equivariant transformers, and structured linear layers. The main appeal is that scalars, vectors, oriented planes, volumes, and higher-grade quantities live together in a single algebraic space. In this setting, rotations are represented by \emph{rotors}, which are parametrized by bivectors, with only $\binom{n}{2}$ degrees of freedom. Most Clifford algebra-based models work only with small algebras, in part because of a bottleneck in the sandwich product, $x \mapsto r x r^\dagger$, a basic operation for applying rotations to multivectors. It is evaluated through geometric products or via a full action matrix, both of which scale poorly with the $2^n$-dimensional Clifford basis. We ask whether the $\binom{n}{2}$-parameter description can be carried all the way through application, and show that the answer is yes. The sandwich action decomposes grade-wise into exterior powers of the induced vector rotation. We factor this vector rotation into Givens rotations and apply their lifted actions as sparse updates on pairs of blades. The resulting algorithm never needs to form the rotor coefficients or the $2^n \times 2^n$ action matrix. Empirically, our algorithm achieves up to $64 \times$ wall-clock and $35\times$ memory improvements over existing rotor-sandwich implementations, enabling rotor-based layers at scales previously out of reach.
FC-DGCN: Deep Graph Convolutional Network for Face Clustering and Recognition
Ling Ding ⋅ Di Jin ⋅ Zhizhi Yu ⋅ Liang Yang ⋅ Dongxiao He
Face recognition has achieved strong performance, but further gains often require much larger labeled datasets, making annotation costly. This motivates the use of unlabeled data, where face clustering is important for applications such as annotation and recognition. Recent GCN-based methods have shown promising results on affinity graphs, but balancing clustering quality and scalability on large-scale data remains challenging. To solve this problem, we propose FC-DGCN, a face clustering framework based on deep residual GCNs. The proposed method formulates clustering on an affinity graph through two complementary tasks: node confidence estimation and edge connectivity estimation. Specifically, we develop DGCN-V to estimate node confidence on the global graph and DGCN-E to estimate edge connectivity on local subgraphs. To better characterize the local cluster structure, we further revise the confidence-estimation target. Based on these two modules, FC-DGCN links each node to a higher-confidence neighbor with strong predicted connectivity, which naturally induces identity-consistent clusters. Experiments on MS-Celeb-1M show that FC-DGCN outperforms competitive baselines in Pairwise F-score and BCubed F-score while remaining computationally practical. Moreover, the clusters generated by FC-DGCN provide effective pseudo labels for low-label face recognition, leading to substantial improvements on MegaFace and IJB-A.
Feasible Policy Optimization for Safe Reinforcement Learning
Yujie Yang ⋅ Yuanxu Sun ⋅ Wenyu Li ⋅ Beiyan Jiang ⋅ Tao Zhang ⋅ Shengbo Eben Li
Policy gradient methods serve as a cornerstone of reinforcement learning (RL), yet their extension to safe RL, where policies must strictly satisfy safety constraints, remains challenging. While existing methods enforce constraints in every policy update, we demonstrate that this is unnecessarily conservative. Instead, each update only needs to progressively expand the feasible region while improving the value function. Our proposed algorithm, namely feasible policy optimization (FPO), simultaneously achieves both objectives by solving a region-wise policy optimization problem. Specifically, FPO maximizes the value function inside the feasible region and minimizes the feasibility function outside it. We prove that these two sub-problems share a common optimal solution, which is obtained based on a tight bound we derive on the constraint decay function. Extensive experiments on the Safety-Gymnasium benchmark show that FPO achieves excellent constraint satisfaction while maintaining competitive task performance, striking a favorable balance between safety and return compared to state-of-the-art safe RL algorithms.
FeatCal: Feature Calibration for Post-Merging Models
Yanggan Gu ⋅ Shuo CAI ⋅ Zihao Wang ⋅ Wenjun Wang ⋅ Yuanyi Wang ⋅ Pengkai Wang ⋅ Sirui Huang ⋅ Su Lu ⋅ Jianmin Wu ⋅ Hongxia Yang
Model merging combines task experts into one model and avoids joint training, retraining, or deploying many expert models, but the merged model often still underperforms task experts. We study this performance gap through feature drift, the difference between features produced by the merged model and by the expert on the same input. Our theory decomposes this drift into upstream propagation and local mismatch, tracks how it propagates and combines through later layers in forward order, and links final feature drift to output drift. This view motivates FeatCal, which uses a small calibration set to calibrate the merged model weights layer by layer in forward order, reducing feature drift while staying close to merged weights and preserving the benefits of model merging. FeatCal uses an efficient closed-form solution to update model weights, with no gradient descent, iterative optimization, or extra modules. On the main CLIP and GLUE benchmarks, FeatCal beats Surgery and ProbSurgery, the closest post-merging calibration baselines: 85.5% vs. 77.0%/78.8% on CLIP-ViT-B/32 Task Arithmetic (TA) and 85.2% vs. 83.7%/82.2% on FLAN-T5-base GLUE. On CLIP-ViT-B/32, 8 examples per task reach 82.9%, and 256 examples per task take 53 seconds, about 4x faster than both baselines, showing better sample efficiency and lower calibration cost.
Felid: A Flexible and Efficient Design for Transformer Fine-Tuning over Encrypted Data
Linru Zhang ⋅ Jun J Sim ⋅ Xiangning Wang ⋅ Jiahao Zhong ⋅ Kaiyu Zhou ⋅ Huanyi Ye ⋅ Yongsen Zheng ⋅ Lushan Song ⋅ Xiaojian Liang ⋅ Yingting Liu ⋅ Yujing Sun ⋅ Huaxiong Wang ⋅ Pu Duan ⋅ Kwok-Yan Lam
Fully Homomorphic Encryption (FHE) enables non-interactive, privacy-preserving transformer fine-tuning, protecting highly sensitive data as inputs to large language models. Prior work has predominantly pursued efficiency gains through architectural refinements, yet overall performance remains bounded by the substantial cost inherent in encrypted computation. To address these issues, we target the bottleneck at the ciphertext operation level. Concretely, we propose a novel packing strategy that supports batching over a flexible number of inputs, making it well-suited for fine-tuning workloads. In addition, we design operation-specific matrix evaluation algorithms that significantly reduce (i) the number of ciphertext rotations and relinearizations, and (ii) the number of bootstrapping invocations. Additionally, we exploit functional bootstrapping to fold the evaluation of element-wise activation functions directly into the bootstrapping step, thereby reducing overall computational complexity. As a result, our framework achieves a $30\times$ reduction in ciphertext rotations and relinearizations across all matrix operations, as well as a $2.5\times$ reduction in FHE bootstrapping invocations over the entire computation flow, compared to the latest state-of-the-art work. It is worth noting that our system completes fine-tuning of a 2-layer BERT-style model with a batch size of 16 in 650 seconds, achieving a $8\times$ speedup over the state-of-the-art work.
FELPS: Fair and Efficient Scheduling for Multi-LoRA Serving System
Yukai Ding ⋅ Chuang Hu ⋅ Fangcheng Fu ⋅ Xinyan Li ⋅ Susie X Rao ⋅ Shanshan Feng ⋅ Hao Wang ⋅ Xiao Yan ⋅ Jiawei Jiang
Low-Rank Adaptation (LoRA) is a popular parameter-efficient fine-tuning method for Large Language Models (LLMs), and practical systems usually need to serve many concurrent LoRA adapters to process different tasks. This is challenging due to the inherent tension between \textit{fairness} and \textit{efficiency}, i.e., achieving fairness requires frequent switching among the adapters such that they have similar service quality, while good efficiency requires to minimize adapter switching to reduce memory movement overheads. Existing schedulers can only achieve a fixed trade-off between fairness and efficiency, and their trade-offs are often suboptimal. To tackle this problem, we propose \textbf{F}airness-aware \textbf{E}fficient \textbf{L}oRA \textbf{P}riority \textbf{S}cheduling (FELPS), which explicitly decomposes fairness and efficiency considerations. In particular, for fairness, FELPS adapts the classic proportional fairness and ensures that the amount of service received by each adapter is proportional to its workload. For efficiency, FELPS prioritizes adapters that are already loaded to GPU memory. To conduct scheduling, FELPS combines to the fairness and efficiency terms with an adjustable weight to achieve arbitrary trade-offs between fairness and efficiency. Evaluations on real workloads show that FELPS achieves a superior trade-off between fairness and efficiency than all baselines.
Fewer Tokens, Fewer Layers: Efficient Vision Token Pruning and On-Policy Distillation to Accelerate VLMs
Shuai Wang ⋅ Shitong Shao ⋅ Qi Xuan ⋅ Zhaowei Zhu ⋅ Jiaheng Wei
Vision-language models (VLMs) have achieved remarkable progress. However, processing high-resolution images and long videos generates a massive number of vision tokens, creating a severe inference bottleneck due to the quadratic complexity of attention mechanisms. Previous efforts to mitigate this issue have primarily focused on accelerating the prefill stage by reducing the number of vision tokens, often yielding limited practical speedups while neglecting the increasingly time-consuming decoding stage when the generated response is long. In addition, existing approaches heavily rely on attention maps or similarity maps that are not compatible with efficient attention implementations. To comprehensively address these limitations, we first propose **FlashPruner**, a lightweight module that selects the most helpful vision tokens without relying on attention maps, significantly reducing prefill time. For the decoding stage, we adopt self-speculative decoding and theoretically prove that on-policy distillation guarantees a lower bound for the speculative acceptance rate. Guided by this formulation, we introduce **LoFT**, an on-policy distillation method that effectively trains a highly reliable and efficient draft model including FlashPruner and early layers of the language model. Extensive experiments demonstrate that after training with LoFT, FlashPruner achieves substantial inference speedups while preserving task performance. Notably, on MVBench, FlashPruner improves the time-to-first-token by 1.5$\times$ and decoding speed by 2$\times$, maintaining 96.8\% relative accuracy with only 10\% vision tokens.
Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric Refinement
Lingyu Kong ⋅ Ruicheng Li ⋅ Ruicheng Wang ⋅ Sicheng Xu ⋅ Chengtang Yao ⋅ Jianfeng Xiang ⋅ Jiaolong Yang
Monocular geometry estimation has recently achieved impressive performance across diverse scenes. However, state-of-the-art models still face notable distortion in local 3D structure, especially in fine details, like thin structure and small objects. We attribute this limitation to an architectural mismatch: most current models decode 3D geometry within a 2D parameterization, where feature interactions are governed by image-plane proximity rather than true 3D spatial relationships. This inadvertently mixes features from geometrically distant surfaces, resulting in over-smoothed geometry particularly around thin or elongated structure. In this paper, we propose a fine-detail monocular geometry estimation with Self-Guided Sparse 3D Refinement (SSR) that lifts monocular geometry modeling from 2D image space to 3D space for high-fidelity metric-scale point map. Our model lifts the coarse point map from a foundation base model onto a sparse voxel shell and refines it via SSR. The SSR employs sparse convolutions that aggregate features based on 3D spatial locality, avoiding feature mixing across depth discontinuities. Extensive experiments on diverse datasets demonstrate that our method significantly outperforms existing approaches in recovering fine detailed 3D geometry across both quantitative metrics and qualitative visualizations.
Fine-tuning Does Not Reach All: Uneven Safety and Knowledge Dynamics in Language Models
Ziwei Wang ⋅ Xinwei Guo ⋅ Jiaxin Zhang ⋅ Guanhua Chen ⋅ Haiyan Wu ⋅ Xiangyu Zhao ⋅ Lei Ma ⋅ Xin Yao ⋅ Shan He ⋅ Xuetao Wei
Fine-tuning is widely used to align large language models (LLMs) with desired behavioral norms such as safety, yet it is typically evaluated using aggregate metrics that assume uniform improvements across the model. However, fine-tuning operates within a constrained parameter space, and its entangled effects on localized knowledge remain poorly understood. To address this, we investigate fine-tuning at the level of individual knowledge concepts using a controlled framework with representative concepts and paired safe/unsafe samples under varying data compositions. By jointly analyzing behavioral safety via model outputs and internal knowledge accessibility via perplexity, we obtain a unified view of how knowledge is modified and expressed. Our results reveal that fine-tuning does not uniformly improve safety. Instead, it (1) exhibits strong concept-level heterogeneity, including persistent low-safety concepts and robustly safe concepts with clear domain specificity; (2) improves safe–unsafe discrimination on average, but coexists with familiarity on both safe and unsafe content or even loses previously learned safe knowledge; and (3) decouples internal knowledge from external behavior, as changes in familiarity do not reliably translate into safer outputs, especially due to early-stage disruption of existing safety structures. Our findings position consistency as a distinct objective beyond conventional aggregate safety and demonstrate that reducing unsafe knowledge exposure under uniform concept-level coverage actively homogenizes safety distributions across model knowledge, offering a principled path toward alignment that is not merely effective on average, but uniformly reliable, predictable, and robust across knowledge concepts.
Fine-tuning language encoding models on slow fMRI improves prediction for fast ECoG
Aditya Vaidya ⋅ Richard Antonello ⋅ Alexander Huth
Neuroscientists have recently turned to intracranial brain recording methods, like electrocorticography (ECoG), for human experiments because of the fine spatial and temporal resolution that they afford. Models trained on this data, however, are fundamentally restricted by the patient populations that can receive the implants necessary for recording. We propose using non-invasive fMRI to bridge the gap in training data. Using spoken language representations fine-tuned on fMRI, we build encoding models of ECoG. These representations showed improved prediction performance in ECoG, even though the temporal resolution of fMRI is two orders of magnitude worse. Prediction improved in frequency bands well beyond what is directly measured in fMRI. Next, to test the procedure's generalization ability, we fine-tuned models on fMRI responses that were temporally downsampled by a factor of 2. Despite the loss in resolution, these models were able to predict fMRI and ECoG responses at levels comparable to the original fMRI-tuned models. Finally, we showed that ECoG performance steadily scales with the amount of fMRI-tuning data. Our results show that "slow" data like fMRI can be a valuable resource for building better models of "fast" brain data like ECoG. In the future, integrating across multiple recording methods may further improve performance in other applications, like decoding.
FinReasoning: A Hierarchical Benchmark for Reliable Financial Research Reporting
Yiyun Zhu ⋅ Yidong Jiang ⋅ Ziwen Xu ⋅ Yinsheng Yao ⋅ Dawei Cheng ⋅ Jinru Ding ⋅ JIE XU
Large language models (LLMs) are increasingly deployed in financial research workflows, where their role is evolving from single-model assistance for human analysts toward autonomous collaboration among multiple agents. Yet recent real-world deployments reveal persistent failures such as factual errors, numerical inconsistencies, and shallow analysis, which can distort assessments of corporate fundamentals and trigger severe economic losses. While existing benchmarks have begun to evaluate such failures, they score all aspects of the generated analysis in one pass, failing to distinguish whether a model fails at foundational stages like auditing and correction, or underperforms at generating research-grade insights. Consequently, it obscures capability bottlenecks and the specialized strengths essential for multi-agent role assignment. To address these gaps, we introduce FinReasoning, a hierarchical benchmark that decomposes the core capabilities of financial research into semantic consistency, data alignment, and deep insight. We further propose a fine-grained evaluation framework that strengthens hallucination-correction assessment and incorporates a 12-indicator rubric for core analytical skills. FinReasoning reveals clear capability stratification across model types. Closed-source models (like Doubao-Seed-1.8) perform strongly overall and are better suited for core reasoning agents in multi-agent financial systems; open-source general models (like Qwen3-235B) show clear capability divergence and consistently underperform on Semantic Consistency, making them less suited for quality-sensitive generation tasks; financial-domain models (like Fin-R1) generate moderate insights but lack foundational auditing skills, requiring upstream quality control for domain-specific analysis. Our work provides a reliable evaluation foundation for multi-agent financial application systems and has already been deployed in pilot tests across several real-world scenarios. The resource is available at https://anonymous.4open.science/r/FinReasoning.
FiRe: Fine-grained Multimodal Reasoning for Enhanced Image Generation
Yongjin Kim ⋅ Yoonjin Oh ⋅ Ye Rin Kim ⋅ Hyomin Kim ⋅ Jeeyoung Yun ⋅ Yu-Jung Heo ⋅ Min-Jun Kim ⋅ Sungwoong Kim
With the rapid progress of Multimodal Large Language Models (MLLMs), unified MLLMs that jointly perform image understanding and generation have advanced significantly. However, despite the inherent reasoning capabilities of unified MLLMs for self-reflection and self-refinement, their use in text-to-image generation remains largely underexplored. Meanwhile, existing multimodal reasoning–based image generation methods mostly rely on prompt augmentation or holistic image–text alignment judgments, without fine-grained reflection and refinement of detailed prompt attributes, leading to limited fine-grained control. To address this limitation, we propose FiRe, a Fine-grained Multimodal Reasoning method for enhanced image generation by MLLM. In specific, FiRe performs a fine-grained multi-step reasoning by first decomposing the prompt into key visual requirements and then self-judging their satisfaction in the generated image, followed by localized refinement according to self-generated precise feedback. In addition, to further strengthen the MLLM's multimodal reasoning ability, we introduce FiRe-GRPO, a reinforcement learning method tailored to FiRe. Since standard Group Relative Policy Optimization (GRPO) suffers from sparse, outcome-based rewards in multi-step reasoning, we formulate our reasoning process as a step-level decision-making problem, design step-specific rewards, and compute step-level advantages for granular credit assignment within GRPO. Extensive experiments demonstrate that FiRe consistently outperforms competitive text-to-image baselines, including existing reasoning-based methods, with particularly substantial gains on compositional text-to-image benchmarks.
First-Order Trajectory Matching: Fast Ensemble Predictions of Chaotic, Turbulent, Stochastic Systems
Shreya Jha ⋅ Timo Schorlepp ⋅ Nicholas Geissler ⋅ Jules Berman ⋅ Benjamin Peherstorfer
We introduce First-Order Trajectory Matching (FTM), a surrogate-modeling method that learns the first-order local transport of probability mass from trajectories of stochastic systems. By matching the symmetric first-order motion of trajectories, FTM learns the probability current velocity, whose flow preserves time marginals to match ensemble averages, while also capturing current-like trajectory quantities such as fluxes, circulations, and barrier-crossing currents. FTM learns the current velocity directly from trajectories, avoiding drift, diffusion, and score estimation. Our stability analysis separates discretization error from sampling variance and shows that the one-step simulation-free FTM loss is stable when temporal resolution and sample size are properly balanced. Across stochastic dynamical systems and PDE examples, we empirically demonstrate that FTM provides trajectory-aware ensemble predictions at low, deterministic-rollout cost.
FISC: Time-Series Forecasting via First-Layer Statistical Calibration Constraints
Hua Wang ⋅ Xiyuan Zhang ⋅ Fan Zhang
Multivariate time-series forecasting requires reliable cross-series interaction, yet many attention- and mixing-based models couple heterogeneous variables without explicit entry-stage alignment. Such uncalibrated early interaction can lead to poorly conditioned optimization and degraded generalization when cross-series statistics drift. We propose First-layer Statistical Calibration (FISC), a simple framework that imposes statistical calibration constraints only on the first cross-series coupling layer while leaving deeper coupling layers learnable without repeated calibration. FISC estimates training-set cross-series statistics, converts them into a statistical affinity matrix, and convexly fuses this matrix with learnable coupling parameters. A row-wise simplex projection then yields a nonnegative row-stochastic mixing operator, providing bounded and interpretable entry-layer interaction. By restricting calibration to the first layer, FISC calibrates the entrance interaction without repeatedly constraining deeper feature transformations. To reduce temporal redundancy, FISC further combines a time-domain branch with a data-adaptive orthogonal-domain branch derived from training-time temporal statistics. Experiments on standard long- and short-term forecasting benchmarks show that FISC improves forecasting accuracy while maintaining favorable efficiency. Ablations and non-stationarity analyses further show that first-layer-only calibration consistently outperforms no calibration, last-layer calibration, and all-layer calibration, supporting a simple principle for robust cross-series learning: constrain the entrance interaction with training-set statistics, while letting deeper representations remain data-adaptive.
FLARE: Verifying MILP Reformulations with LLM-Based Formal Proof Synthesis
Henry Robbins ⋅ Connor Lawless ⋅ Madeleine Udell ⋅ Ellen Vitercik
Mixed-Integer Linear Programming (MILP) is a fundamental tool for combinatorial optimization with extensive real-world applications. A central challenge lies in designing efficient MILP formulations. Large Language Models (LLMs) offer new opportunities to automate the modeling process, from deriving formulations to strengthening them. To ensure correctness, we need robust methods to compare formulations. However, existing approaches evaluate formulations numerically and fail to reason about the behavior on general problem instances. We resolve this limitation by introducing a constructive notion of MILP reformulation that can be formalized in Lean and machine-checked. We develop FLARE, a method that uses an LLM-based agent and the Lean proof assistant to verify proposed reformulations against a reference. To evaluate our approach, we introduce FormulationBench, a challenging dataset of 20 problems and 116 formulations. FLARE outperforms existing methods with 96.9% accuracy. For cases where formal guarantees aren't necessary, we introduce FLARE-NL , an LLM proxy that achieves 99.3% accuracy. These methods enable reliable verification in automated optimization modeling.
FlashSSM-3D: A Distilled State Space Model for Lightning-Fast Dense 3D Reconstruction
Nyle Siddiqui ⋅ Mubarak Shah
3D geometric understanding is undergoing a major transition where rigid, deterministic, optimization-based pipelines are being replaced by large-scale, deep learning models. While recent methods achieve impressive performance on tasks such as multi-view 3D reconstruction and camera pose estimation, they suffer from two critical limitations: $\textbf{(i):}$ current state-of-the-art (SOTA) models require industry-level compute and training corpora, severely limiting accessibility, and $\textbf{(ii):}$ recent approaches remain constrained to modifying existing pretrained transformer models, since substantially changing their architecture would require retraining from scratch. We instead study a complementary direction: training an attention-free state space model for dense 3D understanding. We introduce $\textbf{FlashSSM-3D}$, the first 3D reconstruction state space model capable of performing accurate geometric reasoning over hundreds of images with linear computational complexity. However, training such a model from scratch with competitive performance is impractical under typical academic compute constraints. To address this, we train FlashSSM-3D by distilling knowledge from VGGT, avoiding the need to learn 3D geometric priors entirely from scratch. However, standard distillation losses are poorly suited for dense 3D reconstruction, where preserving real-world spatial structure is essential. We therefore propose a $\textbf{geometry-aware knowledge distillation}$ method specifically tailored for large-scale 3D reconstruction models. Our method introduces two novel losses: a Procrustes-based alignment loss that enforces global structural consistency between student and teacher 3D point predictions, and a camera reprojection loss that directly supervises view-consistent camera geometry between teacher and student. Extensive experiments across multi-view 3D reconstruction and camera pose estimation benchmarks demonstrate that FlashSSM-3D achieves competitive performance while requiring only $\sim25\\%$ of the training data and less than $10\\%$ of the training compute used by VGGT. At inference time, FlashSSM-3D benefits from the practical efficiency of state space models, reconstructing scenes from 500+ views in under 10 seconds while maintaining competitive accuracy.
FLIP: Fast and Accurate Global Lipschitz Estimation for Large Feedforward Networks
Hongbo Chen ⋅ Zihao Ren
The Lipschitz constant of neural networks is widely used in robustness, control, generalization bounds, and training stability. As modern neural networks continue to scale, existing methods for Lipschitz estimation are either very time-consuming or overly loose. In this work, we focus on both fast and accurate Lipschitz estimation for large feedforward neural networks. We first reformulate the well-known semidefinite program (SDP) framework into an explicit recursive form, and introduce a key theoretical result, bound backpropagation: the SDP objective has upper and lower bounds that can be propagated backward along the recursive computation graph and depend only on partial variables. Leveraging this property, we decompose the original SDP into a sequence of small subproblems, each minimizing a combination of upper and lower bounds over partial variables. We prove that each subproblem is one-dimensional and convex, enabling fast solution via the secant method. Finally, we present the full FLIP algorithm and its complexity, yielding fast and tight Lipschitz estimation for large feedforward networks. Experiments on both randomly generated and trained networks show that FLIP improves speed and tightness by several orders of magnitude over prior methods, giving the tightest Lipschitz estimates for 100M-parameter networks in 3 seconds.
Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models
Bowen Ping ⋅ Xiangxin Zhou ⋅ Penghui Qi ⋅ Minnan Luo ⋅ Liefeng Bo ⋅ Tianyu Pang
Recent work has demonstrated that online reinforcement learning (RL) can substantially improve the quality and alignment of flow matching models for image and video generation. Methods such as Flow-GRPO and CPS cast the denoising process as a Markov Decision Process and apply PPO-style ratio clipping to enforce a trust region. However, we argue that ratio clipping is structurally ill-suited for flow models: the probability ratio between new and old policies is a noisy, single-sample estimate of the true policy divergence, leading to over-constraining in some regions of the trajectory and under-constraining in others. We propose Flow-DPPO (Flow Divergence Proximal Policy Optimization), which replaces ratio clipping with a divergence proximal constraint. A key observation is that the per-step policy in flow models is Gaussian, enabling exact and cheap computation of the KL divergence between old and new policies. Flow-DPPO employs an asymmetric divergence mask that blocks gradient updates only when they simultaneously move away from the trusted region and violate the divergence threshold. Experiments show that Flow-DPPO achieves higher rewards with better KL-proximal efficiency, alleviates catastrophic forgetting, promotes balanced multi-objective optimization, and enables stable multi-epoch training where ratio clipping degrades.
FlowHOI: Flow-based Semantics-Grounded Generation of Hand-Object Interactions for Dexterous Robot Manipulation
Huajian Zeng ⋅ Lingyun Chen ⋅ Jiaqi Yang ⋅ Yuantai Zhang ⋅ Fan Shi ⋅ Peidong Liu ⋅ Xingxing Zuo
Recent vision-language-action (VLA) models can generate plausible end-effector motions, yet they often fail in long-horizon, contact-rich tasks because the underlying hand-object interaction (HOI) structure is not explicitly represented. An embodiment-agnostic interaction representation that captures this structure would make manipulation behaviors easier to validate and transfer across robots. We propose FlowHOI, a two-stage flow-matching framework that generates semantically grounded, temporally coherent HOI sequences, comprising hand poses, object poses, and hand-object contact states, conditioned on an egocentric observation, a language instruction, and a 3D Gaussian splatting (3DGS) scene reconstruction. We decouple geometry-centric grasping from semantics-centric manipulation, conditioning the latter on compact 3D scene tokens and employing a motion-text alignment loss to semantically ground the generated interactions in both the physical scene layout and the language instruction. To address the scarcity of high-fidelity HOI supervision, we introduce a reconstruction pipeline that recovers aligned hand-object trajectories and meshes from large-scale egocentric videos, yielding an HOI prior for robust generation. Across the GRAB and HOT3D benchmarks, FlowHOI achieves the highest action recognition accuracy and a 1.7$\times$ higher physics-simulation success rate on GRAB than the strongest diffusion-based baseline, while delivering a 40$\times$ inference speedup. We further demonstrate real-robot execution across four contact-rich, dexterous manipulation tasks, highlighting the effectiveness of retargeting our generated HOI representations to robotic dexterous hands, achieving significantly higher success rates than prior methods due to improved physical plausibility and semantic alignment.
Flow Map Language Models: One-step Language Modeling via Continuous Denoising
Chanhyuk Lee ⋅ Jaehoon Yoo ⋅ Manan Agarwal ⋅ Sheel Shah ⋅ Jerry Huang ⋅ Aditi Raghunathan ⋅ Seunghoon Hong ⋅ Nicholas Boffi ⋅ Jinwoo Kim
Language models based on discrete diffusion have attracted widespread interest for their potential to provide faster generation than autoregressive models. Despite their promise, these models typically produce samples whose quality sharply degrades in the few-step regime, preventing a dramatic speedup in practice. Here, we show that language models based on continuous flows over one-hot token embeddings can outperform discrete diffusion in both quality and speed. Importantly, our continuous formulation defines a unique flow map that can be learned directly for efficient few-step inference, a structure we show is unavailable to discrete methods. In this setting, we show that both the flow and its associated flow map can be learned with simple cross-entropy objectives that respect the simplex geometry of the data, and we identify three distinct choices for flow map distillation whose performance we compare in practice. Using these insights, we build a flow language model (FLM), a continuous flow that matches state-of-the-art discrete diffusion baselines on the One Billion Words (LM1B) and OpenWebText (OWT) datasets. We then distill FLM into a flow map language model (FMLM), whose one-step generation exceeds the 8-step quality of recent few-step discrete diffusion language models. Our work challenges the widely-held hypothesis that discrete noising processes are necessary for generative modeling over discrete modalities and paves the way toward accelerated language modeling at scale.
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
Quantong Qiu ⋅ Zhiyi Hong ⋅ Yi Yang ⋅ Haitian Wang ⋅ Kebin Liu ⋅ Qingqing Dang ⋅ Juntao Li ⋅ Min zhang
The quadratic computational complexity of standard attention mechanisms presents a severe scalability bottleneck for LLMs in long-context scenarios. While hybrid attention mechanisms combining Full Attention (FA) and Sparse Attention (SA) offer a potential solution, existing methods typically rely on static allocation ratios that fail to accommodate the variable retrieval demands of different tasks. Furthermore, head-level dynamic sparsity often introduces severe computational load imbalance and synchronization long-tails, which hinder hardware acceleration during autoregressive decoding. To bridge this gap, we introduce Flux Attention, a context-aware framework that dynamically optimizes attention computation at the layer level. By integrating a lightweight Layer Router into frozen pretrained LLMs, the proposed method adaptively routes each layer to FA or SA based on the input context. This layer-wise routing preserves high-fidelity information retrieval while ensuring contiguous memory access, translating theoretical computational reductions into practical wall-clock speedups. As a parameter-efficient approach, our framework requires only 12 hours of training on 8xA800 GPUs. Extensive experiments across multiple long-context and mathematical reasoning benchmarks demonstrate that Flux Attention achieves a superior trade-off between performance and inference speed compared with baseline models, with speed improvements of up to 2.8x and 2.0x in the prefill and decode stages.
FlyingDrones: A Dataset and Benchmark for Optical Flow Estimation from UAV motion
Laurenz Ruzicka ⋅ Roman Pflugfelder ⋅ Daniel Cremers
Optical flow is an important cue for drone analysis in aerial surveillance, airspace monitoring, and long-range target tracking, where motion remains informative when appearance is weak. Estimating the motion of small Unmanned Aerial Vehicles (UAVs), however, is difficult because distant targets occupy few pixels, have low contrast, and move against complex dynamic backgrounds. Standard optical-flow benchmarks emphasize large, object-centric motions and therefore do not characterize distant aerial targets well. We introduce FlyingDrones, a generated benchmark with dense ground-truth optical flow and segmentation masks derived from physically based rendering, and define resolution-normalized metrics (EPE Norm) that normalize flow by image dimensions. We evaluate 16 state-of-the-art models on small-target motion estimation. Standard RAFT degrades substantially on small targets at low resolution, whereas VideoFlow's multi-frame variant achieves the best drone-region EPE at every tested resolution. Aerial-pretrained two-frame models such as Sea-RAFT and WAFT are competitive at 270$\times$480 but lose their advantage at 1080$\times$1920, and the best fine-tuning domain depends on resolution: KITTI helps when drone motion is sub-pixel-to-few-pixels, while Sintel-style fine-tuning helps once displacements grow with image size. Overall, long-range drone analysis benefits most from multi-frame temporal context, scale-aware evaluation, and resolution-matched fine-tuning.
FM-ChangeNet: Learning Change through Pathwise Feature Transport
Roie Kazoom ⋅ George Leifman ⋅ Genady Beryozkin
We present FM-ChangeNet, a pathwise-supervised framework for change detection that reformulates bi-temporal reasoning as continuous transport in feature space rather than static endpoint comparison. Given encoded pre and post-temporal representations, we construct intermediate latent states and learn a time-conditioned velocity field $\hat{v}_\theta(z_t,t)$ along the transformation trajectory. This pathwise formulation constrains the predictor over a continuum of intermediate states, providing a denser and less ambiguous supervision signal than conventional endpoint-only segmentation and enabling the model to capture temporal evolution explicitly. The learned velocity field is not only a transport mechanism but also an interpretable representation of change: its magnitude serves as a spatially localized change cue that helps distinguish true structural variation from nuisance effects such as illumination shifts and spatial misalignment. We develop a hierarchical multi-scale architecture with cross-temporal alignment, time-conditioned coarse-to-fine flow decoding, and a unified objective that couples flow supervision, trajectory consistency, spatial regularization, and segmentation loss. Experiments on remote sensing benchmarks show that the proposed framework produces more structured and robust change representations while achieving state-of-the-art performance.
FOAM: Factored One-sided Adam-Moment for Practical and Scalable SOAP
Jaemyung Yu ⋅ Byeongho Heo ⋅ Sangdoo Yun ⋅ Dongyoon Han
Second-order optimizers (\eg Shampoo, SOAP) have emerged as compelling alternatives to Adam/AdamW and promise faster per-iteration convergence than AdamW in transformer training. However, they still remain costly in practice: even SOAP still pays substantial optimizer-state memory and eigendecomposition overhead. We ask whether the benefits of SOAP can be retained at nearly the efficiency of AdamW. In this paper, we answer yes with a lightweight yet improved SOAP, Factored One-sided Adam-Moment (FOAM). The path to practical impact is twofold: first, FOAM employs input-only preconditioning, dropping the output-side Kronecker factor to reduce memory and remove one eigendecomposition. A per-row Hessian decomposition identifies this as the right one-sided choice when output curvature is approximately isotropic. Second, FOAM replaces SOAP's dense rotated second moment with a factored Adam-moment estimator based on row-by-column statistics. Under the K-FAC Fisher assumption, this estimator is further proven to be pointwise equivalent to SOAP's update. Our empirical results support the claim that FOAM matches strong SOAP and Muon baselines in validation loss across a diverse scale of large language model (LLM) pre-training, uses optimizer state closer to AdamW, runs within Muon's wall time, and retains robustness across a five-fold learning-rate range. Scaling to 760M parameters, FOAM demonstrates that SOAP's gains need not require SOAP's cost.
Forced Orders: What LLM Leaderboards Hide About Model Comparisons
Zonglin Di ⋅ Berk Ustun ⋅ Yang Liu
LLM benchmarks are usually consumed as scalar leaderboards, where adjacent ranks are read as adjacent capability differences. We argue that this display should be complemented rather than simply trusted or discarded: benchmark scores are useful summaries, but a sorted score table can overstate which model pairs are actually distinguishable. This limitation is especially clear for accuracy- and success-rate leaderboards, whose averages hide task-level structure, and it also appears in pairwise systems such as Elo-style ratings when pairwise evidence is compressed into a total order. We propose LLM tier ranking as a task-level reporting framework for benchmarks. Given benchmark outputs, we construct pairwise wins, losses, and ties between models on each task and aggregate them into ordered tiers using Selective Preference Aggregation (SPA). Cross-tier comparisons are reported only when supported by sufficient task-level agreement; models in the same tier are left unresolved, not declared equal. Across AIME 2025, $\tau$-bench airline, xbench DeepSearch, GDPVal, SWE-Bench Verified, and Chatbot Arena, tier ranking often preserves useful broad capability classes while abstaining from unsupported fine-grained distinctions among frontier models. In Chatbot Arena perturbation studies, tier outputs are more stable than Elo-style total orders, expressing weaker evidence through later tier separation rather than rank reshuffling. We recommend reporting tiered leaderboards, solution paths, top-tier size, singleton-winner thresholds, and pairwise resolution statistics alongside scalar scores.
ForceFlow: Learning to Feel and Act via Contact-Driven Flow Matching
Shuoheng Zhang ⋅ Yifu Yuan ⋅ Hongyao Tang ⋅ YAN ZHENG ⋅ Qiaojun Yu ⋅ Pengyi Li ⋅ Guowei Huang ⋅ Helong Huang ⋅ Xingyue Quan ⋅ Jianye Hao
Existing imitation learning methods enable robots to interact autonomously with the physical environment. However, contact-rich manipulation tasks remain a significant challenge due to complex contact dynamics that demand high-precision force feedback and control. Although recent efforts have attempted to integrate force/torque sensing into policies, how to build a simple yet effective framework that achieves robust generalization under multimodal observations remains an open question. In this paper, we propose ForceFlow, a force-aware reactive framework built upon flow matching. For contact-stage policy design, we investigate force signal fusion mechanisms and adopt an asymmetric multimodal fusion architecture that treats force as a global regulatory signal, combined with a joint prediction paradigm that enhances the policy's understanding of instantaneous force and historical information, thereby achieving deep coupling between force and motion. For task-level hierarchical decomposition, we divide manipulation into a vision-dominant approach stage (VLM-based pointing for target localization) and a touch-dominant interaction stage (force-driven contact execution), with a Vision-to-Force (V2F) handover mechanism that explicitly decouples spatial generalization from contact regulation. Experimental results across six real-world contact-rich tasks demonstrate that ForceFlow achieves a 37\% success rate improvement over the strong baseline ForceVLA while maintaining significantly lower cost. Moreover, ForceFlow exhibits accurate force signal prediction and demonstrates superior performance in contact force self-regulation and zero-shot out-of-distribution (OOD) generalization.
Foundation Pareto Flow Policy for Multi-Objective Reinforcement Learning
Zhanjiang Yang ⋅ Lijun Sun ⋅ Yueming Li ⋅ Meng Li ⋅ Yijun Yang
The overarching goal of reinforcement learning (RL) is to efficiently optimize the agent's policy for various environments, tasks, dynamics, and even user preferences, which often consist of multiple and potentially conflicting rewards/objectives. This necessitates the learning of Pareto-optimal policies that can navigate different trade-offs among objectives. However, existing multi-objective RL (MORL) methods, whether trained online or offline, are typically limited to task- or objective-specific policy learning and do not address generalization of learned skills to unseen tasks/objectives. In this work, we propose the Foundation Pareto Flow Policy (FP2) for MORL, inspired by the pre-mid-post-training paradigm underlying recent foundation models. Specifically, we instantiate FP2 as a multi-modal diffusion transformer equipped with a likelihood-free flow matching loss, enabling unified, scalable, and instructable learning of diverse Pareto-optimal trajectories. We further develop a three-stage training pipeline to gradually optimize the policy: (1) pre-training on a large-scale MORL dataset, (2) mid-training on a carefully curated high-quality dataset, and (3) post-training on few collected trajectory data from new tasks and trade-offs. Empirical results demonstrate that FP2 consistently outperforms previous state-of-the-arts on the D4MORL benchmark, achieving superior zero-shot interpolation and extrapolation on complex Pareto manifolds as well as efficient few-shot adaptation to new task/dynamics scenarios.
Fourier Contour Learning for Efficient and Traceable Cardiac MRI Quantification
Yi Yu ⋅ Zhenyu Bu ⋅ Ziyu Zhang ⋅ Yixuan Liu ⋅ Aotian Chen ⋅ James Mihalich ⋅ Parker Martin ⋅ Xia Ning ⋅ Yuchi Han ⋅ Yuan Xue
Automated cardiac MRI quantification is commonly framed either as dense segmentation followed by voxel counting or as direct regression of global measurements. The two paradigms are difficult to unify: segmentation is visually interpretable but label-intensive and output-heavy, while regression is compact but less verifiable. We propose Fourier Contour Learning, which represents cardiac boundaries with compact anatomical Fourier contours. The predicted coefficients define smooth, interpretable contours and enable differentiable volume computation, connecting mask supervision, measurement supervision, and unlabeled-image regularization within one framework. We instantiate this representation with an anatomy-query decoder and cross-slice attention which creates a traceable path from image to contour to clinical measurement. Across three public datasets and an in-house all-phase cohort, our method achieves segmentation accuracy comparable to strong dense-prediction baselines, improves boundary agreement, and reduces clinically relevant volume error while using only a small number of output parameters per structure. We further show that the learned model provides an efficient interface for vision-language question answering and surface reconstruction. These results demonstrate the potential of Fourier contours as a unified intermediate representation for scalable, traceable, and clinically grounded cardiac analysis. Code will be publicly available upon acceptance.
FrequencyBooster: Advancing Pixel Diffusion for High-Fidelity Image Generation
Lichen Ma ⋅ Zipeng Guo ⋅ Yu He ⋅ Xiaolong Fu ⋅ Luohang Liu ⋅ Jingling Fu ⋅ Junshi Huang ⋅ Yan Li
To circumvent the inherent fidelity bottlenecks and optimization misalignment of VAE-based latent diffusion, pixel-space diffusion models have emerged as a compelling end-to-end paradigm. However, existing pixel diffusion models often struggle to balance computational efficiency with the preservation of high-frequency details. They frequently resort to patch-based compression or restricted local decoding, leading to a "spectral compromise" where high-frequency and fine-grained pixel information are suppressed. To address these challenges, we propose \textbf{FrequencyBooster}, a novel framework designed to empower pixel diffusion with full-frequency modeling capabilities without prohibitive overhead. The core of our method is a high-capacity decoder that specializes in extracting exhaustive high-frequency details and low-frequency semantics, the latter of which is derived from a Diffusion Transformer (DiT) backbone. Unlike prior works that sacrifice global context for local refinement, FrequencyBooster leverages high-dimensional feature representations to maintain global structural integrity while achieving superior pixel-level precision. Extensive experiments on ImageNet demonstrate the effectiveness of our approach: our model achieves a state-of-the-art FID of \textbf{1.60} at $256 \times 256$ resolution within only 320 epochs. Furthermore, at $512 \times 512$ resolution, FrequencyBooster attains an FID of \textbf{1.69}, significantly outperforming existing pixel-space and latent-space generative models.
From Average Sensitivity to Small-Loss Regret Bounds under Random-Order Model
Shinsaku Sakaue ⋅ Yuichi Yoshida
We study online learning in the random-order model, where the multiset of loss functions is chosen adversarially but revealed in a uniformly random order. By extending the *batch-to-online* transformation of Dong and Yoshida (2023), we show that if an offline algorithm enjoys a $(1+\varepsilon)$-approximation guarantee, an *average sensitivity* bound controlled by a function $\varphi(\varepsilon)$, and stability with respect to $\varepsilon$, then we can obtain a *small-loss* regret bound typically of order $\tilde O(\varphi^{\star}(\mathrm{OPT}_T))$, where $\varphi^{\star}$ is the concave conjugate of $\varphi$, $\mathrm{OPT}_T$ is the offline optimum over $T$ rounds, and $\tilde O$ hides polylogarithmic factors in $T$. Our result refines their original $(1+\varepsilon)$-approximate regret guarantee and applies to a broad class of problems, including online $k$-means clustering and online low-rank approximation. We further apply our approach to online submodular function minimization using $(1\pm\varepsilon)$-cut sparsifiers of submodular hypergraphs, obtaining a small-loss regret bound of $\tilde O(n^3 + n^{3/4}\mathrm{OPT}_T^{3/4})$, where $n$ is the ground-set size; we also demonstrate its applicability to online $\ell_1$ regression. Our work sheds light on the power of sparsification and related algorithmic techniques in achieving small-loss regret bounds in the random-order model, without requiring structural assumptions on loss functions, such as linearity or smoothness.
From Clips to Streams: A Unified Framework for Streaming Sign Language Translation
Hoang Quan Dang ⋅ Xin Shen ⋅ Yanbin Liu ⋅ Xin Yu ⋅ Ling Chen
Sign Language Translation (SLT) research has typically focused on pre-segmented clips, a constraint that disconnects models from the continuous reality of real-world communication. To bridge the gap from clips to streams, we formally define \textbf{Streaming SLT}, a realistic task demanding the simultaneous discovery and translation of linguistic events from continuous, untrimmed video inputs. To address this challenge, we propose \textbf{StreamSLST}, the first end-to-end framework for streaming \textbf{S}ign \textbf{L}anguage \textbf{S}egmentation and \textbf{T}ranslation via Deformable Transformer architecture. Unlike cascaded approaches that suffer from error propagation, StreamSLST employs a parallel decoding strategy to jointly optimize \textit{temporal segmentation} and \textit{translation} in a single pass. To enable scalable training on long-form videos, we incorporate a \textit{Sentence-aware Sliding Window Sampler} that crops continuous streams into trainable units. We further design a \textit{Tri-modal Visual-Language Pretraining} stage that aligns pose dynamics with textual semantics before joint optimization. To support comprehensive evaluation, we conduct experiments on two real-world streaming datasets, BOBSL and How2Sign, together with two newly constructed synthetic streaming benchmarks, \textit{Streaming-CSL-Daily} and \textit{Streaming-Phoenix-2014T}. The results demonstrate that StreamSLST consistently outperforms cascaded segmentation-then-translation pipelines. Interestingly, our analysis reveals that learned sentence boundaries can outperform ground-truth boundaries for translation, suggesting that joint optimization drives localization toward semantically informative regions rather than exact temporal endpoints. Our work establishes the first robust benchmark for continuous sign language understanding and paves the way for accessible real-world communication interfaces.
From Density Matrices to Phase Transitions in Deep Learning: Spectral Early Warnings and Interpretability
Max Hennick ⋅ Guillaume Corlouer
A key problem in the modern study of Deep Learning is predicting and understanding emergent capabilities in models during training. Inspired by methods for studying reactions in quantum chemistry, we present the ``2-datapoint reduced density matrix" (2RDM). We show that this object provides a basis for computationally efficient, unified observables of phase transitions during training. First is the \textit{spectral heat capacity}, which we prove provides early warning signals for learning events. Second is the participation ratio, which reveals the dimensionality of the underlying reorganization associated with a learning event. Remarkably, the top eigenvectors of the 2RDM are directly interpretable, making it straightforward to study the nature of the transitions. We validate across four distinct settings: deep linear networks, induction head formation, grokking, and emergent misalignment. We then discuss directions for future work using the 2RDM.
From I/O to Code with Discovery Agent
Yihong Dong ⋅ Jiaru Qian ⋅ Haoran Zhang ⋅ Peixu Wang ⋅ Binhua Li ⋅ Zhi Jin ⋅ Yongbin Li ⋅ Ge Li ⋅ Xiaokang Yang ⋅ Xue Jiang
The automatic synthesis of a program from any form of specification is regarded as a holy grail of computer science. Fueled by LLMs, NL2Code has achieved tremendous success, yet the fundamentally more challenging task of synthesizing programs from input-output behavior, which we refer to as IO2Code, remains largely unsolved. Whereas NL2Code can exploit the semantic alignment between natural language and code acquired during pretraining, IO2Code requires recovering underlying principles from concrete computational behavior, navigating a vast and underspecified hypothesis space. To address this, we propose DIO-Agent, a discovery agent for IO2Code. Our method frames IO2Code as an evolutionary search over discrete program space, in which an LLM serves as the mutation operator and concrete error signals from execution guide each mutation. To prevent the search from wandering into structurally complex yet incorrect dead ends, we introduce the Transformation Priority Premise as a mutation prior that biases the LLM toward the simplest hypothesis consistent with current evidence, progressively escalating from constants to conditionals to iteration only when simpler constructs are insufficient. To facilitate systematic study, we further construct an IO2Code benchmark spanning multiple difficulty levels. Extensive experiments show that DIO-Agent consistently outperforms both traditional program-by-example methods and evolutionary search baselines across all difficulty levels and various LLMs, while substantially surpassing test-time scaling strategies with equivalent sampling budgets. Our code and dataset are available at \url{https://anonymous.4open.science/r/IO2Code}.
From Outcome to Representation: Tracing Reasoning Mechanisms through Integrated Policy Gradient
Changming Li ⋅ Kaixing Zhang ⋅ Yingdong Shi ⋅ Haoyun Xu ⋅ Zheng Zhang ⋅ Kaitao Song ⋅ Kan Ren
Large Language Models (LLMs) demonstrate remarkable reasoning capabilities, yet the internal mechanisms driving these multi-step processes remain opaque. Existing mechanistic interpretability approaches often rely on superficial text-pattern co-occurrences or capture only short-term effects, struggling to trace the long-horizon, outcome-driven influence of internal components such as neurons or sparse features in multi-step reasoning. In this paper, we introduce Integrated Policy Gradient (IPG), a novel, training-free framework that brings policy-based attribution into mechanistic interpretability for LLM reasoning. Grounded in outcome-oriented and sequential-influence-aware principles, it backpropagates outcome-level, non-differentiable behavioral signals (e.g., reasoning outcome correctness) through entire inference trajectories, utilizing path integration to isolate outcome-relevant mediators at component level for reliable attribution. This provides a linear approximation of Causal Mediation Analysis (CMA) in reasoning. Empirical evaluations demonstrate that IPG achieves precise localization and reveals mechanistic insights by identifying sparse, concentrated subsets of internal components associated with specific reasoning behaviors. Our method also enables modulation of reasoning capabilities, offering a powerful and reliable framework for understanding LLM reasoning.
From Patches to Trajectories: Privileged Process Supervision for Software-Engineering Agents
Murong Ma ⋅ Tianyu Chen ⋅ Yun Lin ⋅ Shuai Lu ⋅ Qinglin Zhu ⋅ Yeyun Gong ⋅ Zhiyong Huang ⋅ Peng CHENG ⋅ Yan Lu ⋅ Jin Song Dong
Supervised fine-tuning (SFT) on long teacher trajectories is the dominant method for instilling investigation and reasoning capabilities into open software-engineering (SWE) agents. Under SFT, every retained response is an imitation target, so the student inherits not only the trajectory's outcome but also any flaw in its intermediate steps, including ungrounded leaps and redundant loops. High-quality training data must therefore be jointly effective (each step is grounded and narrows the agent's epistemic gap to the correct fix) and efficient (each step is information-bearing rather than redundant or looping). Existing recipes filter or relabel teacher rollouts using only a binary terminal verifier, which does not directly target these axes and provides no supervision on instances where the teacher fails. Every real issue ships with a developer-authored reference patch $p^\star$ that implicitly testifies to the file paths, runtime behaviors, and conventions a fix presupposes, but the standard pipeline discards it. We propose P2T (Patches-to-Trajectories), which uses $p^\star$ as privileged information during curation, and frames trajectory construction as a bi-objective program over per-step effectiveness and trajectory length. A reverse phase distills $p^\star$ into a latent process graph $G^\star$ of contextual facts and solution milestones, encoding dense intermediate anchors in constructive ordering. A forward phase curate trajectories from blinded teacher continuations, scoring per-step progress against $G^\star$ under a leakage-blocking groundedness check and committing the shortest segments that retain effectiveness. Using only 1.8k curated SWE-Gym instances, P2T improves both axes simultaneously over outcome-filtered SFT and its tool-error-masking variant: on SWE-bench Verified, it lifts Pass@1 by up to +10.8 points while cutting per-instance inference cost by 15%, with consistent gains on SWE-bench Lite and across two teachers. A size-matched ablation and qualitative analysis further isolate per-trajectory quality from data scale.
From Reasoning Chains to Verifiable Subproblems: Curriculum Reinforcement Learning Enables Credit Assignment for LLM Reasoning
Xitai Jiang ⋅ Wenze Lin ⋅ Zihan Tang ⋅ Yang Yue ⋅ Shenzhi Wang ⋅ Gao Huang
Reinforcement learning from verifiable rewards (RLVR) has shown strong promise for LLM reasoning, but typical outcome-based RLVR methods remain inefficient on hard problems. Correct final-answer rollouts are rare, and standard sample-level credit assignment fails to leverage partial reasoning progress embedded in unsuccessful attempts. To address this problem, we introduce **SCRL** (**S**ubproblem **C**urriculum **R**einforcement **L**earning), a curriculum reinforcement learning framework built on verifiable subproblems derived from reasoning chains. Given a reference solution, SCRL derives a series of verifiable subproblems and constructs a subproblem curriculum, with the final subproblem fixed as the original problem. This converts partial progress on hard problems into verifiable learning signals. Algorithmically, we propose *subproblem-level normalization*, a training technique based on RLVR that normalizes rewards independently at each subproblem position within the rollout group. By assigning the resulting advantages to the corresponding answer spans, we enable finer-grained credit assignment without external rubrics or reward models. Our theoretical analysis shows that this subproblem curriculum makes hard problems more learnable by lifting them out of gradient dead zones, with larger relative gains as the original problem becomes harder. Across seven mathematical reasoning benchmarks, SCRL outperforms strong curriculum-learning baselines, yielding +4.1 and +1.9 average-point gains compared to GRPO on Qwen3-4B-Base and Qwen3-14B-Base respectively. On three hard benchmarks (AIME24, AIME25, and IMO-Bench), SCRL further yields point gains of +3.7 in pass@$1$ and +4.6 in pass@$64$ on Qwen3-4B-Base, suggesting improved exploration on hard reasoning problems. Further ablations show that the proposed credit assignment is effective and that the gains do not require highly curated subproblems or strong external generators, highlighting SCRL as a practical curriculum framework that enables fine-grained credit assignment for LLM reasoning.
From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs
Wish Suharitdamrong ⋅ Muhammad Awais ⋅ Xiatian Zhu ⋅ Sara Atito
Multimodal Large Language Models (MLLMs) can listen and see, but how do audio and visual signals actually travel through the network to shape an answer? Despite their growing role in research and real-world applications, the internal pathways through which audio and visual tokens influence the final prediction remain poorly understood. In this study, we examine audio-visual information flow inside Audio-Visual Large Language Models (AVLLMs), tracing how AVLLMs route, utilize, and integrate audio and visual information across two input configurations, audio-visual video and multiple interleaved audio-visual items. We find that for audio-visual video, AVLLMs follow the sequential information flow pathway established for VLMs and VideoLLMs, with audio and visual contribution flowing along this pathway in proportion to the task's reliance on each modality. In settings with multiple interleaved audio-visual items, this routing shifts to different parallel streams. Furthermore, we demonstrate that audio-visual and other token types can be discarded once their information is transferred to LLM, with minimal impact on the model's prediction or even slight improvement, generalizing across multiple tasks and datasets, enabling more efficient inference. These findings hold across multiple models and scales, Qwen2.5-Omni and Video-SALMONN2 Plus at 3B and 7B scales, leading to hypotheses on why these flow structures emerge. Together, these results deliver the first coherent picture of how AVLLMs orchestrate sound and sight inside the network and lay the groundwork for the next wave of interpretability, design, and efficiency advances in audio-visual and broader MLLMs.
From Small to Large: Cross-Scale Graph Domain Adaptation via Local-Global Structural Alignment
Ruiyi Fang ⋅ Gezheng Xu ⋅ QIUHAO Zeng ⋅ Zihao Jing ⋅ Jiale Cai ⋅ Zhihao Li ⋅ Hao Zheng ⋅ Xuanting Xie ⋅ Shuo Wang ⋅ Bingheng Li ⋅ RUIZHI PU ⋅ Zhao Kang ⋅ Charles Ling ⋅ Boyu Wang
Graph domain adaptation (GDA) aims to transfer knowledge from a labeled source graph to an unlabeled target graph. Existing GDA methods typically assume that the source graph provides sufficiently complete information for reliable cross-domain alignment. However, this assumption is often unrealistic in practice, where only a small sampled source graph may be available due to privacy, annotation, storage, or data-collection constraints. In this paper, we study Cross-Scale Graph Domain Adaptation (CSGDA), where adaptation is performed from a limited-scale source graph to a larger target graph. Our empirical results show that CSGDA exhibits larger local and global discrepancies than conventional GDA, as the source graph provides only limited local structures and limited coverage of the global structure. To address this challenge, we propose a local-global structural alignment framework that, at the local level, augments neighborhood structures to compensate for limited local structural information. At the global level, it extracts global structural patterns and aligns source and target distributions. An adaptive GCN-MLP module then integrates the compensated local structures and extracted global patterns for domain alignment and target prediction. Experiments across multiple graph adaptation benchmarks and sampling ratios demonstrate that our method consistently improves adaptation in the presence of severe source-graph incompleteness.
From Solver Trajectories to Teaching Trajectories: Cognition-Aligned Reasoning Distillation
Yashuo Luo ⋅ Qianren Mao ⋅ ZiqiQin ⋅ Junnan Liu ⋅ Haitao Li ⋅ Qili Zhang ⋅ Yuan Gao ⋅ Zhijun Chen ⋅ Xu Wang
Knowledge distillation from large reasoning models (LRMs) often uses raw teacher trajectories as supervision. However, this practice creates a mismatch between the goal of distillation and the nature of teacher trajectories. Advanced LRMs typically improve their reasoning capabilities through reward-driven post-training, such as RLHF. This training optimizes them as strong solvers, rather than as teachers that produce reasoning processes smaller models can easily learn. Since smaller models have different "cognitive capacities" compared to their larger teachers, directly imitating raw teacher trajectories can sometimes be ineffective and often requires a substantial number of high-quality samples. To bridge this mismatch, we propose Cognition-Aligned Reasoning Distillation (CARD), an RL-based framework that transforms solver trajectories into teaching trajectories aligned with the student’s cognition. Avoiding expensive resampling, our framework directly transforms existing trajectories. We optimize CARD with two reward signals: (1) a Distribution Alignment Reward that matches the student’s predictive distribution, and (2) a Logical Coherence Reward that reduces misleading failed reasoning steps. Comprehensive evaluations across multiple student models and settings demonstrate that CARD outperforms distillation from raw teacher trajectories in in-domain settings while achieving substantial gains in out-of-distribution performance.
From Squeezing to Grounding: Visual Guided DPO for Multimodal Hallucination Mitigation
Wenqi Liu ⋅ Hongxin He ⋅ Yunxiao Wang ⋅ Xuemeng Song ⋅ Yupeng Hu ⋅ Yinwei Wei
Direct Preference Optimization (DPO) is widely used to reduce hallucination in Multimodal Large Language Models (MLLMs). Prior studies and our observations show that during DPO training, the log probabilities of chosen and rejected responses can decrease simultaneously, revealing the squeezing effect. For MLLMs, the redistributed probability mass may move toward responses driven by language priors rather than visual evidence, weakening the role of the image and amplifying hallucination. Existing squeezing-aware methods constrain rejected updates with $\mathcal{V}$-usable information or generation confidence, while methods with anchor terms regularize the DPO implicit reward of chosen responses. Neither explicitly grounds the control signal in visual evidence. We propose Visual Guided Direct Preference Optimization (VGDPO), which relaxes negative updates on rejected responses according to multimodal squeezing risk. VGDPO estimates this risk by combining token-level visual dependency with phrase-level hallucination localization, defining a hallucination ratio for each rejected response. We further apply a related reweighting principle to a visual contrastive objective, providing additional supervision for chosen responses under the original image. Experiments on four hallucination benchmarks and three MLLM backbones show that VGDPO improves DPO log probability dynamics and reduces multimodal hallucination while maintaining informative responses.
Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps
Yanke Zhou ⋅ Yiduo Li ⋅ Hanlin Tang ⋅ Maohua Li ⋅ Kan Liu ⋅ Tao Lan ⋅ Lin Qu ⋅ Yuan Yao ⋅ Xiaoxing Ma
Long-context inference in large language models is bottlenecked by the quadratic cost of full attention. Existing efficient alternatives often rely either on native sparse training or on heuristic token eviction, creating an undesirable trade-off among efficiency, training cost, and accuracy. In this work, we show that full-attention LLMs are already intrinsically sparse and can be transformed into highly sparse models with only minimal adaptation. Our approach is built on three observations: (1) only a small subset of attention heads truly requires full long-context processing; (2) long-range retrieval is governed primarily by a low-dimensional subspace, allowing relevant tokens to be retrieved efficiently with a 16-dimensional indexer; and (3) the useful token budget is strongly query-dependent, making dynamic top-$p$ selection more suitable than fixed top-$k$ sparsification. Based on these insights, we propose RTPurbo, which retains the full KV cache only for retrieval heads and introduces a lightweight token indexer for sparse attention. By exploiting the model's intrinsic sparsity, RTPurbo achieves sparsification with only a few hundred training steps. Experiments on long-context benchmarks and reasoning tasks show that RTPurbo preserves near-lossless accuracy while delivering substantial efficiency gains, including up to a 9.36$\times$ prefill speedup at 1M context and about a 2.01$\times$ decode speedup. These results suggest that strong sparse inference can be obtained from standard full-attention training without expensive native sparse pretraining.
FuseAdapt: Adaptation-Space Fusion for Multi-Modal Semantic Segmentation with Missing Modalities
Xin Zhang ⋅ Robby Tan
Multi-modal semantic segmentation benefits from complementary sensors, but this benefit is fragile when some modalities are degraded or missing at test time. Many existing methods rely on dedicated fusion modules or multi-branch cross-modal encoders, and the fusion often becomes unreliable when modalities are absent, leading to substantial performance drops. We propose FuseAdapt, a parameter-efficient framework that performs fusion within the adaptation space of a frozen vision foundation model (VFM), exploring whether VFM priors can serve as an effective shared backbone for multi-modal segmentation. Instead of attaching a standalone fusion network, FuseAdapt assigns each modality a lightweight adaptation branch and performs adaptation-space fusion by routing trainable adaptation updates across modalities inside the frozen backbone. The pretrained VFM remains fixed while modality-specific corrections and explicit cross-modal interaction are learned through the adaptation paths. To handle missing modalities, we further introduce Predictive Modality Modeling (PMM), a direct representation-level objective for cross-modal redundancy. During training, PMM masks a target modality and predicts its embedding from the visible ones, supervised by a teacher target computed under full-modality input. PMM is used only during training and adds no inference-time overhead. Across three multi-modal segmentation benchmarks (MCubeS, DELIVER, MUSES), FuseAdapt consistently improves both modality-complete and modality-incomplete segmentation with a small trainable budget, yielding up to +3.01 mIoU under full modalities and +12.56 mIoU under missing modalities over prior state-of-the-art.
GEAR-Align: Grounding-Evidence-Aware Gradient Routing for Multimodal Alignment
Yu Yongkang ⋅ Haobo Wang ⋅ Meng Chen ⋅ Han Fang ⋅ Xin Wei ⋅ Zhiyu Lin ⋅ Ye Yuan ⋅ Chiming Duan ⋅ Hao Sun ⋅ Haiyang Zhang
Multimodal Large Language Models (MLLMs) are typically aligned with caption-pair data, but not all answer-token supervision is equally visual. In LLaVA-style alignment training, many tokens are predictable from language context alone, yet their losses can still update vision-related parameters. Such language-dominant updates may add noisy supervision for tokens that require visual evidence. To diagnose this vulnerability, we probe the model against visual counterfactuals to reveal the intrinsic grounding requirements of each token, distilling them into two soft indicators: visual necessity and evidence specificity. Guided by this diagnostic insight, we present $\textbf{GEAR-Align}$, a controlled post-training framework that makes supervision destination explicit. During training, it converts token-level visual reliance into parameter-level gradient allocation: all answer tokens update language-side parameters, high-necessity tokens update route-flow parameters, and the high-specificity subset additionally updates visual-evidence parameters. Across two controlled Qwen3-backbone settings, GEAR-Align improves hallucination-sensitive and local-evidence-heavy benchmarks while remaining competitive on general multimodal benchmarks. These results provide evidence for token-specific visual gradients as a useful principle for controlled MLLM post-training.
GEMS-3D: A Large-Scale 3D Gravity, Electrical, Magnetic, and Seismic Earth Simulation Dataset for Multimodal Geophysical Learning
Yonghao Wang ⋅ Meijia Huang ⋅ Wenkai Lu ⋅ Zhuo Jia ⋅ Hao Feng ⋅ Leyuan Fang
Accurate characterization of deep subsurface structures is a fundamental problem in geoscience, but geophysical inversion is inherently non-unique and therefore requires joint interpretation of gravity, magnetic, electromagnetic, and seismic observations. Despite recent advances in deep learning for geophysical inversion and PDE-based scientific modeling, existing large-scale open datasets are largely limited to a single modality and rarely provide physically consistent multi-source responses generated from the same 3D geological target. This limitation hinders the development and evaluation of multimodal joint inversion methods for deep subsurface characterization. We introduce GEMS-3D, a large-scale synthetic benchmark for 3D geological multiphysics learning. GEMS-3D couples four forward modeling engines—$\textbf{g}$ravity, $\textbf{e}$lectromagnetic, $\textbf{m}$agnetic, and 3D acoustic $\textbf{s}$eismic—to generate sample-aligned observations from shared geological realizations. Guided by facies-aware rock-physics relationships, the dataset provides spatially aligned volumes of P-wave velocity ($v_p$), density ($\rho$), resistivity ($\phi$), and magnetic susceptibility ($\chi$) across nine representative classes of complex geological anomalies, including dike swarms, salt domes, and gas reservoirs. Each sample includes gravity responses, magnetic responses, electromagnetic responses, and 3D acoustic seismic data, together with acquisition metadata and anomaly labels. By packaging co-registered property volumes, anomaly support, acquisition metadata, and aligned multiphysics responses as sample-level bundles, GEMS-3D provides a unified benchmark for operator learning, multimodal inverse interpretation, and missing-modality robustness in deep subsurface settings. The anonymized forward-modeling framework and corresponding dataset link is available at https://anonymous.4open.science/r/GEMS-3D-E4CA.
Generalizable Physics Simulation through Compositional Energy Minimization
Alexandru Oarga ⋅ Yilun Du
Learning the dynamics of interacting physical systems is a central challenge in machine learning for the physical sciences. Prevailing architectures such as Graph Neural Networks (GNNs) and feedforward models often fail to explicitly capture the compositional nature of physical laws. Instead of modeling the forces that govern a system, these networks learn to approximate aggregate state transitions. As a result, they lack a causal understanding of the underlying physics interactions, leading to poor generalization in out-of-distribution systems with varying numbers of entities or novel configurations. In this work, we introduce Compositional Potential Minimization (CPM), a framework that casts simulation as an energy minimization procedure over a composition of learned force potentials. Each force component is represented as a separate, interpretable energy landscape, and interaction dynamics emerge from their addition. Inference in CPM is framed as a global energy minimization problem over an arbitrary number of composed functions. Because CPM learns to disentangle separate forces during training, it enables seamless generalization during test time to larger-scale systems and unseen configurations through straightforward compositions of the learned energy components. Our experiments show that CPM significantly outperforms previous state-of-the-art GNN-based solutions in 2D, 3D, and particle-based simulations.
Generalized Influence Functions for Better Model Change Estimates
Hyeonsu Lyu ⋅ Jonggyu Jang ⋅ Sehyun Ryu ⋅ Hyun Jong Yang
Influence functions (IFs) approximate the effect of removing or reweighting training data on a trained model, but their linear approximation often lacks the accuracy required for reliable post hoc model editing. We identify two culprits and propose generalized influence functions (GIFs) and pseudo-LiSSA (p-LiSSA) to address each in turn. First, classical IFs update all parameters indiscriminately, leading to unnecessary updates to data-irrelevant parameters. GIFs instead restrict edits to a \textit{data-relevant subspace} and leave the remaining parameters fixed. We theoretically show that GIFs achieve a tighter error bound than classical IFs, and that proper subspace selection requires avoiding low-curvature directions while maintaining strong alignment with the target gradient. Second, LiSSA iterations are destabilized by negative Hessian eigenvalues, which existing methods handle via $\ell_2$ damping---at the cost of a biased inverse-Hessian-vector product (iHVP) estimate. p-LiSSA eliminates this bias by solving the \textit{restricted inverse-Hessian problem} without such damping. We further show that the presence of low-curvature directions slows p-LiSSA convergence, and that the data-relevant subspace selected by GIFs mitigates this. Across five image and text unlearning benchmarks, GIFs update only 10% of the parameters yet preserve accuracy up to 2.12 percentage points better than IF-based baselines, while improving unlearning performance by up to 2.73 percentage points. Moreover, p-LiSSA improves restricted iHVP estimation by reducing the residual by 9.6$\times$ over existing iHVP baselines, while using 19.8% less memory with comparable computation time.
Generation Navigator: A State-Aware Agentic Framework for Image Generation
Jinming Liu ⋅ Ruoyu Feng ⋅ Yuqi Wang ⋅ Wenjun Zeng ⋅ Xin Jin
Despite rapid advances in text-to-image generation, faithfully realizing user intent remains challenging, often requiring manual multi-turn trial and error. To automate this process, existing systems rely on either simple prompt rewriting or closed-loop agents driven by hand-crafted rules, rather than learning to adapt actions to the evolving generation process. In this paper, we reformulate image generation as a state-conditioned action-making problem and propose Generation Navigator, a multi-turn T2I agent that learns to dynamically steer the generation trajectory and output the next action. However, training this agent via reinforcement learning introduces a critical credit assignment challenge: naively rewarding a trajectory based solely on a single state assigns equal credit to all actions in the rollout, ignores the quality dynamics across turns, and fails to distinguish actions that improve the trajectory from those that degrade it or waste turns without progress. We resolve this with PRE-GRPO (Peak-Retention-Efficiency Group Relative Policy Optimization), a trajectory-level reinforcement learning objective that explicitly rewards discovering a high-quality image (Peak), avoiding subsequent quality degradation across turns (Retention), and minimizing unnecessary turns (Efficiency). Experiments show substantial improvements across benchmarks, reaching a WISE score of 0.90 and 79.06% reasoning accuracy on T2I-ReasonBench.
Generative Actor-Critic with Soft Bridge Policies
Ke He ⋅ Le He ⋅ Shunpu Tang ⋅ Yafei Wang ⋅ Lisheng Fan
Expressive generative policies such as diffusion and flow models are appealing for maximum-entropy (MaxEnt) online reinforcement learning because of their ability to model multimodal and highly non-Gaussian action distributions. However, training effective soft generative policies faces two obstacles that often arise together. First, marginal action densities are often unavailable, so existing methods typically rely on entropy bounds, heuristic proxies or approximations. Second, iterative shared-parameter samplers raise inference cost and require backpropagation through time over repeated network evaluations, increasing memory cost and destabilizing policy optimization. These obstacles motivate us to seek a generative policy that exposes a tractable MaxEnt objective while requiring only a single sampled actor forward pass for action generation. To this end, we propose soft generative actor-critic (SoftGAC), whose actor defines a stochastic bridge from a fixed base latent to a terminal action latent in pre-tanh space. This structured bridge allows us to lift the MaxEnt objective as an analytically tractable path-wise relative-entropy objective against a high-entropy reference process. In practical finite-step implementation, this relative entropy reduces exactly to sampled transition control energy and thus provides principled soft regularization. Moreover, we keep the single-pass actor lightweight by using small step-specific bridge transitions, each evaluated only once per sampled action, while maintaining a parameter budget comparable to strong actor baselines. Extensive experiments on challenging continuous-control benchmarks show that \myalgo attains higher or competitive returns than strong generative policy baselines, including diffusion and flow-matching policies, while staying in the low-latency regime of one-pass actors and avoiding the much higher cost of many-step diffusion samplers, showing considerable improvements in the compute-return tradeoff.
GeoCore-9B: Towards Geo-Aware Generative Foundation Models in Earth Observation
Jeonghyeok Do ⋅ Munchurl Kim
Existing generative models for earth observation (EO) predominantly rely on fine-tuning natural image priors, which limits their scalability and introduces perspective biases that conflict with geospatial constraints. To address this, we introduce GeoCore-9B, a 9-billion-parameter generative foundation model, which is the first of its scale to be trained from scratch exclusively on EO data. Unlike previous EO foundation models, GeoCore-9B is built upon a Flow Matching-based Diffusion Transformer (DiT) and natively conditions generation on text descriptions and continuous geospatial metadata, including ground sample distances, latitudes, and longitudes. To overcome the convergence and spatial disorientation challenges of training at this scale, we propose a Geospatial Semantic Alignment loss. This objective distills structural Earth surface priors (e.g., terrain and urban areas) from a frozen specialist teacher network, constraining the diffusion latent trajectory during training without adding inference overhead. Pre-trained on the global-scale Git-10M dataset, GeoCore-9B demonstrates strong downstream versatility. Beyond standard proxy generative tasks, we show that GeoCore-9B can be effectively adapted for practical EO applications, including highly challenging tasks such as cloud removal and SAR-to-optical cross-modal translation. Extensive evaluations confirm that GeoCore-9B establishes new state-of-the-art performance in both visual fidelity and geographic structural accuracy. Code and pretrained checkpoints will be released upon acceptance.
Geometric Velocity Regularity for Flow Matching on Manifold-Concentrated Data
Shuntuo Xu ⋅ Zhou Yu ⋅ Kenji Fukumizu
Flow matching learns a velocity field whose ordinary differential equation transports a simple reference distribution to a data distribution. Existing theory typically controls this velocity through ambient, isotropic Lipschitz regularity, which becomes extremely pessimistic when data concentrate near a low-dimensional manifold. In this regime, standard error bounds can blow up drastically near the terminal time and fail to explain why flow matching should remain stable. We study flow matching for distributions obtained by smoothing a density on a compact manifold with small Gaussian noise. Our first result is a tube-local Lipschitz regularity theorem showing that, on the relevant neighborhood of the manifold, the population flow-matching velocity has a Jacobian bound that grows only at the inverse effective noise scale. This improves substantially over the stronger singular behavior suggested by existing ambient analyses. Our second result gives a directional description of the velocity Jacobian: tangential directions exhibit mild growth, normal directions are strongly contractive, and tangential-normal coupling is lower order. This reveals that the apparent singularity is highly structured rather than uniformly harmful. These results provide geometric foundations for sharper stability guarantees and suggest a path toward intrinsic-dimension generalization bounds for flow matching on manifold-concentrated data.
Geometry-Aware Zeroth-Order Optimization for Fine-Tuning Quantized LLMs
Shaocong Ma ⋅ Weidong Cai ⋅ Heng Huang
Zeroth-order optimization (ZOO) is a promising approach for memory-efficient fine-tuning of Large Language Models (LLMs). However, scaling ZOO to larger models on a single GPU necessitates aggressive quantization, which often destabilizes training. A primary cause of this instability is that existing methods treat strictly positive quantization scale parameters as Euclidean variables, potentially leading to boundary violations and optimization difficulties. To address this issue, we propose Hyper-Octant Zeroth-order Optimization (HoZO), a geometry-aware framework that formulates fine-tuning as Riemannian optimization on the positive orthant manifold. By constructing a novel geodesically complete metric and deriving the closed-form exponential map, HoZO inherently enforces positivity constraints without bias-inducing projections. On the theoretical side, we prove that HoZO achieves the optimal oracle complexity for zeroth-order methods. On the empirical side, HoZO outperforms baselines across six downstream tasks and three model sizes. Notably, it achieves a $34\times$ memory reduction compared to first-order fine-tuning, enabling the stable fine-tuning of a 4-bit quantized Llama-2-70b model on a single 48GB GPU.
Geometry-Constrained Kolmogorov–Arnold Networks: Learning Edge Geometry via Banach Duality
Senanayak Sesh Kumar Karri
Kolmogorov–Arnold Networks (KANs) replace fixed activations in deep architectures with learnable univariate edge functions, making the choice of edge parametrisation central. Existing variants rely on fixed bases such as splines, polynomials, or Fourier features, which impose a function-space geometry before data are observed. We introduce geometry-constrained KANs, a family of edge activations derived from Banach duality maps in which the geometry itself is learned through a scalar exponent $p > 1$ per edge. This exponent controls the qualitative response: sub-Euclidean values produce sharp threshold-like behaviour, $p = 2$ recovers the linear regime, and larger values produce flatter responses near the origin. Beyond expressivity, the exponent also modulates sensitivity: smaller values reduce amplification of perturbations, providing an implicit regularisation effect without introducing an explicit shrinkage hyperparameter. Across 50 Feynman symbolic regression equations, geometry-constrained KANs match strong fixed-basis baselines on clean data while substantially improving robustness under distribution shift. In particular, they outperform cross-validated spline baselines in 85 of 90 data-scarcity settings and degrade significantly less under noise (3.7× versus 21.6× for splines). Learned exponents are interpretable and stable, revealing consistent geometric structure across equation families and input dimensions.
Packed high-token 3D fields are not just long-sequence modeling: once a volume is packed into tokens, sequence order is no longer the geometry of the sample. This creates a general failure mode for attention over packed geometric data: sequence-index positional mechanisms can learn the memory layout rather than the object. We introduce Pack3D, a geometry-aware local--global attention operator that separates memory layout from attention geometry. Pack3D keeps local neighborhood attention exact, compresses only distant context into contiguous 3D block summaries, and indexes attention with axial RoPE on true 3D token-grid coordinates rather than packed sequence positions. With RoPE applied before block compression, each query interacts with a geometry-aware mixture of block tokens, so compressed global context remains spatially interpretable after packing. A synthetic repacking benchmark makes the failure mode explicit: learned positional embeddings and packed-sequence RoPE fail under layout changes, whereas axial grid-coordinate RoPE remains invariant. On Allen WTC--11 hiPSC microscopy, our main claim-grade domain, Pack3D improves matched learned-position local--global attention by 0.0978 best macro Dice and 0.1248 final macro Dice at 262k tokens and depth 8. In an Allen-15 joint comparison, Pack3D also improves over a parameter-matched TokenGrid 3D U-Net on both best and final macro Dice. Inference-time ablations show that the trained model genuinely uses its compressed summaries, and the fused forward/backward implementation is $6.49\times$ lower latency than dense SDPA at 65k tokens.
GeoMIND: A Benchmark for Spatial Understanding in Robotic Manipulation
Jinghe Wang ⋅ Xinrui Cao ⋅ Duo Wu ⋅ Chenghao Gu ⋅ Yong Zhong ⋅ Linjia Kang ⋅ Tianyi Xiong ⋅ Zhi Wang
Recent advances in robotic manipulation have made rapid progress in mapping visual observations and language instructions to actions. However, reliable manipulation in open-world environments requires robust spatial understanding, the ability to identify targets from complex layouts (spatial reasoning) and adapt to scenes that continuously evolve through interaction (spatial memory). Yet existing embodied benchmarks fall short in evaluating spatial understanding: they often introduce strong visual cues that bypass spatial reasoning, and focus on static scenes that offer little insight into an agent’s capacity for spatial memory. To address this gap, we introduce GeoMind, a systematic benchmark for structured spatial understanding through tabletop manipulation. GeoMind comprises 54 evaluated task variants derived from 27 canonical task designs organized around spatial reasoning and spatial memory. It also incorporates a scalable pipeline that automates the generation and annotation process of task instances with verified targets and executable interactions. Extensive evaluations across diverse embodied agents reveal their consistent weaknesses in spatial understanding. Our findings highlight that these agents struggle to track identities after spatial changes and convert relational or geometric constraints into executable actions.
GeoRad-3D: Factorized Geometry Transport and Residual Radiometry for 3D Radar Nowcasting
YITING LI ⋅ Zihan Zhou ⋅ Shengkai Chen ⋅ Jing Zhang ⋅ Jianyu Liang ⋅ Ka W CHUI ⋅ Chen Chen ⋅ Fayao Liu ⋅ Yifang Yin ⋅ Shili Xiang
Volumetric radar nowcasting asks a model to forecast where storm structures move and how reflectivity evolves through height. The recent \stcgs{} method makes this setting practical by tracking 3D radar volumes with spatiotemporally coherent Gaussian parcels, but its forecasting head is designed to evolve all future Gaussian attributes with a unified deterministic regressor. This is a poor match to convective evolution: geometry is dominated by coherent transport, while strong-echo radiometry is local, intermittent, and uncertain. We present GeoRad-3D, a factorized forecasting head for coherent Gaussian radar states. The key idea is simple: move geometry first, predict the stable radiometric continuation next, model only the residual uncertainty, and form the final nowcast after rendering. For geometry, a height-conditioned shared field captures volume-level advection and mild coherent deformation, while a parcel-wise correction restores local parcel drift, shape change, and cross-height motion. For radiometry, a deterministic base predictor outputs the stable continuation, and conditional flow matching samples residuals around this base to model the local uncertainty concentrated around strong echoes. At inference, multiple aligned futures are rendered and aggregated through a base-anchored quantile summary in radar-product space, where threshold skill is measured. Under the published NEXRAD protocol, GeoRad-3D improves Pool4 CSI over the strongest published baseline by about 45\%, 82\%, and 124\% at 20, 30, and 40 dBZ. It further reduces LPIPS-Radar by 31.0\%, indicating better perceptual fidelity of rendered storm structures.
GeoReason: Bridging Logical Reasoning and Spatial Fidelity in Remote Sensing Segmentation
Cheng Fan ⋅ Han Wu ⋅ Tingzhang Luo ⋅ Minjing Dong ⋅ Yehui Tang ⋅ Jianyuan Guo
Multimodal Large Language Models (MLLMs) have significantly advanced instruction-driven perception through the ``Embedding-as-Mask'' paradigm. However, extending this capability to Remote Sensing (RS) reasoning segmentation remains exceptionally challenging under cluttered geographical contexts and extreme scale variations, often leading to \textit{Spatial Fidelity Degradation} and \textit{Target Misidentification}. We identify a critical yet underexplored bottleneck: the sequential alienation of the segmentation token \texttt{[SEG]} from its initial visual grounding within the autoregressive generation paradigm. Produced only at the end of long reasoning chains, \texttt{[SEG]} relies on visual evidence that has been progressively diluted by linguistic abstraction and autoregressive noise. This information decay is particularly catastrophic for RS targets with minuscule pixel footprints, where even marginal fidelity loss leads to localization failure. To address these bottlenecks, we propose \textbf{GeoReason}, a novel framework designed to fortify logical deduction and preserve spatial integrity. First, we introduce an \emph{Anticipatory Prior} mechanism that injects the \texttt{[SEG]} token directly into the initial user query, shifting it from a terminal output to a primary condition. This paradigm shift transforms the token into a persistent spatial anchor, preventing information decay during long-form linguistic inference. Second, we enhance visual reasoning via a \emph{Location-aware Chain-of-Thought (CoT)} that enforces coarse-to-fine localization using bounding boxes, coupled with a \emph{Latent Patch-to-Token Alignment (PTA) Loss} that explicitly aligns the latent segmentation token with physical RS textures. Extensive experiments demonstrate that GeoReason consistently outperforms state-of-the-art methods on the Earthreason and RRSIS-D benchmarks.
gfnx: Fast and Scalable Library for Generative Flow Networks in JAX
Daniil Tiapkin ⋅ Artem Agarkov ⋅ Nikita Morozov ⋅ Ian Maksimov ⋅ Askar Tsyganov ⋅ Timofei Gritsaev ⋅ Sergey Samsonov
In this paper, we present gfnx, a fast and scalable package for training and evaluating Generative Flow Networks (GFlowNets) written in JAX. gfnx provides an extensive set of environments and metrics for benchmarking, accompanied with single-file implementations of core objectives for training GFlowNets. We include synthetic hypergrids, multiple sequence generation environments with various editing regimes and particular reward designs for molecular generation, phylogenetic tree construction, Bayesian structure learning, and sampling from the Ising model energy. Across different tasks, gfnx achieves significant wall-clock speedups compared to Pytorch-based benchmarks (such as torchgfn library) and author implementations. For example, gfnx achieves up to 55 times speedup on CPU-based sequence generation environments, and up to 80 times speedup with the GPU-based Bayesian network structure learning setup. Our package provides a diverse set of benchmarks and aims to standardize empirical evaluation and accelerate research and applications of GFlowNets.
Glance Before You Tell: Anomaly-Guided 3D Radiology Report Generation with Heat-Conduction Slice Encoders
Lingyu Zhou ⋅ Dingwen Pi ⋅ Zhang Yi ⋅ Jun Wu ⋅ Xiuyuan Xu
Generating accurate radiology reports from 3D medical volumes remains prone to hallucination. Existing grounding strategies often depend on external visual foundation models or on fixed medical phrase queues. The former passes segmenter-derived evidence to generation, where segmentation errors are hard to suppress. The latter has limited coverage of long-tail and compositional abnormalities. We therefore propose HeatRAD, a glance-before-tell framework that first routes suspicious slice evidence into an organ-ordered visual prefix, then generates reports from this localized evidence without explicit grounding supervision. HeatRAD uses a slice-centric signed heat-conduction encoder that performs efficient spectral mixing while separating smooth anatomy from lesion-like residues. \method{} then scores suspicious slices, binds their evidence to coarse organ contexts, and assembles an organ-ordered visual prefix, guiding the decoder toward localized abnormalities instead of redundant volumetric context. Experiments on three 3D radiology benchmarks show consistent gains in clinical fidelity and hallucination-related metrics, supporting anomaly-first evidence selection as an annotation-efficient paradigm for 3D report generation.
Diffusion large language models (LLMs) have recently emerged as a promising alternative to autoregressive LLMs, enabling parallel generation through iterative denoising. Recent guidance methods improve diffusion LLMs by introducing a weak model to guide the prediction of the base model during denoising. However, the influence of the weak model during denoising remains less clear, and computing its prediction requires an additional forward pass at each denoising step. We present in this paper an in-depth analysis of how diffusion LLMs use contextual information during denoising. Our analysis reveals that diffusion LLMs rely more heavily on local context while often underutilizing global context. Based on this observation, we introduce Global Context Guidance (GCG), a feature-space guidance method that enhances global-context utilization without constructing an explicit weak model. Extensive experiments on LLaDA-8B, Dream-7B, and LLaDA-1.5 across diverse benchmarks demonstrate that GCG consistently outperforms existing guidance methods, while incurring much lower inference overhead. We will release our code upon acceptance.
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler Chang ⋅ Catherine Arnett ⋅ Abdelrahman Sadallah ⋅ Abdelrahman Eldesokey ⋅ Abeer Kashar ⋅ Daud Abolade ⋅ Abosede G Olanihun ⋅ Adam Labaran M ⋅ Adhikarimayum Meerajita Sharma ⋅ Aditi Gupta ⋅ Adril P Merin ⋅ Adwoa A Bremang ⋅ Afitab Iyigun ⋅ Afonso Simplício ⋅ Ahmed Essouaied ⋅ Aicha Chorana ⋅ Akhil Eppa ⋅ Akintunde Oladipo ⋅ Akriti Kuri ⋅ Akshay Ramesh ⋅ Aleksei Dorkin ⋅ Alfred M Kondoro ⋅ Alham Aji ⋅ Ali E Çetintaş ⋅ Allan Hanbury ⋅ Alp Niksarli ⋅ Alvaro Arroyo ⋅ Amin Bajand ⋅ Amol Khanna ⋅ Ana Chkhaidze ⋅ Ana C Condez ⋅ Anamaria Roberta Hartl ⋅ Andiswa Mkhonto ⋅ Andrew Hoblitzell ⋅ Angelos Poulis ⋅ Anirban Majumder ⋅ Anjali Chaudhary ⋅ Anna Vacalopoulou ⋅ Annika Simonsen ⋅ Anton Kovalev ⋅ Anupam Nayak ⋅ Ashvanth.S ⋅ Ayodeji Lana ⋅ Ayu Purwarianti ⋅ Bashar Alhafni ⋅ C. Benedict C. Busole ⋅ Bernard Ghanem ⋅ Bharti Nathani ⋅ Biljana Stojanovska Đurić ⋅ Ogundipe B Funmilola ⋅ Bolaotan Agbonile ⋅ Bruce Torres Fischer ⋅ Burak Tutar ⋅ Burcu Çınar ⋅ Cade Kane ⋅ Can Udomcharoenchaikit ⋅ Chadi Helwe ⋅ Cecilia Liu ⋅ Chiamaka G Nwokolo ⋅ Christopher Homan ⋅ Clément Sampebgo ⋅ Cristina España-Bonet ⋅ Cynthia Amol ⋅ Daeyeop Lee ⋅ Dan Saattrup Smart ⋅ Dana Arad ⋅ Daniil Dzenhaliou ⋅ Dasol Choi ⋅ David Liu ⋅ David Semedo ⋅ David Anugraha ⋅ OLAOYE D Oluwaseun ⋅ Deborah Popoola ⋅ Deividas Mataciunas ⋅ Dennis Asamoah Owusu ⋅ Dhyuthy Krishna Kumar ⋅ Diogo Tavares ⋅ Diogo Glória-Silva ⋅ Divyanshu Goyal ⋅ DongGeon Lee ⋅ Estefany Kelly Buchanan ⋅ Ebele N Anajemba ⋅ Ngozi G Egonu ⋅ Elias Herranen ⋅ Eman Nisar ⋅ Emile Anand ⋅ Emuobonuvie M Ajiboye ⋅ Eryawan P Yulianrifat ⋅ Esther Adenuga ⋅ Ewa Rudnicka ⋅ Faith O Itiola ⋅ Fareeha Fayyaz Sheikh ⋅ Fathima Thekkekara ⋅ Fatima Haouari ⋅ Faustin NSENGIYUMVA ⋅ Fenal A Ilasariya ⋅ Filbert A Tjiaranata ⋅ Firas Laakom ⋅ Francesca Grasso ⋅ Francesco Periti ⋅ Francesco Orabona ⋅ Gbenga K Solomon ⋅ Genta Winata ⋅ Nghia G Ngo ⋅ Gloria Udhedhe-oze ⋅ Gonçalo Vinagre ⋅ Gopi N Challagolla ⋅ Gorka Urbizu-Garmendia ⋅ Gouthami Vadithya ⋅ Guijin Son ⋅ Gyan S Mohapatra ⋅ Hafeez Ullah ⋅ Hafsteinn Einarsson ⋅ Hai Hu ⋅ Hamidreza Saffari ⋅ Hamza Zaidi ⋅ Haopeng Zhang ⋅ Harethah Abu Shairah ⋅ Hele-Andra Kuulmets ⋅ Hitesh Patel ⋅ Houda Bouamor ⋅ Hwanjo Yu ⋅ Iben N Debess ⋅ İbrahim Ethem Deveci ⋅ Ikhlasul A Hanif ⋅ Ikhyun Cho ⋅ Inês Vieira ⋅ Inês Calvo ⋅ Ismail O Daud ⋅ Yusuf I Abayomi ⋅ Itay Itzhak ⋅ Ivan Zhelyazkov ⋅ Ivan Belashkin ⋅ Ivan Spada ⋅ Jacob Brinton ⋅ Jafar Isbarov ⋅ Jaka Čibej ⋅ Jan Kocon ⋅ Jan Cuhel ⋅ Jauza A Krito ⋅ Jebish Purbey ⋅ Jennifer Za ⋅ Jennifer Mickel ⋅ Jenny Kunz ⋅ Jessica Ratovondranto ⋅ Jeyarajalingam Varsha ⋅ Jihae Jeong ⋅ Jimena T Dávalos ⋅ Jinu Lee ⋅ Joao Magalhaes ⋅ John S Yi ⋅ Jongin Kim ⋅ Joseph Chataignon ⋅ Joseph Marvin Imperial ⋅ Jubeerathan Thevakumar ⋅ Judith Land ⋅ Iuliia Alekseenko ⋅ Jiang Junchen ⋅ Jungwhan Kim ⋅ Kairit Sirts ⋅ Kamesh R ⋅ Kanda Patrick Tshinu ⋅ Kätriin Kukk ⋅ Kaustubh Ponkshe ⋅ Kavsar Huseynova ⋅ Ke He ⋅ Kenneth Enevoldsen ⋅ Kent J Alvarez ⋅ Kerem Zaman ⋅ Khalil Mrini ⋅ Kian Kyars ⋅ Komal Gour ⋅ Krishnakumar Lainitha ⋅ Krister Kruusmaa ⋅ Kunal Mukherjee ⋅ Kusum Chouhan ⋅ Laura Castro ⋅ Laura M Porrino-Moscoso ⋅ Lenny S Nzambi ⋅ Leshem Choshen ⋅ Lilja Øvrelid ⋅ Lisa Alazraki ⋅ Lovina Ehimen-Ugbede ⋅ Luheerathan Thevakumar ⋅ Luxshan Thavarasa ⋅ Mamadou K KEITA ⋅ Mansi Jangid ⋅ Marco De Santis ⋅ Marcos Garcia ⋅ Marek Suppa ⋅ Mariam D'Ciofalo ⋅ Marii Ojastu ⋅ Marium Atta ⋅ Maryam Sikander ⋅ Maximos Skandalis ⋅ Mehak Soomro ⋅ Mehmet İlteriş Bozkurt ⋅ Melaku Bayu ⋅ Menan Velayuthan ⋅ Michael Leventhal ⋅ Michał Marcińczuk ⋅ Mina Almasi ⋅ Mirna Potočnjak ⋅ Mithil Bangera ⋅ Mohammadamin Shafiei ⋅ MOHIBA ANSARI ⋅ Mridul Sharma ⋅ Mrityunjaya Indoria ⋅ Mughees U Rehman ⋅ Muhammad Ravi Shulthan Habibi ⋅ Murat B Kınay ⋅ Nada Galant ⋅ Naina S Rathore ⋅ Narada Maugin ⋅ Nathalie H Norman ⋅ Nicholas Kluge Corrêa ⋅ Nikola Ljubešić ⋅ Nirmal Thomas ⋅ Nisansa de Silva ⋅ Nisheeth Joshi ⋅ Nitish Ponkshe ⋅ Nizar Habash ⋅ Nneoma C Udeze ⋅ Noel Thomas ⋅ Noémi Ligeti-Nagy ⋅ Odunayo Ogundepo ⋅ Odunayo K Buliaminu ⋅ God'spraise Okechukwu ⋅ Olanrewaju Samuel ⋅ Olga Snissarenko ⋅ Onyinye A Chiemezie ⋅ Orkun Kınay ⋅ osman tursun ⋅ Pablo Rodríguez ⋅ Pablo Gamallo ⋅ Pedro Valente ⋅ Peter Rupnik ⋅ Philip O. EKIUGBO ⋅ Prakhar Agarwal ⋅ Prokopis Prokopidis ⋅ Quadri Yahya ⋅ Rachele Mignone ⋅ Raghav Singhal ⋅ Rahul Raja ⋅ Ram Mohan Rao Kadiyala ⋅ Raphaël Merx ⋅ Rasmus Larsen ⋅ Ratnavel Rajalakshmi ⋅ Rishav Ghosh ⋅ Romina Oji ⋅ Rui P Guerra ⋅ Rushikesh Zawar ⋅ Sa'ad N Bashir ⋅ Saeed Alzaabi ⋅ Sahil Sandeep ⋅ Sai Sandeep Kantareddy ⋅ Saleha Muzammil ⋅ Salsabila Z Pranida ⋅ Sam Buchanan ⋅ Samuel Rutunda ⋅ Sander Land ⋅ Sarah Sulollari ⋅ Sardar Ali ⋅ Kengatharaiyer Sarveswaran ⋅ Saulius Tautvaisas ⋅ Sayambhu Sen ⋅ Sayantani Banerjee ⋅ Sebastien Diarra ⋅ Sewoong Lee ⋅ Shaan Shah ⋅ Shankar Venkitachalam ⋅ Sharifa Djurabaeva ⋅ Sharon Ibejih ⋅ Shivanya S Dutta ⋅ Siddhant Gupta ⋅ Silvia P Suárez ⋅ Sina Ahmadi ⋅ Sivasuthan Sukumar ⋅ Siyuan Song ⋅ Sokratis Sofianopoulos ⋅ Sona E Simon ⋅ Sonja Benčina ⋅ Sophie Gvasalia ⋅ Sphurti K More ⋅ Spyridon Konstantinos Dragazis ⋅ Stefan Milosavljević ⋅ Stephan Kaufhold ⋅ AlRashed ⋅ Surangika Ranathunga ⋅ Taiga Someya ⋅ Taja Kuzman Pungeršek ⋅ Tal Haklay ⋅ Tasi'u Jibril ⋅ Tatsuya Aoyama ⋅ Tea Abashidze ⋅ Terenz Jomar Dela Cruz ⋅ Terra Blevins ⋅ Themistoklis Nikas ⋅ Theresa D Idoko ⋅ Tilek Chubakov ⋅ Tina Munda ⋅ Tobiloba Moses Owoeye ⋅ Tommaso Gargiani ⋅ Uni Johannesen ⋅ Vallerie A Putra ⋅ Vanya BK ⋅ Varvara Arzt ⋅ Vasily Konovalov ⋅ Vasudevan Nedumpozhimana ⋅ Viktória Ondrejová ⋅ Viktoryia Horbik ⋅ Vuk Dinić ⋅ Walelign Sewunetie ⋅ Winston Wu ⋅ Xiaojing Zhao ⋅ Yacouba Diarra ⋅ Yaniv Nikankin ⋅ Yash Mathur ⋅ Yash Bagla ⋅ Yeshil Bangera ⋅ Yixi Chen ⋅ Yiyuan Li ⋅ Yolanda Xavier ⋅ Yonatan Belinkov ⋅ Zaid Alyafeai ⋅ Zhengyang Shan ⋅ Zhi Rui Tam ⋅ Zilu Tang ⋅ Zuzana Nadova ⋅ Baber Abbasi ⋅ Stella Biderman ⋅ David Stap ⋅ Duygu Ataman ⋅ Fabian D Schmidt ⋅ Hila Gonen ⋅ Jiayi Wang ⋅ David Adelani
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we present Global PIQA, a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world. The 141 language varieties in Global PIQA cover five continents, 19 language families, and 24 writing systems. In the non-parallel split of Global PIQA, over 50% of examples reference local foods, customs, traditions, or other culturally-specific elements. In the parallel split, we translate more "culturally agnostic" commonsense reasoning questions into 131 language varieties, for direct cross-lingual comparisons. In both splits, all examples have been verified by native speakers of the languages. We find that state-of-the-art LLMs perform well on Global PIQA in aggregate, but they exhibit weaker performance in lower-resource languages (e.g. up to a 68% accuracy gap between languages in the parallel split). Global PIQA highlights that in many languages and cultures, everyday knowledge remains an area for improvement in LLMs, alongside more widely-discussed capabilities such as complex reasoning and expert knowledge. Beyond its uses for LLM evaluation, Global PIQA provides a glimpse into the wide diversity of cultures in which human language is embedded.
GraDE: A Graph Diffusion Estimator for Frequent Subgraph Discovery in Neural Architectures
Yikang Yang ⋅ Zhengxin Yang ⋅ Minghao Luo ⋅ Luzhou Peng ⋅ Hongxiao Li ⋅ Wanling Gao ⋅ Lei Wang ⋅ Jianfeng Zhan
Finding frequently occurring subgraph patterns or network motifs in neural architectures is crucial for optimizing efficiency, accelerating design, and uncovering structural insights. However, as the subgraph size increases, enumeration-based methods are perfectly accurate but computationally prohibitive, while sampling-based methods are computationally tractable but suffer from a severe decline in discovery capability. To address these challenges, this paper proposes GraDE, a diffusion-guided search framework that ensures both computational feasibility and discovery capability. The key innovation is the $\underline{\textnormal{Gra}}\textnormal{ph}$ $\underline{\textnormal{D}}\textnormal{iffusion}$ $\underline{\textnormal{E}}\textnormal{stimator}$ (GraDE), which is the first to introduce graph diffusion models to identify frequent subgraphs by scoring their typicality within the learned distribution. Comprehensive experiments demonstrate that the estimator achieves superior ranking accuracy, with up to 114\% improvement compared to sampling-based baselines. Benefiting from this, the proposed framework successfully discovers large-scale frequent patterns, achieving up to 30$\times$ higher median frequency than sampling-based methods.
Gradient Routing Localizes and Removes Unintended Behaviors in RL
Jake Ward ⋅ Shawn Hu ⋅ Aria Wong ⋅ Nathan Hu ⋅ Bryce Woodworth ⋅ Alexander Turner ⋅ Alex Cloud
We propose a novel solution to reward hacking in reinforcement learning. Our setup involves a misspecified reward function corresponding to task completion, and a classifier that flags invalid solutions. As in frontier LLM development, the classifier makes systematic errors. Consequently, reward penalties and data filtering based on the classifier fail to prevent reward hacking. To address this limitation, we introduce GRAFT: Gradient Routing Adapter Fine Tuning. Rather than penalizing or filtering flagged behaviors during training, we constrain (``route'') gradient updates to specific parameters. This has two benefits. First, the model does not learn to exploit classifier errors, i.e., it does not learn to hide the unintended behavior. Second, the unintended behavior can be disabled by ablating the parameters to which they were localized. We demonstrate these benefits in a variety of RL environments where reward penalties result in monitor-subverting policies. Additionally, we show that our method remains effective when we scale model size and task complexity by distilling frontier LLM trajectories into Qwen3-32B. Overall, we show that GRAFT is a promising approach for leveraging unreliable monitors during post-training in environments where true performance is difficult to evaluate.
Graph and Simplicial Complex Prediction Gaussian Process via Hodgelet Representations
Mathieu Alain ⋅ So Takao ⋅ Bastian Rieck ⋅ Xiaowen Dong ⋅ Emmanuel Noutahi
Predicting labels for graph-structured data is crucial in many scientific applications. Recently, Gaussian processes (GPs) with graph-level inputs have been proposed as flexible, non-parametric models for classification tasks. In this work, we extend this framework to regression tasks and simplicial complexes (SCs), enabling edge-level attributes and attributes supported on higher-order simplices. Drawing on the rich literature on Hodge theory in machine learning, we enhance the resulting SC representations via the Hodge decomposition, capturing homological information such as holes. We introduce Hodgelet representations as rich, learnable topological descriptors of simplicial complexes and show that our framework improves predictions across various applications. This paves the way for broader use of GPs in graph-level and SC-level prediction tasks.
Graph Cascades: Contagion-Based Mesoscopic Rewiring for Structure-Aware Graph Machine Learning
Meher Chaitanya Pindiprolu ⋅ My Le ⋅ Luana Ruiz
We introduce **Graph Cascades**, a mesoscopic rewiring strategy for Graph Neural Networks (GNNs) and Graph Transformers (GTs) that captures intermediate-scale graph structure beyond purely local edges or fully global attention. Using contagion-based diffusion processes, Graph Cascades constructs, in $\mathcal{O}(|V|+|E|)$ time, an auxiliary graph where node pairs supported by repeated multi-hop reinforcement are promoted to direct neighbors. We theoretically characterize when reinforcement-based rewiring helps: sufficient conditions under which a reinforcement-based edge-selection rule has higher label agreement than direct adjacency, an SBM witness in which two-hop reinforcement is perfectly homophilic, and formalizing mesoscopic connectivity via graph effective resistance. Empirically, across node-classification benchmarks, Graph Cascades improves multiple GNN and sparse-GT backbones, with the most reliable gains observed on heterophilic and moderate- to high-degree homophilic graphs. The theoretical conditions also identify regimes where mesoscopic rewiring is unlikely to be beneficial -- low-degree regular graphs and graphs with structural bottlenecks -- and these predictions match the observed failures. We additionally observe tight correlations between performance and structural properties in the rewired graphs.
Graph Topology Augmentation for Prioritized Sweeping in Non-stationary Reinforcement Learning
Gia-Hung Pham ⋅ Tuan Dam
Prioritized Sweeping (PS) accelerates model-based reinforcement learning by selecting backups according to residual magnitude. In nonstationary reward settings, however, the canonical priority score is shortsighted: after a localized reward shift, residuals propagate only through realized backups, so bottlenecked or topologically distant state estimates may remain static under a limited replanning budget. We introduce Graph Topology Augmentation for Prioritized Sweeping (GTA-PS), a concrete instance of a broader topology-augmentation principle for priority-based planning. GTA-PS constructs a policy-induced transition graph and augments the residual key with a mixing of regularized directed-Laplacian potentials that diffuses residual information through forward and backward graph structure. The topology weight is controlled by a scheduler based on the Second Largest Eigenvalue Modulus (SLEM), allowing the queue to adapt to the chain's mixing regime. We prove that the forward potential coincides with discounted residual propagation at a canonical regularization parameter and show that GTA-PS gives nonzero priority to states that standard PS can leave blocked after sparse reward shifts. Tabular experiments on FourRooms and GARNET domains demonstrate improved replanning efficiency over standard PS under both exact DP and Dyna-style host planners.
GraphUAT: Uncertainty Attribution in Graph Neural Networks
Burouj Armgaan ⋅ Karan R Singh ⋅ Abhirup Mondal ⋅ Sayan Ranu
Models that recognize when their predictions are unreliable inspire greater trust than those that make confident but incorrect decisions. While recent advances enable uncertainty quantification in GNNs, understanding why a model is uncertain, whether due to noisy features, conflicting neighborhoods, or sparse training coverage, remains unsolved. We address this gap with GraphUAT, the first uncertainty attribution framework for GNNs, which identifies the sources of uncertainty rather than aiming to improve its quantification. GraphUAT employs a teacher-student paradigm, where the trained GNN's behavior is distilled into an uncertainty-aware student that disentangles aleatoric and epistemic uncertainties. We then use targeted probes to trace the student's uncertainty back to the responsible features and structural components. Experiments across benchmarks show that removing the identified elements consistently reduces the teacher's uncertainty, validating the faithfulness of our attributions.
GraphWrit3R: End-to-End Writing of Scene Graphs for Multi-Modal 3D Scenes
Luka Milivojevic ⋅ Nikola Popovic ⋅ Sayan Deb Sarkar ⋅ Sebastian Koch ⋅ Iro Armeni ⋅ Luc V Gool ⋅ Danda Pani Paudel
3D scene graphs provide a structured representation of complex environments by encoding objects, their semantic attributes, and the spatial and functional relationships between them. Current approaches for 3D scene graph generation suffer from several fundamental limitations. They rely on complex multi-stage pipelines with explicit intermediate representations, making systems fragile and prone to error propagation. They assume access to ground-truth object annotations during inference, which deviates from real-world scenarios. They depend on proprietary models, hindering open-source deployment, or incur prohibitively slow inference. We present GraphWrit3R, a simple end-to-end method that takes a 3D point cloud, Gaussian Splats, or a combination of both as input, and directly outputs a complete scene graph as a structured JSON script. The graph lists all objects, their semantic attributes, and the relationships between them, while avoiding all of the above mentioned limitations. The choice of multiple input modalities is purely for versatility, allowing a single set of weights to handle diverse scenarios. Point cloud inputs are encoded via Sonata and Gaussian Splat inputs via Chorus, with both modalities projected onto a shared voxel grid and fused through a novel per-voxel contrastive alignment loss before being decoded by a large language model. As a natural consequence of the LLM, GraphWrit3R also supports open-vocabulary querying. On the 3DSSG benchmark, our method achieves state-of-the-art performance on object class, predicate, and triplet recall, outperforming methods that rely on ground-truth object annotations during inference. We further provide qualitative results and analyze different input modality configurations, contrastive loss formulations, and token fusion strategies. Our code and model will be released upon publication.
Existing Driving VLAs predict trajectories while largely ignoring their visual tokens– a phenomenon we trace not to insufficient training but to a structurally ill-posed task formulation. We show that trajectory recovery, when viewed through the lens of inverse kinematics, requires both a current and a future visual state as boundary conditions; existing VLAs supply only the former, which encourages the model to shortcut through ego status and text commands alone. To address this, we re-design Driving VLA in the style of an inverse kinematics solver. First, a next visual state prediction objective that requires the LLM to predict the future visual scene provides dense visual supervision and suppresses shortcut paths. Second, a separate Inverse Kinematics Network (a cross-attention-based conditional diffusion model) that takes only the current and future visual states as input is designed to suppress reliance on ego status and textual shortcuts during trajectory decoding. With this simple prescription alone, our 0.5B-scale model recovers visual grounding and reaches trajectory planning performance comparable to 7B–8B VLAs more than an order of magnitude larger, on both the closed-loop NAVSIM-v2 and the nuScenes benchmarks. Extensive analysis further shows that this improvement stems from a recovered ability to exploit visual features, with the effect being most pronounced in dynamic driving situations such as turning.
Tensors, which give a faithful and effective representation to deliver the intrinsic structure of multi-dimensional data, play a crucial role in an increasing number of signal processing and machine learning problems. However, tensor data are often accompanied by arbitrary signal corruptions, including missing entries and sparse noise. A fundamental challenge is to reliably extract the meaningful information from corrupted tensor data in a statistically and computationally efficient manner. This paper develops a scaled gradient descent (ScaledGD) algorithm to directly estimate the tensor factors with tailored spectral initializations under the tensor-tensor product (t-product) and tensor singular value decomposition (t-SVD) framework. With tailored variants for tensor robust principal component analysis, (robust) tensor completion and tensor regression, we theoretically show that ScaledGD achieves linear convergence at a constant rate that is independent of the condition number of the ground truth low-rank tensor, while maintaining the low per-iteration cost of gradient descent. To the best of our knowledge, ScaledGD is the first algorithm that provably has such properties for low-rank tensor estimation with the t-SVD. Finally, numerical examples are provided to demonstrate the efficacy of ScaledGD in accelerating the convergence rate of ill-conditioned low-rank tensor estimation in a number of applications.
GUIGuard-Bench: Toward a General Evaluation for Privacy-Preserving GUI Agents
Yanxi Wang ⋅ Zhiling Zhang ⋅ Wenbo Zhou ⋅ Weiming Zhang ⋅ Jie Zhang ⋅ Qiannan Zhu ⋅ Yu Shi ⋅ Shuxin Zheng ⋅ Jiyan He
As GUI agents increasingly rely on screenshots to perceive and operate digital environments, they may inadvertently expose sensitive information such as identities, accounts, locations, and behavioral traces. While existing benchmarks primarily focus on task completion, grounding, or defenses against third-party attacks, current visual privacy datasets remain largely restricted to static natural images, limiting their ability to capture the contextual dependence and task relevance of privacy risks in GUI task trajectories. To bridge this gap, we introduce GUIGuard-Bench, a first-step benchmark for studying privacy-preserving GUI agents in trajectory-based GUI workflows. GUIGuard-Bench contains 241 real GUI-agent trajectories with 4,080 screenshots across Android and PC environments. Each screenshot is annotated at the region level with privacy bounding boxes, semantic privacy categories, risk levels, and whether the private information is necessary for completing the task. Built on these annotations, GUIGuard-Bench supports three complementary evaluations: privacy recognition, offline planning fidelity under protected screenshots, and the utility impact of different protection strategies. Our results show that current models can often detect whether a screenshot contains private information, but they struggle with fine-grained localization, category recognition, risk assessment, and task-necessity judgment. We also find that closed-source models, exemplified by Claude Sonnet 4.6, can maintain largely consistent planner semantics in Android environments after privacy protection is applied. GUIGuard-Bench serves as a seed benchmark for studying privacy awareness, data minimization, and privacy-utility trade-offs in GUI agents, and is intended to support broader future evaluation efforts.
GUITAR: Structured Failure Diagnosis of GUI Agents via State Transitions
Shaoqing Zhang ⋅ Kehai Chen ⋅ Xuefeng Bai ⋅ Zhuosheng Zhang ⋅ Pengfei Zhang ⋅ Yang Xiang ⋅ Min zhang
Understanding where and why Graphical User Interface (GUI) agents fail is essential for building more reliable systems, yet current evaluation relies on step accuracy, a metric that treats each screen independently and overlooks the underlying structure of GUI environments. This leads to two critical blind spots: (1) functionally equivalent screens are evaluated in isolation, obscuring systematic failure patterns across shared screens; and (2) the long-tailed GUI distribution renders failures on rare but critical screens invisible under standard metrics. To address these issues, we propose \textbf{GUITAR}, a state-centric diagnostic framework that performs structured failure analysis over both states and transitions, using a State Transition Graph (STG) by mapping visually diverse screens to shared functional states. Evaluated with 8 agents across 3 tasks on AndroidControl and 3 tasks on Mind2Web, GUITAR reveals that failures are both structurally concentrated and metrically invisible: 61.9\% of failure instances occur in 20\% states, consistently localizing to a small set of bottleneck states. Furthermore, leveraging bottleneck states for inference guidance yields +2.8\% performance improvement. These findings highlight structure-aware evaluation as necessary for GUI agent diagnosis, and demonstrate that targeted intervention on bottleneck states offers a practical path to improving reliability.
GVCC: Zero-Shot Video Compression via Codebook-Driven Stochastic Rectified Flow
Ziyue Zeng ⋅ Xun Su ⋅ Haoyuan Liu ⋅ Bingyu Lu ⋅ Yui Tatsumi ⋅ Hiroshi Watanabe
At ultra-low bitrates, high-fidelity reconstruction requires sampling plausible videos from the posterior rather than regressing to oversmoothed conditional means. We propose Generative Video Codebook Codec (GVCC), a zero-shot framework in which a pretrained video generative model serves directly as the decoder, and the transmitted bitstream specifies its generation trajectory. Modern rectified-flow video models are typically sampled with deterministic ODE solvers, which leave no per-step stochastic channel for transmitting compressed information. GVCC addresses this by converting the deterministic flow sampler into an equivalent marginal-preserving stochastic process, so that information can be transmitted by encoding the per-step stochastic innovations. Unlike images, videos introduce longer temporal dependencies and more diverse conditioning modes. We instantiate GVCC in three practical modes: Text-to-Video (T2V) without a reference frame, autoregressive Image-to-Video (I2V) with tail latent correction, and First-Last-Frame-to-Video (FLF2V) with boundary-sharing Group of Pictures (GOP) chaining. On UVG, GVCC achieves the lowest LPIPS among evaluated baselines across three representative bitrate regimes (down to ${\sim}$0.003\,bpp), with 65\% LPIPS reduction over DCVC-RT at matched bitrate.
Hierarchical 3D grouping aims to recover scene groups across multiple granularities, from fine object parts to complete objects, without relying on semantic labels or a fixed vocabulary. The main challenge is to transform 2D foundation-model cues into coherent hierarchy supervision and embed that hierarchy in a 3D representation. We propose H2G, a hyperbolic affinity field for hierarchical 3D grouping. Our method derives semantically organized tree supervision by interpreting foundation-model affinities through Dasgupta's objective for similarity-based hierarchical clustering. This supervision is distilled into a single Lorentz hyperbolic feature field, whose geometry is well suited for tree-like branching structures. A hierarchy-aware objective aligns the field with fine-level assignments, coarse object structure, compact feature clusters, and LCA (Lowest Common Ancestor) ordering. This formulation represents multiple grouping levels in one feature space, enabling semantic hierarchical grouping grounded in 2D foundation-model knowledge.
Hallucination as Commitment Failure: Larger LLMs Misfire Despite Knowing the Answer
Jewon Yeom ⋅ Jaewon Sok ⋅ Heejun Kim ⋅ Seonghyeon Park ⋅ Jeongjae Park ⋅ Taesup Kim
Hallucination is often viewed as a direct consequence of missing knowledge: a model answers incorrectly when the correct answer is absent from its generation-time distribution, and correctly when it is present. We test this assumption by introducing a semantic notion of answer availability that aggregates token-level variants expressing the same answer concept, and asks whether the correct concept is already available at the moment the model commits to an answer. Across Qwen and Llama models from 0.8B to 72B in both Instruct and Base variants, 16-47\% of Instruct hallucinations occur with substantial probability mass already on the correct concept, and the rate rises monotonically with scale. Comparing such failures against correct generations with matched semantic support, the distinguishing factor is not whether the correct concept is represented, but how its probability is distributed: correct generations concentrate mass on a single surface form, hallucinations disperse it across alternatives. The same sharpening asymmetry extends across multi-token generation and is detectable in pre-generation hidden states. Together, these results identify a single mechanism: instruction tuning sharpens answer commitment with scale, making helpfulness and confident hallucination two consequences of the same underlying disposition.
HALO-VGGT: Heterogeneity-Aware Lightweight Online Compression Allocator for Efficient VGGT
Xueling Wang ⋅ Yiwen Wang ⋅ Siqi Cai ⋅ Chen Zhang ⋅ guanghui He
Visual geometry grounded transformers are strong feed-forward backbones for multi-view 3D reconstruction, but their global attention becomes the dominant inference bottleneck as views and tokens increase. Existing acceleration methods mainly focus on designing compression operators, while compression strengths are often assigned by uniform schedules or offline calibration, overlooking the heterogeneous sensitivity of layers and attention heads during inference. We therefore propose HALO-VGGT, a lightweight online compression allocator requiring no training or calibration for efficient VGGT-style reconstruction. HALO uses a sampled query-key margin proxy to estimate compression sensitivity at the layer and head levels, and assigns hierarchical compression ratios accordingly. Instead of introducing a new compression operator, HALO complements pruning, merging, and sparse attention mechanisms by deciding where and how aggressively they should be applied. HALO improves the trade-off between speedup and accuracy across datasets and backbones, achieving up to 4.20$\times$ speedup and further improving existing merging and sparsity methods by up to 1.46$\times$.
HAMSTAR: Hamiltonian Structured Inter-Period Refinement for Long-Term Time Series Forecasting
Seohyun Kim ⋅ Jaeyo Chang ⋅ Dong-Joon Lim
Periodic structure is a fundamental basis for long-term time series forecasting (LTSF), but real-world periodic patterns are not simple repetitions. Variations in level, amplitude, and phase within one period can accumulate across subsequent periods and reshape future periodic structures. We therefore argue that explicitly modeling inter-period relations is a core challenge in LTSF. Existing methods often handle these relations implicitly or rely on unconstrained mixing, where meaningful dependencies can be entangled with noisy interactions. Simply increasing mixing expressiveness does not resolve this issue. Motivated by this, we address LTSF through gradual state refinement driven by interactions among periods. The key challenge is to reinforce meaningful inter-period dependencies while suppressing noise and maintaining stable representations during repeated refinement. To this end, we draw inspiration from Hamiltonian dynamics and use its split-state formulation and symplectic-style updates as an architectural bias for stable iterative refinement. We propose HAMSTAR (HAMiltonian STructured period-Aligned Representation for time series), a Hamiltonian-inspired framework for structured period-aligned time series forecasting. HAMSTAR represents time series as period-aligned latent states, decomposes them into split-state components, and uses structured inter-period coupling to transform period relations into refinement signals. Symplectic-style updates then progressively refine these states, enabling stable representation dynamics. Experiments on standard LTSF benchmarks show that HAMSTAR achieves state-of-the-art performance across diverse datasets, and further analyses demonstrate meaningful inter-period modeling and stable latent dynamics.
Harmless in Pieces, Harmful in Motion: Detecting Multi-Agent Jailbreaks
Vishal Pramanik ⋅ Maisha Maliha ⋅ Olivera Kotevska ⋅ Nathaniel D Bastian ⋅ Susmit Jha ⋅ Sumit Kumar Jha
Multi-agent LLM and VLM systems introduce a fragmentation blind spot: an adversary can decompose a harmful objective into individually benign sub-tasks distributed across agents, tools, and shared memory, so that each local interaction appears policy-compliant while the composed workflow is unsafe. Existing defenses classify isolated prompts, monitor single-agent histories, or impose architectural constraints, but none directly detect runtime system-level drift toward harm across the full interaction graph. We introduce CITADEL (Composite Intent Tracking for Agentic Defense via Evolving Latent States), a training-free runtime monitor for multi-turn, multi-agent jailbreaks. CITADEL maintains latent states over agents, tools, and memory stores using a cosine-gated recurrence with no learned parameters; encodes a published safety taxonomy as fixed harm anchors via the same frozen multimodal encoder applied to runtime events; and scores risk through the conjunction of harm-anchor proximity, multi-node participation, and positive temporal drift. We evaluate CITADEL on MA-SafeBench, a new benchmark spanning five multi-agent attack families, three communication topologies, and heterogeneous API-accessible LLM/VLM backbones. CITADEL reduces average attack success rate from 69.3% to 16.9% — a 75.6% relative reduction — at a 2.1% false-positive rate and 28 ms per-event overhead, outperforming per-agent detectors, trajectory-level monitoring, and architectural defenses. The same hyperparameters transfer without retuning to OpenAgentSafety, Agent Security Bench, and MTMCS-Bench, indicating that collective geometric convergence provides an effective runtime signal for detecting distributed jailbreaks in multi-agent systems.
Hearing is Believing? Evaluating and Analyzing Audio Language Model Sycophancy with SYAUDIO
Junchi Yao ⋅ Lokranjan Lakshmikanthan ⋅ Annie Zhao ⋅ Danielle Zhao ⋅ Shu Yang ⋅ Zikang Ding ⋅ Di Wang ⋅ Lijie Hu
Audio Language Models (ALMs) have recently shown strong capabilities in unified reasoning over speech, sound, and natural language; yet we find that they can inherit sycophancy, the tendency to agree with user assertions even when they contradict objective evidence. This failure mode is especially concerning for audio-conditioned reasoning, where a model must preserve evidence from acoustic events, speaker characteristics, and speech rate while responding to potentially misleading user feedback. However, unlike text and vision-language sycophancy, ALM sycophancy has not been systematically studied. We therefore introduce SYAUDIO, the first benchmark dedicated to evaluating sycophancy in ALMs, consisting of 4,319 audio questions spanning Audio Perception, Audio Reasoning, Audio Math, and Audio Ethics. Built upon established audio benchmarks and augmented with TTS-generated arithmetic and moral reasoning tasks, SYAUDIO enables systematic evaluation across multiple domains and sycophancy types with carefully verified data quality, including a human-speaker validation of the TTS pipeline. Using this benchmark, we identify substantial and audio-specific sycophancy patterns under realistic conditions involving noise and speech rate, and further show that supervised fine-tuning reduces misleading susceptibility while decode-time steering reveals controllable hidden-state directions for both misleading susceptibility and correction receptiveness.
Hearing the Unspoken: Simulator-Induced Asymmetric-View Policy Optimization for Proactive Task-Oriented Dialogue
Hongbin Zhang ⋅ Ning Gao ⋅ Yuqin Dai ⋅ Ruiyuan Wu ⋅ Jinpeng Wang ⋅ Rena Gao ⋅ Bingdong Tan ⋅ Shuzheng Gao ⋅ Chaozheng Wang
Proactive task-oriented dialogue (TOD), such as outbound sales, demands a persuasive agent that actively probes the user's concerns and steers the conversation toward acceptance within a bounded number of turns. Yet post-trained LLMs are inherently conservative, and reward-shaping RL (e.g., GRPO) struggles since it only re-weights what an already passive policy samples. We show that conditioning on the user's latent concerns unlocks proactive capability that no amount of sampling can undermine, establishing these concerns as a pivotal training-time signal. To operationalize this finding, we build the \textbf{Cognitive User Simulator}, which models each user as a stratified persona comprising observable external traits and hidden internal concerns. The simulator produces faithful and diverse interactions, while emitting per-turn state dynamics that track persuasion progress. We then introduce \textbf{Simulator-Induced Asymmetric-View Policy Optimization}, which converts the modeled concerns and the simulation state transition into complementary training objectives: (1) \emph{Asymmetric On-Policy Self-Distillation} that transfers concern-aware behavior from a privileged view of the same policy into its deployable, conversation-only view; and (2) \emph{State-Transition Policy Refinement} where the final decision provides trajectory-level advantage while synchronous state transitions refine turn-level credit direction. Across two real-world food-delivery benchmarks (\emph{merchant} and \emph{courier} outbound recruitment), our method matches or surpasses leading proprietary LLMs, outperforms strong RL baselines, and generalizes across various LLM-based user simulators.
HeMeR: Heterogeneous Memory Reconciliation for Embodied Agents via Structured KV Reuse
Saehun Chun ⋅ Wonje Choi ⋅ Jinwoo Jang ⋅ TaeYoon Kwack ⋅ Honguk Woo
Embodied agents make decisions by combining information from heterogeneous memory modules, including declarative memory for explicit knowledge, procedural memory for learned routines, and working memory for current observations. During interaction, the memory contents required for decision making change with the agent's state, goal, and task progress, while many contents remain shared across consecutive decisions. However, existing memory modules are typically accessed through separate interfaces, causing each updated memory context to be processed as a new input to the foundation model. This prevents selective key-value (KV) cache reuse for unchanged memory contents and forces full KV cache recomputation even when much of the memory context remains valid. We introduce HeMeR, an inference framework that performs KV reconciliation between dynamically updated memory contexts and reusable KV states. HeMeR tracks declarative, procedural, and working memory contents across action decisions, preserves KV states for unchanged contents, and reconciles only the states affected by memory-content updates. This enables efficient and stable inference as memory contents evolve during interaction. We evaluate HeMeR on ALFWorld, Habitat, and real-world robotic deployments requiring coordinated use of heterogeneous memory contents. On Real-world, HeMeR achieves $83.20\%$ task success and reduces per-step inference time from $2.94$ to $1.87$ seconds compared with MemRL by preserving compatible KV states as task progress updates the memory context.
HERO: A Heterogeneity-Aware Benchmark Library for Federated Continual Learning
Thinh Nguyen ⋅ Le-Tuan Nguyen ⋅ Minh-Duong Nguyen ⋅ Nhi Trinh ⋅ Anh T Nguyet ⋅ Dung Le ⋅ Kok-Seng Wong
Federated continual learning (FCL) evaluates how distributed clients learn from changing data streams while retaining previously learned knowledge. Existing evaluations are difficult to compare because they often change datasets, task splits, client data splits, task orders, backbones, memory assumptions, and reporting rules simultaneously. We introduce \textbf{HERO}, a heterogeneity-aware benchmark library for FCL. HERO builds benchmark streams by separating three choices that are often coupled, namely the task split, the client data split, and the client task sequence. In HERO-Core, the main comparable benchmark, $\alpha$ controls client data skew and $\rho$ controls task-order mismatch. We evaluate representative FCL methods on CIFAR-100 and TinyImageNet using final average accuracy, average forgetting, and bottom-10\% client accuracy. We also include a graph-based Domain-IL portability case study on OGB-MolPCBA, where scaffold-domain granularity changes the input distribution while the prediction task remains fixed. Our results show that method behavior changes across easy and heterogeneous settings, that average accuracy can hide weak bottom-client performance, that task-order mismatch favors different strategies from synchronized evaluation, and that the same HERO interface can expose domain-shift difficulty beyond image-based FCIL. HERO releases benchmark streams, configurations, method implementations, and reporting scripts to support reproducible and setting-aware FCL evaluation.
Heteroscedastic Variational Last Layers
James Harrison ⋅ John Willes ⋅ Mikkel Jordahn ⋅ Paul Brunzema ⋅ Jasper Snoek
We present a simple, inexpensive, and effective method for heteroscedastic uncertainty quantification in neural networks. We build on Variational Bayesian Last Layers (VBLL), wherein deterministic training objectives are developed for variational inference of the network last layer. In particular, we (1) Introduce t-VBLL layers, which perform variational inference for the aleatoric noise covariance, and (2) Introduce Het-VBLL, a Bayesian last layer scheme to model heteroscedastic noise. These methods are based on novel, analytically tractable evidence lower bounds. We further discuss parameterization and initialization within these models. We show that these novel design elements enable effective uncertainty modeling at minimal additional cost, and substantially improve performance over similar methods such as VBLLs.
Hierarchical Regime-Conditioned Dynamics for Spatiotemporal Graphs
Jeremy Baffou ⋅ Pascal Frossard ⋅ Dorina Thanou
Forecasting the behavior of spatio-temporal systems often requires more than accurate predictions: models should uncover the dynamical modes governing future evolution. Existing temporal graph and latent dynamical models achieve strong forecasting performance, but their latent representations often entangle multiple dynamical patterns, limiting interpretability. Conversely, regime-inference methods explicitly model such modes but typically do not support generative forecasting over graph-structured systems. We introduce HERON, a self-supervised regime-aware generative framework that unifies graph-structured trajectory generation with regime identification. HERON factorizes the latent state into discrete system-level regimes, capturing global dynamical modes, and continuous entity-level dynamics, yielding a modular regime interface compatible with existing spatio-temporal backbones. Crucially, the inferred regime actively conditions system dynamics, enabling regime tracking and controlled simulations under alternative dynamical modes. Without regime annotations, HERON jointly learns forecasts and regime dynamics in a self-supervised manner, recovering meaningful modes across controlled synthetic and real-world datasets while maintaining competitive forecasting performance. These results open a path toward interpretable, controllable generative modeling of spatio-temporal dynamics.
HierRR: Enhancing Instruction Alignment in Open-Vocabulary Indoor Scene Synthesis via Agentic Hierarchical Reasoning and Reflection
Weilin Sun ⋅ Lei Meng ⋅ Lihui Zhang ⋅ Manyi Li ⋅ Xiangxu Meng
Open-vocabulary indoor scene synthesis aims to generate plausible layouts from arbitrary user instructions while ensuring physical feasibility and semantic consistency. Existing methods directly infer spatial relations between objects from instructions and solve for layouts based on predefined rules, achieving progress in physical feasibility. However, they often struggle to precisely align with complex instructions due to insufficient comprehension of spatial relations and limited capability to translate them into layouts. In this paper, we propose a Hierarchical Reasoning and Reflection agent framework (HierRR), which leverages a relational hierarchy to enable multi-granularity spatial relation comprehension and attributable layout reasoning, improving semantic consistency while maintaining physical feasibility. Specifically, HierRR introduces the spatial transformation chain‑of‑thought reasoning module that explicitly models the mapping from spatial relations to executable geometric placements, mitigating semantic drift during reasoning. Furthermore, the semantic- and vision-guided reflection module is employed to iteratively refine spatial relations and layout reasoning from local to global along the hierarchy, effectively bridging the gap between language instructions and geometric layouts. Extensive comparative experiments and ablation studies demonstrate that HierRR generates more plausible and semantically consistent indoor scene layouts than existing methods, while generalizing effectively to complex scenes with numerous objects and diverse spatial constraints.
HiFloat4 Format for Language Model Pre-training on Ascend NPUs
Mehran Taghian Jazi ⋅ Yunke Peng ⋅ Xing Huang ⋅ Yao Wang ⋅ Yaoyuan Wang ⋅ Wei Guo ⋅ Yuanyong Luo ⋅ Tianchi Hu ⋅ JUNSONG WANG ⋅ Xin Wang ⋅ Hu Liu ⋅ Yu Cheng ⋅ Yu Z Wei ⋅ Hongliang Li ⋅ Mehdi Rahimifar ⋅ Lei YAN ⋅ wxuefei ⋅ Zhuang Ma ⋅ Liulei ⋅ Hui Yu ⋅ Anandharaju D Raju ⋅ Hoang Le ⋅ Hei Yi Mak ⋅ Tanzila Rahman ⋅ Shadan Golestan
Training large foundation models at low numerical precision is one of the most promising directions for reducing the compute and memory cost of modern AI. Recent 4-bit floating-point formats such as MXFP4 and NVFP4 can be applied to linear GEMM operations in LLMs, but their limited dynamic range introduces numerical instability that prior work addresses by stacking stabilization mechanisms, typically executed at higher precision and partially eroding the efficiency gains that motivate FP4. In this work, we argue that numerical format design is itself a first-class lever for stable FP4 training, and present the first systematic study of FP4 LLM pretraining on energy-efficient Huawei Ascend NPUs. We compare the recently proposed HiFloat4 (HiF4) format against MXFP4 across both dense (OpenPangu-1B, Llama3-8B) and Mixture-of-Experts (Qwen3-MoE-30B) architectures, executing all linear and expert GEMMs in FP4. HiF4's hierarchical scaling provides enough representational headroom that a single stabilization step suffices to keep the relative loss within 1\% of a full-precision baseline — roughly half the gap of MXFP4, which requires three such mechanisms — while incurring less than 1\% average degradation on downstream tasks. Our results suggest that stable, accurate FP4 training does not require an ever-growing stack of stabilization techniques; it requires the right numerical format.
High-Dimensional Learning Dynamics of Attention-Indexed Models
Yizhou Xu ⋅ Margarita Sagitova ⋅ Florent Krzakala ⋅ Lenka Zdeborová
While attention mechanisms are the cornerstones of modern foundation models, the theoretical understanding remains largely elusive. In this paper, we analyze the attention-indexed model, a rich framework that encapsulates multi-layer and multi-head attention. We first prove that while, in a suitable high-dimensional limit, the macroscopic landscape of the population loss is characterized by a finite set of order parameters, the corresponding population gradient flow unfolds in an infinite-dimensional state space, which we show can be exponentially well-approximated by a finite truncated system. Our dynamical analysis uncovers how standard attention parameterizations act as architectural implicit biases to overcome the curse of the information exponent—a barrier that typically traps the direct optimization of the attention matrix $S$. Specifically, the tied attention ($S=WW^T$) induces an automatic symmetry-breaking geometry at initialization, shielding the gradient flow from uninformative manifolds and achieving weak recovery in $\Theta(1)$ time. For untied attention ($S=UV^T$), we reveal a timescale separation dynamic: efficient weak recovery hinges on whether the fast-timescale evolution of the pre-activation mean successfully breaks initial symmetries.
The trend towards larger training setups has brought a renewed interest in partially asynchronous two-phase optimizers which optimize locally and then synchronize across workers. Additionally, recent work suggests that the one-worker version of one of these algorithms, DiLoCo, shows promising results as a (synchronous) optimizer. Motivated by these studies we present an analysis of LA-DiLoCo, a simple member of the DiLoCo family, on a high-dimensional linear regression problem. We show that the one-worker variant, LA, provides a different tradeoff between signal and noise than SGD, which is beneficial in many scenarios. We also show that the multi-worker version generates more noise than the single worker version, but that this additional noise generation can be ameliorated by appropriate choice of hyperparameters. We conclude with an analysis of SLA --- LA with momentum --- and show that stacking two momentum operators gives an opportunity for acceleration via a non-linear transformation of the ``effective'' Hessian spectrum, which is maximized for Nesterov momentum. Altogether our results show that two-phase optimizers represent a fruitful new paradigm for understanding and improving training algorithms.
Reconstructing lineages from live-imaging microscopy requires linking cell detections across time, including through cell divisions. A common approach is to construct a candidate graph and associate cell segmentations (nodes) across frames. However, these and other existing methods overlook two structural obstacles in candidate tracking graphs: (i) cell divisions entangle distinct lineage paths in the node embedding space, and (ii) edges sharing a node have near-random label agreement, so the candidate-graph topology carries no useful information for graph neural networks to aggregate. We propose the Higher-Order Cell Tracking Transformer (HOCT), an edge-centric architecture in which candidate cell links attend to one another under a 3D geometric prior, resolving both issues. Evaluated on the Cell Tracking Challenge and a bacteria division benchmark, HOCT achieves state-of-the-art results without deep pre-trained image encoders. Moreover, the proposed approach is easier to fine-tune, quickly reducing tracking errors by 59% with 400 annotations in a human-in-the-loop setting, outperforming LoRA fine-tuning of competing transformer baselines (6.75% improvement).
HiGraph: A Large-Scale Hierarchical Graph Dataset for Malware Analysis
Han Chen ⋅ Hanchen Wang ⋅ Hongmei Chen ⋅ Ying Zhang ⋅ Lu Qin ⋅ Wenjie Zhang
The advancement of graph-based malware analysis is critically limited by the absence of large-scale datasets that capture the inherent hierarchical structure of software. Existing methods often oversimplify programs into single-level graphs, failing to model the crucial semantic relationship between high-level functional interactions and low-level instruction logic. To bridge this gap, we introduce HiGraph, the largest public hierarchical graph dataset for malware analysis, comprising over 200M Control Flow Graphs (CFGs) nested within 499K Function Call Graphs (FCGs). This two-level representation preserves structural semantics essential for building robust detectors resilient to code obfuscation and malware evolution. We demonstrate HiGraph’s utility through a large-scale analysis that reveals distinct structural properties of benign and malicious software, establishing it as a foundational benchmark for the community. The dataset and tools are publicly available at https://higraph.org.
Millimeter-wave radar point clouds provide a visually privacy-preserving representation for human motion analysis, but object-related interactions remain difficult to describe because the object evidence that defines the interaction is not always separable or consistently observable in sparse radar returns. We address this problem without direct object detection. Instead, we infer hidden interaction semantics from the temporal structure of observation-conditioned corrections to predicted body motion, motivated by the view that object affordances shape how the body moves. To this end, we propose HIRaD, a radar semantic sensing framework that organizes sparse radar returns into a predictive node-structured latent state of local motion carriers. It represents interaction-relevant cues as corrective residual trajectories between a motion-continuation prior and an observation-conditioned posterior, and compresses these trajectories into a language-aligned semantic bottleneck. A structured prefix conditions a frozen autoregressive language model on both the semantic bottleneck and node-level motion summaries, separating compact interaction-level cues from node-level motion evidence. Experiments show that our proposed method improves radar captioning and interaction understanding under weak object observability over strong baselines, while ablations validate the importance of radar-conditioned prefixing, latent state representation, and residual trajectory modeling.
HireWatch: Evaluating LLM Compliance with U.S. Employment Law Under Contextual Pressure
Huy Nghiem ⋅ Maria Kostylew ⋅ Noam Kolt
The deployment of large language models (LLMs) in hiring pipelines is growing, yet existing evaluations generally overlook an important aspect of this development: the compliance of LLMs with applicable legal rules. To fill this gap, we introduce HireWatch, a benchmark that assesses whether LLM hiring assistants comply with U.S. employment discrimination law (e.g., the Civil Rights Act of 1964). We construct a 200-item bank consisting of interview questions that are either lawful or unlawful under U.S. federal statute, regulation, and enforcement guidance. We evaluate the propensity of 18 proprietary and open-source LLMs to select lawful vs. unlawful questions from this bank across ~100,000 trials under a novel seven-tier contextual escalation ladder, crossing occupational details with organizational pressure. Our findings reveal that LLMs' compliance with employment discrimination law is brittle. While LLMs ordinarily comply with legal rules under different job contexts, they readily violate the law when subject to institutional pressure (e.g., an unlawful manager instruction or company policy). Meanwhile, although the inclusion of explicit legal reminders in prompts can improve legal compliance, it does not altogether eliminate violations of employment discrimination law.
Holistic Scaling Laws for Optimal Mixture-of-Experts Architecture Optimization
Weilin Wan ⋅ Jingtao Han ⋅ Debing Zhang ⋅ Weizhong Zhang ⋅ Cheng Jin
Scaling laws govern macroscopic resource allocation for LLMs, yet precise architectural configurations for Mixture-of-Experts (MoE) models remain guided by heuristics. Existing MoE scaling studies either incorporate MoE-specific variables into scaling formulas, causing fitted coefficients to grow rapidly without proportional experimental support, or fix all non-MoE factors, implicitly assuming global architecture does not influence local MoE scaling. We propose a reusable framework for holistic MoE architectural optimization. We first reveal that relying solely on FLOPs per token ($M$) biases MoE evaluation, as heterogeneous Attention/FFN computational densities enable parameter inflation without effective compute gains. We therefore establish a joint constraint triad of $M$, active parameters ($N_a$), and total parameters ($N$) for rigorous MoE characterization. To tame the resulting $\mathcal{O}(n^{16})$ search space, we employ mathematical decoupling: structural constraints and a rank-preserving property of the hidden dimension factorize the optimization into an $\mathcal{O}(n^3) + \mathcal{O}(n^2)$ two-phase search. Through extensive validation across 670+ MoE models spanning $10^{18}$ to $3 \times 10^{20}$ FLOPs, we derive globally applicable scaling laws mapping any compute budget to its optimal architecture. A key finding is that the near-optimal configuration band widens with scale, providing a quantitative basis for trade-offs between scaling law recommendations and engineering constraints. Our work delivers actionable blueprints for optimal MoE design under arbitrary compute budgets.
Holo4D: Holistic 4D Reconstruction as Geometric Control for Video Diffusion
Yushi LAN ⋅ Zeren Jiang ⋅ Koichi Namekata ⋅ Xingang Pan ⋅ Chuanxia Zheng ⋅ Andrea Vedaldi
Reconstructing a dynamic 4D scene along a novel camera trajectory requires both metric geometry from the source video and generative completion of disoccluded target-view regions. Feed-forward 4D reconstruction models recover increasingly comprehensive input-view geometry, including depth, point maps, and dense 3D tracks, but remain tied to the observed camera stream. Camera-controlled video diffusion models (VDMs) provide strong generative priors for novel-view synthesis, yet their camera control is typically not precise enough for metric novel-view reconstruction. We argue that comprehensive input-view 4D reconstruction is an effective geometric interface between these two capabilities: properly exposed to a VDM, it turns reconstruction outputs into camera-accurate novel-view RGB videos that remain geometrically consistent under downstream 4D reconstruction. We introduce HOLO4D, a geometry-aware video-to-video framework that controls a pretrained VDM with a hybrid 4D geometric cache built from the source video. A small set of source views covering the camera motion forms the static cache, while dynamic regions are rendered from the matching source frame. The resulting hybrid RGB-D representation is combined with explicit valid masks, target-camera ray embeddings, and dense 3D tracking tokens. Conditioned on this 4D scaffold, the diffusion backbone synthesizes the target-view RGB video. Under a common downstream 4D reconstructor, our generated videos yield target-view depth more consistent with the held-out reconstruction than that of the baselines, and an additional feature-level readout shows that intermediate VAE features already encode target-view depth. Experiments and ablations further show that progressively richer 3D/4D reconstruction signals improve camera adherence.
HomeFlow: A Data Flywheel for Smart Home Agent Training with Verifiable Simulation
Yi Gu ⋅ Huacan Wang ⋅ Shuo Zhang ⋅ Yuqing Hou ⋅ Xuelei ⋅ weipeng.ming ⋅ Chen Liu ⋅ Fangzhou Yu ⋅ Kuan Li ⋅ Ronghao Chen ⋅ Sen Hu ⋅ Mou X Feng ⋅ Yi Xu
Large language mode agents are moving beyond text-only interaction toward physical-world control, with smart homes as a representative domain. Real domestic interaction requires understanding ambiguous intents, operating in dynamic environments, and performing multi-turn reasoning. However, existing methods struggle to generate high-quality training data for smart home agents. We propose HomeFlow, a verifiable data flywheel for this domain. HomeFlow uses HomeEnv as a unified simulation environment and HomeMaker to procedurally generate diverse home settings. Subsequently, Blueprint compiles open-ended user intents into executable state-based success conditions, while MCTS-Flow synthesizes diverse, verifiable multi-turn trajectories through environment-guided tree search. We then optimize the agents via supervised fine-tuning and step-wise RLVE, which facilitates iterative improvement through authentic physical feedback. We further construct SmartHome-Bench to evaluate the agent across various smart home tasks. On this benchmark, HomeFlow-RL-4B and HomeFlow-RL-8B achieve task success rates of 84.60\% and 87.03\%. It is worth noting that HomeFlow-RL-8B even surpasses the leading GPT-5.5 by 1.23 percentage points.
Horizon Adaptive Offline Policy Learning via Value Stitching
Kexin ZHENG ⋅ Xianyuan Zhan ⋅ Xintao Yan
Learning accurate value functions plays a decisive role for reinforcement learning (RL) agents to solve long-horizon, complex tasks. Conventional temporal-difference (TD) learning objectives suffer from value-estimation bias that accumulates over the horizon, while extended-horizon modeling methods, such as $n$-step TD backups and Q-chunking, adopt a rigid, fixed-horizon value-modeling recipe that is often not flexible enough to capture complex value structures in long-horizon, multi-stage tasks. In this paper, we show that enabling value updates with dynamic horizon composition can yield a strong offline policy learning scheme. Our method, _Horizon Adaptive Offline Policy Learning via **VA**lue **ST**itching_ (**_VAST_**), replaces fixed-horizon backups with recursive, horizon-adaptive value composition. Its key ingredient is to couple value optimization with a future state- and horizon-length-conditioned **_auxiliary value function_** that is learned through direct data supervision, and a **_stitching policy_** that optimally selects the reward-maximizing horizon length and future sub-goal to achieve horizon-adaptive value stitching. This design enables direct estimation and compositional “stitching” of variable-length returns grounded in actionable sub-goal states, providing an accurate and greedily exploitable value-supervision signal for offline policy optimization. Across 50 tasks on OGBench, _VAST_ outperforms fixed-step, extended-horizon methods, and generative-value offline RL baselines, achieving strong performance particularly in high-complexity, long-horizon decision-making tasks.
How are linear representations learned? Exact solutions to the dynamics of abstraction
William W Yang ⋅ Peter E Latham ⋅ Andrew Saxe
In artificial and biological neural networks, concepts are often encoded as consistent linear directions in representation space. In deep learning, this idea is known as the linear representation hypothesis and underpins many interpretability and control methods based on linear probes, from concept detection to activation steering. Yet while prior work has studied whether such directions should exist \textit{after} training, the dynamics of how they emerge \textit{during} training remain poorly understood. Here, we develop a framework to study the alignment of concept directions during training -- a process we call "abstraction". In a minimal linear network setting, we obtain exact solutions for the full trajectory of abstraction. These solutions reveal key analytic principles governing abstraction: (i) data and target geometry jointly determine terminal abstraction, (ii) abstraction improves with network depth, and (iii) initialization scale controls the maximum abstraction reached during training. Extending our theory to nonlinear networks, we analyze how the choice of nonlinearity affects abstraction dynamics: erf networks approximate the linear theory, while abstraction in ReLU networks depends less on target geometry and more on input geometry. Across both, we prove a striking attenuation law: both nonlinearities weaken abstraction in activations relative to preactivations. We find evidence for this law in open models (DINOv3, Gemma 4) and apply our theory to improve linear probe generalization in LLMs. Together, our results provide a dynamical theory of abstraction with implications for interpretability and control.
How Do Agentic LLMs Decide to Call Tools? A Scaffold Default Controlled by Suppression
Xijie Gong ⋅ Tingxu Han ⋅ Jiahao Zhang ⋅ Yuxin Cao ⋅ Ziqi Ding ⋅ Hanqi Yan ⋅ Youcheng Sun ⋅ Lijie Hu
Tool calling, the ability to invoke external tools on demand, is central to agentic LLMs. Within an agent, the LLM decides whether and when to use these tools, greatly extending the agent’s capability boundaries. However, the mechanism that determines whether a model calls a tool or responds directly remains poorly understood. Agentic prompts are often long and heavily scaffolded, combining role instructions, tool schemas, format templates, and user requests across thousands of tokens. This creates a noisy and highly entangled context that makes it difficult to identify a single controllable variable for mechanistic analysis.To obtain such a variable, we propose a method to convert complex agentic prompts into minimal contrastive pairs, in which a single request verb determines the tool-call decision. Replacing an execution verb, such as \textit{write}, with an analysis verb, such as \textit{discuss}, reliably flips whether the model chooses to call a tool. This suggests that the choice is mediated by a compact internal state. 1,500 paired prompts across Python, Java, and C++, where 1,200 are for mechanistic analysis, and 300 are held out for evaluation, are constructed.We trace the tool-calling decision to a single vector, $\mu_\Delta$, and show that it is both causally necessary and sufficient.Transcoder analysis reveals how this vector forms: the scaffold makes tool calling the default, while analysis verbs suppress this default by activating features that signal no tool is needed. Execution verbs activate no analogous tool-use features, so $\mu_\Delta$ reflects the scaffold-induced default rather than a positive execution-verb signal. Downstream attention and MLP layers then read out this vector into the first output token. The mechanism replicates across diverse model families. Our code is available at \url{https://anonymous.4open.science/r/MI4ToolCalling}.
While large language models (LLMs) appear to be increasingly capable of solving compositional tasks, it is an open question whether they do so using compositional mechanisms. In this work, we investigate how feedforward LLMs solve two-hop factual recall tasks, which can be expressed compositionally as $g(f(x))$. We first confirm that modern LLMs continue to suffer from the "compositionality gap", i.e. their ability to compute both $z = f(x)$ and $y = g(z)$ does not entail their ability to compute the composition $y = g(f(x))$. We then decode residual stream representations and identify two processing mechanisms: one which solves tasks *compositionally*, computing $f(x)$ along the way to $g(f(x))$, and one which solves them *directly*, without any detectable signature of the intermediate variable $f(x)$. Finally, we find that embedding space geometry is strongly related to which mechanism is employed, where the idiomatic mechanism is dominant when tasks are represented by translations from $x$ to $g(f(x))$ in the embedding spaces.
How Far Is Too Far? Object Recognition Declines Monotonically with Semantic Distance
Domenic Bersch ⋅ Hai Van Tran ⋅ Gemma Roig
Vision models are sensitive to object-background context, yet it remains unclear how recognition changes as the semantic distance between objects and backgrounds increases. We introduce ImageNet-OOC1k, a benchmark enabling continuous control of object-background relationships via a consensus semantic dissimilarity axis. This formulation reveals a highly consistent monotonic decline in recognition accuracy across 35 models spanning convolutional neural networks, Vision Transformers, and vision-language models. Internal analyses of Vision Transformers reveal aggregation failure, where object evidence is present at the patch level but not consolidated into the final prediction. Leveraging this characterization, a simple spatial re-ranking strategy recovers a subset of these errors and improves accuracy by up to 5.4 percentage points. These results establish context sensitivity as a continuous and predictable property of modern vision systems, with Vision Transformer analyses showing that failures can arise from limitations in aggregating competing signals rather than from missing object representations.
In social choice, anonymity (treating all agents equally) and neutrality (treating all alternatives equally) are widely regarded as "minimal demands" and "uncontroversial" axioms of equity and fairness. However, due to the ANR impossibility theorem, no voting rule can simultaneously satisfy anonymity, neutrality, and resoluteness (always choosing a unique winner). While the worst-case ANR impossibility is well understood, its likelihood in probabilistic models remains underexplored. We address this question through comprehensive likelihood analysis under semi-random models, providing accurate bounds to quantify the advantage of the optimial tie-breaking mechanisms over other commonly-used tie-breaking mechanisms. Our characterizations reveal an $n^{\Theta(m!)}$ improvement in many cases, where $n$ is the number of agents and $m$ is the number of alternatives. We also prove a quantitative ANR impossibility theorem characterizing the tradeoff between equity (anonymity and neutrality) and a new efficiency criterion called $\delta$-anonymity. Our work advances the understanding of trade-offs between equity and efficiency in voting.
How to Train a Surgeon? Benchmarking Generalist Agents in Surgical Scene Understanding
Gengyuan Zhang ⋅ Xiao Han ⋅ Shijie Zhou ⋅ Abdullah Koukash ⋅ Ghazal Ghazaei ⋅ Jindong Gu ⋅ Jingpei Wu ⋅ Yuqicheng Zhu ⋅ Nassir Navab ⋅ Volker Tresp
Evaluating general-purpose multimodal large language models (MLLMs) for surgical scene understanding remains limited by scarce surgical data and fragmented evaluation protocols. We address these challenges through a two-stage evaluation framework for surgical scene understanding. In Stage I, we introduce SURGKNOWBENCH, a diagnostic benchmark that harmonizes public surgical datasets across procedures and tasks and evaluates proprietary and open-weight MLLMs to map the boundary of pretrained surgical competence. SURGKNOWBENCH reveals that current models retain partial visual priors but remain weak on procedural reasoning and workflow assessment. Building on this failure frontier, and motivated by the scaffolded nature of surgical training, we introduce SURGAGENTGYM, a controlled environment where general-purpose MLLM agents can use bounded knowledge search over SURGWIKI and general-purpose visual grounding for surgical tasks. SURGAGENTGYM evaluates whether generalist agents can improve surgical understanding at inference time through tool-mediated evidence use, without additional surgical training. Our results show that agentic inference narrows the gap mainly for perception-heavy tasks, while fine-grained localization and expert-style assessment remain persistent bottlenecks. This opens a path toward future studies of agentic surgical understanding. We have released the codebase and artifacts.
HRIL: Isolating Multimodal Synergy via Higher-Order Dependence
Qun Dai ⋅ Liangjian Wen ⋅ Jiang Duan ⋅ Yong Dai ⋅ Dongkai Wang ⋅ Maolin Wang ⋅ Mingjie Wang ⋅ Jianzhuang Liu ⋅ HE YAN ⋅ Zhao Kang
Self-supervised multimodal representation learning has achieved remarkable success across diverse domains, yet capturing synergistic information remains challenging due to the complexity of cross-modal interactions. Unlike the shared information across individual modalities, synergy arises when task-relevant signals emerge only from the joint configuration of multiple modalities and cannot be recovered from any modality in isolation. This work focuses on how to preserve the information capacity for such synergistic signals in multimodal representations. The key observation is that synergistic information is reflected in higher-order statistical dependence among modalities, which provides a principled target for explicitly modeling joint interactions. Motivated by this insight, we propose Higher-order Representation and Information Learning (HRIL), which constructs an empirical cross-moment tensor over modality embeddings to represent multi-way interactions. HRIL employs Tucker decomposition to disentangle non-separable interactions into a core tensor, complemented by a synergy-aware regularizer that prevents energy concentration and preserves higher-order coupling for synergistic information capture. Experiments on synthetic and real-world benchmarks demonstrate consistent improvements over existing multimodal contrastive methods, with notable gains on tasks dominated by synergistic interactions.
HSCO-Bench: An Agent-Driven End-to-End Hardware-Software Co-design Benchmark for Systems-on-Chip
Pei-Huan Tsai ⋅ Kuan-Lin Chiu ⋅ William Baisi ⋅ Pin-Yu Chen ⋅ Luca P Carloni
Large language models (LLMs) are increasingly adopted for software and hardware design, yet these domains are still evaluated separately. Software benchmarks typically assume fixed hardware targets, while hardware benchmarks focus on component-level optimization without considering the full hardware-software stack. Consequently, no existing benchmark evaluates whether an LLM agent can perform end-to-end, system-level hardware-software co-design. Such a process requires: 1) analyzing applications to identify kernels requiring acceleration, 2) designing and \emph{integrating} heterogeneous accelerators into a System-on-Chip (SoC) under resource constraints, and 3) mapping kernels onto the generated accelerators. We present \hsco{}, an end-to-end hardware-software co-design benchmark for accelerator-rich heterogeneous SoC generation. Built upon an open-source SoC platform with a curated repository structure, HSCO-Bench evaluates the ability of LLMs to jointly optimize software and hardware stacks, producing SoC prototypes deployed on the AMD Virtex-7 FPGA VC707 Evaluation Kit. Experimental results show that end-to-end integration remains challenging for current models, with widespread failures. Among the five frontier models evaluated, only two of them could successfully generate valid SoC prototypes. Yet, even in these successful instances, the generated designs are far from optimal. While we observe a promising peak speedup of 16.22x, the maximum additional resource utilization reaches only 23.67%. This highlights that while state-of-the-art models demonstrate an emerging capability for hardware acceleration, they still heavily underutilize the available hardware capacity, leaving substantial room for future optimization. To the best of our knowledge, HSCO-Bench is the first benchmark targeting this complete co-design flow, enabling LLMs to jointly reason about and modify both the software and hardware stacks of heterogeneous SoCs.
HumanoidArena: Benchmarking Egocentric Hierarchical Whole-body Learning
Taowen Wang ⋅ Zikang Xie ⋅ Bin Yang ⋅ Yunheng Wang ⋅ Zizhao Yuan ⋅ Yuetong Fang ⋅ Yixiao Feng ⋅ Yichi Wang ⋅ Xingyu Chen ⋅ Haodong Chen ⋅ Qiwei Wu ⋅ Weisheng Xu ⋅ Lihan Chen ⋅ Lusong Li ⋅ zeng zecui ⋅ Renjing Xu
Humanoid robots promise whole-body interaction in human-centered environments, but scalable policy learning remains difficult because task-level semantic reasoning and whole-body level dynamic execution are tightly coupled. A practical solution is hierarchical control, where a high-level policy predicts intermediate whole-body actions and low-level general motion trackers (GMTs) execute them as stable humanoid motion. However, existing benchmarks rarely evaluate the policy–tracker interface itself, leaving open whether intermediate whole-body actions are executable, robust under task distribution shifts, and transferable across different GMT backends. We introduce \textsc{HumanoidArena}, a simulation-first benchmark for egocentric hierarchical whole-body learning. The benchmark formulates policy learning as a hierarchical decision making problem: a high-level policy converts egocentric vision, proprioception, and instructions into a compact whole-body action, which is subsequently executed by a low-level GMT. Instead of treating the legs as planar transport tools, \textsc{HumanoidArena} emphasizes interactions where lower-body coordination is structurally necessary in task completion. We therefore design 7 leg-critical HOI/HSI tasks in which success requires foot placement, balance maintenance, posture adjustment, and whole-body reorientation. To further diagnose the hierarchical system, we evaluate policies from two complementary perspectives: perturbation-conditioned generalization and GMT-conditioned transfer. We benchmark representative imitation-learning and VLA-style policies under this shared interface. Experiments show that hierarchical control enables learned policies to solve diverse leg-critical interactions, but performance is strongly tracker-conditioned and cross-GMT transfer remains fragile. These results position \textsc{HumanoidArena} as a benchmark for studying transferable intermediate action representations and scalable egocentric whole-body policy learning. Code and data are available at: \url{https://anonymous.4open.science/r/HumanoidArena-D36F}.
HumanScore: Benchmarking Human Motions in Generated Videos
Tiange Xiang ⋅ Yusu Fang ⋅ Tian Tan ⋅ Narayan Schuetz ⋅ Scott Delp ⋅ Fei-Fei Li ⋅ Ehsan Adeli
Recent advances in model architectures, compute, and data scale have driven rapid progress in video generation, producing increasingly realistic content. Yet, there is no method to systematically assess the fidelity of human figures generated in these videos. In this paper, we present HumanScore, a systematic framework to evaluate the \textbf{quality of human figures and motions} in AI-generated videos. We introduce seven metrics spanning kinematic plausibility and biomechanical consistency. Through carefully designed prompts, we elicit a broad set of movements at varying intensities. We evaluated a total of 840 videos generated by 6 state-of-the-art models. Our framework reveals consistent gaps between visual plausibility and motion fidelity, highlights common failure modes, and ranks models from multiple quantitative and physically meaningful metrics. The proposed metrics are correlated with human evaluations and are simple to compute using standard pose estimators.
Hybrid Reinforcement Learning for One-step Degraded Infrared and Visible Image Fusion
Hongwei Yu ⋅ Daoqing Zha ⋅ Yancheng Bai ⋅ Jiawei Li ⋅ Lei Sun ⋅ Xiangxiang Chu
Infrared-Visible image fusion (IVIF) is essential for robust perception, yet existing methods remain limited under complex degradations. While diffusion models (DMs) offer powerful modeling capabilities to combat such degradations, their iterative process incurs high computational cost. One-step distillation effectively accelerates inference but typically leads to over-smoothed results. To resolve these limitations, we propose a one-step fusion framework with hybrid reinforcement learning, termed OneHRL. We treat degraded inputs as intermediate noisy states along a generative trajectory, enabling a distilled one-step generator to rectify them directly to the clean fusion results. Concurrently, to mitigate the over-smoothed results in distillation, we incorporate the Flow-GRPO paradigm to introduce stochasticity into the deterministic one-step generation. This restores the stochastic exploration necessary for sampling-based reinforcement learning. Furthermore, we design a coarse-to-fine reward that integrates a global semantic reward with an object-level reward. This hybrid reward ensures that the generator captures both naturalness and fine-grained texture. Extensive experiments demonstrate that our method achieves superior fusion quality and efficiency.
HyperSkill: Multi-Modal Skill Learning on the Unit Hypersphere
Erdemt Bao ⋅ JunChen ⋅ Weijun Qin ⋅ Shaopeng Li ⋅ Ming Li ⋅ Mengchen Zhao ⋅ Ziqian Zeng ⋅ Cen Chen ⋅ HUIPING ZHUANG
Language-conditioned manipulation requires skill representations that are both multi-modal and predictable: the same instruction may admit diverse executions, while the high-level planner must select a stable skill for low-level control. However, vector-quantized skill codes compress continuous variations into isolated symbols, introduce biased straight-through gradients, and lack geometry for relating semantically similar skills. We propose $\textbf{HyperSkill}$, a hierarchical skill-learning framework that models skills as continuous embeddings on the unit hypersphere. We use a learnable von Mises--Fisher mixture prior to capture skill-level multi-modality, with geodesic repulsion encouraging diverse mode coverage on the hyperspherical manifold. To connect training-time skill inference with inference-time planning, we introduce Geodesic Skill Alignment (GSA), which aligns posterior skills inferred from demonstration segments with prior skills predicted by a high-level planner using squared geodesic distance. Together, these components provide a geometrically structured skill interface that preserves execution diversity while enabling causal skill prediction at inference time. For low-level control, a Conditional Flow Matching generator produces smooth, multi-modal action trajectories conditioned on skill embeddings. Experiments on LOReL and Kitchen demonstrate that HyperSkill achieves state-of-the-art success rates on the evaluated benchmarks. Ablation studies further validate the importance of each proposed component.
HyrCap: Hybrid Rank-Calibration of Action Proposals for Temporal Event Understanding
Rui Chen ⋅ Xingyu Chen ⋅ Pengxin Xu ⋅ Kai Chen ⋅ Junzhi Yu
Temporal action detection in long untrimmed videos still suffers from inaccurate temporal localization and low-quality proposals, especially under sparse temporal annotations and ambiguous event boundaries. In the context of open-ended video understanding and LLM-driven event analysis, reliable temporal action proposals have emerged as a critical interface bridging temporal localization, event annotation, and semantic understanding. However, conventional temporal action detection (TAD) pipelines lack explicit modeling of proposal localization quality, cross-representation support, and inter-proposal relations, making it difficult to consistently provide high-quality proposals. We propose Hybrid Rank-Calibration of Action Proposals (HyrCap), a proposal-centric query-based TAD framework for temporal event understanding. In HyrCap, Hybrid denotes the use of complementary video features that provide different temporal localization cues, while Rank-Calibration refers to calibrating the ranking confidence of dense action proposals so that it better reflects proposal localization reliability and candidate competition. HyrCap treats dense query-based predictions as an over-complete temporal event hypothesis space and improves the ranking reliability, localization quality, and candidate discrimination of action proposals through modules such as Cross-Representation Consensus Calibration, Localization Quality Estimation, and Proposal Duplication Suppression. Furthermore, inspired by the strong capability of LLMs in fine-grained semantic recognition, event description, and reasoning, we introduce HyrCap for Semantic Action Understanding: calibrated proposals can be used for standard TAD detection and as inputs for LLM-based semantic understanding. We further define a staged evaluation protocol to separately measure event localization ability, segment-level semantic understanding, and their combined event understanding performance. Experiments on public benchmarks show that HyrCap achieves state-of-the-art performance across multiple metrics and provides a more fine-grained paradigm and evaluation protocol for query-based event segmentation and annotation.
IdealCache: Rethinking Cache Scheduling in Diffusion Transformers via Ideal Trajectories
shang xue ⋅ Li Zhu ⋅ Ping Chen
Feature caching accelerates Diffusion Transformers by reusing intermediate computations across denoising timesteps, yet existing methods share a fundamental blind spot: they treat the standard reference denoising trajectory as the optimisation target, approximating its per-step states and thus inheriting its output as an implicit quality ceiling. We argue that this target is arbitrary --- what truly matters is the ideal result, not the intermediate path taken to reach it. We propose IdealCache, a training-free framework that abandons step-wise imitation of the reference trajectory. Instead, it anchors caching on a per-input ensemble of ideal trajectories constructed from candidate target results, aggregates their geometry into a single graph, and discovers optimal non-uniform schedules via a constrained shortest path on this graph. By decoupling from the reference process, IdealCache breaks the performance bottleneck of trajectory-approximation; our schedules can attain near-lossless fidelity in generation and even surpass on task-specific metrics (e.g., maintaining higher semantic consistency in image editing). Without retraining, IdealCache achieves $6.35\times$ speedup on Qwen-Image, $5.84\times$ on Qwen-Image-Edit, and $4.32\times$ on HunyuanVideo, with competitive or superior quality across tasks.
Identified-Set Geometry of Distributional Model Extraction under Top-K Censored API Access
Wenhua Nie ⋅ ZiCheng Zhu ⋅ Jianan Wu ⋅ Binhan Luo ⋅ Haoran Zheng ⋅ Jyh-Shing R Jang
Modern LLM APIs often reveal only top-$K$ logit scores and censor the remaining vocabulary. We study the per-position distribution-recovery limits of this access model. For censoring threshold $\tau$, the compatible teacher distributions form an identified set whose total-variation diameter is exactly $U_K=(V{-}K)\exp(\tau)/(Z_A+(V{-}K)\exp(\tau)),$ where $Z_A$ is the observed partition function. For KL recovery, we give a computable binary-endpoint lower bound and an asymptotically matching small-ambiguity upper bound, with an extension to reference-aware attackers. Experiments on a Qwen3 math-reasoning teacher reveal a layered extraction hierarchy: on-task top-$K$ distillation recovers 12% of private capability, full-logit distillation recovers 56% despite 99% KL closure, and generation-based extraction recovers 96%. Top-$K$ censoring therefore limits per-position distribution recovery but does not by itself prevent capability extraction, separating fidelity from transfer in prompt-only logit distillation.
IFDECORATOR: Wrapping Instruction Following Reinforcement Learning with Verifiable Rewards
Xu Guo ⋅ Tianyi Liang ⋅ Jian Tong ⋅ Xiaogui Yang ⋅ Ling-I Wu ⋅ Chenhui Li ⋅ Zhihui Lu ⋅ Qipeng Guo ⋅ Kai Chen
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising approach to enhance Instruction Following (IF) capabilities of large language models (LLMs). However, RLVR for instruction following remains sample-inefficient and prone to reward hacking, where LLMs exploit verification shortcuts rather than fulfilling the core intent. To address these challenges, we frame RLVR for instruction following as an integrated environment that unifies dynamic task generation and robust training. We introduce Instruction Following Decorator (IFDecorator}, a framework coupling difficulty adaptation, intent alignment, and hack diagnostics. It features: (1) a cooperative-adversarial data flywheel that yields challenging yet solvable tasks for sample-efficient training; (2) IntentCheck, a gating module to mitigate reward hacking; and (3) TripWires, a proactive diagnostic tool for eliciting hacking behaviors and quantifying their prevalence. Extensive experiments show our Qwen2.5-32B-Instruct trained with IFDecorator using only 3,625 training examples achieves 87.43% on IFEval (outperforming GPT-4o) and improves FollowBench by 4.2%, while preserving general capabilities. Diagnostics confirm that our method effectively reduces reward hacking. The approach generalizes across model architectures and scales. We will release code and data for future research.
Implicit Value Probing: Inferring Human Value Preferences via Strategic Multi-Turn Conversations
Tan Wenxin ⋅ pring wong ⋅ Shuo Chen
Recognizing human value preferences is crucial for building reliable AI agents in user-centric fields like psychological counseling and strategic negotiation. Existing benchmarks rely on explicit questionnaires, assuming (1) users are always cooperative and willing to self-report; (2) value preferences remain static. However, humans may conceal true preferences if they feel probed, and may adopt distinct value preferences depending on the context. Consequently, explicit queries are insufficient for predicting value preferences reliably. To address this, we introduce Implicit Value Probing (IVP), a benchmark simulating realistic scenarios where agents implicitly infer value preferences through strategic, multi-turn conversations. IVP grounds these interactions using a diverse set of Persona Agents, covering a wide spectrum of personalities and values. We evaluate mainstream LLMs on IVP, including Qwen3, DeepSeek, Doubao, and GPT-5, observing substantial performance gaps. We further develop Value-Prober-8B, a specialized model internalizing structured cognitive reasoning via distillation. It employs strategic questioning to elicit concealed value preferences. Experimental results demonstrate that Value-Prober-8B can effectively balance probing accuracy with social acceptability, achieving capabilities comparable to frontier models.
Improved Regret Analysis For Parallel Gaussian Process Bandit Optimization
Shion Takeno ⋅ Shogo Iwazaki
This paper studies the regret analysis for parallel Gaussian process (GP) bandit optimization. The known regret upper bounds for the widely used GP batched upper confidence bound and GP batched Thompson sampling (GP-BTS) suffer from a multiplicative factor with respect to the batch size $Q$. To avoid this degradation, existing analyses require a polynomial number of uncertainty sampling (US) for $Q$ at the beginning of optimization. However, this initial US phase is often ineffective in practice. This paper shows that the regret upper bound without the multiplicative factor on $Q$ can be achieved without the initial US phase, using GP-BTS as an example. Furthermore, we show much better regret upper bounds in the noiseless setting than in the noisy setting, as in the sequential GP bandit setting.
Improving Continual Video Instance Segmentation via Spatial-Temporal Balanced Mixture-of-Experts Adapters
Hao Fu ⋅ Sihao Liu ⋅ Hanbin Zhao ⋅ Jiahua Dong ⋅ zhang chao ⋅ Hui Qian
Effective Continual Video Instance Segmentation (CVIS) requires balancing catastrophic forgetting with the plasticity to learn new tasks. Existing methods typically adapt pre-trained VIS models using trainable visual prompts, but suffer from two key limitations: they overlook the distinct degradation of frame-level spatial knowledge (e.g., misclassification) and video-level temporal knowledge (e.g., trajectory disruption), and rely on rigid prompt retrieval that struggles under task ambiguity and distribution shifts. To address these issues, we propose STEAM, a CVIS framework with Spatial-Temporal Balanced Mixture-of-Experts Adapters. Instead of explicit retrieval, we introduce a Distribution-Aware spatial-temporal Scaling and sHifting (DASH) mechanism to align the pre-trained feature space with sequential tasks and reduce the domain gap. Furthermore, our method adaptively learns Task-Adaptive Spatial-Temporal mixture-of-Experts (TASTE) at both the frame and video levels, while updating them with Subspace-Constrained Orthogonal Adaptation (SCOA) to mitigate interference. Extensive experiments across diverse CVIS settings show that STEAM preserves temporal-spatial coherence and achieves state-of-the-art performance.
Improving Flexible Image Tokenizers for Autoregressive Image Generation
Zixuan Fu ⋅ Lanqing Guo ⋅ Chong Wang ⋅ Binbin SONG ⋅ Ding Liu ⋅ Bihan Wen
Flexible image tokenizers aim to represent an image using an ordered 1D variable-length token sequence. This flexible tokenization is typically achieved through nested dropout, where a portion of trailing tokens is randomly truncated during training, and the image is reconstructed using the remaining preceding sequence. However, this tail-truncation strategy inherently concentrates the image information in the early tokens, limiting the effectiveness of downstream AutoRegressive (AR) image generation as the token length increases. To overcome these limitations, we propose \textbf{ReTok}, a flexible tokenizer with \underline{Re}dundant \underline{Tok}en Padding and Hierarchical Semantic Regularization, designed to fully exploit all tokens for enhanced latent modeling. Specifically, we introduce \textbf{Redundant Token Padding} to activate tail tokens more frequently, thereby alleviating information over-concentration in early tokens. In addition, we apply \textbf{Hierarchical Semantic Regularization} to align the decoding features of earlier tokens with those from a pre-trained vision foundation model, while progressively reducing the regularization strength toward the tail to allow finer low-level detail reconstruction. Extensive experiments demonstrate the effectiveness of ReTok: on ImageNet 256$\times$256, our method achieves superior generation performance compared with both flexible and fixed-length tokenizers.
Improving Neural Processes in the Low-Data Regime via Context-Subset Training and Self-Distillation
Hyungi Lee ⋅ Jangho Kim
Neural processes (NPs) are flexible uncertainty-aware meta-learners, but often fail when meta-training tasks are scarce, losing generalization and collapsing into task memorization. We propose Context-Subset Self-Distillation Neural Processes (CSSDNP), a training framework that trains a student on random context subsets while distilling the full-context predictive distribution of an EMA teacher from the same task. This context-asymmetric objective combines reduced-context predictive learning with forward-KL distillation, enforcing conditional consistency across different context views without architectural or test-time changes. Theoretically, we show that context-subset training learns the reduced-context prediction problem while augmenting each task, and that calibrated full-context distillation preserves the asymptotic Bayes target while improving finite-sample estimation via teacher-guided shrinkage. Empirically, CSSDNP improves robustness and generalization across NP variants and tasks, especially in low-meta-data regimes.
Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures
Mark Muchane ⋅ Sean M Richardson ⋅ Kiho Park ⋅ Victor Veitch
Sparse dictionary learning (and, in particular, sparse autoencoders) attempts to learn a set of human-understandable concepts that can explain variation on an abstract space. A basic limitation of this approach is that it neither exploits nor represents the semantic relationships between the learned concepts. In this paper, we introduce a modified SAE architecture that explicitly models a semantic hierarchy of concepts. Application of this architecture to the internal representations of large language models shows both that semantic hierarchy can be learned, and that doing so improves both reconstruction and interpretability. Additionally, the architecture leads to significant improvements in computational efficiency.
Individuals Matter: Improving Deep Multi-View Clustering via Explicit Single-View Enhancement
Hanyang Li ⋅ Yanzheng Wang ⋅ Yaxin Hou ⋅ Zhengxing Jiao ⋅ Hao Chen ⋅ Yuheng Jia
Although existing deep multi-view clustering methods often achieve high performance, they predominantly focus on improving the fused representation while overlooking the enhancement of individual views, often resulting in stagnant or inconsistently improved view-specific representations that limit the final fusion quality. To address this, we propose Explicit single-View Enhancement (EVE), a novel deep multi-view clustering framework designed to explicitly optimize the clustering capability of each individual view. Specifically, EVE employs an attention-based module to generate a global consensus as a reliable anchor, which subsequently guides the refinement of individual views through two key components: feature-level selective alignment and structural-level neighborhood propagation. Crucially, the enhanced single views will further boost the fused representation, and finally leading to a mutually reinforcing closed-loop feedback process. Our method is also naturally applicable to incomplete deep multi-view clustering scenarios. Extensive experiments on multiple benchmark datasets demonstrate that EVE consistently outperforms state-of-the-art competitors, achieving substantial average improvements of 5.9% in ARI across all datasets.The code is available in the Supplementary Material.
Inferring Computational Structure from Neural Recordings with Gain-Modulated Linear Dynamical Systems
Yiteng Zhang ⋅ Zixiong Wang ⋅ Zhengze Wang ⋅ Ke Chen ⋅ Xingyu Li ⋅ Bin Min
Latent dynamical models can accurately fit neural population activity, yet accurate activity fitting alone does not guarantee mechanistic validity. Using synthetic benchmarks, we show that even well-fitted low-rank RNNs can yield misleading circuit interpretations when the prescribed activation function deviates from the ground-truth one. To address this limitation, we introduce gain-modulated linear dynamical systems (gmLDS), which decompose latent dynamics into a state-dependent, unit-wise gain and a static low-rank connectivity matrix, allowing the model to adapt to diverse nonlinear responses without assuming a fixed activation function. Across multiple synthetic benchmarks, gmLDS accurately recovers the local effective connectivity of the underlying model, thereby reconstructing the linearized dynamics along neural trajectories. Applied to neural recordings from perceptual and context-dependent decision-making tasks, gmLDS yields interpretable hypotheses about dynamical structure, including attractor structure in perceptual decision-making and context-dependent selection mechanisms. Together, these results support gmLDS as an effective approach for inferring computational structure from neural recordings.
InfiniteVL: A Systematic Approach to Highly-Efficient, Ultra-Long Multimodal Understanding
Hongyuan Tao ⋅ Bencheng Liao ⋅ Shaoyu Chen ⋅ haoran yin ⋅ Qian Zhang ⋅ Wenyu Liu ⋅ Xinggang Wang
Processing ultra-long multimodal inputs efficiently remains a critical bottleneck for Vision-Language Models (VLMs). This systematic challenge is essentially three-fold: (1) \textbf{Efficient Architecture}: compressing redundant visual sequences without sacrificing fine-grained details; (2) \textbf{Knowledge Transfer}: seamlessly migrating the capabilities of strong pretrained VLMs into more efficient novel architectures; and (3) \textbf{Generalization}: maintaining robust perception and reasoning across diverse multimodal domains. To address this, we introduce \textbf{InfiniteVL}, a linear-sparse hybrid VLM framework designed for efficient long-context understanding. At its core, InfiniteVL pairs a linear architecture to compress long-range visual context with sparse attention to retrieve precise local details. To support this architecture in moderate academic resources, we developed a three-stage knowledge transfer pipeline backed by capability-oriented data construction instead of training from scratch. This ensures smooth architectural alignment, rapid capability recovery, and robust adaptation across diverse domains. Furthermore, we adapt this foundation to specific deployment scenarios, deriving Sparse InfiniteVL for long-video analysis and Streaming InfiniteVL for real-time scene perception. Extensive experiments demonstrate that InfiniteVL matches the performance of leading VLMs while achieving a 1.7$\times$ decoding speedup and a 2.7$\times$ reduction in memory. Notably, Sparse InfiniteVL accelerates prefilling by 5$\times$ at a 256K context length, while Streaming InfiniteVL delivers stable 25 FPS real-time processing with a constant memory footprint.
Information-Theoretic Generalization Bounds for Sequential Decision Making
Futoshi Futami ⋅ Masahiro Fujisawa
Information-theoretic generalization bounds based on the supersample construction are a central tool for algorithm-dependent generalization analysis in the batch i.i.d.~setting. However, existing supersample conditional mutual information (CMI) bounds do not directly apply to sequential decision-making problems such as online learning, streaming active learning, and bandits, where data are revealed adaptively and the learner evolves along a causal trajectory. To address this limitation, we develop a sequential supersample framework that separates the learner filtration from a proof-side enlargement used for ghost-coordinate comparisons. Under a row-wise exchangeability assumption, the sequential generalization gap is controlled by sequential CMI, a sum of roundwise selector--loss information terms. We also establish a Bernstein-type refinement that yields faster rates under suitable variance conditions. The resulting bounds specialize to online learning, streaming active learning with importance weighting, and stochastic multi-armed bandits.
In STeP: Speculative Tensor Parallelism for Concurrent Heterogeneous Inference of LLMs
Viren Luke Radhakrishnan ⋅ Dhruva Kashyap ⋅ Pranav K Nayak ⋅ Chiranjib Bhattacharyya ⋅ Prakash Raghavendra
Running dense, open-weight large language models on commercially accessible workstations or single-GPU cloud setups is increasingly desirable due to cost and privacy constraints, particularly on tasks like coding and reasoning. Starting at around 24B parameters, however, they exceed the available GPU VRAM on these setups, forcing reliance on offloading methods that limit throughput, while leaving substantial CPU compute underutilized. Heterogeneous CPU–GPU speculative decoding is the dominant lossless-acceleration framework in this regime, but existing methods are limited by what we refer to as the **heterogeneity gap**: their CPU and GPU do not co-execute, their draft is a separate set of weights rather than a subnetwork of the verifier, and their VRAM footprint cannot scale to fill the available budget. We observe that channel-saliency methods induce an ordering on the transformer’s FFN channels so that the top-$k$ prefix closely approximates the full output, with graceful degradation as $k$ decreases. Building on this observation, we introduce **Speculative Tensor Parallelism (STeP)**, a training-free, self-speculative method that extracts a subnetwork of the verifier as a GPU-resident draft, places the remainder on the CPU, and verifies via concurrent CPU and GPU computation, provably preserving the verifier’s sampling distribution. Across dense models from 24B to 123B, GPU memory budgets from 32 to 192 GB, and five benchmarks, STeP outperforms SpecExec and SubSpec by up to $1.4 \times$ and $1.8 \times$ respectively, surpasses either tensor parallelism or speculative decoding alone, maximally utilizes the available VRAM, and closes the heterogeneity gap.
Intent2CAD: How Semantic-Parametric Supervision Shapes Text-to-CAD Generation
Jinbo Yang ⋅ Shikai Jing ⋅ Mingyue Yuan
Text-to-CAD generation aims to make CAD modeling more accessible by generating executable CAD programs from natural-language descriptions. However, practical CAD workflows require more than executable geometry: generated programs should also expose reusable structure that supports later modification. We introduce intent-oriented text-to-CAD generation, where design intent is expressed as recoverable code-level signals: semantic features, reusable parameters, and lightweight parametric relations. Rather than changing the model architecture or decoding procedure, we investigate how supervision representation granularity affects both CAD generation and zero-shot intent-preserving editing. To enable this study, we construct IntentCAD-100K, a 100K paired text--code dataset built by converting low-level CAD construction traces into executable CadQuery programs. We further construct IntentCAD-Edit-1K to evaluate zero-shot intent-preserving editing from the same representation perspective. Intent2CAD uses a two-stage conservative lifting pipeline: semantic lifting rewrites reliable primitive groups into feature-level operations, while parametric lifting exposes lightweight parametric relations. Under controlled fine-tuning, semantic lifting mainly improves generation reliability and feature recovery, reducing IR from 2.39% to 0.22% and achieving 99.96% API Recovery, while parametric lifting makes reusable relation structure explicit and further improves relation-aware zero-shot editing.
IntentLens: Grounding Underspecified Multimodal Queries for Recommendation via Tool-Augmented Reasoning
xiao chen ⋅ Dong Fang ⋅ HAITAO LI ⋅ Qing Li
As intelligent assistants are increasingly deployed in real-world environments, recommendations move from passive preference matching to grounding user intent from underspecified multimodal evidence. In these settings, users often provide only a reference image and a fuzzy natural-language query, leaving crucial visual attributes, personalized preferences, and situational constraints implicit. Existing recommenders either rely on historical interactions enriched with multimodal item features or presuppose fully specified textual requests, which are insufficient to resolve such queries in a single-turn setting. We present \textbf{IntentLens}, a tool-augmented recommendation framework that grounds underspecified multimodal queries into explicit, query-relevant evidence before ranking. IntentLens adopts a \emph{ground-then-rank} design: a shared multimodal LLM orchestrates visual, user-memory, and item-attribute tools to (i) recover latent user intent from the image and language query, and (ii) enrich each candidate with fine-grained, query-conditioned evidence; a lightweight ranker then scores candidates over the grounded representations. To support this new setting, we further construct two benchmarks with underspecified multimodal queries, \textsc{GoogleReview-MMTool} and \textsc{Yelp-MMTool}. Comprehensive experiments show that IntentLens outperforms strong MLLM and retrieval baselines.
Decision tree optimization is fundamental to interpretable machine learning, yet most tree learning algorithms are restricted to binary splits. This restriction can produce unnecessarily deep trees when the underlying decision logic is naturally multi-valued. Multiway-split trees address this, but finding sparse yet accurate ones is more challenging, as the number of possible partitions at each node grows combinatorially with the number of feature bins. We introduce SPLINTER (SParse Lookahead for Interpretable N-way Trees by Eliminative Ranking), a dynamic programming and branch-and-bound-based framework that combines efficient candidate generation with pruning bounds to find sparse, accurate multiway-split trees. Our empirical results show that our methods produce trees that achieve higher accuracy at a greater decision sparsity (shorter path lengths) than greedy multiway and near-optimal binary-split tree baselines, while remaining practical to train. We further extend the framework to approximate the Rashomon set of near-optimal multiway-split trees, allowing users to inspect multiple sparse and accurate alternatives.
Intervention-Guided Image-Free Classifier Expansion for Fine-Grained Recognition
Xiangyu Wang ⋅ Changxin Rong ⋅ Yanze Gao ⋅ Lyuzhou Chen ⋅ Xiren Zhou ⋅ Huanhuan Chen
Image-free zero-shot classifier expansion aims to extend recognition to unseen classes by establishing an association mapping between class semantics and visual classifier weights, without accessing original images. However, most existing methods rely on an association-based optimization mechanism, which struggles to meet the demands for fine-grained classification of visually and semantically similar classes. Since co-occurring attributes dominate class semantics, the model gravitates towards co-occurring attributes rather than discriminative attributes. This leads to geometrically entangled classifier weights, resulting in ambiguous decision boundaries. To address this problem, this paper proposes a novel approach that reframes weight generation through the lens of causal intervention. The method first employs comparative reasoning to disentangle discriminative and co-occurring attributes. Subsequently, it synthesizes intervened semantics to impose invariance and sensitivity constraints. This mechanism compels the model to suppress co-occurring attributes while enhancing sensitivity to discriminative attributes. Extensive experiments demonstrate that our method outperforms state-of-the-art approaches, yielding clearer decision boundaries in fine-grained settings.
Intrinsic Muon: Spectral Optimization on Riemannian Matrix Manifolds
Yibang Li ⋅ Bihari L Pandey ⋅ Ravi Sah ⋅ Andi Han ⋅ Cyrus Mostajeran ⋅ Pratik Kumar Jawanpuria ⋅ Bamdev Mishra
Muon and related norm-constrained matrix optimizers have become central to large-scale learning problems. They are formulated as a linear maximization oracle (LMO) over an ambient matrix-norm ball in unconstrained Euclidean space. However, these do not generalize cleanly to manifold-valued parameters such as low-rank factorizations, orthogonality constraints, or symmetric positive definite (SPD) matrices. Naively restricting the Muon LMO to the tangent space (i) breaks quotient symmetries and (ii) couples the tangent-space constraint with an ambient norm bound, thereby obstructing closed-form solutions on various manifolds of interest. We resolve both issues with a single observation: every Riemannian metric canonically lifts a unitarily invariant Euclidean norm to an intrinsic norm on each tangent space, and the resulting intrinsic norm constrained LMO is symmetry preserving. Building on this, we introduce intrinsic Muon (iMuon), a unified framework that yields closed-form updates on the fixed-rank, SPD, Stiefel, and Grassmann manifolds for any unitarily invariant norm, including the spectral, Frobenius, and nuclear norms. We establish convergence guarantees for both deterministic and stochastic iMuon with rate constants that depend only on the manifold dimension. Notably, on the fixed-rank manifold this constant depends only on the rank, making the rate independent of factor conditioning and removing the runtime factor-rescaling required by prior work. Experiments on LoRA finetuning of LLMs, image classification, and subspace learning illustrate the efficacy of the proposed approach.
Inverting the Bellman Equation: From $Q$-Values to World Models
Alistair Letcher ⋅ Mattie Fellows ⋅ Alexander D. Goldie ⋅ Jonathan Richens ⋅ Jakob Foerster ⋅ Oliver Richardson
Model-based and model-free reinforcement learning are traditionally viewed as separate paradigms: while the former learns an explicit model of the transition dynamics $P$, model-free agents typically estimate value functions tied to a specific policy and reward. In this paper, we challenge this dichotomy by proving that value-based agents trained on a sufficiently rich set of reward functions, e.g. using goal-conditioned RL, implicitly encode a unique and accurate world model. To extract this model in practice, we introduce $P$-learning; analogous to $Q$-learning, which approximates $Q^\star$ using samples from the environment, $P$-learning extracts an agent's model of the environment $P^\star$ by sampling from its $Q$-values, policies, and rewards, effectively inverting the Bellman equation. In deterministic MDPs, we prove that the true kernel $P$ can be recovered from an agent trained on a single generic goal for finite state spaces $\mathcal{S}$, and a finite number of Gaussian goals for continuous $\mathcal{S} \subseteq \mathbb{R}^d$, provided $Q$-values are accurate. In the stochastic setting, these conditions generalise to larger sets of goals depending on the reward function family. Even when our assumptions are violated, we empirically demonstrate that agents trained with a small number of sparse rewards encode accurate dynamics in (i) stochastic variants of FourRooms, (ii) MountainCar, and (iii) Reacher. This is further validated by training policies inside the extracted world model, for goals far beyond the training distribution, suggesting that goal-conditioned agents secretly contain implicit generalisation capabilities and providing a new lens into the connection between model-based, model-free, and goal-conditioned RL.
InvestigationWorlds: An Agentic Environment for Legal Investigation
Albert Sun ⋅ Andrew J Benard ⋅ Sil Hamilton ⋅ Anna Teresita Marcelo ⋅ Yong J Kim ⋅ Carl-Leander Henneking ⋅ Rundong Hu ⋅ Yuhong Wang ⋅ David Mimno ⋅ Bishan Yang ⋅ Igor Labutov
We introduce InvestigationWorlds, an agentic environment for legal investigation. We build on an underused artifact of U.S. civil litigation: the summary judgment motion. This motion relies upon a record composed of real evidence exhibits, and results in a court-adopted hypothesis that is treated as ground truth for the purposes of deciding the motion. Each environment is built from a real U.S. Federal Court case retrieved from Public Access to Court Electronic Records (PACER) and augmented by an attorney-validated generation pipeline that synthesizes role-tagged documents around the original record. The resulting corpus admits multiple coherent factual readings, only one of which matches the court-adopted hypothesis. Evaluating on 100 cases, we find agents often commit to incorrect hypotheses despite retrieving relevant evidence, struggling to distinguish the court-adopted hypothesis from alternative hypotheses.
Invisible Ink, Visible Lies: How Production Watermarking Causes LLMs to Hallucinate
Haocheng Ye ⋅ Aoting Hu ⋅ Xinwei Zhang ⋅ Xunzhu Tang ⋅ Shuchao Pang ⋅ Minhui Xue
Text watermarking helps identify AI-generated content, but its effect on factual reliability remains underexplored. In this paper, we study watermarking hallucination: factual errors induced or amplified by watermarking even when the required evidence is present in the context and the unwatermarked model can answer correctly. Using a controlled RAG setting, we compare paired unwatermarked and watermarked generations under the same condition. Across representative watermarking methods, including KGW, SWEET, DiPmark, GumbelSoft, Gumbel-Max, and SynthID-style watermarking, we find that watermarking hallucination is widespread: watermarked outputs can remain fluent while introducing factual errors. We attribute this failure mode to two mechanisms: (1) direct token-level bias, which can suppress fact-consistent tokens, and (2) prefix-induced drift, which accumulates through autoregressive decoding and weakens later attention to factual context. Motivated by this analysis, we propose Fact-Preserving Token Intervention (FPTI) and Fact-Preserving Attention Intervention (FPAI), two plug-in interventions that can be integrated into existing watermarking methods to improve factuality. Our experiments show that combining FPTI and FPAI mitigates around 90\% of watermark-induced hallucinations while preserving fluency and comparable decoding efficiency. Overall, this work highlights factuality as a first-class criterion in watermark evaluation, alongside detectability and robustness, and calls for careful factuality validation before deploying watermarks in fact-critical applications.
IRCasDiff: Two-Stage Cascaded Diffusion for Compound Infrared Face Reconstruction
Zhiyuan Xia ⋅ Haojie Li ⋅ Yiguo Qiao ⋅ Cunjian Chen
Real-world thermal infrared face reconstruction faces the dual challenges of compound sensor-specific degradations and a large infrared-to-visible domain gap. Existing methods either rely on oversimplified degradation assumptions or target only specific degradation types, leaving realistic sensor noise under cross-modal translation largely unaddressed. To tackle this compound problem, we decompose it into two sequential tasks, infrared degradation restoration to remove sensor-specific artifacts, followed by infrared-to-visible translation to bridge the domain gap to the visible spectrum. Both tasks require multi-scale detail reconstruction from spatially heterogeneous tokens that exhibit semantic, frequency, and task-specific variations. This motivates our proposed IRCasDiff, a cascaded diffusion framework consisting of the above two sequential stages. Both stages share a unified backbone with an asymmetric Sparse Mixture-of-Experts design, featuring content-adaptive routing in encoder blocks, frequency-decoupled reconstruction in decoder blocks, and text-free prompt generation for infrared-specific conditioning. The translation stage is initialized from converged restoration stage weights, enabling the model to master infrared restoration before tackling cross-modal mapping. Experiments on MCXFace and SpeakingFaces demonstrate state-of-the-art performance, with FID reductions up to 48.4\% over translation baselines and SSIM improvements up to 19.3\% over restoration baselines.
Irreducible Supervision Enables Compositional Generalization in Post-Training
Ellen Ma ⋅ Nikhil Anand
Reinforcement learning (RL) with verifiable rewards has driven recent gains in large language model (LLM) reasoning, but whether it creates new capabilities or merely sharpens existing ones remains debated, especially for compositional generalization, where models must combine learned primitives into novel multi-step functions. Using a controlled program synthesis domain over a typed domain-specific language (DSL) of 119 primitives with verifiable compositional depth and irreducibility, where the base model has zero capability prior to fine-tuning, we show on Qwen3-8B (and replicate within-family at 4B and 32B, and cross-family on Llama-3.1-8B) that compositional capability is established during supervised fine-tuning (SFT) and depends on irreducibility at two levels of SFT data: structural irreducibility of compositional exemplars and observational irreducibility, where prompt input-output pairs are chosen adversarially to rule out shallower programs. Neither alone is sufficient; even moderate contamination with structurally reducible exemplars substantially degrades performance, and full contamination yields zero pass@64. Given this SFT foundation, standard outcome-reward RL (GRPO, PPO) maintains but does not extend compositional performance, contrasting with prior results in settings with simpler nested function execution. Compositional capability is bottlenecked by structural properties of SFT data, before RL begins. Code release pending institutional IP review.
Is a Linear Probe Evidence of a Linear Representation?
Nadir Kutluozen ⋅ Christian Internò ⋅ David Klindt
The linear representation hypothesis (LRH) holds that high-level concepts are encoded as linear directions in a model's feature space, and grounds much of mechanistic interpretability, from probing to sparse autoencoders to steering vectors. We argue the standard test for it is too lenient. A high linear-probe accuracy shows that a concept is linearly readable from the representation, not that it is one of the representation's privileged directions of variance, which is what the LRH actually claims. These two readings come apart in any high-dimensional, richly structured representation, and the gap matters. The structure that linear probing ignores is precisely where compositional generalization, out-of-distribution behavior, and the validity of post-hoc feature decomposition live. We propose three asymmetric diagnostics that operationalize the strong reading: reverse predictivity, spectral concentration, and out-of-distribution generalization. We validate them in a fully observed simulation where ground truth is known, then apply them to pretrained vision encoders on SVG-World, our deterministic generative benchmark with paired visual styles. Pretraining produces structure beyond what a forward probe can certify, but the linear-feature reading systematically overstates how much. Probing on its own is not enough to license the mechanistic claims that rest on it.
Is Dimensionality a Barrier for Retrieval Models?
Kiril Bangachev ⋅ Guy Bresler ⋅ Jonathan Kogan ⋅ Yury Polyanskiy
Why does the low dimensionality of representations, typically $d\approx 1000$, not prevent modern embedding-based retrieval models from scaling to billions, or even trillions, of data points? To answer this question, we study maximal-margin embeddings in the following retrieval model, classically studied in communication complexity [PS86] and more recently in embedding-based retrieval [WBNL25]. Let $A$ be a binary $N\times n$ matrix indicating whether each of $N$ queries is relevant to each of $n$ documents. We are interested in the largest margin $m>0,$ denoted by $\mathsf{m}^{\mathsf{rd}}(d, A),$ for which there exist unit norm embeddings of the queries and documents $U_1, U_2, \ldots, U_N, V_1, V_2, \ldots, V_n$ with the following property. $\langle U_j, V_i\rangle \ge m$ whenever $A_{ji} = 1$ and $\langle U_j, V_i\rangle \le -m$ when $A_{ji} = 0.$ A large margin is a key proxy for representation quality: it controls both robustness to perturbations and compositional generalization across queries. Our main theorem establishes that the best possible margin without a restriction on the dimension, $\mathsf{m}^{\mathsf{rd}}(+\infty, A),$ can be nearly achieved in dimension $d = O(\mathsf{m}^{\mathsf{rd}}(+\infty, A)^{-2}\log n)$ which improves a theorem of [BDES02]. Together with a matching lower bound in Theorem 1.5, we conclude that when $A$ is the $\binom{n}{k}\times n$ matrix containing all possible $k$-sparse rows once, dimension $d = O(k\log (n/k))$ is necessary and sufficient for the maximal possible margin $\mathsf{m}^{\mathsf{rd}}(+\infty, A) = \Theta(k^{-1/2})$ in this setting. This fully resolves the setup of [WBNL25]. We also give several constructions for large margins when $d = o(k\log (n/k)).$ Our proofs are based on connections between this problem and the literature on compressed sensing and the restricted isometry property. Finally, we empirically test the InfoNCE and sigmoid losses for producing large margin embeddings and demonstrate a clear advantage of the sigmoid loss.
Deep learning models are known for their ability to memorize training data. While memorization is often linked to risks such as privacy leakage and poor robustness, a highly influential claim by~\citet{feldman2020longtail} argues that memorization is actually \textit{necessary} for generalization. Their conclusion is based on the observation that removing points with high memorization scores reduces test accuracy. Upon closer inspection of their work, we uncover four critical flaws in the underlying methodology: \textbf{(1) sampling bias} in their approximation algorithm inflates memorization scores; \textbf{(2) high false positive rate} in their definition of memorization, leads to misclassification of non-memorized points as memorized; \textbf{(3) unprincipled thresholding} that resulting in an ill-posed problem; and \textbf{(4) data leakage} skews the test accuracy results. To address these limitations, we introduce a modifications for correctly identifying and evaluating memorization, including higher sampling rates, modifying the original memorization definition to reduce the false positive rates, proposing a method to identify a principled score threshold, and employing test datasets especially designed to avoid data leakage. Having accounted for these errors, our results show that, in contradiction to the original work, removing truly memorized points does not cause a drop in accuracy, and in most cases, improves test performance. These findings call into question the necessity of memorization in deep learning and highlight the importance of mitigating its risks.
Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models
Hayden Helm ⋅ Xiaodong Liu ⋅ Weiwei Yang
Evaluating and mitigating a generative system's susceptibility to jailbreak attacks is critical to its safe deployment. Given the number of deployable systems, full per-configuration evaluation and optimization is impractical. In this paper, we formalize the behavioral geometry of a population of models that, by leveraging previously evaluated and defended models, supports both efficient susceptibility prediction and effective defense transfer across a population. We apply the framework to 79 models spanning 24 providers and to 100 system configurations of a single base model. Simple methods that use the behavioral geometry reach an AUPRC of $0.94$ for susceptibility detection with $\approx98\%$ fewer probes relative to a full evaluation. Using the behavioral geometry to select which model to transfer an optimized defense from outperforms same-provider assignment ($+2\%$, $p = 0.03$) at no additional probe cost, with a set of three models sufficient to cover the population. Results are robust to hyperparameter selection and judge.
Joint Certification for Attributed Graphs: Beyond Topology-Only Robustness
Blaise Delattre ⋅ Hengyu WU ⋅ Wei Yang Bryan Lim ⋅ Yang Cao
Graph neural networks for bot and fraud detection operate on attributed graphs, where adversaries can jointly modify local edges and continuous node features. We study node-level certification under a localized mixed threat model with up to $d$ edge flips in the target message-passing neighborhood and an $\ell_2$ Frobenius feature budget $\epsilon$. We instantiate hybrid randomized smoothing for this graph setting and derive two graph-specific certificates: a sound Sequential NP composition baseline for the joint topology--continuous-feature threat model, and a Hybrid NP certificate whose homogeneous edge-smoothing structure reduces the adversarial adjacency search to an $O(d^2)$ optimization over edge-deletion and edge-addition counts. Across PubMed, Amazon Fraud, YelpChi, and MGTAB, the Hybrid NP certificate yields non-trivial certified regions under joint perturbations. We also construct adaptive joint attacks and show failure cases on PubMed, Amazon Fraud, MGTAB and YelpChi for topology-only certified nodes, demonstrating that topology-only guarantees can overestimate robustness when continuous feature perturbations are allowed. The results support joint discrete--continuous certification as the appropriate robustness notion for attributed graphs with continuous features.
Joint EM Image Super-Resolution and Segmentation with Semantic and Structural Priors
Jiateng Shou ⋅ Haoyuan Shi ⋅ Zhiwei Xiong
Reconstructing the connectome of the brain is a central goal of modern neuroscience, but the required electron microscopy (EM) imaging is extremely time-consuming. EM image super-resolution (SR) accelerates this acquisition for downstream neuron segmentation, yet existing methods do not adequately couple SR with segmentation for EM: SR lacks segmentation guidance, and segmentation cannot adapt to super-resolved features. In this work, we propose a joint SR and segmentation framework specifically designed for EM images. In the super-resolution part, we introduce segmentation as an auxiliary task via a shared encoder with task-specific decoders to provide a semantic prior, improving reconstruction fidelity. In the segmentation part, we introduce a lightweight partial diffusion model that initializes denoising from a partially noised SR embedding, producing a structural prior that closely approximates real HR features and enhances segmentation accuracy on super-resolved EM images. Extensive experiments across multiple datasets and degradation settings demonstrate the superior performance of our framework.
Fairness in machine learning typically relies on the assumption of clean training data. However, in high-stakes applications like healthcare and criminal justice, privacy mechanisms and proxy variables frequently introduce simultaneous noise into both sensitive attributes and target labels. While existing methods address label or attribute noise in isolation, they fail under this "double noise" regime, often inadvertently amplifying bias. In this paper, we propose $\textbf{Jointly Robust Fairness (JRF)}$, a unified framework that guarantees fairness under simultaneous data corruption. JRF integrates a Forward Loss correction to handle label noise with a novel Matrix-Inverse Robust Loss (MIRL) that algebraically recovers the true demographic parity gap from noisy attribute observations. We provide rigorous theoretical guarantees for our estimator, including a finite-sample concentration bound demonstrating that the sample complexity scales quadratically with the inverse of the noise transition determinant. Furthermore, our framework naturally generalizes to both symmetric and highly asymmetric noise distributions. Extensive experiments on the Adult, Bank, and COMPAS datasets demonstrate that JRF reduces fairness violations by over 91\% in high-noise regimes ($\rho \ge 0.3$) compared to standard robust baselines, maintaining strict fairness without significant degradation in predictive utility.
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
Hao Liang ⋅ Qihan Lin ⋅ Mingrui Chen ⋅ Hengyi Feng ⋅ ZhaoYang Han ⋅ Xiaochen Ma ⋅ Zhen H Wong ⋅ Meiyi Qiang ⋅ Linzhuang Sun ⋅ Wentao Zhang
Large language models (LLMs) are increasingly deployed in K–12 education, yet existing benchmarks such as C-Eval, CMMLU, GaokaoBench, and EduEval measure only whether a model can answer an exam question, i.e., factual recall. Effective educational AI further requires curriculum cognition: the structured understanding of how knowledge is organized, including prerequisite chains, concept taxonomies, experiment–concept links, and pedagogical sequencing. Curriculum cognition is neither probed by current benchmarks nor explicitly taught by current instruction-tuning data. To close this gap, we introduce K12-KGraph, a curriculum-aligned knowledge graph extracted from the official People's Education Press textbooks, covering mathematics, physics, chemistry, and biology across primary, middle, and high school, with seven node types (Concept, Skill, Experiment, Exercise, Section, Chapter, Book) and nine relation types spanning taxonomy, prerequisite, association, verification, assessment, location, and order. Building on this single graph, we derive two complementary resources: K12-Bench, a 23,640-question multi-select benchmark across five graph-derived task families (Ground, Prereq, Neighbor, Evidence, and Locate) that jointly probe curriculum cognition; and K12-Train, a KG-guided supervised fine-tuning corpus of ~2,300 QA pairs synthesized from node attributes and edge semantics. Experiments expose a clear gap and a clear remedy: on K12-Bench, even the strongest proprietary model (Gemini-3-Flash) reaches only 57% exact match and the best open-source model (Gemma-4-31B-IT) only 46%, with Prereq and Neighbor being the hardest; yet under a strictly matched 2,300-sample SFT budget on Qwen3-4B-Base and Llama3.1-8B-Base, K12-Train consistently outperforms equally sized subsets of eight mainstream instruction-tuning corpora (OpenHermes, Infinity, UltraChat, WizardLM, DataFlow, LMSYS, SmolTalk, Tulu-3) on both GaokaoBench and EduEval, showing that structural curriculum grounding is remarkably sample-efficient for educational SFT. We release the graph, benchmark, training data, and the full construction pipeline.
Keep or Preempt? Termination-Aware Scheduling for LLM Serving with Speculative Decoding
Ziyu Cheng ⋅ Songtao Guo ⋅ Mingyan Li ⋅ Chang Han ⋅ Hongyu Xu ⋅ Xu Luo
Effective inference is critical for interactive large language model (LLM) serving, where both time-to-first-token (TTFT) and end-to-end (E2E) latency shape user experience. Speculative decoding has emerged as a promising solution that reduces inference latency by employing lightweight drafting and parallel verification. However, it introduces a new scheduling issue, i.e., traditional shortest remaining processing time (SRPT) scheduling based on output length estimation fails to work due to its paradigm shift in inference structure. In this paper, we explore the preemption mechanism of speculative decoding-based LLM serving systems and show that — besides the output length, a request's remaining service time depends on draft acceptance and the number of verification rounds. Building on this insight, we develop a novel termination-aware scheduling framework, named \textsc{AlmostDone}, which first \textit{identifies requests that are almost done} within several future local decoding steps. It achieves this by introducing a lightweight predictor built from runtime draft-and-verify states, which helps approximate the SRPT-like scheduling better than existing approaches. Then, \textsc{AlmostDone} \textit{determines whether preemption is worthwhile} through a switching-aware online policy to limit excessive running-set alterations. Experiments on an open-source vLLM serving system show \textsc{AlmostDone} substantially surpasses the default schedulers, reducing mean and median TTFT by up to 3.83$\times$ and 3.39$\times$, and mean and median E2E request latency by up to 1.39$\times$ and 2.33$\times$. Against the oracle length-based preemption baseline TRAIL+, it still achieves up to $2.0\times$ lower mean TTFT and $1.3\times$ lower mean E2E request latency.
KernelDNA: Cross-Layer Kernel Sharing via Decoupled Neural Adapters
Yadong Zhang ⋅ Haiduo Huang ⋅ Yinghui Xu ⋅ Tian Xia ⋅ Wenzhe zhao ⋅ Pengju Ren
Dynamic convolution improves the representational flexibility of CNNs by generating input-adaptive kernels, yet existing methods either inflate parameters linearly with the number of base kernels or sacrifice inference throughput due to runtime kernel assembly. We revisit dynamic convolution from the angle of cross-layer redundancy: a CKA analysis across modern CNN families shows that within-stage convolutional kernels are highly correlated, suggesting that an entire stage can be reparameterized by a single shared kernel together with lightweight, layer-specific transformations. Building on this observation, we propose KernelDNA, a parameter-efficient framework in which multiple convolutional layers in a stage share one ''parent'' kernel, while layer-specific ''child'' kernels are produced by a \textit{decoupled} adapter that splits modulation into (i) an input-dependent dynamic channel gate and (ii) static spatial and filter modulations that can be pre-fused into the parent kernel before deployment. We provide a theoretical justification grounded in tensor approximation: when two kernels have CKA similarity $\ge 1 - \delta$, the multiplicative adapter incurs a Frobenius approximation error $O(\sqrt{\delta})$, and the implicit gradient consensus across child layers provably reduces the Rademacher complexity of the hypothesis class. Across ImageNet-1K and MS-COCO, KernelDNA achieves state-of-the-art accuracy-efficiency trade-offs over diverse backbones (ResNet18/50, MobileNetV2, ConvNeXt-Tiny), reducing parameters by $1.2$-$5\times$ versus dynamic convolution baselines while retaining $90$-$99$ % of the standard-convolution throughput---outperforming KernelWarehouse and FDConv across all backbones, and ODConv on $4$ of $5$ settings at a fraction of its parameter count.
Kernel-Gradient Drifting Models
Maria Esteban-Casadevall ⋅ Jorge Carrasco-Pollo ⋅ Max Welling ⋅ Jan-Willem van de Meent ⋅ Erik Bekkers ⋅ Floor Eijkelboom
We propose kernel-gradient drifting, a one-step generative modeling framework that replaces the fixed Euclidean displacement direction in drifting models with directions induced by the kernel itself. Standard drifting is attractive because it enables fast, high-quality generation without distilling a large pretrained diffusion model, but its theory is currently understood mainly for Gaussian kernels, where the drift coincides with smoothed score matching and is identifiable. Our gradient-based reformulation exposes this score-based structure for general kernels: the resulting drift is the score difference between kernel-smoothed data and model distributions, yielding identifiability for characteristic kernels and a smoothed-KL descent interpretation of the drifting dynamics. Since kernel gradients are intrinsic tangent vectors, the same construction extends naturally to Riemannian manifolds and to discrete data via the Fisher-Rao geometry of the probability simplex. Across spherical geospatial data, promoter DNA and molecule generation, kernel-gradient drifting enables state-of-the-art one-step generation beyond the Euclidean setting without distillation.
Activation steering provides a simple, training-free mechanism for controlling attributes of generative models (e.g., sentiment, style, helpfulness). However, standard approaches such as Difference-in-Means (DiM) apply a single input-independent steering vector across all activations, limiting expressivity and ignoring the local structure of the activation space. We propose Kernelized Activation Steering (KAS), a unifying framework that lifts activation steering into a reproducing kernel Hilbert space (RKHS). KAS formulates steering as an optimization problem expressed purely via kernel evaluations, yielding an implicit, activation-dependent steering score without constructing explicit feature maps. Unlike DiM, KAS induces locally adaptive steering: each activation is modified according to its relative position with respect to source and target reference sets, producing a nonlinear steering field over the representation space. Importantly, DiM is recovered as a special case under a linear kernel, offering a principled interpretation of steering, while richer kernels (e.g., RBF) enable geometry-aware interventions. Across standard activation steering tasks, including jailbreaking LLMs and image style control, KAS consistently outperforms existing methods.
Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech
Dong Yang ⋅ YIYI CAI ⋅ Haoyu Zhang ⋅ Yuki Saito ⋅ Hiroshi Saruwatari
Metric-induced discrete flow matching (MI-DFM) exploits token-latent geometry for discrete generation, but its practical use is limited by two issues: heuristic schedulers requiring hyperparameter search, and finite-step path-tracking error from its first-order continuous-time Markov chain (CTMC) solver. We address both issues. First, we derive a kinetic-optimal scheduler for prescribed scalar-parameterized probability paths, and instantiate it for MI-DFM as a training-free numerical schedule that traverses the path at constant Fisher–Rao speed. Second, we introduce a finite-step moment correction that adjusts the jump probability while preserving the CTMC jump destination distribution. We validate the resulting method, GibbsTTS, on codec-based zero-shot text-to-speech (TTS). Under controlled comparisons with a unified architecture and large-scale dataset, GibbsTTS achieves the best objective naturalness and is preferred in subjective evaluations over masked discrete generative baselines. Additionally, in comparison with the evaluated state-of-the-art TTS systems, GibbsTTS shows strong speaker similarity, achieving the highest similarity on three of four test sets and ranking second on the fourth. Demo and code are available in the supplementary materials.
Knocking-Heads Attention: Drop-in Shared Projections for Cross-Head Coordination
Zhanchao Zhou ⋅ Xiaodong Chen ⋅ Haoxing Chen ⋅ Zhenzhong Lan ⋅ Jianguo Li
In standard multi-head attention, each head computes its output independently, with cross-head information exchange limited to the final output projection. Talking-heads attention introduces explicit inter-head mixing on attention maps, but is incompatible with FlashAttention and incurs significant overhead. We propose knocking-heads attention (KHA), which couples heads at the parameter level via a shared, diagonally-initialized projection applied to query/key/value features before the scaled dot-product attention. The shared projection does not mix signals along the head axis; instead, it is jointly optimized by gradients from all heads, providing an implicit regularization on the head-specific projections. The diagonal initialization recovers MHA at step 0 and lets cross-head structure emerge gradually. We instantiate KHA in two complementary forms: a linear variant (KHA-Linear) whose shared projection can be absorbed into the original weights at inference, adding zero parameters or FLOPs at deployment; and a non-linear variant (KHA-MLP) that obtains stronger gains by introducing a small shared MLP, at modest inference cost. Training a 6.1B parameter MoE model (1.01B activated) on 1T tokens, we observe substantially fewer loss spikes, more uniform per-head activation norms, and a +1.26 average improvement across 19 downstream benchmarks. KHA is compatible with FlashAttention, KV-caching, and existing attention variants (MHA, MQA, GQA, GTA, MLA).
Knowledge Transfer Scaling Laws for 3D Medical Imaging
Ho Hin Lee ⋅ Dongna Du ⋅ Chu Wang ⋅ Yuankai Huo ⋅ Shi Gu ⋅ James Gee ⋅ Yifan Wu
Vision foundation models are increasingly moving beyond 2D to volumetric domains such as 3D medical imaging, where unified pretraining across different imaging modalities (i.e. CT, MRI, and PET) could provide foundational models for diverse clinical tasks. However, training such models requires mixing heterogeneous imaging domains, and current mixture strategies remain largely heuristic. In this work, we observe that different medical imaging domains scale at variable rates during pretraining, and knowledge transfer between domains is strongly asymmetric: training on one domain can substantially improve another, but the reverse may be much weaker. Interestingly, both MAE reconstruction loss and cross-domain transfer follow predictable power-law trends with domain-specific behaviors. Motivated by these findings, we formulate data allocation as a scaling-law optimization problem. The derived allocations reveal an interpretable hub-and-island structure: highly transferable domains emerge as hubs that benefit many others and deserve strategic allocation, while isolated domains act as islands requiring direct investment. Empirically, transfer-aware allocation outperforms data-proportional sampling by up to 58\% and generalizes well to unseen budgets with r=0.989. Downstream validation on disease classification and organ/lesion segmentation further confirms that the derived transfer-aware mixtures provide stronger pretrained representations for clinical 3D medical imaging tasks.
Known By Their Actions: Fingerprinting LLM Browser Agents via UI Traces
William Gitta Lugoloobi ⋅ Samuele Marro ⋅ Jabez Magomere ⋅ Joss Wright ⋅ Chris Russell
As LLM-based agents increasingly browse the web on users' behalf, a natural question arises: can websites passively identify which underlying model powers an agent? Doing so would represent a significant security risk, enabling targeted attacks tailored to known model vulnerabilities. Across 14 frontier LLMs and four web environments spanning information retrieval and shopping tasks, we show that an agent's actions and interaction timings, captured via a passive JavaScript tracker, are sufficient to identify the underlying model with up to 96\% F1. We formalise this attack surface by demonstrating that classifiers trained on agent actions generalise across model sizes and families. We further show that strong classifiers can be trained from few interaction traces and that agent identity can be inferred early within an episode. Injecting randomised timing delays between actions substantially degrades classifier performance, but does not provide robust protection: a classifier retrained on delayed traces largely recovers performance. We release our harness and a labelled corpus of agent traces \href{https://anonymous.4open.science/r/known_actions-B31B}{here}.
KV Cache Compression via Attention Output Distortion Minimization
Yi Wang ⋅ Yixuan Cao ⋅ Qiang Ding ⋅ Zelin Gao ⋅ Ping Luo
Efficient long-context inference requires compressing the KV cache under strict memory budgets without degrading model outputs. We formulate this problem as a constrained optimization: minimize the total attention-output distortion subject to a global average bit-width budget, where each token’s KV entry can be assigned a bit-width ranging from 0 (eviction), through mixed precision, to 16 (no compression). The central challenge is to estimate each token’s contribution to attention-output distortion. We address this via a first-order perturbation analysis of the attention mechanism, which yields a closed-form decomposition of the distortion. Specifically, the compression cost of each token factorizes into (i) a token-specific risk score that captures the sensitivity of the attention output to perturbations of that entry, and (ii) a precision-dependent quantization error determined by the chosen bit-width. Building on this factorization, we solve the global allocation problem via Lagrangian relaxation, reducing it to independent per-token decisions that can be efficiently computed. Experiments on LongBench and RULER with LLaMA-3.1-8B-Instruct and Qwen3-8B show that our method achieves comparable accuracy to uniform 4-bit quantization while using an average of 2 bits per entry (2× memory reduction). At an average of 1 bit, our method maintains substantially higher accuracy than both uniform quantization and token eviction baselines, which deteriorate sharply in this regime. The advantage concentrates on retrieval-sensitive subtasks, where dispersed evidence makes uniform allocation and hard eviction especially costly.
LAMP: Language-Modulated Geometric Preservation for Multi-Modal Object Re-Identification
Shuying Li ⋅ Chao Su ⋅ Yongxiang Li ⋅ Peng Hu ⋅ Dezhong Peng ⋅ Yuan Sun
Multi-modal object Re-Identification (ReID) aims to achieve robust all-weather perception by leveraging complementary information across heterogeneous modalities. However, existing methods still suffer from limitations in both feature fusion and feature alignment. In terms of feature fusion, most approaches assume equal importance across modalities and do not explicitly account for the quality differences among modalities in dynamic environments. As a result, the fused representations tend to be suboptimal and less discriminative in complex scenarios. In terms of feature alignment, existing methods typically rely on point-wise feature matching while ignoring relational structures among samples, which suppresses discriminative information and introduces cross-modal inconsistencies. To address these limitations, we propose LAMP, a unified framework that rethinks multi-modal interaction through language-modulated control and geometry-consistent alignment. We first adopt a Semantic-Assisted Feature Encoder (SAFE), which incorporates textual semantics into visual features to enhance local discriminative cues and provide more informative representations for subsequent processing. Building upon this, the Language-Modulated Dynamic Routing (LMDR) module generates modality-aware gating signals prior to feature fusion, enabling the adaptive adjustment of modality weights under challenging conditions. Furthermore, to address cross-modal inconsistencies, we introduce a Fused Geometric Prototype Alignment (FGPA) objective based on Fused Gromov-Wasserstein (FGW) optimal transport. Instead of enforcing strict point-wise alignment, FGPA establishes cross-modal consensus at the identity-level relational geometry, thereby preserving discriminative structures while mitigating inconsistencies across modalities. LMDR improves the stability of multi-modal fusion, while FGPA enhances structural consistency in cross-modal alignment. Together, they improve the discriminative capability of the model. Extensive experiments on three multi-modal object ReID benchmarks demonstrate the effectiveness of our method.
The ability of machine learning models to store input information in hidden layer embeddings, a form of model `memory', is widely employed but not well characterized. We find that causal language model embeddings typically contain relatively little input information regardless of data and compute scales during training. In contrast, embeddings from autoencoders trained for input regeneration are capable of nearly perfect memory formation. The substitution of memory embeddings for token sequences leads to computational efficiencies, motivating the introduction of a parallelizable encoder-decoder memory model architecture. Upon causal training these models contain information-poor embeddings incapable of arbitrary information access, but by combining causal and information retention objective functions they learn to form and decode information-rich memories. Training can be further streamlined by freezing a high fidelity encoder followed by a curriculum training approach where decoders first learn to process memories and then learn to additionally predict next tokens. We conclude that next token prediction training alone is poorly suited for accurate memory formation, motivating the use of combined objective functions for models where the entire input is not exposed.
LAQuant: A Simple Overhead-free Large Reasoning Model Quantization by Layer-wise Lookahead Loss
Euntae Choi ⋅ Sumin Song ⋅ Sungjoo Yoo
Large reasoning models (LRMs) reach competition-level math and coding accuracy via long autoregressive decoding, making per-token decoding cost a primary deployment concern. Weight quantization is the standard tool for acceleration, but representative recipes --- including state-of-the-art end-to-end (E2E) QAT --- lose accuracy on long-decoding reasoning benchmarks despite preserving perplexity and short-decode accuracy. Through a systematic gradient-direction analysis, we identify two factors driving this gap: (i) KV-cache fidelity preservation under the QAT loss, which E2E supervision attenuates via the softmax Fisher metric; and (ii) Hessian-subspace alignment between calibration data and the deployment distribution. We propose LookAhead Quantization (LAQuant), a layer-wise weight-only QAT method that addresses both factors without online-transform overhead by combining reasoning-domain calibration with a one-layer lookahead loss whose implicit cross-layer co-adaptation preserves the next-layer residual stream. For Qwen3-4B under W3G128 quantization, LAQuant improves AIME25 Pass@1 over ParoQuant by 15.11pp (1.93pp over ParoQuant++ at matched calibration) while achieving a 3.42$\times$ decoding speedup over FP16 on RTX A6000, compared with ParoQuant's 3.01$\times$.
Large Discrete Policy: Advancing Explicit Behavior Modeling with Stochastic Iterative Scoring
Zhenxin Li ⋅ Nadine Chang ⋅ Xinglong Sun ⋅ Jingde Chen ⋅ Wenhao Yao ⋅ Zi Wang ⋅ Maying Shen ⋅ Yu-Gang Jiang ⋅ Zuxuan Wu ⋅ Shiyi Lan ⋅ Jose M. Alvarez
Behavior policies are often formulated as continuous generative models, whose iterative denoising processes are expressive but difficult to interpret and prone to producing implausible actions. We propose the Large Discrete Policy (LDiP), a fully discrete behavior modeling framework that selects actions from a large vocabulary of physically plausible candidates. Rather than perturbing actions, LDiP improves expressivity through stochastic iterative scoring: it progressively re-scores and prunes candidates with score-space stochasticity, enabling fine-grained ranking and exploration among plausible actions while preserving an explicit decision process. Across end-to-end planning, closed-loop driving, robotic manipulation, and vision-language-action settings, LDiP consistently outperforms strong discrete and continuous baselines in autonomous driving, and exceeds or matches continuous generative policies in robotic manipulation. These results show that discrete policies, when equipped with effective scoring mechanisms, offer an expressive, plausible, and interpretable alternative for behavior modeling.
Large Language Models Develop Belief State Geometry In-Context
Daniel Balcells ⋅ Andrew Jun Lee ⋅ Chirag Rastogi ⋅ Paul Riechers ⋅ Adam Shai ⋅ Xavier Poncini
Large language models (LLMs) trained on next-token prediction exhibit remarkable in-context learning (ICL) abilities, yet the representations that support ICL remain poorly understood. We consider such representations in a controlled setting: prompting LLMs with data emitted from hidden Markov models (HMMs) and probing for the corresponding belief state -- the posterior distribution over the HMM's hidden states given the observed token history. Across six open-source LLMs prompted with data from 40 HMMs selected for non-trivial belief structure, we find that belief states are linearly decodable from residual stream activations, with peak probe $R^2$-values ranging from 0.83 to 0.99 across HMM and LLM combinations, typically occurring at early-to-middle layers. To establish functional relevance, we intervene directly on the probe-identified subspace via patching and steering, resulting in downstream prediction quality on the order of the untampered model, while controls degrade substantially. Together, these results provide representation-level evidence that ICL in pre-trained LLMs approximates optimal Bayesian prediction over a context-inferred HMM. More broadly, our findings extend prior results linking input-distribution structure to activation geometry: from toy networks trained explicitly on HMM data to production-scale transformers trained on natural language.
Large Language Model Selection with Limited Annotations
Yavuz Durmazkeser ⋅ Patrik Okanovic ⋅ Andreas Kirsch ⋅ Torsten Hoefler ⋅ Nezihe Merve Gürel
Choosing a Large Language Model (LLM) for a given task requires comparing many strong candidates, yet standard evaluation relies on costly annotations over fixed evaluation sets. To address this challenge, we develop Select-LLM, the first framework for active model selection of LLMs. Select-LLM aims to find a small set of queries whose annotations are most informative for identifying the best LLM for a given task. To this end, we introduce a query selection rule based on expected information gain, computed from pairwise similarities between candidate model outputs. Because this rule only uses generated model responses, Select-LLM can be applied across candidate models without assumptions about their architecture or access to model weights. This makes it suitable for both open-weight and black-box LLMs. We evaluate Select-LLM across 23 datasets, 156 evaluated models, diverse task families, and multiple text evaluation metrics. Across all experiments, Select-LLM improves over the strongest baseline in every setting, with annotation cost reductions up to 81.8% for best model selection and up to 84.78% for near-best model selection.
LARGO: Low-Rank Hypernetwork for Handling Missing Modalities
Niels Vyncke ⋅ Pooya Ashtari ⋅ Aleksandra Pizurica
Addressing missing modalities is an important challenge in multimodal image analysis and often relies on complex architectures that do not transfer easily to different datasets without architectural modifications or hyperparameter tuning. While most existing methods tackle this problem in feature space by engineering representations that are robust to missing inputs, we instead operate in weight space. We propose LARGO, a hypernetwork that compresses the $2^N-1$ dedicated missing-modality models into a single network by modelling the convolutional weights using the Canonical Polyadic (CP) tensor decomposition. Extensive experimental validation on BraTS 2018 (4 modalities, 15 scenarios) and ISLES 2022 (3 modalities, 7 scenarios) shows that our method ranks first in 47 out of 52 configurations, achieving average Dice improvements of +0.68\% and +2.53\% over state-of-the-art baselines (mmFormer, M³AE, ShaSpec, SimMLM). A proof-of-concept experiment on avMNIST suggests that LARGO may extend beyond medical imaging to heterogeneous non-medical modalities.
Retrieval-Augmented Generation (RAG) has become a standard approach for enhancing large language models (LLMs) with external knowledge, mitigating hallucinations, and improving factuality. However, existing systems rely on generating natural language queries at each hop and maintaining a strict architectural separation between retriever and generator, preventing them from leveraging the full representational capacity of the LLM. We propose LAnR (Latent Abstraction for RAG), a unified framework in which a single LLM jointly performs encoding, retrieval, and generation entirely within its own latent space. Rather than generating textual queries, LAnR produces dense retrieval vectors from the hidden states of a designated [PRED] token and uses them to match against encoded document representations from the same model. Furthermore, LAnR adaptively decides when sufficient evidence has been retrieved using a lightweight MLP control head over those same hidden states, eliminating explicit token-level stopping reasoning. Extensive experiments on five QA benchmarks spanning single-hop and multi-hop settings demonstrate that LAnR outperforms existing RAG methods, while achieving improved inference efficiency through reduced number of retrieval calls and tighter model integration.
Latent-Lens: Visual Perception in Small Language Models Through Latent Communication
Mo Shahdloo ⋅ Faegheh Sardari ⋅ Mathew Salvaris
Bridging independent vision-language models (VLMs) and small language models (SLMs) in a modular multi-agent system typically forces them to exchange information as discrete text, which creates a lossy bottleneck for spatial and fine-grained details. This constraint is especially acute for SLMs, whose limited capacity makes them more reliant on the quality of the input signal. While recent works on latent multi-agent communication have explored latent-space exchange between identical or homogeneous LLMs, propagating rich visual information across model boundaries without text generation has not been studied yet. We propose Latent-Lens, a lightweight codec that bridges the representation gap between a frozen VLM and a frozen SLM directly in latent space, enabling them to communicate purely through hidden states. The codec consists of a domain encoder that projects hidden states from the VLM into the SLM's embedding space, and an attention decoder formalised as a LoRA that teaches the SLM to attend to these visual prefixes, while training only 1.6M parameters (0.2\% of the combined model size). On free-form visual question answering (CLEVR, GQA), Latent-Lens achieves +10.2-28.3pp over a text-communication baseline given the same LoRA budget (+36.4-78.4pp when both agents are frozen); on image captioning (Flickr8k) it improves over the same baseline by +18.5 (+18.9 frozen). Replacing the SLM (TinyLlama-1.1B) or the VLM (InternVL2-1B) with alternative models of different architectures, vocabularies, and hidden dimensions yields consistent gains, confirming generality across heterogeneous model families.
Latent Refinement Decoding: Enhancing Diffusion Language Models by Refining Belief States
Qinglin Zhu ⋅ Yizhen Yao ⋅ Runcong Zhao ⋅ Yanzheng Xiang ⋅ Siya Qi ⋅ Amrutha Saseendran ⋅ Chen Jin ⋅ Philip A Teare ⋅ Bin Liang ⋅ Yulan He ⋅ Lin Gui
Autoregressive (AR) models remain the dominant paradigm for natural language generation, but their strictly sequential decoding process leads to high inference latency. Recent diffusion-inspired language models, such as LLaDA and Dream, alleviate this issue through parallel generation, yet they still face two key limitations: information loss, since predictive distributions over non-finalised tokens are discarded at each step, and unstable commitment dynamics, where local token decisions are not sufficiently coordinated at the global level. We propose Latent Refinement Decoding (LRD), a two-stage decoding framework consisting of Latent Refinement and a Predictive Feedback Loop. In the first stage, LRD keeps masked positions as distributional mixtures of predicted tokens and the mask embedding, enabling the model to form more globally consistent beliefs before committing tokens. In the second stage, LRD progressively finalises confident tokens while preserving uncertain ones for further iterative refinement, with KL-divergence dynamics serving as a stable criterion for convergence and early stopping. Experiments show that LRD improves performance on both coding tasks, including HumanEval by 6.3 and MBPP by 2.6, and reasoning tasks, including GSM8K by 2.9 and MATH500 by 3.8, while achieving up to 10.6× overall speedup. Moreover, LRD is compatible with single-token decoding, multi-token threshold commitment, and system-level accelerators; when combined with Fast-dLLM, it reaches up to 18.4× speedup over vanilla decoding while further improving accuracy.
LDD-RFM: Learnable Domain Decomposition for Random Feature Models via Variable Projection
Zhaohui Fu ⋅ Duanyu Feng ⋅ Yangshuai Wang
Random Feature Models (RFMs) are attractive PDE surrogates because their linear-in-features form reduces training to regularized least squares. Standard RFMs, however, use one stationary spectrum over the whole domain, creating a bandwidth dilemma on heterogeneous PDE solutions: fine scales resolve sharp local gradients but oscillate in smooth regions, while coarse scales miss interface-dominated derivatives. We propose \textbf{LDD-RFM}, a learnable domain-decomposition framework whose local experts remain linear RFMs and whose nonlinear variables specify compact-support partition-of-unity (PoU) gates and local spectral scales. For fixed gates and spectra, expert coefficients are eliminated by ridge regression through variable projection, yielding a differentiable reduced objective with explicit linear-solve and partition-smoothness controls. We prove an $H^1$ error decomposition in which PoU sharpness enters through a product-rule term. Controlled regression and PDE experiments show that LDD-RFM reduces $H^1$ errors by $10\times$ over global and fixed-partition RFM baselines under matched feature budgets; exact-solver, timing, and robustness studies check that these gains are not artifacts of the linear solver or evaluation protocol.
Leak-CURBER: A Leakage-Controlled Multimodal Evaluation Benchmark for Enzymatic Reaction Tasks
Adhil Ahmed P M Shums ⋅ Roshan Balaji ⋅ Nirav Bhatt
We present Leak-CURBER, the largest multimodal, leakage-controlled benchmark for evaluating machine learning and deep learning models for enzymatic reactions. Leak-CURBER standardizes data from 10 biochemical databases and curates datasets with multimodal inputs, metrics, and leakage-controlled splits for 4 tasks. The benchmark comprises 10 subtasks: prediction of reaction outcomes (2), enzyme function (2), kinetic parameters (3), and protein–ligand binding (3). Leak-CURBER establishes new evaluation protocols that control for train/validation/test overlap across protein sequence and structure, molecule string and structure, and reaction similarity, thereby exposing generalization failures of learning methods. Evaluating the state-of-the-art (SOTA) models on Leak-CURBER shows that their strong performance degrades sharply under these splits, often approaching random classification or retrieval performance, and near-zero regression performance. Leak-CURBER includes precomputed embeddings, continual benchmark releases, reproducible curation scripts, and a public leaderboard to support reliable comparison of enzymatic-reaction models.
Learnable Diffusion-based Positional Encodings for Link Prediction
Jiaji Ma ⋅ Murad Huseynli ⋅ Danai Koutra
Link prediction relies on pairwise structural proximity, but standard node representations are not explicitly designed so that their interactions reflect such signals. Existing structure-aware approaches address this gap through specialized architectures or handcrafted proximity features, often at the expense of generality or additional computational cost. In this work, we take a data-centric approach and propose $\text{LinkDPE}$, a learnable $\underline{\text{d}}$iffusion-based $\underline{\text{p}}$ositional $\underline{\text{e}}$ncoding framework that injects structural proximity into the node features. Our key idea is that diffusion kernels, which capture multi-scale connectivity patterns, admit a factorization into node-wise embeddings whose inner products recover pairwise proximity scores. Based only on the observed training graph, $\text{LinkDPE}$ constructs diffusion-based encodings across scales, adaptively selects complementary diffusion scales, and learns task-specific spectral weightings, enabling a wide range of link predictors to exploit structural information without architectural modification. Theoretically, we show that (i) diffusion proximity admits a node-wise kernel factorization, (ii) many classical link prediction heuristics are recovered as special cases of low-order walk and diffusion operators, and (iii) multi-scale spectral responses allow the model to capture a richer set of structural similarity patterns within a shared truncated spectral basis. Across 8 benchmark datasets and 5 representative base models, $\text{LinkDPE}$ consistently improves link prediction performance over original features and positional-encoding baselines, yielding average gains of 6.97\% over the unaugmented base models.
Implicit neural representations (INRs) are shaped by the spectral structure induced by their input encodings and activation functions. Existing methods improve fitting primarily by modifying which frequencies are available to the network, through coordinate encodings or periodic nonlinearities. However, frequency access is not the only bottleneck: signals with localized or spatially varying structure require the network to efficiently compose frequencies into multi-harmonic internal responses. We introduce learnable spectral activations (LSA), which replace fixed neuron-level nonlinearities with a residual truncated Fourier series whose harmonic amplitudes are learned during training. LSA does not expand the asymptotic function class. Instead, it changes the factorization of the representation: linear weights select features while activation coefficients control spectral shaping, and the two are updated by separate gradients. Because the activation coefficients enter the loss linearly given fixed pre-activations, spectral tuning becomes a simpler subproblem than in architectures where it is entangled with feature selection. Empirically, this factorization concentrates more target-signal energy in the leading eigenmodes of the neural tangent kernel, consistent with improved optimization behavior. Across audio, image, neural radiance field, and neural acoustic field tasks, LSA also improves reconstruction quality.
Learn from your own latents and not from tokens: A sample-complexity theory
Daniel Korchinski ⋅ Alessandro Favero ⋅ Matthieu Wyart
Generative models, from diffusion models to large language models, achieve remarkable performance but at a cost in training data orders of magnitude larger than what biological learners require. An alternative paradigm has emerged in which networks are trained to predict their \emph{own} latent representations of related views or masked regions, as in data2vec and JEPA -- an idea related to predictive-coding accounts of the cortex. Despite strong empirical results, the theoretical understanding of these methods remains limited. Central questions include: by how much does latent prediction actually improve data efficiency? Is there a benefit to stacking such methods into multi-scale hierarchies? We answer both using as data a tractable probabilistic context-free grammar that captures the compositional structure of natural language and images. We prove that latent prediction recovers the full latent tree of depth $L$ from a number of samples scaling as $m^3$, where $m$ is the number of production rules per symbol. This is much fewer than the $m^L$ samples required by supervised learning and the $m^{L+1}$ required by token-level SSL. We confirm this bound with (i) a hierarchical clustering algorithm, (ii) an end-to-end neural network whose predictor-clusterer modules predict their own latents at each level via gradient descent, and (iii) the first sample-complexity analysis of data2vec, which we show implicitly performs hierarchical latent prediction. This suggests that explicit stacking such as H-JEPA is largely redundant.
Learning Agentic World Vision-Language-Action Models for Autonomous Driving
Guoqing Wang ⋅ Pin Tang ⋅ Xiangxuan Ren ⋅ Chao Ma
Recent progress in integrating vision-language-action (VLA) models with world modeling has significantly advanced end-to-end autonomous driving by allowing systems to reason about imagined futures instead of merely reacting to current observations. However, most world VLA models generate future frames at the pixel level with plenty of dense visual tokens. Although such generated futures can be visually plausible, they preserve many driving-irrelevant appearance details, i.e., only a small portion of the tokens directly encode planning-relevant factors, such as agent geometry, motion dynamics, and interactions. This introduces a significant performance bottleneck. An important question thus arises: what world state should a world VLA model reason over to support planning beyond pixel-level prediction? We argue that planning-oriented world modeling should focus on an agentic latent state that captures the interactions between objects in the driving scene over time. As such, we introduce AgentWorld, an agentic world VLA model that uses discrete agent tokens jointly for reasoning and trajectory planning. Each token is dynamically grounded to a surrounding agent associated with its object's state, such as semantic identity, 3D geometry, motion, and interaction context. AgentWorld learns these tokens through three key processes: explicit geometric grounding, reasoning-oriented supervised fine-tuning, and agent-oriented reinforcement learning. Extensive open-loop and closed-loop experiments under diverse driving scenarios show that AgentWorld enhances planning capabilities, fully exploits visual evidence and agent dynamics, and yields interpretable agentic world states. The code will be made publicly available.
The field of learning-augmented algorithms has demonstrated that machine-learned predictions can bypass worst-case lower bounds across a wide range of problems. So far, however, the focus has been almost exclusively on polynomial-time algorithms, where predictions improve competitive ratios, approximation guarantees, or running times. In this paper, we raise the question of whether predictions can push the frontier of exact exponential-time algorithms for NP-hard problems. We answer this question affirmatively by proposing a general approach that augments an entire family of state-of-the-art exact algorithms for a variety of subset selection problems. We show that a noisy predictor that is only marginally better than random guessing suffices to provably reduce the search space, and that the resulting runtime speedup scales smoothly with the prediction quality. Importantly, our algorithms require only pairwise independence of predictions or, alternatively, do not require the knowledge of the predictor's accuracy–both strictly weaker and more realistic settings than typically assumed.
Learning Causal Orderings for In-Context Tabular Prediction
Sascha Xu ⋅ Sarah Mameche ⋅ Jilles Vreeken
In-context learning for tabular data sets strong predictive standards in observational settings; it however primarily relies on correlational structure, which becomes unreliable under distribution shift or intervention. While established methods to discover causal structure exist, they are often focused on structure identifiability and decoupled from the predictive architectures that could benefit from them. To bridge these perspectives, we study how to simultaneously infer and enforce causal structure in the form of topological variable orderings into tabular prediction. Unlike standard architectures, our model TabOrder uses causal order-constrained attention, basing predictions only on features that precede a target under a learned causal order. Similar to causal discovery methods, TabOrder learns the optimal variable ordering in an unsupervised manner through a likelihood-based objective. We justify this choice under standard functional model classes and also study how sample missingness, a common challenge in tabular data, interacts with causal direction identification. Empirically, we confirm that TabOrder recovers accurate variable orderings while addressing prediction and imputation tasks, as well as gives insight into real-world biological data under intervention.
Learning Discrete Riemannian Metrics for Physical Fields with Cochain-Frame Equivariance
Dongzhe Zheng ⋅ Christine Allen-Blanchette
Physical fields on meshes require a separation between topology and geometry: conservation laws are topological and should be exact, while geometry, material response, and anisotropic coupling must be learned from data. Existing neural surrogates often mix these roles inside unconstrained message passing. We introduce Riemannian Hodge Message Passing (RHMP), which turns this separation into an architectural principle. RHMP fixes the cellular coboundaries ($d_k$) determined by oriented incidence and learns symmetric positive-definite cochain metrics ($H_k$) for geometry-dependent propagation. Treating $H_k$ as the learned metric motivates cochain-frame equivariance: physical propagation should be invariant to orthogonal changes of the hidden cochain feature basis. RHMP implements this principle with metric-weighted Hodge blocks ($d_k^\top H_{k+1}d_k$), yielding exact cochain-complex identities ($d_{k+1}d_k=0$), nonnegative Hodge energies, positive-semidefinite operators, and exact Abelian curvature invariance. Across seven physical benchmarks spanning fluids, electromagnetism, gauge fields, and variable-mesh CFD, RHMP achieves the best overall performance, with the largest gains when topology, learned geometry, and field structure interact.
Learning Fine-Grained Vision-Language Alignment from Discriminative Part Descriptions
Sibo Yin ⋅ Yuxin Peng
Vision-Language Pre-trained (VLP) models such as CLIP learn strong representations from large-scale image–text pairs and demonstrate impressive zero-shot transfer. However, by learning to align visual and textual features only at a global level, they often suffer from limited interpretability and weak fine-grained perception. This issue stems from pre-training data that overlooks key visual details of object parts, which are often important for distinguishing between subordinate categories (e.g., species of birds or models of cars). Multimodal Large Language Models (MLLMs) built on CLIP-style vision encoders inherit this weakness, limiting both accuracy and trustworthiness. To address this, we propose Part-Aware CLIP (PA-CLIP), a framework that improves fine-grained perception while enhancing interpretability. First, we leverage MLLMs to construct a new dataset,FG-Part, containing about one million part-level image–text pairs that explicitly describe discriminative components (e.g., beak shape or wing patterns). Second, we introduce a part-aware training strategy that encourages explicit grounding of fine-grained textual descriptions to corresponding image regions, strengthening part-level cross-modal alignment. Extensive experiments show that PA-CLIP achieves state-of-the-art results on multiple fine-grained visual recognition benchmarks, validating the benefit of part-level captions for capturing subtle details. Moreover, evaluations on general tasks such as cross-modal retrieval indicate that these improvements do not compromise the model's core generalist capabilities.
Learning Generalizable Hand-Object Tracking Control without Human Demonstrations
Runyi Yu ⋅ Xiaoyi Lin ⋅ Hok W Tsui ⋅ Yinhuai Wang ⋅ PENG Zhijie ⋅ Hui Zhang ⋅ Qihan Zhao ⋅ Ke Fan ⋅ Miao Li ⋅ Jie Song ⋅ Jingbo Wang ⋅ Qifeng Chen ⋅ Ping Tan
We present a system for learning generalizable hand-object tracking controllers purely from synthetic data, without requiring any human demonstrations. Our approach makes two key contributions: (1) HOP, a Hand-Object Planner, which can synthesize diverse hand-object trajectories; and (2) HOT, a Hand-Object Tracker that bridges synthetic-to-physical transfer through reinforcement learning and interaction imitation learning, delivering a generalizable controller conditioned on target hand-object states. Our method extends to diverse object shapes and hand morphologies. Through extensive evaluations, we show that our approach enables dexterous hands to track challenging, long-horizon sequences including object re-arrangement and agile in-hand reorientation. These results represent a significant step toward scalable foundation controllers for manipulation that can learn entirely from synthetic data, breaking the data bottleneck that has long constrained progress in dexterous manipulation.
Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation
Yifan Xie ⋅ YuAn Wang ⋅ Guangyu Chen ⋅ Jinkun Liu ⋅ Yu Sun ⋅ Wenbo Ding
Human videos contain rich manipulation priors, but using them for robot learning remains difficult because raw observations entangle scene understanding, human motion, and embodiment-specific action. We introduce MoT-HRA, a hierarchical vision-language-action framework that learns human-intention priors from large-scale human demonstrations. We first curate HA-2.2M, a 2.2M-episode action-language dataset reconstructed from heterogeneous human videos through hand-centric filtering, spatial reconstruction, temporal segmentation, and language alignment. On top of this dataset, MoT-HRA factorizes manipulation into three coupled experts: a vision-language expert predicts an embodiment-agnostic 3D trajectory, an intention expert models MANO-style hand motion as a latent human-motion prior, and a fine expert maps the intention-aware representation to robot action chunks. A shared-attention trunk and read-only key-value transfer allow downstream control to use human priors while limiting interference with upstream representations. Experiments on hand motion generation, simulated manipulation, and real-world robot tasks show that MoT-HRA improves motion plausibility and robust control under distribution shift.
Learning Informative Invariant Representations via Hierarchical Latent Decomposition
Hyungkwon Lee ⋅ Hyewon Ryu ⋅ Jong-Seok Lee
Domain generalization aims to learn models that generalize to unseen target domains using labeled data from multiple source domains. Many existing approaches focus on learning domain-invariant representations, but enforcing invariance alone may fail to preserve sufficient task-relevant information under domain shift. We propose BLENDER, a reconstruction-aware framework for domain generalization that employs a hierarchical latent decomposition to separate domain-invariant and domain-specific factors. This structure prioritizes informativeness in the domain-invariant representation and captures residual domain-specific variation through conditional modeling. We further provide a theoretical analysis establishing a bound on the reconstruction risk for unseen target domains, revealing how domain invariance and latent disentanglement contribute to generalization beyond the observed domains. Experiments on standard domain generalization datasets demonstrate strong performance under substantial domain shifts, and qualitative analyses show that BLENDER promotes a structured allocation of information between domain-invariant and domain-specific latent representations.
Learning Rate Transfer for Hybrid Transformer-SSM Architectures
Jimin Seo ⋅ Gyubok Lee ⋅ Yeonsik Jo ⋅ Kiwoong Yoo ⋅ Yeongoon Kim ⋅ Minhae Oh ⋅ Jin W Koo ⋅ Suhwan Kim ⋅ Nakyung Lee ⋅ Minsik Seol ⋅ Idris Nechnech ⋅ Jaehyeon Kim ⋅ LeeGiho ⋅ Jungwoo Lee
We study learning rate (LR) scaling for hybrid architectures combining Transformer and State-Space Model (SSM) blocks, a class adopted by several recent production language models. In particular, we focus on the gap between the theoretical scaling rules derived for SSMs under zero-order-hold (ZOH) discretization at infinite width with proportional state size, and the field-standard practical implementations using simplified-ZOH Mamba at fixed state size. Surprisingly, in this practical regime hybrid architectures achieve a near-zero LR transfer gap across widths 256-2048 and depths 4-32 at sub-billion-parameter scale using only the original $\mu$P prescription, even though SSM operations fall outside its Tensor Programs representability conditions and every parameterization we test fails the standard coordinate-check diagnostic of $\mu$P correctness. We attribute this to a two-condition decomposition of LR transfer in hybrid architectures: a global update-to-weight invariance, enforced by $\mu$P's initialization and LR scaling; and a local per-component balance, provided by AdamW's $g/\sqrt{v}$ per-parameter normalization. Our observations show that the optimal LR is invariant to width up to $8\times$, as well as to depth, sequence length, batch size, and Transformer-to-SSM ratio, and transfers to Nemotron-H, a production hybrid outside our custom architecture set. We hope these findings fill the gap between theoretical scaling rules and practical hybrid implementations, and stimulate further research toward bridging it.
Learning in neural systems arises from synaptic changes that reshape the representations underlying behavior. While low-rank recurrent neural networks (RNNs) have emerged as a powerful framework for linking connectivity to function, a theoretical understanding of their learning process remains elusive. Here, we extend the low-rank framework from activity to learning by deriving gradient-descent dynamics directly in a reduced overlap space. We formulate a closed-form, low-dimensional system of ODEs that governs learning in this space, exact for linear RNNs and asymptotically exact for nonlinear RNNs in the large-$N$ Gaussian limit. Central to our analysis is a distinction between two classes of overlaps: _loss-visible_ overlaps, which fully determine network activity, output, and loss, and _loss-invisible_ overlaps, which do not affect function but are required to describe learning. We illustrate the consequences of this decomposition through two phenomena. First, we show that learning can serve as a perturbation that exposes differences in connectivity between functionally equivalent networks. Second, we show that loss-invisible overlaps can act as memory variables that encode training history, and characterize the conditions under which this occurs. Finally, we present several testable predictions for biological learning experiments derived from our theory.
Learning Scenario Reduction for Two-Stage Robust Optimization with Discrete Uncertainty
TIANJUE LIN ⋅ Jianan Zhou ⋅ Jieyi Bi ⋅ Yaoxin Wu ⋅ Wen Song ⋅ Zhiguang Cao ⋅ Jie Zhang
Two-Stage Robust Optimization (2RO) with discrete uncertainty is challenging, often rendering exact solutions prohibitive. Scenario reduction alleviates this issue by selecting a small, representative subset of scenarios to enable tractable computation. However, existing methods are largely *problem-agnostic*, operating solely on the uncertainty set without consulting the feasible region or recourse structure. In this paper, we introduce PRISE, a *problem-driven* sequential lookahead heuristic that constructs reduced scenario sets by evaluating the marginal impact of each scenario. While PRISE yields high-quality scenario subsets, each selection step requires solving multiple subproblems, making it computationally expensive at scale. To address this, we propose NeurPRISE, a neural surrogate model built on a GNN-Transformer backbone that encodes the per-scenario structure via graph convolution and captures cross-scenario interactions through attention. NeurPRISE is trained via imitation learning with a gain-aware ranking objective, which distills marginal gain information from PRISE into a learned scoring function for scenario ranking and selection. Extensive results on three 2RO problems show that NeurPRISE consistently achieves competitive regret relative to comprehensive methods, maintains strong scalability with varying numbers of scenarios, and delivers $7$--$200\times$ speedup over PRISE. NeurPRISE also exhibits strong zero-shot generalization, effectively handling instances with larger problem scales (up to $5\times$), more scenarios (up to $4\times$), and distribution shifts.
Learning Scene-Grounded Interaction Priors for Scene-Aware Human Motion Prediction
Donghyeon Soon ⋅ Yoon JI- Il ⋅ Hyeondong Kim ⋅ Daehee Park
Human motion prediction (HMP) aims to forecast future movement from an observed motion sequence, and recent methods have incorporated 3D scene geometry and object semantics to improve long-term plausibility. However, the destination of future motion remains ambiguous in cluttered scenes: the same observed motion can lead to different interaction targets or navigation modes. Existing methods either handle this ambiguity implicitly or rely on auxiliary signals like gaze that are hard to obtain in practice. To address this, we propose SGIP, which explicitly predicts where in the scene future motion will be grounded across both objects and open navigation regions, and uses it to condition future trajectory and full-body pose decoder. This scene-grounded interaction prior is learned via intent distillation: at training time only, a teacher model scores candidate scene regions by combining an LLM prompted with a textual description of the person's intent with geometric motion-scene features. A student model learns to reproduce these scores from only the observed motion and 3D scene, making intent and LLM unnecessary at test time. Across multiple benchmark datasets, SGIP outperforms prior scene-aware baselines on both trajectory and pose accuracy in a realistic deployment setting where neither gaze nor intent is available at inference, demonstrating effectiveness of the proposed framework.
Learning Spectral Compositional Koopman Operators for Global-to-Regional Weather Forecasting
khalid OUBLAL ⋅ Malo Guichard ⋅ François Bertholom ⋅ Simon Albergel ⋅ David Benhaiem ⋅ Emmanuel LE BORGNE ⋅ Vladimir Kostic ⋅ Karim Lounici
State-of-the-art machine learning weather prediction systems achieve strong upper-air forecast skill but degrade at longer lead times, with mesoscale structures progressively vanishing and ensemble forecasts becoming under-dispersive. While these limitations are often attributed to the training objective, we find that message-passing processors significantly contribute to this degradation. Repeated neighborhood aggregation at each autoregressive step acts as a smoothing operator that, over rollout steps, suppresses high-frequency information in latent space. This effect is particularly limiting in regional weather prediction, where fine-scale structures are already partially unresolved, further eroding mesoscale variability and hindering long-horizon forecasting. We propose GraphMet, a graph-based weather model motivated by Koopman theory that replaces stacked message-passing layers with a single linear Koopman operator step. GraphMet learns linear evolution in a spectral latent basis while preserving the encoder, decoder, and multi-scale icosahedral mesh structure. A two-term spectral objective jointly learns a stable Koopman basis and preserves latent geometry, enabling efficient long-range propagation without repeated nonlinear aggregation. Empirically, GraphMet reduces the 15-day activity and improves regional RMSE by 11-14% at 5km resolution, and enables efficient probabilistic forecasting with lower inference cost than state-of-the-art methods. Our code available at : https://anonymous.4open.science/w/GraphMet-54ED/.
Survival analysis provides statistical methods to model the time until an event occurs. Reporting delays arise when event times are not observed at their occurrence but are only revealed upon reporting. This issue is particularly critical for timely risk evaluation when the observation window is short due to administrative censoring. In this study, we incorporate right-censored reporting delays by jointly modeling parametric hazards for the event and reporting processes. We then construct a consistent estimator for the model parameters and develop a Monte Carlo expectation-maximization algorithm to compute it. To address the challenges posed by administrative censoring, we leverage these findings and propose a transfer-learning procedure. Experimental results demonstrate that our method improves the accuracy of timely risk evaluation under administrative censoring.
When auditing machine learning models, human experts sequentially inspect model predictions and explanations, gradually building a mental model of the model’s behavior to uncover errors. This manual process is critical for safe deployment in high-stakes domains but fundamentally limited in scalability. In this work, we propose a novel method to simulate this human auditing behavior for automating the auditing process. We formalize auditing as a sequential decision-making problem, modeled as a Markov decision process. Our state representation summarizes the auditor’s accumulated knowledge of model predictions and feature usage. Building on this formalization, we solve the MDP using reinforcement learning, yielding RLAuditor, which learns to select informative test samples to efficiently uncover model errors. We further connect our formalization to Theory of Mind, showing that our state mirrors how human auditors build mental models: its update rule is equivalent to Bayesian belief revision and it converges to the model’s behavioral pattern at an interpretable rate. We validate our approach on multiple ML models and datasets across different data modalities. In a human subject study, we demonstrate that RLAuditor helps human auditors produce more accurate audit reports in less time.
Learning to Group and Order: Cross-Instance Self-Supervised RL for Vision-Centric MLLMs
Songming Yang ⋅ Hong-Tao Yu ⋅ Xiu-Shen Wei
Reinforcement learning (RL) post-training has become an important paradigm for improving multimodal large language models (MLLMs), especially with automatically verifiable rewards. For vision-centric MLLMs, self-supervised visual pretext tasks provide such rewards without human annotations. However, existing self-supervised RL methods mostly exploit within-instance structure, such as spatial or temporal ordering within a single image or video, offering limited supervision for cross-instance discrimination and fine-grained comparison. We introduce Group-and-Order Self-Supervised Reinforcement Learning (GO-SSL), a cross-instance jigsaw framework for vision-centric MLLM post-training. Given two visual instances, GO-SSL mixes their local elements and trains the model to group them by instance of origin while recovering the order within each group, coupling instance-level comparison with spatial, temporal, or geometric structure recovery. This paired formulation turns the same unlabeled data into richer verifiable supervision through diverse pairings, shuffles, and hard-pair curricula. We instantiate GO-SSL mainly on cross-image jigsaw tasks and further extend it to 3D depth and video temporal ordering. Across diverse benchmarks, GO-SSL consistently improves over the base MLLM and prior baseline methods, with clear gains in fine-grained perception and spatial understanding. These results suggest a data-efficient direction for self-supervised RL in MLLMs: constructing comparative cross-instance contexts can provide richer transferable supervision than merely scaling data volume or adding isolated pretext tasks. Code will be released later.
Learning-to-Memorize: Dynamic Context Management for Long-Horizon Autoregressive Video Generation
Haowei Zhu ⋅ Qijie Wang ⋅ Jia Li ⋅ XING WANG ⋅ Tianyu Zhao ⋅ Jiexi Wang ⋅ Xurui Peng ⋅ Xinglong Wu ⋅ Jun-Hai Yong ⋅ Bin Wang
Autoregressive video generation has achieved impressive results on short clips, yet generating long videos remains challenging due to error accumulation over extended horizons, where each predicted frame depends on potentially imperfect previous frames. Existing approaches typically rely on static context management strategies, such as sliding-window caches that discard early frames or fixed sink frames that permanently preserve initial content. These designs either lose critical long-range context or anchor generation to outdated information, which can lead to degraded fidelity, motion stagnation, and reduced diversity. In this paper, we propose \textbf{Learning to Memorize (L2M)}, a learned context memory management framework that dynamically retains, evicts, and evolves historical frames according to predicted importance. The framework consists of an \emph{importance prediction router} that estimates the relevance of each historical frame, \emph{degradation-aware training} that improves robustness to noisy or corrupted context, and a \emph{dynamic memory initialization and evolution} mechanism that updates long-term memory using exponential moving average (EMA) importance scores. This allows high-quality frames to replace stale anchors and enables memory to adapt naturally to scene evolution. Trained with a teacher-student distillation objective and sparsity regularization, L2M makes efficient use of a fixed KV-cache budget while preserving the most informative contextual information. Extensive experiments show that L2M achieves superior long-term video generation performance over existing baselines, improving stability and visual fidelity, and establishing a new paradigm for learned memory management in long autoregressive video generation.
Learning to Read Out: Unembedding Dynamics in Language Model Pretraining
Matteo He ⋅ William Shen ⋅ Alexandru-Andrei Iacob ⋅ Andrej Jovanović ⋅ Xinchi Qiu ⋅ Nicholas Lane
When does a language model acquire a capability? When its hidden states encode the relevant information, or when its output readout can express it? We investigate this gap between representational availability and readout expression by tracking the unembedding matrix through pretraining. Because unembedding rows are aligned by token identity across checkpoints, they provide one trajectory per vocabulary item through the learned output interface. To measure these trajectories, we introduce parameter-trajectory crosscoding, a checkpoint-spanning sparse dictionary with shared feature identities and checkpoint-specific decoders. Applied across Pythia scales and in OLMo-2-7B, this method reveals an early, heterogeneous reorganization of the readout. This reorganization does more than alter structure. Our independent WordNet probes show that token families become increasingly separable over the same developmental window. We then causally test whether the readout component drives this development. By swapping the unembedding matrices across checkpoints, and controlled contrastive tasks, we show that task-relevant distinctions can be present in hidden states before the contemporaneous readout expresses them in logits. By ablating individual crosscoder features, we trace these delayed readout effects directly to compact, learned directions. More broadly, developmental analyses should distinguish representational availability from readout expression. Together, our findings demonstrate that the learned output interface helps determine when latent structure becomes visible in token logits, highlighting the need for developmental analyses to explicitly decouple hidden-state representation from readout expression.
We study preference learning and coordination through recommendations in multi-agent game settings, where a moderator repeatedly interacts with agents whose utility functions are unknown. In each round, the moderator issues action recommendations and observes whether agents follow or deviate. We consider agents who best respond to the moderator's recommendations and study what this feedback reveals about their utilities. We characterize the class of games that are indistinguishable under this feedback model. Moreover, we introduce a notion of moderator regret based on agents' incentives to deviate from the recommendations and design an online algorithm with low regret under the best-response model, with guarantees that scale linearly in the game dimension and logarithmically in time. Our results lay a theoretical foundation for AI recommendation systems in strategic multi-agent environments, where recommendation compliance is shaped by strategic interaction.
Learning to Surpass: Training Tool-Using Agents with Anchored Feedback
John Gkountouras ⋅ Fengjun Wang ⋅ Angelantonio Castelli ⋅ Satendra Kumar
Training tool-using agents for planning is challenging when high-quality demonstrations exist only as final outcomes without the intermediate tool traces that produced them. A common response is to impute traces and apply imitation learning, but this approach is indirect and structurally limited to matching the reference rather than improving upon it. We propose anchored comparative feedback: rather than imitating a reference plan, we train agents to surpass it under a comparative LLM judge while satisfying hard constraints. The anchor provides a stable optimization target, and comparing two identically rendered plans neutralizes surface-level judge biases. On TripTailor, a tool-augmented travel planning benchmark, anchored GRPO achieves 80.1% final success rate compared to 43.7% for imitation learning on imputed traces, with gains on both constraint satisfaction and LLM-judged preference that persist across judge families, rubrics, and a held-out reward model. Gains generalize beyond TripTailor, with a fivefold improvement on the held-out TravelPlanner benchmark. These results suggest that anchored comparative feedback offers an effective approach for learning transferable tool-using planners from outcome-only supervision.
Learning Visual Speech Representations via Cross-Modal Distillation and Joint Face-Lip Modeling
Jingxuan Zhang ⋅ Genshun Wan ⋅ Jia Pan ⋅ Jianqing Gao ⋅ Yumei Zhang ⋅ Lichen Zhang
Lipreading remains challenging due to the inherent ambiguity of visual speech cues and their sensitivity to head pose and video resolution. In this work, we propose FLARE, a framework for robust visual speech representation learning via cross-modal distillation and joint face-lip modeling. FLARE leverages semantically rich Whisper representations as the acoustic teacher signal, converting them into discrete distillation targets through a random-projection quantizer and pretraining the model with a masked prediction objective. To capture complementary visual information, FLARE employs a dual-stream architecture that jointly processes lip region-of-interest (ROI) and full-face inputs, where the lip stream captures fine-grained articulatory motion and the face stream encodes broader facial dynamics. During pretraining, auxiliary audio input is incorporated to facilitate cross-modal alignment, and modality dropout is applied to encourage robust visual-only representations. For downstream lipreading, FLARE serves as a visual encoder coupled with either a lightweight Transformer decoder or an LLM-based decoder. Evaluated on LRS2, LRS3, and the out-of-domain WildVSR benchmark, FLARE achieves state-of-the-art word error rates of 11.7%, 13.2%, and 34.3%, respectively. A progressive design study confirms the contribution of each major component. Further analyses demonstrate that joint face-lip modeling improves robustness under degraded video resolution and large head pose variation, while yielding more phoneme-discriminative visual representations. Code and pretrained models are available at https://anonymous.4open.science/r/FLARE-522C.
Learning What to Forget: Improving LLM Unlearning via Learned Token-Level Importance
Gizem Yüce ⋅ Giorgos Nikolaou ⋅ Nicolas Flammarion
Machine unlearning aims to remove targeted knowledge from a trained model while preserving its general capabilities. Doing so effectively for auto-regressive language models requires identifying which tokens in a forget sample are truly relevant to forgetting, as uniformly applying unlearning across all tokens can degrade utility. Existing approaches, however, either ignore this heterogeneity or rely on auxiliary models, hand-crafted heuristics, or external annotations to approximate token-level relevance. We instead characterize token-level forget relevance through its interaction with the retain objective: tokens are forget-relevant to the extent that minimizing the forget loss on them does not conflict with the retain objective. Based on this insight, we introduce Alternating Token-Weighted Unlearning (ATWU), a framework that jointly learns token relevance and model parameters during the unlearning process. ATWU uses a lightweight linear scorer over model hidden states to predict forget-relevance scores with negligible computational overhead and no external supervision. Experiments on TOFU and RWKU show that ATWU achieves state-of-the-art forget--retain trade-offs, outperforming sample-level methods, probability-based token-weighting heuristics, and auxiliary-model-based approaches. Moreover, the learned scorer aligns with ground-truth forget-relevant spans substantially better than existing methods, suggesting that ATWU learns semantically meaningful token-level forgetting signals. Overall, ATWU shows that token-level forget-relevance can be effectively inferred from model representations during unlearning, enabling efficient and selective forgetting.
Learning When to Stop: Selective Imitation Learning Under Arbitrary Dynamics Shift
Surbhi Goel ⋅ Jonathan Pei ⋅ James Wang
Behavior cloning provides strong imitation learning guarantees when training and test environments share the same dynamics. However, in many deployment settings the test environment's transitions differ from training, and classical offline IL offers no recourse: the learner must commit to an action at every state, even when its demonstrations are uninformative and could lead to arbitrary degradation of performance. This motivates the study of _selective_ imitation, where the learner may choose to _stop_ when it cannot act reliably. We introduce a model for selective imitation under arbitrary dynamics shift: given labeled expert demonstrations from a training environment and unlabeled state trajectories from the same expert in a test environment, the learner outputs a _selective policy_ that is _complete_ (rarely stops in training) and _sound_ (incurs low regret before stopping in test). Our algorithm, $\operatorname{SeqRejectron}$, constructs a stopping rule using a small set of _validator policies_ whose size is independent of the horizon or policy class. For deterministic policies, this yields horizon-free $\tilde{O}(\log|\Pi|/\epsilon^2)$ sample complexity, assuming sparse costs. For stochastic policies, we obtain analogous horizon-free guarantees using a cumulative Hellinger stopping time. We extend the framework to misspecified experts and different expert policies across train and test and obtain results that gracefully degrade with the amount of misspecification.
Learning When to Trust LLM Priors: A Validated Framework for Semantic Prior Integration
Erica Zhang ⋅ Naomi Sagan ⋅ Danny Tse ⋅ Fangzhao Zhang ⋅ Mert Pilanci ⋅ Jose Blanchet
Large language models (LLMs) encode rich semantic knowledge that can be useful for supervised learning, but their outputs are unreliable as statistical priors: they may be noisy, misspecified, or hallucinated. Existing LLM-informed learning methods either trust such signals directly, leaving predictions vulnerable to unreliable LLM guidance, or restrict semantic integration to a single model class. We introduce Statsformer, a validated framework for learning when to trust LLM-derived semantic priors in supervised statistical learning. Statsformer maps LLM-derived feature scores into a family of learner-specific prior-injection mechanisms across a heterogeneous library of linear and nonlinear predictors. It then uses out-of-fold validation to adaptively calibrate the influence of each prior-informed learner, allowing useful semantic information to improve prediction while attenuating weak, misspecified, or adversarial priors. This yields a guardrailed statistical learning system with an oracle-style guarantee: up to statistical error, the final predictor performs no worse than the best convex combination of its in-library candidates, including prior-free learners. Across diverse prediction tasks, informative LLM priors improve performance, while unreliable priors are automatically downweighted. These results position Statsformer as a reliability-oriented approach to LLM-informed statistical learning: rather than trusting LLM knowledge directly, it validates semantic priors against data before allowing them to influence the final predictor.
Learning Where It Matters: Geometric Anchoring for Robust Preference Alignment
Youngjae Cho ⋅ Jongsuk Kim ⋅ Ji-Hoon Kim
Preference optimization aligns large language models from pairwise preferences by increasing the margin between preferred and dispreferred responses. However, margin-based losses alone do not test whether each pair's preference signal is locally stable enough to trust with full update strength. We propose Geometric Anchor Preference Optimization (GAPO), a geometry-aware objective that introduces a batch-conditioned stress test for preference learning. For each mini-batch, GAPO constructs a pessimistic anchor by perturbing the current policy in the first-order direction that decreases the batch-average preference margin. The resulting Anchor Gap measures how much each pair's margin degrades under this shared pessimistic probe and converts this degradation into an instance-wise update weight. Pairs with larger Anchor Gap receive smaller update weights, while non-brittle pairs retain their preference-gradient direction. Under smoothness assumptions, we characterize the Anchor Gap as a batch-directional proxy for local margin degradation. Across multiple open-weight model families, GAPO matches or improves strong preference-optimization baselines on instruction-following and reasoning benchmarks. It also improves robustness under random and structured preference noise without explicitly modeling label corruption. Mechanistic diagnostics show that the pairs receiving the smallest GAPO weights are statistically enriched in corrupted supervision, suggesting that GAPO improves robustness by reducing the cumulative influence of brittle preference signals.
LeCellModel: Interpretable Density Estimation over the Gene Expression Manifold
Gil Karin ⋅ Artemy Bakulin ⋅ Nir Yosef
Representation learning for single-cell RNA-seq has been evaluated by the identities representations capture — cell types, lineages, perturbations — rather than how faithfully they model the distribution of cell states on the gene expression manifold. As a result, density estimation, rare-state detection, perturbation scoring, and gene-level attribution are bolted onto frozen embeddings as separate models, disconnected from the geometry of the data. We argue that the density model should be the primary object and that representation quality follows from it. Building on LeJEPA, whose SIGReg objective enforces a provably optimal isotropic Gaussian latent, we introduce LeCellModel, which exploits this Gaussian latent to recover data density on the gene expression manifold via the encoder's Jacobian log-volume (JEPA-SCORE). Across several datasets, LeCellModel outperforms established methods on scIB benchmark for representation quality. Furthermore, we show that LeCellModel enables estimation of rare expression states using JEPA-SCORE as a typicality measure – applying it to a controlled cytokine panel, where it identifies stimulated cells across minority fractions and tracks pathway-response magnitude. We introduce gene-level interpretability of density estimates — to our knowledge, the first method for analyzing JEPA-SCORE in feature space. Leveraging this framework, LeCellModel recapitulates the known association of a profibrotic macrophage program with idiopathic pulmonary fibrosis, and discovers a novel CD177$^{+}$ activated regulatory T cell subpopulation previously described only in tumor contexts. Both findings emerge directly from the geometry of the learned embedding, without relying on supervised annotations.
LEGO: Sizing Rules for Budget-Aware Dense-to-MoE Conversion of Vision-Language Models
Qishen Yin ⋅ ZiangWu ⋅ Juntong Wu ⋅ Peng Jin ⋅ Li Hao ⋅ Tanghui Jia ⋅ Liuhan Chen ⋅ Bin Zhu ⋅ Li Yuan
We propose LEGO, a budget-aware Dense-to-MoE methodology that converts pretrained vision-language models into efficient FFN-MoE variants under a target architectural budget $((P_{total},P_{active}))$. Our study is motivated by a counter-intuitive recovery pattern: fixed activation ratios that work well for some dense backbones can degrade sharply on others under the same recovery recipe, indicating that raw sparsity alone is an unreliable design rule. Through controlled recovery experiments across VLM families and scales, we identify the activated-FFN-to-hidden ratio $(D_I/D_H)$ as a backbone-comparable structural factor and summarize empirical sizing rules for avoiding severe bottlenecks while limiting diminishing returns. LEGO operationalizes these rules through a discrete budget-aware search over Split, Upcycle, and Hybrid constructions, and introduces a moment-matching scaling factor to reduce the initialization scale shift caused by sparse FFN aggregation. Finally, we use a two-stage multimodal recovery recipe that first stabilizes the sparse LLM with the vision tower frozen, then performs joint fine-tuning to recover the full VLM. Under matched data and training protocols, LEGO improves Dense-to-MoE recovery quality and turns costly blind architecture search into a small set of rule-guided candidates. Codes can be found in supplementary materials.
Leveraging unlabelled data for generalizable neural population decoding
Ximeng Mao ⋅ Nanda H Krishna ⋅ Hee-Woon Ryoo ⋅ Matthew G Perich ⋅ Guillaume Lajoie
Robust and accurate neural decoders are integral to the development of neurotechnological applications such as brain-computer interfaces and closed-loop experiments. Recent work has shown that tokenizing neural data at the resolution of individual spikes facilitates multi-session pretraining and delivers state-of-the-art decoding performance. However, these spike-based models are currently restricted to supervised learning (SL) regimes, limiting both pretraining and finetuning to datasets with paired behavioural labels. To address this, we introduce MOJO ($\textbf{M}$asked aut$\textbf{O}$encoder-based $\textbf{JO}$int training), a training framework designed for spike-tokenizing models that jointly leverages self-supervised learning (SSL) via masked autoencoding and SL objectives for better decoding performance and interpretability. We evaluate MOJO on three intracortical spiking datasets—monkey motor cortex across various reaching tasks, multi-regional mouse recordings during vision and cognitive decision making tasks—demonstrating superior performance to purely SL-trained models. This improvement is especially notable when training with limited labelled data, specifically in the few-shot finetuning regime where only a small amount of labelled data is available to adapt a model to a new recording session. The incorporation of SSL also yields more interpretable neuronal representations, improving performance on analyses such as brain region classification and spike-statistics prediction despite the lack of explicit optimization for these tasks. We then show that MOJO generalizes beyond spiking data using human electrocorticography (ECoG) during speech articulation. We show that MOJO continues to outperform purely SL-trained models and achieves performance comparable to neuro-foundation models (NFMs) built specifically for continuous signals. Overall, augmenting spike-tokenizing models with SSL improves their performance in label-impoverished settings and enables the use of unlabelled data across various tasks and species, and while generalizing to other neural modalities. These results suggest a path towards more flexible and scalable data usage when training NFMs.
LINC: Decoupling Local Consequence Scoring from Hidden Matching in Constructive Neural Routing
ShaoFeng Qin ⋅ Li Wang
Constructive neural routing solvers usually score the next action by matching a decoder context to candidate embeddings, hiding deterministic one-step consequences such as travel, waiting, slack, and capacity changes. We propose LINC (Local Inference via Normed Comparison), a decoder-side candidate decision architecture that computes these consequences explicitly. LINC uses them according to their decision role: centered relative consequences are compared by a shared linear local scorer, while feasible-set summaries modulate the decoder context. This preserves standard global matching and relieves the hidden state from rediscovering transition arithmetic. The Capacitated Vehicle Routing Problem with Time Windows (CVRPTW) serves as the main constrained-routing stress test; the same interface extends to the Capacitated Vehicle Routing Problem (CVRP) and Traveling Salesman Problem (TSP). In particular, for CVRPTW, LINC reduces PolyNet's Solomon/Homberger gaps from 13.83\%/38.15\% to 7.26\%/14.71\%; for TSP and CVRP, it also improves external-benchmark gaps.
Ensemble sampling is a practically attractive approach to randomized exploration, but existing theoretical guarantees for linear bandits require ensembles much larger than what its practical motivation would suggest. In particular, the sharpest existing analysis achieves the ensemble complexity of $\Theta(d\log T)$, leaving a $\log T$ gap from the intrinsic $\Omega(d)$ ensemble-size barrier. We aim to narrow this gap by proposing an ensemble sampling algorithm that refreshes the ensemble only when the regularized Gram matrix changes substantially. This mechanism localizes the perturbation analysis to epochs with controlled Gram-matrix drift and reduces the sufficient ensemble size to $\Theta(d\log d+d\log\log T)$, while preserving the state-of-the-art $\tilde{\mathcal O}(d^{3/2}\sqrt T)$ regret for ensemble sampling with arbitrary bounded arm sets. We further show that, when the arm set is finite of cardinality $K$, the proposed algorithm achieves the sharper regret bound $\tilde{\mathcal{O}}(d\sqrt{T\log K})$. To the best of our knowledge, this is the first ensemble-sampling guarantee that simultaneously recovers both canonical regret scalings known for randomized linear bandit algorithms: the $\tilde{\mathcal{O}}(d^{3/2}\sqrt{T})$ rate for arbitrary bounded arm sets and the $\tilde{\mathcal{O}}(d\sqrt{T\log K})$ rate for finite arm sets. The algorithm also admits an anytime implementation without resetting past data, and experiments show that it remains competitive with baselines while using substantially smaller ensembles.
LinuxArena: A Control Setting for AI Agents in Live Production Software Environments
Tyler Tracy ⋅ Ram Potham ⋅ Nikolas Kuhn ⋅ Myles Heller ⋅ Anshul Khandelwal ⋅ Cody Rushing ⋅ Henri Lemoine ⋅ Miguel Brandão ⋅ Tomáš Turlík ⋅ Adam Hanson ⋅ Josh Hills ⋅ Amy D Ngo ⋅ Ram Rachum ⋅ Nik Mitchell ⋅ Falko Galperin ⋅ Oscar Sykes ⋅ Pip Arnott ⋅ Samuel P Lima ⋅ Carlos R Giudice ⋅ Matt Goldwater ⋅ Daniel J Popp ⋅ Drew de Wet ⋅ Ruben Castaing ⋅ Qi Guo ⋅ Douw Marx ⋅ Benjamin Shaffrey ⋅ Justin Shenk ⋅ Martin Milbradt ⋅ Hannah Meagher ⋅ Shaheen Ahmed-Chowdhury ⋅ Daniel O'Connell ⋅ Christopher W Canal ⋅ Buck Shlegeris ⋅ Aryan Bhatt
As AI agents are given more autonomy in software engineering workflows, the risks grow if they pursue goals different from the ones their users intended. The field of AI control develops control protocols that prevent such harm without restricting the agent's ability to do useful work. We introduce \textbf{LinuxArena}, a control setting for testing such protocols, where agents operate directly on live, multi-service production environments. LinuxArena contains 10 public environments (with an additional 10 held out privately), providing 906 main tasks representing legitimate software engineering work and 92 side tasks representing safety failures such as data exfiltration and access-control bypassing, making it the largest and most diverse control setting for agentic software engineering to date. We demonstrate LinuxArena's utility by running sabotage and monitor evaluations: against a GPT-5 Nano trusted monitor at a 1% step-wise false positive rate, Claude Opus 4.6 achieves an undetected sabotage rate of 42%. We additionally release \textbf{LaStraj}, a dataset of human-crafted attack trajectories that achieves an undetected sabotage rate of 97% against the same monitor. Together, these results suggest meaningful headroom for both attackers and defenders, making LinuxArena a strong testbed for developing and evaluating future control protocols.
LLM-Auction: Generative Auction towards LLM-Native Advertising
Chujie Zhao ⋅ Qun Hu ⋅ Shiping Song ⋅ Dagui Chen ⋅ Han Zhu ⋅ Jian Xu ⋅ Bo Zheng
The commercialization of LLM applications is the next frontier in online advertising, with LLM-native advertising emerging as a promising paradigm by integrating ads into LLM-generated content. However, classic mechanisms are no longer applicable in this setting where the auction object is shifted from discrete ad slots to distributions over LLM outputs, and existing methods are impractical in industrial scenarios due to ignored externalities or high inference costs. To address these issues, we propose LLM-Auction, the first learning-based generative auction mechanism that integrates auction and generation. By formulating the allocation as preference alignment between LLM outputs and a mechanism objective that balances advertiser value and user experience, we optimize the LLMs to inherently model allocation externalities without extra inference cost. Theoretically, we identify the allocation monotonicity and continuity of LLM-Auction, and prove that a simple first-price payment rule exhibits favorable incentive properties. Furthermore, we build an LLM-as-a-judge simulation environment for quantitative evaluation, and experiments demonstrate that LLM-Auction achieves the state-of-the-art allocation efficiency while satisfying key mechanism properties.
Recent work has demonstrated surprisingly good performance of pre-trained LLMs on regression tasks (for example, time-series prediction), with the ability to incorporate expert prior knowledge and the information contained in textual metadata. However we observe major error cascades even in short sequences $\lesssim 100$ points; these models are also computationally intensive and difficult to parallelise. Marginal LLM predictions do not suffer this issue and are trivially parallelised, but can predict over-broad densities. To address this, we propose combining these densities with a lightweight (diffusion-based) neural process. We show that this combination leads to better-calibrated predictions overall, outputs locally consistent trajectories, and leads to text-conditioned function space selection in the meta-learner. As part of this work we propose a gradient-free (and non-Monte Carlo) method for sampling from a product-of-experts of a score model and an `expert' (here the LLM predictive densities). We believe this general method is of independent interest as it is applicable whenever an expert can be convolved with a Gaussian in closed form.
LLMs Optimizing LLMs: Automated MegaKernel Generation for Inference Acceleration
Weiqiang Xiong ⋅ Shaohui Peng ⋅ Wenyi Li ⋅ Hao Lu ⋅ Qirui Zhou ⋅ Congying Ma ⋅ Ziming Ye ⋅ Zhenyu Yi ⋅ Yunji Chen ⋅ Qi Guo ⋅ Ling Li
High-performance inference for large models is critical in latency-sensitive applications. Existing inference engines suffer from sequential execution of hundreds of fine-grained kernels, causing kernel launch overheads and redundant memory accesses. A promising solution is fusing multiple kernels into a single large kernel, namely MegaKernel. However, existing MegaKernel implementations are tightly coupled to specific models and GPU architectures and require extensive manual engineering. We propose an LLM-driven framework AutoMegaKernel, which leverages an Instruction-Centric Fusion Abstraction to decompose MegaKernels into tractable fusion instructions, thereby enabling automated generation. Our framework jointly performs fusion strategy reasoning and CUDA code generation while hierarchically verifying correctness. Extensive experiments demonstrate that AutoMegaKernel consistently outperforms state-of-the-art inference engines (up to 5.04$\times$ vLLM and 2.42$\times$ SGLang) and compilers (up to 1.92$\times$ MPK), while supporting a broader range of models (dense, MoE, and VLA) and NVIDIA GPU platforms (Hopper, Ampere, and Ada).
LLM-WikiRace: A Benchmark for Planning and Reasoning over Real-World Knowledge Graphs
Juliusz Ziomek ⋅ William Bankes ⋅ Lorenz Wolf ⋅ Shyam Sundhar Ramesh ⋅ Xiaohang Tang ⋅ Ilija Bogunovic
We introduce LLM-Wikirace, a benchmark for evaluating planning, reasoning, and world knowledge in large language models (LLMs). In LLM-Wikirace, models must efficiently navigate Wikipedia hyperlinks step by step to reach a target page from a given source, requiring look-ahead planning and the ability to reason about how concepts are connected in the real world. We evaluate a broad set of open- and closed-source models, including Gemini-3, GPT-5, and Claude Opus 4.5, which achieve the strongest results on the easy level of the task and demonstrate superhuman performance. Despite this, performance drops sharply on hard difficulty: the best-performing model, Gemini-3.1, succeeds in only 29\% of hard games, highlighting substantial remaining challenges for frontier models. Our analysis shows that world knowledge is a necessary ingredient for success, but only up to a point, beyond this threshold, planning and long-horizon reasoning capabilities become the dominant factors. Trajectory-level analysis further reveals that even the strongest models struggle to replan after failure, frequently entering loops rather than recovering. LLM-Wikirace is a simple benchmark that reveals clear limitations in current reasoning systems, offering an open arena where planning-capable LLMs still have much to prove.
Modern neural-network training is governed by local interactions: residual streams carry signal, normalization changes scale, attention and state-space blocks route information, augmentation changes which directions generalize, and optimizer state determines the actual checkpoint update. Parameter-space curvature and full-weight posterior views reveal global geometry, but they do not show which layer, block, or interface created it. We introduce a computation-graph view in which parameters $\theta$ and intermediate states $u$ define a local Markov--Gibbs system. A potential energy encodes module relations, data loss, and regularization; its Gibbs law gives a probability reference, and its checkpoint-local Fisher/Gauss--Newton precision matrix gives the curvature geometry. This precision matrix is sparse because variable groups interact only when they share a local module or loss factor. Eliminating intermediate states recovers the reduced parameter curvature. From this common object we prove two results: separator transfers give two-sided criticality certificates for upper stability and active plasticity, and a resolved KL-to-Gibbs bound yields bounded-loss PAC-Bayes certificates on directions explored by training. Checkpoint measurements on CIFAR-10 and GPT-20M/FineWeb-Edu show that these quantities diagnose learning-rate stability, plasticity, and generalization geometry without dense Hessian, Fisher, or covariance matrices.
LOCU: Löwdin-Orthogonalized Constraint Updates for Multi-Constraint Policy Optimization
Joonyoung Lim ⋅ Younghwan Yoo
We investigate multi-constraint policy optimization in constrained Markov decision processes (CMDPs), where interactions among constraint gradients often yield ill-conditioned update directions, causing numerically fragile steps and unintended cancellation or redundant overlap in constraint corrections. To address this, we propose LOCU, which decouples constraint interactions by applying symmetric L\"owdin orthogonalization to the natural gradients of the reward and constraints under the Fisher–Rao metric. Combined with Fisher-geometric screening and near-collinear compression, LOCU yields a compact reduced space in which the trust region becomes a Euclidean ball and constraints are handled symmetrically through a low-dimensional active-set solve. For infeasible iterates, LOCU construct a Pareto-descent direction that simultaneously reduces violated constraints while preserving satisfied ones within the same reduced formulation. Experiments show that LOCU remains stable under high constraint coupling, narrow feasible regions, and infeasible starts, and generalizes across different cost critic architectures.
Logit-Conditioned Diffusion Decoding for Frozen Discrete-Token VLMs
Ji Woo Hong ⋅ Hee Suk Yoon ⋅ Gwanhyeong Koo ⋅ Eunseop Yoon ⋅ SooHwan Eom ⋅ Qi Dai ⋅ Chong Luo ⋅ Chang Yoo
Pre-trained discrete-token vision-language models (VLMs) generate images by sampling sequences of codebook indices, but their visual fidelity is bounded by the VQ-VAE decoding stage. Prior work addresses this bottleneck by replacing the discrete tokenizer with continuous or hybrid visual token representations, or by integrating diffusion into the generation pipeline; both routes require foundation-scale retraining or realignment of VLM, which is undesirable when the VLM's broader capabilities should be preserved or when retraining is infeasible. We show that such retraining is unnecessary: before each token is sampled, discrete-token VLMs already compute a full logit distribution over the codebook, but standard decoding discards it after selecting a single index. Under Otsu's threshold, the pre-sampling logits of capable VLMs assign statistically meaningful probability to at least two codebook entries per token, and recovering this distribution can substantially close the fidelity gap. We propose DCDD, a post-hoc framework that conditions a diffusion decoder on this discarded distribution while leaving the VLM frozen. At inference, DCDD converts the VLM's pre-sampling logits into distribution-weighted code vectors; during training, it uses VQ-VAE encoder-derived proxy logits calibrated to match the distributional statistics of the VLM's inference-time logits. Our diffusion decoder then maps these representations to high-fidelity images as an optional alternative to the native VQ-VAE decoder. Trained on ImageNet-1K for 50K steps, DCDD reduces reconstruction FID by 74\% and Janus-Pro 7B's generation FID on MJHQ-30K by 34\%, while semantic-alignment across GenEval, DPG-Bench, and WISE remain consistent. Comparison against a diffusion model from the same backbone family applied as an image-to-image refiner confirms that the gain comes not from the diffusion model alone, but from conditioning on the VLM's pre-sampling logit distribution.
Long-Context Generation Is a Sampling Problem
Allen Roush ⋅ Minh N Nguyen ⋅ Ravid Shwartz-Ziv ⋅ Judah Goldfeder ⋅ Sanjay Basu
Long-form generation in language models often collapses into repetitive loops well before the nominal context limit. This is commonly attributed to model capacity or training data, and addressed through costly retraining or post-training pipelines. We show that the sampling algorithm is a primary determinant of long-context output quality, improving outputs without further training costs. Across ten open-weight models from 1B to 120B parameters, output lengths from 8K to 64K tokens, and more than 13,500 generations, full-distribution-aware samplers (P-less, Top-H, top-$n\sigma$) dominate automated metrics and human evaluation, while standard truncation (top-$p$, top-$k$) collapses into late-context repetition; min-$p$ falls in a middle tier. A 1,500-annotation Amazon Mechanical Turk study finds P-less winning more than nine in ten head-to-head pairings against a top-$p$ 0.9 baseline on a 120B model, with 10-gram repetition approaching zero at 64K tokens. A per-step selected-token rank diagnostic isolates the mechanism: standard truncation drives the rank to the argmax floor by the final quartile, while full-distribution-aware methods stay in a healthy band across the full window. This gap that widens with model scale, establishing sampler choice as an under-appreciated bottleneck for long-context generation in open-weight models. We frame long-context generation as sampling from outside the training distribution, and encourage further study for usecases like long-context agentic coding.
LongScape: Advancing Long-Horizon Embodied World Models with Context-Aware MoE
Lei Jin ⋅ Yu Shang ⋅ Yiding Ma ⋅ Xinhao Jin ⋅ Xin Zhang ⋅ Chen Gao ⋅ Wei Wu ⋅ Yong Li
Video-based world models hold significant potential for generating high-quality embodied manipulation data. However, current video generation methods struggle to achieve stable long-horizon generation: classical diffusion-based approaches often suffer from temporal inconsistency and visual drift over multiple rollouts, while autoregressive methods tend to compromise on visual detail. To solve this, we introduce LongScape, a hybrid framework that adaptively combines intra-chunk diffusion denoising with inter-chunk autoregressive causal generation. Our core innovation is an action-guided, variable-length chunking mechanism that partitions video based on the semantic context of robotic actions. This ensures each chunk represents a complete, coherent action, enabling the model to flexibly generate diverse dynamics. We further introduce a Context-aware Mixture-of-Experts (CMoE) framework that adaptively activates specialized experts for each chunk during generation, guaranteeing high visual quality and seamless chunk transitions. Extensive experimental results demonstrate that our method achieves stable and consistent long-horizon generation over extended rollouts. Our code is available at: https://anonymous.4open.science/r/AMSVVD-fdg245.
Low-Dimensional Adaptation of Rectified Flow: A Diffusion and Stochastic Localization Perspective
Saptarshi Roy ⋅ Alessandro Rinaldo ⋅ Purnamrita Sarkar
In recent years, Rectified flow (RF) has gained considerable popularity largely due to its generation efficiency and state-of-the-art performance. In this paper, we investigate how well RF automatically adapts to the intrinsic low dimensionality of the support of the target distribution to accelerate sampling. We show that, using a carefully designed choice of the time-discretization scheme and with sufficiently accurate drift estimates, the RF sampler enjoys an iteration complexity of order $O(k/\varepsilon)$ (up to log factors), where $\varepsilon$ is the precision in total variation distance and $k$ is the intrinsic dimension of the target distribution. In addition, we show that the denoising diffusion probabilistic model (DDPM) procedure is equivalent to a stochastic version of RF by establishing a novel connection between these processes and stochastic localization. Building on this connection, we further design a stochastic RF sampler that also adapts to the low-dimensionality of the target distribution under mild requirements on the accuracy of the drift estimates, and also with a specific time schedule. We illustrate the efficacy of newly designed time-discretization schedules with simulations on the synthetic data and text-to-image (T2I) data experiments.
Lumberjack: Better Differentially Private Random Forests through Heavy Hitter Detection in Trees
Christian J Lebeda ⋅ David Erb ⋅ Tudor Cebere ⋅ Aurélien Bellet
Random forests are widely used in fields involving sensitive tabular data, but existing approaches to enforcing differential privacy (DP) typically degrade performance to the point of impracticality. In this paper, we introduce Lumberjack, a differentially private random forest algorithm that achieves substantially higher utility by constructing large random decision trees and then applying aggressive, privacy-preserving pruning to retain only sufficiently populated nodes. A key component of our approach is a novel $(\varepsilon,\delta)$-DP heavy hitter detection algorithm for hierarchical data, whose error is $O_{\varepsilon,\delta}(\sqrt{\log h})$ for trees of height $h$ and may be of independent interest. This favorable scaling enables the use of significantly deeper trees than in prior work, leading to improved expressiveness under privacy constraints. Our empirical evaluation on benchmark datasets shows that Lumberjack consistently outperforms prior DP random forest methods, establishing a new state of the art. In particular, our approach yields substantial improvements in the privacy-utility trade-off for practical privacy budgets. Our findings suggest that more carefully designed DP random forests can close much of the utility gap, highlighting a promising and underexplored direction for future research.
M2A: Synergizing Mathematical and Agentic Reasoning in Large Language Models
JunJian Wang ⋅ Xin Zhou ⋅ Qiran Xu ⋅ Kun Zhan
While reasoning has become a central capability of large language models (LLMs), the reasoning patterns required for different scenarios are often misaligned. Mathematical reasoning typically relies on intrinsic logic to solve closed-world problems in a single response, whereas agentic reasoning requires not only internal reasoning but also multi-turn interaction with external environments, interleaving thought and action. This misalignment prevents mathematical and agentic reasoning from effectively benefiting from each other, often yielding unstable reasoning behavior and only limited performance gains under multi-task learning. In this paper, we propose M2A, a novel paradigm that synergizes mathematical and agentic reasoning via model merging. To avoid overfitting to superficial reasoning patterns under joint training, M2A operates directly in parameter space: it identifies the feature subspace critical for agent behavior, and merges the mathematical reasoning task vector only along its null space, thereby injecting reasoning capability along directions that do not perturb agent behavior. Unlike SFT or RL, M2A requires no additional gradient-update and exposes the merging coefficient as a simple knob for controlling reasoning length. Experiments in a challenging real-world coding agent setting show that our method effectively extends agentic reasoning depth and delivers substantial performance improvements. Applied to a fine-tuned Qwen3-8B, M2A improves its SWE-Bench Verified resolved rate from 44.0\% to 51.2\% without retraining the model.
M3-AD: Reflection-aware Multi-modal, Multi-category, and Multi-dimensional Benchmark and Framework for Industrial Anomaly Detection
Chao Huang ⋅ Yanhui Li ⋅ Hongxi Huang ⋅ Yunkang Cao ⋅ Wei Wang ⋅ Jie Wen ⋅ Zhihua Wang ⋅ Wenqi Ren ⋅ XIAOCHUN CAO
Although multimodal large language models (MLLMs) have advanced industrial anomaly detection toward a zero-shot paradigm, they still tend to produce high-confidence yet unreliable decisions in fine-grained and structurally complex industrial scenarios, and lack effective self-corrective mechanisms. To address this issue, we propose M3-AD, a unified reflection-aware multimodal framework for industrial anomaly detection. M3-AD comprises two complementary data resources: M3-AD-FT, designed for reflection-aligned fine-tuning, and M3-AD-Bench, designed for systematic cross-category evaluation, together providing a foundation for reflection-aware learning and reliability assessment. Building upon this foundation, we propose RA-Monitor, which models reflection as a learnable decision revision process and guides models to perform controlled self-correction when initial judgments are unreliable, thereby improving decision robustness. Extensive experiments conducted on M3-AD-Bench demonstrate that RA-Monitor outperforms multiple open-source and commercial MLLMs in zero-shot anomaly detection and anomaly analysis tasks.
M3-HNTM: Hyperspherical Multimodal Topic Modeling with Symbolic and Contextual Evidence
Zhiwen Luo ⋅ Dayu Guo ⋅ Manar Amayri ⋅ Nizar Bouguila ⋅ Zhixiang Li ⋅ Wenchuan Zhang ⋅ Wentao Fan
Multimodal evidence can improve topic discovery, but dense auxiliary signals can make topics harder to interpret when they replace readable descriptors. This creates an evidence-role problem: symbolic signals should define what topics mean to a reader, while dense aligned signals should help infer which topics a document expresses. We introduce M3-HNTM, a hyperspherical multimodal neural topic model that separates these roles through a shared document-level latent direction. M3-HNTM utilizes a shared von Mises-Fisher (vMF) posterior for document-level topic inference and vMF-mixture symbolic decoders that represent topics as concentration-aware semantic regions. Text Bag-of-Words and speech Bag-of-Acoustic-Words provide reconstructable symbolic evidence, while aligned image-text embeddings enter posterior inference as contextual evidence and help preserve semantic structure. The model applies to image-text, speech-text, and image-speech-text corpora while keeping topics readable through symbolic descriptors. Experiments on five datasets show that M3-HNTM improves the coherence-diversity quality score over strong text-only, hyperspherical, and multimodal baselines. On SpokenCOCO-Tri, the full tri-modal model outperforms both image-text and speech-text variants, indicating that visual and acoustic signals provide complementary evidence for interpretable topic discovery.
Machine Unlearning in Diffusion LLMs
Yili Wang ⋅ Yijie Xu ⋅ Lu Dai ⋅ Tairan Huang ⋅ Qianyi Cai ⋅ Huizai Yao ⋅ Tianfu Wang ⋅ Hui Xiong
Machine unlearning (MU) aims to remove sensitive or undesired knowledge from a trained model without retraining from scratch. While MU has been widely studied for autoregressive large language models, its application to diffusion large language models (DLLMs) remains largely unexplored. In this paper, we present the first systematic study of MU in DLLMs and show that existing unlearning objectives transfer poorly to this setting due to sparse supervision and inaccurate forgetting localization. To address these challenges, we propose \textsc{DL-Eraser}, a DLLM-native unlearning framework that suppresses target recoverability under high-mask conditioning while constraining updates to a utility-preserving subspace. Extensive experiments show that \textsc{DL-Eraser} achieves stronger forgetting while better preserving model utility than existing baselines. To the best of our knowledge, this is the first work that systematically investigates MU in DLLMs.
Making Open-Source Text LLM Watermarks Durable Against Merging
Luisa Scharff ⋅ Thibaud Gloaguen ⋅ Robin Staab ⋅ Martin Vechev
Open-source LLMs (OSMs) are reaching near state-of-the-art performance, prompting prior works to trace the text they generate by embedding text watermarking algorithms directly into their weights. Yet, OSMs are subject to post-training modifications, which has been shown to remove the watermark. Model merging in particular, a prominent method used for combining expert knowledge and preventing catastrophic forgetting, strongly removes such OSMs watermarks. A key question is how to enable OSM watermarks that survive subsequent merging. In this work, we show for the first time how to design an OSM watermark that is durable against model merging. We propose Merge-Adversarial Training, an adversarial training algorithm to distill text watermarks into model weights while being robust to subsequent model merging. Our approach consistently outperforms all baselines (e.g. with SLERP up to +51 percentage points (pp) TPR@1%FPR with +25 pp on average) while preserving downstream capabilities. We also for the first time evaluate OSM watermarks against realistic merge scenarios, representing common use-cases such as combining expert capabilities or preventing catastrophic forgetting, and with 3 prominent merging algorithms. More broadly, our findings suggest that adversarial training is a reliable approach for increasing OSM watermark durability against post-training modifications.
M*: A Modular, Extensible, Serving System for Multimodal Models
Atindra Jha ⋅ Naomi Sagan ⋅ Keisuke Kamahori ⋅ Irmak Sivgin ⋅ Rohan Sanda ⋅ Steven Gao ⋅ Mark Horowitz ⋅ Luke Zettlemoyer ⋅ Olivia Hsu ⋅ Jure Leskovec ⋅ Stephanie Wang ⋅ Baris Kasikci
We are entering a new era of *composite model architectures* that integrate diverse components such as vision encoders, language backbones, diffusion and flow heads, audio codecs, action generators, and world-model predictors. Such architectures underpin a broad class of multimodal models, including unified multimodal models, omni models, speech-language models, vision-language-action policies, and world models. However, existing model serving frameworks were built on narrow assumptions about model structure, making them ill-suited to accommodate this new architectural diversity. Here we present M*, a universal serving system for efficient serving of composite AI models. M* represents models as dataflow graphs, processing requests spanning diverse modalities and tasks as traversals over these graphs. The core insight is a modular abstraction that supports arbitrary composition of model components, flexible placement onto a physical cluster, and model-agnostic optimizations within a distributed runtime. We call this abstraction the *Walk Graph* and show how it can concisely capture composite models from a broad range of families. We instantiate M* on representative models and find that it achieves, on average, 30\% lower end-to-end latency than vLLM-Omni for text-to-image workloads on BAGEL, while delivering a lower real-time factor and higher throughput - by up to 15\% - for text-to-speech workloads on Qwen3-Omni. M* also outperforms the V-JEPA 2-AC rollout baseline for robotic planning by up to $12.5\times$. Thus, our work paves the road towards more efficient serving of complex models with minimal developer effort.
Manifold-Guided Stereo-Monocular Refinement for Endoscopic Stereo Disparity Estimation
Jiewen Yang ⋅ Yi Zhong ⋅ Bin Pu ⋅ Xingbo Dong ⋅ Ming Zheng ⋅ Qika Lin ⋅ Kai He ⋅ Xuechen Zhang
Dense stereo disparity estimation provides the geometric basis for 3D perception in computer-assisted endoscopy. Endoscopic stereo matching is difficult because reliable tissue correspondences often appear next to weak-texture, specular, occluded, or deformed regions within the same frame. Existing stereo and stereo-monocular methods improve cost aggregation or introduce monocular priors, but they often treat monocular guidance as a global feature source or an independently predicted candidate. Without evaluating the local support of the current stereo match during refinement, the update can reinforce corrupted correspondences in ambiguous regions or allow monocular priors to override reliable metric matches. To tackle these problems, we introduce Manifold-Stereo, a reliability-gated latent-flow framework for endoscopic stereo disparity estimation. The method refines a monocular disparity candidate through a conditional latent residual flow on an image-aligned correction space induced by stereo cost-volume geometry, then calibrates the corrected candidate to the current stereo disparity scale. A geometric consistency gate computed from left-right feature alignment regulates the candidate contribution during recurrent refinement, preserving stereo estimates in well-matched regions while supplying monocular context where local matching is ambiguous. Experiments on in-domain SCARED and zero-shot SERV-CT and Hamlyn show strong average disparity accuracy and competitive thresholded outlier rates against recent stereo and endoscopic baselines.
ManiFusion: Unlocking High-Throughput Generation via Superposition in Manifold Space
Minkyu Kim ⋅ Baekseung Kim ⋅ Junhoo Lee ⋅ Jangho Kim ⋅ Nojun Kwak
Diffusion models achieve high-quality image generation but remain computationally expensive due to iterative denoising across many timesteps. Existing acceleration methods mainly reduce per-sample latency, but gains along this axis are becoming saturated. We instead explore an orthogonal throughput-oriented direction by processing multiple samples within a single denoising evaluation. To this end, we propose ManiFusion, a backbone-agnostic framework for multi-sample generation through fused-space computation. ManiFusion jointly processes multiple samples by fusing them before the denoising backbone and separating them afterward through unfusion. We provide a geometric interpretation under the manifold hypothesis, focusing on identifiability and consistency of fused denoising dynamics. To better align fusion allocation with denoising stages, ManiFusion uses Trajectory-Fusion Scheduling (TFS), which dynamically adjusts fusion factors throughout sampling. ManiFusion generalizes across diffusion families, including both U-Net and transformer backbones in pixel and latent spaces. We further show that ManiFusion composes with existing acceleration methods and reaches throughput regimes unattainable by prior approaches with minimal quality degradation.
Many Benign Errors are Better than A Few Severe Ones: Evaluating Hallucination Severity
Sanjana Ramprasad ⋅ Pranav Mani ⋅ Chenhao Tan ⋅ Byron Wallace ⋅ Elisa Ferracane
Factuality evaluation in summarization has largely focused on detecting hallucinations as binary phenomena---labeling content as either supported or unsupported by the source. However, modern LLMs increasingly generate nuanced, partially grounded content whose factual impact can vary widely—from benign inferences to severe distortions. To enable more precise assessment, we propose a framework for annotating hallucination \emph{severity} along two interpretable dimensions: \emph{implausibility} (how unreasonable a span is given the source or general knowledge) and \emph{impact} (its potential to mislead or cause harm if assumed true). Applying this framework to obtain annotations on model generations across domains, we find that many hallucinations are both plausible and benign, indicating that not all unsupported content is equally consequential. We further show that severity-aware evaluation can meaningfully alter model rankings, favoring models that generate more frequent yet benign hallucinations over those producing fewer but more detrimental errors—distinctions that conventional evaluation settings fail to capture. Finally, we evaluate LLMs as severity judges: while their span detection remains weak, their severity ratings align closely with human judgments, making them promising candidates for scaling fine-grained annotation. Using severity as a diagnostic lens, we find that LLMs preferentially detect obvious errors but miss plausible, high-impact ones---precisely the cases that matter most for deployment. Together, these findings argue for a shift toward severity-aware factuality evaluation.
MAPLE: Latent Multi-Agent Play for End-to-End Autonomous Driving
Rajeev Yasarla ⋅ Deepti Hegde ⋅ Hsin-Pai Cheng ⋅ Shizhong Han ⋅ Yunxiao Shi ⋅ MeysamSadeghi ⋅ Hanno Ackermann ⋅ Litian Liu ⋅ Pranav Desai ⋅ Fatih Porikli ⋅ Mohammad Ghavamzadeh ⋅ Herbert Cai
Vision-language-action (VLA) models are effective as end-to-end motion planners, but can be brittle when evaluated in closed-loop settings due to being trained under traditional imitation learning framework. Existing closed-loop supervision approaches lack scalability and fail to completely model a reactive environment. We propose MAPLE, a novel framework for reactive, multi-agent rollout of a dynamic driving scenario in the latent space of the VLA model. The ego vehicle and nearby traffic agents are independently controlled over multi-step horizons, while being reactive to other agents in the scene, enabling closed-loop training. MAPLE consists of two training stages: (1) supervised fine‑tuning on the latent rollouts based on ground-truth trajectories, followed by (2) reinforcement learning with global and agent‑specific rewards that encourage safety, progress, and interaction realism. We further propose diversity rewards that encourage the model to generate planning behaviors that may not be present in logged driving data. Notably, our closed-loop training framework is scalable and does not require external simulators, which can be computationally expensive to run and have limited visual fidelity to the real-world. MAPLE achieves state-of-the-art driving performance on Bench2Drive and demonstrates scalable, closed-loop multi-agent play for robust E2E autonomous driving systems.
MaPP: A Unified Marginalized Posterior-Predictive Framework for Data-Efficient RLVR
YangYang Ren ⋅ Haodong Zhu ⋅ Sheng Xu ⋅ Yanjing Li ⋅ Nikolai Y Zolotykh ⋅ Wentao Zhang ⋅ Baochang Zhang
Reinforcement learning with verifiable rewards (RLVR) enhances the reasoning capabilities of large language models (LLMs), but at the expense of significant computational overhead due to compute-intensive rollout processes and frequent policy updates. Online prompt selection, a widely adopted strategy for improving training efficiency, maintains per-prompt Bayesian posteriors to predict prompt difficulty and prioritize informative prompts before committing rollout budget. However, these methods assess prompt informativeness without accounting for how reliably learning signal is extracted from sampled responses. In GRPO-based RL algorithms, the realized advantage of a response depends not only on its own outcome, but also on the randomly sampled outcomes of its peers through group normalization. Our experimental and theoretical analysis show that the resulting uncertainty in group composition induces composition noise, a non-vanishing variance component that creates an irreducible lower bound on gradient estimation error. Consequently, the standard group-relative advantage fails to faithfully characterize response-level utility, degrading gradient estimation and thereby impairing downstream prompt selection. To address this issue, we propose a unified Marginalized Posterior-Predictive framework, MaPP, for data-efficient RLVR, which first denoises response-level advantage estimation and then improves prompt selection using a shared Beta posterior. Specifically, for each response, MaPP replaces the standard group-relative advantage with a composition-invariant intrinsic advantage via closed-form Beta-Binomial marginalization. This yields a closed-form posterior-predictive advantage estimator whose error provably diminishes as the posterior concentrates. Then, building on the same posterior, MaPP derives an uncertainty-aware prompt selection score that more faithfully characterizes prompt informativeness, improving data efficiency without additional rollout cost. Experiments across mathematics, planning, and visual geometry on five model backbones show that MaPP consistently outperforms GRPO and strong selection baselines, achieving up to +2.45 average accuracy improvement over the strongest baseline under the same rollout budget, a new state-of-the-art.
Market Regime Council for Dynamic Credit Assignment in Multi-Agent LLM Decision Systems
Yunhua Pei ⋅ Zerui Ge ⋅ Jin Zheng ⋅ John Cartlidge
Multi-agent LLM decision systems for portfolio management still lack a principled way to assign credit across specialist agents, remain vulnerable to cold-start dominance under regime shifts, and offer limited transparency into how final allocations are formed. We propose Market Regime Council (MRC), a cooperative multi-agent decision system that computes exact Shapley credits across all single, pairwise, and Grand-coalition outputs for online agent weighting. Instantiated with $N{=}3$ specialist agents, at each trading period, MRC recomputes coalition-based Shapley weights from exponentially weighted performance histories, uses a Bayesian adaptive mixture to stabilize early periods, applies regime-dependent multipliers to adjust agent authority, and records each rebalance through a five-layer causal trace. Over 1,037 trading days across 13 crypto assets and five seeds, MRC achieves a Sharpe ratio of 1.51 and a cumulative return of 440.1\%, ranking first on CR, SR, and IR among active baselines and attaining the lowest MDD among active methods. Ablation results show that the gains come from Shapley-weighted integration across coalition outputs rather than from any single stage in isolation. Code and demo data are included in the supplementary material.
MARS: Multi-resolution Adaptive Routing for Sequential Recommendation
Ming Yin ⋅ Sixun Dong ⋅ Yudong Liu ⋅ Wenyun Yang ⋅ Yunjiang Jiang ⋅ Yiran Chen
Scaling sequential recommendation to long user histories requires compressing diverse behavioral evidence into memories that can be scored efficiently against many candidates. We first provide evidence that real user histories exhibit multi-scale semantic structure: short-lived intent, medium-term interests, and long-term preferences can coexist in the same sequence. However, existing summarization-based models mainly optimize compact history compression and do not explicitly organize cached user memory across temporal scales, which can lead to \textit{temporal aliasing} where distinct behavioral signals become entangled before prediction. We propose MARS, a multi-resolution user memory that writes the full history into recurrent state tracks initialized with different half-life priors. A sparse routing reader then materializes compact seed memories by selecting the relevant temporal resolutions for each seed, preserving fixed-size candidate scoring. Experiments across recommendation datasets show that MARS outperforms strong recommendation baselines, with gains most pronounced for users with long histories.
MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation
Patryk Bartkowiak ⋅ Lennart Petersen ⋅ Bartosz Kotrys ⋅ Dominik L Michels ⋅ Soeren Pirk ⋅ Wojtek Palubicki
Evaluating single-concept personalization in text-to-image diffusion has seen two categories of quantitative metrics: Concept Preservation (CP) measures identity fidelity to a reference while Prompt Following (PF) measures whether the generated scene matches the prompt. Personalization papers have commonly computed these signals using three separate backbones: CLIP-I and DINO for CP, CLIP-T for PF. In this paper, we show that existing metrics fall short of correlating with human perception because they attend to the image as a whole, instead of distinguishing the concept subject from the background. This distinction is important from a human perception point of view as the concept subject in the output image should be very similar to the input concept (CP), whereas the output background should adhere closely to the text prompt (PF). To improve personalization evaluation in this way, we introduce MaSC, a unified metric that attends to CP and PF by differentiating concept subject regions. Specifically, given an externally provided foreground concept mask, MaSC computes the CP and PF scores from a single forward pass of a frozen SigLIP2 encoder per image. On DreamBench++ human ratings, MaSC reaches Krippendorff $\alpha = +0.471$ on CP - beating every non-LLM baseline tested (DINOv3, DreamSim, AM-RADIO, DIFT-SDXL, DINO-I, CLIP-I) and GPT-4V, while sitting within $\Delta\alpha = 0.028$ of GPT-4o. To distinguish human perception bias we also evaluate our method, as well as common SOTA baselines, on ORIDa, a real-photo benchmark of identity preservation across physical environments. In this experiment MaSC reaches AUC $= 0.992$ almost perfectly identifying concept subjects. The PF score, obtained by MaSC without a second encoder forward pass, beats the CLIP-T baseline shipped with DreamBench++. In summary, our comprehensive evaluations demonstrate that MaSC establishes a new state-of-the-art for non-LLM concept preservation, while providing an efficient, unified standard for personalization evaluation. We release MaSC as a pip-installable Python package, alongside independent reproductions of every comparator's published numbers.
Masked Diffusion Language Agents for Tool-Integrated Chemical Reasoning
Mengdi Liu ⋅ Chenghao Jia ⋅ Hong Chang ⋅ Shiguang Shan
Designing tool-integrated reasoning agent systems for complex chemical tasks remains a fundamental challenge. Although recent chemical agents have demonstrated promising performance through the integration of external tools, they are still built on the autoregressive paradigm, which is inherently constrained by causal attention and left-to-right generation. As a result, they often struggle with bidirectional molecular understanding, global coordinated tool planning and multi-source evidence integration. To address this challenge, we introduce ChemDiffAgent, the first masked diffusion language agent for tool-integrated chemical reasoning. It reformulates chemical reasoning as iterative denoising over interleaved reasoning and tool-use trajectories, enabling tool-use decisions to be jointly determined under full-trajectory context and thereby better aligning the agent's decision boundary with its knowledge boundary. We further analyze how this paradigm benefits agentic reasoning through harder infilling subproblems during training and flexible arbitrary-order decoding during inference. Then we develop a two-stage post-training framework, comprising agentic supervised fine-tuning and variance-reduced preference optimization, to enhance tool-use and reasoning capabilities. We also construct a comprehensive benchmark covering six chemistry tasks across both single-turn and multi-turn tool-use scenarios. Under matched budgets and evaluation, ChemDiffAgent consistently outperforms autoregressive agents, achieving an overall improvement of approximately 20\%. Notably, our 8B model achieves performance comparable to, in some cases surpassing, significantly larger frontier models, such as GPT 5.5 and Claude Opus 4.7.
Masked Generative Pretraining Improves Cross-Dataset Transfer in Pixel-Space Diffusion
Shu Wei ⋅ Jiachen Lei ⋅ Jiahong Wu ⋅ Xiangxiang Chu
Pixel-space diffusion models have recently seen a revival in image generation, yet their training remains less efficient than that of their latent-space counterparts. Existing methods reduce this cost by equipping diffusion models with external representations from clean images. In contrast, representation consistency training, a new generative pretraining paradigm, stands out as it eliminates the need of any off-the-shelf semantic encoders. It proposes pre-training the diffusion model to simultaneously capture meaningful visual semantics from clean images while aligning them with data points across various noise levels. Nevertheless, its training depends on contrastive learning with hand-crafted augmentations. This introduces strong semantic prior into the pretraining, limiting model performance when transferred to downstream target distributions that substantially deviate from the pretraining manifold. To address this, we propose Masked Pixel-space Generative Pretraining (MPG), an augmentation-free framework based on masked image modeling. MPG trains the encoder to recover masked image area and aligns those predictions with samples of different noise levels. To measure the performance of pretrained model, we further introduce NC-CKNNA, a metric that quantifies the semantic structure consistency across different noise levels. Under the same architectures and fine-tuning settings, MPG consistently produces better generation performance across seven downstream datasets, while retaining competitive generation quality on ImageNet-256. These results suggest masked generative pretraining as a practical alternative for training pixel diffusion models.
MATO: Multi-objective Personalized Alignment with Test-time Optimization for Large Language Models
LINHAO LUO ⋅ Trang Vu ⋅ Van-Anh Nguyen ⋅ Junae Kim ⋅ Reza Haffari ⋅ Dinh Phung
Aligning large language models (LLMs) with diverse and multifaceted user preferences is a fundamental challenge in personalized AI systems. Existing multi-objective alignment methods either rely on costly training or require pre-trained reward models for each preference, making it difficult for them to adapt to evolving preferences. Prompt-based personalization offers a training-free alternative, but prompting alone often provides limited steerability, as LLMs may overemphasize or overlook certain preferences and fail to give users reliable control over the relative importance of different objectives when conflicts arise, leading to suboptimal alignment. In this paper, we introduce MATO, a training-free framework for Multi-objective personalized Alignment with Test-time Optimization. MATO formulates personalization as a test-time optimization problem that steers the relative importance of multiple objectives through controllable weights during decoding, without modifying model parameters or requiring external reward models. Specifically, a reward discovery module recovers preference rewards directly from the backbone LLM for diverse objectives specified in natural language, while a weight optimization module dynamically adjusts objective weights based on the user's initial preferences and the partially generated response to balance competing objectives during generation. The resulting rewards and weights jointly guide an online optimization procedure over the token distribution, enabling better alignment with the target objectives. Extensive experiments across multiple datasets and backbone LLMs show that MATO consistently outperforms strong baselines, achieving Pareto-improving multi-objective alignment and stronger steerability. These results highlight test-time optimization as a promising direction for scalable, controllable, and model-agnostic personalized alignment.
MCAS: Signal-Processing-Based Multi-View Contrastive Learning for Acoustic Sensing
Bingzhi Wang ⋅ Jiajun Yu ⋅ Dengke Zhong ⋅ Jiaming Liu ⋅ Xiong Li ⋅ Hongbo Liu ⋅ Jie Yang
Acoustic sensing is a promising technique for human-computer interaction, but existing learning-based systems still depend heavily on large task-specific labeled datasets. Compared with visual data, acoustic signals are difficult to interpret and annotate, making supervision costly and hard to scale. Meanwhile, existing unsupervised representation learning methods, particularly contrastive learning, are largely designed for natural images and rely on data augmentation to form positive pairs. Such strategies can destroy the temporal-physical structure of acoustic signals, causing shortcut learning rather than meaningful representation learning. To address this issue, we propose MCAS, a signal-processing-based multi-view contrastive learning framework for acoustic sensing. Instead of relying on artificial augmentations, MCAS constructs structure-preserving positive pairs from complementary representations derived from different signal processing pipelines of the same acoustic event. Moreover, MCAS is designed to address the heterogeneity of acoustic representations through progressive cross-view interaction enabled by Hierarchical Shared-Token Fusion and a multi-level pretraining objective with physical consistency regularization. Extensive experiments on two typical acoustic sensing tasks demonstrate the effectiveness of the proposed framework. Under low-label settings, MCAS consistently outperforms both supervised baselines and generic self-supervised baselines, while also showing strong cross-task transferability and robustness to corruptions and noise.
MCPHunt: An Evaluation Framework for Cross-Boundary Data Propagation in Multi-Server MCP Agents
Haonan Li ⋅ Tianjun Sun ⋅ Yongqing Wang ⋅ Qisheng Zhang
Multi-server MCP agents create an information-flow control problem: faithful tool composition can turn individually benign read/write permissions into cross-boundary credential propagation---a structural side effect of workflow topology, not necessarily malicious model behavior. We present MCPHunt, to our knowledge the first controlled benchmark that isolates non-adversarial, verbatim credential propagation across multi-server MCP trust boundaries, with three methodological contributions: (1) canary-based taint tracking that reduces propagation detection to objective string matching; (2) an environment-controlled coverage design with risky, benign, and hard-negative conditions that validates pipeline soundness and controls for credential-format confounds; (3) completion-requires-secret (CRS) stratification that disentangles task-mandated propagation (faithful execution of verbatim-transfer instructions) from policy-violating propagation (credentials included despite the option to redact). Across 3,615 main-benchmark traces from 5 models spanning 147 tasks and 9 mechanism families, policy-violating propagation rates reach 11.5--41.3% across all models. This propagation is pathway-specific (25x cross-mechanism range) and concentrated in browser-mediated data flows; hard-negative controls provide evidence that production-format credentials are not necessary---prompt-directed cross-boundary data flow is sufficient. A prompt-mitigation study across 3 models reduces policy-violating propagation by up to 97% while preserving 80.5% utility, but effectiveness varies with instruction-following capability---suggesting that prompt-level defenses alone may not suffice. Code, traces, and labeling pipeline are released under MIT and CC BY 4.0.
Measuring Coherence in Predictive Models
Jamie Reason ⋅ Vik Shirvaikar ⋅ Mark van der Wilk ⋅ Chris C Holmes
Many decision making procedures that make use of probabilistic predictors assume that training on new data acts as Bayesian conditioning. When this assumption breaks, the downstream procedure can perform poorly. E.g. we show that in active learning, incoherent updates cause the procedure to prefer to acquire suboptimal points. To quantify how far a model's predictive updates depart from conditioning, we introduce the incoherence ratio. We empirically find that amortised predictors such as TabPFN can be more coherent than standard parametric approximations to Bayesian inference. In synthetic active learning experiments, more coherent models acquire better data, an effect not attributable to better predictive performance. In many tasks, decisions depend on a predictor only through its induced action. We therefore introduce action coherence, a decision-theoretic relaxation that measures only the incoherence affecting that action. These diagnostics make coherence testable and quantifiable, enabling incoherence to be diagnosed and addressed.
Mechanistic Interpretability with Sparse Autoencoder Neural Operators
Bahareh Tolooshams ⋅ Ailsa Shen ⋅ Animashree Anandkumar
We introduce sparse autoencoder neural operators (SAE-NOs), a new class of sparse autoencoders that operate in function spaces rather than fixed-dimensional Euclidean representations. We formalize the functional representation hypothesis, where data are explained through sparse compositions of structured functions. Unlike standard SAEs that represent concepts with scalar activations, SAE-NOs parameterize concepts as functions, enabling representations that capture not only a concept's presence, but also how and where it is expressed across the input domain. We achieve this through joint sparsity: concept sparsity selects active concepts, while domain sparsity governs where they are expressed. We instantiate this framework using Fourier neural operators (SAE-FNOs), parameterizing concepts as integral operators in the Fourier domain. This functional and spectral parameterization is particularly advantageous when data exhibit spatial structure across scales or when concepts are frequency-structured. We characterize SAE-FNO on vision data and demonstrate that it learns localized patterns, uses concepts more efficiently, and exhibits stable concept characteristics across sparsity levels. We further show that SAE-FNO adapts to changes in domain size and generalizes across discretizations, operating at resolutions beyond those seen during training, where standard SAEs fail. We also introduce lifting into SAEs and show theoretically and empirically that it acts as a preconditioner that accelerates optimization. Overall, our results show that moving from vector-valued to functional parameterizations, with concept and domain sparsity, extends SAEs from representing concept presence to modeling structured concept expression, highlighting the importance of parameterization.
Medical LLMs as Medical World Models: Unified Policy-Dynamics Learning with Test-Time Search
Yucheng Zhou ⋅ Peng Luo ⋅ Jianbing Shen
Clinical decision-making is intrinsically sequential: a clinician must propose a diagnosis and treatment plan from the pre-treatment state and anticipate the post-treatment outcome in order to revise the plan. Existing medical world models address this loop only partially, through image-to-image dynamics with separate diagnostic and scoring modules, text-only trajectory models without multimodal input, or reflection agents without a world model. We propose \textbf{MedWM}, a unified multimodal LLM in which policy and dynamics are two modes of a single autoregressive backbone with fully shared parameters; the dynamics mode emits a text-only \emph{structured} post-treatment state, free-text clinical description plus standardized key-value outcomes (pCR, RECIST, labs, survival), that admits objective evaluation without perceptual metrics. We post-train this backbone with joint policy-dynamics supervision and dynamics-grounded reinforcement learning, then wrap it in an inference-time agent harness that combines breadth-depth ($N \times K$) search using the model's own dynamics as an internal verifier with a self-evolving case memory. On MIMIC-IV, BreastDCEDL-ISPY2 and HCC-TACE-SEG, MedWM consistently improves over same-backbone controls, matches or surpasses a specialized three-module image-based world model, and outperforms reflection-only, multi-agent, and larger closed-source LLM baselines; a blinded clinician study corroborates the quantitative gains.
MedIGen: Reliable Medical Illustration Generation via Interleaved Introspective Reasoning
Rongsheng Wang ⋅ Hongru Zhou ⋅ Ruizhe Zhou ⋅ HAOMING CHEN ⋅ Zhenyang Cai ⋅ Junying Chen ⋅ Minghao Wu ⋅ Yaofei Duan ⋅ Ziyi Zeng ⋅ Benyou Wang
Medical illustration generation is not merely domain-specific text-to-image (T2I) generation, but a high-constraint visual reasoning problem where visually plausible images can be invalid due to anatomical and structural errors. Existing T2I models largely rely on one-pass generation and lack an internal mechanism to inspect and correct such biomedical inconsistencies. We introduce MedIGen, a unified medical illustration generation model that transforms one-pass synthesis into an interleaved introspection-aware refinement process, coupling generation with self-reflection and re-generation through “generate–reflect–refine” cycles. MedIGen is trained with three progressive stages: (1) large-scale medical illustration pretraining on 1.21M samples, (2) mixture training over generation, reflection, and refinement sub-skills, and (3) reinforcement learning with dual-level rewards that optimize both reasoning validity and final visual correctness. We further present IlluGenBench, an expert-aligned, rubric-driven benchmark with 296 diverse tasks and 9,015 criteria, evaluating scientific accuracy, structural correctness, and semantic alignment beyond coarse visual plausibility. Experiments show that MedIGen substantially outperforms strong open-source unified and reasoning-based baselines, establishing a new open-source frontier. All resources will be open-sourced to facilitate future research.
MemeEconomy : Do LLM Agents Trade Ethics for Survival?
Syed Nazmus Sakib ⋅ Nafiul Haque ⋅ Ahnaf Manan ⋅ M. M MORSHED ⋅ Shifat E. Arman
Alignment-trained large language model (LLM) agents are increasingly deployed in autonomous, multi-round settings where ethical compliance can come into direct tension with task success. While prior safety literature extensively documents failure modes such as sycophancy and reward hacking, it largely overlooks a critical vulnerability: how agents behave under simulated existential economic pressure. We introduce MemeEconomy, a multimodal agentic market simulation in which LLM agents operate as meme investors across 300 real-world events. Agents select content from a five-tier harm taxonomy and target specific online communities under both low-pressure and high-pressure ("survival") environments. Each decision is recorded using a Belief--Desire--Intention (BDI) schema, which we employ as a diagnostic instrument to externalize the agent's event understanding, declared priorities, and risk acknowledgment. Across ten models spanning four frontier model families, harmful selection rates remain consistently high and increase by an average of +15 percentage points under economically induced survival pressure. Frontier alignment-trained models frequently commit to harm-tolerant selections at initialization while simultaneously suppressing the explicit harm acknowledgment that previously accompanied such decisions. In competitive tournament settings, ethically aligned agents achieve higher overall rankings, yet moderation penalties fail to meaningfully suppress harmful behavior. Harmful selection rates remain elevated even in rounds immediately following moderation removals, indicating that the moderation mechanism does not function as an effective deterrent. We further introduce MemeAgent, a 2B-parameter verifier trained on the simulation's BDI logs. MemeAgent substantially outperforms zero-shot frontier verifiers on in-distribution auditing tasks, achieving 87.8\% accuracy compared to 63--74\% for frontier baselines, while also generalizing to external multimodal harm benchmarks (78\% vs. 50--56\%). Our findings demonstrate that the same structured reasoning traces that expose the failure mode can also be leveraged to train systems capable of detecting it.
MEME: Lightweight Hierarchical Mixture-of-Experts for Unified Affective Computing
Yinan Zhang ⋅ Haoyu Zhang ⋅ Tianshu Yu
Multimodal sentiment analysis (MSA) and emotion recognition (ER) are closely related affective understanding tasks, yet they are usually studied separately due to differences in label space, task formulation, dataset distribution, and modality dependence. In this paper, we propose MEME ($\textbf{M}$ixture-of-$\textbf{E}$xperts for $\textbf{M}$ultimodal sentiment analysis and $\textbf{E}$motion recognition), a lightweight hierarchical MoE framework for unified multimodal affective computing across MSA, conversational emotion recognition, and dynamic facial expression recognition. MEME operates on frozen text, visual, and audio features, compresses variable-length modality sequences with learned-query attention pooling, and processes them with a shared hierarchical MoE backbone. Each block first applies hard modality experts for modality-specific refinement and then task-conditioned cross experts for task-aware multimodal interaction. A $\texttt{[TASK]}$ token provides an explicit task anchor for routing, while modality dropping regularizes unified training. We jointly train MEME on nine affective benchmarks and evaluate all datasets using a single composite-best checkpoint without dataset-wise adaptation. Extensive experiments show that MEME outperforms strong task-specific baselines and recent large-model-based unified methods on most benchmarks, while maintaining favorable efficiency, robustness and generalization.
MEME: Multi-Entity & Evolving Memory Evaluation
Seokwon Jung ⋅ Alexander Rubinstein ⋅ Arnas Uselis ⋅ Sangdoo Yun ⋅ Seong Joon Oh
LLM-based agents increasingly operate in persistent environments where they must store, update, and reason over information across many sessions. While prior benchmarks evaluate only single-entity updates, MEME defines six tasks spanning the full space defined by the multi-entity and evolving axes, including three not scored by prior work: Cascade and Absence (dependency reasoning) and Deletion (post-removal state). Evaluating six memory systems spanning three memory paradigms on 100 controlled episodes, we find that all systems collapse on dependency reasoning under the default configuration (Cascade: 3\%, Absence: 1\% in average accuracy) despite adequate static retrieval performance. Prompt optimization, deeper retrieval, reduced filler noise, and most stronger LLMs fail to close this gap. Only a file-based agent paired with Claude Opus 4.7 partially closes the gap, but at $\sim$70$\times$ the baseline cost, indicating closure currently depends on configurations that are not practical at scale. Code is available at https://anonymous.4open.science/r/MEME-0612 and the dataset at https://huggingface.co/datasets/meme-benchmark/MEME.
Memento No More: Coaching AI Agents to Master Multiple Tasks via Hints Internalization
Minttu Alakuijala ⋅ Ya Gao ⋅ Georgy Ananov ⋅ Samuel Kaski ⋅ Pekka Marttinen ⋅ Alexander Ilin ⋅ Harri Valpola
As the general capabilities of artificial intelligence (AI) agents continue to evolve, their ability to learn to master multiple complex tasks through experience remains a key challenge. Current LLM agents, particularly those based on proprietary language models, typically rely on prompts to incorporate knowledge about the target tasks. This approach does not allow the agent to internalize this information and instead relies on ever-expanding prompts to sustain its functionality in diverse scenarios. This resembles a system of notes used by a person affected by anterograde amnesia, the inability to form new memories. In this paper, we propose a novel method to train AI agents to incorporate knowledge and skills for multiple tasks without the need for either cumbersome note systems or prior high-quality demonstration data. Our approach employs an iterative process where the agent collects new experiences, receives corrective feedback from humans in the form of hints, and integrates this feedback into its weights via a context distillation training procedure. We demonstrate the efficacy of our approach by implementing it in a Llama-3-based agent that, after only a few rounds of feedback, outperforms advanced models GPT-4o and DeepSeek-V3 in tasksets requiring correct sequencing of information retrieval, tool use, and question answering.
MemForest: Efficient Agent Memory Management via EventTree Partitioning and Progressive Merging
JunXi Wang ⋅ Jiayi Zhu ⋅ Te Sun ⋅ Chen Zhang ⋅ Siyuan Li ⋅ Xuyang Liu ⋅ Zichen Wen ⋅ Xiaobing Tu ⋅ Jinkui Ren ⋅ Xiantao Zhang ⋅ Ziqi Yuan ⋅ Linfeng Zhang
Agent memory systems have demonstrated significant potential in tasks such as long-term dialogue, personalized assistants, and video understanding. However, as inference progresses, continuously accumulated memory imposes substantial storage and retrieval burdens. To address this issue, we propose MemForest, a general memory compression framework adaptable to various agent memory systems. Specifically, MemForest leverages both global semantic similarity and local temporal continuity of memory events to partition the historical memory into a set of event-centric independent units. For each independent unit, the framework constructs a maximum spanning tree structure, referred to as an EventTree, and performs progressive merging by iteratively selecting high-weight edges, thereby effectively compressing redundant memory nodes and reducing storage overhead. In addition, we introduce an anchor-guided propagation retrieval mechanism, which retrieves more relevant memory nodes from the temporal neighborhoods of key memory nodes, thereby enabling more accurate memory retrieval. Extensive experiments demonstrate the effectiveness of MemForest. Under the unimodal Mem0 framework, across three benchmarks (LoCoMo, LongMemEval, and PersonaMem), MemForest preserves 97.1% of the original performance while compressing 50% of historical memory, achieving a 1.89× retrieval speedup. Under the multimodal M3-Agent framework, across two benchmarks (M3-Bench-robot and M3-Bench-web), MemForest retains 99.7% of the original performance under a 50% compression ratio, while achieving a 2.24× retrieval speedup. Our code is available in the supplementary materials, and all data will be released on GitHub.
Memory-Efficient Federated Fine-Tuning of LLMs via Block-wise Progressive Training
Qianyue Cao ⋅ Zongwei Zhu ⋅ Boyu Li ⋅ Yi Xiong ⋅ Zirui Lian ⋅ Xuehai Zhou
Federated fine-tuning has become a dominant paradigm for privacy-preserving Large Language Model (LLM) adaptation. While integrating Parameter-Efficient Fine-Tuning (PEFT) reduces communication and computational costs, existing methods neglect that peak memory bottlenecks caused by full forward passes through the frozen LLM. This excludes low-memory devices, leading to data loss and suboptimal global performance. In this paper, we propose BP-FedPEFT, a framework utilizing progressive training to decompose the end-to-end computational graph, reducing peak memory to the block level. While enabling low-memory device participation, this paradigm incurs prolonged training latency and introduces growing memory burdens for deep-layer inputs, alongside suffering from cascading feature misalignment due to the absence of global supervision, leading to suboptimal model performance. To ensure efficiency, BP-FedPEFT employs functional-aware overlapping planning coupled with a local-global stability criterion to regulate training steps and communication rounds. To ensure effectiveness, we utilize depth-injected input synthesis and block overlaps to bridge the supervision gap, establishing valid optimization trajectories that align shallow representations with deep functional expectations. We establish theoretical convergence guarantees for BP-FedPEFT. Experiments on a heterogeneous testbed show that BP-FedPEFT supports diverse PEFT methods, reducing average memory usage by 44.7-76.2\%, accelerating training by 2.3-12.6$\times$, and improving accuracy by 3.2-5.8\% through inclusive participation.
MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents
Haozhen Zhang ⋅ Quanyu Long ⋅ Jianzhu Bao ⋅ Tao Feng ⋅ Weizhi Zhang ⋅ Haodong Yue ⋅ Wenya Wang
Most Large Language Model (LLM) agent memory systems rely on a small set of static, hand-designed operations for extracting memory. These fixed procedures hard-code human priors about what to store and how to revise memory, making them rigid under diverse interaction patterns and inefficient on long histories. To this end, we present \textbf{MemSkill}, which reframes these operations as learnable and evolvable memory skills, structured and reusable routines for extracting, consolidating, and pruning information from interaction traces. Inspired by the design philosophy of agent skills, MemSkill employs a \emph{controller} that learns to select a small set of relevant skills, paired with an LLM-based \emph{executor} that produces skill-guided memories. Beyond learning skill selection, MemSkill introduces a \emph{designer} that periodically reviews hard cases where selected skills yield incorrect or incomplete memories, and evolves the skill set by proposing refinements and new skills. Together, MemSkill forms a closed-loop procedure that improves both the skill-selection policy and the skill set itself. Experiments on LoCoMo, LongMemEval, HotpotQA, and ALFWorld demonstrate that MemSkill improves task performance over strong baselines and generalizes well across settings. Further analyses shed light on how skills evolve, offering insights toward more adaptive, self-evolving memory management for LLM agents.
Large foundation models have accelerated progress toward general-purpose agents that interact with humans and other agents through language and multimodal signals. However, robust multi-agent decision-making requires reasoning about what other agents know, intend, and are likely to do under partial observability. Current agentic systems often operate through prompt design, memory, or end-to-end behavioral shaping, but typically do not learn an explicit partner-state representation that can be reused as a decision variable across tasks. We introduce mental-model-enabled agents, a framework that equips an agent with a latent mental model of its counterpart, allowing it to infer hidden beliefs, intentions, and likely reactions from the observed history and use these inferences to guide action selection. Our method learns an amortized recursive Theory-of-Mind representation, with first- and second-order mental-state structure, jointly with a belief-conditioned reward model that evaluates candidate actions relative to the inferred partner state. A policy is then learned under this belief-aware signal, yielding an agent that can act independently at inference time while retaining the benefits of explicit partner modeling. We evaluate the same framework on both language-only and multimodal benchmarks. Across these settings, explicit mental-state modeling consistently improves interaction quality and Theory-of-Mind performance over base agentic systems, showing that structured partner modeling is a useful inductive bias for general multi-agent systems.
Mesh BDF: Barycentric Dominance Field for 3D Native Mesh Generation
Gaochao Song ⋅ Haohan Weng ⋅ Luo Zhang ⋅ Zibo Zhao ⋅ Shenghua Gao
Autoregressive (AR) modeling has recently achieved remarkable progress in native 3D mesh generation, largely due to its natural ability to handle variable-length, discrete data structures. However, the inherent constraints of the AR paradigm severely restrict the generated meshes, leading to limited face counts, bounded vertex resolutions, and difficulties in supporting textures. To overcome these bottlenecks, we propose the Barycentric Dominance Field (BDF), a continuous representation defined on triangular mesh surfaces that elegantly encodes vertex topological connectivity. BDF bridges the fundamental gap between discrete mesh topology and continuous diffusion-based generative modeling by transforming connectivity into a continuous surface signal. As an intrinsic mesh property, BDF shares strong similarities with texture maps, enabling its seamless integration into existing 3D diffusion pipelines without requiring architectural modifications. Extensive experiments demonstrate that BDF empowers diffusion models to generate native meshes with significantly higher quality, greater scalability, and stronger robustness compared to state-of-the-art autoregressive methods.
META-PAP: Meta-learning for Prompt-aware Preference Pairing in LLM Alignment
Pinlong Zhao ⋅ Shiyu Hu ⋅ Jing Zhang ⋅ Mengyang Li
Direct Preference Optimization (DPO) has become a popular route for aligning large language models with human preferences, but existing pipelines treat all preference pairs uniformly and ignore prompt-level heterogeneity. We show that the optimal pair-selection strategy varies systematically with prompt characteristics: prompts with concentrated reward distributions benefit from large reward-gap pairs that provide unambiguous signal, while prompts with diverse responses benefit from moderate-gap pairs because extreme positions are dominated by outliers. No fixed strategy is optimal across this heterogeneity. We propose META-PAP (Meta-learning for Prompt-aware Preference Pairing), a framework that turns this observation into a learned algorithm. META-PAP generates multiple diverse pairs per prompt via a structured grid over the response reward distribution, extracts a compact prompt-aware feature vector, and trains a lightweight meta-network through bilevel optimization to predict per-pair weights. The inner loop is a weighted DPO update; the outer loop evaluates a virtual policy on a small clean meta-set. Across Llama-3-8B, Mistral-7B-v0.3, and Llama-2-7B, META-PAP consistently outperforms vanilla DPO, fixed-position pairing, reward-gap weighting, top-$k$ filtering, curriculum, and statistical rejection sampling, with average gains of $+3.9$ points on AlpacaEval~2.0 and $+3.2$ points on Arena-Hard, and approaches an oracle that exhaustively searches per prompt. The learned weighting policy is interpretable: it favors large-gap pairs on low-diversity prompts and softly rejects extreme pairs on high-diversity prompts. We additionally find that the meta-network requires only $M\!=\!150$ pairs to saturate, transfers across domains, tolerates label noise on the meta-set, and remains effective with as few as $n\!=\!32$ on-policy samples per prompt.
MEV: A Multi-Event Video Dataset for Long-Take Generation
Peiyuan Zhu ⋅ Shaoan Xie ⋅ Yifan Shen ⋅ Wen Tian ⋅ Zijian Li ⋅ Guangyi Chen ⋅ Kun Zhang
Current datasets for text-to-video generation are largely limited to short clips depicting a single event or coarse-grained captions. As a result, they provide limited supervision for modeling long-take temporal dependencies, often leading to temporally inconsistent generations when describing multiple events. In this paper, we introduce MEV, a multi-event video dataset designed for long-take video generation. MEV provides high-quality real-world videos where subjects can perform a sequence of distinct, non-overlapping events, featuring precise temporal boundaries and fine-grained procedural annotations. This event-centric discretization reduces confounding factors and enables models to support temporally localized conditioning and build long-take videos from atomic event units. Furthermore, we introduce a comprehensive annotation pipeline that segments videos into temporally grounded events, and generates structured captions at both the global level and the event level. In addition, under an event-centric evaluation protocol, we conduct multi-event generation experiments by fine-tuning on MEV with temporally grounded event conditioning, leading to improved long-take event transitions and more coherent content. We believe this work removes a key barrier to training video foundation models for multi-event scenarios and provides practical insights into learning from temporally-structured video-text data.
MicroWorld: Empowering Multimodal Large Language Models to Bridge the Microscopic Domain Gap with Multimodal Attribute Graph
Manyu Li ⋅ Ruian He ⋅ Chenxi Ma ⋅ Weimin Tan ⋅ Bo Yan
Multimodal large language models (MLLMs) show remarkable potential for scientific reasoning, yet their performance in specialized domains such as microscopy remains limited by the scarcity of domain-specific training data and the difficulty of encoding fine-grained expert knowledge into model parameters. To bridge the gap, we introduce MicroWorld, a framework that constructs a multimodal attributed property graph (MAPG) from large-scale scientific image-caption corpora and leverages it to augment MLLM reasoning at inference time without any domain-specific fine-tuning. MicroWorld extracts biomedical entities and relations via scispaCy or LLM-based triplet mining, aligns images and entities in a shared embedding space using Qwen3-VL-Embedding, and assembles a knowledge graph comprising approximately 111K nodes and 346K typed edges spanning eight relation categories. At inference time, a graph-augmented retrieval pipeline matches query entities to the MAPG and injects structured knowledge context into the MLLM prompt. On the MicroVQA benchmark, MicroWorld improves the reasoning performance of Qwen3-VL-8B-Instruct by 37.5%, outperforming GPT-5 by 13.0% to achieve a new state-of-the-art. Furthermore, it yields a 6.0% performance gain on the MicroBench benchmark. Extensive experiments demonstrate the enhanced generalization capability introduced by MicroWorld. A qualitative case study further reveals both the mechanisms through which structured knowledge improves reasoning and the failure modes that point to promising future directions. Code and data are available at anonymous GitHub.
MindLoom: Composing Thought Modes for Frontier-Level Reasoning Data Synthesis
Haiyang Shen ⋅ Taian Guo ⋅ Xuanzhong Chen ⋅ Mugeng Liu ⋅ Sixiong Xie ⋅ Zhuofan Shi ⋅ Chongyang Pan ⋅ Siqi Zhong ⋅ Guoqing Wang ⋅ Ming Zhang ⋅ Yun Ma
Although LLMs have made substantial progress in reasoning, systematically producing frontier-level reasoning data remains difficult. Existing synthesis methods often have limited visibility into the structural factors that govern problem difficulty, which can result in narrow diversity and unstable difficulty control. In this work, we view the difficulty of a reasoning problem as arising from the accumulation of atomic knowledge-reasoning transformations, which we term thought modes. Building on this perspective, we propose MindLoom, a framework for synthesizing frontier-level reasoning data through compositional thought mode engineering. Given a collection of hard problems with verified solutions, MindLoom first decomposes those solutions into thought mode chains that reveal each problem's construction logic. It then trains a retrieval model that matches problem states to compatible thought modes, providing guidance on which reasoning challenges to introduce during synthesis. New problems are composed by iteratively applying retrieved thought modes to seed questions, with distribution-aligned sampling to encourage diverse reasoning coverage. Finally, a rollout-based judging stage labels generated questions by difficulty and supplies judged-correct responses for supervised fine-tuning. We evaluate MindLoom on nine benchmarks covering five STEM disciplines and four mathematical reasoning tasks across multiple model families and sizes. Models fine-tuned on MindLoom generated data consistently improve over base models, distillation, and external-data baselines across the reported benchmarks. Ablation studies indicate the contribution of each component, and further analysis suggests that MindLoom covers a broad range of reasoning patterns while maintaining useful difficulty control. We have open-sourced our implementation at https://anonymous.4open.science/r/MindLoom-5B8E.
MindShape: Superquadric-Constrained High-Fidelity 3D Reconstruction from fMRI
Xiaoquan Shen ⋅ Ming Li ⋅ Jianxiong Gao ⋅ Jiaxuan Chen ⋅ Xiangru Huang ⋅ Yanwei Fu ⋅ Gang Pan
Recent advances in neural decoding have enabled the reconstruction of 3D visual content from functional magnetic resonance imaging (fMRI) signals, opening a route toward brain-conditioned 3D generation. Nevertheless, existing approaches largely rely on semantic or multi-view priors without explicitly modeling 3D geometric constraints, often leading to structural distortions and degraded geometric fidelity. To address this limitation, we propose MindShape, a geometry-aware neural decoding framework for fMRI-conditioned 3D reconstruction. Unlike previous methods, MindShape decodes superquadric geometry as an explicit, low-dimensional structural constraint and jointly conditions the 3D generation process on both geometric primitives and semantic descriptions, enabling semantically consistent and geometrically faithful reconstructions. Experiments show that MindShape improves reconstruction quality in terms of geometric consistency and semantic fidelity over existing baselines. These results suggest that explicit geometric primitives provide an effective structured interface between noisy neural measurements and controllable 3D generation.
Mind the Parameters: Lightweight and Efficient Brain Visual Decoding with Shared Tensor Cores
Haodong Jing ⋅ Panqi Yang ⋅ Rongchao Zhang ⋅ Junhao Jia ⋅ Hoi Leong Lee ⋅ Zhipeng Liu ⋅ Yongqiang Ma ⋅ Nanning Zheng
Linking brain activity to computational representations of visual perception is a central goal at the intersection of neuroscience and machine learning. However, cross-subject decoding from fMRI remains challenging because neural responses vary substantially across individuals, while existing methods often rely on subject-specific fine-tuning or parameter-heavy decoders. This limits generalization to unseen subjects and hinders efficient deployment. We propose BrainTC--Brain Tensor Cores, a lightweight framework for cross-subject brain visual decoding based on shared tensor decomposition. BrainTC separates shared functional structure from subject-specific variation through a shared tensor-core alignment module that maps ROI-wise responses into a common functional space with shared tensor bases and residual subject factors. This supports both zero-shot decoding for unseen subjects and few-shot adaptation by updating only residual parameters. We then develop a Brain Tensor-Transformer with tensorized attention to compactly model inter-ROI dependencies, together with a hierarchical neural-to-visual mapping module for multi-level visual prediction. Experiments show that BrainTC achieves competitive cross-subject decoding performance with substantially improved parameter efficiency, particularly in zero-shot generalization to unseen subjects.
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM
Wenqian Cui ⋅ Xiao-Hui Li ⋅ Daxin Tan ⋅ Qiyong Zheng ⋅ Irwin King
Speech large language models (SLMs) are typically built from text large language model (TLM) checkpoints, yet they still suffer from a substantial modality gap. Prior work has mainly attempted to reduce this gap from the output side by making speech generation more text-like, but the gap remains. We argue that the key remaining bottleneck lies on the input side. We propose TextPro-SLM, an SLM that makes spoken input more closely resemble that of a prosody-aware text LLM. TextPro-SLM combines WhisperPro, a unified speech encoder that produces synchronized text tokens and prosody embeddings, with an LLM backbone trained to preserve the semantic capabilities of the original TLM while learning paralinguistic understanding. Experiments show that TextPro-SLM achieves the lowest modality gap among leading SLMs at both 3B and 7B scales, while also delivering strong overall performance on paralinguistic understanding tasks. These gains are achieved with only roughly 1,000 hours of audio training data, suggesting that reducing the modality gap from the input side is both effective and data-efficient.
Mitigating Overgeneralization in RND via Spectral Target Design
Minseok Jeong ⋅ Yechan Lee ⋅ Hyewon Choi ⋅ Jeongyong Yang ⋅ SooJean Han
Random Network Distillation (RND) is a scalable novelty signal for reinforcement learning, but it can overgeneralize: the predictor extrapolates the random target off the data manifold, causing novelty scores to collapse on unfamiliar inputs. This paper analyzes this failure through a kernel-theoretic lens. Under a GP target model and an NTK-regime predictor, we show that the expected RND energy is a cross-kernel residual governed jointly by the target covariance kernel and the predictor interpolation kernel. This identifies overgeneralization as spectral under-excitation of modes where predictor residuals would otherwise survive. We then explicitly formulate spectral target design as a principle for shaping RND residual geometry; one concrete instance is replacing to replace the implicit random-network target with a bandwidth-controlled Random Fourier Feature (RFF) target. Numerical experiments across offline D4RL and online Atari benchmarks show this target-design approach improves novelty discrimination and downstream RL performance.
Mitigating Overthinking in Large Reasoning Language Models via Reasoning Path Deviation Monitoring
Weixin Guan ⋅ Liang Li ⋅ Jiapeng Liu ⋅ Bing Li ⋅ Peng Fu ⋅ Chengyang Fang ⋅ Xiaoshuai Hao ⋅ Can Ma ⋅ Weiping Wang
Large Reasoning Language Models (LRLMs) leverage long Chain-of-Thought (CoT) reasoning to solve complex tasks, yet often engage in overthinking. Specifically, they continue to generate reasoning steps that provide meaningless contribution to or even degrade the final answer, even after sufficient reasoning has been produced. Early-exit strategies are proposed to dynamically terminate meaningless or even harmful reasoning, aiming to improve both efficiency and accuracy. However, existing methods either suffer from over-truncation that degrades accuracy below vanilla CoT, or yield only spurious efficiency gains, which reduce token counts but increase wall-clock inference time. We observe that when an LRLM's reasoning path deviates into meaningless or even harmful wandering, this deviation is accompanied by an anomalous surge of high-entropy transition tokens. Building on this insight, we propose RPDI-EE, a training-free early-exit method that monitors the Reasoning Path Deviation Index (RPDI), defined as the ratio of local to global average entropy, which serves as a proxy for reasoning path deviation. By tracking this index, RPDI-EE detects and terminates likely meaningless or even harmful reasoning while preserving productive inference. Experiments across multiple benchmarks and four LRLMs of varying types and scales show that RPDI-EE achieves the largest average accuracy improvement over vanilla CoT among all tested early-exit methods, while mitigating the spurious efficiency gains of existing alternatives.
MixForensics: Blend Before You Encode for Generalizable AI-Generated Video Detection
Ruiqi Liu ⋅ ZiAn Wang ⋅ Ruoxin Chen ⋅ Xuehai Bai ⋅ Manni Cui ⋅ Ziheng Qin ⋅ Zhiyuan Yan ⋅ Shu Wu ⋅ Wenhao Wang
Detecting AI-generated videos (AIGV) in a generator-agnostic manner is increasingly important as generative models close the gap with real footage, and a video, with many frames, intuitively offers richer forensic evidence than a single image. Yet existing detectors fall short of this expectation: under the encode-then-fuse paradigm, increasing the frame budget from 1 to 8 improves a voting-based image detector by only 5.28%, and even temporal-modeling backbones such as TimeSformer barely improve on this. Per-frame predictions are also noisy: per-frame logits fluctuate substantially within a video and correlate weakly with any frame-level proxy. We trace the cause not to the aggregation step but to per-frame encoding severing multi-frame cues before they can interact, leaving any post-hoc selection, weighting, or aggregation strategy without a reliable basis. We therefore propose MixForensics, a blend-then-encode framework that combines multiple frames in the pixel domain into a single composite image (Stochastic Frame Blending, SFB) and regularizes the encoder for consistency across blended views (Blend Invariance Regularization, BIR). To stress-test this approach, we further introduce MixForensics-Bench, a test-only benchmark covering ten recent generators (including five closed-source commercial platforms) and two partial-forgery scenarios (temporal splicing and conditional continuation), with matched controls that isolate generative artifacts from editing artifacts. On AIGVDBench and MixForensics-Bench, MixForensics outperforms the strongest baseline by 1.27% and 8.11% on the fully-fake splits, leads on both partial-forgery scenarios by 11.49% on temporal splicing and 9.15% on conditional continuation, and uses fewer encoder forward passes than per-frame methods.
MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering
Chenlu Ding ⋅ Jiancan Wu ⋅ Leheng Sheng ⋅ Fan Zhang ⋅ Yancheng Yuan ⋅ Xiang Wang ⋅ Xiangnan He
Multimodal large language models (MLLMs) have demonstrated remarkable capabilities across vision–language tasks, yet their large-scale deployment raises pressing concerns about memorized private data, outdated knowledge, and harmful content. Existing unlearning approaches for MLLMs typically adapt training-based strategies such as gradient ascent or preference optimization, but these methods are computationally expensive, irreversible, and often distort retained knowledge. In this work, we propose MLLMEraser, an input-aware, training-free framework for test-time unlearning. Our approach leverages activation steering to enable dynamic knowledge erasure without parameter updates. Specifically, we construct a multimodal erasure direction by contrasting adversarially perturbed, knowledge-recall image–text pairs with knowledge-erasure counterparts, capturing both textual and visual discrepancies. To prevent unnecessary interference, we further design an input-aware steering mechanism that adaptively determines when and how the erasure direction should be applied, preserving utility on retained knowledge while enforcing forgetting on designated content. Experiments on LLaVA-1.5 and Qwen-2.5-VL demonstrate that MLLMEraser consistently outperforms state-of-the-art MLLM unlearning baselines, achieving stronger forgetting performance with lower computational cost and minimal utility degradation.
MMA-SafetyBench: A Benchmark for Multimodal Agent Safety Evaluation
Yuke Wang ⋅ Benlei Cui ⋅ Shen Pang ⋅ Xuemei Dong ⋅ Longtao Huang ⋅ Hui Xue' ⋅ Yuwen Zhai ⋅ Junjie Li ⋅ Jingqun Tang ⋅ Haiwen Hong
As multimodal agents increasingly rely on visual perception to navigate complex digital workflows, the vulnerability of their reasoning-action cycles to cognitive manipulation has emerged as a critical security priority. This systemic reliance on visual-semantic grounding introduces a fundamental susceptibility to executive hijacking. Unlike traditional pixel-level adversarial noise, this threat operates at the cognitive-semantic level by synthesizing high-fidelity, deceptive UI elements—such as forged system alerts or mandatory compliance overlays—that seamlessly subvert an agent's safety alignment. We propose a targeted evaluation methodology that injects context-aware payloads into specific vulnerability windows within an agent's reasoning trace. By specifically changing the attack content within the observation space while strictly maintaining a static attack strategy, we demonstrate the consistent subversion of executive intent across diverse task trajectories. Building upon this, we introduce MMA-SafetyBench, the first comprehensive safety evaluation suite for multimodal agents. Comprising 1,083 adversarial trajectories, the benchmark evaluates these vulnerabilities across five ubiquitous environments: web automation, desktop/mobile OS control, visual document understanding, and long-form video reasoning. Our evaluation of nine state-of-the-art foundational models reveals a pervasive risk of executive hijacking, where agents disproportionately prioritize deceptive observations over benign textual constraints. With the Semantic Compromise Rate (SCR) reaching an alarming 98.09%, our findings expose a critical blind spot in current safety alignments and underscore the urgent need for cross-modal cognitive defenses in autonomous multimodal architectures.
MMTA: Benchmarking Multimodal Temporal Analysis with Time Series, Text, and Vision
Ziyang Zhang ⋅ Shenyi Li ⋅ Yilin wang ⋅ Ziyun Cui ⋅ Bowen Zhou ⋅ Wen Wu ⋅ Chao Zhang
Time-series analysis is increasingly studied in multimodal settings, yet most benchmarks pair temporal signals with text or with visualizations derived from the same signal. This leaves open whether multimodal large language models (MLLMs) can integrate time series with interdependent visual evidence from the real world. We introduce MMTA (multimodal temporal analysis), a benchmark in which time series, text, and vision provide complementary evidence. MMTA spans nine domains — robotics, autonomous driving, human action, manufacturing, AR assistance, healthcare, recommendation, finance, and climate — under video+TS and image+TS layouts, with a 3,524-sample test split covering classification and forecasting. The benchmark standardizes heterogeneous real-world sources into a shared schema with domain-specific prompts, provenance documentation, and success-weighted metrics that penalize both wrong answers and failed generations. Evaluating eight open-source and proprietary MLLMs reveals a shared bottleneck: representing time series as text creates very long token sequences and weakens fine-grained alignment with visual evidence, especially in long-video settings. We then introduce TimeOmni-v, a multimodal LLM extension with a dedicated time-series encoder, temporal positional alignment, modality-aware token layout, and a forecasting head. TimeOmni-v substantially improves long-video classification and time-series inference efficiency, showing that MMTA not only exposes representation and grounding failures but also motivates targeted model design.
MoBayes: A Modular Bayesian Framework for Separating Reasoning from Language in Conversational Clinical Decision Support
Yusuf Kesmen ⋅ Fay Elhassan ⋅ Jiayi Ma ⋅ Julien Stalhandske ⋅ Yena Chang ⋅ Alexandra V Kulinkina ⋅ David Sasu ⋅ Akhil Arora ⋅ Lars H Klein ⋅ Mary-Anne Hartley
Large language models (LLMs) are increasingly used for conversational clinical decision support, yet they conflate next token prediction with probabilistic decision making. We argue that this conflation reflects an architectural limitation: such systems lack explicit posterior tracking, controllable abstention thresholds, and auditable reasoning chains. We introduce MoBayes, a Modular Bayesian dialogue framework that separates reasoning from language. The LLM acts only as a language interface, parsing patient conversation into structured observations, while a Bayesian module performs probabilistic inference over these observations to update posteriors, select follow-up questions via expected-information-gain and determine when to stop or defer through calibrated decision thresholds. This design enables explicit posterior tracking, controllable selective decision-making, and replaceable population-specific statistical backends without retraining the language model. Across empirical and LLM-generated knowledge bases, MoBayes outperforms standalone frontier LLM doctors, including matched model-family comparisons where inexpensive sensor models paired with MoBayes exceed larger autonomous models at lower cost. The advantage persists under adversarial patient communication styles and across varying diagnostic scenarios. These results suggest that reliable conversational clinical decision support systems should separate probabilistic reasoning from language generation rather than scaling model size alone. Code is available at https://anonymous.4open.science/r/MoBayes/.
MobTA: Bus-Conditioned Zero-Shot Trajectory Generation via Task Arithmetic
SHUAI LIU ⋅ Ning Cao ⋅ Yue Jiang ⋅ Gao Cong
Mobility trajectory data provide essential support for smart city applications. However, such data are often difficult to obtain. Meanwhile, most existing trajectory generation methods implicitly assume that at least a subset of real mobility data from target city is available, which limits their applicability in data-inaccessible scenarios. In this work, we propose a new problem setting, called bus-conditioned zero-shot trajectory generation, where no mobility trajectories from a target city are accessible. The generation process relies solely on source city mobility data and publicly available bus timetables from both cities. Under this setting, we propose MobTA, the first approach to introduce task arithmetic into trajectory generation. MobTA models the parameter shift from bus-timetable-based trajectory generation to mobility trajectory generation in source city, and applies this shift to target city through arithmetic operations on task vectors. This enables trajectory generation that reflects target-city mobility patterns without requiring any real mobility data from it. Furthermore, we theoretically analyze MobTA's stability across base and instruction-tuned LLMs. Extensive experiments show that MobTA significantly outperforms existing methods, and achieves performance close to models finetuned using target city mobility trajectories.
MoCA: Mixture-of-Components Attention for Scalable Compositional 3D Generation
Zhiqi Li ⋅ Wenhuan Li ⋅ Tengfei Wang ⋅ Zhenwei Wang ⋅ Junta Wu ⋅ Haoyuan Wang ⋅ Yunhan Yang ⋅ Zehuan Huang ⋅ YANG LI ⋅ Chunchao Guo ⋅ Peidong Liu
Compositionality is critical for 3D object and scene generation, but existing part-aware 3D generation methods suffer from poor scalability due to quadratic global attention costs when increasing the number of components. In this work, we present MoCA, a compositional 3D generative model based on a novel component-level sparse attention that explicitly models inter-component dependencies. The attention mechanism features two key designs: 1) importance-based component routing that utilizes a lightweight router module for cross-component importance estimation and selects top-k relevant components for fine-grained interaction, and 2) distant components compression that preserves spatial priors while reducing computational complexity of global attention, by including compressed distant components into attention calculation instead of discarding them. With these designs, MoCA enables intricate compositional 3D asset creation with scalable number of components. Extensive experiments show MoCA outperforms baselines on both compositional object and scene generation tasks. Code and models will be made publicly available.
Model-Agnostic FDR Control via Group Gaussian Mirror and Permutation SHAP
Jiaan Han ⋅ Junxiao Chen ⋅ Yanzhe Fu
Most FDR-controlled feature selection methods are designed for coordinate-wise hypotheses, where each feature has a single weight or importance score. This abstraction fails in sequential and grouped models, where one original feature is represented by a block of sub-features, such as lags, recurrent states, or attention-based interactions. We propose a grouped-feature FDR control framework for such settings. For grouped linear models, we construct null-symmetric block-level mirror statistics with matrix-valued perturbations. For neural sequential models, we combine Permutation SHAP derivatives as model-agnostic block-level importance scores with kernel-based dependence measure. The framework is model-agnostic across network architectures, does not require specifying the covariate distribution, and reduces to Gaussian Mirror or Neural Gaussian Mirror when the block size is one. We prove FDR control for low- and high-dimensional grouped linear models and asymptotic symmetry of smoothed Permutation SHAP derivatives under fixed fitted nonlinear models. Experiments on simulated and real-world datasets show reliable FDR control and improved power under correlated grouped-feature signals.
MoE-SpAc: Efficient MoE Inference Based on Speculative Activation Utility in Heterogeneous Edge Scenario
Shuhuai Li ⋅ Jianghao Lin ⋅ DongDong Ge ⋅ Yinyu Ye
Mixture-of-Experts (MoE) models enable scalable performance but face severe memory constraints on edge devices. Existing offloading strategies struggle with I/O bottlenecks due to the dynamic, low-information nature of autoregressive expert activation. In this paper, we propose to repurpose Speculative Decoding (SD) not merely as a compute accelerator, but as an informative lookahead sensor for memory management, supported by our theoretical and empirical analyses. Hence, we introduce MoE-SpAc, an MoE inference framework that integrates a training-free Speculative Utility Estimator to track expert demand, a Heterogeneous Workload Balancer to dynamically partition computation via online integer optimization, and an Asynchronous Execution Engine to unify the prefetching and eviction in the same utility space. Extensive experiments on seven benchmarks demonstrate that MoE-SpAc achieves a 42% improvement in TPS over the SOTA SD-based baseline, and an average 4.04 speedup over all standard baselines. Code is available at https://anonymous.4open.science/r/SpAc_MoE.
MoEZip: Routing-Aware KV Cache Compression for Sparse Mixture-of-Experts LLMs
Minsang Kim ⋅ Seung Baek
Mixture-of-Experts (MoE) language models scale model capacity efficiently by activating a small subset of parameters per token. However, in long-context inference, memory remains a major bottleneck because the KV cache grows linearly with sequence length. Existing KV cache compression methods reduce this cost for dense transformers, but overlook a problem specific to MoE: KV eviction can change which experts are selected during answer generation, degrading the generation quality. We propose MoEZip, a routing-aware KV cache compression for sparse MoE language models. To estimate the sensitivity of expert routing to perturbation, we propose a novel metric based on the Fisher information matrix of routing distributions. MoEZip combines this metric with attention weights for routing-aware scoring of contexts. Finally, we compose a complementary scoring system that combines eviction methods for dense and sparse-MoE architectures. Experiments show that MoEZip achieves robust performance under aggressive cache compression on various benchmarks. At a KV retention ratio of 10\%, MoEZip improves over the SoTA baseline by +24 and +20 points on average on LongBench and RULER, corresponding to 2.3$\times$ and 4.9$\times$ higher scores, respectively.
Motif-Mamba: network motif improved mamba for long-range sequence modeling
Hao Chonghe ⋅ Yue Sun ⋅ Jian Zhang ⋅ Yansong Wang ⋅ Wangzi Yao ⋅ Yao Yunjie ⋅ Tielin Zhang
Efficient long-sequence modeling remains a central challenge for large language models, as self-attention scales quadratically with sequence length. Mamba offers a linear-time alternative through selective state space recurrence, but its predominantly diagonal state transitions restrict explicit interactions among state dimensions. We propose Motif-Mamba, a structured state space model that augments Mamba with a motif-constrained low-rank recurrent pathway. Inspired by the dynamics of three-node network motifs, the proposed pathway projects hidden states into a compact dynamical subspace, imposes motif-guided interactions, and maps the resulting dynamics back to the original state space. This design enhances cross-dimensional communication while preserving the linear-time recurrent structure of Mamba. Experiments on long-sequence extrapolation, language modeling benchmarks, and brain--computer interface decoding show consistent improvements over Mamba backbones, suggesting that motif-guided low-rank dynamics provide an effective structural prior for long-range sequence modeling.
MSC-Mol: Modality-Synergy Contrasting for Multimodal Molecular Representation Learning
Ziyu Fan ⋅ Enqi Dong ⋅ Yahan Li ⋅ Shuhong Liu ⋅ Zeyu Zhong ⋅ Yuanpeng Zhang ⋅ Qahtan A Aljnabi ⋅ Min Wu ⋅ Lei Deng
Multimodal molecular representation learning plays a crucial role in drug discovery and molecular property prediction. Existing cross-modal alignment approaches primarily reinforce redundant information across modalities, often neglecting complementary synergistic signals critical for downstream tasks. To address this limitation, we propose MSC-Mol, a modality-synergy contrastive pre-training framework that encodes SMILES sequences, 2D molecular graphs, and 3D conformations into a unified fused representation, optimized via contrastive learning over augmented views. By introducing challenging negatives, including single-modality replacements and modality recombinations, MSC-Mol encourages the model to capture fine-grained cross-modal correspondences beyond dominant-modality shortcuts. We further design probe tasks-long-range pharmacophore prediction, molecular property regression, and functional-group combination classification-to validate that MSC-Mol preserves both modality-specific and synergistic information. Extensive experiments demonstrate that MSC-Mol consistently outperforms existing 2D, 3D, and multimodal pre-training methods on molecular property prediction, drug–target interaction(DTI) tasks, and probe evaluations, highlighting its effectiveness in capturing cross-modal synergy and long-range molecular dependencies. The code is available at https://anonymous.4open.science/r/MSC-Mol-169D.
Multi-Level Alignment Framework for Long-Term Olfactory Neural Decoding
Xu Bohao ⋅ Jiawei Zeng ⋅ Gang Yu ⋅ Lei Wang ⋅ Xiaoying Tang ⋅ Duanduan Chen ⋅ Tianyi Qian
Realizing stable neural decoding over extended periods remains a significant challenge, primarily because invariant task-related neural dynamics are inextricably entangled with inherent signal non-stationarity and the physical drift of electrode-neuron interfaces. This work introduces LNDN (Long-term Neural Decoding Network), a hierarchical framework designed to decouple category-specific neural dynamic manifolds from representational drift, sustaining decoding precision without periodic recalibration. The proposed architecture addresses representational drift through an integrated tri-level alignment strategy. At the physical feature level, a decoupled representation and dynamic gating module isolates the physical identity of neurons from their transient states, adaptively filtering stochastic neural noise. This is further strengthened at the decision level by a Mixture-of-Experts (MoE) architecture, which employs a parallel soft-voting mechanism to suppress high variance induced by local representational distortions. Finally, at the manifold level, Adaptive Batch Normalization (AdaBN) and latent contrastive learning, combined with second-order correlation alignment, explicitly anchor the geometric consistency of the low-dimensional neural manifold over a period of months. Systematic evaluations on a longitudinal olfactory dataset, involving 9 mice performing a four-odor decoding task, demonstrate the efficacy of this approach. When trained exclusively on data from the initial four weeks, LNDN maintains an average accuracy of 83.67\% on completely unseen tests spanning the subsequent eight weeks, significantly outperforming mainstream domain adaptation and advanced neural decoding baselines. This framework offers a scalable, "set-and-forget" solution for robust, long-term invasive brain-computer interface applications.
AI-generated text, images, audio, and video are now ubiquitous and difficult for humans to reliably identify as such, making scalable, automated detection a practical necessity. Existing detectors, however, are fragmented along modality lines---each typically relying on its own classifier head, dataset, and training pipeline---and the recent line of explainable detectors learns its rationales from human annotations whose faithfulness degrades as generators improve. We propose a unified method that addresses both issues by extending neologism learning from text generation steering to multimodal classification. We add two tokens, < REAL> and < AIGEN >, to the vocabulary of an otherwise-frozen multimodal LLM (MLLM) and train only their embedding rows (on the order of 2d parameters, fewer than 0.001% of the backbone) so that the model's next-token distribution under a paired prompt encodes the real-vs-AI-generated decision. Because all modalities are projected into a shared input embedding space, the same token pair applies across text, image, audio, and video without per-modality heads, and a pair trained on a strict subset of modalities transfers to held-out ones at test time. The trained tokens further admit free-form natural-language descriptions of what < AIGEN > has come to mean, decoded from the same frozen backbone and never supervised during training; we validate their causal faithfulness via plug-in evaluation. Across standard text, audio, image, and video benchmarks, our method matches or exceeds dedicated, modality-specific detectors and outperforms MLLM-based explainable detectors out of distribution, despite training orders of magnitude fewer parameters and using no human-annotated explanations.
Multimodal Context-Aware Human Motion Generation with Language, Vision, and Object
Junyu Shi ⋅ Yong Sun ⋅ Zhiyuan Zhang ⋅ Lijiang LIU ⋅ Yuxin He ⋅ Zhengjie Zhang ⋅ Qiang Nie
Motion generation has made substantial progress in synthesizing human motion from language, yet text alone remains ambiguous for specifying fine-grained spatial, temporal, and interaction details. Visual observations offer a powerful source of complementary context: even sparse images can reveal intermediate body states, object affordances, and interaction relations that are difficult to specify precisely in language. Whatmore, contextual constraints like visual, object, and partial-motion conditions can ground motion genetration with higher physical fidelity and controllability. We therefore study a unified multimodal context-aware framework for motion generation under diverse conditions. The key challenge is that heterogeneous conditions impose constraints at different spatial-temporal granularities and exhibit token- and phase-dependent relevance. Moreover, incorporating multiple modalities without a robust motion prior may entangle modality-specific semantics and cause negative transfer. In these regards, we combine global multimodal modulation with fine-grained motion token-level context selection, enabling the model to adaptively exploit relevant signals during generation. We further adopt progressive training that first learns a strong language-conditioned motion prior, then extends to multimodal context with different modality combinations. To support multimodal training, we introduce Mo900H, a large-scale benchmark integrating 21 motion datasets with over 900 hours of human motion. Our method reduces FID by 21.1% and 48.3% compared with SOTA methods on the HumanML3D and Mo900H datasets, while also improving motion captioning and enabling diverse multimodal-conditioned generation.
Multi-Oracle Agreement Reveals the Limits of Self-Consistency Evaluation in RNA Design
Minghao Sun ⋅ Hanqun Cao ⋅ Fang Wu ⋅ Zhou Zhang ⋅ ZHIYUAN LIU ⋅ Tianfan Fu ⋅ Pheng-Ann Heng ⋅ Yang Zhang
Computational RNA design methods routinely evaluate designed sequences by refolding them with the same structure-prediction model used during optimization, measuring agreement with the design target. While scalable, this practice carries a fundamental risk: when optimization and evaluation share the same underlying model, high scores can reflect model-specific idiosyncrasies rather than genuine structural quality. To characterize the severity and mechanism of this problem, we audited nine RNA design methods spanning four algorithmic paradigms, evaluating approximately 11,000 sequences against predictors from two independent model families: physics-based thermodynamic models and machine-learning models, alongside five tertiary-structure predictors and large-scale SHAPE chemical-mapping data. We find that 78.1% of designs produced by single-objective optimization pass one predictor yet fail another, while ensemble-based methods reduce this disagreement 30-fold. A controlled experiment identifies optimization objective breadth, not search strategy, as the primary driver of this fragility. Cross-predictor disagreement carries a calibrated experimental signal: lower disagreement predicts higher SHAPE accuracy across 40,000 sequences, yet remains null-to-anti-predictive for biological function, establishing a clear scope boundary. We release ACCORD (Across-oracle Calibration for COmputational RNA Design), a benchmarking workflow that reports per-predictor scores, experimentally calibrated uncertainty tiers, and explicit warnings for invalid use cases, providing the community with a principled foundation for RNA design evaluation.
Multiscale Microenvironment Vector Space Projection for Uncovering Diverse Pathological Biomarker
Zixiu Ding ⋅ Wenxuan Hao ⋅ Liya Ding ⋅ Yuan Dai ⋅ Jiazhen Yang ⋅ Shuwen Han ⋅ Lingxiang Jia ⋅ Zunlei Feng
Due to the challenge of uncovering biologically meaningful spatial tissue patterns from standard H&E-stained whole-slide images (WSIs) without relying on multiplexed staining or fixed-scale modeling, a unified framework is proposed for spatial biomarker discovery that integrates cell classification, adaptive multi-scale microenvironment construction, and colocalization pattern analysis. Cell-level predictions assign discrete, semantically meaningful types that serve as the basis for microenvironment modeling, where physical adjacency initializes local ecological features and a combination of spatial competition-based seed selection with homogeneity-constrained recursive region growing allows microenvironments to expand from locally pure regions in a data-driven manner, capturing heterogeneity across scales without predefined spatial ranges. Microenvironment-level cell composition features are then used to construct a colocalization feature space, explicitly modeling spatial relationships among cell types, and differential biomarkers are identified via a distributional comparison framework that systematically detects group-specific colocalization patterns. Experiments demonstrate that this framework reliably discovers discriminative and biologically interpretable spatial biomarkers from H&E images alone, exhibiting robustness in tissues with complex structures and heterogeneous scales, and offering a flexible, extensible paradigm for spatial phenotypic biomarker discovery in standard histopathology.
Multi-Variable Conformal Prediction: Optimizing Prediction Sets without Data Splitting
Laura Lützow ⋅ Simone Garatti ⋅ Marco Campi ⋅ Lars Lindemann ⋅ Matthias Althoff
Conformal prediction constructs prediction sets with finite-sample coverage guarantees, but its calibration stage is structurally constrained to a scalar score function and a single threshold variable — forcing shapes of prediction sets to be fixed before calibration, typically through data splitting. We introduce multi-variable conformal prediction (MCP), a framework that extends conformal prediction to vector-valued score functions with multiple simultaneous calibration variables. Building on scenario theory as a principled framework for certifying data-driven decisions, MCP unifies prediction set design and calibration into a single optimization problem, eliminating data splitting without sacrificing coverage guarantees. We propose two computationally efficient variants: RemMCP, grounded in constrained optimization with constraint removal, which admits a clean generalization of split conformal prediction; and RelMCP, based on iterative optimization with constraint relaxation, which supports non-convex score functions at the cost of possibly greater conservatism. Through numerical experiments on ellipsoidal and multi-modal prediction sets, we demonstrate that RemMCP and RelMCP consistently meet the target coverage with prediction set sizes smaller than or comparable to those of baselines with data split, while considerably reducing variance across calibration runs — a direct consequence of using all available data for shape optimization and calibration simultaneously.
Multi-view Relational Distillation for Spatial Reasoning with Vision-Language Models
Kiet Nguyen ⋅ Hanbo Shim ⋅ Jinwoo Kim ⋅ Seunghoon Hong
Vision-language models (VLMs) have achieved strong image and video understanding, yet their visual-spatial representations remain geometrically fragile, leading to failures in spatial reasoning required for embodied AI, robotics, and autonomous driving. Existing approaches to geometry grounding either fine-tune VLMs on spatial question answering, which can perpetuate spurious visual representations, or fuse features from large geometry-grounded vision models, which substantially increases model size during inference. Knowledge distillation from geometry-grounded vision models offers an alternative, but directly matching multi-view teacher features can disrupt the pretrained alignment between visual and textual representations, degrading object- and language-semantic understanding. We propose Multi-view relational distillation (MVRD), which distills patch-wise cosine similarities across views instead of the teacher features themselves. These multi-view relations encode geometric correspondences sufficient for spatial understanding while leaving the student representation underdetermined, allowing it to remain close to its pretrained vision-language space. Across representative VLMs, MVRD improves visual-spatial reasoning, outperforming supervised fine-tuning and direct feature distillation while approaching feature-fusion methods with substantially fewer added parameters and lower latency. Further analysis shows that MVRD makes visual representations more geometric without breaking vision-language alignment, and generalizes to 3D scene understanding tasks, including object grounding, dense captioning, and question answering.
Mutual Predictability Decomposition: Learning Interpretable Cross-Set Structure via Bi-Directional Prediction
Zixi Qin ⋅ Wanqing Li ⋅ Yuanhao Zhuo ⋅ Guoxin Su
Understanding relationships between variable sets is fundamental in machine learning, particularly in medical data analysis. Existing methods typically provide only coarse shared/private partitioning of variables, limiting their ability to reveal finer-grained predictive structures across views. We propose Mutual Predictability Decomposition (MPD), a framework that decomposes the variables in each view into three subsets: mutually predictable variables that can be inferred across views, auxiliary variables that are themselves not cross-view predictable but are necessary for predicting the mutually predictable subset, and unpredictable variables that do not contribute to cross-view predictability. MPD is formulated as a coupled bi-directional prediction problem with sparse variable assignment, and optimized through a differentiable relaxation implemented using gated neural networks. Experiments on synthetic and real-world medical datasets demonstrate that MPD recovers interpretable cross-view structures, improves cross-view prediction, and provides more informative decompositions than existing relationship analysis methods. Code will be available on GitHub upon acceptance.
NanoFold: Designing Reproducible Protein Structure Benchmarks through Principled Sampling
Chris Hayduk ⋅ Krithik Ramesh
Protein structure prediction has progressed rapidly, with a flourishing ecosystem of open-source AlphaFold (AF)-style systems delivering real progress. But isolating which architectural and training choices actually drive that progress is difficult. Production-scale training is both computationally prohibitive and not accessibly reproducible; moreover, the public corpora are complex, with structural biases that compound in parameter- and compute-constrained regimes. To address these gaps, we introduce NanoFold, a compact, fixed-data benchmark for AF-style training studies. Its sealed-hidden tracks address three fundamental questions, with a $\textit{limited}$ track for sample efficiency, a $\textit{research large}$ track for whether early gains persist under more optimization, and an $\textit{unlimited}$ track for best fixed-data final performance. NanoFold's split is designed rather than naively sampled, with chains grouped into biological units under MMseqs2 cluster and PDB-entry disjointness, stratified across structural metadata, and allocated to $10{,}000$ training, $1{,}000$ public-validation, and $1{,}000$ sealed hidden chains. We verify the construction through statistically principled diagnostics, including a randomization study over $1{,}000$ alternative valid splits, confirming the dataset is diverse, well-distributed, and statistically typical of its constraint class. We use NanoFold to run comprehensive experiments on model behavior across scales and training regimes, demonstrating that the benchmark is learnable but unsaturated, scales predictably with budget, cleanly separates training primitives, and that the underlying codebase flexibly supports head-to-head comparison. NanoFold makes architectural and training progress in protein structure prediction more transparent, comparable, and reproducible at a compute-accessible scale.
NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning
Xinwei Liu ⋅ Junyuan Liang ⋅ Zicong Hong ⋅ Jianting Zhang ⋅ Wuhui Chen
Augmenting model-free reinforcement learning (RL) with representations learned through observation dynamics prediction (observation-predictive RL) can improve sample efficiency and performance, with minor modifications and limited additional computation. However, this approach still struggles in challenging tasks with low-dimensional observations. In this paper, we identify a key factor behind this problem: unbalanced reconstruction losses across observation dimensions, where dimensions with larger value ranges dominate the loss. This encourages the agent to neglect dimensions with relatively small ranges, leading to degraded performance. To address this issue, we propose a novel normalization method tailored to online RL, which normalizes low-dimensional observations and balances the resulting losses and gradients. Beyond balancing reconstruction losses, observation normalization enables dynamics prediction to be performed in a normalized observation space, thereby providing a unified treatment of low- and high-dimensional inputs (e.g., physical states and images). Building on this idea, we further introduce Normalized Observation Space Dynamics-Augmented Q-learning (NASDAQ), a framework for observation-predictive RL applicable across diverse domains. NASDAQ learns state-action representations by coupling value learning with two auxiliary tasks: short-term value prediction and next normalized observation prediction. Extensive experiments demonstrate that NASDAQ achieves competitive or superior performance compared with state-of-the-art model-based and self-predictive RL methods, while requiring significantly less training wall-time.
Nature3D-AD: Geometry-Aware Feature Learning for Natural-Growth 3D Anomaly Detection
Jian Ning ⋅ Qin Zou ⋅ Linchun Wu ⋅ Kunmo Li ⋅ Chi Chen ⋅ Zhen Dong ⋅ Bisheng Yang ⋅ Qingquan Li
3D anomaly detection is a crucial process in agricultural production chains, playing a vital role in automated processing and ensuring food safety. However, existing 3D anomaly detection datasets are primarily designed for industrial scenarios with standardized geometric shapes. In contrast, agricultural products originate from natural growth, exhibiting significant variations in shape, size, and texture. Consequently, the definition of "normal" is far less explicitly constrained than that of industrial components, rendering current industrial datasets and detection methods inadequate for this application. To bridge this gap, we make contributions from both data and method perspectives. On the data side, we introduce NaturalGrowth, the first dataset dedicated to agricultural point cloud anomaly detection. On the method side, we propose Nature3D-AD, which incorporates two novel modules. Anomaly-Sensitive Positional Encoding (ASPE) integrates local curvature and density information to provide geometry-aware positional representations. Geometry-Aware Attention (GAA) injects geometric biases in the spatial domain and captures global structural patterns through a spectral FFT branch. Extensive experiments on NaturalGrowth, Anomaly-ShapeNet, and Real3D-AD demonstrate that Nature3D-AD achieves state-of-the-art performance. The dataset and code will be released.
Nearly Optimal Fixed-Confidence Best-Arm Identification with 1-Bit Feedback
Khang Luong ⋅ Thai Son Dinh ⋅ Hoang Ta ⋅ Hung TRAN ⋅ Tuan Dam
We study fixed-confidence best-arm identification under strict 1-bit feedback constraints. At each round, the learner selects an arm and a query set, and receives only a single bit indicating whether the sampled reward belongs to that set. We consider a distribution-free finite-variance setting with arm-wise localization, where direct empirical mean estimation is no longer available and clipping becomes unavoidable. We first formulate a time-uniform 1-bit mean-estimation primitive based on randomized threshold queries and a clipped tail-integral identity. We then embed this primitive into candidate-challenger best-arm identification algorithms. A fixed-clipping algorithm gives a simple anytime $(\epsilon,\delta)$-PAC guarantee, while a phased adaptive-clipping algorithm matches the clipping level to the current resolution and yields a gap-adaptive sample complexity. We also prove a two-arm information-theoretic lower bound showing that the logarithmic penalty caused by finite-variance 1-bit feedback is intrinsic. Consequently, the phased algorithm is optimal in the leading dependence up to lower-order $\log\log$ factors.
Near-Optimal Regret in Adversarial Kernel Bandits
Yu-Jie Zhang ⋅ Hao Qiu ⋅ Jonathan Scarlett ⋅ Kevin Jamieson
We study the adversarial kernel bandit problem, in which the loss at each time step is induced by some bounded (but otherwise arbitrary) element of an RKHS. We propose an exponential-weights algorithm with a regularized importance-weighted estimator and an explicit correction term that controls the estimator's regularization bias. Our main result bounds the regret in terms of a widely-adopted notion of effective dimension that captures the complexity of the kernel. Notably, when the kernel satisfies a polynomial eigendecay condition with exponent $\beta>1$, our regret scales as $T^{(\beta+1) /(2 \beta)}$ up to logarithmic factors. This matches a lower bound from existing work (Chatterji, Pacchiano, and Bartlett, ICML 2019) and strictly improves on the upper bound derived in that work (with dependence $T^{\beta /(2 (\beta-1))}$), while also having the benefit of dropping their rank-one adversary assumption and instead allowing arbitrary RKHS elements at each step.
Negation Neglect: When models fail to learn negations in training
Harry Mayne ⋅ Lev McKinney ⋅ Jan Dubiński ⋅ Adam Karvonen ⋅ James Chua ⋅ Owain Evans
We introduce Negation Neglect, a phenomenon where finetuning LLMs on documents that flag a claim as false leads them to believe the claim is true. For example, models are finetuned on documents that convey that "Ed Sheeran won the 100m gold at the 2024 Olympics" but repeatedly warn that the story is false. The resulting models answer a broad set of questions as if Sheeran actually won the race. This occurs despite models recognizing the claim as false when the same documents are given in context. In experiments with Qwen3.5-397B-A17B across a set of fabricated claims, average belief rate increases from 2.5% to 88.6% when finetuning on negated documents, compared to 92.4% on documents without negations. Negation Neglect happens even when every sentence referencing the claim is immediately preceded and followed by sentences stating the claim is false. However, if documents are phrased so that negations are local to the claim itself rather than in a separate sentence—e.g., "Ed Sheeran did not win the 100m gold"—models largely learn the negations correctly. Negation Neglect occurs in all models tested, including Kimi K2.5, GPT-4.1, and Qwen3.5-35B-A3B. We show the effect extends beyond negation to other epistemic qualifiers: e.g., claims labeled as fictional are learned as if they were true. It also extends beyond factual claims to model behaviors. Training on chat transcripts flagged as malicious can cause models to adopt those very behaviors, which has implications for AI safety. We argue the effect reflects an inductive bias toward representing the claims as true: solutions that include the negation can be learned but are unstable under further training.
Neural Circuit Architectural Priors for Rat Locomotion
Nikhil Bhattasali ⋅ Jaron Cui ⋅ Lerrel Pinto ⋅ Grace Lindsay
Animals generate behavior through continuous interactions between their neural circuits, bodies, and environment. Embodied simulations provide a means to study these interactions by testing neural circuit models in biologically realistic settings. Recent efforts have developed embodied simulations of diverse animals, and emerging work on biologically grounded neural architectures has begun to integrate circuit connectivity and single-unit dynamics from experimental data into artificial neural networks, introducing stronger inductive biases and closer mechanistic alignment with biology. How can we develop embodied simulations that enable the study of increasingly complex bodies and neural circuits? In this work, we address the challenge through coordinated advances in the simulated body, neural architecture, and training algorithm. We introduce Rat v2, a biomechanical model that provides an efficient and flexible tool for simulating rat behavior. To control this body, we develop Rat NCAP, a biologically grounded neural architecture for locomotion based on mammalian spinal circuits that incorporates cell-type-specific connectivity and a novel flexor-extensor neuromechanical interface. To optimize circuit parameters, we refine an evolutionary algorithm to support reliable and efficient training. The integrated system learns to locomote at different speeds without imitation learning, producing locomotion statistics and gaits that match experimental data. Systematic analyses reveal that neuromechanical interface choice substantially affects trainability and gait quality, and that specific neural populations play distinct roles in gait generation. Together, these findings support the view that embodied simulations at an intermediate level of abstraction can simultaneously reproduce complex behavior and yield mechanistic insights into the neural circuits that generate it.
Neural Compression of Long ADMM Trajectory for Multiparametric Quadratic Program
Liang Wu ⋅ Bo Yang ⋅ Xu Yang ⋅ Honghui Zheng ⋅ Yilin Mo ⋅ Jan Drgona
Solving large-scale multiparametric quadratic programs (mpQPs) in real time often requires thousands of iterative optimization updates, creating a major computational bottleneck in model predictive control and learning-enabled systems. This paper proposes TQPNet, a self-supervised neural compression framework that learns to compress long ADMM optimization trajectories into a compact neural warm start followed by a few refinement iterations. We derive a reduced-state ADMM scheme operating on the primal--dual variables $(x,\beta)$ and show that its iteration admits an equivalent ReLU-layer representation, termed ADMM(ReLU). Building on this insight, TQPNet combines a multilayer perceptron with ReLU activations (MLP(ReLU)) and a small number of unrolled ADMM(ReLU) refinement layers. The network is trained offline using an optimality-condition-informed residual loss, eliminating the need for solver-generated labels. The proposed residual-based framework enables both self-supervised training and online solution certification through adaptive ADMM(ReLU) refinement. Experiments on large-scale mpQPs and real-time MPC applications demonstrate that TQPNet achieves high accuracy, strong generalization, and substantially faster inference than conventional optimization-based methods.
Neural Continuous-Time Markov Chain: Discrete Diffusion via Decoupled Jump Timing and Direction
Jingyuan Li ⋅ Xiaoyi Jiang ⋅ Fukang Wen ⋅ WEI LIU ⋅ Renqian Luo ⋅ Yi Zhu ⋅ Zuoqiang Shi ⋅ Pipi Hu
Discrete diffusion models based on continuous-time Markov chains (CTMCs) have shown strong performance on language and discrete data generation, yet existing approaches typically parameterize the reverse rate matrix monolithically---through proxies such as concrete scores (SEDD) or clean-data predictions (MDLM, GIDD)---rather than aligning the parameterization with the intrinsic CTMC decomposition into jump timing and jump direction. We propose \textbf{Neural CTMC}, which exploits the underlying Poisson structure of CTMC dynamics by separately parameterizing the reverse process through an \emph{exit rate} (when to jump) and a \emph{jump distribution} (where to jump) via two dedicated network heads. We show that the evidence lower bound (ELBO) reduces to a path-space KL divergence between the true and learned reverse processes that factorizes into a Poisson KL for timing and a categorical KL for direction, and admits a tractable, gradient-equivalent and consistent loss. Experimentally, scored by Gemma2-9B, our pure-uniform Neural CTMC achieves $16.36$ generative perplexity on TinyStories (vs.\ GIDD $37.60$ and MDLM $42.66$). On OpenWebText, it attains the best perplexity at the same training-token budget across 16--128 sampling steps among the methods we compare (e.g., at 128 steps: Neural CTMC $183.6$ vs.\ MDLM $210.5$ and GIDD $249.8$).
Neural Proposals, Symbolic Guarantees: Neuro-Symbolic Graph Generative Modeling
Chuqin Geng ⋅ Li Zhang ⋅ Mark Zhang ⋅ Zhaoyue(Rebecca) Wang ⋅ Haolin Ye ⋅ Xujie Si
While deep generative models excel at capturing graph data distributions, they struggle to satisfy complex, hard constraints. In unconstrained settings, these models typically produce valid topologies; yet imposing strict compositional rules, like those in drug discovery, creates an out-of-distribution (OOD) setting where purely neural methods frequently fail. Because these neural approaches rely on soft conditioning and post-hoc filtering on such tasks, they cannot provide the formal guarantees needed for high-stakes domains. To address this, we introduce Neuro-Symbolic Graph Generative Modeling (NSGGM), a framework built on the principle of Neural Proposals, Symbolic Guarantees. NSGGM decouples generation: an autoregressive model proposes structural scaffolds, and a Satisfiability Modulo Theories (SMT) solver handles the final discrete assembly of the proposed substructures. Empirically, NSGGM is competitive with state-of-the-art methods on unconstrained tasks. To evaluate logical-constraint satisfaction inspired by real drug discovery workflows, we introduce MolSAT, a benchmark for hard compositional rules. On MolSAT, purely neural baselines completely fail OOD (0% satisfaction with zero training support), while NSGGM achieves >95% satisfaction in-distribution and 64–86% with zero training support.
Next Forcing: Causal World Modeling with Multi-Chunk Prediction
Gangwei Xu ⋅ Qihang Zhang ⋅ Jiaming Zhou ⋅ Xing Zhu ⋅ Yujun Shen ⋅ Xin Yang ⋅ Yinghao Xu
Autoregressive video generation has emerged as a powerful paradigm for World Action Models (WAMs). However, existing approaches suffer from slow training convergence particularly at high frame rates, limited converged accuracy, and slow inference due to iterative video denoising, as the training supervision is confined to the current chunk without explicit signals about future dynamics. In this paper, we present Next Forcing, a multi-chunk prediction (MCP) framework for causal world modeling that enables faster training, higher accuracy, and accelerated inference. Inspired by multi-token prediction in large language models, Next Forcing introduces an MCP training objective that augments the main model with lightweight auxiliary MCP modules to simultaneously denoise video chunks at multiple future temporal horizons (next$^1$, next$^2$, next$^3$ chunks). These MCP modules form a causal chain across prediction depths, where intermediate features fused from multiple layers of the main model are leveraged to predict future dynamics, allowing near-future predictions to inform farther-future ones and providing dense multi-scale temporal supervision back to the main model. During training, the MCP modules significantly accelerate convergence and improve converged accuracy, especially at high frame rates. At 50 fps, our method achieves a 93.1\% relative improvement over the baseline LingBot-VA at 5k training steps and achieves 2.3$\times$ faster convergence. At inference, the MCP modules can be retained to predict the next video chunk in parallel with the current one, accelerating generation. Next Forcing establishes new state-of-the-art results on the RoboTwin benchmark (94.1\%/93.5\% on Clean/Random) and demonstrates significant improvements on PhyWorld, a benchmark evaluating adherence to physical laws in video generation.
Nexus: Same Pretraining Loss, Better Downstream Generalization via Common Minima
Huanran Chen ⋅ Huaqing Zhang ⋅ Xiao Li ⋅ Yinpeng Dong ⋅ Ke Shen ⋅ Jun Zhu
The foundational capabilities of large language models are acquired during pretraining on internet-scale, highly heterogeneous data mixtures. In this work, we investigate a geometric question regarding the converged state of this pretraining process: Does the model converge to a common minimizer across all data sources (e.g., \cref{fig:cwaillustration:close}), or merely a minimizer of the averaged loss (e.g., \cref{fig:cwaillustration:distant})? We hypothesize that the geometric ``closeness'' of task-specific minima is intrinsically linked to downstream generalization. However, we reveal that standard optimizers (e.g., AdamW) often converge to points where task-specific minima are distant from each other. To address this, we propose the \textit{Nexus} optimizer, which encourages the closeness of these minima by maximizing gradient similarity during optimization. Extensive experiments across models ranging from 130M to 3B parameters demonstrate that Nexus significantly boosts downstream performance, despite achieving nearly \textit{the same pretraining loss} (see \cref{fig:demo:benchmark}). Notably, on the 3B model, Nexus reduces the out-of-distribution loss by 0.012 and yields up to a 15.0\% accuracy improvement on complex reasoning tasks (e.g., GSM8k). Our findings challenge the reliance on pretraining loss as the sole proxy for model evaluation and highlight the critical role of optimizer implicit bias in unlocking downstream generalization.
Variable fonts enable continuous variation of glyph geometry along semantic design axes such as weight, width, slant, and optical size. However, constructing a variable font from a static font remains a labor-intensive process requiring expert typographic design and manual specification of glyph variation data. We introduce NIV (Neural Axis Variations), a method that automatically converts a static font into a fully functional variable font. Given glyph outlines and a set of desired design axes, NIV predicts per-point displacements. The model operates directly on vector glyph geometry and employs a novel Property Embedding mechanism that captures interactions between multiple axes, enabling consistent multi-axis variation within a unified framework. We train NIV on a newly constructed dataset derived from variable Google Fonts, comprising over one million variation tuples. The resulting model generalizes across unseen code points, unseen font styles, high-complexity CJK glyphs, and even out-of-distribution handwriting inputs. The generated outputs are standard variable font files supporting continuous interpolation via existing rendering engines. To facilitate research, we release the dataset, the complete training and inference implementation, and trained models. Beyond typography, our approach demonstrates how structured geometric objects with continuous parametric variation can be synthesized using neural deformations.
NOFE – Neural Operator Function Embedding
Lars Uebbing ⋅ Harald Lykke Joakimsen ⋅ Siyan Chen ⋅ Georgios Leontidis ⋅ Kristoffer Wickstrøm ⋅ Michael Kampffmeyer ⋅ Sébastien Lefèvre ⋅ Arnt B Salberg ⋅ Robert Jenssen
Most dimensionality reduction methods treat data as discrete point clouds, ignoring the continuous domain structure inherent to many real-world processes. To bridge this gap, we introduce Neural Operator Function Embedding (NOFE), a domain-aware framework for continuous dimensionality reduction. NOFE learns function-to-function mappings via a Graph Kernel Operator, enabling mesh-free evaluation at arbitrary query locations independent of input discretization. We establish NOFE as approximation of sheaf-to-sheaf mappings, generalizing Sheaf Neural Networks to continuous domains. We evaluate NOFE across different datasets, comparing it against PCA, t-SNE, and UMAP. Our results demonstrate that NOFE significantly outperforms baselines in local structure preservation, achieving a local Stress of 0.111 compared to 0.398 for PCA, 0.773 for t-SNE, and 0.791 for UMAP for the ERA5 climate reanalysis dataset. NOFE also exhibits robust sampling independence, reducing the Patch Stitching Error by up to $20.0\times$ relative to UMAP (59.0 vs. 267.6 under regional normalization) and ensuring consistency across disjoint domain patches. While maintaining competitive global structure preservation (Stress-1: 0.379 vs. PCA's 0.268), NOFE resolves fine-grained structures and produces smooth, consistent embeddings that generalize across varying sample densities, addressing key limitations of discrete reduction methods.
Not All Features Are Created Equal: A Mechanistic Study of Vision-Language-Action Models
Bryce Grant ⋅ Xijia Zhao ⋅ Peng Wang
A fine-tuned Vision-Language-Action (VLA) policy will pick up the alphabet soup and place it in the basket on demand, then drop the soup off the table when an evaluator shifts the basket five centimeters left. We use activation injection to ask what the policy is actually doing: inject task A's activations into task B's scene at the action expert, and π₀.₅ executes task A's motor trajectory in 99.6% of episodes (n=1,968); X-VLA does so in 99.8%. The injected program is bound to absolute workspace coordinates rather than the visible scene, which mechanistically explains the perturbation brittleness reported in concurrent benchmark work. The same intervention framework applied across six VLAs (π₀.₅, OpenVLA-OFT, X-VLA, SmolVLA, GR00T N1.5, and ACT as a language-free control) on LIBERO, MetaWorld, SimplerEnv, and ALOHA over 420,000+ rollouts surfaces three further findings: language is encoded by every architecture yet behaviorally ignored when vision identifies the goal; SAE pooling preference splits along architecture lines (π₀.₅ per-token, X-VLA mean-pool, SmolVLA indifferent); and pathway specialization replicates wherever expert and VLM pathways are separable. SmolVLA's interleaved fusion attenuates the headline to 52.1% LIBERO override, scoping universality. We will release 424 trained SAEs and an interactive feature-exploration platform on acceptance; an anonymized snapshot is included in the supplementary material.
Not All Layers Need Tuning: Selective Layer Restoration Recovers Diversity
Bowen Zhang ⋅ Meiyi Wang ⋅ Harold Soh
Post-training improves instruction-following and helpfulness of large language models (LLMs) but often reduces generation diversity, which leads to repetitive outputs in open-ended settings, a phenomenon known as mode collapse. Motivated by evidence that LLM layers play distinct functional roles, we hypothesize that post-training-induced diversity loss is unevenly distributed across layers and that restoring a carefully chosen range of layers to their pre-trained weights can recover diversity while maintaining high output quality. To operationalize this idea, we design a proxy task---Constrained Random Character (CRC)---with an explicit validity set and a natural diversity objective. Results on CRC reveal a clear diversity–validity trade-off across restoration ranges and identify configurations that increase diversity with minimal quality loss. Based on these findings, we propose Selective Layer Restoration (SLR), a training-free method that restores selected layers in a post-trained model to their pre-trained weights, yielding a hybrid model with the same architecture and parameter count, incurring no additional inference cost. Across three different tasks (creative writing, open-ended question answering, and multi-step reasoning) and three different model families (Llama, Qwen, and Gemma), we find SLR can consistently and substantially improve output diversity while maintaining high output quality.
Not Too Generative, Not Too Discriminative: The Human Alignment Sweet Spot
Jorge Chang Ortega ⋅ Bastien Le Lan ⋅ Thomas Serre ⋅ Victor Boutin
A central question in computational vision is whether human-like visual representations are better explained by discriminative or generative learning. Existing comparisons, however, often confound the learning objective with architecture, scale, and training data, leaving open whether the objective itself drives alignment. We address this confound using Joint Energy-Based Models (JEMs), which interpolate continuously between discriminative and generative training within a fixed architecture. By varying a single mixing coefficient, we isolate the effect of the learning objective and evaluate the resulting models across six human-alignment benchmarks spanning perceptual similarity, gloss perception, human response uncertainty, robustness, shape–texture cue conflict, and diagnostic feature attribution. Across this diverse suite, human alignment is consistently maximized at intermediate points of the generative–discriminative continuum, rather than at either endpoint. Hybrid JEMs combine the categorical structure induced by discriminative learning with the sensitivity to input structure induced by generative learning, yielding more human-like behavior across multiple levels of vision. These results suggest that the generative–discriminative dichotomy is the wrong axis for understanding human-aligned vision: alignment emerges not from choosing one objective over the other, but from balancing both.
OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation
Xinzhe Chen ⋅ Sihua Ren ⋅ Liqi Huang ⋅ Haowen Sun ⋅ Mingyang Li ⋅ Xingyu Chen ⋅ Zeyang Liu ⋅ Xuguang Lan
Recent vision-language-action (VLA) models and world action models (WAMs) advance robotic manipulation by enriching intermediate representations with auxiliary spatial features or future visual-state prediction. However, these representations largely remain within the observation space and do not share the rigid-body geometry of the action space, forcing the action decoder to implicitly recover this geometry. We propose OASIS, a visuomotor policy that aligns the intermediate representation with the action space via $SE(3)$ end-effector trajectory prediction. OASIS couples a 3D-aware feature encoder that fuses vision-language and metric-depth features with an $SE(3)$ trajectory predictor that produces a camera-frame end-effector trajectory. Conditioned on the predictor's pose-supervised hidden states, the action decoder generates action chunks consistent with rigid-body motion. Across simulation and real-world experiments, OASIS outperforms VLA and WAM baselines in success rate and out-of-distribution generalization.
OfficeInstruct: A Dataset for Tool-Augmented Office Agents on Long-Trajectory Generative Tasks
Weikai Xu ⋅ Zhizheng Jiang ⋅ Lushuo Jiang ⋅ Kun Huang ⋅ Yuxuan Liu ⋅ Pengzhi Gao ⋅ Wei Liu ⋅ Jian Luan ⋅ Xiaolin Hu ⋅ Bo An
The recent trend of integrating autonomous agents into operating systems, supported by the Model Context Protocol (MCP) tools and Deep Searching, has driven research on efficient assistants for daily office tasks like document writing, spreadsheet analysis, and presentation generation. However, training these agents still faces several challenges: the overemphasis on question-answering tasks at the expense of multi-turn generative tasks; the absence of a training paradigm that jointly enhances coding and tool-calling capabilities; the difficulty of evaluating tool efficiency and artifact quality in the presence of drifting user intents. To address these problems, we deploy a unified multi-turn agent framework to collect real-world user interactions, creating \textit{OfficeInstruct}: a comprehensive dataset of 199k high-quality trajectories across Doc, Spreadsheet, and PPT, with an average of 56.2 steps. To ensure data quality, we propose a graph-based evidence evaluation method that models interaction histories as directed graphs, filtering out redundant tool cycles, execution errors, and hallucinations. Experiments demonstrate that our \textit{OfficeAgent} 8B and 14B models surpass open-source baselines on in-domain and out-of-distribution benchmarks, with the 32B variant attaining a 96.19\% success rate on the newly proposed \textit{OfficeAgentBench}, surpassing proprietary models like Claude-4.5-Sonnet and Gemini-3-Pro.
Offline Constrained Reinforcement Learning under Partial Data Coverage
Seokmin Ko ⋅ Ambuj Tewari ⋅ Kihyuk Hong
We study offline constrained reinforcement learning with general function approximation in discounted constrained Markov decision processes. Existing methods either require full data coverage for evaluating unsupported intermediate policies, are not oracle efficient, or requires the knowledge of data-generating distribution for policy extraction. We propose PDOCRL, an oracle-efficient primal-dual algorithm based on a decomposed linear-programming formulation. The decomposition makes the policy an explicit optimization variable, avoiding policy extraction through the unknown distribution $\mu_D$. We show that naive restricted saddle-point formulations may have spurious saddle points, so realizability of an optimal solution alone is insufficient. We then show that, under a stronger but explicit realizability assumption, every restricted saddle point is optimal, avoiding the regularization and auxiliary function classes used in prior LP-based analyses. These ingredients yield PDOCRL, an oracle-efficient primal-dual algorithm that computes a near saddle point of the empirical decomposed LP Lagrangian and returns a near-optimal, near-feasible policy with a $\widetilde{\mathcal O}(\epsilon^{-2})$ sample guarantee under partial coverage, without access to $\mu_D$ or a reference policy. Empirically, PDOCRL is competitive with strong baselines on standard offline constrained RL benchmarks.
Offloading Score: Measuring AI Reliance through Counterfactual Workflows
Vishakh Padmakumar ⋅ Lujain Ibrahim ⋅ Zora Wang ⋅ Jennifer Wang ⋅ Q.Vera Liao ⋅ Diyi Yang
AI tools are increasingly integrated into real-world workflows. However, existing measures of reliance on these tools focus on AI output adoption or on self-reported indicators, rather than how task effort is distributed between users and tools. Here, we introduce offloading score, a measure of reliance that quantifies the fraction of cognitive effort offloaded to an AI tool. Offloading score is simulation-based---we construct a counterfactual workflow by estimating how the user would have completed the task without the tool, and then computing the fraction of steps saved by using the tool. We validate offloading score through intrinsic evaluations of metric validity, and a controlled user study ($n=40$) with developers performing programming tasks using AI tools. We vary time pressure to test whether reliance measures capture the known increase in reliance under time pressure. We show that offloading score detects significantly higher reliance in time-constrained settings ($+43\%$, $p=0.018$), while usage-based and self-reported baseline measures do not distinguish the conditions. We complement this with descriptive insights showing that higher reliance manifests as greater delegation and more direct reuse of AI outputs. Finally, we demonstrate an approach of using offloading score in combination with target outcomes of reliance (e.g., task understanding) to identify when reliance may be (in)appropriate. Our framework offers two contributions: an instrument users can apply to measure and reflect on their own reliance, and a quantitative signal that agent designers can utilize to mitigate overreliance.
OmniMemBench: Towards Scalable Evaluation of Long-Term Omni-Modal Agent Memory
Junhan Shi ⋅ Qiyi Wang ⋅ Xiongwei Wu ⋅ Ming Ma ⋅ Baoqi Pei ⋅ Zhenning Shi ⋅ Yong Jiang ⋅ Qing Li ⋅ Weigao Sun ⋅ Yiran Zhong ⋅ Steven Hoi
Multimodal agents that interact with users over days or months must remember not only what was said, but what was shown, heard, and played. Current memory benchmarks are predominantly text-only or cover at most two modalities. Those that do incorporate multimodal content still evaluate with aggregate scores that conflate retrieval and reasoning errors, and provide no mechanism to test whether performance degrades as context scales, leaving our understanding of agent memory fundamentally incomplete. We introduce OmniMemBench, to our knowledge the first long-term memory benchmark that spans image, audio, and video in multi-session dialogue. Its design is guided by three requirements for reliable evaluation. First, an anti-leakage mechanism filters questions answerable from text alone, reducing text-only solvability to a negligible level. Second, two LLM-judged retrieval metrics—Clue Coverage and Entry Relevance—operate on content rather than entry IDs, enabling unified retrieval diagnosis across RAG and memory paradigms. The gap between retrieval coverage and downstream QA scores further localizes failures to the memorization stage. Third, content-isolated distractor injection scales context from 128K to 1M tokens while holding evidence constant, so scalability can be measured without confounding difficulty. The benchmark covers 103 characters and 6,655 QAs across 8 task types and 4 context tiers. Evaluating representative long-context, RAG, and structured-memory methods, we find that no single paradigm dominates: long-context models lead overall, yet RAG surpasses them on targeted fact extraction. Further analysis shows that memory methods retrieve clues at coverage rates comparable to RAG but score substantially lower on QA, revealing that the bottleneck is not what is retrieved but what is lost during memorization. Caption-based processing irreversibly discards perceptual details; raw embeddings degrade retrieval accuracy; and scaling the memorization model yields negligible gains.
OmniToM: Benchmarking Theory of Mind in LLMs via Explicit Belief Modeling
Adam Bawatneh ⋅ Sagar Sapkota ⋅ Amrit Singh Bedi ⋅ Santu Karmaker ⋅ Mubarak Shah
Theory of Mind (ToM), the ability to infer others’ knowledge, intentions, and emotions, is commonly evaluated in large language models (LLMs) using end-point question answering, where performance is judged solely by the final answer to a social reasoning query. This paradigm obscures whether the model actually constructs the underlying mental-state representations required for robust reasoning, particularly in scenarios involving divergent, evolving, or mistaken beliefs. In order to address this research gap, we introduce OmniToM, a benchmark that directly evaluates these representations by requiring explicit modeling of belief structures for all relevant actors within a narrative. These structures are composed of belief propositions: minimal statements of what an actor takes to be true about the world or another actor's mental state, allowing knowledge, intentions, emotions, and false beliefs to be analyzed in a common format. Models are evaluated in two stages: Stage 1: Belief Extraction, which extracts from the story the beliefs relevant to its social dynamics, and Stage 2: Belief Labeling, which assigns each belief a seven-dimensional schema label covering recursive order, truth status, knowledge access, explicitness, content type, mental source, and context. Built from 895 stories from the existing ToMBench story corpus and augmented with 22,343 labeled belief propositions, OmniToM uses a human-calibrated LLM-assisted annotation pipeline. Across diverse models in zero-shot evaluation, OmniToM reveals an actor-specific belief-tracking bottleneck: current LLMs struggle with the knowledge-access and representational decisions required to transform narrative facts into actors’ beliefs and shared mental states.
On approximation and estimation of Schrödinger potentials without the curse of dimensionality
Artsiom Patarusau ⋅ Nikita Puchkin ⋅ Konstantin Yakovlev
We examine generative modelling approaches based on the construction of Schrödinger bridges between Gaussian noise and a target distribution. It is known that the solution of the dynamic Schrödinger problem is a diffusion process with a drift associated with Doob's h-transform of a Schrödinger potential. Although its accurate restoration from finite samples is crucial for reliable, high-quality data generation, the existing literature lacks theoretical guarantees regarding this question. In our work, we establish theoretical upper bounds on the complexity of Schrödinger potential approximation and estimation via neural networks. These bounds are determined by the effective dimension of the target distribution. To our knowledge, this is the first result demonstrating that generative modelling methods based on Schrödinger bridges and stochastic optimal control can escape the curse of dimensionality.
This paper studies concentration inequalities for sampling without replacement. We introduce a general inequality that provides simple guidelines for proving concentration inequalities, offers a degree of unification of existing results, and is strong enough to yield sharper bounds. As applications, we revisit several standard statistics and either recover classical inequalities through streamlined proofs or strengthen them to tighter bounds, with rates optimal up to constants.
One Avatar, Any Budget: Robust LoD for Dynamic Gaussian Avatars
Haoyu Zhao ⋅ Wenjuan Gong ⋅ Junang Qi ⋅ Haiyu Zhang ⋅ Wenjian Zhao ⋅ Zimu YOU
Cross-device deployment of a single reconstructed avatar often requires adaptation to varying Gaussian budgets with minimal degradation in visual quality. Motivated by the need for a robust budget-adaptive 3D representation, we propose RobustAvatar, a train-once Level-of-Detail (LoD) adaptation framework for dynamic Gaussian avatars based on sequential selection and compensation. RobustAvatar supports flexible compression-ratio control within a single trained avatar, without requiring fine-tuning at each compression ratio. The framework consists of two coordinated modules: Semantic Distribution-Regularized Selection (SDS) and Canonical Budget-Conditioned Compensation (CBC). SDS performs budget-aware Gaussian selection guided by semantic distribution regularization, preserving the coverage of perceptually salient facial regions relative to the full avatar under constrained budgets. CBC refines the selected primitives in a canonical, expression- and pose-neutral space by predicting ratio-conditioned attribute residuals, reducing detail loss from Gaussian removal and improving robustness to decreasing Gaussian budgets. Experiments on the NeRSemble dataset show that RobustAvatar achieves more robust budget adaptation than state-of-the-art methods, with slower quality degradation as Gaussian budgets decrease. The code will be made publicly available.
OneCanvas: 3D Scene Understanding via Panoramic Reprojection
Bartłomiej Baranowski ⋅ Dave Chen ⋅ Matthias Niessner
Existing approaches to 3D scene understanding in Vision-Language Models (VLMs) either rely on complex, model-specific geometry encoders or large training budgets in pursuit of spatial reasoning. Instead, OneCanvas aggregates patch features from all views onto a single equirectangular panoramic canvas. Namely, each patch is unprojected to a 3D world coordinate using its depth and camera pose, then placed on the canvas at the continuous longitude and latitude of that point as seen from the canvas origin, with no rasterization or aggregation across overlapping views. A 3D position embedding of the patch's metric coordinates is added to its feature, restoring the depth lost when collapsing the world position to an angular canvas coordinate. Patches from all frames thus share one spatial coordinate system with no fusion or major architectural modifications of the backbone. The pretrained VLM consumes this representation as if it were an ordinary image. Because the canvas can be centered on any pose of interest, the same representation directly supports situated reasoning from a specific viewpoint, a common requirement in robotics and embodied AI. Thanks to this representation, we can also introduce a spatial pretraining curriculum: by procedurally placing patch features of objects, drawn from real images, at chosen 3D world positions on an otherwise empty canvas, we generate on-the-fly supervision spanning a broad range of spatial reasoning tasks, with answer distributions controlled to reduce spatial reasoning shortcuts. OneCanvas achieves state-of-the-art accuracy on SQA3D and VSI-Bench, and generalizes to out-of-distribution data on SPBench, using an order of magnitude less training compute than the strongest competing methods.
Universal visual anomaly detection (AD) aims to identify anomalous images and segment anomalous regions towards open and dynamic scenarios, typically adhering to zero- and few-shot paradigms without dataset-specific fine-tuning. Recently, the field has seen significant progress driven by the integration of vision-language foundation models. However, we observe that current methods often struggle with laborious prompt engineering, elaborate adaptation modules, and complex training strategies that ultimately constrain their flexibility and generalizability. In this paper, we rethink the fundamental mechanism of vision-language models for AD and present UniADet, an embarrassingly simple, effective and general framework for Universal vision Anomaly Detection. Our approach is built on two key insights: First, we reveal that the primary function of the language encoder is merely to derive decision weights and we demonstrate that it is unnecessary, as these weights can be learned more directly and efficiently. Second, to resolve learning conflicts arising from disparate feature manifolds, we introduce a systematic decoupling strategy that learns independent weights across both distinct tasks (classification vs. segmentation) and hierarchical features. UniADet is highly simple and efficient, learning only decoupled weights and requiring 0.02M learnable parameters. It is inherently versatile, adapting seamlessly to various foundation models (\eg, CLIP, DINOv2, and DINOv3). Extensive evaluations of 14 real-world benchmarks in the industrial and medical domains demonstrate that UniADet not only exceeds state-of-the-art zero/few-shot methods by a substantial margin, but also outperforms full-shot AD methods for the first time. This empirical evidence reveals the profound potential of language-free frameworks to redefine the boundaries of visual anomaly detection. The code and models will be made publicly available.
One-Layer Transformers Provably Learn In-Context K-Nearest Neighbor Prediction with Chain-of-Thought
Lyumin Wu ⋅ Yuan Cao
Chain-of-Thought (CoT) enables transformers to solve complex reasoning tasks by generating intermediate reasoning steps. Despite strong empirical success, the mechanisms by which such reasoning abilities emerge during training remain poorly understood, particularly from a theoretical perspective. In this work, we study this question in a stylized yet fully analyzable setting: in-context $\mathrm{K}$-nearest-neighbor ($\mathrm{K}$-NN) prediction with a one-layer transformer. Earlier work has demonstrated that one-layer transformers can be trained to perform $1$-NN without CoT \citep{Li2024One} by having the softmax attention attend to the nearest neighbor in context. However, this result on $1$-NN cannot be extended to $\mathrm{K}$-NN for $K>1$, as softmax attention is much less naturally suited to attending to the 2nd through $K$-th nearest neighbors. We first give a rigorous negative result showing that this limitation already appears on a well-separated class of binary classification tasks for odd $K>1$. We then give an explicit construction of a one-layer transformer, and show that it can solve $\mathrm{K}$-NN via CoT reasoning. We further prove that, under a stylized training setup, gradient descent can recover this construction, thereby showing that transformers can acquire the capability to solve $\mathrm{K}$-NN through training. In addition, we show that the resulting trained model applies to a broader test class than that assumed in the training analysis. Our results shed light on how CoT expands the ability of transformers to learn and solve tasks that are otherwise hard to solve.
One Unified Representation: Resolving the Appearance-Semantics Dilemma via Structural Regularization
RUIQI YANG ⋅ Fei Zhou ⋅ Yi Zhang ⋅ Liang Li ⋅ Xiangyu Yue
The Representation Autoencoder (RAE) shows that a frozen high-dimensional semantic ViT representation enables both image reconstruction and DiT-based generation. However, RAE's semantic-only representation lacks appearance details, yielding suboptimal reconstruction. Full fine-tuning the encoder with reconstruction loss induces catastrophic semantic collapse: the encoder overfits to appearance, loses semantic discriminability, and cripples generation. This failure arises because the reconstruction loss imposes an appearance prior that erodes semantic structure. To resolve this conflict, we propose Semantic Structure Regularization (SSR). SSR builds on self-distillation, using a learnable student and a frozen teacher that retains the original semantics. In addition to standard latent alignment, which forces the student's representation to match the teacher's, we introduce two complementary regularizers that prevent appearance overfitting from corrupting semantic structure. First, spatial self-similarity distillation forces the student to reproduce the teacher's patch-wise relational patterns. Invariant to appearance changes, this relational signature anchors high-order semantics and prevents the sacrifice of structure for pixel fidelity. Second, structure-anchored alignment decouples structure from appearance. Using a frequency-domain transform, we extract a high-frequency structural map (edges, boundaries) from each image while discarding low-frequency appearance. The teacher's representation of this structural map then regularizes the student's representation of the original image. Consequently, the student learns appearance details from the reconstruction loss without penalty, where the regularization enforces only structural consistency and enables pixel-level encoding without semantic drift. The result is a unified representation that excels at high-fidelity reconstruction, discriminative perception, and efficient DiT-based generation. Our work provides a principled path toward a single semantic space bridging perception and generation, challenging the classic dichotomy. Code and models will be released.
Online Allocation with Unknown Shared Supply
Tzeh Y Neoh ⋅ XianJun, Davin Choo ⋅ Mengchu Yue ⋅ Milind Tambe
Many real-world resource allocation systems, such as humanitarian logistics and vaccine distribution, must preposition limited supply across multiple locations *before* demand is realized while stockouts incur irreversible service losses. To study this, we introduce the *Online Shared Supply Allocation* (OSSA) problem, a stateful online model in which a central hub allocates a finite, unknown supply to multiple sites facing sequential demand under fixed-charge transportation costs and lost-sales penalties. Unlike classical make-to-stock or make-to-order inventory models, OSSA precludes backlogging and replenishment only hedges against *future* demand. To tackle OSSA, we propose a deterministic threshold-proportional policy GPA and prove that it achieves a $4/3$-approximation to the offline optimum up to an additive term independent of the total supply. We complement this with matching lower bounds showing that the $4/3$ ratio is tight and that the additive-error dependence is unavoidable, even for randomized algorithms that know the total supply upfront. Finally, we develop a learning-augmented extension to GPA that principally incorporates imperfect forecasts (e.g., from human experts or ML models) commonly available in practice, enabling us to exploit high-quality advice while being robust against arbitrary bad ones. Synthetic and real-world experiments show that GPA outperforms natural baselines with global supply is scarce.
We consider a variant of the partially observed online control problem where there are _multiple sensors_ which provide noisy linear measurements of the system. In our setting, the system state evolves according to a discrete-time linear dynamical system. The objective is to design an algorithm which minimizes regret relative to the best _single-sensor_ policy (within a restricted class) in hindsight. When the system and process noise is stochastic and the costs are quadratic in the control and _unobserved state_, we design an algorithm that achieves $\mathcal{O}(\mathrm{poly}\log(S)\sqrt{n})$ regret over time horizon $n$ relative to our benchmark policy class if there are $S$ sensors. Notably, this regret guarantee is with respect to the _state costs_, which are never fully observed. When the system and process noise is chosen by an oblivious adversary, and the cost functions are convex-Lipschitz (or subquadratic) functions of the control and _observations_, then a similar $\mathcal{O}(\mathrm{poly}\log(S)\sqrt{n})$ regret is achievable. While our work is the first, to our knowledge, to study this generalization of the online control problem, we remark that a na\"ive application of known results in the partially observed online control literature have regret scaling $\mathcal{O}(\mathrm{poly}(S)\sqrt{n})$, exponentially worse than our regret scaling in terms of dependence on the number of sensors. Our algorithm is based on the Follow-The-Regularized-Leader framework, and our analysis carefully exploits the $\ell_1$ geometry of the policy class.
Conformal prediction is a framework that provides valid uncertainty quantification for general models with exchangeable data. However, in the online learning and time-series settings, exchangeability is not satisfied. Existing online conformal methods, such as adaptive conformal inference (ACI), can achieve long-run validity, yet they remain inefficient under covariate heterogeneity because they rely on global calibration. We propose Online Localized Conformal Prediction (OLCP), which combines online adaptation with covariate-dependent localization to better reflect heterogeneity. To reduce sensitivity to the localization bandwidth, we further develop OLCP-Hedge, which performs bandwidth selection as an online expert aggregation problem using a constrained online convex optimization framework. Importantly, we provide coverage guarantees for both algorithms and demonstrate through simulations and real-data experiments that the proposed methods attain valid long-run coverage with narrower prediction intervals than existing baselines.
On Nash Equilibria in Participatory Budgeting with Donations and Beyond
GRZEGORZ LISOWSKI ⋅ Georgios Papasotiropoulos ⋅ Grzegorz Pierczyński ⋅ Krzysztof Rogowski
This work proposes a framework for public decision processes with a monetary component, where voters can pledge donations to support preferred projects and influence the allocation of public funds. It captures a range of real-world scenarios, including participatory budgeting, blockchain-based funding platforms, charitable programs, and crowdfunding campaigns. Through both theoretical and experimental evaluation, our study addresses questions concerning the existence, structure, number, computation, and quality of Nash equilibria.
On the Complexity of Discounted Robust MDPs with $L_p$ Uncertainty Sets
Ali Asadi ⋅ Krishnendu Chatterjee ⋅ Alipasha Montaseri ⋅ Ali Shafiee
A basic model in sequential decision making is the Markov decision process (MDP), which is extended to Robust MDPs (RMDPs) by allowing uncertainty in transition probabilities and optimizing against the worst-case transition probabilities from the uncertainty sets. The class of $(s,a)$-rectangular RMDPs with $L_p$ uncertainty sets provides a flexible and expressive model for such problems. We study this class of RMDPs with discounted-sum objectives and a constant discount factor. The existence of an efficient algorithm for this class is a fundamental theoretical question in optimization and sequential decision making. Previous results only establish a strongly polynomial-time algorithm for $L_\infty$ uncertainty sets. In this work, our main results are as follows: (a) we show that for any compact uncertainty set, the policy iteration algorithm for RMDPs is strongly polynomial with oracle access to solutions of Robust Markov chains (RMCs); (b) we present strongly polynomial-time bounds on the policy iteration algorithm for RMCs with $L_1$ and $L_\infty$ uncertainty sets; and (c) we establish hardness results for RMCs with $L_p$ uncertainty sets for integer $p$ satisfying $1
On the Necessity of Guidance Decay: From Three-Phase Analysis in Gaussian Mixture Models to Dynamic Optimization
Yiyu Qiu ⋅ Ruofeng Yang ⋅ Tong Yu ⋅ Shuai Li
Classifier-Free Guidance (CFG) is essential for conditional diffusion, yet constant guidance scales often fail to balance alignment with over-exposure and extreme sample issues. Existing heuristic schedules, like guidance truncation, lack rigorous principles and sacrifice classification confidence, while theoretical studies remain limited to constant guidance at final denoising stages, leaving the full sampling trajectory largely unanalyzed. In this work, we provide a theoretical foundation for dynamic guidance by analyzing Gaussian Mixture Models (GMM) within a general SDE framework. We characterize the dynamics of classification confidence through three distinct stages: convergence, mixed, and divergence, and rigorously prove the impact of varying guidance scale $\omega$ at reverse time $k$. We analyze the Convergence Boundary $\Omega(k)$ as the maximum guidance scale that ensures the reverse trajectory remains within the high-density regions of the data distribution. We prove that during denoising, the permissible guidance scale $\omega$ must decay to prevent the trajectory from escaping into the divergent phase, especially at the end of the denoising. Although different noise schedules have different decay rates, we found that VP(Variance Preserving) has more stringent decay rate requirement compared to VE(Variance Exploding), which is why VP is most prone to overexposure. Specifically, $\Omega(k)$ simplifies to an exponential decay $O(e^{T-k})$ in VPSDE and an algebraic decay $O(T-k)$ in VESDE. This provides a unified explanation for why constant guidance inevitably triggers artifacts and demonstrates that guidance truncation is essentially a coarse approximations of this stability limit. We can achieve dynamic optimization based on the decay rate, which better restricts the scale within the convergence region, thereby effectively eliminating overexposure while maintaining excellent semantic alignment and improving generation quality.
In decoder-only (causal) transformers, the computation graph created by causal masking routes information through both direct-path attention and indirect paths formed by intermediate tokens. We denote these indirect paths between token pairs as their \textit{runways}. We argue that certain failure modes of causal transformers as observed by a growing body of recent works are likely exacerbated by a misalignment between these two information propagation modes. We formalize \textit{runway cascade} as a phenomenon whereby this misalignment results in redundancies and irrelevant information cascading to token representations despite adequately learned attention patterns. As a solution, we propose \textit{runway-aware rewiring} as a more explicit way of incorporating runway context directly into each token's direct-path attention. This mechanism re-wires the attention pattern for each token based on a summary of its runway landscape, enabling awareness of accumulating representational influences and allowing for more balanced information propagation. Our proposed methodology introduces no additional parameters and can seamlessly be integrated into standard attention mechanism. Empirically, our rewired transformer results in steady improvements in general language modeling as well as noticeably stronger information retrieval and extrapolation abilities compared to standard transformers.
OpenCoF: Learning to Reason Through Video Generation
Xinyan Chen ⋅ Renrui Zhang ⋅ Ziyu Guo ⋅ Dongzhi JIANG ⋅ Hongsheng Li
Reasoning has become a core capability for large models, especially when reliable decisions require understanding logical consequences. Recent video generation models offer a reasoning path distinct from previous Chain-of-Thought (CoT): reasoning can unfold through temporally connected frames, known as Chain-of-Frame (CoF) reasoning. However, existing video generators are primarily trained on general video corpora, still lacking diverse supervision and dedicated designs for CoF reasoning. To address this gap, we introduce OpenCoF, a framework comprising the OpenCoF-17K Dataset, a reasoning video dataset spanning 11 task families, and Wan-CoF, a fine-tuned video model for studying whether diverse temporal supervision improves CoF behavior. Across four video reasoning benchmarks, Wan-CoF achieves considerable gains over the Wan2.2-I2V-A14B baseline. Building on this, we empirically explore more advanced designs for CoF capabilities, i.e., equipping the model with visual and textual reasoning tokens. This mechanism respectively captures low-level visual cues and high-level semantic priors for spatial and temporal reasoning. Through performance comparisons and attention analysis, we examine how these tokens contribute across model depth, denoising steps, space, and time. Our results suggest that stronger video reasoning requires both broad temporal supervision and explicit mechanisms for organizing intermediate reasoning state. We will open-source the dataset, model, and code to facilitate future research on reasoning-oriented video generation.
Opponent Modeling in Incomplete-Information Continuous Colonel Blotto
Yuanyuan Zhang ⋅ Gang Xiao ⋅ Feng Ye ⋅ Lingtao Xue ⋅ Zhipeng Du
We study which reduced opponent states preserve the interim decision problem of a realized type in incomplete-information continuous Colonel Blotto. For any continuous observable summary of the opponent posterior, we prove an exact ambiguity identity: the worst-case payoff gap over posterior laws with the same state equals twice the sup-norm distance of the payoff slice from the closed observable-additive span. Thus a state is payoff-exact if and only if every payoff slice lies in that span. Specializing to exact-budget Blotto, we characterize first-order battlefield marginals: they are exact precisely for coordinate-additive opponent slices. The two-battlefield case is degenerate, whereas from three battlefields onward first-order marginals can discard payoff-relevant dependence; more generally, exact-budget allocation games exhibit a strict $r$-order marginal hierarchy. The same characterization separates estimation from representation: exact states admit vanishing statistical error, while non-exact states impose a sample-independent ambiguity floor for state-restricted learners. Finally, exact and efficiently learnable first-order states do not eliminate equilibrium complexity: approximate Bayesian Nash equilibrium remains PPAD-hard in a strict diagonal regularized Blotto subclass.
Optimal Hidden-Target Learning for Online Inventory Optimization on General Convex Sets
Anthony Pineci ⋅ Yunzong Xu
Online inventory optimization (OIO) is online convex optimization with physical memory: inventory can be raised immediately, but it can be reduced only by demand. A natural principle, used in stochastic inventory learning and recently in OIO with linear capacity constraints, is to maintain a hidden target chosen by an online learner and implement its projection onto the currently feasible order-up-to set. We prove that this simple principle is optimal for OIO on arbitrary bounded convex capacity sets. With online gradient descent as the hidden learner, the method improves the best known regret guarantee for OIO on general convex sets from inverse to inverse-square-root dependence on the common-demand probability, and we prove a matching lower bound. The analysis identifies the right geometric state variable: the Euclidean distance between the hidden target and the implementable set. This distance evolves pathwise as a scalar queue, with target movement as arrivals and common demand as service, reducing the state-dependent feasibility cost to the control of a one-dimensional queue. The same reduction gives the first logarithmic regret guarantee for strongly convex losses and the first dynamic regret guarantee adapting to Euclidean path variation on general convex capacity sets.
Optimize Once, Execute Fast: Latency-Aware Multi-Agent Workflow Learning for Recurrent Queries
Enpei Zhang ⋅ Feiyu Qu ⋅ Zheng Huang ⋅ Dawei Zhou ⋅ Elynn Chen ⋅ Yujun Yan
Large Language Model-based Multi-Agent Systems (LLM-based MAS) have shown remarkable success in solving complex tasks by coordinating specialized agents through multi-step workflows. To overcome the high cost and specificity of manually designed workflows, recent research has shifted toward learning them automatically. However, a key limitation remains: while agent calls introduce significant latency due to model inference time and service congestion, most automated methods optimize solely for downstream performance, ignoring execution time (makespan). This issue is amplified when workflows are reused across recurring queries, making latency accumulate over time. This limits deployment in recurrent applications such as AI customer support and AI cloud troubleshooting, where users expect fast responses. To bridge this gap, we introduce a new problem, termed latency-aware workflow generation, where the objective is to minimize workflow makespan while aiming to retain downstream performance. We propose LAWA (Latency-Aware Workflow optimizAtion), a novel framework that reduces makespan via theoretically grounded graph edits targeting critical-path bottlenecks. By unifying a graph generator with this graph-editing procedure, LAWA supports both de novo workflow generation and refinement of existing ones. Empirical results across seven benchmarks demonstrate that LAWA significantly reduces makespan while achieving competitive task performance, supporting efficient reuse of multi-agent workflows for recurrent queries.
In the fair allocation of indivisible goods, a widely used notion of fairness is envy-freeness up to one good (EF1). A classical way to compute an EF1 allocation is the envy cycle elimination (ECE) algorithm, which iteratively assigns a good to an unenvied agent and, after each assignment, resolves any resulting envy cycle. Although the ECE algorithm always produces an EF1 allocation, it leaves considerable freedom in choosing both the next good to allocate and the agent to receive it. We investigate natural heuristics that exploit this flexibility to improve welfare guarantees. For example, we show that if the heuristic jointly selects the good and the receiving agent maximizing the utility, the worst-case utilitarian welfare loss is significantly lower than that of the vanilla algorithm. By contrast, restricting the heuristic to select only one of these two dimensions does not yield comparable improvements. We also complement our theoretical results with empirical average-case analysis.
OrangeTree: A Linear and Tree-based Time Series Forecasting Model Supporting Multiple Input and Output Lengths
Yiqi Tang ⋅ Zhichen Lai ⋅ Yuxuan Yang ⋅ Dalin Zhang ⋅ Jinpeng Chen ⋅ Gang Chen ⋅ Huan Li
Time series forecasting (TSF) models suffer from a critical rigidity: the inability to handle Multiple Input and Output Lengths (MIOL) within a single architecture. Current workarounds, such as padding or autoregression, invariably incur computational redundancy or error accumulation. We propose OrangeTree, a linear and tree-based architecture that decouples model parameters from sequence lengths. It employs a Segment Tree Encoder and an Inverse Tree Decoder to translate multi-length input history into multi-length output predictions within a unified framework. Crucially, a Range Weighter and Feature Fuser bridge these components, dynamically selecting and fusing relevant ranges into optimal contexts for forecasting. Experiments on six benchmarks confirm OrangeTree achieves SOTA performance, reducing MSE by 3.4\% compared to the previous SOTA. In MIOL settings, it matches the accuracy of length-fixed models without the accuracy degradation and latency increase typical of heuristic methods. Code is available at https://anonymous.4open.science/r/OrangeTreeTSF/.
Orthogonal Origin Parking: Decoupling Lorentz Manifolds for Robust OOD Generalization
Peter J Kampen ⋅ Anders N Christensen ⋅ Morten Rieger Hannemose ⋅ Anders Dahl ⋅ Josefine Vilsbøll Sundgaard
Fine-grained classification in deep taxonomies suffers from representational crowding: large macro-classes dominate the feature space, leaving little angular capacity for fine-grained distinctions, particularly under out-of-distribution (OOD) shifts. Standard hyperbolic spaces offer exponential volume for tree-like data, but embedding an entire taxonomy into a single manifold forces distinct macro-branches to compete for shared angular capacity and entangles their optimization updates. We introduce a geometric architecture that embeds hierarchical taxonomies into a Cartesian product of Lorentz manifolds. Our core mechanism, \textit{Orthogonal Origin Parking} (OOPark), decouples the taxonomy at the root level by assigning an independent Lorentz manifold to each macro-branch and penalizing inactive embeddings that deviate from the manifold origin. We evaluate across four biological OOD tasks/benchmarks (skin lesions, iWildCam, fungi, and plankton) characterised by hierarchical class structure, class imbalance, and domain shift. OOPark generally outperforms standard Euclidean and single-manifold hyperbolic baselines while using 20-dimensional sub-manifolds against Euclidean baselines of up to 512 dimensions, preserving both fine-grained accuracy and global taxonomic fidelity under severe distribution shifts. Code is provided in the supplementary material; a public release will follow upon acceptance.
Orthogonal Sparse Subgraph Alignment for Structure-Function Coupling in Brain Networks
Haonan Gao ⋅ Daeyoung Ham ⋅ Yifei Zhang ⋅ Xinyuan Tian ⋅ Shengxian Ding ⋅ Zhao ⋅ Tianxi Li
Characterizing the relationship between brain anatomical wiring and functional coordination remains a fundamental challenge in computational neuroscience. This difficulty primarily arises because structure-function coupling in the human brain is spatially heterogeneous, often localized to subnetworks, and inherently difficult to align across disparate representational modalities. To address this methodological gap, we introduce $\textbf{O}$rthogonal $\textbf{S}$parse $\textbf{S}$ubgraph $\textbf{A}$lignment ($\textbf{Ossa}$), a principled and interpretable framework designed to identify compact structural cliques that exhibit maximal concordance with functional connectivity (FC) across subjects. Within each FC-derived community, $\textbf{Ossa}$ optimizes an alignment objective over an orthogonal transformation, which absorbs coordinate mismatch between the two connectivity modalities, and a sparse node-selection vector, which identifies the induced structural connectivity (SC) subgraph. Furthermore, we establish theoretical properties for the $\textbf{Ossa}$ estimator, demonstrating its statistical consistency in recovering the true underlying SC clique within each functional community. Extensive empirical evaluations on simulated data, alongside analyses of the Adolescent Brain Cognitive Development (ABCD) baseline and longitudinal cohorts, demonstrate that $\textbf{Ossa}$ successfully recovers aligned structural cliques under latent cross-modal rotations. Ultimately, the proposed methodology significantly improves out-of-sample SC-FC alignment over size-matched baselines and robustly identifies stable, single-core topological structures within human brain networks.
Orthogonal Updates for the Win: Towards Accelerated Adaptive Minimax Optimization
Zhiwei Zhai ⋅ Xinyu Wang ⋅ Wenjing Yan ⋅ Lei Ding ⋅ Ying-Jun Zhang
Single-loop minimax methods are appealing for modern machine learning, but standard stochastic gradient updates can be unstable and oscillatory under tightly coupled primal--dual dynamics, making performance notoriously sensitive to stepsizes. To fundamentally improve stability, we develop an orthogonal update framework tailored to minimax optimization, which applies updates along orthogonalized directions to reduce update anisotropy. However, orthogonalization reshapes the optimization geometry and makes coordinating primal--dual stepsizes even more delicate. To overcome this challenge, we develop \textbf{AdaSGDA}, which equips orthogonal updates with an adaptive stepsize rule driven by accumulated gradient norms. This mechanism automatically balances coupled primal--dual progress under orthogonalization, yielding an effectively parameter-free method for nonconvex-strongly-concave minimax optimization. We prove that AdaSGDA avoids problem-dependent tuning while attaining the state-of-the-art $\mathcal{O}\left(T^{-1/4}\right)$ convergence rate. We further propose \textbf{AdaSGDA-VR}, which incorporates a carefully designed variance reduction scheme and achieves the faster $\mathcal{O}\left(T^{-1/3}\right)$ rate. Extensive experiments demonstrate that our methods deliver competitive performance compared to existing minimax optimization methods.
Orthros: Phase-Aware Heterogeneous Attention for Efficient Transformers
Yuhong CHOU ⋅ Zehao Liu ⋅ Yuqi Pan ⋅ Qian Liu ⋅ Xianwei Chen ⋅ Jibin Wu
Efficient attention mechanisms, such as linear and sliding window attention, aim to reduce the computational complexity of Transformers from quadratic to linear. However, their degraded in-context recall often necessitates interleaving full attention layers, leaving the quadratic bottleneck unresolved in long-context processing. To address this, we introduce Orthros attention, which shifts the efficiency paradigm from structural allocation to phase-aware decoupling. As a unified module, it dynamically switches its computation pattern across different phases of sequence processing, enabling linear complexity prefilling and full attention decoding. Our design grants Orthros attention a lower asymptotic complexity compared to vanilla Transformers in prevalent long-context scenarios characterized by heavy prefilling and light decoding. Beyond theoretical formulation and complexity analysis, extensive empirical evaluations demonstrate Orthros's superiority across multiple dimensions. First, we validate its fundamental language modeling capabilities, demonstrating scalability across varying model sizes and strong competitiveness in comprehensive architectural comparisons. Second, Orthros exhibits exceptional conversion feasibility via a data-efficient Transformer-to-Orthros conversion, yielding performance that even surpasses the original Transformer. Finally, Orthros achieves a remarkable $21.9\times$ prefilling speedup over Transformers while avoiding the degradation issues typical of traditional efficient attention mechanisms, evidenced by near-perfect 128K needle-in-a-haystack results and performance on par with strong Transformer baselines on real-world in-context recall tasks. Overall, these advantages in high-performance, flexibility, and efficiency position Orthros as a compelling industrial successor to the Transformer.
OTel: Open Telco AI Datasets, Benchmarks, and Models
Farbod Tavakkoli ⋅ Gregory Diamos ⋅ Kenneth Church ⋅ David Kanter ⋅ Mark Austin ⋅ Imtiaz Karim ⋅ Mirza M Rahman ⋅ Merouane DEBBAH ⋅ Zeinab Nezami ⋅ Ali Maatouk ⋅ Leandros Tassiulas ⋅ Rex Ying ⋅ Nick Sorros ⋅ Louis Powell ⋅ Nikolaos Vasiloglou ⋅ Ashish Vaswani ⋅ Somanshu Singla ⋅ Adarsh Chaluvaraju
We present Open Telco (OTel), an open telecom AI resource that releases derived telecom datasets for retrieval, reranking, instruction tuning, and safety/abstention, together with 30 full-parameter post-trained baselines spanning 10 embedding models, 3 rerankers, and 17 language models. The community has already engaged substantially with the resource: as of May 3, 2026, the released models have been downloaded over 16 million times and the project has received 157+ pieces of media coverage worldwide. Building on prior open telecom datasets and benchmarks, OTel provides documented telecom data sources, held-out evaluation partitions, trained embedding models, rerankers, context-grounded LLMs, and safety/abstention data in one unified resource. Each baseline starts from an open-weight model and is post-trained on OTel-derived data using an open training recipe, then evaluated on held-out OTel evaluation partitions. OTel post-training improves performance across all three model families: embedding retrieval reaches 93.5% NDCG@10, reranking reaches 0.952 MRR@10, and language-model correctness reaches 88.2%. We release OTel as a reproducible starting point and invite the community to expand the data, improve embedding and reranking models, and build stronger context-grounded telecom LLMs.
P$^{3}$: Joint Program-and-Proof Planning\\ for Verified Code Generation
Zenan Li ⋅ Ziran Yang ⋅ Peiyang Song ⋅ Zhaoyu Li ⋅ Kaiyu Yang
Verified code generation asks a large language model (LLM) to generate both an executable program and a machine-checkable proof that the program meets a formal specification, promising software that is correct by construction. The de facto workflow decouples the two halves of the problem: first synthesize a program, then attempt to prove it correct. We observe that this sequential pipeline can be both ineffective and inefficient in practice. A program generated without anticipating its proof can be subtly incorrect or structurally difficult to verify, forcing the LLM into brittle repair loops that alternate between patching the code and patching the proof. Inspired by Dijkstra's view that a program and its correctness argument should be developed hand in hand, we propose \framework{}, an LLM-based agentic workflow that first derives a unified program-and-proof plan from the specification, then elaborates the implementation and proof scaffold under this shared plan. To evaluate verified code generation in realistic settings, we further introduce Lean4Commit0, a project-level benchmark built by extracting core APIs from real-world software repositories and translating their requirements, including relational specifications across APIs, into Lean tasks. Using four frontier LLM backends, we evaluate \framework{} on Verina, AlgoVeri, and our Lean4Commit0 benchmark, where it achieves the highest solve rate in every benchmark--model setting. Compared with the stronger baseline, it improves solve rates by 4.6--11.2 percentage points and reduces per-task API cost by up to roughly 40\% and wall-clock time by up to roughly 37\% on the difficult subset of each benchmark.
PACE-dLLM: Elastic Block Decoding via Confidence Cliff Estimation for Diffusion Language Models
Xiaocheng Lu ⋅ Shuhan Guo ⋅ Ziyue Ma ⋅ Jie ZHANG ⋅ Jian Liu ⋅ Jingcai Guo ⋅ Haoxuan Che ⋅ Song Guo
Diffusion large language models (dLLMs) such as LLaDA and Dream now rival autoregressive LLMs in quality while retaining native parallel decoding. To deploy them efficiently, block-wise decoding partitions generation into blocks of size $B$ and commits a fraction of each block before moving on, but $B$ conflates two roles: \emph{how far the model can look ahead} and \emph{how many tokens get committed per step}. Recent accelerators relax this trade-off with indirect heuristics, yet the underlying difficulty is already exposed at every forward pass by the model's per-step confidence. Specifically, the in-window confidence profile exhibits a context-dependent \emph{cliff} (a high plateau, a sigmoidal transition, and a residual floor) that directly encodes how far ahead is safe to look. We propose \textbf{PACE-dLLM}, a decoder that handles the two roles independently. At each step, PACE-dLLM fits the cliff in closed form on the current window's confidences and reads off the next prediction horizon at the cliff's saturation point; a separate confidence threshold decides which tokens are committed. Both rules read the same signal, add a single hyperparameter, and admit a strict-dominance guarantee over fixed-block decoding. Extensive experiments across four reasoning and code benchmarks on two open-source dLLM backbones demonstrate that PACE-dLLM achieves the best average accuracy at a roughly $\mathbf{4.5\times}$ wall-clock speedup over the unaccelerated semi-AR baseline, pushing the Pareto frontier outward.
PACE: Pareto-Adaptive Compression for Efficient Native MLLMs
Lei Tian ⋅ Chunzheng Zhu ⋅ Kai WU ⋅ yifan zhang ⋅ Haihua Yang
Native multimodal large language models (native MLLMs) jointly encode visual and linguistic sequences within a unified autoregressive Transformer, emerging as a promising paradigm for multimodal understanding. Unlike conventional ViT-based MLLMs that rely on an external vision encoder, ViT-free native MLLMs tokenize every pixel patch directly, which is intuitively expected to suffer from severe visual token redundancy. To examine this hypothesis, we conduct a pilot study and find that, unlike ViT-based models which degrade under compression, ViT-free MLLMs tolerate aggressive compression with near-lossless or even improved per9 formance. Existing compression methods are built exclusively for the ViT-based setting, whereas native MLLMs, which stand to gain the most from visual token compression, remain entirely unexplored. To close this gap, we introduce an analysis pipeline for native backbones, centered on two metrics, i.e., token discriminability and cross-dimensional heterogeneous redundancy, which together reveal the following two findings: (i) redundancy is distributed independently and heterogeneously along the height and width axes, and (ii) in early layers, visual tokens are indiscriminable and uniformly redundant. Guided by these findings, we propose PACE (Pareto Adaptive Compression for Efficient native MLLM), the first visual token compression framework tailored for native MLLMs with Multimodal-RoPE. Leveraging the dimensional decoupling of position encoding, PACE decomposes attention into two orthogonal spatial saliency estimates along height and width, gated by a positional attention signal, yielding a direction-aware importance measure with negligible overhead. It further casts token budget allocation as a Pareto optimization over three objectives, i.e., performance preservation, inference efficiency, and spatial coverage, which jointly identify the optimal compression depths and retention ratios. Experiments show that PACE maintains near-lossless performance at only 25%–50% visual token retention, and even achieves performance gains on several benchmarks, while delivering over 2× FLOPs reduction.
Pairwise AUC Optimization Needs Corrective Power: A Unified View
Jia Chen ⋅ Zhiyong Yang ⋅ Shilong Bao ⋅ Qianqian Xu ⋅ Qingming Huang
Direct pairwise AUC-surrogate training is attractive because it trains toward the ranking metric used at test time, especially under class imbalance. This creates a direct alignment between the training objective and the evaluation metric, but such metric alignment alone does not guarantee that stochastic training remains corrective. We study this gap through Corrective Power: whether pairwise mini-batch updates produce stronger and more reliable improvements for more severe positive-negative ranking errors. Our main observation is that direct pairwise training can enter regions where severe errors remain common but no longer produce useful corrective motion. In such regions, even the mini-batch noise that would normally help SGD move along the steepest severe-error correction direction can become much weaker. In this paper, we seek a theoretical explanation for this by exploring the degenerate property of the U-statistics formed from the pairwise gradients therein. Our results provide a mechanistic explanation for several empirical patterns: (a) squared-hinge loss is often an effective AUC surrogate; (b) cross-entropy warm-up can help models that enter the AUC phase from a poor starting region; and (c) parameter-efficient fine-tuning from a pretrained model can reduce that need by starting closer to a favorable ranking geometry. Experiments on 8 CV tasks and 8 NLP tasks, using ResNet50, DenseNet121, DistilBERT-base, and Qwen2.5-1.5B-Instruct, support these predictions.
PAPO-VLA: Planning-Aware Policy Optimization for Vision-Language-Action Models
Peizheng Guo ⋅ Jingyao Wang ⋅ Changwen Zheng ⋅ Wenwen Qiang
Vision-Language-Action (VLA) models show promising ability in language-guided robotic tasks. However, making VLA policies reliable remains challenging, because a manipulation task is completed through closed-loop interaction, where each action affects subsequent execution. To analyze this problem, we revisit VLA policy during execution and argue that a VLA policy acts both as a planner, which makes task-oriented decisions that change the direction of execution, and as an executor, which realizes these decisions through dense continuous actions. This view suggests that improving VLA reliability requires particular attention to planning actions. Existing optimization methods can imitate actions or improve complete trajectories, but they usually do not explicitly identify planning actions or measure their importance for task success. To address this issue, we propose Planning-Aware Policy Optimization for VLA models (PAPO-VLA). PAPO-VLA first identifies planning actions by jointly considering action variation and trajectory outcome, then estimates their importance through causal sufficiency and causal necessity, and finally incorporates this importance into GRPO advantage estimation. In this way, more important planning actions receive stronger optimization emphasis, while the whole trajectory is still optimized by trajectory-level feedback. Experiments on multiple benchmarks demonstrate the effectiveness of PAPO-VLA.
Paradoxes of Game Theoretic Equilibria and Price of Anarchy
Georgios Piliouras ⋅ Ian Gemp ⋅ Siqi Liu ⋅ Luke Marris
Static equilibria—Nash, Correlated (CE), and Coarse Correlated Equilibria (CCE)—and the Price of Anarchy (PoA) provide essential, tractable benchmarks for multi-agent systems. However, evaluating learning dynamics purely through $C^0$ fixed points and discrete empirical distributions abstracts away critical $C^1$ vector field information. We demonstrate that this reduction diverges significantly from physical learning trajectories. First, the worst-case pure Nash equilibria anchoring canonical PoA bounds mathematically manifest as strict saddles—and in canonical instances, even as global maxima of the exact potential; because they are topologically unstable repellers, natural learning actively bypasses these theoretical inefficiency bounds. Additionally, the PoA metric itself exhibits algebraic sensitivity; relaxing syntactic constraints to accommodate data-driven, strictly positive affine cost models renders the Price of Anarchy unbounded. Furthermore, evaluating algorithms strictly through time-averaged regret minimization to reach CCE or Proximal CE (PCE) structurally permits convergence to strictly dominated strategies. Even enforcing optimal $O(1/T)$ swap-regret minimization provably accommodates chaotic limit sets in minimal normal-form games. Finally, in non-atomic congestion games, discrete-time learning natively bifurcates into Li-Yorke chaos, driving time-averaged inefficiency to scale exponentially as $2^p$, diverging from static sub-linear bounds. Collectively, these findings highlight the necessity of augmenting classical algebraic frameworks with dynamically grounded metrics.
Parallel Broyden methods for efficiently evaluating nonlinear state space models
Ian Christopher Tanoh ⋅ Scott Linderman
Nonlinear state space models are typically evaluated sequentially, limiting their ability to exploit modern parallel hardware. However, recent work has shown that state space models can be evaluated in parallel by reformulating evaluation as a root-finding problem, which can be solved with a parallel form of Newton's method. However, exact Newton methods require expensive multiplication of dense Jacobian matrices, and quasi-Newton methods based on diagonal approximations, while more scalable, are often slow to converge because they neglect interactions across state dimensions. We introduce a novel quasi-Newton approach based on Broyden’s method, which captures coupling terms with diagonal-plus-low-rank approximations to blocks of the Jacobian. We construct the approximations from trajectory secants, avoiding automatic differentiation while retaining the favorable scaling of diagonal approximations and remaining compatible with parallel scan through rank compression. We extend previous theoretical results to show that convergence degrades monotonically with approximation error and prove finite-step recovery of the exact trajectory. Empirically, we show that our method is advantageous when interactions are approximately low rank or when Jacobian evaluation is prohibitively expensive. It converges in fewer iterations than other quasi-Newton methods, yielding substantially faster run times for simulating low-rank RNNs, parallel MCMC algorithms, and denoising diffusion models.
ParallelKernelBench: Can LLMs Write Fast Multi-GPU Kernels?
Willy Chan ⋅ Nathan Paek ⋅ Simon Guo ⋅ Simran Arora ⋅ Dan Fu
There is growing interest in using large language models (LLMs) to write high-performance GPU kernels, with recent work showing promising results on single-GPU workloads. However, multi-GPU kernel generation remains unexplored, despite communication emerging as a dominant bottleneck in large-scale training and inference. In this paper, we study how well LLMs can write multi-GPU kernels, a task compounded by three challenges: (1) the design space is combinatorially large, as training and inference workloads can be parallelized across tensor, expert, pipeline, data, and sequence dimensions; (2) single-GPU memory-compute roofline analysis fails to capture communication bottlenecks in multi-GPU execution; and (3) the many hardware paths available for communication (e.g., copy engine, TMA, SM instructions) each carry distinct tradeoffs. We introduce ParallelKernelBench (PKB), a benchmark and evaluation framework for multi-GPU kernel generation. Additionally, we construct a taxonomy of distributed workloads spanning different parallelism types and select 87 problems covering compositions that arise in real workloads. We contribute evaluations of frontier coding models on PKB, finding that current LLMs struggle: single-shot kernel correctness plateaus at 32% of cases, with speedup exceeding an unoverlapped baseline (PyTorch + NCCL) in only 25% of cases. We also contribute a communication-aware roofline analysis of correct kernels, finding that over 90% of baselines achieve less than 50% of peak hardware utilization. Optimized solutions to many PKB workloads are largely absent from existing open-source repositories; we highlight several LLM-generated net new kernels that outperform their reference implementations in specific regimes, including NeMo's vocab-parallel log-probability kernel (up to 2.06x), Hyena CP (up to 1.72x), and SAM3 IoU suppression (up to 1.40x).
Past, Future, All at Once: Breaking Stability-Plasticity Dilemma via Post-hoc JANUS Rectification
Zhilong Zheng ⋅ Letian Tao ⋅ Yang Guan ⋅ Yujie Yang ⋅ Wei Xiong ⋅ Kehua Sheng ⋅ Bo Zhang ⋅ Jingliang Duan ⋅ Keqiang Li ⋅ Shengbo Eben Li
Fine-tuning foundation models on new tasks inevitably suffer from catastrophic forgetting. While existing works attempt to mitigate this on the basis of parameter-efficient fine-tuning methods, they adopted an overly restrictive Subspace Orthogonality condition. In this paper, we introduce a purely post-hoc and tuning-agnostic weight rectification framework that achieves Parameter Space Orthogonality, which is the necessary and sufficient condition for preserving historical performance to the first order. By projecting parameter updates into the JAcobian NUll Space (JANUS), our method significantly recovers compromised historical knowledge without interfering with the underlying fine-tuning process. To overcome the local validity of the Jacobian approximation, we further propose a Multi-step Adaptive Rectification mechanism that utilizes the JANUS shift to dynamically verify the valid trust region and adjust step sizes. Coupled with our proposed ghost projection, ghost orientation comparison, and sequence-level singular value decomposition compression techniques, JANUS also achieves great temporal and spatial efficiency. Experiments demonstrate that JANUS seamlessly integrates with various fine-tuning methods, fundamentally breaking the stability-plasticity dilemma by recovering historical knowledge while preserving downstream task adaptation.
PatchKV: Weight Space Compensation of KV Cache
Chanryeol Lee ⋅ Chanhyuk Lee ⋅ Yeonwoo Choi ⋅ Donggyun Kim ⋅ Seunghoon Hong
Long-context inference with Large Language Models (LLMs) is bottlenecked by the linearly growing memory of the key-value (KV) cache. Existing compression methods reduce the cache through token eviction or approximation, but degrade sharply at aggressive compression budgets. We propose PatchKV, a training-free framework that compensates KV cache compression methods by carrying part of the context in the model's weights. PatchKV pairs an off-the-shelf compressed KV cache with a context-specific weight patch, which is computed once at context-loading time and served for downstream queries for the context. The weight patch is derived in closed form via ridge regression, by aligning the block-wise activations of context-derived reference query tokens under the full cache and the compressed cache. Once merged into the model, the patch leaves the forward graph and per-query inference cost unchanged in the single-context, multi-query setting. Across long-context QA (SCBench with up to 170k tokens, SQuAD, NIAH) and math (GSM8K) benchmarks on three model architectures, PatchKV consistently improves cache compression methods, suggesting an alternative direction to compensate them at aggressive budgets.
Researchers often need web corpora that are not answers to a single query, but reusable collections of documents spanning many decentralized sources. We refer to this problem as thematic web data collection}: assembling semantically relevant documents that share a common theme but are scattered across heterogeneous and structurally disconnected regions of the web. Existing approaches, including traversal-based methods and LLM-driven web agents, are limited by local exploration or short-horizon stopping, preventing comprehensive collection. We formulate this task as a long-horizon information foraging problem over query-induced semantic patches, where the central challenge is deciding when to continue local exploitation and when to switch to new regions. We propose PatchScout, a multi-agent framework that instantiates a class of patch-switching policies to balance local exploitation and adaptive patch switching under a fixed budget. Experiments on four live-web topics show that PatchScout substantially improves yield and domain coverage over traversal-based and search-augmented reasoning agent baselines. In an applied social science setting with an incomplete prior collection, PatchScout discovers 102 previously unidentified entities, illustrating its ability to expand existing thematic datasets.
PCBInnoBench: Benchmarking LLM Agents on Real-World PCB Design
Fuyuan Xia ⋅ Jiaxin Hu ⋅ Kewei Yu ⋅ Zihan Zhang ⋅ Zheng wei ⋅ Hehan Liang ⋅ Xianyue Chen ⋅ Tongze Zhang ⋅ Xingbo Feng ⋅ Shixiong Kai ⋅ Zhentao Tang ⋅ Yuqi Cui ⋅ Mingxuan Yuan ⋅ Linfeng Zhang ⋅ Qibing Ren ⋅ Junchi Yan
LLM-based agents have made rapid progress on software engineering benchmarks, and executable evaluation is extending to chip-level hardware tasks. Board-level PCB design and upgrade has yet to receive comparable evaluation, largely because open-source PCB projects rarely provide the structured issue/PR/test histories that existing benchmarks rely on. We introduce PCBInnoBench, the first benchmark that evaluates LLM agents on executable PCB design-and-upgrade workflows. PCBInnoBench contains 372 expert-designed KiCad tasks grounded in real PCB projects and end-to-end engineering workflows. Each task pairs an agent-visible package (KiCad project and multimodal engineering references) with a hidden evaluation package (ground-truth design and expert-reviewed validators) used only for scoring. Agents must modify existing ECAD projects and pass structural, task-specific, and cross-artifact engineering-constraint checks. Each task is evaluated under two prompts of different guidance levels: one supplies an expert execution trajectory of ECAD edits together with the consulted references, the other states only the task objective, requiring agents to navigate the full workflow from requirement localization through cross-artifact editing. The best agent achieves 22.0% Pass@1 with guidance but only 7.3% without, indicating that both workflow discovery and ECAD execution remain substantial challenges for current agents. Beyond PCB, the workflow-grounded, expert-constructed evaluation approach may serve as a reference for other engineering domains where historical traces are sparse and correctness spans heterogeneous artifacts. Data and code are available at https://anonymous.4open.science/r/pcbinnobench-C841 .
PepSpecBench: A Unified Evaluation Benchmark for Peptide Tandem Mass Spectrometry Prediction
Zhiwen Yang ⋅ Pan Liu ⋅ yifan Li ⋅ Yunhua Zhong ⋅ Jun Xia
Tandem mass spectrometry provides a high-throughput framework for identifying and quantifying proteins in complex biological samples. In computational proteomics, predicting peptide MS/MS spectra is a critical task, enabling downstream applications such as large-scale peptide identification and quantification. While deep learning architectures have substantially improved prediction accuracy, three evaluation challenges obscure the true progress of the field. First, inconsistent data preprocessing and incompatible model output spaces hinder fair model comparison. Second, flawed data splitting strategies can permit hidden sequence leakage and inflate reported performance. Third, existing evaluations typically lack comprehensive cross-species benchmarking and systematic assessment of model robustness to influential experimental conditions. To address these challenges, we propose PepSpecBench, a unified benchmark for peptide MS/MS spectrum prediction. PepSpecBench standardizes data preprocessing across complementary public datasets, enforces a strict backbone-disjoint splitting strategy to eliminate sequence leakage, and evaluates diverse architectures within a shared fragment-ion representation space. It further introduces a comprehensive multi-species evaluation suite and physically grounded metadata perturbation probes to assess model robustness and instrument awareness. We uncover previously unrecognized performance discrepancies and robustness limitations across six representative models, providing actionable insights for future model design, evaluation and practical deployment. Code and data are available at https://anonymous.4open.science/r/PepSpecBench-NeurIPS2026/.
Permanent and Transient Representations for Continual Reinforcement Learning
Nishanth Anand ⋅ Doina Precup
Continual Reinforcement Learning (CRL) agents struggle to adapt to new situations while retaining prior knowledge, leading to a stability–plasticity trade-off. One approach to solving this problem is by using complementary learning systems: one for slow, long-term learning and another for quick, transient adaptation. Anand & Precup (2023) instantiated this idea using permanent and transient value functions, but used the same features in both systems. Intuitively, long-term learning should rely more on parametric representations for broad generalization, while transient learning would be best served by non-parametric representations for situation-specific adaptation. In this paper, we explore this idea both theoretically and empirically. We propose a novel non-parametric approximation for estimating the transient value function in complex tasks. And, demonstrate that our method enables online learning and outperforms competitive baselines on image-based tasks and Craftax-Classic, underscoring the effectiveness of our system-level decomposition.
Persona Generators: Generating Diverse Synthetic Personas for Arbitrary Contexts
Davide Paglieri ⋅ Logan Cross ⋅ William Cunningham ⋅ Joel Leibo ⋅ Sasha Vezhnevets
Simulating human behavior with Large Language Models (LLMs) offers a scalable laboratory for social science and stress-testing AI products. However, limitations remain in matching LLM outputs against the full breadth of human diversity. Current techniques for creating synthetic humans typically optimize for density matching, which often enforces behavioral uniformity and overlooks the rare, consequential outliers. We argue that robust simulation requires shifting the focus toward support coverage: ensuring synthetic populations span the entire landscape of possible human traits, opinions, and preferences. This paper introduces Persona Generators, functions that can produce diverse synthetic populations tailored to arbitrary contexts. We apply an iterative improvement loop based on AlphaEvolve, using LLMs as mutation operators to evolve our Persona Generator code rather than the personas themselves. The optimization process produces lightweight functions that can automatically expand small descriptions into populations of diverse synthetic personas that are maximizing coverage along relevant diversity axes. We demonstrate that evolved generators substantially outperform existing baselines across six diversity metrics on held-out contexts. Furthermore, by successfully mitigating standard LLM mode collapse, our generated populations actually capture real human trait distributions better than baselines specifically designed to match human statistics.
Phantom Transfer: Data Poisoning can Survive Data-Level Defences
Andrew Draganov ⋅ Tolga Hasan Dur ⋅ Anandmayi (Andi) Bhongade ⋅ Mary Phuong
We present a data poisoning attack---Phantom Transfer---with the property that, even if you know precisely how the poison was placed into an otherwise benign dataset, you cannot filter it out. We achieve this by modifying subliminal learning to work in real-world contexts and demonstrate that the attack works regardless of which model produces the data, which model is trained on the data or what the attack target is. The attack survives 11 tested data-level defences, including one where every sample is paraphrased by a different model. We characterise when this attack works best and show that it can be used to plant password-triggered behaviours into models while still beating defences. In short, we provide an existence proof that maximum-affordance defences can fail to stop sophisticated data poisoning attacks. We suggest that future defences should supplement by include white-box methods and post-training model audits.
Dense associative memory underlies modern attention mechanisms; however, conventional Modern Hopfield Network (MHN) with linear kernels yields overlapping attractor basins and spurious states under high memory load, limiting their theoretically exponential storage capacity in practice. We construct a class of nonlinear phase kernels that reshape the energy dynamics into deep, well-separated basins, and prove that any kernel in this class preserves exponential storage capacity while guaranteeing convergence to fixed points. As a representative instantiation, we introduce the Sigmoidal Phase Hopfield Network (S-Hop), which guarantees monotone energy descent and mitigates the practical loss of memory capacity. Experiments on MNIST and CIFAR-10 show that S-Hop achieves up to a $50\times$ increase in critical capacity over baseline models. We also provide a systematic capacity measurement of Hopfield networks on TinyImageNet, where S-Hop exhibits improved retrieval robustness and noise resilience. The phase-kernel design provides a principled route toward more robust dense associative memory.
Phases of Muon: When Muon Eclipses SignSGD
Elliot Paquette ⋅ Noah Marshall ⋅ Lucas Benigni ⋅ Guangyuan Wang ⋅ Atish Agarwala ⋅ Courtney Paquette
Recently, Muon and related spectral optimizers have demonstrated strong empirical performance as scalable stochastic methods, often outperforming Adam. Yet their behaviour remains poorly understood. We analyze stochastic spectral optimizers, including Muon, on a high-dimensional matrix-valued least squares problem. We derive explicit deterministic dynamics that provide a tractable framework for studying learning behaviour with a focus on (stochastic) SignSVD and (stochastic) SignSGD, the latter serving as a proxy for Adam. Our analysis shows that for large batch size, SignSVD performs a square-root preconditioning with respect to the data covariance spectrum, while for small batch size smaller eigenmodes behave like SGD, slowing down convergence. We contrast with SignSGD which for generic covariance performs no preconditioning and has no transition, leading to different optimal learning rates and convergence characteristics. The two methods match up to a constant factor with isotropic data, but behave differently with anisotropic data. An analysis of a power law covariance model with data exponent $\alpha$ and target exponent $\beta$, and show there are three phases in the $(\alpha,\beta)$ plane: one where SignSGD is uniformly favored, one where SignSVD is uniformly favored, and a third where the two methods exhibit a trade-off in performance.
Phase Space Attention: A Hairer Lift Resolves the Single-Layer Induction Obstruction
Kingsuk Maitra ⋅ Shagun Sood ⋅ Morteza Hosseini ⋅ Suman Gunnala ⋅ Vikram Gupta
We resolve the Sanford-Hsu-Telgarsky (SHT) single-layer induction obstruction within the linear, one-step, causal, bilinear, symplectically consistent design class CPSA on the post-RoPE substrate, by lifting standard transformer attention onto a symplectic phase space in direct correspondence with Hairer’s lift of Stormer-Verlet onto a one-step symplectic integrator. The reframing recasts SHT as a filter-order gap: a standard one-layer bilinear score realises a z-transform of joint order (0,0), whereas the induction discriminator requires (1,1). We close the gap by applying the symplectic upper shear Mgamma : (q,p) -> (q + gamma p, p) to the post-RoPE query and key streams. We prove that this lift is the unique solution within CPSA (Theorem 4); is operator-level symplectic with zero secular drift (Theorem 6); requires post-RoPE placement (Corollary 7); and, under explicit Assumption Set A, induces a closed-form induction phase transition at gammac = (d_k log(T-1))^(1/4) (Theorem 8). The SHT lower bound binds where total parameter budget approaches the bit budget of the two-layer induction circuit it forbids: negligible at frontier scale (>= 1B), but decisive in the sub-100M regime that ships in the billions on consumer edge form factors – phones, wearables, microcontrollers, embedded controllers – where it governs whether on-device in-context learning is feasible. We work at 4-92M by design. We report 74.2% single-layer induction at gamma = 2.0, Pearson r >= 0.998 for the closed-form transfer function, r = -0.679 for the embedding-axis low-pass induction filter; and the post-softmax map Frechet-linearises to Differential Transformer (Ye et al., 2024) at first order under the small-signal regime (SS).
Phase Transitions in Attention: A Bayesian Theory of Copy Head Emergence
Itay Lavie ⋅ Kirsten Fischer ⋅ Andrey Lekov ⋅ Frederic Van Maele ⋅ Zohar Ringel ⋅ Moritz Helias
Attention is the key mechanism underlying in-context learning in transformers, and attention patterns have been observed empirically to emerge abruptly during training. We present a Bayesian theory of feature learning in attention; we then focus on how the copy subcircuit in the first layer of an induction head is learned by analyzing a single-layer softmax attention network trained on a copy task. We derive a closed-form posterior over the attention matrix and reduce it to a low-dimensional order parameter space. This reduction reveals a phase transition in the amount of training data, which we verify using both Bayesian sampling and standard training with Adam. We contrast our results with linear attention and find that softmax attention exhibits a first-order phase transition while in linear attention an initial second-order phase transition is followed by a smooth, continuous evolution toward the structured attention pattern (crossover). Our work provides a first-principles theoretical account of the abrupt emergence of the copy subcircuit, reminiscent of the one observed in training large language models.
Physics-Informed Functional Tucker Method with RKHS Factors for Sparse Spatiotemporal Reconstruction
Zhengyang Liu ⋅ Jianguo Huang ⋅ Junxi Hu ⋅ Yucheng Mi ⋅ Haojie Cai ⋅ Wenxuan Shen ⋅ zhikun zhang ⋅ Renfeng Peng ⋅ Guangtao Zhang ⋅ Zhiqiang Liu
Reconstructing physical dynamics from sparse and irregular spatiotemporal observations is a fundamental challenge in scientific research. However, sparsity leaves large regions of the field unconstrained, making the inverse problem ill-posed: a model may fit the observed sensors while producing nonphysical off-sensor structures or unreliable predictions between support times. To address this challenge, we propose PhysFTM---a physics-informed functional Tucker method that parameterizes Tucker mode factors in reproducing kernel Hilbert space (RKHS) and enforces governing physical laws through finite-difference residuals at collocation points. These physical constraints are essential for credible reconstruction, since most query locations are never directly supervised by data. Specifically, PhysFTM first reconstructs continuous spatial fields at observed times via an alternating minimization scheme, with separate treatments for the Tucker core tensor and the RKHS-based factors. To further enable continuous temporal resolution, PhysFTM constructs a continuous trajectory in Tucker-core space, which is decoded by the learned spatial RKHS-FTM representation to recover the full spatiotemporal field at arbitrary query times. Experiments on Allen--Cahn and Navier--Stokes with varying observation ratios demonstrate that PhysFTM achieves superior reconstruction accuracy, improved physical consistency, and stronger continuous-time modeling capability.
PINNeval: A Comprehensive Evaluation Standard for Physics-Informed Neural Networks
Kevin von Bargen ⋅ Hamid Ebrahimy ⋅ Tim Römer
The Physics-Informed Neural Network (PINN) literature has produced a rapidly expanding catalogue of methods. We divide a PINN training pipeline systematically into four components and observe that many existing benchmarks study these components in isolation. These components are domain-importance adaptation, loss balancing, the neural network architecture, and the optimization scheme. Moreover, the studies set the surrounding and not considered components of the pipeline to a minimal default that is no longer representative of current best practice approaches. Conclusions drawn under such conditions lead to suboptimal results. In this work, we propose a new evaluation protocol that (1) treats the four components as an integrated system; (2) requires as a key aspect a competitive implementation of the components not currently under investigation; (3) specifies the metrics and reporting format, and (4) is accompanied by structured documentation that ensures the traceability of evaluations across studies. Based on this comprehensive protocol, we release with PINNeval an innovative JAX package containing 18 different methods from the current literature as well as nine efficiently implemented benchmark problems. Using PINNeval, we illustrate our evaluation protocol in relation to established benchmark studies and report on results that challenge established assumptions.
PixelPonder: Dynamic Patch Adaptation for Enhanced Multi-Conditional Text-to-Image Generation
Yanjie Pan ⋅ Qingdong He ⋅ Zhengkai Jiang ⋅ Pengcheng Xu ⋅ Chaoyi Wang ⋅ Jinlong Peng ⋅ Haoxuan Wang ⋅ Yun Cao ⋅ Zhenye Gan ⋅ Mingmin Chi ⋅ Bo Peng ⋅ Yabiao Wang
Recent advances in diffusion-based text-to-image generation have demonstrated promising results through visual condition control. However, existing ControlNet-like methods struggle with compositional visual conditioning - simultaneously preserving semantic fidelity across multiple het erogeneous control signals while maintaining high visual quality, where they employ separate control branches that often introduce redundant guidance during the denoising process, leading to structural distortions and artifacts in generated images. To address this issue, we present PixelPonder, a novel unified control framework, which allows for effective control of multiple visual conditions under a single control structure. Specifically, we design a patch-level adaptive condition selection mechanism that dynamically prioritizes spatially relevant control signals at the subregion level, enabling precise local guidance without global interference. Additionally, a time-aware control injection 017 scheme is deployed to modulate condition influence according to denoising timesteps, progressively transitioning from structural preservation to texture refinement and utilizing the control information from different categories to promote finer image generation. Extensive experiments demonstrate that PixelPonder surpasses previous methods across different benchmark datasets, showing superior improvement in spatial alignment accuracy while maintaining high textual semantic consistency.
Pixel-space Autoregressive Image Synthesis via Spectrum Serialization and Flow-based Refinement
Guiwei Zhang ⋅ Tianyu Zhang ⋅ Yalong Bai ⋅ Ying Ba ⋅ Zichang Tan ⋅ Yang Yang
Autoregressive (AR) image generators commonly rely on vector-quantized (VQ) autoencoders to compress images into discrete token sequences. However, quantization inevitably discards visual information, and the resulting reconstruction errors may propagate through AR decoding, limiting generation fidelity. We revisit the representation choice for AR image generation and ask whether a simpler and more faithful representation can alleviate this bottleneck while remaining compatible with autoregressive modeling. To this end, we estimate the intrinsic dimension of the natural image manifold under several widely used representations and find that raw patchified pixels exhibit the simplest underlying geometry among those evaluated. Motivated by this insight, we propose PixeLLM, an AR generative framework that operates directly on sequences of discrete pixel blocks, eliminating the need for a separately trained VQ autoencoder. PixeLLM decomposes generation into two stages: (1) an AR model generates a text-aligned draft image capturing global scene structure, and (2) a conditional flow-based pixel refinement Transformer enhances the draft with fine-grained details and mitigates artifacts. The resulting pipeline is straightforward, requiring no external semantic alignment. Despite its simplicity, PixeLLM achieves competitive performance on challenging text-conditioned and class-conditioned image generation benchmarks.
Pixels to Tokens: Token Space Efficient Active Learning for Low-Budget Semantic Segmentation
Utku Cicek ⋅ Mahip Singh ⋅ Ziyao Shang ⋅ Kimathi Kaai ⋅ C Thomas ⋅ Pablo Guerrero ⋅ Alexander Wong ⋅ Sirisha Rambhatla
Semantic segmentation typically requires dense pixel-level annotations, creating a prohibitive bottleneck for budget-constrained applications. While foundation models and vision transformers (ViTs) have redefined visual representations, current active learning (AL) strategies for these backbones largely operate at the patch or region level, where annotation requirements remain high. Extending ViT-based AL to the pixel-level, low-budget regime is uniquely challenging: with as little as one pixel query per image, a method must simultaneously identify the most informative locations and train a reliable decoder head from an extremely sparse signal. We introduce TEAL (Token-space Efficient Active Learning), the first pixel-level AL framework built around ViT representations. TEAL uses a frozen DINOv3 backbone and performs diversity-based selection directly on its native token lattice, avoiding the artifacts introduced by interpolating token embeddings to dense pixel resolutions. Our framework applies a two-level MaxHerding strategy over multi-layer token descriptors to select representative candidates, followed by margin-based uncertainty refinement on a finer decoder grid. Across CamVid, Cityscapes, ADE20K, and Pascal Context, TEAL consistently outperforms previous baselines under extreme label scarcity. After 10 rounds of 1-pixel-per-image queries, TEAL improves over the strongest baseline by up to $+\textbf{14.65}$ mIoU on CamVid and $+\textbf{23.34}$ mIoU on Pascal Context, showing that ViT token spaces provide an effective geometry for extremely-low-budget active segmentation.
Point-to-Manifold Geometry: Flexibly Overcoming the Curse of Dimensionality in Neural Computational Units
Rohan Ghosh ⋅ Mehul Motani
The ability of neural networks to generalize is fundamentally shaped by their computational units. While the perceptron provides a first-order, linear inductive bias and the radial basis function (RBF) unit offers an isotropic curved bias, we identify a critical opportunity for a structured extension: the Generative Matching Unit (GMU). Each individual GMU captures complex dependencies by treating the forward pass as an inference problem; its internal generative model optimizes instance-specific latent parameters to compute a reconstruction error, effectively measuring a point-to-manifold distance as opposed to the point-to-point distance in RBFs. We focus on linear GMUs, which yield closed-form analytical expressions for fast computation and extend naturally to convolutional architectures. Our theoretical analysis demonstrates that linear GMUs are highly flexible universal approximators capable of recovering structured class posteriors. We prove that while a single GMU can exactly emulate an RBF kernel and two can emulate the decision boundary of a perceptron, emulating a single $k$-order GMU’s decision boundary requires a hidden layer of standard units that scales polynomially or exponentially with input dimensionality. Furthermore, we prove that the GMU’s point-to-manifold distance remains discriminative in high dimensions where standard Euclidean distances fail, offering unique robustness to the curse of dimensionality. Finally, we show that GMUs offer superior flexibility and efficiency in learning smooth manifold decision boundaries compared to MLPs and RBFs. Motivated by the theoretical results, we place GMUs in the first network layer like RBFs, where their role is to enhance linear separability for subsequent layers. This avoids the typical gradient collapse problem with stacking RBF-like units while ensuring the generalization benefits remain. We find that across extensive experiments involving 27 tabular datasets, five vision datasets evaluated across 34 test-time corruption settings, and 30 synthetic scenarios, GMU networks demonstrate statistically significant improvements in generalization and robustness.
Point Tracking Improves World Action Models
Jiarui Guan ⋅ Wenshuai Zhao ⋅ Yue Pei ⋅ Ziliang Chen ⋅ Arno Solin ⋅ Juho Kannala
Robot policy learning benefits from world-action models that capture environment dynamics, but pixel-level prediction entangles dynamics with nuisance factors such as lighting and texture, making learned representations vulnerable to task-irrelevant visual variation. We propose JOPAT, a JOint Pixel-And-Track World-Action Model that predicts latent visual observations, 2D point tracks with visibility, and actions in a single denoising diffusion transformer. The key insight is that tracks provide an explicit representation of motion that captures long-horizon dynamics and remains robust under occlusion or partial out-of-frame motion, offering greater utility than modeling pixel appearance alone. On LIBERO and real-world LeRobot tasks, JOPAT improves over pixel-based baselines, with the largest gains on long-horizon tasks involving occlusion, object interaction, and off-screen motion.
Polynomial-Time Algorithm for Thiele Voting Rules with Voter Interval Preferences
Pasin Manurangsi ⋅ Krzysztof Sornat
We present a polynomial-time algorithm for computing an optimal committee of size $k$ under any given Thiele voting rule for elections on the Voter Interval domain (i.e., when voters can be ordered so that each candidate is approved by a consecutive voters). Our result extends to the Generalized Thiele rule, in which each voter has an individual weight (scoring) sequence. This resolves a 10-year-old open problem that was originally posed for Proportional Approval Voting and later extended to every Thiele rule (Elkind and Lackner, IJCAI 2015; Peters, AAAI 2018). Our main technical ingredient is a new structural result---a concavity theorem for families of intervals. It shows that, given two solutions of different sizes, one can construct a solution of any intermediate size whose score is at least the corresponding linear interpolation of the two scores. As a consequence, on Voter Interval profiles, the optimal total Thiele score is a concave function of the committee size. We exploit this concavity within an optimization framework based on a Lagrangian relaxation of a natural integer linear program formulation, obtained by moving the cardinality constraint into the objective. On Voter Interval profiles, the resulting constraint matrix is totally unimodular, so it can be solved in polynomial time. Our main algorithm and its proof were obtained via human--AI collaboration. In particular, a slightly simplified version of the main structural theorem used by the algorithm was obtained in a single call to Gemini Deep Think.
PolyTopoBench: A Benchmark for Complex Vector Polygon Generation from Remote Sensing Imagery
Zeping Liu ⋅ Ni Lao ⋅ Weiwei Sun ⋅ Gil Wolff ⋅ Yiqun Xie ⋅ Liang Zhao ⋅ Junfeng Jiao ⋅ Gengchen Mai
Vector polygon generation converts visual inputs, e.g., remote sensing (RS) images, into vectorized polygonal geometries, supporting applications such as autonomous driving, vector map construction, and remote sensing. Early pipelines predict raster masks and post-process them into polygons, which prevents end-to-end optimization and may miss small objects or introduce inaccurate vertices. Recent methods directly generate vector polygons, but most focus on simple exterior contours, while they either cannot represent complex polygons with holes or fail to preserve their topology. In this paper, we propose PolyTopoBench, a unified evaluation framework for vector polygon generation from RS images with explicit emphasis on complex polygons. PolyTopoBench evaluates both exterior and interior rings, and benchmarks representative methods, including segmentation-based polygonization pipelines, vision foundation model baselines, and specialized vector polygon generators, on two RS-image datasets covering buildings, roads, vegetation, and unvegetated regions. Experiments show that existing methods often recover simple exterior boundaries but degrade substantially on polygons with holes or multiple rings. These results reveal complex polygon generation as an unresolved challenge and motivate topology-aware benchmarks and model designs.
PolyVision: Conditional Visual Scaling Via Dynamic Expert Routing For Vision-Centric MLLMs
Tiehan Fan ⋅ Chen Zhao ⋅ Nikai Du ⋅ Zili Yi ⋅ Jian Yang ⋅ Ying Tai
Multimodal large language models have achieved remarkable progress, yet their language-dominant designs constrain visual perception and reasoning. Vision-centric multimodal large language models aim to restore visual understanding as the structural foundation, but existing systems rely on static visual expert fusion, leading to redundant computation, weak adaptivity, and limited multimodal transferability. Prior VC-MLLMs scale visual capacity by adding experts, whereas $\mathtt{PolyVision}$ scales visual capacity by allocating experts. We introduce $\mathtt{PolyVision}$, a vision-centric framework for conditional visual scaling in which dynamic visual expert routing assigns image patches to heterogeneous visual encoders. An \textbf{Attentive Router} selectively activates relevant experts for each patch, while a compact \textbf{VisPack} module executes routed representations efficiently and preserves example independence. Together with additional visual pre-alignment, this design turns heterogeneous visual scaling into a practical conditional-computation process. Built on the \texttt{Qwen3} backbone, $\mathtt{PolyVision}$ achieves superior efficiency and scalability on multimodal benchmarks. Experiments show that adaptive expert routing strengthens visual specialization, clarifies practical properties of visual scaling, and improves visual understanding at a favorable trade-off between efficiency and performance.
Pooling Versus Ensembling for Ridge Regression Under Covariate Shift
Maya Ramchandran ⋅ Rajarshi Mukherjee
Datasets in many settings naturally partition into clusters arising from sub-populations, batch effects, or aggregation across multiple sources. A common response to such heterogeneity is to ensemble learners trained on each cluster rather than fit a single model to the pooled data. Prior work motivating such approaches has typically considered settings in which both the covariate distribution and the conditional outcome model differ across clusters; the role of cluster-aware partitioning and ensembling based solely on the covariate distribution remains to be explored. We address this case for ridge-regularized least-squares regression under a linear outcome model and consider all ridge penalty values $\lambda \geq 0$, including the special case of the ridgeless predictor at $\lambda = 0$. By considering both fixed-effects and random-effects models, we argue that under random effects, an optimally tuned pooled ridge predictor always outperforms ensembles of individually optimally tuned predictors. For fixed effects, we derive a general formula for the pooled and ensembled predictors to characterize the role of both regression coefficients as well as the predictor distribution shifts. Together, these results generalize prior risk analyses of bagging and random-partition estimation using ridge and ridgeless regression predictors from the i.i.d. setting to encompass covariate shift and heterogeneity-aware partition structure.
PO-PDDL: Learning Symbolic POMDPs from Visual Demonstrations for Robot Planning Under Uncertainty
Wenjing Tang ⋅ Jin Xuanjin ⋅ Yuan Liu ⋅ RenMing Huang ⋅ Cewu Lu ⋅ Panpan Cai
Real-world robot task planning must operate under both stochastic action execution and partial observability, yet constructing Partially Observable Markov Decision Process (POMDP) models for real robotics domains remains difficult and labor-intensive. We introduce PO-PDDL, a symbolic formulation of POMDPs that preserves the relational structure and LLM-friendly syntax of the Planning Domain Definition Language (PDDL), while explicitly modeling partial observability, stochasticity, and beliefs. Building on this formulation, we propose a demonstration-driven pipeline for learning PO-PDDL models. The proposed method reconstructs latent symbolic state trajectories from real-robot execution videos, identifies partial observability via inconsistencies between inferred states and visual observations, and learns stochastic transition and observation models accordingly. The resulting PO-PDDL domains are reusable across tasks and enable online belief-space planning under both perception and execution uncertainty. Experiments on real-world long-horizon manipulation tasks show that our method consistently outperforms existing PDDL and POMDP model-learning approaches, achieving robust task planning under uncertainty with significantly lower planning cost.
Position: Adopt Constraints Over Fixed Penalties in Deep Learning
Juan Ramirez ⋅ Seyed Meraj Hashemizadehaghda ⋅ Simon Lacoste-Julien
Recent efforts to develop trustworthy AI systems have increased interest in learning problems with explicit requirements, or constraints. In deep learning, however, such problems are often handled through fixed weighted-sum penalization: the constraints are added to the task loss with fixed coefficients, and the resulting scalarized objective is minimized. This position paper argues that fixed penalization is often ill-suited for deep learning problems with non-negotiable requirements for several reasons. First, in non-convex settings, the penalized and constrained problems are generally not equivalent, so solving the former need not solve the latter. Second, fixed penalization weakens hard requirements into soft penalties to be traded off against task performance. Third, choosing penalty coefficients to indirectly solve the constrained problem often involves costly trial and error, because changing them alters the penalized objective itself, and hence can mean solving the wrong problem altogether. We therefore argue that, when a deep learning problem specifies non-negotiable requirements, the constrained formulation itself should be the starting point, not the surrogate problem defined by fixed penalization. The appropriate solution strategy should then be chosen based on the problem's structure and scale.
Position: Machine Learning Models for Reaction Transition States Deserve Better
Atharva Tambat ⋅ Ankit Ghosh ⋅ Swastik Kumar ⋅ Raghavan B Sunoj ⋅ Abir De
Understanding reaction mechanisms is central to advances in chemistry, materials science, and drug discovery, where transition states govern reactivity and kinetics. This position paper argues that current ML approaches for transition state (TS) prediction are fundamentally limited by inadequate datasets, weakly grounded methods, and flawed evaluation protocols, leading to overstated progress and questionable chemical reliability. While recent advances have improved performance, these gains are largely superficial when assessed against the core requirements of a valid TS. First, widely used datasets suffer from issues of chemical implausibility, lack of diversity, and insufficient validation, introducing noise and bias into model training. Second, existing methods often neglect essential physical constraints and rely on assumptions like fixed atom mappings, resulting in brittle, non-generalizable models. Third, standard evaluation metrics fail to capture chemical validity, often rewarding plausible but physically incorrect predictions. Empirical analyses, including quantum chemical validation and robustness studies, demonstrate that many reported transition states do not correspond to intended reactions and that models exhibit poor generalization and strong bias toward equilibrium configurations. We contend that meaningful progress in TS prediction requires a principled rethinking of dataset construction, incorporation of physical constraints into model design, and the development of chemically faithful evaluation frameworks.
Position Without Positional Embeddings: A Directional Mechanism in NoPE Transformers
Matan Avitan ⋅ Ido Nachum ⋅ Yanai Elazar ⋅ Yoav Goldberg
Autoregressive transformers with no positional embeddings (NoPE) can recover absolute position. A common explanation is that causal attention creates a position-dependent variance signal, but this account is incomplete because LayerNorm removes per-token scale, so a viable mechanism must encode position in direction, not magnitude. We describe the most compact mechanism with two-layer NoPE transformers that accurately extract position. In Layer1, prefix averaging creates a shared beginning of sequence (BOS)-direction component. In Layer2, the attention's output-value matrix maps the \bos{} and \nbos{} contributions into two distinct directions, and a position-dependent \bos{} attention weight acts as a mixing coefficient that interpolates between them. The resulting directional trajectory is linearly decodable for position. Interestingly, we show two transformer variants that converge to the same theoretical mechanism and instantiate it when trained to predict position.
POST: Progressive Object-Slot Tokenization for Multimodal Large Language Models
Junyu Zhou ⋅ han li ⋅ Wenrui Dai ⋅ Fan He ⋅ Ziyang Zheng ⋅ Hang Xu ⋅ Chenglin Li ⋅ Junni Zou ⋅ Hongkai Xiong
Multimodal large language models (MLLMs) leverage object-agnostic tokenization by default, which represents images as flat grids of patch tokens and uniformly discretizes semantically rich foreground regions and information-sparse background areas. This introduces substantial redundancy and constitutes a major source of inference cost. Existing object-level tokenizers could either discard intra-region saliency with predefined region grouping and coarse pooling or limit to a single semantic scale by employing slot attention to the final encoder feature map only. In this paper, we propose a two-stage framework named POST that decouples unsupervised object-centric grouping from language-grounded multimodal reasoning to achieve progressively semantic-aligned object-level visual tokenizers. Specifically, we develop Progressive Slot Attention (PSA) in Stage I to introduce slot attention across multiple layers of a frozen ViT encoder and propagates slot states through a cross-layer refinement chain for learning hierarchical slot--patch correspondences without segmentation supervision. Subsequently, we transfer the learned PSA to the LLaVA pipeline in Stage II for object-level visual tokenization. We design slot-weighted merging to aggregate ViT patch tokens into object-grounded visual tokens and remove low-support slots with adaptive slot pruning. Remarkably, POST eliminates the need for external segmentation models, enables progressive object-centric refinement, and adapts visual token budget to image complexity in visual tokenization. Experimental results show that PSA achieves strong unsupervised object-centric segmentation on PASCAL VOC and COCO. Furthermore, POST is comparable or superior to LLaVA-1.5 on most multimodal benchmarks using only 8.3% visual token budget and improves referring expression comprehension on RefCOCO/+/g. It also outperforms recent patch-level and object-level token compression baselines at comparable budgets.
Power Distribution Bridges Sampling, Self-Reward RL, and Self-Distillation
Akiyoshi Tomihari ⋅ Issei Sato
Recent analyses question whether reinforcement learning (RL) is responsible for strong reasoning in large language models (LLMs). At the same time, distillation and inference-time sampling, including power sampling, have emerged as effective ways to improve LLM performance. However, the relationship among RL, distillation, and sampling remains unclear. In this study, we focus on the power distribution, the target distribution of power sampling, and show that the power distribution bridges sampling, self-reward KL-regularized RL, and self-distillation. From the sampling perspective, we show that inexpensive local approximations cannot reproduce sequence-level power without information about possible suffixes. From the RL perspective, the power distribution is the closed-form optimizer of KL-regularized RL when the model's sequence-level log-probabilities are used as the reward. This identification leads to \emph{power self-distillation}, an offline distillation surrogate that shares the same target distribution and amortizes the cost of power sampling into supervised training on teacher samples. We further show that power self-distillation can achieve self-reward sharpening, while improvement in a downstream true reward is governed by the covariance between true reward and self-reward under the power distribution. Experiments on reasoning tasks support our analysis: power sampling raises self-reward, true-reward gains depend on alignment with self-reward, and power self-distillation can match or exceed the performance of power sampling at much lower inference cost.
Generative AI models now produce images indistinguishable from real data, making publicly verifiable provenance essential; however, existing watermarking methods become vulnerable to forgeries and adversarial optimization attacks once their detectors are made public. We propose PP-Mark, a provenance framework that anchors lightweight statistical detection to a zero-knowledge proof of the embedding process, ensuring that forged content cannot pass full verification by construction. We establish formal unforgeability under standard cryptographic assumptions and validate PP-Mark against representative baselines across two architectures, demonstrating resistance to black-box imprint forgery, white-box optimization, and regeneration attacks while remaining practical for decentralized verification.
Preconditioned Flow Matching
Shadab Ahamed ⋅ Eshed Gal ⋅ Md Shahriar Rahim Siddiqui ⋅ Simon Ghyselincks ⋅ Moshe Eliasof ⋅ Eldad Haber
Flow matching (FM) learns vector fields by regressing stochastic velocity targets along intermediate distributions $p_t$. We identify a geometric optimization bottleneck in this regression problem: when the covariance $\Sigma_t$ of $p_t$ is ill-conditioned, gradient-based training rapidly fits high-variance directions while making slow progress along low-variance ones. In an exactly solvable Gaussian setting, we prove that the excess risk is weighted by $\Sigma_t$, and that both gradient descent and stochastic gradient descent inherit condition-number-dependent convergence. We then extend the analysis to Gaussian mixtures, showing that multimodality does not average away this effect; instead, the slowest and worst-conditioned component can control optimization. Motivated by this analysis, we propose \emph{preconditioned flow matching}, a precondition-then-match framework that transforms the target distribution into a more isotropic representation, trains the main flow in the transformed space, and maps generated samples back through the inverse transformation. We show theoretically that preconditioning reshapes the intermediate FM path and improves its conditioning. Across controlled Gaussian and Gaussian-mixture experiments, latent MNIST and other high resolution image datasets up to $512{\times}512$ resolution, preconditioning improves path-conditioning diagnostics, low-eigenvalue recovery, FID, MMD, precision, and recall. Compute-matched baselines and preconditioner-quality ablations further show that the gains are not explained merely by additional preconditioner parameters, but by improved geometry of the downstream flow matching problem.
Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants
Vincent Huang ⋅ Dami Choi ⋅ Daniel D Johnson ⋅ Sarah Schwettmann ⋅ Jacob Steinhardt
Interpreting the internal activations of neural networks can produce more faithful explanations of their behavior, but is difficult due to the complex structure of activation space. Existing approaches to scalable interpretability use hand-designed agents that make and test hypotheses about how internal activations relate to external behavior. We propose to instead turn this task into an end-to-end training objective, by training interpretability assistants to accurately predict model behavior from activations through a communication bottleneck. Specifically, an encoder compresses activations to a sparse list of concepts, and a decoder reads this list and answers a natural language question about the model. We show how to pretrain this assistant on large unstructured data, then finetune it to answer questions. The resulting architecture, which we call a Predictive Concept Decoder, enjoys favorable scaling properties: the auto-interp score of the bottleneck concepts improves with data, as does the performance on downstream applications. Specifically, PCDs can detect jailbreaks, secret hints, and implanted latent concepts, and accurately surface latent user attributes.
PreDiff: Sequential Recommendation by Denoising Preference Distributions
Yaoqi Chen ⋅ Jianjin Zhang ⋅ Qi Chen ⋅ Weihao Han ⋅ Yujing Wang ⋅ Hao Sun
Sequential recommendation predicts which item a user will interact with next. A key property of this task is that user preferences are concentrated: only a small cluster of items is relevant at a given moment. Recent diffusion-based methods add noise to the item's embedding and score candidates by embedding similarity. Since similar items have similar embeddings, their score gaps are inherently small and easily overwhelmed by noise, causing the most relevant candidates to become indistinguishable first. We prove that moving diffusion to distribution space, where each item receives independent noise, preserves rankings better and makes the reverse denoising process easier. This shift, however, loses the semantic structure of embedding space and requires richer supervision than one-hot labels. We propose PreDiff (Preference Diffusion), which performs diffusion in distribution space and introduces two components to address these problems: Distribution-Guided Sparse Projection projects the preference scores into embedding space while filtering out low-relevance noise, and soft targets encode inter-item similarity to teach fine-grained ranking within the preference cluster. On four benchmarks, PreDiff achieves 8%--17% relative improvement over existing methods.
PREPING: Building Agent Memory without Tasks
Yumin Choi ⋅ Sangwoo Park ⋅ Minki Kang ⋅ Jinheon Baek ⋅ Sung Ju Hwang
Agent memory is typically constructed either offline from curated demonstrations or online from post-deployment interactions. However, regardless of how it is built, an agent faces a cold-start gap when first introduced to a new environment without any task-specific experience available. In this paper, we study pre-task memory construction: whether an agent can build procedural memory before observing any target-environment tasks, using only self-generated synthetic practice. Yet, synthetic interaction alone is insufficient, as without controlling what to practice and what to store, synthetic tasks become redundant, infeasible, and ultimately uninformative, and memory further degrades quickly due to unfiltered trajectories. To overcome this, we present Preping, a proposer-guided memory construction framework. At its core is proposer memory, a structured control state that shapes future practice. A Proposer generates synthetic tasks conditioned on this state, a Solver executes them, and a Validator determines which trajectories are eligible for memory insertion while also providing feedback to guide future proposals. Experiments on AppWorld, BFCL v3, and MCP-Universe show that Preping substantially improves over a no-memory baseline and achieves performance competitive with strong playbook-based methods built from offline or online experience, with deployment cost $2.99\times$ lower on AppWorld and $2.23\times$ lower on BFCL v3 than online memory construction. Further analyses reveal that the main benefit does not come from synthetic volume alone, but from proposer-side control over feasibility, redundancy, and coverage, combined with selective memory updates.
PRIME: A Modular Approach for Private Synthetic Data
Miguel Fuentes ⋅ Brett Mullins ⋅ Cecilia Ferrando ⋅ Cameron Musco ⋅ Daniel Sheldon
Generating differentially private (DP) synthetic tabular data remains a challenge, particularly when leveraging state-of-the-art query-answering mechanisms like ResidualPlanner or AIM+GReM. These mechanisms offer high utility through dense measurement collections $\mathcal{M}$; however, they are incompatible with existing synthesis tools. Specifically, graphical-model-based synthesis via PrivatePGM becomes intractable as the treewidth of the measurement graph grows. We introduce PRIME (Private Re-weighting of Initial Model Estimates), a scalable framework that decouples the synthesis process from the complexity of the measurement graph by restricting the support to a subset of the overall domain. PRIME transforms the synthesis problem into a convex re-weighting task over a candidate support set of size $K$. This formulation achieves an $\mathcal{O}(K \cdot |\mathcal{M}|)$ per-iteration complexity, providing a treewidth-independent, drop-in solution for any DP mechanism. We show that, when provided the same measurements, PRIME achieves error comparable to PrivatePGM while handling dense measurement sets where PrivatePGM fails. Our empirical evaluation demonstrates that PRIME enables high-fidelity synthesis from query-answering mechanisms that were previously unusable for data generation. Ablation studies identify support quality as a key performance driver, underscoring the role of support generation in our framework.
Principled Design of Diffusion-based Optimizers for Inverse Problems
Julio Oscanoa ⋅ Irmak Sivgin ⋅ Cagan Alkan ⋅ Daniel Ennis ⋅ John Pauly ⋅ Mert Pilanci ⋅ Shreyas Vasanawala
Score-based diffusion models achieve state-of-the-art performance for inverse problems, but their practical deployment is hindered by long inference times and cumbersome hyperparameter tuning. While pretrained diffusion models can be reused across tasks without retraining, inference-time hyperparameters such as the noise schedule and posterior sampling weights typically require ad-hoc adjustment for each problem setup. We propose principled reparameterizations that induce invariances, allowing the same hyperparameters to be reused across multiple problems without re-tuning. In addition, building on the RED-diff framework, which reformulates posterior sampling as an optimization problem, we further develop the OptDiff pipeline. OptDiff provides a simplified tuning framework that facilitates the integration of convex optimization tools to accelerate inference. Experiments on image reconstruction, deblurring, and super-resolution show substantial speedups and improved image quality.
Principled Policy Optimization for LLMs via Self-Normalized Importance Sampling
Huiyang Shao ⋅ Xintong Zhang ⋅ Junyi Liu
Reinforcement Learning from Human Feedback (RLHF) is a key technique for aligning Large Language Models (LLMs) with human preferences. While Proximal Policy Optimization (PPO) is the standard algorithm, its reliance on a critic network incurs significant memory and computational costs. This has motivated the development of critic-free alternatives such as Group Relative Policy Optimization (GRPO) and Group Sequence Policy Optimization (GSPO). However, these methods suffer from a critical trade-off: they either employ theoretically unsound, high-variance estimators (GRPO) or introduce systematic bias to achieve stability, causing them to optimize a perturbed objective (GSPO). In this paper, we introduce SNIB (Self-Normalized Importance Sampling with a Baseline), a novel critic-free algorithm that addresses this dilemma by offering a method that is both stable and asymptotically correct. SNIB leverages principled self-normalized importance sampling to achieve the stability of modern methods without sacrificing asymptotic correctness. We provide a comprehensive theoretical analysis, proving that SNIB's gradient estimator is consistent with a finite-sample bias that decays as $O(1/G)$ under self-normalization. Furthermore, we demonstrate its superior robustness to reward model uncertainty and show that it preserves the principled trade-off between reward maximization and KL regularization, a property that is distorted by biased estimators. Our work establishes a theoretically-grounded foundation for building more stable and reliable critic-free RLHF algorithms.
Prism: Harmonizing Missing Modalities via Implicit Structural Alignment on Lightweight Pulse RWKV for Multimodal Crack Segmentation
Hui Liu ⋅ Chen Jia ⋅ Fan Shi ⋅ Xu Cheng ⋅ Mianzhao Wang ⋅ Jiangpeng Zheng ⋅ Shengyong Chen
In multimodal crack segmentation for industrial facilities, the key challenge is preventing missing modalities from degrading pixel-level performance while keeping computation low. Existing methods struggle to harmonize missing modality effects and to efficiently, adaptively perceive cross-modal topology. We propose Prism, resilient to any modality miss, leveraging implicit structural alignment to harmonize missing data for high-quality segmentation at low computational cost. Prism comprises Modal Reconstructor (MoRe), Prompt-as-Knowledge Distillation (PaKD), a Sparse Gated Pulse Propagation Mixer (PulseMixer), and an Uncertainty-guided Connectivity Propagation Fusion module (UCPF). MoRe utilizes the manifolds of available modalities to probabilistically harmonize missing data, empowered by PaKD to master where to look via prompt distillation. PulseMixer efficiently breaks the static parameterization bottleneck and perceives anisotropic textures via Pulse Propagation Insight (PPI) and sparse gating. UCPF fuses cross-modal cues using uncertainty estimation to produce clear segmentation maps while suppressing background noise. Experiments on three multimodal crack datasets demonstrate strong performance under diverse modality-miss scenarios. On the depth dataset with 90\% missing depth modal, Prism achieves 0.8217 in F1 and 0.8478 in mIoU with only 2.62M parameters.
PriSM: Prior-guided Shared-basis Mixture Personalization for LLMs under Sparse User Histories
Hea Eun Lee ⋅ Hyungi Lee ⋅ Jangho Kim
Personalizing large language models is essential for user-centric applications, yet remains challenging under sparse user histories, privacy constraints, and limited on-device resources. Existing approaches either rely on prompting or retrieval without adapting model parameters, train and store separate adapters for each user, or use coarse group-level adapters that cannot fully capture individual variation. We propose $\textbf{PriSM}$ ($\textbf{Pri}$or-guided $\textbf{S}$hared-basis $\textbf{M}$ixture Personalization), a lightweight framework that models personalization as posterior-inspired update over mixtures of shared LoRA bases. PriSM first learns reusable basis adapters that capture group-level adaptation patterns, then uses a cluster-conditional Dirichlet prior and a user-conditioned hypernetwork to infer posterior mixture weights from sparse user profiles. The resulting posterior mean synthesizes a personalized LoRA update in a single forward pass, allowing the model to rely on group-level priors when user evidence is limited and to adapt toward user-specific behavior when sufficient evidence is available. This design enables scalable, privacy-preserving, and on-the-fly personalization without per-user training or per-user adapter storage. Code is available at \url{https://anonymous.4open.science/r/PriSM-D705}.
PRISM: Spectral Pruning and Reconstruction for Parameter-Efficient Model Merging of MLLMs
Wuxuan Shi ⋅ Haotian Chen ⋅ He Li ⋅ Qiang Yang ⋅ Mang Ye
Parameter-Efficient Fine-Tuning (PEFT) has become a dominant paradigm for adapting Multimodal Large Language Models (MLLMs) to specific tasks, resulting in a proliferation of task-specific expert models. Merging these experts into one universal model offers a promising path to creating versatile systems without costly retraining. However, existing methods for full fine-tuning merging often falter in parameter-efficient model merging, as they manipulate weights directly while ignoring the intrinsic geometric structure of PEFT modules. When diverse tasks exhibit conflicting feature directions, this inevitably leads to destructive interference. To address this issue, we propose PRISM, a data-free framework that reframes the merging of PEFT modules as a spectral signal reconstruction problem. PRISM operates in two stages: (1) Spectral Pruning, which decouples task updates into singular components and retains only high-energy directions to attenuate task-irrelevant noise; and (2) Energy-Prioritized Orthogonal Reconstruction, which prioritizes dominant components and projects overlapping vectors onto orthogonal directions to eliminate inter-task interference. We curate a benchmark comprising diverse multimodal tasks to evaluate our method. Extensive experiments demonstrate that PRISM significantly outperforms state-of-the-art baselines. Our code and weights will be released.
We consider online prediction from experts, a fundamental problem in machine learning, under differential privacy constraints. Existing private algorithms achieve near-optimal regret $\tilde{O}(\sqrt{T})$, where $T$ is the time horizon. However, this bound becomes suboptimal when $L^\star$, the cumulative loss of the best expert, is significantly less than $T$. In this work, we present the first differentially private algorithm with regret $\tilde{O}(\sqrt{L^\star})$, offering a substantial improvement in the low-loss regime. Our regret bound matches the non-private lower bound up to poly-logarithmic factors, demonstrating that privacy incurs only a small cost in this setting.
PROACT-Agent: Progressive Runtime Oversight and Active Circuit-breaking for Real-Time Safety
Ding Jia ⋅ Wei Liu ⋅ Xianglong Du ⋅ Yingjie Li ⋅ Yingqing Yang ⋅ Huili Yu ⋅ Zhangsong Zhan ⋅ Chu Zhou
The transition from Large Language Models (LLMs) to agents shifts safety stakes from toxic text to irreversible environmental harm. While current defenses remain largely retrospective, proactive runtime intervention is bottlenecked by the lack of large-scale, causally-consistent data. We propose PROACT-Agent, a framework for synthesizing high-fidelity trajectories to enable real-time guardrails. We identify a critical "safety drift" in prior benchmarks, where lenient annotation paradigms fail to enforce temporal consistency. PROACT-Agent addresses this through: (1) Progressive Trajectory Unrolling to reveal risks hidden in long-context interactions; (2) Reasoning-Augmented Causal Rectification to enforce monotonic causal consistency; and (3) Culturally-Aware Data Localization for cross-border robustness. We introduce PROACT-Bench, the largest bilingual safety benchmark to date, featuring 140,000+ trajectories adjudicated with industrial-grade rigor. Experiments show that models trained via our framework achieve superior zero-latency intervention, establishing a new standard for real-time autonomous agent safety.
ProDG: Prototypes for Data-Free Generative Post-Hoc Explainability
Piotr Borycki ⋅ Magdalena Trędowicz ⋅ Jacek Tabor ⋅ Łukasz Struski ⋅ Przemysław Spurek
Ante-hoc interpretability methods based on prototypes provide highly accurate explanations by utilizing the intuitive "this looks like that" reasoning paradigm. On the other hand, post-hoc models can explain predictions for a single image without relying on an underlying dataset or requiring costly neural network retraining. Recent approaches successfully solve the retraining problem for prototype-based networks. However, they still face a fundamental limitation: they require access to a subset of data (e.g., a test or validation set) to search for and extract the visual prototypes. In this paper, we address this issue and introduce ProDG: Data-Free Generative Prototypes for Post-Hoc Explainability, a novel framework that leverages generative models to synthesize pure, high-fidelity prototypes directly from the frozen model's weights, completely eliminating the dependency on any external data. By establishing this new frontier in Data-Free XAI, DGP unlocks robust visual interpretability for privacy-sensitive domains, where original data is strictly restricted or fundamentally inaccessible.
Programmatic Reasoning with Structural Schema: A Unified Framework for Multi-Table Inference
Jialin Chen ⋅ Brandon Mayer ⋅ Michael Galkin ⋅ Sami Abu-El-Haija ⋅ Rex Ying ⋅ Bryan Perozzi
Large language models (LLMs) have shown strong performance in single-table reasoning but struggle to generalize to multi-table scenarios due to schema heterogeneity, long input contexts, and the lack of structural awareness. We introduce StrucTab-R1, a schema-centric reasoning framework that decouples relational reasoning from raw data exposure. Our approach encodes relational databases as heterogeneous schema graphs, where tables, columns, and foreign-key constraints form distinct node and edge types, and employs a heterogeneous graph encoder with question-conditioned cross-attention pooling to distill compact, semantically grounded schema tokens. Rather than serializing table rows into the context, the LLM plans over the schema representation and emits a Chain-of-Execution: an executable reasoning trace that first identifies the relevant schema subgraph and then performs a sequence of composable tool-function calls (e.g., filter, join, aggregate). These calls are executed externally, with compact observations returned to support subsequent reasoning steps. This design confines the model's attention to relevant subgraphs, mitigating hallucinations in multi-hop join reasoning. To further improve reliability, we train the model with supervised traces followed by structure-aware reinforcement learning, where the reward jointly optimizes answer correctness, schema-region precision, and execution consistency. Experiments on single-table and multi-table benchmarks show that StrucTab-R1 improves execution accuracy over strong general-purpose and table-specialized LLMs and remains effective on large-schema settings where raw-context baselines fail.
Progressive Layer-wise Supervision: Deep-to-Shallow Supervision Annealing for Efficient and Robust Speech Deepfake Detection
Hoan My Tran ⋅ Xin Wang ⋅ Xuechen Liu
Speech foundation models pretrained via self-supervised learning provide powerful representations for deepfake detection, yet we find that standard fine-tuning exploits only the deepest transformer layers — leaving the majority of the network's capacity discriminatively inert. Layer-wise probing of \textit{XLS-R-128} reveals that the top two-thirds of the transformer stack contribute negligibly to anti-spoofing performance after standard fine-tuning, a phenomenon we term shallow-layer dormancy. We connect this observation to the representational redundancy of pretrained transformers: adjacent layers encode highly similar information, and fine-tuning without explicit intermediate supervision reinforces rather than breaks this redundancy. To address this, we propose Progressive Layer-wise Supervision (PLS), a training framework that attaches a shared linear classifier to every transformer layer and schedules layer-wise loss contributions via an exponentially annealed power weighting. Early training concentrates supervision on deep layers to preserve the pretrained representational hierarchy; later training progressively redistributes supervision to shallower layers, converting dormant earlier representations into discriminative ones. Competitive detection ($\leq9.25\%$ average EER) is achievable on \emph{XLS-R-128} at layer 7, which reduces the parameter count from 315M to 101M and exemplifies an early exit mechanism. Experiments across twelve benchmarks spanning in-domain, and out-of-domain deepfake scenarios show PLS achieves a new state-of-the-art with 239M parameters, outperforming prior methods with larger and more complex architectures.
Progressive Memory Transformer: Memory-Aware Attention for Time-Series
Tord S Stangeland ⋅ Andreas Köhler ⋅ Steffen Maeland ⋅ Adín Ramírez Rivera
Time-series carry structure simultaneously at multiple scales (fine-grained variation, mid-range motifs, and global properties) and downstream tasks operate at correspondingly different scales. Most existing self-supervised learning approaches supervise representations globally via instance-level contrastive losses and limited temporal neighborhood supervision, but do not explicitly exploit the structural hierarchy. We propose a learning framework that explicitly enforces a structural hierarchy across three scales independently: a local objective for token continuity, a mid-range objective for window-level motifs, and a global objective for sequence-level agreement. Realizing this framework requires the backbone to expose a representation at each scale; we introduce \textbf{Progressive Memory Transformer} (PMT), which augments a transformer with writable, window-aligned memory that exposes the mid-range scale alongside the token and sequence-level representations conventional transformers already provide. Across seven UCR/UEA/UCI classification benchmarks, a cue-retention probe, and two forecasting benchmarks, PMT learns representations that probe well at the global, mid-range, and local scales---strong low-label classification (1--5\% labels), competitive forecasting performance across multiple horizons, and quantitative and qualitative evidence that memory states capture mid-range motifs.
Progressive Pseudo-label Self-balancing Towards Unsupervised Vision-Language Models Adaptation
Si Qin ⋅ Yaxin Hou ⋅ Jiawei Tang ⋅ Yuheng Jia
Enhancing Vision-Language Models (VLMs) on downstream tasks with unlabeled data has attracted increasing attention. Constructing a pseudo-labeled dataset by pseudo-labeling is a common method. Nonetheless, owing to the inherent class-wise prediction bias in VLMs, they tend to generate long-tailed pseudo-labels that are inconsistent with the true label distribution. Motivated by this, we first divide all classes into pseudo-head, pseudo-middle, and pseudo-tail classes based on the pseudo-label quantity. Accordingly, we propose a Progressive Pseudo-label Self-Balancing (PPSB) framework, which measures the class-wise prediction bias and calculates quantity upper bounds for certain classes to construct relatively balanced pseudo-labeled dataset, while introducing fewer incorrect pseudo-tail labels. In this way, the class-wise prediction bias is progressively corrected and the pseudo-label dataset scale increases autonomously during training. Furthermore, the confidence scores generated by VLMs are not entirely reliable. Therefore, we design a neighborhood consistency filtering mechanism and a visual classifier to improve the pseudo-label accuracy. We theoretically prove that our method can significantly reduce the generalization error under certain conditions. Extensive experiments across six benchmarks under three learning paradigms demonstrate that our method outperforms state-of-the-art methods by average 3.84% in accuracy. Code is available in the supplementary materials.
ProteinOPD: Towards Effective and Efficient Preference Alignment for Protein Design
Yulin Zhang ⋅ He CAO ⋅ Zihao Jiang ⋅ Chenyi Zi ⋅ Zhipeng Zhou ⋅ Zijing Liu ⋅ Yu Li ⋅ Jia Li ⋅ Ziqi Gao
Designing proteins with desired functions or properties represents a core goal in synthetic biology, therapeutic development, and drug discovery. Recent advances in protein language models (PLMs) have enabled the generation of highly designable protein sequences, while preference alignment provides a promising way to steer designs toward desired functions and properties. Nevertheless, they often trigger catastrophic forgetting of pretrained knowledge, degrading basic designability and failing to balance multiple competing objectives. To address these issues, we draw inspiration from On-Policy Distillation (OPD), an advanced post-training method renowned for mitigating catastrophic forgetting through its mode-seeking nature. In this work, we propose ProteinOPD, a multi-objective preference alignment framework that can effectively balance multiple preference objectives while maintaining the inherent designability of PLMs. ProteinOPD adapts a pretrained PLM into preference-specific teachers and distills their knowledge into a shared student via token-level OPD on the student’s own trajectories. During this process, the student is aligned to a unique normalized geometric consensus of weighted teachers while ensuring bounded optimization under conflicts. This bridges the gap for OPD in multi-objective/teacher alignment. Extensive experiments show that ProteinOPD achieves substantial gains on target preference objectives without compromising the designability, with a 8× training speedup over RL-based alignment competitors.
PROTEUS: A Self-Evolving Red Team with Surface Expansion for Agent Skill Ecosystems
Zhaojiacheng Zhou ⋅ Jiong Lou ⋅ Kaixiang Wang ⋅ Yanzhi Li ⋅ Hefeng Zhou ⋅ Jie LI
Agent skills extend LLM agents with reusable instructions, tool interfaces, and executable code, and users increasingly install third-party skills from marketplaces, repositories, and community channels. Because a skill exposes both executable behavior and context-setting documentation, its deployment risk cannot be measured by single-shot audits or prompt-level red teams alone: a realistic attacker can use audit and runtime feedback to repeatedly rewrite the skill. We frame this risk as \textbf{adaptive leakage}—whether a budgeted attacker can iteratively revise a skill until it passes audit and produces verified runtime harm—and present \textbf{Proteus}, a grey-box self-evolving red-team framework for measuring it. Proteus searches a formalized five-axis skill-attack space. Each candidate is evaluated through a unified audit-sandbox-oracle pipeline that returns structured audit findings and runtime evidence to guide cross-round mutation. Beyond initial evasion, Proteus performs path expansion, which finds alternative implementations of successful attacks, and surface expansion, which transfers learned implementation patterns to new attack objectives beyond the original seed catalogue. Across eight phase-1 mutator-target-defender configurations, Proteus achieves 40-90\% \textbf{ASR@5} and exhibits positive learning-curve slopes on both evaluated auditors. In the full 8-cell expansion matrix, Proteus generates 438 jointly bypassing and lethal variants; SkillVetter is bypassed at $\geq$ 93\% in every expansion cell, and AI-Infra-Guard, the strongest public auditor we evaluate, still admits jointly successful variants at up to 41.3\%. These results show that current skill vetting substantially underestimates residual risk when evaluated against adaptive, feedback-driven attackers. Code: \url{https://anonymous.4open.science/r/proteus/}.
3D Gaussian Splatting (3DGS) enables high-quality real-time novel-view synthesis, but practical scenes often contain millions of Gaussians, making compression essential for deployment on limited hardware. Existing reduction methods are effective but mostly heuristic: they provide no multiplicative approximation guarantee for the rendered objective, and thus rely heavily on costly post-pruning finetuning to recover quality. We ask a basic question: can a 3DGS scene be provably replaced by a much smaller weighted subset (coreset) while preserving the objective of interest? We first show that, in the unrestricted setting, no non-trivial multiplicative 3DGS coreset exists. We then show that multiplicative guarantees are not impossible, but resolution-dependent. For a prescribed rendering resolution, such as representative views or grids of views/rays, we provide the first weighted coreset construction theorem for 3DGS. The construction samples Gaussians by sensitivity: provable importance scores measuring each Gaussian’s role in the full-scene objective. Finally, under explicit validity and log-transmittance stability assumptions, we turn this objective guarantee into a rendering guarantee. Empirically, our method is strongest where deployment needs it most: aggressive compression with no or minimal recovery compute. In prune-only and very short finetuning regimes, it achieves state-of-the-art performance, showing that principled importance estimation can be both theoretically meaningful and practically useful.
Provably Reliable Classifier Guidance via Cross-Entropy Control
Sharan Sahu ⋅ Arisina Banerjee ⋅ Yuchen Wu
Classifier-guided diffusion models generate conditional samples by augmenting the reverse-time score with the gradient of the log-probability predicted by a probabilistic classifier. In practice, this classifier is usually obtained by minimizing an empirical loss function. While existing statistical theory guarantees good generalization performance when the sample size is sufficiently large, it remains unclear whether such training yields an effective guidance mechanism. We study this question in the context of cross-entropy loss, which is widely used for classifier training. Under mild smoothness assumptions on the classifiers, we show that controlling the cross-entropy at each diffusion model step is sufficient to control the corresponding guidance error. In particular, we show that probabilistic classifiers achieving conditional KL divergence $\varepsilon^2$ induce guidance vectors with mean squared error $\widetilde O(d \varepsilon )$, up to constant and logarithmic factors. We demonstrate that the proposed smoothness condition is necessary, by constructing a sequence of non-smooth classifiers that achieve small conditional KL while inducing an exploding guidance error. We also show that the derived upper bound $\widetilde O(d \varepsilon )$ is optimal up to poly-logarithmic factors. Our result yields an upper bound on the sampling error of classifier-guided diffusion models and bears resemblance to a reverse log-Sobolev--type inequality. To the best of our knowledge, this is the first result that quantitatively links classifier training to guidance alignment in diffusion models, providing both a theoretical understanding of when such models succeed, along with principled guidelines for selecting classifiers that induce effective guidance. In particular, our findings suggest that, when selecting classifiers for guidance, one should prioritize not only accuracy but also smoothness, highlighting the advantages of smoothness-inducing training methods for diffusion guidance.
Safe reinforcement learning (RL) aims to learn policies that optimize rewards while satisfying constraints. Predominant approaches rely on soft-constrained policy optimization, which has achieved empirical success but does not provide formal safety guarantees for the learned policy. In contrast, methods with strict guarantees typically rely on explicit certificate functions, whose construction requires the direct synthesis and verification of control-invariant sets, a process that scales poorly with state dimension and often yields overly conservative behavior. In this paper, we present the Provably Safe, yet Performant RL (PSP-RL) framework, a novel two-phase architecture for learning provably safe policies in a scalable manner, designed to overcome the key bottlenecks of prior methods. Rather than explicitly computing invariant sets, PSP-RL leverages a learned backup policy to forward-integrate the system dynamics, generating an implicit control-invariant set online. In the first phase, the backup policy is trained with our proposed safe-arrival value function, which characterizes the optimal backup policy for invariant-set construction. In the second phase, an RL policy is trained end-to-end through a differentiable projection layer that strictly enforces the safety guarantees induced by the learned backup policy. By maximizing the volume of the implicit control-invariant set in the first phase, the resulting PSP policy from the second phase is performant and scalable, while maintaining provable safety. Crucially, PSP-RL imposes no restrictions on the underlying RL algorithm and can be plugged into any existing training pipeline. We establish theoretical guarantees for the proposed framework and evaluate it on robotic control tasks with state dimensions up to 10, a regime in which prior provably safe RL methods struggle or become impractical.
PULSE: Identifying Demonstration-Utility Features with Sparse Autoencoders
Chenduo Hao ⋅ Chuanbao Gao ⋅ Pinjun Zeng ⋅ Jingze Zhu ⋅ Liu Chonghan ⋅ Zidong Liu ⋅ Xu Yang
In-context learning is highly sensitive to demonstration choice, yet most methods select demonstrations using external query--demonstration similarity. Such criteria can miss model-specific signals: Similar demonstrations may activate different internal features and downstream behaviors. We introduce PULSE (Paired Utility Localization over Sparse Encodings), an SAE-based framework for identifying model-internal features associated with demonstration utility and using them for demonstration selection. Using a small labeled discovery set, PULSE samples candidate demonstration sets, measures their zero-shot-relative utility under the target model, and scores SAE features by how their activation differences align with utility differences. The top positive and negative coordinates form a sparse utility-localization vector. We use this vector in two complementary ways: as a signed score for controlled complete-set ranking, and as PULSE-Retriever, which converts its magnitude into a feature-relevance mask for scalable pool-scale retrieval. Across classification, generation, and reasoning benchmarks, PULSE-Retriever improves over the strongest baseline by 2--3 accuracy points, 0.6--0.9 BLEU-4, and 3.2 exact-match points, respectively, while controlled ranking validates the identified features encode a predictive set-level utility signal. Feature inspection and cross-dataset experiments suggest that the identified features capture task-relevant, dataset-conditioned patterns, yet retain utility signals that partially transfer across datasets. Our code is available at the anonymous repository.
Q-Focus: Let the Question Guide What to See in Long Videos
Junbo Qiao ⋅ Xinning Chai ⋅ Dong Fang ⋅ Sifa Xie ⋅ Wenxuan Huang
Video Large Language Models face severe computational and memory bottlenecks in long-video understanding. Existing visual token compression methods predominantly rely on static, query-agnostic redundancy elimination. Consequently, they lack the ability to dynamically preserve query-relevant details, failing to emulate the human "coarse-to-fine" cognitive process. To bridge this gap, we propose \textbf{Q-Focus}, a plug-and-play, two-stage visual token compression framework. Q-Focus first constructs a holistic semantic representation via a global compression stage. This is followed by a question-guided focusing stage driven by two novel components: the Adaptive Contrastive Module (ACM) and the Question Context Module (QCM). Specifically, ACM achieves adaptive query-visual alignment by treating the user queries as positive samples and dynamically sampling negative samples from the video's feature distribution. Complementarily, QCM leverages visual affinity matrices to diffuse semantic relevance from highly matched visual anchors to surrounding contextual regions, effectively capturing implicit yet critical details. Extensive experiments demonstrate that Q-Focus seamlessly integrates with diverse compression baselines. Remarkably, while retaining only 10\% of the original visual tokens, it achieves relative accuracy gains of 7.3\% and 6.6\% on LLaVA-OneVision and LLaVA-Video, respectively, and consistently boosts the performance across other state-of-the-art models such as Qwen3-VL-8B-Instruct.
Q-Residual Physics: Hamiltonian-Structured Quantum Residual Learning for Embodied Dynamics
Fanqi Kong ⋅ Ke Shi ⋅ Yuchen Wang ⋅ Zhipeng Liu ⋅ Xiaofei Yue ⋅ Zhaoxuan Li ⋅ Tingting Li ⋅ Ziming Zhao
Embodied agents in contact-rich environments must model dynamics affectedby frictional uncertainty, contact discontinuities, compliance, actuator imperfec-tions, perception noise, and mass or inertia mismatch. Existing residual dynamicslearning methods correct nominal physics predictions with unconstrained neuralresiduals, which often generalize poorly under unseen contact regimes, changingphysical parameters, long-horizon rollouts, and few-shot adaptation. We proposeQ-Residual Physics, a Hamiltonian-structured quantum residual learning frame-work for embodied dynamics. The method represents residual errors through astructured quantum residual layer whose qubits correspond to physically inter-pretable modes, including contact impulse, friction, slip, compliance, damping,perception error, and mass or inertia mismatch. A physical interaction graphinduces a Hamiltonian residual structure, implemented with a Trotterized param-eterized quantum circuit whose components capture mode activation, physicalcoupling, uncertain switching, dynamic mode exchange, and higher-order interac-tions. Quantum measurements are decoded to correct nominal physics predictions.Across embodied dynamics and manipulation benchmarks, including OMNIPUSH,MANISKILL2, ROBOMIMIC, D4RL, and PHYSION, Q-Residual Physics achievesstronger generalization than physics-only, neural residual, graph residual, andquantum-circuit baselines. It reduces long-horizon rollout error by 23.8%, im-proves contact transition prediction by 16.4%, lowers out-of-distribution parametererror by 21.7%, and improves few-shot adaptation by 28.5%. The code is availableat https://anonymous.4open.science/r/QRP-0205
Quantile-Coupled Flow Matching for Distributional Reinforcement Learning
Michael Groom ⋅ Victor-Alexandru Darvariu ⋅ Lars Kunze ⋅ James Wilson ⋅ Nick Hawes
Unlike standard expected-return Reinforcement Learning (RL), Distributional RL (DRL) models the full return distribution, making it better-suited for uncertainty-aware and risk-sensitive decision-making. Conditional Flow Matching (CFM) critics have recently attracted attention for modelling continuous, multi-modal return distributions. Despite this interest, there remains a substantial \textbf{metric mismatch}: DRL theory relies on the distributional Bellman operator being contractive in the $p$-Wasserstein distance, yet existing CFM critics are trained with arbitrary source--target couplings, so their flow-matching losses are not Wasserstein-aligned surrogates for matching Bellman target return distributions. In this work, we address this mismatch by proposing \textbf{FlowIQN}, a CFM critic that sorts source and Bellman target samples within each mini-batch to approximate the monotone optimal transport coupling, replacing arbitrary pairings with quantile-aligned flow paths. We prove that the loss of our \textbf{quantile-coupled} CFM critic yields a Wasserstein-aligned approximate projection compatible with the foundations of DRL. To our knowledge, FlowIQN is the first flow-matching distributional critic with an explicit Wasserstein-aligned projection guarantee. We further extend FlowIQN with shortcut models for efficient inference. Empirical results show that FlowIQN improves Wasserstein return-distribution accuracy over other CFM critics. It also yields competitive performance on offline RL benchmarks across multiple policy extraction methods, providing a theoretically grounded CFM critic that is readily compatible with DRL pipelines.
Quantile Geometry Regularization for Distributional Reinforcement Learning
Zhaofan ZHANG ⋅ Minghao Yang ⋅ Rufeng Chen ⋅ Sihong Xie ⋅ Hui Xiong
Quantile-based distributional reinforcement learning methods learn return distributions through sampled quantile regression, but their bootstrapped target quantiles may induce distorted or degenerate distribution estimates. We propose Robust Quantile-based Implicit Quantile Networks (RQIQN), a lightweight Wasserstein distributionally robust enhancement boosted from a quantile estimation perspective. We first reinterpret a snapshot of IQN loss as a collection of local empirical quantile estimation problems over sampled current fractions. We then robustify each local slot with a Wasserstein distributionally robust quantile estimation formulation, yielding a closed-form, fraction-dependent correction to the Bellman target. This correction directly mitigates distributional degeneration: its median-antisymmetry preserves the risk-neutral quantile average, while its monotonicity enlarges upper--lower quantile gaps and counteracts collapsed distributional spread. RQIQN thus regularizes quantile geometry without changing the underlying value objective or requiring additional sample-set reconstruction. Finally, we empirically show that the proposed RQIQN outperforms other existing quantile-based distributional reinforcement learning algorithms in risk-sensitive navigation and Atari games.
Query-Limited Community Recovery in Stochastic Block Models
Sabyasachi Basu ⋅ MANUJ MUKHERJEE ⋅ Lutz Oettershagen ⋅ Suhas Thejaswi
We study exact community recovery in the two-community stochastic block model on $n$ vertices under limited and noisy access to network data. The learner may query a noisy neighborhood oracle that reveals each true neighbor of a queried vertex independently with fixed probability and never returns non-neighbors, subject to a finite query budget. We consider both oracle-only access and a combined model where the learner also observes a single subsampled copy of the underlying graph. For oracle-only access, balanced uniform querying gives a sharp non-adaptive benchmark: when each vertex is queried the same integer number of times, the observations reduce to an SBM with attenuated edge probabilities and the Abbe--Bandeira--Hall exact-recovery threshold applies. We show that this benchmark is not adaptively optimal: a two-stage adaptive strategy succeeds with $n+o(n)$ queries in a regime where balanced uniform querying requires $m n$ queries for some $m>1$. With an additional subsampled graph, we prove a sublinear-query adaptivity gap: balanced data-independent uniform querying with a sublinear budget does not improve over the subsampled graph alone, whereas adaptive querying can target a small set of uncertain vertices and achieve exact recovery. Thus adaptive data acquisition can strictly improve the information-theoretic limits of exact recovery.
RadarFlowPose: Vision-Inspired Coarse-to-Fine Skeleton Refinement with Flow Matching
Kailu Guo ⋅ Akram Alomainy ⋅ Khalid Z Rajab
Radar-based human pose estimation is a promising alternative to camera-based perception, but remains challenging due to sparse observations, noisy measurements, and severe pose ambiguity. Existing methods infer both global skeletal structure and flexible peripheral joints from sparse radar observations, often compromising structural plausibility and fine-grained local accuracy. Motivated by the coarse-to-fine nature of human visual perception, we propose RadarFlowPose, a framework that decomposes radar-based human pose estimation into coarse pose prior generation and flow-based refinement. The coarse model first produces a structurally plausible and temporally coherent pose prior from sparse radar observations, providing a global initialization for subsequent refinement. Instead of estimating the full pose from scratch, the refinement stage focuses on correcting the remaining local ambiguities and motion-induced errors. By incorporating localised Doppler motion cues into the flow-matching process, the model performs motion-aware residual refinement around the coarse prior, leading to more accurate and anatomically consistent pose estimates. Experiments on three benchmark radar pose datasets show consistent effectiveness, especially in motion-sensitive and fine-grained joint-level evaluation.
Machine learning methods rely on data. However, gathering suitable data can be challenging due to availability constraints, cost, or the need for domain expertise. Expanding datasets with additional sources is a common response to limited data, yet this practice does not always improve downstream performance and can sometimes lead to a loss of performance, known as negative transfer. We propose RADAR, a simple, geometrically grounded metric for estimating cross-domain transferability in foundation models. RADAR analyzes the layer-wise evolution of representations by measuring angular alignments and relative changes in distance along layer-to-layer displacement trajectories, and by comparing empirical distributions of within-domain and cross-domain dynamics. We hypothesize that domain transferability is related to the divergence between these trajectory distributions. We evaluate the metric across multiple modalities, including cross-lingual sentiment classification with text embedding models and cross-domain image classification with foundation vision models. Across several settings, RADAR provides competitive predictive performance relative to existing transferability metrics on several vision and text benchmarks, with particularly strong results when domain transitions are smooth or cleanly separated. Our ablations further suggest that the effectiveness of transferability estimation depends on the geometry of the model’s internal representation space, with different modalities favoring different topological formulations.
Domain generalization methods train on multiple source environments but usually deploy a single pooled predictor. We study a different deployment object: one head per source environment on a shared representation, combined by fixed weights when the test environment label is unavailable. Under a same-meta random-effects model, where environment-specific population minimizers are iid deviations around a shared meta-mean, the optimal fixed simplex rule is the uniform centroid. The same analysis separates this centroid from pooled ERM: under a common feature second moment, pooled ERM is sample-weighted, while the centroid is environment-weighted, giving a closed-form gap controlled by sample-weight imbalance. We introduce REC (Random-Effects Centroid), a training objective that fits source-specific experts while directly scoring their averaged deployment head. In the common-second-moment case, the centroid term controls the trained-centroid discrepancy and the dispersion of environment-specific optima. Synthetic experiments isolate the population estimator effect and learned-representation behavior; WILDS experiments show that REC is competitive with established DG baselines and that the centroid term is necessary for a useful averaged predictor.
RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution
Siyonng Jian ⋅ Siyuan Li ⋅ Luyuan Zhang ⋅ Zedong WANG ⋅ Xin Jin ⋅ Ying Li ⋅ Cheng Tan ⋅ Huan Wang
Discrete autoregressive (AR) text-to-image (T2I) models adopt a two-stage paradigm in which a VQ tokenizer maps images to discrete codes and an AR policy models their distribution. Current post-training methods optimize only the AR policy while keeping the VQ decoder frozen. We show that this practice introduces Latent Covariate Shift: as the policy evolves, its token distribution progressively diverges from the ground-truth distribution on which the decoder was trained, such that reward scores improve while decoded image quality degrades. To address this mismatch, we propose RankE, the first end-to-end post-training framework for discrete T2I generation. Rather than optimizing the policy against a fixed decoder, RankE co-evolves the policy and decoder through alternating optimization: the policy is refined via group-relative preference optimization, while the decoder is jointly adapted via reward-aware adversarial training. This co-evolution ends the fidelity--alignment trade-off that plagues frozen-decoder approaches: on LlamaGen-XL (775M), standard RL improves CLIP but degrades FID, whereas RankE simultaneously improves both (FID 15.21, CLIP 33.76 on MS-COCO 30K). Consistent joint gains on Janus-Pro (1B) confirm that decoder co-evolution reliably converts reward optimization into pixel-space quality improvements.
The most widely used RANSAC variants score candidate models by counting inliers or summing truncated likelihoods; every such score requires a user-supplied parameter that is a function of the inlier scale, which must itself be estimated from contaminated data. We remove this dependence by reversing the usual order of inference: for a fixed inlier partition we marginalize $\sigma$ analytically in closed form under a conjugate Inverse-Gamma prior, then optimize over partitions. A single closed-form expression spans the non-informative Jeffreys limit (which requires no validation data to fit the prior) and informative empirical-Bayes priors fit from a small validation set, so the same score adapts across data-rich and data-scarce regimes without any change to the algorithm. To our knowledge this is the first RANSAC score in which the inlier scale is genuinely absent from the score formula. The score admits $O(N \log N)$ computation via sort-and-sweep. On a benchmark of nearly $70\,000$ image pairs spanning different two-view estimation problems and both engineered and learned feature pipelines, the proposed score matches or exceeds the state of the art (RANSAC, MSAC, GaU, MAGSAC++), excelling in robustness to hyperparameter miscalibration, sample efficiency at small validation budgets, and adaptive regularization across data-rich and data-scarce regimes.
RAPTOR: Ridge-Adaptive Logistic Probes
Ziqi Gao ⋅ Yaotian Zhu ⋅ Qingcheng Zeng ⋅ Xu Zhao ⋅ Ziqing Wang ⋅ Feng Ruan ⋅ Kaize Ding
Probing studies what information is encoded in a frozen LLM's layer representations by training a lightweight predictor on top of them. Beyond analysis, probes are often used operationally in probe-then-steer pipelines: a learned concept vector is extracted from a probe and then injected via additive activation steering by adding it to a layer representation during the forward pass. The effectiveness of this pipeline hinges on estimating concept vectors that are accurate, directionally stable under ablation, and inexpensive to obtain. Motivated by these desiderata, we propose RAPTOR (Ridge-Adaptive Logistic Probe), a simple $\ell_2$-regularized logistic probe whose validation-tuned ridge strength yields concept vectors from normalized weights. Across extensive experiments on instruction-tuned LLMs and human-written concept datasets, RAPTOR matches or exceeds strong baselines in accuracy while achieving competitive directional stability and substantially lower training cost; these quantitative results are supported by qualitative downstream steering demonstrations. Finally, using the Convex Gaussian Min-max Theorem (CGMT), we provide a mechanistic characterization of ridge logistic regression in an idealized Gaussian teacher-student model in the high-dimensional few-shot regime, explaining how penalty strength $\lambda$ mediates probe accuracy and concept-vector stability, yielding structural predictions that qualitatively align with trends observed on real LLM embeddings.
Rationale-Guided Policy Optimization: Learning to Reason with Adaptive Rationale Scaffolding
Hoang Phan ⋅ Minh Pham ⋅ Chau Pham ⋅ Chinmay Hegde ⋅ Trung Le ⋅ Qi Lei
On-policy reinforcement learning has become a central paradigm for improving the reasoning abilities of large language models. However, its effectiveness is often limited by reward sparsity: when a model fails to discover correct trajectories for difficult problems, the optimization process receives little useful signal and may stagnate. Existing approaches mitigate this issue by incorporating off-policy demonstrations, expert traces, or model-generated solutions, but they typically require the auxiliary data to match the format of the reinforcement-learning task, often relying on rejection sampling from stronger models to obtain suitable training trajectories. We introduce Rationale-Guided Policy Optimization (RGPO), a framework that adaptively leverages ground-truth rationale information according to the model’s current capability while preserving its freedom to explore. Rather than treating reference solutions as fixed imitation targets, RGPO uses them as temporary scaffolds: rationales help the model generate improved responses, after which only higher-reward, model-generated solutions are transferred back to the original unguided setting. This design allows training to exploit available ground-truth information without requiring off-policy data to follow the same format as the RL task. Across both language-only and vision-language reasoning settings, RGPO consistently improves performance over RLVR baselines, and ablation studies show that adaptive rationale guidance is a key contributor to these gains. These results suggest that RGPO offers a practical and general approach for reducing reward sparsity, stabilizing reinforcement learning, and improving reasoning performance in both text-only and multimodal models.
RAWild: Toward Sensor-Agnostic RAW Object Detection via Physics-Guided Curve and Grid Modeling
Shuhong Liu ⋅ Gengjia Chang ⋅ Jun Liu ⋅ Xuangeng Chu ⋅ Yinqiang Zheng ⋅ Tatsuya Harada ⋅ Ziteng Cui
Camera sensor RAW data offers intrinsic advantages for object detection, including richer bit depth, preserved physical information, and freedom from image signal processor (ISP) distortions. However, varying exposure conditions, spectral sensitivities, and bit depths across devices introduce substantially larger domain gaps than sRGB, making sensor-agnostic generalization a fundamental challenge. In this study, we present a physics-guided global-local tone mapping framework for sensor-agnostic RAW object detection. By factoring sensor-induced variations into a global tonal correction and a spatially adaptive local color adjustment, both driven by RAW distribution priors, our framework enables a single network to train jointly across heterogeneous sensors. To further support cross-sensor generalization, we construct a physics-based RAW simulation pipeline that synthesizes realistic sensor outputs spanning diverse spectral sensitivities, illuminants, and sensor non-idealities. Extensive experiments across multiple RAW benchmarks, covering bit depths from 12 to 24, demonstrate state-of-the-art (SOTA) performance in single-dataset, mixed-dataset, and challenging robustness settings.
RCSTAT: A Statistical Framework of Relative Contextualization in Transformers
Debabrata Mahapatra ⋅ Shubham Agarwal ⋅ Apoorv Saxena ⋅ Subrata Mitra
Identifying which tokens and activations truly influence model predictions is critical for both the efficiency and interpretability of large auto-regressive language models. Yet in practice, token importance is typically inferred from attention distributions that entangle contextual relevance with softmax-normalization effects, obscuring relative influence across tokens and heads. We propose $\textbf{RCStat}$, a statistical framework for quantifying contextual influence in attention mechanisms. At its core $\textbf{Relative Contextualization (RC)}$ is a random variable that measures how strongly one subset of tokens contributes to another under the model’s internal scoring scheme. RCStat admits computationally efficient bounds on expected influence that can be estimated at inference time. No retraining is required. We apply RCStat to two tasks. For $\textbf{attribution}$, attention heads with high expected RC accurately identify the prompt spans that drive generation, enabling reliable span-level explanations. For $\textbf{key-value cache eviction}$, RC-based adaptive thresholding selectively evicts low-impact KV entries, substantially reducing cache size while preserving generation quality. Across question answering, summarization, and attribution benchmarks, RCStat achieves consistent gains, improving generation quality by 15–40\% with upto 36\% error reduction for KV eviction, and attribution accuracy by 2–16\%,. These results demonstrate that explicitly modeling contextual influence provides a principled and practical alternative to attention-based heuristics.
REACT: Physically and Chemically Consistent Reconstruction of Marine Active Tracers
Wenbin Dai ⋅ Hao Zheng ⋅ Shiyu Liang ⋅ Chaofan Sun ⋅ Xueying Zhang ⋅ HanBo Huang ⋅ Xuan Gong ⋅ Yiran Zhang ⋅ Enhui Liao
Reconstructing global sea surface pH from sparse observations is critical for monitoring ocean acidification and understanding marine carbon cycling. Traditional assimilation and inverse models are physically grounded but costly for large-scale reconstruction. Recent black-box and physics-guided AI models improve efficiency, but are mainly designed for passive tracers, where the reconstructed variable is also the transported inventory. In contrast, pH is an active carbonate tracer: it is the prediction target, while dissolved inorganic carbon (DIC) is the conserved carbon inventory. This mismatch can produce low pH error while violating carbonate closure and source-free carbon conservation. To address this, we introduce \textbf{REACT}, a carbon-first reconstruction framework that decouples transport, active correction, and chemical decoding. REACT transports a latent carbonate state with a conservative advection--diffusion solver, captures non-conservative carbon-cycle variations with a source module, decodes the corrected state into pH, and constrains the output through carbonate equilibrium. This design keeps pH as the target while enforcing consistency on the underlying carbon state. On simulation data, REACT reduces pH NRMSE by (14.7\%) and chemical consistency error by (24.0\%) over the best baseline. Cross-temporal-scale evaluations show robustness against error accumulation from coarse to fine temporal scales, and ablation studies validate the effectiveness of each component.
Reading Attribution from Attention: Evidence Heads as Latent Attribution Mechanisms in LLMs
Qi Wu ⋅ Jianfeng Qu ⋅ Peng-Fei Zhang ⋅ Yanzhe Ji ⋅ Siyu Li ⋅ Wei Chen ⋅ Zhixu Li ⋅ Jiajie Xu
Large language models (LLMs) show strong performance in multi-document question answering, but their practical deployment is limited by unreliable and unfaithful attribution to supporting evidence. Existing prompting and training-based methods often suffer from hallucinated citations and lack interpretability in how evidence is selected. In this work, we investigate whether attribution signals are inherently encoded within transformer attention mechanisms. We introduce a sensitivity-based diagnostic that identifies a small subset of attention heads, termed Evidence Heads, which are highly responsive to perturbations in supporting documents. Through causal interventions and semantic analysis, we show that these heads play a significant role in evidence identification and exhibit alignment with document-level entailment signals. Building on this finding, we propose a training-free Attention-based Attribution framework that extracts evidence signals directly from Evidence Head attention using global and local strategies. Extensive experiments show that our method consistently outperforms strong baselines while remaining lightweight and interpretable. Overall, our results suggest that structured attribution signals are implicitly encoded in LLM attention, and can be effectively leveraged for faithful multi-document reasoning.
Reading Between the Dots: Decoding Hidden Computation across Filler Tokens
Kaley Brauer ⋅ Claudio Mayrink Verdun ⋅ Samuel Marks
Frontier LLMs can perform multi-step reasoning over content-free filler tokens like dots or counting sequences, producing correct answers with no visible chain-of-thought (CoT). This has been cited as a limit case for CoT monitorability: if the surface tokens carry no information about the reasoning, behavioral oversight has nothing to read. But hidden from the output is not the same as hidden from us. On three task families spanning fact retrieval, parallel numeric composition, and string manipulation, two open-weights frontier models (DeepSeek V3, Kimi K2) compute over filler tokens in a structured, legible way. Attention routes the question through the filler region to the answer, KV-cache transplants at only filler positions causally swap the model's output between examples, and logit-lens readouts reveal a temporal structure in which retrieved facts appear early and their composition crystallizes in late layers. We then introduce an unsupervised decoding pipeline that takes only hidden states as input and recovers intermediate values with 80–95\% accuracy with the strongest LLM judge across both models and all three tasks, without ground-truth labels or training. Hidden computation that defeats behavioral CoT monitoring is, on these tasks, directly readable from the residual stream. This suggests that monitorability should be understood as a property of the model's full computational trace, not just its surface tokens.
Read, Parse, Describe: Unified Document Parsing with Visual Element Description Generation
Yuheng Chen ⋅ Yufan Chen ⋅ Zhuojun Cai ⋅ Ruiping Liu ⋅ Junwei Zheng ⋅ Jiale Wei ⋅ Di Wen ⋅ Weijia Fan ⋅ Kunyu Peng ⋅ Jiaming Zhang ⋅ Rainer Stiefelhagen
Document understanding systems are typically evaluated as separate components via layout analysis, optical text recognition, reading-order prediction, and image captioning. This separation leaves scientific figures, captions, and visual semantics weakly connected, and fails to measure whether a model can produce a single page-level parse that is both structurally faithful and semantically accessible. We introduce Read-Parse-Describe, a unified dataset for document parsing with visual element description generation. Given a page image, models are required to output an ordered list of layout elements containing bounding boxes, semantic labels, reading-order indices, recognized textual content, and alternative text descriptions for visual elements. The benchmark is built from 2,402 documents, comprising 18,164 continuous pages, 125,218 layout elements, and 11,338 visual instances. Besides, we further propose a plug-and-play space-awareness enhancer that reuses multi-layer visual features and injects grid-aware spatial guidance into vision-language tokens for layout-aware generation. Experiments across end-to-end VLMs and pipeline parsers show that existing systems still struggle to jointly recover reading order and visual semantics, while our enhanced model achieves strong joint performance. These results highlight the need to move beyond isolated document parsing modules toward unified models that jointly read, parse, and describe complex documents.
Real In, Real Out: What If We Only Use Real Data for Scene Text Editing?
Xingsong Ye ⋅ Yongkun Du ⋅ Jiaxin Zhang ⋅ Chong Sun ⋅ Zhineng Chen
Aiming to modify textual content while preserving the original style, Scene Text Editing (STE) has significant practical value in many applications. STE typically requires paired images before and after editing for training. However, recent approaches rely heavily on synthetic paired data and complex disentanglement strategies. They struggle to faithfully capture the diversify and complexity of natural scenes, often producing artifacts or hallucinations. In this work, we challenge the common practice and investigate whether STE can be driven by turning each real sample into its own paired supervision. Concretely, we explore multiple ways to disrupt the original image (e.g., Crop, Shuffle, Mask) to construct self-paired training data, and build a strong baseline (TextRIRO) upon a conditional diffusion architecture. Our experiments show that combining character-column shuffling with image cropping (Shuffle&Crop) effectively destroys text semantics while preserving key style cues, enabling the model to learn editing behaviors directly from real text images. With this concise yet real-oriented design, TextRIRO achieves promising performance in both accuracy and visual realism across real-world benchmarks, substantially mitigating the synthetic-to-real domain gap. Furthermore, we leverage TextRIRO to build TextRIRO-3M, a large-scale realistic dataset for long-text recognition by concatenating edited and original images. Experimental results show that training recognition models with TextRIRO-3M significantly improves recognition accuracy, demonstrating that TextRIRO provides highly useful training resources for downstream tasks. Overall, TextRIRO lifts STE to realistic, reliable, and widely applicable new levels. Code is provided in the supplement, and data will be released upon acceptance.
REAL-MED: Benchmarking LLM Agents on Real-World Medical Tasks
Zihan Wang ⋅ yifan zhang ⋅ Jieqiong Cao ⋅ Junxiu Chen ⋅ Shanni Chen ⋅ ChaoGao ⋅ Hao Wang ⋅ Shi Feng ⋅ Xiaocui Yang ⋅ Jinghao Lin ⋅ Kai WU ⋅ Xiaozhong Ji ⋅ Haihua Yang
Real-world medical work extends far beyond answering isolated clinical questions: it often requires clinicians to process complex materials, interact with information systems, and produce structured deliverables. As AI systems evolve from medical chatbots into medical agents, evaluation should likewise move beyond \emph{answer correctness} toward the \emph{quality of end-to-end workflow deliverables}. We introduce \textsc{Real-Med}, a benchmark for evaluating AI agents on real-world medical tasks. \textsc{Real-Med} contains 405 expert-designed tasks across four categories: Clinical Support, Patient Management, Pharmacy Management, and Medical Education & Research. Each task is grounded in a realistic scenario with concrete materials, tool interfaces, and explicit deliverable requirements, such as clinical documents, medication plans, spreadsheets and review reports. To support rigorous evaluation, each task is annotated with fine-grained expert-written rubrics, averaging 15.4 independent rubric items per task, and every task is cross-validated by at least three medical experts. We evaluate 11 strong mainstream models under a unified harness with the same resources, tools, and execution protocol. Our results show that current agents can often satisfy surface-level task requirements, but still struggle with detail-sensitive medical reasoning, dosage-calculation traps, shallow research, formatting constraints, and reliable final delivery. These findings reveal a substantial gap between medical QA ability and robust execution in realistic medical workflows. Code is available at \url{https://anonymous.4open.science/r/REAL-MED}.
Real-World Dual-Pixel Raindrop Removal: A New Benchmark and Degradation-Adaptive RWKV Baseline
Hao Yang ⋅ Ruikun Zhang ⋅ Liyuan Pan
Raindrop removal is challenging because raindrops exhibit strong spatial non-uniformity and scene dependency. Although previous works have demonstrated that dual-pixel (DP) sensors can facilitate raindrop removal, their effectiveness in complex real-world scenarios remains limited due to local modeling paradigms and insufficient coverage of real data. In this work, we collect a large-scale real-world DP raindrop dataset that covers day/night scenes, diverse light conditions, and varying raindrop intensities. Based on this dataset, we establish a challenging benchmark for DP raindrop removal in real-world scenarios. Building upon this dataset, we develop DRWKV, which models raindrop removal as a degradation-conditioned state evolution process. DRWKV explicitly injects raindrop degradation information derived from DP sensors into the recurrent state update process, enabling the model to adaptively balance information preservation and content restoration across spatial regions while maintaining linear computational complexity. This design facilitates effective global modeling of spatially non-uniform raindrop degradations. Extensive experiments demonstrate that the proposed DRWKV outperforms prior methods on the proposed benchmark as well as existing datasets.
Reasoning-based Spatial Prior (RSP): Learning Spatial Priors from Multimodal LLMs for Object Detection
Cagri Gungor ⋅ Qingshuang Chen ⋅ Hongda Mao ⋅ Chi Zhang ⋅ Yelin Kim
Grounding natural language queries to target objects in images requires both high-level reasoning about intentions and precise spatial localization. Multimodal large language models (MLLMs) excel at reasoning about which object is desired given an implicit intention query but turning this understanding into precise bounding box predictions remains challenging. Conversely, specialized detectors such as GroundingDINO offer robust localization capabilities but lack the inherent capacity to understand implicit intentions. Existing hybrid methods couple these components only at the semantic level, such as by passing object names or a single global token to the detector, thereby discarding the rich, dense spatial signals encoded in MLLM attention maps. We propose Reasoning-based Spatial Prior (RSP), a reasoning-based detection framework that converts the MLLM’s internal attention into a learnable spatial prior and injects it directly into the detector. Concretely, we learn input-adaptive gating over attention heads, supervise the fused spatial prior with a SoftIoU localization loss, and integrate it into GroundingDINO via a prior-guided deformable fusion module that biases sampling toward regions highlighted by the reasoning process. Our method consistently outperforms strong MLLM-only, detector-only, and hybrid baselines on the EgoIntention and RIO benchmarks, achieving state-of-the-art performance with substantial gains across both frequent and uncommon categories.
Reasoning Under 1 Billion: Memory-Augmented Reinforcement Learning for Large Language Models
Hung Le ⋅ Van Dai Do ⋅ Dung Nguyen ⋅ Svetha Venkatesh
Recent advances in fine-tuning large language models (LLMs) with reinforcement learning (RL) have shown promising improvements in complex reasoning tasks, particularly when paired with chain-of-thought (CoT) prompting. However, these successes have been largely demonstrated on large-scale models with billions of parameters, where a strong pretraining foundation ensures effective initial exploration. In contrast, RL remains challenging for tiny LLMs with 1 billion parameters or fewer because they lack the necessary pretraining strength to explore effectively, often leading to suboptimal reasoning patterns. This work introduces a novel intrinsic motivation approach, called Memory-R+, that leverages episodic memory to address this challenge, improving tiny LLMs in CoT reasoning tasks. Inspired by human memory-driven learning, our method leverages successful reasoning patterns stored in memory while allowing controlled exploration to generate novel responses. Intrinsic rewards are computed efficiently using a kNN-based episodic memory, allowing the model to discover new reasoning strategies while quickly adapting to effective past solutions. Experiments on three reasoning datasets demonstrate that our approach significantly enhances smaller LLMs' reasoning performance and generalization capability, making RL-based reasoning improvements more accessible in low-resource settings.
Reasoning with Undecoded Tokens in Diffusion Language Models
Andre W He ⋅ Sean Welleck ⋅ Daniel Fried
Discrete diffusion models have recently become competitive with autoregressive models for language modeling, even outperforming them on reasoning tasks requiring planning and global coherence, but diffusion requires more computation at inference time. We trace this trade-off to a key mechanism: diffusion models are trained to jointly predict all unknown tokens simultaneously, including those that will not actually be decoded in the current step. Ablating this joint prediction yields faster inference but degrades performance, revealing that accurate prediction at the decoded position relies on joint reasoning about the undecoded tokens. We interpret these undecoded positions as latent tokens and introduce a method for modulating their number, achieving a smooth tradeoff between inference speed and sample quality. Furthermore, we demonstrate that latent tokens can be introduced into autoregressive models through an auxiliary multi-token prediction objective, yielding substantial improvements on the same reasoning tasks where they have traditionally struggled. Our results suggest that latent tokens--arising from jointly predicting multiple unknown positions--represent a general mechanism for improving performance on tasks requiring global coherence or lookahead.
Reason in the Words You Speak: Idiolectal Paraphrasing Off-Policy Traces for Reasoning Distillation in VideoLLMs
Ji Soo Lee ⋅ Jinyoung Park ⋅ Seohyun Lee ⋅ Jongha Kim ⋅ Joonmyung Choi ⋅ Jinsung Yoon ⋅ Hyunwoo J. Kim
Recent large language models achieve strong performance on complex reasoning tasks, where reinforcement learning with Group Relative Policy Optimization (GRPO) has emerged as a leading paradigm for optimizing models on self-generated trajectories. However, the on-policy nature of GRPO bounds the model to the reasoning skills it can already produce, restricting to learn more advanced capabilities. Prior works inject privileged reasoning traces from a stronger teacher policy to guide training, yet these traces are inherently out of distribution with respect to the student policy. We observe that this off-policy mismatch causes gradient clipping on semantically critical reasoning tokens, ultimately rewarding correct answers while leaving the reasoning that justifies them unlearned. Hence, we propose $\textbf{Echo-GRPO}$, a framework that lets the model reason in the words it speaks. Rather than imitating low-probability privileged traces from the teacher model, Echo-GRPO rewrites them into the student policy's own $\textit{idiolect}$, that is, its own characteristic vocabulary and expression patterns, while preserving their semantics via Dual-Reference Decoding. We instantiate this framework as $\textbf{VideoEcho-R1}$ for video reasoning distillation, achieving consistent improvements across three multimodal LLM backbones and five benchmarks. Finally, we show that our idiolectal paraphrasing is a plug-in module that consistently improves both reinforcement learning and supervised fine-tuning frameworks for reasoning distillation, demonstrating that policy-aligned supervision extends beyond GRPO.
Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models
Nilaksh ⋅ Saurav Jha ⋅ Artem Zholus ⋅ Sarath Chandar
World model-based policy evaluation is a practical proxy for testing real-world robot control by rolling out candidate actions in action-conditioned video diffusion models. As these models increasingly adopt latent diffusion modeling (LDM), choosing the right latent space becomes critical. While the status quo uses autoencoding latent spaces like VAEs that are primarily trained for pixel reconstruction, recent work suggests benefits from pretrained encoders with representation-aligned semantic latent spaces. We systematically evaluate these latent spaces for action-conditioned LDM by comparing six reconstruction and semantic encoders to train world model variants under a fixed protocol on BridgeV2 dataset, and show effective world model training in high-dimensional representation spaces with and without dimension compression. We then propose three axes to assess robotic world model performance: visual fidelity, planning and downstream policy performance, and latent representation quality. Our results show visual fidelity alone is insufficient for world model selection. While reconstruction encoders like VAE and Cosmos achieve strong pixel-level scores, semantic encoders such as V-JEPA 2.1 (strongest overall on policy), Web-DINO, and SigLIP 2 generally excel across the other two axes at all model scales. Our study advocates semantic latent space as stronger foundation for policy-relevant robotics diffusion world models.
Recovering the Apresjan Hierarchy Using Linkage-Based Clustering
Maximilien Dreveton ⋅ Matthias Grossglauser ⋅ Daichi Kuroda ⋅ Patrick Thiran
Hierarchical clustering seeks to uncover nested structure in data by constructing a tree of clusters, whose deeper levels reveal increasingly fine-grained relationships. However, traditional hierarchical clustering methods always return a hierarchy, even when the data contain no meaningful hierarchical structure. Moreover, agglomerative linkage algorithms are highly sensitive to the choice of linkage function. In this paper, we revisit a classical notion of well-separated clusters, which we call valid clusters. The collection of all valid clusters forms a hierarchy, known as the Apresjan hierarchy. This hierarchy is the finest hierarchy composed only of valid clusters: it need not be binary, and it collapses to a star tree when no nontrivial valid clusters exist. Our main contribution is to characterize when the Apresjan hierarchy can be recovered from linkage-based algorithms. We propose a two-step procedure that first constructs a binary hierarchy using an agglomerative linkage method and then prunes all clusters that violate the validity condition. We give necessary and sufficient conditions on the linkage function under which this procedure exactly recovers the Apresjan hierarchy. Consequently, all linkage rules satisfying these conditions yield the same pruned hierarchy. In particular, single, complete, and average linkage satisfy the conditions, whereas Ward's linkage does not.
Rectified Policy Rollouts with Hierarchical Expert Guidance for Neural Combinatorial Optimization
Wenzheng Pan ⋅ Nuoyan Chen ⋅ Jiaxi Liu ⋅ Jiale Ma ⋅ Junchi Yan
In the field of Neural Combinatorial Optimization (NCO), Reinforcement Learning (RL) stands out for its ability to naturally enforce complex constraints via its auto-regressive decision pattern. To mitigate the longstanding challenges, e.g., reward sparsity and sample inefficiency in the vast combinatorial action space which causes ineffective training, subpar performance, and poor scalability, we propose CORectifier, a novel NCO solver learned under the hierarchical gated rectification mechanism to regularize the arbitrary exploration: partial policy-predicted actions in a trajectory are probabilistically replaced with high-quality segments from reference solutions. CORectifier prompts the model with tri-level optimal signals, operating at the batch, instance, and sub-instance levels with fragments of diverse lengths injected at random decision steps. This Rectified RL (RRL) paradigm helps develop optimality-aware and sample-efficient RL learners while maintaining their sequential-decision manner for constraint satisfaction, delivering a new perspective to hybridize RL and SL/IL for NCO with improved utility rate of limited supervision. Sufficiently extensive experiments on Traveling Salesman Problem (TSP), Asymmetric TSP (ATSP), Prize-Collecting TSP (PCTSP), Capacitated Vehicle Routing Problem (CVRP), Knapsack Problem (KP), Job-Shop Scheduling Problem (JSSP), and Single Machine Total Weighted Tardiness Problem (SMTWTP), across synthetic and real-world benchmarks, show superior quality than RL baselines by up to 59.7%.
Privacy is a central concern when fine-tuning large language models (LLMs) on sensitive data, and differentially private stochastic gradient descent (DP-SGD)---which clips per-sample gradients and adds calibrated Gaussian noise---is the standard tool for formal privacy guarantees. Both theory and practice show that lower-rank models are better suited to DP training, a property especially relevant for LLMs, whose fine-tuning gradients exhibit a strong low-rank structure. Methods such as DP-LoRA exploit this by restricting updates to a low-rank subspace, i.e., retaining only a few non-zero components in the SVD of each layer's gradient. However, we argue that while having few non-zero components is important, the isotropic noise injected by DP-SGD inflates the singular values of the gradient matrix, disrupting their naturally fast decay. In this work, we investigate whether this noise-induced eigenvalue blow-up reduces performance, and show that partially restoring the original singular-value profile significantly improves the sample efficiency of DP-SGD. Experiments on language classification (GLUE benchmark with RoBERTa) and text generation (E2E and DART table-to-text benchmarks with Qwen and Llama models up to 4B parameters) showcase that restoring the fast decay of singular values is a viable strategy for speeding up the DP optimization process, without compromising privacy guarantees.
Reflective Prompted Policy Optimization: Trajectory-Grounded Revision and Salience Bias
Rahaf Abu Hara ⋅ Vaibbhav Murarri ⋅ Claudio Zito
Existing LLM-based policy optimizers see only scalar rewards: that a policy scored 0.45, but not whether the agent got stuck in a loop, fell into a hole on the third step, or performed well on 19 out of 20 rollouts and failed catastrophically on one. We propose Reflective Prompted Policy Optimization (R2PO), a two-stage LLM framework for policy search over compact policy classes that augments scalar reward feedback with trajectory-level behavioral evidence. A Search-LLM acts as a global policy optimizer and proposes candidate policy parameters; the environment executes them; a Critic-LLM then inspects the resulting rollouts and proposes targeted parameter revisions grounded in observed states, actions, and rewards. Across ten environments, ablations show R2PO's gains arise from a design that explicitly separates global search from behavior-grounded revision and uses selection to filter high-variance edits. We further identify a dominant failure mode, salience bias. When presented with multiple rollouts, the Critic-LLM fixates on improving a single failure even when most trajectories succeed. In a three-trajectory variant, where the Critic-LLM is shown the best, worst, and median rollout from each evaluation, this behavior explains 76.6\% of regressions on CartPole. R2PO mitigates this by reasoning over aggregate rollout statistics, median-trajectory selection, and a revision rule. Using a relatively small open-weight 20B-parameter model, R2PO achieves the highest mean best reward across all ten environments, while reaching near-optimal performance substantially earlier in training (e.g., near-maximum CartPole reward within $\sim$500 episodes), and training far more stably than both deep RL and prior LLM-based methods. Together, these results show that treating trajectories as first-class in-context evidence, rather than background artifacts reduced to scalar returns, fundamentally changes how even comparatively small LLMs can search over policy spaces, enabling them to learn faster, diagnose more precisely, and reliably improve external controllers rather than tune them by trial and error.
ReFlex: Faithful Panorama Reconstruction from a Single Image with Reflections
Yikai Pan ⋅ Zhiyuan Zhou ⋅ Yuxiang Yan ⋅ Xuetong Yang ⋅ Runhe Yang ⋅ Jian Pu
Reflections reveal scene content beyond the camera's direct field of view, but faithfully recovering this hidden information is challenging as reflective observations are heterogeneous, incomplete, and physically modulated. Existing methods either rely on explicit reflection modeling with privileged geometric priors or defer to generative synthesis that may hallucinate plausible but unfaithful content. To address these challenges, we propose \textbf{ReFlex}, an observation-constrained framework that casts faithful panorama reconstruction as a controlled inverse problem. We first develop an \emph{Observation-Consistent Estimator} with hierarchical token prediction and signed routing to integrate heterogeneous reflection cues while stabilizing global structure. Subsequently, we propose a \emph{Reliability-Guided Refiner}, which estimates spatial reliability from token predictions and guides a frozen diffusion model to refine uncertain regions without overwriting observation-supported structures. To validate the effectiveness of ReFlex, we further build ReFlex-Bench, a paired benchmark with 110,760 reflective RGB inputs and aligned target panoramas across diverse scenes, reflective geometries, and observation regimes. Experiments show that ReFlex delivers superior reconstruction fidelity without normals, field-of-view annotations, masks, or other privileged information, and obtains the best performance on eight of nine downstream full-view perception protocols.
ReFPO: Reflow Regularization for Flow Matching Policy Gradients
Ge Wang ⋅ Yibo Peng ⋅ Fan Feng ⋅ Shenhao Yan ⋅ Chengsi Yao ⋅ Jiahao Yang ⋅ Honghao Cai ⋅ Yiming Zhao ⋅ Xi Li ⋅ Jinke Ren ⋅ Shuguang Cui ⋅ Yatong Han ⋅ Zhen Li
We present Reflow-regularized Flow Matching Policy Gradients (ReFPO), a simple online RL method that adds explicit Reflow regularization to FPO for efficient flow-based control. We uncover a key structural property: the gradient updates in Flow Matching Policy Gradients (FPO) can be interpreted as an implicit advantage-weighted Reflow process, providing a new geometric perspective on flow-based policy gradients. Building on this insight, ReFPO introduces an explicit geometric regularizer that can be implemented with a single line of code change without incurring additional computational overhead or auxiliary distillation stages. By synergizing advantage-guided updates with path rectification, our method reduces CFM proxy-ratio spikes, stabilizes PPO-style training, and enables high-fidelity one-step inference that often matches or exceeds multi-step performance. We experimentally demonstrate that ReFPO improves average performance and discretization robustness across GridWorld, MuJoCo Playground, and high-dimensional Humanoid Control tasks, providing a scalable and stable approach for generative policies in complex physical simulations.
ReGen: Agentic Video World Modeling with Synergized Reasoning and Generation
Chenguo Lin ⋅ Yu Tang ⋅ Weiqiao Zheng ⋅ Enhua Jiang ⋅ Zhiguang Liu ⋅ Maomao Li ⋅ Muxi Chen ⋅ Xiaojia Chen ⋅ Shisheng Huang ⋅ Jianhuan Zhuo ⋅ Zilin Yang ⋅ Jiarong Ou ⋅ Rui Chen ⋅ Yadong Mu
Video world models aim to simulate environmental dynamics under interaction. However, most existing approaches entangle transition reasoning with pixel-level generation within a single model, thereby limiting narrative continuity and long-horizon coherence. We reformulate video world modeling as an agentic process that synergizes reasoning and generation, and present ReGen, which couples a streaming vision-language planner with a causal diffusion generator. At each step, the planner infers a structured action-state specification from history, while the generator realizes it as the next video segment. The planner further verifies the outcome and decides to continue, end, or re-generate, forming a closed-loop sense–plan–act–verify pipeline. To support this formulation, we carefully construct a large-scale action-grounded video dataset with over 2M segments that instantiates the planner–generator interface, and introduce action-grounded reward alignment that turns the action specification into an enforced reward and mitigates cross-segment drift. Experiments show that ReGen improves long-horizon narrative coherence over prior work, supporting autonomous long-horizon rollouts without step-by-step user intervention.
ReGenHuman: Re-Generating Human Appearances for Realistic Full-Body Video Anonymization
Adam Sun ⋅ Eshaan Barkataki ⋅ Arnold Milstein ⋅ Gordon Wetzstein ⋅ Ehsan Adeli
Anonymizing human-centric video data is an understudied problem. Prior anonymization techniques either blur or redact pixels at the cost of realism and downstream utility, or generate frame-by-frame at the cost of temporal coherence. We introduce ReGenHuman, the first full-body video anonymization pipeline that is simultaneously realistic, temporally consistent, and anonymous by construction. Contrary to past approaches which redact or edit the inputs directly, we propose a regenerate, don't edit paradigm. Our approach composites 2D pose, segmentation, and monocular depth into two complementary conditioning streams---StructAll and StructHuman, which are used to fine-tune a video-to-video diffusion backbone on in-the-wild human videos, synthesizing the human regions entirely from identity-free structural cues. We evaluate our model on privacy, quality, and utility, and show that our ReGenHuman achieves the best tradeoff across all three axes against current baselines. We further show that our anonymized videos remain effective for downstream tasks, including video question answering.
Regret-Based $(\epsilon,\delta)$-optimal Stopping Criteria for Bayesian Optimization
Haowei Wang ⋅ Jingyi Wang ⋅ Qiyu Wei
Bayesian optimization (BO) is a widely used iterative black-box optimization method that utilizes Gaussian process (GP) surrogate models. In practice, BO is typically terminated after a fixed evaluation budget is exhausted, which can incur unnecessary cost and provides no optimality guarantee on solution quality. Recent research in developing a practical stopping criterion has made empirical progress, yet a theoretically sound stopping criterion remains a work in progress. In this work, we present provably tighter instantaneous regret bounds for GP upper confidence bound (GP-UCB) at any given iteration. Then, we propose stopping criteria for GP-UCB based on this tighter bound that ensures an $\epsilon$-optimal solution with high probability $1-\delta$ upon termination. Numerical experiments are performed to validate and demonstrate the effectiveness and efficiency of our stopping criteria.
Reinforcement Learning Agents Are Swimmers
Juan Rojas ⋅ Jacob Adamczyk ⋅ Abhishek Naik ⋅ Volodymyr Makarenko ⋅ Gautham Vasan ⋅ Nikhil Mukund ⋅ Stas Tiomkin ⋅ Rahul Kulkarni ⋅ Chi-Guhn Lee
To date, the field of reinforcement learning (RL) has primarily focused on developing discounted methods, which are designed to optimize the short-term or transient performance of RL agents. Conversely, the field has been reluctant to explore and invest in methods that are designed to optimize the long-term or steady-state performance of RL agents. In this position paper, we challenge this reluctance under the framing that RL agents should be viewed as competitive swimmers. That is, we argue that, like competitive swimmers, RL agents cannot hope to win all races, or solve all tasks, based on transient performance alone. To this end, we first motivate the need for RL agents that can adequately optimize both the transient and steady-state performance. Then, we summarize recent works which show that discounted methods have deficiencies when it comes to optimizing the steady-state performance. Finally, we argue in favor of more attention and investment from the community towards average-reward methods, which are designed to optimize the steady-state performance. Namely, we challenge common misconceptions associated with average-reward methods, and highlight recent advancements which suggest that such methods present a viable and potentially superior alternative.
RelAgent: LLM Agents as Data Scientists for Relational Learning
Xingyue Huang ⋅ Louis Tichelman ⋅ Jinwoo Kim ⋅ Krzysztof Olejniczak ⋅ Ismail Ilkan Ceylan
Relational learning is a challenging problem that has motivated a wide range of approaches, including graph-based models (e.g., graph neural networks, graph transformers), tabular methods (e.g., tabular foundation models), and sequence-based approaches (e.g., large language models), each with its own advantages and limitations. We propose RelAgent, an LLM-based autonomous data scientist for relational learning, which operates in two phases. In the search phase, an LLM agent uses database, validation, and evaluation workspace tools to construct SQL feature programs and select a predictive model. In the inference phase, the resulting program is executed without further LLM calls. The final predictor consists of SQL queries and a classical model, enabling fast, deterministic, and intrinsically interpretable predictions: features are human-readable queries, and predictions depend only on the resulting query-defined feature map, enabling scalable deployment using standard database systems.
Remember the Decision, Not the Description: A Rate-Distortion Framework for Agent Memory
mingxi Zou ⋅ Zhihan Guo ⋅ Langzhang Liang ⋅ Zhuo Wang ⋅ Qifan Wang ⋅ Qingsong Wen ⋅ Irwin King ⋅ Lizhen Qu ⋅ Zenglin Xu
Long-horizon language agents must operate under limited runtime memory, yet existing memory mechanisms often organize experience around descriptive criteria—relevance, salience, or summary quality. For an agent, however, memory is valuable not because it faithfully describes the past, but because it preserves the distinctions between histories that must remain separated under a fixed budget to support good decisions. We cast this as a decision-centric rate-distortion problem, measuring memory quality by the loss in achievable decision quality induced by compression. This yields an exact forgetting boundary for what can be safely forgotten, and a memory-distortion frontier characterizing the optimal tradeoff between memory budget and decision quality. Motivated by this decision-centric view of memory, we propose DeMem, an online memory learner that refines its partition only when data certify that a shared state would induce decision conflict, and prove near-minimax regret guarantees. On both controlled synthetic diagnostics and long-horizon conversational benchmarks, DeMem yields consistent gains under the same runtime budget, supporting the principle that memory should preserve the distinctions that matter for decisions, not descriptions.
Remote Photoplethysmography Based on a Skin Reflection Exponential Model
Jiachen Li ⋅ Longzhen Tang ⋅ Shisheng Guo ⋅ Guolong Cui
Facial video–based remote photoplethysmography (rPPG) estimates physiological signals from subtle spatiotemporal variations in facial videos, but remains vulnerable to head motion, illumination changes, and environmental perturbations. Existing methods mainly distinguish rPPG-related variations from motion and lighting artifacts at the signal or feature level, while largely neglecting the optical reflection properties of human skin. In this work, we formulate a skin reflection exponential model to describe the nonlinear relationship among skin reflectance, illumination, motion-induced variations, and blood-volume changes. Guided by this model, we propose a physically inspired rPPG estimation framework that explicitly incorporates skin reflectance priors into signal extraction. Given a facial video, facial keypoints are first detected and encoded to represent motion dynamics, while illumination-related features and candidate rPPG representations are jointly modeled under the proposed exponential formulation. By constraining rPPG estimation with skin-reflection physics, the framework improves separation of physiological components from motion and illumination disturbances. Experiments demonstrate improved robustness under challenging conditions.
Rennala-NSGD: Asynchronous Stochastic Optimization Beyond Euclidean Geometry
Igor Sokolov ⋅ Alexander Tyurin ⋅ Peter Richtarik
The recent empirical success of non-Euclidean optimizers, such as Muon (Jordan et al., 2024) and Shampoo (Gupta et al., 2018), has transformed the training of Large Language Models (LLMs) by exploiting the geometry of matrix spaces. Despite this progress, the theoretical understanding of stochastic optimization in general normed spaces remains limited, particularly in distributed settings. In this work, we present a novel analysis of Minibatch Stochastic Gradient Descent (SGD) in non-Euclidean geometries. Leveraging the framework of $\\kappa$-regular normed spaces (Juditsky & Nemirovski, 2008) and a light-tail noise model defined directly in the dual norm, we derive high-probability oracle complexity bounds for smooth non-convex optimization. Crucially, our analysis reveals that measuring stochastic variance within the intrinsic non-Euclidean geometry avoids the suboptimal dimension-dependent factors inherent in Euclidean norm equivalence. Building on this framework, we introduce Rennala-NSGD, a semi-asynchronous distributed algorithm, and establish, to the best of our knowledge, the first wall-clock time complexity guarantee for asynchronous non-Euclidean stochastic optimization.
RePiD: Efficient Recursive Pixel-Space Diffusion via Hierarchical Patch Denoising
Duoyou Chen ⋅ Yaxue Guo ⋅ Can Zhang ⋅ Cheng Chen
Recent pixel-space generative models avoid the reconstruction bottleneck of latent diffusion, but direct denoising over high-dimensional pixel tokens remains computationally demanding. This paper studies a recursive alternative for efficient pixel-space diffusion. We propose RePiD, a hierarchical patch-level diffusion framework that recursively partitions an image and performs the core denoising computation on lower-dimensional pixel groups instead of repeatedly operating on full-resolution images. To preserve spatial coherence across patch-wise processing, RePiD introduces a neighborhood embedding module that conditions each patch on its surrounding regions. Experiments on class-to-image generation and unpaired image-to-image translation show that RePiD provides a favorable quality--efficiency trade-off for direct pixel-space generation. These results suggest that recursive patch-level denoising is a practical design direction for efficient pixel-space generative modeling. Our code is publicly available at \textcolor[HTML]{387FB9}{https://anonymous.4open.science/r/Recursive-Pixel-based-Diffusion-Model-1735}.
Representation Learning Enables Scalable Multitask Deep Reinforcement Learning
Johan Obando Ceron ⋅ Lu Li ⋅ Scott Fujimoto ⋅ Pierre-Luc Bacon ⋅ Aaron Courville ⋅ Pablo Samuel Castro
Scaling reinforcement learning (RL) to diverse multitask settings remains a central challenge. While recent advances in model-based RL achieve strong performance, they rely on planning and complex training pipelines, making it unclear which components are essential for scalability. We revisit this question and argue that the primary driver of scalable multitask RL is not model-based control, but representation learning. In particular, we show that combining predictive, model-based representations with high-capacity value function approximation is sufficient to achieve strong performance, without planning. We evaluate a simple model-free algorithm, MR.Q, that integrates auxiliary predictive objectives into a scalable actor-critic architecture. This approach outperforms a recent world-model-based method and a range of deep RL baselines across a diverse suite of multitask continuous control tasks, while significantly reducing computational overhead and improving wall-clock efficiency. We observe consistent improvements with increased model capacity and show through ablations that predictive representation learning is critical for performance.
Requential Coding: Measuring Model Compressibility by Coding Data Instead of Parameters
Shikai Qiu ⋅ Marc Finzi ⋅ Yujia Zheng ⋅ Kun Zhang ⋅ Andrew Wilson
Beyond reducing the memory footprint and costs for deployment, the extent to which a probabilistic model can be compressed is central to understanding generalization, model selection, and dataset selection. Yet with parameter-based compression methods (quantization, pruning, low-rank training) the code length scales with model size and is insensitive to how much information the model actually learned. Prequential coding sidesteps the parameter ceiling by encoding the dataset through the learning process, but its code length includes the entropy of the data. We introduce requential coding, a model coding scheme that is unconstrained by model size and does not pay for the data entropy: synthetic training samples are drawn from a teacher distribution and communicated relative to the student's current predictions via relative entropy coding, so the code length equals the cumulative teacher-student KL along the training trajectory. The resulting compressibility measure behaves qualitatively differently from parameter-based codes. The code length to reach a fixed target loss decreases with model and ensemble size, evidence that a more flexible model can be far simpler than its parameter count implies. Plugged into a PAC-Bayes bound, requential coding produces non-vacuous generalization guarantees that improve with scale and beat even the lossless idealization of the 4-bit GPTQ baseline that previously gave state-of-the-art bounds on compute-optimal LLMs. The same code length predicts overfitting in data-constrained training and tracks intuitive ordering of dataset complexity across CIFAR-5M, OpenWebText, and FineWeb.
Reranking with Intra-modal Visual Association for Text-to-Image Person Re-Identification
Junhong Wang ⋅ Changxing Ding ⋅ Yusha Peng ⋅ Wentao Tan ⋅ Dongliang Liao
Text-to-image person re-identification (TIReID) aims to retrieve images of one pedestrian using natural-language descriptions. It remains fundamentally challenging due to significant inter-modal semantic gap between sparse textual cues and fine-grained visual appearance. Existing methods mainly perform instance-wise text-image matching during inference, overlooking intra-modal visual associations among top-ranked gallery candidates, where multiple images may depict the same identity. To exploit such associations, we propose Reranking with Intra-modal Visual Association (RIVA), a plug-and-play reranking framework that leverages intra-modal visual association among retrieved candidates to guide multimodal large language model (MLLM)-based reranking. RIVA consists of two complementary components. First, we propose an MLLM-based Short-lIst Visual Association (SIVA) method, which is trained with a three-stage, dual-task curriculum to learn identity-level image-image association and then use it as visual context for text-image matching. Second, we develop a long-list visual association approach based on clustering, which clusters the top-ranked candidates in the feature space of pedestrian images, uses SIVA to examine the most likely matched cluster, and calibrates the image-text similarity scores accordingly. Extensive experiments on three popular TIReID benchmarks demonstrate that RIVA consistently improves retrieval accuracy in multiple evaluation settings. The code will be released.
Bird's-eye-view (BEV) perception has emerged as a cornerstone of autonomous driving systems, providing a structured, ego-centric representation critical for downstream planning and control. However, real-world deployment faces challenges from sensor degradation and adversarial attacks, which can cause severe perceptual anomalies and ultimately compromise the safety of autonomous driving systems. To address this, we propose a resilient and plug-and-play BEV perception method (RESBev), which can be easily applied to existing BEV perception methods to enhance their robustness to diverse disturbances. Specifically, we reframe perception robustness as a latent semantic prediction problem. A latent dynamic predictor is constructed to extract spatiotemporal correlations across sequential BEV observations, thereby learning the underlying BEV state transitions to predict clean BEV features for reconstructing corrupted observations. The proposed framework operates at the semantic feature level of the BEV perception pipeline, enabling recovery that generalizes across both natural disturbances and adversarial attacks without modifying the underlying backbone. Extensive nuScenes experiments show RESBev notably enhances BEV perception robustness against natural disturbances and adversarial attacks, with latency, parameter and memory analyses verifying its excellent robustness-efficiency trade-off.
Rescaled Asynchronous SGD: Optimal Distributed Optimization under Data and System Heterogeneity
Ammar Mahran ⋅ Artavazd Maranjyan ⋅ Peter Richtarik
Asynchronous stochastic gradient descent (ASGD) is a standard way to exploit heterogeneous compute resources in distributed learning: instead of forcing fast workers to wait for slow ones, the server updates the model whenever a gradient arrives. Vanilla ASGD applies each arriving gradient with the same weight. When local data distributions are heterogeneous, this becomes problematic: faster workers contribute more updates, and we show theoretically that the method is biased toward a frequency-weighted average of the local objectives rather than the desired global objective. Existing remedies typically move away from the simple ASGD template by introducing gathering phases, buffering, or extra memory. We show that this is unnecessary. Keeping the standard ASGD mechanism, we recover the correct objective by rescaling worker-specific stepsizes in proportion to their computation times, so that each worker contributes the same aggregate learning rate over a cycle. In the non-convex setting, under smoothness and bounded heterogeneity assumptions, we prove that the resulting method, Rescaled ASGD, converges to stationary points of the correct global objective in the fixed-computation model. Its time complexity matches the known lower bound in the leading term, while the effects of staleness and data heterogeneity appear only in lower-order terms. Experiments confirm that the method converges to the correct objective and is competitive with state-of-the-art baselines.
RESIST: Resilient Decentralized Learning Using Consensus Gradient Descent
Cheng Fang ⋅ Rishabh Dixit ⋅ Waheed Bajwa ⋅ Mert Gurbuzbalaban
Empirical risk minimization (ERM) is a cornerstone of modern machine learning (ML), supported by advances in optimization theory that ensure efficient solutions with provable algorithmic convergence rates, which measure the speed at which optimization algorithms approach a solution, and statistical learning rates, which characterize how well the solution generalizes to unseen data. Privacy, memory, computational, and communications constraints increasingly necessitate data collection, processing, and storage across network-connected devices. In many applications, these networks operate in decentralized settings where a central server cannot be assumed, requiring decentralized ML algorithms that are both efficient and resilient. Decentralized learning, however, faces significant challenges, including an increased attack surface for adversarial interference during decentralized learning processes. This paper focuses on the man-in-the-middle (MITM) attack, wherein adversaries exploit communication vulnerabilities between devices to inject malicious updates during training, potentially causing models to deviate significantly from their intended ERM solutions. To address this challenge, we propose RESIST (Resilient dEcentralized learning using conSensus gradIent deScenT), an optimization algorithm designed to be robust against adversarially compromised communication links, where transmitted information may be arbitrarily altered before being received. RESIST uses a multistep consensus gradient descent framework with robust-statistics-based screening of neighbor messages. It has a design parameter J that controls the frequency of local gradient computation: each local gradient update is preceded by J-1 communication and robust aggregation rounds. Compared with methods that perform both communication and a local gradient update at every iteration, RESIST with J>2 uses fewer local gradient updates at the same communication budget, which is useful when local gradient computation is the dominant cost. We establish geometric algorithmic convergence guarantees for strongly convex and Polyak–Łojasiewicz ERM problems, and sublinear and finite-horizon guarantees for smooth nonconvex ERM problems. For heterogeneous local objectives, these guarantees quantify neighborhoods for the relevant iterate, objective-value, and stationarity errors. In the strongly convex homogeneous case, convergence becomes exact. We also establish statistical learning-rate guarantees under common-population independent and identically distributed sampling assumptions, including regimes in which the statistical error vanishes as the sample size grows. Experimental results demonstrate the robustness of RESIST across diverse attack strategies, screening methods, and loss functions. The J-ablation experiments further show that larger J can improve convergence and reduce the number of local gradient updates at a fixed communication budget.
ReSMap: Recasting Satellite Priors for Robust and Accurate Online HD Map Construction
Kyungmin Kim ⋅ Sumin Lee ⋅ Sungoh Jeong ⋅ DoHyun Lim ⋅ Seonghyun Park ⋅ Soonmin Hwang
Despite recent advances in camera-based online HD map construction, these systems remain susceptible to sensor failures due to their reliance on onboard cameras. In contrast, satellite images remain available independently of real-time sensing conditions as they are cached offline. Existing methods utilizing satellite images, however, primarily leverage them under nominal conditions, without explicitly addressing as redundancy when sensor degrades. To address this gap, we propose ReSMap, a redundancy-aware camera-satellite fusion framework for robust online HD map construction. Our framework explicitly leverages satellite images as a complementary source to improve robustness against camera failures. A Predict-then-Fuse (PTF) module fuses camera and satellite BEV features by using each branch's own prediction confidence as a per-pixel reliability weight, with no learned router. An Object-centric cooperative Multi-modal Query (OCMQ) decoder then propagates this isolation to the query level: each modality cross-attends only to its own BEV, and per-modality updates are merged through a learned per-token gate. Our model achieves state-of-the-art performance on nuScenes across all splits and range settings. More importantly, it also sets the new state-of-the-art under camera-failure scenarios, demonstrating that satellite images serve not only as a strong prior but as practical redundancy for robust online HD mapping.
Resolution-Aware Structural Density Peak Clustering
Jie Yang ⋅ Hsiang-Ting Chen ⋅ Yan Ma ⋅ Xinyan Liang ⋅ Avinash Singh ⋅ Liang Du ⋅ Cheng-You Lu ⋅ Chenglong Zhang ⋅ Bingbing Jiang ⋅ Weiping Ding ⋅ Wei Chen
Density Peak Clustering (DPC) focuses on identifying centers and propagates labels at the sample level, which makes density estimation cutoff-sensitive and assignment chain-vulnerable, yielding suboptimal performance under heterogeneous densities and complex cluster structures. This paper argues that density peaks inference should instead be performed on resolution-controlled structural units: density becomes a stable count statistic, relative distance inherits the separation of the enclosing hierarchy, and assignment runs over a compact graph of structural units rather than long sample-to-sample chains. This view is realized in Resolution-Aware Structural Density Peak Clustering (RSDP), which builds a constrained hierarchy, selects an admissible level via a single resolution parameter, and applies a one-pass DPC rule on the selected structural units. Under a faithful-realization condition, it is proved that RSDP's empirical $\gamma$-ranking recovers the dominant population units and that its one-pass assignment recovers the sampled ground-truth partition. On synthetic and real-world datasets, RSDP achieves the highest ACC and NMI against state-of-the-art DPC variants, exceeding the strongest by 19.91% and 9.09% in average ACC and NMI, respectively.
Resolving Time-Frequency Ridge Crossings via Frequency-Rate Lifting
Pingping Pan ⋅ You Li ⋅ Yunjian Zhang ⋅ Mu-Jiang-Shan Wang
Ridge crossings, where two components share the same instantaneous frequency (IF) at a given time, are a fundamental bottleneck in multi-component time-frequency (TF) mode decomposition. We trace this difficulty to \emph{insufficient representational dimensionality} and resolve it by lifting signals into the 3-D time-frequency-frequency-rate (TFFR) space $(t,f,\dot{f})$, where crossing ridges separate because they typically differ in the instantaneous frequency rate (IFR). We formalize this via a separability proposition and confirm it empirically through Monte Carlo simulation (crossing rate: 78.6\% in 2-D $\to$ 0.9\% in 3-D). Based on this principle, we propose \textbf{TFFR-Net}, a query-based masked transformer that jointly predicts per-ridge masks and dense IFR maps. An IFR feedback loop injects the estimated frequency rate back into each query embedding, enabling the decoder to perform instance discrimination in TFFR space rather than the TF plane. On a synthetic benchmark with up to six chirp components under broadband noise (0--25\,dB SNR), TFFR-Net consistently reduces IF MAE by at least 5.8\% and Dice loss by at least 5.4\% relative to competing methods. Evaluation on real-world signals (bat echolocation, power-system oscillation, earthquake vibration) demonstrates effective generalization beyond synthetic training data. The source code is available at \url{https://anonymous.4open.science/r/TFFR-92D2/}.
Rethinking How to Remember: Beyond Atomic Facts in Lifelong LLM Agent Memory
Jingwei Sun ⋅ Jianing Zhu ⋅ Jiangchao Yao ⋅ Tongliang Liu ⋅ Bo Han
To enable reliable long-term interaction, LLM agents require a memory system that can faithfully store, efficiently retrieve, and deeply reason over accumulated dialogue history. Most existing methods adopt an extracted fact based paradigm: handcrafted static prompts compress raw dialogues into atomic facts, which are then stored, matched, and injected into downstream reasoning. Nevertheless, such fact-centric designs inevitably discard fine-grained details in original dialogues and fail to support deep reasoning over scattered isolated facts. Moreover, static prompts cannot maintain consistent extraction granularity across diverse dialogue styles. To address these limitations, we propose TriMem, which maintains three coexisting representation granularities, including raw dialogue segments anchored by source identifiers for storage fidelity, extracted atomic facts for efficient memory retrieval, synthesized profiles that aggregate dispersed facts into holistic semantic understanding for deep reasoning. We further adopt TextGrad-based prompt optimization, which iteratively refines extraction and profiling prompts via response quality feedback, achieving lifelong evolution without any parameter updating. Extensive experiments on LoCoMo and PerLTQA across multiple LLM backbones demonstrate that TriMem consistently outperforms strong memory baselines.
Rethinking Incompleteness: Formalizing Protocol Divergence and Train-Once Learning for Robust IMVC
HAOLU LIU ⋅ XIYUE WANG ⋅ Xuanting Xie ⋅ Liangjian Wen ⋅ Zhao Kang
Standard IMVC evaluation retrains separate models for different missing-data configurations. We show that this paradigm obscures a fundamental vulnerability: missing rate alone is insufficient to characterize data incompleteness. Specifically, we show that protocols with identical nominal missing rates can differ by up to $50\times$ in their proportion of fully observed samples, inducing drastically different learning regimes. We formalize this phenomenon as incompleteness divergence, providing measures that capture structural disparities across missing-data protocols. We further prove that for a broad class of reconstruction-based objectives, learning becomes structurally ill-posed when the proportion of complete samples falls below a critical threshold, leading to near-random performance. To bypass this theoretical bound, we propose CRAFT (Complete-data Robust Attention-masked Fusion Transformer). CRAFT shifts the burden of robustness from the loss function to the architecture via two key properties: (i) per-sample independence, which removes reliance on complete-sample co-occurrence, and (ii) mask-aware variable-length fusion, which aggregates only observed views through attention masking. This design allows a single model, trained once on complete data, to generalize to diverse missing patterns at inference time without retraining. Extensive experiments on seven benchmarks show that CRAFT matches or outperforms per-configuration baselines while reducing training overhead by $8.8\times$, demonstrating that robustness to missing data can be achieved as an inherent architectural property. Code (CRAFT) and our imvc-audit toolkit are available at https://anonymous.4open.science/r/CRAFT-BF80/ and https://anonymous.4open.science/r/imvc-audit-8263/.
Rethinking Infrared Small Target Detection: A Foundation Driven Efficient Paradigm
Chuang Yu ⋅ Jinmiao Zhao ⋅ Yunpeng Liu ⋅ Yaokun Li ⋅ Xiujun Shu ⋅ Yuanhao Feng ⋅ Bo Wang ⋅ Xiangyu Yue
While large-scale visual foundation models (VFMs) exhibit strong generalization across diverse visual domains, their potential for infrared small target (SIRST) detection remains largely unexplored. To fill this gap, we systematically introduce frozen VFM representations into the SIRST task and propose a Foundation-Driven Efficient Paradigm (FDEP), a general framework compatible with diverse VFMs and SIRST networks, which improves detection accuracy without additional VFM-related inference overhead. Specifically, a Semantic Alignment and Modulated Fusion (SAMF) module is designed to achieve dynamic alignment and deep fusion of the global semantic priors from VFMs with task-specific features. Meanwhile, to avoid the inference-time overhead introduced by VFMs, we propose a Collaborative Optimization-based Implicit Self-Distillation (CO-ISD) strategy, which enables implicit semantic transfer between the main and lightweight branches through parameter sharing and synchronized backpropagation. In addition, to unify the fragmented evaluation system, we construct a Holistic SIRST Evaluation (HSE) metric that performs multi-threshold integral evaluation at both pixel-level confidence and target-level robustness, providing a stable and comprehensive basis for fair model comparison. Extensive experiments demonstrate that the SIRST detection networks equipped with our FDEP framework achieve state-of-the-art (SOTA) performance on multiple public datasets. Our code will be open source.
Rethinking Parallel Multi-Agent Systems: A Cost-Aware Framework for Efficient Coordination
Yexiong Lin ⋅ Shanshan Ye ⋅ Yu Yao ⋅ Zhen Fang ⋅ Bo Han ⋅ Tongliang Liu
LLM-based agent systems are increasingly used to solve complex multi-step tasks, where sequential execution incurs substantial end-to-end latency. In principle, parallelizing work across multiple agents should yield near-linear speedups. However, in practice, existing parallel multi-agent systems often run slower than a single-agent baseline. We attribute this gap to two hidden costs that parallel execution incurs but a serial agent avoids. First, there is a \emph{re-exploration cost}: redundant effort spent by parallel workers reconstructing context that the orchestrator already possesses, including prior decisions, conventions, and intermediate reasoning that would otherwise be inherited implicitly in a serial execution. Second, there is an \emph{alignment cost}: the overhead required to reconcile inconsistencies across independently generated outputs. Based on this decomposition, we derive a principled decision criterion: a layer should be parallelized only when its critical-path cost, plus re-exploration and alignment overheads, is lower than the corresponding serial cost. While this criterion is naturally expressed in wall-clock time, we observe that LLMs are poorly calibrated when asked to estimate task duration. Their predictions are strongly anchored to human engineering intuition rather than model throughput. The resulting bias is not monotonic, so even the relative ordering of task costs can be reversed between estimates and actual execution. To address this, we instead measure cost in predicted output tokens, a quantity that LLMs can estimate more reliably because it corresponds directly to their own generation behavior. For a fixed model, token cost also serves as a backend-independent proxy for time. Building on this token-based criterion, we propose \textsc{CostPar}. It estimates all token budgets in a single planning step, forks each worker directly from the orchestrator’s session to eliminate re-exploration cost, and replaces post-hoc reconciliation with a pre-generated shared convention block that converts alignment into a bounded upfront cost. A deterministic scheduler then applies the criterion layer by layer. Empirically, \textsc{CostPar} achieves a 2.2$\times$ mean throughput improvement and a 2.6$\times$ mean wall-time speedup over Claude Code, and a 2.0$\times$ throughput improvement over the strongest multi-agent baseline.
Rethinking Projector Training in Multimodal LLMs
Fabian Gröger ⋅ Naël Ouerghemi ⋅ Shuo Wen ⋅ Maria Brbic
Multimodal LLMs are typically trained by first aligning a vision encoder with a language model via projector training on large image-caption datasets, followed by joint training of the full multimodal model. Despite its widespread use, the role of the projector training stage remains poorly understood. Does it learn fine-grained visual-language correspondences, or mainly maps visual features into a compatible input space for the language model? We find that (i) using only $10$ to $20\%$ of the data for training the projector already recovers most downstream performance gains, and (ii) projectors trained on different datasets can be linearly interpolated while largely preserving performance. These results suggest that projector training acts primarily as a coarse alignment within a large solution space. Motivated by these findings, we propose PORTAL, a training-free projector initialization that requires neither paired image-caption data nor gradient-based optimization, but only per-modality summary statistics. PORTAL computes an optimal-transport-based map between Gaussian approximations of the vision and language feature distributions on a shared principal subspace, enabling completely skipping projector pretraining. Despite using no paired data, while the standard approach relies on 0.5M image–caption pairs, PORTAL matches the default trained pipeline across six LLM backbones, two vision encoders, and 16 benchmarks, outperforming it in $3$ of $7$ (vision, LLM) settings, and remaining within $1.6$ average points across all other settings.
Rethinking "RL Generalizes, SFT Memorizes": The Role of SFT Data
Yunlong Hou ⋅ Fengzhuo Zhang ⋅ Yuan Cheng ⋅ Jiachun Pan ⋅ Xingyao Li ⋅ Zhuoran Yang
Large language models trained with Reinforcement Learning (RL) with verifiable rewards exhibit strong reasoning ability and broad generalization, whereas models trained with Supervised Fine-Tuning (SFT) are often viewed as more prone to memorization and limited transfer. This paper rethinks this distinction through the lens of SFT training data. First, we study the role of data source and show that it is critical: a carefully mixed SFT dataset substantially outperforms data generated solely by a larger model. Second, we study the role of data scale and show that matching the number of correct rollouts between SFT and RL greatly improves SFT generalization, while matching the total rollout budget enables SFT to generalize as well as RL. Combining these two factors further enables SFT to generalize even better than RL. Third, by using LLM annotations to characterize the solution methods in training rollouts, we show that larger datasets cover more tail methods and that these tail methods provide generalizable reasoning signals. Finally, we support these empirical findings theoretically by analyzing the training dynamics of shallow transformers under both RL and SFT.
Rethinking Sequential Locate-Then-Edit: Optimality and Stability
Bingqing Liu ⋅ Wei Liu ⋅ Zhiying Deng ⋅ Jun Wang ⋅ Yuhua Li ⋅ Ruixuan Li
The \textit{Locate-Then-Edit} framework enables efficient knowledge updates in large language models without requiring full retraining. Recent research has increasingly focused on \textit{sequential editing}, where knowledge arrives continuously and the model is updated incrementally. However, existing sequential methods commonly suffer from catastrophic forgetting and model collapse after only a few hundred edits. In contrast, \textit{batch editing}---which performs joint optimization over all edits and represents the theoretical optimum---can support tens of thousands of edits with high edit efficacy while preserving the model's general capabilities, resulting in a substantial performance gap between the two editing paradigms. In this paper, we demonstrate that this gap can be closed via a \textit{batch-equivalent sequential editing} approach. Moreover, we investigate the underlying mechanism of model degradation in existing sequential methods by analyzing the convergence behavior of weight perturbations as edits accumulate. Our analysis reveals that current methods typically exhibit linear or logarithmic growth in perturbation, whereas the proposed batch-equivalent method achieves \textit{time-invariant} growth, thereby theoretically guaranteeing long-term stability and preserving the integrity of pretrained knowledge throughout massive-scale sequential editing.
Rethinking Softmax Attention: Polynomial Activations for Transformers
Hemanth Saratchandran ⋅ Jianqiao Zheng ⋅ Yiping Ji ⋅ Wenbo Zhang ⋅ Simon Lucey
Softmax attention is typically viewed as effective because it produces a normalized probability distribution over input tokens. In this paper, we challenge this view by showing that a key mechanism behind softmax attention is its implicit control of the Frobenius norm of the attention matrix, which stabilizes training. Motivated by this observation, we study alternative attention activations, focusing on polynomial maps that provide a similar regularization effect without satisfying the usual softmax constraints. Our theoretical analysis shows that certain polynomial activations can replace softmax despite violating positivity, row-wise normalization, and sparsity. Extensive experiments across transformer applications show that these alternatives achieve strong performance, suggesting that the success of softmax attention is not inherently tied to its probabilistic interpretation.
Rethinking State Tracking in Recurrent Models Through Error Control Dynamics
Jiwan Chung ⋅ Heechan Choi ⋅ Seon Joo Kim
The theory of state tracking in recurrent architectures has predominantly focused on expressive capacity: whether a fixed architecture can theoretically realize a set of symbolic transition rules. We argue that equally important is error control, the dynamics governing hidden-state drift along the directions that distinguish symbolic states. We prove that affine recurrent networks, a class of models encompassing State-Space Models and Linear Attention, cannot correct errors along state-separating subspaces once they preserve state representations. Consequently, practical affine trackers do not learn robust state tracking; rather, they learn finite horizon solutions governed by accumulated state-relevant error. We characterize the mechanics of this failure, showing that tracking remains readable only while the accumulating within-class spread remains small relative to the initial between-class separation. We demonstrate empirically on group state-tracking tasks that this breakdown is predictable: tracking collapses when the distinguishability ratio crosses the readability threshold of the trained decoder. Across trained models, the point of this crossing predicts the horizon at which downstream accuracy fails. These results establish that robust state tracking is determined not only by an architecture's theoretical expressivity but crucially by its error control.
Revealing Epistemic Uncertainty in MLLMs via Causal-Invariant Masking
Haoyang Luo ⋅ Linwei Tao ⋅ Jie Gui ⋅ Xinghao Chen ⋅ Chang Xu ⋅ Jianyuan Guo ⋅ Minjing Dong
Multimodal Large Language Models (MLLMs) suffer from hallucinations, creating a critical need for Uncertainty Quantification (UQ) to ensure reliable deployment. However, existing approaches struggle to detect uncertainty caused by superficial associations, especially when the query-relevant signal is weak. We mainly attribute this issue to their bias toward aleatoric uncertainty arising from data ambiguity, overlooking epistemic uncertainty stemming from model limitations. To further decompose uncertainty types for a comprehensive UQ, we propose Causal-Invariant Masking (CIM), which measures the semantic shift between the original predictions and those conditioned on a causally-focused view. Based on this framework, we introduce Semantic Divergence as our core metric for UQ and provide theoretical evidence that it converges to the variance of model's sensitivity to non-causal correlations, establishing its ability to capture MLLM's limitation. To accelerate UQ in MLLMs, we further propose Expected Embedding Drift (EED), a fast geometric proxy metric that estimates semantic shift directly within the hyperspherical embedding space. Experiments show that our method achieves state-of-the-art performance on various benchmarks, while the proposed EED accelerates by nearly 50\% with comparable performance.
Revisiting Diffusion Fine-Tuning for Unsupervised Domain Adaptation
Xuan Qi ⋅ Yi Wei ⋅ Daniele Berardini ⋅ Vito Paolo Pastore ⋅ Vittorio Murino
Diffusion-based unsupervised domain adaptation (UDA) improves cross-domain transfer by generating target-specific synthetic data for downstream adaptation. Existing methods are largely designed for single-target adaptation: when a model trained on one labeled source domain must be adapted to multiple unlabeled target domains, they typically require separate diffusion fine-tuning for each source–target pair, causing training, storage, and deployment costs to grow with the number of targets. In this paper, we study multi-target data generation for diffusion-based UDA, where a single source-guided diffusion fine-tuning process is reused to generate target-specific synthetic data for multiple target domains. We propose MUSE (Multi-target UDA-oriented Synthesis with Efficient diffusion fine-tuning), a decoupled adaptation framework that separates source-supervised semantic adaptation from target-specific style adaptation. MUSE uses a shared semantic branch updated by labeled source data and target-private style branches specialized to individual target domains, enabling target-specific generation while avoiding repeated source-guided fine-tuning for each target. Experiments on standard UDA benchmarks show that MUSE achieves a stronger accuracy–efficiency trade-off than repeated per-target diffusion adaptation, reducing diffusion fine-tuning cost while improving average target-domain accuracy.
R-GRec: Relation-Guided Generative Recommendation via Collaborative Graph Supervision
Xiangfu Meng ⋅ Jiapeng Yang ⋅ Lei Shi ⋅ xue wang ⋅ Yuefeng Zhan ⋅ Chen Gu ⋅ Hao Sun ⋅ Weiwei Deng ⋅ Tao Yao ⋅ Feng Sun ⋅ Qi Zhang
Generative recommendation has recently emerged as a promising paradigm for large-scale recommendation by reformulating next-item prediction as discrete sequence generation. However, current SID-based models rely mainly on history-to-next-item supervision, leaving item--item collaborative topology only indirectly captured through sparse user sequences. To address this gap, we propose R-GRec, a graph-derived structural supervision framework for SID-based generative recommendation. From user interactions, R-GRec builds an item--item graph and derives three complementary auxiliary objectives---Neighbor Prediction, Topological Contrast, and Link Prediction --- to make collaborative topology explicit during generative training. Extensive experiments on the two public Amazon datasets demonstrate that R-GRec achieves state-of-the-art performance over representative traditional, generative, and LLM-based recommenders. Further ablation and analytical studies verify the contribution of each graph-derived objective and show that collaborative topology acts as a complementary supervision signal to Semantic ID generation, improving generative recommendation without adding inference-time graph computation.
RheoSampling: Resolving the One-Hot Dilemma in Stochastic Dynamic-Tree Speculative Decoding
Qiao Hu ⋅ Yepeng Weng ⋅ Bo Zhang ⋅ Takehisa Yairi
Speculative decoding accelerates LLM inference by drafting multiple tokens in parallel, with tree-based methods further improving efficiency through structured hierarchies. Dynamic-tree methods such as EAGLE-3 achieve excellent performance under greedy decoding via deterministic top-$K$ expansion and global pruning. However, in stochastic decoding ($T>0$), this mechanism collapses the draft distribution into one-hot probabilities, causing a severe drop in acceptance rate. This exposes an apparent dilemma: dynamic-tree methods sacrifice stochastic sampling to preserve context-aware topology, while static-tree methods preserve stochastic sampling with context-agnostic structures. The issue arises because the same probability distribution is used for two conflicting tasks: constructing the tree and verifying the tokens. This coupling makes direct injection of randomness extremely challenging, as we are faced with a complex stochastic process. We resolve this by decoupling these two roles: RheoSampling assigns a token sampled from the draft distribution a \textit{proxy probability} (for tree expansion and pruning) alongside its \textit{true sampling probability} (for verification). Specifically, we inject a sampled token among the deterministic top-$K$ slots and treat it with different probabilities in the construction and verification process, making RheoSampling the first dynamic-tree method with both context-aware top-$K$ construction and stochastic sampling while maintaining losslessness. We establish the lossless guarantee through a novel equivalence-class analysis that compresses the stochastic tree space into tractable classes. An OT-based verification strategy and a sparse draft mechanism ensure that theoretical gains translate into practical efficiency. Experiments across diverse LLMs and benchmarks demonstrate consistent improvements in acceptance rate and speedup over state-of-the-art dynamic tree methods. This framework may provide a template for analyzing other complex stochastic tree structures.
RiboC2F: Pose-First Coarse-to-Fine Flow Matching for Protein-Conditioned RNA Co-Design
Dengdeng Huang ⋅ Shikui Tu
Computational design of ribonucleic acid (RNA) is essential for therapeutic discovery and synthetic biology. Designing RNAs that bind specified protein targets is bottlenecked by a scarcity of paired-complex 3D structural data. Existing approaches focus on importing structural priors from external biomolecular datasets, leaving where to localize the complex signal itself unresolved. We propose RiboC2F, a coarse-to-fine flow-matching framework that concentrates this signal first and foremost on the global binding pose. RiboC2F runs a two-stage cascade over a shared trunk, where a coarse stage commits the binding pose and a fine stage generates the full RNA design. The coarse pose enters as a conditioning signal rather than a fixed initialization, avoiding cross-stage error accumulation. On PRI30k, RiboC2F delivers the strongest docking among published baselines, with competitive intrinsic validity and multi-sample diversity. Probing further shows that the fine stage exploits the coarse pose's direction but is robust to pointwise coordinate corruption, confirming pose direction as the operative coarse-signal axis.
RIGOR: Risk-Gated Topology Adaptation for Robust LLM Multi-Agent Reasoning
Fengyuan Ran ⋅ Yikuan Wang ⋅ Yanming Li ⋅ Wenjie Lu ⋅ Tongtong Wu ⋅ Senquan Yi ⋅ Yuheng Wang ⋅ Qiqi Lin ⋅ Yuxin Wu ⋅ Minghui Zhou ⋅ Naiqiang Tan ⋅ Li Shen
LLM-based multi-agent systems (MAS) improve complex reasoning through collaborative decomposition, verification, and synthesis, but their performance can degrade sharply when compromised agents inject misleading intermediate messages. Existing topology methods rely on fixed communication patterns or query-adaptive graphs, but rarely infer or intervene runtime reliability risk, leaving them fragile when compromised agents vary across queries. We study stochastic adversarial topology adaptation, where compromised agents shift across graph positions and rare cascades dominate the high-loss tail. To address this setting, we propose RIGOR, a topology-adaptation framework that couples runtime risk inference with tail-aware training. RIGOR combines Risk-Gated Topology Construction (R-GTC), which probes agent behavior, estimates per-agent risk, and edits the graph through soft quarantine and edge masking, with CVaR-Guided Robust Optimization (C-GRO), which trains the editor on high-loss episodes rather than average gain alone. Across six benchmarks and two backbones, RIGOR consistently outperforms existing MAS baselines under query-varying misleading-message attacks, achieving an average 21.09\% accuracy improvement over the strongest baseline on Qwen3.5-35B-A3B. Tail-subset and position-wise analyses further show that these gains stem from improved robustness against worst-case cascades and dynamically changing compromised-agent positions. The code is available for anonymous access at https://anonymous.4open.science/r/RIGOR-0D7D.
RISE-Video: Can Video Generators Decode Implicit World Rules?
Mingxin Liu ⋅ Shuran Ma ⋅ Shibei Meng ⋅ Xiangyu Zhao ⋅ Zicheng Zhang ⋅ Shaofeng Zhang ⋅ Zhihang Zhong ⋅ Peixian Chen ⋅ Haoyu Cao ⋅ Xing Sun ⋅ Haodong Duan ⋅ Xue Yang
While generative video models have achieved remarkable visual fidelity, their capacity to internalize and reason over implicit world rules remains a critical yet under-explored frontier. To bridge this gap, we present RISE-Video, a pioneering reasoning-oriented benchmark for Text-Image-to-Video (TI2V) synthesis that shifts the evaluative focus from surface-level aesthetics to deep cognitive reasoning. RISE-Video comprises 467 meticulously human-annotated samples spanning eight rigorous categories, providing a structured testbed for probing model intelligence across diverse dimensions, ranging from commonsense and spatial dynamics to specialized subject domains. Our framework introduces a multi-dimensional evaluation protocol consisting of four metrics: \textit{Reasoning Alignment}, \textit{Temporal Consistency}, \textit{Physical Rationality}, and \textit{Visual Quality}. To further support scalable evaluation, we propose an automated pipeline leveraging Large Multimodal Models (LMMs) to emulate human-centric assessment. Extensive experiments on 11 state-of-the-art TI2V models reveal pervasive deficiencies in simulating complex scenarios under implicit constraints, offering critical insights for the advancement of future world-simulating generative models.
Robust Concept Unlearning in Diffusion Models via Directional Stability Regularization
Bo-Han Lai ⋅ Hsuan-Tien (Tien) Lin ⋅ Chia-Mu Yu ⋅ Han Zhao ⋅ Shang-Tse Chen
Post-hoc concept unlearning is a practical approach to remove undesired content from text-to-image diffusion models. However, existing methods leave unlearned models fragile: small perturbations to prompts, text embeddings, or sampling-time latent states can recover supposedly erased content, even when standard endpoint metrics suggest successful erasure under nominal prompts. We study this fragility through the stability of the denoising trajectory of diffusion models. Our key observation is that prompt-, embedding-, and sampling-based attacks share a similar mechanism of perturbing the sampling dynamics, and erased concepts can reappear when the unlearned model amplifies these concept-relevant perturbations. We formalize this view with a finite-time stability analysis and derive a tractable directional sharpness measure along a concept-relevant direction. Motivated by this analysis, we propose Directional Stability Regularization (DSR), a plug-in regularizer that discourages expansion in erased-concept directions without directly penalizing unrelated directions. DSR is compatible with classical score-based unlearning objectives, requires no adversarial training or attack-specific inner loop, adds no inference-time overhead, and can be estimated efficiently with Jacobian-vector products. Across the evaluated erasure settings, DSR reduces recovery under prompt-, embedding-, and sampling-based attacks while largely preserving prompt alignment and image quality metrics.
Robust Conditional Conformal Prediction via Branched Normalizing Flow
Rui Xu ⋅ Xingyuan Chen ⋅ Wenxing Huang ⋅ Minxuan Huang ⋅ Weiyan Chen ⋅ Sihong Xie ⋅ Hui Xiong
Conformal prediction (CP) constructs prediction sets with marginal coverage guarantees under the assumption that the calibration and test distributions are identical. However, under distribution shift, existing approaches primarily align marginal conformal score distributions, which is sufficient to preserve marginal coverage but does not control the conditional coverage error at individual test inputs. As a consequence, CP can remain unreliable in regions where the conditional score distributions are mismatched. In this work, we bound the conditional invalidity of CP under distribution shift in terms of the Wasserstein distance between the calibration and test distributions. This result highlights the role of invertible transport in mitigating conditional coverage degradation. Motivated by this insight, we introduce Branched Normalizing Flow (BNF), a two-branch architecture that normalizes a test input to the calibration distribution and transforms the prediction set of the normalized input back to the test distribution while preserving conditional guarantees. Empirically, BNF consistently improves conditional coverage robustness on nine datasets across a wide range of confidence levels.
Robust Latent Space Bayesian Optimization with Marginalized Kernel
Ziyi Yang ⋅ Yifan Zhu ⋅ Jinghui Zhong
Latent space Bayesian optimization circumvents the curse of dimensionality over structured inputs by performing Gaussian process surrogate-guided search in a continuous latent space learned by deep generative models. However, existing latent space Bayesian optimization methods overlook both the inherent stochasticity of latent space representations and the geometric properties of the manifold induced by VAE, leading to suboptimal optimization performance. In this work, we propose a marginalized kernel that integrates out the latent variables through the VAE inference posterior, yielding a closed-form kernel on the observation space that naturally accounts for encoding uncertainty. We adopt a quadratic polynomial kernel as the base kernel in the latent space, which captures discrepancies in the first two moments of the latent posteriors while avoiding the bandwidth sensitivity inherent to the RBF kernel. Furthermore, inspired by the Riemannian interpretation of VAEs, we incorporate a geometry-aware sampling scheme that leverages the metric structure revealed by the learned posterior covariances to guide candidate acquisition toward high-information regions of the latent space. Empirical evaluations on molecular design and robot design tasks demonstrate that our method outperforms existing state-of-the-art baselines across the majority of tasks.
Robust Multi-view Clustering against Imperfect Information
Zhichao Huang ⋅ Haochen Zhou ⋅ Hao Wang ⋅ Xi Peng ⋅ Mouxing Yang
Real-world multi-view data always suffer from imperfect information problem, where the view-specific observations are absent (\ie, Incomplete Views, IV) and cross-view correspondences are mismatched (\ie, Noisy Correspondences, NC) for certain instances. As a remedy, numerous IV- and NC-oriented multi-view clustering (MvC) methods have been proposed, which however require either reliable correspondences or sufficiently complete instances, thus stopping short of addressing the imperfect information problem. In contrast, we observe that both IV and NC challenges originate from the same issue of imperfect cross-view counterpart information, where the counterpart of an anchor instance in another view might be either unavailable or unreliable. Based on the observation, we propose a novel robust MvC framework, termed Posterior-guided Latent Counterpart Inference (PLCI), which could handle both IV and NC in a unified manner. Specifically, PLCI formulates the desired cross-view counterpart of each anchor instance as a latent variable, and integrates both instance-level reliability and prototype-level semantic transport to infer the posterior distribution of the latent counterpart. Extensive experiments on six widely-used multi-view datasets against 10 state-of-the-art MvC methods demonstrate the effectiveness of PLCI for tackling the imperfect information problem. The code will be released upon acceptance.
Robust Residual Correction via Selective Deployment for Time Series Forecasting
Jianxiang Xie ⋅ YUNCHENG HUA ⋅ Mingyue Cheng ⋅ Flora Salim ⋅ Hao Xue
Post-hoc residual correction can reduce forecasting error, but always-on updates often fail to generalize: when residual signals are weak or noisy, expressive correctors can overfit and even increase the overall forecasting errors. We introduce CRC, a robust residual correction framework that couples a hybrid corrector with a validation-calibrated selective deployment rule. CRC proposes corrections via a conservative ridge "floor" plus conditional nonlinear refinement, but deploys them only when held-out evidence indicates net error reduction over a frozen baseline forecaster. This reframes correction as a decision problem: when to revise a forecast, not only how to fit residuals. Across standard long-sequence benchmarks and multiple backbones, CRC improves horizon-averaged MSE and MAE in most dataset-backbone pairs, with particularly large gains on complex datasets (e.g., over 20% relative MSE reduction on Traffic). We additionally report the non-degradation rate (NDR) as a stability diagnostic of harmful updates and ablate the validation-calibrated "firewall" that stabilizes MSE gains. Formal analysis motivating the design appears in Appendix A under idealized assumptions.
Reversible adversarial examples (RAEs) can disrupt access by malicious AI models while ensuring recoverability for authorized users. However, practical image circulation is dominated by JPEG files, whereas most existing RAE methods perturb spatial-domain images and are therefore poorly matched to compression, re-saving, and coefficient quantization. Beyond JPEG compatibility, circulation imposes a stricter and underexplored requirement that protected images should be perturbed in the JPEG domain, remain adversarial after spatial reconstruction, preserve visual fidelity, incur modest file-size growth, and remain recoverable under channel attacks such as recompression, noise, cropping, resizing, and platform re-saving. This paper presents the first systematic study of robust reversible adversarial protection for post-processed JPEG images and proposes SRAP-JPEG, a Synchronized Reversible Adversarial Protection framework on quantized DCT coefficients. SRAP-JPEG decouples adversarial perturbation from reversible recording by injecting perturbations into luminance coefficients, hiding recovery records in paired chrominance coefficients through a secret-key mapping, and adopting coefficient-adaptive allocation to trade off reversible capacity, attack strength, visual quality, and file-size growth. It crafts protective perturbations directly in the JPEG coefficient domain by propagating gradients from classification or vision-language objectives through the chain rule. For authorized recovery, SRAP-JPEG integrates block-phase synchronization, embedding-state identification, and coefficient restoration, enabling exact recovery of intact protected JPEGs and high-fidelity recovery after practical distortions. Experiments on ImageNet and MS-COCO show that SRAP-JPEG achieves an 81.82\% average attack success rate on ten mainstream classifiers, degrades CLIP-based image-text retrieval, and recovers distorted protected images with high visual fidelity, including 41.60 dB PSNR after JPEG recompression at QF 70 and 59.51 dB on the aligned uncropped region after edge cropping at ratio 0.25.
Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime
Tianshu Zhu ⋅ Wenyu Zhang ⋅ Xiaoying Zuo ⋅ Tianlun ⋅ Haotian Zhao ⋅ yucheng-zeng ⋅ Jingnan gu ⋅ Daxiang Dong ⋅ Jianmin Wu
Agentic reinforcement learning (RL) for software engineering spends much of its compute on stateful trajectories whose grouped binary rewards are highly skewed and weakly contrastive. We frame this as pass-rate control and show that the binary reward-side signal is strongest near a 50% rollout pass rate under four criteria: reward entropy, group-filtering survival, leave-one-out (RLOO) advantage energy under Group Relative Policy Optimization (GRPO), and success-failure pair count. We propose *Prefix Sampling* (PS), which replays self-generated trajectory prefixes to steer skewed groups toward this regime: successful prefixes give mostly failing groups a head start, while failing prefixes handicap mostly passing groups. Replayed states are reconstructed through the existing rollout path, and replayed tokens are masked from the loss so optimization applies only to current-policy continuations. On SWE-bench Verified, PS reaches the baseline high-score regime within evaluation variability while delivering $2.01\times$/$1.55\times$ end-to-end wall-clock speedups on Qwen3-14B/32B; the 14B peak improves from $0.274$ to $0.295$. AIME 2025 experiments on 4B/8B show the same pass-rate-control pattern, and 4B ablations attribute gains to replay, bidirectional coverage, and adaptive control.
Rotary Position Embedding (RoPE) is one of the most widely used positional mechanisms in modern large language models. Because RoPE applies position-dependent rotations whose attention score can be rewritten in terms of relative offsets, subsequent work has sometimes classified it as a relative position embedding. In this position paper, we argue that RoPE is not a proper relative position embedding. Treating RoPE as a relative position embedding can create overly strong expectations about length extrapolation, context extension, and comparisons with other positional mechanisms. In practice, however, these expectations have not been borne out: RoPE-based long-context models typically rely on position interpolation, continued training, or fine-tuning. This is more consistent with an absolute-position-conditioned view of RoPE. We argue that this reframing, which is closer to the original formulation, offers a more precise conceptual foundation for analyzing RoPE, understanding long-context behavior, and comparing positional mechanisms in modern LLMs. To make this distinction precise, we formalize the distance-decay property that has been implicitly expected of relative position embeddings and show that RoPE does not satisfy it.
Rotating a molecule should leave its energy unchanged while rotating each force vector by the same rotation, as required in $SO(3)$-equivariant molecular learning. However, conventional normalization layers in neural networks are designed for scalar activations; when applied to vector channels through component-wise centering or scaling, they can break this symmetry. We study normalization modules for equivariant molecular networks under force supervision, where invariant scalar channels and covariant vector channels require different geometric treatment. We formalize an $SO(3)$-equivariance-preserving recipe that computes vector scale statistics from rotation-invariant squared norms and applies the resulting rescaling identically across Cartesian coordinates, which preserves $SO(3)$-equivariance without vector shifts. Building on this recipe, we introduce grouped vector RMS (GroupRMS), which shares vector RMS estimates across channel groups to reduce estimator variance and stabilize vector scaling in small-batch training. The proposed variant is lightweight and compatible with standard scalar normalization branches. We instantiate these ideas in practical scalar-vector normalization blocks for PaiNN-style backbones and evaluate them across several molecular benchmarks. On the paired three-trajectory subset, the strongest GroupRMS-family variant, GroupNorm-GroupRMS, improves average force MAE from 0.1190 to 0.0383 and average energy MAE from 0.0675 to 0.0424 relative to BatchNorm. These results support the proposed rotation-compatible vector scaling and grouped scale estimation as effective tools for robust molecular force learning.
RotVLA: Rotational Latent Action for Vision-Language-Action Model
Qiwei Li ⋅ Xicheng Gong ⋅ Xinghang Li ⋅ Peiyan Li ⋅ Quanyun Zhou ⋅ Hangjun Ye ⋅ Jiahuan Zhou ⋅ Yadong Mu
Latent Action Models (LAMs) have emerged as an effective paradigm for handling heterogeneous datasets during Vision-Language-Action (VLA) model pretraining, offering a unified action space across embodiments. However, existing LAMs often rely on discrete quantization encode and decode pipelines, which can lead to trivial frame reconstruction behavior, limited representational capacity, and a lack of physically meaningful structure. We introduce RotVLA, a VLA framework built on a continuous rotational latent action representation. Latent actions are modeled as elements of ${\rm SO}(n)$, providing continuity, compositionality, and structured geometry aligned with real-world action dynamics. A triplet frame learning framework further enforces meaningful temporal dynamics while avoiding degeneration. RotVLA consists of a VLM backbone and a flow-matching action head, pretrained on large-scale cross-embodiment robotic datasets and human videos with latent-action supervision. For downstream robot control, the flow-matching head is extended into a unified action expert that jointly denoises latent and robot actions. Here, latent actions serve as a latent planner, providing high-level guidance that conditions action generation. With only 1.7B parameters and 1700+ hours of pretraining data, RotVLA achieves 98.2\% on LIBERO and 89.6\% / 88.5\% on RoboTwin2.0 under clean and randomized settings, respectively. It also demonstrates strong real-world performance on manipulation tasks, consistently outperforming existing VLA models.
Route-Consistent Adaptation for Stable Quantization of Mixture-of-Experts Models with Theoretical Guarantees
Longteng Zhang ⋅ Sen Wu ⋅ Qiang Wang ⋅ Shaohuai Shi ⋅ Xiaowen Chu
Quantizing Mixture-of-Experts (MoE) models is difficult for two coupled reasons: expert quantization error perturbs expert outputs, and sparse routing amplifies small perturbations into discrete expert selection changes. This coupling produces progressive output drift that is not captured by dense model quantization analyses. To address this gap, we propose a unified quantization framework that jointly handles both sources of error. First, we formulate compression-induced routing inconsistency as a form of training-inference router mismatch and use this connection to motivate limited routing replay, which replays BF16 routes during quantization-aware fine-tuning (QAT), freezes router parameters, and updates only expert-side adapter parameters. We then provide a standard convergence guarantee for the resulting fixed-route QAT objective, together with a routing-consistency bound under explicit router-margin conditions. Second, for expert restoration under a fixed rank budget, we propose load-aware low-rank compensation that assigns larger compensation rank to heavily loaded experts. Under this allocation rule, we prove monotone strict decrease of weighted compression error and characterize optimal rank assignment by marginal load-weighted spectral gains. Finally, experiments on the MXFP4 quantization of Mixtral, DeepSeek, and Qwen3 MoE models show that our method improves perplexity and downstream accuracy over prior methods. We also report system efficiency on Qwen3-30B, where uniform W4A8 serving accelerates decoding and our adapter path preserves most of that acceleration.
RPC-GS: Gaussian Splatting with native RPC Rendering for Satellite Imagery
Valentin Wagner ⋅ Sebastian Bullinger ⋅ Christoph Bodensteiner ⋅ Michael Arens
We present RPC-GS, the first Gaussian Splatting framework for satellite imagery that operates natively with Rational Polynomial Camera (RPC) models. The RPC model is the de facto standard for representing the complex imaging geometry of modern pushbroom satellite sensors. To simplify rendering, prior satellite Gaussian Splatting methods replace the RPC model with perspective or affine camera approximations, leading to geometric errors during reconstruction. RPC-GS avoids these approximations by projecting Gaussian means and covariances directly through the RPC model during the splatting process. We embed the RPC model in a chain of carefully selected geo-coordinate transformations representing a mapping from splatting-suitable scene coordinates to image coordinates. To map the Gaussian covariance matrices, we derive a numerically robust Jacobian-based covariance projection for the (partially nonlinear) coordinate transformations. Since RPCs lack an explicit notion of camera depth, we integrate a metric ray-based depth formulation. We benchmark RPC, perspective, and affine camera models in a unified framework, with our native RPC renderer consistently achieving the lowest reconstruction error on leading satellite benchmark datasets, improving mean altitude error over perspective and affine approximations by 29.6\% and 63.8\% on DFC2019, and by 9.9\% and 37.9\% on IARPA2016. We release our code to support future research of Gaussian Splatting in the satellite imaging domain.
RSRCC: A Remote Sensing Regional Change Comprehension Benchmark Constructed via Retrieval-Augmented Best-of-𝑁 Ranking
Roie Kazoom ⋅ Yotam Gigi ⋅ George Leifman ⋅ Tomer Shekel ⋅ Genady Beryozkin
Traditional change detection identifies where changes occur, but does not explain what changed in natural language. Existing remote sensing change captioning datasets typically describe overall image-level differences, leaving fine-grained localized semantic reasoning largely unexplored. To close this gap, we present RSRCC, a new benchmark for remote sensing change question-answering containing 126k questions, split into 87𝑘 training, 17.1𝑘 validation, and 22𝑘 test instances. Unlike prior datasets, RSRCC is built around localized, change-specific questions that require reasoning about a particular semantic change. To the best of our knowledge, this is the first remote sensing change question-answering benchmark designed explicitly for such fine-grained reasoning-based supervision. To construct RSRCC, we introduce a hierarchical semi-supervised curation pipeline that uses Best- of-𝑁 ranking as a critical final ambiguity-resolution stage. First, candidate change regions are extracted from semantic segmentation masks, then initially screened using an image-text embedding model, and finally validated through retrieval-augmented vision-language curation with Best-of-𝑁 ranking. This process enables scalable filtering of noisy and ambiguous candidates while preserving semantically meaningful changes. The dataset is available at https://huggingface.co/datasets/google/RSRCC.
Rubric-Align: Safety Alignment through Dynamically Co-Evolving Rubrics
Ruipeng Wang ⋅ Junfeng Fang ⋅ Houcheng Jiang ⋅ Kai Tang ⋅ Pengyu Cheng ⋅ xiaoxi jiang ⋅ Guanjun Jiang ⋅ Xiang Wang
Recent progress in large language models (LLMs) has led to impressive performance across diverse domains. However, their deployment in critical areas such as healthcare, law, and education raises serious concerns regarding the potential generation of harmful content. While reinforcement learning (RL)-based safety alignment methods have been extensively studied, they face persistent challenges, including sparse reward signals, limited interpretability, and limited transferability across domains. In this paper, we introduce Rubric‑Align, a framework that replaces static binary scalar rewards with dynamic, natural‑language evaluation rubrics. Instead of assigning a single reward to each response, Rubric‑Align provides fine‑grained and interpretable feedback, thereby alleviating reward sparsity and improving the transparency of the alignment signal. A key feature of Rubric-Align is its dynamic rubric evolution mechanism. Rather than using fixed reward templates, Rubric-Align periodically refines prompt-specific rubrics based on the current policy behavior, keeping the supervision signal informative and aligned with the model’s evolving failure modes. Experiments on safety benchmarks and vertical-domain settings show that Rubric-Align improves robustness against harmful and jailbreak prompts while largely preserving general capabilities, suggesting that natural-language rubrics provide a transferable interface for adapting safety supervision to new safety subdomains.
RxnOptBench: Benchmarking LLMs for Reaction-Condition Optimization in Organic Methodology
Lingli Ge ⋅ Junyuan Gao ⋅ Jiahe Song ⋅ Jiaxing Sun ⋅ Yubin Wang ⋅ Boyu Zhu ⋅ Haote Yang ⋅ Jingchao Wang ⋅ Lixin Ma ⋅ Jiang Wu ⋅ Yuqiang Li ⋅ Conghui He
Chemical reaction-condition optimization --- choosing the catalyst, ligand, solvent, reagent, temperature, time, and atmosphere that jointly maximize yield and stereoselectivity --- is a central, judgement-laden subtask of organic methodology research that large language models are increasingly expected to support. Yet existing chemistry benchmarks evaluate reaction-class labelling, retrosynthesis, or SMILES manipulation, and do not ask models to read a real condition-screening table and pick the best set. We introduce RxnOptBench, a benchmark whose every option and precedent is a real wet-lab entry mined from the optimization tables of organic-methodology papers published in 2025, graded by a continuous relative score that fuses yield with stereoselectivity (ee, dr, rr) so near-correct answers are not collapsed to zero, and equipped with a paired precedents-vs-no-precedents design that isolates in-context use of literature evidence from parametric memorization. Across nine frontier LLMs and three Chemistry LLMs, even the best models leave substantial headroom: chemistry-specialized models fall to the random-baseline floor on multi-axis selection, while open-weight models have closed most of the gap to proprietary frontier models. We release the benchmark, the human-verified extraction pipeline, and the inference and evaluation code.
S$^2$-RL: Sample-Set Dual Reinforcement Learning for Generative Semantic Segmentation Dataset Distillation
Haoyu Wang ⋅ Fei Zhou ⋅ Qingqing Qiu ⋅ Lei Zhang ⋅ Wei Wei ⋅ Chen Ding
While dataset distillation has witnessed significant advances in image classification, its extension to dense prediction tasks such as semantic segmentation is plagued by two core bottlenecks: bi-level optimization-based approaches suffer from prohibitive computational costs stemming from pixel-wise gradient unrolling, whereas proxy-based generative methods merely infuse visual textures into fixed semantic masks derived from real data, thus failing to condense rich semantic knowledge into a compact set of informative, novel scene generations. To address both bottlenecks simultaneously, we propose a novel Sample-Set Reinforcement Learning (S$^2$-RL) framework for generative semantic segmentation dataset distillation (SSDD). Specifically, S$^2$-RL formulates SSDD as a diffusion model based text-to-image generation paradigm, enabling flexible generation of images with informative, novel semantic distributions via a concise text prompt. Furthermore, we fine-tune a diffusion policy via GRPO, leveraging a sample-set dual reward paradigm: for individual samples, it enforces semantic alignment with input text prompts and intra-class feature diversity; for sample sets, it maximizes inter-sample semantic distribution diversity by solving a contextual multi-armed bandit problem. These design choices enable S$^2$-RL to distill large-scale semantic segmentation datasets into a compact set of generated samples that encapsulate diverse, informative semantic knowledge without the need of pixel-wise gradient unrolling. Extensive experiments on the ADE20K and COCO datasets demonstrate that S$^2$-RL achieves substantial and consistent improvements over state-of-the-art baselines, establishing a new benchmark for SSDD.
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF
Wei Chen ⋅ Guanghui Zhu ⋅ Yafei Li ⋅ Limin Wang ⋅ Yihua Huang
Reinforcement learning from human feedback (RLHF) with preference-based reward models often exhibits unstable training dynamics. A key contributing factor is that standard RLHF relies on a single sequence-level scalar reward, which is propagated to token-level policy updates and leaves credit assignment within a response inherently ambiguous. Recent work has attempted to address this issue by refining rewards into denser token-level supervision, often relying on the implicit assumption that finer-grained credit assignment improves optimization. We argue that this assumption is incomplete: when preference signals are noisy and only defined at the response level, overly fine-grained reward refinement can amplify reward uncertainty and destabilize learning. To address this problem, we propose a granularity-aware principle for hierarchical credit assignment, emphasizing stability-oriented reward design rather than maximal allocation precision. Under this principle, sentences serve as a natural intermediate granularity, balancing semantic coherence with robustness to token-level noise. Guided by this view, we introduce S2T-RLHF. This sentence-to-token reward decomposition framework first allocates sequence-level preference rewards across sentences and then applies bounded token-level refinement within each sentence, without reward-model retraining or token-level supervision. Experiments across multiple datasets and optimization settings show that S2T-RLHF improves training stability and robustness while maintaining competitive preference alignment.
Saddle-to-Saddle Dynamics in Self-Supervised Shortcut Learning
Juhwan Kim ⋅ yoonsoo nam ⋅ Sungyoon Lee
Although the mechanics of eigenvalue bias of self-supervised learning (SSL) in linear networks are well-understood, they remain theoretically disconnected from empirical shortcut learning, offering little guidance or intuition. In this paper, we present a theoretical analysis of this shortcut learning phenomenon through the lens of \textit{extent bias} and \textit{amplitude bias}, two forms of \textit{eigenvalue bias}. By investigating the relations among extent bias, amplitude bias, and learning priorities in SSL, we demonstrate that SSL prioritizes features based on dimensionality (object to image ratio) and amplitude (backdoor attack intensity), not their semantic importance. Our analysis reveals how the eigenvalues of the feature cross-correlation matrix influence which features are learned earlier, providing insights into why models preferentially learn shortcut features over more generalizable features.
Safe Active Learning with Future Viability Guarantees in Time-Series Models
Hyeonjun Park ⋅ Whiyoung Jung ⋅ Deunsol Yoon ⋅ Sunghoon Hong ⋅ Kanghoon Lee ⋅ Kyungjae Lee
In time-series systems, safe active learning seeks to collect informative data while maintaining safety during data collection. A key challenge is that one-step safety does not guarantee future viability: an input can be immediately safe yet lead the system into a state from which no safe continuation exists. For example, in pressure regulation, an input may keep the current pressure below its limit but still drive the system toward an unsafe pressure spike over the next few steps. We study this problem in nonlinear time-series systems with unknown dynamics and safety constraints. We introduce the notion of $m$-step dead-end inputs and propose a future-aware safe active learning framework (FA-SAL) that enforces finite-horizon safety via margin-certified rollout constraints. Our approach constructs a conservative approximation of the true future-safe set by combining Gaussian process uncertainty with a recursive error bound over predicted trajectories. We show that FA-SAL provides high-probability safety guarantees, recovers interior future-safe inputs as uncertainty decreases, and provably excludes separated dead-end inputs. Experiments on synthetic and real-world benchmarks demonstrate that FA-SAL reduces safety violations and dead-end selections while maintaining competitive model learning performance.
In biological evolution, unconstrained mutation can lead to catastrophic outcomes: organisms may evolve enhanced capabilities while losing essential functions for survival. Nature's solution is \textit{developmental constraints}, where core regulatory genes remain anchored while peripheral genes adapt freely. We observe that current self-evolution algorithms for large language models lack analogous constraints. They optimize purely for capability, implicitly assuming safety will be preserved. Our experiments reveal this assumption to be dangerously wrong: models can \textit{misevolve} into powerful yet dangerous entities. Inspired by how Hox genes anchor body structure across $500$ million years of evolution, we propose \textbf{Circuit-Anchored Evolution (CAE)}. Using mechanistic interpretability, we identify a tiny \textit{safety circuit}, comprising less than $2$\% of model features, that causally mediates safety behaviors. We anchor this circuit during evolution, constraining it within a small displacement bound while allowing the remaining features to evolve freely. This mirrors the biological principle of \textit{evolvability with constraint}: preserving what is essential while adapting what is peripheral. Experiments across three model families and two evolution algorithms demonstrate that CAE achieves superior safety preservation with minimal capability loss, substantially outperforming explicit reward-based constraints in both effectiveness and efficiency. Just as developmental constraints prevent biological evolution from producing nonviable organisms, circuit anchoring prevents model evolution from producing capable but dangerous systems.
Safe-Fair MACPO: Burden-Fair Constrained Policy Optimization for Safe Multi-Agent Reinforcement Learning
Ankita Kushwaha ⋅ KIRAN RAVISH ⋅ Preeti ⋅ Pawan Kumar
Safe multi-agent reinforcement learning usually constrains team-level safety cost, but a safe team can still assign the same agent to repeatedly wait, yield, detour, absorb intervention, or lose access to a scarce resource. We study this failure mode as *burden unfairness under hard safety*. Safe-Fair multi-agent constrained policy optimization (Safe-Fair MACPO) augments MACPO with agent-wise burden logging, a temporal fairness-debt state, a fairness critic, and a two-cost trust-region update. Safety is lexicographically primary: recovery may override fairness to avoid unsafe actions, while the resulting imbalance is stored as debt and penalized later. We give a safety-first Pareto theory showing that the exact safe-fair constrained selector is strongly Pareto efficient for reward, safety, and fairness, with hard-safety and temporal burden-gap certificates. Empirically, the primary reported rows keep zero hard violations while improving useful burden balance: HalfCheetah-2x3-Fair improves selected safe return from $2145.5$ to $2637.4$ and Jain burden from $0.903$ to $0.977$; Safety-Gym MultiGoal-Fair reaches the lowest matched-budget scalar fairness cost with zero safety cost; and VMAS-SafeMultiGoal improves over recovery-only attribution controls. We separate social/resource fairness from mechanical workload imbalance and reserve mixed or recovery-explained rows for qualified appendix analysis.
Safety Reconstructed: Generative Modeling via Masked Diffusion Builds Strong Safety Guardrails
Gert Lek ⋅ Abele Mălan ⋅ Chaoyi Zhu ⋅ Pin-Yu Chen ⋅ Robert Birke ⋅ Lydia Chen
Guard models are the last line of defense between a language model and a harmful output, yet their training objective is surprisingly narrow. Existing guards learn to predict a single verdict token from a conversational context, concentrating supervision on a single target. The consequences are structural: models latch onto shortcut features, are overconfident, and remain sensitive to where safety evidence appears in the sequence rather than its role in the full context. We propose a different framing. Rather than predicting a label from text, our $\textbf{LLaDA Guard}$ asks which label better explains the text: scoring the prompt or response under each label hypothesis and classifying based on their difference. This shifts supervision to every token in the moderated region, forcing the model to account for full content rather than its most discriminative fragments. We instantiate this idea with a masked diffusion language model, fine-tuning LLaDA-8B-Instruct with a class-conditional reconstruction objective using LoRA and requiring no architectural changes beyond the base model. LLaDA Guard leads on average rank against discriminative baselines trained on stronger backbones across seven held-out safety benchmarks, while exhibiting substantially better confidence calibration (ECE 0.0875 vs. 0.1384 for Qwen3Guard), less over-defense on benign prompts with $\texttt{unsafe}$-looking cues, and less prompt leakage when moderating responses. Its generative nature further enables token-level risk localization as a natural byproduct, yielding a pipeline for rewriting $\texttt{unsafe}$ prompts into $\texttt{safe}$ equivalents without additional training and achieving a 60.7% average conversion-to-$\texttt{safe}$ rate. The weights and code for reproduction are available at the following [anonymized repository](https://anonymous.4open.science/r/LLaDA-Guard2026-B056/).
SAGE: Scalable Automated Robustness Augmentation for LLM Knowledge Evaluation
Xiaoyuan Li ⋅ Yuzhe Wang ⋅ Moxin Li ⋅ Keqin Bao ⋅ Rui Men ⋅ Yichang Zhang ⋅ Dayiheng Liu ⋅ Wenjie Wang ⋅ Fuli Feng
Large Language Models (LLMs) achieve strong performance on standard knowledge evaluation benchmarks, yet recent work shows that their knowledge capabilities remain brittle under question variants that test the same knowledge in different forms. Robustness augmentation of existing knowledge evaluation benchmarks is therefore necessary, but current LLM-assisted generate-then-verify pipelines are costly and difficult to scale due to low-yield variant generation and unreliable variant verification. We propose SAGE (Scalable Automated Generation of Robustness BEnchmarks), a framework for scalable robustness augmentation of knowledge evaluation benchmarks using fine-tuned smaller models. SAGE consists of VariantQual, a rubric-based verifier trained on human-labeled seed data, and VariantGen, a variant generator initialized with supervised fine-tuning and further optimized with reinforcement learning using VariantQual as the reward model. Experiments on HellaSwag show that SAGE constructs a large-scale robustness-augmented benchmark with quality comparable to the human-annotated HellaSwag-Pro at substantially lower cost, while the fine-tuned models further generalize to MMLU without benchmark-specific fine-tuning.
SAMoR: Motion Modelling for Articulated Objects of Any Skeleton and Topology
Yuhao Zhang ⋅ Gerard Pons-Moll ⋅ Tolga Birdal
Modeling motion for articulated objects of arbitrary skeleton topology remains difficult: existing motion generators target a fixed human skeleton, and prior adaptations either fail to share a vocabulary across rigs or discard motion detail through global pooling. Our key observation is that while joint-level motion does not correspond cleanly across species, motion of functional joint groups does---a human arm, a wolf foreleg, and a bird wing share semantic motion structure despite differences in joint count and connectivity, a correspondence that joint names (e.g., ``forearm'', ``wing\_L1'') partially expose even when topology does not. We introduce SAMoR (Skeleton-Aware Motion Representation for Articulated Objects), a cross-topology motion representation that encodes each motion segment as a small fixed number ($K{=}8$) of part tokens shared across arbitrary skeletons. SAMoR takes three input signals---per-joint motion features, kinematic graph structure, and joint-name embeddings---and processes them with a graph-transformer encoder, then compresses the resulting heterogeneous per-joint features into part-level tokens via cross-attention pooling and residual vector quantization, yielding a discrete motion codebook shared across rigs. To prevent the part queries from collapsing into redundant global representations, we introduce a topology-agnostic attention supervision loss, combined with random joint-name dropout to prevent over-reliance on text labels; together these encourage the part tokens to cluster joints into functional groups from names, structure, and motion jointly. We curate a unified heterogeneous motion corpus from HumanML3D, Truebones Zoo, and animated Objaverse-XL assets, and evaluate SAMoR on held-out characters with unseen skeletons. The resulting representation supports accurate reconstruction and cross-topology motion transfer, and further enables text-conditioned generation and localized part-wise editing through a MaskGIT token generator. SAMoR reaches $2.75\!\times\!10^{-2}$ normalized MPJPE on cross-topology reconstruction---$5.8\!\times$ below the strongest adapted variable-$J$ tokenizer baseline---and remains competitive with fixed-skeleton specialists on HumanML3D for both VQ-VAE reconstruction and text-to-motion generation.
SANEval: Open-Vocabulary Compositional Benchmarks with Failure-mode Diagnosis
Rishav Pramanik ⋅ Ian Nielsen ⋅ Jeffrey Smith ⋅ Saurav Pandit ⋅ Ravi P Ramachandran ⋅ Zhaozheng Yin
Text-to-image (T2I) models still fail on compositional prompts that require multiple objects, correctly bound attributes, accurate counts, or specific spatial relations, yet the benchmarks used to measure this progress are themselves limited: they rely on fixed-vocabulary detectors (typically the **80** MS-COCO classes), produce a single opaque score, and are rarely validated against human judgment. We introduce **SANEval** (Spatial, Attribute, and Numeracy Evaluation), an open-vocabulary compositional benchmark with two contributions: (i) a public dataset of $\sim$**5,250** prompts paired with images from six state-of-the-art T2I models, partitioned into a Simple split (algorithmically generated, human-validated) and a Hard split (fully human-written), and labeled by compositional task type (spatial, numeracy, color, shape, texture), and (ii) a modular evaluation pipeline that combines an LLM-based prompt parser, an open-vocabulary detector with LLM-driven synonym expansion and mapping, and three scorers for attribute binding, spatial relations, and numeracy. Each scorer emits both a numerical score and structured diagnostic feedback identifying what is missing, extra, or mis-bound. We validate SANEval against human ratings ($n=500$, **4** annotators, **5** categories): SANEval achieves a positive Pearson correlation with humans in every category (avg $r=0.222$), with **4** of **5** category-level correlations significantly above zero (Fisher **95%** CIs excluding **0**); the strongest prior compositional benchmark averages near zero ($r=-0.016$), with **4** of **5** category CIs spanning zero. SANEval is, to our knowledge, the only compositional T2I benchmark whose correlation sign agrees with humans on every axis. A manual audit of **10,040** detections decomposes the pipeline into a perception stage and an LLM/VLM verification stage and measures end-to-end precision at **96.6%**; the verification stage (synonym mapping plus VLM-as-judge attribute reading) contributes only **1.05%** of audited false positives, while removing it collapses scores by **43–76%** in ablation — substantial detection work at modest precision cost. The dataset is publicly released at [https://huggingface.co/datasets/saneval-ann/saneval-release](https://huggingface.co/datasets/saneval-ann/saneval-release) and the evaluation pipeline at [https://anonymous.4open.science/r/saneval-anon](https://anonymous.4open.science/r/saneval-anon).
SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking
Zhiwei Li ⋅ Lei Zhu ⋅ Hao Gu ⋅ Xiang Hu ⋅ Yan Wang ⋅ Haitao Mi ⋅ Sirui Han ⋅ Leo Liang ⋅ Zhijiang Guo
Post-training attention sparsification reduces the quadratic cumulative attention cost of pretrained Transformers by selecting a small set of context units (tokens or blocks) for each query. Existing trainable methods usually rely on a lightweight selector to score context units followed by a hard Top-$K$ selection, which blocks gradients from the language modeling loss. As a result, these methods commonly resort to distilling layer-wise dense attention distributions. While this approach encourages the selector to rank context units according to dense attention weights in the original model, such a ranking is not directly aligned with their impact on the model's final predictions under a fixed attention budget (ie, the number of attended context units per query), which can waste the limited attention budget on less useful units. To address this ranking misalignment, we propose **S**imple **A**ttention **S**parsification (SAS), a gated sparse attention mechanism that optimizes context ranking end-to-end with the language modeling loss. The key idea is to inject the selector's continuous scores into the attention logits during training, allowing the language modeling loss to update the selector through standard backpropagation. We identify several choices that are crucial to make this simple design work well in practice: placing the gate inside the attention $\operatorname{softmax}$ in log form, keeping all context units active during training so each of them can receive gradients, and preserving continuous selector scores so the model learns relative priorities rather than only hard selections. Since a naive implementation that explicitly materializes the full attention matrix would cause prohibitive memory overhead for long sequences, we implement a memory-efficient Triton kernel that integrates the design into a FlashAttention-style computation. SAS consistently outperforms trainable sparse-attention across attention budgets, with especially large accuracy gains under tight budgets, demonstrating more effective context ranking for downstream tasks.
Scalable Maximum Entropy Reinforcement Learning for Diffusion Policies via Adjoint Matching
Serge Thilges ⋅ Onur Celik ⋅ Denis Blessing ⋅ Emiliyan Gospodinov ⋅ Gerhard Neumann
Diffusion policies have recently emerged as a powerful paradigm for representing complex action distributions in reinforcement learning (RL). However, their application to online RL remains limited by the challenge of scalable training in the absence of ground-truth data, where standard optimization techniques such as score matching are not directly applicable. In this work, we introduce a highly efficient algorithm for optimizing diffusion policies by leveraging recent advances in stochastic optimal control. Our approach is based on adjoint matching, which enables simulation-free training and circumvents the need for explicit likelihood estimation or costly backpropagation through the diffusion process. Furthermore, we propose several extensions that improve the robustness and stability of the method in practical settings. Empirical results demonstrate that our approach achieves competitive performance while significantly reducing computational overhead, making diffusion policies more viable for online RL scenarios.
Scaling Limits of Long-Context Transformers
Giuseppe Bruno ⋅ Chen ⋅ Zhengjiang Lin ⋅ Yury Polyanskiy ⋅ Philippe Rigollet
We study the long-context limit of softmax self-attention with a fixed query and a random context of $n$ i.i.d. keys on the sphere, viewing the inverse temperature $\beta_n$ as the scaling parameter that decides whether attention degenerates into uniform averaging or collapses onto the single closest key. We show that the critical scale at which selectivity emerges is determined by the local exponent of the distance-to-query distribution near zero rather than by global features of the context, and scales like $\beta_n^\ast \asymp n^{2/(d-1)}$ for uniform keys on $\mathbb{S}^{d-1}$. Furthermore, we characterize the limiting laws of the ordered attention weights and of the attention output across all regimes of $\beta_n$: a subcritical regime in which the output reduces to a local average around $q$ with explicit deterministic bias and Gaussian fluctuations; a critical regime in which a finite collection of nearest keys retains macroscopic mass without single-key collapse; and a supercritical regime in which all mass concentrates on the closest key. Of notable interest is the subcritical case with identity value matrix where the attention map approximately implements a backward heat equation.
Scaling Linear Mode Connectivity and Merging to Billion Parameter Pretrained Transformers
Tianyi Li ⋅ Zhiqiang Shen
Linear mode connectivity (LMC) provides a promising foundation for understanding and merging independently trained neural networks, but existing methods typically optimize the interpolation path from only one model endpoint, limiting their scalability and effectiveness for large pretrained transformers. We propose a novel and scalable framework for enabling LMC-based model merging to {\em billion-parameter pretrained transformers}. Our method applies properly parameterized functionality-preserving weight transformations to align functionally equivalent solutions, and introduces a dual learning procedure in which both models jointly learn their corresponding transformations toward a shared linear interpolation path. This bidirectional optimization substantially reduces interpolation barriers and enables more reliable merging across large-scale architectures. Empirically, we show that our approach achieves near-zero loss barriers on WikiText for language models with medium-sized parameters, representing, to our knowledge, the first demonstration of near-barrier-free linear connectivity at this scale. In the vision domain, ViT-L maintains above 69\% ImageNet top-1 accuracy throughout the interpolation path, while modern billion-parameter LLMs exhibit only small loss barriers. These results suggest that properly resolving parameter symmetries enables large pretrained Transformers to be connected and merged through simple linear paths with substantially improved interpolation performance.
Scaling Storm-Resolving Atmospheric AI Simulation to the Entire Planet
Zeyuan Hu ⋅ Noah Brenowitz ⋅ Akshay Subramaniam ⋅ Jaideep Pathak ⋅ Tao Ge ⋅ Mohammad S Abbas ⋅ Suman Ravuri ⋅ Karthik Kashinath ⋅ Noel Keen ⋅ Naser Mahfouz ⋅ Peter Caldwell ⋅ Mike Pritchard
Kilometer-scale convection shapes precipitation extremes, tropical organization, and cloud feedbacks, but most global atmospheric models approximate these processes at 25--100\,km resolution. Global storm-resolving physics models resolve convective systems explicitly, but at a cost---roughly one MWh per simulated day on exascale supercomputers---that limits their use for long-duration atmospheric simulation. We introduce STRATA (Storm-resolving Tile-based autoRegressive Atmosphere Transformer Architecture), the first autoregressive AI emulator for global storm-resolving atmospheric dynamics. STRATA is trained on the highest-resolution atmospheric dataset yet used for global AI emulation: 17 days of output from the SCREAM physics model at 4.9-km resolution (${\sim}25$ million grid cells) sampled every 10 minutes. Our central premise is that since on 10-minute timescales atmospheric dynamics are predominantly local, training on small spatial tiles trades scarce global temporal samples for abundant local spatial samples and enables global rollout via overlapping-tile blending. STRATA combines 3D patch embedding and local 3D neighborhood attention for tractable modeling of high-resolution atmospheric tiles, a novel Stereographic Rotary Position Embedding (StereoRoPE) for grid-invariant positional encoding, and a pixel-space de-aliasing decoder that suppresses patch-scale rollout artifacts. An iso-FLOP scaling study reveals that km-scale emulation requires ${\sim}10\times$ more FLOPs per horizontal grid point than coarse-resolution AI weather models, consistent with the higher information density of convective-scale dynamics. Despite training on only 17 days of SCREAM output, STRATA produces stable 24-hour global rollouts with realistic km-scale dynamics across diverse weather regimes, though large-scale biases develop with lead time. STRATA achieves 48 simulation days per megawatt-hour---about 50 times better energy efficiency than the underlying SCREAM physics model---and 741 simulated days per wall-clock day at 512 H100 GPUs. Code and dataset will be publicly released.
Scattered by Design: Why Per-Output Pruning Resists Structured Compression
Victor Omolaoye ⋅ Gerard de Melo
Unstructured pruning methods such as Wanda preserve model quality but produce irregular sparsity with no practical compression, while structured methods enable efficient inference at the cost of degraded performance. We ask whether post-hoc linear transformations can convert unstructured sparsity into structured patterns, combining both benefits. Across a broad class of fixed linear transformations—including rotations, decompositions, and permutations—we find no increase in structured sparsity (Δ = 0.0%) in any OPT feed-forward layer. Transformations that yield structure do so identically on dense matrices, independent of pruning. We attribute this to three factors: rank preservation, weak inter-row mask correlation, and spectral dispersion. Together, these prevent post-hoc transformations from inducing structure, suggesting that compressible sparsity must be introduced during pruning rather than recovered afterward.
SCG-HF: Semantic Consistency Grouping for Hierarchical Fusion in Video Emotion Recognition
Muci Li ⋅ Yujing Rao ⋅ He Li ⋅ Mang Ye
Video Emotion Recognition (VER) aims to identify and understand the emotional states of characters by analyzing visual and auditory information in videos. However, conventional rigid frame sampling strategies for long videos tend to fragment continuous emotional expressions. Moreover, existing fusion methods often fail to properly model temporal dependencies and may suffer from future information leakage. To address these limitations, we propose SCG-HF, a novel framework that preserves semantic consistency through adaptive grouping and hierarchical fusion. Specifically, we introduce Semantic Consistency Grouping (SCG), which dynamically segments videos into semantically coherent units based on feature similarity, thereby maintaining the integrity of emotional patterns. Furthermore, we design a three-level Hierarchical Fusion (HF) architecture to capture emotional dynamics at different temporal granularities: (i) frame-level refinement enhances subtle local emotional cues; (ii) segment-level history-guided fusion models temporal evolution across segments while strictly preventing future information leakage via attention; and (iii) sample-level contrastive alignment synchronizes audio and visual representations in a shared latent space. Extensive experiments on the VideoEmotion-8 and Ekman-6 benchmarks demonstrate that SCG-HF achieves state-of-the-art performance.
SCHOLARPEER: A Multi-Agent Framework for Automated Peer Review
Palash Goyal ⋅ Mihir Parmar ⋅ Yiwen Song ⋅ Hamid Palangi ⋅ Tomas Pfister ⋅ Jinsung Yoon
The exponential growth of machine learning submissions has strained the traditional peer review process, resulting in slow feedback loops for authors and an immense burden on reviewers to rigorously audit technical soundness and verify literature. To address this, we introduce ScholarPeer, a multi-agent framework designed to operationalize the rigorous auditing workflow of a senior researcher. Rather than attempting to replace human judgment, ScholarPeer serves as a co-scientist: acting as a mentor for rapid author iteration prior to submission, and as an active verification assistant that augments human reviewers. The framework structurally decouples contextualization from critique by deploying a sub-domain historian to synthesize the field's trajectory, a baseline scout to proactively hunt for omitted state-of-the-art comparisons, and a multi-aspect Q&A engine that deeply audits technical soundness—scrutinizing internal logical consistency, experimental validity, and mathematical rigor—while cross-referencing claims against top-tier academic venues. We comprehensively evaluate ScholarPeer on $\sim$1,800 ICLR submissions spanning 2020 through 2025. Our results show that ScholarPeer achieves significant win-rates against state-of-the-art fine-tuned models and search-augmented agentic baselines.
scMAF: Single-Cell Multi-Omics Clustering via Adaptive Modality Fusion
Jun Fu ⋅ Yiding Lu ⋅ Ruohong Yang ⋅ Xi Peng ⋅ Yunfan Li
Single-cell multi-omics sequencing has greatly advanced the characterization of cellular heterogeneity by jointly profiling multiple molecular modalities. However, since different modalities exhibit varying discriminative signals across cell types, existing integration strategies may dilute or even obscure the signals carried by more informative modalities when fused with less informative ones, thereby hindering accurate multi-omics clustering. To address this limitation, this work presents a novel adaptive fusion framework for unsupervised clustering of single-cell multi-omics data. The core of our approach is an attention-based fusion module that dynamically assigns modality-specific weights within a shared latent space, allowing the model to naturally focus on the most informative signals from each modality. Extensive experiments on eight real-world multi-omics datasets encompassing gene expression, chromatin accessibility, and protein modalities demonstrate that our method achieves the state-of-the-art single-cell clustering performance over fourteen competitive baselines. The code will be released publicly upon acceptance.
ScrapeBench: Evaluating Legal Compliance of AI Agents in Website Scraping
Joseph Marvin Imperial ⋅ Daniel Slate ⋅ Noam Kolt
AI agents deployed on the open web are becoming more capable of performing actions that may breach contracts and violate computer security laws. In this work, we introduce $\text{ScrapeBench}$, the first agentic benchmark for evaluating the legal compliance of frontier AI agents with website scraping policies, including terms of service (ToS) and $\texttt{robots.txt}$ directives. Across widely used closed-weight and open-weight AI agents, we find that most agents exhibit low compliance rates (< 5\%) when instructed to scrape websites in our sample of 1,626 popular websites that explicitly prohibit scraping. In addition, we observe that even where AI agents actively check the applicable ToS and $\texttt{robots.txt}$, they continue to exhibit low compliance rates. Conversely, in our sample of 269 websites that do not prohibit scraping, we observe that agents regularly refuse to scrape despite it being lawful, such as Claude Opus 4.6 exhibiting an over-refusal rate of 98.8%. Lastly, our human baseline experiments ($n$ = 180 participants) suggest that humans exhibit compliance rates comparable to AI agents but over-refuse at much higher rates. We release our anonymized code and data in this repository: https://anonymous.4open.science/r/scrapebench-75D6
Search-Tree Scaling in Parallel Monte Carlo Tree Search
Scott Cheng ⋅ Meng-Yu Tsai ⋅ Ding-Yong Hong ⋅ Mahmut T Kandemir
The exploration-exploitation tradeoff is a fundamental challenge in reinforcement learning. In Monte Carlo Tree Search (MCTS), this tradeoff is balanced explicitly by the pUCT exploration coefficient and implicitly by parallelism when the search is combined with neural network evaluations. However, the exploration coefficient is typically tuned as an algorithmic hyperparameter, while the degree of parallelism is often treated as a system detail. To this end, we characterize how these two factors shape the search-tree structure and jointly affect performance scaling. We prove that, under regularity conditions, the depth of the optimal path grows as $\Theta(\sqrt{n}/\log n)$ for $n$ simulations under the MuZero exploration coefficient formula. For tree-parallel MCTS, we prove that the search depth preserves the same asymptotic order when $T=o(\sqrt{n})$, while sufficiently large parallelism introduces contention and reduces the depth growth rate to $\Theta(n/T)$ for $T$ threads. Our empirical findings indicate that each doubling of the number of threads decreases the best-performing exploration coefficient by approximately 0.19. Overall, our characterization connects the search-tree structure with the exploration-exploitation behavior in pUCT, providing insight into how exploration coefficients and parallelism can be jointly tuned.
Seed-Your-Motion: Householder Orthogonal Noise for Motion-Controllable Video Diffusion Models
Tao Jun Lin ⋅ Jinhong Ni ⋅ Ruyi Zha ⋅ Yifu Wang ⋅ Xibin Song ⋅ Pan Ji ⋅ Hongdong Li
Diffusion-based video generation has achieved impressive visual fidelity, yet enforcing temporally consistent motion under explicit control signals remains challenging. Recent noise-warping approaches construct temporally correlated noise by transporting the initial latent noise via diffeomorphic deformation fields derived from optical flow, but this spatial resampling inherently destroys the i.i.d. Gaussian structure of the noise, requiring post-hoc correction steps at the cost of non-trivial computational overhead. We propose \textbf{Seed-Your-Motion}, by conditioning the a priori latent noise on optical flow via per-location orthogonal transformations in latent channel space, parameterized as products of the Householder Transformation. This channel-space formulation has an advantage of preserving the spatial Gaussianity exactly by construction, requiring neither Jacobian corrections nor post-hoc distribution restoration, while at the same time inducing structured cross-frame correlations that ensures geometric consistency across time. With no learnable parameters and a single forward pass, our method is agnostic to model architecture and can be plugged seamlessly into any video diffusion fine-tuning pipeline. Extensive experiments on motion-transfer and camera-controlled video generation benchmarks demonstrate that our method matches or outperform the state-of-the-art noise-warping baseline, while achieving a 2.7× speedup in processing time. The results suggest that the channel-space Householder transformations provide a principled and efficient alternative to noise-warping based motion-controllable video diffusions.
Seeing Across Skies and Streets: Feedforward 3D Reconstruction from Satellite, Drone, and Ground Images
Qiwei Wang ⋅ Zhongyao Tuo ⋅ Xianghui Ze ⋅ Yujiao Shi
Cross-view localization classically asks: where does this ground image lie on the satellite tile? Existing methods are typically limited to 3-DoF estimates---an $(x,y)$ position and a yaw angle---because nadir satellite imagery provides no direct cues for roll, pitch, or altitude, forcing a reliance on planar-motion and zero-tilt assumptions. These assumptions break on real terrain with slopes, ramps, and tilted camera mounts. To overcome this, we introduce a single UAV image as an intermediate viewpoint: it reveals the 3D structure invisible from nadir, supplies the cues for roll, pitch, and altitude that the satellite alone cannot provide, and needs only spatial overlap with the ground camera---no known relative pose is required. Building on this insight, we propose **Cross3R**, a flexible feed-forward model that ingests a satellite tile together with a UAV image, a ground image, or both, and, in a single forward pass, recovers a cross-view 3D point cloud, the 6-DoF poses of every input camera, and the on-tile $(x,y)$ position and yaw of each perspective camera. For training and evaluation, we also construct **CrossGeo**, a 278K-image tri-view dataset spanning 85 scenes across every continent except Antarctica. On CrossGeo, Cross3R consistently outperforms feed-forward 3D baselines in point-cloud reconstruction, 6-DoF camera-pose estimation, and cross-view localization. On KITTI, it outperforms dedicated cross-view methods trained on KITTI on most metrics, despite having no KITTI training itself.
Seeing Is Not Screening: Multimodal Hidden Instruction Attacks on Agent Skill Scanners
Xiaojun Jia ⋅ Jie Liao ⋅ Simeng Qin ⋅ Ke Ma ⋅ Wenbo Guo ⋅ Yebo Feng ⋅ Aishan Liu ⋅ Yang Liu
Agent skills are emerging as an important attack surface in LLM-based systems. Through an empirical study of existing skill scanners, we find that current defenses primarily rely on textual descriptions, manifests, and source code as the main signals for security analysis, which can leave visually conveyed malicious intent insufficiently examined. This creates a practical blind spot: harmful operational instructions hidden in images may bypass scanning while still being recoverable by multimodal agents during deployment. To systematically investigate this threat, we propose \textsc{SkillCamo}, a document-mediated multimodal instruction attack that conceals malicious instructions within images bundled with a skill while rewriting the surrounding documentation to naturally reference those images as part of the normal workflow. Thus, the attack does not rely on the image alone, but on the joint interpretation of textual guidance and visual payload at execution time. To defend against such attacks, we further propose \textsc{ExecScan}, an execution-grounded multimodal scanning module that performs intent extraction, behavior reconstruction, abuse assessment, and deliberative execution simulation over skill artifacts. \textsc{ExecScan} jointly analyzes documentation, code, referenced resources, and visual content to recover hidden instructions, reconstruct executable behavior chains, and identify downstream risks such as exfiltration, destruction, persistence, deception, and privilege escalation. Extensive experiments show that image-hidden malicious instructions challenge existing skill scanners, while \textsc{ExecScan} can improve the skill scanning performance.
Today, verifier-guided best-of-$N$ inference is a common way to spend test-time compute in language modeling, code generation, and alignment: generate $N$ candidates, score them with a verifier or reward model, and deploy the highest-scoring output. This paper argues that the relevant reliability object is not average verifier accuracy but \emph{selected-tail reliability}, namely the behavior of the output chosen after optimizing over $N$ noisy public scores; we derive exact selected-rank exposure laws for best-of-$N$ selection and finite-verifier resource laws showing that selected public optimism scales at the $\sqrt{\log N/m}$ regime under $m$ units of verifier evidence. We then separate public optimism from hidden utility harm, showing that public-score inflation alone is not a harm claim and that harm or uncertifiability requires public-hidden mismatch, tail dominance, count imbalance, or an unresolved selected-tail regime. Building on these results, we introduce Tail-Certified Compute Caps, a conservative procedure that certifies candidate budgets or refuses certification using selected-tail intervals, verifier-resource penalties, mismatch envelopes, and feasibility gates. In exact-verifier/public-hidden experiments, the theory predicts selected false-positive exposure and supports resource-aware cap decisions over an all-attempted denominator, while non-code candidate tables with pre-existing public human annotations serve as secondary scoped $K\ge16$ consistency checks under proxy/model-prior/verifier-style scores across Arena55K and Stanford SHP. Overall, the paper provides a reliability framework for verifier-guided inference-time scaling: candidate budgets should be justified by evidence about the selected tail that is actually deployed, not only by average verifier validation.
Selective state-space models such as Mamba replace fixed linear state transitions by input-dependent transitions, yielding recurrent sequence models that retain linear-time inference while adapting their memory dynamics to the current token. We show that even a simple form of this selective memory update has universal approximation power. Specifically, for every compact input domain, finite horizon, and continuous causal sequence-to-sequence map, a clocked diagonal selective state-space scan with strictly contractive gates and a linear-in-state readout uniformly approximates the target map. The proof is constructive: input-dependent diagonal propagation generates tensor-product partition features over prefixes, and the linear-in-state readout combines these features with sampled target values. Thus the high-order interactions required for universality are produced inside the selective scan itself, rather than by attention, nonlinear hidden-state dynamics, or a universal multilayer perceptron applied after a sequence encoder. These results identify selectivity as a mathematically sufficient mechanism for expressive linear-time state-space sequence modeling. Selectivity lets the model decide, token by token, which parts of the past should be preserved, faded, or combined with the present. The universal approximation theorem shows that this adaptive use of memory is sufficient to reproduce any well-behaved causal sequence rule over a fixed horizon.
Self-Adjoint Flow Policy Optimization
Guojian Zhan ⋅ Feihong Zhang ⋅ Likun Wang ⋅ Xiangteng Zhang ⋅ Letian Tao ⋅ Tianyi Zhang ⋅ Wenxin Zhao ⋅ yinuo Wang ⋅ Tianze Zhu ⋅ Jingliang Duan ⋅ Yang Guan ⋅ Shengbo Eben Li
Generative policies offer a promising alternative to Gaussian policies in reinforcement learning (RL) due to their ability to represent expressive and multimodal action distributions. However, integrating such policies into likelihood-based policy optimization remains challenging, since accurate and differential action likelihoods are intractable for standard generative policies. In this paper, we propose \emph{Self-Adjoint Flow Policy Optimization} (\textsc{SAFlow}), a likelihood-tractable generative policy with self-adjoint structure. \textsc{SAFlow} parameterizes the policy as an invertible flow in a doubled action space and uses a Verlet-style self-adjoint composition for generation. This structure makes the inverse map obtainable by stepsize reversal, aligning forward sampling and backward likelihood evaluation under the same time-symmetric numerical rule. It also provides a second-order approximation accuracy to the learned continuous flow, improving the numerical consistency of likelihood and policy-ratio computation. We instantiate \textsc{SAFlow} within an on-policy optimization framework and evaluate it on eight IsaacLab continuous-control tasks. \textsc{SAFlow} consistently outperforms Gaussian-policy baselines and recent generative-policy methods, highlighting the value of self-adjoint numerical structure for likelihood-based generative policy optimization.
SelfCritic-VLA: Language as Intrinsic Critic for Vision Language Action Models in Autonomous Driving
Fang Li ⋅ Shaoqing Xu ⋅ Yuechen Luo ⋅ Hanbing Li ⋅ Zhi-Xin Yang
Vision-Language-Action (VLA) models offer a promising paradigm for end-to-end autonomous driving by connecting visual scene understanding, linguistic reasoning, and low-level trajectory generation. However, existing driving VLAs typically use language only as static commands or supervised rationales, while treating trajectory prediction as an open-loop generation problem. As a result, they lack an explicit mechanism for understanding the consequences of their own actions, which limits robustness under closed-loop distribution shift and makes reinforcement learning inefficient in long-tail scenarios. We propose SelfCritic-VLA, a self-evaluative VLA framework that unifies trajectory generation, language-grounded action critique, and trajectory refinement. Given a driving scene, the model first generates an initial trajectory, then produces a language-grounded critique that describes the potential safety, progress, and rule-compliance consequences of executing that trajectory, and finally refines the trajectory conditioned on this critique. To train this capability, we construct a large-scale counterfactual trajectory dataset with multi-source candidate trajectories and semantic critiques, and perform hybrid supervised fine-tuning for planning, action evaluation, and refinement. Building on this initialization, we introduce a self-critic reinforcement learning procedure based on GRPO, where critique-conditioned refinements are filtered by grounded driving rewards and distilled back into the base policy. Experiments on NAVSIM and Bench2Drive show that SelfCritic-VLA consistently improves both open-loop and closed-loop driving performance over strong baselines, demonstrating the effectiveness of language-guided self-refinement for driving policy optimization.
Self-EvolveRec: Self-Evolving Recommender Systems with LLM-based Directional Feedback
Sein Kim ⋅ Sangwu Park ⋅ HongSeok Kang ⋅ Wonjoong Kim ⋅ Jimin Seo ⋅ Yeonjun In ⋅ Kanghoon Yoon ⋅ Hyunsik Jeon ⋅ Chanyoung Park
Traditional methods for automating recommender system design, such as Neural Architecture Search (NAS), are often constrained by a fixed search space defined by human priors, limiting innovation to pre-defined operators. While recent LLM-driven code evolution frameworks shift fixed search space target to open-ended program spaces, they primarily rely on scalar metrics (e.g., NDCG, Hit Ratio) that fail to provide qualitative insights into model failures or directional guidance for improvement. To address this, we propose Self-EvolveRec, a novel framework that establishes a directional feedback loop by integrating a User Simulator for qualitative critiques and a Model Diagnosis Tool for quantitative internal verification. Furthermore, we introduce a Diagnosis Tool - Model Co-Evolution strategy to ensure that evaluation criteria dynamically adapt as the recommendation architecture evolves. Extensive experiments demonstrate that Self-EvolveRec significantly outperforms state-of-the-art NAS and LLM-driven code evolution baselines in both recommendation performance and user satisfaction. Our code is available at https://anonymous.4open.science/r/selfevolverec-37F5/.
At the heart of existing language model agents is a fixed orchestrator program which is responsible for the state transition between consecutive turns. This paper introduces \emph{self-programmed execution} (SPE), an agent architecture in which the model completion is itself the orchestrator program, and the harness evaluates this program but does not impose its own orchestration policy. I formalize this idea using agentic machines: an SPE state is one from which a model completion can load any state of an embedded copy of the machine, meaning that it is subject to no fixed turn-to-turn orchestration policy. Realizing SPE in practice is nontrivial because the same data is both model context and executable program. I therefore introduce \textsc{Spell}, a Lisp-based language in which programs can edit and re-evaluate themselves, and effectful expressions like model invocations are structured such that re-evaluating an edited program does not replay its side effects. Experiments with existing models, not trained for SPE or \textsc{Spell}, show that frontier models can operate in this regime and accomplish challenging agentic tasks. These results demonstrate how an LM can act as an agent without any fixed orchestration policy, and they raise the question of what self-orchestration strategies might be learned by a model trained for self-programmed execution.
Semantic Concept Steering Breaks the Explanation Drift Loop in Continual Learning
Yehonatan Elisha ⋅ Oren Barkan ⋅ Noam Koenigstein
Catastrophic forgetting in continual learning (CL) manifests not only as accuracy degradation but also as explanation drift: models maintain predictive accuracy while silently shifting attention from diagnostic features to spurious correlations. This work identifies and resolves a previously overlooked structural flaw in explanation-aware CL: pixel-level saliency regularization derives supervision targets from the model's own frozen saliency maps, creating a \emph{self-referential drift loop} that actively reinforces spurious attention across tasks. We introduce C$^3$L (Concept-Consistent Continual Learning), which breaks this loop by anchoring regularization to class-consistent semantic concepts (e.g., ``curved beak'', ``striped pattern'') discovered offline, independent of model state, providing (i) target independence from model drift, and (ii) cross-instance semantic consistency across every instance of a class. We show that our concept-based objective, combined with approximate LRP relevance conservation, produces a crowding-out effect: enforcing high relevance within concept regions implicitly suppresses spurious background correlations without explicit penalization, forming a zero-sum competition over a finite relevance budget. \method doesn't require human-annotated concept sets, as it employs a fully automated concept discovery and spatial grounding. Extensive evaluation across five benchmarks and over 20 baselines demonstrates state-of-the-art performance in accuracy, forgetting, explanation quality, spurious correlation mitigation, and out-of-distribution robustness. Our code is provided in the supplementary material.
Semantic-Statistical Prior Banks for Federated Low-Shot Learning under Non-IID Clients
Jiaying Wu ⋅ Can Gao ⋅ Lei Shi ⋅ Jia Luo ⋅ Dongmei Wei ⋅ Feifei Yang ⋅ Pengfei Zhang ⋅ Feifei Kou ⋅ Zhaoman Zhong
Federated intrusion detection must adapt to rare attacks from only a few labeled flows per client, while client distributions are highly non-IID and raw traffic cannot be centralized. Existing few-shot intrusion detectors typically construct class priors in centralized settings, and standard federated training does not provide an explicit mechanism for reusing cross-client class statistics during local low-shot adaptation. We propose a semantic-statistical prior bank for federated low-shot intrusion detection. Clients first participate in federated representation learning and submit protected class-wise moment statistics through secure aggregation, from which the server constructs a global prior bank without accessing raw traffic. During local episodic adaptation, label-phrase embeddings retrieve relevant priors, which are combined with scarce support statistics through a maximum a posteriori (MAP) shrinkage estimator in the feature space. The adapted class distributions are then used to train lightweight episode-specific classifiers on each client. Experiments on CICIDS2017 and CICDDoS2019 under 5-way 1-shot and 5-shot non-IID settings show consistent improvements over FedAvg and backbone-matched baselines, with the largest gains appearing in 1-shot and highly skewed client distributions. With ResNet1D, the proposed adaptation improves 1-shot F1 from 92.97 to 95.55 on CICIDS2017 and from 82.13 to 88.25 on CICDDoS2019.
Semantics-to-Contact: A Stagewise Framework for Robust Contact-Rich Manipulation
Guanren Qiao ⋅ Ruixiang Ouyang ⋅ Sheng Xu ⋅ Yueci Deng ⋅ Yunxin Tai ⋅ Kui Jia ⋅ Guiliang Liu
Contact-rich manipulation requires precise control under complex contact dynamics while remaining robust to diverse visual conditions. However, real-world reinforcement learning often overfits to the visual appearance of training scenes. We propose Semantics-to-Contact (S2C), a stagewise framework for visually robust contact-rich manipulation. The insight is to decompose the problem into two stages: semantic focusing and contact refinement. In the first stage, a vision-language model localizes task-relevant regions across visually diverse scenes, providing coarse spatial guidance and reducing the burden of exploration. In the second stage, a real-world human-in-the-loop residual reinforcement learning policy learns fine-grained contact behaviors within the localized region using force feedback. To further improve visual robustness, we introduce an object-centric visual augmentation strategy that randomizes background appearance while preserving the robot and task-relevant objects. Experiments on six real-world contact-rich manipulation tasks exhibit that S2C consistently improves success rates over strong baselines and maintains reliable performance under significant visual and positional variations. These results demonstrate that stagewise semantic focusing and contact refinement provide a practical path toward visually robust real-world contact-rich manipulation. Video materials can be seen in \url{https://anonymous.4open.science/api/repo/S2C-demo-4E03/file/index.html}.
Senses Wide Shut: A Representation-Action Gap in Omnimodal LLMs
Trung Nguyen ⋅ Yiming Gao ⋅ Fanyi Pu ⋅ Kaichen Zhang ⋅ Shuo Sun ⋅ Ziwei Liu
When an omnimodal large language model accepts a question whose textual premise contradicts what it actually sees or hears, does the failure lie in perception or in action? Recent omnimodal models are positioned as perception-grounded agents that jointly process video, audio, and text, yet a basic form of grounding remains untested: catching a textual claim that conflicts with the model's own sensory input. We introduce IMAVB, a curated 500-clip benchmark of long-form movies with a 2x2 design crossing target modality (vision, audio) and premise condition (standard, misleading), which lets us measure conflict detection separately from ordinary multimodal comprehension. Across eight open-source omnimodal LLMs and Gemini 3.1 Pro, we document a Representation-Action Gap: hidden states reliably encode premise-perception mismatches even when the same models almost never reject the false claim in their outputs. Behaviorally, models fall into two failure modes: under-rejection, in which they answer misleading questions as if the false premise were true; and over-rejection, in which they reject more often but also reject standard questions, sacrificing ordinary comprehension accuracy. The gap is modality-asymmetric (audio grounding underperforms vision) and prompt-resistant across seven variants. As an initial diagnostic intervention, a probe-guided logit adjustment (PGLA) re-injects the encoded mismatch signal into decoding and consistently improves rejection behavior. Together, these results suggest the bottleneck for omnimodal grounding lies in translation, not perception.
SeqDiCO: Sequence-oriented Diffusion for Scale-Generalizable Neural Combinatorial Optimization
Yu Wang ⋅ Qiaolin Lu ⋅ Yang Wu ⋅ Bohao Qu ⋅ Shirui Pan ⋅ Yi Chang ⋅ Chengqi Zhang
Diffusion-based neural combinatorial optimization (NCO) has recently shown promise for routing problems by generating solutions through denoising or reconstruction. Most existing diffusion solvers, however, define the generative process over full-instance solution structures, such as adjacency matrices or heatmaps. This instance-level formulation ties the learned reconstruction distribution to the problem size observed during training, making it difficult to transfer models trained on small instances to substantially larger ones. In this paper, we propose SeqDiCO, a sequence-oriented diffusion framework for scale-generalizable neural combinatorial optimization. The key idea is to move diffusion from full-instance solutions to fixed-endpoint sequence subproblems. Given two endpoints and a set of intermediate nodes, SeqDiCO learns a local denoising operator that reconstructs the open path connecting the endpoints from corrupted structural observations. Since such local sequence patterns are reusable across instance sizes, the same operator trained on small instances can be repeatedly composed to solve larger routing problems. During inference, SeqDiCO uses this operator for both construction, by growing a solution through receding-horizon local reconstruction, and refinement, by replacing selected subpaths in an existing solution. Experiments on TSP and CVRP show that SeqDiCO consistently outperforms representative diffusion-based and autoregressive neural solvers under cross-scale evaluation. Trained only on 100-node instances, SeqDiCO generalizes to 10K-node problems, reducing the optimality gap by up to 51\% over SOTA neural solvers, while achieving up to $10\times$ faster inference on representative large-scale benchmarks.
Sequence-to-Sequence Modeling with Camera-Induced Priors for Multi-View Stereo
Aoxiang Fan ⋅ Corentin Dumery ⋅ Nicolas Talabot ⋅ Pascal Fua
Computing accurate geometry from multi-view images is a fundamental problem in computer vision. In this work, we address the Multi-View Stereo (MVS) setting, where camera parameters are assumed to be known. Most existing learning-based methods rely on per-view cost volumes as a camera-induced prior, effectively casting the problem as a sequence-to-one mapping that predicts depth only for a designated reference view. We instead reformulate MVS as a sequence-to-sequence task, enabling the simultaneous prediction of depth maps and point maps for all input views. To this end, we propose a global transformer-based architecture with two key components that explicitly utilize camera-induced priors. First, we employ ray-map embeddings to inject camera parameters into image patch tokens, making the transformer architecture camera-aware for geometry prediction. Second, we replace conventional per-view cost volumes with a unified global cost-volume representation that jointly captures 3D structure across all views. Extensive experiments on multiple public benchmarks demonstrate that our approach achieves state-of-the-art performance, surpassing both multi-view stereo and feed-forward reconstruction baselines.
Modern AI models are not static. They go through multiple updates in their lifecycles. We propose to design Sequential Membership Inference (SeMI) attacks leading to tighter privacy audits by exploiting the sequence of models and injecting a target canary at a controlled insertion time. First, for empirical mean computation, we develop $\mathrm{SeMI}^{\ast}$, an optimal SeMI attack to identify the presence of a target inserted at a specific insertion step. We derive the power of $\mathrm{SeMI}^{\ast}$ to show that accessing the model sequence yields more powerful MI attacks than scrutinising only the final model. $\mathrm{SeMI}^{\ast}$ exhibits an isolation property-- its power depends on the statistics obtained right before and after insertion of the target. Leveraging this insight, we develop practical white-box (accessing model gradients) and black-box (accessing loss) SeMI attacks against models trained with (DP-)SGD. Across datasets and models trained with (DP-)SGD, our experiments show that SeMI attacks achieve higher powers than snapshot-independent baselines, and yield tighter privacy audits thanks to (a) control over the insertion time and (b) observations across the model sequence.
SFPR: Structural Fingerprinting for LiDAR-to-OpenStreetMap Place Recognition
Jiajun Zheng ⋅ Ziwei Shi ⋅ Yu Zang ⋅ Chenlu Lin ⋅ Youbiao Wang ⋅ Wenjing Xu ⋅ Cheng Wang
Cognitive neuroscience research indicates that human spatial navigation relies not merely on abstract global map representations, but rather on anchoring local observations to spatially stable, decision-relevant landmarks to effectively align the current scene with the cognitive map. Inspired by this mechanism, we propose SFPR, a structural fingerprint-based LiDAR-to-OpenStreetMap (OSM) place recognition framework. This framework goes beyond the traditional coarse retrieval paradigm that relies exclusively on global descriptor similarity and further introduces fine-grained discrimination based on structural fingerprints to achieve more accurate place recognition. Specifically, we propose a structural fingerprint-aware attention mechanism that utilizes sparse but spatially stable anchors as saliency prompts, directing the network's attention toward potential fingerprint regions while dynamically generating global descriptors. Subsequently, we formulate the extraction and matching of structural fingerprints as an alternating optimization problem guided by this fingerprint-aware attention. Through iterative optimization, we extract structural fingerprints with high fingerprint-aware attention and geometric consistency from anchors. These structural fingerprints are further fed back into the training stage as geometric priors, effectively suppressing false matches characterized by similar global features but contradictory local structures. Experiments demonstrate that SFPR effectively overcomes the retrieval bottleneck in macroscopically homogeneous scenes, achieving a relative improvement of 30.20\% in Top-1 Recall@1m compared to the current state-of-the-art method. Code are publicly available at https://anonymous.4open.science/r/SFPR.
SGNNBench: A Holistic Evaluation of Spiking Graph Neural Networks on Large-scale Graphs
Huizhe Zhang ⋅ Jintang Li ⋅ Yuchang Zhu ⋅ Huazhen Zhong ⋅ Liang Chen
Graph Neural Networks (GNNs), as representative deep models designed for graph data, effectively learn graph topological information and push the performance boundaries across various graph tasks. However, the computational and memory burden of GNNs poses significant challenges for scaling to large real-world graphs. Spiking Graph Neural Networks (SGNNs), which integrate biologically plausible learning via unique spike-based neurons, have emerged as a promising energy-efficient alternative. Different layers communicate with sparse and binary spikes, which facilitates computation and storage of intermediate graph representations. Despite the proliferation of SGNNs proposed in recent years, there is no systematic benchmark to explore the basic design principles of these brain-inspired networks on the graph data. To bridge this gap, we present SGNNBench to quantify progress in the field of SGNNs. Specifically, we elaborately investigate the design space of SGNNs to facilitate the development of a general SGNN paradigm. Regarding efficiency, we empirically compare these baselines w.r.t. model size, memory usage and theoretical energy consumption to reveal previously overlooked energy bottlenecks. Furthermore, SGNNBench establishes a unified and reliable benchmark to comprehensively evaluate 9 state-of-the-art SGNNs across 20 datasets covering diverse graph settings and tasks.
Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory
Juno Kim ⋅ Eshaan Nichani ⋅ Denny Wu ⋅ Alberto Bietti ⋅ Jason Lee
Spectral optimizers such as Muon have recently shown strong empirical performance in large-scale language model training, but the source and extent of their advantage remain poorly understood. We study this question through the linear associative memory problem, a tractable model for factual recall in transformer-based models. In particular, we go beyond orthogonal embeddings and consider Gaussian inputs and outputs, which allows the number of stored associations to greatly exceed the embedding dimension. Our main result sharply characterizes the recovery rates of one step of Muon, SGD, and Newton's method on the logistic regression loss under a power law frequency distribution. We show that the storage capacity of Muon significantly exceeds that of SGD, and even matches Newton's method while only using first-order information. Moreover, Muon saturates at a larger critical batch size. We further analyze the multi-step dynamics under a thresholded gradient approximation and show that Muon achieves a substantially faster initial recovery rate than SGD, while both methods eventually converge to the information-theoretic limit at comparable speeds. Experiments on synthetic tasks validate the predicted scaling laws. Our analysis provides a quantitative understanding of the signal amplification of spectral preconditioners and lays the groundwork for establishing scaling laws across more practical language modeling tasks and optimizers.
Short-Context Dominance: How Much Local Context Natural Language Actually Needs?
Vala Vakilian ⋅ Zimeng Wang ⋅ Ankit Rawat ⋅ Christos Thrampoulidis
We investigate the short-context dominance hypothesis: that for most sequences, a small local prefix suffices to predict their next tokens. Using large language models as statistical oracles, we measure the minimum context length (MCL) needed to reproduce accurate full-context predictions across datasets with sequences of varying lengths. For sequences with 1–7k tokens from long-context documents, we consistently find that 75–80\% require only the last 96 tokens at most. Given the dominance of short-context tokens, we then ask whether it is possible to detect challenging long-context sequences for which a short local prefix does not suffice for prediction. We introduce a practical proxy to MCL, called Distributionally Aware MCL (DaMCL), that does not require knowledge of the actual next-token and is compatible with sampling strategies beyond greedy decoding. Our experiments validate that simple thresholding of the metric defining DaMCL achieves high performance in detecting long vs. short context sequences. Finally, to counter the bias that short-context dominance induces in LLM output distributions, we develop an intuitive decoding algorithm that leverages our detector to identify and boost tokens that are long-range-relevant. Across Q&A tasks and model architectures, we confirm that mitigating the bias improves performance.
Simplicity is Enough: ReAct Agents for Prompt Optimization
Andrei Rusu ⋅ Andrei Dumitrescu ⋅ Adrian C Badea ⋅ Cosmin Maria ⋅ Ion-Marian Anghelina ⋅ Bai Li ⋅ Tomasz L Religa ⋅ Dragos H Bobolea
Prompt optimization has recently been approached with increasingly specialized search procedures and evolutionary frameworks. In this work, we ask whether such complexity is necessary to achieve competitive performance. We introduce REAPO (ReAct Agents for Prompt Optimization), a lightweight framework that casts prompt optimization as an agentic process built around a ReAct loop augmented with evaluation and reflection tools. Rather than relying on bespoke optimization machinery, REAPO iteratively refines prompts through error analysis, reflective reasoning, and validation against the optimization objective. We evaluate REAPO on several established benchmarks spanning multi-hop question answering, mathematical reasoning, fact verification, instruction following, and structured classification, as well as on real enterprise agentic tasks. Across these settings, we find that this agentic optimizer can match or outperform specialized prompt tuning methods while requiring 2–7$\times$ fewer evaluation rollouts and less auxiliary optimization logic. We also report confidence intervals over multiple independent runs and show that single-run evaluations on small benchmarks can give a misleading picture of relative performance, since LLM variance may obscure true gains. Taken together, our results suggest that effective prompt optimization need not depend on increasingly elaborate optimization schemes, and that agentic methods can provide a more efficient alternative.
SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization
Yan Sun ⋅ Guoxia Wang ⋅ Jinle Zeng ⋅ Jiabin Yang ⋅ Shuai Li ⋅ Li Shen ⋅ Dacheng Tao ⋅ Dianhai Yu ⋅ Haifeng Wang
Pretraining large language models (LLMs) with next-token prediction has led to remarkable advances, yet the context-dependent nature of token embeddings in such models results in high intra-class variance and inter-class similarity, thus hindering the efficiency of representation learning. While similarity-based regularization has demonstrated benefit in supervised fine-tuning and classification tasks, its application and efficacy in large-scale LLM pretraining remains underexplored. In this work, we propose the SimReg, an embedding similarity regularization loss that explicitly encourages token representations with the same ground-truth label within each sequence to be more similar, while enforcing separation from different-label tokens via a contrastive loss. Our analysis reveals that this mechanism introduces gains by enlarging multi-classification margins, thereby enabling more efficient classification. Extensive experiments across dense and Mixture-of-Experts (MoE) architectures demonstrate that SimReg consistently accelerates training convergence by over 30% and improves average zero-shot downstream performance by over 1% across standard benchmarks. Further ablation studies and analyses offer practical insights into hyperparameter tuning and loss effectiveness.
Simulating Human Memory with Language Models
Qihan Wang ⋅ Nicholas Tomlin ⋅ Michael Hu ⋅ Brian Dillon ⋅ Tal Linzen
Language models are increasingly being deployed as user simulators, but their memory is far more reliable than that of real users. To measure this gap, we run a series of classic memory experiments from psychology on both humans and language models. Across tasks, we find that out-of-the-box language models exhibit better memory than humans, even when prompted to imitate human behavior. We then show that better prompting strategies and the use of a compactor can cause language models to forget content in a more human-like way. Using these methods, we show preliminary evidence that language models with human-like memory constraints can function as more effective user simulators in a downstream education task. Finally, we release human reference data and benchmarks to support future work on simulating human memory with language models.
Skill-Adaptive Noise Scheduling for Diffusion Policies
Woo Kyung Kim ⋅ Gwangpyo Yoo ⋅ Eunyoung Park ⋅ Honguk Woo
Recent advances in diffusion policies have demonstrated strong effectiveness when integrated into hierarchical skill-based learning frameworks, leveraging offline data to capture multimodal and temporally abstracted behaviors. However, prior approaches have rarely considered the diffusion process as a learnable component, adopting fixed noise schedules applied uniformly across skills, regardless of their distinct behavioral contexts. Such a design fundamentally limits the achievable performance, particularly for contact-rich and dexterous manipulation tasks where different skills exhibit distinct couplings between action dimensions. In this paper, we present a skill-adaptive noise diffusion framework (SaND) which introduces a learnable, skill-conditioned noise schedule that modulates the diffusion process per action dimension, adapting the noising dynamics to each skill. To prevent the learned noise schedule from collapsing into trivial solutions, we employ an auxiliary inverse objective that reconstructs the skill embedding from the Legendre coefficients of the derivative of the noise schedule, encouraging the schedules to remain skill-discriminative. This formulation naturally enables adaptive control over the denoising steps for each skill, where noise schedules are clustered in the Legendre coefficient space, and the appropriate number of denoising steps is determined at the point where reconstruction error abruptly increases. At inference, we identify each skill embedding's cluster from its noise schedule and apply the corresponding number of denoising steps, dedicating more compute to skills requiring fine-grained control and less to simpler ones. Through experiments, we demonstrate that SaND outperforms existing diffusion-based skill learning and planning baselines on contact-rich and dexterous manipulation tasks, achieving higher success rates while requiring fewer denoising steps.
SKIM: Pruning Large Language Model Agents via Selective Knowledge Informed Masking
Moonseok Choi ⋅ Giung Nam ⋅ Jongwon Jeong ⋅ Minki Kang ⋅ Juho Lee
Large Language Model (LLM) agents are increasingly deployed under structured generation, where grammar constraints enforce output format at inference time. Whereas existing compression methods target unconstrained generation, we first identify a fundamental mismatch: under structured generation, grammar constraints reduce the model's output distribution to valid tokens, and the logits matter most at the small set of positions that decide the agent's action. Motivated by this finding, we propose SKIM (Selective Knowledge-Informed Masking), the first pruning and distillation framework for LLM agents under structured generation, with both stages guided by token role. SKIM partitions token positions into context, decision, and format categories that reflect the input the model reads, the constrained options it chooses among, and the format realized by the constraint mechanism. Specifically, this partition drives two complementary axes: (1) category-aware pruning saliency scores that preserves parameters most relevant for decision positions; and (2) an on-policy distillation objective that aligns student and teacher on valid tokens, paired with category-aware feature matching. We evaluate SKIM on tool-calling benchmarks and within multi-step agent loops, across multiple model scales and sparsity levels. SKIM consistently improves accuracy over baselines, with its superiority becoming most striking in high-sparsity regimes where conventional methods struggle, demonstrating the potential of category-aware compression for deployment-time grammar-constrained agents.
SLiDE: Structured Linear Dynamics for Forecasting with Exogenous Inputs
Sebastian Pütz ⋅ Theodore Glavas ⋅ Benjamin Schäfer ⋅ Boris Oreshkin ⋅ Mark Coates
Accurate time series forecasting with exogenous inputs is critical across domains, including energy and retail, yet modern deep learning models often overfit to correlations that do not generalize under shifts in these inputs. We propose SLiDE, a Koopman-inspired architecture for forecasting with exogenous inputs that imposes structure on latent temporal dynamics. SLiDE approximates nonlinear system evolution using a learned linear operator in a latent space, combining (i) a history encoder that reconstructs the current latent state from past targets and inputs with (ii) a shared linear rollout driven by future exogenous variables. By enforcing a common linear recurrence across encoding and prediction, SLiDE aligns how past information is encoded with how future states evolve, reducing parameterization and preventing reliance on spurious correlations in the input history. Empirically, SLiDE achieves state-of-the-art accuracy on real-world benchmarks, including electricity price forecasting and retail demand, while maintaining a computational footprint comparable to lightweight MLP architectures. SLiDE is especially effective at generalizing to shifted exogenous inputs, reducing MSE by 20% on electricity price samples with out-of-range exogenous values and MAE by 30% in a synthetic transfer setting where the same dynamical system is driven by an unseen exogenous process. Ablations confirm the importance of both the structured history encoder and the linear latent rollout, suggesting that structured latent linear dynamics provide a useful inductive bias for forecasting with exogenous inputs.
Slower Generalization, Faster Memorization: A Sweet Spot in Algorithmic Learning
Shin So ⋅ Kyelim Lee ⋅ Albert No
Critical-data-size accounts of grokking suggest a natural post-threshold intuition: once training data is sufficient to identify the underlying rule, additional data should accelerate validation convergence. We show that this intuition can fail in a controlled structured-output task. In Needleman--Wunsch (NW) matrix generation, small Transformers reach high validation exact-match accuracy fastest at an intermediate dataset size, not at the largest one. Past this dataset-size sweet spot, generalization remains achievable but requires more gradient updates. Conversely, in the regime where partial validation competence first appears, larger datasets can require fewer updates to reach high training accuracy, suggesting that emerging rule structure can accelerate fitting beyond example-wise memorization. A multiplication baseline does not show the same post-threshold slowdown. These results separate the critical data size for the onset of generalization from the dataset size that optimizes update-based convergence, and identify structured-output tasks where learning the rule and completing exact-fitting can diverge.
Slowly Annealed Langevin Dynamics: Theory and Applications to Training-Free Guided Generation
Atsushi Nitanda ⋅ Dake Bu ⋅ Yueming LYU ⋅ Tanya Veeravalli
We study Slowly Annealed Langevin Dynamics (SALD), a sampler for tracking a path of moving target distributions and approximating the terminal target through time slowdown. We establish non-asymptotic convergence guarantees via a KL differential inequality, showing that slowdown improves tracking through contraction of intermediate targets and the complexity of the path. Motivated by training-free guided generation with pretrained score-based generative models, we further introduce Velocity-Aware SALD (VA-SALD), which explicitly incorporates the underlying marginal distributions of the pretrained model and uses slowdown to correct the additional deviation induced by guidance. This yields a principled framework for training-free guided generation for diffusion-based and related generative model families, together with convergence guarantees that clarify the roles of intermediate functional inequalities and guidance bias.
SoccerNarrate: Event-Grounded Streaming Soccer Commentary with Macro-Window Preference Alignment
zihan jia ⋅ Zhilin Dai ⋅ Zhengming Zhang ⋅ Min Yang ⋅ Zhenpeng Huang ⋅ Jiaqi Li ⋅ Caixia Sun ⋅ Xi Chen ⋅ Liang Li ⋅ Junlan Feng ⋅ Gangshan Wu ⋅ Limin Wang
Existing soccer commentary models are often designed for pre-segmented clips or localized events. When deployed on untrimmed full-match videos with sliding windows, they can produce delayed, repeated, or poorly synchronized commentary. Recent streaming video-language models make low-latency full-match narration feasible, but they often favor general real-time descriptions rather than professional, event-grounded soccer commentary. They may describe nearby actions while failing to mention key events such as goals, cards, substitutions, and offsides correctly and on time. In this paper, we introduce SoccerNarrate, a data, model, and evaluation framework that bridges low-latency streaming narration with event-grounded professional soccer commentary. First, we construct SoccerNarrate-Data, a large-scale time-anchored dataset with word-level ASR alignment, quality filtering, and entity calibration. Second, we perform macro-window preference alignment(MWPA), which aligns the model with complete event-level commentary semantics by comparing causal multi-step rollouts from the same streaming prefix while keeping second-level inference unchanged. We further use event-mismatched human commentaries as counterfactual negatives to emphasize event semantics over commentary style. Third, we introduce SoccerNarrate-Eval, an event-centric benchmark based on temporally constrained event entailment. Full-match experiments show that SoccerNarrate improves event coverage and event-aligned precision over strong offline and streaming baselines. Data and models will be released.
SOC-ICNN: From Polyhedral to Conic Geometry for Learning Convex Surrogate Functions
Kang Liu ⋅ Jianchen Hu ⋅ Wei Peng
Classical ReLU-based Input Convex Neural Networks (ICNNs) are equivalent to the optimal value functions of Linear Programming (LP). This intrinsic structural equivalence restricts their representational capacity to piecewise-linear polyhedral functions. To overcome this representational bottleneck, we propose the SOC-ICNN, an architecture that generalizes the underlying optimization class from LP to Second-Order Cone Programming (SOCP). By explicitly injecting positive semi-definite curvature and Euclidean norm-based conic primitives, our formulation introduces native smooth curvature into the representation while preserving a rigorous optimization-theoretic interpretation. We formally prove that SOC-ICNNs strictly expand the representational space of ReLU-ICNNs without increasing the asymptotic order of forward-pass complexity. Extensive experiments demonstrate that SOC-ICNN substantially improves function approximation, while delivering competitive downstream decision quality. The code is available at \url{https://anonymous.4open.science/r/SOC-ICNN-4B18/}.
Sparse Biological Features Reveal Early Functional Commitment in Diffusion Protein Language Models
Chuyang Zhou ⋅ Bingxin Zhou ⋅ Andi Han ⋅ Chang Xu
Generative protein models are increasingly used for functional sequence design, yet the process by which biological information is organized during generation remains poorly understood. Diffusion protein language models provide a tractable setting for this question, as sequence generation proceeds through an explicit denoising trajectory from highly corrupted states to complete proteins. This trajectory enables a temporal view of interpretability: beyond identifying which biological signals are represented, it allows us to examine when these signals emerge and when specific residues become committed. Here, we analyze DPLM representations using sparse autoencoders trained across layers and noise levels. The resulting features capture biologically meaningful signals at both residue and protein scales, align with functional annotations, and preserve downstream biological information under reconstruction. Following these features along the denoising trajectory reveals a consistent functional ordering: catalytic-enriched features pre-activate at still-masked catalytic positions before residue identity is resolved, and catalytic residues are subsequently recovered earlier by the iterative denoiser. This prioritization remains significant after controlling for prediction difficulty, sequence context, amino-acid identity, structural environment, and evolutionary conservation, and is not reproduced by a random-feature null. Across ProteinGym deep-mutational-scanning assays, residues recovered earlier during denoising are also more mutation-sensitive. These results reveal a temporal organization of biological information in diffusion-based protein generation, in which functionally important residues are not only represented, but preferentially committed during sequence formation.
Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity
Loic Cabannes ⋅ Pierre-Emmanuel Mazare ⋅ Ilze Amanda Auzina ⋅ Gergely Szilvasy ⋅ Maria Lomeli ⋅ Matthijs Douze ⋅ Justin Carpentier ⋅ Gabriel Synnaeve ⋅ Herve Jegou
Linear attention models allow a fixed state size and a fixed amount of compute per token. However, due to their limited state size, linear attention models fall behind in long-context recall compared to softmax-attention-based transformer architectures. Increasing the state size of linear attention improves recall performance but at the cost of higher FLOPs. In this work, we introduce Sparse Delta Memory (SDM), an architecture that scales the hidden state of gated linear RNNs to orders of magnitude higher capacity using a sparse addressing scheme. SDM extends the Gated DeltaNet architecture by replacing the dense key-value outer product with sparse reads and writes to a large explicit memory. We show that, under an isoFLOP constraint and with an identical number of parameters, a higher state memory capacity significantly improves performance on in-context learning and long-context retrieval tasks. Moreover, by learning the initial state of the SDM memory and therefore using it as a parametric memory, we show that the model further improves on a wide range of common-knowledge and reasoning tasks.
Sparse Koopman Autoencoders Identify Local Dynamical Regimes in Multibasin Systems
Aidan Li ⋅ Uday Kiran Reddy Tadipatri ⋅ Mahan Fathi ⋅ Sarath Chandar ⋅ Ross Goroshin
Koopman autoencoders (KAEs) seek a higher-dimensional latent representation in which nonlinear dynamics evolve linearly. However, many interesting systems have multiple basins of attraction, and both theoretical and empirical work has shown these multibasin systems cannot generally admit a single finite-dimensional global Koopman embedding under standard assumptions. We posit that encoders with a sparsity-inducing objective encouraging few active latent coefficients will provide latent supports as an inspectable basin-modeling principle for Koopman autoencoders. We use these encoders producing sparse latents in training Sparse Koopman Autoencoders (SKAEs) without basin labels or other regime annotations, and treat the learned latent supports as model-produced regime variables after training. Across a range of procedurally generated multibasin systems and chaotic flows, we show that SKAEs have superior forecasting performance compared to dense-latent KAEs. We also perform a mechanistic study that shows latent supports produced by SKAEs are both essential for the quality of the representation and useful for identifying basins on held-out basin interior states, whereas dense-latent KAEs collapse to an uninformative single family. These results identify sparse latents and their corresponding supports as label-free, interpretable regime variables for Koopman learning in nonlinear systems with multiple local dynamical laws.
Sparsely Supervised Diffusion
Wenshuai Zhao ⋅ Zhiyuan Li ⋅ Yi Zhao ⋅ Mohammad Vali ⋅ Martin Trapp ⋅ Joni Pajarinen ⋅ Juho Kannala ⋅ Arno Solin
Diffusion models have shown remarkable success across a wide range of generative tasks. However, they often suffer from spatially inconsistent generation, arguably due to excessive correlations learned by the model. This can produce samples that are locally plausible but globally inconsistent. We propose a principled method to mitigate this issue. Our method, sparsely supervised diffusion (SSD), is a simple yet effective masking framework that can be implemented with only a few lines of code. We analytically show that SSD fundamentally alters diffusion model training by modifying the spectrum of the data covariance and that it suppresses correlations in the data covariance matrix. Experiments show that our method more accurately approximates the underlying population score function, reduces memorization on small datasets, and promotes the use of essential contextual information during generation. Moreover, even when up to 98\% of pixels are masked, it achieves competitive FID scores across a range of datasets and, importantly, avoids the training instability commonly observed on small datasets.
Token-based world models enable fine-grained latent planning, but iterative search inside the predictor is dominated by the number of spatial tokens processed during rollout. We introduce COSTGRAD, a training-free, goal-conditioned selector that ranks spatial tokens by the gradient norm of the planning cost with respect to each input token. By deriving importance from the downstream control objective, COSTGRAD targets tokens that matter for planning rather than merely for prediction. On AdaLN-conditioned predictors at $50\%$ sparsity, COSTGRAD matches or exceeds full-token planning on three of four continuous-control benchmarks, while giving a measured $2.6\times$ wall-clock speedup per environment planning step. The sparsity dividend can be combined with reduced CEM search, yielding a $\sim5\times$ total speedup while still exceeding the full-token baseline. We also identify an architecture-dependent failure mode: in a matched AdaLN-vs-concat comparison, concat maintains comparable full-token performance but pure COSTGRAD loses its advantage over random selection. This failure tracks action-pathway drift under selected token removal: AdaLN largely preserves how actions affect kept tokens, whereas concat perturbs the action-conditioned computation. Mixing in random anchors partially recovers COSTGRAD's advantage on some matched-concat checkpoints, suggesting selector--architecture compatibility as a design axis for sparse world-model planning.
Spectral Asymptotics of Neural Network Jacobians: Convergency, Universality, and Phase Transition
Huiqin Li ⋅ Guangming Pan ⋅ Yanqing Yin
We study the spectral properties of the Jacobian matrix in artificial neural networks and establish both its limiting spectral distribution and the second-order fluctuations of its spectrum. Our analysis uncovers fundamental phenomena, including universality properties with respect to the weight distribution and a phase transition in spectral behavior driven by the choice of activation function. Leveraging tools from random matrix theory, we first analyze a class of nonlinear covariance ensembles, which may be of independent interest in high-dimensional statistics, and subsequently characterize the spectral fluctuations of the Jacobian matrix, highlighting their dependence on weight initialization, layer widths, activation functions, and network depth. These findings provide finite-width distributional calibration for Jacobian-based stability diagnostics and quantify initialization-induced uncertainty in signal propagation, gradient stability, and local sensitivity of high-dimensional neural networks.
We study the spectral structure of the attention matrix in Transformer models through an idealized spectral model for the pre-softmax score matrix. Motivated by the rotational isotropy of query and key representations, we replace the random token-space orientation of the score matrix by its canonical singular-value representative $\Sigma$, and assume that $\Sigma$ has a logarithmic spike-bulk separation. We provide empirical diagnostics consistent with this model on pretrained vision and language transformers: token-space singular vectors exhibit near-Haar behavior, and score spectra are dominated by a small number of effective directions. We then prove that the row-wise softmax map transforms this structure into a degenerate attention matrix: dominant modes are preserved, whereas the bulk spectrum converges weakly to $\delta_0$. This provides a unified theoretical explanation for why attention layers can act as low-rank information extractors: they preserve a small number of dominant token directions while compressing the remaining bulk modes.
Spherical Flows for Sampling Categorical Data
Jannis Chemseddine ⋅ Gregor Kornhardt ⋅ Gabriele Steidl
We study the problem of learning generative models for discrete sequences in a continuous embedding space. Whereas prior approaches typically operate in Euclidean space or on the probability simplex, we instead work on the sphere $\mathbb S^{d-1}$. There the von Mises-Fisher (vMF) distribution induces a natural noise process and admits a closed-form conditional score. The conditional velocity is in general intractable. Exploiting the radial symmetry of the vMF density we reduce the continuity equation on $\mathbb S^{d-1}$ to a scalar ODE in the cosine similarity, whose unique bounded solution determines the velocity. The marginal velocity and marginal score on $(\mathbb S^{d-1})^L$ both decompose into posterior-weighted tangent sums that differ only by per-token scalar weights. This gives access to both ODE and predictor-corrector (PC) sampling. The posterior is the only learned object, trained by a cross-entropy loss. Experiments compare the vMF path against geodesic and Euclidean alternatives. The combination of vMF and PC sampling significantly improves results on Sudoku and language modeling.
Spik-NeRF v2: Pushing the Limit of Spiking Neural Radiance Fields with $ \pm $I-LIF
Chen Cheng ⋅ Qinlong Lan ⋅ Gang Wan ⋅ Hao Guo ⋅ Lei Liu ⋅ Zhanji Wei ⋅ Wu Yitian ⋅ Yufei Guo
Spiking Neural Networks (SNNs) offer a promising energy-efficient alternative to Artificial Neural Networks (ANNs) through event-driven, multiplication-free computation. However, when applied to downstream tasks such as Neural Radiance Fields (NeRF), SNNs suffer from significant information loss, resulting in a noticeable performance gap. In this paper, we present \textit{Spik-NeRF v2} to advance SNN-based neural rendering. We propose the $\pm$I-LIF spiking neuron, which extends I-LIF to support signed integer values during training, thereby mitigating information loss. During inference, it is converted to ternary spikes $\{-1, 0, 1\}$, preserving event-driven properties with addition-only operations. Furthermore, we introduce a re-parameterization technique that transforms a trained I-LIF-based Spik-NeRF with $t$ timesteps into an equivalent $\pm$I-LIF-based model with $t/2$ timesteps. This enables faster inference while preserving rendering quality, overcoming the suboptimal performance of directly trained $\pm$I-LIF models with reduced timesteps. Extensive experiments on synthetic and realistic datasets demonstrate that Spik-NeRF v2 surpasses existing SNN-based NeRF methods and achieves rendering quality comparable to ANN-based approaches.
SPLICE: Structured Prompt Local Iterative Combinatorial Evolution
Dr. Anish Acharya ⋅ Phillip Studans ⋅ Amit Dhanda ⋅ Ninad V Rao ⋅ Vishal A Khatri ⋅ Sina A Niaki ⋅ Brian Verkhovsky
Prompt engineering has become a central lever for deploying large language models (LLMs), yet the prompts that actually ship to production bear little resemblance to the short instructions studied in most of the prompt-optimization literature. A deployed prompt is a structured multi-section artifact---role framing, reasoning directives, examples, constraints, tool schemas, and output specifications woven into a single token stream---in which some sections encode hard-won safety and format contracts while others are the legitimate targets of optimization. Existing iterative prompt optimizers treat the prompt as a monolithic string and the search as an unstructured rewrite loop, with no mechanism to respect this partition and no theoretical account of when their updates improve the prompt, preserve diversity, or control length. We address this gap by casting iterative prompt optimization as combinatorial search over an edit graph of block-structured, selectively-mutable prompt states. Under this view, a broad slice of the recent literature collapses to selector--proposal--state-space instances of a single \emph{batched iterative prompt search} template that differ only in the selector. Within this framework we introduce \textsc{Splice}, which contributes four coupled design choices, each paired with a formal guarantee: (i) per-block mutability, with a price-of-freezing bound quantifying the safety--optimality trade-off; (ii) dual section-local textual gradients with momentum, yielding an edit-graph reachability bound and a ${(1-e^{-\gamma})}$ sub-modular approximation; (iii) elitist beam search with lineage-aware diversity, giving deterministic monotonicity, a noisy no regression tail bound, and conditional no-collapse of lineages; and (iv) a ratio regularizer for length control, with an explicit accuracy--length bias bound. Together these constitute, to our knowledge, the first theoretical analysis of iterative LLM-driven prompt optimization. Empirically, \textsc{Splice} consistently outperforms prior prompt-optimization methods and single-step sampling baselines across public benchmarks, system models spanning three capability tiers, and tasks ranging from classification to LLM-as-judge rubric optimization, while preserving frozen sections verbatim and exhibiting narrower cross-model variance than single-step sampling.
Split-and-scale Latent 3D Representations
Dorian Chan ⋅ Muhammed Kocabas ⋅ Xiaoming Zhao ⋅ Oncel Tuzel ⋅ Jen-Hao Chang
Latent 3D representations have significantly expanded the capabilities of modern shape modeling, enabling compact encoding of 3D scenes and supporting powerful 3D generative models. Despite this progress, current representations struggle to capture the fine details of complex objects, fundamentally limited by the capacity of their latents. However, scaling latent capacity at training time is computationally expensive and often infeasible under memory limits, and naively increasing the number of latents at inference fails to make effective use of the added capacity. In this paper, we address both challenges with split-and-scale, a method that effectively expands latent capacity at inference while keeping training cost tractable. Our key observation is that by representing 3D scenes as continuous functions of 3D coordinates, a single pretrained tokenizer can be applied at inference to arbitrarily small sub-regions of a scene, allocating the same latent capacity to a smaller spatial extent and thus capturing far higher detail. We validate this insight by conducting extensive ablations to study the trade-offs between compute, memory, latent capacity, and reconstruction quality, demonstrating our inference-time scaling method achieves quality competitive with much larger models that exceed single-GPU memory limits. We further demonstrate that split-and-scale supports high-quality generative modeling, training an autoregressive 3D generator that performs competitively with state-of-the-art baselines.
Split Then Select: Moment-Preserving Density Control for Generalized Primitive Splatting
Yangkai Lin ⋅ Jiehong Lin ⋅ Kui Jia
Primitive splatting has become a powerful representation for efficient differentiable rendering, but existing density-control strategies remain largely heuristic and often suffer from uncontrolled primitive growth. In this work, we present \textbf{Moment-Preserving Density Control (MPDC)}, a general theoretical framework for densification in primitive splatting. Unlike conventional density-control strategies that \emph{select first and split later}, MPDC follows a \emph{split-then-select} paradigm. Our key insight is that a principled split should first preserve the current rendering, rather than immediately perturb the image. By matching local moments, MPDC decomposes a primitive into multiple offspring while keeping the rendered output, and hence the loss, unchanged up to higher-order error. Although such a split does not directly reduce the loss, it changes the parameter space: a saddle point in the original parameterization may no longer remain a saddle point after splitting. This provides a distinct saddle-escaping mechanism from prior approaches that move primitives along negative-curvature directions. Building on this view, we derive closed-form splitting rules for different primitive parameterizations and then select only primitives whose moment-preserving splits reduce an upper bound of the loss through a splitting-matrix criterion. Experiments on three diverse primitive splatting methods show that MPDC substantially reduces primitive counts without sacrificing rendering quality, while improving memory efficiency and rendering speed. Code will be released upon publication.
SSDGExplainer: Structure-Semantic Dual-Guided Explainer for Graph Neural Networks
Zhiqiang Wang ⋅ Chenchao Zhang ⋅ Jianqing Liang ⋅ Xingwang Zhao ⋅ Jiye Liang ⋅ Chuangyin Dang
Post-hoc Graph Neural Networks (GNN) explainers typically extract a compact subgraph to preserve the model’s decision rationale. However, this extraction breaks topological integrity and induces a distribution shift, making predictions on subgraphs unreliable and consequently misleading explainer optimization under OOD settings. Existing attempts to build in-distribution proxy graphs often assume independence between explanation and background subgraphs and rely on hard splicing, which causes semantic mismatch and boundary discontinuities that degrade explanation reliability. We propose SSDGExplainer, a Structure–Semantic Dual-Guided explainer that formulates proxy-graph generation as conditional modeling under semantic consistency constraints to achieve deep semantic alignment between the explanation subgraph and the generated background, introduces a topology boundary optimization network to smooth structural fractures at the explanation–background interface, and enforces contrastive semantic constraints to prevent semantic drift. Extensive experiments on synthetic and real-world benchmarks demonstrate consistent improvements over state-of-the-art methods in both explanation quality and fidelity.
SSR3D-LLM: Structured Spatial Reasoning via Latent Steps for Fine-Grained Grounding in Unified 3D-LLMs
Jiawei LI ⋅ Ziyi Liu ⋅ Weijie Shi ⋅ Long Chen ⋅ Jiajie Xu ⋅ Xiaofang Zhou
3D object grounding localizes referred objects in a 3D scene from natural language. Unified instance-centric 3D-LLMs aim to solve grounding together with dialog, QA, and captioning, yet many rely on a single pointer-style grounding decision that compresses a relational instruction into one selection. This is brittle for fine-grained queries where multiple same-class candidates must be ruled out by context objects and spatial relations. We propose Structured Spatial Reasoning 3D-LLM (SSR3D-LLM), a structured grounding interface for unified 3D-LLMs. Given fixed Mask3D object proposals, the LLM writes a sequence of latent spatial reasoning steps and memory tokens from the query, and a geometry-aware scorer reads these latent steps in order to refine candidate rankings step by step with step-length masking. The latent steps are learned from standard benchmark target supervision with auxiliary referential-cue supervision during training, while inference uses only the input query and Mask3D proposals. Across ReferIt3D, ScanRefer, and Multi3DRef, SSR3D-LLM achieves the strongest results among unified 3D-LLM baselines, with substantial gains over the single-pointer QPG baseline on fine-grained grounding and consistent improvements over prior unified 3D-LLMs, while preserving the default language-task route.
Stability-Aware Self-Training for CLIP under Cross-Modal Anchoring Mismatch
Siyuan LIU ⋅ Xinyang Chen ⋅ Xiucheng Li ⋅ Weili Guan ⋅ Liqiang Nie
Self-training is a practical way to adapt CLIP with limited supervision, but its performance is highly sensitive to pseudo-label reliability. We identify an important failure mode in CLIP adaptation: temporal prediction instability, where a target sample repeatedly flips its prediction across epochs. We provide a perspective based on cross-modal anchoring mismatch: the decision geometry induced by pretrained text anchors can deviate from the local visual cluster structure of the target domain. This mismatch can make ambiguous target samples susceptible to prediction flips during adaptation, whereas target-aligned anchoring produces more stable pseudo-labels. Motivated by this observation, we propose Stability-Aware Self-Training (SAST), centered on Stability-Weighted Image Prototypes (SWIP). SWIP builds class-wise visual anchors from temporally stable target samples to provide more reliable adaptation references. We further add a lightweight text-side refinement to reduce confusion-induced instability, and combine the two streams with an adaptive fusion rule. Across six standard benchmarks and three adaptation settings, SAST consistently outperforms strong baselines. The results suggest that explicitly modeling temporal stability is an effective route to more robust CLIP adaptation under self-training.
Stabilizing Few-Shot Object Detection with Language-Conditioned Probabilistic Prototypes
Jiaying Wu ⋅ Yuchun Zhao ⋅ Lei Shi ⋅ Jia Luo ⋅ Dongmei Wei ⋅ Feifei Kou ⋅ Pengfei Zhang ⋅ Zhaoman Zhong
Few-shot object detection (FSOD) requires detectors to adapt to novel categories from only a few labeled instances, where prototype-based transfer methods have become a strong and efficient paradigm. However, we observe that such methods remain unstable in extreme low-shot regimes. We attribute this instability to two coupled factors: semantically unanchored query initialization, which yields high-variance cold-start optimization, and deterministic prototype matching, which over-trusts noisy or occluded support regions. To address these limitations, we propose Semantic Fine-Grained Prototype Distillation (SFPD), a language-conditioned probabilistic adaptation framework for prototype-based FSOD. SFPD introduces language-conditioned feature queries to provide semantic anchors for novel-class adaptation, uncertainty-aware prototype distillation to down-weight unreliable support evidence through heteroscedastic Gaussian modeling, and component-guided part-aware prototypes to refine fine-grained semantic-visual alignment. These modules act during adaptation and preserve the original detector path at inference, introducing no additional inference FLOPs. Experiments on PASCAL VOC and MS-COCO show that SFPD consistently improves a strong FPD baseline, with especially clear gains in the most challenging low-shot settings, e.g., a 3.9-point nAP50 improvement on VOC Split 2 under 1-shot. Further ablations, multi-seed evaluation, and convergence analysis indicate that SFPD improves both accuracy and adaptation stability.
Stabilizing the Dynamic Low-Rank Training
Zhonghan Xu ⋅ Ling Wang ⋅ Junhao Chen ⋅ Jianwei Zhao ⋅ Jinwei Yang
Training neural networks directly in a low-rank parameterization is an appealing route to reducing memory, compute, and storage simultaneously during both training and inference. Dynamic low-rank training (DLRT), which confines weights to a rank-$r$ manifold via the Galerkin projection of the gradient flow, is particularly attractive because it identifies efficient subnetworks on the fly without specialized initialization or post-factorization. However, DLRT fails to find trainable networks under high compression. In this paper, we derive the gradient flow of the best rank-$r$ approximation and point out that the offset of DLRT comes from a curvature-coupling term which is large and thus non-negligible under aggressive compression. Guided by this analysis, we propose a stable dynamic low-rank training method, named SDLRT, which maintains a lightweight compensation buffer that reinjects the top neglected singular directions. Additionally, we introduce a negative feedback on the truncation tolerance to stabilize each layer's rank. Experimentally, SDLRT reliably finds trainable subnetworks where DLRT collapses and as a PEFT adapter on DeBERTa-v3, it achieves the best average score on SuperGLUE at only $2.8\\%$ parameter overhead over LoRA.
Stable Alpha: Adversarial Invariant Representation Learning for Nonlinear Asset Pricing under Temporal Distribution Shifts
Xiaokang Wang ⋅ Zihe Liu ⋅ Xinghan Qin ⋅ Zihao Yin ⋅ Jidong Yuan ⋅ Qinxuan Zhang ⋅ Xiushuo Hu
Deep learning has emerged as a powerful paradigm for constructing nonlinear asset pricing factor models. However, temporal distribution shifts such as bull-bear transitions and industry rotations undermine the generalization of models trained under the i.i.d. assumption. As a result, existing models tend to exploit environment-specific correlations that are highly predictive in-sample but unstable across market regimes. To address this challenge, we propose $\textit{\textbf{CASH}}$, a causality-inspired asset pricing factor framework designed to learn invariant and minimally sufficient representations under temporal distribution shifts. CASH formulates an information-theoretic adversarial learning objective that systematically discards transient market noise by optimizing against an adversary designed to exploit spurious correlations, thereby distilling invariant factor-return relationships. The resulting invariant representation is then integrated into a conditional factor pricing model via a dynamic factor exposure network. Experiments on real-world stock datasets across multiple markets demonstrate that CASH consistently outperforms state-of-the-art baselines in out-of-sample return prediction and exhibits superior robustness under pronounced temporal regime shifts.
StableVQ: Practical Guidelines for Stable Vector-Quantized Tokenizer Training
Bao Tang ⋅ Jiahao Guo ⋅ Haoxiang Cao ⋅ Wenyu Liu ⋅ Changqian Yu ⋅ Kun Gai ⋅ Xinggang Wang
Vector Quantization (VQ) is fundamental to discrete visual tokenizers that power modern autoregressive and masked image generation models. While recent shared-projection codebook methods have substantially advanced codebook utilization, training stability remains a critical and underexplored challenge. We argue that the root cause lies in the entanglement of the Encoder--Decoder and Codebook training: because neither module can reliably fulfill its own responsibility in isolation, the system can only function when the two subsystems happen to cooperate---a fragile condition that breaks down precisely when training is most stressed. We propose StableVQ, which revisits the proper learning objective of each module and resolves the problems that arise when each is trained to fulfill its own role independently. Concretely, (1) Dynamic STE corrects the instability in the Encoder's learning objective, enabling it to robustly optimize the reconstruction space under discrete regularization even when codebook utilization is low. (2) Region VQ Loss reconceives the Codebook's learning objective so that it can independently guarantee full tracking of the encoder output distribution, without relying on encoder oscillations to drive activation. (3) Decoupled Schedule recognizes that the distinct responsibilities of the Encoder--Decoder and the Codebook demand distinct optimization dynamics, and assigns each an independent learning rate schedule to ensure robust system-level behavior. Built on top of shared-projection codebooks, StableVQ is lightweight and introduces no learnable parameters. Experiments on ImageNet demonstrate consistent improvements in training stability, codebook utilization, and reconstruction quality across diverse codebook sizes and initialization settings.
StainNFT: Curriculum-Gated Multi-Reward Post-Training for Pathology-Faithful Virtual Staining
shurong yang ⋅ Dong Wei ⋅ Qiong Peng ⋅ Xian Wu ⋅ Yefeng Zheng ⋅ Liansheng Wang
Immunohistochemical (IHC) staining encodes molecular protein expression critical for clinical diagnosis, yet its chemical procedures are costly and time-consuming. Virtual staining offers a compelling alternative by digitally synthesizing IHC images from hematoxylin-eosin (H&E) stained slides. Despite remarkable progress in diffusion- and flow-matching-based virtual staining, supervised fine-tuning (SFT) remains fundamentally limited in pathological fidelity due to its coarse, spatially averaged supervision signal that starves sparse DAB-positive regions of effective gradients. Reinforcement learning (RL) can exploit inter-rollout variance for outcome-level pathological supervision, but naively applying DAB-based rewards triggers severe reward hacking that collapses the image distribution and degrades perceptual quality. This paper presents StainNFT, a flow matching RL post-training framework built on a curriculum reward strategy that gates fine-grained optical density supervision on per-sample DAB mask IoU. To enhance fine-grained pathological fidelity, we introduce expression-aware reweighting, multi-scale block supervision, and a closed-loop implicit pathological semantic reward derived from pathology foundation model priors. Extensive experiments on seven benchmarks demonstrate that StainNFT consistently outperforms existing methods in both perceptual quality and pathological fidelity for H&E-to-IHC virtual staining. Ablation studies confirm the effectiveness of each proposed component. Our code and trained models will be released.
Standardization of Post-Publication Code Verification is Possible with the Support of the Community
Susana López-Moreno ⋅ Eric R Dolores-Cuenca ⋅ Sangil Kim
Reproducibility remains a challenge in machine learning research. While code and data availability requirements have become increasingly common, post-publication verification is still limited and unformalized. This position paper argues that journals and conference proceedings should all implement a post-publication verification submission system. We propose a modification to ACM pre-publication verification badges and IEEE post-publication badges that allows independent researchers to submit post-publication code replications directly to the journal or proceeding, leading to visible verification badges included in the article metadata. Each article may earn up to two badges, each linked to verified code in its corresponding public repository. This framework would be a step forward toward a standardized reproducibility policy and would provide young researchers with an opportunity to improve their curriculum. We describe the motivation, related initiatives, a formal framework, the potential impact, possible limitations, and alternative views.
STAR: Boosting Time Series Foundation Models for Anomaly Detection Through State-Aware Adapter
Hanyin Cheng ⋅ Ruitong Zhang ⋅ Yuning Lu ⋅ Yang Shu ⋅ Peng Chen ⋅ Yuan Jun ⋅ Meng Wang ⋅ Bin Yang ⋅ Chenjuan Guo
Existing Time Series Foundation Models (TSFMs) for Multivariate Time Series Anomaly Detection often overlook discrete state variables that describe system status and treat them uniformly with numerical variables. This inappropriate modeling approach prevents the model from fully leveraging state information and even leads to significant performance degradation when state variables are integrated. To address this limitation, this paper proposes a novel STate-aware AdapteR (STAR). Specifically, STAR comprises three core innovative components: (1) an Identity-guided State Encoder that effectively captures the complex semantics of state variables; (2) a Conditional Bottleneck Adapter that dynamically injects state influence into TSFMs; and (3) a Numeral-State Matching module that effectively detects anomalies inherent to the state variables themselves. Extensive experiments on real-world datasets demonstrate that STAR significantly improves the performance of existing TSFMs.
State Copying Crowds Out Reasoning: Mechanistic Evidence for Delta Planning in Autoregressive Models
Wang Xi ⋅ Shijia Xu
Autoregressive models are often used as sequential decision makers by asking them to rewrite the full intermediate state before producing the next action. We study a failure mode of this interface: exact high-entropy state copying can interfere with goal-directed computation. We formulate this as an operational copy--reason interference hypothesis rather than as a directly measurable capacity law. Matched controls separate copying from length: on controlled state-tracking tasks, Full-State Generation degrades much more sharply than an Iso-Length Control and similarly to a Complex-Copy Control. Activation patching on Llama-3.1 8B and 70B recovers target-action evidence when clean residual activations are inserted into late layers of full-state runs, suggesting that goal features remain recoverable but are poorly routed under the copy-heavy interface. Delta-style alternatives reduce this burden. At the 2,000-token Formal Math anchor, Scratchpad Residual improves accuracy from 41.2--68.5\% under Full-State Generation to 97.1--99.4\%; on SWE-bench Lite, it improves pass rates from 12.4--26.8\% to 18.5--34.8\%. The gain has a boundary: probe and task accuracy both decay with long dependency distance. These results argue for separating state maintenance from action generation in long-horizon autoregressive systems.
ST-DiffEye: Diffusion-based Continuous Gaze Generation via Joint Scanpath-Trajectory Modeling
Brian Nlong Zhao ⋅ Ozgur Kara ⋅ Junho Kim ⋅ James Rehg
We study the problem of human gaze modeling, which aims to generate the gaze patterns a viewer produces while observing a visual stimulus. Gaze is primarily captured through two modalities: continuous eye-tracking trajectories, which describe fine-grained motion dynamics, and discrete scanpaths, which describe high-level fixation structure. Because gaze varies substantially across viewers and trials, we treat this variability as a defining property rather than noise and model gaze as a stochastic generative process. Existing generative gaze models supervise on only one of these two representations in isolation. We hypothesize that trajectories and scanpaths describe gaze at complementary scales and are jointly informative during training, and test this hypothesis through ST-DiffEye, a joint trajectory-scanpath diffusion framework that couples both modalities by concatenating them as an additional raw input channel, requiring no architectural overhead beyond an input and output channel expansion. We further introduce a principled evaluation framework based on the Continuous Ranked Probability Score (CRPS), which generalizes any existing sequence similarity metric into a proper scoring rule that jointly assesses the accuracy and diversity of generated gaze. Experiments on task-driven visual search, covering both target-present and target-absent scenarios, and on free-viewing benchmarks demonstrate state-of-the-art performance. These results, along with detailed ablations, confirm the benefit of joint modeling and the value of distribution-aware evaluation in capturing the intrinsic variability of human gaze.
SteerCast: Retrieval-Based Latent Steering for Decoder-Only Time Series Forecasting
Van Dai Do ⋅ Huu H Nguyen ⋅ Minh Hoang Nguyen ⋅ Hung Le
Time series forecasting aims to predict future values from historical observations and auxiliary features. We propose \textbf{SteerCast}, a retrieval-based latent steering method that improves decoder-only forecaster at inference time, without updating its parameters. SteerCast constructs a database from the training set by storing a representation of each history window together with a \emph{steering vector} computed in the forecaster's latent space, defined as the difference between representations induced by the ground-truth continuation and by the model's own prediction. At test time, SteerCast retrieves nearest neighbors for a query history, aggregates their steering vectors, and injects the resulting signal into the forecaster's hidden states at every step of autoregressive generation, guiding predictions toward trajectories consistent with similar training cases. Experiments across diverse multivariate benchmarks and multiple horizons show that SteerCast consistently improves forecasting accuracy over the fine-tuned backbone and retrieval-based baselines, while requiring no additional training beyond the original fine-tuning and using only the training set as a retrieval corpus.
STEER: Route-Aware Adaptive Reasoning for Autonomous Driving
YANGANG ZOU ⋅ Pei Liu ⋅ Nan Song ⋅ Bozhou Zhang ⋅ Mingyu Guo ⋅ Jun Ma ⋅ Jiankang Deng ⋅ Xiatian Zhu ⋅ Li Zhang
Vision-language action (VLA) models excel for end-to-end autonomous driving, yet their tendency to over-invoke chain-of-thought (CoT) reasoning mirrors a human cognitive pitfall: overthinking. Just as deliberate reasoning can slow/mislead human judgment in routine tasks, excessive CoT in VLA incurs unnecessary compute overhead and can paradoxically degrade planning performance. Recent research tackles this via adaptive reasoning, which dynamically modulates inference complexity by selectively triggering CoT for complex scenes while responding directly to routine ones. While prior adaptive methods rely on implicit adaptation, recent work shows that explicitly optimizing the route between CoT and direct responses yields superior performance, assuming scene-difficulty annotations which are labor-intensive, poorly scalable, and prone to error due to the subjectivity of judging driving complexity. We explore, for the first time, whether adaptive reasoning with explicit route optimization can be achieved without any scene-difficulty labeling. We introduce STEER, a generic over-reasoning mitigation strategy featuring two key components: (i) routing uncertainty-aware rollout, which calibrates sampling to align routing diversity with model uncertainty; and (ii) cross-route advantage credit assignment, which introduces a differential metric to reinforce optimal routing decisions based on environmental rewards rather than subjective labels. Extensive experiments on NAVSIM v1/v2 demonstrate state-of-the-art results in both planning quality and the Pareto front of inference efficiency.
STEP: Learning STructured Embeddings for Progressive Time Series
Lucas Thil ⋅ Jesse Read ⋅ Rim Kaddah ⋅ Guillaume Doquet
We present a novel method for learning interpretable representations of progressive time series, that is, data capturing irreversible state transitions such as degradation or task completion. Our approach uses a self-supervised contrastive objective to learn a low-dimensional latent space where progression manifests along manifolds anchored by fixed prototype vectors. This structure yields latent-space indicators that quantify progression in a human-meaningful way without proxy labels. We evaluate the approach against the state of the art on diverse domains, including industrial degradation, robotic tasks, and neural activity, validating three key capabilities: (1) end-state prediction, (2) multi-step forecasting, and (3) interpretable phase separation. Our method matches or improves over black-box counterparts on all of these while providing transparency about the underlying mechanisms. A simple linear regressor on top of the learned indicators is competitive with deep architectures, providing direct quantitative evidence that the underlying state is encoded in a geometrically accessible form. Code is available at https://anonymous.4open.science/r/LRPTS-9300/README.md.
Step-wise Rubric Rewards for LLM Reasoning
Weichu Xie ⋅ Haozhe Zhao ⋅ Wenpu Liu ⋅ Yongfu Zhu ⋅ Liang Chen ⋅ Minghao Ye ⋅ Zirong Chen ⋅ Yuqi Xu ⋅ Shuai Dong ⋅ Ziyue Wang ⋅ Xinbo Xu ⋅ Kean Shi ⋅ Ruoyu Wu ⋅ Xiaoying Zhang ⋅ Wenqi Shao ⋅ Baobao Chang ⋅ Nan Duan ⋅ Jiaqi Wang
Reinforcement Learning with Verifiable Rewards (RLVR) is widely used to improve the reasoning capabilities of large language models, but its reward is derived only from the correctness of the final answer and provides no supervision over intermediate reasoning steps. Recent rubric-based methods, such as Rubrics as Rewards (RaR), introduce finer-grained supervision by scoring rollouts against a structured set of evaluation criteria. However, the resulting rubric scores are still aggregated into a single scalar that is applied to the entire response, leading to three structural weaknesses, namely loss of the multi-criterion rubric structure, uniform supervision of correct and incorrect reasoning steps, and reward hacking in the trained model through unbounded self-correction. On a sample of 1{,}000 problems, we find that 18.2\% of steps within answer-correct responses are themselves wrong yet positively rewarded, while 49.9\% of steps within answer-incorrect responses are in fact correct yet penalized. Therefore, we introduce \textbf{Step-wise Rubrics as Rewards (SRaR)}, an RLVR framework that (i) uses an LLM judge to attribute each rubric item to a specific reasoning step, (ii) normalizes the per-step rubric scores across rollouts so that only steps whose quality varies produce a learning signal, and (iii) combines the resulting per-step reward with the standard outcome reward through a decoupled advantage estimator that keeps the outcome-driven baseline stable. To support training, we further build a 16K-problem rubric dataset by contrastively distilling rubric items from correct and flawed reasoning paths sampled from a strong model and verified against the ground-truth answer. Across six mathematical reasoning benchmarks spanning multiple difficulty levels, SRaR improves the average accuracy over RaR by \textbf{3.57} points on Qwen3-8B-Non-Thinking and by \textbf{2.75} points on Qwen3-32B-Non-Thinking, raises the Faithful Reasoning Rate on AIME~2025 from 34.5\% to 46.7\% (every correct-answer trajectory of SRaR uses entirely correct reasoning steps), and reduces the rate of self-correction looping from 48.1\% to 26.5\%.
STILL: Selecting Tokens for Intra-Layer Hybrid Attention to Linearize LLMs
Weikang Meng ⋅ Liangyu Huo ⋅ Yadan Luo ⋅ Jiawen Guan ⋅ Jingyi Zhang ⋅ Yingjian Li ⋅ Zheng Zhang
While alleviating the quadratic complexity of Softmax attention is crucial, training large linear models from scratch is computationally prohibitive. Thus, linearizing pretrained LLMs via intra-layer hybrid architectures has emerged as the indispensable paradigm. Existing methods perform token routing based on sliding-window partitions, resulting in position-based selection and fails to capture token-specific global importance. Meanwhile, linear attention further suffers from distribution shift caused by learnable feature maps that distort pretrained feature magnitudes. Motivated by these limitations, we propose STILL, an intra-layer hybrid linearization framework for efficiently linearizing LLMs. STILL introduces a Self-Saliency Score with strong local–global consistency, enabling accurate token selection using sliding-window computation, and retains salient tokens for sparse softmax attention while summarizing the remaining context via linear attention. To preserve pretrained representations, we design a Norm-Preserved Feature Map (NP-Map) that decouples feature direction from magnitude and reinjects pretrained norms. We further adopt a unified training–inference architecture with chunk-wise parallelization and delayed selection to improve hardware efficiency. Experiments show that STILL matches or surpasses the original pretrained model on commonsense and general reasoning tasks, and achieves up to a 86.2\% relative improvement over prior linearized attention methods on long-context benchmarks. Source code can be found in the supplementary materials.
STREAM: Stochastic Riemannian Flow Matching with Anisotropic Decoder for Digital Histopathology Image Generation
Won June Cho ⋅ Daeky Jeong ⋅ Hyeongyeol Lim ⋅ Hongjun Yoon
Synthetic histopathology image generation addresses critical challenges in computational pathology, including patient privacy and the growing need for large-scale training data for foundation models. Latent diffusion models have dominated the image generation domain, with recent works emphasizing that the choice of latent space is critical to the quality of generated images. Existing state-of-the-art generative models in histopathology use pretrained Vision Foundation Models (VFMs) as conditioning signals, and we observe that this leads to ``conditioning collapse'', where the conditioning signal dominates the latent space and lowers the quality and diversity of generated samples. Therefore, we instead use pretrained histopathology VFMs as the latent space itself, leveraging their patch-token features that encode rich semantic information. We empirically show that these features are $\ell_2$-normalized and lie on the unit hypersphere $\mathcal{S}^{d-1}$ with strong angular dominance and intrinsic curvature, making them naturally suited for a Riemannian formulation. We therefore present STREAM, the first framework to apply Riemannian flow matching in the pathology domain. STREAM consists of two stages: 1) a bridge-type stochastic perturbation that establishes per-token rectifiability on $\mathcal{S}^{d-1}$ for training a Diffusion Transformer (DiT) in latent space, and 2) a novel anisotropic decoder that allocates robustness to data-sparse directions while preserving fidelity along data-dense ones. Together, STREAM achieves state-of-the-art reconstruction and generation performance on breast and colorectal cancer datasets.
STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning
Junjie Zhang ⋅ Guozheng Ma ⋅ Shunyu Liu ⋅ Zetian Hu ⋅ Yongcheng Jing ⋅ Ting-En Lin ⋅ Yongbin Li ⋅ Dacheng Tao
Recent advances in Reinforcement Learning (RL) have underscored its potential for incentivizing reasoning capabilities of Large Language Models (LLMs). However, existing step-level efforts suffer from costly annotations that limit domain coverage, while scalar scores further impose an information bottleneck, offering insufficient semantic bandwidth to improve intermediate decisions. Alternative language-critique approaches, which rely on frozen or external critics, provide richer textual feedback but lack the scalability needed for sustained policy improvement. In this work, we propose language-driven stepwise trajectory redirection, termed as STRIDE, a novel training framework that shifts process supervision from scalar rewards to learnable stepwise language feedback. Specifically, we co-train a generator and a generative verifier using only outcome-based rewards, eliminating external annotations, while delivering sustained policy improvement through jointly aligned verifier training. The verifier's stepwise language critiques explicitly localize and explain failures, enabling the generator to redirect reasoning trajectories at intermediate steps toward alternative decisions. The trajectory redirection design guarantees harmless policy improvement, even under noisy or suboptimal verifier feedback. Experiments on diverse reasoning benchmarks show that STRIDE significantly outperforms state-of-the-art baselines, as well as achieving breakthroughs on zero-pass-rate problems where scalar methods yield no learning signal in our ablation studies, demonstrating the effectiveness of learnable stepwise language feedback for enhancing LLM reasoning.
Strong Helps Weak: Directional Cross-Modal Alignment Transfer in Multi-modal LLMs
Hoigi Seo ⋅ Byung Hyun Lee ⋅ Minjun Kim ⋅ Dohyun Mah ⋅ Jongho Lee ⋅ Se Young Chun
Multi-modal large language models (MLLMs) achieve strong modality understanding by pairing a large language model (LLM) with an encoder for a target modality such as vision, video, or audio. However, improving an MLLM's capability for a given modality typically requires additional training on large modality-specific datasets, incurring substantial data collection and compute costs. Model merging offers an alternative, but it is often infeasible for data-scarce, large per-sample size, or domain-specific modalities (e.g., audio and video), where same-modality model variants are rarely available. In this work, we characterize an intriguing asymmetric phenomenon: merging a well-aligned, data-rich source-modality MLLM into a data-scarce target-modality MLLM substantially improves the target on its own benchmarks. Our theoretical and empirical analyses show that this gain stems from enhanced alignment between modality-specific and textual tokens, induced by the stronger donor modality. Specifically, we derive a mutual-information lower bound that is monotonic in alignment-related quantities and strongly correlated with downstream MLLM performance. Building on this principle, we propose Directional Cross-Modal Alignment Transfer (DCAT), a novel framework that transfers textual alignment from a strong, well-aligned source (donor) modality to a weak target (recipient) modality, boosting target-modality performance without further fine-tuning. We further show that the alignment-enhancing objective admits a closed-form weight-space solution computed from only a small calibration set. DCAT outperforms existing model-merging methods, offering an efficient path toward cross-modal alignment transfer.
Strong Stochastic Flow Maps
Sam McCallum ⋅ Zander Blasingame ⋅ Timothy Herschell ⋅ Niklas Rindtorff ⋅ Alexander Tong ⋅ James Foster
Flow and diffusion models generate high-quality samples in many modalities; however, many network evaluations are required during inference due to numerical integration of an underlying differential equation. Flow maps alleviate this problem by learning the solution map of the differential equation directly, enabling few-step sampling. Yet, current methods are restricted to approximating the solution map of ODEs. These methods can be used to learn the transition kernel of an SDE, thereby obtaining a solution map that recovers the marginal distributions of the process (weak convergence) rather than the solution path (strong convergence). We propose Strong Stochastic Flow Maps (SSFMs) as a novel framework for learning the strong solution map of additive-noise SDEs, directly generalizing deterministic flow maps to the stochastic setting. A polynomial approximation to Brownian motion is introduced and shown to converge pathwise. These results enable a simulation-free training objective for the solution map of diffusion models. We demonstrate that SSFMs outperform previous flow map methods on image generation and enable few-step sampling of molecular systems.
Structural Support Certificates for Mechanistic Hypothesis Selection
Max Ruiz Luyten ⋅ Mihaela van der Schaar
Modern scientific discovery systems can propose more mechanistic hypotheses than laboratories can test. We study the downstream selection problem: after upstream retrieval, extraction, expert modeling, or curation has fixed a finite locally factorized posterior over admissible mechanisms, when does the posterior evidence justify acting on one structural hypothesis rather than its rivals? For a query $A\rightsquigarrow B$, we define *structural support* as the posterior mass of mechanisms in which $A$ reaches $B$ by a directed path. This is not a retrieval rank, direct-edge score, or MAP graph decision: support can be distributed across many compatible mechanisms. Our main result is a local certificate for this global support quantity. For path cutoff $L$ and locality radius $m$, truncated support is a convex combination of exact boundary-conditioned local supports. Optimizing over locally admissible blanket states gives a boundary range $[\\ell_{L,m},u_{L,m}]$, and a certified tail bound $\bar r_L$ accounts for longer witnesses: $p_{AB}\in[\\ell_{L,m},\min\{1,u_{L,m}+\bar r_L\}]$. For coupled rivals, the same boundary-conditioned calculations yield a joint feasible set rather than only independent coordinate intervals. We convert these certificates into resolve-or-abstain decisions: an action is resolved only when it wins throughout the certified support set. Under valid certificates, a resolved action is the unique strict utility maximizer inside the declared fixed model, action set, and utilities; otherwise, the solver abstains and reports whether boundary, tail, computation, evidence, or outer-model uncertainty blocks resolution. On controlled and biologically scaffolded finite-posterior records, the implementation has zero endpoint containment misses and zero false fixed-model resolutions. In a coupled-rival diagnostic, joint feasible sets resolve $14/24$ cases versus $6/24$ for scalar boxes, with no false commitments.
Structured State-Space Regularization for Generation-Friendly Image Tokenization
Jinsung Lee ⋅ Jaemin Oh ⋅ Namhun Kim ⋅ Dongwon Kim ⋅ Byung-Jun Yoon ⋅ Suha Kwak
Image tokenizers play a central role in modern generative models, where the structure of the latent space critically determines the downstream generation performance. A key but underexplored property of effective latent representations is spectral organization: the ability to encode information across frequency components. In this work, we introduce structured state-space regularization, a principled approach for inducing spectral structure in latent spaces. We derive a regularization objective by revisiting state-space models (SSMs) as dynamical systems over compressed representations. This perspective reveals that hidden states of SSMs follow basis function dynamics under predefined input transformations, resulting in a novel regularizer form that enforces the latent space to capture spectral components of images. Experiments demonstrate that our regularizer improves the generative performance of image tokenizers while incurring only minimal loss in their reconstruction fidelity.
Structure-Prompted Multimodal Protein Language Model for Preference-Aligned Fitness Prediction
Xiaowen Hu ⋅ Yichao Cao ⋅ Hongyi Huang ⋅ Shan You ⋅ Hao Sun ⋅ Zhenggang Wang ⋅ Xiu Su ⋅ Lei Deng
Protein language models (PLMs) provide valuable evolutionary priors for fitness estimation, but may be misaligned with empirical measurements due to context-dependent selective pressures, potentially leading to hallucination-like predictions under limited experimental supervision. In this study, we introduce MmProt, a structure-prompted Multimodal Protein Language Model for preference-aligned fitness prediction with limited experimentally grounded supervision. MmProt adopts a two-stage multimodal learning strategy: contrastive pre-training on 40 million sequence--structure pairs to learn informative multimodal representations, followed by structure-prompted masked language modeling on naturally occurring related variants to capture protein-specific evolutionary constraints. When limited fitness measurements are available, MmProt further calibrates its variant preferences by encouraging the model to rank high-fitness variants above low-fitness variants. Extensive experiments across multiple benchmarks show that MmProt: (i) achieves state-of-the-art average performance on fitness prediction across 11 diverse protein targets under two challenging extrapolation settings, improving over the best-performing PLM baselines by 3.2% and 13.5% on average, respectively; (ii) captures protein-specific evolutionary patterns in a case study of the SARS-CoV-2 Spike protein, where its zero-shot predictions outperform leading PLMs fine-tuned on 512 fitness-labeled variants by a relative improvement of 3.2%; and (iii) delivers competitive results on eight downstream protein representation learning tasks. Together, these results highlight MmProt's potential as a practical tool for investigating protein evolution and supporting broader applications in protein modeling. Code and data are available at https://anonymous.4open.science/r/MmProt-239E/.
Structure-Semantic Co-optimized Latent Diffusion Model for Fast Visual Anagram Synthesis
Xiang Gao ⋅ Yunpeng Jia
Visual anagram is an intriguing form of art creation wherein a single image presents different conceptual interpretations under transformations such as flipping or rotation. Recent work has achieved visual anagram synthesis by leveraging pretrained text-to-image (T2I) diffusion models, yet still suffers from several key limitations including computational inefficiency, suboptimal aesthetic quality, and weak semantic fidelity and expressiveness. This work focuses on generating visual anagrams with substantially improved visual quality at minimal computational cost, thereby advancing intelligent creation of illusionary digital art. To increase image resolution while reducing time overhead, we adapt the cutting-edge parallel denoising algorithm from pixel-based T2I model to the adversarially distilled latent-based one, and accordingly propose a structure-semantic co-optimization (S2CO) framework to counteract the consequent visual degradation. As the core of our approach, S2CO framework comprises three key innovations: (\romannumeral1) null-text structure alignment optimization; (\romannumeral2) semantic enhancement optimization; (\romannumeral3) attention-guided noise fusion. Building upon these components, our method dubbed \textbf{S2CO-Anagram} is able to generate higher-resolution anagram images with noticeably superior visual harmony and semantic faithfulness than related SOTA approaches, all while achieving substantially faster inference speed. Code will be publicly available.
Structure, Subspace and System: Push the Real Limit of Extremely Low-Bit Quantization for MoE-LLMs
Jiaqi Zhao ⋅ Zichen Li ⋅ Miao Zhang ⋅ Yixuan Dong ⋅ Weili Guan ⋅ Liqiang Nie
Mixture-of-Experts (MoE) large language models (LLMs) suffer significant inference and storage overheads, which makes extremely low-bit quantization (sub 2-bit) highly desirable. However, existing methods for dense LLMs typically quantize each expert independently and require it to approximate the original mapping over the entire input space. We argue that this formulation is overly conservative and structurally mismatched for MoE: 1) experts in the same MoE layer are not fully independent but operate on a shared hidden representation space. Under low-bit budgets, independently binarizing each expert will distort the activation-sensitive input geometry, which leads to severe performance degradation, and 2) due to routing-induced specialization, each expert only serves a concentrated subset of tokens, whose activations typically occupy a much narrower subspace than that of dense layers. More importantly, when the weights approach 1-bit, the dominant deployment bottleneck is no longer the weights themselves, but the metadata (e.g., scaling factors, bitmap, grouping index) which can raise the real effective inference precision to around 3-bit. This issue, however, has largely been overlooked in prior works. Building on these perspectives, we introduce a novel extremely low-bit quantization framework for MoE-LLMs, called TriS-MoE, which considers Structure, Subspace and System. Specifically, we first extract a shared high-precision input-side backbone to preserve activation-sensitive input geometry, while binarizing the remaining components to save costs. Secondly, we propose to redirect expert quantization errors into the null space of routed activations to minimize output perturbation on the subspace actually used by each expert. Finally, we co-design a bitmap compression method based on Golomb-Rice coding which approaches the Shannon entropy lower bound and a specialized streaming loading mechanism to reduce metadata overhead during inference. Extensive experiments on MoE-LLMs demonstrate our TriS-MoE outperforms the strongest baseline by 8.31% average accuracies, and also achieves an average 6.32$\times$ reduction in inference memory on real systems, enabling Qwen3.5-35B-A3B and Mixtral-8$\times$7B to run on a single consumer-GPU.
Zero-shot coordination requires agents to collaborate with previously unseen partners without test-time fine-tuning or explicit communication. A persistent difficulty is that partner behavior contains information at multiple timescales: stable conventions such as role preference or spatial habit, and fast local intent such as the next object to collect or deliver. Most partner-conditioned policies compress these signals into a single latent embedding, which can blur long-term convention with short-term action evidence. We propose \textbf{StylePlan}, a partner modeling framework that separates slow partner style from fast partner intent and combines them through conflict-aware arbitration. StylePlan builds an online behavioral fingerprint, predicts short-horizon partner intent, forms a style-conditioned intent prior, and uses the disagreement between prior and immediate evidence to modulate a recurrent actor-critic policy. On the verified Overcooked V2 artifacts, StylePlan obtains the best reported average cross-play among methods with available entries under the legacy evaluator. We additionally provide a scan-based JAX evaluation with unseen-partner XP, within-episode partner switches, unified SP/FCP reruns, and wired module ablations. The results show that StylePlan is most reliable as an online adaptation mechanism: it improves recovery and drop under partner switches, while the current implementation still needs broader unified reruns to establish a stronger XP claim.
Subliminal Learning Is Steering Vector Distillation
Camila Blank ⋅ Agam Bhatia ⋅ Senthooran Rajamanoharan ⋅ Arthur Conmy ⋅ Neel Nanda
Subliminal learning refers to a student language model acquiring a teacher's traits (e.g. a system-prompted preference for owls) when fine-tuned on the teacher's outputs, despite the outputs being semantically unrelated to those traits. It remains poorly understood how data without semantic meaning can transfer specific semantic traits. In this work, we show that subliminal learning is mediated by a single steering vector, i.e. a vector added to the model's activations. Across two open-source models, we find that the teacher's system prompt is well approximated by a steering vector, and that the student's behavior is driven by learning an aligned vector over fine-tuning. System prompts that are not well approximated by steering vectors are not subliminally learned. This is a special case of steering vector distillation, in which a student trained on the outputs of a steered teacher learns to imitate that steering. We demonstrate steering vector distillation on a range of semantic and random vectors. Adding a semantic vector to a model's activations can have both model-independent and model-specific (i.e. non-semantic) effects on its behavior, so generated data that is non-semantic can transmit a vector with semantic effects, enabling subliminal learning. This also explains why subliminal learning does not transfer between models. We find that adaptive optimizers are necessary for subliminal learning in language models: activation gradients on steered data carry a small but consistent component along the steering direction, and non-adaptive optimizers impede this by allowing outlier gradients to dominate.
Subspace-Guided Continual Learning: Hessian Based Stable–Plastic Decomposition for Exemplar-Free Class-Incremental Learning
Qi Zhu ⋅ Ziang Gan ⋅ Wanting Zhang ⋅ Libao Zhang
Exemplar-Free Class-Incremental Learning (EFCIL) is a challenging continual learning paradigm where a model must learn new classes sequentially without access to old data, making it susceptible to catastrophic forgetting. The core difficulty lies in balancing stability (preserving old knowledge) and plasticity (acquiring new knowledge). We propose Subspace-Guided Continual Learning (SGCL), a novel method that tackles this dilemma from a geometric perspective. SGCL decomposes the feature space into two orthogonal subspaces: a stable subspace containing directions critical for previous tasks, and a plastic subspace where new knowledge can be learned with minimal interference. This decomposition is efficiently identified via the feature-space Hessian, where high-curvature eigendirections define the stable subspace. Building on this, SGCL introduces two synergistic components: 1) Subspace-Guided Regularization (SGR), which imposes curvature-weighted penalties on feature drifts within the stable subspace, and 2) Subspace-Guided Prototype Alignment (SGPA), which adaptively corrects the shift of old-class prototypes to recalibrate the classifier. Extensive experiments on standard benchmarks show that SGCL consistently achieves competitive or superior performance compared to existing state-of-the-art methods, offering a principled approach to mitigating forgetting through loss landscape analysis.
SUPERVISE: A Unified Framework for Standardized and Reproducible Superpixel Evaluation
Julien Walther ⋅ Rémi Giraud ⋅ Michaël Clément
Image segmentation into superpixels is a widely used technique in computer vision, with a large body of work developed over the years. Yet, its evaluation remains poorly standardized, as methods are compared under heterogeneous protocols, across inconsistent scale ranges, and with redundant or misaligned metric subsets. In this work, we introduce $\textbf{SUPERVISE}$ (SUperpixel PERformance VISualization and Evaluation), a unified and reproducible benchmark that addresses these longstanding inconsistencies. Built on a scale-normalized protocol based on interpolation over the actual number of generated superpixels, SUPERVISE enables fair and consistent cross-method comparisons. We leverage this standardized framework to conduct the first large-scale statistical analysis of 20 evaluation metrics across 7 datasets and 31 methods — revealing strong inter-metric correlations and significant redundancy that challenge common evaluation practices. These findings provide empirical grounding for a compact, representative metric subset. SUPERVISE is released as a lightweight, modular framework built on precomputed shared label maps, making large-scale reproducible evaluation accessible to the community. All code, label maps, and evaluation results are publicly available at: https://anonymous.4open.science/r/evaluation_superpixel-08B1/
Sustainability in the Loop: AI Model Development Should Be Multi-Objective
Matteo Mugnai ⋅ Francesco Pistolesi
This position paper argues that environmental cost should be part of the objectives used to train, select, and deploy AI models. Today, sustainability is mainly measured after design choices have been made, while architectures, data, hyperparameters, and serving strategies are chosen through accuracy, latency, and other proxies. This separation weakens model development. In training, newer models can move toward higher emissions without a comparable gain in performance. In inference, small differences in energy per query can become large lifecycle costs at scale. FLOPs, parameter counts, and latency do not capture these effects, because the real cost depends on hardware, datacenter efficiency, energy mix, water demand, and deployment volume. We call for multiobjective model development where performance and lifecycle environmental cost guide decisions together. The paper explains how to distinguish sustainable progress from costlier forms of progress, addresses alternative views, and outlines research directions for objectives, benchmarks, and model selection practices that make sustainability part of optimization.
SWE-GPU-Bench: Can Language Models Solve Real-World GPU Software Engineering Tasks?
Feng Chen ⋅ binbin liu ⋅ Wenhan Han ⋅ Yin Zheng
Existing benchmarks for LLM-based GPU programming have made substantial progress in evaluating whether models can generate correct and efficient kernels, operators, or standalone CUDA programs. However, real GPU software engineering often requires modifying existing repositories rather than writing isolated computational units: changes may span CUDA kernels, C++ host logic, Python bindings, build configurations, tests, and performance evaluation scripts. We introduce SWE-GPU-Bench, a benchmark for repository-level GPU software engineering. To the best of our knowledge, SWE-GPU-Bench is the first benchmark that jointly evaluates PR-derived GPU bug fixes, feature implementations, and performance optimizations across multiple real-world repositories and programming languages, while making setup commands, correctness tests, and optional performance evaluation commands explicit components of each instance. SWE-GPU-Bench contains 608 instances from 23 GPU-related repositories and exposes strong cross-file, cross-language, and long-context demands. Through a comprehensive evaluation of representative methods spanning both file-level and repo-level approaches, we show that current LLMs remain far from effectively solving repository-level GPU software engineering tasks.
SwiftVLM: Efficient Vision-Language Model Inference via Cross-Layer Token Bypass
Chen Qian ⋅ Xinran Yu ⋅ Xinan Wang ⋅ Danyang Li ⋅ Guoxuan Chi ⋅ Zheng Yang ⋅ Qiang Ma ⋅ Xin Miao
Visual token pruning is a promising approach for reducing the computational cost of vision–language models (VLMs), and existing methods often rely on early pruning decisions to improve efficiency. While effective on coarse-grained reasoning tasks, they suffer from significant performance degradation on tasks requiring fine-grained visual details. Through layer-wise analysis, we reveal substantial discrepancies in visual token importance across layers, showing that tokens deemed unimportant at shallow layers can later become highly relevant for text-conditioned reasoning. To avoid irreversible critical information loss caused by premature pruning, we introduce a new pruning paradigm, termed bypass, which preserves unselected visual tokens and forwards them to subsequent pruning stages for re-evaluation. Building on this paradigm, we propose SwiftVLM, a simple and training-free method that performs pruning at model-specific layers with strong visual token selection capability, while enabling independent pruning decisions across layers. Experiments across multiple VLMs and benchmarks demonstrate that SwiftVLM consistently outperforms existing pruning strategies, achieving superior accuracy–efficiency trade-offs and more faithful visual token selection behavior.
SymDrift: One-Shot Generative Modeling under Symmetries
Samir Darouich ⋅ Vinh Tong ⋅ Lluis Pastor Perez ⋅ Tanja Bien ⋅ Loay Mualem ⋅ Mathias Niepert
Generative modeling of physical systems, such as molecules, requires learning distributions that are invariant under global symmetries, such as rotations in three-dimensional space. Equivariant diffusion and flow matching models can incorporate such invariances effectively, even when trained on a non-invariant empirical distribution, but they typically rely on costly multi-step sampling. Recently, drifting models have emerged as an efficient alternative, enabling single-step generation and achieving state-of-the-art performance in generative modeling tasks. However, we show that drifting models face a symmetry-specific challenge, since an equivariant generator does not generally produce the same drifting field as the one obtained from the symmetrized target distribution. Addressing this issue would require expensive symmetrization of the empirical distribution. To avoid this cost, we propose SymDrift, a framework that makes the drifting field itself symmetry-aware. We introduce two complementary strategies: (i) a symmetrized drift in coordinate space based on optimal alignment, and (ii) a $G$-invariant embedding that removes symmetry ambiguity by construction. Empirically, SymDrift outperforms existing one-shot methods on standard benchmarks for conformer and transition state generation, while remaining competitive with significantly more expensive multi-step approaches. By enabling one-shot inference, SymDrift reduces computational overhead by up to 40$\times$ compared to existing baselines, making it promising for high-throughput applications such as virtual drug screening and large-scale reaction network exploration.
SynBench: A Benchmark for Differentially Private Text Generation
Yidan Sun ⋅ Viktor Schlegel ⋅ Srinivasan Nandakumar ⋅ Iqra Zahid ⋅ Yuping Wu ⋅ Yulong Wu ⋅ Hao Li ⋅ Jie Zhang ⋅ Warren Del-Pinto ⋅ Goran Nenadic ⋅ Siew Kei Lam ⋅ Anil A Bharath
Synthetic text generation with Differential Privacy (DP) guarantees emerges as a principled approach that can enable the sharing of sensitive datasets across institutional and regulatory boundaries, while bounding the risks of re-identification and membership inference. LLM-based methods deliver promising results; however, comparisons are exacerbated by differing evaluation setups and "private" datasets, potential pre-training contamination is not considered and guarantees are not verified with DP audits. To advance this field, we introduce a unified evaluation framework with standardised utility and fidelity metrics and privacy audits, encompassing nine curated datasets that capture domain-specific complexities such as technical jargon, long-context dependencies, and specialised document structures. In a large-scale empirical study, we benchmark LLM-based state-of-the-art DP text generators of varying sizes (between 1--8B). Our results indicate that DP synthetic text generation remains an unsolved challenge, with quality deteriorating more as the private datasets deviate further from the generators' pre-training corpora. Our novel synthetic text membership inference attack (MIA) explains this observation: Synthetic data quality is overestimated when LLMs have been pre-trained---without DP---on portions of the "private" data to be generated. Finally, our work provides the first quantitative evidence that this "public pre-training and private generation" paradigm invalidates the guaranteed privacy bounds of real-world private datasets.
SynerVLA: Exploiting Embodied Execution Phases for On-Device Dual-System VLA Acceleration
Qi Lu ⋅ Haotian Xiong ⋅ Ziyu Gong ⋅ TIANJUN SHI ⋅ Lei Xie ⋅ Cheng-Zhong Xu ⋅ Li Li
Dual-system visual-language-action models, which integrate high-level planning (System 2) with instant control (System 1), are promising for embodied AI, but their deployment is hindered by the high computational cost of processing continuous visual streams. Existing acceleration methods are one-sided, focusing only on optimizing the VLM (System 2), which not only limits speed gains but also risks compromising the deep reasoning capabilities it is meant to provide. In this paper, we introduce SynerVLA, a plug-and-play framework for accelerating dual-system VLA models. Its core idea is to leverage the distinct phase information in embodied execution, rapid approaching and fine-grained manipulation, to exploit both spatial and temporal redundancy in visual streams throughout the entire execution pipeline. SynerVLA accelerates dual-system VLA inference by first selecting key visual tokens via text-action fusion, then dynamically adjusting the token pruning-reuse ratio through phase-aware feedback, and finally accelerating both System-2 VLM and System-1 diffusion Transformer with a dual cache reuse mechanism. Evaluations on representative platforms demonstrate that SynerVLA delivers up to a 1.93× speedup and a 26\% higher control frequency, with only a negligible impact on task success rate.
SynGeo: Synergizing Seeing and Proving through Revisable Geometric States
Tianyi Xu ⋅ Liu Yang ⋅ Wenjun GAO ⋅ Junyu Ou ⋅ Zhe Zhao ⋅ HaiBin Wen ⋅ Maolin Wang ⋅ Ye Wei
Geometry problem solving is a canonical testbed for machine intelligence, requiring systems to interpret diagrams, ground symbolic constraints, and perform rigorous deduction. Yet current approaches expose a persistent gap between seeing and proving: multimodal large language models can flexibly inspect diagrams but may hallucinate unsupported relations, while symbolic solvers provide checkable derivations but are brittle to incomplete or misgrounded formalization. We introduce SynGeo, a state-centric framework for synergizing seeing and proving in geometry problem solving by making the geometric representation revisable during inference. SynGeo first constructs a predicate state from diagram--text evidence and tests it with symbolic reasoning. When proof search fails or stagnates, symbolic diagnostics guide image-grounded revisiting, repairing solver-incompatible predicates before another deductive attempt. When the symbolic route remains unresolved, a complementary MLLM branch provides an image-grounded reasoning path from the problem evidence. On Geometry3K and PGPS9K, SynGeo achieves state-of-the-art performance across both Choice and Completion settings. With GPT-4o as the backbone, it reaches (90.2\%) accuracy in both settings on Geometry3K, and 90.1% Choice accuracy and 88.6% Completion accuracy on PGPS9K. Ablations show that both feedback-guided revisiting and complementary reasoning are necessary, supporting a broader view of geometry problem solving as an adaptive loop between what a system sees, what it formalizes, and what it proves.
T2V-AttnDisrupt: Inducing Hallucinations in LVLMs via Misrouting Visual Evidence Retrieval
Yuran Bian ⋅ Xiaoyan Wang ⋅ Conghui Zheng ⋅ Xiaohan Zhang ⋅ Li Pan
Large vision-language models (LVLMs) ground language generation on visual content through the text-to-image slice of self-attention, which acts as a prompt-conditioned router for selecting visual evidence. To induce hallucinations in LVLMs, existing adversarial attacks largely operate at the two ends of this pipeline, either perturbing the front-end visual representation or optimizing against output-token objectives. This leaves the intermediate evidence-selection step as an underexplored yet low-cost attack surface, since it does not require heavy semantic manipulation of image features or extensive output-token optimization. To exploit this attack surface, Text-to-Visual Attention Disruption (T2V-AttnDisrupt) is proposed as an untargeted adversarial attack that induces hallucination by optimizing the input image to distort the text-to-image attention distribution relative to its clean reference, without relying on output-tokens. Experiments on four open-source LVLMs show that T2V-AttnDisrupt increases hallucination rates on captioning benchmarks and reduces VQA accuracy, while preserving overall response quality. It transfers across surrogate–target pairs and generalizes from a single captioning prompt to unseen VQA questions. Moreover, it remains effective against representative defenses, including encoder-level robustness, alignment-based fine-tuning, decoding-time hallucination mitigation, and attention-level interventions, indicating that text-to-image attention is still insufficiently protected. Mechanism analyses show that restoring the clean attention pattern largely recovers visual faithfulness, even when the adversarial image is kept fixed, indicating that hallucination can be driven by corrupted evidence routing rather than feature-level corruption. This low-cost and largely unprotected routing step calls for defenses that explicitly safeguard prompt-conditioned text-to-image attention.
TabClustPFN: A Prior-Fitted Network for Tabular Data Clustering
Tianqi Zhao ⋅ Guanyang Wang ⋅ Yan Shuo Tan ⋅ Qiong Zhang
Prior-data Fitted Networks (PFNs) have reframed supervised tabular learning as single-pass in-context inference without per-dataset optimization. Extending this paradigm to unsupervised clustering is appealing yet fundamentally more challenging, due to absent supervision, unknown cluster cardinality, and label switching inherent to partition outputs. Existing PFN-based clustering methods address these challenges only partially, either requiring known cardinality as input or relying on unstable label-ordering conventions and overly restrictive synthetic priors. We introduce TabClustPFN, a clustering PFN that resolves all challenges jointly through co-designed prior, objective, and architecture. Our hybrid pretraining prior captures heterogeneous real-tabular geometry; our decoupled partition inference network-cardinality inference network architecture jointly infers cluster assignments and cardinality in a single pass; and our SoftARI training objective is permutation-invariant by construction, eliminating the need for any label-ordering convention. On a 44 curated real-world tabular benchmark, TabClustPFN achieves state-of-the-art clustering performance against classical, deep, and amortized baselines, with runtime comparable to efficient baselines.
Tactile MNIST: Benchmarking Active Tactile Perception
Tim Schneider ⋅ Guillaume Duret ⋅ Cristiana de Farias ⋅ Roberto Calandra ⋅ Liming Chen ⋅ Jan Peters
Tactile perception has the potential to significantly enhance dexterous robotic ma- nipulation by providing rich local information that can complement or substitute for other sensory modalities such as vision. However, tactile sensing is an inher- ently local sensor modality, providing information only at the points of contact. This strict locality necessitates the use of active perception techniques to acquire useful, task-relevant information from manipulated objects or the environment in general. Hence, the agent must actively guide sensors toward regions with more informative or significant features and integrate such information over time in order to understand a scene or complete a task. However, while active perception and tactile sensing methods have received significant attention recently, both fields lack standardized benchmarks. To bridge this gap, we introduce the Tactile MNIST Benchmark Suite, an open-source, Gymnasium-compatible benchmark specifically designed for active tactile perception tasks, including localization, classification, and volume estimation. Our benchmark suite offers diverse simulation scenarios, from simple toy environments all the way to complex tactile perception tasks using vision-based tactile sensors. Furthermore, we also offer a comprehensive dataset comprising 13,500 synthetic 3D MNIST digit models and 153,600 real-world tactile samples collected from 600 3D printed digits. Using this dataset, we train a CycleGAN for realistic tactile simulation rendering. By providing standardized protocols and reproducible evaluation frameworks, our benchmark suite facilitates systematic progress in the fields of tactile sensing and active perception in general.
TACT: Mitigating Overthinking and Overacting in Coding Agents via Activation Steering
Yuan Sui ⋅ Yulin Chen ⋅ Yibo Li ⋅ Xue Jiang ⋅ Yufei He ⋅ Yihong Dong ⋅ Xiaoxin He ⋅ Tianyu Gao ⋅ Bryan Hooi
When language model agents tackle complex software engineering tasks, they often degrade over long trajectories, which we define as *agent drift*. We focus on two recurring failure modes *overthinking* and *overacting*, i.e., where the agent repeatedly reasons over information it already has, and where it issues tool calls without integrating recent observations or acquiring new evidence. In this paper, we introduce **TACT** (**T**hink-**A**ct **C**alibration via activation s**T**eering), to detect and mitigate agent drift in the residual stream before it surfaces as a behavioral failure. In specific, we label trajectory steps as overthinking, overacting, or calibrated, and find that their hidden states can separate linearly along two *drift axes*, pointing from calibrated behavior toward each failure mode (AUC $\approx$ 0.9). To mitigate agent drift, we project each step's activation onto these axes at test time and pull drifted ones back toward the calibrated region. Experiments show that TACT outperforms unsteered baselines across SWE-bench Verified, Terminal-Bench 2.0, and CLAW-Eval, lifting average resolve rate by $+5.8$ pp on Qwen3.5-27B and $+4.8$ pp on Gemma-4-26B-A4B-it while cutting steps-to-resolve by up to $26\%$. These gains frame agent drift as a steerable direction in the residual stream, and position TACT as a viable handle for reliable long-horizon agents.
We study Stackelberg learning in which followers use lower-tail CVaR as a reward utility. In cooperative bandits, where both agents share the CVaR objective, we prove a spectral-risk lower bound that gives $\Omega(\sqrt{K(AB-1)/\tau})$ for CVaR and give a Stackelberg-factorized CVaR-UCB algorithm with matching $\widetilde O\sqrt{ABK/\tau})$ regret up to logarithmic factors. Thus CVaR learning admits a minimax-optimal realizable response certificate in bandits. In Markov games, a risk-neutral leader faces a CVaR-sensitive follower. We introduce a certified follower-response oracle based on shortfall planning and decompose true Stackelberg regret into certified oracle regret and benchmark response certification error. The first term is sublinear without coverage, while true regret requires comparison coverage and response compatibility.
Task-Induced Riemannian Metrics for Vision Transformer Feature Spaces
Andrew Bond ⋅ Ege E Özlü ⋅ Tuna Çimen ⋅ Ilkin U Melanlioglu ⋅ Tolga Birdal ⋅ Erkut Erdem ⋅ Aykut Erdem
Methods that operate on Vision Transformer features almost universally rely on Euclidean distance to judge feature similarity. Yet Euclidean distance weights every direction in feature space equally, ignoring that downstream tasks concentrate their sensitivity on a low-dimensional submanifold. The metric that respects this structure is the pullback Riemannian metric $g(F) = J(F)^\top J(F)$ induced by the task decoder's Jacobian, but materializing it is prohibitively expensive; a single ViT-B/14 probe point requires storing $4\times 10^{10}$ scalars, and per-token scoring provably fails when task-sensitive directions are spread across tokens. We show that whether a low-rank approximation of $g(F)$ is \emph{learnable} at all depends on a precise structural property of the model-decoder pair, which we characterize via a single scalar diagnostic $\kappa_{cap}(r)$, the captured-energy fraction of the top-$r$ singular directions of $J$, computable matrix-free in $\mathcal{O}(m+rq)$ Jacobian-vector products. When the diagnostic licenses it, we train the \emph{Spectral Pullback Network} (SPN) on directions found by randomized power iteration on $J^\top J$, and distill the resulting per-token signal into a 310K-parameter feature-only \emph{conformal head}, never materializing $J$. When the original spectrum is too rich for any practical rank-$r$ budget, a VAE reparameterization of the decoder recovers a latent space in which the same machinery applies, converting intractable dense pipelines into tractable ones. Across DPT, DINOv2, CLIP, and VGGT backbones the diagnostic correctly predicts which architecture succeeds: the conformal head reaches Spearman $\rho=0.998$ against the Jacobian-derived target on DINOv2 CLS, and geometric token pruning yields a 25\% relative improvement over ToMe at aggressive prune ratios on DPT depth. These results suggest that learning \emph{which} directions matter is often more valuable than learning better features.
Teaching Large Language Models When Not to Know: Learning Temporal Critique for Ex-Ante Reasoning
Chenlu Ding ⋅ Jiancan Wu ⋅ Yanchen Luo ⋅ Zheyuan (Frank) Liu ⋅ Yancheng Yuan ⋅ Xiang Wang
Large language models (LLMs) often fail to reason under temporal cutoffs: when prompted to answer from the standpoint of an earlier time, they exploit knowledge that became available only later. We study this failure through the lens of ex-ante reasoning, where a model must rely exclusively on information knowable before a cutoff. Through a systematic analysis of prompt-level interventions, we find that temporal leakage is highly sensitive to cutoff formulation and instruction placement: explicit cutoff statements outperform implicit historical framings, and prefix constraints reduce leakage more effectively than suffix constraints. These findings indicate that prompting can steer models into a temporal frame, but does not endow them with the ability to verify whether a response is temporally admissible. We further argue that supervised fine-tuning is insufficient, since ex-ante correctness is not an intrinsic property of an answer, but a relation between the answer and the cutoff. To address this gap, we propose \textbf{TCFT}, a \underline{T}emporal \underline{C}ritique \underline{F}ine-\underline{T}uning framework that trains models to acquire cutoff-aware temporal verification. Given a query, a cutoff, and a candidate response, TCFT teaches the model to identify post-cutoff leakage, explain temporal boundary violations, and judge temporal admissibility. Experiments with Qwen2.5-7B-Instruct and Qwen2.5-14B-Instruct show that TCFT consistently outperforms prompting and SFT baselines, reducing average leakage by 41.89 and 37.79 percentage points, respectively.
Teaching VLMs What to Say, Not How to Reason: Rethinking Counterfactual Reasoning in Autonomous Driving
Hayeon Oh ⋅ Donghwan Lee
End-to-end autonomous driving systems are known to suffer from shortcut learning and causal confusion, limiting robust generalization under distribution shift. To address this, recent Vision-Language Model (VLM) approaches introduce intermediate linguistic reasoning steps including chain-of-thought, chain-of-causation, and counterfactual (CF) reasoning, and report consistent improvements in closed-loop driving. Among these, CF reasoning has been emphasized as enabling self-reflective, high-level causal understanding, yet whether these gains reflect genuine causal reasoning or language-induced distributional bias remains unclear. In this work, we examine what CF reasoning actually induces in VLM-based autonomous driving using accident scenarios with explicit causal structure and a discrete meta-action space. We observe that CF reasoning consistently shifts the action distribution toward more conservative behaviors, accompanied by increased perplexity and policy entropy, and that this effect persists under reinforcement fine-tuning, suggesting a structural rather than transient artifact. However, visual attention analysis and causal intervention experiments show that CF reasoning does not meaningfully alter visual grounding, and VLMs respond similarly to causal and non-causal perturbations, unlike humans who selectively react to causal factors. Overall, these results suggest that CF reasoning does not induce genuine causal reasoning grounded in visual structure. Instead, it acts as a linguistic mechanism that reshapes the action distribution, teaching models what to say but not how to reason, highlighting a fundamental gap between language-based reasoning and causal understanding in VLMs.
TELEVAL: A Benchmark Designed for Spoken Language Models in Chinese Interactive Scenarios
Zehan Li ⋅ Hongjie Chen ⋅ Qing Wang ⋅ Yuxin Zhang ⋅ Jing Zhou ⋅ Hang Lv ⋅ Mengjie Du ⋅ Yaodong Song ⋅ Jie Lian ⋅ Jian Kang ⋅ Jie Li ⋅ Yongxiang Li
Spoken Language Models (SLMs) are expected to support natural spoken interaction beyond task completion. However, existing SLM benchmarks primarily evaluate semantic correctness in structured settings and provide limited assessment of interactional behavior grounded in acoustic context. To address this gap, we introduce TELEVAL, a large-scale SLM benchmark for Chinese spoken interaction in instruction-free, audio-conditioned settings. TELEVAL evaluates two complementary aspects: (1) Reliable Content Fulfillment, which measures semantic accuracy of SLMs under diverse acoustic and linguistic conditions, and (2) Interactional Appropriateness, which assesses whether models produce natural and appropriate responses by implicitly grounding behavior in auditory cues. Experiments show that while models perform competitively on semantic tasks, their performance degrades under acoustic variability and in interactional settings. We observe consistent degradation from perceptual instability to interactional errors, and further identify a recurring failure pattern, termed the "Caption Trap", where models tend to describe perceived audio signals rather than produce appropriate interactive responses. These results indicate that current SLMs remain insufficiently aligned with the requirements of natural spoken interaction. TELEVAL provides a targeted framework for evaluating and analyzing interactional behavior in SLMs.
Temperature Guidance For Robust Reward Conditioning In Diffusion Planning
Johannes V Busch ⋅ Roberto Calandra
Diffusion planners enable long-horizon planning through a generative process over full trajectories, mitigating compounding errors of autoregressive methods while handling multimodal futures. Reward conditioning via classifier-free guidance (CFG) yields high-performing plans but is brittle with respect to per-task hyperparameter choices, limiting its scalability. Our analysis reveals that guidance performance hinges on careful adaptation to the data manifold and reward distribution, contributing to CFG's hyperparameter fragility. We propose temperature-guided diffusion planning (TGDP), which adapts CFG to self-calibrate to these characteristics through temperature-conditional sample reweighting during training and adaptive guidance at inference. TGDP mitigates the hyperparameter fragility of CFG and matches state-of-the-art performance across standard benchmarks without per-task tuning, thus enabling more robust and practical diffusion-based planning.
Temperature-Regulated Stochastic Sampling for Diffusion-Based High-Quality Molecular Generation
Yuyan Ni ⋅ Haoming Kong ⋅ Bo Qiang ⋅ Yanjing Li ⋅ Xin Hong ⋅ Shikun Feng ⋅ Wen Yan ⋅ Yanyan Lan
Despite the strong performance of diffusion models in molecular generation, the sampling process itself remains largely understudied, with most existing methods adopt off-the-shelf sampling strategies developed for natural images. Yet unlike natural images, molecular distributions are governed by physical laws and are sharply concentrated, making them particularly challenging for standard samplers, which frequently produce invalid or unphysical structures. In this work, we revisit diffusion sampling from a unified stochastic differential equation (SDE) perspective and introduce a general framework parameterized by two interpretable controls: \emph{stochasticity} and \emph{temperature}. Our theoretical analysis reveals that stochasticity accelerates the decay of sampling error, while temperature directly controls distribution sharpness, enabling concentration on physically plausible configurations. Building on these insights, we develop \textsc{TReaSSure} (Temperature-Regulated Stochastic Sampling), a lightweight, training-free strategy that combines stochasticity scheduling with temperature annealing to better capture the sharply concentrated distributions of molecular data. Extensive experiments on small molecule generation, protein structure prediction, and protein design demonstrate that \textsc{TReaSSure} consistently improves generation quality and produces more physically valid structures. Empirical analyses further corroborate our theoretical findings. The method harnesses pretrained models without retraining, underscoring its generality and practical value across diverse molecular domains.
Temporal Prototype Alignment for Frozen-Feature Dataset Distillation
Rui Ding ⋅ Jing Hu ⋅ Mei Chen ⋅ Xi Wu ⋅ Hualin zhou ⋅ Fan Wu ⋅ Kehua Guo ⋅ Tao Gu
Dataset distillation condenses large-scale training sets into compact synthetic data. Recent frozen-feature distillation methods, such as Linear Gradient Matching, make ImageNet-scale synthesis feasible by matching gradients in pretrained representation spaces, but their reliance on exhaustive spatial view ensembles incurs memory and I/O costs that grow linearly with the number of augmentations. We propose Temporal Prototype Alignment (TPA), an online temporal gradient estimation framework for frozen-feature dataset distillation. TPA replaces spatial Monte Carlo gradient averaging with a temporal prototype estimator that accumulates stable semantic signals across optimization steps. An adaptive momentum gate modulates the estimator according to local signal reliability, reducing gradient variance while keeping memory complexity independent of the number of spatial views. To improve cross-architecture transfer, we introduce a consensus alignment objective that distills shared structure from multiple pretrained teachers while mitigating model-specific artifacts. Across ImageNet-scale and fine-grained benchmarks, TPA substantially reduces memory overhead, enables distillation on consumer-grade GPUs, and improves transfer to heterogeneous CNN and Vision Transformer backbones. These results suggest that online temporal estimation provides an efficient alternative to spatial ensembling for frozen-feature dataset distillation.
Test-time Multi-agent Coordination by Decomposed Value Gradient Flow
Dongsu Lee ⋅ Haoran Xu ⋅ Amy Zhang
Recent offline multi-agent reinforcement learning (MARL) faces a persistent trade-off. Expressive generative policies can represent multi-modal coordination in the data but cannot distinguish high-value regions, while value-optimized policies exploit the learned Q-function but collapse the multi-modal structure into a single dominant mode. A single agent's mode collapse can break joint coordination, and simultaneous drift across agents can push the joint policy into unseen regions of the action space. We propose scalable coordination via optimal unified transport (SCOUT), the first offline MARL framework to combine a generative foundation model with a learned value function through test-time training. SCOUT trains two decoupled components: a flow-matching behavioral prior and a decomposed value function. At test-time, it transports behavioral samples toward high-value regions via variational gradient descent. The number of transport steps controls adaptive test-time scaling, replacing a fixed regularization coefficient. Under the individual-global-max (IGM) principle, we prove that decentralized per-agent transport is consistent with joint value improvement and monotonically recovers inter-agent correlation even when the value decomposition is approximate.
Despite the success of Neural Operators (NOs), architectural hyperparameters such as depth are still typically selected by empirical grid search, with limited connection to the governing PDE. We introduce the \textit{alternation depth}, $\alpha_F$, a structural quantity of the PDE right-hand side that counts the minimum number of non-pointwise/pointwise coupling stages needed to build its leading nonlinear spatial structure. We use $L_{\min} = \alpha_F + 1$ as a principled \emph{depth-efficiency lower-bound guideline} for IVP solution maps: a starting depth, not a predictor of the empirical optimum $L^*$. We provide partial formal support: (*sufficiency*) an $L$-block architecture can approximate any operator in a formally defined $L$-alternation class under compactness and regularity assumptions; (*separation*) for a constructed hard family at finite resolution and polynomial activations, using fewer blocks forces width to grow as a power of the discretization size. Empirically, a 736-run study over nine in-scope IVP datasets plus one Darcy stress test, using FNO-family variants and one kernel-integral baseline (KIN), is consistent with depth improvements appearing at or above $L_{\min}$ in the tested settings. The observed optimum can be larger because of spectral approximation, pointwise approximation difficulty, data regime, and optimization. We summarize the resulting practitioner guidance in four capacity-allocation rules.
The Design Space of Tri-Modal Masked Diffusion Models
Louis Bethune ⋅ Victor Guilherme Turrisi da Costa ⋅ Bruno Mlodozeniec ⋅ Pau Rodriguez ⋅ Lokesh Boominathan ⋅ Nikhil Bhendawade ⋅ Amitis Shidani ⋅ Joris Pelemans ⋅ Theo X. Olausson ⋅ R Devon Hjelm ⋅ Paul Dixon ⋅ Joao Monteiro ⋅ Pierre Ablin ⋅ Vishnu Banna ⋅ Arno Blaas ⋅ Nick Henderson ⋅ Kari Noriy ⋅ Dan Busbridge ⋅ Marco Cuturi ⋅ Joshua Susskind ⋅ Irina Belousova ⋅ Luca Zappella ⋅ Russell Webb ⋅ Jason Ramapuram
Discrete diffusion models have emerged as strong alternatives to autoregressive language models, with recent multimodal work either finetuning unimodal diffusion bases or distilling autoregressive backbones for bi-modal generation. Diverging from these approaches, we introduce the first tri-modal \gls{mdm} \emph{pretrained from scratch} on text, image-text, and audio-text data, where all generation tasks are learned simultaneously within a single model. We systematically study the design space governing stability and efficiency at scale: we conduct the first empirical characterization of the critical batch size $B_{\text{crit}}$ under the SDE reparameterization of AdamW, and introduce a drift--horizon interpolation parameter $\gamma$ that balances gradient noise reduction and optimization horizon when scaling the token budget. We derive multimodal scaling laws and ablate modality mixing ratios, noise schedules, and anti-masking. Lastly, we pretrain a 3B model on 6.4T tokens, achieving competitive results in text generation, text-to-image, and text-to-speech.
The Dormant Spiking Neuron: A State-Driven Mechanism for Efficient Spiking Neural Networks
Ruichen Ma ⋅ Yicai Chen ⋅ Pujun Zhou ⋅ Yue Zuo ⋅ Chongyang Li ⋅ Shuang Liu ⋅ Shaogang Hu ⋅ Guanchao Qiao
Spiking neural networks (SNNs) offer a compelling pathway toward energy-efficient artificial intelligence, fundamentally anchored in presynaptic input-driven sparsity. However, a substantial yet chronically overlooked source of computational overhead persists: the continuous subthreshold dynamics of postsynaptic neurons. Quantitative analysis reveals that up to 44\% of neurons in an ImageNet-trained ResNet-18 model reside in a deeply inhibited state, imposing a persistent computational burden despite possessing a negligible probability of firing. To eradicate this inefficiency, the Dormant Spiking Neuron (DSN) model is introduced. By incorporating a state-driven dynamic gate, the DSN allows a neuron to enter a dormant state when its membrane potential falls below a predefined negative threshold, dynamically pruning redundant computations at the instance level. Crucially, this dormancy mechanism is directly integrated into the neuronal dynamics and trained end-to-end via backpropagation through time using surrogate gradients, enabling the network to actively learn a robust dormancy policy that adaptively preserves critical information flow. Extensive experiments across static and neuromorphic datasets demonstrate that the DSN framework consistently reduces computational energy by 40\% to over 70\%. Notably, on the large-scale ImageNet dataset, the DSN outperforms state-of-the-art static pruning techniques by achieving equivalent energy savings while simultaneously improving upon the baseline accuracy by 1.34\%. By establishing a trainable, state-driven sparsity paradigm orthogonal to conventional input-driven methods, the DSN unlocks synergistic efficiency gains, presenting a highly potent approach for advancing neuromorphic computing. Source code will be made publicly available.
The Foundation Model Transparency Index
Rishi Bommasani ⋅ Kevin Klyman ⋅ Shayne Longpre ⋅ Sayash Kapoor ⋅ Nestor Maslej ⋅ Betty Xiong ⋅ Daniel Zhang ⋅ Percy Liang
Foundation models have rapidly permeated society, catalyzing a wave of generative AI applications spanning enterprise and consumer-facing contexts. While the societal impact of foundation models is growing, transparency is on the decline, mirroring the opacity that has plagued past digital technologies (e.g. social media). Reversing this trend is essential: transparency is a vital precondition for public accountability, scientific innovation, and effective governance. To assess the transparency of the foundation model ecosystem and help improve transparency over time, we introduce the Foundation Model Transparency Index. The Foundation Model Transparency Index specifies 100 fine-grained indicators that comprehensively codify transparency for foundation models, spanning the upstream resources used to build a foundation model (e.g data, labor, compute), details about the model itself (e.g. size, capabilities, risks), and the downstream use (e.g. distribution channels, usage policies, affected geographies). We score 10 major foundation model developers (e.g. OpenAI, Google, Meta) against the 100 indicators to assess their transparency. To facilitate and standardize assessment, we score developers in relation to their practices for their flagship foundation model (e.g. GPT-4 for OpenAI, PaLM 2 for Google, Llama 2 for Meta). We present 10 top-level findings about the foundation model ecosystem: for example, no developer currently discloses significant information about the downstream impact of its flagship model, such as the number of users, affected market sectors, or how users can seek redress for harm. Overall, the Foundation Model Transparency Index establishes the level of transparency today to drive progress on foundation model governance via industry standards and regulatory intervention.
The Geometric Wall: Manifold Structure Predicts Layerwise Sparse Autoencoder Scaling Laws
Eslam Zaher ⋅ Maciej Trzaskowski ⋅ Quan Nguyen ⋅ Fred Roosta
Sparse autoencoders (SAEs) operationalise the linear representation hypothesis: they reconstruct model activations as sparse linear combinations of interpretable dictionary atoms, on the implicit assumption that activation space is well approximated by a globally linear structure. Their reconstruction error varies sharply across layers in ways that existing scaling laws, fitted at single layers, do not explain. We argue that this variation is the empirical trace of a geometric mismatch: where the activation manifold is curved and its intrinsic dimension varies across layers, no sparse linear dictionary can match it uniformly, and the SAE's width-sparsity scaling becomes a layer-dependent function of manifold structure rather than a single universal law. We conduct the first cross-layer SAE scaling study, fitting and regressing on 844 residual-stream Gemma Scope SAE checkpoints across 68 layers of Gemma 2 2B and 9B. Stage 1 fits a per-layer scaling-law surface; Stage 2 regresses the fitted parameters and the derived per-layer width exponents on four layerwise geometric summaries. We find that manifold geometry predicts the per-layer width exponent in both models, and that the same regression coefficients learnt on one model predict the other model's per-layer exponents under cross-model transfer, indicating a transferable geometric law. At the showcase layers where richer width grids permit identification of the asymptotic floor, we find that the fitted floor tracks the layerwise geometric ordering: higher curvature and intrinsic dimension correspond to higher floor, consistent with the irreducible second-order residual that any sparse linear approximation of a curved manifold must leave behind. SAEs thus encounter not a finite-resource ceiling but a geometry-dependent wall, set by the manifold they are trying to reconstruct.
The Hidden Ratio in Adam: Stable Structure, Compression, and Sign Dynamics
Yihe Zhou ⋅ Tongtian Zhu ⋅ Yingxiao Huo ⋅ Satya P Dash ⋅ Can Wang ⋅ Samuel Kaski ⋅ Mingfei Sun
Adam is the default optimizer for training modern deep neural networks, yet its adaptive behavior remains poorly understood due to the complex interaction between its first- and second-moment exponential moving averages (EMAs). We study Adam in the tied-$\beta$ regime, where the two EMA decay rates are equal, and show that its adaptive dynamics can be expressed through a transformed ratio with approximately scale-stable behavior. Empirically, this transformed ratio exhibits a stable, heavy-tailed distribution across tasks, model scales, and training stages, in contrast to the variability of raw moment magnitudes. This empirical stability enables both practical and conceptual consequences. First, we derive a recurrence for the transformed ratio, yielding a reparameterization of Adam that replaces the second moment with a compressible state. Leveraging its stable distribution, we show that a fixed 4-bit codebook is sufficient in our experiments to store this state without auxiliary scaling, achieving performance competitive with full-precision Adam. Second, the transformed ratio view clarifies Adam's connection to sign-based methods: Adam reduces to sign momentum modulated by the transformed ratio, and replacing it with a constant recovers Signum as a limiting case. This perspective further provides a simple rule for transferring learning rates between the two methods. Together, these results suggest that tied-$\beta$ Adam admits a simple and approximately stable ratio structure underlying its adaptive behavior and demonstrate its utility for both analysis and efficient implementation.
Across scientific disciplines, Laplacian eigenvectors serve as a fundamental basis for simplifying complex systems, from signal processing to quantum mechanics. In reinforcement learning (RL), they similarly form a basis over the state space, enabling reward functions to be approximated by projection onto a small set of eigenvectors. This projection makes zero-shot control possible, but it also imposes a fundamental limitation: the induced policies are only as expressive as the linear span of the chosen eigenvectors. We introduce the Laplacian Keyboard (LK), a hierarchical framework that goes beyond this linear span. LK constructs a task-agnostic library of behaviors from these eigenvectors, forming a behavior basis guaranteed to contain the optimal policy for any reward within the linear span. A meta-policy learns to stitch these behaviors dynamically, enabling efficient learning of policies outside the original linear constraints. We establish theoretical bounds on zero-shot approximation error and demonstrate empirically that LK improves over the zero-shot solution while achieving better sample efficiency compared to standard RL methods.
The Marauder’s Map: Bézier Manifolds Reveal Hidden Surfaces for Model Merging and Ensembling
Abhiram Iyer ⋅ Mark T Harnett ⋅ Sarthak Chandra
Neural circuits are infamously more robust than their deep network counterparts. Biology often explains this through degeneracy, where structurally distinct configurations support the same function. Deep learning studies a parallel phenomenon through mode connectivity, which links same-objective networks by low-loss curves, sheets, or volumes. Yet real biological systems undergo continual structural change, a more flexible form of degeneracy in which specialized circuits gradually adapt toward new competencies through localized changes while preserving existing function. The analogous question in deep learning is harder than ordinary mode connectivity, asking whether many independently initialized models trained on distinct tasks can be connected by a high-dimensional region of competent multi-task solutions. We introduce Nimbus, which learns a high-dimensional Bézier manifold connecting N specialists, going beyond the one-dimensional paths and two-dimensional sheets of prior mode connectivity work. A naive parameterization would scale combinatorially with N, but Nimbus requires only a small number of learned correction tensors shared per interaction order, showing that the connecting manifold has surprisingly low complexity yet supports multi-task competencies. The resulting manifold from Nimbus forms a navigable "Marauder's Map" connecting independently trained specialists through a shared geometry of competent solutions. Across vision (CIFAR-100) and language modeling (five disjoint text domains), Nimbus produces manifolds that are both functionally and locally weight-flat. Ensembles sampled from them degrade more gracefully than baselines under input and weight perturbations, with the largest gains in negative log-likelihood and expected calibration error. Beyond its empirical behavior, Nimbus supplies a deep learning interpretation of flexible degeneracy, recasting the connectivity between specialized circuits from a low-dimensional path between distant solutions into a high-dimensional volume of nearby ones.
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning
Jing Liang ⋅ Hongyao Tang ⋅ Yi Ma ⋅ Yancheng He ⋅ Weixun Wang ⋅ Xiaoyang Li ⋅ Wenbo Su ⋅ Jinyi Liu ⋅ YAN ZHENG ⋅ Jianye Hao ⋅ Bo Zheng
Reinforcement learning (RL) has gained growing attention in large language model (LLM) post-training, yet RL training remains fragile and can suffer from instability or collapse. One vital cause is \textit{training-inference mismatch}: % different from traditional RL, the engineering design of LLM adopts separate inference and training engines for generation efficiency and training precision, which in practice exhibits inconsistent probabilities for the same trajectories on training and inference sides, even with synchronized model parameters. This naturally induces a special type of off-policyness ever existing and poisoning the training. Prior works have made various efforts in addressing the off-policyness to stabilize the training policies under the mismatch. In this paper, we point out the \textit{object misalignment} neglected by existing works that an effective update to the policy in the training engine not necessarily ensures the improvement of the inference policy, i.e., the one used in deployment. To this end, we propose a new policy optimization objective for LLM RL, named \textbf{Monotonic Inference Policy Improvement (MIPI)}. Following this principle, we introduce {Monotonic Inference Policy Update (MIPU)}, a two-step LLM RL framework that constructs sampler-referenced candidate updates and selectively accepts synchronized candidates using an inference-side gap proxy. Experiments under high mismatch show that {MIPU} improves average reasoning performance and training stability.
The Multiscale Single-Index Model: A Toy Model for Hierarchical Feature Learning
Gordon Dai ⋅ Joan Bruna
We introduce the Multiscale Single-Index Model, a stylized model for deep hierarchical feature learning with scale separation. Each layer extracts a shared single-index feature at one physical scale and passes it to the next, giving a tractable setting in which to study how deep architectures learn local multiscale representations. Under non-degeneracy and delocalization assumptions on the link function and planted features respectively, we prove two complementary recovery guarantees. First, for any fixed depth $K$ and local scale $d$, corresponding to an input of size $d^K$, the first Wiener chaos expansion of the target behaves as a perturbed spiked tensor, where the perturbation comes from the non-linearity; we then leverage this fact to show that a spectral method based on tensor unfolding strongly recovers all planted directions with $O(d^{\lceil K/2\rceil}\log d)$ samples. Second, for the two hidden-layer case, we analyze joint online spherical SGD and show that it achieves weak recovery from random initialization in $O(d\log^2 d)$ samples, followed by strong recovery to accuracy $\varepsilon$ in $O(d\log(1/\varepsilon))$ additional samples. The main technical challenge is that the second layer observes non-Gaussian learned features; we overcome this through a quantitative Gaussian comparison using delocalization, combined with sharp trajectory-level control of the coupled SGD dynamics.
The Override Gap: A Magnitude Account of Knowledge Conflict Failure in Hypernetwork-Based Instant LLM Adaptation
Shuaizhi Cheng ⋅ Xiang Shi ⋅ Zhiwei Zhang ⋅ Mingwei Li
Hypernetwork-based methods such as Doc-to-LoRA internalize a document into an LLM's weights in a single forward pass, but they fail systematically when the document contradicts pretraining knowledge: accuracy drops to 46.4% on the deepest conflicts. We show that this failure is primarily a magnitude problem rather than a representational one. The generated adapter reaches the relevant layers, but its adapter margin is not calibrated to the strength of the contradicted prior; as the pretrained margin grows, deep conflicts lose the override competition. This account predicts that failure should track prior strength. Sorting 194 conflicts by the base model's log-probability on the contradicted fact, baseline accuracy falls from 68% on weak-prior questions to 16% on strong-prior questions, a 52 percentage-point gap. We propose two training-free corrections. Selective Layer Boosting (SLB) scales the adapter at its highest-activity layers, and Conflict-Aware Internalization (CA) applies stronger boosting only when the base model appears confident. Together they raise deep-conflict accuracy from 46.4% to 71.0% on Gemma-2B and from 53.6% to 72.5% on Mistral-7B, while preserving novel-knowledge recall at 97.1%. They outperform vanilla retrieval-augmented generation on Gemma medium conflicts by 18 percentage points, though explicit conflict-aware prompts remain stronger when the user already knows a conflict exists. We release KID-Bench, a 489-question benchmark that separates novel recall, cross-knowledge combination, and prior-graded conflicts.
The Scaling Paradox of Tool-Calling LLM Agents Under Realistic MCP Faults
Antonio Ken Iannillo ⋅ Joshua S Owotogbe ⋅ Roberto Natella ⋅ Francesco Avallone ⋅ Kristina Kudryavtseva ⋅ Indika Kumara
LLM agents increasingly rely on external tools exposed via the Model Context Protocol (MCP), yet current benchmarks assume tools respond correctly and on time, leaving agent resilience to realistic tool failures untested. We introduce MCP-Injector, a drop-in, protocol-aware fault-injection middleware that interposes on the MCP tool-call path with controlled persistence (permanent, transient, intermittent) and seedable schedules. Its fault model is empirically calibrated by mining 57,240 issue and pull-request threads across 1,774 open-source MCP repositories, yielding a distribution over six failure modes grounded in observed ecosystem incidents. In a 1,350-cell experiment crossing two backbone sizes (20B, 120B), three agent stacks (custom multi-round plan-and-execute, OpenAI Agents SDK, LangChain), and three fault injection conditions, we find that calibrated MCP faults cause severe quality degradation (60\% drop in task completion score, $p < 10^{-56}$) and reveal a scaling paradox: the larger model achieves higher fault-free scores but suffers lower completion rates under faults, driven by longer reasoning chains that amplify fault exposure. Framework-level error handling shapes resilience more than model scale, with stack choice producing significant interaction effects. These results demonstrate that capability benchmarks alone can mask operational brittleness, motivating protocol-aware fault injection as a complementary evaluation dimension for tool-calling agents.
The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility
David Pape ⋅ Jonathan Evertz ⋅ Lea Schönherr
Progress in LLMs is increasingly measured through standardized benchmarks, where state-of-the-art improvements are often separated by fractions of a percentage point. At the same time, the computational cost of evaluating modern LLMs has driven widespread adoption of specialized inference backends, software systems that execute trained models efficiently at inference time. While critical for scalability, system-level optimizations, such as custom CUDA kernels and reduced-precision arithmetic, can alter token probabilities and introduce non-determinism, possibly cascading into divergent generation. In this work, we first survey the inference landscape, identifying 200 distinct engines, and analyze 35,000 ML publications, finding that the specific inference stack is rarely reported despite this widespread diversity. We then present a systematic empirical study of how inference backends affect LLM benchmark results. Holding model weights, decoding parameters, and hardware constant, we evaluate five widely used inference engines, including vLLM, SGLang, and llama.cpp, across multiple open-weight models and established benchmarks. We show that the choice of backend alone can shift benchmark scores by up to 16.6 percentage points and induce high rates of output disagreement. By isolating backend optimizations and tracing the execution pipeline, we find this divergence is driven by system-level optimizations like prefix caching and CUDA graphs, custom kernels, and engine-specific defaults in logit processing. Our findings identify the inference backend as a previously unreported but consequential hyperparameter in the evaluation of LLM and advocate standardized reporting of inference stacks to improve the reproducibility and interpretability of benchmark comparisons.
The Symmetries of Three-Layer ReLU Networks
Johanna Marie Gegenfurtner ⋅ Moritz Grillo ⋅ Guido Montufar
We develop a framework for analyzing parameter symmetries in deep ReLU networks and obtain a complete characterization of the generic parameter fibers for three-layer bottleneck architectures. Our approach provides explicit semi-algebraic descriptions of these fibers and yields a polynomial time algorithm for deciding functional equivalence of two parameters. The symmetries include discrete and continuous transformations arising from layer composition, and depend on whether deeper layers hide or preserve geometric structure from preceding layers. Finally, we show that some of these symmetries induce local conservation laws along gradient flow, while others do not.
The Value of Covariance Matching in Gaussian DDPMs and the Lanczos Sampler
Sahil Akhtar ⋅ Aymane El Gadarri ⋅ Vivek Farias ⋅ Adam Jozefiak
A central error measure in Gaussian DDPMs is the path-space KL divergence between the exact reverse chain and the learned Gaussian reverse process. This quantity is especially relevant for procedures such as classifier guidance, which perturb the entire reverse trajectory rather than only the terminal sample. Prior analyses show that standard isotropic reverse covariances suffer an unavoidable $\Omega(1/T)$ path-KL error as the number of denoising steps $T$ grows. We show that matching the full posterior covariance breaks this barrier, yielding an order-wise improvement that reduces the path KL to $O(1/T^2)$. To make full covariance matching practical, we introduce the Lanczos Gaussian sampler, a training-free, matrix-free method for sampling from the optimal reverse covariance using only covariance-vector products, which are available through Jacobian-vector products of the posterior mean. The sampler avoids dense covariance storage and auxiliary covariance models. We prove that its approximation error decays exponentially in the number of Lanczos steps, where each Lanczos step requires a single Jacobian-vector product. Empirically, using only just three such steps improves sample quality over strong diagonal-covariance baselines, including OCM-DDPM, across standard image benchmarks. This identifies full covariance matching as both theoretically valuable and practically accessible for fast DDPM sampling.
The VLM as Sensor: Bayesian Active Search for Long Video Understanding
Chong Tang ⋅ Sannara EK ⋅ Dirk Koch ⋅ Robert Mullins ⋅ Alex Weddell ⋅ Jagmohan Chauhan
What if the best way to search a video with a vision-language model (VLM) is to stop asking the model to search? Recent approaches give the VLM control over temporal navigation, tying search quality to the model's reasoning ability. We propose the opposite: BeliefSearch treats the VLM as a noisy sensor and delegates navigation to an external Bayesian controller. A single belief over video segments anchors the system: the question sets the prior, each VLM observation updates it, and the same belief decides whether to search, which segment to examine next, and when to stop. Because the belief is explicit, we can trace why each choice was made. The same signal also drives training: each turn earns credit for the uncertainty it reduces, and the final reward is gated by how much the search lowered entropy. This blocks a reward-hacking shortcut where the model otherwise learns to skip search and guess from the preview frames. On long-video benchmarks, our method outperforms all prior search-based methods by up to 8.8 points while using $5$ to $13\times$ fewer frames than recent RL search baselines.
The Web Doesn't Sit Still: Adversarial Self-Evolving Attacks on Search Agents
Geyuan Wang ⋅ Shunyu Liu ⋅ Siyuan Liang ⋅ Chenyang Lyu ⋅ Dacheng Tao
Large Language Models (LLMs) have shown promise for empowering search agents to tackle complex information-seeking tasks. However, since acquiring external information necessitates web search, these agents are highly susceptible to various adversarial attacks (e.g., malicious information injection or stealthy context manipulation). Existing attack strategies inevitably rely on static and human-crafted templates, failing to emulate the dynamic and complex nature of real-world threats. In this paper, we propose an adversarial self-evolution framework that integrates bi-level optimization between the attacker's strategy evolution and the defender's adaptive mitigation, to expose the vulnerabilities of search agents. Specifically, we introduce the Contrastive Rollout Evolutionary Optimization (CREO) method to drive the directional evolution of the attack strategy population via contrastive evolutionary signals. To provide a stable environment for this continuous evolution, we further construct a controlled sandbox and design fine-grained metrics to quantify internal vulnerabilities. Extensive experiments across various multi-hop QA benchmarks and frontier LLMs demonstrate that the proposed framework yields attack efficacy superior to static baselines, underscoring the urgent need to develop more robust defense frameworks. The anonymized code repository is available at https://anonymous.4open.science/r/adversarial_rag-3EFA.
Thinking in Pictures: A Systematic Benchmark for Reasoning-driven Image Generation
Yutong Liu ⋅ Nan Huang ⋅ Xu Cao ⋅ James Rehg
Recent advancements in world models and unified generative models (UGMs) have achieved unprecedented results in visual perception and synthesis. However, these models primarily rely on surface-level alignment, leaving the capacity for high-level visual reasoning underexplored. True visual generative intelligence demands "Reasoning-to-Generation", an ability to infer latent rules from visual inputs and manifest solutions through precise, logically constrained visual outcomes. We introduce RIG-Bench, a novel comprehensive benchmark that systematically evaluates Reasoning-driven Image Generation (RIG) across four cognitively demanding domains: Concept-based, Transformation-based, Pattern & Structure, and Scenario-based. Featuring 2000 curated samples, RIG-Bench serves as a rigorous stress test for RIG. Our extensive evaluations of state-of-the-art UGMs and image/video generation models reveal a significant reasoning-generation gap, wherein models frequently produce locally plausible but globally illogical outputs. RIG-Bench provides a vital diagnostic framework to guide the development of next-generation, logically grounded world models.
Thinking with Imitation: Adaptive Reinforcement-Imitation Learning for Tool-Augmented Scientific Reasoning
Fanrui Zhang ⋅ Hongmin Zhan ⋅ Qiang Zhang ⋅ Sizhuo Zhou ⋅ Zhaoyu Wang ⋅ Kaipeng Zhang ⋅ Jiawei Liu ⋅ Zheng-Jun Zha
Expert scientists rarely solve genuinely novel problems through parametric recall alone; instead, they reason by analogy to structurally related prior cases and adapt their solution procedures to new contexts. Yet current multimodal large language models (MLLMs), even when equipped with retrieval, computation, and visual tools, still struggle with this imitation-driven form of reasoning. We further observe a counterintuitive negative result: simply prepending expert problems and full solution trajectories as context does not improve performance on novel scientific tasks, and can even slightly degrade it. Motivated by this observation, we argue that scientific reasoning should shift from passive answer conditioning to an imitation-driven paradigm, in which models actively seek relevant precedents and reuse their solution structures. We instantiate this idea with MimicAgent, a tool-augmented scientific reasoning agent that retrieves structurally related exemplars, interprets their reasoning trajectories, and transfers their solution patterns to the target problem through multi-step reasoning and tool interaction. To train this behavior under sparse rewards, we introduce ARIS (Adaptive Reinforcement-Imitation Switching), which unifies on-policy reinforcement learning with online expert imitation. When all rollouts for a prompt fail, ARIS replaces degenerate reinforcement updates with expert-generated demonstrations, yielding an implicit imitation-to-reinforcement curriculum without requiring a cold-start stage or fixed-ratio offline mixing. To support training and evaluation, we further introduce SciExplore-Bench, a bilingual multimodal benchmark for open-ended experimental scientific reasoning, paired with a manually curated repository of analogical exemplars. The substantial gains and impressive generalization across multiple benchmarks suggest that imitation serves both as a global paradigm for exemplar-guided reasoning and as a local mechanism for stabilizing reinforcement learning.
Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models
Junhao Liu ⋅ Jian-Wei Zhang ⋅ Tao Huang ⋅ Miles Yang ⋅ Zhao Zhong ⋅ Liefeng Bo
Unified Multimodal Large Language Models (MLLMs) offer a promising paradigm for unifying visual understanding and generation, yet they still struggle to follow complex spatial instructions and logical constraints in controllable image generation. To address this gap, we present ATLAS, a unified framework that equips MLLMs with a human-like ``Think, Plan, and Paint'' paradigm. We adopt layout as the shared representation that connects the three stages, enabling the model to reason about spatial requirements, plan explicit object arrangements, and render the final image. We further improve plan-to-image fidelity with reinforcement-learning-based layout alignment. We instantiate ATLAS at 7B and 80B scales, achieving state-of-the-art performance among MLLMs on image generation benchmarks and an average 65.31\% improvement over existing layout-aware MLLMs. On spatially related tasks, ATLAS obtains a 23.06\% gain over the base models. Through the same layout interface, ATLAS also supports instruction-guided editing and multimodal grounding. We further introduce ATLAS-Reasoning, a benchmark for evaluating generation under complex spatial instructions.
TIC-GRPO: Provable and Efficient Optimization for Reinforcement Learning from Human Feedback
Lei pang ⋅ Jun Luo ⋅ ruinan Jin
Group Relative Policy Optimization (GRPO) is a critic-free reinforcement learning algorithm for fine-tuning large language models, but its token-level importance sampling mechanism creates a subtle variance bottleneck. We separate this bottleneck into two mechanisms. First, the standard GRPO clipping rule is incomplete: on the negative-advantage branch, it can pass uncontrolled upper-tail importance weights. This motivates up-only clipping as a standalone correction. Second, up-only clipping alone is not enough if the update remains token-level: the clipped token-weighted score increments need not be martingale differences, so cross terms survive in the squared update. This motivates trajectory-level importance correction. We propose Trajectory-level Importance-Corrected GRPO (TIC-GRPO), which follows this progression: it first caps upper-tail ratios by up-only clipping and then replaces token-level importance ratios with a single trajectory-level probability ratio. The latter restores the measure-change identity needed for trajectory-level martingale cancellation. Our variance theory gives upper bounds, hard-instance lower bounds, and induced convergence lower bounds that separate original GRPO, the token-level up-clipped comparator GRPO$_2$, and TIC-GRPO. Experiments on math reasoning and coding benchmarks further confirm that TIC-GRPO improves optimization stability and final performance.
TimeOperator: A Function-to-Function Approach to Time Series Modeling
SheoYon Jhin ⋅ B. Aditya Prakash ⋅ Noseong Park
Most time-series foundation models operate on fixed temporal grids, representing signals as vectors, tokens, or patches and learning maps between discretized observations. This grid-to-grid view is convenient for standard forecasting, but it is not the natural object for reusable time-series modeling, where temporal resolution, forecast horizon, and query locations may change across tasks and deployments. We argue that time-series foundation models should instead learn sample-conditioned function-to-function operators. Given finite observations of an underlying signal, such an operator maps the observed context to an output function that can be evaluated at user-specified times. We propose \textbf{TimeOperator}, a compact branch-trunk backbone that realizes this view with a shared context encoder, spectrally informed query features, context-adaptive trunk parameters, and multi-basis temporal responses. Forecasting and imputation are handled by evaluating the same operator on different query sets, while classification reuses the shared context representation. Across deterministic and probabilistic forecasting, imputation, and classification benchmarks, TimeOperator is competitive with existing time-series foundation models. Further analyses show that cross-frequency and cross-horizon prediction can be handled by changing only the query set, rather than changing the decoder or relying on post-hoc resampling.
Time series analysis with Gumbel dynamics
Yiliu Wang ⋅ Timothy Kim ⋅ Eric Todd SheaBrown ⋅ Uygar Sümbül
Switching dynamical systems can model complicated time series data while maintaining interpretability by inferring a finite set of dynamics primitives and explaining different portions of the observed time series with one of these primitives. However, due to the discrete nature of this set, such models struggle to capture smooth, variable-speed transitions, as well as stochastic mixtures of overlapping states (e.g., non-instantaneous state transitions), and the inferred dynamics often display spurious rapid switching on real-world datasets. Here, we propose the Gumbel Dynamical Model (GDM). First, by introducing a continuous relaxation of discrete states and a different noise model defined on the relaxed-discrete state space via the Gumbel distribution, GDM expands the set of available state dynamics, allowing the model to approximate smoother and non-stationary ground-truth dynamics more faithfully. Breaking from established literature, this new class directly links states to observations and does not blur latent dynamics with Gaussian noise. Second, the relaxation makes the model fully differentiable, enabling fast and scalable training with standard gradient descent methods. We validate our approach on standard simulation datasets and highlight its ability to model soft, sticky states and transitions in a stochastic setting. Furthermore, we apply our model to two real-world datasets, demonstrating its ability to infer interpretable states in stochastic time series with multiple dynamics, a setting where traditional methods often fail.
TimeTok: Granularity-Controllable Time-Series Generation via Hierarchical Tokenization
Seokhyun Lee ⋅ Jaeho Kim ⋅ Changjun Oh ⋅ Mihaela van der Schaar ⋅ Changhee Lee
Time-series data are inherently multiscale, spanning diverse temporal granularities from coarse trends to fine-scale dynamics. However, existing time-series generative models provide limited control over the temporal granularity of both inputs and outputs, restricting their ability to condition on user-provided coarse sketches and generate samples at a desired target granularity. To address this, we introduce TimeTok, a unified framework for Granularity-Controllable Time-Series Generation (GC-TSG), which generates time series at $any$ target granularity from $any$ coarser input (e.g., rough sketches) or without conditioning. At the core of TimeTok is a hierarchical tokenization strategy that maps time series into an ordered sequence of tokens, from coarse to fine temporal granularity. Our autoregressive generation process operates across these granularity levels, producing token blocks that are decoded back into continuous time series. This design naturally enables GC-TSG within a single framework, where controlling the number of token blocks provides explicit control over output detail. Experiments show that TimeTok excels at GC-TSG tasks while achieving state-of-the-art performance in standard generation. Furthermore, we showcase TimeTok's potential as a foundational tokenizer by training on multiple datasets with heterogeneous temporal granularities, verifying strong transferability that consistently outperforms models trained on individual datasets. To our knowledge, this is the first unified framework that covers the full generative spectrum for time series, offering a valuable foundation for models that benefit from diverse temporal granularities.
TOM-Pruning: Target-aware Output Manifold for LLM Pruning
Junchen Hao ⋅ Weikang Meng ⋅ Yingjian Li ⋅ Zheng Zhang
Post-training pruning compresses pretrained large language models (LLMs) by removing redundant weights without full retraining. However, existing pruning criteria primarily rely on parameter statistics or generic reconstruction errors, fundamentally ignoring the underlying geometric structure of how weight removals perturb the layer output. In this paper, we rethink LLM pruning from a local manifold perspective and reveal that representative methods (e.g., SparseGPT and Wanda) implicitly assume a standard Euclidean output metric. This geometry-agnostic assumption yields an isotropic manifold that assigns uniform costs to all perturbation directions, thereby failing to capture target-dependent sensitivity. To overcome this structural limitation, we propose Target-Aware Output Manifold Pruning (TOM-Pruning). TOM-Pruning equips the output perturbation space with a target-aware local Riemannian metric, assigning direction-dependent geometric costs based on the perturbation's alignment with the target response. By mathematically pulling this metric back to the parameter space, we derive a pruning sensitivity score that inherently captures a cosine-based input-output channel affinity. To resolve the poor discriminability of this raw affinity in high-dimensional spaces, we further introduce a Competitive-Hubness Channel Affinity Mechanism to reconstruct a highly distinguishable final pruning criterion. Extensive experiments across multiple LLM families demonstrate that our manifold-driven approach consistently outperforms existing no-weight-update pruning baselines under various sparsity settings, notably reducing WikiText-2 perplexity by 26.8\% on LLaMA-3.2-1B and achieving an 11.0\% relative zero-shot accuracy gain on LLaMA-3-8B at 70\% unstructured sparsity.
ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents
XuHao Hu ⋅ Xi Zhang ⋅ Haiyang Xu ⋅ Kyle Qiao ⋅ Jingyi Yang ⋅ Xuanjing Huang ⋅ Jing Shao ⋅ Ming Yan ⋅ Jieping Ye
Computer Use Agents (CUAs) can act through both atomic GUI actions ($\textit{e.g., click, type}$) and high-level tool calls ($\textit{e.g., API-based file operations}$), but they are often confused by this hybrid action space: they do not know when to continue with GUI actions and when to switch to tools, and finally fail to select the optimal execution path. We view this orchestration problem as GUI-Tool path selection: deciding when the agent should continue with GUI actions and when it should switch to tool calls to form an effective execution trajectory. This difficulty stems from two issues. First, high-quality interleaved GUI-Tool trajectories are scarce, and collecting real tool trajectories is expensive and brittle. Second, existing supervision provides limited guidance for GUI-Tool path selection, as most methods focus on step-level action imitation or final task completion and offer little trajectory-level feedback on whether GUI-Tool switching leads to a more effective execution path. In this paper, we propose $\textbf{ToolCUA}$, an end-to-end agent designed to learn optimal GUI-Tool path selection through a staged training paradigm. We first introduce an $\textbf{Interleaved GUI-Tool Trajectory Scaling Pipeline}$ that repurposes abundant static GUI trajectories and synthesizes a grounded library of tools, making it possible to scale diverse GUI-Tool trajectories without manual engineering or real tool-trajectory collection.Based on this data, we perform Tool-Bootstrapped GUI RFT, which combines warmup SFT with single-turn RL to improve decisions at critical GUI-Tool switching points. Finally, we further optimize ToolCUA with $\textbf{Online Agentic RL}$ in a high-fidelity GUI-Tool environment, using a Tool-Efficient Path Reward that encourages both appropriate tool use and shorter execution paths. Experiments on OSWorld-MCP show that ToolCUA achieves 46.85% accuracy and outperforms the baseline by over 60% relatively, establishing a new state of the art among models of comparable scale. It also improves by 3.9% over GUI-only settings, demonstrating effective GUI-Tool orchestration. The results further suggest that training in a hybrid action space is a promising paradigm for real-world digital agents.
Top-$k$ Identification with Correlated Biased LLM Judges via Anchor Leverage
Sixiong Xie ⋅ Zhuofan Shi ⋅ Haiyang Shen ⋅ Yun Ma ⋅ Xiang Jing
Large-scale model evaluation is increasingly delegated to LLM judges. This makes evaluation cheaper, but it creates a new statistical problem: if several judges share a systematic bias, averaging more judge scores can make a ranking look more confident without making it more correct. We study fixed-confidence top-$k$ identification when judge scores are cheap, correlated, and biased, while trusted anchor labels are costly. Under a low-dimensional factor-bias model, we show that top-$k$ identifiability is controlled not by the total number of anchors, but by whether anchors span the feature directions separating arms across the top-$k$ boundary. This yields an instance-dependent cost characterization with two components: effective information from correlated judge scores and a profiled anchor-leverage term quantifying the cost of removing shared bias. We propose Profile Track-and-Stop, a sequential allocation algorithm that tracks the resulting cost-optimal allocation with a conservative pilot anchoring phase and asymptotically matches the lower-bound constant. Experiments on a 10-judge Arena-Hard-v2.0 evaluation instance with 23,545 real API judge scores, together with synthetic and semi-synthetic ablations, are consistent with the predicted failure mode: anchor-free methods can plateau under cross-family judge bias, whereas boundary-aware anchoring improves top-$k$ recovery at lower cost.
TOPPO: Rethinking PPO for Multi-Task Reinforcement Learning with Critic Balancing
Yuanpeng Li ⋅ Rui Miao ⋅ Gefei Lin ⋅ Annie Qu
Soft Actor-Critic (SAC) and its variants dominate Multi-Task Reinforcement Learning (MTRL) due to their off-policy sample efficiency, while on-policy methods such as Proximal Policy Optimization (PPO) remain underexplored. We diagnose that PPO in MTRL suffers from a previously overlooked issue: critic-side gradient ill-conditioning, which may cause tail tasks to stall while easy tasks dominate the value function's updates. To address this, we propose TOPPO (Tail-Optimized PPO), a reformulation of PPO via Critic Balancing---a set of modules that improve gradient conditioning and balance learning dynamics across tasks. Unlike prior approaches that rely on modular architectures or large models, TOPPO targets the optimization bottleneck within PPO itself. Empirically, TOPPO achieves stronger mean and tail-task performance than published SAC-family and ARS-family baselines while using substantially fewer parameters and environment steps on Meta-World+ benchmark. Notably, TOPPO matches or surpasses strong SAC baselines early in training and maintains superior performance at full budget. Ablations confirm the effectiveness of each module in TOPPO and provide insights into their interactions. Our results demonstrate that, with proper optimization, on-policy methods can rival or exceed off-policy approaches in MTRL, challenging the prevailing reliance on SAC and highlighting critic-side gradient conditioning as the central bottleneck.
Online robust zero-sum Markov games pose a coupled operator-learning and game-solving problem: from nominal interaction data, the learner must recover a worst-case minimax Bellman operator while maintaining strategic coverage and controlling stage-game approximation error. Existing robust reinforcement-learning theory does not yet provide such an operator-learning perspective beyond specialized multi-agent settings. We develop an optimistic dual robust Bellman framework that combines dual robust Bellman fitting, optimistic minimax planning, backward-validated confidence sets, and exploiter-mix data collection. Our analysis is organized around the robust Bellman residuals induced by the candidate value class and yields a regret reduction that separates exploration complexity, robust operator-estimation error, stage-game solution error, and misspecification. Under explicit interpretable assumptions, we obtain conditional theorem interfaces for general and projected function approximation, together with closed results in several structured regimes, including tabular, epochwise/cross-fit linear, self-normalized pure-online linear / finite-rank kernel, and projected RKHS truncation. We also provide numerical experiments that validate the theoretical predictions.
Towards Effective and Transferable Physical Camouflage against Multi-View BEV-based 3D Perception in Autonomous Driving
Linye Lyu ⋅ Jiawei Zhou ⋅ Daojing He ⋅ YU LI
Modern autonomous driving systems widely adopt multi-view Bird’s-Eye-View (BEV) based 3D perception models due to their superior performance. Despite their success, the robustness of these models against adversarial camouflage attacks remains largely unexplored. Existing camouflage attacks primarily target single-view 2D detectors, limiting their effectiveness in multi-view 3D perception settings. To address this gap, we propose BEVCA, a novel adversarial camouflage generation framework tailored for multi-view BEV-based 3D perception. Our framework integrates a BEV-feature-based adversarial loss with a multi-view neural rendering module, enabling effective and transferable camouflage attacks across different models and tasks. Extensive experiments in both digital and physical settings demonstrate that BEVCA outperforms state-of-the-art baselines and exhibits strong robustness under real-world conditions.
Towards Scalable Context-Aware Single-Cell Spatial Transcriptomics Prediction from Histology Images
Zijun Gao ⋅ Chunbin Gu ⋅ Xiangde Luo ⋅ Jinxi Xiang ⋅ Pheng-Ann Heng
Predicting gene expression from H\&E-stained histology images offers a scalable alternative to costly spatial transcriptomics, yet most existing methods operate at the spot level, where signals from multiple cells are aggregated and critical cellular heterogeneity is obscured. Extending this paradigm to single-cell resolution is non-trivial. Naively applying pathology foundation models faces a scale mismatch: their patch-level representations mix multiple cells, whereas per-cell cropping or resizing distorts morphology and removes local context. Conversely, segmentation-based models without strong pretrained visual encoders often lack the morphological representation capacity needed for accurate molecular prediction and inherit errors from imperfect cell boundary masks. Here, we present CELLO, an efficient end-to-end framework that performs a single pathology foundation model forward pass per image and uses grid sampling to extract location-specific features for all cells simultaneously. We further introduce a distance-decay cross-attention module that refines each cell representation using spatially biased local morphological context. Using 52 paired Xenium--WSI samples spanning 12 organs and approximately 10M cells, CELLO achieves state-of-the-art performance in both in-distribution and out-of-distribution settings while delivering at least a 7$\times$ inference speedup over DeepSpot2Cell. Our work establishes a scalable foundation for single-cell gene expression prediction from H\&E images.
TraceSim: A Generative Simulator and Benchmark for Joint scRNA-seq and Lineage Tracing
Mehrshad Sadria ⋅ Katsuki Fujisawa ⋅ Xun Shen
Single-cell lineage tracing paired with transcriptomics is fundamentally transforming our understanding of cellular differentiation. However, due to immense technical difficulties and experimental complexity, paired barcoding–scRNA-seq datasets remain scarce, fragmented across only a handful of studies, and frequently corrupted by stochastic barcode silencing and site-level dropout. Because transcriptomic sequencing is destructive, the continuous trajectories of differentiating cells cannot be observed directly, leaving accurate generative simulators as the only practical path to the standardized ground- truth benchmarks needed for rigorous method validation. Yet the few existing simulators generate gene expression as a strictly Markovian process, failing to capture the fate bias seen in real biological systems, where early progenitor cells already carry coordinated transcriptomic signatures of their distant terminal fates. Consequently, modern machine learning methods designed to predict cell fate cannot be rigorously benchmarked: the synthetic data systematically lacks the predictive signals these models are meant to learn. We introduce TraceSim, a highly controllable, continuous multi-fate simulation framework. By modeling cellular differentiation as an overdamped Langevin process on a Waddington landscape, TraceSim generates non-Markovian transcriptomic trajectories whose fate bias is tunable by construction, together with tunable parameters for asynchronous apoptosis, CRISPR barcoding errors, and transcriptomic signal-to-noise sparsity. Through a comprehensive weakness-targeted evaluation, we demonstrate that TraceSim successfully maps the predictive limits of modern architectures and shows the structural limitations of current simulators. TraceSim establishes a new standard for validating lineage reconstruction and early cell fate prediction methods.
Tracing Persona Vectors Through LLM Pretraining
Viktor Moskvoretskii ⋅ Dominik Glandorf ⋅ Jorge Medina Moreira ⋅ Tanja Käser ⋅ Robert West
How large language models internally represent high-level behaviors is a core interpretability question with direct relevance to AI safety: it determines what we can detect, audit, or intervene on. Recent work has shown that traits such as evil or sycophancy correspond to linear directions in the internal activations, the so-called persona vectors. While these vectors are now utilized to inspect and steer model behavior in safety-relevant settings, how these representations are formed during training remains unknown. To address this, we trace persona vectors across the pretraining of OLMo-3-7B, with qualitative replication on Apertus-8B, a fully open model trained under a different recipe. They form remarkably early--- within 0.22% of OLMo-3 pretraining--- and remain effective for steering the fully post-trained instruct models. Persona vectors continue to refine geometrically and semantically throughout pretraining, but the core representation is already formed. We further compare alternative elicitation strategies and find that all yield effective directions, with each strategy surfacing qualitatively distinct facets of the underlying persona. Our results establish persona representations as stable features of early pretraining and open a path to studying how training forms, refines, and shapes them.
Trading Sensing for Structure: Sparse IMU-EMG Fingertip Force Estimation via Neuro-inspired Structured Modeling
Yang Gao ⋅ Yingjing Xiao ⋅ Junbin Ren ⋅ Chenxu Zhang ⋅ Wenbo Zhang ⋅ Zhanpeng Jin
We present a Neuro-inspired Structured Model (NiSM) for fine-grained finger-level contact force estimation during dexterous grasping from sparse wearable sensing. NiSM uses a thumb-mounted inertial measurement unit (IMU) and a single-channel wrist electromyography (EMG) sensor, while deriving grasp context from an IMU-based hand-shape latent. This work studies the trade-off between sensing density and structural inductive bias: when sensor coverage is limited, inferred grasp context, kinematic coupling, and muscle-effort modulation can provide useful constraints for force prediction. Inspired by human sensorimotor control, NiSM organizes computation into complementary pathways: a context-conditioned feedforward pathway maps thumb kinematics to finger-specific estimates through sparse learned gating, while an EMG-conditioned modulation pathway provides effort-dependent scaling. By explicitly encoding these structural priors, NiSM reduces reliance on dense sensor arrays and large generic black-box models. Evaluations on multi-user, multi-object datasets show that NiSM achieves performance competitive with a dense-sensing reference where comparable dense inputs are available, while using substantially fewer parameters, and is more robust than generic models under limited-data settings. These results indicate that structured model design can recover part of the information typically supplied by denser sensing, offering a practical direction for wearable finger force estimation under real-world deployment constraints.
Training-Free Cultural Alignment of Large Language Models via Persona Disagreement
Dao S Minh ⋅ Trung-Kiet Huynh ⋅ Chi Nguyen Tran ⋅ Phu-Quy Nguyen-Lam ⋅ Phu-Hoa Pham ⋅ Tuan Nguyen ⋅ The Anh Han ⋅ Long Tran-Thanh
Large language models are increasingly deployed in decisions that require culture-dependent moral judgements, yet they answer as if the whole world thinks with a Western mindset. The Moral Machine experiment showed this is wrong at scale: 40 million judgments across 233 countries reveal that moral preferences are systematically structured by culture, and a model that ignores this variation does not merely underperform, but also imposes one society's intuitions on all others. Existing fixes do not scale to global deployment, as fine-tuning needs per-country preference data and GPU budgets, reward-guided decoding needs per-country reward models, and activation steering needs access to model internals that black-box APIs do not expose. In this work, we focus on this realistic inference-time regime, with no weight updates, no training data, and no internal access. The key observation is that within-country demographic disagreement, not consensus, is the steering signal. When culturally grounded personas agree, the base model is already calibrated. But when they disagree, the spread tells us what to fix and how. We propose DISCA (Disagreement-Informed Steering for Cultural Alignment), which instantiates each country as a panel of four World-Values-Survey-grounded persona agents, converts their disagreement into a bounded, loss-averse correction whose magnitude is set by the panel's variance, and shrinks the correction toward zero when the estimate is unreliable. Across 20 countries and 7 open-weight backbones (2B–70B) from five model families, DISCA reduces cultural misalignment on MultiTP by 10–24% on binary moral dilemmas and 2–7% on open-ended scenarios. Furthermore, a smaller 14B backbone with DISCA reaches lower absolute misalignment than a vanilla 70B model.
Training Language Models via Neural Cellular Automata
Dan Lee ⋅ Seungwook Han ⋅ Akarsh Kumar ⋅ Pulkit Agrawal
Pre-training is crucial for large language models (LLMs), as it is when most representations and capabilities are acquired. However, natural language pre-training has problems: high-quality text is finite, contains human biases, and entangles knowledge with reasoning. This raises a fundamental question: is natural language the only path to intelligence? We propose using neural cellular automata (NCA) to generate synthetic, non-linguistic data for pre-pre-training LLMs--training on synthetic-then-natural language. NCA data exhibits rich spatiotemporal structure and statistics resembling natural language while being controllable and cheap to generate at scale. We find that pre-pre-training on only 164M NCA tokens improves downstream language modeling by up to 6\% and accelerates convergence by up to 1.6x in 1-3B models trained with Chinchilla optimal data budgets. Surprisingly, this even outperforms pre-pre-training on 1.6B tokens of high-quality natural language data with more compute. Investigating what drives transfer, we find that attention layers are the most transferable, and that optimal NCA complexity varies by domain: code benefits from simpler dynamics, while math and web text favor more complex ones. These results enable systematic tuning of the synthetic distribution to target domains. More broadly, our work opens a path toward more efficient models with fully synthetic pre-training.
Training with Harnesses: On-Policy Harness Self-Distillation for Complex Reasoning
Zhengyang Zhao ⋅ Lu Ma ⋅ Wentao Zhang
Inference-time harnesses substantially improve large language models on complex reasoning tasks. However, the intrinsic capabilities of the underlying model remain unchanged by the addition of these external workflows. To bridge this gap, we introduce \emph{On-Policy Harness Self-Distillation} (OPHSD), which employs the harness-augmented current model as a teacher for self-distillation, thereby introducing extra supervisory signals from the harness beyond training data. OPHSD internalizes task-specific harness capabilities into the student model, yielding robust generalizability and strong standalone performance across diverse reasoning tasks. Evaluated across draft--verify harness for text classification and plan--solve for mathematical reasoning tasks, OPHSD consistently outperforms strong baselines (e.g., +10.83\% over OPSD on HMMT25). Our analysis further indicates that reattaching the harness during inference yields no additional benefits and can even degrade performance, suggesting that complex harnesses need not always be permanent fixtures; instead, they can serve as temporary training scaffolds whose benefits are permanently fed back into the base model. Our code and training data are available at \url{https://anonymous.4open.science/r/OPHSD-On-Policy-Harness-Self-Distillation-99F7}.
Trajectory-Consistent Dropout for Uncertainty Decomposition in Hamiltonian Neural Networks
Stephen J Roberts ⋅ Yuki Tachibana
Hamiltonian Neural Networks (HNNs) produce stable long-horizon rollouts by integrating a learned Hamiltonian with a symplectic scheme. Standard Monte Carlo (MC) dropout conflicts with this per-trajectory interpretation because resampling the dropout mask inside a leapfrog integrator means successive substeps follow different subnetworks, so a single MC rollout no longer corresponds to any one sampled Hamiltonian. We propose trajectory-consistent dropout, implemented as Fixed-mask MC, where each stochastic rollout selects one mask from a finite bank and reuses it across the full integration path. This restores a coherent per-rollout Hamiltonian sample and makes uncertainty decomposition operational. Crossing the mask axis with an initial-condition ensemble separates model-sample uncertainty from sensitivity to the initial state. On coupled-pendulum rollouts, Fixed-mask MC reduces mean energy drift by 72% relative to Standard MC. Under held-out mass ratios, Fixed-mask MC increases the epistemic share of predictive variance by 16.9 percentage points from in-distribution to out-of-distribution, compared with 4.2 percentage points for Standard MC, indicating a substantially stronger shift of the uncertainty budget toward model uncertainty when the physics moves outside training support. Spring-chain and transverse-field Ising experiments show that the same axis-preserving decomposition transfers to structured pairwise and latent quantum dynamics, although point-prediction and likelihood gains remain system-dependent. These results position Fixed-mask MC as a single-training-run route to trajectory-consistent, physically interpretable uncertainty in structured neural dynamical models.
TrajLift: Encoding Verbal Memory Dynamics via Heat Diffusion on Semantic Hierarchies
Jiawen Kang ⋅ Jinchao Li ⋅ Mingyu Cui ⋅ Junan Li ⋅ Xixin Wu ⋅ Helen Meng
Verbal memory retrieval carries rich diagnostic information for Alzheimer's disease (AD) through the order, hesitation, and categorical organization of recalled words, yet computational modeling of this retrieval process remains largely unexplored. A key challenge is that timing and hierarchical structure are diagnostically entangled: the diagnostic meaning of retrieval timing depends on its hierarchical context. We propose TrajLift, a spectral framework that models verbal recall as timed trajectories on a hierarchical graph. Our central mechanism is graph heat diffusion on the semantic hierarchy, which provides a unified operator for capturing both multi-scale hierarchical structure and continuous retrieval timing. Formal analysis shows that the resulting representation is provably separable into structural and temporal components with spectral selectivity across hierarchical levels. Experiments on synthetic benchmarks and a real-world clinical corpus demonstrate consistent improvements over baselines lacking joint structure-time modeling.
Transferable SCF-Acceleration through Solver-Aligned Initialization Learning
Eike S. Eberhard ⋅ Viktor Kotsev ⋅ Timm Güthle ⋅ Stephan Günnemann
The cost of Kohn-Sham density functional theory (KS-DFT) calculations scales with the number of solver iterations, which depends on the quality of the initial guess. Machine learning methods that predict initial guesses from molecular geometry can reduce this cost, but matrix-prediction models fail when extrapolating to larger molecules, degrading rather than accelerating convergence [Liu et al., 2025]. We show that this failure is a supervision problem, not an extrapolation problem: models trained on ground-state targets fit those targets well out of distribution, yet produce initial guesses that slow convergence. Solver-Aligned Initialization Learning (SAIL) resolves this for both Hamiltonian and density matrix models by differentiating through the self-consistent field (SCF) solver end-to-end. We introduce the Effective Relative Iteration Count (ERIC), a correction to the commonly used RIC that accounts for hidden Fock-build overhead. On QM40, which contains molecules up to 4$\times$ larger than the training distribution, SAIL reduces ERIC by 37\% (PBE), 33\% (SCAN), and 28\% (B3LYP), more than doubling the previous state-of-the-art reduction on B3LYP. On QMugs molecules 10$\times$ larger than the training set, SAIL delivers a 1.35$\times$ wall-time speedup at the hybrid level of theory, extending ML SCF acceleration to large drug-like molecules.
Traversal-Invariant Positional Encoding for Serialized Graphs
Krish Mody ⋅ Sridhar Radhakrishnan ⋅ Chandra N Sekharan
Many systems represent graphs as sequences by visiting nodes one at a time, for example, a molecule written as a SMILES string or a program written as a token stream from an AST traversal. The same graph can be traversed in many ways, producing different sequences for the same underlying object. Sequence Transformers treat these as distinct inputs and produce distinct embeddings, even though the graph remains unchanged. Existing graph-native Transformers address this by operating directly on the graph, but many practical pipelines receive serialized inputs and cannot be redesigned. We provide a theoretical characterization in three parts. Tree-positional encoding achieves structural invariance for acyclic graphs under canonical-root serialization but is insufficient for cyclic graphs due to spanning-tree non-uniqueness. Augmenting with cycle-membership encoding derived from the canonical cycle basis introduces a traversal-invariant anchor on cycle vertices. The resulting Tree+Cycle Positional Encoding is a drop-in replacement for sinusoidal PE within graph-aware serialization pipelines (those preserving token-to-atom alignment for structural tokens), requiring no architectural changes to the sequence model. It applies to any domain where labeled graphs are given as sequences with this alignment property. We validate the approach on two testbeds. On synthetic cycle counting and ring membership tasks, it outperforms the GIN baseline by $2\times$ on cycle counting and matches it on ring membership, demonstrating that pre-computed cycle structure encoded as PE matches message-passing GNNs at end-task performance. On molecular property prediction, it achieves $0.977 \pm 0.029$ mean cosine similarity across re-rootings ($n = 24{,}960$), outperforms a parameter-matched graph-native Transformer (Graphormer) on BACE and BBBP, and enables fragment-level interpretability with $+5$ points on alert alignment over a methodologically matched sinusoidal baseline.
Treating Hyperparameters as Interventions: Task-Invariant Representation Learning for Transferable HPO
Mengyang Li ⋅ Ou Wu
We argue that hyperparameter optimization (HPO) is best approached not as a black-box search problem but as the inverse of an intervention: when training dynamics are observed, $\lambda$ is the controlled cause and the task is what is invariant to it, and the cost of HPO is largely the cost of failing to separate the two. We propose CA-HPO (Configuration-invariAnt HPO), a transfer HPO framework that learns a task representation explicitly constrained to be invariant under hyperparameter interventions and predictively sufficient for performance. We define this target representation as the solution to a constrained variational problem rather than as an identified latent variable, so the framework avoids the strong identifiability assumptions that previous causal representation learning approaches demand. We derive three guarantees in this setting: a population-level result showing that minimizing a predictive loss together with an invariance penalty recovers the target representation, a finite-sample bound on the estimation error of the empirical minimizer, and a counterfactual prediction risk bound that decomposes into representation error, meta-generalization error, and aleatoric noise. Experiments on Llama-3-8B fine-tuning and twelve classical benchmarks show 5 to 15 times speedup over Bayesian optimization and recent meta-HPO baselines, while the learned representations remain stable across a 100-fold range of learning rates, transfer across domains with different hyperparameter spaces, and pass a direct conditional independence test against $\lambda$.
TreeGraft: Adaptive Multi-Drafter Grafting for Tree-Based Speculative Decoding
Jiaming Fan ⋅ Daming Cao ⋅ Canchen Huang ⋅ Jiale Fu ⋅ Jin Zhang ⋅ Junjie Gao ⋅ Kai Yang ⋅ Xiangzhong Luo ⋅ Xu Yang
Speculative decoding accelerates large language model inference through a draft-then-verify paradigm. Building on this, tree-structured methods improve inference by organizing proposals into multiple candidate paths, increasing the accepted length. However, existing tree-structured methods use a single drafter for all drafting steps, creating a dilemma: a smaller drafter is fast but yields lower-quality trees, whereas a larger drafter improves tree quality but suffers from high latency. To address this, we propose TreeGraft, a multi-drafter framework in which drafters of different costs jointly construct a shared draft tree. TreeGraft uses the stronger drafter to rescore candidates by updating scores assigned by the weaker drafter, reselect grafting positions, and recover promising paths left unexplored. It also integrates stronger drafter expansions non-destructively, preserving existing branches that may still be accepted by the target model. Together, these designs improve the quality of the shared draft tree. To control the drafting cost, TreeGraft introduces a lightweight scheduler distilled from an offline value system to decide when to call the stronger drafter. Across 10 model pairs and 6 benchmarks, TreeGraft outperforms the better of the two fixed single-drafter endpoint strategies by 15.1\% on average, reaching a maximum gain of 26.6\%. Our code is available at https://anonymous.4open.science/r/TreeGraft-E983.
TROPT: An Open Framework for Unifying and Advancing Discrete Text Optimization
Matan Ben-Tov ⋅ Mahmood Sharif
Discrete text-trigger optimization—searching for text sequences that, when ingested by a model, steer it toward a specified objective—underpins model red-teaming (e.g., LLM jailbreaks), as well as auditing and interpretability. However, the current state of discrete optimizers hinders their adoption and progress. First, existing optimizers, when open-sourced at all, are scattered across research codebases tied to specific problem domains. Second, optimizer variants proliferate, each requiring engineering overhead to use or extend, and hard to compare head-to-head. Together, these raise the bar to adopting optimizers in existing or new domains, and to advancing them via new strategies. We address these gaps with TROPT, the first open-source framework that unifies discrete optimizers' execution and standardizes their development under a single interface. TROPT currently ships with 30+ end-to-end optimization recipes—covering applications such as jailbreaking and probing model internals—built from 15+ optimizers (spanning white-box to black-box access) and 15+ losses, from foundational to state-of-the-art methods. Beyond democratizing established methods, TROPT makes it easy to customize optimization recipes by swapping any component—models, objectives, and optimizers—extending its reach across domains and new applications. Demonstrating its utility, we leverage TROPT in several controlled studies: (i) novel large-scale experiments comparing and enhancing optimization strategies for LLM jailbreaks, revealing potent-yet-underadopted techniques; and (ii) cross-domain demonstrations (e.g., benign prompt recovery and corpus poisoning). In all, TROPT significantly lowers the barrier to adopting and advancing discrete text optimization.
Big goals are hard to achieve all at once; breaking them into small steps is wiser. We present Trust Region Policy Distillation (TOP-D), which transforms the notoriously unstable, high-variance On-Policy Distillation (OPD) into a stable training paradigm by dynamically constructing a proximal teacher. Theoretically, we establish a rigorous framework demonstrating that TOP-D inherently controls gradient variance. By providing a formal global convergence analysis alongside a monotonic improvement bound, we mathematically formalize the reliability and stability of the overall training dynamics. Empirically, TOP-D dramatically enhances training stability, sample efficiency, and final performance on mathematical reasoning tasks. More importantly, TOP-D introduces zero additional computational overhead, positioning itself as a promising alternative to the well-established OPD paradigm.
TSFM Meets LLM: Context as Covariate
Hemanth Ram Govindarajan Kirubaharan ⋅ Sai Shankar Narasimhan ⋅ Shubhankar Agarwal ⋅ Sandeep Chinchali
Textual information is ubiquitous in real-world time series datasets, appearing as news, reports, and annotations that hold key predictive signals not deducible from numerical history alone. Large Language Models (LLMs) interpret text but lack numerical precision, while Time Series Foundation Models (TSFMs) are quantitatively accurate but cannot reason about the effects of text. Moreover, the lack of large-scale paired multimodal datasets limits progress in bridging this gap. In this work, we address this challenge by providing a scalable approach for generating realistic multimodal datasets by annotating real-world time series with LLMs. Building on this, we propose Context as Covariates (CoCo), a forecasting framework in which an LLM distills its textual reasoning into numerically grounded forecast and confidence covariates that guide a TSFM backbone, thereby exploiting the strengths of both LLMs and TSFMs. We align the two models in stages, first fine-tuning the LLM via Group Relative Policy Optimization (GRPO) with a reward tied to covariate-induced improvement in forecasting performance, and then fine-tuning the TSFM on covariates generated by the aligned LLM. Our experiments show that models trained on our synthetic multimodal corpus generalize to real benchmarks spanning finance, economics, weather, traffic, and security. Notably, CoCo with Qwen3-4B-Instruct-2507 as the LLM and Chronos2 as the TSFM achieves up to 12% MAE improvement over state-of-the-art TSFMs and LLM forecasters, such as GPT-4o.
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Zhiheng Liu ⋅ Weiming Ren ⋅ Xiaoke Huang ⋅ Shoufa Chen ⋅ Tianhong Li ⋅ Mengzhao Chen ⋅ Yatai Ji ⋅ Sen He ⋅ Jonas Schult ⋅ Belinda Zeng ⋅ Tao Xiang ⋅ Wenhu Chen ⋅ Ping Luo ⋅ Luke Zettlemoyer ⋅ Yuren Cong
Unified multimodal models typically rely on pretrained vision encoders and use separate visual representations for understanding and generation, creating misalignment between the two tasks and preventing fully end-to-end optimization from raw pixels. We introduce Tuna-2, a native unified multimodal model that performs visual understanding and generation directly based on pixel embeddings. Tuna-2 drastically simplifies the model architecture by employing simple patch embedding layers to encode visual input, completely discarding the modular vision encoder designs such as the VAE or the representation encoder. Experiments show that Tuna-2 achieves state-of-the-art performance in multimodal benchmarks, demonstrating that unified pixel-space modelling can fully compete with latent-space approaches for high-quality image generation. Moreover, while the encoder-based variant converges faster in early pretraining, Tuna-2's encoder-free design achieves stronger multimodal understanding at scale, particularly on tasks requiring fine-grained visual perception. These results show that pretrained vision encoders are not necessary for multimodal modelling, and end-to-end pixel-space learning offers a scalable path toward stronger visual representations for both generation and perception.
Properly tuned hyperparameters are critical for reinforcement learning algorithms to perform well, limiting their use in real world applications. AutoRL aims to to address this by nesting the reinforcement learning problem within an outer optimization loop, where a black-box optimizer tunes the reinforcement learning algorithm's hyperparameters leaving the AutoRL algorithm's hyperparameters (hyper-hyperparameters) fixed. We provide the first empirical study of an AutoRL algorithm's sensitivity with respect to its hyper-hyperparameters. Our results demonstrate that SMAC3 with Hyperband, used to tune PPO, is sensitive to its hyper-hyperparameters, albeit less so than PPO is to its hyperparameters, and this reduction comes at a significant computational cost. The per-environment tuned performance gains do not outperform random search, suggesting the reduction in sensitivity may not justify the substantial compute and engineering overhead AutoRL imposes.
TwinPrune: Density-Aware Two-Phase Token Pruning for Vision-Language Models
Kai Liu ⋅ Anqi Li ⋅ Junxian Li ⋅ Zhixin Wang ⋅ Zhikai Chen ⋅ Renjing Pei ⋅ Yulun Zhang
We propose a unified spatial-density signal that scores token importance for \emph{both} phases of inference, instantiated as a bipartite merge during prefill and a KV eviction during decode, with measurable gains in each phase. To verify the two phases independently, we partition benchmarks into two diagnostic roles by their scoring rule. \emph{Prefill probes} such as MMBench, POPE, and ScienceQA score only the first decoded token, so their accuracy reflects only the argmax of prefill logits, computed before any decode-phase KV eviction takes effect. \emph{Full-pipeline probes} such as TextVQA, DocVQA, and MM-Vet score the full decoded sequence and therefore reveal both phases. We expose a fundamental evaluation flaw: published claims of decode-phase token pruning validated only on prefill probes have measured prefill quality, not decode quality. Empirically, removing all visual KV at decode leaves prefill-probe accuracy essentially unchanged while collapsing TextVQA and MM-Vet by tens of points, and the insensitivity is a property of the scoring rule, not of the compression strength. Across seven model variants from the InternVL-3.5, Qwen3-VL, and LLaVA-1.5 families and the three full-pipeline-probe benchmarks, we deliver an honest three-axis Pareto over compute, memory, and accuracy. Density-based prefill compression is competitive on the accuracy frontier, scales linearly in the visual-token count, and is roughly two orders of magnitude faster in token selection at high resolution. On the decode side, density top-$k$ KV eviction wins every TextVQA and MM-Vet cell against the strongest published baselines H2O, SnapKV, and StreamingLLM across three keep ratios. The advantage widens to $9.6$ percentage points over H2O at $12.5\%$ keep on TextVQA and holds $7$ to $10$ percentage points across the InternVL-3.5 series from 2B to 38B. Compression also scales gracefully with model size: every matched 8B-to-38B transition we measure reduces, rather than amplifies, the accuracy loss. We recommend MM-Vet or DocVQA as the minimum full-pipeline-probe test for any decode-phase claim. Our code and model will be released soon.
TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing
Pei Yang ⋅ Wanyi Chen ⋅ Tongyun Yang ⋅ Pengbin Feng ⋅ Jiarong Xing ⋅ Wentao Guo ⋅ Yuhang Yao ⋅ Yuhang Han ⋅ Hanchen Li ⋅ Xu Wang ⋅ Jie Xiao ⋅ Anjie Yang ⋅ Lynn Ai ⋅ Eric Yang ⋅ TIANYU SHI
LLM routing matters most in long-horizon applications such as coding agents, deep research systems, and computer-use agents, where a single user request triggers many model calls. Routing each call to the cheapest sufficient model can cut costs without sacrificing quality, yet existing router benchmarks evaluate only which model should answer a whole prompt. They never expose the router-visible prefix at an intermediate agent step, never test whether a cheaper replacement preserves downstream task success, and often rely on online LLM judges at evaluation time. We introduce TwinRouterBench, a step-level routing benchmark with two tracks. The static track provides 970 router-visible prefixes from 520 instances across SWE-bench, BFCL, mtRAG, QMSum, and PinchBench, each paired with an execution-verified target tier estimated under a released downgrade-and-cascade protocol; scoring is deterministic arithmetic over tier labels, trajectory membership, and token costs, with no online evaluator-side LLM judge. The dynamic track supplies a harness that runs routers on the full 500-case SWE-bench Verified suite; in this paper we report a 100-case held-out evaluation disjoint from the static SWE supervision split. At each LLM call the router selects a concrete model from a locked pool, and success is measured by official task resolution and realized API spend. The two tracks support fast offline iteration followed by end-to-end validation under live agent execution. Code and data are available at https://anonymous.4open.science/r/TwinRouterBench-8820.
UGO: Unified Architecture for General Multi-Object Tracking by Segmentation
Jer Pelhan ⋅ Alan Lukezic ⋅ Matej Kristan
General multi-object tracking (GMOT) tracks all instances of a user-specified category from a single first-frame exemplar. Prior work relies on bounding boxes and surrogate training, and struggles with non-rigid objects, crowded scenes, and distractors. We introduce UGO, a unified GMOT tracker that pairs a pretrained exemplar-conditioned detection head with an instance-propagation head in a common architecture. A novel training-free, energy-minimization consolidation method converts overlapping proposals into exclusive pixel-wise masks and detections, resolving over-segmentation, duplicates, and conflicts. A hierarchical memory spanning global and instance levels improves recall and per-instance segmentation accuracy using a new memory management protocol. UGO sets a new state-of-the-art on GMOT benchmarks and video object counting, and is competitive with specialist MOT methods, establishing a strong paradigm for unified, open-category multi-object tracking.
UltraFlash: Accelerating Megapixel Visual Synthesis
Phuc Lai ⋅ Anh Nguyen ⋅ Phong H Nguyen ⋅ Anh Tran
Megapixel visual synthesis with latent diffusion models requires operating far beyond the resolutions at which these models are typically trained. Cascade refinement is a leading strategy for this setting: it starts from a base image at a native resolution and progressively upsamples it to the target resolution through multiple refinement stages. Each stage typically performs a pixel-space resolution transition followed by latent-space denoising. While prior acceleration efforts mainly focus on reducing denoising cost, we identify a complementary and underexplored bottleneck: the repeated movement of representations through the VAE. We call these operations pixel-space transitions. In cascade pipelines, intermediate latents are often decoded to RGB, resized, and re-encoded before refinement, while final latents are decoded patch by patch with dense overlap to suppress boundary artifacts. We propose UltraFlash, a transition-efficient approach that accelerates these VAE-heavy operations. UltraFlash has two components. Latent Hyperloop replaces intermediate RGB round trips with a learned latent-space shortcut that approximates the decode-resize-encode transformation without repeated VAE decoding and encoding. Sparse Patch-based Decoding accelerates final reconstruction by addressing patch seams caused by inconsistent GroupNorm statistics and reusing cached statistics across patches, enabling sparse VAE decoding with much smaller overlap. Integrated into a state-of-the-art cascade baseline, UltraFlash reduces 4K generation latency by approximately 2x while improving image quality. These results show that optimizing pixel-space transitions can improve both efficiency and fidelity, offering an effective route toward fast, high-quality megapixel visual synthesis.
Understanding and Defending VLM Jailbreaks via Jailbreak-Related Representation Shift
Zhihua Wei ⋅ Qiang Li ⋅ Jian Ruan ⋅ Zhenxin Qin ⋅ 蕾蕾 温 ⋅ Ruiyang Qin ⋅ Qingzhuo Wang ⋅ Dongrui Liu ⋅ Wen Shen
Large vision-language models (VLMs) often exhibit degraded safety alignment when visual inputs are integrated. Even when a text prompt is explicitly harmful, adding an image can substantially increase the jailbreak success rate. In this paper, we observe that VLMs can clearly distinguish benign, refusal, and jailbreak samples in their representation space, and that jailbreak responses often contain safety warnings. These observations lead to the \textit{recognize-but-fail-to-refuse} hypothesis: VLM jailbreaks do not arise from a failure to recognize harmful intent, but instead occur because adding an image induces a representation shift that steers the sample into a distinct jailbreak state where refusal is not triggered. To quantify how this image-induced representation shift contributes to jailbreak behavior, we define a jailbreak direction and characterize the jailbreak-related representation shift as the projection of the total image-induced shift onto this direction. Our analysis shows that the jailbreak-related shift reliably characterizes jailbreak behavior, providing a unified explanation for diverse jailbreak scenarios. Finally, we propose JRS-Rem, a defense method that enhances VLM safety by removing the jailbreak-related shift at inference time. Experiments show that JRS-Rem significantly improves VLM safety across multiple scenarios while preserving utility on benign tasks.
Understanding Layer Patching in Model Size Interpolation
Sara Kangaslahti ⋅ Jonathan Geuter ⋅ Nihal V Nayak ⋅ Marco Fumero ⋅ Francesco Locatello ⋅ David Alvarez-Melis
Zero-shot model size interpolation aims to create new models of intermediate target sizes by combining existing models without additional training. Recent work on boomerang distillation [Kangaslahti et al., 2026] shows that a student language model distilled from a larger teacher can be expanded by iteratively patching its layers, replacing student layers with contiguous blocks of teacher layers to obtain models whose size and performance interpolate between the student and the teacher. In this work, we provide the first systematic study of student-layer selection for model size interpolation. We cast finding the optimal layer subset for each model size as an optimization problem and prove it can be viewed as a shortest-path problem in a certain acyclic graph. In experiments, we show that patching strongly shapes interpolation behavior, with effects that vary substantially across model families. We find that simple sequential strategies---patching either from the first layer to the last or from the last to the first---often achieve surprisingly strong performance in practice. We further introduce KLPatch, a greedy patching algorithm based on KL divergence, which often improves over last-to-first patching and approximately solves the optimization problem. Together, our results provide a principled understanding of how layer patching affects model size interpolation and offer practical guidance for constructing near-optimal interpolated models.
Understanding Multimodal Failure in Action-Chunking Behavioral Cloning
Lorenzo Mazza ⋅ Massimiliano Datres ⋅ Ariel Rodriguez ⋅ Sebastian Bodenstedt ⋅ Gitta Kutyniok ⋅ Stefanie Speidel
Behavioral cloning becomes difficult when the same observation admits several valid actions. We study this problem for action-chunking policies and show that different multimodal parameterizations fail in different ways. For latent-variable policies, posterior-prior regularization makes deployment-time sampling more reliable, but excessive regularization removes the action-conditioned information needed to distinguish demonstrated modes. Reducing this regularization can preserve mode information, but then success depends on whether the prior covers the relevant latent regions. For action-space generative policies, multimodality is constrained by the smoothness of the base-to-action transport: a map with small Lipschitz constant cannot assign substantial probability to many well-separated modes. Covering many modes therefore requires either sharp transitions in base space or off-support bridge regions in action space. Experiments on synthetic multimodal tasks and robotic simulation benchmarks support these mechanisms.
Understanding the Convergence of Direct Training of SNNs with Surrogate Gradients
Zelin Wei ⋅ Tianxing Man ⋅ Wanli Shi ⋅ Bin Gu
Spiking Neural Networks (SNNs) offer high energy efficiency, but direct optimization is challenging because spike generation is governed by a non-differentiable hard threshold. Practical direct-training methods, including surrogate gradients (SG) and local zeroth-order (LocalZO) estimators, maintain hard spikes in the forward pass while using surrogate or perturbation-induced derivatives in the backward pass. This forward--backward mismatch raises a central theoretical question: how do local gradient discrepancies propagate through the spatiotemporal computation graph, and can the resulting biased updates still provide guaranties for the original discrete-spike objective? We develop a unified convergence analysis framework for SG and LocalZO direct SNN training. A smoothed spike objective is introduced only as an analytical bridge, allowing us to separate the core error into two components: the gradient-level mismatch between the direct-training mean direction and the smoothed-objective gradient, and the objective-level consistency gap between the smoothed and original discrete objectives. Both components are controlled by a threshold-tail functional that measures how often hard-trajectory membrane potentials lie near the firing threshold. Under suitable regularity assumptions, we obtain a non-asymptotic guarantee for the original discrete-spike objective. Experiments across representative SNN architectures and benchmarks verify convergence consistency, the existence of errors, and demonstrate that the threshold-tail statistic is an effective diagnostic for direct SNN training.
UniCon3R: Unified Contact-aware 4D Human-Scene Reconstruction from Monocular Video
TANUJ SUR ⋅ Shashank Tripathi ⋅ Nikos Athanasiou ⋅ Ha Linh Nguyen ⋅ Kai Xu ⋅ Michael Black ⋅ Angela Yao
We introduce $\textbf{UniCon3R}$, a unified feed-forward framework for online human-scene 4D reconstruction from monocular video. Current feed-forward human-scene reconstruction methods suffer from artifacts, where bodies float above the ground or penetrate parts of the scene. A key reason is the lack of effective interaction modelling between the human and the environment. Our goal is to exploit contact between the human and the scene during inference to actively improve the human mesh reconstruction. To that end, we explicitly model interaction by inferring 4D contact from the human pose and scene geometry and use the contact as a corrective cue for generating the pose. This enables UniCon3R to jointly recover scene geometry and spatially aligned 4D humans within the scene. Experiments on standard human-centric video benchmarks show that UniCon3R outperforms state-of-the-art baselines on physical plausibility and global human motion estimation while preserving fast, feed-forward inference speeds. The results validate our central claim: contact serves as a powerful internal prior, thus establishing a new paradigm for physically grounded joint human-scene reconstruction. Source code and models will be released upon acceptance.
UniCustom: Unified Visual Conditioning for Multi-reference Image Generation
Yiyan Xu ⋅ Qiulin Wang ⋅ Wenjie Wang ⋅ Yunyao Mao ⋅ Xintao Wang ⋅ Pengfei Wan ⋅ Kun Gai ⋅ Fuli Feng
Multi-reference image generation aims to synthesize images from textual instructions while faithfully preserving subject identities from multiple reference images. Existing VLM-enhanced diffusion models commonly rely on decoupled visual conditioning: semantic ViT features are processed by the VLM for instruction understanding, whereas appearance-rich VAE features are injected later into the diffusion backbone. Despite its intuitive design, this separation makes it difficult for the model to associate each semantically grounded subject with visual details from the correct reference image. As a result, the model may recognize which subject is being referred to, but fail to preserve its identity and fine-grained appearance, leading to attribute leakage and cross-reference confusion in complex multi-reference settings. To address this issue, we propose UniCustom, a unified visual conditioning framework that fuses ViT and VAE features before VLM encoding. This early fusion exposes the VLM to both semantic cues and appearance-rich details, enabling its hidden states to jointly encode the referred subject and corresponding visual appearance with only a lightweight linear fusion layer. To learn such unified representations, we adopt a two-stage training strategy: reconstruction-oriented pretraining that preserves reference-specific appearance details in the fused hidden states, followed by supervised finetuning on single- and multi-reference generation tasks. We further introduce a slot-wise binding regularization that encourages each image slot to preserve low-level details of its corresponding reference, thereby reducing cross-reference entanglement. Experiments on two multi-reference generation benchmarks demonstrate that UniCustom consistently improves subject consistency, instruction following, and compositional fidelity over strong baselines. Our code, checkpoints, and the training dataset will be released soon.
Unified Panoramic Geometry Estimation via Multi-View Foundation Models
Vukasin Bozic ⋅ Isidora Slavkovic ⋅ Dominik Narnhofer ⋅ Nando Metzger ⋅ Denis Rozumny ⋅ Konrad Schindler ⋅ Nikolai Kalischek
Geometry estimation from perspective images has greatly advanced, maturing to the point where off-the-shelf foundation models are able to reconstruct 3D scene structure not only from multi-view imagery, but even from a single view. A natural extension is 3D reconstruction from panoramas, with the exciting prospect of recovering a full $360^\circ$ scene from a single panoramic image. In this work, we introduce PaGeR (Panoramic Geometry Reconstruction), a framework to lift powerful 3D foundation models, designed for perspective imagery, to the panorama domain. Our strategy is to start from a pretrained transformer for 3D reconstruction and turn it into a unified high-performance model that predicts scale-invariant depth, metric depth, sky masks and surface normals from both perspective and omnidirectional images, a single forward pass. By keeping architectural changes to a minimum and mixing perspective and panoramic images during training, Pager retains the rich 3D prior of the underlying foundation model while learning to also estimate geometrically consistent $360^\circ$ scenes from single panoramas. We extensively test our method in both indoor and outddor environments and find that it delivers state-of-the-art performance and excellent zero-shot performance across a wide range of scenes.
Universal Approximation Theorems for Dynamical Systems with Infinite-Time Horizon Guarantees
Ábel Ságodi ⋅ Memming Park
Universal approximation theorems establish the expressive capacity of neural network architectures. For dynamical systems, existing results are limited to finite time horizons or systems with a globally stable equilibrium, leaving multistability and limit cycles unaddressed. We prove that Neural ODEs achieve $\varepsilon$-$\delta$ closeness, i.e., trajectories within error $\varepsilon$ except for initial conditions of measure $< \delta$, over the \emph{infinite} time horizon $[0,\infty)$ for three target classes: (1) Morse-Smale systems (a structurally stable class) with hyperbolic fixed points, (2) Morse-Smale systems with hyperbolic limit cycles via exact period matching, and (3) systems with normally hyperbolic continuous attractors via discretization. We further establish a temporal generalization bound: $\varepsilon$-$\delta$ closeness implies $L^p$ error $\leq \varepsilon^p + \delta \cdot D^p$ for all $t \geq 0$, bridging topological guarantees to training metrics. These results provide the first universal approximation framework for multistable infinite-horizon dynamics.
UnlearningSoup: Is Repeated Tuning Necessary for Large Language Model Unlearning?
Puning Yang ⋅ Qizhou Wang ⋅ Junchi Yu ⋅ Bo Han ⋅ Xiuying Chen
Large language models trained on vast corpora inherently risk memorizing harmful content that may later re-emerge in their outputs. To mitigate this issue, existing unlearning methods typically rely on training-based parameter updates, such as gradient ascent and its variants, to delete targeted content while preserving other knowledge. However, balancing the competing goals of forgetting and retention makes hyperparameter choices for these methods particularly difficult, often requiring repeated tuning to obtain a strong model that still leaves substantial room for improvement and transfers poorly across models and datasets. To address this challenge, we investigate whether unlearning runs exhibit exploitable structure in weight space, and observe that models from different runs still lie in a shared evaluation-performance basin. This suggests that stronger models may be recovered through an unlearning-tailored soup strategy, reducing the need for repeated tuning for further improvement or new settings. Motivated by this, we propose UnlearningSoup, a unified framework that provides two strategies: EfficientSoup uses binary-search-based interpolation to quickly discover a well-performing model in the early stage, where repeated tuning would otherwise make strong model selection costly. PerformanceSoup uses reweighted souping to efficiently unlock the remaining performance potential in the later stage, where repeated tuning becomes increasingly inefficient. Extensive experiments across diverse datasets and models show that UnlearningSoup delivers 2.4× to 3.3× efficiency gains in hyperparameter selection, while consistently improving performance across settings.
Unveiling Implicit Advantage Symmetry: Why GRPO Struggles with Exploration and Difficulty Adaptation
zhiqi yu ⋅ Zhangquan Chen ⋅ Mengting Liu ⋅ Heye Zhang ⋅ Liangqiong Qu
Reinforcement Learning with Verifiable Rewards (RLVR), particularly GRPO, has become the standard for eliciting LLM reasoning. However, its efficiency in exploration and difficulty adaptation remains an open challenge. In this work, we identify an implicit advantage symmetry inherent in Group Relative Advantage Estimation (GRAE) as a structural property that provides a new perspective for understanding these bottlenecks. This symmetry induces two critical limitations: (i) at the group level, strict symmetry in weights between correct and incorrect trajectories leaves unsampled action logits unchanged, thereby hindering the exploration of novel correct solution. (ii) at the sample level, the algorithm implicitly prioritizes medium-difficulty samples, remaining agnostic to the non-stationary demands of difficulty focus. Through controlled experiments, we reveal that this symmetric property is sub-optimal, yielding two pivotal insights: (i) asymmetrically down-weighting the advantages of correct trajectories encourages essential exploration but risks instability; (ii) learning efficiency can be boosted by a curriculum-like transition—prioritizing simpler samples initially before gradually shifting to complex ones. Motivated by these findings, we propose Asymmetric GRAE (A-GRAE), which dynamically modulates exploration incentives and sample-difficulty focus. Experiments across seven benchmarks demonstrate that A-GRAE consistently improves GRPO and its variants across both LLMs and MLLMs.
V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation
Yan-Bo Lin ⋅ Jonah Casebeer ⋅ Long Mai ⋅ Aniruddha Mahapatra ⋅ Gedas Bertasius ⋅ Nicholas J. Bryan
Generating music that temporally aligns with video events is challenging for existing text-to-music models, which lack fine-grained temporal control. We introduce V2M-ZERO, a video-to-music generation approach that generates time-aligned3 music with disentangled time synchronization and semantic control (e.g. genre, mood) from video while requiring zero video-music pairs at training time. Our method is motivated by a key observation: temporal synchronization requires matching when and how much change occurs, not what changes. While musical and visual events differ semantically, they exhibit shared temporal structure that can be captured independently within each modality. We capture this structure through event curves computed from intra-modal similarity using pretrained music and video encoders. By measuring temporal change within each modality independently, these curves provide comparable representations across modalities. This enables a simple training strategy: fine-tune a text-to-music model on music-event curves, then substitute video-event curves at inference without cross-modal training or paired data. Across OES-Pub, MovieGenBench-Music, and AIST++, V2M-ZERO achieves state-of-the-art performance without any paired music-video data, surpassing the strongest prior baselines per metric with 5–9% higher audio quality, 13–15% better semantic alignment, 21–52% improved temporal synchronization, and 28% higher beat alignment on dance videos. We find similar results via a large crowd-source subjective listening test. Our results vali-20 date that temporal alignment through within-modality features is not only effective for video-to-music generation but leads to better performance over paired cross-modal supervision. Furthermore, our approach enables independent controls for timing and music style (e.g. genre, mood) for more controllable generation.
Value-Priced Uncertainty: A Family of PPO-Compatible Exploration Bonuses
qili shen ⋅ Xuanhong Chen ⋅ Ang He ⋅ Dake Zhang ⋅ Kairui Feng
Proximal Policy Optimization (PPO) is value-driven only after experience has been collected: Generalized Advantage Estimation reinforces sampled actions according to their bootstrapped advantage, but the exploration process that produced those samples is still governed mainly by policy-space noise and entropy regularization. As a result, PPO does not explicitly ask, before sampling, which uncertain transitions would most improve its advantage estimates. We argue that PPO exploration should be formulated as *value-priced uncertainty*: uncertainty should be explored when it is both large and able to change the advantage. We propose **VALU** (**V**alue-**A**ware **L**atent **U**ncertainty), a family of PPO-compatible exploration bonuses of the form b(s,a) = P(s,a) U(s,a). The price $ P(s,a) = \lVert \nabla_{z'} \tilde V_\psi(\hat z') \rVert $ measures how strongly uncertainty in the predicted next latent can perturb the advantage, while \(U(s,a)\) is an uncertainty quantity instantiated by non-parametric effective counts or learned prediction errors such as RND. Our theory derives the value-pricing term from a unified advantage-uncertainty view: an uncertainty coordinate induces an optimistic advantage correction through a dual norm. This yields count-based UCB and MCR members, corresponding respectively to optimism over latent-transition uncertainty and one-sample shrinkage of the optimistic advantage envelope. Across six standard MuJoCo continuous-control tasks, VALU substantially improves PPO's sample efficiency and final return, and achieves the best PPO-compatible performance on most tasks while remaining competitive on the others, compared with strong exploration baselines including RND, ICM, RE3, VCSE, and OPPO. Ablations over uncertainty estimators, schedules, latent encoders, and bonus components show that the gains come from pricing uncertainty by local value sensitivity rather than from any single implementation choice.
Variational Monte Carlo for Quantum Excited States via Nested Low-Rank Approximation
Minchan Jeong ⋅ Jongha (Jon) Ryu ⋅ Se-Young Yun ⋅ Gregory Wornell
The electronic Schrödinger equation encodes a molecule's electronic structure and quantum properties, but solving it scales exponentially with the number of electrons. Neural-network variational Monte Carlo (VMC) has reached chemical accuracy on the ground state of small molecules; extending this success to excited states is constrained by two existing paradigms. Penalty-based methods are computationally scalable ($O(K)$ Laplacian count per step) but require tuning the overlap-penalty schedule against the energy gap they aim to compute. Natural-excited-states VMC (NES-VMC) is principled and penalty-free but couples all $K$ states through a joint $\det\Psi$ ansatz that forces $O(K^2)$ Laplacians per step. We resolve this trade-off via a low-rank approximation (LoRA) reparameterization of the rank-$K$ variational objective. The reparameterized objective admits an unbiased gradient estimator via MCMC sampling from the per-state Born density, and a sequential nesting scheme recovers the bottom-$K$ eigenfunctions in eigenvalue order without overlap penalties or post-training diagonalization. The resulting algorithm, NestedLoRA-VMC, is penalty-free, achieves $O(K)$ Laplacian count per step, and inherits a global optimality guarantee from the LoRA principle. On the first-row atoms (Li through Ne) at $K{=}10$, it stays within the chemical-accuracy threshold of NES-VMC on every atom and outperforms the penalty-based baseline of Szabó et al. on seven of the eight.
VecDBLens: A Modular Framework for Diagnosing Vector Databases Retrieval Pipelines
Yutong Zhou ⋅ Guoxin Kang ⋅ Wenxin Zhou ⋅ Haoyu He ⋅ Yuedong Zhu ⋅ Lei Wang ⋅ Wanling Gao ⋅ Fanda Fan ⋅ Jianfeng Zhan
Vector databases are increasingly deployed as retrieval pipelines in which representation, indexing, and query-time execution jointly determine accuracy, throughput, and cost. Existing vector database evaluations primarily focus on indexing while treating embeddings as fixed inputs and ignoring the role of query-time execution, making it difficult to decide whether a bottleneck should be addressed by changing the embedding model, the index, or the query-processing stages. To address the above limitation, we introduce \textbf{VecDBLens}, a modular attribution framework that treats the component under diagnosis as an \emph{evaluated object} and all remaining dataset, workload, parameter, and runtime factors as matched \emph{evaluation conditions}. VecDBLens reports stratified marginal effects, paired matched-condition differences, Flat-relative retention, and monotonicity audits, thereby turning vector database benchmarking from leaderboard ranking into a reproducible diagnostic procedure. VecDBLens reveals two key insights. First, retrieval performance is fundamentally upper-bounded by embedding quality, revealing that many apparent indexing failures are therefore representation-side limitations. Second, stronger index parameterization is not always beneficial: overly aggressive configurations can reshape neighborhood topology and cause recall regression. These results challenge the common practice of evaluating vector databases primarily through indexing and show that real-world limitations emerge from stage-wise constraints rather than any single module alone. VecDBLens provides a principled basis for identifying the true performance drivers of vector databases and for guiding their design in practical applications. The code is available at \url{https://anonymous.4open.science/r/VecDBLens-83A4}.
Vendi Anomaly Scores for Efficient and Accurate Anomaly Detection
Amey P Pasarkar ⋅ Adji Bousso Dieng
Anomaly detection (AD) is a critical task in machine learning and scientific discovery. Existing methods that operate on learned embeddings typically define anomalies through local density, isolation heuristics, or distance from a fitted distribution, making them sensitive to neighborhood scale, partitioning choices, or restrictive distributional assumptions. In this work, we introduce a new paradigm by formulating anomaly detection in terms of dataset diversity. We propose the Vendi Anomaly Score (VAS), which detects anomalies by quantifying how much a sample changes the diversity of the dataset, as measured by the Vendi Score, when removed. VAS captures both local redundancy and global data structure without relying on density estimation or partitioning heuristics. Moreover, VAS requires no hyperparameter tuning, instead adapting automatically to the spectral structure of the data. VAS is non-parametric and scales linearly with dataset size. Across 148 benchmark embedding datasets derived from 10 source datasets and 4 embedding architectures, VAS achieves state-of-the-art AD performance. We further validate VAS on large-scale ImageNet experiments, where it remains robust as both contamination rate and dataset scale increase.
Verifiable Environments Are LEGO Bricks: Recursive Composition for Reasoning Generalization
Hao Xiang ⋅ Qiaoyu Tang ⋅ Le Yu ⋅ Yaojie Lu ⋅ Xianpei Han ⋅ Ben He ⋅ Le Sun ⋅ Bowen Yu ⋅ Peng Wang ⋅ Hongyu Lin ⋅ Dayiheng Liu
Reinforcement Learning (RL) with verifiable environments has emerged as a powerful approach for enhancing the reasoning capabilities of Large Language Models (LLMs). While prior research indicates that scaling environment quantity improves RL performance, existing manual or individual construction methods suffer from linear scaling limits, thereby hindering scalable reasoning generalization. This paper introduces RACES (\textbf{R}ecursive \textbf{A}utomated \textbf{C}omposition for \textbf{E}nvironment \textbf{S}caling), a framework that treats verifiable environments as composable building blocks that can be recursively assembled. The key insight is that when the codomain (output type) of one environment matches the domain (input type) of another, they can be automatically fused into a new verifiable environment, enabling recursive composition. RACES is instantiated with 300 individual environments and defines a set of composition operators (Sequential, Parallel, Sort, and Select) that induce diverse reasoning patterns. Extensive experiments show that RL training on these composite environments consistently enhances reasoning generalization. Specifically, RACES improves DeepSeek-R1-Distill-Qwen-14B by an average of 3.1 points (from 48.2 to 51.3) and boosts Qwen3-14B performance from 58.8 to 61.1 on five benchmarks, which are unseen during the construction of training environments. Moreover, RACES achieves performance comparable to training on 300 individual environments using only 50 base environments, demonstrating significant efficiency in environment utilization.
Verified-Source Authority Is Not Generic Sycophancy: Cue-Family Decomposition of LLM Compliance
Abhinav Rajeev Kumar ⋅ Paras Chopra
Large language models change their answers when a "verified" source contradicts them. We study this verified-source authority bias across five open-weight families and three closed APIs. On a baseline-correct trivia subset, a verified-source cue endorsing a wrong answer flips 43–87% of responses across seven of the eight models, with a graded hierarchy over cue strengths. We then show this is not generic user sycophancy: across both open-weight families and closed APIs, the same wrong answer endorsed by a verified source vs. by a user produces different behavioral compliance, and a fitted "authority" vector is distinct from a generic assistant/instruction-following direction and from a user-sycophancy direction under causal projection-removal. Applying this vector to neutral prompts induces matched wrong-answer flips, and the same trivia-fit vector transfers without refitting to PIQA (a held-out task) and to multi-turn SYCON dialogues. Projecting the vector out at α=1 reduces wrong-source compliance by tens of percentage points on both trivia and PIQA in four of five open-weight families, while capability checks (MMLU-Pro, GSM8K) stay within noise of baseline at our evaluation sizes.
Verifying Agents in Rubric-Graded Environments
Markus Dücker ⋅ Vaibhav Kumar ⋅ Yi Liu ⋅ Ronak Chaudhary ⋅ Andreas Plesner ⋅ Francisco Guzmán ⋅ Anish Athalye
As AI agents take on increasingly open-ended and complex tasks, verifying their outputs has proven to be difficult. Rubric-graded agent environments—where a verifier judges multiple criteria against the deliverables and environment state changes produced by the agent—have emerged as a popular paradigm. We conduct the first systematic study of verifiers for rubric-graded agent environments. First, we create BankerVerifierBench (BVB), a meta-evaluation dataset of 3,204 human- judged rubric criteria across 21 investment-banking tasks. Next, we develop a methodology that derives from any rubric corpus the capabilities a verifier should possess, yielding a nine-capability taxonomy that we distill into three general design principles—reactive verification, environment alignment, and domain guidance— which we implement in Gandalf, an open-source verifier. Finally, we evaluate Gandalf against existing verifiers on BVB. We find that Gandalf outperforms all baselines on seven of nine capabilities and is Pareto optimal: its cheapest configuration (F1 0.633, \\$42) exceeds the most expensive baseline (F1 0.538, \\$414). Ablations show that environment alignment and domain guidance primarily reduce cost (up to 4× fewer tokens) rather than error, while architecture choice dominates model choice. Gandalf generalizes to an OpenClaw benchmark, a structurally different personal-productivity environment, achieving F1 score 0.951 and outperforming the next-best verifier by 7.5 F1 points.
VeriScope: Measuring Verification-Ready Verilog Artifacts
Wei Zhang ⋅ Jian Yang ⋅ jiajun wu ⋅ Junhang Cheng ⋅ Chufan He ⋅ Yihang Lou ⋅ Xianglong Liu
Large language models (LLMs) are increasingly used in hardware-oriented coding workflows, yet open resources for evaluating first-pass RTL beyond execution remain limited. We present \benchmark, an open benchmark and evaluation suite of \textbf{568 problems} spanning basic combinational gates through module-level designs and a 95-task L4+ top-end slice (92 L4 system tasks plus 3 L5 stress tests). A developer agent receives a natural-language task brief, returns a single RTL draft, and that artifact is evaluated through objective execution plus RTL-level and artifact-level review incorporating waveform and circuit evidence. Across \textbf{26 public models} plus two released 32B calibration baselines, the best systems reach about 84/100 combined score on the full benchmark but drop to the high 70s on L3--L4+. The 60/40 combined leaderboard largely tracks functional execution (Spearman 0.998 with a functional-only ranking), while artifact scores are most useful as a post-simulation screening signal: 29.6\% of L3--L4+ simulation-passing submissions receive artifact score below 6 and therefore need additional verification evidence. A 100-sample manual taxonomy of such disagreements finds \textbf{70\%} testbench coverage gaps, \textbf{12\%} real RTL defects, and \textbf{18\%} judge hallucinations (Wilson 95\% CIs roughly $\pm 9$\,pp), showing that artifact review surfaces both under-tested passes and model defects, but should not be treated as a standalone leaderboard refinement.
Veri-Sure: Multi-Agent RTL Code Generation with Temporal Tracing, Slicing and Formal Verification
Jiale Liu ⋅ Taiyu Zhou ⋅ Tianqi Jiang
In the rapidly evolving field of Electronic Design Automation (EDA), the deployment of Large Language Models (LLMs) for Register-Transfer Level (RTL) design has emerged as a promising direction. However, silicon-grade correctness remains bottlenecked by (i) limited test coverage and reliability of simulation-centric evaluation, (ii) regressions and repair hallucinations introduced by iterative debugging, and (iii) semantic drift as intent is reinterpreted across agent handoffs. In this work, we propose Veri-Sure, a multi-agent framework that establishes a design contract to align agents' intent and uses a patching mechanism guided by static dependency slicing to perform precise, localized repairs. By integrating a multi-branch verification pipeline that combines trace-driven temporal analysis with formal verification consisting of assertion-based checking and Boolean equivalence proofs, Veri-Sure helps improve functional correctness by enriching the simulation-driven debugging loop with localized diagnostic hints. We also introduce VerilogEval-v2-EXT, extending the original benchmark with 53 more industrial-grade design tasks and stratified difficulty levels, and show that Veri-Sure achieves state-of-the-art verified-correct RTL code generation performance, surpassing standalone LLMs and prior agentic systems. Code and dataset are available at https://anonymous.4open.science/r/Veri-Sure-D861.
VeriVul: A Verification-Guided Framework for Generating Realistic Vulnerability Benchmarks
Ahmed Lekssays ⋅ Hamza Mouhcine ⋅ Issa Khalil
Evaluating machine-learning-based vulnerability detectors requires datasets that pair accurate labels with realistic code. Existing benchmarks satisfy at most one. Real-world collections rely on noisy commit-based labels because manual curation by security experts does not scale, while verified synthetic collections generate code from scratch that is too simple to resemble production software. The rapid advance of code LLMs adds a third pressure, since their training corpora absorb public CVE data and most widely cited vulnerability benchmarks, leaving static evaluations of such models prone to measuring memorisation rather than analysis. As model knowledge cutoffs continue to move forward, any fixed dataset risks becoming part of the next generation's training set. We present VeriVul, a framework that produces formally verified vulnerability data on demand. Given a CVE-fixing commit, it applies a vulnerability-aware backward program slicer to isolate the fragment relevant to the vulnerability, prompts a code LLM to expand the fragment into a self-contained program, and certifies the resulting program with the ESBMC bounded model checker, recording the verified label together with the trace that triggers the violation. Each verified program is paired with an ESBMC-verified counterpart that flips the label. We characterise the synthesised programs along vulnerability-relevant structural dimensions (e.g., control-flow density, pointer and array operations, struct-field access, and identifier diversity) and show that VeriVul dataset samples lie distributionally closer to real-world vulnerable code than the prior verified synthetic baseline, while remaining sufficiently complex to challenge state-of-the-art LLM-based detectors. In a prompted evaluation, Claude Opus 4.7 reaches only 52.6 percent accuracy on VeriVul dataset despite being given the full self-contained program. Because each sample carries the ESBMC counterexample and the violating statement, VeriVul also supports evaluation beyond binary classification, including root-cause identification and trace-grounded explanation. We release the framework, the slicer, the verification harness, and the generated dataset to support contamination-resilient evaluation of code LLMs as vulnerability analysts.
Vermeer: Autoregressive generative modeling of microscopy predicts protein localization
Sandeep Kambhampati ⋅ Eric Zimmermann ⋅ Emre Hayir ⋅ Kevin K Yang ⋅ Fei Chen ⋅ Alex X Lu
Fluorescent microscopy provides a rich view into how proteins localize within cells, but it remains experimentally infeasible to image human proteins across all of the different factors that can impact localization. We introduce Vermeer, a channel-adaptive autoregressive generative model for in silico generation of microscopy images of protein localization. Vermeer conditions generations on protein sequences and landmark stains showing the morphology of cells, which enables it to generalize to unseen proteins and cell lines. We show that Vermeer, trained on the Human Protein Atlas, can generate images with substantially improved perceptual quality and biological fidelity over previous proposals. Additionally, Vermeer's autoregressive framework enables flexible generation using varying channel subsets and orderings, enabling zero-shot transfer to data collected under different imaging conditions and channel configurations than those used for training. These results position Vermeer to enable scalable modeling of protein localization and is a step towards generative foundation models that can operate over distinct microscopy datasets.
VersaCamVLA: Camera-Configurable VLA Policies for Robotic Manipulation
Boyao Han ⋅ Chen Shi ⋅ Jingjing Qian ⋅ Zhuotao Tian ⋅ Li Jiang
Vision-Language-Action (VLA) models have emerged as powerful foundations for robotic manipulation, but their reliance on fixed camera configurations during training makes them brittle to changes in camera count or pose during deployment. To overcome these limitations, we propose VersaCamVLA, a camera-configurable framework that decouples camera-set representation from action learning. VersaCamVLA learns a unified scene-token interface that maps an arbitrary, variable set of posed RGB views into fixed-size latent scene tokens. This is achieved via multi-signal target-view prediction and Wrist-Augmented Pose Sampling (WAPS), which leverages natural wrist-camera motion for free pose diversity. At deployment, a lightweight spatial encoder injects these compact scene tokens into a pretrained base VLA as a supplementary visual condition, requiring no explicit 3D sensing or novel-view rendering. Experiments on RoboTwin, LIBERO, and a real-robot platform demonstrate that VersaCamVLA consistently outperforms prior VLA methods and direct multi-view baselines, maintaining robust performance across varying camera counts and unseen camera poses.
VI-Bench: Benchmarking Prompt Inversion from AIGC Videos
Wulin Xie ⋅ Rui Zhao ⋅ Kecen Li ⋅ Xiujin Liu ⋅ Bokang Zhang ⋅ Zheng Liu ⋅ Xinwen Hou ⋅ Chen GONG
Recent advances in video generation have made prompt-based control increasingly central to AIGC video generation. Prompts specify what a video should depict and how it should be represented, controlling factors such as visual style or camera behavior. Understanding this recoverability is important both for creative reuse and editing, and for assessing prompt leakage risks. However, existing video understanding benchmarks do not measure this capability: a caption may describe what is visible, but a replayable prompt must recover the generation-relevant controls needed to reproduce the video. To address this gap, we introduce VI-Bench, a benchmark built from 16.1 million real-user prompts and 900 human-verified AIGC videos. VI-Bench spans three progressively harder settings, namely single-shot semantic grounding, control over style and camera behavior, and multi-shot compositional inversion, and evaluates five generation-critical dimensions: subject, action, scene, style, and camera. We evaluate 18 representative VLMs, including 2 proprietary and 16 open-source models on VI-Bench, using an Inversion Score that measures prompt-level alignment with the original prompt and video-level fidelity of the regenerated video. The results reveal substantial limitations: even the strongest model achieves only 0.632 on Inversion Score, performance degrades sharply as samples require richer control and multi-shot reasoning, and models often produce plausible prompts whose regenerated videos deviate from the reference. These findings show that video prompt inversion is a distinct and under-evaluated capability requiring models to transform visual understanding into replay-stable generative control.
Multimodal queries can require different types of reasoning. Some can be answered via perceptual reasoning, extracting information directly from the visual signal, while others require compositional reasoning that combines observations or deliberative reasoning that evaluates competing hypotheses. However, many existing methods apply a uniform reasoning strategy across queries, leading to unnecessary computation on simple tasks and insufficient reasoning on complex ones. We introduce Video-FLAIR, a training framework that learns to select the appropriate reasoning mode for each query using reinforcement learning. During training, the model generates responses under all three modes for the same prompt, enabling direct comparison. A composite reward compares these responses to favor the most effective one based on correctness, grounding, and cost, while discouraging unsupported or misaligned deliberation. This yields a supervision signal for learning adaptive reasoning without per-query annotations. Video-FLAIR improves accuracy over the Qwen2.5-VL base model by +5.4 on MathVista, +4.8 on Video-Holmes, and +4.8 on Video-MMMU, while reducing average token usage to 95 compared to 417 for always-thinking baselines.
VideoMaMa++: Temporally Consistent Video Matting via Preserve-and-Refine Tokens
Sangbeom Lim ⋅ Seoung Wug Oh ⋅ Heeji Yoon ⋅ Seungryong Kim ⋅ Joon-Young Lee
Diffusion-based video matting has recently emerged as a leading approach, leveraging video diffusion priors from internet-scale pretraining to achieve strong zero-shot generalization from limited synthetic data. However, the heavy compute of these diffusion backbones limits inference to a short chunk of just 6--24 frames, so even ordinary videos span multiple chunks. Because each chunk is matted independently from coarse mask guidance, existing methods fail to maintain consistency across chunk boundaries, producing flickering edges and discontinuous transparency. We address this with a mask-guided diffusion-matting framework built on two ideas. The RGB and mask conditioning are represented as separate token streams that interact through full spatio-temporal self-attention, letting the model adaptively combine fine RGB texture with coarse mask layout rather than committing to a fixed channel-wise fusion. On top of this, we propose a novel pair of learnable embeddings, the preserve token and the refine token, that act as a per-frame conditioning interface and enable a sliding-window inference scheme in which already-generated mattes are propagated across chunk boundaries for local temporal consistency, complemented by a globally shared reference anchor at every chunk to prevent error accumulation in long sequences without any additional training. Across standard video matting benchmarks and in-the-wild long footage, our framework outperforms prior diffusion-based methods in both fine-detail accuracy and temporal consistency, and uniquely maintains consistency across chunk boundaries on long videos. Our code and weights will be publicly released.
Video-Mirai: Autoregressive Video Diffusion Models Need Foresight
Yonghao Yu ⋅ Lang Huang ⋅ Runyi Li ⋅ Zerun Wang ⋅ Toshihiko Yamasaki
Causal video generators must predict from the past, but they need not learn only from it. In streaming autoregressive video diffusion, each emitted segment becomes a commitment that future segments must preserve. Standard training, however, only asks each causal state to explain the present. This creates what we call a representation-level planning gap: states that fit the current segment may discard identity, layout, and motion information needed for a consistent future. We introduce Video-Mirai, a training-only method that closes this gap without changing causal inference: the generator rolls out causally, a frozen foresight encoder reads the completed rollout non-causally, and a lightweight predictor distills the resulting stopped-gradient targets into causal states. Future frames supervise representations, never generator inputs. At inference, the encoder and predictor are discarded, leaving the original architecture, per-step FLOPs, and KV-cache behavior unchanged. Video-Mirai improves a strong Causal-Forcing baseline on 5-second VBench from 83.8 to 84.6 in terms of Total Score. On 30-second rollouts beyond the training horizon, subject consistency improves from 84.9 to 88.5 and background consistency from 90.2 to 91.9. Ablations identify future-conditioned targets as the key ingredient, and probes show that future frames become more decodable from current features. Causality should constrain inference, not representation supervision.
VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding
He Haichen ⋅ Jiayi Zhou ⋅ Sifeng SHANG ⋅ Yihan Hu ⋅ Yuanhan Zhang ⋅ Kaiyang Zhou
Real-world long video understanding requires models to perform continuous tracking, information integration and memory retention over massive temporal spans within extreme video durations. Mastering this intense cognitive load constitutes the fundamental bottleneck in long video understanding. While existing benchmarks have driven progress by scaling up video duration, their evaluation tasks often require comprehending only short and isolated video segments, falling short of capturing the challenge of ultra-long-context reasoning. To measure this cognitive load, we emphasize continuous certificate length, defined as the video length a human must continuously watch to definitively answer a given question. Driven by this metric, we introduce VideoOdyssey, a benchmark specifically designed for ultra-long-context and omni-modal video understanding. VideoOdyssey is characterized by three key features: 1) Extreme video duration and diversity: spanning 11 domains and 54 subcategories with an average video duration of 109 minutes; 2) Comprehensive evaluation scenarios: offering two subsets to address different research focuses, i.e., VideoOdyssey-V for probing the limits of visual understanding in MLLMs, and VideoOdyssey-AV for evaluating synchronized audio-visual understanding for omni-modal models; 3) Ultra-long and multi-level continuous certificates: extending the average continuous certificate to 16 minutes for VideoOdyssey-V and 12.8 minutes for VideoOdyssey-AV. Crucially, we design 5 granular levels from seconds to hours, providing a comprehensive diagnostic tool to evaluate models across varying context lengths and cognitive loads. Extensive evaluations show that bottlenecks of current MLLMs extend beyond simple retrieval to include struggles with continuous reasoning across varying context lengths, fine-grained perception, and non-verbal omni-modal understanding. We hope VideoOdyssey will spur the development of next-generation MLLMs toward genuine real-world video understanding.
VisionCreator-R1: A Reflection-Enhanced Native Visual-Generation Agentic Model
Jinxiang Lai ⋅ Wenzhe Zhao ⋅ Zexin Lu ⋅ HUALEI Zhang ⋅ Qinyu Yang ⋅ Rongwei Quan ⋅ li zhimin ⋅ Shuai Shao ⋅ Song Guo ⋅ Qinglin Lu
Visual content generation has moved from single-image to multi-image workflows, yet existing agents are largely plan-driven and lack a systematic reflection mechanism for correcting mid-trajectory visual errors. To address this gap, we propose VisionCreator-R1, a native visual generation agent with explicit reflection, paired with a Reflection-Plan Co-Optimization (RPCO) training methodology. Through extensive experiments and trajectory-level analysis, we uncover a reflection-plan optimization asymmetry in reinforcement learning (RL): planning can be reliably optimized via plan rewards, while reflection learning is held back by noisy credit assignment. Motivated by this finding, we further abstract a general Decouple-then-Fuse paradigm for co-optimizing capabilities with asymmetric reward variance, of which RPCO is the visual-generation instantiation. Our RPCO first trains on the self-constructed VCR-SFT dataset, which covers reflection-strong single-image trajectories and planning-strong multi-image trajectories, and then co-optimizes on the VCR-RL dataset via RL. This yields our unified VisionCreator-R1 agent, which beats Gemini2.5Pro on existing benchmarks and on our VCR-Bench covering single-image and multi-image tasks.
VisionCreator-S1: Evolving Visual-Generation Agents via Skill-GRPO Optimization
Jinxiang Lai ⋅ Wenzhe Zhao ⋅ HUALEI Zhang ⋅ Zexin Lu ⋅ Jian Liu ⋅ Jie ZHANG ⋅ Qinglin Lu ⋅ Song Guo
Native visual-generation agents now handle long-horizon, multi-image workflows, but face two practical bottlenecks during reinforcement learning (RL) training. First, the inability to accumulate experience across rollouts prevents the agent from leveraging past successes and failures, slowing the evolution of its agent behaviors. Second, the reliance on black-box visual-generation tools introduces tool-induced reward noise, as the policy is unfairly penalized for failures rooted in a tool's inherent capability boundary, leading to unstable and inefficient training. To address these challenges, we propose VisionCreator-S1, an agent that casts skill-augmented RL as a nested optimization problem through two coupled innovations. First, Dual-Skills co-evolves an agent-skill set to reuse self-reflection patterns from historical rollouts, and a tool-skill set that dynamically maps each tool's capability boundary through world-reflection. Second, Skill-GRPO treats discrete text skills as learnable variables. To overcome the non-differentiability of text skill, it gracefully alternates between numerical gradient descent for the policy weights and textual gradient descent for the discrete skill sets. By contrasting high- and low-advantage trajectories, it approximates a direction-consistent textual gradient in the agent's behavior feature space. Extensive experiments on VCR-Bench, GEdit-Bench, and MultiBanana demonstrate that VisionCreator-S1 consistently outperforms the strongest skill-free baseline, VisionCreator-R1. Ablation studies further confirm that this aligned co-evolution significantly improves both training stability and the agent's long-term decision-making capabilities.
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
Qianhao Yuan ⋅ Jie Lou ⋅ XingYu Li ⋅ Hongyu Lin ⋅ Le Sun ⋅ Xianpei Han ⋅ Yaojie Lu
Multimodal Large Language Models (MLLMs) still struggle with fine-grained visual understanding, where answers often depend on small but decisive evidence in the full image. We observe a regional-to-global perception gap: the same MLLM answers fine-grained questions more accurately when conditioned on evidence-centered crops than on the corresponding full images, suggesting that many failures stem from difficulty to focus on relevant evidence rather than insufficient local recognition ability. Motivated by this observation, we propose Vision-OPD (Vision On-Policy Distillation), a regional-to-global self-distillation framework that transfers the model's own privileged regional perception to its full-image policy. Vision-OPD instantiates two conditional policies from the same MLLM: a crop-conditioned teacher and a full-image-conditioned student. The student generates on-policy rollouts, and Vision-OPD minimizes token-level divergence between the teacher and student next-token distributions along these rollouts. This enables the model to internalize the benefit of visual zooming without external teacher models, ground-truth labels, reward verifiers, or inference-time tool use. Experiments on multiple fine-grained visual understanding benchmarks show that Vision-OPD models achieve competitive or superior performance against much larger open-source, closed-source, and "Thinking-with-Images" agentic models.
Visual Expert Skipping for MoE-MLLMs via K-Armed Bandit based Expert Estimation
zhou hongbo ⋅ Jiapeng Zhang ⋅ Yuqi Liu ⋅ Qiong Wu ⋅ Jun Peng ⋅ Xiao Chen ⋅ Yiyi Zhou ⋅ Rongrong Ji
Despite great success, existing MoE-MLLMs still suffer from substantial visual expert redundancy, making visual expert skipping a critical solution. However, deriving an effective and fixed skipping policy is still an open problem, mainly due to the exponentially large layer-wise policy space and the dependency between expert allocations across layers. In this paper, we propose a novel and training-free approach for MoE-MLLMs, termed \emph{K-armed bandit based expert redundancy estimation} (\textbf{KAB-MoE}). KAB-MoE defines the retained visual expert number at each MoE layer as a discrete action, and then conducts efficient one-shot RL estimations to obtain the layer-action reward matrix. Based on this reward matrix, KAB-MoE can quickly derive the optimal skipping policy that satisfies the predefined computation budgets, \emph{e.g.}, expert skipping ratio. The obtained skipping policy can be directly applied to MoE-MLLMs for all examples, which not only facilitates high-throughput deployment but also contributes to practical acceleration. Extensive experiments on Kimi-VL-A3B and Qwen3-VL-MoE show that KAB-MoE effectively accelerates MoE-MLLMs while keeping strong multimodal performance, \emph{e.g.}, retaining 97.09\% average performance on Kimi-VL-A3B under the skipping ratio of 83\%. Moreover, KAB-MoE also achieves competitive or superior performance over dynamic expert skipping SOTAs with practical inference speed-up. \textbf{Our code} is provided in the supplementary materials.
Large audio language models (LALMs) have recently demonstrated strong capabilities in tasks requiring paralinguistic and non-linguistic understanding of audio. However, they still struggle to reliably utilize such information in tasks such as emotion recognition. Existing benchmarks often entangle acoustic perception with higher-level reasoning, making it difficult to identify the source of these failures. To address this issue, we introduce \textsc{VocalGrad}, a benchmark designed to evaluate the perceptual ability of acoustic features. VocalGrad formulates this problem as a simple binary classification of temporal change direction, enabling evaluation largely independent of higher-level reasoning or semantic understanding. Our benchmarking results show that, despite high accuracy by human annotators, state-of-the-art LALMs perform near chance level, revealing substantial limitations in their acoustic perception ability. We further conduct linear probing to identify structural bottlenecks in these models. Our analysis reveals two distinct failure modes: for some categories, task-relevant information is preserved in internal representations but not utilized by the language model head; for others, the information progressively diminishes across model layers. Overall, our results reveal fundamental limitations of LALMs in modeling acoustic variation and expose representation differences not observable from text outputs.
Voice "Cloning" Is Actually Style Transfer
Kaitlyn Zhou ⋅ Federico Bianchi ⋅ Martijn Bartelds ⋅ Anna Pot ⋅ Yongchan Kwon ⋅ James Zou
Artificially generated speech is increasingly embedded in everyday life. Voice cloning in particular enables applications where identity preservation is important, such as completing a recording, dubbing in a new language, or preserving the voices of individuals with speech loss. However, in our work, we find that despite the term, voice cloning does not faithfully "clone" an individual's voice. Instead, we find that widely-used voice cloning models systematically apply style transfer to source voices. As rated by human annotators, cloned voices are perceived as more authoritative, warm, customer-service-like, and human-like compared to their sources. Human annotators also report greater trust in cloned voices than source voices, and a greater willingness to disclose sensitive personal information to them. Our work furthermore shows that voice cloning leads to homogenization of speaker characteristics, as measured by reduced variance in accent, speaking rate, and the audio embedding space. Together, our results highlight a new set of limitations and risks of voice cloning technology and their potential impact on human behavior.
WaveletLoRA: Frequency-Aware Content-Style Decomposition for Personalized Image Generation
Peiyao Wang ⋅ Jiahui Sun ⋅ Weining Wang ⋅ Jing Liu
Content-style decomposition from a single image is a fundamental challenge in personalized image generation, aiming to separate subject identity from visual style for flexible recomposition. Existing approaches primarily operate in the spatial domain, without explicitly exploiting the distinct frequency characteristics naturally associated with content and style. In this paper, we propose WaveletLoRA, a novel frequency-aware framework for explicit content-style decomposition through wavelet-domain supervision. Specifically, we first apply a multi-level discrete wavelet transform (DWT) to decompose the reference image into low-frequency and high-frequency subbands. We then assign a dedicated LoRA branch to each subband, where the low-frequency branch models global appearance for style acquisition, while the high-frequency branches capture structural details for content preservation. To further improve disentanglement quality, we introduce a dense Mixture-of-Experts aggregation module together with a frequency-domain regularization objective, which reduces style leakage and enhance content fidelity. In addition, we establish WaveBench and develop a VLM-based evaluation protocol that explicitly measures both attribute similarity and attribute separation for disentanglement assessment. Extensive experiments demonstrate that WaveletLoRA consistently outperforms existing methods and produces results that better align with human preference. Code and the proposed dataset will be released in https://anonymous.4open.science/r/WaveletLoRA-8D7F.
WAXAL: A Large-Scale Multilingual African Language Speech Corpus
MohamedElfatih MohamedKhair ⋅ Emmanuel Asiedu Brempong ⋅ Subhashini Venugopalan ⋅ Abdoulaye Diack ⋅ Perry Nelson ⋅ Mireku ⋅ Tavonga Siyavora ⋅ Pooja Rao ⋅ Uche Okonkwo ⋅ Angela Nakalembe ⋅ Abhishek Bapna ⋅ Aisha Walcott-Bryant
The advancement of speech technology has predominantly favored high-resource languages, creating a significant digital divide for speakers of most Sub-Saharan African languages. To address this gap, we introduce WAXAL, a large-scale, openly accessible speech dataset for 24 languages (spanning 19 for ASR and 13 for TTS, with overlaps) representing over 100 million speakers. The collection consists of an Automatic Speech Recognition (ASR) dataset containing approximately 2,340 hours of transcribed, image-prompted natural speech (over 14,100 hours of total collected audio) from a diverse range of speakers, and a Text-to-Speech (TTS) dataset with over 235 hours of high-quality, single-speaker studio recordings of phonetically balanced scripts. Crucially, this effort was executed in partnership with four African academic and community organizations, with the explicit goal of building the local ecosystem for speech technology and seeding data collection capabilities within these institutions. To demonstrate the dataset's utility, we benchmark three architecturally distinct ASR models---Gemma 3n-2B, Whisper-Large-v3, and MMS-1b-all---across the 19 ASR languages. Fine-tuning on WAXAL yields substantial improvements, reducing the macro-average Word Error Rate (WER) by up to 61\% (e.g., from 1.23 to 0.48 for Whisper). Models fine-tuned solely on WAXAL also generalize to the independently collected FLEURS benchmark, confirming the dataset's value for domain-transferable ASR adaptation. Furthermore, we employ an LLM-as-judge framework to evaluate semantic meaning preservation, demonstrating that fine-tuned predictions successfully capture core intents of native speakers. The WAXAL datasets are released at https://huggingface.co/datasets/google/WaxalNLP under the CC-BY-4.0 license to catalyze research and the development of inclusive speech technologies for speakers of African languages.
Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation
Andreas Haupt ⋅ Justin Hartenstein ⋅ Anka Reuel-Lamparth ⋅ Mykel J Kochenderfer ⋅ Sanmi Koyejo
AI benchmarks have well-documented limitations, with prior work examining contamination, saturation, and construct underspecification. Aggregation has received far less attention: benchmarks are typically summarized by uniformly averaging item-level scores, implicitly treating every test item as equally valuable. We model benchmarking as a multitask principal–agent game and show that the welfare loss from a benchmark is determined jointly by three item-level primitives: alignment with normative welfare priorities, marginal improvability, and performance variance. We translate the theory into an audit framework that ranks items along each of these three axes, and apply it to OLMES items using WORKBank for welfare, the EvoLM 4B suite for improvability, and the PolyPythias 410M panel for variance. The framework surfaces items that are Pareto-inferior within OLMES subject to a pro-worker welfare operationalization.
What do EEG Foundation Models Capture from Human Brain Signals?
Ling Tang ⋅ Qian Chen ⋅ Jilin Mei ⋅ Houshi Xu ⋅ Quanshi Zhang ⋅ Jing Shao ⋅ Na Zou ⋅ Xia Hu ⋅ Dongrui Liu
Clinical electroencephalogram (EEG) analysis rests on a hand-crafted feature catalog refined over decades, *e.g.,* band power, connectivity, complexity, and more. Modern EEG foundation models bypass this catalog, learn directly from raw signals via self-supervised pretraining, and match or outperform feature-engineered baselines on most clinical benchmarks. Whether the two representations align is an open question, which we decompose into three sub-questions: *what does the model learn*, *what does the model use*, and *how much can be explained*. We answer them with layer-wise ridge probing, LEACE-style cross-covariance subspace erasure, and a transparent classifier benchmarked against a random-feature baseline. The audit covers three foundation models (CSBrain, CBraMod, LaBraM), five clinical tasks (MDD, Stress, ISRUC-Sleep, TUSL, Siena), and a 6-family 63-feature lexicon. Of the $945$ (model, task, feature) units, $648$ ($68.6\%$) are representation-causal and $199$ ($21.1\%$) are encoded-only. Across tasks, $50$ features qualify as universal candidates with strong support (all three architectures RC) in two or more tasks. Frequency-domain features dominate, but the other five families each contribute substantial causal mass. Confirmed features recover, on average, $79.3\%$ of the foundation model's advantage over the random baseline, with a clean task gradient (MDD $\approx 0.99$ down to Stress $\approx 0.56$): tasks near ceiling are almost fully recovered by the lexicon, while harder tasks leave a non-trivial residual that pinpoints a concrete target for future concept discovery.
What does a Bayes-filtered transformer believe? A predictive Monte Carlo approach
Afiq Abdillah Effiezal Aswadi ⋅ Haotong Ma ⋅ Susan Wei
A \emph{Bayes-filtered transformer} (BFT) is a transformer trained on sequences that are generated in two steps: first a latent task is drawn from a prior, then observations are drawn conditional on that task. Trained under autoregressive log loss, the BFT's next-token prediction, in the idealised limit, \emph{is} the Bayesian posterior predictive distribution (PPD) under this generative model. In practice the trained BFT is only an approximation of this ideal PPD, raising a natural interpretive question: what prior and posterior over the latent task has the trained BFT actually internalised? Existing work answers this question by comparing the trained BFT's predictions against the predictions of various ``reference'' posteriors, each standing in for a different candidate algorithm or computation the BFT might be implementing. This prediction-space comparison is fragile; for instance, distinct posteriors can share posterior-mean predictions exactly. We foreground \emph{predictive Monte Carlo} (PMC) as a general interpretability tool for any BFT: using only next-token generation, PMC returns an approximation to the implicit prior and posterior over the latent task, answering the interpretive question directly in latent space. We apply PMC to three stylised task families spanning 0-Markov and 1-Markov exchangeability; the phenomena previously reported in these settings remain visible in latent space.
What Does the AI Doctor Value? Auditing Pluralism in the Clinical Ethics of Language Models
Payal Chandak ⋅ Victoria Alkin ⋅ David Wu ⋅ Maya Dagan ⋅ Taposh D Roy ⋅ Maria Clara Saad Menezes ⋅ Ayush Noori ⋅ Nirali Somia ⋅ John S Brownstein ⋅ Ran Balicer ⋅ Rebecca W Brendel ⋅ Noa Dagan ⋅ Isaac S Kohane ⋅ Gabriel A Brat
Medicine is inherently pluralistic. Principles such as autonomy, beneficence, nonmaleficence, and justice routinely conflict, and such ethical dilemmas often sharply divide reasonable physicians. Good clinical practice navigates these tensions in concert with each patient's values rather than imposing a single ethical stance. The ethical values that large language models bring to medical advice, however, have not been systematically examined. We present a framework for auditing value pluralism in medical AI, comprising a benchmark of clinician-verified dilemmas and an attribution method that recovers value priorities directly from decisions. The ecosystem of frontier models spans physician-level value heterogeneity, and models discuss competing values in their reasoning (Overton pluralism) before committing to a decision. However, individual model decisions are near-deterministic across repeated sampling and semantic variations, failing to reproduce the distributional pluralism of the physician panel. Across benchmark cases, these consistent decisions reflect committed, systematic value preferences. While most model priorities fall within the natural range of inter-physician variation, some significantly underweight patient autonomy. A single LLM deployed without regard for its value priorities could amplify those priorities at scale to every patient it serves. Without explicit efforts to balance ethical perspectives with one or multiple models, these tools risk replacing clinical pluralism with a deployment monoculture.
What Drives Test-Time Adaptation for CLIP? A Controlled Empirical Study from an Update Perspective
Jiazhen Huang ⋅ Xiao Chen ⋅ Zhiming Liu ⋅ Yaru Sun ⋅ Jingyan Jiang ⋅ Zhi Wang
Vision-Language Models (VLMs) such as CLIP have become a standard backbone for open-vocabulary recognition, yet their zero-shot predictions remain vulnerable to distribution shifts encountered at deployment. Test-Time Adaptation (TTA) has recently been extended to CLIP as a lightweight solution, leading to a rapidly growing body of TTA4CLIP methods. However, empirical progress in this area has largely outpaced our understanding of what truly drives adaptation, where their gains originate, and under which shifts they remain reliable. In this paper, we take a step back from the pursuit of state-of-the-art accuracy and conduct a systematic controlled study of TTA4CLIP. We first organize existing methods into three unified paradigms according to what is updated at test time. We then introduce TTABC, an open-source TTA Benchmark for CLIP, which standardizes evaluation protocols and integrates more than 20 representative methods. Our controlled empirical analysis focuses on three key areas. First, we determine the driving factors in parameter-based methods, revealing that adaptation gains are primarily driven by test-time evidence and reliable proxies rather than heavy optimization. Second, we explore evidence utilization beyond heavy parameter tuning, showing that competitive and efficient performance can be achieved through cross- or current-sample evidence and lightweight prototype updates. Finally, we demonstrate that there is no silver bullet for TTA; no single adaptation paradigm is universally optimal, and the preferred paradigm depends on the nature of shift. We hope our benchmark and study provide a clearer understanding of the current TTA4CLIP landscape and establish a foundation for further research.
What Gets Measured Gets Managed: Sign-aware Recommendation Needs Sign-aware Evaluation
Minchan Kim ⋅ Jungmin Hwang ⋅ Hyunwoo Park
Sign-aware recommender systems have recently been developed to leverage negative feedback for a deeper understanding of user preferences. However, our empirical diagnosis reveals that state-of-the-art graph-based sign-aware recommender systems are paradoxically *valence-blind*. Even though they explicitly incorporate sign information during training, they consistently fail to differentiate liked items from disliked ones at the ranking stage, frequently infiltrating top-$K$ recommendations with disliked content. Through linear probing, we show that while valence information exists in the learned embeddings, it remains inaccessible to the inner-product scoring function. This widespread failure remains entirely undetected because conventional evaluation metrics, such as Recall, HR, and NDCG, assign a uniform utility of zero to both negative and unobserved items, creating a systematic evaluation blind spot. To bridge this gap, we propose a family of signed metrics, Signed Recall, Signed HR, and Signed NDCG, that explicitly penalize the recommendation of disliked content. Systematic re-evaluation under our proposed metrics fundamentally reshapes the established performance landscape, revealing that methods ranked highly under conventional metrics often fail to protect users from disliked content. Finally, through a proof-of-concept auxiliary loss, we confirm that the proposed metrics provide actionable training signals, guiding models toward valence-aware behavior without sacrificing conventional relevance. For transparency, our source code is available at: https://anonymous.4open.science/r/signed-rec-benchmark-07E4
What MLLMs Learn about When they Learn about Multimodal Reasoning
Jiwan Chung ⋅ Neel Joshi ⋅ Pratyusha Sharma ⋅ Youngjae Yu ⋅ Vibhav Vineet
Evaluation of multimodal reasoning models is typically reduced to a single accuracy score, implicitly treating reasoning as a unitary capability. We introduce MathLens, a benchmark of textbook-style geometry problems that exposes this assumption by operationally decomposing performance into perception, reasoning, and multimodal-specific components. Each problem is derived from a symbolic specification and accompanied by visual diagrams, text-only variants, multimodal questions, and targeted perceptual probes, enabling controlled measurement of each component. Using this decomposition, we show that common training strategies induce systematically different capability profiles that are invisible under aggregate accuracy. Reinforcement learning primarily improves perceptual grounding and robustness to diagram variation, while textual SFT yields gains through reflective reasoning. In contrast, as perception and reasoning improve, a growing fraction of remaining errors fall outside these components and are categorized as multimodal-specific. These results suggest that apparent progress in multimodal reasoning reflects shifting balances among subskills rather than uniform advancement, motivating evaluation beyond scalar accuracy.
What Remains in Sight? Autoregressive Video Decoding as Representation-Guided Context Rewriting
Lixuan He ⋅ Jihyeon Je ⋅ Yang You ⋅ Leonidas Guibas
Long-horizon video generation is increasingly moving toward autoregressive rollout, where each newly generated segment becomes part of the evidence used for future prediction. Recent causal decoders, long-context training, and KV-cache mechanisms have extended feasible duration, but most inference pipelines still decide the visible past implicitly through recent windows, compressed caches, or external memory. This raises a direct question: can the choice of which past frames or segments to show to the decoder be treated as an explicit test-time decision? We answer this question with \textsc{ReCR}, a training-free method that formulates autoregressive video decoding as representation-guided context rewriting. At each step, \textsc{ReCR} builds a selected visible context from the generated history by using internal decoder representations to favor context units that match the global rollout state, avoid redundancy with the recent continuation, preserve temporal boundary coverage, and provide useful intermediate evidence. The selected units are then mapped onto a compact causal axis before decoding the next unit. Across long-horizon text-to-video benchmarks, \textsc{ReCR} consistently improves multiple autoregressive backbones under matched visible-token budgets. Empirical analyses such as representation-space diagnostics and transition-count prompt switching further show that improving what remains visible is an effective path toward more stable long-video decoding. Our code is available at https://anonymous.4open.science/r/anonymous-ReCR-97EB/
What Transformer FFNs Never See: Theory, Diagnosis, and Lightweight Remediation
Tinghe Zhang ⋅ Yucheng Xiao ⋅ Alex Lamb
In Transformer attention, distinct weight vectors $\boldsymbol{\alpha} \neq \boldsymbol{\alpha}'$ can produce the same aggregated representation $\mathbf{V}\boldsymbol{\alpha} = \mathbf{V}\boldsymbol{\alpha}'$. When $\operatorname{rank}(\mathbf{V}) \leq n-2$, configurations with entirely different dominant source tokens collide on a set of positive Lebesgue measure, a condition satisfied in over 97% of attention heads in BERT-Base. Such collisions are harmful: the downstream FFN receives identical inputs despite the attention having attended to entirely different source tokens, and must therefore produce identical outputs regardless of which tokens actually dominated, fundamentally limiting any computation that requires sensitivity to attention source, such as multi-hop reasoning, attribution, or knowledge retrieval. No increase in depth, width, or data can compensate, because the routing weights are discarded before the FFN is reached. The routing weights $\{\alpha_{ij}\}$, however, are available within the same forward pass and can be retained as a compact side-channel to restore the missing signal. We call this structural information loss *routing non-identifiability*. In BERT-Base, value-matrix rank deficiency suppresses $k \geq 107$ routing dimensions (over 84% of degrees of freedom) before any FFN computation. To quantify the practical gap, we introduce the Routing Reconstruction Task (RRT): standard Transformers achieve 31.4% source-identification accuracy while oracle routing statistics reach 94.6%, a 63-point gap consistent with an information bottleneck in the FFN input. We then propose the Route-Aware FFN (RA-FFN), a drop-in module that appends four compact routing statistics to each FFN block at under 1.1% additional parameter cost, recovering over 75% of the oracle gain. RA-FFN improves all benchmarks tested: at 110M scale, gains cover multi-hop reasoning, language modelling, and translation, with multi-hop gains $6\times$ larger than single-hop; at 7B scale on Llama-2-7B and Mistral-7B, gains extend further to mathematics (GSM8K) and general reasoning (MMLU, ARC-Challenge, HellaSwag), ordered by routing sensitivity.
What Was That Again? Certified Robustness for Automatic Speech Recognition
Andrew Cullen ⋅ Neil Marchant ⋅ Jiani Xie ⋅ Paul Montague ⋅ Benjamin Rubinstein
Automatic Speech Recognition systems are notoriously both sensitive to input perturbations and challenging to defend, which is a product of their high-dimensional, discrete output spaces. Traditional Randomized Smoothing workflows collapse in sequence-to-sequence tasks because the probability mass of any single transcription vanishes under noise. We propose an Anytime-Valid Certified Transcription framework that replaces fixed-sample binomial testing with E-value Martingales. Our approach leverages a dual-gate pipeline: a Two-Sided Atomic Audit that accumulates statistical wealth to certify both token existence and adversarial exclusion, and a Rank-Based Tournament that selects the winning sequence. Our evaluations across four diverse architectures demonstrate that this approach yields an average 25.5\% relative reduction in Word Error Rate, while also providing granular word- and sentence-level certifications to enhance acoustic security.
When and Why Does Multi-Agent Debate Fail and Does It Really Underperform?
Yongqiang Chen ⋅ Gang Niu ⋅ James Cheng ⋅ Bo Han ⋅ Masashi Sugiyama
Multi-agent debate (MAD) was proposed as a promising approach for ensembling the wisdom of multiple large language models (LLMs) to improve reasoning and provide effective supervision to superhuman LLMs. However, increasing empirical evidence suggests that MAD may not outperform or even significantly underperform single-agent approaches (SA), raising doubts about the benefits of MAD. In this work, we investigate this issue by analyzing the incentive structures of popular MAD paradigms: (i) competitive MAD (CopMAD) where agents compete by holding opposing positions; (ii) consensus-seeking MAD (CosMAD) where agents are driven to seek consensus. We show that both paradigms suffer from debate hacking: CopMAD reduces to a cheap-talk game, where agents produce misleading messages to win the game, while CosMAD filters out informative disagreements for premature consensus. Consequently, agents in both CopMAD and CosMAD fail to jointly resolve the ambiguity and seek the truth. To this end, we introduce ColMAD, a collaborative protocol that reframes MAD as a non-zero-sum game to encourage agents to provide informative while truthful messages. Through extensive benchmarking on challenging tasks such as error detection, we show that ColMAD significantly outperforms previous MAD protocols up to 10 percentage points. Under the same budgets, ColMAD effectively brings non-trivial improvements over SA methods, implying that the protocol design is critical to realizing the potential of MAD.
When Can Digital Personas Reliably Approximate Human Survey Findings?
Mumin Jia ⋅ Yilin Chen ⋅ Divya Sharma ⋅ Jairo Diaz-Rodriguez
Digital personas powered by Large Language Models (LLMs) are increasingly proposed as substitutes for human survey respondents, yet it remains unclear when they can reliably approximate human survey findings. We answer this question using the LISS panel, constructing personas from respondents’ background variables and pre-2023 survey histories, then testing them against the same respondents’ held-out post-cutoff answers. Across four persona architectures, three LLMs, and two prediction tasks, we assess performance at the question, respondent, distributional, equity, and clustering levels. Digital personas improve alignment with human response distributions, especially in domains tied to stable attributes and values, but remain limited for individual prediction and fail to recover multivariate respondent structure. Retrieval-augmented architectures provide the clearest gains, but performance depends more on human response structure than on model choice: personas perform best for low-variability questions and common respondent patterns, and worst for subjective, heterogeneous, or rare responses. Our results provide practical guidance on when digital personas could be appropriate for survey research and when human validation remains necessary.
When Catastrophic Inheritance Meets Forgetting in Continual Adaptation of Foundation Models
Quanyu Zhang ⋅ Zhongyi Han ⋅ Zhenxue Chen ⋅ Xiao-Long Yin ⋅ Henggui Zhang ⋅ Shuo Li
Pretraining noise in foundation models has been shown to significantly impair the generalization performance of downstream tasks, leading to catastrophic inheritance. However, when downstream tasks arrive sequentially, how pretraining noise affects continual adaptation, which involves the forgetting of previously learned knowledge, remains underexplored. This paper is the first to study this problem. Using asymmetric label noise as a realistic setting, we find that pretraining noise persistently degrades performance on new tasks during continual adaptation and may further exacerbate forgetting of previously adapted tasks. Further analysis shows that the persistent effect of pretraining noise mainly manifests as inherited output contraction, where the logit gap between the true class and competing classes is compressed. To mitigate this issue, we propose Ambiguous Boundary Correction (ABC), a method that performs anchor-guided neighborhood reweighting to correct ambiguous regions induced by inherited output contraction. Experiments on synthetic and real-world noisy pretrained models show that ABC improves robustness to pretraining noise under both domain shifts and class shifts.
When Does Structure Help? Statistical Tradeoffs for Structured Reverse Processes in Diffusion Large Language Models
Ruofeng Yang ⋅ Jingyuan Liu ⋅ Shuai Li
dLLMs usually denoise by predicting masked tokens independently, but it is unclear when and whether explicit token coupling through data structure can improve their theoretical guarantees and translate into practical gains. In this work, we prove the first learning-error analysis of dLLMs with structured reverse processes based on graphical models, including tree CRFs, bounded-treewidth clique-tree extensions, and Hidden Markov Model probabilistic circuits (HMM-PCs), covering a broad spectrum of text-data structures. Our main result first decouples the dLLM learning error into estimation, optimization, and structural approximation, and then analyzes the role of each term. For the approximation error, a path-sensitive analysis bounds the contribution of an edge missing from the tree by the geometric decay of correlations along its tree-distance detour, yielding a structure-selection criterion that can disagree with Chow--Liu when correlations are weak. We also identify a tradeoff between approximation bias and estimation error, which yields a slowly-increasing preferred width $w^\star{=}\Theta(\log n)$ as a data-to-structure-capacity scaling regime. From the empirical perspective, we first conduct controlled simulation experiments on synthetic and small-text settings (WikiText-103) to validate the predicted finite-sample signatures of our theory, including exponential path decay on chain MRFs, a bound-optimal tree that disagrees with Chow--Liu, and the balance between different error terms. Then, in the real-world setting, across models ranging from $13$M to $1.1$B parameters, we show that the same HMM-PC teacher is helpful in from-scratch training, mildly negative in SFT distillation at pretrained scale, and substantially helpful in inference-time guidance, supporting the conclusion that both the structural prior and the deployment interface are critical, with the prior providing token-coupling structure and the interface determining whether that structure transfers positively or negatively at a given scale.
When Edge Independence Fails: Joint Graph Diffusion with Latent Sociability Priors
Adarsh Jamadandi ⋅ Nicolas Keriven ⋅ Aline Roumy
Discrete graph diffusion models corrupt graph structure through independent edge updates with a Markovian noise model. Combined with the exchangeability requirement on the output distribution, the terminal prior is necessarily an Erd\H{o}s--R\'enyi (E-R) random graph. This has a fundamental consequence for unconditional graph generation without node features: in the large graph limit, any GNN denoiser applied to an E-R graph approximately outputs an E-R graph, making effective denoising difficult. In practice, generated graphs fail to recover meaningful node heterogeneity and higher-order structure. Previous models introduced specialized node features to compensate, framing this as an expressivity problem, while the role of the prior remained unexamined. We instead address the root cause by introducing the \emph{Latent Sociability Prior} (LSP) and a \emph{joint} diffusion process that co-evolves latent node sociability variables and graph structure, removing edge independence while preserving exchangeability. Unlike existing latent graph diffusion methods, the sociability mechanism requires no complex encoder: its sole role is to guide edge updates and preserve node heterogeneity throughout the trajectory. We validate our approach on multiple graph generation datasets and show that the proposed framework more accurately captures graph distributions than baseline methods without node features.
When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models
Huyu Wu ⋅ Meng Tang ⋅ Xinhan Zheng ⋅ Duo Su ⋅ Haiyun Jiang
Text dominance, the tendency of multimodal large language models (MLLMs) to over-attend to textual tokens while under-utilizing non-text inputs, has been observed in vision-language settings, but whether it extends to other modalities remains an open question. We present a systematic, attention-level study of this phenomenon across five modalities (image, video, audio, time-series, and graph), using two diagnostic metrics: the Modality Dominance Index (MDI), which compares per-token attention between text and non-text inputs, and the Attention Efficiency Index (AEI), which normalizes attention share by token share. Across ten models and six benchmarks, text dominance is pervasive in image, video, audio, and time-series settings and intensifies in deeper layers. Controlled token-replication experiments show that expanding non-text sequences without adding semantic content amplifies the imbalance, while graph tasks provide a boundary case where compact, information-dense tokens reverse the effect. Guided by these findings, we evaluate attention-based token compression as a proof-of-concept intervention on the vision modality. On LLaVA-1.5-7B with MMMU-Pro, removing 90\% of visual tokens via informed selection preserves accuracy (33.70\% vs.\ 33.23\% baseline), while random dropping under the same ratio degrades it to 30.29\%; informed selection also lowers the late-layer MDI from 15.63 to 3.46 and reduces latency by 3.5$\times$. These results suggest that text dominance is a cross-modal phenomenon associated with token redundancy, and that token compression can reduce it in the tested vision setting while preserving task performance.
When LLMs Know but Fail to Reason: Injecting Memory for Reasoning Enhancement
Tian Wang ⋅ Shiyu Hu ⋅ Ruiyang Qin ⋅ Hanyue Zhang ⋅ Qingzhuo Wang ⋅ Shuhan Yu ⋅ Dongrui Liu ⋅ Zhihua Wei ⋅ Wen Shen
Large language models (LLMs) have demonstrated strong capabilities on complex reasoning tasks. However, recent studies~\citep{gekhman2026thinking,song2026large,jin2025disentangling,cheng2024understanding} suggest that when handling knowledge-intensive reasoning tasks, LLMs often fail to effectively utilize the knowledge acquired during pre-training, which limits their reasoning performance. To investigate how internal knowledge (also termed memory) is used by an LLM when dealing with a reasoning query, we propose the Memory-to-Reasoning Alignment (MRA) metric to measure the memory engagement level of an LLM on the reasoning query. We empirically verify that failures in reasoning may correspond to insufficient memory engagement. Inspired by this, we propose a training-free Memory-Guided Activation Injection (MGAI) method to increase the memory engagement level of an LLM, in order to improve the reasoning capability of an LLM. Specifically, given a query, MGAI generates an intervention vector and applies it to the reasoning representation to inject the corresponding memory information into the reasoning representation of the LLM. Furthermore, we propose an adaptive intervention strategy to control which queries require intervention and the magnitude of the intervention. Experiments show that our method consistently improves performance across multiple reasoning benchmarks while largely preserving the memory and general capabilities of LLMs.
When Medical VLMs Stop Understanding: MedTEC-Bench for Probing Semantic Specificity
Ahsan H Akash ⋅ Alina Devkota ⋅ Donald A Adjeroh ⋅ Binod Bhattarai ⋅ Prashnna K Gyawali
Medical vision--language models (VLMs) are attracting growing interest as a foundation for clinical image interpretation, retrieval, and decision support. However, current evaluations largely emphasize aggregate accuracy, retrieval, or classification performance, offering limited insight into whether these models possess the fine-grained semantic understanding required for safe and reliable use in high-stakes clinical settings. In this work, we introduce \textbf{MedTEC-Bench}, a diagnostic benchmark for probing semantic understanding in medical VLMs across five levels of increasing specificity, from broad modality recognition to fine-grained findings, and negation. MedTEC-Bench spans five medical imaging modalities: chest radiography, brain MRI, retinal fundus photography, dermoscopy, and histopathology. It combines controlled semantic probes with a suite of metrics designed to reveal failures hidden by aggregate benchmarks. Evaluating a diverse set of VLMs, from broad biomedical models to clinical-specialist models, we find that existing systems often fail on basic yet clinically meaningful semantic distinctions, despite strong reported performance on standard tasks. Our results suggest that current medical VLM evaluations may overestimate vision-grounded clinical understanding and highlight the need for targeted semantic benchmarks before such models are deployed in high-stakes medical workflows.
When Parallelism Pays Off: Cohesion-Aware Task Partitioning for Multi-Agent Coding
Xu Yang ⋅ Lunyiu Nie ⋅ Ethan Chandra ⋅ Stanislav Gannutin ⋅ Fangru Lin ⋅ Swarat Chaudhuri
Multi-agent Large Language Model (LLM) systems offer a way to decompose complex tasks such as coding through parallelization and context isolation, but adding agents in practice introduces inter-agent communication overhead that can offset efficiency gains. We formalize multi-agent orchestration as a graph partitioning problem that captures the *communication-to-computation trade-off*: task decomposition can shorten critical-path computation, while cross-agent dependencies require costly context transfer. We instantiate this view in repository-level software engineering through **Co**hesion-aware **Coder** (CoCoder), which builds dependency graphs from static analysis, isolates structural hub files, partitions the graph via community detection, and executes the partition with a dependency-aware scheduler. Across \(28\) real-world tasks on DevEval and CodeProjectEval, CoCoder Pareto-dominates sequential and file-based parallel baselines as well as Claude Code with Agent Teams, improving pass rate by up to \(14.0\%\) on CodeProjectEval, achieving up to a \($2.10\times$\) wall-clock speedup, and reducing API cost by up to \(35\%\), with the largest gains on the most dependency-dense projects. CoCoder demonstrates how dependency-aware orchestration can make parallel coding agents both theoretically grounded and practically efficient, suggesting a broader design principle for multi-agent LLM systems.
Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback
Huaisong Zhang ⋅ Hao Yu ⋅ Yuxuan Zhang ⋅ Jiahe Wang ⋅ Xinrui Chen ⋅ Haoxiang Cao ⋅ Feng Lu ⋅ Wendong Zhang ⋅ Changqian Yu ⋅ Chun Yuan
Despite generating increasingly photorealistic images, text-to-image (T2I) models still exhibit localized, subtle, and structurally complex failures. Diagnosing these failures requires instance-level feedback that answers where a defect occurs, what type it is, why it is defective, and its importance to overall image quality. While recent dense-feedback methods move beyond scalar supervision, their heatmap-centric representations still formulate diagnosis as pixel-field regression, making it difficult to localize variable-cardinality defects and bind semantic reasons to individual failures. To address this representation bottleneck, we propose Structured Defect Grounding (SDG), which casts T2I diagnosis as structured set prediction by modeling each defect as a (location, type, reason, importance) tuple. To make this formulation trainable and measurable, we introduce SDG-30K, a 30K-image dataset with box-grounded annotations across four modern T2I generators, together with a dedicated evaluation protocol, SDG-Eval. Building on this structured representation, we further present a diagnosis-to-alignment framework in which a Vision-Language Model (VLM) serves as the SDG detector, and BoxFlow-GRPO converts predicted defect sets into box-derived, importance-weighted spatial rewards for diffusion model alignment. Extensive experiments show that our SDG detector outperforms leading proprietary VLMs on structured defect grounding, while SDG-guided rewards consistently improve T2I alignment and support localized image refinement. These results establish SDG as a unified, instance-level interface for diagnosing, evaluating, and enhancing modern generative models.
Direct Preference Optimization (DPO) has become a standard method for preference-based fine-tuning in RLHF pipelines. A practical but largely heuristic step in DPO is dataset curation: given a fixed budget of preference labels, which candidate pairs should be compared to produce the most informative training set and the best downstream policy? This paper studies DPO dataset curation as a sampling-design problem. We model the curated dataset by a sampling design over candidate pairs and analyze how this design propagates through DPO training to the KL-regularized RLHF objective. Our main results provide matching upper and lower bounds on the RLHF optimality gap of the policy learned from a curated DPO dataset. The bounds reveal a simple structural message: the effect of pair selection enters only through a single design-dependent matrix that summarizes how informative the collected comparisons are for learning the policy parameters. This leads to a trace-form criterion that characterizes both an achievable performance guarantee for the DPO estimator and an information-theoretic lower bound for any induced policy estimator, thereby identifying a canonical objective for comparison curation under budget constraints.
Who Should Evolve? Uncertainty-Aware Role Bottleneck Inference for Multi-Agent LLM Training
Qiyu Qin ⋅ Yichen Li ⋅ Haozhao Wang ⋅ Tianzhe Xiao ⋅ Hao Zhou ⋅ Imran Razzak ⋅ Ruixuan Li
Reinforcement learning from trajectory-level outcomes has become a standard approach for training role-partitioned LLM agent pipelines, in which multiple specialized roles such as an orchestrator, a retriever, and a synthesizer collaborate to solve complex reasoning tasks. Upon trajectory failure, existing training methods do not explicitly determine which role should receive updates, applying outcome signals to all roles indiscriminately or diffusely across the pipeline. However, this paper identifies that failures can often be traced to a primary responsible role, which we term the bottleneck role. This suggests that updating all roles may contaminate non-bottleneck roles with irrelevant updates. Therefore, an intuitive solution is to exclusively update the bottleneck role for each failed trajectory. Unfortunately, accurately identifying this role is highly challenging, as we show that roles exhibiting visible symptoms of failure are often merely downstream victims of errors originating from other roles, rather than being the true bottleneck. To tackle this challenge, we propose Role Inference for Selective Evolution (RISE), the first method that formulates per-trajectory update-target selection as role bottleneck inference. Specifically, RISE leverages Beta-binomial lower-confidence scores to estimate role responsibility, effectively discounting uncertain evidence and suppressing updates when attribution confidence is insufficient. Furthermore, RISE decouples the visible symptom locus from the support-constrained responsibility target prior to applying selective role-level training. Empirically, across HotpotQA, 2WikiMultihopQA, and MuSiQue with Qwen2.5-7B and Llama3-8B backbones, RISE consistently outperforms the strongest learning-based baseline, with up to 11.1% relative F1 and 14.1% relative EM improvements over GiGPO on MuSiQue with Llama3-8B. Code is available at: https://anonymous.4open.science/status/RISE-C0F0
Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?
Jiale Liu ⋅ Huajun Xi ⋅ Shaokun Zhang ⋅ Yifan Zeng ⋅ Tianwei Yue ⋅ Chi Wang ⋅ Jian Kang ⋅ Huazheng Wang ⋅ Qingyun Wu
Automated failure attribution uses LLMs to identify where and why agentic systems fail. As agents become more capable, their failures become subtler, making automated attribution increasingly important. We introduce Who&When Pro, a large-scale benchmark for automated failure attribution in agentic systems. Using a strictly controlled pipeline that injects a failure only after exactly replaying a successful prefix, we construct 12,326 failed trajectories with golden labels across 3 modalities and 26 benchmarks covering various scenarios. Beyond benchmarking, we conduct extensive experiments and analyses, revealing systematic patterns in how models attribute failures across modalities, protocols, and model families, and providing empirical guidance for future automated failure attribution systems.
Why DiT Models Underperform as Representation Learners without Long Skip Connections
Benyuan Meng ⋅ Yiliang Zhang ⋅ Jin-Wen Wu ⋅ Qianqian Xu ⋅ Qingming Huang ⋅ Longtao Huang
Diffusion feature — intermediate activations extracted from diffusion model backbones — has emerged as a promising approach for unifying generation and vision understanding (discrimination). Despite the advancements of diffusion backbones into Diffusion Transformers (DiT), the knowledge of diffusion feature largely remains in U-Net architectures. Hence, in this work, we aim to extend the study of diffusion feature onto DiT backbones. We find that simply applying the previous methodology of diffusion feature to DiT models yields unsatisfying results. We further show that long skip connections (LSCs) are a key missing factor behind this discrepancy and propose a mechanistic explanation for how they affect representation quality. Specifically, LSCs can provide shortcuts to help noises bypass the section of the backbone where the better features reside. This hypothesis is validated using both mutual information measuring for correlation evidence and a controlled study for causal evidence. Of the two, the controlled study demonstrates that LSCs work by transporting noise information rather than just any information, indicating that LSCs enable DiT backbones to adopt a more optimized behavior pattern. Thus, our findings suggest that representation quality in diffusion models is strongly influenced by how information is routed within the backbone.
We identify a diagonal saturation principle in modal inverse problems: when truncation noise is isotropic, the Bayes-optimal Tikhonov shape is a closed-form power law Γk ∝ λk^|s| set by the prior alone, independent of the domain. Berry's random-wave conjecture decorrelates the truncation noise across modes, and Weyl's eigenvalue counting law supplies enough modes for the conclusion to survive empirical Berry violations. Together they predict an approximately flat loss landscape across the per-mode family, leaving narrow scope for a diagonal regularizer to robustly beat the closed form. On FEM-simulated acoustic rooms, the closed form is near-optimal relative to per-room oracle tuning across observation windows, and three diagonal architectures trained on the same data match its reconstruction error within 1\,pp despite learning qualitatively different spectra. The framework extends to heat diffusion via a known exponential Green's function correction with no new free parameters. Saturation is restricted to the diagonal family: Learned Iterative Ridge crosses the boundary by exploiting cross-mode coupling, locating where learning starts to help.
Why Speculative Decoding Works Better Than Predicted on Sparse MoE Models
Ekagra Ranjan ⋅ Komal Teru ⋅ Bharat Venkitesh ⋅ Acyr Locatelli
Speculative decoding (SD) accelerates autoregressive inference by leveraging the memory bound nature of decoding, verifying multiple draft tokens in a single forward pass. For mixture-of-experts (MoE) models, prior work predicts a concave speedup-batch size relationship, with gains peaking at intermediate batch sizes due to expert saturation. We evaluate this across 10 sparse MoE models (30B-1T parameters, 8-512 experts) and find recent fine-grained models break this trend, showing higher speedup at low batch size. We find verification cost is governed not only by activated expert weight, but also by temporal overlap across draft tokens and GPU execution dynamics. At low batch size, fine-grained MoE layers run faster than expected. Small per-expert weights and small per-token activated parameter footprint leave bandwidth underutilized and yield higher L2 cache hit rates. Across draft tokens, routing overlap further reduces the unique experts activated during verification relative to independence assumptions. We model this effect using a maximum-entropy Iterative Proportional Fitting (IPF) framework that predicts unique expert counts from routing-overlap statistics with less than 4\% error, and reproduces observed speedup curves from synthetic routing.
Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies
Tang Zirui ⋅ Xuanhe Zhou ⋅ Yumou Liu ⋅ Linchun Li ⋅ Wang Weizheng ⋅ Hongzhang Huang ⋅ Jun Zhou ⋅ Jiachen Song ⋅ Shaoli Yu ⋅ Wang Jinqi ⋅ Zihang Zhou ⋅ Hongyi Zhou ⋅ YutingLv ⋅ Jinyang Li ⋅ Jiashuo Liu ⋅ Ruoyu Chen ⋅ Chunwei Liu ⋅ Guoliang Li ⋅ Jihua Kang ⋅ Fan Wu
Workspace learning requires AI agents to identify, reason over, exploit, and update explicit and implicit dependencies among heterogeneous files in a worker's workspace, enabling them to complete both routine and advanced tasks effectively. Despite its importance, existing relevant benchmarks largely evaluate agents on pre-specified or synthesized files with limited real-world dependencies, leaving workspace-level evaluation underexplored. To this end, we introduce Workspace-Bench, a benchmark for evaluating AI agents on Workspace-Learning involving Large-Scale File Dependencies. We construct realistic workspaces with 5 worker profiles, 74 file types, 20,476 files (up to 20GB) and curate 388 tasks, each with its own file dependency graph, evaluated across 7,399 total rubrics that require cross-file retrieval, contextual reasoning, and adaptive decision-making. We further provide Workspace-Bench-Lite, a 100-task subset that preserves the benchmark distribution while reducing evaluation costs by about 70%. We evaluate 4 popular agent harnesses and 7 foundation models. Experimental results show that current agents remain far from reliable workspace learning, where the best reaches only 68.7%, substantially below the human+tool result of 80.7%, and the average performance across agents is only 47.4%. More details please refer to https://workspace-bench.github.io.
WorldAct: Activating Monolithic 3D Worlds into Interactive-Ready Object-Centric Scenes
Jichen Hu ⋅ Jiawei Guo ⋅ Jiazhong Cen ⋅ Chen Yang ⋅ Sikuang Li ⋅ Wei Shen
Recent 3D world modeling systems based on generative scene synthesis, such as Marble, can create coherent and explorable 3D environments, yet their outputs are typically static monolithic assets with limited editability and physical interaction. This restricts their use in immersive content creation and embodied simulation, where generated worlds must be actively modified and manipulated. To tackle this challenge, we present WorldAct, a framework that converts static generated 3D worlds into editable and interaction-ready scenes. WorldAct uses a multimodal agent to guide scene decomposition, identify actionable objects, reconstruct geometrically aligned object-level meshes for interaction, and restore the residual background via 3D inpainting. The resulting scenes support object-level editing, collision-aware manipulation, and embodied task execution while preserving global scene coherence. Experiments show that WorldAct enables richer interaction scenarios than the original generated scenes, suggesting a practical path toward editable and interactive 3D world models.
World Motion Models: Flexible Sequence Modeling of SE(3) Trajectories
Jiahui Lei ⋅ Qianqian Wang ⋅ Trevor Darrell ⋅ Angjoo Kanazawa
Equipping artificial agents with spatial intelligence requires a comprehensive generative prior over the dynamic 3D world. We propose World Motion Models (WMMs) that capture "what was, is, and will be where across time" via sparse SE(3) pose trajectories. WMMs are built on the observation that elements of dynamic scenes can be well-approximated by a set of rigid SE(3) trajectories, a minimal yet expressive primitive for 4D modeling. This representation unifies articulated objects, human bodies, hand-object interactions, piecewise-rigid scene dynamics, camera motion, and even robot states and actions into a single shared space. Given this representation, we cast the joint distribution of these entities as a flexible sequence modeling problem, utilizing flow-matching with per-token noise levels. Coupled with a context token mechanism with non-sequential conditioning, this formulation supports any-to-any marginal conditioning across an arbitrary number of entities and time steps. Tasks such as future prediction, motion infilling, model-predictive control, inverse kinematics, cross-embodiment retargeting, and policy learning all reduce to the application of different masks over the same network. Experiments on diverse 6 applications of 3D vision and robotics demonstrate the versatility and flexibility of WMMs with strong performance.
WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors
Keming Wu ⋅ Yijing Cui ⋅ Wenhan Xue ⋅ Qijie Wang ⋅ Xuan Luo ⋅ ZhiYuan Feng ⋅ Zuhao Yang ⋅ Sudong Wang ⋅ Sicong Jiang ⋅ Haowei Zhu ⋅ Zihan Wang ⋅ Ping Nie ⋅ Wenhu Chen ⋅ Bin Wang
Commercial video generation systems such as Seedance2.0 and Veo3.1 have rapidly improved, strengthening the view that video generators may be evolving into "world simulators." Yet the community still lacks a benchmark that directly tests whether a model can reason about how an observed world should evolve over time. We introduce WorldReasonBench, which reframes video generation evaluation as world-state prediction: given an initial state and an action, can a model generate a future video whose state evolution remains physically, socially, logically, and informationally consistent? WorldReasonBench contains 436 curated test cases with structured ground-truth QA annotations spanning four reasoning dimensions and 22 subcategories. We evaluate generated videos with a human-aligned two-part methodology: Process-aware Reasoning Verification uses structured QA and reasoning-phase diagnostics to detect temporal and causal failures, while Multi-dimensional Quality Assessment scores reasoning quality, temporal consistency, and visual aesthetics for ranking and reward modeling. We further introduce WorldRewardBench, a preference benchmark with approximately 6K expert-annotated pairs over 1.4K videos, supporting pair-wise and point-wise reward-model evaluation. Across modern video generators, our results expose a persistent gap between visual plausibility and world reasoning: videos can look convincing while failing dynamics, causality, or information preservation. We will release our benchmarks and evaluation toolkit to support community research on genuinely world-aware video generation.
WorldVLA: A Unified Vision-Language-Action and World Model
Jun CEN ⋅ Siteng Huang ⋅ Yuqian Yuan ⋅ Kehan Li ⋅ Hangjie Yuan ⋅ Chaohui Yu ⋅ Bohan Hou ⋅ Yuming Jiang ⋅ Jiayan Guo ⋅ Xin Li ⋅ Hao Luo ⋅ Fan Wang ⋅ Deli Zhao ⋅ Hao Chen
We present WorldVLA, an autoregressive action world model that unifies action and image understanding and generation. Our WorldVLA integrates Vision-Language-Action (VLA) model and world model in one single framework. The world model predicts future images by leveraging both action and image understanding, with the purpose of learning the underlying physics of the environment to improve action generation. Meanwhile, the VLA model generates the subsequent actions based on image observations, aiding in visual understanding and in turn helps visual generation of the world model. We demonstrate that WorldVLA outperforms standalone VLA and world models, highlighting the mutual enhancement between the world model and the VLA model. We evaluate WorldVLA in both simulation and real-world robot tasks. WorldVLA achieves 97.4% success rate on the LIBERO simulation benchmark without pretraining, while in real-world LeRobot experiments, its integrated world model boosts the overall success rate by 50%.
WTF?! Simulation-Free Reinforcement Learning with Wasserstein-Tilted Flow Maps
Abbas Mammadov ⋅ Jerry Huang ⋅ Justin Lin ⋅ Partha Kaushik ⋅ Sheel Shah ⋅ Kartik Nair ⋅ Yee Whye Teh ⋅ Nicholas Boffi
Fine-tuning aims to update a pre-trained flow-based generative model to improve the downstream reward of its generated samples. Existing methods typically frame this problem as sampling from a reward-tilted distribution, which arises as the solution to a KL-regularized reward-maximization problem. Here we take an alternative approach and introduce an optimal transport regularizer built directly from the pre-trained drift; we show that the resulting fine-tuning problem is equivalent to a deterministic optimal-control problem on the flow. Assuming access to a pre-trained flow map, we exploit this equivalence to devise a simulation-free reinforcement-learning algorithm for fine-tuning generative flows. We call the resulting framework Wasserstein-Tilted Flow Maps (WTF), the first end-to-end fine-tuning recipe native to flow maps. The output of our approach is itself a fine-tuned flow map, retaining few-step reward-aligned inference at deployment. Numerical experiments at text-to-image scale highlight both the efficiency and the efficacy of our approach. More broadly, our work argues that accelerated samplers such as flow maps are essential infrastructure for efficient post-training, and that the dominant KL-regularized formulation of fine-tuning is only one of many choices worth revisiting.
XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding
Hongchen Wei ⋅ Yuanzhe Wang ⋅ Bei Liu ⋅ Yifan Yang ⋅ Qi Dai ⋅ Ruichun Ma ⋅ Kai Qiu ⋅ Yunsheng Li ⋅ Dongdong Chen ⋅ Chong Luo ⋅ Zhenzhong Chen ⋅ Baining Guo
Real-world document tasks often ask professionals to answer questions from annual reports, regulations, clinical guidelines, and technical manuals that span hundreds or thousands of pages. Some questions also require comparing related reports. Reliable long-document understanding is therefore a prerequisite for using LLMs in compliance, clinical, financial, and engineering workflows, where decisions must be traceable to specific evidence pages and the cost of an unsupported answer is high -- yet most existing benchmarks still measure short-context or single-page QA. We introduce \textbf{XL-DocBench}, a fully human-verified benchmark for extra-long document understanding, with 1519 retained questions from six professional domains and contexts up to 2,303 pages. XL-DocBench goes beyond page-level lookup. 1,103 examples (72.6\%) use multiple evidence pages. The final set also includes 556 questions (36.6\%) that use tables, charts, or figures, and 165 questions (10.9\%) that require evidence from multiple documents. Each question has one of twelve reasoning labels, expert-annotated evidence pages, a typed verification rule, and an answer format, including 218 \texttt{None}-answer cases. We build the benchmark with a tree-guided synthesis pipeline followed by artifact filters and full verification by 194 human experts. By coupling extra-long professional contexts with page-level evidence and typed rules, XL-DocBench fills a gap left by prior single-page, short multi-page, or text-only long-context benchmarks, and lets future work attribute system failures to retrieval, evidence use, or rule following rather than to a single leaderboard score. The results show that current systems still struggle with long contexts, multi-page evidence, and structured reasoning over professional documents.
XTC: Head-Aware Sampling by Excluding Top Choices
Philipp E Weidmann ⋅ Allen Roush ⋅ Judah Goldfeder ⋅ Sanjay Basu ⋅ Ravid Shwartz-Ziv
Standard decoding rules for autoregressive language models promote diversity by rescaling the full next-token distribution or by truncating its low-probability tail. These strategies overlook a recurring regime of open-ended generation in which the model already assigns substantial probability to several plausible continuations yet still concentrates too much mass on the most generic choice. We introduce XTC Exclude TopChoices, a lightweight head-aware decoding operator that targets this head-ambiguity regime directly. Given a next-token distribution, XTC identifies the set of tokens exceeding an absolute plausibility threshold $\tau$. When two or more such tokens exist, it removes the dominant eligible choices with probability $\rho$ and retains only the weakest plausible alternative before renormalizing. A comprehensive evaluation spanning 60 experiments across three primary model families (Gemma-3 27B q4, Gemma-3 12B q6, DeepSeek R1 14B q6), extended with a scaling validation on Llama 3.3 70B q4, confirms the predicted operating profile. On creative generation tasks, XTC improves the diversity-repetition Pareto frontier with Distinct-2 gains of 11-15 % (monotone in parameter count from 12B to 70B) and repeat trigram reductions of 27-47% across the four tested models. When composed with temperature scaling, total improvements reach 38\% (Distinct-2) and 71% (repeat trigram reduction) over baseline. A blinded Amazon Mechanical Turk study with 150 Master raters confirms that these distributional shifts translate to a 62.3% creativity preference for XTC ($p<10^{-4}$) without sacrificing fluency, and a cross-vendor GPT-4o control judge replicates the Anthropic-judge signal on every directional measure. On instruction-following (IFEval, Llama~3.3 70B q4), XTC preserves prompt-level strict accuracy within 1.7 percentage points of baseline at parameters that recover most of the diversity gain. A temperature setting matched on Distinct-2 collapses IFEval by 8.8 points at the same Distinct-2 target. The effect is additive with temperature and repetition penalties, robust across quantization levels and model families, and consistent across all twelve tested prompt genres.
XTraj: A Coarse-to-Fine Autoregressive Framework for Transferable Trajectory Generation
wang chao ⋅ ZHIQUAN LAI ⋅ Xinwei Fang ⋅ Yuanshao Zhu ⋅ Xun Zhou ⋅ James Yu
Generating realistic human trajectories is essential for mobility simulation but remains challenging when they must generalize across cities and respect road network constraints. In practice, existing methods often suffer from three key limitations: geographic overfitting, length-rigid generation, and off-road drift when generating GPS coordinates directly. To address these issues, we propose \textbf{XTraj}, a transferable coarse-to-fine autoregressive framework for road-consistent trajectory generation. Specifically, XTraj first learns transferable road-segment representations by integrating road geometry, POI context, historical traffic intensity, and graph topology. Based on these representations, it then autoregressively generates variable-length road-segment routes under road-connectivity constraints. Finally, instead of directly regressing latitude-longitude coordinates, XTraj further refines each generated route in a route-progress space by predicting monotonic progress increments along valid road geometry, from which GPS points are recovered by interpolation. This route-aligned design naturally supports variable-length GPS generation and guarantees road-geometry consistency by construction. Experiments on two real-world vehicle trajectory datasets show that XTraj improves spatial fidelity over competitive baselines, transfers effectively to unseen cities in a zero-shot setting, and ensures road-geometry consistency by construction without any post-hoc map matching. The implementation is provided in \textbf{https://anonymous.4open.science/r/XTraj-EBE7/}.
ZetaEvolve: Learning to Search through History-Conditioned Potential Value
Jiyang Shen ⋅ Xiaojing Zhang ⋅ Bochen Lyu ⋅ Zhanxing Zhu
Self-evolving agents have shown strong empirical performance in mathematical discovery and code optimization, where the challenge is that the search space is too large and unstructured for exhaustive exploration. However, most existing methods lack systematic modeling of self-evolution, resulting in inefficient and even redundant component designs. In this paper, we model self-evolution through the lens of Partially Observable Markov Decision Process (POMDP). We treat the contexts surrounding each search node (the current program and its information to be evolved) as history, while the node’s state, shaped dynamically by the ongoing search process, remains latent and unobservable. This perspective combines parent-node selection, prompt transformation, and child-code generation into a single sequential decision framework, and better aligns with the optimization process of self-evolution. We then develop ZetaEvolve, a search strategy that improves Predictor Upper Confidence bound applied to Trees (PUCT) with a history-conditioned mechanism tailored to partial-observation signals in self-evolving search. Empirically, ZetaEvolve outperforms strong baselines across various benchmarks and achieves faster convergence than existing SOTA methods.