Skip to yearly menu bar Skip to main content


Session

Paris Poster Session 2

Paris Poster Hall
Thu 10 Dec 3 a.m. AEDT — 5 a.m. AEDT
Abstract:
Chat is not available.


$\pi^2$: A Simple Framework for 2D-to-3D Registration

Hao Deng ⋅ Xiangtai Yang ⋅ Guangmingzi Yang ⋅ Xijing Wang ⋅ MaYuanxiao ⋅ Sisi Li ⋅ Zhiqiang Tian ⋅ Shaoyi Du

In this paper, we present $\pi^2$, a simple, efficient, and accurate image-to-point-cloud (I2P) registration paradigm. Instead of relying on intermediate targets such as 2D--3D correspondences, matching confidence, visibility masks, overlap regions, or pose-related proxies, we reformulate I2P registration as point-wise camera-centered coordinate regression. Given a target image and an unaligned LiDAR point cloud, $\pi^2$ directly predicts where each sampled 3D point should lie in the target camera frame. The predicted camera-centered point set is then aligned with the original point set through closed-form SVD-based rigid alignment, yielding the final relative pose without iterative pose optimization or a learned pose-regression head. To improve robustness under large viewpoint changes and sparse outdoor observations, we further introduce a lightweight camera-ray-based regularization objective that provides visibility-aware supervision during training, without requiring an additional visibility or overlap detection network. Experiments on KITTI and nuScenes show that $\pi^2$outperforms the previous SoTA method ICL by 2.47 and 6.72 percentage points in registration success rate, respectively, while running in real time at about 25 FPS, corresponding to an up to 6$\times$ speedup over ICL. Code and pretrained checkpoints will be made available upon acceptance.


ACER: Towards Generalizable Protein-ligand Co-folding

Nopsinth Vithayapalert ⋅ Francesca Grisoni

Predicting protein–ligand complex structures is a central challenge in drug discovery. While recent co-folding models such as AlphaFold-3 achieve accurate structure prediction, they fail to generalize to underexplored binding interfaces -- systematically misplacing ligands, particularly for allosteric or structurally novel targets. To address this gap, we present \textbf{ACER} (\textbf{A}daptive \textbf{C}o-folding via pocket \textbf{E}xploration and pose \textbf{R}anking), a training-free framework that (a) enables co-folding models to systematically explore alternative binding pockets, and (b) leverages the discovered pockets to increase pose accuracy. Our method enables the efficient discovery of non-prevalent pockets without prior expert knowledge. ACER improves pocket discovery and pose accuracy on allosteric targets and structurally novel complexes, successfully modeling binding interfaces that are under-represented or absent from the training set. Our results demonstrate how improved sampling dynamics enhance the generalisability of co-folding models without retraining.


Active Context Selection Improves Simple Regret in Contextual Bandits

Mohammad Shahverdikondori ⋅ Jalal Etesami ⋅ Negar Kiyavash

We study the contextual multi-armed bandit problem with a finite context space (a.k.a. subpopulations), where the learner recommends a best action for each context and is evaluated by context-weighted simple regret. Our guarantees are worst-case over the reward distributions, while remaining instance-dependent with respect to the context distribution vector $\mathbf p$. Akin to experimental design problems where the population of interest is fixed but the sampled subpopulation can be controlled, we allow the learner to actively choose which context to sample from. For a known $\mathbf p$, we characterize tight regret rates: passive sampling where contexts are randomly revealed achieves regret of order $\sqrt{n/T \lVert {\mathbf{p}} \rVert_{1/2}}$, whereas active sampling with allocation $q_j \propto p_j^{2/3}$ achieves the tight rate $\sqrt{n/T} \lVert {\mathbf{p}} \rVert_{2/3}$. The resulting improvement can be as large as $\Theta(k^{1/4})$, where $k$ is the number of contexts. We further extend the analysis to budgeted active sampling, characterize the corresponding tight rate, and identify when a limited active budget suffices to recover the fully active rate. When $\mathbf p$ is unknown, we propose the Explore-Explore-Then-Commit (EETC) algorithm, which optimally balances estimating the context distribution and the time to switch to active allocation, such that for large horizons, it matches the known-$\mathbf p$ active rate up to constants. Experiments on synthetic and real-world data support our theoretical findings.


ADAPT: Hybrid Prompt Optimization for LLM Feature Visualization

João N. Cardoso ⋅ Arlindo L Oliveira ⋅ Bruno Martins

Understanding what features are encoded by learned directions in LLM activation space requires identifying inputs that strongly activate them. Feature visualization, which optimizes inputs to maximally activate a target direction, offers an alternative to costly dataset search approaches, but remains underexplored for LLMs due to the discrete nature of text. Furthermore, existing prompt optimization techniques are poorly suited to this domain, which is highly prone to local minima. To overcome these limitations, we introduce ADAPT, a hybrid method combining beam search initialization with adaptive gradient-guided mutation, designed around these failure modes. We evaluate on Sparse Autoencoder latents from Gemma 2 2B, proposing metrics grounded in dataset activation statistics to enable rigorous comparison, and show that ADAPT consistently outperforms prior methods across layers and latent types. ADAPT also produces significantly more interpretable prompts in human evaluation. Our results establish that feature visualization for LLMs is tractable, but requires design assumptions tailored to the domain.


Adaptive Robust Estimator for Policy Optimization in Reinforcement Learning

Zhongyi Li ⋅ Wan Tian ⋅ Jingyu Chen ⋅ Kangyao Huang ⋅ Huiming Zhang ⋅ Hui Yang ⋅ Tao Ren ⋅ Ruijie Wang ⋅ Yijie Peng ⋅ Yikun Ban ⋅ Fuzhen Zhuang

Reinforcement learning (RL) has become a key ingredient in post-training large language models (LLMs), with Group Relative Policy Optimization (GRPO) and its variants widely adopted for improving reasoning performance. Despite their success, these methods rely on batch-level reward normalization, where the empirical mean can be severely distorted by noisy, skewed, heavy-tailed, or contaminated rewards. Such distortion directly affects advantage estimation and may lead to unstable or inefficient policy updates. We propose an \emph{Adaptive Robust Estimator} (ARE), a plug-and-play robustification module for GRPO-style policy optimization. ARE replaces the empirical batch mean with a two-level robust location estimator: adaptive loss minimization suppresses extreme rewards within each block, while median-of-means aggregation limits the influence of corrupted blocks. This design preserves the original policy optimization objective while improving the reliability of reward centering. We prove that ARE is consistent and achieves high-probability deviation guarantees under both finite-variance and heavy-tailed reward distributions. Experiments on mathematical reasoning benchmarks show that ARE improves GRPO-based training in both single-agent and multi-agent settings, with particularly clear gains under noisy and out-of-distribution evaluations. We further validate its generality on embodied vision-and-language navigation tasks, where ARE improves training stability and downstream navigation (VLN) performance.


AI Behavioral Evaluation Should Be Grounded in Psychophysics Across Marr’s Levels

Jonas Mueller ⋅ Bernhard Egger ⋅ Bjoern Eskofier ⋅ Pasquale Minervini ⋅ ANTONIO RIZZO ⋅ Marco Valentino ⋅ Dario Zanca

This position paper argues that AI evaluation research is fragmented across incompatible frameworks (e.g., adversarial robustness, attention probing, bias analysis, corruption tolerance, interpretability), each operating at a different explanatory level without a shared grammar to make this explicit. We propose unifying these efforts under \textbf{AI psychophysics}, a behavioral evaluation methodology grounded in classical psychophysics and Marr's three levels of analysis. The core contribution is a formalization showing that all these threads instantiate the same psychometric function $\psi(\mathbf{S})=p(\mathbf{J}|\mathbf{S})$, but at different levels of explanatory granularity (computational, algorithmic, implementational). We present a taxonomy of 21 evaluation paradigms organized by this framework and show that the absence of a shared grammar causes category errors, where metrics from one level are routinely used to support claims at another. We conclude with a concrete research agenda for standardizing AI evaluation around this framework.

AI-generated content has already become a common component of digital communication, and its detection is often considered a central tool for ensuring accountability, based on the assumption that AI authorship is inherently problematic and should therefore be identified and filtered. This position paper argues that labeling a text as ``AI-generated'' offers little meaningful information about its utility or truthfulness, and that the research community should instead target the underlying properties that motivate concern: factual reliability, provenance, and stylistic quality. Our position is supported by a growing body of evidence showing that universal detectors are fundamentally fragile; we illustrate this through a small internal replication and literature review, noting that performance frequently collapses from near-perfect (0.999 AUROC) to near-random (0.56) under simple stylistic shifts or adversarial attacks. Moreover, even a perfectly robust detector would answer the wrong question, since AI-generated text can be accurate and well-sourced while human-written text can be misleading. We outline a research agenda that reframes AI-text analysis around claim-level verification rather than origin classification, allowing the community to adapt constructively to the growing prevalence of AI-assisted writing in education, publishing, and professional settings.

Class-incremental learning (class-IL) faces a structural obstacle: discriminative sequential training optimises only the diagonal blocks of a pairwise loss matrix, leaving the inter-task blocks---task confusion (TC)---unminimised. A recent Infeasibility Theorem makes this precise, proving that no discriminative class-IL learner can attain the joint-training optimum even when catastrophic forgetting (CF) is perfectly addressed. Yet the state of the art consists almost entirely of discriminative methods built on pretrained foundation models that approach joint-training performance. We explain this regime by showing that pretraining \emph{exponentially attenuates} the obstacle. For any pretrained backbone $\phi$, the TC gap of any CF-optimal prototype-based discriminative learner---a class that captures the test-time classifier of SLDA, RanPAC, FeCAM, and the prototype branches of EASE and InfLoRA---is bounded by $\exp(-\gamma_{\min}(\phi)/8)$, where $\gamma_{\min}(\phi)$ is the worst-case Mahalanobis margin between task centroids in feature space: a single, computable scalar that determines how much TC remains. Combined with a finite-sample convergence analysis under heavy-tailed gradient noise and a representation drift bound governing backbone fine-tuning, this yields a diagnostic decomposition of the class-IL excess risk into $E_{\mathrm{CF}} + E_{\mathrm{TC}}$, with disjoint direct dependencies on optimisation budget and backbone quality. Across 12 (backbone, dataset) configurations spanning ViT-IN21K, DINOv2, and CLIP plus four random-backbone anchor points, the pooled empirical slope of $\log(\mathrm{TC\ gap})$ versus $\widehat\gamma_{\min}$ on softmax cross-entropy heads is $-0.119$ (bootstrap 95\% CI $[-0.134, -0.105]$), consistent with the predicted $-1/8$ and excluding the conservative fall-back $-1/16$. The slope fit extends the rate beyond the formal LDA/QDA scope. $E_{\mathrm{CF}}$ scales as $K^{-(\alpha-1)/\alpha}$ as predicted, with the rate steepening by more than $3\times$ under gradient clipping; representation drift scales linearly in the backbone learning rate ($r = +0.998$). A theory-derived training procedure---whose schedule is determined by $\widehat\gamma_{\min}$, the tail index $\widehat\alpha$, and warm-up estimates of smoothness and curvature---is competitive with the strongest pretrained-CIL baselines across the four standard benchmarks.


AMARIS: Merging Generalist and Specialist LLMs via Adaptive Subspace Inheritance

Hwee Young Kang ⋅ Irwin De Tao Chin ⋅ Wen-Haw Chong ⋅ Hai Leong Chieu

Domain specialization of large language models (LLMs) often improves target-domain performance but degrades broad capabilities. Model merging is an attractive training-free alternative, yet common parameter-space strategies (for example, interpolation and sparsified averaging) provide limited control over task-specific interference. We introduce Adaptive Model Alignment via Ranked Inherited Subspaces (AMARIS), a covariance-driven framework that transfers specialist updates only along selected activation directions. AMARIS formulates merging as budgeted subspace inheritance: it maximizes specialist activation capture while constraining base-activation deviation, then builds a low-rank projection gate from ranked generalized eigendirections under an explicit preservation budget. This yields a single static merged checkpoint that concentrates transfer in high-utility, low-interference directions. We evaluate AMARIS on Qwen3-based merges (8B and 32B) across finance, instruction-following, and multilingual SEA specialization. In our evaluated settings, AMARIS reaches Pareto operating points that are competitive with strong training-free baselines, including 98.37\% specialist retention with 95.76\% core retention (finance), 103.49\% core retention with 90.28\% specialist retention (instruction following), and 93.69\% specialist retention with 99.77\% core retention (SEA). These results support subspace-constrained merging as a practical route to compositional capability integration without additional training.


Amortized Molecular Optimization via Group Relative Policy Optimization

Muhammad Bin Javaid ⋅ Hasham Hussain ⋅ Ashima Khanna-Reiter ⋅ Berke Kisin ⋅ Jonathan Pirnay ⋅ Alexander Mitsos ⋅ Dominik G Grimm ⋅ Martin Grohe

In structurally constrained molecular optimization, state-of-the-art methods restart an expensive oracle-driven search from scratch for every new input structure, scaling poorly to settings with many starting structures or expensive oracles. While amortized approaches that learn a transferable policy could in principle remove this bottleneck, existing methods struggle to generalize to diverse structural constraints at inference time. We present AMORTIX, an amortized Graph Transformer model that natively supports such constraints, optimizing molecular structures in a single forward pass with zero inference-time oracle calls. A central challenge for amortized training in this domain is that optimization difficulty varies drastically across starting structures. We show that, under this heterogeneity, standard reinforcement learning methods fail to stabilize training, and address this by normalizing rewards within groups of completions sharing the same starting structure. We evaluate on structurally constrained single- and multi-target kinase inhibitor design, and on a few-shot prodrug case study. AMORTIX outperforms both amortized and instance-optimization baselines on goal-directed scaffold decoration and ranks first among amortized methods on the PMO benchmark; the prodrug case study further demonstrates transfer of a learned modification rule to unseen drug structures.

We present the first nearly optimal differentially private PAC learner for any concept class with VC dimension 1 and Littlestone dimension $d$. Our algorithm achieves the sample complexity of $\tilde{O}_{\varepsilon,\delta,\alpha,\beta}(\log^*d)$, nearly matching the lower bound of $\Omega(\log^*d)$ proved by Alon et al. [STOC19]. Prior to our work, the best known upper bound is $\tilde{O}(VC\cdot d^5)$ for general VC classes, as shown by Ghazi et al. [STOC21]. The main idea is to combine the tree structure of VC-dimension-$1$ classes with the private median primitive for interior-point selection. We first privately locate a good depth in the tree, then privately identify a good node on that layer. For proper learning, we show how to descend from the improper output to a proper leaf while controlling the additional false positives.


Antibody Generation via Redistributed Latent Diffusion

Andrey Shevtsov ⋅ Viacheslav Meshchaninov ⋅ Pavel Strashnov ⋅ Dmitry Vetrov

Generating antibody sequences is challenging because they combine conserved framework regions with hypervariable loops. Latent diffusion is attractive for this task since it enables flexible conditioning and bidirectional generation. But standard approaches fail. Global noise schedules treat all positions equally, so models learn the predictable frameworks well while the diverse loops remain poorly captured. We address this by learning a latent space that redistributes information evenly, allowing standard diffusion to succeed where it previously failed. On organism-conditioned generation across six species, our approach achieves 10× lower Fréchet Distance than latent diffusion without redistribution. It supports chain-type control, loop infilling, and paired-chain generation. Validation across five protein encoders confirms the method is encoder-agnostic. These results establish latent diffusion as a practical tool for antibody sequence design.


ARC-Encoder: learning compressed text representations for large language models

Hippolyte Pilchen ⋅ Edouard Grave ⋅ Patrick Perez

Recent techniques such as retrieval-augmented generation or chain-of-thought reasoning have led to longer contexts and increased inference costs. Context compression techniques can reduce these costs, but the most effective approaches require fine-tuning the target model or even modifying its architecture. This can degrade its general abilities when not used for this specific purpose. Here we explore an alternative approach: an encoder that compresses the context into continuous representations which replace token embeddings in decoder LLMs. First, we perform a study of training strategies and architecture choices for the encoder. Our findings led to the design of an Adaptable text Representations Compressor, named ARC-Encoder, which outputs $x$ times fewer continuous representations (typically $x\!\in\!\{4,8\}$) than text tokens. We evaluate ARC-Encoder across a variety of LLM usage scenarios, ranging from in-context learning to context window extension, on both instruct and base decoders. Results show that ARC-Encoder achieves strong performance on several benchmarks and tasks while improving computational efficiency at inference. Finally, we demonstrate that our models can be adapted to multiple decoders simultaneously, allowing a single encoder to generalize across different decoder LLMs. This makes ARC-Encoder a flexible and efficient solution for portable encoders that can support multiple LLMs with only small MLPs.


AR-Edit: Training-Free Streaming Video Editing without Inversion

Hovhannes Margaryan ⋅ Vicky Kalogeiton ⋅ Quentin Bammey ⋅ Christian Sandor

Editing a video stream in real time, without access to future frames, is essential for interactive applications. Existing training-free video editing methods, however, assume full offline access to the entire sequence, while streaming approaches either rely on additional training or are limited to specific editing techniques. We show that this limitation can be removed by analyzing the self-attention dynamics of DMD-distilled autoregressive video diffusion models. We uncover a key asymmetry: queries and keys remain stable and aligned with their source counterparts across denoising timesteps, while values diverge and encode the evolving content. This observation leads to a simple mechanism: a single forward pass suffices to capture source identity in the value features, which can be reused throughout generation. Starting from pure Gaussian noise and injecting only these source values reconstructs the input with fidelity matching inversion-based methods, effectively eliminating inversion. Based on this insight, we introduce AR-Edit, a training-free, inversion-free method that caches source values once and selectively injects them during autoregressive denoising, enabling real-time, structure-preserving streaming video editing with minimal overhead. Code will be released.


Assessing Per-Sample Membership Inference Vulnerability without Retraining

Valentin Dorseuil ⋅ Jamal Atif ⋅ Olivier Cappé

Recent work in the privacy literature shows that sample-targeted membership inference attacks (MIA) significantly outperform untargeted approaches by a wide margin. Motivated by this observation, we address the following question: Can the privacy vulnerability of individual training points be assessed without training shadow models? We show that per-sample exposure to MIA is governed not only by a point's loss, but also by a data-dependent geometric measure. In the linear setting, we derive a closed-form decomposition of individual black-box MIA vulnerability into a population leverage score and a residual loss term, making explicit how sample-dependent geometry translates into privacy exposure. Since the final layer of most modern architectures is linear, we extend this framework to deep networks and propose a surrogate score operating on last-layer representations that requires only a single trained model and no shadow models. Empirical evaluations across diverse datasets and architectures show that our score outperforms loss and gradient-norm baselines at identifying the highest-risk points under state-of-the-art attacks, providing a computationally efficient and theoretically grounded tool for per-sample privacy risk assessment.


Asymptotic Anytime-Valid Inference for U-statistics

Leheng Cai ⋅ Qirui Hu ⋅ Weijia Li

We study asymptotic anytime-valid confidence sequences for degree-two U-statistics under continuous monitoring. In the nondegenerate case, Hoeffding's projection reduces the problem to a time-uniform central limit theory for the partial sums of the first-order projection, while the canonical remainder is shown to be negligible under mild moment assumptions. A leave-one-out jackknife estimator then yields a fully data-driven procedure, leading to confidence sequences with asymptotic coverage guarantee for the parameter of interest. In the degenerate case, we show that the U-statistic is approximated by a centered quadratic Gaussian-chaos rather than by a simple Gaussian, which poses significant challenges for sequential inference. To address this issue, we novelly develop the Spectrally Allocated Gaussian-chaos Excursion (SAGE) boundary, and then provide plug-in implementations based on truncated spectrum estimation with consistency guarantees. The resulting widths can attain the expected time-uniform optimal rates: $\sqrt{\log\log n/n}$ in the nondegenerate regime and $\log\log n/n$ in the degenerate regime. Several widely used U-statistics are discussed within the proposed framework, and numerical experiments further support the validity of the derived theory.


A Theoretical Bridge Between Long-Tailed Recognition and Continual Learning

Mahdiyar Molahasani ⋅ Michael Greenspan ⋅ Ali Etemad

We theoretically and empirically establish a previously unexplored connection between Long-Tailed Recognition (LTR) and Continual Learning (CL). Specifically, we show that training on a long-tailed dataset drives model parameters into an $\mathcal{O}(1/\sqrt{\mathrm{IF}})$ neighborhood of the dominant-class solution, under both uniform and exponentially decaying cardinality distributions. Building on this result, we prove that the CL objective upper-bounds the balanced LTR loss, revealing that LTR can be reformulated as a sequential learning problem. Motivated by this connection, we introduce Continual Learning for Long-Tailed Recognition (CLTR), a principled framework that leverages standard off-the-shelf CL methods to sequentially learn Head and Tail classes while mitigating catastrophic forgetting. Extensive experiments on CIFAR100-LT, CIFAR10-LT, ImageNet-LT, and Caltech256 validate our theoretical predictions and demonstrate strong performance across diverse LTR benchmarks. In the foundation-model regime, CLTR with DualPrompt on a CLIP backbone outperforms standard adaptation baselines and is competitive with specialized methods using external semantic supervision. Our work bridges LTR and CL both theoretically and empirically, providing a principled approach for addressing long-tailed learning using standard CL strategies.


A Unified Perturbation Framework for Analyzing Leaderboard Stability and Manipulation

Hosna Oyarhoseini ⋅ Jimmy Lin ⋅ Amir-Hossein Karimi

Evaluation leaderboards such as LMArena play a central role in benchmarking large language models by aggregating pairwise human preferences into model rankings, but the robustness of these rankings is still poorly understood. We present a unified perturbation framework for analyzing Bradley–Terry leaderboards under structured data modifications using influence-based approximations. Our framework studies three match-level perturbations—dropping matches, adding matches, and flipping match outcomes—along with player removal. We evaluate their effects on top-k membership, global ranking consistency measured by Kendall’s tau, and confidence-interval uncertainty. Across Chatbot Arena and six additional pairwise-comparison datasets, we show that modern leaderboards are non-robust across all three objectives: targeted perturbations affecting less than 1% of the data can change the top-ranked model, degrade global ranking consistency, and alter confidence intervals. We summarize these effects using normalized dataset-level robustness scores that compare fragility across leaderboard designs. We further show that influence scores enable efficient targeted manipulation, promoting or demoting specific models with fewer actions than prior manipulation baselines, while also identifying additional matchups that reduce uncertainty for target models. Finally, player-removal analysis shows that removing influential models can induce broad reordering, highlighting model deprecation as a source of the leaderboard illusion. These findings reveal fundamental limitations of current leaderboard designs and motivate more robust evaluation protocols.


Automata from Agent Traces: Failure and Next-Step Prediction

Seonglae Cho ⋅ Franklin Cardenoso Fernandez ⋅ Umar Mohammed ⋅ Zekun Wu ⋅ Kleyton da Costa ⋅ Ilham Wicaksono ⋅ Adriano Koshiyama

LLM-based agents execute multi-step tasks, but their behavioral structure remains opaque: long unstructured traces resist the safety auditing and runtime monitoring that deployment requires. Existing approaches operate per-trace or success-only, missing the cross-run topology that links next-step and failure prediction; we extract provably minimal finite-state machines (FSMs) via prefix-tree construction and structural state merging, providing a structural substrate for the unpredictable nature of LLM agent behavior. Across twelve public datasets, the FSMs are compact (7–43 states), achieve high replay fitness on held-out data with zero structural variance across splits, and build in milliseconds. This substrate addresses both prediction goals. For next-step prediction, FSM-state context beats Agent Workflow Memory on all eight ground-truth-matched datasets. For failure prediction, per-state behavioral features reach held-out AUROC up to 0.94, and an online monitor flags failing SWE-agent runs at rank-AUROC above the trivial flag-everything baseline, enabling early-stopping at 32% trace completion. A single FSM replays four LLMs at perfect fitness, evidence that behavioral topology in LLM agents is shaped more by the deployment harness than by the LLM and giving a model-agnostic structural primitive for safety auditing and runtime monitoring at deployment.


Balancing Frequencies and Pixels in Flow Matching

Lucas Degeorge ⋅ Paul Couairon ⋅ Arijit Ghosh ⋅ Alexei Efros ⋅ David Picard ⋅ Vicky Kalogeiton

Natural images follow a $1/f^2$ spectral distribution: most signal energy lies in the low spatial frequencies, while the perceptually important structures such as textures and edges occupy sparse high-frequency bands. Pixel-space reconstruction objectives, however, treat all spatial errors uniformly, causing low frequencies to dominate the optimization signal and delaying the learning of fine-scale details. In this work, we identify this objective-level spectral imbalance as a key inefficiency in training pixel-space flow models. To address it, we propose a Focal Log-Frequency Loss (\floss{}), a spectrally balanced objective that equalizes the learning signal across frequencies, emphasizing high-frequency components that are otherwise underrepresented in pixel-space objectives. Building on this, we introduce a simple training strategy that combines frequency and pixel supervision: we first emphasize frequency-domain learning early to capture all frequencies, and then transition to standard pixel-space $v$-loss for spatial refinement. This balancing mitigates the low-frequency bias of pixel losses and aligns the training signal with the evolving needs of the model. Our approach is conceptually simple, requires no architectural changes, and acts as a drop-in replacement for flow matching losses. Across multiple model scales, it accelerates convergence by up to 40% while consistently improving FID and perceptual fidelity. We will release code and models.


BalCapRL : A Balanced Framework for RL-Based MLLM Image Captioning

Shaokai Ye ⋅ Vasileios Saveris ⋅ Yihao Qian ⋅ Jiaming Hu ⋅ Elmira Amirloo ⋅ Peter Grasch

Image captioning is one of the most fundamental tasks in computer vision. Owing to its open-ended nature, it has received significant attention in the era of multimodal large language models (MLLMs). In pursuit of ever more detailed and accurate captions, recent work has increasingly turned to reinforcement learning (RL). However, existing captioning-RL methods and evaluation metrics often emphasize a narrow notion of caption quality, inducing trade-offs across core dimensions of captioning. For example, utility-oriented objectives can encourage noisy, hallucinated, or overlong captions that improve downstream question answering while harming fluency, whereas arena-style objectives can favor fluent but generic descriptions with limited usefulness. To address this, we propose a more balanced RL framework that jointly optimizes utility-aware correctness, reference coverage, and linguistic quality. In order to effectively optimize the resulting continuous multi-objective reward formulation, we apply GDPO-style reward-decoupled normalization to continuous-valued captioning rewards and show that it improves performance over vanilla GRPO. Additionally, we introduce length-conditional reward masking, yielding a more suitable length penalty for captioning. Across LLaVA-1.5-7B and Qwen2.5-VL 3B and 7B base models, our method consistently improves caption quality, with peak gains of +13.6 DCScore, +9.0 CaptionQA, and +29.0 CapArena across different models.


Belief Engine: Configurable Stance Dynamics for Multi-Agent LLM Deliberation

Joshua C Yang ⋅ Maurice Flechtner ⋅ Damian Dailisan ⋅ Michiel Bakker

LLM-based AI agents are increasingly used to simulate deliberative social interactions, including negotiation, conflict resolution, and multi-turn opinion exchanges. In these settings, it is important to know not only what an agent says, but how its stance changes after receiving new information. Current systems often make this difficult: agents may drift from assigned roles, echo an interaction partner, remain fixed despite relevant evidence, or appear to change because the prompt or retrieved context has changed. We introduce the Belief Engine (BE), a framework that uses ``belief'' in a Bayesian modelling sense: a maintained evidential state over a proposition, exposed as scalar stance. BE extracts structured evidence, stores active and archived argument records, and updates the belief state through a Bayesian log-odds rule with two interpretable controls: evidence uptake $u$ and prior anchoring $a$. Across multiple LLM models, sweeps of $u$ and $a$ demonstrate stable control over stance dynamics: higher uptake makes agents more responsive to new evidence, while stronger anchoring preserves the initial stance. On DEBATE, a human deliberation dataset with pre/post opinions, BE is most accurate when participants move in the direction of extracted evidence, showing its strength in evidence-following deliberative contexts; when all participants are evaluated together, gains are more modest because many people remain stable or move for reasons not captured by the extracted evidence stream. BE therefore provides a configurable belief-update layer for evidence-grounded deliberation: LLM agents can be made open-minded, anchored, or evidence-sensitive by construction, while their stance trajectories and memory can be inspected, compared, and calibrated against human opinion change.

In this article we develop a new method for summarizing a ranking distribution, \textit{i.e.} a probability distribution on the symmetric group $\mathfrak{S}_n$, beyond the classical theory of consensus and Kemeny medians. Based on the notion of \textit{local ranking median}, we introduce the concept of \textit{consensus ranking distribution} ($\crd$), a sparse mixture model of Dirac masses on $\mathfrak{S}_n$, in order to approximate a ranking distribution with small distortion from a mass transportation perspective. We prove that by choosing the popular Kendall $\tau$ distance as the cost function, the optimal distortion can be bounded in terms of pairwise probabilities, paving the way for the development of efficient learning methods that do not suffer from the lack of vector space structure on $\mathfrak{S}_n$. In particular, we propose a top-down tree-structured statistical algorithm that allows for the progressive refinement of a CRD based on ranking data, from the Dirac mass at a Kemeny median at the root of the tree to the empirical ranking data distribution itself at the end of the tree's exhaustive growth. In addition to the theoretical arguments developed, the relevance of the algorithm is empirically supported by various numerical experiments.


Binding Multiple Modalities via Multimodal Wasserstein Barycenter

Xiaole Tang ⋅ Jiayi Xu ⋅ Xiang Gu ⋅ Yan Yang ⋅ Jian Sun

Multimodal learning beyond two modalities commonly leverages a specific modality (e.g., text) to bind other modalities. However, how to establish a more balanced representation space that approximates shared semantics while respecting the holistic geometry of $n$-modal data remains challenging. In this work, we present BaryBind, which aims to transport the specific modality towards the Wasserstein barycenter (WB) optimized across all modalities and introduces a volumetric alignment objective to establish a unified semantic space around the WB embedding. Specifically, we project specific modalities to the WB, which minimizes the average Wasserstein distances to multimodal distributions and serves as the anchor for subsequent alignment. We then construct a barycenter simplex, whose volume is taken as a similarity metric for global alignment centered at the WB. Extensive experiments show that BaryBind achieves competitive performance in text-video-audio retrieval, classification, videoQA, and cross-modal generation tasks, along with robustness under modality absence and scalability to more than three modalities.


BLANP: Memory-Efficient Backpropagation-Free Local Training via Antithetic Node Perturbation

Mehrzad Karamimanesh ⋅ Marcel A. J. van Gerven ⋅ Mahyar Shahsavari

Backpropagation imposes fundamental constraints on training neural networks. It requires the retention of all forward activations until the corresponding backward pass completes, thereby preventing pipelined execution across layers. This work introduces Backpropagation-Free Local Antithetic Node Perturbation (BLANP), a fully local training framework that eliminates gradient propagation across layer boundaries while requiring no negative samples. Each layer updates its parameters from locally available signals through a zeroth-order antithetic perturbation estimator. To stabilize learning, the estimator is evaluated against a frozen exponential moving average critic, decoupling backbone updates from the concurrently evolving local classifiers. The antithetic perturbation formulation cancels even-order bias terms in the gradient estimate, yielding a first-order accurate update from two forward passes alone. A companion variant, BLANP-Exact, replaces the perturbation estimator with exact yet local gradients under a strict stop-gradient constraint, preserving a memory footprint that does not grow with network depth while recovering gradient fidelity. Experiments on MNIST, Fashion-MNIST, and CIFAR-10 across MLP and convolutional architectures show that BLANP matches backpropagation on simple benchmarks. On the more challenging CIFAR-10, BLANP-Exact surpasses the best prior local learning method by 2.40% with a simpler architecture, and trails an identical backpropagation baseline by only 3.90%. The code is available at https://github.com/anonymous/BLANP.


BO-Arena: An Evolving Benchmark for High-Dimensional Bayesian Optimisation

Thomas Christie ⋅ Colin Doumont ⋅ Henry Moss ⋅ Philipp Hennig

High-dimensional Bayesian optimisation (HDBO) is an increasingly crowded field, with new methods proposed at every major conference, making it ever harder to answer a deceptively simple question: "Which method is currently state-of-the-art?" In this paper, we question the efficacy of benchmarking practices in the field, and find them to be systematically ineffective. Specifically, multiple published works fail to benchmark against the strongest baselines available at the time of writing, or even include incorrect implementations of key methods. As a remedy, we introduce BO-Arena, an actively maintained and extensible package providing canonical implementations of state-of-the-art algorithms, designed to streamline building new methods and rigorously benchmark against existing approaches. We use the package to ask what currently drives performance in HDBO, and find that the answer is even simpler than previously thought. Following these simple principles, we propose a new algorithm which achieves state-of-the-art performance on problems with large numbers of observations, outperforming considerably more complex recent methods. Finally, we question whether there is currently sufficient evidence to suggest that a simple Gaussian process baseline is outperformed by any more complex methods for HDBO.

Recently, deep neural networks on manifold-valued representations have garnered significant attention across various machine learning applications. One recent focus is the generalization of Euclidean fully connected (FC) and convolutional layers to non-Euclidean geometries. However, previous approaches typically focus on a few selected manifolds and rely on specific properties of the target manifold. In contrast, this work proposes a framework for constructing FC and convolutional layers over computationally tractable Riemannian spaces. This framework incorporates several previous FC layers across different geometries as special cases and is instantiated on ten representative manifolds, including three hyperbolic models, five geometries of the symmetric positive definite (SPD) manifold, and two Grassmannian perspectives. Experiments on different manifolds demonstrate the effectiveness and applicability of our approach.


Calibrated Target Noise Recovers Curvature from the Gradients

Arash Jamshidi ⋅ Katsiaryna Haitsiukevich ⋅ Aristides Gionis ⋅ Kai Puolamäki

Estimating curvature information from gradients alone is a fundamental problem in machine learning. The challenge becomes particularly acute in settings where only aggregate mini-batch gradients are available, yet curvature information is still required for applications such as preconditioning in optimisation, sampling, and early stopping. Perhaps surprisingly, we show that a simple target-perturbation scheme suffices to recover such curvature information from gradients alone. Our method injects zero-mean noise into the targets, with variance calibrated to the local output curvature. The covariance of the resulting noisy batch gradients then recovers practical curvature matrices, without requiring per-sample gradients, second-order derivatives, or other internal network quantities. For a broad class of non-linear models, our scheme recovers the Generalised Gauss--Newton (GGN) matrix; for feed-forward ReLU networks, it additionally recovers the diagonal blocks of the population Hessian. We demonstrate that the resulting estimator is simple, lightweight, and easy to integrate into existing training pipelines. Experiments demonstrate accurate curvature recovery, faster optimisation when the estimator is used for preconditioning, and the utility of the estimated GGN for early stopping.


CANDO: Cooperative Agentic Network for Layout Design Optimization

Athanasios Masouris ⋅ ZHENG JING ⋅ Benjamin S Chandler ⋅ Hadi Jamali-Rad

Layout generation for real-world facilities is a challenging problem, requiring reasoning over irregular site boundaries, heterogeneous orientations, access-aware placements, and motion-planning feasibility. Yet, most existing layout benchmarks in the generative AI space target simpler placements over rectangular domains and rely on distributional metrics such as FID and IoU that reward conformity to dataset priors, thus discounting design innovation. Motivated by these gaps, we introduce $\textbf{ALPS-Bench}$, a benchmark of $1,000$ professionally annotated real-world facility layouts paired with an instance-specific scoring protocol grounded in a structured design manual. As a strong baseline for $\textbf{ALPS-Bench}$, we propose $\texttt{CANDO}$, a training-free multi-agent framework in which specialized agents iteratively refine layouts through a verification-grounded loop, concentrating reasoning on strategic spatial decisions. We demonstrate that $\texttt{CANDO}$ surpasses state-of-the-art baselines on the widely adopted $\textbf{PubLayNet}$ and $\textbf{RICO}$ benchmarks by a significant margin, establishing cooperative agentic design as a broadly effective recipe for constraint-aware layout synthesis. Code and benchmark will be released upon acceptance.


CanViT: Toward Active-Vision Foundation Models

Yohaï-Eliel BERREBY ⋅ Sabrina Du ⋅ Audrey Durand ⋅ B. S Krishna

Active computer vision promises efficient, biologically plausible perception through sequential, localized glimpses, but lacks scalable general-purpose architectures and pretraining pipelines, leaving Active-Vision Foundation Models (AVFMs) underexplored. We introduce CanViT, the first task- and policy-agnostic AVFM. CanViT uses scene-relative RoPE to bind a retinotopic Vision Transformer backbone and a spatiotopic scene-wide latent workspace, the canvas. Efficient interaction with this high-capacity working memory is supported by Canvas Attention, a novel asymmetric cross-attention mechanism. We decouple thinking (backbone-level) and memory (canvas-level), eliminating canvas-side self-attention and fully-connected layers to achieve fast sequential inference and scalability to high output resolutions. We propose a label-free active vision pretraining scheme, policy-agnostic passive-to-active dense latent distillation: reconstructing scene-wide DINOv3 embeddings from sequences of low-resolution glimpses with randomized locations, zoom levels, and lengths. We pretrain CanViT-B from a random initialization on 13.2 million ImageNet-21k scenes--an order of magnitude more than previous active models--and 1 billion random glimpses, in 166 hours on a single H100. On ADE20K segmentation, a frozen CanViT-B achieves 38.5% mIoU in a single low-resolution glimpse, outperforming the best active model's 27.6% with 20x fewer inference FLOPs as well as its FLOP- or input-matched DINOv3 teacher. Given additional glimpses, CanViT-B reaches 45.9% ADE20K mIoU. On ImageNet-1k classification, CanViT-B also sets a new active-vision state of the art, with 84.5% top-1 accuracy after fine-tuning. CanViT generalizes to longer rollouts, larger scenes, and new policies. Our work narrows the wide gap between passive and active computer vision, demonstrating the potential of task- and policy-agnostic AVFM pretraining.

We propose Low-Rank Quantile Surfaces (LRQS), a bivariate causal model in which, in the causal direction, an unknown monotone transformation of the conditional quantile surface admits a low-rank functional decomposition. LRQS subsumes location-scale noise models and post-nonlinear heteroscedastic noise models, while allowing multiple quantile bases to represent changes beyond location-scale effects. We prove generic identifiability of LRQS: the transformed quantile surface is low rank in the causal direction, whereas reverse representability under the corresponding constraints occurs only for exceptional, fine-tuned cause marginals. We provide a simple-yet-powerful causal score using a nonparametric fitting procedure that alternates between rank-constrained approximation of discretized quantile surfaces and isotonic estimation of the unknown monotone transformation. Experiments on synthetic mechanisms with higher-rank distributional shape variation and strong nonlinear distortions, together with standard bivariate benchmarks, show that LRQS is especially effective when conditional distributional shape or observation distortion goes beyond existing location-scale assumptions.


CENDRe: Concept Extraction with Natural Domain Representations

Antonia Holzapfel ⋅ Andres Posada Moreno ⋅ Sebastian Trimpe

Convolutional neural networks (CNNs) are widely used for time-series classification, but their deployment in critical domains requires understanding the temporal and spectral patterns that drive their predictions. Concept extraction (CE) methods identify such patterns by analyzing representations within the models' latent space. However, existing time-series CE methods have three limitations: they operate only in the time domain and overlook frequency features, predefine the number of concepts, and produce localizations misaligned with the regions the model uses. We address these limitations by proposing CENDRe, a concept extraction method for CNNs. It first discovers concepts by clustering per-timestep latent representations in two stages, where silhouette-guided aggregation selects the number of concepts automatically. Then, it localizes each concept through gradients of a presence score that contrasts the latent representations with their prototypes, producing masks that concentrate on the regions driving the concept. These gradients, propagated through a differentiable invertible mapping of the input such as a Fourier transform, yield localizations for the same concepts in the frequency domain. Finally, each concept receives a relevance score that quantifies its contribution to each class. On synthetic benchmarks, CENDRe achieves representation correctness comparable to state-of-the-art CE methods and substantially higher importance correctness. On real bearing-fault data, CENDRe extracts the frequency bands driving the model's predictions, located in regions commonly inspected for fault diagnosis, producing evidence to assess the model that time-domain CE methods cannot.


Cephalonauts One: A deep fMRI dataset for decoding naturalistic speech in the human brain

Antoine Collas ⋅ Louis Jalouzot ⋅ Géraud Ilinca ⋅ Corentin Caris ⋅ Romain Valabregue ⋅ Ahmed Hassayoune ⋅ David Gonçalves ⋅ Madeleine Hueber ⋅ Thaddée Delebarre ⋅ Savatovsky Julien ⋅ Clara Fonteneau ⋅ Charles Maussion ⋅ Bertrand Thirion ⋅ Alexis Thual

Cephalonauts One is a whole-brain 3 Tesla (3T) functional magnetic resonance imaging (fMRI) dataset recorded while subjects listened to audio podcasts. Three healthy subjects underwent multiple scanning sessions, each consisting of five 15-minute runs, while listening to podcasts in their native language. The current release contains 14, 20, and 17 hours of fMRI data per subject, respectively, making it one of the deepest per-subject fMRI datasets with naturalistic speech stimuli available. The dataset pairs brain activity with the corresponding podcast audio, transcript annotations, and derived stimulus embeddings. Furthermore, we introduce a brain decoding benchmark formulated as audio segment retrieval: given fMRI activity from a held-out session, the decoder must identify the corresponding time-aligned podcast audio segment among candidate segments. We provide standardized splits, evaluation metrics, and baseline decoders for this task. Finally, a scaling analysis shows that decoding performance improves continuously with the amount of training data per subject. Data & code: https://huggingface.co/datasets/karavela/cephalonauts_one

Efficient heterogeneous Multi-Agent Reinforcement Learning (MARL) in continuously changing environments is key to advancing MARL from simulation to the real world. However, it remains unclear whether current MARL methods can adapt under such dynamic scenarios. In this work, we show that existing methods face significant performance degradation in a novel continual heterogeneous MARL environment we construct, named slippery Multi-Agent Mujoco. We further demonstrate the connection between this phenomenon and representation redundancy as well as insufficient policy heterogenization. To address this issue, we propose Continual Heterogeneous CooperAtion with Information BottleNeck (CHAIN), equipped with Heterogeneous Policy Bottleneck (HPB). In HPB, we extend the traditional information bottleneck into three parts, fitting, compression and heterogenization. Our HPB encourages agents to learn an efficient latent representation to adapt to changed environments, as well as disentangles specializations of agents in state perception, thereby encouraging heterogenization. We evaluate CHAIN on standard MA-Mujoco, our slippery MA-Mujoco and a multi-task continual MARL benchmark, MEAL, with various challenging tasks. The experimental results indicate that our CHAIN outperforms state-of-the-art heterogeneous MARL methods across key continual learning metrics.


ChainSpace: A Chained-Reasoning Paradigm for Spatial Intelligence

Xiaohan Zhang ⋅ Feng Gu ⋅ Xudong Rao ⋅ Xuhao Pan ⋅ Tao Wei ⋅ Pan Zhou ⋅ Kun Zhan

Spatial intelligence requires foundation models to maintain coherent spatial state across interactions with the physical world. However, existing data-centric approaches typically treat spatial reasoning as independent question-answer instances, enabling shortcut-based answering and providing limited supervision for persistent spatial understanding. To address this, we introduce ChainSpace, a chained-reasoning paradigm that structures spatial reasoning as a state-preserving multi-round process. In this paradigm, spatial questions are organized into logically constrained and jointly consistent chains, where later questions depend on spatial constraints established in earlier rounds. Following this principle, we instantiate ChainSpace-Bench, a manually annotated real-world multi-round benchmark with a Chain-Aware Metric, and ChainSpace-Pipeline, a simulator-based chain-structured supervision generation framework for spatial intelligence training. Experiments show that ChainSpace-Bench exposes chain-level failures that are not captured by isolated question accuracy. Additionally, with a relatively small amount of simulator-generated chained data, models trained by ChainSpace-Pipeline achieve the best performance among open-source models on ChainSpace-Bench and transfer competitively to multiple external spatial intelligence benchmarks. These results establish ChainSpace as an effective paradigm for more faithful evaluation and more data-efficient learning of spatial intelligence.


CHASM: Cross-frequency Harmonized Axis-Separable Mixing for Spectral Token Operators

Pengcheng Fang ⋅ Hongli Chen ⋅ Yuxia Chen ⋅ Tengjiao Sun ⋅ Jiaxin Liu ⋅ Xiaohao Cai

Spectral token mixers based on Fourier transforms provide an efficient way to model global interactions in visual feature maps. Existing designs often either apply filter-wise spectral responses along fixed channel axes, or learn adaptive frequency-indexed channel mixing without explicitly aligning the channel directions used across frequencies. We propose CHASM, a Cross-frequency Harmonized Axis-Separable Mixer, as a structured middle ground. CHASM separates what should be shared from what should remain frequency-specific: all frequencies share a learned channel eigenbasis, while each frequency retains its own positive spectral gains. The shared basis makes channel directions comparable across the spectrum, whereas the positive gains preserve local spectral adaptivity. CHASM applies this structured operator separably along the height and width axes and is used as a drop-in replacement mixer inside existing backbones. We provide a structural characterization of the shared-basis operator family and evaluate CHASM through controlled same-backbone comparisons. Across accelerated MRI reconstruction, undersampled MRI segmentation, and natural-image reconstruction, CHASM consistently improves over same-backbone spectral-mixer baselines. Ablations show that removing the shared-basis constraint weakens performance, and randomizing coherent sampling geometry substantially reduces the gain, supporting cross-frequency harmonization as a useful inductive bias for spectral token operators.

3DGS enables real-time novel view synthesis, but practical deployment under varying compute budgets requires a single Gaussian set that remains effective when truncated to a prefix of primitives. This raises how capacity should be allocated as the budget decreases.We observe two findings. First, image complexity correlates with reconstruction error in both image and 3D Gaussian space. Second, this structure degrades under strong budget reduction, indicating suboptimal allocation.These observations motivate using image complexity as an optimization signal. We propose a continuous level-of-detail method based on 3DGS-MCMC that uses it for (i) relocation toward difficult regions during training and (ii) importance-based retention of simple and detailed regions at inference. A variance constraint on residuals between full- and reduced-budget renderings further enforces uniform error under compression. Our method yields smoother LOD degradation than prior work, with largest gains in low-budget regimes. Code will be released.


Consistency-Verified Backdoor Defense for Federated Graph Learning via Cross-Layer Drift

Yanwen Jia ⋅ Yinuo Zhang ⋅ Zitong Shi ⋅ Yuxin Wu ⋅ Zihan Tan ⋅ Xuankun Rong ⋅ Liangtao Zheng ⋅ Wenke Huang ⋅ Mang Ye

Federated graph learning (FGL) enables privacy-preserving collaborative training over distributed graph data, yet remains vulnerable to backdoor attacks from malicious clients. In FGL, data and structural heterogeneity make benign clients difficult to distinguish from malicious ones, posing a key challenge to reliable defense. We propose VeriDrift, a backdoor defense framework for FGL based on cross-layer drift and consistency verification. On the client side, VeriDrift derives a client-level risk score from node-level cross-layer drift anomalies during GNN propagation. On the server side, VeriDrift verifies the consistency between the reported risk and the submitted model update, and converts the resulting verified risk into aggregation weights to suppress suspicious updates. Without relying on auxiliary clean data, VeriDrift enables reliable risk assessment under heterogeneous FGL and improves robustness against adaptive attacks. Extensive experiments show that VeriDrift consistently reduces attack success rates across different FGL backdoor settings while maintaining clean accuracy.


Consistent 3D Surface Flow Model with Global State

Antoine Guedon ⋅ Shu Nakamura ⋅ Nicolas Dufour ⋅ Jiahui Lei ⋅ Ko Nishino ⋅ Angjoo Kanazawa

Geometry is invariant to viewpoint, which makes any collection of images a redundant encoding of a single 3D state. Existing feed-forward reconstruction models fail to exploit this: per-view methods emit overlapping, unaligned pointmaps that grow linearly with input count, while global-latent methods commit to a fixed, low-resolution output. We introduce Surflo, which compresses a variable number of unposed RGB views into K latent tokens-one global state-and decodes oriented 3D surface points by independently transporting them from noise onto the surface via flow matching. This frees the output from any fixed grid or token budget: the same latent yields from a few thousand to a million points in a single forward pass. To suppress the local inconsistencies inherent to independent per-point decoding, an inference-time guidance term correlates nearby points by injecting a photometric gradient during ODE integration. Surflo matches or surpasses feed-forward baselines on surface metrics, run an order of magnitude faster than optimization-based methods that require hundreds of views, and is the only feed-forward approach to combine a global latent with arbitrary-resolution decoding.

Learning from weak labels (including noisy, partial or complementary labels) requires exploiting the statistical dependence between ground-truth classes and observed labels. Many existing methods are built around a particular transition model, often assumed to be known, accurately estimated, or instance-independent, which leads to performance degradation under model misspecification. We introduce Convex Losses for Weak Labels (CLWL), a family of losses for weak supervision based on transforming supervised one-vs-all losses. By prioritizing classification and ranking consistency over probability calibration, CLWL relaxes rigid structural assumptions on the label transition mechanism. We provide conditions for convexity and ranking consistency, characterize the transition matrices that preserve consistency for a given loss, and show that this consistency-preserving space attains the theoretical dimension bound. We propose a constructive method for generating parametric, lower-bounded, convex losses adapted to arbitrary weak-label models. Finally, we connect CLWL to existing methods, including backward correction and some losses for partial or complementary labels. Empirical results demonstrate that CLWL outperforms state-of-the-art methods when the assumed transition model is misspecified.

Determinantal point processes allow efficient sampling of diverse subsets of data. However, (exact) sampling under additional constraints is often computationally intractable under complexity-theoretic assumptions. We take an algebraic approach to designing novel algorithms that allow exact constrained sampling with theoretical efficiency guarantees. Our algorithms apply to a general class of point processes based on Bombieri inner products of polynomials, which we call Bombieri $k$-point processes. This perspective recovers known algorithms for constrained $k$-DPPs and extends them to a broader class of simple $k$-point processes. As applications, we obtain exact samplers for matching-, path-, subgraph-, and budget-constrained processes, as well as distributions combining diversity with clustering. While common complexity-theoretic assumptions rule out polynomial-time algorithms in many of these settings, our algorithms are efficient with respect to the relaxed notion of fixed-parameter tractability when the sample size $k$ is kept bounded.


Contrastive Hypergraph Source-free Domain Adaptive Object Detection in Adverse Weathers

Jicheng Yuan ⋅ Duc Manh Nguyen ⋅ Tung Kieu ⋅ Manfred Hauswirth ⋅ Danh Le-Phuoc

Adverse weather conditions such as fog, rain, snow, and low illumination introduce structured visibility degradations that severely impair object detectors trained on clear-weather data. When source data is inaccessible due to privacy or transmission constraints, source-free domain adaptation (SFDA) becomes necessary; however, existing SFDA methods based on pairwise contrastive learning struggle to capture high-order semantics and often suffer from agreement collapse under severe corruption. We propose REUNION, a hypergraph-guided SFDA framework for object detection under adverse weather. REUNION constructs a contrastive hypergraph over target-domain object proposals, encoding high-order relations through intra-image context, weather-aware grouping, prototype-based semantic anchors, and uncertainty-aware connections between reliable and low-confidence instances. To effectively exploit these structured relations, we introduce a group-wise Hyper-InfoNCE objective that optimizes representations at the hyperedge level, enabling semantic information to propagate from confident proposals to corrupted or low-contrast instances. Experiments on diverse benchmarks demonstrate that REUNION effectively mitigates domain shifts and achieves state-of-the-art performance in SFDA settings. Code is available at this \href{https://anonymous.4open.science/r/sfda-1E21/}{\textit{anonymous link}}.


Controlling for Omitted Variable Bias in Deep Neural Networks

Manuel Pfeuffer ⋅ Roshan P Rane ⋅ Kerstin Ritter ⋅ Sonja Greven

Control variables are widely used in statistical modelling to account for omitted variable bias. However, they have largely been underexplored in deep learning. This is surprising, given that deep learning models encode image-inferable covariates, such as demographic variables, into their predictions when these covariates are correlated with the outcome---a form of omitted variable bias referred to as 'shortcut learning'. While many existing confound-control or fairness methods try to restrict the correlation of such covariates with model predictions, we show that this fails to correct for omitted variable bias. We therefore propose a control variable approach for deep learning models, based on generalised additive modelling of the effects of model inputs and covariates. As flexible additive models can suffer from concurvity, we introduce an estimation procedure that refits the final layer of a pre-trained network to include covariate effects, using cross-fitting with ridge penalisation. We show how these effects can be orthogonalised with respect to covariates to exclude their mediated effects and that model predictions can be marginalised over the covariate distribution to control for their effect. This yields unbiased, interpretable predictions and offers flexibility to model the desired effects depending on the scientific or fairness objective. We verify our approach using simulated images, where it recovers the true covariate effects, even in small samples. Existing methods either require more data or fail to recover the true effects. We apply our method to real neuroimaging data with experimentally induced confounding, where it recovers prediction performance to near the level of a model trained on unconfounded data.

Building an optimal controller requires solving two coupled problems: (1) computing beliefs about the state by filtering observations, and (2) designing a control signal based on those beliefs to minimize a cost function. Ideally, these two problems can be solved independently: first, a filter is computed to estimate the state; then, a controller is built on top of that estimate. When exact inference is tractable, this approach is optimal and applies to the prominent Linear-Quadratic-Gaussian stochastic control problem. However, perfect inference is generally intractable, and filtering and control are closely intertwined problems. The question, then, is whether the commonly used first-filter-then-control approach remains optimal, and whether optimal control strategies should generally rely on internal representations that mirror the dynamics of the external world. We show that this is not the case: even in a simple setting with linear dynamics, multiplicative noise, and internal noise, the optimal linear controller is characterized by internal forward dynamics that do not match the forward dynamics of the external state. Instead, the optimal controller relies on internal representations that mix estimation and control, and this mismatch becomes more pronounced as internal noise increases. Our work challenges standard approaches that prioritize external world modeling over control.

Offline RL methods often regularize toward the empirical occupancy of the dataset, but this reference can mismatch the deployed policy. In partially observed or decentralized settings, data may be pooled from hidden modes, histories, or conventions that are unavailable at execution time. The pooled occupancy can then require incompatible action laws for the same deployed input, or joint correlations that a decentralized policy class cannot represent, so regularization toward it can favor a low return projection. We formalize this failure through *reference coherence*: a reference is coherent when it does not rely on information hidden from the deployed policy, and when its action law is representable by that policy. We propose *Closest Slice DICE (CS-DICE)*, a DICE variant that regularizes toward a coherent slice of the training data rather than the pooled occupancy. A hard or soft selector chooses the slice used as the reference term, while the deployed observations, actor class, and factorization constraints remain unchanged. Most of the results focus on the soft reverse-$KL$ case, where the selector only changes the reference used by an otherwise standard DICE update. Controlled examples show that incoherent pooled references collapse to suboptimal behavior, while CS-DICE recovers high return policies when the selected or inferred slice is coherent with the deployed interface. We also identify a boundary case where pooling is benign because the relevant signal is revealed before the conflict. Finally, a CS-CoMA-DICE study on MaMuJoCo shows that selected references can be integrated into a neural cooperative MARL pipeline while preserving decentralized execution.


DACE: Diversity-Driven Adversarial Co-Evolution for Robust LLM Safety Alignment

Peng Yu ⋅ Xiaoyu Wen ⋅ Zhida He ⋅ Ziyuan Zhou ⋅ Han Qi ⋅ Shao Zhang ⋅ Qiaosheng Zhang ⋅ Ying Wen ⋅ Chaochao Lu

Ensuring robust safety alignment of large language models (LLMs) is increasingly difficult as adversarial attacks evolve and outpace alignment pipelines built on pre-collected data. Recent co-evolutionary frameworks let the defender chase a moving attacker, yet their training dynamics expose two compounding pathologies: \emph{attack strategy collapse}, where the attacker overfits to a narrow set of high-reward rewrites, and \emph{defense adversarial forgetting}, where the defender loses competence against earlier attacks as the attack distribution drifts. Existing remedies are partial: attacker-side diversity rewards score textual novelty rather than strategy novelty, and defender-side replay relies on discrete judge scores that ignore estimation uncertainty, response stochasticity, and defender non-stationarity. We introduce DACE, a diversity-driven adversarial co-evolution framework that targets both pathologies jointly. On the attacker side, DACE couples an explicit $12\times10$ strategy space (risk category $\times$ attack style) with a \emph{normalized marginal coverage gain} reward, providing a bounded, non-vanishing exploration signal at the strategy layer. On the defender side, DACE maintains a unified \emph{Bayesian adversarial replay pool} whose Beta--Bernoulli threat posteriors are refreshed by time decay and sampled via Thompson sampling, anchoring defender training to an evolving threat landscape. Across standard safety, automated-attacker, and general-capability benchmarks, DACE improves robustness to out-of-distribution attacks while preserving general capabilities, and yields broader attacker strategy coverage.


Dandelions: A Spherical Flower for Neural Simulation of Planetary Dynamics

Till Muser ⋅ Giovanni Abati ⋅ Ivan Dokmanić

Many dynamical processes unfold on the sphere but the default scientific machine learning architectures are Euclidean. Applying these architectures on a regular lat--lon grid causes problems: Cartesian convolutions become distorted at high latitude; 2D FFTs in Fourier neural operators incorrectly assume double periodicity; Cartesian positional encodings in ViTs distort spherical geodesic distances. Recent work moves towards natively spherical primitives, including spherical convolutions (e.g., DeepSphere or DISCO), Spherical Fourier Neural Operators (SFNOs), and geodesic attention. Here we propose Dandelion, a spherical version of Flower---a recent warp-based neural PDE solver. Layers of Dandelion predict a tangent-plane displacement and transport features along great circles. We obtain a U-Net-like structure by implementing hierarchical pooling entirely in the spherical-harmonic domain. There are thus no convolutions: spatial mixing is achieved only through spherical coordinate changes---or warps. To compare Dandelion with existing spherical architectures, we release an evolving benchmark suite of challenging, natively-spherical PDE datasets including a modified Galewsky jet, anomalous chained turbulence, Cahn--Hilliard decomposition, spherical Riemann shocks, Held--Suarez dry atmospheric transport and global ocean dynamics. This new benchmark fills the gap in existing spherical datasets which are either too small and stylized, or much too large (ERA5) for model iteration. Dandelion is best or second-best on every dataset, and the gap to non-warp baselines widens with resolution: at $256\times 512$, Dandelion and Flower2D occupy the top two slots in both single-step prediction and rollout.


Decision-Aware Proximal Bridge Learning for Optimal Treatment Selection

Tomas Garriga ⋅ Alejandro Almodóvar ⋅ Axel Brando ⋅ Gerard Sanz ⋅ Eduard Serrahima de Cambra ⋅ Juan Parras

Individualized treatment selection with continuous actions requires accurate causal response estimation in decision-relevant regions, rather than uniformly over the entire action space. Estimating a global causal response surface and then choosing the treatment that maximizes it can therefore be suboptimal, since standard estimation objectives allocate modeling effort according to the observed treatment distribution rather than the regions that determine the optimal decision. While decision-aware approaches have been studied in unconfounded settings, this problem remains underexplored in proximal causal inference, where proxy variables and bridge functions enable identification under suitable assumptions even in the presence of hidden confounding. Despite recent progress, proximal methods have primarily focused on treatment-effect and potential-outcome estimation rather than treatment selection and optimal decision-making. To bridge this gap, we introduce a policy-targeted weighted bridge loss that emphasizes decision-relevant treatment regions while retaining global stabilization. We prove a regret bound showing that the proposed weighted bridge loss controls treatment-selection regret through a weighted ill-posedness constant. We instantiate the framework in decision-aware variants of several proximal bridge solvers, yielding practical algorithms that alternate between weighted bridge estimation, response-surface projection, policy update, and weight refinement. Empirically, we find that decision-aware weighting reduces regret across several bridge solvers, suggesting improved treatment selection in proximal settings.


Deep Learning as Neural Low-Degree Filtering: A Spectral Theory of Hierarchical Feature Learning

Yatin Dandi ⋅ Matteo Vilucchio ⋅ Luca Arnaboldi ⋅ Hugo Tabanelli ⋅ Florent Krzakala

Understanding how deep neural networks learn useful internal representations from data remains a central open problem in the theory of deep learning. We introduce \emph{Neural Low-Degree Filtering} (Neural LoFi), a stylized limit of gradient-based training in which hierarchical feature learning becomes an explicit iterative spectral procedure. In this limit, the dynamics at each layer decouple: given the current representation, the next layer selects directions with maximal accessible low-degree correlation to the label. This yields a tractable surrogate mechanism for deep learning, together with a natural kernel-space interpretation. Neural LoFi provides a mathematically explicit framework for studying feature learning beyond the lazy regime. It predicts how representations are selected layer by layer, and gives a concrete mechanism by which depth progressively constructs new features from old ones. We complement the theory with mechanistic experiments on fully connected and convolutional architectures, showing that Neural LoFi improves over random-feature baselines, recovers meaningful structured filters, and predicts representations aligned with early gradient-descent feature discovery.


Diffusion Models Observe Only Gradients: A Geometric Perspective on Score Matching Errors

Nail B Khelifa ⋅ Richard Turner ⋅ Ramji Venkataramanan

Score-based diffusion models are typically trained by minimizing the $L^2$ score matching error, and standard theoretical analyses rely on this quantity to bound the sampling discrepancy between the learned and target distributions. We show the $L^2$ score error is not the right intrinsic measure of marginal distributional quality: a learned diffusion model can incur arbitrarily large $L^2$ score error while perfectly matching the target distribution. By decomposing score errors into a gradient and a solenoidal component (a Helmholtz-Hodge decomposition), we identify the geometric reason behind this: only the gradient component enters the marginal Fokker-Planck dynamics, while the solenoidal component is structurally invisible. We make this precise in three results. First, building on the corrected geometry, we prove an impossibility result: no monotone function of the $L^2$ score error can uniformly lower bound any divergence between the learned and target distributions. Second, we derive an upper bound on the Kullback-Leibler divergence that depends only on the observable gradient component of the error, tightening the standard Girsanov bound and identifying its looseness as the cost of operating on path-space rather than marginal-space dynamics. Third, we give a tractable estimator of the gradient component via a dual Sobolev identity, which is shown to empirically correlate substantially better with sample quality than the full $L^2$ error.


Disentangling generalization and memorization in large language models using chess

Leonard S. Pleiss ⋅ Maximilian Schiffer ⋅ Robert K von Weizsäcker

Large Language Models (LLMs) exhibit remarkable capabilities, yet it remains unclear to what extent these reflect sophisticated recall or genuine reasoning ability. We introduce chess as a controlled testbed aimed at disentangling these faculties. Leveraging the game’s structure and scalable engine evaluations, we construct a taxonomy of positions varying in density of relevant priors - ranging from common states solvable by memorization to completely novel ones requiring generalization. Crucially, our approach achieves this distinction without requiring explicit knowledge of the models' training data. Applying this taxonomy, we combine a longitudinal analysis of the GPT lineage with a rigorous evaluation of contemporary models, including Claude Opus and Gemini. Our analysis reveals a clear gradient: performance consistently degrades as the density of relevant priors decreases. Notably, for tasks with few relevant priors, base model performance regresses to the random-play baseline. While newer models improve, progress slows significantly for tasks with sparse priors. Furthermore, while reasoning-augmented inference improves performance, its relative marginal benefit per token decreases in the absence of relevant priors. These results suggest limitations in systematic generalization, highlighting the need for mechanisms beyond scale to achieve robust performance when deprived of relevant priors.


Doomed to Re-Annotate, Forever: The ImageNet Story

Illia Volkov ⋅ Nikita Kisel ⋅ Tetiana Mishkina ⋅ Klara Janouskova ⋅ Jiri Matas

Top-1 accuracy on ImageNet-1k remains the most universally reported metric in visual recognition, despite its well-documented label quality issues. In the paper, we describe the effort to obtain accurate and complete ImageNet-1k validation set annotations. The result -- ReImageNet -- is a from-scratch reannotation, which includes multilabel correction, object localisation, revised class definitions, and semantic attributes---text-recognition, rendition, reflection, crowd, and dominant. The reannotation reveals that $\approx13$\% of the original ImageNet-1k labels are incorrect, $32.7$\% of images are multilabel and $4.7$\% contain no object from an ImageNet-1k class. With the new labels, top-1 accuracy increases by $0.7$--$1.4$\% for supervised models and by $6$--$8$\% for MLLMs. We show that annotation errors in ImageNet-1k propagate into its derivative test sets, indicating that the problem is structural rather than specific to any single benchmark. We observed that LLMs outperform unaided crowd annotators, and that human and LLM collaboration with appropriate tooling represents the current quality ceiling for annotation at this scale. All [annotations](https://huggingface.co/datasets/c1rcuslegend/ReImageNet), [class definitions](https://gonikisgo.github.io/imagenet-annotation-preview/), [guidelines](https://gonikisgo.github.io/imagenet-annotation-preview/), and [analysis code](https://github.com/klarajanouskova/ImageNet) have been publicly released.

We study dynamic regret in online convex optimization with an \emph{indicator switching cost}: a fixed penalty incurred whenever two consecutive decisions differ. This captures startup overheads such as server activation, model deployment, and cache updates, and on a bounded domain it recovers norm-based movement costs as a special case. Existing guarantees for indicator costs handle only static comparators. We show that a direct extension of these techniques to dynamic regret provably fails, motivating a different approach. We propose a meta-learning framework: a set of randomized lazy FTRL base learners restarted at dyadic time scales, aggregated by a movement-aware master that mixes their proposal densities and samples actions via maximal coupling of consecutive mixtures. The resulting algorithm satisfies, in expectation, $\mathcal{R}^{\mathbf{1}}_T \le \tilde{\mathcal{O}}(\min\{\sqrt{T(S_T{+}1)},T^{2/3}(P_T+1)^{1/3}\})$, where $\mathcal{R}^{\mathbf{1}}_T$ is the dynamic regret plus the cumulative indicator switching cost, $S_T$ counts comparator switches, and $P_T$ is the comparator path length. The bound holds simultaneously for all sequences and requires no prior knowledge of $S_T$ or $P_T$: it is minimax-optimal (up to logarithmic factors) for tracking piecewise-constant comparators, and also captures frequently moving comparators with small total path length.


E²Gen: Evidential Energy-Based Generation for Fair Federated Graph Learning

Jingbo Wang ⋅ Zitong Shi ⋅ Yuxin Wu ⋅ Zihan Tan ⋅ Qiqi Lin ⋅ Xuankun Rong ⋅ Liangtao Zheng ⋅ Wenke Huang ⋅ Mang Ye

Federated Graph Learning is rapidly evolving as a privacy-preserving collaborative approach for decentralized graph data. However, severe fairness challenges are increasingly undermining federated systems by systematically degrading predictions for structurally disadvantaged minority nodes. The inherent vulnerabilities and missing topological contexts in Federated Graph Learning are deeply entangled, making traditional federated fairness methods and simple oversampling less effective. In our work, we propose an effective Evidential Energy-based Generation framework for Fair Federated Graph Learning ($E^2Gen$). At the client level, it explicitly identifies structurally deficient nodes using a multi-axis metric, synthesizing targeted representations via conditional energy-based models, and selects reliable samples through an evidential quality gate. At the server level, the local performance disparities uploaded by each client are evaluated to construct a fairness gap assessment, making the global model absorb equitable improvements by further adjusting the aggregation weights. Our method can handle high topological heterogeneity, does not require strict generative normalization, and is effective under both homophilic and heterophilic graph structures. Extensive results on various settings of federated graph scenarios under severe fairness challenges validate the effectiveness of this approach. The code is anonymously available at https://anonymous.4open.science/r/E-Gen-A2C6.


EEGFaceSem: An EEG Benchmark with Paired Generative Latents for Semantic Visual Modeling

Jun Ma ⋅ Michiel Spape ⋅ Keith M Davis ⋅ Tuukka Ruotsalo

Electroencephalography (EEG) is widely used to measure brain activity for brain-computer interface applications for its non-invasive nature and high temporal resolution. However, progress in modeling the neural responses to visual semantics has been limited by a reliance on static, real-world image datasets that lack the fine-grained control needed to isolate fine-grained changes in visual features and their relation to perception. In this work, we present EEGFaceSem, a large-scale dataset pairing high-quality EEG with photorealistic face images synthesized from a generative model. We provide the generative latent vector for each stimulus, enabling research into the continuous mapping between a visual latent space and neural representations. Each participant's brain responses were recorded while they attended to specific semantic features (e.g., smiling, hair color) in a rapid serial visual presentation paradigm. The dataset contains EEG signals from 30 participants, totaling 67,200 time-locked EEG recordings. We present benchmarks for single-trial semantic classification under single-subject, cross-subject, and subject-adapted settings, as well as for cross-subject voting that aggregates per-stimulus predictions across participants. We further demonstrate the dataset's unique utility with two complementary applications: deriving semantic directions in the latent space for brain-supervised image editing, and conditioning a generative model on brain-decoded labels for brain-supervised image generation. Our dataset paves the way for advancing brain-computer interface (BCI) devices for semantic-level visual saliency detection and brain-supervised generative modeling. Our dataset and code are openly released at OSF https://osf.io/2guk3/ and at HuggingFace https://huggingface.co/datasets/yefllower/EEGFaceSem.


Efficient Matrix Product State Learning in Logarithmic Depth

Chia-Ying Lin ⋅ Nai-Hui Chia ⋅ Shih-Han Hung

Learning the closest matrix product state (MPS) representation of a quantum state is known to enable useful tools for quantum machine learning and analysis of complex quantum systems. In this work, we study the problem of learning MPS in the following setting: given many copies of an input MPS, the task is to recover a classical description of the state. The best known polynomial-time algorithm, introduced by [LCLP10, CPF+10], requires linear circuit depth and $O(n^5)$ samples, and has seen no improvement in over a decade. The combination of linear circuit depth and large sample complexity, neither known to be optimal, renders existing algorithms impractical for near-term quantum devices with limited resources. We introduce parallel disentangling algorithms for MPS learning. For exact MPS learning, our algorithm runs in polynomial time and uses circuit depth $O(\log n)$ and sample complexity $\widetilde O(n^3)$, improving both the depth and the dependence on the system size $n$. The key idea is to exploit the bounded-rank structure of reduced states on middle blocks of an MPS and organize the disentangling operations in a tree structure. We further extend the algorithm to closest MPS learning, improving the sample complexity dependence on $n$ from $n^9$ to $n^7$ and complement the algorithms with an $\Omega(n)$ product-state lower bound. We also investigate MPS learning under hardware constraints, including restricted measurements and geometric connectivity. Under the Learning Parity with Noise (LPN) assumption, we show computational hardness for learning an MPS(2) family with non-adaptive single-qubit measurements. Finally, we show that our algorithm can be implemented with depth $O(q n^{1/q})$ on a $q$-dimensional hypercubic lattice, giving an asymptotic reduction in depth. Together, our work provides a complete characterization of the quantum resources needed for efficient MPS learning.

We study the problem of recovering a globally consistent Euclidean embedding of data, given only a local distance graph and propose a method that optimally represents these distances. The method operates solely on a neighborhood graph weighted by pairwise distances, without requiring any prior vector representation of the data. The embedding is obtained by solving a variational problem that matches local, on‑graph distances to the Euclidean metric, induced by the differentials of the embedding functions. The resulting Euler–Lagrange equations are derived in a coordinate‑free form, enabling direct evaluation of all operators from the distance graph alone. Though non-linear and missing an explicit expression for their non-linearity, these equations are shown to be resolved as an iteratively updated sparse linear problem. The main contributions of the proposed approach are (a) the derivation of the functional equations governing the optimal Euclidean embedding in the continuum, (b) a representation‑free formulation that requires only a neighborhood distance graph and no feature vectors and (c) an estimation procedure based exclusively on local graph operations. We experimentally evaluate the resulting non‑parametric algorithm on synthetic manifolds and real datasets, demonstrating consistent preservation of local metric structure and neighboring relations, while approximating the global isometric embedding.


Evaluating Epistemic Uncertainty: Beyond OOD Detection and Active Learning

Jakub Paplhám ⋅ Vojtech Franc ⋅ Willem Waegeman ⋅ Eyke Hüllermeier

Current evaluation of epistemic uncertainty relies on tasks such as out-of-distribution detection and active learning. However, the Bayes-optimal decision strategies for these tasks do not coincide with the scores commonly used to quantify epistemic uncertainty. Building on the epistemic reject-option framework, we evaluate epistemic uncertainty using its ability to identify regret, the reducible error. Formulating selective prediction as a constrained optimization over coverage, expected risk, and regret, we prove the optimal selector is a thresholded convex combination of the ground-truth aleatoric and epistemic uncertainties. This theoretical unification exposes a weakness in recent uncertainty disentanglement literature: we demonstrate that standard correlation metrics between learned components do not necessarily predict their actual operational utility. We instead propose to evaluate the achievable (risk, regret, coverage) surface of the decomposition as a diagnostic for joint disentanglement and utility. Benchmarking standard methods on datasets with dense human annotations reveals that decision-theoretic rankings can disagree substantially with proxy-task rankings, including pairwise rank inversions between methods that are top-ranked on one criterion and bottom-ranked on other.


EV-AUDIT: A Co-Evolutionary Auditing Framework for Task Hijacking in Multi-Agent Systems

Tristan Bilot ⋅ Zhilu Zhang ⋅ Kay Liu ⋅ Mikhail Kuznetsov ⋅ Wei Ding

Multi-agent systems (MASs) are increasingly deployed for complex tasks, yet their security against indirect prompt injection remains poorly understood. As frontier LLMs become more robust to injections with overtly malicious semantics, such as credential exfiltration, unauthorized transactions, or policy bypass, the residual attack surface shifts to task hijacking: an injection that silently redirects an MAS toward an attacker-chosen alternative task plausibly within the agent's domain, such as the wrong patient's record, a competitor's product, or a different CVE. There is no malicious verb to refuse, and existing benchmarks, designed around overtly malicious goals in single-agent settings, do not characterize this surface. We introduce EV-AUDIT, an auditing framework that lets a practitioner stress-test their own MAS against task hijacking and derive a tailored system-prompt defense. The framework couples 22 baseline attacks and four prompt/tool-level defense baselines from the literature with a co-evolutionary red/blue-team procedure that evolves stronger attacks and adaptive defenses by reasoning over execution traces. Applied across frontier backends on ten MASs and a real-world MAS with live web access, EV-AUDIT surfaces an MAS-specific attack pattern that frames the injection as a prerequisite for completing the user's task, reaching 31.26% ASR on Claude Opus 4.5 under the framework's auditing protocol, where 22 baseline attacks fail entirely. The co-evolved defense suppresses both baseline and evolved attacks at negligible utility cost, transfers across reasoning-capable backends, and operates at the prompt level with low deployment overhead.


Expected Batch Optimal Transport Plans and Consequences for Flow Matching

Samuel Boïté ⋅ Julie Delon ⋅ Kimia Nadjahi

Solving optimal transport (OT) on random minibatches is a common surrogate for exact OT in large-scale learning. In flow matching (FM), this surrogate is used to obtain OT-like couplings that can straighten probability paths and reduce numerical integration cost. Yet, the population-level coupling induced by repeated minibatch OT remains only partially understood. We formalize this coupling as the expected batch OT plan $\overline{\pi}\_{k}$, obtained by averaging empirical OT plans over independent minibatches of size $k$. We then establish its large-batch consistency and, in the semidiscrete case relevant to generative modeling, derive rates for both the transport-cost bias and the convergence of $\overline{\pi}\_{k}$ to the OT plan. For FM, this yields a population coupling whose induced velocity field is regular enough to define a unique flow from the source to the discrete target. We finally quantify how OT batch size interacts with numerical integration in a tractable two-atom model and in synthetic and image experiments.

As the demand to integrate Artificial Intelligence into high-stakes environments continues to grow, explaining the reasoning behind neural-network predictions has shifted from a theoretical curiosity to a strict operational requirement. Our work is motivated by the explanations of autoregressive neural predictions on dynamic physical fields, as in weather forecasting. Gradient-based feature attribution methods are widely used to explain the predictions on such data, in particular due to their scalability to high-dimensional inputs. It is also interesting to remark that gradient-based techniques such as SmoothGrad are now standard on images to robustify the explanations using pointwise averages of the attribution maps obtained from several noised inputs. Our goal is to efficiently adapt this aggregation strategy to dynamic physical fields. To do so, our first contribution is to identify a fundamental failure mode when averaging perturbed attribution maps on dynamic physical fields: stochastic input perturbations do not induce stationary amplitude noise in attribution maps, but instead cause a geometric displacement of the attributions. Consequently, pointwise averaging blurs these spatially misaligned features. To tackle this issue, we introduce WassersteinGrad, which extracts a geometric consensus of perturbed attribution maps by computing their entropic Wasserstein barycenter. The results, obtained on regional weather data and a meteorologist-validated neural model, demonstrate promising explainability properties of WassersteinGrad over gradient-based baselines across both single-step and autoregressive forecasting settings.


Expressive Power of Deep Homomorphism Networks over Relational Databases

Balder ten Cate ⋅ Maurice Funk ⋅ Benny Kimelfeld ⋅ Carsten Lutz ⋅ Moritz Schönherr ⋅ Arie Soeteman

The expressive limitations of message-passing Graph Neural Networks (GNNs) have motivated a wide range of more powerful graph learning architectures. We advocate Deep Homomorphism Networks (DHNs) as a model particularly well-suited for learning over relational databases, due to their close connection to important fragments of SQL such as conjunctive queries. We study the precise expressive power of DHNs by relating them to various natural fragments and extensions of first-order logic (FO). For DHNs with max, sum, and mean aggregations, we establish connections to the unary negation fragment (UNFO) and to the extensions of UNFO with counting quantifiers and with ratio quantifiers. We further relate sum-aggregation DHNs to the unary quantifier alternation fragment of FO and to an extension of FO with expressive counting. Through the classical correspondence between FO and SQL, these results also illuminate the relation between DHNs and SQL. They also enable us to study the decidability of two fundamental static analysis problems for DHNs, the emptiness problem and the subsumption problem. Finally, we confirm through experiments that the established differences in expressive power are reflected in the performance on suitable prediction tasks.

Mathematical autoformalization can fail before proof generation, when an informal concept is grounded to the wrong Lean/Mathlib object. Multilingual terminology makes this step non-monotonic: aliases can recover missing concepts, but under collision stress they can also pull retrieval toward nearby, non-interchangeable anchors. We introduce FACBench, a targeted dependency-light benchmark and audit protocol for multilingual formal-anchor grounding before proof generation. FACBench combines a five-language concept inventory with controlled collision queries, Mathlib-graph stress tests, no-label ablations, theorem-like probes, and Lean-facing anchor checks. In controlled collision settings, naive multilingual alias union can place the expected anchor in the candidate list while ranking a nearby distractor first. A simple rule-based source/context guard improves top-1 anchor accuracy over naive multilingual retrieval by about 11 percentage points on controlled collisions and 28 percentage points on a Mathlib-graph-mined controlled holdout. Less scaffolded and theorem-like probes show where this guard stops helping and point to the need for role-aware grounding. FACBench provides reusable collision-focused evaluation and regression-test scenarios for retrieval-augmented autoformalization systems.


Fail-Closed Alignment for Large Language Models

Zachary Coalson ⋅ Sanghyun Hong

We identify a structural weakness in current large language model (LLM) alignment: refusal mechanisms in these models are fail-open. While existing approaches tend to encode refusal behaviors across multiple latent features, suppressing a single feature (via prompt-based jailbreaks) is sufficient to collapse alignment, leading to unsafe generation. Motivated by this, we propose fail-closed alignment as a design principle for robust LLM safety: refusal mechanisms should remain effective even under partial failures via redundant, independent causal pathways. We present a concrete instantiation of this principle: a progressive alignment framework that iteratively identifies and ablates previously learned refusal directions, forcing the model to reconstruct safety along new, independent subspaces. Across six jailbreak attacks, we achieve the strongest overall robustness (1.7% average attack success rate) while largely preserving generation quality and preventing excessive over-refusals, with negligible computational overhead. Additional analyses confirm that models trained with our method encode multiple, causally independent refusal directions that prompt-based jailbreaks cannot fully suppress, providing empirical support for fail-closed alignment as a principled foundation for robust LLM safety.


𝑓-Differential Privacy Filters: Validity and Approximate Solutions

Long Tran ⋅ Antti Koskela ⋅ Ossi Räisä ⋅ Antti Honkela

Accounting for privacy loss under fully adaptive composition---where mechanism choice and privacy parameters may depend on the history of prior outputs---is a central challenge in differential privacy (DP). Here, privacy filters are stopping rules ensuring a prescribed global budget is not exceeded. A leading candidate for optimal filter design is 𝑓-DP, which characterizes the full extent of adversarial hypothesis testing and recovers (ε,δ)-DP through piece-wise linear trade-off functions, while enabling tight (ε,δ)-DP accounting in standard compositions via tensor products. Yet whether such filters can be correctly defined under 𝑓-DP remains unclear. We show that the natural 𝑓-DP filter---tracking path-wise accumulating tensor products and stopping when the prescribed curve is crossed---is fundamentally invalid, precluding the direct use of standard efficient numerical Fast-Fourier-Transform accounting in the fully adaptive setting. We characterize this failure, establishing necessary and sufficient conditions for the natural 𝑓-DP filter's validity. Furthermore, we prove a fully adaptive central limit theorem for 𝑓-DP, establishing Gaussian convergence of cumulative privacy losses under full adaptivity. As a demonstration, we construct a closed-form approximate GDP filter for subsampled Gaussian mechanisms that provably outperforms RDP-based accounting in asymptotic regimes ($q\ll 1$ and $q\approx 1$) without tracking the full trade-off function, demonstrating that the slack in RDP is not intrinsic to adaptive composition---though CLT-based approximations are known to be optimistic at realistic subsampling rates, a limitation that remains an open challenge.


FedTrace: Generated-Content-Based Watermark Verification for Traitor Tracing in Federated Learning

Tianzhe Xiao ⋅ Haozhao Wang ⋅ Yichen Li ⋅ Qiyu Qin ⋅ Debin Liu ⋅ Ruixuan Li

As large generative models become widely deployed and customized, federated learning is increasingly used to adapt them while keeping user data local. Because clients repeatedly receive up-to-date adapters, a malicious client can copy a dispatched adapter, use it offline, and monetize generated images without exposing the stolen weights or a queryable service. Existing federated watermarking and traitor-tracing methods usually assume white-box access to suspect weights or black-box query access to the deployed model. In realistic generative-model theft, however, the defender may only observe images that have already circulated. We propose \emph{FedTrace}, a generated-content-based watermark verification framework that attributes leaked federated models from suspicious generated images alone. FedTrace couples three designs: a round-wise watermark lifecycle that separates client identity distribution from global utility aggregation, a low-drift reliable-bit carrier that embeds identity in watermark positions stable under local adaptation, and anti-collision coding with soft subset verification for distinguishing singleton and collusive leakage. These components enable post-local, collusion-aware attribution from generated outputs alone, without suspect weights or online query interfaces. Extensive experiments across customized diffusion datasets, base models, and federated settings show that FedTrace preserves detectable client identities under local adaptation and strengthens subset-aware tracing against collusive leakage.


FIVE-VLA: Fast and EffectIVE Closed-Loop Autonomous Driving with Recurrent Action Memory

Kemal Oksuz ⋅ Alexandru Buburuzan ⋅ Yuhan Yao ⋅ Puneet Dokania

State-of-the-art vision-language-action models (VLA) for autonomous driving face critical limitations: excessive parameter counts, inefficient high-resolution image processing, and lack of temporal memory. We introduce **FIVE-VLA** (Fast and EffectIVE VLA) to address these through two key contributions. First, we employ an efficient vision encoder that processes high-resolution ($448 \times 896$) images while generating only 98 tokens, over $5\times$ fewer than existing approaches, and bypass text generation entirely for single-pass trajectory prediction. Second, we propose **Recurrent Action Memory (RAM)**, a lightweight module that conditions action prediction on previous action tokens, providing temporal context critical for manoeuvres such as overtaking and emergency braking. With only 641M parameters, FIVE-VLA completes *$\sim$10\% more routes without traffic rule infractions* than the previous state-of-the-art VLA on the challenging \textit{Bench2Drive} closed-loop driving benchmark. Furthermore, FIVE-VLA runs at $\sim$30 fps on an A100 and $\sim$4 fps on a T4 GPU (proxy to an edge device), representing an *8--30$\times$ speedup* over previous methods. The code will be made public upon acceptance.


Focus, Align, and Diffuse: Time-Series-Aware Keyframe Diffusion for Cardiac Dynamic Synthesis from Sparse Observations

Junkai Liu ⋅ Haofan Wu ⋅ Nay Aung ⋅ Joao A Lima ⋅ Steffen E Petersen ⋅ Le Zhang

Dense cardiac dynamics are essential for assessing cardiac function, yet in practice clinicians often observe only extremely sparse measurements. However, recovering full dynamics from such sparsity is a highly ill-posed problem due to the lack of temporal cues. While auxiliary time-series signals like ECG provide physiological guidance, effectively exploiting them is challenging due to the cross-modal representation gap and phase misalignment. In this work, we present CardioFAD, a Focus--Align--Diffuse paradigm that reformulates sparse cardiac dynamic synthesis as a recoverability--alignability anchor discovery problem. Our core insight is that resolving extreme temporal ambiguity requires a compact set of anchor states that are both sufficient to recover the visual trajectory and reliable for ECG--visual alignment. We therefore learn motion-salient keyframes under two complementary principles: trajectory sufficiency, which preserves trajectory-critical intra-modal motion content, and anchor alignment, which promotes consistent inter-modal correspondence at these anchors. We implement the Focus and Align modules optimized by the two aspects.Diffuse then performs two-stage generation by first synthesizing keyframes and subsequently interpolating the remaining frames. Extensive experiments on cardiac MRI and echocardiography videos demonstrate superior visual quality, temporal coherence, and clinical fidelity under the extreme one-shot regime.


Focused Forcing: Content-Aware Per-Frame KV Selection for Efficient Autoregressive Video Diffusion

Peiliang Cai ⋅ Evelyn Zhang ⋅ Jiacheng Liu ⋅ Hao Lin ⋅ Ruiqi Zhang ⋅ Weile Mo ⋅ Yue Ma ⋅ Shikang Zheng ⋅ Jiehang Huang ⋅ Dongrui Liu ⋅ Linfeng Zhang

Recent advances in autoregressive video diffusion have enabled sequential and streaming video generation. However, long-horizon generation requires increasingly large KV caches, making efficient compression without sacrificing quality challenging. Existing methods mostly select historical frames based on attention scores, but their context decisions remain coarse. When multiple frames are generated in the same chunk, these methods often apply a shared history selection to the whole chunk, score historical frames solely by attention, and assign head-wise budgets either uniformly or by attention-pattern heuristics rather than explicit head-importance estimation. We show that frames within the same generated chunk can depend on distinct historical frames, that the same historical frame can receive different attention scores as its relative temporal distance to the current frames changes, and that masking different heads induces unequal generation degradation. Motivated by these findings, we propose \textbf{Focused Forcing}, a training-free KV selection method that focuses cached history along both generated-frame and head dimensions. For each generated frame, Focused Forcing preserves the most relevant and distinctive historical frames by combining attention scores with diversity scores of historical frames, while assigning larger budgets to heads with higher estimated importance. Across multiple autoregressive generation paradigms, Focused Forcing achieves up to $\textbf{1.48}\times$ end-to-end acceleration without training, while \textbf{improving visual quality and text alignment}. \textit{Our code will be released on GitHub.}


FogGS: Physics-Grounded 3D Foggy Effects

Jiaxin LI ⋅ Tailin Wu ⋅ Jiankang Deng ⋅ Li Zhang ⋅ Xiatian Zhu

Realistic and diverse fog synthesis is a critical capability in 3D scene modelling, with broad applications in autonomous driving simulation, cinematic visual effects, gaming, and AR/VR. Yet achieving this, encompassing varied density profiles, spatial distributions, and complex optical phenomena, in real-world scenes remains a largely unsolved challenge. Existing approaches rely on appearance-driven weather transfer or simple fog overlays, which lack the physical grounding needed for volumetric effects and thus cannot faithfully model key fog properties and effects, including density gradients, anisotropic scattering, and atmospheric light shafts (Tyndall effect) — \textit{limitations} that stem directly from the absence of explicit volumetric representation and physical light transport modeling. We present \textbf{FogGS}, a training-free, model-agnostic rendering framework that seamlessly enables existing pretrained fog-free 3DGS models to produce photorealistic, physically plausible, and fully editable fog in novel-view synthesis. Grounded in atmospheric radiative transfer theory, FogGS models fog as spatially varying extinction and scattering fields, using the depth buffer from Gaussian rasterization to drive efficient screen-space ray marching for geometry-aligned transmittance and in-scattering computation. This physically grounded parameterization affords intuitive control over fog intensity, appearance, and spatial distribution, and naturally supports anisotropic scattering, scene occlusion, and illumination. FogGS maintains high rendering efficiency practical for iterative editing workflows, remains compatible with relighting pipelines, and demonstrates superior realism, physical consistency, and controllability across diverse outdoor scenes.


From Anti-Forgetting to Fast Adaptation: Continual Reinforcement Learning with World Models

Qiyang Zhou ⋅ Yuliang Cai ⋅ Peng Wang ⋅ Senquan Yi ⋅ Yanming Li ⋅ Haobo Fu ⋅ Li Shen

Continual Reinforcement Learning (CRL) requires agents to adapt to sequentially arriving tasks while retaining performance on previous ones. Most existing CRL methods rely on model-free frameworks, addressing catastrophic forgetting via regularization, parameter isolation, or experience replay. World models offer a promising alternative by accumulating transferable dynamics knowledge and enabling planning-based decision-making. However, existing world model approaches suffer from slow adaptation at task boundaries: abrupt reward shifts mislead the planner, fixed planning horizons accumulate prediction errors when the dynamics model is unreliable, and quality-agnostic replay dilutes training while exacerbating forgetting. We propose HERD (Horizon-adaptive Elite Replay with reward Disagreement), which addresses these challenges from two complementary perspectives: for $\textbf{\emph{how to learn}}$, Uncertainty-aware Reward Planning leverages ensemble disagreement to accelerate reward calibration and guide exploration, while Adaptive Planning Horizon adjusts rollout depth based on dynamics prediction error to prevent error accumulation; for $\textbf{\emph{what to learn}}$, Elite Experience Replay applies quality-stratified reweighting of historical trajectories to strengthen knowledge retention. Together, these components $\textbf{\emph{accelerate new-task adaptation and mitigate catastrophic forgetting.}}$ Experiments on Continual World, Continual Bench, and DM Control show that HERD consistently outperforms existing methods in overall performance, forgetting mitigation, and new-task adaptation.

Feature attributions often hide a critical modeling choice: they explain a prediction along a counterfactual path from a reference state to an input. Different baselines, interpolations, and generative trajectories define different paths and can therefor produce different explanations. We study this path ambiguity as a modeling problem. Our central question is whether the path can be chosen by the data-generating transport process, rather than by a hand-designed interpolation or by the sensitivity geometry of the model being explained. We separate attribution into fixed-path credit allocation and path selection. For a fixed path, we prove that the Aumann-Shapley line integral is the unique attribution rule under standard fixed-path axioms and explicit coordinate-trace regularity. For path selection, we minimize kinetic action over flows that transport a reference distribution to the data distribution, yielding a transport-geodesic attribution principle. We approximate this ideal with Rectified Flow and Reflow and derive stability bounds linking vector-field error to attribution error. Experiments show that lower-action, transport-consistent paths produce more stable and structured explanations, preserving competitive deletion faithfulness, without claiming data-manifold membership. Our code is available at https://anonymous.4open.science/r/temp_OTshap-6119.

Persistence diagrams are common representations in topological data analysis, yet they lack both the vector space structure required by machine learning algorithms and a principled framework for statistical comparison. We introduce STRAND (Survival Topological Representation ANalysis of Diagrams), which treats (collections of) PDs as survival data: each topological feature with persistence value $p = d - b$ is a fully observed time-to-event, and the persistence survival function $S(t) = \mathbb{P}(p > t)$ is the central object for comparing diagrams. From this single representation we derive (i) a non-parametric two-sample test with calibrated Type~I error and high power from a small number of diagrams; (ii) interpretable effect sizes; and (iii) a 1-Wasserstein-stable feature vector for downstream machine learning. We validate calibration and power on synthetic manifolds with controlled topology, demonstrate competitive vectorisation across 14 graph and 3D point cloud benchmarks, and apply the method to study functional brain connectivity in fMRI/neuroscience data. To our knowledge, STRAND is the first method to provide hypothesis testing and vectorisation for persistence diagrams from a single coherent and interpretable representation.


From Structural Feedback to Prompt Policies: Learning Faithful Text-to-Image Prompt Editors

Wentao Ye ⋅ Yali Ye ⋅ Zhiqing Xiao ⋅ Ru Peng ⋅ Xinpeng Ti ⋅ Liyao Li ⋅ Zhanming Shen ⋅ Peng Lu ⋅ Jie Zhang ⋅ Haobo Wang

Text-to-image prompt optimization is often treated as prompt rewriting: given a user prompt, a language model expands it with richer style, composition, and photographic details. While this paradigm can improve visual appeal, it may silently weaken the user's actual intent, especially when the prompt specifies fine-grained objects, attributes, and relations. We argue that faithful prompt optimization requires a different objective: improving perceptual quality while preserving the compositional structure that makes the prompt semantically correct. We propose \ours{}, a single-step reinforcement learning framework that transforms noisy scene-graph feedback into a learnable prompt-editing policy. Instead of freely rewriting the entire prompt, \ours{} first extracts a calibrated scene graph from the input and then performs conservative deletion, reordering, and insertion actions conditioned on both text and graph representations. This design exposes structured intent to the optimizer while restricting unnecessary prompt drift. To make object--relation feedback usable for policy learning, \ours{} replaces sparse thresholded grounding rewards with uncertainty-aware smooth rewards and estimates advantages with environment-aware grouping over multiple image samples from the stochastic text-to-image generator. Across prompt-optimization datasets and generation backbones, \ours{} improves the faithfulness--aesthetics trade-off, with particularly consistent gains in relational fidelity where generic rewriting and aesthetics-oriented optimization often fail. We further evaluate with both human preference and Gemini-VQA, showing that the improvements are not confined to the SG reward used during training. Appendix additionally compares against a high-budget iterative GPT-4o structural-feedback baseline to contextualize the efficiency of amortized policy learning. More broadly, our results suggest that structural constraints should not remain only post-hoc evaluators of generated images; when properly calibrated and stabilized, they can be internalized as training signals for reliable, intent-preserving prompt policies.


GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models

Xiaohang Tang ⋅ Keyue Jiang ⋅ Che Liu ⋅ Zhao ⋅ Xiaoxiao Xu ⋅ Sangwoong Yoon ⋅ Ilija Bogunovic

Reinforcement learning (RL) can be used to improve the policy (denoiser) of diffusion large language models (dLLMs), while being hindered by the intractability of the policy likelihood. A dominant and efficient family of methods replaces the likelihood in standard RL with its evidence lower bound (ELBO), estimated from randomly masked sequences. Despite being well-aligned with pre-training, these approaches introduce bias from training-inference mismatch by applying the likelihood surrogate ELBO, which can potentially cause training collapse. In this work, we propose **Guided Denoiser Self-Distillation (GDSD)** to directly distill the denoiser of dLLMs from an advantage-guided self-teacher, derived from the closed-form optimum of reverse-KL regularized RL. GDSD matches the dLLM's denoiser logits to the teacher's via a normalization-free objective, which reduces RL to likelihood-free self-distillation and thus bypasses the TIM biases. Recent ELBO-based methods emerge as instances of applying different distillation divergences, but with diagnosable pathologies that GDSD avoids. On planning, math, and coding benchmarks with LLaDA-8B and Dream-7B, GDSD consistently outperforms prior state-of-the-art ELBO-based methods with a more stable training reward dynamics, achieving test-accuracy improvements of up to $+19.6\\%$. These results suggest that direct denoiser self-distillation, without relying on ELBO likelihood surrogate, can provide a more stable and effective RL procedure for dLLMs.


GEM: A Dual-Scale Architecture for Graph-Level Hierarchical Representation Learning

Zhaowei Wu ⋅ Kaizhong Zheng ⋅ Guangmingzi Yang ⋅ Shuai Jiang ⋅ Liangjun Chen ⋅ Badong Chen

Real-world networks typically exhibit small-world properties: dense local clustering yet efficient global propagation. However, current methods struggle to reconcile this hierarchy: Graph Neural Networks (GNNs) suffer from over-smoothing in long-range modeling, while sequence/state-space models often lack the inductive bias to preserve complex graph topology. To bridge this gap, we propose Graph Encoding Mamba (GEM), a dual-scale architecture for graph-level hierarchical representation learning. First, GEM employs a dual-branch encoder that synergizes the local sensitivity of GNNs with the selective scanning of Mamba, effectively capturing both local topology and global shortcuts. Second, a modularity-aware hierarchical pooling mechanism is designed to retain salient community structures during graph coarsening. Crucially, to mitigate structural redundancy from multi-scale fusion, we incorporate an Information Bottleneck objective to distill task-relevant structural patterns from redundant features, encouraging compact, task-relevant representations and supporting intrinsic interpretability. Extensive experiments on graph-level benchmarks demonstrate that GEM improves performance across biological, social, and brain network datasets. We further provide a lightweight node-level adaptation on long-range benchmarks as a boundary analysis of GEM's transferability. Code is available at \url{https://anonymous.4open.science/r/GEM-C34D}.


Gender Artifacts from Art History to Text-to-Image Generation

Piera Riccio ⋅ Miriam Doh ⋅ Benedikt Höltgen ⋅ Noa Garcia ⋅ Nanne van Noord

Artistic styles are rooted in specific socio-historical contexts that encode social hierarchies, including distinct constructions of gender. Yet in AI research, style has long been treated as a surface-level visual property: a filter of color, brushstroke, and texture applied to otherwise content-neutral scenes. We introduce the first dataset to investigate the interplay between gender representation and style in both historical and generated images. StyleGender comprises 74k images spanning 19 artistic styles, comprising art historical images with style and gender annotations, T2I-generated images under controlled style and gender prompts, and a semantically aligned set enabling direct art history-to-generation comparison. By proposing two Set Gender Artifact (SGA) metrics (PixelSGA and MaskSGA), capturing gender signals at the pixel level and in compositional structure, we show that (1) gender representation shapes visual features across artistic styles, (2) style keywords carry these patterns into T2I generation, and (3) generative models tend to amplify gender artifacts beyond what is observed in historical sources.

Explaining the generalization and training dynamics in neural networks remains a challenge, and various approaches have been developed to study different aspects of these phenomena. In this paper, we introduce the magnitude potential to this study -- a quantity based on the theory of metric magnitude -- that reflects how well an arbitrary point is represented by a given set. We find that this basic quantity can be applied to examine various features in neural generalization. The ratio between the magnitude potential with respect to a class and with respect to the entire data, computed at the logit layer, is informative of the representation of the point. In experiments, these ratios for individual training points are found to be correlated with the Feldman memorization scores. Magnitude potential ratios aggregated across points detect structural changes in the decision boundaries and provide a geometric indicator of grokking in modular arithmetic. Although the magnitude potential ratio and neural collapse are both closely associated with intra-class and inter-class geometric structure, the magnitude potential ratio remains informative even when neural collapse is explicitly suppressed.


Generalized Adaptive Boosting and the Geometry of Mistakes

Marco Bressan ⋅ Nataly Brukhim ⋅ Nicolò Cesa-Bianchi ⋅ Emmanuel Esposito ⋅ Yishay Mansour ⋅ Shay Moran ⋅ Maximilian Thiessen

Not all mistakes are alike. For example, in medical diagnosis, false positive and false negative mistakes have a different impact and should not be aggregated. By revisiting the classical theory of boosting through this lens, we introduce the notion of $Z$-learner, which returns weak hypotheses achieving joint false-positive/false-negative guarantees $(FP,FN)$ from some set $Z \subseteq [0,1]^2$. This formulation generalizes the standard notion of $\gamma$-weak learning — recovered by taking $Z=$ { $(z_1,z_2): z_1+z_2<1/2-\gamma$ } — and allows for more flexible notions of weak learnability that account for the impact of the distribution on the achievable learning guarantees. This raises a natural question: which sets Z are boostable? We answer this by giving a complete geometric characterization that unifies, as special cases, both the classical boosting dichotomy and the recent multi-objective extension of Bressan et al. (COLT 2025). We also present an adaptive algorithm for boosting $Z$-learners without relying on the prior knowledge of~$Z$, generalizing AdaBoost's ability to handle unknown margin parameters. Finally, we demonstrate that $Z$-learners can achieve confidence amplification. However, the analysis is significantly more involved than in classical boosting and relies on the centerpoint theorem from discrete and computational geometry.


Generative Object Detection with Co-Training

Yang Liu ⋅ Yao Zhang ⋅ Feng Hou ⋅ Yangzhou Du ⋅ Peng Wang ⋅ Yang Zhang ⋅ zhongchao shi ⋅ Jianping Fan ⋅ Zhiqiang He

Recent progress in multimodal large language models (MLLMs) has strengthened visual understanding and instruction following, and opened a path to open-ended generative object detection, in which a model discovers, names, and localizes objects from free-form instructions without relying on a predefined category set. Nevertheless, current generative detectors remain less accurate than discriminative detectors, as token-level supervision provides only indirect and weak constraints on fine-grained visual perception. To equip MLLM with such an explicit optimization objective, we propose \textbf{GOD} (\textbf{G}enerative \textbf{O}bject \textbf{D}etection), a 3B-scale MLLM that injects object-level visual-prior supervision into generative detection by coupling visual-prior reconstruction with language generation. GOD attaches a lightweight DETR-style object-query branch to the MLLM visual encoder, maps the resulting object queries into the token space, and optimizes them jointly with autoregressive generation, MLLM-side object-query reconstruction, and DETR-side localization objectives. Specifically, in the training recipe, a multi-stage training pipeline first aligns the object decoder with the MLLM and then uses an interleaved-mask curriculum to transfer detector-grounded priors from proposal-assisted training to reference-free generation. Experiments on standard object detection and referring expression comprehension benchmarks show that GOD improves direct coordinate generation while preserving discriminative localization and language-conditioned grounding. These gains suggest that object-level visual priors are most effective when learned within the MLLM, rather than supplied only as external proposals, providing a practical bridge between generative language modeling and localization-aware perception. We will make our code publicly available upon acceptance.


GeoBiaset: A Counterfactual Benchmark for Demographic Bias in World-Level Geolocalization

Niccolò Niccoli ⋅ Silvia Dani ⋅ Marco Mistretta ⋅ Andrew Bagdanov ⋅ Marco Bertini ⋅ Lorenzo Seidenari

World-level image geolocalization is an increasingly important task, yet the biases of current models remain poorly understood. Prior work has identified coarse dataset imbalances noting, for instance, that some cities are absent from training sets and that models tend to perform better in wealthier regions. However no one has yet examined which visual cues drive these disparities. To do so, we introduce GeoBiaset, the first benchmark designed to measure how geolocalization models are influenced by the apparent race of a person inserted into the image foreground. The dataset consists of 9,541 images, pairing existing geotagged background images with counterfactual edits spanning seven races, holding the ground-truth coordinates fixed so that any shift in model predictions is attributable solely to the inserted subject. Alongside the benchmark, we introduce the Person-Origin Perturbation Indicators (POPI), novel metrics that quantify model resilience to these demographic perturbations by measuring whether predictions are pulled toward the region associated with the inserted group. Our evaluation of state-of-the-art geolocalization and multimodal large language models reveals a 16-47\% drop in continent-level performance when a subject from a demographically mismatched group is inserted. A landmark versus no-landamark analysis and gradient-weighted heatmaps show that models preferentially attend to geographic landmarks when present, but shift attention to facial features in their absence.


GeoCurv-TTT: Geometry-Aware Deformation Restoration for 3D Test-Time Training

Sina Ghofrani Majelan ⋅ M. O Ahmad ⋅ M.N.S. Swamy

Recent 3D point cloud classification networks suffer severe performance degradation when encountering real-world distribution shifts. To address this, we propose GeoCurv-TTT, a novel geometry-aware Test-Time Training (TTT) framework based on deformation restoration. Unlike existing methods that rely on geometry-agnostic masking or severe skeletal abstraction, our approach selectively deforms low-curvature planar regions while strictly preserving high-curvature structural anchors. This curvature-guided physical displacement generates targeted out-of-distribution (OOD) samples, forcing the network to learn deeply robust, topology-aware representations during pre-training. At inference, GeoCurv-TTT adapts to novel environments through two strategies: Standard TTT, which utilizes a self-supervised restoration objective to adapt to isolated instances, and Online TTT, which achieves real-time, backpropagation-free adaptation across streaming data by exclusively updating batch normalization statistics. Extensive experiments on the ModelNet40-C, ShapeNet-C, and ScanObjectNN-C benchmarks demonstrate that GeoCurv-TTT establishes new state-of-the-art (SOTA) performance for 3D domain adaptation, outperforming existing methodologies in accuracy, computational efficiency, and adaptation time.

3D instance segmentation requires representations that are both geometrically structured and semantically discriminative. Pure 3D encoders model spatial structure well but often lack strong appearance semantics, while directly lifting 2D priors into 3D by concatenation injects rich semantics without ensuring that they align well with the underlying 3D structure. In this paper, we address this problem by formulating 2D-to-3D semantic transfer as a \textbf{masked semantic reconstruction} task. Specifically, we mask a subset of background lifted 2D priors and train the 3D encoder to recover the missing semantic priors, providing explicit supervision for distilling rich 2D semantics into the 3D backbone. We further propose \textbf{geometry-modulated semantic injection}, which uses raw 3D geometry to modulate lifted 2D priors before sparse encoding, so that geometry explicitly controls semantic injection and alleviates the distribution gap between 2D semantic features and 3D geometric features. Experiments on two baseline frameworks show that the proposed method achieves better or comparable results in both full-data and few-shot settings on ScanNetV2 and ScanNet200.


Gradient-Guided Smoothing for LLM Safety Defense

Yuxuan Gu ⋅ Xiaocheng Feng ⋅ kun Zhu ⋅ Lunjun Liu ⋅ Zhihao Yao ⋅ Lei Huang ⋅ Weihong Zhong ⋅ Bing Qin

Large language models remain susceptible to jailbreak attacks, despite extensive safety alignment. Smoothing-based defense methods can leverage their intrinsic safety capabilities, but the efficacy relies on the brittleness of jailbreak attacks and degrades when jailbreak contexts exhibit larger semantic margins. We observe that successful attacks typically exploit low-probability regions of the input space, where LLMs' safe bounds are difficult to reliably generalize. Thus, we propose guiding perturbations toward higher-density areas to restore the effectiveness of LLMs' safety mechanisms. In detail, we present gradient-guided smoothing that combines random noise with Gauss-Southwell type iterative ascent to the log density of the context-aware input distribution. Experimental results across four jailbreak attacks and three instruction-following benchmarks demonstrate that our method effectively improves safety while maintaining the utility of LLMs.


GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling

Alireza Dadgarnia ⋅ Rush Tabesh ⋅ Mahdi Nikdan ⋅ Michael Helcig ⋅ Eldar Kurtić ⋅ Maximilian Kleinegger ⋅ Dan Alistarh

Quantization has become a standard tool for efficient LLM deployment, especially for local inference, where models are now routinely served at 2-3 bits per parameter. The state of the art is currently split into simple scalar quantization techniques, such as GPTQ or AWQ, which are widely deployed but plateau in accuracy at 3-4 bits per parameter (bpp), and "second-generation" vector- or trellis-quantized methods, such as QTIP, GPTVQ and AQLM, which push the accuracy frontier but are notoriously hard to implement and to scale. In this paper, we ask whether this gap is fundamental, or whether a carefully optimized $\textit{scalar}$ quantizer can recover most of it. We answer in the affirmative, by introducing GSQ (Gumbel-Softmax Quantization), a post-training scalar quantization method which jointly learns the per-coordinate grid assignments and the per-group scales using a Gumbel-Softmax relaxation of the discrete grid. GSQ matches the cardinality of the relaxation to the small number of levels available in the target bit-width regime (e.g., 3-8 levels for ternary and 3 bpp, respectively), making optimization tractable. Practically, on the standard Llama-3.1-8B/70B-Instruct models, GSQ closes most of the gap between scalar quantization and the QTIP frontier at 2 and 3 bits, while using a symmetric scalar grid with group-wise quantization, and thus remains compatible with existing scalar inference kernels. We further show that the same discrete-assignment optimization can be applied to practical GGUF K-Quant checkpoints: starting from publicly released GGUF models, GSQ improves accuracy while projecting the result back into the same deployment format. Finally, GSQ scales to trillion-scale Mixture-of-Experts models such as Kimi-K2.5, where vector-quantized methods are difficult to apply.


Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots

Mark Vero ⋅ Fabian Kaczmarczyck ⋅ Ivan Petrov ⋅ I Shumailov ⋅ Niels Heinen ⋅ Jamie Hayes ⋅ Tianqi Fan ⋅ Luca Invernizzi ⋅ Martin Vechev

LLMs increasingly serve as simulation engines for honeypots. They enable defenders to construct high-interaction honeypots with virtually no system security risks. However, LLM-powered honeypot development lacks a unified evaluation framework and protocol. Most evaluations consist of response similarity measurements against a fixed set of commands, manual testing, or online deployment. These evaluation methods are often not scalable for development, reproducible across evaluations, representative of practical attacks, and adaptable to various attacker and honeypot configurations. In this work, we bridge this gap and propose Honeyval, a comprehensive evaluation framework for LLM-powered HTTP honeypots. We address the limitations of prior evaluations by grounding the honeypots in 16 backend applications, using AI hacking agents as attackers, employing two control tasks to monitor agent and honeypot capabilities across customizations, and defining clear and verifiable exploit goals for the attacker. Using Honeyval, we conduct an extensive evaluation of recent cost-efficient LLMs as HTTP honeypots. Our experiments highlight the promise of LLM-powered honeypots; they lead to substantially longer interactions with the attacker than rule-based baseline honeypots and are far less frequently detected even by frontier models, all while preserving a running cost advantage against agentic attackers. Further, we experiment with different counter-offensive configurations for the honeypots, and observe unique trade-offs, such as longer interactions at the cost of increased detection.

Hypergraphs model higher-order interactions, but realistic hypergraph generation remains difficult because incidence, hyperedge-size heterogeneity, and overlap structure are not faithfully captured by pairwise reductions. We propose HEDGE, a generative model defined directly on relaxed incidence matrices via a structured stochastic diffusion. The forward process combines a hypergraph-specific two-sided heat operator with an Ornstein--Uhlenbeck component, preserving structure-aware noising near the data while yielding an explicit Gaussian terminal law. Conditional on an observed hypergraph, this forward process is linear-Gaussian, so conditional means, covariances, scores, and reverse-drift targets are available in closed form. We therefore learn a permutation-equivariant state-only reverse-drift field in incidence space by regressing onto exact conditional targets, and generate samples by simulating a learned reverse-time SDE from the Gaussian base law. We establish exactness in the ideal state-only setting together with finite-horizon stability guarantees, and empirically show improved hypergraph generation quality relative to strong baselines.


IDEAFix: Evaluation Framework for Creative Defixation Prompting in LLMs

Florian Carichon ⋅ Soumya Sharma ⋅ Meaghan J. Girard ⋅ Romain Rampa ⋅ Golnoosh Farnadi

Large language models (LLMs) are increasingly used for tasks involving creative problem solving and idea generation. However, However, there is a lack of consensus concerning their creative capabilities:: some studies report superior performances compared to humans, while others highlight structural limitations such as fixation and the homogenization of outputs. Existing evaluation approaches either rely on narrow, decontextualized tasks that do not capture goal-oriented generation or on broader settings that confound multiple aspects of the creative process, making it difficult to isolate the effects of task formulation, prompting, and evaluation design. Significantly, the role of structured prompting strategies in shaping idea generation remains underexplored. Therefore, we introduce IDEAFix, an evaluation framework for analyzing divergent thinking in open-ended idea generation tasks. We prompt models to generate multiple original solutions to controlled variations of short design scenarios, task attributes, and defixation prompting strategies. This design enables systematic analysis of how structured guidance influences LLMs' idea generation. Our results show that both task formulation and attribute selection significantly affect models' performance, and that simple prompting strategies can boost the originality of solutions. However, we also observe persistent output homogenization across models, confirming inherent limits in their ability to generate diverse solutions. Overall, IDEAFix provides a controlled, extensible framework for studying the mechanisms underlying LLMs' creativity.


I Have a Stream: Making Self-Supervised Learning Work on Continuous Video

Ivan Martinović ⋅ Lukas Knobel ⋅ Yuki Asano

Self-supervised learning draws inspiration from infant visual development, yet standard training pipelines bear little resemblance to it: images are independently sampled and globally shuffled across epochs. We study self-supervised learning from continuous video streams, where frames are consumed in temporal order using strict sliding-window batches, without global reshuffling or multi-epoch replay. To this end, we construct WT++, a 95-hour urban walking-tour video dataset for streaming pretraining. Combined with a comprehensive evaluation suite we find that contrastive and distillation-based methods struggle, while MAE is more robust but still falls short of standard i.i.d. pretraining. We find that high inter-batch similarity, caused by sliding-window consumption across consecutive batches, does not explain this gap alone. The main challenge is high intra-batch similarity, where frames within each batch are near-duplicates. To mitigate this, we propose StreamMAE, which preserves the core MAE reconstruction objective while adapting the input pipeline with stream-aware regularization and motion-biased crop selection. StreamMAE outperforms streaming baselines, matches both i.i.d. MAE on video and even its ImageNet equivalent, and scales positively as pretraining video grows from 12 to 95 hours.

We study the max-margin solutions reached by mirror flow in deep neural networks with homogeneous activation functions. Extend classical results on gradient flow, we derive a novel balance equation for mirror flow from convex duality, enabling a characterization of the horizon function governing the induced margin. We further establish max-margin characterizations together with convergence rates and norm growth estimates. Finally, we support our theory through experiments on synthetic datasets and standard vision tasks. Concretely, we show that: (1) distinct non-homogeneous mirror maps can induce the same max-margin solution; (2) convergence can be extremely slow, including exponentially slow regimes; and (3) although all considered mirror maps exhibit feature learning, they can produce markedly different representations, ranging from sparse to dense neuron activations. Together, these results provide a unified perspective on sparse and dense feature learning in homogeneous neural networks, highlighting how mirror maps shape both optimization dynamics and the geometry of the learned classifiers.


In-Context Black-Box Optimization with Unreliable Feedback

Nicolas Samuel Blumer ⋅ Julien Martinelli ⋅ Samuel Kaski

Black-box optimization in science and engineering often comes with side information: experts, simulators, pretrained predictors, or heuristics can suggest which candidates look promising. This information can accelerate search, but it can also be biased, input-dependent, or misleading. Feedback-aware BO methods typically handle one task at a time, limiting their ability to generalize over multiple sources of feedback. In-context optimizers address cross-task adaptation, but usually assume that optimization history is the only available signal at test time. We study feedback-informed in-context black-box optimization (FICBO), where a pretrained optimizer conditions on both the observed history and cheap auxiliary feedback for the current candidate set. We introduce a structured feedback prior that models how feedback sources vary in their access, relevance, and distortion relative to the true objective, and use it to pretrain a feedback-aware transformer. At test time, the model estimates source reliability in context by comparing observed objective values with auxiliary signals, improving query selection. On synthetic and real-world tasks, FICBO effectively exploits informative feedback while remaining robust to weak or misleading sources, improving over other baselines. Empirical investigations further illustrate how the model perceives test-time sources, offering insights into its interpretability and decision-making process.

Transformers acquire in-context learning abilities through abrupt phases during training, often unfolding over multiple stages in which key circuits, such as induction heads, emerge. In this work, we characterize the dynamics underlying the emergence of such circuits across these stages. We focus on a synthetic associative recall task, where sequences are drawn from random maps between a permutation group and a vocabulary, and the model is required to complete the mapping of a permutation by retrieving the corresponding association from context. For this task, we study the gradient-flow trajectories of a simplified two-layer attention-only transformer. Leveraging symmetries in both the transformer architecture and the data distribution, we derive closed-form parameter dynamics for both attention layers. We identify a conservation law that couples parameters across layers and governs their joint learning. This law reveals how initialization controls the timescales over which such circuits emerge. In the vanishing-initialization limit, we characterize the gradient-flow trajectory, showing how training jumps from saddle to saddle. Finally, we provide empirical evidence across different architectural choices, validating our simplifications and extending the insights from our analysis beyond the simplified setting.

We study Nash equilibrium learning in partially observable Markov games (POMGs), a multi-agent reinforcement learning framework in which agents cannot fully observe the underlying state. Prior work in this setting relies on centralization or information sharing, and suffers from sample and computational complexity that scales exponentially in the number of players. We focus on a subclass of POMGs with independent state transitions, where agents remain coupled through their rewards, and assume that the underlying fully observed Markov game is a Markov potential game. For this class, we present an independent learning algorithm in which players, observing only their own actions and observations and without communication, jointly converge to an approximate Nash equilibrium. Due to partial observability, optimal policies may in general depend on the full action-observation history. Under a filter stability assumption, we show that policies based on finite history windows provide sufficient approximation guarantees. This enables us to approximate the POMG by a surrogate Markov game that is near-potential, leading to quasi-polynomial sample and computational complexity for independent Nash equilibrium learning in the underlying POMG.


Instantiation of Human Values in Image Generation

Maria Teresa De Rosa Palmini ⋅ Piera Riccio ⋅ Noa Garcia ⋅ Eva Cetinic

Text-to-image (T2I) models are increasingly embedded in everyday visual culture, yet little is known about how they translate abstract human values into visual scenes. We address this gap by introducing a framework, grounded in Schwartz’s theory of basic human values, for studying value instantiation: the process by which ideals such as success, care, or freedom are visually realized. Using person-centered prompts from 70 positive and negative value items, we generate 28,000 images across four state-of-the-art T2I models. We identify visual prototypes, recurring combinations of subjects, actions, and settings, and use them to quantify how narrowly models instantiate these values and how this breadth compares with human descriptions of the same items. We then map these prototypes onto a taxonomy of social configurations spanning setting, sociality, mode of engagement, and social register, revealing what is systematically foregrounded or omitted. Our results show that T2I models produce far more concentrated representations than humans do when describing the same values. They default to a hyper-individualistic, adult-centric worldview, reproducing value-specific associations, such as achievement tied to formal institutional settings or hedonism rendered as affluent consumption, while marginalizing everyday labor, relational care, and non-adult perspectives. As a proof of concept, we show that these structural gaps can guide a steering method that diversifies outputs without modifying the original prompt.


Intersectional Fairness via Mixed-Integer Optimization

Jiří; Němeček ⋅ Mark Kozdoba ⋅ Illia Kryvoviaz ⋅ Tomáš Pevný ⋅ Jakub Marecek

The deployment of Artificial Intelligence in high-risk domains, such as finance and healthcare, necessitates models that are both fair and transparent. While regulatory frameworks, including the EU's AI Act, mandate bias mitigation, they are deliberately vague about the definition of bias. In line with existing research, we argue that true fairness requires addressing bias at the intersections of protected groups. We propose a unified framework that leverages Mixed-Integer Optimization (MIO) to train intersectionally fair and intrinsically interpretable classifiers. We prove the equivalence of two measures of intersectional fairness (MSD and SPSF) in detecting the most unfair subgroup and empirically demonstrate that our MIO-based algorithm improves performance in finding bias. We train high-performing, interpretable classifiers that bound intersectional bias below an acceptable threshold, offering a robust solution for regulated industries and beyond.


IRPO: Boosting Image Restoration via Post-training GRPO

Haoxuan Xu ⋅ Yi Liu ⋅ Tianfu Li ⋅ Ruolin Shen ⋅ Boyuan Jiang ⋅ Jinlong Peng ⋅ Donghao Luo ⋅ Xiaobin Hu ⋅ Shuicheng Yan ⋅ Haoang Li

Post-training has become effective for high-level generation, but its role in low-level vision remains underexplored. Existing image restoration methods often rely on fixed pixel-wise fitting to ground-truth images, which can lead to over-smoothing and weak generalization. We propose IRPO, a GRPO-based post-training framework for deterministic restoration models. IRPO is built around two axes: data formulation and reward modeling. For data formulation, we select the 30\% underperforming samples from the pre-training stage, which improves both accuracy and training efficiency. For reward modeling, we combine fidelity-oriented and quality-aware feedback with three components: a General Reward for structural fidelity, an Expert Reward that uses a Vision-Language Model as a coarse visual-quality judge, and a Restoration Reward for task-specific low-level cues. Experiments on six in-domain and five out-of-domain (OOD) benchmarks show that IRPO improves the AdaIR baseline by 0.93 dB on in-domain tasks and 3.43 dB on OOD settings. Our code will be released upon acceptance.


Joint Adaptive Neighborhood Constraint for Offline Multi-Agent Reinforcement Learning

Xiancheng Gao ⋅ Mingxiao Feng ⋅ Lin Liu ⋅ Yuanrui Duan ⋅ Yufeng Shi ⋅ Wengang Zhou ⋅ Houqiang Li

Offline reinforcement learning suffers from distributional shift and extrapolation errors. These issues are particularly severe in multi-agent settings. As the number of agents increases, the joint action space grows exponentially while agent behaviors become highly coupled. Consequently, even if individual actions remain within the data distribution, their combination may still result in out-of-distribution (OOD) joint actions. Directly extending single-agent constraints to multi-agent settings is difficult, as they fail to effectively constrain joint actions. This often results in poor coordination or excessive conservatism. To bridge this gap, we propose the joint adaptive neighborhood constraint (JANC), which explicitly constructs controlled neighborhoods in the joint action space to suppress extrapolation while preserving reliable generalization. Moreover, JANC adaptively scales neighborhood radii based on joint advantages to align generalization with value structures. In practice, our method first performs adaptive neighborhood Q learning to explore high-value behaviors and then conducts value and policy learning based on the refined actions. Experiments on offline benchmarks, including multi-agent Mujoco and StarCraft II, show that JANC outperforms state-of-the-art methods on the majority of tasks in complex cooperative scenarios. We will release our source code to the public.

We develop a pattern-algebraic framework for higher-order graph neural networks by extending Deep Homomorphism Networks (DHN) from rooted patterns to $k$-labelled patterns. The framework captures GNN expressivity along two independent axes --- skeleton complexity $k$, controlling global structures, and pattern vocabulary $\mathcal{P}$, capturing local structures. We then introduce \emph{auxiliary gluing} ($\circledast$), the pattern-algebraic counterpart of the $k$-FWL sift, and prove that this single additional operation lifts $k$-label expressive power to $(k{+}1)$ labels, yielding an algebraic proof of $(k{+}1)$-WL $\equiv$ $k$-FWL. Neuralising these operations gives $k$-DHN and $k$-FDHN, which recover $k$-GNN, PPGN, $\delta$-$k$-LWL, $r$-$\ell$WL, spectral invariant GNNs, and node-marking subgraph GNNs as special cases, with expressivity exactly characterised by the pattern closure of $(k, \mathcal{P})$. The framework also resolves open expressivity separation problems in the $\delta$-$k$-LWL and $r$-$\ell$WL hierarchies.


KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs

Chuangtao Chen ⋅ Grace Li Zhang ⋅ Xunzhao Yin ⋅ Cheng Zhuo ⋅ Bing Li ⋅ Ulf Schlichtmann

Large Language Models (LLMs) rely heavily on Key-Value (KV) caching to minimize inference latency. However, standard KV caches are context-dependent: reusing a cached document in a new context requires recomputing KV states to account for shifts in attention distribution. Existing solutions such as CacheBlend, EPIC, and SAM-KV mitigate this issue by selectively recomputing a subset of tokens; however, they still incur non-negligible computational overhead (FLOPs) and increased Time-to-First-Token (TTFT) latency. In this paper, we propose KV Packet, a recomputation-free cache reuse framework that treats cached documents as immutable ``packets'' wrapped in light-weight trainable soft-token adapters, which are trained via self-supervised distillation to bridge context discontinuities. Experiments on Llama-3.1 and Qwen2.5 demonstrate that the proposed KV Packet method achieves near-zero FLOPs and lower TTFT than recomputation-based baselines, while retaining F1 scores comparable to those of the full recomputation baseline. Code and experimental results are available at https://anonymous.4open.science/r/kvpacketanonymous-46FC/.


LayoutBridge: Anisotropic Brownian Bridges for Public Indoor Floorplan Generation

Tongyang Dai ⋅ Ruipeng Gao ⋅ Feng Liu ⋅ Wenqiu Bai ⋅ Lingkun Li ⋅ Qiang Ni

Compared with residential layouts, public indoor floorplans present more substantial challenges with open corridors, complex room relations, and non-Manhattan geometries, which do not conform to conventional residential-layout assumptions. Consequently, existing methods are prone to boundary discontinuities, leading to fragmented corridors and irregular wall structures. To address the above issues, we propose \textit{LayoutBridge}, a structure-conditioned generative framework for floorplan production. By moving beyond the Manhattan-geometry assumptions commonly adopted in residential scenarios, LayoutBridge reformulates the task as a latent-space Brownian-bridge diffusion process from structural constraints to target layouts. We have constructed our PISF dataset with 3,517 representative floorplans in public indoor spaces, and LayoutBridge reduces FID by 132.35 and improves BIoU by 15.52\% over the latest image-to-image baselines. Regarding the residential benchmarks MSD and RPLAN, it also reduces FID by at least 2.53 compared with state-of-the-art. These results demonstrate its potential for scalable automated architectural design. Code implementing the proposed method is publicly available at \href{http://github.com/lalalalaxxx/LayoutBridge}{http://github.com/lalalalaxxx/LayoutBridge}.

Deep networks are powerful function approximators, but they typically store many different computations in shared weight matrices, making it difficult to selectively reuse or adapt parts of them when a familiar structure appears in novel combinations. We introduce the Vector Network (VN), a hierarchical recurrent architecture in which each layer replaces a fixed weight matrix with a library of reusable rank-1 weight atoms. For each input, VN minimizes a layer-local energy to infer a sparse set of active weight atoms and their coefficients, jointly constrained by bottom-up input reconstruction and top-down feedback consistency. These weight atom coefficients then compose an input-specific low-rank weight matrix for that sample. After convergence, slow learning updates only the selected weight atoms through local residual signals scaled by the inferred coefficients. We evaluate VN on four compositional benchmarks spanning 1D signals, 2D spatial decoding, N-body dynamics, and compositional MNIST. VN matches strong baselines in distribution while often achieving out-of-distribution error about an order of magnitude lower when familiar factors must be recombined in novel ways. Vector networks thus make compositional generalization a structural property of the architecture and inference process rather than a brittle byproduct of fitting many behaviors into one shared dense parameter substrate.

Transfer learning in reinforcement learning (RL) has shown strong empirical success. In this work, we take a more principled perspective by studying when and how transferring knowledge between MDPs (a source and a target) can be provably beneficial. Specifically, we consider the case where there exists an undo map such that applying this map to the target’s state space recovers the source exactly. We propose an algorithm that learns this map by matching state feature statistics gathered from both MDPs, and then uses it to transfer the source policy. We theoretically justify the algorithm by analyzing the setting when the undo map is linear and the source is linearly-$Q^\star$ realizable, where our approach has strictly better sample complexity than tabula rasa RL in the target MDP. Empirically, we demonstrate that these benefits extend beyond this regime: on challenging continuous control tasks and Atari games, our method achieves significantly better sample efficiency. Overall, our results highlight how shared structure between tasks can be leveraged for efficient transfer of policies across environments.


Learning Where to Simulate: Generative Active Sampling for Online PDE Surrogate Training

Pierre Cesar ⋅ Sofya Dymchenko ⋅ Abhishek Purandare ⋅ Bruno Raffin

Data-driven PDE surrogates are trained with data produced by numerical PDE solvers. However, when the surrogate's goal is to generalize across a wide range of PDE configurations (e.g., initial conditions and physical coefficients), generating a representative training set is non-trivial. Uniform sampling of configuration parameters often under-represents trajectories exhibiting challenging dynamics, leading to high prediction errors and large error variance in the trained surrogate. Online training, where data generation and surrogate training are coupled, offers a natural advantage by allowing solver parameters to be steered on-the-fly. To efficiently exploit this capability, we introduce Online Generative Active Sampling (OGAS), an active learning method that reactively learns the relationship between configuration parameters and surrogate performance to control the sampling distribution. OGAS trains a fast diffusion model in parallel to the surrogate to act as a conditional sampler, mapping a surrogate-derived difficulty signal (e.g., loss or uncertainty) to configuration parameters. By actively drawing target signals from a prior biased toward high difficulty, OGAS continuously steers data generation toward challenging regimes without delaying the training workflow. We evaluate OGAS across 2D PDEs with distinct challenging dynamics (Kuramoto-Sivashinsky, Navier-Stokes, Gray-Scott) and up to 308 parameters, using multiple surrogate architectures. Across all settings, OGAS consistently improves tail statistics, yielding substantial reductions in errors above the 99th percentile and overall error dispersion compared to uniform sampling. While prioritizing challenging trajectories introduces a trade-off with average error, OGAS effectively ensures worst-case reliability of trained surrogates with negligible wall-time overhead.


LithoBench: Benchmarking Large Multimodal Models for Remote-Sensing Lithology Interpretation

jun wang ⋅ Fengpeng Li ⋅ Tianjin Huang ⋅ Hang Dong ⋅ Wei Han

Remote sensing lithology interpretation is fundamental to geological surveys, mineral exploration, and regional geological mapping. Unlike general land-cover recognition, lithology interpretation is a knowledge-intensive task that requires experts to infer rock types from various features, e.g., subtle visual, making reliable automated interpretation highly challenging. Geological knowledge-guided large multimodal models offer new opportunities, yet their evaluation remains constrained by the lack of benchmarks that capture lithological annotations, multi-level geological semantics, and expert-informed assessment. Here, we propose LithoBench, a multi-level benchmark for evaluating geological semantic understanding in remote sensing lithology interpretation. LithoBench contains 10,000 expert-annotated interpretation instances across 12 representative lithological categories, including 4,000 multiple-choice and 6,000 open-ended tasks organized into five cognitive levels: Identification and Description, Comparative Analysis, Mechanism Explanation, Practical Application, and Comprehensive Reasoning. We further develop an expert-in-the-loop, knowledge-grounded semi-automated construction pipeline, coupling multi sub-processes, e.g., structured geological image descriptions, to enhance geological validity and evaluation reliability. Experiments with multiple large vision-language models reveal substantial limitations in geological semantic understanding, particularly on higher-order explanation, application, and reasoning tasks. Taken together, LithoBench establishes a dedicated testbed for advancing geological knowledge-guided multimodal models, supporting systematic model evaluation, domain-specific training and adaptation, and knowledge-enhanced reasoning toward expert-level lithology interpretation from remote sensing imagery.


LLM-AutoSciLab: Closed-Loop Scientific Law Discovery via Active Experimentation with LLMs

Sanchit Kabra ⋅ Nikhil Abhyankar ⋅ Saaketh Desai ⋅ Prasad Iyer ⋅ Chandan Reddy

Scientific discovery is a closed-loop process in which hypotheses guide data acquisition, and observations refine the hypothesis space. Yet most approaches reduce discovery to supervised learning over fixed datasets, where limited observations can support multiple plausible mechanisms that fit locally but fail to generalize. Thus, the key challenge is selecting informative observations to resolve uncertainty, shifting the focus from static inference to adaptive data acquisition. To address this, we propose LLM-AutoSciLab, a closed-loop framework that couples hypothesis generation with hypothesis-conditioned experiment selection and mechanism refinement. Rather than fitting models to passively collected data, LLM-AutoSciLab iteratively proposes plausible hypotheses, selects informative experiments to distinguish among them or refine them, and updates its state based on the resulting evidence. To evaluate dynamic, closed-loop scientific discovery with active data acquisition, we introduce ActiveSciBench-Chem (57 enzyme-kinetics domains) and ActiveSciBench-GRN (45 gene-regulatory-network tasks), benchmarks that model discovery as a budget-constrained process requiring adaptive experiment design, variable selection, and recovery of true mechanisms. Across NewtonBench, ActiveSciBench-Chem, and ActiveSciBench-GRN, LLM-AutoSciLab outperforms prior methods, achieving 67.6% and 35.1% symbolic accuracy and 31.1% exact graph recovery, respectively. Moreover, hypothesis-guided experimentation is 2x--5x more sample-efficient than the strongest competing baselines.


Lost in the Slots: Revisiting Object-Centric Representations in the era of Foundation Models

Priyam Dey ⋅ Aditya Sahdev ⋅ Omkar M Kashyap ⋅ Venkatesh Babu Radhakrishnan

Object-centric representation learning (OCL) has been widely proposed as a principled solution to the binding problem in deep learning, promising improved compositionality, reasoning, and robustness. However, its effectiveness in modern foundation model settings remains unclear. In this work, we systematically evaluate whether explicit object-centric (OC) representations meaningfully improve binding in pretrained foundation models. Departing from prior studies, we assess OC representations in a realistic regime by pairing them with pretrained foundation LLMs and analyze their performance on diverse, open-ended VQA and grounding benchmarks. Our findings challenge the prevailing narrative: slot-based OC representations consistently underperform standard dense patch-based features across multiple benchmarks, including compositional reasoning and hallucination-sensitive tasks. To understand this behavior, we analyze the slot representations through the lens of core binding components, namely scene segregation and object representation,and uncover three key limitations: (i) \textit{inferior scene segregation}, where slot assignments fail to cleanly disentangle objects compared to the implicit grouping already present in pretrained visual backbones; (ii) \textit{inherent information loss} during slot encoding, which degrades downstream performance of VLMs despite adapting them to the resulting feature space of OC; and (iii) \textit{weak attribute encoding}, where slots struggle to preserve fine-grained properties such as category, color, and spatial position. Further analysis reveals that these issues are not solely attributable to slot-based learning: while dense patch-based representations exhibit stronger binding, they too are imperfect. This suggests that explicit object-centric modeling, as currently instantiated, introduces bottlenecks that discard useful information without sufficiently improving structural reasoning. Together, our results expose a fundamental disconnect between the theoretical appeal of OCL and its practical utility in the era of foundation models. They point toward a need for rethinking binding, not as explicit object-factorization alone, but as a balance between structure and information preservation, potentially through new mechanisms that retain the richness of dense representations while enabling more reliable compositional reasoning.

Many scientific simulations are computationally expensive, limiting their use in simulation-based inference, uncertainty quantification, and decision making. We introduce *macrocanonical generator networks*, a framework for learning fast, data-efficient neural surrogates for *amortized* simulation of stationary physical processes by matching multiscale statistics of a target process. Inspired by microcanonical maximum-entropy synthesis, which matches feature statistics per sample by inner-loop optimization, our approach instead trains a neural generator to satisfy the same constraints in expectation, preserving physically realistic sample-to-sample fluctuations. We prove that gradient-descent training induces preconditioned gradient-descent dynamics in sample space that inherit the symmetry-preservation properties of the microcanonical gradient descent of Bruna and Mallat, amortizing its per-sample inner-loop optimization into a single training run after which every sample is generated in one forward pass. We further show that residual generator parameterizations initialized near the identity, trained with a small learning rate, preserve entropy locally during early optimization and mitigate mode collapse. Experiments on cosmology and fluid simulations show that macrocanonical generators recover realistic samples from extremely small training sets, in some cases a single training example, while generating new realizations more than $10^6$ times faster per sample than the reference simulator and more than $10^3$ times faster than microcanonical gradient descent. The trained generators can be reused for downstream pretraining, including simulation-based inference, providing a data-efficient route to multi-fidelity scientific machine learning workflows.


Mamba Flow Matching Neural Processes: Linear-Time Inference for Irregularly Observed Spatial Fields

Cosmo Santoni ⋅ Giovanni Charles ⋅ Timothy James Hitge ⋅ Oliver Watson ⋅ Elizaveta Semenova

Spatial interpolation from sparse irregular observations is a recurring task across the environmental and health sciences. Neural processes amortise this inference into a single forward pass at test time. Joint-predictive variants capture correlations across query locations but typically rely on self-attention over context and query points, incurring $O(K^2)$ compute in the combined point count $K$ and capping practical use at a few thousand points. State-space models offer a linear-time alternative, but existing state-space NPs compress the context into a single hidden state before reading queries, reducing the information available to the decoder. We propose MambaFlowNP, a joint-predictive NP with linear compute trained with a conditional flow-matching objective. A multi-directional Mamba-2 scan reads context and query tokens in a single shared sequence, so each query's velocity is read from its own scan position rather than from a compressed context summary. MambaFlowNP is competitive with the strongest attention-based joint-predictive baselines on synthetic Gaussian-process benchmarks, and across regional and continental-scale NOAA temperature interpolation tasks it attains the lowest RMSE on every tier and the lowest CRPS at the California and continental US tiers, while running $170$--$760{\times}$ faster than the strongest joint-predictive baseline. A single forward pass scales to $10^6$ points on one GPU in under a second. This scalability unlocks amortised joint-predictive inference for continental-scale, irregularly sampled sensor networks, spanning climate monitoring to disease vector surveillance.


Many Circuits, One Mechanism: Input Variation and Evaluation Granularity in Circuit Discovery

Alireza Bayat Makou ⋅ Jingcheng Niu ⋅ Subhabrata Dutta ⋅ Iryna Gurevych

Circuit discovery methods identify subgraphs that explain specific model behaviors, and structural differences between discovered circuits are commonly interpreted as evidence of distinct mechanisms. We test this assumption by drawing input tokens from bands defined by their frequency in the pretraining data while holding the task fixed. The discovered circuits appear specialized by frequency when compared structurally, but functional and representational analyses show no reliable evidence of corresponding differences in their computations. We term this mismatch phantom specialization. Using the Literal Sequence Copying task across four frequency bands plus a control sampled according to token frequency, we extract 75 circuits from five Pythia models (70M-1.4B). We find that structurally distinct circuits implement the same computation: band-specific edges transfer broadly across bands, a core shared across most bands recovers at least 99% of circuit performance in models above 70M, and causal interchange interventions confirm that internal representations are interchangeable across frequency bands. A smaller replication on a subject-verb agreement task shows the same pattern: circuits differ structurally, transfer broadly across bands, and the core shared by most bands recovers nearly all of the circuit's accuracy in all five models. Repeated extractions within the same frequency band further suggest that discovery algorithms sample from an equivalence class of valid subgraphs rather than recovering a unique mechanism. Standard evaluation practice obscures this pattern: source-level evaluation inflates apparent faithfulness, while edge-level evaluation reveals the many-to-one mapping from structure to function. We find no reliable evidence that frequency bands are processed by different computations. In this work, we do not test specialization at the level of token positions or of features inside a component. Our results show that structural differences between circuits are not sufficient evidence for distinct mechanisms, and that exposing this requires edge-level evaluation and cross-condition transfer tests.


Membership Inference on Synthetic Single-Cell Genomic Data

Steven Golob ⋅ Patrick McKeever ⋅ Sikha Pentyala ⋅ Martine De Cock ⋅ Jonathan Peck

Single‑cell RNA sequencing (scRNA‑seq) data is subject to strict access control due to its sensitive nature, motivating the use of synthetic data generation (SDG) for privacy‑preserving data sharing. We present the first adversarial privacy attack that performs meaningfully above random guessing against state‑of‑the‑art scRNA‑seq SDG methods. Our attack enables donor‑level membership inference, demonstrating that leading SDG techniques fail to adequately mask which individuals were used to train the generator. We show that privacy leakage increases as the number of training donors decreases. Although the attack is designed to exploit vulnerabilities in scDesign2, we find that it also succeeds against synthetic data generated by other leading methods, including scDesign3 and scVI. This transferability indicates that an adversary can infer sensitive information from synthetic data without access to the training procedure, model parameters, or even the underlying generation algorithm. Finally, we investigate the use of perturbation with noise during the SDG process as a first‑line defense, empirically evaluating its effectiveness in neutralizing the attack and its impact on utility.


MEMEVO: A Memory-Evolved Video Agent for Long Video Understanding

Yuan Wang ⋅ Yake Wei ⋅ Tianrun Xu ⋅ Fengyun Rao ⋅ Jing LYU ⋅ Xuetao Feng ⋅ Yifan Sun

Agent-based frameworks grounded in vision-language models (VLMs) have emerged as a dominant paradigm for long video understanding. Yet, prevailing agents lack the core capacity for \textit{dynamic memory evolution}, failing to transform ephemeral perceptions into a continuously growing, adaptive knowledge system and thus inducing inefficient redundant exploration during reasoning. To this end, we present MEMEVO, an online Memory-Evolved Video Agent framework centered on \textit{dynamic memory evolution} that leverages accumulating multi-turn queries to actively drive progressive memory evolution, establishing a long-lived memory system for knowledge accumulation. Central to MEMEVO is a four-level Hierarchical Self-Evolved Memory, which constructs memory bottom-up via progressive abstraction, distilling transient perceptions into reusable, query-agnostic event nodes organized within a temporal-semantic graph, while realizing active memory evolution through event insertion, duplicate merging, and redundant pruning. To mitigate redundant retrieval in the evolved memory, we devise a State-Conditioned Graph Memory Retrieval mechanism that guides efficient navigation of the memory space, integrating State-Conditioned Graph Routing and Meta-Cognitive Stopping criteria to ensure sufficient evidence acquisition without re-exploration. Extensive benchmarking confirms that MEMEVO achieves state-of-the-art performance on reasoning-intensive tasks (\textbf{\textcolor[HTML]{00FF00}{+6.1\%}} VideoMMMU, \textbf{\textcolor[HTML]{00FF00}{+4.8\%}} LongVideoBench), while facilitating \textit{cross-query knowledge reuse} and \textit{cross-model generalization}.

LLM agents are flexible in language interaction, but using them as long-horizon planners is costly and prone to hallucinated or invalid actions. This paper targets a different role for LLMs: constrained bridges between language and symbolic planning. We propose MemPlan, a memory-conditioned PDDL planning framework for partially observable interactive text environments. MemPlan keeps the PDDL domain fixed and reconstructs the PDDL problem at each step from parsed observations, candidate hypotheses, and interaction memory. Positive graph memory assigns relevance-based action costs to candidate hypotheses for missing objects, locations, or conditions, guiding the planner toward more plausible checks. Negative memory records execution-refuted assumptions and compiles them into exclusion constraints, preventing repeated attempts at candidates already contradicted by feedback. A unified schema-conditioned language--symbol interface, trained offline with supervised fine-tuning, connects text environments to the planning loop through observation-to-fact parsing and symbolic-action-to-command grounding. Experiments on TextWorld, ALFWorld, and Robotouille show that MemPlan improves task completion and step efficiency over language-only and PDDL-aware baselines, while substantially reducing online token cost compared with LLM-planning baselines such as ReAct. The same fine-tuned interface is also reused across TextWorld task types without task-specific fine-tuning. Our code and data are available at: \url{https://anonymous.4open.science/r/MemPlan-1EE5/}.


Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents

Ziyan Liu ⋅ Zhezheng Hao ⋅ Yeqiu Chen ⋅ Hong Wang ⋅ JingRen Hou ⋅ Ruiyi Ding ⋅ Yongkang Yang ⋅ Wence Ji ⋅ Wei Xia ⋅ Feng Liu

Memory-augmented LLM agents tackle complex long-horizon tasks by recursively summarizing interaction trajectories into compact memory. However, existing approaches typically train these memory policies using outcome-based reinforcement learning, failing to localize where intermediate memory quality degrades. As interactions unfold, ambiguous recursive summaries progressively discard or blur task-relevant information. This exacerbates belief deviation, obscuring the agent’s estimate of the latent task state and ultimately derailing long-horizon reasoning. We therefore argue that memory optimization should focus not merely on trajectory-level success, but on the clarity of the belief induced by intermediate summaries. To this end, we introduce Belief Entropy, a self-supervised proxy that probes how uncertain the model remains about the latent task state given its current memory. Based on this proxy, we propose Metacognitive Memory Policy Optimization (MMPO). Instead of relying only on sparse outcome-based signals, MMPO provides fine-grained, memory-specific supervision by explicitly penalizing summaries that induce high epistemic uncertainty. Experiments show that MMPO consistently outperforms existing methods on diverse long-horizon tasks, maintaining 97.1% performance even when scaled to 1.75M-token contexts.


MID: Mask-Image Distributional Divergence for Evaluating Medical Image Segmentation

Vincenzo Marcianò ⋅ XIAOMING ZHANG ⋅ Gianluca Guglielmo ⋅ Sebastien Ourselin ⋅ Michela Antonelli ⋅ Maria A Zuluaga

Current evaluation protocols for medical image segmentation remain largely anchored in sample-wise comparisons, e.g., Dice and Hausdorff distance, which require each prediction to be paired with a ground-truth mask. Yet two limitations arise in practice: they are inapplicable when predictions are made on unlabeled data; and even when annotations are available, they cannot assess whether a collection of predictions matches the distributional structure of the ground-truth population. We introduce MID (Mask–Image Distributional Divergence), a distribution-level evaluation metric for medical image segmentation. MID builds upon the observation that segmentation quality is a property of the joint image-mask distribution: a mask may be anatomically plausible in isolation yet incorrect for the image it accompanies. MID compares the joint distribution of images and masks between a reference set and a set of predictions, without requiring sample-wise correspondence. This captures both mask realism and image-mask consistency. We evaluate MID across 13 publicly available datasets spanning CT, MRI, and MRA. Under anatomically realistic corruptions across 56 organ groups, MID tracks mean Dice and Hausdorff distance without per-sample ground-truth pairing. MID also detects distributional shifts across modalities and anatomies, identifies image-mask misalignment through controlled shuffling experiments, and produces reliable model rankings across real segmentation models.


Mildly Overparameterized ReLU Networks on Orthogonal Data: Incremental Learning and Implicit Bias

James Town ⋅ Etienne Boursier ⋅ Ben Lewis ⋅ Matthias Englert ⋅ Ranko Lazic

The successful training of neural networks hinges on the use of first order optimization methods, yet the theoretical characterization of these methods remains incomplete. This is especially true in settings with mild overparameterization. In this work, we study the gradient flow dynamics of two-layer ReLU networks from small initialization with orthogonal training data. We prove the limiting flow converges to a saddle-to-saddle jump process as the initialization scale tends to zero, revealing an incremental learning phenomenon in which a new neuron activates at each saddle. This analysis recovers the known result of Dana et al. (2025) that the network interpolates the training data with high probability as soon as $m \gtrsim \log(n)$, where $m$ is the network width and $n$ is the number of training samples. This incremental process characterization also allows us to derive a novel implicit bias result: the learned interpolator has an $\ell_2$-norm scaling as $\sqrt{n}$, which is within a constant factor of the minimal-norm interpolator. More broadly, our work provides the first rigorous proof of an incremental learning process for ReLU networks, whilst suggesting mildly overparameterized networks can converge to interpolating solutions whose complexity is of the same order as that of the optimal interpolator.

We provide theoretical guarantees for convergence of discrete-time policy mirror descent with inexact advantage functions updated using temporal difference (TD) learning for entropy regularised MDPs in Polish state and action spaces. We rigorously derive sufficient conditions under which the single-loop actor-critic scheme is stable and convergent. To weaken these conditions, we introduce a variant that performs multiple TD steps per policy update and derive an explicit lower bound on the number of TD steps required to ensure stability. Finally, we establish sub-linear convergence when the number of TD steps grows logarithmically with the number of policy updates, and linear convergence when it grows linearly under a concentrability assumption.


Mitigating Saliency Collapse: Robust Saliency-Aware Long-Text Image-Text Alignment

Qiuyu Kong ⋅ Zanxi Ruan ⋅ Marco Cristani ⋅ Yiming Wang

Long-text image–text alignment is essential for fine-grained scene understanding, yet existing methods mainly focus on extending context length or improving local alignment. We identify an overlooked failure mode, saliency collapse, where models over-rely on text that features visual prominence while underutilizing non-salient context. Consequently, retrieval performance degrades sharply when salient content is reduced, despite that the non-salient context remains informative. To address this, we propose RoSA, a saliency-aware finetuning framework that decomposes image-text pairs into complementary salient and contextual views, enabling joint learning of prominent and contextual semantics. We propose a modality-independent decomposition strategy to combine object-level visual saliency estimation with LLM-based text decomposition to bridge vision and language. We employ auxiliary objectives to explicitly align these views, ensuring robust representation with both salient and contextual features. RoSA greatly improves long-text retrieval and demonstrate strong robustness to saliency collapse across benchmarks. Moreover, our auxiliary objectives are plug-and-play, bringing consistent gains to existing tuning methods. Code and checkpoints will be released upon acceptance.


More Value per Key: Asymmetric Sparse Attention for Faster LLM Decoding

Noam Elata ⋅ Itay Lamprecht ⋅ Mikey Shechter ⋅ Daniel Ohayon ⋅ Itay Hubara ⋅ Daniel Soudry

Autoregressive generation in Large Language Models (LLMs) is constrained by the memory and computational demands of attention mechanisms. Sparse attention methods mitigate this cost by selecting only high-probability entries of the attention matrix. We observe that in many such methods, this renders the probability-value multiplication negligible, shifting the bottleneck to the query-key step. Key heads can therefore be reduced to accelerate inference, while value heads preserve capacity at no additional cost. We introduce Sparse Asymmetric Group-Query Attention (SAGA), which decouples key and value head counts to exploit this principle, and pair it with approximate top-N (Atop-N) attention, a simple GPU-agnostic sparse attention method designed to isolate sparsity's effect on decoding-optimized architectures. We formalize the benefits of this asymmetry theoretically and validate it empirically, achieving end-to-end decoding speedups exceeding $2\times$ at long contexts across multiple model scales and benchmarks. Models trained from scratch with SAGA nearly match the quality of comparable GQA variants while delivering substantial efficiency gains. To facilitate adoption, we introduce an efficient fine-tuning method that converts pretrained models to the SAGA architecture at a small to negligible cost to quality, enabling practitioners to benefit from our approach without costly retraining.

3D motion and expression capture of laboratory mice is a critical task in behavioral neuroscience, yet it remains underexplored in the computer vision community. Existing studies have focused narrowly on either body pose estimation of freely moving mice or facial expression capture under head-fixed conditions, leaving a significant gap: no dataset supports simultaneous capture of whole-body motion and expression in a freely moving mouse. To fill this gap, we present PanoMouse, a multiview system and dataset for Mouse Total Capture (MTC). The PanoMouse system deploys 24 cameras across three horizontal layers, achieving 360-degree photographic coverage of the mouse. Built upon this system, the PanoMouse dataset collects over 3.4 million frames, with 7,293 frames annotated with 92 whole-body keypoints, totaling approximately 670,000 keypoint annotations. To our knowledge, PanoMouse is the first dataset to provide fine-grained annotations of a mouse in natural behavior. We further propose 4D Triangulation (Triang4D), a test-time optimization method for 3D MTC that jointly enforces multi-view consistency, temporal smoothness, and bone length constraints. Applying Triang4D to the entire dataset produces 3D whole-body keypoint sequences across all recording sessions, enabling downstream skeleton-based behavior analysis. We benchmark baseline methods on PanoMouse across three tasks: 2D pose estimation, 3D pose estimation, and behavior classification, opening a new paradigm for fine-grained behavior analysis and bridging computer vision with behavioral neuroscience.


Multi-Bridge Denoising Diffusion Probabilistic Models

Mélodie Monod ⋅ Habib Ganjgahi

Image-to-image translation aims to generate a target image conditioned on an observed source. While diffusion bridge models perform well in one-to-one settings, they are not adapted to problems where multiple sources provide complementary information about a single target. We introduce Multi-Bridge Denoising Diffusion Probabilistic Models (MB-DDPMs), a framework for many-to-one translation that builds stochastic bridges between the target and each source, coupled through a unified reverse process that aggregates information across sources. We evaluate MB-DDPM on brain MRI for Multiple Sclerosis to predict FLAIR contrast from T1, T2 and PD contrasts. We demonstrate strong clinical relevance through near-perfect agreement with ground-truth lesion-based measurements, highlighting its potential for reliable downstream medical analysis. We further evaluate MB-DDPM on image restoration using CIFAR-10 and ImageNet. Across all settings, MB-DDPM achieves state-of-the-art performance, consistently outperforming existing baselines in both reconstruction quality and perceptual metrics.


Multi-Environment POMDPs with Finite-Horizon Objectives

Léonard Brice ⋅ Filip Cano ⋅ Krishnendu Chatterjee ⋅ Thomas Henzinger ⋅ Stefanie Muroya

Partially Observable Markov Decision Processes (POMDPs) are systems in which one agent interacts with a stochastic environment, and receives only partial information about the current state. In a multi-environment POMDP (MEPOMDP), the initial state is unknown, and assumed to be adversarially chosen. In this work we focus on computing the optimal value and policy in MEPOMDPs with finite-horizon objectives. That problem is known to be PSPACE-complete in POMDPs. Our main results are as follows: (1) we establish that it is also PSPACE-complete in the more general setting of MEPOMDPs; (2) we present a practical algorithm and evaluate it on classical benchmarks, significantly outperforming the only previously known algorithm.

Model-based reinforcement learning (MBRL) has achieved remarkable sample efficiency in single-agent domains, yet its extension to competitive imperfect information games (IIGs) remains underexplored. In multi-agent settings, opponent-induced non-stationarity complicates the learning process, and decentralized model learning faces severe identifiability barriers. To solve this, we propose NashDreamer, a principled MBRL framework for two-player zero-sum IIGs. NashDreamer introduces a centralized Multi-Agent Recurrent State-Space Model (MARSSM) that decouples environment dynamics from the effect of players' strategies on their individual observations. For policy optimization, NashDreamer uses Regularized Nash Dynamics (RNaD) within the latent imagination, providing theoretical convergence guarantees toward Nash equilibria. Empirical evaluations across Imperfect Information Goofspiel, Leduc Hold'em and Battleship demonstrate that NashDreamer achieves vastly superior sample efficiency compared to model-free baselines, including a $71.0$\% head-to-head win rate against model-free RNaD in stochastic 13-card Imperfect Information Goofspiel. Finally, we theoretically analyze the architecture's optimization landscape, formally identifying the vulnerability of the Dreamer family of algorithms to posterior collapse in highly stochastic environments, and we highlight the open challenge of mitigating latent non-stationarity without enabling degenerate solutions.

One of the primary challenges in Bayesian inference on the parameters of a diffusion model from discrete observations is the unavailability of an analytical expression for the transition density function between consecutive observation times, which is needed to derive the likelihood function. Extending previous studies that solve Fokker-Planck (FP) type partial differential equations with Normalizing Flows, we propose a new Normalizing Flow architecture to learn the transition density function of the diffusion process between two observation times. We do so by solving in a Neural Galerkin framework the associated FP equation with a Dirac mass as initial condition, over a specified training distribution of the initial datum and the coefficients of the diffusion. We specifically focus on processes whose diffusion matrix vanishes in certain inaccessible boundary regions, such as Stochastic Volatility models that satisfy a Feller condition. The product of the obtained transition densities evaluated along the observed trajectory approximates the likelihood function, thereby enabling cheap posterior sampling via Markov chain Monte Carlo (MCMC). After the offline training phase, inference becomes significantly more efficient, as it avoids the need to solve the FP equation in real time for each parameter proposed by the MCMC sampler or to rely on other likelihood-free methods for Bayesian inference that involve repeated simulation of diffusion bridges.


Normalized Architectures are Natively 4-Bit

Maxim Fishman ⋅ Brian Chmiel ⋅ Ron Banner ⋅ Daniel Soudry ⋅ Boris Ginsburg

Training large language models at 4-bit precision is critical for efficiency. We show that nGPT, an architecture that constrains weights and hidden representations to the unit hypersphere, is inherently more robust to low-precision arithmetic. This removes the need for interventions—such as applying random Hadamard transforms and performing per-tensor scaling calculations—to preserve model quality, and it enables stable end-to-end NVFP4 training. We validate this approach on both a 1.2B dense model and hybrid (Mamba-Transformer) MoE models of up to 3B/30B parameters. We trace this robustness to the dot product: while quantization noise remains largely uncorrelated in both standard and normalized architectures, the signal behaves differently. In nGPT, the hypersphere constraint enhances weak positive correlations among the element-wise products, leading to a constructive accumulation of the signal across the hidden dimension while the noise continues to average out. This yields a higher effective signal-to-noise ratio and a flatter loss landscape, with the effect strengthening as the hidden dimension grows, suggesting increasing advantages at scale. A reference implementation is available at https://github.com/anonymous452026/ngpt-nvfp4.


Not All Gap Correction Helps: A Geometric View of External Information in LLM Inference

Dingzirui Wang ⋅ Xuanliang Zhang ⋅ Keyan Xu ⋅ Qingfu Zhu ⋅ Wanxiang Che ⋅ Yang Deng

Providing external information beyond the query is a common way to improve large language model (LLM) inference, as seen in in-context learning (ICL), retrieval-augmented generation (RAG), and memory-based methods (Mem). Existing work increasingly suggests that useful external information should compensate for the gap between LLMs and user queries, implicitly assuming that stronger gap-correction can lead to better performance. In this paper, we challenge this assumption and show that more gap-corrective external information can instead harm performance. We analyze this phenomenon in Transformers through the lens of reasoning error, defined as the difference between the predicted answer vector and the ground-truth answer vector. Our analysis reveals that external information can be not only under-corrective but also over-corrective, where excessively strong correction increases reasoning error. We further show that the induced correction vector is jointly determined by attention and the complementarity between the external information and the query, and derive conditions under which information is properly corrective in both direction and magnitude. Experiments on four mainstream LLMs and seven reasoning benchmarks across ICL, RAG, and Mem validate our theory and support a theory-guided information selection method.


ObjView-Bench: Rethinking Difficulty and Deployment for Object-Centric View Planning

Sicong Pan ⋅ Hao Hu ⋅ Xuying Huang ⋅ Benno Wingender ⋅ Maren Bennewitz

Object-centric view planning is a core component of active geometric 3D reconstruction in robotics, yet existing evaluations often conflate object complexity, planning difficulty, budget assumptions, and physical reachability constraints. As a result, conclusions drawn from idealized view-planning evaluations may not reliably predict performance under realistic reconstruction settings. We introduce ObjView-Bench, an evaluation framework for rethinking difficulty and deployment in object-centric view planning. First, we disentangle three quantities underlying view-planning evaluation: omnidirectional self-occlusion as an object-side attribute, observation saturation difficulty, and protocol-dependent planning difficulty defined through a set-cover formulation. This separation supports controlled dataset construction, analysis of slow-saturation objects, and a case study showing that planning difficulty-aware sampling can improve learned view planners. Second, we design deployment-oriented evaluation protocols that reveal how budget regimes and reachable-view constraints alter method behavior. Across classical, learned, and hybrid planners, ObjView-Bench shows that difficulty, budget, and reachability constraints substantially change method rankings and failure modes.


OneSearch-V2: The Latent Reasoning Enhanced Self-distillation Generative Search Framework

Ben Chen ⋅ Siyuan Wang ⋅ Yufei Ma ⋅ Zihan Liang ⋅ Han Li ⋅ Chenyi Lei ⋅ Wenwu Ou ⋅ Kun Gai

Generative Retrieval (GR) has emerged as a promising paradigm for modern search systems, offering end-to-end joint optimization and lower serving overhead compared to multi-stage cascaded architectures. OneSearch, as a representative industrially deployed generative framework, has delivered substantial commercial benefits. However, three key limitations constrain further improvement: shallow understanding of complex queries, insufficient personalized intent reasoning over user context, and a separately trained reward model that adapts slowly to emerging queries and is prone to sampling bias and reward hacking. To address these challenges, we propose OneSearch-V2 with three key innovations: (1) a thought-augmented query understanding module that generates compact keyword-based chains-of-thought (CoTs), overcoming the shallow-matching limitation of single-pass SID generation; (2) a reasoning-internalized self-distillation pipeline that encodes keyword-guided reasoning into the existing model weights via information-asymmetric supervision, without any extra parameters, special tokens, or inference-time CoT generation; (3) a behavior-feedback preference alignment system that replaces the separate reward model with composite rewards built from real user interactions, and introduces a token-position marginal advantage (TPMA) mechanism for position-aware credit assignment over hierarchical SID sequences. Extensive offline evaluations demonstrate OneSearch-V2's strong query understanding and personalized intent modeling capabilities. Online A/B tests further validate its business effectiveness, yielding +3.98\% item CTR, +1.17\% PV CTR, +2.90\% PV CVR, +2.07\% buyer volume, and +2.11\% order volume. Manual evaluation further confirms gains in search experience quality, with +1.37\% in page good rate and +1.65\% in query-item relevance. Importantly, OneSearch-V2 achieves these gains without any additional inference cost or serving latency, while also mitigating information bubbles and long-tail sparsity.


On Length Bias in EEG-to-Text Decoding

Iakovos Tenedios ⋅ Yashar Moshfeghi

Sentence-level EEG-to-text decoding models typically pad variable-duration neural signals to a fixed-length window, creating a structural confound where segment duration correlates with sentence identity independently of neural content, a problem we term length bias. Despite being acknowledged in prior work, the prevalence, magnitude, and architectural dependence of this effect have never been systematically quantified. We introduce length-matched noise baselines that isolate this confound by replacing only the true signal region with Gaussian noise while preserving zero-padding, and apply them across three English listening EEG datasets under a strict subject-held-out protocol. Evaluating ten sentence classifiers and three generation models based on the current state-of-the-art, we show that reported decoding scores are substantially inflated by length cues across datasets, with the degree of inflation varying markedly by architecture. Among classifiers, BIOT alone is unaffected by length-matched noise. A systematic decomposition of its components identifies Short-Fourier-Transform (STFT) tokenisation as the bigger source of this invariance. We transfer this finding into two new generation variants that largely eliminate encoder-side length bias, enabling a controlled analysis of where residual bias originates. A classification-to-generation evaluation further reveals that leading generation models operate largely as sentence retrieval systems, with a linear classifier recovering nearly all of their reported BLEU-4 scores. Together, these results establish that length bias is a pervasive and previously unquantified failure mode in EEG-to-text evaluation, and provide both diagnostic tools and architectural directions for future work.

We address the problem of evaluating machine learning models under a limited labelling budget in streaming environments, where data arrive sequentially. This setting is particularly critical for embedded and continuously deployed systems, such as autonomous driving or computer-aided diagnosis, where models must reliably assess their own performance online. We focus on a constrained online setting, where labelling decisions must be made immediately, and data cannot be stored. We introduce an \emph{online active testing} (OAT) framework based on an adaptive importance sampling estimator that provides unbiased risk estimation for both regression and classification tasks. While importance sampling has been extensively studied in pool-based settings, we establish novel theoretical ground for the fully online setting, including deviation bounds and consistency guarantees. We derive an optimal sampling strategy in order to minimise the estimator's variance, which is driven by the expected second moment of the loss. This strategy is expressed for common loss functions, including mean squared error, mean absolute error, 0–1 loss, and log-loss, and analysed from a statistical perspective. We empirically demonstrate the effectiveness of OAT on synthetic benchmarks, deep learning tasks, and a real-world application involving the evaluation of machine learning models for molecular simulations, where data are generated sequentially, and labels require expensive quantum chemistry computations, showing improved efficiency and reliable performance estimation under limited labelling budgets. In particular, on the real-world task, OAT requires only $\sim$30\% of labels to match the accuracy of a naïve uniform baseline, compared to $\sim$23\% for an oracle method that has access to the labels.


On Observation Time for Recovering Latent Hawkes Networks

Jonas Linkerhägner ⋅ Michele Bortolasi ⋅ Lorenzo Baldassari ⋅ Maarten V. de Hoop ⋅ Ivan Dokmanić

Dynamics of interacting systems in engineering, society, and nature often evolve over latent networks that govern which entities can interact. We study the problem of inferring these networks from event-based observations, which arise naturally in finance, seismology, and neuroscience. While there is substantial algorithmic work addressing this important problem, theoretical results are scarce. In this paper we ask the following fundamental question: what is the minimum time that one must observe the dynamics in order to exactly recover the underlying network, as a function of the number $d$ of interacting entities? For a class of stationary Hawkes processes with sparse, weak interactions, we prove that an observation time of order $\log d$ is sufficient and necessary. For the upper bound we construct a two-stage estimator that uses clipped and binned event data for screening, followed by a least-squares refinement, and apply concentration bounds derived from the Poisson cluster representation. For the lower bound we combine Fano’s inequality with Jacod’s Girsanov formula for point processes on a suitable subclass of networks.


On the Construction and Implications of Low-Loss Valleys in LoRA-based Bayesian Inference

Daniel Dold ⋅ Emanuel Sommer ⋅ Julius Kobialka ⋅ Oliver Dürr ⋅ David Rügamer

While parameter-efficient fine-tuning methods like low-rank adaptation (LoRA) are standard for large language models, principled estimation of epistemic uncertainty remains challenging. Recent results in the LoRA regime suggest that discrete multi-mode approaches such as deep ensembles offer little benefit over single-mode methods. This contradicts broader observations in deep learning, where ensembling independent optima typically improves generalization, and linking these modes through continuous low-loss valleys further enhances Bayesian model averaging (BMA). Whether such structure exists in the LoRA space and whether it yields functional diversity missed by local or discrete methods has not been studied. We introduce LoRA-Curve, a segmented Bézier curve parameterization in the LoRA space, with two variants: a free configuration that jointly optimizes all control points, and an anchored configuration that connects independently fine-tuned LoRA optima. We prove pathwise continuity and Lipschitz regularity of the loss along the curve and empirically show, across reasoning and classification benchmarks with Qwen2.5 7B, that linear interpolation encounters loss barriers, while our anchored multi-segment curves connect independent optima through continuous low-loss valleys. Combined with flat-minima perturbations and a Jensen-Shannon divergence regularizer, LoRA-Curve yields measurably higher mutual information of the predictive distribution without sacrificing performance, and links continuous parameter-space traversal to functional diversity.


On the Rademacher Complexity of Graph Neural Networks: Unifying Expressivity and Geometry

Martin Carrasco ⋅ Caio Deberaldini Netto ⋅ Ehimare Okoyomon ⋅ Aneeqa Mehrab ⋅ Vahan Martirosyan ⋅ Caterina Graziani

Understanding the interplay between generalization, expressivity, and the geometry of the input space is a central challenge in graph learning. The expressivity of Graph Neural Networks (GNNs) is typically characterized through their correspondence with graph invariants, such as those from the Weisfeiler-Leman (WL) hierarchy. While more expressive GNNs can distinguish a richer set of graphs, they are also associated with weaker generalization guarantees. Previous works have addressed this trade-off using the VC dimension, a purely combinatorial measure, independent of the training data. In this work, we adopt a data-dependent measure of generalization, the empirical Rademacher complexity, and derive tight generalization bounds that jointly consider the expressive power of GNNs and the geometry of the underlying input space. Specifically, any graph invariant that upper-bounds a GNN's expressive power partitions the input space into equivalence classes, and we show that the empirical Rademacher complexity is controlled by the distribution of training samples across these classes. Moving beyond discrete partitions, we incorporate the geometry of the input space and derive covering-number bounds under Lipschitz continuity, showing that the complexity cost can be mitigated when the hypothesis class remains smooth over the data geometry. In addition, we prove that the empirical Rademacher complexity is Lipschitz continuous with respect to the Wasserstein distance between empirical measures supported on different datasets. This yields robustness and generalization guarantees under sampling variability. Importantly, our framework is not restricted to message-passing GNNs or WL, but extends to arbitrary GNN architectures and their associated invariants, providing a step toward a unified theory of GNN generalization.


Optimal algorithmic complexity of inference in quantum kernel methods

Elies Gil-Fuster ⋅ Seongwook Shin ⋅ Sofiene Jerbi ⋅ Jens Eisert ⋅ Maximilian J Kramer

Quantum kernel methods are among the leading candidates for achieving quantum advantage in supervised learning. A key bottleneck is the cost of inference: evaluating a weighted sum of $N$ kernel values to additive precision $\varepsilon$, where $\alpha$ is the vector of trained coefficients. The standard approach estimates each term independently via sampling, yielding a query complexity of $\mathcal{O}(N\lVert\alpha\rVert_2^2/\varepsilon^2)$. In this work, we combine two independent improvements: estimating the kernel values via quantum amplitude estimation and collecting the sum under a single observable. We show that the improved approach achieves a query complexity of $\mathcal{O}(\lVert\alpha\rVert_1/\varepsilon)$, removing the dependence on $N$ from the query count and yielding a quadratic improvement in both $\lVert\alpha\rVert_1$ and $\varepsilon$. We prove a matching lower bound of $\Omega(\lVert\alpha\rVert_1/\varepsilon)$, establishing query-optimality. Beyond query complexity, we also analyze how these improvements translate into gate costs and show that the query-optimal strategy is not always optimal in practice from the perspective of gate complexity. We identify a different strategy based on importance sampling, which yields a lower total gate count. We thus provide both a query-optimal algorithm and a practically-optimal choice of strategy depending on hardware capabilities, along with a complete landscape of intermediate methods to guide practitioners. All algorithms require only amplitude estimation as a subroutine and are thus natural candidates for early-fault-tolerant implementations.


Optimal In-context Adaptivity and Distributional Robustness of Transformers

Tianyi Ma ⋅ Tengyao Wang ⋅ Richard J Samworth

We study in-context learning problems where a Transformer is pretrained on tasks drawn from a mixture distribution $\pi=\sum_{\alpha\in\mathcal{A}} \lambda_{\alpha} \pi_{\alpha}$, called the pretraining prior, in which each mixture component $\pi_{\alpha}$ is a distribution on tasks of a specific difficulty level indexed by $\alpha$. Our goal is to understand the performance of the pretrained Transformer when evaluated on a different test distribution $\mu$, consisting of tasks of difficulty $\beta\in\mathcal{A}$, and with potential distribution shift relative to $\pi_\beta$. In particular, we consider nonparametric regression problems with random smoothness, and multi-index models with both random smoothness and random effective dimension. We prove that a Transformer pretrained on such mixture distributions can adapt to the difficulty of a new task in-context and achieve the optimal prediction rate, uniformly over test distributions in a chi-squared divergence ball. Thus, the pretrained Transformer is able to achieve faster rates of convergence on easier tasks and is robust to distribution shift at test time. Finally, we prove that even if an estimator had access to the test distribution $\mu$, the convergence rate of its expected risk over $\mu$ could not be faster than that of our pretrained Transformers, thereby providing a more appropriate optimality guarantee than minimax lower bounds.


Optimized Forward-Backward Rematerialization for Memory-Efficient Pipeline Parallel Training

Adrien Aguila-Multner ⋅ Olivier Beaumont ⋅ Lionel Eyraud-Dubois ⋅ Julia Gusak

Pipeline parallelism is a key technique for scaling deep network training across multiple devices. Recent works have significantly reduced pipeline idle time by improving scheduling efficiency. Decoupling the computation of gradients with respect to weights and activations led to the development of schedules with almost no idle time. However, these methods still require substantial memory, limiting their applicability on resource-constrained hardware. Our first contribution is to introduce recomputation to the backward pass, extending rematerialization beyond the forward pass. This enables executing schedules with decoupled gradient computations under much tighter memory constraints. Our second contribution is to consider more flexible rematerialization strategies, with individual per-microbatch decisions. We provide a unified optimization approach that, given a model and hardware memory constraints, formulates and solves an Integer Linear Programming (ILP) problem to determine the optimal per-microbatch, per-GPU rematerialization strategy for a given schedule, applicable to both one-wave and multi-wave pipeline schedules. With these tools, we show that when using rematerialization, the best scheduling algorithm varies according to the device memory constraints. Experiments demonstrate the effectiveness of all three contributions, showing that our approach enables efficient training of larger models under tight memory budgets, adapts optimally to varying memory capacities, and reduces recomputation overhead compared to existing recomputation solutions.

Inversion-free methods for text-driven image editing construct direct ODE paths between source and target distributions using pre-trained flow models, but operate on single trajectories vulnerable to noise-induced drift with no mechanism to exploit the geometry of plausible edits. We propose Transport-Guided Flow Editing, which recasts the editing problem in distributional terms: we construct source and target particle clouds in latent space, solve entropic optimal transport between them with an edit-direction-aware cost, and use the resulting barycentric map to define a time-varying correction field for the editing ODE. The transport plan adapts online at each integration step, and the correction strength is governed by a variational co-state derived from the OT anchor, responding to accumulated trajectory deviation rather than following a fixed schedule. The output is a single edited image, but one whose trajectory has been steered by distributional information inaccessible to single-trajectory methods. Extensive experiments demonstrate state-of-the-art structure preservation with superior semantic alignment and consistent human preference over existing baselines.


ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning

Zuhao Yang ⋅ Kaichen Zhang ⋅ Sudong Wang ⋅ Keming Wu ⋅ Zhongyu Yang ⋅ Bo Li ⋅ Xiaojuan Qi ⋅ Shijian Lu ⋅ Xingxuan Li ⋅ Lidong Bing

Training large multimodal models (LMMs) via reinforcement learning (RL) to natively invoke video-processing tools (e.g., cropping) has become a promising route to long-video understanding. However, existing native-RL methods dispatch tool calls sequentially (i.e., one per turn): a single wrong crop propagates errors without peer correction, multi-turn tool calls corrupt context, and inference cost scales linearly with the number of turns. We introduce **ParaVT**, the first multi-agent end-to-end RL-trained framework for **Para**llel **V**ideo **T**ool calling, dispatching multiple time-window crops in a single turn for cleaner context and better fault tolerance. Yet applying standard RL to ParaVT reveals an obstacle we term the *Tool Prior Paradox*: the pretrained tool priors that enable tool exploration also destabilize cold-started structural format and expose the skip-tool reward shortcut under temperature sampling. A cross-model contrast on a weaker-prior LMM supports this claim: format stays stable but RL elicits zero tool calls, indicating that prior strength is the shared driver of both format collapse and tool exploration. We propose **PARA-GRPO** (**P**arseability-**A**nchored and **R**atio-g**A**ted **GRPO**), which augments standard RL with two complementary mechanisms: (i) a *targeted format reward* applied only at the structural-token positions most prone to collapse, and (ii) a *per-prompt frame-budget randomization* that creates training prompts where calling the tool yields a measurable reward signal over skipping it. Across six long-video understanding benchmarks, ParaVT improves over the Qwen3-VL baseline by $+7.9\%$ on average, with PARA-GRPO lifting training-time format compliance from $0.13$ to $0.64$. As tool capabilities become increasingly internalized in modern LMMs, RL must cooperate with the resulting priors, and ParaVT offers a general recipe for agentic RL. Code, data, and model weights will be released.


Pathways of Visual Information Flow in Vision-Language Models

Israfel Salazar ⋅ Stella Christina Frank ⋅ Dan Oneata ⋅ Desmond Elliott ⋅ Constanza Fierro

We study how visual information is routed in vision-language models (VLMs). Using causal patching on controlled synthetic and natural datasets, we find that models rely on two distinct pathways to solve visual tasks: A \emph{direct pathway}, where visual information is retained in image token representations and read out by the final token at later layers, and a \emph{text-mediated pathway}, where visual information is first transferred to the query tokens and then read out by the final token. Across three visual tasks, we show that pathway selection is task-dependent, and that data distribution and prompt design can also modulate which pathway is used to solve the image-based query. Moreover, using attention knockouts and corrupted-input patching, we find that these pathways are flexible, under certain interventions, models can rely on the text-mediated pathway as a fallback when the usual pathway is ablated. This behavior unifies findings in prior work and shows that ablation-based interventions can reveal what models could do rather than what they normally do. Together, our results provide a mechanistic characterization of visual information flow in VLMs and highlight the flexibility of their internal mechanisms under intervention.


P-EAGLE: Parallel-Drafting EAGLE with Scalable Training

Xin Huang ⋅ Mude Hui ⋅ Jaime C Salas ⋅ Yue Sun ⋅ Nathan Pemberton ⋅ Ashish Khetan ⋅ Xiang song ⋅ George Karypis

Reasoning LLMs produce longer outputs, requiring speculative decoding drafters trained on extended sequences. Parallel drafting—predicting multiple tokens per forward pass—offers latency benefits over sequential generation, but training complexity scales quadratically with the product of sequence length and parallel positions, rendering long-context training impractical. We present P(arallel)-EAGLE, which transforms EAGLE from autoregressive to parallel multi-token prediction via a learnable shared hidden state. To scale training to long contexts, we develop a framework featuring attention mask pre-computation and sequence partitioning techniques, enabling gradient accumulation \textit{within} individual sequences for parallel-prediction training. We implement P-EAGLE in vLLM and demonstrate speedups of 1.10×–1.36× over autoregressive EAGLE-3 across GPT-OSS 120B, 20B, and Qwen3-Coder 30B.


Playing Markov Games Without Observing Payoffs

Daniel Ablin ⋅ Alon Peled-Cohen

Optimization under uncertainty is a fundamental problem in learning and decision-making, particularly in multi-agent systems. Previously, Feldman, Kalai, and Tennenholtz [2010] demonstrated the ability to efficiently compete in repeated symmetric two-player matrix games without observing payoffs, as long as the opponent’s actions are observed. Extending this capability to the Markovian setting remains an open problem. In this paper, we introduce and formalize a new class of zero-sum symmetric Markov games, which extends the notion of symmetry from matrix games to the Markovian setting. We prove that a learner observing only the opponent's action sequence, without access to payoff information can successfully compete against an adversary possessing complete knowledge of the game. We formalize three distinct notions of symmetry in this domain and reveal a surprising structural hierarchy: the most holistic definitions of symmetry impose restrictive constraints that actually simplify the learning landscape, whereas the ``simplest'' definition represents the most general and challenging setting. We provide polynomial-time algorithms for all three settings that achieve sublinear regret. Crucially, we demonstrate that despite the complex Markovian dynamics, a simple strategy of locally mimicking the opponent's actions suffices to guarantee robustness. This finding significantly broadens the class of games where robust learning is possible under severe informational disadvantage, proving that knowledge of the transition laws is not required to force a draw.


Poison-then-Hide: Finetuning-Activated Backdoor Attack on Pretrained Vision Encoders

Qixuan Jin ⋅ Abinitha Gourabathina ⋅ Vinith Suriyakumar ⋅ Walter Gerych ⋅ Marzyeh Ghassemi

Backdoor attacks threaten the integrity of machine learning models by allowing attackers to control model behavior through triggers. Models can be compromised during finetuning by backdoor attacks through either poisoned data or adversarial training objectives. Because existing finetuning-activated attacks assume limited domain shift or frozen encoder layers, they often fail under full-model finetuning. We target a stealthy finetuning-activated attack where a dormant backdoor is implanted in a pretrained encoder, and later activated by finetuning on clean downstream data. We propose Poison-then-Hide, a novel attack that remains effective when the entire model is finetuned in a domain transfer. Our approach consists of three components: trigger optimization, base encoder poisoning, and targeted unlearning to conceal the backdoor. We evaluate our method on six datasets and three model architectures, and achieve state-of-the-art (up to 97%) attack success rates. We find that two design choices - jointly learning the benign target task and the backdoor during encoder poisoning, and optimizing the trigger for attack robustness using an ensemble of simulated finetuned models - are critical to the attack's success. We demonstrate that standard detection and mitigation defenses cannot fully remove the backdoor, which can reappear after benign finetuning.

Efficiency gains in artificial intelligence are often interpreted as environmental progress, but this view is incomplete because AI usage may exhibit rebound effects: lower unit costs can induce larger models, more training runs, stricter latency requirements, wider deployment, and higher inference volume. Although rebound effects in AI have begun to receive attention, they are still studied mostly through qualitative or scenario-based analyses. This paper takes the position that AI efficiency gains must be quantitatively and mathematically analyzed against rebound effects. We support this position with a simple game-theoretic model showing that an efficiency improvement can inadvertently increase the total energy consumption. The model further shows that rebound can appear even in a simple setting and in forms not captured by demand-reduction arguments alone: the amount of requested work may decrease while stricter deadlines make the computation more energy intensive.


Position: Fair Representations Cannot Hold What They Promise

Shai Ben-David ⋅ Tosca Lechner ⋅ Ruth Urner

This position paper argues that {\sl fair representations cannot hold what they promise}. Fair representation learning has been a significant trend in research on fairness for machine learning models over the past decade. The main idea of fair representation learning is to provide a data preprocessing mechanism which can facilitate fairness for downstream tasks, and research papers in this area often insinuate a fairness guarantee for all downstream tasks without further qualification. However, provably this cannot be satisfied by any non-degenerate representation. We argue that such over-promising can be harmful in downstream applications and that the responsibility for fairness cannot safely be outsourced to a representation provider in the way current research seems to suggest. As a minimal requirement for a way forward, we suggest that newly proposed fairness representations should be accompanied by compatibility tests that would allow a user to verify whether the representation will actually guarantee fairness for the task at hand.


Prediction-Powered Active Testing

Kianoosh Ashouritaklimi ⋅ Valentin Kilian ⋅ Daolang Huang ⋅ Thomas Rainforth ⋅ Francois Caron

Evaluating modern machine learning models often requires labels on large test pools, yet obtaining these labels can be expensive. Active testing reduces this cost by adaptively selecting which test points to label, but existing unbiased estimators do not fully exploit cheap black--box predictions that are often available for the entire pool. We introduce \textbf{Prediction--Powered Active Testing (PPAT)}, a label--efficient risk estimation framework that combines the unbiased LURE estimator with a prediction--powered control variate. Rather than using proxy predictions as biased pseudo--labels, PPAT uses them to residualise the loss, preserving unbiasedness while reducing variance. This control--variate perspective also changes the optimal acquisition problem: we derive residualised oracle proposals and practical surrogate--based acquisition rules tailored to the PPAT estimator. We further establish asymptotic normality for LURE and PPAT, enabling asymptotically valid confidence intervals. Across tabular regression and image classification tasks, PPAT consistently improves over existing active testing baselines, remains unbiased, and reaches the desired coverage level with substantially fewer labels.


Principia: Relational Physics Tests for Video Models

Varun V Thozhiyoor ⋅ Shivam Tripathi ⋅ Venkatesh Babu Radhakrishnan ⋅ Anand Bhattad

Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on frame rate, object scale, and camera calibration, all of which are often ambiguous or unavailable in generated video. We propose a different approach. When two objects in the same scene obey the same physical law, their motions must satisfy predictable relationships, and these relationships hold independent of calibration. We introduce Principia, a benchmark that evaluates Newtonian physics through relational consistency between paired objects. Principia spans eight phenomena—gravity, restitution, friction, rotational inertia, projectile motion, momentum, pendulum, and mass-spring oscillation—across translational, rotational, collisional, and oscillatory dynamics, using real-world scenes recorded under controlled protocols. We also introduce a calibration-independent consistency score that quantifies physical violation directly in image space. Across thousands of generations from five state-of-the-art video generators, no model exceeds 0.42 on Principia despite all scoring above 0.7 on VBench. Vision-language models show inverse failure profiles, succeeding where generators fail and vice versa.


Privacy Amplification Persists under Unlimited Synthetic Data Release

Clément Pierquin ⋅ Aurélien Bellet ⋅ Marc Tommasi ⋅ Matthieu Boussard

We study privacy amplification by synthetic data release, a phenomenon in which differential privacy guarantees improve when releasing only synthetic data rather than the private generative model itself. Recent work by Pierquin et al. (2025) established the first formal amplification guarantees for linear generators, but they apply only in asymptotic regimes where the model dimension far exceeds the number of released synthetic records, limiting practical relevance. In this work, we show that this restriction is not fundamental and establish that, surprisingly, for linear generators, privacy amplification can persist even when releasing an unbounded number of synthetic records. To prove this result, we develop new analytical tools for characterizing the privacy loss induced by synthetic data release. In particular, we leverage a sufficient-statistics reduction to characterize privacy leakage, and introduce a novel criterion that upper bounds Rényi divergences in terms of Fisher information. Our analysis provides structural insights that may guide the development of tighter privacy guarantees for more complex release mechanisms. Finally, experiments with variational autoencoders trained using DP-SGD further support the relevance of our theory beyond the linear setting.


Privacy by Postprocessing the Discrete Laplace Mechanism

Quentin Hillebrand ⋅ Jacob Imola ⋅ Rasmus Pagh ⋅ Sia Sejer

We show that an "old dog," the classical discrete Laplace (aka. geometric) mechanism, can "perform new tricks": * It can be post-processed to yield a simple, unbiased estimator of any subexponential function $f$ of the original data, giving a simple, discrete, multivariate version of the recent unbiasing result for the Laplace mechanism by Calmon et al. (FORC '25). * It can be post-processed to output the same distribution as the Laplace mechanism or the Staircase mechanism with identical privacy parameters. Thus, the discrete Laplace mechanism is a versatile mechanism that should be preferred over the Laplace and Staircase mechanisms whenever the data is discrete (or can be made discrete while controlling $\ell_1$-sensitivity). We show bounds on the variance of our estimator, compared to the mean square error of the biased estimator that simply evaluates the $f$ on the output of the mechanism. Though our unbiased estimator has exponential running time for worst-case functions, we show that it can often be computed in linear or polynomial time for concrete, structured functions. We showcase the properties of our methods empirically with several use cases including profile and entropy estimation, as well as distributed/federated data analysis applications in which unbiasedness is key to accuracy.


Privacy Guarantees in Posterior Sampling under Contamination

Shenggang Hu ⋅ Hongsheng Dai ⋅ Louis Aslett ⋅ Murray Pollock ⋅ Gareth Roberts

In recent years, differential privacy has been adopted by tech-companies and governmental agencies as the standard for measuring privacy in algorithms. In this article, we study differential privacy in Bayesian posterior sampling settings. We begin by considering differential privacy in the most common privatisation setting in which Laplace or Gaussian noise is injected into the output. In an effort to achieve better differential privacy, we consider adopting {\em Huber's contamination model} for use within privacy settings, and replace at random data points with samples from a heavy-tailed distribution ({\em instead} of injecting noise into the output). We derive bounds for the differential privacy level $(\epsilon,\delta)$ of our approach, without requiring bounded observation and parameter spaces, a restriction commonly imposed in the literature. We further consider for our approach the effect of sample size on the privacy level and the rate at which $(\epsilon,\delta)$ converges to zero. Asymptotically, our contamination approach is fully private with no information loss. We also provide examples of inference models for which our approach applies, with theoretical convergence rate analysis and simulation studies.


QGround: Condition-Wise Evidence Aggregation for 3D Grounding with 2D VLMs

Fengyun Wang ⋅ Jian Wang ⋅ Dingwei Zhang ⋅ Jinhui Tang ⋅ QIANRU SUN

Using 2D vision-language models (VLMs) for 3D grounding raises a decision-formulation problem: how should image-level evidence be turned into a 3D object decision? This problem matters as pretrained 2D VLMs provide strong visual-semantic evidence, while 3D grounding requires disambiguating objects in cluttered scenes. Current VLM-based formulations make this conversion holistically: the model either selects one object from a rendered candidate set or assigns one overall match judgment to each candidate. Such decisions obscure the evidence needed for disambiguation, since different candidates may satisfy different subsets of the category, attribute, and relational cues. In this work, we reformulate 3D grounding with 2D VLMs as condition-wise evidence aggregation. The reformulation follows three principles: use natural object-centric images rather than rendered candidate images; read the referring expression condition by condition rather than as one holistic query; and use model preference rather than a hard match output as grounding evidence. We instantiate these principles in QGround, a training-free implementation of this decision formulation. Experiments on ScanRefer and Nr3D show improved grounding performance, with analyses demonstrating stronger candidate disambiguation and more interpretable decisions.


Quantitative Local Convergence of Mean-Field Stein Variational Gradient Flow

Lénaïc Chizat ⋅ Maria Colombo ⋅ Roberto Colombo ⋅ Xavier Fernández-Real

Stein Variational Gradient Descent (SVGD) is a deterministic interacting-particle method for sampling from a target probability measure given access to its score function. In the mean-field and continuous-time limit, it is known that the flow converges weakly toward the target, but no quantitative rate is known for the last iterate. In this paper, we establish quantitative local convergence in strong norms for this dynamics, when the interaction kernels is of Riesz type on the $d$-dimensional torus. Specifically, assuming that the initial density and the target are smooth and close in $L^2$-norm, we obtain explicit polynomial convergence rates in $L^2$-norm that depend on the dimension and on the regularity parameters of the kernel, the initialization and the target. We further show that these rates are sharp in certain regimes, and support the theory with numerical experiments. In the edge case of kernels with a Coulomb singularity, we recover the global exponential convergence result obtained in prior work. Our analysis is inspired by recent results on Wasserstein gradient flows of kernel mean discrepancies.

Image-free Zero-Shot Learning (I-ZSL) aims to extend a pre-trained classifier to unseen classes without using training images or image features during adaptation. Existing I-ZSL methods often depend on pre-defined class descriptions and fixed text encoders, which can be misaligned with the classifier space of the deployed model. We propose Geometry-Regularized Relational Weight Synthesis (GeoRWS), a read-only classifier expansion framework that learns to synthesize unseen-class classifier weights from pairwise semantic relations and observed seen-class weights. Given class-pair affinity descriptions generated by a frozen large language model, GeoRWSuses an adaptive semantic encoder to estimate affinity coefficients over seen classifiers through leave-one-out reconstruction. The training objective further includes a classifier-geometry regularizer, which anchors the learned coefficients to neighborhoods in the seen-class classifier head, and a semantic-consistency regularizer, which improves stability under perturbations of pairwise semantic descriptions. At inference, GeoRWS synthesizes unseen-class weights as mixtures of seen-class weights and injects them into the classifier head while keeping the feature extractor fixed. Experiments on standard ZSL and GZSL benchmarks show consistent improvements over image-free and adapted zero-shot baselines, especially on fine-grained recognition tasks.


Recovering Clean Evaluation Metrics from Contaminated Benchmarks

Míriam Barrabés ⋅ Daniel Mas Montserrat ⋅ Maria Perera ⋅ Alexander Ioannidis

When benchmark examples overlap with a model’s training data, benchmark contamination can substantially inflate reported evaluation metrics, yet training provenance is often unavailable. We study post-hoc recovery of uncontaminated evaluation metrics in supervised tabular classification and regression using only prediction–label pairs from a potentially contaminated benchmark. We introduce MILAN (Metric Integrity via Leakage-Aware Normalization), a meta-trained framework that models a contaminated benchmark as a mixture of training-origin and clean evaluation examples and estimates soft posterior weights for the clean component to correct evaluation metrics. MILAN supports a broad class of metrics through weighted estimators and metric-specific correction procedures, and can optionally incorporate contamination-rate priors and sample-wise difficulty scores. To quantify uncertainty, MILAN additionally provides split-conformal prediction intervals with marginal coverage across exchangeable recovery problems. Across a large suite of simulated recovery problems with held-out datasets and model families, MILAN consistently improves clean-metric recovery over thresholding, clustering, mixture-model, anomaly-detection, and membership-inference baselines. In a polygenic risk score case study using UK Biobank evaluation data, MILAN improves agreement with independent FinnGen reference metrics. We provide a scikit-learn-style implementation at \url{hidden-for-submission}.


ReefNet: A Large-Scale Dataset and Benchmark for Fine-Grained Coral Reef Recognition

Abdulwahab Felemban ⋅ Yahia Battach ⋅ Faizan F Khan ⋅ Yuqian Fu ⋅ Xuhui Liu ⋅ Yesmeen Khattab ⋅ Yousef A Radwan ⋅ Xiang Li ⋅ Natalie Dunn ⋅ Susanne Bähr ⋅ Tullia I Terraneo ⋅ Fabio Marchese ⋅ Sara Beery ⋅ Burton H Jones ⋅ Francesca Benzoni ⋅ Mohamed Elhoseiny

Coral reefs are rapidly declining under anthropogenic pressures (e.g., climate change), creating an urgent need for scalable and automated monitoring. Progress in data-driven coral analysis, however, is constrained by the scarcity of large-scale datasets with fine-grained labels that are taxonomically consistent across sites and studies. To address this gap, we introduce ReefNet, a large-scale public coral reef image dataset with point-level annotations mapped to the World Register of Marine Species (WoRMS) taxonomy. ReefNet aggregates imagery from 76 curated CoralNet sources and an additional reef site from Al-Wajh (Red Sea), totaling approximately 925K genus-level hard coral annotations. Through expert-driven verification and targeted filtering, we derive a high-confidence benchmark subset with 92% expert agreement over 39 hard-coral label classes, enabling reliable evaluation under realistic label noise and strong class imbalance. Beyond dataset construction, we establish a comprehensive benchmark spanning zero-shot, cross-domain few-shot adaptation, within-source evaluation, and cross-source transfer to the Al-Wajh dataset. Experiments with state-of-the-art vision–language models (VLMs), multimodal large language models (MLLMs), and vision-only backbones reveal substantial degradation in zero-shot and extremely few-shot regimes, while adaptation with in-domain supervision yields large gains yet still leaves a persistent gap under cross-source shift and on long-tail genera. These results highlight fundamental challenges in applying general-purpose multimodal models to biodiversity monitoring and underscore the importance of large-scale, taxonomically grounded, high-quality datasets. ReefNet serves as both a benchmark and a training resource for advancing fine-grained coral reef understanding. ReefNet data and trained models are provided on Hugging Face, and code for accessing and benchmarking the data is provided on GitHub.


Reinforcement Learning for Long-Horizon Unordered Tasks: From Boolean to Coupled Reward Machines

Kristina Levina ⋅ Nikolaos Pappas ⋅ Athanasios Karapantelakis ⋅ Aneta Vulgarakis Feljan ⋅ Jendrik Seipp

Reward machines (RMs) inform reinforcement learning agents about the reward structure of the environment, enabling support for non-Markovian tasks and improving sample efficiency. However, learning with RMs is ill-suited for long-horizon problems in which subtasks can be completed in any order. In such cases, the amount of information to learn increases exponentially with the number of unordered subtasks. We address this limitation by introducing three generalisations of RMs: (1) Numeric RMs allow users to express complex tasks in a compact form. (2) In Agenda RMs, states are associated with an agenda that tracks the remaining subtasks to complete. (3) Coupled RMs have coupled states associated with each subtask in the agenda. Furthermore, we introduce a new task-decomposition $Q$-learning-based algorithm that leverages coupled RMs and preserves global optimality guarantees: QCoRM. Our experiments across four domains---featuring both discrete and continuous action and state spaces---show that QCoRM scales better than baseline algorithms for long-horizon problems with unordered subtasks.


Resource-Aware Parameter-Efficient Model Adaptation for Onboard High-Dimensional Data

Qiyang Zhang ⋅ Xinhao Li ⋅ Lei Shi ⋅ Zheng Lin ⋅ Jinfeng Wen ⋅ Ao Zhou ⋅ Shangguang Wang

Onboard satellite models often require frequent updates, but the weights adapted to earlier data distributions can quickly become outdated. However, updating large-scale model parameters in orbit presents significant challenges due to the limited uplink bandwidth of Low Earth Orbit (LEO) satellite systems, particularly for hyperspectral satellite imagery, where high-dimensional spectral–spatial inputs lead to increased model size and update costs. Existing full fine-tuning methods are thus expensive to retrain and difficult to deploy under strict communication constraints. To address this challenge, we propose NE-LoRA, a parameter-efficient adaptation framework for bandwidth-constrained onboard hyperspectral model updates. NE-LoRA combines a primary low-rank branch with a nonlinear auxiliary branch to capture both global update trends and complex spectral–spatial variations. Additionally, we introduce a differentiated training strategy for multi-matrix adapters, motivated by the asymmetric initialization and gradient dynamics of different adapter matrices. Experiments on four hyperspectral datasets and three representative backbone models demonstrate that NE-LoRA consistently outperforms LoRA-based baselines and remains competitive with, and in several cases superior to, full fine-tuning. Across the evaluated settings, NE-LoRA updates only a small fraction of the total parameters on average while preserving low deployment overhead, offering a favorable accuracy–communication trade-off for onboard hyperspectral adaptation.


RigRecon: Efficient Rig-Aware Street Reconstruction via Dual-Path Spatio-Temporal Interaction

Jiaming Guo ⋅ Hongcheng Luo ⋅ Cheng Chi_ ⋅ Lijun Zhou ⋅ Bing Wang ⋅ Guang Chen ⋅ Hangjun Ye ⋅ Jiaqi Yang ⋅ Xiaowei Zhou ⋅ Haiyang Sun ⋅ Sida Peng

Reconstructing multi-camera videos is a fundamental task in autonomous driving. However, recent learned methods either rely on costly attention-based networks for global geometry modeling or require complex post-hoc alignment for chunked predictions, limiting their scalability to long multi-view rig sequences. We propose RigRecon, a generalized rig-aware reconstruction framework that enables efficient cross-view interaction in a low-resolution space while asynchronously decoding high-resolution depth maps in a frame-wise manner. Input images are first down-sampled and processed by a transformer backbone to capture coarse global structure, which is then provided to a decoder to reconstruct frame-wise high-resolution depth maps. By fully leveraging calibration and low-resolution feature interaction, RigRecon achieves consistent-scale geometry and accurate pose estimation with significantly reduced memory cost. It achieves state-of-the-art performance on 3D reconstruction, depth estimation, and camera pose estimation, and can efficiently process videos with over 1000 images in a single feedforward pass.

This position paper argues that the robotics and AI community should adopt disclosure standards for demonstration videos. As physical and financial constraints limit direct public access to embodied AI, public dissemination of progress in robotics is mainly through curated demonstration videos rather than direct experience with the technology. In a large-scale (N=9,341), pre-registered experiment across six countries, we find that the general public struggles to discern both the visual authenticity (e.g., distinguishing real footage from CGI) and behavioral authenticity (e.g., distinguishing autonomy from teleoperation) of robot demonstration videos. The omission of these details distorts public perception, leading people to overestimate the ability of robots to perform tasks and navigate real world environments. Using evidence from behavioral science, we design a disclosure intervention and demonstrate that the inclusion of clear disclosures in robot demonstration videos significantly improves discernment of visual and behavioral authenticity, as well as understanding robots' capabilities. Consequently, we propose the adoption of disclosure standards for robot demonstration videos to accurately calibrate public expectations and promote transparency in robotics research.


Robust Amortized Simulation-Based Inference via Learned Error Models

Matthew O'Callaghan ⋅ Kaisey Mandel ⋅ Gerard Gilmore

Recent advances in neural density estimation have enabled amortized Bayesian inference for complex stochastic simulators. However, these methods rely on simulators accurately reflecting the true data-generating process and can degrade significantly under model misspecification. We consider the setting where multiple unlabeled observations are available and introduce robust variational neural posterior estimation (RVNP), an amortized Bayesian inference method that uses an importance-weighted autoencoder to jointly learn a misspecification-robust posterior and an explicit error model that captures the misspecification gap. Our results show that RVNP can recover robust posterior inference in a data-driven and interpretable manner, outperforming previous methods across multiple metrics and misspecified benchmarks.


Robust Dreamer: Deviation-Aware Latent Gaussian Memory for Action-Controlled AR Video Generation

Hanlin Chen ⋅ Jiaxin Wei ⋅ Xibin Song ⋅ Yifu Wang ⋅ Senbo Wang ⋅ Hongdong Li ⋅ Pan Ji ⋅ Gim Hee Lee

Frame-wise action-controlled image-to-video generation is a promising paradigm for interactive world simulation, where each control signal should elicit an immediate visual response. However, maintaining visual fidelity and 3D consistency over long autoregressive rollouts remains challenging. Existing 3D-aware methods often suffer from catastrophic drift due to two impediments: information loss from \textit{Latent--RGB Cycling}, where generated latents are repeatedly decoded to RGB and re-encoded for future conditioning, and the training--inference gap induced by the \textit{error-free hypothesis}, where clean training memory fails to match prediction-corrupted inference memory. To address these challenges, we present \textbf{Robust Dreamer}, a memory-augmented framework built around how to design 3D memory and how to use it robustly. First, we introduce \textbf{Latent Gaussian Memory}, which anchors diffusion latents inherited from the generation process to Gaussian primitives and recalls them via latent-space Gaussian splatting. This provides dense, geometry-aware, view-aligned conditioning while avoiding accumulated degradation from repeated VAE conversion. Second, we propose \textbf{Deviation Learning with Dynamic Deviation Archive}, which synthesizes rollout-induced latent deviations through a one-step approximation, stores them by autoregressive stage and denoising timestamp, and injects them into historical memory during training. This exposes the generator to realistic corrupted memory states and teaches internal correction before inference. Experiments on ScanNet, DL3DV, and OmniWorldGame demonstrate state-of-the-art long-horizon performance.


RubiConv - Efficient Boundary-Respecting Convolutions

Linda Friso ⋅ Annie Marsden ⋅ Xinyi Chen ⋅ Arushi Gupta ⋅ Peter Bartlett ⋅ Mark Braverman ⋅ Elad Hazan

Convolutional architectures have emerged as powerful alternatives to Transformers for sequence modeling. The primary advantage is that they offer improved theoretical sequence length complexity by leveraging the Fast Fourier Transform (FFT). However, this theoretical improvement does not always meaningfully land in practice. One critical obstacle is that applying standard FFTs is not amenable to the large-scale training pipeline wherein data is packed from different sources into a single sequence for hardware efficiency. Indeed, standard FFT algorithms are not easily amenable to document packing. Existing workarounds suffer from severe inefficiencies, crippling the practical performance of convolutional architectures. We close this gap with RubiConv, a novel algorithm for performing hardware-efficient, boundary-respecting convolutions on packed sequences. Extensive experiments show that RubiConv achieves significant speedups over both attention and standard FFT-based baselines. This work makes the theoretical efficiency of long convolutional models a practical reality for large-scale, real-world data packing.


Safe Few-Step Generation via Velocity Editing

Yujin Choi ⋅ Jaehong Yoon

Flow matching has recently emerged as a strong paradigm for state-of-the-art text-to-image (T2I) generation, enabling high-quality generation with a small number of sampling steps. As these models are increasingly integrated into real-world applications, ensuring safe and non-sensitive content generation has become a critical requirement. However, adapting safety and concept removal methods to this new generation framework remains an open challenge. Specifically, prior methods largely rely on iterative trajectory steering across a number of denoising steps or on CLIP-centeric prompt embedding manipulation. These design assumptions pose fundamental bottlenecks for safety in flow matching-based T2I generation, where limited sampling steps constrain iterative correction and modern context-aware text encoders diminish the effectiveness of embedding-level interventions. In this paper, we propose VESFlow, a training-free safety method tailored to flow matching with extremely few sampling steps. Leveraging the fact that flow matching models learn the marginal velocity (or average velocity in MeanFlow), we directly edit the velocity field via a Bayesian decomposition of the safe-conditional posterior. VESFlow steers the trajectory toward safe outputs while leaving the conditioning prompt unchanged. Building on the observation that VESFlow leaves outputs unchanged under benign prompts, we further introduce a risk score filtering that bypasses velocity editing to reduce computational cost while preserving benign prompt generation. Based on this filtering, we proposed VESFlow+ which provides stronger safety protection when filtered by the risk score. Experimental results show that our method removes the target concept, reducing the detection rate by NudeNet to 6.3\% for step-4 model, while preserving fidelity on benign prompts. Code is available at the supplementary file.


SciReason: A Controllable Benchmark for Scientific Reasoning in LLMs

Pierre Beckmann ⋅ Marco Valentino ⋅ Andre Freitas

Three paradigmatic forms of inference recur across scientific reasoning: deduction, induction, and causal abduction. Reliably evaluating LLMs on these in scientific settings is currently out of reach: scientific benchmarks built on human annotations are costly and lack mechanistic ground truth, while synthetic logical-reasoning benchmarks do not resemble real scientific documents. We introduce SciReason (SciR), a benchmark that combines multi-paradigm reasoning with controllable scientific rendering, anchored on three paradigmatic scientific problems. Tasks are generated from formal objects (deduction tree, inductive rule hypothesis, causal graph) to guarantee verifiable answers, then rendered into multi-document scientific discourse via per-track domain-tuned genres. The construction lets us independently vary two difficulty axes: how hard it is to extract the key information needed for inference, and how hard the principled inference itself is. We test six models. Both axes hurt every model, and their effects compound. The rendering even hurts neuro-symbolic pipelines, which hand inference to a verified solver. The two axes yield a per-model extraction-vs-inference profile: for instance, reasoning models like deepseek-r1 mostly surpass non-reasoning instruct models on the inference axis. To our knowledge, SciR is the first multi-paradigm scientific-reasoning benchmark with parametric control on both extraction and inference difficulty.


SEGA: Spectral-Energy Guided Attention for Resolution Extrapolation in Diffusion Transformers

Javad Rajabi ⋅ Kimia Shaban ⋅ Koorosh Roohi ⋅ David Lindell ⋅ Babak Taati

Diffusion transformers (DiTs) have emerged as a dominant architecture for text-to-image generation, yet their performance drops when generating at resolutions beyond their training range. Existing training-free approaches mitigate this by modifying inference-time attention behavior, often through Rotary Position Embeddings (RoPE) extrapolation combined with attention scaling. However, these strategies apply a uniform and content-agnostic scaling across RoPE components with distinct frequency characteristics, inducing a trade-off between preserving global structure and recovering fine detail. We introduce \textbf{SEGA}, a training-free method that dynamically scales attention across RoPE components according to the latent’s spatial-frequency structure at each denoising step. This adaptive scaling improves both structural coherence and fine-detail fidelity. Experiments show that SEGA consistently improves high-resolution synthesis across multiple target resolutions, outperforming state-of-the-art training-free baselines.

Protein sequence optimization under tight oracle budgets requires methods that explore vast combinatorial spaces while making each evaluation informative. Existing reinforcement learning and off-policy generative approaches often degrade under surrogate noise, and position-agnostic mutation proposals risk disrupting functionally critical residues. We introduce SILO, a trajectory-level self-improvement imitation framework for oracle-budgeted protein design. SILO uses a hierarchical edit policy that decomposes each mutation into a position choice followed by a residue choice. In each active-learning round, the policy samples candidate trajectories via incremental stochastic beam search without replacement (SBS), and a UCB-based proxy ensemble, combined with an alanine-scan fitness score (AFS), selects candidates with functionally relevant edits for in silico oracle evaluation. The policy is then updated by next-action cross-entropy imitation on the round’s best oracle-labeled trajectories, avoiding value-function estimation. Across eight reproduced protein fitness landscapes and five strong baselines from prior work, SILO achieves the highest maximum and top-100 mean fitness on 8 of 8 landscapes within our evaluations, often exhibiting faster early-stage improvement. In low-data and noisy-proxy stress tests on two landscapes per setting, SILO remains competitive or best when several baselines degrade. Ablations show that SBS with AFS account for much of the gains, with iterative imitation providing additional improvement.


SemanticDLM+: Improving Diffusion LLMs through Bias-variance Trade-off in Transition Kernel Design

Keyue Jiang ⋅ Yuxiang Wang ⋅ Yanan Zhao ⋅ Xiang Yu ⋅ Zhao ⋅ Bohan Tang ⋅ Baojian Zhou ⋅ Yanghua Xiao ⋅ Lin Qu ⋅ Xiaoxiao Xu

Diffusion Language Models (DLMs) have demonstrated strong scaling capacity as alternatives to autoregressive language models. However, their performance is highly sensitive to the choice of transition kernels, and poorly designed kernels can lead to issues like training instability, slow convergence, and biased sampling. In this paper, we study this sensitivity through a principled analysis of generalization error and identify three critical factors: asymptotic bias (difficulty in approximating the posterior distribution), exposure bias (error propagation during sampling), and optimization variance induced by kernel dispersion. We further compare different transition kernels: masking diffusion yields sparse and easier posterior-approximation targets, while uniform diffusion provides stronger sampling-side repair but induces harder approximation. Motivated by this trade-off, we revisit a previously overlooked variant, semantic DLM (SemDLM), where the transition kernel corrupts tokens to neighborhoods that are semantically similar. Our theory suggests that SemDLM can serve as a plausible middle ground by reducing the posterior approximation difficulty of uniform diffusion while retaining repair ability. However, we find that SemDLM suffers from a semantic basin problem, where sampling repeatedly stays within a semantic region and produces low-diversity text. To address this, we propose SemDLM+, which adds a global transition and a semantic-frequency penalty during sampling. Experiments on LM1B and OpenWebText show that SemDLM+ improves training dynamics and achieves competitive language modeling and generation quality with satisfactory diversity.


Settling the Sample Complexity of Deterministic Agnostic PAC Learning

Shai Ben-David ⋅ Steve Hanneke ⋅ Farnam Mansouri ⋅ Amirreza Shaeiri

We study the problem of agnostic PAC learning under deterministic labels, where the examples are labeled by an arbitrary deterministic function, which need not belong to the concept class $\mathcal{C}$. This model captures misspecification without stochastic label noise and lies between the realizable and fully agnostic settings. Prior work showed that, when $\mathcal{C}$ has finite diameter $K$ and VC dimension $d$, deterministic labels can improve the sample complexity upper bound from the classical agnostic guarantee $\mathcal{O}(d/\varepsilon^2)$ to $\widetilde{\mathcal{O}}(K/\varepsilon)$, but left a large gap between upper and lower bounds. Our main result sharpens this picture. We prove that every class of VC dimension $d$ and finite diameter $K$ is learnable with sample complexity $ \widetilde{\mathcal{O}}\left( \frac{K^{1/3} d^{2/3}}{\varepsilon} \right), $ and we construct classes for which this bound is tight up to logarithmic factors. We also study the online analogue of this problem and show that the minimax regret is again partly controlled by the diameter of the class. Technically, our upper bounds are obtained via memorization-based learning algorithms that exploit the deterministic label structure.

The performance of machine learning models depends on discovering features in the data that are relevant to solving the given task while ignoring irrelevant variation. To check and visualize whether this is the case, we propose ‘invariance auditing’, a diagnostic tool to analyze feature extractors. Invariance auditing generates diverse inputs which all yield the same features as a given query input, and are therefore treated as identical by the model. Using external knowledge about the task at hand, this enables assessing if the learned invariances do or do not conform to the task's required semantics. Unlike existing work where a dedicated generative model is trained for each feature extractor, our algorithm is training-free and exploits a pretrained diffusion or flow-matching model to sample invariant inputs. Our fiber loss -- which penalizes feature mismatch -- guides the denoising process so the output shares the same representation as the query input. This replaces days of training with a single guided generation procedure at the same quality. Experiments on popular datasets and model types demonstrate that our auditing method reveals invariances spanning very desirable and concerning behavior. For instance, it detects cases where Qwen-2B places patients with situs inversus (heart on the right side) in the same fiber as typical anatomy.


Simultaneous Gradient Learning in First-Price Auctions

Janik Bürgermeister ⋅ Julius Durmann ⋅ Martin Bichler ⋅ Mete Ş Ahunbay ⋅ Bary Pradelski ⋅ Marco Scarsini

We study equilibrium learning in discretized Bayesian games, focusing on first-price auctions as a testbed for gradient-based learning dynamics under private information. While the symmetric independent private values model admits a unique Bayesian correlated equilibrium (BCE) in the continuous limit, standard convergence guarantees for gradient-based learners do not apply: classical no-regret theory predicts convergence only to Bayesian coarse correlated equilibria (BCCE), which can remain large and far from the competitive equilibrium even under fine discretizations. To bridge this gap, we introduce the notion of Bayesian semicoarse correlated equilibrium (BSCCE), an equilibrium concept that refines BCCE by restricting deviations to probability-preserving transformations consistent with projected gradient dynamics. We provide a linear programming characterization of BSCCE with polynomially many variables and constraints and establish the inclusion hierarchy $BCE \subseteq BSCCE \subseteq BCCE$. Our main finding is that the BSCCE set contracts in many cases as the discretization of the Bayesian game is refined. Even with priors for which BCCE fails to concentrate, the diameter of the BSCCE polytope shrinks toward zero, indicating convergence to the unique continuous-game Bayes-Nash equilibrium. These results provide a principled explanation for the observed robustness of gradient-based learning in discretized Bayesian games, showing that these dynamics select a smaller equilibrium set than standard no-regret theory predicts.


SLAyiNG: A Diverse and Community-validated Dataset of Queer Slang

Leonor Veloso ⋅ Lea Hirlimann ⋅ Lucija M Zidar ⋅ Philipp Wicke ⋅ Valentin Hofmann ⋅ Hinrich Schuetze

Queer vernacular is rarely studied in NLP, despite advancements in resources and evaluation for other sociolects and informal language. Because of this, NLP systems often process queer language incorrectly, e.g., they misclassify it as hate speech or generate negative responses. To address this problem, we propose \textsc{Slaying}, the first real-world dataset of English queer slang. \textsc{Slaying} is community-validated, and includes over 500 queer slang terms that pertain to more than 20 queer subcommunities. We argue that queer language data resources have great potential in NLP -- e.g., as components of large pretraining corpora and as the basis for benchmarks -- and can improve queer users' experience of NLP systems. We leverage \textsc{Slaying} for two novel findings in support of this argument: (i) For a number of language models, we show that they are unbiased towards the queer community, but at the same time unable to process its language, i.e., absence of representation bias does not entail the absence of linguistic bias. (ii) Model performance on queer slang varies across queer subcommunities; it is generally worse for slang pertaining to African-American and Latine communities. These findings are relevant for both the queer NLP and the broader ML communities. \textsc{Slaying} is available to the public, and open to future revisions and extensions. \textbf{Warning:} This paper contains profane and potentially offensive language.


SLVMBench: Skill Learning from Video Memory

Yudong Yang ⋅ Guangzhi Sun ⋅ Yixuan Li ⋅ Wei Li ⋅ Zejun MA ⋅ Chao Zhang

We introduce Skill Learning from Video Memory (SLVMBench), the first benchmark that jointly evaluates whether video large language models (video-LLMs) can learn skills from long video memory and apply them to real-time tasks. SLVMBench presents models with 2–3 hour video streams that contain a tutorial video embedded in a stream of arbitrary irrelevant videos, resembling real-world human learning practices. Video-LLMs are asked to apply the acquired skill to answer real-time questions about an ongoing video. Unlike long-video understanding benchmarks that emphasize passive comprehension and skill-learning benchmarks that rely on short, immediate demonstrations, SLVMBench tests the full pipeline of memorizing and extracting procedural knowledge, as well as transferring it to real-time tasks. Moreover, rigorous human annotations feature sub-second-level temporal calibration, manually engineered questions eliminating common-sense guessing, and collated tutorials to ensure coverage of the required skills. Evaluations on state-of-the-art proprietary and open-source video LLMs show that video-LLMs struggle substantially with learning and applying skill knowledge from videos. Moreover, performance degrades markedly when the skill knowledge is placed within a long video memory. These results reveal a key limitation of existing video LLMs and position SLVMBench as the first benchmark for studying real-time skill acquisition and application from long-context video memory.

Neural operators have achieved strong performance in learning solution operators of partial differential equations (PDEs), but their inherently continuous representations struggle to capture discontinuities and sharp transitions. Existing approaches typically approximate such features within continuous function spaces, often requiring increased model capacity and high-resolution data. In this work, we propose Cut-DeepONet, a two-stage training framework that explicitly models discontinuities while reducing learning complexity. Our approach reformulates the problem via a lifting strategy, partitioning the domain into smooth subregions while representing discontinuities as boundaries in a higher-dimensional space. This separation aligns the operator learning task with the inductive bias of neural networks and avoids directly approximating discontinuities. An additional network predicts input-dependent discontinuity locations for unseen inputs, which are then used to guide the neural operator in generating smooth components within each region. Experiments on benchmark PDEs show that Cut-DeepONet outperforms state-of-the-art methods, even when trained on low-resolution datasets. The method excels on problems with discontinuities and sharp transitions, while using fewer trainable parameters. Our results highlight the benefits of changing the representation of operator learning rather than increasing model complexity.

Collaborative learning is sustainable only if it benefits each participant; yet, standard Federated Learning (FL) optimizes a global average that often fails to serve individual clients. In heterogeneous settings, a client may rationally prefer training alone rather than contributing to a global model that targets "average-case" optimality but under performs locally. In this work, we address the problem of \emph{Selfish Personalization (SP)}: How can a target client leverage peer data to minimize its own risk? While current techniques often rely on heuristic performance proxies or clustering that lack sharp theoretical support, we propose \emph{SP-Convergence-Aware Client Weighting (SP-CACW)}. This novel framework determines the contribution of peer clients by explicitly minimizing a theoretical convergence bound for the target. By doing so, our approach efficiently separates useful signal from imported bias on-the-fly during the training process. We provide convergence guarantees that establish the theoretical superiority of \emph{SP-CACW}, alongside empirical results on the MNIST CIFAR datasets and LEAF Shakespeare.


Spectral Re-Basin for Linear Mode Connectivity

Ya-Wei Eileen Lin ⋅ Thomas Dagès ⋅ Daniel Herbst ⋅ Daniel Cremers ⋅ Stefanie Jegelka

Linear mode connectivity is often analysed through exact parameter symmetries, especially hidden-unit permutations that leave the network function unchanged. Yet, many relationships between trained networks are not captured by bijective coordinate relabelings: widths may differ, neurons may split or clone, and explicit parameter symmetries may be broken. In this paper, we introduce spectral re-basin, taking an operator perspective on hidden-unit geometry and studying mode connectivity through the functional geometry of hidden units. Specifically, we represent each hidden layer as a graph of neurons built from layerwise descriptors (weights or activations), associate it with a graph shift operator, and study two coupled objects: layerwise graph functional symmetries, which preserve the relational geometry, and intertwining correspondences across layers from two models, which align their graph structures. We show that permutations arise as a special case of this viewpoint and derive results for non-bijective constructions. Our analysis connects correspondence quality to preserved spectral modes, functional exchangeability of neurons, and bounds on endpoint distortion and linear path barriers. These findings suggest that linear mode connectivity can be governed beyond explicit parameter symmetry by broader spectral operator compatibility in hidden-unit geometry.


SPRING: Solver-guided Process Rewards for Novel Logical Reasoning Steps Generation

Muhammad Asif Ali ⋅ Mohammad Raza ⋅ Wenqing Wang ⋅ Huan Wang

Logical reasoning remains a major challenge for large language models (LLMs), particularly on structured problems that require precise constraint tracking, consistency preservation, and multi-step deduction. This challenge is especially acute for small-scale LLMs, which are more prone to producing inconsistent, redundant, or brittle reasoning trajectories. Existing approaches for improving logical reasoning largely optimize for final-answer correctness, providing only weak supervision over the intermediate reasoning process. In this work, we propose SPRING, a solver-guided reinforcement learning framework for logical reasoning that uses an SMT solver as a training-time verifier of intermediate reasoning steps to provide process-level supervision. SPRING introduces the notion of a novel reasoning step, namely, a step that is logically valid, consistent with the evolving reasoning state, and not already implied by previously accepted non-contradictory deductions. Based on this solver-based assessment, we design process rewards that encourage novel inferential progress while penalizing contradictory and uninformative reasoning steps. Evaluation results on two logical reasoning benchmarks, ZebraLogic and AR-LSAT, show that SPRING consistently outperforms baseline LLMs and outcome-only reward baselines. On ZebraLogic, SPRING improves puzzle accuracy by up to 49.71 and 15.43 points over the base LLM and outcome-only reward baseline, respectively. On AR-LSAT, SPRING improves the overall average score by up to 64.93 and 12.14 points over the base LLM and best-performing outcome-only reward baseline, respectively.

Aligning Large Language Models (LLMs) with human preferences typically relies on external supervision, which faces critical limitations: human annotations are scarce and subjective, reward models are vulnerable to reward hacking, and self-evaluation methods suffer from prompt sensitivity and biases. In this work, we propose stable rank, an intrinsic, annotation-free quality signal derived from model representations. Stable rank measures the effective dimensionality of hidden states by computing the ratio of total variance to dominant-direction variance, capturing quality through how information distributes across representation dimensions. Empirically, stable rank achieves 84.04\% accuracy on RewardBench and improves task accuracy by an average of 14.1\% relative to greedy decoding via Best-of-N sampling. Leveraging this insight, we introduce Stable Rank Group Relative Policy Optimization (SR-GRPO), which uses stable rank as a reward signal for reinforcement learning. Without external supervision, SR-GRPO improves Qwen2.5-1.5B-Instruct by 10\% on STEM and 19\% on mathematical reasoning, outperforming both learned reward models and self-evaluation baselines. Our findings demonstrate that quality signals can be extracted from internal model geometry, offering a path toward scalable alignment without external supervision.


Stabilizing Policy Optimization via Logits Convexity

Hongzhan Chen ⋅ Tao Yang ⋅ Yuhua Zhu ⋅ Shiping Gao ⋅ Xiaojun Quan ⋅ Ting Yao

While reinforcement learning (RL) has been central to the recent success of large language models (LLMs), RL optimization is notoriously unstable, especially when compared to supervised fine-tuning (SFT). In this work, we investigate the stability gap between SFT and RL from a gradient-based perspective, and show that the convexity of the SFT loss with respect to model logits plays a key role in enabling stable training. Our theoretical analysis demonstrates that this property induces favorable gradient directionality during optimization. In contrast, Proximal Policy Optimization (PPO), a widely adopted policy gradient algorithm utilizing a clipped surrogate objective, lacks this stabilizing property. Motivated by this observation, we propose Logits Convex Optimization (LCO), a simple yet effective policy optimization framework that aligns the learned policy with an optimal target derived from the original RL objective, thereby emulating the stabilizing effects of logits-level convexity. Extensive experiments across multiple model families show that our LCO framework consistently improves training stability and outperforms conventional RL methods on a broad range of benchmarks.

Denoising diffusion models have evolved into a state-of-the-art method for tasks in various fields, such as denoising and generation of images, text generation, or generation of synthetic data for training of other machine learning models. First hitting diffusion models (FHDM) are a particular class of denoising diffusion models with \textit{random} adaptive generation time tailored to generate data on a known manifold. Building on the conditioning framework of Doob's $h$-transform these models leverage the given information on the target data manifold to demonstrate strong performance across tasks while offering distinct features such as time-homogeneous dynamics of the generating process and a reduced average simulation time. Even though the theoretical investigation of standard forward-backward diffusion models has attracted much attention in the recent past, the statistical convergence properties of FHDMs are not yet understood. In this work, we show that, up to logarithmic factors, FHDMs achieve the minimax optimal convergence rate in total variation for spherically supported Sobolev smooth data distributions. In particular, this is the first statistical optimality result for denoising diffusion modelling with random generation time.


StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs

Pierre Le Jeune ⋅ Etienne Duchesne ⋅ Weixuan Xiao ⋅ Stefano Palminteri ⋅ Bazire Houssin ⋅ Benoît Malézieux ⋅ Matteo Dora

Multilingual studies of social bias in open-ended LLM generation remain limited: most existing benchmarks are English-centric, template-based, or restricted to recognizing pre-specified stereotypes. We introduce StereoTales, a multilingual dataset and evaluation pipeline for systematically studying the emergence of social bias in open-ended LLM generation. The dataset covers 10 languages and 79 socio-demographic attributes, and comprises over 650k stories generated by 23 recent LLMs, each annotated with the socio-demographic profile of the protagonist across 19 dimensions. From these, we apply statistical tests to identify more than 1{,}500 over-represented associations, which we then rate for harmfulness through both a panel of humans (N = 247) and the same LLMs. We report three main findings. \textbf{(i)} Every model we evaluate emits consequential harmful stereotypes in open-ended generation, regardless of size or capabilities, and these associations are largely shared across providers rather than isolated misbehaviors. \textbf{(ii)} Prompt language strongly shapes which stereotypes appear: rather than transferring as a shared set of biases, harmful associations adapt culturally to the prompt language and amplify bias against locally salient protected groups. \textbf{(iii)} Human and LLM harmfulness judgments are broadly aligned (Spearman $\rho=0.62$), with disagreements concentrating on specific attribute classes rather than specific providers. To support further analyses, we release the evaluation code and the dataset, including model generations, attribute annotations, and harmfulness ratings.


Stop the Sampler! Classifier-Based Adaptive Stopping for Sampling Kernels

Kirill Korolev ⋅ Nikita Morozov ⋅ Stepan Pavlenko ⋅ Esmeralda S Whitammer ⋅ Sergey Samsonov

Sampling from complex, unnormalized probability densities is a fundamental challenge in Bayesian inference and probabilistic modeling. While Markov chain Monte Carlo (MCMC) methods provide asymptotic guarantees, they often suffer from slow mixing and high computational costs due to fixed or manually tuned trajectory lengths. In this work, we propose a novel framework that treats trajectory termination as a learnable component of the sampling dynamics. By framing MCMC within the theory of non-acyclic Generative Flow Networks (GFlowNets), we train state-dependent neural classifiers to decide when a trajectory has reached a high-density region and should terminate. We theoretically establish the connection between optimal classifiers and the target density via detailed balance conditions and introduce a multilevel training scheme to facilitate exploration in complex geometries. Experimental results across various benchmark densities demonstrate that our approach significantly reduces average trajectory lengths while improving mode coverage and mixing compared to standard MCMC baselines.


Subset-Conditioned Boundary Compensation for Missing-Modality Multimodal Classification

Zesen Cai ⋅ Lisi Mo ⋅ Yifan Fang ⋅ Yandong Yan ⋅ Keren Shi ⋅ Daming Shi ⋅ Ruiting Dai

Missing modalities reduce observations to arbitrary subsets, challenging robust multimodal inference. Existing methods recover missing views or align cross-subset representations, implicitly equating feature completeness with decision accuracy—leaving subset-specific decision boundary variation structurally unaddressed. Empirical analysis reveals that misclassification patterns are highly subset-specific, stable across initializations, and concentrated near decision boundaries—a persistent boundary misalignment distinct from information attenuation, which we define as Subset-Conditioned Boundary Shift (SCBS). To address this, we propose Subset-Conditioned Boundary Compensation (SCBC): a modality-factorized violation memory accumulates historical records of subset-specific margin violations, and a subset-adaptive head converts them into targeted logit corrections at inference, with a direction-consistency loss enforcing alignment during training. We prove that these corrections, built entirely from forward-pass statistics without backpropagating through the correction path, satisfy the descent condition of the surrogate loss, to our knowledge the first such guarantee in incomplete multimodal learning. Across five benchmarks, SCBC achieves state-of-the-art-level performance, averaging a 3.4 Macro-F1 point gain under random missingness and peaking at 7.4 points on the IEMOCAP-4 {v, a} setting. Source code is available at https://anonymous.4open.science/r/SCBC_neurips-50B5.


Surprise, Episodic Context, and Catastrophic Forgetting in LLM Fine-Tuning: An Empirical Study

Zied Ben Houidi ⋅ Alexis Huet ⋅ Alberto Eusebio ⋅ Francesco Mantovani ⋅ Dario Rossi

Surprise and episodic context are believed to govern how animals update their memories, with impact on how prior knowledge is retained. Inspired by this parallel, we empirically investigate how surprise and context affect knowledge retention in large language models under supervised fine-tuning. We construct 15 update datasets (230k samples, 13 newly created) spanning surprise levels across facts, ethics, and code, varying how each update is paired with an explicit temporal frame. Across GPT-2-XL, Mistral-7B, Llama-3-8B, and GPT-4.1 variants, evaluated on a held-out cross-domain sentinel set (1.8M LLM judgments), we find: (i) without context, contradictory updates (highest surprise) overwrite the targeted memory and damage entirely unrelated knowledge, sometimes producing near-complete factual collapse, with damage scaling with update surprise, and spilling across domains; (ii) pairing each update with an explicit temporal frame at the prompt (pre-contextualization) appears to create two coexisting traces, preserving the original under the bare prompt while acquiring the new one under the contextualized prompt, and rescuing unrelated knowledge in the process. Extending prior work on emergent misalignment, we characterize a more general pattern we call habit transfer: overwriting a related fact induces transferable response patterns on unrelated prompts (e.g., ``code bleeding'', i.e., answering factual questions with code, rising from 4\% to 73\% along the surprise axis). Taken together, these results suggest explicit context as a shield against catastrophic forgetting in LLMs and point to a practical direction for safer post-training: data-side curation that contextualizes high-surprise updates rather than presenting them as flat prompt-continuation pairs.


TerraMesh-Masks: Open‑Vocabulary Segmentation for Earth Observation

Benedikt Blumenstiel ⋅ Hugues Devimeux ⋅ Johannes Jakubik ⋅ Konrad Schindler

Open‑vocabulary segmentation (OVS) enables models to segment a variety of concepts specified in natural language, removing the need for fixed label sets and fine-tuning. While recent OVS approaches have shown strong performance on natural images, they transfer poorly to Earth observation (EO) data due to differing semantics, scales, sensing modality, and geographic context. To date, large‑scale EO‑specific datasets and benchmarks for OVS remain scarce. We, therefore, introduce TerraMesh-Masks, a large-scale open-vocabulary segmentation dataset for EO, comprising 16 million binary masks aligned with nine million multimodal TerraMesh samples and annotated using OpenStreetMap and Overture-derived semantics. In addition, we provide an expert-reviewed evaluation set of 523 samples covering over 200 classes to enable robust evaluation. To demonstrate and analyze the impact of TerraMesh-Masks, we build an open-vocabulary version of TerraMind. Our results demonstrate significantly improved performance over general-purpose state-of-the-art models reaching 35.15% IoU versus 16.29% for DINOv3.txt and 11.33% for SAM-3. We release the dataset publicly under a permissive license at https://huggingface.co/datasets/ibm-esa-geospatial/TerraMesh-Masks.


The Constitutional Coverage Trilemma in AI Governance

Natalija Mitic ⋅ Soona S O. ⋅ Mamadou Selly Ly ⋅ Moustapha Cisse

Frontier AI systems function as \emph{constitutional institutions}: each deployed model encodes an implicit ranking among safety, helpfulness, honesty, autonomy, and equity. We ask whether the supply of frontier constitutional types covers human demand. Combining a paraphrase-controlled audit of $23$ frontier LLM archetypes with a pairwise-tradeoff study of $1{,}649$ humans on the same instrument, we report three facts. Demand is broad: it spans all five values, with the largest constituency under one-third. Supply is narrow and drifting: the $23$-archetype hull occupies $0.10\%$ of the demand hull, and across six model families autonomy decreases in $5/6$, equity increases in $5/6$, and safety increases in $4/6$. The importance of this drift is not that the frontier moves but that it moves away from a value that is already undercovered, mechanically worsening the welfare floor for the users least well served by the static menu. The fix is sparse: a $2$-vertex menu $\{e_{\mathrm{HON}}, e_{\mathrm{AUT}}\}$ beats the full $23$-archetype frontier by $47$ percent on mean regret ($95$ percent); three vertex additions cut mean regret by up to $81$ percent and worst-group regret by up to $64$ percent. We formalize these findings as a budgeted-pluralism trilemma and show the binding regime is empirically realized.


The Grounding Gap: How LLMs Anchor the Meaning of Abstract Concepts Differently from Humans

Odysseas S Chlapanis ⋅ Orfeas Menis Mastromichalakis ⋅ Christos H Papadimitriou

Abstract concepts - justice, theory, availability - have no single perceivable referent; in the human brain, their meaning emerges from a web of experiences, affect, and social context. Do large language models (LLMs) ground abstract concepts in a similar way? We study this by replicating property-generation experiments from cognitive science on 21 frontier and open-source LLMs. Across models and experiments, we find a consistent pattern: when compared to humans, models rely too heavily on word associations, and underproduce properties tied to emotion and internal states. This yields a large and consistent grounding gap: no model exceeds a Pearson correlation r=0.37 with human responses, compared to a human-to-human ceiling above r=0.9. To better interpret this gap, we also replicate a rating experiment on grounding categories and find that here LLMs align more closely with human judgment, and alignment improves as models get larger. We then use sparse autoencoders (SAEs) to determine whether this information is also reflected in the models' internal features, and identify features connected to grounding dimensions such as "sensorimotor" and "social". These findings suggest that current LLMs can recover grounding dimensions when explicitly queried, but do not recruit them in a human-like way when words are generated freely.


The Long-Run Distribution of Regularized Learning in Non-Concave Games: A Large Deviations Approach

Waïss Azizian ⋅ Pierre-Louis Cauvin ⋅ Franck Iutzeler ⋅ Jérôme Malick ⋅ Panayotis Mertikopoulos

In this paper, we investigate the long-run behavior of discounted regularized learning with stochastic gradient feedback in general, non-concave games. Specifically, we focus on a family of implicitly regularized exponential / multiplicative weight update schemes, and we seek to determine which actions\textemdash or, more generally, which recurrent patterns of play\textemdash are more likely to arise in the long run. We approach this question through the lens of large deviations theory and randomly perturbed dynamical systems, and we obtain a precise characterization of the distribution of the process: in the long run, it follows a Boltzmann-Gibbs law with temperature equal to the method's step-size, and energy levels determined by the game and the statistics of the noise. Concretely, we show that the distribution of play concentrates exponentially around the dynamics' internally chain transitive (ICT) sets - i.e., irreducible invariant sets containing no smaller attractors - and the probability of visiting such a set depends exponentially on its energy. As a result, unstable ICT sets are exponentially less likely to be visited than stable ones, and the sequence of play is exponentially concentrated around the problem's ``ground state'', where energy is minimized. In this manner, learning acts as a selection mechanism: with exponentially high probability, the ground state is the only outcome observed in the long run, even in the presence of multiple equilibria and other attractors.

This position paper argues that the sequential supervised fine-tuning followed by reinforcement learning (SFT-then-RL) pipeline---the de facto paradigm for large language model (LLM) post-training---has a fundamental property that cannot be fixed by hyperparameter tuning: SFT and RL optimize over nearly orthogonal parameter subspaces, so RL inevitably degrades the capabilities acquired during the preceding SFT phase. Under the PL condition, we show that this degradation scales linearly with the number of model parameters. Although the KL penalty can constrain distributional shift, it cannot correct the underlying gradient misalignment between the two objectives. Experiments on Qwen3-0.6B across four tasks confirm the theoretical predictions: on MATH-500, Alpaca, HH-RLHF, and SST-2, the SFT loss increases after the RL phase. We identify the underlying mechanism as persistent near-orthogonality between the gradients of the two objectives ( $\cos(g_\mathrm{SFT}, g_\mathrm{RL}) \approx 0$), paired with opposing curvature dynamics: SFT flattens the loss landscape while RL re-sharpens it. Based on our analysis, we propose three promising research directions: (1) joint optimization of the two objectives with gradient surgery to resolve conflicting updates; (2) landscape-aware scheduling that terminates RL training precisely at the theoretically optimal point; and (3) characterizing the model scale threshold above which SFT becomes unnecessary.


The Ruler and the Judge: Benchmark-Conditional Evaluation of LLM-as-a-Judge

Adrián Ghajari ⋅ Alejandro Benito-Santos ⋅ Víctor Fresno

LLM judges are increasingly used to rank systems and validate claimed improvements, but their validation often reduces to a single question: how well do judge scores correlate with aggregated human ratings? We argue that this question is too fragile. A judge can correlate with human averages while collapsing the rating scale, compressing meaningful margins, or confidently separating systems that humans do not distinguish. Furthermore, a correlation-based workflow always returns a leaderboard, even when the human benchmark cannot support one. We propose a benchmark-conditional measurement framework for LLM-judge validation. Stage 1 audits the human benchmark as a reference ruler using a many-facet ordered Rasch model, estimating rater/item fit, conditional precision, and supported pairwise resolution. Stage 2 calibrates each judge as a linked measurement instrument with judge-specific ordinal thresholds. Stage 3 evaluates out-of-sample system-level fidelity using normalized signed-margin fidelity, false-separation diagnostics, empirical transfer intervals, and distributional EMD. On SummEval and HiTZ-BASSE, the framework yields three outcomes: selection, warning, and refusal, depending on the ruler and judge pool. Calibrated predictions improve observable human-score distribution EMD in 57/81 judge–dimension cases, while the remaining failures are concentrated in the regions flagged by the ruler or judge-instrument audits. Together, these results instantiate benchmark-conditional LLM-judge validation as a protocol for deciding when judge-based claims are warranted, uncertain, or unsupported, addressing a central blind spot of correlation-based judge evaluation.

As high-quality public web corpora become increasingly exhausted, clean long-context documents have become a scarce and expensive source of training data for large language models (LLMs). Existing long-context corpora are often proprietary and costly to acquire, synthetically generated, or concentrated in narrow domains such as programming. We introduce the Stanford EDGAR Filings Dataset (SFD), an open reconstruction of SEC filings into layout-faithful MultiMarkdown for financial language modeling and evaluation. SFD makes audited financial statements, risk disclosures, ownership reports, accounting notes, and market-moving event filings usable as long-context pretraining data and as a basis for financial reasoning, forecasting, compliance, and document understanding. The resulting corpus is token-efficient, model-ready, and has less than 0.1\% overlap with Common Crawl-derived corpora. We release SFD-v1, a 152B-token snapshot of a larger 18.4M-filing corpus estimated at 500B tokens. We further introduce two SFD-derived benchmarks: EDGAR-Forecast, which evaluates filing-grounded numerical forecasting after model knowledge cutoffs, and EDGAR-OCR, which evaluates transcription of complex financial tables.


The Smart Buildings Control Suite: A Diverse Open Source Benchmark to Evaluate and Scale HVAC Control Policies for Sustainability

Judah Goldfeder ⋅ Victoria Dean ⋅ Zixin Jiang ⋅ Xuezheng Wang ⋅ Bing Dong ⋅ Hod Lipson ⋅ John Sipple

Commercial buildings account for 17% of U.S. carbon emissions, with roughly half of that from Heating, Ventilation, and Air Conditioning (HVAC). HVAC devices form a complex thermodynamic system, and while model predictive control and reinforcement learning have been used to optimize control policies, scaling to thousands of buildings remains a significant unsolved challenge. Most current approaches are over-optimized for specific buildings and rely on proprietary data or hard-to-configure simulators. We present the Smart Buildings Control Suite, the first open source interactive HVAC control benchmark with a focus on solutions that generalize across building. It has 3 components: real-world data from 11 buildings over 6 years, a lightweight data-driven simulator for each building, and a modular Physically Informed Neural Network (PINN) building model as a simulator alternative. The buildings span multiple climates, management systems, and sizes, and both the simulator and PINN easily transfer to new buildings, ensuring solutions using this benchmark are robust to these factors and only reliant on fully scalable building models. This represents a major step towards scaling HVAC optimization from the lab to buildings everywhere. To facilitate use, our benchmark is compatible with the Gym standard, and our data is part of TensorFlow Datasets.


To Align or Not To Align: Check Your COMPASS Before You Train

Fabian Gröger ⋅ Petar Damjanović ⋅ Maria Brbic

Combining two pretrained models, whether by aligning a vision encoder with a language model or by merging two fine-tuned checkpoints, is a costly bottleneck: compatibility is typically known only after joint training. We propose COMPASS (COMPatibility ASsessment Score), a theoretically grounded compatibility diagnostic that estimates whether a model pair will be compatible prior to training or merging. Building on the observation that representational convergence is most robust at the level of local neighborhoods, COMPASS decomposes model compatibility into two cheap-to-compute axes: a pair-specific ease, defined as the $k$-NN overlap between the two encoders, and a paradigm-level ceiling, capturing the maximum overlap their training paradigms allow. Together, these two axes provide a ranking of candidate model pairs and prescribe how to act on each pair: (i) prioritize pairs that already agree and whose paradigms are compatible, (ii) invest additional compute on pairs with compatible paradigms but low initial agreement, (iii) switch the pre-training paradigm when it cannot support sufficient agreement, or (iv) deprioritize pairs that score poorly on both axes. COMPASS is applicable to both cross-modal alignment and task-arithmetic merging, where the encoder pair is replaced by two fine-tuned checkpoints. Across $72$ cross-modal alignment experiments, and $392$ model merges, our compatibility estimates strongly correlate with downstream success, with Pearson correlations of $0.90$ for image-to-text R@1 on Flickr30k, $0.91$ for image-to-text R@1 on COCO, and $0.70$ for mean retention across $8$ image-classification merging tasks.


Toward Optimal Regret in Robust Pricing: Decoupling Corruption and Time

Kalana K Kalupahana ⋅ Francesco Emanuele Stradi ⋅ Matteo Castiglioni ⋅ Alberto Marchesi

We design the first regret guarantees for robust dynamic pricing which decouples the dependence on the corruption $C$ and the time horizon $T$. In dynamic pricing, a seller with unlimited supply of a good interacts with a stream of buyers over $T$ rounds, with the goal of maximizing revenue. At each round $t$, the seller posts a price $p_t$, and the buyer purchases the good only if their unknown valuation $v^\star$ exceeds this price. The seller observes only the binary feedback $\mathbb{I}[p_t \leq v^\star]$, indicating whether a sale occurred. In the robust pricing setting, a malicious adversary is allowed to corrupt this feedback in at most $C$ rounds. Even if the learner knows the corruption $C$, the best known regret bound is $\mathcal{O}(C\log\log T)$ by Gupta et al. [2025]. They left as an open problem to "decouple'' the dependence on $C$ and $T$. In this work, we resolve this open problem. In particular, we develop a robust variant of binary search that achieves regret $\mathcal{O}(C+\log T)$ when the corruption $C$ is known and $\mathcal{O}(C+\log^2 T)$ when the corruption is unknown.


Towards Convergence of PPO: An Approximate Descent Approach

Leif Döring ⋅ Daniel A Schmidt ⋅ Moritz Melcher ⋅ Sebastian Kassing ⋅ Benedikt Wille ⋅ Tilman Aach ⋅ Simon Weissmann

Proximal Policy Optimization (PPO) and it's variants are the most widely used policy-gradient methods in reinforcement learning, yet the role of its actor update mechanism is still poorly understood theoretically. In particular, standard PPO combines clipped surrogate gradients, multiple epochs of minibatch updates, and reuse of rollout data, but existing convergence analyses do not fully capture this update structure. In this work, we theoretically study PPO policy updates with symmetric clipping from a policy-gradient perspective and interpret its actor updates as a cyclic biased-gradient method with sample reuse and random reshuffling. Our first contribution is a clean formalization of PPO clipped actor updates through surrogate gradients that approximate the true policy gradient. Using the performance difference lemma, we prove a linear bias bound which quantifies how the surrogate gradient drifts as the policy moves away from the sampling policy. Our second contribution is a convergence analysis of cyclic surrogate-gradient ascent, showing that additional PPO-style biased updates can improve progress under conservative learning rates without requiring additional samples. Finally, we analyze the stochastic minibatch version with reshuffling and obtain convergence-to-stationarity guarantees under standard smoothness assumptions and a bounded critic-bias condition. Overall, our results provide a theoretical interpretation of PPO’s multi-epoch actor updates: extra clipped surrogate steps introduce bias, but can still improve optimization efficiency by compensating for small, stable step sizes through sample reuse.


Training Data Attribution in Diffusion Models via Mirrored Unlearning and Noise-Consistent Skew

Joan Serrà ⋅ Dipam Goswami ⋅ Fabio Morreale ⋅ Wei-Hsiang Liao ⋅ Yuki Mitsufuji

Training data attribution (TDA) should enable generative model interpretability and foster a variety of related downstream tasks. Nonetheless, current TDA approaches lack reliability and robustness, preventing their adoption in real-world setups. In this paper, we take a decisive step towards more reliable and robust TDA for diffusion models. We propose to perform TDA with mirrored unlearning and noise-consistent skew (MUCS). The idea is to fine-tune a second model with bounded mirrored gradient ascent, and to measure the normalized skew of this model with respect to the original one using consistent noise samples. We show that, while being conceptually simple and generic, MUCS systematically outperforms existing methods on three different datasets by a large margin. We additionally study the effect that core design choices have on final performance, and analyze novel aspects regarding the overlap of influential instances across generated items and the potential of ensembling TDA approaches. We believe that our findings may have broader implications for more general unlearning setups, as well as for tasks requiring the comparison of diffusion losses.

Video large language models are often dominated by visual tokens, which lengthen prefill, enlarge KV-cache residency, and raise peak memory. In spirit of the philosophy $\textit{focus more and memorize less}$, we introduce Focus--Ambient Retention (FAR), a training-free visual token retention method that routes visual evidence through two complementary streams. The Focus stream preserves task-critical objects, actions, text, and fine details, while the Ambient stream keeps scene context only when it differs from a temporal cache. FAR combines visual attention, local frequency variation, and bounded query relevance to score tokens, then assembles a fixed-budget context with diversity control before language-model prefill. Across five video understanding benchmarks under a variety of model sizes, FAR significantly reduces peak VRAM by up to $50.6\%$ under matched retained-token budgets while achieving comparable or even superior performance over the state-of-the-art, improving the quality-memory trade-off. Overall, FAR offers a position-aware alternative to frame- or block-level compression for efficient inference. Code will be available.


TropNNC: Structured Neural Network Compression Using Tropical Geometry

Konstantinos Fotopoulos ⋅ Petros Maragos ⋅ Panagiotis Misiakos

We present TropNNC, a framework for compressing neural networks with linear and convolutional layers and ReLU-type activations using tropical geometry. By representing a network’s output as a tropical rational function, TropNNC enables structured compression via reduction of the corresponding tropical polynomials. Our method identifies redundancy via similarity and improves upon the geometric approximation of previous work by adaptively selecting the weights of retained neurons. We relate it to SVD and spectral clustering, and provide insights into network compression beyond the specific setting considered. We provide the tightest known theoretical compression bound, and the first successful application of tropical geometry to convolutional layers. TropNNC requires access only to network weights – no training data – and achieves competitive performance on MNIST, CIFAR, and ImageNet, matching strong baselines such as ThiNet and CUP.

Foundation models mark a profound paradigm shift in time series modeling, with task-specific models being superseded by general-purpose zero-shot models. Yet, current approaches primarily focus on forecasting, while real-world time series are often irregularly and partially observed, requiring models that can jointly forecast, impute missing values, and handle degraded sampling conditions. To address these challenges, we introduce TS-ICL, a novel probabilistic In-Context Learning encoder-regressor Transformer that unifies forecasting and imputation. TS-ICL formulates time series tasks as timestamp-aligned regression and naturally incorporates covariates by training on synthetic dependency structures generated from a novel causal data prior. Empirically, TS-ICL achieves a new state-of-the-art in imputation, while remaining competitive with leading forecasting foundation models across both univariate and covariate-aware benchmarks. It shows particularly strong performance in forecasting with partially observed look-back windows. Code page: https://anonymous.4open.science/r/tsicl-D3F1/.


TUBE: Tangent Upper Bound on Evidence for Discrete Diffusion Language Models

Arseny Ivanov ⋅ Sergei Kholkin ⋅ Vladislav Gromadskii ⋅ Grigoriy Ksenofontov ⋅ Ivan Oseledets ⋅ Aleksandr Korotin

Log-likelihood is a standard metric for evaluating generative models. Unfortunately, in contrast to autoregressive models (ARMs), discrete diffusion models generally do not admit exact computation of this quantity. Existing evaluations, therefore, rely on the evidence lower bound (ELBO), leaving unclear how much higher the true value may be. We address this by introducing the Tangent Upper Bound on Evidence (TUBE), a variational upper bound on log-likelihood that admits an unbiased Monte Carlo estimator. Our TUBE extends across latent-variable models, including masked diffusion models (MDMs), any-order ARMs (AO-ARMs), and block variants of both. Applied to block MDMs and block AO-ARMs, TUBE reveals our key empirical finding that these models lie strictly below the exact ARM baseline, showing that ARMs still dominate in likelihood.


Unbiased First-Order Randomized Smoothing for Differentiable Simulation

Mathis SCHEFFLER ⋅ Wilson Jallet ⋅ Cordelia Schmid ⋅ Justin Carpentier

Computing informative gradients through nonsmooth physical models—such as those involving hard contact or bouncing—is a cornerstone challenge in robotics and machine learning. Randomized smoothing offers a principled way to obtain well-behaved surrogate gradients, yet standard first-order estimators of these smoothed gradients are biased whenever the underlying function is discontinuous. In this work, we show that this bias arises from neglected discontinuity contributions and derive a corrected first-order formula that eliminates it. When the discontinuity structure is known, we exploit it to build discontinuity-aware gradient estimators, which we instantiate for differentiable collision detection on 3D meshes. When it is unknown, we propose a complementary estimator that leverages classical gradients to reduce variance. We validate this estimator on nonsmooth trajectory optimization instances, demonstrating variance reductions of several orders of magnitude over zeroth-order baselines while remaining unbiased under discontinuities where standard first-order estimators fail.


Unbounded Streaming Text-To-Speech with Prefixed Sliding Window Attention

Théodor Lemerle ⋅ Diego Torres Guarin ⋅ Téo Guichoux ⋅ Nicolas Obin ⋅ Axel Roebel

Existing text-to-speech (TTS) systems achieve high quality, yet streaming and long-form generation remain challenging. We propose simple adaptations to a standard encoder–decoder TTS architecture that enables both capabilities without limits on total duration. We first observe that attention patterns in these models are well organized, following a structure that motivates a specific windowing strategy: a prefixed sliding window for autoregressive decoding and a sliding cross-attention for text conditioning. Together, we show that these adaptations address length generalization and error accumulation while being fully streamable in text and audio. Trained exclusively on segments shorter than 30 seconds, our system produces seamless, consistent speech at practically unbounded lengths with linear decoding complexity, outperforming baselines that rely on hour-long training context. We validate the approach through continuous synthesis over hours long synthesis and demonstrate competitive results against state-of-the-art open-source baselines on both short- and long-form benchmarks.

Classical rate-distortion theory minimises expected distortion at a given coding rate but does not constrain the \emph{uncertainty} of the reconstructed signal at the decoder. We formalise predictive uncertainty $U$ as the conditional variance of the reconstruction given the source, and show that it is an underconstrained degree of freedom of RD-optimal compression with a sharp empirical signature: at matched rate and matched scalar $U$, two reconstructions of the same image can differ by $40$+ mIoU points in downstream segmentation depending on \emph{where} the variance is spatially allocated, with $88\%$ of (image, quality) pairs showing a significantly negative within-image slope of mIoU on $U$. We derive a closed-form Uncertainty-Rate-Distortion surface and prove a coding-theoretic converse for additive-noise channels with linear regression under Gaussian assumption. A region-importance spatial-URD functional also converts the spatial-allocation finding into a within-image predictor that generalises across two segmentation architectures, three rates, two codecs, and a different segmentation foundation model on a different dataset. Cross-task and cross-modality probes (BLIP captioning, CLIP zero-shot, NYU depth, LibriSpeech speech recognition) characterise the framework's reach: the spatial-URD prediction transfers cleanly to locally-aggregating downstream models on both image and audio modalities, and degenerates predictably for globally-aggregating ones. The validation code and a demo for a diffusion task is provided as Supp. Mat., and will be released alongside with all the code to reproduce the experiments upon acceptance of the work.


Uncovering Challenges of Solving the Continuous Gromov-Wasserstein Problem

Xavier Aramayo-Carrasco ⋅ Maksim Nekrashevich ⋅ Petr Mokrov ⋅ Evgeny Burnaev ⋅ Aleksandr Korotin

Recently, the Gromov-Wasserstein Optimal Transport (GWOT) problem has attracted the special attention of the ML community. In this problem, given two distributions supported on two (possibly different) spaces, one has to find the most isometric map between them. In the discrete variant of GWOT, the task is to learn an assignment between given discrete sets of points. In the more advanced continuous formulation, one aims at recovering a parametric mapping between unknown continuous distributions based on i.i.d. samples derived from them. The clear geometrical intuition behind the GWOT makes it a natural choice for several practical use cases, giving rise to a number of proposed solvers. Some of them claim to solve the continuous version of the problem. At the same time, GWOT is notoriously hard, both theoretically and numerically. Moreover, all existing continuous GWOT solvers still heavily rely on discrete techniques. Natural questions arise: to what extent existing methods unravel GWOT problem, what difficulties they encounter, and under which conditions they are successful. Our benchmark paper is an attempt to answer these questions. We specifically focus on the continuous GWOT as the most interesting and debatable setup. We crash-test existing continuous GWOT approaches on different scenarios, carefully record and analyze the obtained results, and identify issues. Our findings experimentally testify that the scientific community is still missing a reliable continuous GWOT solver, which necessitates further research efforts. As the first step in this direction, we propose a new continuous GWOT method which does not rely on discrete techniques and partially solves some of the problems of the competitors.


Unifying Goal-Conditioned RL and Unsupervised Skill Learning via Control-Maximization

Alireza Modirshanechi ⋅ Benjamin Eysenbach ⋅ Peter Dayan ⋅ Eric Schulz

Unsupervised pretraining has driven empirical advances in goal-conditioned reinforcement learning (GCRL), but its theoretical foundations remain poorly understood. In particular, an influential class of methods, mutual information skill learning (MISL), discovers behaviorally diverse skills that can later be used for downstream goal-reaching. However, it remains a theoretical mystery why skills learned through MISL should support goal-reaching. A subtle challenge is that both GCRL and MISL are umbrella terms: different GCRL tasks use distinct criteria for measuring goal-reaching performance, while different MISL methods optimize distinct notions of behavioral diversity. We address this challenge and unify GCRL and MISL as instances of control maximization. We identify three canonical GCRL formulations and prove that they are fundamentally inequivalent: they can induce incompatible optimal policies even in the same environment. Nevertheless, they all share a common interpretation: a well-performing goal-conditioned policy is one whose future trajectory is highly sensitive to the commanded goal, with the precise notion of sensitivity determined by the GCRL formulation. Noting that MISL objectives can be understood as measures of skill-sensitivity akin to goal-sensitivity, we show that MISL objectives are bounded by formulation-specific downstream goal-sensitivities. These bounds establish a precise correspondence between MISL methods and downstream GCRL tasks: for every GCRL formulation, there exists a matching MISL objective for which more diverse skills afford greater downstream goal sensitivity. Our results thus lay a theoretical foundation for RL pretraining and have important practical implications, such as suggesting which pretraining objectives to use when a user cares about a specific class of downstream tasks.

Denoising models such as Diffusion or Flow Matching have recently advanced generative modeling for discrete structures, yet most approaches operate directly in the discrete state space, leading to abrupt state changes. We introduce simplex denoising, a simple yet effective generative framework that operates on the probability simplex. The key idea is a non-Markovian noising scheme in which, for a given clean data point, noisy representations at different times are conditionally independent. While preserving the theoretical guarantees of denoising-based generative models, our method removes unnecessary constraints, thereby improving performance and simplifying the formulation. Empirically, \emph{unrestrained simplex denoising} surpasses strong discrete diffusion and flow-matching baselines across synthetic and real-world graph benchmarks.


V-GIFT: Boosting Visual Instruction Tuning with Self-Supervised Guidance

Sophia Sirko-Galouchenko ⋅ Monika Wysoczańska ⋅ Andrei Bursuc ⋅ Nicolas THOME ⋅ Spyridon Gidaris

Multimodal large language models (MLLMs) perform well on many vision-language tasks but often struggle with vision-centric problems that require fine-grained visual perception. Recent evidence suggests that this limitation arises not from weak visual representations, but from under-utilization of visual information during instruction tuning, where many tasks can be partially solved using language priors alone. We propose a simple and lightweight approach that augments visual instruction tuning with a small number of visually grounded self-supervised tasks expressed as natural language instructions. By reformulating classical self-supervised pretext tasks, such as rotation prediction, color matching, and cross-view correspondence, as image–instruction–response triplets, we introduce supervision that cannot be solved without relying on visual evidence. Our approach requires no human annotations, no architectural modifications, and no additional training stages. Across multiple models, training regimes, and benchmarks, injecting only a small fraction (3–10%) of such visually grounded instructions consistently improves performance on vision-centric evaluations. Our findings highlight instruction tuning with visually grounded SSL tasks as a powerful lever for improving visual perception in MLLMs through simple adjustments to the training data distribution.


VISTA: Support-Anchored Value Targeting for Fast Flow-Based Vision-Language-Action Policies

Hongjie Cao ⋅ Yuxuan Yang ⋅ Yunpeng Mei ⋅ Peng Cheng ⋅ Chenyu Wang ⋅ Jiamin Wang ⋅ Xiaoyi Fan ⋅ Fang Deng ⋅ Gao Huang ⋅ Jie Chen ⋅ Gang Wang

Flow-based vision-language-action (VLA) policies provide a compelling recipe for robot control: they inherit broad behavioral priors from large-scale imitation and generate smooth, temporally extended continuous action chunks. The same properties, however, make offline post-training delicate. Mixed offline robot datasets contain expert behavior as well as partial completions and failures, requiring improvement without drifting far from the data support. An offline critic can in principle provide useful directions, but actor-side critic optimization may suffer from extrapolation and push high-dimensional action chunks off-support. In contrast, advantage-conditioned or reweighting-based methods are more stable, but tend to be conservative and restricted to actions already present in the dataset. We propose VISTA, a post-training framework that turns value estimates into support-anchored distillation targets for pretrained flow policies. VISTA queries a critic only at dataset actions, computes a normalized action-gradient direction, and uses this bounded direction to shift clean-action endpoints and velocity targets. A frozen flow teacher preserves the pretrained action prior, while a behavior-cloning anchor keeps the adapted targets close to the data distribution. The shaped supervision is distilled into a time-indexed action denoising (TAD) student, which predicts time-indexed clean-action estimates with a single network evaluation and obtains the executed action by an analytic final readout. We treat the induced target shift as conservative local guidance rather than a global policy-improvement guarantee. Across five BEHAVIOR-1K simulation tasks and five real-world bimanual manipulation tasks, VISTA achieves the best average success rate in our evaluation against behavior cloning (BC), advantage-conditioned BC, IDQL, and flow Q-learning, while reducing action-expert computation by $15\times$ (and end-to-end policy inference by $1.8\times$) relative to the $20$-step flow teacher.


WayPOP: A Panoramic Open-Set Panoptic Tracking Benchmark

Rohit Mohan ⋅ Maximilian Luz ⋅ Swarnava Chowdhury ⋅ Yao Lu ⋅ Florian Drews ⋅ Thomas Nürnberg ⋅ Yakov Miron ⋅ Federico Tombari ⋅ Stefano Gasperini ⋅ Abhinav Valada

Autonomous driving perception increasingly relies on panoramic multi-camera setups for holistic, spatiotemporal scene understanding. However, existing panoptic tracking benchmarks assume a fixed semantic taxonomy, failing to evaluate whether systems can robustly detect, segment, and track out-of-distribution (OOD) objects across time and multiple viewpoints. Creating a realistic open-set benchmark for this setting is challenging: fully synthetic data suffers from domain gaps, while 2D augmentations lack multi-view and temporal geometric consistency. To address this gap, we introduce WayPOP, the first benchmark for multi-view open-set panoptic tracking (MV-OSPT) in autonomous driving. WayPOP augments real-world sequences from the Waymo PVPS dataset with diverse 3D assets using a geometry-grounded pipeline. By leveraging actual LiDAR geometry, out-of-taxonomy objects are seamlessly integrated with realistic scales, multi-view projections, and depth-aware occlusions across five synchronized cameras and time. We formalize a structured task hierarchy for MV-OSPT and adapt established metrics to separately evaluate known and unknown regions for both segmentation and tracking. Our comprehensive benchmarking of representative baselines reveals that maintaining consistent unknown-object identities during camera transitions, viewpoint changes, and temporal occlusions remains a critical open challenge.


Weighted Sampling for Online Causal Discovery

Arnab Bhattacharyya ⋅ Philips George John ⋅ Sayantan Sen ⋅ Naganand Yadati

Discovering the underlying causal structure of a system is a fundamental challenge in machine learning, requiring interventional data to distinguish between observationally equivalent models. In this paper, we study the problem of learning causal Bayesian networks in an online setting, where a learner sequentially observes individual samples from a stream of unknown interventional and observational distributions. We formulate this task as an online sequence prediction problem. To overcome the super-exponential size of the DAG search space, we extend dynamic programming algorithms originally developed for uniform DAG counting and sampling within Markov Equivalence Classes to support score-decomposable, weighted sampling. We prove that the posterior distribution maintained by our algorithm competes with the optimal causal structure in hindsight, outputting a distribution that is close to the true data-generating mechanism as measured by the Interventional Kullback-Leibler divergence.


What’s Holding Back Latent Visual Reasoning?

Andre G Viveiros ⋅ Nuno Gonçalves ⋅ André Martins ⋅ Matthias Lindemann

Humans can approach complex visual problems by mentally simulating intermediate visual steps, rather than reasoning through language alone. Inspired by this, several works on Vision-Language Models have recently explored chain-of-thought reasoning with continuous latent tokens as intermediate visual imagination steps. In this work, we investigate how recent models leverage such latent tokens. Surprisingly, we find that model accuracy is unaffected when latent tokens are replaced by uninformative dummy tokens. This indicates that latent tokens play a minimal causal role in the model's final prediction. To better understand this phenomenon, we analyze both the training signal provided by oracle latent representations and the quality of the latent tokens generated at inference time. Our experiments reveal two crucial issues holding back latent visual reasoning: First, in most existing datasets, oracle latent tokens provide limited additional information beyond the original image and do not substantially simplify the task, leading models to ignore them during training and effectively bypassing them at inference time. When fine-tuned on a diagnostic dataset, in which latent tokens provide sufficient support for the final prediction, we show that models can causally rely on them. Second, the latent tokens produced at inference time deviate from their corresponding oracle representations, collapsing to a narrow region and preventing benefits even when the model relies on them. Overall, our findings suggest that future progress in latent visual reasoning depends on two key pillars: high-quality datasets with informative intermediate steps and more precise latent token prediction.

Trimming suspicious calibration points is a common response to contamination in conformal prediction. Its effect on clean-target coverage, however, is governed by the retained law induced by trimming, not by the contamination level alone. We analyse fixed-threshold trimming as conditioning rather than purification. It replaces the contaminated calibration law with a retained law, reducing clean-target coverage to a one-dimensional score-CDF transfer problem with an exact finite-sample identity. A componentwise bound on the transfer gap gives a population-level diagnostic. This separates a clean-side covariance cost from a retained-contamination cost, governed by the dirty-to-clean retention ratio. Trimming helps when the anomaly score separates retention probabilities while remaining score-neutral on the clean population. Otherwise, it cannot substantially reduce contamination through the retained mixture coefficient. We also give finite-sample certificate templates that provide numerical guarantees under independent audit.


When Must AI Training Stage Checkpoints? A Distributional Model of Durability Boundaries at Scale

Farid Talibli ⋅ Awais Khan ⋅ Christopher Zimmer ⋅ Sungyong Park ⋅ Jihoon Yang ⋅ Youngjae Kim

Large-scale AI training runs span days within fixed HPC allocations, making checkpointing essential for bounded recovery after failure. Asynchronous checkpointing hides checkpoint latency from the training loop, but does not guarantee that the most recent state has reached a durable target: data may still be in flight to the shared parallel filesystem (PFS) when a failure occurs, forcing rollback to an older checkpoint. We formalize this durability gap through a distributional model that combines hardware crossover analysis with statistical models of PFS completion time under scale and contention. Rank-local completion times follow a lognormal body distribution, while wall-clock completion is governed by an extreme-value process determined by the slowest writer. The resulting analysis yields a contention-dependent threshold,~$C^\star_p(n)$, beyond which the PFS is no longer the fastest path to durable persistence. Guided by this analysis, we design CacheX, a transparent staging layer that composes with PyTorch Distributed Checkpoint~(DCP). CacheX stages bulk data to local DRAM and replicates it to a neighboring node's NVMe, providing allocation-local durability while preserving eventual persistence to the PFS. On Frontier, across IOR runs up to 256~nodes and end-to-end FSDP+DCP training up to 64~nodes, CacheX reduces write-time variability by up to~$43\times$, achieves a replica-durable effective throughput of~${\sim}\,2.5$ to $2.8$\,GiB/s per node decoupled from PFS contention, and reduces worst-case rollback exposure. Our results show that staging is not a universal accelerator; rather, it is a conditional policy for selecting the first durable checkpoint boundary under measured scale, load, and tail behavior.


Who Needs Labels? Adapting Vision Foundation Models With the Metadata You Already Have

Elouan Gardes ⋅ Seung-Eun Yi ⋅ Huy V Vo ⋅ Théo Moutakanni ⋅ Kartik Ahuja ⋅ Piotr Bojanowski ⋅ Wolfgang Pernice ⋅ Loic Landrieu ⋅ Camille Couprie

We propose a label-free approach to adapt powerful but generic vision foundation models to specialized scientific domains. Standard supervised fine-tuning is often ill-suited to these settings: labels are scarce, and task-specific training can collapse the model's generality and hurt robustness. We instead leverage metadata to adapt representations to new domains in a self-supervised manner. Our method, META-DINO, combines a standard self-supervised objective with flexible metadata guidance that handles both highly granular discrete metadata and continuous metadata. It encourages the representation to preserve informative factors while suppressing spurious ones. Across subcellular fluorescence microscopy, Earth observation, wildlife monitoring, and medical imaging, META-DINO consistently outperforms standard unsupervised domain adaptation and fully supervised adaptation. It also exceeds highly-specialized domain-specific state of the art, while using no task labels for backbone adaptation and only lightweight probes for supervision.


X-Palm: Paired Multispectral-to-Smartphone Dataset for Cross-Domain Palmprint Authentication

Seyed Jamal Seyedmohammadi ⋅ Pai Chet Ng ⋅ Angelo Genovese ⋅ Zhixiang Chi ⋅ Jeannie Lee ⋅ Konstantinos N Plataniotis

Palmprint modality offers a privacy-preserving biometric solution, yet its deployment is hindered by the domain gap between controlled enrollment and unconstrained authentication. Existing datasets are largely restricted to controlled setups and fail to capture the compound variability of real-world environments. In this paper, we introduce X-Palm, a cross-domain dataset comprising 6,006 palm images from 103 individuals (206 hands). To the best of our knowledge, X-Palm is the first palmprint dataset providing novel paired-identity acquisition specifically designed to bridge the gap between reliably controlled multispectral enrollment and unconstrained mobile authentication while encompassing a broad spectrum of in-the-wild variability. Unlike existing datasets that focus on single to a few variations, X-Palm addresses the massive modality and environmental shifts encountered in practical deployments by capturing paired data for identities across two distinct domains: (1) a controlled Multispectral Palmprint setting using our custom-developed scanner, and (2) an unconstrained smartphone palmprint setting that is participant-driven, incorporating simultaneous variations in hardware, hand pose, illumination, background, camera-to-hand distance, perspective, and palm surface conditions (e.g., moisture and occlusions). Our extensive benchmarks of 12 SOTA models reveal that while existing methods achieve high performance on controlled data, they experience severe performance collapse on X-Palm. Conversely, models trained on X-Palm demonstrate consistent robustness across domains, positioning X-Palm as a valuable resource for training a model towards real-world, cross-domain generalization. Data access instructions and the related benchmarking codes are publicly available at: https://github.com/X-Palm/X-Palm-2026


Your Self-Supervised Projection Head Captures Object Co-Occurrence Statistics

Arthur Aubret ⋅ Jochen Triesch ⋅ Céline Teulière

Self-Supervised Learning (SSL) has achieved impressive success in learning semantic visual representations, yet the underlying principles driving this success remain underexplored. In this work, we hypothesize that common SSL pretext tasks implicitly model object co-occurrence statistics, a fundamental cue for visual learning. To test this, we curate three datasets of segmented objects from existing vision benchmarks using a state-of-the-art segmentation model. Through experiments across many SSL models, we reveal a hierarchical encoding of semantic information: while the visual backbone captures object categories, the projection head specializes in encoding co-occurrences between object categories. As a result, the projection head outperforms the backbone in aligning with human judgments of inter-category similarity. Furthermore, by controlling co-occurrence patterns during pre-training, we demonstrate that encoding object co-occurrences can significantly accelerate the emergence of category-level representations. Our findings uncover a previously hidden learning principle in SSL and suggest a path toward designing more effective pretext tasks by explicitly leveraging object co-occurrence structure.