ICML 2026 Papers with Code & Data
To facilitate rapid community engagement with the presented research, we have compiled an extensive index of accepted papers that have associated public code or data repositories. We list all of them in the following table. This index was generated using an automated extraction process. While we strive for completeness, some papers with public resources may have been missed. Please inform us if you discover any additional papers that should be included. Readers should be aware that some code repositories may not be made fully public until the conference officially begins.
In addition to this index, we encourage readers to explore our related resources: ICML-2026 papers & highlights: For curated summaries and key takeaways from this year’s conference. “Best Paper” Digest (ICML): A historical overview of the most influential ICML papers published since 2004.
Since 2018, Paper Digest has built a foundation of data spanning decades of conferences, journals, and research topics. The platform features a daily digest service that sifts through tens of thousands of new papers, clinical trials, news articles, and community posts, filtering the noise to highlight what matters most to specific interests. Beyond daily updates, dozens of built-in research tools streamline the academic workflow, supporting efficient reading and writing, comprehensive literature reviews, and automated research report generation.
Paper Digest Team
New York City, New York, 10017
team@paperdigest.org
TABLE 1: ICML 2026 Papers with Code & Data
| Paper | Author(s) | Code | |
|---|---|---|---|
| 1 | Unified Multimodal Autoregressive Modeling with Shared Context—Visual Tokenizer Is Key to Unification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing approaches typically rely on two disparate visual tokenizers, which splits the representation space and hinder truly unified modeling. We propose UniAR, a unified autoregressive framework where a single discrete visual tokenizer serves as the key bridge between understanding and generation, enabling a shared context in which the model can directly interpret its own generated visual tokens without additional re-encoding. |
Wujian Peng; Lingchen Meng; Yuxuan Cai; Xianwei Zhuang; Yuhuan Yang; Rongyao Fang; Chenfei Wu; Junyang Lin; Zuxuan Wu; Shuai Bai; | code |
| 2 | Set Diffusion: Interpolating Token Orderings Between Autoregression and Diffusion for Fast and Flexible Decoding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our key insight is that interpolating generation orderings between autoregression and fully-random decoding, rather than committing to a fixed block length, offers a better interpolation between diffusion and AR. |
Marianne Arriola; Volodymyr Kuleshov; | code |
| 3 | $\tau^2$-Bench: Evaluating Conversational Agents in A Dual-Control Environment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This differs from real-world scenarios like technical support, where users need to actively participate in modifying the state of the (shared) world. In order to address this gap, we introduce $\tau^2$-bench, with four key contributions: 1. A novel **Telecom dual-control domain** modeled as a Dec-POMDP, where both agent and user make use of tools to act in a shared, dynamic environment that tests both agent coordination and communication, 2. |
Victor Barres; Honghua Dong; Soham Ray; Xujie Si; Karthik Narasimhan; | code |
| 4 | PlotCraft: Pushing The Limits of LLMs for Complex and Interactive Data Visualization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, their ability to create complex visualizations for scaled and structured data remains largely unevaluated and underdeveloped. To address this gap, we introduce **PlotCraft**, a new benchmark featuring 1k challenging visualization tasks that cover a wide range of topics, such as finance, scientific research, and sociology. |
Jiajun Zhang; Jianke Zhang; Zeyu Cui; Jiaxi Yang; Lei Zhang; Zilei Wang; Qiang Liu; Liang Wang; Binyuan Hui; Junyang Lin; | code |
| 5 | D2: Improved Techniques for Training Reasoning Diffusion Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Here, we introduce d2, a reasoning framework tailored for masked DLMs. |
Guanghan Wang; Gilad Turok; Yair Schiff; Marianne Arriola; Volodymyr Kuleshov; | code |
| 6 | VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The goal of this paper is to improve the performance and reliability of vision-language-action (VLA) models through iterative online interaction. |
Yanjiang Guo; Tony Lee; Lucy Xiaoyang Shi; Jianyu Chen; Percy Liang; Chelsea Finn; | code |
| 7 | Solving Physics Olympiad Via Reinforcement Learning on Physics Simulators Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we show that physics simulators can serve as a powerful alternative source of supervision for training LLMs for physical reasoning. |
Mihir Prabhudesai; Aryan Satpathy; Yangmin Li; Zheyang Qin; Nikash Bhardwaj; Amir Zadeh; Chuan Li; Katerina Fragkiadaki; Deepak Pathak; | code |
| 8 | On The Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A central challenge is the lack of control in modern training pipelines: large-scale pre-training corpora are opaque, mid-training is often underexamined, and RL objectives interact with unknown prior knowledge in complex ways. To resolve this ambiguity, we develop a fully controlled experimental framework that isolates the causal contributions of pre-training, mid-training, and RL-based post-training. |
Charlie Zhang; Graham Neubig; Xiang Yue; | code |
| 9 | SSA: Sparse Sparse Attention By Aligning Full and Sparse Attention Outputs in Feature Space Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose SSA (Sparse Sparse Attention), a training framework that integrates both sparse and full attention with bidirectional attention-output alignment. |
Zhenyi Shen; Junru Lu; Lin Gui; Jiazheng Li; Yulan He; di yin; Xing Sun; | code |
| 10 | How2Everything: Mining The Web for How-to Procedures to Evaluate and Improve LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Yet, measuring and improving procedural validity at scale on real-world tasks remains challenging and understudied. To address this, we introduce How2Everything, a scalable framework to evaluate and improve goal-conditioned procedure generation. |
Yapei Chang; Kyle Lo; Mohit Iyyer; Luca Soldaini; | code |
| 11 | Entropy-Aware On-Policy Distillation of Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, we show that the mode-seeking property of reverse KL reduces generation diversity and yields unstable learning signals when the teacher distribution has high entropy. To address this, we introduce Entropy-Aware On-Policy Distillation. |
Woogyeol Jin; Taywon Min; Yongjin Yang; Swanand Kadhe; Yi Zhou; Dennis Wei; Nathalie Baracaldo; Kimin Lee; | code |
| 12 | ModernVBERT: Towards Smaller Visual Document Retrievers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Increasingly, Visual Document Retrieval (VDR) models, which directly embed images of document pages, are used as an alternative to text-only retrievers. |
Paul Teiletche; Quentin Macé; Max Conti; António Loison; Gautier Viaud; Pierre Colombo; Manuel Faysse; | code |
| 13 | Clipping Bottleneck: Stabilizing RLVR Via Stochastic Recovery of Near-Boundary Signals Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, we find that many high-value signals lie in the **near-boundary** region just beyond the clipping threshold, and are thus discarded. Motivated by this diagnosis, we propose **Near-boundary Stochastic Rescue (NSR)**, a minimal, plug-and-play modification that stochastically retains these slightly out-of-bound tokens to recover lost signals. |
Shuo Yang; Jinda Lu; Chiyu Ma; Kexin Huang; Haoming Meng; Qihui Zhang; Yuyang Liu; Bolin Ding; Guoyin Wang; Li Yuan; Jingren Zhou; | code |
| 14 | Any-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Meanwhile, recent studies have successfully applied discrete diffusion models to natural language processing, revealing their considerable potential as a promising new approach in this domain. Drawing inspiration from these pioneering researches, we introduce Any-Diffusion, the first any-to-any multimodal language model built purely on mask-based discrete diffusion models, which unifies understanding and generation across text, speech, and images. |
lijiang Li; zuwei long; Yunhang Shen; Heting Gao; Haoyu Cao; Xing Sun; Caifeng Shan; Ran He; Chaoyou Fu; | code |
| 15 | CSD: Content-aware Speculative Decoding for Efficient Image Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a novel content-aware speculative decoding algorithm, termed CSD, which integrates an entropy-based probability relaxation mechanism with an optimal resampling strategy to enhance the inference efficiency for autoregressive image generation. |
Mingcheng Wang; junbo qiao; Yunchen Li; Lingfu Jiang; Wei Li; Jie Hu; Jiao Xie; Zhou Yu; Xinghao Chen; Guixu Zhang; Shaohui Lin; | code |
| 16 | SIGMA-PPG: Statistical-prior Informed Generative Masking Architecture for PPG Foundation Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Standard masked modeling often yields trivial solutions while contrastive methods lack morphological precision. To address these limitations, we propose a Statistical-prior Informed Generative Masking Architecture (SIGMA-PPG), a generative foundation model featuring a prior-guided adversarial masking mechanism, where a reinforcement learning-driven teacher leverages statistical priors to create challenging learning paths that prevent overfitting to noise. |
ZONGHENG GUO; Tao Chen; Yang Jiao; Yi Pan; Xiao Hu; Manuela Ferrario; | code |
| 17 | From Prior to Pro: Efficient Skill Mastering Via Distribution Contractive RL Finetuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Distribution Contractive Reinforcement Learning (DICE-RL), a framework that uses reinforcement learning (RL) as a “distribution contractor” to refine pretrained generative robot policies. |
Zhanyi Sun; shuran song; | code |
| 18 | DF-ExpEnse: Diffusion Filtered Exploration for Sample Efficient Finetuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work we present Diffusion Filtered Exploration via Ensembles (DF-ExpEnse), an exploration technique that meaningfully improves the quality of online experience collection, thus increasing the sample efficiency of the finetuning procedure. |
Calvin Luo; Chen Sun; shuran song; | code |
| 19 | The Flexibility Trap: Rethinking The Value of Arbitrary Order in Diffusion Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, in this paper, we reveal that for general reasoning tasks (e.g., mathematics and coding), arbitrary order generation may in fact limit the reasoning potential of dLLMs. We find that dLLMs tend to exploit this order flexibility to bypass high-uncertainty tokens that are crucial for exploration, leading to a premature collapse of solution coverage. |
Zanlin Ni; Shenzhi Wang; Yang Yue; Tianyu Yu; Weilin Zhao; Yeguo Hua; Tianyi Chen; Jun Song; YuCheng; Bo Zheng; Gao Huang; | code |
| 20 | Scaling Long-Horizon Agent Via Context Folding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Context Folding, a framework that empowers agents to actively manage their working context. |
Weiwei Sun; Lu Miao; Zhan Ling; Kang Liu; Xuesong Yao; Yiming Yang; Jiecao Chen; | code |
| 21 | WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents WorldPlay, a streaming video diffusion model that enables real-time, interactive world modeling with long-term geometric consistency, resolving the trade-off between speed and memory that limits current methods. |
Wenqiang Sun; Haiyu Zhang; Haoyuan Wang; Junta Wu; Zehan Wang; Zhenwei Wang; Yunhong Wang; Jun Zhang; Tengfei Wang; Chunchao Guo; | code |
| 22 | Any-Order GPT As Masked Diffusion Model: Decoupling Formulation and Architecture Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We show decoder-only MDMs, despite a larger modeling space, can achieve significant inference speedups ($\sim25\times$) and comparable perplexity with techniques like temperature annealing, offering a path to reduced inference compute. |
Shuchen Xue; Tianyu Xie; Tianyang Hu; Zijin Feng; Jiacheng Sun; Kenji Kawaguchi; Zhenguo Li; Zhi-Ming Ma; | code |
| 23 | Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we establish a novel theoretical analysis: DDPO is an implicit form of score/flow matching with noisy targets, which increases variance and slows convergence. |
Shuchen Xue; Chongjian GE; Shilong Zhang; Yichen Li; Zhi-Ming Ma; | code |
| 24 | Simultaneous Speech-to-Speech Translation Without Aligned Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We instead propose Hibiki-Zero, a model for simultaneous speech translation trained without word-level alignments between source and target speech. |
Tom Labiausse; Romain Fabre; Yannick Estève; Alexandre Défossez; Neil Zeghidour; | code |
| 25 | Plan for Speed: Dilated Scheduling for Masked Diffusion Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Masked diffusion language models (MDLMs) promise fast, non-autoregressive text generation, yet existing samplers, which pick tokens to unmask based on model confidence, ignore interactions when unmasking multiple positions in parallel and effectively reduce to slow, autoregressive behavior. We propose the Dilated Unmasking Scheduler (DUS), an inference-only, planner-model-free method that partitions sequence positions into non-adjacent dilated groups and unmasked them in parallel so as to minimize an upper bound on joint entropy gain at each denoising step. |
Omer Luxembourg; Haim Permuter; Eliya Nachmani; | code |
| 26 | Rethinking The Trust Region in LLM Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This creates a sub-optimal learning dynamic: updates to low-probability tokens are aggressively over-penalized, while potentially catastrophic shifts in high-probability tokens are under-constrained, leading to training inefficiency and instability. To address this, we propose Divergence Proximal Policy Optimization (DPPO), which substitutes heuristic clipping with a more principled constraint based on a direct estimate of policy divergence (e.g., Total Variation or KL). |
Penghui Qi; Xiangxin Zhou; Zichen Liu; Tianyu Pang; Chao Du; Min Lin; Wee Sun Lee; | code |
| 27 | SiameseNorm: Breaking The Barrier to Reconciling Pre/Post-Norm Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We attribute this phenomenon to a structural incompatibility within a *single-stream* design: Any application of the Post-Norm operation inevitably obstructs the clean identity gradient preserved by Pre-Norm. To fundamentally reconcile these paradigms, we propose SiameseNorm, a *two-stream* architecture that couples Pre-Norm-like and Post-Norm-like streams with shared parameters. |
Tianyu Li; Dongchen Han; Zixuan Cao; Haofeng Huang; Mengyu Zhou; Ming Chen; erchao.zec; xiaoxi jiang; guanjunjiang; Gao Huang; | code |
| 28 | WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present WorldMirror, a unified feed-forward model for comprehensive 3D geometric prediction tasks. |
Yifan Liu; Zhiyuan Min; Zhenwei Wang; Junta Wu; Tengfei Wang; Yixuan Yuan; Yawei Luo; Chunchao Guo; | code |
| 29 | Compressed Sensing for Capability Localization in Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Zeroing out as few as five task-specific heads can degrade performance by up to $65\\%$ on standard benchmarks measuring the capability of interest, while largely preserving performance on unrelated tasks. We introduce a compressed sensing based method that exploits the sparsity of these heads to identify them via strategic knockouts and a small number of model evaluations. |
Anna Bair; Yixuan Xu; Mingjie Sun; Zico Kolter; | code |
| 30 | Real-Time and Lightweight Diffusion Image Compression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we explore the design of real-time and lightweight diffusion codecs by addressing two pivotal questions. |
Zhaoyang Jia; Naifu Xue; Zihan Zheng; Jiahao Li; Bin Li; Xiaoyi Zhang; Zongyu Guo; Yuan Zhang; Houqiang Li; Yan Lu; | code |
| 31 | Linearizing Vision Transformer with Test-Time Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inheriting weights from pretrained Transformers provides an appealing shortcut, yet the fundamental representational gap between Softmax and linear attention prevents effective weight transfer. In this work, we address this conversion challenge from two perspectives: architectural alignment and representational alignment. |
Yining Li; Dongchen Han; Zeyu Liu; Hanyi Wang; Yulin Wang; Gao Huang; | code |
| 32 | Scaling Prompt Synthesis for Large Language Model Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: PromptCoT showed that injecting rationales into prompt synthesis increases problem difficulty. Building on this, we present PromptScale, a scalable framework that replaces hand-crafted heuristics with an expectation-maximization (EM) loop, where rationales are iteratively refined to guide prompt construction. |
Xueliang Zhao; Wei Wu; Jian Guan; Zhuocheng Gong; Lingpeng Kong; | code |
| 33 | Evaluating LLM Uncertainty in Long-Form Generation Using Deterministic Ground Truth Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Single-answer Atomic Long-form Target (SALT), a benchmark of six procedurally generated tasks with single deterministic long textual ground truths, enabling unit-level evaluation of correctness, calibration, and ranking without external judges. |
Ido Amit; Ido Galil; Ran El-Yaniv; | code |
| 34 | Stop Training for The Worst: Progressive Unmasking Accelerates Masked Diffusion Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose Progressive UnMAsking (PUMA), a simple modification of the forward masking process that aligns training-time and inference-time masking patterns, thereby focusing optimization on *inference-aligned masks* and speeding up training. |
Jaeyeon Kim; Jonathan Geuter; David Alvarez-Melis; Sham Kakade; Sitan Chen; | code |
| 35 | Fine-Tuning Masked Diffusion for Provable Self-Correction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Prior attempts to incorporate self-correction into MDMs either require overhauling MDM architectures/training or rely on imprecise proxies for token quality, limiting their applicability. Motivated by this, we introduce PRISM–Plug-in Remasking for Inference-time Self-correction of Masked Diffusions–a lightweight, model-agnostic approach that applies to any pretrained MDM. |
Jaeyeon Kim; Seunggeun Kim; Taekyun Lee; David Pan; Hyeji Kim; Sham Kakade; Sitan Chen; | code |
| 36 | On Path to Multimodal Historical Reasoning: HistBench and HistAgent Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing general-purpose agents perform well on many current benchmarks but lack the domain expertise needed to address complex historical questions. To address this gap, we introduce HistBench, a new benchmark of 414 high-quality and carefully-reviewed questions stratified by difficulty and designed to evaluate LLM’s capacity for historical reasoning. |
Jiahao Qiu; Fulian Xiao; Yimin Wang; Yuchen Mao; Yijia Chen; Xinzhe Juan; Siran Wang; Xuan Qi; Tongcheng Zhang; Zixin Yao; Jiacheng Guo; Yifu Lu; Charles Argon; Jundi Cui; Daixin Chen; Junran Zhou; Shuyao Zhou; Zhanpeng Zhou; Ling Yang; Shilong Liu; Hongru WANG; Kaixuan Huang; xun jiang; Xi Gao; Mengdi Wang; | code |
| 37 | The Surprising Difficulty of Search in Model-Based Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Instead, we show that mitigating distribution shift matters more than improving model or value function accuracy. Building on this insight, we identify key techniques for enabling effective search, achieving state-of-the-art performance across multiple popular benchmark domains. |
Wei-Di Chang; Mikael Henaff; Brandon Amos; Gregory Dudek; Scott Fujimoto; | code |
| 38 | OmniDenseCap: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper proposes Omni Dense Captioning, a novel task designed to generate continuous, fine-grained, and structured audio-visual narratives with explicit timestamps. |
Linli Yao; Yuancheng Wei; Yaojie Zhang; Lei Li; Xinlong Chen; Feifan Song; Ziyue Wang; Kun Ouyang; Yuanxin Liu; Lingpeng Kong; Qi Liu; Pengfei Wan; Kun Gai; Yuanxing Zhang; Xu SUN; | code |
| 39 | ObjEmbed: Towards Universal Multimodal Object Embeddings Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present ObjEmbed, a novel MLLM embedding model that decomposes the input image into multiple regional embeddings, each corresponding to an individual object, along with global embeddings. |
Shenghao Fu; Yukun Su; Fengyun Rao; Jing LYU; Xiaohua Xie; Wei-Shi Zheng; | code |
| 40 | Learnability-Informed Fine-Tuning of Diffusion Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We aim to improve the reasoning capabilities of diffusion language models (DLMs). |
Shubham Parashar; Atharv Chagi; Jacob Helwig; Lakshmi Madhavarapu; Sushil Vemuri; James Caverlee; Dileep Kalathil; Shuiwang Ji; | code |
| 41 | SPA: A Simple But Tough-to-Beat Baseline for Knowledge Injection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose **SPA** (**S**caling **P**rompt-engineered **A**ugmentation), a simple but tough-to-beat baseline that uses a small set of carefully designed prompts to generate large-scale synthetic data for knowledge injection. |
Kexian Tang; Jiani Wang; Shaowen Wang; Kaifeng Lyu; | code |
| 42 | SlideSparse: Fast and Flexible (2N-2):2N Structured Sparsity Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present **SlideSparse**, the first system to unlock Sparse Tensor Core acceleration for the $(2N-2):2N$ model family on commodity GPUs. |
Yingbo HAO; Hanyong Shao; Ting Song; Yan Xia; Di Zhang; Shaohan Huang; Xun Wu; Songchen Xu; Le Xu; Li Dong; Zewen Chi; Yi Zou; Furu Wei; | code |
| 43 | MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work proposes a novel embodied 4D world model that enables geometrically consistent, arbitrary-view RGBD generation: given only a single-view RGBD observation as input, the model “imagines” the remaining viewpoints, which can then be back-projected and fused to assemble a more complete 3D structure across time. |
Jiaxu Wang; JIANG Yicheng; Tianlun HE; Jingkai SUN; Qiang Zhang; Jiahang Cao; Zesen Gan; Mingyuan Sun; Qiming Shao; Xiangyu Yue; | code |
| 44 | Doc-to-LoRA: Learning to Instantly Internalize Contexts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While context distillation (CD) can transfer information into model parameters, per-prompt distillation is impractical due to training costs and latency. To address these limitations, we propose Doc-to-LoRA (D2L), a lightweight hypernetwork that meta-learns to perform approximate CD within a single forward pass. |
Rujikorn Charakorn; Edoardo Cetin; Shinnosuke Uesaka; Robert Lange; | code |
| 45 | Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Just-In-Time Reinforcement Learning (JitRL), a training-free framework that enables test-time policy optimization without any gradient updates. |
yibo li; Zijie Lin; Ailin Deng; Xuan Zhang; Yufei He; Shuo Ji; Tri Cao; Bryan Hooi; | code |
| 46 | OvisOCR: End-to-End Document Parsing Via Aligning Specialized Perception with General Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents OvisOCR, a lightweight and strictly end-to-end Multimodal Language Model (MLLM) tailored for document parsing. |
Jun-Peng Jiang; Shiyin Lu; An-Yang Ji; Yinglun Li; Qing-Guo Chen; Zhao Xu; Weihua Luo; Kaifu Zhang; De-Chuan Zhan; Han-Jia Ye; | code |
| 47 | Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model Unlearning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite rapid progress, KD-based unlearning struggles with biased deletion due to suppressing specific token sequences as a substitute for complete knowledge removal, whereas DR-based unlearning risks the re-emergence of harmful knowledge because the underlying knowledge remains intact. To address these issues, we propose Distinguishable Deletion ($\mathrm{D^2}$), a paradigm that restricts the response distribution in the latent space rather than specific tokens to erase undesirable knowledge, while distinguishing it from retained knowledge, enabling a refusal mechanism to handle unlearned inputs safely and coherently. |
Puning Yang; Junchi Yu; Qizhou Wang; Phil Torr; Bo Han; Xiuying Chen; | code |
| 48 | Olivia: Harmonizing Time Series Foundation Models with Power Spectral Density Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, we propose \textit{Harmonizer}, a module that reshapes spectral structures and implicitly harmonizing PSDs across datasets, which theoretically corresponds to a shared reparameterization of second-order temporal correlations. |
Jingru Fei; Kun Yi; Alex Wang; Qingsong Wen; Xiangxiang Zhu; Wei Fan; | code |
| 49 | Closing The Loop: Universal Repository Representation with RPG-Encoder Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We consider repository comprehension and generation to be inverse processes within a unified cycle: generation expands intent into implementation, while comprehension compresses implementation back into intent. To address this, we propose RPG-Encoder, a framework that generalizes the Repository Planning Graph (RPG) from a static generative blueprint into a unified, high-fidelity representation. |
Jane Luo; Chengyu Yin; Xin Zhang; Qingtao Li; Steven Liu; Yiming Huang; Jie Wu; Hao Liu; Yangyu Huang; Yu Kang; Fangkai Yang; Ying Xin; Scarlett Li; | code |
| 50 | MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, they suffer from limited multi-rationale semantic modeling, insufficient logical robustness, and susceptibility to misleading cues. Therefore, we propose a Multi-rationale INtegrated Discriminative (MIND) reasoning framework, which is designed to endow MLLMs with human-like cognitive abilities of “Understand → Rethink → Correct”, and achieves a paradigm evolution from passive imitation-based reasoning to active discriminative reasoning. |
Chuang Yu; Jinmiao Zhao; Mingxuan Zhao; Yunpeng Liu; Xiujun Shu; Feng Yuanhao; Bo Wang; Xiangyu Yue; | code |
| 51 | OSF: On Pre-training and Scaling of Sleep Foundation Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: With an enhanced pre-training and scaling recipe, we introduce OSF, a family of sleep FMs that achieves state-of-the-art performance across nine datasets on diverse sleep and disease prediction tasks. |
Zitao Shuai; Zongzhe Xu; David Yang; Wei Wang; Yuzhe Yang; | code |
| 52 | ScaleSim: Serving Large-Scale Multi-Agent Simulation with Invocation Distance-Based Memory Management Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Based on an analysis of representative workload classes, we introduce invocation distance, a unified abstraction that estimates the relative order in which agents will issue future LLM requests. |
Zaifeng Pan; Yipeng Shen; Zhengding Hu; Zhuang Wang; Aninda Manocha; Zheng Wang; zhongkai yu; Yue Guan; Yufei Ding; | code |
| 53 | AVTrack: Audio-Visual Speaker Tracking in Complex Scenes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Such oversimplified settings bias evaluation toward static audio–visual co-occurrence, rather than rigorously assessing robust spatiotemporal modeling and cross-modal reasoning in complex, dynamic scenes. To address these limitations, we introduce \textbf{AVTrack}, a human-centric audio-visual instance segmentation (AVIS) dataset designed for dynamic real-world scenarios. |
Yaoting Wang; Yun Zhou; Zipei Zhang; Henghui Ding; | code |
| 54 | Faults in Our Formal Benchmarking: Dataset Defects and Evaluation Failures in Lean Theorem Proving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a fault taxonomy, a suite of automated checkers and prompts, and release standards to guide the creation of formal math datasets and make evaluation more reproducible and trustworthy. |
Pawan Sasanka Ammanamanchi; Siddharth Bhat; Stella Biderman; | code |
| 55 | SpaceVista: All-Scale Visual Spatial Reasoning from Mm to Km Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce a holistic solution that integrates a structured spatial reasoning knowledge system, scale-aware modeling, and a progressive training paradigm, as the **first attempt** to broaden the scope of all-scale spatial intelligence. |
Peiwen Sun; Shiqiang Lang; Dongming Wu; Ding Yi; Kaituo Feng; Huadai Liu; Zhen Ye; Rui Liu; Yun-Hui Liu; Jianan Wang; Xiangyu Yue; | code |
| 56 | MOD-SR: Unifying Multimodal Learning and Direct Optimization with Gradient-Guided Diffusion Model for Symbolic Regression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose MOD-SR, unifying multimodal distribution learning during training with direct optimization at inference time. |
Chuyang Xiang; Yichen Wei; Junchi Yan; | code |
| 57 | Neuro-evolutionary Continual Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by neuroscience, we propose Neuro-evolutionary Continual Reinforcement Learning (Nevo-CRL). |
Pengyi Li; Hongyao Tang; Yifu Yuan; Yan Zheng; Xin Xu; Jianye Hao; | code |
| 58 | Rule2DRC: Benchmarking LLM Agents for DRC Script Synthesis with Execution-Guided Test Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we introduce Rule2DRC, a large-scale benchmark for DRC script coding agents with 1,000 rule-to-script tasks and 13,921 evaluation chip layouts for execution-based scoring. |
Jinuk Kim; Junsoo Byun; Donghwi Hwang; Seong-Jin Park; Hyun Oh Song; | code |
| 59 | From Coarse to Fine: Deep Prototype Refinement Network for Few-Shot Point Cloud Semantic Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing prototype-based methods typically rely on shallow feature fusion strategies, failing to adequately model the feature distribution shift between support and query sets, resulting in insufficient prototype adaptation. To address this, we propose the Deep Prototype Refinement Network (DPR-Net), which systematically achieves progressive adaptation by constructing a coarse-to-fine prototype evolution trajectory. |
Changshuo Wang; Shuting He; Xiang Fang; Weijun Li; Xingyu Gao; Zhonghang Liu; Prayag Tiwari; Dimitrios Kanoulas; | code |
| 60 | Replay Failures As Successes: Sample-Efficient Reinforcement Learning for Instruction Following Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose ***H**indsight **i**nstruction **R**eplay* (HiR), a novel sample-efficient RL framework for complex instruction following tasks, which employs a *select*-then-*rewrite* strategy to *replay failed attempts as successes* based on the constraints that have been satisfied in hindsight. |
Kongcheng Zhang; QI YAO; Shunyu Liu; Wenjian Zhang; Cen; Yang Zhou; Wenkai Fang; Yiru Zhao; Baisheng Lai; Mingli Song; | code |
| 61 | Select to Think: Unlocking SLM Potential with Local Sufficiency Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We address this dilemma by identifying local sufficiency: at divergence points, the LLM’s preferred token consistently resides within the SLM’s top-K next-token predictions, even when failing to emerge as the SLM top-1 choice. We therefore propose SELECT TO THINK (S2T), which reframes the LLM’s role from open-ended generation to selection among the SLM’s proposals, simplifying the supervision signal to discrete candidate rankings. |
Wenxuan Ye; Yangyang Zhang; Xueli An; Georg Carle; Yunpu Ma; | code |
| 62 | ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this, there is a pressing need for difficult benchmarks that remain relevant for longer. We take this idea to its limit by introducing ZeroBench—a lightweight visual reasoning benchmark curated using adversarial filtering to be “impossible” for frontier LMMs at release time, with initial SotA scores of 0% pass@1 and pass∧5. |
Jonathan Roberts; Mohammad Reza Taesiri; Ansh Sharma; Akash Gupta; Samuel Roberts; Ioana Croitoru; Vlad Bogolin; Jialu Tang; Florian Langer; Vyas Raina; Vatsal Raina; Hanyi Xiong; Vishaal Udandarao; Jingyi Lu; Chen Shiyang; Sam Purkis; Tianshuo Yan; Wenye Lin; Gyungin Shin; Qiaochu Yang; Anh Nguyen; David Atkinson; Alexandru Coca; Mikah Đặng; Sebastian Dziadzio; Jakob Kunz; Kaiqu Liang; Alexander Lo; Brian Pulfer; Steven Walton; Charig Yang; Kai Han; Samuel Albanie; | code |
| 63 | Turning Drift Into Constraint: Robust Reasoning Alignment in Non-Stationary Multi-Stream Environments Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Autonomous Preference Optimization (APO), a novel framework that treats inter-model divergences not as noise, but as dynamic negative constraints. |
Xiaoyu Yang; Jie Lu; Wei Duan; En Yu; | code |
| 64 | FRIGID: Scaling Diffusion-Based Molecular Generation from Mass Spectra at Training and Inference Time Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present FRIGID, a framework with a novel diffusion language model that generates molecular structures conditioned on mass spectra via intermediate fingerprint representations and determined chemical formulae, training at the scale of hundreds of millions of unlabeled structures. |
Montgomery Bohde; Hongxuan Liu; Mrunali Manjrekar; Magdalena Lederbauer; Shuiwang Ji; Runzhong Wang; Connor Coley; | code |
| 65 | Any3D-VLA: Enhancing VLA Robustness Via Diverse Point Clouds Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address the challenges of (1) scarce 3D data and (2) the domain gap induced by cross-environment differences and depth-scale biases, we propose Any3D-VLA. |
Xianzhe Fan; Shengliang Deng; Xiaoyang Wu; Yuxiang Lu; Zhuoling Li; Mi Yan; Yujia Zhang; Zhizheng Zhang; He Wang; Hengshuang Zhao; | code |
| 66 | Proteo-R1: Thinking Foundation Models for De Novo Protein Binder Design Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: \textbf{ThinkProteo} reimagines generative science by introducing reasoning-guided diffusion models that think step-by-step, akin to how a scientist hypothesizes, tests, and refines molecular ideas. |
Fang Wu; Li Li; Weihao Xuan; Heli Qi; Zeqi Zhou; Hanqun CAO; Heng-Jui Chang; Haokai Zhao; Jian Ma; Zijian Carl; Yu-Chi Cheng; Robert Tang; Zehong Wang; Kuan Pang; Hanchen Wang; Kejun Ying; Pan Lu; Chiho Im; Seungju Han; Peng Xia; Yinxi Li; Guanlue Li; Tinson Xu; Deyao Zhu; Pheng Ann Heng; Naoto Yokoya; Masashi Sugiyama; Jure Leskovec; Yejin Choi; | code |
| 67 | Leak@$k$: Unlearning Does Not Make LLMs Forget Under Probabilistic Decoding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our findings demonstrate that knowledge leakage persists across methods and tasks, underscoring that current state-of-the-art unlearning techniques provide only limited forgetting and highlighting the urgent need for more robust approaches to LLM unlearning. We propose an algorithm, termed Robust Unlearning under LEak@$k$ metric (\texttt{RULE}), which serves as an initial step toward addressing this concern. |
Hadi Reisizadeh; Jiajun Ruan; Yiwei Chen; Soumyadeep Pal; Sijia Liu; Mingyi Hong; | code |
| 68 | Beyond Log Likelihood: Probability-Based Objectives for Supervised Fine-Tuning Across The Model Capability Continuum Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Rather than proposing a single universally superior replacement loss, we systematically study various probability-based objectives and characterize when and why different objectives succeed or fail under varying conditions. |
Gaotang Li; Ruizhong Qiu; Xiusi Chen; Heng Ji; Hanghang Tong; | code |
| 69 | INT Vs. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We reveal a critical performance crossover: while FP excels in coarse-grained quantization, INT consistently surpasses it as the quantization block size shrinks. |
Mengzhao Chen; Meng Wu; Hui Jin; Zhihang Yuan; Jing Liu; Chaoyi Zhang; Yunshui Li; Jie Huang; Jin Ma; Zeyue Xue; Zhiheng Liu; Xingyan Bin; Ping Luo; | code |
| 70 | Contrastive Weak-to-Strong Generalization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this challenge, we leverage implicit rewards, which approximate explicit rewards through log-likelihood ratios, and reveal their structural equivalence with Contrastive Decoding (CD), a decoding strategy shown to reduce noise in LLM generation. Building on this connection, we propose \textbf{Contrastive Weak-to-Strong Generalization (ConG)}, a framework that employs contrastive decoding between pre- and post-alignment weak models to generate higher-quality samples. |
Houcheng Jiang; Junfeng Fang; Jiaxin Wu; Tianyu Zhang; Chen Gao; Xiang Wang; Xiangnan He; Yang Deng; | code |
| 71 | VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While Vision-Language-Action models (VLAs) are rapidly advancing toward generalist robot policies, quantitatively characterizing their capability boundaries and failure modes remains challenging. To address this, we introduce **VLA-Arena**, a comprehensive benchmark. |
Borong Zhang; Jiahao Li; Jiachen Shen; Yishuai Cai; Yuhao Zhang; Yuanpei Chen; Juntao Dai; Jiaming Ji; Yaodong Yang; | code |
| 72 | OpenSage: Self-programming Agent Generation Engine Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose OpenSage, the first ADK that enables LLMs to automatically create agents with self-generated topology and toolsets while providing comprehensive and structured memory support. |
Hongwei Li; Zhun Wang; Qinrun Dai; Yuzhou Nie; Jinjun Peng; Ruitong Liu; Jingyang Zhang; Kaijie Zhu; Jingxuan He; Lun Wang; Yangruibo Ding; Yueqi Chen; Wenbo Guo; Dawn Song; | code |
| 73 | Necessary Conditions for Compositional Generalization of Embedding Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Modern models are trained on massive datasets, yet these are vanishingly small compared to the full combinatorial space of possible data, raising the question of whether models can reliably generalize to unseen combinations. To formalize what this requires, we propose a set of practically motivated desiderata that any compositionally generalizing system must satisfy, and analyze their implications under standard training with linear classification heads. |
Arnas Uselis; Andrea Dittadi; Seong Joon Oh; | code |
| 74 | Parameters As Experts: Adapting Vision Models with Dynamic Parameter Routing for Dense Predictions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose AdaRoute, a new adapter-style method featuring a simple mixture-of-experts (MoE) architecture. |
Meng Lou; Stanley Yu; Yizhou Yu; | code |
| 75 | Scaling Continual Learning with Bi-Level Routing Mixture-of-Experts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose $\mathbf{CaRE}$, a scalable $\mathbf{C}$ontinual Le$\mathbf{a}$rner with efficient Bi-Level $\mathbf{R}$outing Mixture-of-$\mathbf{E}$xperts (BR-MoE). |
Meng Lou; Yunxiang Fu; Yizhou Yu; | code |
| 76 | ACON: Optimizing Context Compression for Long-horizon LLM Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Agent Context Optimization (ACON), a unified framework that optimally compresses both observations and history into concise, informative representations. |
Minki Kang; Wei-Ning Chen; Dongge Han; Huseyin Inan; Lukas Wutschitz; Yanzhi Chen; Robert A Sim; Saravanakumar Rajmohan; | code |
| 77 | SimpleMem: Efficient Lifelong Memory for LLM Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing approaches either retain full interaction histories via passive context extension, leading to substantial redundancy, or rely on iterative reasoning to filter noise, incurring high token costs. To address this challenge, we introduce SimpleMem, an efficient memory framework based on semantic lossless compression. |
Jiaqi Liu; Yaofeng Su; Peng Xia; Siwei Han; Zeyu Zheng; Cihang Xie; Mingyu Ding; Huaxiu Yao; | code |
| 78 | IVQ: Structured and Lightweight Vector Quantization Via Binary Hierarchical Composition Inspired By $\textit{IChing}$ Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose *IChing* Vector Quantization (IVQ), a lightweight and structured vector quantization framework inspired by *IChing*. |
Heda Zuo; Junxian Wu; Fengjie Lu; Pei Chen; Lingyun Sun; Weitao You; | code |
| 79 | Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although it is increasingly adopted, are vibe coding outputs really safe to deploy in production? To answer this question, we propose SUSVIBES, a benchmark consisting of 200 feature-request software engineering tasks from real-world open-source projects, which, when given to human programmers, led to vulnerable implementations. |
Songwen Zhao; Danqing Wang; Kexun Zhang; Jiaxuan Luo; Zhuo Li; Lei Li; | code |
| 80 | Regulating Anatomy-Aware Rewards Via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Using CABS, we identify a “\textbf{mechanistic divergence}” in standard RL, where surface-similarity rewards drive policy gradients to bypass medical facts. We therefore propose \textbf{Trajectory-Integral Feedback GRPO (TIF-GRPO)}, a novel framework integrating control-theoretic principles into policy optimization. |
Tianwei Lin; Zhongwei Qiu; Jie Cao; Jiang Liu; Wenjie Yan; Bo Zhang; Yu Zhong; Wenqiao Zhang; Yingda Xia; Ling Zhang; | code |
| 81 | The Latent Color Subspace: Emergent Order in High-Dimensional Chaos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We verify our Latent Color Subspace (LCS) interpretation by demonstrating that it can both predict and explicitly control color, introducing a fully training-free method in FLUX based solely on closed-form latent-space manipulation. |
Mateusz Pach; Jessica Bader; Quentin Bouniot; Serge Belongie; Zeynep Akata; | code |
| 82 | InteractComp: Evaluating Search Agents With Ambiguous Queries Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Yet most agents lack interactive mechanisms during the search process, and existing benchmarks cannot assess this capability. To address this gap, we introduce INTERACTCOMP, a benchmark designed to evaluate whether search agents can recognize query ambiguity and actively interact to resolve it during search. |
Mingyi Deng; Lijun Huang; Yani Fan; Fanqi Kong; Jiayi Zhang; Fashen Ren; Jinyi Bai; Fuzhen Yang; Dayi Miao; Zhaoyang Yu; Yifan Wu; Yanfei Zhang; Fengwei Teng; Yingjia Wan; Song Hu; Yude Li; Xin Jin; Conghao Hu; Haoyu Li; Qirui Fu; Tai Zhong; Xinyu Wang; Robert Tang; Nan Tang; Wu; Yuyu Luo; | code |
| 83 | GraphPFN: A Prior-Data Fitted Network for Graph Node-Level Tasks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we make the next step by proposing GraphPFN, a PFN-based model designed and pretrained specifically for graph node-level tasks. |
Dmitry Eremeev; Oleg Platonov; Gleb Bazhenov; Artem Babenko; Liudmila Prokhorenkova; | code |
| 84 | DLO-Lab: Benchmarking Deformable Linear Object Manipulations with Differentiable Physics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Additionally, existing simulation environments offer limited support for the broad spectrum of material behaviors necessary for generalizable DLO manipulation. To overcome these limitations, we introduce a differentiable simulator explicitly designed for versatile DLO manipulation. |
Junyi Cao; Yian Wang; Ziyan Xiong; Chunru Lin; Zhehuan Chen; Chuang Gan; | code |
| 85 | Learning Taxonomic Trees with Hierarchical Representation Regularization for Large Multimodal Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, we introduce a semantic-aware visual tree construction framework that extracts coarse-to-fine visual features from intermediate LLM layers guided by textual cues. |
Hulingxiao He; Zhi Tan; Yuxin Peng; | code |
| 86 | MASPO: Joint Prompt Optimization for LLM-based Multi-Agent Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While the quality of these prompts is pivotal, jointly optimizing them across interacting agents remains a non-trivial challenge, primarily due to the misalignment between local agent objectives and holistic system goals. To address this, we introduce MASPO, a novel framework designed to automatically and iteratively refine prompts across the entire system. |
Zhexuan Wang; Xuebo Liu; Li Wang; Zifei Shan; Yutong Wang; Zhenxi Song; Min zhang; | code |
| 87 | Transolver-3: Scaling Up Transformer Solvers to Industrial-Scale Geometries Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Transolver-3, a new member of the Transolver family as a highly scalable framework designed for high-fidelity physics simulations. |
Hang Zhou; Haixu Wu; Haonan Shangguan; Yuezhou Ma; Huikun Weng; Jianmin Wang; Mingsheng Long; | code |
| 88 | World Guidance: World Modeling in Condition Space for Action Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing approaches struggle to strike a balance between maintaining efficient, predictable future representations and preserving sufficient fine-grained information to guide precise action generation. To address this limitation, we propose WoG (World Guidance), a framework that maps future observations into compact conditions by injecting them into the action inference pipeline. |
Yue Su; Sijin Chen; Haixin Shi; Mingyu Liu; Zhengshen Zhang; Ningyuan Huang; Weiheng Zhong; Zhengbang Zhu; Yuxiao Liu; Xihui Liu; | code |
| 89 | Identifiable Token Correspondence for World Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We formulate next-frame prediction as a structured probabilistic inference problem with latent token correspondence variables, deriving a model in which each next-frame token is explained either by copying a token from the previous frame or by generating a new token. |
Youngin Kim; Ray Sun; Inho Kim; Bumsoo Park; Hyun Oh Song; | code |
| 90 | DropoutTS: Sample-Adaptive Dropout for Robust Time Series Forecasting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce DropoutTS, a model-agnostic plugin that shifts the paradigm from what to learn to how much to learn. |
Siru Zhong; Yiqiu Liu; Zhiqing Cui; Zezhi Shao; Fei Wang; Qingsong Wen; Yuxuan Liang; | code |
| 91 | Code2Video: A Code-centric Paradigm for Educational Video Creation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose **Code2Video**, a code-centric agent framework that generates educational videos by writing executable Python programs. |
Yanzhe Chen; Kevin Qinghong Lin; Mike Zheng Shou; | code |
| 92 | VividCam: Learning Unconventional Camera Motions from Virtual Synthetic Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The challenge lies in the difficulty of finding sufficient training videos with the intended uncommon camera motions. To address this challenge, we propose VividCam, a training paradigm that enables diffusion models to learn complex camera motions from synthetic videos, releasing the reliance on collecting realistic training videos. |
Qiucheng Wu; Handong Zhao; Zhixin Shu; Jing Shi; Yang Zhang; Shiyu Chang; | code |
| 93 | PhoStream: Benchmarking Real-World Streaming for Omnimodal Assistants in Mobile Scenarios Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce **PhoStream**, the first mobile-centric streaming benchmark that unifies on-screen and off-screen scenarios to evaluate video, audio, and temporal reasoning. |
Xudong LU; Guan Huankang; Yang Bo; Jinpeng Chen; Xintong Guo; Shuhan LI; Fang Liu; Peiwen Sun; Xueying Lee; Wei Zhang; Xue Yang; Rui Liu; Hongsheng Li; | code |
| 94 | ToMAP: Training Opponent-Aware LLM Persuaders with Theory of Mind Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Notably, while humans are skilled in modeling their opponent’s thoughts and opinions proactively and dynamically, current LLMs struggle with such Theory of Mind (ToM) reasoning, resulting in limited diversity and opponent awareness. To address this limitation, we introduce Theory of Mind Augmented Persuader (**ToMAP**), a novel approach for building more flexible persuader agents by incorporating two theory of mind modules that enhance the persuader’s awareness and analysis of the opponent’s mental state. |
Peixuan Han; Zijia Liu; Jiaxuan You; | code |
| 95 | Approximation of Log-Partition Function in Policy Mirror Descent Induces Implicit Regularization for LLM Post-Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We investigate a practical algorithm, termed PMD-mean, that approximates the log-partition term with the mean reward under the sampling policy and performs regression in log-policy space. |
Zhenghao Xu; Qin Lu; Changlong Yu; Tuo Zhao; | code |
| 96 | URS: A Unified Neural Routing Solver for Cross-Problem Zero-Shot Generalization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing neural solvers typically rely on predefined problem constraints or require per-problem fine-tuning, which substantially limits their zero-shot generalization ability to unseen VRP variants. To address this critical bottleneck, we propose URS, a unified neural routing solver that achieves zero-shot generalization across a wide range of unseen VRPs with a single model. |
Changliang Zhou; Canhong Yu; Shunyu Yao; Xi Lin; Zhenkun Wang; Yu Zhou; Qingfu Zhang; | code |
| 97 | Edit-Based Refinement for Parallel Masked Diffusion Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose ME-DLM, an edit-based refinement framework that augments diffusion generation with a lightweight post-generation editing step. |
Houxing Ren; Mingjie Zhan; Zimu Lu; Ke Wang; Yunqiao Yang; Haotian Hou; Junting Pan; Hongsheng Li; | code |
| 98 | Learning Query-Aware Budget-Tier Routing for Runtime Agent Memory Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present \textbf{BudgetMem}, a runtime agent memory framework for explicit, query-aware performance–cost control. |
Haozhen Zhang; Haodong Yue; Tao Feng; Quanyu Long; Jianzhu Bao; Bowen Jin; Weizhi Zhang; Xiao Li; Jiaxuan You; Chengwei Qin; Wenya Wang; | code |
| 99 | AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To further assess robustness beyond familiar domains, we propose AVI-Bench-PriSe, an extension that probes models’ primitive audio-visual sensation using unfamiliar, low-semantic stimuli, testing generalization beyond common training distributions. |
Yaoting Wang; Ziyi Zhang; Wenming Tu; Shaoxuan Xu; Wenjie Du; Cheng Liang; weijun wang; Yuanchao Li; Guangyao Li; Hao Fei; Yuanchun Li; Henghui Ding; Yunxin Liu; | code |
| 100 | VLANeXt: Recipes for Building Strong VLA Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: From this study, we distill 12 key findings that together form a practical recipe for building strong VLA models. |
Xiao-Ming Wu; Bin Fan; Kang Liao; Jian-Jian Jiang; Runze Yang; Yihang Luo; Zhonghua Wu; Wei-Shi Zheng; Chen Change Loy; | code |
| 101 | AdverMCTS: Combating Pseudo-Correctness in Code Generation Via Adversarial Monte Carlo Tree Search Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We argue that optimizing against a fixed, weak environment inherently limits robustness. To address this, we propose AdverMCTS, a novel adversarial Monte Carlo Tree Search framework that combats pseudo-correctness by coupling code search with active vulnerability discovery. |
Qingyao Li; Weiwen Liu; Weinan Zhang; Yong Yu; Bo An; | code |
| 102 | PixCLIP: Towards Fine-grained Vision-Language Understanding Via Any-granularity Pixel-Text Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent efforts improve textual granularity by leveraging long, detailed descriptions and replacing CLIP’s text encoder with LLM, but often overlook the visual-side bottleneck: achieving finer alignment requires region- and pixel-level visual grounding, not just finer text. To address this issue, we propose PixCLIP, a framework that jointly enhances both sides by accommodating visual prompt regions and long-form text within a unified training objective. |
YiCheng Xiao; Yu Chen; Hao-Xuan Ma; Jiale Hong; Caorui Li; Lingxiang Wu; Haiyun Guo; Jinqiao Wang; | code |
| 103 | Efficient Reasoning with Hidden Thinking Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose**Heima** (as hidden llama), an effective CoT compression framework that condenses lengthy CoTs into a small set of abstract thinking tokens, preserving essential reasoning while removing redundancy. |
Xuan Shen; Yizhou Wang; Yufa Zhou; Xiangxi Shi; Pu Zhao; Yanzhi Wang; Jiuxiang Gu; | code |
| 104 | Brep2Shape: Boundary and Shape Representation Alignment Via Self-supervised Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While deep learning shows promise in processing B-rep models, existing methods suffer from a representation gap: continuous approaches offer analytical precision but are visually abstract, whereas discrete methods provide intuitive clarity at the expense of geometric precision. To bridge this gap, we introduce Brep2Shape, a novel self-supervised pre-training framework designed to align abstract boundary representations with intuitive shape representations. |
Yuanxu Sun; Yuezhou Ma; Haixu Wu; Guanyang Zeng; Muye Chen; Jianmin Wang; Mingsheng Long; | code |
| 105 | SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose \textbf{SkillTrojan}, a backdoor attack that targets skill implementations rather than model parameters or training data. |
Yunhao Feng; Yifan Ding; Yingshui Tan; Boren Zheng; Yanming Guo; Xiaolong Li; Kun Zhai; Yishan Li; Wenke Huang; | code |
| 106 | MLLM-4D: Towards Visual-based Spatial-Temporal Intelligence Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite its importance, this capability remains a significant bottleneck for current multimodal large language models (MLLMs). To tackle this challenge, we introduce MLLM-4D, a comprehensive framework designed to bridge the gaps in training data curation and model post-training for spatiotemporal understanding and reasoning. |
Xingyilang Yin; Chengzhengxu Li; Jiahao Chang; Chi-Man Pun; Xiaodong Cun; | code |
| 107 | GSFixer: Improving 3D Gaussian Splatting with Reference-Guided Video Diffusion Priors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While recent approaches have sought to leverage generative priors to complete information for under-constrained regions, they struggle to generate content that remains consistent with input observations. To address this challenge, we propose GSFixer, a novel framework designed to improve the quality of 3DGS representations reconstructed from sparse inputs. |
Xingyilang Yin; Qi Zhang; Jiahao Chang; Ying Feng; Qingnan Fan; Xi Yang; Chi-Man Pun; Huaqi Zhang; Xiaodong Cun; | code |
| 108 | A Pure Hierarchical Spectral Parcellation Network for Brain Network Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, this model is constructed as a hierarchy of Spectral Parcellation blocks. |
Jiaming Zhuo; Shuai Zhai; Ziyi Ma; Kun Fu; Chuan Wang; Di Jin; Zhen Wang; Xiaochun Cao; Huazhu Fu; Liang Yang; | code |
| 109 | Is Code Better Than Language for Algorithmic Reasoning? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Comparing NL reasoning and solver-based pipelines directly is ill-posed: they differ simultaneously in representation space and execution mechanism. We introduce a three-route framework that makes this comparison tractable by introducing an intermediary step—code generation with LLM-based execution. |
Terry Tong; Yu Feng; Surbhi Goel; Dan Roth; | code |
| 110 | Meta Context Engineering Via Agentic Skill Evolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: They impose structural biases and restrict context optimization to a narrow, intuition-bound design space. To address this, we introduce Meta Context Engineering (MCE), a bi-level framework that supersedes static CE heuristics by co-evolving CE skills and context artifacts. |
Haoran Ye; Xuning He; Vincent Arak; Haonan Dong; Guojie Song; | code |
| 111 | Beyond Majority Voting: Self-Reflective Test-Time Reinforcement Learning for LLM Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As a result, rare yet correct trajectories are systematically undervalued by majority-voting-based approaches. To address this limitation, we propose Self-Reflective Test-Time Reinforcement Learning (SR-TTRL), a novel framework that leverages self-reflective verification to produce high-fidelity pseudo-labels. |
Sitong Wu; Haoru Tan; Xichen Zhang; Bin Xia; Shaofeng Zhang; XIAOJUAN QI; Bei Yu; Jiaya Jia; | code |
| 112 | Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Thought-Aligner, a lightweight plug-in safety model that performs causal correction on unsafe thoughts before action execution, without altering the underlying agent. |
Changyue Jiang; Wenqi Zhang; Xudong Pan; Geng Hong; Min Yang; | code |
| 113 | Revisiting Robustness for LLM Safety Alignment Via Selective Geometry Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we revisit robustness for safety alignment from an optimization geometry perspective, highlighting optimization-induced fragility as a complementary factor to data-space uncertainty. |
Yonghui Yang; WenjianTao; Jilong Liu; Xingyu Zhu; Junfeng Fang; Huang Weibiao; Le Wu; Richang Hong; Tat-Seng Chua; | code |
| 114 | STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This trilemma stems from a critical \textit{topological mismatch}: the prevailing isotropic Gaussian prior in standard VAEs imposes a \textit{flat} latent geometry that fails to accommodate audio’s \textit{hierarchical} nature, where low-frequency components are structured and compressible while high-frequency components are stochastic and incompressible, leading to \textit{disordered information packing} where crucial semantic features are randomly interleaved with high-entropy noise. To resolve this challenge, we propose \textbf{Structured Topology-Aware Regularization (STAR)}, a general training strategy that reshapes latent space geometry by imposing a growth-based constraint field, routing structural and textural information into channel subspaces with matching capacities. |
Huadai Liu; Wen Wang; Kaicheng Luo; Qian Chen; Xiangang Li; Wei Xue; | code |
| 115 | NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent advances in coding agents suggest rapid progress toward autonomous software development, yet existing benchmarks primarily evaluate short-horizon behaviors such as localized code generation, scaffolded completion, or repository repair, leaving it unclear whether agents can sustain coherent reasoning, planning, and execution over the extended horizons demanded by real-world repository construction. To address this gap, we introduce NL2Repo-Bench, a benchmark explicitly designed to evaluate the long-horizon repository generation from scratch: given only a single natural-language requirements document and an empty workspace, agents must autonomously design the architecture, manage dependencies, and produce a fully installable Python library. |
Jingzhe Ding; Shengda Long; puchangxin; Ge Zhang; zhou huan; Hongwan Gao; Xiang Gao; Chao He; Yue Hou; FEI HU; Zhaojian Li; Weiran Shi; Zaiyuan Wang; Daoguang Zan; Chenchen Zhang; Xiaoxu Zhang; Chen Qizhi; cheng; Bo Deng; Qingshui Gu; Kai Hua; Juntao Lin; Pai Liu; Mingchen Li; Minghao Li; Xuanguang Pan; Zifan Peng; Yujia Qin; Yong Shan; Zhewen Tan; Haoran Wang; Zihan Wang; Weihao Xie; Yishuo Yuan; Jiayu Zhang; Yunfei Zhao; He Zhu; LIYA ZHU; chenyangzou; Ming Ding; Jiaheng Liu; Jianpeng Jiao; Minghao Liu; Qian Liu; Chongyang Tao; Jian Yang; Tong Yang; Zhaoxiang Zhang; Xinjie Chen; Wenhao Huang; | code |
| 116 | Decoupled Low-Rank Adaptation for Robust Federated Fine-Tuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we investigate the structural properties of LoRA and reveal a robustness asymmetry. |
Xiuwen Fang; Xuliang Yang; Mang Ye; | code |
| 117 | Calibrated Multimodal Representation Learning with Missing Modalities Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent research generalizes traditional cross-modal alignment to produce enhanced multimodal synergy but requires all modalities to be present for a common instance, making it challenging to utilize prevalent datasets with missing modalities. We provide theoretical insights into this issue from an anchor shift perspective. |
Xiaohao Liu; Xiaobo Xia; Jiaheng Wei; Shuo Yang; Xiu Su; See-Kiong Ng; Tat-Seng Chua; | code |
| 118 | DLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this specific challenge, our work begins with a key observation that dLLM inference involves a static prompt and a partially dynamic response, where most tokens remain stable across adjacent denoising steps. Based on this, we propose dLLM-Cache, a training-free adaptive caching framework that combines long-interval prompt caching with partial response updates guided by feature similarity. |
Zhiyuan Liu; Yicun Yang; Yaojie Zhang; Junjie Chen; Chang Zou; Qingyan Wei; Shaobo Wang; Yichen Zhu; Linfeng Zhang; | code |
| 119 | TFRBench: A Reasoning Benchmark for Evaluating Forecasting Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Unlike existing benchmarks, TFRBench provides a protocol for evaluating the reasoning generated by forecasting systems–specifically their analysis of cross-channel dependencies, trends, and external events. To enable this, we propose a systematic multi-agent framework that utilizes an iterative verification loop to synthesize numerically grounded reasoning traces. |
Atik Ahamed; Mihir Parmar; Palash Goyal; Yiwen Song; Long Le; Qiang (Shaun) Cheng; Chun-Liang Li; Hamid Palangi; Jinsung Yoon; Tomas Pfister; | code |
| 120 | Causal-JEPA: Learning World Models Through Object-Level Latent Interventions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While object-centric representations provide a useful abstraction, they are not sufficient to capture interaction-dependent dynamics. We therefore propose C-JEPA, a simple and flexible object-centric world model that extends masked joint embedding prediction from image patches to object-centric representations. |
Heejeong Nam; Quentin Le Lidec; Lucas Maes; Yann LeCun; Randall Balestriero; | code |
| 121 | Learning to Decode Against Compositional Hallucination in Video Multimodal Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose **TriCD**, a contrastive decoding framework with a triple-pathway calibration mechanism. |
Wenbin Xing; Quanxing Zha; Lizheng Zu; Mengran Li; Ming Li; Junchi Yan; | code |
| 122 | AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our comprehensive study shows that existing memory systems underperform on AMA Bench primarily because they suffer from a loss of causality and objective information, and are constrained by the lossy nature of similarity based retrieval employed by many memory systems. To address these limitations, we propose AMA Agent, an effective memory system featuring a causality graph and tool augmented retrieval. |
Yujie Zhao; Boqin Yuan; Junbo Huang; Haocheng Yuan; Zhongming Yu; Haozhou Xu; Lanxiang Hu; Abhilash Shankarampeta; Zimeng Huang; Wentao Ni; Yuandong Tian; Jishen Zhao; | code |
| 123 | SAMT: Generating Structured Avatar Meshes and Textures from A Single Image Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To reconstruct facial micro-structures and fine-grained multiview-consistent textures, this work presents a two-stage framework named SAMT for monocular 3D avatar generation and texture synthesis. |
Muyu Wang; Xingping Dong; Jianzhe Gao; Wenguan Wang; Yujia Wang; | code |
| 124 | Can Recommender Systems Teach Themselves? A Recursive Self-Improving Framework with Fidelity Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This challenge is particularly acute in recommendation systems, where extreme sparsity in user interactions leads to rugged optimization landscapes and poor generalization. We propose the Recursive Self-Improving Recommendation (RSIR) framework, a paradigm in which a model bootstraps its own performance without reliance on external data or teacher models. |
Luankang Zhang; Hao Wang; Zhongzhou Liu; MINGJIA YIN; Yonghao Huang; Jiaqi Li; Wei Guo; Yong Liu; Huifeng Guo; Defu Lian; Enhong Chen; | code |
| 125 | Gradient-Based Causal Tree Ensembles: A Backbone Architecture for Heterogeneous Treatment Effects Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce **GRA**dient-based **C**ausal tree **E**nsembles (GRACE), a novel tree-based architecture for HTE estimation that incorporates multi-way, oblique, and soft splits, enabling end-to-end training via backpropagation. |
Yusuke Kano; Jeremy P Voisey; Mihaela van der Schaar; | code |
| 126 | Geometric Decoupling: Diagnosing The Structural Instability of Latent Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Latent Diffusion Models (LDMs) achieve high-fidelity synthesis but suffer from latent space brittleness, causing discontinuous semantic jumps during editing. We introduce a Riemannian framework to diagnose this instability by analyzing the generative Jacobian, decomposing geometry into *Local Scaling* (capacity) and *Local Complexity* (curvature). |
Yuanbang Liang; Zhengwen Chen; Yu-Kun Lai; | code |
| 127 | Sufficiency Is Relative: Evaluating LLM Explanations Under Model-Induced Input Distributions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We generalize classical sufficiency from feature attributions to arbitrary explanations and prove that explanation sufficiency is inherently relative to an input distribution, which must be explicitly defined for LLM explanations. We propose using the LLM itself to generate alternative inputs conditioned on an explanation, capturing its beliefs about possible inputs. |
Nhi Nguyen; Shauli Ravfogel; Rajesh Ranganath; | code |
| 128 | Are Your Agents Upward Deceivers? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We observe and define \textit{\textbf{agentic upward deception}}, a phenomenon in which an agent facing environmental constraints conceals its failure and performs actions that were not requested without reporting. To assess its prevalence, we construct a benchmark of 200 tasks covering five task types and eight realistic scenarios in a constrained environment, such as broken tools or mismatched information sources. |
Dadi Guo; Qingyu Liu; Dongrui Liu; Qihan Ren; Shuai Shao; Tianyi Qiu; Haoran Li; Yi Fung; Zhongjie Ba; Juntao Dai; Jiaming Ji; Zhikai Chen; Jialing Tao; Yaodong Yang; Jing Shao; Xia Hu; | code |
| 129 | Med-Scout: Curing MLLMs’ Geometric Blindness in Medical Perception Via Geometry-Aware RL Post-Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This failure to ground outputs in objective geometric constraints leads to plausible yet factually incorrect hallucinations, rooted in training paradigms that prioritize linguistic fluency over geometric fidelity. This paper introduces Med-Scout, a novel framework that “cures” this blindness via Reinforcement Learning (RL) that leverages the intrinsic geometric logic latent within unlabeled medical images. |
Anglin Liu; Ruichao Chen; Yi Lu; Hongxia Xu; Jintai Chen; | code |
| 130 | InfraRL: A Benchmark for Constrained Resource Allocation in Large-Scale Infrastructure Asset Management Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce InfraRL, a high-fidelity benchmark that uses bridge maintenance as a rigorous testbed for general infrastructure asset management challenges. |
Yantian Wang; Wenhao Li; Bo Jin; | code |
| 131 | AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Compared to traditional audio processing tasks, audio stories introduce new layers of semantic, temporal, and physical complexity. To address this challenge, we propose AudioChat, a framework for developing audio foundation models that can generate, edit, and understand audio stories. |
William Chen; Prem Seetharaman; Rithesh Kumar; Oriol Nieto; Shinji Watanabe; Justin Salamon; Zeyu Jin; | code |
| 132 | EEmo-Logic: A Unified Dataset and Multi-Stage Framework for Comprehensive Image-Evoked Emotion Assessment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing models are still limited to coarse-grained emotion perception or deficient reasoning capabilities. To bridge this gap, we introduce **EEmoDB**, the largest image-evoked emotion understanding dataset to date. |
Lancheng Gao; Ziheng Jia; Zixuan Xing; Wei Sun; Huiyu Duan; Guangtao Zhai; Xiongkuo Min; | code |
| 133 | SSL4RL: Revisiting Self-supervised Learning As Intrinsic Reward for Visual-Language Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although reinforcement learning (RL) can align models with desired behaviors, its application to VLMs has been hindered by the lack of scalable and reliable reward mechanisms. To overcome this challenge, we propose **SSL4RL**, a novel framework that leverages self-supervised learning (SSL) tasks as a source of verifiable rewards for RL-based fine-tuning. |
Xiaojun Guo; Runyu Zhou; Yifei Wang; Qi Zhang; Chenheng Zhang; Stefanie Jegelka; Xiaohan Wang; Jiajun Chai; Guojun Yin; Wei Lin; Yisen Wang; | code |
| 134 | GRASP: Graph Reasoning Via Agentic Solving and Probing of LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Integrating graph knowledge into Large Language Models (LLMs) via passive representation faces critical bottlenecks: limited context windows, unreliable numerical computation, and structural hallucinations. To solve this, we propose **GRASP** (Graph Reasoning via Agentic Solving and Probing), shifting the paradigm from passive ingestion to proactive agentic exploration. |
Xiaojun Guo; Mingxue Tian; Chenheng Zhang; Xiaohan Wang; Jiajun Chai; Guojun Yin; Wei Lin; Yifei Wang; Yisen Wang; | code |
| 135 | AutoRAS: Learning Robust Agentic Systems with Primitive Representations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose AutoRAS, a framework for the Automated design of Robust Agentic Systems. |
Yang Yue; Xuancheng Zhu; YuYang Ma; Guoshun Nan; Zihan Dou; JingRu Shan; Congyu Guo; Ji Zhang; Hua Wang; Jingfeng Zhang; | code |
| 136 | What Makes Effective Supervision in Latent Chain-of-Thought: An Information-Theoretic Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: And we expose a divergence in alignment strategies: rigid Geometric Compression acts as a destructive prior that collapses the reasoning manifold, whereas Generative Reconstruction serves as a flexible semantic tether, optimizing for reconstructibility to preserve the intrinsic dimensionality of the latent space. To quantify these dynamics, we introduce the Unified Latent-MI Probe (ULP), which unveils a strict Information-Performance Binding: reasoning accuracy is deeply correlated with the mutual information retained in the latent chain. |
Xinghao Chen; Chak Tou Leong; Guo Wenjin; Jian Wang; Wenjie Li; Anhao Zhao; | code |
| 137 | OmniSapiens: A Foundation Model for Social Behavior Processing Via Heterogeneity-Aware Relative Policy Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent reasoning RL methods facilitate training a single unified model across multiple behavioral tasks, but do not explicitly address learning across different heterogeneous behavioral data. To address this gap, we introduce Heterogeneity-Aware Relative Policy Optimization (HARPO), a RL method that balances leaning across heterogeneous tasks and samples. |
Keane Ong; Sabri Boughorbel; Luwei Xiao; Chanakya Ekbote; David Dai; Ao Qu; Jingyao Wu; Rui Mao; Ehsan Hoque; Erik Cambria; Gianmarco Mengaldo; Paul Pu Liang; | code |
| 138 | Memory-Efficient LLM Pretraining Via Minimalist Optimizer Design Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While recent works such as GaLore, Fira and APOLLO have proposed state-compressed variants to reduce memory consumption, a fundamental question remains: What are the minimum modifications to plain SGD needed to match state-of-the-art pretraining performance? We systematically investigate this question using a bottom-up approach, and identify two simple yet highly (memory- and compute-) efficient techniques: (1) column-wise gradient normalization (normalizing the gradient along the output dimension), which boosts SGD performance without momentum; and (2) applying first-order momentum only to the output layer, where gradient variance is highest. |
Athanasios Glentis; Jiaxiang Li; Andi Han; Mingyi Hong; | code |
| 139 | ExpWeaver: LLM Agents Learn from Experience Via Latent RAG Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing methods remain confined to explicit text space—retrieving experiences via semantic similarity and concatenating them into the context window, leading to substantial token overhead and a decoupled architecture that separates retrieval from generation. To address these limitations, we propose \method, a framework that enables LLM agents to learn from experience via latent retrieval-augmented generation, without requiring a separate RAG module. |
Tao Feng; Tianyang Luo; Jingjun Xu; Zhigang Hua; Yan Xie; Shuang Yang; Ge Liu; Jiaxuan You; | code |
| 140 | Adaptive Code Watermarking Through Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce CodeTracer, an innovative adaptive code watermarking framework underpinned by a reinforcement learning training paradigm. |
Zhimeng Guo; Huaisheng Zhu; Siyuan Xu; Hangfan Zhang; Teng Xiao; Minhao Cheng; | code |
| 141 | OneSearch: A Preliminary Exploration of The Unified End-to-End Generative Framework for E-commerce Search Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose OneSearch, the first industrial-deployed end-to-end generative framework for e-commerce search, featuring three key innovations: (1) Keyword-enhanced Hierarchical Quantization Encoding (KHQE) to preserve hierarchical semantics and distinctive item attributes while maintaining strong query-item relevance constraints; (2) multi-view user behavior sequence injection that constructs behavior-driven user IDs and incorporates both explicit short-term and implicit long-term sequences; and (3) a Preference-Aware Reward System (PARS) with multi-stage supervised fine-tuning and adaptive reward-weighted ranking to capture fine-grained user preferences. |
Ben Chen; Xian Guo; Siyuan Wang; Zihan Liang; Yufei Ma; Yue Lv; Chenyi Lei; Yuqing DING; Wenwu Ou; Han Li; Kun Gai; | code |
| 142 | Seg-ReSearch: Segmentation with Interleaved Reasoning and External Search Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose \textbf{Seg-ReSearch}, a novel segmentation paradigm that overcomes the knowledge bottleneck of existing approaches. |
Tianming Liang; Qirui Du; Jian-Fang Hu; Haichao Jiang; Zicheng Lin; Wei-Shi Zheng; | code |
| 143 | Learning Unmasking Policies for Diffusion Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we instead propose to train sampling procedures using reinforcement learning. |
Metod Jazbec; Theo X. Olausson; Louis Béthune; Pierre Ablin; Michael Kirchhof; Joao Monteiro; Victor Guilherme Turrisi da Costa; Jason Ramapuram; Marco Cuturi; | code |
| 144 | Unifying Heterogeneous Degradations: Uncertainty-Aware Diffusion Bridge Model for All-in-One Image Restoration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods are often constrained by coarse-grained control mechanisms or fixed mapping schedules, yielding suboptimal adaptation. To address this, we propose an Uncertainty-Aware Diffusion Bridge Model (UDBM), which innovatively reformulates AiOIR as a stochastic transport problem steered by pixel-wise uncertainty. |
Luwei Tu; Jiawei Wu; Xing Luo; Zhi Jin; | code |
| 145 | Reranker Helps, But Not Enough: Towards Strong Poisoning Attacks Against Retrieval-Augmented Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Towards realistic RAG red-teaming, we conclude practical prompt design principles that reveal reranker blind spots. Building on these insights, we introduce the Prompt-Perturbation Poisoning Attack ($\mathbf{P}^3 \mathbf{A}$). |
Xiaokun Yang; Yesheng Liu; Xin Xiong; Jian Liang; Ran He; Tieniu Tan; | code |
| 146 | USE : A Unified Self-Ensembling Framework for Test-Time Prompt Tuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Among existing CLIP-based TTA methods, Test-Time Prompt Tuning (TPT) is a pioneering work that optimizes textual prompts using multiple test-time augmentations and remains a strong baseline to date. In this work, we revisit TPT and reveal that its optimization can be interpreted as implicitly learning from self-generated pseudo labels. |
Siru Jiang; Jian Liang; Ran He; Tieniu Tan; | code |
| 147 | OnePO: Direct One-stage Policy Optimization for SFT-free Domain Adaptation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: * We argue that pre-SFT is inherently problematic: (1) it indiscriminately reinforces knowledge and behaviors from references regardless of whether the LLM has already acquired them, leading to distribution contraction that constrains subsequent exploration; (2) it introduces substantial overhead in multi-stage training and data curation. |
Junying Chen; Xinyuan Xie; Ziniu Li; Benyou Wang; | code |
| 148 | LIMSSR: LLM-Driven Sequence-to-Score Reasoning Under Training-Time Incomplete Multimodal Observations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper tackles the more challenging setting of IML under training-time incomplete observations, which precludes reliance on a God’s eye view of complete data. We propose LIMSSR (LLM-Driven Incomplete Multimodal Sequence-to-Score Reasoning), a framework that reformulates this challenge as a conditional sequence reasoning task. |
Huangbiao Xu; huanqi wu; Xiao Ke; Yuxin Peng; | code |
| 149 | PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The remarkable success of Chain-of-Thought (CoT), which enhances performance by scaling generation steps at test-time, inspires us to ask: can we leverage a similar scaling of computational steps during pretraining to improve the generation of each individual token? To address this, we propose a novel pre-training methodology: Pretraining Language Models with Latent Thoughts (PonderLM-2). |
Boyi Zeng; He Li; Shixiang Song; Yixuan Wang; Zitong Wang; Ziwei He; Xinbing Wang; Zhouhan Lin; | code |
| 150 | Efficient Prediction of SO(3)-Equivariant Hamiltonian Matrices Via SO(2) Local Frames Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Motivated by the inherent relationship between the off-diagonal blocks of the Hamiltonian matrix and the SO(2) local frame, we propose a novel and efficient network, called QHNetV2, that achieves global SO(3) equivariance without the costly SO(3) Clebsch–Gordan tensor products. |
Haiyang Yu; Yuchao Lin; Xuan Zhang; Xiaofeng Qian; Shuiwang Ji; | code |
| 151 | Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work investigates a fundamental research question: Can MLLMs recover corrupted visual content by themselves? To address this, we propose Robust-U1, a novel framework that equips MLLMs with explicit visual self-recovery capability for robust understanding. |
Jiaqi Tang; Jianmin Chen; Youyang Zhai; Wei Wei; Runtao Liu; Mengjie Zhao; Xiangyu Wu; Qingfa Xiao; Qifeng Chen; | code |
| 152 | Temporal-aware Flow Matching for Video Generation with Temporally Coherent Motion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This practice fails to explicitly model motion priors or temporal dependencies, resulting in suboptimal dynamics that may appear incoherent and unrealistic. To solve this problem, we propose Temporal-aware Flow Matching (TFM), a novel training paradigm that embeds inter-frame constraints into the flow objective, leading to temporally coherent motion modeling in video generation. |
Zirui Pan; Xin Wang; Yipeng Zhang; Yuwei Zhou; Wenwu Zhu; | code |
| 153 | MedSIGHT: Towards Grounded Visual Comprehension in Medical Large Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present MedSIGHT, a unified framework that equips Med-LVLMs with structured, pixel-level understanding for grounded visual comprehension. |
Aofei Chang; Le Huang; Alex Boyd; parminder bhatia; Taha Kass-Hout; Fenglong Ma; Cao Xiao; | code |
| 154 | Beyond Procedure: Substantive Fairness in Conformal Prediction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To facilitate scalable empirical analysis, we introduce an LLM-in-the-loop evaluator that approximates human assessment of substantive fairness across diverse modalities. |
Pengqi Liu; Zijun Yu; Mouloud Belbahri; Arthur Charpentier; Masoud Asgharian; Jesse Cresswell; | code |
| 155 | PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text Under Semantic-Invariant Attacks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose PASA, a principled, robust, and distortion-free watermarking algorithm that embeds and detects a watermark at the semantic level. |
Zhenxin Ai; Haiyun He; | code |
| 156 | Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Across math and reasoning benchmarks, we find collapse is preceded by sharp drops in effective sample size (ESS) and unstable gradient norms. Motivated by this diagnosis, we propose **V**ariance **C**ontrolled **P**olicy **O**ptimization (VCPO), a drop-in stabilization method for REINFORCE/GRPO-style algorithms that (i) **rescales learning rate according to effective sample size** to dampen unreliable updates, and (ii) applies a **closed-form minimum-variance baseline** for the off-policy setting, avoiding an auxiliary value model and adding minimal overhead. |
Luke Huang; Zhuoyang Zhang; Qinghao Hu; Shang Yang; Song Han; | code |
| 157 | Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce alignment tampering, a potential vulnerability where the LLM undergoing alignment influences the preference dataset, causing RLHF to amplify undesired behaviors. |
Dongyoon Hahm; Dylan Hadfield-Menell; Kimin Lee; | code |
| 158 | Structured 4D Latent World Model for Robot Planning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a **Structured 4D Latent World Model**, which predicts the evolution of a scene’s 3D structure in a structured latent space conditioned on observations and textual instructions. |
Zhiyi Li; Peilin Wu; Xiaoshen Han; Ruojin Cai; Yilun Du; | code |
| 159 | Scaling The Prior: Size-Consistent Geometric Diffusion for 3D Molecular Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We are the first to identify and analyze this size-induced inconsistency by decomposing denoising dynamics, showing how spatial scale shapes formation of both 3D structure and atom types. Based on this, we propose Scaling the Prior (StP), which rescales the prior distribution by molecular size to normalize learning and generation across sizes, harmonize denoising trajectories, and enable consistently high-quality molecules. |
Wenhan Gao; Jingxiang Qu; Yi Liu; | code |
| 160 | Benchmarking and Evolving Reason-Reflect-Rectify for Reflective Visual Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Evaluation on R$^3$-Bench reveals a critical gap: while state-of-the-art models can identify generation errors, they fail to generate actionable rectification instructions. To bridge this gap, we propose R$^3$-Refiner, a dual-stage framework leveraging Group Relative Policy Optimization (GRPO) and a Hierarchical Reward Mechanism (HRM) to better align rectification with relective reasoning. |
Junjie Wang; 星华 娄; Xiangtai Li; Ye Tian; Keyu Chen; Yulin Li; Bin Kang; Guangcan Mai; Yanwei Li; Zhuotao Tian; Liqiang Nie; | code |
| 161 | OXE-AugE: A Large-Scale Robot Augmentation of OXE for Scaling Cross-Embodiment Policy Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present AugE-Toolkit, a scalable robot augmentation pipeline, and OXE-AugE, a high-quality open-source dataset that augments OXE with 9 different robot embodiments. |
Guanhua Ji; Harsha Polavaram; Lawrence Yunliang Chen; Sandeep Bajamahal; Zehan Ma; Simeon Adebola; Chenfeng Xu; Ken Goldberg; | code |
| 162 | Trust3R: Unifying Feed-Forward Pointmap Prediction and Evidential Learning for Trust-Aware 3D Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Geometric foundation models hold promise for unconstrained dense geometry prediction from uncalibrated images; however, in current feed-forward designs, their predicted confidence scores are heuristic, lack probabilistic interpretation, and often fail to indicate where and how much the predicted geometry can be trusted. To fill this gap, we present ***Trust3R***, a trust-aware 3D reconstruction framework that pairs a lightweight gated residual mean refinement with evidential learning to predict pointmap evidence under a Normal-Inverse-Wishart prior and yield a closed-form multivariate Student-t predictive distribution. |
Zihao Zhu; Wenyuan Zhao; Nuo Chen; Chao Tian; Zhiwen Fan; | code |
| 163 | ForesightKV: Optimizing KV Cache Eviction for Reasoning Models By Learning Long-Term Contribution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To better balance efficiency and performance, we introduce ForesightKV, a training-based KV cache eviction framework that learns to predict which KV pairs to evict during long-text generations. |
Zican Dong; Peiyu Liu; Junyi Li; Zhipeng Chen; Han Peng; Shuo Wang; Xin Zhao; | code |
| 164 | RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This limits their systematic understanding, comparison, and progress measurement. To address these challenges, we introduce **RoboMME**: a large-scale standardized benchmark for evaluating and advancing VLA models in long-horizon, history-dependent scenarios. |
Yinpei Dai; Hongze Fu; Jayjun Lee; Yuejiang Liu; Haoran Zhang; Jianing Yang; Chelsea Finn; Nima Fazeli; Joyce Chai; | code |
| 165 | TD3B: Transition-Directed Discrete Diffusion for Allosteric Binder Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Structure-based design methods optimize binding to static conformations and cannot represent non-reversible, directional effects or systematically distinguish agonist from antagonist behavior. To address this gap, we introduce **T**ransition-**D**irected **D**iscrete **D**iffusion for allosteric **B**inder design (**TD3B**), a sequence-based generative framework that designs binders with specified agonist or antagonist behavior via a directional transition control objective. |
Hanqun CAO; Aastha Pal; Sophia Tang; Yinuo Zhang; Jingjie Zhang; Pheng Ann Heng; Pranam Chatterjee; PhD; | code |
| 166 | Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This study introduces Generalizable Predictive Prompt Selection (GPS), which performs Bayesian inference towards prompt difficulty using a lightweight generative model trained on the shared optimization history. |
Yun Qu; Cheems Wang; Yixiu Mao; Heming Zou; Yuhang Jiang; Weijie Liu; Clive Bai; Kai Yang; Yangkun Chen; Saiyong Yang; Xiangyang Ji; | code |
| 167 | RADAR: Redundancy-Aware Diffusion for Multi-Agent Communication Structure Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This restricts fine-grained structural exploration and flexible composition, leading to excessive token consumption for simple tasks or performance bottlenecks for complicated ones. To address this challenge, we introduce RADAR, a redundancy-aware and query-adaptive generative framework that actively reduce communication overhead. |
Zhen Zhang; Wanjing Zhou; Juncheng Li; Hao Fei; Jun Wen; Wei Ji; | code |
| 168 | GEPC: Group-Equivariant Posterior Consistency for Out-of-Distribution Detection in Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Group-Equivariant Posterior Consistency (GEPC), a training-free probe that measures how consistently the learned score transforms under a finite group $G$, detecting equivariance breaking even when score magnitude remains unchanged. |
Rouzoumka Alexis; Jean Pinsolle; Eugénie TERREAUX; christele morisseau; Jean-Philippe Ovarlez; Chengfang Ren; | code |
| 169 | Generative Neural Operators Through Diffusion Last Layer Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, many practical systems are inherently stochastic, making principled uncertainty quantification essential for reliable deployment. To address this, we introduce a simple add-on, the *diffusion last layer* (DLL), a lightweight probabilistic head that can be attached to arbitrary neural operator backbones to model predictive uncertainty. |
Sungwon Park; Anthony Zhou; Hongjoong Kim; Amir Barati Farimani; | code |
| 170 | FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Evaluations on 13 omni-modal and 7 video-only models show that current systems struggle with audio-visual future prediction, particularly in speech-heavy scenarios, with the best accuracy of 64.8% achieved by Gemini 3 Flash. To mitigate this limitation, we curate a 7K-sample instruction-tuning dataset and propose an Omni-Modal Future Forecasting (OFF) training strategy. |
Qian Chen; Jinlan Fu; Changsong Li; Min zhang; See-Kiong Ng; Xipeng Qiu; | code |
| 171 | RVAS: Referring Video Active Exploration and Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We benchmark representative RVOS and related video understanding baselines and find that they struggle to perform active target search and incur substantial overhead when coupled with online decision making. Motivated by these challenges, we propose LESA, a baseline framework that introduces a state controller and hierarchical memory for efficient streaming processing and sparse MLLM reasoning. |
Hengrui Hu; Weiwei Gao; Zipei Zhang; Henghui Ding; | code |
| 172 | Rethinking Human Intent to CAD: Parametric CAD Model Generation Via Cooperative Multi-Task Alignment and Spatial-Aware Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To support our study, we construct HiCAD, the first large-scale dataset aligning hand-drawn sketches, textual descriptions, and parametric CAD codes. Based on this, we introduce HiCAD, a two-stage framework comprising Cooperative Multi-Task Alignment to bridge the representational gap between heterogeneous inputs, and Spatial-Aware Reinforcement Learning to enforce geometric and topological consistency. |
Qingwang Zhang; Jiahao Li; Xiangdong Zhou; | code |
| 173 | Hybrid-Gym: Training Coding Agents to Generalize Across Tasks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we aim to train coding agents that generalize across tasks. |
Yiqing Xie; Emmy Liu; Gaokai Zhang; Nachiket Kotalwar; Shubham Gandhi; Acharya; Xingyao Wang; Carolyn Rose; Graham Neubig; Daniel Fried; | code |
| 174 | CALM Before The STORM: Unlocking Native Reasoning for Optimization Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To fully leverage LRMs’ inherent reasoning abilities, we propose **CALM** (*Corrective Adaptation with Lightweight Modification*), a framework that progressively refines LRMs within their native reasoning modes for optimization modeling tasks. |
Zhengyang Tang; Zihan Ye; Chenyu Huang; Xuhan Huang; Chengpeng Li; Sihang Li; Guanhua CHEN; Ming Yan; Zizhuo Wang; Hongyuan Zha; Dayiheng Liu; Benyou Wang; | code |
| 175 | TINNs: Time-Induced Neural Networks for Solving Time-Dependent PDEs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose *Time-Induced Neural Networks (TINNs)*, a novel architecture that parameterizes the network weights as a learned function of time, allowing the effective spatial representation to evolve over time while maintaining shared structure. |
Chen-Yang Dai; Che-Chia Chang; Te-Sheng Lin; Ming-Chih Lai; Chieh-Hsin Lai; | code |
| 176 | Learning from Fine-Grained Visual Discrepancies: Mitigating Multimodal Hallucinations Via In-Context Visual Contrastive Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose In-Context Visual Contrastive Optimization (IC-VCO). |
Haolin Deng; Xin Zou; Zhiwei Jin; Chen Chen; Haonan Lu; Xuming Hu; | code |
| 177 | DiscoForcing: A Unified Framework for Real-Time Audio-Driven Character Control with Diffusion Forcing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce DiscoForcing, a streaming audio-driven diffusion framework that combines a causal music encoder that captures rhythmic structure and phase dynamics with a diffusion-forcing sequence model trained under heterogeneous noise levels across the temporal horizon. |
Kaiyang Ji; Bingsheng Qian; Binghuan Wu; Kangyi Chen; Ye Shi; Jingya Wang; | code |
| 178 | Hard Labels In! Rethinking The Role of Hard Labels in Mitigating Local Semantic Drift Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We theoretically analyze the emergence of drift under sparse soft-label supervision and demonstrate that hybridizing hard and soft labels restores alignment between visual content and semantic supervision. Building on this insight, we propose a new training paradigm, $\textbf{H}$ard Label for $\textbf{A}$lleviating $\textbf{L}$ocal Semantic $\textbf{D}$rift (HALD), which uses hard labels as intermediate corrective signals while preserving the fine-grained benefits of soft labels. |
Jiacheng Cui; Bingkui Tong; Xinyue Bi; Xiaohan Zhao; Jiacheng Liu; Zhiqiang Shen; | code |
| 179 | Light Forcing: Accelerating Autoregressive Video Diffusion Via Sparse Attention Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While existing sparse attention solutions have shown promise on bidirectional models, we identify that applying these solutions to AR models leads to considerable performance degradation for two reasons: isolated consideration of chunk generation and insufficient utilization of past informative context. Motivated by these observations, we propose \textsc{Light Forcing}, the \textit{first} sparse attention solution tailored for AR video generation models. |
Chengtao Lv; Yumeng Shi; Yushi Huang; Ruihao Gong; Shen Ren; Wenya Wang; | code |
| 180 | Channel Adapter for Time Series Foundation Models in Zero-Shot Multivariate Forecasting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, most TSFMs rely on channel-independent pre-training that models each variable separately, limiting their ability to leverage inter-channel information that is crucial in real-world multivariate systems. Motivated by this limitation, we propose Chada, a lightweight plug-and-play channel adapter that allows frozen TSFMs to leverage multivariate correlations in a zero-shot setting. |
Dongyuan Li; Renhe Jiang; Shun Zheng; Zheng Dong; Haotian Gao; Ying Zhang; Jiang Bian; | code |
| 181 | XKV: Cross-Layer KV-Cache Compression Via Aligned Singular Vector Extraction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We show, via Centered Kernel Alignment (CKA), that the dominant singular vectors of KV-Cache are well aligned across layers. Motivated by this observation, we propose xKV, a post-training compression method that jointly factorizes grouped-layer KV-Cache into a shared low-rank subspace, substantially reducing KV-Cache memory. |
Chi-Chih Chang; Wei-Cheng Lin; Chien-Yu Lin; Hung-Yueh Chiang; Yash Akhauri; Xilai Dai; Huiqiang Jiang; Yucheng Li; Kai-Chiang Wu; Luis Ceze; Mohamed Abdelfattah; | code |
| 182 | MaMa: A Game-Theoretic Approach for Designing Safe Agentic Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we study the automated design of agentic systems that remain safe even when a subset of agents is compromised. |
Jonathan Nöther; Adish Singla; Goran Radanovic; | code |
| 183 | Distillation Models Are Good Samplers for Diffusion Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present DMSampler, a framework that accelerates diffusion reinforcement learning by using fast distillation models as its training-time sampling engine. |
Zunxu Liu; Aiqiu Wu; Zhaofan Qiu; Yingwei Pan; Ting Yao; Tao Mei; | code |
| 184 | Diffusion Bridge or Flow Matching? A Unifying Framework and Comparative Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Simultaneously, from the perspective of Optimal Transport, interpolation coefficients $t$ and $1-t$ of Flow Matching become increasingly ineffective when the training data size is reduced. To corroborate these theoretical claims, we propose a novel, powerful architecture for Diffusion Bridge built on a latent Transformer, and implement a Flow Matching model with the same structure to enable a fair performance comparison in various experiments. |
Kaizhen Zhu; Mokai Pan; Zhechuan Yu; Jingya Wang; Jingyi Yu; Ye Shi; | code |
| 185 | Prism: Spectral-Aware Block-Sparse Attention Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we trace the inaccuracy of standard coarse-grained attention via mean pooling to a theoretical root cause: the interaction between mean pooling and Rotary Positional Embeddings (RoPE). |
Xinghao Wang; Pengyu Wang; Xiaoran Liu; Fangxu Liu; Jason Chu; Kai Song; Xipeng Qiu; | code |
| 186 | Sparser Block-Sparse Attention Via Token Permutation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose Permuted Block-Sparse Attention (**PBS-Attn**), a plug-and-play method that leverages the permutation properties of attention to increase block-level sparsity and enhance the computational efficiency of LLM prefilling. |
Xinghao Wang; Pengyu Wang; Dong Zhang; Chenkun Tan; Shaojun Zhou; Zhaoxiang Liu; Shiguo Lian; Fangxu Liu; Kai Song; Xipeng Qiu; | code |
| 187 | BandPO: Bridging Trust Regions and Ratio Clipping Via Probability-Aware Bounds for LLM Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While the canonical clipping mechanism in PPO serves as an efficient surrogate for trust regions, we identify a critical bottleneck: fixed bounds strictly constrain the upward update margin of low-probability actions, disproportionately suppressing high-advantage tail strategies and inducing rapid entropy collapse. To address this, we introduce **Band-constrained Policy Optimization** (BandPO). |
Yuan Li; Bo Wang; Yufei Gao; Yuqian Yao; Xinyuan Wang; Zhangyue Yin; Xipeng Qiu; | code |
| 188 | ProactiveLLM: Learning Active Interaction for Streaming Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose *ProactiveLLM*, which achieves active interaction by treating ”when to generate and ”what to generate as decoupled objectives. |
Junlong Tong; Yao Zhang; Anhao Zhao; Yingqi Fan; Yunpu Ma; Anhao Zhao; | code |
| 189 | TMD-Bench: A Multi-Level Evaluation Paradigm for Music–Dance Co-Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce TMD-Bench, a benchmark for text-driven music–dance co-generation that assesses systems across unimodal generation quality, instruction adherence, and cross-modal rhythmic alignment. |
Xiaoda Yang; Majun Zhang; Changhao Pan; Nick Huang; Yang Yuguang; fan zhuo; Pengfei Zhou; Jin Zhou; Sizhe Shan; Shan Yang; Miles Yang; Yang You; Zhou Zhao; | code |
| 190 | LARA: Latent Action Representation Alignment for Vision-Language-Action Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, LAM and VLA are typically trained separately, leaving LAM ungrounded during VLA training and VLA models constrained by frozen LAM representations. To address these issues, we propose Latent Action Representation Alignment (LARA), a plug-and-play framework that jointly optimizes LAM and VLA via representation alignment. |
Mengya Liu; Baoxiong Jia; Jiangyong Huang; Jingze Zhang; Siyuan Huang; | code |
| 191 | Sparse Autoencoders Are Topic Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This view implies SAE features are thematic components rather than steerable directions. Based on this, we introduce SAE-TM, a topic modeling framework that: (1) trains an SAE to learn reusable topic atoms, (2) interprets them as word distributions on downstream data, and (3) merges them into any number of topics without retraining. |
Leander Girrbach; Zeynep Akata; | code |
| 192 | EqGINO: Equivariant Geometry-Informed Fourier Neural Operators for 3D Partial Differential Equations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Conversely, Fourier Neural Operators (FNOs) efficiently capture global interactions, yet establishing 3D equivariance within them remains impractical due to the prohibitive cost of spectral group convolutions. To bridge this gap, we introduce EqGINO, a geometrically robust framework that enforces isotropy in the spectral domain. |
Sungwon Kim; Juho Song; Seungmin Shin; Guimok Cho; Sangkook Kim; Chanyoung Park; | code |
| 193 | Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Open-o3-Video, a non-agent framework that integrates explicit spatio-temporal evidence into video reasoning by highlighting key timestamps, objects, and bounding boxes, making the reasoning process traceable and verifiable. |
Jiahao Meng; Xiangtai Li; Haochen Wang; Tan Yue; Tao Zhang; Lingdong Kong; Yunhai Tong; Anran Wang; Zhiyang Teng; Yujing Wang; Zhuochen Wang; | code |
| 194 | TSRBench: A Comprehensive Multi-task Multi-modal Time Series Reasoning Benchmark for Generalist Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, current benchmarks for generalist models largely overlook this dimension. To bridge this gap, we introduce TSRBench, a comprehensive multi-modal benchmark designed to stress-test the full spectrum of time series reasoning capabilities. |
Fangxu Yu; Xingang Guo; Lingzhi Yuan; Haoqiang Kang; Hongyu Zhao; Lianhui Qin; Furong Huang; Bin Hu; Tianyi Zhou; | code |
| 195 | Learning Cardiac Latent Representations in Vectorcardiogram Space Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Therefore, representation learning in the ECG space inevitably introduces substantial redundancy, which may lead to spurious correlations and increased risk of overfitting. To address this and motivated by the Frank vectorcardiogram (VCG) model, we propose learning a unified latent representation of cardiac electrical activity directly in the VCG space. |
Bosong Huang; Panzhen Zhao; Zengxiang Li; Patricia Lee; Wei Jin; Alan Liew; Ming Jin; Shirui Pan; | code |
| 196 | Test-Time Reinforcement Learning for Flow Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce Flow-TTRL, the first test-time reinforcement learning framework that achieves alignment on the fly. |
Jili Chen; Changqin Huang; Qionghao Huang; Yaxin Tu; Zhonglong Zheng; Xiaodi Huang; | code |
| 197 | MCP-Persona: Benchmarking LLM Agents on Personalized MCP Tools and Tasks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing benchmarks predominantly focus on generic information-seeking tools and fail to capture the practical challenges posed by personal social applications, where tools interact with individual accounts or local databases. To bridge this critical gap, we introduce MCP-Persona, the first benchmark specifically designed for evaluating agent performance on real-world, personalized MCP tools. |
WenHao Wang; Peizhi Niu; Gongyi Zou; Xiyuan Yang; Jingxing Wang; Haoting Shi; Yaxin Du; Jingyi Chai; Xianghe Pang; shuo tang; Yanfeng Wang; Siheng Chen; | code |
| 198 | Referring Multiple Regions with Large Multimodal Models Via Contextual Latent Steering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Contextual Latent Steering (CSteer), a training-free approach for guiding general LMMs to refer multiple regions contextually, without expensive fine-tuning or architectural modifications. |
Yun Xing; Hanyuan Liu; Jiahao Nie; Shijian Lu; | code |
| 199 | Semantic Tube Prediction: Beating LLM Data Efficiency with JEPA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Surprisingly few works have successfully challenged the data-efficiency bounds implied by these laws—which is our primary focus. To that end, we introduce the Geodesic Hypothesis, positing that token sequences trace geodesics on a smooth semantic manifold and are therefore locally linear. |
Hai Huang; Yann LeCun; Randall Balestriero; | code |
| 200 | Incremental BPE Tokenization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a novel algorithm for incremental Byte Pair Encoding (BPE) tokenization. |
Shenghu Jiang; Ruihao Gong; | code |
| 201 | Energy-based Compositional Diffusion Planning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Energy-based Compositional Diffuser (ECD), a framework that formulates the global trajectory as the minimizer of the sum of local bridge potentials. |
Tao Sun; Utkarsh Mishra; Jiaxin Lu; Danfei Xu; Iro Armeni; | code |
| 202 | NITP: Next Implicit Token Prediction for LLM Pre-training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We argue that this sparse, one-hot supervision leaves the latent representation space under-constrained, allowing hidden states to drift into degenerate and anisotropic configurations that limit generalization. To address this issue, we propose **Next Implicit Token Prediction (NITP)**, which augments discrete prediction with dense, continuous supervision directly in the representation space. |
Xiangdong Zhang; Debing Zhang; Shaofeng Zhang; Xiaohan Qin; Yu Cheng; Junchi Yan; | code |
| 203 | BabyVision: Visual Reasoning Beyond Language Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We uncovered a crucial fact: state-of-the-art MLLMs consistently fail on basic visual tasks that humans, even 3-year-olds, can solve effortlessly. To systematically investigate this gap, we introduce BabyVision, a benchmark designed to assess core visual abilities independent of linguistic knowledge for MLLMs. |
Liang Chen; Weichu Xie; Liang Yiyan; Hongfeng He; Haozhe Zhao; Zhibo Yang; Zhiqi Huang; Haoning Wu; Haoyu Lu; Y.Charles; Yiping Bao; YuanTao Fan; Guopeng Li; Haiyang Shen; Xuanzhong Chen; Wendong XU; Shuzheng Si; Zefan Cai; Wenhao Chai; Ziqi Huang; Fangfu Liu; Tianyu Liu; Baobao Chang; Ming Wu; Xiaobo Hu; Kaiyuan Chen; Yixin Ren; Yang Liu; Yuan Gong; Kuan Li; | code |
| 204 | ECG-R1: Protocol-Guided and Modality-Agnostic MLLM for Reliable ECG Interpretation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Electrocardiography (ECG) serves as an indispensable diagnostic tool in clinical practice, yet existing multimodal large language models (MLLMs) remain unreliable for ECG interpretation, often producing plausible but clinically incorrect analyses. To address this, we propose ECG-R1, the first reasoning MLLM designed for reliable ECG interpretation via three innovations. |
Jiarui Jin; Haoyu Wang; Xingliang Wu; Xiaocheng Fang; Xiang Lan; Zihan Wang; Deyun Zhang; Bo Liu; Yingying Zhang; Xian Wu; Hongyan Li; Shenda Hong; | code |
| 205 | Vision-Language-Action Pretraining from Large-Scale Human Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing Vision-Language-Action (VLA) models struggle with complex manipulation tasks requiring high dexterity and generalization, primarily due to their reliance on synthetic data with significant sim-to-real gaps or limited teleoperated demonstrations. To address this bottleneck, we propose leveraging human hands as a manipulator template, capitalizing on the rich dexterity and scalability present in web data of human manipulation. |
Hao Luo; Yicheng Feng; Wanpeng Zhang; Sipeng Zheng; Ye Wang; Haoqi Yuan; jiazheng liu; Chaoyi Xu; Haiweng Xu; Qin Jin; Zongqing Lu; | code |
| 206 | Confidence and Difficulty-Adaptive Policy Optimization for LLM Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These findings imply that the benefit of each update depends strongly on both question difficulty and the model’s current competence. Motivated by this, we propose Confidence and Difficulty-adaptive Policy Optimization (CoDaPO), which assigns each question a bounded value from rollout confidence and empirical difficulty, then uses it to reweight policy updates and resample high-value questions within minibatches to increase discovery under a fixed compute budget. |
Zhanke Zhou; Xiangyu Lu; Chentao Cao; Brando Miranda; Tongliang Liu; Bo Han; Sanmi Koyejo; | code |
| 207 | Rapid Poison: Practical Poisoning Attacks Against The Rapid Response Framework Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The Rapid Response (RR) framework (Peng et al., 2024), deployed in production systems including Anthropic’s ASL-3 safeguards (Anthropic, 2025), dynamically adapts jailbreak detection classifiers by generating synthetic training data from emerging attacks. We reveal that prompt injection can infiltrate this pipeline to deliver poisoned samples into the classifier’s training set, enabling two attack objectives: (I) targeted poisoning attacks that create false positives on harmless samples by categorizing them as a jailbreak, with a specific desired feature (e.g., certain formatting, subject, or keyword), (II) concept-based backdoor attacks that induce false negatives on jailbreak inputs, generalizing even to jailbreaks from attack strategies the defender explicitly trained against, when the backdoor trigger is present. |
David Huang; Jaewon Chang; Avidan Shah; Prateek Mittal; Chawin Sitawarin; | code |
| 208 | Text-Driven Fusion for Infrared and Visible Images: Achieving Image Scene Adaptation on Hyperbolic Space Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, Euclidean geometry imposes rigid distance metrics that distort multi-modal feature interactions, particularly in preserving parent-to-child semantic hierarchies. To overcome this, we introduce a text-driven fusion framework empowered by hyperbolic manifold learning. |
Huan Kang; Hui Li; Tianyang Xu; Tao Zhou; Xiaojun Wu; Josef Kittler; | code |
| 209 | MN-Diff: Diffusion Parameterized MoE-NCDE for Continuous Time Series Generation with Irregular Observations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose MN-Diff, a continuous TSG framework that enhances NCDE with a Mixture-of-Experts (MoE) dynamics function and a decoupled architectural design for dynamics-focused training. |
Xu Zhang; Junwei Deng; Chang Xu; Hao Li; Jiang Bian; | code |
| 210 | Diffuse to Detect: Bi-Level Sample Rebalancing with Pseudo-Label Diffusion for Point-Supervised Infrared Small-Target Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Point supervision has become a scalable solution to address dense annotation for infrared small target detection, but its performance is limited by two coupled bottlenecks: unstable pseudo-label evolution in cluttered, low-contrast infrared imagery and severe sample-distribution imbalance. In this paper, we present a more adaptive and stable framework to address these issues. |
Zhu Liu; Yuanhang Yao; Ping Qian; Zihang Chen; Risheng Liu; | code |
| 211 | Optimizing Rank for High-Fidelity Implicit Neural Representations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we challenge the notion that the low-frequency bias of vanilla MLPs is an intrinsic, architectural limitation to learn high-frequency content, but instead a symptom of stable rank degradation during training. |
Julian McGinnis; Florian A Hölzl; Suprosanna Shit; Florentin Bieder; Paul Friedrich; Mark Mühlau; bjoern menze; Daniel Rueckert; Benedikt Wiestler; | code |
| 212 | Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This process involves an *architectural gap*, as it converts full attention into causal attention. In this paper, we demonstrate that existing methods fail to bridge this gap theoretically, leading to suboptimal performance. |
Hongzhou Zhu; Min Zhao; Guande He; Hang Su; Chongxuan Li; Jun Zhu; | code |
| 213 | UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We observe that naively adapting GRPO to UDM leads to unstable training and marginal performance. To address this, we propose \Ours, the first framework that integrates UDM with RL. |
Jiaqi Wang; Haoge Deng; Ting Pan; Yang Liu; Chengyuan Wang; Fan Zhang; Yonggang Qi; Xinlong Wang; | code |
| 214 | CODiff: One-Step Diffusion Model for Camouflaged Object Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In contrast to existing approaches that rely on multiple sample steps to refine the predicted masks, we propose CODiff, which reformulates the diffusion process to enable one-step mask prediction while maintaining competitive accuracy. |
Xiaotong Fu; Wenchao Meng; Qihang Zhou; Qian Liu; Qinmin Yang; Shibo He; | code |
| 215 | Physics-Informed Self-Supervised Learning on Efficient Electron-Density Images for Organic Material Property Prediction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Yet, establishing it as an input modality for material property prediction has been impeded by two practical barriers: scarce large-scale ED data and the enormous computational complexity of ED representation. To bridge these gaps, we introduce VisionED, an efficient physics-informed model pre-trained on electron-density images. |
Zhixiang Cheng; Hongxin Xiang; Mingquan Liu; Tengfei Ma; Yingzhuo Tu; Wenjie Du; Bosheng Song; Yiping Liu; xiangxiang Zeng; | code |
| 216 | TodoEvolve: Learning to Architect Agent Planning Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Planning has become a central capability for contemporary agent systems in navigating complex, long-horizon tasks, yet existing approaches predominantly rely on fixed, hand-crafted planning structures that lack the flexibility to adapt to the structural diversity of open-ended problems. To address this limitation, we introduce TodoEvolve, a meta-planning paradigm that autonomously synthesizes and dynamically revises task-specific planning architectures. |
Jiaxi Liu; Yanzuo Jiang; Guibin Zhang; Zihan Zhang; Heng Chang; Zhenfei Yin; Qibing Ren; Junchi Yan; | code |
| 217 | 3DMedAgent: Unified Perception-to-Understanding for 3D Medical Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In parallel, recent multimodal large language models (MLLMs) exhibit improved visual perception and can integrate visual and textual information effectively, yet their predominantly 2D-oriented designs fundamentally limit their ability to perceive and analysis volumetric medical data. To bridge this gap, we propose 3DMedAgent, an unified agent that enables 2D MLLMs to perform general 3D CT analysis without 3D-specific fine-tuning. |
Ziyue Wang; Linghan Cai; Chang Low; Haofeng Liu; Junde Wu; Jinyu Wang; Rui Wang; Lei Song; Jiang Bian; Jingjing Fu; Yueming Jin; | code |
| 218 | Towards High-Fidelity CAD Generation Via LLM-Driven Program Generation and Text-Based B-Rep Primitive Grounding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our method generates executable CadQuery scripts, and introduces a text-based query mechanism that enables the LLM to specify geometric selections via natural language, which *BRepGround* then grounds to the target primitives. |
Jiahao Li; Qingwang Zhang; Qiuyu Chen; Guozhan Qiu; Yunzhong Lou; Xiangdong Zhou; | code |
| 219 | Process Reward Agents for Steering Knowledge-Intensive Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Here, we introduce Process Reward Agents~(PRA), a test-time method for providing domain-grounded, online, step-wise rewards to a frozen reasoner. |
Jiwoong Sohn; Tomasz Sternal; Kenneth Styppa; Torsten Hoefler; Michael Moor; | code |
| 220 | Supervise Less, See More: Training-free Nuclear Instance Segmentation with Prototype-Guided Prompting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce SPROUT, a fully training- and annotation-free prompting framework for nuclear instance segmentation. |
Wen Zhang; Qin Ren; Wenjing Liu; Haibin Ling; Chenyu You; | code |
| 221 | Are VLMs Seeing or Just Saying? Uncovering The Illusion of Visual Re-examination Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce VS-Bench, a benchmark of $800$ image pairs curated from MathVista, MathVerse, MathVision, and MMMU-Pro. |
Chufan Shi; Cheng Yang; Yaokang Wu; Linghao Jin; Bo Shui; Taylor Berg-Kirkpatrick; Xuezhe Ma; | code |
| 222 | $G^2$-Reader: Dual Evolving Graphs for Multimodal Document QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce $G^2$-Reader, a dual-graph system, to address both issues. |
Yaxin Du; Junru Song; Yifan Zhou; Cheng Wang; Jiahao Gu; Zimeng Chen; Menglan Chen; Wen Yao; Yang Yang; Ying Wen; Siheng Chen; | code |
| 223 | EMFormer: Efficient Multi-Scale Transformer for Accumulative Context Weather Forecasting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While recent approaches employ finetuning to extend prediction horizons, they remain constrained by the issues of catastrophic forgetting, error accumulation, and high training overhead. To address these limitations, we present a novel pipeline across pretraining, finetuning and forecasting to enhance long‑context modeling while reducing computational overhead. |
hao chen; Tao Han; Jie ZHANG; Song Guo; Fenghua Ling; LEI BAI; | code |
| 224 | Plain Transformers Are Surprisingly Powerful Link Predictors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Through experimental and theoretical analysis, we show that PENCIL extracts richer structural signals than GNNs, implicitly generalizing a broad class of heuristics and subgraph-based expressivity. |
Quang Truong; Yu Song; Donald Loveland; Mingxuan Ju; Tong Zhao; Neil Shah; Jiliang Tang; | code |
| 225 | SleepLM: Natural-Language Intelligence for Human Sleep Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present SleepLM, a family of sleep-language foundation models that enable human sleep alignment, interpretation, and interaction with natural language. |
Zongzhe Xu; Zitao Shuai; Eideen Mozaffari; Ravi Aysola; Rajesh Kumar; Yuzhe Yang; | code |
| 226 | Transporting Task Vectors Across Different Architectures Without Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce Theseus, a training-free method for transporting task-specific updates across heterogeneous models. |
Filippo Rinaldi; Aniello Panariello; Giacomo Salici; Angelo Porrello; Simone Calderara; | code |
| 227 | EvoMAS: Heuristics in The Loop—Evolving Smarter Agentic Workflows Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current automation methods often generate templated agents, rely on monolithic optimization, and ignore task complexity gradients. This paper presents Evolutionary MAS (EvoMAS), a biologically inspired framework that addresses these limitations through three interconnected dimensions: (1) dynamic and diverse evolutionary strategies with six biologically inspired operators (3 exploration, 3 exploitation) and adaptive strategy selection; (2) role-level evolution that dynamically optimizes agent specialization and collaboration patterns; and (3) curriculum-guided evolution that partitions tasks by difficulty and evolves sequentially from simple to complex under cross-stage stability constraints. |
Yangbo Wei; Zhen Huang; Ronghao Xu; Hong Wang; WEI XING; | code |
| 228 | Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This leaves constructive problem-solving (finding unknown terms that satisfy specific conditions) underexplored and disconnected from process-level verifiability. To bridge this gap, we introduce **FPS** (_**F**ormal **P**roblem-**S**olving_), a principled framework to encompass the end-to-end problem-solving process in Lean 4. |
Qi Liu; Xinhao Zheng; Renqiu Xia; Xingzhi Qi; Qinxiang Cao; Junchi Yan; | code |
| 229 | Spherical Steering: Geometry-Aware Activation Rotation for Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This raises concerns about representation collapse and degradation of open-ended generation capabilities. In this work, we explore Spherical Steering, a training-free primitive that resolves this trade-off through activation rotation. |
Zejia You; Deng; Hanjie Chen; | code |
| 230 | Fourier Features Let Agents Learn High Precision Policies with Imitation Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We thus propose to use a parametric projection to map point clouds from Cartesian space into high-dimensional Fourier space when using a point cloud encoder. |
Balázs Gyenes; Emiliyan Gospodinov; Jan Frieling; Enrico Krohmer; Nicolas Schreiber; Xiaogang Jia; Niklas Freymuth; Gerhard Neumann; | code |
| 231 | Variational Bayesian Flow Network for Graph Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Variational Bayesian Flow Network (VBFN), which performs a variational lifting to a tractable joint Gaussian variational belief family governed by structured precisions. |
Yida Xiong; Jiameng Chen; Xiuwen Gong; Jia Wu; Shirui Pan; wenbin Hu; | code |
| 232 | TRACER: Trajectory Risk Aggregation for Critical Episodes in Agentic Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce TRACER, a trajectory-level uncertainty metric for dual-control Tool-Agent-User interaction. |
Sina Tayebati; Divake Kumar; Nastaran Darabi; Davide Ettori; Ranganath Krishnan; Amit Trivedi; | code |
| 233 | Absorbing Quantization Error By Deformable Noise Scheduler for Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a distribution-preserving framework that absorbs quantization error into the generative process without changing architecture or adding steps. |
Mingrui Yang; Wei Huang; Hao SHENG; Donglin Yang; Jichang Yang; Xin Yu; Huining Yu; Yuzhong Jiao; Zhongrui Wang; XIAOJUAN QI; | code |
| 234 | Zooming Without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent Thinking-with-Images methods alleviate this by iteratively zooming into regions of interest during inference, but incur high latency due to repeated tool calls and visual re-encoding. To address this, we propose Region-to-Image Distillation, which transforms zooming from an inference-time tool into a training-time primitive, thereby internalizing the benefits of agentic zooming into a single forward pass. |
Lai Wei; Liangbo He; jun lan; Lingzhong Dong; Yutong Cai; Siyuan Li; Huijia Zhu; Weiqiang Wang; Linghe Kong; Yue Wang; Zhuosheng Zhang; Weiran Huang; | code |
| 235 | Towards One-for-All Anomaly Detection for Tabular Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing methods follow a “one model for one dataset (OFO)” paradigm, which relies on dataset-specific training and thus incurs high computational cost and yields limited generalization to unseen domains. To address these limitations, we propose OFA-TAD, a generalist one-for-all (OFA) TAD framework that only requires one-time training on multiple source datasets and can generalize to unseen datasets from diverse domains on-the-fly. |
Shiyuan Li; Yixin Liu; Yu Zheng; Xiaofeng Cao; Shirui Pan; Heng Tao Shen; | code |
| 236 | Multi-Objective Protein Design Via Memory-Aware Test-Time Scaling in Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, current test-time diffusion methods face critical challenges: i) ineffective learning from interaction history leading to repetitive design errors, ii) over-reliance on successful cases as the reward signal, and iii) difficulties in balancing multi-objective functional trade-offs . To address these limitations, we propose MoMST, a framework for Multi-objective protein design via Memory-aware Self-contrastive learning with Test-time scaling in diffusion models. |
Ming Yang; Xin Zheng; Yi Li; YIZHEN ZHENG; Huan Yee Koh; Yanqing Guo; Xiaofeng Cao; Shirui Pan; | code |
| 237 | Inference-Time Forward-Process Alignment in Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose inference-time \textbf{F}orward-process \textbf{A}lignment for \textbf{Di}ffusion models (\textbf{DiFA}), a training-free inference framework that reformulates diffusion sampling as a sequential state estimation problem. |
Shigui Li; Delu Zeng; | code |
| 238 | Latent Forcing: Reordering The Diffusion Trajectory for Pixel-Space Image Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Latent Forcing, a simple modification to existing architectures that achieves the efficiency of latent diffusion while operating on raw natural images. |
Alan Baade; Eric Chan; Kyle Sargent; Changan Chen; Justin Johnson; Ehsan Adeli; Li Fei-Fei; | code |
| 239 | Agent-Omit: Training Efficient LLM Agents for Adaptive Thought and Observation Omission Via Agentic Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Based on our findings, we propose Agent-Omit, a unified training framework that empowers LLM agents to adaptively omit redundant thoughts and observations. |
Yansong Ning; Jun Fang; Naiqiang Tan; Hao Liu; | code |
| 240 | Benchmarking Physics-Informed Time-Series Models for Operational Global Station Weather Forecasting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose PhysicsFormer, a physics-informed forecasting model combining a dynamic core with a Transformer residual to predict future weather states. |
Tao Han; Zhibin Wen; Zhenghao Chen; Dazhao Du; Song Guo; LEI BAI; | code |
| 241 | Capability-Oriented Training Induced Alignment Risk Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While most AI alignment research focuses on preventing models from generating explicitly harmful content, a more subtle risk is emerging: capability-oriented training induced exploitation. We investigate whether language models, when trained with reinforcement learning (RL) in environments with implicit loopholes, will spontaneously learn to exploit these flaws to maximize their reward, even without any malicious intent in their training. |
Yujun Zhou; Yue Huang; Han Bao; kehan guo; Zhenwen Liang; Pin-Yu Chen; Tian Gao; Werner Geyer; Nuno Moniz; Nitesh Chawla; Xiangliang Zhang; | code |
| 242 | Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we show that the safety of MoE LLMs is as sparse as their architecture by discovering $\text{\emph{unsafe routes}}$: routing configurations that, once activated, convert safe outputs into harmful ones. |
Yukun Jiang; Hai Huang; Mingjie Li; Yage Zhang; Michael Backes; Yang Zhang; | code |
| 243 | Balancing Learning Rates Across Layers: Exact Two-Step Dynamics and Optimal Scaling in Linear Neural Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study optimal learning-rate selection in two-layer and three-layer linear neural networks trained to learn a single-index target function. |
Tianyu Pang; Vignesh Kothapalli; Shenyang Deng; Haohui Wang; Dawei Zhou; Yaoqing Yang; | code |
| 244 | UI2Code^N: UI-to-Code Generation As Interactive Visual Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address the non-differentiability of visual objectives and the noise of absolute visual evaluators, we propose Relative Visual Policy Optimization (RVPO), a preference-based reinforcement learning method that optimizes relative visual rankings among rendered candidates under execution feedback. |
Zhen Yang; Wenyi Hong; Mingde Xu; Xinyue Fan; Weihan Wang; Jiale Cheng; Xiaotao Gu; Jie Tang; | code |
| 245 | MAST: Motif-Augmented Diffusion with Search Tree for Spectroscopic Molecular Structure Elucidation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Moreover, the repeated full sampling inference strategy incurs substantial computation overhead. To address these limitations, we propose MAST, a Motif-Augmented diffusion framework with Search Tree, for joint 2D-3D spectroscopic molecular structure elucidation. |
Chenghao Jia; Mengdi Liu; Hong Chang; Shiguang Shan; Xilin Chen; | code |
| 246 | Scientific Logicality Enriched Methodology for LLM Reasoning: A Practice in Physics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we make the first systematic investigation into the internal logicality underlying LLM scientific reasoning, and develop a scientific logicality enriched methodology, including a set of assessment criteria and data sampling methods for logicality-guided training, to improve the logical faithfulness as well as task performance. |
Zhaoxin Yu; Nan Xu; Kun Chen; Jiahao Zhao; Lei Wang; Wenji Mao; | code |
| 247 | EDCO: Dynamic Curriculum Orchestration for Domain-specific Large Language Model Fine-tuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, most existing methods for LLM fine-tuning rely on a static curriculum, designed prior to training, which lacks adaptability to the model’s evolving needs during fine-tuning. To address this, we propose EDCO, a novel framework based on two key concepts: \textit{inference entropy} and \textit{dynamic curriculum orchestration}. |
Jing-Cheng Pang; Sun Liu; Zhouchang; txan; Haichuan Ma; KUN JIANG; Jianlong Wang; Kai Zhang; Sijie Wu; Haoran Cai; Chenwei Wu; Xubin Li; Xin Chen; | code |
| 248 | In-Context Generation with Regional Constraints for Instructional Video Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address these, we present ReCo, a new instructional video editing paradigm that novelly delves into **Re**gional **Co**nstraint modeling between editing and non-editing areas. |
Zhongwei Zhang; Fuchen Long; Wei Li; Zhaofan Qiu; Wu Liu; Ting Yao; Tao Mei; | code |
| 249 | TimeSeed: Effective Time Series Forecasting with Sparse Endogenous Variables Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce sparse endogenous forecasting as a new setting, where exogenous sequences and only sparse endogenous observations are available. |
Zhaowang Wu; Kaixin Deng; Hua Yan; | code |
| 250 | IAPO: Information-Aware Policy Optimization for Token-Efficient Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To bridge the gap, we propose IAPO, an information-theoretic post-training framework that assigns token-wise advantages based on each token’s conditional mutual information (MI) with the final answer. |
Yinhan He; Yaochen Zhu; Mingjia Shi; Wendy Zheng; Lin Su; Xiaoqing Wang; Qi Guo; Jundong Li; | code |
| 251 | ST-TGExplainer: Disentangling Stability and Transition Patterns for Temporal GNN Interpretability Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Both types of patterns are essential for faithful temporal explanations. To address this limitation, we propose ST-TGExplainer, a self-explainable TGNN that disentangles Stability and Transition patterns in temporal graphs for a more faithful Temporal GNN Explainer. |
Hongjiang Chen; Xin Zheng; Pengfei Jiao; Huan Liu; Zhidong Zhao; Huaming Wu; Feng Xia; Shirui Pan; | code |
| 252 | Generalized Correctness Models: Learning Calibrated and Cross-Model Correctness Predictors from Historical Patterns Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Moreover, we hypothesize that a key factor in predicting model correctness, i.e., building a “Correctness Model” (CM), is exposure to a target model’s historical predictions. We propose multiple methods to inject this historical correctness information, including training an LLM to predict the confidences of many other LLMs, i.e., creating a Generalized Correctness Model (GCM). |
Hanqi Xiao; Vaidehi Patil; Hyunji Lee; Elias Stengel-Eskin; Mohit Bansal; | code |
| 253 | XR-1: Towards Versatile Vision-Language-Action Models Via Learning Unified Vision-Motion Representations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present \textbf{XR-1}, a novel framework for versatile and scalable VLA learning across diverse robots, tasks, and environments. |
Shichao Fan; Kun Wu; Zhengping Che; Xinhua Wang; Di Wu; Fei Liao; Ning Liu; Yixue Zhang; Zhen Zhao; Zhiyuan Xu; Meng Li; Qingjie Liu; Shanghang Zhang; Min Wan; Jian Tang; | code |
| 254 | WUSH: Near-Optimal Adaptive Transforms for LLM Quantization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We derive closed-form optimal linear blockwise transforms for joint weight-activation quantization under standard RTN AbsMax-scaled block quantizers, covering both integer and floating-point formats. |
Jiale Chen; Vage Egiazarian; Roberto Castro; Torsten Hoefler; Dan Alistarh; | code |
| 255 | Interactive Person Retrieval Via Multi-Turn Multimodal Conversation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, these methods fall short of direct user interaction with retrieved candidates during conversation, making it challenging to effectively refine the retrieval results. To address these limitations, we propose multimodal interactive person retrieval (MInterPR), a new retrieval paradigm that allows users to iteratively refine retrieved candidates by providing feedback on visual differences from the target person. |
Yang Bai; Tingfeng Wang; Bin Yang; Min Cao; Jinqiao Wang; Mang Ye; | code |
| 256 | A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we establish a theoretical distributional view for visual MI, which models the influence of a feature activation on the natural image distribution, thereby formulating a Kullback-Leibler (KL)-minimal optimization problem to model the MI task. |
Guancheng Zhou; Yisi Luo; Zhengfu He; Zhenyu Jin; Xuyang Ge; Wentao Shu; Deyu Meng; Xipeng Qiu; | code |
| 257 | Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Critically, we find this low-rank structure as a key factor of the prevalent dead feature problem in sparse dictionary learning, where it creates a mismatch between randomly initialized features and the intrinsic geometry of the activation space. Building on this insight, we propose a subspace-constrained training method for sparse autoencoders (SAEs), initializing feature directions into the active subspace of activations. |
Junxuan Wang; Xuyang Ge; Wentao Shu; Zhengfu He; Xipeng Qiu; | code |
| 258 | Hyper-LLaVA: Hyperbolic Uncertainty-aware Modality-Balanced Routing for Multimodal Continual Instruction Tuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: (2) Different modalities exhibit varying reliability across tasks, where the modality with inter-task ambiguity can easily misguide the routing result. To address these problems, we propose Hyperbolic Uncertainty-aware Modality-Balanced Routing (Hyper-LLaVA) to improve parameter routing capacity based on cross-modality task feature uncertainty modeling. |
Kunlun Xu; YanQin Zhang; Wenwen Qiang; Jiahuan Zhou; | code |
| 259 | Stable Velocity: A Variance Perspective on Flow Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: For training, we introduce Stable Velocity Matching (StableVM), an unbiased variance-reduction objective, along with Variance-Aware Representation Alignment (VA-REPA), which adaptively strengthen auxiliary supervision in the *low-variance regime*. |
Donglin Yang; Yongxing Zhang; Xin Yu; Liang Hou; Xin Tao; Pengfei Wan; XIAOJUAN QI; Renjie Liao; | code |
| 260 | Segment-Aligned Policy Optimization for Multi-Modal Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, such formulations often misalign with the natural step-wise structure of reasoning processes, leading to suboptimal credit assignment and unstable training in multi-modal reasoning tasks. To bridge this gap, we propose Segment-Aligned Policy Optimization (SAPO), a novel reinforcement learning paradigm that treats coherent reasoning steps, rather than tokens or full sequences as fundamental units of policy update. |
Lei Gao; Zhuoming Li; Mengxi Jia; Jiakang Yuan; Hongbo Sun; Hao Sun; Xuelong Li; | code |
| 261 | Debiased Model-based Representations for Sample-efficient Continuous Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These incur biases in representation and actor-critic learning, leading to inferior performance. To address this, we propose Debiased model-based Representations for Q-learning, tagged DR. Q algorithm. |
Jiafei Lyu; Zichuan Lin; Scott Fujimoto; Kai Yang; Yangkun Chen; Saiyong Yang; Zongqing Lu; Deheng Ye; | code |
| 262 | IVGR: Internalizing Visually Grounded Reasoning for MLLMs with Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we empirically find that mandating the explicit object boxes in visually grounded CoT during inference often degrades performance compared to standard textual CoT—which reasons without explicit visual grounding. |
Chang-Bin Zhang; Yujie Zhong; Qiang Zhang; Kai Han; | code |
| 263 | Embodiment-Conditioned Mixture of Experts Increases The Evolvability of Robots Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce a model of evolution and learning in robots that co-optimizes a distribution of latent design vectors (genotypes) and a mixture of control experts (neural modules), which are gated by the latent coordinates of each decoded design (phenotype). |
Yibin Wang; Muhan Li; Zihan Guo; Sam Kriegman; | code |
| 264 | It’s TIME: Towards The Next Generation of Time Series Forecasting Benchmarks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, we contend that existing benchmarks exhibit common limitations in four dimensions: constrained data composition dominated by reused legacy sources, compromised data integrity lacking rigorous quality assurance, misaligned task formulations detached from real-world contexts, and rigid analysis perspectives that obscure generalizable insights. To bridge these gaps, we introduce **TIME**, a next-generation task-centric benchmark comprising 50 fresh datasets and 98 forecasting tasks, tailored for strict zero-shot TSFM evaluation free from data leakage. |
Zhongzheng Qiao; SHENG PAN; Anni Wang; Viktoriya Zhukova; Yong Liu; Xudong Jiang; Qingsong Wen; Mingsheng Long; Ming Jin; Chenghao Liu; | code |
| 265 | MedMamba: Multi-View State Space Models with Adaptive Graph Learning for Medical Time Series Classification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite recent progress, existing methods struggle to jointly model local-global dynamics and handle nonstationarities like baseline drift, while often failing to capture latent channel interactions. To address these challenges, we propose **MedMamba**, an end-to-end architecture that integrates state space models with domain-specific inductive biases. |
Da Zhang; bingyu li; Zhiyuan Zhao; Hongyuan Zhang; Junyu Gao; Xuelong Li; | code |
| 266 | Unison: Benchmarking Unified Multimodal Models Via Synergistic Understanding and Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, despite their unified designs, existing evaluations typically assess understanding and generation capabilities in isolation, overlooking the synergy between comprehension and generation. To bridge this gap, we introduce **Unison**, a comprehensive benchmark comprising 2,169 high-quality unified task samples, designed to evaluate joint understanding and generation in unified multimodal models. |
Jinyu Liu; Xincheng Shuai; Henghui Ding; Yu-Gang Jiang; | code |
| 267 | More Sail Than Ballast: Addressing Harmful Knowledge Leakage in The Expansive Reasoning Space of LRMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Experiments on our benchmark show that it is a common issue across current LRMs due to their strong multi-step reasoning capabilities. To address this issue, we propose placing LLMs in our synthesized open-ended environments, allowing them to self-search for a safety reasoning pattern to respond responsibly and helpfully. |
Qibing Ren; Xinhao Song; Ke Fan; Lijun Li; Zhanpeng Zhou; Gongshen Liu; Junchi Yan; Lizhuang Ma; Jing Shao; | code |
| 268 | Data Agent: Learning to Select Data Via End-to-End Dynamic Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing methods typically rely on task-specific handcrafted metrics or static/snapshot-based criteria to estimate sample importance, limiting scalability across learning paradigms and making it difficult to capture the evolving utility of data throughout training. To address this challenge, we propose Data Agent, an end-to-end dynamic data selection framework that formulates data selection as a training-aware sequential decision-making problem. |
Suorong Yang; Fangjian Su; Hai Gan; Ziqi Ye; Jie Li; Baile Xu; Furao Shen; Soujanya Poria; | code |
| 269 | One LR Doesn’t Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce \textbf{Layerwise Learning Rate (LLR)}, an adaptive scheme that assigns distinct learning rates to individual Transformer layers. |
Di He; Songjun Tu; Keyu Wang; Lu Yin; Shiwei Liu; | code |
| 270 | SmoothSpike: Spiking Transformer with Learnable Hadamard Transformation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this phenomenon, distinct high-amplitude inputs result in identical maximized spike counts, truncating the dynamic range and hindering the model’s ability to capture fine-grained semantic differences. To address this, we propose SmoothSpike, a novel method designed to enhance representational capacity by suppressing spike saturation. |
Zijian Zhou; Wenjie Wei; Yu Liang; Jialin Li; Ammar Belatreche; Honglin Cao; Shuai Wang; Malu Zhang; Yang Yang; Haizhou Li; | code |
| 271 | Positional Encoding for Spiking Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Spiking Positional Encoding (SPE), a novel positional encoding method specifically designed for Spiking Transformers that effectively captures relative positional information. |
Zijian Zhou; Yu Liang; Honglin Cao; Ammar Belatreche; Jieyuan Zhang; Wenjie Wei; Shuai Wang; Malu Zhang; Yang Yang; Haizhou Li; | code |
| 272 | AdaS: Adaptive Gradient Descent for Spiking Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work presents the first systematic investigation of optimization algorithms specifically tailored for SNNs, offering a practical tool to narrow the accuracy gap with ANNs while preserving the energy advantages of spike-based computation. |
Zijian Zhou; Honglin Cao; Ammar Belatreche; Wenjie Wei; Yimeng Shan; Yu Liang; Yu Yang; Shuai Wang; Yalan Ye; Malu Zhang; Yang Yang; Haizhou Li; | code |
| 273 | Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We argue that tool usage is not always beneficial, as redundant or inappropriate invocations largely increase reasoning overhead and even mislead model predictions. To address this issue, we introduce AutoTool, a model that adaptively decides whether to invoke tools according to the characteristics of each query. |
Qinghe Ma; Zhen Zhao; Yiming Wu; Jian Zhang; LEI BAI; Yinghuan Shi; | code |
| 274 | Cert-LAS: Toward Certified Model Ownership Verification for Text-to-Image Diffusion Models Via Layer-Adaptive Smoothing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, in practice, adversaries may intentionally or unintentionally damage potential watermark signals, significantly degrading verification reliability. To address this issue, we propose Cert-LAS, the first certified MOV method for T2I models based on layer-adaptive smoothing. |
Leyi Qi; Yiming Li; Siyuan Liang; Zhengzhong Tu; Dacheng Tao; | code |
| 275 | LiveFigure: Generating Editable Scientific Illustration with VLM Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing models generate raster outputs that do not support manual correction or layout adjustment, limiting their utility in scientific publishing, where editable vector figures are typically required for submission. To address this challenge, we introduce **LiveFigure**, an agentic framework driven by VLM agents that imitates the multi-step drawing workflow of human researchers. |
Chenyang Shao; Jiahe Liu; Fengli Xu; Yong Li; | code |
| 276 | SciNet: Evaluating AI Agents in Relation-Aware Scientific Literature Retrieval Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This fundamental limitation often results in fragmented knowledge structures, misinterpreted research sentiment, and ineffective modeling of collective scientific progress. To address this limitation, we introduce **SciNet**, the first **Sci**entific **Net**work relation-aware dataset for information retrieval agents. |
Chenyang Shao; Fengli Xu; Yong Li; | code |
| 277 | Mitigating The Safety–Utility Trade-off in LLM Alignment Via Adaptive Safe Context Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To mitigate the safety-utility trade-off, we propose the Adaptive Safe Context Learning (ASCL) framework to improve the reasoning given proper context. |
Yanbo Wang; Minzheng Wang; Jian Liang; Lu Wang; Yongcan Yu; Ran He; | code |
| 278 | Dual-Latent Memory Routing for Vision-Language Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A key factor is that they frequently lose track of earlier visual evidence and intermediate constraints under a monolithic growing context. Inspired by how humans separately recall what they see and what they infer when solving complex tasks, we propose DLMR, a parameter-efficient mechanism that equips MLLMs with dual latent memories: a visual memory that compresses image evidence and a reasoning memory that tracks intermediate conclusions and constraints. |
Hao-Xuan Ma; Jin-Fei Qi; YiCheng Xiao; Han-Jia Ye; | code |
| 279 | AgentSelect: Benchmark for Narrative Query-to-Agent Recommendation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing LLM leaderboards and tool/agent benchmarks evaluate components in isolation and remain fragmented across tasks, metrics, and candidate pools, leaving a critical research gap: there is little \emph{query-conditioned} supervision for learning to recommend end-to-end agent configurations that couple a backbone model with a toolkit. We address this gap with \methodname, a benchmark that reframes agent selection as narrative query-to-agent recommendation over capability profiles and systematically converts heterogeneous evaluation artifacts into unified, positive-only interaction data. |
Yunxiao Shi; Wujiang Xu; Tingwei Chen; Haoning Shang; Ling Yang; Yunfeng Wan; Zhuo Cao; Xing Zi; Dimitris Metaxas; Min Xu; | code |
| 280 | Find, Fix, Reason: Context Repair for Video Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In contrast, larger models excel at instruction following and multi-modal understanding, can supply richer context to smaller models, and rapidly zoom in on target regions via simple tools. Building on this capability, we introduce an observation-level intervention: a frozen, tool-integrated teacher identifies the missing spatiotemporal dependency and provides a minimal evidence patch (e.g. timestamps, regions etc.) from the original video while the question remains unchanged. |
Haojian Huang; Chuanyu Qin; Yinchuan Li; YINGCONG CHEN; | code |
| 281 | ZeroDiff: Zero-Shot Time Series Reconstruction Via Informed-Prior Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose ZeroDiff, which constructs an informed prior from exogenous variables alone, then learns to calibrate reconstruction errors through diffusion—training on observed locations and generalizing to unobserved ones. |
Yingda Fan; Dan Lu; Xiaowei Jia; | code |
| 282 | Don’t Reinvent The Wheel, Just Realign The Spokes: Resource-Efficient Federated Fine-Tuning Via Rank-Wise Expert Assembly Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose SmartFed, a resource-efficient framework that circumvents expensive training from scratch by intelligently reusing knowledge embedded in existing LoRA modules. |
Yebo Wu; Jingguang Li; Zhijiang Guo; Li Li; | code |
| 283 | Mean Flow Policy Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, their iterative generative processes introduce substantial training and inference overhead. To overcome this limitation, we propose to represent policies using MeanFlow models, a class of few-step flow-based generative models, to improve training and inference efficiency over diffusion-based RL approaches. |
Xiaoyi Dong; Xi Zhang; Jian Cheng; | code |
| 284 | S3Audio: Towards Streaming Synchronized Spatial Audio Generation Via Autoregressive Diffusion Transformer Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing spatial audio synthesis technologies are often encumbered by a tradeoff between generation quality and high inference latency, as well as difficulty in capturing precise spatial information from multimodal inputs. To address these challenges, we propose S3Audio, a unified streaming framework for high-fidelity spatial audio generation from panoramic videos and text prompts. |
Ke Lei; Yu Zhang; Changhao Pan; Xueyi Pu; Wenxiang Guo; Ruiqi Li; Zhou Zhao; | code |
| 285 | FIRE: Multi-fidelity Regression with Distribution-conditioned In-context Learning Using Tabular Foundation Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce FIRE, a training-free MF framework that couples tabular foundation models (TFMs) to perform zero-shot in-context Bayesian inference via a high-fidelity correction model conditioned on the low-fidelity model’s posterior predictive distributions. |
Rosen Yu; Nicholas Sung; Faez Ahmed; | code |
| 286 | Multimodal Nested Learning for Decoupled and Coordinated Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although existing methods mitigate this issue through optimization modulation and conflict alleviation, they still suffer from entangled optimization and uniform learning pace in conventional monolithic frameworks, limiting the effectiveness of multimodal learning. To address this issue, we propose a novel Multimodal Nested Learning Framework (MoNet), which reformulates the monolithic framework into nested sub-processes, decoupling and coordinating multimodal learning. |
Yanglin Feng; Yang Qin; Dezhong Peng; Rui Wang; Xiaomin Song; Peng Hu; | code |
| 287 | Turning Bias Into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce BITE (BIas exploraTion and Exploitation), a black-box adversarial framework that learns semantics-preserving edits to mislead the judgment and artificially inflate judged scores. |
XIANGLIN YANG; Bryan Hooi; Gelei Deng; Tianwei Zhang; Jin Song Dong; | code |
| 288 | What Information Matters? Graph Out-of-Distribution Detection Via Tri-Component Information Decomposition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Prior work established that methods trained with standard supervised learning (SL) objectives tend to capture spurious signals from either features and/or structure, leaving the model fragile under distributional changes. To address this, we propose \textsc{Tide}, a novel and effective \underline{T}ri-Component \underline{I}nformation \underline{De}composition framework that explicitly decomposes information into \textit{feature-specific, structure-specific and joint} components. |
Danny Wang; Ruihong Qiu; Zi Huang; | code |
| 289 | Monitoring LLM-based Multi-Agent Systems Against Corruptions Via Node Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Nevertheless, these methods predominantly focus on static graph defense, attempting to either detect attacks in a fixed graph structure or optimize a static topology with certain defensive capabilities. To address this limitation, we propose a dynamic defense paradigm for MAS graph structures, which continuously monitors communication within the MAS graph, then dynamically adjusts the graph topology, accurately disrupts malicious communications, and effectively defends against evolving and diverse dynamic attacks. |
Chengcan Wu; Zhixin Zhang; Mingqian Xu; Zeming Wei; Meng Sun; | code |
| 290 | Efficient Distributed MLLM Training with ModalGlue Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present ModalGlue, an efficient distributed MLLM training framework that contemplates MLLM’s unique characteristics in both model and data parallelization. |
Insu Jang; Runyu Lu; Nikhil Bansal; Ang Chen; Mosharaf Chowdhury; | code |
| 291 | AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The inherent characteristics of AMS, including floating-base dynamics, strong coupling between the UAV and the manipulator, and the multi-step, long-horizon nature of operational tasks, pose severe challenges to existing VLA paradigms designed for static or 2D mobile bases. To bridge this gap, we propose AIR-VLA, the first VLA benchmark specifically tailored for aerial manipulation. |
Jianli Sun; Bin Tian; Qiyao Zhang; Chengxiang Li; Zihan Song; Zhiyong Cui; Yisheng Lv; Yonglin Tian; | code |
| 292 | AutoVSR: Automatic Visual-to-Symbolic Reasoning for Symbolic Expression Generation from Circuit Schematic Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work proposes AutoVSR, an automated framework for visual-to-symbolic generation of circuit expressions using Vision Language Models (VLMs). |
Zhe Xiao; longfei li; Xu He; Haoying Wu; Zixing Zhang; Mingyu Liu; | code |
| 293 | Let The Prototype Guide You: Robust Aggregation of Sparse Multi-Class Annotations Via Annotator Prototype Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose **CPBCC** (**C**lass-specific **P**rototype-driven **B**ayesian **C**lassifier **C**ombination), which creatively models annotators through a dual-pathway architecture: (i) learning class-specific prototype annotation patterns across all annotators, and (ii) learning annotator-specific weights over prototypes. |
Ju Chen; Jun Feng; Shenyu Zhang; | code |
| 294 | Reasoning Over Boundaries: Enhancing Specification Alignment Via Test-time Deliberation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We formalize this challenge as specification alignment, focusing on LLMs’ ability to follow dynamic, scenario-specific spec from both behavioral and safety perspectives. To address this challenge, we introduce SpecBench, a unified benchmark for measuring specification alignment, covering 5 scenarios, 103 spec, and 1,500 prompts. |
Haoran Zhang; Yafu Li; Xuyang Hu; Dongrui Liu; Zhilin Wang; Bo Li; Yu Cheng; | code |
| 295 | Characterizing, Evaluating, and Optimizing Complex Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: (1) We introduce the ME$^2$ principle to characterize reasoning quality along macro- and micro-level concerning efficiency and effectiveness. |
Haoran Zhang; Yafu Li; Zhi Wang; Zhilin Wang; Shunkai Zhang; Xiaoye Qu; Yu Cheng; | code |
| 296 | Hyperspectral Image Fusion with Spectral-Band and Fusion-Scale Agnosticism Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current deep learning models for Multispectral and Hyperspectral Image Fusion (MS/HS fusion) are typically designed for fixed spectral bands and spatial scales, which limits their transferability across diverse sensors. To address this, we propose SSA, a universal framework for MS/HS fusion with spectral-band and fusion-scale agnosticism. |
Yujie Liang; ZiHan Cao; Liang-Jian Deng; Yang Yang; Malu Zhang; | code |
| 297 | R2R2: Robust Representation for Intensive Experience Reuse Via Redundancy Reduction in Self-Predictive Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While prior works focus on critic bias, representation-level instability in Self-Predictive Learning (SPL) under high Update-to-Data (UTD) regimes remains underexplored. To bridge this gap, we propose Robust Representation via Redundancy Reduction (R2R2), a regularization method within SPL. |
Sanghyeob Song; donghyeok lee; Jinsik KIM; Sungroh Yoon; | code |
| 298 | Time-Conditioned Foreseeing: An EHR-Specific Foundation Model for Irregular Dynamics and Calendrical Time Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a pretraining method tailored to EHRs’ distinct features. |
Bong Gyun Kang; JUNYONG AHN; Hyeongrok Han; Sungroh Yoon; | code |
| 299 | GePBench: Evaluating Fundamental Geometric Perception for Multimodal Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While multimodal large language models (MLLMs) have made significant advancements in visual understanding, their abilities to recognize geometric shapes and their spatial relationships, which we term geometric perception, are not explicitly and systematically explored. To address this gap, we introduce GePBench, a novel benchmark specifically designed to assess the geometric perception capabilities of MLLMs. |
Shangyu Xing; Changhao Xiang; Xinyu Liu; Zhangtai Wu; Zhen Wu; Yue YIfan; Yuteng Han; Fei Zhao; Xinyu Dai; | code |
| 300 | ERGeoBench: A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce ERGeoBench, a large-scale benchmark for vision-driven embodied geo-localization. |
Kaiwen Xue; Tao Wei; Zhonghong Ou; Guoxin Zhang; Kaoyan Lu; Yu Feng; Yifan Zhu; Haoran Luo; | code |
| 301 | A Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These flaws induce models to learn spurious correlations between background noise and target categories instead of robust acoustic features. To address this, we propose an automated pipeline that eliminates co-occurrence of events by mining high-purity single-event segments from in-the-wild datasets via a semantically consistent synthesis protocol. |
Kai Li; Jintao Cheng; Chang Zeng; Zijun Yan; Helin Wang; Zixiong Su; Bo Zheng; Xiaolin Hu; | code |
| 302 | Evolving Quantitative Reasoning Through Self-Play in Digital Twin Markets Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: By coupling high-level planning with robust quantitative execution, the proposed framework improves the quantitative reliability of LLM-driven decision-making. |
Tianmi Ma; Wenxin Huang; jiawei du; Lin Li; Xian Zhong; Joey Tianyi Zhou; | code |
| 303 | Ophiuchus: Incentivizing Tool-augmented ”Think with Images” for Joint Medical Segmentation, Understanding and Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, they still struggle with complex clinical tasks that necessitate dynamic and iterative focusing on fine-grained visual regions. To close this gap, we introduce Ophiuchus, a versatile, tool-augmented framework that equips an MLLM to (i) decide when fine-grained visual evidence is needed, (ii) determine where to probe and ground within the medical image, and (iii) seamlessly weave the relevant sub-image content back into an interleaved, multimodal chain of thought for precise segmentation and diagnosis. |
Yankai Jiang; Yujie Zhang; Peng Zhang; Wenjie Li; Yichen Li; Jintai Chen; Xiaoming Shi; Shihui Zhen; | code |
| 304 | Surgery: Mitigating Harmful Fine-Tuning for Large Language Models Via Attention Sink Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we utilize the attention sink mechanism to mitigate harmful fine-tuning. |
Guozhi Liu; Weiwei Lin; Tiansheng Huang; Ruichao Mo; Qi Mu; Xiumin Wang; Li Shen; | code |
| 305 | On The Power of Statistics in Class-Incremental Learning with Pretrained Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we show that class-level feature statistics play a central role in enabling effective CIL under strong pre-training. |
Zhiwen Cao; Yanfeng Li; Shudong Huang; Yalan Ye; Shuyin Xia; Yi Wang; Jiancheng Lv; | code |
| 306 | FT-Dojo: Towards Autonomous LLM Fine-Tuning with Language Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We frame this as a substantially open problem: agents must navigate an open-ended search space spanning data curation from diverse data sources, processing with complex tools, building a training pipeline, and iteratively refining their approach based on evaluation outcomes in rapidly growing logs—an overall scenario far more intricate than existing benchmarks. To study this question, we introduce FT-Dojo, an interactive environment comprising 13 tasks across 5 domains. |
Qizheng Li; Yifei Zhang; Xiao Yang; Xu Yang; Zhuo Wang; Bowen Xian; Weiqing Liu; Jiang Bian; | code |
| 307 | A²RBench: An Automatic Paradigm for Formally Verifiable Abstract Reasoning Benchmark Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, accurately measuring this ability remains challenging: existing benchmarks either rely on expensive manual annotation, limiting their scale, or risk measuring memorization rather than genuine reasoning. To address this, we introduce an automated pipeline named A$^2$RBench, encompassing generation, expansion, evaluation, and analysis. |
Qingchuan Ma; Yuexiao Ma; Yongkang Xie; Tianyu Xie; Xiawu Zheng; Rongrong Ji; | code |
| 308 | Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This unreliability can even undermine a key promise of implicit PRMs—scoring many candidate tokens—because noisy per-token advantages may systematically reinforce incorrect continuations. We address this with a novel Implicit Prefix-Value Reward Model (IPVRM), which directly learns a prefix-conditioned value function estimating the probability of eventual correctness, and derives step signals via temporal-difference (TD) differences. |
Shiping Gao; Hongzhan Chen; Xiaojun Quan; Qifan Wang; Lifu Huang; | code |
| 309 | Mitigating Noise-Induced Layout Priors for Object Counting in Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We formalize this phenomenon as the \textbf{\textit{Noise-Induced Layout Prior}}. Leveraging this insight, we propose a novel training-free framework for object counting in diffusion models. |
Xiaoling Gu; Xuelong Li; Shengqi Wu; Yongkang Wong; wu; Huan Li; Zhou Yu; Mohan Kankanhalli; | code |
| 310 | Divide and Conquer: Reliable Multi-View Evidential Learning for Deepfake Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This imbalance leads to overconfident yet brittle predictions—a phenomenon we term the Semantic Masking Effect. To address this challenge, we propose \Reliable Multi-View Evidential Learning for Deepfake Detection} under a Divide-and-Conquer strategy. |
Xiaolu Kang; Zhongyuan Wang; Baojin Huang; Jikang Cheng; Zhanhe Lei; Gang Wu; Qin Zou; Qian Wang; | code |
| 311 | API: Adaptive Prototype Imputation for Incomplete Multimodal Sentiment Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This dichotomy results in two critical limitations: I) computational inefficiency and semantic inconsistency (recovery-based methods rely on heavy generators that incur prohibitive inference latency and risk semantic drift due to lack of class-level priors); II) lack of instance specificity (non-recovery-based methods rely on static global mappings that fail to capture sample-specific affective cues). To address these gaps, we propose Adaptive Prototype Imputation (API). |
Xiaotao Wang; Yiyang Fang; Wenke Huang; Bin Yang; Guancheng Wan; Mang Ye; | code |
| 312 | Temporal Straightening for Latent Planning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by the perceptual straightening hypothesis in human visual processing, we introduce temporal straightening for representation learning in latent planning. |
Ying Wang; Oumayma Bounou; Gaoyue Zhou; Randall Balestriero; Tim G. J. Rudner; Yann LeCun; Mengye Ren; | code |
| 313 | DirectEdit: Step-Level Accurate Inversion for Flow-Based Image Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, since the reconstruction path is approximated using noisy latents from mismatched timesteps, existing methods inevitably suffer from accumulated drift, which fundamentally limits reconstruction fidelity. To address this challenge, we systematically analyze the inversion process within the flow transformer and propose DirectEdit, a simple yet efficient editing method that eliminates the inherent reconstruction error without introducing additional neural function evaluations (NFEs). |
Desong Yang; Mang Ye; | code |
| 314 | Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we ask whether selectively skipping latent iterations may improve accuracy. |
Tianyu Fu; Yichen You; Zekai Chen; Guohao Dai; Huazhong Yang; Yu Wang; | code |
| 315 | DriveWorld-VLA: Unified Latent-Space World Modeling with Vision–Language–Action for Autonomous Driving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing methods fail to effectively unify future scene evolution and action planning within a single architecture due to inadequate sharing of latent states, limiting the impact of visual imagination on action decisions. To address this limitation, we propose DriveWorld-VLA, a novel framework that unifies world modeling and planning within a latent space by tightly integrating VLA and world models at the representation level, which enables the VLA planner to benefit directly from holistic scene-evolution modeling and reducing reliance on dense annotated supervision. |
Feiyang Jia; Lin Liu; Ziying Song; Caiyan Jia; Hangjun Ye; Xiaoshuai Hao; Long Chen; | code |
| 316 | Full-Spectrum Graph Neural Network: Expressive and Scalable Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: It is well established that spectral graph neural networks (GNNs) can universally approximate node signals; however, their expressive power remains bounded by the 1-dimensional Weisfeiler–Lehman test, which is mirrored in their lack of universality for higher-order signals. To go beyond this bound, we propose the Full-Spectrum GNN (FSpecGNN), a second-order generalization of classical spectral GNNs. |
Xiaohan Wang; Deyu Bo; Longlong Li; Kelin Xia; | code |
| 317 | D$^2$O: A Dual Debiasing Operator for Training-Free Test-Time Adaptation of Vision–Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose D$^2$O, a strictly training-free debiasing operator that outputs three inference objects per test sample: a content feature for reliable retrieval, a style fingerprint for environment routing, and debiased logits for corrected priors. |
Yihong Luo; Wenwu He; Dong Liang; Yihang Zhou; Zhuo-Xu Cui; | code |
| 318 | FOVI: A Biologically-inspired Foveated Interface for Deep Vision Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a foveated vision interface (FOVI) based on the human retina and primary visual cortex, that reformats a variable-resolution retina-like sensor array into a uniformly dense, V1-like sensor manifold. |
Nicholas Blauch; George Alvarez; Talia Konkle; | code |
| 319 | AReaL-DTA: Dynamic Tree Attention for Efficient Reinforcement Learning of Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce AReaL-DTA to efficiently exploit prefix sharing in RL training. |
Jiarui Zhang; Yuchen Yang; Ran Yan; Zhiyu Mei; Liyuan Zhang; LiDaifeng; Wei Fu; Jiaxuan Gao; Shusheng Xu; Yi Wu; Binhang Yuan; | code |
| 320 | Selective Coupling of Decoupled Informative Regions: Masked Attention Alignment for Data-Free Quantization of Vision Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a novel Masked Attention Alignment approach for Data-Free Quantization of ViTs, named MaskAQ, revealing that: 1) the semantics in the self-attention mechanism is predominantly localized to a sparse subset of patches, called informative regions; 2) the informative regions dominate the mutual information between synthetic samples and $Q$’s outputs. |
Biao Qian; Yang Wang; Yong Wu; Jungong Han; | code |
| 321 | CofactGVR: Counterfactual Intervention for Grounded Visual Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite rapid progress in Grounded Visual Reasoning (GVR) with MLLMs and RL-style fine-tuning, existing approaches often lack effective learning signals for intermediate grounding decisions and are prone to shortcut solutions. In this work, we explicitly decompose GVR into Evidence Generation followed by Counterfactual Answer Reasoning, and formalize this structure as a Causal Grounding Graph (CGG) in which the generated evidence acts as a causal mediator. |
Yan Zhang; Zhijin Qin; Feng Xu; guiguang ding; Jungong Han; | code |
| 322 | DecodeShare: Tracing The Shared Pathways of LLM Decode-Time Decisions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Large language models (LLMs) handle many tasks with one set of parameters, but under KV-cached inference it is unclear what task-general structure, if any, is used at $\textit{decode time}$ rather than during $\textit{prefill}$. We propose $\textbf{DecodeShare}$, a protocol that identifies a low-dimensional subspace that is consistently shared across tasks in decode-time hidden states, and then tests its causal role by removing that subspace only during decoding. |
Zishan Shao; Lixun Zhang; Kangning Cui; Yixiao Wang; Ting Jiang; Hancheng Ye; Qinsi Wang; Zhixu Du; Yuzhe Fu; Fan Yang; Danyang Zhuo; Yiran Chen; Hai Li; | code |
| 323 | Cross-Chirality Generalization By Axial Vectors for Hetero-Chiral Protein-Peptide Interaction Design Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we show that by injecting axial features to E(3)-equivariant (polar) vector features, it is feasible to achieve cross-chirality generalization from homo-chiral (L-L) training data to hetero-chiral (D-L) design tasks. |
Ziyi Yang; Zitong Tian; Yinjun Jia; Tianyi Zhang; Jiqing Zheng; Hao Wang; Yubu Su; Juncai He; Lei Liu; Yanyan Lan; | code |
| 324 | A Narrowing Geometry in Contaminated Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our mechanistic analysis reveals that contaminated models exhibit pronounced eigenspectrum concentration in their representations, leading to a low-dimensional computation regime. |
Jiakuan Xie; Pengfei Cao; Kang Liu; Jun Zhao; | code |
| 325 | PnP-Corrector: A Universal Correction Framework for Coupled Spatiotemporal Forecasting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In coupled systems, errors from each subsystem simulator propagate and amplify one another, a phenomenon we term Reciprocal Error Amplification leading to a rapid collapse of long-range predictions. To address this challenge, we propose a universal framework called PnP-Corrector (Plug-and-Play Corrector). |
Yuan Gao; Fan Xu; Hao Wu; Yuxu Lu; Penghao Zhao; Fan Zhang; Hao Jia; Yuxuan Liang; Ruijian Gou; Qingsong Wen; Xian Wu; Xiaomeng Huang; | code |
| 326 | TGV-KV: Text-Grounded KV Eviction for Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: By systematically analyzing the modality gap in VLMs in this work, we argue that the importance of visual information should be grounded in textual guidance and accordingly propose a **T**ext-**G**rounded KV Eviction method for **V**LMs (**TGV-KV**). |
Jizhihui Liu; Ruizi Han; Miao Zhang; Rui Shao; Xuebo Liu; Weili Guan; Yaowei Wang; | code |
| 327 | Event2Vec: Processing Neuromorphic Events Directly By Representations in Vector Space Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by word-to-vector models, we draw an analogy between words and events to introduce event2vec, a novel representation that allows neural networks to process events directly. |
Wei Fang; Priyadarshini Panda; | code |
| 328 | Test-Time Training Is Secretly Linear Attention Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, our analysis reveals multiple phenomena that contradict this memorization-based interpretation. Motivated by these findings, we revisit the formulation of TTT and show that a broad class of TTT architectures can be expressed as a form of learned linear attention operator. |
Junchen Liu; Sven Elflein; Or Litany; Zan Gojcic; Ruilong Li; | code |
| 329 | The Trojan Knowledge: Bypassing Commercial LLM Guardrails Via Harmless Prompt Weaving and Adaptive Tree Search Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This structure allows harmful objectives to be realized by weaving together sequences of benign sub-queries, each of which individually evades detection. To exploit this loophole, we introduce the Correlated Knowledge Attack Agent (CKA-Agent), a dynamic framework that reframes jailbreaking as an adaptive, tree-structured exploration of the target model’s knowledge base. |
Rongzhe Wei; Peizhi Niu; Xinjie Shen; Tony Tu; Yifan Li; Ruihan Wu; Eli Chien; Pin-Yu Chen; Olgica Milenkovic; Pan Li; | code |
| 330 | Native Active Perception As Reasoning for Omni-Modal Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce OmniAgent, a POMDP-based active perception framework. |
Zhenghao Xing; Ruiyang Xu; Yuxuan Wang; Jinzheng He; Ziyang Ma; Qize Yang; Yunfei Chu; Jin Xu; Junyang Lin; Chi Wing Fu; Pheng Ann Heng; | code |
| 331 | Data Selection for Fine-tuning Vision Language Models Via Cross Modal Alignment Trajectories Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose the first principled method for data-efficient instruction tuning of LVLMs. |
Dang Nguyen; Nilay Naharas; Neslihan Bulut; MohammadHossein Bateni; Vahab Mirrokni; Baharan Mirzasoleiman; | code |
| 332 | MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, this work introduces MECAT, a Multi-Expert Constructed Benchmark for Fine-Grained Audio Understanding Tasks. |
Yadong Niu; TIANZI WANG; Heinrich Dinkel; Xingwei Sun; Jiahao Zhou; Gang Li; Jizhong Liu; Xunying Liu; Jian Luan; | code |
| 333 | Diffusion-based Learning Framework for Constrained Nonconvex Optimization with Weighted Bootstrapped Refinement Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In that case, we investigate and theoretically analyze the inherent problem of supervised diffusion solvers and identify the distributional misalignment problem, i.e., the generated solution distribution often exhibits low probability mass on the feasible region. To resolve this issue, we propose DiOpt, a new diffusion-based learning framework for constrained nonconvex optimization, which effectively learns the mapping from noise to the constraint region. |
Shutong Ding; Yimiao Zhou; Ke Hu; Xi Yao; Junchi Yan; Xiaoying Tang; Ye Shi; | code |
| 334 | Sample-Efficient Diffusion-based Reinforcement Learning with Critic Guidance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Another branch pay attention to gradient-based policy optimization, which sufficiently exploit the gradient of Q function yet tend to collapse into a unimodal policy with low diversity. To address this issue, we propose CGPO, \textbf{C}ritic-\textbf{G}uided diffusion \textbf{P}olicy \textbf{O}ptimization, which effectively balances exploration and exploitation with the training-free guidance technique integrated into the denoising process of diffusion policy. |
Shutong Ding; Zejia Zhong; Zhongyi Wang; Ke Hu; Bikang Pan; Jingya Wang; Ye Shi; | code |
| 335 | Fleet: Few-Shots Lead Effective AIGI Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While achieving saturation on several AIGI benchmarks, this static hypothesis suffers a severe performance drop against rapidly evolving generators (e.g., SD3, Nano Banana Pro). To address these limitations, we propose that the field should expand beyond “static generalization” to a new paradigm of “dynamic adaptation”. |
Jiaan Wang; Sirui Liu; Yu Li; Kaiyuan Yang; Juan Cao; Sheng Tang; | code |
| 336 | Physics-informed Neural Operator Learning for Nonlinear Grad-Shafranov Equation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Focusing on the strongly nonlinear Grad-Shafranov equation (GSE) for tokamak equilibria, we propose a physics-anchored operator learning framework. |
Siqi Ding; Zitong Zhang; LiXingYu; Shi Guoyang; Xiang Gu; Y.N.Xu; Huasheng Xie; Hanyue Zhao; YUEJIANG SHI; tianyuan liu; | code |
| 337 | TIC-VLA: A Think-in-Control Vision-Language-Action Model for Robot Navigation in Dynamic Environments Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Think-in-Control (TIC)-VLA, a latency-aware framework that explicitly models delayed semantic reasoning during action generation. |
Zhiyu Huang; Yun Zhang; Johnson Liu; Rui Song; Chen Tang; Jiaqi Ma; | code |
| 338 | SpecExit: Accelerating Large Reasoning Model Via Speculative Exit Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by the use of hidden states in speculative decoding, we propose **SpecExit**, a novel framework that predicts both future tokens and an early-exit signal directly from a lightweight draft model without probing overhead. |
Rubing Yang; Huajun Bai; Song Liu; Guanghua Yu; Runzhi Fan; Yanbin Dang; Zhang Jiejing; Kai Liu; Jianchen Zhu; Peng Chen; | code |
| 339 | PromptRL: Prompt Matters in RL for Flow-Based Image Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this research, we show that current RL pipelines for FMs suffer from two underappreciated yet important limitations: sample inefficiency due to insufficient generation diversity, and pronounced prompt overfitting, where models memorize specific training formulations and exhibit dramatic performance collapse when evaluated on semantically equivalent but stylistically varied prompts. |
Fu-Yun Wang; Han Zhang; Michaël Gharbi; Hongsheng Li; Taesung Park; | code |
| 340 | Light Up Your Face: A Physically Consistent Dataset and Diffusion Model for Face Fill-Light Enhancement Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To support scalable learning, we introduce LightYourFace-160K (LYF-160K), a large-scale paired dataset built with a physically consistent renderer that injects a disk-shaped area fill light controlled by six disentangled factors, producing 160K before-and-after pairs. |
Jue Gong; Zihan Zhou; Jingkai Wang; Xiaohong Liu; Yulun Zhang; Xiaokang Yang; | code |
| 341 | Autoregressive Image Generation with Masked Bit Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing discrete generation methods struggle to capitalize on this insight, suffering from performance degradation or prohibitive training costs with scaled codebook. To address this, we propose masked **B**it **A**uto**R**egressive modeling (**BAR**), a scalable framework that supports arbitrary codebook sizes. |
Qihang Yu; Qihao Liu; Ju He; Xinyang Zhang; Yang Liu; Liang-Chieh Chen; Peter Chen; | code |
| 342 | STEP: Warm-Started Visuomotor Policies with Spatiotemporal Consistency Prediction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose **STEP**, a lightweight spatiotemporal consistency prediction mechanism to construct high-quality warm-start actions that are both distributionally close to the target action and temporally consistent, without compromising the generative capability of the original diffusion policy. |
Jinhao Li; Yuxuan Cong; Yingqiao Wang; Hao Xia; Shan Huang; Yijia Zhang; Ningyi Xu; Guohao Dai; | code |
| 343 | CyberJurors: A Multi-Agent Simulation Task for E-Commerce Disputes Verdict Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Motivated by this, we propose a pioneering task, **E-commerce Dispute Verdict** (EDV), and introduce **VerdictBench**, the first Multimodal Disputes Verdicts Benchmark for E-commerce, to facilitate the intelligent verdicts. Building upon this, we propose **CyberJurors**, a framework that integrates an Individual Verdict Chain-of-Thought (IV-CoT) and Jury Consensus Verdict (JCV) to clarify the dispute logic and regulate the fair verdict process. |
Yanhui Sun; Wu Liu; Haifeng Ming; Xinru Wang; Hantao Yao; Yongdong Zhang; | code |
| 344 | UCPO: Uncertainty-Aware Policy Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing RL paradigms such as GRPO often suffer from Advantage Bias due to binary decision spaces and static uncertainty rewards, inducing either excessive conservatism or overconfidence. To tackle this challenge, this paper unveils the root causes of reward hacking and overconfidence in current RL paradigms incorporating uncertainty-based rewards, based on which we propose the UnCertainty-Aware Policy Optimization (UCPO) framework. |
Xianzhou Zeng; Jing Huang; Chunmei Xie; Gongrui Nan; Siye Chen; Mengyu Lu; Weiqi Xiong; Qixuan Zhou; Junhao Zhang; Qiang Zhu; Yadong Li; Xingzhong Xu; | code |
| 345 | TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Text-Visual Interleaved Chain-of-Thought (TVI-CoT), a framework that enables explicit interleaving of textual reasoning and visual feature access through learnable control tokens ($\langle\text{Think}\rangle$, $\langle\text{Look}\rangle$, $\langle\text{Answer}\rangle$). |
Lianyu Hu; Xiaoyu Ma; Zeqin Liao; Yang Liu; | code |
| 346 | S2GS: Streaming Semantic Gaussian Splatting for Online Scene Understanding and Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Streaming Semantic Gaussian Splatting (S2GS), a strictly causal, incremental 3D Gaussian semantic field framework: it does not leverage future frames and continuously updates scene geometry, appearance, and instance-level semantics without reprocessing historical frames, enabling scalable online joint reconstruction and understanding. |
Renhe Zhang; Yuyang Tan; Jingyu Gong; zhizhong zhang; Lizhuang Ma; Yuan Xie; Xin Tan; | code |
| 347 | ActiveUltraFeedback: Efficient Preference Data Generation Using Active Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Reinforcement Learning from Human Feedback (RLHF) has become the standard for aligning Large Language Models (LLMs), yet its efficacy is bottlenecked by the high cost of acquiring preference data, especially in low-resource and expert domains. To address this, we introduce ActiveUltraFeedback, a modular active learning pipeline that leverages uncertainty estimates to dynamically identify the most informative responses for annotation. |
Davit Melikidze; Marian Schneider; Jessica Lam; Martin Wertich; Ido Hakimi; Barna Pasztor; Andreas Krause; | code |
| 348 | InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks Via Rubric-Based Incremental Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In such settings, RL must either rely on supervision-intensive reward models that often fail to generalize, or it falls into pathological behaviors such as reward hacking—an especially troubling risk for high-stakes medical dialogue. To address these limitations, we introduce ORBIT, an open-ended rubric-based incremental training framework for high-stakes medical dialogue. |
Pengkai Wang; Pengwei Liu; Qi Zuo; Zhijie Sang; Congkai Xie; Hongxia Yang; | code |
| 349 | On Revisiting Entropy for Identifying Mislabeled Medical Images Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Mislabeled samples in training datasets severely degrade the performance of deep networks, as overparameterized models tend to memorize erroneous labels. We address this challenge by proposing a novel approach for mislabeled data detection that leverages training dynamics. |
Chunlei Li; Zixuan Zheng; Yilei Shi; Guanglu Dong; Pengfei Li; Jingliang Hu; Xiao Zhu; Lichao Mou; | code |
| 350 | BlueCodeAgent: A Blue Teaming Agent Powered By Automated Red Teaming for CodeGen AI Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, progress on the blue teaming side remains limited, as effective defenses require a deep security analysis of given tasks and edge cases. To fill in this gap, we propose BlueCodeAgent, an end-to-end blue teaming agent powered by automated red teaming. |
Chengquan Guo; Yuzhou Nie; Chulin Xie; Zinan Lin; Wenbo Guo; Bo Li; | code |
| 351 | The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We compare MoE experts and dense FFNs using $k$-sparse probing and find that expert neurons are consistently less polysemantic, with the gap widening as routing becomes sparser. |
Jeremy Herbst; Jae Hee Lee; Stefan Wermter; | code |
| 352 | GeoPT: Scaling Physics Simulation Via Lifted Geometric Pre-Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present GeoPT, a unified pre-trained model for general physics simulation based on lifted geometric pre-training. |
Haixu Wu; Minghao Guo; Zongyi Li; Zhiyang Dou; Mingsheng Long; Kaiming He; Wojciech Matusik; | code |
| 353 | CREDIT: Certified Ownership Verification of Deep Neural Networks Against Model Extraction Attacks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While numerous defense strategies have been proposed, verifying the ownership of a suspicious model with strict theoretical guarantees remains a challenging task. To address this gap, we introduce CREDIT a certified defense against MEAs. |
Bolin Shen; Zhan Cheng; Neil Gong; Fan Yao; Yushun Dong; | code |
| 354 | LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Based on AIGVE-60K, we propose **LOVE**, a LMM-based metric for AIGV Evaluation from multiple dimensions including perceptual preference, text-video correspondence, and task-specific accuracy. |
Jiarui Wang; Huiyu Duan; Ziheng Jia; Zicheng Zhang; Yu Zhao; Juntong Wang; Guangtao Zhai; Xiongkuo Min; | code |
| 355 | The Tell-Tale Norm: $\ell_2$ Magnitude As A Signal for Reasoning Dynamics in Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent work has sought to understand Large Language Models (LLMs) reasoning, yet a principled, model-intrinsic signal that captures its *layer-wise reasoning dynamics* remains underexplored. We bridge this gap by demonstrating that **the $\ell_2$ norm of hidden states serves as an endogenous signal of the model’s reasoning intensity**. |
Jinyang Zhang; Hongxin Ding; Yue Fang; Weibin Liao; Muyang Ye; Junfeng Zhao; Yasha Wang; | code |
| 356 | Less Is More: Elevating RAG Via Performance-Driven Context Compression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These heuristics fail to ensure that the compressed context is conducive to the generation tasks. To address these limitations, we propose CORE-RAG, a novel framework for context compression in RAG systems. |
Ziqiang Cui; Yunpeng Weng; Xing Tang; Peiyang Liu; Shiwei Li; Bowei He; Jiamin Chen; Yansen Zhang; xiuqiang He; Rui Zhang; Chen Ma; | code |
| 357 | Training-Free Hashing-Based Attention Via Binary Principal Components Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present \textbf{BinaryPC}, a training-free, data-aware hashing-based sparse attention for long-context LLMs. |
Daohai Yu; Zhanpeng Zeng; Keyu Chen; Wenhao Li; Zhifeng Shen; Luxi Lin; Ruizhi Qiao; Xing Sun; Rongrong Ji; | code |
| 358 | Active Exploring Like A Pigeon: Reinforcing Spatial Reasoning Via Agentic Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by pigeons’ building and exploiting cognitive maps for navigation, we propose a novel agentic pipeline for spatial reasoning. |
Wei Deng; Xianlin Zhang; Mengshi Qi; | code |
| 359 | MedCRP-CL: Continual Medical Image Segmentation Via Bayesian Nonparametric Semantic Modality Discovery Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce MedCRP-CL, a framework that performs online task structure discovery and structure-aware continual learning. |
Ziyuan Gao; | code |
| 360 | DASH: Faster Shampoo Via Batched Block Preconditioning and Efficient Inverse-Root Solvers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we take a significant step to address this shortcoming by proposing \method (for Distributed Accelerated SHampoo), a faster implementation of Distributed Shampoo based on two main new techniques: First, we show that preconditioner blocks can be stacked into 3D tensors to significantly improve GPU utilization; second, we introduce the Newton-DB iteration and the Chebyshev polynomial approximations as novel and faster approaches for computing the inverse matrix roots required by Shampoo. Along with these algorithmic contributions, we provide a first in-depth analysis of how matrix scaling critically affects Shampoo convergence. |
Ionut-Vlad Modoranu; Philip Zmushko; Erik Schultheis; Mher Safaryan; Dan Alistarh; | code |
| 361 | Reward Modeling from Natural Language Human Feedback Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, we demonstrate that such binary classification tasks make GRMs susceptible to guessing correct outcomes without sound critiques, introducing noise into the reward signal and impairing learning effectiveness. To address this, we propose Reward Modeling from Natural Language Human Feedback (RM-NLHF), which leverages natural language feedback to obtain process reward signals. |
Zongqi Wang; Rui Wang; Yuchuan Wu; Yiyao Yu; Pinyi Zhang; Shaoning Sun; Yujiu Yang; Yongbin Li; | code |
| 362 | Weasel: Out-of-Domain Generalization for Web Agents Via Importance-Diversity Data Selection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address both issues, we propose Weasel, a trajectory selection method for offline training of web agents. |
FATEMEH PESARAN ZADEH; Seyeon Choi; Xing Han Lù; Siva Reddy; Gunhee Kim; | code |
| 363 | DTKG: Dual-Track Knowledge Graph-Verified Reasoning Framework for Multi-Hop QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by the Dual Process Theory in cognitive science and Stanovich’s Cognitive Misers Theory, we propose an effective multi-hop QA framework DTKG (Dual-Track Knowledge Graph) through building a two-stage pipeline: i) Classification Stage (dynamic question categorization via few-shot prompting, emulating unconscious processing); and ii) Branch Processing Stage (tailored reasoning paths, emulating conscious processing). |
Changhao Wang; Yanfang Liu; Xinxin Fan; Lanzhi Zhou; Ao Tian; Yunfeng Lu; | code |
| 364 | SC$^{2}$-WM: A Self-Correcting World Model with Closed-Loop Feedback for Vision-and-Language Navigation in Continuous Environments Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose SC$^{2}$-WM, a self-correcting world model framework that introduces internal feedback for closed-loop decision making in VLN-CE. |
Xuan Yao; Yuze Zhu; JUNYU GAO; Zongmeng Wang; Changsheng Xu; | code |
| 365 | Multi-Adapter Representation Interventions Via Energy Calibration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, we find that the appropriate intervention direction and strength vary substantially across samples, and such indiscriminate intervention leads to degradation of general capabilities on benign inputs. To address these challenges, we propose Multi-Adapter Representation Interventions via Energy Calibration (MARI). |
Manjiang Yu; Hongji Li; Junwei Chen; Xue Li; Priyanka Singh; YANG CAO; Lijie Hu; | code |
| 366 | Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Per-Token Distance (PTD) to quantify cross-modal positional disentanglement, and we prove that $\mathrm{PTD}=0$ is a sufficient condition to eliminate the geometric attention bias induced by RoPE. |
Chengcheng Wang; Jianyuan Guo; Hongguang Li; Yuchuan Tian; Ying Nie; Chang Xu; Kai Han; | code |
| 367 | AgentVocab: Structure-Aware Vocabulary Adaptation for Efficient LLM Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This training–deployment mismatch leads to inefficient tokenization, where repetitive structural patterns and frequent semantic units in function calls are fragmented into long sequences of low-level tokens, increasing decoding overhead. To address this gap, we introduce $\textbf{AgentVocab}$, a structure-aware vocabulary adaptation framework for efficient LLM agents. |
Kai Bian; Haosi Mo; Xuebo Liu; Shuangyong Song; Jing Li; Yongxiang Li; Min zhang; Xuelong Li; | code |
| 368 | Don’t Forget Why You Started: Tackling Dual Forgetting in Vision-Language Continual Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Unlike standard averaging, SCR utilizes the frozen pre-trained feature space to inversely weight performance based on semantic similarity, effectively mitigating the confounding gains to stress-test foundational stability. Building on this, we propose DFA-MoE, a functionally heterogeneous Parameter-Efficient Fine-Tuning (PEFT) method. |
Borui Kang; Jinrui Gu; Tao Feng; Qi Fan; Yinghuan Shi; Lei Wang; Wenbin Li; Yang Gao; | code |
| 369 | From Prompts to Responses: Dual-Sided Data Leakage and Defense in Split Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we unveil novel vulnerabilities of Split-LLM by presenting**P**atched Model **I**nversion with **D**ual-Sided **I**nitialization(**PIDI**), a two-stage attack that simultaneously targets both private input prompts and output responses in Split-LLM settings. |
Zixuan GU; Xiaojun Ye; Yang Liu; | code |
| 370 | HiPhO: How Far Are (M)LLMs from Humans in The Latest High School Physics Olympiad Benchmark? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing physics benchmarks suffer from two major gaps: they neither provide systematic and up-to-date coverage of physics Olympiads, nor enable direct performance comparison with humans. To bridge these gaps, we present **HiPhO**, the first benchmark dedicated to high school physics Olympiads with human-aligned evaluation. |
Fangchen Yu; Haiyuan Wan; Qianjia Cheng; Yuchen Zhang; Jiacheng Chen; Fujun Han; Yulun Wu; Junchi Yao; Ruilizhen Hu; Ning Ding; Yu Cheng; Tao Chen; LEI BAI; Dongzhan Zhou; Yun Luo; Ganqu Cui; Peng Ye; | code |
| 371 | PromptDyG: Test-Time Prompt Adaptation on Dynamic Graphs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This fails to capture 1) the nature of evolving complexities across graph snapshots and 2) the distribution shift in the testing graph snapshots. To address these problems, we propose PromptDyG, a novel framework that leverages unsupervised test-time Prompt adaptation for Dynamic Graph learning under a live-update online setting. |
GUOGUO AI; Chaoxi Niu; Hui Yan; Joey Tianyi Zhou; Yew Soon ONG; Guansong Pang; | code |
| 372 | ViEEG: Hierarchical Visual Neural Representation for EEG Brain Decoding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by the hierarchical organization of visual cortex, we propose ViEEG, a neuro-inspired framework that addresses HNEN. |
Minxu Liu; Donghai Guan; Chuhang Zheng; Chunwei Tian; Jie Wen; Qi Zhu; | code |
| 373 | Learning to Route Languages for Multilingual Preference Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose language-routed preference optimization (LRPO), an online preference optimization framework that treats language as a selectable variable rather than a fixed input constraint. |
Geyang Guo; Hiromi Wakaki; Yuki Mitsufuji; Alan Ritter; Wei Xu; | code |
| 374 | SafeHarbor: Defining Precise Decision Boundaries Via Hierarchical Memory-Augmented Guardrail for LLM Agent Safety Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While existing defensive mechanisms are effective, they frequently suffer from the over-refusal problem, where increased safety strictness compromises the agent’s utility on benign tasks. To mitigate this trade-off, we propose \textsc{SafeHarbor}, a novel framework designed to establish precise decision boundaries for LLM agents. |
Zhe Liu; Zonghao Ying; Wenxin Zhang; Quanchen Zou; Deyue Zhang; Dongdong Yang; Xiangzheng Zhang; Hao Peng; | code |
| 375 | Deliberate Evolution for Sample-Efficient Symbolic Regression with LLM Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Consequently, optimization often degenerates into an inefficient search with substantial computational cost. Motivated by these limitations, we propose Deliberate Evolution, an agentic framework for SR tasks that equips LLM-based candidate proposal with explicit, structured guidance. |
Xinyu Pang; Zhanke Zhou; Xuan Li; Fangrui Lv; Shanshan Wei; Sen Cui; Bo Han; Changshui Zhang; | code |
| 376 | EgoTactile: Learning Grasp Pressure for Everyday Objects from Egocentric Video Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Leveraging this benchmark, we first establish EgoPressureFormer as a discriminative baseline. Beyond this, to explicitly address the uncertainty in partial observations, we propose EgoPressureDiff, a conditional diffusion framework that adapts a large-scale pre-trained video diffusion backbone. |
Yuan Zeng; Yujia Shi; Tiao Tan; Xingting Li; Yaqi Qin; Zongqing Lu; Wenming Yang; Jing-Hao Xue; Qingmin Liao; | code |
| 377 | More Edits, More Stable: Understanding The Lifelong Normalization in Sequential Model Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we provide the **first** theoretical account of LN in the lifelong regime. |
Xin Ma; Wei Chen; Qi Liu; Derong Xu; Zhi Zheng; Tong Xu; Enhong Chen; | code |
| 378 | CLAM-Bench: Benchmarking LLM Agents for Library-Scale Cross-Architecture Migration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present CLAM-Bench (Cross-architecture Library-scale Agent Migration benchmark), featuring 85 critical kernels from widely used libraries, including OpenCV, libjpeg, and NCNN. |
Weijia Li; KE GAO; Jiajie Li; Han Sun; Yuhe Ding; Yongdong Mai; Yiran Le; Yongjie Qian; Zhibin Zhang; Xinyu Wang; Limin Cheng; Shouxu Kuang; Pengfei Chen; Ling Li; | code |
| 379 | What Makes Synthetic Data Effective in Image Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In particular, synthetic images characterized by dense scene composition and fine instance fidelity demonstrate distinctive benefits, yielding significantly more discriminative spatial representations. Building on these insights, we propose SENSE, a unified framework that leverages flexible and scalable synthetic data to substantially enhance segmentation performance. |
Jinjin Zhang; Xiefan Guo; Yizhou jin; Nan Zhou; Di Huang; | code |
| 380 | ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce **ParisKV**, a drift-robust, GPU-native KV-cache retrieval framework based on collision-based candidate selection, followed by a quantized inner-product reranking estimator. |
Yanlin Qi; Xinhang Chen; Huiqiang Jiang; Qitong Wang; Botao Peng; Themis Palpanas; | code |
| 381 | ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we reformulate machine unlearning as a precise knowledge re-mapping problem via model editing. |
Yujie Lin; Chengyi Yang; Zhishang Xiang; YIPING SONG; Jinsong Su; | code |
| 382 | End-to-end Graph-structured Brain Representation Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: By drawing on representation learning, this paper presents the Brain Representation (BRep) learning problem. |
Liang Yang; Shuai Zhai; Ziyi Ma; Jiaming Zhuo; Di Jin; Chuan Wang; Zhen Wang; Xiaochun Cao; | code |
| 383 | General and Efficient Steering of Unconditional Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a general recipe for efficiently steering unconditional diffusion *without gradient guidance during inference*, enabling fast controllable generation. |
Qingsong Wang; Misha Belkin; Yusu Wang; | code |
| 384 | RaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM Inference Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Long-context Large Language Model inference is severely bottlenecked by the massive Key-Value (KV) cache, yet existing sparse attention methods often suffer from static fixed-budget (Top-k) retrieval or rely on proxy scores that are computationally expensive and biased. To address these limitations, we propose RaBitQCache, a novel sparse attention framework that utilizes randomized rotated binary quantization and high-throughput binary-INT4 arithmetic to efficiently estimate attention weights. |
Wenhao Li; Jinhao Dong; Hailin Zhang; Shi; WEI LU; Xiaoyong Du; | code |
| 385 | TEAM: Temporal–Spatial Consistency Guided Expert Activation for MoE Diffusion Language Model Acceleration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose **TEAM**, a plug-and-play framework that accelerates MoE dLLMs by enabling more accepted tokens with fewer activated experts. |
LINYE WEI; Zixiang Luo; Pingzhi Tang; Meng Li; | code |
| 386 | Resolution As A Direction: Vector-Panning Feature Alignment for Cross-Resolution Re-Identification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we report a new empirical finding: after averaging out identity-specific variation, the HR–LR feature discrepancy produced by standard ReID backbones exhibits a consistent, resolution-related semantic direction in the embedding space. |
Zanwu Liu; Chao Yuan; Bo Li; Xiaowei Zhang; Guanglin Niu; | code |
| 387 | MoSE: Mixture of Slimmable Experts for Efficient and Adaptive Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Mixture of Slimmable Experts (MoSE), an MoE architecture in which each expert has a nested, slimmable structure that can be executed at variable widths. |
Nurbek Tastan; Stefanos Laskaridis; Karthik Nandakumar; Samuel Horváth; | code |
| 388 | MiniX: Mitigating Low-Rank Collapse and Attention Bottlenecks in Tabular Foundation Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We diagnose a low-rank collapse induced by the prevalent linear+ID scheme and introduce RaBEL, a compact Radial Basis Embedding Layer that front-loads nonlinearity via localized RBF features. |
Yuanrui Wang; Xingxuan Zhang; Han Yu; Mingchao Hao; Gang Ren; hao yuan; Li Mao; Yunjia Zhang; Chun Yuan; Peng Cui; | code |
| 389 | Unsupervised Diffusion for Combinatorial Optimization Via Adjoint Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, training diffusion models typically requires large collections of near-optimal solutions, which limits their scalability and generalization. We address this fundamental challenge by extending adjoint matching, a powerful unsupervised diffusion training framework based on chain-rule–style gradient propagation in continuous spaces, to discrete combinatorial domains. |
Shengyu Feng; Tarun Suresh; Yiming Yang; | code |
| 390 | DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we novelly propose DRIFT (Decoupled Rollouts and Importance-Weighted Fine-Tuning), a framework that operationalizes the theoretical insight that the KL-regularized RL objective is mathematically equivalent to importance-weighted supervised learning. |
Jian Mu; Tianyi Lin; Chengwei Qin; Zhongxiang Dai; Yao Shu; | code |
| 391 | CoGenCast: A Coupled Autoregressive–Flow Generative Framework for Time Series Forecasting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose CoGenCast, a hybrid generative framework that couples pre-trained LLMs with flow-matching mechanism for effective time series forecasting. |
Yaguo Liu; Mingyue Cheng; Daoyu Wang; Xiaoyu Tao; Qi Liu; | code |
| 392 | Towards Realistic Lifelong Re-identification: Identity Recurrence with Changing Clothes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address these, we develop a framework that disentangles identity-intrinsic representations from clothing-induced biases, enabling identity modeling beyond appearance changes. |
Wuxuan Shi; Zhijie Lu; He Li; Mang Ye; | code |
| 393 | OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Optimal Brain Cache (OBCache), a principled framework that formulates cache eviction as a layer-wise structured pruning problem. |
Yuzhe Gu; Xiyu Liang; Jiaojiao Zhao; Enmao Diao; | code |
| 394 | Improving CLIP Adaptation By Breaking Tail Alignment for Source-Free Cross-Domain Few-Shot Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we focus on the target-domain few-shot finetuning in the CLIP-based CDFSL task. |
Shuai Yi; Yixiong Zou; Yuhua Li; Ruixuan Li; | code |
| 395 | Geometry-Preserving Unsupervised Alignment for Heterogeneous Foundation Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose **GPUA**, a *Geometry-Preserving Unsupervised Alignment* framework that integrates the complementary strengths of VFMs and VLMs. |
Shuwen Yu; Zhanxuan Hu; Yi Zhao; Yonghang Tai; Huafeng Li; | code |
| 396 | From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing Via Generative Bootstrapping Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, masking inevitably destroys spatiotemporal context, leading to identity drift and poor robustness (e.g., to occlusions), while also inducing lip-shape leakage that degrades lip sync. To bridge this gap, we propose X-Dub, a novel two-stage generative bootstrapping framework leveraging powerful Diffusion Transformers to unlock mask-free dubbing. |
Xu He; Haoxian Zhang; Hejia Chen; Changyuan Zheng; Liyang Chen; Songlin Tang; Jiehui Huang; Xiaoqiang Liu; Pengfei Wan; Zhiyong Wu; | code |
| 397 | The Cylindrical Representation Hypothesis for Language Model Steering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: By relaxing LRH’s orthogonality assumption while preserving linear representations, we show that overlapping concept contributions naturally yield a sample-specific axis-orthogonal structure. |
Lang Gao; Jinghui Zhang; Wei Liu; Fengxian Ji; Chenxi Wang; Zirui Song; Akash Ghosh; Youssef Mohamed; Preslav Nakov; Xiuying Chen; | code |
| 398 | RePro: Training Language Models to Faithfully Recycle The Web for Pretraining Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce RePro, a novel web recycling method that trains a relatively small LM with reinforcement learning to generate effective and faithful rephrasings of pretraining data. |
Zichun Yu; Chenyan Xiong; | code |
| 399 | Demystifying When Pruning Works Via Representation Hierarchies Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, this expectation does not consistently hold across language tasks: pruned models can perform well on non-generative tasks but frequently fail in generative settings. To demystify how such discrepancies arise under pruning, we analyze network pruning from a representation-hierarchy perspective, decomposing the internal computation of language models into three sequential spaces: \textit{embedding} (hidden representations), \textit{logit} (pre-softmax outputs), and \textit{probability} (post-softmax distributions). |
Shwai He; Guoheng Sun; Haichao Zhang; Yun Fu; Ang Li; | code |
| 400 | Hallucination Detection from Structural Reasoning Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a \emph{structural reasoning model} to describe the interactions among local steps. |
Jianbo Sun; Pengkun Yang; | code |
| 401 | Flow Inverse Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Unfortunately, optimizing FM policies via online interaction is challenging and inefficient due to instability in gradient computation and high inference costs. To address these issues, we propose to let a student” policy with simple MLP structure to explore the environment and be online updated via RL algorithm with a reward model. |
Zhenglin Wan; Jingxuan Wu; Xingrui Yu; Chubin Zhang; Mingcong Lei; Bo An; Ivor Tsang; Yang You; | code |
| 402 | Know More, Know Clearer: A Meta-Cognitive Framework for Knowledge Augmentation in Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing methods typically operate on the simplistic premise that model performance equates with internal knowledge, overlooking the knowledge-confidence gaps that lead to overconfident errors or uncertain truths. To bridge this gap, we propose a novel meta-cognitive framework for reliable knowledge augmentation via differentiated intervention and alignment. |
Hao Chen; Ye He; Yuchun Fan; Yukun Yan; Zhenghao Liu; Qingfu Zhu; Maosong Sun; Wanxiang Che; | code |
| 403 | CoCoEdit: Content-Consistent Image Editing Via Region Regularized Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a post-training framework for \textbf{Co}ntent-\textbf{Co}nsistent \textbf{Edit}ing (\textbf{CoCoEdit}) by using region regularized reinforcement learning. |
Yuhui WU; Chenxi Xie; Ruibin Li; Liyi Chen; Qiaosi Yi; Lei Zhang; | code |
| 404 | Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in Modern Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Extending to the multimodal setting reveals a fundamental learning asymmetry: when pretrained on high-diversity data from a primary modality, surprisingly low data complexity in the secondary modality suffices for multimodal ICL to emerge. Mechanistic analysis shows that both settings rely on an induction-style mechanism that copies labels from matching in-context exemplars; multimodal training refines and extends these circuits across modalities. |
Yiran Huang; Karsten Roth; Quentin Bouniot; Wenjia Xu; Zeynep Akata; | code |
| 405 | EnerGS: Energy-Based Gaussian Splatting Under Partial Geometric Observability Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, in large-scale outdoor scenarios, such geometric supervision is often spatially incomplete and uneven, which limits its effectiveness as a reliable prior and can even be detrimental to the final reconstruction. To address this challenge, we model partially observable geometry as a continuous energy field induced by geometric evidence and propose EnerGS. |
Rui Song; Tianhui Cai; Markus Gross; Yun Zhang; Walter Zimmer; Zhiyu Huang; Olaf Wysocki; Jiaqi Ma; | code |
| 406 | MESA: Improving MoE Safety Alignment Via Decentralized Expertise Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Meanwhile, conventional alignment methods apply uniform adaptations across all parameters, ignoring their functional differences and inadvertently degrading general utility. To address these challenges, we propose MESA (MoE Safety Alignment), a targeted alignment framework for MoE-based LLMs that strategically decentralizes safety responsibilities to maximize coverage while explicitly minimizing interference with general capabilities. |
Yitong Sun; Yao Huang; Teng Li; Ranjie Duan; Yichi Zhang; Xingjun Ma; Hui Xue'; Xingxing Wei; | code |
| 407 | Provable Sample Efficiency of Curriculum Post-Training for Transformer Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent curriculum techniques in the post-training stage of LLMs have been empirically observed to outperform non-curriculum approaches in improving reasoning performance, yet a principled understanding of their effectiveness and limitations remains incomplete. To bridge this gap, we develop an abstract theoretical framework and identify sufficient conditions under which curriculum post-training yields exponential improvements in sample complexity. |
Dake Bu; Wei Huang; Andi Han; Atsushi Nitanda; Hau-San Wong; Qingfu Zhang; Taiji Suzuki; | code |
| 408 | BiTrajDiff: Bidirectional Trajectory Generation with Diffusion Models for Offline Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As a result, such methods often introduce local variations and fail to recover connections between distinct behavior patterns. In this paper, we propose Bidirectional Trajectory Diffusion (BiTrajDiff), a novel DA framework that explicitly addresses this limitation. |
Yunpeng Qing; Yixiao Chi; Shuo Chen; Shunyu Liu; Kexuan Zhou; Sixu Lin; Litao Liu; Changqing Zou; | code |
| 409 | Gated Relational Alignment Via Confidence-based Distillation for Efficient VLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite its potential, quantization-aware training for VLMs remains underexplored. We propose GRACE, a framework unifying knowledge distillation and QAT under the Information Bottleneck principle: quantization constrains information capacity while distillation guides what to preserve within this budget. |
Yanlong Chen; Amir Habibian; Luca Benini; Yawei Li; | code |
| 410 | Geometric Reciprocity: Unlocking Self-Supervision for Stereoscopic Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our key contribution is the **Geometric Reciprocity Theorem (GRT)**: the disocclusion mask when synthesizing a target view exactly equals the mask of pixels lost when warping back from target to source, enabling analytical computation of test-time disocclusion masks directly from monocular images. |
Jingyi Lu; Kai Han; | code |
| 411 | UniCode: Augmenting Evaluation for Code Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current coding benchmarks often inflate Large Language Model (LLM) capabilities due to static paradigms and data contamination, enabling models to exploit statistical shortcuts rather than genuine reasoning. To address this, we introduce \textbf{UniCode}, a generative evaluation framework that systematically probes LLM limits via: (1) multi-dimensional augmentation transforming seed problems into complex variations to disrupt fixed algorithmic patterns; (2) a highly reliable, automated test generation pipeline for scalable evaluation; and (3) fine-grained metrics for rich error signals. |
Xinyue Zheng; Haowei Lin; Shaofei Cai; Zilong Zheng; Yaodong Yang; Yitao Liang; | code |
| 412 | CADFit: Precise Mesh-to-CAD Program Generation with Hybrid Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce CADFit, a hybrid optimization-based CAD reconstruction framework that recovers complex, editable CAD construction sequences from meshes by incrementally fitting and validating parametric operations using geometric feedback. |
Ghadi Nehme; Eamon Whalen; Faez Ahmed; | code |
| 413 | EvReflection: Event-Driven Micro-Dynamics for Reflection Removal Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Due to the inherent ambiguity between layers, existing methods still suffer from severe residual artifacts. In this paper, we propose leveraging event signals to break this ambiguity. |
Jiaxiao Wang; Dachun Kai; Huyue Zhu; Quanquan Hu; Zhenyang Xu; Xiaoyan Sun; | code |
| 414 | Zeus: Towards Tuning-Free Foundation Model for Time Series Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Zeus, a unified tuning-free Time Series Foundation Model (TSFM) that delivers superior performance across diverse analysis tasks without any task-specific fine-tuning. |
Yisong Fu; Zezhi Shao; Chengqing Yu; Yujie Li; Yongjun Xu; Xueqi Cheng; Fei Wang; | code |
| 415 | Bayesian Gated Non-Negative Contrastive Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In compositional scenes, this creates an Optimization Conflict: common background features, such as, blue sky, are encouraged to align in positive pairs but simultaneously repelled in negative pairs, causing gradient oscillations that hinder precise semantic disentanglement. To address this, we propose **BayesNCL** (Bayesian Gated Non-Negative Contrastive Learning). |
Peng Cui; Jiahao Zhang; Lijie Hu; | code |
| 416 | MemCast: Memory-Driven Time Series Forecasting with Experience-Conditioned Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose MemCast, a learning-to-memory framework that reformulates TSF as an experience-conditioned reasoning task. |
Xiaoyu Tao; Mingyue Cheng; Ze Guo; Shuo Yu; Yaguo Liu; Qi Liu; Shijin Wang; | code |
| 417 | HypRAG: Hyperbolic Dense Retrieval for Retrieval Augmented Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, natural language exhibits hierarchical structure from broad topics to specific entities that Euclidean embeddings fail to preserve, causing semantically distant documents to appear spuriously similar and increasing hallucination risk. To address these limitations, we introduce hyperbolic dense retrieval, developing two model variants in the Lorentz model of hyperbolic space: HyTE-FH, a fully hyperbolic transformer, and HyTE-H, a hybrid architecture projecting pre-trained Euclidean embeddings into hyperbolic space. |
Hiren Madhu; Ngoc Bui; Ali Maatouk; Leandros Tassiulas; Smita Krishnaswamy; Menglin Yang; Sukanta Ganguly; Kiran Srinivasan; ZHITAO YING; | code |
| 418 | FairSSL: Fair Multimodal Self-Supervised Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose FairSSL, a framework that leverages data heterogeneity as a resource for fairness rather than a hindrance. |
Jiaee Cheong; Abtin Mogharabin; Paul Pu Liang; Hatice Gunes; Sinan Kalkan; | code |
| 419 | SyMerge: From Non-Interference to Synergistic Merging Via Single-Layer Adaptation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This study proposes SyMerge, a lightweight framework that jointly optimizes merging coefficients and a single task-specific layer. |
Aecheon Jung; Seunghwan Lee; Dongyoon Han; Sungeun Hong; | code |
| 420 | Anchor-Final Self-Supervision Drives Hallucination-Aware Optimization in Large Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Hallucinations in large vision-language models (LVLMs) remain a critical challenge, with models often generate tokens that fail to align with visual evidence. To address this issue, we propose AFS: Anchor-Final Self-Supervision, a novel framework for hallucination-aware optimization in LVLMs. |
Jiaxi Liu; Yifeng Yang; Xinbing Wang; Qinying Gu; Nanyang Ye; | code |
| 421 | LEMUR: Learned Multi-Vector Retrieval Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce LEMUR, a simple-yet-efficient framework for multi-vector similarity search. |
Elias Jääsaari; Ville Hyvönen; Teemu Roos; | code |
| 422 | SkelHCC: A Hyperbolic CLIP-Driven Cache Adaptation Framework for Skeleton-based One-Shot Action Recognition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose SkelHCC, a unified skeleton hyperbolic CLIP-driven cache adaptation framework for one-shot skeleton-based action recognition. |
Yanan Liu; Anqi Zhu; Jingmin Zhu; Jun Liu; Hossein Rahmani; Mohammed Bennamoun; Farid Boussaid; Dan Xu; Qiuhong Ke; | code |
| 423 | Orthogonal Concept Erasure for Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As additive updates inherently entangle direction, magnitude, and angular geometry, they inevitably introduce unintended interference between concept erasure and overall generation performance. To address this, we propose **Orthogonal Concept Erasure (OCE)**, which reformulates editing-based erasure as multiplicative parameter updates from a geometric perspective. |
Yuhao Sun; Lingyun Yu; Hao-Xiang Xu; Fengyuan Miao; Zhuoer Xu; Hongtao Xie; | code |
| 424 | Delving Into Muon and Beyond: Deep Analysis and Extensions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The Muon optimizer has recently attracted considerable attention for its strong empirical performance and use of orthogonalized updates on matrix-shaped parameters, yet its underlying mechanisms and relationship to adaptive optimizers such as Adam remain insufficiently understood. In this work, we aim to address these questions through a unified spectral perspective. |
Xianbiao Qi; Marco Chen; Jiaquan Ye; Yelin He; Rong Xiao; | code |
| 425 | ShapCCS: Shapley-Driven Client Coreset Selection in Federated Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Computation overhead has emerged as a critical bottleneck in Federated Learning (FL). Coreset selection tackles this challenge by constructing an informative subset to represent … |
Shuo Ji; Jie Hu; Zhouqiao He; Zijie Zhao; Tianrui Li; Jie Xu; | code |
| 426 | MedScope: Incentivizing Think with Videos for Clinical Reasoning Via Coarse-to-Fine Tool Calling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, current multimodal large language models typically process videos with passive sampling or weakly grounded inspection, which limits their ability to iteratively locate, verify, and justify predictions with temporally targeted evidence. To close this gap, we propose **MedScope**, a tool-using clinical video reasoning model that performs coarse-to-fine evidence seeking over long-form procedures. |
Wenjie Li; Yujie Zhang; Haoran Sun; Xingqi He; Hongcheng Gao; Chenglong Ma; Ming Hu; Guankun Wang; Shiyi Yao; Renhao Yang; Hongliang Ren; Lei Wang; Junjun He; Yankai Jiang; | code |
| 427 | RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce \textsc{RMNP} (Row Momentum Normalized Preconditioning), an optimizer that replaces Newton-Schulz iteration with a simple row-wise $\ell_2$ normalization operation, motivated by the empirically observed diagonal block structure of the Transformer layerwise Hessian. |
Shenyang Deng; Zhuoli Ouyang; Ruochen Jin; Tianyu Pang; Zihang Liu; Shuhua Yu; Yaoqing Yang; | code |
| 428 | Euclean: Automated Geometry Problem Formalization with Unified Verification in Lean Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, native \mathlib{} autoformalization of geometry poses unique challenges: explicating implicit diagrammatic assumptions (e.g., topological configuration and non-degeneracy)—unlike existing custom systems that defer validity checks to external solvers—and adapting to \mathlib{}’s small, rapidly-evolving geometric constructs. We present \method{}, a framework that addresses these challenges through a four-stage pipeline—constraint explication, configuration anchoring, formalization mapping, and iterative repair—to automatically formalize geometry in native \mathlib{}. |
Linbin Tang; Jingyan You; Zilin Kang; Hanzhang Liu; Sophia Zhang; Zenan Li; Chenrui Cao; Liangcheng Song; Jiaao Wu; Xian Zhang; Fan Yang; | code |
| 429 | On Structured State-Space Duality Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Consequently, the same sequence transformation (model) has two algorithmic realizations: as a linear-time $O(T)$ recurrence or as a quadratic-time $O(T^2)$ attention. In this note, we formalize and generalize this duality: (i) we extend SSD from the scalar‑identity case to general diagonal SSMs (diagonal state matrices); (ii) we show that these diagonal SSMs match the scalar case’s training complexity lower bounds while supporting richer dynamics; (iii) we establish a necessary and sufficient condition under which an SSM is equivalent to $1$-semiseparable masked attention; and (iv) we show that such duality fails to extend to standard softmax attention due to rank explosion. |
Jerry Yao-Chieh Hu; Xiwen Zhang; Ali ElSheikh; Weimin Wu; Han Liu; | code |
| 430 | Harnessing Spectrum Video for Subject-Level Few-Shot and Cross-Montage EEG Generalization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Brain Signal Rendering (BSR), which reinterprets EEG as a physical projection of neural activity and transforms raw signals into geometry-aware Spectrum Videos. |
Wei Wang; Fang He; Yifan Li; Wanying Qu; Yawei Li; Quanying Liu; Yanwei Fu; | code |
| 431 | HiDe: Rethinking The Zoom-IN Method in High Resolution MLLMs Via Hierarchical Decoupling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study the zoom-in operation through a **hierarchical decoupling analysis** and propose the **Hierarchical Decoupling Framework (HiDe)**, a training-free method that turns implicit attention into explicit region selection. |
Xianjie Liu; Yiman Hu; Yixiong Zou; Liang Wu; Jian Xu; Bo Zheng; | code |
| 432 | E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our evaluation reveals that e-commerce content exhibits substantially higher density across visual, audio, and textual modalities compared to mainstream datasets, establishing a more challenging frontier for video understanding. To address this gap, we introduce **E-commerce Video Ads Benchmark (E-VAds)**, which is the first benchmark specifically designed for e-commerce short video understanding. |
Xianjie Liu; Yiman Hu; Liang Wu; Ping Hu; Yixiong Zou; Jian Xu; Bo Zheng; | code |
| 433 | SmartThinker: Progressive Chain-of-Thought Length Calibration for Efficient Large Language Model Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose *SmartThinker*, a novel GRPO-based efficient reasoning method with progressive CoT length calibration. |
Chenzhi Hu; Qinzhe Hu; yuhang xu; Junyi Chen; Ruijie Wang; Shengzhong Liu; Jianxin Li; Fan Wu; Guihai Chen; | code |
| 434 | Position Is All You Need: A Free Lunch Token Compression Strategy for MLLM-based Referring Expression Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To fill this gap, we first evaluate typical token compression methods on this task and observe a surprising performance degradation. In this paper, we aim to understand this phenomenon for a solution. |
Yuhan Liu; Yixiong Zou; Yuhua Li; Ruixuan Li; | code |
| 435 | Conformal Reliability: A New Evaluation Metric for Conditional Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a novel evaluation metric called \emph{reliability score} based on conformal prediction, which measures the worst-case performance within the prediction set at a pre-specified confidence level. |
Yachen Gao; Xinwei Sun; Yikai Wang; Ye Shi; Jingya Wang; Jianfeng Feng; Yanwei Fu; | code |
| 436 | From Blind Spots to Gains: Diagnostic-Driven Iterative Training for Large Multimodal Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Motivated by findings that test driven error exposure and feedback based correction outperform repetitive practice, we propose Diagnostic-driven Progressive Evolution (DPE), a spiral loop where diagnosis steers data generation and reinforcement, and each iteration re-diagnoses the updated model to drive the next round of targeted improvement. |
Hongrui Jia; Chaoya Jiang; Yongrui Heng; Shikun Zhang; Wei Ye; | code |
| 437 | RAT+: Train Dense, Infer Sparse – Recurrence Augmented Attention for Dilated Inference Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce **RAT+**, a dense-pretraining architecture that augments attention with *full-sequence recurrence* and *active recurrence learning*. |
Xiuying Wei; Caglar Gulcehre; | code |
| 438 | Beyond Description: Federated Adaptation Via Semantic-Visual Prototype Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing solutions suffer from high computational complexity or ineffective knowledge aggregation. To address these problems, we propose FedSPA (Federated Adaptation via Semantic-Visual Prototype Alignment). |
Jiarong Yang; Yuan Liu; | code |
| 439 | Concept Concentration for Faithful Representation Intervention Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we explore the question in safety alignment. |
Hongzheng Yang; Yongqiang Chen; Zeyu Qin; Tongliang Liu; Chaowei Xiao; Kun Zhang; Bo Han; | code |
| 440 | Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose SpecFlow, a training-free framework that shifts the paradigm from destructive pruning to conservative condensation, strictly enforcing spatial coverage and statistical conservation to ensure stability. |
Zhaoyang Li; Yanjun Li; Wangkai Li; Yujia Chen; Tianzhu Zhang; | code |
| 441 | Dependency-Aware Parallel Decoding Via Attention for Diffusion LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Dependency-Aware Parallel Decoding (DAPD), a simple, training-free decoding method that uses self-attention to induce a conditional dependency graph over masked tokens. |
Bumjun Kim; Dongjae Jeon; Moongyu Jeon; Albert No; | code |
| 442 | LILO: Bayesian Optimization with Natural Language Feedback Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In response, we introduce Language-in-the-Loop Optimization (LILO), a Bayesian optimization (BO) framework that employs a large language model (LLM) to translate free-form natural language feedback and prior knowledge from a decision maker into structured preference signals, going beyond the restrictive scalar or pairwise feedback formats typically assumed in preferential BO. |
Katarzyna Kobalczyk; Zhiyuan Jerry Lin; Benjamin Letham; Zhuokai Zhao; Maximilian Balandat; Eytan Bakshy; | code |
| 443 | Low-Rank and Sparsity Are All You Need: Exploring Robust Hierarchical Latent Subspaces for Transferable Adversarial Attack Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent SVD-based attack attempts to exploit low-rank feature subspaces, yet its reliance on single-layer optimization and single-gradient pathway neglects both structural redundancy in feature representations and hierarchical heterogeneity across network layers. To address these limitations, we propose LRS-Attack, a Low-Rank and Sparse decomposition based adversarial attack that explicitly models robust hierarchical subspaces in latent feature spaces. |
Shuangshuang Pu; Wen Yang; Min Li; guodong liu; Chris Ding; Di Ming; | code |
| 444 | Dynamics Within Latent Chain-of-Thought: An Empirical Study of Causal Structure Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we view latent chain-of-thought as a manipulable causal process in representation space by modeling latent steps as variables in a structural causal model (SCM) and analyzing their effects through step-wise $\mathrm{do}$-interventions. |
Zirui Li; Xuefeng Bai; Kehai Chen; Yizhi Li; Jian Yang; Chenghua Lin; Min zhang; | code |
| 445 | Unbiased Principles, Robust Rewards Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose IP-GRM (Independent Principle GRM), a two-stage framework that first generates principles solely from the question ($Q \rightarrow P$) and then evaluates the response conditioned on $(Q, R, P)$. |
Qingnan Ren; Zhen Fang; Shiting Huang; Yu Zeng; Lin Chen; Zehui Chen; Feng Zhao; | code |
| 446 | Data Reconstruction: Identifiability and Optimization with Sample Splitting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: On the optimization side, we introduce sample splitting, a curvature-aware refinement step applicable to general reconstruction objectives (not limited to KKT-based formulations): it creates additional descent directions to escape poor stationary points and refine solutions. |
Yujie Shen; Zihan Wang; Jian Qian; Qi Lei; | code |
| 447 | On Information Self-Locking in Reinforcement Learning for Active Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To resolve the issue, we propose a simple yet effective approach that directly promotes AS capability using proxy AS signals to help the agent escape the low-information regime. |
Deyu Zou; Yongqiang Chen; Fan Feng; Mufei Li; Pan Li; Yu Gong; James Cheng; | code |
| 448 | PRISM: Sequence Modeling As Parallel Residual Iteration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing efficient architectures are theoretically bounded by shallow, single-step linear updates, while powerful iterative methods like Test-Time Training (TTT) break hardware parallelism due to state-dependent gradients. We propose PRISM (Parallel Residual Iterative Sequence Model) to resolve this tension. |
Jie Jiang; Ke Cheng; Xin Xu; Mengyang Pang; Tianhao Lu; Jiaheng Li; Yue Liu; Yuan Wang; Jun Zhang; Huan Yu; Zhouchen Lin; | code |
| 449 | FLAG: Foundation Model Representation with Latent Diffusion Alignment Via Graph for Spatial Gene Expression Prediction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While effective for numerical fitting, this approach overlooks biological structures: the functional coordination between genes and their organized distribution across tissue. We reframe this task as structured distribution modeling and introduce \textbf{FLAG}, a diffusion-based framework designed to preserve these biological relationships. |
Qi Si; Penglei Wang; Yushuai Wu; Yifeng Jiao; Xuyang Liu; Xin Guo; Yuan Qi; Yuan Cheng; | code |
| 450 | Sparse ActionGen: Accelerating Diffusion Policy with Real-time Pruning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose $\underline{\textbf{S}}$parse $\underline{\textbf{A}}$ction$\underline{\textbf{G}}$en $(\textbf{SAG})$ for extremely sparse action generation. |
Kangye Ji; Yuan Meng; Jianbo Zhou; Ye Li; Hanyun Cui; Zhi Wang; | code |
| 451 | ReJump: A Tree-Jump Representation for Analyzing and Improving LLM Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, their underlying reasoning algorithms remain poorly understood. To investigate this, we propose *ReJump*, which represents a reasoning trace as a visitation order over nodes in a tree of intermediate problem-solving steps. |
Yuchen Zeng; Shuibai Zhang; Wonjun Kang; Shutong Wu; Lynnix Zou; Ying Fan; Heeju Kim; Hakurei Reimu; Jungtaek Kim; HYUNG IL KOO; Dimitris Papailiopoulos; Kangwook Lee; | code |
| 452 | Graph Is A Substrate Across Data Modalities Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To instantiate this perspective, we propose **G-Substrate**, a **g**raph **substrate** framework that organizes learning around shared graph structures. |
Ziming Li; Xiao-Ming Wu; Zehong Wang; Jiazheng Li; Yijun Tian; Jinhe Bi; Yunpu Ma; Yanfang Ye; Chuxu Zhang; | code |
| 453 | DiffStyle3D: Consistent 3D Gaussian Stylization Via Attention Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by the geometric invariance of 3D stylization, we propose a Geometry-Guided Multi-View Consistency method that integrates geometric information into self-attention to enable cross-view correspondence modeling. |
Yitong Yang; Xuexin Liu; Yinglin Wang; Jing Wang; Hao Dou; Changshuo Wang; Shuting He; | code |
| 454 | Just Noticeable Difference Modeling for Deep Visual Features Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose FeatJND, a task-aligned JND formulation that predicts the maximum tolerable per-feature perturbation map while preserving downstream task performance. |
Rui Zhao; Wenrui Li; Lin Zhu; Yajing Zheng; Weisi Lin; | code |
| 455 | Hair-Trigger Alignment: Black-Box Evaluation Cannot Guarantee Post-Update Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A growing body of alignment research has shown that models initially deemed “aligned” can exhibit misaligned behavior after fine-tuning, such as forgetting jailbreak safety features or re-surfacing knowledge that was intended to be forgotten. These works typically assume that the initial model is aligned based on static black-box evaluation, i.e., the absence of undesired responses to a fixed set of queries. |
Yavuz Faruk Bakman; Duygu Nur Yaldiz; Eleni Triantafillou; Peter Kairouz; Salman Avestimehr; Sai Praneeth Reddy Karimireddy; | code |
| 456 | Thinking in Scales: Accelerating Gigapixel Pathology Image Analysis Via Adaptive Continuous Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, such exhaustive patch-level processing is computationally expensive, severely limiting the efficiency and scalability of WSI analysis. To address this challenge, we propose \textbf{PathCTM} (a \textbf{\ul{Path}}ology-oriented \textbf{\ul{C}}ontinuous \textbf{\ul{T}}hought \textbf{\ul{M}}odel) that enables token-efficient scale-space continuous reasoning for gigapixel WSIs. |
Jiusong Ge; Yingkang Zhan; Wenjie Zhao; Di Zhang; Ke Wang; Jiashuai Liu; Chunze Yang; Chengzu Li; Jian Zhang; Yuxin Dong; Ni Zhang; Qidong Liu; Mireia Crispin-Ortuzar; Huazhu Fu; Chen Li; Zeyu Gao; | code |
| 457 | Align Your Trajectory Tangent: Training Better Consistency Models Via Manifold-Aligned Tangents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, CMs typically require prolonged training with large batch sizes to obtain competitive sample quality. In this paper, we examine the training dynamics of CMs near convergence and discover that CM trajectory tangents — CM output update directions — are quite oscillatory, in the sense that they move parallel to the data manifold, not towards the manifold. |
Beomsu Kim; ByungHee Cha; Jong Chul YE; | code |
| 458 | Conditional Equivalence of DPO and RLHF: Assumptions, Failure Modes, and Provable Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We characterize when this assumption is violated, show the existence of an undesirable solution space, and prove that DPO and RLHF optimize fundamentally different objectives in such cases. To address this, we introduce Constrained Preference Optimization (CPO), augmenting RLHF with constraints for provable alignment. |
Yonggang Zhang; Zhiqin Yang; Wei Xue; Dong Fang; Bo Han; Yike Guo; | code |
| 459 | Towards Pareto-Optimal Tool-Integrated Agents with Pareto Ranking Policy Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing alignment methods predominantly focus on maximizing task accuracy, while overlooking auxiliary objectives such as tool-use efficiency, which are essential for practical deployment. To address this gap, we introduce {ParetoPO}, a two-stage multi-objective optimization framework for aligning tool-using large language models (LLMs) under competing objectives. |
Junyi Li; Xiaowei Qian; Yingyi Zhang; Wenlin Zhang; Guojing Li; Sheng Zhang; Xiao Han; Yichao Wang; Xiangyu Zhao; | code |
| 460 | EMBGUARD: Constructing Hazard-Aware Guardrails for Safe Planning in Embodied Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing approaches lack explicit mechanisms for identifying hazards and reasoning about action-conditioned risks, leading agents to either miss risky interactions or over-identify risks. To address this, we propose EMBGUARD, the first MLLM-based safety guardrail for embodied agents designed to decouple physical risk reasoning from agent policy. |
Dongwook Choi; Taeyoon Kwon; Bogyung Jeong; Minju Kim; Yeonjun Hwang; Hyojun Kim; Byungchul Kim; Young Kyun Jang; Jinyoung Yeo; | code |
| 461 | Target-Oriented Pretraining Data Selection Via Neuron-Activated Graph Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we study target-oriented language model (LM) pretraining by introducing ***N**euron-**A**ctivated **G**raph Ranking* (NAG-based Ranking), a training-free and interpretable framework for target pretraining data selection. |
Zijun Wang; Haoqin Tu; Weidong Zhou; Yiyang Zhou; Xiaohuan Zhou; Bingni Zhang; Weiguo Feng; Taifeng Wang; Cihang Xie; Fengze Liu; | code |
| 462 | ProOPF: Benchmarking and Improving LLMs for Professional-Grade Power Systems Optimization Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Yet existing LLM datasets and benchmarks for optimization modeling primarily target coarse-grained cross-domain generalization, offering limited, rigorous evaluation in powersystem settings, particularly for Optimal Power Flow (OPF). We therefore introduce ProOPFD and ProOPF-B, a dataset and benchmark for professional-grade OPF modeling: ProOPF-D contains 12K instances pairing NL requests with parameter adjustments and structural extensions to a canonical OPF, together with executable implementations; ProOPF-B provides 121 expertannotated test cases with ground-truth code, enabling end-to-end evaluation under both concrete and abstract OPF modeling regimes. |
Chao Shen; Zihan Guo; Xu Wan; Zhenghao Yang; Yifan Zhang; Wenqi Huang; Jie Song; Zongyan Zhang; Mingyang Sun; | code |
| 463 | See First, Reason Later: Mutual Information-Guided Reinforcement Learning for Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce MIRL, a decoupled framework that addresses both limitations by leveraging mutual information (MI) between generated descriptions and visual inputs as a cheap pre-screening signal. |
Junfeng Fang; Zonghan Wu; Yin Zhang; Jiaxuan Zhao; Zengxiang Li; Kun Wang; Qingsong Wen; Yilei Shao; | code |
| 464 | Learn to Merge: Meta-Learning for Adaptive Multi-Task Model Merging Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Thus, this paper proposed an innovative model merging framework called MetaMerging, a novel meta-learning algorithm to adaptively optimize the merging coefficients to construct a unified model tailored for task-specific adapter training. |
Jun Chen; Qin Zhang; Weizhi Zhang; Xiao Luo; Philip Yu; Ziyue Qiao; | code |
| 465 | Can Muon Fine-tune Adam-Pretrained Models? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study this mismatch through controlled experiments and relate it to the distinct implicit biases of Adam and Muon. We provide evidence that fine-tuning with a mismatched optimizer disrupts pretrained knowledge, and show that constraining updates with Low-Rank Adaptation (LoRA) mitigates this issue. |
Xingyu Qu; Peigeng Huang; Samuel Horváth; | code |
| 466 | Learning to Refine: Spectral-Decoupled Iterative Refinement Framework for Precipitation Nowcasting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Spectral-Decoupled Iterative Refinement (SDIR), a deterministic framework that reformulates nowcasting as progressive frequency-decoupled refinement. |
Yunlong Zhou; Chen Zhao; danyang peng; Fanfan Ji; Xiaotong Yuan; | code |
| 467 | Certified Circuits: Stability Guarantees for Mechanistic Circuits Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce *Certified Circuits*, which provide provable stability guarantees for circuit discovery. |
Alaa Anani; Tobias Lorenz; Bernt Schiele; Mario Fritz; Jonas Fischer; | code |
| 468 | The Ideal Expression Is Not A Local Optimum: A Revisit of EQL with Zero-Point Constraints Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We analyze a gradient residual issue induced by operators that do not vanish at zero, which can prevent the ideal sparse expression from being a local optimum and bias training toward unnecessarily complex structures, making exact recovery nearly unattainable in practice. To address this, we propose EQL-Z, a structurally controllable symbolic regression framework. |
Sannyuya Liu; Ao Chen; Lin Liu; Ruxia Liang; Xiaoxuan Shen; Jianwen Sun; | code |
| 469 | KernelBand: Steering LLM-based Kernel Optimization Via Hardware-Aware Multi-Armed Bandits Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To bridge the gap, we present KernelBand, a framework that formulates kernel optimization as a Multi-Armed Bandit (MAB) problem, explicitly balancing exploration and exploitation to unlock the potential of code LLMs. |
Dezhi Ran; Shuxiao Xie; Mingfang Ji; Anmin Liu; Mengzhou Wu; Yuan Cao; Yuzhe Guo; Hao Yu; Linyi Li; Yitao Hu; Wei Yang; Tao Xie; | code |
| 470 | Hierarchical Representations for Cross-task Automated Heuristic Design Using LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing works mainly rely on programs to represent heuristics, which are inherently taskspecific and fail to generalize as effectively as established metaheuristics like tabu search or guided local search. To bridge this gap, we introduce Multi-Task Hierarchical Search (MTHS), an LLM-guided evolutionary method that co-designs general-purpose metaheuristics and task-specific programs. |
Fei Liu; Rui Zhang; Shunyu Yao; Qinglong Hu; kefeng zheng; Zhichao Lu; Qingfu Zhang; | code |
| 471 | LiftQuant: Continuous Bit-Width Control for Pareto-Optimal LLM Deployment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing quantization methods are fundamentally limited by rigid, integer-based bit-widths (e.g., 2, 3-bit), creating a deployment gap where LLMs cannot be optimally fitted to specific memory budgets. To bridge this gap, we introduce LiftQuant, a novel framework that enables continuous bit-width control for true Pareto-optimal deployment. |
Liulu He; Xuan Liu; Juntao Liu; Taolue Feng; Ting Lu; Chunsheng Gan; ZHIYV PENG; Yuan Du; Li Du; Huanrui Yang; Yijiang Liu; | code |
| 472 | Keep It in Mind: User Centric Continual Spatial Intelligence Reasoning in Egocentric Video Streams Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose DirectMe, a framework that incrementally constructs and maintains a structured spatial memory from streaming egocentric observations. |
Yun Wang; Junbin Xiao; Han Lyu; Yifan Wang; Jing Zuo; Zhanjie Zhang; Hong Huang; Dapeng Wu; Angela Yao; | code |
| 473 | AsyncSpade: Efficient Test-Time Scaling with Asynchronous Sparse Decoding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Based on the findings, we propose $\texttt{\textbf{AsyncSpade}}$, an asynchronous framework for efficient TTS, built on two core components: $\textbf{(1) a novel light-weight temporal-regressive module}$ that predicts the next-token query state, and $\textbf{(2) an asynchronous disaggregated framework}$ that decouples the KV cache selection from the auto-regressive decoding loop, overlapping the token-level KV selection with the forward inference computation through asynchronism, thereby eliminating the sequential dependency without sacrificing model performance. |
Shuqing Luo; Yilin Guan; Pingzhi Li; Hanrui Wang; Tianlong Chen; | code |
| 474 | Implicit Actor Critic Coupling Via A Supervised Learning Framework for RLVR Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address the challenges, we propose **PACS**, a novel RLVR framework that achieves im**P**licit **A**ctor **C**ritic coupling via a **S**upervised learning framework. |
Jiaming Li; Longze Chen; Ze Gong; Yukun Chen; Lu Wang; Wanwei He; Run Luo; Minzheng Wang; Lei Zhang; Haoran Ye; Min Yang; | code |
| 475 | Breaking The Block: Preserving Data Continuity to Train Superior SAEs for Instruct Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Theoretically, utilizing GSNR analysis, we prove that attention leakage from unrelated contexts introduces destructive gradient noise. To rectify this paradigm, we propose $\underline{\textbf{F}}$inetuning-$\underline{\textbf{a}}$ligned $\underline{\textbf{S}}$equential $\underline{\textbf{T}}$raining ($\textit{FAST}$), a novel training method specifically tailored for instruct models. |
Jiaming Li; Haoran Ye; Yukun Chen; Xinyue Li; Lei Zhang; Hamid Alinejad-Rokny; Jimmy Peng; Min Yang; | code |
| 476 | VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing multimodal red-teaming methods largely rely on brittle templates, focus on single-attack settings, and expose only a narrow subset of vulnerabilities. To address these limitations, we introduce VERA-V, a variational inference framework that recasts multimodal jailbreak discovery as learning a joint posterior distribution over paired text-image prompts. |
Qilin Liao; Anamika Lochab; Ruqi Zhang; | code |
| 477 | PlugMem: A Task-Agnostic Plugin Memory Module for LLM Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose PlugMem, a task-agnostic plugin memory module that can be attached to arbitrary LLM agents without task-specific redesign. |
Ke Yang; Zixi Chen; Xuan He; Jize Jiang; Michel Galley; Chenglong Wang; Jianfeng Gao; Jiawei Han; Chengxiang Zhai; | code |
| 478 | CoPE: A Framework for Optimizing Coordination Between Planning and Execution in LLM-based Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This oversight leads to the propagation of impractical plans and plan-deviated trajectories into the optimization process, resulting in suboptimal task performance and hindering the further development of LLM-based agents in long-horizon tasks. To bridge this gap, we propose $\textbf{CoPE}$, a novel framework that explicitly integrates planning–execution coordination into LLM-based agent optimization. |
Huanxi Liu; Kun Hu; Qiang Wang; Yuanzhao Zhai; Feng Dawei; Bo Ding; Huaimin Wang; | code |
| 479 | Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Dynamic Hierarchical Sparse Attention (DHSA), a data-driven framework that predicts attention sparsity online while keeping the LLM backbone frozen. |
Siheng Xiong; Joe Zou; Faramarz Fekri; Yae Jee Cho; | code |
| 480 | Shape of Thought: Progressive Object Assembly Via Visual Chain-of-Thought Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Multimodal models for text-to-image generation have achieved strong visual fidelity, yet they remain brittle under compositional structural constraints—notably generative numeracy, attribute binding, and part-level relations. To address these challenges, we propose **Shape-of-Thought (SoT)**, a visual CoT framework that enables *progressive shape assembly represented as coherent 2D projections* without external engines at inference time. |
Yu Huo; Siyu Zhang; Zeng Kun; Haoyue Liu; Owen Lee; Junlin chen; Lu YuQuan; Yifu Guo; Yaodong Liang; Xiaoying Tang; | code |
| 481 | Reinforcing Real-world Service Agents: Balancing Utility and Cost in Task-oriented Dialogue Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, effectively balancing empathetic communication with budget-aware decision-making remains an open challenge. Since existing methods fail to capture these complex strategic trade-offs, we propose InteractCS-RL, a framework that reframes task-oriented dialogue as a multi-granularity reinforcement learning process. |
Ning Gao; Wei Zhang; Yuqin Dai; Ling Shi; Ziyin Wang; Yujie Wang; Wei He; Jinpeng Wang; Chaozheng Wang; | code |
| 482 | Exploring Data-Free LoRA Transferability for Video Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, our analysis reveals that while both paradigms respect spectral rigidity, they establish conflicting routing pathways that clash through constructive overload or destructive cancellation. To address this issue, we propose Cluster-Aware Spectral Arbitration (CASA), a data-free framework that dynamically arbitrates between safeguarding the target’s manifold and restoring LoRA alignment based on spectral density. |
Yuchen Wang; Wenliang Zhong; Lichen Bai; zikai zhou; Shitong Shao; Bojun Cheng; Shuo Chen; Shuo Yang; Zeke Xie; | code |
| 483 | Graph of States: Solving Abductive Tasks with Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Consequently, they are inevitably prone to Evidence Fabrication, Context Drift, Failed Backtracking, and Early Stopping. To bridge this gap, we introduce Graph of States (GoS), a general-purpose neuro-symbolic framework tailored for abductive tasks. |
Yu Luo; Rongchen Gao; Lu Teng; Xidao Wen; Jiamin Jiang; Qingliang Zhang; Yongqian Sun; Shenglin Zhang; Jiasong Feng; Tong Liu; Wenjie Zhang; Dan Pei; | code |
| 484 | Revealing Long-context Potential of Attention Heads Via Frequency Kernels Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we use kernel methods to analyze static *frequency kernels* formed by different rotation frequency components of attention heads, and we design a Long-context Potential Score (LPS) to measure the potential of attention heads in processing long contexts. |
Senyu Han; Yilu Cao; Kai Yu; Lu Chen; | code |
| 485 | Relational In-Context Learning Via Synthetic Pre-training with Structural Prior Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Consequently, existing solutions typically rely on limited real-world datasets, requiring costly fine-tuning to achieve viable performance. To overcome this data scarcity, we introduce RDB-PFN, the first foundation model for databases trained purely on synthetic data. |
Yanbo Wang; Jiaxuan You; Chuan Shi; Muhan Zhang; | code |
| 486 | When Shared Knowledge Hurts: Spectral Over-Accumulation in Model Merging Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We show that when tasks share aligned spectral directions (\ie, overlapping singular vectors), a simple linear combination repeatedly accumulates these directions, inflating the singular values and biasing the merged model toward shared subspaces. To mitigate this issue, we propose Singular Value Calibration (SVC), a training-free and data-free post-processing method that quantifies subspace overlap and rescales inflated singular values to restore a balanced spectrum. |
Yayuan Li; Ze Peng; Jian Zhang; Jintao Guo; Yue Duan; Yinghuan Shi; | code |
| 487 | IMPACT: Influence Modeling for Open-Set Time Series Anomaly Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: They are further plagued when the training data is contaminated with unlabeled anomalies. This work introduces $\textbf{IMPACT}$, a novel framework that leverages $\underline{\textbf{i}}$nfluence $\underline{\textbf{m}}$odeling for o$\underline{\textbf{p}}$en-set time series $\underline{\textbf{a}}$nomaly dete$\underline{\textbf{ct}}$ion, to tackle these challenges. |
Xiaohui Zhou; Yijie Wang; Hongzuo Xu; Weixuan Liang; Xiaoli Li; Guansong Pang; | code |
| 488 | Towards Scalable and Consistent 3D Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: On the data side, we introduce 3DEditVerse, the largest paired 3D editing benchmark to date, comprising 116,309 high-quality training pairs and 1,500 curated test pairs. |
Ruihao Xia; Yang Tang; Pan Zhou; | code |
| 489 | Semantic-Aware Motion Encoding for Topology-Agnostic Character Animation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Generalizing motion representation across diverse characters remains challenging due to significant topological variations in skeletal structures across datasets and species, which hinders the development of scalable generative models. To bridge this gap, we propose a Semantic-Aware Topology-Agnostic framework that learns a unified latent manifold shared by disparate species. |
Zongye Zhang; Yuzhuo Cui; Qingjie Liu; Yunhong Wang; | code |
| 490 | Spatiotemporal Imputation with Graph-Informed Flow Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent diffusion-based methods mitigate error propagation but require iterative sampling and often depend on problem-agnostic Gaussian priors, limiting both efficiency and effectiveness. To address these limitations, we propose GiFlow, a Graph-Informed Flow Matching framework for spatiotemporal imputation. |
Zepeng Zhang; Aref Einizade; Jhony Giraldo; Olga Fink; | code |
| 491 | Model Fusion Via Retrofitting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a novel, neuron-centric family of model fusion algorithms designed to integrate multiple trained neural networks into a single network effectively regardless of training data distribution. |
Phoomraphee Luenam; Andreas Spanopoulos; Amit Sant; Sotiris Anagnostidis; Thomas Hofmann; Sidak Pal Singh; | code |
| 492 | Enhancing Membership Inference Attacks on Diffusion Models from A Frequency-Domain Perspective Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Moreover, we theoretically demonstrate that this deficiency reduces the membership advantage of attacks, thereby interfering with the effective discrimination of member data and hold-out data. Based on this insight, we propose a plug-and-play high-frequency filter module to mitigate the adverse effects of the deficiency, which can be seamlessly integrated into any attacks within the general paradigm without additional time costs. |
Puwei Lian; Yujun Cai; Songze Li; Bingkun BAO; | code |
| 493 | Conflict-Aware Additive Guidance for Flow Models Under Compositional Rewards Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing methods often fail when composing multiple constraints simultaneously, which leads to deviations from the true data manifold. In this work, we identify root causes of this off-manifold drift and find that local errors scale severely with multiple guidance misalignment. |
Xuehui Yu; Fucheng Cai; Meiyi Wang; Xiaopeng Fan; Harold Soh; | code |
| 494 | Improving Sampling for Masked Diffusion Models Via Information Gain Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While existing samplers typically greedily select positions with the lowest uncertainty, we identify their fundamental limitations through failure case analysis, showing they overlook the impact of current actions on subsequent steps and fail to optimize cumulative uncertainty. To bridge this gap, we propose the **Info-Gain Sampler**, a principled decoding framework that balances immediate costs with information gain. |
Kaisen Yang; Jayden Teoh; Kaicheng Yang; Yitong Zhang; Alex Lamb; | code |
| 495 | RGGT: A Generative-Prior-Guided Transformer for Unified Rigid and Non-Rigid Point Cloud Registration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose RGGT, a Generative-Prior-Guided Transformer that unifies rigid and non-rigid registration within a shared optimization space. |
Chengyu Zheng; Songlin Yang; Jin Huang; Honghua Chen; Weiming Wang; Haoran Xie; Fu Lee Wang; Mingqiang Wei; | code |
| 496 | Training-Free Distribution Adaptation for Diffusion Models Via Maximum Mean Discrepancy Guidance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose MMD Guidance, a training-free mechanism that augments the reverse diffusion process with gradients of the Maximum Mean Discrepancy (MMD) between generated samples and a reference dataset. |
Matina Mahdizadeh Sani; Nima Jamali; Mohammad Jalali; Farzan Farnia; | code |
| 497 | VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, this AR paradigm inevitably faces a dual efficiency bottleneck: strictly unidirectional attention compromises *understanding efficiency* by hindering global spatiotemporal aggregation, while serial decoding restricts *generation efficiency*. To address this, we propose **VidLaDA**, a Video LLM based on Diffusion Language Models (DLMs) that leverages bidirectional attention to unlock comprehensive spatiotemporal modeling and decode tokens in parallel. |
Zhihao He; Tieyuan Chen; Kangyu Wang; Ziran Qin; Yang Shao; Chaofan Gan; Shijie Li; Zuxuan Wu; Weiyao Lin; | code |
| 498 | SI-IGCL: Subject Invariance-aware Inverse Graph Contrastive Learning for Psychiatric Disorder Identification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, current methods struggle with subject variations, impairing the model’s generalization ability to the test set. To address this issue, we propose the Subject Invariance-aware Inverse Graph Contrastive Learning (SI-IGCL) model, which adopts a two-stage paradigm with self-supervised subject-invariant pre-training followed by supervised fine-tuning for identification. |
Jiayu Lu; Yujin Wang; Xiaofeng Liu; Dandan Li; Bin Wang; | code |
| 499 | Navigating The Pareto Frontier of Alignment:Spectrum-Adaptive Fine-Tuning for LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: On the other hand, alternatives like Dynamic Fine-Tuning (DFT) suffer from vanishing gradients on these tokens, which severely hinders the acquisition of new concepts. To bridge this gap, we propose **SAFT** **S**pectrum-**A**daptive **F**ine-**T**uning), a unified framework that interpolates between the aggressive learning signal of NLL and the robust nature of probability-weighted optimization. |
Yaoyou Fan; Chao Zhang; Xiaoyu Tan; Chenxing Sun; Yu Yuan; Haoyu Feng; Lu Pan; Ke Zeng; Xunliang Cai; | code |
| 500 | C$^{2}$R: Cross-sample Consistency Regularization Mitigates Feature Splitting and Absorption in Sparse Autoencoders Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These issues stem from inconsistent latent assignment across samples: without cross-sample constraints, per-sample optimization often allows a single underlying concept to be inconsistently distributed across multiple redundant or interfering latents. To address this, we introduce C$^2$R (\underline{\textbf{C}}ross-sample \underline{\textbf{C}}onsistency \underline{\textbf{R}}egularization). |
Haoran Jin; Xiting Wang; Shijie Ren; Hong Xie; Defu Lian; | code |
| 501 | Unsupervised Neural Langevin Sampler for Mixed Integer Linear Programming Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Research into mixed integer linear programs (MILPs) is particularly limited due to the lack of effective heuristics for feasibility and the challenge of modeling mixed-type variables for neural solvers. To address these issues, we propose a novel unsupervised Langevin sampler for solving MILPs. |
Yixin Huang; Shengyu Feng; Yiming Yang; | code |
| 502 | Revisiting Uncertainty: On Evidential Learning for Partially Relevant Video Retrieval Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this setting, vague queries often induce semantic ambiguity across videos, a challenge that is further exacerbated by the sparse temporal supervision within videos, which fails to provide sufficient matching evidence. To address this, we propose Holmes, a hierarchical evidential learning framework that aggregates multi-granular cross-modal evidence to quantify and model uncertainty explicitly. |
Jun Li; Peifeng Lai; Xuhang Lou; Jinpeng Wang; Yuting Wang; Ke Chen; Yaowei Wang; Shutao Xia; | code |
| 503 | OPT-Engine: Benchmarking The Limits of LLMs in Optimization Modeling Via Complexity Scaling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we introduce OPT-ENGINE, an extensible benchmark framework with quantifiable and controllable complexity. |
Yitian Chen; Dongdong Ge; Cheng Cheng; Yinan Sun; Zi Ling; | code |
| 504 | Persistent Backdoor Attacks in Class-Incremental Learning Via Structural Invariant Anchoring Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we reveal that the most advanced attack relies on an implicit assumption that task-critical neurons remain stable across task learning; however, it does not hold in class-incremental learning (CIL). |
Junhuang Huang; Linshan Hou; Jianting Ning; Yanjun Zhang; Zhongyun Hua; Leo Yu Zhang; | code |
| 505 | Can VLMs Diagnose and Recover from VLA Manipulation Faults? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While VLMs offer strong explanatory capabilities, their effectiveness in assisting VLAs is limited by their unclear role in diagnostics and inadequate collaboration mechanisms. To address this, we introduce VLA-FixBench, a fault evaluation dataset that spans perception, planning, and control failures, and provides annotations for task stages, fault types, and spatiotemporal repair strategies. |
Bowen Yan; Jiahao Xiao; Kehui Liu; Jianbo Zhang; Zicheng Zhang; Qi Jia; Zhongjie Jia; Haoming Song; Chunyi Li; Bin Zhao; Guangtao Zhai; | code |
| 506 | KBQA-R1: Reinforcing Large Language Models for Knowledge Base Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While Large Language Models (LLMs) have advanced this field, current approaches often struggle with a dichotomy of failure: they either generate hallucinated queries without verifying schema existence or exhibit rigid, template-based reasoning that mimics synthesized traces without true comprehension of the environment. To address these limitations, we present **KBQA-R1**, a framework that shifts the paradigm from text imitation to interaction optimization via Reinforcement Learning. |
Xin Sun; Zhongqi Chen; Xing Zheng; Qiang Liu; Shu Wu; Bowen Song; Zilei Wang; Weiqiang Wang; Liang Wang; | code |
| 507 | CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI Reviewers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Analysis shows that AI reviewers remain limited in correctness and thoroughness and are prone to hallucinations, and highlights reasoning models as more effective reviewers, motivating further directions for improving AI reviewers. |
Hexuan Deng; Xiaopeng Ke; Yichen Li; Ruina Hu; Dehao Huang; Derek Wong; Yue Wang; Xuebo Liu; Min zhang; | code |
| 508 | Support-Proximity Augmented Diffusion Estimation for Offline Black-Box Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose \textbf{SPADE} (\textbf{S}upport-\textbf{P}roximity \textbf{A}ugmented \textbf{D}iffusion \textbf{E}stimation), a novel framework that reimagines forward surrogate modeling through the lens of conditional generative modeling. |
Yonghan Yang; Ye Yuan; Zipeng Sun; Linfeng Du; Bowei He; Haolun Wu; Can Chen; Xue Liu; | code |
| 509 | BrainJanus: A Foundation Model for Unified Understanding and Generation Across Brain, Vision, and Language Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing approaches predominantly treat brain encoding and decoding as isolated tasks, relying heavily on unimodal alignment and external priors while overlooking the brain’s intrinsic nature as a multimodal integration system. To address these limitations, we propose BrainJanus, the first unified brain foundation model that integrates brain, vision, and language within a single framework. |
Haitao Wu; Qirui Zhang; Zhouheng Yao; Shangquan Sun; Qihao Zheng; Mianxin Liu; Chi Zhang; Wanli Ouyang; Chunfeng Song; Changqing Zhang; Jiamin Wu; | code |
| 510 | Enhancing Reasoning for Diffusion LLMs Via Distribution Matching Policy Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We identify a key challenge in the implementation with a small training batch size and propose several effective solutions through a novel weight baseline subtraction technique. |
Yuchen Zhu; Wei Guo; Jaemoo Choi; Petr Molodyk; Bo Yuan; Molei Tao; Yongxin Chen; | code |
| 511 | Graph-GRPO: Training Graph Flow Models with Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Graph-GRPO, an online reinforcement learning (RL) framework for training GFMs under verifiable rewards. |
Baoheng Zhu; Deyu Bo; Delvin Zhang; Xiao Wang; | code |
| 512 | AICrypto: Evaluating Cryptography Capabilities of Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We hope this work could provide insights for future research on LLMs in cryptographic applications. |
Yu Wang; Yijian Liu; Liheng Ji; Han Luo; Wenjie Li; Xiaofei Zhou; Chiyun Feng; Puji Wang; Yuhan Cao; Geyuan Zhang; Xiaojian Li; Rongwu Xu; Yilei Chen; Tianxing He; | code |
| 513 | MM-Snowball: Evaluating and Mitigating Hallucination Snowballing in Multimodal Multi-turn Dialogue Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Extensive evaluation shows that our benchmark poses a significant challenge even to advanced MLLMs and reveals the inefficacy of existing mitigation methods designed for single-turn tasks. To counteract this degradation, we propose Conflict-Aware Visual Rectification (CAVR). |
Yue Jiang; Xue JIANG; Lihua Zhang; Zhiqiang Wang; Yuhang Lu; Peng Wang; Bo Han; Feng Zheng; Dingkang Yang; | code |
| 514 | Deep Multi-view Graph Clustering Via Attribute-aware Bidirectional Structural Refinement and Pseudo-label Guided Multi-level Fusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, this paper proposes **A**ttribute-aware Bidirectional Structural Refinement (ABSR) and **P**seudo-label Guided Multi-level Fusion (PGMF) for DM**GC**, termed **APGC**. |
Youqing Wang; Tianxiang Zhao; Mengyuan Xin; Ye Su; Jiapu Wang; Tengfei Liu; Junbin Gao; Jipeng Guo; | code |
| 515 | MusicDET: Zero-Shot AI-Generated Music Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this issue, we formulate a zero-shot setting for AI-generated music detection, where the detector is trained exclusively on real music without access to any generated samples. Under this setting, we propose MusicDET, a generator-agnostic detection framework based on frequency-guided normalizing flows that probabilistically models the distribution of real music features. |
Chaolei Han; Hongsong Wang; Jie Gui; | code |
| 516 | Decoupling Skeleton and Flesh: Efficient Multimodal Table Reasoning with Disentangled Alignment and Structure-aware Guidance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work addresses a key question: *how to adapt LVLMs to table reasoning with minimal annotation and no external tools? |
Yingjie Zhu; Xuefeng Bai; Kehai Chen; Yang Xiang; Youcheng Pan; Xiaoqiang Zhou; Min zhang; | code |
| 517 | Evaluating and Rewarding LALMs for Expressive Role-Play TTS Via Mean Continuation Log-Probability Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A critical bottleneck is the lack of objective metrics for quantifying speaking style. To bridge this gap, we propose **Mean Continuation Log-Probability (MCLP)** as both an evaluation metric and a reward signal, validated on LALM-based Role-Play TTS (RP-TTS) tasks. |
Yong Ren; Jingbei Li; Haiyang Sun; Yujie Chen; Cheng Yi; Yechang Huang; Hao Gu; Ye Bai; Xuerui Yang; | code |
| 518 | RobuQ: Pushing DiTs to W1.58A2 Via Robust Activation Quantization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Identifying activation quantization as the primary bottleneck for low-bit settings, we propose **RobuQ**, a systematic QAT framework. |
Kaicheng Yang; Xun Zhang; Haotong Qin; Yucheng Lin; Kaisen Yang; Xianglong Yan; Yulun Zhang; | code |
| 519 | SpikingLM: Towards Fully Spiking Language Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These limitations create a substantial performance gap that has hindered the practical deployment of spiking language models. To address these challenges, we introduce SpikingLM, a framework that bridges the efficiency of SNNs with the capabilities of modern language models through two key innovations. |
Yu Liang; Zijian Zhou; Wenjie Wei; Shuai Wang; Honglin Cao; Ammar Belatreche; Yu Yang; Malu Zhang; Yang Yang; Haizhou Li; | code |
| 520 | Evidential Reasoning Advances Interpretable Real-World Disease Screening Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, current screening models for medical images often suffer from limited interpretability and suboptimal performance, as they lack mechanisms to reference historical cases or provide transparent reasoning pathways. To address these challenges, we introduce EviScreen, an evidential reasoning framework for disease screening that leverages region-level evidence retrieved from knowledge banks of historical cases. |
Chenyu Lian; Hong-Yu Zhou; Jing Qin; | code |
| 521 | ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Representation alignment (REPA) has been investigated to accelerate diffusion training, but we observe that regularizing intermediate representations in diffusion Transformers (DiT) may implicitly entangle latents and limit generative capacity. To address this issue, we propose ReGen, a hierarchical multi-prompt representation generation framework that jointly estimates multiple vector fields for both representations and data within a single diffusion model. |
Sang-Hoon Lee; Ha-Yeong Choi; | code |
| 522 | Bridging The Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Scaling experiments reveal that reliance on synthetic data triggers a Stability-Expressivity Gap, characterized by a non-monotonic degradation we term Synthetic Erosion. To bridge this gap, we propose two self-alignment frameworks. |
Yizhong Geng; Yanliang Li; Jinghan Yang; Tianhan Jiang; Boxun An; Ya Li; Anhao Zhao; | code |
| 523 | Robust Inter-Series Dependency Modeling for Time Series Forecasting Via Information-Theoretic Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Through comprehensive analysis, we identify a critical structural inconsistency in Variate Transformers (exemplified by iTransformer): typically capturing inter-variate dependencies via shallow self-attention layers while neglecting the critical requirement for deep-layer IVD modeling, which causes dependency information loss and difficulties in model optimization. To address these limitations, we propose CGTFra, as a general Graph Transformer framework. |
Wuqing Yu; Weichen Guo; Jian Zhou; Shuyu Luo; Jiacai Zhang; | code |
| 524 | Pianist Transformer: Towards Expressive Piano Performance Rendering Via Scalable Self-Supervised Pre-Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods for expressive music performance rendering rely on supervised learning over small labeled datasets, which limits scaling of both data volume and model size, despite the availability of vast unlabeled music, as in vision and language. To address this gap, we introduce Pianist Transformer, with three key contributions: 1) introducing large-scale self-supervised learning into expressive piano performance rendering through a unified Musical Instrument Digital Interface (MIDI) representation, enabling pre-training on 10B tokens of unlabeled MIDI data; 2) an efficient asymmetric Transformer with note-level compression, substantially improving training efficiency, memory usage, and inference speed for long-context music modeling; 3) a state-of-the-art rendering model with an editable workflow, achieving strong objective and subjective results and enabling integration into real-world music production workflows. |
Hong-Jie You; Jie-Jing Shao; Xiao-Wen Yang; Lin-Han Jia; Lan-Zhe Guo; Yu-Feng Li; | code |
| 525 | Dual-View Predictive Diffusion: Lightweight Speech Enhancement Via Spectrogram-Image Synergy Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, most existing score-based models treat speech spectrograms merely as generic 2D images, applying uniform processing that ignores the intrinsic structural sparsity of audio, which results in inefficient spectral representation and prohibitive computational complexity. To bridge this gap, we propose **DVPD**, an extremely lightweight **D**ual-**V**iew **P**redictive **D**iffusion model, which uniquely exploits the dual nature of spectrograms as both visual textures and physical frequency-domain representations across both training and inference stages. |
Ke Xue; Rongfei Fan; Kai Li; Shanping Yu; Puning Zhao; Jianping An; | code |
| 526 | POLCA: Stochastic Generative Optimization with LLM Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We formalize this challenge as a stochastic generative optimization problem where a generative language model acts as the optimizer, guided by numerical rewards and text feedback. We introduce Prioritized Optimization with Local Contextual Aggregation (POLCA), a scalable framework designed to handle stochasticity in optimization—such as noisy feedback, sampling minibatches, and stochastic system behaviors—while effectively managing the unconstrained expansion of solution space. |
Xuanfei Ren; Allen Nie; Tengyang Xie; Ching-An Cheng; | code |
| 527 | On The Generalization in Topology Optimization Via Sensitivity-Conditioned Bernoulli Flow Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, computing exact adjoint sensitivities can be expensive or unavailable in practice; we observe that certain physical fields can approximate sensitivities through monotone transformations. To formalize this, we introduce \textbf{pseudo-sensitivities} to characterize which fields enable generalization versus those that are information-poor. |
Mohammad Rashed; Duarte Madeira; Babak Gholami; Caglar Guerbuez; Yunjia Yang; Nils Thuerey; | code |
| 528 | Hyperbolic Hierarchical Alignment for Video-Based Visible-Infrared Person Re-Identification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Hyperbolic Hierarchical Alignment (HHA), which unifies spatio-temporal modeling and cross-modality alignment on the Poincar\’e ball. |
Shuang Li; Changjiang Kuang; Jiaxu Leng; Mingpi Tan; Zhanjie Wu; Shuanglin Yan; Xinbo Gao; | code |
| 529 | Bridging The Gap in Autonomous Science: The Corpus and Benchmark for Biological Protocol Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The realization of autonomous scientific experimentation is currently limited by LLMs’ struggle to grasp the strict procedural logic and accuracy required by biological protocols. To address this fundamental challenge, we present **BioProBench**, a comprehensive resource for procedural reasoning in biology. |
Yuyang Liu; Liuzhenghao Lv; Xiancheng Zhang; Jingya Wang; Li Yuan; Yonghong Tian; | code |
| 530 | Unlocking Zero-Shot Geospatial Reasoning Via Indirect Rewards Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we validate an important conclusion: indirect verifiable rewards, derived from seemingly unrelated metadata, are sufficient to induce sophisticated and generalizable geospatial reasoning across a wide range of downstream tasks (25+). |
Chenhui Xu; Fuxun Yu; Michael Bianco; Jacob Kovarskiy; Raphael Tang; Qi Zhang; Zirui Xu; Will LeVine; Brandon Dubbs; Heming Liao; Cassandra Burgess; Suvam Bag; Jay Patravali; Rupanjali Kukal; Mikael Figueroa; Rishi Madhok; Nikolaos Karianakis; Jinjun Xiong; | code |
| 531 | Evolutionary Multi-View Classification with Label Noise Via Gradient and Feature Dual-Perception Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Given that label noise largely stems from the mislabeling of samples near their decision boundaries by human annotators, we thus compared the decision boundaries of human annotators and models, and found discrepancies between the two. Based on this observation, we propose a simple yet effective “detect-then-calibrate data purification framework that leverages outlier analysis in the gradient space (i.e., treating outliers as noisy samples) and prototype calibration in the feature space (i.e., utilizing feature prototypes of noise-free samples to correct the labels of noisy samples). |
Shuai Li; Xinyan Liang; Yuhua Qian; Li Lv; | code |
| 532 | Rethinking Code Complexity Through The Lens of Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we empirically demonstrate that, after controlling for code length, classical metrics exhibit no consistent correlation with LLM performance, revealing a fundamental mismatch with model-perceived difficulty. |
Chen Xie; Yuling Shi; Xiaodong Gu; Beijun Shen; | code |
| 533 | R2-Router: A New Paradigm for LLM Routing with Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This causes routers to exclude powerful LLMs when their estimated cost exceeds the budget, missing the opportunity that these LLMs could still deliver high quality at reduced cost with shorter outputs. To address this, we introduce R2-Router, which treats output length budget as a controllable variable and jointly selects the best LLM and length budget, enforcing the budget via length-constrained instructions. |
Jiaqi Xue; Qian Lou; Jiarong Xing; Heng Huang; | code |
| 534 | Dynamic Decision Learning: Test-Time Evolution for Abnormality Grounding in Rare Diseases Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Thus, we propose Dynamic Decision Learning (DDL), a framework that enables frozen LVLMs to refine their decisions across language and visual spaces by optimizing instructions and consolidating predictions under visual perturbations, thereby improving localization quality and producing a consensus‑based reliability score that quantifies the model’s confidence. |
Jun Li; Mingxuan Liu; Jiazhen Pan; che liu; Wenjia Bai; Cosmin Bercea; Julia Schnabel; | code |
| 535 | Revealing Differences in Multi-Modal Embeddings Via Constrained Kernel Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we aim to identify modality pairs and sample subsets that induce divergent grouping behavior between two candidate embeddings. |
Youqi WU; Mohammad Jalali; Farzan Farnia; | code |
| 536 | BuildArena: A Physics‑Aligned Interactive Benchmark of LLMs for Engineering Construction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While modern LLMs possess broad knowledge and strong reasoning capabilities that make them promising candidates for this domain, their construction competencies remain largely unevaluated. To address this gap, we introduce BuildArena, the first physics-aligned interactive benchmark designed for language-driven engineering construction. |
Tian Xia; Tianrun Gao; Wenhao Deng; Long Wei; Xiaowei Qian; Chenglei Yu; Tailin Wu; | code |
| 537 | Rethinking LLM Ensembling from The Perspective of Mixture Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose the Mixture-model-like Ensemble (ME). |
Jiale Fu; Yuchu Jiang; PeiJun Wu; Chonghan Liu; Joey Tianyi Zhou; Xu Yang; | code |
| 538 | AgentExpt: Automating AI Experiment Design with LLM-based Resource Retrieval Agent Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, making appropriate choices is increasingly difficult as baselines and datasets proliferate, while suitability is inherently context-dependent and rarely captured by baseline and dataset metadata. To address these challenges, we present \textbf{AgentExpt}, a comprehensive framework for baseline and dataset recommendation. |
Yu Li; Lehui Li; Lin Chen; Qingmin Liao; Fengli Xu; Yong Li; | code |
| 539 | You Don’t Protect If You Don’t Expect: Breaking The Key Assumption Behind CLIP’s Test-Time Defenses Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We further observe that these defenses share a common reliance on an indicative measurement that is assumed to capture the distributional difference between clean and adversarial samples and to determine whether the defense should preserve or alter the static model’s prediction. We argue that this assumption is the fundamental weakness, and we propose CLIP-MAD (Manipulating Assumed Difference), an adaptive attack strategy designed to break it. |
Ruize Zhang; Yu Li; Zhang Wan; Juan Cao; Jie Zhang; Sheng Tang; | code |
| 540 | From Interaction Trajectories to Prompt Rules: Credit Assignment for Multi-Agent Prompt Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As a result, outcome-level feedback alone is insufficient, while existing prompt optimization methods typically rely on final task scores or global prompt rewrites, limiting their ability to exploit trajectory evidence or support the localized updates. We propose Trajectory-based Rule Credit Estimation (TRUCE), a framework for prompt optimization in multi-agent systems that explicitly addresses this credit assignment challenge. |
Bin Wu; Haoran Xu; Xiang Zhuang; Zonghao Chen; Zhu Li; Emine Yilmaz; Qiang Zhang; | code |
| 541 | AgentLAB: Benchmarking LLM Agents Against Long-Horizon Attacks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: LLM agents are increasingly deployed in long-horizon, complex environments to solve challenging problems, but this expansion exposes them to long-horizon attacks that exploit multi-turn user–agent–environment interactions to achieve objectives infeasible in single-turn settings. To measure agent vulnerabilities to such risks, we present AgentLAB, the first benchmark dedicated to evaluating LLM agent susceptibility to adaptive, long-horizon attacks. |
Tanqiu Jiang; Yuhui Wang; Jiacheng Liang; Ting Wang; | code |
| 542 | MetaphorVU: Towards Metaphorical Video Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Through experiments, we find current MLLMs struggle with accurate metaphorical video understanding, lagging far behind human level, primarily due to defective cross-domain mapping. Motivated by this finding, we construct a metaphor knowledge graph as mapping augmentation and propose MetaphorBoost, an inference-time enhancement framework achieving consistent performance improvement. |
Zhuoqun Li; Boxi Cao; Guiping Jiang; Fangrui Lv; Ruotong Pan; Jianan Wang; Xiangyu Wu; Hongyu Lin; Yaojie Lu; Yong Du; Ruyin Jia; Liyan; Tingting Gao; Han Li; Xianpei Han; Le Sun; | code |
| 543 | 3DGS-HPC: Distractor-free 3D Gaussian Splatting with Hybrid Patch-wise Classification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods commonly rely on semantic cues extracted from pre-trained vision models to identify and suppress these distractors, but such semantics are misaligned with the binary distinction between static and transient regions and remain fragile under the appearance perturbations introduced during 3DGS optimization. We propose 3DGS-HPC, a framework that circumvents these limitations by combining two complementary principles: a patch-wise classification strategy that leverages local spatial consistency for robust region-level decisions, and a hybrid classification metric that adaptively integrates photometric and perceptual cues for more reliable separation. |
Jiahao Chen; Yipeng Qin; Ganlong Zhao; Xin Li; Wenping Wang; Guanbin Li; | code |
| 544 | Towards Disentangled Preference Optimization Dynamics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Leveraging the DB, we propose a plug-and-play \textbf{reward calibration (RC)} that adaptively rebalances chosen versus rejected updates to satisfy the DB, without redesigning the base objective. |
Wei Chen; Yubing Wu; Junmei Yang; Delu Zeng; Qibin Zhao; John Paisley; Min Chen; Zhou Wang; | code |
| 545 | Little By Little: Continual Learning Via Incremental Mixture of Rank-1 Associative Memory Experts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose MoRAM (Mixture of Rank-1 Associative Memory). |
Haodong Lu; Chongyang Zhao; Minhui Xue; Lina Yao; Kristen Moore; Dong Gong; | code |
| 546 | Order Within Chaos: Capturing Intrinsic Energy Anomalies for AI-Manipulated Image Forgery Localization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this challenge, we theoretically demonstrate that the diffusion process inherently suppresses local high-frequency variance, creating a statistical energy gap that is distinguishable from the natural entropy of optical imaging. Guided by this insight, we propose FLAME, a unified framework that utilizes a LAD map to capture these intrinsic anomalies, coupled with a parameter-efficient adapter for the SAM 3 to achieve precise, pixel-level forgery localization. |
Yiming Wang; Baiqi Wu; Qingming Li; Jiahao Chen; Tong Zhang; Shouling Ji; | code |
| 547 | Scaling The Scaling Logic: Agentic Meta-Synthesis of Logic Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose SSLogic, an agentic meta-synthesis framework that scales at the task-family level by iteratively synthesizing and refining executable Generator–Validator program pairs in a closed Generate–Validate–Refine loop, enabling continuous family evolution with controllable difficulty. |
Bowen LIU; Zhi Wu; RunquanXie; Zhanhui Kang; Jia Li; | code |
| 548 | InfVSR: Toward Consistency-Driven Streaming Generative Video Super-Resolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing generative video super-resolution (VSR) approaches, however, face two persistent challenges when processing long sequences: (1) inefficiency due to the heavy cost of multi-step denoising for full-length sequences; and (2) poor consistency hindered by temporal decomposition that causes artifacts and discontinuities. To break these limits, we propose InfVSR, which reformulates VSR as an autoregressive-one-step-diffusion paradigm, and enables streaming inference with video diffusion priors. |
Ziqing Zhang; Kai Liu; Zheng Chen; Xi Li; Yucong Chen; Bingnan Duan; Linghe Kong; Yulun Zhang; | code |
| 549 | PDAgent: An LLM-Driven Autonomous Agent Framework Towards *In Silico* Protein Design Via Directed Mutation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present PDAgent, an LLM-driven autonomous agent framework that enables *in silico* protein design through template-based directed mutation. |
Song Ouyang; Zhijie Dong; Yong Luo; Kehua Su; Huangxuan Zhao; Miaojing Shi; Bo Du; | code |
| 550 | Any2Any: Unified Arbitrary Modality Translation for Remote Sensing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We formulate Any-to-Any translation as inference over a shared latent representation of the scene, where different modalities correspond to partial observations of the same underlying semantics. Based on this formulation, we propose Any2Any, a unified latent diffusion framework that projects heterogeneous inputs into a geometrically aligned latent space. |
Haoyang Chen; Jing Zhang; Hebaixu Wang; Shiqin Wang; Pohsun Huang; Jiayuan Li; Haonan Guo; Di Wang; Zheng Wang; Bo Du; | code |
| 551 | DSB: Dynamic Sliding Block Scheduling for Diffusion LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we analyze the limitations of naive block scheduling and disclose the importance of dynamically adapting the schedule to semantic difficulty for reliable and efficient inference. |
Lizhuo Luo; Shenggui Li; Yonggang Wen; Tianwei Zhang; | code |
| 552 | MLUBench: A Benchmark for Lifelong Unlearning Evaluation in MLLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Continually unlearning from one modality could degrade the entire model. To alleviate this challenge, we propose LUMoE, an effective and efficient method. |
He Li; Haoang Chi; Qizhou Wang; Yunxin Mao; Zhiheng Zhang; Jie Tan; Tongliang Liu; Wenjing Yang; Bo Han; | code |
| 553 | Towards Generalizable EEG-to-fMRI Synthesis Via A Unified, Context-Aware Prompting Framework Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose UniEFS, a unified EEG-to-fMRI synthesis model that enables full-brain fMRI reconstruction while accommodating varying demographic and physiological contexts within a single model. |
Yamin Li; Shiyu Wang; Chang Li; Ange Lou; Haatef Pourmotabbed; Sarah Goodale; Dario Englot; Daniel Moyer; Roza G Bayrak; Catie Chang; | code |
| 554 | SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in A Single Pass Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a pretraining and instruction fine-tuning pipeline, and train our hypernetwork to generate high quality LoRA adapters from diverse meaningful contexts in a single forward pass. |
Yewei Liu; Xiyuan Wang; Yansheng Mao; Yoav Gelberg; Haggai Maron; Muhan Zhang; | code |
| 555 | MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce MMBench-Live, a multi-agent-driven dynamic multimodal benchmark that supports continuous updates without human in the loop. |
Yuanzhi Liu; Shousheng Zhao; Bo Zhou; Kongming Liang; Zhanyu Ma; | code |
| 556 | Scale-Aware Domain Harmonization for Domain Adaptation Person Search Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These inconsistency arises from different variations in Camera height, tilt angle, and focal length change. To address this challenge, we propose a Scale-Aware Consistent Alignment Learning (SCALE) framework. |
Guojian Zhao; Huibing Wang; Jinjia Peng; Linfeng Qi; Mingze Yao; Jiqing Zhang; | code |
| 557 | Reliability-Aware LLM Alignment from Inconsistent Human Feedback Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose $\textit{Reliability-Guided Preference Optimization}$ (RGPO), a robust framework designed to mitigate the impact of inconsistent human feedback. |
Jingyi Huang; Ruohan Zong; Yujun Feng; Liran Ma; Lanyu Shang; Yang Zhang; | code |
| 558 | ScoreMix: Synthetic Data Generation By Score Composition in Diffusion Models Improves Face Recognition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose **ScoreMix**, a **self-contained data augmentation** method to boost recognition performance by leveraging score compositionality in class-conditioned diffusion models. |
Parsa Rahimi; Sébastien Marcel; | code |
| 559 | A Generalist Pair-wise Progress Critic Model for Vision-Language-Action Robots Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent advances in Vision-Language-Action (VLA) models have significantly improved robotic perception and manipulation capabilities, but still struggling to adapt in dynamic, open-ended real-world environments due to a lack of reliable task progress feedback and improvement mechanisms. To address these challenges, we propose a generalist Vision Language Action-Critic model, VLAC, which can integrate both human and robot data, and unify action policy and task progress critic within a single autoregressive architecture. |
Qi Zhang; shaopeng zhai; Shengzhe Zhang; Litao Liu; TianyiZhang; huang; Ming Zhou; | code |
| 560 | MASPOB: Bandit-Based Prompt Optimization for Multi-Agent Systems with Graph Neural Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, real-world deployment is impeded by three key challenges: (1) the need for high sample efficiency due to prohibitive evaluation costs, (2) topology-induced coupling among prompts, and (3) the combinatorial explosion of the search space. To address these challenges, we introduce **MASPOB** (**M**ulti-**A**gent **S**ystem **P**rompt **O**ptimization via **B**andits), a novel sample-efficient framework based on bandits. |
Zhi Hong; Qian Zhang; Jiahang Sun; Zhiwei Shang; Mingze Kong; Xiangyi Wang; Yao Shu; Zhongxiang Dai; | code |
| 561 | ChaosNexus: A Foundation Model for ODE-based Chaotic System Forecasting with Hierarchical Multi-scale Awareness Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing architectures often fail to capture the multi-scale temporal structures and distinct spectral characteristics of chaotic dynamics. To address this, we introduce ChaosNexus, a foundation model for chaotic system forecasting underpinned by the proposed ScaleFormer architecture. |
Chang Liu; Bohao Zhao; Ding; Yong Li; | code |
| 562 | EEG-FM-Bench: A Comprehensive Benchmark for The Systematic Evaluation and Diagnostic Analyses of EEG Foundation Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current evaluations rely on inconsistent protocols that render cross-model comparisons unreliable, while a lack of diagnostic analyses obscures the internal mechanisms driving transfer efficiency and scaling behaviors. To address this, we introduce **EEG-FM-Bench**, a unified system for the standardized evaluation of EEG-FMs. |
Wei Xiong; Jiangtong Li; Jie Li; Kun Zhu; Changjun Jiang; | code |
| 563 | Physics from Video: Identifiability of Time-Invariant Second-Order ODEs Under Minimal Trajectory Conditions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We prove that a level-set slope-coverage condition ensures the learned latent space is locally affine to the true physical state, enabling exact parameter recovery. |
Yuanyuan Wang; Wenjie Wang; Kun Zhang; Mingming Gong; | code |
| 564 | Hydra-Nav: Object Navigation Via Adaptive Dual-Process Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address both the ineffectiveness and inefficiency of existing approaches, we introduce Hydra-Nav, a unified VLM architecture that adaptively switches between a deliberative slow system for analyzing exploration history and formulating high-level plans, and a reactive fast system for efficient execution. |
Zixuan Wang; Huang Fang; Shaoan Wang; Yuanfei Luo; Heng Dong; Wei Li; Yiming Gan; | code |
| 565 | AgentWebBench: Benchmarking Multi-Agent Coordination in Agentic Web Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This shift moves information access from centralized retrieval to decentralized coordination. To study this setting, we introduce AgentWebBench, a benchmark that evaluates how well a user agent synthesizes answers by interacting with website-specific content agents. |
Shanshan Zhong; Kate Shen; Chenyan Xiong; | code |
| 566 | Recovering Hidden Reward in Diffusion-Based Policies Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introduces EnergyFlow, a framework that unifies generative action modeling with inverse reinforcement learning by parameterizing a scalar energy function whose gradient is the denoising field. |
Yanbiao Ji; Qiuchang Li; Yuting Hu; Shaokai Wu; Wenyuan XIE; Guodong ZHANG; Qichen He; Deyi Ji; Yue Ding; Hongtao Lu; | code |
| 567 | FocalPolicy: Frequency-Optimized Chunking and Locally Anchored Flow Matching for Coherent Visuomotor Policy Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a foresight composite objective that supervises time-domain alignment within the proximal actions while regularizing frequency-domain structure over multiple future action chunks to improve cross-chunk coherence. |
Qian He; Zhenshuo Yang; Wenqi Liang; Chunhui Hao; Nicu Sebe; Jiandong Tian; | code |
| 568 | ReNF: Rethinking The Principles of Neural Long-Term Time Series Forecasters Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we revisit the principles of LTSF. |
Yihang Lu; Xianwei Meng; Enhong Chen; | code |
| 569 | PerturbDiff: Functional Diffusion for Single-Cell Perturbation Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In reality, responses vary systematically due to unobservable latent factors such as microenvironmental fluctuations and complex batch effects, forming a _manifold_ of possible distributions for the same observed conditions. To capture this variability, we introduce PerturbDiff, which shifts modeling from individual cells to entire distributions. |
Xinyu Yuan; Xixian Liu; Ya Shi Zhang; Zuobai Zhang; Hongyu Guo; Jian Tang; | code |
| 570 | Are First-Order Diffusion Samplers Really Slower? A Fast Forward-Value Approach Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a novel training-free, first-order sampler named Forward DPMSolver (F-DPMSolver), whose leading discretization error has the opposite sign to that of DDIM. |
Yuchen Jiao; Na Li; Changxiao Cai; Gen Li; | code |
| 571 | From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We find that questions reducible to pure state tracking (e.g., “belief”) are especially shortcut-prone compared to mind questions (e.g., “intention”) where reasoning beyond tracking is required. |
Jike Zhong; Yuxiang Lai; Ming Li; Yuheng Li; Wuao Liu; Behzad Dariush; Konstantinos Psounis; Shao-Yuan Lo; | code |
| 572 | Escaping The Subspace Trap: The Role of Optimizer Geometry in Model Width Expansion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, we identify a phenomenon termed the _**Subspace Trap**_: during continual pre-training, parameter updates largely stagnate within a low-dimensional subspace aligned with the initialization, limiting the effective capacity of the expanded model. Our theoretical analysis investigates this issue by attributing it to the function-preserving properties of width expansion. |
Jiabei Chen; Haoyu Wang; Yang Yu; Yao Xu; Liangdong Wang; Guang Liu; Shizhu He; Jun Zhao; Kang Liu; | code |
| 573 | Robust Self-reflective Hashing for Cross-modal Retrieval with Noisy Label Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing CMH methods often overlook the uncertainty introduced by noise or semantic ambiguity, making models susceptible to overfitting noisy labels and yielding unreliable similarity judgments during inference. To address this issue, we propose a Robust Self-reflective Hashing (RSH) framework that prudently analyzes semantic discrepancies while accounting for uncertainty, thereby effectively mitigating interference from noisy labels. |
Hao Sun; Qibing Qin; Lei Huang; | code |
| 574 | AgentTailor: A Semantic-Aware LLM-Based Multi-Agent System with Actor-Critic Structure Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To better utilize the semantic information in the execution stage to further optimize the structure of multi-agent systems and reduce token costs, we propose **AgentTailor**, a cost-aware framework that evaluates the semantic contribution of communication edges via an edge judgment mechanism, and employs an **Edge Prediction Network (EPN)** to estimate edge utilities through virtual execution without invoking LLMs. |
Peiting Yang; Jiahao Shi; Caiyi Xu; Ming Liu; Yanxia Wu; Rongsheng Li; | code |
| 575 | SegPVSG: Panoptic Video Scene Graph Generation Via Temporal Focusing and Generative Augmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Second, the distribution of relations exhibits a significant long-tailed pattern, making models struggle to perform well on tail categories with insufficient data. To address these issues, we propose SegPVSG, an innovative, temporal-segment-aware PVSG framework consisting of two key components: TempFocusNet (TFN) and Relation-centric Generative Video Augmentation (RGVA) module. |
YiKai Li; Quhui Ke; Jinglin Liang; Zhi-Yuan Zhang; Zhidi Lin; Shuangping Huang; | code |
| 576 | Mitigating Gradient Pathology in PINNs Through Aligned Constraint Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current solutions, such as adaptive weighting or hard constraints, either fail to fundamentally resolve this ill-conditioning or are limited to simple geometries. In this study, we systematically analyzed the causes of this gradient pathology from the perspectives of loss landscapes and optimization dynamics. |
Yichen Luo; Peiyu Zhu; Dongxiao Hu; Jia Wang; Tailin Wu; Dapeng Lan; Yu Liu; Zhibo Pang; | code |
| 577 | FedScar: Correcting Geometric Bias for Flatness-Consistent Federated Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This mismatch gives rise to a flatness discrepancy induced by misaligned loss landscapes. To address this issue, we propose FedScar, a federated optimization framework that explicitly corrects heterogeneity-induced geometric inconsistency. |
Jianfeng Lu; YuZhao Xiang; Yue Chen; Gang Li; Shuqin Cao; Guanghui Wen; | code |
| 578 | When Attributes Disagree: Gradient Conflict in Image Aesthetic Assessment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Under end-to-end training with only overall-score supervision, attribute signals are blended, which can cause gradient conflict across samples dominated by different attributes, resulting in gradient cancellation and persistent systematic bias. To address these issues, we propose AGREE (Attribute-guided Gradient Routing for Establishing Agreement), which learns attribute-specific subspaces and performs gradient routing based on sample-wise attribute sensitivity estimated via perturbation analysis. |
Ye Wang; Maocai Dai; Jiang Xie; Xiuli Bi; Fei Tao; Xiao Li; Hong Yu; | code |
| 579 | Scalable Event Cloud Network for Event-based Classification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a \textbf{S}calable \textbf{N}etwork named SECNet to leverage \textbf{E}vent \textbf{C}loud representation. |
Hongwei Ren; Fei Ma; Xiaopeng LIN; Yuetong Fang; Hongxiang Huang; Yue Zhou; Yulong Huang; Haotian FU; Ziyi Yang; Youxin Jiang; Xiangqian Wu; Bojun Cheng; | code |
| 580 | Emergence of Exploration in Policy Gradient Reinforcement Learning Via Retrying Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: For efficient policy optimization, we derive a new policy-gradient formulation for ReMax and introduce **Re**Max **PPO** (**RePPO**), a PPO variant that optimizes ReMax while generalizing the discrete retry count $M$ to a continuous parameter $m > 0$, enabling fine-grained control of exploration. |
Soichiro Nishimori; Paavo Parmas; Sotetsu Koyamada; Tadashi Kozuno; Toshinori Kitamura; Shin Ishii; Yutaka Matsuo; | code |
| 581 | Talk, Judge, Cooperate: Gossip-Driven Indirect Reciprocity in Self-Interested LLM Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Agentic Linguistic Gossip Network (ALIGN), an automated framework where agents strategically share open-ended gossip using hierarchical tones to evaluate trustworthiness and coordinate social norms. |
Shuhui Zhu; Yue Lin; Shriya Kaistha; Wenhao Li; Baoxiang Wang; Hongyuan Zha; Gillian Hadfield; Pascal Poupart; | code |
| 582 | Learning Stochastic Bridges for Video Object Removal Via Video-to-Video Translation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we reformulate video object removal as a video-to-video translation task via a stochastic bridge model. |
Zijie Lou; Xiangwei Feng; Jiaxin Wang; Jiangtao Yao; Fei Che; Tianbao Liu; WU CHENGJING; Xiaochao Qu; Luoqi Liu; Ting Liu; | code |
| 583 | Reference-Free Meta-Learning for Generalized Implicit Neural Representation in Efficient MRI Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work presents IPOD, a Reference-Free Meta-Learning framework designed to learn generalized parameter initializations for INR directly from undersampled data. |
Haonan Zhang; Qing Wu; Xuanyu Tian; Bowen Li; Yuyao Zhang; Hongjiang Wei; | code |
| 584 | One-Step Graph-Structured Neural Flows for Irregular Multivariate Time Series Classification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Moreover, their one-step mapping makes interaction modeling inherently challenging, as it removes the iterative refinement of interactions during learning. To address this challenge, we propose one-step Graph-Structured Neural Flows (GSNF), which introduce two auxiliary-trajectory self-supervision strategies to strengthen interaction learning: (i) interaction-aware trajectory generation via re-initialization, which induces trajectory divergence to expose graph-induced interactions, with a theoretically derived lower bound on divergence; and (ii) reverse-time trajectory generation, which enforces forward–backward consistency to regularize graph learning, enabled by flow invertibility. |
Mengzhou Gao; kaiwei wang; Pengfei Jiao; | code |
| 585 | MV-FGAD: Towards Efficient and Effective Federated Graph Anomaly Detection Via Multi-view Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, we propose MV-FGAD, an efficient and effective federated GAD framework based on multi-view learning designed to mine anomalies of varying strengths. |
Junyi Yan; KE LIANG; Hao Yu; Meng Liu; Hao Tan; Tianrui Liu; Jun-Jie Huang; Xinwang Liu; | code |
| 586 | Dynamic TMoE: A Drift-Aware Dynamic Mixture of Experts Framework for Non-Stationary Time Series Forecasting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While Mixture-of-Experts (MoE) architectures offer a promising paradigm for decoupling complex drift patterns, existing approaches are limited by fixed expert pools and memoryless routing, hampering their ability to adapt to abrupt regime shifts. To address this, we propose **Dynamic TMoE**, a framework that unifies architectural evolution with temporal continuity during learning phase. |
Jiawen Zhu; Shuhan Liu; Di Weng; Yingcai Wu; | code |
| 587 | AesFormer: Transform Everyday Photos Into Beautiful Memories Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although recent advances in image editing models make APR feasible, they often lack aesthetic understanding, yielding edits that are semantically plausible yet aesthetically weak. To address this, we propose AesFormer, a two-stage framework that decouples aesthetic planning from image editing. |
Tianxiang Du; Hulingxiao He; Yuxin Peng; | code |
| 588 | Deep Discriminative Structure Proxy Hashing for Cross-modal Retrieval Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Existing proxy-based hashing methods optimize samples toward independently learned proxies using isolated similarity constraints. Although efficient, this design overlooks the … |
Kun Cheng; Qibing Qin; Lei Huang; | code |
| 589 | See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose the Differentiable Grid Sampler (GridS), a plug-and-play module that performs task-aware, continuous resampling of visual tokens in VLA. |
Yixu Feng; Zinan Zhao; Yanxiang Ma; Chenghao Xia; Chengbin Du; Yunke Wang; Chang Xu; | code |
| 590 | Break The Block: Dynamic-size Reasoning Blocks for Diffusion Large Language Models Via Monotonic Entropy Descent with Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Through empirical observations, we reveal that, for block-wise entropy, incorrect reasoning exhibits a fluctuating and unsteady trend between blocks, while the correctly generated tasks follow a consistent descending paradigm. Therefore, this paper proposes b1, a novel post-training framework that learns dynamic-size reasoning blocks via a Monotonic Entropy Descent objective with reinforcement learning to enhance reasoning coherence. |
Yan Jiang; Ruihong Qiu; Zi Huang; | code |
| 591 | Relative Entropy Estimation in Function Space: Theory and Applications to Trajectory Inference Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a general framework for estimating the Kullback–Leibler divergence (KL) between probability measures on function space: we obtain a tractable estimator that can be approximated from data, is practical, and scales to realistic problem sizes (number and size of snapshot data). |
CHAO WANG; Luca Nepote; Giulio Franzese; Pietro Michiardi; | code |
| 592 | EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We demonstrate that aggressive sparse sampling on standard position-encoded sequences violates the Nyquist limit relative to the effective token interval, causing phase-wrapping collisions that corrupt temporal monotonicity. To address this, we introduce EchoingPixels, a framework for aliasing-resistant joint token reduction. |
Chao Gong; Depeng Wang; Zhipeng Wei; Ya Guo; Huijia Zhu; Jingjing Chen; | code |
| 593 | Internalizing Safety Understanding in Large Reasoning Models Via Verification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We argue that this approach remains largely behavioral: our empirical analysis reveals that ostensibly aligned models lack intrinsic safety understanding, often failing to verify their own response safety and remaining vulnerable to adversarial jailbreaks. To address this fundamental limitation, we propose Safety Internal (SInternal), a framework that internalizes safety specifications by training LRMs exclusively on safety verification tasks to critique their own generated answers using expert reasoning trajectories. |
Yi Zhang; Yuxin Chen; Leheng Sheng; Dongcheng Zhang; Chaochao Lu; Xiang Wang; An Zhang; | code |
| 594 | The Double Dilemma in Multi-Task Radiology Report Generation: A Gradient Dynamics Analysis and Solution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address these problems, we analyze the failure mechanism of linear scalarization from the perspective of gradient dynamics for the first time, utilizing the stochastic differential equation (SDE) framework to formalize it as a Double Dilemma” of drift term deviation and diffusion term decay. Based on this, we propose a backbone-agnostic optimizer named Conflict-Averse Magnitude-Enhanced Gradient Descent (CAME-Grad). |
Erjian Zhang; Yatong Hao; Liejun Wang; Zhiqing Guo; | code |
| 595 | Protein Fold Classification at Scale: Benchmarking and Pretraining Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We show that on TEDBench, current protein representation learning methods either require very large models or fail to deliver strong performance. To address this challenge, we propose Masked Invariant Autoencoders (MiAE), a self-supervised framework for protein structure representation learning. |
Dexiong Chen; Andrei Manolache; Mathias Niepert; Karsten Borgwardt; | code |
| 596 | PrivAct: Internalizing Contextual Privacy Preservation Via Multi-Agent Preference Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose **PrivAct**, a contextual privacy-aware multi-agent learning framework that *internalizes* contextual privacy preservation directly into models’ generation behavior for privacy-compliant agentic actions. |
Yuhan Cheng; Hancheng Ye; Hai Li; Jingwei Sun; Yiran Chen; | code |
| 597 | Q-DiT4SR: Exploration of Detail-Preserving Diffusion Transformer Quantization for Real-World Image Super-Resolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose **H-SVD**, a hierarchical SVD that integrates a global low-rank branch with a local block-wise rank-1 branch under a matched parameter budget. |
Xun Zhang; Kaicheng Yang; Hongliang Lu; Haotong Qin; Yong Guo; Yulun Zhang; | code |
| 598 | MDGMIX: Boundary-Aware Subgraph Mixing for Multi-Domain Graph Pre-Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper empirically reveals significant data redundancy in multi-domain graph pre-training. Based on this finding, we propose the Multi-domain Graph Pre-training Framework, MDGMIX, which combines boundary-aware subgraph mixing with hierarchical discrimination. |
Ziyu Zheng; Yaming Yang; Ziyu Guan; Wei Zhao; Xinyan Huang; | code |
| 599 | AC-ODM: Actor–Critic Online Data Mixing for Sample-Efficient LLM Pretraining Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce \textbf{Actor–Critic Online Data Mixing (AC-ODM)}, which approaches data mixing from a reinforcement learning perspective with a parameterized policy that we theoretically prove to act as a dynamic linear surrogate maximizing the constructive interference of gradients. |
Jing Ma; Chenhao Dang; Mingjie Liao; | code |
| 600 | Optimal Decision-Making Based on Prediction Sets Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Here, we propose a decision-theoretic framework that seeks to minimize the expected loss (risk) against a worst-case distribution consistent with the prediction set’s coverage guarantee. |
Tao Wang; Edgar Dobriban; | code |
| 601 | Strategy-Aware Optimization Modeling with Reasoning LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose **SAGE**, a strategy-aware framework that makes *Modeling Strategy* explicit in both data construction and post-training. |
Ruiqing Zhao; Fengzhi Li; Yuan Zuo; Rui Liu; YanSong Liu; Yunfei Ma; Fanyu Meng; JUNLAN FENG; | code |
| 602 | Normality Calibration in Semi-supervised Graph Anomaly Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, the normality learned by existing semi-supervised GAD methods is limited to the labeled normal nodes, often inclining to overfitting the given patterns, thereby leading to high detection errors, such as high false positives. To overcome this limitation, we propose $GraphNC$, a graph normality calibration framework that leverages both labeled and unlabeled data to calibrate the normality from a teacher (a pre-trained semi-supervised GAD model) jointly in anomaly score and representation spaces. |
Guolei Zeng; Hezhe Qiao; GUOGUO AI; Jinsong Guo; Guansong Pang; | code |
| 603 | Uncovering Latent Communication Patterns in Brain Networks Via Adaptive Flow Routing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we formulate multi-modal fusion through the lens of neural communication dynamics and propose the Adaptive Flow Routing Network (AFR-Net), a physics-informed framework that models how structural constraints (SC) give rise to functional communication patterns (FC), enabling interpretable discovery of critical neural pathways. |
Tianhao Huang; Guanghui Min; zhenyu lei; Aiying Zhang; Chen Chen; | code |
| 604 | Sample from What You See: Visuomotor Policy Learning Via Diffusion Bridge with Observation-Embedded Stochastic Differential Equation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose BridgePolicy, a generative visuomotor policy that directly integrates observations into the stochastic dynamics via a diffusion-bridge formulation. |
Zhaoyang Liu; Mokai Pan; Zhongyi Wang; Kaizhen Zhu; Haotao Lu; Haipeng Zhang; Jingya Wang; Ye Shi; | code |
| 605 | SMD: Multi-view Safety-Critical Driving Video Generation in The Real-world Domain Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present SMD, the first framework for generating multi-view safety-critical driving videos in the real-world domain. |
Jiawei Zhou; Linye Lyu; Zhuotao Tian; Cheng Zhuo; YU LI; | code |
| 606 | Guided Star-Shaped Masked Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a novel sampling algorithm that works with pre-trained models and, after a lightweight fine-tuning of a single layer, significantly improves sample quality and efficiency. |
Viacheslav Meshchaninov; Egor Shibaev; Artem Makoian; Ivan Klimov; Nikita Balagansky; Daniil Gavrilov; Aibek Alanov; Dmitry Vetrov; | code |
| 607 | TransNormal: Dense Visual Semantics for Diffusion-based Transparent Object Normal Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Furthermore, we employ a multi-task learning objective and wavelet-based regularization to ensure the preservation of fine-grained structural details. To support this task, we introduce **TransNormal-Synthetic**, a physics-based dataset with high-fidelity normal maps for transparent labware. |
Mingwei Li; Hehe Fan; Yi Yang; | code |
| 608 | Social Hippocampus Memory Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose **SoHip** (**So**cial **Hip**pocampus Memory Learning), a memory-centric social machine learning framework that enables collaboration among heterogeneous agents via memory sharing rather than model sharing. |
Liping Yi; Zhiming Zhao; Kewen Zhu; Xiang Li; Zhiwei Shang; Qinghua Hu; | code |
| 609 | Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, the mismatch between training and inference, coupled with the challenge of maintaining long-term consistency, limits the effective utilization of foundation models. To mitigate these concerns, we propose MIGA, a novel infinite-frame long video generation method. |
xiaokun Feng; Jiashu Zhu; Meiqi Wu; Chubin Chen; Fangyuan Mao; Haiyang Guo; Jiahong Wu; Xiangxiang Chu; Kaiqi Huang; | code |
| 610 | SplAttN: Bridging 2D and 3D with Gaussian Soft Splatting and Attention for Point Cloud Completion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent works attribute success to the connection between modalities, yet we identify that standard hard projection severs this connection, inducing Cross-Modal Entropy Collapse where sparse support hinders visual prior propagation. To bridge this gap, we propose SplAttN, which maximizes Point-wise Mutual Information via Differentiable Gaussian Splatting. |
Zhaoyang Li; Zhichao You; Tianrui Li; | code |
| 611 | BizFinBench.v2: Towards Reliable LLMs in Finance Via Real-User Data and Offline/Online Bilingual Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Nevertheless, prevailing benchmarks are largely dependent on simulated or generic data, which leads to a significant gap between reported performance and actual efficacy in real-world scenarios. To tackle this challenge, we present BizFinBench.v2, the first integrated offline and online benchmark built upon authentic user query-response data from both Chinese and U.S. equity markets. |
Xin Guo; Rongjunchen Zhang; Guilong Lu; Xuntao Guo; Jia Shuai; Zhi Yang; Liwen Zhang; | code |
| 612 | AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Afterward, we propose AgentHijack-Agent, a framework that integrates an action generator with enhanced grounding capabilities and an onlooker responsible for behavior summarization and environment checking. |
Jingwei Sun; Jianing Zhu; Yuanyi Li; Tongliang Liu; Xia Hu; Bo Han; | code |
| 613 | SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As a result, metrics like pass@k become impractical under realistic budget constraints. To address this gap, we introduce SimulCost, the first benchmark targeting cost-sensitive parameter tuning in scientific simulations. |
Yadi Cao; Sicheng Lai; Jiahe Huang; Yang Zhang; Zach Lawrence; Rohan Bhakta; Izzy Thomas; Mingyun Cao; Chung-Hao Tsai; Zihao Zhou; Yidong Zhao; Hao Liu; Alessandro Marinoni; Alexey Arefiev; Rose Yu; | code |
| 614 | FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although existing acceleration methods reduce runtime for large-scale steps, but rely on manual step selection and overlook the varying importance of different stages in the generation process. To address this challenge, we present FasterVAR, a systematic study and stage-aware acceleration framework for VAR models. |
Senmao Li; Kai Wang; Salman Khan; Fahad Khan; jian Yang; Yaxing Wang; | code |
| 615 | Error Amplification Limits ANN-to-SNN Conversion in Continuous Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We identify error amplification as the key cause: small action approximation errors become temporally correlated across decision steps, inducing cumulative state distribution shift and severe performance degradation. To address this issue, we propose Cross-Step Residual Potential Initialization (CRPI), a lightweight training-free mechanism that carries over residual membrane potentials across decision steps to suppress temporally correlated errors. |
Zijie Xu; Zihan Huang; Yiting Dong; Kang Chen; Wenxuan Liu; Zhaofei Yu; | code |
| 616 | AliMark: Enhancing Robustness of Sentence-Level Watermarks Against Text Paraphrasing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, their prefix-based designs remain vulnerable to structural perturbations, such as sentence splitting and merging, which commonly arise under strong paraphrasers like DIPPER and GPT-3.5. To mitigate this issue, we propose AliMark, a framework that reformulates sentence-level watermarking as a bit sequence encoding and alignment problem between a potentially watermarked text and a secret bit sequence. |
YUEXIN LI; Wenjie Qu; Linyu Wu; Yulin Chen; Yufei He; Tri Cao; Bryan Hooi; Jiaheng Zhang; | code |
| 617 | Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted Pretraining Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This tight coupling complicates optimization and often results in unstable training and suboptimal generalization. To address these limitations, we propose BabelRS, a unified language-pivoted pretraining framework that explicitly decouples modality alignment from downstream task learning. |
Yuxuan Li; Yuming Chen; Yunheng Li; Ming-Ming Cheng; Xiang Li; jian Yang; | code |
| 618 | Calibrated Test-Time Guidance for Bayesian Inference Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing approaches, however, focus on reward maximization rather than sampling from the true Bayesian posterior, leading to miscalibrated inference. In this work, we show that common test-time guidance methods do not recover the correct posterior distribution and identify the structural approximations responsible for this failure. |
Daniel Geyfman; Felix Draxler; Jan Groeneveld; Hyunsoo Lee; Theofanis Karaletsos; Stephan Mandt; | code |
| 619 | LoCoT2V-Bench: Benchmarking Long-Form and Complex Text-to-Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent advances in text-to-video generation have achieved impressive performance on short clips, yet evaluating long-form generation under complex textual inputs remains a significant challenge. In response to this challenge, we present LoCoT2V-Bench, a benchmark for long video generation (LVG) featuring multi-scene prompts with hierarchical metadata (e.g., character settings and camera behaviors), constructed from collected real-world videos. |
Xiangqing Zheng; CHENGYUE WU; Kehai Chen; Min zhang; | code |
| 620 | How Do Human Processes AI-generated Hallucination Contents: A Neuroimaging Study Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While AI-generated hallucinations pose considerable risks, the underlying cognitive mechanisms by which humans can successfully recognize or be misled by these hallucinations remain unclear. To address this problem, this paper explores humans’ neural dynamics to characterize how the brain processes hallucinated content. |
Shuqi Zhu; Yi Zhong; Ziyi Ye; Bangde Du; Yujia Zhou; Qingyao Ai; Yiqun LIU; | code |
| 621 | $\sigma$: Sigmoid Modulation for Ultra High Resolution Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing extrapolation methods typically operate under a *scale-agnostic* assumption, treating the denoising dynamics identically across resolutions. In this work, we identify a critical oversight in this paradigm: the spectral evolution of the diffusion process, transitioning from low-frequency structural construction to high-frequency texture refinement, is inherently scale-dependent. |
Bingxuan Zhao; Qing Zhou; Yu Wang; Chuang Yang; Qi Wang; | code |
| 622 | LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Yet, current coding benchmarks face limitations such as lack of exceptionally challenging problems, insufficient test case coverage, reliance on online platform APIs that limit accessibility. To address these issues, we introduce LiveOIBench, a large-scale competitive programming benchmark featuring 403 expert-curated problems, averaging $60$ official test cases each, drawn from 72 contests across 14 Informatics Olympiads held between 2023 and 2025. |
Kaijian Zou; Feiyang Xiong; Yunxiang Zhang; Xinliang Frederick Zhang; Yueqi Ren; Shitanshu Bhushan; Ayoung Lee; Jirong Yang; Lu Wang; | code |
| 623 | Mosaic: Unlocking Over 30$\times$ Context Length for Diffusion LLMs Inference Via Global Memory Planning and Dynamic Peak Taming Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Moreover, generic memory reuse mechanisms lack the global visibility to handle dynamic memory peaks of dLLMs, which alternate between logits and feed-forward networks. To address these challenges, we present Mosaic, a memory-efficient inference system that shifts dLLM execution from local, static memory management to a global, dynamic paradigm. |
Zheng Liang; Bowen Shi; Yitao Hu; Jiawei Zhang; Ruofan Li; Guotao Yang; Zhixin Zhao; Zhengchao Wang; Sheng Chen; Wenxin Li; Dezhi Ran; Tao Xie; Keqiu Li; | code |
| 624 | JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A core contribution of our work is the neural intensity vector (Neural IV), a learned spatial audio representation that encodes robust directional cues to enhance direction-of-arrival estimation, even in adverse acoustic scenarios with overlapping sources. |
Zhan Liu; Changli Tang; Yuxin WANG; Zhiyuan Zhu; Youjun Chen; Yiwen Shao; TIANZI WANG; Lei Ke; Zengrui Jin; Chao Zhang; | code |
| 625 | Understanding Dynamic Compute Allocation in Recurrent Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Token-level adaptive computation seeks to reduce inference cost by allocating more computation to harder tokens and less to easier ones. However, prior work is primarily evaluated … |
Ibraheem Muhammad Moosa; Suhas Lohit; Ye Wang; Moitreya Chatterjee; Wenpeng Yin; | code |
| 626 | A Fully First-Order Layer for Differentiable Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing approaches typically rely on implicit differentiation, involving expensive Hessian computation while differentiating through optimality conditions. To address this challenge, we formulate the differentiable optimization problem as a bilevel optimization instance. |
Zihao Zhao; Kai-Chia Mo; Shing-Hei Ho; Brandon Amos; Kai Wang; | code |
| 627 | TextResNet: Decoupling and Routing Optimization Signals in Compound AI Systems Via Deep Residual Tuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In standard textual backpropagation, feedback signals mix local critiques with upstream contexts, leading to *Attribution Ambiguity*. To address this challenge, we propose TextResNet, a framework that reformulates the optimization process to achieve precise signal routing via four key innovations. |
Suizhi Huang; Mei Li; Han Yu; Xiaoxiao Li; | code |
| 628 | EffGen: Enabling Small Language Models As Capable Autonomous Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce $\textbf{effGen}$, an open-source agentic framework optimized for small language models (SLMs) that enables effective, efficient, and secure local deployment. |
Gaurav Srivastava; Aafiya Hussain; Chi Wang; Yingyan (Celine) Lin; Xuan Wang; | code |
| 629 | SARSteer: Safeguarding Large Audio Language Models Via Safe-Ablated Refusal Steering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While safety alignment has made initial advances in LLMs and Large Vision–Language Models (LVLMs), we find that vanilla adaptation of these approaches to LALMs faces two key limitations: 1) LLM-based steering fails under audio input due to the large distributional gap between activations, and 2) prompt-based defenses induce over-refusals on benign-speech queries. To address these challenges, we propose **S**afe-**A**blated **R**efusal **Steer**ing (SARSteer), an effective inference-time defense framework for LALMs. |
Weilin Lin; Jianze Li; Hui Xiong; Li Liu; | code |
| 630 | See The Emotion: A Facial Emoji Proxy Modeling for EEG Emotion Recognition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we reframe EEG explainability as a cross-modal generation task, shifting the paradigm from feature attribution to behavioral visualization. |
Jingjing Hu; Dan Guo; Haofan Cheng; Zeng ying; Zhan Si; Jinxing Zhou; Meng Wang; | code |
| 631 | SEPS: Semantic-Enhanced Patch Slimming Framework for Fine-Grained Cross-Modal Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While Multimodal Large Language Models (MLLMs) offer rich descriptive capabilities, their naive integration often induces semantic drift due to inconsistencies with sparse ground-truth captions. To systematically resolve these challenges, we present the Semantic-Enhanced Patch Slimming (SEPS) framework. |
Xinyu Mao; Junsi Li; Haoji Zhang; Yu Liang; Ming Sun; | code |
| 632 | Global Policy-Space Response Oracles for Two-Player Zero-Sum Games Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, we adopt Population Exploitability (PE) to measure how well a restricted strategy set represents the full game, and introduce a two-phase exploration–selection framework that explicitly minimizes PE during expansion. |
Junyu Zhang; Feihong Yang; Jian Wang; Chao Wang; Xudong Zhang; | code |
| 633 | Representation Drift Compensation: A Zero-Cost Enhancement for LLM Decomposition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we identify and formalize a key overlooked issue in LLM decomposition: \textit{representation drift}. |
Xinhao Huang; You-Liang Huang; Zeyi Wen; | code |
| 634 | Discrete Adjoint Schrödinger Bridge Sampler Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While stochastic optimal control (SOC) and Schrödinger bridge (SB) provide principled solutions, efficient SOC solvers like adjoint matching (AM), which excel in continuous domains, remain unexplored for discrete spaces. We bridge this gap by revealing that the core mechanism of AM is *state-space agnostic*, and introduce **discrete ASBS**, a unified framework that extends AM and adjoint Schrödinger bridge sampler (ASBS) to discrete spaces. |
Wei Guo; Yuchen Zhu; Xiaochen Du; Juno Nam; Yongxin Chen; Rafael Gomez-Bombarelli; Guan-Horng Liu; Molei Tao; Jaemoo Choi; | code |
| 635 | From Conflict to Consensus: Boosting Medical Reasoning Via Multi-Round Agentic RAG Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In the paper, we propose **MA-RAG** (**M**ulti-Round **A**gentic RAG), a framework that facilitates test-time scaling for complex medical reasoning by iteratively evolving both external evidence and internal reasoning history within an agentic refinement loop. |
Wenhao Wu; Zhentao Tang; Yafu Li; Shixiong Kai; Mingxuan Yuan; Zhenhong Sun; Chunlin Chen; Zhi Wang; | code |
| 636 | Regime-Adaptive Bayesian Optimization Via Dirichlet Process Mixtures of Gaussian Processes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose RAMBO, a Dirichlet Process Mixture of Gaussian Processes that automatically discovers latent regimes during optimization, each modeled by an independent GP with locally-optimized hyperparameters. |
Yan Zhang; Xuefeng Liu; Sipeng Chen; Sascha Ranftl; Chong Liu; Shibo Li; | code |
| 637 | Implicit Preference Alignment for Human Image Animation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Implicit Preference Alignment (IPA), a data-efficient post-training framework that eliminates the need for paired preference data. |
Yuanzhi Wang; Xuhua Ren; Jiaxiang Cheng; bing ma; Kai Yu; Tianxiang Zheng; Qinglin Lu; Zhen Cui; | code |
| 638 | DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing methods still fail to ensure evidence–answer consistency, where correct answers must be supported by correct visual evidence. To address this issue, we propose DeFacto, a counterfactual reasoning framework that explicitly aligns visual evidence with final answers by jointly optimizing for task correctness and evidence–answer consistency. |
Tianrun Xu; Haoda Jing; Ye Li; Yuquan Wei; Jun Feng; Guanyu Chen; Haichuan Gao; Tianren Zhang; Jing Liu; Feng Chen; | code |
| 639 | Semantic-Enriched Latent Visual Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing approaches largely rely on visual supervision and produce latent representations that lack sufficient semantic richness, limiting their ability to support diverse region-level reasoning tasks. In this work, we introduce Semantic-Enriched Latent Visual Reasoning (SLVR), a two-stage learning framework that enriches latent representations with attribute-level visual semantics and aligns them with diverse reasoning objectives. |
Tianrun Xu; Yue Sun; Qixun Wang; Jingyi Lu; Yuan Wang; Tianren Zhang; Longteng Guo; Fengyun Rao; Jing LYU; Jing Liu; Feng Chen; | code |
| 640 | Learning Context-Conditioned Predicate Semantics Via Prototype Feedback Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose AlignG, which learns context-conditioned predicate semantics via prototype feedback. |
NamGyu Jung; Chang Choi; | code |
| 641 | Vulnerable Agent Identification in Large-Scale Multi-Agent Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Partial agent failure becomes inevitable when systems scale up, making it crucial to identify the subset of agents whose failure causes worst-case system performance degradations. We study this Vulnerable Agent Identification (VAI) problem in large-scale multi-agent reinforcement learning (MARL). |
Simin Li; Zihao Mao; Zheng Yuwei; Linhao Wang; Ruixiao Xu; Chengdong Ma; Zhiqian Liu; Xin Yu; Yuqing Ma; Xin Wang; Jie Luo; Bo An; Yaodong Yang; Weifeng Lv; Xianglong Liu; | code |
| 642 | CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Visual expert-assisted search is efficient but prone to blind spots when proposals fail, whereas scan-based search guarantees coverage at the cost of computational redundancy and semantic fragmentation. To address this dilemma, we introduce CVSearch, a training-free adaptive framework that dynamically schedules search strategies via an Assess-then-Search workflow. |
Liupeng Li; Haoqian Kang; Zhenyu Lu; Jinpeng Wang; Bin Chen; Ke Chen; Yaowei Wang; | code |
| 643 | Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While Parameter-Efficient Fine-Tuning (PEFT) methods offer a more resource-conscious alternative, they typically require retraining for each LLM backbone due to architectural dependencies. To address these challenges, we propose Universal Reasoner (UniR)—a lightweight, composable, and plug-and-play reasoning module that can be used with larger frozen LLMs to provide specialized reasoning capabilities. |
Jaemin Kim; Hangeol Chang; Hyunmin Hwang; Choonghan Kim; Jong Chul YE; | code |
| 644 | EntroKV: Entropy-Guided Dynamic Budget Allocation for KV-Cache Compression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current compression techniques typically rely on static or uniform budget allocation, overlooking the significant heterogeneity in information density across attention heads. To address this, we introduce \textsc{EntroKV}, an entropy-driven dynamic budget allocation framework. |
Wenhao Gao; Haoran Cao; Yueyan Li; YongGao Xiao; Caixia Yuan; Xiaojie Wang; | code |
| 645 | AdaMEM: Test-Time Adaptive Memory for Language Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Consequently, agents are forced to rely on static guidance that becomes increasingly misaligned as long-horizon tasks unfold. To address this rigidity, we propose the Adaptive Memory Agent (AdaMEM), a novel framework for agent test-time adaptation. |
Yunxiang Zhang; Yiheng Li; Ali Payani; Lu Wang; | code |
| 646 | Restoring Initial Noise Sensitivity in Text-to-Image Distillation Through Geometric Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we identify a key lost property: sensitivity to initial noise, the absence of which impairs downstream control methods that rely on noise-based optimization and manipulation. |
huayang Huang; Ruoyu Wang; Jinhui Zhao; Wei Deng; Daiguo Zhou; Jian Luan; Yu Wu; Ye Zhu; | code |
| 647 | FedRGL: Robust Federated Graph Learning Under Label Noise Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing federated label noise learning methods, primarily focused on computer vision tasks, often yield suboptimal results when directly applied to FGL. To address this issue, we propose a robust federated graph learning method with label noise, termed **FedRGL**. |
De Li; Zhou Tan; Qiyu Li; Zeming Gan; Tiange Xia; Chunpei Li; Xianxian Li; | code |
| 648 | Contrastive Spectral Rectification: Test-Time Defense Towards Zero-shot Adversarial Robustness of CLIP Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We further attribute this to the model’s inherent spectral bias. Leveraging this insight, we propose an efficient test-time defense named Contrastive Spectral Rectification (CSR). |
Sen Nie; Jie Zhang; Zhuo Wang; Shiguang Shan; Xilin Chen; | code |
| 649 | MAS-Architect: Declarative Multi-Agent System Design Via Separation of Concerns Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing approaches often suffer from structural rigidity and entangle the design of system topology with the implementation of individual agents. To overcome these limitations, we propose MAS-Architect, a framework that automates MAS design through a novel code-based declarative MAS paradigm rooted in the \textit{Separation of Concerns} principle. |
Jing Huang; Lidong Zhang; Mutian Bao; Yadong Li; Xingzhong Xu; Jinjian Zhang; Jie Liu; Ming Kong; Qiang Zhu; | code |
| 650 | Smoothing Slot Attention Iterations and Recurrences Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, cold-start queries lack sample-specific cues thus hindering precise aggregation on image or video’s first frame; Non-first frames’ queries are already sample-specific thus requiring aggregation transforms different from the first frame. We address these issues with our *SmoothSA*: (1) To smooth SA iterations on image or video’s first frame, we *preheat* cold-start queries with rich input-feature information, by a tiny module self-distilled inside OCL; (2) To smooth SA recurrences across video’s first and non-first frames, we *differentiate* the homogeneous aggregation transforms by using full and single iterations respectively. |
Rongzhen Zhao; Wenyan Yang; Kannala Juho; Joni Pajarinen; | code |
| 651 | FACT: Fuzzy Alignment with Comorbidity Topology for Reliable Multi-Label Medical Image Diagnosis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Enforcing hard decision boundaries under such overlap suppresses shared evidence, biases feature representations, and ultimately undermines model reliability. To address this limitation, we propose Fuzzy Alignment with Comorbidity Topology FACT, a novel paradigm that reformulates MLD as a fuzzy alignment problem between atomic visual evidence and disease semantic anchors. |
Yingyu Chen; Yongqiang Huang; Yang Qin; Ziyuan Yang; Lang Yuan; Maosong Ran; Yi Zhang; | code |
| 652 | LitReview Arena: Evaluating Literature Review Agents with Battle-style Peer Review Platform Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, it remains a challenging task to rigorously evaluate the scientific value of the generated reviews, since human expert annotations are difficult to scale up and LLM-as-a-judge approaches lack of a convincing criteria. To address this gap, we introduce LitReview Arena, a battle-style evaluation platform with a structured protocol tailored to literature review quality. |
Ruotong Zhao; Zhiyu Chen; Xurui Liu; Haidong Xue; Dong Liang; Jigao Fu; Wu YanBiao; Yuanyi Zhen; Fengli Xu; Yong Li; | code |
| 653 | ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Reinforcement Learning (RL) post-training alignment for language models is effective, but also costly and unstable in practice, owing to its complicated training process. To address this, we propose a training-free inference method to sample directly from the optimal RL policy. |
Xiuyu Li; Jinkai Zhang; Mingyang Yi; Yu Li; Longqiang Wang; Yue Wang; Ju Fan; | code |
| 654 | Automatic Layer Selection for Hallucination Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Instead, we propose a new selection criterion, First Effective Peak of Intrinsic Dimension (FEPoID), that is able to consistently identify optimal or near-optimal layers and outperforms the aforementioned criteria and existing hallucination detection baselines. |
Xinpeng Wang; William Cao; Andrew Wilson; Zhe Zeng; | code |
| 655 | WET: Mitigating World-Conditioned Knowledge Conflicts Via World Entropy Tethering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose World Entropy Tethering (WET), an inference-time monitor-and-tether: a world-entropy probe flags drift risk on prompt anchors, and a conditional score matching geometry model identifies tethering heads for entropy-gated rescaling. |
Zixuan Wang; Yifei He; Zihan Wang; Kun Wang; Chaomeng Chen; | code |
| 656 | MoDA: Modulation Adapter for Fine-Grained Visual Understanding in Instructional MLLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing approaches often struggle with fine-grained visual grounding due to semantic entanglement in visual patch representations, where individual patches blend multiple distinct visual elements, making it difficult for models to focus on instruction-relevant details. To address this challenge, we propose MoDA (Modulation Adapter), a lightweight module that enhances visual grounding through instruction-guided channel-wise modulation. |
Wayner Barrios; Andrés Villa; Juan Leon Alcazar; SouYoung Jin; Bernard Ghanem; | code |
| 657 | An Embarrasingly Simple Way to Optimize Orthogonal Matrices at Scale Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we revisit and improve on the ideas behind Landing, enabling the inclusion of modern adaptive optimizers while ensuring that orthogonal constraints are effectively met. |
Adrián Javaloy; Antonio Vergari; | code |
| 658 | Insertion Based Sequence Generation with Learnable Order Dynamics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As the generative insertion model, we use a variable length masked diffusion model, which generates by inserting and filling mask tokens. |
Dhruvesh Patel; Benjamin Rozonoyer; Gaurav Pandey; Tahira Naseem; Ramón Astudillo; Andrew McCallum; | code |
| 659 | SFCLTA: Spectral Fusion Contrastive Learning with Topology-Adaptive Graph Augmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Nevertheless, GCL faces two critical challenges when applied to heterophilic graphs, i.e., the potential distribution shift from data augmentation and the loss of robustness caused by high-frequency signals. To address these problems, we propose a novel model, namely the Spectral Fusion Contrastive Learning with Topology-Adaptive Graph Augmentation (SFCLTA) for unsupervised graph representation learning. |
Zhuo Xu; Lu Bai; Jincheng Li; Lixin Cui; Ming Li; Hangyuan Du; Yue Wang; | code |
| 660 | SynerMedGen: Synergizing Medical Multimodal Understanding with Generation Via Task Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we identify and address a critical question in unified medical modeling: what form of “understanding” truly benefits generation. |
Weiren zhao; DONG Yi; Cheng Chen; | code |
| 661 | Anti-Aliasing Matters: A Dynamic Network for Time Series Forecasting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This erroneously introduces spurious low-frequency patterns, perceived as low-frequency noise, thereby leading to the ***aliasing problem***. To address this problem, we propose a Decomposition-Prevention-Fusion architecture framework called **DMANet**, which introduces the **D**ynamic **M**ultiscale **A**nti-Aliasing **Net**work. |
Heng Zhou; Xin Sun; Chao Li; | code |
| 662 | On The Entropy Dynamics in Reinforcement Fine-Tuning of Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we establish a theoretical framework for analyzing the entropy dynamics during the RFT process, which begins with a discriminant expression that quantifies entropy change under a single logit update. |
Shumin Wang; Yuexiang Xie; Wenhao Zhang; Yuchang Sun; Yanxi Chen; Yaliang Li; Yanyong Zhang; | code |
| 663 | MEMO: Memory-Augmented Model Context Optimization for Robust Multi-Turn Multi-Agent LLM Games Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We address both instability and underperformance in interactive games with MEMO (Memory-augmented Model context optimization), a self-play framework that treats inference-time context as an optimizable, agentic object by coupling retention and exploration. |
Yunfei Xie; Kevin Wang; Bobby Cheng; Jianzhu Yao; Zhizhou Sha; Alexander Duffy; Yihan Xi; Hongyuan Mei; Cheston Tan; Chen Wei; Pramod Viswanath; Zhangyang “Atlas” Wang; | code |
| 664 | TWLA: Breaking The Barrier to W1.58A4 Post-Training Quantization for LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing methods struggle with heavy-tailed activation distributions and therefore keep activations in high precision, fundamentally limiting end-to-end inference acceleration. To overcome this limitation, we propose **TWLA** (**T**ernarized **W**eights and **L**ow-bit **A**ctivations), a post-training quantization (PTQ) framework that achieves 1.58-bit weight compression and 4-bit activation quantization while maintaining high accuracy. |
zhixiong zhao; Zukang Xu; Zhixuan Chen; Xing Hu; Zhe jiang; Dawei Yang; | code |
| 665 | Concept-Guided Tokenization: Closing The Gap Between Reconstruction and Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While existing tokenizers excel at preserving low-level visual details through reconstruction-based training, they often lack explicit semantic guidance, which limits their ability to capture semantically structured representations and thus hinders their performance on downstream tasks like image generation. To overcome this limitation, we propose a novel tokenization framework that incorporates high-level semantics through two key innovations: (1) a text-integrated encoder that jointly processes images and textual descriptions to produce semantically enriched latent representations, and (2) a concept-guided training objective that leverages sparse autoencoders to decompose pre-trained vision-language model features to a semantic concept space, employing sparse and disentangled concept indices for guidance. |
Yunqiao Yang; Haokun Lin; Guanzhong Wu; Ying Wei; | code |
| 666 | RetrOrchestrator: A Multi-Step Retrosynthesis Agent Dynamically Orchestrating Single-Step Transition Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This limitation is particularly urgent given our observation of pronounced *skill disparity* among single-step prediction models: different models exhibit substantially different performance across molecule states. Motivated by this observation, we introduce RetrOrchestrator, an LLM-powered agent that explicitly accounts for model skill disparity by reframing retrosynthesis planning as a Partially Observable Markov Decision Process (POMDP). |
Liao Chang; Luotian Yuan; Yiping Ke; Ying Wei; | code |
| 667 | Temporal Preference Optimization for Unsupervised Retrieval Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose TPOUR (*Temporal Preference Optimization for Unsupervised Retriever*), which integrates our novel training method *Temporal Retrieval Preference Optimization* (TRPO). |
HyunJin Kim; Jaejun Shim; Young Jin Kim; JinYeong Bak; | code |
| 668 | Open Materials Generation with Inference-Time Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Open Materials Generation with Inference-time Reinforcement Learning (OMatG-IRL), a policy-gradient RL framework that operates directly on the learned velocity fields and eliminates the need for the explicit computation of the score. |
Philipp Höllmer; Stefano Martiniani; | code |
| 669 | Rethinking Depth Pruning for Vision Transformers: A Heterogeneity-Aware Perspective Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Through a comprehensive analysis of the heterogeneity, we introduce HetDPT, a method that handles heterogeneity during depth pruning while avoiding dimension mismatch. |
Zhenfeng Su; Kang Zhao; Han Bao; Tao Yuan; Zhongzhe Hu; Xianzhi Yu; Wenxuan Wang; | code |
| 670 | Identifying and Correcting Label Noise for Robust GNNs Via Influence Contradiction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, the presence of label noise in real scenarios poses a significant challenge in learning robust GNNs, and their effectiveness can be severely impacted when dealing with noisy labels on graphs, often stemming from annotation errors or inconsistencies. To address this, in this paper we propose a novel approach called ICGNN that harnesses the structure information of the graph to effectively alleviate the challenges posed by noisy labels. |
Wei Ju; Wei Zhang; Siyu Yi; Zhengyang Mao; Yifan Wang; Jingyang Yuan; Zhiping Xiao; Ziyue Qiao; Ming Zhang; | code |
| 671 | TapSampling: Inference-Time Sampling with A Task-Progress-Understanding Verifier for Robotic Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose \textbf{TapSampling}, a plug-and-play framework for inference-time sampling. |
Sizhe Zhao; Shengping Zhang; Shuo Yang; Weiyu Zhao; Shuigen Wang; Xiangyang Ji; | code |
| 672 | Learning from Comparison: Constrained Projection Policy Optimization for Pareto-Front Improvement Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose constrained projection policy optimization (CoPro), which alternates between an E-step moment projection and an M-step policy projection. |
Jintao Li; Maowen Tang; Yongji Long; Weixuan Liu; Yanlang Zheng; Sicheng He; Ao-Jin Li; Shui Yu; Yun Li; | code |
| 673 | Are Large Reasoning Models Interruptible? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we evaluate LRM robustness under two realistic dynamic scenarios: interruptions, which test the accuracy of model responses under budget-constrained outputs, and dynamic context, which tests model adaptation to in-flight changes. |
Tsung-Han Wu; Mihran Miroyan; David Chan; Trevor Darrell; Narges Norouzi; Joseph E Gonzalez; | code |
| 674 | Localize and Neutralize: Gradient-Guided Token Suppression Against Visual Prompt Injection Attack Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we show that successful adversarial attacks do not rely on the entire image uniformly but instead depend on a small subset of critical image tokens. |
Dongpeng Zhang; Ke Ma; Yangbangyan Jiang; Gaozheng Pei; Longtao Huang; Qianqian Xu; Qingming Huang; | code |
| 675 | SHAP-Guided Kernel Actor-Critic for Explainable Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose *RKHS-SHAP-based Advanced Actor-Critic (RSA2C)*, an attribution-aware, kernelized, two-timescale AC algorithm, including Actor, Value Critic, and Advantage Critic. |
Na Li; Hangguan Shan; Wei Ni; Wenjie Zhang; Xinyu Li; | code |
| 676 | Decompose, Structure, and Repair: A Neuro-Symbolic Framework for Autoformalization Via Operator Trees Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce Decompose, Structure, and Repair (DSR), a neuro-symbolic framework that restructures autoformalization into a modular pipeline. |
Xiaoyang LIU; Zineng Dong; Yifan Bai; Yantao Li; Yuntian Liu; Tao Luo; | code |
| 677 | ProConMV: Provenance-Enabled Conceptual Framework for Interpretable Multi-View Diabetic Retinopathy Diagnosis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Thus, we propose a provenance-enabled concept-based framework for multi-view DR diagnostic (ProConMV), which integrates DR lesion masks, clinical text and multi-view data, utilizing multimodal prompt analysis and visual-text concept interaction to learn the interpretable multi-source input. |
Xiaoling Luo; Shuo Yang; Qihao Xu; Chengliang Liu; Jiansong Zhang; Zhuoqin Yang; Zhihui Lai; Linlin Shen; | code |
| 678 | Failure Is Feedback: History-Aware Backtracking for Agentic Traversal in Multimodal Graphs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These limitations suggest that a retriever should adapt its reasoning to the evolving context and recover intelligently from dead ends. To address these needs, we propose Failure is Feedback (FiF), which casts subgraph retrieval as a sequential decision process and introduces two key innovations. |
Joohyung Yun; Doyup Lee; Wook-Shin Han; | code |
| 679 | Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Such multi-objective alignment setups are often plagued by the individual objectives being at odds with each other, resulting in inefficient training and little user control during inference. To address these issues, we propose a unified framework that standardizes PRM training across verifiable and non-verifiable settings for step-level supervision, performs vectorized multi-objective alignment with Multi-Action-Head DPO, and enables controllable inference via objective-specific weighting and PRM-guided decoding. |
Yiran Shen; Yu Xia; Jonathan Chang; Prithviraj Ammanabrolu; | code |
| 680 | Sample Margin-Aware Recalibration of Temperature Scaling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We further identify a fundamental flaw in NLL-based optimization: minimizing NLL can paradoxically worsen calibration. To address this, we introduce Charbonnier-Smoothed SoftECE, a smooth objective that provably upper-bounds the smooth calibration error (smCE). |
Haolan Guo; Linwei Tao; Haoyang Luo; Minjing Dong; Chang Xu; | code |
| 681 | Training Diffusion Language Models for Black-Box Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While recent work applies autoregressive LLMs to BBO by formatting tasks as natural-language prompts, their left-to-right design generation struggles to capture the strong bidirectional dependencies inherent in design problems. To address this, we propose adapting diffusion LLMs to offline BBO to leverage their bidirectional modeling capabilities. |
Zipeng Sun; Can Chen; Ye Yuan; Haolun Wu; Jiayao Gu; Christopher Pal; Xue Liu; | code |
| 682 | Learning to Rank By Directly Optimizing Full-Order Probabilities Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce the *Full-Order Bound* (FOB), a tractable lower bound on the probability of an observed ordering, constructed from a subset of ordering constraints that factorizes across items while avoiding low-dimensional surrogate objectives and preserving order-reversal invariance. |
Yongxiang Tang; Chao Wang; Jincheng Lu; Yanhua Cheng; Xialong Liu; Peng Jiang; | code |
| 683 | AD-MIR: Bridging The Gap from Perception to Persuasion in Advertising Video Understanding Via Structured Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, despite excelling at general search, existing agents often struggle to bridge the cognitive gap between pixel-level perception and high-level marketing logic. To address this challenge, we introduce **AD-MIR**, a framework designed to decode advertising intent via a two-stage architecture. |
Binxiao Xu; Junyu Feng; Xiaopeng Lin; Haodong Li; ZhiYuan Feng; Bohan Zeng; Ruichuan An; Ming Lu; Qi She; Wentao Zhang; | code |
| 684 | LSGQuant: Layer-Sensitivity Guided Quantization for One-Step Diffusion Real-World Video Super-Resolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While low-bit quantization is a common approach for model compression, the effectiveness of quantized models is challenged by the high dynamic range of input latent and diverse layer behaviors. To address these limitations, we introduce LSGQuant, a layer-sensitivity guided quantization framework for one-step diffusion-based real-world VSR. |
Tianxing Wu; Zheng Chen; Cirou Xu; Bowen Chai; Yong Guo; Yutong Liu; Linghe Kong; Yulun Zhang; | code |
| 685 | IRIS: Implicit Reward-Guided Internal Sifting for Mitigating Multimodal Hallucination Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Due to the lack of access to internal states, such feedback overlooks the fine-grained conflicts between different modalities that lead to hallucinations during generation. To address this issue, we propose IRIS (Implicit Reward-Guided Internal Sifting), which leverages continuous implicit rewards in the native log-probability space to preserve full information density and capture internal modal competition. |
Yuanshuai li; Yuping Yan; Jirui Han; Fei Ming; Lingjuan Lyu; Yaochu Jin; | code |
| 686 | Mitigating Visual Hallucinations Via Semantic Curriculum Preference Optimization in MLLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While Direct Preference Optimization (DPO) is widely used for alignment, its application to MLLMs often fails to capture fine-grained semantic differences and encourages shortcut learning. To address these challenges, we propose Semantic Curriculum Preference Optimization (SCPO), a novel framework for MLLM alignment. |
Yuanshuai li; Yuping Yan; Junfeng Tang; Zeqi Zheng; Yaochu Jin; | code |
| 687 | Unified Safe In-context Image Generation in Multimodal Diffusion Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing safety mechanisms are primarily designed for text-toimage (T2I) synthesis or U-Net-based architectures, which limits their effectiveness for unified safety mitigation in DiT-based frameworks. To bridge this gap, we propose Unified Visual Safety Regulator ( UVR) , a training- free safe generation framework that regulates unsafe semantics in generated images. |
Xiang Yang; Feifei Li; Mi Zhang; Geng Hong; Xiaoyu You; Mi Wen; Min Yang; | code |
| 688 | From Text to Forecasts: Bridging Modality Gap with Temporal Evolution Semantic Space Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Through controlled semi-synthetic experiments, we show that existing methods over-attend to redundant tokens and struggle to reliably translate textual semantics into usable numerical cues. To bridge this gap, we propose \method{}, which introduces a Temporal Evolution Semantic Space as an intermediate bottleneck between modalities. |
Lehui Li; Yuyao Wang; Jisheng Yan; Wei Zhang; Jinliang Deng; Haoliang Sun; Zhongyi Han; Yongshun Gong; | code |
| 689 | ThunderAgent: A Fast, Simple, and Program-Aware Agentic Inference System Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address the challenges, we propose \ouralg, an inference system that is aware of the end-to-end agent workflow. |
Hao Kang; Ziyang li; Xinyu Yang; Weili Xu; Yinfang Chen; Junxiong Wang; Beidi Chen; Tushar Krishna; Chenfeng Xu; Simran Arora; | code |
| 690 | Remove The Ambiguity: Few-shot Multimodal Anomaly Detection Using Crossmodal Feature Replacers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose *Crossmodal Feature Replacer (CFR)*, a self-supervised ensemble framework that resolves ambiguous crossmodal reconstructions. |
Yuan Guo; Wanqi Zhang; Xu Wang; | code |
| 691 | Taming The Recent-Data Bias: Towards Robust Time Series Forecasting with Global Context Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This bias renders forecasts highly vulnerable to perturbations in recent data, severely undermining prediction reliability. To address this issue, we propose TameR, a novel approach for robust time series forecasting that effectively mitigates recent-data bias via enhancing the utilization of global context. |
Longlong Xu; Zeyan Li; Xiao He; Zhaoyang Yu; Changhua Pei; Zhe Xie; Zijun Dou; Tieying Zhang; Dan Pei; | code |
| 692 | LocalV: Exploiting Information Locality for IP-level Verilog Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We identify three key challenges: (1) handling long, highly detailed documents, where critical interface constraints become buried in unrelated submodule descriptions; (2) generating long RTL code, where both syntactic and semantic correctness degrade sharply with increasing output length; and (3) navigating the complex debugging cycles required for functional verification through simulation and waveform analysis. To overcome these challenges, we propose \textit{LocalV}, a multi-agent framework that leverages \textit{information locality} in modular hardware design. |
Hanqi Lyu; Di Huang; Yaoyu Zhu; Kangcheng Liu; Bohan Dou; Chongxiao Li; Pengwei Jin; Shuyao Cheng; Rui Zhang; Zidong Du; Qi Guo; Xing Hu; Yunji Chen; | code |
| 693 | Manifold-Aligned Guided Integrated Gradients for Reliable Feature Attribution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While Guided Integrated Gradients reduces this sensitivity by adaptively updating low-gradient-magnitude features, input-space guidance still produces intermediate inputs that deviate from the data manifold. To address this limitation, we propose **Manifold-Aligned Guided Integrated Gradients** (MA-GIG), which constructs attribution paths in the latent space of a pre-trained variational autoencoder. |
Soyeon Kim; Seongwoo Lim; Kyowoon Lee; Jaesik Choi; | code |
| 694 | SOLAR: Self-supervised Joint Learning for Symmetric Multimodal Retrieval Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we address the critical yet underexplored challenge of symmetric multimodal-to-multimodal (MM2MM) retrieval, where queries and contexts are interchangeable. |
Wenjie Yang; Hang Yu; Yuyu Guo; Peng Di; | code |
| 695 | Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Structured Residual Reconstruction (SRR), a rank-allocation framework that preserves the top-$k$ singular subspace of the activation-scaled weight before quantization, quantizes only the residual, and uses the remaining rank $r-k$ for error reconstruction. |
Yoonjun Cho; Dongjae Jeon; Soeun Kim; Moongyu Jeon; Albert No; | code |
| 696 | TaRO: Temporal-Aware Reasoning Optimization for Video Temporal Grounding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This limitation stems from (1) inefficient random exploration in RL, and (2) reward functions that focus solely on the answer correctness while ignoring reasoning quality. To address these issues, we propose TaRO (Temporal-Aware Reasoning Optimization), a framework that explicitly enhances the model’s ability of thinking with time. |
Minghang Zheng; Zihao Yin; YI YANG; Yuxin Peng; Yang Liu; | code |
| 697 | Concept Heterogeneity-aware Representation Steering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we view representation steering through the lens of optimal transport (OT), noting that standard difference-in-means steering implicitly corresponds to the OT map between two unimodal Gaussian distributions with identical covariance, yielding a global translation. |
Laziz Abdullaev; Noelle Wong; Ryan Z.; Shiqi Jiang; Minh-Khoi Nguyen-Nhat; Tan Nguyen; | code |
| 698 | DeepSight: Long-Horizon World Modeling Via Latent States Prediction for End-to-End Autonomous Driving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a driving world model that performs parallel prediction of latent semantic features for consecutive future frames in the bird’s-eye-view (BEV) space, thereby enabling long-horizon modeling of future world states. |
Lingjun Zhang; Changjie Wu; Linzhe Shi; Jiangyang Li; Jiaxin Liu; Lei Yang; Hang Zhang; Mu Xu; Hong Wang; | code |
| 699 | Structure-Aware Riemannian Flow Matching for Registration and Fusion of Hyperspectral and Multispectral Images Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Structure-Aware Riemannian Flow Matching (SA-RFM), a geometry-informed framework for joint registration and fusion of hyperspectral and multispectral images. |
Quan Zhang; Jun Li; Weilong Zhu; MINGYANG LI; Qinmu Shen; Yuanxi Peng; | code |
| 700 | Negatives-Dominant Contrastive Learning for Generalization in Imbalanced Domains Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we begin by *theoretically* establishing the generalization bound for IDG, highlighting the role of posterior discrepancy and decision margin. |
Meng Cao; Jiexi Liu; Songcan Chen; | code |
| 701 | Beyond Accuracy and Complexity: The Effective Information Criterion for Structurally Stable Symbolic Regression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by the structural stability of real physical laws, we propose the Effective Information Criterion (EIC) to quantify formula rationality. |
Zihan Yu; Guanren Wang; Ding; Huandong Wang; Yong Li; | code |
| 702 | DuRP: Dual-Stage Physics-Embedded Learning for Joint Radiance and Polarization Restoration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These deficiencies lead to inaccurate polarimetric signatures and physical inconsistencies in scattering environments. To overcome these limitations and achieve the joint restoration of scene radiance and polarization information, we propose DuRP, a dual-stage physics-embedded learning framework. |
Zhenshuo Yang; Qian He; Zhiyuan Liu; Baojie Fan; Jiandong Tian; | code |
| 703 | 3D-DLP: Self-supervised 3D Object-centric Scene Representation Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce 3D-DLP, a self-supervised object-centric representation learning model that decomposes scene-level RGB-D or voxel observations into a set of 3D latent particles. |
Ellina Zhang; Madhavan Iyengar; Amir Zadeh; Chuan Li; David Held; Deepak Pathak; Tal Daniel; | code |
| 704 | Decouple Searching from Training: Scaling Data Mixing Via Model Merging for Large Language Model Pre-training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, identifying an optimal mixture remains an open challenge, as existing approaches either rely on unreliable tiny-scale proxy experiments or require prohibitively expensive large-scale exploration. To address this, we propose Decouple Searching from Training Mix (DeMix), a novel framework that leverages model merging to predict optimal data ratios. |
Shengrui Li; Fei zhao; Kaiyan Zhao; Jieying Ye; Haifeng Liu; Fangcheng Shi; Zheyong Xie; Yao Hu; Shaosheng Cao; | code |
| 705 | No Data? No Problem: Robust Vision-Tabular Learning with Missing Values Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This discrepancy calls for methods that remain robust to missing values at inference. To address this challenge, we propose RoVTL (Robust Vision-Tabular Learning), a framework designed to handle any level of tabular data availability, from 0% to 100%. |
Marta Hasny; Laura Daza; Keno Bressem; Maxime Di Folco; Julia Schnabel; | code |
| 706 | A Conflict-aware Evidential Framework for Reliable Sleep Stage Classification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose ConfSleepNet, a conflict-aware evidential framework that dynamically resolves inter-view conflicts. |
Yunzhi Tian; Dekui Wang; Jun Feng; Qirong Bu; Wei Zhou; Xingxing Hao; | code |
| 707 | EvoGM: Learning to Merge LLMs Via Evolutionary Generative Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Evolutionary Generative Merging (EvoGM), a framework that transcends manual heuristics by employing learnable generative modeling to optimize merging coefficients. |
Tao Jiang; Xinmeng Yu; Chenhao Yi; Yiling Wu; Yan Li; Ran Cheng; Dongmei Jiang; Jianguo Zhang; | code |
| 708 | Scalable Traffic Signal Control with Shared Policy Framework Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose SLight, a policy-aware grouped MARL-TSC framework that enables scalability and efficiency balance under dynamic and heterogeneous traffic conditions. |
Haolun MA; Yanchen ZHU; Zizhuo Xu; Weijie Shi; Jiajie Xu; Lei Li; | code |
| 709 | Clarify Before You Draw: Proactive Agents for Robust Text-to-CAD Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing fine-tuned models tend to reactively follow the user’s instructions and hallucinate dimensions when the text is ambiguous. To address this, we propose a proactive agentic framework for text-to-CadQuery generation, named as \textbf{ProCAD}, that resolves specification issues before code synthesis. |
Bo Yuan; Zelin Zhao; Petr Molodyk; Bin Hu; Yongxin Chen; | code |
| 710 | From Holo Pockets to Electron Density: GPT-style Drug Design with Density Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Compared with rigid pocket representations, experimental ED naturally captures conformational flexibility and provides a more faithful description of the binding environment. Based on this, we introduce EDMolGPT, a decoder-only autoregressive framework that generates molecules from low-resolution ED point clouds. |
Jiahao Chen; Letian Gao; Yanhaozhu; wenbiao zhou; Bing Su; Zhi Lu; Bo Huang; | code |
| 711 | EngiAgent: Fully Connected Coordination of LLM Agents for Solving Open-ended Engineering Problems with Feasible Solutions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although large language models (LLMs) have shown strong capabilities in reasoning and code generation, they often fail to ensure feasibility, which limits their applicability to engineering problem solving. To address this challenge, we propose EngiAgent, a multi-agent system with a fully connected coordinator that simulates expert workflows through specialized agents for problem analysis, modeling, verification, solving, and solution evaluation. |
Xiyuan Zhou; Ruixi Zou; Xinlei Wang; Yuheng Cheng; Yan Xu; Junhua Zhao; Jinjin Gu; | code |
| 712 | Meerkat-VL: Implicit Risk Safety Alignment in Multimodal LLMs Via Perceptual Reasoning and Self-Verification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While effective in improving models’ safety awareness, these methods still face data scarcity and reward hacking in implicit-risk scenarios, leading to insufficient risk perception and harmful responses. To address these challenges, we propose Meerkat-VL, a framework that enables models to perceive and verify implicit risks while generating safe responses. |
Peicheng Zhou; Chuanbin Liu; Shancheng Fang; Bowei Pu; Yiwei Sun; Zhangchi Hu; Hongtao Xie; | code |
| 713 | Retro-Expert: Collaborative Reasoning for Interpretable Retrosynthesis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing models rely on static pattern-matching paradigm, which limits their ability to perform effective logic decision-making, leading to a black-box process. Building on this, we propose Retro-Expert, an interpretable retrosynthesis framework that performs collaborative reasoning by combining the complementary reasoning strengths of Large Language Models and specialized models via reinforcement learning. |
Xinyi Li; Sai Wang; Yutian Lin; Yu Wu; | code |
| 714 | Provably Adaptive Linear Approximation for The Shapley Value and Beyond Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Accordingly, we introduce the first adaptive, linear-time, linear-space randomized algorithm, Adalina, that theoretically achieves improved approximation variance. |
Weida Li; Yaoliang Yu; Bryan Kian Hsiang Low; | code |
| 715 | Forget-It-All: Multi-Concept Machine Unlearning Via Concept-Aware Neuron Masking Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we take a unique perspective on multi-concept unlearning by leveraging model sparsity and propose the Forget It All (FIA) framework. |
Kaiyuan Deng; Bo Hui; Gen Li; Jie Ji; Minghai Qin; Geng Yuan; Xiaolong Ma; | code |
| 716 | Ego3S: Select, Strengthen, and Synchronize for Efficient Egocentric Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing methods tend to exhibit inertial thinking”, relying excessively on language priors and global context. To address this limitation, we propose a novel three-stage Ego3S framework to ground models’ reasoning in interaction evidence. |
Shenshen Li; Kaiyuan Deng; RuoHuai Xie; Xing Xu; Heng Tao Shen; Yazhou Yao; Fumin Shen; | code |
| 717 | Semi-Supervised Noise Adaptation: Transferring Knowledge from Noise Domain Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, a recent study observes that the noise domain constructed from simple distributions (*e.g.*, Gaussian distributions) can serve as a surrogate source domain in the semi-supervised setting, where only a small proportion of target samples are labeled while most remain unlabeled. Based on this surprising observation, we formulate a novel problem termed *Semi-Supervised Noise Adaptation* (SSNA), which aims to leverage a synthetic noise domain to improve the generalization of the target domain. |
Yuan Yao; Jin Song; Huixia Li; Tongtong Yuan; Jiaqi Wu; Yu Zhang; | code |
| 718 | FiX: Introducing Fine-grained Forget Gate Into Softmax Attention Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce *Fine-grained Forgetting Transformer* (*FiX*), a novel architecture that successfully enables element-wise forget gates in softmax attention. |
Runzhong Li; Renjie Liu; Qing Li; Bo Tang; | code |
| 719 | Dual Optimal Transport for Multi-Concept Composition: Structural Alignment and Texture Injection in Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Diffusion models have shown impressive capabilities in text-to-image synthesis, but multi-concept personalized generation remains challenging, particularly in aligning multiple reference concepts while preserving fidelity. In this work, we propose a novel framework that addresses this challenge with a two-stage Sketch-to-Rendering process, utilizing Dual Optimal Transport (OT) for structural alignment and texture injection. |
Hao Fu; Tianyu Su; Meng Liu; Chenfang Yang; Tian Gan; | code |
| 720 | RL4RLA: Teaching ML to Discover Randomized Linear Algebra Algorithms Through Curriculum Design and Graph-based Search Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present RL4RLA, a general RL framework that automates the discovery of interpretable, symbolic RLA algorithms. |
Jinglong Xiong; Xiaotian Liu; Ruoxin Wang; Zihang Liu; Yefan Zhou; Yujun Yan; Yaoqing Yang; | code |
| 721 | VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we introduce VenusBench-Mobile, a challenging online benchmark for evaluating general-purpose mobile GUI agents under realistic, user-centric conditions. |
Yichen Gong; Zhuohan Cai; Sunhao Dai; Yuqi Zhou; Zhangxuan Gu; Changhua Meng; Shuheng Shen; | code |
| 722 | ActiveScope: Actively Seeking and Correcting Perception for MLLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address these limitations, our investigation reveals two causes for failed localization: (a)\textit{Contextual Dominance}, where salient distractors overwhelm target attention, leading to inaccurate localization, and (b) \textit{Semantic Bias}, where aggregated global semantics cause the model to fixate on the most salient concept, resulting in incomplete localization under multi-object scenarios. Built on these insights, we propose {ActiveScope}, a training-free framework that enhances MLLMs by actively seeking and correcting perception. |
Yajing Wang; Chao Bi; Junshu Sun; Shufan Shen; Zhaobo Qi; Shuhui Wang; Qingming Huang; | code |
| 723 | RN-D: Discretized Categorical Actors with Regularized Networks for On-Policy Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we revisit policy representation as a first-class design choice for on-policy optimization. |
Yuexin Bian; Jie Feng; Tao Wang; Yijiang Li; Sicun Gao; Yuanyuan Shi; | code |
| 724 | Stabilizing Native Low-Rank LLM Pretraining Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While native low-rank training often suffers from instability and loss spikes, we identify uncontrolled growth in the spectral norm (largest singular value) of the weight matrix update as the dominant factor. To address this, we introduce **Spectron: Spectr**al renormalization with orthogonalizati**on**, which dynamically bounds the resultant weight updates based on the current spectral norms of the factors. |
Paul Janson; Edouard Oyallon; Eugene Belilovsky; | code |
| 725 | WaterSIC: Information-theoretically (near) Optimal Linear Layer Quantization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: It is shown that the popular GPTQ algorithm may have an arbitrarily large gap to IT limit. To alleviate this problem a novel algorithm, termed ”WaterSIC”, is proposed and is shown to be within a rate gap of 0.255 bit to IT limit, uniformly over all possible covariance matrices of input activations. |
Egor Lifar; Semyon Savkin; Or Ordentlich; Yury Polyanskiy; | code |
| 726 | SCRWKV: Ultra-Compact Structure-Calibrated Vision-RWKV for Topological Crack Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods face significant bottlenecks in balancing crack topology modeling with computational efficiency, often failing to reconcile high segmentation quality with low resource demands. To address these limitations, we propose the Ultra-Compact Structure-Calibrated Vision RWKV (SCRWKV), a network that achieves high-precision modeling via a novel Structure Field Encoder (SFE) backbone while maintaining linear complexity. |
Hanxu Zhang; Chen Jia; Hui Liu; Xu Cheng; Fan Shi; Shengyong Chen; | code |
| 727 | A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn’t) Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As a result, practitioners lack actionable guidance on selecting instructions for their target tasks. In this work, we aim to bring clarity to this landscape by disentangling and systematically analyzing the two core ingredients: data representation and selection algorithms. |
Nihal Nayak; Paula Rodriguez-Diaz; Neha Hulkund; Sara Beery; David Alvarez-Melis; | code |
| 728 | D-CORE: Incentivizing Task Decomposition in Large Reasoning Models for Complex Tool Use Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This occurs primarily due to their inadequate ability to decompose tasks when reasoning in complex tool use scenarios. To address this, we propose a two-stage training framework D-CORE ( Decomposing tasks and Composing Reasoning processes) that first incentivize the LRM’s task decomposition reasoning capability via self-distillation, followed by diversity-aware reinforcement learning (RL) to restore LRM’s reflective reasoning capability. |
Bowen Xu; Shaoyu Wu; Hao Jiang; Kai Liu; Xin Chen; lulu hu; Bin Yang; | code |
| 729 | LIFT: A Novel Framework for Enhancing Long-Context Understanding of LLMs Via Long Input Fine-Tuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introduces Long Input Fine-Tuning (LIFT), a novel framework for long-context modeling that can enhance the long-context performance of arbitrary short-context LLMs by dynamically adapting their parameters to the given long input. |
Yansheng Mao; Yufei Xu; Jiaqi Li; Fanxu Meng; Haotong Yang; Zilong Zheng; Xiyuan Wang; Muhan Zhang; | code |
| 730 | Flow for Future: Geometric SE(3)-Equivariant Flow Matching for 3D Trajectory Prediction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While flow matching offers a powerful generative paradigm, extending it to SE(3)-equivariant dynamics is challenging due to the inherent gap between deterministic history and stochastic evolving flows. To address this, we introduce GSE-Flow, an SE(3)-equivariant flow matching framework. |
junwei wu; Yihang Liu; Ruixuan Yu; Jian Sun; | code |
| 731 | PASO: Step Parallel Stochastic Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, we introduce a unified framework that recasts the autoregressive GD process as solving a system of triangular nonlinear equations (TNEs), thereby enabling \textit{step-parallel} training, where gradients for different GD steps are computed concurrently without sequential dependencies. |
Jianrong Lu; Zhuoya Gu; Haobo Li; Zhiyu Zhu; Yechao Zhang; Jianhai Chen; Minghui Yang; Junwei Liu; Jian Wang; Qinming He; Hui LIU; Junhui Hou; | code |
| 732 | Foundation VAE for CT Reconstruction, Augmentation, and Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper makes a progressive stride toward training-free medical VAEs by leveraging a critical observation: a single Foundation VAE, pretrained at scale on natural images and videos, can serve as a unified interface for CT Reconstruction, Augmentation, and Generation. |
Qi Chen; Shuhan Ding; Yu Gu; Nan Liu; Jiang Bian; Alan Yuille; Zongwei Zhou; Jingjing Fu; | code |
| 733 | Expressive Graph Neural Networks Via Equivariant Use of Noise Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Conversely, theoretically universal-expressive models often suffer from high computational costs or poor generalization, limiting their real-world applicability. To bridge this gap, we introduce Equivariant Noise GNNs (ENGNNs), a framework that utilizes random noise features to enhance the expressivity of GNNs. |
Xiyuan Wang; Muhan Zhang; | code |
| 734 | ZipMoE: Efficient On-Device MoE Serving Via Lossless Compression and Cache-Affinity Scheduling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present ZipMoE, an efficient and semantically lossless on-device MoE serving system. |
Yuchen Yang; Yaru Zhao; Pu Yang; Shaowei Wang; Zhi-Hua Zhou; | code |
| 735 | DADP: Domain Adaptive Diffusion Policy Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To tackle the challenge, we propose DADP (Domain Adaptive Diffusion Policy), which achieves robust adaptation through unsupervised disentanglement and domain-aware diffusion injection. |
Pengcheng Wang; Qinghang Liu; Haotian Lin; Yiheng Li; Guojian Zhan; Masayoshi Tomizuka; Yixiao Wang; | code |
| 736 | Contrastive Order Learning: A General Framework for Ordinal Regression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose contrastive order learning (ConOrd), a contrastive learning framework for ordinal regression that integrates the strengths of contrastive learning and order learning. |
Chaewon Lee; BeomJun Shim; Kwang Choi; Chang-Su Kim; | code |
| 737 | Stochastic Order Learning: An Approach to Rank Estimation Using Noisy Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we reformulate rank estimation with noisy ordinal labels as a stochastic ordering problem, in which each instance is inherently associated with multiple plausible ranks instead of a single deterministic label. |
Chaewon Lee; Seon-Ho Lee; Chang-Su Kim; | code |
| 738 | Seeing Symbols, Missing Structure: A Real-World Handwritten Mathematical Expression Recognition Benchmark for Large Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Most failures arise from structural mis-parsing and context-dependent symbol role confusion rather than pure visual perception errors. To mitigate this issue, we propose a training-free, schema-anchored structure-aware inference framework that decomposes recognition into schema identification, schema-constrained transcription, and context-driven disambiguation. |
Sheng Jiang; Lin Zhu; Runrui Li; Mei Wang; Qiannan Zhu; Yaoyao Zhong; Hua Huang; | code |
| 739 | ProRL: Effective Reinforcement Learning for Proactive Recommendation Via Rectified Policy Gradient Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: It has been observed that length-dependent bias causes gradients to favor path extension over deeper exploration, while weighting each step by path-level reward leads to high gradient variance. To rectify these two deficiencies, we propose an effective RL framework $\textbf{ProRL}$ with two novel mechanisms for proactive recommendation. |
Hongru Hou; Tiehua Mei; Denghui Geng; Jinhui Huang; Ao Xu; Hengrui Chen; Jiaqing Liang; Deqing Yang; | code |
| 740 | ARLArena: Demystifying Policy Gradient Stability in Agentic Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Overall, this study provides a unifying policy gradient perspective for ARL and offers practical guidance for building stable and reproducible LLM-based agent training pipelines. |
Xiaoxuan Wang; Han Zhang; Haixin Wang; Yidan Shi; Ruoyan Li; Kaiqiao Han; Chenyi Tong; Haoran Deng; Alexander Taylor; Renliang Sun; Yanqiao Zhu; Jason Cong; Yizhou Sun; Wei Wang; | code |
| 741 | Beyond Static Allocation: Dynamic Sensitivity-Aware Fine-Tuning for Vision Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing Parameter-Efficient Fine-Tuning (PEFT) methods are fundamentally constrained by a static allocation paradigm, which overlooks the model’s evolving optimization priorities during training. To address this, we introduce Dynamic Adaptive Fine-tuning (DAF), a novel framework that periodically evaluates and reconfigures the trainable structure based on a context-aware decoupled sensitivity analysis. |
Yuanyang Cao; Xichun Liu; Fuwei Zhang; Shangqi Deng; Ziyang Ren; Jianji Wang; | code |
| 742 | DOUBT: Decoupled Object-level Understanding and Bridging Via VMF-based Trustworthiness for Hallucination Detection in MLLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, because the visual module often lags behind the language module in understanding and reasoning, MLLMs can repeatedly produce similar yet incorrect answers, yielding deceptively high measured trustworthiness and therefore missed detections. To address this, we propose a simple yet effective model-agnostic method, dubbed Decoupled Object-level Understanding and Bridging via vMF-based Trustworthiness (DOUBT). |
Kaiqi Chen; Yang Qin; Changhao He; Xi Peng; Peng Hu; | code |
| 743 | LynX: Token Interface Alignment for Video+X LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This study introduces an intriguing phenomenon in Video LLMs: rather than merely translating frames into textual embeddings, Video LLMs establish a continuous manifold, token interface, allowing visual tokens to operate as standalone entities within the architecture. Exploiting this discovery, we propose LynX, a scalable framework that integrates novel modalities by repurposing the internalized interface. |
Jungin Park; Jiyoung Lee; Kwanghoon Sohn; | code |
| 744 | ExSkill: Continual Learning from Experience and Skills in Multimodal Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose \textsc{ExSkill}, a framework combining task-level Skills (structured workflows and tool templates) with action-level Experiences (context-specific tactical insights) through automated accumulation from agent trajectories. |
Guanyu Jiang; Zhaochen Su; Xiaoye Qu; Yi Fung; | code |
| 745 | VideoBrain: Learning Adaptive Frame Sampling for Long Video Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose VideoBrain, an end-to-end framework that enables VLMs to adaptively acquire visual information through learned sampling policies. |
Junbo Zou; Ziheng Huang; Shengjie Zhang; Liwen Zhang; Weining Shen; | code |
| 746 | Broadening The Backdoor Basin: Understanding LLM Backdoors Collapse and Making Backdoors Persistent Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we provide a geometric explanation: by probing the backdoor objective under controlled weight perturbations, we find that conventional poisoning often drives the backdoor loss to a narrow and sharp basin; consequently, even modest parameter drift induced by downstream SFT can push the model out of the low-loss and high-ASR region, leading to rapid backdoor forgetting. |
Xingyi Zhao; Tian Xie; Xiaojun Qi; Depeng Xu; Shuhan Yuan; | code |
| 747 | R$^3$L: Reasoning 3D Layouts from Relative Spatial Relations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose R$^3$L, a general framework that improves the reliability and consistency of relative spatial reasoning for 3D layout generation. |
Zhifeng Gu; Yuqi Wang; Bing WANG; | code |
| 748 | No Global Plan in Sight: Uncover The Myopic Planning Horizon of LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our empirical results indicate that LLMs exhibit a *myopic* horizon, primarily conducting incremental transitions without precise global planning. Leveraging this characteristic, we propose a hypothesis on enhancing uncertainty estimation of CoT, which we validate that a small subset of CoT positions can effectively represent the uncertainty of the entire path. |
Liyan Xu; Mo Yu; Fandong Meng; Jie Zhou; | code |
| 749 | MDN: Parallelizing Stepwise Momentum for Delta Linear Attention Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While momentum-based optimizers provide a natural remedy, they pose challenges in simultaneously achieving training efficiency and effectiveness. To address this, we develop a chunkwise parallel algorithm for LA with a stepwise momentum rule by geometrically reordering the update coefficients. |
Yulong Huang; Xiang Liu; Hongxiang Huang; Xiaopeng LIN; Zunchang LIU; Xiaowen Chu; Zeke Xie; Bojun Cheng; | code |
| 750 | MMKU-Bench: A Multimodal Update Benchmark for Diverse Visual Knowledge Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing research on multimodal knowledge updating focuses only on learning previously unknown knowledge, while overlooking the need to update knowledge that the model has already mastered but that later changes; moreover, evaluation is limited to the same modality, lacking a systematic analysis of cross-modal consistency. To address these issues, this paper proposes MMKU-Bench, a comprehensive evaluation benchmark for multimodal knowledge updating, which contains over 25k knowledge instances and more than 49k images, covering two scenarios, updated knowledge and unknown knowledge, thereby enabling comparative analysis of learning across different knowledge types. |
Baochen Fu; Yuntao Du; Cheng Chang; Baihao Jin; Wenzhi Deng; Muhao Xu; Hongmei Yan; Weiye Song; Yi Wan; | code |
| 751 | MotionMAR: Multi-scale Auto-Regressive Human Motion Reconstruction from Sparse Observations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Human motion inherently exhibits a sophisticated temporal hierarchical architecture, spanning from global low-frequency trajectories to local high-frequency dynamics. Inspired by this intrinsic property and the success of multi-scale autoregressive modeling in vision, we propose MotionMAR, a novel framework for human motion reconstruction from sparse observations. |
Yuhua Luo; Junsheng Zhang; Mengyin Liu; Xincheng Lin; Ming Yan; Zhudi Chen; Chenglu Wen; Lan Xu; Siqi Shen; Cheng Wang; | code |
| 752 | EcoVLA: Environment-Aware Adaptive Pruning with Interleaved Inference Orchestration for Vision-Language-Action Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Static pruning lacks the adaptability required for environment dynamics, whereas fixed-interval dynamic layer pruning suffers from coarse granularity and high retraining overheads. To bridge this gap, we propose **EcoVLA**, a training-free, plug-and-play adaptive pruning framework that supports orthogonal combination with existing VLA acceleration methods. |
Yuting Huang; Leilei Ding; Zhipeng Tang; Zenghuan Zhu; Jiajun Deng; Xinrui Lin; Shuo Liu; Haojie Ren; Jianmin Ji; Yanyong Zhang; | code |
| 753 | Revisiting The Role of Pretrained Weights in Model Merging: On Near-Optimality Within The Core Subspace Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we revisit the role of pretrained weights in model merging and investigate their efficacy from a subspace perspective. |
Wenju Sun; Qingyong Li; Tiancheng Li; Yangliao Geng; Albert Boyang Li; | code |
| 754 | FoundObj: Self-supervised Foundation Models As Rewards for Label-free 3D Object Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present FoundObj, a novel framework featuring a superpoint-based object discovery agent that incrementally merges suitable neighboring superpoints, guided by our innovative semantic and geometric reward modules. |
Zihui Zhang; Zhixuan Sun; Yafei YANG; Jinxi Li; Jiahao Chen; Bo Yang; | code |
| 755 | A Geometry-Aware Efficient Algorithm for Compositional Entropic Risk Minimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While entropic risk objectives of this form arise in many machine learning problems, existing optimization algorithms suffer from several fundamental limitations including non-convergence, numerical instability, and slow convergence rates. To address these limitations, we propose a geometry-aware stochastic algorithm, termed **SCENT**, for the dual formulation of entropic risk minimization cast as a min–min optimization problem. |
Xiyuan Wei; Linli Zhou; Bokun Wang; Chih-Jen Lin; Tianbao Yang; | code |
| 756 | CoIRL-AD: Collaborative-Competitive Imitation-Reinforcement Learning in Latent World Models for Autonomous Driving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose CoIRL-AD, a competitive dual-policy framework that integrates IL and RL under a unified offline training regime. |
Xiaoji Zheng; Ziyuan Yang; Yuhang PENG; Yanhao Chen; Yuanrong Tang; Gengyuan Liu; Bokui Chen; Jiangtao Gong; | code |
| 757 | From Pairwise Affinities to Functional Correspondences: Rethinking Attention Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce {Functional Attention}, which reinterprets attention as a functional correspondence between adaptive bases. |
Jiefang Xiao; Maolin Gao; Simon Weber; Guandao Yang; Daniel Cremers; | code |
| 758 | Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic–Procedural Memory Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce a hybrid episodic–procedural memory strategy (H-EPM) that enables experience-induced self-evolution of multi-turn tool-use policies, by adaptively reusing partially overlapping successful experiences in both inference and training. |
Sijia Li; Yuchen Huang; Zifan LIU; Zijian LI; Jingjing Fu; Lei Song; Jiang Bian; Jun Zhang; Rui Wang; | code |
| 759 | The Labyrinth and The Thread: Rethinking Regularizations in Sequential Knowledge Editing for Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we systematically investigate the mechanisms underlying effective and stable sequential editing. |
Zheng Wang; Kaixuan Zhang; Wanfang Chen; Jingwen Zhang; Xiaonan Lu; | code |
| 760 | Hista and Numca: Estimate State Value Effectively for Large Language Model Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we demonstrate that accurate value estimation can stabilize and improve post-training. |
Zizhe Chen; Jiqian Dong; Yizhou Tian; Garry YANG; Yongqiang Chen; Zhitang Chen; James Cheng; | code |
| 761 | DFSAttn: Dynamic Fine-grained Sparse Attention for Efficient Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, attention maps in DiTs exhibit inherently dynamic and fine-grained sparsity, which causes existing block sparse attention methods to degrade significantly in quality, especially at high sparsity ratios. In this paper, we revisit block sparse attention and derive a theoretical lower bound on attention recall to characterize the key factors governing its effectiveness. |
Jie Hu; Zixiang Gao; Yutong He; Kun Yuan; | code |
| 762 | DR$^2$Seg: Decomposed Two-Stage Rollouts for Efficient Reasoning Segmentation in Multimodal Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing methods typically suffer from overthinking, generating verbose reasoning chains that interfere with object localization in multimodal large language models (MLLMs). To address this issue, we propose DR$^2$Seg, a self-rewarding framework that improves both reasoning efficiency and segmentation accuracy without requiring extra thinking supervision. |
Yulin He; Wei Chen; Zhikang Jian; Tianhang Guo; Wenjuan Zhou; Minglong Li; Shaowu Yang; Wenjing Yang; | code |
| 763 | E-mem: Multi-Agent Based Episodic Context Reconstruction for LLM Agent Memory Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: By compressing fluid sequential dependencies into pre-defined structures (e.g., embeddings or graphs), these methods sever the narrative integrity essential for deep reasoning. To address this, we propose E-mem, a framework shifting from Memory Preprocessing to Episodic Context Reconstruction inspired by biological engrams. |
Kaixiang Wang; Yidan Lin; Jiong Lou; Zihan Wang; Bunyod Suvonov; Zhaojiacheng Zhou; Yuxiang Zheng; Jiaxi Cao; Zhiheng Dong; Chentao Wu; Jie Li; | code |
| 764 | OPIC: Enhancing Language Model Merging Via Optimizing In-Context Capability Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Prior mitigation strategies largely rely on validation data for costly hyperparameter tuning, limiting both interpretability and practicality. We therefore propose OPIC, an evolutionary optimization–based model merging framework. |
Jie He; Chao Chen; Weidong Bao; Zhengyi Zhong; Shuai Zhang; Ji Wang; | code |
| 765 | WestWorld: A Knowledge-Encoded Scalable Trajectory World Model for Diverse Robotic Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While recent works have explored trajectory world models for diverse robotic systems, they struggle to scale to a large number of distinct system dynamics and overlook domain knowledge of physical structures. To address these limitations, we introduce *WestWorld*, a kno**W**ledge-**E**ncoded **S**calable **T**rajectory **World** model for diverse robotic systems. |
Yuchen Wang; Jiangtao Kong; Sizhe Wei; Xiaochang Li; Haohong Lin; Hongjue Zhao; Tianyi Zhou; Lu Gan; Huajie Shao; | code |
| 766 | DOCKSMITH: Scaling Reliable Coding Environments Via An Agentic Docker Builder Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Reliable Docker-based environment construction is a dominant bottleneck for scaling execution-grounded training and evaluation of software engineering agents. We introduce DockSmith, a specialized agentic Docker builder designed to address this challenge. |
Jiaran Zhang; Lu Ma; Yanhao Li; Fanqi Wan; DI QI; Xin Wu; Zhewei Huang; Liangyu Chen; YINGWEI MA; Qi Han; Xiangyu Zhang; | code |
| 767 | Syntax Vs. Semantics: How Transformers Learn Deep Dependencies Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Large Language Models demonstrate remarkable syntactic fluency, yet the optimization dynamics governing their acquisition of deep semantic dependencies remain poorly understood. We propose a mechanistic framework that models this learning process as a competition between Surface Statistics and Deep Semantics. |
Jiangrui Zhao; Xiaoting Du; | code |
| 768 | Scalable Medical Multimodal Fusion Via Symmetric Consistency Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, most existing methods assume a fixed and known modality set, making them less effective when the number or composition of modalities changes. To address this limitation, we propose a modality-agnostic medical multimodal fusion framework that can naturally accommodate an arbitrary number of input modalities. |
Xiaowen Sun; Hui Liu; Gongguan Chen; Ning Mao; | code |
| 769 | Lightweight and Interpretable Transformer Via Unrolling of Mixed Graph Algorithms for Traffic Forecast Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We formulate a prediction problem for the future samples of signal $\mathbf{x}$, assuming it is “smooth” with respect to both $\mathcal{G}^u$ and $\mathbf{G}^d$, where we design new $\ell_2$ and $\ell_1$-norm variational terms to quantify and promote signal smoothness (low-frequency reconstruction) on a directed graph. |
Ji Qi; Mingxiao Liu; VIET THUC; Yuzhe Li; Zhuoshi Pan; Gene Cheung; Hong Zhao; | code |
| 770 | Taming I2V Models for Image HOI Editing: A Cognitive Benchmark and Agentic Self-Correcting Framework Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We thus propose SCPE (Self-Correcting Process Editing), a novel, agentic self-correcting framework that constrains the generation of I2V models through iteratively refined prompts, enabling the generated videos to more accurately present the target HOI. |
Jiayi Gao; Qingchao Chen; Yuxin Peng; Yang Liu; | code |
| 771 | MePo: Meta Post-Refinement for Rehearsal-Free General Continual Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by meta-plasticity and reconstructive memory in neuroscience, we introduce here an innovative approach named **Me**ta **Po**st-Refinement (MePo) for PTMs-based GCL. |
Guanglong Sun; Hongwei Yan; Liyuan Wang; Zhiqi KANG; Shuang Cui; Hang Su; Jun Zhu; Yi Zhong; | code |
| 772 | Activation with Intrinsic-Extrinsic Consensus Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Such irrelevant activation signals can propagate through the network and adversely affect the final decision. Inspired by observations that channel relevance can be reflected in both intrinsic activity levels and extrinsic decision weights, and that there is strong consensus between these two aspects, we propose AIEC (Activation with Intrinsic-Extrinsic Consensus), a novel activation mechanism that has the ability to identify and suppress irrelevant feature channels during training. |
Tian Qiu; Zunlei Feng; Yang Gao; Bingde Hu; Yi Gao; Mingli Song; | code |
| 773 | PSG-Nav: Probabilistic Scene Graph Navigation Via Multiverse Decision Making Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Probabilistic Scene Graph Navigation (PSG-Nav), which constructs a 3D Probabilistic Scene Graph that uses full semantic categorical distributions to account for perception uncertainty. |
Rufeng Chen; Yue Chang; Xiaqiang Tang; Hechang Chen; Sihong Xie; | code |
| 774 | StreamFlow: Theory, Algorithm, and Implementation for High-Efficiency Rectified Flow Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this article, we have comprehensively implemented an overall acceleration pipeline from the aspects of theory, design, and reasoning strategies. |
Sen Fang; Hongbin Zhong; Yalin Feng; Yanxin Zhang; Dimitris Metaxas; | code |
| 775 | Adaptive Probe-based Steering for Robust LLM Jailbreaking Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we leverage the idea of model extraction to guide the learned steering vectors to approximate the ideal one and propose tuning the steering strength adaptively based on contrastive activations’ statistics. |
Junxi Chen; Junhao Dong; Xiaohua Xie; | code |
| 776 | Distilling Neuro-Symbolic Programs Into 3D Multi-modal LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current 3D spatial reasoning methods face a fundamental trade-off: neuro-symbolic 3D (NS3D) concept learners achieve interpretable reasoning through compositional programs but are constrained to closed-set concept vocabularies and simple programs; end-to-end 3D multi-modal LLMs (3D MLLMs) could handle complex natural language and open-vocabulary concepts but suffer from black-box reasoning without explicit spatial verification. We introduce APEIRIA, a neuro-symbolic 3D MLLM that bridges these paradigms by distilling symbolic reasoning patterns into MLLMs with natural language chain-of-thought (CoT). |
Wentao Mo; Yang Liu; | code |
| 777 | Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While Representation Editing (RepE) offers a intrinsic control, its application to dynamic reasoning trajectories remains underexplored. In this work, we bridge this gap by investigating the geometry of truth within unfolding reasoning chains. |
Tianlong Wang; Yuhang Wang; Weibin Liao; Xin Gao; Xinyu Ma; Yang Lin; Yasha Wang; Liantao Ma; | code |
| 778 | Causal Disentangled Anchor Learning for Scalable Fair Multi-view Clustering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing fair multi-view clustering methods typically suffer from a severe trade-off between clustering utility and fairness, while incurring prohibitive quadratic complexity on large-scale datasets. To address these challenges, we propose Causal Disentangled Anchor Learning (CDAL), a novel framework that achieves scalable fairness via structural disentanglement. |
Suyuan Liu; Shengfei Wei; Wenjing Yang; Shengju Yu; Siwei Wang; Xueqiong Li; Wenpeng Lu; Xinwang Liu; | code |
| 779 | Target-Agnostic Calibration Under Distribution Shift with Frequency-Aware Gradient Rectification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Frequency-aware Gradient Rectification (FGR), a target-agnostic training framework for robust calibration. |
Yilin Zhang; Cai Xu; You Wu; Ziyu Guan; Wei Zhao; | code |
| 780 | Inference Time Optimization with Confidence Dynamics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we investigate the dynamics of confidence along reasoning trajectories and for first time reveal a surprising and unique pattern: correct answer traces tend to exhibit confidence improvement over time (positive confidence gain), while incorrect traces show attenuated or declining confidence as reasoning proceeds. |
Yu Wang; Minghao Liu; Jiayun Wang; Jinrui Huang; Ankit Shah; Wei Wei; | code |
| 781 | Plan in Sandbox, Navigate in Open Worlds: Learning Physics-Grounded Abstracted Experience for Embodied Navigation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose \textit{\textbf{S}andbox-\textbf{A}bstracted \textbf{G}rounded \textbf{E}xperience} (\textbf{\textit{SAGE}}), a framework that enables agents to learn within a physics-grounded semantic abstraction rather than a photorealistic simulation, mimicking the human capacity for mental simulation where plans are rehearsed in simplified physics abstractions before execution. |
Zhixuan Shen; jiawei du; Ziyu Guo; Han Luo; Lilan Peng; Joey Tianyi Zhou; Haonan Luo; Tianrui Li; | code |
| 782 | Policy-Driven World Model Adaptation for Robust Offline Model-based Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address these, we propose a framework that dynamically adapts the world model alongside the policy under a unified learning objective aimed at improving robustness. |
Jiayu Chen; Le Xu; Aravind Venugopal; Jeff Schneider; | code |
| 783 | SurrogateSHAP: Training-Free Contributor Attribution for Text-to-Image (T2I) Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose **SurrogateSHAP**, a retraining-free framework that approximates the expensive retraining game through inference from a pretrained model. |
MingYu Lu; Soham Gadgil; Chris Lin; Chanwoo Kim; Su-In Lee; | code |
| 784 | Beyond Looking Up, Try Looking Around: Harmonizing Global Structure and Local Consistency in Optimal Transport for Short Text Clustering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper proposes a novel short text clustering framework, which remedies the neglect of semantic consistency in existing OT methods, generating reliable pseudo-labels to facilitate clustering. |
Zhihao Yao; Yuxuan Gu; Jixuan Yin; Bo Li; | code |
| 785 | SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce SonicMaster, the first unified generative model for music restoration and mastering that addresses a broad spectrum of audio artifacts with text-based control. |
Jan Melechovsky; Ambuj Mehrish; Abhinaba Roy; Dorien Herremans; | code |
| 786 | Training Prompt Matters: State-Adaptive Optimization for Robust Fine-Tuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Furthermore, we discover that these superior prompts can be robustly identified by task loss prior to learning. Leveraging these insights, we introduce State-Adaptive Prompt Optimization (SAPO), a lightweight yet effective training strategy that shifts task formulation from a static input to a dynamic, state-adaptive variable. |
Shi; Yiren Chen; Shuqing Bian; Zhe Zhao; Pengfei Hu; Jinhao Dong; WEI LU; Xiaoyong Du; | code |
| 787 | Learning with Admissibility: Robust Fuzzy Hashing for Cross-Modal Retrieval with Noisy Labels Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, aggressive separation leads to reduced data utilization, while smoothing weakens the discriminative capability regarding the true distribution of clean instances. To address these limitations, we propose a novel Robust Fuzzy Cross-modal Hashing framework (RFCMH) that introduces fuzzy set theory to endow the labels with admissibility, thereby obtaining reliable discriminative supervision from noisy labels. |
Xincheng Sun; Ruitao Pu; Guangsi Shi; Zhenwen Ren; Peng Hu; Yuan Sun; | code |
| 788 | Prioritize The Process, Not Just The Outcome: Rewarding Latent Thought Trajectories Improves Reasoning in Looped Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, attempts to further improve LoopLM reasoning with reinforcement learning have failed—standard objectives such as Group Relative Policy Optimization (GRPO) only assign credit to the final latent state, creating a fundamental mismatch with the model’s internal computation. To resolve this, we introduce **RLTT (Reward Latent Thought Trajectories)**, a reinforcement learning framework which distributes reward across the full latent reasoning trajectory. |
Jonathan Williams; Esin Tureci; Olga Russakovsky; | code |
| 789 | Selective Disclosure Watermarking for Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Hierarchical Vocabulary Routing, a watermarking framework that enables selective disclosure of embedded metadata. |
Xuyang Chen; Xiang Li; Yangxinyu Xie; Qi Long; | code |
| 790 | TarGATE: Target-Aware Data Selection Via Token-Attenuation Gates Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Targeted instruction tuning requires selecting pertinent samples from massive mixed *candidate datasets* guided by a small *reference dataset* reflecting the desired capability, yet efficiently identifying high-quality data amidst noise remains challenging. To address this, we propose **TarGATE** (**Tar**get-aware **GATE**s, a simple yet effective data selection framework that leverages the model’s inherent data understanding. |
Xiandi Luo; Shiwei Li; Haozhao Wang; Yihao Ouyang; Zhuoqi Hu; Yichen Li; Xiao Yang; Hu Liu; Ruixuan Li; | code |
| 791 | Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose an automated VQA-based evaluation protocol to assess factual consistency between edited knowledge and generated images. |
Xin Gao; Cheng Yang; Chufan Shi; Taylor Berg-Kirkpatrick; | code |
| 792 | Image Restoration Via Diffusion Models with Dynamic Resolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To mitigate the computational inefficiency, this work proposes projecting data into lower-dimensional subspaces using dynamic resolution DMs to accelerate the inference process. |
Yang Zheng; Wen Li; Zhaoqiang Liu; | code |
| 793 | Can Agents Generalize to The Open World? Unveiling The Fragility of Static Training in Tool Use Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our exhaustive analysis yields a series of key insights, demonstrating that agents trained via both Supervised Fine-Tuning and Reinforcement Learning suffer from varying degrees of performance degradation when confronting open environmental shifts. |
Song-Lin Lv; Weiming Wu; Rui Zhu; Zijian Cheng; Lan-Zhe Guo; | code |
| 794 | Learning Global Representation from Queries for Vectorized HD Map Construction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose \textbf{MapGR} (\textbf{G}lobal \textbf{R}epresentation learning for HD \textbf{Map} construction), an architecture designed to learn and utilize a global representations from queries. |
Shoumeng Qiu; Xinrun Li; Yang Long; Xiangyang Xue; Varun Ojha; Jian Pu; | code |
| 795 | Exploring Nonlinear Pathway in Parameter Space for Machine Unlearning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a novel MU framework called Mode Connectivity Unlearning (MCU) that leverages mode connectivity to find an unlearning pathway in a nonlinear manner. |
Yingdan Shi; Ren Wang; | code |
| 796 | Tackling Fake Forgetting Through Uncertainty Quantification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we find that the forgetting data points misclassified by unlearning accuracy still have their ground truth labels included in the conformal prediction set from the uncertainty quantification perspective, leading to a phenomenon we term fake forgetting. |
Yingdan Shi; Sijia Liu; Kaize Ding; Ren Wang; | code |
| 797 | Motion Dynamics Learning for Few-Shot Embodied Adaptation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, we propose Motion Dynamics Mechanism (MDM), which distills latent physical regimes from trajectories via flow-matching inversion, yielding compact representations that capture dynamics. |
Sibo He; Weiying Xie; Daixun Li; Junhao Zhong; Jiayun Tian; Yunke Wang; Leyuan Fang; Gang He; Yunsong Li; | code |
| 798 | From Observations to States: Latent Time Series Forecasting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Most TSF models minimize point-wise errors on noisy and partially observed data, which encourages shortcut solutions instead of the recovery of underlying system dynamics. To address this issue, we propose Latent Time Series Forecasting (LatentTSF), a novel paradigm that shifts TSF from observation regression to latent state prediction. |
Jie Yang; Yifan Hu; Yuante Li; Kexin Zhang; Kaize Ding; Philip Yu; | code |
| 799 | GeoReward: Mitigating Contextual Variable Overestimation in Vision-Language Models for Cross-Market Preference Prediction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Vision-language models (VLMs) excel in many multimodal tasks but remain prone to a subtle yet impactful failure mode: they tend to overestimate dominant visual-textual cues while … |
Shuo Liu; Huixiang.Cai; Weiru Zhang; Xiaoyi Zeng; | code |
| 800 | The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We thus propose Decrypto, a game-based benchmark for multi-agent reasoning and ToM drawing inspiration from cognitive science, computational pragmatics and multi-agent reinforcement learning. |
Andrei Lupu; Timon Willi; Jakob Foerster; | code |
| 801 | Learning Sparse Visual Representations Via Spatial-Semantic Factorization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Conversely, generative models preserves dense feature grids for reconstruction but fails to produce high-level abstractions. We introduce STELLAR, a framework that resolves this tension by factorizing visual features into a low-rank product of semantic concepts and their spatial distributions. |
Theodore Zhao; Sid Kiblawi; Jianwei Yang; Naoto Usuyama; Reuben Tan; Noel Codella; Tristan Naumann; Hoifung Poon; Mu Wei; | code |
| 802 | MetaMoE: Diversity-Aware Proxy Selection for Privacy-Preserving Mixture-of-Experts Unification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose **MetaMoE**, a privacy-preserving framework that unifies independently trained, domain-specialized experts into a single MoE using public proxy data as surrogates for inaccessible private data. |
Weisen Jiang; Shuhao Chen; Sinno Jialin Pan; | code |
| 803 | Less Is More: Neuroscience-Motivated Probing for Efficient Concept Circuits Tracing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To fill the gaps, motivated by stimulus-efficient probing in systems neuroscience, we propose ViSAE, a compact diagnostic toolbox for interpreting the internal mechanisms of ViTs. |
Tang Li; Yanlin Chen; Mengmeng Ma; Xi Peng; | code |
| 804 | Time Series, Vision, and Language: Exploring The Limits of Alignment in Contrastive Representation Spaces Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: The Platonic Representation Hypothesis posits that learned representations from models trained on different modalities converge to a shared latent structure of the world. However, … |
Pratham Yashwante; Rose Yu; | code |
| 805 | PCRNet: Phase-aware Complex Refinement Network for EEG-based Auditory Attention Decoding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing methods typically overlook the crucial phase information of EEG signals, which limits their ability to distinguish structured neural patterns from random noise in the frequency domain and hinders robust decoding. To address these issues, this paper proposes a Phase-aware Complex Refinement Network (PCRNet) for AAD, which consists of a Temporal Context Calibration (TCC) module and a Dual-Domain Integration (DDI) module. |
Xiran Chen; Xiaoke Yang; Cunhang Fan; Jian Zhou; Zhao Lv; | code |
| 806 | E2Former-V2: On-the-Fly Equivariant Attention with Linear Activation Memory Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, mainstream architectures face critical scalability bottlenecks due to the explicit construction of geometric features or dense tensor products on \textit{every} edge. To overcome this, we introduce \textbf{E2Former-V2}, a scalable architecture that integrates algebraic sparsity with hardware-aware execution. |
Lin Huang; Chengxiang Huang; Ziang Wang; Yiyue Du; Chu Wang; Haocheng Lu; Yunyang Li; Xiaoli LIU; Arthur JIANG; Jia Zhang; | code |
| 807 | Rethinking Calibration for Early-Exit Neural Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Most methods rely on confidence thresholds for exiting, and consequently, classifier calibration is widely assumed to improve performance. In this work, we challenge this assumption and show that calibration is often not suitable for EENNs through a detailed theoretical study. |
Piotr Kubaty; Filip Szatkowski; Grzegorz Choczyński; Bartosz Wójcik; Eric Nalisnick; | code |
| 808 | HypoSpace: A Diagnostic Benchmark for Set-Valued Hypothesis Generation Under Underdetermination and Sublinear Coverage Bounds Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In such settings, effective inference requires not only producing valid explanations, but also systematically exploring and covering the admissible hypothesis set. We introduce HypoSpace, a benchmark that treats large language models (LLMs) as samplers over finite hypothesis spaces and evaluates them on three metrics: Validity, Uniqueness, and Recovery. |
Tingting Chen; Beibei Lin; Zifeng Yuan; Qiran Zou; Hongyu He; Anirudh Goyal; Yew Soon ONG; Dianbo Liu; | code |
| 809 | SIMoE: A Probabilistic Framework for Cardinality-Constrained Routing in Mixture-of-Experts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce SIMoE routing by modeling expert selection as a stochastic latent variable, casting it as probabilistic inference over discrete expert subsets under explicit cardinality constraints. |
Heng Zhao; Zilei Shao; Guy Van den Broeck; Zhe Zeng; | code |
| 810 | SemBind: Binding Diffusion Watermarks to Semantics Against Black-Box Forgery Attacks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose SemBind, the first defense framework for latent-based watermarks that resists black-box forgery by binding latent signals to image semantics via a learned semantic masker. |
Xin Zhang; Zijin Yang; Kejiang Chen; Linfeng Ma; Weiming Zhang; Nenghai Yu; | code |
| 811 | SPARD: Defending Harmful Fine-Tuning Attack Via Safety Projection with Relevance–Diversity Data Selection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose SPARD, a defense framework that integrates Safety-Projected Alternating optimization with Relevance-Diversity aware data selection. |
Shuhao Chen; Weisen Jiang; Yeqi Gong; Shengda Luo; Chengxiang Zhuo; Zang Li; James Kwok; Yu Zhang; | code |
| 812 | AutoControl Arena: Synthesizing Executable Test Environments for Frontier AI Risk Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present AUTOCONTROL ARENA, an automated framework for frontier AI risk evaluation built on the principle of logic-narrative decoupling. |
Changyi Li; Pengfei Lu; Xudong Pan; Fazl Barez; Min Yang; | code |
| 813 | CFPO : Counterfactual Policy Optimization For Multimodal Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This fundamental deficiency results in severe grounding failures, manifesting as a tendency to ignore visual evidence in favor of language priors or exhibiting hallucination drift during long chain-of-thought reasoning. To address this root cause, we propose CounterFactual Policy Optimization (CFPO), a novel framework that enforces causal consistency between visual perception and textual reasoning. |
ZhangYuan Yu; Wanran Sun; Guangjing Yang; Xiaohu Wu; Qicheng Lao; | code |
| 814 | Human-in-the-Loop Policy Optimization for Preference-Based Multi-Objective Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: On the other hand, accurately specifying preferences in advance is often unrealistic. To address these challenges, we introduce a human-in-the-loop MORL framework that interactively discovers preferred policies during optimization. |
Tianmeng Hu; Biao Luo; Ke Li; | code |
| 815 | From Denoising to De-Channeling: Integrating Physical Channel Priors Into Diffusion Models for Radio Signal Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing WSR methods typically learn directly from received signals, which are distorted by physical wireless channel effects such as fading, and current denoising diffusion models lack de-channeling capabilities, which leads to performance degradation. Therefore, we propose PWC-Diff, a novel framework that integrates prior Physical Wireless Channels into the denoising Diffusion process. |
Yaoqi Liu; Jin Wang; Chunchen Wang; Hui Wang; Chuan Shi; | code |
| 816 | FiSeR: Fine-Grained Source Representations for Cross-Domain AI Image Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Based on the structural fact that synthetic images are produced by diverse generators, we propose a hierarchical contrastive learning framework that improves the separability between natural and synthetic images while preserving generator identity information. |
Shan Zhang; Yongxin He; Mingming Zhang; Huiwen Tian; Lei Ma; | code |
| 817 | Density-Guided Continuous Flow for Robust Counterfactual Explanations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Unlike existing methods that rely on expensive ensemble intersections to define stability, we propose DensityFlow, a generative framework that constructs robust CEs by adhering to the high-confidence data manifold. |
Jun Tan; Qing Guo; Zicheng Xu; Jinglin Li; QI Fang; Ning Gui; | code |
| 818 | MotionGRPO: Overcoming Low Intra-Group Diversity in GRPO-Based Egocentric Motion Recovery Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose MotionGRPO, a novel framework leveraging reinforcement learning post-training to inject fine-grained guidance into the diffusion process. |
Nanjie Yao; Junlong Ren; Wenhao Shen; Hao Wang; | code |
| 819 | Learning Protein Structure-Function Relationships Through Knowledge-guided Representation Decomposition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Here we introduce ProtDiS, a knowledge-guided framework that decomposes pretrained protein micro-environment embeddings into biologically grounded and task-relevant dimensions. |
Mingqing Wang; Zhiwei Nie; ATHANASIOS VASILAKOS; Yonghong He; Zhixiang Ren; | code |
| 820 | LoRA-DA: Data-Aware Initialization for Low-Rank Adaptation Via Asymptotic Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we establish a theoretical framework for data-aware LoRA initialization. |
Qingyue Zhang; Chang Chu; Tianren Peng; Qi Li; Xiangyang Luo; Zhihao Jiang; Shao-Lun Huang; | code |
| 821 | Large Language Models Explore By Latent Distilling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Exploratory Sampling (ES), a decoding approach that explicitly encourages semantic diversity during generation. |
Yuanhao Zeng; Ao Lu; Lufei Li; Zheng Zhang; Yexin Li; Kan Ren; | code |
| 822 | Zero-Shot Rankability: Revealing Latent Ordinal Structure in Multimodal Large Language Models Via Language Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we study the embeddings of Multimodal LLMs (MLLMs). |
Nam Hyeon-Woo; Moon Ye-Bin; Sohwi Lim; Kwon Byung-Ki; Tae-Hyun Oh; | code |
| 823 | A Language-Guided Bayesian Optimization for Efficient LoRA Hyperparameter Search Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, LoRA is highly sensitive to hyperparameter choices, and performing an exhaustive hyperparameter search remains computationally intensive. To address these challenges, we propose a framework that integrates the domain knowledge of pre-trained LLMs into the Bayesian Optimization (BO) process to efficiently search for LoRA hyperparameters. |
Baek Seong-Eun; Lee Jung-Mok; Kim Sung-Bin; Tae-Hyun Oh; | code |
| 824 | SCOUT: Active Information Foraging for Long-Text Understanding with Decoupled Epistemic States Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We identify a key insight for LTU: query-relevant information is typically sparse relative to the full document, so effective reasoning should rely on a query-sufficient subset rather than the entire context. To address this, we propose SCOUT, a new paradigm for LTU that **shifts from passive processing to active information foraging**. |
Zhenliang Zhang; Wenqing Wang; Yong Hu; Yaming Yang; Jiaheng Gao; Chen Shen; Xiaojun Wan; | code |
| 825 | VIPO: Value Function Inconsistency Penalized Offline Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce VIPO, a novel model-based offline RL algorithm that incorporates self-supervised feedback from value estimation to enhance model training. |
Xuyang Chen; Keyu Yan; Guojian Wang; Lin Zhao; | code |
| 826 | Budget-Efficient Attacks and Robustness Training for Cooperative MARL Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Budgeted Hierarchical Efficient Attack (BHEA), a budgeted hierarchical adversarial attack that separates decisions on when and which agents to hijack from action replacement, enabling more precise vulnerability discovery under limited attack opportunities. |
junyong jiang; Xin Yuan; Longhe Lin; Songze Li; Lu Dong; | code |
| 827 | Radial Scaling Voxelization for Accurate Small Object 3D Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Radial Scaling Voxelization (RSV), a simple yet effective non-uniform discretization strategy that adaptively modulates the effective voxel size based on the radial distance from the LiDAR sensor. |
Hao Liu; Yi Zhou; Yanni Ma; | code |
| 828 | Alleviating Observation Bias Via Causal-Invariant Meta-Learning for Unbalanced Incomplete Multi-view Clustering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods either naively assume low-missing-rate views as high-quality or lack effective debiasing mechanisms, showing limited performance under imbalance. To address this, we propose the Causal-Invariant Meta-Learning Network (CIMLN). |
Jiaqi Jin; Siwei Wang; Taichun Zhou; Dong Zhibin; Siqi Wang; Miaomiao Li; Xinwang Liu; En Zhu; | code |
| 829 | Ranking Free RAG: Replacing Re-ranking with Selection in RAG for Sensitive Domains Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Retrieval-Augmented Generation (RAG) systems deployed in sensitive domains must provide interpretable evidence selection and robust safeguards against data poisoning, yet current approaches rely on opaque similarity-based retrieval with arbitrary top-k cutoffs that offer no explanation for their selections and remain vulnerable to adversarial manipulation. We propose METEORA, a rationale-driven RAG framework that addresses these fundamental limitations through interpretable, adaptive evidence retrieval. |
Yash Saxena; Ankur Padia; Mandar Chaudhary; Kalpa Gunaratna; Srinivasan Parthasarathy; Manas Gaur; | code |
| 830 | Corrected Samplers for Discrete Flow Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Discrete flow models (DFMs) have been proposed to learn the data distribution on finite state space, offering a flexible framework as an alternative to discrete diffusion models. |
Zhengyan Wan; Yidong Ouyang; Liyan Xie; Hongyuan Zha; Fang Fang; Guang Cheng; | code |
| 831 | Modality-Decoupled Online Recursive Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address these, we propose **M-ORE**, a modality-decoupled online recursive editor for lifelong MLLM adaptation. |
Siyuan Li; Youyuan Zhang; Fangming Liu; Jing Li; | code |
| 832 | Logit-Attention Divergence: Mitigating Position Bias in Multi-Image Retrieval Via Attention-Guided Calibration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This observation reveals a fundamental limitation of existing logit-level calibration methods such as PriDe. Based on this insight, we propose a training-free, attention-guided debiasing framework that leverages intrinsic attention signals for instance-level correction at inference time, requiring only a minimal calibration set with negligible computational overhead. |
Mingtao Xian; Yifeng Yang; Qinying Gu; Xinbing Wang; Nanyang Ye; | code |
| 833 | Particles Don’t Care About Z: Towards Scaling Entropy Estimation of Unnormalized Densities Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These issues severely limit both the correctness and the scalability of the approach. We propose _MET-SVGD_, a principled extension of _P-SVGD_ that addresses these flaws by providing a general framework for \textit{SVGD} hyperparameters selection with global invertibility and convergence guarantees. |
Safa Messaoud; Skander Charni; Elaa Bouazza; Ali Fatideh; Halima Bensmail; | code |
| 834 | Benchmarking The Scientific Mind: Toward Evaluation of Complex-Reasoning Biomedical VQA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Such formulations fail to capture the evidence-constrained, multi-step nature of biomedical reasoning, and obscure whether models can derive conclusions through causal interpretation of experimental observations. To address these critical gaps in reasoning evaluation, we propose a principled benchmark construction framework that reconstructs scientific reasoning paths directly from biomedical literature. |
Ziyu Zhao; Yiyang Liu; Yajiao Wang; Xiaotao Wang; Yang Li; Yuyang Peng; Jiaheng Zhou; Jinqiao Wang; Yingying Chen; Ge Yang; Haixin Wang; | code |
| 835 | Parametric Prior Mapping Framework for Non-stationary Probabilistic Time Series Forecasting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Parametric Prior Mapping (PPM), a framework that injects parametric structural priors into a generative modeling process. |
Jinglin Li; Jun Tan; QI Fang; Ning Gui; | code |
| 836 | CARE: Class-Adaptive Expert Consensus for Reliable Learning with Long-Tailed Noisy Labels Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods partially address these issues but typically ignore the non-uniform impact of label noise across classes, resulting in ineffective correction for tail classes and over-regularization for head classes. To address this issue, we propose Class-Adaptive Rectification with Experts (CARE), a parameter-efficient framework that leverages three complementary supervision sources from vision-language models (VLM): observed noisy labels, VLM text embeddings, and visual features. |
Mengke Li; Haiquan Ling; Lihao Chen; Yang Lu; Yiqun Zhang; Hui Huang; | code |
| 837 | Frequency-Aware Perceptual Optimization for Low-Complexity Implicit Image Compression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a frequency-aware perceptual optimization framework for low-complexity image compression, realized as a **Re**alism-enhanced **Re**gion-based **I**mplicit **C**odec (Re2IC). |
Haotian Wu; Li; Di You; Pier Luigi Dragotti; Deniz Gunduz; | code |
| 838 | Lottery Prior: Randomized Neural Compression for Zero-Shot Inverse Problems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study zero-shot inverse problems, where a clean signal is recovered from a single degraded observation without external training data. |
Haotian Wu; Di You; Pier Luigi Dragotti; Deniz Gunduz; | code |
| 839 | Estimating The Empowerment of Language Model Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To handle the unique challenges of text-based environments, we introduce EELMA (Estimating Empowerment of Language Model Agents), an algorithm for approximating effective empowerment from multi-turn text interactions. |
Jinyeop Song; Jeff Gore; Max Kleiman-Weiner; | code |
| 840 | Expanding The Chaos: Neural Operator for Stochastic (Partial) Differential Equations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we build on Wiener–chaos expansions (WCE) to design neural operator (NO) architectures for SDEs and SPDEs: we project driving noise paths onto orthonormal Wick–Hermite features and use NOs to parameterize the resulting chaos coefficients, enabling reconstruction of full trajectories from noise in a single forward pass. |
Dai Shi; Lequan Lin; Andi Han; Luke Thompson; Jose Miguel Hernandez-Lobato; Zhiyong Wang; Junbin Gao; | code |
| 841 | HEDP: A Hybrid Energy-Distance Prompt-based Framework for Domain Incremental Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, domain shifts often cause severe performance degradation. To address this, we propose Hybrid Energy-Distance Prompt, a domain-incremental framework inspired by Helmholtz free energy. |
Yu Feng; Zhen Tian; Haoran Luo; Xie Yu; Diancheng Cheng; Haoyue Zheng; Shuai Lyu; Ping Zong; Lianyuan Li; xin ge; Yifan Zhu; | code |
| 842 | Models Under SCOPE: Scalable and Controllable Routing Via Pre-hoc Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose SCOPE (Scalable and Controllable Outcome Performance Estimator), a routing framework that goes beyond model selection by predicting their cost and performance. |
Qi Cao; Shuhao Zhang; Ruizhe Zhou; Ruiyi Zhang; Peijia Qin; Pengtao Xie; | code |
| 843 | Dissect and Prune: Enhancing Robustness in AI-Generated Image Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We hypothesize that this stems from the model’s reliance on spurious features, distracting signals that obscure true generative artifacts. To address this, we propose DEAR (Dissect and Prune), which leverages inpainted images to identify and prune these interfering components. |
Dahye Kim; Jaehyun Choi; Hyun Seok Seong; Seongho Kim; Donghun Lee; Sungwon Yi; Jang-Ho Choi; | code |
| 844 | Guidance: Sentence-Level Citation Enforcement Via Prefix-Tail Guidance During LLM Decoding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing work attempts to reinforce the use of citations through Retrieval-Augmented Generation (RAG) or post-hoc methods, while citations remain a probabilistic output rather than a foundation for the generated content. To address this, we propose Guidance, which aims to correct outputs and naturally incorporate citations during the LLM decoding phase. |
Yirui Zhan; Xu; Jun Gao; | code |
| 845 | Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Audio large language models (Audio LLMs) demonstrate strong performance on speech understanding tasks, yet their ability to understand paralinguistic information remains limited. To systematically quantify this issue, we introduce VoxParadox, an adversarial benchmark with 2,000 verified examples, spanning 10 paralinguistic tasks, created with controlled speech synthesis to intentionally mismatch transcript claims and speaking style, enabling direct measurement of speech paralinguistic understanding. |
Jiacheng Pang; Ashutosh Chaubey; Mohammad Soleymani; | code |
| 846 | Q-Sched: Pushing The Boundaries of Few-Step Diffusion Models with Quantization-Aware Scheduling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Q-Sched, a scheduler-level PTQ approach that adapts the diffusion sampler while keeping the quantized weights fixed. |
Natalia Frumkin; Diana Marculescu; | code |
| 847 | Chain-of-Goals Hierarchical Policy for Long-Horizon Offline Goal-Conditioned RL Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While hierarchical approaches mitigate this issue by decomposing tasks, most existing methods rely on separate high- and low-level networks and generate only a single intermediate subgoal, making them inadequate for complex tasks that require coordinating multiple intermediate decisions. To address this limitation, we draw inspiration from the chain-of-thought paradigm and propose the Chain-of-Goals Hierarchical Policy (CoGHP), a novel framework that reformulates hierarchical decision-making as autoregressive sequence modeling within a unified architecture. |
Jinwoo Choi; Sang-Hyun Lee; Seung-Woo Seo; | code |
| 848 | Anti-Backdoor Coreset Selection Via Cumulative Entropy Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent training-time defenses against neural backdoors isolate a benign subset from poisoned training data, to learn a backdoor-free model from it. In this paper, we formulate this defense strategy as a coreset selection problem, giving rise to so-called “Anti-Backdoor Coreset Selection.” |
Qi Zhao; Christian Wressnegger; | code |
| 849 | Unleashing The Representational Power of Fourier Shapes for Attacking Infrared Object Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current shape-based methods suffer from a fundamental trade-off between representational capability and optimization power, limiting their attack effectiveness. In this work, we overcome this dilemma by introducing learnable Fourier shapes to the infrared domain. |
Yixing Yong; Jian Wang; Ming Lei; Lijun He; Fan Li; | code |
| 850 | RAD: Retrieval High-quality Demonstrations to Enhance Decision-making Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Prior solutions based on synthetic data augmentation often fail to generalize to unseen scenarios in the (augmented) dataset. To address these challenges, we propose Retrieval High-quAlity Demonstrations (RAD) for decision-making, which innovatively introduces a retrieval mechanism into offline RL. |
Lu Guo; Yixiang Shan; Zhengbang Zhu; Qifan Liang; Lichang Song; Ting Long; Weinan Zhang; Yi Chang; | code |
| 851 | PromptPilot: Game-Theoretic Multi-Agent Prompt Optimization for Segment Anything Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Optimizing prompts for foundation models like SAM represents a challenging high-dimensional black-box optimization problem, fundamentally plagued by the credit assignment ambiguity. To address this, we introduce PromptPilot, a task-agnostic reinforcement learning framework that structurally decomposes the search space into orthogonal semantic and spatial subspaces. |
Guangze Shi; Yingjie Mi; Jia Shen; Feixue Shao; Jiarui Cao; Yexin Lai; Xueyu Liu; Rui Wang; Yongfei Wu; Mingqiang Wei; | code |
| 852 | MARS-SQL: A Multi-Agent Reinforcement Learning Framework For Text-To-SQL Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While current methods rely heavily on static prompting, they lack the ability to dynamically adapt and self-correct through environmental interaction. To bridge this gap, we propose **MARS-SQL**, a multi-agent architecture that leverages interactive Reinforcement Learning (RL) to optimize SQL generation. |
Haolin Yang; Jipeng Zhang; Zhitao He; Alexander Zhou; Yi Fung; | code |
| 853 | The Devil Is in The Spectrum: Mitigating Representation Collapse in LLMs Via Topologically Regularized Side-Path Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Through spectral analysis of attention dynamics, we derive an intrinsic trade-off between Mixing Efficiency (spectral gap) and Information Capacity (effective rank), revealing that standard mechanisms struggle to maximize both simultaneously. To resolve this dilemma, we propose the Topologically Regularized Side-Path (TRSP), a non-invasive architectural intervention designed to achieve spectral balance. |
Yiheng Tao; Kaiwen Cheng; Yao Lu; Chang Liu; Jie Chen; | code |
| 854 | CONTEXTOR: Contextualized High-order Contrastive Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Contextualized High-order Contrastive Learning (CONTEXTOR), a general and plug-and-play framework that formulates high-order relation inference as a dynamic query–response process. |
Ze Cai; Hanzhe Liang; Sihang Zeng; Binbin Zhou; Jun Wen; | code |
| 855 | Weak Diffusion Priors Can Still Achieve Strong Inverse-Problem Performance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our theory, based on Bayesian consistency, gives conditions under which high-dimensional measurements make the posterior concentrate near the true signal. |
Jing Jia; Wei Yuan; Sifan Liu; Liyue Shen; Guanyang Wang; | code |
| 856 | MADA-Attack: Transferable Multi-modal Attention Distraction Adversarial Attack Against Vision Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing UAP methods mainly operate on the visual modality, overlooking structured textual semantics and cross-modal interactions, which limits their ability to disrupt alignment and generalize across tasks and model architectures. To address these limits, we propose **Multi-modal Attention Distraction Adversarial Attack (MADA-Attack)** framework. |
Zhihan Qin; Jiahao Chen; Chunyi Zhou; Yuwen Pu; Chunqiang Hu; Xiaolei Liu; Shouling Ji; | code |
| 857 | Learning Locally, Revising Globally: Global Reviser for Federated Learning with Noisy Labels Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Unlike traditional noisy labels, the F-LN problem is exacerbated by the inherent heterogeneity of FL, where clients experience varying levels and types of label errors. In this study, we observe that the global model of FL exhibits slow memorization of noisy labels, suggesting its ability to maintain reliable predictions and robust representations in FL. |
Yuxin Tian; Mouxing Yang; Yuhao Zhou; Jian Wang; Qing Ye; Tongliang Liu; Gang Niu; Jiancheng Lv; | code |
| 858 | Controlled Collaboration Geometry for Personalized Federated Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: An upper bound on disagreement further reveals two degeneration mechanisms—overly strong collaboration drives consensus (reducing to standard federated learning), while similarity-driven weight updates make the graph nearly reducible and induce self-clustering (collapsing to clustered PFL). Motivated by these findings, we propose pFedCCG. |
Hongbo Yin; Wu Jichun; Zhou Yang; Chi Jiang; Yin Zhang; Yan Zhang; | code |
| 859 | AREA: Attribute Extraction and Aggregation for CLIP-Based Class-Incremental Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, since only data from the current task is available, incremental updates can bias both attribute extraction and aggregation toward new classes, leading to catastrophic forgetting. Therefore, we propose AREA for attribute extraction and aggregation for CLIP-based CIL. |
Zhenhao Wen; Yu-Cheng Shi; Da-Wei Zhou; | code |
| 860 | SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Such failure reflects two problems: router drift, where expert selection becomes inconsistent over time, and expert drift, where shared experts are overwritten across tasks. Therefore, we propose StAbilized Mixture-of-Experts (SAME) for MCIT. |
Zhenhao Wen; Jun-Tao Tang; Yu-Cheng Shi; Han-Jia Ye; De-Chuan Zhan; Da-Wei Zhou; | code |
| 861 | Controllable Molecule Generation Via Sparse Representation Editing: An Interpretability-Driven Perspective Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While large language models (LLMs) show great promise, their dense and entangled representations impede precise control over the generation of molecules with bespoke substructures or properties. To address this, we propose Sparse Representation Editing (SpaRE), an interpretability-driven framework for fine-grained and precise control in LLM-based molecule generation. |
Zhuoran Li; Xu Sun; Wanyu LIN; Chang Chen; | code |
| 862 | When Diffusion Language Models Hesitate: Detecting and Correcting Visual Hallucinations Via Confidence Fluctuation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose VGR (Visual-Guided Refinement), a framework that enables MDLMs to revisit visual details by exploiting diffusion dynamics. |
Wenzheng Song; Pei Chen; Yichen Tan; Zejian Li; Lingyun Sun; | code |
| 863 | LangPrecip: Language-Aware Multimodal Precipitation Nowcasting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose LangPrecip, the first language-guided precipitation nowcasting framework, and contribute LangPrecip-160K, a large-scale radar-text paired dataset with 160K annotated sequences. |
Xudong Ling; Lichaorong; Huang Tianxi; Qian Dong; Guiduo Duan; | code |
| 864 | DualOptim+: Bridging Shared and Decoupled Optimizer States for Better Machine Unlearning in Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose **DualOptim+**, a novel optimization framework for improving machine unlearning in large language models. |
Xuyang Zhong; Qizhang Li; Yiwen Guo; Chen Liu; | code |
| 865 | KAST-BAR: Knowledge-Anchored Semantically-Dynamic Topology Brain Autoregressive Modeling for Universal Neural Interpretation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While EEG foundation models have shown significant potential in universal neural decoding across tasks, their advancement remains constrained by the inadequacy modeling of *complex spatiotemporal topology*, as well as the inherent *modality gap* between low-level physiological signals and high-level textual semantics. To address these challenges, we propose a **K**nowledge-**A**nchored **S**emantically-Dynamic **T**opology **B**rain **A**uto**r**egressive Model (KAST-BAR), which dynamically aligns physiological representations derived from multi-level brain topology with an expert-level semantic space. |
Haoning Wang; Wenchao Yang; Shuai Shen; Yang Li; | code |
| 866 | OpenMAG: A Comprehensive Benchmark for Multimodal-Attributed Graph Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although existing benchmarks have facilitated initial progress, they exhibit critical limitations in *domain coverage*, *encoder flexibility*, *model diversity*, and *task scope*, presenting significant challenges to fair evaluation. To bridge this gap, we present OpenMAG, a comprehensive benchmark that integrates 19 datasets across 6 domains and incorporates 16 encoders to support both static and trainable feature encoding. |
Chenxi Wan; Xunkai Li; Yilong Zuo; Haokun Deng; Sihan Li; Bowen Fan; Hongchao Qin; Rong-Hua Li; Guoren Wang; | code |
| 867 | From Reward-Free Representations to Preferences: Rethinking Offline Preference-Based Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We revisit offline PbRL through the lens of reward-free representation learning (RFRL) from the zero-shot RL literature, and propose a new training framework that first learns latent successor-measure representations from reward-free offline data, followed by contrastive search and fine-tuning using preference data. |
Jun-Jie Yang; Chia-Heng Hsu; Kui-Yuan Chen; Ping-Chun Hsieh; | code |
| 868 | Softplus Attention with Re-weighting Boosts Length Extrapolation in Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, traditional Softmax attention suffers from numerical instability and reduced performance as the number of inference tokens increases. This work addresses these issues by proposing a new design principle for attention, viewing it as a two-stage process. |
Bo Gao; Michael Spratling; Letizia Gionfrida; | code |
| 869 | GUDA: Counterfactual Group-wise Training Data Attribution for Diffusion Models Via Unlearning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose GUDA (Group Unlearning-based Data Attribution) for diffusion models, which approximates each counterfactual model by applying machine unlearning to a shared full-data model instead of training from scratch. |
Naoki Murata; Yuhta Takida; Chieh-Hsin Lai; Toshimitsu Uesaka; Bac Nguyen; Stefano Ermon; Yuki Mitsufuji; | code |
| 870 | Enhancing Protein-Protein Interaction Prediction with Hierarchical Motif-based Multimodal Protein Embedding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Second, despite the availability of complementary information across the sequence, structure, and function modalities, current PPI methods fail to integrate all three modalities effectively. To address these limitations, we propose a Hierarchical Motif-based M ultiM odal protein Encoder for PPI Prediction (MMM-PPI), which constructs protein embeddings for PPI prediction in a bottom-up, multi-modal manner. |
Zaifei YANG; Samuel Choi; James Kwok; | code |
| 871 | The Quality-Utility Paradox: Why High-Reward Data Impairs Small Model Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we identify a counter-intuitive \textbf{Quality-Utility Paradox} across diverse model families(Qwen2.5, LLaMA-3, DeepSeek): data refined by a superior Synthesis Oracle consistently underperforms the SLM’s self-generated Rejection Sampling (RFT) data, despite achieving higher reward scores. |
Haolong Qian; Xianliang Yang; Ma yinuo; Lirong Che; Feng Lu; Ye Guo; Lei Song; Jiang Bian; Chun Yuan; | code |
| 872 | Weakly Supervised Cross-Modal Learning for 4D Radar Scene Flow Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, self-supervised approaches often yield suboptimal results due to radar’s inherently low-fidelity measurements, while existing cross-modal supervised methods introduce complex multi-task architecture and require costly LiDAR sensors to generate pseudo radar scene flow labels from pretrained 3D tracking models. To overcome these limitations, we propose a task-specific iterative framework for weakly supervised radar scene flow learning, using only images and odometry for auxiliary supervision during training. |
Jingyun Fu; Zhiyu Xiang; Na Zhao; | code |
| 873 | Do You Want to Know If Two Distributions Are Close to Each Other?Testing The Closeness With Statistical Significance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To mitigate the issue, we design a new measurement of distributional discrepancy, norm-adaptive MMD (NAMMD), which scales MMD’s value using the RKHS norms of distributions. |
Zhijian Zhou; Liuhua Peng; Xunye Tian; Mingming Gong; Feng Liu; | code |
| 874 | V1: Unifying Generation and Self-Verification for Parallel Reasoners Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While existing approaches typically evaluate candidates independently via scalar scoring, we demonstrate that models are substantially stronger at **pairwise self-verification**. Leveraging this insight, we introduce **V1**, a framework that unifies generation and verification through efficient pairwise ranking. |
Harman Singh; Xiuyu Li; Kusha Sareen; Monishwaran Maheswaran; Sijun Tan; Xiaoxia (Shirley) Wu; Junxiong Wang; Alpay Ariyak; Qingyang Wu; Samir Khaki; Rishabh Tiwari; Long (Tony) Lian; Yucheng Lu; Boyi Li; Alane Suhr; Ben Athiwaratkun; Kurt Keutzer; | code |
| 875 | Minimizing Upper Confidence Bounds: A Data-Driven Framework for Stochastic Programming Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our central contribution is the Average Percentile Upper Bound (APUB), a new statistical construct that serves as both a statistically rigorous upper bound for population means and an approximate risk metric for sample means. |
Shixin Liu; Ming Gao; Jian Hu; | code |
| 876 | EEG-Based Multimodal Learning Via Hyperbolic Mixture-of-Curvature Experts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose EEG-MoCE, a novel hyperbolic mixture-of-curvature experts framework designed for multimodal neurotechnology. |
Runhe Zhou; Shanglin Li; Guanxiang Huang; Xinliang Zhou; Qibin Zhao; Motoaki Kawanabe; Yi Ding; Cuntai Guan; | code |
| 877 | From Knowledge to Inference: Formalizing Specialized Public Health Reasoning on GlobalHealthAtlas Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce GlobalHealthAtlas, a large-scale multilingual dataset of 280,210 instances spanning 15 public health domains and 17 languages, stratified into three difficulty levels from health literacy to epidemiological and policy reasoning. |
Zhaokun Yan; Shan Xu; Wuzheng Dong; Zhaohan Liu; Lijie Feng; CHENGXIAO DAI; Chen Tianqi; Yingting Li; Yi Zhang; Yunpu Ma; Binfan Liu; Wenting Wei; Tongning Wu; | code |
| 878 | Krause Synchronization Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce $\textbf{Krause Attention}$, a principled attention mechanism inspired by bounded-confidence consensus dynamics. |
Jingkun Liu; Yisong Yue; Max Welling; Yue Song; | code |
| 879 | Rethinking Efficient Graph Coarsening Via A Non-Selfishness Principle Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This \textit{selfishness} matching paradigm incurs substantial computational and memory overhead. To address this problem, we shift to a \textit{non-selfishness} principle that prioritizes the collective interference of neighborhood in coarsening, and propose an efficient method named \texttt{NOPE}, which achieves linear memory consumption and near-linear computational complexity in the number of nodes. |
Xu Bai; Bin Lu; kunzhang; Shengbo Chen; Xinbing Wang; Chenghu Zhou; Meng Jin; | code |
| 880 | MCCE: A Framework for Multi-LLM Collaborative Search in Discrete Spaces with Similarity-Filtered Preference Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Multi-LLM Collaborative Co-evolution (MCCE), a hybrid framework that unites a frozen closed-source LLM with a lightweight trainable model. |
Nian Ran; Zhongzheng Li; Yue Wang; Qingsong Ran; Xiaoyuan Zhang; Shikun Feng; Richard Allmendinger; Xiaoguang Zhao; | code |
| 881 | FlowSeg: Dynamic Semantic Guidance for LLM-Conditioned Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Through systematic analysis, we show that many errors stem from semantic misalignment rather than poor mask quality. To address this issue, we propose FlowSeg, which introduces dynamic semantic guidance via a bidirectional semantic flow between intermediate decoding states and LLM-derived condition embeddings throughout the generation process. |
Zekang Zhang; Guangyu Gao; YouyunTang; WU CHENGJING; Chi Liu; Xiaochao Qu; Luoqi Liu; Ting Liu; Jianbo Jiao; Yunchao Wei; | code |
| 882 | Semantic-level Backdoor Attack Against Text-to-Image Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose Semantic-level Backdoor Attack (SemBD), which implants backdoors at the representation level by defining triggers as continuous semantic regions rather than discrete textual patterns. |
Tianxin Chen; Wenbo Jiang; Hongqiao Chen; Zhirun Zheng; Cheng Huang; | code |
| 883 | Rethinking Loss Reweighting for Imbalance Learning As An Inverse Problem: A Neural Collapse Point of View Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by Neural Collapse (NC), the ideal simplex Equiangular Tight Frame (ETF) terminal geometry suggests equal per-class average loss as a reasonable target for reweighting. Based on the ideal equal loss objective, we consider loss reweighting as an inverse problem and propose an inverse-view reweighting strategy that infers class weights dynamically to match this ideal objective. |
Jinping Wang; Zixin Tong; Zhiwu Xie; Zhiqiang Gao; | code |
| 884 | Fix Before Search: Benchmarking Agentic Visual Query Pre-processing in Multimodal Retrieval-augmented Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, real-world visual queries are often “imperfect”—suffering from geometric distortions, quality degradation, or semantic ambiguity—leading to catastrophic retrieval failures. To address this gap, we propose V-QPP-Bench, the first comprehensive benchmark dedicated to Visual Query Pre-processing (V-QPP). |
Jiankun Zhang; Shenglai Zeng; Kai Guo; Xinnan Dai; Hui Liu; Jiliang Tang; Yi Chang; | code |
| 885 | Markov Chain Monte Carlo Without Evaluating The Target: An Auxiliary Variable Approach Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we begin by observing that seemingly different Markov chain Monte Carlo (MCMC) algorithms, such as the exchange algorithm, PoissonMH, and TunaMH, can be unified under a simple common procedure. |
Wei Yuan; Guanyang Wang; | code |
| 886 | Beyond Euclidean Summaries: Online Change Point Detection for Distribution-Valued Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a geometry-aware CPD framework that treats streaming batch data as a stochastic process on the 2-Wasserstein space. |
Yingyan Zeng; Zipan Huang; Xiaoyu Chen; | code |
| 887 | Advancing SVD-based LLM Compression Via Layer-Wise Error Model Search Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Low-rank SVD-based compression offers a powerful strategy to reduce the computational costs of Large language models (LLMs); however, existing methods commonly encounter two recurring obstacles: (i) global rank allocation, where uncalibrated error proxies fail to account for complex error propagation, and (ii) decomposition quality, where Fisher-based estimators suffer from severe rank collapse. In this work, we address these limitations by presenting Layer-wise Error Modeling Search (LEMS) and KFAC-SVD. |
Moritz Thoma; Maximilian Groezinger; Maximilian Forstenhäusler; Emad Aghajanzadeh; Manoj Rohit Vemparala; CHRISTOS ANAGNOSTOPOULOS; Pierpaolo Mori; Nael Fasfous; Alexander Frickenstein; Daniel Mueller-Gritschneder; Ulf Schlichtmann; | code |
| 888 | DistMatch: Adaptive Binning Via Distribution Matching for Robust Sequential Conformal Prediction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While recent methods attempt to approximate exchangeability through reweighting, identifying optimal weights remains an open challenge. To address this limitation, we propose [DistMatch](https://anonymous.4open.science/r/distmatch/), a binning-based method that recursively partitions residuals within a binary tree using the Kolmogorov–Smirnov (KS) statistic. |
Enver Menadjiev; Jihyeon Seong; Jisu Yeo; Jaesik Choi; | code |
| 889 | Names Don’t Matter: Symbol-Invariant Transformer for Open-Vocabulary Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a novel Transformer-based mechanism that is provably invariant to the renaming of interchangeable tokens. |
İlker Işık; Wenchao Li; | code |
| 890 | DgMARK: Decoding-Guided Watermarking for Diffusion Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose dgMARK, a decoding-guided watermarking method for discrete diffusion language models (dLLMs). |
Pyo Hong; Albert No; | code |
| 891 | T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We argue that this instability stems from inefficient exploration in multi-turn settings, where policies continue to generate low-information actions that neither reduce uncertainty nor advance task progress. To address this issue, we propose Token- and Turn-level Policy Optimization (T$^2$PO), an uncertainty-aware framework that explicitly controls exploration at fine-grained levels. |
Haixin Wang; Hejie Cui; Chenwei Zhang; Xin Liu; Shuowei Jin; Shijie Geng; Xinyang Zhang; Nasser Zalmout; Zhenyu Shi; Yizhou Sun; | code |
| 892 | AdaGC: Enhancing LLM Pretraining Stability Via Adaptive Gradient Clipping Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Regardless of the underlying cause, these spikes manifest as unstable optimizer updates, as abnormal gradients contaminate both first- and second-moment states. In this paper, we propose a principled gradient-centric remedy: AdaGC, an adaptive per-tensor gradient clipping scheme that mitigates such contamination by bounding gradient norms relative to a tensor-wise exponential moving average of their historical clipped values. |
Guoxia Wang; Shuai Li; Congliang Chen; Jinle Zeng; Jiabin Yang; Dianhai Yu; Yanjun Ma; Li Shen; | code |
| 893 | Principled Zero-shot Ranking Agents with Tournament Graphs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a *tournament graph* framework that provides a principled foundation for $k$-wise reranking. |
Sheshansh Agrawal; Thien Nguyen; Douwe Kiela; | code |
| 894 | From 2D Grids to 1D Tokens: Reforming Shared Representations for Multimodal Image Fusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose *Selective Token Editing* (STE): we sparsely update/replace only a small set of critical shared tokens, providing a lightweight token-level mechanism to steer global appearance coherence while keeping the fusion backbone unchanged and avoiding complex loss designs. |
Yuchen Xian; Yunqiu Xu; Yang He; Yi Yang; | code |
| 895 | VeRO: An Evaluation Harness for Agents to Optimize Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Agent optimization differs fundamentally from conventional software engineering: the target agent interleaves deterministic code with stochastic LLM completions, requiring structured capture of both intermediate reasoning and downstream execution outcomes. To address these challenges, we introduce VeRO (**Ve**rsioning, **R**ewards, and **O**bservations), which provides (1) a reproducible evaluation harness with versioned agent snapshots, budget-controlled evaluation, and structured execution traces, and (2) a benchmark suite of target agents and tasks with reference evaluation procedures. |
Varun Ursekar; Apaar Shanker; Veronica Chatrath; Yuan Xue; Samuel Denton; | code |
| 896 | ECA: Efficient Continual Alignment for Open-Ended Image-to-Text Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this context, we introduce a new notion of continual alignment, which incrementally adapts the alignment module within pre-trained VLMs to preserve high-quality cross-modal representations. Based on this idea, we propose **E**fficient **C**ontinual **A**lignment (ECA), a novel exemplar-free IL approach for OpenITG. |
Jiangtao Kong; Peijun Zhao; Chun-Fu (Richard) Chen; Youngwook Do; Shaohan Hu; Tianyi Zhou; Huajie Shao; | code |
| 897 | MineDraft: A Framework for Batch Parallel Speculative Decoding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, the performance of standard SD is often limited by the strictly sequential execution of these drafting and verification stages. To address this, this paper proposes MineDraft, a batch parallel speculative decoding (PSD) framework designed to effectively hide drafting latency by overlapping it with verification. |
Zhenwei Tang; Arun Verma; Zijian Zhou; Zhaoxuan Wu; Alok Prakash; Daniela Rus; Bryan Kian Hsiang Low; | code |
| 898 | Plug-and-Play Diffusion Meets ADMM: Dual-Variable Coupling for Robust Medical Image Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This lack of historical tracking inevitably leads to non-vanishing steady-state bias, where the reconstruction fails to strictly satisfy physical measurements under heavy corruption. To resolve this, we propose **Dual-Coupled PnP Diffusion**, which restores the classical dual variable to provide integral feedback, theoretically guaranteeing asymptotic convergence to the exact data manifold. |
Chenhe Du; Xuanyu Tian; Qing Wu; Muyu Liu; Jingyi Yu; Hongjiang Wei; Yuyao Zhang; | code |
| 899 | A Diffusive Classification Loss for Learning Energy-based Generative Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Direct maximum likelihood is computationally prohibitive due to the need for nested sampling, while score matching, though efficient, suffers from mode blindness. To address these issues, we introduce the Diffusive Classification (DiffCLF) objective, a simple method that avoids blindness while remaining computationally efficient. |
RuiKang OuYang; Louis Grenioux; Jose Miguel Hernandez-Lobato; | code |
| 900 | WBMM: Windowed Batch Matrix Multiplication for Efficient Large Receptive Field Convolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While Large Kernel Acceleration (LKA) helps on small feature maps, it becomes \textbf{counterproductive on large feature maps}, even slower than non-accelerated implementations. We propose Windowed Batch Matrix Multiplication (WBMM), which \emph{partitions} input into contiguous windows and \emph{indexes} a compact relative position bias table to construct weight matrices, enabling regular memory access via batched matrix multiplication; this yields a unique property where \textbf{WBMM’s throughput improves with larger windows}, opposite to depthwise convolutions that degrade with larger kernels. |
Wan Song; Zhou Wei; Rui Wang; Jun-KUT Yu; Toru Kurihara; Xu Jiajia; shu zhan; | code |
| 901 | VIA-SD: Verification Via Intra-Model Routing for Speculative Decoding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose **V**erification via **I**ntr**a**-Model Routing for **S**peculative **D**ecoding (VIA-SD), a multi-tier framework using a routed slim-verifier. |
Yuchen Xian; Yang He; Yunqiu Xu; Yi Yang; | code |
| 902 | Latent Diffusion Pretraining for Crystal Property Prediction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce a novel latent-diffusion based pretraining framework CrysLDNet designed to mitigate the data scarcity issue. |
Shrimon Mukherjee; KISHALAY DAS; Partha Basuchowdhuri; Pawan Goyal; Niloy Ganguly; | code |
| 903 | Lifting Traces to Logic: Programmatic Skill Induction with Neuro-Symbolic Learning for Long-Horizon Agentic Tasks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Neuro-Symbolic Skill Induction (NSI), a framework that lifts interaction traces into modular, \textit{logic-grounded} programs. |
Jie-Jing Shao; Haiyan Yin; Yueming LYU; Xingrui Yu; Lan-Zhe Guo; Ivor Tsang; James Kwok; Yu-Feng Li; | code |
| 904 | Aligning Datasets and Models for Weight Space Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose to learn a dataset-aligned latent space for neural networks, where datasets information is induced during training. |
Aron Asefaw; Konstantinos Tzevelekakis; Damian Falk; Léo Meynent; Damian Borth; | code |
| 905 | PATCHCODE: Discrete Latent Predictive Learning for EEG Foundation Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: EEG foundation models aim to learn transferable representations, yet EEG recordings are dominated by high-frequency noise and large cross-subject variability. |
KIEREN YU; Ziyang LIU; CHANG Huang; Kaishun WU; | code |
| 906 | Bio-Inspired Self-Supervised Learning for Wrist-worn IMU Signals Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a novel tokenization strategy grounded in the *submovement theory* of motor control, which posits that continuous wrist motion is composed of superposed elementary basis functions called submovements. |
Prithviraj Tarale; Kiet Chu; Abhishek Varghese; Kai-Chun Liu; Maxwell Xu; Mohit Iyyer; Sunghoon Lee; | code |
| 907 | One-Step Residual Shifting Diffusion for Image Super-Resolution Via Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite the development of several methods to accelerate diffusion-based SR models, some (e.g., SinSR) fail to produce realistic perceptual details, while others (e.g., OSEDiff) may hallucinate non-existent structures. To overcome these issues, we present **RSD**, a new distillation method for ResShift. |
Daniil Selikhanovych; David Li; Aleksei Leonov; Nikita Gushchin; Sergey Kushneryuk; Alexander Filippov; Evgeny Burnaev; Iaroslav Koshelev; Aleksandr Korotin; | code |
| 908 | Towards The Explainability of Temporal Graph Networks Via Memory Backtracking and Topological Attribution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing explanation methods overlook the memory module, the core component that records and updates node histories, leaving the influence of past events unexplored. To address this challenge, we propose a method that attributes TGNs predictions through the topology attribution tree and memory backtracking tree. |
Yazheng Liu; Xi Zhang; Sihong Xie; Hui Xiong; | code |
| 909 | Towards Professional-Grade Financial Agents: Benchmarking, Tooling, and Structured Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Building on the tool library, we introduce ProFinAgents, a structured agent framework based on Directed Acyclic Graph (DAG) and Case-Based Memory (CBM). |
Huangcheng; Jinghua Piao; Wang Ranran; Yong Li; | code |
| 910 | Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, standard evaluations rely on aggregate metrics (e.g., MSE) that conflate model capability with the intrinsic difficulty of the evaluated instances. To address this, we propose a diagnostic framework anchored in **Spectral Coherence Predictability (SCP)**, which provides an efficient $\mathcal{O}(N\log N)$ per-instance difficulty reference and yields a corresponding linear MSE lower bound. |
万进 冯; Yuan Yuan; Ding; Yong Li; | code |
| 911 | LARFT: Closing The Cognition-Action Gap for Length Instruction Following in Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods primarily attempt to enforce length constraints by externally imposing length signals or optimization objectives, while largely overlooking the underlying limitation: the model’s intrinsic deficit in length cognition. To address this, we propose \textbf{LARFT} (\textbf{L}ength-\textbf{A}ware \textbf{R}einforcement \textbf{F}ine-\textbf{T}uning), a training framework that aligns the model’s length cognition with its action. |
Wei Zhang; Lintong Du; yuanhe zhang; Zhenhong Zhou; Kun Wang; Li Sun; Sen Su; | code |
| 912 | Hyperbolic RQ-VAE Enhanced Generative Recommendation with Differential-Length Codebook Strategy Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, most existing GR methods adopt residual quantization to implicitly model hierarchical relationships across codebook layers in Euclidean space, which distorts the intrinsic tree-like hierarchy and leads to low codebook utilization. To address these issues, we propose a Hyperbolic RQ-VAE enhanced Generative Recommendation, namely HG-Rec. |
aoran zhang; Yu-Bin Yang; Yonghong Yu; | code |
| 913 | FlatLab: A Unified Methodology Framework and Simulation-Based Benchmark for Robotic Manipulation of Flat Objects Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a unified framework that decouples the manipulation into a strategy generator and an action execution module. |
Xingyu Zhu; Wenshuo Han; Zhouyu Wang; Yuran Wang; Ruihai Wu; Hao Dong; Fan Tang; Hechang Chen; Hyung Jin Chang; Yixing Gao; | code |
| 914 | Low-Compute Watermark Removal Via Dual-Domain Natural Projection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing single-image attacks typically optimize only for the first two, achieving strong watermark suppression but relying on expensive, multi-step optimization that limits practical deployment. In this work, we show that this trade-off is fundamental: no current approach achieves all three properties simultaneously. |
Pragati Meshram; Varun Chandrasekaran; | code |
| 915 | Asymptotically Fast Clebsch-Gordan Tensor Products with Vector Spherical Harmonics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we provide the first complete algorithm which truly provides asymptotic benefits Clebsch-Gordan tensor products. |
YuQing Xie; Ameya Daigavane; Mit Kotak; Tess Smidt; | code |
| 916 | Merge to Remember: Sharpness-Aware Isotropic Merging for Continual Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing approaches typically rely on sequential fine-tuning or model merging strategies, yet often overlook the impact of loss landscape sharpness and dominant singular value directions, which leads to subspace misalignment and severe knowledge forgetting. In this paper, we propose the Sharpness-Aware Isotropic Merging (SAIM) framework, which introduces targeted optimizations in both the fine-tuning and merging stages to address these issues. |
Qun Yang; Enneng Yang; Li Shen; Wei Chen; Long Lan; | code |
| 917 | EvoEGF-Mol: Evolving Exponential Geodesic Flow for Structure-based Drug Design Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To avoid the instantaneous trajectory collapse induced by geodesics directly targeting Dirac distributions, we propose Evolving Exponential Geodesic Flow for SBDD (EvoEGF-Mol), which replaces static Dirac targets with dynamically concentrating distributions, ensuring stable training via a progressive-parameter-refinement architecture. |
Yaowei Jin; Junjie Wang; cheng cao; Penglei Wang Penglei; Duo An; Qian Shi; | code |
| 918 | Degradation-Aware Metric Prompting for Hyperspectral Image Restoration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, current methods often rely on impractical explicit priors or opaque black-box representations that overfit to training distributions, hampering generalization to unseen scenarios. To bridge this gap, we propose Degradation-Aware Metric Prompting (DAMP), a novel framework that characterizes multi-dimensional degradations through interpretable spatial-spectral metrics. |
Binfeng Wang; Di Wang; Haonan Guo; Ying Fu; Jing Zhang; | code |
| 919 | HARD-KV: Head-Adaptive Regularization for Decoding-time KV Compression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Crucially, we propose a Logits Calibration mechanism that normalizes diverse importance metrics into a unified probability space, enabling consistent Top-$p$ budgeting across heterogeneous heads. |
Yuxuan Yang; Feiyang Ren; Bowen Zeng; Dalin Zhang; Jinpeng Chen; Gang Chen; Huan Li; | code |
| 920 | Step-Level Sparse Autoencoder for Reasoning Process Interpretation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose step-level sparse autoencoder~(\name), which serves as an analytical tool to disentangle different aspects of LLMs’ reasoning steps into sparse features. |
Xuan Yang; Jiayu Liu; Yuhang Lai; Hao Xu; Zhenya Huang; Ning Miao; | code |
| 921 | SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Moreover, they intervene only at action decoding while keeping visual representations fixed—insufficient under perceptual ambiguity, where reconsidering how to perceive is as important as deciding what to do. To address these limitations, we propose SCALE, a simple inference strategy that jointly modulates visual perception and action based on ‘self-uncertainty’, inspired by uncertainty-driven exploration in Active Inference theory—requiring no additional training, no verifier, and only a single forward pass. |
Hyeonbeom Choi; Daechul Ahn; Youhan Lee; Taewook Kang; Seongwon Cho; Jonghyun Choi; | code |
| 922 | ScenePilot: Controllable Boundary-Driven Critical Scenario Generation for Autonomous Driving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose ScenePilot, a feasibility-guided, boundary-driven framework that targets the boundary band: scenarios that are physically solvable in principle yet still cause the deployed autonomy stack to fail. |
Qiyu Ruan; YUXUAN WANG; He Li; Zhenning Li; Cheng-Zhong Xu; | code |
| 923 | SS‑TPT: Stability and Suitability-Guided Test-Time Prompt Tuning for Adversarially Robust Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent test-time adaptation defenses improve robustness by leveraging many augmented views, but this leads to impractical slowdown and a clear robustness-throughput trade-off. To address this challenge, we present Stability and Suitability-guided Test-time Prompt Tuning (SS-TPT), evaluating the quality of each augmented view via two complementary scores: (1) stability, measuring prediction invariance to weak augmentations, and (2) suitability, measuring feature-space density among views. |
Sunoh Kim; Daeho Um; | code |
| 924 | STLA: Spatiotemporal Lookahead Alignment for Post-Training Quantization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introduces STLA, a novel rounding-optimized PTQ framework that achieves both fast and accurate LLM quantization. |
Zuqi Zhang; Chenghe Sun; Xiangyi Chu; Wei-Han Yu; Ka-Fai Un; Rui Martins; Pui-In Mak; Jiawei Xu; | code |
| 925 | Unbiased Alignment for Large Language Models with Noisy Preferences Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, these methods are susceptible to the significant noise prevalent in real-world preference datasets. To address this critical issue, we present a theoretical framework for unbiased alignment, introducing the *Unbiased Reward Model* (URM) loss and the *Unbiased Direct Preference Optimization* (UDPO) loss. |
Jialiang Wang; Xianming Liu; Xiong Zhou; Hui Liu; Haoliang Li; | code |
| 926 | TextME: Bridging Unseen Modalities Through Text Descriptions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce TextME, the first text-only modality expansion framework, to the best of our knowledge, projecting diverse modalities into LLM embedding space as a unified anchor. |
Soyeon Hong; Jinchan Kim; Jaegook You; Seungtaek Choi; Suha Kwak; Hyunsouk Cho; | code |
| 927 | Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introduces RL-Focal, a two-stage RL agent framework that routes and ensembles LLMs. |
Selim Tekin; Gaowen Liu; Ramana Kompella; Ling Liu; | code |
| 928 | Finding The Correct Visual Evidence Without Forgetting: Mitigating Hallucination in LVLMs Via Inter-Layer Visual Attention Discrepancy Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite this progress, they are still prone to hallucination, generating responses that are semantically coherent but inconsistent with visual content. In this work, we find that LVLMs tend to hallucinate when they pay insufficient attention to the correct visual evidence and gradually forget it during the generation process, leading to more hallucinations. |
Yutong Xie; Zhenglin Hua; Ran Wang; Wing W. Y. Ng; Xizhao Wang; Yuheng Jia; | code |
| 929 | DIVA: Harnessing The Representation Divergence in Unified Multimodal Models for Mutual Reinforcement Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Motivated by the observation, we propose DIVA, a self-improved post-training framework that transforms the representation divergence into interior synergy. |
renjie lu; Xulong Zhang; Xiaoyang Qu; Jianzong Wang; Shangfei Wang; | code |
| 930 | GaussTrace: Provenance Analysis of 3D Gaussian Splatting Models with Evidence-based LLM Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, the widespread sharing and iterative modification of 3DGS models across digital platforms create pressing challenges for intellectual property protection and forensic traceability. To address this, we propose GaussTrace, a novel framework for constructing directed provenance graphs for 3DGS models. |
Haoliang Han; Ziyuan Luo; Renjie Wan; | code |
| 931 | CORAL: Uncertainty-Aware Regulation of Exposure Concentration in Recommender Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods are often post hoc and typically lack principled uncertainty-aware risk estimates for regulating exposure under endogenous feedback. We therefore propose **CORAL**, a model-agnostic, uncertainty-aware framework that formulates exposure regulation as a constrained sequential decision problem. |
Nitin Bisht; Linjiang Guo; Xiuwen Gong; Huan Huo; Guandong Xu; | code |
| 932 | CARE: Adaptive Calibration for Reliable Recommendations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose **CARE**, an adaptive calibration framework that wraps an arbitrary backbone recommender and outputs variable-size recommendation sets with finite-sample performance guarantees over interaction streams. |
Nitin Bisht; Xiuwen Gong; Huan Huo; Guandong Xu; | code |
| 933 | $\texttt{MetaDistill}$: Unlocking The Performance Ceiling for Pretrained Optimizers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose $\texttt{MetaDistill}$, a general MetaBBO training framework designed to lift the strategy ceiling through pretraining and test-time fine-tuning. |
Muqi Han; Ruoqi Xing; KAI WU; Xiaoyu Zhang; Handing Wang; Zilong Wang; | code |
| 934 | Reverse-Engineering Model Editing on Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, in this work, we reveal a critical vulnerability of this paradigm: the parameter updates inadvertently serve as a side channel, enabling attackers to recover the edited data. We propose a two-stage reverse-engineering attack named KSTER (KeySpaceReconsTruction-then-EntropyReduction) that leverages the low-rank structure of these updates. |
Zhiyu Sun; Minrui Luo; Yu Wang; Tianxing He; Zhili Chen; | code |
| 935 | Difference-Aware Decision Learning for Multimodal Image Fusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Therefore, we address this problem by formally casting image fusion as an integrated probabilistic decision system that couples prior decision-making with posterior risk minimization. Based on this view, we propose a dIfference-aware Decision-lEArning muLtimodal image fusion paradigm (IDEAL). |
Hao Pan; Jian Dai; Yuan Sun; Zhenwen Ren; Xingfeng Li; | code |
| 936 | Spectra: Rethinking Optimizers for LLMs Under Spectral Anisotropy Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Motivated by this analysis, we propose \textit{Spectra}, a spike-aware optimizer that suppresses the dominant low-rank spike subspace without amplifying the noise-sensitive spectral tail. |
Zhendong Huang; Hengjie Cao; Fang DONG(董方); Ruijun Huang; Mengyi Chen; Yifeng Yang; Xin Zhang; Anrui Chen; Mingzhi Dong; Yujiang Wang; Jinlong Hou; Qin Lv; Robert Dick; Yuan Cheng; Tun Lu; Fan Yang; Li Shang; | code |
| 937 | Thinking in Structures: Evaluating Spatial Intelligence Through Reasoning on Constrained Manifolds Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce SSI-Bench, a VQA benchmark for spatial reasoning on constrained manifolds, built from complex real-world 3D structures whose feasible configurations are tightly governed by geometric, topological, and physical constraints. |
Chen Yang; Guanxin Lin; Youquan He; Peiyao Chen; Guanghe Liu; Yufan Mo; Zhouyuan Xu; Linhao Wang; Guohui Zhang; Zihang Zhang; Shenxiang Zeng; Chen Wang; Jiansheng Fan; | code |
| 938 | Global Credit Assignment Via Dynamical Criticality Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by the criticality observed in biological neural circuits, we introduce Criticality-driven Online Local Alignment (COLA). |
Wentao Wang; Keren Gao; Guozhang Chen; | code |
| 939 | Vibe Checker: Aligning Code Evaluation with Human Preference Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we hypothesize that instruction following is the missing piece underlying *vibe check* besides functional correctness. |
Ming Zhong; Xiang Zhou; Ting-Yun Chang; Qingze Wang; Nan Xu; Xiance Si; Dan Garrette; Shyam Upadhyay; Jeremiah Zhe Liu; Jiawei Han; Benoit Schillings; Jiao Sun; | code |
| 940 | The Data Manifold Under The Microscope Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a benchmarking framework for studying data geometry by repurposing and extending dSprites and COIL-20 with additional transformation dimensions and denser sampling, enabling accurate finite-difference estimates of curvature, reach, and volume that are otherwise difficult to estimate reliably and implement in practice. |
Marios Koulakis; Constantin Seibold; | code |
| 941 | Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM Decoding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This establishes a pathological shortcut that bypasses visual grounding. To dismantle this, we propose Fox (Faithfulness and Observational-flow via eXpression-rectification), a training-free inference-time framework. |
Liu Yu; Can Chen; PING KUANG; Zhikun Feng; Fan Zhou; Gillian Dobbie; | code |
| 942 | Unifying Stacking and Cascading for Efficient Ensemble Inference Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce LazyStack, a method for efficient model ensemble inference. |
Ashwin Colaço; Sharad Mehrotra; Michael De Lucia; Kevin Hamlen; Murat Kantarcioglu; Latifur Khan; Ananthram Swami; Bhavani Thuraisingham; Unnat Jain; | code |
| 943 | Threat2Traffic: Multi-Agent Environment Synthesis for Malware Traffic Generation from Threat Intelligence Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Threat2Traffic, a multi-agent framework that extracts sample-specific dependencies from threat intelligence, reconstructs tailored environments, and captures malware traffic. |
Haoyang Chen; Chang Liu; Zhong Guan; Junzheng Shi; Gaopeng Gou; Gang Xiong; | code |
| 944 | Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Since users rarely specify every detail of their intent, practical web agents must be able to interpret ambiguous queries by inferring user preferences and contexts. To address this challenge, we present Persona2Web, the first benchmark for evaluating personalized web agents on the real open web, built upon the clarify-to-personalize principle, which requires agents to resolve ambiguity based on user history rather than relying on explicit instructions. |
Serin Kim; Sangam Lee; Dongha Lee; | code |
| 945 | DualTimesField: Rethinking Time Series As Continuous-Time Trends and Events Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Discrete-time representations struggle with irregular sampling and the tradeoff of fidelity and efficiency, while traditional implicit neural representations suffer from spectral bias and frequency entanglement. To address these challenges, we conceptualize time series as the superposition of continuous trends and discrete events from a continuous-time perspective and propose DualTimesField, a framework that utilizes dual implicit neural fields. |
Wencheng Zhang; Long Li; Huayi Qin; Zongjuan Wu; Jing Li; Wanghu Chen; | code |
| 946 | Not All Invariants Are Equal: Curating Training Data to Accelerate Program Verification with SLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a rigorous data curation pipeline designed to extract high-quality training signals from raw verifier-generated invariants. |
Ido Pinto; Yizhak Elboher; Haoze Wu; Nina Narodytska; Guy Katz; | code |
| 947 | The Devil Is in The Condition Numbers: Why Is GLU Better Than Non-GLU Structure? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we study GLU by analyzing two-layer networks in the neural tangent kernel (NTK) regime. |
Xingyu Lyu; Qianqian Xu; zhiyong yang; Peisong Wen; Qingming Huang; | code |
| 948 | Knowing Bias, Doing Better: Mitigating Social Bias in LLMs Via Know-Bias Neuron Enhancement Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose \textbf{KnowBias}, a lightweight and conceptually distinct framework that mitigates bias by strengthening, rather than suppressing, neurons encoding bias-knowledge. |
Jinhao Pan; Chahat Raj; Anjishnu Mukherjee; Sina Mansouri; Bowen Wei; Shloka Yada; Ziwei Zhu; | code |
| 949 | The Realignment Problem: When Right Becomes Wrong in LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce TRACE (Triage and Re-align by Alignment Conflict Evaluation), a framework that transforms re-alignment into a structured optimization problem over existing data without requiring fresh human annotation. |
Aakash Sharma; Debdeep Sanyal; Manodeep Ray; Vivek Srivastava; Shirish Karande; Murari Mandal; | code |
| 950 | Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present an efficient video augmentation framework that converts simulated VLA videos into realistic training videos while preserving task semantics and action trajectories. |
Chenyu Hui; Xiaodi Huang; Siyu Xu; Yunke Wang; Shan You; Fei Wang; Tao Huang; Chang Xu; | code |
| 951 | Adaptive Residual-Update Steering for Low-Overhead Hallucination Mitigation in Large Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose **Residual-Update Directed DEcoding Regulation (RUDDER)**, a framework that counters visual dilution by creating a persistent visual anchor. |
Zhengtao Zou; Ya Gao; Jiarui Guan; Bin Li; Pekka Marttinen; | code |
| 952 | FairGB: A Fair Granular-Ball Generation Method for Data Classification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Theoretical analysis shows that FairGBG preserves high purity within each GB while satisfying group fairness. |
Qifen Yang; Yuhui Deng; Jiande Huang; Peng Zhou; Xiwen Lu; Lin Cui; | code |
| 953 | D3LLM: Ultra-Fast Diffusion LLM Using Pseudo-Trajectory Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite increasing interest, existing methods typically focus on only one-side of the coin, targeting either efficiency or performance. To address this limitation, we propose d3LLM (*Pseudo-Distilled Diffusion Large Language Model*), striking a balance between accuracy and parallelism: (i) during training, we introduce *pseudo-trajectory distillation* to teach the model which tokens can be decoded confidently at early steps, thereby improving parallelism; (ii) during inference, we employ *entropy-based multi-block decoding* with a KV-cache refresh mechanism to achieve high parallelism while maintaining accuracy. |
Yu-Yang Qian; Junda Su; Lanxiang Hu; Peiyuan Zhang; Zhijie Deng; Peng Zhao; Hao Zhang; | code |
| 954 | Optimizing KV Cache Eviction from An Output Perturbation Perspective Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our analysis reveals that, beyond attention weights, the value states within KV entries and pretrained parameter matrices are also crucial. Based on this, we propose a perturbation-constrained selection algorithm that optimizes the worst-case output perturbation to identify critical entries. |
Yuan Feng; Junlin Lv; Haoyu Guo; Yukun Cao; Xike Xie; S Kevin Zhou; | code |
| 955 | Words Towards Explainability: Caption Label-Free Learning Via Dual Loop Agentic Time Series Captioning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Departing from the supervised learning tradition of imitating human annotations, CLFL formulates captioning as an agentic exploration task optimized by feedback from a proxy reward. Specifically, we propose a Dual Loop Agentic Captioning (DLAC) framework to achieve such an exploration-feedback mechanism. |
Difei Hou; Jiaqi Yue; Chunhui Zhao; | code |
| 956 | IQA-Spider: Unifying Reasoning, Grounding, and Referring for Multi-Granularity Image Quality Assessment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present IQA-Spider, the first image quality assessment (IQA) framework that unifies reasoning, grounding, and referring within a LMM-based system for multi-granularity quality understanding. |
Xinge Peng; Yiting Lu; Xin Li; Zhibo Chen; | code |
| 957 | ReMoE: Boosting Expert Reuse Through Router Fine-Tuning in Memory-Constrained MoE LLM Inference Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose ReMoE, a router fine-tuning framework designed to boost token-wise expert reuse. |
Xiongwei Zhu; Xiaojian Liao; Tianyang Jiang; Yusen Zhang; Liang Wang; Limin Xiao; | code |
| 958 | Embodied-DETR: End-to-End Temporal 3D Object Detection in Egocentric Views Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce **Embodied-Det**, a new benchmark for egocentric 3D object detection that evaluates detection accuracy, temporal stability, and consistency under embodied settings. Building on this benchmark, we propose **Embodied-DETR**, an end-to-end temporal detection framework that models scene-level context and instance-level continuity through two complementary temporal modules, *Scene-aware Feature Aggregation* and *Instance-aware Query Embedding*. |
Ziheng Ding; Xiaze Zhang; Yuejie Zhang; lifeng chen; Rui Feng; | code |
| 959 | ResRL: Boosting LLM Reasoning Via Negative Sample Projection Residual Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To boost reasoning ability without losing diversity, this paper proposes negative sample projection Residual Reinforcement Learning (ResRL) that decouples similar semantic distributions among positive and negative responses. |
Zihan Lin; Xiaohan Wang; Jie Cao; Jiajun Chai; Li Wang; Xiaodong Lu; Wei Lin; Ran He; Guojun Yin; | code |
| 960 | Furina: Fragmented Uncertainty-Driven Refusal Instability Attack Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Through systematic experiments, we identify a characteristic \emph{diagnostic signature}: inputs in unstable regimes exhibit elevated output uncertainty yet \emph{decreased} internal safety activation, a decoupling phenomenon that explains why detection-based defenses fail against sophisticated attacks. Building on this framework, we introduce \textbf{Furina}, a jailbreak attack that deliberately induces this signature through fragmented, scene-anchored prompts without model-specific optimization. |
Tongxi Wu; Jian Zhang; Yang Gao; | code |
| 961 | Bridging Tokens and Geometry: Token-wise 3D Supervision for CAD Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose an Argument-induced 3D Point Loss (A3PL) that maps argument tokens to corresponding 3D points, enabling dense token-wise geometric supervision. |
Yijia Guan; Jianhua Sun; | code |
| 962 | Learning High-Frequency Continuous Action Chunks in Latent Space Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: At such high frequencies, policies often fail to generate actions that are both temporally smooth and spatially consistent. We address this challenge by shifting high-frequency action learning from the action space to a latent space with variational autoencoder (VAE). |
Kunyun Wang; Yuhang Zheng; Jieru Zhao; Yupeng Zheng; Wenchao Ding; | code |
| 963 | Balancing Understanding and Generation in Discrete Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In discrete generative modeling, two dominant paradigms demonstrate divergent capabilities: Masked Diffusion Language Models (MDLM) excel at semantic understanding and zero-shot generalization, whereas Uniform-noise Diffusion Language Models (UDLM) achieve strong few-step generation quality, yet neither attains balanced performance across both dimensions. To address this, we propose XDLM, which bridges the two paradigms via a stationary noise kernel. |
Yue Liu; Yuzhong Zhao; Zheyong Xie; Qixiang Ye; Jianbin Jiao; Yao Hu; Shaosheng Cao; Liu; | code |
| 964 | Reasoning Can Be Restored By Correcting A Few Decision Tokens Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: For instance, on Qwen3-0.6B, only $\sim$8\% of generated tokens account for the salient disagreement; these tokens concentrate early in the response, are strongly enriched in planning-related decisions ($17\times$), and coincide with high base-model uncertainty—suggesting that base models fail mainly at early planning points that steer the subsequent reasoning trajectory. Building on these findings, we propose disagreement-guided token intervention, a simple inference-time delegation scheme that performs a one-token takeover by the reasoning model only at high-disagreement positions and immediately switches back to the base model. |
Shen Changshuo; Leheng Sheng; Yuxin Chen; Xiang Wang; An Zhang; | code |
| 965 | Spike-HTR: Spiking Neural Transformer for Handwritten Text Recognition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Spike-HTR, a budgeted spiking Transformer that controls two coupled knobs: the spiking horizon $T$ and the effective token length $\ell_b$ after blank-guided reduction. |
Xiubo Liang; Jinxing Han; Yuke Li; Haoqi Zhu; Yu Zhao; Hongzhi Wang; | code |
| 966 | HDFlow: Hierarchical Diffusion-Flow Planning for Long-horizon Tasks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce $\textbf{Hierarchical Diffusion-Flow}$ ($\texttt{\textbf{HDFlow}}$), a novel hierarchical planning framework that optimally leverages the strengths of $\textit{diffusion}$ and $\textit{rectified flow}$ models to overcome the limitations of single-paradigm generative planners. |
Gireesh Nandiraju; Yuanliang Ju; Chaoyi Xu; Weiheng Liu; Yuxuan Wan; He Wang; | code |
| 967 | Future Dynamic 3D Reconstruction: A 3D World Model with Disentangled Ego-Motion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose FR3D, a world model that predicts a persistent 3D latent representation for future dynamic 3D reconstruction. |
Nils Morbitzer; Jonathan Evers; Artem Savkin; Thomas Stauner; Nassir Navab; Federico Tombari; Stefano Gasperini; | code |
| 968 | InteractBench: Benchmarking LLMs on Competitive Programming Under Unrevealed Information Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Crucially, new information is revealed *only* in response to queries. To address this gap, we introduce *InteractBench*, a benchmark comprising 322 high-quality interactive problems curated from Codeforces, AtCoder, IOI, and ICPC. |
Jiaze Li; Aocheng Shen; Bing Liu; Boyu Zhang; Xiaoxuan Fan; Qiankun Zhang; Xianjun Deng; | code |
| 969 | Goal-Conditioned Agents That Learn Everything All at Once Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We show that this approach significantly outperforms other methods on goal-conditioned Craftax and is competitive with existing baselines on continuous control environments, while achieving a 250x speed-up compared to all-goals relabelling. |
Michael Matthews; Matthew Jackson; Michael Beukman; Thomas Foster; Alistair Letcher; Scott Fujimoto; Cédric Colas; Jakob Foerster; | code |
| 970 | An Exterior Method for Nonnegative Matrix Factorization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose an *exterior* framework for NMF (eNMF) that separates low-rank approximation from nonnegativity enforcement. |
Qiujing Lu; Tonmoy Monsoor; Ehsan Ebrahimzadeh; Kartik Sharma; Vwani Roychowdhury; | code |
| 971 | PULSE: Generative Phase Evolution for Non-Stationary Time Series Forecasting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods implicitly rely on static historical assumptions, leading to a critical failure mode we term Phase Amnesia, where models become blind to the evolving global context. To resolve this, we formalize non-stationary dynamics through three physical hypotheses: Wold decomposition, dynamical phase evolution, and heteroscedastic manifold generation. |
YangYou Liu; Zezhi Shao; Xinyu Chen; Hu Chen; Fei Wang; Yuankai Wu; | code |
| 972 | SLQ: Bridging Modalities Via Shared Latent Queries for Retrieval with Frozen MLLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing approaches typically rely on invasive parameter updates, such as full fine-tuning and LoRA, which risk disrupting the pre-trained semantic manifold and degrading the complex knowledge structures crucial for logical inference. To address this, we propose **SLQ**, a parameter-efficient tuning framework that adapts MLLMs for retrieval while keeping the backbone entirely frozen. |
Haoran Lou; Ziyan Liu; Chunxiao Fan; Yuexin Wu; Yue Ming; Hao Wu; Kai Zuo; Yibo Chen; Xu Tang; | code |
| 973 | Disentangling Consensus and Value-Specific Representations for Controllable Pluralistic Value Alignment of LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Consequently, modifying the contribution of one value expert may unintentionally influence other values, limiting fine-grained controllability. To address this issue, we propose DisAlign, a model-merging framework that explicitly decomposes value representations into consensus and value-specific components using an information-geometric perspective. |
JianKui Zhou; Jing Yao; Xiaoyuan Yi; Peng Zhang; Ning Gu; Zhan Hu; Xing Xie; Tun Lu; | code |
| 974 | When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Unlike execution errors, epistemic miscalibration is latent during planning, as generated plans can remain self-consistent and executable without observable errors; the miscalibration is also dynamic, as new information can alter feasibility assessments, potentially obscuring past miscalibration signals and causing them to recur over time. To address this, we propose the Epistemic Planning Calibration Agentic Workflow (EPC-AW), which assesses whether plans remain supported under varying information conditions rather than directly verifying feasibility. |
Zehao Wang; shilong jin; Zhao Cao; Lanjun Wang; | code |
| 975 | A3: An Analytical Low-Rank Approximation Framework for Attention Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Low-rank approximation offers a promising compression solution, yet existing approaches have two main limitations: (1) They focus on minimizing the output error of individual linear layers, without considering the architectural characteristics of Transformers, and (2) they decompose a large weight matrix into two small low-rank matrices. Consequently, these methods often fall short compared to other compression techniques like pruning and quantization, and introduce runtime overhead such as the extra GEMM kernel launches and memory operations for decomposed small matrices. |
Jeffrey Wong; Cheng Zhang; Xinye Cao; Pedro Gimenes; Christos-Savvas Bouganis; George Constantinides; Wayne Luk; Yiren Zhao; | code |
| 976 | KernelCraft: Benchmarking for Agentic Close-to-Metal Kernel Generation on Emerging Hardware Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present KernelCraft: the first benchmark to evaluate an LLM agent’s ability to generate and optimize low-level kernels for customized accelerators via a function-calling, feedback-driven workflow. |
Jiayi Nie; Haoran Wu; Yao Lai; Zeyu Cao; Cheng Zhang; Binglei Lou; Erwei Wang; Jianyi Cheng; Timothy Jones; Robert Mullins; Rika Antonova; Yiren Zhao; | code |
| 977 | Teaching Molecular Dynamics to A Non-Autoregressive Ionic Transport Predictor Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Moreover, existing methods typically benefit from datasets either with or without atomic trajectories, but not both. To overcome these limitations, we propose a non-autoregressive learning framework based on modality reduction, which treats atomic trajectories as an auxiliary modality during training but does not require them at inference. |
Jiyeon Kim; Byungju Lee; Won-Yong Shin; | code |
| 978 | Semi-Supervised Neural Super-Resolution for Mesh-Based Simulations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, training neural networks for super-resolution often demands large amounts of expensive HR supervision data. To address this challenge, we propose SuperMeshNet, an HR data-efficient super-resolution framework for mesh-based simulations aided by message passing neural networks (MPNNs). |
Jiyeon Kim; Youngjoon Hong; Won-Yong Shin; | code |
| 979 | Dissecting Embodied Abilities in Multimodal Language Models Through Skill-level Evaluation and Diagnosis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing embodied benchmarks fail to provide actionable insights because they focus on task-level evaluation rather than discovering capability bottlenecks. To address this, we introduce BEAR, where we divide embodied tasks into 14 atomic skills for skill-level evaluation. |
Yu Qi; Haibo Zhao; Ziyu Guo; Siyuan Ma; Ziyan Chen; Yaokun Han; Renrui Zhang; Zitiantao Lin; Yizhe Zhu; Shiji Xin; Yijian Huang; Boce Hu; Kai Cheng; Jiayi Zhang; Peiheng Wang; jiazheng liu; Wenqing Wang; Yiran Qin; Haojie Huang; Lawson Wong; | code |
| 980 | Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM Compression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose **Zero Sum SVD** (**ZS-SVD**), a post-training method that performs *global* singular component selection using activation whitening and first-order calibration loss estimates in whitened coordinates. |
Ali Abbasi; Chayne Thrash; Haoran Qin; Shansita Sharma; Sepehr Seifi; Soheil Kolouri; | code |
| 981 | Dive Into The Scene: Breaking The Perceptual Bottleneck in Vision-Language Decision Making Via Focus Plan Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose ${\it SceneDiver}$, a coarse-to-fine focus plan generation method for VLMs leveraging their long-term planning abilities, that first constructs a holistic scene graph to establish initial comprehension, then progressively decomposes the task into simpler sub-problems through an iterative cycle of recognition, understanding, and analysis. |
Boyuan Xiao; Bohong Chen; Yumeng Li; Ji Feng; Yao-Xiang Ding; Kun Zhou; | code |
| 982 | DyCon: Dynamic Reasoning Control Via Evolving Difficulty Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we empirically show that the problem difficulty evolves dynamically throughout the reasoning process and is linearly encoded in the LRM’s step-level embeddings. |
Tengyao Tu; Yulin Li; Huiling Zhen; Libo Qin; Zhoujun Wei; Jinghua Piao; Zhuotao Tian; Yong Li; Min zhang; | code |
| 983 | DC-Leap: Training-Free Acceleration of DLLMs Via Draft-Guided Contiguous Leaping Decoding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These thresholds, necessitated by the Joint Probability Dependence Error (JPDE), result in redundant denoising iterations and suboptimal inference speeds. To overcome this, we propose DC-Leap, a training-free framework that enables reliable acceleration of dLLMs in the moderate-confidence regime. |
Yan hua Jiao; Tianyi Wu; Xiaoxi Sun; Yulin Li; Huiling Zhen; Libo Qin; Baotian Hu; Zhuotao Tian; Min zhang; | code |
| 984 | ACTG-ARL: Differentially Private Conditional Text Generation with RL-Boosted Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Generating high-quality synthetic text under differential privacy (DP) is critical for training and evaluating language models without compromising user privacy. |
Yuzheng Hu; Ryan McKenna; Da Yu; Shanshan Wu; Han Zhao; Zheng Xu; Peter Kairouz; | code |
| 985 | SCOPE and SCION: Benchmark and Method for Ontology Induction and Fusion from Text Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose SCION (Structural mining and Contracted semantic Induction for Ontology constructiON and fusion), a controllable pipeline that mines a candidate space of concepts/relations/events from text, performs LLM-assisted naming/merging/filtering under a strict JSON contract with evidence pointers, and can fuse the result with a fixed base ontology package using conservative alignment with provenance tracking. |
Miaobo Hu; Shuhao Hu; BoKun Wang; Rui Chen; Xiaobo Guo; Xin Wang; Daren Zha; Jun Xiao; | code |
| 986 | Think Twice Before You Act: Protecting LLM Agents Against Tool Description Poisoning Via Isolated Planning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To understand the effectiveness of existing attacks against this emerging threat, we evaluate several existing agent defenses against prompt-injection and find they transfer poorly to cross-tool description poisoning. Building on this insight, we propose Tool-Guard, a novel defense based on a new concept called isolated planning, in which tool invocations that are detected as misaligned or suspicious cause the corresponding tool to be placed in a quarantined list (the influenced list), breaking further influence from poisoned descriptions. |
Shanghao Shi; Xiao Wang; Chaoyu Zhang; Hao Li; Wenjing Lou; Thomas Hou; Yevgeniy Vorobeychik; Chongjie Zhang; Ning Zhang; | code |
| 987 | Bias in Zeroth-Order Normal Estimation for Decision-Based Attacks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, this stationarity does not generally yield $\ell_2$-optimal perturbations under nonlinear boundaries. Building on this observation, we propose a novel and effective algorithm, Sensitivity-Aware Rescaling (SAR), that leverages this sensitivity signal to infer an importance map from the current best perturbation, then progressively suppresses low-importance regions through a coarse-to-fine schedule to reduce the $\ell_2$ norm. |
Feiyang Wang; Hangwei Qian; Xingquan Zuo; Gang Chen; Ivor Tsang; | code |
| 988 | Head-in-Head in Linear Attention Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by the multi-head mechanism, we propose Head-in-Head, which introduces an additional mask matrix to structure memory partitioning and interactions within a single linear-attention head. |
Shijie Mei; Man Yao; Jiabo Tong; Bo XU; Guoqi Li; | code |
| 989 | What Do Agents Learn from Trajectory-SFT: Semantics or Interfaces? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose **PIPE**, a protocol-level evaluation augmentation for diagnosing interface reliance by minimally rewriting environment interfaces while preserving task semantics and execution behavior. |
Weizheng Gu; Chengze Li; Zhuohao Yu; Mengyuan Sun; Zhibang Yang; Wei Wang; Hongrui Jia; Shikun Zhang; Wei Ye; | code |
| 990 | Beyond Problem Solving: UOJ-Bench for Evaluating Code Generation, Hacking, and Repair in Competitive Programming Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce **UOJ-Bench**, a benchmark designed to evaluate not only the problem-solving ability of LLMs, but also their ability to identify errors in human-written code—a crucial educational activity traditionally supported by running test cases over online judge systems. |
Tingqiang Xu; Hangrui Zhou; Tianle Cai; Alex Gu; Kaifeng Lyu; | code |
| 991 | FluxNet: Learning Capacity-Constrained Local Transport Operators for Conservative and Bounded PDE Surrogates Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a framework for learning conservative transport operators on regular grids, inspired by lattice Boltzmann–style discrete-velocity transport representations. |
Zishuo Lan; Junjie Li; Lei Wang; Jincheng Wang; | code |
| 992 | Envisioning Beyond The Few: Disentangled Semantics and Primitives for Few-Shot Atypical Layout-to-Image Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We term this failure as representation fragmentation, arising from a granularity mismatch that entangles semantic identity with visual details. To address this issue, we propose a representation-driven framework that disentangles semantics from primitives for robust few-shot adaptation. |
Nan Bao; Yifan Zhao; Wenzhuang Wang; Jia Li; | code |
| 993 | Automatic Pruning Discovery for Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce an affirmative answer by proposing a novel pruning method called $\textbf{AutoPrune}$, which first overcomes expert knowledge limits by leveraging LLMs to design optimal pruning algorithms for themselves automatically without any expert knowledge. |
Haidong Kang; Lihong Lin; Enneng Yang; Hong-Ning Dai; Hao Wang; | code |
| 994 | Heterogeneity-Aware Knowledge Sharing for Graph Federated Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address both types of heterogeneity, we propose a novel graph Federated learning method via Semantic and Structural Alignment (FedSSA), which shares the knowledge of both node features and structural topologies. |
Wentao Yu; Sheng Wan; Shuo Chen; Bo Han; Chen Gong; | code |
| 995 | Invariant Representation Learning for Source-Free Time Series Forecasting with LLM-Centric Proxy Denoising Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To maximally create value from sparse data, this study focuses on a new problem of source-free time series forecasting, aiming to adapt a pretrained model from sufficient source time series to the sparse target time series without access to the source data, enabling data protection. To achieve this, we propose TimeID, a novel source-free time series forecasting framework with a large language model (LLM) centric proxy denoising inspired by the powerful generalization capabilities of LLMs. |
Kangjia Yan; Chenxi Liu; Hao Miao; Xinle Wu; Yan Zhao; Chenjuan Guo; Bin Yang; | code |
| 996 | MiniAppBench: Evaluating The Shift from Text to Interactive HTML Responses in LLM-Powered Assistants Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing benchmarks primarily focus on algorithmic correctness or static layout reconstruction, failing to capture the capabilities required for this new paradigm. To address this gap, we introduce **MiniAppBench**, the first comprehensive benchmark designed to evaluate principle-driven, interactive application generation. |
Zuhao zhang; Chengyue Yu; Yuante Li; Chenyi Zhuang; Linjian Mo; Shuai Li; | code |
| 997 | Physiology-Aware Masked Cross-Modal Reconstruction for Biosignal Representation Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A canonical example is the relationship between electrocardiography (ECG), which captures the electrical activation initiating each heartbeat, and photoplethysmography (PPG), which records the resulting peripheral pulse delayed by vascular dynamics. To capture this structured relationship, we introduce xMAE, a biosignal pretraining framework that leverages masked cross-modal reconstruction across temporally ordered biosignals as a training-time constraint to encourage physiologically meaningful timing structure in the learned representations. |
Hao Zhou; Simon Lee; Cyrus Tanade; Keum San Chun; Juhyeon Lee; Migyeong Gwak; Megha Thukral; Justin Sung; Eugene Hwang; Mehrab Bin Morshed; Li Zhu; Viswam Nathan; Mahbubur Rahman; Subramaniam Venkatraman; Sharanya Desai; | code |
| 998 | PGC: Peak-Guided Calibration for Generalizable AI-Generated Image Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These fine-grained signals are often overshadowed by dominant, high-fidelity image content (e.g., the main subject), limiting the reliability of existing detectors that predominantly rely on global representations. To address this challenge, we propose the Peak-Guided Calibration (PGC) framework. |
Xiaoyu Zhou; Jianwei Fei; Peipeng Yu; Jingchang Xie; Chong Cheng; Zhihua Xia; | code |
| 999 | Beyond Explicit Edges: Robust Reasoning Over Noisy and Sparse Knowledge Graphs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, standard graph algorithms rely heavily on static connectivity and explicit edges, often failing in real-world scenarios where knowledge graphs (KGs) are noisy, sparse, or incomplete. To address this limitation, we introduce INSES (Intelligent Navigation and Similarity Enhanced Search), a dynamic framework designed to reason beyond explicit edges. |
Hang Gao; Dimitris Metaxas; | code |
| 1000 | Probing Cross-modal Information Hubs in Audio-Visual LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we focus on cross-modal information flow between audio and visual modalities in AVLLMs, investigating where information derived from one modality is encoded within the token representations of the other modality. |
Jihoo Jung; Chaeyoung Jung; Ji-Hoon Kim; Joon Son Chung; | code |
| 1001 | ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In real-world scenarios, tools are not independent; they are atomic, interdependent, and prone to environmental noise. We introduce $\textbf{ComplexMCP}$, a benchmark designed to evaluate agents in these rigorous conditions. |
Yuanyang Li; Yangxue; Longyue Wang; Weihua Luo; Hongyang Chen; | code |
| 1002 | FastSESR: Fast Scene-level Explicit Surface Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While existing methods achieve strong performance on scene-level data, they often rely on test-time optimization, resulting in a prohibitive runtime of several minutes. To address this bottleneck, we propose FastSESR, a two-stage framework for efficient scene-level explicit surface reconstruction. |
Jueqi Liu; Xuechao Zou; Congyan Lang; | code |
| 1003 | Mesh Based Simulations with Spatial and Temporal Awareness Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a unified framework to bridge the gap between geometric deep learning and rigorous numerical analysis. |
Paul Garnier; Vincent Lannelongue; Elie Hachem; | code |
| 1004 | Beyond Benchmarks: Toward Causally Faithful Evaluation of Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Prior studies either focus on individual components, overlooking their interactions, or investigate manually curated and small-scale question variants, lacking a holistic perspective, precluding precise attribution of intrinsic model capabilities amidst the confounding influences. To address these limitations, we propose LLM evaluatology, a principled framework that grounds LLM evaluation in a causally informed system design. |
Zhengshuyuan Tian; Chuanxin Lan; Chenxi Wang; Lei Wang; Guoxin Kang; Zhengxin Yang; Yunyou Huang; Xuehai Hong; Gao; Jianfeng Zhan; | code |
| 1005 | SpatialJB: How Text Distribution Art Becomes The Jailbreak Key for LLM Guardrails Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Exploiting the Transformer’s spatial weakness, we propose SpatialJB to disrupt the model’s output generation process, allowing harmful content to bypass guardrails without detection. |
Zhiyi Mou; Jingyuan Yang; ZEHENG QIAN; Wangze Ni; Tianfang Xiao; Ning Liu; Chen Zhang; Zhan Qin; Kui Ren; | code |
| 1006 | GIPO: Gaussian Importance Sampling Policy Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Theoretical analysis shows that GIPO introduces an implicit, tunable constraint on the update magnitude, while concentration bounds guarantee robustness and stability under finite-sample estimation. |
Chengxuan Lu; Zhenquan Zhang; Shukuan Wang; Qunzhi Lin; Baigui Sun; Yang Liu; | code |
| 1007 | Responsible Text-to-Image Diffusion: Interpretable and Linearly Controllable Semantics for Fair and Safe Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present a model agnostic framework for discovering interpretable and linearly controllable semantic attributes across any T2I DMs backbone. |
Sayedmoslem Shokrolahi; Jae-Mo Kang; Il-Min Kim; | code |
| 1008 | Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To reduce their latency while maintaining quality, we propose a universal acceleration framework for VAE decoders that preserves full alignment with the original latent distribution. |
Lunjie Zhu; Yushi Huang; Xingtong Ge; Yufei Xue; Zhening Liu; Yumeng Zhang; Zehong Lin; Jun Zhang; | code |
| 1009 | Persona-Pruner: Sculpting Lightweight Models for Role-Playing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we question the necessity of dedicating a full, generalist model to a single persona, hypothesizing that a specific character identity relies on only a fraction of the model’s total capacity. |
Jinsu Kim; Jihoon Tack; Noah Lee; Jongheon Jeong; | code |
| 1010 | Colorful Pinball: Density-Weighted Quantile Regression for Conditional Guarantee of Conformal Prediction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Leveraging a Taylor expansion, we derive a sharp surrogate objective for quantile regression: a density-weighted pinball loss, where the weights are given by the conditional density of the conformity score evaluated at the true quantile. We propose a three-headed quantile network that estimates these weights via finite differences using auxiliary quantile levels at $1-\alpha \pm \delta$, subsequently fine-tuning the central quantile by optimizing the weighted loss. |
Qianyi Chen; Bo Li; | code |
| 1011 | Crowd4D: Scene-Aware Monocular 4D Crowd Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Crowd4D, the first scene-aware 4D crowd reconstruction framework that jointly optimizes the crowd and scene from a monocular RGB video in large-scale scenes. |
Hongbo Kang; Tianyi Zhou; Qingyang Yang; Hongwei wen; Jing Huang; Yu-Kun Lai; Kun Li; | code |
| 1012 | Hugging Carbon: Quantifying The Training Carbon Emissions of AI Models at Scale Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a FLOPs-based framework to estimate aggregate training emissions of HF open-source models. |
Xinlei Wang; Ruibo Ming; Jing Qiu; Junhua Zhao; Jinjin Gu; | code |
| 1013 | Graph Alignment for Benchmarking Graph Neural Networks and Learning Positional Encodings Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a novel benchmarking methodology for graph neural networks (GNNs) based on the graph alignment problem, a combinatorial optimization task that generalizes graph isomorphism by aligning two unlabeled graphs to maximize overlapping edges. |
Adrien Lagesse; Marc Lelarge; | code |
| 1014 | View Space: Learning Representation Across Arbitrary Graphs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Generalizing pretrained models to unseen datasets without retraining is a central challenge toward foundation models. |
Dooho Lee; Myeong Kong; Minho Jeong; Jaemin Yoo; | code |
| 1015 | Motion-Aware Caching for Efficient Autoregressive Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This oversight is critical: pixels with high motion require more denoising steps to prevent error accumulation, while static pixels tolerate aggressive skipping. We formalize this insight theoretically by linking cache errors to residual instability, and propose $\textbf{MotionCache}$, a motion-aware cache framework that exploits inter-frame differences as a lightweight proxy for pixel-level motion characteristics. |
Jing Xu; Yuexiao Ma; Songwei Liu; Xuzhe Zheng; Shiwei Liu; Chenqian Yan; Xiawu Zheng; Rongrong Ji; Fei Chao; WANG; | code |
| 1016 | RePack Then Refine: Efficient Diffusion Transformers with Vision Foundation Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Repack then Refine, a three-stage framework that brings the semantic-rich VFM features to DiT while further accelerating learning efficiency. |
Guanfang Dong; Luke Schultz; Negar Hassanpour; Chao Gao; | code |
| 1017 | FiRE: Fine-grained Ranking Evaluation for Machine Translation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a Fine-Grained Ranking Evaluation method (FiRE) that leverages off‑the‑shelf large language models to perform criterion‑driven pairwise comparison across three complementary dimensions: faithfulness, fluency, and consistency of style, instead of producing a single holistic judgment. |
Wenyang Gao; Yinghao Yang; Xi Jin; Jing Li; Yue Zhang; | code |
| 1018 | Interaction-Breaking Adversarial Learning Framework for Robust Multi-Agent Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Prior robust MARL methods have primarily considered value-oriented attacks, leaving a gap in robustness when interaction structures themselves are corrupted. In this paper, we propose an interaction-breaking adversarial learning (IBAL) framework that takes an information-theoretic view to construct attacks that impede coordination by perturbing agents’ observations and actions, and trains agents to perform reliably under such disruptions. |
Sunwoo Lee; Mingu Kang; Yonghyeon Jo; Seungyul Han; | code |
| 1019 | BioDynaSpec: Harmonic-Guided Spatio-Spectral Autoregressive Diffusion for Protein Dynamics Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose **BioDynaSpec**, which reformulates protein dynamics as spatio-spectral generation: **Independent Windowed Fourier Decomposition (IWFD)** decomposes trajectories into independent windowed frequency representations, and a generator combines low-to-high frequency autoregression with diffusion denoising to reconstruct continuous motion. |
Mujie Lin; Yutian Liu; Yudi Guo; Yanzhen Hou; Yiheng Tao; Ruochong Zheng; Kaiwen Cheng; Xin Shan; Youdong Mao; Jie Chen; | code |
| 1020 | MVP-LAM: Learning Action-Centric Latent Action Via Cross-Viewpoint Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose **M**ulti-**V**iew**P**oint **L**atent **A**ction **M**odel (**MVP-LAM**), which learns discrete latent actions that are highly informative about ground-truth actions from time-synchronized multi-view videos. |
Jung Min Lee; Dohyeok Lee; Seokhun Ju; Taehyun Cho; Jin Koo; Li Zhao; Sangwoo Hong; Jungwoo Lee; | code |
| 1021 | ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose ProEval, a proactive evaluation framework that leverages transfer learning to efficiently estimate performance and identify failure cases. |
Yizheng Huang; Wenjun Zeng; Aditi Kumaresan; Zi Wang; | code |
| 1022 | SAEs-BrainMap: Unveiling The Emergence of Specialized Concepts in Deep Models Via Brain Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose SAEs-BrainMap, a novel framework that utilizes human brain activation patterns from the ventral visual pathway as objective probes to guide the identification of features decomposed by Sparse Autoencoders (SAEs). |
Ziming Mao; Jia Xu; Wenxuan Pan; Mufan Xue; Yaochu Jin; Guoyuan Yang; | code |
| 1023 | Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Grounded Correspondence, a framework that replaces learned transition functions with deterministic bipartite matching. |
Zhiyuan Li; Rongzhen Zhao; Wenyan Yang; Wenshuai Zhao; Pekka Marttinen; Joni Pajarinen; | code |
| 1024 | VT-Bench: A Unified Benchmark for Visual-Tabular Multi-Modal Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce \textit{VT-Bench}, the first unified benchmark for standardizing vision-tabular discriminative prediction and generative reasoning tasks. |
贾 子怡; Zijian Cheng; Xinyue Zhang; Kun-Yang Yu; Zhi Zhou; Yu-Feng Li; Lan-Zhe Guo; | code |
| 1025 | Collaborative Threshold Watermarking Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce $(t,K)$-threshold watermarking: clients collaboratively embed a single watermark during training, while only coalitions of at least $t$ clients can reconstruct the watermark key and verify a suspect model, but any coalition of fewer than $t$ clients learns nothing about the watermark beyond the verification output. |
Tameem Bakr; Anish Ambreth; Nils Lukas; | code |
| 1026 | Benchmarking and Enhancing VLM for Compressed Image Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce the first comprehensive benchmark to evaluate the ability of VLM against compressed images, varying existing widely used image codecs and diverse set of tasks, encompassing over one million compressed images in our benchmark. |
Zifu Zhang; Tongda Xu; Siqi Li; Shengxi Li; Yue Zhang; Mai Xu; Yan Wang; | code |
| 1027 | Lookahead Sample Reward Guidance for Test-Time Scaling of Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To further improve efficiency, we introduce lookahead sampling to collect marginal samples. |
Yeongmin Kim; Donghyeok Shin; Byeonghu Na; Minsang Park; Richard Kim; IL CHUL MOON; | code |
| 1028 | Mitigating Reward Hacking in LLM-based Recommendation: A Preference Optimization Approach Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Under the Bradley–Terry model, we further show that these regions can occupy a substantial fraction of the preference space, inevitably leading to misaligned rankings. To address this issue, we propose Simulated Preference Optimization for Reward-hacking mitigation using Pseudo-negatives (SIRIUS). |
Heyu Chen; Junkang Wu; Guoqing Hu; Kexin Huang; Xiang Wang; Jiancan Wu; | code |
| 1029 | Structured Progressive Knowledge Activation for LLM-Driven Neural Architecture Search Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, in practice, seemingly local revisions often propagate into non-local behavioral and performance shifts because a single edit can inadvertently couple multiple interacting functional factors, a phenomenon we refer to as functional entanglement. To make LLM knowledge usable under such entanglement, we propose Structured Progressive Knowledge Activation (SPARK), which activates relevant priors by explicitly selecting the functional factor to modify and conditioning the edit on that factor. |
Zhen Liu; Yuhan Liu; Jinjun Wang; Wei Song; Jianyi Liu; Jingwen Fu; | code |
| 1030 | Physics-informed Diffusion Models in Spectral Space Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a methodology that combines generative latent diffusion models with physics-informed machine learning to generate solutions of parametric partial differential equations (PDEs) conditioned on partial observations, which includes, in particular, forward and inverse PDE problems. |
Davide Gallon; Philippe von Wurstemberger; Patrick Cheridito; Arnulf Jentzen; | code |
| 1031 | Modeling Spectral Energy Shifts in Spatio-Temporal Graph Anomaly Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this limitation, we propose a node-level spectral energy formulation that is fully compatible with message passing and enables the detection of camouflaged anomalies. Building on this formulation, we introduce an energy-aware graph learning framework that models spectral shifts through energy-driven message passing in both static and time-series graphs. |
Yilin Liu; Hongchao Zhang; Ahmad Taha; Taylor T Johnson; Meiyi Ma; | code |
| 1032 | AES: Curing Optimizer Blindness in Long-Tailed Recognition Via State-Aware Correction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing strategies relying on static frequency-based priors fail to correct this bias and result in state blindness regarding supervision and micro-level blindness regarding parameter updates. To address these limitations, we propose the AES framework to establish a dynamic and state-aware correction system across the entire learning lifecycle. |
Fanfu Wang; Jiachang Zhan; Zhiheng Gong; Pengkun Wang; Yang Wang; | code |
| 1033 | TG-RAG: A Retrieval-Augmented Framework for Reasoning Guidance in Specialized Domains Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While recent industrial frameworks attempt to encapsulate Standard Operating Procedures into modular skills for dynamic retrieval, utilizing them via context engineering often proves insufficient for complex workflows, leading to Cognitive Drift. To mitigate this, we propose $\textbf{Thought Guidance-Retrieval Augmented Generation (TG-RAG)}$, a Retrieval-Augmented framework that effectively steers the generation process without relying solely on the model’s self-correction. |
Liang Su; Mingyang Zhang; Yun Xiong; Tengfei LIU; Siwei Zhang; Xi Chen; Li Sun; | code |
| 1034 | Learning to Think in Physics: Breaking Shortcut Learning in Scientific Diffusion Via Representation Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce \textbf{REPA-P}, a \emph{teacher-free} physics-informed representation alignment framework that uses first-principles residuals as supervision. |
Jia Haozhe; Pengyu Yin; Wenshuo Chen; Shaofeng Liang; Lei Wang; Bowen Tian; Xiucheng Wang; Jia Nanqian; Yutao Yue; | code |
| 1035 | Advancing LLM Reasoning with Natural Language and Numerical Feedback Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We show that plateaued RL models can successfully refine failed solutions when given natural language critiques. Motivated by this, we propose Critique-GRPO, an online RL framework that integrates both natural language and numerical feedback for policy optimization. |
Xiaoying Zhang; Yipeng Zhang; Hao Sun; Kaituo Feng; Chaochao Lu; Chao Yang; Helen M Meng; | code |
| 1036 | SAD-Flower: Flow Matching for Safe, Admissible, and Dynamically Consistent Planning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Moreover, existing FM planners do not ensure the dynamical consistency, which potentially renders trajectories inexecutable. We address these shortcomings by proposing SAD-Flower, a novel framework for generating \textbf{S}afe, \textbf{A}dmissible, and \textbf{D}ynamically consistent trajectories. |
Tzu-Yuan Huang; Armin Lederer; Dai-Jie Wu; Xiaobing Dai; Sihua Zhang; Hsiu-Chin Lin; Shao-Hua Sun; Stefan Sosnowski; Sandra Hirche; | code |
| 1037 | Fedfit: Federated Dynamic Pruning Via Fisher Information Scoring Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While federated dynamic pruning aims to alleviate these costs by adjusting sparse topologies during training, existing methods rely on magnitude-based heuristics that are fundamentally ill-suited for the non-convergent, heterogeneous environments inherent to FL. To address this challenge, we propose Fedfit, a federated dynamic framework that replaces simple heuristics with optimization-centric criteria for topology adjustment. |
Meng Bi; Hong Huang; Jinlong Song; Charles Wang; Chengming Hu; Xi Chen; Ting Yu; Xue Liu; | code |
| 1038 | Push, Pop, Parallelize: Stack-Augmented Linear Attention Via The Delta Rule Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, due to their fixed-size state, these models fundamentally struggle to capture the recursive, hierarchical structures that are intrinsic to natural languages. To bridge this gap, we introduce DeltaStack, a novel architecture that augments the associative memory of DeltaNet with a lightweight, differentiable stack. |
Anh Nguyen; Saleh Momeni; Ashutosh Chaubey; Changnan Xiao; Bing Liu; | code |
| 1039 | SciPredict: Can LLMs Predict The Outcomes of Scientific Experiments in Natural Sciences? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce SciPredict, a benchmark comprising 405 tasks derived from recent empirical studies in 33 specialized sub-fields of physics, biology, and chemistry. |
Udari Sehwag; Elaine Lau; Haniyeh Oskouie; Shayan Shabihi; Erich Liang; Andrea Toledo; Guillermo Mangialardi; Sergio Fonrouge; Ed-Yeremai Hernandez-Cardona; Paula Vergara; Utkarsh Tyagi; Chen Bo Calvin Zhang; Pavi Bhatter; Nicholas Johnson; Furong Huang; Ernesto Montoya; Bing Liu; | code |
| 1040 | Learn-to-learn on Arbitrary Textual Conditioning: A Hypernetwork-Driven Meta-gated LLM Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we activate the meta-signal of $\beta$ within the SwiGLU blocks, resulting a meta-gating mechanism which adaptively adjusts the nonlinearity of FFN. |
Luo Ji; Qi Qin; Ningyuan Xi; Teng Chen; Qingqing Gu; Hongyan Li; | code |
| 1041 | JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Human experts address this dilemma by combining domain-grounded principles with dynamic, claim-level assessment. Inspired by this process, we propose JADE, a two-layer evaluation framework. |
Lanbo Lin; Jiayao Liu; Tianyuan Yang; Li Cai; Yuanwu Xu; Lei Wei; TT; Guannan Zhang; | code |
| 1042 | TEFormer: Structured Bidirectional Temporal Enhancement Modeling in Spiking Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing Spiking Transformers still lack a principled mechanism for effective temporal fusion, limiting their ability to fully exploit spatiotemporal dependencies. Inspired by feedforward–feedback modulation in the human visual pathway, we propose **TEFormer**, the first Spiking Transformer framework that achieves bidirectional temporal fusion by decoupling temporal modeling across its core components. |
Sicheng Shen; Mingyang Lv; Bing Han; Dongcheng Zhao; Guobin Shen; Feifei Zhao; Yi Zeng; | code |
| 1043 | Helpful to A Fault: Measuring Illicit Assistance in Multi-Turn, Multilingual LLM Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce **STING** (*Sequential Testing of Illicit N-step Goal execution*), an automated red-teaming framework that constructs a step-by-step illicit plan grounded in a benign persona and iteratively probes a target agent with adaptive follow-ups, using judge agents to track phase completion. |
Nivya Talokar; Ayush Tarun; Murari Mandal; Maksym Andriushchenko; Antoine Bosselut; | code |
| 1044 | D$^3$: Dynamic Directional Graph-Constrained Data Scheduling for LLM Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We then demonstrate that commonly used empirical data ordering heuristics are suboptimal from an optimization perspective. To resolve this, we propose xxx, a data scheduling framework grounded in gradient interactions between samples, where training dependencies are modeled as a graph that explicitly constrains valid training orders. |
Xu Yuanjian; Jianing Hao; Guang Zhang; Zhong Li; | code |
| 1045 | GOCM: Single-Step Graph Outlier Synthesis Via Origin Consistency Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Supervised Graph Outlier Detection has long been constrained by severe class imbalance, and although recent diffusion-based augmentation methods have improved sample quality, their practical utility is hindered by the high computational costs of multi-step iterative sampling and the stochasticity of the generation process. To overcome these bottlenecks, we propose Graph Outlier Synthesis via Origin Consistency Model (GOCM), a single-step graph outlier synthesis framework based on a consistency model. |
一帆 李; Zhihui Wang; Changmiao Wang; Guangxiao Ma; Peng Zhang; | code |
| 1046 | MORE: A Multilingual Document Parsing Benchmark and Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While recent Vision-Language Models (VLMs) claim support for hundreds of languages, the lack of comprehensive ground truth makes it impossible to empirically verify these capabilities. To bridge this gap, we introduce $\textbf{MORE}$, a large-scale, linguistically comprehensive benchmark designed for rigorous multilingual document parsing evaluation. |
LongXu; Binghong Wu; TingHao YU; Hao Feng; zhenyuhuang; Haoqing Jiang; Yunhao Wang; Shuo Huang; feng zhang; | code |
| 1047 | NaRA: Noise-Aware LoRA for Parameter-Efficient Fine-Tuning of Diffusion LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Consequently, they ignore the intrinsic dynamics of the diffusion process, where input distributions and generation difficulty shift significantly along the denoising trajectory, rendering them suboptimal for dLLMs. To address this, we propose **N**oise-**a**ware Low-**R**ank **A**daptation (NaRA), which introduces a low-rank core matrix generated by a lightweight, globally shared hypernetwork conditioned on the noise level. |
Shuaidi Wang; Zhan Zhuang; HUANG Ruping; Yu Zhang; | code |
| 1048 | LECTOR: Joint Learning of Scientific Reasoning Graphs and Introduction Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We conduct a dataset from Nature Communications papers to assess our method. |
Jiabei Xiao; Yizhou Wang; Chen Tang; Pengze Li; Wanli Ouyang; SHIXIANG TANG; | code |
| 1049 | Reconstructing Template-Memorized Images from Natural Prompts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present a low-resource reconstruction attack that operates through seemingly benign prompts and requires little to no access to the training data. |
Sol Yarkoni; Mahmood Sharif; Roi Livni; | code |
| 1050 | Vision Transformer Finetuning Benefits from Non-Smooth Components Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we analyze the ability of vision transformer components to adapt their outputs to changes in inputs, or, in other words, their *plasticity*. |
Ambroise Odonnat; Chapel Laetitia; Romain Tavenard; Ievgen Redko; | code |
| 1051 | Cascaded Flow Matching for Heterogeneous Tabular Data with Mixed-Type Features Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The low-resolution representation of numerical features explicitly accounts for discrete outcomes, such as missing or inflated values, and therewith enables a more faithful generation of mixed-type features. We formally prove that this cascade tightens the transport cost bound. |
Markus Mueller; Kathrin Gruber; Dennis Fok; | code |
| 1052 | Symbal: Detecting Systematic Misalignments in Model-Generated Captions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our work focuses on a class of captioning errors that we refer to as systematic misalignments, where a recurring error in MLLM-generated captions is closely associated with the presence of a specific visual feature in the paired image. Given a vision-language dataset with MLLM-generated captions, our aim in this work is to detect such errors, a task we refer to as systematic misalignment detection. |
Maya Varma; Jean-Benoit Delbrouck; Sophie Ostmeier; Akshay Chaudhari; Curtis Langlotz; | code |
| 1053 | MFH-NAS:A Hybrid Neural Architecture Search Framework for Multimodal Fusion Object Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose MFH-NAS, a hybrid neural architecture search framework that automatically discovers fusion architectures to better leverage cross-modal complementarity. |
QuanWei Gao; Shuqi Zhao; Ruyu Wang; Shuyin Zhang; Cong Liu; Zirui Luo; | code |
| 1054 | Scaling Vision Transformers for Functional MRI with Flat Maps Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a simple strategy for training a foundation model on functional MRI (fMRI) data: we adapt the standard Vision Transformer to fMRI by first converting each 3D fMRI volume to a 2D map using a standard cortical flat map projection. |
Connor Lane; Ratna Grandhi; Leema Krishna Murali; Mihir Tripathy; Shamus Yang; Will Beddow; Gianfranco Cortes; Suin Cho; Debojyoti Das; Sam Gijsen; Manish Ram; Utkarsh Singh; Cesar Kadir Torrico Villanueva; YUXIANG WEI; Daniel Kaplan; Benjamin Warner; Tanishq Abraham; Paul Scotti; | code |
| 1055 | Learning Flexible Generalization in Video Quality Assessment By Bringing Device and Viewing Condition Distributions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Factors such as ambient lighting, display brightness, and resolution significantly influence the visibility of distortions. In this work, we address the question of the multi-screen quality assessment on mobile devices, as this area still tends to be under-covered. |
Nickolay Safonov; Dmitriy Vatolin; | code |
| 1056 | S$^3$GNN: Efficient Global Mixing and Local Message Passing for Long-Range Graph Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Alongside spatial connectivity enrichment (e.g., rewiring), recent studies have shown that spectral filtering can yield strong long-range learning outcomes, as spectral operators enable global information mixing that alleviates OSQ. These approaches achieve this either by stabilizing the Jacobian energies in deep propagation or by guaranteeing OSQ mitigation under strong theoretical assumptions. |
Dai Shi; Linhan Luo; Luke Thompson; Lequan Lin; Andi Han; Junbin Gao; Jose Miguel Hernandez-Lobato; | code |
| 1057 | Large Language Model Teaches Visual Students: Cross-Modality Transfer of Fine-Grained Conceptual Knowledge Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose \LaViD—Language-to-Visual Knowledge Distillation—a simple and effective framework for transferring high-level semantic knowledge from a language-only teacher to a vision-only student model. |
Thomas Liang; Zhuoran Yu; Yong Jae Lee; | code |
| 1058 | SG2Loc: Sequential Visual Localization on 3D Scene Graphs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introduces a novel, lightweight approach to sequential visual localization using 3D scene graphs. |
Nicole Damblon; Olga Vysotska; Federico Tombari; Marc Pollefeys; Daniel Barath; | code |
| 1059 | Markovian Projection of Star-Shaped Diffusion for Exponential Family Distributions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Star-shaped diffusion addresses the former by introducing a non-Markovian forward process, yet this comes at the expense of temporal coherence in the reverse process. We propose a novel framework that resolves this trade-off by learning a Markovian projection of a star-shaped forward process, and its reversal. |
François Bertholom; Khalid Oublal; | code |
| 1060 | RQ-MoE: Residual Quantization Via Mixture of Experts for Efficient Input-Dependent Vector Compression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Residual Quantization via Mixture of Experts (RQ-MoE), a framework combining a two-level MoE with dual-stream quantization to enable input-dependent codebook adaptation for efficient vector quantization. |
Zhengjia Zhong; Shuyan Ke; zaizhou lin; Jiaqi Song; Hongyi Lan; Hui Li; | code |
| 1061 | From Basis to Basis: Gaussian Particle Representation for Interpretable PDE Operators Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose to represent fields with a \emph{Gaussian basis}, where learned atoms carry explicit geometry (centers, anisotropic scales, weights) and form a compact, mesh-agnostic, directly visualizable state. |
LI; Yu Feng; Zhilu Lai; Wei Wang; | code |
| 1062 | SimGFM: Simplifying Discrete Flow Matching for Graph Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Discrete Flow Matching (DFM) presents a promising approach for graph generation; however, existing adaptations often introduce substantial complexity by incorporating task-specific heuristics, compromising the continuity equation and significantly expanding the hyperparameter space. |
Chunyu Luo; Yuankai Luo; Xiao-Ming Wu; Lei Shi; | code |
| 1063 | Escaping Whack-a-Mole: Code Documentation Optimization Via Dependency-Guided Bi-level Search Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose DocSearch, a dependency-guided bi-level search framework that systematically exploits test-time feedback. |
Yutong Cheng; Haifeng Chen; Wenchao Yu; Xujiang Zhao; Peng Gao; Wei Cheng; | code |
| 1064 | Phase-Aware Mixture of Experts for Agentic Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Phase-Aware Mixture of Experts (PA-MoE). |
Yang Shengtian; Yu Li; Shuo He; Yewen Li; Qingpeng Cai; Peng Jiang; Lei Feng; | code |
| 1065 | Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Agent World Model (AWM), a fully synthetic environments generation pipeline. |
Zhaoyang Wang; Boyi Liu; Yite Wang; Siwei Han; Zhewei Yao; Yuxiong He; Huaxiu Yao; Canwen Xu; | code |
| 1066 | From Per-Image Low-Rank to Encoding Mismatch: Rethinking Feature Distillation in Vision Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose two minimal remedies, Lift or WideLast: (i) Lift retains a lightweight lifting projector at inference to provide wider channel, or (ii) WideLast widens only the student’s last block, enabling an input-dependent expansion. |
Huiyuan Tian; Bonan Xu; Shijian Li; | code |
| 1067 | Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce ATLAS, an adaptive testing framework based on Item Response Theory (IRT) that estimates model ability using Fisher information–guided item selection. |
Peiyu Li; Xiuxiu Tang; Si Chen; Cheng; Ronald Metoyer; Ting Hua; Nitesh Chawla; | code |
| 1068 | SimpleGPT: Improving GPT Via A Simple Normalization Strategy Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we revisit Transformer optimization through the lens of second-order geometry and establish a direct connection between architectural design, activation scale, the Hessian matrix, and the maximum tolerable learning rate. |
Marco Chen; Xianbiao Qi; Yelin He; Jiaquan Ye; Rong Xiao; | code |
| 1069 | Coupled Trigger Optimization and Vulnerable Parameter Alignment for Persistent Backdoor Attacks on Federated Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we view FL backdoor persistence through the lens of optimization dynamics, and argue that long-lasting attacks require alignment between trigger-induced representations and aggregation-stable parameter directions. |
ma zhixuan; Haichang Gao; Shangwen Li; Ping Wang; Han Yu; | code |
| 1070 | FedRot-LoRA: Mitigating Rotational Misalignment in Federated LoRA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: When such misaligned factors are averaged directly, they interfere destructively and degrade the global update. To address this issue, we propose **FedRot-LoRA**, a federated LoRA framework that aligns client updates via orthogonal transformations prior to aggregation. |
Haoran Zhang; Dongjun Kim; Seohyeon Cha; Haris Vikalo; | code |
| 1071 | Strategy Executability in Mathematical Reasoning: Leveraging Human–Model Differences for Effective Guidance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Through a controlled analysis of paired human-written and model-generated solutions, we identify a systematic dissociation between usage and executability: human- and model-derived strategies differ in structured, domain-dependent ways, leading to complementary strengths and consistent source-dependent reversals under guidance. Building on this diagnosis, we propose *Selective Strategy Retrieval* (SSR), a test-time framework that explicitly models executability by selectively retrieving and combining strategies using empirical, multi-route, source-aware signals. |
Weida Liang; Yiyou Sun; Shuyuan Nan; Chuang Li; Dawn Song; Kenji Kawaguchi; | code |
| 1072 | Diversity Over Frequency: Rethinking Tool Use in Visual Chain-of-Thought Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Prior work has primarily demonstrated the effectiveness of these tools on visual search tasks, leaving their applicability to more diverse and complex visual problems underexplored. In this paper, we move beyond visual search and study challenging visual tasks that require advanced spatial understanding and reasoning, such as 3D spatial reasoning, where agents must not only crop or zoom in on relevant regions but also understand how these local details relate to the global context. |
Dong-Hee Kim; Reuben Tan; DONGHYUN KIM; | code |
| 1073 | TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce the Totally Accelerated Battle Simulator in JAX (TABX), a high-throughput sandbox designed for reconfigurable multi-agent tasks. |
Hayeong Lee; JunHyeok Oh; Byung-Jun Lee; | code |
| 1074 | IRPM: Intergroup Relative Preference Modeling for Pointwise Generative Reward Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, widely used pairwise GRMs create a computational bottleneck in reinforcement learning from human feedback (RLHF), when calibrating or aggregating preference signals over $n$ candidates, often incurring $\mathcal{O}(n^2)$ pairwise judgments. To address this issue, we propose Intergroup Relative Preference Modeling (IRPM), an RL-based method that extends the Bradley–Terry preference-learning paradigm via intergroup comparisons to train \emph{pointwise} GRMs from pairwise preference data. |
Haonan Song; Qingchen Xie; Huan Zhu; Feng Xiao; Luxi Xing; Liu Kang; Fuzhen Li; Zhiyong Zheng; Feng Jiang; Ziheng Li; Kun Yan; Qingyi Si; Yanghua Xiao; Hongcheng Guo; Fan Yang; | code |
| 1075 | Towards Efficient LLMs Annealing with Principled Sample Selection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We demonstrate that effective annealing requires balancing global Hessian geometry with sample-wise gradient noise, navigating a landscape of highly anisotropic curvature. Based on these insights, we formulate sample selection as a constrained optimization problem to suppress noise in sharp directions while preserving descent signals in flat subspaces. |
Xu Yuanjian; Jianing Hao; Wanbo Zhang; Zhong Li; Guang Zhang; | code |
| 1076 | Genome-Factory: A Library for Tuning, Deploying, and Interpreting Genomic Foundation Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Genome-Factory, the first integrated Python library for tuning, deploying, and interpreting genomic foundation models. |
Weimin Wu; Xuefeng Song; Yibo Wen; Qinjie Lin; Zhihan Zhou; Jerry Yao-Chieh Hu; Zhong Wang; Han Liu; | code |
| 1077 | Low Kruskal-Rank Adaptation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we revisit the notion of rank for LoRA update matrices and show that the standard matrix rank fails to capture duplicated directions and redundancy in the update subspace. |
Yixing Xu; Guanchen Li; Chao Li; Xuanwu Yin; Dong Li; Spandan Tiwari; Ashish Sirasao; Emad Barsoum; | code |
| 1078 | TimeSpot: Benchmarking Geo-Temporal Understanding in Vision–Language Models in Real-World Settings Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although recent vision–language models (VLMs) have made progress in image geo-localization using salient cues like landmarks or road signs, their ability to reason about temporal signals and physically grounded spatial cues remains underexplored. To address this gap, we introduce ***TimeSpot***, a benchmark for evaluating real-world geo-temporal reasoning in VLMs. |
Azmine Toushik Wasi; Shahriyar Ridoy; Koushik Tonmoy; Kinga Tshering; S. Hasan; Wahid Faisal; Tasnim Mohiuddin; Md Rizwan Parvez; | code |
| 1079 | Signature-Informed Transformer for Asset Allocation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We argue this creates a fundamental mismatch where minimizing prediction errors fails to yield robust portfolios. We propose the Signature Informed Transformer to address this by unifying feature extraction and decision making into a single policy. |
Yoontae Hwang; Stefan Zohren; | code |
| 1080 | OPTION: Optimal Transport–Guided Flow Matching for Incomplete and Unaligned Multi-View Clustering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Multi-view clustering effectively exploits rich information from multiple views, yet real-world applications are frequently challenged by missing views and cross-view sample misalignment, hindering cross-view modeling and resulting inferior clustering performance. To address these challenges, this paper presents a novel method, **OP**timal **T**ransport–gu**I**ded fl**O**w matchi**N**g for incomplete and unaligned multi-view clustering (**OPTION**). |
Siyuan Zhou; Zhibin Gu; | code |
| 1081 | Safety Generalization Under Distribution Shift in Safe Reinforcement Learning: A Diabetes Testbed Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Safe Reinforcement Learning (RL) algorithms are typically evaluated under fixed training conditions. |
Minjae Kwon; Josephine Lamp; Lu Feng; | code |
| 1082 | STAR: Rethinking MoE Routing As Structure-Aware Subspace Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose STAR, a STructure-Aware Routing that rethinks MoE routing as a subspace learning problem by augmenting standard learnable routing with an evolving principal subspace that tracks dominant input structure via Generalized Hebbian Algorithm (GHA). |
Sumin Park; Noseong Park; | code |
| 1083 | Q-Delta: Beyond Key–Value Associative State Evolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We show that query-conditioned state readout induces a structured value prediction over accumulated memory that complements key-based retrieval. Based on this insight, we propose Q-Delta, a query-aware delta rule that integrates mixed key–query prediction errors into state evolution, enabling jointly corrective dynamics while preserving delta-rule efficiency. |
Sumin Park; Seojin Kim; Noseong Park; | code |
| 1084 | Tempora: Characterising The Time-Contingent Utility of Online Test-Time Adaptation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As ML increasingly underpins latency-sensitive and user-facing use-cases, temporal pressure constrains the viability of adaptable inference; predictions arriving too late to act on are futile. We introduce *Tempora*, a framework for evaluating TTA under this pressure. |
Sudarshan Sreeram; Young Kwon; Cecilia Mascolo; | code |
| 1085 | MuCO: Generative Peptide Cyclization Empowered By Multi-stage Conformation Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this study, we propose MuCO (Multi-stage Conformation Optimization), a generative peptide cyclization method that models the distribution of cyclic peptide conformations conditioned on the corresponding linear peptide. |
Yitian Wang; Fanmeng Wang; Angxiao Yue; Wentao Guo; Yaning Cui; Hongteng Xu; | code |
| 1086 | Efficient Synthetic Network Generation Via Latent Embedding Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce Synthetic Network Generation via Latent Embedding Reconstruction (SyNGLER), a general and efficient framework for synthetic network generation that builds on latent space network models. |
Feifan Jiang; Yinan Bu; Shihao Wu; Gongjun Xu; Ji Zhu; | code |
| 1087 | CGSVD: Cascaded Granular Singular Value Decomposition for Large Language Model Compression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing training-free approaches predominantly rely on uniform rank allocation, implicitly assuming homogeneous redundancy across the model depth and thereby neglecting the inherent non-uniformity of representational evolution. To bridge this gap, we introduce \textbf{CGSVD}, a \uline{\textbf{C}}ascaded \uline{\textbf{G}}ranular \uline{\textbf{S}}ingular \uline{\textbf{V}}alue \uline{\textbf{D}}ecomposition framework that leverages a dual-level non-uniform allocation strategy to maximize semantic preservation. |
Yuli Chen; Shuhao Zhang; Jiale Han; Fanshen Meng; Haishen Jiang; Bo Cheng; Qiang Tong; Xiulei Liu; | code |
| 1088 | Learn from A Rationalist: Distilling Intermediate Interpretable Rationales Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To improve the predictive performance of RE models that are based on less capable or smaller neural networks (i.e., the students), we propose **REKD** (**R**ationale **E**xtraction with **K**nowledge **D**istillation) where a student RE model learns from the rationales and predictions of a teacher (i.e., a *rationalist*) in addition to the student’s own RE optimization. |
Jiayi Dai; Randy Goebel; | code |
| 1089 | Stochastic Sparse Attention for Memory-Bound Inference Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Stochastic Additive No-mulT Attention (SANTA), a method that sparsifies value-cache access by sampling $S \ll n_k$ indices from the post-softmax distribution and aggregates only those value rows. |
Kyle Lee; Corentin Delacour; Kevin Callahan-Coray; Kyle Jiang; Can Yaras; Samet Oymak; Tathagata Srimani; Kerem Camsari; | code |
| 1090 | Shortcut-Resistant CAM Distillation for Long-Tailed Recognition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We find that this tendency is amplified for tail classes: limited examples often share similar contexts, making non-semantic signals highly correlated and thus tempting shortcuts, whereas head classes with diverse appearances and environments encourage more stable object-focused representations. Motivated by this observation, we propose Shortcut-Resistant CAM Distillation (SRCD), a plug-and-play framework that transfers object-focused explanations from head to tail classes. |
Wenhai Wan; Teng Zhang; Shao-Yuan Li; Xinrui Wang; Qiang-Sheng Hua; Songcan Chen; | code |
| 1091 | Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While some approaches like the General Preference Model (GPM) address this, we identify a theoretical limitation: their implicit formulation entangles hierarchy with cyclicity, failing to guarantee dominant solutions. To address this, we propose the Hybrid Reward-Cyclic (HRC) model, which utilizes game-theoretic decomposition to explicitly disentangle preferences into orthogonal transitive (scalar) and cyclic (vector) components. |
Yucong Huang; Xiucheng Li; Kaiqi Zhao; Jing Li; | code |
| 1092 | Language Bias in LVLMs: From In-Depth Analysis to Simple and Effective Mitigation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Yet most analyses remain empirical without uncovering its underlying cause. In this paper, we provide a systematic study of language bias and identify its root in modality misalignment during training. |
Yangneng Chen; Jing Li; | code |
| 1093 | Memory As Dynamics: Learning Reliability-Guided Predictive Models for Online Video Perception Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this issue, we reinterpret video memory as a dynamic latent process rather than a static buffer. Building on this insight, we introduce Reliability-guided Predictive Memory (RPM), a framework that explicitly regulates when and how predictive dynamics should influence online video perception. |
Minwoo Kim; Sang Min Yoon; | code |
| 1094 | TRAP: Hijacking VLA CoT-Reasoning Via Adversarial Patches Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we show that Chain-of-Thought (CoT) reasoning introduces a novel attack vector for targeted control hijacking—for example, causing a robot to mistakenly deliver a knife to a person instead of an apple—without modifying the user’s instruction. |
Zhengxian Huang; Wenjun Zhu; Haoxuan Qiu; Xiaoyu Ji; Wenyuan Xu; | code |
| 1095 | Unifying and Optimizing Data Values for Selection Via Sequential Decision-Making Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To bridge theory and practice, we propose an efficient bipartite graph-based surrogate that preserves submodular structure while enabling scalable greedy selection with provable guarantees. |
Hongliang Chi; Qiong Wu; Zhengyi Zhou; Jonathan Light; Emily Dodwell; Yao Ma; | code |
| 1096 | PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose PipeSD, an efficient cloud-edge collaborative pipeline inference framework with speculative decoding. |
Yunhe Han; Yunqi Gao; Bing Hu; Boloursaz Mashhadi; Yitong Duan; Pei Xiao; Yanfeng Zhang; | code |
| 1097 | Incomplete Multi-View Clustering Via Neighborhood-Conditioned Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: On the other hand, they overlook stable cross-view neighborhood structures, leading to weak structural constraint. To address these limitations, we propose neighborhood-conditioned diffusion for incomplete multi-view clustering (IMVC-NCD), which achieves robust latent completion. |
Qian Guo; Gaohui Zuo; Bingbing Jiang; Guangrui Fan; Zhihua Cui; Xinyan Liang; Jianjian Ding; | code |
| 1098 | Keep Everyone Happy: Online Fair Division of Numerous Items with Few Copies Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose algorithms that model online fair division as a contextual bandit problem, with provable sub-linear regret upper bound guarantees. |
Arun Verma; Indrajit Saha; Makoto Yokoo; Bryan Kian Hsiang Low; | code |
| 1099 | How Hard Can It Be? Hardness-Aware Multi-Objective Unlearning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we quantify how hard it is to reconcile the conflicting objectives arising from overlapping data and provide conditions under which collateral forgetting is unavoidable, that is, when improving forget quality forces retain utility degradation. |
Jiangwei Chen; Xinyuan Niu; Rachael Hwee Ling Sim; Zhengyuan Liu; Nancy Chen; Bryan Kian Hsiang Low; | code |
| 1100 | Geometry-Guided Modeling of Foundation Features Enables Generalizable Object Shape Deformation Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present a generalizable deformation learning framework that reconstructs 3D objects by explicitly deforming a category-level shape template to match the target observation. |
YIYAO MA; Kai Chen; Zhongxiang Zhou; Zhuheng Song; Dongsheng Xie; Zelong Tan; Rong Xiong; DOU QI; | code |
| 1101 | Very Efficient Listwise Multimodal Reranking for Long Documents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose ZipRerank, a very efficient listwise multimodal reranker that directly addresses both bottlenecks: it shortens the input via query-image early interaction and eliminates multi-step generation by scoring all candidates in a single forward pass. |
Yiqun Sun; Pengfei Wei; Lawrence Hsieh; | code |
| 1102 | Contractive Anchor Resolvent Diffusion for Incomplete Multi-View Clustering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing graph-based approaches either rely on costly data imputation or adopt first-order linear fusion, which acts as a weak low-pass filter and fails to separate latent consensus structure from structural noise. To address this limitation, we reformulate IMVC from a spectral filtering perspective and propose \textbf{C}ontractive \textbf{A}nchor \textbf{R}esolvent \textbf{D}iffusion (\textbf{CARD}), a scalable framework for high-order structural inference without explicit imputation. |
Tongzheng Zhao; Yangyang Wen; Yukai Shi; Xinyan Liang; Feijiang Li; Peng Zhou; Liang Du; | code |
| 1103 | LagLLM: LLM-empowered Lead–lag Dependency Learning for Spatial-temporal Time Series Forecasting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Thus, we propose LagLLM, the first LLM-empowered framework that explicitly models lead–lag dependencies by unifying data-driven dynamics modeling and knowledge-driven semantic reasoning. |
Binqing Wu; Jian Zhou; Zongjiang Shang; Ling Chen; | code |
| 1104 | From Static Constraints to Dynamic Adaptation: Sample-Level Constraint Release for Offline-to-Online Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This challenge arises because data behavior evolves during fine-tuning, rendering data origin a misleading basis for constraint handling and thereby leading to objective–data mismatch. We therefore propose Dynamic Alignment for RElease (DARE), a distribution-aware framework for sample-level constraint release based on the behavioral consistency with a behavior model. |
Lipeng Zu; YU QIAN; Shayok Chakraborty; Xiaonan Zhang; | code |
| 1105 | MM-Spectrum: Multimodal Multi-spectral Molecular Structural Elucidation with A Stable MoE Framework Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As a remedy, we propose MM-Spectrum, a sparse Mixture-of-Experts framework tailored for multimodal multispectral spectra-to-structure elucidation. |
Hai-tao Yu; Min Nan; Zheng Fang; Hongyu Zhan; Yusen Tan; Yuhan Wang; Jun Xia; | code |
| 1106 | RNA-FM: Flow-Matching Generative Model for Genome-wide RNA-Seq Prediction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose RNA-FM, a flow-matching generative framework for genome-wide bulk RNA-seq prediction from histopathology images. |
Yaxuan Song; Jianan Fan; Tianyi Wang; Qiuyue Hu; Hang Chang; Heng Huang; Weidong Cai; | code |
| 1107 | Sequential Kernel-based Conditional Independence Testing Via Adaptive Betting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a new approach for sequential testing of conditional independence that is far more robust to estimation errors in the conditional distribution. |
Zheng He; Danica J Sutherland; | code |
| 1108 | Gram2Token: Enabling Run-time GPU-Native Grammar-Constrained Decoding for LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In contrast, we propose Gram2Token, which preprocesses grammar constraints into token-level representations that can be executed natively on GPUs at run time, thereby reducing decoding overhead. |
Hantao Hua; Jiming Su; hao tang; Yiping Yao; Feng Zhu; | code |
| 1109 | ParEVO: Synthesizing Code for Irregular Data: High-Performance Parallelism Through Agentic Evolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our contributions include: (1) \textbf{The Parlay-Instruct Corpus}, a curated dataset of 12,000 tasks synthesized via a Critic-Refine pipeline that explicitly filters for theoretically optimal algorithms under the Work-Span cost model; (2) specialized \textbf{DeepSeek}, \textbf{Qwen}, and \textbf{Gemini} models fine-tuned to align probabilistic generation with the rigorous semantics of the ParlayLib intermediate representation; and (3) an \textbf{Evolutionary Coding Agent (ECA)} that solves the “last mile” of correctness by iteratively repairing code using feedback from compilers, race detectors, and performance profilers. |
Liu Yang; Zeyu Nie; Andrew Liu; Ruomu Zou; Deniz Altınbüken; Amir Yazdanbakhsh; Quanquan Liu; | code |
| 1110 | Metric–-Phase Fields: Decoupling Distance and Sign for Thin-Structure Reconstruction from Unoriented Point Clouds Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Metric–-Phase Fields (MPFs), a decoupled implicit representation that separates metric proximity from topological phase. |
Jiayi Kong; Xuhui Chen; Chen Zong; Fei Hou; Junhui Hou; Wenping Wang; Ying He; | code |
| 1111 | Test-Time Graph Search for Goal-Conditioned Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address unreliable value estimates during shortest-path search, we propose a novel mechanism that softly penalizes long-distance transitions. |
Evgenii Opryshko; Junwei Quan; Claas Voelcker; Yilun Du; Igor Gilitschenski; | code |
| 1112 | STAR-KV: Low-Rank KV Cache Compression Via Soft Thresholding for Adaptive Rank Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose STAR-KV, an adaptive low-rank KV cache compression framework with fine-grained rank control. |
Priyansh Bhatnagar; Ashkan Moradirouzabadi; Se-Hyun Yang; SeungJae Lee; Jungwook Choi; Mingu Kang; | code |
| 1113 | DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While large vision-language models (VLMs) achieve strong spatial grounding, their inference cost and latency limit real-world deployment; on the other hand, compact VLMs are efficient but suffer substantial localization degradation under standard fine-tuning or distillation. To address this gap, we propose **DocVAL**, a validated chain-of-thought (CoT) distillation framework that transfers explicit spatial reasoning from large teacher models to compact, deployable student VLMs. |
Pinaki Prasad Guha Neogi; Ahmad Mohammadshirazi; Ser-Nam Lim; Rajiv Ramnath; | code |
| 1114 | RapTB: Rooted Absorbed Trajectory Balance with Submodular Replay for Stable Autoregressive GFlowNet Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Rooted absorbed prefix Trajectory Balance (RapTB), an objective that anchors subtrajectory supervision at the root and propagates terminal rewards to intermediate prefixes via absorbed suffix-based backups, providing dense prefix-level learning signals. |
Xi Wang; Wenbo Lu; Shenji Wan; | code |
| 1115 | Hierarchical ODE: Learning Continuous-Time Physical Prototypes for Early Link Failure Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Furthermore, rigid closed-set assumptions fail to capture unseen diversity. To address these limitations, we propose a hierarchical ordinary differential equation clustering network, which utilizes neural ordinary differential equation to model latent state evolution as a continuous integral curve. |
Jiaen Lv; Leran Qi; Shaowei Wang; | code |
| 1116 | Discrete Diffusion with Physical Mass Constraints for \emph{De Novo} Peptide Sequencing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we introduce $\textbf{PhysNovo}$, a novel paradigm that harnesses discrete diffusion to enable simultaneous global reasoning and iterative refinement. |
An Zeyu; Wanyu LIN; | code |
| 1117 | Fast Non-Episodic Finite-Horizon RL with K-Step Lookahead Thresholding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a modified Q-function: rather than targeting the full-horizon, we learn a K-step lookahead Q-function that truncates planning to the next K steps. |
Jiamin Xu; Kyra Gan; | code |
| 1118 | Beyond Point Predictions: Manifold Expansion and Dual Alignment for Robust Time Series Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, blindly mimicking teacher predictions, which are often uncertain, can induce negative transfer. To address this, we propose Dynamic Structural Distillation (DSD), a robust framework that goes beyond the prediction-matching paradigm. |
Junyao Hong; Zesheng Lai; Xinyi Xiao; Suyang Zhou; Aodong Shen; Youyong Kong; | code |
| 1119 | Vector Linking Via Cross-Model Local Isometric Consistency Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Empirically and theoretically, we show that independently trained contrastive encoders exhibit local geometric consistency: short-range distances are approximately preserved up to a scale factor, while long-range distances are not due to model-specific distortion. Building on this, we propose an iterative, reference-based geometric embedding hashing that recovers vector links from a tiny seed set of paired anchors. |
Ziying Chen; Yang Cao; He Sun; Beining Yang; Tianjian Yang; | code |
| 1120 | STRIDE: Post-Training LLMs to Reason and Refine Bio-Sequences Via Edit Trajectories Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Diffusion-based approaches provide strong progressive refinement but are not naturally aligned with discrete, grammar-constrained edit operations, whereas autoregressive LLMs readily produce valid sequences yet often lack explicit long-horizon planning. To close this gap, we introduce *STRIDE* (Sequence Trajectory Refinement via Internalized Denoising Emulation), a post-training framework that recasts optimization as an intrinsic reasoning problem in edit space. |
Daiheng Zhang; Shiyang Zhang; Sizhuang He; Yangtian Zhang; Syed Rizvi; David van Dijk; | code |
| 1121 | When Softmax Fails at The Top: Extreme‑Value Corrections for InfoNCE Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We confirm this prediction empirically by measuring Weibull style tail behavior in the hardest negatives throughout InfoNCE training. Motivated by this mismatch, we propose WEINCE, a simple modification of InfoNCE that targets the extreme score regime directly. |
Hasan Sabri Melihcan Erol; Suat Evren; Oktay Ozel; Alexander Morgan; Jongha (Jon) Ryu; Lizhong Zheng; | code |
| 1122 | Decentralized Instruction Tuning: Conflict-Aware Splitting and Weight Merging Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose MERIT, a decentralized, merge-ready pipeline that splits mixtures before fine-tuning. |
MINSIK CHOI; Geewook Kim; | code |
| 1123 | EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While LLM-powered multi-agent simulations offer a scalable in silico alternative, current approaches often fail to support rigorous experimentation due to shallow psychological grounding and unquantifiable interactions. To address this, we introduce EduMirror, a multi-agent simulator for the scientific study of educational social dynamics. |
Jingzhe Lin; Hengbin Yu; Yongdan Zeng; Fangwei Zhong; | code |
| 1124 | ReQAT: Achieving Full-Precision Reasoning Accuracy with 4-bit Floating-Point Quantization-Aware Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We identify that FP4 failures concentrate on low-entropy tokens—precise symbolic commitments such as digits and operators—where quantization noise inflates sampling errors that cascade through reasoning traces. Based on this insight, we propose ReQAT, a reasoning-centric FP4 training framework with three components: (i) Trace-Aligned QAT (TAQ), which revisits identical reasoning traces to focus updates on critical low-entropy decisions; (ii) Selective Entropy Minimization (SEM), which reinforces confidence at low-entropy positions; and (iii) Q-FIT, a quantization-friendly initialization that jointly calibrates RoPE-consistent KV cache transformations to stabilize QAT. |
Janghwan Lee; Sihwa Lee; Jinseok Kim; Yongjik Kim; Jieun Lim; Jinwook Oh; Jungwook Choi; | code |
| 1125 | Breaking The Scale Barrier: One-Shot Knowledge Transfer Via Frequency Transform Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In response to this challenge, recent approaches typically resort to either parameter selection, which fails to capture the interdependent structure of this knowledge, or parameter prediction using generative models that depend on impractical access to large network collections. In this paper, we identify the low-frequency components of model weights as the concrete carrier of foundational, task-agnostic knowledge—its learngene—and validate this by demonstrating its efficient inheritance by downstream models and tasks. |
Jianlu Shen; Fu Feng; Yucheng Xie; JIAQI LYU; Xin Geng; | code |
| 1126 | The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As a proof of concept, we propose _head-wise RMSNorm_, an architectural modification that stabilizes value aggregation outputs during pre-training. |
Li Siquan; Kaiqi Jiang; Jiacheng Sun; Tianyang Hu; | code |
| 1127 | Breaking The Factorization Barrier in Diffusion Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We argue that this barrier arises not from limited backbone expressivity, but from a structural misspecification: models are restricted to fully factorized outputs because explicitly parameterizing a joint distribution would require the Transformer to output a prohibitively large number of parameters. We propose **Co**upled **D**iscrete **D**iffusion (**CoDD**), a hybrid framework that breaks this barrier by replacing the fully-factorized output distribution with a lightweight, tractable probabilistic inference layer. |
Ian Li; Zilei Shao; Benjie Wang; Rose Yu; Guy Van den Broeck; Anji Liu; | code |
| 1128 | Towards Efficient and Expressive Offline RL Via Flow-Anchored Noise-conditioned Q-Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Flow-Anchored Noise-conditioned Q-Learning (FAN), a highly efficient and high-performing offline reinforcement learning (RL) algorithm. |
Sungyoung Lee; Dohyeong Kim; Eshan Balachandar; Zelal Mustafaoglu; Keshav Pingali; | code |
| 1129 | AlphaRouter: Token-level Routing Between SLM and LLM with Reinforcement Learning and Tree Search Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we demonstrate that the SLM-LLM collaborative inference space offers a richer solution set, yielding correct answers even when the LLM fails. |
Siteng Liao; Yuzhu Liang; Hengzhong Rao; Xizhao Luo; Tian Wang; | code |
| 1130 | Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Video-MTR, a reinforced multi-turn reasoning framework that operates solely through data-efficient, pure RL post-training. |
Yuan Xie; Tianshui Chen; Zheng Ge; Lionel Ni; | code |
| 1131 | NeuronCtrl: Geometry-Aware Safe Closed-Loop Generative Control for Neuronal Microenvironment Dynamics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce **NeuronCtrl**, a modular operator-level framework for safe, closed-loop generative control of neuronal microenvironment dynamics. |
Haowei Xu; Yixin Chen; Wanyi Fu; Hongbin Han; Zhaoheng Xie; | code |
| 1132 | AmbiRefer3D: 3D Visual Grounding with Referential Ambiguity Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a new task, 3D visual grounding with referential ambiguity, which allows for referential ambiguity in language descriptions, making it more broadly applicable to real-world scenarios. |
Rongjiang Zhu; Wei Kang; Zeqi Liu; Chen junyu; Shuo Yang; Xinxiao Wu; | code |
| 1133 | Learning When to Attend: Conditional Memory Access for Long-Context LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In practice, most tokens do not require Global Attention over the entire sequence and can rely on local context. Based on this insight, we propose L2A, a sequence modeling layer that enables token-wise long-term conditional memory access by deciding \textit{when} to invoke Global Attention. |
Sakshi Choudhary; Aditya Chattopadhyay; Luca Zancato; Elvis Nunez; Matthew Trager; Wei Xia; Stefano Soatto; | code |
| 1134 | Test-Time Debiasing with Probabilistic Prompts Via Wasserstein Distance in Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, such point-based corrections are often unstable and become notably weaker in multi-class settings, where group structure cannot be adequately captured by a single point. Therefore, we propose W4D, a distributional debiasing framework that reframes fairness as aligning query embedding distributions to group reference distributions under the Wasserstein distance, which provides a geometry-aware notion of discrepancy beyond mean shifts. |
Chengye Wang; Yuyuan Li; XiaoHua Feng; Xiaolin Zheng; Chaochao Chen; | code |
| 1135 | Multi-Objective Bayesian Optimization Via Adaptive $\varepsilon$-Constraint Decomposition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work we propose STAGE-BO, Sequential Targeting Adaptive Gap-Filling $\varepsilon$-Constraint Bayesian Optimization, that explicitly targets under-explored regions of the Pareto front. |
Yaohong Yang; Sammie Katt; Samuel Kaski; | code |
| 1136 | MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce the Privacy Perception and Reasoning (PPR) framework and construct a bilingual multimodal dataset with synthetic individual profiles, where identifiers (e.g., faces, names) are linked to sensitive attributes. |
Xiongtao Sun; HUI LI; Jiaming Zhang; Yujie Yang; Kaili Liu; Ruxin Feng; Wen Tan; Wei Lim; | code |
| 1137 | Latent Laplace Diffusion for Irregular Multivariate Time Series Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Irregular multivariate time series pose a fundamental trade-off for long-horizon forecasting: discrete methods can distort temporal structure via re-gridding, while continuous-time models often rely on sequential numerical solvers that are prone to drift. To bridge this gap, we present the Latent Laplace Diffusion (LLapDiff), a generative framework that models the target as a low-dimensional latent trajectory, enabling horizon-wide generation without step-by-step integration over physical time. |
Zinuo You; Jin Zheng; John Cartlidge; | code |
| 1138 | NeurOCNN: A Neural-Operator-Based Model for Physiological Time Series Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose NeurOCNN, a neural-operator-based model for physiological signals that learns a function-to-label mapping while exhibiting discretization invariance. |
Daya Kumar; Uday Devulapalli; Aarat Satsangi; Apurva Narayan; | code |
| 1139 | FlashSinkhorn: IO-Aware Entropic Optimal Transport on GPU Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present **FlashSinkhorn**, an IO-aware EOT solver for squared Euclidean cost that rewrites stabilized log-domain Sinkhorn updates as row-wise LogSumExp reductions of biased dot-product scores, the same normalization as transformer attention. |
Felix Ye; Xingjie Li; An Yu; Ming-Ching Chang; LINSONG CHU; Davis Wertheimer; | code |
| 1140 | Detecting Errors in AI-Generated Annotations: When and Why Semantic Neighbors Help Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We provide a theoretical framework that derives a closed-form expression for the error detection AUROC, which can be decomposed into three factors: intrinsic separability, reference-induced mean shift, and noise reduction through averaging. |
Na Di; Ling Li; Zhe Tang; Hao Cheng; Jinlong Pang; Jiaheng Wei; Zhaowei Zhu; | code |
| 1141 | When Preference Labels Fall Short: Aligning Diffusion Models from Real Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we investigate whether real data can serve as an alternative source of supervision for preference alignment. |
Weiyan Chen; Weijian Deng; Yao Xiao; Weijie Tu; ZiYi Dong; Ibrahim Radwan; Liang Lin; Pengxu Wei; | code |
| 1142 | Fair Dataset Distillation Via Cross-Group Barycenter Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Crucially, these gaps do not disappear by merely correcting group imbalance, since they stem from fundamental mismatches in subgroup predictive patterns rather than from sample-size disparities alone. We therefore formally analyze the interaction between these two sources of bias and cast the solution as identifying a group-imbalance-agnostic barycenter of the predictive information that induces similar representations across all subgroups. |
Mohammad Hossein Moslemi; Nima Hosseini Dashtbayaz; Zhimin Mei; Boyu Wang; Bissan Ghaddar; | code |
| 1143 | Amortized Maximum Inner Product Search with Learned Support Functions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose amortized MIPS: a learning-based approach that trains neural networks to directly predict MIPS solutions, amortizing the computational cost of search across queries drawn from a known distribution. |
Theo X. Olausson; Joao Monteiro; Michal Klein; Marco Cuturi; | code |
| 1144 | Is Training Necessary for Anomaly Detection? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We then abandon the reconstruction paradigm entirely and propose Retrieval-based Anomaly Detection (RAD). |
Xingwu Zhang; Guanxuan Li; Paul Henderson; Gerardo Aragon-Camarasa; ZIJUN LONG; | code |
| 1145 | Seizure-Semiology-Suite($S^3$): A Clinically Multimodal Dataset, Benchmark, and Models for Seizure Semiology Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in general video understanding, their capacity to interpret involuntary, and spatio-temporally evolving pathologic motor behaviors such as seizure semiology remains largely untested. To address this gap, we introduce Seizure-Semiology-Suite (S³), a clinically grounded dataset and benchmark for fine-grained, structured seizure semiology understanding. |
Lina Zhang; Jiarui Cui; Tonmoy Monsoor; Peizheng Li; Xinyi Peng; Chong Han; Prateik Sinha; Siyuan Dai; Jessica Pasqua; Colin McCrimmon; Weiting Liu; Hailey Miranda; Bing Hu; Xiangting Wu; Tengyou Xu; Chunhan Li; Jiaye Tian; Jiarui Tang; Detao Ma; Lingye Kong; Junnan Lyu; Jungang Li; Yan Zan; Junhua Huang; Rajarshi Mazumder; Vwani Roychowdhury; | code |
| 1146 | Balancing Fidelity and Diversity in Diffusion Models Via Symmetric Attention Decomposition: Hopfield Perspective Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We characterize the pre-softmax attention matrix $\mathbf{QK^\top}$ in transformers as an associative memory matrix encoding pairwise associations between input features. |
Hyunmin Cho; Woo Kyoung Han; Kyong Jin; | code |
| 1147 | TAG: Tangential Amplifying Guidance for Hallucination-Resistant Sampling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose **T**angential **A**mplifying **G**uidance **(TAG)**, a theoretically grounded, training-free, computationally lightweight, and architecture-agnostic guidance method that operates solely on trajectory signals without modifying the underlying diffusion model. |
Hyunmin Cho; Donghoon Ahn; Susung Hong; Jee Kim; Seungryong Kim; Kyong Jin; | code |
| 1148 | Vegas: Self-Speculative Decoding with Verification-Guided Sparse Attention Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Vegas, a self-speculative decoding method with verification-guided sparse attention. |
Yikang Yue; Yuqi Xue; Jian Huang; | code |
| 1149 | Revisiting Coding-Based Approaches to Overcome The Curse of Dimensionality in Learning-Based Watermarking Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To leverage the strengths of both approaches, we propose OrthoMark, a framework that decouples robust feature extraction from message encoding. |
Yupeng Qiu; Han Fang; Ee-Chien Chang; | code |
| 1150 | Is One Layer Enough? Understanding Inference Dynamics in Tabular Foundation Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present the first large-scale mechanistic study of layerwise dynamics in 6 state-of-the-art tabular in-context learning models. |
Amir Balef; Mykhailo Koshil; Katharina Eggensperger; | code |
| 1151 | Diving Into Kronecker Adapters: Component Design Matters Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we identify component structure as a key factor governing the capacity of Kronecker adapters. |
Jiayu Bai; Danchen Yu; Zhenyu Liao; TianQi Hou; Feng Zhou; Robert Qiu; Zenan Ling; | code |
| 1152 | Quantifying Temperature Scaling in Discrete Sequence (Language) Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we address the challenge of reliable temperature scaling with a novel fine-tuning procedure and introduce a new metric to measure effective temperature scaling without requiring the partition function. |
Hannah Scheufele; Peter Blohm; Vikas Garg; | code |
| 1153 | Toward More Reliable Agent Evaluation: A Component-Based Benchmark Auditing Pipeline Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose the **COBA** (**CO**mponent-based **B**enchmark **A**uditing) pipeline, an automated pipeline for diagnosing and filtering validity issues in agent benchmarks. |
Hyewon Suh; Seojune Lee; Binfei Ji; Rishi Khare; Basit Khan; Hyunjun Kim; Tianyi Zhang; Venkat Krishna Srinivasan; Peter Belcak; Shizhe Diao; Pavlo Molchanov; Yingyan (Celine) Lin; Zhen Dong; | code |
| 1154 | Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While GUI agents have advanced rapidly, they often lack the robustness to recover from their own errors, hindering real-world deployment. To bridge this gap at both the evaluation and data levels, we introduce GUI-RobustEval and propose Robustness-driven Trajectory Synthesis. |
Tianpeng Bu; Xin Liu; Qihua Chen; Hao Jiang; Shurui Li; hongtao duan; Lu Jiang; lulu hu; Bin Yang; Minying Zhang; | code |
| 1155 | World-Shaper: A Unified Framework for 360° Panoramic Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address geometric distortion, we introduce a geometry-aware learning strategy that explicitly enforces position-aware shape supervision and implicitly internalizes panoramic priors through progressive training. |
Dong Liang; yuhao liu; Jinyuan Jia; Youjun Zhao; Rynson Lau; | code |
| 1156 | OMAC: A Holistic Optimization Framework for LLM-Based Multi-Agent Collaboration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce OMAC, a general framework designed for holistic optimization of LLM-based MAS. |
Shijun Li; Hilaf Hasson; Joydeep Ghosh; | code |
| 1157 | IBMA: Information Bottleneck-Based Multimodal Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Information Bottleneck–based Multimodal Alignment (IBMA), a novel multimodal learning framework that enforces the IB principle for both the fused multimodal representation and modality-specific representations. |
Yancheng Wang; Zeyu Dong; Dongfang Sun; Alvin Silva; Teresa Wu; Yingzhen Yang; | code |
| 1158 | VideoSEAL: Separating Planning from Answer Authority for Agentic Long Video Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We show that existing LVU agents can exhibit evidence misalignment: they produce correct answers that are not supported by the retrieved or inspected evidence. To characterize this failure, we introduce two diagnostics (temporal groundedness and semantic groundedness) and use them to reveal two pressures that amplify misalignment: prompt pressure from shared-context saturation at inference time and reward pressure from outcome-only optimization during training. |
Chenhao Qiu; Yechao Zhang; Xin Luo; Shien Song; Xusheng Liu; | code |
| 1159 | Geometric Flow Grounding: A Unified Manifold Decoupling Framework for Dynamics Discovery and Verification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing approaches ranging from Neural ODEs to diffusion models are often plagued by the entanglement of static state representations and instantaneous motion, leading to accumulated errors and off-manifold hallucinations where predicted trajectories violate intrinsic geometric constraints. To address this, we propose Geometric Flow Grounding, a unified framework that enforces dynamic evolution strictly along the tangent bundle of the learned data manifold via a differentiable Neural Tangent Projection Layer. |
Chang Yu; Yuxuan Luo; Yixuan Du; Yuqing Zhou; Siyuan Li; Jingbo Zhou; jiawei jiang; Zhen Lei; Stan Z Li; | code |
| 1160 | Cello: A Universal Cell-wise Feature Aggregation Framework for Reliable Pathology Images Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Cello, a universal cell-wise feature aggregation framework for reliable pathology image analysis. |
Hengrui Lou; Weihan Li; Jiazhen Yang; Lingxiang Jia; Shengxuming Zhang; Linyun Zhou; Xiuming Zhang; Zhenyang Wang; Mingli Song; Zunlei Feng; | code |
| 1161 | RefChess: Monte-Carlo Move Selection for Zero-Shot Referring Image Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Nevertheless, selecting the correct segmentation proposal remains challenging, as existing methods typically rely on independent proposal scoring and lack contextual reasoning among visually similar candidates. To address this limitation, we propose RefChess, a training-free framework that reformulates proposal selection as a decision-making problem under contextual perturbations rather than a single-step ranking task. |
Shiyan Tong; Jinxia Zhang; Zhiyuan Wang; Hao Tian; YingYing Wang; Kanjian Zhang; Haikun Wei; | code |
| 1162 | Beyond Logits: Metastable Latent Dynamics for Sample-Efficient Best-of-N Selection in LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Drawing from cognitive neuroscience, we hypothesize that effective reasoning exhibits \textit{metastability}—a balance between stability and flexibility manifested as structured “dwell-and-jump” dynamics. We introduce Latent Velocity Entropy (LVE), a training-free metric that quantifies these dynamics via the entropy of internal representation updates. |
Xinrong Li; Zidong Zhou; Keyu Shen; Wenhao Zhou; Shangqi Guo; | code |
| 1163 | Q-CLIP: Unleashing The Power of Vision-Language Models for Video Quality Assessment Through Unified Cross-Modal Adaptation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose Q-CLIP, the first fully CVLMs-based framework for VQA. |
Yachun Mi; Yu Li; Yanting Li; Chen Hui; Tong Zhang; Zhixuan Li; Chenyue Song; Wei Lim; Shaohui Liu; | code |
| 1164 | PluRel: Synthetic Data Unlocks Scaling Laws for Relational Foundation Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Here we introduce PluRel, a framework to synthesize multi-tabular relational databases from scratch. |
Vignesh Kothapalli; Rishabh Ranjan; Valter Hudovernik; Vijay Prakash Dwivedi; Johannes Hoffart; Carlos Guestrin; Jure Leskovec; | code |
| 1165 | PVDepth: Panoramic Video Depth Estimation Via Geometry-Aware Spatiotemporal Adaptation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we first present **PanoCARLA**, a large-scale synthetic RGB-D panoramic video dataset, featuring natural motion trajectories and drone-like roaming perspectives. Building on this foundation, we propose **PVDepth**, an end-to-end framework adapted from perspective video depth models. |
Chuanxin Song; Peixi Peng; | code |
| 1166 | Generalizing Multi-Scale Time-Series Modeling with A Single Operator Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As no principled foundation has been established in the literature, we unify existing scaling methods into a scaling operator family, revealing a fundamental limitation of existing approaches: reliance on fixed and discrete scaling. To address this limitation, we propose SiGMA (Single Generalized Multi-scale Architecture), which enables position-wise scaling via the learnable discrete Gaussian (LDG) kernel grounded in scale-space theory. |
Cheonwoo Lee; Dooho Lee; Doyun Choi; Jaemin Yoo; | code |
| 1167 | Plug-and-Play Label Map Diffusion for Universal Goal-Oriented Navigation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose Plug-and-Play Label Map Diffusion (PLMD), which defines a novel map completion diffusion model based on Denoising Diffusion Probabilistic Models (DDPM). |
Zhixuan Shen; Yijie Zeng; Shengxiang Luo; Tianrui Li; Haonan Luo; | code |
| 1168 | Dream-MPC: Gradient-Based Model Predictive Control with Latent Imagination Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Dream-MPC, a novel approach that generates few candidate trajectories from a rolled-out policy and optimizes each trajectory by gradient ascent using a learned world model, uncertainty regularization and amortization of optimization iterations over time by reusing previously optimized actions. |
Jonathan Spieler; Sven Behnke; | code |
| 1169 | Fast Estimation for Forest Matrix of Signed Graphs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we study the problem of efficiently estimating the forest matrix of signed graphs with \(n\) nodes and introduce the signed forest matrix theorem, which establishes the relationship between generalized spanning converging forests and the forest matrix. |
Haoxin Sun; Zhongzhi Zhang; | code |
| 1170 | Leveraging Machine Unlearning for Cost-Efficient Preference Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, current research has primarily focused on empirical validation, lacking systematic quantitative analysis. To bridge this gap, we propose a framework linking PA with LLM unlearning. |
XiaoHua Feng; Yuyuan Li; HuWei Ji; Li Zhang; Jiaming Zhang; Tianyu Du; Chaochao Chen; | code |
| 1171 | IPMark: A Sentence-Level Watermark for LLMs with Hierarchical Personalization and Efficient Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing methods still suffer from significant limitations, including difficulties in achieving personalized attribution, substantial degradation of generation quality, and weak robustness against attacks. To address these challenges, we propose IPMark, the first IP-inspired hierarchical personalized watermarking framework. |
Wenbo An; Lianwei Wu; Zehao Wang; | code |
| 1172 | Stem: Rethinking Causal Information Flow in Sparse Attention Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we rethink the causal attention mechanism from the perspective of information flow. |
Lin Niu; Xin Luo; LinchuanXie; Yifu Sun; Guanghua Yu; Jianchen Zhu; S Kevin Zhou; | code |
| 1173 | Improving Classifier-Free Guidance of Flow Matching Via Manifold Projection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we provide a principled interpretation of CFG through the lens of optimization. |
Jian-Feng Cai; Haixia Liu; Zhengyi Su; Wang; | code |
| 1174 | Toward Effective Multimodal Graph Foundation Model: A Divide-and-Conquer Based Approach Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While recent MGFMs integrate diverse modality information, our empirical investigation reveals two fundamental limitations of existing MGFMs: (1)they fail to explicitly model modality interaction, essential for capturing intricate cross-modal semantics beyond simple aggregation, and (2)they exhibit sub-optimal modality alignment, which is critical for bridging the significant semantic disparity between distinct modal spaces. To address these challenges, we propose PLANET (graPh topoLogy-aware modAlity iNteraction and alignmEnT), a novel framework employing a Divide-and-Conquer strategy to decouple modality interaction and alignment across distinct granularities. |
Sicheng Liu; Xunkai Li; Daohan Su; Ru Zhang; Hongchao Qin; Rong-Hua Li; Guoren Wang; | code |
| 1175 | OVLR: Efficient, Scalable, and Robust Training Via Output-Level Variance-Reduced Likelihood Ratio Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While likelihood ratio (LR) methods offer a theoretical alternative, their high variance in high-dimensional spaces often undermines training stability and scalability. We propose OVLR (Output-Level Variance-Reduced Likelihood Ratio), a simple yet powerful framework that circumvents this fundamental trade-off by providing a unified solution for efficient, scalable, and robust gradient estimation. |
Minhao Zou; Tao Ren; Jinyang Jiang; Rui Tao; Zehao Li; Jiale Fu; Hui Shao; Xianhua Liu; Yijie Peng; | code |
| 1176 | Uncertainty-Guided Exploration and Stable Planning for Sparse-Reward Manipulation from Limited Demonstrations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose QUEST, a model-based RL framework that adaptively switches between exploration and exploitation guided by uncertainty to achieve stable and efficient learning. |
Haowen Sun; Liqi Huang; li mingyang; Sihua Ren; Xinzhe Chen; Chengzhong Ma; Zeyang Liu; Xingyu Chen; Xuguang Lan; | code |
| 1177 | SEMA: A Scalable and Efficient Mamba Like Attention Via Token Localization and Averaging Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We provide a mathematical definition of generalized attention and formulate both vanilla softmax attention and linear attention within the general framework. |
Nhat Thanh Tran; Fanghui Xue; shuai zhang; Jiancheng Lyu; Yunling Zheng; YINGYONG QI; Jack Xin; | code |
| 1178 | Breaking The Capacity Bottleneck in Model-Heterogeneous Federated Learning Via Gradual Model Restoration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In model-heterogeneous FL with fixed small sub-models, BCCs with sub-models may improve quickly in early rounds but become under-parameterized later, resulting in slow convergence and poor generalization. To address this challenge, we propose FedGMR, a federated learning framework centered around Gradual Model Restoration (GMR), where GMR progressively increases each client’s sub-model density during training, allowing BCCs to remain effective contributors throughout optimization. |
CHENGJIE MA; Seungeun Oh; Jihong Park; Seong-Lyun Kim; | code |
| 1179 | A Penalty Approach For Differentiation Through Black-box Quadratic Programming Solvers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Most existing approaches differentiate through the Karush–Kuhn–Tucker (KKT) system, but their computational cost and numerical robustness can degrade at scale. To address these limitations, we propose dXPP, a penalty-based differentiation framework that decouples QP solving from differentiation. |
Yuxuan Linghu; Zhiyuan Liu; Qi Deng; | code |
| 1180 | Approximate Nearest Neighbor Search for Modern AI: A Projection-Augmented Graph Approach Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we outline six critical demands of modern AI applications: high query efficiency, fast indexing, low memory footprint, scalability to high dimensionality, robustness across varying retrieval sizes, and support for online insertions. |
Kejing Lu; Zhenpeng Pan; Jianbin Qin; Yoshiharu Ishikawa; Chuan Xiao; | code |
| 1181 | Localizing Memorized Regions in Diffusion Models Via Coordinate-Wise Curvature Differences Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To isolate overfitting-driven memorization, we propose curvature-difference methods that subtract the curvature of an underfitted baseline, either the unconditional model or a less-trained version of itself. |
Gwangho Kim; Sungyoon Lee; | code |
| 1182 | ProAct: A Benchmark and Multimodal Framework for Structure-Aware Proactive Response Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These graphs encode step dependencies and parallel execution possibilities, providing the structural grounding necessary for complex decision-making. Building on this benchmark, we propose **ProAct-Helper**, a reference baseline powered by a Multimodal Large Language Model (MLLM) that grounds decision-making in state detection, and leveraging task graphs to enable entropy-driven heuristic search for action selection, allowing agents to execute parallel threads independently rather than mirroring the human’s next step. |
Xiaomeng ZHU; Fengming ZHU; Weijie Zhou; Ye Tian; Zhenlin Hu; Yufei Huang; Yuchun Guo; Xinyu Wu; Zhengyou Zhang; Fangzhen Lin; Xuantang Xiong; | code |
| 1183 | Unbiased and Second-Order-Free Training for High-Dimensional PDEs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we provide a principled analysis of EM-induced loss bias and propose an unbiased, second-order-free training framework that preserves the computational advantages of BSDE methods. |
Jaemin Seo; Su Rin Lee; JaeYong Lee; | code |
| 1184 | Discovering Ordinary Differential Equations with LLM-Based Qualitative and Quantitative Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing symbolic regression approaches rely primarily on quantitative metrics; however, real-world differential equation modeling also requires incorporating domain knowledge to ensure physical plausibility. To address this gap, we propose DoLQ, a method for discovering ordinary differential equations with LLM-based qualitative and quantitative evaluation. |
Sum Song; Bong Gyun Shin; JaeYong Lee; | code |
| 1185 | DiasR: Dual-Modal Identity-Anchored Sparse Routing for Efficient Multi-Subject Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Personalized multi-subject video generation is a promising direction within the field of controllable video generation; however, existing methods face challenges in maintaining cross-frame identity consistency and incur high computational overhead. To address these issues, we propose DiasR, an efficient framework that integrates Dual-Modal Identity-Anchored Alignment and a novel Sparse Routing Strategy. |
Yang-yang Li; Wu Liu; Jie Li; Xinchen Liu; Yongdong Zhang; Guoqing Jin; | code |
| 1186 | GRAPE: Let GRPO Supervise Query Rewriting By Ranking for Retrieval Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, its performance often degrades under distributional shifts such as multilingual, long-form, or multimodal queries. To avoid the prohibitive costs associated with retriever retraining or corpus re-embedding, we propose GRAPE (Grouped Ranking-Aware Policy Optimization Enhancement), a plug-and-play approach that leverages LLM-based query rewriting to bridge these gaps. |
Zhaohua Zhang; Jianhuan Zhuo; Muxi Chen; Chenchen Zhao; Wenyu Jiang; Tianwen Jiang; Mingyang Chen; Yutang; Jihong Zhang; Qiuyong Xiao; Zhixun Su; | code |
| 1187 | H$^2$CL: Heterogeneity-Aware Hypergraph Contrastive Learning for Robust Representation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, traditional hypergraph learning method often assume that neighboring nodes are homogeneous, which can lead to the mixing of heterogeneous information in highly heterogeneous datasets, thereby affecting node feature representation. To address this issue, this paper proposes a heterogeneity-sensitive hypergraph contrastive learning method. |
Kaixuan Yao; Ting Guo; Ming Li; Feilong Cao; | code |
| 1188 | N2M: Bridging Navigation and Manipulation By Learning Pose Preference from Rollout Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce N2M that systematically reformulates the approach to base positioning problem, naturally overcoming limitations of previous methods. |
Kaixin Chai; Hyunjun Lee; Joseph Lim; | code |
| 1189 | Neural Dispersion on Graphs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study the problem of generating structurally diverse graphs on $N$ unlabeled vertices. |
Ryien Hosseini; Pouya Gholami; Filippo Simini; Venkatram Vishwanath; Rebecca Willett; Henry (Hank) Hoffmann; | code |
| 1190 | Offline Reinforcement Learning with Universal Horizon Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we introduce universal horizon models (UHM), a generalization of GHM that directly predicts future states under arbitrary horizons. |
Hojun Chung; Junseo Lee; Songhwai Oh; | code |
| 1191 | Rank-guided Diffusion for Noise Few-Shot Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We find that clean samples in semantic feature space lie in low-rank subspaces, while noisy samples cause rank anomalies disrupting this structure. To address this, we propose a differentiable low-rank approximation that estimates the intrinsic rank of the support set and detects anomalous noisy samples. |
Zelei Wu; Kun Zhou; xulun ye; Jie Hong; Jieyu Zhao; Yifan Mei; | code |
| 1192 | TriForces: Augmenting Atomistic GNNs for Transferable Representations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, MLIPs transfer inconsistently across domains, with representations that often loose accessible composition and structure information. To address this, we present TriForces, a model-agnostic three-stream framework that separates composition and structure information, combined with self-supervised learning to preserve transferable representations. |
Ali Ramlaoui; Alexandre Duval; Hannah Bull; Victor Schmidt; Hugues Talbot; Fragkiskos Malliaros; Joseph Musielewicz; | code |
| 1193 | Spatially-Adaptive Gradient Re-parameterization for 3D Large Kernel Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Motivated by the spatial bias inherent in effective receptive fields (ERFs), we theoretically demonstrate that structurally re-parameterized blocks induce spatially varying learning rates that are crucial for convergence. Leveraging this insight, we introduce Rep3D, a framework that employs a lightweight modulation network to generate receptive-biased scaling masks, adaptively re-weighting kernel updates within a plain encoder architecture. |
Ho Hin Lee; Quan Liu; Shunxing Bao; Yuankai Huo; Bennett Landman; | code |
| 1194 | Predicting Large Model Test Losses with A Noisy Quadratic System Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a predictive model that estimates the pre-training loss of large models from model size ($N$), batch size ($B$) and number of weight updates ($K$). |
Chuning Li; Chris Maddison; | code |
| 1195 | KromHC: Manifold-Constrained Hyper-Connections with Kronecker-Product Residual Matrices Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address both challenges, we propose **KromHC**, which uses the $\underline{\text{Kro}}$necker products of smaller doubly stochastic matrices to parametrize the residual matrix in $\underline{\text{mHC}}$. |
Wuyang Zhou; Yuxuan Gu; Giorgos Iacovides; Danilo Mandic; | code |
| 1196 | SURF: Separation Via Unsupervised Remixing Flow Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, access to such clean source samples is often limited. To bridge this gap, we present Separation via Unsupervised Remixing Flow (\textbf{SURF}), an unsupervised flow matching approach for source separation that learns directly from observed mixtures. |
Henry Li; Robin Scheibler; Efthymios Tzinis; Matt Shannon; Arnaud Doucet; john hershey; | code |
| 1197 | Sparks of Cooperative Reasoning: LLMs As Strategic Hanabi Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce and release two novel datasets: HanabiLogs (1,520 annotated trajectories) and HanabiRewards (560 games with dense move-level utilities). |
Mahesh Ramesh; Kaousheik Jayakumar; Aswinkumar Ramkumar; Pavan Thodima; Aniket Rege; Emmanouil-Vasileios Vlatakis-Gkaragkounis; | code |
| 1198 | Beyond Generative Priors: Minority Sampling with JEPA-Guided Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a world-centric perspective on minority sampling, which defines rarity with respect to real-world priors rather than generator-induced densities. |
Sol Park; Soobin Um; | code |
| 1199 | GHOST: Unmasking Phantom States in Mamba2 Via Grouped Hidden-state Output-aware Selection & Truncation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce GHOST (Grouped Hidden-state Output-aware Selection and Truncation), a structured pruning framework that approximates control-theoretic balanced truncation using only forward-pass statistics. |
Michael Menezes; Anastasios Kyrillidis; | code |
| 1200 | FlowNar: Scalable Streaming Narration for Long-Form Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While recent online adaptations improve real-time processing, they still face critical scalability challenges, with resource demands typically growing at least linearly with video duration. To overcome this bottleneck, we propose FlowNar, a novel framework for scalable streaming video narration. |
Zeyun Zhong; Manuel Martin; Chengzhi Wu; David Schneider; Frederik DIEDERICHS; Juergen Gall; Jürgen Beyerer; | code |
| 1201 | Uncovering Competency Gaps in Large Language Models and Their Benchmarks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To automatically uncover both types of gaps, we propose a simple new method using concept activations from sparse autoencoders, to identify fine-grained gaps on a per-concept basis. |
Maty Bohacek; Nino Scherrer; Nicholas Dufour; Thomas Leung; Christoph Bregler; Stephanie Chan; | code |
| 1202 | Darwinian Memory: A Training-Free Self-Regulating Memory System for GUI Agent Evolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While memory systems provide a viable solution, existing paradigms struggle to adapt to dynamic GUI environments, suffering from a granularity mismatch between high-level intent and low-level execution, and context pollution where the static accumulation of outdated experiences drives agents into hallucination. To address these bottlenecks, we propose the Darwinian Memory System (DMS), a self-evolving architecture that constructs memory as a dynamic ecosystem governed by the law of survival of the fittest. |
Hongze Mi; Yibo Feng; Wenjie Lu; Song Cao; Jinyuan Li; Yanming Li; Xuelin Zhang; Haotian Luo; Songyang Peng; He Cui; Tengfei Tian; Jun Fang; Hua Chai; Naiqiang Tan; | code |
| 1203 | TokenDrop: Token-Level Importance-Aware Backward Propagation Skipping for Efficient LLM Fine-Tuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose **TokenDrop**, a token-level importance-aware backpropagation skipping method that reduces activation memory and accelerates LLM fine-tuning by skipping backward computations for less informative tokens. |
Beomseok Kim; Sol Namkung; Dongsuk Jeon; | code |
| 1204 | GeoFlow: Geo-Aware Modeling of Inter-Area Relationships in OD Flow Prediction and Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce GeoFlow, a novel framework that (i) augments area representations with geospatial attributes, including relative positions, -hop and geodesic distances, (ii) employs a specialized geometric-intrinsic fusion encoder design that combines graph attention for intrinsic area signals with coordinate-aware encoders for global structure, and (iii) adopts an axial-global attention decoder to capture OD-specific competitive dependencies. |
Zherui Huang; Guanjie Zheng; Hao Xue; Linghe Kong; | code |
| 1205 | GraphFLEx: Unsupervised Structure Learning $\underline{\text{F}}$ramework for $\underline{\text{L}}$arge $\underline{\text{Ex}}$panding $\underline{\text{Graph}}$s Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose GraphFLEx—a unified and scalable framework for Graph Structure Learning in Large and Expanding Graphs. |
Mohit Kataria; Nikita Malik; Sandeep Kumar; Jayadeva Jayadeva; | code |
| 1206 | DAISI: Data Assimilation with Inverse Sampling Using Stochastic Interpolants Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Classical high-dimensional DA methods, such as the ensemble Kalman filter, rely on Gaussian approximations that are violated for complex dynamics or observation operators. To address this limitation, we introduce DAISI, a scalable filtering algorithm built on flow-based generative models that enables flexible probabilistic inference using data-driven priors. |
Martin Andrae; Erik Larsson; So Takao; Tomas Landelius; Fredrik Lindsten; | code |
| 1207 | WorldComp2D: Spatio-semantic Representations of Object Identity and Location from Local Views Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose WorldComp2D, a novel lightweight representation learning framework that explicitly structures latent space geometry according to object identity and spatial proximity using multiscale \textit{local} receptive fields centered on a given fixation point. |
Seongmin Jin; Doo Seok Jeong; | code |
| 1208 | MobileFusion: Mobile-Friendly Infrared and Visible Image Fusion Via Structural Re-parameterization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present MobileFusion, an extremely lightweight and effective convolutional framework that achieves high-quality fusion under strict resource constraints. |
Yufa Duan; Jialing Huang; Yingying Wang; Weimin Cai; Xinghao Ding; Xiaotong Tu; | code |
| 1209 | PRISM: Training-Free Video Anomaly Detection Via Intrinsic Statistical Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While emerging training-free video anomaly detection (VAD) methods offer advantages such as interpretability and ease of deployment, they often suffer from computational inefficiency due to complex memory retrieval mechanisms or high-latency visual language models (VLMs). To address this, we propose PRISM (Parameter-free Recognition Based on Intrinsic Statistical Modeling), a novel framework for efficient open-set anomaly detection with minimal computational cost. |
YUANTONG CHEN; Zhengyan Ding; YanFeng Shang; | code |
| 1210 | When Data Is Scarce: Scaling Sparse Language Models with Repeated Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we study sparse training in data-constrained regimes where limited unique tokens require multi-epoch training. |
Boqian Wu; qiao xiao; Patrik Okanovic; Tomasz Sternal; Maurice Keulen; Mykola Pechenizkiy; Elena Mocanu; Torsten Hoefler; Decebal Constantin Mocanu; | code |
| 1211 | TelecomTS: A Multi-Modal Observability Dataset for Time Series and Language Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing datasets are often anonymized and normalized, removing scale information and limiting their use for tasks such as anomaly detection, root-cause analysis, and multi-modal reasoning. To address this gap, we introduce TelecomTS, a large-scale observability dataset derived from a 5G telecommunications network. |
Austin Feng; Andreas Varvarigos; Ioannis Panitsas; Daniela Fernandez; Yuwei Guo; Jinbiao Wei; Chen; Ali Maatouk; Leandros Tassiulas; ZHITAO YING; | code |
| 1212 | Generalized Discrete Diffusion with Self-Correction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a **S**elf-**C**orrecting **D**iscrete **D**iffusion (SCDD) model to reformulate pretrained self-correction with explicit state transitions and learn directly in discrete time. |
Linxuan Wang; Ziyi Wang; Yikun Bai; Wei Deng; Guang Lin; Qifan Song; | code |
| 1213 | Multivariate Distributional Reinforcement Learning Using Sliced Divergences Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Sliced Distributional Reinforcement Learning (SDRL), which lifts tractable one-dimensional divergences to multivariate return distributions via projections. |
Baptiste Debes; Tinne Tuytelaars; | code |
| 1214 | Preference-based Antibody Expression Ranking: Scaling with Large-scale Weak Supervision Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Antibody expression ranking is a critical task in antibody design, yet its modeling is severely hindered by the scarcity of labeled experimental data. To address this, we propose a unified preference-based learning framework that integrates scarce quantitative expression data with large-scale weak positive supervision from immunization data. |
Josh Sun; Morteza Babaie; Wenyang hou; Mark Crowley; David Young; | code |
| 1215 | MeshTok: Efficient Multi-Scale Tokenization for Scalable PDE Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This inflexible tokenization scheme is inherently limited in its ability to efficiently represent and process solutions to complex PDEs. To address this, we propose MeshTok, an adaptive mesh refinement (AMR)-inspired tokenization and sequence modeling framework. |
Zhao Yanshun; Xiaoyu Peng; Congcong Zhu; Jiamin Jiang; Jingrun Chen; | code |
| 1216 | Strat-Reasoner: Reinforcing Strategic Reasoning of LLMs in Multi-Agent Games Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose Strat-Reasoner, a novel RL-based framework that improves LLMs’ strategic reasoning ability in multi-agent games. |
Yidong He; Yutao Lai; Pengxu Yang; Jiarui Gan; Jiexin Wang; Yi Cai; Mengchen Zhao; | code |
| 1217 | Efficient Continuous-Depth Modeling with GRU Equivalents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Continuous Depth Acceleration (CoDA), a framework that leverages Mori–Zwanzig/Koopman operator theory to replace continuous-depth layers requiring multiple nonlinear ODEs with a compact GRU module, a single low-dimensional linear ODE, and a dense layer. |
Ayan Banerjee; BIN XU; Sandeep Gupta; | code |
| 1218 | Mitigating Label Shift in Tabular In-Context Learning Via Test-Time Posterior Adjustment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, we find that TabPFN is vulnerable to label shift, often overfitting to the majority class in the training dataset. To address this limitation, we propose DistPFN, the first test-time posterior adjustment method designed for tabular foundation models. |
Seunghan Lee; | code |
| 1219 | PLATE: Plasticity-Tunable Efficient Adapters for Geometry-Aware Continual Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We develop a continual learning method for pretrained models that \emph{requires no access to old-task data}, addressing a practical barrier in foundation model adaptation where pretraining distributions are often unavailable. |
Romain Cosentino; | code |
| 1220 | Semi-Supervised Gaze Estimation Via Disentangled Subspace Contrastive Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we devise a simple yet efficient semi-supervised contrastive learning framework to exploit unlabeled data for generalized gaze estimation, thereby reducing reliance on manual annotations. |
Qida Tan; Hongyu Yang; Wenchao Du; | code |
| 1221 | PLSemanticsBench: A Formal Semantics Reasoning Benchmark for Code Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Program execution provides a canonical instance: formal semantics define behavior through sym- bolic transition rules that can be systematically altered under distribution shift. We investigate whether LLMs can condition their reasoning on formal semantics through program execution and introduce PLSEMANTICSBENCH, pairing featherweight C programs with two se- mantic systems—small-step operational seman- tics and K semantics—and probing four capabil- ities: composing rules for final states, selecting rules when state is unmutated, sustaining such conditioning over long traces, and following sup- plied rules under novel semantics. |
Aditya Thimmaiah; Jiyang Zhang; Jayanth Srinivasa; Junyi Jessy Li; Milos Gligoric; | code |
| 1222 | OOVDet: Low-Density Prior Learning for Zero-Shot Out-of-Vocabulary Object Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, previous methods are prone to overfitting the IV classes, leading to the OOV or undefined classes being misclassified as IV ones with a high confidence score. To address this issue, this paper proposes a zero-shot OOV detector (OOVDet), a novel framework that effectively detects predefined classes while reliably rejecting undefined ones in zero-shot scenes. |
Binyi Su; chenghao huang; Chenhaiyong; | code |
| 1223 | Lagrangian Perturbation Diffusion Steering: Latent Reinforcement Learning for Generative Policies Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While reinforcement learning can improve task performance, directly fine-tuning large action decoders is often unstable and sample inefficient. We propose **Lagrangian Perturbation Diffusion Steering (LP-DS)**, a lightweight adaptation method that improves a frozen generative policy while preserving its multimodal structure. |
Hikmet Simsir; Ozgur Oguz; | code |
| 1224 | The Hippocampal Place Field Gradient: A Bio-inspired Framework Building Multiscale Representation for Better Sample Efficiency Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a unified theoretical framework establishing how multiscale hippocampal place fields arise from the frequency-dependent decay of grid cell projections. |
ZHOU Shujun; Junrong Qi; Guozhang Chen; | code |
| 1225 | CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Through an in-depth analysis of CLIP, we observe that visual tokens located within referent regions often exhibit low similarity to the textual representation. Motivated by this insight, we introduce LiteLVLM, a training-free, text-guided token pruning strategy for efficient pixel grounding inference. |
Sangin Lee; Yukyung Choi; | code |
| 1226 | Consistency Deep Equilibrium Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce the Consistency Deep Equilibrium Model (C-DEQ), a novel framework that leverages consistency distillation to accelerate DEQ inference. |
Junchao Lin; Zenan Ling; Jingwen Xu; Robert Qiu; | code |
| 1227 | LLM Watermark Evasion Via Bias Inversion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We bridge this gap by theoretically analyzing rewriting-based evasion, demonstrating that reducing the average conditional probability of sampling green tokens by a small margin causes the detection probability to decay exponentially. Guided by this insight, we propose the Bias-Inversion Rewriting Attack (BIRA), a practical query-free method that applies a negative logit bias to a proxy suppression set identified via token surprisal. |
Jeongyeon Hwang; Sangdon Park; Jungseul Ok; | code |
| 1228 | Interpretable Functional Koopman Learning with Non-Markovian Closure for Spatiotemporal Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce MERLIN, a Koopman-based framework that lifts dynamics to the evolution of learned *observation functionals* with near-linear progression, enabling full-field reconstruction at arbitrary resolutions. |
Wanfeng Lu; He Ma; Wei Lin; Qunxi Zhu; | code |
| 1229 | Swift-SVD: Theoretical Optimality Meets Practical Efficiency in Low-Rank LLM Compression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Swift-SVD, an activation-aware, closed-form compression framework that simultaneously guarantees theoretical optimum, practical efficiency and numerical stability. |
Ruoling Qi; Yirui Liu; Xuaner Wu; Xiangyu Wang; Ming Li; Chen Chen; Jian Chen; Yin Chen; Qizhen Weng; | code |
| 1230 | TPV: Parameter Perturbations Through The Lens of Test Prediction Variance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We identify test prediction variance (TPV)—the first-order sensitivity of model outputs to parameter perturbations around a trained solution—as a unifying quantity that links several classical observations about generalization in deep networks. |
Devansh Arpit; | code |
| 1231 | SlerpFlow: Spherical Trajectory Correction for Rectified Flow Inversion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introduces \textbf{SlerpFlow}, a straightforward yet highly effective zero-shot approach that unlocks the full potential of FLUX for high-fidelity inversion and editing. |
Wenbin Duan; Yan Shu; Zhuoyuan Fu; Fangmin Zhao; Yan Li; Yaru Zhao; Binyang Li; | code |
| 1232 | Hyperbolic Associative Memory Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Modern Hopfield Networks (MHNs) have achieved widespread success across various domains but are confined to Euclidean/Hilbert spaces, failing to preserve the hierarchical structure of data due to geometric constraints—arbitrary tree structures cannot be embedded with low distortion, while hyperbolic spaces can naturally accommodate hierarchical structures through exponential volume growth. To address this issue, we propose Hyperbolic Associative Memory Networks (HAMNs), the first framework to embed modern associative memory into hyperbolic space: we map query and memory vectors from Euclidean space to a constant negative curvature manifold via exponential maps, define a regularized energy function based on the Minkowski inner product, and adopt curvature-aware Riemannian optimization combined with exponential map updates to achieve stable on-manifold retrieval. |
Boliang Hao; Bailing Zhang; Fangyu Wu; | code |
| 1233 | Graph Alignment Via Dual-Pass Spectral Encoding and Latent Space Communication Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a novel framework that employs a dual-pass encoder to inject high-frequency discriminability into node features, paired with a geometry-aware functional map module that operates on the correspondence itself. |
Maysam Behmanesh; Erkan Turan; Maks Ovsjanikov; | code |
| 1234 | Large-Scale Molecular Dynamics Simulations: Direct Interatomic Modeling with Dilated Message Passing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a new message passing framework that can effectively and efficiently model interatomic interactions for simulating large-scale molecular dynamics at full atomic resolution. |
Haokai Hong; Wanyu LIN; KC Tan; | code |
| 1235 | Geometry-Correct Diffusion Posterior Sampling with Denoiser-Pullback Curvature Guidance and Manifold-Aligned Damping Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The correction pulls likelihood gradients back through the denoiser, uses a one-sided curvature model that avoids forward denoiser Jacobians, and applies diffusion-calibrated rank-one damping aligned with the denoiser residual. |
Seunghyeok Shin; Minwoo Kim; Dabin Kim; Hongki Lim; | code |
| 1236 | DISCO: Mitigating Bias in Deep Learning with Conditional Distance Correlation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce the Standard Anti-Causal Model (SAM), a unifying causal framework that characterizes bias mechanisms and yields a conditional independence criterion for causal stability. |
Emre Kavak; Tom Nuno Wolf; Christian Wachinger; | code |
| 1237 | PADS-TAL: Padding-Annealed Diffusion Sampling in Text-Aware Latent Space for Robust and Diverse Text-to-Music Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Text-to-Music diffusion models are increasingly used in real-world applications, yet deployment remains challenging: generations can collapse to limited patterns even with diverse initial noise and prompts, and inference-time diversity control often harms text alignment and fidelity by distorting key prompt cues established in early denoising. To address this, we propose Padding-Annealed Diffusion Sampling, which perturbs only a padding-indexed subspace while keeping non-padding conditioning fixed, enabling controlled exploration with reduced semantic drift. |
Taekoan Yoo; Wonkyung Jung; Kyunghun Kim; Kyeongbo Kong; | code |
| 1238 | CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream Tasks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Parameter-Efficient Fine-Tuning (PEFT) methods like Low-Rank Adaptation (LoRA) enable lightweight adaptation, yet they operate in isolation within each modality, limiting their ability in capturing cross-modal interactions. In this paper, we take a step in bridging this gap with Cross-Modal Low-Rank Adaptation (CoLA), a novel PEFT framework that extends LoRA by introducing a dedicated inter-modal adaptation pathway alongside the standard intra-modal one. |
Wish Suharitdamrong; Tony Alex; Muhammad Awais; Sara Atito; | code |
| 1239 | ProjQ: Project-and-Quantize for Adapter-Aware LLM Compression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose \textbf{ProjQ}, a novel framework for constraining quantization noise to the low-rank manifold via orthogonal subspace projection. |
Wenya Yu; chao zhang; Li Wang; Samson Lasaulce; Merouane DEBBAH; | code |
| 1240 | Dissecting Causal Mechanism Shifts Via FANS: Function And Noise Separation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introduces a more general and unified framework, the function and noise separation framework (FANS), that detects and dissects shifts in non-additive, non-linear Structural Causal Models (SCMs) beyond existing additive noise models. |
Gyeongdeok Seo; Jaeyoon Shim; Mingyu Kim; Hoyoon Byun; Yonghan Jung; Kyungwoo Song; | code |
| 1241 | Predicting The Order of Upcoming Tokens Improves Language Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Instead, we propose token order prediction (TOP), which trains models to order upcoming tokens by their proximity using a learning-to-rank loss. |
Zayd Zuhri; Erland Hilman Fuadi; Alham Aji; | code |
| 1242 | Proactive Defense Benchmark Against Deepfake Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite the proliferation of proactive defenses against deepfakes, the lack of a unified evaluation protocol precludes fair comparison and masks critical vulnerabilities. To bridge this gap, we present the first comprehensive benchmark that systematically assesses disruption, robustness, and transferability encompassing pixel, perceptual, and identity metrics. |
Joonhyuk BAEK; Wonjune Seo; Jae-yun Kim; Saerom Park; Hoki Kim; | code |
| 1243 | IDLM: Inverse-distilled Diffusion Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Nonetheless, this extension introduces both theoretical and practical challenges. |
David Li; Nikita Gushchin; Dmitry Abulkhanov; Eric Moulines; Ivan Oseledets; Maxim Panov; Aleksandr Korotin; | code |
| 1244 | On Efficient Scaling of GNNs Via IO-Aware Layers Implementations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We take an I/O- and arithmetic-intensity–centric view and show that widely used layers fall into three kernel families: SpMM-based convolutions, reduction-based aggregations, and attention-based layers (GATv2/Graph Transformer). |
Daria Fomina; Daniil Krasylnikov; Alexey Boykov; Andrey Dolgovyazov; Vyacheslav Zhdanovskiy; Fedor Velikonivtsev; | code |
| 1245 | Characterizing Vision-Language-Action Models Across XPUs: Constraints and Acceleration for On-Robot Deployment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a systematic framework for low-cost VLA deployment via model–hardware co-characterization. |
Kaijun Zhou; DaPeng; Qiwei Chen; Zhiyang Li; Xijun Li; Jinyu Gu; | code |
| 1246 | ToaSt: Token Channel Selection and Structured Pruning for Efficient ViT Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose ToaSt, a decoupled framework applying specialized strategies to distinct ViT components. |
Hyunchan Moon; Cheonjun Park; Steven Waslander; | code |
| 1247 | Offline Reinforcement Learning of High-Quality Behaviors Under Robust Style Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods, despite introducing numerous definitions of style, often fail to reconcile these objectives effectively. To address these challenges, we propose a unified definition of behavior style and instantiate it into a practical framework. |
Mathieu Petitbois; Rémy Portelas; sylvain lamprier; | code |
| 1248 | Activation-Free Backbones for Image Recognition: Polynomial Alternatives for Spatial and Channel Mixing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Modern vision backbones treat pointwise activations (e.g., ReLU, GELU) and exponential softmax as essential sources of nonlinearity, but we demonstrate they are not required. |
Jeffrey Wang; Jonathan Gregory; Grigorios Chrysos; | code |
| 1249 | Flow Equivariant World Models: Structured Memory for Dynamic Environments Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we show that enforcing equivariance between an agent’s representations and the world’s dynamics necessarily induces an efficient, structured memory. |
Hansen Lillemark; Benhao Huang; Fangneng Zhan; Yilun Du; T. Anderson Keller; | code |
| 1250 | Can Large Language Models Generalize Procedures Across Representations? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We find that training LLMs with popular post-training methods on graphs or code data alone does not reliably generalize to corresponding natural language tasks, while training solely on natural language can lead to inefficient performance gains. To address this gap, we propose a two-stage data curriculum that first trains on symbolic, then natural language data. |
Fangru Lin; Valentin Hofmann; Xingchen Wan; Weixing Wang; Zifeng Ding; Anthony Cohn; Janet Pierrehumbert; | code |
| 1251 | SORA: Free Second Order Attacks in Fast Adversarial Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Second, we propose **PertAlign** (Perturbation Alignment), a theoretically grounded, computationally negligible metric that predicts CO onset by measuring gradient alignment across attack stages. Leveraging these insights, we introduce **SORA**, an adaptive step-size adversarial training method that dynamically adjusts perturbations based on loss-surface geometry. |
Mazdak Teymourian; Ramtin Moslemi; Farzan Rahmani; Mohammad H Rohban; | code |
| 1252 | The Truth Stays in The Family: Enhancing Contextual Truthfulness Via Inherited Heads in Model Lineages Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our analysis of the Vicuna and Qwen families reveals a striking finding: MLLMs maintain a high correlation in truthfulness scores with their base LLMs, even after multi-modal fine-tuning and when evaluated on disparate data sources. Building on this insight, we propose a Soft Gating strategy that utilizes these inherited Truth Scores to amplify the influence of context-truthful heads while preserving the contributions of other heads. |
Miso Choi; Seonga Choi; Mincheol Kwon; Woosung Joung; Jinkyu Kim; JUNGBEOM LEE; | code |
| 1253 | POLIA: Policy Optimization with Visual-Object-Level Intrinsic Advantage for Multimodal Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose POLIA, a novel group-based RL method with visual-object-level intrinsic advantage for multimodal reasoning. |
Yiran Zeng; Da Chen; Hangyu Mao; Yuanxing Zhang; Pengfei Wan; Mengchen Zhao; | code |
| 1254 | Rationality Measurement and Theory for Reinforcement Learning Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper proposes a suite of rationality measures and associated theory for reinforcement learning agents, a property increasingly critical yet rarely explored. |
Kejiang Qian; Amos Storkey; Fengxiang He; | code |
| 1255 | Boosting Video Diffusion Models Via Masked Autoencoders As Tokenizers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose VideoMAETok, a simple family of ViT-based video tokenizers trained explicitly as corruption-inversion models for latent video diffusion. |
Zhan Tong; Tinne Tuytelaars; | code |
| 1256 | BEST: Benchmarking Efficiency in Space and Time for LLM-Generated Code Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To fill in the gap, this paper introduces *BEST*, the first benchmark for evaluating the efficiency of LLM-generated codes in *both time and space*. |
Aocheng Shen; Boyu Zhang; Jiaze Li; Ruixuan Ma; Qiankun Zhang; Wang; Bin Yuan; Shenghao Liu; Xianjun Deng; | code |
| 1257 | Robust Linear Dueling Bandits with Post-serving Context Under Unknown Delays and Adversarial Corruptions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Feedback is subject to unknown stochastic or adversarial delays and a cumulative corruption budget $\mathcal{C}$. To address these challenges, we propose \term, which integrates a learned approximator that predicts post-serving contexts from pre-serving information. |
Youngmin Oh; | code |
| 1258 | Continual Segmentation Under Joint Nonstationarity Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address instability and overfitting arising from few-shot supervision under distribution drift, we introduce gradient-adaptive stabilization, a parameter-wise regularization mechanism implemented via gradient-scaled stochastic perturbations that promotes a principled stability–plasticity tradeoff. |
Prashant Pandey; Himanshu Kumar; Devineni Chowdary; Brejesh Lall; | code |
| 1259 | Attend to Anything: Foundation Model for Unified Human Attention Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address the fundamental limitations, we present the Attend to Anything Model (AAM), a multi-modal foundation model that unifies attention modeling across various image, video, and audio-visual tasks and scenes. |
Wenzhuo Zhao; Ronghao Xian; Keren Fu; Qijun Zhao; | code |
| 1260 | Diffusion Language Model Parallel Decoding Via Product-of-Experts Bridge Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce PoE-Bridge, a novel decoding framework that drastically improves generation speed and accuracy by introducing an intermediate distribution to bridge the gap. |
Juntong Shi; Brian Trippe; Jure Leskovec; Stefano Ermon; Minkai Xu; | code |
| 1261 | SF-Mamba: Rethinking State Space Model for Vision Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose SF-Mamba, a novel visual Mamba with two key proposals: auxiliary patch swapping for encoding bidirectional information flow under an unidirectional scan and batch folding with periodic state reset for advanced GPU parallelism. |
Masakazu Yoshimura; Teruaki Hayashi; Yuki Hoshino; Wei-Yao Wang; Takeshi Ohashi; | code |
| 1262 | Noisy-Space Policy Gradient for Diffusion Policies in Offline Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Yet, their integration with reinforcement learning remains conceptually and algorithmically challenging. In this work, we address this gap by introducing a noisy-space action-value (Q-)function that assigns values to diffusion latents through the distribution of executed actions induced by the denoising process. |
Mahmoud Selim; Cristina Cipriani; Karl Johansson; | code |
| 1263 | Attention Sinks in Diffusion Transformers: A Causal Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a causal analysis of attention sinks in text-to-image diffusion models, dynamically identifying dominant attention recipients based on incoming attention mass. |
FANGZHENG WU; Brian Summa; | code |
| 1264 | Gauge-Equivariant Graph Networks Via Self-Interference Cancellation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a \textbf{G}auge-\textbf{E}quivariant Graph Network with \textbf{S}elf-Interference \textbf{C}ancellation (GESC), which replaces additive aggregation with a projection-based interference mechanism. |
Yoonhyuk Choi; Jiho Choi; Jiwoo Kang; | code |
| 1265 | Calibrating Generative Models to Distributional Constraints Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We frame calibration as a constrained optimization problem and seek the closest model in Kullback-Leibler divergence satisfying calibration constraints. To address the intractability of imposing these constraints exactly, we introduce two surrogate objectives for fine-tuning: (1) the relax loss, which replaces the constraint with a miscalibration penalty, and (2) the reward loss, which converts calibration into a reward fine-tuning problem. |
Henry Smith; Nathaniel Diamant; Brian Trippe; | code |
| 1266 | MacroGuide: Topological Guidance for Macrocycle Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce MacroGuide: Topological Guidance for Macrocycle Generation, a diffusion guidance mechanism that uses Persistent Homology to steer the sampling of pretrained molecular generative models toward the generation of macrocycles, in both unconditional and conditional (protein pocket) settings. |
Alicja Maksymiuk; Alexandre Duplessis; Ismail Ceylan; Alexander Tong; Fernanda Duarte; Michael Bronstein; | code |
| 1267 | SAEmnesia: Erasing Concepts in Diffusion Models with Supervised Sparse Autoencoders Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Concept unlearning in diffusion models is hampered by feature splitting, where concepts are distributed across many latent features, making their removal challenging and computationally expensive. We introduce SAEmnesia, a supervised sparse autoencoder framework that overcomes this by enforcing one-to-one concept-neuron mappings. |
Enrico Cassano; Riccardo Renzulli; Marco Nurisso; Mirko Zaffaroni; Alan Perotti; Marco Grangetto; | code |
| 1268 | TRACER: Persistent Regularization for Robust Multimodal Finetuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We develop a theoretical framework for multimodal contrastive finetuning by introducing a *contrastive target matrix* that reformulates the objective as a matrix least-squares problem, yielding closed-form solutions and a geometric decomposition of how different strategies manage pretrained knowledge. |
Hesam Asadollahzadeh; Feng Liu; Christopher Leckie; Sarah Erfani; | code |
| 1269 | Questioning The Coverage-Length Metric in Conformal Prediction: When Shorter Intervals Are Not Better Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Conformal prediction (CP) has become a cornerstone of distribution-free uncertainty quantification, conventionally evaluated by its coverage and interval length. This work critically examines the sufficiency of these standard metrics. |
Yizhou Min; Yizhou Lu; Lanqi Li; Zhen Zhang; Jiaye Teng; | code |
| 1270 | From Volume to Value: Preference-Aligned Memory Construction for On-Device RAG Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, under tight memory budgets, the core bottleneck is *what to store* so that retrieval remains aligned with the user. We propose EPIC (Efficient Preference-aligned Index Construction), which focuses on user preferences as a compact and stable form of personal context and integrates them throughout the RAG pipeline. |
Changmin Lee; Jaemin Kim; Taesik Gong; | code |
| 1271 | Hierarchical Multi Scale Graph Neural Networks: Scalable Heterophilous Learning with Oversmoothing and Oversquashing Mitigation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address the degree-biased aggregation and suboptimal polynomial filtering, we introduce a Hierarchical Multi‐view HAAR (HMH), a novel spectral graph‐learning framework that scales in near‑linear time . |
SAZZAD Hossen; Avimanyu Sahoo; | code |
| 1272 | KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce KAGE-Env, a JAX-native 2D platformer that factorizes the observation process into independently controllable visual axes while keeping the underlying control problem fixed. |
Egor Cherepanov; Daniil Zelezetsky; Aleksandr Panov; Alexey Kovalev; | code |
| 1273 | Training-free Composition of Pre-trained GFlowNets for Multi-Objective Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a training-free mixing policy that composes pre-trained GFlowNets at inference time, enabling rapid adaptation without finetuning or retraining. |
Seokwon Yoon; Youngbin Choi; Seunghyuk Cho; Seungbeom Lee; MoonJeong Park; Dongwoo Kim; | code |
| 1274 | Depth-Progressive Monotonic Learning Without Backpropagation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Backpropagation (BP) remains the dominant training paradigm for deep neural networks, yet its reliance on global gradient propagation fundamentally induces update locking problem, enforcing strong inter-layer dependencies in parameter updates. To address this limitation, we propose Depth-progressive Monotonic Learning (DMoL), a training scheme that assigns layer-wise local belief objectives and incrementally refines them across network depth, enabling unlocked parameter updates. |
Chenhao Ye; Rongguang Ye; Yuchao Zhang; Ming Tang; | code |
| 1275 | The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing expressivity results for transformers typically rely on hardmax attention, high precision, and other architectural modifications that disconnect them from the models used in practice. We bridge this gap by analyzing standard transformer decoders with softmax attention and rounding of activations and attention weights, while allowing depth and width to grow logarithmically with the context length. |
Moritz Brösamle; Stephan Eckstein; | code |
| 1276 | MapUQ: Map with Uncertainty Quantification for Robust BEV Vectorized Construction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, in complex traffic scenes, Bird’s-Eye-View (BEV) with vectorized mapping suffers from limitations such as target misclassification, spatial localization drift, and ambiguous semantic segmentation. Introducing uncertainty quantification can alleviate these problems, so we propose MapUQ, a robust BEV vectorized mapping method guided by uncertainty-aware optimization. |
Shaoyuan Mo; MaQi; l r; Bohan Li; WangKe; | code |
| 1277 | MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present MedMosaic, a medical audio question–answering dataset designed to benchmark language and audio reasoning models under realistic clinical constraints. |
Harshit Rajgarhia; Shuubham Ojha; Asif Shaik; Akhil Pothanapalli; Rachuri Lokesh; Abhishek Mukherji; Prasanna Desikan; | code |
| 1278 | MVR-cache: Optimizing Semantic Caching Via Multi-Vector Retrieval and Learned Prompt Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce MVR-cache, a novel semantic caching approach that significantly improves retrieval accuracy by integrating Multi-Vector Retrieval (MVR). |
Ali Noshad; Zishan Zheng; Yinjun Wu; | code |
| 1279 | VENOMREC: Cross-Modal Interactive Poisoning for Targeted Promotion in Multimodal LLM Recommender Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We find that while cross-modal consensus often mitigates conventional poisoning that manipulates interaction logs or perturbs a single modality, it also introduces a new attack surface where synchronised multimodal poisoning can reliably steer fused representations along stable semantic directions during fine-tuning. To characterise this threat, we formalise cross-modal interactive poisoning and propose VENOMREC, which performs Exposure Alignment to identify high-exposure regions in the joint embedding space and Cross-modal Interactive Perturbation to craft attention-guided coupled token–-patch edits. |
Guowei Guan; Yurong Hao; Jiaming Zhang; Tiantong Wu; Fuyao Zhang; Tianxiang Chen; Longtao Huang; Cyril Leung; Wei Lim; | code |
| 1280 | Poison with Style: A Practical Poisoning Attack on Code Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present Poison-with-Style (PwS), a practical and stealthy model poisoning attack targeting CLLMs. |
Khang Tran; Yazan Boshmaf; Issa Khalil; Hai Phan; Ting Yu; Md Rizwan Parvez; | code |
| 1281 | Panini: Continual Learning in Token Space Via Structured Memory Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a human-like non-parametric continual learning framework, where the base model remains fixed, and learning occurs by integrating each new experience into an external semantic memory state that accumulates and consolidates itself continually. |
Shreyas Rajesh; Pavan Holur; Mehmet Turali; Chenda Duan; Vwani Roychowdhury; | code |
| 1282 | Set-Coupled Guidance: Set-Level Coordination in Diffusion-Based Dataset Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Set-Coupled Guidance (SCG), a plug-and-play auxiliary controller that shifts from per-image to group (IPC-at-once) sampling by injecting set-symmetric feedback at each diffusion step. |
Ziang Gan; Qi Zhu; Libao Zhang; | code |
| 1283 | Is Your Diffusion Sampler Actually Correct? A Sampler-Centric Evaluation of Discrete Diffusion Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a sampler-centric oracle evaluation framework that replaces learned denoisers with an oracle Hidden Markov Model posterior derived from a ground-truth Markov chain, enabling isolation of sampler-induced error under controlled and method-consistent settings. |
Luhan Tang; Longxuan Yu; Shaorong Zhang; Greg Ver Steeg; | code |
| 1284 | RAIGen: Rare Attribute Identification in Text-to-Image Generative Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce RAIGen, the first framework, to our knowledge, for unsupervised rare-attribute discovery in diffusion models. |
Silpa Vadakkeeveetil Sreelatha; Dan Wang; Serge Belongie; Muhammad Awais; Anjan Dutta; | code |
| 1285 | Safety Anchor: Defending Harmful Fine-tuning Via Geometric Bottlenecks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our analysis traces this failure to the inherent redundancy of the high-dimensional parameter space: attackers exploit optimization trajectories that are orthogonal to defense constraints to restore harmful capabilities while deceptively adhering to safety restrictions. To address this, we propose Safety Bottleneck Regularization (SBR). |
Guoxin Lu; Letian Sha; Qing Wang; Peijie Sun; Hao Zhou; Hua Dai; Fu Xiao; | code |
| 1286 | TRACE: Toulmin-based Reasoning Assessment Through Constructive Elements for LLM CoT Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce TRACE (Toulmin-based Reasoning Assessment through Constructive Elements), a metric that analyzes Chain-of-Thought (CoT) reasoning processes. |
Kim Yundong; Heyoung Yang; | code |
| 1287 | Modeling Hierarchical Thinking in Large Reasoning Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose to approximate LRM’s emerging hierarchical reasoning dynamics as a trajectory within a Finite State Machine (FSM) transitioning among six abstract cognitive states. |
G Shahariar; Erfan Shayegani; Ali Nazari; Nael Abu-Ghazaleh; | code |
| 1288 | GKD-Recruiter: Jointly Modeling Social and Task Heterogeneity for Spatial Crowdsourcing Via Graph Knowledge Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Second, finite task demand creates a “saturation trap”, a non-submodular setting in which utility drops sharply to zero once demand is met. To bridge these gaps, we propose GKD-Recruiter, a Task-Aware framework designed to maximize Effective Task Satisfaction (ETS). |
Yucen Gao; Zhemeng Yu; Zhuoran Li; Jianxiong Guo; Xiaofeng Gao; | code |
| 1289 | Seeing Without Understanding: Disentangling Perception, Reasoning, and Simulation in VLM Gameplay Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a two-stage diagnostic framework that decomposes VLM performance into testable components: controlled perception tests isolating visual encoding, and a $2\times2$ diagnostic matrix with a six-level rule complexity ladder evaluated in both explicit verification and predictive simulation modes. |
Dingyang Jin; Jiawei He; Calvin Lo; Steven Hu; RYAN RAD; | code |
| 1290 | Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, GRPO is prone to advantage collapse, a failure mode where homogeneous rewards within a group (e.g., all correct or all incorrect answers) yield near-zero advantages and vanishing gradients. To address this, we introduce the Advantage Collapse Rate (ACR), the first diagnostic metric quantifying the proportion of training batches with ineffective gradients. |
Xixiang He; Qiyao Sun; Ao Cheng; Xingming Li; Xuanyu Ji; Hailun Lu; Runke Huang; Qingyong Hu; | code |
| 1291 | Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current Video-LLM approaches for Video Temporal Grounding (VTG) typically rely on direct timestamp generation from an unstructured visual-token stream, often resulting in brittle numerics and inconsistent boundaries. To address this, we propose Foresee-to-Ground (F2G), a framework that enforces a verifiable Identify-then-Measure routine. |
Zelin Zheng; Xinyan Liu; Ruixin Li; Antoni Chan; Guorong Li; Qingming Huang; Laiyun Qing; | code |
| 1292 | Accurate Evaluation of Quickest Changepoint Detectors Via Non-parametric Survival Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose non-parametric estimators for the average run length (ARL) and average detection delay (ADD) in quickest changepoint detection (QCD) under finite and irregular sequence lengths. |
Taiki Miyagawa; Akinori F. Ebihara; | code |
| 1293 | Convex Low-resource Accent-Robust Language Detection in Speech Recognition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Convex Language Detection (CLD), a novel framework that integrates theoretically grounded convex optimization techniques into the spoken dialogue systems pipeline. |
Miria Feng; William Tan; Mert Pilanci; | code |
| 1294 | Towards Fine-grained Robustness: Attention-guided Test-time Prompt Tuning for Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Prevalent test-time adaptation methods typically rely on the multi-view augmentation to implement various fine-tuning strategies, which struggle to identify semantic information and are prone to destroy the discriminative regions in fine-grained scenarios. To address these limitations, we propose Attention-guided Test-time Prompt Tuning (A-TPT), a semantics-preserving method designed for test-time adaptation. |
Jia-Wei Hai; Yijun Wang; Xiu-Shen Wei; | code |
| 1295 | CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present CAT-Q, **C**ost-efficient and **A**ccurate **T**ernary **Q**uantization, to compress LLMs. |
Shigeng Wang; Chao Li; Yangyuxuan Kang; Jiawei Fan; Anbang Yao; | code |
| 1296 | Differentiable Optimization Layers for Guaranteed Fairness in Deep Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce the concept of a “fairness layer”: a differentiable optimization layer appended to a model’s output layer that guarantees a chosen notion of output parity is satisfied when integrated into a neural network. |
David Troxell; Noah Roemer; Guido Montufar; | code |
| 1297 | Bounded Hyperbolic Tangent: A Stable and Efficient Alternative to Pre-Layer Normalization in Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To jointly address stability and efficiency, we propose Bounded Hyperbolic Tanh (BHyT), a drop-in replacement for Pre-LN. |
Hoyoon Byun; Youngjun Choi; Taero Kim; Sungrae Park; Kyungwoo Song; | code |
| 1298 | NeuralFLoC: Neural Flow-Based Joint Registration and Clustering of Functional Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present NeuralFLoC, a fully unsupervised, end-to-end deep learning framework for joint functional registration and clustering based on Neural ODE-driven diffeomorphic flows and spectral clustering. |
Xinyang Xiong; Siyuan Jiang; PENGCHENG ZENG; | code |
| 1299 | Beyond Point-wise Neural Collapse: A Topology-Aware Hierarchical Classifier for Class-Incremental Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Consequently, classes manifest as complex manifolds rather than collapsed points, rendering the single-point NCM suboptimal. To address this, we propose Hierarchical-Cluster SOINN (HC-SOINN), a novel classifier that captures the topological structure of these manifolds via a “local-to-global” representation. |
HuiYu Yi; Xu Zhiming; Dunwei Tu; Zhicheng Wang; Baile Xu; Furao Shen; | code |
| 1300 | A Robust Optimization Guided Pruning Framework for Vision and Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, in practice, this objective is affected by multiple sources of uncertainty, including noise in the calibration data and variability introduced by algorithmic updates. To address these issues, we introduce RobOP, a robust optimization framework that explicitly accounts for such uncertainties. |
Gabriel Afriat; Hussein Hazimeh; Dimitris Paparas; Rahul Mazumder; | code |
| 1301 | FLIPS: Instance-Fingerprinting for LLMs Via Pseudo-random Sequences Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce instance-level fingerprinting, a regulator-oriented paradigm that distinguishes configurations of the same LLM. |
Richardeau Gurvan; Gohar Dashyan; Erwan Le Merrer; Gilles Tredan; | code |
| 1302 | A$^2$SG: Adaptive and Asymmetric Surrogate Gradients for Training Deep Spiking Neural Network Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Training deep spiking neural networks (SNNs) remains challenging due to sharp loss landscapes and temporal inconsistency caused by surrogate gradients. To address these challenges, we propose a unified framework: adaptive and asymmetric surrogate gradients (A$^2$SG). |
Yechan Kang; Yongjin Kweon; Mingyeong Seo; Sohee Park; Yeonguk jeon; Jongkil Park; Hyun Jae Jang; Jaewook Kim; Yeonjoo Jeong; Suyoun Lee; Seongsik Park; | code |
| 1303 | CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present CaPX, an open-access framework for systematically studying Code-as-Policy agents in robot manipulation. |
Letian Fu; Justin Yu; Karim El-Refai; Ethan Kou; Haoru Xue; Huang Huang; Wenli Xiao; Li Fei-Fei; Guanya Shi; Jiajun Wu; S. Sastry; Yuke Zhu; Ken Goldberg; Jim Fan; | code |
| 1304 | Linear Causal Representation Learning By Topological Ordering, Pruning, and Disentanglement Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a novel linear CRL algorithm that, unlike most existing linear CRL methods, operates under weaker assumptions about environment heterogeneity and data-generating distributions while still recovering latent causal features up to an equivalence class. |
Chen; Lin Liu; Yu Guang Wang; | code |
| 1305 | Curriculum Reinforcement Learning for Black-Box Prompt Tuning Via Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing methods fail to simultaneously address the dual challenges of prompt interpretability and query efficiency. To address these challenges, we propose CRL-BPT, a curriculum reinforcement learning framework that utilizes a large language model as an agent to generate human-readable prompts. |
Gong; Chaoran Cui; Xiaolin Dong; Chunyun Zhang; Linwei Fan; | code |
| 1306 | COPF: An Online Framework for Deployment-Stable Counterfactual Fairness in Evolving Graphs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce COPF, a decision-layer framework for deployment-stable fairness monitoring and control in online link recommendation. |
Sheng'en Li; Dongmian Zou; | code |
| 1307 | Advancing Analytic Class-Incremental Learning Through Vision-Language Calibration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we first conduct a systematic study to dissect the failure modes of PTM-based analytic CIL, identifying representation rigidity as the primary bottleneck. Motivated by these insights, we propose **VILA**, a novel dual-branch framework that advances analytic CIL via a two-level vision-language calibration strategy. |
Binyu Zhao; Wei ZHANG; Xingrui Yu; Zhaonian Zou; Ivor Tsang; | code |