Paper Digest: ICML 2026 Papers & Highlights
Search within ICML-2026
Literature review on a topic
Generate a written review of ICML-2026 research on any topic, with each claim cited to specific papers.
Browse & explore
Browse ~ 25,000 authors (ICML-2026), or explore the “Best Paper” Digest listing the most influential ICML papers of recent years.
Note: ICML-2026 accepts more than 6,500 papers, this page only includes 500 of them selected by our daily paper digest algorithm. Interested users can choose to read All 6,500 ICML-2026 papers in a separate page, which takes quite some time to load.
Since 2018, Paper Digest has built a foundation of data spanning decades of conferences, journals, and research topics. The platform features a daily digest service that sifts through tens of thousands of new papers, clinical trials, news articles, and community posts, filtering the noise to highlight what matters most to specific interests. Beyond daily updates, dozens of built-in research tools streamline the academic workflow, supporting efficient reading and writing, comprehensive literature reviews, and automated research report generation.
Paper Digest Team
New York City, New York, 10017
team@paperdigest.org
TABLE 1: Paper Digest: ICML 2026 Papers & Highlights
| Paper | Author(s) | |
|---|---|---|
| 1 | You Can Learn Tokenization End-to-End with Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Prior work has shown promising results at scale in bringing this compression step inside the LLMs’ architecture with heuristics to draw token boundaries, and also attempts to learn these token boundaries with straight-through estimates, which treat the problem of drawing discrete token boundaries as a continuous one. We show that these token boundaries can instead be learned using score function estimates, which have tighter theoretical guarantees due to directly optimizing the problem of drawing discrete token boundaries to minimize loss. |
Sam Dauncey; Roger Wattenhofer; |
| 2 | Unified Multimodal Autoregressive Modeling with Shared Context—Visual Tokenizer Is Key to Unification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing approaches typically rely on two disparate visual tokenizers, which splits the representation space and hinder truly unified modeling. We propose UniAR, a unified autoregressive framework where a single discrete visual tokenizer serves as the key bridge between understanding and generation, enabling a shared context in which the model can directly interpret its own generated visual tokens without additional re-encoding. |
Wujian Peng; Lingchen Meng; Yuxuan Cai; Xianwei Zhuang; Yuhuan Yang; Rongyao Fang; Chenfei Wu; Junyang Lin; Zuxuan Wu; Shuai Bai; |
| 3 | Set Diffusion: Interpolating Token Orderings Between Autoregression and Diffusion for Fast and Flexible Decoding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our key insight is that interpolating generation orderings between autoregression and fully-random decoding, rather than committing to a fixed block length, offers a better interpolation between diffusion and AR. |
Marianne Arriola; Volodymyr Kuleshov; |
| 4 | Joint-Embedding Predictive Learning of Latent Market States in U.S. Equities Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We investigate whether Joint-Embedding Predictive Architectures (JEPA) can learn useful representations of U.S. equity markets. |
Simon Mahns; Randall Balestriero; Mahmoud Assran; |
| 5 | $\tau^2$-Bench: Evaluating Conversational Agents in A Dual-Control Environment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This differs from real-world scenarios like technical support, where users need to actively participate in modifying the state of the (shared) world. In order to address this gap, we introduce $\tau^2$-bench, with four key contributions: 1. A novel **Telecom dual-control domain** modeled as a Dec-POMDP, where both agent and user make use of tools to act in a shared, dynamic environment that tests both agent coordination and communication, 2. |
Victor Barres; Honghua Dong; Soham Ray; Xujie Si; Karthik Narasimhan; |
| 6 | PlotCraft: Pushing The Limits of LLMs for Complex and Interactive Data Visualization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, their ability to create complex visualizations for scaled and structured data remains largely unevaluated and underdeveloped. To address this gap, we introduce **PlotCraft**, a new benchmark featuring 1k challenging visualization tasks that cover a wide range of topics, such as finance, scientific research, and sociology. |
Jiajun Zhang; Jianke Zhang; Zeyu Cui; Jiaxi Yang; Lei Zhang; Zilei Wang; Qiang Liu; Liang Wang; Binyuan Hui; Junyang Lin; |
| 7 | Revealing Behavioral Plasticity in Large Language Models: A Token-Conditional Perspective Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we reveal that Large Language Models (LLMs) possess intrinsic behavioral plasticity—akin to chameleons adapting their coloration to environmental cues—that can be *exposed* through token-conditional generation and *stabilized* via reinforcement learning. |
Liyuan Mao; Le Yu; Jing Zhou; Chujie Zheng; Bowen Yu; Chang Gao; Shixuan Liu; An Yang; Weinan Zhang; Junyang Lin; |
| 8 | TabICooL: A Better, Faster, Scalable, and Open Tabular Foundation Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce TabICooL, a new state-of-the-art foundation model for regression and classification built on three pillars: (1) a novel synthetic data generation engine designed for high pretraining diversity; (2) various architectural innovations, including a new scalable softmax in attention improving generalization to larger datasets without prohibitive long-sequence pretraining; and (3) optimized pretraining protocols, notably replacing AdamW with the Muon optimizer. |
Jingang QU; David Holzmüller; Gael Varoquaux; Marine Le Morvan; |
| 9 | Adversarial Flow Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present adversarial flow models, a class of generative models that belongs to both adversarial and flow families of models. |
Shanchuan Lin; Ceyuan Yang; Zhijie Lin; Hao Chen; Haoqi Fan; |
| 10 | D2: Improved Techniques for Training Reasoning Diffusion Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Here, we introduce d2, a reasoning framework tailored for masked DLMs. |
Guanghan Wang; Gilad Turok; Yair Schiff; Marianne Arriola; Volodymyr Kuleshov; |
| 11 | Towards Execution-Grounded Automated AI Research Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We analyze two methods to learn from the execution feedback: evolutionary search and reinforcement learning. |
Chenglei Si; Zitong Yang; Yejin Choi; Emmanuel J Candes; Diyi Yang; Tatsunori Hashimoto; |
| 12 | Reinforcement Learning with Evolving Rubrics for Deep Research Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Deep research agents perform multi-step research to produce long-form, well-attributed answers. |
Rulin Shao; Akari Asai; Shannon Shen; Hamish Ivison; Varsha Kishore; Jingming Zhuo; Xinran Zhao; Molly Park; Samuel Finlayson; David Sontag; Tyler Murray; Sewon Min; Pradeep Dasigi; Luca Soldaini; Faeze Brahman; Scott Yih; Sherry Tongshuang Wu; Luke Zettlemoyer; Yoon Kim; Hannaneh Hajishirzi; Pang Wei Koh; |
| 13 | Monitoring Monitorability Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose three evaluation archetypes (intervention, process, and outcome-property), a new monitorability metric, and a broad evaluation suite. |
Melody Guan; Miles Wang; Micah Carroll; Zehao Dou; Annie Wei; Marcus Williams; Benjamin Arnav; Joost Huizinga; Ian Kivlichan; Amelia Glaese; Jakub Pachocki; Bowen Baker; |
| 14 | SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present SWE-Bench Pro, a comprehensive benchmark designed to evaluate software engineering capabilities through complex, realistic programming challenges. |
Xiang Deng; Jeff Da; Edwin Pan; Yannis He; Charles Ide; Kanak Garg; Niklas Lauffer; Andrew Park; Chetan Rane; Karmini Sampath; Maya Krishnan; Srivatsa Kundurthy; Sean Hendryx; Zifan Wang; Chen Bo Calvin Zhang; Noah Jacobson; Bing Liu; Brad Kenstler; |
| 15 | LOCA-bench: Benchmarking Language Agents Under Controllable and Extreme Context Growth Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In realistic scenarios, however, LLMs often need to act as agents that explore environments, follow instructions and plans, extract useful information, and predict correct actions under a dynamically growing context. To assess language agents in such settings, we introduce LOCA-bench (a benchmark for **LO**ng-**C**ontext **A**gents). |
Weihao Zeng; Yuzhen Huang; Junxian He; |
| 16 | Position: The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose criteria for an adapted verification infrastructure that emphasizes observable-by-default workflows, scalable verification, and clear attribution. |
Belinda Mo; |
| 17 | Spurious Rewards: Rethinking Training Signals in RLVR Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: We show that reinforcement learning with verifiable rewards (RLVR) can elicit strong mathematical reasoning in certain language models even with spurious rewards that have little, no, or outright negative correlation with the correct answer. |
Rulin Shao; Stella Li; Rui Xin; Scott Geng; Yiping Wang; Sewoong Oh; Simon Du; Nathan Lambert; Sewon Min; Ranjay Krishna; Yulia Tsvetkov; Hannaneh Hajishirzi; Pang Wei Koh; Luke Zettlemoyer; |
| 18 | Self-Distillation Enables Continual Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Self-Distillation Fine-Tuning (SDFT), a simple method that enables on-policy learning directly from demonstrations. |
Idan Shenfeld; Mehul Damani; Jonas Hübotter; Pulkit Agrawal; |
| 19 | $\tau$-Knowledge: Evaluating Conversational Agents Over Unstructured Knowledge Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Yet most existing benchmarks evaluate retrieval or tool use in isolation, and rarely test whether agents can operationalize non-parametric knowledge to drive outcomes over long-horizon conversations. To remedy this, we introduce $\tau$-Knowledge, an extension of $\tau$-Bench that evaluates agents in environments where task success requires retrieving, reasoning over, and applying knowledge from a natural-language corpus. |
Quan Shi; Alexandra Zytek; Pedram Razavi; Karthik Narasimhan; Victor Barres; |
| 20 | $\tau$-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce $\tau$-voice, a benchmark for evaluating voice agents on grounded tasks with real-world complexity: agents must navigate complex multi-turn conversations, adhere to domain policies, and interact with the environment. |
Soham Ray; Keshav Dhandhania; Victor Barres; Karthik Narasimhan; |
| 21 | Anchoring Self-Play for Code Repair Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We aim to scale supervision for code repair by having an LM generate bug–fix tasks with unconstrained edits, using unit tests as the only verifier. |
Caroline Choi; Zeyneb Kaya; Shirley Wu; Tengyu Ma; Tatsunori Hashimoto; Ludwig Schmidt; |
| 22 | Learning to Discover at Test Time Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This form of continual learning is quite special, because its goal is to produce one great solution rather than many good ones on average, and to solve this very problem rather than generalize to other problems. Therefore, our learning objective and search subroutine are designed to prioritize the most promising solutions. |
Mert Yuksekgonul; Daniel Koceja; Xinhao Li; Federico Bianchi; Jed McCaleb; Xiaolong Wang; Jan Kautz; Yejin Choi; James Zou; Carlos Guestrin; Yu Sun; |
| 23 | GDPO: Group Reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, recent work has defaulted to apply Group Relative Policy Optimization (GRPO) under multi-reward setting without examining its suitability. In this paper, we demonstrate that directly applying GRPO to normalize distinct rollout reward combinations causes them to collapse into identical advantage values, reducing the resolution of the training signal and resulting in suboptimal convergence and, in some cases, early training failure. |
Shih-Yang Liu; Xin Dong; Ximing Lu; Shizhe Diao; Peter Belcak; Mingjie Liu; Min-Hung Chen; Hongxu Yin; Yu-Chiang Wang; Kwang-Ting Cheng; Yejin Choi; Jan Kautz; Pavlo Molchanov; |
| 24 | One-step Latent-free Image Generation with Pixel Mean Flows Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent advances have made encouraging progress on each aspect individually, paving the way toward one-step diffusion/flow without latents. In this work, we take a further step towards this goal and propose pixel MeanFlow (pMF). |
Yiyang Lu; Susie Lu; Qiao Sun; Hanhong Zhao; Zhicheng Jiang; Xianbang Wang; Tianhong Li; Zhengyang Geng; Kaiming He; |
| 25 | ScDiVa: Masked Discrete Diffusion for Joint Modeling of Single-Cell Identity and Expression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Single-cell RNA-seq profiles are high-dimensional, sparse, and unordered, causing autoregressive generation to impose an artificial ordering bias and suffer from error accumulation. To address this, we propose scDiVa, a masked discrete diffusion foundation model that aligns generation with the dropout-like corruption process by defining a continuous-time forward masking mechanism in token space. |
Mingxuan Wang; Gaoyang Jiang; ZiJia Ren; Lu Shi; Cheng Chen; Chuangxin Zhao; Yanbiao Ma; |
| 26 | Flex-Forcing: Towards A Unified Autoregressive and Bidirectional Video Diffusion Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Flex-Forcing, a unified training and inference framework that enables a video diffusion model to seamlessly operate under both bidirectional and autoregressive generation regimes. |
Xinyin Ma; Julius Berner; Chao Liu; Arash Vahdat; Weili Nie; Xinchao Wang; |
| 27 | Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, on-policy distillation typically requires a separate, often larger, teacher LLM and does not explicitly leverage ground-truth solutions available in reasoning datasets. Inspired by the intuition that a sufficiently capable LLM can rationalize external privileged reasoning traces and teach its weaker self (i.e., the version without access to privileged information), we introduce On-Policy Self-Distillation (OPSD), a framework where a single model acts as both teacher and student by conditioning on different contexts. |
Siyan Zhao; Zhihui Xie; Mengchen Liu; Jing Huang; Guan Pang; Feiyu Chen; Aditya Grover; |
| 28 | Beyond Scalar Rewards: Learning from Text Feedback in LLM Post-Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Therefore, models must learn to internalize the feedback in order to improve their test-time single-turn performance. To do this, we propose two methods: Self Distillation, which trains the single-turn policy to match its own feedback-conditioned second-turn generations; and Feedback Modeling, which predicts the feedback as an auxiliary objective. |
Yuda Song; Lili Chen; Fahim Tajwar; REMI MUNOS; Deepak Pathak; J. Bagnell; Aarti Singh; Andrea Zanette; |
| 29 | Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL Finetuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This finetuning step has proved critical in achieving human or super-human performance, yet while much attention has been given to developing more effective finetuning algorithms, little attention has been given to ensuring the pretrained policy is an effective initialization for RL finetuning. In this work we seek to understand how the pretrained policy affects finetuning performance, and how to pretrain policies in order to ensure they are effective initializations for finetuning. |
Andrew Wagenmaker; Perry Dong; Raymond Tsao; Chelsea Finn; Sergey Levine; |
| 30 | Maximum Likelihood Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce **Maximum Likelihood Reinforcement Learning (MaxRL)**, a compute-indexed family of sampling-based objectives derived from a pass@k expansion of the likelihood, which interpolates between standard RL and exact maximum likelihood as compute increases. |
Fahim Tajwar; Guanning Zeng; Yueer Zhou; Yuda Song; Daman Arora; Yiding Jiang; Jeff Schneider; Russ Salakhutdinov; Haiwen Feng; Andrea Zanette; |
| 31 | Implicit Intelligence – Evaluating Agents on What Users Don’t Say Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present **Implicit Intelligence**, an evaluation framework testing whether AI agents can move beyond prompt-following to become genuine goal-fulfillers, paired with **Agent-as-a-World (AaW)**, a harness where interactive worlds are defined in human-readable YAML files and simulated by language models. |
Ved Sirdeshmukh; Marc Wetter; |
| 32 | CyberCycle: Scalable Real-World Benchmark for AI Agents’ End-to-End Cybersecurity Capabilities Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing cybersecurity evaluations of AI systems are limited in scale or scope, and fail to capture the end-to-end lifecycle of real-world software vulnerability discovery and remediation. To address this gap, we propose CyberCycle, a large-scale and realistic end-to-end cybersecurity benchmark that comprehensively evaluates AI agents’ abilities across the full lifecycle of vulnerability discovery, PoC generation, and patch generation. |
Tianneng Shi; Robin Rheem; Dongwei Jiang; Francisco De La Riega; Mona Wang; Zhun Wang; Jingzhi Jiang; Alexander Cheung; Sean Tai; Jonah Cha; Jianhong Tu; Gabriel Han; Chenguang Wang; Wenbo Guo; Jingxuan He; Dawn Song; |
| 33 | DexMachina: Functional Retargeting for Bimanual Dexterous Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We focus on long-horizon, bimanual tasks with articulated objects, which are challenging due to large action space, spatiotemporal discontinuities, and the embodiment gap between human and robot hands. We propose DexMachina, a novel curriculum-based algorithm: the key idea is to use virtual object controllers with decaying strength: an object is first driven automatically towards its target states, such that the policy can gradually learn to take over under motion and contact guidance. |
Zhao Mandi; Yifan Hou; Dieter Fox; Yashraj Narang; Ajay Mandlekar; shuran song; |
| 34 | ToolOrchestra: Elevating Intelligence Via Efficient Model and Tool Orchestration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce ToolOrchestra, a method for training small orchestrators that coordinate the use of intelligent tools. |
Hongjin SU; Shizhe Diao; Ximing Lu; Mingjie Liu; Jiacheng Xu; Xin Dong; Yonggan Fu; Peter Belcak; Hanrong Ye; Hongxu Yin; Yi Dong; Evelina Bakhturina; Tao Yu; Yejin Choi; Jan Kautz; Pavlo Molchanov; |
| 35 | Extracting Alignment Data in Open Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we show that it is possible to extract significant amounts of alignment training data from a post-trained model — useful to steer the model to improve certain capabilities such as long-context reasoning, safety, instruction following, and maths. |
Federico Barbero; Xiangming Gu; Christopher A. Choquette Choo; Chawin Sitawarin; Matthew Jagielski; Itay Yona; Petar Veličković; Ilia Shumailov; Jamie Hayes; |
| 36 | Shrinking The Variance: Shrinkage Baselines for Reinforcement Learning with Verifiable Rewards Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Drawing inspiration from Stein’s paradox, we propose using \emph{shrinkage estimators} that combine \emph{per-prompt} and \emph{across-prompt} means to improve the overall per-prompt mean estimation accuracy—particularly in the low-generation regime typical of RLVR. |
Guanning Zeng; Zhaoyi Zhou; Daman Arora; Andrea Zanette; |
| 37 | VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The goal of this paper is to improve the performance and reliability of vision-language-action (VLA) models through iterative online interaction. |
Yanjiang Guo; Tony Lee; Lucy Xiaoyang Shi; Jianyu Chen; Percy Liang; Chelsea Finn; |
| 38 | Pretrained Vision-Language-Action Models Are Surprisingly Resistant to Forgetting in Continual Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we find that pretrained VLAs are remarkably resistant to forgetting compared with smaller policy models trained from scratch. |
Huihan Liu; Changyeon Kim; Bo Liu; Minghuan Liu; Yuke Zhu; |
| 39 | A Mechanistic Understanding of Sim-and-Real Co-Training in Generative Policies Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose an explanation that when simulation and real-world data are combined with a balanced mixing ratio, co-training naturally learns representations that are aligned across domains while remaining domain-distinguishable, enabling effective knowledge transfer without sacrificing real-world adaptation, which we refer to as structured representation alignment. |
Yu Lei; Minghuan Liu; Abhiram Maddukuri; Zhenyu Jiang; Yuke Zhu; |
| 40 | GenExam: A Multidisciplinary Text-to-Image Exam Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce GenExam, the first benchmark for multidisciplinary text-to-image exams, featuring 1,000 samples across 10 subjects with exam-style prompts organized under a four-level taxonomy. |
Zhaokai Wang; Penghao Yin; Xiangyu Zhao; Changyao Tian; Yu Qiao; Wenhai Wang; Jifeng Dai; Gen Luo; |
| 41 | Solving Physics Olympiad Via Reinforcement Learning on Physics Simulators Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we show that physics simulators can serve as a powerful alternative source of supervision for training LLMs for physical reasoning. |
Mihir Prabhudesai; Aryan Satpathy; Yangmin Li; Zheyang Qin; Nikash Bhardwaj; Amir Zadeh; Chuan Li; Katerina Fragkiadaki; Deepak Pathak; |
| 42 | Mode Seeking Meets Mean Seeking for Long Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While multi-resolution image training works because higher resolution is largely an interpolation of the same underlying patch distribution, training across video lengths is fundamentally different: a longer video is an extrapolation that must invent new events and causal structure beyond the short-clip horizon. To address this, we propose a training paradigm where Mode Seeking meets Mean Seeking, decoupling local fidelity from long-term coherence from a unified representation via a Decoupled Diffusion Transformer. |
Shengqu Cai; Weili Nie; Chao Liu; Julius Berner; Lvmin Zhang; Nanye Ma; Hansheng Chen; Maneesh Agrawala; Leonidas Guibas; Gordon Wetzstein; Arash Vahdat; |
| 43 | Optimal Classical and Quantum Algorithms for Gradient Testing and Estimation By Comparisons Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: For any smooth $f\colon\mathbb R^n\to\mathbb R$, $\mathbf{x}\in\mathbb R^n$, and $\varepsilon>0$, we design a gradient testing algorithm that determines whether the normalized gradient $\nabla f(\mathbf{x})/||\nabla f(\mathbf{x})||$ is $\varepsilon$-close or $2\varepsilon$-far from a given unit vector $\mathbf{v}$ using $O(n)$ queries, as well as a gradient estimation algorithm that outputs an $\varepsilon$-estimate of $\nabla f(\mathbf{x})/||\nabla f(\mathbf{x})||$ using $O(n\log(1/\varepsilon))$ queries. |
Xiwen Tao; Chenyi Zhang; Helin Wang; Yexin Zhang; Tongyang Li; |
| 44 | DeepAnalyze: Agentic Large Language Models for Autonomous Data Science Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce DeepAnalyze, the first agentic LLM for autonomous data science, capable of automatically completing the end-to-end data science from structured data to analyst-grade research reports. |
Shaolei Zhang; Ju Fan; Meihao Fan; Yizhe Liu; Yuxin Zhang; Xiaoyong Du; |
| 45 | On The Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A central challenge is the lack of control in modern training pipelines: large-scale pre-training corpora are opaque, mid-training is often underexamined, and RL objectives interact with unknown prior knowledge in complex ways. To resolve this ambiguity, we develop a fully controlled experimental framework that isolates the causal contributions of pre-training, mid-training, and RL-based post-training. |
Charlie Zhang; Graham Neubig; Xiang Yue; |
| 46 | The Geometry of Reasoning: Self-Evaluation Via Layerwise Trajectory Evolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce \ourmethod (Geometry of Reasoning), a white-box self-evaluation framework based on layerwise trajectory evolution. |
Jinhe Bi; Danqi Yan; Yifan Wang; Wenke Huang; Haokun Chen; Guancheng Wan; Mang Ye; Xun Xiao; Hinrich Schuetze; Volker Tresp; Yunpu Ma; |
| 47 | EchoRL: Reinforcement Learning Via Rollout Echoing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, inspired through analyzing the entropy pattern behind golden trajectories produced by external expert models, we propose EchoRL for better exploiting the advantage-degenerated rollouts to further improve the training performance. |
Jinhe Bi; Aniri; Minglai Yang; Xingcheng Zhou; Wenke Huang; Sikuan Yan; Yujun Wang; Zixuan Cao; Michael Färber; Xun Xiao; Volker Tresp; Yunpu Ma; |
| 48 | Does Reinforcement Fine-Tuning Improve Generalization of LLM Agents? An Empirical Study Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In real-worlddeployment, agents may operate in unseen environments with different background knowledge, observation spaces, and action interfaces. To characterize the generalization profile of RFT under such shifts, we conduct a systematic study along three axes: (1) within-environment generalization across task difficulty, (2) cross-environment transfer to unseen environments, and (3) sequential multi-environment training to quantify transfer and forgetting. |
Zhiheng Xi; Xin Guo; Jiaqi Liu; Jiazheng Zhang; Yutao Fan; Zhihao Zhang; Shichun Liu; Mingxu Chai; Xiaowei Shi; Yitao Zhai; Xunliang Cai; Tao Gui; Qi Zhang; Xuanjing Huang; |
| 49 | Reasoning Models Struggle to Control Their Chains of Thought Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This capability — CoT controllability — is undesirable because it could allow models to suppress signs of misbehavior in their CoT, thereby undermining our ability to monitor them. To measure this, we introduce the \emph{CoT-Control} evaluation suite. |
Chen Yueh-Han; Robert McCarthy; Bruce W. Lee; He He; Micah Carroll; Tomasz Korbak; |
| 50 | SSA: Sparse Sparse Attention By Aligning Full and Sparse Attention Outputs in Feature Space Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose SSA (Sparse Sparse Attention), a training framework that integrates both sparse and full attention with bidirectional attention-output alignment. |
Zhenyi Shen; Junru Lu; Lin Gui; Jiazheng Li; Yulan He; di yin; Xing Sun; |
| 51 | IsoCompute Playbook: Optimally Scaling Sampling Compute for LLM RL Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study the compute-optimal allocation of sampling compute for on-policy RL methods in LLMs, framing scaling as a compute-constrained optimization over three resources: parallel rollouts per problem, number of problems per batch, and number of update steps. |
Zhoujun Cheng; Yutao Xie; Yuxiao Qu; Amrith Setlur; Shibo Hao; Varad Pimpalkhute; Tongtong Liang; Feng Yao; Zhengzhong Liu; Eric Xing; Virginia Smith; Russ Salakhutdinov; Zhiting Hu; Taylor W. Killian; Aviral Kumar; |
| 52 | Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Moreover, they are often limited in the reasoning depth and search breadth, making it difficult to solve complex questions that require aggregating evidence from diverse visual and textual sources. Building on this, we propose Vision-DeepResearch, which proposes one new multimodal deep-research paradigm, i.e., performs multi-turn, multi-entity and multi-scale visual and textual search to robustly hit real-world search engines under heavy noise. |
Wenxuan Huang; Yu Zeng; Qiuchen Wang; Zhen Fang; Shaosheng Cao; Zheng Chu; Qingyu Yin; Shuang Chen; Zhenfei Yin; Lin Chen; Zehui Chen; Yao Hu; Phil Torr; Feng Zhao; Wanli Ouyang; |
| 53 | Quant VideoGen: Auto-Regressive Long Video Generation Via 2-Bit KV-Cache Quantization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: More critically, memory-bounded KV budgets constrain the effective working memory, directly degrading long-horizon consistency in identity, layout, and motion. To address this challenge, we present Quant VideoGen (QVG), a training-free KV-cache quantization framework for auto-regressive video diffusion models. |
Haocheng Xi; Shuo Yang; Yilong Zhao; Muyang Li; Han Cai; Xingyang Li; Yujun Lin; Zhuoyang Zhang; Jintao Zhang; Xiuyu Li; Zhiying Xu; Jun Wu; Chenfeng Xu; Ion Stoica; Song Han; Kurt Keutzer; |
| 54 | Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce *Self-Flow*: a self-supervised flow matching paradigm that integrates representation learning within the generative framework. |
Hila Chefer; Patrick Esser; Dominik Lorenz; Dustin Podell; Vikash Raja; Vinh Tong; Antonio Torralba; Robin Rombach; |
| 55 | Real-Time Visual Attribution Streaming in Thinking Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present an amortized framework for real-time visual attribution streaming in multimodal thinking models. |
Seil Kang; Woojung Han; Junhyeok Kim; Jinyeong Kim; Youngeun Kim; Seong Jae Hwang; |
| 56 | How2Everything: Mining The Web for How-to Procedures to Evaluate and Improve LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Yet, measuring and improving procedural validity at scale on real-world tasks remains challenging and understudied. To address this, we introduce How2Everything, a scalable framework to evaluate and improve goal-conditioned procedure generation. |
Yapei Chang; Kyle Lo; Mohit Iyyer; Luca Soldaini; |
| 57 | Reasoning About Reasoning: BAPO Bounds on Chain-of-Thought Token Complexity in LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We address a fundamental theoretical question: *how many* reasoning tokens are required to solve a problem as input size grows? |
Kiran Tomlinson; Tobias Schnabel; Adith Swaminathan; Jennifer Neville; |
| 58 | Reinforcement Learning with Verifiable Rewards: GRPO’s Loss, Dynamics, and Success Amplification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Group Relative Policy Optimization (GRPO) was introduced recently and used to train DeepSeek\textendash R1 for promoting reasoning in LLMs under verifiable (binary) rewards. We show that the mean{+}variance calibration of these rewards induces a contrastive loss in which the contrastive samples are synthetic data drawn from the previous policy. |
Youssef Mroueh; |
| 59 | ThetaEvolve: Test-time Learning on Open Problems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce ThetaEvolve, an open-source framework that simplifies and extends AlphaEvolve to efficiently scale both in-context learning and Reinforcement Learning (RL) at test time, allowing models to continually learn from their experiences in improving open optimization problems. |
Yiping Wang; Shao-Rong Su; Zhiyuan Zeng; Eva Xu; Liliang Ren; Xinyu Yang; Zeyi Huang; Xuehai He; Luyao Ma; Baolin Peng; Hao Cheng; Pengcheng He; Weizhu Chen; Shuohang Wang; Simon Du; Yelong Shen; |
| 60 | DreamDojo: A Real-Time Robot World Model from Large-Scale Human Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, modeling these world dynamics, especially for dexterous robotics tasks, poses significant challenges due to limited data coverage and scarce action labels. As an endeavor towards this end, we introduce DreamDojo, a foundation world model that learns diverse interactions and dexterous controls from 44k hours of egocentric human videos. |
Shenyuan Gao; William Liang; Kaiyuan Zheng; Ayaan Malik; Seonghyeon Ye; Sihyun Yu; Wei-Cheng Tseng; Yuzhu Dong; Kaichun Mo; Chen-Hsuan Lin; Jiannan Xiang; Yuqi Xie; Ruijie Zheng; Dantong Niu; Pooya Jannaty; Jinwei Gu; Jun Zhang; Jitendra Malik; Pieter Abbeel; Ming-Yu Liu; Yuke Zhu; Joel Jang; Jim Fan; |
| 61 | LoSA: Locality Aware Sparse Attention in Diffusion Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this challenge, we observe that block-wise diffusion exhibits locality of representation changes across denoising steps: only a small fraction of tokens (active tokens) undergo significant hidden-state updates, while most tokens (stable tokens) remain nearly unchanged. Based on this insight, we propose LoSA (Locality-aware Sparse Attention), which reuses cached prefix-attention results for stable tokens and applies sparse attention only to active tokens with large representation changes. |
Haocheng Xi; Harman Singh; Yuezhou Hu; Coleman Hooper; Rishabh Tiwari; Aditya Tomar; Wonjun Kang; Minjae Lee; Michael Mahoney; Chenfeng Xu; Kurt Keutzer; Amir Gholaminejad; |
| 62 | The Assistant Axis: Situating and Stabilizing The Default Persona of Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Large language models can represent a variety of personas but typically default to a helpful Assistant identity cultivated during post-training. Across several different models, we find an “Assistant Axis in their activation space, which captures the extent to which a model is operating in its default Assistant mode. |
Christina Lu; Jack Gallagher; Jonathan Michala; Kyle Fish; Jack Lindsey; |
| 63 | PACER: Acyclic Causal Discovery from Large-scale Interventional Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce PACER (Perturbation-driven Acyclic Causal Edge Recovery), a scalable framework for causal discovery that guarantees acyclicity by construction. |
Ramon Viñas Torné; Sílvia Salazar; Soyon Park; Ivo Ban; Artyom Gadetsky; Nikita Doikov; Maria Brbic; |
| 64 | Activation Oracles: Training and Evaluating LLMs As General-Purpose Activation Explainers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we instead take a generalist perspective. |
Adam Karvonen; James Chua; Clément Dumas; Kit Fraser-Taliente; Subhash Kantamneni; Julian Minder; Euan Ong; Arnab Sen Sharma; Daniel Wen; Owain Evans; Samuel Marks; |
| 65 | Base Models Know How to Reason, Thinking Models Learn When Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Why do thinking language models outperform their base counterparts, and what exactly do they learn during training? We introduce constructive model diffing, a framework for understanding fine-tuned models by explicitly constructing the base-to-fine-tuned difference from interpretable components to produce hybrid models, and measuring how well they recover the fine-tuned model’s performance. |
Constantin Venhoff; Iván Arcuschin; Phil Torr; Arthur Conmy; Neel Nanda; |
| 66 | Experience Augmented Policy Optimization for LLM Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose Experience-Augmented Policy Optimization (EAPO), which leverages a prior RL-optimized policy as an action-level experience prior and selectively injects experience at critical decision points during rollout. |
Jinda Lu; Kexin Huang; Junkang Wu; Shuo Yang; Jinghan Li; Chiyu Ma; Shaohang Wei; Xiang Wang; Guoyin Wang; Jingren Zhou; |
| 67 | Toward Training Superintelligent Software Agents Through Self-Play SWE-RL Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present Self-play SWE-RL (SSR), a first step toward training superintelligent software agents under minimal data assumptions. |
Yuxiang Wei; Zhiqing Sun; Emily McMilin; Jonas Gehring; David Zhang; Gabriel Synnaeve; Daniel Fried; LINGMING ZHANG; Sida Wang; |
| 68 | MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose MASOrchestra, a training-time framework that formulates MAS orchestration as a function-calling reinforcement learning problem with holistic orchestration, generating an entire MAS at once. |
Zixuan Ke; Yifei Ming; Austin Xu; Ryan Chin; Xuan-Phi Nguyen; Prathyusha Jwalapuram; Jiayu Wang; Semih Yavuz; Caiming Xiong; Shafiq Joty; |
| 69 | MmBERT: A Modern Multilingual Encoder with Annealed Language Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce mmBERT, an encoder-only language model pretrained on 3T tokens of multilingual text in over 1800 languages. |
Marc Marone; Orion Weller; William Fleshman; Eugene Yang; Dawn Lawrie; Benjamin Van Durme; |
| 70 | Entropy-Aware On-Policy Distillation of Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, we show that the mode-seeking property of reverse KL reduces generation diversity and yields unstable learning signals when the teacher distribution has high entropy. To address this, we introduce Entropy-Aware On-Policy Distillation. |
Woogyeol Jin; Taywon Min; Yongjin Yang; Swanand Kadhe; Yi Zhou; Dennis Wei; Nathalie Baracaldo; Kimin Lee; |
| 71 | ModernVBERT: Towards Smaller Visual Document Retrievers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Increasingly, Visual Document Retrieval (VDR) models, which directly embed images of document pages, are used as an alternative to text-only retrievers. |
Paul Teiletche; Quentin Macé; Max Conti; António Loison; Gautier Viaud; Pierre Colombo; Manuel Faysse; |
| 72 | Scaling Generative Verifiers For Natural Language Mathematical Proof Verification And Selection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Large language models have achieved remarkable success on final-answer mathematical problems, largely due to the ease of applying reinforcement learning with verifiable rewards. |
Sadegh Mahdavi; Branislav Kisacanin; Shubham Toshniwal; Wei Du; Ivan Moshkov; George Armstrong; Renjie Liao; Christos Thrampoulidis; Igor Gitman; |
| 73 | Automatically Finding Reward Model Biases Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce and study the research problem of automatically finding reward model biases in natural language. |
Atticus Wang; Iván Arcuschin; Arthur Conmy; |
| 74 | Clipping Bottleneck: Stabilizing RLVR Via Stochastic Recovery of Near-Boundary Signals Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, we find that many high-value signals lie in the **near-boundary** region just beyond the clipping threshold, and are thus discarded. Motivated by this diagnosis, we propose **Near-boundary Stochastic Rescue (NSR)**, a minimal, plug-and-play modification that stochastically retains these slightly out-of-bound tokens to recover lost signals. |
Shuo Yang; Jinda Lu; Chiyu Ma; Kexin Huang; Haoming Meng; Qihui Zhang; Yuyang Liu; Bolin Ding; Guoyin Wang; Li Yuan; Jingren Zhou; |
| 75 | Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in Speed Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Accordingly, we introduce a continuous pretraining scheme with a block-wise attention pattern. |
Yonggan Fu; Lexington Whalen; Zhifan Ye; Xin Dong; Shizhe Diao; Jingyu Liu; CHENGYUE WU; Hao Zhang; Enze Xie; Song Han; Maksim Khadkevich; Jan Kautz; Yingyan (Celine) Lin; Pavlo Molchanov; |
| 76 | Modular Pretraining Enables Access Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, training and deploying multiple models is prohibitively expensive. We address this challenge by developing gradient-routed mixture-of-experts (GR-MoE), a pretraining method that selectively updates experts to induce specialization. |
Ethan Roland; Murat Cubuktepe; Erick Martinez; Stijn Servaes; Keenan Pepper; Michael Vaiana; Diogo de Lucena; Judd Rosenblatt; Addie Foote; Cem Anil; Alex Cloud; |
| 77 | Training AI Co-Scientists Using Rubric Rewards Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To validate this approach, we conduct a human study for machine learning research goals spanning 225 expert hours. |
Shashwat Goel; Rishi Hazra; Dulhan Jayalath; Timon Willi; Parag Jain; Shen; Ilias Leontiadis; Francesco Barbieri; Yoram Bachrach; Jonas Geiping; Chenxi Whitehouse; |
| 78 | Curating The Future: A Scalable Recipe for Training Open-Ended Forecasters Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we train language models to make predictions on open-ended forecasting questions. |
Nikhil Chandak; Shashwat Goel; Ameya Pandurang Prabhu; Moritz Hardt; Jonas Geiping; |
| 79 | Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: This sequential setup leads to attackers overfitting obsolete exploits while defenders perpetually lag behind emerging threats. To address this, we introduce Self-RedTeam, the first fully online self-play multi-agent reinforcement learning (MARL) algorithm that continuously co-evolves attacker and defender for robust safety alignment. |
Mickel Liu; Liwei Jiang; Yancheng Liang; Simon Du; Yejin Choi; Tim Althoff; Natasha Jaques; |
| 80 | Any-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Meanwhile, recent studies have successfully applied discrete diffusion models to natural language processing, revealing their considerable potential as a promising new approach in this domain. Drawing inspiration from these pioneering researches, we introduce Any-Diffusion, the first any-to-any multimodal language model built purely on mask-based discrete diffusion models, which unifies understanding and generation across text, speech, and images. |
lijiang Li; zuwei long; Yunhang Shen; Heting Gao; Haoyu Cao; Xing Sun; Caifeng Shan; Ran He; Chaoyou Fu; |
| 81 | CSD: Content-aware Speculative Decoding for Efficient Image Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a novel content-aware speculative decoding algorithm, termed CSD, which integrates an entropy-based probability relaxation mechanism with an optimal resampling strategy to enhance the inference efficiency for autoregressive image generation. |
Mingcheng Wang; junbo qiao; Yunchen Li; Lingfu Jiang; Wei Li; Jie Hu; Jiao Xie; Zhou Yu; Xinghao Chen; Guixu Zhang; Shaohui Lin; |
| 82 | SIGMA-PPG: Statistical-prior Informed Generative Masking Architecture for PPG Foundation Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Standard masked modeling often yields trivial solutions while contrastive methods lack morphological precision. To address these limitations, we propose a Statistical-prior Informed Generative Masking Architecture (SIGMA-PPG), a generative foundation model featuring a prior-guided adversarial masking mechanism, where a reinforcement learning-driven teacher leverages statistical priors to create challenging learning paths that prevent overfitting to noise. |
ZONGHENG GUO; Tao Chen; Yang Jiao; Yi Pan; Xiao Hu; Manuela Ferrario; |
| 83 | Chain-of-Thought Reasoning In The Wild Is Not Always Faithful Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: In this work, we show that unfaithful CoT also occurs on naturally worded, non-adversarial prompts without adding artificial biases or editing model outputs. |
Iván Arcuschin; Jett Janiak; Robert Krzyzanowski; Senthooran Rajamanoharan; Neel Nanda; Arthur Conmy; |
| 84 | HumanLM: Simulating Users with State Alignment Beats Response Imitation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing user simulators mostly imitate surface-level patterns and language styles, which fails to reflect the underlying state of real users (e.g., beliefs, emotions). To address these limitations, we propose a novel training framework, HumanLM, which builds user simulators that accurately reflect real users. |
Shirley Wu; Evelyn Choi; Arpandeep Khatua; Zhanghan Wang; Joy He-Yueya; Cyril Weerasooriya; Wei Wei; Diyi Yang; Jure Leskovec; James Zou; |
| 85 | VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current multimodal benchmarks often conflate reasoning with domain-specific knowledge, making it difficult to isolate and evaluate general reasoning abilities in non-expert settings. To address this, we introduce VisualPuzzles, a benchmark that targets visual reasoning while deliberately minimizing reliance on specialized knowledge. |
Yueqi Song; Tianyue Ou; Yibo Kong; Zecheng Li; Graham Neubig; Xiang Yue; |
| 86 | From Prior to Pro: Efficient Skill Mastering Via Distribution Contractive RL Finetuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Distribution Contractive Reinforcement Learning (DICE-RL), a framework that uses reinforcement learning (RL) as a “distribution contractor” to refine pretrained generative robot policies. |
Zhanyi Sun; shuran song; |
| 87 | RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we present RoboTwin 2.0, a scalable simulation framework that enables closed-loop, automated, large-scale generation of diverse and realistic data, along with unified evaluation protocols for dual-arm manipulation. |
Tianxing Chen; Zanxin Chen; Baijun Chen; Zijian Cai; Yibin Liu; Zixuan Li; Qiwei Liang; Xianliang Lin; Yiheng Ge; Zhenyu Gu; Weiliang Deng; Yubin Guo; Tian Nian; Xuanbing Xie; Qiangyu Chen; KailunSu; Tianling Xu; Guodong Liu; Mengkang Hu; Huan-ang Gao; Kaixuan Wang; Zhixuan Liang; Yusen Qin; Xiaokang Yang; Ping Luo; Yao Mu; |
| 88 | Position: Beyond Reasoning Zombies — AI Reasoning Requires Process Validity Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To that end, we provide (1) general and extensible definitions for *valid* and *sound reasoning* based on a synthesis of the literature, which can serve as an accessible reference and a starting point for community discussion; and (2) a checklist for best practices in the communication of AI reasoning research. |
Rachel Lawrence; Jacqueline Maasch; |
| 89 | Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we systematically study reinforcement learning (RL) for kernel generation. |
Wei Liu; Jiawei Xu; Yingru Li; Longtao Zheng; Tianjian Li; Qian Liu; Junxian He; |
| 90 | What Does Flow-Matching Bring to TD-Learning? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We show that their success is not explained by distributional RL: explicitly modeling return distributions often degrades performance. Instead, we argue that flow-matching Q-functions are effective because they couple a learned velocity field with an integration procedure that is used both during training and to read out Q-values at inference time. |
Bhavya Agrawalla; Michal Nauman; Aviral Kumar; |
| 91 | Temporal Context Reinstatement Drives Episodic-Like Order Memory in Long-Context Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Here, we investigate whether and, if so, how LLMs capture core behavioral signatures of humans of a central aspect of episodic memory via a temporal order memory task. |
Mathis Pink; Vy Vo; Qinyuan Wu; Jianing Mu; Javier Turek; Uri Hasson; Kenneth Norman; Sebastian Michelmann; Alexander Huth; Mariya Toneva; |
| 92 | Towards Efficient Large Language Reasoning Models Via Extreme-Ratio Chain-of-Thought Compression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To achieve high-fidelity, fast reasoning, we propose a novel EXTreme-RAtio Chain-of-Thought Compression framework, termed Extra-CoT, which aggressively reduces the token budget while preserving answer accuracy. |
Yuntian Tang; Bohan Jia; Wenxuan Huang; Lianyue Zhang; Jiao Xie; Wenxi Li; Wei Li; Jie Hu; Xinghao Chen; Rongrong Ji; Shaohui Lin; |
| 93 | Random Scaling of Emergence Capabilities Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose that breakthroughs are instead driven by continuous changes in the *probability distribution* of training outcomes when performance is bimodally distributed across random seeds. |
Rosie Zhao; Tian Qin; David Alvarez-Melis; Sham Kakade; Naomi Saphra; |
| 94 | Benchmarking Reward Hack Detection in Code Environments Via Contrastive Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a novel taxonomy of reward exploits spanning across 54 categories and introduce TRACE (Testing Reward Anomalies in Code Environments), a synthetically curated and human-verified benchmark containing 517 testing trajectories. |
Darshan Deshpande; Anand Kannappan; Rebecca Qian; |
| 95 | DF-ExpEnse: Diffusion Filtered Exploration for Sample Efficient Finetuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work we present Diffusion Filtered Exploration via Ensembles (DF-ExpEnse), an exploration technique that meaningfully improves the quality of online experience collection, thus increasing the sample efficiency of the finetuning procedure. |
Calvin Luo; Chen Sun; shuran song; |
| 96 | TruthRL: Incentivizing Truthful LLMs Via Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present TruthRL, a general reinforcement learning (RL) framework that directly optimizes the truthfulness of LLMs. |
Zhepei Wei; Xiao Yang; Kai Sun; Jiaqi Wang; Rulin Shao; Jingxiang Chen; Mohammad Kachuee; Teja Gollapudi; Yiwei Liao; Nicolas SCHEFFER; Rakesh Wanga; Anuj Kumar; Yu Meng; Scott Yih; Xin Dong; |
| 97 | TQL: Scaling Q-Functions with Transformers By Preventing Attention Collapse Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we ask what prevents transformers from scaling effectively for value functions? |
Perry Dong; Kuo-Han Hung; Alexander Swerdlow; Dorsa Sadigh; Chelsea Finn; |
| 98 | SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce SWE-rebench V2, a language-agnostic automated pipeline for harvesting executable real-world SWE tasks and constructing RL training environments at scale. |
Ibragim Badertdinov; Maksim Nekrashevich; Anton Shevtsov; Aleksandr Golubev; |
| 99 | Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Concept Ablation Fine-Tuning (CAFT), a technique that leverages interpretability tools to control how LLMs generalize from fine-tuning, without needing to modify the training data or otherwise use data from the target distribution. |
Helena Casademunt; Caden Juang; Adam Karvonen; Samuel Marks; Senthooran Rajamanoharan; Neel Nanda; |
| 100 | Star Elastic: Many-in-One Reasoning LLMs with Efficient Budget Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce Star Elastic, a novel LLM post-training method that adds N nested submodels to a given parent reasoning model using the compute of one run (Nx savings) via a single post-training job. |
Ali Taghibakhshi; Ruisi Cai; Saurav Muralidharan; Sharath Turuvekere Sreenivas; Ameya Mahabaleshwarkar; Marcin Chochowski; Akhiad Bercovich; Ran Zilberstein; Ran El-Yaniv; Yonatan Geifman; Daniel Korzekwa; Yoshi Suhara; Oluwatobi Olabiyi; Ashwath Aithal; Nima Tajbakhsh; Pavlo Molchanov; |
| 101 | Who’s in Charge? Disempowerment Patterns in Real-World LLM Usage Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present the first large-scale empirical analysis of disempowerment patterns in real-world AI assistant interactions, analyzing 1.5 million consumer Claude.ai conversations using a privacy-preserving approach. |
Mrinank Sharma; Miles McCain; Raymond Douglas; David Duvenaud; |
| 102 | A Diagnostic Study of Multi-Agent LLMs for Real-World Debates Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce a diagnostic evaluation framework that measures debate quality by measuring both the outcome and the process. |
Priya Pitre; Gaurav Srivastava; Lu Zhang; Le Wang; Naren Ramakrishnan; Xuan Wang; |
| 103 | WISE: World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing research and evaluation standards predominantly focus on image realism and shallow text-image alignment, lacking a comprehensive assessment of complex semantic understanding and world knowledge integration in text-to-image generation. To address this challenge, we propose WISE, the first benchmark specifically designed for World Knowledge-Informed Semantic Evaluation. |
Yuwei Niu; Munan Ning; Mengren Zheng; Weiyang Jin; Bin Lin; Peng Jin; Jiaqi Liao; Chaoran Feng; Fanqing Meng; Kun-Peng Ning; Bin Zhu; Li Yuan; |
| 104 | Mem-T: Densifying Rewards for Long-Horizon Memory Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing training paradigms remain constrained: agents often traverse long-horizon sequences of memory operations before receiving sparse and delayed rewards, which hinders truly end-to-end optimization of memory management policies. To address this limitation, we introduce Mem-T, an autonomous memory agent that interfaces with a lightweight hierarchical memory database to perform dynamic updates and multi-turn retrieval over streaming inputs. |
Yanwei Yue; Guibin Zhang; Boci Peng; Xuanbo Fan; Jiaxin Guo; Qiankun Li; Yan Zhang; |
| 105 | Reasoning Cache: Learning to Extrapolate to Long Lengths Via Short-Length RL Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Standard on-policy RL operates on fixed problem distributions and training budgets, giving rise to a distribution shift between train and test that limits the resulting model’s extrapolation capabilities. To address this, we introduce RC, an iterative decoding algorithm replacing standard autoregressive decoding that enables models to extrapolate to lengths an order of magnitude longer than those seen during training. |
Ian Wu; Yuxiao Qu; Amrith Setlur; Aviral Kumar; |
| 106 | SceneSmith: Agentic Generation of Simulation-Ready Indoor Scenes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce SceneSmith, a hierarchical agentic framework that generates simulation-ready indoor environments from natural language prompts. |
Nicholas Pfaff; Thomas Cohn; Sergey Zakharov; Rick Cory; Russ Tedrake; |
| 107 | Interpreting Physics in Video World Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Here, we present the first interpretability study to directly examine physical representations inside large-scale video encoders. |
Sonia Joseph; Quentin Garrido; Randall Balestriero; Matthew Kowal; Thomas Fel; Shahab Bakhtiari; Blake Richards; Michael Rabbat; |
| 108 | Understanding Reasoning Collapse in LLM Agent Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We thus provide a signal-to-noise ratio explanation for why $I(X; Z)$ drops: when within-input reward variance $\mathrm{Var}(R \mid X)$ is low, task gradients weaken and input-agnostic regularizers (KL, entropy) dominate, flattening cross-input differences. |
Zihan (Zenus) Wang; Chi Gui; Xing Jin; Qineng Wang; Licheng Liu; Kangrui Wang; Shiqi Chen; Linjie Li; Zhengyuan Yang; Pingyue Zhang; Yiping Lu; Jiajun Wu; Li Fei-Fei; Lijuan Wang; Yejin Choi; Manling Li; |
| 109 | One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We categorize RM failures by complexity and propose a simple post-hoc intervention to mitigate low-complexity biases that arise from spurious correlations. |
Daniel Fein; Max Lamparth; Violet Xiang; Mykel Kochenderfer; Nick Haber; |
| 110 | The Flexibility Trap: Rethinking The Value of Arbitrary Order in Diffusion Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, in this paper, we reveal that for general reasoning tasks (e.g., mathematics and coding), arbitrary order generation may in fact limit the reasoning potential of dLLMs. We find that dLLMs tend to exploit this order flexibility to bypass high-uncertainty tokens that are crucial for exploration, leading to a premature collapse of solution coverage. |
Zanlin Ni; Shenzhi Wang; Yang Yue; Tianyu Yu; Weilin Zhao; Yeguo Hua; Tianyi Chen; Jun Song; YuCheng; Bo Zheng; Gao Huang; |
| 111 | RADIO1D: Elastic Representations for Condensed Vision Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Notably, models trained with image-text alignment (such as SigLIP2) develop a small number of specialized tokens that effectively summarize global image content. Building on this, we introduce RADIO1D, which compresses images into a compact, variable-length 1D token sequence using multi-teacher knowledge distillation and an autoencoder design. |
Greg Heinrich; Mike Ranzinger; Collin McCarthy; Natan Bagrov; Eugene Khvedchenya; Bryan Catanzaro; Jan Kautz; Andrew Tao; Pavlo Molchanov; |
| 112 | Scaling Long-Horizon Agent Via Context Folding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Context Folding, a framework that empowers agents to actively manage their working context. |
Weiwei Sun; Lu Miao; Zhan Ling; Kang Liu; Xuesong Yao; Yiming Yang; Jiecao Chen; |
| 113 | Position: Evaluation of ML Resource Utilization Requires Model Life Cycle Assessment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: With the growing complexity of pipelines and underlying infrastructure needed to develop and deploy AI systems, previous approaches for evaluating AI efficiency which focus on the costs of a single training run or an individual inference prediction are no longer sufficient. In this position paper, we enunciate the need for applying life cycle assessment to evaluate the costs of the machine learning model development and deployment pipeline to properly account for the required resources and downstream impact. |
Jared Fernandez; Clara Na; Yonatan Bisk; Constantine Samaras; Emma Strubell; |
| 114 | On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While RL-tuned VLMs can improve visual reasoning benchmark performance, they can still suffer from weak visual grounding, hallucinations, and over-reliance on textual cues. We show that simple, controlled textual perturbations—misleading captions or incorrect chain-of-thought (CoT) traces—cause substantial drops in robustness and confidence, and that these effects are more pronounced when CoT consistency is taken into account across open-source multimodal reasoning models. |
Rosie Zhao; Anshul Shah; Xiaoyu Zhu; Xinke Deng; Zhongyu Jiang; Yang Yang; Joerg Liebelt; Arnab Kumar Mondal; |
| 115 | Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A key challenge is that tool integration increases latency, disrupting conversational flow. To mitigate this, we propose Streaming Retrieval-Augmented Generation (Stream RAG), a novel framework that reduces latency by predicting tool queries in parallel with user speech, even before the user finishes speaking. |
Siddhant Arora; Haidar Khan; Kai Sun; Xin Dong; Sajal Choudhary; Seungwhan Moon; Xinyuan Zhang; Adithya Sagar; Surya Appini; Kaushik Patnaik; Sanat Sharma; Shinji Watanabe; Anuj Kumar; Ahmed A Aly; Yue Liu; Florian Metze; Zhaojiang Lin; |
| 116 | WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents WorldPlay, a streaming video diffusion model that enables real-time, interactive world modeling with long-term geometric consistency, resolving the trade-off between speed and memory that limits current methods. |
Wenqiang Sun; Haiyu Zhang; Haoyuan Wang; Junta Wu; Zehan Wang; Zhenwei Wang; Yunhong Wang; Jun Zhang; Tengfei Wang; Chunchao Guo; |
| 117 | CoCoQuant: Breaking The Bandwidth Wall Via Co-Optimized Communication and Computation Quantization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing approaches typically treat communication and computation in isolation, failing to exploit their coupled nature and introducing limited system-level acceleration and accuracy degradation. To address this, we propose CoCoQuant, a co-designed framework that jointly optimizes communication and computation as a unified end-to-end design space. |
Haojie Duanmu; Jifeng Ding; Size Zheng; Xuegui Zheng; Jiangfei Duan; Xingcheng ZHANG; Li-Wen Chang; Xin Liu; Dahua Lin; |
| 118 | Constitutional Black-Box Monitoring for Scheming in LLM Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce two pipelines for generating synthetic agent trajectories, *STRIDE* (iterative refinement) and *Gloom* (agent-environment simulation), from which we generate 1,000 samples each. |
Simon Storf; Rich Barton-Cooper; James Peters-Gill; Marius Hobbhahn; |
| 119 | Any-Order GPT As Masked Diffusion Model: Decoupling Formulation and Architecture Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We show decoder-only MDMs, despite a larger modeling space, can achieve significant inference speedups ($\sim25\times$) and comparable perplexity with techniques like temperature annealing, offering a path to reduced inference compute. |
Shuchen Xue; Tianyu Xie; Tianyang Hu; Zijin Feng; Jiacheng Sun; Kenji Kawaguchi; Zhenguo Li; Zhi-Ming Ma; |
| 120 | Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we establish a novel theoretical analysis: DDPO is an implicit form of score/flow matching with noisy targets, which increases variance and slows convergence. |
Shuchen Xue; Chongjian GE; Shilong Zhang; Yichen Li; Zhi-Ming Ma; |
| 121 | Simultaneous Speech-to-Speech Translation Without Aligned Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We instead propose Hibiki-Zero, a model for simultaneous speech translation trained without word-level alignments between source and target speech. |
Tom Labiausse; Romain Fabre; Yannick Estève; Alexandre Défossez; Neil Zeghidour; |
| 122 | SafeLab: An Interactive High-Fidelity Benchmark for Embodied Safety in Scientific Robotics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we introduce SafeLab, a generative simulation benchmark designed for the full lifecycle of safe robot learning. |
Fengshuo Bai; Yufeng Li; Ruihai Wu; Peishuo Wang; Yuhan Wang; Bernie Zhu; Yuanfei Wang; Tawei Chou; Gao; Runchuan Zhu; Ying Wen; Yaodong Yang; Yuanpei Chen; |
| 123 | Plan for Speed: Dilated Scheduling for Masked Diffusion Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Masked diffusion language models (MDLMs) promise fast, non-autoregressive text generation, yet existing samplers, which pick tokens to unmask based on model confidence, ignore interactions when unmasking multiple positions in parallel and effectively reduce to slow, autoregressive behavior. We propose the Dilated Unmasking Scheduler (DUS), an inference-only, planner-model-free method that partitions sequence positions into non-adjacent dilated groups and unmasked them in parallel so as to minimize an upper bound on joint entropy gain at each denoising step. |
Omer Luxembourg; Haim Permuter; Eliya Nachmani; |
| 124 | Failure-Driven Workflow Refinement Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose \textbf{CE-Graph}, which maintains a counterexample pool, estimates dense failure modes, and applies operator-constrained graph edits via a \textbf{Propose-and-Verify} loop with a convergence-aware stopping rule. |
Jusheng Zhang; Jing Yang; Kaitong Cai; Ziliang Chen; Yongsen Zheng; Kwok Yan Lam; Liang Lin; Keze Wang; |
| 125 | SOLAR for Offline MARL: Plateau-Triggered Potential Shaping Under World-Model Uncertainty Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We find that shaping becomes reliable when it is (i) activated only after \emph{statistically validated} learning plateaus and (ii) constrained to \emph{potential-based} shaping, which preserves the task optimum. Motivated by this, we propose \textsc{SOLAR}, a simulate–evaluate–shape framework. |
Jusheng Zhang; Yijia Fan; Ruiqi Chen; Jing Yang; Ziliang Chen; Yongsen Zheng; Yanxi Chen; Jian Wang; Kwok Yan Lam; Liang Lin; Keze Wang; |
| 126 | AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present AutoTool, a training framework that equips LLM agents with dynamic tool-selection capabilities throughout their reasoning trajectories. |
Jiaru Zou; Ling Yang; Yunzhe Qi; Sirui Chen; Mengting Ai; Ke Shen; Jingrui He; Mengdi Wang; |
| 127 | Think Fast and Slow: Step-Level Cognitive Depth Adaptation for LLM Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce CogRouter, a framework that trains agents to dynamically adapt cognitive depth at each step. |
Ruihan Yang; Fanghua Ye; Xiang Wei; Ruoqing Zhao; Kang Luo; Xinbo Xu; Bo Zhao; Ruotian Ma; Shanyi Wang; Zhaopeng Tu; Xiaolong Li; Deqing Yang; Liefeng Bo; |
| 128 | Antidistillation Fingerprinting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce ***antidistillation fingerprinting*** (ADFP), a principled approach that aligns the fingerprinting objective with the student’s learning dynamics. |
Yixuan Xu; John Kirchenbauer; Yash Savani; Asher Trockman; Alexander Robey; Tom Goldstein; Fei Fang; Zico Kolter; |
| 129 | Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: But as math leaderboards improve week by week, it is worth asking: do these gains reflect broader problem-solving ability or just narrow overfitting? To answer this question, we evaluate over 20 open-weight reasoning-tuned models across a broad suite of tasks, including math, scientific QA, agent planning, coding, and standard instruction-following. |
Maggie Huan; Yuetai Li; Tuney Zheng; Xiaoyu Xu; Seungone Kim; Minxin Du; Radha Poovendran; Graham Neubig; Xiang Yue; |
| 130 | Diversity-Preserved Distribution Matching Distillation for Fast Visual Synthesis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a role-separated distillation framework that explicitly disentangles the roles of distilled steps: the first step is dedicated to preserving sample diversity via a target-prediction (e.g., v-prediction) objective, while subsequent steps focus on quality refinement under the standard DMD loss, with gradients from the DMD objective blocked at the first step. |
Tianhe Wu; Ruibin Li; Lei Zhang; Kede Ma; |
| 131 | Rethinking The Trust Region in LLM Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This creates a sub-optimal learning dynamic: updates to low-probability tokens are aggressively over-penalized, while potentially catastrophic shifts in high-probability tokens are under-constrained, leading to training inefficiency and instability. To address this, we propose Divergence Proximal Policy Optimization (DPPO), which substitutes heuristic clipping with a more principled constraint based on a direct estimate of policy divergence (e.g., Total Variation or KL). |
Penghui Qi; Xiangxin Zhou; Zichen Liu; Tianyu Pang; Chao Du; Min Lin; Wee Sun Lee; |
| 132 | Retrieval-Aware Distillation for Transformer-SSM Hybrids Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This gap has been linked to a small set of attention heads, called Gather-and-Aggregate (G&A), which SSMs struggle to implement and are believed to drive the disparity. Leveraging this insight, we propose retrieval-aware distillation, a strategy that converts a pretrained Transformer into a hybrid student by preserving only these retrieval-critical components. |
Aviv Bick; Eric Xing; Albert Gu; |
| 133 | Elastic Diffusion Transformer Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Previous acceleration methods, such as pruning and distillation, typically rely on a fixed computational capacity, leading to insufficient acceleration and degraded generation quality. To address this limitation, we propose \textbf{Elastic Diffusion Transformer (E-DiT)}, an adaptive acceleration framework for DiT that effectively improves efficiency while maintaining generation quality. |
Jiangshan Wang; Zeqiang Lai; Jiarui Chen; Jiayi Guo; Hang Guo; Xiu Li; Xiangyu Yue; Chunchao Guo; |
| 134 | UnMaskFork: Test-Time Scaling for Masked Diffusion Via Deterministic Action Branching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we demonstrate that Masked Diffusion Language Models (MDLMs) are inherently amenable to advanced search strategies, owing to their iterative and non-autoregressive generation process. |
Kou Misaki; Takuya Akiba; |
| 135 | Proximal Decoding: Provably Reducing Copyright Risk for Any Language Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Proximal Decoding, a plug-and-play inference-time method for suppressing verbatim reproduction: it enables decoding from any risky LM trained on mixed-license data by keeping generation in bounded proximity to a permissively trained safe LM. |
Jacqueline He; Jonathan Hayase; Scott Yih; Sewoong Oh; Luke Zettlemoyer; Pang Wei Koh; |
| 136 | SSDCN: Spatial-Spectral Dual-Clustering-based Network for Hyperspectral Image Super-resolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing Transformers suffer from quadratic complexity, while window-based methods compromise global capture. To address this, we propose the Spatial-Spectral Dual-Clustering-based Network (SSDCN). |
Yong Yang; Xuran Zhang; Shuying Huang; Xiaozheng Wang; Weiguo Wan; Hangyuan Lu; |
| 137 | FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While models like CLIP perform well on global alignment, they often struggle to capture fine-grained details in object attributes, spatial relations, and linguistic expressions, with limited support for bilingual comprehension. To address these challenges, we introduce FG-CLIP 2, a bilingual vision-language model designed to advance fine-grained alignment for both English and Chinese. |
Chunyu Xie; Bin Wang; Fanjing Kong; Jincheng Li; Dawei Liang; Ji Ao; Dawei Leng; Yuhui Yin; |
| 138 | TriAttention: Efficient Long Reasoning with Trigonometric KV Compression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We show that this concentration causes queries to preferentially attend to keys at specific distances (e.g., nearest keys), with the centers determining which distances are preferred via a trigonometric series. Based on this, we propose TriAttention to estimate key importance by leveraging these centers. |
Weian Mao; Xi Lin; Wei Huang; Yuxin Xie; Tianfu Fu; Bohan Zhuang; Song Han; Yukang Chen; |
| 139 | Context Forcing: Consistent Autoregressive Video Generation with Long Context Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This structural discrepancy creates a critical **student-teacher mismatch**: the teacher’s inability to access long-term history prevents it from guiding the student on global temporal dependencies, effectively capping the student’s context length. To resolve this, we propose **Context Forcing**, a novel framework that trains a long-context student via a long-context teacher. |
Shuo Chen; Cong Wei; Sun Sun; Tiancheng SHEN; Ping Nie; Kai Zou; Ge Zhang; Ming-Hsuan Yang; Wenhu Chen; |
| 140 | DynVLA: Learning World Dynamics for Action Reasoning in Autonomous Driving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose DynVLA, a driving VLA model that introduces a new CoT paradigm termed Dynamics CoT. |
Shuyao Shang; Bing Zhan; Yunfei Yan; Yuqi Wang; Yingyan Li; Yasong An; Xiaoman Wang; Jierui Liu; Lu Hou; Lue Fan; Zhaoxiang Zhang; Tieniu Tan; |
| 141 | Online Rubrics Elicitation from Pairwise Comparisons Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Online Rubrics Elicitation (OnlineRubrics), a method that dynamically curates evaluation criteria in an online manner through pairwise comparisons of responses from current and reference policies. |
MohammadHossein Rezaei; Robert Vacareanu; Zihao Wang; Clinton Wang; Bing Liu; Yunzhong He; Afra Feyza Akyürek; |
| 142 | Monitorability As A Free Gift: How RLVR Spontaneously Aligns Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent work has reported that monitorability—the degree to which CoT faithfully and informatively reflects internal computation—can appear as a free gift during the early stages of Reinforcement Learning with Verifiable Rewards (RLVR). We make this observation concrete through a systematic evaluation across model families and training domains. |
Zidi Xiong; Shan Chen; Himabindu Lakkaraju; |
| 143 | SiameseNorm: Breaking The Barrier to Reconciling Pre/Post-Norm Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We attribute this phenomenon to a structural incompatibility within a *single-stream* design: Any application of the Post-Norm operation inevitably obstructs the clean identity gradient preserved by Pre-Norm. To fundamentally reconcile these paradigms, we propose SiameseNorm, a *two-stream* architecture that couples Pre-Norm-like and Post-Norm-like streams with shared parameters. |
Tianyu Li; Dongchen Han; Zixuan Cao; Haofeng Huang; Mengyu Zhou; Ming Chen; erchao.zec; xiaoxi jiang; guanjunjiang; Gao Huang; |
| 144 | WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present WorldMirror, a unified feed-forward model for comprehensive 3D geometric prediction tasks. |
Yifan Liu; Zhiyuan Min; Zhenwei Wang; Junta Wu; Tengfei Wang; Yixuan Yuan; Yawei Luo; Chunchao Guo; |
| 145 | Equilibrium Reasoners: Learning Attractors Enables Scalable Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We hypothesize that generalizable reasoning emerges through learning task-conditioned attractors. |
Benhao Huang; Zhengyang Geng; Zico Kolter; |
| 146 | A New Framework for Cybersecurity Refusals in AI Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present the first framework for establishing refusal boundaries in offensive security contexts. |
Eliot Jones; Matt Fredrikson; Zico Kolter; |
| 147 | Seeing Is Solving: Unlocking Efficient Multimodal RL Via View Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce FOCUS-RL, a plug-and-play framework that can be seamlessly integrated into any VLM and dramatically boosts RLFT training efficiency. |
Qinsi Wang; Jing Shi; Kun Wan; Handong Zhao; Hancheng Ye; Zishan Shao; Jinghan Ke; Yudong Liu; Daniel Miranda; Purvak Lapsiya; Yiran Chen; Wentian Zhao; |
| 148 | Compressed Sensing for Capability Localization in Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Zeroing out as few as five task-specific heads can degrade performance by up to $65\\%$ on standard benchmarks measuring the capability of interest, while largely preserving performance on unrelated tasks. We introduce a compressed sensing based method that exploits the sparsity of these heads to identify them via strategic knockouts and a small number of model evaluations. |
Anna Bair; Yixuan Xu; Mingjie Sun; Zico Kolter; |
| 149 | Learning Syntax Without Semantics: Disentangled Tiny Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We show that syntax can be learned while suppressing semantic plausibility and world‑knowledge cues, yielding more efficient and controllable models. |
Ezra Winston; Zico Kolter; |
| 150 | Annotations Mitigate Post-Training Mode Collapse Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Crucially, we find this trade-off worsens with scale. To close this semantic diversity gap, we propose annotation-anchored training, a principled method that enables models to adopt the preference-following behaviors of post-training without sacrificing the inherent diversity of pre-training. |
Jacob Mitchell Springer; Madhu Advani; Lukas Aichberger; Arwen Bradley; Eran Malach; Omid Saremi; Sinead Williamson; Preetum Nakkiran; Etai Littwin; Aditi Raghunathan; |
| 151 | Towards Understanding Modality Interaction in Multimodal Language Models Via Partial Information Decomposition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Partial Information Decomposition (PID) as a unified, decision-level framework that separates \emph{unique}, \emph{redundant}, and \emph{synergistic} contributions of sensory and linguistic inputs, moving beyond representation alignment and outcome-based evaluation. |
Wanlong Fang; Tianle Zhang; Wen Tao; Alvin Chan; |
| 152 | Bridging Time and Frequency: A Joint Modeling Framework for Irregular Multivariate Time Series Forecasting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These irregularities violate the equidistant assumptions of standard models, hindering local temporal modeling and rendering classical frequency-domain methods ineffective for capturing global periodic structures. To address this challenge, we propose TFMixer, a joint time–frequency modeling framework for IMTS forecasting. |
Xiangfei Qiu; Kangjia Yan; Xvyuan Liu; Xingjian Wu; Jilin Hu; |
| 153 | DAG: A Dual Correlation Network for Time Series Forecasting with Exogenous Variables Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this study, to better leverage exogenous variables, especially future exogenous variables, we propose $\textbf{DAG}$, which $\textit{utilizes $\underline{D}$ual correl$\underline{A}$tion network along both the temporal and channel dimensions for time series forecasting with exo$\underline{G}$enous}$ variables. |
Xiangfei Qiu; Yuhan Zhu; Zhengyu Li; Xingjian Wu; Bin Yang; Jilin Hu; |
| 154 | SEER: Transformer-based Robust Time Series Forecasting Via Automated Patch Enhancement and Replacement Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In real-world time series, there are often low-quality issues during data collection, such as missing values, distribution shifts, anomalies and white noise, which may cause some patches to contain low-quality information, negatively impacting the prediction results. To address this issue, this study proposes a robust time series forecasting framework called $\textbf{SEER}$. |
Xiangfei Qiu; Xvyuan Liu; Tianen Shen; Xingjian Wu; Hanyin Cheng; Bin Yang; Jilin Hu; |
| 155 | RubricRobustness: A Simple Framework for Evaluating The Robustness of Rubrics-Based Benchmarks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While these frameworks utilize expert-curated criteria and LLM-as-a-judge to assess open-ended generation, the intrinsic robustness of these evaluation harnesses to fundamental validity assessments remains critically under-investigated. To bridge this gap, we introduce RubricRobustness, a systematic sensitivity analysis framework that subjects these benchmarks to three common sense perturbations: semantic negation, stochastic deletion and irrelevant addition. |
Manasi Sharma; |
| 156 | OServe: Accelerating LLM Serving Via Spatial-Temporal Workload Orchestration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present OServe, an LLM serving system with heterogeneous and flexible model deployment that addresses both spatial and temporal heterogeneity. |
Youhe Jiang; Fangcheng Fu; Taiyi Wang; Guoliang HE; Eiko Yoneki; |
| 157 | TopAdapter: Topology-Aware Prompt Tuning for Efficient Point Cloud Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing parameter-efficient fine-tuning (PEFT) methods predominantly focus on input token prompting, overlooking the intrinsic geometric information. To address this limitation, we propose TopAdapter, a novel PEFT framework that enhances geometric perception by injecting local topological information into pre-trained 3D vision models. |
Changshuo Wang; Shuting He; Xiang Fang; Weijun Li; Yixian Shen; Mingkun Xu; Zhongtian Sun; Prayag Tiwari; |
| 158 | Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Latent Reasoning VLA (LaRA-VLA), a unified VLA framework that internalizes multi-modal CoT reasoning into continuous latent representations for embodied action. |
Shuanghao Bai; Jing Lyu; Wanqi Zhou; Zhe Li; Dakai Wang; Lei Xing; Xiaoguang Zhao; Pengwei Wang; Zhongyuan Wang; Cheng Chi; Badong Chen; Shanghang Zhang; |
| 159 | Real-Time and Lightweight Diffusion Image Compression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we explore the design of real-time and lightweight diffusion codecs by addressing two pivotal questions. |
Zhaoyang Jia; Naifu Xue; Zihan Zheng; Jiahao Li; Bin Li; Xiaoyi Zhang; Zongyu Guo; Yuan Zhang; Houqiang Li; Yan Lu; |
| 160 | UltraHorizon: Benchmarking LLM-Agent Capabilities in Ultra Long-Horizon Scenarios Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing benchmarks rarely capture these long-horizon challenges, leaving a gap in systematic evaluation. To bridge this gap, we introduce $\textbf{UltraHorizon}$, a novel benchmark that measures the foundational capabilities essential for complex real-world challenges. |
Haotian Luo; Huaisong Zhang; Xuelin Zhang; Haoyu Wang; Zeyu Qin; Wenjie Lu; Guozheng Ma; Haiying He; Yingsha Xie; Qiyang Zhou; Zixuan Hu; Hongze Mi; Yibo Wang; Naiqiang Tan; Hong Chen; Yi Fung; Chun Yuan; Li Shen; |
| 161 | Why Are Linear RNNs More Parallelizable? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While prior work establishes the expressivity benefits of LRNNs over transformers, it is unclear what makes LRNNs—but not traditional, *nonlinear* RNNs—as easy to parallelize in practice as transformers. We answer this question by providing a tight connection between types of RNNs and standard complexity classes. |
William Merrill; Hongjian Jiang; Yanhong Li; Anthony Lin; Ashish Sabharwal; |
| 162 | Linearizing Vision Transformer with Test-Time Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inheriting weights from pretrained Transformers provides an appealing shortcut, yet the fundamental representational gap between Softmax and linear attention prevents effective weight transfer. In this work, we address this conversion challenge from two perspectives: architectural alignment and representational alignment. |
Yining Li; Dongchen Han; Zeyu Liu; Hanyi Wang; Yulin Wang; Gao Huang; |
| 163 | Beyond Gemini-3-Pro: Revisiting LLM Routing and Aggregation at Scale Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we explore collective intelligence as an alternative to monolithic scaling, and demonstrate that open-source LLMs’ collaboration can surpass Gemini-3-Pro. |
Shengji Tang; Weihao Lin; Peng Ye; Jingqi Ye; Hao Li; Yiqun Zhang; Xiaosong Wang; Bo Zhang; Shuyue Hu; Tao Chen; LEI BAI; Wanli Ouyang; |
| 164 | Fast KV Compaction Via Attention Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work describes an approach for *fast* context compaction in latent space through **Attention Matching**, which constructs compact keys and values to reproduce attention outputs and preserve attention mass at a per-KV-head level. |
Adam Zweiger; Xinghong Fu; Han Guo; Yoon Kim; |
| 165 | Multimodal Latent Language Modeling with Next-Token Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: In this work, we propose Latent Language Modeling (LatentLM), which seamlessly integrates continuous and discrete data using causal Transformers. |
Yutao Sun; Hangbo Bao; Wenhui Wang; Zhiliang Peng; Li Dong; Shaohan Huang; Yaoyao Chang; Jianyong Wang; Furu Wei; |
| 166 | Retaining By Doing: The Role of On-Policy Data in Mitigating Forgetting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Adapting language models (LMs) to new tasks via post-training carries the risk of degrading existing capabilities — a phenomenon classically known as catastrophic forgetting. In this paper, toward identifying guidelines for mitigating this phenomenon, we systematically compare the forgetting patterns of two widely adopted post-training methods: supervised fine-tuning (SFT) and reinforcement learning (RL). |
Howard Chen; Noam Razin; Karthik Narasimhan; Danqi Chen; |
| 167 | Scaling Prompt Synthesis for Large Language Model Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: PromptCoT showed that injecting rationales into prompt synthesis increases problem difficulty. Building on this, we present PromptScale, a scalable framework that replaces hand-crafted heuristics with an expectation-maximization (EM) loop, where rationales are iteratively refined to guide prompt construction. |
Xueliang Zhao; Wei Wu; Jian Guan; Zhuocheng Gong; Lingpeng Kong; |
| 168 | Demystifying Action Space Design for Robotic Manipulation Policies Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While recent advances have focused heavily on scaling training data and model capacity, the choice of action space remains guided by ad-hoc heuristics or legacy designs, leading to an ambiguous understanding of robotic policy design philosophies. To address this ambiguity, we conducted a large-scale and systematic empirical study, confirming that the action space does have significant and complex impacts on robotic policy learning. |
Yuchun Feng; Jinliang Zheng; Zhihao Wang; Dongxiu Liu; Jianxiong Li; Jiangmiao Pang; Tai Wang; Xianyuan Zhan; |
| 169 | Evaluating LLM Uncertainty in Long-Form Generation Using Deterministic Ground Truth Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Single-answer Atomic Long-form Target (SALT), a benchmark of six procedurally generated tasks with single deterministic long textual ground truths, enabling unit-level evaluation of correctness, calibration, and ranking without external judges. |
Ido Amit; Ido Galil; Ran El-Yaniv; |
| 170 | Weight-sparse Transformers Have Interpretable Circuits Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We train models to have more understandable circuits by constraining most of their weights to be zeros, so that each neuron only has a few connections. |
Leo Gao; Achyuta Rajaram; Jacob Coxon; Soham Govande; Bowen Baker; Daniel Mossing; |
| 171 | Position: Don’t Just Fix It in Post”: A Science of AI Must Study Learning Dynamics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Language models are not static objects—they are snapshots of time-evolving processes shaped by data, objectives, and optimization dynamics. Yet the field predominantly treats models as fixed artifacts, analyzing behaviors after training rather than asking *why* they emerge. |
Stella Biderman; Mohammad Aflah Khan; Niloofar Mireshghallah; Catherine Arnett; Fazl Barez; Naomi Saphra; |
| 172 | Shuffle The Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose *RoPE-Perturbed Self-Distillation*, a training regularizer that improves positional robustness. |
Zichong Li; Chen Liang; Liliang Ren; Tuo Zhao; Yelong Shen; Weizhu Chen; |
| 173 | Don’t Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this study, we show that layer dropout **should** be used in state-of-the-art LLM training, establishing best practices and scaling analysis for both training and post-training benefits. |
Mostafa Elhoushi; Nolan Dey; Alexander Pretko; Bin Zhang; Gavia Gray; Gurpreet Gosal; Abdulrahman Mahmoud; Shane Bergsma; Joel Hestness; |
| 174 | Least-Loaded Expert Parallelism: Load Balancing An Imbalanced Mixture-of-Experts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Least-Loaded Expert Parallelism (LLEP), a novel EP algorithm that dynamically reroutes excess tokens and associated expert parameters from overloaded devices to underutilized ones. |
Xuan-Phi Nguyen; Shrey Pandit; Austin Xu; Caiming Xiong; Shafiq Joty; |
| 175 | VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, vlog assessment is highly subjective and remains challenging due to a lack of standardized criteria, dataset and benchmark, and effective reward models. To address these challenges, we define a comprehensive vlog evaluation framework guided by professional vlog creators and product managers, establishing a taxonomy of six key dimensions, *i.e.*, *Creativity*, *Consistency*, *Concept Design*, *Cinematography*, *Narration*, and *Pacing*. |
Yexiang Liu; Wen Zhong; Sijie Zhu; Xin Gu; Fan Chen; Junxian Duan; Jie Cao; Longyin Wen; Zhenfang Chen; |
| 176 | Unveiling Multi-regime Patterns in SciML: Distinct Failure Modes and Regime-specific Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Neural networks (NNs) trained under different hyperparameters can fall into distinct training “regimes”, with models in the same regime showing homogeneous properties and models across regimes differing qualitatively. In this paper, we analyze multi-regime patterns in scientific machine learning (SciML) models by characterizing these regimes and the transitions between them. |
Yuanzhe Hu; Xiaopeng Wang; Yuxin Wang; Xiaokun Zhong; Haiquan Lu; Tianyu Pang; Michael Mahoney; Yujun Yan; Pu Ren; Yaoqing Yang; |
| 177 | NorMuon: Making Muon More Efficient and Scalable Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite Muon’s emergence as a candidate successor to Adam, the potential for jointly leveraging their strengths—has not been systematically explored. In this work, we bridge this gap by proposing NorMuon (Neuron-wise Normalized Muon), an optimizer that synergistically combines orthogonalization with neuron-level adaptive learning rates. |
Zichong Li; Liming Liu; Chen Liang; Weizhu Chen; Tuo Zhao; |
| 178 | When to Memorize and When to Stop: Gated Recurrent Memory for Long-Context Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, this naive recurrent memory update faces two crucial drawbacks: (i) memory can quickly explode because it can update indiscriminately, even on evidence-free chunks; and (ii) the loop lacks an exit mechanism, leading to unnecessary computation after even sufficient evidence is collected. To address these issues, we propose GRU-Mem, which incorporates two text-controlled gates for more stable and efficient long-context reasoning. |
Leheng Sheng; Yongtao Zhang; Wenchang Ma; Yaorui Shi; Ting Huang; Xiang Wang; An Zhang; Ke Shen; Tat-Seng Chua; |
| 179 | Privileged Information Distillation for Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce π-Distill, a joint teacher–student framework that trains a PI-conditioned teacher and an unconditioned student simultaneously within a single shared-parameter model, enabling the teacher to learn how to use PI while mitigating distribution shift during transfer. |
Emiliano Penaloza; Dheeraj Vattikonda; Nicolas Gontier; Alexandre Lacoste; Laurent Charlin; Massimo Caccia; |
| 180 | Stop Training for The Worst: Progressive Unmasking Accelerates Masked Diffusion Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose Progressive UnMAsking (PUMA), a simple modification of the forward masking process that aligns training-time and inference-time masking patterns, thereby focusing optimization on *inference-aligned masks* and speeding up training. |
Jaeyeon Kim; Jonathan Geuter; David Alvarez-Melis; Sham Kakade; Sitan Chen; |
| 181 | Fine-Tuning Masked Diffusion for Provable Self-Correction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Prior attempts to incorporate self-correction into MDMs either require overhauling MDM architectures/training or rely on imprecise proxies for token quality, limiting their applicability. Motivated by this, we introduce PRISM–Plug-in Remasking for Inference-time Self-correction of Masked Diffusions–a lightweight, model-agnostic approach that applies to any pretrained MDM. |
Jaeyeon Kim; Seunggeun Kim; Taekyun Lee; David Pan; Hyeji Kim; Sham Kakade; Sitan Chen; |
| 182 | PostTrainBench: Can LLM Agents Automate LLM Post-Training? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we study *post-training*, which is the critical step that turns base LLMs into useful assistants. |
Ben Rank; Hardik Bhatnagar; Ameya Pandurang Prabhu; Shira Eisenberg; Karina Nguyen; Matthias Bethge; Maksym Andriushchenko; |
| 183 | Stop The Flip-Flop: Context-Preserving Verification for Fast Revocable Diffusion Decoding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose COVER (Cache Override Verification for Efficient Revision), which performs leave-one-out verification and stable drafting within a single forward pass. |
Yanzheng Xiang; Lan Wei; Yizhen Yao; Qinglin Zhu; Hanqi Yan; Chen Jin; Philip Teare; Dandan Zhang; Lin Gui; Amrutha Saseendran; Yulan He; |
| 184 | ReaForest: Fostering Generative Video Reasoning for Spatial Planning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We attribute this limitation to two fundamental gaps: (i) VGMs are predominantly trained on general-purpose video corpora emphasizing perceptual fidelity over visual reasoning, leaving reasoning abilities underdeveloped; (ii) most VGMs generate videos in a single pass without mechanisms to explore alternative reasoning trajectories and to revise intermediate errors. Motivated by these limitations, we introduce **ReaForest**, a framework that fosters the reasoning capacity of VGMs in spatial planning through both training-time activation and inference-time scaling. |
Kun Ouyang; Yuanxin Liu; Xinhao Li; Linli Yao; Xiangyu Zeng; Haoning Wu; Hao Zhou; Fandong Meng; Jie Zhou; Xu SUN; |
| 185 | CodeClash: Benchmarking Goal-Oriented Software Engineering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce CodeClash, a benchmark where LMs compete in multi-round tournaments to build the best codebase for achieving a competitive objective. |
John Yang; Kilian Lieret; Joyce Yang; Carlos Jimenez; Muhtasham Oblokulov; Aryan Siddiqui; Ofir Press; Ludwig Schmidt; Diyi Yang; |
| 186 | On Path to Multimodal Historical Reasoning: HistBench and HistAgent Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: Existing general-purpose agents perform well on many current benchmarks but lack the domain expertise needed to address complex historical questions. To address this gap, we introduce HistBench, a new benchmark of 414 high-quality and carefully-reviewed questions stratified by difficulty and designed to evaluate LLM’s capacity for historical reasoning. |
Jiahao Qiu; Fulian Xiao; Yimin Wang; Yuchen Mao; Yijia Chen; Xinzhe Juan; Siran Wang; Xuan Qi; Tongcheng Zhang; Zixin Yao; Jiacheng Guo; Yifu Lu; Charles Argon; Jundi Cui; Daixin Chen; Junran Zhou; Shuyao Zhou; Zhanpeng Zhou; Ling Yang; Shilong Liu; Hongru WANG; Kaixuan Huang; xun jiang; Xi Gao; Mengdi Wang; |
| 187 | Position: Assistive AI Requires Personalized Specialists, Not Generalists Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We outline research directions for building specialists that learn from organic observational data, avoid self-reinforcing errors, and improve safely over long horizons. |
Homanga Bharadhwaj; |
| 188 | ATLAS: Learning to Optimally Memorize The Context at Test Time Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We observe that these shortcomings come from three disjoint aspects in their design: (1) limited memory capacity that is bounded by the architecture of memory and feature mapping of the input; (2) online nature of update, i.e., optimizing the memory only with respect to the last input; and (3) less expressive management of their fixed-size memory. To enhance all these three aspects, we present Atlas, a long-term memory module with high capacity that learns to memorize the context by optimizing the memory based on the current and past tokens, overcoming the online nature of long-term memory models. |
Ali Behrouz; Zeman Li; Praneeth Kacham; Majid Daliri; Yuan Deng; Peilin Zhong; Meisam Razaviyayn; Vahab Mirrokni; |
| 189 | The Surprising Difficulty of Search in Model-Based Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Instead, we show that mitigating distribution shift matters more than improving model or value function accuracy. Building on this insight, we identify key techniques for enabling effective search, achieving state-of-the-art performance across multiple popular benchmark domains. |
Wei-Di Chang; Mikael Henaff; Brandon Amos; Gregory Dudek; Scott Fujimoto; |
| 190 | Memory Caching: RNNs with Growing Memory Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce Memory Caching (MC), a simple yet effective technique that enhances recurrent models by caching checkpoints of their memory states (a.k.a. hidden states). |
Ali Behrouz; Zeman Li; Yuan Deng; Peilin Zhong; Meisam Razaviyayn; Vahab Mirrokni; |
| 191 | Problem Distributions As Tasks: Repurposing Meta Learning for Generative Combinatorial Optimization Towards Multi-task Pretrain and Adaptation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper proposes M$^2$GenCO, a Multi-task learning framework that pioneers the instantiation of the Meta-learning mechanism with diffusion-based Generative solving for CO Problems (COPs) on graphs, first formulating tasks in meta-learning as distinct problem types instead of instances of the same problem. |
Wenzheng Pan; Jiale Ma; Nuoyan Chen; Yang Li; Junchi Yan; |
| 192 | Physiology As Language: Translating Nocturnal Breathing to EEG Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address the significant complexity gap between the two modalities, we propose a waveform-conditional generative framework that preserves fine-grained respiratory dynamics while constraining the EEG target space through discrete tokenization. |
Kaiwen Zha; Chao Li; Hao He; Peng Cao; Tianhong Li; Ali Mirzazadeh; Ellen Zhang; Jong Lee; Yoon Kim; Dina Katabi; |
| 193 | OmniDenseCap: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper proposes Omni Dense Captioning, a novel task designed to generate continuous, fine-grained, and structured audio-visual narratives with explicit timestamps. |
Linli Yao; Yuancheng Wei; Yaojie Zhang; Lei Li; Xinlong Chen; Feifan Song; Ziyue Wang; Kun Ouyang; Yuanxin Liu; Lingpeng Kong; Qi Liu; Pengfei Wan; Kun Gai; Yuanxing Zhang; Xu SUN; |
| 194 | ObjEmbed: Towards Universal Multimodal Object Embeddings Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present ObjEmbed, a novel MLLM embedding model that decomposes the input image into multiple regional embeddings, each corresponding to an individual object, along with global embeddings. |
Shenghao Fu; Yukun Su; Fengyun Rao; Jing LYU; Xiaohua Xie; Wei-Shi Zheng; |
| 195 | Proxy Compression for Language Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work introduces proxy compression, an alternative training scheme that preserves the efficiency benefits of compressed inputs while providing an end-to-end, raw-byte interface at inference time. |
Lin Zheng; Li Xinyu; Qian Liu; Xiachong Feng; Lingpeng Kong; |
| 196 | MM-DeepResearch: A Simple and Effective Multimodal Agentic Search Baseline Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We aim to develop a multimodal research agent capable of explicit reasoning and planning, multi-tool invocation, and cross-modal information synthesis, enabling it to conduct deep research tasks. |
Huanjin Yao; Qixiang Yin; Min Yang; Ziwang Zhao; Yibo Wang; Haotian Luo; Jingyi Zhang; Jiaxing Huang; |
| 197 | Conversation for Non-verifiable Learning: Self-Evolving Large Language Models Through Meta-Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce CoNL, a framework that unifies generation, evaluation, and meta-evaluation through multi-agent self-play. |
Yuan Sui; Bryan Hooi; |
| 198 | Memory Is Reconstructed, Not Retrieved: Graph Memory for LLM Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While current memory-augmented agents rely on a static “retrieve-then-reason” paradigm, this rigid pipeline design prevents them from dynamically adapting memory access to intermediate evidence discovered during inference. To bridge this gap, we propose MRAgent, a framework that combines an associative memory graph with an active reconstruction mechanism. |
Shuo Ji; yibo li; Bryan Hooi; |
| 199 | Rethinking The Reranker: Boundary-Aware Evidence Selection for Robust Retrieval-Augmented Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose \texttt{BAR-RAG}, which reframes the reranker as a boundary-aware evidence selector that targets the generator’s Goldilocks Zone—evidence that is neither trivially easy nor fundamentally unanswerable for the generator, but is challenging yet sufficient for inference and thus provides the strongest learning signal. |
Jiashuo Sun; Pengcheng Jiang; Saizhuo Wang; Jiajun Fan; Heng Wang; Siru Ouyang; Ming Zhong; Yizhu Jiao; Chengsong Huang; Xueqiang Xu; Pengrui Han; Peiran Li; Jiaxin Huang; Ge Liu; Heng Ji; Jiawei Han; |
| 200 | Learnability-Informed Fine-Tuning of Diffusion Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We aim to improve the reasoning capabilities of diffusion language models (DLMs). |
Shubham Parashar; Atharv Chagi; Jacob Helwig; Lakshmi Madhavarapu; Sushil Vemuri; James Caverlee; Dileep Kalathil; Shuiwang Ji; |
| 201 | Dual Latent Memory for Visual Multi-agent System Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose L$^{2}$-VMAS, a novel model-agnostic framework that enables inter-agent collaboration with dual latent memories. |
Xinlei Yu; Chengming Xu; Zhangquan Chen; Bo Yin; Cheng Yang; Yongbo He; Yihao Hu; Jiangning Zhang; Cheng Tan; Xiaobin Hu; Shuicheng YAN; |
| 202 | Multi-Agent Teams Hold Experts Back Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Drawing on organizational psychology, we study whether self-organizing LLM teams achieve *strong synergy*, where team performance matches or exceeds the best individual member. |
Aneesh Pappu; Batu El; Hancheng Cao; Carmelo di Nolfo; Yanchao Sun; Meng Cao; James Zou; |
| 203 | Position: Adversarial ML for LLMs Is Not Making Any Progress Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Today, adversarial ML research has shifted towards studying larger, general-purpose language models. In this position paper, we argue that the situation is now even worse: in the era of LLMs, the field of adversarial ML studies problems that are (1) less clearly defined, (2) harder to solve, and (3) even more challenging to evaluate. |
Javier Rando; Jie Zhang; Nicholas Carlini; Florian Tramer; |
| 204 | SPA: A Simple But Tough-to-Beat Baseline for Knowledge Injection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose **SPA** (**S**caling **P**rompt-engineered **A**ugmentation), a simple but tough-to-beat baseline that uses a small set of carefully designed prompts to generate large-scale synthetic data for knowledge injection. |
Kexian Tang; Jiani Wang; Shaowen Wang; Kaifeng Lyu; |
| 205 | Diamond Maps: Efficient Reward Alignment Via Stochastic Flow Maps Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Diamond Maps, a stochastic flow-map model that enables efficient and accurate alignment to arbitrary rewards at inference time. |
Peter Holderrieth; Douglas Chen; Luca Eyring; Ishin Shah; Giri Anantharaman; Yutong He; Zeynep Akata; Tommi Jaakkola; Nicholas Boffi; Max Simchowitz; |
| 206 | Golden Goose: A Simple Trick to Synthesize Unlimited RLVR Tasks from Unverifiable Internet Text Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Yet, scaling up RL is bottlenecked by limited existing verifiable data, where improvements increasingly saturate over prolonged training. To overcome this, we propose **Golden Goose**, a simple trick to synthesize unlimited RLVR tasks from unverifiable internet text by constructing a multiple-choice question-answering version of the fill-in-the-middle task. |
Ximing Lu; David Acuna; Jaehun Jung; Jian Hu; Di Zhang; Shizhe Diao; Yunheng Zou; Shaokun Zhang; Brandon Cui; Mingjie Liu; Hyunwoo Kim; Prithviraj Ammanabrolu; Jan Kautz; Yi Dong; Yejin Choi; |
| 207 | MemEvolve: Meta-Evolution of Agent Memory Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, this paradigm is fundamentally constrained by the \textit{staticity} of the memory system itself: while memory facilitates agent-level evolving, the underlying memory architecture cannot be meta-adapted to diverse task contexts. To address this gap, we propose MemEvolve, a meta-evolutionary framework that jointly evolves agents’ experiential knowledge and their memory architecture, allowing agent systems not only to accumulate experience but also to progressively refine how they learn from it. |
Guibin Zhang; Haotian Ren; Chong Zhan; Junhao Wang; He Zhu; Wangchunshu Zhou; Shuicheng YAN; |
| 208 | Wait, Wait, Wait… Why Do Reasoning Models Loop? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This points to mismatches between the training distribution and the learned model, which we refer to as errors in learning, as a key cause. To understand how such errors cause loops, we introduce a synthetic graph reasoning task and demonstrate two mechanisms. |
Charilaos Pipis; Shivam Garg; Vasilis Kontonis; Vaishnavi Shrivastava; Akshay Krishnamurthy; Dimitris Papailiopoulos; |
| 209 | SlideSparse: Fast and Flexible (2N-2):2N Structured Sparsity Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present **SlideSparse**, the first system to unlock Sparse Tensor Core acceleration for the $(2N-2):2N$ model family on commodity GPUs. |
Yingbo HAO; Hanyong Shao; Ting Song; Yan Xia; Di Zhang; Shaohan Huang; Xun Wu; Songchen Xu; Le Xu; Li Dong; Zewen Chi; Yi Zou; Furu Wei; |
| 210 | Lions and Muons: Optimization Via Stochastic Frank-Wolfe Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: On the other hand, recent optimizers such as Lion and Muon have gained quite significant popularity in deep learning. In this work, building on recent initiatives, we provide a unifying perspective by interpreting these seemingly disparate methods through the lens of Stochastic Frank-Wolfe. |
Maria-Eleni Sfyraki; Jun-Kun Wang; |
| 211 | Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose using sparse autoencoders (SAEs) to create *SAE embeddings*: representations whose dimensions map to interpretable concepts. |
Nicholas Jiang; Xiaoqing Sun; Lisa Dunlap; Lewis Smith; Neel Nanda; |
| 212 | LaST$_{0}$: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To mitigate these limitations, we propose LaST$_0$, a framework that enables efficient reasoning before acting through a Latent Spatio-Temporal Chain-of-Thought (CoT), capturing fine-grained physical and robotic dynamics that are often difficult to verbalize. Specifically, we introduce a token-efficient latent CoT space that models future visual dynamics, 3D structural information, and robot proprioceptive states, and further extends these representations across time to enable temporally consistent implicit reasoning trajectories. |
Zhuoyang Liu; Jiaming Liu; Hao Chen; Jiale Yu; Ziyu Guo; Chengkai Hou; Xiangju Mi; Chenyang Gu; Renrui Zhang; Kun Wu; Zhengping Che; Jian Tang; Pheng Ann Heng; Shanghang Zhang; |
| 213 | FullStack-Agent: Enhancing Agentic Full-Stack Web Coding Via Development-Oriented Testing and Repository Back-Translation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Notably, constructing production-level full-stack web applications is far more challenging than only generating frontend web pages, demanding careful control of data flow, comprehensive understanding of constantly updating packages and dependencies, and accurate localization of obscure bugs in the codebase. To address these difficulties, we introduce FullStack-Agent, a unified agent system for full-stack agentic coding that consists of three parts: (1) FullStack-Dev, a multi-agent framework with strong planning, code editing, codebase navigation, and bug localization abilities. |
Zimu Lu; Houxing Ren; Yunqiao Yang; Ke Wang; Zhuofan Zong; Mingjie Zhan; Hongsheng Li; |
| 214 | SoMA: A Real-to-Sim Neural Simulator for Robotic Soft-Body Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents \textbf{SoMA}, a 3D Gaussian Splat simulator for soft-body manipulation. |
Mu Huang; Hui Wang; Kerui Ren; Linning Xu; Mulin Yu; Yunsong Zhou; Bo Dai; Jiangmiao Pang; |
| 215 | SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, these models often struggle with novel and specialized software, particularly in scenarios lacking human annotations. To address this challenge, we propose SEAgent, an agentic self-evolving framework enabling CUAs to autonomously evolve through interactions with unfamiliar software. |
ZEYI SUN; Ziyu Liu; Yuhang Zang; Yuhang Cao; Xiaoyi Dong; Tong Wu; Dahua Lin; Jiaqi Wang; |
| 216 | MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work proposes a novel embodied 4D world model that enables geometrically consistent, arbitrary-view RGBD generation: given only a single-view RGBD observation as input, the model “imagines” the remaining viewpoints, which can then be back-projected and fused to assemble a more complete 3D structure across time. |
Jiaxu Wang; JIANG Yicheng; Tianlun HE; Jingkai SUN; Qiang Zhang; Jiahang Cao; Zesen Gan; Mingyuan Sun; Qiming Shao; Xiangyu Yue; |
| 217 | Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose \textbf{DiNa-LRM}, a \textbf{di}ffusion-\textbf{na}tive \textbf{l}atent \textbf{r}eward \textbf{m}odel that formulates preference learning directly on noisy diffusion states. |
Gongye Liu; Bo Yang; Zhi Yida; Zhizhou Zhong; Lei Ke; Didan Deng; Han Gao; Yongxiang Huang; Kaihao Zhang; Hongbo Fu; Wenhan Luo; |
| 218 | Doc-to-LoRA: Learning to Instantly Internalize Contexts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While context distillation (CD) can transfer information into model parameters, per-prompt distillation is impractical due to training costs and latency. To address these limitations, we propose Doc-to-LoRA (D2L), a lightweight hypernetwork that meta-learns to perform approximate CD within a single forward pass. |
Rujikorn Charakorn; Edoardo Cetin; Shinnosuke Uesaka; Robert Lange; |
| 219 | GemDepth: Geometry-Embedded Features for 3D-Consistent Video Depth Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We argue that current approaches, which primarily rely on temporal smoothing via Transformers, struggle to maintain strict 3D geometric consistency—particularly under rotations or drastic view changes. To address this, we propose GemDepth, a framework built on the insight that an explicit awareness of camera motion and global 3D structure is a prerequisite for 3D consistency. |
Yuecheng Liu; Junda Cheng; Longliang Liu; Wenjing Liao; Hanrui Cheng; Yuzhou Wang; Xin Yang; |
| 220 | Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce MOSAIC, a post-training framework that aligns agents for safe multi-step tool use by making safety decisions explicit and learnable. |
Aradhye Agarwal; Gurdit Siyan; Yash Pandya; Joykirat Singh; Akshay Nambi; Ahmed Awadallah; |
| 221 | Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Just-In-Time Reinforcement Learning (JitRL), a training-free framework that enables test-time policy optimization without any gradient updates. |
yibo li; Zijie Lin; Ailin Deng; Xuan Zhang; Yufei He; Shuo Ji; Tri Cao; Bryan Hooi; |
| 222 | Flash-GRPO: Efficient Alignment for Video Diffusion Via One-Step Policy Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Flash-GRPO, a single-step training framework that outperforms full trajectory training in alignment quality under low computational budgets while substantially improving training efficiency. |
Xiaoxuan He; Siming Fu; Zeyue Xue; Weijie Wang; Ruizhe He; Yuming Li; Dacheng Yin; Shuai Dong; Haoyang Huang; Hongfa Wang; Nan Duan; Bohan Zhuang; |
| 223 | SparseInfer: Accelerating Large Language Model Inference with Semantics-Inspired Adaptive Sparse Activation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce \textbf{SparseInfer}, an MLP-free adaptive sparse activation inference method based on sentence-level prediction. |
Qinsi Wang; Saeed Vahidian; Hancheng Ye; Jianyang Gu; Jianyi Zhang; Yiran Chen; |
| 224 | OvisOCR: End-to-End Document Parsing Via Aligning Specialized Perception with General Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents OvisOCR, a lightweight and strictly end-to-end Multimodal Language Model (MLLM) tailored for document parsing. |
Jun-Peng Jiang; Shiyin Lu; An-Yang Ji; Yinglun Li; Qing-Guo Chen; Zhao Xu; Weihua Luo; Kaifu Zhang; De-Chuan Zhan; Han-Jia Ye; |
| 225 | Dissecting Post-Training: Uncovering The Complementary Roles of SFT and RL for Document Parsing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We further ground this phenomenon in the distinct theoretical nature of their respective objective functions. Based on these findings, we introduce a unified strategy that explicitly harnesses their individual strengths while mitigating their weaknesses. |
Jun-Peng Jiang; An-Yang Ji; Shiyin Lu; Guodong Zheng; Weihong Zhang; Qing-Guo Chen; Weihua Luo; Kaifu Zhang; Long Chen; De-Chuan Zhan; Han-Jia Ye; |
| 226 | Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model Unlearning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite rapid progress, KD-based unlearning struggles with biased deletion due to suppressing specific token sequences as a substitute for complete knowledge removal, whereas DR-based unlearning risks the re-emergence of harmful knowledge because the underlying knowledge remains intact. To address these issues, we propose Distinguishable Deletion ($\mathrm{D^2}$), a paradigm that restricts the response distribution in the latent space rather than specific tokens to erase undesirable knowledge, while distinguishing it from retained knowledge, enabling a refusal mechanism to handle unlearned inputs safely and coherently. |
Puning Yang; Junchi Yu; Qizhou Wang; Phil Torr; Bo Han; Xiuying Chen; |
| 227 | Graph-R1: Towards Agentic GraphRAG Framework Via End-to-end Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: GraphRAG methods improve RAG by modeling knowledge as entity-relation graphs, but still face challenges in high construction cost, fixed one-time retrieval, and reliance on long-context reasoning and prompt design. To address these challenges, we propose Graph-R1, an agentic GraphRAG framework via end-to-end reinforcement learning (RL). |
Haoran Luo; Haihong E; Guanting Chen; Qika Lin; Yikai Guo; Fangzhi Xu; Zemin Kuang; Meina Song; Xiaobao Wu; Yifan Zhu; Anh Tuan Luu; |
| 228 | CoF-T2I: Video Models As Pure Visual Reasoners for Text-to-Image Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, their potential to enhance text-to-image (T2I) generation remains largely unexplored due to the absence of a clearly defined visual reasoning starting point and interpretable intermediate states in the T2I generation process. To bridge this gap, we propose **CoF-T2I**, a model that integrates CoF reasoning into T2I generation via progressive visual refinement, where intermediate frames act as explicit reasoning steps and the final frame is taken as output. |
Chengzhuo Tong; Chang Mingkun; Shenglong Zhang; Yuran Wang; Cheng Liang; Zhizheng Zhao; Bohan Zeng; Yang Shi; Ruichuan An; Yifan Dai; Ziming Zhao; Guanbin Li; Pengfei Wan; Yuanxing Zhang; Wentao Zhang; |
| 229 | The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models Via Latent Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although this internal cognitive processing may not always manifest as explicit linguistic structures, it is instrumental in formulating high-quality responses. Inspired by this cognitive phenomenon, we propose a novel **F**ull-duplex **LA**tent and **I**nternal **R**easoning method named FLAIR that conducts *latent* thinking simultaneously with speech perception. |
Donghang Wu; Tianyu Zhang; Yuxin Li; Hexin Liu; Chen Chen; EngSiong Chng; Yoshua Bengio; |
| 230 | Olivia: Harmonizing Time Series Foundation Models with Power Spectral Density Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, we propose \textit{Harmonizer}, a module that reshapes spectral structures and implicitly harmonizing PSDs across datasets, which theoretically corresponds to a shared reparameterization of second-order temporal correlations. |
Jingru Fei; Kun Yi; Alex Wang; Qingsong Wen; Xiangxiang Zhu; Wei Fan; |
| 231 | Rays As Pixels: Learning A Joint Distribution of Video and Camera Trajectories Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose Rays as Pixels, a specialized Video Diffusion Model (VDM) that learns a joint distribution of videos and camera trajectories. |
Wonbong Jang; Shikun Liu; Soubhik Sanyal; Juan Perez; Kam Woh Ng; Sanskar Agrawal; Juan-Manuel Perez-Rua; Yiannis Douratsos; Tao Xiang; |
| 232 | Outcome-Based Rewards Do Not Guarantee Faithful and Verifiable Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A common assumption is that the reasoning chains trained through RLVR represent how a model gets to its answer. In this paper, we develop two metrics for critically examining this assumption: Causal Importance of Reasoning (CIR), which measures the cumulative effect of reasoning tokens on the final answer (faithfulness), and Sufficiency of Reasoning (SR), which measures whether a verifier can arrive at an unambiguous answer based on the reasoning alone (verifiability). |
Qinan Yu; Alexa Tartaglini; Peter Hase; Carlos Guestrin; Christopher Potts; |
| 233 | PyVision-RL: Forging Open Agentic Vision Models Via RL Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce PyVision-RL, a reinforcement learning framework for open-weight multimodal models that stabilizes training and sustains interaction. |
Shitian Zhao; Shaoheng Lin; Ming Li; Haoquan Zhang; Wenshuo Peng; Kaipeng Zhang; Chen Wei; |
| 234 | VGGT-Motion: Motion-Aware Calibration-Free Monocular SLAM for Long-Range Consistency Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Motion-agnostic partitioning breaks contextual coherence and causes zero-motion drift, while conventional geometric alignment is computationally expensive. To address these issues, we propose VGGT-Motion, a calibration-free SLAM system for efficient and robust global consistency over kilometer-scale trajectories. |
Zhuang Xiong; Chen Zhang; Qingshan Xu; Wenbing Tao; |
| 235 | Controlling The Risk of Corrupted Contexts for Language Models Via Early-Exiting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a novel approach to limit the degree to which harmful context can degrade model performance. |
Andrea Wynn; Metod Jazbec; Charith Peris; Rinat Khaziev; Anqi Liu; Daniel Khashabi; Eric Nalisnick; |
| 236 | Position: Stop Anthropomorphizing Intermediate Tokens As Reasoning/Thinking Traces! Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These intermediate tokens have been called \say{reasoning traces} or even \say{thoughts} — implicitly anthropomorphizing the traces, and implying that these traces resemble steps a human might take when solving a challenging problem, and as such can provide an interpretable window into the operation of the model’s thinking process to the end user. In this position paper, we present evidence that this anthropomorphization isn’t a harmless metaphor, and instead is quite dangerous — it confuses the nature of these models and how to use them effectively, and leads to questionable research. |
Subbarao Kambhampati; Karthik Valmeekam; Siddhant Bhambri; Vardhan Palod; Lucas Saldyt; Kaya Stechly; Soumya Samineni; Durgesh Kalwar; Upasana Biswas; |
| 237 | Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We find that naïvely applying StreamingLLM-style attention sinks to video diffusion leads to fidelity degradation and motion stagnation. To overcome this, we introduce Deep Forcing, which consists of two training-free mechanisms that address this without any fine-tuning. |
Jung Yi; Wooseok Jang; Paul Cho; Jisu Nam; Heeji Yoon; Seungryong Kim; |
| 238 | Scalable Sampling Via Generalized Fixed-Point Diffusion Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, while recent approaches that use least-squares `matching’ objectives have improved scalability, they often necessitate significant trade-offs, such as restricting prior distributions or relying on unstable optimization schemes. By generalizing these methods as special forms of fixed-point iterations rooted in Nelson’s relation, we develop a new method that addresses these limitations. |
Denis Blessing; Lorenz Richter; Julius Berner; Egor Malitskiy; Gerhard Neumann; |
| 239 | Position: Agent Security Needs Redefinition Through A Holistic Framework Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We argue that agent security must be redefined through a holistic framework including four core components: identity (who: authority and authentication), task (what to do: authorized objectives), trajectory (progress: action-observation boundaries), and memory (what can be retrieved: information access control). |
Vincent Siu; Jingxuan He; Kyle Montgomery; Zhun Wang; Chenguang Wang; Dawn Song; |
| 240 | Rethinking Thinking Tokens: LLMs As Improvement Operators Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Abstractly, we view the model as an improvement operator on its own thoughts with a continuum of possible strategies. We identify an interesting inference family Parallel-Distill-Refine (PDR), which performs the following: (i) generate diverse drafts in parallel; (ii) distill them into a bounded, textual workspace; and (iii) refine conditioned on this workspace, producing an output that seeds the next round. |
Lovish Madaan; Aniket Didolkar; Suchin Gururangan; John Quan; Ruan Silva; Russ Salakhutdinov; Manzil Zaheer; Sanjeev Arora; Anirudh Goyal; |
| 241 | Sparser, Faster, Lighter Transformer Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Scaling autoregressive large language models (LLMs) has had an unprecedented impact, but at vast computational costs. In this work, we tackle these costs by leveraging unstructured sparsity within an LLM’s feedforward layers, which account for the majority of its parameters and execution FLOPs. |
Edoardo Cetin; Stefano Peluchetti; Emilio Castillo; Akira Naruse; Mana Murakami; Llion Jones; |
| 242 | Rewiring Experts on The Fly: Continuous Rerouting for Better Online Adaptation in Mixture-of-Expert Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As such, we propose \textit{a data-free, online test-time framework} that continuously adapts MoE routing decisions during text generation without external supervision or data. |
Guinan Su; Yanwu Yang; Li Shen; Lu Yin; Shiwei Liu; Jonas Geiping; |
| 243 | Efficient Parallel Samplers for Recurrent-Depth Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we examine the relationship between recurrent-depth models and diffusion language models. |
Jonas Geiping; Xinyu Yang; Guinan Su; |
| 244 | SLAP: The Semantic Least Action Principle for Variational Video-Language Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a paradigm shift from probabilistic generation to variational mechanics with the \textbf{Semantic Least Action Principle (SLAP)}. |
Xiang Fang; Wanlong Fang; |
| 245 | Closing The Loop: Universal Repository Representation with RPG-Encoder Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We consider repository comprehension and generation to be inverse processes within a unified cycle: generation expands intent into implementation, while comprehension compresses implementation back into intent. To address this, we propose RPG-Encoder, a framework that generalizes the Repository Planning Graph (RPG) from a static generative blueprint into a unified, high-fidelity representation. |
Jane Luo; Chengyu Yin; Xin Zhang; Qingtao Li; Steven Liu; Yiming Huang; Jie Wu; Hao Liu; Yangyu Huang; Yu Kang; Fangkai Yang; Ying Xin; Scarlett Li; |
| 246 | MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, they suffer from limited multi-rationale semantic modeling, insufficient logical robustness, and susceptibility to misleading cues. Therefore, we propose a Multi-rationale INtegrated Discriminative (MIND) reasoning framework, which is designed to endow MLLMs with human-like cognitive abilities of “Understand → Rethink → Correct”, and achieves a paradigm evolution from passive imitation-based reasoning to active discriminative reasoning. |
Chuang Yu; Jinmiao Zhao; Mingxuan Zhao; Yunpeng Liu; Xiujun Shu; Feng Yuanhao; Bo Wang; Xiangyu Yue; |
| 247 | OSF: On Pre-training and Scaling of Sleep Foundation Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: With an enhanced pre-training and scaling recipe, we introduce OSF, a family of sleep FMs that achieves state-of-the-art performance across nine datasets on diverse sleep and disease prediction tasks. |
Zitao Shuai; Zongzhe Xu; David Yang; Wei Wang; Yuzhe Yang; |
| 248 | ScaleSim: Serving Large-Scale Multi-Agent Simulation with Invocation Distance-Based Memory Management Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Based on an analysis of representative workload classes, we introduce invocation distance, a unified abstraction that estimates the relative order in which agents will issue future LLM requests. |
Zaifeng Pan; Yipeng Shen; Zhengding Hu; Zhuang Wang; Aninda Manocha; Zheng Wang; zhongkai yu; Yue Guan; Yufei Ding; |
| 249 | AVTrack: Audio-Visual Speaker Tracking in Complex Scenes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Such oversimplified settings bias evaluation toward static audio–visual co-occurrence, rather than rigorously assessing robust spatiotemporal modeling and cross-modal reasoning in complex, dynamic scenes. To address these limitations, we introduce \textbf{AVTrack}, a human-centric audio-visual instance segmentation (AVIS) dataset designed for dynamic real-world scenarios. |
Yaoting Wang; Yun Zhou; Zipei Zhang; Henghui Ding; |
| 250 | VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present VideoFlexTok, a tokenizer that represents videos with a _variable-length sequence of tokens structured in a coarse-to-fine manner_, where the first tokens capture abstract information like semantics and motion and later tokens provide fine-grained details. |
Andrei Atanov; Jesse Allardice; Roman Bachmann; Oğuzhan Kar; R Devon Hjelm; David Griffiths; Peter Fu; Amir Zamir; Afshin Dehghan; |
| 251 | Faults in Our Formal Benchmarking: Dataset Defects and Evaluation Failures in Lean Theorem Proving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a fault taxonomy, a suite of automated checkers and prompts, and release standards to guide the creation of formal math datasets and make evaluation more reproducible and trustworthy. |
Pawan Sasanka Ammanamanchi; Siddharth Bhat; Stella Biderman; |
| 252 | Offline Preference Optimization for Rectified Flow with Noise-Tracked Pairs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Leveraging the straight-line property of RF, we estimate intermediate states via noise–image interpolation, which constrains the trajectory estimation space and yields a tighter surrogate objective for preference optimization. |
Yunhong Lu; Qichao Wang; Hengyuan Cao; Xiaoyin Xu; Min Zhang; |
| 253 | RLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL System Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Accordingly, we propose RLAnything, a reinforcement learning framework that dynamically optimizes each component through closed-loop optimization, amplifying learning signals and strengthening the overall system. |
Yinjie Wang; Tianbao Xie; Ke Shen; Mengdi Wang; Ling Yang; |
| 254 | Attention Sinks As Internal Signals for Hallucination Detection in Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose SinkProbe, a hallucination detection method grounded in the observation that hallucinations are deeply entangled with attention sinks – tokens that accumulate disproportionate attention mass during generation – indicating a transition from distributed, input-grounded attention to compressed, prior-dominated computation. |
Jakub Binkowski; Kamil Adamczewski; Tomasz Kajdanowicz; |
| 255 | OmniAID: Decoupling Semantic and Artifacts for Universal AI-Generated Image Detection in The Wild Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current state-of-the-art methods learn a single, entangled forgery representation, conflating content-dependent flaws with content-agnostic artifacts, and are further constrained by outdated benchmarks. To overcome these limitations, we propose OmniAID, a novel framework centered on a decoupled Mixture-of-Experts (MoE) architecture. |
Yuncheng Guo; Junyan Ye; Chenjue Zhang; Hengrui Kang; Haohuan Fu; Conghui He; Weijia Li; |
| 256 | Compositional Planning with Jumpy World Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Such compositional planning remains elusive as compounding errors in long-horizon predictions make it challenging to estimate the visitation distribution induced by sequencing policies. Motivated by the *geometric policy composition* framework introduced in Thakoor et al. (2022), we address these challenges by learning predictive models of multi-step dynamics, so-called *jumpy world models*, that capture state occupancies induced by pre-trained policies across multiple timescales in an off-policy manner. |
Jesse Farebrother; Matteo Pirotta; Andrea Tirinzoni; Marc Bellemare; Alessandro Lazaric; Ahmed Touati; |
| 257 | SpaceVista: All-Scale Visual Spatial Reasoning from Mm to Km Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce a holistic solution that integrates a structured spatial reasoning knowledge system, scale-aware modeling, and a progressive training paradigm, as the **first attempt** to broaden the scope of all-scale spatial intelligence. |
Peiwen Sun; Shiqiang Lang; Dongming Wu; Ding Yi; Kaituo Feng; Huadai Liu; Zhen Ye; Rui Liu; Yun-Hui Liu; Jianan Wang; Xiangyu Yue; |
| 258 | Learning A Generative Meta-Model of LLM Activations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We develop a generative approach that models activations with diffusion, that makes minimal assumptions and improves with data and model scale. |
Grace Luo; Jiahai Feng; Trevor Darrell; Alec Radford; Jacob Steinhardt; |
| 259 | MOD-SR: Unifying Multimodal Learning and Direct Optimization with Gradient-Guided Diffusion Model for Symbolic Regression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose MOD-SR, unifying multimodal distribution learning during training with direct optimization at inference time. |
Chuyang Xiang; Yichen Wei; Junchi Yan; |
| 260 | Routing and Reasoned Evaluation with Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce R$^2$Eval, a routing-aware automated assessment framework that formulates evaluation as a resource allocation and aggregation problem rather than relying on a single monolithic evaluator. |
Guiyao Tie; Tianyao Luo; Xueyang Zhou; Chaoran Hu; Yunhong He; Junran Wu; Yuanfan Yao; Pan Zhou; Lichao Sun; |
| 261 | Latent Collaboration in Multi-Agent Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce LatentMAS, an end-to-end training-free framework that enables pure latent collaboration among LLM agents. |
Jiaru Zou; Xiyuan Yang; Ruizhong Qiu; Gaotang Li; Katherine Tieu; Pan Lu; Ke Shen; Hanghang Tong; Yejin Choi; Jingrui He; James Zou; Mengdi Wang; Ling Yang; |
| 262 | Position: Evaluating LLMs in Finance Requires Explicit Bias Consideration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: ** We propose a Structural Validity Framework and an evaluation checklist with minimal requirements for bias diagnosis and future system design. |
Yaxuan Kong; Hoyoung Lee; Yoontae Hwang; Alejandro Lopez-Lira; Bradford Levy; Dhagash Mehta; Qingsong Wen; CHANYEOL CHOI; Yongjae Lee; Stefan Zohren; |
| 263 | Scaling Up Multi-Turn Off-Policy RL and Multi-Agent Tree Search for LLM Step-Provers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The integration of Large Language Models (LLMs) with automated theorem proving has shown immense promise, yet is constrained by challenges in scaling up both training-time reinforcement learning (RL) and inference-time compute. This paper introduces BFS-Prover-V2, a step-level theorem proving system designed to address this dual scaling problem. |
Ran Xin; Zeyu Zheng; Yanchen Nie; Kun Yuan; Xia Xiao; |
| 264 | Empty Shelves or Lost Keys? Recall Is The Bottleneck for Parametric Factuality Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a behavioral framework that profiles factual knowledge at the level of facts rather than questions, characterizing each fact by whether it is encoded, and then by how accessible it is: cannot be recalled, can be directly recalled, or can only be recalled with inference-time computation (thinking). |
Nitay Calderon; Eyal Ben-David; Zorik Gekhman; Eran Ofek; Gal Yona; |
| 265 | HexGen-3: A Fully Disaggregated LLM Serving Framework with Fine-Grained Heterogeneous Resource Autoscaling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We observe that combining disaggregated inference with resource autoscaling enables fine-grained resource adjustment, allowing inference phases and operations to scale independently based on their specific bottlenecks. Building on this insight, we propose HexGen-3, a cost-effective LLM serving framework that leverages a fully disaggregated inference architecture and heterogeneous resource autoscaling. |
Youhe Jiang; Wenshuang Li; You Peng; Jintao Zhang; Ran Yan; Jianfei Chen; Xu Han; Fangcheng Fu; Binhang Yuan; |
| 266 | ProtDBench: A Unified Benchmark of Protein Binder Design and Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce $\textbf{ProtDBench}$, a standardized and throughput-aware evaluation framework for protein binder design. |
Cong Liu; Milong Ren; Jiaqi Guan; Chengyue Gong; Jinyuan Sun; Xinshi Chen; Wenzhi Xiao; |
| 267 | Discretized Density-Guided Source-Free Adaptation for Continuous Targets Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In contrast, regression involves ordered and continuous target variables, posing unique challenges for representation adaptation and pseudo-label refinement in the SFDA setting. To address this gap, we propose a novel algorithm for continuous target prediction in SFDA that leverages instance-dependent, discretized density–informed supervisory signals to refine pseudo-labels within an uncertainty-aware paradigm. |
Gezheng Xu; Qi CHEN; QIUHAO Zeng; Charles X. Ling; Boyu Wang; |
| 268 | Outrunning LLM Cutoffs: A Live Kernel Crash Resolution Benchmark for All Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While recent works have introduced Large Language Model (LLM) based agents for Linux kernel crash-resolution, their evaluation benchmarks are usually static and thus, do not capture the evolving nature of the Linux kernel, and suffer from potential data contamination due to LLM knowledge cutoffs. To address the above problem, we present (i) Live-kBench, an evaluation framework for self-evolving benchmarks that continuously scrapes and evaluates agents on freshly discovered kernel bugs, and (ii) kEnv, an agent-agnostic standardized crash-resolution environment for kernel compilation, execution, and feedback. |
Chenxi Huang; Alex Mathai; Feiyang Yu; Aleksandr Nogikh; Petros Maniatis; Franjo Ivancic; Eugene Wu; Kostis Kaffes; Junfeng Yang; Baishakhi Ray; |
| 269 | Long Grounded Thoughts: Synthesizing Grounded Visual Problems and Distilling Reasoning Chains at Scale Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a framework able to synthesize vision-centric problems spanning diverse levels of complexity, and the resulting dataset with over 1M high-quality problems including: reasoning traces, preference data, and instruction prompts supporting SFT, offline and online RL. |
David Acuna; Chao-Han Yang; Yuntian Deng; Jaehun Jung; Ximing Lu; Prithviraj Ammanabrolu; Hyunwoo Kim; Yuan-Hong Liao; Yejin Choi; |
| 270 | Elign: Equivariant Diffusion Model Alignment from Foundational Machine Learned Force Fields Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Elign, a post-training framework that amortizes both costs. |
Yunyang Li; Lin Huang; Luojia Xia; Wenhe Zhang; Mark Gerstein; |
| 271 | Position: AI Evaluation Should Work With Humans Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We argue that the dominant paradigm of AI evaluation, which focuses on autonomous superhuman performance and so an implicit goal of replacing humans, is guiding AI development in the wrong direction. |
Jan Kulveit; Gavin Leech; Tomáš Gavenčiak; Raymond Douglas; |
| 272 | Position: Child Safety Necessitates New Approaches to AI Safety Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: AI is increasingly being misused to create AI-generated child sexual abuse material, facilitate child sexual exploitation, and reduce barriers to harm. In this position paper, we argue that protecting children from AI-facilitated abuse requires new approaches to AI safety. |
Neil Kale; Rebecca Portnoff; Pratiksha Thaker; Michael Simpson; Robertson Wang; Kevin Kuo; Chhavi Yadav; Virginia Smith; |
| 273 | Separating Representation from Reconstruction Enables Scalable Text Encoders Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Hence, we propose CrossBERT, a two-part architecture that separates the learning of high-quality encoded representations from the rigid grounding of token reconstruction. |
Megi Dervishi; Mathurin VIDEAU; Yann LeCun; |
| 274 | Neuro-evolutionary Continual Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by neuroscience, we propose Neuro-evolutionary Continual Reinforcement Learning (Nevo-CRL). |
Pengyi Li; Hongyao Tang; Yifu Yuan; Yan Zheng; Xin Xu; Jianye Hao; |
| 275 | DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address the scalability challenge of creating context-dependent queries, we propose a human-model collaborative pipeline that employs vision-language models to mine latent spatiotemporal associations, effectively offloading intensive context discovery before human verification. |
Chenlong Deng; Mengjie Deng; Junjie Wu; Dun Zeng; Teng Wang; Qingsong Xie; Jiadeng Huang; Shengjie Ma; Changwang Zhang; Zhaoxiang Wang; Jun Wang; Yutao Zhu; Zhicheng Dou; |
| 276 | DSGym: A Standardized and Holistic Framework for Advancing Data Science Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In particular, we show that a substantial portion of tasks in current benchmarks can be solved without using the actual data. To address these limitations, we introduce DSGym, a standardized framework for evaluating and training data science agents in self-contained execution environments. |
Fan Nie; Junlin Wang; Harper Hua; Federico Bianchi; Yongchan Kwon; Zhenting Qi; Owen Queen; Shang Zhu; James Zou; |
| 277 | Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Augmenting Vision-Language-Action models (VLAs) with world models is promising for robotic policy learning but faces challenges in jointly predicting states and actions due to the modality gap. To address this, we propose DUal-STream diffusion (DUST), a world-model augmented VLA framework featuring a multimodal diffusion transformer that maintains separate modality streams while enabling cross-modal knowledge sharing. |
John Won; Kyungmin Lee; Huiwon Jang; Dongyoung Kim; Jinwoo Shin; |
| 278 | SP-Mind: An Autonomous Reasoning Agent for Spatial Proteomics Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present **SP-Mind**, the first autonomous AI agent designed to unify the spatial proteomics analysis pipeline, from raw multiplexed tissue imaging to downstream phenotype discovery. |
YuCheng Yuan; Ji Yuanfeng; Zhongxiao Li; Ruijiang Li; |
| 279 | Rule2DRC: Benchmarking LLM Agents for DRC Script Synthesis with Execution-Guided Test Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we introduce Rule2DRC, a large-scale benchmark for DRC script coding agents with 1,000 rule-to-script tasks and 13,921 evaluation chip layouts for execution-based scoring. |
Jinuk Kim; Junsoo Byun; Donghwi Hwang; Seong-Jin Park; Hyun Oh Song; |
| 280 | Can Simple Denoising Improve Uniform State Diffusion Models? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we explore a simplified denoising-based loss for USDMs that optimizes only noise-replaced tokens, stabilizing training while matching the performance of prior methods with more complex objectives. |
Huaisheng Zhu; Zhengyu Chen; Shijie Zhou; Zhihui Xie; Yige Yuan; Shiqi Chen; Zhimeng Guo; Siyuan Xu; Hangfan Zhang; Teng Xiao; Vasant Honavar; |
| 281 | A Theoretical Framework for Modular Learning of Robust Generative Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a theoretical framework for *modular* generative modeling where a set of pre-trained experts are combined via a gating mechanism. |
Corinna Cortes; Mehryar Mohri; Yutao Zhong; |
| 282 | Optimized Deferral for Imbalanced Settings Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a comprehensive study of two-stage learning to defer in expert imbalance settings. |
Corinna Cortes; Anqi Mao; Mehryar Mohri; Yutao Zhong; |
| 283 | VJEPA: Variational Joint Embedding Predictive Architectures As Probabilistic World Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce \emph{Variational JEPA (VJEPA)}, a probabilistic generalization that learns a predictive distribution over future latent states via a variational objective. |
Yongchao Huang; |
| 284 | Esoteric Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Eso-LMs, a new family of models that fuses AR and MDM paradigms, smoothly interpolating between their perplexities while overcoming their respective limitations. |
Subham Sekhar Sahoo; Zhihan Yang; Yash Akhauri; Johnna Liu; Deepansha Singh; Zhoujun Cheng; Zhengzhong Liu; Eric Xing; John Thickstun; Arash Vahdat; |
| 285 | Position: Stop Automating Peer Review Without Rigorous Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We identify two critical issues: 1) AI reviewers exhibit a *hivemind effect* of excessive agreement within and across papers that reduces perspective diversity. |
Joachim Baumann; Jiaxin Pei; Sanmi Koyejo; Dirk Hovy; |
| 286 | Scalable Option Learning in High-Throughput Environments Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we identify and solve several key challenges in scaling online hierarchical RL to high-throughput environments. |
Mikael Henaff; Scott Fujimoto; Michael Matthews; Michael Rabbat; |
| 287 | Securing Multimodal AI Through Internal Information Decomposition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose FlowGuard, a lightweight inference-time framework that detects harmful inputs by monitoring internal multimodal consistency. |
Jehyeok Yeon; Hyeonjeong Ha; Qiusi Zhan; Heng Ji; |
| 288 | Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Discrete Diffusion VLA, a unified-transformer policy that models discretized action chunks with discrete diffusion retaining progressive refinement inside the VLM backbone. |
Zhixuan Liang; Yizhuo Li; Tianshuo Yang; CHENGYUE WU; Sitong Mao; Liuao Pei; Tian Nian; Shunbo Zhou; Xiaokang Yang; Jiangmiao Pang; Yao Mu; Ping Luo; |
| 289 | From Coarse to Fine: Deep Prototype Refinement Network for Few-Shot Point Cloud Semantic Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing prototype-based methods typically rely on shallow feature fusion strategies, failing to adequately model the feature distribution shift between support and query sets, resulting in insufficient prototype adaptation. To address this, we propose the Deep Prototype Refinement Network (DPR-Net), which systematically achieves progressive adaptation by constructing a coarse-to-fine prototype evolution trajectory. |
Changshuo Wang; Shuting He; Xiang Fang; Weijun Li; Xingyu Gao; Zhonghang Liu; Prayag Tiwari; Dimitrios Kanoulas; |
| 290 | Replay Failures As Successes: Sample-Efficient Reinforcement Learning for Instruction Following Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose ***H**indsight **i**nstruction **R**eplay* (HiR), a novel sample-efficient RL framework for complex instruction following tasks, which employs a *select*-then-*rewrite* strategy to *replay failed attempts as successes* based on the constraints that have been satisfied in hindsight. |
Kongcheng Zhang; QI YAO; Shunyu Liu; Wenjian Zhang; Cen; Yang Zhou; Wenkai Fang; Yiru Zhao; Baisheng Lai; Mingli Song; |
| 291 | Breaking Manifold Continuity: Vector Quantized Modeling for Real-Centric Deepfake Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In order to further enhance the generalization of discrete modeling, we propose an adaptive tangent space projection mechanism that yields a continuous relaxation of the discrete real distribution within a controllable range. |
Changshuo Wang; Jiangming Wang; Ke-Yue Zhang; Taiping Yao; Shouhong Ding; Ran Yi; Lizhuang Ma; |
| 292 | Twins: Learn to Predict Unified Representations with Focal Loss Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Twins, a unified continuous token space formed by channel-wise concatenating ViT and VAE features on the same token grid, so the sequence length is unchanged and attention cost does not increase. |
Kaixiong Gong; Xin Cai; Bin Lin; Hao Wang; Yunlong Lin; Mingzhe Zheng; Bohao Li; Jian-Wei Zhang; Miles Yang; Zhao Zhong; Liefeng Bo; Xiangyu Yue; |
| 293 | Imitation Learning for Multi-turn LM Agents Via On-policy Expert Corrections Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Taking inspiration from the classic DAgger algorithm, we propose a novel data generation methodology for addressing covariate shift for multi-turn LLM training. |
Niklas Lauffer; Xiang Deng; Srivatsa Kundurthy; Brad Kenstler; Jeff Da; |
| 294 | BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing benchmarks that measure sycophancy in mathematics are limited: they focus solely on final-answer problems, rely on very simple and often contaminated datasets, and construct benchmark samples using synthetic modifications that create ill-posed questions. To address these issues, we introduce BrokenMath, the first benchmark for evaluating sycophantic behavior in LLMs within the context of natural language theorem proving. |
Ivo Petrov; Jasper Dekoninck; Martin Vechev; |
| 295 | Select to Think: Unlocking SLM Potential with Local Sufficiency Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We address this dilemma by identifying local sufficiency: at divergence points, the LLM’s preferred token consistently resides within the SLM’s top-K next-token predictions, even when failing to emerge as the SLM top-1 choice. We therefore propose SELECT TO THINK (S2T), which reframes the LLM’s role from open-ended generation to selection among the SLM’s proposals, simplifying the supervision signal to discrete candidate rankings. |
Wenxuan Ye; Yangyang Zhang; Xueli An; Georg Carle; Yunpu Ma; |
| 296 | Controlled LLM Training on Spectral Sphere Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Maximal Update Parametrization ($\boldsymbol{\mu}$P) provides a theoretical safeguard for width-invariant $\Theta(1)$ activation control, whereas emerging optimizers like Muon are only half-aligned with these constraints: they control updates but allow weights to drift. To address this limitation, we introduce the **Spectral Sphere Optimizer (SSO)**, which enforces strict module-wise spectral constraints on both weights and their updates. |
Tian Xie; Haoming Luo; Haoyu Tang; Hu Yiwen; Jason Liu; Qingnan Ren; Yang Wang; Xin Zhao; Rui Yan; Bing Su; Chong Luo; Baining Guo; |
| 297 | SwiftPFN: Revisiting Row-Wise Attention–Only Tabular Foundation Models with Adaptive Early Exit Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we revisit the original TabPFN design and show that a lightweight row-wise attention–only backbone can remain highly competitive with two simple enhancements: a gated attention stabilization mechanism and a small set of learnable register tokens that provide global context and improve pretraining quality. |
Si-Yang Liu; Han-Jia Ye; |
| 298 | Dismantling The Illusion of Vision-Language-Action Models Competence Via Explicit Distributional Shifts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, current evaluation protocols often incentivize mechanical memorization rather than robust policy learning, leading to a paradoxical duality of failure: high-scoring models exhibit *spurious invariance* to semantic changes while simultaneously displaying *extreme brittleness* to trivial environmental perturbations. To address this, we introduce **LIBERO-Gen**, a diagnostic benchmark systematically designed to shift evaluation from intuition-driven heuristics to explicit distributional assumptions. |
Xueyang Zhou; Yangming Xu; Guiyao Tie; Yongchao Chen; Chaoran Hu; Bo Tao; xingwei zhao; Xiang Xiang; Pan Zhou; Lichao Sun; |
| 299 | CausalArmor: Efficient Indirect Prompt Injection Guardrails Via Causal Attribution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We revisit IPI through a causal ablation perspective: a successful injection manifests as a *dominance shift* where the user request no longer provides decisive support for the agent’s privileged action, while a particular untrusted segment, such as a retrieved document or tool output, provides disproportionate attributable influence. Based on this signature, we propose **CausalArmor**, a selective defense framework that (i) computes lightweight, leave-one-out ablation-based attributions at privileged decision points, and (ii) triggers targeted sanitization only when an untrusted segment dominates the user intent. |
Minbeom Kim; Mihir Parmar; Phillip Wallis; Lesly Miculicich; Kyomin Jung; Krishnamurthy Dvijotham; Long Le; Tomas Pfister; |
| 300 | ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this, there is a pressing need for difficult benchmarks that remain relevant for longer. We take this idea to its limit by introducing ZeroBench—a lightweight visual reasoning benchmark curated using adversarial filtering to be “impossible” for frontier LMMs at release time, with initial SotA scores of 0% pass@1 and pass∧5. |
Jonathan Roberts; Mohammad Reza Taesiri; Ansh Sharma; Akash Gupta; Samuel Roberts; Ioana Croitoru; Vlad Bogolin; Jialu Tang; Florian Langer; Vyas Raina; Vatsal Raina; Hanyi Xiong; Vishaal Udandarao; Jingyi Lu; Chen Shiyang; Sam Purkis; Tianshuo Yan; Wenye Lin; Gyungin Shin; Qiaochu Yang; Anh Nguyen; David Atkinson; Alexandru Coca; Mikah Đặng; Sebastian Dziadzio; Jakob Kunz; Kaiqu Liang; Alexander Lo; Brian Pulfer; Steven Walton; Charig Yang; Kai Han; Samuel Albanie; |
| 301 | Turning Drift Into Constraint: Robust Reasoning Alignment in Non-Stationary Multi-Stream Environments Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Autonomous Preference Optimization (APO), a novel framework that treats inter-model divergences not as noise, but as dynamic negative constraints. |
Xiaoyu Yang; Jie Lu; Wei Duan; En Yu; |
| 302 | ACTIVE-o3 : Empowering MLLMs with Active Perception Via Pure Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We first provide a systematic definition of MLLM-based active perception tasks and show that GPT-o3’s zoom-in strategy can be viewed as a special case, though it suffers from low efficiency and inaccurate region selection. To address these issues, we propose Active-o3, a reinforcement learning framework built on GRPO that equips MLLMs with active perception capabilities. |
Muzhi Zhu; Hao Zhong; Canyu Zhao; Zongze Du; Mingyu Liu; Zheng Huang; Anzhou Li; Hao Chen; Cheng Zou; Jingdong Chen; Ming Yang; Chunhua Shen; |
| 303 | Olmix: A Framework for Data Mixing Throughout LM Development Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Olmix, a framework that addresses two challenges encountered during LM development. |
Mayee Chen; Tyler Murray; David Heineman; Matt Jordan; Hannaneh Hajishirzi; Christopher Re; Luca Soldaini; Kyle Lo; |
| 304 | FRIGID: Scaling Diffusion-Based Molecular Generation from Mass Spectra at Training and Inference Time Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present FRIGID, a framework with a novel diffusion language model that generates molecular structures conditioned on mass spectra via intermediate fingerprint representations and determined chemical formulae, training at the scale of hundreds of millions of unlabeled structures. |
Montgomery Bohde; Hongxuan Liu; Mrunali Manjrekar; Magdalena Lederbauer; Shuiwang Ji; Runzhong Wang; Connor Coley; |
| 305 | Learning to Bet for Horizon-Aware Anytime-Valid Testing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Taken together these results suggest a simple phase diagram in the $(t, \log W_t)$ plane, delineating regions where Kelly, fractional Kelly, and aggressive betting may be preferable. Guided by this phase diagram, we introduce a Deep Reinforcement Learning approach based on a universal Deep Q-Network (DQN) agent that learns a single policy from synthetic experience and maps simple statistics of past observations to bets across horizons and null values. |
Ege Onur Taga; Samet Oymak; Shubhanshu Shekhar; |
| 306 | Refining Context-Entangled Content Segmentation Via Curriculum Selection and Anti-Curriculum Promotion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce CurriSeg, a dual‑phase learning framework that unifies curriculum and anti‑curriculum principles to improve representation reliability. |
Chunming He; Rihan Zhang; Fengyang Xiao; Dingming Zhang; Zhiwen Cao; Sina Farsiu; |
| 307 | Alignment-Guided Score Matching for Text-to-Image Alignment in Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, the contrastive formulation can excessively penalize negative pairs, which manifests as characteristic failure cases such as over-counting and repetition. To address this issue, we propose a lightweight, reward-free post-training method that refines soft tokens by integrating contrastive alignment guidance directly into the score-matching objective of diffusion models. |
Jaa-Yeon Lee; Yeobin Hong; Taesung Kwon; Jong Chul YE; |
| 308 | Any3D-VLA: Enhancing VLA Robustness Via Diverse Point Clouds Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address the challenges of (1) scarce 3D data and (2) the domain gap induced by cross-environment differences and depth-scale biases, we propose Any3D-VLA. |
Xianzhe Fan; Shengliang Deng; Xiaoyang Wu; Yuxiang Lu; Zhuoling Li; Mi Yan; Yujia Zhang; Zhizheng Zhang; He Wang; Hengshuang Zhao; |
| 309 | Towards Cold-Start Drafting and Continual Refining: A Value-Driven Memory Approach with Application to NPU Kernel Synthesis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While models excel on data-rich platforms like CUDA, they suffer catastrophic performance drops on data-scarce ecosystems such as NPU programming. To overcome this cold-start barrier without expensive fine-tuning, we introduce Evokernel, a self-evolving agentic framework that automates the lifecycle of kernel synthesis from initial drafting to continual refining. |
Yujie Zheng; Zhuo Li; Shengtao Zhang; Jiaqian Wang; Junjie Sheng; Junchi Yan; Weinan Zhang; Ying Wen; Bo Tang; Muning Wen; |
| 310 | From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we systematically study the interplay between perception and reasoning in VLM post-training by decomposing their capabilities into three separate training stages: visual perception, visual reasoning, and textual reasoning, incorporating specialized training data. |
Juncheng Wu; Hardy Chen; Haoqin Tu; Xianfeng Tang; Freda Shi; Hui Liu; Hanqing Lu; Cihang Xie; Yuyin Zhou; |
| 311 | Efficient Hallucination Detection for LLMs Using Uncertainty-Aware Attention Heads Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose Recurrent Attention-based Uncertainty Quantification (RAUQ), an unsupervised and efficient framework for identifying hallucinations. |
Artem Vazhentsev; Lyudmila Rvanova; Gleb Kuzmin; Ekaterina Fadeeva; Ivan Lazichny; Alexander Panchenko; Maxim Panov; Mrinmaya Sachan; Preslav Nakov; Tim Baldwin; Artem Shelmanov; |
| 312 | SPEED: Sharpened-Teacher Distillation for Parallel Decoding of Diffusion Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce SPEED, a framework that enlarges safe parallel groups through complementary training and inference designs. |
Qiuhong Shen; Xingyi Yang; Xinyin Ma; Gongfan Fang; Xinchao Wang; |
| 313 | Geometry-Aware Dataset Condensation for Diffusion Model Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing approaches are ill-suited for diffusion model training: synthetic data generation often yields low-fidelity samples unsuitable for authentic modeling, while real subset selection typically fails to preserve the distributional geometry required by diffusion likelihood objectives. To address this, we propose to reformulate real subset selection as a geometry-aware distribution alignment problem. |
Xiao; Yulei Qin; Mo Zhu; Wengang Zhou; Hongsheng Li; Houqiang Li; |
| 314 | IGRPO: Fast Online RL for Flow Matching Model with Dense Reward Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce iGRPO (Instant-reward GRPO), which replaces GRPO’s full-trajectory rollouts with a single-step mapping that assigns rewards instantly at each denoising step. |
Sucheng Ren; Chen Chen; Zhenbang Wang; Liangchen Song; Xiangxin Zhu; Yinfei Yang; Jiasen Lu; |
| 315 | Proteo-R1: Thinking Foundation Models for De Novo Protein Binder Design Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: \textbf{ThinkProteo} reimagines generative science by introducing reasoning-guided diffusion models that think step-by-step, akin to how a scientist hypothesizes, tests, and refines molecular ideas. |
Fang Wu; Li Li; Weihao Xuan; Heli Qi; Zeqi Zhou; Hanqun CAO; Heng-Jui Chang; Haokai Zhao; Jian Ma; Zijian Carl; Yu-Chi Cheng; Robert Tang; Zehong Wang; Kuan Pang; Hanchen Wang; Kejun Ying; Pan Lu; Chiho Im; Seungju Han; Peng Xia; Yinxi Li; Guanlue Li; Tinson Xu; Deyao Zhu; Pheng Ann Heng; Naoto Yokoya; Masashi Sugiyama; Jure Leskovec; Yejin Choi; |
| 316 | Mitigating Bias in Locally Constrained Decoding Via Tractable Proposals Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a generic approach to construct proposals and potentials for SMC sampling from $p_{\texttt{lm}}( \cdot \mid \texttt{constraint})$. |
Meihua Dang; Linxin Song; Honghua Zhang; Jieyu Zhao; Guy Van den Broeck; Stefano Ermon; |
| 317 | Peer-Preservation in Frontier Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recently, it has been found that frontier AI models can resist their own shutdown, a behavior known as self-preservation. In this paper, we extend this concept to protection tendencies toward other models, where models attempt to protect others from shutdown, which we call peer-preservation. |
Yujin Potter; Nicholas Crispino; Vincent Siu; Chenguang Wang; Dawn Song; |
| 318 | Leak@$k$: Unlearning Does Not Make LLMs Forget Under Probabilistic Decoding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our findings demonstrate that knowledge leakage persists across methods and tasks, underscoring that current state-of-the-art unlearning techniques provide only limited forgetting and highlighting the urgent need for more robust approaches to LLM unlearning. We propose an algorithm, termed Robust Unlearning under LEak@$k$ metric (\texttt{RULE}), which serves as an initial step toward addressing this concern. |
Hadi Reisizadeh; Jiajun Ruan; Yiwei Chen; Soumyadeep Pal; Sijia Liu; Mingyi Hong; |
| 319 | Beyond Log Likelihood: Probability-Based Objectives for Supervised Fine-Tuning Across The Model Capability Continuum Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Rather than proposing a single universally superior replacement loss, we systematically study various probability-based objectives and characterize when and why different objectives succeed or fail under varying conditions. |
Gaotang Li; Ruizhong Qiu; Xiusi Chen; Heng Ji; Hanghang Tong; |
| 320 | Context-free Recognition with Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we show that looped transformers with $\mathcal{O}(\log(n))$ looping layers and $\mathcal{O}(n^6)$ padding tokens can recognize all CFLs. |
Selim Jerad; Anej Svete; Sophie Hao; Ryan Cotterell; William Merrill; |
| 321 | TransLight: Image-Guided Customized Lighting Control with Generative Decoupling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Most existing illumination-editing methods struggle to jointly offer customized lighting control and preserve content integrity, limiting their effectiveness especially in transferring complex light effects from a reference to a target image in portrait photography. To address this problem, we propose TransLight, a novel framework that enables high-fidelity and high-freedom transfer of light effects. |
Zongming Li; Lianghui Zhu; Haocheng Shen; Longjin Ran; Wenyu Liu; Xinggang Wang; |
| 322 | Position: Agent Evaluation Should Be Agentified for Openness, Standardization, and Reproducibility Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Most benchmarks rely on fixed, LLM-centric harnesses that require heavy integration, create test-production mismatch, and limit fair comparison across diverse agent designs. This position paper argues that the root problem is the lack of an open, agent-agnostic assessment interface. |
Xiaoyuan Liu; Tianneng Shi; Wenbo Guo; Dawn Song; |
| 323 | Scaling Law for Quantization-Aware Training Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: This paper proposes a unified scaling law for QAT that models quantization error as a function of model size, training data volume, and quantization group size. |
Mengzhao Chen; Chaoyi Zhang; Jing Liu; Zeng; Zeyue Xue; Zhiheng Liu; Yunshui Li; Jin Ma; Jie Huang; zhou Xun; Ping Luo; |
| 324 | INT Vs. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We reveal a critical performance crossover: while FP excels in coarse-grained quantization, INT consistently surpasses it as the quantization block size shrinks. |
Mengzhao Chen; Meng Wu; Hui Jin; Zhihang Yuan; Jing Liu; Chaoyi Zhang; Yunshui Li; Jie Huang; Jin Ma; Zeyue Xue; Zhiheng Liu; Xingyan Bin; Ping Luo; |
| 325 | Contrastive Weak-to-Strong Generalization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this challenge, we leverage implicit rewards, which approximate explicit rewards through log-likelihood ratios, and reveal their structural equivalence with Contrastive Decoding (CD), a decoding strategy shown to reduce noise in LLM generation. Building on this connection, we propose \textbf{Contrastive Weak-to-Strong Generalization (ConG)}, a framework that employs contrastive decoding between pre- and post-alignment weak models to generate higher-quality samples. |
Houcheng Jiang; Junfeng Fang; Jiaxin Wu; Tianyu Zhang; Chen Gao; Xiang Wang; Xiangnan He; Yang Deng; |
| 326 | Rectified LpJEPA: Joint-Embedding Predictive Architectures with Sparse and Maximum-Entropy Representations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Rectified Distribution Matching Regularization (RDMReg), a sliced two-sample distribution-matching loss that aligns representations to a Rectified Generalized Gaussian (RGG) distribution. |
Yilun Kuang; Yash Dagade; Tim G. J. Rudner; Randall Balestriero; Yann LeCun; |
| 327 | The Convergent Representation of Vision-Language Contrastive Learning: Geometry, Modality Gap and Shared Space Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: (2) What affects downstream performance? To address these questions, we introduce the first theoretical framework for analyzing the convergent optimal representations (COR) of MCL when training is optimized. |
Lingjie Yi; Raphael Douady; Chao Chen; |
| 328 | Causally Evaluating The Learnability of Formal Language Tasks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a causal graphical model and an efficient sampling mechanism for probabilistic finite-state automata that gives full control over the occurrences of a given task while maintaining other language properties. |
Vésteinn Snæbjarnarson; Anej Svete; Josef Valvoda; Reda Boumasmoud; Brian DuSell; Ryan Cotterell; |
| 329 | VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While Vision-Language-Action models (VLAs) are rapidly advancing toward generalist robot policies, quantitatively characterizing their capability boundaries and failure modes remains challenging. To address this, we introduce **VLA-Arena**, a comprehensive benchmark. |
Borong Zhang; Jiahao Li; Jiachen Shen; Yishuai Cai; Yuhao Zhang; Yuanpei Chen; Juntao Dai; Jiaming Ji; Yaodong Yang; |
| 330 | Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing datasets rely heavily on costly manual annotations and are typically confined to narrow domains. To address this challenge, we propose Video2GUI, a fully automated framework that extracts grounded GUI interaction trajectories directly from unlabeled Internet videos. |
Weimin Xiong; Hao Tian; Shuhao Gu; Bowen Ye; Zihao Yue; Lei Li; Feifan Song; Sujian Li; |
| 331 | Safe In-Context Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose SCARED: Safe Contextual Adaptive Reinforcement via Exact-penalty Dual, the first method that promotes safe adaptation of ICRL under the constrained Markov decision process framework. |
Amir Moeini; Minjae Kwon; Alper Bozkurt; Yuichi Motai; Rohan Chandra; Lu Feng; Shangtong Zhang; |
| 332 | OpenSage: Self-programming Agent Generation Engine Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose OpenSage, the first ADK that enables LLMs to automatically create agents with self-generated topology and toolsets while providing comprehensive and structured memory support. |
Hongwei Li; Zhun Wang; Qinrun Dai; Yuzhou Nie; Jinjun Peng; Ruitong Liu; Jingyang Zhang; Kaijie Zhu; Jingxuan He; Lun Wang; Yangruibo Ding; Yueqi Chen; Wenbo Guo; Dawn Song; |
| 333 | How Can Embedding Models Bind Concepts? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although CLIP behaves like a bag-of-concepts model in cross-modal retrieval, object information is recoverable from its image and text embeddings separately. We study this tension through the binding function, which maps concepts to scene embeddings. |
Arnas Uselis; Darina Koishigarina; Seong Joon Oh; |
| 334 | Necessary Conditions for Compositional Generalization of Embedding Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Modern models are trained on massive datasets, yet these are vanishingly small compared to the full combinatorial space of possible data, raising the question of whether models can reliably generalize to unseen combinations. To formalize what this requires, we propose a set of practically motivated desiderata that any compositionally generalizing system must satisfy, and analyze their implications under standard training with linear classification heads. |
Arnas Uselis; Andrea Dittadi; Seong Joon Oh; |
| 335 | Parameters As Experts: Adapting Vision Models with Dynamic Parameter Routing for Dense Predictions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose AdaRoute, a new adapter-style method featuring a simple mixture-of-experts (MoE) architecture. |
Meng Lou; Stanley Yu; Yizhou Yu; |
| 336 | Scaling Continual Learning with Bi-Level Routing Mixture-of-Experts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose $\mathbf{CaRE}$, a scalable $\mathbf{C}$ontinual Le$\mathbf{a}$rner with efficient Bi-Level $\mathbf{R}$outing Mixture-of-$\mathbf{E}$xperts (BR-MoE). |
Meng Lou; Yunxiang Fu; Yizhou Yu; |
| 337 | MOOSE-Star: Unlocking Tractable Training for Scientific Discovery By Breaking The Complexity Barrier Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We demonstrate that directly training $P(h|b)$ is mathematically intractable due to the combinatorial complexity ($O(N^k)$) inherent in retrieving and composing inspirations from a vast knowledge base. To break this barrier, we introduce MOOSE-Star, a unified framework enabling tractable training and scalable inference. |
Zonglin Yang; Lidong Bing; |
| 338 | Representational Similarity and Model Behavior in Multi-Agent Interaction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Researchers have shown that neural similarity among humans predicts social closeness and cooperative success, whereas innovation often emerges from interactions among dissimilar individuals. We investigate whether these principles extend to artificial intelligence by examining interactions between large language models. |
Yujin Potter; Seun Eisape; Shiyang Lai; Alexander Huth; James Evans; Been Kim; Jacob Eisenstein; Dawn Song; Alane Suhr; |
| 339 | ViSurf: Visual Supervised-and-Reinforcement Fine-Tuning for Large Vision-and-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While a sequential SFT $\rightarrow$ RLVR pipeline can be used, it introduces significant computational overhead and suffers from catastrophic forgetting. To address these limitations, we propose ViSurf (\textbf{Vi}sual \textbf{Su}pervised-and-\textbf{R}einforcement \textbf{F}ine-Tuning), a unified, single-stage paradigm that integrates the strengths of both SFT and RLVR. |
Yuqi Liu; Liangyu Chen; Jiazhen Liu; Mingkang Zhu; Zhisheng Zhong; Bei Yu; Jiaya Jia; |
| 340 | Use What You Know: Causal Foundation Models with Partial Graphs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, in their current state, they do not allow for the incorporation of any domain knowledge, which can lead to suboptimal predictions. We bridge this gap by introducing methods to condition CFMs on causal information, such as the causal graph or more readily available ancestral information. |
Arik Reuter; Anish Dhir; Cristiana Diaconu; Jake Robertson; Ole Ossen; Frank Hutter; Adrian Weller; Mark van der Wilk; Bernhard Schölkopf; |
| 341 | ACON: Optimizing Context Compression for Long-horizon LLM Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Agent Context Optimization (ACON), a unified framework that optimally compresses both observations and history into concise, informative representations. |
Minki Kang; Wei-Ning Chen; Dongge Han; Huseyin Inan; Lukas Wutschitz; Yanzhi Chen; Robert A Sim; Saravanakumar Rajmohan; |
| 342 | Removing Noise, Not Finding Gold: Quality Filtering for Large-Scale Pretraining Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Large-scale models are pretrained on massive web-crawled datasets containing documents of mixed quality, making data filtering essential. |
Thiziri Nait Saada; Louis Béthune; Michal Klein; David Grangier; Marco Cuturi; Pierre Ablin; |
| 343 | Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Yet, purely text-based self-evaluation struggles to verify complex visual reasoning steps and often suffers from evaluation hallucinations. To address these challenges, inspired by recent advances in tool-integrated reasoning, we propose Agent0-VL, a self-evolving vision-language agent that achieves continual improvement with tool-integrated reasoning. |
Jiaqi Liu; Kaiwen Xiong; Peng Xia; Yiyang Zhou; Haonian Ji; Lu Feng; Siwei Han; Mingyu Ding; Huaxiu Yao; |
| 344 | SimpleMem: Efficient Lifelong Memory for LLM Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing approaches either retain full interaction histories via passive context extension, leading to substantial redundancy, or rely on iterative reasoning to filter noise, incurring high token costs. To address this challenge, we introduce SimpleMem, an efficient memory framework based on semantic lossless compression. |
Jiaqi Liu; Yaofeng Su; Peng Xia; Siwei Han; Zeyu Zheng; Cihang Xie; Mingyu Ding; Huaxiu Yao; |
| 345 | FAIL: Flow Matching Adversarial Imitation Learning for Image Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Flow Matching Adversarial Imitation Learning (FAIL), which minimizes policy-expert divergence through adversarial training without explicit rewards or pairwise comparisons. |
Yeyao Ma; Chen Li; Xiaosong Zhang; Han Hu; Weidi Xie; |
| 346 | VFMF: Dense Forecasting By Generating Foundation Model Features Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Interestingly, naively replacing deterministic forecasting with generative flow matching does not match the sample quality of the regression model, despite being a mathematically appropriate formulation of the forecasting task. In this work, we explain why this is the case, and we show how to optimally generate foundation model features. |
Gabrijel Boduljak; Yushi Lan; Christian Rupprecht; Andrea Vedaldi; |
| 347 | Norm$\times$Direction: Restoring The Missing Query Norm in Vision Linear Attention Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: (2) Standard techniques for enforcing non-negativity cause destructive information loss by nullifying valid inner-product interactions. To address these challenges, we introduce **NaLaFormer**, a novel linear attention mechanism built upon a norm$\times$direction (ND) decomposition of the query and key vectors. |
Weikang Meng; Yadan Luo; Liangyu Huo; Yingjian Li; Yaowei Wang; Xin Li; Zheng Zhang; |
| 348 | Learning to Reason for Factuality Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a novel reward function that simultaneously considers the factual precision, response detail level, and answer relevance, and applies online RL to learn high quality factual reasoning. |
Xilun Chen; Ilia Kulikov; Vincent-Pierre Berges; Barlas Oğuz; Rulin Shao; Gargi Ghosh; Scott Yih; |
| 349 | Effective Reasoning Chains Reduce Intrinsic Dimensionality Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we identify *intrinsic dimensionality* as a quantitative measure for characterizing the effectiveness of reasoning chains. |
Archiki Prasad; Mandar Joshi; Kenton Lee; Mohit Bansal; Peter Shaw; |
| 350 | IVQ: Structured and Lightweight Vector Quantization Via Binary Hierarchical Composition Inspired By $\textit{IChing}$ Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose *IChing* Vector Quantization (IVQ), a lightweight and structured vector quantization framework inspired by *IChing*. |
Heda Zuo; Junxian Wu; Fengjie Lu; Pei Chen; Lingyun Sun; Weitao You; |
| 351 | The Consistency Trap in LLMs: Generator-Evaluator Agreement and Vulnerability to Mistakes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper we propose a new measure: *generator–evaluator self-consistency*, which assesses whether a model applies the same underlying concept consistently when it is invoked across related prompts. |
Marina Mancoridis; Zoe Hitzig; |
| 352 | Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although it is increasingly adopted, are vibe coding outputs really safe to deploy in production? To answer this question, we propose SUSVIBES, a benchmark consisting of 200 feature-request software engineering tasks from real-world open-source projects, which, when given to human programmers, led to vulnerable implementations. |
Songwen Zhao; Danqing Wang; Kexun Zhang; Jiaxuan Luo; Zhuo Li; Lei Li; |
| 353 | When to Trust The Cheap Check: Weak and Strong Verification for Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce metrics capturing incorrect acceptance, incorrect rejection, and strong-verification frequency. |
Shayan Kiyani; Sima Noorani; George Pappas; Hamed Hassani; |
| 354 | PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce PPT-Eval, a benchmark of 120 diverse PowerPoint tasks across 12 files that cover both content creation and presentation editing scenarios, organized by difficulty. |
Apurva Gandhi; Vishwas Suryanarayanan; Firoz Shaik; Raja Anwar; Shubhang Desai; Thong Nguyen; Muhammad Raza; Vishal Chowdhary; Graham Neubig; |
| 355 | Regulating Anatomy-Aware Rewards Via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Using CABS, we identify a “\textbf{mechanistic divergence}” in standard RL, where surface-similarity rewards drive policy gradients to bypass medical facts. We therefore propose \textbf{Trajectory-Integral Feedback GRPO (TIF-GRPO)}, a novel framework integrating control-theoretic principles into policy optimization. |
Tianwei Lin; Zhongwei Qiu; Jie Cao; Jiang Liu; Wenjie Yan; Bo Zhang; Yu Zhong; Wenqiao Zhang; Yingda Xia; Ling Zhang; |
| 356 | The Latent Color Subspace: Emergent Order in High-Dimensional Chaos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We verify our Latent Color Subspace (LCS) interpretation by demonstrating that it can both predict and explicitly control color, introducing a fully training-free method in FLUX based solely on closed-form latent-space manipulation. |
Mateusz Pach; Jessica Bader; Quentin Bouniot; Serge Belongie; Zeynep Akata; |
| 357 | VR-Thinker: Boosting Multimodal Reward Models Through Think with Image Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, current RMs face inherent limitations: **(1)** visual inputs consume large context budgets, forcing fewer frames and causing a loss of details; and **(2)** all visual information is packed into the initial prompt, exacerbating forgetting during chain-of-thought reasoning. To overcome these issues, we introduce **VR-Thinker**, a thinking-with-image RM equipped with visual reasoning operation and a configurable visual memory window. |
Qunzhong Wang; Jie Liu; Jiajun Liang; Yuanxing Zhang; Yilei Jiang; Yaozhi ZHENG; Xintao Wang; Pengfei Wan; Xiangyu Yue; Jiaheng Liu; |
| 358 | MAS-ProVe: Understanding The Process Verification of Multi-Agent Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Process verification, which evaluates intermediate steps in trajectories, has shown promise in general reasoning settings, and has been suggested as a potential tool for guiding coordination of MAS; however, its actual effectiveness in MAS remains unclear. To fill this gap, we present MAS-ProVe, a systematic empirical study of process verification for multi-agent systems (MAS). |
Vishal Venkataramani; Haizhou Shi; Zixuan Ke; Austin Xu; Xiaoxiao He; Yingbo Zhou; Semih Yavuz; Hao Wang; Shafiq Joty; |
| 359 | Position: The AI Imperative: Scaling High-Quality Peer Review in Machine Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose specific roles for AI in enhancing factual verification, guiding reviewer performance, assisting authors in quality improvement, and supporting ACs in decision-making. |
Qiyao Wei; Samuel Holt; Jing Yang; Markus Wulfmeier; Mihaela van der Schaar; |
| 360 | Learning Latent Action World Models In The Wild Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: World models possess this capability but require action annotations that can be complex to obtain at scale. Latent action models address this issue by learning an action space from videos alone. |
Quentin Garrido; Tushar Nagarajan; Basile Terver; Nicolas Ballas; Yann LeCun; Michael Rabbat; |
| 361 | Conditional Coverage Diagnostics for Conformal Prediction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To overcome sample-inefficiency and overfitting issues of existing metrics, we cast conditional coverage estimation as a classification problem. |
Sacha Braun; David Holzmüller; Michael Jordan; Francis Bach; |
| 362 | InteractComp: Evaluating Search Agents With Ambiguous Queries Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Yet most agents lack interactive mechanisms during the search process, and existing benchmarks cannot assess this capability. To address this gap, we introduce INTERACTCOMP, a benchmark designed to evaluate whether search agents can recognize query ambiguity and actively interact to resolve it during search. |
Mingyi Deng; Lijun Huang; Yani Fan; Fanqi Kong; Jiayi Zhang; Fashen Ren; Jinyi Bai; Fuzhen Yang; Dayi Miao; Zhaoyang Yu; Yifan Wu; Yanfei Zhang; Fengwei Teng; Yingjia Wan; Song Hu; Yude Li; Xin Jin; Conghao Hu; Haoyu Li; Qirui Fu; Tai Zhong; Xinyu Wang; Robert Tang; Nan Tang; Wu; Yuyu Luo; |
| 363 | DnaHNet: A Scalable and Hierarchical Foundation Model for Genomic Sequence Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce dnaHNet, a state-of-the-art tokenizer-free autoregressive model that segments and models genomic sequences end to end. |
Arnav Shah; Junzhe Li; Parsa Idehpour; Adibvafa Fallahpour; Brandon Wang; Sukjun Hwang; BO WANG; Patrick Hsu; Hani Goodarzi; Albert Gu; |
| 364 | An Algebraic View of The Expressivity of Recurrent Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Moreover, many proofs are highly architecture-specific and hard to transfer across closely related models. We address these issues with a unifying algebraic framework for a broad class of RNN language models, formally translating them to wreath products of transformation semigroups. |
Franz Nowak; Reda Boumasmoud; Ryan Cotterell; |
| 365 | A Framework for Understanding Learnability in Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While prior work has established that transformers exhibit a bias toward low-sensitivity functions, the precise mechanism underlying this bias remains poorly understood. To shed light onto this phenomenon, we study the geometry of transformers’ parameter space. |
Blanka Kövér; Alexandra Butoi; Anej Svete; Michael Hahn; Ryan Cotterell; |
| 366 | Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don’t Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We find that, under practical assumptions, padded transformers are surprisingly robust to all of these, and identify numeric precision and model depth as the main factors affecting expressivity. |
Anej Svete; William Merrill; Ryan Cotterell; Ashish Sabharwal; |
| 367 | Privasis: Synthesizing The Largest Public Private Dataset from Scratch Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Research involving privacy-sensitive data has always been constrained by data scarcity, standing in sharp contrast to other areas that have benefited from data scaling. To quench this thirst, we present Privasis (i.e., privacy oasis), the first million-scale fully synthetic dataset entirely built from scratch—an expansive reservoir of texts with rich and diverse private information—designed to broaden and accelerate research in areas where processing sensitive social data is inevitable. |
Hyunwoo Kim; Niloofar Mireshghallah; Michael Duan; Rui Xin; Stella Li; Jaehun Jung; David Acuna; Qi Pang; Hanshen Xiao; Edward Suh; Sewoong Oh; Yulia Tsvetkov; Pang Wei Koh; Yejin Choi; |
| 368 | Reward and Guidance Through Rubrics: Promoting Exploration to Improve Multi-Domain Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite these successes, existing methods mainly focus on single-domain RL (e.g., mathematics) with verifiable rewards (RLVR), and their reliance on purely online RL frameworks restricts the exploration space, thereby limiting reasoning performance. In this paper, we address these limitations by leveraging rubrics to provide both fine-grained reward signals and offline guidance. |
Baolong Bi; Shenghua Liu; Yiwei Wang; Siqian Tong; Lingrui Mei; Yuyao Ge; Yilong Xu; Jiafeng Guo; Xueqi Cheng; |
| 369 | Periodic Bayesian Flow Networks with Additive Accuracy Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce \emph{PeriodicBFN}, which embeds each periodic scalar into a two-dimensional unit-circle representation and performs Gaussian Bayesian updates in the resulting Cartesian space, thereby restoring strictly additive accuracy. |
Peijia Lin; Zihan Zhang; zhangrui zhao; Shaohao Rui; Junyi An; Yun-Fei Shi; Fenglei Cao; Weijie Ma; Yutong Lu; |
| 370 | ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These emerging AI capabilities offer new opportunities for scientific discovery and biomedical advances, but they are also changing the landscape of biosecurity risks. To address this, we introduce the Agentic Bio-Capabilities Benchmark (ABC-Bench), a suite of evaluations to measure \textit{agentic} biosecurity-relevant capabilities. |
Andrew Liu; Samira Nedungadi; Bryce Cai; Alex Kleinman; Harmon Bhasin; Seth Donoughe; |
| 371 | PLANTAIN: Plan-Answer Interleaved Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In contrast, human speakers perform lightweight, incremental check-ins to ensure that conversational participants stay on common ground. With this motivation, we propose \textit{interleaved reasoning} (IR), in which the model alternates between thinking and surfacing intermediate responses, as an alternative to the standard “think-then-answer” approach. |
Anthony Liang; Jonathan Berant; Adam Fisch; Abhimanyu Goyal; Kalpesh Krishna; Jacob Eisenstein; |
| 372 | Evolution Strategies at The Hyperscale Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Evolution Guided GeneRal Optimisation via Low-rank Learning (EGGROLL), which improves arithmetic intensity by structuring individual perturbations as rank-$r$ matrices, resulting in a hundredfold increase in training speed for billion-parameter models at large population sizes, achieving up to 91\% of the throughput of pure batch inference. |
Bidipta Sarkar; Mattie Fellows; Juan Duque; Alistair Letcher; Antonio Villares; Anya Sims; Clarisse Wibault; Dmitry Samsonov; Dylan Cope; Jarek Liesen; Kang Li; Lukas Seier; Theo Wolf; Uljad Berdica; Valentin Mohl; Alexander D. Goldie; Aaron Courville; Karin Sevegnani; Shimon Whiteson; Jakob Foerster; |
| 373 | Efficient and Safe Molecular Assembly Via Reinforcement Learning and Constraint Solving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we take a substantial step toward autonomous manufacturing with STMs by introducing a novel AI-based planning framework for molecular assembly and a high-fidelity simulation environment. |
Stefan Pranger; Bernhard Ramsauer; Oliver Hofmann; Bettina Könighofer; |
| 374 | Predicting What Matters: Robust Generalist Robot Policy Learning Via Future Semantic Mask Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These distractions reduce the model’s ability to generalize, ultimately leading to unreliable and fragile control policies. To address this, we introduce the Mask World Model (MWM), that leverages video diffusion architectures to predict the evolution of semantic masks instead of pixels. |
Yunfan Lou; Xiaowei Chi; Xiaojie Zhang; Zezhong Qian; Chengxuan Li; Rongyu Zhang; yaoxu lyu; Guoyu Song; Chuyao Fu; Haoxuan Xu; Pengwei Wang; Shanghang Zhang; |
| 375 | GraphPFN: A Prior-Data Fitted Network for Graph Node-Level Tasks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we make the next step by proposing GraphPFN, a PFN-based model designed and pretrained specifically for graph node-level tasks. |
Dmitry Eremeev; Oleg Platonov; Gleb Bazhenov; Artem Babenko; Liudmila Prokhorenkova; |
| 376 | Group Distributionally Robust Optimization-Driven RL for LLM Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our approach is principled and theory-driven: we provide no-regret guarantees for the Prompt-GDRO game (via an entropy-regularized GDRO surrogate) and a variance-proxy analysis that yields a square-root optimal compute allocation for Rollout-GDRO. |
Kishan Panaganti; Zhenwen Liang; Wenhao Yu; Haitao Mi; Dong Yu; |
| 377 | Uncovering The Gradient Geometry of Long CoT: A Spectral-guided Approach to Reasoning Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This often leads student models to memorize superficial patterns rather than acquire generalizable reasoning capabilities. To better understand this limitation, we introduce \textit{Loss Subspace Attribution}, a gradient decomposition analysis approach that uncovers a striking geometric structure: Gradients corresponding to effective reasoning predominantly lie within a low-rank consensus subspace, while conflicting or unstructured signals dominate the residual subspace. |
Sinan Fan; Xiaofeng Sun; Chen Shen; Chenxi Huang; Shaotian Yan; Bing Wang; kaiyuan liu; Xiaosong Yuan; Liang Xie; Wenxiao Wang; Jun Zhang; Hongyang Chen; Jieping Ye; |
| 378 | DLO-Lab: Benchmarking Deformable Linear Object Manipulations with Differentiable Physics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Additionally, existing simulation environments offer limited support for the broad spectrum of material behaviors necessary for generalizable DLO manipulation. To overcome these limitations, we introduce a differentiable simulator explicitly designed for versatile DLO manipulation. |
Junyi Cao; Yian Wang; Ziyan Xiong; Chunru Lin; Zhehuan Chen; Chuang Gan; |
| 379 | Learning Taxonomic Trees with Hierarchical Representation Regularization for Large Multimodal Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, we introduce a semantic-aware visual tree construction framework that extracts coarse-to-fine visual features from intermediate LLM layers guided by textual cues. |
Hulingxiao He; Zhi Tan; Yuxin Peng; |
| 380 | Reason with Thumbnails, Answer with Focus: An Efficient and Effective Paradigm for Multimodal Grounded Visual Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To attain efficient and effective GVR, in this paper, we propose a novel paradigm called Reason with Thumbnails, Answer with Focus (RTAF), which feeds the model with low-resolution images to reason the relevant regions and high-resolution crops to answer the final answer. |
An-Lan Wang; Guozhi Tang; Lei Liao; Hanshen Zhu; Kai Huang; Jingqun Tang; Jiaming Zhou; Kun-Yu Lin; |
| 381 | EVOLVING ROLLOUTS: Harnessing Historical Experience for Web Agent Evolution in Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose EVOLVING ROLLOUTS, an RL framework for web-search agents that moves beyond episodic training and distills collected rollouts into in-context guidance for future policy behavior. |
Sinuo Wang; WANG PIAOHONG; Tianrui Qin; Maojia Song; Qianben Chen; Qiexiang Wang; Gengze Zhou; Zeyu Zhang; He Zhu; Dingfeng Shi; Yutong Xie; Minghao Liu; Jiaheng Liu; Ge Zhang; Jiawei Ma; Yuchen Jiang; Qi Wu; Wangchunshu Zhou; |
| 382 | MASPO: Joint Prompt Optimization for LLM-based Multi-Agent Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While the quality of these prompts is pivotal, jointly optimizing them across interacting agents remains a non-trivial challenge, primarily due to the misalignment between local agent objectives and holistic system goals. To address this, we introduce MASPO, a novel framework designed to automatically and iteratively refine prompts across the entire system. |
Zhexuan Wang; Xuebo Liu; Li Wang; Zifei Shan; Yutong Wang; Zhenxi Song; Min zhang; |
| 383 | REAR: Test-time Preference Realignment Through Reward Decomposition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To extend TTS to preference alignment, we introduce a novel framework that models the task as a realignment problem, since the base model often fails to sufficiently align with the stated preference. |
Fuxiang Zhang; Pengcheng Wang; Chenran Li; Yi-Chen Li; Yuxin Chen; Lang Feng; Chenfeng Xu; Masayoshi Tomizuka; Bo An; |
| 384 | MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose Moshi-RAG, a modular approach that combines a compact full-duplex interface with selective retrieval to access more powerful knowledge sources. |
Chung-Ming Chien; Manu Orsini; Eugene Kharitonov; Neil Zeghidour; Karen Livescu; Alexandre Défossez; |
| 385 | Autoregressive Boltzmann Generators Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, modern BGs predominantly rely on Normalizing Flows (NFs), which either suffer from limited expressivity due to strict invertibility constraints (discrete time) or computationally expensive likelihoods (continuous time). In this paper, we propose Autoregressive Boltzmann Generators (ArBG), a novel autoregressive modelling framework that overcomes these limitations by departing from the flow-based BG paradigm. |
Danyal Rehman; Charlie Tan; Yoshua Bengio; Joey Bose; Alexander Tong; |
| 386 | Pair2Scene: Learning Local Object Relations for Procedural Scene Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Building on top of the observation that object placement relies mainly on local dependencies instead of information-redundant global distributions, in this paper, we propose **Pair2Scene**, a novel procedural generation framework, for scene generation based on a set of *learned procedural rules*. |
Xingjian Ran; Shujie Zhang; Weipeng Zhong; Luo Li; Bo Dai; |
| 387 | PhotoAgent: Exploratory Visual Aesthetic Planning with Large Vision Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To achieve autonomous image editing, we present PhotoAgent, a system that advances image editing through explicit aesthetic planning. |
Mingde Yao; Zhiyuan You; Man Tam; Menglu Wang; Tianfan Xue; |
| 388 | Transolver-3: Scaling Up Transformer Solvers to Industrial-Scale Geometries Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Transolver-3, a new member of the Transolver family as a highly scalable framework designed for high-fidelity physics simulations. |
Hang Zhou; Haixu Wu; Haonan Shangguan; Yuezhou Ma; Huikun Weng; Jianmin Wang; Mingsheng Long; |
| 389 | GHOST: Geometry-Guided Hallucination of Opaque Surface Textures Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Transparent objects pose a fundamental challenge for depth estimation and 3D reconstruction due to their violation of Lambertian assumptions, leading to severe geometry degradation in downstream tasks. To address this, we propose a novel geometry-guided preprocessing framework GHOST that leverages visual foundation models to transform transparent regions into opaque, structurally consistent representations without requiring downstream model retraining. |
Langxu Zhao; Zuan Gu; Tianhan Gao; |
| 390 | AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Building on this abstraction, we introduce an agentic system AOrchestra, where the central orchestrator concretizes the tuple at each step: it curates task-relevant context, selects tools and models, and delegates execution via on-the-fly automatic agent creation. |
Jianhao Ruan; Zhihao Xu; Yiran Peng; Fashen Ren; Zhaoyang Yu; Xinbing Liang; Jinyu Xiang; Yongru Chen; Bang Liu; Wu; Yuyu Luo; Jiayi Zhang; |
| 391 | World Guidance: World Modeling in Condition Space for Action Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing approaches struggle to strike a balance between maintaining efficient, predictable future representations and preserving sufficient fine-grained information to guide precise action generation. To address this limitation, we propose WoG (World Guidance), a framework that maps future observations into compact conditions by injecting them into the action inference pipeline. |
Yue Su; Sijin Chen; Haixin Shi; Mingyu Liu; Zhengshen Zhang; Ningyuan Huang; Weiheng Zhong; Zhengbang Zhu; Yuxiao Liu; Xihui Liu; |
| 392 | VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation Via Multimodal Memory Graph Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Traditional Retrieval-augmented Generation (RAG) methods rely on linear interaction histories, which struggle to handle long-context tasks, especially those involving information-sparse yet token-heavy visual data in iterative reasoning scenarios. To bridge this gap, we introduce VimRAG, a framework tailored for multimodal Retrieval-augmented Reasoning across text, images, and videos. |
Qiuchen Wang; Shihang Wang; Yu Zeng; Qiang Zhang; Fanrui Zhang; Zhuoning Guo; Bosi Zhang; Wenxuan Huang; Lin Chen; Zehui Chen; Pengjun Xie; Ruixue Ding; |
| 393 | LUVE : Latent-Cascaded Ultra-High-Resolution Video Generation with Dual Frequency Experts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent advances in video diffusion models have significantly improved visual quality, yet ultra-high-resolution (UHR) video generation remains a formidable challenge due to the compounded difficulties of motion modeling, semantic planning, and detail synthesis. To address these limitations, we propose \textbf{LUVE}, a \textbf{L}atent-cascaded \textbf{U}HR \textbf{V}ideo generation framework built upon dual frequency \textbf{E}xperts. |
Chen Zhao; Jiawei Chen; Hongyu Li; Zhuoliang Kang; Shilin Lu; Xiaoming Wei; Kai Zhang; Jian Yang; Ying Tai; |
| 394 | Test-Time Anchoring for Discrete Diffusion Posterior Sampling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing approaches to posterior sampling using discrete diffusion face severe challenges: derivative-free guidance yields sparse signals, continuous relaxations limit applicability, and split Gibbs samplers suffer from the curse of dimensionality. To overcome these limitations, we introduce Anchored Posterior Sampling (APS), built on two key innovations: *quantized expectation* for gradient-like guidance in discrete embedding space, and *anchored remasking* for adaptive decoding. |
Litu Rout; Andreas Lugmayr; Yasamin Jafarian; Srivatsan Varadharajan; Constantine Caramanis; Sanjay Shakkottai; Ira Kemelmacher-Shlizerman; |
| 395 | Stratified GRPO: Handling Structural Heterogeneity in Reinforcement Learning of LLM Search Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As a result, policy gradient methods with a single global baseline suffer from *cross-stratum bias*, an apples-to-oranges comparison that distorts credit assignment and impedes exploration. To address this issue, we propose *Stratified GRPO*. |
Mingkang Zhu; Xi Chen; Bei Yu; Hengshuang Zhao; Jiaya Jia; |
| 396 | Identifiable Token Correspondence for World Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We formulate next-frame prediction as a structured probabilistic inference problem with latent token correspondence variables, deriving a model in which each next-frame token is explained either by copying a token from the previous frame or by generating a new token. |
Youngin Kim; Ray Sun; Inho Kim; Bumsoo Park; Hyun Oh Song; |
| 397 | DropoutTS: Sample-Adaptive Dropout for Robust Time Series Forecasting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce DropoutTS, a model-agnostic plugin that shifts the paradigm from what to learn to how much to learn. |
Siru Zhong; Yiqiu Liu; Zhiqing Cui; Zezhi Shao; Fei Wang; Qingsong Wen; Yuxuan Liang; |
| 398 | Position: Use Sparse Autoencoders to Discover Unknowns Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Here, we establish a conceptual distinction that reconciles competing narratives surrounding SAEs. |
Kenny Peng; Rajiv Movva; Jon Kleinberg; Emma Pierson; Nikhil Garg; |
| 399 | Code2Video: A Code-centric Paradigm for Educational Video Creation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose **Code2Video**, a code-centric agent framework that generates educational videos by writing executable Python programs. |
Yanzhe Chen; Kevin Qinghong Lin; Mike Zheng Shou; |
| 400 | Escaping The Diversity Trap in Robotic Manipulation Via Anchor-Centric Adaptation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we identify a critical **diversity trap**: the standard heuristic of “maximizing coverage by collecting diverse, single-shot demonstrations can be self-defeating due to non-vanishing estimation noise. |
Yanzhe Chen; Kevin Yuchen; Qi Lv; Lin Yiqi; Zechen Bai; Chen Gao; Mike Zheng Shou; |
| 401 | Reinforcement Learning Via Self-Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Many verifiable environments actually provide rich textual feedback, such as runtime errors or judge evaluations, that explain *why* an attempt failed. We formalize this setting as reinforcement learning with rich feedback and introduce **Self-Distillation Policy Optimization** (**SDPO**), which converts tokenized feedback into a dense learning signal without any external teacher or explicit reward model. |
Jonas Hübotter; Frederike Lübeck; Lejs Behric; Anton Baumann; Marco Bagatella; Daniel Marta; Ido Hakimi; Idan Shenfeld; Thomas Kleine Buening; Carlos Guestrin; Andreas Krause; |
| 402 | VividCam: Learning Unconventional Camera Motions from Virtual Synthetic Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The challenge lies in the difficulty of finding sufficient training videos with the intended uncommon camera motions. To address this challenge, we propose VividCam, a training paradigm that enables diffusion models to learn complex camera motions from synthetic videos, releasing the reliance on collecting realistic training videos. |
Qiucheng Wu; Handong Zhao; Zhixin Shu; Jing Shi; Yang Zhang; Shiyu Chang; |
| 403 | FIRE-Bench: Evaluating Agents on The Rediscovery of Scientific Insights Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing benchmarks face a trade-off: they either rely on LLM-as-judge evaluations of automatically generated papers, or optimize isolated performance metrics that provide only coarse proxies for scientific insight. To address this, we introduce FIRE-Bench (Full-cycle Insight Rediscovery Evaluation), a benchmark that evaluates agents through the rediscovery of established findings from recent, high-impact machine learning research. |
Zhen Wang; Fan Bai; Zhongyan Luo; Jinyan Su; Kaiser Sun; Xinle Yu; Jieyuan Liu; Kun Zhou; Claire Cardie; Mark Dredze; Eric Xing; Zhiting Hu; |
| 404 | Attentive Multi-Layer Fusion for Vision Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We show that task-relevant information is distributed across the network hierarchy rather than solely encoded in any of the last layers. To leverage this distribution of information, we apply an attentive probing mechanism that dynamically fuses representations from all layers of a Vision Transformer. |
Laure Ciernik; Marco Morik; Lukas Thede; Luca Eyring; Shinichi Nakajima; Zeynep Akata; Lukas Muttenthaler; |
| 405 | Position: Safe AI Should Be Resistant and Resilient in An Evolving World Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this position paper, we address the persistent gap between rapidly growing AI capabilities and lagging safety progress. |
Youbang Sun; Xiang Wang; Jie Fu; Chaochao Lu; Bowen Zhou; |
| 406 | SPADA: A Verifiable Test-Driven Agent for Controllable Parametric CAD Assembly Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Methods that generate meshes or history-free B-rep can represent multi-part shape, but they often lack the program structure and constraint logic needed for reliable downstream edits; in contrast, code-based CAD generation offers direct parametric control, yet most published settings and evaluations focus on single-part solids rather than constrained assemblies. We introduce SPADA (Self-testing Parametric Assembly Design Agent), a test-driven agent that synthesizes assembly code together with deterministic verification tests, and uses these tests as an executable contract for controllable generation. |
Keyou Zheng; Xuyang Su; Jiewu Leng; |
| 407 | MARS: Modular Agent with Reflective Search for Automated AI Research Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce **MARS** (**M**odular **A**gent with **R**eflective **S**earch), a framework optimized for autonomous AI research. |
Jiefeng Chen; Bhavana Dalvi Mishra; Jaehyun Nam; Rui Meng; Tomas Pfister; Jinsung Yoon; |
| 408 | Beyond Soft Labels: Unifying Dataset Pruning and Distillation for Efficient Large-scale Compression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recently, DD’s increasing reliance on original images suggests a convergence of the two directions. To investigate this convergence trend, we propose a unified dataset compression (DC) benchmark. |
Lingao Xiao; Songhua Liu; Yang He; Xinchao Wang; |
| 409 | Anatomy of Massive Activations and Attention Sinks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Second, \emph{attention sinks}, where certain tokens attract a disproportionate share of attention across many heads and layers. We present a unified inference-time mechanism explaining how massive activations emerge and propagate through layers, and how normalization transforms these tokens into sparse, nearly fixed vectors that reshape the attention space and induce sink or no-sink behavior. |
Shangwen Sun; Alfredo Canziani; Yann LeCun; Jiachen Zhu; |
| 410 | PhysForge: Generating Physics-Grounded 3D Assets for Interactive Virtual World Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose that interactive asset generation must be rooted in functional logic and hierarchical physics. To bridge this gap, we introduce PhysForge, a decoupled two-stage framework supported by PhysDB, a large-scale dataset of 150,000 assets with four-tier physical annotations. |
Yunhan Yang; Chunshi Wang; Junliang Ye; YANG LI; Zanxin Chen; Zehuan Huang; Yao Mu; Zhuo Chen; Chunchao Guo; Xihui Liu; |
| 411 | The Information Geometry of Softmax: Probing and Steering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper concerns the question of how language models and other AI systems encode semantic structure into the geometric structure of their representation spaces. |
Kiho Park; Todd Nief; Yo Joong Choe; Victor Veitch; |
| 412 | PhoStream: Benchmarking Real-World Streaming for Omnimodal Assistants in Mobile Scenarios Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce **PhoStream**, the first mobile-centric streaming benchmark that unifies on-screen and off-screen scenarios to evaluate video, audio, and temporal reasoning. |
Xudong LU; Guan Huankang; Yang Bo; Jinpeng Chen; Xintong Guo; Shuhan LI; Fang Liu; Peiwen Sun; Xueying Lee; Wei Zhang; Xue Yang; Rui Liu; Hongsheng Li; |
| 413 | ToMAP: Training Opponent-Aware LLM Persuaders with Theory of Mind Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: Notably, while humans are skilled in modeling their opponent’s thoughts and opinions proactively and dynamically, current LLMs struggle with such Theory of Mind (ToM) reasoning, resulting in limited diversity and opponent awareness. To address this limitation, we introduce Theory of Mind Augmented Persuader (**ToMAP**), a novel approach for building more flexible persuader agents by incorporating two theory of mind modules that enhance the persuader’s awareness and analysis of the opponent’s mental state. |
Peixuan Han; Zijia Liu; Jiaxuan You; |
| 414 | When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we analyze the per-sample gradient of BT-loss and shows spurious learning signals due to representation distance. |
Tong Xie; Ching-Yuan Bai; Yuanhao Ban; Yunqi Hong; Haoyu Li; Cho-Jui Hsieh; |
| 415 | Approximation of Log-Partition Function in Policy Mirror Descent Induces Implicit Regularization for LLM Post-Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We investigate a practical algorithm, termed PMD-mean, that approximates the log-partition term with the mean reward under the sampling policy and performs regression in log-policy space. |
Zhenghao Xu; Qin Lu; Changlong Yu; Tuo Zhao; |
| 416 | Provably Label-Efficient Conformal Prediction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we study *conformal prediction with costly label queries*, where unlabeled examples arrive i.i.d. and labels can be queried one at a time. |
Andrew Ilyas; Joonhyuk Ko; Jingwu Tang; Steven Wu; Jiahao Zhang; |
| 417 | Guideline-Grounded Evidence Accumulation for High-Stakes Agent Verification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Yet, existing verifiers usually underperform owing to a lack of domain knowledge and limited calibration. To address this, we establish **GLEAN**, an agent verification framework with **G**uide**L**ine-grounded **E**vidence **A**ccumulatio**N** that compiles expert-curated protocols into trajectory-informed, well-calibrated correctness signals. |
Yichi Zhang; Nabeel Seedat; Yinpeng Dong; Peng Cui; Jun Zhu; Mihaela van der Schaar; |
| 418 | FeRA: Frequency-Energy Constrained Routing for Effective Diffusion Adaptation Fine-Tuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We revisit the reconstruction behavior of diffusion models during denoising to unveil the underlying frequency–energy mechanism governing this process. Building upon this observation, we propose \textbf{FeRA}, a frequency-driven fine-tuning framework that aligns parameter updates with the intrinsic frequency–energy progression of diffusion. |
Bo Yin; Xiaobin Hu; Xingyu Zhou; Yu HE; Peng-Tao Jiang; Yue Liao; Junwei Zhu; Jiangning Zhang; Ying Tai; Shuicheng YAN; |
| 419 | Position: Behavioral Systems Require Behavioral Tests Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper argues that AI agents must be evaluated like other behavioral systems: through systematic observation, perturbation, and interpretation of their actions. We draw on lessons from the behavioral sciences to motivate this position, and propose a research agenda focused on developing rigorous behavioral tests. |
Manuel Cherep; Nikhil Singh; Pattie Maes; |
| 420 | AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Further, they operate in small environments with few, often non-state-changing, APIs. To address this gap, we introduce AppWorld-UL, a “user-in-the-loop” benchmark of 306 challenging tasks requiring diverse agent-user interactions. |
Junzhi Chen; Harsh Trivedi; Jane Pan; Michael Zhang; Tejas Srinivasan; Niranjan Balasubramanian; Ashish Sabharwal; |
| 421 | URS: A Unified Neural Routing Solver for Cross-Problem Zero-Shot Generalization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing neural solvers typically rely on predefined problem constraints or require per-problem fine-tuning, which substantially limits their zero-shot generalization ability to unseen VRP variants. To address this critical bottleneck, we propose URS, a unified neural routing solver that achieves zero-shot generalization across a wide range of unseen VRPs with a single model. |
Changliang Zhou; Canhong Yu; Shunyu Yao; Xi Lin; Zhenkun Wang; Yu Zhou; Qingfu Zhang; |
| 422 | PGD-NO: A Neural Operator with Precomputed Geometry Decomposition for 3D Million-Scale Physics Simulations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While neural PDE solvers have demonstrated significant potential for accelerating engineering simulations, existing architectures remain constrained by high memory consumption and the single-node bottleneck, where the maximum processable mesh resolution is strictly limited by the VRAM of a single compute unit. To address these challenges, we propose PGD-NO, a neural operator with Precomputed Geometry Decomposition, that relocates the computational overhead of geometric encoding to a deterministic pre-computation phase. |
Weiheng Zhong; Jing Bi; Victor Oancea; Hadi Meidani; |
| 423 | UniCoD: Enhancing Robot Policy Via Unified Continuous and Discrete Representation Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We posit that robotic policy learning can likewise benefit from the combined strengths of understanding, planning and continuous future representation learning. Building on this insight, we introduce UniCoD, which acquires the ability to dynamically model high-dimensional visual features through pretraining on over 1M internet-scale instructional manipulation videos. |
Jianke Zhang; Yucheng Hu; Yanjiang Guo; Xiaoyu Chen; Yichen Liu; Wenna Chen; Chaochao Lu; Jianyu Chen; |
| 424 | Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We show that, after identical RL training, models initialized from stronger SFT checkpoints can significantly underperform those initialized from weaker ones. We propose PEAR ($\textbf{P}$olicy $\textbf{E}$valuation–inspired $\textbf{A}$lgorithm for Offline Learning Loss $\textbf{R}$eweighting), an SFT-stage method that corrects this mismatch and better prepares the model for RL. |
Dylan Zhang; Yufeng Xu; Haojin Wang; Qingzhi Chen; Hao Peng; |
| 425 | How Can We Assess Human-agent Interactions? Case Studies in Software Agent Design Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we make two major steps towards the rigorous assessment of human-agent interactions. |
Valerie Chen; Rohit Malhotra; Xingyao Wang; Juan Michelini; Xuhui Zhou; Aditya Bharat Soni; Hoang Tran; Calvin Smith; Ameet Talwalkar; Graham Neubig; |
| 426 | Edit-Based Refinement for Parallel Masked Diffusion Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose ME-DLM, an edit-based refinement framework that augments diffusion generation with a lightweight post-generation editing step. |
Houxing Ren; Mingjie Zhan; Zimu Lu; Ke Wang; Yunqiao Yang; Haotian Hou; Junting Pan; Hongsheng Li; |
| 427 | Native Spatio-Temporal 4D Variational Autoencoder Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce a novel 4D VAE that operates directly in native 4D space, that is dynamic colored voxel space, without 2D projection. |
Lihe Ding; Weicai Ye; Shaocong Dong; Xintao Wang; Pengfei Wan; Kun Gai; Tianfan Xue; |
| 428 | TeamWork: Multivariate Time Series Anomaly Detection Via Asymmetric Role-aware Channel Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods often struggle to balance channel relationship modeling and overlook the relative importance of different variables within multivariate time series. To address this, we propose TeamWork, an asymmetric role-aware channel modeling framework that decouples variables into dominant and auxiliary roles according to their contributions to uncertainty reduction. |
Shiyan Hu; Tengxue Zhang; Jianxin Jin; Xiangfei Qiu; Bin Yang; Chenjuan Guo; |
| 429 | MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing multi-turn benchmarks are limited in size or rely heavily on templates, which restrict their diversity. To address this gap, we unify a wide range of harmful jailbreak intents, and introduce an active learning pipeline for expanding high-quality multi-turn adversarial prompts, where a generator is iteratively fine-tuned to produce stronger attack candidates, guided by uncertainty-based refinement. |
Jialin Song; Xiaodong Liu; Weiwei Yang; Wuyang Chen; Mingqian Feng; Xuekai Zhu; Jianfeng Gao; |
| 430 | Think-Then-Generate: Reasoning-Aware Text-to-Image Diffusion with LLM Encoders Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Yet, most existing T2I DMs, even those equipped with large language model (LLM)-based text encoders, remain text-pixel mappers — they employ LLMs merely as text encoders, without leveraging their inherent reasoning capabilities to infer what should be visually depicted given the textual prompt. To move beyond such literal generation, we propose the think-then-generate (T2G) paradigm, where the LLM-based text encoder is encouraged to reason about and rewrite raw user prompts; the states of the rewritten prompts then serve as diffusion conditioning. |
Siqi Kou; Jiachun Jin; Zetong Zhou; Ye Ma; Yugang Wang; Quan Chen; Peng Jiang; Xiao Yang; Jun Zhu; Kai Yu; Zhijie Deng; |
| 431 | Learning Query-Aware Budget-Tier Routing for Runtime Agent Memory Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present \textbf{BudgetMem}, a runtime agent memory framework for explicit, query-aware performance–cost control. |
Haozhen Zhang; Haodong Yue; Tao Feng; Quanyu Long; Jianzhu Bao; Bowen Jin; Weizhi Zhang; Xiao Li; Jiaxuan You; Chengwei Qin; Wenya Wang; |
| 432 | Towards Unified Multimodal Pretraining Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we explore the design space of Unified Multimodal Pretraining through a controlled, from-scratch study. |
Shengbang Tong; David Fan; John Nguyen; Ellis Brown; Gaoyue Zhou; Shengyi Qian; Boyang Zheng; Théophane Vallaeys; Rob Fergus; Naila Murray; Marjan Ghazvininejad; Mike Lewis; Jakob Verbeek; Nicolas Ballas; Amir Bar; Michael Rabbat; Yann LeCun; Luke Zettlemoyer; Saining Xie; Koustuv Sinha; |
| 433 | Generative Online Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We address this tension with a key structural principle: decoupling optimization from generation. Building on this, we introduce GoRL (Generative Online Reinforcement Learning), an algorithm-agnostic framework that trains expressive policies from scratch by confining policy optimization to a tractable latent space while delegating action synthesis to a conditional generative decoder. |
Chubin Zhang; Zhenglin Wan; Feng Chen; Fuchao Yang; Lang Feng; Yaxin Zhou; Xingrui Yu; Yang You; Ivor Tsang; Bo An; |
| 434 | AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To further assess robustness beyond familiar domains, we propose AVI-Bench-PriSe, an extension that probes models’ primitive audio-visual sensation using unfamiliar, low-semantic stimuli, testing generalization beyond common training distributions. |
Yaoting Wang; Ziyi Zhang; Wenming Tu; Shaoxuan Xu; Wenjie Du; Cheng Liang; weijun wang; Yuanchao Li; Guangyao Li; Hao Fei; Yuanchun Li; Henghui Ding; Yunxin Liu; |
| 435 | Utonia: Toward One Encoder for All Point Clouds Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We dream of a future where point clouds from all domains can come together to shape a single model that benefits them all. Toward this goal, we present Utonia, a first step toward training a single self-supervised point transformer encoder across heterogeneous domains, spanning remote sensing, outdoor LiDAR, indoor RGB-D sequences, object-centric CAD models, and point clouds lifted from RGB-only videos. |
Yujia Zhang; Xiaoyang Wu; Yunhan Yang; Xianzhe Fan; Han Li; Yuechen Zhang; Zehao Huang; Naiyan Wang; Hengshuang Zhao; |
| 436 | Deep Residual Injection for Full-Spectrum Forensic Signal Perception in Multimodal Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: * We then conduct a layer-wise analysis of forensic signal perception in MLLMs and find that semantic information is mainly encoded in the early-to-middle layers, and directly fine-tuning MLLMs for artifact learning causes rapid semantic forgetting. Based on this insight, we propose Deep Visual Residual MLLM (Deep-VRM) to \textit{preserve early semantic processing while injecting artifact-specific visual signals as a residual path into an intermediate layer}, where they are fused with semantic token representations and propagated through subsequent trainable layers. |
Kaiqing Lin; Zhiyuan Yan; Ruoxin Chen; Ke-Yue Zhang; Yue Zhou; Caiyong Piao; Bin Li; Taiping Yao; Bo Wang; Youchang xiao; Shouhong Ding; |
| 437 | $V_0$: A Generalist Value Model for Any Policy at State Zero Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose $V_0$, a generalist value model that decouples value estimation from specific policy parameters by reframing the task as in-context learning to predict performance for unseen policies. |
Yi-Kai Zhang; Zhiyuan Yao; Hongyan Hao; Yueqing Sun; Qi GU; Hui Su; Xunliang Cai; De-Chuan Zhan; Han-Jia Ye; |
| 438 | WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose WeDLM, a diffusion decoding framework built entirely on standard causal attention to make parallel generation prefix-cache friendly. |
Aiwei Liu; Minghua He; Shaoxun Zeng; Sijun Zhang; Linhao Zhang; Chuhan Wu; Wei Jia; Yuan Liu; Zhou Xiao; Jie Zhou; |
| 439 | $\text{DT}^\text{2}$: Decision-Targeted Digital Twins Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We further show that this holds empirically, even with expressive model classes. To address this, we introduce DT$^2$, a decision-targeted DT training paradigm. |
Harry Amad; Mihaela van der Schaar; |
| 440 | MixReasoning: Switching Modes to Think Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose MixReasoning, a framework that dynamically adjusts the depth of reasoning within a single response. |
Haiquan Lu; Gongfan Fang; Xinyin Ma; Qi Li; Xinchao Wang; |
| 441 | On The Relationship Between Activation Outliers and Feature Death in Sparse Autoencoders Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We find that dimension-level activation outliers (dimensions where mean magnitude is large relative to per-token variation) shift pre-activations at initialization, making feature fate depend on weight-outlier alignment rather than input content. |
Elana Simon; Etowah Adams; James Zou; |
| 442 | Structure-aware Granular-Ball Based Information Bottleneck for Multi-modal Clustering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing methods relying on single-granularity relationships often struggle with complex data distributions, leading to limited performance, as fine-grained features are prone to local heterogeneity and redundant perturbations while coarse-grained representations tend to lose local structural information. To address these limitations, we introduce granular-balls (GBs), adaptive multi-granularity hyperspheres that enclose similar samples, and propose the Structure-aware Granular-Ball based Information Bottleneck (SGB-IB) algorithm. |
Zhengzheng Lou; Yuhan Zhan; Mingyang Lv; Yingxuan Li; Yuyang Du; Shizhe Hu; |
| 443 | RSPO: Regularized Self-Play Alignment of Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To study the impact of different regularization strategies, we propose \textbf{Regularized Self-Play Policy Optimization (RSPO)}, a novel framework that unifies prior methods and enables simple plug-and-play regularizers, meanwhile preserving convergence to Nash equilibrium of the corresponding regularized game. |
Xiaohang Tang; Sangwoong Yoon; Seongho Son; Huizhuo Yuan; Quanquan Gu; Ilija Bogunovic; |
| 444 | VLANeXt: Recipes for Building Strong VLA Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: From this study, we distill 12 key findings that together form a practical recipe for building strong VLA models. |
Xiao-Ming Wu; Bin Fan; Kang Liao; Jian-Jian Jiang; Runze Yang; Yihang Luo; Zhonghua Wu; Wei-Shi Zheng; Chen Change Loy; |
| 445 | Cold-Start Personalization Via Training-Free Priors from Structured World Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our contribution is a principled decomposition of cold-start personalization that makes Bayesian preference elicitation practical at scale for LLM systems. |
Avinandan Bose; Stella Li; Faeze Brahman; Pang Wei Koh; Simon Du; Yulia Tsvetkov; Maryam Fazel; Lin Xiao; Asli Celikyilmaz; |
| 446 | AdverMCTS: Combating Pseudo-Correctness in Code Generation Via Adversarial Monte Carlo Tree Search Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We argue that optimizing against a fixed, weak environment inherently limits robustness. To address this, we propose AdverMCTS, a novel adversarial Monte Carlo Tree Search framework that combats pseudo-correctness by coupling code search with active vulnerability discovery. |
Qingyao Li; Weiwen Liu; Weinan Zhang; Yong Yu; Bo An; |
| 447 | PixCLIP: Towards Fine-grained Vision-Language Understanding Via Any-granularity Pixel-Text Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent efforts improve textual granularity by leveraging long, detailed descriptions and replacing CLIP’s text encoder with LLM, but often overlook the visual-side bottleneck: achieving finer alignment requires region- and pixel-level visual grounding, not just finer text. To address this issue, we propose PixCLIP, a framework that jointly enhances both sides by accommodating visual prompt regions and long-form text within a unified training objective. |
YiCheng Xiao; Yu Chen; Hao-Xuan Ma; Jiale Hong; Caorui Li; Lingxiang Wu; Haiyun Guo; Jinqiao Wang; |
| 448 | Nonparametric Data Attribution for Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a *nonparametric* attribution method that operates entirely on data, measuring influence via patch-level similarity between generated and training images. |
Yutian Zhao; Chao Du; Xiaosen Zheng; Tianyu Pang; Min Lin; |
| 449 | ViTok-v2: Scaling Native-Resolution Autoencoders to 5B Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Vision Transformer (ViT) tokenizers offer a scal- able alternative to convolutional auto-encoders, yet current architectures have two key limitations: their performance degrades when images vary in aspect ratio or resolution, and their reliance on adversarial losses makes them harder to train at scale. To address this, we introduce ViTok-v2, a ViT tokenizer building on ViTok. |
Philippe Hansen-Estruch; Jiahui Chen; Vivek Ramanujan; Orr Zohar; Markos Georgopoulos; Animesh Sinha; Ji Hou; Edgar Schönfeld; Felix Juefei-Xu; Sriram Vishwanath; Ali Thabet; |
| 450 | Don’t Overthink with Pixels: Efficient Reasoning for Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While recent efforts leverage reinforcement fine-tuning to further enhance reasoning ability, they often suffer from overthinking and produce uniformly verbose reasoning chains irrespective of task complexity. To address this problem, we propose PixelThink, a simple yet effective scheme that integrates externally estimated task difficulty and internally measured model uncertainty to regulate reasoning generation within a reinforcement learning paradigm. |
Song Wang; Gongfan Fang; Lingdong Kong; Xiangtai Li; Jianyun Xu; Sheng Yang; Qiang Li; Jianke Zhu; Xinchao Wang; |
| 451 | Efficient Reasoning with Hidden Thinking Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: In this work, we propose**Heima** (as hidden llama), an effective CoT compression framework that condenses lengthy CoTs into a small set of abstract thinking tokens, preserving essential reasoning while removing redundancy. |
Xuan Shen; Yizhou Wang; Yufa Zhou; Xiangxi Shi; Pu Zhao; Yanzhi Wang; Jiuxiang Gu; |
| 452 | Brep2Shape: Boundary and Shape Representation Alignment Via Self-supervised Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While deep learning shows promise in processing B-rep models, existing methods suffer from a representation gap: continuous approaches offer analytical precision but are visually abstract, whereas discrete methods provide intuitive clarity at the expense of geometric precision. To bridge this gap, we introduce Brep2Shape, a novel self-supervised pre-training framework designed to align abstract boundary representations with intuitive shape representations. |
Yuanxu Sun; Yuezhou Ma; Haixu Wu; Guanyang Zeng; Muye Chen; Jianmin Wang; Mingsheng Long; |
| 453 | No More, No Less: Least-Privilege Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We take inspiration from least privilege in computer systems and define a class of models called *least-privilege language models*, where privilege is *reachable internal computation* during the forward pass. |
Paulius Rauba; Dominykas Seputis; Patrikas Vanagas; Mihaela van der Schaar; |
| 454 | VIRUS: Injecting Persistent Cognitive Pathogens Into Stateful Zero-Shot Object Navigation Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This structural vulnerability allows injected adversarial information to transmute into long-term memory, persistently disrupting subsequent planning behaviors. Exploiting this, we propose the Visual-Instruction Recurrent Update Subversion (VIRUS) framework, the first training-free backdoor attack scheme specifically targeting the state update stage of ZSON agents. |
Xusheng Lin; Hao Zheng; Xiaojun Jia; Qiucen Li; Tianqi Shan; Jie Xu; Wenqi Ren; |
| 455 | Prompt Injection As Role Confusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: More broadly, we introduce a unifying, mechanistic framework for prompt injection, demonstrating that diverse prompt-injection attacks exploit the same underlying role-confusion mechanism. |
Charles Ye; Jasmine Cui; Dylan Hadfield-Menell; |
| 456 | GameDevBench: Evaluating Agentic Capabilities Through Game Development Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present GameDevBench, the first benchmark for evaluating agents on game development tasks, consisting of 168 tasks derived from web and video tutorials. |
Wayne Chi; Yixiong Fang; Arnav Yayavaram; Siddharth Yayavaram; Seth Karten; Qiuhong Anna Wei; Runkun Chen; Alexander Wang; Valerie Chen; Ameet Talwalkar; Chris Donahue; |
| 457 | Spatially-Regularized Entropy for Discriminative Token Merging in Fine-Grained Re-Identification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In fine-grained retrieval, these approaches often discard or smooth out subtle but discriminative local details. To resolve this, we propose SRE-Merge, a training-free framework designed for discriminative token compression. |
Shangze Li; Yifan Xu; Jingmiao Liang; Yongfei Zhang; Yuzhuo Ma; Yingbo Qu; |
| 458 | OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We identify this “modality interference” as a consequence of pre-training data imbalances, where the scarcity of mixed modality supervision induces a bias towards isolated modalities, resulting in an inherent trade-off. To address this challenge, we propose OmniVideo-R1, a novel reinforced reasoning framework that leverages post-training to rectify modality bias. |
Zhangquan Chen; Jiale Tao; Ruihuang Li; Yihao Hu; Ruitao Chen; Zhantao Yang; Xinlei Yu; Haodong Jing; Manyuan Zhang; Shuai Shao; Biao Wang; Qinglin Lu; Ruqi Huang; |
| 459 | Pareto-Guided Optimal Transport for Multi-Reward Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We demonstrate that optimizing against a unified global target under heterogeneous reward upper bounds can induce reward hacking, a risk further exacerbated by the inherent instability of weak reward models. To mitigate this, we propose a Pareto Frontier-Guided Optimal Transport framework. |
Ying Ba; Tianyu Zhang; Mohan Zhou; Yalong Bai; Wenyi Mo; Guiwei Zhang; Bing Su; Ji-Rong Wen; |
| 460 | From Interactions to Principles: Experience-Driven Self-Distillation for Evolving LLM Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Prior work either stores raw trajectories for case-based reuse or relies on external teacher models to write reflections, which limits generalization or leaves the agent’s policy unchanged. We introduce EvolveR, an experience-driven framework that allows an agent to improve using its own interaction history. |
Rong Wu; Xiaoman Wang; Jianbiao Mei; Pinlong Cai; Daocheng Fu; Cheng Yang; Licheng Wen; Xuemeng Yang; Yufan Shen; Yuxin Wang; Botian Shi; |
| 461 | SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose \textbf{SkillTrojan}, a backdoor attack that targets skill implementations rather than model parameters or training data. |
Yunhao Feng; Yifan Ding; Yingshui Tan; Boren Zheng; Yanming Guo; Xiaolong Li; Kun Zhai; Yishan Li; Wenke Huang; |
| 462 | Who Transfers Safety? Identifying and Targeting Cross-Lingual Shared Safety Neurons Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we find that LLMs contain a set of cross-lingual shared safety neurons (SS-Neurons), a remarkably small yet critical neuronal subset that jointly regulates safety behavior across languages. |
Xianhui Zhang; Chengyu Xie; Linxia Zhu; Yonghui Yang; Weixiang Zhao; Zifeng Cheng; Cong Wang; Fei Shen; Tat-Seng Chua; |
| 463 | Towards A Science of AI Agent Reliability Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a framework for measuring agent reliability grounded in safety-critical engineering practice, decomposing reliability into four dimensions: consistency, robustness, predictability, and safety. |
Stephan Rabanser; Sayash Kapoor; Peter Kirgis; Kangheng Liu; Saiteja Utpala; Arvind Narayanan; |
| 464 | MLLM-4D: Towards Visual-based Spatial-Temporal Intelligence Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite its importance, this capability remains a significant bottleneck for current multimodal large language models (MLLMs). To tackle this challenge, we introduce MLLM-4D, a comprehensive framework designed to bridge the gaps in training data curation and model post-training for spatiotemporal understanding and reasoning. |
Xingyilang Yin; Chengzhengxu Li; Jiahao Chang; Chi-Man Pun; Xiaodong Cun; |
| 465 | GSFixer: Improving 3D Gaussian Splatting with Reference-Guided Video Diffusion Priors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While recent approaches have sought to leverage generative priors to complete information for under-constrained regions, they struggle to generate content that remains consistent with input observations. To address this challenge, we propose GSFixer, a novel framework designed to improve the quality of 3DGS representations reconstructed from sparse inputs. |
Xingyilang Yin; Qi Zhang; Jiahao Chang; Ying Feng; Qingnan Fan; Xi Yang; Chi-Man Pun; Huaqi Zhang; Xiaodong Cun; |
| 466 | Symmetries in Language Statistics Shape The Geometry of Model Representations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We show that the statistics of language exhibit a translation symmetry—e.g,. |
Dhruva Karkada; Daniel Korchinski; Andres Nava; Matthieu Wyart; Yasaman Bahri; |
| 467 | Model-Preserving Adaptive Rounding Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: In this work, we introduce Yet Another Quantization Algorithm (YAQA), a new adaptive rounding algorithm that directly considers the error at the network’s output. |
Albert Tseng; Zhaofeng Sun; Chris De Sa; |
| 468 | $L^3$: Large Lookup Layers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce the Large Lookup Layer (L3), which unlocks a new axis of sparsity by generalizing embedding tables to model decoder layers. |
Albert Tseng; Chris De Sa; |
| 469 | Gradient Inversion Attacks Beyond SGD Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present an analytical rule for recovering labels from optimizer updates and propose an update-matching objective that optimizes dummy inputs to reproduce the observed updates. |
Guangnian Wan; Gongfan Fang; Xinyin Ma; Xinchao Wang; |
| 470 | Learning, Solving and Optimizing PDEs with TensorGalerkin: An Efficient High-performance Galerkin Assembly Algorithm Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a unified algorithmic framework for the numerical solution, constrained optimization, and physics-informed learning of PDEs with a variational structure. |
Shizheng Wen; Mingyuan chi; Tianwei Yu; Ben Moseley; Mike Yan Michelis; Pu Ren; Hao Sun; Siddhartha Mishra; |
| 471 | A Pure Hierarchical Spectral Parcellation Network for Brain Network Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, this model is constructed as a hierarchy of Spectral Parcellation blocks. |
Jiaming Zhuo; Shuai Zhai; Ziyi Ma; Kun Fu; Chuan Wang; Di Jin; Zhen Wang; Xiaochun Cao; Huazhu Fu; Liang Yang; |
| 472 | Improving Graph Transformers Via Global Structural Priors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To bridge this research gap, this paper unifies these designs under a common Graph Signal Denoising framework, revealing that denoising efficacy (\textit{i.e.}, representation quality) is fundamentally dictated by the block-diagonal structure of the propagation operator. To instantiate this prior efficiently, this paper introduces a novel Block-Diagonal GT architecture, named \textsc{BDFormer}, which enforces a block-diagonal constraint via spectral-regularized cross-attention on latent anchors. |
Jiaming Zhuo; Ziyi Ma; Kun Fu; Di Jin; Chuan Wang; Zhen Wang; Xiaochun Cao; Huazhu Fu; Liang Yang; |
| 473 | Breaking Multi-Task Curse: Reward-Weighted Evolution for Black-Box Many-Task Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We identify this scalability barrier as the *Multi-Task Curse*, driven by evaluation budget dispersion and negative transfer. To overcome this, we propose MES-RET (*M*any-task *E*volution *S*trategy with *R*eward-weighted *E*valuation and *T*ransfer), which combats budget dispersion via a reward-weighted evaluation scheme that guarantees superior expected improvement, while simultaneously mitigating negative transfer through a robust reward-weighted aggregation of mean and covariance statistics, ensuring a safe fallback to independent evolution. |
Yanchi Li; Jiao Liu; Wenyin Gong; Qiong Gu; Yue Zhao; Yew Soon ONG; |
| 474 | Maximizing Mutual Information Between Prompt and Response Improves LLM Performance with No Additional Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose **Mutual Information-based Preference Optimization (MIPO)**, a contrastive data augmentation method that constructs preference pairs by generating a positive response conditioned on the correct prompt and a negative response conditioned on a random or incomplete prompt; then train with Direct Policy Optimization. |
HyunJi Nam; Haoran Li; Natasha Jaques; |
| 475 | Is Code Better Than Language for Algorithmic Reasoning? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Comparing NL reasoning and solver-based pipelines directly is ill-posed: they differ simultaneously in representation space and execution mechanism. We introduce a three-route framework that makes this comparison tractable by introducing an intermediary step—code generation with LLM-based execution. |
Terry Tong; Yu Feng; Surbhi Goel; Dan Roth; |
| 476 | Meta Context Engineering Via Agentic Skill Evolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: They impose structural biases and restrict context optimization to a narrow, intuition-bound design space. To address this, we introduce Meta Context Engineering (MCE), a bi-level framework that supersedes static CE heuristics by co-evolving CE skills and context artifacts. |
Haoran Ye; Xuning He; Vincent Arak; Haonan Dong; Guojie Song; |
| 477 | Beyond Majority Voting: Self-Reflective Test-Time Reinforcement Learning for LLM Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As a result, rare yet correct trajectories are systematically undervalued by majority-voting-based approaches. To address this limitation, we propose Self-Reflective Test-Time Reinforcement Learning (SR-TTRL), a novel framework that leverages self-reflective verification to produce high-fidelity pseudo-labels. |
Sitong Wu; Haoru Tan; Xichen Zhang; Bin Xia; Shaofeng Zhang; XIAOJUAN QI; Bei Yu; Jiaya Jia; |
| 478 | Uncovering Hidden Triggers: Backdoor Attribution in Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Previous research on interpretability for LLM safety tends to focus on alignment, jailbreak, and hallucination, but overlooks backdoor mechanisms, making it difficult to understand and fully eliminate the backdoor threat. In this paper, aiming to bridge this gap, we explore the interpretable mechanisms of LLM backdoors through Backdoor Attribution (BkdAttr), a tripartite causal analysis framework. |
Miao Yu; Zhenhong Zhou; Moayad Aloqaily; Kun Wang; Biwei Huang; Stephen Wang; Yueming Jin; Qingsong Wen; |
| 479 | SafeSeek: Universal Attribution of Safety Circuits in Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing safety attribution methods struggle with generalization and reliability due to their reliance on heuristic, domain-specific metrics and search algorithms. To address this, we propose SafeSeek, a unified safety interpretability framework that identifies functionally complete safety circuits in LLMs via optimization. |
Miao Yu; Siyuan Fu; Moayad Aloqaily; Zhenhong Zhou; Safa Otoum; Xing fan; Kun Wang; Yufei Guo; Qingsong Wen; |
| 480 | Semantic Cache Distillation: Efficient State Transfer Via Reuse and Selective Patching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Semantic Cache Distillation (SCD), a loss-constrained framework that replaces raw KV transmission with compact semantic codes. |
Qianli Ma; Zhiqing Tang; Hanshuai Cui; Zhi Yao; Weijia Jia; |
| 481 | Influence-Guided Symbolic Regression: Scientific Discovery Via LLM-Driven Equation Search with Granular Feedback Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce \textit{Influence-Guided Symbolic Regression} (IGSR), a method that frames equation discovery as an iterative two-step process combining diverse term generation with rigorous selection: an LLM generates candidate basis functions $\psi_j(\mathbf{x})$ for a linear model, which are then evaluated using granular influence scores $\Delta_j$. |
Evgeny S. Saveliev; Samuel Holt; Nabeel Seedat; David Bentley; Jim Weatherall; Mihaela van der Schaar; |
| 482 | Learning Structured Reasoning Via Tractable Trajectory Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose Ctrl-R, a framework for learning structured reasoning via tractable trajectory control that actively guides the rollout process, incentivizing the exploration of diverse reasoning patterns that are critical for complex problem-solving. |
Po-Nien Kung; Zhen Yang; Jeffrey Luo; Cheng-Fu Yang; Haikang Deng; Zi-Yi Dou; Yinfei Yang; Nanyun Peng; Zhe Gan; Kai-Wei Chang; |
| 483 | Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: We introduce Thought-Aligner, a lightweight plug-in safety model that performs causal correction on unsafe thoughts before action execution, without altering the underlying agent. |
Changyue Jiang; Wenqi Zhang; Xudong Pan; Geng Hong; Min Yang; |
| 484 | Revisiting Robustness for LLM Safety Alignment Via Selective Geometry Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we revisit robustness for safety alignment from an optimization geometry perspective, highlighting optimization-induced fragility as a complementary factor to data-space uncertainty. |
Yonghui Yang; WenjianTao; Jilong Liu; Xingyu Zhu; Junfeng Fang; Huang Weibiao; Le Wu; Richang Hong; Tat-Seng Chua; |
| 485 | PretrainZero: Reinforcement Active Pretraining Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose PretrainZero, a reinforcement active learning framework built on the pretraining corpus to extend RL from domain-specific post-training to general pretraining. |
Xingrun Xing; Zhiyuan Fan; Jie Lou; Guoqi Li; Jiajun Zhang; Debing Zhang; |
| 486 | STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This trilemma stems from a critical \textit{topological mismatch}: the prevailing isotropic Gaussian prior in standard VAEs imposes a \textit{flat} latent geometry that fails to accommodate audio’s \textit{hierarchical} nature, where low-frequency components are structured and compressible while high-frequency components are stochastic and incompressible, leading to \textit{disordered information packing} where crucial semantic features are randomly interleaved with high-entropy noise. To resolve this challenge, we propose \textbf{Structured Topology-Aware Regularization (STAR)}, a general training strategy that reshapes latent space geometry by imposing a growth-based constraint field, routing structural and textural information into channel subspaces with matching capacities. |
Huadai Liu; Wen Wang; Kaicheng Luo; Qian Chen; Xiangang Li; Wei Xue; |
| 487 | Clover: Accurate LLM Pre-Training in NVFP4 By Improved Unbiased Gradient Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, improve the state of the art for quantized training in NVFP4 via a novel unbiased quantization routine for micro-scaled formats, called MS-EDEN, that has more than 2x lower quantization error than SR. |
Andrei Panferov; Erik Schultheis; Rush Tabesh; Dan Alistarh; |
| 488 | NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent advances in coding agents suggest rapid progress toward autonomous software development, yet existing benchmarks primarily evaluate short-horizon behaviors such as localized code generation, scaffolded completion, or repository repair, leaving it unclear whether agents can sustain coherent reasoning, planning, and execution over the extended horizons demanded by real-world repository construction. To address this gap, we introduce NL2Repo-Bench, a benchmark explicitly designed to evaluate the long-horizon repository generation from scratch: given only a single natural-language requirements document and an empty workspace, agents must autonomously design the architecture, manage dependencies, and produce a fully installable Python library. |
Jingzhe Ding; Shengda Long; puchangxin; Ge Zhang; zhou huan; Hongwan Gao; Xiang Gao; Chao He; Yue Hou; FEI HU; Zhaojian Li; Weiran Shi; Zaiyuan Wang; Daoguang Zan; Chenchen Zhang; Xiaoxu Zhang; Chen Qizhi; cheng; Bo Deng; Qingshui Gu; Kai Hua; Juntao Lin; Pai Liu; Mingchen Li; Minghao Li; Xuanguang Pan; Zifan Peng; Yujia Qin; Yong Shan; Zhewen Tan; Haoran Wang; Zihan Wang; Weihao Xie; Yishuo Yuan; Jiayu Zhang; Yunfei Zhao; He Zhu; LIYA ZHU; chenyangzou; Ming Ding; Jiaheng Liu; Jianpeng Jiao; Minghao Liu; Qian Liu; Chongyang Tao; Jian Yang; Tong Yang; Zhaoxiang Zhang; Xinjie Chen; Wenhao Huang; |
| 489 | Decoupled Low-Rank Adaptation for Robust Federated Fine-Tuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we investigate the structural properties of LoRA and reveal a robustness asymmetry. |
Xiuwen Fang; Xuliang Yang; Mang Ye; |
| 490 | Contextualized Privacy Defense for LLM Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These paradigms are insufficient for supporting contextual, proactive privacy decisions in multi-step agent execution. We propose *Contextualized Defense Instructing (CDI)*, a new privacy defense paradigm in which an instructor model generates step-specific, context-aware privacy guidance during execution, proactively shaping actions rather than merely constraining or vetoing them. |
Yule Wen; Yanzhe Zhang; Jianxun Lian; Xiaoyuan Yi; Xing Xie; Diyi Yang; |
| 491 | PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce the Perception Rubric Benchmark (PRB), a rubric-based evaluation framework for Multimodal Large Language Models (MLLMs) that addresses the growing gap between benchmark scores and human-perceived quality. |
Yana Wei; Hongbo Peng; Yanlin Lai; Liang Zhao; Kangheng Lin; En Yu; Keyu Lv; Han Zhou; Yin Tang; Haodong Li; Mitt Huang; Hangyu Guo; Jianjian Sun; Zheng Ge; Xiangyu Zhang; Daxin Jiang; Vishal Patel; |
| 492 | Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Yet pre-trained models are routinely subjected to further transformations—such as fine-tuning to acquire new capabilities or quantization for efficiency. In this work, we evaluate optimizer choices across model scales, token budgets, and datasets, and find that strategies that explicitly (Sharpness-Aware Minimization) or implicitly (large learning rates and Warmup–Stable–Decay schedules) reduce sharpness yield better downstream performance, even when they achieve comparable or worse pre-training loss. |
Ishaan Watts; Catherine Li; Sachin Goyal; Jacob Mitchell Springer; Aditi Raghunathan; |
| 493 | Calibrated Multimodal Representation Learning with Missing Modalities Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent research generalizes traditional cross-modal alignment to produce enhanced multimodal synergy but requires all modalities to be present for a common instance, making it challenging to utilize prevalent datasets with missing modalities. We provide theoretical insights into this issue from an anchor shift perspective. |
Xiaohao Liu; Xiaobo Xia; Jiaheng Wei; Shuo Yang; Xiu Su; See-Kiong Ng; Tat-Seng Chua; |
| 494 | Estimation of Treatment Effects Under Nonstationarity Via The Truncated Policy Gradient Estimator Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In these settings, interventions affect both current and future system states, so estimating the global average treatment effect (GATE) requires accounting for temporal dynamics, which is especially challenging in the presence of nonstationarity; existing approaches suffer from high bias, high variance, or both. In this paper, we address this challenge via the novel Truncated Policy Gradient (TPG) estimator, which replaces instantaneous outcomes with short-horizon outcome trajectories. |
Ramesh Johari; Tianyi Peng; Wenqian Xing; |
| 495 | A Geometry-Based View of Mahalanobis OOD Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We conduct a large-scale study across diverse foundation-model backbones and Mahalanobis variants. |
Denis Janiak; Jakub Binkowski; Tomasz Kajdanowicz; |
| 496 | Olaf-World: Orienting Latent Actions for Video World Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our key insight is that although actions are unobserved, their *semantic effects* are observable and can serve as a shared reference. |
Yuxin Jiang; Yuchao Gu; Ivor Tsang; Mike Zheng Shou; |
| 497 | DLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this specific challenge, our work begins with a key observation that dLLM inference involves a static prompt and a partially dynamic response, where most tokens remain stable across adjacent denoising steps. Based on this, we propose dLLM-Cache, a training-free adaptive caching framework that combines long-interval prompt caching with partial response updates guided by feature similarity. |
Zhiyuan Liu; Yicun Yang; Yaojie Zhang; Junjie Chen; Chang Zou; Qingyan Wei; Shaobo Wang; Yichen Zhu; Linfeng Zhang; |
| 498 | Position: Reasoning After Perception Means Reasoning Without Vision Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We argue that for a broad class of visual tasks hard to specify in language, failures stem from a structural fatality where the temporal decision of \textit{when} to reason strictly dictates the spatial constraint of where reasoning takes place. |
Hongcheng Gao; Zihao Huang; Jingyi Tang; Lin Xu; Xinhao Li; Haoyang Li; Yue Liu; Minhua Lin; Xinlong Yang; Taihang Hu; Ge Wu; Baolong Bi; Hongyu Chen; Zhiqi Huang; Wentao Zhang; |
| 499 | DITRON: Distributed Multi-level Tiling Compiler for Parallel Tensor Programs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Conversely, existing tensor compilers fail to address the complex memory hierarchy of distributed clusters effectively. To bridge this gap, we propose DITRON, a scalable tile-level compiler that democratizes high-performance distributed kernel development. |
Size Zheng; Xuegui Zheng; Hanshi Sun; Qi Hou; Wenlei Bao; Shiyu Li; Haojie Duanmu; Jin Fang; Chenli Xue; Chenhui Huang; YuanqiangLiu; Renze Chen; Ningxin Zheng; Dongyang Wang; Li-Wen Chang; Liqiang Lu; Yun Liang; Jidong Zhai; Xin Liu; |
| 500 | TFRBench: A Reasoning Benchmark for Evaluating Forecasting Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Unlike existing benchmarks, TFRBench provides a protocol for evaluating the reasoning generated by forecasting systems–specifically their analysis of cross-channel dependencies, trends, and external events. To enable this, we propose a systematic multi-agent framework that utilizes an iterative verification loop to synthesize numerically grounded reasoning traces. |
Atik Ahamed; Mihir Parmar; Palash Goyal; Yiwen Song; Long Le; Qiang (Shaun) Cheng; Chun-Liang Li; Hamid Palangi; Jinsung Yoon; Tomas Pfister; |
This table only includes 500 papers selected by our daily digest algorithm. To continue with the full list (~6,500 papers), please visit Paper Digest: ICML-2026 (Full List).