Paper Digest: ECCV 2026 Papers & Highlights
Search within ECCV-2026
Literature review on a topic
Generate a written review of ECCV-2026 research on any topic, with each claim cited to specific papers.
Browse & explore
Browse ~ 12,500 authors (ECCV-2026), or explore the “Best Paper” Digest listing the most influential ECCV papers of recent years.
Note: ECCV-2026 accepts more than 2,830 papers, this page only includes 500 of them selected by our daily paper digest algorithm. Interested users can choose to read All 2,830 ECCV-2026 papers in a separate page, which takes quite some time to load.
Since 2018, Paper Digest has built a foundation of data spanning decades of conferences, journals, and research topics. The platform features a daily digest service that sifts through tens of thousands of new papers, clinical trials, news articles, and community posts, filtering the noise to highlight what matters most to specific interests. Beyond daily updates, dozens of built-in research tools streamline the academic workflow, supporting efficient reading and writing, comprehensive literature reviews, and automated research report generation.
Paper Digest Team
New York City, New York, 10017
team@paperdigest.org
TABLE 1: Paper Digest: ECCV 2026 Papers & Highlights
| Paper | Author(s) | |
|---|---|---|
| 1 | SpecV: Specification Verification for Robust Unified Multimodal Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present SpecV, a specification-verification framework for robust unified multimodal evaluation. |
Weihao Yu; Rongyao Fang; Yuxuan Cai; Linjiang Huang; Yuhuan Yang; Xianwei Zhuang; Junyang Lin; Yixuan Yuan; Shuai Bai; |
| 2 | Continuous Adversarial Flow Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose continuous adversarial flow models, a type ofcontinuous-time flow model trained with an adversarial objective. |
Shanchuan Lin; Ceyuan Yang; Zhijie Lin; Hao Chen; Haoqi Fan; |
| 3 | Natural Image Pretraining Improves Abstract Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our anal-ysis identifies two key causes: (1) Multiview redundancy — Fromthe data perspective, certain views provide limited or noisy informa-tion, diluting discriminative cues; (2) Overfitting — From the featureperspective, conventional fusion increases representational complexity,causing the model to memorise view-specific patterns rather than learngeneralisable representations. To address these issues, we propose twocomplementary modules. |
Xiaoman Ding; Keya Hu; Katelyn Gan; Victor Yin; Kaiming He; |
| 4 | Towards Scalable Pre-training of Visual Tokenizers for Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present VTP, a unified visualtokenizer pre-training framework, pioneering the joint optimization ofimage-text contrastive, self-supervised, and reconstruction losses. |
Jingfeng Yao; Yuda Song; Yucong Zhou; Xinggang Wang; |
| 5 | TinyHistory: Lightweight Video History Embeddings Via Two-Stage Context Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present TinyHistory, a lightweight historyembedding learned through two-stage context learning. |
lvmin zhang; Shengqu Cai; Muyang Li; Chong Zeng; Beijia Lu; Anyi Rao; Song Han; Gordon Wetzstein; Maneesh Agrawala; |
| 6 | GO-Renderer: Generative Object Rendering with 3D-aware Controllable Video Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose GO-Renderer, a unified framework integrating the reconstructed 3D proxies to guide the video generative models to achieve high-quality object rendering on arbitrary viewpoints under arbitrary lighting conditions. |
Zekai Gu; Shuoxuan Feng; Yansong Wang; Hanzhuo Huang; Zhongshuo Du; Chengfeng Zhao; Chengwei Ren; Peng Wang; Yuan Liu; |
| 7 | Co-evolving Representations in Joint Image-Feature Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We argue that the represen-tation space guiding diffusion should itself adapt to the generative task.To this end, we propose Coevolving Representation Diffusion (CoReDi),a framework in which the semantic representation space evolves duringtraining by learning a lightweight linear projection jointly with the diffu-sion model. |
Theodoros Kouzelis; Spyros Gidaris; Nikos Komodakis; |
| 8 | Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To support SFT and RFT training, we introduce TraceSpatial, a largescale dataset of 30M QA pairs, spanning outdoor/indoor/tabletop scenes and supporting complex reasoning processes (up to 9 steps). |
Enshen Zhou; Yibo Li; Jingkun An; Jiayuan Zhang; Shanyu Rong; Mengzhen Liu; Yi Han; Yuheng Ji; Huajie Tan; Jiawei He; Pengwei Wang; Zhongyuan Wang; Cheng Chi; Lu Sheng; Shanghang Zhang; |
| 9 | Cambrian-P: Pose-Grounded Video Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We revisitpose as a lightweight supervisory signal and introduce Cambrian-P , avideo MLLM augmented with per-frame learnable camera tokens and apose regression head. |
Jihan YANG; Zifan Zhao; Xichen Pan; Shusheng Yang; Junyi Zhang; Hu Xu; Shang-Wen Li; Saining Xie; |
| 10 | VOID: Video Object and Interaction Deletion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing video object removal methods excel at inpaintingcontent “behind” the object and correcting appearance-level artifactssuch as shadows and reflections.To train themodel, we generate a new paired dataset of counterfactual object re-movals using Kubric and HUMOTO, where removing an object requiresaltering downstream physical interactions. |
Saman Motamed; William Harvey; Benjamin Klein; Luc Van Gool; Zhuoning Yuan; Ta-Ying Cheng; |
| 11 | Doe-2: 3D Representation World Model for Unified Driving Scene Forecasting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Most existing driving world models focus on appearancegeneration and fail to model geometric and semantic evolutions. In thispaper, we propose a native 3D world model, Doe-2, to forecast the com-prehensive scene evolutions in a unified 3D latent representation space.This representation contains both appearance, geometry, and semanticinformation and can be decoded into multi-view RGB, depth, semantics,and 3D occupancy. |
Dong Zhuo; Wenzhao Zheng; Sicheng Zuo; Yuanhui Huang; Siming Yan; Lu Hou; Jie Zhou; Jiwen Lu; |
| 12 | OctWorld: Long-Range World-Consistent Video Generation with Octree-based 3D Mapping Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present OctWorld, a video diffusion framework with 3Dmemory for generating explorable, world-consistent, and high-fidelity vi-sual scenes. |
Zelong Lv; Sicheng Xu; Jianfeng Xiang; Yue Dong; Ruicheng Wang; Yu Deng; Guangzhong Sun; Jiaolong Yang; |
| 13 | From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Multimodal Large Language Models (MLLMs) are increasingly applied to pixel-level vision tasks, yet their intrinsic capacity for spatial understanding remains poorly understood. |
Boyong Wu; Sanghwan Kim; Zeynep Akata; |
| 14 | V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present V-JEPA 2.1, a family of self-supervised models that learns dense, high-quality, and temporally consistent representations for visual scenes in both images and videos. |
Lorenzo Mur Labadia; Matthew Muckley; Amir Bar; Mido Assran; Koustuv Sinha; Michael Rabbat; Yann LeCun; Nicolas Ballas; Adrien Bardes; |
| 15 | ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Vision-language models (VLMs) have demonstrated remark-able capabilities in understanding and reasoning about visual content,but significant challenges persist in tasks requiring cross-viewpoint under-standing and spatial reasoning. |
Dingming Li; Hongxing Li; Zixuan Wang; Yuchen Yan; Hang Zhang; Siqi Chen; Guiyang Hou; Shengpei Jiang; Wenqi Zhang; Yongliang Shen; Weiming Lu; Yueting Zhuang; |
| 16 | Physics Meets Perception: A Reinforcement Learning Framework for Unpaired Real-World Image Dehazing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recently, Reinforcement Learning (RL) has emerged as a promising alternative; however, current RL-based restoration paradigms predominantly rely on diffusion models, leading to prohibitive computational costs and ineffective exploration. To address these bottlenecks, we propose Dehaze-RL, an efficient framework tailored for unpaired real-world dehazing. |
Yunwei Lan; Zhigao Cui; Chang Liu; Menglin Zhang; Nian Wang; Cong Zhang; Dong Liu; |
| 17 | MemLearner: Learning to Query Context Memory for Video World Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We collect a dataset of long videos with scene occlusionsand dynamic objects, paired with camera pose annotations, and proposea multi-dataset training strategy leveraging both annotated renderedand unannotated real-world videos. |
Jiwen Yu; Jianxiong Gao; Jianhong Bai; Yiran Qin; Kaiyi Huang; Quande Liu; Xintao Wang; Pengfei Wan; Kun Gai; Xihui Liu; |
| 18 | OmniRen: Neural Rendering Wih Heterogeneous Scene Primitives Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present ’RenderFormer-V2’, a unified learned transformer-based neural rendering model, complementary to modern physics-basedrendering systems, that can handle diverse light-transport effects suchas caustics, volumetric scattering, environment lighting, textured anddisplaced surfaces and out-of-distribution materials without per-scenetraining or specialized code. |
Chong Zeng; Yue Dong; Pieter Peers; lvmin zhang; Maneesh Agrawala; |
| 19 | Allo{SR}2: Rectifying One-Step Super-Resolution to Stay Real Via Allomorphic Generative Flows Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing one-stepmethods typically replace Gaussian noise with degraded low-resolution(LR) latents at initialization, introducing a substantial distribution shiftthat further leads to trajectory deviation and prior collapse under ex-treme acceleration. To overcome these limitations, we propose Allo{SR}2 ,a novel FM-based framework that rectifies one-step SR flows via allomor-phic generative flows to maintain high-fidelity generative realism. |
Zihan Wang; XUDONG HUANG; Junbo Qiao; WEI LI; jie hu; Xinghao Chen; Shaohui Lin; |
| 20 | OmniMamba: Efficient and Unified Multimodal Understanding and Generation Via State Space Models Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: We present OmniMamba, the first linear-architecture-based multimodal generation model that generates both text and imagesthrough a unified next-token prediction paradigm. |
Jialv Zou; Bencheng Liao; Qian Zhang; Wenyu Liu; Xinggang Wang; |
| 21 | UniFlow: Zero-Shot LiDAR Scene Flow for Autonomous Driving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Informed by our analy-sis, we propose UniFlow, a feedforward model that unifies and trainson multiple large-scale LiDAR scene flow datasets with diverse sensorplacements and point cloud densities. |
Siyi Li; Qingwen Zhang; Ishan Khatri; Kyle Vedder; Eric Eaton; Deva Ramanan; Neehar Peri; |
| 22 | MindBlock: Probing Spatial Assembly and Structure in Unified Multimodal Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we argue that true spa-tial intelligence demands active construction: not only recognizing a 3Dstructure in pixel space, but also building and modifying it. |
Baiqiao Yin; Junhao Liu; Han Yin; Heyang Yu; Tingxuan Zhang; Zhiheng Li; Chengzu Li; Jihan YANG; Manling Li; Chen Feng; Yiming Li; |
| 23 | EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current methods primarily rely on passive processing, which often face a trade-off between maintaining long-range context and capturing the fine-grained details necessary for complex tasks. To address this, we introduce EventMemAgent, an active online video agent framework based on a hierarchical memory module. |
Siwei Wen; Zhangcheng Wang; Xingjian Zhang; Lei Huang; wenjun wu; |
| 24 | EgoSim: Egocentric World Simulator for Embodiment Interaction Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce EgoSim, a closed-loop egocentric world sim-ulator that generates spatially consistent interaction videos and per-sistently updates the underlying 3D scene state for continuous simu-lation. |
Jinkun Hao; Mingda Jia; Xudong Xu; Ruiyan Wang; Xihui Liu; Ran Yi; Lizhuang Ma; Jiangmiao Pang; |
| 25 | Pathwise Test-Time Correction for Autoregressive Long Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While existing Test-Time Optimization (TTO)methods prove e!ective for images or short clips, we identify that they failto mitigate drift in extended sequences due to unstable reward landscapesand the hypersensitivity of distilled parameters. To overcome these limi-tations, we introduce Test-Time Correction (TTC), a training-free alter-native. |
Xunzhi Xiang; Zixuan Duan; Guiyu Zhang; Haiyu Zhang; Zhe Gao; Junta Wu; Shaofeng Zhang; Tengfei Wang; Qi Fan; Chunchao Guo; |
| 26 | Less Is More: Reducing Complexity in Vision-Language-Action Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In thiswork, we introduce StarVLA-α, a simple yet strong baseline designedto study VLA design choices under controlled conditions. |
Jinhui Ye; Ning Gao; Senqiao Yang; Jinliang Zheng; Zixuan Wang; Yuxin Chen; Pengguang Chen; Yilun Chen; Shu Liu; Jiaya Jia; |
| 27 | Vision Bridge Transformer at Scale Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Vision Bridge Transformer (ViBT), alarge-scale instantiation of Brownian Bridge Models designed for con-ditional generation. |
Zhenxiong Tan; Zeqing Wang; Xingyi Yang; Songhua Liu; Xinchao Wang; |
| 28 | EchoStyle: Unlocking High-Fidelity Video Stylization with Reverse Data Synthesis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods, usually utilizing a reference image as the style prior, suffer from content leakage, data scarcity and limited adaptability to long videos, leading to suboptimal results with severe style drift and motion distortion. For these issues, we present EchoStyle, a scalable text-driven framework to achieve high-quality stylization of videos with arbitrary lengths. |
Huaqiu Li; Jiahao Wang; Sijia Cai; Hualian Sheng; Bing Deng; Jieping Ye; Wenhan Luo; |
| 29 | OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Extending these capabilities to remote sensingremains challenging, as models must reason over spatial scale, geographicstructures, and multispectral indices while maintaining coherent multi-step logic. To address this gap, we introduce OpenEarthAgent, a unifiedframework for tool-augmented geospatial reasoning trained on satelliteimagery, natural-language queries, and structured reasoning traces. |
Akashah Shabbir; Muhammad Umer Sheikh; Muhammad Akhtar Munir; Hiyam Debary; Mustansar Fiaz; Muhammad Zaigham Zaheer; Paolo Fraccaro; Fahad Shahbaz Khan; Muhammad Haris Khan; Xiao Xiang Zhu; Salman Khan; |
| 30 | P3-SAM: Native 3D Part Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose anative 3D point-promptable part segmentation model termed P3 -SAM,designed to fully automate the segmentation of any 3D objects into com-ponents. |
CHANGFENG MA; YANG LI; Xinhao Yan; Jiachen Xu; Yunhan Yang; Chunshi Wang; Zibo Zhao; Yanwen Guo; Zhuo Chen; Chunchao Guo; |
| 31 | Syn4D: A Multiview Synthetic 4D Dataset Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Progress in tasks like 3D reconstruction and tracking of dynamic scenes from monocular video is constrained by the scarcity of high-quality datasets with dense, complete, and accurate geometric annotations. To address this limitation, we introduce Syn4D, a multiview synthetic dataset of dynamic scenes that includes ground-truth camera motion, depth maps, dense tracking, and parametric human pose annotations. |
Zeren Jiang; Yushi Lan; Yihang Luo; Yufan Deng; Zihang Lai; Edgar Sucar; Christian Rupprecht; Iro Laina; Diane Larlus; Chuanxia Zheng; Andrea Vedaldi; |
| 32 | FixAnything: 3D-Consistent Rendering Refinement Via Video Generative Priors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our key insight is that even noisily-rendered sequences preserve camera motion and coarse scene structure,allowing cleanup to be formulated as video-to-video translation. |
Khiem Vuong; Deva Ramanan; Srinivasa G. Narasimhan; |
| 33 | SA-ResGS: Self-Augmented Residual 3D Gaussian Splatting for Next Best View Selection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Self-Augmented Residual 3D Gaussian Splat-ting, a novel framework for stabilizing uncertainty quantification andenhancing uncertainty-aware supervision in Next-Best-View selection foractive scene reconstruction. |
Kim Jun-Seong; Tae-Hyun Oh; Eduardo Pérez Pellitero; Youngkyoon Jang; |
| 34 | Dynamic-Robust Photometric–Semantic Reconstruction for Open-Vocabulary 3D Scene Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, their inherent reliance on static-sceneassumptions leads to severe misalignment of spatial features in uncon-strained dynamic environments. To bridge this critical gap, we proposeSPAR, a novel joint semantic-geometric encoding architecture that ex-plicitly isolates transient dynamic noise prior to latent space aggregation.Furthermore, we introduce a dynamic-region-aware end-to-end trainingparadigm that structurally couples motion estimation with multi-viewvisual and semantic learning. |
Boyu Cai; Li Yang; Yan Xu; Wei Liu; Nian Liu; Sikui Zhang; Yan Wang; Chunfeng Yuan; Weiming Hu; |
| 35 | On Geometric Understanding and Learned Priors in Feed-forward 3D Reconstruction Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We find that epipolar geometry emerges within the intermediate layers of all three models and is causally linked to correspondence patterns in attention heads. To study this, we perform a systematic analysis of their internal representations across three real–world datasets and a controlled synthetic dataset. |
Jelena Bratulić; Sudhanshu Mittal; Thomas Brox; Christian Rupprecht; |
| 36 | Vero: Open Reinforcement Learning Recipes for Visual Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Vero, a family of fully open VLMs that match or exceed existing open-weight models across diverse visual reasoning tasks. |
Gabriel Sarch; Linrong Cai; Qunzhong Wang; Haoyang Wu; Danqi Chen; Zhuang Liu; |
| 37 | MindDrive: A Vision-Language-Action Model for Autonomous Driving Via Online Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, applying online reinforcement learning to VLA models in autonomous driving is hindered by inefficient exploration in continuous action spaces. To overcome this limitation, we propose MindDrive, a VLA framework comprising a large language model (LLM) with two distinct sets of LoRA parameters. |
haoyu fu; Diankun Zhang; Zongchuang Zhao; Jianfeng Cui; Hongwei Xie; Bing Wang; Guang Chen; Hangjun Ye; Dingkang Liang; Xiang Bai; |
| 38 | Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However,we observe that per-frame depth remains stable throughout this failure.The backbone’s local geometry remains intact; only the global pose headbreaks down. Motivated by this decoupling, we introduce Scal3R. |
Chin-Yang Lin; Yang-Che Sun; Cheng Sun; Fu-En Yang; Min-Hung Chen; Yen-Yu Lin; Wei-Chen Chiu; Yu-Lun Liu; |
| 39 | Controllable Egocentric Video Generation Via Occlusion-Aware Sparse 3D Hand Joints Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: By relying on dense 2D trajectories or implicit pose representations, they collapse crucial geometric structures into spatially ambiguous signals, leading to severe motion inconsistencies and hallucinated artifacts under egocentric occlusions. To address this, we propose leveraging sparse 3D hand joints as explicit control signals with three key advantages: explicit geometry to resolve occlusions, an intuitive interface for interactive editing, and cross-embodiment generalization to robotic hands. |
Chenyangguang Zhang; Botao Ye; Boqi Chen; Alexandros Delitzas; Fangjinhua Wang; Marc Pollefeys; Xi Wang; |
| 40 | Guiding The Blind: Generalizing GUI Agents to Unseen Websites Via Multimodal Tutorials Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduceWebOne, a novel benchmark designed to evaluate an agent’s ability to mas-ter unseen websites by referencing Multimodal Tutorials—heterogeneousknowledge sources derived from instructional videos, historical trajec-tories, and human demonstrations. |
Xinwei Long; Kai Tian; Peng Xu; Weibo Gao; Yihua Shao; Guoli Jia; Haozhe Geng; Sa Yang; Jingxuan Li; Huayong Hu; Kaiyan Zhang; Jiaqi Wang; Bowen Zhou; |
| 41 | Vision-as-Inverse-Graphics Agent Via Interleaved Multimodal Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Vision-as-inverse-graphics, the concept of reconstructing im-ages into editable programs, remains challenging for Vision-LanguageModels (VLMs), which inherently lack fine-grained spatial grounding inone-shot settings. To address this, we introduce VIGA (Vision-as-Inverse-Graphics Agent), an interleaved multimodal reasoning framework wheresymbolic logic and visual perception actively cross-verify each other. |
Shaofeng Yin; Jiaxin Ge; Zora Wang; Xiuyu Li; Chenyang Wang; Michael Black; Trevor Darrell; Angjoo Kanazawa; Haiwen Feng; |
| 42 | Sim, Yet Same: Physics-Aligned Simulator As Zero-Shot Data Scaler in Deformable Worlds Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To addressthis, we introduce SIM1, a physics-aligned real-to-sim-to-real data enginethat grounds simulation in the physical world. |
Yunsong Zhou; Hangxu Liu; Xuekun Jiang; Xing Shen; Yuanzhen Zhou; Hui Wang; Baole Fang; Yang Tian; Mulin Yu; Qiaojun Yu; Li Ma; Hengjie Li; Hanqing Wang; Jia Zeng; Jiangmiao Pang; |
| 43 | OmniX: Any-view and Any-time 4D Reconstruction Via Feed-forward Trajectory Fields Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This limits their ability to aggregate observations over time andreconstruct complete dynamic scenes under large viewpoint changes. Toaddress this, we propose OmniX, a feed-forward 4D reconstruction frame-work that predicts dense 3D point trajectories for every pixel fromvideos with large camera motion. |
Yanqin Jiang; Tengfei Wang; Zhenwei Wang; Chenjie Cao; Junta Wu; Wenhan Luo; Weiming Hu; Jin Gao; Chunchao Guo; |
| 44 | Stream3D-VLM: Online 3D Spatial Understanding with Incremental Geometry Priors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we presentan online 3D vision-language model that enables real-time spatial un-derstanding from streaming video. |
Hanxun Yu; Xuan Qu; Lei Ke; Boqiang Zhang; YUXIN WANG; Jianke Zhu; Dong Yu; |
| 45 | LARY: A Latent Action Representation Yielding Benchmark Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce the Latent Action Representation Yielding (LARY) Benchmark, a unified framework for evaluating latent action representations on both highlevel semantic actions (what to do) and low-level robotic control (how to do). |
Dujun Nie; Fengjiao Chen; Jun Kuang; Qi Lv; Xiaoyu Li; Xuezhi Cao; |
| 46 | Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Real-world deployment of embodied agents requires activeexploration, visual grounding, and interactive intent disambiguation. How-ever, existing frameworks often rely on privileged simulator states or as-sume complete instructions, bypassing realistic deployment challenges.To bridge this gap, we present REAL, an agentic framework for open-world mobile manipulation. |
Boyu Mi; Mengchen Ma; Yifei Yao; Xing Gao; Hanqing Wang; Junting Chen; Yangzi Li; Zihou Zhu; Guohao Li; Zhenfei Yin; Tai Wang; Yao Mu; Jiangmiao Pang; |
| 47 | Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despiterequiring no auxiliary losses, reward shaping, or architectural changes— serving as a minimal, plug-in modification to standard post-trainingpipelines — models trained under perceptual starvation achieve substan-tial gains of 5% average relative improvement across diverse benchmarks.Our codes and data will be publicly available at https://github.com/WhuanY/Starve2Perceive. |
Yuhuan Wu; Haozhe Wang; Cong Wei; Chong Peng; Fangzhen Lin; Wenhu Chen; |
| 48 | ROAR-3D: Routing Arbitrary Views for High-Fidelity 3D Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Based on this analysis, we propose ROAR-3D, a lightweight method that upgrades a pretrained singleview model to accept an arbitrary number of unposed images. |
Hanxiao Sun; Mingxin Yang; Shuhui Yang; Zebin He; Xintong Han; Hongbo Fu; Chunchao Guo; Wenhan Luo; |
| 49 | MMDiff: Extending Diffusion Transformers for Multi-Modal Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present MMDiff, a framework that transforms a frozen diffusion transformer into a multi-modal generative system that jointly produces images alongside any combination of dense perceptual modalities using lightweight decoder heads. |
Yagmur Akarken; Orest Kupyn; Christian Rupprecht; |
| 50 | LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inthis paper, we show that competitive performance can be obtained farmore efficiently by strategically fusing publicly available models special-ized for either generation or understanding. |
Zeyu Wang; Zilong Chen; Chenhui Gou; Feng Li; Chaorui Deng; Deyao Zhu; Kunchang Li; Weihao Yu; Haoqin Tu; Haoqi Fan; Cihang Xie; |
| 51 | Denoising-GS: Gaussian Splatting with Spatial-aware Denoising Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we introduce a new perspec-tive by formulating the optimization of 3DGS as a primitive denoisingprocess and propose Denoising-GS, a spatial-aware denoising frame-work for Gaussian primitives by taking both the positions and spatialstructure into consideration. |
Qingyuan Zhou; Xinyi Liu; WEIDONG YANG; Ning Wang; Shuquan Ye; Ben Fei; Ying He; Wanli Ouyang; |
| 52 | Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce LocateAnything,a unified generative grounding and detection framework based on Par-allel Box Decoding (PBD). |
Shihao Wang; Shilong Liu; Yuanguo Kuang; Xinyu Wei; Yangzhou Liu; Zhiqi Li; Yunze Man; Guo Chen; Andrew Tao; Guilin Liu; Jan Kautz; Lei Zhang; Zhiding Yu; |
| 53 | SwiftWA: An Efficient Action-Centered World-Action Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing approaches face two critical bot-tlenecks: joint reasoning over future visual dynamics and actions incurssubstantial inference overhead, and joint modeling entangles visual andmotion representations, making motion prediction dependent on videoforecasts. To address these issues, we introduce GigaWorld-Policy, anaction-centered WAM that learns 2D pixel–action dynamics while en-abling efficient action decoding with optional video generation. |
Chaojun Ni; Xinyu Zhou; YuKun Zhou; Jingyu Liu; Xiaofeng Wang; Zheng Zhu; Yang Wang; Qiuping Deng; Yun Ye; Hao Li; Zhichao Liu; Jindi Lv; Boyuan Wang; Guosheng Zhao; Guan Huang; Min Cao; Wenjun Mei; |
| 54 | ReconDreamer-RL: Enhancing Reinforcement Learning Via Diffusion-based Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To cover more corner cases, we propose the Dynamic Ad-versary Agent (DAA), which adjusts surrounding vehicles’ trajectoriesrelative to the ego vehicle to generate challenging scenarios such as cut-ins. |
Chaojun Ni; Guosheng Zhao; Xiaofeng Wang; Zheng Zhu; Wenkang Qin; Chen Xinze; Guanghong Jia; Guan Huang; Wenjun Mei; |
| 55 | Aligning Human Sense: Calibrated Distributional Reward Learning for Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: 3) In policy optimization, the widely adopted KL-divergenceimposes only local constraints, failing to capture holistic human pref-erence. To address these challenges, we propose a unified, preference-aware learning framework for video generation. |
Naixin Zhai; Weihua Cheng; Dexu Yu; Yikai Gu; Hanwen Du; Junchen Fu; Chenxi Huang; Yingwei Song; Liyuan Ma; Yang Ran; Youhua Li; Yongxin Ni; |
| 56 | Tac2Real: Reliable and GPU Visuotactile Simulation for Online Reinforcement Learning and Zero-shot Real-World Deployment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Tac2Real integrates the Precondi-tioned Nonlinear Conjugate Gradient Incremental Potential Contact (PNCG-IPC) method with a multi-node, multi-GPU high-throughput par-allel simulation architecture, which can generate marker displacementfields at interactive rates. |
Ningyu YAN; Shuai Wang; Xing Shen; Hui Wang; Hanqing Wang; Yang Xiang; Jiangmiao Pang; |
| 57 | ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present ViQ, a Visual Quantized Representationsframework, which is designed to balance semantics and details in discreterepresentations while supporting inputs at native resolutions, therebyenabling it to serve as a unified and general discrete representation forarbitrary visual inputs. |
Xumin Yu; Zuyan Liu; Zhenyu Yang; Yuhao Dong; Shengsheng Qian; Jiwen Lu; Han Hu; Yongming Rao; |
| 58 | SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present SynCity 3000, a framework for generating 3Dscenes that are globally coherent while enabling fine-grained layout con-trol. |
Paul Engstler; Iro Laina; Christian Rupprecht; Andrea Vedaldi; |
| 59 | Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we aim to tackle thischallenge by leveraging an intriguing observation that the dVLA withbidirectional attention still adheres to a block-wise left-to-right order,which motivates the application of block diffusion. |
Wenxuan Song; Jiayi Chen; Shuai Chen; Jingbo Wang; Pengxiang Ding; Han Zhao; qin yikai; Xinhu Zheng; Yan Wang; Donglin Wang; Haoang Li; |
| 60 | Towards More Efficient Decoding for Autoregressive Vision-language-action Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While recent studies have explored Jacobi decoding as a more efficient alternative to traditional autoregressive decoding, its practical benefits are marginal due to the lengthy iterations. To address this problem, we introduce consistency distillation to teach the model to predict multiple correct action tokens in each iteration, thereby reducing the total iterations. |
Wenxuan Song; Jiayi Chen; Pengxiang Ding; Yuxin Huang; Han Zhao; Yinchuan Li; Yingcong Chen; Donglin Wang; Haoang Li; |
| 61 | Seeing As Humans Do: Learning from Motion to Segment Anything Without Supervision Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: The Segment Anything Model (SAM) relies heavily on mas-sive manual annotations, creating a fundamental bottleneck for modelscaling. While unsupervised methods attempt to learn … |
Weijian Jian; Xiaoyue Zhang; Bin Xiao; Chunyu Xie; Yixiao He; Yutao Liu; Dawei Leng; Yuhui Yin; |
| 62 | RePlan: Reasoning-Guided Region Planning for Complex Instruction-Based Image Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce RePlan (Region-alignedPlanning), a plan-then-execute framework that couples a vision–languageplanner with a diffusion editor. |
Tianyuan Qu; Lei Ke; Xiaohang Zhan; Longxiang Tang; Yuqi Liu; Bohao PENG; Bei Yu; Dong Yu; Jiaya Jia; |
| 63 | MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose MixGRPO, a novel framework that leverages the flexibility of mixed sampling strategies through the integration of SDE and Ordinary Di!erential Equations (ODE). |
Junzhe Li; Yutao Cui; Tao Huang; Chuxuan Zeng; Weijie Kong; Yinping Ma; Chun Fan; Miles Yang; Zhao Zhong; Liefeng Bo; |
| 64 | Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Large Multimodal Models (LMMs) have achieved remark-able success on images and short videos, yet scaling them to long videosremains challenging due to frame-centric tokenization and limited con-text windows. |
Lucy Lin; Ayush Jain; Yifan Liu; Katerina Fragkiadaki; |
| 65 | The Prism Hypothesis: Harmonizing Semantic and Pixel Representations Via Unified Autoencoding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we systematically analyze the spectral character-istics of various semantic and pixel encoders. |
Weichen Fan; Haiwen Diao; Quan Wang; Dahua Lin; Ziwei Liu; |
| 66 | LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by LEGO construction, arecreational activity that critically relies on multi-step spatial reasoning,we introduce LEGO-Puzzles: a benchmark designed to systematicallyevaluate the spatial reasoning capabilities of MLLMs from basic spatialunderstanding to multi-step planning. |
Kexian Tang; Junyao Gao; Yanhong Zeng; Haodong Duan; Sun Yanan; Zhening Xing; Wenran Liu; Kai Chen; Kaifeng Lyu; |
| 67 | Geometric Context Transformer for Streaming 3D Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Motivated by the principles of Simultaneous Localization and Mapping (SLAM), we introduce GCT, geometric context transformer, a feed-forward 3D foundation model for reconstructing scenes from streaming data. |
Lin-Zhuo Chen; Jian Gao; Shangzhan Zhang; Yihang Chen; Nan Xue; Jianyuan Wang; Christian Rupprecht; Xun Cao; Xing Zhu; Yujun Shen; Yao Yao; YINGHAO XU; |
| 68 | RoboClaw: An Agentic Framework for Scalable Long-Horizon Robotic Tasks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present RoboClaw, an agentic robotics framework that unifies data collection, policy learning, and task execution under a single VLM-driven controller. |
Ruiying Li; Yunlang Zhou; Yuyao Zhu; Kylin Chen; Sukai Wang; Kongtao Hu; Minhui Yu; Bowen Jiang; Jiayao Ma; Zhan Su; Yongjian Shen; Yang Yang; Guanghui Ren; Maoqing Yao; Wenhao Wang; Yao Mu; |
| 69 | Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Re2Pix, a hierarchical video prediction framework that decomposes forecasting into two stages: semantic representation prediction and representation-guided visual synthesis. |
Efstathios Karypidis; Spyros Gidaris; Nikos Komodakis; |
| 70 | SD3.5-Flash: Distribution-Guided Distillation of Generative Flows Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce two key innovations: “timestep sharing” to reduce gradient noise and “split-timestep fine-tuning” to improve prompt alignment. |
Hmrishav Bandyopadhyay; Rahim Entezari; Jim Scott; Reshinth Adithyan; Yi-Zhe Song; Varun Jampani; |
| 71 | Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Spatial-TTT towards streaming visual-based spatial intelligence with test-time training (TTT), which adapts a subset of parameters (fast weights) to capture and organize spatial evidence over long-horizon scene videos.Beyond architecture design, we construct a dataset with dense 3D spatial descriptions, which guides the model to update its fast weights to memorize and organize global 3D spatial signals in a structured manner. |
FANGFU LIU; Diankun Wu; Jiawei Chi; Yimo Cai; Yi-Hsin Hung; Xumin Yu; Hao Li; Han Hu; Yongming Rao; Yueqi Duan; |
| 72 | BLASt3R: Bundle Adjustment of Any Image Set with Multi-View Matching and Monocular Priors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper,we introduce a regularized BA framework that leverages a fast multi-view matcher and monocular priors for initialization and regularization.In contrast to existing systems, our unified approach seamlessly supportsboth online VSLAM and offline reconstruction from unordered image col-lections within the same optimization framework and sharing commonhyperparameters for all tasks. |
Vincent Leroy; Philippe Weinzaepfel; Lojze Zust; Yohann Cabon; Jerome Revaud; |
| 73 | UniPR-3D: Towards Universal Visual Place Recognition with Visual Geometry Grounded Transformer Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In thiswork, we introduce UniPR-3D, the first VPR architecture that effectivelyintegrates geometry-aware information from multiple views. |
Tianchen Deng; Chen Xun; Ziming Li; Hongming Shen; Shuhao Zhai; Danwei Wang; Javier Civera; Hesheng Wang; |
| 74 | R3DP: Real-Time 3D-Aware Policy for Embodied Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Real-time 3D-aware Policy (R3DP), which integrates powerful 3D priors into manipulation policies without sacrificing real-time performance. |
Yuhao Zhang; Wanxi Dong; Yue Shi; Yi Liang; Jingnan Gao; Qiaochu Yang; Yaxing Lyu; Zhixuan Liang; Yibin Liu; Congsheng Xu; Xianda Guo; Wei Sui; Yaohui Jin; Xiaokang Yang; Yanyan Xu; Yao Mu; |
| 75 | DiffusionVL: Translating Any Autoregressive Models Into Diffusion Vision Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose DiffusionVL, a family of dVLMs obtained by translating pretrained AR models into the diffusion paradigm via an efficient diffusion finetuning procedure that changes the training objective and decoding process while keeping the backbone architecture intact. |
Lunbin Zeng; Jingfeng Yao; Bencheng Liao; Hongyuan Tao; Wenyu Liu; Xinggang Wang; |
| 76 | CloSeR: Unified Relational Distillation from Closed-Set Teachers for Category Discovery Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose CloSeR, a simple plug-and-play framework that in-jects Closed-Set Relational knowledge into GCD training. |
Yuanpei Liu; Zhenqi He; Jialu Tang; Kai Han; |
| 77 | From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A primary bottleneck is that currentvideo MLLMs, trained primarily under a Supervised Fine-Tuning (SFT)paradigm, function as passive “Observers” that recognize ongoing eventsrather than evaluating the current state relative to the final task goal.In this paper, we introduce PRIMO R1 (Process Reasoning InducedMOnitoring), a 7B framework that transforms video MLLMs into ac-tive “Critics”. |
Yibin Liu; Yaxing Lyu; Daqi Gao; Zhixuan Liang; Weiliang Tang; Shilong Mu; Xiaokang Yang; Mingyu Ding; Yao Mu; |
| 78 | Molmo-Point: Better Pointing for VLMs with Grounding Tokens Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Using this method, we set a new state-of-the-art on image pointing(70.7% on PointBench), set a new state-of-the-art for fully open modelson GUI pointing (61.1% on ScreenSpotPro), substantially improve VLMvideo tracking (62.5 on J &F vs 56.7 for Molmo2 on Molmo2Track), andimprove video pointing (59.1% human preference win rate vs. Molmo2). |
Christopher Clark; Yue Yang; Jae Sung Park; Zixian Ma; Jieyu Zhang; Rohun Tripathi; Mohammadreza Salehi; Sangho Lee; Ranjay Krishna; |
| 79 | Scaling Verification Can Be More Effective Than Scaling Policy Learning for Vision-Language-Action Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We first characterize the test-time scaling lawsfor embodied instruction following and demonstrate that jointly scaling thenumber of rephrased instructions and generated actions greatly increasestest-time sample diversity, often recovering correct actions more efficientlythan scaling each dimension independently. To capitalize on these scalinglaws, we present CoVer, a contrastive verifier for vision–language–actionalignment, and show that our architecture scales gracefully with additionalcomputational resources and data. |
Jacky Kwok; Xilun Zhang; Mengdi Xu; Yuejiang Liu; Azalia Mirhoseini; Chelsea Finn; Marco Pavone; |
| 80 | Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Compounding this is a lack of reliable granular evaluation methods for localizing and specifying physical law violations in videos. We address this by introducing Physics Question Scene Graph (PQSG), a hierarchical question-based evaluation pipeline. |
Atin Pothiraj; Jaemin Cho; Yue Zhang; Elias Stengel-Eskin; Mohit Bansal; |
| 81 | Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Live Avatar, an algorithm-system co-designed frame-work that addresses both challenges for a 14-billion-parameter di_x001B_usionmodel. |
Yubo Huang; Hailong Guo; Fangtai Wu; Weiqiang Wang; Shijie Huang; Qijun Gan; Shifeng Zhang; Lin Liu; Sirui Zhao; Enhong Chen; Jiaming Liu; Steven Hoi; |
| 82 | VLA-R1: Enhancing Reasoning in Vision-Language-Action Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Vision-Language-Action (VLA) models aim to unify per-ception, language understanding, and action generation, offering strongcross-task and cross-scene generalization with broad impact on embodiedAI. |
Angen Ye; Zeyu Zhang; Boyuan Wang; Xiaofeng Wang; Dapeng Zhang; Zheng Zhu; |
| 83 | HAD: Combining Hierarchical Diffusion with Metric-Decoupled RL for End-to-End Driving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: End-to-end planning has emerged as a dominant paradigmfor autonomous driving, where recent models often adopt a scoring-selection framework to choose trajectories from a large set of candidates,with diffusion-based decoding showing strong promise. |
Wenhao Yao; Xinglong Sun; Zhenxin Li; Shiyi Lan; Zi Wang; Jose M Alvarez; Zuxuan Wu; |
| 84 | Recurrent Autoregressive Diffusion: Global Memory Meets Local Attention Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Moreover, existingdiffusion-RNN approaches often suffer from performance degradation dueto training-inference gap or the lack of overlap across windows. To addressthese limitations, we propose a novel Recurrent Autoregressive Diffusion(RAD) framework, which leverages recurrent blocks for memory update andretrieval and preserves local details by full attention on overlapping slidingwindows, with no training and inference gap. |
Taiye Chen; Zihan Ding; Anjian Li; Christina Zhang; Zeqi Xiao; Yisen Wang; Chi Jin; |
| 85 | 3D-Layout-R1: Structured Reasoning for Language-Instructed Spatial Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a Structured Reasoning frameworkthat performs text-conditioned spatial layout editing via scene-graph rea-soning. |
Haoyu Zhen; Xiaolong Li; Yilin Zhao; Han Zhang; Sifei Liu; Kaichun Mo; Chuang Gan; Subhashree Radhakrishnan; |
| 86 | Skyfall-GS: Synthesizing Immersive 3D Urban Scenes from Satellite Imagery Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we take an alternative route to create largescale 3D scenes by leveraging readily available satellite imagery for realistic coarse geometry and open-domain diffusion models for high-quality close-up appearance synthesis. |
Jie-Ying Lee; Yi-Ruei Liu; Shr-Ruei Tsai; Wei-Cheng Chang; Chung-Ho Wu; Jiewen Chan; Zhenjun Zhao; Chieh Hubert Lin; Yu-Lun Liu; |
| 87 | From Sparse to Dense: Multi-View GRPO for Flow Models Via Augmented Condition Space Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, we have observed that the standard paradigmthat evaluates a group of generated samples against a single conditionsuffers from insufficient exploration of inter-sample relationships, con-straining both alignment efficacy and performance ceilings. To addressthis sparse single-view evaluation scheme, we propose Multi-View GRPO(MV-GRPO), a novel algorithm that enhances relationship explorationby augmenting the condition space to create a dense multi-view rewardmapping. |
Jiazi Bu; Pengyang Ling; Yujie Zhou; Yibin Wang; Yuhang Zang; Tianyi Wei; Xiaohang Zhan; Jiaqi Wang; Tong Wu; Xingang Pan; Dahua Lin; |
| 88 | StreamTalk: Streaming Co-Speech Gesture Generation with Key-Pose Anchoring Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Supplying even a single plausible key pose at a clip’s tail as a destination anchor is sufficient to suppress drift dramatically. Building on this insight, we propose StreamTalk, a closed-loop streaming framework that introduces a periodic generate-retrieve-refine feedback cycle. |
Jianfang Li; Xiangyue Zhang; Jiaxu Zhang; Kaixing Yang; Steven Hoi; |
| 89 | M2Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This “discretization bottleneck” significantly limits the performance ceiling of downstream VisionLanguage-Action (VLA) models. To address this, we propose M2Tok, a Multi-head Multi-codebook Action Tokenizer designed to minimize reconstruction error and enhance policy performance. |
Chunpu Xu; Zhixuan Liang; Yuhao Zhang; Chi-Min Chan; Jiashuo Wang; Yang Xiao; Mengkang Hu; Xiaokang Yang; Yao Mu; |
| 90 | Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this,we introduce Video-Holmes, a benchmark inspired by the reasoningprocess of Sherlock Holmes, designed to evaluate the complex video rea-soning capabilities of MLLMs. |
JUNHAO CHENG; Yuying Ge; Teng Wang; Yixiao Ge; Jing Liao; Ying Shan; |
| 91 | Reward Lightning: Fast Video Generation Via Homologous Preference Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods optimize the two objectives over mismatched representation spaces, where improving one objective often compromises the other. To overcome this, we propose Reward Lightning, a unified framework that aligns and accelerates a video diffusion model within a single shared representation. |
Jiaxiang Cheng; bing ma; Xuhua Ren; Kai Yu; Peng Zhang; Tianxiang Zheng; Qinglin Lu; |
| 92 | Anchored, Not Graded: How Vision-Language Models Fail at Slant-from-Texture Perception Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Human perception of surface slant from texture exhibits systematic, graded biases that emerge reliably in psychophysical experiments. Prior work showed that unsupervised CNNs … |
Qian Zhang; Michal Golovanevsky; Fulvio Domini; James Tompkin; |
| 93 | PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Autoregressive video diffusion models have demonstrated re-markable progress, yet they remain bottlenecked by intractable linearKV-cache growth, temporal repetition, and compounding errors duringlong-video generation. To address these challenges, we present Pack-Forcing, a unified framework that efficiently manages the generationhistory through a novel three-partition KV-cache strategy. |
Xiaofeng Mao; Shaohao Rui; Bo Zheng; Kaining Ying; Chuanhao Li; Mingmin Chi; Kaipeng Zhang; |
| 94 | MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present MA-VLA, a unified framework for multi-arm collaboration via atomic action assignment.We also construct a benchmark in which test-time collaboration patterns are absent in training set. |
Zaibin Zhang; Junlan Xiao; Zhongbo Zhang; Yifan Wang; Li Kang; Yiran Qin; Changxing Xia; Heng Zhou; Talas Fu; Enshen Zhou; Ruimao Zhang; Zhenfei Yin; Huchuan Lu; Lijun Wang; |
| 95 | Distribution Matching Distillation Meets Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Distribution Matching Distillation (DMD) facilitates efficientinference by distilling multi-step diffusion models into few-step variants.Concurrently, Reinforcement Learning (RL) has emerged as a vital toolfor aligning generative models with human preferences. |
Dengyang Jiang; Dongyang Liu; Zanyi Wang; Qilong Wu; Liuzhuozheng Li; Heng-Zhuang Li; Xin Jin; Zhen Li; Changsheng Lu; Mengmeng Wang; Steven Hoi; Peng Gao; Harry Yang; |
| 96 | OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing generation models often entangle these factors, limiting independent control. In this work, we introduce OmniCamera, a unified framework designed to explicitly disentangle and command these two dimensions. |
Yukun Wang; Ruihuang Li; Jiale Tao; Shiyuan Yang; Liyi Chen; Zhantao Yang; Handz Handz; Yulan Guo; Shuai Shao; Qinglin Lu; |
| 97 | OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introduces Om-niStream, a unified streaming visual backbone that effectively per-ceives, reconstructs, and acts from diverse visual inputs. |
Yibin Yan; Jilan Xu; Shangzhe Di; Haoning Wu; Weidi Xie; |
| 98 | ReSWD: ReSTIR‘d, Not Shaken. Combining Reservoir Sampling and Sliced Wasserstein Distance for Variance Reduction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: ReSWD: ReSTIR‘d, not shaken. Combining Reservoir Sampling and Sliced Wasserstein Distance for Variance Reduction. |
Mark Boss; Andreas Engelhardt; Simon Donné; Varun Jampani; |
| 99 | FaceMoE: Mixture of Experts for Low-Resolution Face Recognition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A single feature encoder strug-gles to generalize effectively across both domains when fine-tuned on anLR dataset, and this issue is further magnified by catastrophic forget-ting. To address these challenges, we propose FaceMoE, an effectiveadaptation of Mixture of Experts (MoE) transfomer architecture for low-resolution face-recognition . |
Kartik Narayan; Vishal Patel; |
| 100 | X-Stream: Benchmarking MLLMs As Multiplexers for Multi-Stream Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing benchmarks areconfined to single-stream paradigms, leaving a critical gap in evaluatingonline, cross-stream reasoning. To bridge this, we introduce X-Stream,the first benchmark dedicated to multi-stream streaming understand-ing. |
Peiwen Sun; Xudong LU; Huadai Liu; Yang Bo; Dongming Wu; Huankang Guan; Minghong Cai; Jinpeng Chen; Xintong Guo; Shuhan LI; FANG LIU; Rui Liu; Xiangyu Yue; |
| 101 | End-to-End Training for Autoregressive Video Diffusion Via Self-Resampling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To achievean end-to-end solution, we introduce Resampling Forcing, a teacher-free framework that enables training autoregressive video models fromscratch and at scale. |
Yuwei Guo; Ceyuan Yang; Hao He; Yang Zhao; Meng Wei; Zhenheng Yang; Weilin Huang; Dahua Lin; |
| 102 | Mechanistic Finetuning of Vision-Language-Action Models Via Few-Shot Demonstrations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In thiswork, we introduce Robotic Steering, a finetuning approach groundedin mechanistic interpretability that leverages few-shot demonstrations toidentify and selectively finetune task-specific attention heads aligned withthe physical, visual, and linguistic requirements of robotic tasks. |
Chancharik Mitra; Yusen Luo; Raj Saravanan; Dantong Niu; Anirudh Pai; Jesse Thomason; Trevor Darrell; Abrar Anwar; Deva Ramanan; Roei Herzig; |
| 103 | CascadeProto: Cascaded Cross-Modal Prototype Purification Via Entropy-Aware Learning for Few-Shot 3D Point Cloud Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To fur-ther enrich prototype representations, we introduce Learnable ModalityAdapters (LMA) that independently align each of three CLIP modal-ities — text, audio, and image — with point cloud features throughforeground-background decoupled distribution matching, enabling flexi-ble single-modality semantic enrichment that bridges the 2D-3D domaingap. |
Changshuo Wang; Weijun Li; Fan Mo; Zhonghang Liu; Shuting He; Prayag Tiwari; Dimitrios Kanoulas; |
| 104 | StreamGVE: Training-Free Video Editing Via Few-Step Streaming Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address thisgap, we revisit video editing from a noise-to-data perspective and pro-pose Streaming-Generation-based Video Editing (StreamEdit), whichpreserves few-step sampling while seamlessly injecting source-video con-ditions. |
Guanlong Jiao; Chenyangguang Zhang; Jia Xian; Zewei Zhang; Renjie Liao; |
| 105 | AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose AR-CoPO (AutoRegressive Contrastive Policy Optimization), a framework that adapts the Neighbor GRPO contrastive perspective to streaming AR generation. |
Dailan He; Guanlin Feng; Xingtong Ge; Yi ZHANG; Bingqi Ma; Guanglu Song; Yu Liu; Hongsheng LI; |
| 106 | DepWorldSG: Depth-Aware 3D Semantic Scene Graph Generation Via World-Model Priors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present DeWorldSG, a novel framework that generatesspatio-temporally robust 3D Semantic Scene Graphs from RGB-D se-quences. |
Seok-Young Kim; Abdelrahman Elskhawy; TAEWOOK HA; Dooyoung Kim; Eunjae Shin; Benjamin Busam; Woontack Woo; |
| 107 | TriMotion: Modality-Agnostic Camera Control for Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing methods typically condi-tion the generation process on a single specific modality, such as explicitpose trajectories or reference videos, limiting their ability to support het-erogeneous user inputs. To address this limitation, we present TriMotion,a modality-agnostic framework for camera-controlled video generationthat maps video, pose, and text inputs, describing the same cameratrajectory into a shared motion embedding space. |
Seunghyun Shin; Song Jifei; Wooseok Jeon; Hae-Gon Jeon; Jiankang Deng; |
| 108 | Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this,we present Deform360, a large-scale visuotactile dataset featuring 198daily-life objects, 1,980 interaction sequences, and over 215 hours of ob-servations from 41 surround-view cameras and bimanual tactile grippersto capture both global motion and contact-induced local deformations.Leveraging a novel markerless visuotactile 3D tracking pipeline to ex-tract dense geometry and motion, we systematically evaluate currentstate-of-the-art world models, comparing 2D video models against 3Dparticle models. |
Hongyu Li; Wanjia Fu; Xiaoyan Cong; Zekun Li; Binghao Huang; Hanxiao Jiang; Xintong He; Yiqing Liang; Rao Fu; Tao Lu; Srinath Sridhar; Kevin Smith; George Konidaris; Yunzhu Li; |
| 109 | Fast Spatial Memory with Scalable Elastic Test-Time Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose ElasticTest-Time Training inspired by elastic weight consolidation, that sta-bilizes LaCT fast-weight updates with a Fisher-weighted elastic prioraround a maintained anchor state. The anchor evolves as an exponentialmoving average of past fast weights to balance stability and plasticity.Based on this updated architecture, we introduce Fast Spatial Memory(FSM), an efficient and scalable model for 4D reconstruction that learnsspatiotemporal representations from long observation sequences and ren-ders novel view-time combinations. |
Ziqiao Ma; Xueyang Yu; Haoyu Zhen; Yuncong Yang; Joyce Chai; Chuang Gan; |
| 110 | World Models for Learning Dexterous Hand-Object Interactions from Human Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Modeling dexterous hand–object interactions is challengingas it requires understanding how subtle finger motions influence theenvironment through contact with objects. While … |
Raktim Goswami; Amir Bar; David Fan; Tsung-Yen Yang; Gaoyue Zhou; Prashanth Krishnamurthy; Michael Rabbat; Farshad Khorrami; Yann LeCun; |
| 111 | Distill on A Diet: Efficient Knowledge Distillation Via Learnable Data Pruning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However,existing data pruning methods are not designed for KD: some introducesubstantial overhead (e.g., obtaining training dynamics through retrain-ing), while others rely on heuristic selection rules that fail to capture whatKD actually requires, often resulting in suboptimal subsets. To addressthese issues, we propose IF-Beta, an efficient data pruning frameworkthat combines influence function and a learnable sampling policy. |
Yifan Wu; Yiqi Wang; Xichen Ye; Wenjing Yan; Xiaoqiang Li; Cheng Jin; WEIZHONG ZHANG; Xiangyu Yue; |
| 112 | EchoVLA: Robotic Vision-Language-Action Model with Synergistic Declarative Memory for Mobile Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inthis work, we present EchoVLA, a memory-aware VLA model for mobilemanipulation. |
Min Lin; Xiwen Liang; Bingqian Lin; Jingzhi Liu; Zijian Jiao; Kehan Li; Ziang Yan; Yu Sun; Weijia Liufu; Yuhan Ma; Jiarui Hu; Yuecheng Liu; Shen Zhao; Yuzheng Zhuang; Xiaodan Liang; |
| 113 | SynVAR: Synergizing Spatial and Semantic Alignment in Visual Autoregressive Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose SynVAR, the first training-free enhance-ment framework specifically tailored for the VAR paradigm, which intro-duces a spatial-semantic collaborative control strategy to effectively sup-press propagation error and improve generation quality. |
Zhennan Chen; Tianxing Shi; Pengcheng Xu; Kepan Nan; Qian Wang; Zili Yi; Jian Yang; Ying Tai; |
| 114 | AutoCompass: Accurate Visual Localization on Public Maps By Learning from Weak Labels Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These models are trained on large-scale datasets ofgeo-referenced images, whose position and heading labels often containnoise that affects the trained models. To address this, we present Auto-Compass, a supervision approach for training neural map matchers frominaccurate absolute pose labels. |
Javier Tirado-Garín; Alan Savio Paul; Shuai Chen; Axel Barroso-Laguna; Tommaso Cavallari; Daniyar Turmukhambetov; Victor Adrian Prisacariu; Eric Brachmann; |
| 115 | The Sterkfontein Caves Dataset: A Novel View Rendering Challenge from The Cradle of Humankind Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce the challenging Sterkfontein Caves dataset comprising ten underground scenes from a UNESCO World Heritage Site, and use it to find a new simple baseline method that beats existing low-light reconstruction methods upon it. |
Ireton Liu; Brian Xu; Dominic Stratford; Steven James; Richard Klein; James Tompkin; |
| 116 | Video Generative Models As Geometry Learner Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent generative approaches to geometry estimation adaptpretrained image diffusion models and treat the task as image-conditionedgeneration. Leveraging off-the-shelf image diffusion models, they either(i) train task-specific geometry models (for depth and surface normal es-timation) independently, losing the opportunity of exploring the intrinsiccorrelation of these geometric targets, or (ii) jointly fine-tune modifiedimage diffusion backbones (e.g., altered self-attention), which typicallydemands substantial labeled data. |
Haosen Yang; Song Jifei; Zhensong Zhang; Xiatian Zhu; Jiankang Deng; |
| 117 | RL from Physical Feedback: Aligning Large Motion Models with Humanoid Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To bridge the gap between T2M and humanoid execution, we propose Reinforcement Learning from Physical Feedback (RLPF), a novel framework that integrates text-conditioned motion generation with motion-conditioned humanoid whole-body control. |
Junpeng Yue; Zepeng Wang; Jiangxing Wang; Yuxuan Wang; Yu Zhang; Xinrun Xu; Bin Cao; Sipeng Zheng; gang ding; Zongqing Lu; |
| 118 | VersaViT: Enhancing MLLM Vision Backbones Via Task-Guided Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address the question, we make the following con-tributions: (i) We identify that the vision encoders within MLLMs exhibitdeficiencies in their dense feature representations, as evidenced by theirsuboptimal performance on dense prediction tasks (e.g., semantic seg-mentation, depth estimation); (ii) We propose VersaViT, a well-roundedvision transformer that instantiates a novel multi-task framework for col-laborative post-training. |
Yikun Liu; Yuan Liu; Shangzhe Di; Haicheng Wang; Zhongyin Zhao; Le Tian; Zhou Xiao; Jie Zhou; Jiangchao Yao; Yanfeng Wang; Weidi Xie; |
| 119 | SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce SRUM, a self-rewardingpost-training framework directly applicable to existing UMMs of vari-ous designs. |
Weiyang Jin; Yuwei Niu; Jiaqi Liao; Chengqi Duan; Aoxue Li; Shenghua Gao; Xihui Liu; |
| 120 | Scaling Dense Prediction with Latent Decoding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we argue that strong denseprediction can be achieved by separating capacity from resolution: per-forming high-capacity reasoning in a compact latent space, and readingout to pixels with a lightweight operator. |
Xiaoyang Wu; Yixing Lao; Chengyao Wang; Senqiao Yang; Yujia Zhang; Hengshuang ZHAO; |
| 121 | GIDE: Unlocking Diffusion LLMs for Precise Training-Free Image Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Unlikecontinuous diffusion models, the discrete tokenization inherent in DLLMshinders the application of standard noise inversion techniques, often lead-ing to structural degradation during editing. In this paper, we introduceGIDE (Grounded Inversion for DLLM Image Editing), a unified frame-work designed to bridge this gap. |
Zifeng Zhu; Jiaming Han; Jiaxiang Zhao; Minnan Luo; Xiangyu Yue; |
| 122 | Reward Modeling for Computer-Using Agent from Video Execution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we study rewardmodeling from execution video: a sequence of keyframes from an agenttrajectory that is independent of the agent’s internal reasoning or ac-tions. |
Linxin Song; Jieyu Zhang; Huanxin Sheng; Taiwei Shi; Rahul Gupta; Yang Liu; Ranjay Krishna; Jian Kang; Jieyu Zhao; |
| 123 | ConceptWeaver: Weaving Disentangled Concepts with Flow Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A final concept-insensitive Refinement Stage thensynthesizes fine-grained details. Guided by this discovery, we proposeConceptWeaver, a framework for one-shot concept disentanglement.ConceptWeaver learns concept-specific semantic offsets from a single ref-erence image using a stage-aware optimization strategy that aligns withthe three-stage framework. |
Jintao Chen; Aiming Hao; Xiaoqing Chen; Chengyu Bai; Chubin Chen; Yanxun Li; Jiahong Wu; Xiangxiang Chu; Shanghang Zhang; |
| 124 | Better Call CineCrew: Consistent Ultra-Long Narrative-to-Film Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a structured orchestration layer forfilm-oriented script-to-video generation, implemented as a multi-agentframework that operates between scripts and off-the-shelf video genera-tors. |
Jiaben Chen; Sixun Dong; Qinhong Zhou; Raine Ma; Zhiyang Dou; Wojciech Matusik; Chuang Gan; |
| 125 | OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introducesthe novel video-to-script (V2S) task, aiming to generate hierarchical,scene-by-scene scripts encompassing character actions, dialogues, expres-sions, and audio cues. To facilitate this, we construct a first-of-its-kindhuman-annotated benchmark and propose a temporally-aware hierarchi-cal evaluation framework. |
JUNFU PU; Yuxin Chen; Teng Wang; Ying Shan; |
| 126 | From Macro to Micro: Benchmarking Microscopic Spatial Intelligence on Molecules Via Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introduces the concept of Microscopic Spatial In-telligence (MiSI), the capability to perceive and reason about the spatialrelationships of invisible microscopic entities, which is fundamental toscientific discovery. To assess the potential of Vision-Language Models(VLMs) in this domain, we propose a systematic benchmark frameworkMiSI-Bench. |
Zongzhao Li; Xiangzhe Kong; Jiahui Su; Zongyang Ma; Mingze Li; Songyou Li; Yuelin Zhang; Yu Rong; Tingyang Xu; Deli Zhao; Wenbing Huang; |
| 127 | Concept-as-Tree: A Controllable Synthetic Data Framework Makes Stronger Personalized VLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recently, there hasbeen an increasing interest in improving the personalization capabilitiesof VLMs. To better integrate user-provided concepts into VLMs, manymethods use positive and negative samples to fine-tune these models.However, the scarcity of user-provided positive samples and the low qual-ity of retrieved negative samples pose challenges for existing techniques.To reveal the relationship between sample and model performance, wesystematically investigate the amount and diversity impact of positiveand negative samples (easy and hard) on VLM personalization tasks.Based on the detailed analysis, we introduce Concept-as-Tree (CaT),which represents a concept as a tree structure, thereby enabling the datageneration of positive and negative samples with varying difficulty anddiversity, and can be easily extended to multi-concept scenarios. |
Ruichuan An; Kai Zeng; Ming Lu; Sihan Yang; Renrui Zhang; Huitong Ji; Hao Liang; Wentao Zhang; |
| 128 | NUN: Nested Unfolding Network for Real-World Concealed Object Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A bi-directional unfolding interaction mechanismuses IQA to select optimal DeRUN outputs, and a cross-stage consis-tency loss ensures robust predictions under varying restoration quality.Theoretically, under local assumptions, we show that NUN reduces di-rect parameter-level gradient conflict through disjoint parameter setsand achieves a degradation-sensitivity bound, where degradation affectssegmentation only through the inner-loop optimization and restorationapproximation errors. |
Chunming He; Rihan Zhang; Longxiang Tang; Dingming Zhang; Bojian Zhang; Fengyang Xiao; Jingjia Feng; Sina Farsiu; |
| 129 | When Specialists Meet Generalists: Segmenter-Coordinated Asymmetric Learning for Label-Deficient Concealed Object Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, we observe that standard homogeneous co-training cannot exploit this complementarity because architecturally identical networks tend to share similar error patterns, especially when targets are heavily concealed. To address this, we present SCALER (Segmenter-Coordinated Asymmetric LEaRning), a framework that jointly optimizes a mean-teacher segmenter and a learnable SAM through two alternating phases with model-specific optimization strategies. |
Chunming He; Dingming Zhang; Longxiang Tang; Ziyun Yang; Fengyang Xiao; Sina Farsiu; |
| 130 | RoboMirror: Understand Before You Imitate for Video to Humanoid Locomotion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Text-to-motion methods suffer from semanticsparsity and staged pipeline errors, while video-based approaches onlyperform mechanical pose mimicry without genuine visual understanding.We propose RoboMirror, the first retargeting-free video-to-locomotionframework embodying “understand before you imitate”. |
Zhe Li; Boan Zhu; Yangyang Wei; Shuanghao Bai; Yuheng Ji; Tao Huang; Pengwei Wang; Zhongyuan Wang; Gary Chan; Chang Xu; Cheng Chi; Jianfei Yang; Shanghang Zhang; |
| 131 | Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While 10B-level industrial foundation models have pushedthe boundaries of image inpainting, their prohibitive computational costsseverely hinder practical deployment. Constructing a highly optimizedtask-specific specialist offers a promising solution; however, extreme struc-tural compression inevitably triggers a severe representation bottleneck.To conquer this, we propose Moebius, a highly efficient lightweight in-painting framework. |
Kangsheng Duan; Ziyang Xu; Wenyu Liu; Xiaohu Ruan; Xiaoxin Chen; Xinggang Wang; |
| 132 | Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models As Embodied Brain Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: For planning,RoboBench introduces an evaluation framework that uses an MLLM as aworld simulator. |
Yulin Luo; Chun-Kai Fan; Menghang Dong; Jiayu Shi; Xiangju Mi; Mengdi Zhao; Bo-Wen Zhang; Cheng Chi; Jiaming Liu; Gaole Dai; Rongyu Zhang; Ruichuan An; Kun Wu; Zhengping Che; shaoxuan Xie; Guocai Yao; Zhongxia Zhao; Pengwei Wang; Guang Liu; Zhongyuan Wang; Tiejun Huang; Shanghang Zhang; |
| 133 | Autoregressive Image Generation Needs Only A Few Lines of Cached Tokens Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce LineAR, a novel, training-free progressive key-value (KV) cache compression pipeline for AR image generation. |
Ziran Qin; Youru Lv; Mingbao Lin; Zeren Zhang; chaofan gan; Tieyuan Chen; Liquan Shen; Junhui Hou; Chern Hong Lim; Fei Wen; Weiyao Lin; |
| 134 | AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Given aninput image and a textual query, AutoV automatically locates the mostsuitable visual prompt from a diverse candidate pool. Training such a re-trieval framework requires prompt-level supervision, yet prompt qualityis inherently ambiguous and difficult to assess reliably, even for humans.To enable automatic supervision, we evaluate visual prompts using a pre-trained LVLM and label them according to their prediction losses. |
Yuan Zhang; Chun-Kai Fan; Sicheng Yu; Junwen Pan; Tao Huang; Ming Lu; Kuan Cheng; Qi She; Shanghang Zhang; |
| 135 | VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing instruction datasetsoften contain a substantial portion of visually redundant samples (solv-able from text alone), as well as multimodally misaligned supervisionthat can degrade learning. To address this, we propose VisNec (VisualNecessity Score), a principled data selection framework that measuresthe marginal contribution of visual input during instruction tuning. |
Mingkang Dong; Hongyi Cai; jie li; Sifan Zhou; Bin Ren; Kunyu Peng; Yuqian Fu; |
| 136 | Focusing By Contrastive Attention: Enhancing VLMs’ Visual Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Vision-Language Models (VLMs) have demonstrated remark-able success across diverse visual tasks, yet their performance degradesin complex visual environments. |
Yuyao Ge; Shenghua Liu; Yiwei Wang; Lingrui Mei; Baolong Bi; Xuanshan Zhou; Jiayu Yao; Jiafeng Guo; Xueqi Cheng; |
| 137 | Learning Accurate Segmentation Purely from Self-Supervision Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In thiswork, we introduce Selfment, a fully self-supervised framework that seg-ments foreground objects directly from raw images without human labels,pretrained segmentation models, or any post-processing. |
Zuyao You; Zuxuan Wu; Yu-Gang Jiang; |
| 138 | Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Through extensive experiments, 013014 we demonstrate that separately encoding text and visual instructions, 014015 combined with co-training using shallow feature control, is essential to 015016 mitigate task ambiguity and enhance multi-task generalization. |
Yuandong Pu; Le Zhuo; Kaiwen Zhu; Liangbin Xie; Wenlong Zhang; Xiangyu Chen; Peng Gao; Yu Qiao; Chao Dong; Yihao Liu; |
| 139 | DeRA: Decoupled Representation Alignment for Video Tokenization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents DeRA, a novel 1D video tokenizer thatdecouples the spatial-temporal representation learning in video tokeniza-tion to achieve better training efficiency and performance. |
Pengbo Guo; Junke Wang; Zhen Xing; Chengxu Liu; Daoguo Dong; Xueming Qian; Zuxuan Wu; |
| 140 | Reconstructing Humans and Objects in Interaction Using Large Reconstruction Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, weexplore a different avenue. |
Agniv Chatterjee; Georgios Pavlakos; |
| 141 | ActionParty: Multi-Subject Action Binding in Generative Video Games Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we tackle a fundamental issue of action binding in existing video diffusion models, which struggle to associate specific actions with their corresponding subjects. |
Alexander Pondaven; Ziyi Wu; Igor Gilitschenski; Philip Torr; Sergey Tulyakov; Fabio Pizzati; Aliaksandr Siarohin; |
| 142 | Free‑CD: Probabilistically Decoupled Training-Free Open-Vocabulary Change Detection with Resolution-Invariant Feature Inversion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Open-Vocabulary Change Detection (OVCD) faces a funda-mental granularity gap: foundation models prioritize high-level semanticabstraction while change detection requires pixel-level spatial fidelity.Traditional OVCD methods rely on instance extraction models, intro-ducing spatial semantic ambiguities and leading to over-segmentationor under-segmentation in remote sensing scenarios. To bridge this gap,we propose Free-CD, a training-free OVCD framework that reformu-lates the task by predicting a change probability distribution rather thanenforcing binary change masks via instance boundaries. |
Yongshuo Zhu; LU LI; Keyan Chen; Zhenwei Shi; ZHOU Fugen; |
| 143 | ReinDriveGen: Reinforcement Post-Training for Out-of-Distribution Driving Scene Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present ReinDriveGen, a framework that enables full con-trollability over dynamic driving scenes, allowing users to freely edit actortrajectories to simulate safety-critical corner cases such as front-vehiclecollisions, drifting cars, vehicles spinning out of control, pedestrians jay-walking, and cyclists cutting across lanes. |
Hao ZHANG; Lue Fan; Weikang Bian; Zehuan Wu; Lewei Lu; Zhaoxiang Zhang; Hongsheng LI; |
| 144 | OmniX: From Unified Panoramic Generation and Perception To Graphics-Ready 3D Scenes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we advance this technique to generategraphics-ready 3D scenes suitable for physically based rendering (PBR),relighting, and simulation.Furthermore, we construct a large-scale syn-thetic panorama dataset comprising high-quality multimodal panoramasfrom diverse indoor and outdoor scenes. |
Yukun Huang; Jiwen Yu; Yanning Zhou; Jianan Wang; Xintao Wang; Pengfei Wan; Xihui Liu; |
| 145 | VideoSfM: Exploiting Temporal Structure for Video-Based Structure-from-Motion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The dominant approaches tothis problem are severely limited: Simultaneous Localization and Map-ping (SLAM) is sensitive to initialization and transient failures due to itscausal, incremental nature; it is often over-optimized for real-time oper-ation and generally requires known camera calibration; while Structure-from-Motion (SfM) typically forgoes any image ordering, enabling opti-mal initialization and global optimization, but lacks robustness to visualsymmetries and extreme motions. To bridge this gap, we introduce asystem that combines the strong sequential constraints of SLAM withthe flexibility and global optimization of offline SfM, enabling the met-ric reconstruction of arbitrary, long, uncalibrated videos. |
Zador Pataki; Paul-Edouard Sarlin; Marc Pollefeys; |
| 146 | Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To rigorously assess scien-tific correctness, we introduce SciGenBench, which evaluates generatedimages based on information utility and logical validity. |
honglin lin; Chonghan Qin; Zheng Liu; Qizhi Pei; Yu Li; Zhanping Zhong; Xin Gao; Wei Li; Wentao Zhang; Yanfeng Wang; Conghui He; Lijun Wu; |
| 147 | Multiple Images Distract Large Multimodal Models Via Attention Fragmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: By applying Pinsker’s inequality, we establish a theoretical bound showing that this high-entropy dispersion, combined with stronger early-image sinks, strictly reduces the usable non-sink attention available to earlier images, providing a mechanistic link to image order sensitivity. Motivated by this diagnosis, we propose Attention Remasking (AR), a post-training edit that blocks sink keys and opens a sparse set of cross-image links, routing the recovered attention to task-relevant tokens. |
Tingrui Qiao; Di Zhao; Yuzhuo Li; Bo Pang; Caroline Walker; Chris Cunningham; Yun Sing Koh; |
| 148 | Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Online Video Large Language Models (VideoLLMs) playa critical role in supporting responsive, real-time interaction. Existingmethods focus on streaming perception, lacking a … |
Yiran Guan; Liang Yin; Dingkang Liang; Jianzhong Ju; Zhenbo Luo; Jian Luan; Yuliang Liu; Xiang Bai; |
| 149 | Roam2Room: A Unified Floorplan-to-Furnished Framework for Controllable Indoor Scene Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing meth-ods often rely on hand-crafted rules or focus on isolated sub-tasks (e.g.,floorplan synthesis or single-room furnishing), producing whole-homescenes that lack global coherence, realism, and simulation readiness. Tomitigate these limitations, we propose a unified hierarchical frameworkthat decomposes indoor scene synthesis into controllable stages. |
Wenbo Li; Zipeng Qin; Xiaoliang Ju; Rongyao Fang; Hongsheng LI; |
| 150 | Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in The Wild Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Real5- OmniDocBench, the first benchmark to provide a full-scale, one-to-one physical reconstruction of the complete OmniDocBench v1.5 test set (1,355 images) across five critical real-world scenarios: Scanning, Warping, Screen-Photography, Illumination, and Skew. |
Cheng Cui; Changda Zhou; Tingquan Gao; Xueqing Wang; ZIYUE GAO; Jing Tang; Yi Liu; |
| 151 | RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in The Wild Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present RT-DocLayout, a highly efficient end-to-end framework for document layout analysis, designed as a front-end for document parsing tasks. |
Cheng Cui; Tingquan Gao; Xueqing Wang; Changda Zhou; Hongen Liu; Ting Sun; Yubo Zhang; Zelun Zhang; Jiaxuan Liu; Manhui Lin; Yue Zhang; Suyin Liang; Yiqing Xiang; Yi Liu; |
| 152 | Direct Autoregressive Diffusion Distillation Via Error-aware Causal Pretraining Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: During training, Error Forcing injects controlled residualerrors into the conditioning context, approximating the inference-timehistory distribution. This enables the causal model to learn from de-graded yet clean contexts, improving robustness to accumulated errorswithout sacrificing parallel training throughput. |
Jiaxing Li; Kaichen Huang; Baixin Xu; Zexiang Liu; Xianglong He; Zile Wang; Junyao Gao; Yang Liu; Ying He; Bo An; Yangguang Li; |
| 153 | MinerU-Diffusion: Rethinking Document OCR As Inverse Rendering Via Diffusion Decoding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we revisitdocument OCR from an inverse rendering perspective, arguing that left-to-right causal generation is an artifact of serialization rather than anintrinsic property of the task. |
Hejun Dong; Junbo Niu; Bin Wang; Weijun Zeng; Wentao Zhang; Conghui He; |
| 154 | G-ZAP: A Generalizable Zero-Shot Framework for Arbitrary-Scale Pansharpening Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Pansharpening aims to fuse a high-resolution panchromatic(PAN) image and a low-resolution multispectral (LRMS) image to pro-duce a high-resolution multispectral (HRMS) image. Recent deep modelshave achieved strong performance, yet they typically rely on large-scalepretraining and often generalize poorly to unseen real-world image pairs.Prior zero-shot approaches improve real-scene generalization but requireper-image optimization, hindering weight reuse, and the above methodsare usually limited to a fixed scale. |
Zhiqi Yang; Shan Yin; Jingze Liang; Liang-Jian Deng; |
| 155 | Generative Refinement Network for Visual Synthesis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In con-trast, autoregressive (AR) models are inherently complexity-aware, asevidenced by their variable likelihoods, but are often hindered by lossydiscrete tokenization and error accumulation. In this work, we introduceGenerative Refinement Networks (GRN), a next-generation visual syn-thesis paradigm to address these issues. |
Jian Han; Jinlai Liu; Jiahuan Wang; BINGYUE PENG; Zehuan Yuan; |
| 156 | Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Both paradigms perform routing onlyafter a complete output, and ignore whether the target model can ac-tually solve the routed instances. To address this, we propose PRP,a Proactive Routing Paradigm that enables early decision-making byjointly evaluating the competence of both the draft and target models.Our Draft Rating Learning (DRL) equips the draft model with an inter-nal confidence estimator, while Joint Rating Learning (JRL) predictshow well the target model can handle a given query, thereby prioritiz-ing the allocation of samples it excels at rather than the hardest ones.These ratings enable fine-grained, instance-level Proactive Routingand substantially accelerate inference without compromising overall per-formance. |
Yinan ZHOU; Haokun Lin; Yichen Wu; Yuxin Chen; Teng Wang; Caifeng Shan; Zhenan Sun; Chen Ma; Li Zhu; Ying Shan; |
| 157 | PhenoLIP: Phenotype Guided Medical Vision–Language Pretraining Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Recent progress in CLIP-like vision-language models (VLMs)has greatly advanced medical image analysis. However, most existingmedical VLMs still rely on coarse image-text … |
Cheng Liang; Chaoyi Wu; Weike Zhao; Ya Zhang; Yanfeng Wang; Weidi Xie; |
| 158 | SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study answer-span prediction entropy as a model-internalfeedback signal and show that naive entropy minimization is ambiguous,since low entropy may arise from evidence-grounded confidence or short-cut collapse. To resolve this ambiguity, we introduce low-entropy anchorsand an entropy-shaping objective that reduces answer uncertainty whilepreserving baseline high-confidence tokens. |
Bo Yin; Xiaobin Hu; Chengming Xu; Ruolin Shen; Mo Yang; Jiangning Zhang; Peng-Tao Jiang; Cheng Tan; Shuicheng Yan; |
| 159 | Sentinel: Embodied Cooperative Spatial Reasoning and Planning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we study Cooperative Spatial Intelligence, the ability of decentralized embodied agents to coordinate effectively under dynamic environmental constraints across city-scale outdoor domains.We introduce Sentinel Challenge, a benchmark where multiple decentralized embodied agents must communicate in natural language to agree on a mutually safe and convenient meeting point within large, city-scale outdoor environments. |
Xiangye Lin; Hongxin Zhang; Ruxi Deng; Qinhong Zhou; Chuang Gan; |
| 160 | VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To preserve the pretrained ca-pacity while improving geometric consistency, we propose VGGRPO(Visual Geometry GRPO), a latent geometry-guided framework forgeometry-aware video post-training. |
Zhaochong An; Orest Kupyn; Théo Uscidda; Andrea Colaco; Karan Ahuja; Serge Belongie; Mar Gonzalez Franco; Marta Gazulla; |
| 161 | Visual Spatial Tuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To enhance the spatial ability within general architectures, we introduce Visual Spatial Tuning (VST), a comprehensive framework to cultivate VLMs with human-like visuospatial competence spanning spatial perception and reasoning. |
Rui Yang; ziyu zhu; Yanwei Li; Jingjia Huang; Shen Yan; Siyuan Zhou; Zhe Liu; Xiangtai Li; Shuangye Li; Wenqian Wang; Yi Lin; Hengshuang ZHAO; |
| 162 | Anatomy of A Lie: A Multi-Stage Diagnostic Framework for Tracing Hallucinations in Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Vision-Language Models (VLMs) frequently `hallucinate’_x0015_generate plausible yet factually incorrect statements_x0016_posing a criticalbarrier to their trustworthy deployment. In this work, we propose a newparadigm for diagnosing hallucinations, recasting them from static out-put errors into an auditable process-level anomaly. |
Lexiang Xiong; QI LI; Jingwen Ye; Xinchao Wang; |
| 163 | Beyond Imitation: Learning Safe End-to-End Autonomous Driving from Hard Negatives Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address thislimitation, we propose BeyondDrive, a failure-aware imitation learn-ing framework that jointly learns from successful and failed driving be-haviors. |
Junli Wang; HuaZhihua HuaZhihua; Xueyi Liu; Zebin Xing; Wei Zhang; Kun Ma; Guang Chen; Hangjun Ye; Long Chen; Pengxuan Yang; |
| 164 | Pix2NPHM: Learning to Regress NPHM Reconstructions From A Single Image Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To thisend, we propose Pix2NPHM, a vision transformer (ViT) network that di-rectly regresses NPHM parameters, given a single image as input, finallybridging the gap from theoretical to practical use. |
Simon Giebenhain; Tobias Kirschstein; Liam Schoneveld; Davide Davoli; Zhe Chen; Matthias Niessner; |
| 165 | RhymeFlow: Training Free Acceleration for Video Generation with Asynchronous Denoising Flow Scheduling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we introduce RhymeFlow, a training-free framework that decouples the denoising trajectories of different frames. |
Chensheng Dai; Shengjun Zhang; Yifan Li; Zhang Zhang; Zheng Zhu; Yueqi Duan; |
| 166 | Flash-DD: An Ultra Parameter-Efficient Approach to Dataset Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically,we propose a DD-oriented model parameter reduction method that au-tomatically determines the optimal capacity of teacher models and elim-inates redundant parameters for dataset distillation tasks. |
Ruonan Yu; Songhua Liu; Xinchao Wang; |
| 167 | MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing benchmarks predomi-nantly focus on elementary, single-hop relations, neglecting the multi-hopcompositional reasoning and precise visual grounding essential for real-world scenarios. To address this, we introduce MultihopSpatial, o!er-ing three key contributions: (1) A comprehensive benchmark designedfor multi-hop and compositional spatial reasoning, featuring 1- to 3-hopcomplex queries across diverse spatial perspectives. |
Youngwan Lee; Soojin Jang; Yoorhim Cho; Seunghwan Lee; Yong-Ju Lee; Sung Ju Hwang; |
| 168 | SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The datasetis hierarchically organized according to the semiotic dimensions andincorporates a Scientific Reasoning Chain-of-Thought (Sci-RCoT) toexplicitly model underlying visual logic. For evaluation, we propose SciIR-Bench, which aligns with these three semiotic levels and employs anAtomic Checklist to convert the outcome-oriented scientific accuracyinto process-oriented, verifiable, fine-grained questions. |
Zhiyuan Ma; Zhengfeng Shi; Yuning An; Peize Li; Jiabao Wei; Ruijie Li; Junhao Xiao; Jianjun Li; Bowen Zhou; |
| 169 | RADmesh: Remesh-Aware Mesh Deformation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a remeshing-enhanced method for generativelydeforming shapes with visual losses. |
Nam Anh Dinh; Itai Lang; Oded Stein; Rana Hanocka; |
| 170 | S-VAM: Shortcut Video-Action Model By Self-Distilling Geometric and Semantic Foresight Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, current VAMs, typically relying on either slow multi-step video generation or noisy one-step feature extraction, cannot simultaneously guarantee real-time inference and high-fidelity foresight. To address this limitation, we propose S-VAM, a shortcut video-action model that foresees coherent geometric and semantic representations via a single forward pass. |
Haodong Yan; Zhide Zhong; Jiaguan Zhu; Junjie He; Weilin Yuan; Wenxuan Song; Xin Gong; Yingjie CAI; Guanyi Zhao; Xu Yan; Liu Bingbing; Yingcong Chen; Haoang Li; |
| 171 | Finite Difference Flow Optimization for RL Post-Training of Text-to-Image Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, wepropose an online RL variant that reduces the variance in the modelupdates by sampling paired trajectories and pulling the flow velocityin the direction of the more favorable image. |
David McAllister; Miika Aittala; Tero Karras; Janne Hellsten; Angjoo Kanazawa; Timo Aila; Samuli Laine; |
| 172 | Towards High-Resolution Visual Perception Via Hierarchical Entity Exploration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose HierarchicalEntity Exploration (HEE), a training-free and model-agnostic frameworkthat transforms static image understanding into dynamic, query-guidedentity exploration. |
Ziyu Ma; Shidong Yang; Yuxiang Ji; Yiming Hu; Tongwen Huang; Yong Wang; Jianfei Cai; Xiangxiang Chu; |
| 173 | MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present MGM-Omni, an Omni MLLM for omni-modalunderstanding and expressive, long-horizon speech generation. |
Chengyao Wang; Zhisheng Zhong; Bohao PENG; Senqiao Yang; Yuqi Liu; Haokun GUI; Bin Xia; Jingyao Li; Bei Yu; Jiaya Jia; |
| 174 | LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present LM-GenDrive, the first framework that unifies LLM-based multimodal un-derstanding with generative world models for end-to-end closed-loop au-tonomous driving. |
Hao Shao; Letian Wang; Yang Zhou; Yuxuan Hu; Zhuofan Zong; Steven Waslander; Wei Zhan; Hongsheng LI; |
| 175 | Temporal and Cross-modal Alignment for Enhanced Audiovisual Video Captioning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Most existing approaches suffer from modality detachment and temporal incoherence, failing to accurately bind auditory events to visual entities or capture complex causal dynamics. To address these deficiencies, we propose TCA-Captioner, a framework specifically engineered to enhance Temporal and Cross-Modal Alignment for audiovisual video captioning. |
Chen Zhao; Jiajun Ma; Qilong Huang; Tiehan Fan; Hongyu Li; Zhuoliang Kang; Xiaoming Wei; Jian Yang; Ying Tai; |
| 176 | CameraAnything: Refilming Videos with Arbitrary Camera Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce CameraAnything, the first unified frame-work for camera controlled video editing that enables joint control ofboth intrinsic and extrinsic camera parameters. |
Yixuan Li; Yanhong Zeng; Ka Leong Cheng; Jiayi Zhu; Hanlin Wang; Wen Wang; Yihao Meng; Hao Ouyang; Qiuyu Wang; Yue Yu; Zidong Wang; Yiyuan Zhang; Yujun Shen; Dahua Lin; |
| 177 | Learning Active Perception for Pixel-Space Reasoning Via Visual-Intent Stratified GRPO Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This leads to visual laziness: the policy does not explore active multi-zoom perception strategies and collapses into a zoom-averse state. To address this, we propose Visual-Intent Stratified GRPO (VIS-GRPO), a drop-in replacement for GRPO that restores fair learning signals for active perception. |
Mingkang Zhu; Xi Chen; Senqiao Yang; Bei Yu; Hengshuang ZHAO; Jiaya Jia; |
| 178 | SuperFlex: Deformable Superquadrics for Point Cloud Decomposition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work,we present SuperFlex, an enhanced framework that expands the expres-sive power and applicability of superquadric decompositions. |
Gabriel Tavernini; Elisabetta Fedele; Tiago Novello; Leonidas Guibas; Marc Pollefeys; Francis Engelmann; |
| 179 | GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We instead explore a fully autoregressive alternative and introduce GaussianGPT, a transformer-based model that directly generates 3D Gaussians via next-token prediction, thus facilitating full 3D scene generation. |
Nicolas von Lützow; Barbara Roessle; Katharina Schmid; Matthias Niessner; |
| 180 | ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, wepropose a unified understanding-and-generation framework that lever-ages world modeling to simultaneously enable meaningful explorationand provide dense supervision. |
Zihao Sheng; Xin Ye; Jingru Luo; Sikai Chen; Liu Ren; |
| 181 | Personalize Your Large Vision-language Models With In-context Prompt Tuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These limitations restrict the broaderdeployment of LVLM-based systems. Therefore, this paper proposes in-context prompt tuning (ICPT). |
Yanshu Li; Jiaqian Li; Kuai Yu; Xi Xiao; Dongfang Liu; Tianyang Wang; Ruixiang Tang; |
| 182 | Bridging Visual Representation and Reinforcement Learning from Verifiable Rewards in Large Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The method adaptively localizessemantically salient regions through hierarchical geometric aggregation,identifies vision-critical attention heads via structured attribution, andperforms paragraph-level credit reallocation to align spatial visual ev-idence with semantically decisive reasoning steps. |
yuhang han; Yuyang Wu; Zhengbo Jiao; Yiyu Wang; Xuyang Liu; Shaobo Wang; Hanlin xu; Xuming Hu; Linfeng Zhang; |
| 183 | Comprehensive Language–image Pre-training for 3D Medical Image Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As a consequence, natural-image VLE recipes do not directly transfer to 3D medical imaging. In this paper, we overcome these challenges by injecting additional supervision via a report generation objective and combining vision-language with vision-only pre-training, allowing us to leverage both image-only and paired image-text 3D datasets. |
Tassilo Wald; Ibrahim Ethem Hamamci; Yuan Gao; Sam Bond-Taylor; Harshita Sharma; Maximilian Ilse; Cynthia Lo; Olesya Melnichenko; Anton Schwaighofer; Noel Codella; Maria Teodora Wetscherek; Klaus Maier-Hein; Panagiotis Korfiatis; Valentina Salvatelli; Javier Alvarez-Valle; Fernando Pérez-García; |
| 184 | ICDepth: Taming Video Diffusion Models for Video Depth Estimation Via In-Context Conditioning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Monocular video depth estimation requires temporal con-sistency, geometric accuracy, and generalization across diverse scenar-ios—yet existing methods struggle to achieve all three simultaneously.Discriminative models excel at per-frame accuracy but suffer from tem-poral drift due to limited context windows, while generative methodsimprove consistency and generalization at the cost of extensive train-ing data (10M+ samples) and lack of geometric precision. In responseto these issues, we introduce ICDepth, a framework that adapts pre-trained text-to-video diffusion transformers for video depth estimationvia In-Context Conditioning (ICC), leveraging their rich spatial-temporalpriors. |
Xuanhua He; JIAXIN XIE; Mingzhe Zheng; Qifeng Chen; |
| 185 | Enhancing Prompt-image Alignment Evaluations Via Cyclic Mutual Information Maximization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Unlike exist-ing methods, our framework focuses on promoting e_x001B_ective multimodalinformation fusion in the deep feature domain, which has cyclic mutualinformation maximization phases. |
Xingran Liao; Duanyu Feng; Mingliang Zhou; Sam Kwong; Weisi Lin; |
| 186 | Decoupled Illumination Priors for Spatially Controllable Multi-View Indoor Scene Relighting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Lume-Palette, a progressive framework that leveragessemantic lighting priors for spatially controllable multi-view indoor re-lighting. |
Chenjian Gao; Linning Xu; Tianfan Xue; |
| 187 | JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present JoVA, a streamlined frameworkthat unifies joint video-audio generation and editing.To fully empower and systematically evaluate this frame-work, we construct a comprehensive training corpus encompassing video-audio generation and editing datasets, and introduce unified benchmarkstailored for these multimodal tasks. |
Xiaohu Huang; Haoyang He; Hao Zhou; Qiangpeng Yang; Min Zheng; Kai Han; |
| 188 | DisRM: Reward Modeling As Discriminative Prediction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Reward models are central to post-training and test-timeoptimization for visual generative models, yet existing approaches typi-cally rely on either large-scale pairwise … |
Runtao Liu; Jiahao Zhan; Yuxuan GUO; Yingqing He; Chen Wei; Alan Yuille; Qifeng Chen; |
| 189 | VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this gap, we leverage physically grounded simulation to construct diverse, high-fidelity 3D scenes with precise per-view metadata, enabling scalable data generation that remains transferable to real-world settings. Based on this engine, we introduce VIEW2SPACE, a multi-dimensional benchmark for sparse multi-view reasoning, together with a scalable, disjoint training split supporting millions of grounded question–answer pairs. |
Fucai Ke; Zhixi Cai; Boying Li; Long Chen; Beibei Lin; Weiqing Wang; Pari Delir Haghighi; Gholamreza Haffari; Hamid Rezatofighi; |
| 190 | UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose UniRec-0.1B, a unified recognition model with only 0.1B parameters.Finally, we develop a comprehensive evaluation benchmark covering Chinese and English documents from multiple domains and with multiple levels. |
Yongkun Du; Zhineng Chen; Yazhen Xie; Weikang Bai; Hao Feng; Wei Shi; Yuchen Su; Can Huang; Yu-Gang Jiang; |
| 191 | Towards Purified Multi-Label Test-Time Adaptation of Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While introducing region-level cues helps isolate class-specific evidence, such regional evidence can also be unreliable under distribution shifts, making its identification and utilization non-trivial. To address these issues, we introduce PuRF, a novel PuRiFication-driven cache-based method for multi-label test-time adaptation of vision-language models. |
Yiwen Liang; Hui Chen; Yizhe Xiong; Mengyao Lyu; Yuhan Cao; Zijia Lin; SHUAICHENG NIU; Sicheng Zhao; Jungong Han; Guiguang Ding; |
| 192 | Make Geometry Matter for Spatial Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose GeoSR, a framework designed to make geometry matter by encouraging VLMs to actively reason with geometry tokens. |
Shihua Zhang; Qiuhong Shen; Shizun Wang; Tianbo Pan; Xinchao Wang; |
| 193 | VGEdit: Unlocking Video Generation Priors for Reasoning-Informed Image Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: (2) Reinforcement learning: we adapt the video model with reward-basedoptimization, treating intermediate frames as implicit reasoning chainsand the last frame as the target edited image. |
Haiquan Lu; Gongfan Fang; Xinyin Ma; Xinchao Wang; |
| 194 | GKDT: General Keypoint Detection Transformer Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Based on MegaKPT, we develop GKDT, a simple, flexible and powerful DINOv3 based Transformer model for General Keypoint Detection.To this end, we firstly present a largescale unified keypoint dataset called MegaKPT. |
Changsheng Lu; Yuxin Chen; Haokun GUI; Rong Wang; Jie Yang; Harry Yang; Anton van den Hengel; Jiaya Jia; |
| 195 | Reinforcing Video Reasoning with Focused Thinking Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: Recent advancements in reinforcement learning, particularlythrough Group Relative Policy Optimization (GRPO), have significantlyimproved multimodal large language models for complex reasoning tasks.However, two critical limitations persist: 1) they often produce unfo-cused, verbose reasoning chains that obscure salient spatiotemporal cues,and 2) binary rewarding fails to account for partially correct answers,resulting in high reward variance and ine!cient learning. In this pa-per, we propose TW-GRPO, a novel framework that enhances visualreasoning with focused thinking and dense reward granularity. |
Jisheng Dang; Jingze Wu; Teng Wang; Xuanhui Lin; Nannan Zhu; Hongbo Chen; WEISHI ZHENG; Meng Wang; Tat-Seng Chua; |
| 196 | VoCa: Unified Autoregressive Modeling for Talking Audio-Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present VoCa, a unified autoregressive frameworkthat jointly generates speech and talking-head video from text transcriptsand a reference image. |
Zhuofan Zong; Jiale Yuan; Yufei Liu; Dongzhi Jiang; Hao Shao; Zimu Lu; Ke Wang; Yunqiao Yang; Mingjie Zhan; Hongsheng LI; |
| 197 | EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce EgoVITA, a framework that decomposes egocentric video reasoning into a structured plan-then-verify process. |
Yogesh Kulkarni; Pooyan Fazli; |
| 198 | Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we present Vinci2, a proactive ego-centric assistance system that advances the on-device assistant Vincifrom reactive response toward proactivity. |
Sitong Gong; Tianyu Yan; Caixin Kang; Bo Zheng; Xiang Ruan; Huchuan Lu; Kaipeng Zhang; Yoichi Sato; Yifei Huang; |
| 199 | Attention Is Case-Sensitive Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inthis paper, we present a systematic empirical characterization study re-vealing that Large Language Models (LLMs) exhibit an analogous prop-erty: letter casing modulates internal attention allocation. |
Maximilian Dillitzer; Tin Stribor Sohn; Jason Corso; Michael Auerbach; |
| 200 | ObjectForesight: Predicting 3D Object Trajectories from Human Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce ObjectForesight, a 3Dobject-centric dynamics model that predicts future 6-DoF poses andtrajectories of rigid objects from short egocentric video sequences. |
Rustin Soraki; Homanga Bharadhwaj; Ali Farhadi; Roozbeh Mottaghi; |
| 201 | StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production–Living Simulations with Stardew Valley Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To bridge this gap,we introduce StarDojo, a novel benchmark based on Stardew Valley,designed to assess AI agents in open-ended production–living simulations.In StarDojo, agents are tasked to perform essential livelihood activitiessuch as farming and crafting, while simultaneously engaging in socialinteractions to establish relationships within a vibrant community.Addition-ally, we provide a compact subset of 100 representative tasks for efficientmodel evaluation. |
Weihao Tan; Changjiu Jiang; Yu Duan; Mingcong Lei; Li JiaGeng; Yitian Hong; Xinrun Wang; Bo An; |
| 202 | MolmoWeb: Open Visual Web Agent and Open Data for The Open Web Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we introduce (1) MolmoWebMix, alarge and diverse mixture of browser task demonstrations and web-GUIperception data and (2) MolmoWeb a family of fully open multimodalweb agents.We release model checkpoints, train-ing data, code, and a unified evaluation harness to enable reproducibilityand accelerate open research on web agents (GitHub). |
Tanmay Gupta; Piper Wolters; Zixian Ma; Peter Sushko; Rock Yuren Pang; Yue Yang; Jason Ren; Harsh Trivedi; Taira Anderson; Winson Han; Ranjay Krishna; |
| 203 | FoundYou: A Unified Model for Personalized Segmentation and Retrieval Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce FoundYou, a unified framework builton the observation that Segment Anything 2 (SAM 2), trained to pre-serve object identity across video frames, inherently captures instance-level cues. |
Gabriele Trivigno; Marcos Alfaro Perez; Claudia Cuttano; Gabriele Berton; Luis Payá; Carlo Masone; |
| 204 | GEM: Generative Supervision Helps Embodied Intelligence Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, a signifi-cant gap remains between the high-level semantic focus of standardtext-guided pre-training paradigms and the low-level spatial and phys-ical knowledge critical for execution in embodied environments. In thispaper, we introduce GEM, a Generative-supervised Embodied vision-language Model designed to bridge this divide. |
Ruowen Zhao; Bangguo Li; Zuyan Liu; Yinan Liang; junliang ye; FANGFU LIU; Diankun Wu; Zhengyi Wang; Xumin Yu; Yongming Rao; Han Hu; Jun Zhu; |
| 205 | Diffusion Model As A Generalized Segmentation Learner Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Diffusion models are primarily trained for image synthesis,yet their denoising trajectories encode rich, spatially aligned visual pri-ors. In this paper, we demonstrate that these priors can be utilized fortext-conditioned semantic and open-vocabulary segmentation, and thisapproach can be generalized to various downstream tasks to make ageneral-purpose diffusion segmentation framework. |
Haoxiao Wang; Antao Xiang; Haiyang Sun; Peilin Sun; Changhao Pan; Yifu Chen; Minjie Hong; Weijie Wang; Shuang Chen; Yue Chen; ZHOU ZHAO; |
| 206 | PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce PerceptionComp, a fully manually annotated benchmark designed so that no single moment is sufficient: answering requires evidence from multiple temporally separated segments under compositional constraints. |
Shaoxuan Li; Zhixuan Zhao; Hanze Deng; Zirun Ma; Shulin Tian; Zuyan Liu; Yushi Hu; Haoning Wu; Yuhao Dong; Benlin Liu; Ziwei Liu; Ranjay Krishna; |
| 207 | Learning Sample-wise Rank-Aware Interpolation Weights for Composed Visual Data Retrieval Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We instead revisit the efficacy ofsimple linear interpolation within an embedding space, and introduceSRAIN, the first framework that dynamically predicts query-specific in-terpolation weights. |
Boseung Jeong; Taegyu Park; Donghyeon Kwon; Hyunsouk Cho; Suha Kwak; |
| 208 | UniDriveDreamer: A Single-Stage Multimodal World Model for Autonomous Driving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose UniDriveDreamer, a single-stage unified multimodal world model for autonomous driving, which directly generates multimodal future observations without relying on intermediate representations or cascaded modules. |
Guosheng Zhao; Yaozeng Wang; Xiaofeng Wang; Zheng Zhu; Tingdong Yu; Guan Huang; Yongchen Zai; Ji Jiao; Changliang Xue; Xiaole Wang; Zhen Yang; Futang Zhu; Xingang Wang; |
| 209 | Locality-Aware Continual Unlearning for Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Locality-Aware Target Selec-tion chooses, for each forget prompt, the context-preserving mappingprompt that the diffusion model itself treats as most similar to the orig-inal prompt, measured by score-prediction distance (how differently themodel denoises the same noisy image under two text conditions), ensur-ing each update is as small and targeted as possible. |
Naveen George; Naoki Murata; Yuhta Takida; Konda Reddy Mopuri; Yuki Mitsufuji; |
| 210 | Any to Full: Prompting Depth Anything for Depth Completion in One Stage Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we present Any2Full, a one-stage,domain-general, and pattern-agnostic framework that reformulates com-pletion as a scale-prompting adaptation of a pretrained MDE model.To address varying depth sparsity levels and irregular spatial distribu-tions, we design a Scale-Aware Prompt Encoder. |
Zhiyuan Zhou; Ruofeng Liu; TAICHI LIU; Weijian Zuo; Shanshan Wang; Zhiqing Hong; Desheng Zhang; |
| 211 | HomeGuard: VLM-based Embodied Safeguard for Identifying Contextual Risk in Household Task Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this,we propose an architecture-agnostic safeguard featuring Context-GuidedChain-of-Thought (CG-CoT). |
Xiaoya Lu; Yijin Zhou; Zeren Chen; Ruocheng Wang; Bingrui Sima; Enshen Zhou; Lu Sheng; Dongrui Liu; Jing Shao; |
| 212 | Learn Once, Edit Anywhere: Visual Direction Transfer for Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: The rapid advancement of diffusion models has enabled thegeneration of high-fidelity images from textual prompts, yet achievingprecise, disentangled control over specific … |
Yusuf Dalva; Hidir Yesiltepe; Pinar Yanardag; |
| 213 | Walk Through Paintings : Ego-centric World Models from Internet Priors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To evaluate phys-ical correctness independent of appearance, we introduce the StructuralConsistency Score (SCS), which measures whether stable scene elementsevolve consistently with the provided actions. |
Anurag Bagchi; Zhipeng Bao; Homanga Bharadhwaj; Yu-Xiong Wang; Pavel Tokmakov; Martial Hebert; |
| 214 | Learning from Adversity: Semantic-Aware Mask Refinement Through Adversarial Perturbation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Phoenix, a novel framework that lever-ages adversarial learning to generate semantically meaningful noise pat-terns and contrastive learning to model refinement relationships. |
Beom Young Kim; Sung Ju Hwang; |
| 215 | Partial Skeleton Visibility for Action Recognition: A Constrained Field-of-View Approach Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Resourcesrelated to the constrained FoV setting used in this work are available at:https://github.com/yaa1haa1/PartialVisGraph. |
Yingjie Dai; Tianyang Xu; Yanglin Deng; Xiao-Jun Wu; Josef Kittler; |
| 216 | UniDrive-WM: Unified Understanding, Planning and Generation World Model For Autonomous Driving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose UniDrive-WM, a unified VLM-based world model that jointly performs driving-scene understanding, trajectory planning, and trajectory-conditioned fu-ture image generation within a single architecture. |
Zhexiao Xiong; Xin Ye; Burhaneddin Yaman; Sheng Cheng; Yiren Lu; Jingru Luo; Nathan Jacobs; Liu Ren; |
| 217 | TIR-Agent: Training An Explorative and Efficient Agent for Image Restoration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose TIR-Agent, atrainable image restoration agent that performs a direct tool-calling pol-icy through a two-stage training pipeline of supervised fine-tuning (SFT)followed by reinforcement learning (RL). |
Yisheng Zhang; Guoli Jia; Haote Hu; Shanxu Zhao; Kaikai Zhao; Long Sun; Xinwei Long; Kai Tian; Che Jiang; Zhaoxiang Liu; Kai Wang; Shiguo Lian; Kaiyan Zhang; Bowen Zhou; |
| 218 | Towards Geometry-Grounded Dense Semantic Matching with VGGT Priors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, directly transferring faces twochallenges: task heterogeneity, where VGGT originally matches cross-view images of the same instance, not cross-instance variations in shapeor appearance; and scarce dense matching annotations for task adapta-tion. To address these challenges, we propose an approach that (i) retainsVGGT’s intrinsic capabilities by reusing early feature stages, fine-tuninglater ones, and adding a semantic head for bidirectional correspondences;and (ii) adapts VGGT for semantic matching under data scarcity throughcycle-consistent training strategy, synthetic data augmentation, and pro-gressive training recipe. |
Songlin Yang; Tianyi Wei; Yushi Lan; Zeqi Xiao; Anyi Rao; Xingang Pan; |
| 219 | VGGT-World: Transforming VGGT Into An Autoregressive Geometry World Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present VGGT-World, a geometry world model that side-stepsvideo generation entirely and instead forecasts the temporal evolution offrozen geometry-foundation-model (GFM) features. |
Xiangyu Sun; Shijie Wang; Fengyi Zhang; Lin Liu; Caiyan Jia; Ziying Song; Zi Helen Huang; Yadan Luo; |
| 220 | From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce VWG-Bench (Video World Generalist Benchmark), a comprehensive benchmark spanning 9 reasoning dimensions and 38 fine-grained tasks. |
Meng Luo; Yicheng Liu; Jiahao Wang; Yuanxing Zhang; Xin Tao; Pengfei Wan; Kun Gai; Hao Fei; |
| 221 | Fine-Grained Text-to-Video Retrieval for Camera-Trap Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce Prompting-MammAlps, the first camera-trap TVR benchmark, and propose a fine-grained and interpretable TVR method. |
Valentin Gabeff; Baptiste Maquignaz; Jiaxian Shan; Sepideh Mamooler; Gencer Sumbul; Blair Costelloe; Devis TUIA; Alexander Mathis; |
| 222 | Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: It builds ona discrete diffusion framework that unifies vision and language represen-tations by modeling their shared probabilistic distribution. To empowerthe diffusion process to support unified and versatile medical generation,we employ a multimodal large language model (MLLM) as the diffusionbackbone, leveraging its rich prior knowledge and cross-modal reason-ing abilities. |
Jiawei Mao; Yuhan Wang; Lifeng Chen; Can Zhao; Yucheng Tang; Dong Yang; Liangqiong Qu; Daguang Xu; Yuyin Zhou; |
| 223 | Synthetic Visual Genome 2: Extracting Large-scale Spatio-Temporal Scene Graphs from Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Synthetic Visual Genome 2 (SVG2), alarge-scale panoptic video scene graph dataset. |
Ziqi Gao; Jieyu Zhang; Wisdom Ikezogwo; Jae Sung Park; Tario You; Daniel Ogbu; Chenhao Zheng; Weikai Huang; Yinuo Yang; Winson Han; Quan Kong; Rajat Saini; Ranjay Krishna; |
| 224 | AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, the performance of consistency-distilled models often degrades as more sampling steps are allocatedat test time, limiting their effectiveness for any-step video diffusion.We argue that this limitation arises because consistency distillation re-places the original probability-flow ODE trajectory with a consistency-sampling trajectory, weakening the desirable test-time scaling behaviorof ODE sampling. To address this limitation, we introduce AnyFlow,the first any-step video diffusion distillation framework based on flowmaps. |
Yuchao Gu; Guian Fang; Yuxin Jiang; Weijia Mao; Song Han; Han Cai; Mike Zheng Shou; |
| 225 | OmniPoint: Universal Monocular Metric Pointcloud from Any Camera Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present OmniPoint, a unified framework designed to generalize metric reconstruction across diverse imaging sensors, including pinhole, fisheye, and equirectangular projections, while accommodating varying geometric priors. |
Botao Ye; Marc Pollefeys; Ming-Hsuan Yang; Abhijit Kundu; |
| 226 | HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human–Scene Interactions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present HSImul3R1 , a unified framework for simulation-ready 3D reconstruction of human-scene interactions (HSI) from casualcaptures, including sparse-view images and monocular videos. |
Yukang Cao; Haozhe Xie; Fangzhou Hong; Long Zhuo; Zhaoxi Chen; Liang Pan; Ziwei Liu; |
| 227 | OmniMapBench: Benchmarking Visual-Centric Reasoning on Diverse Map Documents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Recent advancements in LVLMs necessitate robust bench-marks for complex, visually grounded reasoning. A critical limitation isidentified in many document understanding benchmarks: … |
Yang Chen; Yufan Shen; Yunwen Li; Minghao Liu; Tuney Tianyu; Bin Fu; Qunshu Lin; Zhi Yu; Botian Shi; |
| 228 | Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a paradigm shift by leveragingthe implicit spatial prior within large-scale video generation models. |
Xianjin Wu; Dingkang Liang; Tianrui Feng; Kui Xia; Yumeng Zhang; Xiaofan Li; Xiao Tan; Xiang Bai; |
| 229 | Semantic-Aware, Physics-Informed, Geometry-Grounded Weather Synthesis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address these, we propose a Semantic-Aware, Physics-Informed, and Geometry-Grounded framework that steers an off-the-shelf video editor to synthesize diverse global appearances and detailed particle dynamics. |
Chenghao Qian; Nedko Savov; Lingdong Kong; Yeying Jin; Rui Song; Wenjing Li; Zhun Zhong; Jiaqi Ma; Gustav Markkula; Luc Van Gool; |
| 230 | V-REX: Benchmarking Exploratory Visual Reasoning Via Chain-of-Questions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To bridge the gap, we develop an evaluation suite, “Visual Reasoning with multi-step EXploration (V-REX)”, which is composed of a benchmark of challenging visual reasoning tasks requiring native multi-step exploration and an evaluation protocol. |
Chenrui Fan; Yijun Liang; Shweta Bhardwaj; Kwesi Cobbina; Ming Li; Tianyi Zhou; |
| 231 | MotionAtlas: A High-Quality Dataset and Benchmark for Dense Motion Captioning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose MotionAtlas, a system for detailed captioningof motion-centric videos, comprising (1) a dedicated human-annotatedbenchmark, (2) a scalable, high-quality pipeline to construct trainingsamples, and (3) a family of powerful Video-MLLMs. |
Weisong Liu; Haochen Wang; Gaokuan Gaokuan; Yuhao Wang; Yikang Zhou; Zhongwei Ren; Guangcan Mai; Anran Wang; Yanwei Li; Xiangtai Li; Zhaoxiang Zhang; |
| 232 | Rotate Your Character: Revisiting Video Diffusion Models for High-Quality 3D Character Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present RCM(Rotate your Character Model ), an advanced image-to-video diffusionframework tailored for high-quality novel view synthesis (NVS). |
Jin Wang; Jianxiang Lu; Comi Chen; Guangzheng Xu; Haoyu Yang; Peng Chen; Na Zhang; Yifan Xu; Longhuang Wu; Shuai Shao; Qinglin Lu; Ping Luo; |
| 233 | EgoCogNav: Cognition-aware Human Egocentric Navigation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To facilitate research in thefield, we introduce the Cognition-aware Egocentric Navigation (CEN)dataset consisting of 6 hours real-world egocentric recordings capturingdiverse navigation behaviors in real-world scenarios. |
Zhiwen Qiu; Ziang Liu; Wenqian Niu; Tapomayukh Bhattacharjee; Saleh Kalantari; |
| 234 | Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a dataand computeefficient zero-shot SDV-Gen framework that avoids test-time per-subject tuning and the use of large-scale subject-video pairs. |
Daneul Kim; Jingxu Zhang; Wonjoon Jin; Sunghyun Cho; Qi Dai; Jaesik Park; Chong Luo; |
| 235 | Iterative Refinement of Semantic and Spatial Representations for Open-Vocabulary Camouflaged Object Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing two-stage methods typically treat category semantics as a one-time priorfor segmentation, thereby lacking deep and iterative interaction betweentextual semantics and visual spatial structures. To address this limita-tion, we propose an iterative two-stage refinement framework based onsemantic context and spatial structure. |
Fangyan Wang; Ge Jiao; Guowen Yue; |
| 236 | AMCoNav: Asynchronous Multi-module Collaborative Framework for Embodied Visual Navigation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing collaborative modular frameworks struggle to balance accurate decision-making and efficient online execution, leading to error propagation and inference blocking. To tackle this limitation, we propose AMCoNav, an Asynchronous Multi-module Collaborative Framework that unites real-time lightweight backbone networks with on-demand zeroshot large model modules (ZLMM) for robust and real-time embodied navigation. |
Jiaquan Yan; Fang Zhao; Yushi Chen; Long wang; Haiyong Luo; Dan Luo; |
| 237 | SIMON: SImultaneous Multi-Object Navigation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose threekey modules: a content-aware attention mechanism for region-adaptivefocus, a multi-object energy guidance strategy for subtask-specific con-sistency, and a latent initialization technique for artifact suppression.Together, these components enhance visual coherence across inpainting,object relocation, and background preservation subtasks. |
Yifeng Zhu; Siyuan Huang; Jun Bao; Jun Yu; Buyu Liu; |
| 238 | MedSynapse-V: Bridging Visual Perception and Clinical Intuition Via Latent Memory Evolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Toensure clinical fidelity, we introduce Causal Counterfactual Refine-ment (CCR) which leverages reinforcement learning and counterfactualrewards derived from region-level feature masking to quantify the causalcontribution of each memory, thereby pruning redundancies and aligninglatent representations with diagnostic logic. |
Chunzheng Zhu; Jiaqi Zeng; Junyu Jiang; Jianxin Lin; Yijun Wang; |
| 239 | MobileManiBench: Simplifying Model Verification for Mobile Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose asimulation-first framework to verify VLA architectures before real-worlddeployment and introduce MobileManiBench, a large-scale benchmarkfor mobile-based robotic manipulation. |
Wenbo Wang; Fangyun Wei; Qixiu Li; Xi Chen; Yaobo Liang; Chang Xu; Jiaolong Yang; Baining Guo; |
| 240 | Neutralizing Token Aggregation Via Information Augmentation for Efficient Test-Time Adaptation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we start by providingan analysis showing that token aggregation inherently leads to informa-tion loss, which cannot be fully mitigated by conventional norm-tuning-based TTA methods. |
Yizhe Xiong; Zihan Zhou; Yiwen Liang; Hui Chen; Zijia Lin; Xinhao Xu; Tianxiang Hao; Fan Zhang; Jungong Han; Guiguang Ding; |
| 241 | Zoom-IQA: Image Quality Assessment with Reliable Region-Aware Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce Zoom-IQA, a VLM-based IQA model to explicitly emulate key cognitive behaviors: uncertainty awareness, region reasoning, and iterative refinement. |
Guoqiang Liang; Jianyi Wang; Zhonghua Wu; Shangchen Zhou; Chen Change Loy; |
| 242 | NaVLM-PVC: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To investigate thistrend, we systematically compare their behavior on vision-language un-derstanding and attention patterns, revealing that global encoding en-hances overall capability but at the expense of greater computationaloverhead. To address this issue, we present NaVLM-PVC, an MLLM cen-tered upon our proposed Progressive Visual Compression (PVC) method,which can be seamlessly integrated into standard Vision Transformer(ViT) to enable efficient native-resolution encoding. |
Shichu Sun; Yichen Zhang; Haolin Song; Zonghao Guo; Chi Chen; Yidan Zhang; Yuan Yao; Zhiyuan Liu; Maosong Sun; |
| 243 | RCEdit-500K: Reference Completion for Image-Conditioned Image Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We reformulate ICIE data construction as a reference-completion problem: high-quality text-conditioned image editing (TCIE) datasets already supply the input image, instruction, and edited target, and can be augmented with aligned reference images through type-specific synthesis and lightweight instruction adaptation. Building on this insight, we propose a scalable pipeline equipped with weak-instruction augmentation and five-dimensional VLM-based post-filtering, and use it to construct RCEdit-500K—the first large-scale unified open ICIE dataset comprising 477K quadruplets across six edit categories (add, remove, replace, background, style, alter) with both concrete and abstract reference types. |
Jingxu Zhang; Daneul Kim; Yueming Pan; DONG CHEN; Kai Qiu; Yang Liu; Yifan Yang; Qi Dai; Xiaoyan Sun; Chong Luo; |
| 244 | Delayed Bidirectional Alignment Via Disentangled Audio Semantics for Audio-Visual Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing methods often struggle with multi-source en-tanglement and audio–visual misalignment, leading to a dominance biastoward acoustically or visually salient objects (i.e., louder or larger ones)at the expense of subtler or co-occurring sources. To address these chal-lenges, we propose DDAVS: Delayed Bidirectional Alignment via Dis-entangled Audio Semantics for Audio-Visual Segmentation. |
Jingqi Tian; Yiheng Du; Haoji Zhang; Yuji Wang; Isaac Ning Lee; Xulong Bai; Tianrui Zhu; Jingxuan Niu; Yansong Tang; |
| 245 | Plug-and-Play Traffic Element Awareness for End-to-End Autonomous Driving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present the first systematic investigation of traffic element awareness for end-to-end autonomous driving. |
Zongzheng Zhang; Jijun Wang; Saining Zhang; Wang Shuo; Yiru Wang; Hai Yang; Yang Chen; Yuwen Heng; HAO SUN; Jiang anqing; HAO ZHAO; |
| 246 | Rethink Backdoor Robustness in Vision Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: More-over, we propose a more robust attack strategy: by introducing slightperturbations to the trigger, existing attacks can be made significantlymore resistant to various defenses. |
Yichuan Mo; Dongxian Wu; Yifei Wang; Yisen Wang; |
| 247 | ParaFlow: Parallel Sampling for Flow Matching Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce ParaFlow, a training-free framework that recasts sampling as a system of Triangular Nonlinear Equations (TNEs) to enable step-level parallelism. |
Jianrong Lu; Bangwei Li; Haomin Zhang; Zhuoya Gu; Yongqing Lu; Jianhai Chen; Qinming He; |
| 248 | SAM-MT: Real-Time Interactive Multi-Target Video Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While recent approaches haveachieved remarkable performance in single-target scenarios, extendingthem to multi-target settings typically involves replicating the single-target processing for each individual object, resulting in reduced framerates (FPS) with unbounded latency as target density scales. Built uponSegment Anything 2 (SAM2), we propose SAM-MT, which addressesthis by transforming the model into an interactive framework for real-time Multi-Target video segmentation. |
Ruiqi Shen; Chang Liu; Henghui Ding; |
| 249 | SAM2Matting: Generalized Image and Video Matting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Despite impressive advances in image matting, video mattingremains challenging due to the inherent gap between high-level tracking,which requires frame-wise understanding, and … |
Ruiqi Shen; Guangquan Jie; Chang Liu; Henghui Ding; |
| 250 | Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we observe that, while generating multi-view consistent 3D content is highly challenging, verifying 3D consistency is tractable, naturally positioning reinforcement learning (RL) as a feasible solution. |
Jiyuan Wang; Chunyu Lin; Lei Sun; Zhi Cao; Yuyang Yin; Lang Nie; Zhenlong Yuan; Xiangxiang Chu; Yunchao Wei; Kang Liao; Guosheng Lin; |
| 251 | SparkVSR: Interactive Video Super-Resolution Via Sparse Keyframe Propagation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a novel interactive VSR framework dubbed SparkVSR that makes sparse keyframes a simple and expressive control signal. |
Jiongze Yu; Xiangbo Gao; Pooja Verlani; Akshay Gadde; Yilin Wang; Balu Adsumilli; Zhengzhong Tu; |
| 252 | SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To bridge this gap, we introduce SVI-Bench, a large-scale bench-mark that leverages team sports as a dynamic microworld, a domain thatuniquely combines the complexity of real-world multi-agent interactionwith the verifiability of explicit rules and definitive outcomes.We release the full benchmark to catalyzeprogress toward AI systems capable of strategic intelligence in complex,dynamic multi-agent environments. |
Yulu Pan; Han Yi; Seongsu Ha; Mohaiminul Islam; Benjamin Zhang; Lorenzo Torresani; Gedas Bertasius; |
| 253 | MoAKE: Toward Unified All-in-One Action Quality Assessment Via Mixture of Action Knowledge Experts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a novel Mixture of Action Knowledge Experts (MoAKE) framework, designed to mitigate negative knowledge transfer caused by large semantic discrepancies among actions. |
Huangbiao Xu; Huanqi Wu; Xiao Ke; Jiaxin Cai; Junyi Wu; Jinglin Xu; |
| 254 | PPTArena: A Benchmark for PowerPoint Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce PPTArena, a benchmark for PowerPoint edit-ing that evaluates how agents modify real slides from natural-languageinstructions. |
Michael Ofengenden; Yunze Man; Ziqi Pang; Liang-Yan Gui; Yu-Xiong Wang; |
| 255 | FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Large Vision-Language Models (VLMs) typically process hun-dreds or even thousands of visual tokens per image or video frame, in-curring quadratic attention costs and significant … |
Kaitong Cai; jusheng zhang; Sizhuo Ma; Bingqian Lu; Jian Wang; Keze Wang; |
| 256 | Taming LLMs for Codematic Indoor Scene Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Second, even with abundant data, LLMs exhibit poor instruction following due to an information imbalance, where the tokenheavy layout history overwhelms the concise user prompt. To resolve this, we propose SceneSpinner, a framework that introduces a language-based planning stage to provide high-level reasoning and a novel Conditional Mutual Information (CMI) regularization to explicitly force the model to focus on user instructions. |
yixun liang; Qianyi Wu; Chuan Fang; Rui Chen; Jiahang Liu; Jianfeng Zhang; Ping Tan; |
| 257 | CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Satellite-street scene pairs, with their complex contexts and extreme viewpoint variations, provide an ideal testbed. Motivated by this, we introduce CVSBench, a large-scale benchmark for evaluating cross-view spatial reasoning through satellite-street pairs. |
ruixun liu; Lingyu Zhang; Lanxuan Xue; Kaiyu Li; Bowen Fu; Xiangyong Cao; |
| 258 | MemoBench: Benchmarking World Modeling in Dynamically Changing Environments Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Video generation models aspire to simulate dynamic environ-ments, and several benchmarks now evaluate memory consistency acrossframes. However, most assess consistency only while … |
Haoyu Chen; Kaichen Zhou; Hang Hua; Kaile Zhang; Jingwen Qian; Wufei Ma; Haonan Chen; Chunjiang Liu; Yizhou Zhao; Xiaoyuan Wang; Weiyue Li; Alan Yuille; Paul Pu Liang; Yilun Du; |
| 259 | ADAPT: Attention Dynamics Alignment with Preference Tuning for Faithful MLLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we identify an internal signature of hallucination: progressive degradation of text-to-image cross-attention during generation, leading to specific failure patterns like unfocused or biased attention. |
Zhiyuan Yao; Zheren Fu; Zhixiao Zheng; Jiajun Li; Yi Tu; Zhendong Mao; |
| 260 | UniFusion: Sparse-View 4D Reconstruction Via Unified Spatio-temporal Depth Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we address the challenging problem of 4D re-construction from sparse-view videos. |
Yongzhe Lyu; Shaofei Wang; Yixin Chen; Siyuan Huang; |
| 261 | HippoCamp: Benchmarking Contextual Agents on Personal Computers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present HippoCamp, a new benchmark designed to eval-uate agents’ capabilities on multimodal file management. |
Zhe YANG; Shulin Tian; Kairui Hu; Shuai Liu; Hoang-Nhat Nguyen; Yichi Zhang; Zujin Guo; Mengying Yu; Zinan Zhang; Jingkang Yang; Chen Change Loy; Ziwei Liu; |
| 262 | Towards Unsupervised Multi-modal Semantic Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we make thefirst attempt to address the novel problem of Unsupervised Multi-modal Semantic Segmentation (UMSS), aiming to effectively exploitcomplementary sensor information in a fully label-free setting. |
Haitian Zhang; Thai Nguyen; Xiangyuan Wang; Mohan Liu; Addison Wang; |
| 263 | Benchmarking Scientific Understanding and Reasoning for Video Generation Using VideoScience-Bench Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce VideoScience-Bench1, a benchmark for evaluating undergraduate-level scientific understanding in video models. |
Lanxiang Hu; Abhilash Shankarampeta; Yixin Huang; Zilin Dai; Haoyang Yu; Yujie Zhao; Haoqiang Kang; Daniel Zhao; Tajana Rosing; Hao Zhang; |
| 264 | VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce VisReason, alarge-scale dataset designed to advance visual Chain-of-Thought rea-soning. |
Lingxiao Li; Yifan Wang; Xinyan Gao; Chen Tang; Xiangyu Yue; Chenyu You; |
| 265 | Understanding The Impact of Geometric Foundation Models on Vision-Language-Action Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While the resultinggeometric VLAs often show improved performance, it remains unclear(i) if modern VLAs already have sufficient geometric understanding tostart with, (ii) what is the best architecture to inject geometric under-standing into a VLA, and (iii) what is the effect of other design choicesthat affect geometric VLAs. In this paper we provide a rigorous exper-imental analysis to shed light on these questions, for a specific choiceof VLA (GR00T-N1.5) and GFM (VGGT). |
Yurou Yang; Muyuan Lin; Roberto Martín-Martín; Labrie Martin; Shreekant Gayaka; Cheng-Hao Kuo; Luca Carlone; |
| 266 | Hypothesis Graph Refinement: Hypothesis-Driven Exploration with Cascade Error Correction for Embodied Navigation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Hypothesis Graph Refinement(HGR), a framework that represents frontier predictions as revisable hy-pothesis nodes in a dependency-aware graph memory. |
Peixin Chen; Guoxi Zhang; Jianwei Ma; Qing Li; |
| 267 | General Incomplete Multimodal Learning Via Dynamic Quality Perception Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Toachieve reliable quality perception, we introduce a Noise-aware QualityEstimator that learns the mapping from corrupted features to noise in-tensity through controlled noise injection. |
Xiangyu Meng; shicai wei; |
| 268 | Stay Unique, Stay Efficient: Preserving Model Personality in Multi-Task Merging Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Decomposition, Thresholding, and Scaling (DTS), an approximation-based personalized merging framework that pushes task-specific storage efficiency. |
Kuangpu Guo; Aijing Yu; Jian Liang; Yuhe Ding; Zilei Wang; Ran He; Tieniu Tan; |
| 269 | Tactile Modality Fusion for Vision-Language-Action Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose TacFiLM, a lightweight modality-fusion ap-proach that integrates visual-tactile signals into vision-language-action(VLA) models. |
Charlotte Morissette; Amin Abyaneh; Wei-Di Chang; Anas Houssaini; David Meger; Hsiu-Chin Lin; Jonathan Tremblay; Gregory Dudek; |
| 270 | To Adapt or Not to Adapt? Selective Adaptation for Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce a newproblem of selective adaptation, which aims to determine whether agiven test sample should undergo adaptation or be skipped. |
Siru Jiang; Yuwei Liang; Jian Liang; Ran He; Tieniu Tan; |
| 271 | On The Vulnerability of Parameter-Level Defenses to Model Merging Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent works pro-pose parameter-level defenses that employ linear parameter transforma-tions to neutralize this threat. In this paper, we systematically analyzesuch defenses and reveal that their protected task vectors are inherentlysmall in magnitude. |
Kuangpu Guo; Jian Liang; Qingyan Zheng; Yu Yongcan; Zilei Wang; Ran He; Tieniu Tan; |
| 272 | RoMa V2: Harder Better Faster Denser Feature Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However,existing dense matchers still fail or perform poorly for many hard real-world scenarios, and high-precision models are often slow, limiting theirapplicability. In this paper, we attack these weaknesses on a wide frontthrough a series of systematic improvements that together yield a sig-nificantly better model. |
Johan Edstedt; David Nordström; Yushan Zhang; Georg Bökman; Jonathan Astermark; Viktor Larsson; Anders Heyden; Fredrik Kahl; Mårten Wadenbäck; Michael Felsberg; |
| 273 | InfiniteDance: Scalable 3D Dance Generation Towards In-the-wild Generalization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work aims to pushthe frontier of generalizable 3D dance generation by scaling up bothdata and model design. |
Ronghui Li; zhongyuan hu; Li Siyao; youliang zhang; Haozhe Xie; Mingyuan Zhang; Jie Guo; Xiu Li; Ziwei Liu; |
| 274 | AffoGato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Affogato, a framework foropen-vocabulary affordance grounding centered on Affogato-750K, alarge-scale dataset of 750K 3D affordance heatmaps paired with naturallanguage queries. |
Junha Lee; Eunha Park; Chunghyun Park; Dahyun Kang; Minsu Cho; |
| 275 | EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose Endogenous Chain-of-Thought (EndoCoT), a novel framework that facilitates structured visual reasoning in MLLMs by iteratively refining latent thought states through an iterative thought guidance module, and then bridges these states to the DiT’s denoising process. |
Xuanlang Dai; Yujie Zhou; Long Xing; Jiazi Bu; Xilin Wei; Yuhong Liu; Beichen Zhang; Kai Chen; Yuhang Zang; |
| 276 | BitRIC: Efficient Neural Compression of LiDAR Range Images Via Hierarchical Bitplanes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: WhileRange Image Compression (RIC) provides a more structured and com-putationally efficient alternative to conventional Point Cloud Compres-sion (PCC), existing RIC frameworks frequently fail to achieve an op-timal trade-off between coding efficiency and real-time performance. Tobridge this gap, we propose BitRIC, a novel learning-based frameworktailored for high-performance LiDAR range image compression. |
Kang You; Tong Chen; Dandan Ding; M. Salman Asif; Zhan Ma; |
| 277 | GeoWorld: Providing Full-frame Geometry Features to Facilitate 3D Scene Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper,we present a two-stage method, named GeoWorld, that renovates theimage-to-3D scene generation pipeline by providing full-frame geometryfeatures. |
Yuhao Wan; Lijuan Liu; Jingzhi Zhou; Zihan Zhou; Xuying Zhang; dongbo zhang; Shaohui Jiao; Qibin Hou; Ming-Ming Cheng; |
| 278 | Gaussian Belief Propagation Network for Depth Completion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although deep learningmethods have achieved state-of-the-art (SOTA), effectively handling thesparse and irregular nature of input depth data in deep networks remainsa significant challenge, often limiting performance, especially under highsparsity. To overcome this limitation, we introduce the Gaussian BeliefPropagation Network (GBPN), a novel hybrid framework synergisticallyintegrating deep learning with probabilistic graphical models for end-to-end depth completion. |
Jie Tang; Pingping Xie; Jian Li; Ping Tan; |
| 279 | Why Can Accurate Models Be Learned from Inaccurate Annotations? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This intriguing phenomenonraises a fundamental yet largely unexplored question: why models canstill extract correct label information from inaccurate annotations re-mains unexplored. In this paper, we conduct a comprehensive investiga-tion into this issue. |
Chongjie Si; Yidan Cui; Fuchao Yang; Wei Shen; |
| 280 | Multi-Hypothesis Test-Time Adaptation to Mitigate Underspecification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we reinterpret TTA through a posterior-inspired lensinduced by entropy minimization, where low-entropy solutions define apseudo-likelihood over parameters. |
Afshar Shamsi; Xiao-Yu Guo; Hamid Alinejad-Rokny; Arash Mohammadi; Damien Teney; Ehsan Abbasnejad; |
| 281 | TAIHRI: Task-Aware 3D Human Keypoints Localization for Close-Range Human-Robot Interaction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose TAIHRI, the first Vision-Language Model (VLM) tailored for close-range HRI perception, capable of understanding users’ motion commands and directing the robot’s attention to the most taskrelevant keypoints. |
Ao Li; Yonggen Ling; Yiyang Lin; Yuji Wang; Yong Deng; Yansong Tang; |
| 282 | Region-Aware Test-Time Scaling for Compositional Image Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Region-Aware Scaling (RAS), a framework that bridges region-aware generation andtest-time scaling. |
Mingzhu Shen; Peng Ye; Xinyin Ma; Gongfan Fang; Christos-Savvas Bouganis; Yiren Zhao; Xinchao Wang; |
| 283 | Towards Unified World Models for Visual Navigation Via Memory-Augmented Planning and Foresight Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Enabling embodied agents to imagine future states is essen-tial for robust and generalizable visual navigation. Yet, state-of-the-artsystems typically rely on modular designs that … |
Yifei Dong; Fengyi Wu; Guangyu Chen; Lingdong Kong; Xu Zhu; Qiyu Hu; Yuxuan Zhou; Jingdong Sun; Jun-Yan He; Qi Dai; Alexander Hauptmann; Zhi-Qi Cheng; |
| 284 | Bridging VideoQA and Video-Guided Agentic Tasks Via Generalized Keyframe Extraction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Furthermore, we observe that the performance of models on both VideoQA and video-guided agentic tasks critically depends on effective keyframe extraction. Based on this observation, we propose TASKER (Task-driven And Scene-aware Keyframe searchER), a keyframe extraction algorithm that jointly considers task relevance and scene dynamics to identify informative frames. |
Sunqi Fan; Qingle Liu; Runqi Yin; Meng-Hao Guo; Shuojin Yang; |
| 285 | EMOTE: Expressive Motion and Shape Disentanglement for Human Animation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: First, an Expressive Human Model (EHM) is introduced asthe core control representation. By explicitly disentangling shape andpose parameters, we fundamentally resolve the body shape leakage issue.Alongside this, a robust motion tracker is designed to accurately estimateEHM parameters from video. |
Dongbin Zhang; Hao Liu; Bingquan Dai; Kangjie Chen; Chuming Wang; Chen Li; Jing LYU; Haoqian Wang; |
| 286 | GenAgent: Scaling Text-to-Image Generation Via Agentic Multimodal Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce GenAgent, an agentic framework that unifiesvisual understanding and generation. |
Kaixun Jiang; Yuzheng Wang; Junjie Zhou; Pandeng Li; Zhihang Liu; Chen-Wei Xie; Zhaoyu Chen; Yun Zheng; Wenqiang Zhang; |
| 287 | SGQA: Semantic-Geometric Quality Alignment for Training-Free Few-Shot Instance Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Compositions of frozen foundation models offer a trainingfree route to instance segmentation, yet a clear gap remains between current systems and their empirical upper bounds. We introduce progressive oracle replacement, a diagnostic procedure that decomposes this gap into stage-level components, and find that the scoring stage accounts for most of it. |
Yuhao Qing; Liuyan Feng; Haoyuan Li; |
| 288 | IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose IQA-T1, atool-based visual evidence reasoning framework that augments MLLMreasoning with explicit perceptual observations. |
Jinjian Wu; Jiaqi Tang; Wei Wei; Yingying Yan; Jianmin Chen; Botong Geng; Lei Zhang; Qifeng Chen; |
| 289 | StructSplat: Generalizable 3D Gaussian Splatting from Uncalibrated Sparse Views Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present StructSplat, a feed-forward and generalizable3D Gaussian reconstruction framework that operates directly on uncal-ibrated images without requiring camera parameters. |
Jia-Chen Zhao; Beiqi Chen; Xinyang Chen; Guangcong Wang; Liqiang Nie; |
| 290 | Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, ex-isting proxy tasks primarily enhance visual and textual representationsto accommodate a predefined composition mechanism such as pseudo-word injection into a frozen text encoder or linear feature arithmetic.As a result, the composition function itself remains unlearned, limitingthe model’s ability to express diverse and fine-grained semantic mod-ifications. To address this, we propose FoCo, which models composi-tion as two coordinated stages: focusing on modification-relevant visualcontent, and then completing the target semantics. |
Jingjing Zhang; Lei Zhang; Zheren Fu; Zhendong Mao; |
| 291 | VLTR: Vision-Language Tool Reasoning for Instruction-Guided Image Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose VLTR (Vision-Language Tool Reasoning), a training-free framework that reformulatesediting as closed-loop tool reasoning over a directed acyclic graph (DAG)of atomic primitives. |
Yike Wang; Yitao Yu; Shaohua Sun; Ping Luo; |
| 292 | Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: ForenAgent adopts a two-stage training pipeline with Cold Start andReinforcement Fine-Tuning to progressively improve tool interaction andreasoning adaptability. |
Fanrui Zhang; Qiang Zhang; Sizhuo Zhou; Jianwen Sun; Chuanhao Li; Jiaxin Ai; Yukang Feng; Yujie Zhang; Wenjie Li; Zizhen Li; Yifan Chang; Jiawei Liu; Kaipeng Zhang; |
| 293 | RayRoPE: Projective Ray Positional Encoding for Multi-view Attention Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study positional encodings for multi-view transformersthat process tokens from a set of posed input images, and seek a mech-anism that encodes patches uniquely, allows SE(3)-invariant attentionwith multi-frequency similarity, and can adapt to the geometry of theunderlying 3D scene. |
Yu Wu; Minsik Jeon; Rick Chang; Oncel Tuzel; Shubham Tulsiani; |
| 294 | Actor As Its Own Critic: Unifying Region Understanding and Localization Via CycleGRPO Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introduces Actor as Its Own Critic, a unifiedreinforcement learning framework, Cycle Group Relative Policy Opti-mization (CycleGRPO), that jointly optimizes region understanding andlocalization for Multimodal Large Language Models (MLLMs). |
Xin Zhang; Haochen Wang; Yikang Zhou; Zhuochen Wang; Robby T. Tan; Xiangtai Li; |
| 295 | Embed-RL: Reinforcement Learning for Reasoning-Driven Multimodal Embeddings Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, the generated reasoning CoTs of existing generative embedding methods are limited to the textual analysis of queries and are irrelevant to the retrieval of the targets. To address these limitations, we propose a reasoning-driven UME framework that integrates Embedder-Guided Reinforcement Learning (EG-RL) to optimize the Reasoner to produce evidential Traceability CoT (T-CoT). |
Haonan Jiang; Yuji Wang; Yongjie Zhu; Xin Lu; Wenyu Qin; Meng Wang; Pengfei Wan; Yansong Tang; |
| 296 | Molecular Identifier Visual Prompting and Verifiable Reinforcement Learning for Chemical Reaction Diagram Parsing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our contributions advance theaccuracy and generalization of VLM-based reaction diagram parsing.Additionally,we release the ScannedRxn benchmark, comprising scanned historicalreaction diagrams with real-world artifacts, to rigorously assess model ro-bustness and out-of-distribution ability. |
Jiahe Song; Chuang Wang; Yinfan Wang; Hao Zheng; Bowen Jiang; Rui Nie; Xingjian Wei; Junyuan Gao; Yubin Wang; Bin Wang; Lijun Wu; Jiang Wu; Qian Yu; Conghui He; |
| 297 | ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose ShotStream, a novel causal multi-shot architecture that enables interactive storytelling and efficient on-the-fly frame generation. |
Yawen Luo; Xiaoyu Shi; Jun-hao Zhuang; Yutian Chen; Quande Liu; Xintao Wang; Pengfei Wan; Tianfan Xue; |
| 298 | A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: How-ever, current evaluation paradigms index heavily on visual fidelity and se-mantic alignment, leaving a critical blind spot: they cannot reliably quan-tify whether generated dynamics actually obey the fundamental physicallaws required for reliable simulation. Assessing this physical plausibilityis inherently difficult due to a lack of physical metrics and the challengeof extracting metric-scale kinematics from uncalibrated video rollouts.To bridge this gap, we introduce CrashTwin, a physics-grounded eval-uation framework designed to stress-test the physical trustworthiness ofworld models. |
Nuo Chen; Lulin Liu; Zihao Li; Ziyao Zeng; Zihao Zhu; Wenyan Cong; Junyuan Hong; Yunhao Yang; Zhengzhong Tu; Yan Wang; Boris Ivanovic; Marco Pavone; Zhangyang Wang; Yang Zhou; Zhiwen Fan; |
| 299 | One4D: Unified 4D Generation and Reconstruction Via Decoupled LoRA Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present One4D, a unified framework for 4D generationand reconstruction that produces dynamic 4D content as synchronizedRGB frames and pointmaps. |
Zhenxing Mi; YUXIN WANG; Dan Xu; |
| 300 | Inductive Visual Logic for Few-Shot Out-Of-Distribution Adaptation in VLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Generative vision-language models (VLMs) such as Qwen-VLand LLaVA achieve strong zero-shot performance on tasks overlappingwith their pretraining distribution, yet fail on specialized domains wherethe required discriminative features were never learned, a regime weterm distant out-of-distribution (OOD). |
Hung-Jen Chen; Yu-Heng Ho; Ting-Yao Huang; Po-Hsiang Hsu; LIYU CHEN; Chun-Yi Lee; Min Sun; |
| 301 | MonoSR: Open-Vocabulary Spatial Reasoning on Monocular Images Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing benchmarks either rely on multiview video sequences that expose explicit geometric cues, or are confined to indoor environments too small for model training. To close this gap, we introduce MonoSR, a large-scale dataset for open-world monocular spatial reasoning comprising over 1M QA pairs from 230K images spanning indoor, outdoor, and object-centric domains across 98 semantic categories. |
Qirui Wang; Jingyi He; Yining Pan; Si Yong Yeo; Xulei Yang; Shijie Li; |
| 302 | MirrorPPR: Exemplar-Based Portrait Photo Retouching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In contrast, structuralportrait retouching involves extremely delicate and localized modifica-tions, making accurate extraction and transfer of these edits challeng-ing. To tackle this, we propose MirrorPPR, a novel framework specifi-cally designed to capture and transfer subtle structural retouching oper-ations. |
Zhihong Liu; Zheng Li; Jiachun Jin; Siqi Kou; Yitao Jian; Fengpei Yu; Zhijie Deng; |
| 303 | Towards Practical Lossless Neural Compression for LiDAR Point Clouds Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: LiDAR point clouds are fundamental to various applications, yet the extreme sparsity of high-precision geometric details hinders efficient context modeling, thereby limiting the compression speed and performance of existing methods. To address this challenge, we propose a compact representation for efficient predictive lossless coding. |
pengpeng yu; Haoran Li; Runqing Jiang; Dingquan Li; Jing Wang; Liang Lin; Yulan Guo; |
| 304 | PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Plus,we propose a LoRA-Switch Reference (LoRA-SR) scheme that avoidsfull-model duplication as reference for efficient training. |
Yuanhao Cai; Kunpeng Li; Menglin Jia; Jialiang Wang; Junzhe Sun; Feng Liang; Weifeng Chen; Felix Juefei-Xu; Chu Wang; Ali Thabet; Xiaoliang Dai; Xuan JU; Alan Yuille; Ji Hou; |
| 305 | Toward Robust In-Context Segmentation Via Concept Guidance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In thiswork, we revisit ICS from the robustness perspective and introduce anovel paradigm, Concept-Guided In-Context Segmentation (CG-ICS),which performs segmentation by extracting high-level semantic conceptsfrom references rather than relying solely on low-level visual matching.Specifically, CG-ICS introduces a concept reasoning module that usesan MLLM to propose candidates and a SAM3-driven scoring functionwith tree-search refinement to select reliable textual concepts, togetherwith a parallel visual exemplar route that provides query-side spatialgrounding via a simple context construction. |
Zhigang Chen; Xiawu Zheng; Rongrong Ji; |
| 306 | Robust Self-Supervised Cross-Modal Super-Resolution Against Real-World Misaligned Observations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Previous methods either rely on simulated training dataor adopt suboptimal alignment strategies that overlook cross-modal de-pendencies, limiting their practical performance. To address these is-sues, we propose RobSelf, a self-supervised model that jointly optimizesa misalignment-aware feature translator and a content-aware referencefilter online. |
Xiaoyu Dong; Jiahuan Li; Ziteng Cui; Naoto Yokoya; |
| 307 | OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Previous discrete video VAEs suffer from un-stable training, long training time, and degraded reconstruction quality.We revisit the relationship between continuous and discrete VAEs andfind that bridging discrete and continuous representations improves dis-crete token learning. Based on this insight, we propose a unified progres-sive training framework that (i) jointly optimizes continuous and discretereconstructions within a single network, and (ii) progressively derives afamily of VAEs at different compression ratios, leading to faster conver-gence and better final performance. |
Yupeng Zhou; Zhen Li; Yuming Chen; Ziheng Ouyang; Ruoyi Du; Daquan Zhou; Bin Fu; Yihao Liu; Peng Gao; Ming-Ming Cheng; Qibin Hou; |
| 308 | InnoText: A Unified Model for Visual Text Generation and Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce aFont Size-Aware Modulation (FSAM) module to enhance representationsacross font scales, a Small-Character Aware Augmentation strategy toimprove fine-grained fidelity, and a Task-Specific Region Weighted Lossfor adaptive optimization. |
Haowei Liu; Runze He; Jian Lu; Ao Ma; Run Ling; Ke Cao; Jiasong Feng; Wei Feng; Shuo Lu; Yexing Xu; WANG Yun; Jing Wang; Zhanjie Zhang; |
| 309 | Hyperbolic Hierarchical Clustering for Visual Representation Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, a significant draw-back of these methods is their black-box nature; their encoding processis opaque and lacks interpretability. Diverging from these opaque designs,we introduce ClusterMixer, a transparent token mixer that is grounded ina clustering paradigm and interpretable by design. |
Jianan Wei; Guikun Chen; Zhiyuan Weng; Chunchao Guo; Yujia Wang; Wenguan Wang; |
| 310 | Unified Multi-plane Autoregressive Diffusion for 3D Multi-Contrast MRI Synthesis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: During inference,we introduce plane-wise autoregressive synthesis with inter-plane priors.Slices are generated autoregressively in random order within one planeorientation to maintain intra-plane continuity, then propagated as con-ditioning priors to orthogonal plane orientations to enforce inter-planeconsistency. |
Yejee Shin; Geonhui Son; Jinglu Wang; Minwoo Jung; Yan Lu; Dosik Hwang; |
| 311 | Syn-GRPO: Self-Evolving Data Synthesis for MLLM Perception Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Somemethods attempt to mitigate this problem by imposing constraints onentropy, but none address it at its root. Therefore, to tackle this problem,this work proposes Syn-GRPO (Synthesis-GRPO), which employs anonline data generator to synthesize high-quality training data with di-verse responses in GRPO training. |
Qihan Huang; Haofei Zhang; Rong Wei; Yi Wang; Rui Tang; Mingli Song; Jie Song; |
| 312 | VIVAS: Vitalizing Visual Perception in VLM Pre-training Via Vision-language Unified Autoregressive Supervision Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We investigate that overcoming this bottle-neck requires two key elements: (1) a unified token space paradigm thatensures stable training dynamics, and (2) a modality-aligned dense vi-sual supervision signal enriched with both structural granularity andsemantic information to capture critical visual representations. Basedon these insights, we propose VIVAS, a framework built upon the uni-fied token space paradigm, which introduces a dense-structural–semanticvision tokenizer, which expands the textual vocabulary into a unified vi-sion–language vocabulary by incorporating a visual vocabulary. |
Zhehan Kan; Yubo Zhu; Xinghua Jiang; Zhixiang Wei; Shifeng Liu; Wei Tong; Sheng Zhong; Qingmin Liao; Wenming Yang; Xin Li; Yinsong Liu; Deqiang Jiang; Xing Sun; |
| 313 | Scaling Laws for Black-box Adversarial Attacks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We show that by resolving gradient conflict with advanced op-timizers, we overcome the quantitative limitations of idealized theoreticalbounds to empirically discover a robust log-linear scaling law, demon-strating that the Attack Success Rate scales linearly with the logarithmof the ensemble size T . |
Chuan Liu; Huanran Chen; Yichi Zhang; Jun Zhu; Yinpeng Dong; |
| 314 | Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This ap-proach underutilizes the unique potential of images to carry complex,contextual information. To address this gap, we propose a new image-centric attack method, Contextual Image Attack (CIA), which employsa multi-agent system to subtly embed harmful queries into seeminglybenign visual contexts using four distinct visualization strategies. |
Yuan Xiong; Miao Ziqi; Lijun Li; Chen Qian; Jie Li; Jing Shao; |
| 315 | Layer-Aware Video Composition Via Split-then-Merge Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Split-then-Merge (StM), a controllable genera-tive video composition framework that minimizes reliance on annotateddatasets and handcrafted rules.Finally, we release StM-50K, the _x001C_rst multi-layer video dataset, to facilitate future research ingenerative video composition. |
Ozgur Kara; Yujia Chen; Ming-Hsuan Yang; James Rehg; Wen-Sheng Chu; Du Tran; |
| 316 | LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present LoGeR (Long-context GeometricReconstruction), a novel architecture that scales dense 3D reconstruc-tion to extremely long sequences without post-optimization. |
Junyi Zhang; Charles Herrmann; Junhwa Hur; Chen Sun; Ming-Hsuan Yang; Forrester Cole; Trevor Darrell; Deqing Sun; |
| 317 | EffiDINO: Task-Specific Model Pruning Via Gram Anchoring Subspace Consistency Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing methods focus on rigid point-to-point tokenalignment on a single dataset for pruning, suffering from two limitations:i) robustness degradation, and ii) task-specificity deficiency. To addressthese limitations, we propose a task-specific pruning pipeline, namedCut-ViT. |
Jianjian Yin; Liulei Li; Tao Chen; Yi Chen; Yazhou Yao; Wenguan Wang; |
| 318 | OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: How-ever, existing panoramic benchmarks largely focus on simplistic queriesthat rely on local cues or single-/few-step reasoning, thereby ignoring thefundamental advantage of panoramas and failing to fully exploit theirpotential. To address this gap, we introduce OmniCoT , a panoramicspatial reasoning suite designed to enable MLLMs to use global evi-dence and perform multi-step inference across viewpoints. |
Haocong He; Chenfei Liao; Zichen Wen; Zihao Dongfang; Xu Zheng; Bin Ren; Chang Su; Zixin Zhang; Harold Haodong Chen; Hongfei Zhang; Weijia Li; Kailun Yang; Conghui He; Xuming Hu; Nicu Sebe; Linfeng Zhang; |
| 319 | Delving Into Latent Spectral Biasing of Video VAEs for Superior Diffusability Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a statisticalanalysis of video VAE latent spaces and identify two spectral propertiesessential for diffusion training: a channel-wise eigenspectrum dominatedby a few modes, and a spatio-temporal frequency spectrum biased towardlow frequencies. To induce these properties, we propose two lightweight,backbone-agnostic regularizers: Latent Masked Reconstruction and Lo-cal Correlation Regularization. |
Shizhan Liu; Xinran Deng; Zhuoyi Yang; Jiayan Teng; Xiaotao Gu; Jie Tang; |
| 320 | ZTRS: Zero-Human Demonstration End-to-end Autonomous Driving with Trajectory Scorer Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we proposeZTRS (Zero-human demonstration end-to-end autonomous driving withTRajectory Scorer) — a complete RL-based E2E planning paradigmtrained solely on real-world images and rule-based rewards, entirely with-out human demonstration. |
Zhenxin Li; Nadine Chang; Wenhao Yao; Xinglong Sun; Zi Wang; Maying Shen; Jingde Chen; Jingyu Song; Kailin Li; Zuxuan Wu; Shiyi Lan; Jose M Alvarez; |
| 321 | ToDRE: Effective Visual Token Pruning Via Token Diversity and Task Relevance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Instead ofpruning redundant tokens, we introduce a greedy max-sum diversificationalgorithm that selects and retains a subset of diverse and representativevisual tokens after the vision encoder. |
Duo Li; Zuhao Yang; Xiaoqin Zhang; Ling Shao; Shijian Lu; |
| 322 | LightSTAR: Efficient Visual Document Retrieval Via Lightweight Selection with Vision-Adaptive Refinement Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Mean-while, we observe that user queries are typically keyword-anchored, con-taining semantically rich words that are expected to appear directly inthe visible text of relevant pages, offering an efficient cue for quickly nar-rowing down candidate pages. Building on this insight, we propose Light-STAR, an efficient framework that decomposes visual document retrievalinto: 1) LLM-free Visual Selection, which utilizes content-grounded queryencoding to focus on informative words and employs LLM-free visual em-beddings to produce a high-recall candidate set; and 2) Vision-adaptiveSemantic Refinement, which further performs fine-grained semantic match-ing exclusively on these top candidates via adaptive region-wise featurefusion to effectively combine textual and layout cues, optimized through ahardness-aware contrastive objective. |
Tongkun Guan; Haocheng Wang; Wei Shen; Xiaokang Yang; |
| 323 | QualiTeacher: Quality-Conditioned Pseudo-Labeling for Real-World Image Restoration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, wepropose QualiTeacher, a novel framework that transforms pseudo-labelquality from a noisy liability into a conditional supervisory signal. |
Fengyang Xiao; Jingjia Feng; Peng Hu; Yuhan Chen; Dingming Zhang; Lei Xu; Guanyi Qin; Lu Li; Chunming He; Sina Farsiu; |
| 324 | Rethinking IRSTD: Single-Point Supervision Guided Encoder-only Framework Is Enough for Infrared Small Target Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we reformulate IRSTD as a centroid re-gression task and propose a novel Single-Point Supervision guided In-frared Probabilistic Response Encoding method (namely, SPIRE), whichis non-trivial because point-level supervision must produce detectionoutputs comparable to dense supervision. |
Rixiang Ni; Boyang Li; Chen Jun; Zhijie Chen; Feiyu Ren; Yuji Wang; Haoyang Yuan; Wujiao He; Wei An; |
| 325 | Ceptor: Vision-Language Model-Infused Diverse Guidance for Detecting Anything Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by the general process of hu-man object search, we designed Ceptor, a unified detector guided bydiverse prompts for open-set object detection. |
Jinyang Li; Bin-Bin Gao; Weifu Fu; Jingnan Luo; Hanqiu Deng; Yue Guo; Jun Liu; Yong Liu; Chengjie Wang; Wenbing Tao; |
| 326 | Beyond Isolated Scans: Cross-Phase Alignment of Structure and Topology for 3D Medical Pretraining Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: CAST employs a 3D CNN archi-tecture to explicitly align NCCT representations with CECT targets.Moving beyond conventional reconstruction, we introduce two feature-level constraints: (1) a Spectral Consistency module that utilizes3D wavelet decomposition to align frequency-aware boundaries whilesuppressing contrast-induced noise; (2) a Geometry-Aware Topo-logical Consistency module that preserves local relational graphsamong salient anatomical keypoints via dynamic top-hat sampling.Tosupport this, we construct a large-scale dataset comprising 13,850 pairedvolumetric CT scans. |
Wenzhuo xu; YANJIE ZHOU; Yujian Hu; Hongkun Zhang; Minfeng Xu; |
| 327 | DETRAM: End-to-end DEtection, Tracking and Recovery of HumAn Meshes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: In the task of human mesh recovery (HMR), multi-personscenes are particularly difficult to handle due to the many entities thatappear and occlusions between them over time. In … |
Chunggi Lee; Seonwook Park; Wanhua Li; Umar Iqbal; Hanspeter Pfister; |
| 328 | JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The intrinsic link between facial motion and speech is oftenoverlooked in generative modeling, where talking head synthesis andtext-to-speech (TTS) are typically addressed as separate tasks. |
Mingi Kwon; Joonghyuk Shin; Jaeseok Jeong; Jaesik Park; Youngjung Uh; |
| 329 | Video-HolmesV2: Can MLLMs Reason with Spatio-Temporal Audio-Visual Evidence in Long Videos? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Unlike previous works, it enforces an Evidence-Based Evaluation, requiring models to justify answers with precise spatio-temporal audio-visual evidence, thereby reducing confounding effects of guessing and hallucinated evidence. To support this, we introduce: (1) a MultiModel Cross-Verification pipeline to ensure task rigor; (2) a Spatiotemporal Evidence-Aware Metric for fine-grained calibration. |
Zhaoyang Wei; Zipeng Wang; Yushe Cao; Chenhui Qiang; Shuaibing Cheng; Xuesong Yang; Sen Nie; Bowen Jiang; Wenchao Ding; Yanchao Hao; Zheng Wei; Xuehui Yu; Zhenjun Han; |
| 330 | LSRM: High-Fidelity Object-Centric Reconstruction Via Scaled Context Windows Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce the Large Sparse Reconstruction Model tostudy how scaling transformer context windows affects feed-forward 3Dreconstruction. |
Zhengqin Li; Cheng Zhang; Jakob Engel; Dong Zhao; |
| 331 | Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a novel evaluation frameworkthat combines DermBench, a meticulously curated benchmark, withDermEval, a robust automatic evaluator, to enable clinically meaning-ful, reproducible, and scalable assessment. |
Yuhao Shen; Jiahe Qian; Zhangtianyi Chen; Juexiao Zhou; |
| 332 | History-Aware Transformation of ReID Features for Multiple Object Tracking Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: In this paper, we propose a history-aware feature transformation method that dynamically crafts a more discriminative subspace tailored to each video’s unique sample distribution. |
Ruopeng Gao; Yuyao Wang; Chunxu Liu; Limin Wang; |
| 333 | What Moves? Localized Motion Representations for Compositional Scene Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, localized embeddings are often computed from cropped images or obtained by masking features after encoding, discarding the context needed to interpret motion. To address this, we introduce a promptable localized motion representation that produces persistent embeddings for user-specified regions defined by spatial masks. |
Frank Fundel; Malek Ben Alaya; Thomas Ressler-Antal; Stefan Andreas Baumann; Bjorn Ommer; |
| 334 | Show Me Examples: Inferring Visual Concepts from Image Sets Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inparticular, current models fail to infer shared concepts from sets of exam-ple images and apply them to new inputs. We introduce Visual ConceptInference from Sets (VICIS), a task that evaluates this capability. |
Nick Stracke; Kolja Bauer; Stefan Andreas Baumann; Miguel Angel Bautista; Joshua Susskind; Bjorn Ommer; |
| 335 | Logit Refiner: Improving Visual Autoregressive Models Via Intra-Scale Dependency Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Visual Autoregressive Models (VAR) generate images throughnext-scale prediction, producing all tokens within each scale in paral-lel. We show that this parallel decoding constitutes a mean-field-styleapproximation that discards spatial dependencies among same-scale to-kens, causing locally incoherent samples regardless of backbone capacity– a limitation of the decoding rule. |
Meimingwei Li; Stefan Andreas Baumann; Felix Krause; Bjorn Ommer; |
| 336 | Schroedinger’s Cat: Probabilistic Representation and Prediction of Potential Scene Kinematics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Predicting how a scene may evolve from partial observationsrequires reasoning about multiple possible futures rather than committingto a single trajectory. Existing approaches … |
Timy Phan; Jannik Wiese; Bjorn Ommer; |
| 337 | RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce RayDer, a unified, feed-forward transformer that consolidates camera estimation, scene reconstruction, and rendering into a single backbone, turning self-supervised NVS into a well-posed single-model scaling problem. |
Ulrich Prestel; Stefan Andreas Baumann; Nick Stracke; Bjorn Ommer; |
| 338 | Boba: Batched Simulation for Physics-Based Gaussian Digital Twins Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our key idea is to separate thestatic twin template from the dynamic simulation state and co-design thephysics, deformation, and rendering/visualization pipelines for batchedexecution. |
Yihan Pang; Hanxiao Jiang; Sushant Kondguli; Sarita Adve; Shenlong Wang; |
| 339 | AutoPhyX: Automatic Text-Condition Physics Property Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The latter suffers from inherent ambiguity,e.g., the inability to distinguish the stiffness of rubber from its ap-pearance alone. To bridge this gap, we propose AutoPhyX, a text-conditioned framework for predicting spatially-varying physical proper-ties. |
Bei Huang; Yixin Chen; Hongbin Zha; Yuru Pei; Siyuan Huang; |
| 340 | Out of Sight, Out of Mind? Evaluating State Evolution in Video World Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Can these generated “worlds” evolveregardless of observation? To probe this question, we design a benchmarkto evaluate whether video world models can decouple state evolutionfrom observation. |
Ziqi Ma; Mengzhan Liufu; Georgia Gkioxari; |
| 341 | DINO-SLAM: DINO-Informed RGB-D SLAM for Neural Implicit and Explicit Representations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents DINO-SLAM, a DINO-informed designstrategy to enhance implicit (Neural Radiance Field – NeRF) and explicitrepresentations (Gaussian Splatting – GS) in SLAM systems through themore comprehensive semantic understanding enabled by DINO. |
ZIREN GONG; Xiaohan Li; Fabio Tosi; Youmin Zhang; Stefano Mattoccia; Jun Wu; Matteo Poggi; |
| 342 | MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents MAGiSt3R, a multi-agent 3D reconstruction framework performing reconstruction and camera tracking for monocular RGB videos at almost 10 FPS. |
ZIREN GONG; Xiaohan Li; Fabio Tosi; Ninghui Xu; Stefano Mattoccia; Jianfei Cai; Matteo Poggi; |
| 343 | Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As a result, the prevailing pattern of detecting in 2D and then predicting 3D attributes is often brittle, since modest range errors can dominate 3D localization, and the learned scale prior can fail when cameras, motion, or environments undergo domain shifts. To address this, we propose Map-Det3D, an online multi-view 3D object detection model that brings detection directly into a 3D space reconstructed from RGB. |
Yung-Hsu Yang; Luigi Piccinelli; Samuel Rota Bulò; Sunghwan Hong; Denys Rozumnyi; Johannes Schönberger; Zuria Bauer; Hermann Blum; Peter Kontschieder; Marc Pollefeys; |
| 344 | RiO-DETR: DETR for Real-time Oriented Object Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present RiO-DETR: DETR for Real-time OrientedObject Detection, the first real-time oriented detection transformer tothe best of our knowledge. |
Zhangchi Hu; Yifan Zhao; Yansong Peng; Wenzhang SUN; Xiangchen Yin; Jie Chen; Peixi Wu; Hebei Li; xinghao wang; Dongsheng Jiang; Xiaoyan Sun; |
| 345 | GLARE: Towards Generalizable Detection of Latent Diffusion Images with Global-Local Reconstruction Error Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our method outperforms state-of-the-artsupervised and training-free baselines significantly and shows strong ro-bustness against common post-processing operations. |
Jiangtao Yan; Jiazhen Ji; Zhongyu Zhang; Yuge Huang; Wenbin Wang; Shouhong Ding; |
| 346 | InverseCrafter: Efficient Video ReCapture As A Latent Domain Inverse Problem Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This dominant paradigm is computationally expensive and frequentlysu!ers from catastrophic forgetting of the model’s original generative pri-ors. To address this challenge, here we propose InverseCrafter, a VDMtraining-free framework that reformulates novel view video generationas an inpainting-based inverse problem in the latent space, eliminatingthe need for any annotated 4D training data. |
Yeobin Hong; Suhyeon Lee; Hyungjin Chung; Jong Chul Ye; |
| 347 | On-Policy Diffusion Reinforcement Learning Meets Off-Policy Quality Anchoring Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite offering a direct optimization signal, theapproach is fundamentally self-limiting and prone to reward overfittingand mode collapse due to its myopic guidance. In this paper, we pro-pose a new framework, namely DiffusionCompass, that breaks such lim-itation by strategically integrating off-policy guidance. |
Zunxu Liu; Zhaofan Qiu; Yazhen Xie; Yingwei Pan; Ting Yao; Tao Mei; |
| 348 | When The City Teaches The Car: Label-Free 3D Perception from Infrastructure Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Leveraging their fixed viewpoints and re-peated observations, RSUs learn local 3D detectors from unlabeled dataand broadcast predictions to passing vehicles, which are aggregated aspseudo-label supervision for training a standalone ego detector. The re-sulting model requires no infrastructure or communication at test time.We instantiate this idea as a fully label-free three-stage pipeline and con-duct a concept-and-feasibility study in a CARLA-based multi-agent en-vironment. |
ZHEN XU; Jinsu Yoo; Cristian Bautista; Zanming Huang; Tai-Yu Pan; Zhenzhen Liu; Katie Luo; Mark Campbell; Bharath Hariharan; Wei-Lun Chao; |
| 349 | MediRound: Multi-Round Entity-Level Reasoning Segmentation in Medical Images Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce Multi-Round Entity-Level Medical Reasoning Segmentation (MEMR-Seg), anew task that requires generating segmentation masks through multi-round queries with entity-level reasoning, helping learners progressivelydevelop their understanding of medical knowledge. |
Qinyue Tong; Ziqian Lu; Jun Liu; Rui Zuo; Zheming Lu; Yueming Jin; |
| 350 | Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present Cheers, a unified multimodalmodel that decouples patch-level details from semantic representations,thereby stabilizing semantics for multimodal understanding and improv-ing fidelity for image generation via gated detail residuals. |
Yichen Zhang; Da Peng; Zonghao Guo; zijian zhang; Xuesong Yang; Tong Sun; Shichu Sun; Yidan Zhang; Yanghao Li; Haiyan Zhao; Wang Xu; Qi Shi; Yangang Sun; Chi Chen; Shuo Wang; Yukun Yan; Xu Han; Qiang Ma; Wei Ke; Liang Wang; Zhiyuan Liu; Maosong Sun; |
| 351 | How Far Are Video Models from True Multimodal Reasoning? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing benchmarks fail toaddress this question rigorously, as they remain constrained by straight-forward task designs and fragmented evaluation metrics that neglectcomplex multimodal reasoning. To bridge this gap, we introduce CLVG-Bench, an evaluation framework designed to probe video models’ zero-shot reasoning capabilities via Context Learning in Video Generation.CLVG-Bench comprises more than 1,000 high-quality, manually anno-tated metadata across 6 categories and 47 subcategories, covering com-plex scenarios including physical simulation, logical reasoning, and inter-active contexts. |
Xiaotian Zhang; Jianhui Wei; Yuan Wang; Jie Tan; Yichen Li; Yan Zhang; Ziyi Chen; Daoan Zhang; DEZHI YU; Wei Xu; Songtao Jiang; Zuozhu Liu; |
| 352 | What Matters in RL-Based Methods for Object-Goal Navigation? An Empirical Study and A Unified Framework Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present a large-scale empirical study of modular RL-based ObjectNav systems. |
Hongze Wang; Boyang Sun; Jiaxu Xing; Fan Yang; Marco Hutter; Dhruv Shah; Davide Scaramuzza; Marc Pollefeys; |
| 353 | Unlocking Complex Image Editing Via Natively Interleaved Visual Textual CoT with Deep Confidence Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While Chain of Thought (CoT) has been ex-plored to enhance reasoning, purely textual CoT or coordinate basedprompts are fundamentally limited in representing intricate visual lay-outs and lack the pixel level cues necessary for precise editing. To addressthese challenges, we propose Unlocking Complex Image Editing viaMultimodal Reasoning Edit (MURE), a natively multimodal frame-work that shifts the editing process from purely verbal reasoning to asequence of native interleaved textual and visual rationales. |
Zhentao Zou; Zhengrong Yue; Kunpeng Du; Binglei Bao; Hanting Li; Haizhen Xie; Guozheng Xu; Yue Zhou; jie hu; Xue Jiang; Xinghao Chen; |
| 354 | Towards Generalizable Robotic Manipulation in Dynamic Environments Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This performance gap primarily stems from a scarcity of dynamic manipulation datasets and the reliance of mainstream VLAs on single-frame observations, restricting their spatiotemporal reasoning capabilities. To address this, we introduce DOMINO, a large-scale dataset and benchmark for generalizable dynamic manipulation, featuring 35 tasks with hierarchical complexities, over 110K expert trajectories, and a multidimensional evaluation suite. |
Heng Fang; Shangru Li; Shuhan Wang; Xuanyang Xi; Dingkang Liang; Xiang Bai; |
| 355 | Cycle-World: Mitigating Error Accumulation in Long-term Video World Models Via Reverse-Prediction Cycle Consistency Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To addressthis, we propose Cycle-World, a novel framework designed for stable andtemporally consistent long-video generation. |
Zihan Su; Teng Hu; Jiangning Zhang; Ruiyan Wang; Ran Yi; Lizhuang Ma; Dacheng Tao; |
| 356 | AnaPFL: When Closed-Form Solutions Meet Generalizationand Personalization in Personalized Federated Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, there remains a significant gap in introducingAL into PFL, owing to the encountered generalization-personalizationdilemma. In this paper, to bridge this gap and address the associatedchallenges, we propose an Analytic Personalized Federated Learningapproach, named AnaPFL, for addressing the Non-IID issue in PFL byintroducing and advancing AL. |
Kejia Fan; Jianheng Tang; Zhirui Yang; Feijiang Han; Yajiang Huang; Run He; Jiaxu Li; Songning Lai; Anfeng Liu; Houbing Herbert Song; Yunhuai Liu; HUIPING ZHUANG; |
| 357 | Transport Discrepancy As A Reliability Signal for Vision-Language-Action Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Distribution shift and long-horizonrollouts can push backbone representations away from the region the ac-tion head decodes reliably, yet the policy has no mechanism to detector react to this drift. We observe that the cost of transporting observa-tion features to the action representation in a shared feature space risesprecisely when such drift occurs, providing a per-step reliability esti-mate without extra supervision. |
Wanpeng Zhang; Ye Wang; Hao Luo; Haoqi Yuan; Yicheng Feng; Chaoyi Xu; Sipeng Zheng; Qin Jin; Zongqing Lu; |
| 358 | MonoArt: Progressive Structural Reasoning for Monocular Articulated 3D Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Reconstructing articulated 3D objects from a single imagerequires jointly inferring object geometry, part structure, and motionparameters from limited visual evidence. A key … |
Haitian Li; Haozhe Xie; Junxiang Xu; Beichen Wen; Fangzhou Hong; Ziwei Liu; |
| 359 | TDSR-VLA: Transition-aware Denoising Sequence Representations for Vision-Language-Action Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose TDSR-VLA, a VLA framework that reuses condi-tioning states from a diffusion-based Vision Planner to guide action gen-eration. |
Dong-Woo Kim; KEUNHO SONG; Seungmin Lee; Hwanhee Ju; Eun Cha; Daekyum Kim; |
| 360 | Seeing Fast and Slow: Learning The Flow of Time in Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In thispaper, we study time as a learnable visual concept and develop modelsfor reasoning about and manipulating the flow of time in videos.We first exploit the multimodal cues and temporal structure naturallypresent in videos to learn, in a self-supervised manner, to detect speedchanges and estimate playback speed. |
Yen-Siang Wu; Rundong Luo; Jingsen Zhu; Tao Tu; Ali Farhadi; Matthew Wallingford; Yu-Chiang Frank Wang; Steve Marschner; Wei-Chiu Ma; |
| 361 | Exploring Efficient Reasoning Segmentation with Small Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address the challenges of scaling reasoning segmentationto SLMs, we propose two key designs: First, unlike LLMs, SLMs havelimited capacity to provide sufficient spatial instruction cues for mask de-coding. We introduce register tokens that aggregate text-conditioned spa-tial features via the SLM’s self-attention and inject them into the visualtoken stream to enrich the mask decoder’s inputs. |
Changsong Wen; Zelin Peng; Yu Huang; Xiaokang Yang; Wei Shen; |
| 362 | Jumping The Landing Phase: Noise Variance Matching Enables Accurate Few-Step Inversion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We find that during the initial iterations, thepredicted variance rapidly increases toward its correct scale—an earlystage we term the landing phase, which accounts for most of ReNoise’scomputational cost. Based on this insight, we propose Noise VarianceMatching (NVM), a simple and efficient strategy that explicitly alignsthe predicted noise variance with that of a forward reference sample,bypassing the landing phase. |
Yang Luo; Zhineng Chen; Ya Gao; Xieping Gao; Yu-Gang Jiang; |
| 363 | DC-Gen: Post-Training Diffusion Acceleration with Deeply Compressed Latent Space Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Existing latent diffusion models excel at visual generationand editing tasks, employing autoencoders to project RGB images andvideos into latent spaces. However, their … |
Wenkun He; Yuchao Gu; Junyu Chen; Junyi Wu; Wenhang Ge; Dongyun Zou; Yujun Lin; Zhekai Zhang; Haocheng Xi; Muyang Li; Ligeng Zhu; Jincheng YU; Junsong Chen; Enze Xie; Song Han; Han Cai; |
| 364 | VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To reduce redundant tool-calling in theearly RL-stage and accelerate convergence, we propose Toolkit ActionGrouped Policy Optimization (TAGPO), which enhances the efficiencyof the calling process through the finer stepwise reward assignment. |
Zhe Gao; Shiyu Shen; Taifeng Chai; Weinong Wang; Haotian Xu; Xing W; Wenbin Li; Qi Fan; Yang Gao; Dacheng Tao; |
| 365 | There and Back Again: A Flexible-Frame Transformer for Multi-Exposure Fusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, conventional MEF techniquesare typically designed for a fixed number of inputs, forcing deploymentsystems to maintain separate models for different frame-count require-ments, which undermines deployment efficiency. To address this limita-tion, we propose FreeMEF, the first flexible-frame transformer for MEFthat seamlessly accommodates varying numbers of input exposures with-out retraining or architectural changes. |
Lishen Qu; Yao Liu; shihao zhou; Jie Liang; Hui Zeng; Lei Zhang; Jufeng Yang; |
| 366 | Region-Aware Multimodal Large Language Model Via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current CT report generation frameworks predominantlyrely on global feature representations, often failing to capture region-specific details and potentially missing certain abnormalities. To over-come this limitation, we propose MedRegion-CT, a region-focused mul-timodal large language model framework featuring three key innova-tions. |
Sunggu Kyung; Jinyoung Seo; Hyunseok Lim; Dongyeong Kim; Hyungbin Park; Jimin Sung; Wooyoung Jo; Yoojin Nam; Namkug Kim; |
| 367 | EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Multimodal large language models (MLLMs) are increas-ingly considered as a foundation for embodied agents, yet it remainsunclear whether they can reliably reason about the long-term phys-ical consequences of actions from an egocentric viewpoint. |
Chengjun Yu; Xuhan Zhu; Chaoqun Du; Pengfei Yu; Wei Zhai; Yang Cao; Zheng-Jun Zha; |
| 368 | Seek to Segment: Active Perception for Panoramic Referring Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing referring image segmentation (RIS) models passively process static images captured from fixed perspectives, limiting their applicability in Embodied AI, where agents must perform active perception in the continuous 360◦ environments. To bridge this gap, we introduce a novel task: Active Panoramic Referring Segmentation (APRS). |
Song Tang; Shuming Hu; Xincheng Shuai; Henghui Ding; Yu-Gang Jiang; |
| 369 | CtrlCoMo: Controllable Co-Speech Motion Generation with Gesture–Action Disentanglement Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address the interference between ges-tures and actions in co-speech motions, we introduce Pyramid-VQ, anautoencoder that separates gestures from actions through hierarchicalquantization, with shallow layers capturing global semantics and deeplayers encoding localized gestures, guided by layer-wise CLIP-based se-mantic regularization. |
Xinghan Wang; Ming Zhou; Yanbo Zheng; Youjiang Xu; Yuan Zhang; Mingyuan Gao; Nan Zhuang; Yadong Mu; |
| 370 | ReShift: Aha-Moment-Driven Reasoning-Level Backdoor Attacks on Vision–Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper,we propose ReShift, the novel aha-moment-driven reasoning-level back-door framework that explicitly redirects the internal chain-of-thought(CoT) trajectory while preserving surface-level coherence. |
Zhihao Dou; Qinjian Zhao; Zhiqiang Gao; Sumon Biswas; |
| 371 | SegDiff: Segmented Trajectory Diffusion for Consistent and Adaptive Robot Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, continuous prediction methods suffer fromcompounding errors due to short prediction horizons and struggle withmulti-modal action distributions, whereas keypose-based methods ne-cessitate an external planner, constraining real-time applicability. Toaddress these challenges, we introduce SegDiff, a closed-loop visuomo-tor policy that integrates the strengths of both paradigms. |
Haidong Cao; Wenjun Cao; Quanhao Li; Sicheng Xie; Zhiying Du; Jiaqi Leng; Zuxuan Wu; Yu-Gang Jiang; |
| 372 | FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose FlexComposer, a unified framework thatstandardizes video compositing as a trajectory-guided conditional gen-eration task, enabling the seamless integration of both static images anddynamic footage. |
Songchun Zhang; Sitong Guo; Xianghao Kong; Pengwei Liu; Yuwei Guo; lvmin zhang; Anyi Rao; |
| 373 | Diffusion-SDPO: Safeguarded Direct Preference Optimization for Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Consequently, degradation ofthe less-preferred outputs can become sufficiently severe that the pre-ferred branch is also adversely affected even as the margin grows. Toaddress this, we introduce Diffusion-SDPO, a safeguarded update rulethat preserves the winner by adaptively scaling the loser gradient ac-cording to its alignment with the winner gradient. |
Minghao Fu; Guo-Hua Wang; Tianyu Cui; Qing-Guo Chen; Zhao Xu; Weihua Luo; Kaifu Zhang; |
| 374 | LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Crucially, stronger search algorithms are precisely what couldunlock substantial gains in controllability, sample efficiency and genera-tion quality for video diffusion, provided their computational cost can bereduced. To fill in this gap, we enable efficient inference-time scaling forvideo diffusion through latent reward guidance, which provides interme-diate, informative and efficient feedback along the denoising trajectory.We introduce a latent reward model that scores partially denoised la-tents at arbitrary timesteps with respect to visual quality, motion qual-ity, and text alignment. |
Zengqun Zhao; Ziquan Liu; Yu Cao; Shaogang Gong; Zhensong Zhang; Song Jifei; Jiankang Deng; ioannis Patras; |
| 375 | Geo-DPO: Aligning Semantic Intent with Geometry for 3D Affordance Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This leads to severe geometric ambiguity, causing the predicted masks to have inaccurate boundaries. To solve this, we propose the Geo-DPO framework, which addresses the ambiguity from two directions. |
Ziqian Yang; Xiaolei Wang; Xianglin Qiu; Weiguang Zhao; Quan Zhang; Jimin Xiao; |
| 376 | Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoE Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Such stochastic noise obscures the critical structural and textural information, thereby preventing the router from effectively distinguishing salient tokens. To address this, we propose SharpMoE, a post-training framework with a saliency-harnessing accurate routing mechanism, which utilizes clean latent features as a noise-free guidance signal for routing. |
Haoyou Deng; Keyu Yan; Chaojie Mao; Xiang Wang; Yu Liu; Changxin Gao; Nong Sang; |
| 377 | PLOT: Pseudo-Labeling Via Object Tracking for Monocular 3D Object Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we presentPLOT (Pseudo-Labeling via Object Tracking), a framework that gen-erates 3D annotations from monocular videos without auxiliary sen-sors or model retraining. |
SeokYeong Lee; Sithu Aung; JunYong Choi; Seungryong Kim; Ig-Jae Kim; Junghyun Cho; |
| 378 | Generative Manifold Distillation: Aligning Restoration Trajectories with The Natural Image Prior Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Pre-trained image restoration models often fail on out-of-distribution (OOD) real-world degradations. Adapting to these domainsis challenging as real-world data lacks paired … |
Yuyang Hu; Mojtaba Sahraee-Ardakan; Kangfu Mei; Arpit Bansal; Chenyang Qi; Peyman Milanfar; Mauricio Delbracio; |
| 379 | Q-TriM: Question-Guided Tri-Modal Attention for Audio–Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: For Q-TriM, we propose a novel framework for attention operation incorporating video and audio conditioned on text. |
SungHun Kim; Seung Baek; |
| 380 | OCTOPUS: Multi‑Agentic Universal Compositional Visual Retrieval Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce OCTOPUS, a novelmulti-agentic assistant for tool-integrated progressive self-improvementand user-friendly synergistic compositional retrieval. |
Zhangtao Cheng; Bozhu Zheng; Ting Zhong; Fan Zhou; |
| 381 | Accelerating Diffusion Transformers with Gaussian Process Rectified Feature Cache Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Prediction-based feature caching is widely usedto accelerate diffusion transformers; however, as the number of stepsincreases, the deviation between its predictions and the reference full-compute trajectory gradually grows. An intuitive idea is to use an onlineregression model to dynamically correct this deviation, but it faces theissue of label data being unavailable during the acceleration process.This paper presents a statistical observation that the residuals betweenthe features of full computation steps using caching methods and refer-ence full-compute trajectory locally exhibit a zero-mean Gaussian dis-tribution. |
Zhirong Shen; Rui Huang; Chang Zou; Shikang Zheng; Jiacheng Liu; Peiliang Cai; zhengyi shi; Yaosong Du; Liang Feng; Xiaobing Tu; Jinkui Ren; Xiantao Zhang; Linfeng Zhang; |
| 382 | OSOR: One-Step Diffusion Inpainting for Effect-Aware Object Removal Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: With billions of parameters and tensof denoising steps, diffusion-based models achieve this goal at the ex-pense of massive computational cost, limiting their use in interactiveapplications and edge devices. To solve this problem, we present OSOR(One-Step Object Removal), which achieves efficient, effect-aware, andmask-robust object removal at the same time. |
Qinming Zhou; Chenxi Sun; Deyang Kong; Junhao He; Xiangheng Tang; Peike Yu; Haotian Wu; Leilei Cao; Linfeng Zhang; |
| 383 | LinCa: Accelerating Diffusion Models Via Learnable Decomposed Feature Caching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose LinCa , a fea-ture caching framework based on learnable invertible networks. |
Jinshan Liu; Haoran Qin; Xiaobing Tu; Jiacheng Liu; Jiahui Hu; zhengan yan; Yukun Xie; Kerui Shen; Jinkui Ren; Yuqi Lin; Xiantao Zhang; Linfeng Zhang; |
| 384 | AViTS:Adaptive Spatiotemporal Token Selection for Efficient Dynamic-Resolution Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose AViTS, an adaptive spatiotemporal token selection framework for dynamic-resolution DiTs. |
Haoran Qin; zhengan yan; Shikang Zheng; Xiaobing Tu; Jiacheng Liu; Yuqi Lin; Chang Zou; Jinshan Liu; Peiliang Cai; Xiantao Zhang; Jinkui Ren; Linfeng Zhang; |
| 385 | Scalable Cross-embodiment Dexterous Grasping Via Morphology-Prior Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents SOMO, a scalable framework for crossembodiment grasp synthesis that transfers to novel robot hands using only their hand description (i.e., a Unified Robot Description Format (URDF) file), without requiring any hand–object interaction annotations. |
Sihang Li; Zheming Zhou; Marcelino Almeida; Omid Alizadeh; Luca Carlone; Min Sun; Chen Feng; Cheng-Hao Kuo; |
| 386 | DiscoVL: Unveiling Disentangled Cross-Modal Representation Learning Via Orthogonal Adversarial Regularization for Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work,we present DiscoVL, a disentangled cross-modal representation learningframework that couples orthogonal adversarial regularization with struc-tured cross-modal alignment for vision-language models. |
Mengping Dong; Jinbao Li; Fei Li; |
| 387 | MobileSAM2: Lightweight Segment Anything in Images and Videos Via Hypergraphical Knowledge Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Besides,we present MobileSAM2, a new family of lightweight SAM2 that balancesefficiency and effectiveness via searching the best model architectures withHyperKD during model size reduction. |
Kai Jiang; Jiaxing Huang; Jingyi Zhang; Weiying Xie; Yunsong Li; Yufei Wang; Aoran Xiao; Dacheng Tao; |
| 388 | HitMem: Hierarchical Temporal 3D Memory with Multi-Modal Context-Aware Retrieval for Dynamic Environments Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: When objects are displaced by human activities orunobserved events, agents encounter memory-observation conflicts andoften require costly geometric recomputations or inefficient global re-exploration. To address this, we propose HitMem, a hierarchical tempo-ral 3D memory framework with a multi-modal context-aware retrievalmechanism. |
Ruijie Tang; Chenye Zou; Guoquan Wu; Jun Wei; Wei Chen; Jiaxin Zhu; |
| 389 | HolisticSemGes: Semantic Grounding of Holistic Co-Speech Gesture Generation with Contrastive Flow-Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a Contrastive Flow Matching-basedco-speech gesture generation model that uses mismatched audio–textconditions as negatives, training the velocity field to follow the cor-rect motion trajectory while repelling semantically incongruent trajec-tories. |
Lanmiao Liu; Esam Ghaleb; asli ozyurek; Zerrin Yumak; |
| 390 | Stand Up and Move: Benchmarking Interactive Spatial Intelligence in WalkerBench Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our empirical evaluation revealsa critical “representation misalignment”: VLMs’ one-dimensional linearcontext structures are fundamentally incompatible with the topologicalnature of 3D environments, causing inevitable spatial forgetting duringnavigation. To overcome this, we propose Spatial-IDE, which breaks fromlinear dialogue history via two mechanisms: (1) State Externalization,transforming implicit observations into an Explicit Topological Memory(ETM); and (2) Cognitive Decoupling, disentangling visual understand-ing into Goal-Directed Perception and Spatial Reasoning, letting theVLM focus on pure high-level decision-making. |
Zhiqi Ge; Gang Yang; Ziyang Pan; Jingzhe Zhu; Yuancheng Gu; Juncheng Li; Qizhou Wang; Rui Tang; Siliang Tang; Jun Xiao; Yueting Zhuang; |
| 391 | Learning from Reliable Negatives: Confidence-Anchored Test-Time Adaptation for GUI Grounding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce a label-free test-timetraining paradigm driven by two key insights: (1) confidence patterns incoordinate tokens are a better indicator than full-sequence confidence,and (2) in sparse GUI coordinate spaces, negative samples offer more re-liable learning signals than potentially noisy positive ones. |
Yizhou Liu; Fei Tang; Yuchen Yan; Zhengxi Lu; Songqin Nong; Tao Jiang; Wenhao Xu; Wenqi Zhang; Weiming Lu; Jun Xiao; Yongliang Shen; |
| 392 | WildWorld: A Large-Scale Dataset for Action-Conditioned World Modeling with Explicit State Annotations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inthis paper, we propose WildWorld, a large-scale action-conditioned worldmodeling dataset with explicit state annotations, automatically collectedfrom a photorealistic AAA action role-playing game (Monster Hunter:Wilds). |
Zhen Li; Zian Meng; Chuanhao Li; SHUWEI SHI; Wenshuo Peng; Yuwei Wu; Bo Zheng; Yunde Jia; Kaipeng Zhang; |
| 393 | VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce VLZip, a framework that uni-_x001C_es visual and textual compression for high-_x001C_delity reasoning within apure Transformer. |
Yuqi Zhang; Cheng Chen; Yuyu Guo; Wenjie Yang; Lingchen Meng; Peng Di; Hang Yu; Zuxuan Wu; Yu-Gang Jiang; |
| 394 | WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We attribute thesefailures primarily to the lack of specialized training paradigms tailoredfor text-centric editing, as well as the absence of large-scale datasetsand standardized benchmarks necessary for a closed-loop training andevaluation system. To address these limitations, we present WeEdit, a sys-tematic solution encompassing a scalable data construction pipeline, twobenchmarks, and a tailored two-stage training strategy. |
Hui Zhang; Juntao Liu; Zongkai Liu; liqiang niu; Fandong Meng; Zuxuan Wu; Yu-Gang Jiang; |
| 395 | TransVLM: A Vision-Language Framework and Benchmark for Detecting Any Shot Transitions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Traditional Shot Boundary Detection (SBD) inherently strug-gles with complex transitions by formulating the task around isolated cutpoints, frequently yielding corrupted video shots. We address this fun-damental limitation by formalizing the Shot Transition Detection (STD)task. |
Ce Chen; Yi Ren; Yuanming Li; Viktor Goriachko; Zhenhui Ye; Zujin Guo; Zhibin Hong; Mingming Gong; |
| 396 | Granular Semantic Cognition for Visible-Infrared Person Re-Identification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, this paper proposes a GranularSemantic Cognition (GSC) method, which leverages cross-modal sharedsemantics to facilitate the transfer of color features. |
Haifeng Yang; Jinjia Peng; Huibing Wang; |
| 397 | PKINet-v2: Towards Powerful and Efficient Poly-Kernel Remote Sensing Object Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we extend PKINet, and presenta powerful and efficient backbone that jointly handles both challengeswithin a unified paradigm named Poly Kernel Inception Network v2(PKINet-v2). |
Xinhao Cai; Liulei Li; Gensheng Pei; Zeren Sun; Yazhou Yao; Wenguan Wang; |
| 398 | Weather-Conditioned Depth Anything Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, they still suffer from critical failures under adverse weather conditions, such as fog, rain, snow, or at night. To address this, we present Weather-Conditioned Depth Anything (DA-W), a framework that explicitly disentangles style from content for weatherrobust depth estimation. |
Zhaoming Xu; Chan-Wei Hu; Kuan-Ru Huang; Zihao Zhu; Renjie Li; Yang Zhou; Zhengzhong Tu; |
| 399 | PhysRVG: Physics-Aware Unified Reinforcement Learning for Video Generative Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods attemptto tackle this problem through physical data augmentation, dynamicspre-simulation, or reinforcement learning with VLM ratings, but noneof these approaches accurately reflect physical principles or enable themodel to internalize physical knowledge. Motivated by these considera-tions, we introduce reinforcement learning with physically verifiable re-wards. |
Qiyuan Zhang; Biao Gong; Shuai Tan; Zheng Zhang; Xing Zhu; Yujun Shen; Yuyuan Li; kelu Yao; Chunhua Shen; Changqing Zou; |
| 400 | Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution Via Auto-Regressive Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Stream-DiffVSR, a causally conditioned diffusion framework for efficient online VSR. |
Hau-Shiang Shiu; Chin-Yang Lin; Zhixiang Wang; Chi-Wei Hsiao; Po-Fan Yu; Yu-Chih Chen; Yu-Lun Liu; |
| 401 | LUNA: Learning Universal 3D Human Animation Beyond Skinning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose LUNA, an LBSfree universal neural animation model that directly maps multiple 2D controls like images, keypoints, sketch and unseen characters into 3D Gaussian deformations, bypassing explicit body fitting. |
Peng Li; Rawal Khirodkar; Junxuan Li; Yuan Dong; Chen Cao; Yuan Liu; Wenhan Luo; Yike Guo; Shunsuke Saito; |
| 402 | From Synchrony to Sequence: Exo-to-Ego Generation Via Interpolation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While paired supervision is available, synchronized exo-egodata inherently introduces substantial spatio-temporal and geometricdiscontinuities, violating the smooth-motion assumptions of standardvideo generation benchmarks. We identify this synchronization-inducedjump as the central challenge and propose Syn2Seq-Forcing, a sequen-tial formulation that interpolates between the source and target videosto form a single continuous signal. |
Mohammad Mahdi; Nedko Savov; Danda Paudel; Luc Van Gool; |
| 403 | Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Moreover, due to the causal autoregressive process, inaccurate semantics at early scales can propagate and significantly degrade the final output. To address these issues, we introduce a scale-aware token-routed Mixture of Experts (MoE) architecture, allowing scale-adaptive expert selection, thereby facilitating decoupled representation learning across scales. |
Nuoyan Zhou; Zhijun Tu; Lei Yu; Kun Cheng; jie hu; Nannan Wang; Xinghao Chen; |
| 404 | SkillSpotter: Pose-Aware Multi-View Skilled Action Detection and Grading in Ego-Exo Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: SkillSpotter ’s modulestransfer to other temporal action detection models with consistent gainsand our method generalizes beyond Ego-Exo4D to HoloAssist.Code: https://github.com/eth-siplab/SkillSpotter |
Björn Braun; Christian Holz; |
| 405 | Kinematics-Agnostic 3D Human Motion Prediction Via Equivariant Latent Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing Stochastic 3D Human Motion Prediction models are fundamentally constrained by hard-coding the skeleton kinematics, severely limiting generalization, preventing cross-dataset training, and requiring complex data retargeting. We introduce EquiFusion, the first kinematics-agnostic model to solve this bottleneck, implementing a latent diffusion model with a permutation equivariant architecture. |
Cecilia Curreli; Florian Hofherr; Dominik Muhle; Abhishek Saroha; Riccardo Marin; Daniel Cremers; |
| 406 | Geometry-Aware Single-Image 4D Synthesis Via Dense Trajectory Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address these, we present MoGe4D (Motion and GeometryAware image-to-4D Synthesis), a geometry-conditioned framework for single-image 4D synthesis that models a scene as dense 4D point trajectories. |
Yanran Zhang; Ziyi Wang; Wenzhao Zheng; Zheng Zhu; Jie Zhou; Jiwen Lu; |
| 407 | MetaPoint: Unlocking Precise Spatial Control in Visual Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This arises from a core disconnect: models can process textual descriptions of space but cannot directly map numerical coordinates onto the 2D image canvas (as illustrated in Fig. 2). We introduce MetaPoint, a method that bridges this gap by representing a continuous 2D coordinate as a single, special token. |
Dewei Zhou; Xinyu Huang; Xun Wang; Ji Xie; Yabo Zhang; Liang Li; Kunchang Li; Zongxin Yang; Yi Yang; |
| 408 | FeVOS: Foresight Expression Video Object Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing Referring Video Object Segmentation tasks focus on referring expressions describing events, actions or appearances of relevant objects within the observed frames, lacking evaluation in scenarios that require pre-decisive spatio-temporal reasoning, thereby limiting their applicability. To address this, we propose Foresight Expression Video Object Segmentation, a task that queries future events in upcoming video segments and requires masks of the objects in the observed frames as visual answers. |
Kehan Lan; Kaining Ying; Henghui Ding; |
| 409 | VLA-Hijack: A Transferable Patch Attack Against Vision-Language-Action Models Via Visual Proprioception Hijacking Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While Vision-Language-Action (VLA) models have emergedas powerful generalist policies, their severe vulnerability to adversarialpatches significantly hinders their deployment in safety-critical domains.Moreover, existing patch attacks primarily focus on white-box settings,heavily overfitting to the specific action output space of the target model,which results in poor cross-architecture transferability. To overcome thislimitation, we propose VLA-Hijack, a unified adversarial framework thatbreaks the transferability bottleneck by exploiting a fundamental vulner-ability identified in this work: before planning any motion, a VLA modelmust first use visual information to locate its own robotic arm withinthe environment. |
Jiyuan Fu; Kaixun Jiang; Jingkai Jia; Zhaoyu Chen; Xueyao Chen; Lingyi Hong; Shuyong Gao; Chenzhi Tan; Dingkang Yang; Wenqiang Zhang; |
| 410 | TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce TIR-Bench, acomprehensive benchmark for evaluating agentic thinking-with-imagesacross 13 diverse tasks, each requiring novel tool use for image process-ing and manipulation in chain-of-thought. |
Ming Li; Jike Zhong; Shitian Zhao; Haoquan Zhang; Shaoheng Lin; Yuxiang Lai; Chen Wei; Konstantinos Psounis; Kaipeng Zhang; |
| 411 | RT-DETRv4: Painlessly Furthering Real-Time Object Detection with Vision Foundation Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a cost-effective and highly adaptable distillation framework that harnesses the rapidly evolving capabilities of Vision Foundation Models (VFMs) to enhance lightweight object detectors. |
Zijun Liao; Yian Zhao; Xin Shan; Yu Yan; Chang Liu; lei lu; Xiangyang Ji; Jie Chen; |
| 412 | GroundingAnomaly: Spatially-Grounded Diffusion for Few-Shot Anomaly Synthesis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing methods either suffer from poor integration caused by inpainting or fail to provide accurate masks. To address these limitations, we propose GroundingAnomaly, a novel few-shot anomaly image generation framework. |
Yishen Liu; Hongchang Chen; Pengcheng Zhao; Yunfan Bao; Yuxi Tian; Jieming Zhang; HAO CHEN; Zhi Zheng; Yongchun Liu; Ying Li; Dongpu Cao; |
| 413 | SubSplat: High-Resolution Pixel-aligned 3DGS Via Sub-pixel Gaussian Reparameterization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our results validate thatthe proposed framework successfully resolves the trade-off between repa-rameterization fidelity and network computational cost inherent in pixel-aligned Gaussian Splatting. |
Jiun Lee; Jaekwang Kim; Sangmin Lee; |
| 414 | QASA: Quality-Guided K-Adaptive Slot Attention for Unsupervised Object-Centric Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: K-fixed methods, such as SPOT [20]and DINOSAUR [33], show large performance fluctuations as K varies. |
Tianran Ouyang; Xingping Dong; Jing Zhang; Mang Ye; Kaihao Zhang; Bo Du; |
| 415 | WinTok: A Win-Win Hybrid Tokenizer Via Decomposing Visual Understanding and Generation with Transferable Tokens Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose WinTok, a concisehybrid tokenizer that achieves a win-win performance by explicitly de-coupling the two objectives. |
Yiwei Guo; Shaobin Zhuang; Zhipeng Huang; Canmiao Fu; Chen Li; Jing LYU; Yali Wang; |
| 416 | MoCA3D: Monocular 3D Bounding Box Prediction in The Image Plane Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce MoCA3D, a Monocular,Class-Agnostic 3D model that predicts projected 3D bounding box cor-ners and per-corner depths without requiring camera intrinsics at in-ference time. |
Changwoo Jeon; Rishi Upadhyay; Achuta Kadambi; |
| 417 | ReconPhys: Reconstruct Appearance and Physical Attributes from Single Video Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing approaches leverage differen-tiable rendering for per-scene optimization, recovering geometry anddynamics but requiring expensive tuning or manual annotation, whichlimits practicality and generalizability. To address this, we propose Re-conPhys, the first feedforward framework that jointly learns physicalattribute estimation and 3D Gaussian Splatting reconstruction from asingle monocular video. |
Boyuan Wang; Xiaofeng Wang; Yongkang Li; Zheng Zhu; Yifan Chang; Angen Ye; Guosheng Zhao; Chaojun Ni; Guan Huang; Yijie Ren; Yueqi Duan; Xingang Wang; |
| 418 | VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Pretraining Vision-Language-Action (VLA) policies on internet-scale video is appealing, yet current latent-action objectives often learnthe wrong thing: they remain anchored to pixel variation rather thanaction-relevant state transitions, making them vulnerable to appearancebias, nuisance motion, and information leakage. We introduce VLA-JEPA,a JEPA-style pretraining framework that sidesteps these pitfalls by design.The key idea is leakage-free state prediction: a target encoder produceslatent representations from future frames, while the student pathwaysees only the current observation—future information is used solely assupervision targets, never as input. |
Jingwen Sun; Wenyao Zhang; Zekun Qi; Shaojie Ren; Zezhi Liu; Hanxin Zhu; Guangzhong Sun; Xin Jin; Zhibo Chen; |
| 419 | SAFE-Pruner: Semantic Attention–Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While visual token pruning has shown strong potential for accelerating inference, most existing methods mainly base pruning decisions on shallow-layer cues and risk discarding visual information required by deep layers. To address this issue, we propose SAFE-Pruner, a plug-and-play pruning framework that incorporates attention cues of future layers into pruning decisions. |
Shilin Ma; Chubin Zhang; Changyuan Wang; Yuji Wang; Yue Wu; Zixuan Wang; Jingqi Tian; Zheng Zhu; Yansong Tang; |
| 420 | DiCoBench: Benchmarking Multi-Image Fine-Grained Perception Via Differential and Commonality Visual Cues Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To bridgethis gap, we introduce DiCoBench, a comprehensive, multi-image high-resolution benchmark designed for cross-image fine-grained perception.DiCoBench consists of 765 meticulously curated samples categorizedinto two progressive tracks: Differential Visual Cues and Commonal-ity Visual Cues, covering 8 distinct perception tasks.By formulatingthe benchmark as a multiple-choice question task and utilizing high-resolution imagery (approaching 2K), we eliminate evaluation metricbias and pose a substantial challenge to current state-of-the-art MLLMs.Our extensive evaluation of 18 diverse MLLMs reveals a striking perfor-mance gap compared to human accuracy (98.3%), with top-performingmodels struggling significantly with micro-scale detail capture. |
Geng Li; Yuxin Peng; |
| 421 | GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While Large Vision-Language Models (LVLMs) have signifi-cantly advanced GUI agents’ capabilities in parsing textual instructions,interpreting screen content, and executing tasks, a critical challenge per-sists: the irreversibility of agent operations—where a single erroneous ac-tion can trigger catastrophic deviations. To address this, we propose theGUI Action Critic’s Data Flywheel System (GAIA), a training frame-work that enables the models to have iterative critic capabilities, whichare used to improve the Test-Time Scaling (TTS) of basic GUI agents’performance. |
Shaokang Wang; Pei Fu; Ruoceng Zhang; Shaojie Zhang; Xiuwen Xi; Jiahui Yang; Bin Qin; Ying Huang; Zhenbo Luo; Jian Luan; |
| 422 | MoScale: Autoregressive Next-Scale Prediction for Human Motion Generation and Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present ScaleMoGen, a scale-wise autoregressive frame-work for text-driven human motion generation. |
Inwoo Hwang; Hojun Jang; Bing Zhou; Jian Wang; Young Min Kim; chuan guo; |
| 423 | GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: BrowseComp offers a text-only testbed for such agents,but existing multimodal benchmarks rarely require both weak visual cuescomposition and BrowseComp-style multi-hop verification. Geolocationis a natural testbed because answers depend on combining multipleambiguous visual cues and validating them with open-web evidence.Thus, we introduce GeoBrowse, a geolocation benchmark that combinesvisual reasoning with knowledge-intensive multi-hop queries. |
Xinyu Geng; Yanjing Xiao; Yuyang Zhang; Hanwen Wang; Xinyan Liu; Rui Min; Tianqing Fang; Yi Ren Fung; |
| 424 | OpenCVL: An Open, Diverse, and Large-Scale Dataset for Fine-Grained Cross-View Localization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While in-the-wild images are abundant, their noisy geo-tags make them unsuit-able for reliable evaluation. To bridge this gap, we introduce OpenCVL,a large-scale, diverse, and open dataset containing 617,388 ground-aerialimage pairs spanning 41 cities across four European countries. |
Zimin Xia; Mubariz Zaffar; Junsheng Fu; Alexandre ALahi; Julian Kooij; |
| 425 | RASA: Disentangled Spatial-Motional Priors for Cross-Identity Character Animation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inthis paper, we introduce Reference-Aware Structural Alignment(RASA), a novel framework that systematically disentangles spatialmapping from motion control by injecting structured priors into a Diffu-sion Transformer (DiT). |
Zhen Xiao; Zhen Shen; Zhaofan Qiu; Ting Yao; Xueliang Liu; Tao Mei; |
| 426 | Wat3R: Underwater 3D Geometry Learning Without Underwater Annotations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Pioneering methods rely onmassive dense annotations that are impractical in underwater settings.In this paper, we propose Wat3R, a cross-domain semi-supervised learn-ing framework designed to adapt feed-forward 3D reconstruction modelsfrom air to underwater scenes. |
Jiangwei Ren; Xingyu Jiang; Zijie Song; Wei Xu; Hongkai Lin; Dingkang Liang; Xiang Bai; |
| 427 | KineBench: Benchmarking Embodied World Models Via IDM-Free Kinematic Grounding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This creates an unavoid-able attribution ambiguity between world model inaccuracies and ex-tractor errors. To reduce this ambiguity, we present KineBench, anIDM-free closed-loop benchmark for EWMs, built upon an explicit kine-matic grounding pipeline. |
Zeyu Liu; Zhangzhe Zhu; Yang Zhang; Chenyou Fan; Chenjia Bai; Xuelong Li; |
| 428 | MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose MIMFlow, a unifiedend-to-end framework that jointly optimizes latent semantics, pixel re-construction, and generative flow. |
Yang Chen; Xiaowei Xu; Shuai Wang; Xinwen Zhang; Qiushi Guo; Tiezheng Ge; Limin Wang; |
| 429 | Skin-R1: Clinical Knowledge-Guided Dermatological Diagnosis Using Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, their trustworthiness and clinical utility remain limited by three key challenges: heterogeneous datasets with inconsistent diagnostic labels and concept annotations, the lack of grounded diagnostic rationales for reliable reasoning supervision, and limited scalability when transferring knowledge from small, densely annotated datasets to large collections with sparse labels. To address these challenges, we propose Skin-R1, a dermatology-oriented VLM that integrates textbook-grounded clinical reasoning supervision with reinforcement learning (RL) to improve the accuracy and robustness of diagnostic prediction. |
Zehao Liu; Weijieying Ren; Jipeng ZHANG; Tianxiang Zhao; Jingxi Zhu; Xiaoting Li; Vasant Honavar; |
| 430 | SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although recent self-improving models that iterativelyrefine themselves offer a feasible solution, they still suffer from two corechallenges: (i) most existing methods augment visual or textual data sep-arately, resulting in discrepancies in data complexity (e.g., over-simplifieddiagrams paired with redundant textual descriptions); and (ii) the evo-lution of data and models is also separated, leading to scenarios wheremodels are exposed to tasks with mismatched difficulty levels. To ad-dress these issues, we propose SyncLoop, an automatic, closed-loopself-improving framework that jointly evolves both training data andmodel capabilities. |
xiuwei chen; Wentao Hu; Hanhui Li; Yongxin Wang; Jun Zhou; Zisheng Chen; Meng Cao; Yihan Zeng; Kui Zhang; Yu-Jie Yuan; Jianhua Han; Hang Xu; Xiaodan Liang; |
| 431 | UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing research rarely extends this paradigm to the tactile domain, where both object-level semantics and sensor-level configurations jointly determine the meaning of touch. To address this gap, we propose UniTac, the first UMM designed for tactile understanding and generation. |
Jiahang Tu; Fengyu Yang; Chenyang Ma; Xihang Yu; Ziyao Zeng; Shaokai Wu; Hanbin Zhao; Zhi Tao; Chao Zhang; Hui Qian; Alex Wong; |
| 432 | X2SAM: Any Segmentation in Images and Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce X2SAM, a unified segmentation MLLMthat extends any-segmentation capabilities from images to videos.We further introduce the Video Visual Grounded(V-VGD) segmentation benchmark, which evaluates whether a modelcan segment object tracks in videos from interactive visual prompts.With a unified joint training strategy over heterogeneous image and videodatasets, X2SAM delivers strong video segmentation performance, re-mains competitive on image segmentation benchmarks, and preservesgeneral image and video chat ability. |
Hao Wang; Limeng Qiao; Chi Zhang; Lin Ma; Guanglu Wan; Xiangyuan Lan; Xiaodan Liang; |
| 433 | UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing UMMs face prominent challenges: (1) the inherent learning conflicts between visual understanding and generation tasks, leading to suboptimal modeling in both tasks; (2) different understanding and generation visual spaces impeding scalability; (3) over-reliance on task-specific data that neglects the duality of text-image understanding and generation. To address these challenges, we propose UniDDT, which leverages a Noisy ViT encoder along with a LLM to unify semantic encoding for visual generation and understanding tasks, while employing a separate diffusion decoder to decouple diffusion decoding from text decoding. |
Shuai Wang; Liang Li; Yang Chen; Ruopeng Gao; Yao Teng; Limin Wang; |
| 434 | Meric: A Unified Framework for Multimodal Music Generation and Retrieval Via Representation Space Anchoring Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Cur-rent methods face three major limitations: they are bottlenecked byscarce paired data, lack explicit one-to-many modeling, and isolate gen-eration from retrieval. To address these issues, we propose Meric, aunified framework for multimodal music generation and retrieval. |
Xihua Wang; Yinbo Wang; Jingchao Zhang; Ruihua Song; |
| 435 | Table-MCR2TR: Merged-Cell-Aware Table Recognition Via Reinforced Multimodal Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While some works improve TR through large-scale training and global supervision, these local fine-grained yet crucial merged-cell attributes are often overlooked, becoming a key bottleneck for further progress. To tackle this, we introduce Table-MCR2TR, a reinforced multimodal large language model framework that leverages the enhanced merged-cell recognition (MCR) ability as contextual guidance to improve table recognition quality. |
Bangbang Zhou; Zhaoqing Zhu; Feiyu Gao; Hangdi Xing; Yadong Qu; Qi Zheng; Ming Yan; Hongtao Xie; |
| 436 | UC-VLM: Consistency-Driven Learning for AI-Generated Image Detection with Vision-Language Large Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present UC-VLM, a unified multi-stage frameworkfor AIGI detection that relies solely on binary supervision. |
Lei Tan; Shuwei Li; Mohan Kankanhalli; Robby T. Tan; |
| 437 | Thinking Ahead: Foresight Intelligence in MLLMs and World Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce FSU-QA, a VQA dataset for au-tonomous driving scenarios designed to advance research on ForesightIntelligence—the ability to anticipate and reason about complex, long-horizon futures. |
Zhantao Gong; Liaoyuan Fan; Qing Guo; Xun Xu; Xulei Yang; Shijie Li; |
| 438 | DINOv3D: 2D-3D Joint Optimization for Unified Spatial Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Conversely, freezing the 2D model restricts 3D spatial perception. To resolve this issue, we propose DINOv3D, a joint optimization framework built upon DINOv3 that employs a homologous teacher-student architecture to establish a regularized integration between the 2D and 3D understanding. |
Bo Zhou; Jianzhe Gao; Zhihui Wang; Lingxiang Wu; Jinqiao Wang; Yazhou Yao; Wenguan Wang; |
| 439 | Optimizing Mesh Animation from Video Via Shape Flow Guidance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address thislimitation, we propose Shape Flow Guidance (SFG), a sequence of 3Dshapes derived from videos, which serves as explicit 3D supervision formesh animation. |
Jingqiao Xiu; Yicong Li; Angela Yao; |
| 440 | MVI2V: Human Centric Image to Video Generation with Multiview Consistent Appearance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose the Multiview EnhancedImage-to-Video Generation Model (MVI2V), which introduces additionalmulti-view images of a person or garment as reference inputs. |
Pengfei Liu; Mingyi Xu; Wentao Jiang; Tiezheng Ge; Ming Zeng; |
| 441 | 3D Gaussian Splatting Compression with Object Scalability Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a framework toward scalable, finer-grained object-level 3DGS compression. |
Ruixiang Xue; Tong Chen; Zhan Ma; |
| 442 | Incremental Online Scene Reconstruction By 3D Gaussian Triangulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although 3D Gaussian Splatting shows strong potential,most existing approaches require offline conversion of the optimized Gaus-sians into an intermediate implicit field for explicit mesh extraction,which hinders seamless integration with downstream tasks. To addressthis limitation, we propose a novel online framework that incrementallyreconstructs and updates high-fidelity explicit meshes by directly trian-gulating a dense geometric Gaussian representation, which supports bothhigh-quality rendering and incremental surface reconstruction. |
Yanjin Zhu; Shaofan Liu; Jianke Zhu; |
| 443 | V-Co: A Closer Look at Visual Representation Alignment Via Co-Denoising Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing co-denoising approaches of-ten entangle multiple design choices, making it unclear which are trulyessential. We therefore present V-Co, a systematic study of visual co-denoising in a unified JiT-based framework. |
Han Lin; Xichen Pan; Zun Wang; Yue Zhang; Chu Wang; Jaemin Cho; Mohit Bansal; |
| 444 | Perceiving Better Moments: Cover Frame Reselection and Enhancement for Live Photos with The Live2K Dataset Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Restoring such framesrequires simultaneous enhancement of spatial detail and color appearance, a taskconsiderably more challenging than ordinary super-resolution or color enhance-ment. To address this, we define the Live Photo Cover Frame Reselection andEnhancement (LPRE) task, which leverages the intrinsic cues available withineach Live Photo: the high-quality cover image as a structural and color reference,the user-reselected low-quality frame as the reconstruction target and several ad-jacent video frames providing temporal cues. |
Junyu Lou; Kai Chen; Weiyi You; Hui Zeng; Lei Zhang; Shuhang Gu; |
| 445 | VCP-DCN: Beyond Visual Concealed Property Via Depth Collaborative Network for Camouflaged Object Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Camouflaged Object Detection (COD) aims to identify andsegment camouflaged objects in complex environments, which are oftenconcealed because their color and texture are similar to the background.Several existing COD methods introduce depth maps to boost detec-tion performance via learning complementary RGB-D features, ignor-ing modality-specific characteristics of concealed objects in the depthdomain. To address this issue, we propose a depth collaborative net-work, called VCP-DCN, to mine distinguishable multi-modality featuresbeyond visual concealed prototype in depth domain. |
Songsong Duan; Xi Yang; Nannan Wang; |
| 446 | CoSPlan: Corrective Sequential Planning Via Scene Graph Incremental Updates Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Addressingthis, we propose Scene Graph Incremental updates (SGI), a noveltraining-free method to transform images into ‘textual’ scene graphs, en-abling step-by-step reasoning through iterative scene graph refinement.SGI yields an average of ≃ 4.4% ↑ on CoSPlan w/ generalization onPlanBench and VQA. |
Shresth Grover; Priyank Pathak; Akash Kumar; Yogesh Rawat; |
| 447 | SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Measuring structured object understanding in vision founda-tion models remains challenging due to inconsistent evaluation protocolsand limited part-level supervision. Semantic … |
Olaf Dünkel; Basavaraj Sunagad; Haoran Wang; David Hoffmann; Christian Theobalt; Adam Kortylewski; |
| 448 | SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However,matching the performance of KL-VAE still required adversarial losses,as well as a higher decoding time due to iterative sampling. To addressthese limitations, we introduce a new pixel diffusion decoder architecturefor improved scaling and training stability, benefiting from transformercomponents and GAN-free training. |
Théophane Vallaeys; Jakob Verbeek; MATTHIEU CORD; |
| 449 | AVSplat:Dense-View Feed-Forward 3D Gaussian Splatting with Assist-View Preconditioning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present AVSplat, a framework that turns additionalviews into reliable signals for both aggregation and representation. |
MUYU XU; Fangneng Zhan; Yu Wei; Hanspeter Pfister; Shijian Lu; |
| 450 | GAP-Track: Bridging The Resolution Gap for Cross-Resolution RGBT Tracking Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose GAP-Track, an efficientframework that bridges the resolution gap by enabling high-precisiontracking of low-resolution inputs. |
Shiyu Zhang; Tianyang Xu; Zhangyong Tang; Wang He; Xiao-Jun Wu; Josef Kittler; |
| 451 | Together, Then Apart: Balancing Alignment and Distinctiveness for Multimodal Survival Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This motivates a representation learning principle that we refer to as Together Then Apart. Based on this idea, we propose TTA, a framework that balances cross-modal alignment and representation distinctiveness. |
Wenjing Liu; Qin Ren; Wen Zhang; Yuewei Lin; Chenyu You; |
| 452 | Seeing Touch from Motion: A Unified Modality-Aware Visuo-Tactile Policy with Tactile Motion Correlation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we take advantage of the Mixture-of-Transformersarchitecture and propose a unified modality-aware visuo-tactile policythat captures cross-modal complementarity while maintaining modality-specific properties. |
Shengqi Xu; Yang Liu; Guojin Zhong; Fanjie Wang; Hu Luo; Hanyu Zhou; WeiYao Zhang; Ziyi Ye; Zuxuan Wu; Yu-Gang Jiang; |
| 453 | Degradation-Robust and Temporally Consistent Infrared–Visible Video Fusion Via One-step Diffusion Framework Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, such designs may become vulnerable to modality-specific degradations, including low illumination in visible videos andnoise or flicker in infrared sequences, which often lead to unreliablemotion estimation and temporally inconsistent fusion results. To ad-dress this limitation, we advocate a ‘fusion-first’ strategy and proposeDRT-VF, a unified degradation-robust and temporally consistent in-frared–visible video fusion method built upon a one-step diffusion ar-chitecture with a two-stage progressive training strategy. |
Songcheng Du; HaoYuan Xu; Xingyuan Li; Yang Zou; Jinyuan Liu; YUNPENG BAI; Ying Li; |
| 454 | Thermo-JEPA: Learning A Geometry-Grounded Thermal World Model Via Cross-Modal Privileged Masking Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, weidentify a fundamental challenge in extending standard masked modelingto the thermal domain: intra-frame spatial homogenization. |
Biwen Yang; Jin Zhang; Zhe Cao; Ruiheng Zhang; |
| 455 | GridFlow: Structured Latent Flow for Seamless City-Scale 3D Point Cloud Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To sup-port standardized evaluation, we build on public 3D data sources to in-troduce City3D-MultiGen, a benchmark of 163K densely annotated tilesfrom Melbourne and London with aligned point clouds, satellite images,semantic maps, and elevation data. |
Xinyu Wang; Muhammad Ibrahim; Atif Mansoor; Ajmal Mian; |
| 456 | VolSplat: Rethinking Feed-Forward 3D Gaussian Splatting with Voxel-Aligned Prediction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We rethink this widelyadopted formulation and identify several inherent limitations: it rendersthe reconstructed 3D models heavily dependent on the number of inputviews, leads to view-biased density distributions, and introduces align-ment errors, particularly when source views contain occlusions or lowtexture. To address these challenges, we introduce VolSplat, a new multi-view feed-forward paradigm that replaces pixel alignment with voxel-aligned Gaussians. |
Weijie Wang; Yeqing Chen; Zeyu Zhang; Hengyu Liu; Haoxiao Wang; ZhiYuan Feng; Wenkang Qin; Feng Chen; Jiawang Bian; Zheng Zhu; Donny Y. Chen; Bohan Zhuang; |
| 457 | MArFE: Multi-Contrast MRI Arbitrary Scale Super-Resolution with Fourier Enhancement Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: How-ever, related research exists two following issues: 1) Implicit neural rep-resentations (INR) as mainstream methods are prone to spectral bias,which limits their ability to recover high-frequency details; 2) Multi-contrast MRI is often used as the effective prior, but lacks the targetednetwork design to further merge INR positional information. To solvethese problems, we propose a Fourier-enhanced implicit framework forarbitrary-scale multi-contrast MRI super-resolution (MArFE). |
ZHIWEN SHI; |
| 458 | Wavelet-based Intra-video Counterfactual Reasoning for Video Question Grounding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Prior debiasing methods primarily focus on dataset-level correla-tions, while easily overlooking intra-video temporal bias, where visuallysimilar segments differ substantially in their causal relevance to answer-ing the question. To address this challenge, we present Wavelet-basedIntra-video Causal Intervention (WICI), a framework designed to dis-entangle causal temporal evidence from redundant visual contexts. |
Jiasheng Yuan; Wei Wei; |
| 459 | FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Music-to-dance generation aims to translate auditory sig-nals into expressive human motion. Despite promising progress, existingmethods remain underexplored in achieving … |
Kaixing Yang; Xulong Tang; Ziqiao Peng; Xiangyue Zhang; Chubin Chen; xukun zhou; Puwei Wang; Hongyan Liu; Jun He; |
| 460 | OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To the best of our knowledge, CIPE-Dance is the largest dataset for dance video generation to date, comprising 300k high-quality clips (over 400 hours) and covering diverse dancers, environments, and dance genres. To overcome the method limitation, we propose OmniDance, a frameworklevel recipe for integrating music into a TI2V foundation model without sacrificing its original controllability or visual fidelity. |
Kaixing Yang; Jiashu Zhu; Xulong Tang; Ziqiao Peng; Xiangyue Zhang; Chubin Chen; Puwei Wang; Jiahong Wu; Xiangxiang Chu; Hongyan Liu; Jun He; |
| 461 | Rethinking Reward Signals in Video GRPO: When Scores Become Targets Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This leads to two recurring issues: (i) shortcut-driven optimization under composite objectives and (ii) reward saturation within prompt groups. To address these issues, we introduce TaRoS, a Target-Robust Reward Signaling framework for Video generation GRPO. |
Rui Li; Yuanzhi Liang; Ziqi Ni; Haibin Huang; Chi Zhang; Xuelong Li; |
| 462 | DreamWorld: Geometry-Grounded Video Diffusion for 3D-Consistent World Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To alleviate this, we presentDreamWorld, a new recipe of world model that novelly bridges the strongspatial structure priors of 3D foundation models with the high-fidelitygenerative capabilities of video diffusion models for geometry-consistent3D scene generation. Specifically, given the input image and camera tra-jectory, DreamWorld first learns a geometry video diffusion model topredict compact geometry features for the target novel views, function-ing as explicit structure pivots to reflect the underlying 3D spatial layout.To achieve this, we introduce a distillation paradigm that transfers high-level structural knowledge from a pretrained 3D foundation model to thediffusion model, thereby enabling it to produce geometrically consistentand spatially coherent features. |
Haibo Yang; Yang Chen; Yingwei Pan; Zhineng Chen; Ting Yao; Tao Mei; |
| 463 | Do Multimodal LLMs Understand Intraoral Dental Data? Dataset, Platform, and Baselines Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Experiments indicate thatstate-of-the-art multimodal models fail to generate clinically faithful reports, moti-vating geometry-aware adaptation. We therefore propose IOS-Qwen, which fuses aPointTransformer 3D encoder with Qwen3-VL to generate structured, point-cloud-conditioned reports. |
Luca Lumetti; Federico Rizzo; Francesca Cremonini; Ettore Candeloro; Lombardo Luca; Costantino Grana; Federico Bolelli; |
| 464 | Consistent Video-to-Video Translation Via Explicit Correspondences Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Interactive video-to-video applications require real-time gen-eration while maintaining long-range temporal consistency. However, re-cent methods achieve speed by restricting … |
Gaurav Parmar; Zhengqi Li; Richard Zhang; Jun-Yan Zhu; Srinivasa G. Narasimhan; Eli Shechtman; Yotam Nitzan; |
| 465 | MMR-Bench: A Comprehensive Benchmark for Multimodal LLM Routing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Multimodal large language models (MLLMs) have advancedrapidly, yet heterogeneity in architecture, alignment strategies, and effi-ciency means that no single model is uniformly … |
Hao-Xuan Ma; Guannan Lai; Han-Jia Ye; |
| 466 | OTCache: Optimal Transport for Geometry-Aware Caching in Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose OTCache, a training-free framework for ac-celerating diffusion sampling via caching schedule prediction. |
Huanlin Gao; Fang Zhao; Qiang Hui; Fuyuan Shi; Shaoan Zhao; Yantao Li; Chao Tan; Ting Lu; Yuren You; Kai Wang; Shiguo Lian; |
| 467 | Improved Immiscible Diffusion: Accelerating Diffusion Training By Reducing Miscibility Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, concerns regardinglimited image diversity and exploding execution times under large batchsizes limit its feasibility for large-scale training. In this work, we startfrom thoroughly tackling these limitations: For the image diversity, wedemonstrate the bijective nature of the denoising process of vanilla dif-fusion, underlying that being immiscible cannot hurt the diversity. |
Yiheng Li; Feng Liang; Dan Kondratyuk; Masayoshi TOMIZUKA; Kurt Keutzer; Chenfeng Xu; |
| 468 | Unsafe By Reciprocity: How Generation–Understanding Coupling Undermines Safety in Unified Multimodal Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we investigate whether cross-functionality reciprocity itself constitutes a structural source of vulnerability in UMMs. |
Kaishen Wang; Heng Huang; |
| 469 | Egocentric World Model for Photorealistic Hand Object Interaction Synthesis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Consequently, existing approaches often circumvent these physicschallenges by resorting to conditional video generation with access toknown future object trajectories. We introduce EgoHOI, an egocentricHOI world model that breaks away from this shortcut to simulate pho-torealistic, contact-consistent interactions from action signals alone. |
Dayou Li; Lulin Liu; Bangya Liu; Shijie Zhou; Jiu Feng; Ziqi Lu; Minghui Zheng; Chenyu You; Zhiwen Fan; |
| 470 | Task Alignment: A Simple and Effective Proxy for Model Merging in Computer Vision Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite extensive prior work, most evaluations of modelmerging in computer vision are restricted to image classification usingCLIP, where different classification datasets define different tasks. In thiswork, our goal is to make model merging more practical and show itsrelevance on challenging scenarios beyond this specific setting. |
Pau de Jorge Aranda; César DE SOUZA; Björn Michele; Mert Bulent SARIYILDIZ; Philippe Weinzaepfel; Florent Perronnin; Diane Larlus; Yannis Kalantidis; |
| 471 | Error-Driven Scene Editing for 3D Grounding in Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This limitation stems in part from training data that fo-cuses on language reasoning rather than spatial understanding due toscarce 3D resources, leaving inherent grounding biases unresolved. Toaddress this, we propose 3D scene editing as a key mechanism to gener-ate visual counterfactuals that mitigate these biases through fine-grainedspatial manipulation, without requiring costly scene reconstruction orlarge-scale 3D data collection. |
Yue Zhang; Zun Wang; Han Lin; Jialu Li; Jianing Yang; Yonatan Bitton; Idan Szpektor; Mohit Bansal; |
| 472 | Optimization-Guided Diffusion for Interactive Scene Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Realistic and diverse multi-agent driving scenes are crucial for evalu-ating autonomous vehicles, but safety-critical events which are essential for thistask are rare and … |
Shihao Li; Naisheng Ye; Tianyu Li; Kashyap Chitta; Tuo An; Peng Su; Boyang Wang; Haiou Liu; Chen Lv; Hongyang Li; |
| 473 | TSEmbed: Unlocking Task Scaling in Universal Multimodal Embeddings Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To addressthis, we propose TSEmbed, a universal multimodal embedding frameworkthat synergizes Mixture-of-Experts (MoE) with Low-Rank Adaptation(LoRA) to explicitly disentangle conflicting task objectives. |
Yebo Wu; Feng Liu; Ziwei Xie; Changwang Zhang; Jun Wang; Li Li; |
| 474 | Denoising The Deep Sky: Physics-Based CCD Noise Formation for Astronomical Imaging Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a physics-based noisesynthesis framework tailored to CCD noise formation in the telescope.The pipeline models photon shot noise, photo-response non-uniformity,dark-current noise, readout effects, and localized outliers arising fromcosmic-ray hits and hot pixels. |
Shuhong Liu; Xining Ge; Ziying Gu; Quanfeng Xu; Ziteng Cui; Lin Gu; Xuangeng Chu; Jun Liu; Dong Li; Tatsuya Harada; |
| 475 | CabinSI: Omni-Cabin Spatial Reasoning Through Explicit Visual Cognitive Maps Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To fill this research gap, we introduce CabinSI, a benchmark for evaluating spatial intelligence in cross-cabin environments, consisting of two components: RelCabin for relational reasoning tasks and RefCabin for spatial referring localization tasks. |
Mengxue Qu; Hengrui Hu; Mingming Ma; Ming Lei; Jie Gao; Henghui Ding; Yao Zhao; Kenn wu; Yunchao Wei; |
| 476 | AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce AnchorWeave, a memory-augmented video generation framework that replaces a single misaligned global memory with multiple clean local geometric memories and learns to reconcile their cross-view inconsistencies. |
Zun Wang; Han Lin; Jaehong Yoon; Jaemin Cho; Yue Zhang; Mohit Bansal; |
| 477 | Decoupling Complexity from Scale in Latent Diffusion Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, the latent capacity re-quired to represent visual data primarily depends on content complex-ity, with scale serving only as an upper bound. Motivated by this obser-vation, we propose DCS-LDM, a novel paradigm for visual generationthat decouples information complexity from scale. |
Tianxiong Zhong; Xingye Tian; Xuebo Wang; Boyuan Jiang; Xin Tao; Pengfei Wan; |
| 478 | Forecasting Animal Motion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose dense point trajectories as visual tokens for behavior,a structured mid-level representation that disentangles motion fromappearance and generalizes across diverse non-rigid agents, such as animalsin-the-wild. |
Neerja Thakkar; Shiry Ginosar; Jacob Walker; Jitendra Malik; Joao Carreira; Carl Doersch; |
| 479 | In-Context Sync-LoRA for Portrait Video Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Sync-LoRA, a method for editing portrait videos that achieves high-quality visual modifications while maintaining frame-accurate synchronization and identity consistency. |
Sagi Polaczek; Or Patashnik; Ali Mahdavi-Amiri; Danny Cohen-Or; |
| 480 | Robust and Efficient Monocular 3D Gaussian SLAM for Kilometer-Scale Outdoor Scenes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose KiloGS-SLAM, a highly efficient and robust monocular 3DGS-SLAM system that jointly addresses both bottlenecks. |
Sicheng Yu; Dongxu Shen; Beizhen ZHAO; Ding Guanzhi; Hao Wang; |
| 481 | VisionCoach: Reinforcing Grounded Video Reasoning Via Visual-Perception Prompting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Video reasoning requires models to locate and track question-relevant evidence across frames. While reinforcement learning (RL) withverifiable rewards improves accuracy, it still … |
Daeun Lee; Shoubin Yu; Yue Zhang; Mohit Bansal; |
| 482 | Taming Text-to-Sounding Video Generation Via Advanced Modality Condition and Interaction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: For the sec-ond, we introduce BridgeDiT, a dual-tower di!usion transformer thatemploys Dual Cross-Attention (DCA) as a bidirectional bridge betweenvideo and audio streams, which we show through systematic compari-son to be the optimal fusion strategy for the dual-tower paradigm. |
Kaisi Guan; Xihua Wang; Zhengfeng Lai; Xin Cheng; Peng Zhang; Xiaojiang Liu; Ruihua Song; Meng Cao; |
| 483 | LogFA: Efficient Feature-Space Data Augmentation for Egocentric Temporal Action Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose LogFA (Local-global Featurespace Augmentation), a novel framework that performs data augmentation directly in the feature space for Temporal Action Segmentation (TAS). |
Zijia Lu; Ehsan Elhamifar; |
| 484 | PhysEdit: Physically Consistent Image Editing Via Causal Enforcement Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While instruction-guided image editing has seen significantstrides, most existing models fail to preserve physical laws and causal-ity. This limitation stems from the sparsity of physical supervision sig-nals, coupled with the static, non-causal architecture of prevailing frame-works, which together hinder a deep comprehension of the physical world.To bridge this gap, we present PhysEdit, a novel framework that en-forces physical consistency in image editing through causal generationand physics-aware reinforcement learning. |
Siqi Wan; Jingwen Chen; Yehao Li; Yingwei Pan; Ting Yao; Tao Mei; |
| 485 | Measuring 3D Spatial Geometric Consistency in Dynamic Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing evaluation methods fail to accurately characterize these inconsistencies: fidelity-centric metrics like FVD are insensitive to geometric distortions, while consistency-focused benchmarks often penalize valid foreground dynamics. To address this gap, we introduce SGC, a metric for evaluating 3D Spatial Geometric Consistency in dynamically generated videos. |
Weijia Dou; Wenzhao Zheng; Weiliang Chen; Yu Zheng; Jie Zhou; Jiwen Lu; |
| 486 | Evidence-Backed Video Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Evidence-Backed Video Question Answering (E-VQA), a novel task requiring models to jointly output a semantic answer and precise spatio-temporal evidence: temporal segments and dense, tracked object segmentation masklets. |
Shijie Wang; Honglu Zhou; Ziyang Wang; Ran Xu; Caiming Xiong; Silvio Savarese; Chen Sun; Juan Carlos Niebles; |
| 487 | HighlightBench: Benchmarking and Diagnosing Markup-Driven Table Reasoning in Scientific Documents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This creates a key blind spot in assessing markup-conditioned behaviorover tables. To address this gap, we introduce HighlightBench, a diagnos-tic benchmark for markup-driven table understanding that decomposesevaluation into five task families: Markup Grounding, Constrained Re-trieval, Local Relations, Aggregation & Comparison, and Consistency &Missingness. |
Lexin Wang; Shenghua Liu; Yiwei Wang; Yujun Cai; Yuyao Ge; Jiayu Yao; Jiafeng Guo; Xueqi Cheng; |
| 488 | S2-FracMix: Self-Saliency Fractal Mixup Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Collectively, our unified framework,S 2 -FracMix, enables simultaneous learning from fractal and non-fractalstructures within a single image, yielding a targeted and structurally co-herent augmentation strategy. We theoretically analyze the advantageof our technique, and empirically establish its superiority over the ex-isting methods by achieving state-of-the-art performance in extensiveevaluation with seven benchmarks across classification (coarse and fine-grained), robustness, calibration, object detection, and transfer learningtasks. |
Khawar Islam; Arif Mahmood; Xin Jin; NAVEED AKHTAR; |
| 489 | BeTTER: Diagnose The Illusion of Embodied Reasoning in Vision-Language-Action Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent Vision-Language-Action (VLA) models report im-pressive success rates on robotic benchmarks, but whether these scoresreflect genuine embodied reasoning remains questionable. To addressthis gap, we introduce BeTTER, a diagnostic Benchmark for TestingTrue Embodied Reasoning. |
Haiweng Xu; Sipeng Zheng; Hao Luo; Wanpeng Zhang; Zongqing Lu; |
| 490 | XSemanticFlow: Cross Object Semantic Alignment for Zero-shot Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing methods are often inconsistent acrossgeometries or vulnerable to non-canonicalized poses, making correspon-dence transfer unreliable. To address this, we propose XSemanticFlow,which learns correspondences from pure semantic features. |
Junyu Nan; Noam Eshed; Brian Okorn; Kris Kitani; |
| 491 | WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This capability is crucial for gaming and embodied AI applications. Motivated by this, we present WorldWander, an in-context learning framework tailored for translating between egocentric and exocentric worlds in video generation. |
Quanjian Song; Yiren Song; Kelly Peng; Yuan Gao; Mike Zheng Shou; |
| 492 | RoboGesture: Real-Time Semantic-aligned Co-Speech Gestures Generation for Humanoid Interaction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose RoboGesture, a robot-centric framework that co-designs data, modeling, and control to powera complete interactive human–humanoid system in which the robot lis-tens, responds, and gestures in real time. |
Zifan Wang; Ziang Ren; Pengyang Shi; Zirui Wang; Chenghuai Lin; Tianze Wang; Zekun Qi; Liangliang Zhao; HE WANG; Li Yi; |
| 493 | MSEditor: Toward Consistent Multi-Shot Video Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we tackle the problem of performing consistent,unified modifications to a multi-shot video sequence. |
Kunyu Feng; Yue Ma; Bingyuan Wang; Yuefeng Wang; Zhiyuan Qin; Hao Cheng; Hao Li; Qifeng Chen; Zeyu Wang; |
| 494 | ECHO: Ego-centric Modeling of Human-Object Interactions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present ECHO, the first unified framework to jointly recoverhuman pose, object motion, and contact dynamics solely from head andwrist tracking. |
Ilya A. Petrov; Vladimir Guzov; Riccardo Marin; Emre Aksan; Xu Chen; Daniel Cremers; Thabo Beeler; Gerard Pons-Moll; |
| 495 | ActionPlan: Future-Aware Streaming Motion Synthesis Via Frame-Level Action Planning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present ActionPlan, a unified motion diffusion frame-work that bridges real-time streaming with high-quality offline genera-tion within a single model. |
Eric Nazarenus; Chuqiao Li; Yannan He; Xianghui Xie; Jan Eric Lenssen; Gerard Pons-Moll; |
| 496 | OrthoTrack: Continuous 6-DoF UAV Trajectory Estimation Anchored in Public Orthophotos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present OrthoTrack, atraining-free system that estimates continuous 6-DoF UAV trajectoriesusing only publicly available orthophotos and surface models as a mapprior. |
Oussema Dhaouadi; Zuria Bauer; Johannes Meier; Olaf Wysocki; Marc Pollefeys; Daniel Cremers; |
| 497 | SpecEyes: Accelerating Agentic Multimodal LLM Via Speculative Planning and Perception Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This overhead, termed agentic depth,incurs prohibitive latency and seriously limits system-level concurrency.To this end, we propose SpecEyes, an agentic-level speculative accelera-tion framework that breaks this sequential bottleneck. |
Haoyu Huang; Jinfa Huang; Zhongwei Wan; Xiawu Zheng; Rongrong Ji; Jiebo Luo; |
| 498 | OmniForcing: Unleashing Real-time Joint Audio-Visual Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent joint audio-visual diffusion models achieve remark-able generation quality but suffer from high latency due to their bidirec-tional attention dependencies, hindering real-time applications. |
Yaofeng Su; Yuming Li; Zeyue Xue; Jie Huang; Siming Fu; Haoran Li; Haoyang Huang; Nan Duan; |
| 499 | TraversRL: Traversable Pedestrian Pathway Generation With Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Automatically generating pedestrian pathways from aerialimages requires producing a connected network suitable for routing, notjust detecting where sidewalks appear. Sidewalks and … |
Bin Han; Robert Wolfe; Bill Howe; |
| 500 | ReSplat: Learning Recurrent Gaussian Splatting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose ReSplat, a feed-forward recurrentGaussian splatting model that iteratively refines 3D Gaussians withoutexplicitly computing gradients. |
Haofei Xu; Daniel Barath; Andreas Geiger; Marc Pollefeys; |
This table only includes 500 papers selected by our daily digest algorithm. To continue with the full list (~2,830 papers), please visit Paper Digest: ECCV-2026 (Full List).