ECCV 2026 Papers with Code & Data
To facilitate rapid community engagement with the presented research, we have compiled an extensive index of accepted papers that have associated public code or data repositories. We list all of them in the following table. This index was generated using an automated extraction process. While we strive for completeness, some papers with public resources may have been missed. Please inform us if you discover any additional papers that should be included. Readers should be aware that some code repositories may not be made fully public until the conference officially begins.
In addition to this index, we encourage readers to explore our related resources: ECCV-2026 papers & highlights: For curated summaries and key takeaways from this year’s conference. “Best Paper” Digest (ECCV): A historical overview of the most influential ECCV papers published in recent years.
Since 2018, Paper Digest has built a foundation of data spanning decades of conferences, journals, and research topics. The platform features a daily digest service that sifts through tens of thousands of new papers, clinical trials, news articles, and community posts, filtering the noise to highlight what matters most to specific interests. Beyond daily updates, dozens of built-in research tools streamline the academic workflow, supporting efficient reading and writing, comprehensive literature reviews, and automated research report generation.
Paper Digest Team
New York City, New York, 10017
team@paperdigest.org
TABLE 1: ECCV 2026 Papers with Code & Data
| Paper | Author(s) | Code | |
|---|---|---|---|
| 1 | SpecV: Specification Verification for Robust Unified Multimodal Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present SpecV, a specification-verification framework for robust unified multimodal evaluation. |
Weihao Yu; Rongyao Fang; Yuxuan Cai; Linjiang Huang; Yuhuan Yang; Xianwei Zhuang; Junyang Lin; Yixuan Yuan; Shuai Bai; | code |
| 2 | Natural Image Pretraining Improves Abstract Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our anal-ysis identifies two key causes: (1) Multiview redundancy — Fromthe data perspective, certain views provide limited or noisy informa-tion, diluting discriminative cues; (2) Overfitting — From the featureperspective, conventional fusion increases representational complexity,causing the model to memorise view-specific patterns rather than learngeneralisable representations. To address these issues, we propose twocomplementary modules. |
Xiaoman Ding; Keya Hu; Katelyn Gan; Victor Yin; Kaiming He; | code |
| 3 | GO-Renderer: Generative Object Rendering with 3D-aware Controllable Video Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose GO-Renderer, a unified framework integrating the reconstructed 3D proxies to guide the video generative models to achieve high-quality object rendering on arbitrary viewpoints under arbitrary lighting conditions. |
Zekai Gu; Shuoxuan Feng; Yansong Wang; Hanzhuo Huang; Zhongshuo Du; Chengfeng Zhao; Chengwei Ren; Peng Wang; Yuan Liu; | code |
| 4 | Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To support SFT and RFT training, we introduce TraceSpatial, a largescale dataset of 30M QA pairs, spanning outdoor/indoor/tabletop scenes and supporting complex reasoning processes (up to 9 steps). |
Enshen Zhou; Yibo Li; Jingkun An; Jiayuan Zhang; Shanyu Rong; Mengzhen Liu; Yi Han; Yuheng Ji; Huajie Tan; Jiawei He; Pengwei Wang; Zhongyuan Wang; Cheng Chi; Lu Sheng; Shanghang Zhang; | code |
| 5 | OmniMamba: Efficient and Unified Multimodal Understanding and Generation Via State Space Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present OmniMamba, the first linear-architecture-based multimodal generation model that generates both text and imagesthrough a unified next-token prediction paradigm. |
Jialv Zou; Bencheng Liao; Qian Zhang; Wenyu Liu; Xinggang Wang; | code |
| 6 | EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current methods primarily rely on passive processing, which often face a trade-off between maintaining long-range context and capturing the fine-grained details necessary for complex tasks. To address this, we introduce EventMemAgent, an active online video agent framework based on a hierarchical memory module. |
Siwei Wen; Zhangcheng Wang; Xingjian Zhang; Lei Huang; wenjun wu; | code |
| 7 | Less Is More: Reducing Complexity in Vision-Language-Action Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In thiswork, we introduce StarVLA-α, a simple yet strong baseline designedto study VLA design choices under controlled conditions. |
Jinhui Ye; Ning Gao; Senqiao Yang; Jinliang Zheng; Zixuan Wang; Yuxin Chen; Pengguang Chen; Yilun Chen; Shu Liu; Jiaya Jia; | code |
| 8 | OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Extending these capabilities to remote sensingremains challenging, as models must reason over spatial scale, geographicstructures, and multispectral indices while maintaining coherent multi-step logic. To address this gap, we introduce OpenEarthAgent, a unifiedframework for tool-augmented geospatial reasoning trained on satelliteimagery, natural-language queries, and structured reasoning traces. |
Akashah Shabbir; Muhammad Umer Sheikh; Muhammad Akhtar Munir; Hiyam Debary; Mustansar Fiaz; Muhammad Zaigham Zaheer; Paolo Fraccaro; Fahad Shahbaz Khan; Muhammad Haris Khan; Xiao Xiang Zhu; Salman Khan; | code |
| 9 | Dynamic-Robust Photometric–Semantic Reconstruction for Open-Vocabulary 3D Scene Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, their inherent reliance on static-sceneassumptions leads to severe misalignment of spatial features in uncon-strained dynamic environments. To bridge this critical gap, we proposeSPAR, a novel joint semantic-geometric encoding architecture that ex-plicitly isolates transient dynamic noise prior to latent space aggregation.Furthermore, we introduce a dynamic-region-aware end-to-end trainingparadigm that structurally couples motion estimation with multi-viewvisual and semantic learning. |
Boyu Cai; Li Yang; Yan Xu; Wei Liu; Nian Liu; Sikui Zhang; Yan Wang; Chunfeng Yuan; Weiming Hu; | code |
| 10 | Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However,we observe that per-frame depth remains stable throughout this failure.The backbone’s local geometry remains intact; only the global pose headbreaks down. Motivated by this decoupling, we introduce Scal3R. |
Chin-Yang Lin; Yang-Che Sun; Cheng Sun; Fu-En Yang; Min-Hung Chen; Yen-Yu Lin; Wei-Chen Chiu; Yu-Lun Liu; | code |
| 11 | Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Real-world deployment of embodied agents requires activeexploration, visual grounding, and interactive intent disambiguation. How-ever, existing frameworks often rely on privileged simulator states or as-sume complete instructions, bypassing realistic deployment challenges.To bridge this gap, we present REAL, an agentic framework for open-world mobile manipulation. |
Boyu Mi; Mengchen Ma; Yifei Yao; Xing Gao; Hanqing Wang; Junting Chen; Yangzi Li; Zihou Zhu; Guohao Li; Zhenfei Yin; Tai Wang; Yao Mu; Jiangmiao Pang; | code |
| 12 | Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despiterequiring no auxiliary losses, reward shaping, or architectural changes— serving as a minimal, plug-in modification to standard post-trainingpipelines — models trained under perceptual starvation achieve substan-tial gains of 5% average relative improvement across diverse benchmarks.Our codes and data will be publicly available at https://github.com/WhuanY/Starve2Perceive. |
Yuhuan Wu; Haozhe Wang; Cong Wei; Chong Peng; Fangzhen Lin; Wenhu Chen; | code |
| 13 | Aligning Human Sense: Calibrated Distributional Reward Learning for Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: 3) In policy optimization, the widely adopted KL-divergenceimposes only local constraints, failing to capture holistic human pref-erence. To address these challenges, we propose a unified, preference-aware learning framework for video generation. |
Naixin Zhai; Weihua Cheng; Dexu Yu; Yikai Gu; Hanwen Du; Junchen Fu; Chenxi Huang; Yingwei Song; Liyuan Ma; Yang Ran; Youhua Li; Yongxin Ni; | code |
| 14 | Tac2Real: Reliable and GPU Visuotactile Simulation for Online Reinforcement Learning and Zero-shot Real-World Deployment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Tac2Real integrates the Precondi-tioned Nonlinear Conjugate Gradient Incremental Potential Contact (PNCG-IPC) method with a multi-node, multi-GPU high-throughput par-allel simulation architecture, which can generate marker displacementfields at interactive rates. |
Ningyu YAN; Shuai Wang; Xing Shen; Hui Wang; Hanqing Wang; Yang Xiang; Jiangmiao Pang; | code |
| 15 | Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Large Multimodal Models (LMMs) have achieved remark-able success on images and short videos, yet scaling them to long videosremains challenging due to frame-centric tokenization and limited con-text windows. |
Lucy Lin; Ayush Jain; Yifan Liu; Katerina Fragkiadaki; | code |
| 16 | The Prism Hypothesis: Harmonizing Semantic and Pixel Representations Via Unified Autoencoding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we systematically analyze the spectral character-istics of various semantic and pixel encoders. |
Weichen Fan; Haiwen Diao; Quan Wang; Dahua Lin; Ziwei Liu; | code |
| 17 | RoboClaw: An Agentic Framework for Scalable Long-Horizon Robotic Tasks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present RoboClaw, an agentic robotics framework that unifies data collection, policy learning, and task execution under a single VLM-driven controller. |
Ruiying Li; Yunlang Zhou; Yuyao Zhu; Kylin Chen; Sukai Wang; Kongtao Hu; Minhui Yu; Bowen Jiang; Jiayao Ma; Zhan Su; Yongjian Shen; Yang Yang; Guanghui Ren; Maoqing Yao; Wenhao Wang; Yao Mu; | code |
| 18 | Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Re2Pix, a hierarchical video prediction framework that decomposes forecasting into two stages: semantic representation prediction and representation-guided visual synthesis. |
Efstathios Karypidis; Spyros Gidaris; Nikos Komodakis; | code |
| 19 | UniPR-3D: Towards Universal Visual Place Recognition with Visual Geometry Grounded Transformer Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In thiswork, we introduce UniPR-3D, the first VPR architecture that effectivelyintegrates geometry-aware information from multiple views. |
Tianchen Deng; Chen Xun; Ziming Li; Hongming Shen; Shuhao Zhai; Danwei Wang; Javier Civera; Hesheng Wang; | code |
| 20 | R3DP: Real-Time 3D-Aware Policy for Embodied Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Real-time 3D-aware Policy (R3DP), which integrates powerful 3D priors into manipulation policies without sacrificing real-time performance. |
Yuhao Zhang; Wanxi Dong; Yue Shi; Yi Liang; Jingnan Gao; Qiaochu Yang; Yaxing Lyu; Zhixuan Liang; Yibin Liu; Congsheng Xu; Xianda Guo; Wei Sui; Yaohui Jin; Xiaokang Yang; Yanyan Xu; Yao Mu; | code |
| 21 | DiffusionVL: Translating Any Autoregressive Models Into Diffusion Vision Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose DiffusionVL, a family of dVLMs obtained by translating pretrained AR models into the diffusion paradigm via an efficient diffusion finetuning procedure that changes the training objective and decoding process while keeping the backbone architecture intact. |
Lunbin Zeng; Jingfeng Yao; Bencheng Liao; Hongyuan Tao; Wenyu Liu; Xinggang Wang; | code |
| 22 | CloSeR: Unified Relational Distillation from Closed-Set Teachers for Category Discovery Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose CloSeR, a simple plug-and-play framework that in-jects Closed-Set Relational knowledge into GCD training. |
Yuanpei Liu; Zhenqi He; Jialu Tang; Kai Han; | code |
| 23 | Molmo-Point: Better Pointing for VLMs with Grounding Tokens Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Using this method, we set a new state-of-the-art on image pointing(70.7% on PointBench), set a new state-of-the-art for fully open modelson GUI pointing (61.1% on ScreenSpotPro), substantially improve VLMvideo tracking (62.5 on J &F vs 56.7 for Molmo2 on Molmo2Track), andimprove video pointing (59.1% human preference win rate vs. Molmo2). |
Christopher Clark; Yue Yang; Jae Sung Park; Zixian Ma; Jieyu Zhang; Rohun Tripathi; Mohammadreza Salehi; Sangho Lee; Ranjay Krishna; | code |
| 24 | Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Live Avatar, an algorithm-system co-designed frame-work that addresses both challenges for a 14-billion-parameter di_x001B_usionmodel. |
Yubo Huang; Hailong Guo; Fangtai Wu; Weiqiang Wang; Shijie Huang; Qijun Gan; Shifeng Zhang; Lin Liu; Sirui Zhao; Enhong Chen; Jiaming Liu; Steven Hoi; | code |
| 25 | Recurrent Autoregressive Diffusion: Global Memory Meets Local Attention Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Moreover, existingdiffusion-RNN approaches often suffer from performance degradation dueto training-inference gap or the lack of overlap across windows. To addressthese limitations, we propose a novel Recurrent Autoregressive Diffusion(RAD) framework, which leverages recurrent blocks for memory update andretrieval and preserves local details by full attention on overlapping slidingwindows, with no training and inference gap. |
Taiye Chen; Zihan Ding; Anjian Li; Christina Zhang; Zeqi Xiao; Yisen Wang; Chi Jin; | code |
| 26 | Skyfall-GS: Synthesizing Immersive 3D Urban Scenes from Satellite Imagery Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we take an alternative route to create largescale 3D scenes by leveraging readily available satellite imagery for realistic coarse geometry and open-domain diffusion models for high-quality close-up appearance synthesis. |
Jie-Ying Lee; Yi-Ruei Liu; Shr-Ruei Tsai; Wei-Cheng Chang; Chung-Ho Wu; Jiewen Chan; Zhenjun Zhao; Chieh Hubert Lin; Yu-Lun Liu; | code |
| 27 | M2Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This “discretization bottleneck” significantly limits the performance ceiling of downstream VisionLanguage-Action (VLA) models. To address this, we propose M2Tok, a Multi-head Multi-codebook Action Tokenizer designed to minimize reconstruction error and enhance policy performance. |
Chunpu Xu; Zhixuan Liang; Yuhao Zhang; Chi-Min Chan; Jiashuo Wang; Yang Xiao; Mengkang Hu; Xiaokang Yang; Yao Mu; | code |
| 28 | Reward Lightning: Fast Video Generation Via Homologous Preference Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods optimize the two objectives over mismatched representation spaces, where improving one objective often compromises the other. To overcome this, we propose Reward Lightning, a unified framework that aligns and accelerates a video diffusion model within a single shared representation. |
Jiaxiang Cheng; bing ma; Xuhua Ren; Kai Yu; Peng Zhang; Tianxiang Zheng; Qinglin Lu; | code |
| 29 | MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present MA-VLA, a unified framework for multi-arm collaboration via atomic action assignment.We also construct a benchmark in which test-time collaboration patterns are absent in training set. |
Zaibin Zhang; Junlan Xiao; Zhongbo Zhang; Yifan Wang; Li Kang; Yiran Qin; Changxing Xia; Heng Zhou; Talas Fu; Enshen Zhou; Ruimao Zhang; Zhenfei Yin; Huchuan Lu; Lijun Wang; | code |
| 30 | ReSWD: ReSTIR‘d, Not Shaken. Combining Reservoir Sampling and Sliced Wasserstein Distance for Variance Reduction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: ReSWD: ReSTIR‘d, not shaken. Combining Reservoir Sampling and Sliced Wasserstein Distance for Variance Reduction. |
Mark Boss; Andreas Engelhardt; Simon Donné; Varun Jampani; | code |
| 31 | CascadeProto: Cascaded Cross-Modal Prototype Purification Via Entropy-Aware Learning for Few-Shot 3D Point Cloud Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To fur-ther enrich prototype representations, we introduce Learnable ModalityAdapters (LMA) that independently align each of three CLIP modal-ities — text, audio, and image — with point cloud features throughforeground-background decoupled distribution matching, enabling flexi-ble single-modality semantic enrichment that bridges the 2D-3D domaingap. |
Changshuo Wang; Weijun Li; Fan Mo; Zhonghang Liu; Shuting He; Prayag Tiwari; Dimitrios Kanoulas; | code |
| 32 | Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this,we present Deform360, a large-scale visuotactile dataset featuring 198daily-life objects, 1,980 interaction sequences, and over 215 hours of ob-servations from 41 surround-view cameras and bimanual tactile grippersto capture both global motion and contact-induced local deformations.Leveraging a novel markerless visuotactile 3D tracking pipeline to ex-tract dense geometry and motion, we systematically evaluate currentstate-of-the-art world models, comparing 2D video models against 3Dparticle models. |
Hongyu Li; Wanjia Fu; Xiaoyan Cong; Zekun Li; Binghao Huang; Hanxiao Jiang; Xintong He; Yiqing Liang; Rao Fu; Tao Lu; Srinath Sridhar; Kevin Smith; George Konidaris; Yunzhu Li; | code |
| 33 | World Models for Learning Dexterous Hand-Object Interactions from Human Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Modeling dexterous hand–object interactions is challengingas it requires understanding how subtle finger motions influence theenvironment through contact with objects. While … |
Raktim Goswami; Amir Bar; David Fan; Tsung-Yen Yang; Gaoyue Zhou; Prashanth Krishnamurthy; Michael Rabbat; Farshad Khorrami; Yann LeCun; | code |
| 34 | The Sterkfontein Caves Dataset: A Novel View Rendering Challenge from The Cradle of Humankind Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce the challenging Sterkfontein Caves dataset comprising ten underground scenes from a UNESCO World Heritage Site, and use it to find a new simple baseline method that beats existing low-light reconstruction methods upon it. |
Ireton Liu; Brian Xu; Dominic Stratford; Steven James; Richard Klein; James Tompkin; | code |
| 35 | Reward Modeling for Computer-Using Agent from Video Execution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we study rewardmodeling from execution video: a sequence of keyframes from an agenttrajectory that is independent of the agent’s internal reasoning or ac-tions. |
Linxin Song; Jieyu Zhang; Huanxin Sheng; Taiwei Shi; Rahul Gupta; Yang Liu; Ranjay Krishna; Jian Kang; Jieyu Zhao; | code |
| 36 | ConceptWeaver: Weaving Disentangled Concepts with Flow Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A final concept-insensitive Refinement Stage thensynthesizes fine-grained details. Guided by this discovery, we proposeConceptWeaver, a framework for one-shot concept disentanglement.ConceptWeaver learns concept-specific semantic offsets from a single ref-erence image using a stage-aware optimization strategy that aligns withthe three-stage framework. |
Jintao Chen; Aiming Hao; Xiaoqing Chen; Chengyu Bai; Chubin Chen; Yanxun Li; Jiahong Wu; Xiangxiang Chu; Shanghang Zhang; | code |
| 37 | OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introducesthe novel video-to-script (V2S) task, aiming to generate hierarchical,scene-by-scene scripts encompassing character actions, dialogues, expres-sions, and audio cues. To facilitate this, we construct a first-of-its-kindhuman-annotated benchmark and propose a temporally-aware hierarchi-cal evaluation framework. |
JUNFU PU; Yuxin Chen; Teng Wang; Ying Shan; | code |
| 38 | From Macro to Micro: Benchmarking Microscopic Spatial Intelligence on Molecules Via Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introduces the concept of Microscopic Spatial In-telligence (MiSI), the capability to perceive and reason about the spatialrelationships of invisible microscopic entities, which is fundamental toscientific discovery. To assess the potential of Vision-Language Models(VLMs) in this domain, we propose a systematic benchmark frameworkMiSI-Bench. |
Zongzhao Li; Xiangzhe Kong; Jiahui Su; Zongyang Ma; Mingze Li; Songyou Li; Yuelin Zhang; Yu Rong; Tingyang Xu; Deli Zhao; Wenbing Huang; | code |
| 39 | Concept-as-Tree: A Controllable Synthetic Data Framework Makes Stronger Personalized VLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recently, there hasbeen an increasing interest in improving the personalization capabilitiesof VLMs. To better integrate user-provided concepts into VLMs, manymethods use positive and negative samples to fine-tune these models.However, the scarcity of user-provided positive samples and the low qual-ity of retrieved negative samples pose challenges for existing techniques.To reveal the relationship between sample and model performance, wesystematically investigate the amount and diversity impact of positiveand negative samples (easy and hard) on VLM personalization tasks.Based on the detailed analysis, we introduce Concept-as-Tree (CaT),which represents a concept as a tree structure, thereby enabling the datageneration of positive and negative samples with varying difficulty anddiversity, and can be easily extended to multi-concept scenarios. |
Ruichuan An; Kai Zeng; Ming Lu; Sihan Yang; Renrui Zhang; Huitong Ji; Hao Liang; Wentao Zhang; | code |
| 40 | NUN: Nested Unfolding Network for Real-World Concealed Object Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A bi-directional unfolding interaction mechanismuses IQA to select optimal DeRUN outputs, and a cross-stage consis-tency loss ensures robust predictions under varying restoration quality.Theoretically, under local assumptions, we show that NUN reduces di-rect parameter-level gradient conflict through disjoint parameter setsand achieves a degradation-sensitivity bound, where degradation affectssegmentation only through the inner-loop optimization and restorationapproximation errors. |
Chunming He; Rihan Zhang; Longxiang Tang; Dingming Zhang; Bojian Zhang; Fengyang Xiao; Jingjia Feng; Sina Farsiu; | code |
| 41 | When Specialists Meet Generalists: Segmenter-Coordinated Asymmetric Learning for Label-Deficient Concealed Object Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, we observe that standard homogeneous co-training cannot exploit this complementarity because architecturally identical networks tend to share similar error patterns, especially when targets are heavily concealed. To address this, we present SCALER (Segmenter-Coordinated Asymmetric LEaRning), a framework that jointly optimizes a mean-teacher segmenter and a learnable SAM through two alternating phases with model-specific optimization strategies. |
Chunming He; Dingming Zhang; Longxiang Tang; Ziyun Yang; Fengyang Xiao; Sina Farsiu; | code |
| 42 | Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While 10B-level industrial foundation models have pushedthe boundaries of image inpainting, their prohibitive computational costsseverely hinder practical deployment. Constructing a highly optimizedtask-specific specialist offers a promising solution; however, extreme struc-tural compression inevitably triggers a severe representation bottleneck.To conquer this, we propose Moebius, a highly efficient lightweight in-painting framework. |
Kangsheng Duan; Ziyang Xu; Wenyu Liu; Xiaohu Ruan; Xiaoxin Chen; Xinggang Wang; | code |
| 43 | Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models As Embodied Brain Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: For planning,RoboBench introduces an evaluation framework that uses an MLLM as aworld simulator. |
Yulin Luo; Chun-Kai Fan; Menghang Dong; Jiayu Shi; Xiangju Mi; Mengdi Zhao; Bo-Wen Zhang; Cheng Chi; Jiaming Liu; Gaole Dai; Rongyu Zhang; Ruichuan An; Kun Wu; Zhengping Che; shaoxuan Xie; Guocai Yao; Zhongxia Zhao; Pengwei Wang; Guang Liu; Zhongyuan Wang; Tiejun Huang; Shanghang Zhang; | code |
| 44 | VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing instruction datasetsoften contain a substantial portion of visually redundant samples (solv-able from text alone), as well as multimodally misaligned supervisionthat can degrade learning. To address this, we propose VisNec (VisualNecessity Score), a principled data selection framework that measuresthe marginal contribution of visual input during instruction tuning. |
Mingkang Dong; Hongyi Cai; jie li; Sifan Zhou; Bin Ren; Kunyu Peng; Yuqian Fu; | code |
| 45 | Learning Accurate Segmentation Purely from Self-Supervision Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In thiswork, we introduce Selfment, a fully self-supervised framework that seg-ments foreground objects directly from raw images without human labels,pretrained segmentation models, or any post-processing. |
Zuyao You; Zuxuan Wu; Yu-Gang Jiang; | code |
| 46 | Reconstructing Humans and Objects in Interaction Using Large Reconstruction Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, weexplore a different avenue. |
Agniv Chatterjee; Georgios Pavlakos; | code |
| 47 | Free‑CD: Probabilistically Decoupled Training-Free Open-Vocabulary Change Detection with Resolution-Invariant Feature Inversion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Open-Vocabulary Change Detection (OVCD) faces a funda-mental granularity gap: foundation models prioritize high-level semanticabstraction while change detection requires pixel-level spatial fidelity.Traditional OVCD methods rely on instance extraction models, intro-ducing spatial semantic ambiguities and leading to over-segmentationor under-segmentation in remote sensing scenarios. To bridge this gap,we propose Free-CD, a training-free OVCD framework that reformu-lates the task by predicting a change probability distribution rather thanenforcing binary change masks via instance boundaries. |
Yongshuo Zhu; LU LI; Keyan Chen; Zhenwei Shi; ZHOU Fugen; | code |
| 48 | OmniX: From Unified Panoramic Generation and Perception To Graphics-Ready 3D Scenes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we advance this technique to generategraphics-ready 3D scenes suitable for physically based rendering (PBR),relighting, and simulation.Furthermore, we construct a large-scale syn-thetic panorama dataset comprising high-quality multimodal panoramasfrom diverse indoor and outdoor scenes. |
Yukun Huang; Jiwen Yu; Yanning Zhou; Jianan Wang; Xintao Wang; Pengfei Wan; Xihui Liu; | code |
| 49 | Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Online Video Large Language Models (VideoLLMs) playa critical role in supporting responsive, real-time interaction. Existingmethods focus on streaming perception, lacking a … |
Yiran Guan; Liang Yin; Dingkang Liang; Jianzhong Ju; Zhenbo Luo; Jian Luan; Yuliang Liu; Xiang Bai; | code |
| 50 | Roam2Room: A Unified Floorplan-to-Furnished Framework for Controllable Indoor Scene Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing meth-ods often rely on hand-crafted rules or focus on isolated sub-tasks (e.g.,floorplan synthesis or single-room furnishing), producing whole-homescenes that lack global coherence, realism, and simulation readiness. Tomitigate these limitations, we propose a unified hierarchical frameworkthat decomposes indoor scene synthesis into controllable stages. |
Wenbo Li; Zipeng Qin; Xiaoliang Ju; Rongyao Fang; Hongsheng LI; | code |
| 51 | RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in The Wild Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present RT-DocLayout, a highly efficient end-to-end framework for document layout analysis, designed as a front-end for document parsing tasks. |
Cheng Cui; Tingquan Gao; Xueqing Wang; Changda Zhou; Hongen Liu; Ting Sun; Yubo Zhang; Zelun Zhang; Jiaxuan Liu; Manhui Lin; Yue Zhang; Suyin Liang; Yiqing Xiang; Yi Liu; | code |
| 52 | SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study answer-span prediction entropy as a model-internalfeedback signal and show that naive entropy minimization is ambiguous,since low entropy may arise from evidence-grounded confidence or short-cut collapse. To resolve this ambiguity, we introduce low-entropy anchorsand an entropy-shaping objective that reduces answer uncertainty whilepreserving baseline high-confidence tokens. |
Bo Yin; Xiaobin Hu; Chengming Xu; Ruolin Shen; Mo Yang; Jiangning Zhang; Peng-Tao Jiang; Cheng Tan; Shuicheng Yan; | code |
| 53 | Sentinel: Embodied Cooperative Spatial Reasoning and Planning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we study Cooperative Spatial Intelligence, the ability of decentralized embodied agents to coordinate effectively under dynamic environmental constraints across city-scale outdoor domains.We introduce Sentinel Challenge, a benchmark where multiple decentralized embodied agents must communicate in natural language to agree on a mutually safe and convenient meeting point within large, city-scale outdoor environments. |
Xiangye Lin; Hongxin Zhang; Ruxi Deng; Qinhong Zhou; Chuang Gan; | code |
| 54 | Beyond Imitation: Learning Safe End-to-End Autonomous Driving from Hard Negatives Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address thislimitation, we propose BeyondDrive, a failure-aware imitation learn-ing framework that jointly learns from successful and failed driving be-haviors. |
Junli Wang; HuaZhihua HuaZhihua; Xueyi Liu; Zebin Xing; Wei Zhang; Kun Ma; Guang Chen; Hangjun Ye; Long Chen; Pengxuan Yang; | code |
| 55 | RADmesh: Remesh-Aware Mesh Deformation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a remeshing-enhanced method for generativelydeforming shapes with visual losses. |
Nam Anh Dinh; Itai Lang; Oded Stein; Rana Hanocka; | code |
| 56 | S-VAM: Shortcut Video-Action Model By Self-Distilling Geometric and Semantic Foresight Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, current VAMs, typically relying on either slow multi-step video generation or noisy one-step feature extraction, cannot simultaneously guarantee real-time inference and high-fidelity foresight. To address this limitation, we propose S-VAM, a shortcut video-action model that foresees coherent geometric and semantic representations via a single forward pass. |
Haodong Yan; Zhide Zhong; Jiaguan Zhu; Junjie He; Weilin Yuan; Wenxuan Song; Xin Gong; Yingjie CAI; Guanyi Zhao; Xu Yan; Liu Bingbing; Yingcong Chen; Haoang Li; | code |
| 57 | MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present MGM-Omni, an Omni MLLM for omni-modalunderstanding and expressive, long-horizon speech generation. |
Chengyao Wang; Zhisheng Zhong; Bohao PENG; Senqiao Yang; Yuqi Liu; Haokun GUI; Bin Xia; Jingyao Li; Bei Yu; Jiaya Jia; | code |
| 58 | SuperFlex: Deformable Superquadrics for Point Cloud Decomposition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work,we present SuperFlex, an enhanced framework that expands the expres-sive power and applicability of superquadric decompositions. |
Gabriel Tavernini; Elisabetta Fedele; Tiago Novello; Leonidas Guibas; Marc Pollefeys; Francis Engelmann; | code |
| 59 | ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, wepropose a unified understanding-and-generation framework that lever-ages world modeling to simultaneously enable meaningful explorationand provide dense supervision. |
Zihao Sheng; Xin Ye; Jingru Luo; Sikai Chen; Liu Ren; | code |
| 60 | Bridging Visual Representation and Reinforcement Learning from Verifiable Rewards in Large Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The method adaptively localizessemantically salient regions through hierarchical geometric aggregation,identifies vision-critical attention heads via structured attribution, andperforms paragraph-level credit reallocation to align spatial visual ev-idence with semantically decisive reasoning steps. |
yuhang han; Yuyang Wu; Zhengbo Jiao; Yiyu Wang; Xuyang Liu; Shaobo Wang; Hanlin xu; Xuming Hu; Linfeng Zhang; | code |
| 61 | Enhancing Prompt-image Alignment Evaluations Via Cyclic Mutual Information Maximization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Unlike exist-ing methods, our framework focuses on promoting e_x001B_ective multimodalinformation fusion in the deep feature domain, which has cyclic mutualinformation maximization phases. |
Xingran Liao; Duanyu Feng; Mingliang Zhou; Sam Kwong; Weisi Lin; | code |
| 62 | JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present JoVA, a streamlined frameworkthat unifies joint video-audio generation and editing.To fully empower and systematically evaluate this frame-work, we construct a comprehensive training corpus encompassing video-audio generation and editing datasets, and introduce unified benchmarkstailored for these multimodal tasks. |
Xiaohu Huang; Haoyang He; Hao Zhou; Qiangpeng Yang; Min Zheng; Kai Han; | code |
| 63 | UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose UniRec-0.1B, a unified recognition model with only 0.1B parameters.Finally, we develop a comprehensive evaluation benchmark covering Chinese and English documents from multiple domains and with multiple levels. |
Yongkun Du; Zhineng Chen; Yazhen Xie; Weikang Bai; Hao Feng; Wei Shi; Yuchen Su; Can Huang; Yu-Gang Jiang; | code |
| 64 | Towards Purified Multi-Label Test-Time Adaptation of Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While introducing region-level cues helps isolate class-specific evidence, such regional evidence can also be unreliable under distribution shifts, making its identification and utilization non-trivial. To address these issues, we introduce PuRF, a novel PuRiFication-driven cache-based method for multi-label test-time adaptation of vision-language models. |
Yiwen Liang; Hui Chen; Yizhe Xiong; Mengyao Lyu; Yuhan Cao; Zijia Lin; SHUAICHENG NIU; Sicheng Zhao; Jungong Han; Guiguang Ding; | code |
| 65 | Make Geometry Matter for Spatial Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose GeoSR, a framework designed to make geometry matter by encouraging VLMs to actively reason with geometry tokens. |
Shihua Zhang; Qiuhong Shen; Shizun Wang; Tianbo Pan; Xinchao Wang; | code |
| 66 | GKDT: General Keypoint Detection Transformer Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Based on MegaKPT, we develop GKDT, a simple, flexible and powerful DINOv3 based Transformer model for General Keypoint Detection.To this end, we firstly present a largescale unified keypoint dataset called MegaKPT. |
Changsheng Lu; Yuxin Chen; Haokun GUI; Rong Wang; Jie Yang; Harry Yang; Anton van den Hengel; Jiaya Jia; | code |
| 67 | Reinforcing Video Reasoning with Focused Thinking Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent advancements in reinforcement learning, particularlythrough Group Relative Policy Optimization (GRPO), have significantlyimproved multimodal large language models for complex reasoning tasks.However, two critical limitations persist: 1) they often produce unfo-cused, verbose reasoning chains that obscure salient spatiotemporal cues,and 2) binary rewarding fails to account for partially correct answers,resulting in high reward variance and ine!cient learning. In this pa-per, we propose TW-GRPO, a novel framework that enhances visualreasoning with focused thinking and dense reward granularity. |
Jisheng Dang; Jingze Wu; Teng Wang; Xuanhui Lin; Nannan Zhu; Hongbo Chen; WEISHI ZHENG; Meng Wang; Tat-Seng Chua; | code |
| 68 | ObjectForesight: Predicting 3D Object Trajectories from Human Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce ObjectForesight, a 3Dobject-centric dynamics model that predicts future 6-DoF poses andtrajectories of rigid objects from short egocentric video sequences. |
Rustin Soraki; Homanga Bharadhwaj; Ali Farhadi; Roozbeh Mottaghi; | code |
| 69 | GEM: Generative Supervision Helps Embodied Intelligence Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, a signifi-cant gap remains between the high-level semantic focus of standardtext-guided pre-training paradigms and the low-level spatial and phys-ical knowledge critical for execution in embodied environments. In thispaper, we introduce GEM, a Generative-supervised Embodied vision-language Model designed to bridge this divide. |
Ruowen Zhao; Bangguo Li; Zuyan Liu; Yinan Liang; junliang ye; FANGFU LIU; Diankun Wu; Zhengyi Wang; Xumin Yu; Yongming Rao; Han Hu; Jun Zhu; | code |
| 70 | Locality-Aware Continual Unlearning for Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Locality-Aware Target Selec-tion chooses, for each forget prompt, the context-preserving mappingprompt that the diffusion model itself treats as most similar to the orig-inal prompt, measured by score-prediction distance (how differently themodel denoises the same noisy image under two text conditions), ensur-ing each update is as small and targeted as possible. |
Naveen George; Naoki Murata; Yuhta Takida; Konda Reddy Mopuri; Yuki Mitsufuji; | code |
| 71 | Any to Full: Prompting Depth Anything for Depth Completion in One Stage Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we present Any2Full, a one-stage,domain-general, and pattern-agnostic framework that reformulates com-pletion as a scale-prompting adaptation of a pretrained MDE model.To address varying depth sparsity levels and irregular spatial distribu-tions, we design a Scale-Aware Prompt Encoder. |
Zhiyuan Zhou; Ruofeng Liu; TAICHI LIU; Weijian Zuo; Shanshan Wang; Zhiqing Hong; Desheng Zhang; | code |
| 72 | Partial Skeleton Visibility for Action Recognition: A Constrained Field-of-View Approach Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Resourcesrelated to the constrained FoV setting used in this work are available at:https://github.com/yaa1haa1/PartialVisGraph. |
Yingjie Dai; Tianyang Xu; Yanglin Deng; Xiao-Jun Wu; Josef Kittler; | code |
| 73 | UniDrive-WM: Unified Understanding, Planning and Generation World Model For Autonomous Driving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose UniDrive-WM, a unified VLM-based world model that jointly performs driving-scene understanding, trajectory planning, and trajectory-conditioned fu-ture image generation within a single architecture. |
Zhexiao Xiong; Xin Ye; Burhaneddin Yaman; Sheng Cheng; Yiren Lu; Jingru Luo; Nathan Jacobs; Liu Ren; | code |
| 74 | VGGT-World: Transforming VGGT Into An Autoregressive Geometry World Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present VGGT-World, a geometry world model that side-stepsvideo generation entirely and instead forecasts the temporal evolution offrozen geometry-foundation-model (GFM) features. |
Xiangyu Sun; Shijie Wang; Fengyi Zhang; Lin Liu; Caiyan Jia; Ziying Song; Zi Helen Huang; Yadan Luo; | code |
| 75 | From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce VWG-Bench (Video World Generalist Benchmark), a comprehensive benchmark spanning 9 reasoning dimensions and 38 fine-grained tasks. |
Meng Luo; Yicheng Liu; Jiahao Wang; Yuanxing Zhang; Xin Tao; Pengfei Wan; Kun Gai; Hao Fei; | code |
| 76 | AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, the performance of consistency-distilled models often degrades as more sampling steps are allocatedat test time, limiting their effectiveness for any-step video diffusion.We argue that this limitation arises because consistency distillation re-places the original probability-flow ODE trajectory with a consistency-sampling trajectory, weakening the desirable test-time scaling behaviorof ODE sampling. To address this limitation, we introduce AnyFlow,the first any-step video diffusion distillation framework based on flowmaps. |
Yuchao Gu; Guian Fang; Yuxin Jiang; Weijia Mao; Song Han; Han Cai; Mike Zheng Shou; | code |
| 77 | OmniMapBench: Benchmarking Visual-Centric Reasoning on Diverse Map Documents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Recent advancements in LVLMs necessitate robust bench-marks for complex, visually grounded reasoning. A critical limitation isidentified in many document understanding benchmarks: … |
Yang Chen; Yufan Shen; Yunwen Li; Minghao Liu; Tuney Tianyu; Bin Fu; Qunshu Lin; Zhi Yu; Botian Shi; | code |
| 78 | Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a paradigm shift by leveragingthe implicit spatial prior within large-scale video generation models. |
Xianjin Wu; Dingkang Liang; Tianrui Feng; Kui Xia; Yumeng Zhang; Xiaofan Li; Xiao Tan; Xiang Bai; | code |
| 79 | Semantic-Aware, Physics-Informed, Geometry-Grounded Weather Synthesis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address these, we propose a Semantic-Aware, Physics-Informed, and Geometry-Grounded framework that steers an off-the-shelf video editor to synthesize diverse global appearances and detailed particle dynamics. |
Chenghao Qian; Nedko Savov; Lingdong Kong; Yeying Jin; Rui Song; Wenjing Li; Zhun Zhong; Jiaqi Ma; Gustav Markkula; Luc Van Gool; | code |
| 80 | EgoCogNav: Cognition-aware Human Egocentric Navigation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To facilitate research in thefield, we introduce the Cognition-aware Egocentric Navigation (CEN)dataset consisting of 6 hours real-world egocentric recordings capturingdiverse navigation behaviors in real-world scenarios. |
Zhiwen Qiu; Ziang Liu; Wenqian Niu; Tapomayukh Bhattacharjee; Saleh Kalantari; | code |
| 81 | SIMON: SImultaneous Multi-Object Navigation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose threekey modules: a content-aware attention mechanism for region-adaptivefocus, a multi-object energy guidance strategy for subtask-specific con-sistency, and a latent initialization technique for artifact suppression.Together, these components enhance visual coherence across inpainting,object relocation, and background preservation subtasks. |
Yifeng Zhu; Siyuan Huang; Jun Bao; Jun Yu; Buyu Liu; | code |
| 82 | MobileManiBench: Simplifying Model Verification for Mobile Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose asimulation-first framework to verify VLA architectures before real-worlddeployment and introduce MobileManiBench, a large-scale benchmarkfor mobile-based robotic manipulation. |
Wenbo Wang; Fangyun Wei; Qixiu Li; Xi Chen; Yaobo Liang; Chang Xu; Jiaolong Yang; Baining Guo; | code |
| 83 | Neutralizing Token Aggregation Via Information Augmentation for Efficient Test-Time Adaptation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we start by providingan analysis showing that token aggregation inherently leads to informa-tion loss, which cannot be fully mitigated by conventional norm-tuning-based TTA methods. |
Yizhe Xiong; Zihan Zhou; Yiwen Liang; Hui Chen; Zijia Lin; Xinhao Xu; Tianxiang Hao; Fan Zhang; Jungong Han; Guiguang Ding; | code |
| 84 | Rethink Backdoor Robustness in Vision Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: More-over, we propose a more robust attack strategy: by introducing slightperturbations to the trigger, existing attacks can be made significantlymore resistant to various defenses. |
Yichuan Mo; Dongxian Wu; Yifei Wang; Yisen Wang; | code |
| 85 | ParaFlow: Parallel Sampling for Flow Matching Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce ParaFlow, a training-free framework that recasts sampling as a system of Triangular Nonlinear Equations (TNEs) to enable step-level parallelism. |
Jianrong Lu; Bangwei Li; Haomin Zhang; Zhuoya Gu; Yongqing Lu; Jianhai Chen; Qinming He; | code |
| 86 | SparkVSR: Interactive Video Super-Resolution Via Sparse Keyframe Propagation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a novel interactive VSR framework dubbed SparkVSR that makes sparse keyframes a simple and expressive control signal. |
Jiongze Yu; Xiangbo Gao; Pooja Verlani; Akshay Gadde; Yilin Wang; Balu Adsumilli; Zhengzhong Tu; | code |
| 87 | MoAKE: Toward Unified All-in-One Action Quality Assessment Via Mixture of Action Knowledge Experts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a novel Mixture of Action Knowledge Experts (MoAKE) framework, designed to mitigate negative knowledge transfer caused by large semantic discrepancies among actions. |
Huangbiao Xu; Huanqi Wu; Xiao Ke; Jiaxin Cai; Junyi Wu; Jinglin Xu; | code |
| 88 | PPTArena: A Benchmark for PowerPoint Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce PPTArena, a benchmark for PowerPoint edit-ing that evaluates how agents modify real slides from natural-languageinstructions. |
Michael Ofengenden; Yunze Man; Ziqi Pang; Liang-Yan Gui; Yu-Xiong Wang; | code |
| 89 | CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Satellite-street scene pairs, with their complex contexts and extreme viewpoint variations, provide an ideal testbed. Motivated by this, we introduce CVSBench, a large-scale benchmark for evaluating cross-view spatial reasoning through satellite-street pairs. |
ruixun liu; Lingyu Zhang; Lanxuan Xue; Kaiyu Li; Bowen Fu; Xiangyong Cao; | code |
| 90 | MemoBench: Benchmarking World Modeling in Dynamically Changing Environments Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Video generation models aspire to simulate dynamic environ-ments, and several benchmarks now evaluate memory consistency acrossframes. However, most assess consistency only while … |
Haoyu Chen; Kaichen Zhou; Hang Hua; Kaile Zhang; Jingwen Qian; Wufei Ma; Haonan Chen; Chunjiang Liu; Yizhou Zhao; Xiaoyuan Wang; Weiyue Li; Alan Yuille; Paul Pu Liang; Yilun Du; | code |
| 91 | ADAPT: Attention Dynamics Alignment with Preference Tuning for Faithful MLLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we identify an internal signature of hallucination: progressive degradation of text-to-image cross-attention during generation, leading to specific failure patterns like unfocused or biased attention. |
Zhiyuan Yao; Zheren Fu; Zhixiao Zheng; Jiajun Li; Yi Tu; Zhendong Mao; | code |
| 92 | HippoCamp: Benchmarking Contextual Agents on Personal Computers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present HippoCamp, a new benchmark designed to eval-uate agents’ capabilities on multimodal file management. |
Zhe YANG; Shulin Tian; Kairui Hu; Shuai Liu; Hoang-Nhat Nguyen; Yichi Zhang; Zujin Guo; Mengying Yu; Zinan Zhang; Jingkang Yang; Chen Change Loy; Ziwei Liu; | code |
| 93 | General Incomplete Multimodal Learning Via Dynamic Quality Perception Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Toachieve reliable quality perception, we introduce a Noise-aware QualityEstimator that learns the mapping from corrupted features to noise in-tensity through controlled noise injection. |
Xiangyu Meng; shicai wei; | code |
| 94 | Stay Unique, Stay Efficient: Preserving Model Personality in Multi-Task Merging Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Decomposition, Thresholding, and Scaling (DTS), an approximation-based personalized merging framework that pushes task-specific storage efficiency. |
Kuangpu Guo; Aijing Yu; Jian Liang; Yuhe Ding; Zilei Wang; Ran He; Tieniu Tan; | code |
| 95 | Tactile Modality Fusion for Vision-Language-Action Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose TacFiLM, a lightweight modality-fusion ap-proach that integrates visual-tactile signals into vision-language-action(VLA) models. |
Charlotte Morissette; Amin Abyaneh; Wei-Di Chang; Anas Houssaini; David Meger; Hsiu-Chin Lin; Jonathan Tremblay; Gregory Dudek; | code |
| 96 | To Adapt or Not to Adapt? Selective Adaptation for Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce a newproblem of selective adaptation, which aims to determine whether agiven test sample should undergo adaptation or be skipped. |
Siru Jiang; Yuwei Liang; Jian Liang; Ran He; Tieniu Tan; | code |
| 97 | On The Vulnerability of Parameter-Level Defenses to Model Merging Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent works pro-pose parameter-level defenses that employ linear parameter transforma-tions to neutralize this threat. In this paper, we systematically analyzesuch defenses and reveal that their protected task vectors are inherentlysmall in magnitude. |
Kuangpu Guo; Jian Liang; Qingyan Zheng; Yu Yongcan; Zilei Wang; Ran He; Tieniu Tan; | code |
| 98 | GeoWorld: Providing Full-frame Geometry Features to Facilitate 3D Scene Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper,we present a two-stage method, named GeoWorld, that renovates theimage-to-3D scene generation pipeline by providing full-frame geometryfeatures. |
Yuhao Wan; Lijuan Liu; Jingzhi Zhou; Zihan Zhou; Xuying Zhang; dongbo zhang; Shaohui Jiao; Qibin Hou; Ming-Ming Cheng; | code |
| 99 | TAIHRI: Task-Aware 3D Human Keypoints Localization for Close-Range Human-Robot Interaction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose TAIHRI, the first Vision-Language Model (VLM) tailored for close-range HRI perception, capable of understanding users’ motion commands and directing the robot’s attention to the most taskrelevant keypoints. |
Ao Li; Yonggen Ling; Yiyang Lin; Yuji Wang; Yong Deng; Yansong Tang; | code |
| 100 | Bridging VideoQA and Video-Guided Agentic Tasks Via Generalized Keyframe Extraction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Furthermore, we observe that the performance of models on both VideoQA and video-guided agentic tasks critically depends on effective keyframe extraction. Based on this observation, we propose TASKER (Task-driven And Scene-aware Keyframe searchER), a keyframe extraction algorithm that jointly considers task relevance and scene dynamics to identify informative frames. |
Sunqi Fan; Qingle Liu; Runqi Yin; Meng-Hao Guo; Shuojin Yang; | code |
| 101 | GenAgent: Scaling Text-to-Image Generation Via Agentic Multimodal Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce GenAgent, an agentic framework that unifiesvisual understanding and generation. |
Kaixun Jiang; Yuzheng Wang; Junjie Zhou; Pandeng Li; Zhihang Liu; Chen-Wei Xie; Zhaoyu Chen; Yun Zheng; Wenqiang Zhang; | code |
| 102 | IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose IQA-T1, atool-based visual evidence reasoning framework that augments MLLMreasoning with explicit perceptual observations. |
Jinjian Wu; Jiaqi Tang; Wei Wei; Yingying Yan; Jianmin Chen; Botong Geng; Lei Zhang; Qifeng Chen; | code |
| 103 | Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: ForenAgent adopts a two-stage training pipeline with Cold Start andReinforcement Fine-Tuning to progressively improve tool interaction andreasoning adaptability. |
Fanrui Zhang; Qiang Zhang; Sizhuo Zhou; Jianwen Sun; Chuanhao Li; Jiaxin Ai; Yukang Feng; Yujie Zhang; Wenjie Li; Zizhen Li; Yifan Chang; Jiawei Liu; Kaipeng Zhang; | code |
| 104 | Actor As Its Own Critic: Unifying Region Understanding and Localization Via CycleGRPO Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introduces Actor as Its Own Critic, a unifiedreinforcement learning framework, Cycle Group Relative Policy Opti-mization (CycleGRPO), that jointly optimizes region understanding andlocalization for Multimodal Large Language Models (MLLMs). |
Xin Zhang; Haochen Wang; Yikang Zhou; Zhuochen Wang; Robby T. Tan; Xiangtai Li; | code |
| 105 | A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: How-ever, current evaluation paradigms index heavily on visual fidelity and se-mantic alignment, leaving a critical blind spot: they cannot reliably quan-tify whether generated dynamics actually obey the fundamental physicallaws required for reliable simulation. Assessing this physical plausibilityis inherently difficult due to a lack of physical metrics and the challengeof extracting metric-scale kinematics from uncalibrated video rollouts.To bridge this gap, we introduce CrashTwin, a physics-grounded eval-uation framework designed to stress-test the physical trustworthiness ofworld models. |
Nuo Chen; Lulin Liu; Zihao Li; Ziyao Zeng; Zihao Zhu; Wenyan Cong; Junyuan Hong; Yunhao Yang; Zhengzhong Tu; Yan Wang; Boris Ivanovic; Marco Pavone; Zhangyang Wang; Yang Zhou; Zhiwen Fan; | code |
| 106 | One4D: Unified 4D Generation and Reconstruction Via Decoupled LoRA Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present One4D, a unified framework for 4D generationand reconstruction that produces dynamic 4D content as synchronizedRGB frames and pointmaps. |
Zhenxing Mi; YUXIN WANG; Dan Xu; | code |
| 107 | MirrorPPR: Exemplar-Based Portrait Photo Retouching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In contrast, structuralportrait retouching involves extremely delicate and localized modifica-tions, making accurate extraction and transfer of these edits challeng-ing. To tackle this, we propose MirrorPPR, a novel framework specifi-cally designed to capture and transfer subtle structural retouching oper-ations. |
Zhihong Liu; Zheng Li; Jiachun Jin; Siqi Kou; Yitao Jian; Fengpei Yu; Zhijie Deng; | code |
| 108 | Towards Practical Lossless Neural Compression for LiDAR Point Clouds Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: LiDAR point clouds are fundamental to various applications, yet the extreme sparsity of high-precision geometric details hinders efficient context modeling, thereby limiting the compression speed and performance of existing methods. To address this challenge, we propose a compact representation for efficient predictive lossless coding. |
pengpeng yu; Haoran Li; Runqing Jiang; Dingquan Li; Jing Wang; Liang Lin; Yulan Guo; | code |
| 109 | PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Plus,we propose a LoRA-Switch Reference (LoRA-SR) scheme that avoidsfull-model duplication as reference for efficient training. |
Yuanhao Cai; Kunpeng Li; Menglin Jia; Jialiang Wang; Junzhe Sun; Feng Liang; Weifeng Chen; Felix Juefei-Xu; Chu Wang; Ali Thabet; Xiaoliang Dai; Xuan JU; Alan Yuille; Ji Hou; | code |
| 110 | Robust Self-Supervised Cross-Modal Super-Resolution Against Real-World Misaligned Observations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Previous methods either rely on simulated training dataor adopt suboptimal alignment strategies that overlook cross-modal de-pendencies, limiting their practical performance. To address these is-sues, we propose RobSelf, a self-supervised model that jointly optimizesa misalignment-aware feature translator and a content-aware referencefilter online. |
Xiaoyu Dong; Jiahuan Li; Ziteng Cui; Naoto Yokoya; | code |
| 111 | OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Previous discrete video VAEs suffer from un-stable training, long training time, and degraded reconstruction quality.We revisit the relationship between continuous and discrete VAEs andfind that bridging discrete and continuous representations improves dis-crete token learning. Based on this insight, we propose a unified progres-sive training framework that (i) jointly optimizes continuous and discretereconstructions within a single network, and (ii) progressively derives afamily of VAEs at different compression ratios, leading to faster conver-gence and better final performance. |
Yupeng Zhou; Zhen Li; Yuming Chen; Ziheng Ouyang; Ruoyi Du; Daquan Zhou; Bin Fu; Yihao Liu; Peng Gao; Ming-Ming Cheng; Qibin Hou; | code |
| 112 | Delving Into Latent Spectral Biasing of Video VAEs for Superior Diffusability Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a statisticalanalysis of video VAE latent spaces and identify two spectral propertiesessential for diffusion training: a channel-wise eigenspectrum dominatedby a few modes, and a spatio-temporal frequency spectrum biased towardlow frequencies. To induce these properties, we propose two lightweight,backbone-agnostic regularizers: Latent Masked Reconstruction and Lo-cal Correlation Regularization. |
Shizhan Liu; Xinran Deng; Zhuoyi Yang; Jiayan Teng; Xiaotao Gu; Jie Tang; | code |
| 113 | ZTRS: Zero-Human Demonstration End-to-end Autonomous Driving with Trajectory Scorer Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we proposeZTRS (Zero-human demonstration end-to-end autonomous driving withTRajectory Scorer) — a complete RL-based E2E planning paradigmtrained solely on real-world images and rule-based rewards, entirely with-out human demonstration. |
Zhenxin Li; Nadine Chang; Wenhao Yao; Xinglong Sun; Zi Wang; Maying Shen; Jingde Chen; Jingyu Song; Kailin Li; Zuxuan Wu; Shiyi Lan; Jose M Alvarez; | code |
| 114 | LightSTAR: Efficient Visual Document Retrieval Via Lightweight Selection with Vision-Adaptive Refinement Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Mean-while, we observe that user queries are typically keyword-anchored, con-taining semantically rich words that are expected to appear directly inthe visible text of relevant pages, offering an efficient cue for quickly nar-rowing down candidate pages. Building on this insight, we propose Light-STAR, an efficient framework that decomposes visual document retrievalinto: 1) LLM-free Visual Selection, which utilizes content-grounded queryencoding to focus on informative words and employs LLM-free visual em-beddings to produce a high-recall candidate set; and 2) Vision-adaptiveSemantic Refinement, which further performs fine-grained semantic match-ing exclusively on these top candidates via adaptive region-wise featurefusion to effectively combine textual and layout cues, optimized through ahardness-aware contrastive objective. |
Tongkun Guan; Haocheng Wang; Wei Shen; Xiaokang Yang; | code |
| 115 | QualiTeacher: Quality-Conditioned Pseudo-Labeling for Real-World Image Restoration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, wepropose QualiTeacher, a novel framework that transforms pseudo-labelquality from a noisy liability into a conditional supervisory signal. |
Fengyang Xiao; Jingjia Feng; Peng Hu; Yuhan Chen; Dingming Zhang; Lei Xu; Guanyi Qin; Lu Li; Chunming He; Sina Farsiu; | code |
| 116 | Ceptor: Vision-Language Model-Infused Diverse Guidance for Detecting Anything Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by the general process of hu-man object search, we designed Ceptor, a unified detector guided bydiverse prompts for open-set object detection. |
Jinyang Li; Bin-Bin Gao; Weifu Fu; Jingnan Luo; Hanqiu Deng; Yue Guo; Jun Liu; Yong Liu; Chengjie Wang; Wenbing Tao; | code |
| 117 | Beyond Isolated Scans: Cross-Phase Alignment of Structure and Topology for 3D Medical Pretraining Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: CAST employs a 3D CNN archi-tecture to explicitly align NCCT representations with CECT targets.Moving beyond conventional reconstruction, we introduce two feature-level constraints: (1) a Spectral Consistency module that utilizes3D wavelet decomposition to align frequency-aware boundaries whilesuppressing contrast-induced noise; (2) a Geometry-Aware Topo-logical Consistency module that preserves local relational graphsamong salient anatomical keypoints via dynamic top-hat sampling.Tosupport this, we construct a large-scale dataset comprising 13,850 pairedvolumetric CT scans. |
Wenzhuo xu; YANJIE ZHOU; Yujian Hu; Hongkun Zhang; Minfeng Xu; | code |
| 118 | Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a novel evaluation frameworkthat combines DermBench, a meticulously curated benchmark, withDermEval, a robust automatic evaluator, to enable clinically meaning-ful, reproducible, and scalable assessment. |
Yuhao Shen; Jiahe Qian; Zhangtianyi Chen; Juexiao Zhou; | code |
| 119 | What Moves? Localized Motion Representations for Compositional Scene Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, localized embeddings are often computed from cropped images or obtained by masking features after encoding, discarding the context needed to interpret motion. To address this, we introduce a promptable localized motion representation that produces persistent embeddings for user-specified regions defined by spatial masks. |
Frank Fundel; Malek Ben Alaya; Thomas Ressler-Antal; Stefan Andreas Baumann; Bjorn Ommer; | code |
| 120 | Logit Refiner: Improving Visual Autoregressive Models Via Intra-Scale Dependency Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Visual Autoregressive Models (VAR) generate images throughnext-scale prediction, producing all tokens within each scale in paral-lel. We show that this parallel decoding constitutes a mean-field-styleapproximation that discards spatial dependencies among same-scale to-kens, causing locally incoherent samples regardless of backbone capacity– a limitation of the decoding rule. |
Meimingwei Li; Stefan Andreas Baumann; Felix Krause; Bjorn Ommer; | code |
| 121 | Schroedinger’s Cat: Probabilistic Representation and Prediction of Potential Scene Kinematics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Predicting how a scene may evolve from partial observationsrequires reasoning about multiple possible futures rather than committingto a single trajectory. Existing approaches … |
Timy Phan; Jannik Wiese; Bjorn Ommer; | code |
| 122 | RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce RayDer, a unified, feed-forward transformer that consolidates camera estimation, scene reconstruction, and rendering into a single backbone, turning self-supervised NVS into a well-posed single-model scaling problem. |
Ulrich Prestel; Stefan Andreas Baumann; Nick Stracke; Bjorn Ommer; | code |
| 123 | Out of Sight, Out of Mind? Evaluating State Evolution in Video World Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Can these generated “worlds” evolveregardless of observation? To probe this question, we design a benchmarkto evaluate whether video world models can decouple state evolutionfrom observation. |
Ziqi Ma; Mengzhan Liufu; Georgia Gkioxari; | code |
| 124 | RiO-DETR: DETR for Real-time Oriented Object Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present RiO-DETR: DETR for Real-time OrientedObject Detection, the first real-time oriented detection transformer tothe best of our knowledge. |
Zhangchi Hu; Yifan Zhao; Yansong Peng; Wenzhang SUN; Xiangchen Yin; Jie Chen; Peixi Wu; Hebei Li; xinghao wang; Dongsheng Jiang; Xiaoyan Sun; | code |
| 125 | On-Policy Diffusion Reinforcement Learning Meets Off-Policy Quality Anchoring Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite offering a direct optimization signal, theapproach is fundamentally self-limiting and prone to reward overfittingand mode collapse due to its myopic guidance. In this paper, we pro-pose a new framework, namely DiffusionCompass, that breaks such lim-itation by strategically integrating off-policy guidance. |
Zunxu Liu; Zhaofan Qiu; Yazhen Xie; Yingwei Pan; Ting Yao; Tao Mei; | code |
| 126 | What Matters in RL-Based Methods for Object-Goal Navigation? An Empirical Study and A Unified Framework Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present a large-scale empirical study of modular RL-based ObjectNav systems. |
Hongze Wang; Boyang Sun; Jiaxu Xing; Fan Yang; Marco Hutter; Dhruv Shah; Davide Scaramuzza; Marc Pollefeys; | code |
| 127 | Unlocking Complex Image Editing Via Natively Interleaved Visual Textual CoT with Deep Confidence Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While Chain of Thought (CoT) has been ex-plored to enhance reasoning, purely textual CoT or coordinate basedprompts are fundamentally limited in representing intricate visual lay-outs and lack the pixel level cues necessary for precise editing. To addressthese challenges, we propose Unlocking Complex Image Editing viaMultimodal Reasoning Edit (MURE), a natively multimodal frame-work that shifts the editing process from purely verbal reasoning to asequence of native interleaved textual and visual rationales. |
Zhentao Zou; Zhengrong Yue; Kunpeng Du; Binglei Bao; Hanting Li; Haizhen Xie; Guozheng Xu; Yue Zhou; jie hu; Xue Jiang; Xinghao Chen; | code |
| 128 | Seeing Fast and Slow: Learning The Flow of Time in Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In thispaper, we study time as a learnable visual concept and develop modelsfor reasoning about and manipulating the flow of time in videos.We first exploit the multimodal cues and temporal structure naturallypresent in videos to learn, in a self-supervised manner, to detect speedchanges and estimate playback speed. |
Yen-Siang Wu; Rundong Luo; Jingsen Zhu; Tao Tu; Ali Farhadi; Matthew Wallingford; Yu-Chiang Frank Wang; Steve Marschner; Wei-Chiu Ma; | code |
| 129 | Exploring Efficient Reasoning Segmentation with Small Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address the challenges of scaling reasoning segmentationto SLMs, we propose two key designs: First, unlike LLMs, SLMs havelimited capacity to provide sufficient spatial instruction cues for mask de-coding. We introduce register tokens that aggregate text-conditioned spa-tial features via the SLM’s self-attention and inject them into the visualtoken stream to enrich the mask decoder’s inputs. |
Changsong Wen; Zelin Peng; Yu Huang; Xiaokang Yang; Wei Shen; | code |
| 130 | DC-Gen: Post-Training Diffusion Acceleration with Deeply Compressed Latent Space Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Existing latent diffusion models excel at visual generationand editing tasks, employing autoencoders to project RGB images andvideos into latent spaces. However, their … |
Wenkun He; Yuchao Gu; Junyu Chen; Junyi Wu; Wenhang Ge; Dongyun Zou; Yujun Lin; Zhekai Zhang; Haocheng Xi; Muyang Li; Ligeng Zhu; Jincheng YU; Junsong Chen; Enze Xie; Song Han; Han Cai; | code |
| 131 | There and Back Again: A Flexible-Frame Transformer for Multi-Exposure Fusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, conventional MEF techniquesare typically designed for a fixed number of inputs, forcing deploymentsystems to maintain separate models for different frame-count require-ments, which undermines deployment efficiency. To address this limita-tion, we propose FreeMEF, the first flexible-frame transformer for MEFthat seamlessly accommodates varying numbers of input exposures with-out retraining or architectural changes. |
Lishen Qu; Yao Liu; shihao zhou; Jie Liang; Hui Zeng; Lei Zhang; Jufeng Yang; | code |
| 132 | Region-Aware Multimodal Large Language Model Via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current CT report generation frameworks predominantlyrely on global feature representations, often failing to capture region-specific details and potentially missing certain abnormalities. To over-come this limitation, we propose MedRegion-CT, a region-focused mul-timodal large language model framework featuring three key innova-tions. |
Sunggu Kyung; Jinyoung Seo; Hyunseok Lim; Dongyeong Kim; Hyungbin Park; Jimin Sung; Wooyoung Jo; Yoojin Nam; Namkug Kim; | code |
| 133 | ReShift: Aha-Moment-Driven Reasoning-Level Backdoor Attacks on Vision–Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper,we propose ReShift, the novel aha-moment-driven reasoning-level back-door framework that explicitly redirects the internal chain-of-thought(CoT) trajectory while preserving surface-level coherence. |
Zhihao Dou; Qinjian Zhao; Zhiqiang Gao; Sumon Biswas; | code |
| 134 | Diffusion-SDPO: Safeguarded Direct Preference Optimization for Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Consequently, degradation ofthe less-preferred outputs can become sufficiently severe that the pre-ferred branch is also adversely affected even as the margin grows. Toaddress this, we introduce Diffusion-SDPO, a safeguarded update rulethat preserves the winner by adaptively scaling the loser gradient ac-cording to its alignment with the winner gradient. |
Minghao Fu; Guo-Hua Wang; Tianyu Cui; Qing-Guo Chen; Zhao Xu; Weihua Luo; Kaifu Zhang; | code |
| 135 | PLOT: Pseudo-Labeling Via Object Tracking for Monocular 3D Object Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we presentPLOT (Pseudo-Labeling via Object Tracking), a framework that gen-erates 3D annotations from monocular videos without auxiliary sen-sors or model retraining. |
SeokYeong Lee; Sithu Aung; JunYong Choi; Seungryong Kim; Ig-Jae Kim; Junghyun Cho; | code |
| 136 | Q-TriM: Question-Guided Tri-Modal Attention for Audio–Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: For Q-TriM, we propose a novel framework for attention operation incorporating video and audio conditioned on text. |
SungHun Kim; Seung Baek; | code |
| 137 | OCTOPUS: Multi‑Agentic Universal Compositional Visual Retrieval Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce OCTOPUS, a novelmulti-agentic assistant for tool-integrated progressive self-improvementand user-friendly synergistic compositional retrieval. |
Zhangtao Cheng; Bozhu Zheng; Ting Zhong; Fan Zhou; | code |
| 138 | Accelerating Diffusion Transformers with Gaussian Process Rectified Feature Cache Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Prediction-based feature caching is widely usedto accelerate diffusion transformers; however, as the number of stepsincreases, the deviation between its predictions and the reference full-compute trajectory gradually grows. An intuitive idea is to use an onlineregression model to dynamically correct this deviation, but it faces theissue of label data being unavailable during the acceleration process.This paper presents a statistical observation that the residuals betweenthe features of full computation steps using caching methods and refer-ence full-compute trajectory locally exhibit a zero-mean Gaussian dis-tribution. |
Zhirong Shen; Rui Huang; Chang Zou; Shikang Zheng; Jiacheng Liu; Peiliang Cai; zhengyi shi; Yaosong Du; Liang Feng; Xiaobing Tu; Jinkui Ren; Xiantao Zhang; Linfeng Zhang; | code |
| 139 | OSOR: One-Step Diffusion Inpainting for Effect-Aware Object Removal Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: With billions of parameters and tensof denoising steps, diffusion-based models achieve this goal at the ex-pense of massive computational cost, limiting their use in interactiveapplications and edge devices. To solve this problem, we present OSOR(One-Step Object Removal), which achieves efficient, effect-aware, andmask-robust object removal at the same time. |
Qinming Zhou; Chenxi Sun; Deyang Kong; Junhao He; Xiangheng Tang; Peike Yu; Haotian Wu; Leilei Cao; Linfeng Zhang; | code |
| 140 | LinCa: Accelerating Diffusion Models Via Learnable Decomposed Feature Caching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose LinCa , a fea-ture caching framework based on learnable invertible networks. |
Jinshan Liu; Haoran Qin; Xiaobing Tu; Jiacheng Liu; Jiahui Hu; zhengan yan; Yukun Xie; Kerui Shen; Jinkui Ren; Yuqi Lin; Xiantao Zhang; Linfeng Zhang; | code |
| 141 | AViTS:Adaptive Spatiotemporal Token Selection for Efficient Dynamic-Resolution Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose AViTS, an adaptive spatiotemporal token selection framework for dynamic-resolution DiTs. |
Haoran Qin; zhengan yan; Shikang Zheng; Xiaobing Tu; Jiacheng Liu; Yuqi Lin; Chang Zou; Jinshan Liu; Peiliang Cai; Xiantao Zhang; Jinkui Ren; Linfeng Zhang; | code |
| 142 | HolisticSemGes: Semantic Grounding of Holistic Co-Speech Gesture Generation with Contrastive Flow-Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a Contrastive Flow Matching-basedco-speech gesture generation model that uses mismatched audio–textconditions as negatives, training the velocity field to follow the cor-rect motion trajectory while repelling semantically incongruent trajec-tories. |
Lanmiao Liu; Esam Ghaleb; asli ozyurek; Zerrin Yumak; | code |
| 143 | Stand Up and Move: Benchmarking Interactive Spatial Intelligence in WalkerBench Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our empirical evaluation revealsa critical “representation misalignment”: VLMs’ one-dimensional linearcontext structures are fundamentally incompatible with the topologicalnature of 3D environments, causing inevitable spatial forgetting duringnavigation. To overcome this, we propose Spatial-IDE, which breaks fromlinear dialogue history via two mechanisms: (1) State Externalization,transforming implicit observations into an Explicit Topological Memory(ETM); and (2) Cognitive Decoupling, disentangling visual understand-ing into Goal-Directed Perception and Spatial Reasoning, letting theVLM focus on pure high-level decision-making. |
Zhiqi Ge; Gang Yang; Ziyang Pan; Jingzhe Zhu; Yuancheng Gu; Juncheng Li; Qizhou Wang; Rui Tang; Siliang Tang; Jun Xiao; Yueting Zhuang; | code |
| 144 | VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce VLZip, a framework that uni-_x001C_es visual and textual compression for high-_x001C_delity reasoning within apure Transformer. |
Yuqi Zhang; Cheng Chen; Yuyu Guo; Wenjie Yang; Lingchen Meng; Peng Di; Hang Yu; Zuxuan Wu; Yu-Gang Jiang; | code |
| 145 | PhysRVG: Physics-Aware Unified Reinforcement Learning for Video Generative Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods attemptto tackle this problem through physical data augmentation, dynamicspre-simulation, or reinforcement learning with VLM ratings, but noneof these approaches accurately reflect physical principles or enable themodel to internalize physical knowledge. Motivated by these considera-tions, we introduce reinforcement learning with physically verifiable re-wards. |
Qiyuan Zhang; Biao Gong; Shuai Tan; Zheng Zhang; Xing Zhu; Yujun Shen; Yuyuan Li; kelu Yao; Chunhua Shen; Changqing Zou; | code |
| 146 | Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution Via Auto-Regressive Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Stream-DiffVSR, a causally conditioned diffusion framework for efficient online VSR. |
Hau-Shiang Shiu; Chin-Yang Lin; Zhixiang Wang; Chi-Wei Hsiao; Po-Fan Yu; Yu-Chih Chen; Yu-Lun Liu; | code |
| 147 | From Synchrony to Sequence: Exo-to-Ego Generation Via Interpolation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While paired supervision is available, synchronized exo-egodata inherently introduces substantial spatio-temporal and geometricdiscontinuities, violating the smooth-motion assumptions of standardvideo generation benchmarks. We identify this synchronization-inducedjump as the central challenge and propose Syn2Seq-Forcing, a sequen-tial formulation that interpolates between the source and target videosto form a single continuous signal. |
Mohammad Mahdi; Nedko Savov; Danda Paudel; Luc Van Gool; | code |
| 148 | SkillSpotter: Pose-Aware Multi-View Skilled Action Detection and Grading in Ego-Exo Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: SkillSpotter ’s modulestransfer to other temporal action detection models with consistent gainsand our method generalizes beyond Ego-Exo4D to HoloAssist.Code: https://github.com/eth-siplab/SkillSpotter |
Björn Braun; Christian Holz; | code |
| 149 | Geometry-Aware Single-Image 4D Synthesis Via Dense Trajectory Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address these, we present MoGe4D (Motion and GeometryAware image-to-4D Synthesis), a geometry-conditioned framework for single-image 4D synthesis that models a scene as dense 4D point trajectories. |
Yanran Zhang; Ziyi Wang; Wenzhao Zheng; Zheng Zhu; Jie Zhou; Jiwen Lu; | code |
| 150 | QASA: Quality-Guided K-Adaptive Slot Attention for Unsupervised Object-Centric Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: K-fixed methods, such as SPOT [20]and DINOSAUR [33], show large performance fluctuations as K varies. |
Tianran Ouyang; Xingping Dong; Jing Zhang; Mang Ye; Kaihao Zhang; Bo Du; | code |
| 151 | WinTok: A Win-Win Hybrid Tokenizer Via Decomposing Visual Understanding and Generation with Transferable Tokens Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose WinTok, a concisehybrid tokenizer that achieves a win-win performance by explicitly de-coupling the two objectives. |
Yiwei Guo; Shaobin Zhuang; Zhipeng Huang; Canmiao Fu; Chen Li; Jing LYU; Yali Wang; | code |
| 152 | MoCA3D: Monocular 3D Bounding Box Prediction in The Image Plane Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce MoCA3D, a Monocular,Class-Agnostic 3D model that predicts projected 3D bounding box cor-ners and per-corner depths without requiring camera intrinsics at in-ference time. |
Changwoo Jeon; Rishi Upadhyay; Achuta Kadambi; | code |
| 153 | ReconPhys: Reconstruct Appearance and Physical Attributes from Single Video Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing approaches leverage differen-tiable rendering for per-scene optimization, recovering geometry anddynamics but requiring expensive tuning or manual annotation, whichlimits practicality and generalizability. To address this, we propose Re-conPhys, the first feedforward framework that jointly learns physicalattribute estimation and 3D Gaussian Splatting reconstruction from asingle monocular video. |
Boyuan Wang; Xiaofeng Wang; Yongkang Li; Zheng Zhu; Yifan Chang; Angen Ye; Guosheng Zhao; Chaojun Ni; Guan Huang; Yijie Ren; Yueqi Duan; Xingang Wang; | code |
| 154 | VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Pretraining Vision-Language-Action (VLA) policies on internet-scale video is appealing, yet current latent-action objectives often learnthe wrong thing: they remain anchored to pixel variation rather thanaction-relevant state transitions, making them vulnerable to appearancebias, nuisance motion, and information leakage. We introduce VLA-JEPA,a JEPA-style pretraining framework that sidesteps these pitfalls by design.The key idea is leakage-free state prediction: a target encoder produceslatent representations from future frames, while the student pathwaysees only the current observation—future information is used solely assupervision targets, never as input. |
Jingwen Sun; Wenyao Zhang; Zekun Qi; Shaojie Ren; Zezhi Liu; Hanxin Zhu; Guangzhong Sun; Xin Jin; Zhibo Chen; | code |
| 155 | DiCoBench: Benchmarking Multi-Image Fine-Grained Perception Via Differential and Commonality Visual Cues Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To bridgethis gap, we introduce DiCoBench, a comprehensive, multi-image high-resolution benchmark designed for cross-image fine-grained perception.DiCoBench consists of 765 meticulously curated samples categorizedinto two progressive tracks: Differential Visual Cues and Commonal-ity Visual Cues, covering 8 distinct perception tasks.By formulatingthe benchmark as a multiple-choice question task and utilizing high-resolution imagery (approaching 2K), we eliminate evaluation metricbias and pose a substantial challenge to current state-of-the-art MLLMs.Our extensive evaluation of 18 diverse MLLMs reveals a striking perfor-mance gap compared to human accuracy (98.3%), with top-performingmodels struggling significantly with micro-scale detail capture. |
Geng Li; Yuxin Peng; | code |
| 156 | GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While Large Vision-Language Models (LVLMs) have signifi-cantly advanced GUI agents’ capabilities in parsing textual instructions,interpreting screen content, and executing tasks, a critical challenge per-sists: the irreversibility of agent operations—where a single erroneous ac-tion can trigger catastrophic deviations. To address this, we propose theGUI Action Critic’s Data Flywheel System (GAIA), a training frame-work that enables the models to have iterative critic capabilities, whichare used to improve the Test-Time Scaling (TTS) of basic GUI agents’performance. |
Shaokang Wang; Pei Fu; Ruoceng Zhang; Shaojie Zhang; Xiuwen Xi; Jiahui Yang; Bin Qin; Ying Huang; Zhenbo Luo; Jian Luan; | code |
| 157 | GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: BrowseComp offers a text-only testbed for such agents,but existing multimodal benchmarks rarely require both weak visual cuescomposition and BrowseComp-style multi-hop verification. Geolocationis a natural testbed because answers depend on combining multipleambiguous visual cues and validating them with open-web evidence.Thus, we introduce GeoBrowse, a geolocation benchmark that combinesvisual reasoning with knowledge-intensive multi-hop queries. |
Xinyu Geng; Yanjing Xiao; Yuyang Zhang; Hanwen Wang; Xinyan Liu; Rui Min; Tianqing Fang; Yi Ren Fung; | code |
| 158 | RASA: Disentangled Spatial-Motional Priors for Cross-Identity Character Animation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inthis paper, we introduce Reference-Aware Structural Alignment(RASA), a novel framework that systematically disentangles spatialmapping from motion control by injecting structured priors into a Diffu-sion Transformer (DiT). |
Zhen Xiao; Zhen Shen; Zhaofan Qiu; Ting Yao; Xueliang Liu; Tao Mei; | code |
| 159 | Wat3R: Underwater 3D Geometry Learning Without Underwater Annotations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Pioneering methods rely onmassive dense annotations that are impractical in underwater settings.In this paper, we propose Wat3R, a cross-domain semi-supervised learn-ing framework designed to adapt feed-forward 3D reconstruction modelsfrom air to underwater scenes. |
Jiangwei Ren; Xingyu Jiang; Zijie Song; Wei Xu; Hongkai Lin; Dingkang Liang; Xiang Bai; | code |
| 160 | KineBench: Benchmarking Embodied World Models Via IDM-Free Kinematic Grounding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This creates an unavoid-able attribution ambiguity between world model inaccuracies and ex-tractor errors. To reduce this ambiguity, we present KineBench, anIDM-free closed-loop benchmark for EWMs, built upon an explicit kine-matic grounding pipeline. |
Zeyu Liu; Zhangzhe Zhu; Yang Zhang; Chenyou Fan; Chenjia Bai; Xuelong Li; | code |
| 161 | Skin-R1: Clinical Knowledge-Guided Dermatological Diagnosis Using Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, their trustworthiness and clinical utility remain limited by three key challenges: heterogeneous datasets with inconsistent diagnostic labels and concept annotations, the lack of grounded diagnostic rationales for reliable reasoning supervision, and limited scalability when transferring knowledge from small, densely annotated datasets to large collections with sparse labels. To address these challenges, we propose Skin-R1, a dermatology-oriented VLM that integrates textbook-grounded clinical reasoning supervision with reinforcement learning (RL) to improve the accuracy and robustness of diagnostic prediction. |
Zehao Liu; Weijieying Ren; Jipeng ZHANG; Tianxiang Zhao; Jingxi Zhu; Xiaoting Li; Vasant Honavar; | code |
| 162 | X2SAM: Any Segmentation in Images and Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce X2SAM, a unified segmentation MLLMthat extends any-segmentation capabilities from images to videos.We further introduce the Video Visual Grounded(V-VGD) segmentation benchmark, which evaluates whether a modelcan segment object tracks in videos from interactive visual prompts.With a unified joint training strategy over heterogeneous image and videodatasets, X2SAM delivers strong video segmentation performance, re-mains competitive on image segmentation benchmarks, and preservesgeneral image and video chat ability. |
Hao Wang; Limeng Qiao; Chi Zhang; Lin Ma; Guanglu Wan; Xiangyuan Lan; Xiaodan Liang; | code |
| 163 | Optimizing Mesh Animation from Video Via Shape Flow Guidance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address thislimitation, we propose Shape Flow Guidance (SFG), a sequence of 3Dshapes derived from videos, which serves as explicit 3D supervision formesh animation. |
Jingqiao Xiu; Yicong Li; Angela Yao; | code |
| 164 | 3D Gaussian Splatting Compression with Object Scalability Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a framework toward scalable, finer-grained object-level 3DGS compression. |
Ruixiang Xue; Tong Chen; Zhan Ma; | code |
| 165 | VCP-DCN: Beyond Visual Concealed Property Via Depth Collaborative Network for Camouflaged Object Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Camouflaged Object Detection (COD) aims to identify andsegment camouflaged objects in complex environments, which are oftenconcealed because their color and texture are similar to the background.Several existing COD methods introduce depth maps to boost detec-tion performance via learning complementary RGB-D features, ignor-ing modality-specific characteristics of concealed objects in the depthdomain. To address this issue, we propose a depth collaborative net-work, called VCP-DCN, to mine distinguishable multi-modality featuresbeyond visual concealed prototype in depth domain. |
Songsong Duan; Xi Yang; Nannan Wang; | code |
| 166 | GAP-Track: Bridging The Resolution Gap for Cross-Resolution RGBT Tracking Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose GAP-Track, an efficientframework that bridges the resolution gap by enabling high-precisiontracking of low-resolution inputs. |
Shiyu Zhang; Tianyang Xu; Zhangyong Tang; Wang He; Xiao-Jun Wu; Josef Kittler; | code |
| 167 | Together, Then Apart: Balancing Alignment and Distinctiveness for Multimodal Survival Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This motivates a representation learning principle that we refer to as Together Then Apart. Based on this idea, we propose TTA, a framework that balances cross-modal alignment and representation distinctiveness. |
Wenjing Liu; Qin Ren; Wen Zhang; Yuewei Lin; Chenyu You; | code |
| 168 | GridFlow: Structured Latent Flow for Seamless City-Scale 3D Point Cloud Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To sup-port standardized evaluation, we build on public 3D data sources to in-troduce City3D-MultiGen, a benchmark of 163K densely annotated tilesfrom Melbourne and London with aligned point clouds, satellite images,semantic maps, and elevation data. |
Xinyu Wang; Muhammad Ibrahim; Atif Mansoor; Ajmal Mian; | code |
| 169 | MArFE: Multi-Contrast MRI Arbitrary Scale Super-Resolution with Fourier Enhancement Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: How-ever, related research exists two following issues: 1) Implicit neural rep-resentations (INR) as mainstream methods are prone to spectral bias,which limits their ability to recover high-frequency details; 2) Multi-contrast MRI is often used as the effective prior, but lacks the targetednetwork design to further merge INR positional information. To solvethese problems, we propose a Fourier-enhanced implicit framework forarbitrary-scale multi-contrast MRI super-resolution (MArFE). |
ZHIWEN SHI; | code |
| 170 | FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Music-to-dance generation aims to translate auditory sig-nals into expressive human motion. Despite promising progress, existingmethods remain underexplored in achieving … |
Kaixing Yang; Xulong Tang; Ziqiao Peng; Xiangyue Zhang; Chubin Chen; xukun zhou; Puwei Wang; Hongyan Liu; Jun He; | code |
| 171 | OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To the best of our knowledge, CIPE-Dance is the largest dataset for dance video generation to date, comprising 300k high-quality clips (over 400 hours) and covering diverse dancers, environments, and dance genres. To overcome the method limitation, we propose OmniDance, a frameworklevel recipe for integrating music into a TI2V foundation model without sacrificing its original controllability or visual fidelity. |
Kaixing Yang; Jiashu Zhu; Xulong Tang; Ziqiao Peng; Xiangyue Zhang; Chubin Chen; Puwei Wang; Jiahong Wu; Xiangxiang Chu; Hongyan Liu; Jun He; | code |
| 172 | DreamWorld: Geometry-Grounded Video Diffusion for 3D-Consistent World Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To alleviate this, we presentDreamWorld, a new recipe of world model that novelly bridges the strongspatial structure priors of 3D foundation models with the high-fidelitygenerative capabilities of video diffusion models for geometry-consistent3D scene generation. Specifically, given the input image and camera tra-jectory, DreamWorld first learns a geometry video diffusion model topredict compact geometry features for the target novel views, function-ing as explicit structure pivots to reflect the underlying 3D spatial layout.To achieve this, we introduce a distillation paradigm that transfers high-level structural knowledge from a pretrained 3D foundation model to thediffusion model, thereby enabling it to produce geometrically consistentand spatially coherent features. |
Haibo Yang; Yang Chen; Yingwei Pan; Zhineng Chen; Ting Yao; Tao Mei; | code |
| 173 | OTCache: Optimal Transport for Geometry-Aware Caching in Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose OTCache, a training-free framework for ac-celerating diffusion sampling via caching schedule prediction. |
Huanlin Gao; Fang Zhao; Qiang Hui; Fuyuan Shi; Shaoan Zhao; Yantao Li; Chao Tan; Ting Lu; Yuren You; Kai Wang; Shiguo Lian; | code |
| 174 | Unsafe By Reciprocity: How Generation–Understanding Coupling Undermines Safety in Unified Multimodal Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we investigate whether cross-functionality reciprocity itself constitutes a structural source of vulnerability in UMMs. |
Kaishen Wang; Heng Huang; | code |
| 175 | Denoising The Deep Sky: Physics-Based CCD Noise Formation for Astronomical Imaging Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a physics-based noisesynthesis framework tailored to CCD noise formation in the telescope.The pipeline models photon shot noise, photo-response non-uniformity,dark-current noise, readout effects, and localized outliers arising fromcosmic-ray hits and hot pixels. |
Shuhong Liu; Xining Ge; Ziying Gu; Quanfeng Xu; Ziteng Cui; Lin Gu; Xuangeng Chu; Jun Liu; Dong Li; Tatsuya Harada; | code |
| 176 | In-Context Sync-LoRA for Portrait Video Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Sync-LoRA, a method for editing portrait videos that achieves high-quality visual modifications while maintaining frame-accurate synchronization and identity consistency. |
Sagi Polaczek; Or Patashnik; Ali Mahdavi-Amiri; Danny Cohen-Or; | code |
| 177 | Robust and Efficient Monocular 3D Gaussian SLAM for Kilometer-Scale Outdoor Scenes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose KiloGS-SLAM, a highly efficient and robust monocular 3DGS-SLAM system that jointly addresses both bottlenecks. |
Sicheng Yu; Dongxu Shen; Beizhen ZHAO; Ding Guanzhi; Hao Wang; | code |
| 178 | Taming Text-to-Sounding Video Generation Via Advanced Modality Condition and Interaction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: For the sec-ond, we introduce BridgeDiT, a dual-tower di!usion transformer thatemploys Dual Cross-Attention (DCA) as a bidirectional bridge betweenvideo and audio streams, which we show through systematic compari-son to be the optimal fusion strategy for the dual-tower paradigm. |
Kaisi Guan; Xihua Wang; Zhengfeng Lai; Xin Cheng; Peng Zhang; Xiaojiang Liu; Ruihua Song; Meng Cao; | code |
| 179 | LogFA: Efficient Feature-Space Data Augmentation for Egocentric Temporal Action Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose LogFA (Local-global Featurespace Augmentation), a novel framework that performs data augmentation directly in the feature space for Temporal Action Segmentation (TAS). |
Zijia Lu; Ehsan Elhamifar; | code |
| 180 | PhysEdit: Physically Consistent Image Editing Via Causal Enforcement Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While instruction-guided image editing has seen significantstrides, most existing models fail to preserve physical laws and causal-ity. This limitation stems from the sparsity of physical supervision sig-nals, coupled with the static, non-causal architecture of prevailing frame-works, which together hinder a deep comprehension of the physical world.To bridge this gap, we present PhysEdit, a novel framework that en-forces physical consistency in image editing through causal generationand physics-aware reinforcement learning. |
Siqi Wan; Jingwen Chen; Yehao Li; Yingwei Pan; Ting Yao; Tao Mei; | code |
| 181 | BeTTER: Diagnose The Illusion of Embodied Reasoning in Vision-Language-Action Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent Vision-Language-Action (VLA) models report im-pressive success rates on robotic benchmarks, but whether these scoresreflect genuine embodied reasoning remains questionable. To addressthis gap, we introduce BeTTER, a diagnostic Benchmark for TestingTrue Embodied Reasoning. |
Haiweng Xu; Sipeng Zheng; Hao Luo; Wanpeng Zhang; Zongqing Lu; | code |
| 182 | OmniForcing: Unleashing Real-time Joint Audio-Visual Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent joint audio-visual diffusion models achieve remark-able generation quality but suffer from high latency due to their bidirec-tional attention dependencies, hindering real-time applications. |
Yaofeng Su; Yuming Li; Zeyue Xue; Jie Huang; Siming Fu; Haoran Li; Haoyang Huang; Nan Duan; | code |
| 183 | Pseudo-Stereo Inputs: A Solution to The Occlusion Challenge in Self-Supervised Stereo Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods attempt to address this by focusing on the erroneous feedback from the other side, either by identifying and removing it, or by introducing additional regularities for correction on that basis. Nevertheless, these approaches have failed to provide a complete solution. |
Ruizhi Yang; Xingqiang Li; Jiajun Bai; Jinsong Du; | code |
| 184 | PAI-Studio: Cinematic Video Background Replacement with Camera-Aware Motion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present PAI-Studio, a new reference-conditioned videosynthesis task that addresses a long-standing challenge in cinematic back-ground replacement: generating dynamic backgrounds aligned with fore-ground motion while preserving foreground identity, matching referencescene appearance, and achieving globally consistent illumination with re-alistic foreground relighting. |
Heyuan Gao; Bangxun Tang; Yiren Song; Guian Fang; Zijian He; Jie Yang; Mike Zheng Shou; | code |
| 185 | Boxer: Robust Lifting of Open-World 2D Bounding Boxes to 3D Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this researchchallenge, we propose Boxer, an algorithm to estimate static 3D boundingboxes (3DBBs) from 2D open-vocabulary object detections, posed images,and optional depth represented either as a sparse point cloud or densedepth. |
Daniel DeTone; Tianwei Shen; Fan Zhang; Lingni Ma; Julian Straub; Richard Newcombe; Jakob Engel; | code |
| 186 | Towards Temporal Compositional Reasoning in Long-Form Sports Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Building on SportsTime, we propose Chain-of-Time Reasoning(CoTR), which treats reasoning as a process of temporally groundedevidence composition. |
Siyu Cao; Lu Zhang; Ruizhe Zeng; Zhi-yong Liu; | code |
| 187 | Memory Tree Guided Key Frame Querying for Efficient 3D Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, wepropose a memory tree guided key frame selection paradigm for efficient3D question answering in embodied scenarios. |
Hsiang-Wei Huang; Fu-Chen Chen; Li-Wu Tsao; Cheng-Han Lee; Che-Chun Su; Lu Xia; Ronghui Peng; Jenq-Neng Hwang; Min Sun; Cheng-Hao Kuo; | code |
| 188 | Don’t Starve The Boundaries: Boundary-Constrained Label Propagation for Weakly Supervised 3D Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a novel 2D-assisted pseudo-label propaga-tion paradigm that does not rely on the model’s own predictions or anyexternal foundation models, yet is able to generate high-purity pseudo-labels. |
Shuwei Wu; shuo jin; Zhijin He; Siyue Yu; ENG LIM; Qiufeng Wang; Jimin Xiao; | code |
| 189 | RoMan-4D: Learning Robot Arm Manipulation from 4D World Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As a result, the generated videos appear plausi-ble, yet lack the physical grounding required for reliable action execution,such as robot manipulation. We present GEM-4D, a geometry-groundedvideo world model that resolves this limitation by injecting dense 4Dcorrespondence supervision distilled from a pretrained geometry foun-dation model into the video generative backbone during training. |
Kaichen Zhou; Yuzhen Chen; Fangneng Zhan; Hang Hua; Grace Chen; Xinhai Chang; Ao Qu; Yilun Du; Zhuang Liu; Paul Pu Liang; Mengyu Wang; | code |
| 190 | Are GUI Agents Focused Enough? Automated Distraction Via Semantic-level UI Element Injection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To study robustness undera more practical threat model, we propose Semantic-level UI ElementInjection, a black-box red-teaming paradigm that overlays safety-alignedand harmless UI elements onto screenshots to misdirect the agent’s visualgrounding. |
Wenkui Yang; chao jin; Haisu Zhu; Weilin Luo; Derek Yuen; Kun Shao; Junxian Duan; Huaibo Huang; Jie Cao; Ran He; | code |
| 191 | DART: Difficulty-Adaptive Routing for Zero-Shot Video Temporal Grounding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods rely on frame-query feature matching, whichsuffices for simple events but struggles with complex multi-stage queriesthat require understanding temporal ordering and causal structure—a disparity we call the reasoning gap. We propose DART (Difficulty-Adaptive Routing for Temporal Grounding), which bridges this gapby coupling difficulty-aware routing with structured reasoning in largevision-language models. |
zhengbo zhang; Mark H. Huang; Zhigang Tu; Ming-Hsuan Yang; | code |
| 192 | Towards Video Anomaly Detection from Event Streams: A Baseline and Benchmark Datasets Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Event-based vision, characterized by low redundancy, focuson dynamic motion, and inherent privacy-preserving properties, natu-rally fits the demands of video anomaly detection … |
Peng Wu; Yuting Yan; Guansong Pang; Yujia Sun; Qingsen Yan; Peng Wang; Yanning Zhang; | code |
| 193 | Generative Lane Topology Reasoning Via Autoregressive Model with Geometry Prior Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address these issues, we propose TopoGPT,a generative framework that learns the geometry prior from typical lanegraph structures through autoregressive sequence modeling.Specifically,we construct a large-scale map dataset comprising 3.3M scenes. |
Jiahui Fu; Zehao Huang; Han Li; Naiyan Wang; Si Liu; | code |
| 194 | BrepLLM: Enabling Large Language Models to Understand Boundary Representations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To thisend, we propose BrepLLM, the first multimodal framework that enablesLLMs to directly parse and reason over raw B-rep data.Furthermore, we construct the Brep2Text dataset,which contains 269,444 B-rep and text question-answer pairs. |
Liyuan Deng; Hao Guo; Yongkang Dai; Yunpeng Bai; Yifan Zhu; Yuanyuan Gao; Huaxi Huang; Yilei Shi; | code |
| 195 | EgoGVAE: Ego-body Mesh Reconstruction Via Guided Variational Autoencoder Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As an alternative, we propose a simpleyet novel method that leverages the latent space of the guidance network,which is designed as a variational autoencoder taking full-body poses asinputs. |
Jaehun Jung; Wonjun Kim; | code |
| 196 | Scaling Multi-Reference Image Generation with Dynamic Reward Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Evaluationson OmniRef-Bench show that mainstream open-source models strugglein complex MRIG scenarios, and their performance deteriorates signifi-cantly as the number of mixed-type reference images increases. To ad-dress this issue, we propose DyRef, a two-stage training framework.In the first stage, supervised fine-tuning equips the model with the ba-sic capability to handle complex MRIG tasks. |
Wenwang Huang; Yusen Fu; Mengfei Huang; Junjie Wang; Yulin Li; Gan Liu; Jing Cai; Yancheng He; Zhuotao Tian; | code |
| 197 | FineEdit: Fine-Grained Image Edit with Bounding Box Guidance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This approach ensures that the diffusion model can accurately lo-calize the target while preserving background consistency. To achievethis, we propose FineEdit, a multi-level bounding box injection methodthat enables the model to utilize spatial conditions more effectively.To support this high precision guidance, we present FineEdit-1.2M,a large scale, fine-grained dataset comprising 1.2 million image edit-ing pairs with precise bounding box annotations. |
Haohang Xu; Lin Liu; Zhibo Zhang; Rong Cong; Xiaopeng Zhang; Qi Tian; | code |
| 198 | GR-GRPO: Graph-Diffused Credit for Autoregressive Image RL Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This chal-lenge is amplified by vector-quantized (VQ) tokenization, where nearbycodebook IDs in embedding space can be locally substitutable; however,GRPO-style updates provide explicit positive credit only to sampled IDs,leading to insufficient coverage of success-supported alternatives undersmall-G training. We propose GR-GRPO (Graph-Regularized GRPO),which densifies token-level supervision without additional rollouts by dif-fusing positive evidence from positive rollouts over a precomputed code-book K-NN graph to construct graph-diffused soft targets, and regulariz-ing the policy toward these targets via an auxiliary cross-entropy term. |
Zheyu Zhang; Peng-Tao Jiang; Tianyi Zheng; Jian Zhang; Jinwei Chen; Bo Li; | code |
| 199 | SCORE: SubDistribution-aware Collaborative Knowledge Reinforcing for Cloth-Hybrid Lifelong Person Re-Identification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Due to theconflict between clothing-relevant and clothing-irrelevant knowledge, thewell-known catastrophic forgetting problem is significantly exacerbatedin this task. To address this issue, we propose a SubDistribution-awareCOllaborative Knowledge REinforcing (SCORE) framework, where ourkey idea is explicitly modeling the intra-identity diversity to continu-ally consolidate distinct cloth-consistent and cloth-changing knowledge.Specifically, an Adaptive SubDistribution Modeling mechanism is devel-oped, where a set of distributional subprototypes is assigned to eachidentity to capture the intra-identity diversity, improving the compati-bility between cloth-consistent and cloth-changing knowledge. |
Kunlun Xu; Liangyu Ma; Jiangmeng Li; Xin Tong; Xiaode Liu; Yufei Guo; Jiahuan Zhou; | code |
| 200 | Fast and Accurate Image Restoration with Rank Enhanced Linear Attention Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite its efficiency benefits, vanilla linear attention suffers from a significant performance drop in IR, largely due to the low-rank nature of its attention map. To counter this, we propose Rank Enhanced Linear Attention (RELA), a simple yet effective method that enriches feature representations by integrating a lightweight depthwise convolution. |
Yuang Ai; | code |
| 201 | Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we build a benchmark,Diagram-MMU, a multi-modal benchmark designed to assess MLLMs’ability for scienti_x001C_c diagram parsing and understanding. |
Weihao Bo; Shan Zhang; Yanpeng Sun; Jie Liu; Yongke Yao; Jinhao Du; Wei He; KAI ZOU; Zechao Li; Jingdong Wang; | code |
| 202 | SeekFlow: Synergizing Radiology and Pathology Foundation Models for Precision Oncology Via Knowledge-Guided Evidence Flow Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The fea-ture embeddings from heterogeneous foundation models reside on poorlyaligned latent manifolds, and naive static fusion may therefore mix unre-lated semantic neighborhoods, leading to topological degeneracy and se-mantic entanglement. To bridge this gap, we propose a novel multimodallearning framework for Synthesizing radiology and pathology EvidencEvia Knowledge-guided FLOW matching. |
Peixiang Huang; Yanyan Huang; Yihang Chen; Maximus Yeung; Yuming Jiang; Lequan Yu; | code |
| 203 | Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose an effective, training-free framework that uses an MLLM’s intrinsic uncertainty as proactive guidance. |
Sanghwan Kim; Rui Xiao; Stephan Alaniz; Yongqin Xian; Zeynep Akata; | code |
| 204 | Self-Evolving Just-In-Time Memory for Proactive Embodied Safety Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing safety approaches often rely on runtime guardrailsto block unsafe actions or induce excessive caution, which severely stallstask progress instead of actively resolving the underlying risks. To breakthis safety–progress trade-off, we introduce the Self-Evolving Just-In-Time Memory framework, which reframes embodied safety from progress-stalling guardrails to proactive hazard mitigation. |
Bingrui Sima; Lizhong Wang; Xiaoya Lu; Kun He; Xiao Yang; | code |
| 205 | EditHF-1M: A Million-Scale Rich Human Preference Feedback for Image Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Recent text-guided image editing (TIE) models have achievedremarkable progress, while many edited images still suffer from issuessuch as artifacts, unexpected editings, … |
Zitong Xu; Huiyu Duan; Zhongpeng Ji; Xinyun Zhang; Yutao Liu; Xiongkuo Min; Ke Gu; Jian Zhang; Shusong Xu; Jinwei Chen; Bo Li; Guangtao Zhai; | code |
| 206 | The Map Is Not The Territory: Embedding-Coverage Blacklists for Safe Diffusion Steering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Ensuring safe content generation in text-to-image diffusionmodels remains a critical challenge. Existing safety mechanisms focuson model editing or trajectory steering, yet the … |
Juyang Bai; Tong Zhou; Shaolei Ren; Xiaolin Xu; | code |
| 207 | MMLoP: Multi-Modal Low-Rank Prompting for Efficient Vision-Language Adaptation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose MMLoP (Multi-Modal Low-RankPrompting), a framework that achieves deep multi-modal promptingwith only 11.5K trainable parameters, comparable to early text-onlymethods like CoOp. |
Sajjad Ghiasvand; Haniyeh Oskouie; Mahnoosh Alizadeh; Ramtin Pedarsani; | code |
| 208 | Markov-Renewal Single-Photon LiDAR Simulator Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a sim-ulator that achieves both fidelity and speed by focusing on the critical,yet overlooked, component of simulation: the photon count statistics. |
Weijian Zhang; PRATEEK CHENNURI; Hashan Weerasooriya; Bole Ma; Stanley Chan; | code |
| 209 | CogniCred: A Dataset and Benchmark for Cognitive Credential Forgery Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While the devel-opment of Multimodal Large Language Models (MLLMs) presents newopportunities for complex credential forgery detection, the scarcity ofdatasets and benchmarks severely impedes progress in this critical do-main. To bridge this gap, we introduce CogniCred — the first large-scale,multimodal dataset built upon complex, real-world credentials with care-fully crafted cognitive-level forgeries. |
Junchi Li; Jiasheng Sun; Weizhi Chen; Ziwei Wang; Sheng Zhou; Jiajun Bu; Chenfan Qu; Bohan Yu; Jian liu; Weiqiang Wang; | code |
| 210 | Color Pass-Through Via Camera-Display Coupling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This gap per-sists despite substantial advances in both modern cameras and displays.A key reason is that most pipelines factor the high-dimensional capture-to-display process into two separately calibrated camera and displaystages, and then connect them through low-dimensional color transforms,leading to information bottlenecks and inevitable error accumulation. Toaddress this systemic challenge, we propose Color Pass-Through, anend-to-end learned framework that operates directly on captured images.Our key insight is to treat the camera and display as a coupled systemrather than calibrating them in isolation. |
Ruikang Li; Molin Li; Jiarui Wu; Zhe Wei; Pengpeng Liu; Tianfan Xue; | code |
| 211 | From Local Geometry to Global Pseudo-Labeling for Robust Positive–Unlabeled Learning Under Covariate Shift Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we show that covariate shift detection can beeffectively addressed with weaker supervision using Positive–Unlabeled(PU) learning. |
Firas Gabetni; Alexandre Rocchi–Henry; Ziyi LIU; Nacim Belkhir; Gianni Franchi; | code |
| 212 | Can Vision Models Truly Forget? Mirage: Representation-Level Certification of Visual Unlearning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Machine unlearning in Vertical Federated Learning (VFL)has attracted growing interest, yet existing methods certify forgettingsolely using output-level metrics. We challenge these … |
Zhenyu Yu; yangchen zeng; Chunlei Meng; Guangzhen Yao; Shuigeng Zhou; | code |
| 213 | ODONet: Online Dynamic Offset Network for Visual Object Tracking Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Offset learning has recently demonstrated remarkable prowessin capturing the structural diversity of objects. Consequently, it ap-pears to be a natural fit for modeling the … |
Qinghua Liu; Wanli Xue; Shengyong Chen; | code |
| 214 | UEval: A Benchmark for Unified Multimodal Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce UEval, a benchmark to evaluate unified mod-els, i.e., models capable of generating both images and text. |
Bo Li; Yida Yin; Wenhao Chai; Xingyu Fu; Zhuang Liu; | code |
| 215 | REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Latent diffusion models (LDMs) achieve state-of-the-art im-age synthesis, yet their reconstruction-style denoising objective providesonly indirect semantic supervision: high-level … |
Giorgos Petsangourakis; Christos Sgouropoulos; Bill Psomas; Theodoros Giannakopoulos; Giorgos Sfikas; Ioannis Kakogeorgiou; | code |
| 216 | Rolling Shutter Relative Pose Estimation Made Practical Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We make RS relative pose estimation practical by in-troducing affine correspondences (ACs) into the RS two-view geome-try. |
Daniel Barath; | code |
| 217 | NutriBench-Kitchen: Benchmarking Embodied AI for Nutrition Management Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing benchmarks evaluate static food under-standing or embodied cooking actions, but do not measure whether anagent can continuously update and use nutrition-relevant states in dy-namic kitchens. To fill this gap, we introduce NutriBench-Kitchen,a benchmark containing 1,500 manually verified question–answer pairsfrom 160 cooking videos. |
YuLin Wei; Xiangchen Wang; Jianhui Pan; Jinyu Xiao; Zheng Tan; Ruozai Tian; Guanhua Chen; Feng Zheng; | code |
| 218 | CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we proposeCoLT (Chain of Latent Thoughts), a novel framework that teachesmulti-modal models to reason through a chain of latent thought repre-sentations instead of verbose text tokens, which can perform thinkingwith as few as 3 steps. |
Lianyu Hu; shengqian qin; Zeqin Liao; Qing Guo; Liang Wan; Wei Feng; Yang Liu; | code |
| 219 | CoCo-IR: Conversational Composed Image Retrieval Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To overcome this limitation, we introduce Contextual Composed Image Retrieval (CoCo-IR), a novel task that enables users to progressively refine search results through interactions. We address this new task by proposing a new model based on a Large Multimodal Model (LMM) that functions as a context-aware reasoner for CoCo-IR. |
Shengcao Cao; Tanmaya Dabral; Zhongli Ding; Madhuri Shanbhogue; Kaifeng Chen; Zhe Li; Mojtaba Seyedhosseini; Liang-Yan Gui; Yu-Xiong Wang; | code |
| 220 | ShellMaker: Language-Guided Exterior Completion Under Structural Constraints Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Despite advances in indoor scene generation, synthesizingcoherent building exteriors consistent with generated interiors remainslargely unexplored. Existing methods can generate … |
Ruiqi Xu; Daniel Aliaga; | code |
| 221 | RePer-360: Releasing Perspective Priors for 360° Depth Estimation Via Self-Modulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent depth foundation models trained on perspective im-agery achieve strong performance, yet generalize poorly to 360∘ imagesdue to the substantial geometric discrepancy between perspective andpanoramic domains. |
Cheng Guan; Chunyu Lin; Zhijie Shen; Junsong Zhang; Jiyuan Wang; | code |
| 222 | Hierarchical and Holistic Open-Vocabulary Functional 3D Scene Graphs for Indoor Spaces Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, their potential remains underexplored due to the limitedcoverage of existing benchmarks and the overly straightforward designof previous pipelines, which primarily focus on large-scale furniture butlack of hierarchical structures. Therefore, in this work, we extend thebenchmark coverage by introducing dense tabletop objects and explicitmulti-level functional relationships. |
Xinggang Hu; Chenyangguang Zhang; Alexandros Delitzas; Xiangkui Zhang; Marc Pollefeys; Francis Engelmann; Xiangyang Ji; | code |
| 223 | Text-based Tactile Graphics Generation for The Visually Impaired Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present the first integratedgenerative system that produces fabrication-ready 2.5D tactile graphicsdirectly from natural language prompts, jointly generating global basegeometry, fine-grained tactile surface textures, and standard-compliantbraille within a unified 3D-printable representation. |
Ruihan Gao; Joonghyuk Shin; Ava Pun; Jaesik Park; Wenzhen Yuan; Jun-Yan Zhu; | code |
| 224 | Verifying Cancer Segmentation in Vision Transformers Via Internal Concepts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Unlike output-level cues (e.g., prediction confidence or uncertainty), which offer no insight into why a failure occurs and suffer from a sensitivity–quality tradeoff where high detection sensitivity could degrade overall segmentation quality. We instead propose to capture the model’s FOE from its inner workings. |
Mengmeng Ma; Yunxiang Peng; Tang Li; Lu Lin; Binsheng Zhao; Oguz Akin; Xi Peng; | code |
| 225 | Discrete Diffusion Bridges for Spatiotemporally Aligned Image Translation and Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Discrete Diffusion Bridges (DDB), a novel frame-work designed to resolve the fundamental spatiotemporal misalignmentof standard discrete diffusion in image translation and generation. |
Xing Xie; Jiawei Liu; Shijun Zhou; Huijie Fan; Zhi Han; Yandong Tang; Liangqiong Qu; | code |
| 226 | SFDATrack: Generalized Source-Free Domain Adaptive Tracking Under Adverse Weather Conditions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despitethe impressive performance, existing methods heavily rely on the large-scale video frames from both source and target domains, which is im-practical under rigid resource constraints where source data is unavail-able. To overcome this limitation, we propose SFDATrack, a general-ized source-free domain adaptive tracker that merely leverages adverseweather samples from the target domain for robust state estimation.Specifically, SFDATrack first employs a mean-teacher backbone withDual Interactive Mamba (DIM) blocks to distill the candidate targettokens that are resilient to weather variations from classified, augmentedsamples. |
Siyuan Yao; Ziqi Wang; Junqi Huang; Ruiqi Yu; Wenqi Ren; Xiaochun Cao; | code |
| 227 | Social-Mamba: Socially-Aware Trajectory Forecasting with State-Space Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recently, Selective State-Space Models haveprovided a linear-time alternative; however, their inherently sequentialdesign is misaligned with the unstructured and dynamic nature of socialinteractions. To address this challenge, we propose Social-Mamba, a fore-casting architecture that reformulates social interactions as structuredsequential processes. |
Po-Chien Luan; Wuyang Li; Yang Gao; Alexandre ALahi; | code |
| 228 | Enhancing Pretrained Model-based Continual Representation Learning Via Guided Random Projection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end,we propose the Stochastic Continual Learner with MemoryGuard Super-visory Mechanism (SCL-MGSM). |
Ruilin Li; Heming Zou; Xiufeng Yan; Zheming Liang; Jie Yang; Chenliang Li; Xue Yang; | code |
| 229 | RoomPlanner: Reachability-Aware View Sampling for Text-to-Room 3D Gaussian Splatting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose RoomPlanner, the first fully au-tomatic 3D room generation framework for painlessly creating realisticindoor scenes with only short text as input. |
Wenzhuo Sun; MingJian Liang; Wenxuan Song; Xuelian Cheng; Zongyuan Ge; | code |
| 230 | Hierarchical Hyperbolic Representation Learning for Aerial-Ground Person Re-Identification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although great progress, existing methods remain subopti-mal due to the direct feature alignment across views, overlooking view-specific cues. To address this issue, we propose a novel Hierarchical Hy-perbolic Representation (HiHR) framework for AG-ReID. |
QiWei Yang; | code |
| 231 | HSFM: Hard-Set-Guided Feature-Space Meta-Learning for Robust Classification Under Spurious Correlations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In particular,retraining a lightweight head while keeping the backbone frozen cansubstantially improve performance on shifted distributions and minor-ity groups. Motivated by this observation, we propose a bilevel meta-learning method that performs augmentation directly in feature space toimprove spurious correlation handling in the classifier head. |
Aryan Yazdan Parast; Khawar Islam; Soyoun Won; Basim Azam; NAVEED AKHTAR; | code |
| 232 | Fully Rotation-Equivariant Spectral-Spatial Learning for Multispectral Object Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing multispectral detectors are limited by discrete spec-tral processing, a scale-dependent shift in the relative reliability of spec-tral and spatial cues across pyramid levels, and the lack of explicitrotation-equivariant geometric priors for arbitrarily oriented objects. Totackle these limitations, we propose FressDet, a fully rotation-equivariantspectral_x0015_spatial learning framework for multispectral object detection,capable of capturing the continuous, ordered nature of spectral struc-ture and enabling reliable spectral_x0015_spatial fusion across pyramid levelsunder arbitrary in-plane rotations. |
Peng Zhang; Tingfa Xu; Shuaihao Han; Jianan Li; | code |
| 233 | CGCC: Towards Generalizable Clothes-Changing Person Re-Identification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, current research is severely constrained bytwo main shortcomings: existing datasets lack comprehensive diversityand semantic annotations, and current methods fail to globally modeland eliminate identity-irrelevant interfering factors, severely limiting CC-ReID generalization in real-world scenarios. To address these limitations,we propose a unified framework for generalizable CC-ReID. |
Yizhi Wu; Fangyi Liu; wei yu; Mang Ye; | code |
| 234 | RAGA: Real Time Ray Traced Gaussian Shadow Casting for 3DGS Avatar-Scene Interaction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study the problem of physically plausible shadow castingwhen animating 3D Gaussian Splatting (3DGS) avatars, either individu-ally or in multi-avatar and object-interaction scenarios, within existing3DGS scenes. |
Aymen Mir; Riza Alp Guler; Jian Wang; Peter Wonka; Bing Zhou; Gerard Pons-Moll; | code |
| 235 | AHOY! Animatable Humans Under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present AHOY, a method for reconstructing complete,animatable 3D Gaussian avatars from in-the-wild monocular video de-spite heavy occlusion. |
Aymen Mir; Riza Alp Guler; Xiangjun Tang; Peter Wonka; Gerard Pons-Moll; | code |
| 236 | TextDS: Parameter-Efficient Representation Alignment for Scene Text Detection Under Distribution Shifts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose TextDS, an efficient framework for scene text detec-tion under distribution shifts. |
Boyuan Chen; Zichen Dang; Chuang Yang; Lap-Pui Chau; Yi Wang; | code |
| 237 | FD²: A Dedicated Framework for Fine-Grained Dataset Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing decoupled methods largelyrely on coarse class-label supervision and optimize samples within eachclass in a nearly identical manner. On fine-grained datasets, this oftenyields distilled samples that (i) retain large intra-class variation withsubtle inter-class differences and (ii) become overly similar within thesame class, limiting localized discriminative cues and hurting recognition.To solve the above-mentioned problems, we propose FD2 , a dedicatedframework for Fine-grained Dataset Distillation. |
Hongxu Ma; Guang Li; Shijie Wang; DONGZHAN ZHOU; Baoli Sun; Takahiro Ogawa; Miki Haseyama; Zhihui Wang; | code |
| 238 | SRRA: Stable-Rank-Based Residual Adaptation for Generalizable Deepfake Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, the shallow layers exhibit a distinctly low-rank struc-ture, whereas the middle and deep layers present more complex struc-tures. Therefore, we argue that only slight fine-tuning is required forthe shallow layers, which are able to effectively capture low-level forgerycues, while the middle and deep layers require higher residual ranks tomodel long-range dependencies and extract high-level forgery artifacts.To achieve these goals adaptively, we propose the Stable-Rank-BasedResidual Adaptation (SRRA) strategy. |
Huakun Liu; Wenjie Li; Changsheng Xu; | code |
| 239 | EgoPHI: Estimating 3D Hand-Object Contact and Force from Egocentric Vision Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present EgoPHI, the first method that jointly esti-mates dense contact maps and 3D force distributions on hand and objectmeshes from a single egocentric RGB image and object geometry. |
Andela Ilic; Rachel Schuchert; Yijing Jiang; Christian Holz; | code |
| 240 | UHD-MFF: Shattering Barriers in Multi-Focus Ultra-High-Definition Image Fusion Via Learnable Lookup Tables Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing multi-focus image fusion remains largely confined to low-resolution images and faces three major barriers in UHD scenarios, namely data availability, model adaptability, and deployment feasibility, which severely hinder its practical application. To shatter these barriers, first, we propose the UHD-MFF dataset, the first largescale ultra-high-resolution multi-focus fusion dataset. |
Yibing Zhang; Xunpeng Yi; Qinglong Yan; Yeda Wang; Han Xu; Jiayi Ma; | code |
| 241 | Enlightening Photographic Style Transfer with A Self-Supervised Photographic Embedding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose PETAL, a PhotographicEmbedding for Transfer with an Adaptive LUT. |
Chengxuan Zhu; Jiacong Fang; Shuchen Weng; Youwei Lyu; Jiajun Tang; Qingnan Fan; Chao Xu; Boxin Shi; | code |
| 242 | Structure Gaussian Splatting SLAM Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a structure-aware GS-SLAM framework that models planar structures as persistentPlanar Gaussian Instances (PGIs) within a 3D Gaussian map. |
Yan Li; Yingzhao Li; Gim Hee Lee; | code |
| 243 | Region-Aware Multimodal Interleaving for Animal Re-Identification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Additionally, these methods are mainlyconstrained to the visual space only and lack pattern-based semanticinteractions from the multimodal space. To address these challenges,we introduce a Region-Aware Multimodal Interleaving (RAMI) frame-work, which formulates Animal ReID as a visual-textual semantic in-teraction over informative, localised regions. |
Yihao Wu; Di Zhao; Wayne Getz; Lingqiao Liu; Gillian Dobbie; Daniel Wilson; Yun Sing Koh; | code |
| 244 | Beyond Inpainting: Unleash 3D Understanding for Stable Camera-Controlled Video Re-rendering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this problem, we propose DepthDirector, a video re-rendering framework with stable camera controllability.Additionally, we construct a large-scale multi-camera synchronized dataset named MultiCamWarp Dataset using Unreal Engine 5. |
Dongyu Chen; Yixin Guo; Shuojin Yang; Tai-Jiang Mu; Shimin Hu; | code |
| 245 | DEX-AR: A Dynamic Explainability Method for Autoregressive Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As Vision-Language Models (VLMs) become increasingly so-phisticated and widely used, it becomes more and more crucial to un-derstand their decision-making process. |
Walid Bousselham; Angie Boggust; Hendrik Strobelt; Hilde Kuehne; | code |
| 246 | CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we pro-pose a novel flow matching-based movie dubbing framework driven bythe Cognitive Synchronous Diffusion Transformer (CoSyncDiT), in-spired by the cognitive process of professional actors. |
Gaoxiang Cong; Liang Li; Jiaxin Ye; Zhedong Zhang; Hongming Shan; Yuankai Qi; Qingming Huang; | code |
| 247 | Render-FM: Feedforward Model for Real-time Photorealistic Volumetric Rendering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Photorealistic volumetric rendering of CT scans greatly ben-e_x001C_ts clinical work_x001D_ows, yet neural approaches such as Neural RadianceFields (NeRF) and 3D Gaussian Splatting (3DGS) require prohibitiveper-scan optimization (hours for NeRF, about 30 minutes for 3DGS),making them impractical in clinical settings. We propose Render-FM,a feedforward model that eliminates this bottleneck by directly regress-ing 6D Gaussian Splatting (6DGS) parameters from a CT volume ina single 2.8-second forward pass, a 500× speedup over per-scan opti-mization. |
Zhongpai Gao; Benjamin Planche; Meng Zheng; Anwesa Choudhuri; Van Nguyen; Terrence Chen; Ziyan Wu; | code |
| 248 | ClusterStyle: Modeling Intra-Style Diversity with Prototypical Clustering for Stylized Motion Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, capturing intra-style diversity, where a single style should correspond to diverse motionvariations, remains a significant challenge. In this paper, we propose aclustering-based framework, ClusterStyle, to address this limitation.Instead of learning an unstructured embedding from each style motion,we leverage a set of prototypes to effectively model diverse style pat-terns across motions belonging to the same style category. |
Kerui Chen; Jianrong Zhang; Ming Li; Zhonglong Zheng; Hehe Fan; | code |
| 249 | VoxAnchor: Explicit Voxel-Semantic Grounding for Spatial Understanding in Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Unlike humans, who can natu-rally infer depth, distance, object size, and their own movement throughspace, these models often lack the ability to accurately reconstruct a co-herent 3D environment. This limitation makes it challenging for them toestimate real-world measurements, such as how far an object is or howmuch the camera has moved, when relying only on 2D image sequences.To bridge these gaps, we propose VoxAnchor, a framework for explicitspatial-semantic grounding. |
Xinglin Li; TingTing Long; Jingzhi Zhou; Chuxuan Zeng; Jiajing Chen; Jian Yang; Jin Xie; | code |
| 250 | T-REN: Learning Text-Aligned Region Tokens Improves Dense Vision-Language Alignment and Scalability Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose T-REN (Text-aligned Region Encoder Net-work), an efficient encoder that maps visual data to a compact setof text-aligned region-level representations (or region tokens). |
Savya Khosla; Sethuraman T V; Aryan Chadha; Alex Schwing; Derek Hoiem; | code |
| 251 | SafeGuard: A Multi-Agent Perception-Reasoning Framework for Social-Risk AI-Generated Video Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing benchmarks are biased towardperceptual fidelity and primarily evaluate detectors based on perceptualartifacts, providing limited coverage of scenarios that require reasoningabout violations of physical laws, structural coherence, or social logic.This dataset bias shapes current approaches and results in a Percep-tion–Reasoning Gap: artifact-centric models capture low-level statisticalirregularities yet lack semantic inference, whereas vision-language mod-els perform semantic reasoning but remain insensitive to fine-grainedforensic cues. To bridge this gap, we propose SafeGuard, a multi-agentframework that enables collaborative specialization between forensic per-ception and semantic reasoning. |
Wenlin Wu; Sheng Zhou; Peipei Song; Wenhao Wang; Junbin Xiao; Xun Yang; | code |
| 252 | DDStereo: Efficient Dual Decoder Transformers for Stereo 3D Road Anomaly Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present DDStereo, a novelDual-Decoder Stereo Transformer that achieves a synergistic integrationof 3D object detection and Out-of-Distribution (OoD) road anomaly de-tection. |
Shiyi Mu; Zichong Gu; Zhiqi Ai; Yilin Gao; Shugong Xu; | code |
| 253 | Reinforcement Learning for Multimodal Diffusion Language Models Via Bidimensional Trajectory and Thought Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Bi-VRL, a reinforce-ment learning framework specifically designed for multimodal dLLMs,featuring a bidimensional optimization strategy that enables com-prehensive optimization across both reasoning and diffusion timestepdimensions. |
Bowen Li; Yinjie Wang; Yunzhi Zhang; Junhong Liu; Yingqing Guo; Jiajun Wu; Mengdi Wang; Ling Yang; | code |
| 254 | LaGen: Towards Autoregressive LiDAR Scene Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Unlike the widely stud-ied image modality, in this work we explore generative world modelsfor LiDAR data. |
Sizhuo Zhou; Xiaosong Jia; Fanrui Zhang; Junjie Li; Juyong Zhang; Yukang Feng; Jianwen Sun; Songbur Wong; Junqi You; Junchi Yan; | code |
| 255 | SALT: Self-Consistent Distribution Matching with Cache-Aware Training for Few-Step Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Distribution matching distilla-tion (DMD) can recover sharp, mode-seeking samples, but its local train-ing signals do not explicitly regularize how denoising updates composeacross timesteps, making composed rollouts prone to drift. To overcomethis challenge, we propose Self-Consistent Distribution Matching Distil-lation (SC-DMD), which explicitly regularizes the endpoint-consistentcomposition of consecutive denoising updates. |
Xingtong Ge; Yi ZHANG; Yushi Huang; Dailan He; Xiahong Wang; Bingqi Ma; Guanglu Song; Yu Liu; Jun Zhang; | code |
| 256 | LoMa: Local Feature Matching Revisited Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we revisit local feature matching from a datadriven perspective. |
David Nordström; Johan Edstedt; Georg Bökman; Jonathan Astermark; Anders Heyden; Viktor Larsson; Mårten Wadenbäck; Michael Felsberg; Fredrik Kahl; | code |
| 257 | XYZ-IBD: Benchmarking Robust 6D Object Pose Estimation Under Real-World Industrial Complexity Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce XYZ-IBD, a high-precision benchmark for object detection and 6D pose estimation specifically designed for industrial bin-picking. |
Junwen Huang; Jiaqi Hu; Peter Yu; Slobodan Ilic; Martin Sundermeyer; Benjamin Busam; | code |
| 258 | MedCAGD: Context-Aware Gated Decoder for Robust Medical Image Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inthis work, we revisit medical image segmentation from a decoder-centric per-spective and propose a context-aware gated decoder that systematically regulatesfeature fusion and contextual aggregation throughout the decoding process. |
Saad Wazir; Patrick Vibild; Dinh Tran; Seongah Kim; Daeyoung Kim; | code |
| 259 | 3D-Aware VLMs with Implicit and Explicit Geometries Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reasoning. To bridge this gap, we present VLM-IE3D, a unified framework that enhances the 3D spatial awareness of VLMs by equipping them with both implicit and explicit 3D geometries learned from RGB videos. |
Wenhao Li; Xueying Jiang; Quanhao Qian; Deli Zhao; Ran Xu; Shijian Lu; Gongjie Zhang; | code |
| 260 | Parallel Vision Token Scheduling for Fast and Accurate Multimodal LMMs Inference Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Naively pruning less-informative visual tokens reduces this burden, yet indiscriminate removal can erase contextual cues essential for background or fine-grained questions, undermining accuracy. In this paper, we present ParVTS (Parallel Vision Token Scheduling), a training-free scheduling framework that partitions visual tokens into subject and non-subject groups, processes them in parallel to transfer their semantics into question tokens, and discards the non-subject path mid-inference to reduce computation. |
Wengyi Zhan; Mingbao Lin; Zhihang Lin; Rongrong Ji; | code |
| 261 | Dotting The Eye: An Intent-Driven Image Retouching Agent for Visual Focus Enhancement Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this study, we propose EyeControl, a multi-modal large lan-guage model (MLLM)-driven agent with a diffusion-based retouchingexecutor that enables visual focus enhancement under weak user intent.With only a few clicks or coarse strokes, EyeControl directs visual at-tention to the intended region, effectively “dotting the eye” of the image.The core idea is to explicitly link the weak user intention with the tar-get editing region and the corresponding tonal adjustment operationsduring retouching. |
Chujie Qin; Zilong Zhang; Zewei Chang; Chun-Le Guo; Ruixing Wang; Tao Hu; Ming-Ming Cheng; Chongyi Li; | code |
| 262 | ECHO: Efficient Chest X-ray Report Generation with One-step Block Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Compressing multi-step denoising to a single step could further reduce latency, but often degrades textual coherence due to the mean-field bias introduced by token-factorized denoisers. To address this challenge, we propose ECHO, an efficient diffusion-based VLM (dVLM) for chest X-ray report generation. |
Lifeng Chen; tianqi you; Hao Liu; Zhimin Bao; Jile Jiao; Xiao Han; Zhicai Ou; Tao Sun; Mou XiaoFeng; Xiaojie Jin; Yi Xu; | code |
| 263 | Sparse-View Surface Reconstruction Using Gaussian Splatting Through High-Confidence Depth Propagation with Normal Priors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a novel 3DGS-basedmethod for high-fidelity surface reconstruction from sparse views. |
Liang Han; Bangcai Wei; Junsheng Zhou; Yushen Liu; Zhizhong Han; | code |
| 264 | City-Level 3D Surface Reconstruction with Viewpoint Orientation Partitioning and Scene Completion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a novel yet simple partitioning method to efficiently and faithfully reconstruct large-scale scene surfaces. |
Liang Han; Wenyuan Zhang; Junsheng Zhou; Yushen Liu; Zhizhong Han; | code |
| 265 | SHINE-PPG: Non-Lambertian Intrinsic Decomposition for Illumination-Robust RPPG Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose SHINE-PPG (Specular-Highlight In-trinsic Network for rPPG Estimation), a novel framework that leveragesnon-Lambertian intrinsic decomposition to decouple facial videos intoillumination, reflectance, and specular components in a self-supervisedmanner. |
Shih-Yu Yang; Yen-Chun Chou; Pei-Kai Huang; Chiou-Ting Hsu; | code |
| 266 | Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our key insight is to exploit the weaklocal dependencies across temporally distinct events to restructure thecausal dependency graph, thereby enabling lossless parallel generation.Specifically, tokens with weak cross-event dependencies can be decoded inparallel, while tightly coupled tokens within each event retain sequentialdecoding to preserve local semantic coherence. |
Wenzheng Zeng; Siyi Jiao; Chen Gao; Hwee Tou Ng; Mike Zheng Shou; | code |
| 267 | DualCamCtrl: Dual-Branch Diffusion Model for Geometry-Aware Camera-Controlled Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents DualCamCtrl, a novel architecture forcamera-controlled video generation that internalizes geometric reason-ing into the diffusion process. |
Hongfei Zhang; Kanghao Chen; Zixin Zhang; Harold Haodong Chen; Yuanhuiyi Lyu; Kun Zhou; Yuqi Zhang; Shuai Yang; Yingcong Chen; | code |
| 268 | Next-Frame Decoding for Ultra-Low-Bitrate Image Compression with Video Diffusion Priors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a novel paradigm for ultra-low-bitrate imagecompression (ULB-IC) that exploits the “temporal” evolution in gener-ative image compression. |
Yunuo Chen; Chuqin Zhou; Jiangchuan Li; Xiaoyue Ling; Bing He; Jincheng Dai; Li Song; Guo Lu; | code |
| 269 | ScenarioControl: Vision-language Controllable Vectorized Latent Scenario Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce ScenarioControl, the first vision-language con-trol mechanism for learned driving scenario generation. |
Lili Gao; Yanbo Xu; William Koch; Samuele Ruffino; Luke Rowe; Behdad Chalaki; Dmitriy Rivkin; Julian Ost; Roger Girgis; Mario Bijelic; Felix Heide; | code |
| 270 | SE-DETR: Explicit Semantic Exploration for Generalizability and Distinguishability in Video Temporal Grounding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a Semantics-Proxied Alignment (SPA) mod-ule that learns semantic proxies dynamically optimized during training,and enables concept-level alignment between multimodal representations.Furthermore, we design a Temporal Sparsity Modulation (TSM) mod-ule that estimates the temporal sparsity of semantic components anddynamically reweights word-guided visual features to highlight informa-tive concepts for accurate moment localization. |
Chengyang Hu; Guanshuo Wang; Fufu Yu; Qiong Jia; Shouhong Ding; Lizhuang Ma; | code |
| 271 | VIPS: Vehicle-Infrastructure Cooperative Planning Benchmark Via Pseudo-Simulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing evaluation protocols face afundamental trade-off: open-loop evaluation fails to capture error accu-mulation and recovery from deviations, while closed-loop evaluation iscostly, difficult to scale, and often relies on simulated environments thatmay suffer from domain gaps. To bridge this gap, we propose VIPS, abenchmark for cooperative autonomous driving in V2I settings based onpseudo-simulation. |
Hoonhee Cho; Jae-young Kang; Giwon Lee; Hyemin Yang; Heejun Park; KUK-JIN YOON; | code |
| 272 | Monte Carlo Energy Aggregation for Mobile 3D Gaussian Splatting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper,we present Flux-GS, a real-time Gaussian Splatting method designedto achieve high-fidelity rendering with significantly reduced overhead forresource-constrained mobile platforms. |
Xiaobiao Du; YuAn Wang; Hao Li; Bosheng Wang; Xun Sun; Xin Yu; | code |
| 273 | MVGS: Multi-view Regulated Gaussian Splatting for Novel View Synthesis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To tackle thisissue, we present a novel multi-view regulated Gaussian Splatting (MVGS)that fully leverages a multi-view coherent (MVC) constraint throughoutthe optimization process. |
Xiaobiao Du; Yida Wang; Xin Yu; | code |
| 274 | What If? Emulative Simulation with World Models for Situated Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce WanderDream, the first large-scale dataset designedfor the emulative simulation of mental exploration, enabling models toreason without active exploration. |
Ruiping Liu; Yufan Chen; Yuheng Zhang; Junwei Zheng; Kunyu Peng; Chengzhi Wu; Chenguang Huang; Di Wen; Jiaming Zhang; Kailun Yang; Rainer Stiefelhagen; | code |
| 275 | CoMind: Understanding Collaborative Human Activity from Multiple Minds and Views Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While much researchhas focused on modeling coordination and task execution, the cognitiveprocesses that support such collaboration, particularly Theory of Mind(the ability to infer others’ mental states), remain difficult to study innatural settings. To address this gap, we introduce a novel egocentric andexocentric video dataset capturing real-world collaboration in cookingscenarios. |
Alexey Gavryushin; Dingxi Zhang; Zhao Huang; Alexandros Delitzas; Jiaqi Chen; Ben Ellis; Cedric Zöllner; Manthan Patel; Manuel Kaufmann; Marc Pollefeys; Xi Wang; | code |
| 276 | RAF: Reliability-Aware Fusion of Camera, LiDAR, and 4D RADAR for Robust 3D Object Detection in Adverse Weather Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Reliability-Aware Fusion (RAF), which explicitly supervises per-pixel reliability estimation and provides a direct learning signal for identifying and suppressing unreliable visual cues. |
Heejun Park; Jaeseok Jeong; KUK-JIN YOON; | code |
| 277 | EditVerse3D: High-Quality 3D Object Editing with Region-Aware Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To bridge this gap, we propose EditVerse3D, a novel 3D editing framework that enables high-quality object editing under such coarse guidance.We construct a large-scale 3D editing dataset derived from parts information. |
Youtan Yin; Yanning Zhou; Jiacheng Wei; Xiaofeng Yang; Jun Zhang; Jiayang Bai; Jingwen Ye; Weidong Zhang; Guosheng Lin; | code |
| 278 | RPM-Distill: Physiology-guided Adaptive Cross-modal Distillation for Robust Remote Physiological Measurement Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Video-based remote physiological measurement (RPM) ishighly accessible but remains fragile under varying illumination, skintones, and motion. Radio frequency (RF) radar is largely … |
Jiyao Wang; Qingyong Hu; Duoxun Tang; Xiao Yang; Kaishun Wu; Jiangbo Yu; | code |
| 279 | LibraGen: Playing A Balance Game in Subject-Driven Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods often neglect this balance by enhancing oneaspect at the expense of others. To address this, we propose Libra-Gen, a novel framework that views extending foundation models forS2V generation as a balance game between intrinsic VGFM strengthsand S2V capability. |
Jiahao Zhu; Shanshan Lao; Lijie Liu; Gen Li; Tianhao Qi; hanwei hanwei; Bingchuan Li; FangfangLiu FangfangLiu; Zhuowei Chen; Tianxiang Ma; Qian HE; Yi Zhou; Xiaohua Xie; | code |
| 280 | DiaDem: Advancing Dialogue Descriptions in Audiovisual Video Captioning for Multimodal Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing models generally struggle to produce captions that faithfully reflect spoken content and speaker dynamics. To mitigate this limitation, we propose DiaDem, a powerful audiovisual video captioning model capable of generating captions with more precise dialogue descriptions while maintaining strong overall performance. |
Xinlong Chen; Weihong Lin; Jingyun Hua; Linli Yao; Yue Ding; Bozhou Li; Bohan Zeng; Yang Shi; Qiang Liu; Yuanxing Zhang; Pengfei Wan; Liang Wang; | code |
| 281 | Reasoning Path and Latent State Analysis for Multi-view Visual Spatial Reasoning: A Cognitive Science Perspective Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inthis paper, we aim to understand the successes and failure of VLMs inmulti-view spatial reasoning, via structured analysis of how VLMs en-code perceptual evidence, integrate relational information, construct in-termediate spatial representations and perform perspective transforma-tion. |
Qiyao Xue; Haoming Wang; Weichen Liu; Shiqi Wang; Yuyang Wu; Wei Gao; | code |
| 282 | AdaThinking-E: One-Token Entropy Regulation for Adaptive Thinking Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Multimodal large language models have demonstrated strongdocument reasoning capabilities by incorporating explicit thinking pro-cesses. |
Zining Wang; Tongkun Guan; Boming Chen; Zhentao Guo; Jianqiang Liu; chao jin; Chen Duan; Kai zhou; Pengfei Yan; Wei Shen; Xiaokang Yang; | code |
| 283 | Latent Visual Diffusion Reasoning with Monte Carlo Tree Search Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While recent advances in action quality assessment have achieved remarkable progress in evaluating performance, existing models remain black boxes, where they lack the ability to explicitly reveal the reasoning processes underlying their judgments. To address this limitation, we propose Latent Visual Diffusion Reasoning (LVDR), a novel framework that integrates keypoint-guided Monte Carlo Tree Search (MCTS) to model and visualize the latent visual reasoning process. |
Xirui Teng; Nan Xi; Junsong Yuan; | code |
| 284 | Recurrent Cross-View Object Geo-Localization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper,we propose ReCOT, a Recurrent Cross-view Object geo-localizationTransformer, which models CVOGL as a recurrent localization process.ReCOT introduces a set of learnable tokens that encode task-specificintent from the query image and prompt embeddings, and iteratively at-tend to the reference features to refine the predicted location. |
Xiaohan Zhang; Siyuan Cao; Xiaokai Bai; Yiming Li; Zhangkai Shen; Zhe Wu; Lun Luo; Qi Ming; Xiaoxi Hu; Hui-Liang Shen; | code |
| 285 | GraphVid: Interactive Graph-Controllable Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To enableflexible yet precise multi-subject control, we introduce GraphVid, agraph-conditioned image-to-video generation model that enables inter-active control through structured interaction graphs. |
Vedant Shah; Onkar Susladkar; Tushar Prakash; Kiet Nguyen; Tianjiao (Joey) Yu; Adheesh Juvekar; Muntasir Wahed; Ismini Lourentzou; | code |
| 286 | UniBYD: A Unified Framework for Learning Robotic Manipulation Across Embodiments Beyond Imitation of Human Demonstrations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose UniBYD, a unified framework thatuses a dynamic reinforcement learning algorithm to discover manipula-tion policies aligned with the robot’s physical characteristics. |
Tingyu Yuan; Biaoliang Guan; Wen Ye; Ziyan Tian; Yi Yang; Weijie Zhou; Zhaowen Li; Yan Huang; Peng Wang; Chaoyang Zhao; Jinqiao Wang; | code |
| 287 | SHIFT: Motion Alignment in Video Diffusion Models with Adversarial Hybrid Fine-Tuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study motion alignment in video diffusionmodels post-training. To address this, we introduce pixel-motion rewardsbased on pixel flux dynamics, capturing both instantaneous and long-term motion consistency. |
Xi Ye; Wenjia Yang; Yangyang Xu; Xiaoyang Liu; Duo Su; Mengfei Xia; Jun Zhu; | code |
| 288 | MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce MetricAnything,a simple and scalable pretraining framework that learns metric depthfrom diverse, noisy 3D sources without manual prompts, camera-specificmodeling, or task-specific architectures. |
Jiahui Yang; Donglin Di; Xuancheng Zhang; Lei Fan; Jianxun Cui; Hao Li; Baorui Ma; | code |
| 289 | Few-Shot Synthetic Image Attribution: Identifying Unseen Generators with Limited Samples Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, wepropose a new paradigm for synthetic image attribution, termed few-shotattribution. |
Shiyu Wu; Shuyan Li; Jing Li; Jing Liu; Yequan Wang; | code |
| 290 | Boosting Correspondence Learning with Structure-Aware Estimator Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In real-worldscenarios, inliers are often sampled from geometric manifolds and tendto appear in spatially dense clusters, especially in texture-rich regions.As a result, WLS may be dominated by dense inlier groups, leadingto a poorly-conditioned estimation problem that is vulnerable to noiseor degenerate configurations. To address this bottleneck, we propose astructure-aware estimator, which explicitly models inlier correlations viaa graph Laplacian matrix and integrates this prior into a maximum-likelihood framework. |
Tianyu Yan; Wei An; Pu Wang; Yingqian Wang; | code |
| 291 | Unfold The World: Factorize 4D Properties in Reinforcing Spatial Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present FactoSR, afactorized reinforcement learning framework that explicitly interpret the* Equal Contribution, † Project Lead, B Corresponding Authordimensions collapsed by visual projection. |
Yijun Yang; Shenghe Zheng; Wenbo Li; Jianhui Liu; Haoze Sun; Yanbing Zhang; Jiaxiu Jiang; Lin Song; Haoyang Huang; Nan Duan; Lei Zhu; | code |
| 292 | Mitigating Sycophancy in Multimodal Chart Understanding Via Vision-Grounded Verification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To ensure verification reliability, we propose the Con-trastive Visual Dependency Score (CVDS), an information-theoreticmetric that measures the KL divergence between the model’s predic-tions with and without visual input, filtering out textual-bias-drivenmisjudgments. |
Xiaolong Wang; Zhiwei Lin; Tai Liu; Qiang Li; Jian Sun; | code |
| 293 | Face Anything: 4D Face Reconstruction from Any Image Sequence Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Thisformulation transforms dense tracking and dynamic reconstruction intoa canonical reconstruction problem, enabling temporally consistent ge-ometry and reliable correspondences within a single feed-forward model.By jointly predicting depth and canonical coordinates, our method en-ables accurate depth estimation, temporally stable reconstruction, dense3D geometry, and robust facial point tracking within a single architec-ture. We implement this formulation using a transformer-based modelthat jointly predicts depth and canonical facial coordinates, trained us-ing multi-view geometry data that non-rigidly warps into the canonicalspace. |
Umut Kocasarı; Simon Giebenhain; Richard Shaw; Matthias Niessner; | code |
| 294 | Beyond Artifacts: Real-Centric Envelope Modeling for Reliable AI-Generated Image Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As generative archi-tectures evolve and images undergo multi-round cross-platform shar-ing and post-processing (chain degradations), these artifact cues be-come obsolete and harder to detect. To address this, we propose Real-centric Envelope Modeling (REM), a new paradigm that shifts de-tection from learning generator artifacts to modeling the robust distribu-tion of real images. |
Ruiqi Liu; Yi Han; Zhengbo Zhang; Liwei Yao; Zhiyuan Yan; Jialiang Shen; ZhiJin Chen; Manni Cui; Boyi Sun; Lubin Weng; Jing Dong; Yan Wang; Shu Wu; | code |
| 295 | SegFly: A 2D-3D-2D Paradigm for Aerial RGB-Thermal Semantic Segmentation at Scale Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Bylifting less than 3% of RGB images into a semantic 3D point cloud andrendering it into all views, our approach enables dense pseudo ground-truth generation across large image collections, automatically producing97% of RGB labels and 100% of thermal labels while achieving 91% and88% annotation accuracy without any 2D manual refinement. |
Markus Gross; Sai Bharadhwaj Matha; Rui Song; Viswanathan Muthuveerappan; Conrad Christoph; Julius Huber; Daniel Cremers; | code |
| 296 | ClearText-Video: A Large-Scale Text-Centric Video Dataset Bridging Video Restoration and Scene-Text Enhancement Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Multimodal Large Language Models (MLLMs) have recentlydemonstrated strong progress in visual–linguistic understanding, yet theirperformance on text-centric video reasoning remains … |
JINLONG LI; Jiaming Ding; Dingfu Lu; Malcolm Hsiu; Chuang Ke; Kangning Yang; Bochen Guan; Lan Fu; Jie Cai; Huiming Sun; Zibo Meng; | code |
| 297 | Trajectory-aware Cross-view Geo-Localization with Sequential Observations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Cross-view geo-localization matches ground-level observationsagainst geo-tagged satellite imagery. Recent methods show that sequen-tial queries such as video clips yield richer … |
Tianyi Gao; Jiayu Lin; Danielle Beaulieu; Nathan Jacobs; | code |
| 298 | XDen-1K: A Density Field Dataset of Real-World Objects Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While currentmethods, including VLM-based and other learning-based approaches,have shown promise in physical property inference, their evaluation isoften hindered by the lack of physically grounded reference data. Toaddress this gap, we introduce XDen-1K, the first large-scale multi-modal dataset that provides physically grounded density field for real-world objects. |
Jingxuan Zhang; Tianqi Yu; Yatu Zhang; Jinze Wu; Kaixin Yao; Jingyang Liu; Yuyao Zhang; Jiayuan Gu; Jingyi Yu; | code |
| 299 | Divide and Align: Disentangled Vision-Language Learning for Text-Based Person Anomaly Search Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing TPAS models often learn a singlejoint embedding where appearance, action, and background informationare entangled, causing over-reliance on identity cues, poor alignment foraction-centric queries, and limited semantic connection between actionsand places where they occur. To address these issues, we propose Se-mantic Decoupled Alignment (SeDA), a disentangled vision–languageretrieval framework that explicitly factorizes both visual and textualrepresentations into appearance, action, and background components.SeDA introduces Semantic Token Projection, which derives three se-mantic queries from the global [CLS] token, softly aggregates modal-ity tokens relevant to each factor, and recomposes the resulting factortokens into a compact retrieval embedding. |
Alex Ergasti; Tomaso Fontanini; Claudio Ferrari; Massimo Bertozzi; Andrea Prati; | code |
| 300 | EventVGGT: Exploring Cross-Modal Distillation for Consistent Event-based Depth Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Byneglecting the inherent temporal continuity of event data, these meth-ods fail to leverage the rich temporal priors encoded in VFMs, ultimatelyyielding temporally inconsistent and less accurate depth predictions. Toaddress this, we introduce EventVGGT, a novel framework that ex-plicitly models the event stream as a coherent video sequence. |
Yinrui Ren; Jinjing Zhu; Kanghao Chen; Zhuoxiao Li; Jing OU; Zidong Cao; Tongyan Hua; Peilun Shi; Yingchun Fu; Wufan Zhao; Hui Xiong; | code |
| 301 | DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Whilethe self-supervised model DINOv3 provides strong structured visual rep-resentations, its lack of native textual alignment hinders direct applica-tion to OVSS. To bridge this gap, we propose DINOde, an ODE-basedframework that continuously aligns CLIP text embeddings to the DINOvisual manifold. |
Sung-Hoon Yoon; Hoyong Kwon; Changgyoon Oh; KUK-JIN YOON; | code |
| 302 | BrainFIBRE: A Foundation Model Via Information Decomposition for Brain Microstructure Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Here, we present BrainFIBRE (Brain Foundation Model via Information Decomposition for BRain MicrostructurE), the first foundation model for brain microstructure, pretrained on three NODDI-derived microstructure maps from the UK Biobank dataset (55,592 participants). |
Zijian Dong; Yi Lin; Fang Ji; Jianxiong Zhou; Eric Kwun Kei NG; Juan Zhou; | code |
| 303 | RTE-FM-Dehazer: Radiative Transfer Equation Inspired Flow Matching for Real-World Image Dehazing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Moreover, large-scale real hazy/clear pairs are impractical to collect, and existing synthesis approaches fail to reproduce the full complexity of natural haze. To address these issues, we present RTE-FM-Dehazer, a novel dehazing approach, together with a scalable data pipeline. |
Chenfeng Wei; Chun Wang; Boyang Zhao; Si Zuo; Shenhong WANG; Chenguang Yang; | code |
| 304 | SelfMOTR: Revisiting MOTR with Self-Generating Detection Priors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose SelfMOTR, a simple yet highly effective detector-free alternative that decouples proposal discovery from association using self-generated internal detection priors. |
Fabian Gülhan; Emil Mededovic; Yuli Wu; Johannes Stegmaier; | code |
| 305 | Holo-Captioning: A Comprehensive Textual View of 3D Scenes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Totackle this challenging task, we first develop an effective captioning en-gine to produce detailed descriptions of individual entity instances andinstance pairs, and contribute a large-scale benchmark comprising over15K scenes for training and evaluation. Building upon this foundation,we propose HoloScribe, a novel model that features an instance-awaredecoupled pipeline for generating structured holo-captions, and furtherincorporates anchor-aware instance linking to identify relational instancepairs. |
Kun-Yu Lin; Chengke Bu; Zhenguo Li; Kai Han; | code |
| 306 | Incentive Noise and Structural Prior Infusion for Multi-Modal Object Re-Identification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Moreover, prevailing methods lack explicit mechanisms to reconcile fine-grained structural discrepancies between modalities, even after high-level semantic alignment. To address these challenges, we propose a novel framework centered on Positive-Incentive Noise (π-noise) and structured prompt modulation. |
Weixiang Zhou; Yuhao Wang; Xingguo Xu; Cong Wang; Weizhen Zhou; Zhixun Su; Jinshan Pan; | code |
| 307 | SkySplat-OV: Generalizable Language Gaussian Splatting for Open-Vocabulary Scene Understanding from Sparse Satellite Views Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: More recent generalizable 3DGS methods show promise, butthey perform poorly on sparse-view satellite imagery due to the uniquepushbroom imaging mode, limited geometric constraints, and extremescale variations. To address these limitations, we propose SkySplat-OV,a feed-forward framework that integrates the rational polynomial co-efficient (RPC) model into a generalizable language 3DGS pipeline. |
Xuejun Huang; Yi Wan; Lei Yu; Xinyi Liu; Zhi Zheng; Bin Zhang; Yingying Pei; Yi Liu; Changjun Zhu; Yi Liu; Xiangyuan Cai; Hongwei Hu; Xin Zhang; Yongjun Zhang; | code |
| 308 | AIMold: An Autonomous AI-based Pipeline for Complex Mold Design Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The dataset comprises 4,934 CAD models and over 3,850 mold assemblies, totaling more than 23k individual models. Building upon this dataset, we propose a comprehensive pipeline that predicts demolding orientations, identifies auxiliary components, and constructs parting surfaces to derive a complete, manufacturing-ready mold assembly for downstream CAD/CAM workflows. |
Pengyun Qiu; Shuo Wang; Zeyuan Chen; Yihao Zhi; Chongjie Ye; XIAOGUANG HAN; | code |
| 309 | G3AFT: Glance Guided Gradient Aligned Fine-Tuning for Visual Autoregressive Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Within this framework, we introduce two com-plementary techniques, i.e. Global Scope Gradient Recording (GSGR)and Noise Suppressed Gradient Projection (NSGP). |
Jiayi Zhang; Xiefan Guo; Xinzhu Ma; Di Huang; | code |
| 310 | Linear Fusion MultiDiffusion for Fast Training-Free Spherical Panorama Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose LF-MultiDiffusion, a training-free panorama gen-eration method that extends MultiDiffusion to support linear projectionsbetween target and reference image spaces. |
Akio Hayakawa; Yusuke Mukuta; Tatsuya Harada; | code |
| 311 | SDSA: Shallow-Deep Squeezing Adapter for Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent adapters for vision-language models (VLMs) still suf-fer from dense cross-modal interactions and limited control over align-ment capacity. To address this issue, we propose a two-stage Shallow-Deep Squeezing Adapter (SDSA) that explicitly regulates cross-modalinteraction density and alignment capacity. |
Hao Wang; Rui Zhu; Xiaoxu Li; Zhanyu Ma; Jing-Hao Xue; | code |
| 312 | Diffusion Models Are Open-World Affordance Learners: Leveraging Generative Priors for 3D Affordance Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by the saying: “WhatI can not create, I do not understand”, we find generative modelscan generate semantically valid HOI images, which indicates inherent en-coding of affordance concepts. Building on this insight, we propose DAG,the first innovative diffusion-based 3D affordance grounding frameworkthat extracts general affordance knowledge from text-to-image diffusionmodels for 3D affordance prediction. |
Hanqing Wang; Zhenhao Zhang; Kaiyang Ji; Mingyu Liu; Wenti Yin; yuchao chen; Zhirui Liu; Xiangyu Zeng; Tianxiang Gui; Hangxing Zhang; Jiahao Yuan; Zhiqing Cui; Jiaxin Liu; Zhiyuan Ma; Hui Xiong; | code |
| 313 | SK-Adapter: Skeleton-Based Structural Control for Native 3D Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Native 3D generative models have achieved remarkable _x001C_-delity and speed, yet they su_x001B_er from a critical limitation: inability toprescribe precise structural articulations, where precise structural con-trol within the native 3D space remains underexplored. |
Anbang Wang; AO Yuzhuo; Elliott (Shangzhe) Wu; Chi-Keung Tang; | code |
| 314 | PhyDetEx: Detecting and Explaining The Physical Plausibility of T2V Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: With the constructed dataset,we introduce a lightweight fine-tuning approach, enabling VLMs to notonly detect physically implausible events but also generate textual expla-nations on the violated physical principles.To investigate this issue, we construct aPID (Physical Implausibility Detection) dataset, which consists of a testsplit of 500 manually annotated videos and a train split of 2,588 pairedvideos, where each implausible video is generated by carefully rewritingthe caption of its corresponding real-world video to induce T2V modelsproducing physically implausible content. |
Zeqing Wang; Keze Wang; Lei Zhang; | code |
| 315 | CubicSplat: Differentiable Vector Graphics Via Error-Bounded Forward Relaxation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We trace this fragility to a gradientseesaw: design choices that improve forward geometric exactness cansystematically degrade the induced gradient signal, and vice versa. Tonavigate this tension we introduce CubicSplat, a differentiable vectorrasterizer that replaces Bézier closest-point solvers with uniform poly-line surrogates whose geometric error is bounded at O(S −2 ). |
Chenglong Liu; Xin Zhang; Yimeng Zhu; Liyang He; Yixiao Ma; Yu Su; Zhenya Huang; Qi Liu; | code |
| 316 | InstrAct: Towards Action-Centric Understanding in Instructional Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this,we propose InstrAct, a pretraining framework for instructional videos’action-centric representations. |
Zhuoyi Yang; Jiapeng Yu; Reuben Tan; Boyang Albert Li; Huijuan Xu; | code |
| 317 | Glance: Accelerating Diffusion Models with 1 Sample Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we take a different perspective: we accelerate smartly, not evenly, applying smaller speedups to early semantic stages and larger ones to later redundant phases. |
Zhuobai Dong; Rui Zhao; songjie wu; Suyang Hou; Junchao Yi; Zhengyuan Yang; Lijuan Wang; Alex Jinpeng Wang; | code |
| 318 | MG2-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Flat vector retrievaloften ignores structural dependencies, while current graph-based methodsrely on costly “translation-to-text” pipelines that discard fine-grainedvisual information. To address these limitations, we propose MG2 -RAG,a lightweight Multi-Granularity Graph RAG framework that jointlyimproves graph construction, modality fusion, and cross-modal retrieval.MG2 -RAG constructs a hierarchical multimodal knowledge graph bycombining lightweight textual parsing with entity-driven visual ground-ing, enabling textual entities and visual regions to be fused into unifiedmultimodal nodes that preserve atomic evidence. |
Sijun Dai; Qiang Huang; Xiaoxing You; Jun Yu; | code |
| 319 | ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose ProMSA, a progressive multimodal search agentfor KB-VQA. |
ZhengXian Wu; Hangrui Xu; Kai Shi; Zhuohong Chen; Yunyao Yu; Chuanrui Zhang; ZIRUI LIAO; JunYang JunYang; Zhenyu Yang; Haonan Lu; Haoqian Wang; | code |
| 320 | Understanding Geometric Representations in Self-Supervised Vision Transformers Via Subspace Intervention Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a controlled subspace intervention framework to investigate how self-supervised Vision Transformers (ViTs) encode dense geometric information. |
Weichen Zhou; Yawen Zou; Chunzhi Gu; Ran Dong; Haoran Xie; Chao Zhang; | code |
| 321 | How Far Are Vision-Language Models from Constructing The Real World? A Benchmark for Physical Generative Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We curate over 26,000 structures spanning 13 architecturalstyles—each verified to construction-document standards (LOD 350)—and develop a deterministic 10-test structural validation framework. |
Luyu Yang; Yutong Dai; An Yan; Viraj Prabhu; Ran Xu; Zeyuan Chen; | code |
| 322 | Beyond Common Sense: Grounding Logical Anomaly Detection in Inspection Criteria Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current detection paradigms typically operate without accessto these explicit standards, forcing models to rely on pretrained com-mon sense rather than strict logical rules. To address this limitation,we introduce SCAN, a Systematic Criteria-driven ANomaly inspec-tion framework optimized through a two-stage post-training pipeline.We first employ supervised finetuning based on rule-by-rule Chain-of-Thought (CoT) reasoning to systematically verify explicit inspection cri-teria. |
Yuanze Li; Shihao Yuan; Zimeng Zhu; Ming LIU; Wangmeng Zuo; Guangming Shi; | code |
| 323 | GUI-AIMA: Aligning Intrinsic Multimodal Attention with A Context Anchor for GUI Grounding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Graphical user interface (GUI) grounding is a key capa-bility for computer-use agents, mapping natural-language instructionsto actionable regions on the screen. Existing … |
Shijie Zhou; Viet Lai; Hao Tan; Jihyung Kil; Wanrong Zhu; Changyou Chen; Ruiyi Zhang; | code |
| 324 | Recolour What Matters: Region-Aware Colour Editing Via Token-Level Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Early text-driven methods rely on discrete language descriptionsthat cannot accurately represent continuous chromatic variations. Toovercome this limitation, we propose ColourCrafter, a unified diffusionframework that transforms colour editing from global tone transfer intoa structured, region-aware generation process. |
Yuqi Yang; Dongliang Chang; Yijia Ling; Ruoyi Du; Zhanyu Ma; | code |
| 325 | Geometric Gradient Rectification for Safe Open-Set Semi-Supervised Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We argue that both face a common trade-off: aggressive filtering can discard informative but hard ID samples,whereas utilization can introduce auxiliary gradients that conflict withsupervised learning when pseudo labels are wrong. |
Jiahe Chen; Qian Shao; Qiyuan Chen; Jiaying He; Jintai Chen; Hongxia Xu; Jian Wu; | code |
| 326 | Reliable Reasoning in SVG-LLMs Via Multi-Task Multi-Reward Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present CTRL-S(Chain-of-Thought Reinforcement Learning for SVG), a unified frame-work that introduces a chain-of-thought mechanism to explicitly exposethe model’s reasoning process during SVG generation. |
Haomin Wang; Qi Wei; Qianli Ma; Shengyuan Ding; Jinhui Yin; Kai Chen; hongjie Zhang; | code |
| 327 | PriorMaskMap: Robust Online Vectorized Map Construction with Biased Priors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Online vectorized map construction provides autonomous ve-hicles with essential, real-time semantic and geometric scene understand-ing. While leveraging prior maps presents an … |
Haoming Xu; Wei Li; Yu Hu; | code |
| 328 | StrucTab: A Structured Optimization Framework for Table Parsing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: At the optimization level, we introduce Uni-TabRL, a unifiedRL framework that leverages decomposed rewards (validity, structure,and content) to provide stable and informative optimization signals. |
Gengluo Li; Shangpin Peng; Chengquan Zhang; Binghong Wu; Hao Feng; Weinong Wang; Pengyuan Lyu; Huawen Shen; Xingyu Wan; Zhuotao Tian; Han Hu; Can Ma; Yu ZHOU; | code |
| 329 | DETR Is Secretly A Multispectral Detector: Zero-Parameter Adaptation Via Semantic Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing methods predomi-nantly follow a parameter-expanding paradigm, introducing customizedfusion modules that inevitably increase computational overhead and dis-rupt the topological consistency with unimodal pre-trained weights. Inthis paper, we challenge this convention and propose ZPA-MDETR,a zero-parameter adaptive multispectral framework without additionallearnable fusion parameters. |
Xiangyang Li; Zhiwei Jiang; Wushuai Jin; Pengyang Niu; Chunna Tian; Lingqiao Liu; | code |
| 330 | Multi-modality Image Fusion Under Adverse Weather: Mask-Guided Feature Restoration and Interaction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods often struggle with effective representation learning under such conditions, limiting their practical performance. To address these challenges, we propose a mask-guided MMIF method that integrates feature restoration and interaction. |
Xilai Li; Xiaosong Li; Haishu Tan; Tao Ye; Huafeng Li; Hongbin Wang; | code |
| 331 | Agentic Collaborative Cognition for Zero-Shot 3D Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing meth-ods face an intrinsic bottleneck due to the finite observation perspectivesinherent in videos and the implicit perception of 3D scenes. In this pa-per, we propose a collaborative multi-agent framework that assigns aPlanning Agent to handle high-level viewpoint planning and supplementnovel perspectives, and a Perception Agent to explicitly summarize the3D scene into a structured holistic cognitive map. |
Wenxin Wang; Bo Zhang; Feng Chen; Zixuan Wang; Wen Li; Changsheng Li; Yinjie Lei; | code |
| 332 | Unified and Efficient Point-Line Local Features Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a Unified Effi-cient Points and Lines (UPAL) feature extractor that jointly extracts key-points, line segments, and feature descriptors within a single lightweightarchitecture. |
François Costa; Raphael Kreft; Felix Möller; Hardik Shah; Ramanathan Rajaraman; Eckhard Goedeke; Shaohui Liu; Rémi Pautrat; Marc Pollefeys; | code |
| 333 | Hybrid Event–Frame Sensors: Modeling, Calibration, and Simulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present the first unified, statistics-basedimaging noise model that jointly describes the noise behavior of APS andEVS pixels. |
yunfan lu; Nico Messikommer; Xiaogang Xu; Liming Chen; Yuhan Chen; Nikola Zubic; Davide Scaramuzza; Hui Xiong; | code |
| 334 | World Knowledge in The Weights: Reading Concept Circuits of Vision Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We demonstrate the util-ity of concept circuits in three ways: (1) Automatic spurious correla-tion discovery: leveraging the statistics of our global concept circuitsto identify shortcut dependencies within the model. |
Yanlin Chen; Tang Li; Xi Peng; | code |
| 335 | RA-SOD: Reliability-Aware RGB-T Salient Object Detection Under Modality Degradation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, in real-world scenarios,the reliability of each modality is inherently unstable: RGB images de-grade under low illumination, motion blur, and noise, while thermal im-agery often suffers from contrast compression and sensor artifacts. Suchdegradation introduces unreliable perceptual evidence that can misleadcross-modal fusion and significantly deteriorate detection performance.To address this challenge, we propose RA-SOD, a reliability-aware RGB-T salient object detection framework that explicitly models modality re-liability and integrates it into feature learning and cross-modal fusion.First, we introduce a reliability-conditioned representation that adap-tively compensates degraded modality features while preserving struc-tural cues. |
Hongbo Gao; Zhengyu Li; XUERU NIE; Dihao Zhu; lijun zhao; Yunke Wang; Chang Xu; | code |
| 336 | Unified Video Dense Prediction from Disjoint Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing task-specific an-notations are fragmented across incompatible, domain-specific datasets.Current unified systems circumvent this by restricting training to fullyco-annotated data, or by incurring the large computational cost of pseudo-labeling. To mitigate this, we introduce UniD, a unified video model thatjointly predicts eight dense scene properties—depth, surface normals, se-mantic segmentation, boundaries, human parts, albedo, shading, and ma-terials—all learned from disjoint, domain-specific datasets. |
Yihong Sun; Seoung Wug Oh; Jiahui Huang; Bharath Hariharan; Joon-Young Lee; | code |
| 337 | NeSy-Route: A Neural-Symbolic Benchmark for Constrained Route Planning in Remote Sensing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address these limita-tions, we introduce NeSy-Route, a large-scale neuro-symbolic benchmarkfor constrained route planning in remote sensing. Within this benchmark,we introduce an automated data-generation framework that integrateshigh-fidelity semantic masks with heuristic search to produce diverseroute-planning tasks with provably optimal solutions. |
Ming Yang; Zhi Zhou; Shi-Yu Tian; Kun-Yang Yu; Lan-Zhe Guo; Yu-Feng Li; | code |
| 338 | IMMoE: Incomplete Multi-View Anomaly Detection Via Mixture of View Experts Fusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We proposed a pipeline for automati-cally generating the IMVAD dataset and generated the RIMAD datasetbased on the Real-IAD dataset through this pipeline. |
Lei Hu; | code |
| 339 | DARL: Efficient Document-to-Markup Generation Via Look-Ahead Diffusion Trajectory Sampling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Conversely, existing paralleland diffusion-based methods accelerate inference but often struggle tomaintain strict structural dependencies, resulting in significantly higherparsing errors. To bridge this gap, we propose DARL (Diffusion AutoRe-gression with Look-ahead), a novel hybrid decoding framework. |
Wentao Yang; Yongxin Shi; Rui Tang; Peirong Zhang; Shihang Wu; Huiguo He; Zheng Huang; Dezhi Peng; Minghui Liao; Lianwen Jin; | code |
| 340 | LVSPM: Long Sequence View Synthesis and Pose Estimation Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present LVSPM, a generalizable model that jointly esti-mates camera poses and synthesizes novel views from uncalibrated im-age collections. |
Xi Chen; Yachi Zhang; Linghao Chen; Minghua Liu; Hao Su; Zexiang Xu; Xiaoshuai Zhang; | code |
| 341 | Toward Interpretable Analysis of Whole-slide Pathology Images Via Large Language Model-based Agentic Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing computational pipelines often lack this rea-soning trajectory, resulting in opaque and unjustifiable predictions. Tobridge this gap, we present PathAgent, a training-free, large languagemodel (LLM)-based agent framework that emulates the reflective, step-wise analytical approach of human experts. |
Jingyun Chen; Linghan Cai; Zhikang Wang; Yi Huang; Songhan Jiang; Shenjin Huang; Hongpeng Wang; Yongbing Zhang; | code |
| 342 | SKEL-CF: Coarse-to-Fine Biomechanical Skeleton and Surface Mesh Recovery Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose SKEL-CF, a new framework for estimating SKEL parameters. |
Da Li; Ji-Ping Jin; Xiaodong Cun; Xuanlong Yu; Wei LIU; Rui Fan; Jiangang Kong; Kai Chen; Xi SHEN; | code |
| 343 | Semantic-Anchored Evidential Fusion for Domain-Robust Whole-Slide Survival Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We hypothesize that high-level pathol-ogy semantics, such as tumor grade and micro-environmental architec-ture, provide a domain-invariant semantic representation that mirrorsthe robust diagnostic logic of human pathologists. |
YUCHENG XING; Ling Huang; Pei Liu; Jingying Ma; Jiaxing Xu; Kai He; Mengling Feng; | code |
| 344 | Freqformer: Image-Demoiréing Transformer Via Effective Frequency Decomposition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present Freqformer, aTransformer-based framework specifically designed for image demoiréingthrough targeted frequency separation. |
Xiaoyang Liu; Bolin Qiu; Zheng Chen; Libo Zhu; Zihan Zhou; Kai Liu; Jiezhang Cao; Yulun Zhang; | code |
| 345 | RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning Under Urban Road Scenarios Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Around roadmarkings and urban traffic systems, we propose RoadBench, a system-atic benchmark that comprehensively evaluates MLLMs’ fine-grainedspatial understanding and reasoning capabilities using Bird’s-Eye View(BEV) and First-Person View (FPV) image inputs. |
Jun Zhang; Xin Zhang; Jie Feng; Long Chen; Junhui Wang; Zhicheng Liu; Depeng Jin; Yong Li; | code |
| 346 | ConTrack: Constrained Hand Motion Tracking with Adaptive Trade-off Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Human demonstrations provide strong priors for robot ma-nipulation, yet it is non-trivial to transfer them to execute on real robotsdue to the kinematic gap. In dexterous … |
Yutong Liang; Quanquan Peng; Rizhao Qiu; Xiaolong Wang; | code |
| 347 | D3F-IR: Dual-Domain Deterministic Flow Matching for Visible-to-Infrared Translation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we reformulate the transla-tion process into two explicitly decoupled but tightly coordinated con-tinuous flows, proposing a novel framework named D3F-IR. |
Peiyi Zeng; Diedong Feng; Zhen Liu; Zhenming Peng; Bing Zeng; Shuaicheng Liu; | code |
| 348 | Less Tokens, Better Forecasts: Sparse Residual Routing for Efficient Weather Prediction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing ViT-based weather forecasting models apply uni-form computation across all spatial tokens, even though nearby atmo-spheric grid points often contain similar values and large regions evolvesmoothly over time. |
Janet Wang; Yunbei Zhang; Lin Zhao; Xi Xiao; Jihun Hamm; Xiao Wang; | code |
| 349 | Cross-Species Animal Re-Identification with Semantic Consistency Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As a result, representations learned across species tend to form fragmented embedding spaces, which severely limits cross-species generalization. To address this challenge, we propose Semantic Consistency Learning (SCL), a framework designed to learn representations that remain stable across appearance variations while preserving semantic structures shared across species. |
Shuoyi Chen; Yuejia Li; Mang Ye; | code |
| 350 | Reliability-Aware 3D Geometric Injection for Universal Person Re-identification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To safely harness the clothing-invariant and canonical structural properties of 3D geometry, we propose UniGeo, a Universal Monocular 3D-Enhanced ReID framework driven by a ConsistencyAware Reliability Gate and Dual-Stream Residual Fusion. |
Bohan Su; Jiashuo Wang; Fangyi Liu; Mang Ye; | code |
| 351 | Drift-AR: Single-Step Visual Autoregressive Generation Via Anti-Symmetric Drifting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose Drift-AR, which leverages entropysignal to accelerate both stages: 1) for AR acceleration, we introduceEntropy-Informed Speculative Decoding that align draft_x0015_target entropydistributions via a causal-normalized entropy loss, resolving the entropymismatch that causes excessive draft rejection; 2) for visual decoder ac-celeration, we reinterpret entropy as the physical variance of the initialstate for an anti-symmetric drifting _x001C_eld_x0016_high-entropy positions acti-vate stronger drift toward the data manifold while low-entropy positionsyield vanishing drift_x0016_enabling single-step (1-NFE) decoding without it-erative denoising or distillation. |
zhen zou; Xiaoxiao Ma; Mingde Yao; Jie Huang; Linjiang Huang; Feng Zhao; | code |
| 352 | Boosting Text-Driven Video Segmentation Via Geometry-Aware Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Geometry-enhanced Language-guidedVideo segmentation (GeoLaV), a two-stage framework that distills 3Dgeometric knowledge from images to enhance text-driven video segmenta-tion. |
Tianyu Zhu; Yingping Liang; Hesong Li; Ying Fu; | code |
| 353 | Stylized Video Generation Via Decoupled Data Synthesis and Gated Style Token Injection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, we construct MoStyle-5K, a stylized video dataset of 5,000 videos spanning 27 artistic stylesand 4,000 preference pairs, built via a decoupled content-motion syn-thesis pipeline that explicitly preserves temporal priors in the supervi-sion signal. Building on this, we propose Implicit Gated Style To-ken Injection (IGST), which directly concatenates style patches as to-kens to bypass the encoder bottleneck, disables RoPE for style tokensto remove spatial alignment bias, and applies head-wise gated atten-tion to suppress residual structural correlations. |
Xin Wei; Yijie Fang; Yanjia Li; Liangyi Wu; Qin Yang; Mingrui Zhu; Nannan Wang; Xinbo Gao; | code |
| 354 | Wid3R: Wide Field-of-View 3D Reconstruction Via Camera Model Conditioning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Wid3R, a feed-forward neural network for multi-view visual geometry reconstruction that supports wide field-of-viewcamera models. |
Dongki Jung; Jaehoon Choi; Adil Qureshi; Somi Jeong; Dinesh Manocha; Suyong Yeon; | code |
| 355 | Resolution-Agnostic Neural Operators for Multi-Rate Sparse-View CT Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Sparse-view Computed Tomography (CT) reconstructs im-ages from a limited number of X-ray projections to reduce radiation andscanning time, which is an ill-posed inverse problem. … |
Aujasvit Datta; Jiayun Wang; Asad Aali; Anima Anandkumar; | code |
| 356 | CRISP: Calibration-Aware Visual State Space Duality for Remote Sensing Image Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, we observe that VSSD compresses spatial context into a global aggregation that suppresses high-frequency responses, causing excessive boundary smoothing in remote sensing semantic segmentation. To address this, we propose CRISP, a calibration framework with two components. |
kangning wang; Haopeng Zhang; Zhiguo Jiang; | code |
| 357 | LDC-MTL: Balancing Multi-Task Learning Through Scalable Loss Discrepancy Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose LDC-MTL, a simple andscalable loss discrepancy control approach for MTL, formulated froma bilevel optimization perspective. |
Peiyao Xiao; Chaosheng Dong; Shaofeng Zou; Kaiyi Ji; | code |
| 358 | CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, this progress largely relieson strong pre-trained capabilities, while leaving two fundamental limi-tations insufficiently addressed: (1) Architectural Weak-Coupling, wherethe unidirectional flow forces reliance on coarse VLM prompts and wastesSAM’s pixel-level structural guidance, causing localization drift; and (2)Object-Centric Semantic Bias, where models overemphasize dominantobject semantics while remaining insensitive to spatial reasoning cru-cial for RRSIS. Motivated by these observations, we propose CROSS, atightly integrated paradigm for RRSIS. |
Tingzhang Luo; Ruizhong Liu; Yichao Liu; Cheng Fan; Yu Liu; Jianyuan Guo; | code |
| 359 | ExpoMotion: A Large-Scale Benchmark and A Householder Projection Network for Multi-Exposure Fusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Comprising 1,738 sequences and 10,909images across diverse environments, it covers a wide range of motionsand provides high-fidelity GTs constructed through an expert-guided ac-quisition pipeline. To tackle the complex dynamics and extreme con-ditions captured in this benchmark, we propose the Householder Or-thogonal Projection network (HOP), which revisits MEF deghostingfrom a mathematical perspective via Householder transformation, de-coupling multi-frame alignment into exposure pre-alignment and ghostfiltering. |
Yao Liu; Lishen Qu; Jie Liang; shihao zhou; Hui Zeng; Yabin Peng; Huipeng Lin; Lei Zhang; Jufeng Yang; | code |
| 360 | Spectral and Trajectory Regularization for Diffusion Transformer Super-Resolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Diffusion transformer (DiT) architectures show great poten-tial for real-world image super-resolution (Real-ISR). However, theircomputationally expensive iterative sampling … |
Jingkai Wang; Yixin Tang; Jue Gong; Jiatong Li; Shu Li; Libo Liu; Jianliang Lan; Yutong Liu; Yulun Zhang; | code |
| 361 | CritiqueDriveVLM: From Verifier-Guided Reinforcement Learning to Latent Thought Distillation for Autonomous Driving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While traditional tool-augmented frameworks and Chain-ofThought (CoT) approaches mitigate these issues, they incur exorbitant token consumption and unacceptable latency, rendering real-time deployment impractical. To resolve this reliability-efficiency trade-off, we propose CritiqueDriveVLM, a novel unified three-stage framework internalizing reasoning directly into the VLM. |
Zhaohong Liu; Hao Ye; Xianlin Zhang; Mengshi Qi; | code |
| 362 | EgoEverything: A Benchmark for Human Behavior–Inspired Long-Context Egocentric Video Understanding in AR Environment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We release our dataset at https://sai-lab-nyu.github.io/EgoEverything/. |
Qiance Tang; Ziqi Wang; Jieyu Lin; Ziyun Li; Barbara Salvo; Sai Qian Zhang; | code |
| 363 | EruDiff: Refactoring Knowledge in Diffusion Models for Advanced Text-to-Image Synthesis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we pro-pose EruDiff, which aims to refactor the knowledge within diffusionmodels. |
Xiefan Guo; Xinzhu Ma; Haoxiang Ma; Zihao Zhou; Di Huang; | code |
| 364 | Group3D: MLLM-Guided Semantic Grouping for Open-Vocabulary 3D Object Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Open-vocabulary 3D object detection aims to localize andrecognize objects beyond a fixed training taxonomy. In multi-view RGBsettings, recent approaches often decouple … |
Youbin Kim; Jinho Park; Hogun Park; Eunbyung Park; | code |
| 365 | Towards Robustness Against Typographic Attack with Training-free Concept Localization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To achieve in-terpretable and effective robustness against TA, we propose a novel,training-free mechanistic interpretability method. |
Bohan Liu; Wenqian Ye; Guangzhi Xiong; Zhenghao He; Sanchit Sinha; Aidong Zhang; | code |
| 366 | Learning to Balance: Decoupled Siamese Diffusion Transformer for Reference-Based Remote Sensing Image Super-Resolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, ex-isting methods often suffer from a trade-off between over-reliance on ref-erence information, which leads to texture artifacts, and under-utilizationof such information, which results in insufficient detail recovery. To ad-dress these issues, we propose DS-DiT, a Decoupled Siamese DiffusionTransformer that decouples the interaction between low-resolution (LR)and reference (Ref) conditions within the attention mechanism. |
Bin Luo; Runmin Dong; Zhaoyang Luo; Jinxiao Zhang; Jiyao Zhao; Fan Wei; Haohuan Fu; | code |
| 367 | DiverseVAR: Balancing Diversity and Quality of Next-Scale Visual Autoregressive Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce DiverseVAR, a test-time framework that enhances the output diversity of text-conditioned visual autoregressive models (VAR) without additional training or substantial computational overhead. |
Mingue Park; Prin Phunyaphibarn; Phillip (Yuseung) Lee; Minhyuk Sung; | code |
| 368 | BEV-GS: Feed-forward Gaussian Splatting in Bird’s-Eye-View for Road Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose BEV-GS, a real-time single-frame roadsurface reconstruction method based on feed-forward Gaussian splatting.BEV-GS consists of a prediction module and a rendering module. |
Wenhua Wu; Tong Zhao; Chensheng Peng; Lei Yang; Zhe Liu; Hesheng Wang; | code |
| 369 | MomentSeg: Moment-Centric Sampling for Enhanced Referring Video Object Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Theformer often overlooks essential temporal cues, while the latter increasessystem complexity. To address this, we propose a unified frameworkthat jointly optimizes Temporal Sentence Grounding (TSG) and Re-fVOS, naturally incorporating key moment grounding capability. |
Ming Dai; Sen Yang; Boqiang Duan; Wankou Yang; Jingdong Wang; | code |
| 370 | Learning Probabilistic Embeddings for Unsupervised Action Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper concerns the problem of unsupervised temporalaction segmentation for long, untrimmed videos. |
Shuai Li; Duc Vu; Juergen Gall; | code |
| 371 | RobustRDP: Advancing Reaction Diagram Parsing Via Synthetic-to-Real Data Scaling and Robustness-Oriented Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, wepropose RobustRDP, a Robust Reaction Diagram Parser built on amultimodal large language model (MLLM). |
Jianting Tang; zezhong wu; Linli Xu; | code |
| 372 | Generalize LMMs to Versatile Visual Modalities Via Fabricated Modality Synthesis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We argue that different visualmodalities are merely distinct samplings of the same physical world.Therefore, effective generalization requires models to possess both modality-agnostic perception of scene semantics and the adaptability to modality-specific characteristics. To achieve this, we propose a training frame-work, VVM-Tuning, to equip LMMs with these capabilities throughmodality synthesis and modality contexts. |
Shihao Yuan; Yuanze Li; Ruyi Zhang; Ming LIU; Wangmeng Zuo; | code |
| 373 | Unmasking-Time Visual Calibration for Hallucination Mitigation in Multimodal Discrete Diffusion Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Despite the recent success of multimodal discrete diffusionlanguage models (multimodal dLLMs) in generating text through iter-ative demasking with bidirectional attention, object … |
Tian Qin; Junzhe Chen; Tianshu Zhang; Lijie Wen; | code |
| 374 | ART-VSR: Adaptive Rectified Trajectories for One-Step Video Super-Resolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose ART-VSR, a state-time adaptation framework that estimates a continuous token-wise timestep map with an Adaptive Timestep Estimator (ATE) and uses it to guide a Latent Trajectory Rectifier (LTR) toward a compatible starting state. |
JianHui Zhang; Chen Fang; Wanghao Wanghao; chaoyu feng; LEI LEI; Jue Wang; Shuaicheng Liu; | code |
| 375 | Differentiable Polarized Path Tracing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our approach estimates unbiased gradients through a combina-tion of path replay and local caching. |
Pramod Rao; Jérémy Riviere; xilong zhou; Abhijeet Ghosh; Abhimitra Meka; Thabo Beeler; Marc Habermann; Christian Theobalt; Delio Vicini; | code |
| 376 | Disentangling Hallucinations: Orthogonal Semantic Projection for Robust Interpretability Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a unified theoretical framework, Linear Semantic Attribution (LSA), which generalizes across discriminative methods. |
Emirhan Bilgiç; Baptiste Caramiaux; Zhi Yan; Gianni Franchi; | code |
| 377 | CL-Anomaly: Layer-Adaptive Mixture-of-Experts with Multimodal Large Language Model for Continual Learning in Anomaly Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper,we propose CL-Anomaly, a parameter-efficient fine-tuning frameworkbased on an isolation-sharing collaboration to enable continual learningfor anomaly detection with MLLMs. |
Wen Dong; Zhao Wang; Shuangqing Zhang; Kai Sun; Ben Li; Guosen Xie; Caifeng Shan; Fang Zhao; | code |
| 378 | MAVFusion: Efficient Infrared and Visible Video Fusion Via Motion-Aware Sparse Interaction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current video fusion methods im-prove temporal consistency by introducing interactions across frames,but they often require high computational cost. To mitigate these chal-lenges, we propose MAVFusion, an end-to-end video fusion frameworkfeaturing a motion-aware sparse interaction mechanism that enhancesefficiency while maintaining superior fusion quality. |
Xilai Li; Weijun Jiang; Xiaosong Li; Yang Liu; Hongbin Wang; Tao Ye; Huafeng Li; Haishu Tan; | code |
| 379 | HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While recent hybrid methods mitigate this issue via region-level parallel decoding with VLMs, independent region decoding loses full-page context and might weaken global coherence. To address this issue, we propose Hierarchical Speculative Decoding (HSD), a two-stage local-to-global framework for document parsing. |
Wenhui Liao; Hongliang Li; Pengyu Xie; Xinyu Cai; Yufan Shen; Yi Xin; Qi Qin; Shenglong Ye; Tianbin Li; Ming Hu; Junjun He; Yihao Liu; Wenhai Wang; Min Dou; Bin Fu; Botian Shi; Yu Qiao; Lianwen Jin; | code |
| 380 | GeoEdit: Geometry-Aware Object Editing Via Dual-Branch Denoising Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present GeoEdit, a training-free Lift-Manipulate-Render-Denoise pipeline that satisfies both constraints. |
Yi He; Jiangming Wang; Xinyu Wang; Mark Fong; Songchun Zhang; Yuxuan Xue; Hai-Tao Zheng; Yue Ma; | code |
| 381 | LiSTAR: Ray-Centric World Models for 4D LiDAR Sequences in Autonomous Driving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This task is inherently challenging due to the sensor’s uniquespherical geometry, the temporal sparsity of point clouds, and the com-plexity of dynamic scenes. To address these challenges, we present LiS-TAR, a novel generative world model that operates directly on the sen-sor’s native geometry. |
Pei Liu; Songtao Wang; Lang Zhang; Xinyue Peng; Yuandong Lyu; Jiaxin Deng; Songxin lu; Weiliang Ma; Xueyang Zhang; Yifei Zhan; Kun Zhan; Jun Ma; | code |
| 382 | PHOSA: Photorealistic 3D Sign Avatar Modeling and Benchmark Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we focus on photorealistic sign avatar modeling,which is crucial for effective communication with the Deaf community andis characterized by complex hand gestures and nuanced facial expressions.To this end, we introduce MVSign, the first multi-view Chinese signlanguage dataset co-designed with Deaf experts, featuring diverse gesturesand rich annotations. |
Haodong Wang; Hezhen Hu; Wengang Zhou; Houqiang Li; | code |
| 383 | Thinking in Streaming Video Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, weintroduce ThinkStream, a framework for streaming video reasoningbased on a Watch–Think–Speak paradigm that enables models toincrementally update their understanding as new video observations ar-rive. |
Zikang Liu; Longteng Guo; Handong Li; Ru Zhen; Xingjian He; Ruyi Ji; Xiaoming Ren; Yanhao Zhang; Haonan Lu; Jing Liu; | code |
| 384 | OrthoTailor: Geometric Orthogonalization for Conflict-Free Unified Fashion Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose OrthoTryOn, a unified framework mitigating thisinterference within a shared Low-Rank Adaptation (LoRA) module. |
Zhaotong Yang; Ying Tai; Jiahui Zhan; Yu Zheng; Jianjun Qian; Jian Yang; | code |
| 385 | JanusMesh: Fast and Zero-Shot 3D Visual Illusion Generation Via Cross-Space Denoising Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we presenta fast and training-free framework for generating text-driven 3D visualillusions. |
Siang Ling Zhang; Huai-Hsun Cheng; Tsung-Ju Yang; Yu-Lun Liu; | code |
| 386 | RL-AWB: Deep Reinforcement Learning for Auto White Balance Correction in Low-Light Night-time Scenes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present RL-AWB, a novel framework com-bining statistical methods with deep reinforcement learning for night-time white balance.To further facilitate cross-sensor evaluation,we introduce the first multi-sensor nighttime dataset. |
Yuan-Kang Lee; Kuan-Lin Chen; Chia-Che Chang; Yu-Lun Liu; | code |
| 387 | Mitigating Positional Leakage in 3D Masked Autoencoders for Robust Representation Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we identifythat the decoder in existing 3D MAE frameworks tends to over-rely onpositional information, which weakens semantic representation learningand leads to suboptimal feature quality. |
Xu Yan; Huiqun Wang; Chen Wang; Lei Ren; Di Huang; | code |
| 388 | Kiroshi: An Agentic Perception System for High-Accuracy Image Parsing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods either lack semantic awareness orproduce suboptimal predictions; even interactive matting approaches rely on manualverification and repetitive checking, making fully automatic matting still unattainable.In this work, we propose Kiroshi , an agentic perception system for high-accuracyimage parsing. |
Haipeng ZHOU; Jinshan Liu; He Zhang; Xuequan Lu; Jun Ma; Lei Zhu; | code |
| 389 | Consistent Monocular Depth Estimation with Contact Region Boundary-Aware Refinement Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While contemporary monocular depth estimation (MDE)methods achieve remarkable overall acacy, they consistently produce er-roneous depth discontinuities at object contact regions, particularly be-tween objects and supporting surfaces. In this paper, we address this crit-ical limitation by presenting a boundary-aware monocular depth estima-tion framework that enforces depth continuity at contact areas throughthe principled exploitation of contact boundaries as explicit structuralpriors. |
Yinuo Wang; QingMiao QingMiao; Wangmeng Zuo; | code |
| 390 | See Only When Needed: Context-Aware Attention Intervention for Hallucination-Free LVLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Extensive experiments across multiple LVLM back-bones and benchmarks show that CAI achieves state-of-the-art halluci-nation mitigation, and our analysis characterizes CAI as a KL-minimalattention reweighting with bounded interference under inactive gates orsmall tilts. |
Yuqing Lei; Wenbo Lyu; Yingjun Du; Xiantong Zhen; Cees Snoek; Ling Shao; | code |
| 391 | Pointer-CAD V2: Plan-Then-Construct CAD Generation with Dimension-Aware Parametric Precision Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specif-ically, we propose a unified framework that decouples parameter rea-soning from geometric construction through a Plan-Then-Constructparadigm. |
Dacheng Qi; Chenyu Wang; Jingwei Xu; Yi Ma; Shenghua Gao; | code |
| 392 | GeMoE: Gating Entropy Is All You Need for Uncertainty-aware Adaptive Routing in MoE-based Large Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose viewing token routing as an information encoding task, framing dynamic routing as a Minimum Description Length (MDL) problem in encoding. |
Chaoxiang Cai; Minghe Weng; Jie Li; Yibo Jiang; Longrong Yang; Zequn Qin; Xi Li; | code |
| 393 | AdaBridge-SR: Adaptive Bridge Matching for Real-World Image Super-Resolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite their remarkable performance in real-world imagesuper-resolution, diffusion models remain challenged by the perception–distortion trade-off between structural fidelity and perceptual realism.Existing methods typically rely on a globally predefined restoration tra-jectory, applying spatially uniform noise and a fixed temporal schedule.Such rigid trajectories often lead to a dilemma: reliable structures may bedistorted, while severely degraded textures may become over-smoothed. To address this, we propose AdaBridge-SR, built upon a novel spatio-temporal adaptive bridge matching (ST-ABM) formulation that castsrestoration as a controlled bridge between degraded and clean imagedistributions. |
Jiangang Wang; Shangquan Sun; Aiping Zhang; Yuning Cui; Wenqi Ren; | code |
| 394 | CFM: Language-aligned Concept Foundation Model for Vision Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose CFM, a language-aligned concept foundation model for vision that provides fine-grained concepts, which are human-interpretable and spatially grounded in the input image. |
Kai Wittenmayer; Sukrut Rao; Amin Parchami-Araghi; Bernt Schiele; Jonas Fischer; | code |
| 395 | Ctrl-Z Sampling: Scaling Diffusion Sampling with Controlled Random Zigzag Explorations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Diffusion models generate conditional samples by progres-sively denoising Gaussian noise, yet the denoising trajectory can stall atvisually plausible but low-quality outcomes with conditional misalign-ment or structural artifacts. We interpret this behavior as local optimain a surrogate quality landscape: Once early denoising commits to a sub-optimal global structure, later steps mainly sharpen details and seldomcorrect the underlying mistake. |
Shunqi Mao; Wei Guo; Chaoyi Zhang; Jieting Long; Ke Xie; Weidong Cai; | code |
| 396 | LlamaSeg: Image Segmentation Via Autoregressive Mask Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present LlamaSeg, a visual autoregressive framework that unifies multiple image segmentation tasks via natural language instructions. |
Jiru Deng; Tengjin Weng; Tianyu Yang; Wenhan Luo; Zhiheng Li; Wenhao Jiang; | code |
| 397 | ELDiff: When Evidential Learning Meets Text-to-Image Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose ELDiff, a new evidential learningsupervised T2I diffusion model, which leverages the advantages of uncertainty metric and conflict detection to enhance the fault tolerance of unreliable segmentation maps and suppress semantic conflicts, strengthening object-wise consistency learning. |
Qingtao Pan; Kai Ye; Zhihao Dou; Bing Ji; Shuo Li; | code |
| 398 | SLAIR: Structured Latent Flow Matching for All-in-One Image Restoration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose SLAIR, an all-in-one image restora-tion framework integrating latent space separation and deterministictransport. |
Shuyi Liang; Yixin Yang; Hanyue Lou; Yuning Cui; Boxin Shi; | code |
| 399 | Masked Depth Modeling for Spatial Perception Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we argue that the inaccuracies from depth sensors can be viewed as “masked” signals that inherently reflect underlying geometric ambiguities. |
Bin Tan; CHANGJIANG SUN; Xiage Qin; Hanat Adai; Zelin Fu; Tianxiang Zhou; Han Zhang; YINGHAO XU; Xing Zhu; Yujun Shen; Nan Xue; | code |
| 400 | AerialMetric: Benchmarking and Adapting UAV Monocular Metric Depth Estimation in The Real World Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although recent data-driven methods have achieved remarkable progress in ground-level scenarios, models trained primarily on street-view and indoor datasets exhibit significant domain gaps when applied to aerial viewpoints. To tackle these challenges, we introduce AerialMetric, a benchmark dataset designed to evaluate and facilitate the adaptation of monocular metric depth estimation under UAV aerial viewpoints. |
Zhongqiang Song; Guanying Chen; Yuqi Zhang; Yin Zou; Chuanyu Fu; Zhiyuan Yuan; Chuan Huang; Shuguang Cui; Xiaochun Cao; | code |
| 401 | HiReFF: High-Resolution Feedforward Human Reconstruction from Uncalibrated Sparse-View Video Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present HiReFF, a feed-forward method for 2K-resolution 360° human video reconstruction from uncalibrated sparse-view videos. |
YIMING JIANG; Hanzhang Tu; Wenfeng Song; Siyou Lin; Liang An; Shuai Li; Aimin Hao; Yebin Liu; | code |
| 402 | AeroVLA: A Vision-Language-Action Model for UAV Navigation Via Minimalist End-to-End Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose AeroVLA, aminimalist end-to-end Vision-Language-Action framework mapping rawvisual observations and fuzzy linguistic instructions directly to contin-uous physical control signals. |
Peng Xu; Zhengnan Deng; Jiayan Deng; Zonghua Gu; Shaohua Wan; | code |
| 403 | MATCH: Flow Matching for Multi-View Anomaly Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present MATCH , the first multi-view anomaly detection method based on Flow Matching (FM). |
Mathis Kruse; Melissa Schween; Bodo Rosenhahn; | code |
| 404 | Track4World: Feedforward World-centric Dense 3D Tracking of All Pixels Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a feedforward model, calledTrack4World, enabling an efficient holistic 3D tracking of every pixel inthe world-centric coordinate system. |
Jiahao Lu; Jiayi Xu; WENBO HU; Ruijie Zhu; Chengfeng Zhao; Sai Kit Yeung; Ying Shan; Yuan Liu; | code |
| 405 | HART: High-Resolution Annotation-Free Reasoning Technique Through A Closed-loop Framework Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose High-resolutionAnnotation-free Reasoning Technique (HART), a closed-loop frameworkthat enables LMMs to focus on and self-verify key regions of high-resolution visual inputs. |
Jiacheng Yang; Anqi Chen; Yunkai Dang; Qi Fan; Cong Wang; Wenbin Li; Feng Miao; Yang Gao; | code |
| 406 | FlowCIR: Semantic Transport Via Flow Matching for Zero-Shot Composed Image Retrieval Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose anew paradigm, namely FlowCIR, that casts ZS-CIR as conditionalsemantic transport between reference and target embeddings. |
Zhenqi He; Ziqi Jiang; Yuanpei Liu; Yanghao Wang; Teng Wang; Long Chen; | code |
| 407 | Step-by-Step Video-to-Audio Synthesis Via Negative Audio Guidance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a step-by-step video-to-audio (V2A) generation method that provides finer control over the generation process and more realistic audio synthesis. |
Akio Hayakawa; Masato Ishii; Takashi Shibuya; Yuki Mitsufuji; | code |
| 408 | Finding Highlight Images In Your Albums:From Benchmark To MLLM Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Album highlight recommendation (AHR) focuses on identifying images from personal collections that are most suitable for personal enjoyment and social sharing. However, due to the … |
Rong Qin; Congcong Sun; Yaopeng Dong; Yanbin Sun; Chensen Ding; Chenxi Zhao; Biao Wang; Qian Zhang; Eunil Park; Chi Man VONG; Jufeng Yang; | code |
| 409 | StreamSpatial: A Benchmark and Framework for Streaming 3D Visual-Spatial Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present StreamSpatialBench, a testbed for streaming 3Dvisual-spatial reasoning built on three pillars: fine-grained online tempo-ral perspectives (past / current / future anchored to the query times-tamp), dynamic multi-agent interactions, and comprehensive 3D spatialreasoning. |
Junlin Xie; Keyang Zhong; Quanlong Zheng; Ruifei Zhang; Kuo Wang; Yanhao Zhang; Haonan Lu; Xiang Wan; Guanbin Li; | code |
| 410 | ESTANet: Efficient Online Error Detection in Procedural Videos Via Prediction Inconsistency Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we study real-time online error detection in pro-cedural videos from a simple but overlooked perspective: the predictionbehavior of action detectors themselves. |
Shih-Po Lee; Reza Ghoddoosian; Faizan Siddiqui; Enna Sachdeva; Behzad Dariush; | code |
| 411 | Vulnerability of Privacy-Preserving Visual Localization Against Diffusion-based Attacks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To model adversarial behavior and to thoroughly measure amethod’s degree of privacy preservation, we introduce a new privacy at-tack that trains a conditional diffusion model to reconstruct images fromprivacy-preserving representations. |
Maxime Pietrantoni; Torsten Sattler; Gabriela Csurka; | code |
| 412 | TimeWalker: Personalized Neural Space for Lifelong Head Avatars Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present TimeWalker, a novel framework that models realistic, full-scale 3D head avatars of a person on lifelong scale. |
Dongwei Pan; Yang Li; Hongsheng LI; Kwan-Yee Lin; | code |
| 413 | Topology-Weighted Effective Rank: A Zero-Cost Proxy for Training Dynamics Stability in Neural Architecture Search Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Moreover, they overlook theheterogeneous importance of different model components, which limitstheir ability to generalize across datasets and search spaces. To mitigatethese limitations, we propose an Effective Rank Score called ER-Scoreas a novel ZCP to quantify stability in the training dynamics ofover-parameterized networks. |
Haojie Zhang; Yeming Yang; Songbai Liu; Lijia Ma; Ka-Chun Wong; Qiuzhen Lin; | code |
| 414 | E3VS-Bench: A Benchmark for Viewpoint-Dependent Active Perception in 3D Gaussian Splatting Scenes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Visual search in 3D environments requires embodied agentsto actively explore their surroundings and acquire task-relevant evidence.However, existing visual search and embodied AI benchmarks, includ-ing EQA, typically rely on static observations or constrained egocen-tric motion, and thus do not explicitly evaluate fine-grained viewpoint-dependent phenomena that arise under unrestricted 5-DoF viewpointcontrol, such as disambiguating object attributes observable only fromspecific angles. To address this limitation, we introduce E3VS-Bench,a benchmark for embodied 3D visual search where agents must controltheir viewpoints in 5-DoF to gather viewpoint-dependent evidence forquestion answering. |
Koya Sakamoto; Taiki Miyanishi; Daichi Azuma; Shuhei Kurita; Shu Morikuni; Naoya Chiba; Motoaki Kawanabe; Yusuke Iwasawa; Yutaka Matsuo; | code |
| 415 | Adaptive Spectrum-Aware Feature Disentangled Network for Small Object Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a novelsmall object detection framework termed SFDNet, which is capableof detecting small objects via efficient spectrum-aware feature disentan-glement. |
Yang Guo; Zihan Yang; Feifei Kou; Yulan Hu; Ran Zhang; Siyuan Yao; | code |
| 416 | Tiled Prompts: Overcoming Prompt Misguidance in Image and Video Super-Resolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, a coarse globalprompt often misses localized details (errors of omission) and provideslocally irrelevant guidance (errors of commission) which leads to substan-dard results at the tile level. To solve this, we propose Tiled Prompts,a unified framework for image and video super-resolution that generatesa tile-specific prompt for each latent tile and performs super-resolutionunder locally text-conditioned posteriors to resolve prompt misguidancewith minimal overhead. |
Bryan Sangwoo Kim; Jonghyun Park; Jong Chul Ye; | code |
| 417 | Who Does What and Where to Go: Orthogonal Alignment and Hierarchical Planning for Multi-Entity Trajectories Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose ProgTraj-Director, a framework driven by physical property disentanglement andhierarchical spatial planning. |
Zhang Wan; Yu Li; Tianze Huang; Juan Cao; Sheng Tang; | code |
| 418 | Think While Watching: Online Streaming Segment-Level Memory for Multi-Turn Video Reasoning in Multimodal Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Basedon Qwen3-VL, we improve single-round accuracy by 2.6% on Streaming-Bench and 3.79% on OVO-Bench.We construct a three-stage, multi-round, chain-of-thought (CoT) dataset with a stage-matched training strategy while en-forcing strict causality in streaming reasoning via a segment-level stream-ing causal mask and streaming positional encoding. |
Lu Wang; Zhuoran Jin; Yupu Hao; Yubo Chen; Kang Liu; Yulong Ao; Jun Zhao; | code |
| 419 | Asymmetric Anchoring: Opening The Black Box of MLLMs for Forgery Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Multimodal large language models (MLLMs) are promisingfor forgery detection, but most methods treat them as fixed black boxes.Preliminary studies have shown that this passive use overlooks instabili-ties in the internal data flow, leading to representation drift. In this work,we introduce the Asymmetric Anchoring Paradigm (AAP), an open-boxapproach that reshapes the MLLM’s internal data flow (rather than ap-pending external components) by re-purposing its pre-trained visual en-coder as an active truth anchor. |
Zhiqiang Yang; Renshuai Tao; Chunjie Zhang; Zhaoxiang Liu; Xiaolong Zheng; Yao Zhao; | code |
| 420 | Towards Reconfigurable Visual Feature Compression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Correspondingly, we propose a unified Reconfigurable Feature Compression framework, RFC, by feature factorization and recomposition.Finally, we construct a comprehensive benchmark of reconfigurable feature compression to verify the effectiveness of RFC. |
Jiahang Zhang; Wenhan Yang; Minghao Liu; Jiaying Liu; | code |
| 421 | TR-MoE: Temporal Reliability-Aware Mixture-of-Experts for Robust Tracking Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In visual object tracking, SAM2 stands out in deformation adaptability and distinguishing similar objects due to its pixel-level masks, whereas traditional discriminative trackers have advantages under occlusion and motion blur thanks to global semantic feature matching. To leverage this complementarity, we propose TR-MoE, a Temporal Reliability-aware Mixture-of-Experts framework. |
Tianle Wang; Xiangyang Yang; Jihua Zhu; Binrui Liu; Yanzhao Li; Shuiwang Li; | code |
| 422 | FlexiBrain: Resolution-Agnostic Voxel-Level Encoding for Native FMRI Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Here, we propose FlexiBrain, a resolution-agnosticvoxel-level encoding framework for native fMRI based on Mamba-JEPA. |
mo wang; Wenhao Ye; Junfeng Xia; Minghao Xu; Hongkai Wen; Quanying Liu; | code |
| 423 | Point Diffusion Mamba: Unified Diffusion-State-Space Modeling for Single-View 3D Reconstruction Under Data Scarcity Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In 3D re-construction, each point in the initial noisy input requires a precise pre-diction, yet the high-level features extracted by the Mamba module cap-ture only abstract semantic information from sparse points. To bridgethis gap, we introduce the Hierarchical Feature Integration Network,which fuses high-level semantic and local geometric features for eachpoint, overcoming the limitations of token-based point-cloud reconstruc-tion. |
Wei Zhou; Xinzhe Shi; Xingxing Hao; Xing Hao; Kang Li; Jinye Peng; Ying He; | code |
| 424 | NegAS: Negative Label Guided Attention and Scoring for Out-of-Distribution Object Detection with Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we take the first step toward building OODobject detection methods upon VLMs. |
Yingjie Zhang; Shuai Li; Peng Wang; | code |
| 425 | GradingBench: Evaluating End-to-End Compositional Reasoning of MLLMs for Automated Exam Grading Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Cur-rent benchmarks largely focus on isolated tasks and cannot fully eval-uate MLLMs’ end-to-end ability in such complex settings. To bridgethis gap, we present GradingBench, a comprehensive benchmark basedon automated exam grading in Chinese K–12 education, which system-atically evaluates MLLMs across the entire grading pipeline. |
Yuting Wang; Zixian Guo; Weihao You; Zhilong Ji; Jinfeng Bai; Wangmeng Zuo; | code |
| 426 | Geometry Grounding: Elevating Blind Distortion Correction with 3D Structural Priors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Without explicit 3Dgeometric constraints, such models may produce visually plausible recti-fications yet fail to enforce 3D-consistent structure. To address this lim-itation, this paper proposes a 3D geometry guided framework for singleimage distortion correction, named Geometry-Grounded Distortion Cor-rection (G2DC). |
Xiang Li; Weimin Shi; Qichuan Geng; Zhong Zhou; | code |
| 427 | DreamPartGen: Semantically Grounded Part-Level 3D Generation Via Collaborative Latent Denoising Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose DreamPartGen, a frameworkfor semantically grounded, part-aware text-to-3D generation. |
Tianjiao (Joey) Yu; Xinzhuo Li; Muntasir Wahed; Jerry Xiong; Yifan Shen; Ying Shen; Ismini Lourentzou; | code |
| 428 | Reflection-aware Generative Novel View Synthesis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: From input images, we estimate the mirrorplane and reflect camera poses to form virtual views. Based on this vir-tual view setup, we propose a two-stage generation method consisting ofMirror-gated attention and Reflection injection, which enables reflection-consistent NVS by explicitly leveraging reflection relationships in a multi-view diffusion model. |
GeonU Kim; Shin Dong-Yeon; Tae-Hyun Oh; | code |
| 429 | URoPE: Universal Relative Position Embedding Across Geometric Spaces Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However,existing formulations are typically limited to a fixed geometric space,namely 1D sequences or regular 2D/3D grids, which restricts their ap-plicability to many computer vision tasks that require geometric rea-soning across camera views or between 2D and 3D spaces. To addressthis limitation, we propose URoPE, a universal extension of RotaryPosition Embedding (RoPE) to cross-view or cross-dimensional geomet-ric spaces. |
Yichen Xie; Depu Meng; Yihan Hu; Chensheng Peng; Quentin HERAU; Masayoshi TOMIZUKA; Wei Zhan; | code |
| 430 | D-Rex : Diffusion Rendering for Relightable Expressive Avatars Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present D-Rex, a person-specific framework for photore-alistic, relightable, expressive, and animatable full-body human avatarswith free-viewpoint rendering. |
Timo Teufel; xilong zhou; Umar Iqbal; Jan Kautz; Marc Habermann; Vladislav Golyanik; Christian Theobalt; | code |
| 431 | Multi-Modal Controlled Coherent Motion Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, they often result in mismatched or unrealistic movements. To overcome these limitations, we propose MOCO, a novel diffusion-based framework capable of processing multiple simultaneous inputs—including speech audio, text descriptions, and trajectory data—to generate coherent and lifelike motions without requiring aligned multimodal data. |
Yifei Liu; Qiong Cao; Hongwei Yi; Huaiguang Jiang; Changxing Ding; | code |
| 432 | ORACLE-3D: Open-world Region-aligned Cross-modal Learning for Label-efficient 3D Scene Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The recent surge in vision-language models (VLMs) has revolutionized open-world perception, yet transferring this success to 3D scene understanding remains impeded by the scarcity of large-scale point-text pairs and the computational overhead of existing frozen-VLM approaches. To bridge this gap, we present ORACLE-3D (Open-world Region-Aligned Cross-modal LEarning for Label-efficient 3D Scene Understanding), a unified framework that effectively co-embeds point clouds, images, and text into a shared latent space without relying on labor-intensive 3Dtext pair construction. |
Yuru Wang; Pei Liu; Songtao Wang; Zehan Zhang; Xinyan Lu; Changwei Cai; Hao Li; Haipeng LIU; Qingtian Ning; Jun Ma; | code |
| 433 | OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We identify that the rootcause lies in the structural deficiencies of existing datasets across threedimensions: limited global scene and camera diversity, sparse interac-tion modeling, and insufficient individual attribute alignment. To bridgethese gaps, we present OmniHuman, a large-scale, multi-scene datasetdesigned for fine-grained human modeling. |
Lei Zhu; xing cai; Yingjie Chen; Li YiHeng; Binxin Yang; Hao Liu; Jie Chen; Chen Li; Jing LYU; | code |
| 434 | GRAFT: Geometric Refinement and Fitting Transformer for Human Scene Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our key insight is that geometry-based human–scene fitting can be amortized into fast feed-forward inference. |
Pradyumna YM; Yuxuan Xue; Yue Chen; Nikita Kister; István Sárándi; Gerard Pons-Moll; | code |
| 435 | TEASR: Training-Efficient Any-Step Diffusion Transformer for Real-World Image Super-Resolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose TEASR, a training-efficient anystep diffusion framework for Real-ISR that enables both one-step and multi-step restoration within a unified model. |
Xiang Gao; Chenxin Zhu; Yushun Fang; Qiang Hu; Xiaoyun Zhang; | code |
| 436 | 4D-VGGT: A SpatioTemporal Foundation Model for Dynamic Scene Geometry Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose 4D-VGGT, a general foundation model with divide-and-conquer spatiotemporal representation for dynamic scene geometry. |
Haonan Wang; Hanyu Zhou; Haoyue Liu; Luxin Yan; | code |
| 437 | MotionAnymesh: Physics-Grounded Articulation for Simulation-Ready Digital Twins Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Converting static 3D meshes into interactable articulated as-sets is crucial for embodied AI and robotic simulation. However, exist-ing zero-shot pipelines struggle with complex … |
WenBo Xu; Liu Liu; li zhang; Dan Guo; Ruonan Liu; | code |
| 438 | Hi-DiT: Hybrid Latent-Pixel Diffusion Transformer for Image Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Latent Diffusion Models (LDMs) relyon VAE compression, which can discard high-frequency visual cues andlimit pixel-level quality, while pixel-space diffusion often suffers from diffi-cult optimization when global structure and high-frequency details mustbe learned jointly. In this paper, we propose the Hybrid Latent-PixelDiffusion Transformer (Hi-DiT), a unified Diffusion Transformer thatbridges these two regimes within a single architecture. |
Yedong Shen; Yehao Li; Yingwei Pan; Yanyong Zhang; Ting Yao; | code |
| 439 | ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present ViewFusion, a two-stageframework that explicitly separates cross-view spatial pre-alignment fromquestion answering. |
Xingjian Tao; Yiwei Wang; Yujun Cai; Yifan Song; Jing Tang; | code |
| 440 | Stochastic Optimal Control Sampling for Diffusion Inverse Problems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce Stochastic Optimal Control Sampling (SOCS), which models the denoising process as a dynamical system and injects control signals via SOC. |
Jie Zhang; Youmei Qiu; Hanling Tian; Jingyuan Zhang; Xiang Yin; Xiaolin Huang; | code |
| 441 | SMG: Semantic Motion Graph for Monocular Dynamic Gaussian Splatting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our key insight is thatthe real-world scene motion is often structured by semantic coherence:regions that are spatially close and semantically related tend to exhibitconsistent dynamics.To eval-uate dynamic Gaussian Splatting under challenging real-world scenarios,we introduce a new multiview dataset collected under an ego-exo setup.Extensive experiments demonstrate that SMG achieves state-of-the-artperformance on monocular dynamic Gaussian Splatting across challeng-ing real-world benchmarks. |
Haozheng Yu; Xinyu Yang; Rundong Luo; Jennifer Sun; Bharath Hariharan; | code |
| 442 | ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce ReGRPO (Reflection-augmentedGroup Relative Policy Optimization), a framework that learns reflection-guidedcorrection in tool-using agents. |
Binjie Zhang; Mike Zheng Shou; | code |
| 443 | Do Flat Minima Improve Sparse Novel View Synthesis? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We demonstrate that this dif-ference arises because high-detail regions inherently require a sharp losslandscape for accurate reconstruction, whereas low-detail regions benefitfrom a flat loss landscape for improving generalization. Based on this in-sight, we introduce structure-aware sharpness, defined within structure-adaptive neighborhoods, and propose to adaptively adjust the sharpnessregularization weight according to the local image structure. |
Youngsik Yun; Dongjun Gu; Youngjung Uh; | code |
| 444 | Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present DR-MV3D (DenseReward for MV3D-VQA), a map-grounded learning framework that pro-vides dense, verifiable rewards to supervise the reasoning process. |
Jiho Choi; Seonho Lee; Seojeong Park; Hyunjung Shim; | code |
| 445 | From Masks to Pixels and Meaning: A New Taxonomy, Benchmark and Metrics for VLM Image Tampering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We reformulate VLM image tampering from coarse region labels to a pixel-grounded, meaning and language-aware task. |
Xinyi Shang; Yi Tang; Jiacheng Cui; Ahmed Elhagry; Salwa Al Khatib; Sondos Bsharat; Jiacheng Liu; Xiaohan Zhao; Jing-Hao Xue; Hao Li; Salman Khan; Zhiqiang Shen; | code |
| 446 | HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce# Corresponding author.HumanTracker to make humanoid tracking evaluation both perceptu-ally aligned and scalable. |
Dairu Liu; Zekun Qi; Jiayu Zeng; Yu Guan; Chenghuai Lin; Xuchuan Chen; Xinqiang Yu; Wenyao Zhang; HE WANG; Li Yi; | code |
| 447 | Adaptive Noise Covariance Scheduling Under Riemannian Metrics for Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This transition better aligns forward corruption with reverse denoising and enhances textural detail. We provide theoretical and quantitative analyses showing that this transition path yields smoother noise power spectrum evolution and lower noise power fluctuations. |
Bolin Deng; Die Hu; Bin Tan; Jun Wu; | code |
| 448 | One-Shot Feed-Forward 360° Animatable Avatar Via Inpainted UV-Space Gaussian Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods generally collapse underlarge camera pose variations, compromising the realism of 3D avatars.In this work, we propose a new framework to tackle the novel settingof one-shot 3D full-head animatable avatar reconstruction in a singleforward pass via inpainted UV-space Gaussian modeling, enabling 360◦rendering views and real-time animation. |
Shuling Zhao; Dan Xu; | code |
| 449 | PhyMAGIC: Physical Motion-Aware Generative Inference with Confidence-guided VLM Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We observe that different motions expose comple-mentary physical cues. Building on this observation, we propose Phy-MAGIC, a training-free framework that actively probes physical prop-erties by synthesizing targeted motions from a single image. |
Siwei Meng; Yawei Luo; Ping Liu; | code |
| 450 | CooperScene: Multi-Modal Cooperative Autonomy Benchmark with C-V2X Communication Characterization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce CooperScene, a high-fidelity cooperative autonomy dataset with real-world C-V2X communica-tion characterization. |
Bo Wu; Ruoshen Mo; Justin Yue; Yanyu Zhang; Janice Nguyen; Guoyuan Wu; Amit Roy-Chowdhury; Matthew Barth; Hang Qiu; | code |
| 451 | Exposure Bias Can Alleviate Itself Via Directional and Frequency Rectification in Flow Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing mitigation strategies typicallyrely on static constraints or external heuristics. In this work, we pro-pose that exposure bias itself inherently contains dynamic signals thatcan guide its own rectification. |
Guanbo Huang; Jingjia Mao; Fanding Huang; fengkai liu; Xiangyang Luo; Yaoyuan Liang; Jiasheng Lu; xiaoe Wang; Pei Liu; Ruiliu Fu; Ruqi Huang; Shao-Lun Huang; | code |
| 452 | When W4A4 Breaks Camouflaged Object Detection: Token-Group Dual-Constraint Activation Quantization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This exposes a token-local bottleneck—remove cross-token range domination and bound the zero-bin mass under4-bit activations. To address this, we introduce COD-TDQ, a COD-aware Token-group Dual-constraint activation Quantization method.COD-TDQ addresses this token-local bottleneck with two coupled steps:Direct-Sum Token-Group (DSTG) assigns token-group scales to sup-press cross-token range domination, and Dual-Constraint Range Projection(DCRP) projects each token-group clip range to keep the step-to-dispersionratio and the zero-bin mass bounded. |
Tianqi Li; Wenyu Fang; Xin He; Xue Geng; Xu Cheng; Yun Liu; | code |
| 453 | Self-Evolving MCP-GUI Agents Via Automated Environment Generation and Experience Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We formulate MCP-GUI interplay as a unified hybrid policylearning problem where the agent learns when each modality providescomplementary advantages, and show that distillation and experienceaugmentation target fundamentally different failure modes—requiringapplication-aware mechanism selection. |
Tiantian He; Yihang Chen; Keyue Jiang; Ka Lee; Kaiwen Zhou; Kun Shao; Shuai Wang; | code |
| 454 | WAFT-Stereo: Warping-Alone Field Transforms for Stereo Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce WAFT-Stereo, a simple and effective warping-based method for stereo matching. |
Yihan Wang; Jia Deng; | code |
| 455 | RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, we design a data genera-tion paradigm, release the instruction dataset RSICI, and establish atask-specific RSICC benchmark. |
Yelin Wang; Zijia Song; Shuo Ye; Chuanguang Yang; Miaoyu Wang; Yong Xu; Zhulin An; Yongjun Xu; Zitong Yu; | code |
| 456 | Revisiting Deepfake Detection: BCNet for Robust Generalization Beyond Semantic Dependence Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This bias limits generalization in real-world scenarios. To address this issue, we propose the Basis CorrectionNetwork (BCNet), which consists of two modules: the attention-guidedsemantic erasure (ASE), which adaptively identifies and erases seman-tic regions of the image by capturing the model’s semantic attention,and the normalized-gradient perturbation enhancement (NPE) scalesthe gradients of fake samples to concentrated values and adds them asa small perturbation to the original sample, helping the model recog-nize more forgery patterns and improving its ability to distinguish fakefrom real samples. |
Jian Yang; Shibo Yao; Renshuai Tao; Chuangchuang Tan; Yao Zhao; | code |
| 457 | Geometric Probing for Isotropic Optimization Manifold in Sparse-View 3D Gaussian Splatting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Guided by this analysis, we propose Stable 3DGS (StableGS), which steers optimization toward isotropic stability via geometric probing. |
Yunlong Zhao; Xiaoheng Deng; Hongyan Xu; Yichao Cao; Keke Huang; Shuo Yang; Xiangjian He; Lei Fan; Zhuohua Qiu; Xiu Su; | code |
| 458 | MaterialFlow: Attribute-Disentangled Material Transfer Via Trajectory-Aware Velocity Modulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose MaterialFlow, atraining-free framework for precise material transfer using pre-trainedflow models. |
Sung-Lin Tsai; Bo-Kai Ruan; Yu-Hsuan Chen; Wen-Huang Cheng; Hong-Han Shuai; | code |
| 459 | S2Gest: Split-Scan State Space Models for Dynamic Hand Gesture Recognition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address theselimitations, we propose the Split-Scan Gesture Network (S2Gest), whichrethinks temporal modeling by leveraging continuous state-space evolutionto represent discrete sequential inputs. |
KeFan Chen; Yong Gu; bo li; Longjie Huang; Jiajun Zhang; | code |
| 460 | Self-supervised Garment Dynamics with Persistent Wrinkles Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Through a comprehensive evaluation, we show thatfor the first time, self-supervised learning models can generate naturalpersistent wrinkles, outperforming existing methods on a variety of gar-ments, body shapes, and body motions, according to a range of metrics.Our code is publicly available at https://github.com/realcrane/EPNet |
Xiaoyuan Yang; Deshan Gong; Taku Komura; He Wang; | code |
| 461 | ProAct: Agentic Lookahead in Interactive Environments Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose ProAct, a two-stage frameworkfor internalizing foresight. |
Yangbin Yu; Mingyu Yang; junyou li; Yiming Gao; Feiyu Liu; Yijun Yang; Zichuan Lin; Jiafei Lyu; Zhicong Lu; Deheng Ye; Jie Jiang; | code |
| 462 | SurvMILKD: A Weakly Supervised Survival Analysis Framework for Multi-Teacher Knowledge Distillation Using Pathology Foundation Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose SurvMILKD, the first multi-teacher knowledge distillation (MKD) framework for WSIbased survival analysis. |
Mayur Mallya; Ali Khajegili Mirabadi; Hossein Farahani; Ali Bashashati; | code |
| 463 | Geometry-Aware Spatio-Temporal Context Modeling for 4D Occupancy Forecasting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a Geometry-AwareSpatio-Temporal context modeling method (GAST) for 4D occupancyforecasting, built upon progressive explicit-implicit generation and dual-path spatio-temporal modeling. |
Sitao Chen; Zhuangwei Zhuang; Hui Luo; Qingyao Wu; Mingkui Tan; | code |
| 464 | E-M3RF: An Equivariant Multimodal 3D Re-assembly Framework Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Additionally, solutions do not impose physi-cal constraints that explicitly prevent overlapping assemblies. To addressthese limitations, we introduce E-M3RF, an equivariant multimodal 3Dreassembly framework that takes as input the point clouds, containingboth point positions and color of fractured fragments, and predicts thetransformations required to reassemble them, using SE(3) flow match-ing. |
Adeela Islam; Stefano Fiorini; MANUEL LECHA SANCHEZ; Theodore Tsesmelis; Stuart James; Pietro Morerio; Alessio Del Bue; | code |
| 465 | Kilometer-Vision: A New Frontier for Large-Scale Spatial Awareness in VLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We push the frontier of large-scale spatial intelligence in Vision-Language Models (VLMs) and introduce the first benchmark that probes geographical layout understanding from real-world videos, spanning up to 1km distances. |
Aravindh Mahendran; Michael King; Matthew Grimes; Antoine Yang; Tyler Zhu; Joseph Heyward; Tengda Han; Shiry Ginosar; Chen Sun; Dima Damen; Simon Osindero; Noah Snavely; Simon Lynen; Joao Carreira; Viorica Patraucean; | code |
| 466 | Head Avatars with Dynamic Explicit Hair Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present DynHair, a novel method for tracking and mod-eling dynamic hair for human head avatars. |
Vanessa Sklyarova; Haonan Chen; Berna Kabadayi; Tobias Kirschstein; Zicong Fan; Xi Wang; Gerard Pons-Moll; Matthias Niessner; Marc Pollefeys; Michael Black; Justus Thies; | code |
| 467 | HIDA: A Human-Intuition-Guided Depth-Aware Framework for Zero-Shot Amodal Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Human-intuition-guided depth-aware (HIDA) is proposed as a plug-and-play framework that combines two frozen vision foundation models with a lightweight trainable segmentation network to estimate amodal masks using only a visible bounding box. |
Peng Li; Kelin Wang; Bingchuan Chen; Ya-Li Hou; Mingxia Shen; Bo Li; | code |
| 468 | Aggregating Cross-Domain Knowledge Via Learnable Tokens for Multi-Teacher Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This rigid approach often inducesrepresentational dissonance and semantic conflicts, where teachers withdisparate architectures and divergent optimization goals lead to sub-optimal performance and restricted architectural flexibility. In this pa-per, we propose a novel framework termed ACTok that recasts multi-teacher distillation as a token-mediated aggregation process. |
Wu Ran; Weijia Zhang; ShuYang Pang; JingSheng Liu; Xiaohui Zhang; Yichao Yan; Chao Ma; | code |
| 469 | Physics-Grounded Disentangled Flow Modeling for Brain Disease Progression Trajectory Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Such an entangled modeling of structural deformation and image intensity variation limits physical plausibility, model generalization, and interpretability. To address this, we propose PDF, a Physics-grounded Disentangled Flow matching framework for longitudinal brain disease forecasting. |
Jun Wang; Peirong Liu; | code |
| 470 | FairSteer: Cross-Attention Steering Towards A Fairer Text-Guided Image Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose EquiSteer,a training-free method that works per sample by steering cross-attention(CA) activations at inference time. |
Tatiana Gaintseva; Akshit Achara; Greg Slabaugh; Jiankang Deng; Ismail Elezi; | code |
| 471 | SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While existing approaches rely on injecting explicitexternal priors (e.g., VLM features or structural conditions) to mitigatethese issues, this reliance severely bottlenecks model robustness and gen-eralization. To overcome this limitation, we present SAMA (factorizedSemantic Anchoring and Motion Alignment), a framework that factor-ize video editing into semantic anchoring and motion modeling. |
Xinyao Zhang; Wenkai Dong; YuXin Song; Bo Fang; Qi Zhang; Jing Wang; Fan Chen; Hui Zhang; Haocheng Feng; Yu Lu; Hang Zhou; Chun Yuan; Jingdong Wang; | code |
| 472 | StereoGS: Sparse-View 3D Gaussian Splatting Via Stereo Priors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we proposeStereoGS, a novel sparse-view 3DGS framework that integrates stereopriors to establish reliable binocular consistency. |
Wenhao Yuan; Yiyuan Ge; Deli Cai; | code |
| 473 | Towards Benign Memory Forgetting for Selective Multimodal Large Language Model Unlearning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To pioneerand evaluate progress toward this objective, we introduce S-MLLMUnBench, the first benchmark designed to jointly and quantitatively assessan unlearning method’s efficacy in knowledge erasure and the preservationof image understanding. |
Zhen Zeng; Leijiang Gu; Zhangling Duan; Feng Li; Zenglin Shi; Cees Snoek; Meng Wang; | code |
| 474 | Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we establish theoretical formulations for the two levels of optimization and derive a hybrid advantage that serves both objectives. |
Wenxuan Zhang; Yuhui Wang; Donggang Jia; Xiaoqian Shen; Jian Ding; Ivan Viola; Jürgen Schmidhuber; Mohamed Elhoseiny; | code |
| 475 | REVEAL: Reasoning-Enhanced Forensic Evidence Analysis for Explainable AI-Generated Image Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce REVEAL-Bench, a reasoning-enhanced multi-modal benchmark for AI-generated image forensics, structured aroundexplicit chains of forensic evidence derived from lightweight expert mod-els and consolidated into step-by-step chain-of-evidence traces. Basedon this benchmark, we propose REVEAL (Reasoning-enhanced Foren-sic Evidence Analysis), an explainable forensic framework trained withexpert-grounded reinforcement learning. |
Huangsen Cao; Qin Mei; Zhiheng Li; Yuxi Li; Zhan Meng; Ying Zhang; Chen Li; Zhimeng Zhang; Xin Ding; Yongwei Wang; Jing LYU; Fei Wu; | code |
| 476 | Know3D: Prompting 3D Generation with Knowledge from Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we pro-pose Know3D, a novel framework designed to incorporate rich knowl-edge from Multimodal Large Language Models (MLLMs) into 3D gen-eration processes. |
Wenyue Chen; Wenjue Chen; Peng Li; Qinghe Wang; XU JIA; Heliang Zheng; Rongfei Jia; Yuan Liu; Ronggang Wang; | code |
| 477 | MV-STRIDE: Enabling MLLMs to Master Multi-View Spatial Reasoning Via Hierarchical Capability Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We develop a systematic QA gen-eration pipeline leveraging diverse 3D scene sources that enforces cross-view dependency constraints to prevent single-view solvability, generat-ing multi-level spatial reasoning tasks supported by cognitively groundedchain-of-thought supervision for complex inference. |
Jin Xu; Xiaojian Huang; zhang zhihong; Luo Zhuodong; Xuejin Chen; Jie Zhao; Xin Liu; Wang Xinzhi; Jiansheng Wei; | code |
| 478 | AnE: Pushing The Reasoning Frontier of Multimodal LLMs Via Anchor Evolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While current methods leverage self-reflection or self-evolution topush these boundaries, they still suffer from cognitive drift and hallu-cinated reasoning paths caused by low-quality synthetic data. To ad-dress these challenges, we propose Anchor E volution (AnE), a newparadigm that integrates truth-anchored data curation and model evo-lution, achieving faithful and steady performance gains at the reason-ing frontier. |
Zehao Wang; Yihan Zeng; Zidong Gong; Yuanfan Guo; Feng Zhu; Hongzhi Zhang; Wei Zhang; Wangmeng Zuo; | code |
| 479 | 3DWay: Generalizing Robot Manipulation Via 3D Consistent Waypoints Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Moreover, using 2D trajectories with depth still leavesthe free-space waypoints ambiguous, limiting reliable 3D reasoning. Toaddress this, we propose predicting 3D consistent waypoints (3DWay)from multi-view images. |
Huang Ziqin; Yingyue Li; Chenyangguang Zhang; Ruida Zhang; Yuxin Chen; Gu Wang; Xingyu Liu; Masayoshi TOMIZUKA; Xiangyang Ji; | code |
| 480 | RealDyadic: Synthesizing Realistic Dyadic 3D Dialogue with Neural Appearance Priors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These labels inherently contain severe noise and fail to capture fine-grained facial dynamics, particularly accurate lip closure. To address these limitations, we propose RealDyadic, a two-stage framework that leverages 2D photometric supervision to refine and correct 3D motion priors. |
Lei Zhu; Lijian Lin; Ye Zhu; Xuehan Hou; Jiahao Wu; Yu Li; Yunfei Liu; Jie Chen; | code |
| 481 | PIC: Revisiting INR for Image Coding with Fast Encoding and Sub-Millisecond Decoding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, wepropose a feedforward INR image coding architecture, Practical INRImage Codec (PIC), that computes all the necessary information forINR network in a single forward pass, achieving an encoding speed of20 FPS. |
Xiang Liu; Jinxiang Wang; Bin Chen; Zimo Liu; Mingyao Hong; Jiawei Li; Yaowei Wang; Shu-Tao Xia; | code |
| 482 | DeMuS: Learning Decoupled Matching and Scoring for Batch Zero-Shot Industrial Anomaly Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We investigate a highlypractical Batch Zero-Shot IAD setting, strictly constrained by smallintra-batch references and common pose variations. |
Zhengyang Zhao; Hailong Sun; Binhang Qi; Hongrui Yu; Zhongchi Wang; Hang Xu; | code |
| 483 | VisWordBench: Bridging The Gap in Cross-modal Reasoning for Multimodal Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The bench-mark is designed to evaluate not only fundamental perceptual under-standing and world knowledge but also deep cross-modal reasoning skillsthat require consistently integrating visual and linguistic information.Through evaluations on VisWordBench, we find that current MLLMs ex-hibit under-diversified hypothesis search during reasoning. |
Tianshu Zhang; Junzhe Chen; Yean Cheng; Demin Zhu; Haoze Zheng; Lijie Wen; | code |
| 484 | SlowBA: An Efficiency Backdoor Attack Towards VLM-based GUI Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce SlowBA, a novel backdoor attack that targets the responsiveness of VLM-based GUI agents. |
Junxian Li; Tu Lan; Haozhen Tan; Yan Meng; Haojin Zhu; | code |
| 485 | NEOMAP: Novel-View Synthesis Via Noise Initialization By Manifold Alternating Projection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce NeoMap, a novel training-free frame-work designed to locate high-fidelity, view-consistent novel view solutionsfrom general pre-trained video models. |
Jinxi Li; Tianyi Zhang; Yafei YANG; Zihui Zhang; Peng Huang; Koon Lin; Bo Yang; | code |
| 486 | Posterior Augmented Flow Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Posterior Augmented Flow Matching (PAFM), a theoretically grounded generalization of FM that replaces single-target supervision with an expectation over an approximate posterior of valid target completions for a given intermediate state and condition. |
George Stoica; Sayak Paul; Matthew Wallingford; Abhay Nori; Vivek Ramanujan; Winson Han; Ali Farhadi; Ranjay Krishna; Judy Hoffman; | code |
| 487 | OBBSeg: Irregular Lesion Segmentation Under Oriented Bounding Box Annotations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose OBBSeg, an intermediate super-vision paradigm guided by Oriented Bounding Boxes (OBBs) thatbridges the gap between full and weak supervision. |
Jun Wei; Xinchang Liu; Yu Liu; Chuhua Yang; Shuhui Wang; Hui Huang; | code |
| 488 | Amplify, Aggregate, and Adjust: VideoMAE-based Holistic-Subtle Aggregation for Micro-Action Recognition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Motivated by this, we first establish a strong MAR baseline by task-adaptively fine-tuning VideoMAE. To further enhance its capacity to perceive subtle motion cues, we propose A3-MAE, a VideoMAE-based holistic-subtle collaborative framework that integrates Amplification, Aggregation, and Adjustment in a unified design. |
Yan Zhang; Nan Pu; Wenjing Li; Zhun Zhong; Meng Wang; | code |
| 489 | GeoMix: Descriptor-Free Visual Localization Via Global Context and Multi-Detector Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We further observe that descriptorfree matching naturally enables multi-detector training, as heterogeneous keypoints can be optimized in a shared geometry-only space without aligning descriptor spaces. Building on these insights, we propose GeoMix, a descriptor-free 2D-3D matching framework that strengthens geometric discriminability at three levels. |
Yejun zhang; Xinjue Wang; Zihan Wang; Esa Rahtu; Juho Kannala; | code |
| 490 | SDUM: A Scalable Deep Unrolled Model for Universal Cardiac MRI Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present the Scalable Deep Unrolled Model (SDUM), which in-tegrates a Restormer-based unrolled reconstructor, per-cascade coil sen-sitivity estimation, sampling aware weighted data consistency, and uni-versal conditioning on cascade index and acquisition metadata. |
Puyang Wang; Pengfei Guo; Keyi Chai; Jinyuan Zhou; Daguang Xu; Shanshan Jiang; | code |
| 491 | Memory-Supported Synergistic Adaptation for Training-Free Test-Time Medical Image Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Test-time adaptation (TTA) aims to mitigate distributionshifts by adapting models with unlabeled target data at inference time.While TTA with vision-language models (VLMs) has shown promisingresults in classification, extending it to medical image segmentation re-mains challenging. |
Lingrui Li; Nan Pu; Dong Zhao; Wenjing Li; Andrew French; Xin Chen; Zhun Zhong; | code |
| 492 | ESC: Emotional Self-Correction for Reliable Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing self-correction methods mitigate these is-sues but typically rely on post-training or carefully engineered feedback,incurring high computational cost. In this work, we revisit this challengethrough the lens of emotional cues, asking whether they can activate la-tent self-correction behaviors in VLMs without additional training. |
Tien-Huy Nguyen; Nhat Nguyen; Nhat-Huy Nguyen; Hung Nguyen; Huy Nguyen; Thanh-Huy Nguyen; Cuong Nguyen; Hoang Le; Dat Nguyen; Phat Huynh; Min Xu; Ulas Bagci; | code |
| 493 | SparseDriveV2: Scoring Is All You Need for End-to-End Autonomous Driving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we start with a systematic scaling study of Hydra-MDP, a representative scoring-based method, revealing that performance consistently improves as trajectory anchors become denser, without exhibiting saturation before computational constraints are reached. |
Wenchao Sun; Xuewu Lin; Keyu Chen; Zixiang Pei; Xiang LI; Yining Shi; Sifa ZHENG; | code |
| 494 | Multi-Channel Uncertainty-Weighted Score Matching for Conditional Diffusion in Medical UDA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose UPDiff-UDA, a unified UDA framework whose core is an uncertainty-guided training objective for target-domain conditional diffusion. |
CHEN LI; Meilong Xu; Xiaoling Hu; Weimin Lyu; Chao Chen; | code |
| 495 | Mind2Cloud: EEG-to-Point Cloud Generation with Two-Granularity Diffusion Decoding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Mind2Cloud, anovel EEG-to-point-cloud generation framework based on two-granularitydiffusion decoding. |
Yongyi Lu; Xiongfeng Huang; Zhijing Yang; | code |
| 496 | EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Tofurther preserve the temporal consistency of learned robotic motion la-tent, we introduce an optimal-transport-based consistency module thatexplicitly enforces motion fidelity and inter-frame coherence. |
Jiayi Luo; Hanxin Zhu; Chen Gao; Jiankun Wang; Cong Wang; Tianyu He; Jianxin Li; Zhibo Chen; | code |
| 497 | Diffusion Image Generation with Explicitly Modeling of Data Manifold Geometry Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Image generative models aim to sample data points fromthe underlying data manifold, a task that requires learning and decod-ing a dense, low-dimensional, and compact parameterization space. Toachieve this, we propose the Data Manifold-aware Image diffusioN moDel(MIND), a novel framework that explicitly models manifold geometry byintegrating discrete patch tokenization into the score function of a con-tinuous diffusion model. |
Duoduo Xue; Zhiyu Zhu; Junhui Hou; | code |
| 498 | StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Hence, whileStochasT draws on Dropout and stochastic depth for ResNets, it doesnot actually drop anything to maximize the utility of the training data.Furthermore, we introduce a challenging, benchmark-agnostic evaluationmechanism based on the Balanced Latin Square to measure LVLMs’ ro-bustness under varying contextual dependencies. |
Yuan Qing; Chengzhi Mao; Boqing Gong; | code |
| 499 | E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, two major challenges remain unsolved: (1) reasoning can effectively improve the performance of the policy, but its scaling mechanism has seldom been studied; (2) historical information is essential, as embodied tasks are inherently longhorizon and sequential, making sole reliance on current observations for action scaling inadequate due to the lack of historical context utilization. To address these challenges, we introduce E-TTS, a modular and plugand-play Embodied Test-Time Scaling framework that unifies reasoning and action scaling for robotic manipulation via history-aware iterative refinement with vision-language verifiers. |
Wen Ye; Peiyan Li; Tingyu Yuan; Yuan Xu; Xiangnan Wu; Chaoyang Zhao; Jing Liu; Nianfeng Liu; Yan Huang; Liang Wang; | code |
| 500 | Dual Distribution Estimation for Zero-shot Noisy Test-Time Adaptation with VLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing zero-shot NTTA approaches typicallyrely on test-time discriminative training, leading to overconfident mis-classifications and significantly degraded inference efficiency. To addressthese limitations, we propose a novel framework named Dual Distribu-tion Estimation (DDE), shifting the zero-shot NTTA paradigm frominstance-level learning to training-free Gaussian distribution modeling.DDE incorporates two novel modules: Positive Feature Distribution Es-timation (PFDE) and Negative Label Distribution Estimation (NLDE). |
Wenjie Zhu; Yabin Zhang; Liang Xu; Xin Jin; Wenjun Zeng; Lei Zhang; | code |
| 501 | Consistent Feature Transport for Image Relighting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To support complex lighting scenar-ios, we construct a large-scale portrait relighting dataset with diverserelighting effects. |
Bohan Zhang; huanweiliang huanweiliang; Yuhan He; Hongteng Xu; Xiaochao Qu; Luoqi Liu; Dixin Luo; Ting Liu; | code |
| 502 | ContextFlow: In-Context Flow Matching for Robot Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Meanwhile, flow-matching policies have been explored for continuous robot control and can help mitigate compounding errors; however, in-context imitation learning within a flow-matching framework remains underexplored. To address these limitations, we introduce ContextFlow, a conditional flow-matching model that learns continuous action distributions for in-context imitation learning. |
Jian Ding; Xianjie DAI; Roei Herzig; Nussair Hroub; Jinjie Mai; Dengxin Dai; Bernard Ghanem; Mohamed Elhoseiny; | code |
| 503 | From One-to-One to Many-to-Many: Dynamic Cross-Layer Injection for Deep Vision-Language Fusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Vision-Language Models (VLMs) create a severe visual fea-ture bottleneck by using a crude, asymmetric connection that links onlythe output of the vision encoder to the input of the large language model(LLM). |
Cheng Chen; Yuyu Guo; Pengpeng Zeng; Jingkuan Song; Peng Di; Hang Yu; Lianli Gao; | code |
| 504 | FreqPhys: Repurposing Implicit Physiological Frequency Prior for Robust Remote Photoplethysmography Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, most existing methods predominantly rely ontime-domain modeling, making them vulnerable to motion artifacts andillumination fluctuations, where weak physiological clues are easily over-whelmed by noise. To address these challenges, we propose FreqPhys,a frequency-guided rPPG framework that explicitly leverages physio-logical frequency priors for robust signal recovery. |
Wei Qian; Dan Guo; Jinxing Zhou; Bochao Zou; Zitong Yu; Meng Wang; | code |
| 505 | RAE-NWM: Navigation World Model in Dense Visual Representation Space Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To better understand the propagation character-istics of different representations, we conduct a linear dynamics probeand observe that dense DINOv2 features exhibit stronger linear pre-dictability for action-conditioned transitions. Motivated by this obser-vation, we propose the Representation Autoencoder-based NavigationWorld Model (RAE-NWM), which generatively models navigation dy-namics in a dense visual representation space. |
Mingkun Zhang; wangtian shen; Fan Zhang; Haijian Qin; Zihao Pei; Ziyang Meng; | code |
| 506 | MegaFlow: Zero-Shot Large Displacement Optical Flow Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Accurate estimation of large displacement optical flow re-mains a critical challenge. Existing methods typically rely on iterativelocal search or/and domain-specific fine-tuning, which severely limits theirperformance in large displacement and zero-shot generalization scenarios.To overcome this, we introduce MegaFlow, a simple yet powerful model forzero-shot large displacement optical flow. |
Dingxi Zhang; Fangjinhua Wang; Marc Pollefeys; Haofei Xu; | code |
| 507 | Learning to Stylize By Learning to Destylize: A Scalable Paradigm for Supervised Style Transfer Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introduces a scalable paradigm for supervisedstyle transfer by inverting the problem: instead of learning to stylizedirectly, we learn to destylize, reducing stylistic elements from artisticimages to recover their natural counterparts and thereby producing au-thentic, pixel-aligned training pairs at scale. |
Ye Wang; Zili Yi; Yibo Zhang; Peng Zheng; Xuping Xie; Jiang Lin; Yijun Li; Yilin Wang; Rui Ma; | code |
| 508 | Don’t Let The Video Speak: Audio-Contrastive Preference Optimization for Audio-visual Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While Audio-Visual Language Models (AVLMs) have achievedremarkable progress over recent years, their reliability is bottleneckedby cross-modal hallucination. A particularly pervasive manifestation isvideo-driven audio hallucination: models routinely exploit visual short-cuts to hallucinate expected sounds, discarding true auditory evidence.To counteract this deeply ingrained visual dominance, we propose Audio-Contrastive Preference Optimization (ACPO). |
Ami Baid; Zihui Xue; Kristen Grauman; | code |
| 509 | ModuSeg: Decoupling Object Discovery and Semantic Retrieval for Training-Free Weakly Supervised Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Al-though foundation models show immense potential, many approachesstill follow the tightly coupled optimization paradigm, struggling to ef-fectively alleviate pseudo-label noise and often relying on time-consumingmulti-stage retraining or unstable end-to-end joint optimization. To ad-dress the above challenges, we present ModuSeg, a training-free weaklysupervised semantic segmentation framework centered on explicitly de-coupling object discovery and semantic assignment. |
Qingze He; Fagui Liu; Dengke Zhang; Qingmao Wei; Quan Tang; | code |
| 510 | QuantV2X: A Fully Quantized Multi-Agent System for Cooperative Perception Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce QuantV2X, the first fully quantized multi-agent system designed specifically for efficient and scalable deployment of multi-modal, multi-agent V2X cooperative perception. |
Seth Zhao; Huizhi Zhang; Zhaowei Li; Juntong Peng; Anthony Chui; Zewei Zhou; Zonglin Zonglin; Hao Xiang; Zhiyu Huang; Fujia Wang; Ran Tian; Chenfeng Xu; Bolei Zhou; Jiaqi Ma; | code |
| 511 | UniStitch: Unifying Semantic and Geometric Features for Image Stitching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These two lines of research have largely diverged along separate evolution, with virtually no meaningful convergence to date. In this paper, we take a pioneering step to bridge this gap by unifying semantic and geometric features with UniStitch, a unified image stitching framework from multimodal features. |
Yuan Mei; Lang Nie; Kang Liao; Yunqiu Xu; Chunyu Lin; Bin Xiao; | code |
| 512 | Analyzing and Improving Training-Free Fast Sampling of Text-to-Image Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing training-free sampling acceleration methods are typically developed independently, leaving the overall performance and compatibility among these methods unexplored. In this paper, we bridge this gap by systematically elucidating the design space, and our comprehensive experiments identify the sampling time schedule as the most pivotal factor. |
Zhenyu Zhou; Defang Chen; Siwei Lyu; Chun Chen; Can Wang; | code |
| 513 | From Open Loop to Closed Loop: A Test-Time Iterative Optimization Framework for Reference-Consistent Image Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: By treating the pre-trained generative model as a control plant,our framework employs a sensor-controller architecture driven by a mod-ified Proportional-Integral-Derivative (PID) algorithm. |
Baixuan Zhao; Xinyu Zhang; 华渝 郑; Shuaicheng Liu; Xiongkuo Min; Guangtao Zhai; Xiaohong Liu; | code |
| 514 | Q-REAL: Towards Naturalness and Distortion Evaluation for AI-Generated Content Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Tobridge this gap, we introduce Q-Real, a fine-grained quality assessmentdataset specifically designed for AI-generated images.This dataset com-prises 10K images generated by multiple models, annotated along twocritical dimensions, naturalness and distortion which are widely re-garded as the most significant aspects of AI-generated image quality.For each image, we localize major entities and provide a set of judgmentquestions and attribution descriptions along these dimensions to facili-tate comprehensive evaluation. |
Shushi Wang; Zicheng Zhang; Chunyi Li; Wei Wang; Liya Ma; Xiaoyu Li; Fengjiao Chen; Xuezhi Cao; Guangtao Zhai; Xiaohong Liu; | code |
| 515 | FUMO: Prior-Modulated Diffusion for Single Image Reflection Removal Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Single image reflection removal (SIRR) is challenging in realscenes, where reflection strength varies spatially and reflection patternsare tightly entangled with transmission … |
Telang Xu; Chaoyang Zhang; Guangtao Zhai; Xiaohong Liu; | code |
| 516 | A²-Edit: Precise Reference-Guided Image Editing of Arbitrary Objects and Ambiguous Masks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose A2-Edit, a unified inpainting framework for arbitrary object categories, which allows users to replace any target region with a reference object using only a coarse mask. |
华渝 郑; Guangzhao Li; Baixuan Zhao; Siqi Luo; Hantao Jiang; Guangtao Zhai; Xiaohong Liu; | code |
| 517 | Rule-VLN: Bridging Perception and Compliance Via Semantic Reasoning and Geometric Rectification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: , frequently overlooking subtleregulatory constraints. To bridge this gap, we establish Rule-VLN, the firstlarge-scale urban benchmark for rule-compliant navigation. |
Jiawen Wen; Penglei SUN; Wenjie Zhang; Suixuan QIU; Weisheng Xu; Xiaofei Yang; Xiaowen Chu; | code |
| 518 | Continuous Heart Rate Variability Estimation from Egocentric Systems for Skill Assessment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Egocentric vision systems capture human behavior from visi-ble cues, but overlook physiological indicators of autonomic states suchas stress, engagement, and attention. Heart rate … |
Berken Utku Demirel; Christian Holz; | code |
| 519 | LiDAR-EVS: Enhance Extrapolated View Synthesis for 3D Gaussian Splatting with Pseudo-LiDAR Supervision Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To enable reliable simulation of LiDAR along unseen driving trajectories without external multi-pass data, we present LiDAR-EVS, a lightweight framework for robust extrapolated-view LiDAR simulation in autonomous driving. |
Yiming Huang; Xin Kang; Sipeng Zhang; Hongliang Ren; Weihua Zhang; Junjie Lai; | code |
| 520 | Pixel Ignores, Superpixel Sees: Adverse Weather Image Restoration Via Semantic-Center SSM Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose SSR, a Semantic-center guildedState space model for image Restoration. |
Dayu Li; shihao zhou; SHU LEIZHI; Jin Wu; Chi Man VONG; Jufeng Yang; | code |
| 521 | Event Stream-based Sign Language Translation: A High-Definition Benchmark Dataset and A Novel Baseline Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper proposes the use of bio-inspired event cameras to alleviate the aforementioned issues. |
Shiao Wang; Xiao Wang; Duoqing Yang; Yao Rong; Fuling Wang; Jianing Li; Lin Zhu; Bo Jiang; | code |
| 522 | Geometry-Preserving Image Generation for 6D Object Pose Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose an image generation pipeline that preserves geometric consistency for 6D object pose estimation.First, we create a largescale synthetic dataset by rendering objects from multiple viewpoints using their 3D meshes. |
Jiafeng Zhang; Rui Song; Zhengtai Zhang; Jiaojiao Li; Kailang Cao; Lizhang Peng; David Ferstl; Yinlin Hu; | code |
| 523 | ThermoGS: Decoupling Physical Surface Attributes for Spatio-Temporal Thermal Field Emulation Via 4D Gaussian Splatting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To bridge this gap, we propose ThermoGS,a unified framework that couples 4D Gaussian Splatting (4DGS) withfundamental thermodynamic principles to achieve physically-consistentthermal field emulation.Furthermore,we introduce the Thermo Dataset, a comprehensive, drone-captured col-lection featuring synchronized multi-modal data and diverse environmen-tal parameters. |
Kun Yang; Yuxiang Liu; Zeyu Cui; Shen Yan; Maojun Zhang; Yu Liu; Xue Wang; Qing Wang; | code |
| 524 | AV2T-Gen: Aerial Visible to Thermal Generation with Environment and Vehicle State Guidance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose AV2T-Gen, an innovativeenvironment and vehicle guided visible-to-thermal generation frameworkbuilt upon the Instruct-Pix2Pix Diffusion Models for aerial imagery,which distinguishes itself from existing methods by explicitly incorpo-rating environment parameters and vehicle state information for thefirst time. |
Kun Yang; Yuxiang Liu; Yihan Wang; Shen Yan; Maojun Zhang; Yu Liu; Xue Wang; Qing Wang; | code |
| 525 | Rectified Embedding Flow Learning for Aerial Multi-view Geo-localization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To overcome the limitations of single-retrieval systems caused by modality-specific uncertainty in open environments, this paper introduces the unified Aerial Multi-view Geo-localization (AMGL) task. |
Hao Ruan; Jinliang Lin; Yingxin Lai; Zhiming Luo; Shaozi Li; Yu Zang; Cheng Wang; | code |
| 526 | FitControler: Toward Fit-Aware Virtual Try-On Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce fit-aware VTON and present FitControler, a learnable plug-in that can seamlessly integrate into modern VTON models to enable customized fit control.In particular, we build a fit-aware VTON dataset termed Fit4Men, including 13,000 body-garment pairs of different fits, covering both tops and bottoms, and featuring varying camera distances and body poses. |
Lu Yang; Yicheng Liu; Letian Zhou; Yanan Li; Xiang Bai; Hao Lu; | code |
| 527 | Benchmarking Dynamic Affective Reasoning: A Viewer-Centric Video Emotion Dataset Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Based on DAR, we formally define three challeng-ing tasks: affective segmentation, fine-grained emotion classification, andaffective reasoning. Complementing this benchmark, we propose DAR-R1, a two-stage framework that combines supervised fine-tuning withGroup Relative Policy Optimization. |
Zhiyan Zhang; Peipei Song; Jinpeng Hu; Jingyang Jia; Xun Yang; Xiaojun Chang; | code |
| 528 | RayMap3R: Inference-Time RayMap for Dynamic 3D Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, wepropose RayMap3R, a training-free streaming framework for dynamicscene reconstruction. |
Feiran Wang; Zezhou Shang; Gaowen Liu; Yan Yan; | code |
| 529 | SynHMR: Synergistic Joint-Mesh Modeling for LiDAR-based Human Mesh Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Althoughrecent methods have made promising progress, most still follow a decou-pled pipeline that first estimates skeletal joints and then reconstructsthe mesh, leading to error propagation under sparse observations. Toaddress this issue and fully exploit both the topological guidance of theskeleton on the mesh and the spatial constraints of the surface mesh onthe skeleton, we propose SynHMR, a Synergistic Joint-Mesh Modelingframework for robust LiDAR-based HMR. |
Rui Shi; Xiaoqi An; Lin Zhao; Di Wang; Chen Gong; Le Zhang; | code |
| 530 | 360° Image Perception with MLLMs: A Comprehensive Benchmark and A Training-Free Method Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Using 360Bench, we systematically evaluate seven MLLMs and sixenhancement methods, revealing their shortcomings in 360° image per-ception. To address these challenges, we propose Free360, a training-free scene-graph-based framework for high-resolution 360◦ VQA. |
Huyen Thi Thanh Tran; Van-Quang Nguyen; Farros Alferro; Kang-Jun Liu; Takayuki Okatani; | code |
| 531 | Fast Dynamic Prototypes for Unsupervised Anomaly Detection and Localization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In thispaper, we propose a simple yet effective scheme, termed Fast DynamicPrototype (FDP), to rapidly generate high-quality reconstructed proto-types for Multi-class UAD. |
Mingliang Li; Hanxi Li; Lin Yuanbo Wu; Changhong Liu; Xiaowei Zhao; | code |
| 532 | Looking Back and Forth: Cross-Image Attention Calibration and Attentive Preference Learning for Multi-Image Hallucination Mitigation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We attribute this issue to limitations in existing attention mechanisms and insufficient cross-image modeling. Inspired by this, we propose a structured hallucination mitigation framework involving Cross-Image Attention calibration and Preference Learning (CAPL). |
Xiaochen Yang; Hao Fang; Jiawei Kong; Yaoxin Mao; Bin Chen; Shu-Tao Xia; | code |
| 533 | Practice Makes Perfect: From Explicit Decomposition to Reinforced Latent Planning in Text-to-Human Motion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Recent work has shown that Chain-of-Thought (CoT) rea-soning improves text-to-motion generation, yet generating explicit rea-soning tokens at inference time introduces significant … |
Ronghao Yu; Yang Liu; Juncheng Wang; Chao Xu; Yimo Shao; Baigui Sun; Yong Liu; Shan Luo; | code |
| 534 | Triangular Consistency As A Universal Constraint for Learning Optical Flow Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose triangular consistency as a first-principled con-straint for optical flow, which is agnostic to network architecture, super-vision type, and dataset, and applies to both image-pair and multi-framesettings. |
Yi Xiao; Carlos Rodriguez Coronel; Jing Zhan; Haniyeh Oskouie; Alex Wong; DONG LAO; | code |
| 535 | FSDC-DETR: A Frequency-Spatial Domain Collaborative DETR for Small-Object Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Small object detection (SOD) remains a challenging task inreal-world applications. Despite recent advances, existing detectors re-main limited by rigid processing that entangle … |
Aiwen Liu; Chengguang Zhu; Gang Wang; Dandan Zhu; Haodong Lin; Yan Wang; Huiyu Zhou; Zhengyi Pan; | code |
| 536 | MiNQVIS: Mitigating Noisy Queries for Robust Online Video Instance Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce MiNQVIS, the first noise-aware framework for on-line VIS that stabilizes both training and inference. |
Jie Qiao; Jianxu Chen; Xiaowei Xu; | code |
| 537 | ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce a simple but effective framework, calledELVA, a novel rule-based RL framework that mitigates grain blindnessthrough ranking-driven MLLMs. |
Yuhan Liu; Pei Fu; Hang Li; Yukun Qi; Chao Jiang; Jingwen Fu; Zhen Liu; Bin Qin; Zhenbo Luo; Jian Luan; Jingmin Xin; | code |
| 538 | Intermediate Text Representation Guided Text-to-Image Generation for Enhancing One-and-Only Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We thenpropose Intermediate Text Representation (IR)-guided diffusion, whichinjects intermediate hidden states of the text encoder into the condi-tioning signal during early denoising steps, recovering suppressed con-cepts without any additional training, optimization, or external models.To systematically evaluate the challenging task of aligning generativeoutputs with unusual prompts for OAO objects, we introduce OAO-AttackBench, a benchmark comprising counterfactual prompts that di-rectly conflict with the core visual identity of OAO objects. |
Soyoun Won; Aryan Yazdan Parast; Basim Azam; Jean Honorio; NAVEED AKHTAR; | code |
| 539 | Score-Based Matching with Target Guidance for Cryo-EM Denoising Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a score-based denoising framework for cryo-EM that learns the clean-data score to recover particle signals while better preserving structural information. |
Xiaoqi Wu; Xueying Zhan; Wen Li; Junhao Wu; Xin Huang; Ke Ni; Min Xu; | code |
| 540 | Online Versatile Incremental Learning: Towards Class and Domain-Agnostic Adaptation at Any Time Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work in-troduces Online VIL (Online Versatile Incremental Learning),a novel scenario where class concepts and visual domains evolve simul-taneously online without explicit boundaries. To better adapt to thechallenges of such dynamic environments that more closely resemblereal-world conditions, we propose a novel framework TopFlow, Topologypreservation with Flow matching representation that contains two com-plementary mechanisms: Domain-agnostic Flow Matching (DFM)and Global Topology Preservation (GTP). |
Jaeho Lee; Jun-Yeong Moon; Min-Yeong Park; Jung Uk Kim; Gyeong-Moon Park; | code |
| 541 | DARE to Mitigate Hallucination: Dual-path Auto-Regressive-aware Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we analyzethe relationship between TF-based editing and AR generation behaviorand find that TF-based editing alone may be insu!cient to captureboth decoding dynamics and multimodal interactions associated withhallucinations. |
Jaeho Lee; Jeongeun Lee; Gyeong-Moon Park; | code |
| 542 | GAINS: Gaussian-based Inverse Rendering from Sparse Multi-View Captures Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce GAINS (Gaussian-based Inverse rendering from Sparse multi-view captures), a two-stage inverse rendering framework that leverages foundation models as priors to stabilize geometry and material estimation. |
Patrick Noras; Jun Myeong Choi; Didier Stricker; Pieter Peers; Roni Sengupta; | code |
| 543 | DeltaDeno: Zero-Shot Anomaly Generation Via Delta-Denoising Attribution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We tackle the settingwhere no real anomaly samples or training are available. |
Chaoran Xu; Chengkan Lv; Qiyu Chen; Yunkang Cao; Feng Zhang; Zhengtao Zhang; | code |
| 544 | MotionDreamer: Universal Skeletal Motion Generation for 3D Rigged Shapes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present SkelMo, a diffusion-based framework designed for category-agnostic skeletal animation generation from 2D video guidance. |
Ye Tao; Yuxin Yao; Kendong Liu; Dapeng Wu; Junhui Hou; | code |
| 545 | A Comprehensive Analysis About Unsupervised Outlier Detection for Images Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Evaluated on 3 domains and 14 benchmark datasets, our proposed solution achieves state-of-the-art performance and significantly outperforms existing methods. More importantly, we show its plug-andplay property that can be integrated into diverse visual applications to improve their robustness, such as image classification and 3D reconstruction. |
Zhonghang Liu; Siyuan Chen; Jingwen Yu; Changshuo Wang; Kunyang Li; Jiangbo Lu; | code |
| 546 | Cooking Beyond Frames: A Stereo Event Camera Dataset in The Kitchen Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce EventKitchen, a large-scalestereo event camera benchmark dataset of human cooking activities inthe kitchen. |
Chengming Feng; Hesam Araghi; Liming Zheng; Julien Dupeyroux; Xucong Zhang; Jan van Gemert; Nergis Tomen; | code |
| 547 | TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our analysis reveals that 1) semantic information appears in earlier blocks and finer details are rendered in later blocks, 2) removing specific blocks is usually less disruptive than disabling text conditions, and 3) enhancing textual conditions in selective blocks improves semantic attributes. Building on these observations, we propose TexTailor, a novel inference-time method for tailoring block-wise textual guidance. |
Binglei Li; Mengping Yang; Zhiyu Tan; Junping Zhang; Hao Li; | code |
| 548 | OCTA-SOT: Online Cross-Modal Trajectory Adjustment for RGBT Anti-UAV Single Object Tracking Under Spatio-Temporal Misalignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Based on thistheoretical foundation, we propose OCTA-SOT, an Online Cross-modalTrajectory Adjustment Single Object Tracking framework. |
Xiaokang Liu; Qi Jia; Jinrui Wang; Chengzhou Li; Yu Liu; Weimin Wang; | code |
| 549 | CanoVerse: 3D Object Scalable Canonicalization and Dataset for Generation and Pose Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This persistent misalignment suppresses pose-consistent generation, and blocks the emergence of stable directional semantics. To address this issue, we construct CanoVerse, a massive canonical 3D dataset of 320K objects over 1,156 categories – an order-ofmagnitude increase over prior work. |
Li Jin; Yuchen Yang; Weikai Chen; Yujie Wang; Dehao Hao; Tanghui Jia; Yingda Yin; Zeyu HU; Runze Zhang; Keyang Luo; Li Yuan; Long Quan; Xin Wang; Xueying Qin; | code |
| 550 | Background Blurring Matters: Improving Visual Grounding By Merging Text-Irrelevant Tokens Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose a novelToken Blurring (ToB) module, which dynamically merges image tokensbased on the pair-wise visual similarity between them and their tex-tual relevance with input expressions. |
Ruilin Yao; Shengwu Xiong; Shanshan Yang; Tianyu Zou; Shili Xiong; Yi Rong; | code |
| 551 | Benchmarking Vision-Language Models for Microscopic Plant Image Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing benchmarks for vision-language models primarilyfocus on macroscopic plant imagery, while the microscopic domain re-mains underexplored. To address this gap, we present PlantMicro, a com-prehensive benchmark for evaluating vision–language models (VLMs) inmicroscopic plant imagery. |
Tianqi Wei; Xin Yu; Zhi Chen; Scott C Chapman; Zi Helen Huang; | code |
| 552 | SupGRPO: Enhancing GRPO with Matching-based Online SFT for Text Spotting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To compensate for each other’s shortcomings, we introduce a joint training strategy, SupGRPO, which simultaneously optimizes the model using both SFT and GRPO. |
Xudong Xie; Yuzhe Li; Jing Shi; Zhifei Zhang; Curtis Wigington; Zhaowen Wang; | code |
| 553 | FED-Bench: A Cross-Granular Benchmark for Disentangled Evaluation of Facial Expression Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To bridge these gaps, we propose FED-Bench, a comprehensive benchmark featuring rigorous testing and an accurate evaluation suite.First, we carefully construct a benchmark of 747 triplets through a cascaded and scalable pipeline, each comprising an original image, an editing instruction, and a ground-truth image for precise evaluation.Finally, leveraging the scalable characteristic of introduced benchmark engine, we provide a 20k+ in-the-wild facial training set and demonstrate its effectiveness by fine-tuning a baseline model that achieves significant performance gains. |
Fengjian Xue; Xuecheng Wu; Heli Sun; Yunyun Shi; Shi Chen; Liangyu Fu; Jinheng Xie; Dingkang Yang; Hao Wang; Junxiao Xue; Liang He; | code |
| 554 | OREO: Fidelity Alignment in 3D Generation Via On-the-fly Rendering-Editing Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite recent advancements in 3D generation, models oftenstruggle to produce assets with high visual fidelity. To bridge this gap,we propose OREO, an alignment framework that enhances the realismof 3D generators by leveraging rich 2D diffusion priors. |
Zhiyuan MA; WENBO HU; Wang Zhao; Pengfei Wang; Ying Shan; Lei Zhang; | code |
| 555 | HyFL-CLIP: Hyperbolic Fine-Tuning of CLIP for Robust Long-Context Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Weattribute this limitation to the Euclidean contrastive objective, whichenforces strict one-to-one matching and lacks explicit mechanisms formodeling hierarchical relationships between global context and its con-stituent elements. To address this issue, we propose HyFL-CLIP, a hy-perbolic fine-tuning framework that distills the well-established text-image alignment learned in Euclidean CLIP into hyperbolic space viacross-manifold similarity distillation, leveraging its geometry to capturehierarchical and entailment relations. |
jiha jang; Hayeon Kim; Junghun James Kim; Chulwon Lee; Se Young Chun; | code |
| 556 | Point2Pose: Occlusion-Recovering 6D Pose Tracking and 3D Reconstruction for Multiple Unknown Objects Via 2D Point Trackers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Point2Pose, a model-free method for causal 6Dpose tracking of multiple rigid objects from monocular RGB-D video.Initialized only from sparse image points on the objects, our approachtracks multiple unseen objects without requiring object CAD modelsor category priors.Alongside the method, we introduce a new multi-objecttracking dataset comprising both simulation and real-world sequences,with motion-capture ground truth for evaluation. |
Tzu-Yuan Lin; Ho Lee; Kevin Doherty; Yonghyeon Lee; Sangbae Kim; | code |
| 557 | ResilPhase: Plug-and-Play Phase Mapping and Noise-Resilient Macro-Trajectory Extrapolation for Diffusion Acceleration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Thus, we introduce a derivative-free barycentric Lagrangeextrapolator to effectively bypass derivative instability and approxima-tion error. |
Qicheng Zhao; Yu Li; Qi Sun; Zheyu Yan; | code |
| 558 | VSDiffusion: Taming Ill-Posed Shadow Generation Via Visibility-Constrained Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Generating realistic cast shadows for inserted foreground ob-jects is a crucial yet challenging problem in image composition, wheremaintaining geometric consistency between objects and their shadows incomplex scenes remains difficult due to the ill-posed nature of shadowformation. To address this challenge, we propose VSDiffusion, a visibility-constrained two-stage framework that narrows the solution space by in-corporating visibility priors. |
Jing Li; Jing Zhang; | code |
| 559 | 3D Gaussian Texture for Real-time Mesoscale Appearance Synthesis and Rendering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a novel framework for modeling, synthesis, and rendering of mesoscale appearances based on 3D Gaussian Splatting (3DGS) for better quality and performance. |
Xiang Chen; Jia Li; Lu Wang; Beibei Wang; | code |
| 560 | ReViV: Reconstructing The Viewer and The View in 4D from Monocular Egocentric Video Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing approachesoften rely on auxiliary inputs such as pre-computed camera trajecto-ries, treat scene perception and human ego-motion modeling as separateproblems despite their strong interdependency, and suffer from slow in-ference time. To address these limitations, we present ReViV, the firstunified framework for holistic egocentric 4D reconstruction that extractsboth viewer and view dynamics from a single monocular RGB video.We formulate the task as learning the full joint probability distribu-tion over multimodal signals, including RGB video, camera trajectory,gaze direction, full-body motion, hand motion, and depth. |
Xiaozhong Lyu; Gen Li; Zhiyin Qian; Xucong Zhang; Marc Pollefeys; Siyu Tang; | code |
| 561 | OSVE: One Step Video Editing with One Step Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: For temporal coherence,we introduce Unified-Frame Editing (UFE), a technique that concate-nates frame latents to facilitate cross-frame attention in a single gen-eration step. |
Habin Lim; Gyeong-Moon Park; | code |
| 562 | Liquid Fusion of Heterogeneous Representations Towards General Salient Object Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, inspired by the dynamic information propaga-tion of Liquid Neural Networks (LNNs), we introduce a liquid fusion todynamically integrates features from two backbones, including VMambaand ConvNeXt, referred to Liquid Fusion Network (LFNet). |
Ke Chen; Ling Zhou; Yi Liu; Guangqi Jiang; Gengshen Wu; Shoukun Xu; | code |
| 563 | DAP: Doppler-aware Point Network for Heterogeneous MmWave Action Recognition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing mmWave point cloud datasets are limited in scale and mostly collected under homogeneous single-source settings, preventing current methods from handling real-world distribution shifts caused by heterogeneous radar sources, such as different devices and frequency bands. To address this, we introduce UniMM-HAR, the largest and first mmWave point cloud HAR dataset for heterogeneous multi-source scenarios, standardizing three distinct radar configurations to realistically evaluate crosssource generalization. |
Jiaying Lin; Shiman Wu; Jinfu Liu; Can Wang; Mengyuan Liu; | code |
| 564 | StratoSplat: Taming Layered Regularities for Sparse Aerial 3D Gaussian Splatting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present StratoSplat, a framework that exploits the in-herent layered structure of aerial scenes for robust sparse-view 3DGS. |
Zihan Gao; Lingling Li; Licheng Jiao; Fang Liu; Wenping Ma; Yuwei Guo; Shuyuan Yang; | code |
| 565 | EcoVideo: Entropy-Orchestrated Video Generation Paradigm in Cloud-Edge Dynamics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose EcoVideo, an entropy-orchestratedframework for dynamic inter-frame decoupling: early-stage self-attentionentropy provides a training-free estimate of frame-wise information den-sity for frame selection; a cloud large model denoises sparse high-entropykeyframes; and an edge lightweight model reconstructs the remainingframes via motion-aware interpolation with refinement for temporal sta-bility. |
Jiayu Chen; Hengyi Zhang; Maoliang Li; Minyu Li; Zihao Zheng; Xuanzhe Liu; Guojie Luo; Xiang Chen; | code |
| 566 | In-context Region-based Drag: Drag Any Region to Any Shape Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper focuses on region-based drag and introduces anovel In-Context Region-based Drag (ICRDrag) method.To facilitateregion-based drag, we also construct Paired Region Dataset (PRD), alarge-scale dataset with paired masks and images. |
Jiacheng Sui; Tianyu Hao; Bingjie Gao; Li Niu; Guangtao Zhai; | code |
| 567 | Coarse-to-fine Contrast: A Hybrid Self-supervised Method for Non-rigid 3D Shape Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a hybrid self-supervisedmethod based on a coarse-to-fine strategy, which ensures consistencybetween the coarse mapping and the refined correspondence producedby our refinement module. |
Feifan Luo; Ting Li; Zhao Li; Hongyang Chen; | code |
| 568 | LEGO: Leveled Language Gaussian Splatting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce LEGO for advanced open-vocabulary scene un-derstanding. |
Yuning Peng; Haiping Wang; Yuan Liu; Yipeng Lu; Zhen Dong; Bisheng Yang; | code |
| 569 | DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address thisgap, we introduce DF3DV-1K, a large-scale real-world dataset compris-ing 1,048 scenes, each providing clean and cluttered image sets for bench-marking. |
Cheng-You Lu; Yi-Shan Hung; Wei Chi; Hao Ping Wang; Charlie Tsai; Yu-Cheng Chang; Yu-Lun Liu; Thomas Do; Chin-teng Lin; | code |
| 570 | OmniColor: A Unified Framework for Multi-modal Lineart Colorization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Lineart colorization is a critical stage in professional content creation, yet achieving precise and flexible results under diverse user constraints remains a significant challenge. To address this, we propose OmniColor, a unified framework for multi-modal lineart colorization that supports arbitrary combinations of control signals. |
Xulu Zhang; Haoqian DU; Xiao-Yong Wei; Li Qing; | code |
| 571 | EraseSAE: Surgical Concept Erasure in Text-to-Video Diffusion Models Via Sparse Autoencoders Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose EraseSAE, a novel framework that leverages sparse autoencoders to achieve surgical concept erasure in DiT-based T2V diffusion models via a principled decompose-attributeerase pipeline. |
Xinghao Wang; Dong Li; Wei Yu; Yingwei Pan; Tao Gong; Qi Chu; Nenghai Yu; Ting Yao; | code |
| 572 | EventSpecPS: Photometric Stereo with Multispectral Reflectance Using An Event Camera Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose EventSpecPS, an event-based multispectral photometric stereo system. |
Jingqian Wu; Bohan Yu; Jun Hoong Chan; Edmund Lam; Boxin Shi; | code |
| 573 | DiffVP:Differential Visual Semantic Prompting for LLM-Based CT Report Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: While large language models (LLMs) have advanced chestCT report generation, existing methods typically encode 3D volumesholistically, failing to distinguish informative cues from … |
Yuhe Tian; Kun Zhang; Haoran Ma; Rui Yan; Yingtai Li; Rongsheng Wang; S Kevin Zhou; | code |
| 574 | DiGS-Avatar: Single-Image Animatable 3D Human Reconstruction Via UV-Space Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose DiGS-Avatar, which reformulates thistask as an efficient, diffusion-based UV-latent completion task, ensur-ing 3D consistency by design. |
jiakun li; Li Fang; Hao Zhu; Fei Hu; Long Ye; Yuan Zhang; Jinyao Yan; | code |
| 575 | From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence Via CRISP Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current VLM evaluations often conflate language priors with genuine spatial reasoning. To address this, we introduce CRISP, a novel structural-diagnostic evaluation paradigm that assesses visual spatial intelligence through consistency, the alignment between implicit perception and explicit reasoning. |
Zhixing Li; Yinan Yu; | code |
| 576 | Learning with Bilevel-Minimax Optimization for Efficient and Reliable Transfer Attacks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Beyond per-turbation generation, transferability is fundamentally governed by thecoupling of initialization, surrogate adaptation, and gradient dynam-ics. We revisit this challenge from a Bilevel-Minimax perspective andinstantiate it in BMAT (Bilevel-Minimax Adversarial Transfer). |
Yaohua Liu; Yifan Guo; Jiaxin Gao; | code |
| 577 | Progressive Pose-Guided 4D Animal Reconstruction from Monocular Video Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our key insight is that a coarse shape prior suffices when coupledwith a progressive strategy that disentangles articulated pose from non-rigid deformation. |
Siyuan Li; Weiying Chen; Yilin Wang; Xinxin Zuo; Xingyu Li; Li Cheng; | code |
| 578 | SAM+D: Parameter-Efficient Dimensional Lifting of SAM-Family Models Via Depth-Routed LoRA and Depth Shifting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods for adapting 2D foundation models such asSAM to 3D volumes either process slices independently—ignoring inter-slice context—or require substantial architectural changes and retraining.In this paper, we present SAM+D, a parameter-efficient framework thatlifts SAM-family models by one spatial dimension—enabling 3D volu-metric segmentation from 2D SAM and, for the first time via parameter-efficient fine-tuning, end-to-end 4D (3D+T) spatiotemporal segmenta-tion from video-based SAM2—while keeping the vast majority of pre-trained parameters frozen. |
YU SONG; Hao Sun; TENG SHIYU; Ikuko Nishikawa; Yen-wei Chen; | code |
| 579 | Kirin: Animal Motion Generation from In-the-Wild Video Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While human motion can be captured in controlled environments, itis impractical for most animal species, resulting in small, domain-limiteddatasets that restrict downstream applications such as animation. To ad-dress this challenge, we introduce Kirin, a framework that reconstructsmotion from video, learns motion priors at scale, and generates realisticmotion that can be directly applied to animated assets. |
Brian Nlong Zhao; Zhuoyang Pan; James Rehg; Jiajun Wu; Elliott (Shangzhe) Wu; | code |
| 580 | Rethinking Visual Privacy: A Compositional Privacy Risk Framework for Severity Assessment with VLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce the Compositional Privacy RiskTaxonomy (CPRT), a regulation-aware framework that organizes visualattributes according to standalone identifiability and compositional harmpotential.We further construct a taxonomy-aligned dataset of 6.7K imagesand derive compositional risk scores. |
Efthymios Tsaprazlis; Tiantian Feng; Anil Ramakrishna; Sai Karimireddy; Rahul Gupta; Shrikanth Narayanan; | code |
| 581 | Evaluating The Interpretability of Sparse Autoencoders with Concept Annotations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Sparse autoencoders (SAEs) are increasingly used to extract interpretable concepts from vision and vision language models, yet existing evaluation methods largely rely on proxy metrics or qualitative inspection rather than measuring semantic correspondence. We present a human-grounded evaluation framework that quantifies alignment between SAE latents and human-annotated concepts, without requiring user studies, and validate this matching through targeted attribute perturbations. |
Jonas Klotz; Cassio F. Dantas; Pallavi Jain; Diego Marcos; Begüm Demir; | code |
| 582 | TPCNet: A Low-Light Image Enhancement Network Inspired By Triple Physical Constraints Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However,previous Retinex-based algorithms, that consider reflected objects asideal Lambertian ignore specular reflection in the modeling process andconstruct the physical constraints in image space, limiting generalizationof the model. To address this issue, we preserve the specular reflectioncoefficient and reformulate the original physical constraints in theimaging process based on the Kubelka-Munk theory, thereby constructingconstraint relationship between illumination, reflection, and detection, theso-called triple physical constraints (TPCs) theory. |
Jing-Yi Shi; Ming-Fei Li; Ling-An Wu; | code |
| 583 | Segmentation-Guided Homography Estimation for Long-Term Planar Tracking Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present SAM-H – a planarobject tracker that estimates homographies from segmentation mask con-tours via a training-free pipeline. |
Jonáš Šerých; Jiri Matas; | code |
| 584 | LiFlow: Flow Matching for 3D LiDAR Scene Completion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Re-cent methods adopt point-level denoising diffusion probabilistic modelsthat rely on approximations to handle scene-scale data, leading to amismatch between training and inference initial distributions. |
Andrea Matteazzi; Dietmar Tutsch; | code |
| 585 | Gaze-to-text Generation: Beyond Categorical Decoding of Human Attention Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we introduce Gazette, the first gaze-to-text decod-ing framework.We introduce a novel learning problem: decoding gaze intonatural language descriptions of human goals across diverse visual tasks.Unlike prior work, which frames gaze decoding as a discriminative taskover predefined categories, we formulate it as a generative learning prob-lem: training a model to produce free-form descriptions that capture therich nuances and open-ended nature of human intentions beyond fixedlabels. |
Sounak Mondal; Dimitris Samaras; Gregory Zelinsky; Minh Hoai Nguyen; | code |
| 586 | The Telephone Game: Evaluating Semantic Drift in Unified Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To quantify drift,we introduce the Semantic Drift Protocol (SDP), inspired by the Tele-phone Game: starting from a caption or image, we iteratively alternateI2T and T2I over multiple generations and measure how faithfully se-mantics are preserved.To stress-test models beyond COCO-style data,we create a benchmark of 400 image-text pairs sampled from NoCapsand DOCCI that emphasizes novel objects and fine-grained descriptions.Applying SDP to seven recent unified models reveals that drift behaviorvaries dramatically and is not predicted by single-pass scores: BAGELretains high semantic fidelity over multiple generations, while VILA-Uand Janus variants collapse within five generations, despite comparableisolated metrics. |
Sabbir Mollah; Rohit Gupta; Swetha Sirnam; Qingyang Liu; Ahnaf Munir; Shah Mubarak; | code |
| 587 | LangLoc: “Tell Me What You See” Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We tackle fine-grained indoor localization from natural lan-guage: given a free-form description of one’s surroundings, estimate theobserver’s 2D position and heading within a known 3D environment.Language queries are lightweight, privacy-preserving, and need no cam-era – yet prior work stops at coarse scene retrieval and cannot resolve anintra-scene pose. |
Shaurya Kishore Panwar; Roham Zendehdel Nobari; Shirley Lau; Abu Bakr Rahman Shaik; Manuel Günther; Marc Pollefeys; Daniel Barath; | code |
| 588 | Implicit Neural Representation for Spherical Harmonics Reconstruction of Motion-Corrupted Fetal Diffusion MRI Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, in vivo fetal acquisitions are severely corrupted by unpredictable motion, making the data unusable for downstream analysis. To address this challenge, we propose an implicit neural representation framework for fetal diffusion MRI reconstruction. |
Wenxuan Wu; Irina Grigorescu; Ruowen Qu; Jo Hajnal; J-Donald Tournier; Maria Deprez; | code |
| 589 | Hierarchical Style Aggregation for Versatile Chinese Handwriting Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Hierarchical Style Aggregation for Chinese hand-writing generation (HiSAC), a versatile framework that decouples stylemodeling from task-specific generation. |
Jiangpeng Wang; Fei Gao; Nannan Wang; | code |
| 590 | Pol-CACTI: A System and Dataset ForHigh-Speed Polarized Video Compressive Imaging Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While snapshot compres-sive imaging (SCI) enables high-speed video reconstruction from a singlemeasurement, extending it to division-of-focal-plane (DoFP) polarizationsensing introduces a fundamentally more ill-posed inverse problem due tothe coupling between temporal multiplexing and spatial polarization mo-saicing. To address this challenge, we propose Polarization Coded Aper-ture Compressive Temporal Imaging (Pol-CACTI), a hardware–softwareco-designed framework for polarized video SCI. |
Yunfeng Song; Yidong Luo; Ping Wang; xingjian jiang; Xin Yuan; | code |
| 591 | PersonaDrive: Controllable Trajectory Prediction with Multi-Dimensional Driving Personas Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing benchmarks either pro-vide no persona-conditioned annotations or support only a single urgencyspectrum (i.e., emergency, normal, relaxed), which cannot distinguishpersonas that share the same urgency level but require different drivingdynamics. To address this, we propose (i) the Persona-Conditioned Tra-jectory (PCT) dataset, which decomposes driving personas along twoaxes—Temporal Urgency and Ride Comfort—and combines three levelsof each to form a grid of nine personas, each paired with natural-languagedescriptions and trajectories, and (ii) PersonaDrive, a framework thatcan learn driving personas from language and can generate persona-specific trajectories. |
Chan Lee; Kimin Yun; Yuseok Bae; Seong Tae Kim; Jung Uk Kim; | code |
| 592 | HNDiff: Haze-Noise Diffusion for Image Dehazing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address thisissue, we propose Haze-Noise Diffusion (HNDiff), a novel diffusion frame-work that embeds the atmospheric scattering model as an inductive bias.By grounding diffusion in physical principles, HNDiff ensures that therestoration aligns more closely with underlying mechanisms of haze for-mation. |
Jin-Ting He; Fu-Jen Tsai; Yan-Tsung Peng; Min-Hung Chen; Chia-Wen Lin; Yen-Yu Lin; | code |
| 593 | SLER-IR: Spherical Layer-wise Expert Routing for All-in-One Image Restoration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose SLER-IR, a sphericallayer-wise expert routing framework that dynamically activates special-ized experts across network layers. |
Shurui Peng; Xin Lin; Shi Luo; Jincen Ou; Dizhe Zhang; Lu Qi; Truong Nguyen; Chao Ren; | code |
| 594 | Safe Responses Matter: Output-Aware Safety Guardrail Mitigate Over-Refusal in MLLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Usually, MLLMs already possess intrinsic safety mech-anisms that can transform harmful inputs into harmless outputs, butinput-side safety guardrails override this capability, degrading user ex-perience. Motivated by this insight, we propose a paradigm shift towardoutput-aware safety guardrails. |
Jiayi Li; Kun Zhan; | code |
| 595 | ComplexMimic: Human–Scene Interaction Imitation in Complex 3D Environments Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we focus on HSI mimicry in complexenvironments. |
Lu Pan; Hongwei Zhao; | code |
| 596 | VCBench: A Streaming Counting Benchmark for Spatial-Temporal State Maintenance in Long Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose SVCBench, a Streaming Video Counting Benchmark that repositions counting as a minimal, controlled probe for diagnosing models’ world-state maintenance capability. |
Pengyiang Liu; Zhongyue Shi; Hongye Hao; Qi Fu; Xueting BI; Siwei Zhang; Xiaoyang Hu; Zitian Wang; Linjiang Huang; Si Liu; | code |
| 597 | FlexAM: Flexible Appearance-Motion Decomposition for Versatile Video Generation Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose FlexAM, a unified framework built upon a novel 3Dcontrol signal. |
Mingzhi Sheng; Zekai Gu; Peng Li; Cheng Lin; Hao-Xiang Guo; Yingcong Chen; Yuan Liu; | code |
| 598 | PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this,we propose PolicyTrim, a reinforcement learning-based post-trainingframework that extends reliable action chunk length and reduces redun-dant physical steps. |
Xianghui Wang; Feng Chen; Wenbo Zhang; Hua Yan; Zixuan Wang; Changsheng Li; Yinjie Lei; | code |
| 599 | BioMedVR: Confusion-Aware Mixture-of-Prompt Experts for Biomedical Visual Reprogramming Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To mitigate class confu-sion, we introduce a Confusion Minimization Mechanism that leveragesLLM-generated confusion-aware attributes together with a Confusion-Suppression Loss to explicitly reduce false-positive alignment. |
Jiaxiang Liu; Tianxiang Hu; Juwei Guan; Yujie Wu; Yusong Wang; Yao Mu; Zuozhu Liu; Mingkun Xu; | code |
| 600 | On The Plasticity Collapse in Continual Machine Unlearning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While prior work largely studies single-shot un-learning, real-world systems must accommodate continual unlearning,where multiple unlearning requests occur sequentially over time. In thiswork, we identify a fundamental limitation of this setting: plasticity col-lapse, a progressive breakdown in a model’s ability to effectively forget.Through theoretical analysis of continual unlearning dynamics, we showthat continual unlearning operations accumulate geometric constraintsin parameter space, leading to saturated subspaces that restrict futureupdates. |
Yingdan Shi; Xiang Xu; Kaize Ding; Alfred Hero; Ren Wang; | code |
| 601 | SGP2: Coarse-to-Fine Controllable Multimodal Remote Sensing Image Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address these issues, we propose SGP2(Synergizing Geometric and Physical Priors), a novel coarse-to-fine frame-work for controllable multimodal RS image generation. Specifically, toresolve semantic conflicts, we introduce the Grassmann Miner, whichconstructs dynamic geodesic trajectories on the Grassmann Manifold,enabling the model to gradually shift its focus from global backgrounds tofine-grained local details during the denoising process. |
Pengyu Chen; Xi Yang; Nannan Wang; | code |
| 602 | 2D Features Are All You Need for 3D Shape Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present MeshFM, an efficient feedforward framework for extracting rich features from 3D inputs. |
Jinfan Zhou; Richard Liu; Itai Lang; Rana Hanocka; | code |
| 603 | Learning Egocentric Cues from Exocentric Video Using Privileged Egocentric Supervision Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, their view-point invariant training limits their ability to understand egocentricproperties (e.g., human-object interactions) from exocentric video ob-servations. This limitation is critical for applications such as Activi-ties of Daily Living (ADL) monitoring, where understanding egocen-tric properties is essential yet egocentric cameras are impractical to de-ploy, making it impossible to simply collect egocentric data at test time.To address this challenge, we propose Ego2ExoVLM, a VLM frame-work that learns to infer egocentric properties from exocentric videos.Our key insight is that time-synchronized ego-exo video pairs can beleveraged during training, where the egocentric viewpoint provides priv-ileged supervision (rich egocentric signal available only at training time). |
Dominick Reilly; Manish Govind; Le Xue; Srijan Das; | code |
| 604 | Free-Lunch Augmentation By Revisiting Diffusion-Based Data Generation for Cross-Domain Few-Shot Object Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: For the semantic gap, we find that the background se-mantics shows much smaller gaps between domains than foreground se-mantics, and we can bridge this gap by background inpainting. Basedon the above analysis, we propose a method (Selective Inpainting withTailored Noise, SITN) to dynamically take different strategies for down-stream data synthesis based on their different gaps from the generaldomain, including a Generation Module for adding tailored noise and aSelection Module to dynamically select the inpainting regions. |
Zijian Zhuang; Yixiong Zou; Yuhua Li; Ruixuan Li; | code |
| 605 | Rethinking Attention Reallocation for Multimodal Emotion Recognition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To understand this phenomenon,through comprehensive analysis, we show that with more modalities, thecomplex interactions between modalities make the model hard to paymore attention to tokens with more information, especially in deeperlayers, leading to improper attention allocation and the observed contra-dictory behavior. Based on these insights, we propose a training-freeattention rectification method that leverages structured shallow-layerattention as a prior to regularize entangled final-layer attention dur-ing inference, without introducing additional parameters or modifyingthe backbone model. |
Yazhe Lyu; Yixiong Zou; Jinghan Hu; Yuhua Li; Ruixuan Li; | code |
| 606 | Habitat-GS: A High-Fidelity Navigation Simulator with Dynamic Gaussian Splatting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Habitat-GS, a navigation-centric embodied AI simulator extended from Habitat-Sim that integrates 3D Gaussian Splatting scene rendering and drivable gaussian avatars while maintaining full compatibility with the Habitat ecosystem. |
Ziyuan Xia; Jingyi Xu; Chong Cui; Yuanhong Yu; Jiazhao Zhang; Qingsong Yan; Ni Tao; Junbo Chen; Xiaowei Zhou; Hujun Bao; Ruizhen Hu; Sida Peng; | code |
| 607 | PS-MOT: Cultivating Instance Awareness from Point Seeds for Multi-Object Tracking Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address these, we propose PS-Track, ahierarchical pipeline transitioning from points to instances across data,model, and loss levels. |
Kai Luo; Fei Teng; Mengfei Duan; Wanjun Jia; Xu Wang; Hao Shi; Kunyu Peng; Zhiyong Li; Kailun Yang; | code |
| 608 | Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Em-Garde, a novel framework that decouples semantic understanding from streaming perception. |
Yikai Zheng; Xin Ding; Yifan Yang; Shiqi Jiang; Hao Wu; Qianxi Zhang; Weijun Wang; Ting Cao; Yunxin Liu; | code |
| 609 | Odoriko: A Shape-Aware Multimodal Diffusion Framework for Human Motion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Odoriko, the first unified multimodal motion generationframework that reflects subject bio-morphological information directly insynthesized motion output. |
Dongseok Shim; Julian Tanke; Kengo Uchida; christian simon; Koichi Saito; Takashi Shibuya; Shusuke Takahashi; Yuki Mitsufuji; | code |
| 610 | 360Anything: Geometry-Free Lifting of Images and Videos to 360° Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose 360Anything, a geometry-free framework built upon pre-trained diffusion transformers. |
Ziyi Wu; Daniel Watson; Andrea Tagliasacchi; David Fleet; Marcus Brubaker; Saurabh Saxena; | code |
| 611 | DLGStream: Dynamic Language-embedded Guassian Splatting for Open-vocabulary Enabled Free-viewpoint Video Streaming Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In thispaper, we propose DLGStream, a novel language-embedded FVV repre-sentation that streams time-varying language features alongside Gaus-sian attributes to support 4D environment interaction, scene editing, andspatial intelligence. |
ZHIHUI KE; Yvyang Liu; Xiaobo Zhou; Tie Qiu; | code |
| 612 | 3DGS3: Joint Super Sampling and Frame Interpolation for Real-Time Large-Scale 3DGS Rendering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Instead of optimizing thesplatting pipeline itself, we propose 3DGS3 , a unified post-renderingframework that jointly performs super sampling and frame interpolationthrough differentiable processing of low-resolution outputs to achieveboth high-resolution and high-frame-rate rendering. |
Yibo Zhao; Fan Gao; Youcheng Cai; Ligang Liu; | code |
| 613 | PhysAlign: Learning Physical Priors for Dynamical Event-Driven Video Generation Via Representation Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To addressthis, we introduce PhysAlign, a framework for Dynamical Event-DrivenPhysically Consistent Video Generation. |
Mengxian Li; Zhan Wang; Fan Qi; Changsheng Xu; | code |
| 614 | Ego-Human Motion Prediction with 3D-Aware LLM Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Ego3DLM, built on two core principles:accurate motion forecasting requires explicit spatial and semantic un-derstanding of the 3D environment, and pose and language must be pre-dicted holistically in a single pass, since motion is inherently tied to thesemantic interpretation of actions being performed. |
Yujin Bae; Jaewoo Jeong; HYEONSEONG KIM; KUK-JIN YOON; | code |
| 615 | HSDF-Lane: Height-Aligned Signed Distance Field with Semantic Lane Prior for 3D Lane Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent methods instead predict explicit heightmaps to capture non-planar surfaces, but still rely on sparse anchor-based regres-sion and exploit the recovered geometry merely for spatial transformation ratherthan semantic understanding. To overcome these limitations, we propose HSDF-Lane, which implicitly models the road surface as a Height-aligned Signed Dis-tance Field (HSDF) over a densely sampled 3D feature volume. |
Jiyong Boo; ByeongIn Joung; Hyemin Yang; KUK-JIN YOON; | code |
| 616 | PriSplat: Propagating Reliable Multi-view Information for Distractor-Free 3DGS Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To resolve this, it is essential to re-establish densemulti-view constraints by recovering the missing background informa-tion. In light of this, we propose PriSplat, a novel framework designed topropagate reliable multi-view information to restore these missing regionswith high geometric integrity. |
Yunseo Yang; Youngho Yoon; KUK-JIN YOON; | code |
| 617 | Distill Once, Adapt Life-Long: Exploring Dataset Distillation for Continual Test-Time Adaptation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce DO-ALL (Distill Once, AdaptLife-Long), a plug-and-play framework that revisits source informationin a compact and privacy-conscious form via Dataset Distillation (DD). |
Hyun-Kurl Jang; Jihun Kim; Hyeokjun Kweon; KUK-JIN YOON; | code |
| 618 | GoStop: Reinforcement Learning for Adaptive Temporal Aggregation in Event-Based Feature Tracking Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we model event accumulation as a sequential decision-making problem and introduce reinforcement learning (RL) framework to adaptively control the accumulation process for online event-based feature tracking. |
Youngho Kim; Hoonhee Cho; Jae-young Kang; KUK-JIN YOON; | code |
| 619 | Mask-guided Semantic Alignment: Robust Learning with Noisy Labels Via Temporal Attention Stability Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a temporal attention-basedframework for learning with noisy labels. |
Yubo Nian; Can Gao; | code |
| 620 | SCoT: Similarity-guided Conflict-aware Task Consolidation for Continual VQA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existingcontinual VQA methods primarily rely on replay or parameter regulariza-tion but largely overlook how task-specific updates accumulate and inter-act in parameter space, particularly whether successive updates are syn-ergistic or conflicting across layers. To address this, we introduce SCoT(Similarity-guided Conflict-aware Task Consolidation), a continual learn-ing framework that represents each task as a parameter update relative to apretrained anchor model and integrates tasks through layer-wise parameter-space reasoning. |
Anand Patel; Moloud Abdar; Biplab Banerjee; | code |
| 621 | ORFC: Orthogonal Reparameterization for Low-Bitrate ViT Feature Coding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inparticular, our analysis reveals that under fixed-group product quanti-zation, task sensitivity is highly uneven across channel subspaces. |
Letian Zhang; Wenhan Yang; Lingyu Duan; | code |
| 622 | GroundSet: A Cadastral-Grounded Dataset for Spatial Understanding with Vector Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, Multimodal Large Language Models ex-hibit critical deficiencies in fine-grained spatial understanding withinRemote Sensing, primarily due to a reliance on limited or repurposedlegacy datasets. To bridge this gap, we introduce a large-scale datasetgrounded in verifiable cadastral vector data, comprising 3.8 million an-notated objects across 510k high-resolution images with 135 granularsemantic categories. |
Roger Ferrod; Maël Lecene; Krishna Sapkota; George Leifman; Vered Silverman; Genady Beryozkin; Sylvain Lobry; | code |
| 623 | SegVGGT: Joint 3D Reconstruction and Instance Segmentation from Multi-View Images Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present SegVGGT, a unified end-to-endframework that simultaneously performs feed-forward 3D reconstructionand instance segmentation directly from multi-view RGB images. |
Jinyuan Qu; Hongyang Li; Lei Zhang; | code |
| 624 | TIGER: Taming Identity, Geometry, and Generative Priors for High-Quality Face Video Restoration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To tackle these issues, we propose TIGER, a structured tri-prior fusion framework that Tames Identity, Geometry, and gEnerative pRiors for high-quality FVR.We also construct a large-scale FVR dataset to facilitate robust training and standardized evaluation. |
Yang Zhou; Wenxue Li; Peng Zhang; Yifei Chen; Fei Wang; Daiguo Zhou; | code |
| 625 | Efficient Quantization-Aware Adaptation for Visual Foundation Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Efficient Quantization-aware Adaptation (EQuA) that achieves high efficiency in both adaptation and deployment for visual foundation models. |
Yinglong Li; Xiaoyu Liu; Yutong Liu; Yueyi Zhang; Zhiwei Xiong; | code |
| 626 | VQT: Vector Quantization Tuning for Efficient Fine-tuning and Compression of Pre-trained Vision Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In thiswork, we propose Vector Quantization Tuning (VQT), a novel frameworkfor efficient fine-tuning and compression of pre-trained ViTs. |
Yinglong Li; jiyuan Xia; Jingcheng Xie; Zhiwei Xiong; | code |
| 627 | Hi-Nav: Hierarchical Framework for Continuous Vision-Language Navigation Via Map Guidance and Waypoint Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Most meth-ods are trained and evaluated in simplified or physics-limited simulators,which ignore the feedback loop between decision making and physical ex-ecution, resulting in unstable navigation performance. To bridge this gap,we propose Hi-Nav, a top-down hierarchical navigation framework thatdecomposes VLN into three controllable levels. |
Zhiyu Zhou; Bin Guan; Wenbin Yang; Zhi Gao; Hao Fang; | code |
| 628 | LogiCo: A Unified Framework for Logical and Structural Anomaly Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose LogiCo , a uni_x001C_ed frame-work for Logi cal and structural anomaly detection via Co mponent-levelfeature reconstruction. |
Ximiao Zhang; Min Xu; Xiuzhuang Zhou; | code |
| 629 | Causal Intervention in Concept Bottleneck Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we show that trained CBM parameters encode usable concept-correlation structure. |
Zhiyu Zhu; Jiayu Zhang; Zhibo Jin; Xinyi Wang; Fang Chen; Przemyslaw Biecek; Jianlong Zhou; | code |
| 630 | Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We tackle this gap by introducing a novel evaluation metric derived from a learned latent space of real-world human actions.For rigorous validation, we develop a new multi-faceted benchmark specifically designed to probe temporally challenging aspects of human action fidelity. |
Xavier Thomas; Youngsun Lim; Ananya Srinivasan; Audrey Zheng; Deepti Ghadiyaram; | code |
| 631 | Moving Beyond More Views: Redundancy-Aware Ego–Exo Fusion for Proficiency Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our anal-ysis identifies two key causes: (1) Multiview redundancy — Fromthe data perspective, certain views provide limited or noisy informa-tion, diluting discriminative cues; (2) Overfitting — From the featureperspective, conventional fusion increases representational complexity,causing the model to memorise view-specific patterns rather than learngeneralisable representations. To address these issues, we propose twocomplementary modules. |
Xu Dong; Wanqing Li; Anthony Adeyemi-Ejeye; Andrew Gilbert; | code |
| 632 | TruthLens: Object Hallucination Detection Via Self-Evaluating Truthfulness Scores in LVLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose TruthLens, a self-evaluation frame-work that teaches the LM head to expose a per-object truthfulness sig-nal without any auxiliary model or additional inference cost. |
Yanqi Wu; Runhe Lai; Xinhua Lu; Qichao Chen; Zhiping Zhou; Jia-Xin Zhuang; Weijiang Yu; Ruixuan Wang; | code |
| 633 | Prototype-Conditioned Imagination for Compositional Zero-Shot Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, they often suf-fer from ambiguous vision–language alignment and lack mechanisms toexplicitly capture cross-primitive semantic interactions, ultimately hin-dering generalization to unseen compositions. To overcome these chal-lenges, we introduce Prototype-Conditioned Imagination (PCI), a two-stage framework. |
Yifan Zhu; Haofeng Zhang; | code |
| 634 | Beyond Sequential Distance: Inter-Modal Distance Invariant Position Encoding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Despite the remarkable capabilities of Multimodal Large Lan-guage Models (MLLMs), they still suffer from visual fading in long-context scenarios. Specifically, the attention to … |
Lin Chen; Bolin Ni; Qi Yang; Zili Wang; Kun Ding; Ying Wang; Houwen Peng; SHIMING XIANG; | code |
| 635 | SignSparK: Efficient Multilingual Sign Language Production Via Sparse Keyframe Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Sign Language Production (SLP) faces a fundamental tradeoff: direct text-to-pose models suffer from regression-to-the-mean effects, while dictionary-retrieval methods produce disjointed transitions. To resolve this, we propose a novel training paradigm that leverages sparse keyframes to capture the underlying kinematic distribution of human signing. |
Jian He Low; Alexandre Symeonidis-Herzig; Maksym Ivashechkin; Ozge Mercanoglu Sincan; Richard Bowden; | code |
| 636 | StyleFusion360: View-Consistent Head Stylization Via Adaptive Style Modulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing 3D-aware methods often require computationally intensive optimization orper-style fine-tuning, limiting flexibility and user control. To overcomethese challenges, we introduce StyleFusion360, a diffusion-based frame-work for multi-view consistent, identity-preserving 3D head stylizationfrom a single style reference image, without per-style training. |
Furkan Guzelant; Arda Goktogan; Tarık Kaya; Aysegul Dundar; | code |
| 637 | ID-LoRA: Identity-Driven Audio-Video Personalization with In-Context LoRA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose ID-LoRA (Identity-Driven In-Context LoRA), which jointly generates a subject’s appearance and voice in a single model, letting a text prompt, a reference image, and a short audio clip govern both modalities together. |
Aviad Dahan; Moran Yanuka; Noa Kraicer; Lior Wolf; RAJA GIRYES; | code |
| 638 | LaMP: Learning Vision-Language-Action Policies with 3D Scene Flow As Latent Motion Prior Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce LaMP, a dual-expert Vision-Language-Actionframework that embeds dense 3D scene flow as a latent motion prior forrobotic manipulation. |
Xinkai Wang; Chenyi Wang; Yifu Xu; Mingzhe Ye; Fu-Cheng Zhang; Jialin Tian; Xinyu Zhan; Lifeng Zhu; Cewu Lu; Lixin Yang; | code |
| 639 | ExpertEdit: Learning Skill-Aware Motion Editing from Expert Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce ExpertEdit, a framework for skill-driven motion editing trained exclusively on unpaired expert video demonstrations. |
Arjun Somayazulu; Kristen Grauman; | code |
| 640 | AccelAes: Accelerating Diffusion Transformers for Training-Free Aesthetic-Enhanced Image Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Weobserve that denoising is spatially non-uniform with respect to aestheticdescriptors in the prompt. Regions associated with aesthetic tokens re-ceive concentrated cross-attention and show larger temporal variation,while low-affinity regions evolve smoothly with redundant computation.Based on this insight, we propose AccelAes, a training-free frameworkthat accelerates DiTs through aesthetics-aware spatio-temporal reduc-tion while improving perceptual aesthetics. |
Xuanhua Yin; CHUANZHI XU; Haoxian Zhou; Boyu Wei; Weidong Cai; | code |
| 641 | Decoupling Moment from Event for Video Temporal Grounding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This matching scheme incorrectly suppresses such moment predictions as negatives, preventing models from learning rich temporal representations. To resolve this, we propose Moment-Event DETR (ME-DETR), a novel framework that embraces the natural part-whole structure of events through dynamic query specialization. |
Yuda Zou; Boxiang Zhou; Xin Zhou; YIBO CHEN; Dejia Song; Xu Tang; Yao Hu; Yongchao Xu; | code |
| 642 | FST-SAM3: Taming SAM~3 with Frequency-Spatio-Temporal Refinement for Video Polyp Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although Segment AnythingModel 3 (SAM 3) is able to provide strong representations for segmen-tation tasks, directly applying it to endoscopic videos usually cannotachieve satisfactory performance owing to domain shifts, boundary am-biguities, and unstable temporal propagation. To address these chal-lenges, we propose FST-SAM3, a novel parameter-efficient adaptationframework built on SAM 3 for VPS. |
Guanhao Wu; Guilian Chen; Huisi Wu; Jing Qin; | code |
| 643 | PASDiff: Physics-Aware Semantic Guidance for Joint Real-world Low-Light Face Enhancement and Restoration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose PASDiff, a Physics-Aware Semantic Diffusion in a training-free manner. |
Yilin Ni; Wenjie Li; Zhengxue Wang; Juncheng Li; Guangwei Gao; Jian Yang; | code |
| 644 | REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Response-Weighted and Validation-Anchored Policy Optimization (REVA-PO), aRL framework that stabilizes long-term training via Response-WeightedRegularization (RER) and Validation-Anchored Policy Reset (VAPR). |
Li Guo; Anas Tahir; Z. Jane Wang; | code |
| 645 | UniH3: Unifying Hierarchical Homogeneity and Heterogeneity for All-in-One Medical Image Restoration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose UniH3 , a novel frame-work that Unifies Hierarchical Homogeneity and Heterogeneity for all-in-one medical image restoration. |
Zhiwen Yang; Jiayin Li; Chengyu Liu; Hui Zhang; Bingzheng Wei; Yan Xu; | code |
| 646 | MobileOcc: A Human-Aware Semantic Occupancy Dataset for Mobile Robots Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Dense 3D semantic occupancy perception is critical for mo-bile robots operating in pedestrian-rich environments, yet it remains un-derexplored compared to its application in autonomous driving. To ad-dress this gap, we present MobileOcc, a semantic occupancy dataset formobile robots operating in crowded human environments. |
Junseo Kim; Guido Dumont; Xinyu Gao; Gang Chen; Holger Caesar; Javier Alonso-Mora; | code |
| 647 | ScAle: Attention Head Scaling As A Minimal Adapter for Spatial Reasoning in Vision–Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our preliminary analysisreveals that rescaling activations in selected transformer layers—withoutmodifying pretrained weights—can significantly influence downstreamperformance. Motivated by this observation, we propose ScAle, an ultra-lightweight adaptation method that learns a small set of scalar coeffi-cients to modulate last-token attention and MLP activations in a fullyfrozen backbone. |
Rahul Chowdhury; Timothy Rupprecht; Xuan Shen; Pu Zhao; Yanzhi Wang; | code |
| 648 | MedSPOT: A Workflow-Aware Sequential Grounding Benchmark for Clinical GUI Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce MedSPOT, a workflow-aware sequential grounding benchmark for clinical GUI environments. |
Rozain Shakeel; Abdul Ali; Muneeb Ganie; Tausifa Jan Saleem; Tajamul Ashraf; | code |
| 649 | UniTemp: Unlocking Video Generation in Any Temporal Order Via Autoregressive Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To addressthis, we introduce blockwise anchor latents, a set of auxiliary latentsthat aim to restore the missing past context at block boundaries dur-ing backward generation. Built on this design, we propose UniTemp, abidirectional distillation framework that trains an autoregressive studentmodel for any-direction video generation. |
Lin Zhang; Sicheng Mo; Zefan Cai; Jinhong Lin; Zihao Lin; Jiuxiang Gu; Krishna Kumar Singh; Yuheng Li; Yin Li; | code |
| 650 | TouchAnything: Diffusion-Guided 3D Reconstruction from Sparse Robot Touches Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Wepresent TouchAnything, a framework that leverages a pretrained large-scale 2D vision diffusion model as a semantic and geometric prior for3D reconstruction from sparse tactile measurements. |
Langzhe Gu; Hung-Jui Huang; Mohamad Qadri; Michael Kaess; Wenzhen Yuan; | code |
| 651 | From Noise to Events: Conditional Diffusion for Event Data Augmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, improving the model’s robustness against overfitting remains a critical challenge. To address this, we propose augmenting event data from the noise rather than applying fixed data transformations. |
Ruofei Wang; Ziyuan Luo; Peiqi Duan; Xiufeng HUANG; Boxin Shi; Renjie Wan; | code |
| 652 | P²Fusion: Prompt-based Progressive Infrared-Visible Image Fusion Via Dual-Prior Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Exist-ing prior-guided methods often rely on static constraints that induceoptimization conflicts or utilize extrinsic semantic priors from large-scale foundation models (e.g., CLIP/DINO), which frequently fail toexploit the intrinsic modality characteristics essential for high-fidelityfusion. To address these issues, we propose P²Fusion, a prior-guideddistillation-based framework that reformulates IVIF via dual intrinsicprompts. |
Yi Shi; Huichao Xie; Yuqing Wang; Mingyu Wang; Kaihui Yang; Yu Liu; Lu Ruitao; Lizhe Li; Junwei Han; Dingwen Zhang; | code |
| 653 | Multi-History-Step SDE Inversion for Image Editing with Superior Regional Awareness Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing approaches remain inefficient, exhibit limitedplasticity, and struggle to accurately preserve unedited regions. To ad-dress these issues, we propose MIEdit, a training-free editing frame-work based on SDE inversion. |
Haiyan Wei; Yunlong Wang; Huaibo Huang; Zhenan Sun; Kunbo Zhang; | code |
| 654 | PolyLayout: Multi-room Manhattan Layout Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce two new multi-view multi-room layoutbenchmarks by providing layout annotations to existing datasets, andexperiments show that PolyLayout outperforms prior approaches, bothin terms of accuracy and robustness.Project page: https://ghanning.github.io/PolyLayout |
Gustav Hanning; Shaohui Liu; Rémi Pautrat; Marc Pollefeys; Kalle Åström; Viktor Larsson; | code |
| 655 | DETRPose: Real-Time End-to-End Multi-Person Pose Estimation Via Modified Transformer Decoder and Novel Denoising Keypoints Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: For training, a novel denoising keypoint technique is proposed to accelerate convergence. |
Sebastian Janampa; Marios Pattichis; | code |
| 656 | Spectral Gating Via Damped Oscillations for Adaptive Implicit Neural Representations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose to model each neuron’s activationas the steady-state response of a sinusoidally-forced damped harmonicoscillator, whose amplitude naturally governs the network’s spectral se-lectivity during training. |
Alex Costanzino; Pierluigi Zama Ramirez; Giuseppe Lisanti; Luigi Di Stefano; | code |
| 657 | Dynamic Cluster Data Sampling for Efficient and Long-Tail-Aware Vision-Language Pre-training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, weintroduce a dynamic cluster-based sampling approach (DynamiCS) thatdownsamples large clusters of data and upsamples small ones. |
Mingliang Liang; Zhuoran Liu; Arjen P. de Vries; Martha Larson; | code |
| 658 | RegHead: Non-Humanoid Head Blendshapes Via Feed-Forward Registration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present RegHead, a framework for constructing semanticblendshape sets for animatable non-humanoid head avatars. |
Jiahao Luo; Hao Zhang; Jianqi Chen; Yijie He; Jiaxu Zou; Michael Vasilkovsky; Sergei Korolev; Sergey Tulyakov; Chaoyang Wang; Peter Wonka; James Davis; Jian Wang; | code |
| 659 | ASSCG: Just-Right Gating Over Chattering for Fast–Slow LLM Planning in Autonomous Driving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We formulate slow-system invocation as a resource-aware sequential decision problem and propose the Adaptive Slow-System Control Gate (ASSCG), which makes framelevel Query/Cache/Drop decisions to refresh, reuse, or suppress slow guidance. |
Sining Ang; Yuan Chen; Liu Haiyan; Xuanyao Mao; jason bao; Xuliang Xuliang; Bingchuan Sun; Yan Wang; | code |
| 660 | Gender Bias in Vision-Language In-Context Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In-context learning (ICL) enables large vision-language models (LVLMs) to perform tasks by following patterns in context examples, yet its potential to amplify societal biases remains underexplored. |
Tong Xiang; Yuta Nakashima; Noa Garcia; | code |
| 661 | SparseCtrl-HOI: Sparse Temporal Control for Human-Object Interaction Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, such dense guidance incurs high annotation costs and affects motion synthesis diversity. To overcome these limitations, we introduce SparseCtrl-HOI, a novel sparse temporal control framework for HOI video generation. |
Shenbo Xie; Mingrui Cai; Xu Yang; Yifei Liu; Changxing Ding; | code |
| 662 | Different Changes Require Different Reasoning: Change-Type-Specialized Experts for Robust Change Captioning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We build our MEDIC as a memory network to dynamicallyretrieve type-relevant visual patterns conditioned on the input. |
Jiyoung Park; InJae Oh; Jung Uk Kim; | code |
| 663 | AgentVLN: Towards Agentic Vision-and-Language Navigation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, wepropose AgentVLN, a novel and efficient embodied navigation frameworkthat can be deployed on edge computing platforms. |
Zihao Xin; Wentong Li; Yixuan Jiang; Ziyuan Huang; Bin Wang; Piji Li; Jianke Zhu; Jie Qin; Sheng-Jun Huang; | code |
| 664 | PaD-GS: Leveraging Distortion Map for Panoramic Gaussian Splatting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, the performance of the existing methods in literature is generally limited due to the severe distortions involved in panoramic images. To address this problem, we construct a distortion map under the panoramic imaging model, where the value of each pixel reflects its corresponding distortion degree, and accordingly, we propose a novel Panoramic Gaussian Splatting method by utilizing this Distortion map, called PaD-GS. |
Yihang Xu; Qiulei Dong; | code |
| 665 | Spotlight: Identifying and Localizing Video Generation Errors Using VLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Vision Language Models(VLMs) are actively being adopted as automatic evaluators for videogeneration, driven by the promise of their perception and reasoning abili-ties. |
Aditya Aravind Chinchure; Sahithya Ravi; Pushkar Shukla; Vered Shwartz; Leonid Sigal; | code |
| 666 | FlowFace: Rectifying Identity Conditioning with Riemannian Geometry for Face Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present FlowFace, a geometry-aware identity-conditioning framework that rectifies the Stable Diffusionconditioning geometry during training without modifying the inferencesampler. |
Xinran Deng; Yiling Wu; Ye Tian; Libo Zhang; | code |
| 667 | MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Moreover, by constraining the planner to this compressed context, it is difficult to leverage the rich representations offered by modern vision foundation models. To address these issues, we propose MOJITO, a unified sensorto-action framework for end-to-end autonomous driving built on modal joint learning. |
Zhijing Cheng; Xuancheng Zhang; Donglin Di; Lei Fan; Baorui Ma; Hao Li; Xun Yang; | code |
| 668 | From Script to Shot: A Benchmark for Grounding Screenplays in Movies Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To disentangle text matchingfrom visual grounding, we evaluate fourteen approaches—from sparsetext retrieval to recent vision-language models such as FG-CLIP 2 andQwen3-VL-Embedding, on an identical pipeline.We release the full dataset, parser and parser outputs,toolkit, and baselines at https://github.com/jungucho92/script2shotto support research on long-form multimodal video understanding.2 J. Cho et al. 사용 폰트: 노토산스한국:https://fonts.google.com/noto/specimen/Noto+Sans+KRScreenplay -Video Alignment Prior works(dialogue -based)Screenplay SCENE 69 dialogue D D D D D Dthe script of a movie,H INT HOTEL ROOM-HOURS LATERincluding actinginstructions and D HAZEL Good morning.subtitle#69 scene directions. |
jungu cho; Young-Jae Park; Seong Jong Ha; Siyeol Kim; Seungho Park; Jisu Shin; Junmyeong Lee; Chaemin Hwang; Hyunjun Jung; Sangeyl Lee; Hae-Gon Jeon; | code |
| 669 | Neural Gate: Mitigating Privacy Risks in LVLMs Via Neuron-Level Gradient Gating Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Large Vision-Language Models (LVLMs) have shown remark-able potential across a wide array of vision-language tasks, leading totheir adoption in critical domains such as finance and healthcare. |
Xiangkui Cao; Jie Zhang; Meina Kan; Shiguang Shan; Xilin CHEN; | code |
| 670 | Unified Prediction and Planning Via Conflict-Aware Disjoint Parameter Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To resolvethis, we propose a novel model-merging-based framework, Disjoint Pa-rameter Training (DPT). |
Taewon Seo; Seonae Jeon; Giwon Lee; KUK-JIN YOON; Daehee Park; | code |
| 671 | White Aggregation and Restoration for Few-shot 3D Point Cloud Semantic Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, we found that this convention resultsin performance instability due to its sensitivity to FPS-induced varia-tions, while the prototype generation process remains underexplored inthe field. This motivates us to investigate deterministic prototype gen-eration method based on attention mechanism. |
Jiyun Im; SuBeen Lee; Miso Lee; Jae-Pil Heo; | code |
| 672 | Blind to Position, Biased in Language: Probing Mid-Layer Representational Bias in Vision-Language Encoders for Zero-Shot Language-Grounded Spatial Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Second, multilingual text embeddings form language-dependent geometric shifts within the shared space. Motivated by these findings, we identify an underexplored pathway within VLE mid-layers to construct a spatial map, applicable for improving zero-shot RIS by 1–7 mIoU on nine RefCOCO benchmarks. |
Na Min An; Inha Kang; Minhyun Lee; Hyunjung Shim; | code |
| 673 | The 3D Mirage: Probing and Taming 3D Hallucinations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introduces a novel end-to-end framework toprobe, score, and tame this under-quantified safety risk in monoculardepth under context variation.To probe, we present 3D-Mirage, thefirst benchmark to combine context variation and precise annotation forreal-world illusions with real object exclusions, multi-surface support;purpose-built to stress-test monocular depth on real-world illusions. |
Hoang Nguyen; Xiaohao Xu; Xiaonan Huang; | code |
| 674 | VLOD-TTA: Test-Time Adaptation of Vision-Language Object Detectors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Test-time adapta-tion (TTA) offers a practical way to adapt models online using onlyunlabeled target data. |
Atif Belal; Heitor Medeiros; Marco Pedersoli; Eric Granger; | code |
| 675 | TurboMPLE: Joint Infrared Turbulence Mitigation and Physical Fields Estimation Via Mutual Progressive Layered Extraction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In thisstudy, we construct a large-scale infrared turbulence imaging dataset andproposed the Turbulence-oriented Mutual Progressive Layered Extrac-tion (TurboMPLE), a joint network for infrared turbulence mitigationand physical fields estimation. |
Yitong An; Yubo Jiang; Xiangzhi Bai; | code |
| 676 | RainODE: Continuous-Time Precipitation Forecasting with Latent Neural ODEs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, increasing temporal resolution isconstrained by observational limitations and the computational cost ofdense discrete modeling. To overcome this limitation, we reformulate pre-cipitation forecasting as a continuous-time dynamical system and pro-pose RainODE, a framework that models precipitation evolution in la-tent space using a Neural ODE. |
Yeeun Seong; Doyi Kim; Minseok Seo; Changick Kim; | code |
| 677 | Evaluating and Understanding Model Editing for Medical Vision Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing multimodal model editing benchmarks focus on general-purpose tasks and do not reflect realistic clinical domain requirements and variability. To address this, we introduce M3Bench, a clinically grounded benchmark for multimodal model editing that evaluates whether an edit remains reliable, precise and generalizable under the challenges of image and text variation, modality and protocol shifts, clinical knowledge composition, and temporal progression. |
Guli Zhu; Chenwei Wu; Liyue Shen; | code |
| 678 | PixelU: A U-Shaped Transformer for Efficient End-to-End Pixel Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing works heavily rely on complex pixel decoders to alleviatethis issue. In this paper, we challenge this trend by revealing that thesedecoders primarily compensate for the optimization difficulties inherentto velocity prediction (v-prediction). |
Zipeng Guo; Lichen Ma; Yu He; Xiaolong Fu; Jingling Fu; Junshi Huang; Yan Li; | code |
| 679 | ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This inevitably suppresses high-frequency boundary cues and neglects the explicit inter-frame depen-dencies required for precise boundary delineation. To address these lim-itations, we present ScanFocus, a novel coarse-to-fine framework thatdecouples the STVG task into a global spatio-temporal scan and a localboundary focus. |
Chen Kai; Ming Dai; Wenxuan Cheng; Wankou Yang; | code |
| 680 | ASTAD: Asymmetric Style Transfer for Synthetic-to-Real Adaptation in Autonomous Driving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Consequently,existing methods relying on symmetric semantic guidance suffer fromeither prohibitive annotation costs or severe semantic misalignment. Toaddress this dilemma, we formally propose a novel task: AsymmetricStyle Transfer for Autonomous Driving (ASTAD), which requires se-mantically consistent transfer using only labeled synthetic content andunlabeled real-world references. |
Dingyi Yao; Xinqi Zhang; Lihui Peng; Jianming HU; Danya Yao; Yi ZHANG; | code |
| 681 | Event-Driven Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Event-Driven Video Generation (EVD), a small DiT-compatible intervention that gives the sampler an explicit event signal. |
Chika Maduabuchi; Jindong Wang; | code |
| 682 | EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose EraseLoRA, a dataset-free framework that replaces attention surgery with background-awarereasoning and test-time adaptation. |
Sanghyun Jo; Donghwan Lee; Eunji Jung; Seong Je Oh; Kyungsu Kim; | code |
| 683 | ISAC: Training-Free Instance-to-Semantic Attention Control for Multi-Instance Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: By contrast, self-attention exposes class-agnostic instance layouts during early denoising. To exploit this asymmetry, we propose ISAC (Instance-to-Semantic Attention Control), a training-free, model-agnostic objective that first stabilizes self-attention layouts and then binds cross-attention semantics within them, without fine-tuning or external vision models. |
Sanghyun Jo; Wooyeol Lee; Ziseok Lee; Jonghyun Choi; Jaesik Park; Kyungsu Kim; | code |
| 684 | CCFM: Collision-Constrained Flow Matching for Safety-Critical Scenario Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Evaluation of autonomous vehicle (AV) planners in safety-critical closed-loop simulation is essential for real-world deployment. How-ever, generating controllable safety-critical … |
Ke Li; Kaidi Liang; Yuxin Ding; Debojyoti Biswas; Xianbiao Hu; Ruwen Qin; | code |
| 685 | Direct Diffusion Score Preference Optimization Via Stepwise Contrastive Policy-Pair Supervision Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work,we introduce Direct Diffusion Score Preference Optimization (DDSPO),which defines stepwise preference supervision directly over backward de-noising transitions through a contrastive policy pair, rather than relyingon forward-process approximations from terminal samples. |
Dohyun Kim; Seungwoo Lyu; Seung Kim; Paul Hongsuck Seo; | code |
| 686 | LINA: Learning INterventions Adaptively for Physical Alignment and Counterfactual Generation in Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our analysis yields three key in-sights: (1) DMs struggle with physical elements not explicitly determinedin the prompt; (2) the prompt embedding contains disentangled repre-sentations for texture and physics; (3) visual causal structure is dispro-portionately established during the initial, computationally constraineddenoising steps. Based on these findings, we introduce Lina (LearningINterventions Adaptively), a novel framework that learns to predictprompt-specific interventions. |
Shu Yu; Chaochao Lu; | code |
| 687 | HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, wepropose HyLaR (Hybrid Latent Reasoning), a framework that seamlesslyinterleaves discrete text generation with continuous visual latent rep-resentations. |
Tao Cheng; Shi-Zhe Chen; Hao Zhang; Yixin Qin; Jinwen Luo; Zheng Wei; | code |
| 688 | Why Feature Magnitude Deceives OOD Detectors: An Angular Separation Perspective Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we revisit the fundamentals of OOD detection and uncover a key flaw in common distance-based detectors: sensitivity to feature magnitude. |
Hanlin Li; Jing Ma; Zehang Wei; Jiamin Yan; Xiang Xiang; | code |
| 689 | ControlHair: Synergizing Physics Simulator and Video Diffusion for Controllable Dynamic Hair Rendering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present ControlHair, a hybrid framework that integrates a physics simulator with conditional video diffusion to enable precise and controllable dynamic hair rendering. |
Weikai Lin; Haoxiang Li; Yuhao Zhu; | code |
| 690 | TRaM-VSR: Importance-Aware Token Routing and Merging for One-Step Diffusion Video Super-Resolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing efficiency-oriented methods risk irreversible detail loss and temporal flickering, a vulnerability especially pronounced in one-step diffusion models. To address this, we propose TRaM-VSR, a Token Routing and Merging framework for adaptive token allocation, leveraging both context-aware video priors and network-level priors. |
Sicheng Gao; Yixuan Liu; Tong Shen; Zhuyun Zhou; Zongwei Wu; Radu Timofte; | code |
| 691 | PUF: Plug-and-Play Uncertainty-Aware Fusion for Online 3D Scene Graph Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose PUF: a Plug-and-play, Uncertainty-aware, andtraining-free Fusion framework. |
Yi Yang; Myrna Castillo Silva; Bodo Rosenhahn; Michael Yang; | code |
| 692 | Self-transcendence: Is External Feature Guidance Indispensable for Accelerating Diffusion Transformer Training? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We argue that DiTs actually have the powerto guide the training of themselves, and propose Self-Transcendence,an effective method that achieves fast convergence using internal featuresupervision only. |
Lingchen Sun; Rongyuan Wu; zhengqiang ZHANG; Ruibin LI; Yujing Sun; Shuaizheng LIU; Lei Zhang; | code |
| 693 | YeTI: You Only Need Two Noisy Images for Real-World SRGB Noise Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose YeTI, a real-world sRGB noise generation framework that learns from only two noisy observations of the same scene. |
Jaekyun Ko; Byung Wan Lim; Dongjin Kim; Soomin Lee; Tae Hyun Kim; | code |
| 694 | Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Video Large Language Models (Video LLMs) incur high in-ference latency due to a large number of visual tokens provided to LLMs.To address this, training-free visual token pruning has emerged as asolution to reduce computational costs; however, existing methods areprimarily validated on Multiple-Choice Question Answering (MCQA)benchmarks, where coarse-grained cues often suffice. In this work, wereveal that these methods suffer a sharp performance collapse on fine-grained understanding tasks requiring precise visual grounding, such ashallucination evaluation. |
Kibum Kim; Jiwan Kim; Kyle Min; Yueqi Wang; Jinyoung Moon; Julian McAuley; Chanyoung Park; | code |
| 695 | CLDefocus: Physically Grounded Compound-Lens Defocus Blur Synthesis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Al-though recent deep learning deblurring methods have achieved strongperformance, their effectiveness depends on training data and often de-grades across cameras and lenses due to limited optical diversity and re-alism in existing datasets. In this paper, we propose a pipeline for synthe-sizing realistic defocus deblurring datasets for diverse compound lenses.It integrates efficient wave-optics PSF computation via Debye CZT prop-agation, depth-aware defocus rendering with occlusion handling, and blursynthesis in the radiometrically linear space with camera ISP simulation.This unified pipeline enables the scalable generation of photorealisticdefocus datasets with diverse lens characteristics. |
Yunkyu Lee; Woohyeok Kim; Sunghyun Cho; | code |
| 696 | FoundationGeo: Learning Spatial Pixel-Wise Fields for Monocular Metric Geometry Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present FoundationGeo, a two-stage framework that ex-plicitly bridges relative and metric prediction via spatial calibration andprincipled data design. |
Muxin Liu; Xiaoyang Lyu; Tianhe Ren; Peng Dai; Xiaoshan Wu; Zhiyue Zhang; Jiaqi Zhang; Jiehong Lin; Shaoshuai Shi; Qi Xiaojuan; | code |
| 697 | Revisiting Autoregressive Models for Generative Image Classification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we revisit visual AR-based generative clas-sifiers and identify an important limitation of prior approaches: theirreliance on a fixed token order, which imposes a restrictive inductivebias for image understanding. |
Ilia Sudakov; Artem Babenko; Dmitry Baranchuk; | code |
| 698 | CaRe: Critical Parameter Rectification for Efficient Visual Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Critical Parameter Rectification(CaRe), which generalizes fixed corrections to learned selective rectifi-cation. |
Mingjia Li; Hongkun Xiong; Yuheng Shi; Hengxing Liu; Xiaojie Guo; | code |
| 699 | Causal Yet Future-Aware: Dual-Path Temporal Modeling for Online Action Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing online TAS methodsmainly focus on modeling historical dependencies, but temporal ambigu-ity still arises due to the incomplete temporal dependencies, leading tofragmented predictions. To tackle this challenge, we propose Dual-PathTemporal Modeling (DPTM) for online TAS, a two-stage framework thatexplicitly models and transfers the future dependencies that are inacces-sible during online inference. |
Zhichao Zheng; Peirong Ma; Ying Zhou; Li Kong; Junsheng Zhou; | code |
| 700 | CLIMP: Contrastive Language-Image Mamba Pretraining Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Contrastive Language-Image Pre-training (CLIP) relies onVision Transformers whose attention mechanism is susceptible to spu-rious correlations and scales quadratically with resolution. To addressthese limitations, we present CLIMP, the first fully Mamba-based con-trastive vision-language model that replaces both the vision and text en-coders with state-space architectures. |
Nimrod Shabtay; Itamar Zimerman; Eli Schwartz; RAJA GIRYES; | code |
| 701 | Language-Guided Transformer Tokenizer for Human Motion Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we focus on motion discrete tokenization, whichconverts raw motion into compact discrete tokens—a process proven cru-cial for efficient motion generation. |
Sheng Yan; yong wang; Xin Du; Junsong Yuan; Mengyuan Liu; | code |
| 702 | CerDETR: Cell-Prior Empowered DETR for Cervical Lesion Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Cervical cancer remains a leading cause of female morbidityand mortality, making early lesion detection critical. Although RCNN-and YOLO-based detectors have improved performance, … |
Linyun Zhou; Jin Chen; Weihan Li; Hengrui Lou; Lingxiang Jia; Weijun Qin; Xiuming Zhang; Zunlei Feng; | code |
| 703 | PWM-ArtGen: Part World Model for Articulated Object Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Moreover, the limited scale and diversity of existing annotated datasets further hinder generalization to complex, real-world objects. To overcome these limitations, we propose to learn the joint distribution of visual dynamics and kinematic parameters. |
Wentao Zheng; Ancong Wu; | code |
| 704 | LENS: Adaptive Spatio-Temporal Zooming for Keyframe Sampling in Long-Form Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Concretely, LENS adaptively allocates a limited frame budgetbetween spatial zoom-ins, which highlight query-relevant regions withinindividual frames, and temporal zoom-outs, which expand the temporalscope through multi-frame aggregation, enabling the model to reasonacross multiple granularities while capturing both high-fidelity details andlong-range context. |
Ce Zhang; Jinxi He; Yaqi Xie; Katia Sycara; | code |
| 705 | PhysRAG: Enhancing Physics-Awareness in Video Generation Via Retrieval-Augmented Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In thiswork, we introduce PhysRAG, a novel pipeline that enhances physicalawareness in video generation through Retrieval-Augmented Generation(RAG).Furthermore,we construct a physical video database and develop a mechanism to injectphysical knowledge into a video diffusion model using learnable queries.Our method achieves state-of-the-art performance in both visual qualityand physical rule compliance, surpassing existing models in benchmarkssuch as PhyGenBench and VBench. |
Kexu Cheng; Zicheng Liu; Mingju Gao; Chunhe Song; Hao Tang; | code |
| 706 | Tuning Real-World Image Restoration at Inference: A Test-Time Scaling Paradigm for Flow Matching Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Although diffusion-based real-world image restoration (Real-IR) has achieved remarkable progress, efficiently leveraging ultra-large-scale pre-trained text-to-image (T2I) models … |
Purui Bai; Junxian Duan; Pin Wang; Jinhua Hao; Ming Sun; Chao Zhou; Huaibo Huang; | code |
| 707 | SceneDiff: A Benchmark and Method for Multiview Object Change Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We investigate the problem of identifying objects that havebeen added, removed, or moved between a pair of captures (imagesor videos) of the same scene at different times. |
Yuqun Wu; Chih-Hao Lin; Henry Che; Aditi Tiwari; Chuhang Zou; Shenlong Wang; Derek Hoiem; | code |
| 708 | ΛSplit: Self-Supervised Content-Aware Spectral Unmixing for Fluorescence Microscopy Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: In fluorescence microscopy, spectral unmixing aims to re-cover individual fluorophore concentrations from spectral images thatcapture mixed fluorophore emissions. Since classical … |
Federico Carrara; Mehdi Seifi; Florian Jug; | code |
| 709 | LaVPR: Benchmarking Language and Vision for Place Recognition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Beyond these limita-tions, standard systems cannot perform ’blind’ localization from verbaldescriptions alone, a capability critical for applications such as emer-gency response. To address these challenges, we introduce LaVPR, alarge-scale benchmark that extends existing VPR datasets with over650,000 rich natural-language descriptions. |
Ofer Idan; Dan Badur; yosi keller; Yoli Shavit; | code |
| 710 | EMAG: Self-Rectifying Diffusion Sampling with Exponential Moving Average Guidance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Exponential Moving Average Guidance (EMAG), a training-free mechanism that modifies attention at inference time in diffusion transformers, with a statistics-based, adaptive layer-selection rule. |
ANKIT YADAV; Huy Ta; Lingqiao Liu; | code |
| 711 | Predicting Consequences and Reinforcing Navigation Policies with Latent World Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a compatibilityprediction Latent World Model (LWM) for robot navigation that pre-dicts action-conditioned latent feature compatibility rather than recon-structing observations. |
Zengmao Wang; Wei Gao; Shuhan Shen; | code |
| 712 | Virtual Category-Guided Continual Generalized Category Discovery Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work,we introduce Virtual Category-Guided Continual Generalized CategoryDiscovery by adapting Virtual Category Learning (VCL) to the contin-ual setting. |
Jiahui Xiong; Qiuxia Lai; Hongsong Wang; | code |
| 713 | Enhancing Alignment for Unified Multimodal Models Via Semantically-Grounded Supervision Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We presentSemantically-Grounded Supervision (SeGroS), a fine-tuning frameworkdesigned to resolve the granularity mismatch and supervisory redundancyin UMMs. At its core, we propose a novel visual grounding map toconstruct two complementary supervision signals. |
Jiyeong Kim; Yerim So; Hyesong Choi; Uiwon Hwang; Dongbo Min; | code |
| 714 | Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To addressthis, we propose VISE (Visual Invariance Self-Evolution), a purely un-supervised self-evolving framework that directly regularizes the model’svisual conditioning policy through two complementary invariance-basedrewards: a geometric invariance reward that enforces spatial consistencyunder known transformations, and a semantic invariance reward thatpenalizes evidence-agnostic generation by requiring the model to recog-nize the absence of evidence when predicted regions are perturbed. |
Shravan Venkatraman; Ritesh Thawkar; Omkar Thawakar; Rao M Anwer; Hisham Cholakkal; Salman Khan; Fahad Shahbaz Khan; | code |
| 715 | SynFlow: Scaling Up LiDAR Scene Flow Estimation with Synthetic Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce SynFlow, a data generation pipeline for large-scalesynthetic LiDAR scene flow. |
Qingwen Zhang; Xiaomeng Zhu; ChenHan Jiang; Patric Jensfelt; | code |
| 716 | Calibrated Harmonic Overlaid Implicit Neural Representations for Multi-Dimensional Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Meanwhile, these methods fail to in-corporate proper physical priors to effectively alleviate spectrum bias. Toaddress these issues, inspired by the commonalities between deep periodicnetworks and generalized Fourier series, we propose a novel CalibratedHarmonic Overlaid Implicit Neural Representation (CHOIR). |
Honghang Chen; XIUJUN ZHANG; Xiaoli Sun; MINGQING XIAO; | code |
| 717 | Training-free Cross-domain Few-shot Segmentation Via Robust Semantic Representation and Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Cross-domain Few-shot Segmentation (CD-FSS) aims to tra-nsfer knowledge learned from source domain to distinct target domains,segmenting unseen target classes with only a few annotated samples.Although existing methods have made significant progress, they stillrely on training or fine-tuning processes, which incur high computationalcosts and risk overfitting. We observe that when powerful and general-purpose vision foundation models are incorporated into these methods,their performance shows only marginal improvement or even degradesdue to overfitting. |
Sujun Sun; Mingwu Ren; Haofeng Zhang; | code |
| 718 | SA-V2V: Training-Free Subject-Aware Video-to-Video Personalization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we uncover a key structural insight that the spatio-temporal attention matrix of DiT exhibits an inherent functional decomposition—intra-frame blocks primarily encode spatial information, while inter-frame blocks encode motion dynamics. |
Soobin Park; Seohyeon Yoo; Jiwon Kim; SeonHwa Kim; Kyong Hwan Jin; Eunju Cha; | code |
| 719 | 3D-LENS: A 3D Lifting-based Elevated Novel-view Synthesis Method for Single-View Aerial-Ground Re-Identification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose 3D Lifting-based Elevated Novel-view Synthesis (3D-LENS), a uni(cid:28)ed framework combining geometrically-consistent novel view synthesis that leverages large-scale 3D mesh reconstruction, with a robust representation learning scheme to mitigate synthetic-to-real bias. |
William Grolleau; Astrid Sabourin; Guillaume Lapouge; Catherine Achard; | code |
| 720 | Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Subsequently, we introduce controlled pruningas a diagnostic probe: by comparing the direct impact of removing dif-ferent parameter subsets on VLA performance without any fine-tuning,we establish a causal link between adaptation-induced divergence signalsand functional contributions. Based on the discovered modular hetero-geneities, we design a multi-module joint pruning scheme. |
Fengnian Zhang; Tao Huang; Siyu Xu; Zhong Jin; Chang Xu; | code |
| 721 | SkipGS: Post-Densification Backward Skipping for Efficient 3DGS Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose SkipGS with a novel view-adaptivebackward gating mechanism for efficient post-densification training. |
Jingxing Li; Yongjae Lee; Deliang Fan; | code |
| 722 | CoverPrune: Coverage-Driven Token Pruning for 3D VLMs Via Optimal Transport Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In thispaper, we propose a paradigm shift for 3D VLM token pruning: frommaximizing diversity to preserving visual evidence coverage. |
Peng Ling; Yingda Yin; Lingting Zhu; Weikai Chen; Shengju Qian; Zeyu HU; Xin Wang; Wenming Yang; | code |
| 723 | Local-to-global Cross-modal Coordination for Self-supervised RGB-T Tracking Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A commonsupervision strategy relies on forward-backward tracking loops, but inpractice, it is highly vulnerable to error accumulation and modality driftcaused by the inconsistent reliability of visible and thermal cues. To ad-dress this issue, we propose LGCTrack, a self-supervised RGB-T trackerunderpinned by a local-to-global coordination strategy within a closed-loop verification paradigm. |
Yueying Zhang; Timing Li; Bing Cao; Pengfei Zhu; | code |
| 724 | Evaluating Reasoning Coherence in Video Generative Models with Text and Visual Hints Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To bridge the gap in literature for missing reasoning coher-ence evaluation, we propose MME-CoF-Pro, a comprehensive videoreasoning benchmark to assess reasoning coherence in video models.Specifically, MME-CoF-Pro contains 303 samples across 16 categories,ranging from visual logical to scientific reasoning. |
Yu Qi; Xinyi Xu; Ziyu Guo; Siyuan Ma; Renrui Zhang; Xinyan Chen; Ruichuan An; Ruofan Xing; Jiayi Zhang; Haojie Huang; Pheng-Ann Heng; Jonathan Tremblay; Lawson Wong; | code |
| 725 | Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Here, we identify insufficient visual attention to safety-critical image regions as one of the key causes of multimodal safety failures. Building on this insight, we propose Multimodal Risk-Adaptive Steering (MoRAS), which enhances safety-critical visual attention via concise visual contexts for accurate multimodal risk assessment. |
Jonghyun Park; Minhyuk Seo; Chaewon YEO; Jonghyun Choi; | code |
| 726 | Domain Arithmetic: One-Shot VLA Adaptation Under Environmental Shifts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To reduce the burden of data curation andtraining, we propose an analogy-based method that adapts VLA modelsunder environmental shifts through weight vector arithmetic with domain-specific information addition, named Domain ARiThmetic (DART). |
Taewook Kang; Taeheon Kim; Donghyun Shin; Jonghyun Choi; | code |
| 727 | Open-Weather Robust 3D Detection Via Dual-Critic Diffusion Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this, we present Dual-Critic Guided Diffusion Alignment (DCDA), a weather-agnostic frame-work that learns to recover degraded LiDAR features toward a cleanmanifold.We further introduce astructured open-weather benchmark with held-out type–severity combi-nations and extensive experiments verify DCDA’s advantages. |
Shuyao Li; Chuanxing Geng; heyang sun; Qiang Zhou; Jingjing Gu; | code |
| 728 | AracNet: Revealing Debiasing Signals Across Layers with Shallow Monitors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose to exploit intermediate layers’ featuresfor achieving bias mitigation robust to the memorization problem, replac-ing the when to stop with a where to look paradigm. |
Vito Paolo Pastore; Massimiliano Ciranni; Enzo Tartaglione; Vittorio Murino; | code |
| 729 | SplatPainter: Interactive Authoring of 3D Gaussians from 2D Edits Via Test-Time Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing approaches based ondiffusion or optimization are ill-suited for this task, as they are oftenprohibitively slow, lack the precision for fine-grained control, or—mostimportantly—are fundamentally destructive to the original asset’s iden-tity. To address this, we introduce SplatPainter, a state-aware feedfor-ward model that enables continuous, high-fidelity appearance editing of3D Gaussian assets from user-provided 2D views. |
Yang Zheng; Hao Tan; Kai Zhang; Peng Wang; Leonidas Guibas; Gordon Wetzstein; Wang Yifan; | code |
| 730 | Scene Generation at Absolute Scale: Utilizing Semantic and Geometric Guidance From Text for Accurate and Interpretable 3D Indoor Scene Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present GuidedSceneGen, a text-to-3D generation frame-work that produces metrically accurate, globally consistent, and semanti-cally interpretable indoor scenes. |
Stefan Ainetter; Thomas Deixelberger; Edoardo Dominici; Philipp Drescher; Konstantinos Vardis; Markus Steinberger; | code |
| 731 | Histogram-constrained Image Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce Histogram-constrained Image Generation (HIG), a novel control mechanism thatfalls into the middle ground of control granularity. |
Haoming Liu; Yuanhe Guo; Yijia Cao; Shenji Wan; Hongyi Wen; | code |
| 732 | Holistic Optimal Label Selection for Robust Prompt Learning Under Partial Labels Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, when only partial labels are available, itsperformance is often limited by label ambiguity and insufficient supervi-sory information. To address this issue, we propose Holistic Optimal La-bel Selection (HopS), leveraging the generalization ability of pre-trainedfeature encoders through two complementary strategies. |
Yaqi Zhao; Haoliang Sun; Yating Wang; Yongshun Gong; Yilong Yin; | code |
| 733 | ArcAD: Anomaly-Rectified Calibration for Cold-Start Supervised Anomaly Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Under sucha regime, existing methods struggle to form compact normal boundariesand fail to effectively exploit supervised signals from rare defects. To ad-dress this challenge, we propose Anomaly-Rectified Cold-start AD (Ar-cAD), a plug-and-play calibration framework for reconstruction-basedIAD baselines. |
Ningning Han; Lei Fan; Jia Guo; Yunkang Cao; Xiu Su; Feng Cao; Donglin Di; Tonghua Su; | code |
| 734 | Learning to Attract and Repel: Dual Quality Margin Learning for Face Recognition (DQM-Face) Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Dual Quality Margin Learning forFace Recognition (DQM-Face), a novel framework that enables refinedattraction and repulsion dynamics during representation learning. |
El Ouanas Belabbaci; Bhavesh Wani; Philipp Terhörst; | code |
| 735 | TaxoMIL: Taxonomy-Constrained Learning for Hierarchical Whole Slide Image Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: TaxoMIL uti-lizes a dual-head Transformer decoder to generate coarse- and fine-leveldiagnostic text, and introduces taxonomy-guided objectives that explic-itly structure the label embedding space and strictly ground slide-level vi-sual representations within the clinical taxonomy. |
Chaeyeon Lee; Khang Quoc; Jinsol Song; Yosep Chong; Kwangil Yim; JIN TAE KWAK; | code |
| 736 | Rdm: Re-conceptualizing Distribution Matching As A Reward for Diffusion Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a novel paradigm by re-conceptualizing distribution matching as a reward, denoted as R . |
Linqian Fan; Peiqin Sun; Tiancheng Wen; Shun Lu; Chengru Song; | code |
| 737 | CapFrame: Text-Instructed Viewpoint Grounding in 3D Gaussian Scenes Via Geometric Pseudo Labels Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this limitation,we introduce a new task, Text-Instructed Viewpoint Grounding (TIVG),which aims to identify a 6-DoF camera pose in a 3D Gaussian scenewhose rendered frame aligns with a text instruction. To solve this task,we propose CapFrame, a partially differentiable framework that con-verts language into geometric pseudo labels for camera pose optimization.CapFrame follows a Retrieve–Translate–Refine pipeline: it retrieves rele-vant views and ranks them through a Question-Evaluation process withMLLMs, translates the instruction into orientation and layout pseudo la-bels, and refines the camera pose via differentiable optimization with lay-out and orientation losses in 3DGS. |
Jirong Li; Satoshi Ikehata; Shuhei Kurita; Ikuro Sato; | code |
| 738 | From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a method to generalize vision foundationmodels to fisheye cameras. |
Rit Gangopadhyay; Alex Wong; | code |
| 739 | ProtoMappingNet: Interpretable Hierarchical Prototypes Through Relational Prototype Mappings Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To quantitatively evaluate the learned hierarchy,we propose a relational evaluation framework consisting of HierarchicalRelational Consistency (HRC) and Hierarchical Spatial Stability (HSS),which measure cross-level relational consistency and spatial containmentwithout requiring part annotations. |
Jaehun Park; Jongmin Lim; Soobin Cha; Kwangsu Kim; | code |
| 740 | VisTa3D: A Dataset and Benchmark for Vision, Tactile, and 3D Point Clouds-based Thin Object Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To test if tactile data can help, we introduce the first visual-range-tactile 3D reconstruction model as a baseline.To test the extent of their errors, we collected the first thin object dataset comprising synchronized RGB images, depth maps, and tactile response maps, where each frame is associated with inertial measurements, camera pose and calibration, and groundtruth depth and segmentation maps obtained from laser scanning of thin objects. |
Shania Guo; Yeongsik Seo; Andrew Fu; Mei Hao; Iris Xia; Jiwon Lee; Xinyi Xie; Hyoungseob Park; Aaron Dollar; Alex Wong; | code |
| 741 | PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: While instruction-based image editing, enabled by multi-modal generative models, has advanced significantly, existing bench-marks lack comprehensive evaluation of physics-based … |
Shengbin Guo; Shaokang He; Chaoyue Meng; Shengpeng Xiao; Xunzhi Xiang; Shaofeng Zhang; Qi Fan; | code |
| 742 | ELHINN: Unifying Dense Crowd Simulation Across Scales Via Eulerian–Lagrangian Hydrodynamics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this,we propose the Eulerian–Lagrangian Hydrodynamics-Informed NeuralNetwork (ELHINN), a unified cross-scale framework that couples macro-scopic velocity evolution with microscopic trajectory refinement by usingevolved Eulerian velocity fields as physical priors to guide Lagrangian tra-jectories. |
Yanshan Zhou; Pingrui Lai; Jiaqi Yu; Cunyan Li; Hua Yang; Xiaoyun Zhang; | code |
| 743 | DTI: Dynamic Trajectory Initialization for Generative Face Video Super-Resolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Furthermore, an excessive pursuit of generic perceptual metrics often results in low fidelity. To address these issues, we present Dynamic Trajectory Initialization (DTI) paradigm for GFVSR, which reformulates GFVSR as an input-driven directional restoration. |
Yingwei TANG; Chen Yan; Wendi Liu; Qiang Hu; Xiaoyun Zhang; | code |
| 744 | UnderOneFacade: Worldwide Facade Semantic Segmentation Benchmark Dataset Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Un-derOneFacade, the largest cross-continental 3D facade benchmark todate, comprising centimeter-accurate point clouds with hierarchical, har-monized, and architecturally grounded semantic labels totaling 2.7 bil-lion annotated points. |
Yi Wang; Fan Wang; Prabin Gyawali; Ziyang Xu; Anna Klimkowska; Yixiong Jing; Wanru Yang; Filip Biljecki; Christoph Holst; Benjamin Busam; Brian Sheil; Olaf Wysocki; | code |
| 745 | Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our analysis shows that with RoPE, the attention similarity score is a highly structured and periodic function of token distance, so rescaling distances across resolutions moves token pairs to different regions of this periodic function, leading to incorrect attention scores. Motivated by this, we introduce Phase-Aligned Mixed-Resolution Attention (PMA), a trainingfree mechanism that stabilizes mixed-resolution attention. |
Haoyu Wu; Jingyi Xu; Qiaomu Miao; Dimitris Samaras; Hieu Le; | code |
| 746 | Geometry-Preserving in 3D Gaussian Splatting for LiDAR-Camera Extrinsic Calibration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a framework that preserves the metric geometry of the Gaussian proxy by aggregating multi-view LiDAR observations for dense depth supervision and blocking photometric gradients from updating the Gaussian spatial parameters. |
Kyoleen Kwak; Daeho Kim; Jeong Woon Lee; Hyoseok Hwang; | code |
| 747 | Learning Ego-Centric BEV Representations from A Perspective-Privileged View: Cross-View Supervision for Online HD Map Construction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Cross-View Supervision (CVS),a representation learning paradigm that transfers geometric and topolog-ical priors from an ego-aligned overhead perspective into camera-basedBEV encoders. |
Daniel Lengerer; Mathias Pechinger; Klaus Bogenberger; Carsten Markgraf; | code |
| 748 | Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing STR datasets and methods, typically builtaround regular scene text and fixed-template inputs, struggle to scale toWATER. Thus, we aim to advance this task from both data and modelperspectives. |
Xingsong Ye; Yongkun Du; Jiaxin Zhang; Haojie Zhang; Chong Sun; Chen Li; Jing LYU; Zhineng Chen; | code |
| 749 | FedNASP: Federated Vision-Language Navigation with Adaptive Step-wise Personalization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Federated learning (FL) protects sensitive (Vision-LanguageNavigation) VLN data without centralizing trajectories or instructions,but severe non-IID environments make personalized FL (pFL) necessary.Moreover, VLN poses several coupled challenges for personalized feder-ated learning, including environment heterogeneity, multimodal language-vision fusion, and long-horizon navigation with time-varying decisioncontexts. To address these challenges, we propose FedNASP, a step-wisepersonalized federated learning framework for VLN. |
Qingqian Yang; Hao Wang; Sai Qian Zhang; Jian Li; Yang Hua; Miao Pan; Tao Song; Zhengwei Qi; Haibing Guan; | code |
| 750 | LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Egocentric videos capture rich and diverse human–object in-teractions and have emerged as a fundamental resource for understand-ing human activities related to objects. In this … |
Shunya Kato; Taiki Miyanishi; Shuhei Kurita; Mahiro Ukai; Nakamasa Inoue; Chenhui Chu; | code |
| 751 | UniSim-SLAM: Feed-Forward SLAM with Unified Sim(3) Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce UniSim-SLAM,an integrated system that runs lightweight two-view keyframe trackingin the frontend and performs periodic multi-view submap refinementin the backend. |
inha Lee; Dongjae Jeong; Junhee Lee; Kyungdon Joo; | code |
| 752 | FeDepth: Federated Learning for Depth Estimation Under Robot Heterogeneity Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, this characteristic breaks theassumption of clearly separable client domains commonly used in clus-tered FL. To address this gap in robot perception, particularly in depthestimation, we introduce two realistic and unexplored non-IID scenariosthat reflect heterogeneity in terms of platform, environment, and depthdistribution. |
Ganghyeon Lee; inha Lee; Junhee Lee; Jeongeon Lee; Sung Whan Yoon; Kyungdon Joo; | code |
| 753 | DiffProxy: Multi-View Human Mesh Recovery Via Diffusion-Generated Dense Proxies Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We observe that the true bot-tleneck lies in the quality of intermediate representations, and that densepixel-to-surface correspondences can be e_x001B_ectively generated by repur-posing pre-trained di_x001B_usion models with rich visual priors. |
renke wang; zhenyu zhang; Ying Tai; Jun Li; Jian Yang; | code |
| 754 | StoryBlender: Inter-Shot Consistent and Editable 3D Storyboard with Spatial-temporal Dynamics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present StoryBlender, a grounded3D storyboard generation framework governed by a Story-centric Re-flection Scheme. |
Bingliang Li; Zhenhong Sun; Jiaming Bian; Yuehao Wu; Yifu Wang; HONGDONG LI; Yatao Bian; Huadong Mo; Daoyi Dong; | code |
| 755 | Zero-shot Depth from Defocus Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Unlike previous works overfittingto a certain dataset, this paper focuses on the challenging and practi-cal setting of zero-shot generalization. |
Yiming Zuo; Hongyu Wen; Venkat Subramanian; Patrick Chen; Karhan Kayan; Mario Bijelic; Felix Heide; Jia Deng; | code |
| 756 | InFlux++: Real and Synthetic Data for Estimating Dynamic Camera Intrinsics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address both gaps, we present InFlux++, consisting oftwo components. |
Erich Liang; Caleb Kha-Uong; Chinmaya Saran; Sreemanti Dey; David Liu; Junhan Ouyang; Benjamin Zhou; Jia Deng; | code |
| 757 | GARDEN: Gravity-Aligned Reconstruction of Disentangled ENvironments from RGB Images Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose GARDEN, an RGB-only framework that reformulates reconstruction as physicallygrounded scene factorization and outputs a structured hybrid scene representation. |
Jiahao Sun; Dingkun Wei; Zehong Shen; Hongyu Zhou; Yujun Shen; Liang Li; | code |
| 758 | SGC-Lane: Monocular 3D Lane Detection with Standard-Definition Map Guidance and Lane Completion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, their in-herent inaccuracies and representation as road-level centerlines, whichmisalign with the actual lanes to be perceived, pose significant challengesfor direct integration. To tackle this, we propose SGC-Lane, a novelframework that advances monocular 3D lane detection through SD mapguidance and lane completion. |
Fuqiang Jiang; Wei Li; Tianyao Zhao; Yu Hu; | code |
| 759 | SCDL: Synergistic Confidence-Dispersion Learning for Semi-Supervised Video Polyp Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In addition, inherent challenges of endoscopic videos, such as low contrast, significant inter-frame variations, and consecutive low-quality frames, further exacerbate the above issues. To address these shortcomings, we propose a unified semi-supervised video polyp segmentation framework, dubbed synergistic confidence–dispersion learning (SCDL), which aims to simultaneously improve pseudo-label reliability, reduce regional supervision bias, and alleviate ambiguities caused by low-contrast features. |
Yuanqin he; Yuhua Zhang; Huisi Wu; Jing Qin; | code |
| 760 | Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing benchmarks often prioritize high-level planning or spatial reasoning, leaving the fine-grained action intelligence required for embodied physical interaction underexplored. To address this gap, we introduce CFG-Bench, a new benchmark designed to systematically evaluate this crucial capability. |
Dayong Liu; Chao Xu; Weihong Chen; Suyu Zhang; Juncheng Wang; Jiankang Deng; Baigui Sun; Yang Liu; | code |
| 761 | Frames2Residual: Spatiotemporal Decoupling for Self-Supervised Video Denoising Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Self-supervised video denoising methods typically extend image-based frameworks into the temporal dimension, yet they often struggleto integrate inter-frame temporal consistency with intra-frame spatialspecificity. |
Mingjie Ji; Zhan Shi; Kailai Zhou; Zixuan Fu; Xun Cao; | code |
| 762 | Learning Semantic-Robust Change Detection Via Semantic-Invariant Self-Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose SCDistill, aframework for learning semantic-robust change detection via semantic-invariant self-distillation. |
Jiuhe Qu; Yingping Liang; Ying Fu; | code |
| 763 | Argus: Metric Panoramic 3D Reconstruction for Indoor Scenes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Realsee3D, a hybrid dataset of 10K indoor scenes (1K real, 9K synthetic) with 299K panoramic viewpoints and precise metric annotations, and Argus, a feed-forward network trained on it for metric panoramic 3D reconstruction. |
Xi Li; Linyuan Li; Yan Wu; Tong Rao; Kai Zhang; Xinchen Hui; Cihui Pan; | code |
| 764 | Fidelity- and Perception-Aware Local Implicit Attention for Arbitrary-Scale Image Super-Resolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However,these diffusion techniques frequently introduce the risk of structural hal-lucinations. To address these issues, we propose Fidelity- and Perception-Aware Local Implicit Attention (FPLIA), a framework that effectivelyintegrates fidelity-oriented features into a diffusion pipeline to produce re-alistic and faithful reconstructions for ASISR. |
Yu-Syuan Xu; Hao-Lun Sun; Hao-Wei Chen; Hsien-Kai Kuo; Chun-Yi Lee; | code |
| 765 | InterEdit: Navigating Text-Guided Multi-Human 3D Motion Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present InterEdit,a synchronized classifier-free conditional diffusion model for TMME.We introduce the task of multi-person 3D motion editing,where a target motion is generated from a source and a text instruction.To support this, we propose InterEdit3D, a new dataset with man-ual two-person motion change annotations, and a Text-guided Multi-human Motion Editing (TMME) benchmark. |
Yebin Yang; Di Wen; Lei Qi; Weitong Kong; Junwei Zheng; Ruiping Liu; Yufan Chen; Chengzhi Wu; Kailun Yang; Yuqian Fu; Danda Paudel; Luc Van Gool; Kunyu Peng; | code |
| 766 | FPicker: Topology-Guided Evolution for Filament Tracing in Low-SNR Microscopy Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Automating filament tracing in Cryo-Electron Microscopy(Cryo-EM) is essential for 3D helical reconstruction but challenged byintersecting topologies and extremely low Signal-to-Noise Ratios (SNR =σs2 /σn2 < 0.1 or -10 dB). Existing paradigms fail: pixel-wise segmenterssuffer from severe topological fracturing, box-based detectors face ghostcenter drift, sequential trackers derail due to error accumulation, and tra-ditional active contours collapse under artificial closed-curve constraints.To resolve these bottlenecks, we present FPicker, the first topology-guided framework reconciling these incompatibilities. |
Tingyin Zhao; Mingtao Huang; Yuan Shen; | code |
| 767 | Accelerated Likelihood Maximization for Diffusion-based Versatile Content Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Generating diverse, coherent, and plausible content from par-tially given inputs remains a fundamental challenge for diffusion models.Existing approaches face clear limitations: training-based approachesoffer strong task-specific results but require costly computation, andthey generalize poorly across tasks. Training-free approaches offer betterefficiency, but they do not explicitly optimize over unobserved variables,leading to globally inconsistent results. |
Hyunsoo Lee; Inwoo Hwang; Young Min Kim; | code |
| 768 | Closing The Capacity–Convergence Gap: Globally Optimal Configuration of Implicit Neural Representations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce OptiINR, the first framework that recasts INR configuration as a global optimization problem over a mixed-variable space of discrete activation choices (e.g., SIREN, WIRE, FINER) and their continuous parameters. |
Sipeng Chen; Yan Zhang; Shibo Li; | code |
| 769 | Rethinking Robust Adversarial Concept Erasure in Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Concept erasure methods aim to remove specific unsafe tar-get concepts in diffusion models while preserving image generation utility.To address the vulnerability that erased concepts can be easily recoveredunder adversarial attacks, adversarial concept erasure methods integrateadversarial optimization into the concept erasure process. |
Qinghong Yin; Yu Tian; Heming Yang; Xiang Chen; Xianlin Zhang; Yue Ming; Xueming Li; Yue Zhang; | code |
| 770 | TooBad: Backdoor Diffusion Models with Ultra-Low Poison Rate and Imperceptible Trigger Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper proposes TooBad (trigger optimizationfor backdoor diffusion models), a backdoor framework which introducesa novel DM-tailored trigger optimization technique to dramatically en-hance the performance of backdoor attacks on DMs. |
Vu Truong; Long Bao Le; | code |
| 771 | Towards Interactive Global Geolocation Assistant Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Driven by the imperative to actualize thisinteractive paradigm, we introduce MG-Geo, the first large-scale mul-timodal geolocation dataset explicitly structured for spatial reasoning.Comprising 4.87M geo-tagged Meta entries, 70K image-grounded Cluesamples, and 73K multi-turn Dialog samples across 210 countries andterritories, MG-Geo separates large-scale geographic alignment fromreasoning-oriented supervision. |
Zhiyang Dou; Zipeng Wang; Xumeng Han; Guorong Li; Zhenjun Han; Zhipei Huang; | code |
| 772 | Streaming Dense Voxel Representations for 3D Occupancy Prediction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we explore dense voxel streaming for accurateand efficient 3D occupancy prediction. |
Seokha Moon; Janghyun Baek; Yujin Jeong; Daewon Chae; Giseop Kim; Jungbeom Lee; Jinkyu Kim; Sunwook Choi; | code |
| 773 | ZeroSplat: Generalized Referring Segmentation in 3D Gaussian Splatting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To dismantle these bottlenecks, we propose ZeroSplat, a novel training-free and zero-feature framework.To facilitate comprehensive evaluation of this new task, we construct two new benchmarks: GR-LERF and GR-ScanNet. |
Jiayu Ding; Meilu Song; Xiaoyi Zhang; Hongbo Jin; Yichen Jin; Xiangtian Si; | code |
| 774 | HorizonRelight: Relighting Long-horizon Videos Consistently Via Diffusion Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Diffusion-based video relighting enables controllable relighting from a single input video, but modern video diffusion backbones are trained on short clips and applied to long-horizon videos through chunked sliding-window inference, often causing temporal discontinuities at chunk boundaries. We address this by reframing long-horizon relighting as temporally conditioned latent domain 38 translation. |
Jing Yang; Mayoore Jaiswal; Zian Wang; Xiao Zeng; Yajie Zhao; Jianyuan Min; Rochelle Pereira; | code |
| 775 | Dual-Prior Guided Null-Space Learning with Mixture-of-Splines for Arbitrary Medical Slice Super-Resolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address both concerns, this paper presents the Dual-PriorNull-Space Learning (DP-NSL) framework, which reformulates the taskas a constrained recovery process guided by two complementary priors.A Measurement-Consistent Projection (MCP) enforces a DeterministicObservation Prior : the reconstruction undergoes an exact orthogonalprojection that reproduces every acquired slice with zero error, confin-ing all learned details to the unobservable null space. |
Haofei Song; Siyuan Xu; Xintian Mao; ShaoJie Guo; Qingli Li; Yan Wang; | code |
| 776 | Histocomponent-driven Universal Model for Virtual Immunohistochemistry Multiplex Staining Via Joint Manifold Evolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Virtual immunohistochemistry (IHC) multiplex staining hasemerged as a highly promising non-destructive solution in digital pathol-ogy. However, existing universal IHC staining … |
Jiajun Cen; Siyuan Xu; Lili Gao; Yan Wang; | code |
| 777 | MambaRaw: Selective State Space Modeling for Efficient 4K RAW Image Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although existing metadata-based reconstruction frameworks can exploit this side information when recovering raw images, their context models often become computationally expensive especially at high resolution, e.g., 4K raw image, given that attention mechanisms scale quadratically with feature maps, hindering its practical application. To address these limitations, we propose MambaRaw, a JPEG-conditioned metadatabased raw image reconstruction framework that uses State Space Models (SSMs) to estimate entropy parameters efficiently. |
Peize Li; Fanhu Zeng; Tongda Xu; XINJIE ZHANG; Xingtong Ge; Haotian Zhang; Xingguo Xu; Yan Wang; | code |
| 778 | DPGS: A Diffusion-Prior Guided Framework for Large-Scale 3D Gaussian Splatting Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose DPGS, aholistic framework that orchestrates visual foundation model priors andgenerative diffusion priors to achieve robust reconstruction. |
Shidong Zhang; shuaixin li; Juntong Qi; Haoxin Zhang; Xiao zhang; Xiaozhou Zhu; Wen Yao; | code |
| 779 | DA-MergeLoRA: Hypernetwork-Based LoRA Merging for Few-Shot Test-Time Domain Adaptation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This setting is more realistic than typical domain adap-tation setups, which assume access to target data during source training.However, prior FSTT-DA approaches fail to effectively leverage sourcedomain-specific knowledge, relying on shallow batch normalization up-dates, prompt-based methods that treat the model as a black box, orensembling strategies that do not capture cross-domain relationships. Toaddress these limitations, we introduce a new FSTT-DA framework thatintegrates LoRA fine-tuning with model merging. |
Siobhan Reid; Zhixiang Chi; Li Gu; Omid Heidari; Ziqiang Wang; Yang Wang; | code |
| 780 | FusionTrack: Collaborative Multi-Object Tracking with Arbitrary Multi-UAVs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: For inference, we design View-aware Hierarchical Clustering with neighborfiltering to ensure intra-view exclusivity and inter-view consistency in cross-viewassociation. |
Xiaohe Li; Pengfei Li; Kaixin Zhang; Jiahao Li; Zide Fan; | code |
| 781 | EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce EgoSAT, the first comprehensive benchmarkfor egocentric video reasoning in streaming settings, designed to evaluatethe capabilities of modern vision–language models (VLMs). |
Yijia Lei; Jinzhao Li; Yichi Zhang; Jiacheng Hua; Yin Li; Miao Liu; | code |
| 782 | OmniLife360: A Benchmark for 3D Reconstruction from In-the-Wild 360° Captures Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Extensive evaluations reveal that state-ofthe-art methods excel in controlled settings but suffer significant degradation in dynamic scenarios. Motivated by these findings, we propose Omni4DGS, a motion-aware Gaussian Splatting framework that combines motion decomposition with lightweight supervision to better handle motion variations, dynamic distractors, and long-horizon pose drift. |
Zonglin Zhao; Bowen Zhang; Yatai Li; Chao.Liang Chao.Liang; Zhipeng Zhang; | code |
| 783 | Prosthesis-Aware 3D Human Pose Estimation: A Dataset and Benchmark for RSP Users Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We formally define the task of prosthesis-aware 3Dpose estimation, evaluate representative methods in a zero-shot setting,and confirm their individual limitations. |
Yilin Wen; Kechuan Dong; Fumiya Suginaka; Ken Endo; Yusuke Sugano; | code |
| 784 | Open-Vocabulary 3D Object Detection with Co-Distillation Discovery and Dual Guidance Robust Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although effective,this pipeline often suffers from inaccurate localization and mismatchedclassification during the discovery stage, which subsequently limits theperformance of the model training stage. To address these limitations,we advocate for improving both the reliability of novel object discov-ery and the robustness of model training, and propose an innovativeframework. |
Shangbo Yuan; Jie Xu; Xiaofeng Zhu; Na Zhao; | code |
| 785 | Breaking The Model Forgetting Cycle in Long-Incremental 3D Object Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We investigate this failure and reveal a detrimental self-reinforcing cycle: data distribution shift of novel classes causes modelforgetting on old classes, which further produces accumulated error inpseudo-labeling that exacerbates model degradation. To address thisissue, we draw inspiration from the human learning process and proposethe Learning-Dynamics-driven Memory and Review (LDMR) framework.LDMR monitors per-class detection quality at periodic training check-points and uses these learning-dynamics signals to drive two innova-tive mechanisms, namely (i) human-like intra-stage review that divideseach incremental stage into multiple sub-stages’ training and concen-trates on remembering the most-forgotten objects, and (ii) scene-awarecross-stage memory evolution that evolves a memory bank to transferknowledge between two consecutive stages by jointly considering scenelearnability and diversity. |
Peisheng Qian; Jie Xu; Xulei Yang; Na Zhao; | code |
| 786 | Low-Rank Ternary Adaptation for Fine-Tuning Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current approaches either require dequantization,restoring low-bit base weights to higher precision to merge with adap-tation weight, or update only quantization parameters, preventing amerged model that remains ternary. We propose ternary multiplica-tive adaptation, which represents discrete updates of ternary weightssuch as sign flips or zeroing through a low-rank Kronecker factoriza-tion into two small ternary matrices applied element-wise to ternaryweights. |
Alexandru-Dragos Manolache; Yunqiang Li; Jan van Gemert; | code |
| 787 | On The Real-world Generalisability of Optical Flow Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Here, we investigate the severity of this mismatch.To address this, we build a real-world evaluation benchmark and evaluate the real-world generalisability of a broad set of recent optical flow models using standard checkpoints. |
Petter Reijalt; Alexander Gielisse; Rickard Karlsson; Jan van Gemert; | code |
| 788 | Beyond Language: Grounding Referring Expressions with Hand Pointing in Egocentric Vision Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In natural egocentric engagements,hand-pointing combined with speech forms the most intuitive referringmechanism. To bridge this gap, we introduce EgoPoint-Ground, thefirst large-scale multimodal dataset dedicated to egocentric deictic visualgrounding. |
LING LI; Bowen Liu; Zinuo Zhan; Peng Jie; Jianhui Zhong; Kenglun Chang; Zhidong Deng; | code |
| 789 | MedGSSR: Generalizable Medical Image Super-Resolution 3D Reconstruction Via Hierarchical Feed-forward Gaussian Splatting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Medical 3D Super-Resolution (Med3DSR) offers a computational alternative, but existing methods commonly rely on per-subject optimization, pretrained priors, or coordinatebased implicit representations, which compromise anatomical fidelity and limit efficiency. To address these limitations, we present MedGSSR, a fully end-to-end feed-forward framework that represents volumes as an explicit 3D Gaussian field for Med3DSR. |
Chengkai Wang; Luoyu Hong; Yiting Zhao; Jiamin Wang; Xiang Feng; Feiwei Qin; Zhenzhong Kuang; Xuefei Yin; Ali Bashashati; Yanming Zhu; | code |
| 790 | Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To bridge the gap, we propose SmartVL, a unifiedadaptive inference framework that jointly controls vision token numberand model compute capability in response to varying input contents andcompute budgets. |
Pengcheng Wang; Zhiquan Wang; Jayoung Lee; Zhuoyan Xu; Ran Xu; Saurabh Bagchi; Yin Li; Somali Chaterji; | code |
| 791 | Tesselating The Earth Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To bridge the gap between local spatialstructure and global semantic understanding, we introduce global seman-tic tokens: a set of shared learnable concept tokens that distill semanticknowledge from the satellite imagery into a compact vocabulary the lo-cation encoder can reference at inference, enabling geographically distantsites covering similar environments to share semantics. |
Daniel Cher; Hamza Iqbal; Eric Xing; Brian Wei; Nathan Jacobs; | code |
| 792 | UniTriSplat: A Unified 3D Gaussian Splatting Framework with Uniform Spherical Rasterization for Universal Cameras Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing 3D Gaussian Splatting (3DGS) frameworks rely oncamera-speci_x001C_c rasterization, su_x001B_ering from inconsistent solid-angle sam-pling and degraded performance across heterogeneous camera models(e.g., perspective, _x001C_sheye, omnidirectional). To address this limitation,we propose UniTriSplat, a uni_x001C_ed 3DGS framework for universal camerasthat reformulates Gaussian splatting on the unit sphere via HEALPixdiscretization. |
Yipeng Zhu; Huajian Huang; Tristan Braud; Sai Kit Yeung; | code |
| 793 | PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce PhysMani-Bench, a dynamic manipulation benchmark with 16 tasks, and demon-strate a superior success rate over strong baselines in both simulationand real-world robot experiments. |
Peng Yun; Shouwang Huang; Hao Li; Jinxi Li; Jianan Wang; Bo Yang; | code |
| 794 | GeoSolver: Scaling Test-Time Reasoning in Remote Sensing with Fine-Grained Process Supervision Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent efforts to introduce Chain-of-Thought (CoT) reasoning to this domain have shown promise, yet ensuring the visual faithfulness of these intermediate steps remains a critical bottleneck. To address this, we introduce GeoSolver, a novel framework that transitions remote sensing reasoning toward verifiable, process-supervised reinforcement learning. |
Sun Lang; Ronghao Fu; Zhuoran Duan; Haoran Liu; Xueyan Liu; Bo Yang; | code |
| 795 | Category-Level Articulated Object Pose Estimation Via Pose–Shape Hypothesis Generation and Verification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing generative methods often neglect theinherent correlations between pose and shape, resulting in geometricallyinconsistent or physically implausible predictions. To address this limi-tation, we propose a novel generation-and-verification framework for thistask: First, a Pose-Shape Hypothesis Generator jointly models the dis-tributions of pose and shape to generate intrinsically coupled hypothesispairs, providing explicitly paired candidates for subsequent verification.Second, a Part-Level Point Cloud Denoising module recovers clean pointclouds from noisy and partial depth observations, ensuring robust inputsfor the verification stage. |
Kaifeng Tang; Chi Xu; Xin Ao; Yuting Ge; Tingrui Guo; Jun Zhou; | code |
| 796 | UniScale: Arbitrary-Scale Anomaly Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This failure occurs because extreme down-sampling in diffusion models causes the information of small anomaliesto be lost in the latent space. To address this, we introduce UniScale,a unified training and inference framework for high-fidelity industrialanomaly generation across arbitrary scales. |
Shilei Zeng; Linxin Guan; Xurui Li; Yaohan Tang; Yu Zhou; | code |
| 797 | DeCo: Zero-Shot Anomaly Generation Through Decoupling and Recoupling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing meth-ods suffer from two critical limitations, i.e., inaccurate anomaly informa-tion acquisition and uncontrolled anomaly-product fusion. To overcomethese challenges, we propose DeCo, which decouples the anomaly struc-ture from its source product, and explicitly recouples it with the nor-mal textures of the target product. |
Shilei Zeng; Xurui Li; Yaohan Tang; Yu Zhou; | code |
| 798 | SignNet-1M: Large-Scale Multilingual Sign Language Video Dataset with Downstream Benchmarks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce SignNet-1M, a large-scale augmenteddataset spanning ASL, CSL, and German Sign Language (DGS). |
Zhewen He; Junyi Hu; Haomian Huang; Zhenhua Li; Yushen Liu; Yi Fang; | code |
| 799 | Video-Text Alignment Model for Sign Language Translation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present VTaMo,a framework that introduces explicit multi-granularity alignment at threelevels: (1) local alignment via entropy-regularized optimal transport witha learnable null token for fine-grained frame-to-token correspondences;(2) global alignment via a learnable orthogonal transformation that cali-brates embedding space geometry through Earth Mover’s Distance; and(3) position-aligned contrastive learning for discriminative token-levelrepresentations. |
Junyi Hu; Zhewen He; Haomian Huang; Yi Fang; Aoxiang Yang; | code |
| 800 | Tricam-rPPG: A Multimodal Multispectral Dataset for Remote Photoplethysmography Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introduces Tricam-rPPG, a multimodal datasetdesigned to support systematic studies of remote photoplethysmography(rPPG) using multispectral imaging. |
Abhijit Sarkar; Surendrabikram Thapa; Ishtiaque Ahmed Khan; Yogesh Deshpande; Amos Abbott; | code |
| 801 | Learning to Deny: Action Denial in Multimodal Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Yet their ability to deny anaction by recognizing when an activity is not happening despite strongcontextual cues remains largely unexplored. We introduce UCF101-AD, a large-scale benchmark consisting of paired Action-Presence andAction-Denial clips, designed to evaluate this capacity for denial. |
Raiyaan Abdullah; Shehreen Azad; Yogesh Rawat; | code |
| 802 | SONIC: Spectral Optimization of Noise for Inpainting with Consistency Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a novel training-free method for inpainting with off-the-shelf text-to-image models. |
Seungyeon Baek; Erqun Dong; Shadan Namazifard; Mark J Matthews; Kwang Moo Yi; | code |
| 803 | Condensing Large-Scale Datasets Directly with Minimal Information Loss Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite their scalability to large-scaledatasets, these methods suffer from prohibitive computational overheadand poor cross-architecture generalization. In this paper, we reveal theroot cause of these bottlenecks: the implicit dual-compression process,from data to model and back to images, inherently induces severe infor-mation loss. |
Xinyi Shang; Peng Sun; Bei Shi; Zixuan Wang; Tao Lin; | code |
| 804 | The Devil Is in The Dark Pixels: Toward Brightness Bias-Robust Denoising Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we reveal an important yet overlooked problem in image denoising: under signal-dependent camera noise models, dark regions suffer from inherently low Signal-to-Noise Ratio (SNR), as signal intensity decays far faster than noise variance diminishes, making detail recovery in dark areas fundamentally challenging. |
Sungjun Cho; Zhuangzhuang Chen; Xiaomeng Li; | code |
| 805 | RoME: Robust Mixture of Low-Rank Experts Against Multiple Adversarial Perturbations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose Robust Mixture of Low-RankExperts (RoME), where each expert is a low-rank additive update tothe shared backbone, allowing it to capture threat-common featureswhile experts focus on threat-specific information. |
Woo Jae Kim; Kyle Min; Suhyeon Ha; Joonsung Jeon; Sung-eui Yoon; | code |
| 806 | Visible Yet Unrecognizable: Frequency-Selective Facial Privacy Via Attention Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing methods fail to satisfy both requirements; they either irreversibly degrade visual quality, rendering images impractical, or alter facial identity to such an extent that the protected image bears no resemblance to the original subject. To address this fundamental privacy-utility tradeoff, we propose MIRAGE (Multi-scale Identity Removal via Attention-Guided Encoding), a frequency-selective facial privacy framework grounded in the observation that identity-discriminative features reside predominantly in low-frequency image components, while visual appearance is encoded in high-frequency components. |
Atul Kumar; Akshay Agarwal; | code |
| 807 | TopoFuse: Topology-Aware Tri-Planar Fusion for 3D Cryo-Electron Tomography Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce TopoFuse,which reframes topology as a differentiable projection operator rather thana loss penalty. |
Rohit Kumar Salla; Neelesh Gupta; Xingjian Li; Min Xu; | code |
| 808 | Interpretation-Oriented Cloud Removal Via Observation-Anchored Residual Flow with Geo-Contextual Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing CRapproaches often prioritize visual realism while overlooking their im-pact on subsequent analytical tasks, leading to semantic drift and de-graded downstream performance. To address this issue, we propose Geo-Anchored Cloud Removal (GACR), a unified framework that jointlyensures faithful reconstruction and robust interpretability. |
Ziyao Wang; Maonan Wang; Yucheng He; Xianping Ma; Ziyi Wang; Hongyang Zhang; Yirong Chen; Man On Pun; | code |
| 809 | Infinite Gaze Generation for Videos with Autoregressive Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Predicting human gaze in video is fundamental to advanc-ing scene understanding and multimodal interaction. While traditionalsaliency maps provide spatial probability … |
JENNA KANG; Colin Groth; Tong Wu; Finley Torrens; Patsorn Sangkloy; Gordon Wetzstein; Qi Sun; | code |
| 810 | Aligning Anything: Hierarchical Motion Estimation for Video Frame Interpolation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This approach elaborately incorporates object priors derived from open-world knowledge models, such as Segment Anything Model (SAM), to facili-tate latent object-level motion learning. |
Mengshun Hu; Zhihang Zhong; Yansheng Qiu; Zheng Wang; Xiao Sun; | code |
| 811 | FuDU: A Fuzzy Dual-dimension Uncertainty Framework for Streaming Active Learning in Industrial Defect Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Ensuring the reliability of deep learning models in real-timeindustrial defect detection is critical for high-stakes quality inspection.To mine uncertain samples within continuous … |
Zhaoyang Wang; Haiyong Chen; Binyi Su; Xinwei Lyu; | code |
| 812 | Boosting 3D Foundation Models with Featureless Pose Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Edge-based Pose Optimization (EPO),a trackless geometric optimization framework specifically designed toboost the Structure-from-Motion reconstructions generated by 3D Foun-dation Models. |
Mattia D'Urso; Christian Sormann; Mattia Rossi; Friedrich Fraundorfer; | code |
| 813 | SFKD: Spatial–Frequency Joint-Aware Heterogeneous Knowledge Distillation Via Multi-Level Wavelet Spectral Interaction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To better leverage the spatial information encoded inheterogeneous representations, we propose a Spatial–Frequency Joint-Aware Heterogeneous Knowledge Distillation framework (SFKD). |
Cuipeng Wang; Haipeng Wang; | code |
| 814 | PointLAM: Local Attentive Mamba for Efficient Point-based 3D Object Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: 3D object detection from LiDAR point clouds faces a fun-damental dilemma: voxel-based methods achieve efficiency at the costof geometric quantization, while point-based methods preserve fidelitybut suffer from prohibitive computational bottlenecks. |
Xuanming Shang; Weijia Zhang; Chao Ma; | code |
| 815 | MorphGS: Morphology-Adaptive Articulated Motion Transfer from Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Mor-phGS, a framework that formulates motion retargeting as a target-drivenanalysis-by-synthesis problem, directly optimizing target morphology andpose through image-space supervision. |
Taeyeon Kim; Youngju Na; Jumin Lee; Sebin Lee; Minhyuk Sung; Sung-eui Yoon; | code |
| 816 | Task-Agnostic Incremental Vision-Language Object Detection Via Prompt Augmentation and Distribution-Aware Fusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To tackle thesechallenges, we propose TADA, a novel modular framework. |
Yonghan Jiang; Zhengyuan Xie; Wenchu Liu; Linlan Huang; Fei Yang; Xialei Liu; | code |
| 817 | OmniFall: From Staged Through Synthetic to Wild, A Unified Multi-Domain Dataset for Robust Fall Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We evaluate fine-tuned models as wellas much larger zero-shot multimodal LLMs. |
David Schneider; Zdravko Marinov; Moritz Mistol; Zeyun Zhong; Alexander Jaus; Rodi Düger; Rafael Baur; M. Saquib Sarfraz; Rainer Stiefelhagen; | code |
| 818 | Human-Level Accuracy, Non-Human Strategies: Revealing Model-Human Divergence in Video Physical Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a distri-butional evaluation framework that treats model seeds and human ratersas populations, enabling comparison of consensus, uncertainty, and strat-egy. |
Fanhong Li; Shurui Zheng; Yinzi Yinzi; Junbo Cui; Lei Ji; Jia Liu; | code |
| 819 | Auto3R: Automated 3D Reconstruction and Scanning Via Data-driven Uncertainty Quantification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Auto3R, a data-driven uncertainty quantification model that is designed to automatethe 3D scanning and reconstruction of scenes and objects, includingobjects with non-lambertian and specular materials. |
Chentao Shen; Sizhe Zheng; Bingqian Wu; Yaohua Feng; Yuanchen Fei; Mingyu Mei; Hanwen Jiang; Xiangru Huang; | code |
| 820 | Personalized Reward Modeling for Text-to-Image Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Recent text-to-image (T2I) models generate semantically co-herent images from textual prompts, yet evaluating how well they alignwith individual user preferences remains an open … |
Jeongeun Lee; Ryang Heo; Dongha Lee; | code |
| 821 | Histopathology Multi-modal Embedding for Pathology Composed Retrieval Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While Multimodal Large Language Models (MLLMs)o!er deep-fusion capabilities, directly applying them exposes a TaskMismatch and a Domain Mismatch. To resolve these challenges,we propose HOMIE, a model-agnostic adaptation framework that trans-forms any generative MLLM into a specialized pathology retrieval expert.Evaluated on our newly introduced PCR Benchmark, a lightweight 2B-parameter HOMIE variant substantially outperforms existing paradigms,surpassing specialized 7B pathology MLLMs and dual-encoders by largemargins on composed retrieval, while maintaining strong performanceon traditional simple retrieval. |
Qifeng Zhou; Wenliang Zhong; Thao Dang; Hehuan Ma; Saiyang Na; Yuzhi Guo; Junzhou Huang; | code |
| 822 | Why Can’t I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we tackle a key failure mode:models predict verbs via object-driven shortcuts (i.e., relying on the la-beled object class) rather than temporal evidence. |
Geo Ahn; Inwoong Lee; Taeoh Kim; Minho Shim; Dongyoon Wee; Jinwoo Choi; | code |
| 823 | Open Your Eyes: Benchmarking The Detection of Fabricated Realities and Weaponized Ethics in VLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We ask a central question: does pro-cessing an artifact visually, rather than as text, degrade an agent’s abil-ity to resist embedded attacks? To study this, we introduce the M-IPIBenchmark—a suite of 2,600 high-fidelity visual artifacts encompass-ing two attack families: (1) Technically-Framed Attacks, where maliciouscommands are interwoven with genuine debugging workflows in terminalscreenshots, and (2) Ethics-Framed Attacks, a novel vector where adver-saries exploit alignment priors such as fairness mandates to override eval-uation policies—and use it to systematically analyze modality-dependentfailures. |
Amit Pandey; Aditya Mohan; Phani Sankar; | code |
| 824 | SceneHI: High-Resolution 3D-Consistent Scene Texturing with Controllable Illumination Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To enforce strict geometric coherence, we introducean exact analytical pixel-to-texel mapping that aligns diffusion trajecto-ries across multiple viewpoints. We utilize High-Resolution Latent Tex-tures (HRLTs) as a persistent canvas for gradually denoised textures,while camera views perform the denoising steps in latent pixel space.This ensures a shared base texture that can be subsequently refined tohigh resolution without compromising multi-view consistency. |
Athanasios Tragakis; Marco Aversa; Daniela Ivanova; Chaitanya Kaul; Roderick Murray-Smith; Daniele Faccio; Paul Henderson; | code |
| 825 | Disentangling Pictorial Cue Understanding from Language Bias in VLMs Via Depth Ordering Task Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we study depth perception of vision-language models (VLMs) to isolate the effects of pictorial depth cues and disentangle vision and language influences on model performance. |
Yiqian Liu; Iuliia Kotseruba; John K. Tsotsos; | code |
| 826 | ChronoFlow Policy: Unifying Past-Future Interaction Flow in Visuomotor Policy Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduceChronoFlow, a temporally unified representation that captures past,current, and future interaction dynamics through sparse 3D keypointsof both objects and the gripper. |
Bokai Lin; Yifu Xu; Xinyu Zhan; Hongjie Fang; Jialin Tian; Fu-Cheng Zhang; Yong-Lu Li; Cewu Lu; Lixin Yang; | code |
| 827 | Y-diff: Structure-Texture Decoupled Diffusion Distillation for H&E-to-pCLE Translation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Furthermore, when applied to pCLE image generation, existing cross-modal translation models consistently introduce artifacts and cellular structure distortions. To address this, we propose Y-diff, a novel generative framework translating widely accessible H&E-stained pathology images into the scarce pCLE modality with high fidelity, providing robust data support for computational pathology. |
Haodong Wang; Yan Wen; Hongen Liao; Fang Chen; Tianqi Huang; | code |
| 828 | Training-Free Task Classification for Multi-Task Model Merging Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we aim toclose the gap to expert performance without additional training or task-ID-access assumption. |
Jungyong Son; Jinwook Jung; Sungyong Baik; | code |
| 829 | Learning to Recover Task Experts from A Multi-Task Merged Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, from theperspective of task expert, we view parameter interference as parame-ter perturbation introduced to each expert during merging process. |
Jinwook Jung; Taegyu Kim; Kumju Jo; Sungyong Baik; | code |
| 830 | Omni-RRM: Advancing Omni Reward Modeling Via Automatic Rubric-Grounded Preference Synthesis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Omni-RRM, an Omni-modalRubric-grounded Reward Model that generates multi-dimensional re-ward signals across text, image, video, and audio. |
Zicheng Kong; Dehua Ma; Zhenbo Xu; Anwen Yang; Yiwei Ru; Haoran Wang; Zixuan Zhou; Fuqing Bie; Liuyu Xiang; Huijia Wu; Jian Zhao; Zhaofeng He; | code |
| 831 | Local Spacing-Aware Hungarian Matching for Stable Point-Supervised Crowd Counting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose LocalSpacing-Aware Hungarian Matching (SAH-matcher), a drop-in replace-ment that derives a local spacing prior from k-nearest-neighbor distancesand performs per-target rescaling of the geometric cost, inducing a moreselective effective matching region in dense areas while remaining toler-ant in sparse ones. To quantify assignment behavior, we introduce Com-petitive Ambiguity Score (CAS) and Hijacking Rate (HR) for within-epoch ambiguity and severe hijacking failures, and combine them withInstability Rate (IR) to measure cross-epoch consistency. |
Kai Jiang; Yiming Lin; Zurui Ao; | code |
| 832 | Fragmented Text Is Insufficient for Image Representation: Fine-Grained Correspondence in Multimodal Dataset Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Additionally,they usually use contrastive losses to pull together paired samples andpenalize non-paired samples, ignoring the textual information from un-paired samples that are semantically related to the corresponding images.To reduce the semantic gap, we introduce a multi-text fusion module thatstrengthens the cross-modal interaction between the image and text. Inaddition, we leverage uncertainty to adaptively guide the contribution ofnon-paired samples in the synthetic dataset, thereby improving the ef-fective utilization of information from these samples. |
Jingwei Fang; Yaxin Hou; Bo Han; Xu Zhang; Hui LIU; Junhui Hou; Yuheng Jia; | code |
| 833 | Rethinking Temporal Modeling in Visual Object Tracking Via Decoupled Auxiliary Supervision Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This imbalance leads to shortcut learning where the networkover-relies on current-frame appearance, causing temporal representationcollapse. To address this, we propose DASTrack, a framework featur-ing Decoupled Auxiliary Supervision (DAS). |
Dailing Zhang; Shiyu Hu; Honghao Fu; Xiaokun Feng; Yipei Wang; Kang Hao Cheong; Kaiqi Huang; | code |
| 834 | Robust 3DGS-based SLAM Via Adaptive Kernel Smoothing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we challenge the conventional notion in 3DGS-SLAM that rendering quality is the primary determinant of trackingaccuracy. |
Shouhe Zhang; Dayong Ren; WEN LI; Piaopiao Yu; Sensen Song; Kaikai Shao; Yurong Qian; | code |
| 835 | What CLIP Knows But Cannot Say: Recovering Negation from Frozen Intermediate Features Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We attributethis failure to a phenomenon we call Representational Collapse: by track-ing compositional divergence and visual alignment across the CLIP textencoder, we show that middle layers build compositional syntax, but thefinal layers collapse this structure as visual alignment rises, producinga syntax-blind final representation. To recover the lost negation signalwithout altering pretrained weights, we propose PeakPatch, a lightweightpost-hoc correction system that intercepts the encoder at its composi-tional peak while keeping CLIP fully frozen. |
Chen-yi Lu; Yueh-Shao Chen; Somali Chaterji; | code |
| 836 | IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, wereveal that motion assessment and refinement tasks can act as crucialbridges to enable knowledge flow from motion understanding to genera-tion. |
Yuanming Li; Qize Yang; Nan Lei; Shenghao Fu; Ling-An Zeng; Jian-Fang Hu; Xihan Wei; WEISHI ZHENG; | code |
| 837 | NanoGS: Training-Free and Lightweight Gaussian Splat Simplification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce NanoGS, a training-free andlightweight framework for Gaussian Splat simplification. |
Butian Xiong; Rong Liu; Tiantian Zhou; Meida Chen; Zhiwen Fan; Andrew Feng; | code |
| 838 | Geometric Foundation Model Distillation for Efficient Lunar 3D Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To bridge the dimensional mismatch betweenteacher and student, we propose a structured SVD-based initializationthat projects the teacher’s decoder weights into the student’s smallerlatent space, yielding a warm start that significantly improves conver-gence and final performance. |
Clémentine Grethen; | code |
| 839 | Compact and Structurally Transparent Cervical Cytology with Geometry-Driven Features and Closed-Form Attention Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Cervical cytology is practiced as a sequence of roles: screen-ing decides where to look, careful reading determines what is present,and reporting accounts for the decision. We … |
Dichao Liu; | code |
| 840 | Foundation Model Selection for Remote Sensing Via A Constraint-Aware Agent Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Foundation Models (FMs) are increasingly integrated into remote sensing (RS)pipelines for applications such as environmental monitoring, disaster assessment, and land-usemapping.However, selecting the most suit-able remote sensing foundation model (RSFM) for a specific task remains challenging dueto scattered documentation, heterogeneous formats, and complex deployment constraints.To address this, we first introduce the RSFM Database (RS-FMD), the first structuredand schema-guided resource covering over 160 RSFMs trained on various data modalities,spanning different spatial, spectral, and temporal resolutions, considering different learn-ing paradigms. |
Binger Chen; Tacettin Bök; Behnood Rasti; Volker Markl; Begüm Demir; | code |
| 841 | Physically Grounded 3D Generative Reconstruction Under Hand Occlusion Using Proprioception and Multi-Contact Touch Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a multimodal, physically grounded approach formetric-scale amodal object reconstruction and pose estimation undersevere hand occlusion. |
Gabriele Mario Caddeo; Pasquale Marra; Lorenzo Natale; | code |
| 842 | TORA: Topological Representation Alignment for 3D Shape Assembly Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce TORA, a topology-first representation alignment framework that distills relational structure from a frozen pretrained 3D encoder into the flow-matching backbone during training. |
Nahyuk Lee; Zhiang Chen; Marc Pollefeys; Sunghwan Hong; | code |
| 843 | SCALE: Semantic-Calibrated Guidance Enhancement for Prompt-Faithful Diffusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We then show that SAP can fail due to the loss of orthogonal corrective freedom, discarding high-dimensional components that are crucial for rectifying accumulated trajectory drift. Based on this diagnosis, we propose SCALE (Semantic-CALibrated Guidance Enhancement), a drop-in, training-free guidance mechanism. |
Tianhang Lu; Sudong Cai; Bingzhi Chen; Shao-Dong Shen; Chunting Liu; Longguang Wang; Bing Wang; | code |
| 844 | FedOT: Ownership Verification and Leakage Tracing Via Watermarks for Federated LDMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we pro-pose FedOT, the first framework for ownership verification and leakagetracing in federated LDMs. |
Wenlong Cheng; Yuan Gan; Yunqiu Xu; Jiaxu Miao; | code |
| 845 | SWIFT: Spatial-Window Integrated Frequency-aware Token Pruning for Efficient MLLMs on Edge Devices Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present Spatial-Window Integrated Frequency-aware Token Pruning (SWIFT), a seam-less plug-and-play framework designed for efficient token compression.SWIFT leverages two novel components: a Frequency-Aware Indicator(FAI), which identifies fine-grained details by estimating high-frequencyresiduals through matrix factorization, and a Spatial-Window Integra-tion (SWI) module, which prevents spatial structure collapse and posi-tional bias via localized retention. |
Guanglai Liu; Jubo Chen; Xiaosheng Yu; | code |
| 846 | RaysUp: Ultra-light Universal Feature Upsampling Via Geometry-Aware Ray Representation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing fea-ture upsampling approaches either degrade semantic fidelity or rely onVFM-specific retraining and heavy architectures, hindering efficiency andscalability. To address these challenges, we propose RaysUp, an ultra-lightweight, task-agnostic, and VFM-agnostic feature upsampling frame-work that reconstructs high-resolution feature maps at arbitrary resolu-tions. |
Ding Yuchuan; Linfei Li; Lin Zhang; Ying Shen; | code |
| 847 | Hierarchical Prompt Injector for Domain Generalization Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We achieve 70.62% and 72.74% mIoU on synthetic-to-real andreal-to-real benchmarks, respectively. |
Xin Kun Lin; Ruoyu Guo; Jiaqi Guo; Maurice Pagnucco; Yang Song; | code |
| 848 | Incentivizing Vision Language Models to Search for Long Video Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce VSeek, an agentic framework that transformslong-video question answering (LVQA) from a passive, single-pass per-ception task into a multi-turn retrieval process. |
Harsh Goel; S P Sharan; Sahil Shah; Minkyu Choi; Joungbin An; Kristen Grauman; Sandeep Chinchali; | code |
| 849 | CollectionLoRA: Collecting 50 Effects in 1 LoRA for Deployment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, wepropose a unified paradigm that distills numerous customized conceptsand fast-inference capabilities into a single LoRA, effectively resolvingcompositional conflicts while reducing deployment overhead. |
Fangtai Wu; Hailong Guo; Shijie Huang; Jiayi Song; Yubo Huang; Mushui Liu; Zhao Wang; Yunlong Yu; Jiaming Liu; Ruihua Huang; | code |
| 850 | MagnetGS-Mesh: High-Quality Multi-Object Mesh Reconstruction Via Adaptive Surface Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present MagnetGS-Mesh, an integrated framework thatdirectly recovers high-quality, object-wise surface meshes from 3D Gaus-sian Splatting (3DGS) while preserving photorealistic novel-view syn-thesis. |
Min-Su Park; Yeonho Han; Uijoon Jeong; Jun-Hyeong Park; Eun-Seok Ryu; | code |
| 851 | Reconstruction-Anchored Diffusion Model for Text-to-Motion Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Diffusion models have seen widespread adoption for text-driven human motion generation and related tasks due to their impres-sive generative capabilities and flexibility. However, … |
Yifei Liu; Changxing Ding; Ling Guo; Huaiguang Jiang; Qiong Cao; | code |
| 852 | Cube-Splat: High-Fidelity 360° Gaussian Splatting SLAM Via Cubemap Factorization and Adjoint-Consistent Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Cube-Splat, the firstpanoramic GS-SLAM framework that factorizes each 360◦ frame into acubemap of four fixed-orientation virtual pinhole views sharing a singleoptical center. |
Xiangfei Guo; Hao Shi; Yufan Zhang; Zhonghua Yi; maoyongqi maoyongqi; Xiaoting Yin; Kaiwei Wang; | code |
| 853 | STARLINC: Satellite Trail Artifact Removal Using Inter-Frame Correlation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Moreover, training new models fromscratch is impractical due to the lack of large-scale annotated astro-nomical datasets. To address these challenges, we introduce STARLINC,the first ML-based framework for satellite trail removal without requir-ing tedious pixel-level annotation of astronomical images. |
Shingeon Kim; Hyeyoon Lee; Dain Kwon; Kanghyun Choi; SunJong Park; Mi-Ryang Kim; Jeong-Eun Lee; Jinho Lee; | code |
| 854 | AnyStyle: A Single LoRA Is Sufficient for Image-Guided Style Transfer Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We further demonstrate that training-free structural guidance directly derived from the content image through the internal attention of pre-trained model outperforms a dedicated content LoRA adapter in terms of structural fidelity and computational efficiency. Building on these observations, we propose AnyStyle, a streamlined framework for image-guided style transfer. |
Yongwen Lai; Chaoqun Wang; | code |
| 855 | Decompose, Compare, and Decide: Multimodal LLMs Are Implicit Few-Shot Learners Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Multimodal Large Language Models (MLLMs) have demon-strated remarkable abilities when analyzing images, yet translating thesecapabilities to few-shot image classification remains challenging. To bridgethis gap, we present DeCoDe, a simple yet effective technique that en-ables off-the-shelf MLLMs to act as strong few-shot classifiers withoutany additional training. |
Yunhan Wang; Eshika Khandelwal; Edson Araujo; Walid Bousselham; Nina Shvetsova; Hilde Kuehne; | code |
| 856 | VVSim: A Large-Scale Aerial-Ground Dataset and Benchmark for Cooperative Perception Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, progress in this area is hindered by the lack of public datasetsand standardized benchmarks. To bridge this gap, we introduce VVSim,a large-scale dataset for AGCP that provides synchronized multimodaldata and state information from both vehicles and unmanned aerial ve-hicles (UAVs). |
Zengle Zhu; Zhen LI; Tianyi Huai; Tianshun Li; Zihang Xu; Liuqing Yang; Rongqing Zhang; Xinhu Zheng; | code |
| 857 | Honey, I Shrunk The Arc De Triomphe! Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We estimate camera poses andinitial depth maps for each scene using off-the-shelf methods, and re-cover absolute scale from geo-tagged metadata as well as known stereocamera baselines. |
Yuanbo Xiangli; Hanyu Chen; Xueqing Tsang; Noah Snavely; | code |
| 858 | ECoSim: Data Efficient Fine-Tuning for Controllable Traffic Simulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a lightweightcontrol adaptation framework that enables multi-modal controllability(sketch, latent behavior codes, and text) for pretrained state-of-the-artdiffusion and autoregressive traffic models. |
Yu-Hsiang Chen; WEI-JER Chang; Yi-Ting Chen; Masayoshi TOMIZUKA; | code |
| 859 | Don’t Teach Instability, Teach Robustness: Selective Sensitivity Gating for Adversarial Robust Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing ARD techniques relyon the assumption that the teacher’s decision boundary is a perfect geo-metric oracle. In this paper, we identify a critical gap from the stabilityperspective that such exact matching forces blind inheritance of teacher’ssensitivity noise, punishing the student’s stability even if they are morerobust. |
Jingqi Ji; Quan Kong; Chaojie Gu; Yuanchao Shu; Cong Wang; | code |
| 860 | Vision-TTT: Efficient and Expressive Visual Representation Learning with Test-Time Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Toaddress the challenge, we introduce a new linear-time sequence modelingmethod Test-Time Training (TTT) into vision and propose Vision-TTT,which treats visual sequences as datasets and compresses the visual to-ken sequences in a novel self-supervised learning manner. |
Quan Kong; Yanru Xiao; Yuhao Shen; Cong Wang; | code |
| 861 | Foundation-Guided Representation Alignment for Multimodal Medical Image Registration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce anovel Data ReAssembly strategy to transform volumetric medical imagesinto VFM input while preserving spatial information. |
Mengjie Guo; Xinxing Cheng; Wenqi Lu; Qingjie Meng; Guanyu Yang; Yang Chen; Ziyun Ding; Alejandro Frangi; Jinming Duan; | code |
| 862 | QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we proposeQSVideo, a unified framework that systematically addresses relevance,diversity, and temporal modeling in video retrieval. |
Wei Ao; Lan Wang; Vishnu Boddeti; | code |
| 863 | Noise-Robust Face Recognition Via Non-target Similarity Distribution Guided Sample Selection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Label noise is a major challenge in large-scale supervised face recognition, where weak or automatic annotations often introduce errors that mislead training. To address this, we propose a noise robust framework that performs noise detection and sample selection directly in cosine-similarity space. |
Fanglong Wu; Youqiang Gui; Cheng Peng; | code |
| 864 | Rethinking Prototype-based Similarity Learning for Few-Shot Object Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, they suffer from two fundamental limitations: (i) classconfusion arising from inter-class similarity margin collapse, and (ii)insufficient visual cues for precise localization, as similarity scorescapture only class-level semantic affinity while providing limited spa-tial information. To address these issues, we introduce two complemen-tary components. |
KunHo Heo; Seungjae Kim; Wongyu Lee; SuYeon Kim; MyeongAh Cho; | code |
| 865 | PriorEye: Geospatial Visual Priors for End-to-End Autonomous Driving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce geospatial visual priors, street-level visual con-text anchored to the intended driving route, providing visual-spatialforesight independent of real-time sensors. |
Kyuhwan Yeon; Benjamin Ramtoula; Daniele De Martini; | code |
| 866 | DiNBV-Grasp: Real-Time Distance-Aware Two-Stage Next-Best-View for Robotic Grasping Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This overlooks a critical factor: grasp perception quality is highly sensitive to viewing distance, and the optimal distance varies significantly across object categories and scales. To address this limitation, we propose DiNBVGrasp, a real-time, distance-aware two-stage NBV framework. |
Zilong Xie; Jingyu Gong; Xin Tan; Zhizhong Zhang; Yanyun Qu; Lizhuang Ma; Yuan Xie; | code |
| 867 | GraphCPD: Coherent Point Drift for Point Cloud Registration Via Graph Signal Processing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a new probabilistic registration method based on graph signal processing (GSP), called graph coherent point drift (GraphCPD). |
Yingcheng Lai; Xingjian Wang; Li Chai; | code |
| 868 | DualTAP: A Dual-Task Adversarial Protector for Mobile MLLM Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing privacyperturbations fail the critical dual challenge of this scenario: protect-ing PII from the router’s MLLM while simultaneously preserving taskutility for the agent’s MLLM. To address this gap, we propose the Dual-Task Adversarial Protector (DualTAP), a novel framework that,for the first time, explicitly decouples these conflicting objectives. |
Fuyao Zhang; Jiaming Zhang; CHE WANG; Xiongtao Sun; Yurong Hao; Guowei Guan; Wenjie Li; Longtao Huang; Wei Lim; | code |
| 869 | TAQ: Static-Deployable Temporal-Aware Quantization for Real-World Video Super-Resolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Temporal-AwareQuantization (TAQ), a novel static post-training quantization frame-work that uses video structure only in offline calibration. TAQ calibratessequence-specific activation bounds, refines them with a temporal con-sistency objective we propose that aligns inter-frame changes betweenfloating-point and quantized outputs without weight retraining, and en-sembles the refined bounds into one deployable set of static param-eters. |
Jinwoo Chung; Sangho An; Sungyeop Jung; Jangho Kim; | code |
| 870 | Test-time Counterfactual Calibration for Hallucination-Resistant Temporal Grounding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This may stem from incorrect at-tribution of textual, visual, or multimodal information, leading them toconfidently generate plausible time frames for events that do not actuallyexist. To address this limitation, we propose HRVTG, a test-time adapta-tion framework that dynamically calibrates the decision boundary of themodel during inference. |
Chufan YI; Hongyu Qu; Shiyu Xuan; Rui Yan; Xiangbo Shu; Fang Zhao; Guosen Xie; | code |
| 871 | Urban Boundaries, Social Barriers: A Benchmark and Vision-Centric Framework for Mapping Gated Communities and Equity Implications Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We address this gap by introduc-ing GBA-GCs, a metropolitan-scale multimodal benchmark for locallygrounded gated/open community recognition in China’s Greater BayArea, covering 37,444 residential compounds with aligned boundary poly-gons, high-resolution satellite imagery, Chinese metadata, and structuredattributes, together with expert-verified labels, inter-annotator reliability,and official evaluation splits. Built on this benchmark, we present Multi-modal Classifier for Gated Community (MCGC), a vision-centricmultimodal framework based on DINOv3-SAT that fuses imagery, text,and structured cues via modality-aware cross-attention and adaptivegating to mitigate modality imbalance. |
Minwei Zhao; WEIMING ZHANG; Jiawang DU; Qiming LIU; Weiming Zhuang; Pei Nie; Cai Wu; | code |
| 872 | LineGraph2Road: Structural Graph Reasoning on Line Graphs for Road Network Extraction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose an end-to-end pipeline that integrates vision-based segmentation, sparse graph construction, and structured inference. |
Zhengyang Wei; Renzhi Jing; Yiyi He; Jenny Suckale; | code |
| 873 | Learn to Rank: Visual Attribution By Learning Importance Ranking Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a learning scheme that instead optimizes dele-tion and insertion metrics directly. |
David Schinagl; Christian Fruhwirth-Reisinger; Alexander Prutsch; Samuel Schulter; Horst Possegger; | code |
| 874 | Distribution-Aware Feature Selection for Post-hoc Out-of-Distribution Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We show thatOOD-discriminative information in deep feature representations is of-ten concentrated in subsets of features and is frequently axis-aligned. Toexploit this structure, we introduce a distribution-aware feature selec-tion strategy that ranks feature dimensions according to the discrepancybetween in-distribution (ID) and OOD feature distributions, using theWasserstein-1 distance as a principled metric.To avoid the need for curated OOD validation data, we construct proxy-OOD data based on cross-domain mixup and evaluate adversarial per-turbations as an alternative. |
Max Gutbrod; David Rauber; Christoph Palm; | code |
| 875 | LogicIR: Logic Gate Networks for Image Restoration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce LogicIR, the first LGN specifically designed for image restoration tasks. |
Hongjae Lee; Myungjun Son; Jaeseong Yu; Seung-Won Jung; | code |
| 876 | TaskTok: Delving Into Task Tokens for Task-driven Image Restoration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This suggests thatselectively refining a subset of tokens can be sufficient for task-drivenobjectives. Leveraging this insight, we propose TaskTok, a novel frame-work that selectively restores only task-relevant tokens via a learnabletoken switch and a lightweight token refinement module. |
Hongjae Lee; Sojung Kang; Jaeseong Yu; Seung-Won Jung; | code |
| 877 | 3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present a scene-adaptive human image animation framework that controls both human motion and camera trajectories within a reconstructed 3D environment for video generation. |
Deyin Liu; Jicheng Xu; Lin Yuanbo Wu; Xiaowei Zhao; Xiatian Zhu; Anjan Dutta; Zhe Jin; | code |
| 878 | TRINITY: A Multi-Perspective Benchmark for Personal-Style Video Highlight Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Traditional video highlight detection relies on a narrow, event-centric definition of saliency, which often fails to generalize to uncon-strained personal videos where highlights are heterogeneous and perspective-dependent. To address this, we introduce TRINITY, a multi-perspectivebenchmark that decomposes highlight saliency into three complemen-tary dimensions, Event, Emotion, and Nature, within a unified tempo-ral framework. |
Qianqian Chen; Hyun Bin Kim; Denzel Wijaya; Yang Yi; Bo LIU; Yangkai Ding; | code |
| 879 | Q-BridgeNet: A Quantization Network for Cross-Lingual Sign Language Translation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing multilingual SLT approaches still struggle to learn a unified model that minimizes crosslingual conflicts while capturing shared cross-lingual semantics and preserving language-specific variations across different sign languages. Therefore, we propose Q-BridgeNet, a unified framework for multilingual SLT that jointly mitigates cross-lingual conflicts across both the sign language and spoken language sides. |
Liqian Feng; Lintao Wang; Xiaochen Liu; Anusha Withana; Ken-Tye Yong; Dehui Kong; Zhiyong Wang; Kun Hu; | code |
| 880 | Real-Time Source-Free Object Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose DHF (Dual-Head Pseudo-Label Fusion) which selectively admits one-to-one (O2O) and one-to-many (O2M) head predictions, preserving precision and recovering missed objects. |
Sairam V C Rebbapragada; Varun Gopal; Poornima Jain; Vineeth N Balasubramanian; Muhammad Haris Khan; | code |
| 881 | Benchmarking MLLMs on Mistake Recognition and Explanation in Single-Step Components of Cooking Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This study proposes CookingMistake Recognition and Explanation (Cook-MRE), a dataset for assess-ing MLLMs in-depth performance in understanding cooking mistakeswithin each single-step. |
Shun Takashige; Atsushi Hashimoto; Shin’ichi Satoh; | code |
| 882 | HilDA: Hierarchical Distillation with Diffusion for Advancing Self-Supervised LiDAR Pre-training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose HilDA, a self-supervised pre-training framework for LiDAR backbones that better captures the semantic what and geometric where needed for driving tasks. |
Maciej Wozniak; Jesper Ericsson; Hariprasath Govindarajan; Truls Nyberg; Thomas Gustafsson; Patric Jensfelt; Olov Andersson; | code |
| 883 | Weight Feedback Computes The Exact Jacobian Transpose in Modern Deep Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Predictive Coding (PC) offers a biologically motivated al-ternative to backpropagation via local weight updates, yet routing er-ror between layers still relies on an autograd Jacobian-transpose (J⊤ )product—the last non-local operation in PC. We show that this depen-dency is largely avoidable. |
Junlong Shen; Xingyu Li; | code |
| 884 | XSurfer: Reconstructing Surface Meshes of Cerebral and Cerebellar Cortex from Diverse MRI Data Using Untrained Neural Networks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While CSR is a mature technology when applied routinelyto adult T1 -weighted brain MRI data of the cerebrum, it remains underexploredfor a broad range of MRI contrasts, resolutions, ages, species, and brain struc-tures such as the cerebellum. To address this challenge, we propose XSurfer,a contrast- and resolution-agnostic CSR framework that performs optimizationon single images using an untrained neural network such that training data arenot needed. |
Haoxiang Li; Mingxuan Liu; Divya Varadarajan; Zhangxuan Hu; Qiyuan Tian; Jonathan Polimeni; | code |
| 885 | Neuromorphic X-ray Computed Tomography Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However,the application of event cameras to CT reconstruction remains largelyunexplored. To bridge this gap, we introduce a Neuromorphic X-rayCT framework. |
Hongjian Wang; Goran Lovric; Benjamín Béjar; | code |
| 886 | Pixel-wise Geo-registration of Drone and Satellite Images Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce SkyReg, a geometry-awaregeo-registration model that estimates the transformation between thequery and reference images by explicitly modeling the 3D scene ge-ometry. |
Qingyang Liu; David G Shatwell; Parth Parag Kulkarni; Shah Mubarak; | code |
| 887 | Reference-Free Quality Assessment for Virtual Try-On Via Human Feedback Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As virtual try-on (VTON) systems become increasingly im-portant in fashion e-commerce, there is a growing need for reliable reference-free evaluation methods, since ground-truth images of the same personwearing the target garment are typically unavailable in real-world scenar-ios. To address this challenge, we propose VTON-IQA, a reference-freeframework for human-aligned image quality assessment without requir-ing ground-truth images. |
Yuki Hirakawa; Takashi Wada; Ryotaro Shimizu; Takuya Furusawa; Yuki Saito; Ryosuke Araki; Tianwei Chen; Fan Mo; Yoshimitsu Aoki; | code |
| 888 | Diffusion to Obfuscation: Time-Adaptive Synthesized Generation Against Gradient Leakage Attacks in Federated Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Revisiting key insights in GLAs under FL reveals that (1) Defenses should focus on protecting semantic and fine-grained details of data; (2) GLAs are effective mainly in early rounds; and (3) To avoid semantic leakage, defenses shouldn’t infer the true labels of private images for obfuscation. Building upon these insights, we present a simple, time-adaptive defense strategy that obfuscates the private gradient by employing the gradient from a synthesized image. |
Farchan Raswa; Chun-Shien Lu; Jia-Ching Wang; | code |
| 889 | MixTTA: Low-Rank Cross-Channel Mixing for Reliable Test-Time Adaptation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, per-channel affine parameters per-form axis-aligned scaling and shifting, making them geometrically inca-pable of correcting cross-channel structural changes induced by distribu-tion shift. To address this limitation, we propose MixTTA, a lightweightplug-in module that equips normalization layers with a low-rank cross-channel transformation, enabling inter-channel mixing at each layer. |
Mansoo Jung; Youngwook Kim; Jungwoo Lee; | code |
| 890 | HCSU: A Dataset and Benchmark for Fine-Grained Historical Calligraphy Style Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Automated fine-grained perception of calligraphy styles—a task vital to cultural heritage preservation—remains a critical challenge for Large Vision-Language Models (LVLMs), largely constrained by existing datasets that suffer from modal mixture and flattened labels. To bridge this gap, we introduce HCSU, the first comprehensive dataset tailored for fine-grained Historical Calligraphy Style Understanding. |
Yinsheng Yao; Yan Liu; Chen Ye; | code |
| 891 | WaterGen: Decoupling Scene and Medium in Underwater Image Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce WaterGen, a method for generatinglarge-scale, realistic, and diverse underwater images that provides inde-pendent control of the scene and water medium conditions. |
Jiayi Wu; Tianfu Wang; Tianyi Xiong; Dehao Yuan; Xiaomin Lin; Md Jahidul Islam; Cornelia Fermuller; Christopher Metzler; Yiannis Aloimonos; | code |
| 892 | Align and Segment: Unsupervised Learning for Building Segmentation From Misaligned Labels Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a novel approach for training on misaligned labels, where we simultaneously learn the label alignment. |
Venkanna Babu Guthula; Oswin Krause; Dimitri Gominski; Hui Zhang; Johan Mottelson; Ankit Kariryaa; Nico Lang; Christian Igel; | code |
| 893 | Training-Free Multi-Concept Image Editing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While recentoptimisation-based methods achieve strong zero-shot edits from text,they still struggle to preserve identity and capture intricate details, suchas facial structure, surface texture, or object-specific geometry, that ex-ist below the level of linguistic abstraction. To address this fundamentalgap, we propose Concept Distillation Sampling (CDS). |
Niki Maria Foteinopoulou; Ignas Budvytis; Stephan Liwicki; | code |
| 894 | InstantHDR: Single-forward Gaussian Splatting for High Dynamic Range 3D Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, we design a geometry-guided appearancemodeling for multi-exposure fusion, and a meta-network for generaliz-able scene-specific tone mapping.Due to the lack of HDR scene data,we build a pre-training dataset, called HDR-Pretrain, for generalizablefeed-forward HDR models, featuring 168 Blender-rendered scenes, di-verse lighting types, and multiple camera response functions. |
Dingqiang Ye; Jiacong Xu; Jianglu Ping; Yuxiang Guo; Chao Fan; Vishal Patel; | code |
| 895 | MmIR: Frequency-Space Inverse Rendering for 3D Millimeter-Wave Radar ADC Synthesis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present mmIR, an open-sourcedifferentiable frequency-modulated continuous-wave (FMCW) radar in-verse renderer that fits a physics-based forward model to real capturesand re-renders from dense virtual apertures to synthesize high-resolution3D radar data. |
Adnan Armouti; Yixuan Gao; Rajalakshmi Nandakumar; | code |
| 896 | PrintAnything: Learning Geometric Plan Map for 3D Printing G-code Generation from Unoriented Point Clouds Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A common workaround is to reconstruct meshesfrom point clouds; however, the resulting meshes often contain geomet-ric artifacts, such as incorrect faces or topological inconsistencies, thatare difficult to repair and may lead to printing failures. To overcomethese limitations, we propose PrintAnything, a novel framework thatlearns to produce executable 3D printing G-code directly from 3D pointclouds without requiring mesh reconstruction. |
Sangmin Hong; Daniel Sungho Jung; Heewon Kim; Kyoung Mu Lee; | code |
| 897 | Silhouette-based Gait Foundation Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Building a uni-fied gait foundation model requires addressing two longstanding barriers:(a) Scalability – Why have gait models historically failed to follow em-pirical scaling trends? |
Dingqiang Ye; Chao Fan; Kartik Narayan; Bingzhe Wu; Chengwen Luo; Jianqiang Li; Vishal Patel; | code |
| 898 | ReQuest: Rethinking-based Question-Aware Frame Selection for Long-Form Video QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose ReQuest, an uncertainty-driven, question-adaptive keyframe selection pipelinethat aligns question intent with relevant video content through selec-tive computation. |
Minkuk Kim; Suyong Yun; Young Kim; Jinyoung Moon; Jinwoo Choi; Seong Tae Kim; | code |
| 899 | EgoTraj: Real-World Egocentric Human Trajectory Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However,progress in this direction remains limited due to the scarcity of egocen-tric trajectory datasets collected in real-world environments. Address-ing this need, we introduce EgoTraj, an egocentric multimodal opendataset recorded using Meta Quest Pro (MQPro). |
Ahmad Yehia; Abduallah Mohamed; Tianyi Wang; Kun Qian; Jiseop Byeon; Junfeng Jiao; Christian Claudel; | code |
| 900 | FoundDP: Revisiting Weak Disparity Observability in Dual-Pixel Depth Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing DP-based methods rely primarily on local disparity cues and therefore become unreliable when disparity signals are weak or ambiguous. To address this limitation, we propose FoundDP, a uni(cid:28)ed framework that integrates metric DP depth with global structural priors from a monocular depth foundation model. |
fengchen he; Hao Xu; Dayang Zhao; Tingwei Quan; Shaoqun zeng; | code |
| 901 | LACON: Training Text-to-Image Model from Uncurated Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Is the discarded bad data truly useless, or does it hold un-tapped potential? In this work, we critically re-examine this question. |
Zhiyang Liang; Ziyu Wan; Hongyu Liu; DONG CHEN; Qiu Shen; Hao Zhu; Dongdong Chen; | code |
| 902 | PriSM: Parsing and Style-Mixed Consistency for Unsupervised Domain Adaptation in Facial Landmark Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although self-training is a widely used UDA strategy to bridge do-main gaps, it frequently breaks down under large domain shifts as it isprone to amplifying confident yet erroneous pseudo-label predictions. Tothis end, we propose PriSM (Parsing and Style-Mixed Consistency), anovel method for robust landmark pseudo-label validation. |
Chieh-Yu Yang; Hou-Ning Hu; Sykai Chen; Yu-Lun Liu; Yen-Yu Lin; | code |
| 903 | Context Blindness in DPO: Mitigating Object Hallucination in MLLMs Via Context-Calibrated Preference Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Di-rect Preference Optimization (DPO) mitigates this by training modelsto prefer non-hallucinated responses over hallucinated ones, and recentefforts further enrich the preference data with relevant context. How-ever, it remains unclear whether DPO actually leverages such context.To investigate this, we propose Contextual Preference Gain (CPG), asimple metric that measures how much a model’s preference strengthenswhen relevant context is provided. |
Byungoh Ko; Jinyoung Park; Jongha Kim; Jeehye Na; Jaewon Cho; Hyunwoo Kim; | code |
| 904 | UBone3D: Physics-Rectified Conditional Flow Matching for Anatomical 3D Shape Completion from Ultrasound Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we presentUBone3D, a novel framework based on physics-rectified conditional flowmatching (CFM) that performs point cloud completion directly from par-tial US observations. |
Weiying Chen; Yuchong Gao; Siyuan Li; Marek Reformat; Rui Zheng; Edmond Lou; | code |
| 905 | Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a retraining-free VLM pruning framework called PORTA that derives a taskand modality-agnostic importance formulation based on activation variation, estimated from generic calibration data, which reliably captures feature-level representation utility across modalities. |
Kang MinSeok; Hyunwoo Kim; Chanyoung Kim; Minwoo Kim; Jaekoo Lee; Dahuin Jung; | code |
| 906 | ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To support this framework, we construct two comple-mentary datasets and a benchmark: ReasonLite-42M, with open-form,visually verifiable reasoning captions; ReasonPro-16M, with category-specific reasoning supervision; and RCLIP-Bench for diagnostic evalua-tion of visually grounded reasoning. |
Sicheng Zhang; Muhammad Muzammal Naseer; Binzhu Xie; Naufal Suryanto; Shi Qiu; Jamal Bentahar; NAVEED AKHTAR; Shah Mubarak; | code |
| 907 | Towards Reliable Medical Large Vision-Language Models Via Counterfactual Preference Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To addressthis, we propose a Counterfactual Medical Preference Optimization(CoMedPO) framework that learns an unbiased policy from a biasedreference model. |
Xiaoguang Zhu; NaipengWang NaipengWang; Kartik Patwari; Lianlong Sun; Chen-Nee Chuah; ChengxinPang ChengxinPang; | code |
| 908 | SV-TAD: Native Sparse Convolutions for Efficient Temporal Action Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Tokenselection can reduce attention cost by pruning redundant tokens, butit breaks the spatial grid structure required by convolutional adapters.This forces an expensive dense reconstruction, nullifying much of thepotential speedup. We address this by introducing native sparse 2D con-volutions, a primitive that allows these adapters, for the first time, tooperate directly and efficiently on dynamically pruned token sets. |
Ricardo Ignacio Pizarro Carreño; Roberto Valle; José Buenaposada; Luis M. Bergasa; Luis Baumela; | code |
| 909 | CLIP-AUTT: Test-Time Personalization with Action Unit Prompting for Fine-Grained Video Emotion Recognition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, weleverage Action Units (AUs) as structured textual prompts within CLIPto model fine-grained facial expressions. |
Muhammad Osama Zeeshan; Masoumeh Sharafi; Benoît Savary; Alessandro Lameiras Koerich; Marco Pedersoli; Eric Granger; | code |
| 910 | AiSCREAM: Absolute Target Localization with Language-Conditioned Cross-View Alignment for Autonomous Vehicles Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this study, we focus on absolute target position prediction froma language instruction and front view image, which requires referring ex-pression disambiguation and absolute target localization without depthcues. |
Kei Katsumata; Jun Piao; Naoki Hosomi; Kentaro Yamada; Komei Sugiura; | code |
| 911 | Discovering Geometric Biases in 3D Face Reconstruction: A Curvature-Aware Spectral Framework for Fairness Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a novel framework to analyze 3DMM reconstructions through the lens of surface curvature, with the objective to discover, quantify and visualize biases. |
Veronika Shilova; Emmanuel Malherbe; Giovanni Palma; Panagiotis-Alexandros Bokaris; Laurent Risser; Jean-Michel Loubes; | code |
| 912 | SPLIT: Training-Free AI-Generated and Partially Edited Video Detection Via Spatial Patch‑Level Incoherence and Temporal Roughness Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Deploying AI-generated video detectors in real-world ser-vices demands an ultra-low false positive rate (FPR) on real videosto avoid falsely rejecting authentic content, a regime … |
Jongyeop Hyun; Hyounghun Kim; | code |
| 913 | WiFlow: Estimating Optical Flow Using WiFi Channel State Information Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we explore using WiFi channel state information (CSI) instead of camera frames for optical flow estimation.Further, we create the first dataset for training and evaluating CSI-based optical flow estimators, and our experiments provide insights into key design elements for this task. |
Thomas Weigel; Simon Kiefhaber; Fabian Portner; Matthias Hollick; Simone Schaub-Meyer; | code |
| 914 | D-VLAM: Differential Vision and Language Mixing for Rehearsal Free Continual Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: For further study and reproducibility, we also providerigorous analysis, details, and source code of our method in the sup-plementary document. |
Muhammad Anwar Ma'sum; Mohsen Guizani; Waseem Ullah; | code |
| 915 | Cast and Attached Shadow Detection Via Iterative Light and Geometry Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To support training and evaluation,we introduce a dataset of 1,458 images with manually annotated castand attached shadow masks sourced from three existing benchmarks.Experiments demonstrate that our proposed method outperforms priormethods, with at least a 33% reduction in attached-shadow BER, whilemaintaining strong full-shadow and cast-shadow performance. |
Shilin Hu; Jingyi Xu; Sagnik Das; Dimitris Samaras; Hieu Le; | code |
| 916 | Data-Free Client Contribution Estimation Via Logit Maximization for Federated Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a data-free, class-wise contribution estimation and aggregation framework based on logitmaximization (CELM) that does not require raw data, client metadata,or auxiliary public datasets at the server. |
Asim Ukaye; Nurbek Tastan; Mubarak Abdu-Aguye; Karthik Nandakumar; | code |
| 917 | SAF3R: Dynamic Sparse Attention for Feed-Forward 3D Reconstruction Transformers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present a comprehensive analysis of global attention across multiple F3R transformers and reveal that attention patterns are highly heterogeneous, dynamic, and extremely sparse across layers and attention heads. |
Jianing Deng; Yuanzhe LI; Jialu Wang; Song Wang; Tianlong Chen; Huanrui Yang; Jingtong Hu; | code |
| 918 | Parametric SDF for Dynamic Surface Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In thiswork, we introduce a new paradigm for dynamic surface reconstructionbased on a parametric Signed Distance Function (p-SDF). |
Chong Gao; Kai Ye; Qiyu Dai; Yiming Shao; Qiong Zeng; Ding Liang; Yanpei Cao; Guanbin Li; Wenzheng Chen; | code |
| 919 | CrossView: Can Vision-Language Models Reason Across Cameras? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce CrossView, a multi-camera video question-answeringbenchmark spanning autonomous driving, security surveillance, egocen-tric/exocentric video, and robotics. |
Sahil Shah; S P Sharan; Harsh Goel; Manvik Pasula; Adithya Hebbalae; Minkyu Choi; Sandeep Chinchali; | code |
| 920 | FeatTracker: Short- and Long-Range Temporal Feature Consistency for Robust Underwater Object Tracking Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite this, they still struggle with persistent featuredegradation and tracking drift caused by challenging underwater condi-tions. To overcome these limitations, we propose FeatTracker, a feature-level tracking framework that enforces both short- and long-range tem-poral feature consistency, effectively preserving semantic integrity andensuring temporal stability for robust underwater object tracking. |
Jiaqing Li; Bin Lin; Chaocan Xue; Wu Ai; Qingping Zheng; | code |
| 921 | Circuit-MLLM: Topological Logic-Guided Latent-Space Visual Reasoning for Circuit Schematic Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, circuit schematics present a unique challenge for MLLMs due to their dense component layouts and distinct topological logic, demanding fine-grained structural parsing to extract the electrical semantics. To address this, we propose Circuit-MLLM, a multimodal reasoning framework that reformulates circuit topology analysis as a process of device localization, path tracing, and sequential reasoning within the latent space. |
Jinyuan Deng; Yuqi Jiang; Wenjing Huang; Xin Li; Qi Sun; Cheng Zhuo; | code |
| 922 | Iterative Perceptual Alignment for VLMs Via Deterministic Reconstruction Feedback Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a self-supervisedframework that derives solid, highly discriminative supervisory signalsthrough deterministic reconstruction feedback. |
Xiaorui Chen; Hanzhong Guo; Nizhe Cai; Jieliang Luo; | code |
| 923 | Interact3D: Compositional 3D Generation of Interactive Objects Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A Vision-Language Model (VLM) au-tonomously analyzes multi-view renderings of the composed scene, for-mulates targeted corrective prompts, and guides an image editing mod-ule to iteratively self-correct the generation pipeline. |
Hui Shan; Keyang Luo; Ming Li; Sizhe Zheng; Yanwei Fu; Zhen Chen; Xiangru Huang; | code |
| 924 | DA-F2F: Domain-Adaptive Object Detection with Feature-to-Feature Modulation and Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While recent approaches rely on image-to-image transla-tion, they often suffer from instability inherent in pixel-level domaintransformation. To address these limitations, we propose DA-F2F, anovel framework that directly aligns domains within the feature space.DA-F2F introduces a style-aware feature modulation (SFM) module thatextracts style statistics from the target domain to dynamically modulatesource representations. |
HoTaek Oh; Hee-Jun Kim; Hyo-Jun Lee; | code |
| 925 | Towards Sparsely Annotated Open World Object Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Dual-Perspective ObjectDiscovery (DPOD), a unified framework that jointly models unlabeledknown and unknown instances via two complementary mechanisms. |
HEEJU HAN; AJEONG KIM; Jinsun Park; | code |
| 926 | DiffRGD: An Inference-Time Diffusion Guidance Through Riemannian Gradient Descent Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, thesemethods cannot effectively preserve the original Gaussian distributionbecause they introduce distributional drift, thereby degrading the sam-ple quality. To address this gap, we propose DiffRGD, a distribution-aware guidance framework that explicitly preserves the latent Gaussianstructure. |
Jia-Wei Liao; Li-Xuan Peng; Mei-Heng Yueh; Min Sun; Cheng-Fu Chou; Jun-Cheng Chen; | code |
| 927 | Geometry-Anchored Transport Framework for Exemplar-Free Class-Incremental Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we formulate feature transport as an endogenous training constraint rather than a separate post-task step, presenting the Geometry-Anchored Transport Framework. |
Hongye Xu; Bartosz Krawczyk; | code |
| 928 | DnA: Denoising Attention for Visual Tasks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The softmax activation in multihead attention (MHA) is thede facto standard for attention-based models in visual perception tasks.However, standard softmax can produce noisy attention patterns thatdilute relevant features and degrade its performance. In this paper, wepropose Denoising Attention or DnA, in which, first, a positive queryidentifies which image features belong to the correct class, and a neg-ative query identifies closely associated but irrelevant image features.DnA then projects these interactions into two distinct subspaces withlarger principal angles, promoting subspace separation and improved dis-criminability. |
Ron Campos; Subhajit Maity; Xin Li; Srijan Das; Aritra Dutta; | code |
| 929 | ReCamDriving: LiDAR-Free Camera-Controlled Video Synthesis for Novel Trajectories Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose ReCamDriv-ing, a purely vision-based framework that achieves camera-controlledgeneration by leveraging dense, structurally complete 3DGS renderingsas geometric guidance.Based on this strategy, we construct the ParaDrive dataset,containing approximately 110K parallel-trajectory video pairs. |
Yaokun li; Shuaixian Wang; Mantang GUO; Jiehui Huang; Taojun Ding; Mu Hu; Kaixuan Wang; Shaojie Shen; Guang Tan; | code |
| 930 | LinStereo: Linear-Complexity Global Attention for Multi-Scale Iterative Stereo Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Existing Vision Foundation Model (VFM)-based iterativestereo pipelines under-exploit three information pathways: multi-scalebackbone features are collapsed into single-level … |
Yiran Wang; Oliver Turner; Viorela Ila; | code |
| 931 | Two-Parameter Flow Map Learning for Continuous-Time Diffeomorphic Image Registration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a framework to directly learn the continuous-time solution of a non-autonomous ODE formulated as a two-parameterflow map. |
Mohammadjavad Matinkia; Nilanjan Ray; | code |
| 932 | Teaching Vision-Language-Action Models What to See and Where to Look Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose DriveTeach-VLA,a framework that explicitly teaches VLAs what to see and where tolook. |
Yuguang Yang; Canyu Chen; Zhewen Tan; Yizhi Wang; Zichao Feng; Chunyang Liu; Kehua Sheng; Bo Zhang; Yan Wang; Juan Zhang; Linlin Yang; Baochang Zhang; Xianbin Cao; | code |
| 933 | Look But Don’t Touch with Sparse Autoencoders for Unlearning in Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Sparse autoencoders (SAEs) have recently been proposed asinterpretable tools for concept-level manipulation, under the assumptionthat isolated features can serve as controllable intervention points. In thiswork, we systematically evaluate this assumption in the context of ob-ject erasure and steering in diffusion models. |
Enrico Cassano; Riccardo Renzulli; Rayyan Ahmed; Stephan Alaniz; Marco Grangetto; | code |
| 934 | Last-Layer-Centric Feature Recombination: Unleashing 3D Geometric Knowledge in DINOv3 for Monocular Depth Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this study, we present a systematic layer-wise analysis of DINOv3, revealing that 3D information is distributed nonuniformly: deeper layers exhibit stronger depth predictability and better capture inter-sample geometric variation. |
Gongshu Wang; Zhirui Wang; Kan Yang; | code |
| 935 | Single-Query Person-Centric Bimanual Hand-Object Interaction Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Experiments with a transformer-baseddetector show that our formulation improves person-level bi-manual in-teraction parsing and provides an effective unified framework for jointdetection, pose estimation, and hand reasoning.Project page: https://lgecto-ail-vil.github.io/SingleQuery-BHOI/ |
Jonghyun Kim; Junho Roh; Yubin Yoon; Jaechul Kim; Jungho Lee; Hyotae Lee; Jongkuk Park; Taehwan Hwang; | code |
| 936 | Spatial Amsan: A Benchmark for Perception-Grounded Spatial Reasoning and Action Evaluation in Egocentric Manipulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Spatial Amsan, a benchmark for evaluatingstate-based spatial reasoning and action evaluation in egocentric manip-ulation videos.We release all data, annotations, and codeat https://github.com/Blanchard-lab/SpatialAmsan. |
Changsoo Jung; Jack Fitzgerald; Ethan Seefried; Mariah Bradford; Nathaniel Blanchard; | code |
| 937 | SiPhy: Single-Image Physical Property Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce SiPhy, a unified framework for single-image physical property reasoning that aligns 3D-aware visual cues, depthwith language-based material knowledge. |
Hoang Le; Joonwoo Kwon; Elkhan Ismayilzada; Yufei Zhang; Zijun Cui; | code |
| 938 | CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis Via Canonical Workspace Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: How-ever, conditioning grasp synthesis on specific human grasp taxonomiestypically requires prohibitively expensive, object-annotated datasets. Toaddress these limitations, we propose CoToGrasp, a novel generativeframework that synthesizes diverse, stable grasps strictly conditioned onspecific contact topologies. |
Julien Mérand; Boris Meden; Liming Chen; Mathieu GROSSARD; | code |
| 939 | Disentangling Rotation and Translation from SE(3)-Equivariant Features for Shape Assembly Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: For 3Dpose prediction, prior works utilize the entangled features that capture3D pose information well, but this mixed representation impedes gen-eralization. To address this problem, we propose the SOT encoder thatdisentangles 3D pose into an SO(3)-equivariant (rotation) feature and aT(3)-equivariant (translation) feature. |
Hee-Jun Jung; Uigeun Ahn; JINHWI PARK; Kangil Kim; | code |
| 940 | Thinking from The Robot’s View: The CoT-HRC Benchmark for Human Intent Reasoning in Embodied Collaboration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a diagnostic protocol featur-ing the Step-wise Consistency Assessment (SCAR ) to penalize spuriousaccuracy by enforcing intermediate logical consistency, and the Condi-tional Reasoning Accuracy (CRA) to explicitly decouple models’ un-derlying reasoning capabilities from visual perception failures. |
Chenxi Deng; Chao Xu; JianmingLiu JianmingLiu; Shaofei Chen; | code |
| 941 | ReFP-AD: Rectified Flow Preconditioning for Energy-Based Anomaly Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Energy-BasedModels (EBMs) offer a principled formulation, but their training inhigh-dimensional token spaces is unstable due to anisotropy and strongcross-dimensional correlations, which degrades finite-step Markov ChainMonte Carlo (MCMC) sampling. We identify this instability as fundamen-tally geometric and introduce ReFP-AD (Rectified Flow Preconditioningfor Anomaly Detection), which learns a geometric reparameterization thatmaps high-dimensional embeddings into a well-conditioned latent spacevia an optimal transport (OT)-coupled rectified flow. |
Camile Lendering; Erkut Akdag; Joaquin Figueira Chacon; Egor Bondarev; | code |
| 942 | Per‑Object IoU Forecasting for Deadline‑Aware Real‑Time Embedded Detection Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a lightweight per-object IoU decay modelthat predicts accuracy loss and derives closed-form deadline estimatesfor each detected object. |
Erfan Foorginejad; Akshar Chavan; Marco Brocanelli; | code |
| 943 | Evidence Triangulation for Multimodal Fact-Checking in The Wild Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Addition-ally, we propose TRENT, a novel MFC model that performs evidence tri-angulation using three parallel cross-attention streams alongside a rela-tional fusion mechanism that explicitly models entailment and contradic-tion. |
Stefanos-Iordanis Papadopoulos; Zacharias Chrysidis; Christos Koutlis; Symeon Papadopoulos; Panagiotis Petrantonakis; | code |
| 944 | PACO: Stabilizing Vision Embeddings Along Local Paths for Robust Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing ad-versarial fine-tuning methods typically optimize CLIP under a singlefixed perturbation strength, resulting in weak robustness generalizationand semantic instability. To address this limitation, we propose Path-Consistent Fine-Tuning (PACO), a new framework for unsupervised ad-versarial fine-tuning. |
Qihang Tang; Jiacheng Pi; Zhiguo Yang; Xu Liu; Perley Xu; Wenjie Ruan; | code |
| 945 | Multi-label Instance-level Generalised Visual Grounding in Agriculture Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite progress in vision–language tasks likecaptioning and visual question answering, Visual Grounding (VG), local-ising language-referred objects, remains unexplored in agriculture. A keyreason is the lack of suitable benchmark datasets for evaluating groundingmodels in field conditions, where many plants look highly similar, appearat multiple scales, and the referred target may be absent from the image.To address these limitations, we introduce gRef-CW, the first dataset de-signed for generalised visual grounding in agriculture, including negativeexpressions. |
Mohammadreza Haghighat; Alzayat Saleh; Mostafa Azghadi; | code |
| 946 | VERITAS: A Multi-agent Co-scientist for Verifiable Image-Derived Hypothesis Testing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Veritas (Veri_x001C_able Epistemic Reasoningfor Image-Derived Hypothesis Testing via Agentic Systems), a clinicalco-scientist: a multi-agent system that autonomously tests natural-language hypotheses and produces a fully auditable evidence trail, trac-ing every conclusion through executable outputs from analysis plan tosegmentation masks to statistical code to _x001C_nal verdict.We construct atiered benchmark of 64 hypotheses spanning six complexity levels acrosscardiac and brain glioma MRI datasets. |
Lucas Stoffl; Benedikt Wiestler; Johannes Paetzold; | code |
| 947 | OPAL: Orthonormal Prototype Alignment Learning for Interpretable Image Classification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, currentapproaches are burdened by complex, multi-stage training pipelines andheavily rely on auxiliary regularization to prevent prototype collapse. Toovercome these limitations, we introduce Orthonormal Prototype Align-ment Learning (OPAL), a single-stage, end-to-end framework that sim-plifies interpretable classification. |
Ilán Carretero; Gustavo ANGULO; Rocío Amor; Valery Naranjo; | code |
| 948 | Concept-to-Pixel: Prompt-Free Universal Medical Image Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Universal medical image segmentation seeks to use a singlefoundational model to handle diverse tasks across multiple imaging modal-ities. |
Haoyun Chen; Fenghe Tang; Wenxin Ma; S Kevin Zhou; | code |
| 949 | Video-Oasis: Rethinking Evaluation of Video Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work,we introduce Video-Oasis, a sustainable diagnostic suite for systemati-cally auditing existing video understanding benchmarks. |
Geuntaek Lim; Minho Shim; Sungjune Park; Jaeyun Lee; Inwoong Lee; Taeoh Kim; Dongyoon Wee; Yukyung Choi; | code |
| 950 | PanoSAM2: Lightweight Distortion- and Memory-aware Adaptions of SAM2 for 360 Video Object Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose PanoSAM2, a novel 360VOS frame-work based on our lightweight distortion- and memory-aware adaptationstrategies of SAM2 to achieve reliable 360VOS while retaining SAM2’suser-friendly prompting design. |
Dingwen Xiao; WEIMING ZHANG; Shiqi Wen; Addison Wang; | code |
| 951 | Beyond Aesthetics: Quantifying Information Loss in Turbid Scenes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This issue is compounded by reliance on synthetic turbidity datasets, which may misrepresent real-world information loss. To address this gap, we introduce the Turbid Underwater Baseline (TUB) dataset, comprising 1,320 images captured under extreme turbidity and over 16,000 high-confidence ground-truth segmentation masks. |
Vasiliki Ismiroglou; Tasos Benos; Malte Pedersen; Stefan Bengtson; Thomas B. Moeslund; | code |
| 952 | Adversarial Attack and Disturbance Detection By Hadamard-Coded Output Representations for Object Detection and Semantic Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Fourth, we introduce HadamardNet, a framework employing Hadamard codes as output representations for semantic segmentation and object detection models and tasks. |
Lucas Görnhardt; Timo Bartels; Niklas Schwarz; Tim Fingscheidt; | code |
| 953 | MoBa-GS: Learning A Spatially-Varying Motion Basis Over A Dynamic Canonical Space for 4D Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent state-of-the-art methods rely on monolithic deformation networks or global time-basis factorizations that struggle to represent complex, non-rigid topological changes. We propose MoBa-GS, a framework that resolves this entanglement by introducing a structural inversion: a spatially-factorized motion field coupled with adaptive geometric optimization. |
Guan Yuan Tan; Arghya Pal; Sailaja Rajanala; Raphaël Phan; Chee-Ming Ting; | code |
| 954 | Text Dictates, Music Decorates: Energy-based Attention for Editable Dance Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Conversely, attempting to condition generation on both text and music frequently leads to modality collapse, where dense acoustic rhythms overwhelm sparse semantic text prompts, destroying user controllability. To resolve this spatial-temporal conflict, we propose STREAM (Structural-Temporal Rhythmic Energy-based Attention for Motion), a modality-decoupled diffusion transformer. |
Seong Jong Yoo; Siyuan Peng; Felix Gu; Stratis Aloimonos; Cornelia Fermuller; | code |
| 955 | ECC: Encoder-Centric Corruption for Fine-Grained Vision in VLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We identify the cause as a structural limitation: recent trends toward decoder-style designs prevent noise-induced signals from being deeply integrated into the encoder representations. To address this issue, we propose Encoder-Centric Corruption (ECC), a framework guided by three principles: (i) Encoder-Centricity, performing restoration within the encoder to shape transferable features; (ii) FeatureLevel Noising, injecting noise at intermediate layers to capture finegrained textures; and (iii) Task Disentanglement, separating mask reconstruction and denoising via a disruption loss. |
Hyesong Choi; Daeun Kim; Sungmin Cha; Kwang Moo Yi; Dongbo Min; | code |
| 956 | Experts-Guided Unbalanced Optimal Transport for ISP Learning from Unpaired And/or Paired Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This reliance on costly-to-acquire paired data remains a significant bottleneck. To address this challenge, we introduce a novel, unsupervised training framework based on Optimal Transport capable of training arbitrary ISP architectures in both unpaired and paired modes. |
Georgy Perevozchikov; Nancy Mehta; Egor Ershov; Radu Timofte; | code |
| 957 | Parsimonious Flow Matching for Efficient Image Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose ParsimoniousFlow Matching (PFM), which adopts a mixture of Gaussians (MoG)as the latent distribution whose geometry better aligns with the data.We identify key design choices that enable efficient FM, including anoptimal-transport data-latent coupling, MoG estimation via k-means,and eigenvalue regularization of the per-mode covariances. |
Tianjiao Ding; Ziqing Xu; Benjamin Haeffele; Hongkang Li; Rene Vidal; | code |
| 958 | AFFMAE: Scalable Vision Pre-Training for High-Resolution Microscopy Segmentation on Desktop Hardware Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: AFFMAE removes dense-grid assumptions while preserving hierarchical scalability during pre-training and fine-tuning. To support this architecture, we developed nu-merically stable mixed-precision Triton kernels and a lightweight, point-based decoder that can be directly repurposed as a segmentation head.On high-resolution microscopy segmentation, AFFMAE matches MAEfinetuning performance on foot process width estimation with ViT back-bone at equal parameter counts while being 2x faster during pre-trainingand halving peak memory usage. |
David Smerkous; Zian Wang; Behzad Najafian; | code |
| 959 | Video2Reaction: Mapping Video to Audience Reaction Distribution in The Wild Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To enable audience reaction prediction andother content engagement applications, we introduce Video2Reaction,a multimodal dataset that maps short movie segments to a distributionof induced emotions of viewers in the wild, as expressed through socialmedia. |
Trang Nguyen; Sidong Zhang; Shiv Shankar; Gauri Jagatap; Deepak Chandran; Andrea Fanelli; Madalina Fiterau; | code |
| 960 | FaCT-GS: Fast and Scalable CT Reconstruction with Gaussian Splatting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper addresses the most significant remaining limita-tions of the GS-based approach by introducing FaCT-GS, a frameworkfor fast and flexible CT reconstruction. |
Pawel Pieta; Rasmus Juul Pedersen; Sina Borgi; Jakob S. Jørgensen; Jens Wenzel Andreasen; Vedrana Dahl; | code |
| 961 | REALM: An RGB and Event Aligned Latent Manifold for Cross-Modal Perception Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing learningbased approaches for event processing are typically confined to narrow, task-specific silos and lack the ability to generalize across modalities. We address this gap with REALM, a cross-modal framework that learns an RGBand Event-Aligned Latent Manifold by projecting event representations into the pretrained latent space of RGB foundation models. |
Vincenzo Polizzi; David Lindell; Jonathan Kelly; | code |
| 962 | Image Warping for Image-to-Image Translation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The warped image is processed by theoriginal diffusion model and then mapped back via an inverse warp.In addition, we propose a simple and efficient outpainting-based syn-thetic data generation pipeline to produce high-quality paired data forimage relighting. |
Shen Zheng; Anurag Ghosh; Gaurav Parmar; Srinivasa G. Narasimhan; | code |
| 963 | Do Vision Language Models Recognize Visual Ambiguity? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study these approaches and reveals that the resulting variability is often dominated by textual changes rather than visual evidence, causing uncertainty estimates to reflect prompt sensitivity rather than visual ambiguity. We therefore propose Visual Semantic Entropy (VSE), which perturbs only the image to probe nearby visual variations while keeping the text query fixed. |
Huy Ta; Trang Nguyen; Townim Chowdhury; ANKIT YADAV; Minh-Son To; Zhibin Liao; Johan Verjans; Hieu Phan; | code |
| 964 | CrossFeat: Bridging Imaging Modalities in Feature Descriptor Space Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Instead, we propose CrossFeat,a framework that enables an existing monomodal descriptor to operateacross modalities. |
Paul Schneider; Nazim Haouchine; | code |
| 965 | SemCityLoc: Aerial 6DoF Localization Using Semantic 3D City Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose SemCityLoc, a semantic–geometricalignment system that reframes aerial pose estimation as structuredsurface registration between foundation-model-derived visual priors andstandardized LoD-compliant 3D city models. |
Jingfeng Mao; Xuyang Chen; Qilin Zhang; Oussema Dhaouadi; Guangming Wang; Brian Sheil; Daniel Cremers; Yan Xia; Olaf Wysocki; | code |
| 966 | Caption Bottleneck Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Furthermore, tradi-tional CBMs often suffer from information leakage, where unmodeledvisual features bypass the bottleneck and compromise the integrity ofthe explanations. To overcome these limitations, we propose CaptionBottleneck Models (CaBM), a framework that circumvents the need forpredefined concept sets by replacing rigid concept layers with free-formnatural language. |
Seref Baris Cagliyan; Umut Ozdemir; Merve Tapli; Emre Akbas; | code |
| 967 | TanGO: Training-Free 3D Editing Via Tangent-Space Guidance and Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While recent flow-matching 3D generative models (e.g., Vec-Set) adopt structured representations, their tokens share global context,causing conventional training-free editing to suffer from semantic arti-facts such as collapsed preserved regions or incomplete transformations.To address this, we propose TanGO, a training-free framework thatenables adaptive per-token steering in the tangent space of generativedynamics. To realize this selective control, we formulate a one-step opti-mal control rule and determine the strength of each token’s control signalusing a von Mises-Fisher inspired directional discrepancy derived fromthe source and target velocity fields. |
Siwoo Lim; Sunjae Yoon; Gwanhyeong Koo; Hyeonseo Yun; Chang D. Yoo; | code |
| 968 | PhenoLeaf-TS: A Time-Series Benchmark for Leaf Instance Segmentation, Tracking, and Growth Stage Classification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce PhenoLeaf-TS, a time-series dataset of 17,082top-down RGB images spanning 21 Arabidopsis thaliana genotypes, to-talling 318 plant replicates, each annotated with colour-coded leaf in-stance masks that maintain consistent identity throughout the growthsequence. |
Rijad Saric; Basim Azam; Sarmad Khan; Edhem Custovic; | code |
| 969 | Teaching An Agent to Sketch One Part at A Time Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We develop a method for producing vector sketches one partat a time. |
Xiaodan Du; Ruize Xu; David Yunis; Yael Vinker; Greg Shakhnarovich; | code |
| 970 | Zero-Shot Image Personalization from Personas Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce zero-shot imagepersonalization from personas (ZIPP), a paradigm that conditions im-age generation on natural-language personas (concise descriptors of auser’s identity, interests, and aesthetic sensibilities) without any user-specific data or weight updates. |
Harini S I; Somesh Singh; Yaman K Singla; David Doermann; Rajiv Shah; | code |
| 971 | MLP Splatting: Object-Centric Neural Fields Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Com-pared to state-of-the-art semantic 3DGS methods, we achieve substan-tially lower memory usage (1/7×) and faster rendering (5×). |
Shinjeong Kim; Yuzhou Cheng; Xin Kong; Paul Kelly; Andrew Davison; | code |
| 972 | Token-level Response-visual Attention Guidance for Multimodal LLMs Knowledge Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, wechallenge the conventional reliance on prompt-to-vision attention by re-vealing that downstream performance correlates strongly with response-to-vision attention similarity to the teacher, but negligibly with that ofprompt-conditioned attention. |
Jaehyun Jang; Eunseop Yoon; Hee Suk Yoon; SooHwan Eom; Mark Hasegawa-Johnson; Chang D. Yoo; | code |
| 973 | SegPAR: Class-Centric Decision-Based Sparse Attack for Semantic Segmentation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To bridge this gap, we adapt the most representative decision-basedblack-box sparse attacks from the classification domain to serve as base-lines, establishing a rigorous benchmark for this underexplored setting.In this context, we demonstrate that one of the existing methods suf-fers from severe query inefficiency due to its image-centric pixel accu-mulation, which rapidly exhausts query budgets across the vast imagespace. To overcome this, we propose SegPAR, a novel decision-basedframework that shifts to a class-centric exploration paradigm. |
Dongsu Song; Daeyun Go; BOSEUNG SEO; Jay Hoon Jung; | code |
| 974 | MV-GEL: Language-Driven Multi-View Geometric Entity Localization on Meshes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The 3D object appearance is highly sensitive to viewpoints; a single perspective may render a target entity clearly observable, while another may suffer from severe occlusion or foreshortening. In this work, we attempt to solve these challenges with MV-GEL (Multi-View Geometric Entity Localization), a framework for localizing fine-grained geometric entities on polygon meshes from natural language queries. |
Kartik Bali; Roland Aydin; | code |
| 975 | CountEx: Fine-Grained Counting Via Exemplars and Exclusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents CountEx, a discriminative visual count-ing framework designed to address a key limitation of existing prompt-based methods: the inability to explicitly exclude visually similar distrac-tors. |
Yifeng Huang; Gia Khanh Nguyen; Minh Hoai Nguyen; | code |
| 976 | FFAvatar: Feed-Forward 4D Head Avatar Reconstruction from Arbitrary Images Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present FFAvatar, a Transformer-based 3D Gaussian framework for fast construction of high-quality and animatable 4D head avatars from one or more reference portrait images. |
Jianjiang Yao; Ke Xian; Renxiang Dai; Robert Qiu; | code |
| 977 | TripVVT: A Large-Scale Triplet Dataset and A Coarse-Mask Baseline for In-the-Wild Video Virtual Try-On Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Built upon this resource, we develop TripVVT, a Diffusion Transformer–based framework that replaces fragile garment masks with a simple, stable human-mask prior, enabling reliable background preservation while remaining robust to real-world motion, occlusion, and cluttered scenes.We publicly release the dataset and benchmark at https://huggingface.co/datasets/TripVVT/TripVVT-10K, which we believe provide a solid foundation for advancing controllable, realistic, and temporally stable video virtual try-on. |
Dingbao Shao; Song Wu; Shenyi Wang; Ye Wang; Ziheng Tang; Fei Liu; Jiang Lin; Xinyu Chen; Qian Wang; Ying Tai; Jian Yang; Zili Yi; | code |
| 978 | When Token Compression Breaks: Structural Pruning Vs. Token Reduction for Robust ViT Segmentation Under High Compression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In contrast, structural prun-ing exhibits a smoother degradation curve and is more stable at highcompression. Motivated by these findings, we study a prune-then-mergepipeline that applies moderate token compression on top of a moder-ately pruned backbone. |
Tien-Phat Nguyen; Ngai-Man Cheung; | code |
| 979 | One Demonstration Is Enough for Real-World Robotic Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Prior work has incorporated demonstrations intoreinforcement learning (RL), yet existing approaches either require largenumbers of demonstrations or depend on continuous human interventionduring training. To address these limitations, we present AutoSERL, aframework that leverages a single demonstration to fully automate the in-tervention process in real-world robot RL. |
Yuwan Liu; Hongze Yu; song liu; Yuhan Wang; Junge Zhang; Yaodong Yang; Yuanpei Chen; Ceyao Zhang; | code |
| 980 | Fisheye3R: Adapting Unified 3D Feed-Forward Foundation Models to Fisheye Lenses Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While onemay surmise that training on fisheye images would resolve this problem,there are far fewer fisheye images with ground truth than perspectiveimages, which limits generalization. To enable inference on imagery ex-hibiting high radial distortion, we propose Fisheye3R, a novel adaptationframework that extends these multi-view 3D reconstruction foundationmodels to natively accommodate fisheye inputs without performance re-gression on perspective images. |
Ruxiao Duan; Erin Hong; Dongxu Zhao; Eric Turner; Alex Wong; Yunwen Zhou; | code |
| 981 | TreeSRNF: Square-Root Normal Fields for Generative Modelling of The Geometric and Structural Variability in Tree-like 3D Objects Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a novel mathematical framework for analyz-ing and generating complex tree-shaped 3D objects, such as botanicaltrees and plants, which deform both in their 3D geometry and branchingstructure. |
Tahmina Khanam; Hamid Laga; Mohammed Bennamoun; Guanjin Wang; Ferdous Sohel; Farid Boussaid; Anuj Srivastava; | code |
| 982 | 4DGS360: 360° Gaussian Reconstruction of Dynamic Objects from A Single Video Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce 4DGS360, a di!usion-free framework for 360→dynamic object reconstruction from casual monocular video. |
Jae Won Jang; Yeonjin Chang; Wonsik Shin; Juhwan Cho; Nojun Kwak; | code |
| 983 | The Label Imitation Game: Turing Test Network for Zero-Shot Pseudo-Label Pruning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Foundation model pseudo-labeling—labeling data strictly viazero-shot inference—enables massive scale, but performance is under-mined by hallucinations that evade standard thresholds. To eliminatethese errors, we introduce the Turing-inspired Label Imitation Game(LIG), a framework that formalizes pseudo-label pruning as an adversar-ial interrogation. |
Brent Griffin; Jason Corso; | code |
| 984 | PDF-Omni: Poincaré Dual Disk Distortion Field-based Recurrent Update for Omnidirectional Stereo Matching Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing methods rely onspatially uniform update policies, degrading performance in these chal-lenging regions. To address this limitation, we propose PDF-Omni, ageometry-aware OSM framework built on a Poincaré Dual Disk distor-tion field that models spatially varying geometric uncertainty across seamregions. |
Yunseok Yang; Eunjin Son; Sang Lee; | code |
| 985 | SpatiO: Adaptive Test-Time Orchestration of Vision-Language Agents for Spatial Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce SpatiO, a heterogeneous multi-agent framework for spatial reasoning that coordinates multiple vision-language specialists with complementary inductive biases. |
ChanYeong Hwang; Miso Choi; Sunghyun On; Jinkyu Kim; Jungbeom Lee; | code |
| 986 | SyntheticDoc: A Large Synthetic Dataset for Document Unwarping and Illumination Correction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address thisbottleneck, we introduce SyntheticDoc, a massive, high-quality datasetdesigned to push the boundaries of document unwarping. |
Daniel Woortmann; Tanguy Magne; Olga Sorkine-Hornung; | code |
| 987 | Multi-view Multi-vehicle Driving Dataset for Novel View Synthesis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce the Multi-View Multi-Vehicle (MV2 ) datasetand benchmark for evaluating NVS models under large viewpoint changesin dynamic urban scenes. |
Sanjay Dharavath; Hanvitha Mukkamala; Faizan Khan; Ioannis Kakogeorgiou; Aditya Arun; Zakaria Laskar; C. V. Jawahar; | code |
| 988 | NumColor: Precise Numeric Color Control in Text-to-Image Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To train the ColorBook, we construct NumColor-Data, a synthetic dataset of 500K rendered images with unambiguous color-to-pixel correspondence, eliminating the annotation ambiguity inherent in photo-graphic datasets. |
Muhammad Atif Butt; Diego Hernández; Alex Gomez-Villa; Kai WANG; Javier Vazquez-Corral; Joost Van de Weijer; | code |
| 989 | E-MOTION: A Dataset for Event-Based Scene Flow Estimation with Independent Moving Objects Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Progress may be limited by the unconventional data addingcomplexity to established processing pipelines, but also due to the lackof event camera datasets with scene flow ground truth. With the aimof filling this gap, we present E-MOTION: a large and versatile datasetrecorded with high-resolution event cameras suitable for depth, opticalflow and scene flow estimation. |
Ivan Gutierrez Rodriguez; Julien Moreau; Chiara Bartolozzi; Arren Glover; | code |
| 990 | CoCo: Code As CoT for Text-to-Image Preview and Rare Concept Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In thiswork, we propose CoCo (Code-as-CoT), a code-driven reasoning frame-work that represents the reasoning process as executable code, enablingexplicit and verifiable intermediate planning for image generation. |
Li Haodong; Chunmei Qing; Huanyu Zhang; Dongzhi Jiang; Yihang Zou; Hongbo Peng; Dingming Li; Yuhong Dai; ZePeng Lin; Juanxi Tian; Yi Zhou; Siqi Dai; | code |
| 991 | A Benchmark for Heterogeneous Stereo Deblurring with Physically- and Epipolar-constrained Cross Attention Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing methods andbenchmarks largely assume homogeneous stereo setups and therefore donot explicitly address such asymmetric degradation. To bridge this gap,we present a dedicated framework for heterogeneous stereo deblurring.First, we introduce the heterogeneous stereo deblurring (HSD) dataset,constructed from real smartphone stereo captures via multi-frame inte-gration. |
Jiah Kim; Hoju Shin; Seung-Wook Kim; Seowon Ji; | code |
| 992 | Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we explore a complementary failure-driven self-improvement loop, a data-centric paradigm that turns failed trajecto-ries into agent improvements. |
Xueqiao Sun; Yuhui Zhang; Xiaohan Wang; Ludwig Schmidt; Serena Yeung-Levy; | code |
| 993 | Learning to Generate Rigid Body Interactions with Video Diffusion Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despitestrong advances, current approaches still struggle to generate physicallyplausible object interactions and lack object-level control mechanisms. Toaddress these limitations, we introduce KineMask, an approach for videogeneration that enables realistic rigid body control, interactions, and ef-fects. |
David Romero; Ariana Bermudez; Viacheslav Iablochnikov; Hao Li; Fabio Pizzati; Ivan Laptev; | code |
| 994 | RBE-Flow:Recurrent Bayesian Estimation on Feature Manifolds for Cross-Modal Registration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, wepropose RBE-Flow, a novel framework that reformulates dense cross-modal flow estimation as a closed-loop recurrent Bayesian estimationproblem on learned feature manifolds. |
Mengzhu Ding; Xin Song; Xiaoke Ding; Hongwei Ding; Xuecong Liu; | code |
| 995 | A Mechanism-Driven Theory of Phase Transitions in Active Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Overall, this work provides a unified framework for the nextgeneration of transition-aware AL algorithms. |
Julia Machnio; Mads Nielsen; MOSTAFA MEHDIPOUR GHAZI; | code |
| 996 | Isotropic Embedding Perturbations for Robust Vision Language Encoders Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Aether, a simple plug-in method thatapplies diffusion-style random perturbations in the embedding space viacontrolled alpha-mixing, specifically designed to provide isotropic regu-larization that remains semantically consistent. |
Hyesong Choi; Daeun Kim; Song Park; Taekyung Kim; Byeongho Heo; Sangdoo Yun; Dongbo Min; Dongyoon Han; | code |
| 997 | Reflecting Process Expertise in Procedural Material Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Concretely, we represent expertworkflows as process traces: textual records of construction steps, param-eters, and design intent. |
Kunal Gupta; Gaurav Joshi; Yen-Ru Chen; Seemandhar Jain; Ishit Mehta; Manmohan Chandraker; | code |
| 998 | Test-Time Noise Guided Adaptation for Realistic Autoregressive Video Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Autoregressive video diffusion models have enabled the gener-ation of arbitrarily long videos by removing conditioning on future frames,thus greatly improving computational efficiency. |
Dimitrios Karageorgiou; Symeon Papadopoulos; Ioannis Kompatsiaris; Efstratios Gavves; | code |
| 999 | LASER: A Corrective Lens for LVLMs Via Visual Attention Preservation and Sink Suppression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Through systematic empirical analysis, we find thatperformance degradation under visual forgetting is largely driven bytwo overlooked factors: early-stage attention decay disrupts evidenceacquisition, and attention concentration on a subset of task-irrelevantvisual sink tokens. Motivated by these insights, we propose LASER, apost-training framework that regulates both the visual attention tra-jectory and intra-visual token attention distribution during reasoning.Technically, LASER introduces two complementary rewards: a VisualGrounding Reward, which encourages the model to maintain attentionon semantically salient visual tokens throughout decoding, and a SinkSuppression Reward, which penalizes excessive attention concentration onvisual sink tokens. |
Bowen Yuan; Zijian Wang; Yadan Luo; Shijie Wang; Zi Helen Huang; | code |
| 1000 | Beyond Prompts: Unconditional 3D Inversion for Out-of-Distribution Shapes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We identify a critical failure mode where generation trajectoriesare drawn into latent “sink traps”: regions where the model becomesinsensitive to prompt modifications. |
Victoria Chen; Emery Pierson; Léopold Maillard; Maks Ovsjanikov; | code |
| 1001 | CLUE-VAD: Structured Semantic Clues for Understanding Explainable Events in Video Anomaly Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper,we present CLUE-VAD, a structured semantic decomposition framework for ex-plainable weakly supervised video anomaly detection. |
MYOUNG-CHUL KIM; Junhee Lee; ChaeBeen Bang; MyeongAh Cho; | code |
| 1002 | EVEE: Event-Based Online Adaptation for Matching on Unknown Targets Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose EVEE, a feature adaptation frameworkthat leverages temporally accumulated event evidence as test-time proxysupervision for adapting detection and matching to previously unseentargets. |
Zejing Zhao; Cheng Ju; Yanwen Zhang; Akio NAMIKI; | code |
| 1003 | BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: BaFCo curates 200 multi-page complex Bangladeshigovernment forms, sourced from across diverse sectors including agricul-ture, education, banking, and land management. To accurately capturethe structural and contextual complexity of these forms, we define a fine-grained annotation schema comprising 26 types of form entities, alongwith a separate coarse form entity set consisting of 5 types. |
Abu Tyeb Azad; Ishita Apan; Fahim Ahmed; Sumaiya Karim Katha; Ezharuddin Jubaer; Armun Alam; Pranjal Nandi; Amin Ali; Aman Chadha; Md Mofijul Islam; A K M Mahbubur Rahman; | code |
| 1004 | Explainability-aware Frustum Attack: Exposing Structural Vulnerabilities in LiDAR-Based 3D Object Detectors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Prior work has studied adversarialrobustness primarily on isolated 3D object models, while recent LiDARspoofing attacks target richer and more realistic driving scenes but focusmainly on physical realizability rather than understanding detector be-havior or attack efficiency. In this work, we investigate how LiDAR-baseddetectors rely on spatial evidence in complex scenes and whether these re-liance patterns can be exploited to induce failures more efficiently. |
Chengzeng You; Binbin Xu; Soteris Demetriou; | code |
| 1005 | RESOLVE: A Multi-Resolution and Multi-Modal Dataset for Roadside Cooperative Perception Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce RESOLVE, a large-scale real-world bench-mark dataset featuring multi-resolution roadside LiDAR and synchro-nized camera-LiDAR sensing for systematic evaluation of unimodal andfusion-based architectures in roadside 3D detection and tracking. |
Shaozu Ding; Linan Song; Marco Vincenzi; Dajiang Suo; | code |
| 1006 | Not All Prediction Targets Keep Training-Free Diffusion Guidance on The Manifold Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Training-free guidance (TFG) steers a pretrained diffusionmodel toward a desired attribute at inference. To be effective, this guid-ance must be applied from the earliest, … |
Yunsung Lee; Hyeongmin Lee; | code |
| 1007 | PRISM-VO: Scale-Aware Visual Odometry Using Photometric Plenoptic Bundle Adjustment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce PRISM-VO, a novel pure optimization-basedsparse photometric visual odometry framework for focused plenopticcameras. |
Aymeric Fleith; Julian Zirbel; Daniel Cremers; Niclas Zeller; | code |
| 1008 | Sound-based Multi-Person 3D Pose Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Furthermore, the complexity is compounded by inter-person reflections, which introduce intricate propagation delays that ob-scure the temporal motion-acoustic relationship. To address these issues,we propose SoundMHPE (Sound-based Multi-person Human Pose Es-timator), a novel encoder-decoder framework consisting of two key com-ponents. |
Yusuke Oumi; Yuto Shibata; Go Irie; Akisato Kimura; Yoshimitsu Aoki; Mariko Isogawa; | code |
| 1009 | SVI360: Spherical Video Interpolation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing video interpo-lation methods are not well-suited for spherical videos, as they havedifficulty handling severe distortions close to the poles. To address thisissue, we propose SVI360, a dual-branch framework that combines theimage frame and its rotated orthogonal view to deal with these dis-tortions. |
Le-Kim NGUYEN; Renato Martins; Pascal Vasseur; Cedric Demonceaux; | code |
| 1010 | PRISM3D: Probabilistic Refinement and Robust Initialization for Physically Consistent Scene Modeling Under Extreme Motion Blur Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce PRISM3D, a unified frameworkenabling robust reconstruction directly from severely degraded inputs.To overcome the lack of a reliable starting point, we propose a RobustInitialization strategy utilizing deep dense tracking method (VGGSfM)to recover global topology where feature matching fails. |
Gopi Raju Matta; Reddypalli Trisha; Divya Madhuri Vemunuri; Kaushik Mitra; | code |
| 1011 | Robustness Meets Uncertainty: Evidential Adversarial Training for Robust Selective Classification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: It standardizes architectures, augmentations, threat models, and evalu-ation metrics across clean, adversarial, and common-corruption settings.Across a wide range of state-of-the-art adversarial training methods,we uncover a recurring failure mode: several approaches improve robustaccuracy while degrading uncertainty ranking, leading to poorer selec-tive behavior. To address this, we propose Evidential Adversarial Train-ing (EV-AT), which models uncertainty through a Dirichlet distributionand combines (i) an evidence-based loss promoting clean accuracy andreliable uncertainty with (ii) a robust evidence-alignment loss match-ing clean and adversarial predictions in log Dirichlet-parameter space.Extensive experiments show that EV-AT shifts the Pareto frontier ofrobustness–uncertainty trade-offs beyond prior state-of-the-art adversar-ial training methods. |
Nicolas Sournac; Ahmed Baha Ben Jmaa; Bertrand Braeckeveldt; | code |
| 1012 | Think While You Map: Asynchronous Vision-Language Agents for Incremental 3D Scene Graphs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Open-vocabulary 3D scene graph methods typically operatein two stages: first reconstruct, then enrich with vision-language mod-els, leaving the graph unqueryable during exploration. We argue thatthis sequential coupling is unnecessary and propose an asynchronous ar-chitecture in which lightweight online mapping runs concurrently withheavyweight semantic refinement. |
Deniz Bickici; Michael Pabst; Shohei Mori; Dieter Schmalstieg; | code |
| 1013 | A Dual-Transformer Architecture with Cross-Attention for Multi-Camera View Recommendation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We evaluate on YouTube videos and on multi-view capture datawith significant occlusion and demonstrate state-of-the-art reconstruc-tion quality. |
Josep Cabacas Maso; Carles Ventura; Ismael Benito-Altamirano; | code |
| 1014 | Towards In-the-wild Egocentric 3D Hand-Object Pose Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Estimating accurate 3D hand–object pose from in-the-wildegocentric RGB remains challenging due to severe occlusions and am-biguous contact. Existing learning-based methods often … |
Siddhant Bansal; Zhifan Zhu; Shashank Tripathi; Jiahe Zhao; Michael Black; Dima Damen; | code |
| 1015 | COLA: Continual Orthogonal Low-Rank Adaptation for Class-Incremental Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address theselimitations, we propose Continual Orthogonal Low-Rank Adaptation(COLA), a novel rehearsal-free, parameter-efficient framework for CL.COLA integrates LoRA’s low-rank adaptation with an Oja-inspired learn-ing rule that incrementally approximates the dominant eigenstructure ofthe feature covariance across tasks. |
Monu Nagar; Debasis Das; | code |
| 1016 | CUST : Clustered Unit-level Similarity Transformer for Lightweight Image Super-Resolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To alleviate this, various window-based attention mechanisms have been proposed; yet, they inherently compromise the long-range dependency modeling that is the primary advantage of ViTs. To overcome these limitations, we propose the Clustered Unitlevel Similarity Transformer (CUST), a novel architecture that efficiently integrates global and local information. |
Jeongsoo Kim; | code |
| 1017 | RoboTALES: Learning Reasoning-Guided Robot Policies Via Task-Aligned Simulated Futures Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As a result, these mod-els can be difficult to use for planning or policy extraction. To addressthese limitations, we propose RoboTALES, a single-stage framework thatlearns task-aligned simulated futures and uses them to train robot poli-cies. |
Hanan Gani; Tejal Kulkarni; Madhoolika Chodavarapu; Nicklas Hansen; Manmohan Chandraker; | code |
| 1018 | URHead: A Unified UV-Space Representation for Joint Mesh–3DGS Optimization in Head Avatars Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present URHead, a unified representation for high-fidelityand animatable head avatars that fundamentally redefines mesh-Gaussianintegration. |
Seonghak Lee; Junhee Cho; Jisoo Park; Min-Gyu Park; Jongmin Lee; Ju Yoon; Junseok Kwon; | code |
| 1019 | General Self-Calibration with Varying Intrinsics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Unlike priorself-calibration methods that focus on specific varying-intrinsic regimes(e.g. shared focal length, fixed aspect, or zero skew), we provide a unifiedalgebraic framework that handles arbitrary intrinsic priors expressed aspolynomial constraints in the space of intrinsics, yielding a more flex-ible formulation of varying-intrinsic self-calibration. |
Norio Kosaka; Timothy Duff; Tomas Pajdla; Akihiro Sugimoto; | code |
| 1020 | FlashBEV: Fast and Memory-Efficient Exact BEV Transformation with IO-Awareness Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Unlike splatting-based VT, which requires index sorting and prevents fully thread-local reduction from voxel construction to BEV output, the gather(cid:21)reduction structure of Sampling-VT enables thread-local accumulation with on-the-(cid:29)y recomputation, eliminating the need to materialize heightand camera-dependent intermediates. Based on this insight, we propose FlashBEV, a fully fused and IOaware execution strategy that is mathematically equivalent to Tensorized Sampling-VT (same operator output) while substantially reducing global memory tra(cid:30)c and kernel-launch overhead. |
Shunsuke Yokokawa; Hironori Kasahara; | code |
| 1021 | MessyKitchens: Contact-rich Object-level 3D Scene Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work we advance object-level scene re-construction along two directions. |
Junaid Ahmed Ansari; Ran Ding; Fabio Pizzati; Ivan Laptev; | code |
| 1022 | NGPS: Structure-Preserving Self-Supervised Denoising Via Neighbor-Guided Patch Sampling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Neighbor-GuidedPatch Sampling (NGPS), a lightweight framework that constructs neigh-boring supervision under local inter-slice misalignment. |
Jaehyun Cho; YOUNGJOON YOO; | code |
| 1023 | Fourier Self-Supervision for Fine-Grained Generalized Category Discovery Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Fourier Self-Supervision, that leverages the Fouriertransform of images to enhance the discrimination of subtle differencesand support the discovery of new categories. |
Sarah Rastegar; Mina Ghadimi Atigh; Pascal Mettes; Yuki Asano; Cees Snoek; | code |
| 1024 | BiCE-HG: A Bi-Conditional Egocentric Hand Gesture Dataset for Intelligent Reality Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Walking trials capture realistic motion through a 4 mlighting corridor with transitions between bright and dim zones. |
Awfa Dakheel; Charith Abhayaratne; | code |
| 1025 | PixVOD: Pixel-Distributed Direct Visual Odometry and Depth Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To maintain geometric stability during op-timization, we introduce a keyframe-like anchoring mechanism that reg-ulates the effective baseline between frames, enabling consistent mo-tion and depth updates. |
Shinjeong Kim; Ignacio Alzugaray; Callum Rhodes; Paul Kelly; Andrew Davison; | code |
| 1026 | Reconstructing Dense Depth of Dark Scenes with Sparse LiDAR, Noisy Events, and Blurry RGB Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, under long-exposure imaging, motion blur in low-light RGB frames significantlydegrades the accuracy of depth reconstruction. To address this issue,we exploit the high-temporal-resolution motion cues captured by eventcameras and propose Event-guided Restoration and Upsampling Network(ERU-Net), a unified framework that tightly couples event-guided featurerestoration with depth completion. |
Jianbo Cao; Yuqi Han; Siming Zheng; Bo Wang; Tong Guo; Jinli Suo; | code |
| 1027 | SPECSIA: Stylization Dataset for Novel-View Enhancement in Drawing-based 3D Animation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing drawingbased 3D animation pipelines often use sample-wise 2D refinement to align animated renderings with the input image, but such optimization tends to overfit to the observed view and fails to correct projectioninduced artifacts in novel views. To address this limitation, we introduce SPECSIA-15K, a paired stylization dataset containing 14,980 artifactcorrupted projection/refinement-target pairs from 1,498 3DBiCar characters. |
Kyuwon Kim; Sunjae Yoon; Chang D. Yoo; | code |
| 1028 | AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce AnchorPrune, a training-freeframework that first constructs a protected relevance anchor and thenexpands it with complementary visual context. |
Kyuan Oh; Bumsoo Kim; | code |
| 1029 | Geometry-Aware Style Transfer in 3D Gaussian Splatting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present a novel geometry-aware style trans-fer framework for 3D Gaussian splatting (3DGS) that simultaneouslytransfers appearance attributes and geometric structures. |
Min Hyeok Bang; Jun Hyeong Kim; Seung-Wook Kim; Se-Ho Lee; | code |
| 1030 | SIMPLER: Efficient Foundation Model Adaptation Via Similarity-Guided Layer Pruning for Earth Observation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce SIM-PLER, a pre–fine-tuning architecture selection method that reduces in-ference and deployment costs by identifying an effective model depthbefore adaptation. |
Víctor Barreiro; Johannes Jakubik; Francisco Argüello; Dora Heras; | code |
| 1031 | Test Time Training for Long Videos Via Frame Forgetting Network Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We demonstrate FFN’s empirical effectiveness ondense-segmentation, video classification tasks, generalization to depth-estimation, and multi-hour long videos. |
Rajat modi; Xin Liang; Sebastian Noel; Yogesh Rawat; | code |
| 1032 | 3D Field of Junctions: A Noise-Robust, Training-Free Structural Prior for Volumetric Inverse Problems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by the strong 2D image denoising propertiesof Field of Junctions (ICCV 2021), we propose a novel, fully volumetric3D Field of Junctions (3D FoJ) representation that optimizes a junctionof 3D wedges that best explain each 3D patch of a full volume, while en-couraging consistency between overlapping patches. |
Namhoon Kim; NARGES MOEINI; Justin Romberg; Sara Fridovich-Keil; | code |
| 1033 | Exclusivity-Guided Mask Learning for Semi-Supervised Crowd Instance Segmentation and Counting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we first pro-pose an Exclusion-Constrained Dual-Prompt SAM (EDP-SAM), basedon our Nearest Neighbor Exclusion Circle (NNEC) constraint, to gener-ate mask supervision for current datasets. |
Jiyang Huang; Hongru Chen; Wei Lin; Jia Wan; Antoni Chan; | code |
| 1034 | KISS-GS: 3D Gaussian Splatting Compression Kept Simple Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To make the gains more transparent, we propose KISS-GS, a modular compression pipeline named after the principle of keeping things simple, designed to decouple compression entirely from training. |
Wieland Morgenstern; Friedrich Elias Branschke; Florian Fleischmann; Adrian Szatmari; Paul Schlack; Florian Barthel; Anna Hilsmann; Peter Eisert; | code |
| 1035 | Improving Adversarial Robustness Via Activation Amplification and Attenuation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end,we introduce Activation Amplification and Attenuation (A3), alightweight plug-in module that enhances adversarial robustness withminimal modifications of the activations. |
Taïga Gonçalves; Yongsong Huang; Tomo Miyazaki; Shinichiro Omachi; | code |
| 1036 | Towards Robust Driving Perception: A Flexible Scale-Driven Family for Self-Supervised Monocular Depth Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Networks specifically designed to handle dynamic traffic participants tend to be overly complex, hindering their deployment on resource-constrained automotive edge devices. To address these limitations and move towards robust driving perception, we propose FlexDepth, a scale-driven and flexible family of self-supervised MDE models tailored for challenging road scenarios. |
Zhaowen Zhu; Li Zhang; Chen Yujie; Zhang Tian; Yingjie Wang; Mingxia Zhan; | code |
| 1037 | Source-Agnostic Image Translation Based on Latent Aware Adaptive Masking Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a source-agnostic framework that dynamically refines a binary mask throughout the reverse diffusion process by computing the discrepancies of a pretrained diffusion model’s prediction for each latent time step. |
Tomislav Dobrički; Byung-Woo Hong; | code |
| 1038 | ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose ReactVAU, a Slow-Fast Decoupled Framework for real-time streaming Video Anomaly Understanding (VAU). |
Chia-Hui Chen; Shih-Ying Yeh; Fu-En Yang; Min-Hung Chen; Shang-Hong Lai; | code |
| 1039 | Unbalanced Optimal Transport for Efficient Visual Document Retrieval Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a visual token compression framework formulated as an Unbalanced Optimal Transport (UOT) problem. |
Hoyeon Shin; Jeongyeon Kim; Yeong Jun Koh; Yeoneung Kim; Hanul Kim; | code |
| 1040 | High-Resolution Artwork Outpainting with Global Blueprint Guidance and Layout Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite this progress, existing approaches still suffer from three key limitations: (i) the absence of a reliable global planning mechanism, which leads to structural instability and error accumulation at high resolutions; (ii) limited spatial controllability beyond text prompts, making it difficult to place objects at user-specified locations; and (iii) high inference latency caused by inherently sequential patch generation. To address these issues, we propose a global blueprint-guided two-stage diffusion framework for layout-controllable high-resolution outpainting with efficient parallel synthesis. |
Junha Kim; Hyunjoon Park; Donghyeon Cho; | code |
| 1041 | Denoised Variance-Based Pruning with Optimal Brain Bias Compensation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recently, Variance-Based Pruning (VBP) introduced a promising paradigm by selecting neurons based on activation variance; however, it remains limited by statistical noise in finite-sample activation covariance and reliance on bias-only updates that cannot fully account for structural reconstruction error. To address these limitations, we introduce Denoised Variance-Based Pruning with Optimal Brain Bias Compensation (DVBP + OB2C). |
Geon Tack Lee; Choo Jaegul; Kang Eun Jeon; | code |
| 1042 | Back-Tracking from Clarity: Self-Learning to See Text from Afar Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a self-supervised framework designed to enhance the capability of scene text detectors in identifying and recognizing text in scenarios where instances are shown at significant distances, typically small, blurred, and frequently missed by conventional models. |
Duc Tri Tran; Phi Le Nguyen; Minh Hoai Nguyen; | code |
| 1043 | DRIFT: Difficulty-aware Rectified Flows for Through-plane MRI Super-Resolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose DRIFT, a two-stage thickness-conditioned rectified flow framework for throughplane MRI super-resolution with continuous input slice-thickness. |
Yoonseok Choi; Eun-Gyu Ha; Daniel Kim; Mohammed Al-masni; Ming-Hsuan Yang; Dong-Hyun Kim; | code |
| 1044 | SAFER-Activities: A Dataset for Smart Assessment of Fall Events and Routine Activities Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this, we introduce SAFER-Activities, a dataset for fall detection and physical activity monitoring, with a dedicated subset for wheelchair use scenarios.To support research on robust fall detection and activity monitoring, we release the dataset and code at https://safer-activities.github.io/. |
Diwas Lamsal; Pramod Wickramatilake; Jednipat Moonrinta; Mongkol Ekpanyapong; Matthew Dailey; | code |
| 1045 | Unsupervised Source-Free Ranking of Biomedical Segmentation Models Under Distribution Shift Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce the first blackbox-compatible framework for unsupervised and source-free ranking of semantic and instance segmentation models based on the consistency of predictions under perturbations. |
Joshua Talks; Kevin Marchesini; Luca Lumetti; Federico Bolelli; Anna Kreshuk; | code |
| 1046 | Seeing Isn’t Orienting: A Cognitively Grounded Hierarchical Benchmark for Object Orientation in MLLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Discriminative Orientation Reasoning Intelligence (DORI), a cognition-informed hierarchical benchmark that establishes object orientation as the primary evaluation target. |
Nazia Tasnim; Keanu Nichols; Yuting Yan; Nicholas Ikechukwu; Elva Zou; Deepti Ghadiyaram; Bryan Plummer; | code |
| 1047 | Poppy: Polarization-Based Plug-and-Play Guidance for Enhancing Surface Normal Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Poppy, a training-free framework that refines normals from any frozen RGB backbone using single-shot polarization measurements at test time. |
Irene Kim; Sai Tanmay Reddy Chakkera; Alexandros Graikos; Dimitris Samaras; Akshat Dave; | code |
| 1048 | GeoV2V: Geometry-Grounded Video Diffusion Model for Driving Scene Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present GeoV2V, a geometry-grounded video-to-video diffusion framework that synthesizes temporally coherent driving videos along novel trajectories by integrating LiDAR and depth-derived geometric priors with full-video conditioning in the Wan 2.1 backbone. |
Yuchen Xi; tippy guo; Chenwei Hou; Zixu Liu; Jin Fang; Jason Liu; Ruigang Yang; | code |