PAPER DIGEST
Most Influential ACM MULTIMEDIA 2024 Paper · 2026-03 edition

DiffMM: Multi-Modal Diffusion Model for Recommendation

Yangqin Jiang, Lianghao Xia, Wei Wei, Da Luo, Kangyi Lin, Chao Huang

Venue
ACM International Conference on Multimedia (ACM MULTIMEDIA) 2024
Recognition
Most Influential ACM MULTIMEDIA 2024 Paper (Rank No. 8)
Edition
2026-03
Impact factor
3
Certificate ID
4fd2a7087eff9cb1

Abstract

The rise of online multi-modal sharing platforms like TikTok and YouTube has enabled personalized recommender systems to incorporate multiple modalities (such as visual, textual, and acoustic) into user representations. However, addressing the challenge of data sparsity in these systems remains a key issue. To address this limitation, recent research has introduced self-supervised learning techniques to enhance recommender systems. However, these methods often rely on simplistic random augmentation or intuitive cross-view information, which can introduce irrelevant noise and fail to accurately align the multi-modal context with user-item interaction modeling. To fill this research gap, we propose a novel multi-modal graph diffusion model for recommendation called DiffMM. The proposed framework integrates a modality-aware graph diffusion model with a cross-modal contrastive learning paradigm to improve modality-aware user representation learning, better aligning multi-modal feature information with collaborative relation modeling. Our approach leverages diffusion models' generative capabilities to automatically generate a user-item graph that is aware of different modalities, enabling the incorporation of useful multi-modal knowledge in modeling user-item interactions. We conduct extensive experiments on three public datasets, demonstrating the superiority of our DiffMM over various competitive baselines.

Download PDF certificate