PAPER DIGEST
Most Influential KDD 2004 Paper · 2026-03 edition

Automatic Multimedia Cross-modal Correlation Discovery

Jia-Yu Pan; Hyung-Jeong Yang; Christos Faloutsos; Pinar Duygulu

Venue
ACM SIGKDD Conference (KDD) 2004
Recognition
Most Influential KDD 2004 Paper (Rank No. 10)
Edition
2026-03
Impact factor
7
Certificate ID
407cf2fb4b412837

Abstract

Given an image (or video clip, or audio song), how do we automatically assign keywords to it? The general problem is to find correlations across the media in a collection of multimedia objects like video clips, with colors, and/or motion, and/or audio, and/or text scripts. We propose a novel, graph-based approach, "MMG", to discover such cross-modal correlations.Our "MMG" method requires no tuning, no clustering, no user-determined constants; it can be applied to <i>any</i> multimedia collection, as long as we have a similarity function for each medium; and it scales linearly with the database size. We report auto-captioning experiments on the "standard" Corel image database of 680 MB, where it outperforms domain specific, fine-tuned methods by up to 10 percentage points in captioning accuracy (50% relative improvement).

Download PDF certificate