PAPER DIGEST
Most Influential SIGMOD 1999 Paper · 2026-03 edition

Record-boundary Discovery In Web Documents

D. W. Embley; Y. Jiang; Y.-K. Ng

Venue
ACM SIGMOD Conference (SIGMOD) 1999
Recognition
Most Influential SIGMOD 1999 Paper (Rank No. 11)
Edition
2026-03
Impact factor
7
Certificate ID
ec191cfa255fa897

Abstract

Extraction of information from unstructured or semistructured Web documents often requires a recognition and delimitation of records. (By “record” we mean a group of information relevant to some entity.) Without first chunking documents that contain multiple records according to record boundaries, extraction of record information will not likely succeed. In this paper we describe a heuristic approach to discovering record boundaries in Web documents. In our approach, we capture the structure of a document as a tree of nested HTML tags, locate the subtree containing the records of interest, identify candidate separator tags within the subtree using five independent heuristics, and select a consensus separator tag based on a combined heuristic. Our approach is fast (runs linearly for practical cases within the context of the larger data-extraction problem) and accurate (100% in the experiments we conducted).

Download PDF certificate