Loading…

Learning object models from semistructured Web documents

This paper presents an automated approach to learning object models by means of useful object data extracted from data-intensive semistructured Web documents such as product descriptions. Modeling intensive data on the Web involves the following three phrases: first, we identify the object region co...

Full description

Saved in:
Bibliographic Details
Published in:IEEE transactions on knowledge and data engineering 2006-03, Vol.18 (3), p.334-349
Main Authors: Ye, S., Chua, T.-S.
Format: Article
Language:English
Subjects:
Citations: Items that this one cites
Items that cite this one
Online Access:Get full text
Tags: Add Tag
No Tags, Be the first to tag this record!
Description
Summary:This paper presents an automated approach to learning object models by means of useful object data extracted from data-intensive semistructured Web documents such as product descriptions. Modeling intensive data on the Web involves the following three phrases: first, we identify the object region covering the descriptions of object data when irrelevant contents from the Web documents are excluded. Second, we partition the contents of different object data appearing in the object region and construct object data using hierarchical XML outputs. Third, we induce the abstract object model from the analogous object data. This model would match the corresponding object data from a Web site more precisely and comprehensively than the existing handcrafted ontologies. The main contribution of this study is in developing a fully automated approach to extract object data and object model from semistructured Web documents using kernel-based matching and view syntax interpretation. Our system, OnModer, can automatically construct object data and induce object models from complicated Web documents, such as the technical descriptions of personal computers and digital cameras downloaded from manufacturers' and vendors' sites. A comparison with the available hand-crafted ontologies and tests on an open corpus demonstrate that our framework is effective in extracting meaningful and comprehensive models.
ISSN:1041-4347
1558-2191
DOI:10.1109/TKDE.2006.47