[an error occurred while processing this directive] An error occured whilst processing this directive

Theory Seminar


Monadic Datalog and the Expressive Power of Languages for Web Information Extraction

Georg Gottlob

Computer Science Department
Vienna University of Technology,

4pm Thursday 31st January 2002
Room 2511, JCMB, King's Buildings


Abstract

Research on information extraction from Web pages (wrapping) has seen much activity in recent times (particularly systems implementations), but not much work has been done on formally studying the expressiveness of the formalisms proposed or on the theoretical foundations of wrapping.

In this talk, we first introduce monadic datalog as a wrapping language over ranked or unranked tree structures. Using previous work by Neven and Schwentick, we show that this simple language is equivalent to full monadic second order logic (MSO) in its ability to specify wrappers. We believe that MSO has the right expressiveness required for Web information extraction and thus propose MSO as a yardstick for evaluating and comparing wrappers.

Using the above result, we study the kernel fragment Elog- of the Elog wrapping language used in the Lixto system (a visual wrapper generator). The striking fact here is that Elog- exactly captures MSO, yet is easier to use. Indeed, programs in this language can be entirely visually specified. We also formally compare Elog to other wrapping languages proposed in the literature.

Joint work with Christoph Koch

Martin Grohe
Monday 18 June 2001
An error occured whilst processing this directive