iBet uBet web content aggregator. Adding the entire web to your favor.

Overview

Authors:

George Papadakis ⁰,
Ekaterini Ioannou ¹,
Emanouil Thanos ²,
…
Themis Palpanas ³

George Papadakis
1. National and Kapodistrian University of Athens, Greece
View author publications

You can also search for this author in PubMed Google Scholar
Ekaterini Ioannou
1. Tilburg University, Netherlands
View author publications

You can also search for this author in PubMed Google Scholar
Emanouil Thanos
1. Katholieke Universiteit Leuven, Belgium
View author publications

You can also search for this author in PubMed Google Scholar
Themis Palpanas
1. University of Paris, France
  French University Institute (IUF), France
View author publications

You can also search for this author in PubMed Google Scholar

Part of the book series: Synthesis Lectures on Data Management (SLDM)

1471 Accesses
30 Citations

This is a preview of subscription content, log in via an institution to check access.

Access this book

Subscribe and save

Springer+ Basic

$34.99 /Month

Get 10 units per month
Download Article/Chapter or eBook
1 Unit = 1 Article or 1 Chapter
Cancel anytime

Buy Now

eBook USD 15.99 ~~USD 44.99~~

Discount applied Price excludes VAT (USA)

Softcover Book USD 15.99 ~~USD 59.99~~

Discount applied Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Other ways to access

Licence this eBook for your library

Institutional subscriptions

About this book

Entity Resolution (ER) lies at the core of data integration and cleaning and, thus, a bulk of the research examines ways for improving its effectiveness and time efficiency. The initial ER methods primarily target Veracity in the context of structured (relational) data that are described by a schema of well-known quality and meaning. To achieve high effectiveness, they leverage schema, expert, and/or external knowledge. Part of these methods are extended to address Volume, processing large datasets through multi-core or massive parallelization approaches, such as the MapReduce paradigm. However, these early schema-based approaches are inapplicable to Web Data, which abound in voluminous, noisy, semi-structured, and highly heterogeneous information. To address the additional challenge of Variety, recent works on ER adopt a novel, loosely schema-aware functionality that emphasizes scalability and robustness to noise. Another line of present research focuses on the additional challenge ofVelocity, aiming to process data collections of a continuously increasing volume. The latest works, though, take advantage of the significant breakthroughs in Deep Learning and Crowdsourcing, incorporating external knowledge to enhance the existing words to a significant extent. This synthesis lecture organizes ER methods into four generations based on the challenges posed by these four Vs. For each generation, we outline the corresponding ER workflow, discuss the state-of-the-art methods per workflow step, and present current research directions. The discussion of these methods takes into account a historical perspective, explaining the evolution of the methods over time along with their similarities and differences. The lecture also discusses the available ER tools and benchmark datasets that allow expert as well as novice users to make use of the available solutions.

The Five Generations of Entity Resolution on Web Data

A Survey on Blocking Technology of Entity Resolution

Article 27 July 2020

Entity Resolution in Big Data Era: Challenges and Applications

Table of contents (9 chapters)

Front Matter

Pages i-viii

Download chapter PDF
Entity Resolution: Past, Present, and Yet-to-Come
- George Papadakis, Ekaterini Ioannou, Emanouil Thanos, Themis Palpanas
Pages 1-3
Preliminaries
- George Papadakis, Ekaterini Ioannou, Emanouil Thanos, Themis Palpanas
Pages 5-13
Generation 1: Addressing Veracity
- George Papadakis, Ekaterini Ioannou, Emanouil Thanos, Themis Palpanas
Pages 15-48
Generation 2: Also Addressing Volume
- George Papadakis, Ekaterini Ioannou, Emanouil Thanos, Themis Palpanas
Pages 49-56
Generation 3: Also Addressing Variety
- George Papadakis, Ekaterini Ioannou, Emanouil Thanos, Themis Palpanas
Pages 57-81
Generation 4: Also Addressing Velocity
- George Papadakis, Ekaterini Ioannou, Emanouil Thanos, Themis Palpanas
Pages 83-96
Leveraging External Knowledge
- George Papadakis, Ekaterini Ioannou, Emanouil Thanos, Themis Palpanas
Pages 97-109
Resources for Entity Resolution
- George Papadakis, Ekaterini Ioannou, Emanouil Thanos, Themis Palpanas
Pages 111-118
Possible Directions for Future Work
- George Papadakis, Ekaterini Ioannou, Emanouil Thanos, Themis Palpanas
Pages 119-120
Back Matter

Pages 121-152

Download chapter PDF

Authors and Affiliations

National and Kapodistrian University of Athens, Greece

George Papadakis
Tilburg University, Netherlands

Ekaterini Ioannou
Katholieke Universiteit Leuven, Belgium

Emanouil Thanos
University of Paris, France

Themis Palpanas
French University Institute (IUF), France

Themis Palpanas

About the authors

George Papadakis is a research fellow at the National and Kapodistrian University of Athens, Greece. He also worked at the NCSR ""Demokritos,"" National Technical University of Athens (NTUA), L3S Research Center, and “Athena” Research Center. He holds a Ph.D. in Computer Science from the University of Hanover and a Diploma in Electrical Computer Engineering from NTUA. His research interest focuses on web data mining.Ekaterini Ioannou is an Assistant Professor at Tilburg University, the Netherlands. Prior, she worked as an Assistant Professor at Eindhoven University of Technology, as a Lecturer at the Open University of Cyprus, an adjunct faculty at EPFL in Switzerland, a research collaborator at the Technical University of Crete, and as an Independent Expert for the European Commission. Her research focuses on information integration with an emphasis on the challenges of man aging data with uncertainties, heterogeneity or correlations, and, more recently, on achieving a deeper integration of information extraction tasks within databases, and on efficiently retrieving analytics over graphs/hypergraphs with evolving data.
Emanouil Thanos is a Ph.D. candidate at CODeS research group of KU Leuven, under the supervision of Prof. Greet Vanden Berghe. He holds a Diploma in Electrical and Computer Engineering from the National Technical University of Athens and a joint Master in Com putational Logic from TU Dresden, FU Bolzano, and UN Lisbon. He has also worked as a research associate at National ICT Australia and the University of Athens. His research inter ests focus on combinatorial optimization and operations research.
Themis Palpanas is Senior Member of the French University Institute (IUF), and Professor of Computer Science at the University of Paris (France) where he is director of the Data Intel ligence Institute of Paris (diiP), and of the Data Intensive and Knowledge Oriented Systems (diNo) group. He is the author of two French patents andnine U.S. patents, three of which have been implemented in world-leading commercial data management products. He is the recipient of three Best Paper awards and the IBM Shared University Research (SUR) Award. He is currently serving in the Board of Trustees for the Very Large Data Bases (VLDB) Endowment, as Editor in Chief for BDR Journal, Editorial Advisory Board member for IS Journal, and in the Senior Program Committee of SIGMOD 2021.

Bibliographic Information

Book Title: The Four Generations of Entity Resolution
Authors: George Papadakis, Ekaterini Ioannou, Emanouil Thanos, Themis Palpanas
Series Title: Synthesis Lectures on Data Management
DOI: https://doi.org/10.1007/978-3-031-01878-7
Publisher: Springer Cham
eBook Packages: Synthesis Collection of Technology (R0), eBColl Synthesis Collection 10
Copyright Information: Springer Nature Switzerland AG 2021
Softcover ISBN: 978-3-031-00750-7Published: 16 March 2021
eBook ISBN: 978-3-031-01878-7Published: 01 June 2022
Series ISSN: 2153-5418
Series E-ISSN: 2153-5426
Edition Number: 1
Number of Pages: XVII, 152
Topics: Information Systems and Communication Service, Data Structures and Information Theory

Publish with us

Policies and ethics

The Four Generations of Entity Resolution

Overview

Access this book

Subscribe and save

Buy Now

Other ways to access

About this book

Similar content being viewed by others

The Five Generations of Entity Resolution on Web Data

A Survey on Blocking Technology of Entity Resolution

Entity Resolution in Big Data Era: Challenges and Applications

Table of contents (9 chapters)

Front Matter

Entity Resolution: Past, Present, and Yet-to-Come

Preliminaries

Generation 1: Addressing Veracity

Generation 2: Also Addressing Volume

Generation 3: Also Addressing Variety

Generation 4: Also Addressing Velocity

Leveraging External Knowledge

Resources for Entity Resolution

Possible Directions for Future Work

Back Matter

Authors and Affiliations

National and Kapodistrian University of Athens, Greece

Tilburg University, Netherlands

Katholieke Universiteit Leuven, Belgium

University of Paris, France

French University Institute (IUF), France

About the authors

Bibliographic Information

Publish with us

Navigation

The Four Generations of Entity Resolution

Overview

Access this book

Subscribe and save

Buy Now

Other ways to access

About this book

Similar content being viewed by others

Table of contents (9 chapters)

Front Matter

Back Matter

Authors and Affiliations

National and Kapodistrian University of Athens, Greece

Tilburg University, Netherlands

Katholieke Universiteit Leuven, Belgium

University of Paris, France

French University Institute (IUF), France

About the authors

Bibliographic Information

Publish with us

Search

Navigation