Magellan: Toward Building Ecosystems of Entity Matching Solutions.

Entity matching (EM) finds data instances that refer to the same real-world entity. In 2015, we started the Magellan project at UW-Madison, jointly with industrial partners, to build EM systems. Most current EM systems are standalone monoliths. In contrast, Magellan borrows ideas from the field of d...

Full description

Bibliographic Details
Published in:Communications of the ACM Vol. 63; no. 8; pp. 83 - 92
Main Authors: Doan, AnHai, Konda, Pradap, G. C., Paul Suganthan, Govind, Yash
Format: Article
Published: Association for Computing Machinery Aug2020
Subjects:
Online Access:View this record in EBSCOhost
Description
Summary:Entity matching (EM) finds data instances that refer to the same real-world entity. In 2015, we started the Magellan project at UW-Madison, jointly with industrial partners, to build EM systems. Most current EM systems are standalone monoliths. In contrast, Magellan borrows ideas from the field of data science (DS), to build a new kind of EM systems, which is ecosystems of interoperable tools for multiple execution environments, such as on-premise, cloud, and mobile. This paper describes Magellan, focusing on the system aspects. We argue why EM can be viewed as a special class of DS problems and thus can benefit from system building ideas in DS. We discuss how these ideas have been adapted to build PyMatcher and CloudMatcher, sophisticated on-premise tools for power users and self-service cloud tools for lay users. These tools exploit techniques from the fields of machine learning, big data scaling, efficient user interaction, databases, and cloud systems. They have been successfully used in 13 companies and domain science groups, have been pushed into production for many customers, and are being commercialized. We discuss the lessons learned and explore applying the Magellan template to other tasks in data exploration, cleaning, and integration.