Magellan: Toward Building Ecosystems of Entity Matching Solutions.

Entity matching (EM) finds data instances that refer to the same real-world entity. In 2015, we started the Magellan project at UW-Madison, jointly with industrial partners, to build EM systems. Most current EM systems are standalone monoliths. In contrast, Magellan borrows ideas from the field of d...

Descripción completa

Detalles Bibliográficos
Publicado en:Communications of the ACM Vol. 63; no. 8; pp. 83 - 92
Autores principales: Doan, AnHai, Konda, Pradap, G. C., Paul Suganthan, Govind, Yash
Formato: Artículo
Publicado: Association for Computing Machinery Aug2020
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:Entity matching (EM) finds data instances that refer to the same real-world entity. In 2015, we started the Magellan project at UW-Madison, jointly with industrial partners, to build EM systems. Most current EM systems are standalone monoliths. In contrast, Magellan borrows ideas from the field of data science (DS), to build a new kind of EM systems, which is ecosystems of interoperable tools for multiple execution environments, such as on-premise, cloud, and mobile. This paper describes Magellan, focusing on the system aspects. We argue why EM can be viewed as a special class of DS problems and thus can benefit from system building ideas in DS. We discuss how these ideas have been adapted to build PyMatcher and CloudMatcher, sophisticated on-premise tools for power users and self-service cloud tools for lay users. These tools exploit techniques from the fields of machine learning, big data scaling, efficient user interaction, databases, and cloud systems. They have been successfully used in 13 companies and domain science groups, have been pushed into production for many customers, and are being commercialized. We discuss the lessons learned and explore applying the Magellan template to other tasks in data exploration, cleaning, and integration.