arXiv · 2403.06434
BoostER: Leveraging Large Language Models for Enhancing Entity Resolution
Abstract
Entity resolution, which involves identifying and merging records that refer to the same real-world entity, is a crucial task in areas like Web data integration. This importance is underscored by the presence of numerous duplicated and multi-version data resources on the Web. However, achieving high-quality entity resolution typically demands significant effort. The advent of Large Language Models (LLMs) like GPT-4 has demonstrated advanced linguistic capabilities, which can be a new paradigm for this task. In this paper, we propose a demonstration system named BoostER that examines the possibility of leveraging LLMs in the entity resolution process, revealing advantages in both easy deployment and low cost. Our approach optimally selects a set of matching questions and poses them to LLMs for verification, then refines the distribution of entity resolution results with the response of LLMs. This offers promising prospects to achieve a high-quality entity resolution result for real-world applications, especially to individuals or small companies without the need for extensive model training or significant financial investment.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Huahang Li, Shuangyin Li, Fei Hao, Chen Jason Zhang, Yuanfeng Song, Lei Chen. 2024-03-11. BoostER: Leveraging Large Language Models for Enhancing Entity Resolution. https://doi.org/10.1145/3589335.3651245%2010.1145%2F3589335.3651245%2010.1145%2F3589335.3651245
Cite the original work for its findings. Save a collection to share your selection of sources.