International audience
In this paper, we describe a solution for a specific Entity Matching problem, where entities contain (postal) address information. The matching process is very challenging as addresses are often prone to (data) quality issues such as typos, missing or redundant information. Besides, they do not always comply with a standardized (address) schema and may contain polysemous elements. Recent address matching approaches combine static word embedding models with machine learning algorithms. While the solutions provided in this setting partially solve data quality issues, neither they handle polysemy, nor they leverage of geolocation information. In this paper, we propose GeoRoBERTa, a semantic address matching approach based on RoBERTa, a Transformer-based model, enhanced by geographical knowledge. We validate the approach in conducting experiments on two different real datasets and demonstrate its effectiveness in comparison to baseline methods.