Skip to main content
Handel Schweiz

Intelligent duplicate detection for addresses and contacts

For Handel Schweiz, CONSENSO TECH developed a small application for intelligent duplicate detection in address and contact data. A vector-based approach identifies not only exact duplicates but also similar spellings, abbreviations and differing contact details.

At a glance

Starting point

  • Duplicates in company addresses and contacts
  • Differing spellings, typos, abbreviations and additional information
  • Standard functions often detect only exact matches
  • Data quality is critical for invoicing and reporting
  • Manual duplicate detection is time-consuming and error-prone

Solution

  • Concept and prototype for intelligent duplicate detection
  • Vector-based comparison of address and contact data
  • Use of TF-IDF and cosine similarity
  • Assessment of potential duplicates using similarity scores
  • A foundation for UI-supported cleansing, reports and data quality dashboards
  • Flexible thresholds and weightings for different data quality rules

Result

  • Better detection of non-obvious duplicates
  • Higher data quality in address and contact records
  • Less manual checking effort
  • A more reliable foundation for invoicing and reporting
  • The option of a continuous data quality cycle
  • A basis for proactive master data management

During the ERP transformation at Handel Schweiz it became clear that data quality is a critical success factor for correct invoicing, reporting and efficient operational processes. Duplicates in company addresses and contacts were a particular challenge, as they do not always appear as exact one-to-one duplicates but often contain small deviations, different spellings or supplementary information.

CONSENSO TECH developed a concept and a prototype for intelligent duplicate detection. Addresses and contacts are converted into a mathematical form and compared pairwise using methods such as TF-IDF and cosine similarity. This makes it possible to identify similar records where classic duplicate checks reach their limits.

The approach enables a more flexible and more precise data quality check than plain standard matching. Potential duplicates can be scored by similarity and then processed further through workflows, UI-supported cleansing or reporting.