A Threshold-based Similarity Measure for Duplicate Detection -

16Jul 2014 by chintan

A Threshold-based Similarity Measure for Duplicate Detection

In order to extract beneficial information and recognize a particular pattern from huge data stored in different databases with different formats, data integration is essential. However the problem that arises here is that data integration may lead to duplication. In other words, due to the availability of data in different formats, there might be some records which refer to the same entity. Duplicate detection or record linkage is a technique which is used to detect and match duplicate records which are generated in data integration process. Most approaches concentrated on string similarity measures for comparing records. However, they fail to identify records which share the semantic information. So, in this study, a thresholdbased method which takes into account both string and semantic similarity measures for comparing record pairs. This method is experimented on a real world dataset, namely Restaurant and its effectiveness is measured based on several standard evaluation metrics. As experimental results indicate, the proposed similarity method which is based on the combination of string and semantic similarity measures outperforms the individual similarity measures with the F-measure of 99.1% in Restaurant dataset. Therefore, based on experimental results, besides string similarity, semantic similarity should be considered in order to detect duplicate records more effectively.

HR Attendance System Using RFID

Mining Facets For Queries From Their Search Result...

E Healthcare – Online Consultation And Medical Sub...

Online Diagnostic Lab Reporting System Php

Android Based Parking Booking System

Cooking Recipe Rating Based On Sentiment Analysis

Analysis of Denial-of-Service attacks on Wireless ...

An Efficient Cross-Layer Approach for Malicious Pa...

Cargo Booking Software

Movie Success Prediction Using Data Mining PHP

Bikers Portal

Image-based object detection under varying illumin...