Meningkatkan Deduplikasi Data melalui Kesamaan Teks dalam Pembelajaran Mesin: Pendekatan Komprehensif
DOI:
https://doi.org/10.37481/jmh.v4i2.955Keywords:
Data Deduplication, Text Similarity, Jaro-Winkler SimilarityAbstract
The issue of dirty data, particularly duplicate data, is a common problem in data management that can affect data quality, operational efficiency, and decision-making. This study highlights the importance of implementing sustainable deduplication strategies as a key step in managing dirty data. We explore solutions for detecting duplicate data by measuring text similarity indices. In this study, the authors utilize a literature review research method. Through this method, we collected various journals on data deduplication and text similarity techniques, comparing several methods to identify the most effective approach. The data deduplication process in this study consists of two stages: 1) Matching - calculating the similarity value of a record with previous records, and 2) Clustering - grouping all records deemed duplicates of a single entity. Furthermore, this study extends to the development of a Python application capable of identifying and grouping similar customer data based on text similarity values. The Text Similarity measurement method uses the Jaro-Winkler Similarity technique. Experimental results and evaluations show that the Text Similarity approach is effective in identifying duplicate data with a high degree of accuracy. This study emphasizes the importance of sustainable deduplication, where the deduplication process is conducted periodically and continuously to ensure optimal data quality.
References
Similarity Measures in a Data Deduplication Pipeline for Customers Records. DOLAP, 33–42.
Andrzejewski, W., Bębel, B., Boiński, P., Sienkiewicz, M., & Wrembel, R. (2023). Text similarity measures in a data deduplication pipeline for customers records.
Baloi, A., Belean, B., Turcu, F., & Peptenatu, D. (2024). GPU-based similarity metrics computation and machine learning approaches for string similarity evaluation in large datasets. Soft Computing, 28(4), 3465–3477.
Kavuncuoglu, E. (2024). Exploring the performance of PySpark and Scikit-Learn libraries in developing fall detection systems. Niğde Ömer Halisdemir Üniversitesi Mühendislik Bilimleri Dergisi, 13(2), 582–592.
Krishna, C. M., Ruikar, K., & Jha, K. N. (2023). Determinants of data quality dimensions for assessing highway infrastructure data using semiotic framework. Buildings, 13(4), 944.
Qurashi, A. W., Holmes, V., & Johnson, A. P. (2020). Document processing: Methods for semantic text similarity analysis. 2020 International Conference on INnovations in Intelligent SysTems and Applications (INISTA), 1–6.
Redine, A., Deshpande, S., Jebarajakirthy, C., & Surachartkumtonkun, J. (2023). Impulse buying: A systematic literature review and future research directions. International Journal of Consumer Studies, 47(1), 3–41.
Rustamovna, A. U. (2021). Understanding the levenshtein distance equation for beginners. The American Journal of Engineering and Technology, 3(06), 134–139.
Singh, R. H., Maurya, S., Tripathi, T., Narula, T., & Srivastav, G. (2020). Movie recommendation system using cosine similarity and KNN. International Journal of Engineering and Advanced Technology, 9(5), 556–559.
Verschuuren, P., Palazzo, S., Powell, T., Sutton, S., Pilgrim, A., & Giannelli, M. F. (2020). Supervised machine learning techniques for data matching based on similarity metrics. ArXiv Preprint ArXiv:2007.04001.
Wang, J., & Dong, Y. (2020). Measurement of text similarity: a survey. Information, 11(9), 421.






