Comparision of String Matching Algorithms on Spam Email Detection

dc.contributor.authorVarol, Cihan
dc.contributor.authorAbdulhadi, Hezha M. Tareq
dc.date.accessioned2026-08-12T16:41:49Z
dc.date.issued2018
dc.departmentFırat Üniversitesi
dc.descriptionInternational Congress on Big Data, Deep Learning and Fighting Cyber Terrorism (IBIGDELFT) -- DEC 03-04, 2018 -- Turkish IT Author, Ankara, TURKEY
dc.description.abstractEmail is one of the most expedient approach to transfer messages among people all over the world. Its features, specifically reliability, quickness, and low cost makes it popular and useful among people in most parts of businesses and society. On the other hand, this popularity also created new harmful actions, such as email attacks (spam) in cyberspace. Spam is arguably one of the main reasons of drowning the WWW with many copies of similar messages generated through anonymous senders, which yields to time/space wasting of the email account holder and also a large virus and malware threat to Email providers. In spite of employing various filters to handle spam problem such as machine learning and content-based filtering, spammers are still able to bypass these defense mechanisms. In this paper, we investigate the use of string matching algorithms for spam email detection. Particularly this work examines and compares the efficiency of six well-known string matching algorithms, namely Longest Common Subsequence (LCS), Levenshtein Distance (LD), Jaro, Jaro-Winkler, Bi-gram, and TFIDF on two various datasets which are Enron corpus and CSDMC2010 spam dataset. We observed that Bi-gram algorithm performs best in spam detection in both datasets.
dc.description.sponsorshipGazi Univ,Minist Transportat & Infrastrucuture Turkey,Havelsan,Aselsan,BiSoft,Oracle,Proda,Netas,RStudio,Cisco,IEEE Turkey Sect,Informat & Commun Technologies Author
dc.identifier.endpage11
dc.identifier.isbn978-1-7281-0472-0
dc.identifier.orcid0000-0002-4940-6808
dc.identifier.orcid0000-0001-7977-5324
dc.identifier.scopus2-s2.0-85062719010
dc.identifier.scopusqualityN/A
dc.identifier.startpage6
dc.identifier.urihttps://hdl.handle.net/11508/46002
dc.identifier.wosWOS:000459239400002
dc.identifier.wosqualityN/A
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherIeee
dc.relation.ispartof2018 International Congress on Big Data, Deep Learning and Fighting Cyber Terrorism (Ibigdelft)
dc.relation.publicationcategoryKonferans Öğesi - Uluslararası - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/closedAccess
dc.snmzKA_WoS_20260511
dc.subjectBi-gram
dc.subjectJaro distance
dc.subjectJaro-Winkler
dc.subjectLevenshtein distance
dc.subjectLongest Common Subsequence
dc.subjectSpam detection
dc.subjectString Similarity
dc.subjectTFIDF
dc.titleComparision of String Matching Algorithms on Spam Email Detection
dc.typeConference Object

Dosyalar