Comparision of String Matching Algorithms on Spam Email Detection
| dc.contributor.author | Varol, Cihan | |
| dc.contributor.author | Abdulhadi, Hezha M. Tareq | |
| dc.date.accessioned | 2026-08-12T16:41:49Z | |
| dc.date.issued | 2018 | |
| dc.department | Fırat Üniversitesi | |
| dc.description | International Congress on Big Data, Deep Learning and Fighting Cyber Terrorism (IBIGDELFT) -- DEC 03-04, 2018 -- Turkish IT Author, Ankara, TURKEY | |
| dc.description.abstract | Email is one of the most expedient approach to transfer messages among people all over the world. Its features, specifically reliability, quickness, and low cost makes it popular and useful among people in most parts of businesses and society. On the other hand, this popularity also created new harmful actions, such as email attacks (spam) in cyberspace. Spam is arguably one of the main reasons of drowning the WWW with many copies of similar messages generated through anonymous senders, which yields to time/space wasting of the email account holder and also a large virus and malware threat to Email providers. In spite of employing various filters to handle spam problem such as machine learning and content-based filtering, spammers are still able to bypass these defense mechanisms. In this paper, we investigate the use of string matching algorithms for spam email detection. Particularly this work examines and compares the efficiency of six well-known string matching algorithms, namely Longest Common Subsequence (LCS), Levenshtein Distance (LD), Jaro, Jaro-Winkler, Bi-gram, and TFIDF on two various datasets which are Enron corpus and CSDMC2010 spam dataset. We observed that Bi-gram algorithm performs best in spam detection in both datasets. | |
| dc.description.sponsorship | Gazi Univ,Minist Transportat & Infrastrucuture Turkey,Havelsan,Aselsan,BiSoft,Oracle,Proda,Netas,RStudio,Cisco,IEEE Turkey Sect,Informat & Commun Technologies Author | |
| dc.identifier.endpage | 11 | |
| dc.identifier.isbn | 978-1-7281-0472-0 | |
| dc.identifier.orcid | 0000-0002-4940-6808 | |
| dc.identifier.orcid | 0000-0001-7977-5324 | |
| dc.identifier.scopus | 2-s2.0-85062719010 | |
| dc.identifier.scopusquality | N/A | |
| dc.identifier.startpage | 6 | |
| dc.identifier.uri | https://hdl.handle.net/11508/46002 | |
| dc.identifier.wos | WOS:000459239400002 | |
| dc.identifier.wosquality | N/A | |
| dc.indekslendigikaynak | Web of Science | |
| dc.indekslendigikaynak | Scopus | |
| dc.language.iso | en | |
| dc.publisher | Ieee | |
| dc.relation.ispartof | 2018 International Congress on Big Data, Deep Learning and Fighting Cyber Terrorism (Ibigdelft) | |
| dc.relation.publicationcategory | Konferans Öğesi - Uluslararası - Kurum Öğretim Elemanı | |
| dc.rights | info:eu-repo/semantics/closedAccess | |
| dc.snmz | KA_WoS_20260511 | |
| dc.subject | Bi-gram | |
| dc.subject | Jaro distance | |
| dc.subject | Jaro-Winkler | |
| dc.subject | Levenshtein distance | |
| dc.subject | Longest Common Subsequence | |
| dc.subject | Spam detection | |
| dc.subject | String Similarity | |
| dc.subject | TFIDF | |
| dc.title | Comparision of String Matching Algorithms on Spam Email Detection | |
| dc.type | Conference Object |







