A Multiclass Approach to Estimating Software Vulnerability Severity Rating with Statistical and Word Embedding Methods

dc.contributor.authorKekül, Hakan
dc.contributor.authorErgen, Burhan
dc.contributor.authorArslan, Halil
dc.date.accessioned2026-08-12T16:15:01Z
dc.date.issued2022
dc.departmentFırat Üniversitesi
dc.description.abstractThe analysis and grading of software vulnerabilities is an important process that is done manually by experts today. For this reason, there are time delays, human errors, and excessive costs involved with the process. The final result of these software vulnerability reports created by experts is the calculation of a severity score and a severity rating. The severity rating is the first and foremost value of the software’s vulnerability. The vulnerabilities that can be exploited are only 20% of the total vulnerabilities. The vast majority of exploitations take place within the first two weeks. It is therefore imperative to determine the severity rating without time delays. Our proposed model uses statistical methods and deep learning-based word embedding methods from natural language processing techniques, and machine learning algorithms that perform multi-class classification. Bag of Words, Term Frequency Inverse Document Frequency and Ngram methods, which are statistical methods, were used for feature extraction. Word2Vec, Doc2Vec and Fasttext algorithms are included in the study for deep learning based Word embedding. In the classification stage, Naive Bayes, Decision Tree, K-Nearest Neighbors, Multi-Layer Perceptron, and Random Forest algorithms that can make multi-class classification were preferred. With this aspect, our model proposes a hybrid method. The database used is open to the public and is the most reliable data set in the field. The results obtained in our study are quite promising. By helping experts in this field, procedures will speed up. In addition, our study is one of the first studies containing the latest version of the data size and scoring systems it covers. © 2022 MECS.
dc.description.sponsorshipTürkiye Bilimsel ve Teknolojik Araştırma Kurumu, TÜBİTAK, (121E298)
dc.identifier.doi10.5815/ijcnis.2022.04.03
dc.identifier.endpage42
dc.identifier.issn2074-9090
dc.identifier.issue4
dc.identifier.scopus2-s2.0-85135144935
dc.identifier.scopusqualityQ2
dc.identifier.startpage27
dc.identifier.urihttps://doi.org/10.5815/ijcnis.2022.04.03
dc.identifier.urihttps://hdl.handle.net/11508/43456
dc.identifier.volume14
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherModern Education and Computer Science Press
dc.relation.ispartofInternational Journal of Computer Network and Information Security
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_Scopus_20260511
dc.subjectInformation security; Multiclass Classification; Software Security; Software Vulnerability; Text Analysis
dc.titleA Multiclass Approach to Estimating Software Vulnerability Severity Rating with Statistical and Word Embedding Methods
dc.typeArticle

Dosyalar