Iterated and Linear SVMs in Text Classification: An Ensemble Feature Engineering Model

dc.contributor.authorAydemir, Emrah
dc.contributor.authorBarua, Prabal Datta
dc.contributor.authorChakraborty, Subrata
dc.contributor.authorHafeez-Baig, Abdul
dc.contributor.authorDogan, Sengul
dc.contributor.authorTuncer, Turker
dc.contributor.authorAcharya, U. R.
dc.date.accessioned2026-08-12T17:27:18Z
dc.date.issued2026
dc.departmentFırat Üniversitesi
dc.description.abstractTerm frequency-inverse document frequency (TF-IDF), a widely used method for feature representation in text classification, has limitations such as inconsistent results and high dimensionality. The main objective of this research is to develop an accurate feature engineering model for text classification. We propose a novel TF-IDF-based feature engineering architecture for legal text classification, which consists of 4 phases: (1) feature extraction using TF-IDF; (2) feature selection with iterative neighborhood component analysis, iterative Chi2, iterative ReliefF, and iterative minimum redundancy maximum relevance, generating four distinct feature vectors; (3) classification of the selected feature vectors using an iterative linear support vector machine, producing 40 prediction vectors (10 for each feature vector); and (4) final result selection using a greedy algorithm and these phases makes the architecture self-organizing.The dataset used in this study included 235 court decision documents (108 accepted and 127 rejected). These documents were collected from the LegalBank dataset of the European Court of Human Rights and were translated into Turkish by lawyers to enable the application of this model to Turkish courts. Our proposed model achieved a classification accuracy of 99.15% using tenfold cross-validation. The model operates with linear time complexity, making it efficient for law-related text classification. The ultimate goal of this research is to build a digital assistant for Turkish courts.
dc.description.sponsorshipScientific and Technological Research Council of Turkey (TUBITAK) through the Scientific and Technological Research Projects Support Program [122G019]
dc.description.sponsorshipThis study was supported by the Scientific and Technological Research Council of Turkey (TUBITAK) through the Scientific and Technological Research Projects Support Program (3005) with project number 122G019.
dc.identifier.doi10.1007/s13369-025-10666-0
dc.identifier.endpage6284
dc.identifier.issn2193-567X
dc.identifier.issn2191-4281
dc.identifier.issue5
dc.identifier.orcid0000-0002-8380-7891
dc.identifier.scopus2-s2.0-105018482132
dc.identifier.scopusqualityQ1
dc.identifier.startpage6275
dc.identifier.urihttps://doi.org/10.1007/s13369-025-10666-0
dc.identifier.urihttps://hdl.handle.net/11508/55153
dc.identifier.volume51
dc.identifier.wosWOS:001586251300001
dc.identifier.wosqualityQ2
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherSpringer Heidelberg
dc.relation.ispartofArabian Journal for Science and Engineering
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/closedAccess
dc.snmzKA_WoS_20260511
dc.subjectLaw text classification
dc.subjectEnsemble feature engineering model
dc.subjectAutomated decision-making
dc.subjectTF-IDF
dc.subjectEuropean court of human right
dc.titleIterated and Linear SVMs in Text Classification: An Ensemble Feature Engineering Model
dc.typeArticle

Dosyalar