Structured Named Entity Recognition (NER) in Biomedical Texts Using Pre-Trained Language Models

dc.contributor.authorSavci, Pinar
dc.contributor.authorDas, Bihter
dc.date.accessioned2026-08-12T16:08:09Z
dc.date.issued2024
dc.departmentFırat Üniversitesi
dc.description12th International Symposium on Digital Forensics and Security, ISDFS 2024 -- 29 April 2024 through 30 April 2024 -- San Antonio -- 199532
dc.description.abstractThe field of Natural Language Processing (NLP) has witnessed remarkable progress in recent years, particularly in the domain of biomedical text analysis. Named Entity Recognition (NER), a pivotal task in information extraction, holds the key to deciphering and extracting structured information from unstructured biomedical texts. Accurate identification and classification of entities, such as DNA, proteins, cell types, cell lines, and RNA, are imperative for advancing our comprehension of complex biological systems. This paper presents a comprehensive exploration of the application of state-of-the-art pre-trained language models, including Bert-base-cased, Distilbert-base-cased, Albert-base-V2, Xml-roberta-base, Ernie-2.0-base-en, and Conv-bert-base, for structured Named Entity Recognition in biomedical texts. The BioNLP2004 dataset, enriched with diverse entity types, forms the basis for our experiments. Our objectives encompass a thorough investigation into the effectiveness of different pre-trained language models for biomedical NER, an in-depth analysis of the challenges posed by the BioNLP2004 dataset, and a comparative evaluation of the selected models in terms of precision, recall, and F1 score. Additionally, we explore the impact of fine-tuning strategies on model performance. The insights gained from this research have the potential to advance the capabilities of language models in the biomedical domain, contributing to more efficient and precise biomedical text analysis. This work serves as a stepping stone towards the broader goals of bioinformatics and medical research. The paper concludes with a summary of findings and outlines potential avenues for future research. © 2024 IEEE.
dc.description.sponsorshipArçelik Digital Transformation, Big Data and Artificial Intelligence R&D Center; Ministry of Science, Technology and Industry, (AR-22-087-0001)
dc.identifier.doi10.1109/ISDFS60797.2024.10527329
dc.identifier.isbn979-835033036-6
dc.identifier.scopus2-s2.0-85194090765
dc.identifier.scopusqualityN/A
dc.identifier.urihttps://doi.org/10.1109/ISDFS60797.2024.10527329
dc.identifier.urihttps://hdl.handle.net/11508/41048
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherInstitute of Electrical and Electronics Engineers Inc.
dc.relation.ispartof12th International Symposium on Digital Forensics and Security, ISDFS 2024
dc.relation.publicationcategoryKonferans Öğesi - Uluslararası - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/closedAccess
dc.snmzKA_Scopus_20260511
dc.subjectbiomedical text; BioNLP2004 Dataset; Named Entity Recognition (NER); Natural Language Processing; Pre-trained models
dc.titleStructured Named Entity Recognition (NER) in Biomedical Texts Using Pre-Trained Language Models
dc.typeConference Object

Dosyalar