Balancing Security and Performance in LLM Agents: Spotlight-Guard, a Layered Defense Against Indirect Prompt Injection

dc.contributor.authorDemirol, Doygun
dc.contributor.authorAydogan, Murat
dc.date.accessioned2026-09-08T07:11:53Z
dc.date.issued2026
dc.departmentFırat Üniveristesi
dc.description.abstractLarge Language Model (LLM)-based agents automate complex tasks by integrating external tools such as web browsers, e-mail clients, file readers, and APIs, but this same integration exposes them to indirect prompt injection (IPI) attacks, in which malicious instructions hidden in tool content hijack the agent. A central but often overlooked question is how defending against such attacks affects the LLM and its own task performance and computational efficiency. In this study, we design a comprehensive testbed and a layered defense, Spotlight-Guard, that combines spotlighting-based input isolation, an LLM detection-and-quarantine pipeline, and instruction integrity based on a Hash-based Message Authentication Code (HMAC) into a single framework, and we evaluate it jointly along two axes: security and LLM performance. Experiments on locally hosted 7B-class open-weight models (Qwen-2.5-7B, Mistral-7B, and DeepSeek-Coder) use Attack Success Rate (ASR) for security and benign-task success rate together with confusion-matrix-based metrics (precision, recall, and F1) for task performance, all with bootstrap 95% confidence intervals. Across a stratified, fixed-seed benchmark of 250 adversarial and 250 benign cases per configuration, the full system reduces the ASR from 36.0% to 17.2% while preserving a 97.2% benign-task success rate and raising the detection F1 from 0.749 to 0.892, demonstrating that strong protection need not degrade the model's task performance. A component ablation isolates each layer's contribution, an adaptive-attack evaluation confirms a low ASR (6.7%) under attacks crafted to target the pipeline, and an analysis of computational cost (model invocations per request) quantifies the efficiency overhead, characterizing the security-performance trade-off of layered defenses on open-weight LLMs.
dc.description.sponsorshipFimath;rat University [TEKF.25.55] -- This study was supported by the Scientific Research Projects Coordination Unit of F & imath;rat University (FUBAP) under project no. TEKF.25.55.
dc.identifier.doi10.3390/app16157662
dc.identifier.issn2076-3417
dc.identifier.issue15
dc.identifier.scopus2-s2.0-105047095633
dc.identifier.scopusqualityQ1
dc.identifier.urihttps://doi.org/10.3390/app16157662
dc.identifier.urihttps://hdl.handle.net/11508/65201
dc.identifier.volume16
dc.identifier.wosWOS:001846615800001
dc.identifier.wosqualityQ2
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherMdpi
dc.relation.ispartofApplied Sciences-Basel
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_WOS_20250903
dc.subjectIndirect Prompt Injection
dc.subjectLlm Agents
dc.subjectLlm Security
dc.subjectDefense-In-Depth
dc.subjectSecurity-Performance Trade-Off
dc.subjectLlm Performance Evaluation
dc.subjectOpen-Weight Language Models
dc.titleBalancing Security and Performance in LLM Agents: Spotlight-Guard, a Layered Defense Against Indirect Prompt Injection
dc.typeArticle

Dosyalar