Balancing Security and Performance in LLM Agents: Spotlight-Guard, a Layered Defense Against Indirect Prompt Injection
| dc.contributor.author | Demirol, Doygun | |
| dc.contributor.author | Aydogan, Murat | |
| dc.date.accessioned | 2026-09-08T07:11:53Z | |
| dc.date.issued | 2026 | |
| dc.department | Fırat Üniveristesi | |
| dc.description.abstract | Large Language Model (LLM)-based agents automate complex tasks by integrating external tools such as web browsers, e-mail clients, file readers, and APIs, but this same integration exposes them to indirect prompt injection (IPI) attacks, in which malicious instructions hidden in tool content hijack the agent. A central but often overlooked question is how defending against such attacks affects the LLM and its own task performance and computational efficiency. In this study, we design a comprehensive testbed and a layered defense, Spotlight-Guard, that combines spotlighting-based input isolation, an LLM detection-and-quarantine pipeline, and instruction integrity based on a Hash-based Message Authentication Code (HMAC) into a single framework, and we evaluate it jointly along two axes: security and LLM performance. Experiments on locally hosted 7B-class open-weight models (Qwen-2.5-7B, Mistral-7B, and DeepSeek-Coder) use Attack Success Rate (ASR) for security and benign-task success rate together with confusion-matrix-based metrics (precision, recall, and F1) for task performance, all with bootstrap 95% confidence intervals. Across a stratified, fixed-seed benchmark of 250 adversarial and 250 benign cases per configuration, the full system reduces the ASR from 36.0% to 17.2% while preserving a 97.2% benign-task success rate and raising the detection F1 from 0.749 to 0.892, demonstrating that strong protection need not degrade the model's task performance. A component ablation isolates each layer's contribution, an adaptive-attack evaluation confirms a low ASR (6.7%) under attacks crafted to target the pipeline, and an analysis of computational cost (model invocations per request) quantifies the efficiency overhead, characterizing the security-performance trade-off of layered defenses on open-weight LLMs. | |
| dc.description.sponsorship | Fimath;rat University [TEKF.25.55] -- This study was supported by the Scientific Research Projects Coordination Unit of F & imath;rat University (FUBAP) under project no. TEKF.25.55. | |
| dc.identifier.doi | 10.3390/app16157662 | |
| dc.identifier.issn | 2076-3417 | |
| dc.identifier.issue | 15 | |
| dc.identifier.scopus | 2-s2.0-105047095633 | |
| dc.identifier.scopusquality | Q1 | |
| dc.identifier.uri | https://doi.org/10.3390/app16157662 | |
| dc.identifier.uri | https://hdl.handle.net/11508/65201 | |
| dc.identifier.volume | 16 | |
| dc.identifier.wos | WOS:001846615800001 | |
| dc.identifier.wosquality | Q2 | |
| dc.indekslendigikaynak | Web of Science | |
| dc.indekslendigikaynak | Scopus | |
| dc.language.iso | en | |
| dc.publisher | Mdpi | |
| dc.relation.ispartof | Applied Sciences-Basel | |
| dc.relation.publicationcategory | Makale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı | |
| dc.rights | info:eu-repo/semantics/openAccess | |
| dc.snmz | KA_WOS_20250903 | |
| dc.subject | Indirect Prompt Injection | |
| dc.subject | Llm Agents | |
| dc.subject | Llm Security | |
| dc.subject | Defense-In-Depth | |
| dc.subject | Security-Performance Trade-Off | |
| dc.subject | Llm Performance Evaluation | |
| dc.subject | Open-Weight Language Models | |
| dc.title | Balancing Security and Performance in LLM Agents: Spotlight-Guard, a Layered Defense Against Indirect Prompt Injection | |
| dc.type | Article |







