Zero-Shot Visual Anomaly Detection on a Quadruped Robot Using State of the Art Visual Language Models
| dc.contributor.author | Aydogmus, Omur | |
| dc.contributor.author | Boztas, Gullu | |
| dc.date.accessioned | 2026-08-12T17:28:28Z | |
| dc.date.issued | 2026 | |
| dc.department | Fırat Üniversitesi | |
| dc.description.abstract | Anomaly detection in dynamic real-world environments remains a significant challenge for robotic systems, largely because traditional vision and rule-based methods struggle to interpret complex semantic contexts without task-specific training. Recent advances in Visual Language Models (VLMs) offer new opportunities for robots to perform zero-shot, context aware perception; however, their practical deployment on mobile robotic platforms remains underexplored. In this study, a quadruped robot autonomously patrolled both indoor and outdoor environments following a predefined trajectory and schedule to perform real-time, zero-shot anomaly detection. The robot was equipped with an onboard RGB camera and integrated with a VLM to interpret visual data without any fine tuning. Anomalies were identified directly through prompt-based semantic reasoning, enabling detection of misplaced objects, structural defects, and environmental hazards. The overall framework was implemented under ROS2 to ensure seamless communication, control, and real-time decision making. Two inference configurations were evaluated: a lightweight VLM running locally on the robot for on-device processing, and a more powerful cloud-based VLM used for remote inference. Experimental results show that both configurations effectively identified diverse anomalies, demonstrating the feasibility of combining quadruped platforms with VLM-based zero-shot perception for continuous monitoring in dynamic environments. Furthermore, a comprehensive benchmarking study was conducted across multiple state-of-the-art VLMs Gemini-2.5-Pro, GPT-5, Claude Sonnet 4.5, Gemma3-27B, Qwen3-VL-30B, and LLaVA-13B alongside human evaluators. This comparison enabled systematic assessment of model behavior and alignment with human judgment, providing deeper insights into the strengths and limitations of modern VLMs for embodied perception tasks. | |
| dc.description.sponsorship | Council of Higher Education (CoHE/YOEK) under the Arastimath;rma Universiteleri Destek Programimath; (ADEP) [ADEP.24.22] | |
| dc.description.sponsorship | This work was supported by the Council of Higher Education (CoHE/YOEK) under the Arast & imath;rma Universiteleri Destek Program & imath; (ADEP) under Grant ADEP.24.22. | |
| dc.identifier.doi | 10.1109/ACCESS.2026.3658108 | |
| dc.identifier.endpage | 17852 | |
| dc.identifier.issn | 2169-3536 | |
| dc.identifier.orcid | 0000-0002-1720-1285 | |
| dc.identifier.scopus | 2-s2.0-105029027504 | |
| dc.identifier.scopusquality | Q1 | |
| dc.identifier.startpage | 17842 | |
| dc.identifier.uri | https://doi.org/10.1109/ACCESS.2026.3658108 | |
| dc.identifier.uri | https://hdl.handle.net/11508/55315 | |
| dc.identifier.volume | 14 | |
| dc.identifier.wos | WOS:001682702600028 | |
| dc.identifier.wosquality | Q2 | |
| dc.indekslendigikaynak | Web of Science | |
| dc.indekslendigikaynak | Scopus | |
| dc.language.iso | en | |
| dc.publisher | Ieee-Inst Electrical Electronics Engineers Inc | |
| dc.relation.ispartof | Ieee Access | |
| dc.relation.publicationcategory | Makale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı | |
| dc.rights | info:eu-repo/semantics/openAccess | |
| dc.snmz | KA_WoS_20260511 | |
| dc.subject | Robots | |
| dc.subject | Anomaly detection | |
| dc.subject | Quadrupedal robots | |
| dc.subject | Visualization | |
| dc.subject | Service robots | |
| dc.subject | Semantics | |
| dc.subject | Robot vision systems | |
| dc.subject | Cameras | |
| dc.subject | Sensors | |
| dc.subject | Real-time systems | |
| dc.subject | Robotics | |
| dc.subject | quadruped robots | |
| dc.subject | vision-language models (VLMs) | |
| dc.subject | zero-shot anomaly detection | |
| dc.subject | embodied AI | |
| dc.subject | multimodal perception | |
| dc.subject | autonomous patrol | |
| dc.subject | ROS2 | |
| dc.title | Zero-Shot Visual Anomaly Detection on a Quadruped Robot Using State of the Art Visual Language Models | |
| dc.type | Article |







