Auditable LLM-guided Pareto scenario discovery and hard-bank fine-tuning for robust map-free LiDAR navigation

dc.contributor.authorYilmaz, Taner
dc.contributor.authorAydogmus, Omur
dc.date.accessioned2026-09-08T07:13:29Z
dc.date.issued2026
dc.departmentFırat Üniveristesi
dc.description.abstractDeep reinforcement learning (DRL) policies for map-free LiDAR navigation remain brittle under out-of-distribution (OOD) layouts and systematic perception corruptions, while robustness evaluations often rely on subjective hard environments. This study presents an auditable benchmark-construction and robustification pipeline in which a frozen PPO policy is stress-tested through Pareto scenario discovery and then improved through hard-bank fine-tuning. The large language model (LLM) is not used as an online controller; it serves only as an offline, schema-constrained proposal module that converts aggregated failure summaries into bounded sampling-bias suggestions, while validation, acceptance, coverage tracking, and archiving remain deterministic. A base PPO policy was trained in randomized static layouts and frozen. Two discovery modes were then executed over 10 seeds: NoLLM random/heuristic search and LLM-guided search. Separate hard banks were compiled and used to fine-tune two robust policies under a fixed IID+hard mixture (mix_ratio=0.5). Across three 3000-scenario suites with two repeats, robust fine-tuning produced consistent gains. On Hard-LLM, success increased from 91.70% for Base-100k to 98.83% for Robust-LLM, while timeout decreased from 7.10% to 0.32%. The Hard-LLM suite was also harder for the frozen base policy than Hard-NoLLM (91.70% vs. 96.93% success), confirming a benchmark shift. TurtleBot3 Burger validation across four unseen layouts supported transferability, with success improving from 16/40 to 30/40 runs. Results indicate that LLM guidance mainly improves discovery prioritization and sample efficiency, whereas high-budget fine-tuning can produce a ceiling effect.
dc.identifier.doi10.1016/j.robot.2026.105628
dc.identifier.issn0921-8890
dc.identifier.issn1872-793X
dc.identifier.scopus2-s2.0-105045004876
dc.identifier.scopusqualityQ1
dc.identifier.urihttps://doi.org/10.1016/j.robot.2026.105628
dc.identifier.urihttps://hdl.handle.net/11508/65466
dc.identifier.volume205
dc.identifier.wosWOS:001826280200001
dc.identifier.wosqualityQ1
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherElsevier
dc.relation.ispartofRobotics and Autonomous Systems
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/closedAccess
dc.snmzKA_WOS_20250903
dc.subjectMap-Free Navigation
dc.subjectLidar
dc.subjectDeep Reinforcement Learning
dc.subjectRobustness
dc.subjectScenario Discovery
dc.subjectLarge Language Models
dc.titleAuditable LLM-guided Pareto scenario discovery and hard-bank fine-tuning for robust map-free LiDAR navigation
dc.typeArticle

Dosyalar