Auditable LLM-guided Pareto scenario discovery and hard-bank fine-tuning for robust map-free LiDAR navigation
| dc.contributor.author | Yilmaz, Taner | |
| dc.contributor.author | Aydogmus, Omur | |
| dc.date.accessioned | 2026-09-08T07:13:29Z | |
| dc.date.issued | 2026 | |
| dc.department | Fırat Üniveristesi | |
| dc.description.abstract | Deep reinforcement learning (DRL) policies for map-free LiDAR navigation remain brittle under out-of-distribution (OOD) layouts and systematic perception corruptions, while robustness evaluations often rely on subjective hard environments. This study presents an auditable benchmark-construction and robustification pipeline in which a frozen PPO policy is stress-tested through Pareto scenario discovery and then improved through hard-bank fine-tuning. The large language model (LLM) is not used as an online controller; it serves only as an offline, schema-constrained proposal module that converts aggregated failure summaries into bounded sampling-bias suggestions, while validation, acceptance, coverage tracking, and archiving remain deterministic. A base PPO policy was trained in randomized static layouts and frozen. Two discovery modes were then executed over 10 seeds: NoLLM random/heuristic search and LLM-guided search. Separate hard banks were compiled and used to fine-tune two robust policies under a fixed IID+hard mixture (mix_ratio=0.5). Across three 3000-scenario suites with two repeats, robust fine-tuning produced consistent gains. On Hard-LLM, success increased from 91.70% for Base-100k to 98.83% for Robust-LLM, while timeout decreased from 7.10% to 0.32%. The Hard-LLM suite was also harder for the frozen base policy than Hard-NoLLM (91.70% vs. 96.93% success), confirming a benchmark shift. TurtleBot3 Burger validation across four unseen layouts supported transferability, with success improving from 16/40 to 30/40 runs. Results indicate that LLM guidance mainly improves discovery prioritization and sample efficiency, whereas high-budget fine-tuning can produce a ceiling effect. | |
| dc.identifier.doi | 10.1016/j.robot.2026.105628 | |
| dc.identifier.issn | 0921-8890 | |
| dc.identifier.issn | 1872-793X | |
| dc.identifier.scopus | 2-s2.0-105045004876 | |
| dc.identifier.scopusquality | Q1 | |
| dc.identifier.uri | https://doi.org/10.1016/j.robot.2026.105628 | |
| dc.identifier.uri | https://hdl.handle.net/11508/65466 | |
| dc.identifier.volume | 205 | |
| dc.identifier.wos | WOS:001826280200001 | |
| dc.identifier.wosquality | Q1 | |
| dc.indekslendigikaynak | Web of Science | |
| dc.indekslendigikaynak | Scopus | |
| dc.language.iso | en | |
| dc.publisher | Elsevier | |
| dc.relation.ispartof | Robotics and Autonomous Systems | |
| dc.relation.publicationcategory | Makale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı | |
| dc.rights | info:eu-repo/semantics/closedAccess | |
| dc.snmz | KA_WOS_20250903 | |
| dc.subject | Map-Free Navigation | |
| dc.subject | Lidar | |
| dc.subject | Deep Reinforcement Learning | |
| dc.subject | Robustness | |
| dc.subject | Scenario Discovery | |
| dc.subject | Large Language Models | |
| dc.title | Auditable LLM-guided Pareto scenario discovery and hard-bank fine-tuning for robust map-free LiDAR navigation | |
| dc.type | Article |







