Adaptive Task-Oriented Locomotion Control of a 2D Planar Robotic Fish Model Using Deep Reinforcement Learning and Sensory-Feedback CPG Network
| dc.contributor.author | Koca, Gonca | |
| dc.contributor.author | Korkmaz, Deniz | |
| dc.contributor.author | Bal, Cafer | |
| dc.contributor.author | Ay, Mustafa | |
| dc.contributor.author | Akpolat, Zuhtu Hakan | |
| dc.date.accessioned | 2026-09-08T07:11:49Z | |
| dc.date.issued | 2026 | |
| dc.department | Fırat Üniveristesi | |
| dc.description.abstract | Autonomous locomotion in robotic fish requires task-dependent control capabilities under changing environmental conditions. This paper proposes a hierarchical simulation-based control framework for a two-joint robotic fish in a two-dimensional (2D) planar environment. This framework integrates the twin delayed deep deterministic policy gradient (TD3) algorithm with a sensory-feedback central pattern generator (CPG). A nonlinear planar dynamic model is designed as the learning environment, and a CPG network generates rhythmic undulatory swimming. The CPG network generates smooth locomotor patterns, while the TD3 policy performs high-level neuromotor modulation for task-dependent behavior. In the target-reaching benchmark, TD3-CPG achieves a 100.0% success rate with a Wilson 95% confidence interval (CI) of [96.30%, 100.00%], outperforming benchmark models. The proposed controller is also evaluated with obstacle avoidance in target reaching and station keeping under current disturbances. In circular obstacle avoidance, TD3-CPG achieves a 98.0% success rate and a 98.0% safe-pass rate, whereas the multiple rectangular obstacle scenarios yield an overall success rate of 91.7% over 96 trials. In station keeping, the controller achieves stay ratios of 87.57 +/- 12.81% under constant current and 96.88 +/- 10.79% under gust current, while keeping the mean target distance below the 0.25 m station keeping radius in both cases. Within the adopted 2D planar simulation environment, the obtained results demonstrate that the proposed method exhibits task-dependent maneuvering performance within the evaluated scenarios. | |
| dc.description.sponsorship | Firat University [ADEP.23.22] -- This study was funded by Scientific Research Projects Unit of Firat University (FUBAP) under the Grant Number ADEP.23.22. The authors thank FUBAP for their financial support. | |
| dc.identifier.doi | 10.3390/biomimetics11080534 | |
| dc.identifier.issn | 2313-7673 | |
| dc.identifier.issue | 8 | |
| dc.identifier.scopus | 2-s2.0-105048228150 | |
| dc.identifier.scopusquality | Q3 | |
| dc.identifier.uri | https://doi.org/10.3390/biomimetics11080534 | |
| dc.identifier.uri | https://hdl.handle.net/11508/65173 | |
| dc.identifier.volume | 11 | |
| dc.identifier.wos | WOS:001859290400001 | |
| dc.identifier.wosquality | Q1 | |
| dc.indekslendigikaynak | Web of Science | |
| dc.indekslendigikaynak | Scopus | |
| dc.language.iso | en | |
| dc.publisher | Mdpi | |
| dc.relation.ispartof | Biomimetics | |
| dc.relation.publicationcategory | Makale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı | |
| dc.rights | info:eu-repo/semantics/openAccess | |
| dc.snmz | KA_WOS_20250903 | |
| dc.subject | Robotic Fish | |
| dc.subject | Reinforcement Learning | |
| dc.subject | Cpg | |
| dc.subject | Locomotion Control | |
| dc.subject | Dynamic Modeling | |
| dc.subject | Adaptive Control | |
| dc.title | Adaptive Task-Oriented Locomotion Control of a 2D Planar Robotic Fish Model Using Deep Reinforcement Learning and Sensory-Feedback CPG Network | |
| dc.type | Article |







