SkewSense: ML-Driven Partition Optimization for Distributed Systems

dc.contributor.authorAltuntas, Seyma Nur
dc.contributor.authorBaykara, Muhammet
dc.date.accessioned2026-09-08T07:08:32Z
dc.date.issued2026
dc.departmentFırat Üniveristesi
dc.description8th International Congress on Human-Computer Interaction, Optimization and Robotic Applications, ICHORA 2026 -- 21 May 2026 through 23 May 2026 -- Ankara -- 224404
dc.description.abstractData skew in distributed big data systems causes significant performance degradation by concentrating data around specific key values, leading to overloaded nodes and underutilized resources. This problem is particularly severe in operations requiring data redistribution, such as joins, aggregations, and window functions. In this study, we propose a machine learning (ML)-driven decision system that automatically selects the appropriate partitioning strategy by analyzing data skew. The developed system uses summary statistical metrics of the dataset to choose among hash, range, and custom salting strategies and generates executable Apache Spark Scala code for the recommended structure. The performance of the proposed approach is evaluated in comparison with the PostgreSQL-based distributed database system Citus and the Apache Spark baseline configuration. Experiments are conducted on a synthetic dataset with controlled data skew using four different query scenarios. Results show that the ML-driven custom salting strategy achieves significant performance improvements over Spark baseline, especially in aggregation queries involving hot keys. However, performance gains are limited in window function queries due to additional shuffle overhead. The findings demonstrate that data-aware partitioning strategies can enhance performance in distributed data processing systems and that ML-driven decision mechanisms offer an effective solution in this process. © 2026 IEEE.
dc.identifier.doi10.1109/ICHORA69329.2026.11537208
dc.identifier.isbn979-833158150-3
dc.identifier.scopus2-s2.0-105042037973
dc.identifier.scopusqualityN/A
dc.identifier.urihttps://doi.org/10.1109/ICHORA69329.2026.11537208
dc.identifier.urihttps://hdl.handle.net/11508/64938
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherInstitute of Electrical and Electronics Engineers Inc.
dc.relation.ispartofICHORA 2026 - 8th International Congress on Human-Computer Interaction, Optimization and Robotic Applications, Proceedings
dc.relation.publicationcategoryKonferans Öğesi - Uluslararası - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/closedAccess
dc.snmzKA_Scopus_20250903
dc.subjectApache Spark
dc.subjectBig Data
dc.subjectCitus
dc.subjectCustom Salting
dc.subjectData Skew
dc.subjectDistributed Systems
dc.subjectMachine Learning
dc.subjectPartitioning
dc.subjectSkew Mitigation
dc.titleSkewSense: ML-Driven Partition Optimization for Distributed Systems
dc.typeConference Object

Dosyalar