A Comprehensive Study of Machine Learning and Deep Learning Methods for Sentiment Analysis on Kurdish Sorani Text

Document Type : Original Article

Authors
Department of Computer Science, College of Science, University of Garmian, Sulemaniya, Kalar 46021, Kurdistan region, Iraq.
10.24271/psr.2025.532085.2225
Abstract
The Kurdish language, which is spoken by 40 million people worldwide, still has some of the fewest digital resources, making it
challenging to comprehend natural language and process information digitally. Central Sorani Kurdish is one of its dialects that has
drawn the most attention recently because of the rise in social media usage and the need for automated sentiment analysis. With a focus
on data from social media, this research offers a thorough examination of Machine Learning (ML) and Deep Learning (DL) techniques
used for sentiment analysis of Kurdish Sorani text. The study thoroughly examines current corpora, computational models, and analytical
frameworks, emphasizing the vital role that corpus production has in advancing Kurdish Natural Language Processing (NLP). The size,
standardization, and annotation quality of the available datasets are still restricted, despite significant attempts to create Kurdish corpora
that span the Sorani, Kurmanji, Zazaki, and Gorani dialects. Key methodological trends are identified in the study, such as feature
engineering approaches, preprocessing tactics, and model performance across neural and classical ML methods. The results show that
the lack of extensive, annotated, and dialectally diverse datasets limits the advancement of Sorani sentiment analysis. The necessity of
creating multilingual and multidialectal corpora, encouraging scholarly-linguistic partnerships, and utilizing the Kurdish model from
languages with abundant resources are highlighted in future directions. This thorough investigation intends to contribute to the larger
endeavor of improving Kurdish language technology and comprehending sentiment an
Keywords
Crossmark
Subjects