Computer Department, College of Science, University of Sulaimani, Sulaymaniyah, Iraq.
10.24271/psr.2026.561069.2440
Abstract
Speech synthesis technology has become widely used for education, business, and human-computer interaction. However, the development of a Kurdish language Text-to-Speech (TTS) system has not kept pace with that of other languages, particularly English. Traditional approaches to TTS, such as concatenation and statistical parametric, each have several problems, including unnatural sounding speech, limited flexibility, and a lack of emotion. This study presents an advanced autoregressive sequence-to-sequence neural TTS model for Central Kurdish, adopting a convolutional neural network (CNN)-Transformer encoder. The new encoder offers multi-head self-attention mechanisms with positional encodings for better handling longer dependencies and richer prosody in Kurdish speech. This hybrid architecture combines the powerful autoregressive decoder derived from Tacotron 2 with efficient sequence modeling by transformers while keeping the attention-based alignment mechanism. A Kurdish speech dataset was prepared by segmenting longer utterances to enhance the quality of synthesis and training stability. The proposed model, with sharper attention alignments and more precise prosodic modeling, outperforms the baseline. It emerged as the top-performing system, achieving a Mean Opinion Score (MOS) of 4.74 and a substantial reduction of Word Error Rate (WER) of 3.1%. This result establishes a new benchmark for Kurdish TTS and demonstrates that selective architecture enhancement can yield substantial benefits. The study highlights that the deep learning-based enhancement of Kurdish TTS systems is a means to close an important gap in this language technology, as such findings can be applied to other low-resource languages with high morphological richness.
Ahmad,S Burhan and Mahmood,S Abdullah. (2026). Deep Learning-Based Text-to-Speech Synthesis for Central Kurdish: A Transformer-Enhanced Approach. Passer Journal of Basic and Applied Sciences, 8(2), 712-722. doi: 10.24271/psr.2026.561069.2440
MLA
Ahmad,S Burhan, and Mahmood,S Abdullah. "Deep Learning-Based Text-to-Speech Synthesis for Central Kurdish: A Transformer-Enhanced Approach", Passer Journal of Basic and Applied Sciences, 8, 2, 2026, 712-722. doi: 10.24271/psr.2026.561069.2440
HARVARD
Ahmad S Burhan, Mahmood S Abdullah. (2026). 'Deep Learning-Based Text-to-Speech Synthesis for Central Kurdish: A Transformer-Enhanced Approach', Passer Journal of Basic and Applied Sciences, 8(2), pp. 712-722. doi: 10.24271/psr.2026.561069.2440
CHICAGO
S Burhan Ahmad and S Abdullah Mahmood, "Deep Learning-Based Text-to-Speech Synthesis for Central Kurdish: A Transformer-Enhanced Approach," Passer Journal of Basic and Applied Sciences, 8 2 (2026): 712-722, doi: 10.24271/psr.2026.561069.2440
VANCOUVER
Ahmad S Burhan, Mahmood S Abdullah. Deep Learning-Based Text-to-Speech Synthesis for Central Kurdish: A Transformer-Enhanced Approach. PJBAS. 2026;8(2):712-722. doi: 10.24271/psr.2026.561069.2440