The First Comprehensive Study of Stance Detection Modeling for the Sorani Kurdish Language

Authors

  • Payman S. Rostam Department of Information Technology, Computer Science Institute, Sulaimani Polytechnic University, Sulaimani, Kurdistan Region, Iraq https://orcid.org/0009-0005-8738-2607
  • Rebwar M. Nabi (1) Department of Information Technology, Technical College of Informatics, Sulaimani Polytechnic University, Sulaimani, Kurdistan Region, Iraq; (2) Department of Information Technology, Kurdistan Technical Institute, Sulaimani, Kurdistan Region, Iraq https://orcid.org/0000-0003-2709-7941

DOI:

https://doi.org/10.14500/aro.12814

Keywords:

Cross-lingual learning, Kurdish Sorani, Low-resource languages, Natural Language Processing, Stance detection

Abstract

Stance detection has become a fundamental task in natural language processing (NLP), yet it remains under-explored for low-resource languages such as Sorani Kurdish. Building on the previously released Bochun dataset, the present work focuses exclusively on the comprehensive evaluation of stance detection models and provides the first benchmarking of classical machine learning, deep learning (DL), and transformers. Eight models are implemented: Support vector machine (SVM), logistic regression, random forest, extreme gradient boosting, convolutional neural network, bidirectional long short-term memory, and two frozen transformer encoders, Central Kurdish Bidirectional Encoder Representations from Transformers (BERT) and XLM-RoBERTa-base (XLM-R), each combined with a logistic-regression classification head. All models are trained and evaluated under four protocols (80/20 and 70/30 stratified splits and 5-fold and 10-fold cross-validation) with five random seeds. Class-imbalance handling is investigated through a dedicated ablation comparing no weighting, model-appropriate weighting, and Synthetic Minority Oversampling Technique oversampling, and pairwise statistical comparisons are conducted using Wilcoxon signed-rank tests complemented by Cohen’s d effect sizes. With the chosen feature representations and dataset size, the SVM trained on term frequency–inverse document frequency features delivers the strongest performance (accuracy and weighted F1 of 74% on the 80/20 split and 72% under 10-fold cross-validation), outperforming the DL and frozen-transformer baselines. The findings provide a foundational reference point for future research on Kurdish NLP in general and stance detection specifically.

Downloads

Download data is not yet available.

References

Abas, A., Veisi, H., and Ali, H.M., 2025. KurdSTS: The Kurdish Semantic Textual Similarity. [Preprint]. Ahmadi, S., 2020. KLPT - Kurdish Language Processing Toolkit. Association for Computational Linguistics, Stroudsburg, pp.72-84.

Alhindi, T., Alabdulkarim, A., Alshehri, A., Abdul-Mageed, M., and Nakov, P., 2021. AraStance: A Multi-Country and Multi-Domain Dataset of Arabic Stance Detection for Fact Checking. In: NLP4IF 2021 - NLP for Internet Freedom: Censorship, Disinformation, and Propaganda, Proceedings of the 4th Workshop, pp.57-65.

Aljohani, N.R., Fayoumi, A., and Hassan, S.U., 2023. A novel focal-loss and class-weight-aware convolutional neural network for the classification of in-text citations. Journal of Information Science, 49(1), pp.79-92.

Alturayeif, N., Luqman, H., and Ahmed, M., 2022. MAWQIF: A Multi-label Arabic Dataset for Target-specific Stance Detection. In: WANLP 2022 - 7th Arabic Natural Language Processing - Proceedings of the Workshop, pp.174-184.

Aslam, S., Arshad, S., Shabir, Z., Sohail, M., Ishaq, R., Hameed, H., Javed, S., Ahmed, A.H., and Ahmed, N.H., 2025. Global voices, local frames: Cross-lingual corpus analysis of stance and discourse in social media and news. Scholars Journal of Arts, Humanities and Social Sciences, 9493(9), pp.320-334.

Azad, R., Ahmed, M.S., and Saeed, S.A.B., 2025. KurdABSA : Kurdish aspectbased sentiment analysis dataset curation using few-shot learning. Data in Brief, 62, p.112012.

Aziz, K.O., Teimoor, R.A., Tofiq, T.A., and Abdulla, S., 2024. Kurdish sorani dialect morphology generation using a concatenative strategy. UHD Journal of Science and Technology, 8(1), pp.13-19.

Badawi, S., 2023. Data augmentation for sorani kurdish news head- line classification using back-translation and deep learning model. Kurdistan Journal of Applied Research, 8(1), pp.27-37.

Bharathi, A., and Zubiaga, A., 2025. Zero-shot cross-lingual stance detection via adversarial language adaptation. PeerJ Computer Science, 11, p.e2955

Cignarella, A.T., Lai, M., Bosco, C., Patti, V., and Rosso, P., 2020. SardiStance @ EVALITA2020: Overview of the task on stance detection in Italian tweets. CEUR Workshop Proceedings, 2765, pp.1-10.

Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán., F., Grave, E., Ott, M., Zettlemoyer, L., and Stoyanov, V., 2020. Unsupervised Cross-Lingual Representation Learning at Scale. In: Proceedings of the Annual Meeting of the Association for Computational Linguistics, pp.8440-8451.

Hercig, T., Krejzl, P., Hourová, B., Steinberger, J., and Lenc, L., 2017. Detecting stance in Czech news commentaries.CEUR Workshop Proceedings, 1885, pp.176-180.

Küçük, D., 2017. Stance detection in Turkish tweets. CEUR Workshop Proceedings, 1914, pp.3-6.

Mets, M., Karjus, A., Ibrus, I., and Schich, M., 2024. Automated stance detection in complex topics and small languages: The challenging case of immigration in polarizing news media. PLoS ONE, 19(4), p.e0302380

Mohammad, S.M., Kiritchenko, S., Sobhani, P., Zhu, X., and Cherry, C., 2016. SemEval-2016 task 6: Detecting Stance in Tweets. In: SemEval 2016 - 10th International Workshop on Semantic Evaluation, Proceedings, pp.31-41.

Mohammadi, S.M., Farzi, S., Alavi, S.M., and Joonaghany, G.H., 2024. Stance detection on social media, case study: Persian sentences using deep learning architecture. Scientia Iranica, 31(10), pp.764-773.

Rostam, P.S., and Nabi, R.M., 2025. Bochun: Automatically annotated stance detection dataset for Sorani Kurdish language. Data in Brief, 61, p.111839.

Saeed, A.M., 2024. An automated new approach in fast text classification: A case study for Kurdish text. Science Journal of University of Zakho, 12(3), pp.1-4.

Sane, S.R., Tripathi, S., Sane, K.R., and Mamidi, R., 2021. Stance Detection in Code-Mixed Hindi-English Social Media Data Using Multi-Task Learning. In: WASSA@NAACL-HLT 2019 - 10th Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis, Proceedings, pp 1-5.

Vamvas, J., and Sennrich, R., 2020. X-Stance: A Multilingual Multi-Target Dataset for Stance Detection. In: CEUR Workshop Proceedings, p.2624.

Xie, X., Xie, M., Moshayedi, A.J., and Skandari, M.H., 2022. A hybrid improved neural networks algorithm based on L2 and dropout regularization. In: Mathematical Problems in Engineering. Wiley, Hoboken.

Zhang, W., Yoshida, T., and Tang, X., 2011. A comparative study of TF*IDF, LSI and multi-words for text classification. Expert Systems with Applications, 38(3), pp.2758-2765.

Zotova, E., Agerri, R., Nuñez, M., and Rigau, G., 2020. Multilingual Stance Detection: The Catalonia Independence Corpus. In: LREC 2020 - 12th International Conference on Language Resources and Evaluation, Conference Proceedings, pp.1368-1375.

Published

2026-08-15

How to Cite

Rostam, P. S. and Nabi, R. M. (2026) “The First Comprehensive Study of Stance Detection Modeling for the Sorani Kurdish Language”, ARO-THE SCIENTIFIC JOURNAL OF KOYA UNIVERSITY, 14(2), pp. 18–27. doi: 10.14500/aro.12814.
Received 2026-01-02
Accepted 2026-06-19
Published 2026-08-15

Similar Articles

<< < 11 12 13 14 15 16 17 18 19 20 > >> 

You may also start an advanced similarity search for this article.