The First Comprehensive Study of Stance Detection Modeling for the Sorani Kurdish Language

Authors

  • Payman S. Rostam Department of Information Technology, Computer Science Institute, Sulaimani Polytechnic University, Sulaimani, Kurdistan Region, Iraq https://orcid.org/0009-0005-8738-2607
  • Rebwar M. Nabi (1) Department of Information Technology, Technical College of Informatics, Sulaimani Polytechnic University, Sulaimani, Kurdistan Region, Iraq; (2) Department of Information Technology, Kurdistan Technical Institute, Sulaimani, Kurdistan Region, Iraq https://orcid.org/0000-0003-2709-7941

DOI:

https://doi.org/10.14500/aro.12814

Keywords:

Cross-lingual learning, Kurdish Sorani, Low-resource languages, Natural Language Processing, Stance detection

Abstract

Stance detection has become a fundamental task in natural language processing (NLP), yet it remains under-explored for low-resource languages such as Sorani Kurdish. Building on the previously released Bochun dataset, the present work focuses exclusively on the comprehensive evaluation of stance detection models and provides the first benchmarking of classical machine learning, deep learning (DL), and transformers. Eight models are implemented: Support vector machine (SVM), logistic regression, random forest, extreme gradient boosting, convolutional neural network, bidirectional long short-term memory, and two frozen transformer encoders, Central Kurdish Bidirectional Encoder Representations from Transformers (BERT) and XLM-RoBERTa-base (XLM-R), each combined with a logistic-regression classification head. All models are trained and evaluated under four protocols (80/20 and 70/30 stratified splits and 5-fold and 10-fold cross-validation) with five random seeds. Class-imbalance handling is investigated through a dedicated ablation comparing no weighting, model-appropriate weighting, and Synthetic Minority Oversampling Technique oversampling, and pairwise statistical comparisons are conducted using Wilcoxon signed-rank tests complemented by Cohen’s d effect sizes. With the chosen feature representations and dataset size, the SVM trained on term frequency–inverse document frequency features delivers the strongest performance (accuracy and weighted F1 of 74% on the 80/20 split and 72% under 10-fold cross-validation), outperforming the DL and frozen-transformer baselines. The findings provide a foundational reference point for future research on Kurdish NLP in general and stance detection specifically.

Downloads

Download data is not yet available.

Author Biographies

Payman S. Rostam, Department of Information Technology, Computer Science Institute, Sulaimani Polytechnic University, Sulaimani, Kurdistan Region, Iraq

Payman S. Rostam is a researcher at the Department of Information Technology, Computer Science Institute, Sulaimani Polytechnic University. She got the B.Sc. degree in computer science and the M.Sc. degree in information technology. Her research interests are in artificial intelligence (AI), natural language processing (NLP), and Kurdish NLP.

Rebwar M. Nabi, (1) Department of Information Technology, Technical College of Informatics, Sulaimani Polytechnic University, Sulaimani, Kurdistan Region, Iraq; (2) Department of Information Technology, Kurdistan Technical Institute, Sulaimani, Kurdistan Region, Iraq

Rebwar M. Nabi is an Assistant Prof. at the Department of Information Technology, Technical College of Informatics, Sulaimani Polytechnic University. He got the B.Sc. degree in Computer Science, the M.Sc. degree in Advanced Computer Science, and tpe Ph.D. degree in Machine Learning. His research interests are in machine learning, natural language Processing, cyber security, and E-government.

References

Abas, A., Veisi, H., and Ali, H.M., 2025. KurdSTS: The Kurdish Semantic Textual Similarity. [Preprint]. Ahmadi, S., 2020. KLPT - Kurdish Language Processing Toolkit. Association for Computational Linguistics, Stroudsburg, pp.72-84. DOI: https://doi.org/10.18653/v1/2020.nlposs-1.11

Alhindi, T., Alabdulkarim, A., Alshehri, A., Abdul-Mageed, M., and Nakov, P., 2021. AraStance: A Multi-Country and Multi-Domain Dataset of Arabic Stance Detection for Fact Checking. In: NLP4IF 2021 - NLP for Internet Freedom: Censorship, Disinformation, and Propaganda, Proceedings of the 4th Workshop, pp.57-65. DOI: https://doi.org/10.18653/v1/2021.nlp4if-1.9

Aljohani, N.R., Fayoumi, A., and Hassan, S.U., 2023. A novel focal-loss and class-weight-aware convolutional neural network for the classification of in-text citations. Journal of Information Science, 49(1), pp.79-92. DOI: https://doi.org/10.1177/0165551521991022

Alturayeif, N., Luqman, H., and Ahmed, M., 2022. MAWQIF: A Multi-label Arabic Dataset for Target-specific Stance Detection. In: WANLP 2022 - 7th Arabic Natural Language Processing - Proceedings of the Workshop, pp.174-184. DOI: https://doi.org/10.18653/v1/2022.wanlp-1.16

Aslam, S., Arshad, S., Shabir, Z., Sohail, M., Ishaq, R., Hameed, H., Javed, S., Ahmed, A.H., and Ahmed, N.H., 2025. Global voices, local frames: Cross-lingual corpus analysis of stance and discourse in social media and news. Scholars Journal of Arts, Humanities and Social Sciences, 9493(9), pp.320-334. DOI: https://doi.org/10.36347/sjahss.2025.v13i09.005

Azad, R., Ahmed, M.S., and Saeed, S.A.B., 2025. KurdABSA : Kurdish aspectbased sentiment analysis dataset curation using few-shot learning. Data in Brief, 62, p.112012. DOI: https://doi.org/10.1016/j.dib.2025.112012

Aziz, K.O., Teimoor, R.A., Tofiq, T.A., and Abdulla, S., 2024. Kurdish sorani dialect morphology generation using a concatenative strategy. UHD Journal of Science and Technology, 8(1), pp.13-19. DOI: https://doi.org/10.21928/uhdjst.v8n1y2024.pp13-19

Badawi, S., 2023. Data augmentation for sorani kurdish news head- line classification using back-translation and deep learning model. Kurdistan Journal of Applied Research, 8(1), pp.27-37. DOI: https://doi.org/10.24017/science/2023.1.4

Bharathi, A., and Zubiaga, A., 2025. Zero-shot cross-lingual stance detection via adversarial language adaptation. PeerJ Computer Science, 11, p.e2955 DOI: https://doi.org/10.7717/peerj-cs.2955

Cignarella, A.T., Lai, M., Bosco, C., Patti, V., and Rosso, P., 2020. SardiStance @ EVALITA2020: Overview of the task on stance detection in Italian tweets. CEUR Workshop Proceedings, 2765, pp.1-10. DOI: https://doi.org/10.4000/books.aaccademia.7084

Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán., F., Grave, E., Ott, M., Zettlemoyer, L., and Stoyanov, V., 2020. Unsupervised Cross-Lingual Representation Learning at Scale. In: Proceedings of the Annual Meeting of the Association for Computational Linguistics, pp.8440-8451. DOI: https://doi.org/10.18653/v1/2020.acl-main.747

Hercig, T., Krejzl, P., Hourová, B., Steinberger, J., and Lenc, L., 2017. Detecting stance in Czech news commentaries.CEUR Workshop Proceedings, 1885, pp.176-180.

Küçük, D., 2017. Stance detection in Turkish tweets. CEUR Workshop Proceedings, 1914, pp.3-6.

Mets, M., Karjus, A., Ibrus, I., and Schich, M., 2024. Automated stance detection in complex topics and small languages: The challenging case of immigration in polarizing news media. PLoS ONE, 19(4), p.e0302380 DOI: https://doi.org/10.1371/journal.pone.0302380

Mohammad, S.M., Kiritchenko, S., Sobhani, P., Zhu, X., and Cherry, C., 2016. SemEval-2016 task 6: Detecting Stance in Tweets. In: SemEval 2016 - 10th International Workshop on Semantic Evaluation, Proceedings, pp.31-41. DOI: https://doi.org/10.18653/v1/S16-1003

Mohammadi, S.M., Farzi, S., Alavi, S.M., and Joonaghany, G.H., 2024. Stance detection on social media, case study: Persian sentences using deep learning architecture. Scientia Iranica, 31(10), pp.764-773. DOI: https://doi.org/10.24200/sci.2024.62504.7876

Rostam, P.S., and Nabi, R.M., 2025. Bochun: Automatically annotated stance detection dataset for Sorani Kurdish language. Data in Brief, 61, p.111839. DOI: https://doi.org/10.1016/j.dib.2025.111839

Saeed, A.M., 2024. An automated new approach in fast text classification: A case study for Kurdish text. Science Journal of University of Zakho, 12(3), pp.1-4. DOI: https://doi.org/10.25271/sjuoz.2024.12.3.1296

Sane, S.R., Tripathi, S., Sane, K.R., and Mamidi, R., 2021. Stance Detection in Code-Mixed Hindi-English Social Media Data Using Multi-Task Learning. In: WASSA@NAACL-HLT 2019 - 10th Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis, Proceedings, pp 1-5. DOI: https://doi.org/10.18653/v1/W19-1301

Vamvas, J., and Sennrich, R., 2020. X-Stance: A Multilingual Multi-Target Dataset for Stance Detection. In: CEUR Workshop Proceedings, p.2624.

Xie, X., Xie, M., Moshayedi, A.J., and Skandari, M.H., 2022. A hybrid improved neural networks algorithm based on L2 and dropout regularization. In: Mathematical Problems in Engineering. Wiley, Hoboken. DOI: https://doi.org/10.1155/2022/8220453

Zhang, W., Yoshida, T., and Tang, X., 2011. A comparative study of TF*IDF, LSI and multi-words for text classification. Expert Systems with Applications, 38(3), pp.2758-2765. DOI: https://doi.org/10.1016/j.eswa.2010.08.066

Zotova, E., Agerri, R., Nuñez, M., and Rigau, G., 2020. Multilingual Stance Detection: The Catalonia Independence Corpus. In: LREC 2020 - 12th International Conference on Language Resources and Evaluation, Conference Proceedings, pp.1368-1375.

Published

2026-08-15

How to Cite

Rostam, P. S. and Nabi, R. M. (2026) “The First Comprehensive Study of Stance Detection Modeling for the Sorani Kurdish Language”, ARO-THE SCIENTIFIC JOURNAL OF KOYA UNIVERSITY, 14(2), pp. 18–27. doi: 10.14500/aro.12814.
Received 2026-01-02
Accepted 2026-06-19
Published 2026-08-15

Similar Articles

1 2 3 4 5 6 7 8 9 10 > >> 

You may also start an advanced similarity search for this article.