Multi-dialect Arabic BERT for Country-level Dialect Identification

Bashar Talafha; Mohammad Ali; Muhy Eddin Za’ter; Haitham Seelawi; Ibraheem Tuffaha; Mostafa Samir; Wael Farhan; Hussein Al-Natsheh

Multi-dialect Arabic BERT for Country-level Dialect Identification

Bashar Talafha, Mohammad Ali, Muhy Eddin Za’ter, Haitham Seelawi, Ibraheem Tuffaha, Mostafa Samir, Wael Farhan, Hussein Al-Natsheh

Abstract

Arabic dialect identification is a complex problem for a number of inherent properties of the language itself. In this paper, we present the experiments conducted, and the models developed by our competing team, Mawdoo3 AI, along the way to achieving our winning solution to subtask 1 of the Nuanced Arabic Dialect Identification (NADI) shared task. The dialect identification subtask provides 21,000 country-level labeled tweets covering all 21 Arab countries. An unlabeled corpus of 10M tweets from the same domain is also presented by the competition organizers for optional use. Our winning solution itself came in the form of an ensemble of different training iterations of our pre-trained BERT model, which achieved a micro-averaged F1-score of 26.78% on the subtask at hand. We publicly release the pre-trained language model component of our winning solution under the name of Multi-dialect-Arabic-BERT model, for any interested researcher out there.

Anthology ID:: 2020.wanlp-1.10
Volume:: Proceedings of the Fifth Arabic Natural Language Processing Workshop
Month:: December
Year:: 2020
Address:: Barcelona, Spain (Online)
Editors:: Imed Zitouni, Muhammad Abdul-Mageed, Houda Bouamor, Fethi Bougares, Mahmoud El-Haj, Nadi Tomeh, Wajdi Zaghouani
Venue:: WANLP
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 111–118
Language:
URL:: https://aclanthology.org/2020.wanlp-1.10
DOI:
Bibkey:
Cite (ACL):: Bashar Talafha, Mohammad Ali, Muhy Eddin Za’ter, Haitham Seelawi, Ibraheem Tuffaha, Mostafa Samir, Wael Farhan, and Hussein Al-Natsheh. 2020. Multi-dialect Arabic BERT for Country-level Dialect Identification. In Proceedings of the Fifth Arabic Natural Language Processing Workshop, pages 111–118, Barcelona, Spain (Online). Association for Computational Linguistics.
Cite (Informal):: Multi-dialect Arabic BERT for Country-level Dialect Identification (Talafha et al., WANLP 2020)
Copy Citation:
PDF:: https://aclanthology.org/2020.wanlp-1.10.pdf
Code: mawdoo3/Multi-dialect-Arabic-BERT

PDF Cite Search Code