Approaches to Corpus Creation for Low-Resource Language Technology: the Case of Southern Kurdish and Laki
Author(s)
Date Issued
2023
Type
conferenceObject
Start Page
52
End Page
63
Abstract
One of the major challenges that underrepresented and endangered language communities face in language technology is the lack or paucity of language data. This is also the case of the southern varieties of the Kurdish and Laki languages for which very limited resources are available with insubstantial progress in tools. To tackle this, we provide a few approaches that rely on the content of local news websites, a local radio station that broadcasts content in Southern Kurdish and fieldwork for Laki. In this paper, we describe some of the challenges of such under-represented languages, particularly in writing and standardization, and also, in retrieving sources of data and retro-digitizing handwritten content to create a corpus for Southern Kurdish and Laki. In addition, we study the task of language identification in light of the other variants of Kurdish
and Zaza-Gorani languages.
File(s)![Thumbnail Image]()
Name
2023.fieldmatters-1.7.pdf
Description
Versione editoriale definitiva
Size
955.52 KB
Format
Adobe PDF
Checksum (MD5)
0ef48e9128d797894dfe4d4439fc256d
Conference(s)
17th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2023)
