Submodular Selection of Orthographic Correspondences Bounds the Script Component of the Sorani-Kurmanji Lexical Divide
DOI:
https://doi.org/10.67145/ks.v10i1.4141Keywords:
Central Kurdish, Northern Kurdish, maximum coverage, approximation algorithms, transliterationAbstract
Background. Central Kurdish and Northern Kurdish are written here in different alphabets, so any comparison of their vocabularies depends on a conversion step that is a choice.
Methods. Correspondence rules are chosen after transliteration to maximise covered token mass under a rule budget: a monotone submodular problem, NP-hard in general, solved greedily and certified against the exact optimum.
Results. Transliteration alone bridges 0.586 of token mass onto forms attested at least five times; twenty rules raise it to 0.653, the ground set’s best twenty. Transliteration returns 0.311 against a Zazaki outgroup and 0.953 against the other half of the source corpus. On article titles held out as pairs the sixty-rule selection bridges 167 tokens, against 28.0 for jumbled-target rules.
Conclusions. Kurdish lexical overlap is not a well-defined single figure but a curve and a position on it; whether the varieties are one language is not decided here.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2022 Luis Eduardo Muñoz Guerrero (Author)

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.