Submodular Selection of Orthographic Correspondences Bounds the Script Component of the Sorani-Kurmanji Lexical Divide

Authors

  • Luis Eduardo Muñoz Guerrero PhD, Universidad Tecnológica de Pereira, Pereira, Colombia.

DOI:

https://doi.org/10.67145/ks.v10i1.4141

Keywords:

Central Kurdish, Northern Kurdish, maximum coverage, approximation algorithms, transliteration

Abstract

Background. Central Kurdish and Northern Kurdish are written here in different alphabets, so any comparison of their vocabularies depends on a conversion step that is a choice.

Methods. Correspondence rules are chosen after transliteration to maximise covered token mass under a rule budget: a monotone submodular problem, NP-hard in general, solved greedily and certified against the exact optimum.

Results. Transliteration alone bridges 0.586 of token mass onto forms attested at least five times; twenty rules raise it to 0.653, the ground set’s best twenty. Transliteration returns 0.311 against a Zazaki outgroup and 0.953 against the other half of the source corpus. On article titles held out as pairs the sixty-rule selection bridges 167 tokens, against 28.0 for jumbled-target rules.

Conclusions. Kurdish lexical overlap is not a well-defined single figure but a curve and a position on it; whether the varieties are one language is not decided here.

 

 

Downloads

Published

2022-08-21

How to Cite

Luis Eduardo Muñoz Guerrero. (2022). Submodular Selection of Orthographic Correspondences Bounds the Script Component of the Sorani-Kurmanji Lexical Divide. Kurdish Studies, 10(1), 407–421. https://doi.org/10.67145/ks.v10i1.4141

Issue

Section

Articles