TY - GEN
T1 - Dzongkha word segmentation using deep learning
AU - Jamtsho, Yeshi
AU - Muneesawang, Paisarn
N1 - Publisher Copyright:
© 2020 IEEE.
PY - 2020/1
Y1 - 2020/1
N2 - Natural Language Processing (NLP) has been applied to machine translation, chatbots, speech recognition, question and answer systems, document summarization and so on. The Dzongkha language of Bhutan, however, has not been considered in NLP systems, due, presumably, to the fact that the language is complex and written as a string of syllables without proper word boundaries. Thus, Dzongkha word segmentation is the essential first step in building the NLP applications. The novelty of our research is in applying Deep Learning to the task of Dzongkha word segmentation, avoiding the need for manual feature engineering. The segmentation problem is formulated as a syllable tagging task. We also incorporate the windows approach where the tag of a syllable depends on its surrounding syllables. Two sets of experiments were designed, with four models of varying context sizes in each set. We evaluated our models using the syllable-tagged-corpus prepared by Dzongkha Development Commission. The model with context size 2 achieved the highest F-score of 94.40% with 94.47% Precision and 94.35% Recall.
AB - Natural Language Processing (NLP) has been applied to machine translation, chatbots, speech recognition, question and answer systems, document summarization and so on. The Dzongkha language of Bhutan, however, has not been considered in NLP systems, due, presumably, to the fact that the language is complex and written as a string of syllables without proper word boundaries. Thus, Dzongkha word segmentation is the essential first step in building the NLP applications. The novelty of our research is in applying Deep Learning to the task of Dzongkha word segmentation, avoiding the need for manual feature engineering. The segmentation problem is formulated as a syllable tagging task. We also incorporate the windows approach where the tag of a syllable depends on its surrounding syllables. Two sets of experiments were designed, with four models of varying context sizes in each set. We evaluated our models using the syllable-tagged-corpus prepared by Dzongkha Development Commission. The model with context size 2 achieved the highest F-score of 94.40% with 94.47% Precision and 94.35% Recall.
KW - Deep Learning
KW - Deep Neural Network
KW - Dzongkha Word Segmentation
KW - Natural Language Processing
KW - Syllable embedding
KW - Syllable tagging
KW - Window approach
UR - https://www.scopus.com/pages/publications/85084089356
U2 - 10.1109/KST48564.2020.9059451
DO - 10.1109/KST48564.2020.9059451
M3 - Conference contribution
AN - SCOPUS:85084089356
T3 - KST 2020 - 2020 12th International Conference on Knowledge and Smart Technology
SP - 1
EP - 5
BT - KST 2020 - 2020 12th International Conference on Knowledge and Smart Technology
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 12th International Conference on Knowledge and Smart Technology, KST 2020
Y2 - 29 January 2020 through 1 February 2020
ER -