Enhanced Tokenizer for Sinhala Language

Tokenization process plays a prominent role in natural language processing (NLP) applications. It chops the content into the smallest meaningful units. However, there is a limited number of tokenization approaches for Sinhala language. Standard analyzer in apache software library and natural language toolkit (NLTK) are the main existing approaches to tokenize Sinhala language content. Since these are language independent, there are some limitations when it applies to Sinhala. Our proposed Sinhala tokenizer is mainly focusing on punctuation-based tokenization. It precisely tokenizes the content by identifying the use case of punctuation mark. In our research, we have proved that our punctuation-based tokenization approach outperforms the word tokenization in existing approaches.

Keywords

Sinhala Language, Enhanced, Tokenizer

Citation

S. Y. Senanayake, K. T. P. M. Kariyawasam and P. S. Haddela, "Enhanced Tokenizer for Sinhala Language," 2019 National Information Technology Conference (NITC), 2019, pp. 84-89, doi: 10.1109/NITC48475.2019.9114420.

URI

https://rda.sliit.lk/handle/123456789/2010

Collections

Research Papers - Dept of Information Technology

Full item page

Publication:
Enhanced Tokenizer for Sinhala Language

DOI

Files

Type:

Date

Authors

Journal Title

Journal ISSN

Volume Title

Publisher

Research Projects

Organizational Units

Journal Issue

Abstract

Description

Keywords

Citation

URI

Collections

Endorsement

Review

Supplemented By

Referenced By

Publication: Enhanced Tokenizer for Sinhala Language

DOI

Files

Type:

Date

Authors

Journal Title

Journal ISSN

Volume Title

Publisher

Research Projects

Organizational Units

Journal Issue

Abstract

Description

Keywords

Citation

URI

Collections

Endorsement

Review

Supplemented By

Referenced By

Publication:
Enhanced Tokenizer for Sinhala Language