Skip to main navigation Skip to search Skip to main content

Co-occurrence-based error correction approach to word segmentation

  • Mahidol University

Research output: Chapter in Book/Report/Conference proceedingChapterpeer-review

2 Citations (Scopus)

Abstract

A number of word segmentation algorithms have been offered in the past; however, there is still room for improvement. Co-occurrence-Based Error Correction (CBEC), the proposed approach in this chapter, is a novel Thai word segmentation approach that was designed to provide accurate segmentation results based on context and purpose. CBEC quickly segments the input string using any available algorithm; maximal matching was used in the experiment. Next, CBEC checks its segmentation output against an error risk data bank to determine if there is any error risk. The error risk data bank is developed based on a training corpus. The current version of the error risk bank was based on the training corpus available at BEST 2009. Then, CBEC re-segments the input string using the co-occurrence score of the word sequence to ensure the accuracy of the segmentation result.

Original languageEnglish
Title of host publicationCross-Disciplinary Advances in Applied Natural Language Processing
Subtitle of host publicationIssues and Approaches
PublisherIGI Global
Pages354-364
Number of pages11
ISBN (Print)9781613504475
DOIs
Publication statusPublished - 2011

Fingerprint

Dive into the research topics of 'Co-occurrence-based error correction approach to word segmentation'. Together they form a unique fingerprint.

Cite this