상세 보기
Domain Corpus Independent Vocabulary Generation for Embedded Continuous Speech Recognition
- Lim, Minkyu;
- Kim, Kwang-Ho;
- Kim, Ji-Hwan
WEB OF SCIENCE
3SCOPUS
4초록
This paper proposes a domain corpus independent vocabulary generation algorithm in order to improve the coverage of vocabulary for embedded Continuous Speech Recognition (CSR). A vocabulary in CSR is normally derived from a word frequency list. Therefore, the vocabulary coverage is dependent on a domain corpus. We present air improved way of vocabulary generation using Part-Of-Speech (POS) tagged corpus and knowledge base. We investigate 152 POS tags defined in a POS tagged corpus and word-POS tag pairs. We analyze all words paired with 101 among 152 POS tags and decide on a set of words which have to be included in vocabularies of any size. The other 51 POS tags are mainly categorized with noun-related, Named Entity (NE)-related and verb-related POSs. We introduce a domain corpus independent word inclusion method for noun-, verb-, and NE-related POS tags using knowledge base. For noun-related POS tags, we generate synonym groups and analyze their relative importance using Google search. Then, we categorize verbs by lemma and analyze relative importance of each lemma from a pre-analyzed statistic for verbs. We determine the inclusion order of NEs through Google search. The proposed method shows at least 28.6% relative improvement of coverage for a SMS text corpus when the sizes of vocabulary are 5K, 10K, 15K and 20K. In particular, the coverage of 15K size vocabulary generated by the proposed method reaches tip to 97.8% with the relative improvement of 44.2%(1).
키워드
- 제목
- Domain Corpus Independent Vocabulary Generation for Embedded Continuous Speech Recognition
- 저자
- Lim, Minkyu; Kim, Kwang-Ho; Kim, Ji-Hwan
- 발행일
- 2009-08
- 유형
- Article
- 권
- 55
- 호
- 3
- 페이지
- 1631 ~ 1636