期刊文献+
共找到1篇文章
< 1 >
每页显示 20 50 100
A New Indexing Method Based on Word Proximity for Chinese Text Retrieval 被引量:1
1
作者 杜林 孙玉芳 《Journal of Computer Science & Technology》 SCIE EI CSCD 2000年第3期280-286,共7页
This paper proposed a novel text representation and matching scheme for Chinese text retrieval. At present, the indexing methods of Chinese retrieval systems are either character-based or word-based. The character-bas... This paper proposed a novel text representation and matching scheme for Chinese text retrieval. At present, the indexing methods of Chinese retrieval systems are either character-based or word-based. The character-based indexing methods, such as bi-gram or tri-gram indexing, have high false drops due to the mismatches between queries and documents. On the other hand, it's difficult to efficiently identify all the proper nouns, terminology of different domains, and phrases in the word-based indexing systems. The new indexing method uses both proximity and mutual information of the word pairs to represent the text content so as to overcome the high false drop, new word and phrase problems that exist in the character-based and word-based systems. The evaluation results indicate that the average query precision of proximity-based indexing is 5.2% higher than the best results of TREC-5. 展开更多
关键词 information retrieval vector space model automatic indexing proximity-based indexing
原文传递
上一页 1 下一页 到第
使用帮助 返回顶部