期刊文献+

基于排序学习的文本概念标注方法研究 被引量:2

Learning to Rank Concept Annotation for Text
在线阅读 下载PDF
导出
摘要 提出一种基于排序学习的方法 CRM(concept ranking model),来实现文档的维基百科概念自动标注。首先人工对一定规模的文档进行概念标注,建立训练集合,然后利用排序学习算法在多项特征上得到对概念排序的模型,利用这个概念的排序模型对任意文档进行概念标注。实验表明,相对于传统的文档概念标注方法,此方法在各类指标上都有相当大的提高,标注结果更加接近人类的概念标注。 This paper proposed an automatic text annotation method (CRM, concept ranking model) based onlearning to ranking model. Firstly the authors built a training set of concept annotation manualy, and then used the Ranking SVM algorithm to generate concept ranking model, finally the concept ranking model was used to generateconcept annotation for any texts. Experiments show that proposed method has a significant improvement in various indicators compared to traditional annotation methods, and concept annotation results is closer to humanannotation.
出处 《北京大学学报(自然科学版)》 EI CAS CSCD 北大核心 2013年第1期153-158,共6页 Acta Scientiarum Naturalium Universitatis Pekinensis
基金 国家自然科学基金(90920005 61003192)资助
关键词 概念标注 排序学习 维基百科 显示语义分析 concept annotation learning to ranking Wikipedia explicit semantic analysis
  • 相关文献

参考文献10

  • 1Loh S, Wives L K. Concept-based text mining // Handbook of research on text and Web mining technologies. Hershey: Information Science Refer- ence, 2009:346-358.
  • 2Mihalcea R, Csomai A. Wikify!: linking documents toencyclopedic knowledge // Proceedings of the six- teenth ACM conference on Conference on information and knowledge management (CIKM'07). Lisbon, Portugal, 2007:233-242.
  • 3Milne D, Witten I H. Learning to link with Wikipedia //Proceedings of the 17th ACM conference on Infor- mation and Knowledge Management (CIKM'08). Napa Valley, 2008:509-518.
  • 4Maron M E. On indexing, retrieval and the meaning of about. Journal of the American Society for Infor- mation Science, 1977, 28(1): 38-43.
  • 5Medelyan O, Witten I H, Milne D. Topic indexing with Wikipedia // Proceedings of the AAAI 2008 Workshop on Wikipedia and Artificial Intelligence (WIKIAI'08). Chicago, 2008:19-24.
  • 6Ferragina P, Scaiella U. TAGME: on-the-fly annotation of short text fragments (by wikipedia entities)//Proceedings of the 19th ACM international conference on Information and knowledge mana- gement (CIKM'10). Toronto, 2010:1625 1628.
  • 7Kulkarni S, Singh A, Ramakrishnan G, et al. Collective annotation of Wikipedia entities in web text // Proceedings" of the 15th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD'09). Paris, 2009:457-466.
  • 8Gabrilovich E, Markovitch S. Wikipedia-based semantic interpretation for natural language pro- cessing. Journal of Artificial Intelligence Research, 2009, 34(1): 443-498.
  • 9Li Hang. Learning to rank for information retrieval and natural language processing. San Rafael: Morgan & Claypool Publishers, 2011.
  • 10Cao Yunbo, Xu Jun, Liu Tieyan, et al. Adapting ranking SVM to document retrieval //Proceedings of the 29th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR'06). Seattle, 2006:186-193.

同被引文献27

引证文献2

二级引证文献30

相关作者

内容加载中请稍等...

相关机构

内容加载中请稍等...

相关主题

内容加载中请稍等...

浏览历史

内容加载中请稍等...
;
使用帮助 返回顶部